Implemented: offline V·J reference build (5 organisms) with a per-organism
locus-coverage manifest (loci_manifest.tsv; warns on empty loci / unreachable
D germlines), MMseqs2 runtime mapping, C++ markup transfer, spec-valid AIRR output
(per-segment V/D/J/C cigars, sequence/germline alignments, read-as-submitted
orientation via rev_comp), reverse-complement handling, all-loci single-DB
querying, streaming/bounded-memory FASTQ I/O (optional quality retention),
out-of-frame junction translation, extended V/J-position markup, D-segment mapping
(incl. D-D fusions across all D loci), offline GenBank-vs-IgBLAST test fixtures.
-
D-segment mapping. After V/J transfer, the V..J interior of the junction (between the projected
v_sequence_endandj_sequence_start) is aligned against the per-organism D germline set by gapless local alignment in the C++_markup.d_local_alignprimitive — mmseqs is unreliable on ~8-31 nt D — and the best hit is emitted asd_call(a comma-separated allele ambiguity list on a score tie) +d_sequence_start/d_sequence_end(AIRR, query coords). D germlines ship indatabase/vdj/<org>/d_germlines.fasta(VDJ loci only); VJ loci are skipped automatically. The call is gated by a Karlin–Altschul E-value (_D_MAX_EVALUE = 0.2), which replaced four hand-tuned per-locus score floors. Concordance vs IgBLAST where both call a D: TRB/TRD ~97% gene agreement, IGH 94-98% across the five organisms (recall 52-67%).- Genomic order constrains D×J. TRBD2 lies 3' of the entire TRBJ1 cluster, so
deletional joining can never produce a TRBD2–TRBJ1 pair. Unenforced, TRBD2 (16 nt)
outscored TRBD1 (12 nt) on noise and took 17% of real human TRB J1-cluster D calls.
transfer._allowed_dmasks the candidate set (holding the E-value'snat the full locus size, so a J1 record can only lose an impossible call, never gain a weak one), andscripts/build_d_priors.pyzeroes the same cells before renormalising. IGH and TRD place every D 5' of every J: nothing is masked there. - Double D-D junctions. For D-D loci (IGH/TRB/TRD — every locus with a D
germline set) a second non-overlapping D is sought above a stricter threshold
and emitted as
d2_call(also a comma-separated ambiguity list) +d2_sequence_*;np1/np2/np3partition the junction between V, the D(s), and J. Recall is 61% on injected IGH tandems but only 13-15% on TRB, where trimming usually leaves one D under the ~7 nt needed to see it; falsed2_callon true single-D is 0-1%.
- Genomic order constrains D×J. TRBD2 lies 3' of the entire TRBJ1 cluster, so
deletional joining can never produce a TRBD2–TRBJ1 pair. Unenforced, TRBD2 (16 nt)
outscored TRBD1 (12 nt) on noise and took 17% of real human TRB J1-cluster D calls.
-
Multi-node sharding.
arda splitround-robins a huge FASTA/FASTQ into N shards (one pass);arda mergeconcatenates per-shard AIRR TSVs (single header);arda slurmrenders/submits asubmit.shchaining split →sbatch --arrayannotate → merge via anafterokdependency (arda.cluster). Split/merge/script are unit-tested; the cluster run is pending a live SLURM test. -
Nucleotide junction re-mapping → recombination-scenario counts.
arda markupandarda.cdr3fixcurrently take an amino-acid junction (the VDJdb case); add thecdr3ntpath. With nucleotides the germline boundaries are exact rather than codon-quantised, so a(cdr3nt, V, J)record yields the full recombination scenario:delV,insVD,D+delDl/delDr,insDJ,delJ. Half of this already exists —annotate.dmap.map_d_junctionderivesv_sequence_end/j_sequence_startin nt (longest common prefix/suffix vs the shippedcdr3_anchors.tsvgermlines) and calls D with an E-value gate.The point is the sufficient statistics: counting scenarios over a repertoire gives exactly the E-step tallies needed to re-estimate a generative model by EM in the style of Murugan et al. (PNAS 2012) / IGoR — `P(V)`, `P(D,J)`, `P(delV|V)`, `P(delDl,delDr|D)`, `P(delJ|J)`, `P(insVD)`, `P(insDJ)` and the insertion Markov matrices. arda would then generate its own priors instead of borrowing OLGA's (see `arda.dpost`, whose `d_prior.tsv` tables are currently OLGA/vdjrearm-derived). Note a scenario is not unique — several `(delV, insVD, ...)` explain one junction — so the E-step must sum over scenarios, not take the MAP. -
arda.hmm— a hidden semi-Markov model of V→N1→D→N2→J. The E-step above is a forward-backward pass, so this is the same project, not a second one. Semi-Markov because insert lengths and germline deletions have explicit non-geometric durations — exactly the shape of whatd_prior.tsvalready ships. Condition on the mmseqs V/J call so the DP sums only over(delV, insVD, D, delDl, delDr, insDJ, delJ): an O(L²·|D|) recursion on a ~60 nt junction, cheap per clonotype (post-correct), not per read. Prior art: partis (Ralph & Matsen, arXiv:1503.04224; GPL-3.0, so compatible), which pairs a Smith-Waterman stage with the HMM exactly as arda pairs mmseqs with_map_d, and whose central finding is that per-allele categorical transitions and per-allele-per-position mutation probabilities beat parametric ones.Three things arda hand-rolls become HMM primitives, which is the real argument: `_allowed_d` (TRBD2×TRBJ1 = 0) is a transition-probability zero, currently enforced in three separate places; `_D_MAX_EVALUE` stands in for a likelihood ratio; and `dpost`'s `beta` exists *only* because it multiplies a raw aa score rather than a log-likelihood. `_anchored_vj_bounds`' longest-common-prefix is Viterbi with the insert-length prior pinned to a point mass at 0 — which is why it overshoots 1–2 nt when the first N base happens to match germline. **Two measured negatives, so nobody re-runs them.** (i) Re-ranking nt D candidates by `λ·S + log P(insVD) + log P(dlen) + log P(insDJ) + log P(D|J)` changes *nothing*: gene accuracy 98.9→97.8 % (IGH), 94.2→94.5 % (huTRB), flat elsewhere, identical call rate and `d_start` error. With 10–18 matched nt, `λ·S` is 11–20 nats and the prior moves ±3. (ii) Replacing the E-value gate with a present/absent Bayes factor buys IGH ~+3 pp recall at matched FP (93.6 % vs 90.7 % @ ~2 % FP) and nothing for TRD — but needs a *per-locus* threshold (BF>6 for IGH, BF>10 for TRD at the same FP), reintroducing the four knobs the E-value removed, and only exists for 4 of the 15 (organism, D-locus) pairs arda ships. So: the HMM's value is posteriors, uncertainty, one home for the constraint, and the EM E-step — **not** D-call accuracy. Prerequisite for IGH: a per-allele-per-position SHM model. Without one the HMM will explain mutated germline as N-region and do *worse* than the current exact-match anchors, which at least fail safely (SHM truncates the match, widening the interior, never clipping the D). Also fix `_map_d`'s aa path while there: it searches the three translated D frames as independent database entries, tripling `n`, when the prior over `insVD` induces a prior over frame (as `dpost` already knows). -
cdr3fixrepairs to canonical, always.cdr3_repairedis accepted only when it opens with Cys104 and closes with Phe/Trp118 (_canonicalise); otherwise the submission comes back untouched.goodimplies canonical by rule. Trimming flanking framework (_MAX_TRIM = 3) is budgeted separately from inventing germline (_MAX_FIX = 2) — removing residues the germline never explained is a much smaller risk than adding ones never observed. Each trimmed residue costs_TRIM: a free trim could tie the untrimmed alignment, win the tie-break, and eat the conserved Phe of clean short IGK junctions. Result on the 250-row fixture: VDJdb's repair reproduced 100/100 (was 98), zero novel rewrites, andgood/vCanonical/jCanonicalagree with VDJdb on every row. OLGA false repairs unchanged at 0.35 %.family*01was dropped from the allele ladder: dead for all five shipped organisms.FailedReplaceturned out to be reachable after all (three substitutions beside a long J anchor,max_replace >= 3), and is now tested. -
Full AIRR productivity.
productiveis currently a heuristic (in-frame + stop-free V..J span); align it with the complete AIRR productivity rules (start codon, stop-codon scan over the whole VDJ, frame of the junction). -
Performance. Optional per-chunk process-pool for inputs where mmseqs is not the bottleneck (mostly-non-receptor bulk RNA-seq); mmseqs index reuse.
-
RNA-seq mode (
arda rnaseq). Staged pipeline for bulk RNA-seq (~1-5% receptor reads).map: recall-first, paired-FASTQ (--r1/--r2), streams and writes only mapped reads (keyed by read id = read-id → junction map) + optional candidate FASTA + a JSON report (arda.rnaseq.map, reusesannotate.mapperwith amapped_onlyfast-path).correct: CDR3 error correction, a port of vdjtoolsCorrector(parent:child count ratio, ≤2 mismatch, 20×) overseqtreeneighbour search (arda.rnaseq.correct; core depseqtree).arda igblast: all-loci gold-standard AIRR (refbuild.gold). Benchmarked in thearda-benchmarkrepo vs assembly-based extractors (speed) and IgBLAST (accuracy).- Seamless pipeline integration. A one-shot
arda rnaseq run(map + correct →<prefix>.clones.tsv/.airr.tsv/.arda.json), a top-level--version, and a drop-in nf-core-style Nextflow module (integrations/nextflow/arda/: process + conda env + Dockerfile + per-process config + README). Bulk RNA-seq pipelines (nf-core/rnaseq and friends) get per-sample AIRR clonotype tables from the same trimmed-FASTQ channel their aligners consume, published to${params.outdir}/arda/. Seedocs/pipeline_integration.rst. - Stage 3 — contig assembly. Reconstruct full-length V(D)J contigs from the
candidate reads (interface stub in
arda.rnaseq.assemble; the role assembly-based extractors play). Deferred — a de-novo assembler, out of scope for the filter-first goal.- Contig cigars (
arda.annotate.contig). Once a contig exists, give it valid V/J/C cigars + AIRR alignments two ways that produce the same record:reannotate_contigs(re-align the contig throughannotate_records) andmerge_contig/merge_contigs(stitch the reads' existing alignments via the C++_markup.merge_alignment, no second mmseqs pass — ~9× faster at ~10^5 contigs/sample). This annotates a contig; the de-novo assembler above is still what would produce one.
- Contig cigars (
- C k-mer prefilter (contingency). If the MMseqs2 prefilter is
the throughput bottleneck vs assembly-based extractors, add a parallel spaced-seed germline
index (new
src/_vjprefilter/pybind11 ext) that rejects non-receptor reads and emits V/J allele hints to prune alignment. Gated on measured need.
- Seamless pipeline integration. A one-shot