This document describes the input and output formats for olmsted-cli: what data is expected, where fields are found, how they map to the output, and what is enforced by validation.
See also:
- README.md: Command usage and examples
- ARCHITECTURE.md: Processing pipeline and data flow
- DEVELOPMENT.md: Setup guide, testing, how-to guides
- CLAUDE.md: Code quality rules, terminology
- PCP Input Format
- AIRR Input Format
- Mutations CSV Format
- Olmsted JSON Output Format
- Field Metadata
- Validation
- Field Mapping: Input → Output
PCP (Parent-Child Pair) format uses one or two CSV files.
Three columns identify the row's clonal-family membership. Each role
auto-detects across a small family of names; pass --<role>-col to
override.
| Role | Required? | Accepted column names | CLI override |
|---|---|---|---|
| Sample | required (PCP CSV); optional (trees CSV) | sample, sample_id, sample_name |
--sample-col |
| Family | required | family, family_id, family_name |
--family-col |
| Tree | optional | tree, tree_id, tree_name |
--tree-col |
Preference order when multiple variants are present: _id > bare >
_name. If two variants appear with conflicting values on the same
row, processing fails with a Column conflict error — pass the
override to disambiguate.
The output clone_id is synthesized as {sample}_{family} so families
that repeat across samples no longer silently merge.
The tree role enables multiple per-row alternate reconstructions of the
same clonal family. When absent, every (sample, family) collapses to a
single tree (today's common case). When present, rows are grouped by
(sample, family, tree) and each tree value becomes a separate entry
in clone.trees[].
Each row represents one parent-child edge in a phylogenetic tree.
Required columns (processing will fail without these):
| Column | Description |
|---|---|
sample role (one of sample / sample_id / sample_name) |
Sample identifier |
family role (one of family / family_id / family_name) |
Clonal family identifier |
parent_name |
Parent node name |
child_name |
Child node name |
Standard optional columns (parsed by the standard pipeline):
| Column | Maps to | Level | Notes |
|---|---|---|---|
tree role (tree / tree_id / tree_name) |
tree.tree_name |
tree | Disambiguates rows that belong to alternate reconstructions of the same (sample, family). Omit for single-tree-per-family data. |
parent_heavy / child_heavy |
sequence_alignment |
node | Heavy chain DNA sequences |
parent_light / child_light |
sequence_alignment (light clone) |
node | Light chain DNA sequences (paired data) |
branch_length or edge_length |
length |
branch | Branch length |
distance |
distance |
node | Cumulative distance from root |
depth |
— | — | Depth in tree (not carried to output) |
sample_count |
multiplicity |
node | Sequence abundance |
v_gene_heavy |
v_call |
family | V gene assignment |
d_gene_heavy |
d_call |
family | D gene assignment |
j_gene_heavy |
j_call |
family | J gene assignment |
v_gene_light / j_gene_light |
v_call_light / j_call_light |
family | Light chain gene calls |
cdr1_codon_start_heavy / _end |
cdr1_alignment_start / _end, cdr1_length |
family | CDR1 positions + length; cdr1_length = end − start (nucleotides) |
cdr2_codon_start_heavy / _end |
cdr2_alignment_start / _end, cdr2_length |
family | CDR2 positions + length; cdr2_length = end − start (nucleotides) |
cdr3_codon_start_heavy / _end |
cdr3_alignment_start / _end, cdr3_length |
family | CDR3 (junction) positions + length (nucleotides). Despite the _codon_ column name, the input values are nucleotide coordinates. |
parent_is_naive |
Node type: "root" |
node | Boolean |
child_is_leaf |
Node type: "leaf" |
node | Boolean |
light_chain_type |
light_chain_type |
family | "kappa" or "lambda" |
Column aliases (automatically mapped to canonical names):
| Alias | Maps to |
|---|---|
v_gene, v_call |
v_gene_heavy |
d_gene, d_call |
d_gene_heavy |
j_gene, j_call |
j_gene_heavy |
parent_seq, parent_sequence |
parent_heavy |
child_seq, child_sequence |
child_heavy |
Extra columns: Any column not listed above is captured as a custom node-level field on the child node. Values are auto-coerced: int → float → JSON → bool → string.
Chain suffix convention (paired data): Columns ending with _heavy are applied only to the heavy chain clone/nodes. Columns ending with _light are applied only to the light chain clone/nodes. Columns without a suffix are shared between both chains. Suffixes are stripped in the output (e.g., score_heavy → score).
Each row provides a Newick tree for one clonal family. Role columns (sample / family / tree) follow the same auto-detection rules as the main PCP CSV (see Role columns above).
Required columns:
| Column | Description |
|---|---|
family role (family / family_id / family_name) |
Clonal family ID (matches the main CSV's family role) |
newick_tree or newick |
Newick format tree string |
Optional columns:
| Column | Maps to | Level |
|---|---|---|
sample role (sample / sample_id / sample_name) |
Composite-key component with family role | — |
tree role (tree / tree_id / tree_name) |
tree.tree_name, drives multi-tree-per-clone grouping |
tree |
reconstruction_method |
tree.reconstruction_method |
tree |
rate_scale_heavy |
rate_scale_heavy |
family |
rate_scale_light |
rate_scale_light |
family |
Extra columns: Any column not listed above is auto-classified by intra-clone variance. A column whose values vary across the trees of at least one clone is classified at the tree level (drives the tree dropdown's color/filter/sort controls); a column that is constant within every clone is classified at the clone level. Same chain suffix convention applies for clone-level extras.
Multiple trees per family: Rows are grouped by composite key
(sample, family, tree). Supply a distinct tree-role value
(tree_name recommended) on each row that belongs to an alternate
reconstruction of the same (sample, family). Each composite key
becomes a separate entry in clone.trees[] and in the top-level
trees[]. Without a tree role column, every row attaching to the same
(sample, family) is still treated as an alternate reconstruction
(back-compatible with the older "list-per-key" trees CSV convention).
To produce a valid tree, you need:
sample_id,parent_name,child_namecolumns- At least one parent-child edge with a root node (detected automatically as the node that never appears as a child)
- Sequence data (
parent_heavy/child_heavy) on at least the root node (required formean_mut_freqcalculation)
Missing gene calls, CDR positions, and alignment positions are handled gracefully (empty strings/zeros).
AIRR (Adaptive Immune Receptor Repertoire) format is a single JSON file following the AIRR Community standards.
{
"dataset_id": "...", // Required
"ident": "...",
"subjects": [...],
"samples": [...],
"seeds": [...],
"clones": [...] // Required: array of clone objects
}| Field | Required | Maps to | Notes |
|---|---|---|---|
clone_id |
Yes | clone_id |
Unique family identifier |
sample_id |
Yes | sample_id |
Links to samples array |
subject_id |
No | subject_id |
Defaults to "unknown" |
v_call |
No | v_call |
V gene assignment |
d_call |
No | d_call |
D gene assignment |
j_call |
No | j_call |
J gene assignment |
germline_alignment |
No | germline_alignment |
Germline sequence |
unique_seqs_count |
Yes* | unique_seqs_count |
Schema-required |
mean_mut_freq |
Yes* | mean_mut_freq |
Schema-required |
v_alignment_start |
No | v_alignment_start |
1-based → 0-based conversion |
d_alignment_start |
No | d_alignment_start |
1-based → 0-based conversion |
j_alignment_start |
No | j_alignment_start |
1-based → 0-based conversion |
junction_start |
No | cdr3_alignment_start |
1-based → 0-based conversion, then renamed from the AIRR-standard junction_start |
junction_length |
No | cdr3_length |
Renamed from the AIRR-standard junction_length; value (nucleotides) carried through unchanged |
trees |
Yes | trees |
Array of tree objects |
*Required by schema validation, but processing won't crash without them.
Extra fields: Any additional fields on clone objects are preserved in the output and auto-detected by field_metadata generation.
| Field | Required | Notes |
|---|---|---|
newick |
Yes | Newick tree string (schema-required) |
tree_id |
No | Tree identifier |
nodes |
Yes | Dict or array of node objects |
| Field | Required | Notes |
|---|---|---|
sequence_id |
Yes | Node identifier (schema-required) |
sequence_alignment |
Yes | DNA sequence (schema-required) |
sequence_alignment_aa |
Yes | Amino acid sequence (schema-required) |
parent |
No | Parent node ID (null for root) |
type |
No | "root", "internal", or "leaf" |
length |
No | Branch length to parent |
distance |
No | Cumulative distance from root |
multiplicity |
No | Sequence abundance |
lbi, lbr |
No | Computed with --compute-metrics |
affinity |
No | |
timepoint_id |
No | Sampling timepoint |
mutations |
No | Array of per-mutation records |
Extra fields: Preserved in output and auto-detected.
AIRR uses 1-based closed intervals. olmsted-cli converts *_start positions to 0-based (subtracting 1) during processing. Missing positions are skipped gracefully.
External mutation-level annotations consumed by the merge command and the process --mutations flag. Each row describes one substitution.
| Column | Description |
|---|---|
family |
Clonal family identifier — joined against clone_id in the Olmsted JSON |
site |
Integer amino acid position (0-based, matching the sequence_alignment_aa index) |
parent_aa |
Single-character parent amino acid |
child_aa |
Single-character child amino acid |
If site is non-numeric, parsing fails with a clear ValueError pointing at the offending row.
These are read for context but not added as mutation-level fields on the output:
| Column | Purpose |
|---|---|
sample_id |
Optional sample identifier for cross-checking (not enforced) |
pcp_index |
Optional integer index into the source PCP CSV |
depth |
Optional tree depth where the mutation occurs |
Any column not listed above becomes a mutation-level field on matching nodes. Common examples produced by upstream pipelines:
| Column | Output type | Auto-detected label |
|---|---|---|
surprise_mutsel |
continuous | Surprise (MutSel) |
surprise_neutral |
continuous | Surprise (Neutral) |
surprise_mutsel_theoretical |
continuous | Surprise (MutSel, Theoretical) |
selection_contribution |
continuous | Selection Contribution |
log_selection_factor |
continuous | Log Selection Factor |
num_codon_changes |
continuous | Number of Codon Changes |
These are pre-registered in KNOWN_MUTATION_FIELDS. Any other score column will be auto-detected (continuous if numeric, categorical if string) and appear in field_metadata.mutation with a generated label.
For each tree whose clone_id matches a CSV family:
- The CSV rows for that family are indexed by
(site, parent_aa, child_aa). - For each tree node, mutations are derived by diffing
node.sequence_alignment_aaagainst its parent's (or read directly if amutationsarray already exists). Gap characters (-,.,X,*,?) are skipped. - Each derived mutation is looked up in the CSV index. On match, the score columns are merged onto the mutation dict.
Rows whose family doesn't appear in the JSON, or whose (site, parent_aa, child_aa) doesn't appear on any node in the matched tree, are reported as warnings. The merge still completes; warnings include counts at normal verbosity and per-family detail at -v 2.
family,site,parent_aa,child_aa,surprise_mutsel,selection_contribution,sample_id,depth
clone-abc,12,K,R,4.21,0.77,s1,3
clone-abc,57,A,T,3.06,1.31,s1,3
clone-xyz,9,G,D,5.21,0.94,s1,4The consolidated output format produced by olmsted process.
{
"metadata": {
"format": "olmsted",
"format_version": "1.0",
"schema_version": "2.0.0",
"created_at": "ISO 8601 timestamp",
"source_format": "pcp" | "airr",
"source_files": ["filename.csv"],
"processing_info": {
"datasets_count": 1,
"total_clones_count": 8,
"total_trees_count": 8,
"total_leaf_nodes_count": 186
},
"generated_by": {
"tool": "olmsted-cli",
"version": "0.2.0",
"git_hash": "abc1234"
},
"name": "My Dataset",
"processing_options": { ... }
},
"datasets": [ ... ],
"clones": { "dataset_id": [ ... ] },
"trees": [ ... ]
}| Field | Description |
|---|---|
dataset_id |
Unique identifier (required) |
name |
User-provided name |
clone_count |
Number of clonal families |
field_metadata |
Describes available fields (see below) |
subjects, samples, timepoints |
Metadata arrays |
Contains all family-level data: gene calls, alignment positions, CDR positions, gene support probabilities, and any extra fields from input data.
| Field | Description |
|---|---|
ident |
CLI-minted primary key (tree-{uuid}) |
tree_id |
Semantic identifier: from PCP trees.csv tree_id column if present, otherwise synthesized as tree-{family_id} (paired: -heavy / -light suffix). AIRR: passed through from input, or falls back to ident. |
clone_id |
Links to parent clone |
reconstruction_method |
(optional) Method used to build the tree (e.g. "dnapars", "raxml_ng"). Only present when the input provided one; absent means unknown. |
newick |
Newick tree string |
nodes |
Array of node objects |
A clone can carry multiple alternate-reconstruction trees in its clone.trees[] list; each gets its own entry in the top-level trees[] with a full nodes array.
Contains sequence data, tree topology (parent, type), metrics (lbi, lbr, affinity), and any extra fields from input data.
The field_metadata object on each dataset describes available data fields for the web app's visualization controls.
{
"field_metadata": {
"clone": {
"unique_seqs_count": {
"type": "continuous",
"display": "dropdown",
"label": "Unique Sequences Count"
},
"v_call": {
"type": "categorical",
"display": "dropdown",
"label": "V Gene"
}
},
"node": { ... },
"branch": { ... },
"mutation": {
"child_aa": {
"type": "aa",
"display": "dropdown",
"label": "Child Amino Acid"
},
"parent_aa": {
"type": "aa",
"display": "tooltip",
"label": "Parent Amino Acid"
},
"selection_contribution": {
"type": "continuous",
"display": "dropdown",
"label": "Selection Contribution",
"range": [-2.5, 5.1]
}
}
}
}| Key | Values | Description |
|---|---|---|
type |
continuous, categorical, aa, dna |
What the data is |
display |
dropdown, tooltip |
How the web app uses it |
label |
String | Human-readable display name |
range |
[min, max] |
Value bounds (continuous mutation fields) |
- Known field registries (
KNOWN_CLONE_FIELDS, etc. inconstants.py) — matched by name, provides type/display/label - Auto-inference — values sampled from data, type inferred (numeric → continuous, string → categorical, single-char AA → aa, single-char DNA → dna)
- Suggestions (
SUGGESTED_SKIP_FIELDS,SUGGESTED_DISPLAY_MODES) — removes non-visualization fields, overrides display for context fields - Custom fields (from YAML config) — overrides everything above
| Level | Internal key | Source data | Used for |
|---|---|---|---|
| family | clone |
Clone/family objects | Scatterplot axes, color, shape, facet |
| tree | tree |
Per-tree refs in clone.trees[] |
Tree-dropdown color/filter/sort (multi-tree clones) |
| node | node |
Tree node objects | Tree node tooltips, properties |
| branch | branch |
Branch length on nodes | Tree branch coloring, width |
| mutation | mutation |
mutations[] on nodes, or derived by web app |
Alignment mutation coloring |
A clone-level extra is auto-promoted to tree level when its value varies across the trees of at least one clone — single-tree-per-clone datasets classify everything as clone-level, preserving today's output shape.
When nodes have sequence_alignment_aa, the mutation level includes child_aa and parent_aa even though they aren't in the data — the web app derives them at render time by diffing parent/child sequences.
| Object | Required fields |
|---|---|
| Dataset | dataset_id |
| Clone | unique_seqs_count, mean_mut_freq |
| Tree | newick |
| Node | sequence_id, sequence_alignment, sequence_alignment_aa |
All schemas allow additionalProperties: true — extra fields are preserved.
| Missing field | Behavior |
|---|---|
v/d/j_alignment_start |
Skipped (not adjusted), notification at verbose ≥ 2 |
subject_id |
Left unset; webapp renders its own unknown marker |
timepoint_id |
Left unset; webapp renders its own unknown marker |
Sample not found for sample_id (AIRR) |
clone.sample left unset, notification at verbose ≥ 1 |
Gene calls (v_call, d_call, j_call) |
Empty string; locus is inferred from V-gene prefix when possible, else left unset |
| CDR/alignment positions | Zero values |
| Tree file (PCP) | Trees built from parent-child edges |
tree.reconstruction_method |
Left unset when not supplied by input |
tree.type / dataset.type |
Not synthesized — passed through from input only |
Node length / distance |
olmsted merge backfills these from the tree's newick branch lengths when absent (no-clobber); olmsted validate warns when a branch-length newick has nodes that lack them |
*_id fields that the webapp uses to cross-reference objects must be unique within their natural scope. olmsted process, olmsted tag, and olmsted merge all check this before writing output and fail fast on collisions:
| Scope | Field |
|---|---|
| Within output | dataset.dataset_id |
| Within a dataset | clone.clone_id |
| Within a clone | tree.tree_id |
Within dataset.samples[] |
sample.sample_id |
Within dataset.subjects[] |
subject.subject_id |
Pass --allow-duplicate-ids to downgrade these to warnings and let the data pass through unchanged. sequence_id uniqueness within a tree is always enforced upstream by the Newick parser (duplicate leaf names get _1, _2 suffixes).
detect_file_format() identifies input format:
| Check | Result |
|---|---|
.csv extension |
pcp |
JSON with metadata.format == "olmsted" |
olmsted (explicit tag) |
JSON with datasets + metadata keys |
olmsted (heuristic) |
JSON with dataset_id or clones key |
airr |
| Otherwise | unknown |
| PCP Column | Output Location | Output Field |
|---|---|---|
sample_id |
clone | sample_id |
family |
clone | clone_id |
parent_heavy/child_heavy |
node | sequence_alignment |
| (translated) | node | sequence_alignment_aa |
branch_length |
node | length |
distance |
node | distance |
v_gene_heavy |
clone | v_call |
d_gene_heavy |
clone | d_call |
j_gene_heavy |
clone | j_call |
cdr1_codon_start_heavy |
clone | cdr1_alignment_start |
parent_is_naive |
node | type: "root" |
child_is_leaf |
node | type: "leaf" |
| (computed) | clone | mean_mut_freq |
| (computed) | clone | unique_seqs_count |
| (tree CSV extra cols) | clone | (field name preserved) |
| (PCP CSV extra cols) | node | (field name preserved) |
AIRR fields are mostly passed through directly. Key transformations:
| Transformation | Description |
|---|---|
*_start positions |
1-based → 0-based (subtract 1) |
Matching dataset.samples[] entry |
Denormalized as clone.sample (webapp reads clone.sample.locus etc.) |
| Tree nodes | Extracted from clones, stored in top-level trees[] |
clone.trees |
Reduced to metadata references (nodes removed) |
tree.tree_id |
Passed through from input; when absent, filled with the CLI-minted tree.ident to satisfy the AIRR Community schema's required-field contract. |
Last updated: 2026-06-09