Genome-taxonomy and manuscript-support repository for Strain Pseudo, a candidate novel taxon within Oscillospiraceae.
Generate a publication-grade genome assembly for Strain Pseudo and evaluate whether it represents a novel species and potentially a novel genus using genome-based taxonomy workflows aligned with IJSEM expectations.
This repository is now finalized as a manuscript-support and provenance repository for the current Pseudo genome taxonomy analysis. All major analysis needed for the current draft has been processed and staged.
Species-level novelty is evaluated primarily with TYGS/dDDH comparisons to type-strain genomes. TYGS classified the query genome as a potential new species, and all dDDH values were far below the 70% species boundary.
Genus-level placement is evaluated using EzAAI plus a curated bac120/IQ-TREE phylogeny. The final genus-claim tree will use a reduced type-strain/type-species anchor set to compare Pseudo against the nomenclatural anchors of neighboring validly published genera, while retaining "Vermiculatibacterium agrestimuris" as biologically important provisional lineage context.
The broader TYGS-neighbor bac120 tree is retained as nearest-neighbor context, not as the primary formal genus-claim tree.
The current assembly used for downstream taxonomy comparisons is assembly.polished.fasta.
Selected taxonomy assembly
- Selected as the current best assembly for taxonomy comparisons
- Size: 2,575,873 bp
- GC: 59.68%
- TYGS proteins: 2,593
QC summary for the selected assembly
- Contigs: 6
- Largest contig: 1,329,041 bp
- N50: 1,329,041
- N90: 1,018,847
- auN: 1,098,850.5
- L50: 1
- L90: 2
- Ns per 100 kbp: 0.00
- CheckM completeness: 98.66%
- CheckM contamination: 0.67%
Important assembly provenance note
This repository uses the 2025-09-22 workflow as the assembly/QC/taxonomy workflow backbone, but the final assembly selected for taxonomy comparisons is documented as the Flye-derived assembly.polished.fasta based on overall downstream use and QC. Raven produced a strong alternative assembly with fewer contigs in exploratory comparison, but the current downstream taxonomic summaries are anchored to the Flye-selected draft.
- GTDB-Tk places Strain Pseudo in Oscillospiraceae
- GTDB-Tk places it near placeholder genus
g__Marseille-P3106
TYGS GBDP/dDDH comparisons to type strains are all far below the usual species boundary (~70%), supporting candidate novel species status.
Examples:
- 26.4% dDDH(d4) vs Flintibacter porci P01025
- 20.8% dDDH(d4) vs Vermiculatibacterium agrestimuris DSM 112226T
EzAAI currently shows the strongest proteome-level similarity to Vermiculatibacterium agrestimuris:
- AAI = 74.109%
- Proteome coverage = 0.5537
- Approximate POCP context: 55.37%
Barrnap-extracted 16S support files are included for context only:
- 2 copies
- 1,529 bp each
- supportive, but not treated as the primary delimiter for species/genus claims
An initial FastANI screen against GTDB-derived neighbor genomes recovered the same Marseille-P3106 neighborhood indicated by GTDB-Tk, with a top ANI of approximately 97.44 in that exploratory neighbor panel.
However, this ANI result is retained as supporting neighborhood evidence, not as the primary taxonomic anchor, for two reasons:
- the older ANI workflow produced a raw
fastani.tsvagainst a GTDB-derived neighbor set rather than a finalized curated nomenclatural panel; - a later stricter ANI workflow that attempted to enforce type-material neighbors failed during panel construction and is retained in the repository for provenance only.
In addition, manual inspection of the downloaded metadata for the top Marseille-P3106 RefSeq neighbor indicated that the record was flagged as contaminated/suppressed, which is important context for why the exploratory ANI screen is not centered in the final taxonomic argument.
As a result, final taxonomic interpretation in this repository is based primarily on:
- TYGS/dDDH
- EzAAI
- phylogenomic reconstruction
rather than the initial Marseille-based ANI screen alone.
The genome is not fully closed, but the current assembly is still a high-quality draft suitable for taxonomy-focused analysis.
Current evidence indicates that closure is being limited by long repeat structure rather than by poor overall genome quality. Based on graph inspection and downstream review, two major contigs appear to be separated by repeat regions of approximately ~70 kb and ~140 kb. Resolving those repeats would require multiple long ONT reads that span the repeat plus unique flanking sequence on both sides.
Although the ONT dataset has high total depth, the read-length distribution is not ideal for confidently bridging these repeat regions:
- Reads: 610,420
- Total ONT bases: 1.432 Gb (~550× coverage relative to a ~2.6 Mb genome)
- Read length N50: 3,389 bp
- Longest read: 177,575 bp
- Second longest read: 145,754 bp
In practice, this means overall depth is high, but most reads are much shorter than the repeat regions that need to be resolved. As a result, repeat-spanning evidence is weak and can be confounded by ambiguous repeat mapping.
Bridging diagnostics were also not strongly supportive of confident closure:
- 20 kb end-window comparison: 0 shared reads across orientations
- 50-100 kb windows: some shared reads were observed, but the signal was interpreted as weak/ambiguous and likely influenced by repeat-associated mappings rather than unique bridging evidence
Multiple assembly strategies were explored:
- Unicycler hybrid: fragmented (16 contigs)
- Unicycler bold: improved but still fragmented (12 contigs)
- Raven: strongest contiguity in exploratory comparison (5 contigs)
- Flye: selected as the current best assembly for taxonomy comparisons (6 contigs)
- Autocycler/Trycycler-style workflow: limited benefit in this dataset because the input assemblies were not close enough to complete and the repeat structure still exceeded typical read span
The selected assembly used for taxonomy comparisons remains strong by the metrics most relevant to IJSEM-style genome quality assessment:
- Assembly size: 2,617,187 bp
- GC: 59.52%
- Contigs: 6
- Largest contig: 1,329,041 bp
- N50: 1,329,041
- L50: 1
- CheckM completeness: 98.66%
- CheckM contamination: 0.67%
- Barrnap-extracted 16S copies: 2, each 1,529 bp
So while the genome is not closed, the current evidence suggests that this is a case of the read length being too short in compareison to the repeat regions rather than an assembly that is unusable for taxonomic inference. For that reason:
- a high-quality draft is already sufficient for the current taxonomy work
- closure-support experiments are preserved in this repository as support history
.
|-- README.MD
|-- assembly_qc
| |-- README_2025-09-22.md
| |-- results
| | `-- 2025-09-22
| |-- support_history
| | `-- 2025-12-15-insilico
| `-- workflow_backbone
|-- docs
| `-- PROVENANCE.md
|-- genus_identity
| |-- aai
| | |-- results
| | `-- workflow
| `-- phylogeny
| |-- results
| `-- workflow
`-- species_identity
`-- ani
|-- README.md
|-- archive
| |-- current_failed_type_enforced
| `-- old_success
`-- results
`-- 2025-09-30
Main assembly/QC workflow backbone and the selected summary artifacts used for manuscript-relevant assembly evaluation.
workflow_backbone/= main workflow scripts from the IJSEM-oriented assembly/QC backboneresults/2025-09-22/= selected summary artifacts (QUAST, CheckM, BUSCO, Yak, 16S)finalized_assembly_taxonomy= what was used for all downstream modeling needssupport_history/2025-12-15-insilico/= closure-support experiments and exploratory repeat-resolution work
ANI provenance and workflow history.
archive/old_success/= older ANI workflow that produced the rawfastani.tsvarchive/current_failed_type_enforced/= stricter later ANI prototype that failed during type-material panel constructionresults/2025-09-30/= raw ANI result, refs list, and metadata panel
EzAAI workflow and outputs.
workflow/= AAI pipeline filesresults/=Pseudo_vs_refs.aai.tsv,aai.tsv,aai.nwk
Genus-placement phylogeny workflow and outputs.
workflow/= bac120 / IQ-TREE genus phylogeny workflow filesresults/= current FastTree / IQ-TREE tree artifacts
Repository-level provenance notes.
Final follow up analysis added.
support_history/2026-06-26-final-followup contains the following things:
- Bakta functional annotation screen across Psuedo and selected reference genomes.
- POCPu / conspot genus neighborhood comparison.
- Expanded bac120 phylogeny including Lawsonbacter hominis NSJ-51.
These analysis support the final interpretation that strain Pseudo is a candidate novel species in the Vermiculatibacterium lineage, with additional functional and genus boundary checks retained for provenance.
- workflow backbone organization
- selected taxonomy assembly summary
- TYGS/dDDH evidence
- EzAAI module
- ANI provenance split into old-success vs current-failed
- genus phylogeny workflow/results staging
- All finished on my end.
- assembly/QC workflow organization and final taxonomy-facing assembly summary
- TYGS/dDDH species-level evidence
- EzAAI genus-neighbourhood evidence
- strict type-strain/type-species bac120/IQ-TREE genus-placement phylogeny
- ANI provenance split into old-success versus current-failed/type-enforced attempts
- TYGS genome and 16S GBDP phylograms as supplementary context
- software version and conda-environment provenance records
- Methods/Results draft prepared for collaborator review
This repository is organized around the functional modules that produced the manuscript-relevant results, not as a direct mirror of all original dated run directories. Original run provenance is preserved in docs/PROVENANCE.md and in archived workflow subdirectories where needed.