shahlab/demux separates ATAC and WGS reads from the DATAC single-cell co-assay: it demultiplexes raw fastqs by barcode, filters the sample's aligned BAM down to ATAC-only reads, and runs ArchR-based QC to produce an interactive HTML report (fragment size distribution, TSS enrichment) per sample.
This is a Nextflow translation of an existing Snakemake pipeline (originally written by Andrew McPherson), built following nf-core conventions.
- Read QC (
FastQC) - Demultiplex reads by barcode into ATAC, WGS, and no-match (
Ultraplex) - Extract ATAC read IDs per fastq
- Merge per-fastq read IDs into one deduplicated list per sample
- Filter the sample's aligned BAM down to ATAC-only reads
- Index the filtered BAM (
samtools index) - Convert the filtered BAM to a fragments file (
SnapATAC2) - Sort, bgzip, and tabix-index the fragments file
- Run ArchR QC (
createArrowFiles) to compute per-cell TSS enrichment and fragment counts (ArchR) - Build an interactive HTML QC report (fragment size distribution, TSS enrichment vs. unique fragments) per sample
Note
If you are new to Nextflow and nf-core, please refer to this page on how to set up Nextflow.
First, prepare a samplesheet with your input data:
samplesheet.csv:
sample,fastq_r2,markdup_bam
SAMPLE1,/path/to/sample1_L001_R2.fastq.gz,/path/to/sample1_markdup.bam
SAMPLE1,/path/to/sample1_L002_R2.fastq.gz,/path/to/sample1_markdup.bam
SAMPLE2,/path/to/sample2_L001_R2.fastq.gz,/path/to/sample2_markdup.bamEach row represents one R2 fastq file (one per sequencing lane). markdup_bam is the sample's already-aligned, duplicate-marked BAM (produced upstream by mondrian) — repeat the same path across every row for a given sample; the pipeline merges these lanes back together internally.
Now, you can run the pipeline using:
nextflow run shahlab/demux \
-profile singularity,slurm \
--input samplesheet.csv \
--outdir <OUTDIR> \
--barcodes_csv <path/to/barcodes.csv>--barcodes_csv is required — it points to the barcode reference file used for demultiplexing.
All tool dependencies are supplied by containers that Nextflow pulls automatically, so no local images or conda environments need to be built first.
Outputs for each sample are published together under <OUTDIR>/<sample>/, including the filtered BAM, its index, the fragments file, ArchR's QC PDFs, and the interactive HTML report.
Most processes use ready-made BioContainers images. Two need software combinations that no public image provides, so they use images built on the Seqera Containers community registry:
| Module | Contents |
|---|---|
archr_qc |
ArchR plus the hg19 annotation packages addArchRGenome("hg19") requires |
build_qc_report |
python, pandas, numpy, plotly |
Each of those modules has an environment.yml that is the single source of truth for its image. Wave derives the image tag from a hash of the build request, so an unchanged environment.yml always resolves to the same image. If you edit one, regenerate the image and update the module's container directive:
scripts/build_wave_containers.py --waitSeqera commits to retaining community images for a minimum of five years. The script is there so the images can be rebuilt if that ever lapses, or if you would rather host them yourself.
Warning
Please provide pipeline parameters via the CLI or Nextflow -params-file option. Custom config files including those provided by the -c Nextflow option can be used to provide any configuration except for parameters; see docs.
shahlab/demux was originally written by Alex Radu, translating an existing Snakemake pipeline written by Andrew McPherson, with guidance from Eli Havasov.
If you would like to contribute to this pipeline, please see the contributing guidelines.
An extensive list of references for the tools used by the pipeline can be found in the CITATIONS.md file.
This pipeline uses code and infrastructure developed and maintained by the nf-core community, reused here under the MIT license.
The nf-core framework for community-curated bioinformatics pipelines.
Philip Ewels, Alexander Peltzer, Sven Fillinger, Harshil Patel, Johannes Alneberg, Andreas Wilm, Maxime Ulysse Garcia, Paolo Di Tommaso & Sven Nahnsen.
Nat Biotechnol. 2020 Feb 13. doi: 10.1038/s41587-020-0439-x.