Single entry point for all agent-driven and AI-coding work in this repository. Start here. This file holds the agent-operational guidance; for the rules and rationale behind it, follow the links to the canonical human-facing docs below — those are the source of truth, so do not restate them here.
Opening, updating, or merging pull requests is never an agent's job. Do the work on a branch and push it when asked; the human opens and manages the PR. Also confirm a branch still exists before pushing to it — pushing to a deleted (e.g. post-merge) branch silently recreates it.
Debugging a failed Terra/WDL pipeline test? See FISS_MCP_RUNBOOK.md — submission URL → the exact failing task/metric, via the
fiss-mcptools (subworkflow log drill-down, GCS traversal, stale-run checks).
| Topic | Document |
|---|---|
| Changelog format and language (entry structure, active voice, sample entries) | website/docs/contribution/contribute_to_warp/changelog_style.md |
| Semantic-versioning classification (major/minor/patch) and release packaging | website/docs/About_WARP/VersionAndReleasePipelines.md |
| Pipeline best practices (automation, testability, portability, scaling, maintainability) | website/docs/About_WARP/BestPractices.md |
| Requirements a released pipeline must meet | website/docs/About_WARP/PipelineRequirements.md |
| Testing: branch model (develop/staging/master) and how the smart test system selects pipelines | website/docs/About_WARP/TestingPipelines.md |
| Docusaurus markdown style (admonitions, tables, tabs, code-block highlighting, cross-refs) | website/docs/contribution/contribute_to_warp_docs/doc_style.md |
| How to build / serve / validate the docs site | website/docs/contribution/contribute_to_warp_docs/docsite_maintenance.md |
The human-facing docs are canonical for rules and rationale. Anything agent-specific or operational lives in this file — do not duplicate human-facing content here; link to it instead.
- Validate all modified WDLs — and every WDL that imports them (transitively). See WDL Validation (MANDATORY).
- Changelog and versioning — follow the agent automation rules below; for entry format see the changelog style guide, and for what makes a change major/minor/patch see VersionAndReleasePipelines.md.
- Cascading version bumps — when modifying a shared task (
tasks/wdl/), every pipeline that imports it (directly or transitively) needs either a patch bump + changelog note or an explicit "no functional impact" entry. See Cascading version bumps. Validation passing is not sufficient — changelogs must also be updated. Scope: an edit to a pipeline WDL (pipelines/wdl/**) or a sharedtasks/wdl/task takes a patch+0.0.1bump; a change confined toverification/**(test wrappers and compare WDLs) takes no bump and no changelog — it ships no pipeline artifact. See What requires a version bump. - After merging develop, grep for duplicate
pipeline_versionlines — keep the higher one. See Git merge conflict resolution. - Sub-workflow contract — removing an input from a shared WDL requires removing it from every caller. See Sub-workflow input contract.
- Stale example/test inputs — when you rename a workflow or remove/rename inputs, audit
pipelines/wdl/<name>/example_inputs/*.jsonandtest_inputs/**/*.json; they break silently because they are not checked by womtool. - Touching a pipeline's interface — also update the pipeline's docs page under
website/docs/Pipelines/<Name>_Pipeline/README.md. Ensure version information is linked from the changelog, not hardcoded. Runyarn --cwd=website buildto catch broken links. - Documented defaults must match the code — if a comment,
parameter_meta, or README says an optional input "defaults to X when unspecified", make that default declarative (select_first([input, "X"])or a defaulted declaration) and verify the consuming code's unset path actually yields X. AString?interpolates to an empty string in bash, so a downstreamif/elsecan silently encode the opposite default (as happened with Optimustenx_chemistry_subversion); womtool checks types, not this, and tests that pin explicit values never exercise the default. Treat every prose "default" as a claim to verify against the code.
WARP is a collection of cloud-optimized WDL (Workflow Description Language) pipelines for biological data processing. It uses a flattened directory structure:
pipelines/wdl/— all WDL workflow definitions, organized by category (arrays/,atac/,dna_seq/,genotyping/,multiome/,optimus/,peak_calling/,reprocessing/,rna_seq/, …)pipelines/nextflow/— placeholder for future Nextflow workflows.tasks/wdl/— all reusable WDL tasks (e.g.Alignment.wdl,BamProcessing.wdl,FastqProcessing.wdl,StarAlign.wdl,Utilities.wdl)structs/— WDL struct definitions for type safetyverification/— test workflows that validate pipeline outputs (verification/test-wdls/for test implementations)scripts/— build and validation automationwebsite/— source for the Docusaurus documentation site.all_of_us/— workflows supporting the All of Us Research Program.
History: Content was previously split between
pipelines/broad/(DNA-seq, arrays, reprocessing) andpipelines/skylab/(single-cell RNA-seq, ATAC-seq, multiome)--any remnants you should offer to clean up. Both are now unified underpipelines/wdl/andtasks/wdl/. Docker images are maintained separately in warp-tools.
All WDL files use relative imports:
import "../../../tasks/wdl/UnmappedBamToAlignedBam.wdl" as ToBam
import "../../../structs/dna_seq/DNASeqStructs.wdl"Import path errors are the most common validation failure.
womtool does not flag an import "..." as NS whose namespace is never used, so dead imports accumulate (they make a reader chase a dependency that isn't there). Scan for them:
scripts/find_unused_wdl_imports.shIt skips structs/ imports — WDL references struct types by bare name, so those legitimately have no NS. usage.
Removing a dead import from a pipeline or shared tasks/wdl/ WDL is still a WDL change, so it forces a patch bump + changelog (+ cascade for shared tasks) — see Cascading version bumps. Dead imports in verification/** are free to drop. So don't do a repo-wide sweep in one PR: clean the verification/** ones anytime, and fold each pipeline/task import removal into that pipeline's next real change so it shares an already-required bump. Known debt as of this writing (run the scan for the current list): atac.wdl (Merge), UltimaGenomicsWholeGenomeGermline.wdl (InternalTasks, QC, UltimaGenomicsWholeGenomeGermlineAlignmentMarkDuplicates, UltimaGenomicsWholeGenomeGermlineQC), UltimaGenomicsWholeGenomeCramOnly.wdl (VariantDiscoverTasks), Optimus.wdl (FastqProcessing), PairedTag.wdl (H5adUtils), SlideSeq.wdl (OptimusInputChecks), tasks/wdl/SplitLargeReadGroup.wdl (Utils).
Future CI wiring (not done — deliberately): the test system is currently flaky, so this stays a manual scan rather than a gate. When CI is trusted, add it as an advisory (non-blocking, continue-on-error) job first so it reports without failing PRs; promote to a blocking check only once the known debt above is cleared and the signal is proven quiet.
Each .github/workflows/test_*.yml has a paths: filter enumerating the WDLs whose changes should trigger that pipeline's test. The filter is maintained by hand, so it drifts out of sync with what the Test WDL actually imports — a watched file gets renamed/deleted (the filter then watches a dead or wrong file), or a new sub-workflow/task import is added but never watched. Either way the test silently stops firing on the edits that matter. Scan for both:
scripts/find_stale_ci_path_filters.shIt reports DEAD (watched path missing / points at the wrong file) and BLIND (imported by the test but not covered by any watched path or dir/** glob) per workflow, and flags any workflow that doesn't resolve to exactly one Test WDL.
These are CI-only changes — editing a workflow's paths: filter touches no WDL, so there's no version bump or changelog cascade. Fix them anytime; they don't need to ride a pipeline's next real change the way unused imports do.
Future CI wiring (not done — deliberately): same rationale as the unused-import scan above — the test system is flaky, so this stays a manual scan. When CI is trusted, wire it as an advisory (continue-on-error) job before ever making it blocking.
A WDL file defines a single input contract for all callers. You cannot expose an input to one workflow but hide it from another that imports the same WDL. To remove a parameter from one consumer, you must remove it from the shared task/sub-workflow and from every caller.
Womtool validation is REQUIRED for all WDL changes, and must pass for every modified file and its dependents. The womtool JAR is expected at a user-specific location such as ~/womtool-90.jar (with a compatible Java available); a convenience wrapper exists at scripts/validate_wdls.sh.
Validate a single file:
java -jar ~/womtool-90.jar validate path/to/pipeline.wdlAfter modifying a task, validate every workflow that imports it — a task can validate in isolation but break a caller's argument list. Use a loop:
for f in tasks/wdl/StarAlign.wdl tasks/wdl/CheckInputs.wdl \
pipelines/wdl/optimus/Optimus.wdl \
pipelines/wdl/multiome/Multiome.wdl \
pipelines/wdl/paired_tag/PairedTag.wdl \
pipelines/wdl/slidetags/SlideTags.wdl \
pipelines/wdl/slideseq/SlideSeq.wdl \
pipelines/wdl/smartseq2_single_nucleus_multisample/MultiSampleSmartSeq2SingleNucleus.wdl \
verification/test-wdls/TestOptimus.wdl \
verification/test-wdls/TestMultiome.wdl \
verification/test-wdls/TestPairedTag.wdl \
verification/test-wdls/TestSlideTags.wdl \
verification/test-wdls/TestSlideSeq.wdl; do
printf '%-90s ' "$f"
java -jar ~/womtool-90.jar validate "$f" 2>&1 | tail -1
doneTest JSON inputs live at two locations per pipeline:
pipelines/wdl/<name>/test_inputs/{Plumbing,Scientific}/— CI test inputs (Plumbing = fast/cheap, Scientific = end-to-end accuracy)pipelines/wdl/<name>/example_inputs/— example inputs for documentation
When adding a new test input for a regression case, also wire it into the relevant GitHub Actions workflow under .github/workflows/ so CI picks it up.
Stale inputs break silently — they are not validated by womtool. When adding, renaming, or removing/renaming inputs, audit example_inputs/*.json and test_inputs/**/*.json and thread the change through the Test<Name>.wdl wrapper (see Registering a CI test step 2 — forward every testable input). Mind the framework's UpdateTestInputs.py key rewrite: a test-JSON key with more than two dotted parts (e.g. Pipeline.Sub.input, reaching into a sub-workflow) is rewritten to TestPipeline.Pipeline.Sub.input, which the wrapper rejects as an extra input. Prefer exposing the input at the pipeline's top level (Pipeline.input) and setting it that way — a two-part key rewrites cleanly to TestPipeline.input.
A WARP "Plumbing" or "Scientific" test is a CI-integrated test that runs the pipeline on Terra and compares its outputs to a stored truth set — not a local demonstration or a notebook. Standing one up for pipeline <Pipeline> requires this whole set of artifacts; a missing piece typically fails only at CI run time (womtool does not catch most of it). Canonical reference: TestingPipelines.md. There is also an internal SOP, "How to Port a Pipeline Test to the New Testing Framework" (Optimus example) — useful for the Dockstore-publish click-path, but its file paths predate the directory flatten (it shows pipelines/skylab/, tasks/broad/; use today's pipelines/wdl/, tasks/wdl/), and its "TestX.wdl modifications" section is Vault→Terra migration only, not relevant to a net-new test. The fastest way to author the Actions YAML and the test wrapper is to copy an existing pipeline's test_<name>.yml / Test<Name>.wdl and adapt.
- Test input JSON(s) —
pipelines/wdl/<name>/test_inputs/{Plumbing,Scientific}/<Case>.json, keyed by the pipeline's input names (<Pipeline>.foo). Host the input data undergs://pd-test-storage-public/<Pipeline>/input/{plumbing,scientific}/.... Plumbing = tiny/fast/cheap (every PR); Scientific = realistic end-to-end (PRs to master, or on demand). - Test wrapper —
verification/test-wdls/Test<Pipeline>.wdl: imports the pipeline,Verify<Pipeline>,Utilities, andTerraCopyFilesFromCloudToCloud; declares every pipeline input a test JSON might set plus the framework-injectedtruth_path,results_path,update_truth(andcloud_providerif the pipeline takes it); setsmeta { allowNestedInputs: true }. It calls the pipeline, gathers outputs into anArray[String], copies them toresults_path, copies totruth_pathwhenupdate_truth, else runsGetValidationInputsand callsVerify<Pipeline>. Forward every testable input — an input the pipeline has but the wrapper omits crashes the moment a test JSON sets it (see Test inputs above). - Verification —
verification/Verify<Pipeline>.wdlcompares each output to truth (tolerantly). Keep pipeline-specific compare tasks here, not in the sharedverification/VerifyTasks.wdl, so editing them doesn't trip every other pipeline's path filter. When aVerify*comparison fails on tiny float or "1-of-N" differences, it is usually known run-to-run nondeterminism, not a regression — compare with a tolerance by extending an existing tolerant task (e.g.CompareGeneMetricsWithTolerance, or the tolerantCompareSAMsinVerifyRNAWithUMIs.wdl) rather than adding a strict check. Setting a nondeterminism tolerance: once a human contributor confirms an observed drift is acceptable (not a real regression), set the threshold to double the observed drift, not just above it. A threshold tuned to barely clear one run's drift (≤2×) flakes intermittently, because the drift varies run-to-run; 2× buys margin while staying orders of magnitude below any real regression, so you lose no sensitivity. Any metric absent from a percentage-threshold table (e.g.CompareAtacLibraryMetricsinVerifyTasks.wdl) defaults to exact match — a newly-nondeterministic metric must be added explicitly or it fails on the first sub-unit drift. - GitHub Actions entry point ("the buttons") —
.github/workflows/test_<pipeline>.yml:on: pull_requestwith apaths:filter scoped to the pipeline's files (pipeline dir, its tasks,Verify<Pipeline>.wdl,Test<Pipeline>.wdl,TerraCopyFilesFromCloudToCloud.wdl, the two workflow files,firecloud_api.py) andworkflow_dispatchwith inputsuseCallCache,updateTruth,testType(choice Plumbing/Scientific),truthBranch. The jobuses: ./.github/workflows/warp_test_workflow.ymlwithpipeline_name: Test<Pipeline>,dockstore_pipeline_name: <Pipeline>,pipeline_dir, those inputs, andsecrets: { PDT_TESTER_SA_B64, DOCKSTORE_TOKEN };permissions: { contents: read, id-token: write, actions: write }. (warp_test_workflow.ymlis the shared reusable workflow: it creates the Terra method config, submits, polls, copies results, and runs verification when not updating truth.) - Dockstore registration and publish — add a
Test<Pipeline>entry (name,subclass: WDL,primaryDescriptorPath: /verification/test-wdls/Test<Pipeline>.wdl) to.dockstore.yml(the pipeline itself is usually already registered). Then publish it in the Dockstore UI — a manual step that's easy to forget and not caught by anything in-repo: wait ~5 min for Dockstore to ingest the new descriptor, go to the Dockstore dashboard, findwarp/Test<Pipeline>via the Search Workflows field, then Versions → Actions → Set as Default Version → Publish. The CI resolves the workflow bydockstore_pipeline_name, so an unpublished / no-default-version workflow makes the Terra method-config step fail. - Seed truth FIRST — a brand-new test has no golden files. Run the Actions workflow manually (
workflow_dispatch) withupdateTruth: trueand the righttruthBranchto populate the truth bucket; only then do compare-runs pass. A compare-run before truth exists fails with nothing to diff. The truth path is keyed by the test JSON's filename, notinput_id:gs://pd-test-storage-public/<Pipeline>/truth/<plumbing|scientific>/<truthBranch>/<JSON-basename>/(e.g.mouse_v4_snRNA_example.json→.../mouse_v4_snRNA_example/). So changing onlyinput_idoverwrites the same truth key; renaming the JSON file starts a fresh one. - Validate every new/modified WDL with womtool. Note what womtool does not check: the Actions YAML, the JSON keys, the
.dockstore.ymlentry, and whether truth exists — all of which only surface at CI run time.
If a pipeline ships only one test kind, pin it in the caller (e.g. scANVI is Scientific-only → test_type: ${{ github.event.inputs.testType || 'Scientific' }}, and warp_test_workflow.yml honors an explicit caller override over the branch-derived default).
Every pipeline carries a String pipeline_version = "major.minor.patch" and a cumulative <PipelineName>.changelog.md. For entry format and language follow the changelog style guide; for what makes a change major / minor / patch follow VersionAndReleasePipelines.md. The agent-operational rules below are not in those docs:
- Version bumps are once per branch/PR. If the top changelog entry on the current branch already has an unreleased version bump (i.e. its version is higher than what is in
develop), append new bullet points to that entry — do not create a new version section. pipeline_versions.txtis automatically generated by a GitHub Action (.github/workflows/update_versions_file.yml) on pull requests todevelop. Do not edit this file manually. The action derives its content from the version and date at the top of each pipeline's.changelog.md.- Linked versions — some pipelines reference a sub-workflow's version (e.g. ImputationBeagle references the ArrayImputationQuotaConsumed version); update both together.
- What requires a version bump. A change to a pipeline WDL (
pipelines/wdl/**) or a sharedtasks/wdl/task it imports requires a patch+0.0.1bump on every affected pipeline (see Cascading version bumps) and a changelog entry. A change confined toverification/**— theTest<Pipeline>.wdlwrappers and theVerify*/CompareMetrics-style compare tasks — requires nopipeline_versionbump and no changelog, because it is test infrastructure and ships no pipeline artifact. (Still validate it with womtool.)
Editing a shared task (tasks/wdl/*.wdl) forces a version bump on every pipeline that imports it, directly or transitively. For each such pipeline update two things: the WDL's String pipeline_version, and its .changelog.md (new entry — a plain "No functional impact" bullet is fine when the change can't affect it).
Get the exact list from the tool — don't guess and don't rely on the chains below being current: run scripts/validate_release.sh -g origin/staging (add -i true to also check changelogs). It prints X.wdl has not been changed and needs updating for every importer still missing a bump; fix and re-run until it says all are valid. This is exactly what the WARP Validate Version / WARP Validate Changelog CI checks run, so a local pass means those checks pass. womtool does not catch this — it validates call signatures, not changelogs.
Commit first. The script diffs committed
HEADagainst the ref (git diff HEAD <ref>), so uncommitted working-tree edits are invisible — with changes unstaged it will falsely report all-valid. Commit, then run it. (It does trace task→pipeline imports transitively, so once committed a shared-task change correctly flags every importer.)
Known chains (a starting point; the script above is authoritative):
tasks/wdl/CheckInputs.wdl→ Optimus (checkOptimusInput) + its wrappers Multiome, PairedTag, SlideTags; MultiSampleSmartSeq2SingleNucleus (checkInputArrays); SlideSeq imports it but calls nothing (→ "no functional impact").tasks/wdl/StarAlign.wdl→ Optimus, SlideSeq, Multiome, PairedTag, SlideTags, MultiSampleSmartSeq2SingleNucleus.tasks/wdl/Metrics.wdl→ most Skylab-origin pipelines.
- Version declaration: WDL files declare
version 1.0. - Formatting: 2-space indentation; blank lines to separate logical sections; no strict line-length limit.
- Naming: tasks and call aliases use
UpperCamelCase; variables uselowercase_underscore(Python style). - File naming: never bake a lifecycle qualifier —
Updated,New,Old,Final,V2,Deprecated, ... — into a WDL filename or itsworkflowname; that's changelog language, not identity. It shows up when a WDL is replaced but the replacement keeps a disambiguating suffix instead of being renamed once the original is deleted. Fix: delete the dead file and rename the replacement in the same change, updating every reference — theimportpath, theasalias, any call-site namespace, and CIpaths:filters (see Stale CI path filters). Example:verification/VerifyCramToUnmappedBamsUpdated.wdloutlived theVerifyCramToUnmappedBams.wdlit had replaced; a stale-migration artifact caught and fixed in #1913. - Workflow input block order: required inputs first, optional inputs with defaults second, runtime-configuration parameters last.
- Task section order: input → command → output → runtime. In
commandblocks, put one input argument per line for clarity. meta { allowNestedInputs: true }— include for Terra compatibility.- GPU runtime keys: use camelCase
gpuType,gpuCount, andnvidiaDriverVersioninruntimeblocks — universally portable across Cromwell/Terra/GCP. The snake_case aliases (hardware_gpu_type,nvidia_driver_version) are not portable. - GPU-optional tasks need TWO task definitions. Cromwell rejects
gpuCount: 0("Expecting gpuCount runtime attribute value greater than 0") and rejects optional-typed GPU attributes (gpuCount: Int?→ "Expected positive Int but got Int? null";gpuType: String?→ "String value required but got String?"), and WDL 1.0 cannot conditionally omit aruntimekey. So a single task is either always-GPU or never-GPU. To offer both GPU and CPU-only runs, write two tasks — identical except one has the GPUruntimeattributes and the other omits them — and route with a workflowif (gpu_count > 0)/if (gpu_count == 0), gathering outputs withselect_first. Keep the container device-agnostic (auto-detect the accelerator) so both tasks run the same command and only VM provisioning differs. Example:pipelines/wdl/scanvi/scANVI.wdl(MultiomeLabelTransfer+MultiomeLabelTransferCpu). - Preemptibility: tasks wire a
preemptible_triesinput topreemptible:inruntime. Keep it preemptible for short/retryable steps, but passpreemptible_tries = 0(non-preemptible) for long single-threaded steps whose runtime can exceed the preemption window — e.g. a large-BAM gather/merge — because a preemption wastes the whole run. Signature in a failed test: the task hits repeatedRetryableFailureand can end on a VM-teardownexit 125("job was stopped before the command finished") while the tool itself had actually completed. Precedent:Qc.wdl,JointGenotypingTasks.wdl,UnmappedBamToAlignedBam.wdl(GatherBamFiles). - WARP-tools vs inline WDL: package complex or reusable processing scripts (Python, R, specialized deps, data transforms) into warp-tools Docker images. Use inline WDL only for simple file operations, basic string manipulation, parameter validation, and straightforward conditional logic.
commandblock style: For multi-linecommandblocks, use a heredoc. The convention depends on the script content:- For bash scripts, use the
<<</>>>style.command <<< set -e echo "Hello, ~{name}!" >>>
- For inline Python, use the
<<CODE/CODEstyle. This is the established convention for tasks that primarily wrap a Python script.command <<CODE set -e python3 -c ' # Python code here. # Use ~{wdl_var} for WDL interpolation. ' CODE
- General guidance: Do not use quoted heredoc delimiters for the main WDL command block (e.g.,
command <<'EOF'), as this disables WDL variable interpolation. In Python, avoid f-strings combined with~{}interpolation since both use curly braces and can conflict (prefer'~{prefix}_output.txt'overf'~{prefix}_output.txt').
- For bash scripts, use the
Pipelines support GCP/Azure via a cloud_provider input. The pattern varies by pipeline origin:
Broad-origin pipelines (DNA-seq, arrays, genotyping) maintain docker registries and validate cloud_provider at runtime:
String gatk_docker = "us.gcr.io/broad-gatk/gatk:4.6.1.0"Error handling: use the ErrorWithMessage task for input validation. For broad-origin pipelines validating cloud_provider:
if ((cloud_provider != "gcp") && (cloud_provider != "azure")) {
call utils.ErrorWithMessage as ErrorMessageIncorrectInput {
input: message = "cloud_provider must be supplied with either 'gcp' or 'azure'."
}
}Expect a pipeline's documentation to be split across two places. This isn't an enforced policy — it's just how pipelines evolve: a README starts in the pipeline directory, and the full write-up later lands on the docs site.
-
In-repo
README.md(pipelines/wdl/<name>/README.md) — usually slim and mostly a pointer to the docs-site page: a one-sentence summary, a link to the full docs page on the WARP site, a minimal "Running the pipeline" snippet, required inputs at a glance, and a link to<Pipeline>.changelog.md. Keep it that way — don't mirror the full docs page into it. -
Docs-site page (
website/docs/Pipelines/<Name>_Pipeline/README.md) — the full documentation. Required Docusaurus frontmatter:--- sidebar_position: 1 slug: /Pipelines/<Name>_Pipeline/README ---
Plus a sibling
_category_.jsonwhen creating a new category.
Do not hardcode version information. Pipeline documentation, especially the header table on the docs-site page, should not contain hardcoded version numbers or dates. This information becomes stale. Instead, link to the pipeline's .changelog.md file as the single source of truth for versioning.
When you change a pipeline's interface, update the docs-site page so its inputs/outputs stay accurate, and make sure the slim in-repo README still points correctly.
Validate docs changes with a full build — it catches broken links and bad frontmatter:
yarn --cwd=website install # first time only
yarn --cwd=website buildyarn --cwd=website start is fine for previewing but does not fail on broken links — always use build before marking docs work complete.
Cross-page links: prefer relative paths to other pipeline pages (e.g. [Multiome](../Multiome_Pipeline/README.md)). When the target lacks a docs page, link to the GitHub source rather than a non-resolving relative path.
For Docusaurus markdown conventions follow doc_style.md; for site build/serve commands follow docsite_maintenance.md.
The branch model (develop/staging/master) and how the smart test system selects which pipelines to test are documented in TestingPipelines.md. Operational specifics:
-
Path-based triggers — GitHub Actions workflows trigger on the paths a pipeline depends on:
paths: - 'pipelines/wdl/optimus/**' - 'tasks/wdl/StarAlign.wdl' - 'tasks/wdl/Metrics.wdl' - 'tasks/wdl/Utilities.wdl'
When adding a regression test input, wire its path into the relevant workflow so CI picks it up.
-
Test WDL naming & contract — see Registering a CI test for a pipeline above for the full
Test<Pipeline>.wdlcontract (naming, imports,GetValidationInputs,update_truth). -
Release builds —
./scripts/build_pipeline_release.sh -w pipeline.wdl -v version -o output_dir.
When resolving merge conflicts in WARP:
- WDL files first — resolve WDL pipeline conflicts before configuration files.
- Version precedence — use the higher version number from develop, then increment by 0.0.1.
- Import path updates — check and fix import statements after resolving conflicts.
- Validation required — run womtool validation on all modified WDL files.
# 1. Resolve all .wdl files first
# 2. Handle .dockstore.yml
# 3. Process changelog files
# 4. Validate all WDL files with womtool
# 5. Commit and pushAfter merging develop, grep each touched WDL for duplicate pipeline_version declarations, which merge commits can silently introduce:
grep -n 'pipeline_version' pipelines/wdl/*/*.wdl | grep -v '\.changelog' | awk -F: '{print $1}' | sort | uniq -dKeep the higher version and delete the duplicate line.
- No legacy
tasks/broad/ortasks/skylab/import paths remain - No duplicate
pipeline_versiondeclarations in any WDL - All WDL files pass womtool validation (run the loop above)
- Changelogs updated with current date and description
-
.dockstore.ymlpaths reflect current structure
pipeline_versions.txt— master version tracking.dockstore.yml— Dockstore registry configuration.github/workflows/— CI/CD workflowswreleaser/— CLI tool for querying releasesscripts/common.sh— shared build functionsscripts/validate_wdls.sh— convenience womtool wrapperverification/test-wdls/scripts/— test generation utilities
When a pipeline emits per-sample artifacts, accept a String input_id workflow input and prefix every output filename with ~{input_id}_. This matches the convention used by scANVI, Optimus, Multiome, and other Skylab-origin pipelines.
When adding support for a new 10x chemistry in Optimus:
- Use
tenx_chemistry_version(Integer, e.g.4) for the major chemistry version andtenx_chemistry_subversion(optional String, e.g."v4_TRU") for whitelist variant selection within that version. - Optimus whitelist files live at
gs://gcp-public-data--broad-references/optimus_whitelists/. Do not use the oldRNA/resources/path (which contained afebrarytypo and is incorrect). - Validate that inputs required by a given chemistry (e.g.
i1_fastqfor v4) are enforced viaErrorWithMessagewhen the chemistry version demands them.
Shrink Plumbing/Scientific inputs by subsetting reads to a genomic region (e.g. one chromosome), which keeps the called-cell count high. Do not subsample down to a handful of barcodes: the library-level doublet step sets its KNN neighbor count to k = round(min(100, n_called_cells × 0.01)), so below 50 cells 500 called cells. See the Minimum cell count warning.k becomes 0 and the run dies with sklearn's n_neighbors ... Got 0 error. Keep ≥
Never hand-write .ipynb JSON. Build notebooks programmatically with nbformat (markdown + code cells as plain strings), then execute them with jupyter nbconvert --to notebook --execute --inplace … so the delivered notebook has its figures and tables rendered inline. If the (often containerized) environment lacks nbconvert/ipykernel, install them there first — a scientific report container with scanpy/matplotlib usually has network for pip. Deliver an executed notebook or stop and ask — never ship an unexecuted notebook or one assembled by hand.
Record it here, under Agent-Specific Notes — this file is the only home for agent guidance. Keep entries short and link out to the canonical reference rather than duplicating it. Before adding a new note, check whether the information is already covered above or in a linked doc — prefer a one-line reference over a duplicate section.
If the user requests updating the AGENTS.md file with what you have learned, review this file and propose concise additions for future agents — see the reflection prompt in AGENT_reflections.md.