Evidence-grounded extraction, grounding, and analytics for organoid culture protocols.
Turn organoid-culture papers into structured, queryable, evidence-grounded protocol records — the axes along which protocols actually differ (source cells, matrix, base media, signaling cocktail, timeline, passaging, endpoints) — where every extracted value that can carry an evidence span is backed by a verbatim quote from the source and a DOI.
This is not a scraper. The work is in what's hard: entity normalization, distinguishing not reported from not extracted, cross-modal confirmation from figures, and detecting protocols stated by reference to a cited paper. The guiding rule throughout is missing evidence beats false evidence — values without grounding are dropped or flagged, never fabricated.
Everything runs on a local A100 (models via ollama) — no API, no keys. Demo
- 582 papers · 25 organoid types (intestinal, cerebral, cardiac, kidney, liver, lung, pancreatic, thyroid, retinal, gastric, and more)
- 5,458 grounded reagent records across the full corpus
- Schema v0.4:
FailureMode,ProtocolModification,Evidence.sentence_id - Public exports:
exports/public/protocols.jsonl,exports/public/reagents.jsonl,exports/public/manifest.json
paper (PMC)
└─ Tier 0 Europe PMC JATS XML → methods + supplement + tables + figures + refs (deterministic)
└─ Tier 1 local LLM (gemma3:12b) → OrganoidProtocol JSON, each value verbatim-grounded
└─ Tier 2 local vision (gemma3:12b) on figure schematics → cross-modal figure-confirmed factors
└─ Tier 3 detect protocols delegated to a citation ("…as previously described (Sato 2011)")
→ resolve + verify the cited paper (human-review queue; never auto-attribute provenance)
└─ S1 live SRI / Cellosaurus grounding → CURIE resolution with three-state reporting
└─ S2 Biolink-validated KGX export (nodes.tsv + edges.tsv + kgx_manifest.json)
└─ entity normalization (bFGF≡FGF2, RSPO1≡R-spondin1, …)
└─ analytics pipeline → coverage · quality · consensus · failure modes · lineage · assay endpoints
└─ REST API (Datasette plugin, 62 routes)
All endpoints return JSON and degrade gracefully (404 + hint when not yet computed).
Use these 8 first — they cover the most common protocol-intelligence queries:
| Endpoint | Purpose |
|---|---|
/analytics/summary |
Corpus stats, quality distribution, top types |
/analytics/coverage |
Per-type coverage and completeness |
/analytics/reporting-gaps |
Field reporting rates — transparency audit of systematic gaps |
/analytics/consensus/{type} |
Consensus concentrations + reagents for one type |
/analytics/reagent?q= |
Cross-corpus reagent lookup with evidence quotes |
/analytics/concentration-by-type?q= |
Per-type dose stats for one canonical reagent |
/analytics/within-type-dose-range |
Intra-type dose disagreement ranking |
/analytics/protocol-completeness |
Per-paper completeness scores (0–6) |
Agents and crawlers: use /llms.txt, /analytics/*, and public exports (exports/public/).
Do not scrape Datasette table pages or JS-rendered dashboard/consensus/heatmap pages row-by-row.
GET /analytics index of all endpoints + generate commands
GET /analytics/summary dashboard: corpus stats, quality distribution, top types
GET /analytics/status live system health (corpus + artifact inventory)
GET /analytics/consensus list available per-type consensus files
GET /analytics/consensus/{type} consensus concentrations + reagents + timeline for one type
GET /analytics/coverage per-type corpus coverage and completeness report
GET /analytics/coverage/{type} coverage for one organoid type
GET /analytics/quality per-paper quality scores (gold ≥ 0.80 / silver ≥ 0.55 / bronze)
GET /analytics/reagent?q=TERM cross-corpus reagent lookup: usage, concentrations, evidence quotes
GET /analytics/reagent-network?q=TERM reagent co-occurrence: which reagents most often appear in the same papers
GET /analytics/type-similarity pairwise organoid type Jaccard similarity on canonical reagent sets
GET /analytics/type-timeseries organoid type publication counts by year — growth trends + first-appearance dates
GET /analytics/universal-reagents canonical reagents in >= 50% of protocols per type + cross-type universals
GET /analytics/species-breakdown species distribution per organoid type (human / mouse / other) from protocols.jsonl
GET /analytics/matrix-breakdown extracellular matrix usage per organoid type (Matrigel / Geltrex / Vitronectin / ...) with alias normalisation
GET /analytics/base-media-breakdown base media usage per organoid type (DMEM/F12 / mTeSR1 / Advanced DMEM/F12 / ...) with alias normalisation
GET /analytics/source-cell-breakdown source cell type distribution per organoid type (iPSC / adult_stem_cell / primary_tissue / ESC)
GET /analytics/protocol-complexity per-type protocol complexity: avg signaling factors / supplements / figure-confirmation / grounding rate
GET /analytics/reporting-gaps field reporting rates (species/matrix/base_media/passaging/timeline) — transparency audit of systematic gaps
GET /analytics/year-trend yearly trends: paper count, avg signaling factors, avg grounding rate, field reporting rates by publication year
GET /analytics/grounding-quality reagent grounding coverage: grounding_rate, evidence_quote_rate, suspect_unit_count by type and kind; top ungrounded canonical names
GET /analytics/concentration-stats aggregate concentration stats per canonical reagent: median/min/max/std by unit; top 50 by n_with_value; ?q= for one reagent
GET /analytics/temporal-reagent-adoption per-reagent temporal adoption: fraction of papers per year using each canonical reagent; ?q= for year-by-year data, ?type= for one type
GET /analytics/kgx-summary KGX graph state: node/edge counts by category, resolution rate, review queue breakdown (needs_review/not_found), top unresolved entities
GET /analytics/concentration-by-type per-organoid-type concentration stats for one canonical reagent — median/min/max/n per unit per type; ?q=EGF required
GET /analytics/journal-breakdown journal contribution counts: cross-corpus top 50 + per-type top 5; ?type=kidney for full breakdown of one type
GET /analytics/type-comparison side-by-side organoid type comparison: shared/unique canonical reagents, Jaccard, per-kind breakdown; ?a=intestinal&b=cerebral
GET /analytics/concentration-deviation dose inconsistency ranking: canonical reagents sorted by coefficient of variation (std/mean); most_variable + most_consistent lists; ?min_n= threshold
GET /analytics/reagent-prevalence type-breadth ranking: canonicals sorted by n_organoid_types they appear in; cross_field + specialist sub-lists; ?q=EGF for per-type breakdown; ?min_types= threshold
GET /analytics/protocol-outliers per-type outlier detection on n_signaling_factors: complex/minimal protocols with z-scores; ?type=kidney for one type; ?z_thresh= sensitivity (default 1.5)
GET /analytics/grounding-distribution per-paper grounding rate histogram (10 buckets), per-type mean ranking, top/bottom 20 papers; ?type=kidney for one type; live from protocols.jsonl
GET /analytics/type-maturity field maturity per organoid type: first_year, n_years_active, trajectory (accelerating/stable/slowing), maturity_tier (established/developing/emerging)
GET /analytics/reagent-cooccurrence pairwise signaling-factor co-occurrence: top pairs by n_papers + Jaccard; ?q=EGF for all partners; ?type= filter; ?min_papers= threshold
GET /analytics/supplement-breakdown per-type and cross-type breakdown of supplement canonicals: global top 50, cross-type list, per-type top 10; ?q= and ?type= filters
GET /analytics/role-breakdown normalized functional role distribution for signaling reagents: signaling_factor/growth_factor/differentiation/inhibitor/agonist etc.; ?q= for top canonicals per role; ?type= filter
GET /analytics/type-reagent-heatmap type × canonical reagent usage matrix (top_n canonicals × all types, cell = n_papers); ?kind=signaling|supplement|all; ?top_n= (default 20, max 50)
GET /analytics/canonical-name-variants normalization complexity: canonical → all raw names, top 30 by n_variants; ?q= for one canonical; ?min_variants= threshold
GET /analytics/concentration-unit-distribution unit inconsistency: canonicals using multiple unit systems, top 30 by n_units; ?q= for one canonical with min/median/max per unit
GET /analytics/protocol-size-distribution full histogram of protocol sizes: n_signaling_factors and n_supplements per paper; global + per-type mean/median/std; ?type=kidney for one type
GET /analytics/evidence-quote-coverage per-type and per-kind rate of verbatim evidence quotes in reagent records; overall_coverage_rate + by_kind breakdown; per_type sorted by coverage_rate; ?type= for top canonicals; ?kind=signaling|supplement
GET /analytics/concentration-value-rate canonicals ranked by fraction of records with numeric dose value; highest_reporters + lowest_reporters (top 30 each); ?q=EGF for per-type breakdown; ?min_n= threshold; ?kind= filter
GET /analytics/kind-ambiguity canonicals appearing in both signaling and supplement kinds; sorted by minority_fraction; ?q=Y-27632 for per-type kind breakdown; ?min_n= threshold (default 3)
GET /analytics/canonical-type-adoption reagent diffusion: n distinct organoid types using each canonical by year; first_year, n_types_current, year_peak; ?q=EGF for per-year type list + cumulative; ?min_types= threshold (default 5)
GET /analytics/unit-normalization-report audit of raw unit → canonical_unit clusters; sorted by n_raw_strings (most ambiguous first); ?q=uM for detailed raw strings + top canonicals using that unit
GET /analytics/source-cell-reagent-profile characteristic reagents by source_cell_type; top 20 per source + pairwise Jaccard; ?source=iPSC for top 30 + exclusive_to_source flags; ?min_papers= threshold (default 3)
GET /analytics/protocol-completeness per-paper completeness scores (0-6) across species/matrix/base_media/passaging/timeline/assay_endpoints; histogram + per-type ranking + top/bottom 20 papers; ?type= for one type
GET /analytics/cross-type-concentration-variance canonicals where dose differs most across organoid types; sorted by max/min per-type-median ratio; ?q= for per-canonical detail; ?min_n= (default 3)
GET /analytics/reagent-type-enrichment enrichment ratio (type-rate / global-rate) per canonical per type; ?type=retinal top enriched (taurine 14.85x, blebbistatin 11.83x); ?q= all types; global=top 50; ?min_n= (default 3)
GET /analytics/grounding-inconsistency S1 target: 91 canonicals grounded in some papers but ungrounded in others (Y-27632 32%, GlutaMAX 98%); ?sort=total|rate|n_ungrounded; ?min_n= (default 5)
GET /analytics/canonical-merge-candidates canonical names likely referring to the same entity (top 100 groups by combined record count); ?min_records= threshold (default 3)
GET /analytics/grounding-by-kind grounding rate by reagent kind (signaling vs supplement): reveals 0% supplement grounding gap; ?type= ?kind=
GET /analytics/temporal-variance concentration CV trends over time for a canonical reagent; ?q=CHIR99021 required; ?min_n= (default 3)
GET /analytics/base-media-cooccurrence conditional P(base_media | source_cell_type) + imputation candidates; ?source= for one cell type
GET /analytics/assay-endpoints assay endpoint cluster summary (12 clusters, per-type + cross-type)
GET /analytics/failure-modes failure mode cluster summary across the corpus
GET /analytics/lineage DOI→DOI protocol lineage graph (ProtocolModification data)
GET /analytics/compare/{a}/{b} protocol diff between two papers (pre-computed cache)
GET /analytics/substitutions?q=TERM search ProtocolModification records for reagent substitutions
GET /analytics/mior MIOR completeness per paper + corpus (12 items, 5 modules)
GET /analytics/candidates OA/license verification status of candidate pool (issue #14)
GET /analytics/within-type-dose-range within-type fold-range per canonical × organoid type: min/max/fold_range/CV; ranks intra-type dose disagreement (e.g. FGF2 kidney 6–200 ng/mL = 33×); ?q= ?type= ?unit= ?min_n=
GET /analytics/convergence-leaders canonicals ranked by temporal CV trend: converging (consensus emerging) vs diverging (dose disagreement growing); surfaces BMP4 retinal-style success stories; ?min_years= ?min_n=
POST /trapi/query single-hop Biolink query over the committed KGX graph
GET /trapi/meta_knowledge_graph KGX summary: node categories, predicates, edge counts
GET /trapi HTML explainer + interactive try-it console
The TRAPI endpoint serves the committed exports/kgx/{nodes,edges}.tsv as a live-queryable
TRAPI 1.5 graph. Nodes carry Biolink CURIEs
resolved via SRI; edges use biolink:mentions predicates. See serve/plugins/trapi_endpoint.py.
Generate all analytics outputs:
make all-analytics # regenerate everything in dependency order
# or individually:
python pipeline/generate_coverage_report.py # → outputs/analysis/coverage_report.json
python pipeline/score_protocol_quality.py # → outputs/analysis/protocol_quality_scores.json
python pipeline/compute_consensus.py --all # → outputs/analysis/consensus_*.json
python pipeline/aggregate_failure_modes.py # → outputs/analysis/failure_mode_summary.json
python pipeline/build_lineage.py # → outputs/analysis/protocol_lineage.json
python pipeline/aggregate_assay_endpoints.py # → outputs/analysis/assay_endpoint_summary.json
python pipeline/score_mior.py # → outputs/analysis/mior_completeness.json
python pipeline/check_concentration_consistency.py # → outputs/validation/concentration_consistency.json
python pipeline/system_status.py # check what's missing- Every reagent's
evidence_quoteis a verbatim substring of the source — never paraphrased. - Three-state
grounding_status:resolved(real CURIE from SRI/Cellosaurus),not_found,not_attempted. Resolved means a real service response was cached as a fixture. - Tier 2 adds figure factors only as a
figure_confirmedannotation — vision corroborates, it doesn't inject unverified data. - Tier 3 emits a review queue, never auto-ingested records.
- Implausible units (e.g. a growth factor in mg/mL) are flagged, not silently "fixed".
- No metric, count, or rate appears in docs unless generated by a committed artifact.
reported/not_reported/not_extracted/not_applicabledistinctions are preserved.
pipeline/
tier0_extract.py Tier 0: Europe PMC JATS → evidence bundles
tier1_extract.py Tier 1: local-LLM structured extraction + verbatim grounding
tier2_vision.py Tier 2: local vision on figure schematics (cross-modal)
fetch_figures.py figure-image acquisition (PMC OA AWS S3 mirror; local-only)
tier3_detect.py Tier 3: detect delegated-citation protocols
tier3_resolve.py Tier 3: resolve + verify cited paper (review queue)
ground.py S1: live SRI Name Resolver + Cellosaurus grounding (cached fixtures)
export_kgx.py S2: Biolink-validated KGX export (nodes/edges TSV)
normalize.py reagent entity canonicalization
ingest_orchestrator.py discovery → QC → ingestion pipeline
ingestion_auth.py R4: ingestion authorization gate (who can ingest what tier)
citation_expand.py citation expansion (expand references of accepted papers)
hybrid_discover.py semantic + lexical hybrid discovery
semantic_index.py dense semantic index (sentence-transformers, A100)
compute_consensus.py consensus reagents/concentrations per organoid type
aggregate_failure_modes.py failure mode cluster aggregation
build_lineage.py DOI→DOI protocol lineage graph
generate_coverage_report.py per-type coverage + completeness scoring
score_protocol_quality.py per-paper quality scorer (gold/silver/bronze)
aggregate_assay_endpoints.py assay endpoint cluster analysis (12 clusters)
reagent_lookup.py cross-corpus reagent search with concentration stats
compare_protocols.py pairwise protocol diff
find_substitutions.py ProtocolModification substitution search
system_status.py system health CLI (corpus + analytics artifact inventory)
trapi.py minimal TRAPI responder shape
export_public.py export public protocols/reagents JSONL snapshots
validate_evidence.py evidence fidelity validator (verbatim substring checks)
validate_predictions.py prediction file schema validator (v0.4, offline, pre-PR gate)
relabel_organoid_type.py rescue corpus.tsv 'other' rows to discovery-CSV type (idempotent, --dry-run)
audit_units.py unit plausibility audit (R2: concentration vs. in-vivo/volume/percent)
check_concentration_consistency.py cross-paper concentration outlier detection (≥10x median)
score_mior.py MIOR completeness scorer (12 items, 5 modules, per-paper + corpus)
score_protocol_quality.py per-paper quality scorer (gold/silver/bronze)
ground_predictions.py S1→S2 handoff: ground prediction entities, write sidecars
discover_candidates.py keyword-based candidate discovery
serve/
run.sh serve the atlas (Datasette + plugins)
metadata.yaml facets + canned queries
plugins/
analytics_endpoint.py 62-route analytics REST API (pure handlers + thin Datasette wrappers)
ask.py grounded Q&A (RAG over FTS → local model)
templates/ landing, recipe cards, /heatmap, /consensus
static/atlas.css|js theme + dark-mode toggle
exports/
public/
protocols.jsonl 582 papers, 25 organoid types (public snapshot)
reagents.jsonl 5,458 public reagent rows (DOI-linked evidence where available; missing evidence preserved, not invented)
manifest.json counts + schema version
kgx/
nodes.tsv KGX nodes (Biolink categories + CURIEs)
edges.tsv KGX edges (biolink:mentions predicates)
kgx_manifest.json counts + validation report
data/
corpus/corpus.tsv PMC corpus manifest
corpus/incoming/ candidate CSVs (QC-gated, not yet ingested)
predictions/local/ Tier 1/2 predictions (A100, git-ignored)
predictions/local/grounded/ S1 grounding sidecars (git-ignored)
outputs/
analysis/ pre-computed analytics (coverage, quality, consensus, etc.)
kgx/ KGX graph export
comparison/ pre-computed protocol diffs
tests/ offline test suite (1317 tests, no network, no GPU)
docs/ SUPERVISOR_CHECKLIST.md, PLAN, RESEARCH_BRIEF
Full text, figure images, and predictions are local-only (git-ignored); only metadata, count-level summaries, public snapshots, and grounded CURIEs are committed.
# serve the atlas
./serve/run.sh # → http://localhost:8002
# pipeline (local A100 + ollama)
python pipeline/tier0_extract.py # evidence bundles
python pipeline/tier1_extract.py # extraction + grounding
python pipeline/tier2_vision.py # figure confirmation
python pipeline/ground.py # S1: live SRI/Cellosaurus grounding
python pipeline/export_kgx.py # S2: KGX graph export
python pipeline/system_status.py # check health / what needs generating
# analytics pipeline (no GPU needed)
python pipeline/compute_consensus.py --all
python pipeline/generate_coverage_report.py
python pipeline/score_protocol_quality.py
python pipeline/aggregate_failure_modes.py
python pipeline/build_lineage.py
python pipeline/aggregate_assay_endpoints.py
make test # run offline test suite (1317 tests)
make validate-batch # pre-PR check: tests + prediction schema + evidence
# or: pytest -q