Purpose. Build or extend a bibliography for an academic review topic, and optionally render it as a figure and a written review. This is ONE tool with two front-ends — topic mode (start from a query) and lab mode (start from a lab's corpus) — that share the entire downstream pipeline; only the front-end differs. Per topic, aim for ~50-70 high-impact and recent papers, classified and summarized. Decide the mode first (Phase 0), then gather that mode's inputs.
These are the load-bearing invariants. Everything below the contract is reference detail that elaborates them; when in doubt, obey this list. Section pointers are in parentheses. Treat bold emphasis elsewhere in this file as ordinary guidance — the genuinely inviolable rules are only the eight here.
- Verify EVERY citation before it enters a deliverable (Phase 3). About 1 in 4 agent-returned refs has a fabricated author list, wrong year, reversed conclusion, or bad DOI. No exceptions — preprints included.
- Every reference is canonical (Phase 3f). Rebuild each
apafrom the verified DOI/arXiv withreferences.py; never ship an agent-typed or OpenAlex-typed string.references.py --auditis a hard gate (exit 1) — run it before every deliverable. - One row per DOI — global dedup; a paper appears once in
rows.json. Bites hardest at the lab-mode merge, where one paper surfaces under several theme-searches and a lab paper can resurface as "field" (Phase L4c). (Distinct from the one-family-per-paper rule, whichfamilies.pyenforces automatically — Phase 6b. This contract item is only about row-level deduplication.) - Run the antecedents pass on every review, both modes (Phase 2b). The forward search misses the topic's methodological, empirical, and theoretical roots; without it the field looks ~10 years old.
- Audit the temporal order of ideas before delivering any written review (Phase 7). Origin claims must cite the EARLIEST deserving paper, oldest-first — not whichever ref fits the sentence.
rows.jsonis the live table after Phase 3f. Edit it by hand for any later change; never re-run the row-emitter (it wipes canonicalapa+ citation counts).- Don't ask before fetching from PubMed/PMC/CrossRef/OpenAlex/Unpaywall/arXiv/
publishers — these are read-only academic GETs; do confirm destructive or
shared-state actions. Every link is a bare
https://doi.org/<doi>(never a libproxy URL). Set a contact email (LITREVIEW_EMAILor--email) for the API User-Agent. - PDFs are opt-in (Phase 4) — default no, and never ask whether to fetch them.
Default tier criteria. Pre-2021: only highly cited / foundational. 2022+: promiscuous (no citation-count gate — too recent to have accrued cites). The boundary is "today minus ~5 years"; advance it as the calendar moves.
There is a public documentation website built from docs/ (MkDocs + Material),
live at https://gallantlab.org/literature-review-toolkit/. It is a superset
of this PLAYBOOK and the README, not a fork. When you change the toolkit — a
tool, a phase, a command/flag, a guardrail or lesson — update the matching page
under docs/ in the same change (fastest to drift: docs/phases.md,
docs/pipeline.md). The tool index in docs/tools.md, tools/README.md and
this file is generated — run python3 tools/gen_docs.py after adding a tool
or a flag; .github/workflows/tests.yml fails on a stale copy, and also runs
ruff check . and tools/tests/test_formatting.py. The site auto-deploys via
.github/workflows/docs.yml on push to main (build runs mkdocs build --strict). Full editing/figure/snippet details live in docs/maintaining.md.
The repo is public (that's what enables free Pages); site_url uses the org's
gallantlab.org custom domain, not github.io.
One tool, two front-ends. Everything after the front-end — verify, citation counts, families, figure, spreadsheet — is the SAME shared machinery, run the same way. There are no mode-specific shortcuts.
| Mode | Start from | User says… | Front-end | Then gather |
|---|---|---|---|---|
| Topic | a query/topic | "lit review on X", "extend the bibliography for Y" | Phase 1 (scope) → 2 (search) → 2b (antecedents) | topic name + 1-paragraph definition, source doc if any, target spreadsheet path, tier criteria |
| Lab | a lab's publications | "review lab Z's work", "how has Z's research evolved" | Phase L1–L3 (ingest corpus → derive themes) → L4c (+ 2b antecedents) | the lab/author ids, the inclusion filter (e.g. human-only), target paths |
Phase 2b (antecedents) is required in both modes — the forward search misses a topic's methodological, empirical, and theoretical roots; do not skip it.
Both then converge on the shared pipeline: Phase 3 verify → 3f canonicalize refs → 5 spreadsheet → 5b citation counts → 6 cross-citation → 6b families → 7 review article (optional) → 8 hand-off. Lab mode's outward/contextualize layer (L4c) is not a lighter pass — it runs the topic-mode front-end (Phases 2–6) once per theme, with the identical verify/count/dedup guardrails. Topic mode is the next section; lab mode is under "Lab mode" below.
- New rows appended to
<spreadsheet>.xlsxwith columns:Topic | Ref# | APA reference | Link | Summary | Tag | Family | Cite (OpenAlex) | Cite (S2) | PDF (local) | Xref.Linkis always the DOI URL (https://doi.org/<doi>).Family(Phase 6b) and the twoCitecolumns (Phase 5b) are auto-added byspreadsheet.pywhenever rows carry them. citation_counts.json— per-paper OpenAlex + Semantic Scholar counts (Phase 5b)families.json+families.md, and<topic>_families.{html,svg,png,pdf}— the theoretical grouping and its interactive figure (Phase 6b, optional)- Cross-reference index at
xref_<topic_slug>.json(after Phase 6) - Only if Phase 7 was opted into: a narrative review article
<Topic>_review.docx(prose authored intocontent.json, rendered withtools/review_paper.py; APA-7 reference list pulled fromrows.json). - Only if Phase 4 was opted into:
- PDFs at
papers/<topic_slug>/<paper_slug>.pdf - Browser-helper page
papers/<topic_slug>/_download_helper.htmlfor paywalled / bot-blocked papers
- PDFs at
(The query-driven front-end. Lab mode reuses Phases 3–7 verbatim; see "Lab mode" below.)
1a. Read source if provided. If the user has a source doc (.docx/.pdf),
extract text. For docx: unzip -p X.docx word/document.xml | python3 strip_xml.py.
Identify which references are actually cited in the main text (not just in
the bibliography). The bibliography may have hundreds of refs the doc never
discusses; only main-text-cited ones are baseline.
1b. Define the topic precisely. Write 3-5 sentences of what counts as relevant. Include the contested theoretical positions, the methods / sub-areas / populations involved, and the boundary with adjacent topics. The search agent will use this verbatim.
1c. List "already-known" papers. Pull from the existing spreadsheet
(filter by Topic). The search agent must not re-find these.
Use the general-purpose Agent (or any web-enabled subagent). Give it a
self-contained prompt — it has no context from this conversation. Use the
template in tools/search_prompt_template.md and fill in:
{TOPIC_NAME}and{TOPIC_DEFINITION}{ALREADY_HAVE}— bullet list of existing papers (don't rediscover){TODAY}— current date (gives the agent a recency anchor){TIER_BOUNDARY_YEAR}{TARGET_COUNT}— usually 25-40 papers
The agent should return a numbered list with: APA citation, DOI link in
https://doi.org/<doi> form (not PubMed/PMC URLs), PMCID if available,
3-5 sentence summary, tag (classic/recent-review/recent-empirical/
recent-method/recent-LLM/recent-theory/recent-clinical), and year.
Do not act on the agent's output yet. It will contain errors. Proceed to Phase 2b, then Phase 3.
The Phase-2 search is biased toward recent work and the topic's current framing, so it systematically misses the literature the topic was built on. A review that omits its antecedents reads as if the field began ~10 years ago. Run a dedicated antecedents pass in both modes (contract rule 4), after the main search and before verifying.
Spawn a separate search agent per axis for the topic's intellectual roots:
- Measurement / methodology origins — the instrument, signal, or technique the work depends on, and the papers that established and validated it (e.g. for human-fMRI work: the BOLD mechanism, the first functional studies, what the signal actually measures).
- Foundational empirical results — the classic findings the topic builds on, including older work in adjacent methods, species, or eras that the forward search's recency bias skips (e.g. single-unit neurophysiology, psychophysics, the first description of an effect or region).
- Theory / computational framework — the conceptual claims that motivate the work (e.g. efficient coding, a normative principle, a levels-of-analysis framing).
Reuse tools/search_prompt_template.md, but flip the tier emphasis: the
target here is foundational / highly-cited / classic work that PRE-DATES the
modern literature, not recent papers. Give each agent the already-have list
(now including the Phase-2 results) so it does not re-find them, and have it tag
each paper with the best-fit existing theme/family. Antecedents fold into the
existing lanes by default — do NOT spin up new lanes for them unless the user
asks. Feed every returned paper through Phase 3 → 3f → 5b like any other.
Old classics often have no DOI (pre-2000 papers, books, book chapters). Keep
them as hand-written canonical APA no-source rows (references.py flags them;
the audit gate allows them) and exclude them from citations.json — the same
pattern as any DOI-less item. Beware reissue DOIs for old books (they re-date the
work to the reprint year); prefer a hand APA citing the original edition. Verify
each by title/author against the publisher or a library record before trusting
the agent's APA.
Lab mode: the antecedents include the lab's OWN pre-paradigm work — the
earlier-method, other-species, or pre-tool publications that the inclusion filter
(Phase L2) drops. Reconsider that filter: a lab's foundational pre-paradigm papers
are usually the most direct antecedent of its current program, and belong in the
corpus as source=lab (starred) rather than excluded.
Effect on the figure: antecedents widen the time span (often back to the
mid-20th century) while most papers cluster in the last decade. Set the figure's
--min-year to the earliest antecedent and add --time-warp so the sparse early
decades compress and the dense recent years expand — otherwise the modern
literature collapses into an unreadable clump at the right. See Phase 6b.
In a previous run, the search agent fabricated 5 author lists, reversed one paper's conclusion, and invented a bioRxiv DOI that didn't exist. About 1 in 4 citations had errors. Always verify before adding.
verify.py returns one verdict per citation: OK, MISMATCH (author/year),
NOT-FOUND, or ERROR. NOT-FOUND and ERROR are NOT the same and must be
handled differently: NOT-FOUND means every lookup completed and none matched
(chase it down — likely fabricated); ERROR means a lookup could not complete
(rate-limit / network), so re-run those rather than treating them as missing.
On a big run this matters — arXiv rate-limits hard, so the tool prefetches all
arXiv ids in batches (many per id_list call); a genuinely real preprint that
would otherwise 429 into a false NOT-FOUND now comes back OK (or, if the batch
still fails, ERROR to re-run).
Feed it the live table directly — python3 tools/verify.py --rows rows.json --out verify_report.json derives the label, DOI and the expected first author / year /
title from each row's apa; do not write a per-project converter script (fourteen
projects did, each with its own first-author regex).
For each paper the agent returned:
3a. If a PMCID was given: call NCBI esummary
(https://eutils.ncbi.nlm.nih.gov/entrez/eutils/esummary.fcgi?db=pmc&id=<num>&retmode=json).
Confirm first author, year, and title match. See tools/verify.py.
3b. If no PMCID but a title is given: call PubMed esearch
(https://eutils.ncbi.nlm.nih.gov/entrez/eutils/esearch.fcgi?db=pubmed&term=<title>&retmode=json)
then esummary. Confirm match.
3c. If only an arxiv ID: WebFetch https://arxiv.org/abs/<id> and read
title/authors from the page.
3d. If a publisher landing page (Nature, Springer, OUP): WebFetch the page and confirm author/year. Don't rely on the agent's claim.
3e. Drop / fix:
- Citation completely fabricated (URL doesn't resolve, no PubMed match) → drop.
- Wrong first author / wrong year → fix using the verified metadata.
- Title matches but agent's summary contradicts the abstract → fix summary.
- Suspicious DOI (e.g. unusual prefix, no resolution) → drop unless you can confirm via web search.
Common fabrication patterns to flag:
- Author name that doesn't appear in any of the paper's actual authors.
- Conclusion that is the OPPOSITE of the paper's actual finding.
- A DOI that does not resolve, or that resolves to an unrelated paper (agents invent plausible DOIs and also mis-copy real ones — confirm the target, not just that it resolves).
- arxiv preprint IDs that don't resolve.
Don't treat an unfamiliar DOI shape as fabrication evidence — resolve it.
Prefixes and suffix formats drift, so a DOI not matching the shape you expect is
often just a newer pattern, not a fake. Seen in real builds: bioRxiv now issues
10.64898/... DOIs alongside the older 10.1101/...; Imaging Neuroscience uses
10.1162/imag.a.NNNN (dots, not the imag_a_NNNNN underscores you might guess —
the wrong shape 404s). Always judge a DOI by what it resolves to, never by its
string.
Verification (3a–3e) confirms a citation is real; this makes its apa string
perfect (contract rule 2). Never ship a reference typed from an agent's memory
(topic mode) or OpenAlex's light metadata (lab mode) — rebuild every apa from the
verified DOI against the authoritative source. references.py is the single
canonical formatter and a hard gate, used identically in both modes:
python3 tools/references.py --rows rows.json --out rows.json # rebuild + report
python3 tools/references.py --rows rows.json --audit # gate: exit 1 on any defect
It pulls CrossRef (DOIs) or the arXiv API (arXiv ids / 10.48550/arXiv.* DOIs),
then builds APA-7 with: full author list (>20 → 19 + ellipsis + last), correct
initials and nobiliary particles (de Heer, Dupré la Tour), fixed name casing
(ANDERSON→Anderson, zhang→Zhang), HTML-unescaped + sentence-cased
all-caps titles, and a real venue — including preprint servers CrossRef leaves
bare (bioRxiv, PsyArXiv, arXiv). The --audit gate fails the build on any
defect (missing author/year, et al., HTML entity, U+FFFD replacement-char
mojibake, truncated/empty venue, uppercase title, JATS/HTML markup left in a
title (<scp>, <i>), a ?. or !. double terminal punctuation, and a
U+2010/U+2011 Unicode hyphen in a name). The ONLY allowed non-fatal
case is a DOI-less item (book, report, old proceedings) — it keeps its
hand-written apa and is reported as a
manual ref; verify those by hand. Run the gate before every deliverable.
Two things the gate reports as warnings, because neither can be decided
automatically: a near-duplicate row pair, and a multi-word surname that may be
a mis-split given name. Lambon Ralph is a real compound surname and Thomas Yeo
is CrossRef folding B. T. T. Yeo's given names into the family field; they are
indistinguishable to a machine, so each needs a human verdict. (A leading initial
in a family field — CrossRef's family="A. Moffat" — IS unambiguous and is now
repaired automatically, since no surname begins with an initial.)
Sentence-case titles after canon with tools/sentence_case.py. Strict APA-7
wants sentence case, and canon deliberately does not impose it (see Lessons). The
tool proposes, you review, then --apply. Keep the corpus's proper nouns in a
per-project --proper allowlist file so a generic word lowercases while a named
entity does not (yoga practitioners but Sahaja Yoga). On a large corpus use
--vocab to review the ~N distinct token changes rather than 150 title diffs — a
mis-cased proper noun is obvious there and invisible in a long diff.
Skip this phase by default. Run only if the user explicitly asks for PDFs. The default workflow is Phase 1 → 2 → 3 → 5 → 6 → 7. PDF acquisition will eventually be replaced by a separate dedicated tool; treat the machinery below as legacy that still works on demand.
If opted in, try sources in this order (tools/download.py does this
automatically):
-
arxiv direct —
https://arxiv.org/pdf/<id>.pdf. Always works for arxiv preprints. Only one risk: rate-limit (429) if you hit too fast; use 2s sleeps between calls. -
Unpaywall API —
https://api.unpaywall.org/v2/<doi>?email=<user_email>. Returnsoa_locationswith PDF URLs. Prefer non-PMC URLs first, since PMC has aggressive bot blocking. Often gives author institutional repos (.edu/.ac.ukpages) that work with simple curl. -
Direct journal URL via Unpaywall's
best_oa_location.url_for_pdf—https://www.nature.com/articles/<id>.pdftypically works for OA Nature, Nat Commun, Nat Neuro, Nat Hum Behav, Sci Rep. -
Europe PMC —
https://europepmc.org/articles/<PMCID>?pdf=renderworks for many NIH-funded papers. -
Manual fallback via browser-helper page (preferred over
_needs_manual.txt). For papers that fail the auto-download, generatepapers/<topic>/_download_helper.html: one row per failed paper with author/year/slug/title and an Open link to the journal landing page (use plainhttps://doi.org/<doi>for paywalled — the user has institutional access; do not wrap in libproxy URLs, those land on a generic library page). Use direct PMC/articles/<PMCID>/URLs for OA-on-PMC papers andhttps://www.biorxiv.org/content/<doi>v1for bioRxiv preprints. Open the helper withopen <path>so it loads in the user's browser. The user clicks through, downloads each via the publisher's own PDF button, PDFs land in~/Downloadswith publisher-chosen filenames. Then runtools/reconcile_downloads.py --manifest <topic>/_manifest.json --out-dir papers/<topic>/to read each PDF's first-page title via pdftotext, fuzzy-match to the manifest, and move into place with the right slug name.
Verify each download is actually a PDF (first 4 bytes == %PDF).
A 200 response can still return an HTML challenge page.
Do NOT attempt these sources — they all reliably fail to bots:
- PMC direct PDF URLs (
https://pmc.ncbi.nlm.nih.gov/articles/<PMCID>/pdf/): Cloudflare Proof-of-Work challenge. - bioRxiv / medRxiv direct: Cloudflare bot mitigation (403).
- PNAS direct PDF (
pnas.org/doi/pdf/...): 403 via curl. - OUP
academic.oup.com/.../article-pdf/...: 403. - MIT Press
direct.mit.edu/imag/article-pdf/...: 403. - Elsevier ScienceDirect
.../pdfft: 403. - Wiley
onlinelibrary.wiley.com/doi/pdfdirect/...: 403.
These all work fine in a real browser, so route them to the helper page described in step 5 — don't keep retrying programmatically.
Use xlsxwriter (no install if already present; if not, write CSV instead
and tell the user). Schema:
| Topic | Ref # | APA reference | Link | Summary | Tag | Family | Cite (OpenAlex) | Cite (S2) | PDF (local) | Xref |
|---|
(Family appears only after Phase 6b, and the two Cite columns only when
Phase 5b has populated them.)
Topic: one of the project's topic categories (e.g. "Multimodal networks").Ref #: numeric for source-document refs; use<topic-letter><n>for added refs (e.g.M1-M40for first multimodal batch,M41-M70for xref batch). Keep numbering monotonically increasing across batches.APA reference: the canonicalapafrom Phase 3f — full author list (APA-7: up to 20; 19 + ellipsis + last beyond that). Neveret al.; the audit gate fails on it.Link: DOI URL inhttps://doi.org/<doi>form — verified to resolve. PubMed/PMC URLs are NOT used as the primary link. If a paper has only a PMID/PMCID, look up its DOI before adding the row.Summary: 3-5 sentences. State what the paper did and why it matters for the topic. Don't just paraphrase the abstract.Tag: see Phase 2 list.PDF (local): relative path if downloaded, else empty.Xref: citation count from cross-reference analysis (Phase 6), else empty.
Color-code rows so origin is visible (source field; the rules live in
spreadsheet.py's COLORS, and an unknown value renders white with a warning):
- White: refs from the source paper (
source-doc). - Cream
#FFF7E0: refs added in the search passes (search). - Green
#E2F0D9: refs added via cross-citation analysis, Phase 6 (xref). - Blue
#DDEBF7: the lab's own papers in lab mode (lab). - Lilac
#F3E6F5: Phase-2b antecedents (anteced;anteced-nosrcfor hand-cited classics with no DOI).
tools/spreadsheet.py does the rebuild from a JSON of rows: it freezes the
header, sets the column widths and 110-pt row heights, and adds the Family and
Cite columns when the rows carry them.
Add per-paper citation counts. Google Scholar is not usable — it has no
API and CAPTCHA-blocks automated queries after a handful of requests, so it
cannot be pulled for a whole bibliography. Use tools/citations.py, which
queries two databases by DOI:
- OpenAlex — primary source. Free, no key, reliable, near-complete by DOI, batchable. (Undercounts arXiv-only preprints, which it often files under a separate record from the published version — cross-check those with S2.)
- Semantic Scholar — secondary. Often higher for CS/AI venues and gives an
influentialCitationCount. Its free endpoints rate-limit hard (HTTP 429/400) from shared IPs and silently drop papers; treat as best-effort. SetS2_API_KEYin the environment to make it reliable.
python3 tools/citations.py --rows rows.json --out citation_counts.json \
--email you@inst.edu --asof <YYYY-MM-DD>Then attach the counts to each row (cite_openalex / cite_s2 keys) in your
build_data.py/rows pipeline and rebuild — spreadsheet.py auto-adds the two
Cite columns when it sees them. Counts are a snapshot at run time; re-run to
refresh. Papers with no DOI (books, blog/tech-report releases) stay blank.
Per-version data scripts: if you split batch data across importing Python
files, guard the xlsx-writing block under if __name__ == "__main__": so an
import doesn't rewrite the spreadsheet as a side effect — or just keep all rows in
one JSON and rebuild via tools/spreadsheet.py (the simpler path; see Lessons →
On the spreadsheet).
Run after Phase 5 is committed. The point: find high-impact papers the initial search missed by looking at what the papers we DO have cite repeatedly.
6a. Fetch reference lists. For each paper with a DOI, call CrossRef:
https://api.crossref.org/works/<doi>. The message.reference[] field has
the cited refs. Most have a DOI field; some have only unstructured strings.
For papers without DOIs (arxiv-only), fall back to extracting DOIs from the
PDF text via pdftotext -layout <pdf> - | grep -oE '10\.\d+/...'. This is
crude but recovers some.
6b. Build the frequency table. For each cited DOI, count how many of
your N papers cite it. tools/xref.py does this.
6c. Resolve unknowns. Many cited refs have only a DOI in the CrossRef
response, no title/author. Look these up via CrossRef metadata
(api.crossref.org/works/<doi> again, but for the cited DOI).
6d. Filter and select. Take refs cited by ≥4 of your papers
(definite-include) plus selected ≥3-cited foundational classics. Filter
out:
- Refs already in the spreadsheet (check by DOI normalized to lowercase).
- Methods/software citations (SciPy, NumPy, FreeSurfer, fMRIPrep, etc.) unless the topic is methods.
- Off-topic refs that just happened to be popular (e.g. a stats paper).
Aim for ~25-35 additions. More than that and the spreadsheet becomes unwieldy; less and you've under-mined.
6e. Repeat Phases 3-5 for the new batch. Verify every citation, attempt PDF download, append to spreadsheet (with the green color and Xref column populated).
Group the finished bibliography into a few theoretical families — a
conceptual axis orthogonal to the Topic column (Topic captures method/sub-area;
families capture what each paper is fundamentally for). Adds a Family column
and a families.md (grouped tables + a family×topic cross-tab). Run after the
bibliography is assembled, verified, and counted.
This phase has two judgment gates with a human checkpoint between them; the
rest is mechanical, owned by tools/families.py:
- Propose (agent, reading the corpus via
tools/families.py --digest): propose ~3-8 families, each{key, name, claim, lineage}, and state the one organizing principle. The hard constraint: families must cut across the Topic lanes — a good family unites textually-dissimilar papers and splits similar ones. Do NOT cluster embeddings to make families; that yields surface-similarity groups, not theoretical ones. Use the prompt intools/family_prompt_template.md. - Confirm — show the user just the ~6 family definitions for approval/edit. This is the cheap, high-leverage checkpoint: iterating on six definitions is free; redoing the assignment is not.
- Assign — against the frozen spec, assign every paper to one family
(dominant commitment). Assign in batches for large corpora; never one rushed
250-paper pass. Write
families_input.json({principle, families, assignments:{ref:key}}). - Validate + render:
It enforces exhaustive / exclusive / balanced (fails loud otherwise), stamps
python3 tools/families.py --rows rows.json --assign families_input.json \ --out families.jsonfamilyonto rows.json, writesfamilies.json(the reproducible cache, likecitation_counts.json) +families.md, andspreadsheet.pyauto-adds theFamilycolumn on the next rebuild. Re-run only when the taxonomy changes.
The figure is an interactive HTML (not a static png), produced by
tools/families_figure.py from rows.json + families.json:
# first emit the within-review citation graph (criterion 2 below); reuses the xref pass:
python3 tools/xref.py --rows rows.json --out xref_<topic>.json \
--exclude xref_exclude.json --internal-out internal_citations.json --email you@inst.edu
python3 tools/families_figure.py --rows rows.json --families families.json \
--internal internal_citations.json \
--out-prefix <topic>_families --title "<Topic> — theoretical families"It writes a self-contained .html (family lanes with their defining sentences,
every paper as a dot beeswarm-packed by year, landmark studies as big labeled dots;
hover any node for its full reference, click for citation + DOI, hover a family name
to spotlight its lineage) plus a standalone .svg and — if rsvg-convert/inkscape
is present — .png + .pdf for slides/papers. This replaces the old static figure.
Landmark labeling is AUTOMATIC — do not hand-build a labels overlay. A paper is
labeled as a landmark (big dot) if ANY of: (1) it is among the most-cited in its
family (top --per-family, default 4, by max(OpenAlex, S2)); (2) it is foundational
within this review — cited by ≥ --motif-min (default 3) of the corpus's own papers
(this is criterion (2) and needs internal_citations.json from xref.py --internal-out;
silently skipped if absent — so always pass --internal); or (3) it is a home-lab
paper — an author surname listed in --lab-author or the LITREVIEW_LAB_AUTHOR env
var, or a row with source=="lab" — these are starred (★) and gold-ringed so the
lab's own work stands out. Total labels are capped at --max-labels (default 28); what
survives the cap is the home-lab papers plus the top-2 most-cited per family, with the
rest of the budget filled by within-review in-degree.
--motif-min does not scale with corpus size, so watch the drop count. The default
of 3 is tuned for a ~50-paper review. On a 396-paper corpus whose papers cite each other
heavily, 175 papers cleared it and the cap silently discarded 147 of them — a figure
that reads as "here are the landmarks" when it is really "here are 28 of 175". The tool
now prints how many qualified and how many were dropped on every run. If that number
is large, raise --motif-min (25 was right for 396 papers) rather than letting the cap
choose for you.
Home-lab favoring is OFF by default — this is a shared, lab-neutral toolkit, so criterion (3) does nothing until you opt in. Turn it on per project by passing
--lab-author Surname(repeatable), or set it once for your environment withexport LITREVIEW_LAB_AUTHOR=Surname(comma-separated for several surnames). The CLI flag overrides the env var. Rows taggedsource=="lab"(from Lab mode) are always starred regardless of the switch.
Time axis. --min-year clamps the axis start (older papers pin to the left edge).
When the corpus spans many decades but is recency-heavy — the usual shape after a
Phase-2b antecedents pass — add --time-warp <0–1>. It blends the linear axis with
the empirical CDF of all paper years, GLOBALLY (not per-region): sparse early spans
compress, dense recent spans expand. 0 = linear, 1 = full density-equalizing; ~0.85
keeps old foundations legible while decluttering the modern clump. Faint gridlines mark the
labeled years so the nonlinear scale stays readable. Always note the nonlinear axis in the
figure caption (independence principle).
Only the editorial arrows/notes remain a human checkpoint (cross-family convergence
arrows and annotations are judgment). Curate those via an optional --spec figure_spec.json
({arrows:[{from,to,color,label}], notes:[{at,text,color}], order, subtitle}); a labels
map there still overrides auto-selection if you ever need to force a specific set. Don't
expect a good arrow set auto-generated.
Turn the finished corpus into a narrative review article as a .docx. Run only when
the user asks for a written review (not for the bibliography itself). Prerequisites: Phase 3f
(canonical apa) and 5b (counts) are done; ideally Phase 6b families + figure exist too, since
the families are the natural section structure.
Authorship and honesty (non-negotiable when an LLM writes it). If the article is
AI-authored, say so plainly. Put the model's name in authors, add an author_note that
identifies it as an AI, and include a disclosure paragraph stating that the bibliography was
machine-assembled and machine-verified and that the author has read only abstracts/metadata, not
full texts. Language models fabricate citations; the Phase-3/3f verification is what makes an
AI-written review trustworthy, and the disclosure must make that provenance explicit.
Prose. Author the prose with the scientific-writing skill (one idea per sentence, forward
flow, reserve "represent" for brain representations). Organize sections by the Phase-6b
families — the theoretical axis orthogonal to the topic lanes makes a better narrative than the
method/region lanes. The title should convey the question, the answer, and why it matters. Every
in-text citation is APA author–date ((Huth et al., 2016)) and MUST name a paper that exists
in rows.json, so the reference list backs it.
Respect the temporal order of ideas — distil the intellectual history, do not force refs into the narrative (contract rule 5; confirmed by user 2026-06-13). The single most common failure of an AI-written review is crediting the wrong paper for an idea: it picks whichever citation fits the sentence it wants to write, rather than the paper that actually established the idea first. When a sentence makes an origin claim — signalled by emerges, first, established, identified, introduced, was mapped, showed that, had been, early work, foundational, began, demonstrated, discovery — it MUST cite the earliest paper that deserves priority, and order multiple citations oldest-first. Four recurring inversions to watch for (each example is a real miss caught 2026-06-13):
- crediting a later review for a finding an earlier primary paper made (e.g. Tanaka 1996 review vs. Desimone et al. 1984 for object/face selectivity in IT);
- crediting a later model/normalization/synthesis for a phenomenon earlier empirical work established (e.g. Reynolds & Heeger 2009 model vs. McAdams & Maunsell 1999 for attention changing gain/tuning);
- crediting a later, narrower paper while ignoring an earlier, more general one from the same year (e.g. Dumoulin & Wandell 2008 pRF vs. Kay et al. 2008's more general per-voxel model);
- in lab mode, relegating the lab's OWN foundational paper to a later section while a follow-up from another group gets the priority slot (e.g. Hegdé & Van Essen 2000 vs. Gallant et al. 1993 for complex-form selectivity in V4). The lab's antecedents (Phase 2b / L4c) exist precisely so the review can assign priority correctly — use them.
Priority audit (before delivering the review; contract rule 5). After drafting content.json, run a
dedicated audit pass — analogous to the Phase-3 citation verify and the Phase-2b antecedents pass.
Dispatch one agent with the draft prose plus rows.json (which carries every candidate paper and
its year) and instruct it to: scan every origin-claim sentence; for each, check whether an
earlier paper in rows.json (or an undisputed classic) deserves priority for that specific
idea; and report each inversion as claim → currently cites (year) → earlier source (year) → fix.
Apply the confirmed fixes (reorder citations oldest-first, add the originating paper, demote the
later review/model to "later", and adjust wording so the sentence reads as history not narrative).
This pass catches what self-review misses because the drafting model is biased toward its own story.
Mechanics — tools/review_paper.py. The tool owns only the mechanical render; it does not
write prose. It reads rows.json and builds the APA-7 reference list straight from the
canonical apa strings (deduped, alphabetized, hanging indent, with DOI links), embeds the
families figure with a standalone caption, and lays out the title/author/disclosure block + the
abstract + sections. Keep the prose in a small per-project emitter that dumps content.json
(see the schema in review_paper.py); render with the shared tool:
python3 write_review.py # project file: authors prose -> content.json
python3 tools/cite_check.py --rows rows.json --content content.json # GATE: exits 1
python3 tools/review_paper.py --rows rows.json --content content.json \
--figure <topic>_families.png --out <Topic>_review.docxcite_check.py is a gate, not a nicety. The renderer prints whatever prose it
is given, so a citation naming no row in rows.json ships silently and the reader
cannot follow it. The tool also warns when one author-year matches TWO references —
on a 396-row corpus that happened five times (two Hölzel 2011s, two Kral 2022s, two
Yang 2025s, two Haudry 2025s, two Gusnard 2001s). Fix those with APA-7 §8.19:
name enough subsequent authors to distinguish them, (Kral, Davis, et al., 2022).
The other APA disambiguator, a 2025a/2025b year suffix, is accepted by the audit
gate but means editing the canonical apa strings, so it usually costs more.
The reference list comes from rows.json, so it is automatically canonical and complete; verify
by opening the .docx and confirming the figure renders and the section/citation structure reads
correctly. Worked example: distributed_conceptual_network/ (write_review.py + content.json
→ an AI-authored review with all 370 refs in APA-7). Output artifact:
<Topic>_review.docx (plus the project's content.json).
Tell the user:
- Total rows in spreadsheet, broken down (source / search / xref).
- Any verification corrections you made (e.g. fabricated PMCIDs, wrong first authors).
- Only if Phase 4 was run: PDFs downloaded vs. failed, and the path to
the browser-helper page or
_needs_manual.txtfor paywalled papers.
The workflow above is topic mode: it starts from a query and searches outward. Lab mode inverts the front end — it starts from a known body of work (a lab's publications), derives the lab's research themes and how they shifted over time, then searches outward to place that work in the field. Everything downstream (verify, count, families, figure) is the same machinery.
Phase L1 — ingest the corpus. tools/lab_corpus.py pulls the lab's full
publication list from OpenAlex by author id (use --search to find it; pass
several --author ids for PI + key lab members, or for one person whose record
is split across ids). Output lab_papers.json. --search prints each
candidate's ORCID and publication year span and warns on the failure shapes that
are visible from the listing alone — read those rather than picking by
institution, which is wrong more often than it is right (see L2).
Phase L1b — enrich abstracts (REQUIRED). OpenAlex metadata is not enough:
its abstracts are missing for a sizable minority of papers and its topics tags
are coarse, so classifying from them alone mislabels papers. Fill missing
abstracts from Semantic Scholar (/paper/batch, by DOI) and/or PubMed first.
Phase L2 — define the lab & verify the corpus (HUMAN CHECKPOINT #1). The
load-bearing gate: author-id disambiguation is the #1 correctness risk (OpenAlex
ids split / merge / collide; trainees move between labs). Have an agent classify
every paper from its actual content — not database topic tags — into the
buckets the user wants (e.g. for "human work only": human / primate /
other), and web-verify (PubMed / publisher) every paper without an abstract
and every ambiguous call. Prune false-positives; keep what the user asked for.
But do not let the inclusion filter discard the lab's foundational pre-paradigm
work (the macaque physiology, the pre-tool methods papers): those are the lab's
own antecedents and should re-enter as source=lab (starred) in the Phase-2b
pass even when the headline filter is, say, "human fMRI only." Flag them for the
user at this checkpoint rather than silently dropping them.
Resolving the author id — five real bootstraps, four distinct failure shapes. The institution label is the obvious discriminator and it was misleading in three of the five. Do not pick by it.
| Shape | What it looked like | What decided it |
|---|---|---|
| Wrong-university | The id labeled with the right university had 3 works; the correct one showed an unrelated institution and had 130 | Works count, then reading titles |
| Moved lab | Two candidates carried the university being searched for and neither was the person; one was a glaciologist. The correct id was still labeled with the PI's previous university | Distinct ORCIDs settle it instantly; otherwise the raw affiliation strings on the newest works |
| Merged | A single candidate, spanning 1976–2026, holding at least five different people | The year span — --search warns above 45 years |
| Split | Three candidates, all the same person, one holding a single high-profile paper | Same ORCID / same institutions / adjacent years; pass every id to --author |
Two rules follow, and they are easy to get backwards:
- A single candidate is the DANGEROUS case, not the safe one. When namesakes collide OpenAlex frequently merges them into one id rather than splitting them, so "only one match" can mean "all the contamination is in here".
- ORCID confirms authorship; it does not refute it. Use it to establish that a surprising paper is the PI's — one lab's 18 papers of psychiatric neuroimaging looked exactly like a collision and were her own pre-PhD work, and pruning on topic would have deleted a third of a real record. But profiles go stale: another PI's ORCID listed 18 works against OpenAlex's 39, so absence from ORCID is not evidence a paper belongs to someone else.
L2 is therefore a completeness check as well as a purity check. Every bootstrap before the split case only ever needed records removed, and nothing fails when a record is merely missing — so check explicitly for a split id, and say in writing how many records were added as well as pruned.
Keeping papers that are genuinely the PI's but not the lab's program. Early career work (a PhD in another field, a postdoc in another lab) is real authorship and L2 has no basis to prune it — but its vocabulary will skew any downstream step that reads titles. In one corpus such papers were 35% of the total and the keyword derivation duly proposed schizophrenia for a speech lab. Keep them, and record which groups are off-program in the notes so the next step knows to reject their vocabulary rather than rediscovering the problem.
Phase L3 — derive themes (HUMAN CHECKPOINT #2). Run the families step
(family_prompt_template.md → user approves the ~N themes → assign every kept
paper). The "families" are now the lab's research programs; tools/families.py
validates/stamps and emits families.json + families.md.
Phase L4 — render the lab's trajectory.
- Trajectory figure:
tools/families_figure.py— themes × year, the lab's papers as the spine (milestones labeled, the rest dots). This is "the lab's topics and how they changed over time." - Bibliography:
tools/spreadsheet.py(lab papers get thelabrow color).
Phase L4c — contextualize: a FULL topic-mode review, once per theme. This is NOT an optional or "lighter" pass. Placing the lab in its field means running the entire topic-mode workflow (Phases 2–6) for each theme — same rigor, same guardrails, no shortcuts. The recurring failure mode is treating this as a quick "context" add-on and skipping the machinery; that is exactly how a sloppy, half-fabricated field set sneaks into an otherwise careful review. For each theme:
- Search (Phase 2) with
search_prompt_template.md— precise theme definition, the lab's papers in that theme as the "already have"/exclude list, two-tier criteria (foundational vs recent), a capped target (~30–40), multiple query angles. One agent per theme. - Verify EVERY citation (Phase 3) with
tools/verify.py. It resolves arXiv/conference papers against the arXiv API (batched) — so a NOT-FOUND is a real failure to investigate, never "an arXiv paper, skip it." Re-run any ERROR verdicts (transient rate-limit/network — distinct from NOT-FOUND). Expect ~1 in 4 to need a fix (fabricated author lists, wrong arXiv ids, garbage DOIs). - Citation counts (Phase 5b) with
tools/citations.pyfor every field paper — bibliographies always carry counts. - Consolidate with two guarded steps (fail loud):
(a) cross-theme dedup — a paper found by several themes must be assigned to
exactly ONE (
families.pychecks ref-level exclusivity but NOT duplicate DOIs, so dedup by DOI yourself first); (b) exclude the reviewed lab's own DOIs — the agents only excluded each theme's seed list, so a lab paper from theme A can resurface as "field" in theme B. Assert zero field↔lab DOI collisions and zero cross-theme duplicate DOIs before merging. - Merge into the context corpus (lab rows
source=lab, field rowssource=search) and re-run families + figure (--emphasize-source labto keep the lab papers as the labeled spine over the field dots).
The three human checkpoints mirror topic mode: (1) the corpus (not a topic), (2) the themes, (3) the figure.
The lab's foundational pre-paradigm papers (gathered in L2 + Phase 2b) exist so a Phase-7 review can assign priority to the lab's own earlier work over later follow-ups from other groups — enforce that with the priority audit (contract rule 5; Phase 7).
- Always verify. ~25% of agent-returned citations have errors. Wrong first authors are the most common; the agent confuses similar-titled papers and mixes up author lists.
- The conclusion can be reversed. Read the abstract before trusting any summary — agents have been observed to invert a paper's headline finding (e.g., describing "X > Y" when the paper says the opposite).
- Cap the request. Asking for 40 papers gives 40-44; asking for "as many as possible" gives sprawl with more fabrications.
- Give exhaustive "do not include" lists. Without these, the agent re-finds papers already in the spreadsheet (3 of 44 in the first run).
- NOT-FOUND ≠ unverifiable ≠ fine. arXiv/conference papers used to slip
through because CrossRef/PubMed can't see them, so they returned NOT-FOUND and
got waved through.
verify.pynow hits the arXiv API directly; a NOT-FOUND is a real problem to chase, never a license to skip. (One run: a "Vo et al." attention paper was actually Foster et al.; two Jain & Huth arXiv ids pointed at unrelated papers — all caught only because every citation, preprint included, was verified.) - NOT-FOUND ≠ ERROR — a transient failure must never read as "missing." On a
~90-paper world_models run, 8 real preprints came back NOT-FOUND purely because
arXiv rate-limited (429) a per-paper loop into a temporary ban; trusting that
verdict would have dropped real papers.
verify.pynow (a) prefetches arXiv ids in batches (many perid_listcall, ~3s apart) so the ban doesn't happen, and (b) reports a distinct ERROR verdict when a lookup can't complete, kept separate from NOT-FOUND. Re-run ERRORs; only NOT-FOUND means "does not exist." Passexpect_yearas a string OR int — either is accepted (a JSON int no longer crashes the run), and one malformed row degrades to ERROR instead of aborting. - Reference titles are strict APA-7 sentence case; DOIs are the ground truth
for location, the APA string is just the bibliography display. CrossRef and
arXiv return titles in inconsistent casing (arXiv and many publishers use Title
Case; Nature deposits sentence case), so
references.pynormalizes ALL-CAPS titles but deliberately does NOT auto-transform Title Case → sentence case: correct sentence-casing needs the proper-noun judgment APA bakes in (Bayesian,Atari,Weber,Tolman-Eichenbaumstay capitalized;Active,World,Modellowercase), and a mechanical caser silently mis-cases proper nouns — which the audit can't catch. So after canon, sentence-case new preprint titles in a reviewed pass (auto-protect all-caps acronyms, camelCase/digit model names, and hyphen parts; keep a small proper-noun allowlist; eyeball every changed title, e.g. a product name likeMatrix-Game). This is a post-canon hand-fix like the mojibake and compound-surname fixes below.
- The outward search is a FULL topic-mode review, not a "context" add-on. Framing it as optional/lighter is precisely how a sloppy, half-fabricated field set sneaks into an otherwise careful review. Run Phases 2–6 per theme with the same verify/count/dedup guardrails — no shortcuts.
- Dedup by DOI and exclude the lab's own papers before merging. Multiple
theme-agents find the same landmark (one paper turned up under three themes),
and an agent only excludes its own theme's seeds — so a lab paper resurfaces as
"field."
families.pyenforces ref-level exclusivity but not duplicate DOIs; assert zero cross-theme dup DOIs and zero field↔lab collisions yourself.
- Order ideas by publication date, not by your narrative. The drafting model reliably presents a later paper as an idea's origin because that ref fits the sentence it wants to write. The rule, the four inversion patterns, the real misses caught 2026-06-13, and the required pre-delivery priority audit are all in Phase 7 — run the audit. Self-review misses these because the author is biased toward its own story; an independent pass with the publication years catches them.
Phase 4 has the full source order and the bot-blocked list. Three things that recur:
institutional-repo URLs (.edu/.ac.uk) from Unpaywall almost always work;
PMC/Cell/Elsevier/Wiley/OUP/MIT Press/PNAS/bioRxiv all block bots (route to the
helper page, don't retry); always validate the first 4 bytes are %PDF (a 200 can
be an HTML challenge page).
- Use a "source" column or color code so a future you (or the user) knows where each ref came from and how confident to be.
- Ref ids must be globally unique, and stay unique across merges. A
single-character lane prefix plus a multi-digit counter is ambiguous: lane
4with refs41…49then410…423parses the same as lane41, and a naive merge re-emitted the xref batch as410…on top of the existing410…, silently duplicating ids. Two failures followed:families.pyassigned the same paper twice, and — worse —citation_counts.jsonkeyed by ref got cross-contaminated (two rows sharing id415shared one count, so Carvalho 2024 inherited Elman 1990's 10,838 cites). After any merge, assertlen(refs)==len(set(refs)); attach citation counts only AFTER ids are final (or key the counts by DOI, not by ref); and prefer zero-padded or separator'd ids (4-01, or two-char lanes) so the counter can't collide with the prefix. - Keep summaries to 3-5 sentences. Long ones become unreadable in a row.
- Don't try to read existing xlsx with xlsxwriter — it's write-only. If appending, regenerate the whole file from a JSON of accumulated rows.
- Keep all rows in
rows.jsonand rebuild the xlsx viatools/spreadsheet.py. A project's row emitter (start fromtemplates/build_rows_template.py) runs ONCE, guards its writing block underif __name__ == "__main__":, imports the toolkit'scommon, and writes throughcommon.write_rows— which refuses to overwrite a canonical table, so re-running it later cannot wipe Phase 3f.
- When
--auditflagsU+FFFDmojibake, hand-fix it LAST. CrossRef stores a few names/venues with broken encoding (the original glyph is unrecoverable), soreferences.pypulls the broken character back in on every run. Fix it directly inrows.jsonAFTER the finalreferences.pypass and before building the spreadsheet/figure — fixing it earlier just gets it overwritten on the next canon, and you loop on the gate. Real cases:Bürgel,Zeitschrift für Anatomie. - CrossRef mis-splits compound / particle surnames (
Lambon Ralph→Ralph, M. A. L.;de Heer→Heer, W. A. D.). The shared formatter handles the common particles, but novel ones slip through and re-canon reintroduces the bad split — so correct these inrows.jsonafter the last canon, same as mojibake.--auditnow warns on every multi-word surname so they get looked at; expect a handful of legitimate ones (Spanish, Vietnamese and Italian double surnames) alongside the real errors. - What CrossRef deposits is sometimes simply wrong, not just mis-parsed. Seen in
one 396-row corpus:
family="A. Moffat"withgiven="Bradford"(a middle initial folded into the surname — now auto-repaired, since no surname starts with an initial);family="(Bud) Craig"for A. D. Craig (parenthetical nickname — now stripped); andSprbyfor Terje Sparby, a plain misspelling of a living author's name, verifiable only by finding the same author spelled correctly on a sibling paper. The first two are handled; the third can only be caught by reading. - Four formatter defects that shipped in five delivered bibliographies before the
gate caught them, all now hard defects: JATS markup left inside a title
(
<i>Generalization and Differentiation</i>), a?.where a question-mark title got an extra period, and U+2010/U+2011 Unicode hyphens in surnames (Fischer‐Baum,Kabat‐Zinn,low‐frequency) that look identical to ASCII but break every string match. Re-run--auditover old projects after a formatter change; these had been sitting in finished work for months. - The audit's
no-yearcheck used to reject APA year suffixes.\(\d{4}\)fails on(2025a), so the standard way to disambiguate two same-author/same-year works could not be expressed. It now accepts\(\d{4}[a-z]?\). rows.jsonis the live table; the row-emitter script is destructive once you pass Phase 3f. A per-projectbuild_data.py(or equivalent) only knows the original search rows — re-running it after canon/xref/families drops the xref rows and wipes the canonicalapa+ citation counts. After the first build, editrows.jsondirectly (or splice via targeted canon); don't regenerate it from the emitter.references.pyprefers the journal DOI over arXiv when a row has both (the version of record), and falls back to arXiv only for preprint-only rows (no journal DOI) or rows whose DOI is itself an arXiv DOI. So you no longer need to hand-clear anarxivfield to stop a published paper from being cited as its preprint — just keep both ids and let canon pick the journal version. (Agents routinely return a journal DOI and the preprint id for the same paper.)- BUT: not every "published" DOI is the version of record — check the count
before you swap. Curran/Proceedings.com registers DOIs (
10.52202/*) for the printed NeurIPS volumes. They resolve, and a CrossRef title search happily returns them, but they are shadow records of the paper the community actually cites: on a 2026 gallant_lab pass, moving 6 rows to their10.52202DOIs would have cut OpenAlex counts by ~3-4× (MindEye 40 → 11, Toneva-style rows 34 → 10) while adding nothing. ACL Anthology (10.18653/*), IEEE/CVF (10.1109/*) and real journal DOIs are the opposite — genuine upgrades (MindBridge 2 → 41). Rule: before promoting a preprint row to a "published" DOI, query the candidate DOI's OpenAlex count; if it is much LOWER than the arXiv record's, keep arXiv and note the venue in prose. Preprint-heavy CS/AI reviews are where this bites. - The same paper can enter a review twice, and per-row canon will never notice.
One agent finds the arXiv preprint, another finds the journal version; two
different DOIs, so the one-row-per-DOI rule passes and both rows canonicalize
perfectly. Three such pairs sat undetected in a 362-row corpus for six weeks.
references.py --auditnow runs a corpus-levelduplicate_scan()and prints⚠ A ~ B: possible duplicatefor near-identical titles (labeling the preprint-vs-published case explicitly). It is a warning, not a defect — real distinct papers do collide (a 2014 toolbox paper and its 2026 successor; successive years of the same challenge), so every pair needs a human verdict. Keep the version of record, drop the preprint, and re-check any in-text citation whose YEAR moves as a result (a preprint→journal promotion can shift 2025 → 2026).
- CrossRef coverage varies by publisher. Nature, Cell, OUP, JNeurosci have excellent coverage. Some smaller journals deposit no refs.
- ≥4 citations across 40 papers is a strong signal. ≥3 is borderline; only pick if the paper is clearly foundational.
- The xref pass typically finds 25-35 papers per topic that the initial search missed — almost half as many again.
- Google Scholar can't be automated. No API; CAPTCHA after ~10-20 requests. Don't try to scrape it for a whole bibliography — use OpenAlex + S2.
- OpenAlex is the reliable workhorse (~95%+ coverage by DOI). It undercounts arXiv-only preprints (separate record from the published version), so for preprint-heavy reviews lean on the S2 number for those rows.
- OpenAlex's BATCH filter can return a low-count duplicate record for a DOI
that has both a merged primary work and a stub (e.g. Whittington 2020 TEM came
back as 6 from the batch but 667 from the canonical
/works/doi:endpoint; Tolman 1948 as 1 vs 6656).citations.pynow (a) keeps the MAX count when a DOI appears more than once in a batch, and (b) re-queries the single-work endpoint whenever the OpenAlex count is < half the S2 count (with S2 ≥ 50). Still spot-check landmark/foundational counts against the S2 column before delivering — a famous old paper showing single-digit OpenAlex is the tell. - Semantic Scholar's free endpoints are flaky: the
/paper/batchendpoint 429s and sometimes 400s (one malformed id poisons the whole batch); the single/paper/{id}endpoint 404s valid papers under load. SetS2_API_KEYto fix it. Without a key, accept partial S2 coverage — OpenAlex stands alone. - Counts are a snapshot; record the
--asofdate. Don't expect the two columns to match — GS-style totals (which neither gives) run higher than both.
| Service | URL pattern | Returns |
|---|---|---|
| PubMed esearch | https://eutils.ncbi.nlm.nih.gov/entrez/eutils/esearch.fcgi?db=pubmed&term=<q>&retmode=json |
List of PMIDs |
| PubMed esummary | https://eutils.ncbi.nlm.nih.gov/entrez/eutils/esummary.fcgi?db=pubmed&id=<id>&retmode=json |
Title/authors/year |
| PMC esummary | ...?db=pmc&id=<numeric_pmc> |
Same, for PMC |
| Unpaywall | https://api.unpaywall.org/v2/<doi>?email=<email> |
OA PDF URLs |
| CrossRef metadata | https://api.crossref.org/works/<doi> |
Title, authors, references |
| OpenAlex (counts) | https://api.openalex.org/works?filter=doi:<d1>|<d2>...&mailto=<email> |
cited_by_count, batchable 50/req |
| Semantic Scholar (counts) | POST https://api.semanticscholar.org/graph/v1/paper/batch?fields=citationCount,influentialCitationCount body {"ids":["DOI:..","ARXIV:.."]} |
citation + influential counts; 429s without S2_API_KEY |
| EuropePMC search | https://www.ebi.ac.uk/europepmc/webservices/rest/search?query=<q>&format=json |
Full search |
| EuropePMC PDF | https://europepmc.org/articles/<PMCID>?pdf=render |
PDF (often) |
| arxiv API | http://export.arxiv.org/api/query?search_query=all:<q> |
Atom XML |
| arxiv PDF | https://arxiv.org/pdf/<id>.pdf |
|
| Nature direct | https://www.nature.com/articles/<id>.pdf |
PDF (if OA) |
Rate limits worth knowing:
- arxiv API: ~1 req/3s; bursts trigger 429.
- NCBI eutils: 3 req/s without API key, 10 req/s with key. Use 0.4s sleep.
- CrossRef: polite pool with
mailto:in User-Agent gives unlimited; without, ~50/s. - Unpaywall: 100k req/day per email.
Always include a User-Agent header with your email for these APIs.
All in <project>/tools/. Each is standalone, takes input via JSON/CLI,
outputs JSON/files. Run python3 tools/<script>.py --help for flags. The index
below is generated from the modules by python3 tools/gen_docs.py (CI fails if
it is stale); the per-tool detail is in tools/README.md and docs/tools.md.
| Script | Phase | Purpose | Flags |
|---|---|---|---|
verify.py |
3 | Verify a list of citations against PMC / PubMed / CrossRef / arXiv. | --citations --email --key --out --rows --sleep |
references.py |
3f | Canonical reference builder — make EVERY reference perfect, in both modes. | --asof --audit --email --key --out --repair --rows --sleep |
sentence_case.py |
3f | Post-canon pass — propose strict APA-7 sentence case for reference titles. | --apply --out --proper --rows --vocab |
download.py |
4 (opt-in) | Multi-source PDF downloader (Phase 4 — OPT-IN, not run by default). | --email --manual-list --out-dir --papers --sleep |
reconcile_downloads.py |
4 (opt-in) | Reconcile manually-downloaded PDFs against a slug+title+doi manifest. | --downloads-dir --dry-run --manifest --out-dir --since-hours |
spreadsheet.py |
5 | Build/rebuild the bibliography xlsx from a JSON of accumulated rows. | --out --rows --sheet-name |
citations.py |
5b | Fetch citation counts for a bibliography from OpenAlex + Semantic Scholar. | --asof --email --key --out --rows --sources |
xref.py |
6 | Build a cross-citation index from a list of papers. | --email --exclude --internal-out --key --min-cites --out --papers --resolve-unknown --rows --sleep |
families.py |
6b | Phase 6b — validate an LLM-proposed family taxonomy against the bibliography, stamp family onto rows.json, and emit families.json (the reproducible cache) + families.md (grouped tables + a family x topic cross-tab). |
--asof --assign --digest --md --out --rows |
families_figure.py |
6b | Phase 6b — render the interactive HTML lineage figure of the theoretical families. | --emphasize-source --families --internal --lab-author --max-labels --min-year --motif-min --no-auto-landmarks --no-raster --out-prefix --per-family --rows --spec --time-warp --title --xlsx |
cite_check.py |
7 | Phase 7 gate — every in-text citation must name a paper in rows.json. | --content --key --quiet --rows |
review_paper.py |
7 | Phase 7 — build a review ARTICLE (.docx) from a finished review corpus. | --content --figure --out --rows |
lab_corpus.py |
L1 | Lab mode — Phase L1: ingest a lab's full publication corpus from OpenAlex. | --author --email --from-year --out --search --to-year |
common.py |
— | Shared helpers for the literature-review toolkit. | — |
Notes the index cannot carry:
verify.pyandxref.pyaccept--rows rows.jsondirectly, so a project needs no converter script to feed them.references.py --repairretrofits an old corpus offline (no re-fetch, so post-canon hand fixes survive); canon and repair stamp rowscanonical_at.reconcile_downloads.pymatches by filename ↔ DOI substring first, then by first-author + year + title overlap on the first page (pdftotext), and refuses to move a PDF it is unsure about.- Project scripts should
import common(see tools/README.md, "Using the toolkit from a project script") and writerows.jsonthroughcommon.write_rows, which refuses to overwrite a canonical table.
Each helper is small and meant to be read + adapted. They are not a framework — they're scaffolding to keep the LLM judgment work fast.
1. Read this playbook.
2. Read the existing `rows.json` — after Phase 3f it is the live table and
the xlsx is only a rendering of it (xlsxwriter is write-only).
3. Confirm topic + criteria with the user. Do NOT ask whether to download
PDFs — the default is no (Phase 4 is opt-in only).
4. Phase 1: collect baseline (source-doc citations).
5. Phase 2: spawn search agent using tools/search_prompt_template.md.
6. Phase 3: verify EVERY citation (tools/verify.py).
6b. Phase 3f: canonicalize (tools/references.py), pass --audit, then
sentence-case titles in a reviewed pass (tools/sentence_case.py --proper).
7. Phase 5: update spreadsheet (tools/spreadsheet.py) with DOI URLs as Link.
8. Phase 5b: citation counts (tools/citations.py); attach to rows, rebuild.
9. Phase 6: cross-citation pass (tools/xref.py); verify and append xref
batch via Phases 3 + 5 again.
9b. Phase 6b (OPTIONAL): thematic families — propose → confirm with user →
assign → tools/families.py validates/stamps/renders; tools/families_figure.py
draws the figure and auto-labels landmarks (only arrows/notes are editorial).
9c. Phase 7 (OPTIONAL): review article — author prose into content.json,
gate it with tools/cite_check.py, then render with tools/review_paper.py
(APA-7 refs from rows.json). If AI-authored, state the AI author + a
verification disclosure. REQUIRED before delivery: run the priority audit
(origin claims must cite the EARLIEST paper, oldest-first) — see Phase 7.
10. Phase 8: report to user.
11. Phase 4 (PDF download) is OPTIONAL. Only run if the user explicitly
asks for PDFs.
12. If you changed any tool/phase/command, update the matching docs/ page
(see "Documentation site — keep it in sync" above).
For a topic with ~40 search-added + ~30 xref-added papers, the default no-PDF workflow takes Claude roughly 1-2M tokens and 5-10 minutes wall-clock — Phase 6 (xref) is the slowest step (~3-5 minutes for CrossRef calls). With Phase 4 turned on, add 10-20 more minutes for downloads.