Guidance for AI agents (and humans) working in this repository. CLAUDE.md points
here; this is the canonical file.
A pipeline that scrapes academic conference / journal programs and publishes a
static, searchable website (Netlify) for browsing papers. Venues are configured in
config/venues.yaml, scraped through pluggable adapters, enriched from bibliographic
metadata APIs, normalized to one schema, and shown in a single site with a category
sidebar. Enabled venues now span EDA, computer architecture, software engineering,
testing, programming languages, security/privacy, systems/networking, AI/ML, NLP, and
journals; new venues are added by config + adapter, never by changing the site.
Stack: a monorepo — a Python scraper (scraper/, package confer) that
emits unified JSON, consumed by an Astro static site (web/) that renders a single
page driven by a client-side script (scripts/app.ts) reading the embedded manifest.
The site builds to static assets deployed on Netlify.
Status legend used below: [now] = exists today · [target] = planned, not yet built. Keep this file honest — when you land something, move it from [target] to [now].
Three decoupled layers joined by one unified Paper schema:
config/venues.yaml ─▶ scraper + enrichers (Python) ─▶ unified JSON per venue ─▶ Astro site ─▶ Netlify
- Config names which scraper adapter a venue uses and passes only the source
locator the adapter cannot infer, such as
program_url,base_url, ortoc_url. - Adapters know one platform each and all emit the same
Papershape. - Enrichers merge DOI, abstracts, publication metadata, keywords, and open-access links from Crossref/OpenAlex by default without changing venue-specific adapters.
- Site consumes only the unified data — it never knows which platform data came from.
config/
venues.yaml [now] registry of venues to publish; read by config.py
scraper/ [now] Python project, package `confer`
pyproject.toml [now] console script: `confer`
src/confer/
cli.py [now] `build [--venue ID] [--refresh] [--limit N]`, `list`
config.py [now] load + validate ../config/venues.yaml (PyYAML)
models.py [now] unified Paper dataclass + schema
fetcher.py [now] HTTP + disk cache
paths.py [now] cache / output path helpers
pipeline.py [now] per-venue orchestration
enrichers.py [now] Crossref/OpenAlex metadata enrichment
export.py [now] write web/public/data/<venue>.json + venues.json
util.py [now] shared helpers
scrapers/
__init__.py [now] SCRAPERS registry
aaai.py [now] AAAI OJS archive / issue / article-detail adapter
acl_anthology.py [now] ACL Anthology event-page adapter
base.py [now] Scraper ABC
dateconf.py [now] DATE official programme adapter
dblp.py [now] DBLP bibliography TOC adapter
ieeesp.py [now] IEEE S&P accepted-papers page adapter (affiliations)
linklings.py [now] DAC (Linklings program) adapter
ndss.py [now] NDSS accepted-paper / detail-page adapter
openreview.py [now] OpenReview notes API adapter
researchr.py [now] Researchr program / timeline / accepted-list adapter
sigarch.py [now] SIGARCH-style static program adapter
sigchi.py [now] SIGCHI program platform (CHI/UIST/CSCW/…) JSON adapter
sosp.py [now] SOSP (SIGOPS) accepted-papers page adapter (affiliations)
... [target] ieee.py, acm_dl.py
tests/fixtures/ [now] small sample of cached pages for offline parse tests
web/ [now] Astro static site (Netlify)
package.json astro.config.mjs tsconfig.json
src/pages/index.astro [now] the whole single-page app shell (sidebar + content)
src/layouts/Layout.astro [now] html shell; sets theme/sidebar state before paint
src/lib/data.ts [now] read public/data at build (venues, generatedAt)
src/core/ [now] framework-agnostic query core (shared with mcp/)
text.ts [now] searchBlob, eventList, parseAff/authorAff/instList, normKey, authorResolver, tfidfTokenize, fieldText
query.ts [now] Term, FIELD_ALIASES, parseQuery, matchQuery(row, terms, QueryContext)
↳ QueryContext = { venueById, tagsOf? }. app.ts passes tagsOf at BOTH
call sites (matches + facetBase) so tag: queries work in the browser.
The MCP omits tagsOf by design — tag: no-ops (users have no tag store).
similar.ts [now] buildTfidfIndex(rows) → TfidfIndex { similar, recommend }
insights.ts [now] computeInsights(rows) → InsightsData, topN
*.test.ts [now] vitest unit tests (query/similar/insights); run: cd web && npm test
src/scripts/
app.ts [now] client island: search / filter / sort / favorites / export (imports from core/)
export.ts [now] BibTeX + CSV serialization (also used by mcp/)
types.ts [now] Paper / Venue / SavedSearch types (shared with mcp/)
src/styles/global.css [now] Claude-style CSS + sidebar/card layout
public/data/
venues.json [now] sidebar manifest (written by scraper export)
<venue_id>.json [now] unified papers per venue (written by scraper export)
mcp/ [now] MCP precompute artifacts (written by `confer build`)
vectors.json [now] per-paper TF-IDF vectors + slim records (find_similar)
stats.json [now] global top authors / institutions / tracks (top_*)
dist/ [now] Astro build output, gitignored (Netlify publishes it)
mcp/ [now] MCP stdio server (TypeScript, Node 22)
package.json [now] name: confer-mcp (publishable; bin → npx confer-mcp);
deps: @modelcontextprotocol/sdk, zod, undici
tsconfig.json [now] typecheck only (bundler resolution); tsup does the build
src/
corpus.ts [now] data loader: local dir (CONFER_DATA_DIR / repo checkout) OR
remote fetch from CONFER_DATA_URL (default confer.repus.me/data,
disk-cached, HTTPS_PROXY-aware). Manifest loads first so
venueById is sync; per-venue files load lazily + async.
loadVectors()/loadStats() read the precomputed artifacts.
server.ts [now] McpServer + StdioServerTransport; 8 tools (search, similar, stats, bibtex).
find_similar + unfiltered top_* use the precompute (with a
live-over-corpus fallback when artifacts are absent).
precompute.ts [now] build-time: reuses web/src/core to write public/data/mcp/*;
run by `confer build` (Node subprocess) and `npm run precompute`
README.md [now] npx quick-start + per-client setup + precompute + remote-not-planned
dist/ [now] tsup bundle (gitignored); `npm run build` → server.js + precompute.js
data/cache/ [now] cached raw HTML, gitignored. Per-venue subdirs.
netlify.toml [now] Netlify build config (base=web, publish=dist)
The site consumes this shape (see web/public/data/dac2026.json). Generalize, do
not reinvent it. Every adapter must produce records with these keys:
Sidebar manifest web/public/data/venues.json:
{ "generatedAt": "ISO-8601", "venues": [
{ "id": "dac2026", "name": "DAC 2026", "series": "DAC", "year": 2026,
"kind": "conference", "count": 543 }
]}class Scraper(ABC):
name: str
def __init__(self, venue: VenueConfig, fetcher: Fetcher): ...
def scrape(self) -> list[Paper]: ... # returns unified PapersAdapters are selected by venue.scraper via a registry
(SCRAPERS = {"linklings": LinklingsScraper, ...}). Adding a platform = one new
file in scrapers/ + one registry entry. Do not branch on platform anywhere else.
Prefer a venue-native adapter. When a venue publishes its own program on its
own site in its own format, implement a dedicated adapter for that source — even
when no existing adapter fits — rather than defaulting to a generic one such as
DBLP. A first-party source is usually richer (abstracts, author affiliations,
DOIs) and more authoritative than a bibliographic fallback, so the up-front cost
of a new adapter pays off in data quality and fewer enrichment gaps. Reach for an
existing adapter only when the venue has no usable first-party source of its own.
Known venues still missing author institutions or abstracts — with root cause and
fix direction — are tracked in docs/metadata-gaps.md.
- Add an entry to
config/venues.yaml(see the seed file for fields). - Ensure its
scraper:matches a registered adapter — writing a new one first if the venue's own format warrants it (see the principle above). - Provide only the adapter's required source locator. Tracks, event types, default labels, and normal metadata enrichers are inferred by the adapter and pipeline.
- Run
confer build --venue <id>and checkweb/public/data/<id>.json.
- Create
scraper/src/confer/scrapers/<platform>.pyimplementingScraper. - Register it in
scrapers/__init__.pySCRAPERS. - Add a fixture (a cached page) under
scraper/tests/fixtures/and a parse test. - All output must already be normalized to the
Paperschema — normalization lives in the adapter, not the site.
# Scraper (run inside scraper/)
uv run confer list # show configured venues
uv run confer build # all enabled venues → web/public/data/
uv run confer build --venue dac2026 # a single venue
uv run confer build --venue date2026 # DATE official programme venue
uv run confer build --venue asplos2026 # SIGARCH-style static programme venue
uv run confer build --venue hpca2026 # Researchr accepted-list venue
uv run confer build --venue icse2026 # ICSE 2026 via Researchr
uv run confer build --venue fse2026 # Researchr detailed timeline venue
uv run confer build --venue tosem2026 # DBLP journal TOC + default metadata enrichment
uv run confer build --venue popl2026 # Researchr + default metadata enrichment
uv run confer build --refresh # ignore cache, refetch over the network
uv run confer build --venue dac2026 --limit 5 # debug: only a few detail pages (skips precompute)
uv run confer build --no-precompute # skip regenerating web/public/data/mcp/*
uv run --extra dev pytest # offline parser tests (tests/fixtures/)
# Astro site (run inside web/)
npm install
npm run dev # local dev server
npm run build # static build → web/dist/ (what Netlify publishes)
npm test # vitest unit tests for web/src/core
# MCP server (run inside mcp/) — needed once so `confer build` can precompute
npm install
npm run build # → dist/server.js (+ dist/precompute.js)
npm run precompute # regenerate web/public/data/mcp/* (also run by confer build)
confer buildregenerates the MCP precompute artifacts (web/public/data/mcp/ vectors.json+stats.json) as its final step by invokingmcp/dist/precompute.js— so buildmcp/once first. The artifacts are committed and served like the rest of the data.
find_similar and the unfiltered top_* tools would otherwise need the whole corpus
(~84 MB) plus a TF-IDF build in every server process. mcp/src/precompute.ts reuses
web/src/core (buildTfidfModel, computeInsights) to emit two static artifacts at
build time:
vectors.json— per-paper TF-IDF vectors (+ slim records).find_similarloads it once and scores a single query against it — exact cosine (same math as the site), no corpus download. ~47 MB raw, ~15 MB gzipped over the wire.stats.json— global top authors / institutions / tracks for the unfilteredtop_*.
The server falls back to live computation over the corpus when the artifacts are absent
(e.g. an older deploy). Venue/query-scoped top_* and search_papers never use them.
ChatGPT's web connectors need a public HTTPS MCP URL; everything else runs npx confer-mcp (stdio, zero-install). A first-party hosted endpoint is deliberately out
of scope because: (1) the audience is narrow — only ChatGPT web needs a URL, and the
stdio→HTTP bridge in mcp/README.md covers it; (2) it would turn a free, stateless
static deploy into an operated service needing uptime, rate-limiting, and abuse/cost
control, since corpus-wide queries are comparatively expensive; (3) low marginal value —
the precompute already makes the heavy tools cheap. If that changes, the same core/ +
precompute could back a thin stateless Streamable-HTTP function later.
There are two independently-versioned artifacts — don't conflate them:
- confer (the site/repo) — versioned by
web/package.json, git tagsvX.Y.Z, and GitHub Releases. The site footer readsweb/package.jsonversion and links to/releases/tag/v<version>. - confer-mcp (the npm package in
mcp/) — versioned bymcp/package.jsonand published to npm (npm publish, needs the maintainer's 2FA OTP). Bump + publish it only whenmcp/source changes; it is not tied to confer site versions.
The version number, the git tag, and the GitHub Release are interdependent: the footer
badge links to /releases/tag/v<version>, so a version bump without a matching tag
and Release yields a dead link. Never do them piecemeal.
- CHANGELOG: rename
## [Unreleased]→## [X.Y.Z] - YYYY-MM-DDand leave a fresh empty## [Unreleased]above it. Entries are user-facing (impl details → commits). - Version: set
web/package.json"version"toX.Y.Z(semver: new feature → minor). - Commit:
chore(release): confer vX.Y.Z(CHANGELOG + web/package.json together). - Push:
git push origin main. - Tag:
git tag vX.Y.Z && git push origin vX.Y.Z. - GitHub Release:
gh release create vX.Y.Z --title "vX.Y.Z" --notes "<this version's CHANGELOG section>".
After step 6, confirm the footer link resolves (the Release for v<web/package.json version> must exist). Netlify redeploys from main on push and serves the new footer.
- Determinism: sort outputs (by id, then session) and serialize JSON with
ensure_ascii=False, indent=2, sort_keys=True. Stable diffs matter — the data is committed and served. - Pipeline must be a pure function of source + APIs — no backfill.
pipeline.pymust not read previously written output JSON to compensate for enrichment misses (DOI/abstract backfill). A build that depends on its own historical products is non-reproducible, masks real regressions, and produces subtly stale data. Root-cause any enrichment miss at the right layer: transient 429/5xx → add retry/backoff toFetcher; structural miss → fix the adapter or enricher match logic. The committedFetcheralready has exponential backoff withRetry-Aftersupport; lean on it. If a cold build cannot achieve the same quality as one that reads old JSON, that is a bug in the scraper/enricher, not a reason to add backfill. - Caching: never refetch when a cache file exists unless
--refresh. Be polite: reuse the sharedFetcher, keep a saneUser-Agent, support--delay. data/cache/is gitignored raw HTML — do not commit it. It is regenerable and kept locally as the offline parser test corpus; copy a small sample intotests/fixtures/for committed tests.- One schema: adapters may use the optional publication metadata keys above. If a
source has data that still does not fit the
Paperkeys, put it inextra; do not add untyped top-level keys the site does not read. - Site aesthetic: keep the minimal, Claude-like style (warm paper bg
#faf9f5, clay accent#c96442, Source Serif 4 headings, flat thin-bordered cards). The tokens live inweb/src/styles/global.css. - No left-edge accent bars on text/content boxes (e.g.
border-left: … solid var(--accent)on note previews, blockquotes, or read-only text areas). They look unfinished. Use a background tint (var(--panel-soft)) or spacing instead. The status-stripe on.mini-card--toread/reading/doneis a deliberate state indicator on a card row and is exempt. - Fade fixed/scroll seams. Any layout that pins a header, search bar, or
toolbar over a scrolling area (modal bodies, popovers, sidebar facets, settings)
must soften the boundary with a short (~10 px) gradient fade on the scroll
container so rows don't clip hard against the pinned element. Use
mask-image/-webkit-mask-imagewith a linear gradient fromtransparentto#000at both top and bottom edges. Never let scrolling content collide edge-to-edge with a sticky title or search input. - Favorites are client-side
localStorage. When multi-venue lands, key them asvenueId:paperIdso they do not collide across venues. - No backend: the site is fully static. All filtering/search is client-side.
- Changelog: every bug fix, feature, or notable behaviour change must add a
concise, user-facing one-liner to
CHANGELOG.mdunder## [Unreleased]in the appropriate### Added / Changed / Fixedsub-section. Keep entries terse — describe the visible effect, not the implementation. When tagging a release, rename[Unreleased]to a version + date section (e.g.## [1.2.0] - 2026-06-15).
Done:
- Scaffolding — AGENTS.md, CLAUDE.md, seed
config/venues.yaml, untracked cache. - Scraper monorepo —
src/dac26→scraper/src/confer; renamed package + console script (confer); split intofetcher/models/config/util/paths/pipeline/export/cliscrapers/{base,linklings}; readsconfig/venues.yamlvia PyYAML;confer build --venue dac2026reproduces 543 papers byte-for-byte; offline pytest fromtests/fixtures/.
- Export + Astro site —
export.pywritesweb/public/data/{venues.json,<venue>.json}; the Astro app reads it at build, ported Claude-style CSS. Deploys on Netlify (netlify.toml, buildsweb/, publishesdist/); the olddocs/site was removed and build artifacts are not committed. - Single-page site — one
index.astroshell driven by a client island (scripts/app.ts) reading the embedded manifest; category sidebar toggles venues; client-side search / track + event filters / sort / per-venue favorites (venueId:paperId) / BibTeX + CSV export / saved searches. - UX polish — collapsible sidebar & filter groups (animated), responsive mobile layout
- drawer, sidebar footer (last update, repo + commit links).
- Researchr venues —
researchradapter parses Researchr program tables, detailed timelines, and accepted-paper track pages; infers paper-bearing tracks and event types from the linked page; fetches cached detail modals for abstracts and official detail URLs; can merge several track pages into one venue viasource.program_urls(used by ASE 2026, whose full program page is not public yet); and publishes ICSE 2023/2024/2025/2026, FSE 2023/2024/2025/2026, ASE 2023/2024/2025/2026, ISSTA 2023/2024/2025, OOPSLA 2025/2026, HPCA 2026, POPL 2025/2026, and PLDI 2025/2026. - DATE 2025/2026 —
dateconfadapter parses the DATE official detailed programme, keeps downloadable paper rows, normalizes session metadata / author affiliations / PDF links, and publishes DATE 2025 and DATE 2026 alongside DAC 2025/2026. - Computer architecture venues —
sigarchadapter parses SIGARCH-style static program pages with session metadata and author affiliations, including malformed nested institution strings seen in live pages; publishes ASPLOS 2025/2026, ISCA 2026, and MICRO 2025. HPCA 2026 is published through the Researchr adapter. - Publication metadata enrichment —
enrichers.pymerges Crossref/OpenAlex data after the primary scrape by default, filling DOI, abstracts, publication date, publisher, container, volume/issue/pages, keywords, PDF/open-access links, and source provenance. Matching uses a broader conference-year window and can replace visibly truncated titles with richer bibliographic titles. - DBLP journals and proceedings —
dblpadapter publishes TOSEM 2025/2026 and TSE 2025/2026 journal articles, plus security/privacy and systems/networking conference proceedings from DBLP TOC XML. It follows linked USENIX presentation pages to fill abstracts, author-affiliation display text, pages, publisher, and PDF links. Researchr publishes POPL 2025/2026 and PLDI 2025/2026. - Official AI/NLP/security sources —
openreviewpublishes ICLR 2026, ICML 2025, and NeurIPS 2025 from accepted OpenReview notes;aaaipublishes AAAI 2026 from the official OJS archive;acl_anthologypublishes ACL 2025 from ACL Anthology event pages; andndsspublishes NDSS 2026 from the official accepted-paper pages.
Planned:
- More adapters — IEEE and ACM DL for sources that are not exposed through official
program pages, DBLP, or OpenReview. Adding a platform stays "one file in
scrapers/+ one registry entry".
- Scraper: Python
>=3.10,uv. Deps:beautifulsoup4,requests, andPyYAML(config is YAML, decided). - Site: Astro (Node ≥ 18). Static output (
astro build) deployed on Netlify; a single pre-rendered page with interactivity in a client island (scripts/app.ts). - Tests:
pytestinscraper/, driven by cached HTML intests/fixtures/so they run offline.
- iOS Safari ghost layer (unsolved). Toggling the theme, opening/closing the
mobile sidebar drawer, or opening a modal can leave a stale color/shadow band at
the bottom of the screen until a reload. It's a Safari compositor bug: it fails
to repaint fixed full-screen layers / the safe-area canvas when colors change or a
layer is created/destroyed. Things tried that did not fully fix it:
viewport-fit=coverenv()safe-area insets, atheme-colormeta, JS repaint nudges (background /color-scheme/backdrop-filteroff-for-a-frame), a real fixed.app-bgelement, and switching scrim/modals fromdisplaytoggling toopacity/visibility. We mirrored the Astro docs setup (noviewport-fit=cover;color-schemekeyed bydata-theme) which is the cleanest baseline, but the band can still appear. Treat as a known Safari limitation; revisit if Safari fixes the underlying repaint behavior.
{ "id": "RESEARCH123", // unique within a venue "title": "...", "abstract": "...", "authors": ["Jane Doe"], // array of names "authorInstitutions": "...", // display string "tracks": ["EDA", "Security"], "eventType": "Research Manuscript", "sessionTitles": ["..."], "sessions": ["sess123"], "dates": ["..."], "locations": ["..."], "urls": ["https://..."], "doi": "10.xxxx/...", // optional publication metadata "publicationDate": "2026-01-01", "publisher": "ACM", "container": "Proceedings / journal name", "volume": "35", "issue": "1", "pages": "1-20", "pdfUrls": ["https://..."], "artifactUrls": ["https://..."], "keywords": ["Software engineering"], "extra": { } // optional adapter-specific passthrough; never required }