cv-datasheets is the data supplier for dataset documentation in Common Voice. It maintains Jinja2 templates, community-written content, and static metadata that are compiled into a JSON file consumed by the bundler pipeline at release time.
cv-datasheets (compile-time) Bundler (runtime)
───────────────────────────────── ─────────────────────────
API snapshot ─────┐
Jinja2 templates ─┤
content/ files ───┤── compile ──> datasheets.json ──> datasheetsFetcher.ts
metadata/ files ──┘ │
DatasheetLocalePayload
{ template, community_fields, metadata }
│
datasheets.ts fills {{KEY}} with
live stats + community data
│
README.md -> tar.gz + GCS /datasheets/cv-datasheets/
├── templates/ Jinja2 templates
│ ├── base.md.j2 Shared skeleton
│ ├── scripted.md.j2 SCS child template
│ ├── spontaneous.md.j2 SPS child template
│ ├── i18n/ Section title translations (auto-discovered)
│ └── _legacy/ Old plain markdown templates (reference)
│
├── content/ Community content
│ ├── _field_map.json Field-to-bundler-key mapping
│ ├── _defaults/ Fallback content per template language
│ ├── _template/ Empty file structure for contributors
│ ├── _example/ Filled-in example (Klingon)
│ └── locales/{code}/ Per-locale content (shared/, scripted/, spontaneous/)
│
├── metadata/ Static data files
│ ├── api-snapshots/ Timestamped API snapshots (language names, variants, accents)
│ ├── locale-extras.json Locales not in API (el-CY, ms-MY)
│ ├── template-languages.json Non-"en" template overrides
│ └── funding.tsv OMSF-funded locales
│
├── scripts/ Utilities
│ ├── fetch_api_metadata.py Fetch SCS + SPS API -> snapshot
│ ├── preview_datasheets.py Preview datasheets with dummy stats
│ └── extract_community_data.py One-time extraction from legacy datasheets
│
├── compile_datasheets.py Main compile script
├── schema/ JSON Schema for output validation
├── docs/ Documentation
│ ├── ARCHITECTURE.md System design and bundler integration
│ ├── CONTRIBUTING.md Community contribution guide
│ └── COMPILING.md Compile script usage and release workflow
│
├── .github/workflows/ CI/CD
│ ├── preview.yml PR preview (auto-generates preview comment)
│ └── compile-latest.yml Auto-compile datasheets-latest.json on merge
│
├── releases/ Compiled output (datasheets-{snapshot_date}.json)
├── previews/ Local preview output (gitignored)
│
└── _legacy/ Deprecated scripts, metadata, and generated datasheetsJinja2 runs at compile time only inside this repository. It resolves template inheritance (base.md.j2 -> child templates) and produces flat markdown strings with {{KEY}} markers. The bundler never touches Jinja2.
How it works:
base.md.j2defines the shared skeleton (header, demographics, links, licence)scripted.md.j2andspontaneous.md.j2extend the base with modality-specific blocks- The compile script renders each child template with placeholder values from three namespaces:
stats.*-- title/header fields ({{ stats.native_name }}->{{NATIVE_NAME}})auto.*-- bundler-generated tables and blocks ({{ auto.gender_table }}->{{GENDER_TABLE}})community.*-- community content fields ({{ community.description }}->{{LANGUAGE_DESCRIPTION}})i18n.*-- literal i18n text with optional{variable}->{{KEY}}conversion{% block %}/{% extends %}are resolved (inheritance flattened)
- Result: flat markdown with
{{KEY}}markers -- same format the bundler already consumes
Some i18n strings contain {variable} placeholders that reference bundler statistics (sentence counts, clip counts, durations). At compile time, render_inline_vars() converts these to {{KEY}} bundler markers.
Example: "The dataset contains {validated_clips} validated clips" becomes "The dataset contains {{VALIDATED_CLIPS}} validated clips" in the compiled template.
The mapping is defined in INLINE_VAR_MAP in compile_datasheets.py. All i18n string values are processed through this conversion -- not just header intros.
Available inline variables:
| Variable | Bundler key | Description |
|---|---|---|
{version} |
VERSION |
Release version |
{english_name} |
ENGLISH_NAME |
Language name in English |
{locale} |
LOCALE |
Locale code |
{clips} |
CLIPS |
Total clips |
{hours_recorded} |
HOURS_RECORDED |
Hours of recorded speech |
{hours_validated} |
HOURS_VALIDATED |
Hours of validated speech |
{speakers} |
SPEAKERS |
Number of speakers |
{total_sentences} |
TOTAL_SENTENCES |
Total sentences in corpus |
{avg_duration_secs} |
AVG_DURATION_SECS |
Average clip duration |
{validated_clips} |
VALIDATED_CLIPS |
Validated clip count |
{invalidated_clips} |
INVALIDATED_CLIPS |
Invalidated clip count |
{other_clips} |
OTHER_CLIPS |
Unresolved clip count |
{validated_sentences} |
VALIDATED_SENTENCES |
Validated sentence count |
{unvalidated_sentences} |
UNVALIDATED_SENTENCES |
Unvalidated sentence count |
{rejected_sentences} |
REJECTED_SENTENCES |
Rejected sentence count |
{pending_sentences} |
PENDING_SENTENCES |
Pending review sentence count |
{reported_sentences} |
REPORTED_SENTENCES |
Reported sentence count |
i18n strings that use these variables: header_intro_scs, header_intro_sps, data_splits_detail, text_corpus_detail.
Language metadata comes from a snapshot of the Common Voice APIs, fetched by scripts/fetch_api_metadata.py:
| Source | Endpoint | Data |
|---|---|---|
| SCS languagedata | /api/v1/languagedata |
Names, text direction, variants, predefined accents, contributable status |
| SPS locales | /spontaneous-speech/beta/api/v1/locales |
Contributable SPS locale codes |
The snapshot is stored in metadata/api-snapshots/languagedata-{YYYYMMDD}.json and passed to the compile script via --api-snapshot.
Locale extras: Regional codes not present in the API (e.g. el-CY, ms-MY) are defined in metadata/locale-extras.json and merged into the snapshot at load time.
Template languages: Auto-discovered from templates/i18n/*.json -- the compile script checks for header_intro_scs/header_intro_sps keys to determine modality support. Default locale mapping is en; non-English overrides are in metadata/template-languages.json (currently 13 locales using es).
The compile script auto-generates content in two ways:
Compile-time auto-injection (community_fields fallback):
| Field | Source | Condition |
|---|---|---|
funding_description |
OMSF funding text | Locale in funding.tsv, no funding.md in content |
Community content always takes precedence over auto-injected defaults.
Template-embedded defaults (always present):
Three link fields have default links embedded directly in the compiled template via i18n keys. These defaults always appear alongside any community content:
| Field | Default links | i18n key |
|---|---|---|
community_links_list |
Pontoon + Communities page | community_links_defaults |
discussion_links_list |
Matrix, Discourse, Discord, Telegram | discussion_links_defaults |
contribute_links_list |
Speak/Write/Listen/Review (SCS) | scs_contribute_links_defaults |
contribute_links_list |
Question/Transcription links (SPS) | sps_contribute_links_defaults |
Bundler-generated mergeable data (runtime):
For mergeable fields (variants, accents, corpus, sources, text domains, transcriptions), the bundler generates statistical data at runtime. Community content is optional -- if provided, it appears before the bundler's auto-generated data. If not provided, only the bundler's data appears.
When the compile script loads community content for a locale, it checks four locations in order:
1. content/locales/{code}/{modality}/{field}.md <- modality-specific (highest)
2. content/locales/{code}/shared/{field}.md <- explicit shared (opt-in)
3. content/_defaults/{template_lang}/{field}.md <- template language default
4. content/_defaults/en/{field}.md <- English default
5. "" <- empty (bundler handles)The compile script produces a single JSON file per release:
{
"schema_version": "2.0.0",
"generated_at": "2026-02-26T...",
"snapshot_date": "2025-12-05",
"templates": {
"scs": { "en": "flat markdown...", "es": "...", "zh-TW": "..." },
"sps": { "en": "flat markdown...", "es": "..." }
},
"locales": {
"scs": {
"{locale}": {
"template_language": "en",
"metadata": {
"native_name": "...", "english_name": "...",
"text_direction": "LTR", "funding": "..."
},
"community_fields": {
"language_description": "...",
"variant_description": "...",
"accents_description": "...",
...
}
}
},
"sps": { "..." }
}
}The SCS bundler fetches the JSON by release name:
https://raw.githubusercontent.com/common-voice/cv-datasheets/main/releases/datasheets-{releaseName}.jsonIt then:
- Picks the correct template using
entry.template_language - Builds a replacement map from auto-generated stats + community fields + metadata
- Replaces
{{KEY}}placeholders via regex - Writes
README.mdper locale into the dataset tar
Two GitHub Actions workflows automate the pipeline:
Preview (preview.yml): Triggers on every PR that touches content/, templates/, or metadata/. Runs preview_datasheets.py --changed to compile affected locales from the working tree with dummy statistics, then posts the rendered preview as a comment on the PR. Also detects template-language remappings in metadata/template-languages.json.
Auto-compile (compile-latest.yml): Triggers on push to main (after merge). Runs compile_datasheets.py with the latest committed API snapshot and commits releases/datasheets-latest.json if it changed. The commit message (chore: auto-compile) prevents re-triggering.
| Content type | Handled by | Bundler sees |
|---|---|---|
| Community-written content | compile_datasheets.py | Filled community_fields values |
| API-derived names & direction | compile_datasheets.py (API snapshot) | metadata.native_name, metadata.english_name, metadata.text_direction |
| Link defaults (Pontoon, Discourse, Speak/Write, etc.) | i18n keys -> template-embedded | Literal text in compiled template (always present) |
| OMSF funding | compile_datasheets.py (metadata/funding.tsv) |
Pre-filled funding_description |
| Auto-generated stats (clips, hours, demographics) | Bundler at runtime | {{KEY}} placeholders in template |
| Inline stats (sentence counts, split counts, durations) | Bundler at runtime | {{KEY}} placeholders inside i18n text |
| Sentence/question samples | Bundler at runtime | {{KEY}} placeholders in template |
| Data splits table (SCS only) | Bundler at runtime | {{DATA_SPLITS_TABLE}} placeholder in template |
| Mergeable field stats (corpus, sources, domains, variants, accents, transcriptions) | Bundler at runtime | {{KEY}} placeholders adjacent to community fields |