Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
105 commits
Select commit Hold shift + click to select a range
e924aca
Increase limits for LLM output tokens [skip ci]
mckeea Jun 3, 2026
74713aa
dont edit date field on public release [skip ci]
mckeea Jun 3, 2026
ee7da4b
Update CLC+Core_User_Manual_Issue_4.0_v1.qmd
EEA-Tommy-Andersen Jun 3, 2026
b20dd4a
Render index.qmd serially (-P1) to avoid shared-file race (xargs 123)
mckeea Jun 4, 2026
87918bc
Update CLC+Core_User_Manual_Issue_4.0_v1.qmd
EEA-Tommy-Andersen Jun 8, 2026
45b3426
pdf2qmd tool implementation [skip ci]
mckeea Jun 26, 2026
25bcc01
feat: add 15 newly converted products documents
MatMatt Jun 27, 2026
a0e15ba
Merge pull request #35 from eea/feat-new-docs
MatMatt Jun 27, 2026
596be08
Update LLM cache and version history [skip ci]
actions-user Jun 27, 2026
4013bde
fix: set categories to products and clean double nested quotes in titles
MatMatt Jun 28, 2026
ac3cd43
fix: simplify changelog list parsing filter to prevent panflute versi…
MatMatt Jun 28, 2026
58d1096
fix: build container runner with panflute>=2.3.1 for Pandoc3 Typst AS…
MatMatt Jun 28, 2026
92c4f59
Merge pull request #36 from eea/fix-panflute-container
MatMatt Jun 28, 2026
97a8744
fix: force update container panflute library from inline action step
MatMatt Jun 28, 2026
51eee96
fix: globally monkeypatch panflute RAW_FORMATS to support typst in as…
MatMatt Jun 28, 2026
a6eb02e
fix: add ruff noqa E402 to ignore top level monkeypatch imports
MatMatt Jun 28, 2026
4641de3
fix: resolve hardcoded FIG_3 token reference in 2018 Technical Guidel…
MatMatt Jun 28, 2026
7b4a41d
fix: temporarily move Ice products ATBD out of build to bypass Jupyte…
MatMatt Jun 28, 2026
bc8bd06
fix: correct Ice products grouped filename with products_ prefix in b…
MatMatt Jun 28, 2026
88c68e2
fix: correct relative origin path inside DOCS directory in build move
MatMatt Jun 28, 2026
224fb46
fix: sanitize stray at-sign handle to prevent bibtex bibliography err…
MatMatt Jun 28, 2026
17eb5ee
fix: escape at-sign handle with a backslash to prevent bibtex citatio…
MatMatt Jun 28, 2026
4c4cab6
feat: update metadata validation parameters and assign category produ…
MatMatt Jun 29, 2026
b444cac
chore: safely remove JRC files and clean up products staging folder
MatMatt Jun 29, 2026
cd324d1
fix: restore the official 1.0.0 audited small landscape features manu…
MatMatt Jun 29, 2026
c685628
Update LLM cache and version history [skip ci]
actions-user Jun 29, 2026
c2e0bc8
style: rename CLC+ to CLCplus in clcplus-core user manuals and media …
MatMatt Jun 29, 2026
47a8d68
Update LLM cache and version history [skip ci]
actions-user Jun 29, 2026
04fce1b
chore: rename and structure vegetation layers report to hrl_vegetatio…
MatMatt Jun 29, 2026
35613ad
Update LLM cache and version history [skip ci]
actions-user Jun 29, 2026
b45de30
fix: correct unused footnote key 7 reference link in tree cover and f…
MatMatt Jun 29, 2026
0626705
chore: remove the Coastal Zones Quality Assessment Report D3.2
MatMatt Jun 29, 2026
98289ef
fix: format spatial coverage projcs coordinate table cell with a clea…
MatMatt Jun 29, 2026
c4fb96d
DOCS and building tools cleanup
mckeea Jul 1, 2026
ef12b97
versions cleanup for new documents [skip ci]
mckeea Jul 1, 2026
6eec0de
html table simplifier for .llms.md files
mckeea Jul 1, 2026
cda6756
misc tools fixes [skip ci]
mckeea Jul 1, 2026
597a315
feat: map reports category in build and packaging script
MatMatt Jun 29, 2026
420dec0
docs: modernize and modularize README for generic fork-ready usage
MatMatt Jul 1, 2026
eadb91f
feat: agent guidance, category single-source, and local preview
mckeea Jul 8, 2026
8ddb1c9
fix: single quote normalization in .qmd files [skip ci]
mckeea Jul 10, 2026
a660c56
CDSE Migration Status — weekly dashboard [non-browsable]
MatMatt Jul 14, 2026
3e562bc
Update LLM cache and version history [skip ci]
actions-user Jul 14, 2026
d4ff063
CDSE Migration Status — fix Typst font-weight for PDF build
MatMatt Jul 14, 2026
e6d4315
Update LLM cache and version history [skip ci]
actions-user Jul 14, 2026
32b7d0c
Update non-browsable doc map [skip ci]
actions-user Jul 14, 2026
4a64bd7
CDSE Migration Status — Hot Spots → Global, remove 0%/100% sections
MatMatt Jul 14, 2026
af36deb
Update LLM cache and version history [skip ci]
actions-user Jul 14, 2026
d0052f6
CDSE Migration Status — Component/Total headers, % in progress bar, n…
MatMatt Jul 14, 2026
136359e
docs: add CLMS → CDSE migration status dashboard
MatMatt Jul 16, 2026
e7f31a8
chore: weekly migration data update
MatMatt Jul 16, 2026
7365f6b
chore: remove data JSON — hosted externally on copernicus-land/clms-c…
MatMatt Jul 16, 2026
f240574
Merge pull request #43 from eea/feature/dashboard-update-2026-07-16
MatMatt Jul 16, 2026
46f73a2
Update LLM cache and version history [skip ci]
actions-user Jul 16, 2026
1f7205a
fix: date must be YYYY-MM-DD, not 'today' for CI validator
MatMatt Jul 16, 2026
81b2461
Update non-browsable doc map [skip ci]
actions-user Jul 16, 2026
0070357
Merge pull request #44 from eea/fix/date-format-dashboard
MatMatt Jul 16, 2026
2f8f8ef
fix: hide OJS code blocks (echo: false at document level)
MatMatt Jul 16, 2026
e883bda
fix: hide OJS code blocks in dashboard
MatMatt Jul 16, 2026
bf7b33a
fix: disable code-fold for OJS cells
MatMatt Jul 16, 2026
566437e
fix: disable code-fold for OJS cells
MatMatt Jul 16, 2026
628c49d
fix: hide code blocks via CSS (frontmatter gets stripped by pipeline)
MatMatt Jul 16, 2026
0f1de50
fix: hide code blocks via CSS
MatMatt Jul 16, 2026
e39cf78
fix: hide code blocks via JS (CSS gets stripped by pipeline)
MatMatt Jul 16, 2026
a1449df
fix: hide code blocks via JS
MatMatt Jul 16, 2026
8b04425
fix: hide code blocks via OJS (only content that survives pipeline)
MatMatt Jul 16, 2026
e9f2faa
fix: hide code blocks via OJS
MatMatt Jul 16, 2026
657c617
chore: remove old CLMS_CDSE_Migration_Status.qmd (replaced by Dashboard)
MatMatt Jul 17, 2026
083b8a0
chore: remove old status report
MatMatt Jul 17, 2026
1766f6e
update: CLMS→CDSE Migration Dashboard
MatMatt Jul 18, 2026
b380506
Merge pull request #54 from eea/feat/dashboard-update-2026-07-18
MatMatt Jul 18, 2026
e0e9c1e
Update LLM cache and version history [skip ci]
actions-user Jul 18, 2026
95ee932
feat: add type: dashboard to migration dashboard
MatMatt Jul 18, 2026
6e0e14b
Merge pull request #55 from eea/feat/dashboard-type
MatMatt Jul 18, 2026
6c91bd0
feat: update migration dashboard with daily scraper & embedded previo…
MatMatt Jul 19, 2026
10d2f9c
Merge pull request #56 from eea/feat/dashboard-update-2026-07-19
MatMatt Jul 19, 2026
c3499d8
CDSE Migration Status — weekly data update
MatMatt Jul 20, 2026
aa8b3c9
Update LLM cache and version history [skip ci]
actions-user Jul 20, 2026
8e5bb58
chore: remove stale CDSE_Migration files (status.qmd, DATA/)
MatMatt Jul 20, 2026
c1adc03
Revert "chore: remove stale CDSE_Migration files (status.qmd, DATA/)"
MatMatt Jul 20, 2026
9042b33
chore: remove stale CDSE_Migration files (status.qmd, DATA/)
MatMatt Jul 20, 2026
c5431bd
docs: fix grammar in migration status table sentence
MatMatt Jul 20, 2026
8ec3381
Merge pull request #57 from eea/docs/fix-migration-status-grammar
MatMatt Jul 20, 2026
b903070
feature: render configuration per doc-type added
mckeea Jul 22, 2026
71f4fc8
fix: for dashboard document, handle data fetching error
mckeea Jul 22, 2026
593fd08
fix: for contact field in document-type config
mckeea Jul 22, 2026
897dab8
docs: add MRVPP ATBD v2 document
MatMatt Jul 24, 2026
80ccf64
Merge pull request #58 from eea/docs/mrvpp-atbd-v2
MatMatt Jul 24, 2026
27e5036
feat: add CLMS filenaming guidelines (tree + design principles)
MatMatt Jul 24, 2026
811f1dc
Merge pull request #59 from eea/docs/clms-filenaming
MatMatt Jul 24, 2026
f5d4862
fix: rename RELEASE to REVISION in filenaming docs
MatMatt Jul 25, 2026
1ca50f2
Merge pull request #60 from eea/fix/clms-filenaming-revision
MatMatt Jul 25, 2026
4749286
docs: add MR-VPP Product User Manual v1
MatMatt Jul 27, 2026
1dfc478
Merge pull request #61 from eea/docs/mrvpp-pum
MatMatt Jul 27, 2026
1d72f4a
docs: tidy MR-VPP ATBD v2 formatting
MatMatt Jul 27, 2026
a0ce524
docs: add author to MR-VPP PUM v1
MatMatt Jul 27, 2026
f969c5b
Merge pull request #62 from eea/docs/mrvpp-pum
MatMatt Jul 27, 2026
8cf4fb6
fix: rename MR-VPP ATBD folder/file to underscore separator
MatMatt Jul 27, 2026
fd3daec
Merge pull request #63 from eea/fix/mrvpp-atbd-filename-separator
MatMatt Jul 27, 2026
a407fe4
docs: tidy MR-VPP table captions and code fence
MatMatt Jul 27, 2026
be3dff4
fix: crop baked-in captions from MR-VPP figures
MatMatt Jul 27, 2026
3d8bcb5
Merge pull request #64 from eea/fix/mrvpp-figure-captions
MatMatt Jul 27, 2026
75cb098
fix: drop duplicate Figure 8 image in MR-VPP PUM
MatMatt Jul 27, 2026
841a051
fix: separate MR-VPP PUM Figures 4 and 5 into own paragraphs
MatMatt Jul 27, 2026
51b6c3d
Merge pull request #65 from eea/fix/mrvpp-pum-figures
MatMatt Jul 27, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
30 changes: 30 additions & 0 deletions .Rprofile
Original file line number Diff line number Diff line change
@@ -0,0 +1,30 @@
# CLMS Technical Library — local PDF preview font setup.
#
# Makes the bundled Typst fonts (_meta/theme/typst-fonts/) discoverable so that
# rendering a DOCS/*.qmd to PDF — e.g. via the RStudio "Render" button — uses the
# CLMS fonts (Lato, Liberation Sans, JetBrains Mono) without installing anything
# system-wide. Quarto's typst format has no font-path option, so we set Typst's
# native TYPST_FONT_PATHS env var here; RStudio runs this file on session start,
# and the quarto render it launches inherits the variable.
#
# Preview only — the CI build (build-docs.sh) does not use this.
local({
font_dir <- normalizePath(
file.path(getwd(), "_meta", "theme", "typst-fonts"),
mustWork = FALSE
)
if (dir.exists(font_dir)) {
existing <- Sys.getenv("TYPST_FONT_PATHS")
paths <- if (nzchar(existing)) {
strsplit(existing, .Platform$path.sep, fixed = TRUE)[[1]]
} else {
character(0)
}
if (!(font_dir %in% paths)) {
Sys.setenv(
TYPST_FONT_PATHS = paste(c(font_dir, paths), collapse = .Platform$path.sep)
)
message("CLMS: TYPST_FONT_PATHS set to bundled fonts for local PDF preview.")
}
}
})
10 changes: 10 additions & 0 deletions .github/non_browsable_doc_map.json
Original file line number Diff line number Diff line change
Expand Up @@ -15,6 +15,16 @@
"source": "Copernicus_Land_Data_Store_CLDS/Survey_Evaluation_v1.qmd",
"base": "dr3yejtzxollhygq0lwfgjxf755jvzyly2lemmvpgiqrzf5own2074qpdivug54c",
"url": "/dr3yejtzxollhygq0lwfgjxf755jvzyly2lemmvpgiqrzf5own2074qpdivug54c.html"
},
{
"source": "CDSE_Migration/CLMS_CDSE_Migration_Status.qmd",
"base": "54ab2b07e82f82143780f24e2704ef211394c00216e83839e7c1dc42ba8e3a91",
"url": "/54ab2b07e82f82143780f24e2704ef211394c00216e83839e7c1dc42ba8e3a91.html"
},
{
"source": "CDSE_Migration/CLMS_CDSE_Migration_Dashboard.qmd",
"base": "a3e44009d6b7375f60a58b11278cbe079d93407249c0c5e1c97f00e12e98c09a",
"url": "/a3e44009d6b7375f60a58b11278cbe079d93407249c0c5e1c97f00e12e98c09a.html"
}
]
}
2 changes: 1 addition & 1 deletion .github/runners/Dockerfile.quarto-doc-builder
Original file line number Diff line number Diff line change
Expand Up @@ -55,7 +55,7 @@ RUN arch=$(dpkg --print-architecture) && \
# files.upload(file=), see gemini_client.py)
# tiktoken -> token counting for the LLM rate limiter
RUN pip install --no-cache-dir --break-system-packages \
panflute PyYAML google-genai==2.6.0 tiktoken==0.13.0 && \
panflute>=2.3.1 PyYAML google-genai==2.6.0 tiktoken==0.13.0 && \
python3 -m pip cache purge 2>/dev/null || true && \
rm -rf /root/.cache /tmp/* /var/tmp/*

Expand Down
11 changes: 6 additions & 5 deletions .github/scripts/ai/update_versions_and_changelogs.py
Original file line number Diff line number Diff line change
Expand Up @@ -816,9 +816,8 @@ def _do_call():
print("=" * 70)
print(result_text)
print("=" * 70)
raise Exception(
f"Batch {batch_num}/{total_batches} failed: Invalid JSON response"
)
# Empty -> batch_with_retry splits and retries the halves instead of aborting.
return {}
except Exception as e:
print(f"\n❌ ERROR: Batch {batch_num}/{total_batches} processing failed")
print(f" Error: {e}")
Expand Down Expand Up @@ -862,6 +861,8 @@ def _process(sub_batch, _i=i):

def calculate_new_version(current_version, bump_type, major_from_filename):
"""Calculate new version based on bump type"""
# First published version is 1.0.0, so a _v0 filename floors to major 1.
major_from_filename = max(major_from_filename, 1)
try:
parts = current_version.split(".")
major = int(parts[0])
Expand Down Expand Up @@ -947,7 +948,7 @@ def initialize_first_release(all_files):
print(f"[ERROR] {filepath}: {e}")
continue

initial_version = f"{major_version}.0.0"
initial_version = f"{max(major_version, 1)}.0.0"

update_qmd_version_only(filepath, initial_version)

Expand Down Expand Up @@ -1073,7 +1074,7 @@ def main():
file_info[filepath] = {
"major_version": major_version,
"current_version": versions_metadata.get(filepath, {}).get(
"current_version", f"{major_version}.0.0"
"current_version", f"{max(major_version, 1)}.0.0"
),
}

Expand Down
17 changes: 16 additions & 1 deletion .github/scripts/build/apply_cached_intros.py
Original file line number Diff line number Diff line change
Expand Up @@ -15,8 +15,9 @@
SCRIPTS_ROOT = SCRIPT_DIR.parent
sys.path.insert(0, str(SCRIPTS_ROOT))

from helpers.qmd_utils import find_qmd_files # noqa: E402
from helpers.qmd_utils import find_qmd_files, read_qmd_frontmatter # noqa: E402
from helpers.file_updater import apply_all_updates # noqa: E402
from helpers.doc_types import element_off # noqa: E402

ROOT_DIR = (SCRIPT_DIR / "../../..").resolve()
CACHE_DIR = (ROOT_DIR / ".llm_cache").resolve()
Expand Down Expand Up @@ -52,6 +53,20 @@ def main() -> int:
print(f"[apply_cached_intros] no .qmd files under {input_dir}")
return 0

# Skip types that take neither intro nor keywords (dashboard) so nothing is
# injected, not even into the origin_DOCS copy. They come as one bundle,
# hence the "both off" check — split it if a type ever wants only one.
def _wants_intro(qmd: Path) -> bool:
fm, _ = read_qmd_frontmatter(qmd)
dtype = fm.get("type")
return not (element_off(dtype, "keywords") and element_off(dtype, "description"))

kept = {q for q in qmd_files if _wants_intro(q)}
skipped = len(qmd_files) - len(kept)
if skipped:
print(f"[apply_cached_intros] skipping {skipped} file(s) whose type omits intros/keywords")
qmd_files = kept

stats = apply_all_updates(
qmd_files,
CACHE_DIR,
Expand Down
27 changes: 24 additions & 3 deletions .github/scripts/build/build-docs.sh
Original file line number Diff line number Diff line change
Expand Up @@ -31,6 +31,11 @@ step() {
_STEP_PREV=$now; _STEP_NAME="$*"
}

# Snapshot the pristine source before the build mutates it, so a local build is
# easy to undo. After testing, restore your working tree with:
# rm -rf DOCS origin_DOCS && mv source_DOCS DOCS
rm -rf source_DOCS && cp -rp DOCS source_DOCS

# Apply cached intros/keywords before the rename - the cache is keyed by original path.
echo "Injecting cached intros & keywords (no API)..."
python3 .github/scripts/build/apply_cached_intros.py DOCS
Expand Down Expand Up @@ -73,6 +78,13 @@ python3 ../.github/scripts/build/fill_version.py .
echo "Balancing table column widths..."
python3 ../.github/scripts/qmd-tools/fix_table_colwidths.py .

# Promote bare "Table N:"/"Figure N:" captions (left as plain text by the
# converters) into .tbl-caption divs / image alt text, so they render as styled
# captions instead of body text. Runs after the grid->pipe conversion above so a
# caption next to a (now pipe) table is recognised. Build-copy only; idempotent.
echo "Promoting bare table/figure captions..."
python3 ../.github/scripts/qmd-tools/promote_bare_captions.py .

# Bake image descriptions into the qmd source now, so the render doesn't re-hash
# every image once per format (see the script). This was a Lua filter.
echo "Baking image descriptions into qmd source..."
Expand All @@ -83,9 +95,12 @@ python3 ../.github/scripts/build/inject_image_descriptions.py .
cp _quarto-no-headers.yml _quarto.yml

step "[3/6] Rendering all documents (HTML + Typst PDF + gfm) in one pass..."
# Render every format in the config - html (site), typst (PDFs), gfm
# (the .llms.md sidecars). One pass over the files instead of one per format.
# # Temporary move out of docs before render to avoid Jupyter engine selections crashes
# echo " [BUILD BYPASS] Temporarily moving Ice products out of the build context..."
# mv products/products_Algorithm_theoretical_basis_document_-_High_Resolution_Ice_products_Europe.qmd ../origin_DOCS/
quarto_render --no-clean
# echo " [BUILD BYPASS] Restoring Ice products..."
# mv ../origin_DOCS/products_Algorithm_theoretical_basis_document_-_High_Resolution_Ice_products_Europe.qmd products/

# Back up sitemap.xml and llms.txt - the index.qmd renders below regenerate
# them, and we want to keep the values from this first render.
Expand All @@ -99,8 +114,10 @@ python3 ../.github/scripts/build/generate_index_all.py
step "[5/6] Rendering index.qmd files..."
mv _quarto.yml _quarto_not_used.yml
mv _quarto-index.yml _quarto.yml
# Serial (-P1): parallel renders raced on shared _site files (sitemap/search/
# listings) and intermittently failed the whole build with xargs exit 123.
find ./ -type f -name index.qmd -print0 | \
xargs -0 -P4 -I{} \
xargs -0 -P1 -I{} \
bash -c 'quarto render "$1" --profile index --to html --no-clean --quiet 2>&1 | grep -v -e "Unknown meta key .* specified in a metadata Shortcode" -e "^Output created:"; exit ${PIPESTATUS[0]}' _ {}
mv _quarto.yml _quarto-index.yml
cp _quarto_not_used.yml _quarto.yml && rm _quarto_not_used.yml
Expand All @@ -110,6 +127,10 @@ cp _site/sitemap.xml.bkp _site/sitemap.xml
rm -f _site/sitemap.xml.bkp
mv _site/llms.txt.bkp _site/llms.txt

# Drop .llms.md for llms-off types (dashboard), before the llm sitemap so it
# falls back to the HTML URL.
python3 ../.github/scripts/build/strip_llms_sidecars.py . _site

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Run the new post-render LLMS cleaner

The newly added clean_llms_md.py documents that Quarto injects residual <div> wrappers after Pandoc filters run, but the inspected build-docs.sh post-render sequence only strips disabled sidecars and never invokes that cleaner. Consequently every deployed .llms.md companion still contains the wrapper noise that the new script was introduced to remove before RAG ingestion; call it on _site before generating the LLM sitemap.

Useful? React with 👍 / 👎.


# Remove non-browsable links from sitemap.xml and llms.txt
python3 ../.github/scripts/build/remove_non_browsable.py _site/sitemap.xml
python3 ../.github/scripts/build/remove_non_browsable.py --format llms _site/llms.txt
Expand Down
65 changes: 65 additions & 0 deletions .github/scripts/build/clean_llms_md.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,65 @@
#!/usr/bin/env python3
"""Strip residual block-level HTML wrappers from generated .llms.md sidecars.

The gfm writer + the simplify_tables_gfm.lua filter already linearize tables and
remove styling, but Quarto injects crossref float wrappers (`<div id="tbl-…">` …
`</div>`) and the occasional layout `<div style="overflow-x:…">` AFTER the pandoc
filters run, so they can't be removed at filter level. These bare wrapper lines are
pure noise in the plain-text companion the RAG ingests. This post-render pass removes
standalone block-level `<div>`/`</div>` lines (inline tags like <sub>/<sup> in prose
are left intact).

Idempotent. Usage:
python3 clean_llms_md.py _site # walk *.llms.md under a directory
python3 clean_llms_md.py a.llms.md b.md # specific files
"""

import re
import sys
from pathlib import Path

# A line that is ONLY an opening/closing <div …> (optionally indented). Block-level
# wrapper noise — never matches inline tags mid-prose.
_BARE_DIV = re.compile(r"^[ \t]*</?div\b[^>]*>[ \t]*$")


def clean_text(text: str) -> str:
out = [ln for ln in text.split("\n") if not _BARE_DIV.match(ln)]
cleaned = "\n".join(out)
# collapse the blank-line runs a removed wrapper can leave behind (3+ → 2)
cleaned = re.sub(r"\n{3,}", "\n\n", cleaned)
return cleaned


def clean_file(path: Path) -> bool:
original = path.read_text(encoding="utf-8")
cleaned = clean_text(original)
if cleaned != original:
path.write_text(cleaned, encoding="utf-8")
return True
return False


def _iter_targets(args):
for a in args:
p = Path(a)
if p.is_dir():
yield from p.rglob("*.llms.md")
elif p.exists():
yield p


def main(argv) -> int:
if not argv:
print(__doc__)
return 1
changed = 0
for path in _iter_targets(argv):
if clean_file(path):
changed += 1
print(f"clean_llms_md: cleaned {changed} file(s)")
return 0


if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
54 changes: 41 additions & 13 deletions .github/scripts/build/fill_version.py
Original file line number Diff line number Diff line change
Expand Up @@ -6,17 +6,22 @@
version, so we fill it each build, like intros and changelogs.

The value is the tracked version from .llm_cache/versions.json, looked up via the
original-filename field. No record yet -> {major}.0.0 from the _vN.qmd name. No
original-filename field. The doc's own `version:` frontmatter is never trusted as
input - it is output we overwrite. No record yet -> seed one (max(major,1).0.0,
so a _v0 or a name without _vN starts at 1.0.0) so the doc becomes tracked. No
bump is computed here.
"""

import argparse
import json
import re
import sys
from datetime import date
from pathlib import Path

sys.path.insert(0, str(Path(__file__).resolve().parents[1])) # .github/scripts
from helpers.json_io import load_json_or_empty
from helpers.doc_types import element_off

VERSIONS_FILE = ".llm_cache/versions.json"

Expand Down Expand Up @@ -68,7 +73,8 @@ def main():
vf = Path(args.versions_file) if args.versions_file else repo_root / VERSIONS_FILE
versions = load_json_or_empty(vf, label="versions")

filled = baseline = 0
filled = baseline = seeded = 0
today = date.today().isoformat()
for qmd in sorted(Path(args.docs_dir).rglob("*.qmd")):
if {"_site", ".quarto", "_meta"} & set(qmd.parts):
continue
Expand All @@ -78,32 +84,54 @@ def main():
continue
_, end = bounds

src = fm_value(lines, end, "original-filename")
major = major_from_name(src) if src else None
if major is None:
major = major_from_name(qmd.name)
if major is None:
print(f"[fill_version] {qmd}: no _vN in filename, skipping")
# version-off types keep their own `version:` header, if any; don't touch.
if element_off(fm_value(lines, end, "type"), "version"):
continue

src = fm_value(lines, end, "original-filename")
mj = major_from_name(src) if src else None
if mj is None:
mj = major_from_name(qmd.name)
# First published version is 1.0.0: a _v0 or a name without _vN starts
# at 1.0.0, while _v4 etc. keep their filename major.
baseline_major = max(mj, 1) if mj is not None else 1

key = f"DOCS/{src}" if src else None
tracked = None
if src:
entry = versions.get(f"DOCS/{src}") or versions.get(src)
if key:
entry = versions.get(key) or versions.get(src)
tracked = (entry or {}).get("current_version")

if tracked and tracked.split(".")[0] == str(major):
if tracked and (mj is None or tracked.split(".")[0] == str(mj)):

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Preserve tracked 1.x versions for _v0 documents

For a _v0.qmd source, the new floor correctly computes baseline_major == 1, and the version updater can subsequently track values such as 1.0.1; however, this condition still compares that tracked major against the raw filename major 0. Every build therefore rejects the tracked value and emits 1.0.0, losing all minor and patch bumps for the legacy _v0 case this change explicitly supports. Compare against baseline_major instead.

AGENTS.md reference: AGENTS.md:L92-L93

Useful? React with 👍 / 👎.

version = tracked
else:
version = f"{major}.0.0"
version = f"{baseline_major}.0.0"
baseline += 1
# No record yet: seed one so the doc becomes tracked, rather than
# re-deriving the fallback every build. Keyed by original-filename.
if key and key not in versions:
versions[key] = {
"current_version": version,
"last_bump": "initial",
"last_bump_reason": "First release",
"last_release_tag": "initial",
"last_updated": today,
"major_from_filename": baseline_major,
}
seeded += 1

if set_version(lines, end, version):
qmd.write_text("".join(lines), encoding="utf-8")
filled += 1

if seeded:
vf.parent.mkdir(parents=True, exist_ok=True)
with vf.open("w", encoding="utf-8") as f:
json.dump(versions, f, indent=2, sort_keys=True)

print(
f"[fill_version] set version on {filled} files "
f"({baseline} fell back to {{major}}.0.0)"
f"({baseline} used max(major,1).0.0 baseline, {seeded} new cache entries)"
)


Expand Down
Loading