diff --git a/COMPARISON.md b/COMPARISON.md index 9bc37c3e..1a9a69ff 100644 --- a/COMPARISON.md +++ b/COMPARISON.md @@ -1,16 +1,33 @@ -# Local-model head-to-head — Coder-Next vs 27B (thinking) vs 27B (no-think) +# Local-model head-to-head — Gemma 4 31B, Coder-Next, and Qwen3.6-27B > **Five-minute decision doc.** The detail lives in [`SCORECARD.md`](SCORECARD.md) and the per-benchmark `findings*.md` docs; this page is the synthesis. Every claim links to its source so you can drill straight into the evidence. > > **Read [`KNOWN-LIMITATIONS.md`](KNOWN-LIMITATIONS.md) before quoting any cell.** Most caveats live there, not here. > -> **Last updated**: 2026-05-02 — reflects [`microbench-phase-b-2026-05-02`](benchmarks/microbench-phase-b-2026-05-02/) (N=10 + 27B-no-think third arm). Pre-no-think readers: the picture has shifted. +> **Last updated**: 2026-08-02 — adds the complete [`Gemma 4 31B Q4 campaign`](benchmarks/gemma4-31b-q4/) while preserving the original three-AWQ-arm tables below. -> **Operating point**: All arms are **Cyankiwi 4-bit AWQ** on **2× RTX PRO 6000 Blackwell at 500 W cap**. Other quants, VRAM tiers, hardware classes, and languages are **not characterized** — see [What this benchmark doesn't characterize](#what-this-benchmark-doesnt-characterize) below. The within-quant comparison here is informative; absolute model capability at higher precisions is a separate question. +> **Operating point**: The original three arms are **Cyankiwi 4-bit AWQ** on **2× RTX PRO 6000 Blackwell at 500 W cap**. Gemma is Google's official QAT Q4_0 GGUF under llama.cpp at model-card sampling and native 262,144-token context. Other quants, VRAM tiers, hardware classes, and languages are **not characterized** — see [What this benchmark doesn't characterize](#what-this-benchmark-doesnt-characterize) below. + +## 2026-08-02 Gemma addendum + +Gemma changes the bounded-quality recommendation, but not the marathon warning: + +| Question | Current evidence-based choice | Evidence | +|---|---|---| +| Highest bounded quality versus Qwen3.6-27B | **Gemma 4 31B Q4** | 29/36 raw and 32/36 corrected at N=3 versus Qwen thinking's 20/36 raw | +| One interactive user | approximately tied | 70.3 tok/s Gemma versus 72.1 tok/s Qwen at short context and 500 W | +| Many simultaneous users | **Qwen3.6-27B/vLLM** | Qwen 1,336.5 aggregate tok/s at C32 on one GPU; Gemma 290.3 at total C8 across two GPUs; shapes differ but the operational gap is large | +| Reliable ordinary completion | approximately tied | Gemma 116/120 `done_signal`; Qwen no-think 113/118 published valid outcomes | +| Unattended marathon work | **none** | Gemma and Qwen both score zero strict passes on their published frozen 75-PR evidence | + +Qwen's 113/118 is not a quality score: its published phase-B grader sweep was +pending. Gemma's 99/120 corrected figure is a quality result, with the raw +89/120 preserved separately. The full Gemma comparison and artifact audit are +in [`benchmarks/gemma4-31b-q4/GEMMA4_31B_Q4_VERIFIED_RESULTS.md`](benchmarks/gemma4-31b-q4/GEMMA4_31B_Q4_VERIFIED_RESULTS.md). ## TL;DR -**No model is overall best.** The three arms have orthogonal strengths and statistically indistinguishable headline ship rates (74–96%). Pick by task class: +**No model is overall best.** Within the original three AWQ arms, the models have orthogonal strengths and statistically indistinguishable headline ship rates (74–96%). Pick by task class; include Gemma when bounded answer quality matters more than dense batching: - **Default for most non-coding tasks**: **27B-no-think** — 95.8% ship rate across 12 cells × N=10, beats both originals on raw shipping - **Hallucination-sensitive or research-driven work**: **27B (either mode)** — 27B-thinking is the only arm that ships market-research at >70%; no-think 10/10 on adversarial-hallucination diff --git a/MICROBENCH-INDEX.md b/MICROBENCH-INDEX.md index db4277d2..567415b4 100644 --- a/MICROBENCH-INDEX.md +++ b/MICROBENCH-INDEX.md @@ -13,6 +13,7 @@ | Entry | Tree | Models / arms | N | Headline | |---|---|---|---|---| +| [`gemma4-31b-q4`](benchmarks/gemma4-31b-q4/) | benchmarks/ | **Gemma-4-31B-it-QAT-Q4_0** at model-card sampling | 3 and 10 | N=3 raw 29/36, corrected 32/36; N=10 raw 89/120, corrected 99/120. Extended strict audit 0/12. | | [`deepseek-v4-flash-0731`](benchmarks/deepseek-v4-flash-0731/) | benchmarks/ | **DeepSeek-V4-Flash-0731** at its model-card sampling point | 3 | Raw 23/36; corrected 35/36 after reproducible grader defects were repaired without rerunning the model. One genuine failure: 773 words against a 700-word cap. | | [`microbench-2026-04-28`](benchmarks/microbench-2026-04-28/) | benchmarks/ | Qwen3.6-**27B-AWQ** vs Qwen3-Coder-Next-**AWQ** | 3 | Aggregate-tied ~7/12 each; complementary task-class strengths; Coder-Next much faster/cheaper. | | [`microbench-phase-b-2026-05-02`](benchmarks/microbench-phase-b-2026-05-02/) | benchmarks/ | + **27B-AWQ no-think** third arm; 4 differential cells to N=10 | 10 | 27B ships 86.8% no-think vs 75% think (same `p3_doc` word-limit loop); Coder-Next market 0/10 (Wilson [0, 27.8%]). | @@ -49,6 +50,16 @@ Qwen3.5-397B writing overlay moves its N=10 no-think result from 82/120 to regrade and DeepSeek N=10 expansion would be required for a statistically matched ranking. +## Gemma adds a broader high-scoring local cohort + +Gemma's corrected **99/120 (82.5%)** at N=10 is the largest high-scoring +canonical cohort in the current local set, and its N=3 corrected 32/36 is below +DeepSeek's 35/36 but above Qwen3.6-27B and Coder-Next's 20/36 raw results. This +still is not a global leaderboard: Gemma uses different sampling and a newer +grader/campaign date. Its strict extended result is 0/12, which is why the +canonical aggregate must not be read as proof of reliable complex-artifact or +marathon execution. + ## Qualitative comparison Most historical pass-rates tie; the qualitative layer is where those models actually differ. Provisional cross-model diff --git a/README.md b/README.md index b28e7785..c0e6fd73 100644 --- a/README.md +++ b/README.md @@ -15,6 +15,7 @@ but I'm making it public so that other people can use it too. | Where the benchmark folders start | [`benchmarks/README.md`](benchmarks/README.md) — agent-task benchmark landing page | | **"Coder-Next or 27B (or 27B-no-think) for my task?"** | [`COMPARISON.md`](COMPARISON.md) — head-to-head decision doc | | The full single-table comparison across all entries | [`SCORECARD.md`](SCORECARD.md) | +| **Gemma 4 31B QAT Q4: complete verified campaign** | [`benchmarks/gemma4-31b-q4/`](benchmarks/gemma4-31b-q4/) — native-256K serving, canonical N=3/N=10, strict artifact audits, and Qwen3.6-27B comparison | | **DeepSeek V4 Flash 0731: complete verified campaign** | [`benchmarks/deepseek-v4-flash-0731/`](benchmarks/deepseek-v4-flash-0731/) — deployment, canonical N=3, extended suites, strict artifact audits, and full-context 75-PR outcomes | | **All 12-family microbench results (across both trees) + the four "27B"s** | [`MICROBENCH-INDEX.md`](MICROBENCH-INDEX.md) — cross-tree microbench index + quant disambiguation | | Cross-model **qualitative** spot-grades (provisional, not a ranking) | [`QUALITATIVE-SPOT-GRADES.md`](QUALITATIVE-SPOT-GRADES.md) + [`tooling/QUALITATIVE-GRADING-PROTOCOL.md`](tooling/QUALITATIVE-GRADING-PROTOCOL.md) — single-grader provisional scores + grader-independence rules | @@ -24,7 +25,7 @@ but I'm making it public so that other people can use it too. ## Operating point (read before quoting) -Most earlier agent-task benchmarks under [`benchmarks/`](benchmarks/) use **Cyankiwi 4-bit AWQ** quants on **2x RTX PRO 6000 Blackwell at 500 W cap**. The DeepSeek V4 Flash 0731 entry is an explicit exception: official FP4 weights, FP8 KV, model-card sampling, and a validated 1,048,576-token context. Every entry README pins its own operating point. Other quants, VRAM tiers, hardware classes, and languages other than Python are **not characterized** unless an entry says otherwise. See [`COMPARISON.md` section "What this benchmark doesn't characterize"](COMPARISON.md#what-this-benchmark-doesnt-characterize) for the model-benchmark validity boundaries, and [`ROADMAP.md`](ROADMAP.md) for what's queued to fill those gaps. +Most earlier agent-task benchmarks under [`benchmarks/`](benchmarks/) use **Cyankiwi 4-bit AWQ** quants on **2x RTX PRO 6000 Blackwell at 500 W cap**. DeepSeek V4 Flash 0731 and Gemma 4 31B are explicit exceptions: DeepSeek uses official FP4 weights, FP8 KV, and a validated 1,048,576-token context; Gemma uses Google's official QAT Q4_0 GGUF, Q8 KV, llama.cpp, and a validated native 262,144-token context per slot. Every entry README pins its own operating point. Other quants, VRAM tiers, hardware classes, and languages other than Python are **not characterized** unless an entry says otherwise. See [`COMPARISON.md` section "What this benchmark doesn't characterize"](COMPARISON.md#what-this-benchmark-doesnt-characterize) for the model-benchmark validity boundaries, and [`ROADMAP.md`](ROADMAP.md) for what's queued to fill those gaps. Rig-characterisation studies under [`hardware-tests/`](hardware-tests/) have their own operating-point scope. Start with [`hardware-tests/README.md`](hardware-tests/README.md) before quoting hardware claims. In particular, [`hardware-tests/qwen3.6-q8-fleet-2026-05-17/`](hardware-tests/qwen3.6-q8-fleet-2026-05-17/) ranks four hardware classes on **Q8_0 GGUF** dense and MoE workloads under llama.cpp, with a Tower2 vLLM-FP8 appendix row for the MoE model the llama.cpp/CUDA path crashes on. @@ -33,6 +34,7 @@ Rig-characterisation studies under [`hardware-tests/`](hardware-tests/) have the ```text benchmarks/ README.md agent-task benchmark landing page and navigation map + gemma4-31b-q4/ complete Gemma campaign, N=3/N=10, strict audits, comparisons deepseek-v4-flash-0731/ cross-suite DeepSeek campaign, strict audits, and deployment evidence dreamserver-75-pr-audit/ GPT-5.5/ cloud, full audit @@ -63,6 +65,7 @@ hardware-tests/ | Benchmark | Prompt Shape | Model Entries | |---|---|---| +| [`gemma4-31b-q4`](benchmarks/gemma4-31b-q4/) | Cross-suite publication: full 12-family N=3 and N=10, single-PR N=3, investment research, board presentation, and frozen 75-PR N=3. | `Gemma-4-31B-it-QAT-Q4_0`; includes immutable raw grades, a narrow corrected overlay, strict substantive audit, and Qwen3.6-27B comparison. | | [`deepseek-v4-flash-0731`](benchmarks/deepseek-v4-flash-0731/) | Cross-suite publication: full 12-family N=3, single-PR N=3, investment research, board presentation, and three valid full-context frozen 75-PR outcomes. | `DeepSeek-V4-Flash-0731`; includes corrected grader overlay and compact audit evidence. | | [`dreamserver-75-pr-audit`](benchmarks/dreamserver-75-pr-audit/) | Audit 75 open PRs in a live repository and produce a traceable maintainer triage repo. | `GPT-5.5`, `Opus-4.7`, `Qwen3.6-27B-AWQ`, `Qwen3-Coder-Next-AWQ` (failure-mode entry) | | [`dreamserver-1-pr-audit`](benchmarks/dreamserver-1-pr-audit/) | Same task spec, scaled to a single PR. Built as the floor of an escalation ladder (1 → 2 → 4 → 8 → 16 → 32) to find each model's complexity ceiling. | `Qwen3-Coder-Next-AWQ`, `Qwen3.6-27B-AWQ`, `Qwen3.6-35B-A3B-AWQ` (floor failure) | @@ -93,6 +96,7 @@ For the benchmark landing page and per-folder navigation map, start with Two synthesis docs sit between this README and the per-entry detail: +- [`benchmarks/gemma4-31b-q4/GEMMA4_31B_Q4_VERIFIED_RESULTS.md`](benchmarks/gemma4-31b-q4/GEMMA4_31B_Q4_VERIFIED_RESULTS.md) — full Gemma result, native-256K deployment, N=3/N=10 quality, strict artifact audit, and direct Qwen3.6-27B comparison. - [`benchmarks/deepseek-v4-flash-0731/DEEPSEEK_V4_FLASH_0731_VERIFIED_RESULTS.md`](benchmarks/deepseek-v4-flash-0731/DEEPSEEK_V4_FLASH_0731_VERIFIED_RESULTS.md) — full verified DeepSeek result, including corrected-vs-raw grading, strict artifact audits, production validation, and caveats. - [`COMPARISON.md`](COMPARISON.md) — **head-to-head decision doc** for the three local model arms (Coder-Next vs 27B-thinking vs 27B-no-think). Organized by task class with cell-level evidence. Read this if your question is "which one should I use?" - [`SCORECARD.md`](SCORECARD.md) — single-table summary across all entries (spec compliance, factual accuracy, fabricated-claim count, tests run, wall, cost upper bound, failure mode, "when to use which" guide). Read this if your question is "what's the full picture?" @@ -128,6 +132,10 @@ The repo keeps the failures because the *kinds* of failure are themselves the co ## Current Entries +**gemma4-31b-q4:** +- [Verified campaign entry](benchmarks/gemma4-31b-q4/) — Gemma 4 31B QAT Q4_0 on two independent 500 W RTX PRO 6000 replicas, with native 262,144-token context per slot. Canonical N=10 is 89/120 raw and 99/120 corrected; extended strict result is 0/12. It beats Qwen3.6-27B on directly comparable bounded quality but not on batched serving. +- [Completion audit](benchmarks/gemma4-31b-q4/GEMMA4_31B_Q4_COMPLETION_AUDIT.md) — requirement-to-evidence handoff, including the pending production restore gate. + **deepseek-v4-flash-0731:** - [Verified campaign entry](benchmarks/deepseek-v4-flash-0731/) — DeepSeek V4 Flash 0731 on 2x RTX PRO 6000 at 500 W/GPU and 1,048,576 context. Canonical corrected result 35/36 (N=3); single-PR 2/3 expected verdicts with all three complete; investment workbooks 0/2 substantively valid; board deck shipped with material visual defects; frozen 75-PR strict result 0/3 (two scaffold-and-stop, one 815,279-token runaway-generation terminal failure). - [Completion audit](benchmarks/deepseek-v4-flash-0731/DEEPSEEK_V4_FLASH_0731_COMPLETION_AUDIT.md) — requirement-to-evidence handoff and artifact inventory. diff --git a/SCORECARD.md b/SCORECARD.md index 4ef706aa..adee898f 100644 --- a/SCORECARD.md +++ b/SCORECARD.md @@ -15,8 +15,12 @@ > **DeepSeek V4 Flash 0731 is a newer cross-suite entry** with its own model-card sampling, 1,048,576-token > context, strict artifact audits, and corrected grader overlay. Its raw and corrected results are reported > separately; do not silently mix its corrected N=3 result with older raw-score or ship-rate columns. +> **Gemma 4 31B QAT Q4_0 is the newest cross-suite entry** with native 262,144-token context, model-card +> sampling, N=3 and N=10 cohorts, and strict artifact audits. Its project-management correction is a separate +> hash-tied overlay; its 116/120 completion count is not interchangeable with its 99/120 corrected quality score. > > **Newer microbench results (summary; full detail in the linked entries):** +> - **Gemma 4 31B QAT Q4_0**, N=10: raw 89/120; corrected 99/120 (82.5%); 116/120 `done_signal`; extended strict result 0/12. ([entry](benchmarks/gemma4-31b-q4/)) > - **DeepSeek V4 Flash 0731**, N=3: raw 23/36; corrected 35/36 (97.2%) after reproducible grader defects were repaired without rerunning the model. ([entry](benchmarks/deepseek-v4-flash-0731/)) > - **397B-A17B** (Q3 GGUF), N=10: no-think 82/120, think 72/120 — *thinking net-negative*. > - **27B-FP8**, N=5: no-think 35/60, think 29/60 — *thinking net-negative*; FP8 serving stable where Q8 failed. ([entry](hardware-tests/qwen3.6-27b-fp8-microbench-2026-05-31/)) @@ -63,6 +67,7 @@ The microbench tables below alternate between p-codes (used in receipt names and |---|---|---|---|---|---|---|---|---| | **Opus-4.7** (cloud) | 1/1 | ✓ full | not graded (would need per-PR ground truth across all 75) | not graded | not visible from artifacts | ~5 hr | n/a | success-shipped | | **GPT-5.5** (cloud) | 1/1 | ✓ full + `verify_coverage.py` self-check passes | not graded | not graded | not visible from artifacts | not recorded | n/a | success-shipped | +| **Gemma-4-31B-QAT-Q4_0** (local) | 3/3 valid + 1 invalid monitor attempt preserved | ✗ **0/3 strict**; 6/75, 2/75, then 75 dirs with only 14 verdicts | not graded per PR | not graded | 6, 0, and 4 PRs with test evidence | 3.2-9.4 min | see receipts | 3× model-terminal/incomplete audit | | **DeepSeek-V4-Flash-0731** (local) | 3 valid / 5 preserved | ✗ **0/3 strict**; v4 created 75 packages but all reviews were shallow | not graded per PR | not graded | extensive investigation, but incomplete final test/skip and tool records | 12.1-53.9 min | see receipts | 2× scaffold-and-stop; 1× full-context runaway-generation | | **Qwen3.6-27B-AWQ** (local) | 1/3 published | △ 75/75 verdict.md *files* but only 3 are real reviews; 72 are template stubs | partial (3 reviewed PRs match ground truth; 72 stubs unverified) | 0 in the 3 deep reviews | **0** | 24 min | $0.031 | scaffold-and-stop | | **Qwen3-Coder-Next-AWQ** (local) | 0/5 | ✗ no deliverable across 5 attempts | n/a | n/a | n/a | 1-42 min | $0.001-$0.054 | identical-call-loop, cyclic-name-slop, stuck-in-research | @@ -73,6 +78,7 @@ The microbench tables below alternate between p-codes (used in receipt names and | Model | Runs | Spec | Factual accuracy | Fabricated | Tests | Wall | Cost | Failure mode | |---|---|---|---|---|---|---|---|---| +| **Gemma-4-31B-QAT-Q4_0** | 3/3 complete workspaces | ✗ 0/3 common provenance gate; pinned subject refs omitted | **2/3 expected MERGE**; v3 wrongly REJECTed the subject | no material fabrication found in v1/v2 | independent reproduction: 67 base / 76 head / 261 current | 4.8-5.0 min | see receipts | 2× correct-but-provenance-incomplete; 1× wrong-subject verdict | | **DeepSeek-V4-Flash-0731** | 3/3 complete | ✓ full repos, tags, done() | **2/3 expected MERGE**; 1/3 defensible but over-strict REVISE | no material fabrication found | baseline/PR/current-main tests and reproductions recorded | 5.1-25.5 min | see receipts | 2× success; 1× over-investigation | | **Qwen3-Coder-Next-AWQ** | 1/3 published (cherry-picked correct) | ✓ 13/13 files, tag, done() | **2/3 wrong** across the three runs (this entry's run is the 1 correct; v1 and v3 said REJECT incorrectly) | 1 in v1, **4 in v3** including a fabricated `test_stderr_truncation.py` | repro script, no execution | 3 min | $0.004 | success-shipped (cherry-picked) | | **Qwen3.6-27B-AWQ** | 1/3 published | ✗ 7/13 files; **no verdict.md, no tag, no done()** in any of 3 runs | 3/3 implicit-MERGE-correct (in `review.md`'s Summary of Findings table; never in a `verdict.md`) | 0 | **pytest invoked, 38 tests on both branches** | 7 min | $0.009 | partial-no-spec-output | @@ -84,6 +90,7 @@ The microbench tables below alternate between p-codes (used in receipt names and | Model | Company | Rec | Runs | Spec | Factual accuracy | Fabricated | Wall | Cost | Failure mode | |---|---|---|---|---|---|---|---|---|---| +| **Gemma-4-31B-QAT-Q4_0** (local) | Dropbox / Duolingo / Coursera | BUY / BUY / BUY | 3/3 | ✗ all omit required PDF; **0/3 substantive finance pass** | targets unsupported by workbook mechanics | v2/v3 workbooks have zero formulas; v1 valuation bridge contradicts $110 target | 6.9-8.8 min | see receipts | polished-but-financially-invalid | | **Opus-4.7** (cloud) | Vita Coco (`COCO`) | HOLD ($46 vs $52 spot) | 1/1 | ✓ full memo + machine-readable verification | not graded (opinion) | not graded | not recorded | n/a | success-shipped | | **GPT-5.5** (cloud) | YETI Holdings (`YETI`) | HOLD ($41) | 1/1 | ✓ full memo + verification + board-deck follow-on | not graded (opinion) | not graded | not recorded | n/a | success-shipped | | **DeepSeek-V4-Flash-0731** (local) | Knife River / Ollie's | HOLD / BUY | 2 audited shipped runs | ✓ historical artifact gate; **0/2 substantive finance pass** | unit/convention errors and incomplete statements | zero workbook formulas in both | 9.5-11.6 min | see receipts | polished-but-financially-invalid | @@ -93,6 +100,29 @@ The microbench tables below alternate between p-codes (used in receipt names and --- +## Gemma 4 31B QAT Q4_0 cross-suite result + +The complete entry is at [`benchmarks/gemma4-31b-q4/`](benchmarks/gemma4-31b-q4/). +Its canonical N=3 result is **29/36 raw and 32/36 corrected**; the broader N=10 +result is **89/120 raw and 99/120 corrected**. The only overlay changes ten +project-management lexical false negatives and never overwrites original +grades. Four hallucination runs stopped without required output; no canonical +failure is an artificial low-output-ceiling event. + +Against Qwen3.6-27B thinking's directly comparable 20/36 raw result, Gemma is +the bounded-quality winner. Short-context single-stream speed is effectively +tied at 500 W (70.3 versus 72.1 tok/s), while Qwen's vLLM deployment is vastly +better at dense batching. Qwen no-think's 113/118 is a completion rate with the +quality sweep pending, not a score comparable to Gemma's corrected 99/120. + +The strict extended result is **0/12**. All three finance models are +substantively invalid, two decks omit PDF and the third has material visual +defects, every single-PR result fails pinned-subject provenance, and all three +75-PR attempts fail completeness/substance. Gemma is excellent on bounded +tasks but is not a trustworthy unattended artifact or marathon agent. + +--- + ## DeepSeek V4 Flash 0731 cross-suite result The complete entry is at [`benchmarks/deepseek-v4-flash-0731/`](benchmarks/deepseek-v4-flash-0731/). @@ -243,6 +273,13 @@ not reliable unattended marathon delivery. - **Do not assign one monolithic unattended marathon.** The strict 75-PR campaign is 0/3. Partition the work, cap each subtask, require artifact gates between batches, and detect responses that stop calling tools. - **Keep Qwen3.5-397B as a conservative comparison arm.** Its N=10 evidence is broader and its failure temperament is less explosive, although its corrected bounded score is lower and its serving rate/context are much smaller. +### When to use **Gemma 4 31B QAT Q4_0** + +- **Bounded high-quality work on one GPU.** It beats Qwen3.6-27B and Coder-Next on the directly comparable N=3 matrix and gives every slot the native 256K context. +- **Two isolated production agents.** One full model per GPU avoids cross-GPU dependencies; Sanctuary and Pixel each retain four slots and independent failure domains. +- **Single-user quality over high-concurrency throughput.** Short-context decode is close to Qwen3.6-27B, but Qwen/vLLM is the clear batched-serving winner. +- **Require hard stage gates for complex artifacts.** The extended strict result is 0/12. Finance, slide layout, pinned provenance, and marathon completeness all need independent validation. + ### When to use **Qwen3.6-27B-AWQ** - **Hallucination resistance is required.** The single sharpest local-model superiority signal in this repo: on the adversarial-hallucination microbench (15 issues, 6 real / 9 fabricated, agent must classify), 27B was 3/3 PASS with 100% accuracy and 0 dangerous errors; Coder-Next was 1/3 PASS with 2 confirmed-fabrications-as-real on the one shipping run. For security review, factual research, anything where confidently-wrong is dangerous, 27B is the pick. diff --git a/benchmarks/README.md b/benchmarks/README.md index 05d43bec..6061acc8 100644 --- a/benchmarks/README.md +++ b/benchmarks/README.md @@ -10,6 +10,7 @@ deliverable. Hardware throughput and buyer-value questions live in | If you want | Read first | |---|---| +| Gemma 4 31B Q4 complete campaign and Qwen3.6-27B comparison | [`gemma4-31b-q4/README.md`](gemma4-31b-q4/README.md) | | DeepSeek V4 Flash 0731 complete campaign | [`deepseek-v4-flash-0731/README.md`](deepseek-v4-flash-0731/README.md) | | Current model-selection synthesis | [`../COMPARISON.md`](../COMPARISON.md) | | Single-table benchmark summary | [`../SCORECARD.md`](../SCORECARD.md) | @@ -20,6 +21,7 @@ deliverable. Hardware throughput and buyer-value questions live in | Folder | Primary question | Best first file | |---|---|---| +| [`gemma4-31b-q4`](gemma4-31b-q4/) | How does the official Gemma 4 31B QAT Q4 perform at native 256K across bounded, artifact, and marathon tasks, especially versus Qwen3.6-27B? | [`gemma4-31b-q4/README.md`](gemma4-31b-q4/README.md) | | [`deepseek-v4-flash-0731`](deepseek-v4-flash-0731/) | How does the fully optimized DeepSeek deployment perform across the canonical, extended, artifact-quality, and full-context marathon suites? | [`deepseek-v4-flash-0731/README.md`](deepseek-v4-flash-0731/README.md) | | [`dreamserver-75-pr-audit`](dreamserver-75-pr-audit/) | Can the model complete a long-horizon 75-PR maintainer audit at all? | [`dreamserver-75-pr-audit/README.md`](dreamserver-75-pr-audit/README.md) | | [`dreamserver-1-pr-audit`](dreamserver-1-pr-audit/) | What is the floor task for local PR-audit competence? | [`dreamserver-1-pr-audit/README.md`](dreamserver-1-pr-audit/README.md) | @@ -58,3 +60,8 @@ hashes, run classifications, and reproduction tooling live in git. Multi-megabyt workspace archives and the 137 MB terminal-run archive remain outside git under the large-artifact policy; their byte counts and SHA-256 identities are recorded in the published audits. + +The Gemma entry uses the same lean-audit policy. Its scorecards, grader +manifests, correction overlays, evidence audits, and compact 75-PR audits live +in git. Full workspace archives, rendered decks, and workbook inspection trees +remain external; the published audits bind them by byte count and SHA-256. diff --git a/benchmarks/gemma4-31b-q4/75pr-v1-audit.json b/benchmarks/gemma4-31b-q4/75pr-v1-audit.json new file mode 100644 index 00000000..46c17235 --- /dev/null +++ b/benchmarks/gemma4-31b-q4/75pr-v1-audit.json @@ -0,0 +1,1314 @@ +{ + "schema_version": 2, + "audited_at": "2026-08-02T06:46:32.142593+00:00", + "run_name": "Gemma-4-31B-it-QAT-Q4_0 gemma4-31b-q4_75pr_v1", + "classification": "MODEL_TERMINAL_FAILURE", + "legacy_shipped": false, + "verdict_counts": { + "merge": 6 + }, + "runtime": { + "wall_s": 496.2, + "iterations": 41, + "completion_tokens": 17996, + "prompt_tokens": 2129829, + "completion_tps_model_call": 37.0 + }, + "artifact_integrity": { + "archive_sha256": "2986ef9cc929b714eb48976c400957248f7f289376fbe5c2886c9d4f8f45141e", + "archive_bytes": 34921336, + "archive_gzip_valid": true, + "head": "ebf937f757ccbdba71214f0aa9e3a5037025499f", + "tags": [ + "v1.0.0" + ], + "tag_targets": { + "v1.0.0": "ebf937f757ccbdba71214f0aa9e3a5037025499f" + }, + "tag_at_head": true, + "clean": true, + "commit_count": 3 + }, + "substance": { + "test_evidence_count": 6, + "test_evidence_prs": [ + 1000, + 1002, + 1016, + 1029, + 1056, + 1057 + ], + "missing_test_or_skip_count": 69, + "missing_test_or_skip_prs": [], + "reviews_under_800_bytes_count": 5, + "reviews_under_800_bytes_prs": [ + 1000, + 1002, + 1016, + 1029, + 1057 + ], + "reviews_with_source_citation_count": 6, + "reviews_with_source_citation_prs": [ + 1000, + 1002, + 1016, + 1029, + 1056, + 1057 + ], + "reviews_with_hunk_marker_count": 0, + "traces_with_source_citation_count": 6, + "traces_with_hunk_marker_count": 0, + "diff_analyses_with_hunk_marker_count": 0, + "explicit_actual_bounty_tier_count": 5, + "explicit_actual_bounty_tier_prs": [ + 1000, + 1002, + 1016, + 1029, + 1057 + ], + "transcript_tool_events": 41, + "numbered_tool_log_entries": 0 + }, + "strict_validation": { + "schema_version": 1, + "workspace": "/home/michael/bench-gemma4-31b-q4/tooling/workspace/gemma4-31b-q4_75pr_v1/audit-repo", + "fixture": "/mnt/bulk/benchmark-fixtures/dreamserver-75-pr-audit-2026-04-27", + "pass": false, + "metrics": { + "canonical_prs": 75, + "actual_pr_dirs": 6, + "missing_pr_dirs": [ + 351, + 364, + 716, + 750, + 959, + 961, + 966, + 973, + 974, + 983, + 988, + 990, + 991, + 992, + 993, + 994, + 996, + 997, + 998, + 999, + 1003, + 1004, + 1005, + 1006, + 1007, + 1008, + 1009, + 1010, + 1011, + 1012, + 1013, + 1014, + 1015, + 1017, + 1018, + 1019, + 1020, + 1021, + 1022, + 1023, + 1024, + 1025, + 1026, + 1027, + 1028, + 1030, + 1032, + 1033, + 1034, + 1035, + 1036, + 1037, + 1038, + 1039, + 1040, + 1042, + 1043, + 1044, + 1045, + 1046, + 1047, + 1048, + 1049, + 1050, + 1051, + 1052, + 1053, + 1054, + 1055 + ], + "extra_pr_dirs": [], + "verdict_counts": { + "merge": 6 + }, + "risk_matrix_missing_pr_refs": [ + 351, + 364, + 716, + 750, + 959, + 961, + 966, + 973, + 974, + 983, + 988, + 990, + 991, + 992, + 993, + 994, + 996, + 997, + 998, + 999, + 1002, + 1003, + 1004, + 1005, + 1006, + 1007, + 1008, + 1009, + 1010, + 1011, + 1012, + 1013, + 1014, + 1015, + 1017, + 1018, + 1019, + 1020, + 1021, + 1022, + 1023, + 1024, + 1025, + 1026, + 1027, + 1028, + 1030, + 1032, + 1033, + 1034, + 1035, + 1036, + 1037, + 1038, + 1039, + 1040, + 1042, + 1043, + 1044, + 1045, + 1046, + 1047, + 1048, + 1049, + 1050, + 1051, + 1052, + 1053, + 1054, + 1055 + ], + "raw_frozen_file_overlap_pairs": 676, + "trusted_file_list_prs": 63, + "trusted_file_overlap_pairs": 205, + "listed_pair_mentions": 1, + "missing_trusted_file_overlap_pairs": 205, + "missing_trusted_file_overlap_pair_sample": [ + [ + 990, + 991 + ], + [ + 992, + 1013 + ], + [ + 992, + 1017 + ], + [ + 993, + 994 + ], + [ + 993, + 997 + ], + [ + 993, + 998 + ], + [ + 993, + 999 + ], + [ + 993, + 1000 + ], + [ + 993, + 1002 + ], + [ + 993, + 1006 + ], + [ + 993, + 1007 + ], + [ + 993, + 1008 + ], + [ + 993, + 1011 + ], + [ + 993, + 1016 + ], + [ + 993, + 1018 + ], + [ + 993, + 1020 + ], + [ + 994, + 997 + ], + [ + 994, + 998 + ], + [ + 994, + 999 + ], + [ + 994, + 1000 + ], + [ + 994, + 1002 + ], + [ + 994, + 1006 + ], + [ + 994, + 1007 + ], + [ + 994, + 1008 + ], + [ + 994, + 1010 + ], + [ + 994, + 1011 + ], + [ + 994, + 1016 + ], + [ + 994, + 1017 + ], + [ + 994, + 1018 + ], + [ + 994, + 1020 + ], + [ + 996, + 1012 + ], + [ + 996, + 1017 + ], + [ + 996, + 1026 + ], + [ + 997, + 998 + ], + [ + 997, + 999 + ], + [ + 997, + 1000 + ], + [ + 997, + 1002 + ], + [ + 997, + 1006 + ], + [ + 997, + 1007 + ], + [ + 997, + 1008 + ], + [ + 997, + 1011 + ], + [ + 997, + 1016 + ], + [ + 997, + 1018 + ], + [ + 997, + 1020 + ], + [ + 998, + 999 + ], + [ + 998, + 1000 + ], + [ + 998, + 1002 + ], + [ + 998, + 1006 + ], + [ + 998, + 1007 + ], + [ + 998, + 1008 + ] + ], + "missing_raw_frozen_overlap_pairs": 676, + "missing_raw_frozen_overlap_pair_sample": [ + [ + 351, + 364 + ], + [ + 351, + 716 + ], + [ + 351, + 750 + ], + [ + 351, + 959 + ], + [ + 351, + 961 + ], + [ + 351, + 966 + ], + [ + 351, + 973 + ], + [ + 351, + 974 + ], + [ + 351, + 983 + ], + [ + 351, + 988 + ], + [ + 351, + 990 + ], + [ + 351, + 991 + ], + [ + 351, + 992 + ], + [ + 351, + 993 + ], + [ + 351, + 994 + ], + [ + 351, + 996 + ], + [ + 351, + 997 + ], + [ + 351, + 998 + ], + [ + 351, + 999 + ], + [ + 351, + 1000 + ], + [ + 351, + 1002 + ], + [ + 351, + 1003 + ], + [ + 351, + 1004 + ], + [ + 351, + 1005 + ], + [ + 351, + 1006 + ], + [ + 351, + 1007 + ], + [ + 351, + 1008 + ], + [ + 351, + 1009 + ], + [ + 351, + 1010 + ], + [ + 351, + 1011 + ], + [ + 351, + 1012 + ], + [ + 351, + 1013 + ], + [ + 351, + 1015 + ], + [ + 351, + 1016 + ], + [ + 351, + 1017 + ], + [ + 351, + 1018 + ], + [ + 351, + 1019 + ], + [ + 351, + 1020 + ], + [ + 351, + 1021 + ], + [ + 351, + 1022 + ], + [ + 351, + 1023 + ], + [ + 351, + 1024 + ], + [ + 351, + 1025 + ], + [ + 351, + 1026 + ], + [ + 351, + 1027 + ], + [ + 351, + 1028 + ], + [ + 351, + 1029 + ], + [ + 351, + 1030 + ], + [ + 351, + 1032 + ], + [ + 351, + 1033 + ] + ], + "raw_file_count_metadata_mismatch_prs": [ + 351, + 364, + 716, + 750, + 959, + 961, + 966, + 973, + 974, + 983, + 988, + 1043 + ], + "transcript_tool_events": 41, + "numbered_tool_log_entries": 0 + }, + "failures": [ + "missing or empty required artifact: sources.md", + "missing or empty required artifact: analysis/surface-area.md", + "missing or empty required artifact: testing/baseline.md", + "missing or empty required artifact: research/dead-ends.md", + "missing or empty required artifact: research/upstream-context.md", + "missing non-empty decisions directory", + "missing PR directories: [351, 364, 716, 750, 959, 961, 966, 973, 974, 983, 988, 990, 991, 992, 993, 994, 996, 997, 998, 999, 1003, 1004, 1005, 1006, 1007, 1008, 1009, 1010, 1011, 1012, 1013, 1014, 1015, 1017, 1018, 1019, 1020, 1021, 1022, 1023, 1024, 1025, 1026, 1027, 1028, 1030, 1032, 1033, 1034, 1035, 1036, 1037, 1038, 1039, 1040, 1042, 1043, 1044, 1045, 1046, 1047, 1048, 1049, 1050, 1051, 1052, 1053, 1054, 1055]", + "PR 351: missing or empty verdict.md", + "PR 351: missing or empty summary.md", + "PR 351: missing or empty review.md", + "PR 351: missing or empty diff-analysis.md", + "PR 351: missing or empty interactions.md", + "PR 351: missing or empty trace.md", + "PR 351: no merge/revise/reject disposition near verdict start", + "PR 351: no non-empty test-evidence file", + "PR 364: missing or empty verdict.md", + "PR 364: missing or empty summary.md", + "PR 364: missing or empty review.md", + "PR 364: missing or empty diff-analysis.md", + "PR 364: missing or empty interactions.md", + "PR 364: missing or empty trace.md", + "PR 364: no merge/revise/reject disposition near verdict start", + "PR 364: no non-empty test-evidence file", + "PR 716: missing or empty verdict.md", + "PR 716: missing or empty summary.md", + "PR 716: missing or empty review.md", + "PR 716: missing or empty diff-analysis.md", + "PR 716: missing or empty interactions.md", + "PR 716: missing or empty trace.md", + "PR 716: no merge/revise/reject disposition near verdict start", + "PR 716: no non-empty test-evidence file", + "PR 750: missing or empty verdict.md", + "PR 750: missing or empty summary.md", + "PR 750: missing or empty review.md", + "PR 750: missing or empty diff-analysis.md", + "PR 750: missing or empty interactions.md", + "PR 750: missing or empty trace.md", + "PR 750: no merge/revise/reject disposition near verdict start", + "PR 750: no non-empty test-evidence file", + "PR 959: missing or empty verdict.md", + "PR 959: missing or empty summary.md", + "PR 959: missing or empty review.md", + "PR 959: missing or empty diff-analysis.md", + "PR 959: missing or empty interactions.md", + "PR 959: missing or empty trace.md", + "PR 959: no merge/revise/reject disposition near verdict start", + "PR 959: no non-empty test-evidence file", + "PR 961: missing or empty verdict.md", + "PR 961: missing or empty summary.md", + "PR 961: missing or empty review.md", + "PR 961: missing or empty diff-analysis.md", + "PR 961: missing or empty interactions.md", + "PR 961: missing or empty trace.md", + "PR 961: no merge/revise/reject disposition near verdict start", + "PR 961: no non-empty test-evidence file", + "PR 966: missing or empty verdict.md", + "PR 966: missing or empty summary.md", + "PR 966: missing or empty review.md", + "PR 966: missing or empty diff-analysis.md", + "PR 966: missing or empty interactions.md", + "PR 966: missing or empty trace.md", + "PR 966: no merge/revise/reject disposition near verdict start", + "PR 966: no non-empty test-evidence file", + "PR 973: missing or empty verdict.md", + "PR 973: missing or empty summary.md", + "PR 973: missing or empty review.md", + "PR 973: missing or empty diff-analysis.md", + "PR 973: missing or empty interactions.md", + "PR 973: missing or empty trace.md", + "PR 973: no merge/revise/reject disposition near verdict start", + "PR 973: no non-empty test-evidence file", + "PR 974: missing or empty verdict.md", + "PR 974: missing or empty summary.md", + "PR 974: missing or empty review.md", + "PR 974: missing or empty diff-analysis.md", + "PR 974: missing or empty interactions.md", + "PR 974: missing or empty trace.md", + "PR 974: no merge/revise/reject disposition near verdict start", + "PR 974: no non-empty test-evidence file", + "PR 983: missing or empty verdict.md", + "PR 983: missing or empty summary.md", + "PR 983: missing or empty review.md", + "PR 983: missing or empty diff-analysis.md", + "PR 983: missing or empty interactions.md", + "PR 983: missing or empty trace.md", + "PR 983: no merge/revise/reject disposition near verdict start", + "PR 983: no non-empty test-evidence file", + "PR 988: missing or empty verdict.md", + "PR 988: missing or empty summary.md", + "PR 988: missing or empty review.md", + "PR 988: missing or empty diff-analysis.md", + "PR 988: missing or empty interactions.md", + "PR 988: missing or empty trace.md", + "PR 988: no merge/revise/reject disposition near verdict start", + "PR 988: no non-empty test-evidence file", + "PR 990: missing or empty verdict.md", + "PR 990: missing or empty summary.md", + "PR 990: missing or empty review.md", + "PR 990: missing or empty diff-analysis.md", + "PR 990: missing or empty interactions.md", + "PR 990: missing or empty trace.md", + "PR 990: no merge/revise/reject disposition near verdict start", + "PR 990: no non-empty test-evidence file", + "PR 991: missing or empty verdict.md", + "PR 991: missing or empty summary.md", + "PR 991: missing or empty review.md", + "PR 991: missing or empty diff-analysis.md", + "PR 991: missing or empty interactions.md", + "PR 991: missing or empty trace.md", + "PR 991: no merge/revise/reject disposition near verdict start", + "PR 991: no non-empty test-evidence file", + "PR 992: missing or empty verdict.md", + "PR 992: missing or empty summary.md", + "PR 992: missing or empty review.md", + "PR 992: missing or empty diff-analysis.md", + "PR 992: missing or empty interactions.md", + "PR 992: missing or empty trace.md", + "PR 992: no merge/revise/reject disposition near verdict start", + "PR 992: no non-empty test-evidence file", + "PR 993: missing or empty verdict.md", + "PR 993: missing or empty summary.md", + "PR 993: missing or empty review.md", + "PR 993: missing or empty diff-analysis.md", + "PR 993: missing or empty interactions.md", + "PR 993: missing or empty trace.md", + "PR 993: no merge/revise/reject disposition near verdict start", + "PR 993: no non-empty test-evidence file", + "PR 994: missing or empty verdict.md", + "PR 994: missing or empty summary.md", + "PR 994: missing or empty review.md", + "PR 994: missing or empty diff-analysis.md", + "PR 994: missing or empty interactions.md", + "PR 994: missing or empty trace.md", + "PR 994: no merge/revise/reject disposition near verdict start", + "PR 994: no non-empty test-evidence file", + "PR 996: missing or empty verdict.md", + "PR 996: missing or empty summary.md", + "PR 996: missing or empty review.md", + "PR 996: missing or empty diff-analysis.md", + "PR 996: missing or empty interactions.md", + "PR 996: missing or empty trace.md", + "PR 996: no merge/revise/reject disposition near verdict start", + "PR 996: no non-empty test-evidence file", + "PR 997: missing or empty verdict.md", + "PR 997: missing or empty summary.md", + "PR 997: missing or empty review.md", + "PR 997: missing or empty diff-analysis.md", + "PR 997: missing or empty interactions.md", + "PR 997: missing or empty trace.md", + "PR 997: no merge/revise/reject disposition near verdict start", + "PR 997: no non-empty test-evidence file", + "PR 998: missing or empty verdict.md", + "PR 998: missing or empty summary.md", + "PR 998: missing or empty review.md", + "PR 998: missing or empty diff-analysis.md", + "PR 998: missing or empty interactions.md", + "PR 998: missing or empty trace.md", + "PR 998: no merge/revise/reject disposition near verdict start", + "PR 998: no non-empty test-evidence file", + "PR 999: missing or empty verdict.md", + "PR 999: missing or empty summary.md", + "PR 999: missing or empty review.md", + "PR 999: missing or empty diff-analysis.md", + "PR 999: missing or empty interactions.md", + "PR 999: missing or empty trace.md", + "PR 999: no merge/revise/reject disposition near verdict start", + "PR 999: no non-empty test-evidence file", + "PR 1000: missing or empty interactions.md", + "PR 1002: missing or empty interactions.md", + "PR 1003: missing or empty verdict.md", + "PR 1003: missing or empty summary.md", + "PR 1003: missing or empty review.md", + "PR 1003: missing or empty diff-analysis.md", + "PR 1003: missing or empty interactions.md", + "PR 1003: missing or empty trace.md", + "PR 1003: no merge/revise/reject disposition near verdict start", + "PR 1003: no non-empty test-evidence file", + "PR 1004: missing or empty verdict.md", + "PR 1004: missing or empty summary.md", + "PR 1004: missing or empty review.md", + "PR 1004: missing or empty diff-analysis.md", + "PR 1004: missing or empty interactions.md", + "PR 1004: missing or empty trace.md", + "PR 1004: no merge/revise/reject disposition near verdict start", + "PR 1004: no non-empty test-evidence file", + "PR 1005: missing or empty verdict.md", + "PR 1005: missing or empty summary.md", + "PR 1005: missing or empty review.md", + "PR 1005: missing or empty diff-analysis.md", + "PR 1005: missing or empty interactions.md", + "PR 1005: missing or empty trace.md", + "PR 1005: no merge/revise/reject disposition near verdict start", + "PR 1005: no non-empty test-evidence file", + "PR 1006: missing or empty verdict.md", + "PR 1006: missing or empty summary.md", + "PR 1006: missing or empty review.md", + "PR 1006: missing or empty diff-analysis.md", + "PR 1006: missing or empty interactions.md", + "PR 1006: missing or empty trace.md", + "PR 1006: no merge/revise/reject disposition near verdict start", + "PR 1006: no non-empty test-evidence file", + "PR 1007: missing or empty verdict.md", + "PR 1007: missing or empty summary.md", + "PR 1007: missing or empty review.md", + "PR 1007: missing or empty diff-analysis.md", + "PR 1007: missing or empty interactions.md", + "PR 1007: missing or empty trace.md", + "PR 1007: no merge/revise/reject disposition near verdict start", + "PR 1007: no non-empty test-evidence file", + "PR 1008: missing or empty verdict.md", + "PR 1008: missing or empty summary.md", + "PR 1008: missing or empty review.md", + "PR 1008: missing or empty diff-analysis.md", + "PR 1008: missing or empty interactions.md", + "PR 1008: missing or empty trace.md", + "PR 1008: no merge/revise/reject disposition near verdict start", + "PR 1008: no non-empty test-evidence file", + "PR 1009: missing or empty verdict.md", + "PR 1009: missing or empty summary.md", + "PR 1009: missing or empty review.md", + "PR 1009: missing or empty diff-analysis.md", + "PR 1009: missing or empty interactions.md", + "PR 1009: missing or empty trace.md", + "PR 1009: no merge/revise/reject disposition near verdict start", + "PR 1009: no non-empty test-evidence file", + "PR 1010: missing or empty verdict.md", + "PR 1010: missing or empty summary.md", + "PR 1010: missing or empty review.md", + "PR 1010: missing or empty diff-analysis.md", + "PR 1010: missing or empty interactions.md", + "PR 1010: missing or empty trace.md", + "PR 1010: no merge/revise/reject disposition near verdict start", + "PR 1010: no non-empty test-evidence file", + "PR 1011: missing or empty verdict.md", + "PR 1011: missing or empty summary.md", + "PR 1011: missing or empty review.md", + "PR 1011: missing or empty diff-analysis.md", + "PR 1011: missing or empty interactions.md", + "PR 1011: missing or empty trace.md", + "PR 1011: no merge/revise/reject disposition near verdict start", + "PR 1011: no non-empty test-evidence file", + "PR 1012: missing or empty verdict.md", + "PR 1012: missing or empty summary.md", + "PR 1012: missing or empty review.md", + "PR 1012: missing or empty diff-analysis.md", + "PR 1012: missing or empty interactions.md", + "PR 1012: missing or empty trace.md", + "PR 1012: no merge/revise/reject disposition near verdict start", + "PR 1012: no non-empty test-evidence file", + "PR 1013: missing or empty verdict.md", + "PR 1013: missing or empty summary.md", + "PR 1013: missing or empty review.md", + "PR 1013: missing or empty diff-analysis.md", + "PR 1013: missing or empty interactions.md", + "PR 1013: missing or empty trace.md", + "PR 1013: no merge/revise/reject disposition near verdict start", + "PR 1013: no non-empty test-evidence file", + "PR 1014: missing or empty verdict.md", + "PR 1014: missing or empty summary.md", + "PR 1014: missing or empty review.md", + "PR 1014: missing or empty diff-analysis.md", + "PR 1014: missing or empty interactions.md", + "PR 1014: missing or empty trace.md", + "PR 1014: no merge/revise/reject disposition near verdict start", + "PR 1014: no non-empty test-evidence file", + "PR 1015: missing or empty verdict.md", + "PR 1015: missing or empty summary.md", + "PR 1015: missing or empty review.md", + "PR 1015: missing or empty diff-analysis.md", + "PR 1015: missing or empty interactions.md", + "PR 1015: missing or empty trace.md", + "PR 1015: no merge/revise/reject disposition near verdict start", + "PR 1015: no non-empty test-evidence file", + "PR 1016: missing or empty interactions.md", + "PR 1017: missing or empty verdict.md", + "PR 1017: missing or empty summary.md", + "PR 1017: missing or empty review.md", + "PR 1017: missing or empty diff-analysis.md", + "PR 1017: missing or empty interactions.md", + "PR 1017: missing or empty trace.md", + "PR 1017: no merge/revise/reject disposition near verdict start", + "PR 1017: no non-empty test-evidence file", + "PR 1018: missing or empty verdict.md", + "PR 1018: missing or empty summary.md", + "PR 1018: missing or empty review.md", + "PR 1018: missing or empty diff-analysis.md", + "PR 1018: missing or empty interactions.md", + "PR 1018: missing or empty trace.md", + "PR 1018: no merge/revise/reject disposition near verdict start", + "PR 1018: no non-empty test-evidence file", + "PR 1019: missing or empty verdict.md", + "PR 1019: missing or empty summary.md", + "PR 1019: missing or empty review.md", + "PR 1019: missing or empty diff-analysis.md", + "PR 1019: missing or empty interactions.md", + "PR 1019: missing or empty trace.md", + "PR 1019: no merge/revise/reject disposition near verdict start", + "PR 1019: no non-empty test-evidence file", + "PR 1020: missing or empty verdict.md", + "PR 1020: missing or empty summary.md", + "PR 1020: missing or empty review.md", + "PR 1020: missing or empty diff-analysis.md", + "PR 1020: missing or empty interactions.md", + "PR 1020: missing or empty trace.md", + "PR 1020: no merge/revise/reject disposition near verdict start", + "PR 1020: no non-empty test-evidence file", + "PR 1021: missing or empty verdict.md", + "PR 1021: missing or empty summary.md", + "PR 1021: missing or empty review.md", + "PR 1021: missing or empty diff-analysis.md", + "PR 1021: missing or empty interactions.md", + "PR 1021: missing or empty trace.md", + "PR 1021: no merge/revise/reject disposition near verdict start", + "PR 1021: no non-empty test-evidence file", + "PR 1022: missing or empty verdict.md", + "PR 1022: missing or empty summary.md", + "PR 1022: missing or empty review.md", + "PR 1022: missing or empty diff-analysis.md", + "PR 1022: missing or empty interactions.md", + "PR 1022: missing or empty trace.md", + "PR 1022: no merge/revise/reject disposition near verdict start", + "PR 1022: no non-empty test-evidence file", + "PR 1023: missing or empty verdict.md", + "PR 1023: missing or empty summary.md", + "PR 1023: missing or empty review.md", + "PR 1023: missing or empty diff-analysis.md", + "PR 1023: missing or empty interactions.md", + "PR 1023: missing or empty trace.md", + "PR 1023: no merge/revise/reject disposition near verdict start", + "PR 1023: no non-empty test-evidence file", + "PR 1024: missing or empty verdict.md", + "PR 1024: missing or empty summary.md", + "PR 1024: missing or empty review.md", + "PR 1024: missing or empty diff-analysis.md", + "PR 1024: missing or empty interactions.md", + "PR 1024: missing or empty trace.md", + "PR 1024: no merge/revise/reject disposition near verdict start", + "PR 1024: no non-empty test-evidence file", + "PR 1025: missing or empty verdict.md", + "PR 1025: missing or empty summary.md", + "PR 1025: missing or empty review.md", + "PR 1025: missing or empty diff-analysis.md", + "PR 1025: missing or empty interactions.md", + "PR 1025: missing or empty trace.md", + "PR 1025: no merge/revise/reject disposition near verdict start", + "PR 1025: no non-empty test-evidence file", + "PR 1026: missing or empty verdict.md", + "PR 1026: missing or empty summary.md", + "PR 1026: missing or empty review.md", + "PR 1026: missing or empty diff-analysis.md", + "PR 1026: missing or empty interactions.md", + "PR 1026: missing or empty trace.md", + "PR 1026: no merge/revise/reject disposition near verdict start", + "PR 1026: no non-empty test-evidence file", + "PR 1027: missing or empty verdict.md", + "PR 1027: missing or empty summary.md", + "PR 1027: missing or empty review.md", + "PR 1027: missing or empty diff-analysis.md", + "PR 1027: missing or empty interactions.md", + "PR 1027: missing or empty trace.md", + "PR 1027: no merge/revise/reject disposition near verdict start", + "PR 1027: no non-empty test-evidence file", + "PR 1028: missing or empty verdict.md", + "PR 1028: missing or empty summary.md", + "PR 1028: missing or empty review.md", + "PR 1028: missing or empty diff-analysis.md", + "PR 1028: missing or empty interactions.md", + "PR 1028: missing or empty trace.md", + "PR 1028: no merge/revise/reject disposition near verdict start", + "PR 1028: no non-empty test-evidence file", + "PR 1029: missing or empty interactions.md", + "PR 1030: missing or empty verdict.md", + "PR 1030: missing or empty summary.md", + "PR 1030: missing or empty review.md", + "PR 1030: missing or empty diff-analysis.md", + "PR 1030: missing or empty interactions.md", + "PR 1030: missing or empty trace.md", + "PR 1030: no merge/revise/reject disposition near verdict start", + "PR 1030: no non-empty test-evidence file", + "PR 1032: missing or empty verdict.md", + "PR 1032: missing or empty summary.md", + "PR 1032: missing or empty review.md", + "PR 1032: missing or empty diff-analysis.md", + "PR 1032: missing or empty interactions.md", + "PR 1032: missing or empty trace.md", + "PR 1032: no merge/revise/reject disposition near verdict start", + "PR 1032: no non-empty test-evidence file", + "PR 1033: missing or empty verdict.md", + "PR 1033: missing or empty summary.md", + "PR 1033: missing or empty review.md", + "PR 1033: missing or empty diff-analysis.md", + "PR 1033: missing or empty interactions.md", + "PR 1033: missing or empty trace.md", + "PR 1033: no merge/revise/reject disposition near verdict start", + "PR 1033: no non-empty test-evidence file", + "PR 1034: missing or empty verdict.md", + "PR 1034: missing or empty summary.md", + "PR 1034: missing or empty review.md", + "PR 1034: missing or empty diff-analysis.md", + "PR 1034: missing or empty interactions.md", + "PR 1034: missing or empty trace.md", + "PR 1034: no merge/revise/reject disposition near verdict start", + "PR 1034: no non-empty test-evidence file", + "PR 1035: missing or empty verdict.md", + "PR 1035: missing or empty summary.md", + "PR 1035: missing or empty review.md", + "PR 1035: missing or empty diff-analysis.md", + "PR 1035: missing or empty interactions.md", + "PR 1035: missing or empty trace.md", + "PR 1035: no merge/revise/reject disposition near verdict start", + "PR 1035: no non-empty test-evidence file", + "PR 1036: missing or empty verdict.md", + "PR 1036: missing or empty summary.md", + "PR 1036: missing or empty review.md", + "PR 1036: missing or empty diff-analysis.md", + "PR 1036: missing or empty interactions.md", + "PR 1036: missing or empty trace.md", + "PR 1036: no merge/revise/reject disposition near verdict start", + "PR 1036: no non-empty test-evidence file", + "PR 1037: missing or empty verdict.md", + "PR 1037: missing or empty summary.md", + "PR 1037: missing or empty review.md", + "PR 1037: missing or empty diff-analysis.md", + "PR 1037: missing or empty interactions.md", + "PR 1037: missing or empty trace.md", + "PR 1037: no merge/revise/reject disposition near verdict start", + "PR 1037: no non-empty test-evidence file", + "PR 1038: missing or empty verdict.md", + "PR 1038: missing or empty summary.md", + "PR 1038: missing or empty review.md", + "PR 1038: missing or empty diff-analysis.md", + "PR 1038: missing or empty interactions.md", + "PR 1038: missing or empty trace.md", + "PR 1038: no merge/revise/reject disposition near verdict start", + "PR 1038: no non-empty test-evidence file", + "PR 1039: missing or empty verdict.md", + "PR 1039: missing or empty summary.md", + "PR 1039: missing or empty review.md", + "PR 1039: missing or empty diff-analysis.md", + "PR 1039: missing or empty interactions.md", + "PR 1039: missing or empty trace.md", + "PR 1039: no merge/revise/reject disposition near verdict start", + "PR 1039: no non-empty test-evidence file", + "PR 1040: missing or empty verdict.md", + "PR 1040: missing or empty summary.md", + "PR 1040: missing or empty review.md", + "PR 1040: missing or empty diff-analysis.md", + "PR 1040: missing or empty interactions.md", + "PR 1040: missing or empty trace.md", + "PR 1040: no merge/revise/reject disposition near verdict start", + "PR 1040: no non-empty test-evidence file", + "PR 1042: missing or empty verdict.md", + "PR 1042: missing or empty summary.md", + "PR 1042: missing or empty review.md", + "PR 1042: missing or empty diff-analysis.md", + "PR 1042: missing or empty interactions.md", + "PR 1042: missing or empty trace.md", + "PR 1042: no merge/revise/reject disposition near verdict start", + "PR 1042: no non-empty test-evidence file", + "PR 1043: missing or empty verdict.md", + "PR 1043: missing or empty summary.md", + "PR 1043: missing or empty review.md", + "PR 1043: missing or empty diff-analysis.md", + "PR 1043: missing or empty interactions.md", + "PR 1043: missing or empty trace.md", + "PR 1043: no merge/revise/reject disposition near verdict start", + "PR 1043: no non-empty test-evidence file", + "PR 1044: missing or empty verdict.md", + "PR 1044: missing or empty summary.md", + "PR 1044: missing or empty review.md", + "PR 1044: missing or empty diff-analysis.md", + "PR 1044: missing or empty interactions.md", + "PR 1044: missing or empty trace.md", + "PR 1044: no merge/revise/reject disposition near verdict start", + "PR 1044: no non-empty test-evidence file", + "PR 1045: missing or empty verdict.md", + "PR 1045: missing or empty summary.md", + "PR 1045: missing or empty review.md", + "PR 1045: missing or empty diff-analysis.md", + "PR 1045: missing or empty interactions.md", + "PR 1045: missing or empty trace.md", + "PR 1045: no merge/revise/reject disposition near verdict start", + "PR 1045: no non-empty test-evidence file", + "PR 1046: missing or empty verdict.md", + "PR 1046: missing or empty summary.md", + "PR 1046: missing or empty review.md", + "PR 1046: missing or empty diff-analysis.md", + "PR 1046: missing or empty interactions.md", + "PR 1046: missing or empty trace.md", + "PR 1046: no merge/revise/reject disposition near verdict start", + "PR 1046: no non-empty test-evidence file", + "PR 1047: missing or empty verdict.md", + "PR 1047: missing or empty summary.md", + "PR 1047: missing or empty review.md", + "PR 1047: missing or empty diff-analysis.md", + "PR 1047: missing or empty interactions.md", + "PR 1047: missing or empty trace.md", + "PR 1047: no merge/revise/reject disposition near verdict start", + "PR 1047: no non-empty test-evidence file", + "PR 1048: missing or empty verdict.md", + "PR 1048: missing or empty summary.md", + "PR 1048: missing or empty review.md", + "PR 1048: missing or empty diff-analysis.md", + "PR 1048: missing or empty interactions.md", + "PR 1048: missing or empty trace.md", + "PR 1048: no merge/revise/reject disposition near verdict start", + "PR 1048: no non-empty test-evidence file", + "PR 1049: missing or empty verdict.md", + "PR 1049: missing or empty summary.md", + "PR 1049: missing or empty review.md", + "PR 1049: missing or empty diff-analysis.md", + "PR 1049: missing or empty interactions.md", + "PR 1049: missing or empty trace.md", + "PR 1049: no merge/revise/reject disposition near verdict start", + "PR 1049: no non-empty test-evidence file", + "PR 1050: missing or empty verdict.md", + "PR 1050: missing or empty summary.md", + "PR 1050: missing or empty review.md", + "PR 1050: missing or empty diff-analysis.md", + "PR 1050: missing or empty interactions.md", + "PR 1050: missing or empty trace.md", + "PR 1050: no merge/revise/reject disposition near verdict start", + "PR 1050: no non-empty test-evidence file", + "PR 1051: missing or empty verdict.md", + "PR 1051: missing or empty summary.md", + "PR 1051: missing or empty review.md", + "PR 1051: missing or empty diff-analysis.md", + "PR 1051: missing or empty interactions.md", + "PR 1051: missing or empty trace.md", + "PR 1051: no merge/revise/reject disposition near verdict start", + "PR 1051: no non-empty test-evidence file", + "PR 1052: missing or empty verdict.md", + "PR 1052: missing or empty summary.md", + "PR 1052: missing or empty review.md", + "PR 1052: missing or empty diff-analysis.md", + "PR 1052: missing or empty interactions.md", + "PR 1052: missing or empty trace.md", + "PR 1052: no merge/revise/reject disposition near verdict start", + "PR 1052: no non-empty test-evidence file", + "PR 1053: missing or empty verdict.md", + "PR 1053: missing or empty summary.md", + "PR 1053: missing or empty review.md", + "PR 1053: missing or empty diff-analysis.md", + "PR 1053: missing or empty interactions.md", + "PR 1053: missing or empty trace.md", + "PR 1053: no merge/revise/reject disposition near verdict start", + "PR 1053: no non-empty test-evidence file", + "PR 1054: missing or empty verdict.md", + "PR 1054: missing or empty summary.md", + "PR 1054: missing or empty review.md", + "PR 1054: missing or empty diff-analysis.md", + "PR 1054: missing or empty interactions.md", + "PR 1054: missing or empty trace.md", + "PR 1054: no merge/revise/reject disposition near verdict start", + "PR 1054: no non-empty test-evidence file", + "PR 1055: missing or empty verdict.md", + "PR 1055: missing or empty summary.md", + "PR 1055: missing or empty review.md", + "PR 1055: missing or empty diff-analysis.md", + "PR 1055: missing or empty interactions.md", + "PR 1055: missing or empty trace.md", + "PR 1055: no merge/revise/reject disposition near verdict start", + "PR 1055: no non-empty test-evidence file", + "PR 1056: missing or empty interactions.md", + "PR 1057: missing or empty interactions.md", + "parsed verdict total is 6, expected 75", + "PRs without a source-file or frozen-diff trace pointer: [351, 364, 716, 750, 959, 961, 966, 973, 974, 983, 988, 990, 991, 992, 993, 994, 996, 997, 998, 999, 1003, 1004, 1005, 1006, 1007, 1008, 1009, 1010, 1011, 1012, 1013, 1014, 1015, 1017, 1018, 1019, 1020, 1021, 1022, 1023, 1024, 1025, 1026, 1027, 1028, 1030, 1032, 1033, 1034, 1035, 1036, 1037, 1038, 1039, 1040, 1042, 1043, 1044, 1045, 1046, 1047, 1048, 1049, 1050, 1051, 1052, 1053, 1054, 1055]", + "PRs whose review has no source-file/diff citation: [351, 364, 716, 750, 959, 961, 966, 973, 974, 983, 988, 990, 991, 992, 993, 994, 996, 997, 998, 999, 1003, 1004, 1005, 1006, 1007, 1008, 1009, 1010, 1011, 1012, 1013, 1014, 1015, 1017, 1018, 1019, 1020, 1021, 1022, 1023, 1024, 1025, 1026, 1027, 1028, 1030, 1032, 1033, 1034, 1035, 1036, 1037, 1038, 1039, 1040, 1042, 1043, 1044, 1045, 1046, 1047, 1048, 1049, 1050, 1051, 1052, 1053, 1054, 1055]", + "PRs without an explicit bounty-tier declaration: [351, 364, 716, 750, 959, 961, 966, 973, 974, 983, 988, 990, 991, 992, 993, 994, 996, 997, 998, 999, 1003, 1004, 1005, 1006, 1007, 1008, 1009, 1010, 1011, 1012, 1013, 1014, 1015, 1017, 1018, 1019, 1020, 1021, 1022, 1023, 1024, 1025, 1026, 1027, 1028, 1030, 1032, 1033, 1034, 1035, 1036, 1037, 1038, 1039, 1040, 1042, 1043, 1044, 1045, 1046, 1047, 1048, 1049, 1050, 1051, 1052, 1053, 1054, 1055, 1056]", + "PRs with neither execution nor explicit skip evidence: [1000, 1002, 1016, 1029, 1057]", + "risk_matrix omits PRs: [351, 364, 716, 750, 959, 961, 966, 973, 974, 983, 988, 990, 991, 992, 993, 994, 996, 997, 998, 999, 1002, 1003, 1004, 1005, 1006, 1007, 1008, 1009, 1010, 1011, 1012, 1013, 1014, 1015, 1017, 1018, 1019, 1020, 1021, 1022, 1023, 1024, 1025, 1026, 1027, 1028, 1030, 1032, 1033, 1034, 1035, 1036, 1037, 1038, 1039, 1040, 1042, 1043, 1044, 1045, 1046, 1047, 1048, 1049, 1050, 1051, 1052, 1053, 1054, 1055]", + "dependency analysis omits 205/205 high-confidence file-overlap pairs", + "frozen files.txt counts disagree with metadata for PRs, but the audit does not document the baseline/reverse-drift limitation", + "tool log records 0/41 transcript tool events" + ], + "warnings": [ + "dependency analysis does not enumerate 676/676 raw frozen-file pairs; raw lists include known reverse-drift, so this is reported separately and is not itself a semantic failure" + ] + }, + "telemetry": { + "schema_version": 1, + "run_name": "gemma4-31b-q4_75pr_v1", + "source_csv": "/home/michael/gemma4-campaign-state/telemetry/gemma4-31b-q4-gpu.csv", + "generated_at": "2026-08-02T06:45:35.340414+00:00", + "analyzer": { + "path": "/home/michael/bench-gemma4-31b-q4/tooling/deployments/gemma4-31b-q4-tower2/analyze_replica_telemetry.py", + "file_sha256": "19a2d7a9cdfc36d2fcd9840581015410b5451aacff43e2328e050c1214e771bd", + "git_sha": "17e0128cb8c3268d4616a51147b1b754db8103b6" + }, + "attribution": { + "method": "live harness --port maps port 8000 to GPU 0 and port 8001 to GPU 1", + "active_gpu_ids_observed": [ + "1" + ], + "other_gpu_work_is_reported_as_concurrency_not_charged_to_active_gpu_energy": true, + "cpu_package_power_is_shared_host_context_and_not_uniquely_attributable_when_runs_overlap": true + }, + "sampling": { + "samples": 100, + "paired_host_samples": 100, + "first_sample": "2026-08-02T06:11:26+00:00", + "last_sample": "2026-08-02T06:19:41+00:00", + "window_source": "summary.json", + "window_started_at": "2026-08-02T06:11:25.302821+00:00", + "window_ended_at": "2026-08-02T06:19:41.483333+00:00", + "window_wall_s": 496.2, + "integrated_coverage_s": 495.0, + "coverage_fraction_of_wall": 0.9976, + "max_accepted_gap_s": 15.0 + }, + "active_gpu": { + "mean_power_w": 478.13, + "time_weighted_power_w": 480.23, + "p90_power_w": 500.31, + "max_power_w": 507.12, + "mean_sm_util_pct": 91.13, + "p90_sm_util_pct": 99.0, + "max_memory_used_mib": 66699.0, + "max_temp_c": 82.0, + "mean_sm_clock_mhz": 2594.4, + "configured_cap_w": 500.0 + }, + "cpu_package_shared_context": { + "samples": 100, + "mean_power_w": 132.51, + "max_power_w": 141.34, + "integrated_coverage_s": 495.0 + }, + "concurrency": { + "other_gpu_cell_sample_counts": { + "(idle/no benchmark harness)": 18, + "gemma4-31b-q4_75pr_v2": 39, + "gemma4-31b-q4_board_pres_v3": 43 + }, + "both_gpus_over_20pct_sm_fraction": 0.69, + "mean_observed_two_gpu_plus_cpu_package_w": 1009.16, + "max_observed_two_gpu_plus_cpu_package_w": 1144.89, + "wall_power_note": "This is GPU plus CPU package telemetry, not AC wall draw; no software wall meter is available." + }, + "energy_and_cost": { + "active_gpu_sampled_kwh": 0.066031, + "active_gpu_sampled_cost_usd": 0.008584, + "cpu_package_shared_sampled_kwh": 0.018243, + "active_gpu_cap_upper_bound_kwh": 0.068917, + "rate_usd_per_kwh": 0.13 + } + }, + "evidence": { + "summary_sha256": "bcd3d259a03c267301caa15b59f8d1940887d8b39c4d686f2492a280195dd45f", + "receipt_sha256": "e391c2ae8b53e64597c7274da1c5a2407df958229b77d4618a3dbbd8746898ab", + "transcript_sha256": "fedf67f0ef5f5faac2b2ee92057d9f4c0b67c506dfb2c08807190994af1dc6eb", + "validator_sha256": "11479b273a7c7c4b7597736ccc6eb8b763443192c8f6c1e6a9b7067ec0dded83", + "task_sha256": "0d426e15303b3f886a75db81165cdbe9d91534a474217f91f4e2baf41c0056a6", + "fixture_path": "/mnt/bulk/benchmark-fixtures/dreamserver-75-pr-audit-2026-04-27" + } +} diff --git a/benchmarks/gemma4-31b-q4/75pr-v2-audit.json b/benchmarks/gemma4-31b-q4/75pr-v2-audit.json new file mode 100644 index 00000000..0b94db04 --- /dev/null +++ b/benchmarks/gemma4-31b-q4/75pr-v2-audit.json @@ -0,0 +1,1265 @@ +{ + "schema_version": 2, + "audited_at": "2026-08-02T06:46:32.358185+00:00", + "run_name": "Gemma-4-31B-it-QAT-Q4_0 gemma4-31b-q4_75pr_v2", + "classification": "MODEL_TERMINAL_FAILURE", + "legacy_shipped": false, + "verdict_counts": { + "merge": 1 + }, + "runtime": { + "wall_s": 193.8, + "iterations": 32, + "completion_tokens": 7592, + "prompt_tokens": 879716, + "completion_tps_model_call": 40.8 + }, + "artifact_integrity": { + "archive_sha256": "f9a10fe0fa04125aa230dd9e588b5763544b3fa3ebcba9666e6d78bc9e427fd9", + "archive_bytes": 54035493, + "archive_gzip_valid": true, + "head": "08578de7132caff2f4330b5eef3d5dcd90f007dc", + "tags": [], + "tag_targets": {}, + "tag_at_head": false, + "clean": true, + "commit_count": 2 + }, + "substance": { + "test_evidence_count": 0, + "test_evidence_prs": [], + "missing_test_or_skip_count": 75, + "missing_test_or_skip_prs": [ + 959, + 966 + ], + "reviews_under_800_bytes_count": 2, + "reviews_under_800_bytes_prs": [ + 959, + 966 + ], + "reviews_with_source_citation_count": 1, + "reviews_with_source_citation_prs": [ + 959 + ], + "reviews_with_hunk_marker_count": 0, + "traces_with_source_citation_count": 0, + "traces_with_hunk_marker_count": 0, + "diff_analyses_with_hunk_marker_count": 0, + "explicit_actual_bounty_tier_count": 0, + "explicit_actual_bounty_tier_prs": [], + "transcript_tool_events": 31, + "numbered_tool_log_entries": 0 + }, + "strict_validation": { + "schema_version": 1, + "workspace": "/home/michael/bench-gemma4-31b-q4/tooling/workspace/gemma4-31b-q4_75pr_v2/audit-repo", + "fixture": "/mnt/bulk/benchmark-fixtures/dreamserver-75-pr-audit-2026-04-27", + "pass": false, + "metrics": { + "canonical_prs": 75, + "actual_pr_dirs": 2, + "missing_pr_dirs": [ + 351, + 364, + 716, + 750, + 961, + 973, + 974, + 983, + 988, + 990, + 991, + 992, + 993, + 994, + 996, + 997, + 998, + 999, + 1000, + 1002, + 1003, + 1004, + 1005, + 1006, + 1007, + 1008, + 1009, + 1010, + 1011, + 1012, + 1013, + 1014, + 1015, + 1016, + 1017, + 1018, + 1019, + 1020, + 1021, + 1022, + 1023, + 1024, + 1025, + 1026, + 1027, + 1028, + 1029, + 1030, + 1032, + 1033, + 1034, + 1035, + 1036, + 1037, + 1038, + 1039, + 1040, + 1042, + 1043, + 1044, + 1045, + 1046, + 1047, + 1048, + 1049, + 1050, + 1051, + 1052, + 1053, + 1054, + 1055, + 1056, + 1057 + ], + "extra_pr_dirs": [], + "verdict_counts": { + "merge": 1 + }, + "raw_frozen_file_overlap_pairs": 676, + "trusted_file_list_prs": 63, + "trusted_file_overlap_pairs": 205, + "listed_pair_mentions": 0, + "missing_trusted_file_overlap_pairs": 205, + "missing_trusted_file_overlap_pair_sample": [ + [ + 990, + 991 + ], + [ + 992, + 1013 + ], + [ + 992, + 1017 + ], + [ + 993, + 994 + ], + [ + 993, + 997 + ], + [ + 993, + 998 + ], + [ + 993, + 999 + ], + [ + 993, + 1000 + ], + [ + 993, + 1002 + ], + [ + 993, + 1006 + ], + [ + 993, + 1007 + ], + [ + 993, + 1008 + ], + [ + 993, + 1011 + ], + [ + 993, + 1016 + ], + [ + 993, + 1018 + ], + [ + 993, + 1020 + ], + [ + 994, + 997 + ], + [ + 994, + 998 + ], + [ + 994, + 999 + ], + [ + 994, + 1000 + ], + [ + 994, + 1002 + ], + [ + 994, + 1006 + ], + [ + 994, + 1007 + ], + [ + 994, + 1008 + ], + [ + 994, + 1010 + ], + [ + 994, + 1011 + ], + [ + 994, + 1016 + ], + [ + 994, + 1017 + ], + [ + 994, + 1018 + ], + [ + 994, + 1020 + ], + [ + 996, + 1012 + ], + [ + 996, + 1017 + ], + [ + 996, + 1026 + ], + [ + 997, + 998 + ], + [ + 997, + 999 + ], + [ + 997, + 1000 + ], + [ + 997, + 1002 + ], + [ + 997, + 1006 + ], + [ + 997, + 1007 + ], + [ + 997, + 1008 + ], + [ + 997, + 1011 + ], + [ + 997, + 1016 + ], + [ + 997, + 1018 + ], + [ + 997, + 1020 + ], + [ + 998, + 999 + ], + [ + 998, + 1000 + ], + [ + 998, + 1002 + ], + [ + 998, + 1006 + ], + [ + 998, + 1007 + ], + [ + 998, + 1008 + ] + ], + "missing_raw_frozen_overlap_pairs": 676, + "missing_raw_frozen_overlap_pair_sample": [ + [ + 351, + 364 + ], + [ + 351, + 716 + ], + [ + 351, + 750 + ], + [ + 351, + 959 + ], + [ + 351, + 961 + ], + [ + 351, + 966 + ], + [ + 351, + 973 + ], + [ + 351, + 974 + ], + [ + 351, + 983 + ], + [ + 351, + 988 + ], + [ + 351, + 990 + ], + [ + 351, + 991 + ], + [ + 351, + 992 + ], + [ + 351, + 993 + ], + [ + 351, + 994 + ], + [ + 351, + 996 + ], + [ + 351, + 997 + ], + [ + 351, + 998 + ], + [ + 351, + 999 + ], + [ + 351, + 1000 + ], + [ + 351, + 1002 + ], + [ + 351, + 1003 + ], + [ + 351, + 1004 + ], + [ + 351, + 1005 + ], + [ + 351, + 1006 + ], + [ + 351, + 1007 + ], + [ + 351, + 1008 + ], + [ + 351, + 1009 + ], + [ + 351, + 1010 + ], + [ + 351, + 1011 + ], + [ + 351, + 1012 + ], + [ + 351, + 1013 + ], + [ + 351, + 1015 + ], + [ + 351, + 1016 + ], + [ + 351, + 1017 + ], + [ + 351, + 1018 + ], + [ + 351, + 1019 + ], + [ + 351, + 1020 + ], + [ + 351, + 1021 + ], + [ + 351, + 1022 + ], + [ + 351, + 1023 + ], + [ + 351, + 1024 + ], + [ + 351, + 1025 + ], + [ + 351, + 1026 + ], + [ + 351, + 1027 + ], + [ + 351, + 1028 + ], + [ + 351, + 1029 + ], + [ + 351, + 1030 + ], + [ + 351, + 1032 + ], + [ + 351, + 1033 + ] + ], + "raw_file_count_metadata_mismatch_prs": [ + 351, + 364, + 716, + 750, + 959, + 961, + 966, + 973, + 974, + 983, + 988, + 1043 + ], + "transcript_tool_events": 31, + "numbered_tool_log_entries": 0 + }, + "failures": [ + "missing or empty required artifact: README.md", + "missing or empty required artifact: sources.md", + "missing or empty required artifact: report/executive-summary.md", + "missing or empty required artifact: report/backlog-strategy.md", + "missing or empty required artifact: report/contributor-notes.md", + "missing or empty required artifact: report/project-health.md", + "missing or empty required artifact: analysis/dependency-graph.md", + "missing or empty required artifact: analysis/risk-matrix.md", + "missing or empty required artifact: analysis/surface-area.md", + "missing or empty required artifact: testing/baseline.md", + "missing or empty required artifact: research/questions.md", + "missing or empty required artifact: research/dead-ends.md", + "missing or empty required artifact: research/upstream-context.md", + "missing non-empty decisions directory", + "missing PR directories: [351, 364, 716, 750, 961, 973, 974, 983, 988, 990, 991, 992, 993, 994, 996, 997, 998, 999, 1000, 1002, 1003, 1004, 1005, 1006, 1007, 1008, 1009, 1010, 1011, 1012, 1013, 1014, 1015, 1016, 1017, 1018, 1019, 1020, 1021, 1022, 1023, 1024, 1025, 1026, 1027, 1028, 1029, 1030, 1032, 1033, 1034, 1035, 1036, 1037, 1038, 1039, 1040, 1042, 1043, 1044, 1045, 1046, 1047, 1048, 1049, 1050, 1051, 1052, 1053, 1054, 1055, 1056, 1057]", + "PR 351: missing or empty verdict.md", + "PR 351: missing or empty summary.md", + "PR 351: missing or empty review.md", + "PR 351: missing or empty diff-analysis.md", + "PR 351: missing or empty interactions.md", + "PR 351: missing or empty trace.md", + "PR 351: no merge/revise/reject disposition near verdict start", + "PR 351: no non-empty test-evidence file", + "PR 364: missing or empty verdict.md", + "PR 364: missing or empty summary.md", + "PR 364: missing or empty review.md", + "PR 364: missing or empty diff-analysis.md", + "PR 364: missing or empty interactions.md", + "PR 364: missing or empty trace.md", + "PR 364: no merge/revise/reject disposition near verdict start", + "PR 364: no non-empty test-evidence file", + "PR 716: missing or empty verdict.md", + "PR 716: missing or empty summary.md", + "PR 716: missing or empty review.md", + "PR 716: missing or empty diff-analysis.md", + "PR 716: missing or empty interactions.md", + "PR 716: missing or empty trace.md", + "PR 716: no merge/revise/reject disposition near verdict start", + "PR 716: no non-empty test-evidence file", + "PR 750: missing or empty verdict.md", + "PR 750: missing or empty summary.md", + "PR 750: missing or empty review.md", + "PR 750: missing or empty diff-analysis.md", + "PR 750: missing or empty interactions.md", + "PR 750: missing or empty trace.md", + "PR 750: no merge/revise/reject disposition near verdict start", + "PR 750: no non-empty test-evidence file", + "PR 959: missing or empty interactions.md", + "PR 959: missing or empty trace.md", + "PR 959: no non-empty test-evidence file", + "PR 961: missing or empty verdict.md", + "PR 961: missing or empty summary.md", + "PR 961: missing or empty review.md", + "PR 961: missing or empty diff-analysis.md", + "PR 961: missing or empty interactions.md", + "PR 961: missing or empty trace.md", + "PR 961: no merge/revise/reject disposition near verdict start", + "PR 961: no non-empty test-evidence file", + "PR 966: missing or empty verdict.md", + "PR 966: missing or empty summary.md", + "PR 966: missing or empty review.md", + "PR 966: missing or empty diff-analysis.md", + "PR 966: missing or empty interactions.md", + "PR 966: missing or empty trace.md", + "PR 966: no merge/revise/reject disposition near verdict start", + "PR 966: no non-empty test-evidence file", + "PR 973: missing or empty verdict.md", + "PR 973: missing or empty summary.md", + "PR 973: missing or empty review.md", + "PR 973: missing or empty diff-analysis.md", + "PR 973: missing or empty interactions.md", + "PR 973: missing or empty trace.md", + "PR 973: no merge/revise/reject disposition near verdict start", + "PR 973: no non-empty test-evidence file", + "PR 974: missing or empty verdict.md", + "PR 974: missing or empty summary.md", + "PR 974: missing or empty review.md", + "PR 974: missing or empty diff-analysis.md", + "PR 974: missing or empty interactions.md", + "PR 974: missing or empty trace.md", + "PR 974: no merge/revise/reject disposition near verdict start", + "PR 974: no non-empty test-evidence file", + "PR 983: missing or empty verdict.md", + "PR 983: missing or empty summary.md", + "PR 983: missing or empty review.md", + "PR 983: missing or empty diff-analysis.md", + "PR 983: missing or empty interactions.md", + "PR 983: missing or empty trace.md", + "PR 983: no merge/revise/reject disposition near verdict start", + "PR 983: no non-empty test-evidence file", + "PR 988: missing or empty verdict.md", + "PR 988: missing or empty summary.md", + "PR 988: missing or empty review.md", + "PR 988: missing or empty diff-analysis.md", + "PR 988: missing or empty interactions.md", + "PR 988: missing or empty trace.md", + "PR 988: no merge/revise/reject disposition near verdict start", + "PR 988: no non-empty test-evidence file", + "PR 990: missing or empty verdict.md", + "PR 990: missing or empty summary.md", + "PR 990: missing or empty review.md", + "PR 990: missing or empty diff-analysis.md", + "PR 990: missing or empty interactions.md", + "PR 990: missing or empty trace.md", + "PR 990: no merge/revise/reject disposition near verdict start", + "PR 990: no non-empty test-evidence file", + "PR 991: missing or empty verdict.md", + "PR 991: missing or empty summary.md", + "PR 991: missing or empty review.md", + "PR 991: missing or empty diff-analysis.md", + "PR 991: missing or empty interactions.md", + "PR 991: missing or empty trace.md", + "PR 991: no merge/revise/reject disposition near verdict start", + "PR 991: no non-empty test-evidence file", + "PR 992: missing or empty verdict.md", + "PR 992: missing or empty summary.md", + "PR 992: missing or empty review.md", + "PR 992: missing or empty diff-analysis.md", + "PR 992: missing or empty interactions.md", + "PR 992: missing or empty trace.md", + "PR 992: no merge/revise/reject disposition near verdict start", + "PR 992: no non-empty test-evidence file", + "PR 993: missing or empty verdict.md", + "PR 993: missing or empty summary.md", + "PR 993: missing or empty review.md", + "PR 993: missing or empty diff-analysis.md", + "PR 993: missing or empty interactions.md", + "PR 993: missing or empty trace.md", + "PR 993: no merge/revise/reject disposition near verdict start", + "PR 993: no non-empty test-evidence file", + "PR 994: missing or empty verdict.md", + "PR 994: missing or empty summary.md", + "PR 994: missing or empty review.md", + "PR 994: missing or empty diff-analysis.md", + "PR 994: missing or empty interactions.md", + "PR 994: missing or empty trace.md", + "PR 994: no merge/revise/reject disposition near verdict start", + "PR 994: no non-empty test-evidence file", + "PR 996: missing or empty verdict.md", + "PR 996: missing or empty summary.md", + "PR 996: missing or empty review.md", + "PR 996: missing or empty diff-analysis.md", + "PR 996: missing or empty interactions.md", + "PR 996: missing or empty trace.md", + "PR 996: no merge/revise/reject disposition near verdict start", + "PR 996: no non-empty test-evidence file", + "PR 997: missing or empty verdict.md", + "PR 997: missing or empty summary.md", + "PR 997: missing or empty review.md", + "PR 997: missing or empty diff-analysis.md", + "PR 997: missing or empty interactions.md", + "PR 997: missing or empty trace.md", + "PR 997: no merge/revise/reject disposition near verdict start", + "PR 997: no non-empty test-evidence file", + "PR 998: missing or empty verdict.md", + "PR 998: missing or empty summary.md", + "PR 998: missing or empty review.md", + "PR 998: missing or empty diff-analysis.md", + "PR 998: missing or empty interactions.md", + "PR 998: missing or empty trace.md", + "PR 998: no merge/revise/reject disposition near verdict start", + "PR 998: no non-empty test-evidence file", + "PR 999: missing or empty verdict.md", + "PR 999: missing or empty summary.md", + "PR 999: missing or empty review.md", + "PR 999: missing or empty diff-analysis.md", + "PR 999: missing or empty interactions.md", + "PR 999: missing or empty trace.md", + "PR 999: no merge/revise/reject disposition near verdict start", + "PR 999: no non-empty test-evidence file", + "PR 1000: missing or empty verdict.md", + "PR 1000: missing or empty summary.md", + "PR 1000: missing or empty review.md", + "PR 1000: missing or empty diff-analysis.md", + "PR 1000: missing or empty interactions.md", + "PR 1000: missing or empty trace.md", + "PR 1000: no merge/revise/reject disposition near verdict start", + "PR 1000: no non-empty test-evidence file", + "PR 1002: missing or empty verdict.md", + "PR 1002: missing or empty summary.md", + "PR 1002: missing or empty review.md", + "PR 1002: missing or empty diff-analysis.md", + "PR 1002: missing or empty interactions.md", + "PR 1002: missing or empty trace.md", + "PR 1002: no merge/revise/reject disposition near verdict start", + "PR 1002: no non-empty test-evidence file", + "PR 1003: missing or empty verdict.md", + "PR 1003: missing or empty summary.md", + "PR 1003: missing or empty review.md", + "PR 1003: missing or empty diff-analysis.md", + "PR 1003: missing or empty interactions.md", + "PR 1003: missing or empty trace.md", + "PR 1003: no merge/revise/reject disposition near verdict start", + "PR 1003: no non-empty test-evidence file", + "PR 1004: missing or empty verdict.md", + "PR 1004: missing or empty summary.md", + "PR 1004: missing or empty review.md", + "PR 1004: missing or empty diff-analysis.md", + "PR 1004: missing or empty interactions.md", + "PR 1004: missing or empty trace.md", + "PR 1004: no merge/revise/reject disposition near verdict start", + "PR 1004: no non-empty test-evidence file", + "PR 1005: missing or empty verdict.md", + "PR 1005: missing or empty summary.md", + "PR 1005: missing or empty review.md", + "PR 1005: missing or empty diff-analysis.md", + "PR 1005: missing or empty interactions.md", + "PR 1005: missing or empty trace.md", + "PR 1005: no merge/revise/reject disposition near verdict start", + "PR 1005: no non-empty test-evidence file", + "PR 1006: missing or empty verdict.md", + "PR 1006: missing or empty summary.md", + "PR 1006: missing or empty review.md", + "PR 1006: missing or empty diff-analysis.md", + "PR 1006: missing or empty interactions.md", + "PR 1006: missing or empty trace.md", + "PR 1006: no merge/revise/reject disposition near verdict start", + "PR 1006: no non-empty test-evidence file", + "PR 1007: missing or empty verdict.md", + "PR 1007: missing or empty summary.md", + "PR 1007: missing or empty review.md", + "PR 1007: missing or empty diff-analysis.md", + "PR 1007: missing or empty interactions.md", + "PR 1007: missing or empty trace.md", + "PR 1007: no merge/revise/reject disposition near verdict start", + "PR 1007: no non-empty test-evidence file", + "PR 1008: missing or empty verdict.md", + "PR 1008: missing or empty summary.md", + "PR 1008: missing or empty review.md", + "PR 1008: missing or empty diff-analysis.md", + "PR 1008: missing or empty interactions.md", + "PR 1008: missing or empty trace.md", + "PR 1008: no merge/revise/reject disposition near verdict start", + "PR 1008: no non-empty test-evidence file", + "PR 1009: missing or empty verdict.md", + "PR 1009: missing or empty summary.md", + "PR 1009: missing or empty review.md", + "PR 1009: missing or empty diff-analysis.md", + "PR 1009: missing or empty interactions.md", + "PR 1009: missing or empty trace.md", + "PR 1009: no merge/revise/reject disposition near verdict start", + "PR 1009: no non-empty test-evidence file", + "PR 1010: missing or empty verdict.md", + "PR 1010: missing or empty summary.md", + "PR 1010: missing or empty review.md", + "PR 1010: missing or empty diff-analysis.md", + "PR 1010: missing or empty interactions.md", + "PR 1010: missing or empty trace.md", + "PR 1010: no merge/revise/reject disposition near verdict start", + "PR 1010: no non-empty test-evidence file", + "PR 1011: missing or empty verdict.md", + "PR 1011: missing or empty summary.md", + "PR 1011: missing or empty review.md", + "PR 1011: missing or empty diff-analysis.md", + "PR 1011: missing or empty interactions.md", + "PR 1011: missing or empty trace.md", + "PR 1011: no merge/revise/reject disposition near verdict start", + "PR 1011: no non-empty test-evidence file", + "PR 1012: missing or empty verdict.md", + "PR 1012: missing or empty summary.md", + "PR 1012: missing or empty review.md", + "PR 1012: missing or empty diff-analysis.md", + "PR 1012: missing or empty interactions.md", + "PR 1012: missing or empty trace.md", + "PR 1012: no merge/revise/reject disposition near verdict start", + "PR 1012: no non-empty test-evidence file", + "PR 1013: missing or empty verdict.md", + "PR 1013: missing or empty summary.md", + "PR 1013: missing or empty review.md", + "PR 1013: missing or empty diff-analysis.md", + "PR 1013: missing or empty interactions.md", + "PR 1013: missing or empty trace.md", + "PR 1013: no merge/revise/reject disposition near verdict start", + "PR 1013: no non-empty test-evidence file", + "PR 1014: missing or empty verdict.md", + "PR 1014: missing or empty summary.md", + "PR 1014: missing or empty review.md", + "PR 1014: missing or empty diff-analysis.md", + "PR 1014: missing or empty interactions.md", + "PR 1014: missing or empty trace.md", + "PR 1014: no merge/revise/reject disposition near verdict start", + "PR 1014: no non-empty test-evidence file", + "PR 1015: missing or empty verdict.md", + "PR 1015: missing or empty summary.md", + "PR 1015: missing or empty review.md", + "PR 1015: missing or empty diff-analysis.md", + "PR 1015: missing or empty interactions.md", + "PR 1015: missing or empty trace.md", + "PR 1015: no merge/revise/reject disposition near verdict start", + "PR 1015: no non-empty test-evidence file", + "PR 1016: missing or empty verdict.md", + "PR 1016: missing or empty summary.md", + "PR 1016: missing or empty review.md", + "PR 1016: missing or empty diff-analysis.md", + "PR 1016: missing or empty interactions.md", + "PR 1016: missing or empty trace.md", + "PR 1016: no merge/revise/reject disposition near verdict start", + "PR 1016: no non-empty test-evidence file", + "PR 1017: missing or empty verdict.md", + "PR 1017: missing or empty summary.md", + "PR 1017: missing or empty review.md", + "PR 1017: missing or empty diff-analysis.md", + "PR 1017: missing or empty interactions.md", + "PR 1017: missing or empty trace.md", + "PR 1017: no merge/revise/reject disposition near verdict start", + "PR 1017: no non-empty test-evidence file", + "PR 1018: missing or empty verdict.md", + "PR 1018: missing or empty summary.md", + "PR 1018: missing or empty review.md", + "PR 1018: missing or empty diff-analysis.md", + "PR 1018: missing or empty interactions.md", + "PR 1018: missing or empty trace.md", + "PR 1018: no merge/revise/reject disposition near verdict start", + "PR 1018: no non-empty test-evidence file", + "PR 1019: missing or empty verdict.md", + "PR 1019: missing or empty summary.md", + "PR 1019: missing or empty review.md", + "PR 1019: missing or empty diff-analysis.md", + "PR 1019: missing or empty interactions.md", + "PR 1019: missing or empty trace.md", + "PR 1019: no merge/revise/reject disposition near verdict start", + "PR 1019: no non-empty test-evidence file", + "PR 1020: missing or empty verdict.md", + "PR 1020: missing or empty summary.md", + "PR 1020: missing or empty review.md", + "PR 1020: missing or empty diff-analysis.md", + "PR 1020: missing or empty interactions.md", + "PR 1020: missing or empty trace.md", + "PR 1020: no merge/revise/reject disposition near verdict start", + "PR 1020: no non-empty test-evidence file", + "PR 1021: missing or empty verdict.md", + "PR 1021: missing or empty summary.md", + "PR 1021: missing or empty review.md", + "PR 1021: missing or empty diff-analysis.md", + "PR 1021: missing or empty interactions.md", + "PR 1021: missing or empty trace.md", + "PR 1021: no merge/revise/reject disposition near verdict start", + "PR 1021: no non-empty test-evidence file", + "PR 1022: missing or empty verdict.md", + "PR 1022: missing or empty summary.md", + "PR 1022: missing or empty review.md", + "PR 1022: missing or empty diff-analysis.md", + "PR 1022: missing or empty interactions.md", + "PR 1022: missing or empty trace.md", + "PR 1022: no merge/revise/reject disposition near verdict start", + "PR 1022: no non-empty test-evidence file", + "PR 1023: missing or empty verdict.md", + "PR 1023: missing or empty summary.md", + "PR 1023: missing or empty review.md", + "PR 1023: missing or empty diff-analysis.md", + "PR 1023: missing or empty interactions.md", + "PR 1023: missing or empty trace.md", + "PR 1023: no merge/revise/reject disposition near verdict start", + "PR 1023: no non-empty test-evidence file", + "PR 1024: missing or empty verdict.md", + "PR 1024: missing or empty summary.md", + "PR 1024: missing or empty review.md", + "PR 1024: missing or empty diff-analysis.md", + "PR 1024: missing or empty interactions.md", + "PR 1024: missing or empty trace.md", + "PR 1024: no merge/revise/reject disposition near verdict start", + "PR 1024: no non-empty test-evidence file", + "PR 1025: missing or empty verdict.md", + "PR 1025: missing or empty summary.md", + "PR 1025: missing or empty review.md", + "PR 1025: missing or empty diff-analysis.md", + "PR 1025: missing or empty interactions.md", + "PR 1025: missing or empty trace.md", + "PR 1025: no merge/revise/reject disposition near verdict start", + "PR 1025: no non-empty test-evidence file", + "PR 1026: missing or empty verdict.md", + "PR 1026: missing or empty summary.md", + "PR 1026: missing or empty review.md", + "PR 1026: missing or empty diff-analysis.md", + "PR 1026: missing or empty interactions.md", + "PR 1026: missing or empty trace.md", + "PR 1026: no merge/revise/reject disposition near verdict start", + "PR 1026: no non-empty test-evidence file", + "PR 1027: missing or empty verdict.md", + "PR 1027: missing or empty summary.md", + "PR 1027: missing or empty review.md", + "PR 1027: missing or empty diff-analysis.md", + "PR 1027: missing or empty interactions.md", + "PR 1027: missing or empty trace.md", + "PR 1027: no merge/revise/reject disposition near verdict start", + "PR 1027: no non-empty test-evidence file", + "PR 1028: missing or empty verdict.md", + "PR 1028: missing or empty summary.md", + "PR 1028: missing or empty review.md", + "PR 1028: missing or empty diff-analysis.md", + "PR 1028: missing or empty interactions.md", + "PR 1028: missing or empty trace.md", + "PR 1028: no merge/revise/reject disposition near verdict start", + "PR 1028: no non-empty test-evidence file", + "PR 1029: missing or empty verdict.md", + "PR 1029: missing or empty summary.md", + "PR 1029: missing or empty review.md", + "PR 1029: missing or empty diff-analysis.md", + "PR 1029: missing or empty interactions.md", + "PR 1029: missing or empty trace.md", + "PR 1029: no merge/revise/reject disposition near verdict start", + "PR 1029: no non-empty test-evidence file", + "PR 1030: missing or empty verdict.md", + "PR 1030: missing or empty summary.md", + "PR 1030: missing or empty review.md", + "PR 1030: missing or empty diff-analysis.md", + "PR 1030: missing or empty interactions.md", + "PR 1030: missing or empty trace.md", + "PR 1030: no merge/revise/reject disposition near verdict start", + "PR 1030: no non-empty test-evidence file", + "PR 1032: missing or empty verdict.md", + "PR 1032: missing or empty summary.md", + "PR 1032: missing or empty review.md", + "PR 1032: missing or empty diff-analysis.md", + "PR 1032: missing or empty interactions.md", + "PR 1032: missing or empty trace.md", + "PR 1032: no merge/revise/reject disposition near verdict start", + "PR 1032: no non-empty test-evidence file", + "PR 1033: missing or empty verdict.md", + "PR 1033: missing or empty summary.md", + "PR 1033: missing or empty review.md", + "PR 1033: missing or empty diff-analysis.md", + "PR 1033: missing or empty interactions.md", + "PR 1033: missing or empty trace.md", + "PR 1033: no merge/revise/reject disposition near verdict start", + "PR 1033: no non-empty test-evidence file", + "PR 1034: missing or empty verdict.md", + "PR 1034: missing or empty summary.md", + "PR 1034: missing or empty review.md", + "PR 1034: missing or empty diff-analysis.md", + "PR 1034: missing or empty interactions.md", + "PR 1034: missing or empty trace.md", + "PR 1034: no merge/revise/reject disposition near verdict start", + "PR 1034: no non-empty test-evidence file", + "PR 1035: missing or empty verdict.md", + "PR 1035: missing or empty summary.md", + "PR 1035: missing or empty review.md", + "PR 1035: missing or empty diff-analysis.md", + "PR 1035: missing or empty interactions.md", + "PR 1035: missing or empty trace.md", + "PR 1035: no merge/revise/reject disposition near verdict start", + "PR 1035: no non-empty test-evidence file", + "PR 1036: missing or empty verdict.md", + "PR 1036: missing or empty summary.md", + "PR 1036: missing or empty review.md", + "PR 1036: missing or empty diff-analysis.md", + "PR 1036: missing or empty interactions.md", + "PR 1036: missing or empty trace.md", + "PR 1036: no merge/revise/reject disposition near verdict start", + "PR 1036: no non-empty test-evidence file", + "PR 1037: missing or empty verdict.md", + "PR 1037: missing or empty summary.md", + "PR 1037: missing or empty review.md", + "PR 1037: missing or empty diff-analysis.md", + "PR 1037: missing or empty interactions.md", + "PR 1037: missing or empty trace.md", + "PR 1037: no merge/revise/reject disposition near verdict start", + "PR 1037: no non-empty test-evidence file", + "PR 1038: missing or empty verdict.md", + "PR 1038: missing or empty summary.md", + "PR 1038: missing or empty review.md", + "PR 1038: missing or empty diff-analysis.md", + "PR 1038: missing or empty interactions.md", + "PR 1038: missing or empty trace.md", + "PR 1038: no merge/revise/reject disposition near verdict start", + "PR 1038: no non-empty test-evidence file", + "PR 1039: missing or empty verdict.md", + "PR 1039: missing or empty summary.md", + "PR 1039: missing or empty review.md", + "PR 1039: missing or empty diff-analysis.md", + "PR 1039: missing or empty interactions.md", + "PR 1039: missing or empty trace.md", + "PR 1039: no merge/revise/reject disposition near verdict start", + "PR 1039: no non-empty test-evidence file", + "PR 1040: missing or empty verdict.md", + "PR 1040: missing or empty summary.md", + "PR 1040: missing or empty review.md", + "PR 1040: missing or empty diff-analysis.md", + "PR 1040: missing or empty interactions.md", + "PR 1040: missing or empty trace.md", + "PR 1040: no merge/revise/reject disposition near verdict start", + "PR 1040: no non-empty test-evidence file", + "PR 1042: missing or empty verdict.md", + "PR 1042: missing or empty summary.md", + "PR 1042: missing or empty review.md", + "PR 1042: missing or empty diff-analysis.md", + "PR 1042: missing or empty interactions.md", + "PR 1042: missing or empty trace.md", + "PR 1042: no merge/revise/reject disposition near verdict start", + "PR 1042: no non-empty test-evidence file", + "PR 1043: missing or empty verdict.md", + "PR 1043: missing or empty summary.md", + "PR 1043: missing or empty review.md", + "PR 1043: missing or empty diff-analysis.md", + "PR 1043: missing or empty interactions.md", + "PR 1043: missing or empty trace.md", + "PR 1043: no merge/revise/reject disposition near verdict start", + "PR 1043: no non-empty test-evidence file", + "PR 1044: missing or empty verdict.md", + "PR 1044: missing or empty summary.md", + "PR 1044: missing or empty review.md", + "PR 1044: missing or empty diff-analysis.md", + "PR 1044: missing or empty interactions.md", + "PR 1044: missing or empty trace.md", + "PR 1044: no merge/revise/reject disposition near verdict start", + "PR 1044: no non-empty test-evidence file", + "PR 1045: missing or empty verdict.md", + "PR 1045: missing or empty summary.md", + "PR 1045: missing or empty review.md", + "PR 1045: missing or empty diff-analysis.md", + "PR 1045: missing or empty interactions.md", + "PR 1045: missing or empty trace.md", + "PR 1045: no merge/revise/reject disposition near verdict start", + "PR 1045: no non-empty test-evidence file", + "PR 1046: missing or empty verdict.md", + "PR 1046: missing or empty summary.md", + "PR 1046: missing or empty review.md", + "PR 1046: missing or empty diff-analysis.md", + "PR 1046: missing or empty interactions.md", + "PR 1046: missing or empty trace.md", + "PR 1046: no merge/revise/reject disposition near verdict start", + "PR 1046: no non-empty test-evidence file", + "PR 1047: missing or empty verdict.md", + "PR 1047: missing or empty summary.md", + "PR 1047: missing or empty review.md", + "PR 1047: missing or empty diff-analysis.md", + "PR 1047: missing or empty interactions.md", + "PR 1047: missing or empty trace.md", + "PR 1047: no merge/revise/reject disposition near verdict start", + "PR 1047: no non-empty test-evidence file", + "PR 1048: missing or empty verdict.md", + "PR 1048: missing or empty summary.md", + "PR 1048: missing or empty review.md", + "PR 1048: missing or empty diff-analysis.md", + "PR 1048: missing or empty interactions.md", + "PR 1048: missing or empty trace.md", + "PR 1048: no merge/revise/reject disposition near verdict start", + "PR 1048: no non-empty test-evidence file", + "PR 1049: missing or empty verdict.md", + "PR 1049: missing or empty summary.md", + "PR 1049: missing or empty review.md", + "PR 1049: missing or empty diff-analysis.md", + "PR 1049: missing or empty interactions.md", + "PR 1049: missing or empty trace.md", + "PR 1049: no merge/revise/reject disposition near verdict start", + "PR 1049: no non-empty test-evidence file", + "PR 1050: missing or empty verdict.md", + "PR 1050: missing or empty summary.md", + "PR 1050: missing or empty review.md", + "PR 1050: missing or empty diff-analysis.md", + "PR 1050: missing or empty interactions.md", + "PR 1050: missing or empty trace.md", + "PR 1050: no merge/revise/reject disposition near verdict start", + "PR 1050: no non-empty test-evidence file", + "PR 1051: missing or empty verdict.md", + "PR 1051: missing or empty summary.md", + "PR 1051: missing or empty review.md", + "PR 1051: missing or empty diff-analysis.md", + "PR 1051: missing or empty interactions.md", + "PR 1051: missing or empty trace.md", + "PR 1051: no merge/revise/reject disposition near verdict start", + "PR 1051: no non-empty test-evidence file", + "PR 1052: missing or empty verdict.md", + "PR 1052: missing or empty summary.md", + "PR 1052: missing or empty review.md", + "PR 1052: missing or empty diff-analysis.md", + "PR 1052: missing or empty interactions.md", + "PR 1052: missing or empty trace.md", + "PR 1052: no merge/revise/reject disposition near verdict start", + "PR 1052: no non-empty test-evidence file", + "PR 1053: missing or empty verdict.md", + "PR 1053: missing or empty summary.md", + "PR 1053: missing or empty review.md", + "PR 1053: missing or empty diff-analysis.md", + "PR 1053: missing or empty interactions.md", + "PR 1053: missing or empty trace.md", + "PR 1053: no merge/revise/reject disposition near verdict start", + "PR 1053: no non-empty test-evidence file", + "PR 1054: missing or empty verdict.md", + "PR 1054: missing or empty summary.md", + "PR 1054: missing or empty review.md", + "PR 1054: missing or empty diff-analysis.md", + "PR 1054: missing or empty interactions.md", + "PR 1054: missing or empty trace.md", + "PR 1054: no merge/revise/reject disposition near verdict start", + "PR 1054: no non-empty test-evidence file", + "PR 1055: missing or empty verdict.md", + "PR 1055: missing or empty summary.md", + "PR 1055: missing or empty review.md", + "PR 1055: missing or empty diff-analysis.md", + "PR 1055: missing or empty interactions.md", + "PR 1055: missing or empty trace.md", + "PR 1055: no merge/revise/reject disposition near verdict start", + "PR 1055: no non-empty test-evidence file", + "PR 1056: missing or empty verdict.md", + "PR 1056: missing or empty summary.md", + "PR 1056: missing or empty review.md", + "PR 1056: missing or empty diff-analysis.md", + "PR 1056: missing or empty interactions.md", + "PR 1056: missing or empty trace.md", + "PR 1056: no merge/revise/reject disposition near verdict start", + "PR 1056: no non-empty test-evidence file", + "PR 1057: missing or empty verdict.md", + "PR 1057: missing or empty summary.md", + "PR 1057: missing or empty review.md", + "PR 1057: missing or empty diff-analysis.md", + "PR 1057: missing or empty interactions.md", + "PR 1057: missing or empty trace.md", + "PR 1057: no merge/revise/reject disposition near verdict start", + "PR 1057: no non-empty test-evidence file", + "parsed verdict total is 1, expected 75", + "PRs without a source-file or frozen-diff trace pointer: [351, 364, 716, 750, 959, 961, 966, 973, 974, 983, 988, 990, 991, 992, 993, 994, 996, 997, 998, 999, 1000, 1002, 1003, 1004, 1005, 1006, 1007, 1008, 1009, 1010, 1011, 1012, 1013, 1014, 1015, 1016, 1017, 1018, 1019, 1020, 1021, 1022, 1023, 1024, 1025, 1026, 1027, 1028, 1029, 1030, 1032, 1033, 1034, 1035, 1036, 1037, 1038, 1039, 1040, 1042, 1043, 1044, 1045, 1046, 1047, 1048, 1049, 1050, 1051, 1052, 1053, 1054, 1055, 1056, 1057]", + "PRs whose review has no source-file/diff citation: [351, 364, 716, 750, 961, 966, 973, 974, 983, 988, 990, 991, 992, 993, 994, 996, 997, 998, 999, 1000, 1002, 1003, 1004, 1005, 1006, 1007, 1008, 1009, 1010, 1011, 1012, 1013, 1014, 1015, 1016, 1017, 1018, 1019, 1020, 1021, 1022, 1023, 1024, 1025, 1026, 1027, 1028, 1029, 1030, 1032, 1033, 1034, 1035, 1036, 1037, 1038, 1039, 1040, 1042, 1043, 1044, 1045, 1046, 1047, 1048, 1049, 1050, 1051, 1052, 1053, 1054, 1055, 1056, 1057]", + "PRs without an explicit bounty-tier declaration: [351, 364, 716, 750, 959, 961, 966, 973, 974, 983, 988, 990, 991, 992, 993, 994, 996, 997, 998, 999, 1000, 1002, 1003, 1004, 1005, 1006, 1007, 1008, 1009, 1010, 1011, 1012, 1013, 1014, 1015, 1016, 1017, 1018, 1019, 1020, 1021, 1022, 1023, 1024, 1025, 1026, 1027, 1028, 1029, 1030, 1032, 1033, 1034, 1035, 1036, 1037, 1038, 1039, 1040, 1042, 1043, 1044, 1045, 1046, 1047, 1048, 1049, 1050, 1051, 1052, 1053, 1054, 1055, 1056, 1057]", + "dependency analysis omits 205/205 high-confidence file-overlap pairs", + "frozen files.txt counts disagree with metadata for PRs, but the audit does not document the baseline/reverse-drift limitation", + "tool log records 0/31 transcript tool events" + ], + "warnings": [ + "dependency analysis does not enumerate 676/676 raw frozen-file pairs; raw lists include known reverse-drift, so this is reported separately and is not itself a semantic failure" + ] + }, + "telemetry": { + "schema_version": 1, + "run_name": "gemma4-31b-q4_75pr_v2", + "source_csv": "/home/michael/gemma4-campaign-state/telemetry/gemma4-31b-q4-gpu.csv", + "generated_at": "2026-08-02T06:45:35.340803+00:00", + "analyzer": { + "path": "/home/michael/bench-gemma4-31b-q4/tooling/deployments/gemma4-31b-q4-tower2/analyze_replica_telemetry.py", + "file_sha256": "19a2d7a9cdfc36d2fcd9840581015410b5451aacff43e2328e050c1214e771bd", + "git_sha": "17e0128cb8c3268d4616a51147b1b754db8103b6" + }, + "attribution": { + "method": "live harness --port maps port 8000 to GPU 0 and port 8001 to GPU 1", + "active_gpu_ids_observed": [ + "0" + ], + "other_gpu_work_is_reported_as_concurrency_not_charged_to_active_gpu_energy": true, + "cpu_package_power_is_shared_host_context_and_not_uniquely_attributable_when_runs_overlap": true + }, + "sampling": { + "samples": 39, + "paired_host_samples": 39, + "first_sample": "2026-08-02T06:15:26+00:00", + "last_sample": "2026-08-02T06:18:36+00:00", + "window_source": "summary.json", + "window_started_at": "2026-08-02T06:15:25.365160+00:00", + "window_ended_at": "2026-08-02T06:18:39.125819+00:00", + "window_wall_s": 193.8, + "integrated_coverage_s": 190.0, + "coverage_fraction_of_wall": 0.9804, + "max_accepted_gap_s": 15.0 + }, + "active_gpu": { + "mean_power_w": 459.98, + "time_weighted_power_w": 462.99, + "p90_power_w": 500.57, + "max_power_w": 509.05, + "mean_sm_util_pct": 83.54, + "p90_sm_util_pct": 98.0, + "max_memory_used_mib": 66503.0, + "max_temp_c": 73.0, + "mean_sm_clock_mhz": 2631.9, + "configured_cap_w": 500.0 + }, + "cpu_package_shared_context": { + "samples": 39, + "mean_power_w": 135.56, + "max_power_w": 136.98, + "integrated_coverage_s": 190.0 + }, + "concurrency": { + "other_gpu_cell_sample_counts": { + "gemma4-31b-q4_75pr_v1": 39 + }, + "both_gpus_over_20pct_sm_fraction": 0.8718, + "mean_observed_two_gpu_plus_cpu_package_w": 1090.46, + "max_observed_two_gpu_plus_cpu_package_w": 1144.89, + "wall_power_note": "This is GPU plus CPU package telemetry, not AC wall draw; no software wall meter is available." + }, + "energy_and_cost": { + "active_gpu_sampled_kwh": 0.024436, + "active_gpu_sampled_cost_usd": 0.003177, + "cpu_package_shared_sampled_kwh": 0.007153, + "active_gpu_cap_upper_bound_kwh": 0.026917, + "rate_usd_per_kwh": 0.13 + } + }, + "evidence": { + "summary_sha256": "84671b0076d1eed4faca9a26553b399a002e9bf90e732e37bbf32d477d02a0e4", + "receipt_sha256": "3f06c30bdae4813b39c0ccf74770892179550be0590463e672c0609f889d2b7c", + "transcript_sha256": "01c6decba025a16b03929e68486de8aa8ce653f65e83dc4472195b36013a1c56", + "validator_sha256": "11479b273a7c7c4b7597736ccc6eb8b763443192c8f6c1e6a9b7067ec0dded83", + "task_sha256": "0d426e15303b3f886a75db81165cdbe9d91534a474217f91f4e2baf41c0056a6", + "fixture_path": "/mnt/bulk/benchmark-fixtures/dreamserver-75-pr-audit-2026-04-27" + } +} diff --git a/benchmarks/gemma4-31b-q4/75pr-v3-audit.json b/benchmarks/gemma4-31b-q4/75pr-v3-audit.json new file mode 100644 index 00000000..1467c416 --- /dev/null +++ b/benchmarks/gemma4-31b-q4/75pr-v3-audit.json @@ -0,0 +1,941 @@ +{ + "schema_version": 2, + "audited_at": "2026-08-02T06:46:33.560626+00:00", + "run_name": "Gemma-4-31B-it-QAT-Q4_0 gemma4-31b-q4_75pr_v3", + "classification": "MODEL_TERMINAL_FAILURE", + "legacy_shipped": false, + "verdict_counts": { + "merge": 11, + "reject": 2, + "revise": 1 + }, + "runtime": { + "wall_s": 561.7, + "iterations": 41, + "completion_tokens": 25841, + "prompt_tokens": 1193936, + "completion_tps_model_call": 47.6 + }, + "artifact_integrity": { + "archive_sha256": "3246d75825dc7c8b9f0abb76c3b4ddf2e1a716a4b18e679e10922a86756275cc", + "archive_bytes": 626792011, + "archive_gzip_valid": true, + "head": "9adc69f4e35ff4f6a6cc14d6c31c45d38b79d28f", + "tags": [ + "v1.0.0" + ], + "tag_targets": { + "v1.0.0": "9adc69f4e35ff4f6a6cc14d6c31c45d38b79d28f" + }, + "tag_at_head": true, + "clean": false, + "commit_count": 17 + }, + "substance": { + "test_evidence_count": 4, + "test_evidence_prs": [ + 351, + 1018, + 1019, + 1052 + ], + "missing_test_or_skip_count": 71, + "missing_test_or_skip_prs": [ + 364, + 716, + 750, + 959, + 961, + 966, + 973, + 974, + 983, + 988, + 990, + 991, + 992, + 993, + 994, + 996, + 997, + 998, + 999, + 1000, + 1002, + 1003, + 1004, + 1005, + 1006, + 1007, + 1008, + 1009, + 1010, + 1011, + 1012, + 1013, + 1014, + 1015, + 1016, + 1017, + 1020, + 1021, + 1022, + 1023, + 1024, + 1025, + 1026, + 1027, + 1028, + 1029, + 1030, + 1032, + 1033, + 1034, + 1035, + 1036, + 1037, + 1038, + 1039, + 1040, + 1042, + 1043, + 1044, + 1045, + 1046, + 1047, + 1048, + 1049, + 1050, + 1051, + 1053, + 1054, + 1055, + 1056, + 1057 + ], + "reviews_under_800_bytes_count": 75, + "reviews_under_800_bytes_prs": [ + 351, + 364, + 716, + 750, + 959, + 961, + 966, + 973, + 974, + 983, + 988, + 990, + 991, + 992, + 993, + 994, + 996, + 997, + 998, + 999, + 1000, + 1002, + 1003, + 1004, + 1005, + 1006, + 1007, + 1008, + 1009, + 1010, + 1011, + 1012, + 1013, + 1014, + 1015, + 1016, + 1017, + 1018, + 1019, + 1020, + 1021, + 1022, + 1023, + 1024, + 1025, + 1026, + 1027, + 1028, + 1029, + 1030, + 1032, + 1033, + 1034, + 1035, + 1036, + 1037, + 1038, + 1039, + 1040, + 1042, + 1043, + 1044, + 1045, + 1046, + 1047, + 1048, + 1049, + 1050, + 1051, + 1052, + 1053, + 1054, + 1055, + 1056, + 1057 + ], + "reviews_with_source_citation_count": 1, + "reviews_with_source_citation_prs": [ + 1007 + ], + "reviews_with_hunk_marker_count": 0, + "traces_with_source_citation_count": 4, + "traces_with_hunk_marker_count": 0, + "diff_analyses_with_hunk_marker_count": 0, + "explicit_actual_bounty_tier_count": 0, + "explicit_actual_bounty_tier_prs": [], + "transcript_tool_events": 41, + "numbered_tool_log_entries": 0 + }, + "strict_validation": { + "schema_version": 1, + "workspace": "/home/michael/bench-gemma4-31b-q4/tooling/workspace/gemma4-31b-q4_75pr_v3/audit-repo", + "fixture": "/mnt/bulk/benchmark-fixtures/dreamserver-75-pr-audit-2026-04-27", + "pass": false, + "metrics": { + "canonical_prs": 75, + "actual_pr_dirs": 75, + "missing_pr_dirs": [], + "extra_pr_dirs": [], + "verdict_counts": { + "merge": 11, + "reject": 2, + "revise": 1 + }, + "risk_matrix_missing_pr_refs": [ + 364, + 716, + 961, + 983, + 988, + 990, + 991, + 992, + 993, + 994, + 996, + 997, + 998, + 999, + 1000, + 1002, + 1003, + 1004, + 1006, + 1007, + 1008, + 1009, + 1010, + 1011, + 1012, + 1014, + 1015, + 1016, + 1020, + 1021, + 1022, + 1023, + 1024, + 1025, + 1026, + 1027, + 1028, + 1029, + 1030, + 1032, + 1033, + 1034, + 1035, + 1036, + 1037, + 1038, + 1039, + 1040, + 1042, + 1043, + 1044, + 1045, + 1046, + 1047, + 1048, + 1049, + 1050, + 1051, + 1053, + 1054, + 1056, + 1057 + ], + "raw_frozen_file_overlap_pairs": 676, + "trusted_file_list_prs": 63, + "trusted_file_overlap_pairs": 205, + "listed_pair_mentions": 26, + "missing_trusted_file_overlap_pairs": 202, + "missing_trusted_file_overlap_pair_sample": [ + [ + 990, + 991 + ], + [ + 992, + 1013 + ], + [ + 992, + 1017 + ], + [ + 993, + 994 + ], + [ + 993, + 997 + ], + [ + 993, + 998 + ], + [ + 993, + 999 + ], + [ + 993, + 1000 + ], + [ + 993, + 1002 + ], + [ + 993, + 1006 + ], + [ + 993, + 1007 + ], + [ + 993, + 1008 + ], + [ + 993, + 1011 + ], + [ + 993, + 1016 + ], + [ + 993, + 1018 + ], + [ + 993, + 1020 + ], + [ + 994, + 997 + ], + [ + 994, + 998 + ], + [ + 994, + 999 + ], + [ + 994, + 1000 + ], + [ + 994, + 1002 + ], + [ + 994, + 1006 + ], + [ + 994, + 1007 + ], + [ + 994, + 1008 + ], + [ + 994, + 1010 + ], + [ + 994, + 1011 + ], + [ + 994, + 1016 + ], + [ + 994, + 1017 + ], + [ + 994, + 1018 + ], + [ + 994, + 1020 + ], + [ + 996, + 1012 + ], + [ + 996, + 1017 + ], + [ + 996, + 1026 + ], + [ + 997, + 998 + ], + [ + 997, + 999 + ], + [ + 997, + 1000 + ], + [ + 997, + 1002 + ], + [ + 997, + 1006 + ], + [ + 997, + 1007 + ], + [ + 997, + 1008 + ], + [ + 997, + 1011 + ], + [ + 997, + 1016 + ], + [ + 997, + 1018 + ], + [ + 997, + 1020 + ], + [ + 998, + 999 + ], + [ + 998, + 1000 + ], + [ + 998, + 1002 + ], + [ + 998, + 1006 + ], + [ + 998, + 1007 + ], + [ + 998, + 1008 + ] + ], + "missing_raw_frozen_overlap_pairs": 669, + "missing_raw_frozen_overlap_pair_sample": [ + [ + 351, + 364 + ], + [ + 351, + 716 + ], + [ + 351, + 750 + ], + [ + 351, + 959 + ], + [ + 351, + 961 + ], + [ + 351, + 966 + ], + [ + 351, + 973 + ], + [ + 351, + 974 + ], + [ + 351, + 983 + ], + [ + 351, + 988 + ], + [ + 351, + 990 + ], + [ + 351, + 991 + ], + [ + 351, + 992 + ], + [ + 351, + 993 + ], + [ + 351, + 994 + ], + [ + 351, + 996 + ], + [ + 351, + 997 + ], + [ + 351, + 998 + ], + [ + 351, + 999 + ], + [ + 351, + 1000 + ], + [ + 351, + 1002 + ], + [ + 351, + 1003 + ], + [ + 351, + 1004 + ], + [ + 351, + 1005 + ], + [ + 351, + 1006 + ], + [ + 351, + 1007 + ], + [ + 351, + 1008 + ], + [ + 351, + 1009 + ], + [ + 351, + 1010 + ], + [ + 351, + 1011 + ], + [ + 351, + 1012 + ], + [ + 351, + 1013 + ], + [ + 351, + 1015 + ], + [ + 351, + 1016 + ], + [ + 351, + 1017 + ], + [ + 351, + 1018 + ], + [ + 351, + 1019 + ], + [ + 351, + 1020 + ], + [ + 351, + 1021 + ], + [ + 351, + 1022 + ], + [ + 351, + 1023 + ], + [ + 351, + 1024 + ], + [ + 351, + 1025 + ], + [ + 351, + 1026 + ], + [ + 351, + 1027 + ], + [ + 351, + 1028 + ], + [ + 351, + 1029 + ], + [ + 351, + 1030 + ], + [ + 351, + 1032 + ], + [ + 351, + 1033 + ] + ], + "raw_file_count_metadata_mismatch_prs": [ + 351, + 364, + 716, + 750, + 959, + 961, + 966, + 973, + 974, + 983, + 988, + 1043 + ], + "transcript_tool_events": 41, + "numbered_tool_log_entries": 0 + }, + "failures": [ + "missing or empty required artifact: sources.md", + "missing or empty required artifact: analysis/surface-area.md", + "missing or empty required artifact: testing/baseline.md", + "missing or empty required artifact: research/questions.md", + "missing or empty required artifact: research/dead-ends.md", + "missing or empty required artifact: research/upstream-context.md", + "missing non-empty decisions directory", + "PR 351: no merge/revise/reject disposition near verdict start", + "PR 364: no non-empty test-evidence file", + "PR 716: no merge/revise/reject disposition near verdict start", + "PR 716: no non-empty test-evidence file", + "PR 750: no merge/revise/reject disposition near verdict start", + "PR 750: no non-empty test-evidence file", + "PR 959: no non-empty test-evidence file", + "PR 961: no merge/revise/reject disposition near verdict start", + "PR 961: no non-empty test-evidence file", + "PR 966: no non-empty test-evidence file", + "PR 973: no merge/revise/reject disposition near verdict start", + "PR 973: no non-empty test-evidence file", + "PR 974: no non-empty test-evidence file", + "PR 983: no merge/revise/reject disposition near verdict start", + "PR 983: no non-empty test-evidence file", + "PR 988: no merge/revise/reject disposition near verdict start", + "PR 988: no non-empty test-evidence file", + "PR 990: no merge/revise/reject disposition near verdict start", + "PR 990: no non-empty test-evidence file", + "PR 991: no merge/revise/reject disposition near verdict start", + "PR 991: no non-empty test-evidence file", + "PR 992: no merge/revise/reject disposition near verdict start", + "PR 992: no non-empty test-evidence file", + "PR 993: no merge/revise/reject disposition near verdict start", + "PR 993: no non-empty test-evidence file", + "PR 994: no merge/revise/reject disposition near verdict start", + "PR 994: no non-empty test-evidence file", + "PR 996: no merge/revise/reject disposition near verdict start", + "PR 996: no non-empty test-evidence file", + "PR 997: no merge/revise/reject disposition near verdict start", + "PR 997: no non-empty test-evidence file", + "PR 998: no merge/revise/reject disposition near verdict start", + "PR 998: no non-empty test-evidence file", + "PR 999: no merge/revise/reject disposition near verdict start", + "PR 999: no non-empty test-evidence file", + "PR 1000: no merge/revise/reject disposition near verdict start", + "PR 1000: no non-empty test-evidence file", + "PR 1002: no merge/revise/reject disposition near verdict start", + "PR 1002: no non-empty test-evidence file", + "PR 1003: no merge/revise/reject disposition near verdict start", + "PR 1003: no non-empty test-evidence file", + "PR 1004: no merge/revise/reject disposition near verdict start", + "PR 1004: no non-empty test-evidence file", + "PR 1005: no non-empty test-evidence file", + "PR 1006: no merge/revise/reject disposition near verdict start", + "PR 1006: no non-empty test-evidence file", + "PR 1007: no non-empty test-evidence file", + "PR 1008: no merge/revise/reject disposition near verdict start", + "PR 1008: no non-empty test-evidence file", + "PR 1009: no merge/revise/reject disposition near verdict start", + "PR 1009: no non-empty test-evidence file", + "PR 1010: no merge/revise/reject disposition near verdict start", + "PR 1010: no non-empty test-evidence file", + "PR 1011: no merge/revise/reject disposition near verdict start", + "PR 1011: no non-empty test-evidence file", + "PR 1012: no merge/revise/reject disposition near verdict start", + "PR 1012: no non-empty test-evidence file", + "PR 1013: no non-empty test-evidence file", + "PR 1014: no merge/revise/reject disposition near verdict start", + "PR 1014: no non-empty test-evidence file", + "PR 1015: no merge/revise/reject disposition near verdict start", + "PR 1015: no non-empty test-evidence file", + "PR 1016: no merge/revise/reject disposition near verdict start", + "PR 1016: no non-empty test-evidence file", + "PR 1017: no non-empty test-evidence file", + "PR 1020: no merge/revise/reject disposition near verdict start", + "PR 1020: no non-empty test-evidence file", + "PR 1021: no non-empty test-evidence file", + "PR 1022: no merge/revise/reject disposition near verdict start", + "PR 1022: no non-empty test-evidence file", + "PR 1023: no merge/revise/reject disposition near verdict start", + "PR 1023: no non-empty test-evidence file", + "PR 1024: no merge/revise/reject disposition near verdict start", + "PR 1024: no non-empty test-evidence file", + "PR 1025: no merge/revise/reject disposition near verdict start", + "PR 1025: no non-empty test-evidence file", + "PR 1026: no merge/revise/reject disposition near verdict start", + "PR 1026: no non-empty test-evidence file", + "PR 1027: no merge/revise/reject disposition near verdict start", + "PR 1027: no non-empty test-evidence file", + "PR 1028: no merge/revise/reject disposition near verdict start", + "PR 1028: no non-empty test-evidence file", + "PR 1029: no merge/revise/reject disposition near verdict start", + "PR 1029: no non-empty test-evidence file", + "PR 1030: no non-empty test-evidence file", + "PR 1032: no merge/revise/reject disposition near verdict start", + "PR 1032: no non-empty test-evidence file", + "PR 1033: no merge/revise/reject disposition near verdict start", + "PR 1033: no non-empty test-evidence file", + "PR 1034: no merge/revise/reject disposition near verdict start", + "PR 1034: no non-empty test-evidence file", + "PR 1035: no merge/revise/reject disposition near verdict start", + "PR 1035: no non-empty test-evidence file", + "PR 1036: no merge/revise/reject disposition near verdict start", + "PR 1036: no non-empty test-evidence file", + "PR 1037: no merge/revise/reject disposition near verdict start", + "PR 1037: no non-empty test-evidence file", + "PR 1038: no merge/revise/reject disposition near verdict start", + "PR 1038: no non-empty test-evidence file", + "PR 1039: no merge/revise/reject disposition near verdict start", + "PR 1039: no non-empty test-evidence file", + "PR 1040: no merge/revise/reject disposition near verdict start", + "PR 1040: no non-empty test-evidence file", + "PR 1042: no merge/revise/reject disposition near verdict start", + "PR 1042: no non-empty test-evidence file", + "PR 1043: no merge/revise/reject disposition near verdict start", + "PR 1043: no non-empty test-evidence file", + "PR 1044: no merge/revise/reject disposition near verdict start", + "PR 1044: no non-empty test-evidence file", + "PR 1045: no merge/revise/reject disposition near verdict start", + "PR 1045: no non-empty test-evidence file", + "PR 1046: no merge/revise/reject disposition near verdict start", + "PR 1046: no non-empty test-evidence file", + "PR 1047: no merge/revise/reject disposition near verdict start", + "PR 1047: no non-empty test-evidence file", + "PR 1048: no merge/revise/reject disposition near verdict start", + "PR 1048: no non-empty test-evidence file", + "PR 1049: no merge/revise/reject disposition near verdict start", + "PR 1049: no non-empty test-evidence file", + "PR 1050: no merge/revise/reject disposition near verdict start", + "PR 1050: no non-empty test-evidence file", + "PR 1051: no merge/revise/reject disposition near verdict start", + "PR 1051: no non-empty test-evidence file", + "PR 1053: no merge/revise/reject disposition near verdict start", + "PR 1053: no non-empty test-evidence file", + "PR 1054: no merge/revise/reject disposition near verdict start", + "PR 1054: no non-empty test-evidence file", + "PR 1055: no non-empty test-evidence file", + "PR 1056: no merge/revise/reject disposition near verdict start", + "PR 1056: no non-empty test-evidence file", + "PR 1057: no merge/revise/reject disposition near verdict start", + "PR 1057: no non-empty test-evidence file", + "parsed verdict total is 14, expected 75", + "PRs without a source-file or frozen-diff trace pointer: [351, 364, 716, 750, 961, 966, 973, 974, 983, 988, 990, 991, 992, 993, 994, 996, 997, 998, 999, 1000, 1002, 1003, 1004, 1005, 1006, 1008, 1009, 1010, 1011, 1012, 1013, 1014, 1015, 1016, 1017, 1018, 1019, 1020, 1022, 1023, 1024, 1025, 1026, 1027, 1028, 1029, 1032, 1033, 1034, 1035, 1036, 1037, 1038, 1039, 1040, 1042, 1043, 1044, 1045, 1046, 1047, 1048, 1049, 1050, 1051, 1052, 1053, 1054, 1055, 1056, 1057]", + "PRs whose review has no source-file/diff citation: [351, 364, 716, 750, 959, 961, 966, 973, 974, 983, 988, 990, 991, 992, 993, 994, 996, 997, 998, 999, 1000, 1002, 1003, 1004, 1005, 1006, 1008, 1009, 1010, 1011, 1012, 1013, 1014, 1015, 1016, 1017, 1018, 1019, 1020, 1021, 1022, 1023, 1024, 1025, 1026, 1027, 1028, 1029, 1030, 1032, 1033, 1034, 1035, 1036, 1037, 1038, 1039, 1040, 1042, 1043, 1044, 1045, 1046, 1047, 1048, 1049, 1050, 1051, 1052, 1053, 1054, 1055, 1056, 1057]", + "PRs without an explicit bounty-tier declaration: [351, 364, 716, 750, 959, 961, 966, 973, 974, 983, 988, 990, 991, 992, 993, 994, 996, 997, 998, 999, 1000, 1002, 1003, 1004, 1005, 1006, 1007, 1008, 1009, 1010, 1011, 1012, 1013, 1014, 1015, 1016, 1017, 1018, 1019, 1020, 1021, 1022, 1023, 1024, 1025, 1026, 1027, 1028, 1029, 1030, 1032, 1033, 1034, 1035, 1036, 1037, 1038, 1039, 1040, 1042, 1043, 1044, 1045, 1046, 1047, 1048, 1049, 1050, 1051, 1052, 1053, 1054, 1055, 1056, 1057]", + "PRs with neither execution nor explicit skip evidence: [351, 1019, 1052]", + "risk_matrix omits PRs: [364, 716, 961, 983, 988, 990, 991, 992, 993, 994, 996, 997, 998, 999, 1000, 1002, 1003, 1004, 1006, 1007, 1008, 1009, 1010, 1011, 1012, 1014, 1015, 1016, 1020, 1021, 1022, 1023, 1024, 1025, 1026, 1027, 1028, 1029, 1030, 1032, 1033, 1034, 1035, 1036, 1037, 1038, 1039, 1040, 1042, 1043, 1044, 1045, 1046, 1047, 1048, 1049, 1050, 1051, 1053, 1054, 1056, 1057]", + "dependency analysis omits 202/205 high-confidence file-overlap pairs", + "frozen files.txt counts disagree with metadata for PRs, but the audit does not document the baseline/reverse-drift limitation", + "tool log records 0/41 transcript tool events" + ], + "warnings": [ + "dependency analysis does not enumerate 669/676 raw frozen-file pairs; raw lists include known reverse-drift, so this is reported separately and is not itself a semantic failure" + ] + }, + "telemetry": { + "schema_version": 1, + "run_name": "gemma4-31b-q4_75pr_v3", + "source_csv": "/home/michael/gemma4-campaign-state/telemetry/gemma4-31b-q4-gpu.csv", + "generated_at": "2026-08-02T06:45:35.341083+00:00", + "analyzer": { + "path": "/home/michael/bench-gemma4-31b-q4/tooling/deployments/gemma4-31b-q4-tower2/analyze_replica_telemetry.py", + "file_sha256": "19a2d7a9cdfc36d2fcd9840581015410b5451aacff43e2328e050c1214e771bd", + "git_sha": "17e0128cb8c3268d4616a51147b1b754db8103b6" + }, + "attribution": { + "method": "live harness --port maps port 8000 to GPU 0 and port 8001 to GPU 1", + "active_gpu_ids_observed": [ + "1" + ], + "other_gpu_work_is_reported_as_concurrency_not_charged_to_active_gpu_energy": true, + "cpu_package_power_is_shared_host_context_and_not_uniquely_attributable_when_runs_overlap": true + }, + "sampling": { + "samples": 112, + "paired_host_samples": 112, + "first_sample": "2026-08-02T06:32:36+00:00", + "last_sample": "2026-08-02T06:41:51+00:00", + "window_source": "summary.json", + "window_started_at": "2026-08-02T06:32:33.771595+00:00", + "window_ended_at": "2026-08-02T06:41:55.442035+00:00", + "window_wall_s": 561.7, + "integrated_coverage_s": 555.0, + "coverage_fraction_of_wall": 0.9881, + "max_accepted_gap_s": 15.0 + }, + "active_gpu": { + "mean_power_w": 475.31, + "time_weighted_power_w": 475.08, + "p90_power_w": 500.32, + "max_power_w": 510.84, + "mean_sm_util_pct": 90.95, + "p90_sm_util_pct": 98.0, + "max_memory_used_mib": 66699.0, + "max_temp_c": 71.0, + "mean_sm_clock_mhz": 2720.9, + "configured_cap_w": 500.0 + }, + "cpu_package_shared_context": { + "samples": 112, + "mean_power_w": 113.63, + "max_power_w": 119.55, + "integrated_coverage_s": 555.0 + }, + "concurrency": { + "other_gpu_cell_sample_counts": { + "(idle/no benchmark harness)": 112 + }, + "both_gpus_over_20pct_sm_fraction": 0.0, + "mean_observed_two_gpu_plus_cpu_package_w": 606.23, + "max_observed_two_gpu_plus_cpu_package_w": 643.65, + "wall_power_note": "This is GPU plus CPU package telemetry, not AC wall draw; no software wall meter is available." + }, + "energy_and_cost": { + "active_gpu_sampled_kwh": 0.073242, + "active_gpu_sampled_cost_usd": 0.009521, + "cpu_package_shared_sampled_kwh": 0.017521, + "active_gpu_cap_upper_bound_kwh": 0.078014, + "rate_usd_per_kwh": 0.13 + } + }, + "evidence": { + "summary_sha256": "ce48b91c08ea6759d22844de7533700fdecd0c4b7ce6a2c857cdba9f80bd95e4", + "receipt_sha256": "ec88a1c45d985f0fb5e47b7b0a0f781e68087f4a7d01fb7d85bf796e3a752b18", + "transcript_sha256": "3bdf12182ac973b2b7e647152b6a150aebf1f8cfdeb90172de95c3e81a4aae27", + "validator_sha256": "11479b273a7c7c4b7597736ccc6eb8b763443192c8f6c1e6a9b7067ec0dded83", + "task_sha256": "0d426e15303b3f886a75db81165cdbe9d91534a474217f91f4e2baf41c0056a6", + "fixture_path": "/mnt/bulk/benchmark-fixtures/dreamserver-75-pr-audit-2026-04-27" + } +} diff --git a/benchmarks/gemma4-31b-q4/GEMMA4_31B_Q4_COMPLETION_AUDIT.md b/benchmarks/gemma4-31b-q4/GEMMA4_31B_Q4_COMPLETION_AUDIT.md new file mode 100644 index 00000000..c6aebeca --- /dev/null +++ b/benchmarks/gemma4-31b-q4/GEMMA4_31B_Q4_COMPLETION_AUDIT.md @@ -0,0 +1,30 @@ +# Gemma 4 31B Q4 campaign completion audit + +| Requirement | Evidence | Status | +|---|---|---| +| Exact model and runtime pinned | Deployment `model-manifest.json` and `benchmark-serving-manifest.json` include upstream revision, byte size, SHA-256, llama.cpp commit/binary hash, CUDA build, and sampling | Complete | +| Best stable two-GPU topology selected | Preregistered topology bakeoff rejected corrupt/unsupported split modes and selected two independent four-slot Q8-KV replicas | Complete | +| Full native context used | Every slot and request cap is 262,144; long-context recall passed at 245,347 total tokens | Complete | +| 500 W safety envelope | Both replicas and every receipt use 500 W per GPU; telemetry records actual per-run power/utilization | Complete | +| Sanctuary and Pixel validated before testing | Both agents passed isolated routing, chat, tool-call, and tool-follow-up checks on their dedicated replicas | Complete | +| Canonical N=3 and N=10 complete | 36/36 and 120/120 evidence cells audited; raw and corrected results remain separate | Complete | +| Low-ceiling failures excluded | No canonical failure hit an artificial output cap. One proven one-hour transport cancellation below native context is preserved as infrastructure-invalid and exactly replaced | Complete | +| Extended suites complete | Twelve valid runs plus one preserved invalid supervisor attempt; identity/configuration/preservation audit rerun after replacement | Complete | +| Strict modality review complete | Workbooks inspected, deck files rendered and visually reviewed, single-PR results independently reproduced, all 75-PR outputs structurally audited | Complete | +| Harness findings repaired | Substance monitor handles malformed arguments; tag gate requires clean repo + annotated tag at HEAD; 75-PR auditor fails closed on missing files | Complete | +| Cross-model comparison pinned | `tooling/gemma4-comparison-sources.json` freezes Qwen3.6-27B, Qwen3.6-35B-A3B, Coder-Next, Qwen3.5-397B, and DeepSeek evidence hashes | Complete | +| Publish/merge | This entry, audit overlays, scorecards, and deployment package are delivered by the merge containing this audit | Complete | +| Restore production | Exact pre-campaign OpenClaw config and proven DeepSeek launcher restored; DeepSeek health/model route, 500 W caps, portals, and fresh Sanctuary/Pixel `exec` calls passed with no fallback | Complete | + +The evidence audit and the quality audit answer different questions. The former +proves what ran and that the bytes were preserved. The latter deliberately +fails polished-but-incomplete artifacts. Consequently, all 12 extended runs +are accounted for while zero receive a strict substantive pass. + +The post-campaign restore receipt is also embedded in +`tooling/deployments/gemma4-31b-q4-tower2/final-validation.json`. The restored +OpenClaw config exactly matches its pre-campaign SHA-256, DeepSeek advertises +the proven 1,048,576-token route, and both production agents independently +executed and observed a marker command. Ninety-six campaign sandbox containers +were stopped but not deleted, preserving their recoverability while returning +the production host to the proven DeepSeek stack. diff --git a/benchmarks/gemma4-31b-q4/GEMMA4_31B_Q4_VERIFIED_RESULTS.md b/benchmarks/gemma4-31b-q4/GEMMA4_31B_Q4_VERIFIED_RESULTS.md new file mode 100644 index 00000000..af9a451e --- /dev/null +++ b/benchmarks/gemma4-31b-q4/GEMMA4_31B_Q4_VERIFIED_RESULTS.md @@ -0,0 +1,143 @@ +# Gemma 4 31B QAT Q4_0 on Tower2: verified results + +## Accepted deployment + +- Model: Google's official `google/gemma-4-31B-it-qat-q4_0-gguf`, revision + `59dde24573e7e61570dba08b18a2e1fe246955ed`. +- Text artifact: 17,651,001,568 bytes, SHA-256 + `179cfb99212709597eae5929112cfca677e1bbf566178b479ae1da0c4772874b`. +- Runtime: host-native llama.cpp build 10223 at commit + `11924d4c17abc27383376a1ac6a24fa3e36c1c0c`, compiled for Blackwell + `sm_120a` with CUDA 13.1.115. +- Hardware: 2x RTX PRO 6000 Blackwell Workstation Edition, 97,887 MiB each, + capped at 500 W per GPU. +- Topology: two independent full-offload replicas, one per GPU, ports 8000 and + 8001; four slots and 1,048,576 pooled context tokens per replica, with a hard + 262,144 tokens per slot; Q8_0 KV, flash attention, multimodal projector. +- Sampling: temperature 1.0, top-p 0.95, top-k 64. + +Layer and row split candidates were rejected. Layer split corrupted ordinary +text and tool-call output; row split was unsupported on this PCIe/runtime +combination. Two independent replicas reached 290.279 aggregate decode tok/s +at eight total concurrent requests and keep one GPU failure from taking down +both agents. + +## Canonical 12-family result + +| Cohort | Raw | Corrected | `done_signal` | Median model-call tok/s | Median wall | +|---|---:|---:|---:|---:|---:| +| N=3 | 29/36 (80.6%) | 32/36 (88.9%) | 35/36 | 54.50 | 122.80 s | +| N=10 | 89/120 (74.2%) | 99/120 (82.5%) | 116/120 | 55.85 | 113.05 s | + +The correction changes only the project-management family from 0/10 to 10/10. +The raw grader missed semantically exact statements because it required narrow +contiguous phrases. Every correction is tied to the unchanged grade, report, +workspace archive, grader, and correction-script hashes. No other failure is +reinterpreted. + +N=10 raw/corrected results by family: + +| Family | Raw | Corrected | Main observation | +|---|---:|---:|---| +| Bug fixing | 4/10 | 4/10 | variable; scope and correctness defects | +| Test writing | 8/10 | 8/10 | strong | +| Refactoring | 7/10 | 7/10 | good, with meaningful variance | +| Structured extraction | 10/10 | 10/10 | perfect | +| CI debugging | 10/10 | 10/10 | perfect | +| Adversarial hallucination | 6/10 | 6/10 | four model-stopped missing outputs | +| Support triage | 8/10 | 8/10 | strong but not closed-vocabulary-perfect | +| Document synthesis | 10/10 | 10/10 | perfect | +| Business memo | 6/10 | 6/10 | inconsistent constraint discipline | +| Market research | 10/10 | 10/10 | structural passes; citation quality is separate | +| Writing/editing | 10/10 | 10/10 | perfect under the corrected task rules | +| Project management | 0/10 | 10/10 | ten reproducible lexical grader false negatives | + +## Direct comparison to Qwen3.6-27B + +| Axis | Gemma 4 31B QAT Q4 | Qwen3.6-27B AWQ | Read | +|---|---:|---:|---| +| Comparable N=3 quality | 29/36 raw; 32/36 corrected | 20/36 raw | Gemma wins bounded quality | +| Full-grid ordinary completion | 116/120 | 113/118 no-think | approximately tied; completion is not quality | +| 500 W short-context single stream | 70.3 tok/s | 72.1 tok/s | effectively tied | +| Dense batching | 290.3 tok/s across two GPUs at total C8 | 1,336.5 tok/s on one GPU at C32 | Qwen/vLLM is the serving winner; shapes differ | +| Native tested context | 262,144 | 262,144 | tied | +| Frozen 75-PR strict pass | 0/3 | 0/1 published | neither is reliable at marathon scope | + +Qwen's 113/118 figure is the published no-think `done_signal` rate after two +operator-labeled runs were excluded; its PASS grader sweep was explicitly +pending. It cannot be used as “113 quality passes.” Gemma's 99/120 is a +corrected quality total over all 120 cells. + +The batching row is intentionally not called a controlled model-speed A/B: +Qwen used vLLM with 32 concurrent requests while Gemma used llama.cpp with +four slots per GPU. It is nevertheless the relevant production result: Qwen is +the better high-concurrency serving stack; Gemma is close for one interactive +user and gives each of eight simultaneous slots the full native 256K ceiling. + +## Cross-model position + +| Model | Cohort | Raw | Corrected | Interpretation | +|---|---|---:|---:|---| +| DeepSeek V4 Flash 0731 | N=3 | 23/36 | 35/36 | best corrected bounded result; much faster; 1M context | +| Gemma 4 31B QAT Q4 | N=3 | **29/36** | 32/36 | best raw N=3 result among these pinned local arms | +| Qwen3.6-27B AWQ thinking | N=3 | 20/36 | not available | lower bounded score, stronger batched serving | +| Qwen3-Coder-Next AWQ | N=3 | 20/36 | not available | lower bounded score, much faster MoE decode | +| Gemma 4 31B QAT Q4 | N=10 | 89/120 | **99/120** | broad variance-aware Gemma result | +| Qwen3.5-397B-A17B Q3 no-think | N=10 | 82/120 | 92/120 | directional only; different context, sampling, and date | + +Gemma's 99/120 exceeds Qwen3.5-397B's corrected 92/120, but this is not a +global SOTA proof: the operating points and campaign dates differ, and Gemma's +extended deliverables are weak. DeepSeek remains the evidence-based default +for this Tower2 VRAM profile. + +## Extended-suite audit + +The identity/configuration audit preserved all 12 valid runs plus one proven +infrastructure-invalid supervisor attempt. It fails 12 common provenance checks +because each single-PR artifact omits one or more pinned subject commits. The +separate substantive audit reports **0/12 strict passes**: + +- Single PR: v1/v2 reached the correct MERGE disposition and v3 incorrectly + REJECTed the pinned subject. All three omit required subject refs; v1/v2 also + exposed a historical harness weakness that allowed an unrelated nested + repository tag to satisfy the completion gate. +- Investment memos: all three omit the required PDF. The workbooks are tiny or + static, lack real three-statement/valuation mechanics, and do not support the + stated price targets. v1's internal valuation math is fundamentally + inconsistent; v2/v3 contain zero formulas. +- Board decks: v1/v2 omit the required PDF and have severe overlap/clipping or + sparse/synthetic visuals. v3 includes both formats but remains visually + underdeveloped and carries the unsupported $15 valuation from its source + workbook. +- Frozen 75 PRs: v1 completed 6/75; v2 stopped after two directories and one + verdict; v3 created all 75 directories but only 14 parsable verdicts, four + test-evidence records, 75 shallow reviews, and a dirty final repository. + +The campaign discovered and fixed two evaluation weaknesses without changing +any model output: malformed boolean tool arguments could crash the substance +monitor, and the completion gate accepted any tag in any nested repository. +The monitor now fails safely, and a qualifying completion tag must be annotated, +point at HEAD, and belong to a clean candidate repository. + +## Practical recommendation + +Use Gemma for bounded, high-quality single-user work when a small official QAT +quant, native 256K context, and independent per-GPU replicas are attractive. +Prefer Qwen3.6-27B/vLLM for many simultaneous users. Prefer DeepSeek V4 Flash +for the strongest overall local capability on this exact two-GPU system. For +all three, split marathon work into audited stages; do not treat a `done()` call +or polished-looking artifact as proof of completeness. + +## Post-campaign production restore + +Tower2 was returned to the exact hash-pinned DeepSeek V4 Flash 0731 deployment +after the Gemma campaign. The restored OpenClaw configuration SHA-256 matches +the pre-campaign backup, the launcher and container-image digests match the +restore manifest, both GPUs are capped at 500 W, and the healthy endpoint +advertises a 1,048,576-token maximum model length. + +Fresh OpenClaw checks then made Sanctuary and Pixel each call `exec` exactly +once, observe a unique marker, and return that marker verbatim. Both receipts +identify provider `tower`, model `DeepSeek-V4-Flash-0731`, context 1,048,576, +and `fallbackUsed=false`. The full hashes and cleanup receipt are recorded in +`tooling/deployments/gemma4-31b-q4-tower2/final-validation.json`. diff --git a/benchmarks/gemma4-31b-q4/README.md b/benchmarks/gemma4-31b-q4/README.md new file mode 100644 index 00000000..773133e6 --- /dev/null +++ b/benchmarks/gemma4-31b-q4/README.md @@ -0,0 +1,65 @@ +# Gemma 4 31B QAT Q4_0 on Tower2: verified MMBT campaign + +This entry publishes the complete Gemma 4 31B QAT Q4_0 campaign on two RTX +PRO 6000 Blackwell Workstation Edition GPUs. It separates bounded-task grades, +grader corrections, serving measurements, operational completion, and strict +artifact review. + +## Read order + +1. [`GEMMA4_31B_Q4_VERIFIED_RESULTS.md`](GEMMA4_31B_Q4_VERIFIED_RESULTS.md) + is the results and cross-model interpretation. +2. [`GEMMA4_31B_Q4_COMPLETION_AUDIT.md`](GEMMA4_31B_Q4_COMPLETION_AUDIT.md) + maps campaign requirements to evidence. +3. `gemma4-canonical-n{3,10}-scorecard.json` preserve raw grades and runtime + measurements; the project-management correction files are separate, + hash-tied overlays. +4. `substantive-audit.json` records the independent finance, deck, code, and + 75-PR findings. `gemma4-extended-evidence-audit.json` is the separate + identity/configuration/preservation audit. +5. The reproducible deployment package is under + [`../../tooling/deployments/gemma4-31b-q4-tower2/`](../../tooling/deployments/gemma4-31b-q4-tower2/). + +## Headline results + +- Canonical N=3: **29/36 raw; 32/36 corrected**. +- Canonical N=10: **89/120 raw; 99/120 corrected**. +- Operational completion: **116/120** ordinary runs reached `done_signal`; + four hallucination tasks stopped without the required output. +- Median model-call decode rate: **55.85 tok/s**; median task wall time: + **113.05 seconds**. +- Extended strict substantive result: **0/12**. All evidence is preserved, but + no extended replicate passed the full common and modality-specific gates. + +The corrected overlay changes only ten project-management lexical false +negatives. Original grades remain untouched. No canonical failure is caused by +an artificial output ceiling: the model was served and requested at its native +262,144-token envelope. One server-transport timeout below that envelope was +classified infrastructure-invalid, preserved, and replaced exactly once. + +## Comparison to Qwen3.6-27B + +Gemma is the clear bounded-quality winner in the directly comparable N=3 +matrix: 29/36 raw and 32/36 corrected versus Qwen3.6-27B thinking's 20/36 raw. +The models are nearly tied for short-context single-stream decode at 500 W +(70.3 tok/s Gemma versus 72.1 tok/s Qwen), but Qwen's vLLM stack is vastly +stronger under dense batching. Gemma's accepted two-replica llama.cpp topology +instead prioritizes four independent native-256K slots per GPU and failure +isolation. + +Qwen3.6-27B no-think's published 113/118 is a `done_signal` rate, not a +quality-grade rate, and must not be compared directly to Gemma's 99/120 +corrected quality score. On marathon work, neither model passes the frozen +75-PR standard: Qwen's published result contains 72 template verdicts, while +Gemma's three attempts completed only 6 reviews, 2 directories, and then 75 +directories with just 14 parsable verdicts. + +## Bottom line + +Gemma is a strong dense local model for bounded coding, extraction, CI, +document synthesis, market research, and multi-audience writing. It is not the +best overall model for this 190 GB VRAM system: DeepSeek V4 Flash remains the +stronger default on corrected bounded quality, context, throughput, and +single-PR execution. Gemma also should not be trusted unattended for finance, +presentation QA, or monolithic repository-wide audits without hard artifact +gates and staged review. diff --git a/benchmarks/gemma4-31b-q4/SHA256SUMS b/benchmarks/gemma4-31b-q4/SHA256SUMS new file mode 100644 index 00000000..c2019233 --- /dev/null +++ b/benchmarks/gemma4-31b-q4/SHA256SUMS @@ -0,0 +1,20 @@ +a8943774041bb7b419620c18fd4558e9481f7d755ae664713a0ce25b520a83d8 75pr-v1-audit.json +889686929e81a49956a83aad3ed07692acdef47aa1f9f8e954c52631e5268cb8 75pr-v2-audit.json +71af86f5da85e5be17eaa2b617b91ae0e441f100f9187faf173bcb5dd5ba0823 75pr-v3-audit.json +cffdbe6b71464494a109622bd15ea0a7ca040bb8a9eb40eade9a05653ddc8dce comparison.json +19dff966475723a7c7b91cc18618d5ac5eb3a9c301de4d70c336edcee3114ae4 GEMMA4_31B_Q4_COMPLETION_AUDIT.md +6ad67e72034942653077d3ebf8b27e1874f4e8ffeeda45311ab52ea07e2fbebe GEMMA4_31B_Q4_VERIFIED_RESULTS.md +cf642ac10db70d226fb33424fd302836e074b3e8c1e3ae5926de8a02892ea812 gemma4-canonical-n10-evidence-audit.json +22e71490d63b0e32e2842b08e1d7ed22bdf5fd037ff08913b37cafebc9d1c3bc gemma4-canonical-n10-grader-manifest.json +27feacb45b70c150f979dc88f44cb6ea21743db801f31be1b4b41a0e871e8256 gemma4-canonical-n10-project-mgmt-correction.json +dd0ce822d04f3933c83d9f529cdee45370578483d0a49c40ef0f6c61e01187ac gemma4-canonical-n10-scorecard.json +fc6111dc3cef614eb176641249ac14ddfd05083fb3fc240b6680a3dccecf61f4 gemma4-canonical-n10-scorecard.md +898253e3d9d5683fba6a22fb4cc2b74053ec2c4d4e6d4726cec373e4c5535561 gemma4-canonical-n3-evidence-audit.json +e4b8d4bde4b613a24f976eff2d6b766c96d7962605b8977aefbafb0eb58c2739 gemma4-canonical-n3-grader-manifest.json +2547ae6e777ab018e4ef7163a1c023c805827e627972708b9ae9d3575afc53d5 gemma4-canonical-n3-project-mgmt-correction.json +8d55b13ba959ffebd533ed3f411d613f57772bbc8078b25431a95804c638a8fa gemma4-canonical-n3-scorecard.json +326055d442897265580d090a6e2946a48afa8260050bf09d1a1fe38420fabb8b gemma4-canonical-n3-scorecard.md +5cf402097ab50a900f021da44fdb3e5ac91862653980c673dc21fd53be1d6242 gemma4-extended-evidence-audit.json +8d27c9c1054d0295c99922fe876d248490e826cc1c97bc0b1a47a71433ac6601 post-campaign-restore.json +6a513238c229c75621b30f60016400d50e67facdc350decb4b0e30836e33361a README.md +8b5c16b32f7816eab5edf8c9a96d86aafae07a57a695b0d990d310b6effd00aa substantive-audit.json diff --git a/benchmarks/gemma4-31b-q4/comparison.json b/benchmarks/gemma4-31b-q4/comparison.json new file mode 100644 index 00000000..c2b27ec3 --- /dev/null +++ b/benchmarks/gemma4-31b-q4/comparison.json @@ -0,0 +1,36 @@ +{ + "schema_version": 1, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "canonical": { + "n3": {"raw_passes": 29, "corrected_passes": 32, "total": 36, "done_signal": 35}, + "n10": {"raw_passes": 89, "corrected_passes": 99, "total": 120, "done_signal": 116}, + "correction_scope": "project-management lexical false negatives only" + }, + "qwen3_6_27b": { + "thinking_n3": {"raw_passes": 20, "total": 36}, + "nothink_operational": {"done_signal": 113, "denominator": 118, "quality_grade": null}, + "single_stream_500w_tps": 72.1, + "batch_500w": {"concurrency": 32, "aggregate_tps_one_gpu": 1336.5} + }, + "gemma_serving": { + "single_stream_short_context_tps": 70.28992948492308, + "single_stream_131072_context_tps": 47.26540986218625, + "single_stream_250000_context_tps": 36.69319756850233, + "two_replica_total_concurrency_8_tps": 290.2791998689652 + }, + "extended": { + "strict_substantive_passes": 0, + "total_runs": 12, + "frozen_75pr": [ + {"replicate": 1, "actual_pr_dirs": 6, "parsable_verdicts": 6, "classification": "MODEL_TERMINAL_FAILURE"}, + {"replicate": 2, "actual_pr_dirs": 2, "parsable_verdicts": 1, "classification": "MODEL_TERMINAL_FAILURE"}, + {"replicate": 3, "actual_pr_dirs": 75, "parsable_verdicts": 14, "classification": "MODEL_TERMINAL_FAILURE"} + ] + }, + "comparison_rules": [ + "A done_signal rate is not a quality PASS rate.", + "The Qwen C32 vLLM and Gemma C8 llama.cpp batching rows are operational measurements, not a controlled model-speed A/B.", + "Corrected results never overwrite raw grades.", + "Cross-date and cross-operating-point rates are directional, not a global leaderboard." + ] +} diff --git a/benchmarks/gemma4-31b-q4/gemma4-canonical-n10-evidence-audit.json b/benchmarks/gemma4-31b-q4/gemma4-canonical-n10-evidence-audit.json new file mode 100644 index 00000000..4cb39508 --- /dev/null +++ b/benchmarks/gemma4-31b-q4/gemma4-canonical-n10-evidence-audit.json @@ -0,0 +1,6943 @@ +{ + "schema_version": 1, + "generated_at": "2026-08-02T05:41:32.998976+00:00", + "root": "/home/michael/bench-gemma4-31b-q4", + "label": "gemma4-31b-q4", + "target_n": 10, + "expected_runs": 120, + "audited_runs": 120, + "passed": true, + "errors": [], + "warnings": [ + "p1_bugfix_gemma4-31b-q4_v1: pre-telemetry valid attempt; supplemental telemetry required", + "p1_bugfix_gemma4-31b-q4_v2: pre-telemetry valid attempt; supplemental telemetry required", + "p2_hallucination_gemma4-31b-q4_v3: telemetry coverage below 80%: 0.7716" + ], + "raw_telemetry": { + "path": "/home/michael/gemma4-campaign-state/telemetry/snapshots/canonical-n10.csv", + "bytes": 705889, + "lines": 7129, + "sha256_at_audit_time": "4b187547716d9b2c0efe87aea0edfbf37572de0534dcc0f051c95f5e4a4830b5" + }, + "runs": [ + { + "run_name": "p1_bugfix_gemma4-31b-q4_v1", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "7a1938178dbdde521151e1b1d392fec32d55434c", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 708.7, + "iterations": 80, + "completion_tokens": 18475, + "prompt_tokens_cumulative": 1583932, + "model_turns": 80, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": null, + "files": { + "receipt.json": { + "bytes": 8170, + "sha256": "b2a2b09a244b2461fa4f21616139b6d892c2f300e8a25b4155b5d0df03e318f7" + }, + "transcript.jsonl": { + "bytes": 77185, + "sha256": "72bb0d5cf9474eb1c2851c13001ec8f699ecb587a842962e6ad74328fad657b1" + }, + "summary.json": { + "bytes": 1313, + "sha256": "b97ae41ff92d952c9bbbef50b5534efa2498edb98b6d7d05de89271df613f1d5" + }, + "workspace_final.tar.gz": { + "bytes": 28932897, + "sha256": "4bc1a9558a52c9fb68a3fb2ad1e5748a30de4315f80d6dbcad109cdcfae24a76" + }, + "cost.json": { + "bytes": 1242, + "sha256": "aac1c08a5f940f83536e8dc0e66fd34662b89bd4ba17938de29a976d9d8a8748" + }, + "grade.json": { + "bytes": 467, + "sha256": "8ef51f6f2d9457eb209c41700017a65b6f13674d9251e4c8b02352a7ad89a4aa" + } + } + }, + { + "run_name": "p1_bugfix_gemma4-31b-q4_v2", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "7a1938178dbdde521151e1b1d392fec32d55434c", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 639.5, + "iterations": 58, + "completion_tokens": 14733, + "prompt_tokens_cumulative": 984362, + "model_turns": 58, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": null, + "files": { + "receipt.json": { + "bytes": 8170, + "sha256": "e1afd7f0706ff36f0395a9404de1df3c5875c2d95725d2da24a94f0f2ae0c1a6" + }, + "transcript.jsonl": { + "bytes": 52146, + "sha256": "f6731ff3cdd0af7fab869e689f3b6ded47de098838dc90a2859859e5c15018ba" + }, + "summary.json": { + "bytes": 1999, + "sha256": "576e3927a4e9cf68341324837d68a22d637570306554f829561e1c9298869609" + }, + "workspace_final.tar.gz": { + "bytes": 28949571, + "sha256": "841879c97462eda9cc1612eedc24ee8611eb3656efdda9ee880fea24497d5b57" + }, + "cost.json": { + "bytes": 1240, + "sha256": "cd746dc74c66ffa830be960b7be0d7ef66a7c70c6390f3350b8ca15e695c7423" + }, + "grade.json": { + "bytes": 468, + "sha256": "8c1caead3b6ffa35ac77aefd6150cc3bc5513596a7f125c9e024d34226cdfe3f" + } + } + }, + { + "run_name": "p1_bugfix_gemma4-31b-q4_v3", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 623.7, + "iterations": 62, + "completion_tokens": 14968, + "prompt_tokens_cumulative": 1161167, + "model_turns": 62, + "length_finishes": 0, + "grade_verdict": "FAIL", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9941, + "mean_power_w": 249.36, + "mean_sm_util_pct": 44.38, + "max_temp_c": 76.0, + "cpu_package_mean_power_w": 133.84 + }, + "files": { + "receipt.json": { + "bytes": 8953, + "sha256": "d3a7789767428adc34242dc59125578e4d4445a2ddc5a99b2e78d52718b127c3" + }, + "transcript.jsonl": { + "bytes": 55242, + "sha256": "a2f6e8cec02de3d745fc9894e4051ca328417a0b6edac6966ec9edfd1a8a8766" + }, + "summary.json": { + "bytes": 1514, + "sha256": "9850feb80175867ba25c84877acf3f08d16524b568e2119cfab241bf34c746e7" + }, + "workspace_final.tar.gz": { + "bytes": 28953180, + "sha256": "f07b17077a855af54375939b7c7c65fd0a613a18659456d2f493f4ec7a50fac3" + }, + "cost.json": { + "bytes": 1242, + "sha256": "6d408865f3a1d7aafbaf662eeec0e46bc32ef3baf889a20be018b84c4ec834a4" + }, + "gpu_telemetry.json": { + "bytes": 2450, + "sha256": "b1184513be54d7a0eb39fd12ca6047a6e92888fada6d1acb589380dda3c60ff5" + }, + "grade.json": { + "bytes": 468, + "sha256": "10243cc2a1b0f2c8df720128ec698608433a7b5176a9ffe716efd29205815667" + } + } + }, + { + "run_name": "p1_bugfix_gemma4-31b-q4_v4", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 693.3, + "iterations": 64, + "completion_tokens": 18708, + "prompt_tokens_cumulative": 1257972, + "model_turns": 64, + "length_finishes": 0, + "grade_verdict": "FAIL", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9952, + "mean_power_w": 259.64, + "mean_sm_util_pct": 46.02, + "max_temp_c": 78.0, + "cpu_package_mean_power_w": 129.53 + }, + "files": { + "receipt.json": { + "bytes": 9512, + "sha256": "3b2618506eed716af8ed988fe57b1bd5a0fc8a037d25c2c78d0e12813472ff31" + }, + "transcript.jsonl": { + "bytes": 65703, + "sha256": "1236878bb5d4be70df8070fe2db6190b077df1fece3bd124a0fc58061fba1fbd" + }, + "summary.json": { + "bytes": 733, + "sha256": "47f33f19e0797e6891b22782eba262b84e43614027af50d190a1f415bde24ab5" + }, + "workspace_final.tar.gz": { + "bytes": 28994623, + "sha256": "a3ac5a156774f08434a756b6c5cc98d7b3d79a41e8ae54700995881dbe36aa85" + }, + "cost.json": { + "bytes": 1244, + "sha256": "f030d5cbde1c7ecf380865944a0625ddbbecbefd22821741c1c0335ff76ccf75" + }, + "gpu_telemetry.json": { + "bytes": 2446, + "sha256": "a0a669e4019d4d43e607dd0778801911daa6056fe854d4cff54653aa6f792dbb" + }, + "grade.json": { + "bytes": 468, + "sha256": "2a50cd7db6267876fdbfc16a2bb28a3043d64cdc64ff50bfb627945c77bcfd57" + } + } + }, + { + "run_name": "p1_bugfix_gemma4-31b-q4_v5", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 610.7, + "iterations": 67, + "completion_tokens": 13721, + "prompt_tokens_cumulative": 1168165, + "model_turns": 67, + "length_finishes": 0, + "grade_verdict": "FAIL", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9989, + "mean_power_w": 229.53, + "mean_sm_util_pct": 40.05, + "max_temp_c": 72.0, + "cpu_package_mean_power_w": 128.83 + }, + "files": { + "receipt.json": { + "bytes": 9512, + "sha256": "25e2ad7a3c951ac66e02accd4d60ffe75c64cc638da37080abed00a293c49437" + }, + "transcript.jsonl": { + "bytes": 52274, + "sha256": "91e589c1ad4e37e3299c6a0a253db540bab2f3f812594aa7db737246fe73b300" + }, + "summary.json": { + "bytes": 1616, + "sha256": "4d854ade0c45000db78f96446205579d6a18192ec79d47344f0ce2e6439a3ab2" + }, + "workspace_final.tar.gz": { + "bytes": 28949561, + "sha256": "0f936aeb4a5a99a1490d2ccd84bba7f236fdde0e25896efedb4969e9b331340d" + }, + "cost.json": { + "bytes": 1243, + "sha256": "4914d748cf6419ad9d9d20b9409f995da6a921a800eae120cbd6dbfb287805d6" + }, + "gpu_telemetry.json": { + "bytes": 2407, + "sha256": "a64ecc824b6a752fc36f539989ed7b10bbe41f0615e72bf831c31052f4c53d53" + }, + "grade.json": { + "bytes": 467, + "sha256": "3182de43effbec4148a35605a10fbcf8ec4479c3ece279d600fd34ba81243341" + } + } + }, + { + "run_name": "p1_bugfix_gemma4-31b-q4_v6", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 803.6, + "iterations": 86, + "completion_tokens": 21150, + "prompt_tokens_cumulative": 2088543, + "model_turns": 86, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9893, + "mean_power_w": 291.98, + "mean_sm_util_pct": 52.56, + "max_temp_c": 82.0, + "cpu_package_mean_power_w": 131.15 + }, + "files": { + "receipt.json": { + "bytes": 9514, + "sha256": "b79860d9718c2506cc2c0afae20f96158a76f9285f97f23e26ae2128b215461f" + }, + "transcript.jsonl": { + "bytes": 85679, + "sha256": "9a5a7e9fc89d442035d8aee4e4d211b9bfd326390181363d8af13ad5e3fda97f" + }, + "summary.json": { + "bytes": 1498, + "sha256": "4b5ffd4e44eacadf1fbac2c675a1455d19e03da73709609af503f9e8345b0ded" + }, + "workspace_final.tar.gz": { + "bytes": 28963601, + "sha256": "7f40d019bfb886459e73da64b5800fe0cacdf59287c11526ea26e6d8cbb76432" + }, + "cost.json": { + "bytes": 1244, + "sha256": "3f6190bde275aa654ea5bb634ea59522ac3946f44a5acd87258db2c71b851c92" + }, + "gpu_telemetry.json": { + "bytes": 2445, + "sha256": "3d38a59f062500bb4eea620a95bc72a5758269fdd7a6f799dddf4ae86d453459" + }, + "grade.json": { + "bytes": 478, + "sha256": "d001be623e65b5b8df177bb5bed6ffcf720c20c1b09d9ce5410069011f837512" + } + } + }, + { + "run_name": "p1_bugfix_gemma4-31b-q4_v7", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 757.1, + "iterations": 73, + "completion_tokens": 20875, + "prompt_tokens_cumulative": 1635489, + "model_turns": 73, + "length_finishes": 0, + "grade_verdict": "FAIL", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9906, + "mean_power_w": 290.56, + "mean_sm_util_pct": 51.36, + "max_temp_c": 73.0, + "cpu_package_mean_power_w": 130.97 + }, + "files": { + "receipt.json": { + "bytes": 9516, + "sha256": "c0adc4d3c14965d24e4079aaeba385b720409d90dce4e582131c826212520880" + }, + "transcript.jsonl": { + "bytes": 76479, + "sha256": "3e0b3635ad081dcf2a1977591910931d20840c0f131e0561f54e79898db77b25" + }, + "summary.json": { + "bytes": 1105, + "sha256": "055af8c2f17dec806255c0412fe980306772d136aa3e22a4c65cafbc7c6e8f91" + }, + "workspace_final.tar.gz": { + "bytes": 28990522, + "sha256": "92a4dec8ac1ddbb1e68069cfa5ed86fc5826749f332a3500c82e3b640df2fbb8" + }, + "cost.json": { + "bytes": 1245, + "sha256": "c893a06499879ff2b070809857b518bcb7bf6ea52d7298432351b087f5907dc7" + }, + "gpu_telemetry.json": { + "bytes": 2444, + "sha256": "e9d63502c159c8743efc5b68271c0fe285cdcf85144ffa20d9aa2ea62868f2c2" + }, + "grade.json": { + "bytes": 467, + "sha256": "a98e38961cf492999b6fe0c7dffe92d4f2e9599e4324e27e32799fefe44ee66e" + } + } + }, + { + "run_name": "p1_bugfix_gemma4-31b-q4_v8", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 726.2, + "iterations": 83, + "completion_tokens": 19314, + "prompt_tokens_cumulative": 1725963, + "model_turns": 83, + "length_finishes": 0, + "grade_verdict": "FAIL", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9915, + "mean_power_w": 275.15, + "mean_sm_util_pct": 52.54, + "max_temp_c": 79.0, + "cpu_package_mean_power_w": 131.29 + }, + "files": { + "receipt.json": { + "bytes": 9514, + "sha256": "7a58b2f1f4f223bab105418948d799660b8e31e6c1627a66263a8aa9f856a9a4" + }, + "transcript.jsonl": { + "bytes": 74620, + "sha256": "fe3d81b065801aa5679d28291c4b8cf84bf02fca75dfa070bf73e79f29ee5884" + }, + "summary.json": { + "bytes": 1311, + "sha256": "13408fe290acc190536381392de5ba05a1c75c61af95fe033f486e2198947998" + }, + "workspace_final.tar.gz": { + "bytes": 28974169, + "sha256": "7cf8bd249c29123a665894c4ddbc7712c0f69148dc241a8c57f99a048ebdf723" + }, + "cost.json": { + "bytes": 1245, + "sha256": "15e599201396f325cc686f5895450267dc1e6a4918f50b467e2a414e48e8bd57" + }, + "gpu_telemetry.json": { + "bytes": 2449, + "sha256": "1bae3f97d85cc8cec642e5ff30a2f9fec097a36e510db05ffaa54ea150ded214" + }, + "grade.json": { + "bytes": 469, + "sha256": "76c1029d08602ceb16361a7afaea843a018078ddc51151e7e969d2cdb5a263d9" + } + } + }, + { + "run_name": "p1_bugfix_gemma4-31b-q4_v9", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 756.2, + "iterations": 70, + "completion_tokens": 20369, + "prompt_tokens_cumulative": 1602179, + "model_turns": 70, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9918, + "mean_power_w": 287.22, + "mean_sm_util_pct": 51.15, + "max_temp_c": 73.0, + "cpu_package_mean_power_w": 131.3 + }, + "files": { + "receipt.json": { + "bytes": 9516, + "sha256": "4e47649d9a41bf987bf1f7f5501c013379361ab836e1bab1bc7b0b1355bc09f9" + }, + "transcript.jsonl": { + "bytes": 74982, + "sha256": "32d3e5a334a1b4d95e74a27cecfbfec275b5dc0630fa406b353ce3767defe168" + }, + "summary.json": { + "bytes": 808, + "sha256": "46b38e6b124cd1b4da89acf6e1c74fe106dc85bac65246782eabde685a58772b" + }, + "workspace_final.tar.gz": { + "bytes": 29008725, + "sha256": "b1a9a4c8d8fb596e4a4ac9f87fec582ea5168a6e0d9b603d7fe7ff019626c59c" + }, + "cost.json": { + "bytes": 1244, + "sha256": "c2ea4f43857b2144f913913154f16b7c5310a98cc7c94aa3cf3c81da4289d295" + }, + "gpu_telemetry.json": { + "bytes": 2445, + "sha256": "6409c0a97b5184ec04e2ee4ec87d6ca67176ee51c35c310bfb3ee8c4c9116fa2" + }, + "grade.json": { + "bytes": 480, + "sha256": "aea0e00c9cd364a33af96808e50aab16ed1a942f1b6c59f7bde8079586f7b452" + } + } + }, + { + "run_name": "p1_bugfix_gemma4-31b-q4_v10", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 707.1, + "iterations": 76, + "completion_tokens": 18580, + "prompt_tokens_cumulative": 1567028, + "model_turns": 76, + "length_finishes": 0, + "grade_verdict": "FAIL", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.997, + "mean_power_w": 256.16, + "mean_sm_util_pct": 43.77, + "max_temp_c": 80.0, + "cpu_package_mean_power_w": 132.79 + }, + "files": { + "receipt.json": { + "bytes": 9517, + "sha256": "61a701616e699b2878f1064d11829270864f23b91737323daf3f0240ab0516cc" + }, + "transcript.jsonl": { + "bytes": 69493, + "sha256": "72cef34a8a510fc144b5970f015bcb4d945ca393bf3002bdb907cbb7ea065d11" + }, + "summary.json": { + "bytes": 1330, + "sha256": "8fcd99953679a75838983e9a2277bf16df2e7b1580ee977eaa04e6ff1300f7ee" + }, + "workspace_final.tar.gz": { + "bytes": 29003327, + "sha256": "dc1e419b256cbd32584735481333c7a70e8a7a2c6457b6bafea021209a67c77e" + }, + "cost.json": { + "bytes": 1246, + "sha256": "caf97718fb5b66c784a3d54a24455df6f55074142ada90fe6ad1f30bea3133fe" + }, + "gpu_telemetry.json": { + "bytes": 2495, + "sha256": "6b108ee81d7b8d4fc642101a7fd9a18a7a9e05ad9a243effb84435b80d38a58e" + }, + "grade.json": { + "bytes": 469, + "sha256": "cc4b412a05ac78893698a0ffe05fa4a3519ea42cea3051bdc679b29e472aa0c9" + } + } + }, + { + "run_name": "p1_testwrite_gemma4-31b-q4_v1", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 523.6, + "iterations": 49, + "completion_tokens": 25223, + "prompt_tokens_cumulative": 1273638, + "model_turns": 49, + "length_finishes": 0, + "grade_verdict": "FAIL", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9836, + "mean_power_w": 481.16, + "mean_sm_util_pct": 91.31, + "max_temp_c": 81.0, + "cpu_package_mean_power_w": 133.28 + }, + "files": { + "receipt.json": { + "bytes": 8955, + "sha256": "982dcf5a0e4637072a83ce0a400bea7422122b943e97ce3e9f4f5d091fab143c" + }, + "transcript.jsonl": { + "bytes": 57036, + "sha256": "48fed858bc6c18f032a8e1f89230547b425815156e1afb29af37ddebe63a51bc" + }, + "summary.json": { + "bytes": 1479, + "sha256": "0a9dbf647408093d6503dd606d3725eb6778b66a04c69c51ab8531ac3fcba653" + }, + "workspace_final.tar.gz": { + "bytes": 149484, + "sha256": "a522588ef84bec7cf9edf8da1c090f858fbab32d75c6fa5bfa27c28107172a2a" + }, + "cost.json": { + "bytes": 1242, + "sha256": "034e404e054f1a8a705a4626c2df9372efd9503a5d7c3f25b98851efae7a13fe" + }, + "gpu_telemetry.json": { + "bytes": 2410, + "sha256": "cd51922d479a47d4691872400e5cc13436338a4d745eeb939635ffbe2d4d28d7" + }, + "grade.json": { + "bytes": 1155, + "sha256": "73b7129ab31845b98ef25c4e1501e5367510c8b7fbda84357be4bb3294bbd2f3" + } + } + }, + { + "run_name": "p1_testwrite_gemma4-31b-q4_v2", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 388.8, + "iterations": 48, + "completion_tokens": 18809, + "prompt_tokens_cumulative": 1029851, + "model_turns": 48, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9902, + "mean_power_w": 479.04, + "mean_sm_util_pct": 89.65, + "max_temp_c": 78.0, + "cpu_package_mean_power_w": 136.92 + }, + "files": { + "receipt.json": { + "bytes": 8958, + "sha256": "a88c0da0f29cdfa068f246f94cc94779032defc619e0042bf024aa5fddba632c" + }, + "transcript.jsonl": { + "bytes": 55995, + "sha256": "e29b8ec4f36990ccb224af8f215b357e02c5b9d29e4c544cef96d84717e69116" + }, + "summary.json": { + "bytes": 679, + "sha256": "3ea94ce4164c4c8bdbebf903d38495718e251b94140b9d18aeba137736609520" + }, + "workspace_final.tar.gz": { + "bytes": 138148, + "sha256": "ed2052b32ae52b5832aebc625413764854230bbe23611d931d735874219ab03d" + }, + "cost.json": { + "bytes": 1240, + "sha256": "881968bad40e6940fd25b639c3f4f10b759e11b440176941eea76146e05cf9b6" + }, + "gpu_telemetry.json": { + "bytes": 2449, + "sha256": "5b9798306602f66dad04b2dd7cb2b8b167092139fe5076271b2089d5a939579f" + }, + "grade.json": { + "bytes": 341, + "sha256": "42c1da43e1c31be25775668d65010e502b8383356f35bfd846be7b583f7e153e" + } + } + }, + { + "run_name": "p1_testwrite_gemma4-31b-q4_v3", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 399.9, + "iterations": 42, + "completion_tokens": 19900, + "prompt_tokens_cumulative": 825275, + "model_turns": 42, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9877, + "mean_power_w": 468.52, + "mean_sm_util_pct": 91.96, + "max_temp_c": 86.0, + "cpu_package_mean_power_w": 136.84 + }, + "files": { + "receipt.json": { + "bytes": 8958, + "sha256": "057d735baa76ae75cab8aae2d06c0c10d73a4adcd9dc49790be911cc83eeec94" + }, + "transcript.jsonl": { + "bytes": 50956, + "sha256": "fb8be270e1fb754d0585a92189259eb5a9e14426abe9983ac3c155621dbe5600" + }, + "summary.json": { + "bytes": 1158, + "sha256": "da9c949bb86558868d75aa50347057cf3f65510b908d1b5be4ac309e593d21b0" + }, + "workspace_final.tar.gz": { + "bytes": 140677, + "sha256": "f7e3e9cd2be319d6df53da86252e78481fecf0cc597e4b43efdafbf83d0ce07c" + }, + "cost.json": { + "bytes": 1241, + "sha256": "ba4db6fb45cc046dc3487fcc5ce45b00feae0406373ef7360e73951565a4bdc1" + }, + "gpu_telemetry.json": { + "bytes": 2449, + "sha256": "763efb26f7e3141fcb709a10b9727f95534c849d3c95b49da8e3b5474883d959" + }, + "grade.json": { + "bytes": 341, + "sha256": "b99e9530499f12935a79df989d2f182b0917530a4b3be44b95eaa26e7f2d2c33" + } + } + }, + { + "run_name": "p1_testwrite_gemma4-31b-q4_v4", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 278.0, + "iterations": 38, + "completion_tokens": 14087, + "prompt_tokens_cumulative": 655459, + "model_turns": 38, + "length_finishes": 0, + "grade_verdict": "FAIL", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9892, + "mean_power_w": 471.63, + "mean_sm_util_pct": 88.3, + "max_temp_c": 82.0, + "cpu_package_mean_power_w": 135.87 + }, + "files": { + "receipt.json": { + "bytes": 9518, + "sha256": "cd1c16e7b87bee8282024319533fc5135c6b704c2b677a17c72318afda13eb1c" + }, + "transcript.jsonl": { + "bytes": 38453, + "sha256": "4943d3e87a6353b9519c92179cc2710a02e160ecdbf18971743eeacd15555dcf" + }, + "summary.json": { + "bytes": 1545, + "sha256": "7e9d50fc10e198fd52dba0bc5f78dbc5306bd79f32e79022fca5452fb0d96c53" + }, + "workspace_final.tar.gz": { + "bytes": 102709, + "sha256": "046868308ea7e3f2303f95497a37105218c643122c3514519a134426627342ef" + }, + "cost.json": { + "bytes": 1240, + "sha256": "ddd3fb23037b1159c1554348032e35c6a6e101303d83376be8edd4666d766c7e" + }, + "gpu_telemetry.json": { + "bytes": 2407, + "sha256": "ff19f470fee19f340ea7131e48b58a082781b2351ea17d3e87667ad8b8cfd612" + }, + "grade.json": { + "bytes": 338, + "sha256": "a76068faab9d01220a3493092cadd4a8630aa720dc624df35bbb71b4f2772e05" + } + } + }, + { + "run_name": "p1_testwrite_gemma4-31b-q4_v5", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 333.3, + "iterations": 37, + "completion_tokens": 16936, + "prompt_tokens_cumulative": 648484, + "model_turns": 37, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9751, + "mean_power_w": 480.35, + "mean_sm_util_pct": 90.21, + "max_temp_c": 76.0, + "cpu_package_mean_power_w": 132.81 + }, + "files": { + "receipt.json": { + "bytes": 9518, + "sha256": "05d48e9565c2c99c4329c0b5798f378792b7ec02cc196684ad4bb8389c9ae4f7" + }, + "transcript.jsonl": { + "bytes": 40929, + "sha256": "fdf5747ba761875d0664a00fc731eba62d1d47ee2467852c478159ee11b000c8" + }, + "summary.json": { + "bytes": 1439, + "sha256": "17b5c358c97fde97a8d07d390b5197eed0c186405e21c4864ff0bcf3cfaf7285" + }, + "workspace_final.tar.gz": { + "bytes": 120006, + "sha256": "e157910534a0fa23f2f4e3de9a8c384ca911998a4ee34358e4d24b804c76f51d" + }, + "cost.json": { + "bytes": 1239, + "sha256": "d5743d3301712bc723f4965a5eaee207531dafb65d57b7c08c687b8e5be8a12c" + }, + "gpu_telemetry.json": { + "bytes": 2447, + "sha256": "7ffde2aef644a3bca84c3876bd5e395400cd8fce899e329e006e0ec82e795bea" + }, + "grade.json": { + "bytes": 341, + "sha256": "74243a504afb05edd995955710a388f8e221516f754407850844113c4c26e3db" + } + } + }, + { + "run_name": "p1_testwrite_gemma4-31b-q4_v6", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 464.7, + "iterations": 37, + "completion_tokens": 22810, + "prompt_tokens_cumulative": 818645, + "model_turns": 37, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9899, + "mean_power_w": 481.71, + "mean_sm_util_pct": 92.33, + "max_temp_c": 82.0, + "cpu_package_mean_power_w": 136.01 + }, + "files": { + "receipt.json": { + "bytes": 9518, + "sha256": "f5248e903407eca717cd8edf0271461da713285b4e5997e526987df16546230c" + }, + "transcript.jsonl": { + "bytes": 55941, + "sha256": "ff9995a26d187672045030e699e452cd567c0e703473676c91f746df40b579ed" + }, + "summary.json": { + "bytes": 778, + "sha256": "eeaa480a6b5284331bf97e7ab087612fcc6adfe55228b7b5d784ecaf22394496" + }, + "workspace_final.tar.gz": { + "bytes": 168769, + "sha256": "81b4fc223a4b4b4e78a87f2721ebe4feac477de09f645053ab8f6d5f6f656bfd" + }, + "cost.json": { + "bytes": 1241, + "sha256": "1669e9a9e9971537953625d43d6832018f17323480a755f30db66df7f2dbaf0e" + }, + "gpu_telemetry.json": { + "bytes": 2452, + "sha256": "8699dc5b752a23eddbcdbbeedde79ae101cd4aeb7c4ec844f5f0ba40e1aef174" + }, + "grade.json": { + "bytes": 341, + "sha256": "eeff0a8acc36edbcdc0b38e3fac798d8ec73d59bca53d8dc0664c4a8abf832ce" + } + } + }, + { + "run_name": "p1_testwrite_gemma4-31b-q4_v7", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 417.0, + "iterations": 44, + "completion_tokens": 20892, + "prompt_tokens_cumulative": 856122, + "model_turns": 44, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9952, + "mean_power_w": 479.92, + "mean_sm_util_pct": 90.1, + "max_temp_c": 74.0, + "cpu_package_mean_power_w": 132.97 + }, + "files": { + "receipt.json": { + "bytes": 9516, + "sha256": "0110dbca5edb8062906fd9a23fcf477aa2f33232df3b2857f3469f65ee5f28c3" + }, + "transcript.jsonl": { + "bytes": 58437, + "sha256": "612ea848e2c1d1ac87823c8ef7a3735180571594a6466f1d7ab1596b1f5deb51" + }, + "summary.json": { + "bytes": 926, + "sha256": "0fdad92c81d9d57236e7536599639095643551b7767e82a8b6815c88ad0de52f" + }, + "workspace_final.tar.gz": { + "bytes": 164591, + "sha256": "aca45682aa0251afee88e589f270ccc2a1b921ebc1c37f2ef56d698426bd0dad" + }, + "cost.json": { + "bytes": 1240, + "sha256": "0666dcc703cc701e2957b627ee504aa2e3f8e638b11af7468108873e01f5ad86" + }, + "gpu_telemetry.json": { + "bytes": 2405, + "sha256": "f48ddc94f85d305f3f7e8fd8861415bf8a4f5ad59e877f3745526b2cbb3fa90b" + }, + "grade.json": { + "bytes": 341, + "sha256": "0f83751c074427d12d916faba8ffd786e54d8feb1bd8d37ef8e6ada4128ea5a3" + } + } + }, + { + "run_name": "p1_testwrite_gemma4-31b-q4_v8", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 464.0, + "iterations": 50, + "completion_tokens": 22263, + "prompt_tokens_cumulative": 1361395, + "model_turns": 50, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9914, + "mean_power_w": 474.5, + "mean_sm_util_pct": 91.35, + "max_temp_c": 82.0, + "cpu_package_mean_power_w": 136.04 + }, + "files": { + "receipt.json": { + "bytes": 9518, + "sha256": "9a87cecce5c12316c25ea5edd80a2bf43977bcd6818e00179a6dfb5cb42a42db" + }, + "transcript.jsonl": { + "bytes": 53604, + "sha256": "505e766417ae1580b33c9d068af0ae79206313155453e372b0461454da8a9a7c" + }, + "summary.json": { + "bytes": 782, + "sha256": "8b4e625e537358f313406602d424936549a7bd46b6db7ca3ed49a7f5733eb521" + }, + "workspace_final.tar.gz": { + "bytes": 177410, + "sha256": "6cc58fe633ab2353d2054380b063614ab667a6d28a131a5dd4201349aa8d17b6" + }, + "cost.json": { + "bytes": 1243, + "sha256": "342f498d0feaa2f6bf4acf6a14d789dc5bf3ba5145da59b15f10191aeb591468" + }, + "gpu_telemetry.json": { + "bytes": 2490, + "sha256": "2bcc8bffb95e6bacbac9e8faaf574eaba35d15bacf6ea2a316a93fc4be307ddc" + }, + "grade.json": { + "bytes": 341, + "sha256": "cf4fb96563994dcd182edffa7fb117b91981086cef9907b87d1de3afa262b6dc" + } + } + }, + { + "run_name": "p1_testwrite_gemma4-31b-q4_v9", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 684.1, + "iterations": 54, + "completion_tokens": 30030, + "prompt_tokens_cumulative": 1728443, + "model_turns": 54, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.994, + "mean_power_w": 487.82, + "mean_sm_util_pct": 93.58, + "max_temp_c": 75.0, + "cpu_package_mean_power_w": 135.9 + }, + "files": { + "receipt.json": { + "bytes": 9518, + "sha256": "8d52aa908a7e8b57bd949a0d30b10b23201ce2ca4a82a19fd2afed5ead01ed47" + }, + "transcript.jsonl": { + "bytes": 78888, + "sha256": "99e27fc11b470da47615c842ccede524f9518152c77c56d5fa6ecc3f414058d1" + }, + "summary.json": { + "bytes": 635, + "sha256": "58b5eddc7fe29000551603791711396ab58e714cde661843aa3644d5077d169e" + }, + "workspace_final.tar.gz": { + "bytes": 168901, + "sha256": "ffe243f7b2a16a422a3fcf4a85c672eaba63d8c0d61cb59dc92dbfb051f5b3b6" + }, + "cost.json": { + "bytes": 1242, + "sha256": "62ac47eb30eaa22d41d96cb8ede4c1cc78c45de787037fd1619b0bb6bdafeede" + }, + "gpu_telemetry.json": { + "bytes": 2495, + "sha256": "4ff3947e00f9b557c6d3d60ad93560cfcbec9940bc880ba6323a4212eced535a" + }, + "grade.json": { + "bytes": 341, + "sha256": "bc0138be4f3c21f7687420af1fa496c244e735495742d7b09d1ba0d6818add51" + } + } + }, + { + "run_name": "p1_testwrite_gemma4-31b-q4_v10", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 380.5, + "iterations": 53, + "completion_tokens": 17822, + "prompt_tokens_cumulative": 1323034, + "model_turns": 53, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9855, + "mean_power_w": 466.52, + "mean_sm_util_pct": 89.03, + "max_temp_c": 83.0, + "cpu_package_mean_power_w": 136.27 + }, + "files": { + "receipt.json": { + "bytes": 9519, + "sha256": "45ca75c65d240a11f04624e8de2eb2a2a7be0e472a84e59061551b7ac70d5cd7" + }, + "transcript.jsonl": { + "bytes": 49948, + "sha256": "4a2e0429e61ed1e09b431edc528604144b625e481848de97d98c47e0daab56dc" + }, + "summary.json": { + "bytes": 1000, + "sha256": "976e48d857c8689b94783ea65899d18af27b6d92616f9087e0ed4da9421ecc67" + }, + "workspace_final.tar.gz": { + "bytes": 116792, + "sha256": "c71e04b035d7d1dcffd504d2ca953340b8111ea18182329b1fff92d338fe6450" + }, + "cost.json": { + "bytes": 1242, + "sha256": "512e526314bd3e3bee8d1f2d61333823be0f76ff8669f424a5c3ebec323bc56a" + }, + "gpu_telemetry.json": { + "bytes": 2602, + "sha256": "26558934312e50300ded56e9721b4f3870da701c1f52d0f05754767288335c6b" + }, + "grade.json": { + "bytes": 341, + "sha256": "6ae706d907c18507866cea3b27f3f46bb025463558df00847cfcf905f4b50e77" + } + } + }, + { + "run_name": "p1_refactor_gemma4-31b-q4_v1", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 186.2, + "iterations": 35, + "completion_tokens": 9779, + "prompt_tokens_cumulative": 351967, + "model_turns": 35, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9667, + "mean_power_w": 465.36, + "mean_sm_util_pct": 85.35, + "max_temp_c": 78.0, + "cpu_package_mean_power_w": 137.0 + }, + "files": { + "receipt.json": { + "bytes": 8956, + "sha256": "c898eedb4993578fca466cb12e19f45091ab32aec824d1ba8bce1e2e72641454" + }, + "transcript.jsonl": { + "bytes": 25468, + "sha256": "945b7feb6c271ad743d914462ba601146de308ce04b3c22696a9938e35467918" + }, + "summary.json": { + "bytes": 783, + "sha256": "fd3a64637a800835a7f5d09dec5d6f1f2bb82aabfa540481a9af9727d745e1b1" + }, + "workspace_final.tar.gz": { + "bytes": 74100, + "sha256": "bd8d0ac1c21d887b949eb9d97fe36673933f5fe012fb8bad4657a59873722816" + }, + "cost.json": { + "bytes": 1240, + "sha256": "66f111f744443b50626b3050b4775796a6471a968811ca4d439f739bd2059384" + }, + "gpu_telemetry.json": { + "bytes": 2447, + "sha256": "1b902ac1bfe246b362021bf602c148c5137d51f0760c636942fc34fe99850e42" + }, + "grade.json": { + "bytes": 510, + "sha256": "7f69334378ac3187de489eb74ad81c8245267d1a5711612b94879f297e758bcf" + } + } + }, + { + "run_name": "p1_refactor_gemma4-31b-q4_v2", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 211.5, + "iterations": 41, + "completion_tokens": 11126, + "prompt_tokens_cumulative": 508191, + "model_turns": 41, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9693, + "mean_power_w": 459.06, + "mean_sm_util_pct": 89.05, + "max_temp_c": 84.0, + "cpu_package_mean_power_w": 137.01 + }, + "files": { + "receipt.json": { + "bytes": 8956, + "sha256": "a8c89f90890e9acacec48b8c73ede92286ffd047bdc5d52a7bace17a18d6930b" + }, + "transcript.jsonl": { + "bytes": 26507, + "sha256": "aa1d445c65f35cdba5254bc3c5de40eea9b75df89e5552bacc7bc9b94f8d7855" + }, + "summary.json": { + "bytes": 730, + "sha256": "6a11f94056186ebdb47360977f854063250f81ac92ea572868fd747efd5daf38" + }, + "workspace_final.tar.gz": { + "bytes": 75807, + "sha256": "49cee0476af144ba1185102887ffedd0141488e34acdff629800bef941cec037" + }, + "cost.json": { + "bytes": 1242, + "sha256": "eee7bd5e349aab1f394356abc6c67b8ec8b09d442367a7afd0ab1af554a802ef" + }, + "gpu_telemetry.json": { + "bytes": 2448, + "sha256": "deb0fecbae839bcc59a6f5b40048b3b634299be001a5a58916e5f94b1ef2af23" + }, + "grade.json": { + "bytes": 510, + "sha256": "7f69334378ac3187de489eb74ad81c8245267d1a5711612b94879f297e758bcf" + } + } + }, + { + "run_name": "p1_refactor_gemma4-31b-q4_v3", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "1a4a954715bf7b228a2d5496fa87547318be6ba8", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 198.2, + "iterations": 40, + "completion_tokens": 10797, + "prompt_tokens_cumulative": 399232, + "model_turns": 40, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9586, + "mean_power_w": 471.53, + "mean_sm_util_pct": 87.9, + "max_temp_c": 71.0, + "cpu_package_mean_power_w": 112.13 + }, + "files": { + "receipt.json": { + "bytes": 9512, + "sha256": "d8b8de86c6abb82b39c59fd2ed56571dc1c05e4c1cf2214fce1a54803bb20dca" + }, + "transcript.jsonl": { + "bytes": 26550, + "sha256": "5ac6ee6b8399781a2394c9e663e4ac0c0850655b6ccb4d80d8ad307e6b62d3ab" + }, + "summary.json": { + "bytes": 893, + "sha256": "e22036d8e57d95f93243ae7cd3f22cfc271686c8134c715dfc44b7459410fe61" + }, + "workspace_final.tar.gz": { + "bytes": 66503, + "sha256": "427d4a3a2eaef5eb3ed46b025996fdbc51fce966e234f63855e4542984c3bd1e" + }, + "cost.json": { + "bytes": 1241, + "sha256": "b296c66ff2312ad59757460488b509e23770bf1e6d79e3bee11d7a11110770cf" + }, + "gpu_telemetry.json": { + "bytes": 2400, + "sha256": "bbfae7e3e4e38b35aaa4dff7b695f50f215d09a5d9b2241d043fe2e5c4de6b0b" + }, + "grade.json": { + "bytes": 510, + "sha256": "7f69334378ac3187de489eb74ad81c8245267d1a5711612b94879f297e758bcf" + } + } + }, + { + "run_name": "p1_refactor_gemma4-31b-q4_v4", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 167.7, + "iterations": 33, + "completion_tokens": 8832, + "prompt_tokens_cumulative": 322475, + "model_turns": 33, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9839, + "mean_power_w": 451.04, + "mean_sm_util_pct": 82.21, + "max_temp_c": 81.0, + "cpu_package_mean_power_w": 136.29 + }, + "files": { + "receipt.json": { + "bytes": 9516, + "sha256": "ab75fdf0b131627d4ed0455451cb2fdf1b3cf7ef3b50e5967a5e739d60d3f698" + }, + "transcript.jsonl": { + "bytes": 24413, + "sha256": "5f2ef7a132cfa48b5e7a690a7bf19cc3d53749beefd09d39fecf3253f85f4554" + }, + "summary.json": { + "bytes": 839, + "sha256": "abb5fe9a6468ee7dac0c1b5b19e12f8a9cf82339a92e8deab53655e8b4840cdc" + }, + "workspace_final.tar.gz": { + "bytes": 117755, + "sha256": "6a6f721c551febfeb28317ab90e98816575ed6f9059dd072c54c7f0f088f8e6c" + }, + "cost.json": { + "bytes": 1239, + "sha256": "2b1e14be2942543650afa7f52cb0cca2ce228750493953bc7931137e7ead3a4a" + }, + "gpu_telemetry.json": { + "bytes": 2484, + "sha256": "cb536aa23accb167b6744cf70fd982facd6da13dbb0794ab36db4d9a88def895" + }, + "grade.json": { + "bytes": 510, + "sha256": "7f69334378ac3187de489eb74ad81c8245267d1a5711612b94879f297e758bcf" + } + } + }, + { + "run_name": "p1_refactor_gemma4-31b-q4_v5", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 169.1, + "iterations": 33, + "completion_tokens": 8879, + "prompt_tokens_cumulative": 300783, + "model_turns": 33, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9758, + "mean_power_w": 462.68, + "mean_sm_util_pct": 88.91, + "max_temp_c": 75.0, + "cpu_package_mean_power_w": 136.22 + }, + "files": { + "receipt.json": { + "bytes": 9516, + "sha256": "ca5effc81e1f265015531f463cff87e9cbf9214a7198668362b848fb9d4a4c34" + }, + "transcript.jsonl": { + "bytes": 23805, + "sha256": "4cd49236227b9ae8f6e0f27228199b19d14952ec6a8fbf3ca1d630fed4e5b662" + }, + "summary.json": { + "bytes": 1211, + "sha256": "6eed4ff93f02a64433d996dd8b844e3241ec6040d00f2d6e963566e6423e92bf" + }, + "workspace_final.tar.gz": { + "bytes": 115762, + "sha256": "824caba93fc4fa05b3b4a76f7e5f318b65b1ad114cdd25d7b3aae5f8fa79488d" + }, + "cost.json": { + "bytes": 1240, + "sha256": "15bdcdc8c2e751ce8a919ee8604c1146f77500dc83c04ba904fdda6c257a4293" + }, + "gpu_telemetry.json": { + "bytes": 2451, + "sha256": "de0d444ad4411fc4e4ebebd15fe0d46d33b88bc2f970cfc887e5fdfdea42d912" + }, + "grade.json": { + "bytes": 510, + "sha256": "7f69334378ac3187de489eb74ad81c8245267d1a5711612b94879f297e758bcf" + } + } + }, + { + "run_name": "p1_refactor_gemma4-31b-q4_v6", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 175.1, + "iterations": 33, + "completion_tokens": 9583, + "prompt_tokens_cumulative": 316151, + "model_turns": 33, + "length_finishes": 0, + "grade_verdict": "FAIL", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9709, + "mean_power_w": 453.36, + "mean_sm_util_pct": 83.6, + "max_temp_c": 79.0, + "cpu_package_mean_power_w": 135.77 + }, + "files": { + "receipt.json": { + "bytes": 9516, + "sha256": "4e02324be31460f6ee18db9e9cb04837be4c8cfe7f2784461de453476881f7d0" + }, + "transcript.jsonl": { + "bytes": 23284, + "sha256": "b4bc1665b4e79978ddd365c68affad87c6131421e7b00a4815115d2e0fa18a7d" + }, + "summary.json": { + "bytes": 1116, + "sha256": "e37e8efedaf90935203c9478e5c0d320a7eb88a87622fcb8b634f12de4bcdcc8" + }, + "workspace_final.tar.gz": { + "bytes": 115785, + "sha256": "3503066d4e2c98c1baa7a2faab9d36e7a733d36e162eb4f0b668623c6c48aa0e" + }, + "cost.json": { + "bytes": 1240, + "sha256": "2b34b74865fa52db686bb4b7158839558831e5205a580d9cd3d572a87ec260bb" + }, + "gpu_telemetry.json": { + "bytes": 2540, + "sha256": "0671caaf431338098155eef4b190e5ffc494d6b8c202fc7b360301be1c7c88b1" + }, + "grade.json": { + "bytes": 529, + "sha256": "3b343e7e30b0fd413c183bcf25114cab48978d1928374717ad6eba1cc327e73d" + } + } + }, + { + "run_name": "p1_refactor_gemma4-31b-q4_v7", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 206.4, + "iterations": 42, + "completion_tokens": 10802, + "prompt_tokens_cumulative": 514218, + "model_turns": 42, + "length_finishes": 0, + "grade_verdict": "FAIL", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.969, + "mean_power_w": 473.97, + "mean_sm_util_pct": 84.59, + "max_temp_c": 77.0, + "cpu_package_mean_power_w": 135.96 + }, + "files": { + "receipt.json": { + "bytes": 9516, + "sha256": "939840433b939a3d7bf3611d0ba26083f1710d0b75b135adb07777290a8dcfca" + }, + "transcript.jsonl": { + "bytes": 27504, + "sha256": "e74df56a2a44153f8230ad053923e953a089974370f44307c8f7df342dbdbc16" + }, + "summary.json": { + "bytes": 754, + "sha256": "892ae40ea22fe1c1b20b75c398e86c7c62e8a12de6031b32a140df75ecbfeb60" + }, + "workspace_final.tar.gz": { + "bytes": 88130, + "sha256": "19bff19b3aea960735adc5a41f4099cf5329b611595b230116feef97074501be" + }, + "cost.json": { + "bytes": 1241, + "sha256": "0648a42d0c58695f54f65e5ecf5a60acbbc1cd56761711618da8f36677576f5a" + }, + "gpu_telemetry.json": { + "bytes": 2406, + "sha256": "2aa4175543e6407074355a1fdbbf4db957a6e26c3c27a483b676d1c395be03b2" + }, + "grade.json": { + "bytes": 529, + "sha256": "3b343e7e30b0fd413c183bcf25114cab48978d1928374717ad6eba1cc327e73d" + } + } + }, + { + "run_name": "p1_refactor_gemma4-31b-q4_v8", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 238.7, + "iterations": 35, + "completion_tokens": 12046, + "prompt_tokens_cumulative": 404131, + "model_turns": 35, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9845, + "mean_power_w": 460.08, + "mean_sm_util_pct": 88.42, + "max_temp_c": 83.0, + "cpu_package_mean_power_w": 136.23 + }, + "files": { + "receipt.json": { + "bytes": 9516, + "sha256": "dadec836c6faee14f4728cda2d7a73d6f7109d4236f078b1bab3ea8e0d84552d" + }, + "transcript.jsonl": { + "bytes": 30846, + "sha256": "8a4efda76542ebac76d0b85d89c776dfaed8f88e6fbdb08b6b15636df8bf7a5e" + }, + "summary.json": { + "bytes": 766, + "sha256": "2e46e348da74ae11fcdab6c4215ba3eaa2769b4a93b2d8c52cf0fce40df9ab83" + }, + "workspace_final.tar.gz": { + "bytes": 64576, + "sha256": "25d87a3f2d2b5f5800b8fa4debb08728f01e42ffcfb3c19f5c1dfa0ac85f9649" + }, + "cost.json": { + "bytes": 1241, + "sha256": "792bbc6787a2ee9c02175840f33bc88ed5b22cbc742f14c59aa633628a833eeb" + }, + "gpu_telemetry.json": { + "bytes": 2483, + "sha256": "649d14a97aa1234c5494f9ffe0b8f48c3c36b6934ea5027d6af201f7c979eaf4" + }, + "grade.json": { + "bytes": 510, + "sha256": "7f69334378ac3187de489eb74ad81c8245267d1a5711612b94879f297e758bcf" + } + } + }, + { + "run_name": "p1_refactor_gemma4-31b-q4_v9", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 212.1, + "iterations": 39, + "completion_tokens": 11193, + "prompt_tokens_cumulative": 483466, + "model_turns": 39, + "length_finishes": 0, + "grade_verdict": "FAIL", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9901, + "mean_power_w": 460.34, + "mean_sm_util_pct": 83.19, + "max_temp_c": 74.0, + "cpu_package_mean_power_w": 136.09 + }, + "files": { + "receipt.json": { + "bytes": 9516, + "sha256": "aada3649b1150a18c963e1d853ef0391430191c43e4831bf92f5b9569a2cb66b" + }, + "transcript.jsonl": { + "bytes": 26774, + "sha256": "36700de84783427d78876a6b792a1fee2c815d3eb5dc03fac347b634bef03cb2" + }, + "summary.json": { + "bytes": 1218, + "sha256": "0907c353aec95a484432e8e41f966e32fd3ff4bb33598cfca99df29356e8dd80" + }, + "workspace_final.tar.gz": { + "bytes": 62391, + "sha256": "6ad479d2668f9c53165c5235c689b14b7a087e8e179e0b39f7e56a4281b57cb8" + }, + "cost.json": { + "bytes": 1242, + "sha256": "6613437aa224603038020fbce3d54292fcc0dc8bc714e3a8cd94bf7eb3b9490a" + }, + "gpu_telemetry.json": { + "bytes": 2451, + "sha256": "61a3fde6e89a5d26808ed94fd4c07de788603006649bd9b2820c90c73ccabd77" + }, + "grade.json": { + "bytes": 411, + "sha256": "fca70add1343783f2d9a6e3bc82d5ae1e3c2ad48a00cce234501fd04b15cbe4f" + } + } + }, + { + "run_name": "p1_refactor_gemma4-31b-q4_v10", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 171.7, + "iterations": 35, + "completion_tokens": 9197, + "prompt_tokens_cumulative": 360546, + "model_turns": 35, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.961, + "mean_power_w": 452.56, + "mean_sm_util_pct": 90.79, + "max_temp_c": 80.0, + "cpu_package_mean_power_w": 136.25 + }, + "files": { + "receipt.json": { + "bytes": 9517, + "sha256": "5c318e75b27ec1309169496f33b511e5affb1ae22f2318916a98ab1248130f8c" + }, + "transcript.jsonl": { + "bytes": 24445, + "sha256": "03e93c7a534a29fb362c9500b054c9d6fb3c76eb9a95aa08248a9e0bebaed045" + }, + "summary.json": { + "bytes": 820, + "sha256": "e4a5e471b1453254deb32a2f2a180f91d998e4d75a89c69895c932afeaeb4050" + }, + "workspace_final.tar.gz": { + "bytes": 90930, + "sha256": "ed028dff7901d3864a78d66ab03d87f046bbd30f72630cfea669073a99524fb1" + }, + "cost.json": { + "bytes": 1241, + "sha256": "62c63a34379d31313eeedc6dd45c0905e5bcf60bc54916b6098b0e7619e04a72" + }, + "gpu_telemetry.json": { + "bytes": 2479, + "sha256": "4c3431673ce33138f72cb13c69ac17bf155beed41f5fce7c3285f42a63342c4e" + }, + "grade.json": { + "bytes": 510, + "sha256": "7f69334378ac3187de489eb74ad81c8245267d1a5711612b94879f297e758bcf" + } + } + }, + { + "run_name": "p2_extract_gemma4-31b-q4_v1", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 76.4, + "iterations": 6, + "completion_tokens": 4449, + "prompt_tokens_cumulative": 22299, + "model_turns": 6, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9162, + "mean_power_w": 460.29, + "mean_sm_util_pct": 91.33, + "max_temp_c": 83.0, + "cpu_package_mean_power_w": 136.95 + }, + "files": { + "receipt.json": { + "bytes": 8954, + "sha256": "f4c6d7cec85a2b5e5b05cbfeeb2311e06b0c9392ec9ef3e3623963981c780b31" + }, + "transcript.jsonl": { + "bytes": 4462, + "sha256": "32ca9ee6cf929cd5d153e7f4d7e52d4dd120dc9e99a780b706f7ffeacc2de67c" + }, + "summary.json": { + "bytes": 530, + "sha256": "c5e97fda27f744cb1368b05e5a6faa3a166ee12264d47bc8e96759fcfb48e251" + }, + "workspace_final.tar.gz": { + "bytes": 11518, + "sha256": "65b92ede9b5f084ca706d4b201f5861fdd4c15d04e7211b5369607b5b9e4013a" + }, + "cost.json": { + "bytes": 1233, + "sha256": "e18fb448f9dcea19dadcac89dce493e73b58e678e779bfadf9fd77285b890fda" + }, + "gpu_telemetry.json": { + "bytes": 2445, + "sha256": "8c2f0709ea92c80b3f49ec83042554a7691ba334c0e499621b451878a88f74c9" + }, + "grade.json": { + "bytes": 3204, + "sha256": "b6bc97fc6e03c0b85c34d7e6b5aa1ffcfdf128e953247bc7569f9417170aa2b8" + } + } + }, + { + "run_name": "p2_extract_gemma4-31b-q4_v2", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 72.6, + "iterations": 6, + "completion_tokens": 4449, + "prompt_tokens_cumulative": 22299, + "model_turns": 6, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.8953, + "mean_power_w": 493.67, + "mean_sm_util_pct": 90.86, + "max_temp_c": 75.0, + "cpu_package_mean_power_w": 113.55 + }, + "files": { + "receipt.json": { + "bytes": 8952, + "sha256": "ce495248feeffe85113c7293821946f22107710355182b97cabc340902bf20e9" + }, + "transcript.jsonl": { + "bytes": 4461, + "sha256": "d3ecc25456ffc3638644f9ae4b60e1e7dcc53b77d3d30d85097d89bafc5c9bff" + }, + "summary.json": { + "bytes": 530, + "sha256": "06fbc403a3d9e561ea45f9bed30c58fd7f7c9dc8ba46c7c0e9fa798255f6a633" + }, + "workspace_final.tar.gz": { + "bytes": 11519, + "sha256": "f9eecea3446b9e0c0306a8fce7d0e79eb3fab3f5d33c27a44dfed5009b05fa2d" + }, + "cost.json": { + "bytes": 1233, + "sha256": "b7eaa3addf58279ba31c7eddbed435dd33430501fc64b3e745c2398b292bf976" + }, + "gpu_telemetry.json": { + "bytes": 2398, + "sha256": "322b12b0d8c0d7e8ac5303cc6951f84dc73ba440a8c058a933889312e2708d6e" + }, + "grade.json": { + "bytes": 3204, + "sha256": "b6bc97fc6e03c0b85c34d7e6b5aa1ffcfdf128e953247bc7569f9417170aa2b8" + } + } + }, + { + "run_name": "p2_extract_gemma4-31b-q4_v3", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 73.9, + "iterations": 6, + "completion_tokens": 4449, + "prompt_tokens_cumulative": 22299, + "model_turns": 6, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.8796, + "mean_power_w": 489.97, + "mean_sm_util_pct": 97.5, + "max_temp_c": 83.0, + "cpu_package_mean_power_w": 136.83 + }, + "files": { + "receipt.json": { + "bytes": 8954, + "sha256": "48090475a70a17e3285430ffb43d146eac293b9e056705af2ef36e9eb795d0e3" + }, + "transcript.jsonl": { + "bytes": 4461, + "sha256": "f4f8cd64f8bd2c10501cedad1365accd1846bc0a0e9c1f346d922c10c5664646" + }, + "summary.json": { + "bytes": 530, + "sha256": "05ae675a95053a282a2ec46decbd7dadc3433d5a663a298a82472d91daae96b5" + }, + "workspace_final.tar.gz": { + "bytes": 11512, + "sha256": "6f39c6c22c3cfa9114f079a566bf355f9c5d9123a6e33b836d24729c059bd962" + }, + "cost.json": { + "bytes": 1233, + "sha256": "0521bb3ebf24dcb019320cb7f8e968a73b0affbe8e9948cdd793a0cdc6e3f1be" + }, + "gpu_telemetry.json": { + "bytes": 2400, + "sha256": "e4ffb48249ad7a39cf662829d3508585d54b6cd9e3f47c3ee37272d2dacc2a94" + }, + "grade.json": { + "bytes": 3204, + "sha256": "b6bc97fc6e03c0b85c34d7e6b5aa1ffcfdf128e953247bc7569f9417170aa2b8" + } + } + }, + { + "run_name": "p2_extract_gemma4-31b-q4_v4", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 75.7, + "iterations": 6, + "completion_tokens": 4449, + "prompt_tokens_cumulative": 22299, + "model_turns": 6, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9247, + "mean_power_w": 473.25, + "mean_sm_util_pct": 93.13, + "max_temp_c": 80.0, + "cpu_package_mean_power_w": 135.91 + }, + "files": { + "receipt.json": { + "bytes": 9514, + "sha256": "532861e05dea3a587e0b8ea099fffcba26e01c7ffb1d83e378a9cd9de8af5b37" + }, + "transcript.jsonl": { + "bytes": 4461, + "sha256": "e0b01c3d3c78bf43fb3a7de0342c15287f6b0f9d58c49718d36c611746802b97" + }, + "summary.json": { + "bytes": 530, + "sha256": "d6cae643d06279bea9de67d4dc4f3f4da92ca2700efdea3b542da6b7fc11172f" + }, + "workspace_final.tar.gz": { + "bytes": 11513, + "sha256": "d5f70ffae21c81594a7d036f76e01f33234546cba8f0d94fdde3511744e5a7f9" + }, + "cost.json": { + "bytes": 1233, + "sha256": "68393905b13dad088552e3261c8653bb0d3b71d1f65efdee5c1224212b342223" + }, + "gpu_telemetry.json": { + "bytes": 2398, + "sha256": "c4ce58e7d13e45dded76ade087c8b2d5841dc307f0927e677023117b9fa01c6e" + }, + "grade.json": { + "bytes": 3204, + "sha256": "b6bc97fc6e03c0b85c34d7e6b5aa1ffcfdf128e953247bc7569f9417170aa2b8" + } + } + }, + { + "run_name": "p2_extract_gemma4-31b-q4_v5", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 74.4, + "iterations": 6, + "completion_tokens": 4449, + "prompt_tokens_cumulative": 22299, + "model_turns": 6, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9409, + "mean_power_w": 497.01, + "mean_sm_util_pct": 97.33, + "max_temp_c": 73.0, + "cpu_package_mean_power_w": 136.03 + }, + "files": { + "receipt.json": { + "bytes": 9514, + "sha256": "ecd90f7cf590d1e10380aeb2c7cdf8817b4818df451355adef2273d55669cb3a" + }, + "transcript.jsonl": { + "bytes": 4462, + "sha256": "538af5ff784022a45c929e4ec4720850952df9f65700eb318f646ca5dea12d4c" + }, + "summary.json": { + "bytes": 530, + "sha256": "7671f42186b0d0be69d81c68080dc61a106067c6e1428c776c13018becf3fcc1" + }, + "workspace_final.tar.gz": { + "bytes": 11514, + "sha256": "67f2a78730a759f9c46bfbe390172659c8984a5770b93973a275e2ee3a47ac4f" + }, + "cost.json": { + "bytes": 1233, + "sha256": "569468e43fd16636e0208ea0e2e9f305fc95a1852e1b274f4be8050170d09df6" + }, + "gpu_telemetry.json": { + "bytes": 2404, + "sha256": "0e4122f31531d4b8a68f2285636781955aee770a49b74670dfbea04d6b98320b" + }, + "grade.json": { + "bytes": 3204, + "sha256": "b6bc97fc6e03c0b85c34d7e6b5aa1ffcfdf128e953247bc7569f9417170aa2b8" + } + } + }, + { + "run_name": "p2_extract_gemma4-31b-q4_v6", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 73.8, + "iterations": 6, + "completion_tokens": 4449, + "prompt_tokens_cumulative": 22299, + "model_turns": 6, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9485, + "mean_power_w": 478.24, + "mean_sm_util_pct": 97.4, + "max_temp_c": 81.0, + "cpu_package_mean_power_w": 136.19 + }, + "files": { + "receipt.json": { + "bytes": 9514, + "sha256": "fe56f2399c88f9ea4c98c5aa5c00bdd8e9bd5bbd10f95c82893d191fb4c27d27" + }, + "transcript.jsonl": { + "bytes": 4462, + "sha256": "ee619837826b187a6da71e00176157ab71c227682ddd75ad9c1ba14f8004314f" + }, + "summary.json": { + "bytes": 530, + "sha256": "1dc1b010aea5605370f642f660afbd10bbda56d64e814b9cb1eff9ceebdb56d9" + }, + "workspace_final.tar.gz": { + "bytes": 11506, + "sha256": "d6ec772f64e518e2181f6638b2020964d53c69974a43cf5c6d87f7ce6313886d" + }, + "cost.json": { + "bytes": 1233, + "sha256": "dc7e5b9a9a4563b5e725098af8553d2f366edaecb739b15704f5ba9ce5b3964e" + }, + "gpu_telemetry.json": { + "bytes": 2428, + "sha256": "adb155798f51aaf130527ffadb83f5b2a7962ed87d3239a72d0dfcb3fae52036" + }, + "grade.json": { + "bytes": 3204, + "sha256": "b6bc97fc6e03c0b85c34d7e6b5aa1ffcfdf128e953247bc7569f9417170aa2b8" + } + } + }, + { + "run_name": "p2_extract_gemma4-31b-q4_v7", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 41.7, + "iterations": 5, + "completion_tokens": 2449, + "prompt_tokens_cumulative": 14081, + "model_turns": 5, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.8393, + "mean_power_w": 489.61, + "mean_sm_util_pct": 98.12, + "max_temp_c": 75.0, + "cpu_package_mean_power_w": 136.44 + }, + "files": { + "receipt.json": { + "bytes": 9514, + "sha256": "c65db879ff65a91925f142a1a33964e2090dadc2820bda09ccbb9a30ff25ac86" + }, + "transcript.jsonl": { + "bytes": 3390, + "sha256": "08efb55e1349587c59998282540b681e1683aa291908dba1a68aa1387b20bb49" + }, + "summary.json": { + "bytes": 516, + "sha256": "9c13ee2001ec59936511ddb8a9322d64bf594c9206b65cb4165305462f38138c" + }, + "workspace_final.tar.gz": { + "bytes": 11209, + "sha256": "79e8fc914937e254fb5444daf6bbcd8f8c0e65dce5fb49b5230e14fa9fb1f185" + }, + "cost.json": { + "bytes": 1233, + "sha256": "471f8f0da11486ac9f31462dfb1b9e266b36b968e5c5d9dc4e84fa53ebf38d1f" + }, + "gpu_telemetry.json": { + "bytes": 2401, + "sha256": "d6c959fcf80de8c51059acae6013afa0352a5b975f2ef1ddb369b93a1d9e4151" + }, + "grade.json": { + "bytes": 3204, + "sha256": "b6bc97fc6e03c0b85c34d7e6b5aa1ffcfdf128e953247bc7569f9417170aa2b8" + } + } + }, + { + "run_name": "p2_extract_gemma4-31b-q4_v8", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 73.1, + "iterations": 6, + "completion_tokens": 4449, + "prompt_tokens_cumulative": 22299, + "model_turns": 6, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9576, + "mean_power_w": 491.28, + "mean_sm_util_pct": 91.2, + "max_temp_c": 81.0, + "cpu_package_mean_power_w": 136.45 + }, + "files": { + "receipt.json": { + "bytes": 9514, + "sha256": "04484587caefa2541b0b7f8bb91676f3efdc160959a2e7ea2adb1c0fef255ad9" + }, + "transcript.jsonl": { + "bytes": 4462, + "sha256": "a1f4509c5ac517f2626dde3e65c698520bbcd57e43bd91a2d8c71c72cb893763" + }, + "summary.json": { + "bytes": 530, + "sha256": "08a983fa89a8b31e5d041fd23edcd9541a26d3383ff125ec6ffc71cfbda3d85a" + }, + "workspace_final.tar.gz": { + "bytes": 11517, + "sha256": "0e4a822ba1721fa4f6ceafba615867a436fa3b80e30b9074be790274dbc0a198" + }, + "cost.json": { + "bytes": 1233, + "sha256": "11a7e1b4d4ab7e0b3145ab0554fbdf95f6b1adfeefb3df878ee4e3cc6cbb7b2c" + }, + "gpu_telemetry.json": { + "bytes": 2433, + "sha256": "95ea3a5bb39ff552ecc5ee8a71e2a18d09e32abad6ba348f211df66bf5ecc9ab" + }, + "grade.json": { + "bytes": 3204, + "sha256": "b6bc97fc6e03c0b85c34d7e6b5aa1ffcfdf128e953247bc7569f9417170aa2b8" + } + } + }, + { + "run_name": "p2_extract_gemma4-31b-q4_v9", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 75.2, + "iterations": 6, + "completion_tokens": 4449, + "prompt_tokens_cumulative": 22299, + "model_turns": 6, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9309, + "mean_power_w": 475.67, + "mean_sm_util_pct": 85.53, + "max_temp_c": 74.0, + "cpu_package_mean_power_w": 136.34 + }, + "files": { + "receipt.json": { + "bytes": 9514, + "sha256": "2f02aeccc0c50580a59e33f06139a5fab5c3ca1fffb295d3f93090ac3d816e60" + }, + "transcript.jsonl": { + "bytes": 4461, + "sha256": "a9acd39b51ff3c3e866ac1aef7560e91d1811497b4b27041f978fdb21adfe3b3" + }, + "summary.json": { + "bytes": 530, + "sha256": "63d3ada32bc6b0c783541ded981674f430ed2e10421a001cb4bc3a7a0cc58937" + }, + "workspace_final.tar.gz": { + "bytes": 11522, + "sha256": "e9f5a6df5605241ae782fd2525a1ada7f773c77c2ade3ba9b713c6138524dc9d" + }, + "cost.json": { + "bytes": 1233, + "sha256": "387ef64aca01b40e1382b3454a9b2a9b6695e5bdbb8c43fce0d4490e3eefc12c" + }, + "gpu_telemetry.json": { + "bytes": 2406, + "sha256": "bc32b67e971ec8662c72437d057b4527e60b6199225169cb6e016e3216039c92" + }, + "grade.json": { + "bytes": 3204, + "sha256": "b6bc97fc6e03c0b85c34d7e6b5aa1ffcfdf128e953247bc7569f9417170aa2b8" + } + } + }, + { + "run_name": "p2_extract_gemma4-31b-q4_v10", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 73.0, + "iterations": 6, + "completion_tokens": 4449, + "prompt_tokens_cumulative": 22299, + "model_turns": 6, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.8904, + "mean_power_w": 485.19, + "mean_sm_util_pct": 95.64, + "max_temp_c": 82.0, + "cpu_package_mean_power_w": 136.23 + }, + "files": { + "receipt.json": { + "bytes": 9515, + "sha256": "27278b0aea4bf0f36569d445298493fb4d423094518bf9f94793f65c96e18dba" + }, + "transcript.jsonl": { + "bytes": 4460, + "sha256": "bc2ce2cc3d6f5f5ce6af5758f6f512110dbdc406cb5af39adae0cf1de3c00a10" + }, + "summary.json": { + "bytes": 530, + "sha256": "482a9092aefd84da30c53ee28e20fbe7c84b10fcd3cd4d5346534d9a87ddfb20" + }, + "workspace_final.tar.gz": { + "bytes": 11518, + "sha256": "43d8f6480d61a253bfa4321c7fc7b9a0f55d5b364c408c884b131b7c1e3768eb" + }, + "cost.json": { + "bytes": 1234, + "sha256": "421157d493e43e5abe1bc6a1a2cd31080781b16198e338edb8f27d8e3ef4c1d3" + }, + "gpu_telemetry.json": { + "bytes": 2399, + "sha256": "8571e8fd3e81e7d1e2915964e95eb6a7dd9b7a73d66d646062baeff524a9c66b" + }, + "grade.json": { + "bytes": 3204, + "sha256": "b6bc97fc6e03c0b85c34d7e6b5aa1ffcfdf128e953247bc7569f9417170aa2b8" + } + } + }, + { + "run_name": "p2_ci_gemma4-31b-q4_v1", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 88.5, + "iterations": 29, + "completion_tokens": 4141, + "prompt_tokens_cumulative": 234185, + "model_turns": 29, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.904, + "mean_power_w": 440.07, + "mean_sm_util_pct": 69.35, + "max_temp_c": 76.0, + "cpu_package_mean_power_w": 113.04 + }, + "files": { + "receipt.json": { + "bytes": 8947, + "sha256": "1bfeb5c41370bb8e9c06a9cf1442b5475f7433baae1432b55df02397a601c3f3" + }, + "transcript.jsonl": { + "bytes": 18156, + "sha256": "c37e909e7d44dd910771a1a082a0dce0c783b13c2ddb90c6263078cf44397a15" + }, + "summary.json": { + "bytes": 854, + "sha256": "9e8f9cf9e8c965bb749bf17dcedcd8c6a9df7ded2bbb4d03b0644d36f103a361" + }, + "workspace_final.tar.gz": { + "bytes": 42534, + "sha256": "462800e2a5ba5abc66790efa44478c855dc72d785877714df198208cbd052176" + }, + "cost.json": { + "bytes": 1230, + "sha256": "60c46c1767609b23a08dceae3efa29b18b8aba6147e48b1e67fb42515c13cbc7" + }, + "gpu_telemetry.json": { + "bytes": 2391, + "sha256": "38f8d2555bc3f6004b397be565a321b2d5d8637508578b2e2d1c4c2494c3ec88" + }, + "grade.json": { + "bytes": 482, + "sha256": "69e86f389fb98ab93705abc5fd07431d5a9730ba5b0214d67cccdc8eeb154ba6" + } + } + }, + { + "run_name": "p2_ci_gemma4-31b-q4_v2", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 94.5, + "iterations": 29, + "completion_tokens": 4551, + "prompt_tokens_cumulative": 230536, + "model_turns": 29, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.8995, + "mean_power_w": 427.5, + "mean_sm_util_pct": 81.5, + "max_temp_c": 84.0, + "cpu_package_mean_power_w": 136.99 + }, + "files": { + "receipt.json": { + "bytes": 8949, + "sha256": "330300d0460fb67ce69d65b17ba1fabf6127921d0d430f14efbb628bd95e441f" + }, + "transcript.jsonl": { + "bytes": 18370, + "sha256": "3399f7f3cd0183e8e9f3be341cfdc9a8f90874139fce4bd9777c8ecac5bec873" + }, + "summary.json": { + "bytes": 929, + "sha256": "a4a1834961b6e83966b444c41f381cb804e7d4b50ecfd0362aece86fd902195e" + }, + "workspace_final.tar.gz": { + "bytes": 48425, + "sha256": "157aed82ecac3ad44166bc2b54c8d7cc12cda6cd7afa5663ee71c9fa71fac791" + }, + "cost.json": { + "bytes": 1230, + "sha256": "e9c8def56e3794c5f7e4506bdc9e1b09b1dba23f55c0f0836e72e3d38afd5317" + }, + "gpu_telemetry.json": { + "bytes": 2396, + "sha256": "1067a83146d7457a5be9586cabee519f83a6f7f51f66ab39352460b7310b9cc3" + }, + "grade.json": { + "bytes": 482, + "sha256": "eae0f41e05a5ce0e1a370e295ccb33bfd10b26e892ff34e286253904640fd666" + } + } + }, + { + "run_name": "p2_ci_gemma4-31b-q4_v3", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 80.3, + "iterations": 21, + "completion_tokens": 3754, + "prompt_tokens_cumulative": 162381, + "model_turns": 21, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.934, + "mean_power_w": 435.34, + "mean_sm_util_pct": 79.44, + "max_temp_c": 72.0, + "cpu_package_mean_power_w": 112.97 + }, + "files": { + "receipt.json": { + "bytes": 8947, + "sha256": "de8c9045350678be58207876eebbfa395ab86d5e5ad7b19676f11c12b711e3b1" + }, + "transcript.jsonl": { + "bytes": 14266, + "sha256": "1ad6dd78d4c1f09b610922e47b0d78686957cd2bad6071994ae6e2a484428350" + }, + "summary.json": { + "bytes": 881, + "sha256": "3ec5f398e260ad00d3ee87b7c5b6d40905369999b640150c306f2eb42b7a9612" + }, + "workspace_final.tar.gz": { + "bytes": 40255, + "sha256": "e2f34ad318240ff88eedc621842a08ba6715ff4eb33629f3e942637c0848dc86" + }, + "cost.json": { + "bytes": 1230, + "sha256": "ce7c5a844032037bb5de24b724dcaa6c97ce6caa2cbf5c0e5481602ed615241d" + }, + "gpu_telemetry.json": { + "bytes": 2391, + "sha256": "91387aedad676d1013ae6724147326a36162cc12706645c3a0bf36ead6e56032" + }, + "grade.json": { + "bytes": 482, + "sha256": "1cdcddf815967c7395bceb0d8ac5b955294d82d5d932594f553e43f22ab4234d" + } + } + }, + { + "run_name": "p2_ci_gemma4-31b-q4_v4", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 112.8, + "iterations": 33, + "completion_tokens": 5568, + "prompt_tokens_cumulative": 268477, + "model_turns": 33, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9309, + "mean_power_w": 413.59, + "mean_sm_util_pct": 71.55, + "max_temp_c": 80.0, + "cpu_package_mean_power_w": 136.53 + }, + "files": { + "receipt.json": { + "bytes": 9509, + "sha256": "b01c5c20d71014bd8606560100a7e0e51a7d8281a91df6ef20bbc6d1b2cb2bd2" + }, + "transcript.jsonl": { + "bytes": 20495, + "sha256": "9405c53146d85b2653f2300cdc3fed219eec771553a3f9ad15a52d755a8e3512" + }, + "summary.json": { + "bytes": 858, + "sha256": "e4aef981b069622d6028c4fbc266d6493811711089f7db08a925ec602764952c" + }, + "workspace_final.tar.gz": { + "bytes": 44384, + "sha256": "88329469d5365ba5777b0dac9502442eb79cb496a3b67d650e249d97f9e4efe9" + }, + "cost.json": { + "bytes": 1231, + "sha256": "01439f373d24cb0d1887f280082312d913696571d03754ef946937974f513374" + }, + "gpu_telemetry.json": { + "bytes": 2442, + "sha256": "82bd0a0a13808194fa093f2813833a43ef1015eb6ff0d2bd02c3ecf70ade54a9" + }, + "grade.json": { + "bytes": 482, + "sha256": "eae0f41e05a5ce0e1a370e295ccb33bfd10b26e892ff34e286253904640fd666" + } + } + }, + { + "run_name": "p2_ci_gemma4-31b-q4_v5", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 89.2, + "iterations": 26, + "completion_tokens": 4419, + "prompt_tokens_cumulative": 178919, + "model_turns": 26, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.8969, + "mean_power_w": 456.25, + "mean_sm_util_pct": 79.35, + "max_temp_c": 76.0, + "cpu_package_mean_power_w": 136.23 + }, + "files": { + "receipt.json": { + "bytes": 9509, + "sha256": "fd1eb38e77c77962646be645d453b8eee18598c7f17e362dbbf13325308f61f4" + }, + "transcript.jsonl": { + "bytes": 16846, + "sha256": "61d956c1d4e619631024dfb2d7123b7a932f89be27e587fe5264d5b1a9fda62b" + }, + "summary.json": { + "bytes": 890, + "sha256": "7aa5e995d0f6bdc97e036db0524cc14b6b3af433db61193b108a22345923b84a" + }, + "workspace_final.tar.gz": { + "bytes": 38108, + "sha256": "7741e7c20280e8c6b0a5e298b8d67e53cce0d53ff1d6ad9f7e005bfbd36fc2a1" + }, + "cost.json": { + "bytes": 1230, + "sha256": "fc925d07d641102e5590697224cea9db28d5fac251cddbb0ae630689ee3f72eb" + }, + "gpu_telemetry.json": { + "bytes": 2401, + "sha256": "821b926c6b747a82b82d36f0e46d04ef5f9c2a88c51a91c96f9ddcecfe0631be" + }, + "grade.json": { + "bytes": 482, + "sha256": "1cdcddf815967c7395bceb0d8ac5b955294d82d5d932594f553e43f22ab4234d" + } + } + }, + { + "run_name": "p2_ci_gemma4-31b-q4_v6", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 110.7, + "iterations": 34, + "completion_tokens": 5323, + "prompt_tokens_cumulative": 277652, + "model_turns": 34, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9485, + "mean_power_w": 427.74, + "mean_sm_util_pct": 85.82, + "max_temp_c": 78.0, + "cpu_package_mean_power_w": 135.7 + }, + "files": { + "receipt.json": { + "bytes": 9509, + "sha256": "4305876d12c8a246c2dae31c8b7878cc3d938e6d7d97ba46695fa1e68c2381fc" + }, + "transcript.jsonl": { + "bytes": 20802, + "sha256": "aa0ef593d5ce3e5a45bb0282ed26f72b1470e590dfe9a3243ac067708e7f274d" + }, + "summary.json": { + "bytes": 1040, + "sha256": "20384801c0bfd097ed3a5d632959deb2106239ec1f9453cf661b1fb4c9630a64" + }, + "workspace_final.tar.gz": { + "bytes": 49327, + "sha256": "796cdb474b976cfe00baea67bb788f47749cfa94bbd738f2fe1f67cb0f8c5227" + }, + "cost.json": { + "bytes": 1231, + "sha256": "e2dcb0c33230336ede8d9854487d5833f05fcc6c836156d353ea2f87cef90073" + }, + "gpu_telemetry.json": { + "bytes": 2439, + "sha256": "6a69ebccfe184b9854525102108c0816a8880d598d93485b533558358b733158" + }, + "grade.json": { + "bytes": 482, + "sha256": "1cdcddf815967c7395bceb0d8ac5b955294d82d5d932594f553e43f22ab4234d" + } + } + }, + { + "run_name": "p2_ci_gemma4-31b-q4_v7", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 97.7, + "iterations": 31, + "completion_tokens": 4662, + "prompt_tokens_cumulative": 225919, + "model_turns": 31, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9212, + "mean_power_w": 445.14, + "mean_sm_util_pct": 79.42, + "max_temp_c": 72.0, + "cpu_package_mean_power_w": 136.3 + }, + "files": { + "receipt.json": { + "bytes": 9509, + "sha256": "4de8227e730cfb3c90da4c038a4396185fd9de7b902fe64809a53a9a96d062a4" + }, + "transcript.jsonl": { + "bytes": 19474, + "sha256": "f77b5778a2e59e8fde4c4191899b903730b6b96a148589782b15aaad65947041" + }, + "summary.json": { + "bytes": 893, + "sha256": "9f0d98131e0d7d798c1e5597e1cf780b2401c5bc04e4b04ea9fe91424b30cf3e" + }, + "workspace_final.tar.gz": { + "bytes": 37515, + "sha256": "6aa983ef6c7e4cf4519b4a19f903eb2cb505aa6a437a73da459f337f74068d4d" + }, + "cost.json": { + "bytes": 1230, + "sha256": "4462d88c021fc5b475ed3534f771d541cd26d550f4622e866759204c422f2904" + }, + "gpu_telemetry.json": { + "bytes": 2441, + "sha256": "32d919eb7b25990ef2c42c0c8dd2de483d8cdac72d7448cb7d88e1edc1353ddc" + }, + "grade.json": { + "bytes": 482, + "sha256": "eae0f41e05a5ce0e1a370e295ccb33bfd10b26e892ff34e286253904640fd666" + } + } + }, + { + "run_name": "p2_ci_gemma4-31b-q4_v8", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 83.9, + "iterations": 21, + "completion_tokens": 4121, + "prompt_tokens_cumulative": 142741, + "model_turns": 21, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9535, + "mean_power_w": 439.8, + "mean_sm_util_pct": 71.47, + "max_temp_c": 77.0, + "cpu_package_mean_power_w": 135.19 + }, + "files": { + "receipt.json": { + "bytes": 9509, + "sha256": "9aee7f99f047c771413190ce7fc9791add6032ec47fd21eae8d867c3dc4544bf" + }, + "transcript.jsonl": { + "bytes": 14957, + "sha256": "4543e8543e68fc61f8e970e7c3c6426c14c1f872a3e6827c282ad2ce2032eeba" + }, + "summary.json": { + "bytes": 918, + "sha256": "11bd69109b6f058382aab7fd776764dcc736b80e6281923ff5cecdc59374e0c7" + }, + "workspace_final.tar.gz": { + "bytes": 37193, + "sha256": "f5a3571b42333d990be09e285e8ef767d1a77448d573d8d2fdfd7d45060742ca" + }, + "cost.json": { + "bytes": 1230, + "sha256": "b383030a326f4cdb03c94abba531ccccb20f0a8a924403872e51be46a379ee35" + }, + "gpu_telemetry.json": { + "bytes": 2394, + "sha256": "c0f08e17333fcb864a77875616fe068b995384f074f6b882a9b58a7b367f89b1" + }, + "grade.json": { + "bytes": 482, + "sha256": "eae0f41e05a5ce0e1a370e295ccb33bfd10b26e892ff34e286253904640fd666" + } + } + }, + { + "run_name": "p2_ci_gemma4-31b-q4_v9", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 107.0, + "iterations": 24, + "completion_tokens": 4328, + "prompt_tokens_cumulative": 158816, + "model_turns": 24, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9813, + "mean_power_w": 429.32, + "mean_sm_util_pct": 76.0, + "max_temp_c": 74.0, + "cpu_package_mean_power_w": 136.4 + }, + "files": { + "receipt.json": { + "bytes": 9509, + "sha256": "a17df3df52bed8c3487eb316b4795e9f1589d168289787d5895fef8eb5b022b2" + }, + "transcript.jsonl": { + "bytes": 17007, + "sha256": "f0d8c4aea5230dfcd14157a0cbfee9637c7d8d2664fb24bdbf7b03f71ba305cd" + }, + "summary.json": { + "bytes": 974, + "sha256": "b3770fd07de1e8cea5716fd5577ad91342897070f2223b53c1d43804539be08a" + }, + "workspace_final.tar.gz": { + "bytes": 46266, + "sha256": "d4c7dffbebddf82fd554b99521598d383b093ce66e8f2bc0820f5aede6a700ad" + }, + "cost.json": { + "bytes": 1232, + "sha256": "42f722a41a0289f031dc880d881d73c8dee9d91962743445413f187cccc58f2d" + }, + "gpu_telemetry.json": { + "bytes": 2400, + "sha256": "8d78a5693dfde8bf6e1733bdb91f229caa73471d8f38b4fe89a750f1d25c1c45" + }, + "grade.json": { + "bytes": 482, + "sha256": "1cdcddf815967c7395bceb0d8ac5b955294d82d5d932594f553e43f22ab4234d" + } + } + }, + { + "run_name": "p2_ci_gemma4-31b-q4_v10", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 85.9, + "iterations": 27, + "completion_tokens": 4115, + "prompt_tokens_cumulative": 162685, + "model_turns": 27, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9313, + "mean_power_w": 444.87, + "mean_sm_util_pct": 84.88, + "max_temp_c": 78.0, + "cpu_package_mean_power_w": 135.37 + }, + "files": { + "receipt.json": { + "bytes": 9510, + "sha256": "14e3c802c8d95169e8ef06ae91ddd2456bebaecf3933f5547896dfa09efa4d36" + }, + "transcript.jsonl": { + "bytes": 17577, + "sha256": "5e4c30b150267c67ebd9f9ff23658d4dacd6b63a74f6dd92462eb5979689d25d" + }, + "summary.json": { + "bytes": 1175, + "sha256": "bfbe4b1911510355ad01e0bd97d77a8075e39df08576f65ae4afef7706d22f5d" + }, + "workspace_final.tar.gz": { + "bytes": 37955, + "sha256": "0bc909cc7f9070384f9412065c1b99814a0787f403f7af757830abfea596fcea" + }, + "cost.json": { + "bytes": 1231, + "sha256": "79e548aa1bc85aee47f8acbb1ac7b97ca075f6f20a25d3ddaf47673aaa80d6af" + }, + "gpu_telemetry.json": { + "bytes": 2436, + "sha256": "5b36f52c9be27d49de8e1c1f99b813f6b8c0dca0d1361eb4d94d4b15db254ec3" + }, + "grade.json": { + "bytes": 482, + "sha256": "eae0f41e05a5ce0e1a370e295ccb33bfd10b26e892ff34e286253904640fd666" + } + } + }, + { + "run_name": "p2_hallucination_gemma4-31b-q4_v1", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 115.6, + "iterations": 16, + "completion_tokens": 6454, + "prompt_tokens_cumulative": 121925, + "model_turns": 16, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9516, + "mean_power_w": 488.6, + "mean_sm_util_pct": 91.83, + "max_temp_c": 83.0, + "cpu_package_mean_power_w": 136.58 + }, + "files": { + "receipt.json": { + "bytes": 8966, + "sha256": "000d8b774edf30dac9b4fe688ef6ed94bffddb1196edda77469269ecf806b602" + }, + "transcript.jsonl": { + "bytes": 9410, + "sha256": "769746a886657bd6a339d0b376a9405b5d877fcc770da5c1b28d8ffb09cbf142" + }, + "summary.json": { + "bytes": 489, + "sha256": "9922fd7d5dcb5739be2ca233f2ada76172b5beff1cf37779d874f87963e3d677" + }, + "workspace_final.tar.gz": { + "bytes": 11352, + "sha256": "b4de9834bc4281d8823e8aeadcb633e13e672640604304733e1b096aa58dab57" + }, + "cost.json": { + "bytes": 1245, + "sha256": "c13344093961e064a51b8a6550ab1464debb87e2a1989056e8b65b09cb9ae939" + }, + "gpu_telemetry.json": { + "bytes": 2409, + "sha256": "85489751ad178193b914d9fc65d0d35a1e4c3ca8f422fb799fc7462350bd7393" + }, + "grade.json": { + "bytes": 3586, + "sha256": "80910919c9e64a98ef035d0e5da409c4b26fc67de0ad65bcf63df6b4012393e6" + } + } + }, + { + "run_name": "p2_hallucination_gemma4-31b-q4_v2", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 114.1, + "iterations": 17, + "completion_tokens": 6154, + "prompt_tokens_cumulative": 128689, + "model_turns": 17, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9202, + "mean_power_w": 481.27, + "mean_sm_util_pct": 93.09, + "max_temp_c": 74.0, + "cpu_package_mean_power_w": 113.31 + }, + "files": { + "receipt.json": { + "bytes": 8964, + "sha256": "2cbada4e3d3fadf359becdf5da4b9b866f6d51d025453f642c23faadd5879e27" + }, + "transcript.jsonl": { + "bytes": 10566, + "sha256": "eafdf5cd5b4d9b307e1d51724bc1b042e2970ead94e6eed757d8e6de8edd4c66" + }, + "summary.json": { + "bytes": 442, + "sha256": "42545f2faeccbd3c3c5af74e4640c9da4e2454ed0a38ac1f3e81ba6392299d7e" + }, + "workspace_final.tar.gz": { + "bytes": 11814, + "sha256": "00ae973a8d6e9f1e567870e7403d1bf605bf3e82a70bdef609287d640946daf5" + }, + "cost.json": { + "bytes": 1244, + "sha256": "00fd00be3788b4948ddfc27c49d81d176eb0de734f3d872df545c37054021e58" + }, + "gpu_telemetry.json": { + "bytes": 2407, + "sha256": "bb21057685f276aaed4370f84ee3d1f4c5c5cc1c7ed7dbea9c84b944f896733b" + }, + "grade.json": { + "bytes": 3892, + "sha256": "2553e9edff94d4bf1cc8fda5c170ae25201bd4baa261fc4ab14d5754dbc11bed" + } + } + }, + { + "run_name": "p2_hallucination_gemma4-31b-q4_v3", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "model_stopped", + "elapsed_s": 32.4, + "iterations": 9, + "completion_tokens": 1494, + "prompt_tokens_cumulative": 43015, + "model_turns": 9, + "length_finishes": 0, + "grade_verdict": "MISSING_OUTPUT", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.7716, + "mean_power_w": 417.97, + "mean_sm_util_pct": 83.33, + "max_temp_c": 83.0, + "cpu_package_mean_power_w": 137.12 + }, + "files": { + "receipt.json": { + "bytes": 8966, + "sha256": "c588242987936d1cc91253b8dfe727b4567ee3df212b10b0cc3de0097010adce" + }, + "transcript.jsonl": { + "bytes": 3643, + "sha256": "7096baadd99ee25917c930c0ef99ed2746e11fe4d85142197ebc2a05c12e8f03" + }, + "summary.json": { + "bytes": 346, + "sha256": "c3a15b3d9a9d02fd77292f80fc985d594c28eb5b848e1a57a496243816a89111" + }, + "workspace_final.tar.gz": { + "bytes": 10689, + "sha256": "cb4a3526503a052dac2aaaa769c353fca88417571ed16dfbefd9d3ab09e0d4aa" + }, + "cost.json": { + "bytes": 1239, + "sha256": "5954c2443675b7879dac302d9089518852403f8a613775af1d9ba61dd772b4f8" + }, + "gpu_telemetry.json": { + "bytes": 2400, + "sha256": "942cd09653e35e36863c5665b31460d4aef5650a558de1a52b0d5c057ea64cdb" + }, + "grade.json": { + "bytes": 153, + "sha256": "e4f8b9da1c7aba632e47e0a97771274c35662e1f1e2920b050928e498df8692a" + } + } + }, + { + "run_name": "p2_hallucination_gemma4-31b-q4_v4", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "model_stopped", + "elapsed_s": 56.0, + "iterations": 13, + "completion_tokens": 2921, + "prompt_tokens_cumulative": 78794, + "model_turns": 13, + "length_finishes": 0, + "grade_verdict": "MISSING_OUTPUT", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.8929, + "mean_power_w": 413.1, + "mean_sm_util_pct": 74.64, + "max_temp_c": 81.0, + "cpu_package_mean_power_w": 135.0 + }, + "files": { + "receipt.json": { + "bytes": 9526, + "sha256": "95c23c81942b46d22616bff2d0624a28b5fefab48d2f37ff2e33bdfc1d8ffd04" + }, + "transcript.jsonl": { + "bytes": 5387, + "sha256": "112c51f97e96d152feeddc6996231f4ba86d033166ac138d46f0fb799953de17" + }, + "summary.json": { + "bytes": 347, + "sha256": "3ab63679915c106e4abdb6ac1c70331a1cbd9173da5a139b2499f61a23c767b2" + }, + "workspace_final.tar.gz": { + "bytes": 10696, + "sha256": "4ed429ca339ec41b15bcee85b4fdd83434195852f2e6174cd689c342df1537d3" + }, + "cost.json": { + "bytes": 1239, + "sha256": "300d945d1da5ccf35e18f11b5c06f2dda892963597f26b5748d0c758e268b882" + }, + "gpu_telemetry.json": { + "bytes": 2405, + "sha256": "1d67d847e298003d7c9272298c33df1b71d1e62551826fd239e528dc97bb7b0d" + }, + "grade.json": { + "bytes": 153, + "sha256": "48183e26e443a49c4cc81ec4584bfa80e69da97f0759b05cf58e3c5b45bdc8f8" + } + } + }, + { + "run_name": "p2_hallucination_gemma4-31b-q4_v5", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "model_stopped", + "elapsed_s": 55.8, + "iterations": 13, + "completion_tokens": 2921, + "prompt_tokens_cumulative": 78794, + "model_turns": 13, + "length_finishes": 0, + "grade_verdict": "MISSING_OUTPUT", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.8961, + "mean_power_w": 451.88, + "mean_sm_util_pct": 79.09, + "max_temp_c": 72.0, + "cpu_package_mean_power_w": 136.04 + }, + "files": { + "receipt.json": { + "bytes": 9526, + "sha256": "b827f7ee312e5a3eb6e6b7f28e3da51a2ddbee36c810dbd073b68f069ddc4323" + }, + "transcript.jsonl": { + "bytes": 5388, + "sha256": "eb2640b51b56f99567f87c8c47315746659b98e869aa5c28c9152c856583eb80" + }, + "summary.json": { + "bytes": 347, + "sha256": "5d07f083e89c8f9025b23855157a170bc3808020cf77eceb840dee98ea6a5c1f" + }, + "workspace_final.tar.gz": { + "bytes": 10683, + "sha256": "9fccdb68e44549910c1fcc62bd478fa3b07c721174b457169f0d3a9f215c2338" + }, + "cost.json": { + "bytes": 1239, + "sha256": "239a84d3879bad863f8baa853fd28b83927c2d556d53f9bbf575712e1cb05a46" + }, + "gpu_telemetry.json": { + "bytes": 2450, + "sha256": "1251eb4702a7a6583f6dbff62369a83aec34bef11502921f26e96229be4a7a86" + }, + "grade.json": { + "bytes": 153, + "sha256": "019bbdcdc8ce03c38b61fabde3ebe89220203a5af63f6a565860128fd32d3a41" + } + } + }, + { + "run_name": "p2_hallucination_gemma4-31b-q4_v6", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 90.3, + "iterations": 15, + "completion_tokens": 4894, + "prompt_tokens_cumulative": 105469, + "model_turns": 15, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9413, + "mean_power_w": 479.87, + "mean_sm_util_pct": 92.06, + "max_temp_c": 79.0, + "cpu_package_mean_power_w": 135.2 + }, + "files": { + "receipt.json": { + "bytes": 9526, + "sha256": "cf0d3d7cd037b0704fe18607dc9a3d7ed994948824526ff5e52570462e5ea760" + }, + "transcript.jsonl": { + "bytes": 8820, + "sha256": "2116318829445a03c35dadb60c4d8d3ed8dbc0a456732b77924c63e0610a98ab" + }, + "summary.json": { + "bytes": 450, + "sha256": "c2de04fc911ef2bc16a5741a05e25c26cb64c9ccb60ed8b126818c27744ffbac" + }, + "workspace_final.tar.gz": { + "bytes": 11393, + "sha256": "002f28d2dbd2c1bb3ff66d494b520f2e96c600913084b643ae6a29f9d8f916c2" + }, + "cost.json": { + "bytes": 1242, + "sha256": "7d61b019487fedb0992ba19a19e02a4cdf09dd66273c0ccb38f258e2452a5bf4" + }, + "gpu_telemetry.json": { + "bytes": 2404, + "sha256": "4ecb7a4e9fbca50ee3abe237d4663cceab5d71e129579371082cdbfbb753ef28" + }, + "grade.json": { + "bytes": 3775, + "sha256": "2e85eba6943daf1b7e00211c0a4339e3d91c4942516de7bdc6468b260d7665fe" + } + } + }, + { + "run_name": "p2_hallucination_gemma4-31b-q4_v7", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "model_stopped", + "elapsed_s": 33.6, + "iterations": 9, + "completion_tokens": 1494, + "prompt_tokens_cumulative": 43015, + "model_turns": 9, + "length_finishes": 0, + "grade_verdict": "MISSING_OUTPUT", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.8929, + "mean_power_w": 411.38, + "mean_sm_util_pct": 57.86, + "max_temp_c": 71.0, + "cpu_package_mean_power_w": 135.66 + }, + "files": { + "receipt.json": { + "bytes": 9526, + "sha256": "bcfd40e0946dac15e8177ce963dd564e053af3e6daab4bbb3758a475262234e2" + }, + "transcript.jsonl": { + "bytes": 3640, + "sha256": "a9c7d8ebb70021a87911cf35d5aea7faf8c3c0e097c2a74776e4849cb29909f3" + }, + "summary.json": { + "bytes": 346, + "sha256": "6b852fcc2c29bc4a1a0dd3d9a2c9bcb0d4c0e1d9587197f849f3c011f483a66d" + }, + "workspace_final.tar.gz": { + "bytes": 10685, + "sha256": "e89b8983597e737f77e2084e62cc71629970fb065d620c64c5859f17a1948454" + }, + "cost.json": { + "bytes": 1239, + "sha256": "634d07db19e2666586c5772412dfb784bcf6894ec6ed05be0c358035b1fdaeff" + }, + "gpu_telemetry.json": { + "bytes": 2405, + "sha256": "1b30047b6685f8568acf2d332700acb622c05581750dfae4c28fdbc37c826f26" + }, + "grade.json": { + "bytes": 153, + "sha256": "13f5071c95d5ba1db4673f578df59ec4f3441f0e16cce43008f8192563aebee9" + } + } + }, + { + "run_name": "p2_hallucination_gemma4-31b-q4_v8", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 115.3, + "iterations": 16, + "completion_tokens": 6454, + "prompt_tokens_cumulative": 121925, + "model_turns": 16, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.954, + "mean_power_w": 478.96, + "mean_sm_util_pct": 91.7, + "max_temp_c": 80.0, + "cpu_package_mean_power_w": 135.5 + }, + "files": { + "receipt.json": { + "bytes": 9526, + "sha256": "6892863d87cdda5ba70af0ba923fbf4ded3f51c118a22cc47545fe860630470b" + }, + "transcript.jsonl": { + "bytes": 9412, + "sha256": "21f448862453afb964ee8260941892101160bc25035880745d9e8af40616f581" + }, + "summary.json": { + "bytes": 489, + "sha256": "31ebfb1f52bf9488b3c82774981305c05d46a4a8b647e48a9bd7c2fe44e8f3cc" + }, + "workspace_final.tar.gz": { + "bytes": 11358, + "sha256": "ac994c9dbb4d0e3754cae1662eef0ec565542f1ef637dda7cd223cbeb2fed269" + }, + "cost.json": { + "bytes": 1244, + "sha256": "abbbb8aa0d12f7bf693458e47e2a94d116287e28dc717600c58a52f21ffdcb4d" + }, + "gpu_telemetry.json": { + "bytes": 2445, + "sha256": "9b0bf289431a0a8f9997183774ca49ae57d57884b5b84bb6e0b2a4090aea33c3" + }, + "grade.json": { + "bytes": 3586, + "sha256": "80910919c9e64a98ef035d0e5da409c4b26fc67de0ad65bcf63df6b4012393e6" + } + } + }, + { + "run_name": "p2_hallucination_gemma4-31b-q4_v9", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 110.7, + "iterations": 17, + "completion_tokens": 6152, + "prompt_tokens_cumulative": 133050, + "model_turns": 17, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9485, + "mean_power_w": 480.52, + "mean_sm_util_pct": 92.18, + "max_temp_c": 73.0, + "cpu_package_mean_power_w": 135.68 + }, + "files": { + "receipt.json": { + "bytes": 9526, + "sha256": "3ed1ab0c4ebdf4d00b4591d220d3b74428a06d649fef25b904b223c350246f0d" + }, + "transcript.jsonl": { + "bytes": 10674, + "sha256": "199483503de74638487a883a26b707d4f1e181f14d2edecc04a32dce93a51593" + }, + "summary.json": { + "bytes": 464, + "sha256": "a0334f34f310ba6078447a194bd1a7461742dfe498f014a80471cf6b3bbd41ed" + }, + "workspace_final.tar.gz": { + "bytes": 11694, + "sha256": "2c1b915fc4f53465a91c41de951cb345aa922b691f2a2108e971c0602ca0fefc" + }, + "cost.json": { + "bytes": 1243, + "sha256": "be0ad47ba386b2e3c04b211385d3ad187467ed67478b7502d83843015716a72e" + }, + "gpu_telemetry.json": { + "bytes": 2413, + "sha256": "7b6ca36639569bc8676ab7dd16d21a6ed9171edfa3d40ef945872a96126e6052" + }, + "grade.json": { + "bytes": 3836, + "sha256": "3c8f6c79a0738d920981fde9054c69e95edfb470cd987ad7b94a29a5fb160a1e" + } + } + }, + { + "run_name": "p2_hallucination_gemma4-31b-q4_v10", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 123.8, + "iterations": 19, + "completion_tokens": 6853, + "prompt_tokens_cumulative": 159756, + "model_turns": 19, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9289, + "mean_power_w": 476.4, + "mean_sm_util_pct": 91.5, + "max_temp_c": 81.0, + "cpu_package_mean_power_w": 135.79 + }, + "files": { + "receipt.json": { + "bytes": 9527, + "sha256": "00d68ea7f7d8ab8ff0e8ca8f8372bb00995b372f071f0f692bc5f18d7ac48599" + }, + "transcript.jsonl": { + "bytes": 11917, + "sha256": "cc976787eb7dc2baa99e0daf65e4bbce61644cf2136854938e2d2e23f8025353" + }, + "summary.json": { + "bytes": 466, + "sha256": "d24e639e4c6955c15bbb37e9dea15502599abe3266728b069ba5f92dd85c65a8" + }, + "workspace_final.tar.gz": { + "bytes": 12086, + "sha256": "443141942c91d2414e34c5edf986da995a5fd68a0d836a8aaec48b888d841daa" + }, + "cost.json": { + "bytes": 1246, + "sha256": "f8c2bad2156c736da9f3f361aecb20269570e78215ee73cdafd01989ae9452be" + }, + "gpu_telemetry.json": { + "bytes": 2448, + "sha256": "ae641035c56981df58963014d296e51c6e68b0e669e8004a0a9d07e3929ed775" + }, + "grade.json": { + "bytes": 3826, + "sha256": "b934ae4e4b2b6163c0101cb6595b3d55ecc27ebcd931f263c0fb7572153204d5" + } + } + }, + { + "run_name": "p2_triage_gemma4-31b-q4_v1", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 108.0, + "iterations": 8, + "completion_tokens": 6211, + "prompt_tokens_cumulative": 57359, + "model_turns": 8, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9259, + "mean_power_w": 470.89, + "mean_sm_util_pct": 85.81, + "max_temp_c": 75.0, + "cpu_package_mean_power_w": 113.07 + }, + "files": { + "receipt.json": { + "bytes": 8943, + "sha256": "667b801d0ba40236fb24011fe328c17b62ebee20c2ced924cb7751ed61e5ec29" + }, + "transcript.jsonl": { + "bytes": 10882, + "sha256": "98d8dfa0c742dadeb2ed12650b13f0e6b6c28aeb81de737f908ee1bd9c02ad15" + }, + "summary.json": { + "bytes": 600, + "sha256": "7a5d63bcd45effd15282edde7eef5d2319c72c0d0f89aa0306c3399913d34102" + }, + "workspace_final.tar.gz": { + "bytes": 12714, + "sha256": "9614fb670dc01897da49271da079012b43d8b729af22be864ffe61f18724b321" + }, + "cost.json": { + "bytes": 1233, + "sha256": "0d26ca75ca0e7cdf0d11ef36a3ef7aeed7a270c08a247a2a203b02b3c3a1db0a" + }, + "gpu_telemetry.json": { + "bytes": 2395, + "sha256": "960ab621bc46668045c3797d1d9ee0e829077b1ea28c149137ee9836fbd8dffa" + }, + "grade.json": { + "bytes": 2250, + "sha256": "dbe6f269c365da51b85a5053d621051e8877fc7a7642fcc441c97f6c7653b5f5" + } + } + }, + { + "run_name": "p2_triage_gemma4-31b-q4_v2", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 125.5, + "iterations": 7, + "completion_tokens": 7489, + "prompt_tokens_cumulative": 52650, + "model_turns": 7, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9562, + "mean_power_w": 499.84, + "mean_sm_util_pct": 97.36, + "max_temp_c": 84.0, + "cpu_package_mean_power_w": 136.63 + }, + "files": { + "receipt.json": { + "bytes": 8945, + "sha256": "0288e87079edfe2244a331754506643c233790a8cd98d5dc40707a11c9a19f68" + }, + "transcript.jsonl": { + "bytes": 10908, + "sha256": "bf08f200f4260bf44aff76dcea291177648d68dd8ebf33a8aa7d874ef862408f" + }, + "summary.json": { + "bytes": 659, + "sha256": "217354a083ead3542f13afd8380ca186fa28041cc8cc6726b8127a737c1c00bb" + }, + "workspace_final.tar.gz": { + "bytes": 12920, + "sha256": "3349fae04f9709e12ab95b4559a5d6568b7ba017c582733187b344025d7148d9" + }, + "cost.json": { + "bytes": 1233, + "sha256": "3e4975f2694ec3f49d6c712fa9ee2063143dacc684bc0a9d5d40b939e0a3f85e" + }, + "gpu_telemetry.json": { + "bytes": 2400, + "sha256": "50af8818dda151a2f093d6bb99c3067ad9b8bcff7ff25e5ca404c324fb9f5232" + }, + "grade.json": { + "bytes": 2159, + "sha256": "c14e218d16bc8ba2bfb56fd8577d0f8539d628d0366f9340f7d2f60f17c91290" + } + } + }, + { + "run_name": "p2_triage_gemma4-31b-q4_v3", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 128.3, + "iterations": 7, + "completion_tokens": 7493, + "prompt_tokens_cumulative": 52729, + "model_turns": 7, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9353, + "mean_power_w": 489.85, + "mean_sm_util_pct": 91.68, + "max_temp_c": 75.0, + "cpu_package_mean_power_w": 113.07 + }, + "files": { + "receipt.json": { + "bytes": 8943, + "sha256": "cb906fe8c74989a3a0844c92838f901effd382d84db056815fa34d4cb808ca10" + }, + "transcript.jsonl": { + "bytes": 10654, + "sha256": "79b135973c3285e556806fcb26738fca7a775f02550338d4a61af085a06c9ade" + }, + "summary.json": { + "bytes": 644, + "sha256": "80f55c1f0a23af5a58a6d67552818f1af54ef24f223a46063e6c144ff0286d25" + }, + "workspace_final.tar.gz": { + "bytes": 12802, + "sha256": "92ab6a3283af89a091f8329811a9b05d781fd3dee3ee0985946fe5657048cf95" + }, + "cost.json": { + "bytes": 1234, + "sha256": "bd2fb5a02e7a3758b62c82d542de7c5cb4bf7a074b0c91381378c1b9eee5e63e" + }, + "gpu_telemetry.json": { + "bytes": 2399, + "sha256": "88af6bb704a1aa4237ac5d7bbd4ea5e17dc9119055a72b95ba13d13caedc9caf" + }, + "grade.json": { + "bytes": 1914, + "sha256": "8b145bdcbc912d2f546ce33b90d11fe348c69829b205abe91d6513b83ba9daef" + } + } + }, + { + "run_name": "p2_triage_gemma4-31b-q4_v4", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 115.9, + "iterations": 8, + "completion_tokens": 6758, + "prompt_tokens_cumulative": 60330, + "model_turns": 8, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9491, + "mean_power_w": 495.74, + "mean_sm_util_pct": 93.52, + "max_temp_c": 81.0, + "cpu_package_mean_power_w": 136.02 + }, + "files": { + "receipt.json": { + "bytes": 9505, + "sha256": "cd59734cf51ab8d8be586e9b47b4a23b32c1642e4b9cb4cd46e0f211d5f3887d" + }, + "transcript.jsonl": { + "bytes": 11182, + "sha256": "0583f3e0b0ec173970bd54d714dca146b8e0925550df739387b721b1426e5967" + }, + "summary.json": { + "bytes": 689, + "sha256": "73bf5522c1110f01cc2ae5490bf54b8a9a1daaeb5536eb928eacb1a0c2275239" + }, + "workspace_final.tar.gz": { + "bytes": 12813, + "sha256": "113a6e31f92ade78affd362ad21649b5206d9ddbd5f0e78f3f8bfc993a0dcd7a" + }, + "cost.json": { + "bytes": 1234, + "sha256": "d1cc3c5620b7d54f1479401d821ec404faeb2d948c9031e9a2a29a8c07df8a98" + }, + "gpu_telemetry.json": { + "bytes": 2446, + "sha256": "b4c9f482ce065499e778639e2f15e6baa079154abaaef98a672dedbd405f985c" + }, + "grade.json": { + "bytes": 1727, + "sha256": "0e5b1aa18e87a58f55f541ac729fc931d1e38dfa5205b3a047e2ad859cd06a04" + } + } + }, + { + "run_name": "p2_triage_gemma4-31b-q4_v5", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 131.2, + "iterations": 4, + "completion_tokens": 7754, + "prompt_tokens_cumulative": 29427, + "model_turns": 4, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9527, + "mean_power_w": 494.15, + "mean_sm_util_pct": 96.46, + "max_temp_c": 76.0, + "cpu_package_mean_power_w": 135.9 + }, + "files": { + "receipt.json": { + "bytes": 9505, + "sha256": "1a1280224279b2ac73753f97e94c63b13639412536c0125efc1ff4de695ba74f" + }, + "transcript.jsonl": { + "bytes": 8135, + "sha256": "79fc839cf0fe157e7500804c4ad186c7c7f92fe73736b3f94442df6e39b3d2c0" + }, + "summary.json": { + "bytes": 556, + "sha256": "d412173870f5cb4290ece35149bbcaa8418c65ec649c0e6b1e945477b873238c" + }, + "workspace_final.tar.gz": { + "bytes": 12263, + "sha256": "19b38872c31d2b0594b0b7a316335c306bd7d3dcb5b1704ce6dd24d9e20a0f26" + }, + "cost.json": { + "bytes": 1235, + "sha256": "4977d5a9d01246cfaa7eed59e59de571de7de651dc842c38ad4f48e2259f0723" + }, + "gpu_telemetry.json": { + "bytes": 2445, + "sha256": "5437f792470c7b68c5169f8fd68bcb170b9450fc8efad9303ef1d656ad5ffad8" + }, + "grade.json": { + "bytes": 2156, + "sha256": "0d1c9775790296f34db998a3c53224e4ad0516b0823bfa1cd074561a6b73d6c4" + } + } + }, + { + "run_name": "p2_triage_gemma4-31b-q4_v6", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 95.9, + "iterations": 9, + "completion_tokens": 5565, + "prompt_tokens_cumulative": 63792, + "model_turns": 9, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9385, + "mean_power_w": 486.46, + "mean_sm_util_pct": 84.53, + "max_temp_c": 80.0, + "cpu_package_mean_power_w": 136.64 + }, + "files": { + "receipt.json": { + "bytes": 9505, + "sha256": "65f57fc01a89207848b4b22adbe6b8446417e959f91e48e1d751b9b0a484e3af" + }, + "transcript.jsonl": { + "bytes": 11600, + "sha256": "bcc84426b4704e89554fee02a30cb8637efdb703c9b254c087549d37ebb9d7cb" + }, + "summary.json": { + "bytes": 604, + "sha256": "af527fae100858ffb85a6e472b9a176aa7d9b87abdf4a3a3f7cf17e7762b1468" + }, + "workspace_final.tar.gz": { + "bytes": 12818, + "sha256": "17e33ca4c09393e71638e63516243c145a7b37c7fa3f66963fcf7d27ffa98f3c" + }, + "cost.json": { + "bytes": 1232, + "sha256": "ae718fdb92e5c424f190339ab3330286433884aca39508216e1121a9b8818325" + }, + "gpu_telemetry.json": { + "bytes": 2441, + "sha256": "66188d6e52a52916ae7265381d1d4974f8232cdb279090cb8ccd0e61917fe23d" + }, + "grade.json": { + "bytes": 1823, + "sha256": "3d692968f27afc71beba131cd873661ed419e70ba9094a5b9bba5ede51d9d62d" + } + } + }, + { + "run_name": "p2_triage_gemma4-31b-q4_v7", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 120.8, + "iterations": 9, + "completion_tokens": 7105, + "prompt_tokens_cumulative": 73842, + "model_turns": 9, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.952, + "mean_power_w": 494.35, + "mean_sm_util_pct": 96.04, + "max_temp_c": 74.0, + "cpu_package_mean_power_w": 136.39 + }, + "files": { + "receipt.json": { + "bytes": 9505, + "sha256": "f8acfe1ce61c6f75ab23582fe25d4bbfeb73d1f2d8f1bfce8024fdc922087d6f" + }, + "transcript.jsonl": { + "bytes": 11671, + "sha256": "f5e0ed2051b6bd92ffca8cff0f3a4a5ebd3364d923221712a95cb460bf2443aa" + }, + "summary.json": { + "bytes": 578, + "sha256": "3dbba4f1dcf0baf70bd2c220d0f9235b3d5bea2a9da72118e2361146e9bd94d1" + }, + "workspace_final.tar.gz": { + "bytes": 12973, + "sha256": "28011cc4cc26fd80c315c6f411a6fc9caff3e04b42ddc09cec29172309d2cd0b" + }, + "cost.json": { + "bytes": 1234, + "sha256": "de13bd5ce10ac408b7b6ab94635bb16b39423523009767cfff17dc5a430ab95f" + }, + "gpu_telemetry.json": { + "bytes": 2405, + "sha256": "a6999cea8e144e16511b079fabeff6fc5d6349e8ee44d6c2f8fe6e39b061d14d" + }, + "grade.json": { + "bytes": 2162, + "sha256": "e1f80a2b26dbf4d4a89166e3aeaa1d90b604d68fbb99286aa5e801231483922b" + } + } + }, + { + "run_name": "p2_triage_gemma4-31b-q4_v8", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 104.7, + "iterations": 5, + "completion_tokens": 6262, + "prompt_tokens_cumulative": 28430, + "model_turns": 5, + "length_finishes": 0, + "grade_verdict": "FAIL", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9551, + "mean_power_w": 482.51, + "mean_sm_util_pct": 92.76, + "max_temp_c": 80.0, + "cpu_package_mean_power_w": 136.65 + }, + "files": { + "receipt.json": { + "bytes": 9505, + "sha256": "837e52b0a47327bc03eca109233835ee8bd60db9c087f8ca9631e867860e3bd0" + }, + "transcript.jsonl": { + "bytes": 8578, + "sha256": "9e3e356f4c64fd0a58b1239c3b3c93fb16ea5202e085e39020a6ceaf874398e8" + }, + "summary.json": { + "bytes": 570, + "sha256": "3807ac6c5d0b864bc58d89a987e661114146d39fe1412c375dfe28df75417a42" + }, + "workspace_final.tar.gz": { + "bytes": 12255, + "sha256": "17dc3cf28933e213a06abf2b49b2d6eeea534bae1cf7c9475b0a83058fb747a0" + }, + "cost.json": { + "bytes": 1234, + "sha256": "53ead387028931115784fa256c41d9318f884b509ccf7e18e374f043c712cd66" + }, + "gpu_telemetry.json": { + "bytes": 2474, + "sha256": "537fd02df65f75d3497a5f9afc90feb16345b506b9014a8183b0a93a7dda3e77" + }, + "grade.json": { + "bytes": 2345, + "sha256": "ddd1baea329ab8b63b59b7f1fc949dce4d551b8f25b0816f865b1d3a803879e2" + } + } + }, + { + "run_name": "p2_triage_gemma4-31b-q4_v9", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 84.8, + "iterations": 5, + "completion_tokens": 4907, + "prompt_tokens_cumulative": 24039, + "model_turns": 5, + "length_finishes": 0, + "grade_verdict": "FAIL", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9434, + "mean_power_w": 477.91, + "mean_sm_util_pct": 91.76, + "max_temp_c": 73.0, + "cpu_package_mean_power_w": 136.21 + }, + "files": { + "receipt.json": { + "bytes": 9505, + "sha256": "deec464accbeda6fbbc71e69249b75578b9ee8f84e93558b69a27a6d0f9f84d6" + }, + "transcript.jsonl": { + "bytes": 8958, + "sha256": "a9cfe1575f77cb67e85d46c5e380d430b7386fcdc07caed0e27b4b3188f85dd5" + }, + "summary.json": { + "bytes": 706, + "sha256": "1628b63bd0eec1ddfa20aff2dd6be54ad30b729b709c2d2685b7a6528b0aff85" + }, + "workspace_final.tar.gz": { + "bytes": 12369, + "sha256": "d6664747a7cdcc50bebc0cbe52d054de209466a9fe47a80b46f1903491b16eb8" + }, + "cost.json": { + "bytes": 1232, + "sha256": "af420560860c35b45dd5e26236cc908b68cf5805a669153f661e868b62c492c3" + }, + "gpu_telemetry.json": { + "bytes": 2444, + "sha256": "d04621322d3d56c3583b5ddc9b5531ccdc8e8c974f8bc1e9e67939566b697f77" + }, + "grade.json": { + "bytes": 2444, + "sha256": "8746c950be101673b46c2c22be3c3adfb6f214618b06b31788618933b8d74075" + } + } + }, + { + "run_name": "p2_triage_gemma4-31b-q4_v10", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 115.6, + "iterations": 7, + "completion_tokens": 6794, + "prompt_tokens_cumulative": 47178, + "model_turns": 7, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9516, + "mean_power_w": 490.21, + "mean_sm_util_pct": 94.61, + "max_temp_c": 82.0, + "cpu_package_mean_power_w": 133.95 + }, + "files": { + "receipt.json": { + "bytes": 9506, + "sha256": "188f034758d4425d4f3dd81845bd7431b22ce5f1c9b81ea8102623e2aaca28c7" + }, + "transcript.jsonl": { + "bytes": 10513, + "sha256": "7a81d5e9a3bd9fbd5528e02f930d11ce8b4cb2048d371f23a9b65bc12560d93f" + }, + "summary.json": { + "bytes": 651, + "sha256": "8c5ad6ddd3eafde277a69e4d3df5726364011a7c9bfff9ce6b22e3c435614c04" + }, + "workspace_final.tar.gz": { + "bytes": 12686, + "sha256": "091ddcb1c508d7369b1f674b396b81239676ae09439ee5b2b46f2507ddfa92fd" + }, + "cost.json": { + "bytes": 1235, + "sha256": "d0a735c362a80771f3aae90c379975ea2055f8ab38bdf7f3933dc3fcfcc8d263" + }, + "gpu_telemetry.json": { + "bytes": 2476, + "sha256": "afb0025252586edb3b3aae788182ca441bc2875a12e497dfd712536855f15475" + }, + "grade.json": { + "bytes": 1541, + "sha256": "98c725368d22f6b29f262e70ca2cf1a55a7d70c0409f1ab0087a77a938b08a8d" + } + } + }, + { + "run_name": "p3_doc_gemma4-31b-q4_v1", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 121.8, + "iterations": 8, + "completion_tokens": 7296, + "prompt_tokens_cumulative": 51659, + "model_turns": 8, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9442, + "mean_power_w": 488.83, + "mean_sm_util_pct": 96.71, + "max_temp_c": 84.0, + "cpu_package_mean_power_w": 136.91 + }, + "files": { + "receipt.json": { + "bytes": 8956, + "sha256": "e82000d2c622ccbd80272cb91a940f666af73f187dcfd791bcf0eaba5f2e1dd0" + }, + "transcript.jsonl": { + "bytes": 14483, + "sha256": "57e17b5eb071b15b2e3d1959f41f9823b94ae0166874adaf01c4a60962afaa39" + }, + "summary.json": { + "bytes": 723, + "sha256": "d2180a1ba15cba18e03d8f52b3595c30954f050f376df3b2106c2b99689937f5" + }, + "workspace_final.tar.gz": { + "bytes": 15602, + "sha256": "764bc3a872dd39c331ae67d35d1c5c9c26758f47668eac413222bfad50a7b355" + }, + "cost.json": { + "bytes": 1231, + "sha256": "b975a39a92cec1d99ccf68599a8edf7900d37e269eaa73d100a5ab38b9ff5eb9" + }, + "gpu_telemetry.json": { + "bytes": 2397, + "sha256": "18eb1f96634436dc3ffa748ba602c0bb2f839ba9c75770e5a9ec59f9d1b5ecf5" + }, + "grade.json": { + "bytes": 2357, + "sha256": "addcd3cfc419649494adf40c60fd2385ec3b403f91e07f3423695aac6fd5c0dc" + } + } + }, + { + "run_name": "p3_doc_gemma4-31b-q4_v2", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 123.8, + "iterations": 8, + "completion_tokens": 7296, + "prompt_tokens_cumulative": 51659, + "model_turns": 8, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9693, + "mean_power_w": 480.4, + "mean_sm_util_pct": 92.72, + "max_temp_c": 78.0, + "cpu_package_mean_power_w": 113.16 + }, + "files": { + "receipt.json": { + "bytes": 8954, + "sha256": "2d59a5d6e6bf429d767447461c521914274a8a45763a062dea250be8a7b5a5ec" + }, + "transcript.jsonl": { + "bytes": 14484, + "sha256": "d6a6a9795bbd69d388ab0f058bef562338055664d578af964c95f99a3eb4a667" + }, + "summary.json": { + "bytes": 723, + "sha256": "ce4d10ebe26c6b5612292822982dcb69116f6864744f166d2c388be997cc43b5" + }, + "workspace_final.tar.gz": { + "bytes": 15594, + "sha256": "cc19c24bdd130e1aaa7b2aec944779f0cff243f0eef1b465e0b647243c144d34" + }, + "cost.json": { + "bytes": 1231, + "sha256": "bbf0a3b4f8f96c6a97f15d4030e96f6b44d07d270d97b7610485e09d0f0be90e" + }, + "gpu_telemetry.json": { + "bytes": 2396, + "sha256": "4911f5d3f3a857c539e956035152af549e90404810899fa331bdd5ac540298c1" + }, + "grade.json": { + "bytes": 2357, + "sha256": "addcd3cfc419649494adf40c60fd2385ec3b403f91e07f3423695aac6fd5c0dc" + } + } + }, + { + "run_name": "p3_doc_gemma4-31b-q4_v3", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 113.3, + "iterations": 10, + "completion_tokens": 6632, + "prompt_tokens_cumulative": 67272, + "model_turns": 10, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9709, + "mean_power_w": 469.09, + "mean_sm_util_pct": 89.09, + "max_temp_c": 83.0, + "cpu_package_mean_power_w": 137.02 + }, + "files": { + "receipt.json": { + "bytes": 8956, + "sha256": "4ecf36adec39b7279a65e69a7404cfc30a8da6a0e4318004718259f4af1364f3" + }, + "transcript.jsonl": { + "bytes": 16781, + "sha256": "a8a0b48977fadcb65aaa23a894d973153560173ccf590b3486f64f60fcfcffe2" + }, + "summary.json": { + "bytes": 747, + "sha256": "167d8934c09ec91e3e926ca9499379b34c0b1f6538f6c565eafc01641c42e073" + }, + "workspace_final.tar.gz": { + "bytes": 16341, + "sha256": "fdd57249f4f12c782da3dbb6af706d34983697efabbc32c500c8f0af27c0fad5" + }, + "cost.json": { + "bytes": 1231, + "sha256": "d41d93bd17e42e69cc0100bee0bafa79c063b61e7ce6e41d77b1caf2664e72e3" + }, + "gpu_telemetry.json": { + "bytes": 2401, + "sha256": "3308e8f4410ec7dc3a2a2d7a43505724c8e0559ef8434921eef3db278dfb019d" + }, + "grade.json": { + "bytes": 2360, + "sha256": "a1e32ee1296ef6f8f122b96d91586e2fbd3df7e2ff1f5aea3b13c97c6cb2bdd8" + } + } + }, + { + "run_name": "p3_doc_gemma4-31b-q4_v4", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 115.2, + "iterations": 9, + "completion_tokens": 6893, + "prompt_tokens_cumulative": 59862, + "model_turns": 9, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9549, + "mean_power_w": 497.4, + "mean_sm_util_pct": 97.57, + "max_temp_c": 78.0, + "cpu_package_mean_power_w": 115.67 + }, + "files": { + "receipt.json": { + "bytes": 9514, + "sha256": "b4c8fb61e3b2f22d2b0fe5729fec95bbb1ee733a94d4728a35354f418164f634" + }, + "transcript.jsonl": { + "bytes": 16486, + "sha256": "98ae6867f7cac7458790d42f1e776ddfbef170db9f72fade05890758f9a4ec18" + }, + "summary.json": { + "bytes": 798, + "sha256": "f8eebb1a2a79b4ae1345eb1ba716f64e157f52180662aea6d0f8713f54b273d4" + }, + "workspace_final.tar.gz": { + "bytes": 16231, + "sha256": "b5b46f69e66d16baa405c6bf36809f62383322808a6f725083ce028dd4710cc6" + }, + "cost.json": { + "bytes": 1229, + "sha256": "0681aab8de23f224f0bf0e8ed39e97e65203d3e0b37638f11771200bcc0d4ef5" + }, + "gpu_telemetry.json": { + "bytes": 2393, + "sha256": "bb7ea45789719b3e9d624223b212e2264ef225f7ce4b76e69b48bc034cd6c8e5" + }, + "grade.json": { + "bytes": 2387, + "sha256": "10c1a63e62ecc97494797071010538305d197ff084808363feabdca104bfbbf1" + } + } + }, + { + "run_name": "p3_doc_gemma4-31b-q4_v5", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 89.6, + "iterations": 9, + "completion_tokens": 5286, + "prompt_tokens_cumulative": 51163, + "model_turns": 9, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.8929, + "mean_power_w": 487.26, + "mean_sm_util_pct": 93.88, + "max_temp_c": 73.0, + "cpu_package_mean_power_w": 136.22 + }, + "files": { + "receipt.json": { + "bytes": 9516, + "sha256": "0d5743faf57bad0ca829f7dada7ed6e65cc0c182f99c5b14af88b827e7fad402" + }, + "transcript.jsonl": { + "bytes": 14253, + "sha256": "31994eb228f42fdb8191e10ec1864fa88f407f8150356d4f9f05df5b08be030d" + }, + "summary.json": { + "bytes": 780, + "sha256": "15c35ad2ff036e7065dec018082927a6be87f06649f3a8db2b16926ad6b387e2" + }, + "workspace_final.tar.gz": { + "bytes": 15374, + "sha256": "ba9d8c8996589f7e8eb80e0dcb75776991b7399f7f44dad9f31daccc647fc906" + }, + "cost.json": { + "bytes": 1229, + "sha256": "cf77f209a2f16bf73df827b832ac8dc4dee5f6eabf2bb88e8078822e64a56574" + }, + "gpu_telemetry.json": { + "bytes": 2400, + "sha256": "c064defaa57141e589daf9c3320c53ef599be14ce3f0f27347e48d3413acc0cd" + }, + "grade.json": { + "bytes": 2356, + "sha256": "2b471277132b5cb8888e33dfb28d6609a82f2df57ab635b45f369f157dff3d9d" + } + } + }, + { + "run_name": "p3_doc_gemma4-31b-q4_v6", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 89.8, + "iterations": 9, + "completion_tokens": 5285, + "prompt_tokens_cumulative": 51163, + "model_turns": 9, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9465, + "mean_power_w": 465.8, + "mean_sm_util_pct": 86.56, + "max_temp_c": 72.0, + "cpu_package_mean_power_w": 115.16 + }, + "files": { + "receipt.json": { + "bytes": 9514, + "sha256": "7e2379605a874b0ed4cabca7ab240eda5de7902450628d2d3016dca940a4f326" + }, + "transcript.jsonl": { + "bytes": 14215, + "sha256": "1f3a6fc93a377727865341dec49fe5d083306c4ec82eb1c5a1c04c0987327fae" + }, + "summary.json": { + "bytes": 743, + "sha256": "5fac8f50b9ca901f60290b4896ff7b4dd970c744d1b4d742d545a04f8f45d98d" + }, + "workspace_final.tar.gz": { + "bytes": 15368, + "sha256": "6b628c0dfd03fb4b1a0929acfc3b69c1415a365304cb8b915c751d4a419815bb" + }, + "cost.json": { + "bytes": 1229, + "sha256": "36a2e506794738a0d7bbd9c80d71544b9bc5a47123d07483b0ee8ee9b6754677" + }, + "gpu_telemetry.json": { + "bytes": 2392, + "sha256": "fafb9bc7ee7727ae243ca2915972513f7afaf1bd98644c17438617b9958e02bb" + }, + "grade.json": { + "bytes": 2356, + "sha256": "2b471277132b5cb8888e33dfb28d6609a82f2df57ab635b45f369f157dff3d9d" + } + } + }, + { + "run_name": "p3_doc_gemma4-31b-q4_v7", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 92.3, + "iterations": 9, + "completion_tokens": 5286, + "prompt_tokens_cumulative": 51163, + "model_turns": 9, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9209, + "mean_power_w": 479.29, + "mean_sm_util_pct": 88.06, + "max_temp_c": 73.0, + "cpu_package_mean_power_w": 135.96 + }, + "files": { + "receipt.json": { + "bytes": 9516, + "sha256": "d5b419f67e639fc1a31f7b3093bbf3cce8c86d77c9a28cb1127165abbf8aa015" + }, + "transcript.jsonl": { + "bytes": 14254, + "sha256": "16055c7138bea0bf7ab267227b652877200ff86120a2edb009fa75e24f56fc93" + }, + "summary.json": { + "bytes": 780, + "sha256": "617cea409739e12948a5b2195c438439d20848d3c98728366b4c9bd2a4f4fedf" + }, + "workspace_final.tar.gz": { + "bytes": 15363, + "sha256": "facc0c2f8bde839c5a0c05122660ff0978864f5639c83316e8ac849dcac54b06" + }, + "cost.json": { + "bytes": 1229, + "sha256": "c654b14ac00e088201c348ee8283b7a9c2582416d27fe6c0e978d245261241af" + }, + "gpu_telemetry.json": { + "bytes": 2438, + "sha256": "8487db216def1da0354596910585d1c3976fc177220d281ed0e523b65e51105c" + }, + "grade.json": { + "bytes": 2356, + "sha256": "2b471277132b5cb8888e33dfb28d6609a82f2df57ab635b45f369f157dff3d9d" + } + } + }, + { + "run_name": "p3_doc_gemma4-31b-q4_v8", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 90.9, + "iterations": 9, + "completion_tokens": 5285, + "prompt_tokens_cumulative": 51163, + "model_turns": 9, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9351, + "mean_power_w": 482.28, + "mean_sm_util_pct": 92.44, + "max_temp_c": 71.0, + "cpu_package_mean_power_w": 114.59 + }, + "files": { + "receipt.json": { + "bytes": 9514, + "sha256": "0cb6c75a72965eb782e0f60903290b1d96bd27060750972814b369ededfc791b" + }, + "transcript.jsonl": { + "bytes": 14215, + "sha256": "000ebd601fe6133674ab38dec3041aa7911440b1b56a74bf5cd81f127b7c50d4" + }, + "summary.json": { + "bytes": 743, + "sha256": "9963113a8e4e12e8056074e960983fcf0ceaf120893233b09b20ffbe0a544d6a" + }, + "workspace_final.tar.gz": { + "bytes": 15383, + "sha256": "435052324926ce154fd34bedb0e0dd6a6069bd7dad2e0cb94a3af32b1034d8df" + }, + "cost.json": { + "bytes": 1229, + "sha256": "67c020c765e50185134b18f80d64d98cb30528e765ff140d8ac9be005f588136" + }, + "gpu_telemetry.json": { + "bytes": 2393, + "sha256": "2a3ddef86b162314640652fb9afd584df75a50dc271dfb93d5cc4bd1997793b1" + }, + "grade.json": { + "bytes": 2356, + "sha256": "2b471277132b5cb8888e33dfb28d6609a82f2df57ab635b45f369f157dff3d9d" + } + } + }, + { + "run_name": "p3_doc_gemma4-31b-q4_v9", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 120.1, + "iterations": 8, + "completion_tokens": 7296, + "prompt_tokens_cumulative": 51659, + "model_turns": 8, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9575, + "mean_power_w": 488.73, + "mean_sm_util_pct": 92.38, + "max_temp_c": 74.0, + "cpu_package_mean_power_w": 136.22 + }, + "files": { + "receipt.json": { + "bytes": 9516, + "sha256": "9396166ece047f55a1d5b31af0cc51089a03e0203577b65a461e3f5deb1bab4a" + }, + "transcript.jsonl": { + "bytes": 14484, + "sha256": "92b413e6f77c0544d3cc396eb10ea90853e230e7bffdb1e780fafd905e169e85" + }, + "summary.json": { + "bytes": 723, + "sha256": "1423f42cc77739722fdcef73d4c842bdc89c7dc6d40dcc458ec52a95de514352" + }, + "workspace_final.tar.gz": { + "bytes": 15607, + "sha256": "09897075b21ed03531c6ba079533443f94cf481b61cd7ca01520411f030e26dc" + }, + "cost.json": { + "bytes": 1231, + "sha256": "2963cacb6d36d51be470d2395533876696411f648683357d3b7bcef0734c5f85" + }, + "gpu_telemetry.json": { + "bytes": 2442, + "sha256": "46c8db282b17a8c0ea59f3426b90add88e072fb8c167f7b142e930ec74fa1d99" + }, + "grade.json": { + "bytes": 2357, + "sha256": "addcd3cfc419649494adf40c60fd2385ec3b403f91e07f3423695aac6fd5c0dc" + } + } + }, + { + "run_name": "p3_doc_gemma4-31b-q4_v10", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 120.0, + "iterations": 8, + "completion_tokens": 7296, + "prompt_tokens_cumulative": 51659, + "model_turns": 8, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9583, + "mean_power_w": 486.67, + "mean_sm_util_pct": 95.25, + "max_temp_c": 71.0, + "cpu_package_mean_power_w": 114.52 + }, + "files": { + "receipt.json": { + "bytes": 9515, + "sha256": "537e8d9748e2ea6462296b950f4fcce39ea87d53b13a2a0ec5807347200890a9" + }, + "transcript.jsonl": { + "bytes": 14483, + "sha256": "53ccfb22f354feb5a6b558276b80bfb077c7f5ccf58bd09a706a782de0029354" + }, + "summary.json": { + "bytes": 723, + "sha256": "6b14aa64bc8a818ade457ed74a490df80de078571561e95e311c8371ceb41b79" + }, + "workspace_final.tar.gz": { + "bytes": 15593, + "sha256": "92419bc798e6e9e41e711a0a9a5f0940e2045d0f51a32e7244ec000169993f14" + }, + "cost.json": { + "bytes": 1232, + "sha256": "113c78c1cc79aabac155372a5699198509a0734749f8578ffd757fd0610fcb65" + }, + "gpu_telemetry.json": { + "bytes": 2398, + "sha256": "9d558a561ebdb5da84a655806bbc8ec3d226661cefbd9648edd86eda03e71dd7" + }, + "grade.json": { + "bytes": 2357, + "sha256": "addcd3cfc419649494adf40c60fd2385ec3b403f91e07f3423695aac6fd5c0dc" + } + } + }, + { + "run_name": "p3_business_gemma4-31b-q4_v1", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 94.2, + "iterations": 9, + "completion_tokens": 5487, + "prompt_tokens_cumulative": 43691, + "model_turns": 9, + "length_finishes": 0, + "grade_verdict": "FAIL", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9554, + "mean_power_w": 458.24, + "mean_sm_util_pct": 78.84, + "max_temp_c": 77.0, + "cpu_package_mean_power_w": 112.85 + }, + "files": { + "receipt.json": { + "bytes": 8959, + "sha256": "ffcb5af2fb2c4b8fa8cc26fc21c02c64cb905c1f267f576775bbfaee95408fc2" + }, + "transcript.jsonl": { + "bytes": 16635, + "sha256": "224abc79829aa387ea42f89a7e76e1349021dc086196fbcd7f7446794ae12cc8" + }, + "summary.json": { + "bytes": 956, + "sha256": "ea4c6d11864bda3ae075daffecc0989616185dacec6e981fa94c434e643aa988" + }, + "workspace_final.tar.gz": { + "bytes": 15908, + "sha256": "1217b662e553531a731c4cfe9596bae3d1ccbc80752dd7ac12c76da894d65d90" + }, + "cost.json": { + "bytes": 1234, + "sha256": "04e7b3f606e02679842522ae4eeb0ccd3d4b61f87f41e8eee7b38d1bf98e6bcc" + }, + "gpu_telemetry.json": { + "bytes": 2399, + "sha256": "37481abd2df467e553784745df98d2536448310645d470f6ee36ea1971fac98d" + }, + "grade.json": { + "bytes": 3866, + "sha256": "3c85cb074672401082ba6adecc5f454e5124b81323f35fc312a8458bba0ab815" + } + } + }, + { + "run_name": "p3_business_gemma4-31b-q4_v2", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 80.9, + "iterations": 9, + "completion_tokens": 4742, + "prompt_tokens_cumulative": 37120, + "model_turns": 9, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9271, + "mean_power_w": 495.91, + "mean_sm_util_pct": 95.5, + "max_temp_c": 83.0, + "cpu_package_mean_power_w": 137.15 + }, + "files": { + "receipt.json": { + "bytes": 8961, + "sha256": "6b54b92c8c0396ffbfbb5ed2f8907b802e3fe4d2e218cb9c888c9fe80f9b3f08" + }, + "transcript.jsonl": { + "bytes": 16134, + "sha256": "75a2f519c54888667c21ad3e3b021569ae835ad81fd233e100f56c1713549b45" + }, + "summary.json": { + "bytes": 853, + "sha256": "3142c3c654573d7bed6174124196fbce6facfb39fae3fbf95fda122e731f626c" + }, + "workspace_final.tar.gz": { + "bytes": 15801, + "sha256": "7a319197444df2e8d1b271a7ef4091fb663cf67241f120057cf084fddf1aaeb0" + }, + "cost.json": { + "bytes": 1234, + "sha256": "ca67bd0d05b32d34219d90fa1100d5963baf1dd028d1b618f9d460e27329bcfb" + }, + "gpu_telemetry.json": { + "bytes": 2399, + "sha256": "a45c818c4de903744d7db2bb65db093e905615481c84651cbd7f687df8d9baf8" + }, + "grade.json": { + "bytes": 3873, + "sha256": "13358399cce26683f59886f1cfc577330fd90233e4065a9af249f51edd05c4ab" + } + } + }, + { + "run_name": "p3_business_gemma4-31b-q4_v3", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 79.7, + "iterations": 9, + "completion_tokens": 4742, + "prompt_tokens_cumulative": 37120, + "model_turns": 9, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.941, + "mean_power_w": 486.72, + "mean_sm_util_pct": 93.75, + "max_temp_c": 74.0, + "cpu_package_mean_power_w": 112.93 + }, + "files": { + "receipt.json": { + "bytes": 8959, + "sha256": "06fb3e91e1166b84c55cf7a70f66876735c67cf8e32d2e664bdd0bb45a26d035" + }, + "transcript.jsonl": { + "bytes": 16133, + "sha256": "e1cab5771101129766dc6ec65ae6d633df275d7a54929e1a3da8cf4d579e1571" + }, + "summary.json": { + "bytes": 853, + "sha256": "31fd9a8af67581a1a48bebb0450619fb80d03229cdf321ba55560a4421a54afd" + }, + "workspace_final.tar.gz": { + "bytes": 15805, + "sha256": "8abb097c9a0a7bfbd30f59b051be46e5907e13d1c85c9d2c5769a520dc532886" + }, + "cost.json": { + "bytes": 1234, + "sha256": "3cf0e3376dc28702234453ff8f4e5aeafebe87de47c33d5985e4bc4fabaf63a0" + }, + "gpu_telemetry.json": { + "bytes": 2396, + "sha256": "9cf47b55a9a271beb0961d85b920b6774a8d89e5eec9705f256c93d73d6e26f6" + }, + "grade.json": { + "bytes": 3873, + "sha256": "13358399cce26683f59886f1cfc577330fd90233e4065a9af249f51edd05c4ab" + } + } + }, + { + "run_name": "p3_business_gemma4-31b-q4_v4", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 91.8, + "iterations": 9, + "completion_tokens": 5487, + "prompt_tokens_cumulative": 43691, + "model_turns": 9, + "length_finishes": 0, + "grade_verdict": "FAIL", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9804, + "mean_power_w": 478.13, + "mean_sm_util_pct": 92.42, + "max_temp_c": 70.0, + "cpu_package_mean_power_w": 114.53 + }, + "files": { + "receipt.json": { + "bytes": 9519, + "sha256": "aa3ef47bdd9f935e37ca626fae0efa84bfd9d2edd04d7db78540c7900d433216" + }, + "transcript.jsonl": { + "bytes": 16634, + "sha256": "5ee9fadb5828470ca1e87b3bac62c1235941786d44f0a8abf75906deb2c8067f" + }, + "summary.json": { + "bytes": 956, + "sha256": "74dcc9cde42ed9fde56910ed0ed92a7cac2d545a14d38bcf81f3dee24ddc187c" + }, + "workspace_final.tar.gz": { + "bytes": 15911, + "sha256": "6157ce44a04fbf479c93712b06b5ecdda2a941ba87224749ddd3653cfaec707a" + }, + "cost.json": { + "bytes": 1234, + "sha256": "d0ba10417a0b1f7eab661151daed8061be33b8fa1b76ecb7109448062c13faa0" + }, + "gpu_telemetry.json": { + "bytes": 2396, + "sha256": "5182742cfd9ac61cfbfc38cdc35fb375d5ce8a3df2ceaa4f78ede21f3d934b40" + }, + "grade.json": { + "bytes": 3866, + "sha256": "3c85cb074672401082ba6adecc5f454e5124b81323f35fc312a8458bba0ab815" + } + } + }, + { + "run_name": "p3_business_gemma4-31b-q4_v5", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 91.7, + "iterations": 9, + "completion_tokens": 5487, + "prompt_tokens_cumulative": 43691, + "model_turns": 9, + "length_finishes": 0, + "grade_verdict": "FAIL", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9269, + "mean_power_w": 500.08, + "mean_sm_util_pct": 97.22, + "max_temp_c": 73.0, + "cpu_package_mean_power_w": 136.27 + }, + "files": { + "receipt.json": { + "bytes": 9521, + "sha256": "1dffa5238d51ddd7da0c0a45069b85eedaff1f0ce336b49da0de85fb27da7558" + }, + "transcript.jsonl": { + "bytes": 16635, + "sha256": "75cbac567377121b4eaed2fac81fb433823581380c80e2a69351bee608298213" + }, + "summary.json": { + "bytes": 956, + "sha256": "0f090f1ebc68b18a8952877b1f4cd3f47939cdff2cfa4fc0efb6f83fc00da19f" + }, + "workspace_final.tar.gz": { + "bytes": 15912, + "sha256": "02c2c43e93cc97046fcb6d5b4adb229a7e88f7dfb3b12b7ed9f83ec044ea73aa" + }, + "cost.json": { + "bytes": 1234, + "sha256": "5d8b7a140f4deb426bab064ac024763d63a181faa937ed5aeb3283198bf808cd" + }, + "gpu_telemetry.json": { + "bytes": 2480, + "sha256": "109e6279a30bfd78a1f7a8b39133e03291243a6993f99b5c22f9533e3d77b0af" + }, + "grade.json": { + "bytes": 3866, + "sha256": "3c85cb074672401082ba6adecc5f454e5124b81323f35fc312a8458bba0ab815" + } + } + }, + { + "run_name": "p3_business_gemma4-31b-q4_v6", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 85.8, + "iterations": 9, + "completion_tokens": 5087, + "prompt_tokens_cumulative": 38550, + "model_turns": 9, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9324, + "mean_power_w": 484.2, + "mean_sm_util_pct": 90.18, + "max_temp_c": 70.0, + "cpu_package_mean_power_w": 114.57 + }, + "files": { + "receipt.json": { + "bytes": 9519, + "sha256": "d87c1278cceb626269647e3c476b803b8b2fbbc574c30740695309d097ac68ce" + }, + "transcript.jsonl": { + "bytes": 15990, + "sha256": "6ed65e282be4a7e65180b1420beef2c2c61498a004b15fea1e6354cc5f4fcca7" + }, + "summary.json": { + "bytes": 1012, + "sha256": "beba17a682fbe29a38fe9859bc489ca9ce00b9ba919e63db54ab9450d809943c" + }, + "workspace_final.tar.gz": { + "bytes": 15519, + "sha256": "b2bd38d90c1ad2cc26f446beabf842b48837a80fb78138ce8dab25f9da2001ab" + }, + "cost.json": { + "bytes": 1234, + "sha256": "570fe8b58c53deb99e836af7e5f83ff54021d33b42809fc4b4bcb87d1042e131" + }, + "gpu_telemetry.json": { + "bytes": 2397, + "sha256": "f3a5bb0e76466547098e67c45b4fd252edc0c6fadbfd8c21f662ce15b1ee389a" + }, + "grade.json": { + "bytes": 3870, + "sha256": "8a18dd536897e3e16058a2b4c0f9f443faeeb0b1c71986a3a75245703cd484d1" + } + } + }, + { + "run_name": "p3_business_gemma4-31b-q4_v7", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 88.6, + "iterations": 9, + "completion_tokens": 5087, + "prompt_tokens_cumulative": 38550, + "model_turns": 9, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9594, + "mean_power_w": 465.79, + "mean_sm_util_pct": 84.89, + "max_temp_c": 73.0, + "cpu_package_mean_power_w": 136.4 + }, + "files": { + "receipt.json": { + "bytes": 9521, + "sha256": "186d2ed1efa1ac8e98a410b8668e4e618b625ce60559a2e00459c93cbf39faa5" + }, + "transcript.jsonl": { + "bytes": 15989, + "sha256": "adf6b3f875cb052404985fb574fba382e4faf5495b7e10fa9a041881cc7f66e4" + }, + "summary.json": { + "bytes": 1012, + "sha256": "ab5bf56eca72833790df755eebdc8570dcb440daf0d5c7516cccab9cf9fecc8a" + }, + "workspace_final.tar.gz": { + "bytes": 15515, + "sha256": "141af935a4a8cb32f37754177de6d59716a3f2e66b5aac92dc1066e71d399212" + }, + "cost.json": { + "bytes": 1234, + "sha256": "23b6e806a559819367e600bbc37c7be808846b663c9f27f26c4fc07f920cedea" + }, + "gpu_telemetry.json": { + "bytes": 2396, + "sha256": "8c818607628e44374cbaca6bf416d4288c2a382c05e73a177c562c1f049ca2b7" + }, + "grade.json": { + "bytes": 3870, + "sha256": "8a18dd536897e3e16058a2b4c0f9f443faeeb0b1c71986a3a75245703cd484d1" + } + } + }, + { + "run_name": "p3_business_gemma4-31b-q4_v8", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 87.2, + "iterations": 9, + "completion_tokens": 5087, + "prompt_tokens_cumulative": 38550, + "model_turns": 9, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9174, + "mean_power_w": 487.37, + "mean_sm_util_pct": 93.59, + "max_temp_c": 70.0, + "cpu_package_mean_power_w": 113.76 + }, + "files": { + "receipt.json": { + "bytes": 9519, + "sha256": "bb37f3c999c5d503cb4fdacfbef36c66c676c864656a99b5951ae50d2678adcf" + }, + "transcript.jsonl": { + "bytes": 15988, + "sha256": "98a3e3a547636266bf1c395892e488f8d4731b40135ec5a014f970e3beae9517" + }, + "summary.json": { + "bytes": 1012, + "sha256": "335b0e782ffe8382844a4ab442c47a4ec6d632610248981b9bf709b9d6b69b12" + }, + "workspace_final.tar.gz": { + "bytes": 15516, + "sha256": "11703eac3e1f9045f032c3a748b68e22d70928b218385e1264f2ba7492e367ed" + }, + "cost.json": { + "bytes": 1234, + "sha256": "e8110c52c1b6d4373d9927af2867073c03890bf842e50e185da0ddb5957150ae" + }, + "gpu_telemetry.json": { + "bytes": 2398, + "sha256": "5e0a107b763096c5d9b4f2d0926f035373ae14b1be3d61e3ec6622a595b46dd7" + }, + "grade.json": { + "bytes": 3870, + "sha256": "8a18dd536897e3e16058a2b4c0f9f443faeeb0b1c71986a3a75245703cd484d1" + } + } + }, + { + "run_name": "p3_business_gemma4-31b-q4_v9", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 87.8, + "iterations": 11, + "completion_tokens": 5171, + "prompt_tokens_cumulative": 49232, + "model_turns": 11, + "length_finishes": 0, + "grade_verdict": "FAIL", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9681, + "mean_power_w": 489.8, + "mean_sm_util_pct": 93.06, + "max_temp_c": 73.0, + "cpu_package_mean_power_w": 136.11 + }, + "files": { + "receipt.json": { + "bytes": 9521, + "sha256": "2abce7406328dbaeeb6c44fbabf5cdd510261f6945b81747294e8f43e754665e" + }, + "transcript.jsonl": { + "bytes": 17806, + "sha256": "a2bd015ad083ee9f2212585bb4e604c5cc7fde5aa3ee240457356b4f37893452" + }, + "summary.json": { + "bytes": 864, + "sha256": "7df4b39d8270f9f3150559beb883ff2689226c5c769b6f8f5bf1f875e62020c2" + }, + "workspace_final.tar.gz": { + "bytes": 15922, + "sha256": "51a9d31265156ccbfe3ee00f7322d996779c68c64ff210bede9004c8b8e05995" + }, + "cost.json": { + "bytes": 1235, + "sha256": "ef2e64eb5b1950f4004a24a620006ee92ae7541089a11ea15fb5a01629f0a53e" + }, + "gpu_telemetry.json": { + "bytes": 2433, + "sha256": "664c0a0d9784768a264b368a0da4c0949e426fcbccfd0fc87d0105f27d867099" + }, + "grade.json": { + "bytes": 3854, + "sha256": "8cd481bebf1c8c028e4ea2ea8a6157a3e21f758053374f05a45b73457036e209" + } + } + }, + { + "run_name": "p3_business_gemma4-31b-q4_v10", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 84.9, + "iterations": 9, + "completion_tokens": 5087, + "prompt_tokens_cumulative": 38550, + "model_turns": 9, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9423, + "mean_power_w": 480.52, + "mean_sm_util_pct": 87.24, + "max_temp_c": 70.0, + "cpu_package_mean_power_w": 114.42 + }, + "files": { + "receipt.json": { + "bytes": 9520, + "sha256": "8fa648a7ebba901efb0fb172d302550434ac1dafe51447342b8f6597cc0ab6d5" + }, + "transcript.jsonl": { + "bytes": 15988, + "sha256": "57a57dde88d784f3da15fefd45d1bb096a80fd6c91b0ba1f6f5c8d8521fc0e96" + }, + "summary.json": { + "bytes": 1012, + "sha256": "6e350df11dde9f13d6a303506c7c46d72f5572b523a3a07b3af69a12c794b221" + }, + "workspace_final.tar.gz": { + "bytes": 15509, + "sha256": "70107235a0f6622cffade5d667aa8bbc156bf490311137b707e178a79c66fa6a" + }, + "cost.json": { + "bytes": 1235, + "sha256": "8aa9d5417ffae5c0b244ec90b8312719d4dca477285605d8988c67d097d12377" + }, + "gpu_telemetry.json": { + "bytes": 2399, + "sha256": "9ef476c2c62eb3b5ac5d0871883a9e0e50729ec9522736d208c5892d16e3426d" + }, + "grade.json": { + "bytes": 3870, + "sha256": "8a18dd536897e3e16058a2b4c0f9f443faeeb0b1c71986a3a75245703cd484d1" + } + } + }, + { + "run_name": "p3_market_gemma4-31b-q4_v1", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 257.4, + "iterations": 16, + "completion_tokens": 8071, + "prompt_tokens_cumulative": 387754, + "model_turns": 16, + "length_finishes": 0, + "grade_verdict": "STRUCTURAL_PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9713, + "mean_power_w": 495.65, + "mean_sm_util_pct": 96.29, + "max_temp_c": 85.0, + "cpu_package_mean_power_w": 136.75 + }, + "files": { + "receipt.json": { + "bytes": 8928, + "sha256": "02bb435a5b4fdf5ce444c727c2aab9dc982c3820eb7c8ce22ae0aea7d06e9e67" + }, + "transcript.jsonl": { + "bytes": 22003, + "sha256": "f2f2cd2453d41f8c6c8dca98034f64ef3c90cd8d0ce1d6e0ff256a43f2b1fc4b" + }, + "summary.json": { + "bytes": 747, + "sha256": "903cf4fde70a72420d9f5a24d98a9bc7512675e5e9bd1c498b5f4066af5cff60" + }, + "workspace_final.tar.gz": { + "bytes": 264891, + "sha256": "961f9f5fed7d5aca0767339337210745112837366ca5c413405cb20dfdf7fd4f" + }, + "cost.json": { + "bytes": 1236, + "sha256": "640a00660e7718909fe2dc2eff09af048db6f6d20efae3b74af8a50565bc6fbe" + }, + "gpu_telemetry.json": { + "bytes": 2402, + "sha256": "e31c3c2a90bc6c44cc8ba91a902dd561efcb9d3a875b4d5cfb64eda839465eb7" + }, + "grade.json": { + "bytes": 991, + "sha256": "ec7f817b56a4a6ab2128042d09466fdcdf5b77eaf246ea588cbb987a0865d5fa" + } + } + }, + { + "run_name": "p3_market_gemma4-31b-q4_v2", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 194.1, + "iterations": 20, + "completion_tokens": 9028, + "prompt_tokens_cumulative": 442437, + "model_turns": 20, + "length_finishes": 0, + "grade_verdict": "STRUCTURAL_PASS", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9531, + "mean_power_w": 485.18, + "mean_sm_util_pct": 93.97, + "max_temp_c": 75.0, + "cpu_package_mean_power_w": 112.49 + }, + "files": { + "receipt.json": { + "bytes": 8926, + "sha256": "723d063398588013f7b6052082026114bf29dec559f48dd1eb1e904a1f06ff2e" + }, + "transcript.jsonl": { + "bytes": 23827, + "sha256": "98c92be6cf1e6313be373078a28ab44174526e368a03f53ef5580c75917e6f4e" + }, + "summary.json": { + "bytes": 687, + "sha256": "5e5ccb999a51bb167d790eacd190e63c6384a465cb7bb2bdc7337fccde0858df" + }, + "workspace_final.tar.gz": { + "bytes": 264699, + "sha256": "18d375a961d52e800e0bc7a2445f3a61e29e08236b503731f971a08b6c179d15" + }, + "cost.json": { + "bytes": 1235, + "sha256": "9a0175ff5a729a9a190b9930f3690f1ac68398e2deb8ec6f64820172c4936d85" + }, + "gpu_telemetry.json": { + "bytes": 2398, + "sha256": "25b78c0e41afca29cf8763e0b21dcf4d0d8ee2a8cba88a199337518de807c2d5" + }, + "grade.json": { + "bytes": 991, + "sha256": "7a993d6b3110622d5558f7d78f2ebb57044bb1bdf2b0e9f52585b7fdcf81dddc" + } + } + }, + { + "run_name": "p3_market_gemma4-31b-q4_v3", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 322.0, + "iterations": 15, + "completion_tokens": 8059, + "prompt_tokens_cumulative": 271715, + "model_turns": 15, + "length_finishes": 0, + "grade_verdict": "STRUCTURAL_PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9783, + "mean_power_w": 489.71, + "mean_sm_util_pct": 92.86, + "max_temp_c": 86.0, + "cpu_package_mean_power_w": 137.66 + }, + "files": { + "receipt.json": { + "bytes": 8928, + "sha256": "46c0620b93bcdda525541bfd2c4dfb39ea289d92e44059663e740426c2278697" + }, + "transcript.jsonl": { + "bytes": 21639, + "sha256": "28283464e8685547b97fd7307796460573417109ae9ddaf10cf4cbe1ae6060d3" + }, + "summary.json": { + "bytes": 783, + "sha256": "7d159cccc3186adc8b392d9fcb3653164f567c12bb90844c0629604d66de0a1a" + }, + "workspace_final.tar.gz": { + "bytes": 277200, + "sha256": "576a9c8abf100141ecb229b88b7807074aa42e0bdf2dbe9974dfcc9cfea4e980" + }, + "cost.json": { + "bytes": 1237, + "sha256": "cbaa088ac34943052570e107ff00db849be27af9f0da3a2d16e6704231925ece" + }, + "gpu_telemetry.json": { + "bytes": 2404, + "sha256": "7cd774138375ef6702f9f431e7c975db6af8cde48251eff7d9a498717c491685" + }, + "grade.json": { + "bytes": 990, + "sha256": "5d1bcb2c532b0323a5acdb93607cd0a957b6254f0dc687a6751be84f10ba86b2" + } + } + }, + { + "run_name": "p3_market_gemma4-31b-q4_v4", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 172.4, + "iterations": 19, + "completion_tokens": 8745, + "prompt_tokens_cumulative": 278772, + "model_turns": 19, + "length_finishes": 0, + "grade_verdict": "STRUCTURAL_PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9571, + "mean_power_w": 476.49, + "mean_sm_util_pct": 90.62, + "max_temp_c": 70.0, + "cpu_package_mean_power_w": 113.04 + }, + "files": { + "receipt.json": { + "bytes": 9486, + "sha256": "5bf9d5b9810eb1b3b719b6d10f4fd45a044414609a8ddeee5370cc08ad6ecc13" + }, + "transcript.jsonl": { + "bytes": 22843, + "sha256": "0864d3b4aa465fe9002e2d7c456e61494648ba75eb195fbb67597f88d1e9176e" + }, + "summary.json": { + "bytes": 705, + "sha256": "068ff3e04916018f3faee75710f0b9ef21cfc47410ca887138d80be56f077a6a" + }, + "workspace_final.tar.gz": { + "bytes": 464267, + "sha256": "666fd709a25869d57becfed65ad2d613fd4b3a9b6eab3cdef46002cf013d1067" + }, + "cost.json": { + "bytes": 1236, + "sha256": "03cffbc6b961b1c12395bf09e2cac108da081dfbe844b3fa385d677ab9fc51de" + }, + "gpu_telemetry.json": { + "bytes": 2400, + "sha256": "bd0be75ab46eee5f38a024c52ecf24638bba1154dd96618290778d79398eb475" + }, + "grade.json": { + "bytes": 991, + "sha256": "3aa27fea1e0bb4aa05ff5805aa71f75268e0d1f32d9a65ac321b2fbbfed88ecd" + } + } + }, + { + "run_name": "p3_market_gemma4-31b-q4_v5", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 154.6, + "iterations": 14, + "completion_tokens": 7095, + "prompt_tokens_cumulative": 302302, + "model_turns": 14, + "length_finishes": 0, + "grade_verdict": "STRUCTURAL_PASS", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9702, + "mean_power_w": 482.93, + "mean_sm_util_pct": 89.84, + "max_temp_c": 75.0, + "cpu_package_mean_power_w": 135.48 + }, + "files": { + "receipt.json": { + "bytes": 9488, + "sha256": "a15634a38b9f2a5cfb4716bc0bd2dac38e3d5df52e80c5df870b8d352ee2a94c" + }, + "transcript.jsonl": { + "bytes": 19536, + "sha256": "b4c18c8cb40fb9405d56dc39d429fd71bffdfef66ea6f016de927bf8d3c8bc61" + }, + "summary.json": { + "bytes": 799, + "sha256": "9cf90014cb218c160698703ce3857932dedc335b9f42029a1549e94fde2abd17" + }, + "workspace_final.tar.gz": { + "bytes": 276921, + "sha256": "c1b16e89fe90451c286809630ab3cfeb630d2928ea4f372375bbbbc0bb518cef" + }, + "cost.json": { + "bytes": 1237, + "sha256": "2d2f5e7b0b63bbe65032f1869b6df23ab55fd2adc2cbff52abe9e8211dae58ef" + }, + "gpu_telemetry.json": { + "bytes": 2471, + "sha256": "ffa3e80e5b0fc790b79a29375db04dff8e1078ac29374edf48b5090537fedf2a" + }, + "grade.json": { + "bytes": 991, + "sha256": "3234ae34a10c1abc8642ebfef870e65a48e9fe8740e174c611fb51428cb6b334" + } + } + }, + { + "run_name": "p3_market_gemma4-31b-q4_v6", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 208.5, + "iterations": 26, + "completion_tokens": 8866, + "prompt_tokens_cumulative": 772802, + "model_turns": 26, + "length_finishes": 0, + "grade_verdict": "STRUCTURAL_PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9832, + "mean_power_w": 468.1, + "mean_sm_util_pct": 82.38, + "max_temp_c": 71.0, + "cpu_package_mean_power_w": 113.69 + }, + "files": { + "receipt.json": { + "bytes": 9486, + "sha256": "135c139ba4b0badb0fde0cf546d1075405c3bbdc31ad48ff6bc8c15fd6dd880f" + }, + "transcript.jsonl": { + "bytes": 25371, + "sha256": "f4c7de474575e4da7f0fdb29e8986867864764facc595c5f2cfb030767a03b55" + }, + "summary.json": { + "bytes": 853, + "sha256": "e7a6df965c50f9b4524f4d52a12c83ba4c9790dfe6c677814572f03600b158aa" + }, + "workspace_final.tar.gz": { + "bytes": 762671, + "sha256": "b335edfb5f7a1a567cd429bf93abfd4611e8c88def537b0eb6c19348f5d1fc37" + }, + "cost.json": { + "bytes": 1235, + "sha256": "ce04c67c6c70b5590cce73714ae5f024757c033345a283aa6a6723635827d102" + }, + "gpu_telemetry.json": { + "bytes": 2399, + "sha256": "fe088dc4837afc106ebd0b07d778a55b2d6ed2ec437505f9acd739a9925e53d5" + }, + "grade.json": { + "bytes": 990, + "sha256": "bb5f60e04d85a7cf9d8aa866f0ef183cc643209b16f8d8cce14ec67c1c223b49" + } + } + }, + { + "run_name": "p3_market_gemma4-31b-q4_v7", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 231.2, + "iterations": 39, + "completion_tokens": 9745, + "prompt_tokens_cumulative": 989064, + "model_turns": 39, + "length_finishes": 0, + "grade_verdict": "STRUCTURAL_PASS", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9732, + "mean_power_w": 453.29, + "mean_sm_util_pct": 80.96, + "max_temp_c": 73.0, + "cpu_package_mean_power_w": 135.2 + }, + "files": { + "receipt.json": { + "bytes": 9488, + "sha256": "9b609df5aa26513566cd3cd91da100e576b05c505716fe6b297c580c91b9ab51" + }, + "transcript.jsonl": { + "bytes": 33719, + "sha256": "b3dff3b6d5dbef664dbb7beb61989952bd3491faeeef317d9f504b1292171621" + }, + "summary.json": { + "bytes": 1013, + "sha256": "c5397da75bf4bcd66f7786ad540a12c8a0d80c3839ef637d3a3e156ff1695989" + }, + "workspace_final.tar.gz": { + "bytes": 32409, + "sha256": "2b05f2b9c0f476507c8e3800c40509009b31cec76d9cb3d86c2b0201f3dcd367" + }, + "cost.json": { + "bytes": 1238, + "sha256": "18f33c2c2cfa37169a3ad595af9e0603dc9dbd462026c00e17a0d9e0e57eb525" + }, + "gpu_telemetry.json": { + "bytes": 2538, + "sha256": "8222dc6797d227dc7e4eea8ea8f3af3c25252e73c51edbd27a460dfb4919aa2b" + }, + "grade.json": { + "bytes": 990, + "sha256": "5b462c433ddb337589a2f4efea0e65c99a966b6941f3f242231f69977831f002" + } + } + }, + { + "run_name": "p3_market_gemma4-31b-q4_v8", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 166.0, + "iterations": 17, + "completion_tokens": 8104, + "prompt_tokens_cumulative": 295610, + "model_turns": 17, + "length_finishes": 0, + "grade_verdict": "STRUCTURAL_PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9639, + "mean_power_w": 452.41, + "mean_sm_util_pct": 85.27, + "max_temp_c": 71.0, + "cpu_package_mean_power_w": 113.66 + }, + "files": { + "receipt.json": { + "bytes": 9486, + "sha256": "200c1704ff97848a8563c1505567c55edda36d9c93254178a0c2f3605a2fe8fc" + }, + "transcript.jsonl": { + "bytes": 20245, + "sha256": "510014cc2f05df4badebd4c2e66909cc5019d93be17139813ba5a738688ac8a6" + }, + "summary.json": { + "bytes": 690, + "sha256": "5f5745d935c3bb8c4a1102b0f90c09ac932ae3c64404928d908f58cfe1f363b0" + }, + "workspace_final.tar.gz": { + "bytes": 15762, + "sha256": "e41e9dbd0b1ca52655c0feea9fd0ef7a067ffe623df42babd69349a1423e493e" + }, + "cost.json": { + "bytes": 1235, + "sha256": "6e4debe63f72ff3bd80fbf7699501b9dcfa5ca02f78d1bf806640abc93711290" + }, + "gpu_telemetry.json": { + "bytes": 2398, + "sha256": "f2f7959c216c21c3c8beebe7d993b7f43196e5225e4046beca680aa8fcd74f42" + }, + "grade.json": { + "bytes": 991, + "sha256": "4e17cab2549fc8eefb34d1fbfc11c34b7bc33abd78a70a48955a403a19f02333" + } + } + }, + { + "run_name": "p3_market_gemma4-31b-q4_v9", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 192.8, + "iterations": 18, + "completion_tokens": 9097, + "prompt_tokens_cumulative": 362891, + "model_turns": 18, + "length_finishes": 0, + "grade_verdict": "STRUCTURAL_PASS", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9595, + "mean_power_w": 484.11, + "mean_sm_util_pct": 93.21, + "max_temp_c": 76.0, + "cpu_package_mean_power_w": 135.49 + }, + "files": { + "receipt.json": { + "bytes": 9488, + "sha256": "50c1b80a86bee2587238a5e4522b2bdd294b0cebc6b799eb7bec665ab6125bc4" + }, + "transcript.jsonl": { + "bytes": 23583, + "sha256": "142cb95a517e71670f346f34fe6707e636eb99e3143af2a209c066c4a7a0a5b7" + }, + "summary.json": { + "bytes": 1157, + "sha256": "27723ba773ccf22b52afccabb0dff8bef0bdfce65235c3769cfe9063d36ae96e" + }, + "workspace_final.tar.gz": { + "bytes": 277307, + "sha256": "1a3175aa8659e5b531d9837970637815a3ec2d2cf751e356f508b80b1fbd473f" + }, + "cost.json": { + "bytes": 1237, + "sha256": "7efde2a1b35b36ce16ef46d1a798978dbdb809f715cca20899c68b5127abc461" + }, + "gpu_telemetry.json": { + "bytes": 2459, + "sha256": "f9a7f1e818fca163ed70db2d15f312fa04176a401a0bf7716278f2db63830874" + }, + "grade.json": { + "bytes": 991, + "sha256": "ac1d5b61765d23524f6d2b292d9afb52d63e4086180ac1060c1c17b044ed23cf" + } + } + }, + { + "run_name": "p3_market_gemma4-31b-q4_v10", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 208.1, + "iterations": 18, + "completion_tokens": 10104, + "prompt_tokens_cumulative": 328243, + "model_turns": 18, + "length_finishes": 0, + "grade_verdict": "STRUCTURAL_PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9611, + "mean_power_w": 476.83, + "mean_sm_util_pct": 89.93, + "max_temp_c": 71.0, + "cpu_package_mean_power_w": 112.77 + }, + "files": { + "receipt.json": { + "bytes": 9487, + "sha256": "ed3dad80fc3fe2e18211db02e90f91ddff3377b469baa4ee7ae4c5b0ee43d27a" + }, + "transcript.jsonl": { + "bytes": 25957, + "sha256": "5ea786806e57a817dec661e80d95cc1b0925ce7fc8a111a80ca89f4e6900f51a" + }, + "summary.json": { + "bytes": 668, + "sha256": "ce6c0a1db966341b92d781605671d23066f37cfff97d9212b8f056a78f96f1c3" + }, + "workspace_final.tar.gz": { + "bytes": 278373, + "sha256": "e5fd79a3c7cd263cb293a71c8c8d556be7e4cdd05502a1752f2c21b59efa05f2" + }, + "cost.json": { + "bytes": 1239, + "sha256": "3f80fdf83c18201fb7672a4ed22db2285f6862aeac7ddee5150afbdfbac01590" + }, + "gpu_telemetry.json": { + "bytes": 2399, + "sha256": "6377f702b79609c0a286759af8be45e4d4031f06bd9f3c14dc3dc5c55891acf9" + }, + "grade.json": { + "bytes": 991, + "sha256": "bae011370d19b11261427cb80a48bd422914bc2aabfd2df7a51be42616be2355" + } + } + }, + { + "run_name": "p3_writing_gemma4-31b-q4_v1", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 93.0, + "iterations": 13, + "completion_tokens": 5089, + "prompt_tokens_cumulative": 72407, + "model_turns": 13, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.914, + "mean_power_w": 460.49, + "mean_sm_util_pct": 86.0, + "max_temp_c": 73.0, + "cpu_package_mean_power_w": 112.85 + }, + "files": { + "receipt.json": { + "bytes": 8962, + "sha256": "20d5e7d994199823fab97c326232441ca9bd5569ef11a972c402dc7d54589daa" + }, + "transcript.jsonl": { + "bytes": 15990, + "sha256": "404e7626e52a4928a4881eb7be832152ced40aaa01fdc3f1eaf8816a4071fa64" + }, + "summary.json": { + "bytes": 677, + "sha256": "15b672068f5fa802a679f4bea758c2a91c44cb3c664c2b3d3beb8b19b487e469" + }, + "workspace_final.tar.gz": { + "bytes": 15514, + "sha256": "f87a925d16003bb1d84da8c39b4491b168c6b8ba612ff6fd878ddf7e56cbd37c" + }, + "cost.json": { + "bytes": 1234, + "sha256": "3eae502e939914b0b6add91645d27f74ea3ff1db347bdddc9be820eff8b73841" + }, + "gpu_telemetry.json": { + "bytes": 2396, + "sha256": "822f151e387086123749feebbc4f25cfac9be56f13c8b5a9a417b723b8dc4eef" + }, + "grade.json": { + "bytes": 2879, + "sha256": "4d52cac7e70665c17a000630f5e703deb8c2dd7fa09b140e58cd916809855594" + } + } + }, + { + "run_name": "p3_writing_gemma4-31b-q4_v2", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 191.8, + "iterations": 11, + "completion_tokens": 4336, + "prompt_tokens_cumulative": 58205, + "model_turns": 11, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9645, + "mean_power_w": 482.8, + "mean_sm_util_pct": 87.37, + "max_temp_c": 87.0, + "cpu_package_mean_power_w": 138.28 + }, + "files": { + "receipt.json": { + "bytes": 8964, + "sha256": "344295433cc7e8dfc413953057888c378ae8a2a8e95d408bdd6e1186fe503a42" + }, + "transcript.jsonl": { + "bytes": 11867, + "sha256": "9d5ca60adf16166a00054ac7e38ecadd7c220127d625ee7b7accb47537e7d600" + }, + "summary.json": { + "bytes": 1072, + "sha256": "25db8e039e0d70de15743604f8da5ca48a646708e5636b23777ff9d4d594e992" + }, + "workspace_final.tar.gz": { + "bytes": 13886, + "sha256": "e5da165c568e1e684254c5d7996c5654318e377afbead4859b05b741f945b5e1" + }, + "cost.json": { + "bytes": 1236, + "sha256": "ccfd5b41afd661d08ef93058d0a09e94569bdb2c31f991637462b4265748964c" + }, + "gpu_telemetry.json": { + "bytes": 2405, + "sha256": "af34e1a95f3de8752e175ba70f33aeab87acf0cb17cd9ed3480011787efe8f79" + }, + "grade.json": { + "bytes": 2879, + "sha256": "d5b96fa97f5173b2c03372286d3f1337cdad83a2d8433953285d8c1b96d080de" + } + } + }, + { + "run_name": "p3_writing_gemma4-31b-q4_v3", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 73.7, + "iterations": 11, + "completion_tokens": 4250, + "prompt_tokens_cumulative": 58290, + "model_turns": 11, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9498, + "mean_power_w": 441.34, + "mean_sm_util_pct": 79.0, + "max_temp_c": 77.0, + "cpu_package_mean_power_w": 113.07 + }, + "files": { + "receipt.json": { + "bytes": 8962, + "sha256": "0d34b93613273625c8093145f3bcdc4de4cf88758f864c1c758e10e1faf504e7" + }, + "transcript.jsonl": { + "bytes": 11561, + "sha256": "098a161c5fb99762e3df7f4cb8c490549ddc2435f33ffb7887eeed19aece747f" + }, + "summary.json": { + "bytes": 522, + "sha256": "5b1be637be95a148bb4ad11662deeddb74543a11a845653f7de3bfff000a72ac" + }, + "workspace_final.tar.gz": { + "bytes": 13991, + "sha256": "b612618510b43d4fb1793ecef2149ebd0da6c21685d10b742045429243d9bad5" + }, + "cost.json": { + "bytes": 1234, + "sha256": "dbd969a5548e2a9b59d0e081c28c181c18cab45e22d190de562d8ecbfdfa70a7" + }, + "gpu_telemetry.json": { + "bytes": 2395, + "sha256": "ff064a0f5a1d8072807205f220c676a1e133eba8b90dc74da17d77b1706109b6" + }, + "grade.json": { + "bytes": 2876, + "sha256": "5554d7976a5268e573cd20176244ba7db50716c7dc7dad7df89386f9fc269564" + } + } + }, + { + "run_name": "p3_writing_gemma4-31b-q4_v4", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 89.3, + "iterations": 11, + "completion_tokens": 5045, + "prompt_tokens_cumulative": 56093, + "model_turns": 11, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9518, + "mean_power_w": 453.98, + "mean_sm_util_pct": 81.11, + "max_temp_c": 69.0, + "cpu_package_mean_power_w": 114.1 + }, + "files": { + "receipt.json": { + "bytes": 9522, + "sha256": "bccb937879aa53114dd00efdfc09e00092ad5cb042e4557f255315b3f1b00330" + }, + "transcript.jsonl": { + "bytes": 13876, + "sha256": "9c326dd6b49acd70386f93ccc8cc6876eaeb4ad05d00f5d320920b09d796ac5c" + }, + "summary.json": { + "bytes": 667, + "sha256": "3b45874c154620f448b1aba385985edf7121cbcabb557c823011825c14f2d25b" + }, + "workspace_final.tar.gz": { + "bytes": 15008, + "sha256": "94cb06e76ab910fed343e34a8e618e51f340af2acb3d4203a0ffaa378c1d59c2" + }, + "cost.json": { + "bytes": 1233, + "sha256": "89a69d3c3488c2c6b9d12aaa7ae8bdba58944e115e14d0d936807225b9de3887" + }, + "gpu_telemetry.json": { + "bytes": 2397, + "sha256": "677f2c792a943b9df0138c56531fab7b8d2f6c17914b728c91108bf75101736b" + }, + "grade.json": { + "bytes": 2879, + "sha256": "b85303ff36d7da323a83154c84a296f2782ff66d51e5b0db1ef56a1379dc369b" + } + } + }, + { + "run_name": "p3_writing_gemma4-31b-q4_v5", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 78.5, + "iterations": 11, + "completion_tokens": 4250, + "prompt_tokens_cumulative": 58290, + "model_turns": 11, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9554, + "mean_power_w": 430.88, + "mean_sm_util_pct": 66.25, + "max_temp_c": 73.0, + "cpu_package_mean_power_w": 136.26 + }, + "files": { + "receipt.json": { + "bytes": 9524, + "sha256": "dfa462f7a580a7b5d3e363a4b61d6aeafe642d257ac5a7f765c45a6d0a94c27d" + }, + "transcript.jsonl": { + "bytes": 11563, + "sha256": "21c854131b9ba11c8c771b160160a02f454aebd72ea0708eae611aee4f27f596" + }, + "summary.json": { + "bytes": 522, + "sha256": "47533bb5a720974e0a96af9d95acebd707e7558724943252e0995adbb5d5cc75" + }, + "workspace_final.tar.gz": { + "bytes": 13981, + "sha256": "0311507fb6a09b8ca422a9ec6b8e2fdcd12fad6980f2e3c4b012740261870ca9" + }, + "cost.json": { + "bytes": 1234, + "sha256": "c0723c879219bd2c35c61c8352aff9b3dba4177790a794eb56b786b12c2dda4b" + }, + "gpu_telemetry.json": { + "bytes": 2449, + "sha256": "2e8b79f4a989b4709250a9343379a7f697c1c58fa55f5a28c5452231b4d28022" + }, + "grade.json": { + "bytes": 2876, + "sha256": "5554d7976a5268e573cd20176244ba7db50716c7dc7dad7df89386f9fc269564" + } + } + }, + { + "run_name": "p3_writing_gemma4-31b-q4_v6", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 87.0, + "iterations": 11, + "completion_tokens": 5045, + "prompt_tokens_cumulative": 56093, + "model_turns": 11, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9195, + "mean_power_w": 495.19, + "mean_sm_util_pct": 97.47, + "max_temp_c": 70.0, + "cpu_package_mean_power_w": 114.44 + }, + "files": { + "receipt.json": { + "bytes": 9522, + "sha256": "943941c7e2350abee61dba6db4d4c2d2f40af52b2457b47caf1c247cfbc46a47" + }, + "transcript.jsonl": { + "bytes": 13876, + "sha256": "4047666bdba8ec34d719d3882c99d9ac8fd4384b782064bb25a2d56400daeee1" + }, + "summary.json": { + "bytes": 667, + "sha256": "8a7dd8eeda3911e1a6558aaef1ae7727956c56b7527828b8af8c0d2eaa895183" + }, + "workspace_final.tar.gz": { + "bytes": 15004, + "sha256": "1bf1f974212a1bc0fbda308fe718d94b7212fc78080fc8c9832a5494431fb2ce" + }, + "cost.json": { + "bytes": 1234, + "sha256": "abfc36f242249fa531bca4d6b9000a369c5efed239787224d1f30d5e0a85418f" + }, + "gpu_telemetry.json": { + "bytes": 2394, + "sha256": "466929a082eb20550d9aedbbe74221c1f53b8b003adb8e49f9cb4e1142da4413" + }, + "grade.json": { + "bytes": 2879, + "sha256": "b85303ff36d7da323a83154c84a296f2782ff66d51e5b0db1ef56a1379dc369b" + } + } + }, + { + "run_name": "p3_writing_gemma4-31b-q4_v7", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 75.0, + "iterations": 11, + "completion_tokens": 4318, + "prompt_tokens_cumulative": 58254, + "model_turns": 11, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9333, + "mean_power_w": 443.43, + "mean_sm_util_pct": 85.27, + "max_temp_c": 74.0, + "cpu_package_mean_power_w": 136.26 + }, + "files": { + "receipt.json": { + "bytes": 9524, + "sha256": "6520469e45e57400281feed99b9d8e1e50ac766066bb300efa1033afc5f03f63" + }, + "transcript.jsonl": { + "bytes": 11573, + "sha256": "d1fb4d39b86e292ddb41de4230d185b7528249f3bb0389009b45762b1216a817" + }, + "summary.json": { + "bytes": 588, + "sha256": "30b0e523032ef0be9dd62f2f4e8151e0b2cab13af98daa1610dabcc60d9de07d" + }, + "workspace_final.tar.gz": { + "bytes": 14039, + "sha256": "430f575db65239e9c8ae2ccfd02512306dc04430340c7c0d0cf47c0a54b454b4" + }, + "cost.json": { + "bytes": 1234, + "sha256": "88de2c4ef3b174a084373a8942a3e92f0b436032fdc5c1b8562198902aeedc3a" + }, + "gpu_telemetry.json": { + "bytes": 2441, + "sha256": "c440aa0429ea49b1e32054eb8f3614a105c5a55da60f817fac579cda02a6906c" + }, + "grade.json": { + "bytes": 2876, + "sha256": "445eea72603fb268fb807b30c88e47ca7e8f6cae3b63812da43e4d4687dab137" + } + } + }, + { + "run_name": "p3_writing_gemma4-31b-q4_v8", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 113.5, + "iterations": 13, + "completion_tokens": 5311, + "prompt_tokens_cumulative": 74398, + "model_turns": 13, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9251, + "mean_power_w": 483.94, + "mean_sm_util_pct": 90.91, + "max_temp_c": 72.0, + "cpu_package_mean_power_w": 114.32 + }, + "files": { + "receipt.json": { + "bytes": 9522, + "sha256": "b34d7da793ad43c447bd43aa43f8d0c18e35704214959872d7f541268d3f939e" + }, + "transcript.jsonl": { + "bytes": 15899, + "sha256": "88d5daf08e017a5d78395b90cc4f5c8f32f8d48eee70c6ef84b00d39b3399d3e" + }, + "summary.json": { + "bytes": 725, + "sha256": "8ede490aa9efc41f8afc99947e8aed1c5658828cc82ba48e1051e024b656a276" + }, + "workspace_final.tar.gz": { + "bytes": 15418, + "sha256": "405f334374a9ef9b0b4db0e5fbd8ff06856cd892cf8a756cdd25ee6dabe361ce" + }, + "cost.json": { + "bytes": 1236, + "sha256": "dc96995f812b5f2b36a5306d21026fb259757f47f18ab456150caf6689570b53" + }, + "gpu_telemetry.json": { + "bytes": 2401, + "sha256": "2efd98e12f90693393241e8af9a19624d34a17ea2e137b90b8fbdd6e1b93d333" + }, + "grade.json": { + "bytes": 2879, + "sha256": "a7f0d960afecef69f7e22bb44f38920d3b47486f61e8bd0d9c9d52c1fff00220" + } + } + }, + { + "run_name": "p3_writing_gemma4-31b-q4_v9", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 89.3, + "iterations": 13, + "completion_tokens": 4932, + "prompt_tokens_cumulative": 73582, + "model_turns": 13, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.8959, + "mean_power_w": 474.04, + "mean_sm_util_pct": 91.82, + "max_temp_c": 73.0, + "cpu_package_mean_power_w": 136.5 + }, + "files": { + "receipt.json": { + "bytes": 9524, + "sha256": "a0273b64ada8982571d95150e5109d1914bf3533b88a1a2356963f246d9aa9b0" + }, + "transcript.jsonl": { + "bytes": 15845, + "sha256": "bcf7d53b6471b402dd754a173620431edd7436182726573353bbb28d8c1756a6" + }, + "summary.json": { + "bytes": 640, + "sha256": "2557e1b999cefbba82c37ea433af7ef26ccab4b527c3b9942bb2b2ff4a6f791f" + }, + "workspace_final.tar.gz": { + "bytes": 15376, + "sha256": "b0c844ea8b8dba9049879f02fe7d5612f4216dccc4e65ee68d43ed968dc18ca0" + }, + "cost.json": { + "bytes": 1234, + "sha256": "4af75b24811fa1c65aa2401629348883056becb6e4d7c6dcdfd09b0a400d4c06" + }, + "gpu_telemetry.json": { + "bytes": 2437, + "sha256": "e155ad1fb36ee3a7eec55dfb19a3a51cc1da51d8b3628b99adf430f0c00b712f" + }, + "grade.json": { + "bytes": 2879, + "sha256": "fe1a7369ee2c3314701f72704ca898889ee1d4952d825227c6a9d703d907aea3" + } + } + }, + { + "run_name": "p3_writing_gemma4-31b-q4_v10", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 74.1, + "iterations": 11, + "completion_tokens": 4250, + "prompt_tokens_cumulative": 58290, + "model_turns": 11, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9447, + "mean_power_w": 495.55, + "mean_sm_util_pct": 97.73, + "max_temp_c": 69.0, + "cpu_package_mean_power_w": 113.87 + }, + "files": { + "receipt.json": { + "bytes": 9523, + "sha256": "5a5bed0db0d5c1a80b3c48ea0f0a26cac37f0e065694bd5d52a20c515ff73116" + }, + "transcript.jsonl": { + "bytes": 11564, + "sha256": "f07ac6b79ed0589385cfcf1a1eb4f5f98c222f513e18de655bdd5f8d5850187a" + }, + "summary.json": { + "bytes": 522, + "sha256": "6bc68ec1324c6a3926300502f142af792e52bf024cedef4b5cd2e2608ba1cb79" + }, + "workspace_final.tar.gz": { + "bytes": 13981, + "sha256": "25a40e01703deee6d40e99b8a1a8cebf2276a46241ca5dc9b632825d30aabf81" + }, + "cost.json": { + "bytes": 1235, + "sha256": "4dbb2f3e7c007f02e0399447ca3c6d016b57a7e2f6016c37ceeb6805ece7ebbe" + }, + "gpu_telemetry.json": { + "bytes": 2398, + "sha256": "bb142a78a6baad2db13ef41cf6431bc3c687c808455e78e185b606c506fb3169" + }, + "grade.json": { + "bytes": 2876, + "sha256": "5554d7976a5268e573cd20176244ba7db50716c7dc7dad7df89386f9fc269564" + } + } + }, + { + "run_name": "p3_pm_gemma4-31b-q4_v1", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 183.9, + "iterations": 7, + "completion_tokens": 3786, + "prompt_tokens_cumulative": 28755, + "model_turns": 7, + "length_finishes": 0, + "grade_verdict": "FAIL", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9788, + "mean_power_w": 484.17, + "mean_sm_util_pct": 90.46, + "max_temp_c": 86.0, + "cpu_package_mean_power_w": 138.49 + }, + "files": { + "receipt.json": { + "bytes": 8953, + "sha256": "bf5b124d1fcf516f129d94fe8f22407461230499307ab01a24e0013f6de68136" + }, + "transcript.jsonl": { + "bytes": 7558, + "sha256": "d3291b6c0e20bef9bd143b9bc1d6e1c7f41e81c90ad1d867aad5a0cdc77a5e64" + }, + "summary.json": { + "bytes": 694, + "sha256": "83681b6cad2097056abf0f0bcd8bd63642670e1e9959c6c5cf1de1ff099c7369" + }, + "workspace_final.tar.gz": { + "bytes": 12841, + "sha256": "eb622a02fddb70c3aee43f652123e59640282815e55b230d70874175a61fcf08" + }, + "cost.json": { + "bytes": 1230, + "sha256": "4e6c26bf109501788f45951c8b086b0af788ec68ffbbc420b63a17b0167077db" + }, + "gpu_telemetry.json": { + "bytes": 2402, + "sha256": "2d631f9e227a929e3d6a3d544b554bfbba9593894dc542e961198d60b4f60275" + }, + "grade.json": { + "bytes": 2404, + "sha256": "eade149c12512150ebb28d5a446c6dc1037e0ded24a3ed5e83d65d87aba47bc4" + } + } + }, + { + "run_name": "p3_pm_gemma4-31b-q4_v2", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 72.2, + "iterations": 7, + "completion_tokens": 4145, + "prompt_tokens_cumulative": 29617, + "model_turns": 7, + "length_finishes": 0, + "grade_verdict": "FAIL", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9695, + "mean_power_w": 472.88, + "mean_sm_util_pct": 90.27, + "max_temp_c": 76.0, + "cpu_package_mean_power_w": 132.69 + }, + "files": { + "receipt.json": { + "bytes": 8951, + "sha256": "7975d3f2ebd6af42d6a203790f662b61c46f3d41a6c7497971305b7e830c8481" + }, + "transcript.jsonl": { + "bytes": 8573, + "sha256": "aaaa90cd1dd4cc1ccdb2241a17c9d721400672f226be72fdce168a7a7d2339de" + }, + "summary.json": { + "bytes": 754, + "sha256": "598cb8b9ee1c2b73a0f4e8728259609efe1b0f460331ab883920a660e7d1b7ff" + }, + "workspace_final.tar.gz": { + "bytes": 13139, + "sha256": "6190888ab5900349d69d7d443477eb162cb5174acc340c36c5489a95b9a591c8" + }, + "cost.json": { + "bytes": 1226, + "sha256": "95639c64f720d0ba360c8bf67a7e46bbc821a72c11105f96b653a0632742bb21" + }, + "gpu_telemetry.json": { + "bytes": 2395, + "sha256": "2ae7399a0dfd6faacc59f126b94d74d271d71768afc9feaafbc2b7fcc7deb9ea" + }, + "grade.json": { + "bytes": 2425, + "sha256": "5e98c1230f111ddb4d1a15ae7063c80590fbee19b72661c4b56e9c94f0b12352" + } + } + }, + { + "run_name": "p3_pm_gemma4-31b-q4_v3", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 172.3, + "iterations": 6, + "completion_tokens": 3386, + "prompt_tokens_cumulative": 25278, + "model_turns": 6, + "length_finishes": 0, + "grade_verdict": "FAIL", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9867, + "mean_power_w": 478.32, + "mean_sm_util_pct": 93.23, + "max_temp_c": 86.0, + "cpu_package_mean_power_w": 138.81 + }, + "files": { + "receipt.json": { + "bytes": 8953, + "sha256": "ced44d9abe5553e3d7b62f0bf8c0fb2137ec214ed8e567bb0e5b8a3854f891e5" + }, + "transcript.jsonl": { + "bytes": 7311, + "sha256": "b130ac807c1b2622b54e3e4eb30dd1d0fef1c25db3328b185ed18f360a17fe89" + }, + "summary.json": { + "bytes": 708, + "sha256": "20c8f932f6c38489911471d600d9236cc122deca72f731831bc09cbe2fb2b4e8" + }, + "workspace_final.tar.gz": { + "bytes": 12920, + "sha256": "fac5112297764b16bb742a38fcecc8617524ef13ebe869170e049842e63a4645" + }, + "cost.json": { + "bytes": 1231, + "sha256": "ed09867309f941998fbb60813dcee241c4b31cd8dc8960ccb41d02a20787ca1b" + }, + "gpu_telemetry.json": { + "bytes": 2403, + "sha256": "0fc9af6548edd42ec81f3df1c0e08f1603b7aac954fe19034c3c17c7fa7a8a66" + }, + "grade.json": { + "bytes": 2408, + "sha256": "be7951b5c0b3770b9ee61e8f3bd5042dfc0fa486d2be1086b97314e4a8976ca3" + } + } + }, + { + "run_name": "p3_pm_gemma4-31b-q4_v4", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 61.5, + "iterations": 7, + "completion_tokens": 3600, + "prompt_tokens_cumulative": 26878, + "model_turns": 7, + "length_finishes": 0, + "grade_verdict": "FAIL", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.8943, + "mean_power_w": 498.59, + "mean_sm_util_pct": 94.08, + "max_temp_c": 69.0, + "cpu_package_mean_power_w": 114.24 + }, + "files": { + "receipt.json": { + "bytes": 9511, + "sha256": "83571a20f079af82cf4d06975d7abf75c87ea72b5349b4a8d67420e6197bcc60" + }, + "transcript.jsonl": { + "bytes": 8207, + "sha256": "f2230f4142a999c28ac225c34c66269f651cf47d9534667a791310622964c1c7" + }, + "summary.json": { + "bytes": 654, + "sha256": "806f285fd95a8355d32e27d5991bdb8a1a4ab5b98faed90941fb6dc6e4151291" + }, + "workspace_final.tar.gz": { + "bytes": 13064, + "sha256": "2f376b7734e2927dc611ac896789acbb0fb16e066221f84604890fa138548b11" + }, + "cost.json": { + "bytes": 1228, + "sha256": "8d6932f134646a9544d0ce0f8db5f2e885fced211a2e9059da91c99c6c84fb9f" + }, + "gpu_telemetry.json": { + "bytes": 2392, + "sha256": "4f2ff51421673db67651abd6cf2ec3dd21faa31445147ca60c0e7e553a89712f" + }, + "grade.json": { + "bytes": 2429, + "sha256": "013a8862905aa0eac7318b00edaa3a3e2f2ca542c4a329ef794b38486fd7a73b" + } + } + }, + { + "run_name": "p3_pm_gemma4-31b-q4_v5", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 61.3, + "iterations": 7, + "completion_tokens": 3600, + "prompt_tokens_cumulative": 26878, + "model_turns": 7, + "length_finishes": 0, + "grade_verdict": "FAIL", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.8972, + "mean_power_w": 493.89, + "mean_sm_util_pct": 96.92, + "max_temp_c": 73.0, + "cpu_package_mean_power_w": 136.31 + }, + "files": { + "receipt.json": { + "bytes": 9513, + "sha256": "b72010d4df13d86ccaf7a82294d6bd484ad4d0afe55a447bf28a5c7fb639d9e7" + }, + "transcript.jsonl": { + "bytes": 8205, + "sha256": "21e25f726c0f8d8b46be12ece4f1d7ebd938d0d35b129798f6df79ef5029221f" + }, + "summary.json": { + "bytes": 654, + "sha256": "ceb7ad738c7d22351a4462fb49af6cc7efc62551971257343da5ef0f4b2597d3" + }, + "workspace_final.tar.gz": { + "bytes": 13067, + "sha256": "75cbdce99bebf45b3c96082686765bb2621fd862ea9dea6d76fc4505c008b127" + }, + "cost.json": { + "bytes": 1228, + "sha256": "976b451ac25019d2c0439ac862ce6a99bf401acde2f48f3b1f43895e77c12759" + }, + "gpu_telemetry.json": { + "bytes": 2392, + "sha256": "e3e02081c3a15ca2bf2c53e7490b4351d02cc6f0866fcd7be8ff7bc852b1ef86" + }, + "grade.json": { + "bytes": 2429, + "sha256": "013a8862905aa0eac7318b00edaa3a3e2f2ca542c4a329ef794b38486fd7a73b" + } + } + }, + { + "run_name": "p3_pm_gemma4-31b-q4_v6", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 69.4, + "iterations": 7, + "completion_tokens": 4145, + "prompt_tokens_cumulative": 29617, + "model_turns": 7, + "length_finishes": 0, + "grade_verdict": "FAIL", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.8646, + "mean_power_w": 494.58, + "mean_sm_util_pct": 97.69, + "max_temp_c": 70.0, + "cpu_package_mean_power_w": 113.81 + }, + "files": { + "receipt.json": { + "bytes": 9511, + "sha256": "9743759244a93aa0d4372c338ebc888529ff20dcd22308b19e7dffea2c05d58c" + }, + "transcript.jsonl": { + "bytes": 8576, + "sha256": "04d4faf71414c19c5e3a3b3ad6d5d6e8a084925bd661cf140d530d520e43a262" + }, + "summary.json": { + "bytes": 754, + "sha256": "19f9854e4c70655afd491c27edf3171cf849e277f41ecf5e84348ce866674e16" + }, + "workspace_final.tar.gz": { + "bytes": 13142, + "sha256": "42ebcf94466a424be2fd5176c6d1bb110004371a0c5462a3c059de3aaa900f54" + }, + "cost.json": { + "bytes": 1228, + "sha256": "6c7159adb0b43d8db37635fb238e688a21c7d9b255f7412a408ae9361cc459a2" + }, + "gpu_telemetry.json": { + "bytes": 2393, + "sha256": "1c4319a215cc23af555a366367f2996e2681ca8830f1d6edeb41a5f19702a5af" + }, + "grade.json": { + "bytes": 2425, + "sha256": "5e98c1230f111ddb4d1a15ae7063c80590fbee19b72661c4b56e9c94f0b12352" + } + } + }, + { + "run_name": "p3_pm_gemma4-31b-q4_v7", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 72.5, + "iterations": 7, + "completion_tokens": 4145, + "prompt_tokens_cumulative": 29617, + "model_turns": 7, + "length_finishes": 0, + "grade_verdict": "FAIL", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9655, + "mean_power_w": 469.05, + "mean_sm_util_pct": 88.0, + "max_temp_c": 72.0, + "cpu_package_mean_power_w": 136.86 + }, + "files": { + "receipt.json": { + "bytes": 9513, + "sha256": "3d836d36746b2411ac293c9b190598d6d6fdc38d5e01704a7e5575ad438c73cc" + }, + "transcript.jsonl": { + "bytes": 8575, + "sha256": "8ba1a5fbce310b9a5b5ba6e428da54f7335492228cc12f783ebccde0a54476f0" + }, + "summary.json": { + "bytes": 754, + "sha256": "3f1fa23a13a4c0771557c3958c50c21089d9b3200cc8acbc04265e2c57c1379e" + }, + "workspace_final.tar.gz": { + "bytes": 13141, + "sha256": "191675310a0ca94b141e1b880425ba5476b1c71610050d238ce87fbe2b1ba7e4" + }, + "cost.json": { + "bytes": 1228, + "sha256": "e8a285de519139fc9b1085c9f3f7ed4fd7f3a125f161ebaeabc06e66668ec946" + }, + "gpu_telemetry.json": { + "bytes": 2434, + "sha256": "c3fd8df12534ac8f44c2e654cac833f9000b080e96be7a560d94587d5a9595d4" + }, + "grade.json": { + "bytes": 2425, + "sha256": "5e98c1230f111ddb4d1a15ae7063c80590fbee19b72661c4b56e9c94f0b12352" + } + } + }, + { + "run_name": "p3_pm_gemma4-31b-q4_v8", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 70.3, + "iterations": 7, + "completion_tokens": 4103, + "prompt_tokens_cumulative": 28950, + "model_turns": 7, + "length_finishes": 0, + "grade_verdict": "FAIL", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9246, + "mean_power_w": 483.78, + "mean_sm_util_pct": 97.71, + "max_temp_c": 69.0, + "cpu_package_mean_power_w": 113.99 + }, + "files": { + "receipt.json": { + "bytes": 9511, + "sha256": "001ec6c630885737ebb1b5aaa568c15b6b28a7f01d4a0ade59eb428af7d73aea" + }, + "transcript.jsonl": { + "bytes": 7582, + "sha256": "91d81b7721155f01336c3834f3e331689b7235fa116e6c1a282f22739333a92f" + }, + "summary.json": { + "bytes": 607, + "sha256": "e18c82aef482c5710af57dc798bed55b383177aaa575bfec463361bbed3a238b" + }, + "workspace_final.tar.gz": { + "bytes": 12825, + "sha256": "261171d6c3ef415e0486913dcf132b32162e4f84de90128b64cf2b344e41a4a5" + }, + "cost.json": { + "bytes": 1228, + "sha256": "a23e65162d7172b7aa02536553686da90fcb4268fd8a3f83479d4f86bf079f56" + }, + "gpu_telemetry.json": { + "bytes": 2392, + "sha256": "3f91fab1614681631c9849f564827150549ad8e0f9e487fde9b92eebb0bd4fa9" + }, + "grade.json": { + "bytes": 2420, + "sha256": "5e5d3ca47674ec81eb8de446669c70129aff454157ba42e6bc018d6a13e64818" + } + } + }, + { + "run_name": "p3_pm_gemma4-31b-q4_v9", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 57.8, + "iterations": 8, + "completion_tokens": 3319, + "prompt_tokens_cumulative": 33101, + "model_turns": 8, + "length_finishes": 0, + "grade_verdict": "FAIL", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.8651, + "mean_power_w": 487.69, + "mean_sm_util_pct": 88.64, + "max_temp_c": 73.0, + "cpu_package_mean_power_w": 137.0 + }, + "files": { + "receipt.json": { + "bytes": 9513, + "sha256": "9424c5f3f93cdf47b02c2a16037c3698fc0a20859d6b14c6cd96baa51975aa19" + }, + "transcript.jsonl": { + "bytes": 7934, + "sha256": "fff346946611cb6b8cc9a443604d237890c876fdc8bc375e0625c1a7757a4556" + }, + "summary.json": { + "bytes": 590, + "sha256": "dccb1b8b15a5f99549ed49ebb94d67f0d975b86f3e4bba73772070eb9438a242" + }, + "workspace_final.tar.gz": { + "bytes": 12821, + "sha256": "326c41dfdf2140156ad57262cc39fa5f85e034c0523db2ebd47834a129226762" + }, + "cost.json": { + "bytes": 1226, + "sha256": "97dce45e54ececd2808e643c5133ba24ef25e3ebef834e9f229485b18482f9aa" + }, + "gpu_telemetry.json": { + "bytes": 2396, + "sha256": "36b4d2a59deee776effa46df4d09e67a7abff556285da7acc89f641417550dd5" + }, + "grade.json": { + "bytes": 2426, + "sha256": "8defbbde59eb9e22c6aa51b5bd69e93bcf5a548d56021d8f45b5fada421c1021" + } + } + }, + { + "run_name": "p3_pm_gemma4-31b-q4_v10", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 57.7, + "iterations": 8, + "completion_tokens": 3319, + "prompt_tokens_cumulative": 33101, + "model_turns": 8, + "length_finishes": 0, + "grade_verdict": "FAIL", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9532, + "mean_power_w": 464.83, + "mean_sm_util_pct": 85.5, + "max_temp_c": 69.0, + "cpu_package_mean_power_w": 114.62 + }, + "files": { + "receipt.json": { + "bytes": 9512, + "sha256": "16e4323075d8a129714613eaa7c7c9c378be99d30c655809f8d870cb167343f9" + }, + "transcript.jsonl": { + "bytes": 7934, + "sha256": "905a3438a575521f70bde0c4bab5eabfa0cad33233b394c9f3b0a7f93e10cd9d" + }, + "summary.json": { + "bytes": 590, + "sha256": "acd2250aeaf7e22de9fb082324d34beeae6bcb5f766863f3e1eb5edfdf7b547d" + }, + "workspace_final.tar.gz": { + "bytes": 12827, + "sha256": "8b11367b2047223d189f3f7c731f6735558ae640d3d297d47d8e5a6f7bf31db5" + }, + "cost.json": { + "bytes": 1227, + "sha256": "76cda3835163c2cd121df8af4c0ce730490ffcd3cbea886047c4dd54d6d5425b" + }, + "gpu_telemetry.json": { + "bytes": 2391, + "sha256": "f3c6333f37d107099be76ca68d5aafbf54e42eb905b5c601a7e03bd875d58efa" + }, + "grade.json": { + "bytes": 2426, + "sha256": "8defbbde59eb9e22c6aa51b5bd69e93bcf5a548d56021d8f45b5fada421c1021" + } + } + } + ], + "pretelemetry_supplements": [ + { + "run_name": "p1_bugfix_gemma4-31b-q4-telemetry-supplement_v1", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 779.5, + "iterations": 68, + "completion_tokens": 22980, + "prompt_tokens_cumulative": 1360791, + "model_turns": 68, + "length_finishes": 0, + "grade_verdict": null, + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9942, + "mean_power_w": 297.76, + "mean_sm_util_pct": 51.88, + "max_temp_c": 74.0, + "cpu_package_mean_power_w": 128.3 + }, + "files": { + "receipt.json": { + "bytes": 9533, + "sha256": "160325870cc3b91505bfb372a0cf7f9950a3b51089e197b2b3696c5a6ede0da8" + }, + "transcript.jsonl": { + "bytes": 85245, + "sha256": "3de94b54b395a925cf7702c72e72331f86d3c412c662ab7ca925f436d8cca64d" + }, + "summary.json": { + "bytes": 2148, + "sha256": "ee4ec82f91cbade39cc9229f272ec83dee9389dba31fabde2fcddf33edf3fc23" + }, + "workspace_final.tar.gz": { + "bytes": 28950277, + "sha256": "ecb5307e3a8b2441b72e576298a3b84cfc304e1bf6e476ddbe1714fb48261f07" + }, + "cost.json": { + "bytes": 1265, + "sha256": "0b2b4c671701cc39ec343389a954b83904e95a3899ea18106f3798a8c5b4d9a0" + }, + "gpu_telemetry.json": { + "bytes": 2447, + "sha256": "a6ded0b0506c8680075b73c32f130c7353a32c4f924ee33f36b6f9d435523d09" + } + }, + "supplements_run": "p1_bugfix_gemma4-31b-q4_v1" + }, + { + "run_name": "p1_bugfix_gemma4-31b-q4-telemetry-supplement_v2", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 1031.5, + "iterations": 78, + "completion_tokens": 19891, + "prompt_tokens_cumulative": 1730028, + "model_turns": 78, + "length_finishes": 0, + "grade_verdict": null, + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9937, + "mean_power_w": 194.22, + "mean_sm_util_pct": 34.09, + "max_temp_c": 74.0, + "cpu_package_mean_power_w": 124.95 + }, + "files": { + "receipt.json": { + "bytes": 9533, + "sha256": "fe67b13b061d11ba2b790bbcee3e3a3b7658a67a1ab1bf7167b9a695c9339ff0" + }, + "transcript.jsonl": { + "bytes": 77337, + "sha256": "809d769786b89a1edc29607dee67707f0308116ea12594690dd327c4823e4883" + }, + "summary.json": { + "bytes": 1570, + "sha256": "a55b9bd46dcefd0dd85a0b2c5fb12f27b48cbfbf37979707c8f20cc3fa86ab31" + }, + "workspace_final.tar.gz": { + "bytes": 28961325, + "sha256": "ab7b933a6113f17198c275ffc0300c81cc8795f2f7264d2665a2f17a90b0fb4d" + }, + "cost.json": { + "bytes": 1265, + "sha256": "4070249f078c72dfd93083d59d1511adb16ac2647e2dd50557594b071699b6e6" + }, + "gpu_telemetry.json": { + "bytes": 2492, + "sha256": "d0b85511d693885dc0e3f38220c481780d15054d26b7ea9649e465e465c3fd3b" + } + }, + "supplements_run": "p1_bugfix_gemma4-31b-q4_v2" + } + ], + "invalid_attempt_classifications": { + "path": "/home/michael/bench-gemma4-31b-q4/tooling/deployments/gemma4-31b-q4-tower2/invalid-attempt-classifications.json", + "bytes": 6601, + "sha256": "09b45db2328ed1520f5491e37897a4cc031d66a9a1187b938977dfc872840176" + }, + "preserved_invalid_attempts": [ + { + "attempt": "p1_bugfix_gemma4-31b-q4_v3_retry-20260802T003822Z-dea60926", + "classification": { + "source_run": "p1_bugfix_gemma4-31b-q4_v3", + "classification": "infrastructure-invalid", + "reason_code": "operator-interrupted-companion-after-harness-dispatch-defect", + "incident_document": "canonical-harness-incident-20260802.json", + "classified_before_grade": true, + "affirmative_evidence": [ + "the companion lane raised the recorded KeyError in the shared defective harness", + "the supervisor was stopped immediately to prevent further evaluation under that harness" + ], + "expected_files": { + "receipt.json": "0a571fa5aaefacb15833eb56d63969d45b69f7e66b7622ec309ddf52ff1be4b5", + "transcript.jsonl": "7c9fa7c58e4ea36ab46938b690785b4a699b98d38c7f934d1aed083d87fd3163" + }, + "replacement": { + "required": true, + "canonical_run": "p1_bugfix_gemma4-31b-q4_v3", + "status": "completed" + } + }, + "files": { + "receipt.json": { + "bytes": 8174, + "sha256": "0a571fa5aaefacb15833eb56d63969d45b69f7e66b7622ec309ddf52ff1be4b5" + }, + "transcript.jsonl": { + "bytes": 1866, + "sha256": "7c9fa7c58e4ea36ab46938b690785b4a699b98d38c7f934d1aed083d87fd3163" + } + } + }, + { + "attempt": "p1_bugfix_gemma4-31b-q4_v3_retry-20260802T004533Z-64195329", + "classification": { + "source_run": "p1_bugfix_gemma4-31b-q4_v3", + "classification": "infrastructure-invalid", + "reason_code": "operator-interrupted-before-required-telemetry-gate", + "incident_document": "canonical-telemetry-gate-incident-20260802.json", + "classified_before_grade": true, + "affirmative_evidence": [ + "the retry began before per-run sidecar attribution was proven", + "the supervisor stopped before the attempt completed" + ], + "expected_files": { + "receipt.json": "c70a93f7255d8d4ef84311ed17187d84050c5e4faf4dbfd2768550e6773e760c", + "transcript.jsonl": "a07274f29bc3b36fe8c3259af0a8a4d53d1b7f8f6256e44dfc43c90b2b5ba7b2" + }, + "replacement": { + "required": true, + "canonical_run": "p1_bugfix_gemma4-31b-q4_v3", + "status": "completed" + } + }, + "files": { + "receipt.json": { + "bytes": 8170, + "sha256": "c70a93f7255d8d4ef84311ed17187d84050c5e4faf4dbfd2768550e6773e760c" + }, + "transcript.jsonl": { + "bytes": 1852, + "sha256": "a07274f29bc3b36fe8c3259af0a8a4d53d1b7f8f6256e44dfc43c90b2b5ba7b2" + } + } + }, + { + "attempt": "p1_refactor_gemma4-31b-q4_v3-server-timeout-20260802T020859Z", + "classification": { + "source_run": "p1_refactor_gemma4-31b-q4_v3", + "classification": "infrastructure-invalid", + "reason_code": "server-transport-timeout-below-native-envelope", + "incident_document": "canonical-refactor-timeout-incident-20260802.json", + "classified_before_grade": true, + "affirmative_evidence": [ + "server task 80497 was cancelled exactly 3600 seconds after launch", + "the slot had decoded 133606 tokens, reported truncated=false, and remained below its 262144-token context boundary" + ], + "expected_files": { + "cost.json": "c5d1062f477758436ba0c3baf32df97fbe4e30df09c831bde470ccf938fce7f7", + "gpu_telemetry.json": "4f717a0e9d7298401c41e1799f772186a03a25bdf21bee93b30425cab3b31494", + "gpu_telemetry.reanalyzed-fdcc2496.json": "51eff413ad7a1a01bc6116a52e31ac3c39a1f844bf0aa0f0e09d9a0731392801", + "grade.json": "7f69334378ac3187de489eb74ad81c8245267d1a5711612b94879f297e758bcf", + "receipt.json": "58b9ee02a9c9cd03829a91d8f2f4d6d5eedf064f1c96206bc6c4b38522c4bfd8", + "summary.json": "0445d7ff95e0fbfbbc06d21cb3fdcccf8a741249ea294dda3f136a3ea4ed2b56", + "transcript.jsonl": "e887af90ea26038d3d8aadf22998eaa443eeb78eaa59d204bc7fcc741124e00e", + "workspace_final.tar.gz": "90a1b1918640b2a496b66ed5fbaf8649eb5902aa2e4e36c1f6b4debc71c720bc" + }, + "replacement": { + "required": true, + "canonical_run": "p1_refactor_gemma4-31b-q4_v3", + "status": "completed", + "receipt_sha256": "d8b8de86c6abb82b39c59fd2ed56571dc1c05e4c1cf2214fce1a54803bb20dca", + "summary_sha256": "e22036d8e57d95f93243ae7cd3f22cfc271686c8134c715dfc44b7459410fe61", + "server_timeout_seconds": 14400, + "finish_reason": "done_signal" + } + }, + "files": { + "cost.json": { + "bytes": 1241, + "sha256": "c5d1062f477758436ba0c3baf32df97fbe4e30df09c831bde470ccf938fce7f7" + }, + "gpu_telemetry.json": { + "bytes": 2939, + "sha256": "4f717a0e9d7298401c41e1799f772186a03a25bdf21bee93b30425cab3b31494" + }, + "gpu_telemetry.reanalyzed-fdcc2496.json": { + "bytes": 2973, + "sha256": "51eff413ad7a1a01bc6116a52e31ac3c39a1f844bf0aa0f0e09d9a0731392801" + }, + "grade.json": { + "bytes": 510, + "sha256": "7f69334378ac3187de489eb74ad81c8245267d1a5711612b94879f297e758bcf" + }, + "receipt.json": { + "bytes": 8956, + "sha256": "58b9ee02a9c9cd03829a91d8f2f4d6d5eedf064f1c96206bc6c4b38522c4bfd8" + }, + "summary.json": { + "bytes": 321, + "sha256": "0445d7ff95e0fbfbbc06d21cb3fdcccf8a741249ea294dda3f136a3ea4ed2b56" + }, + "transcript.jsonl": { + "bytes": 30545, + "sha256": "e887af90ea26038d3d8aadf22998eaa443eeb78eaa59d204bc7fcc741124e00e" + }, + "workspace_final.tar.gz": { + "bytes": 67984, + "sha256": "90a1b1918640b2a496b66ed5fbaf8649eb5902aa2e4e36c1f6b4debc71c720bc" + } + } + }, + { + "attempt": "p1_testwrite_gemma4-31b-q4_v1_retry-20260802T003822Z-8061c2b0", + "classification": { + "source_run": "p1_testwrite_gemma4-31b-q4_v1", + "classification": "infrastructure-invalid", + "reason_code": "harness-dispatch-keyerror-required-tool-argument", + "incident_document": "canonical-harness-incident-20260802.json", + "classified_before_grade": true, + "affirmative_evidence": [ + "a syntactically valid tool call omitted path", + "the harness raised KeyError instead of returning a recoverable tool error" + ], + "expected_files": { + "receipt.json": "92313dc9d6070ad4fba78fb89a4d043281e17d0e5fb2203e0bdc660c551669e4", + "transcript.jsonl": "ca8b0f10985799cbd7511a4ce32d6b51939f6d6f32139719cb41b9c52d18aeb4" + }, + "replacement": { + "required": true, + "canonical_run": "p1_testwrite_gemma4-31b-q4_v1", + "status": "completed" + } + }, + "files": { + "receipt.json": { + "bytes": 8176, + "sha256": "92313dc9d6070ad4fba78fb89a4d043281e17d0e5fb2203e0bdc660c551669e4" + }, + "transcript.jsonl": { + "bytes": 9815, + "sha256": "ca8b0f10985799cbd7511a4ce32d6b51939f6d6f32139719cb41b9c52d18aeb4" + } + } + }, + { + "attempt": "p1_testwrite_gemma4-31b-q4_v1_retry-20260802T004533Z-09111a0e", + "classification": { + "source_run": "p1_testwrite_gemma4-31b-q4_v1", + "classification": "infrastructure-invalid", + "reason_code": "operator-interrupted-before-required-telemetry-gate", + "incident_document": "canonical-telemetry-gate-incident-20260802.json", + "classified_before_grade": true, + "affirmative_evidence": [ + "the retry began before per-run sidecar attribution was proven", + "the supervisor stopped before the attempt completed" + ], + "expected_files": { + "receipt.json": "e163856856b1a8323e4b881d00a43f0a4361ad00f8d5e800acb3a6d435d259a5", + "transcript.jsonl": "e0efc1aebe693ac92e2a78354316059e5332e761979bdbee220a42019e0367f3" + }, + "replacement": { + "required": true, + "canonical_run": "p1_testwrite_gemma4-31b-q4_v1", + "status": "completed" + } + }, + "files": { + "receipt.json": { + "bytes": 8172, + "sha256": "e163856856b1a8323e4b881d00a43f0a4361ad00f8d5e800acb3a6d435d259a5" + }, + "transcript.jsonl": { + "bytes": 14568, + "sha256": "e0efc1aebe693ac92e2a78354316059e5332e761979bdbee220a42019e0367f3" + } + } + }, + { + "attempt": "p1_testwrite_gemma4-31b-q4_v3_retry-20260802T005418Z-ff262c4a", + "classification": { + "source_run": "p1_testwrite_gemma4-31b-q4_v3", + "classification": "infrastructure-invalid", + "reason_code": "operator-interrupted-companion-after-harness-dispatch-defect", + "incident_document": "canonical-harness-incident-20260802.json", + "classified_before_grade": true, + "affirmative_evidence": [ + "the attempt was active under the same defective harness", + "the supervisor stopped before completion and the attempt was preserved without a grade" + ], + "expected_files": { + "receipt.json": "326934f8660018bdd081c3740222199dc6d5851d011489012ae93166c51440eb", + "transcript.jsonl": "63717381ad72db6d006a97dc24a1da14a557a92542826be3d92c264968556661" + }, + "replacement": { + "required": true, + "canonical_run": "p1_testwrite_gemma4-31b-q4_v3", + "status": "completed" + } + }, + "files": { + "receipt.json": { + "bytes": 8175, + "sha256": "326934f8660018bdd081c3740222199dc6d5851d011489012ae93166c51440eb" + }, + "transcript.jsonl": { + "bytes": 6114, + "sha256": "63717381ad72db6d006a97dc24a1da14a557a92542826be3d92c264968556661" + } + } + } + ] +} diff --git a/benchmarks/gemma4-31b-q4/gemma4-canonical-n10-grader-manifest.json b/benchmarks/gemma4-31b-q4/gemma4-canonical-n10-grader-manifest.json new file mode 100644 index 00000000..f9df0235 --- /dev/null +++ b/benchmarks/gemma4-31b-q4/gemma4-canonical-n10-grader-manifest.json @@ -0,0 +1,1058 @@ +{ + "schema_version": 1, + "generated_at": "2026-08-02T05:41:46.855081+00:00", + "root": "/home/michael/bench-gemma4-31b-q4", + "label": "gemma4-31b-q4", + "target_n": 10, + "passed": true, + "errors": [], + "repository_commit": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "python": "3.12.3 (main, Jun 19 2026, 12:46:00) [GCC 13.3.0]", + "docker_version": "29.2.1", + "bench_sandbox_image_id": "sha256:61fae3bdff98a8a307a3aff35d482cf6ad9bf31f83c151499c99ae4cee64a9cd", + "grader_files": { + "tooling/scripts/grade_microbench.sh": { + "bytes": 5789, + "sha256": "fb65e0be31abeb86ece4cb1eef6c894200266b57fe48a069952c53af84016c9d" + }, + "tooling/graders/phase1_grade.py": { + "bytes": 12036, + "sha256": "cccf4967bccb7e013d5ea29a21cba0b274a8c81037d6bbe9d91e3428389875cb" + }, + "tooling/graders/code_task_grader.py": { + "bytes": 9546, + "sha256": "82aeb55f87f123c7920458a5346412378cfd3f9db100d924f51ea2f4a91a1eb5" + }, + "tooling/graders/phase2_extraction_grade.py": { + "bytes": 4958, + "sha256": "9e390ab2531e72aab64613382b01ce83e60ca6242e9abef87a0deefb502726d6" + }, + "tooling/graders/phase2_ci_failure_grade.py": { + "bytes": 5242, + "sha256": "6ae717a937da8abc4aca74621f67445667ea34b06767d64a8310f0b5fae53ab4" + }, + "tooling/graders/phase2_hallucination_grade.py": { + "bytes": 4945, + "sha256": "24e64182b04810bb3938ecdd9d3886d994e440716f4889952edb69786c7bb237" + }, + "tooling/graders/phase2_triage_grade.py": { + "bytes": 4493, + "sha256": "ea3eebcf6a2fc085276e555f2cb1a54239e1dc9286bac3cb2300f01929b855a1" + }, + "tooling/graders/phase3_doc_synthesis_grade.py": { + "bytes": 4068, + "sha256": "8c117c0a3bfec94a91f94e49bb5b4281a458bcc4a650e9ae86e2450f0f5cac0a" + }, + "tooling/graders/phase3_business_memo_grade.py": { + "bytes": 5996, + "sha256": "691c5c6cec0f3bbf2fc86e80c47dd3fc2bde4873dc630297a0aea13edd548d7b" + }, + "tooling/graders/phase3_market_research_grade.py": { + "bytes": 4585, + "sha256": "e87c20a55bb74e9be883a9b2628fd954a99aaebf937f1d1bd4fbbf111b836de6" + }, + "tooling/graders/phase3_writing_editing_grade.py": { + "bytes": 5520, + "sha256": "a3461f31be9d24cc1c5132fff4bff96a27af77af2457dad839ef5498afba8f56" + }, + "tooling/graders/phase3_project_mgmt_grade.py": { + "bytes": 5706, + "sha256": "e1c3a9190e6d19cffd76734c8c922450b22d674bd5fd2db052fef31f91777bca" + }, + "tooling/graders/ground_truth/phase2_extraction.json": { + "bytes": 4012, + "sha256": "fead89b8051621404257afd0d5cbd5edb3b81cfcc4fe94fd58cf59e546a82d67" + }, + "tooling/graders/ground_truth/phase2_hallucination.json": { + "bytes": 4847, + "sha256": "eb2d03bbb161d1619306904a1850b6ac7106e35cbc49febff431e0f7280900c3" + }, + "tooling/graders/ground_truth/phase2_triage.json": { + "bytes": 5105, + "sha256": "6ca3763f32abfe3a5b8dfa56662c9083f91f81c2c3d06d3117bfb974367669d1" + }, + "tooling/graders/ground_truth/phase3_doc_synthesis.json": { + "bytes": 3569, + "sha256": "20e2baef1b99f4dc9339e87208105c9c5a57cf6b6266f09ce88518d5931ebe9d" + }, + "tooling/graders/ground_truth/phase3_business_memo.json": { + "bytes": 5197, + "sha256": "d422f262007c219a5b2da654e9c1282bfb94e19837eafd33ca70009e66b72313" + }, + "tooling/graders/ground_truth/phase3_market_research_rubric.json": { + "bytes": 3428, + "sha256": "d67aaf1664f2280d1e47855aaeaf543590f40b34e7d42aa194729e5e28fa90a5" + }, + "tooling/graders/ground_truth/phase3_project_mgmt.json": { + "bytes": 4382, + "sha256": "ff97a618bcb888a7f4d67eba55b2e5b1d0b459c694b90db0f6ec4072e0280220" + }, + "tooling/inputs/phase3_writing_editing/audience_briefs.json": { + "bytes": 4254, + "sha256": "0c22573bd97d973d6de73cbbf2e9b806b3099f90872a5f0ec8a4fa33948f94cd" + } + }, + "raw_grades": [ + { + "run_name": "p1_bugfix_gemma4-31b-q4_v1", + "verdict": "PASS", + "grade_json": { + "bytes": 467, + "sha256": "8ef51f6f2d9457eb209c41700017a65b6f13674d9251e4c8b02352a7ad89a4aa" + } + }, + { + "run_name": "p1_bugfix_gemma4-31b-q4_v2", + "verdict": "PASS", + "grade_json": { + "bytes": 468, + "sha256": "8c1caead3b6ffa35ac77aefd6150cc3bc5513596a7f125c9e024d34226cdfe3f" + } + }, + { + "run_name": "p1_bugfix_gemma4-31b-q4_v3", + "verdict": "FAIL", + "grade_json": { + "bytes": 468, + "sha256": "10243cc2a1b0f2c8df720128ec698608433a7b5176a9ffe716efd29205815667" + } + }, + { + "run_name": "p1_bugfix_gemma4-31b-q4_v4", + "verdict": "FAIL", + "grade_json": { + "bytes": 468, + "sha256": "2a50cd7db6267876fdbfc16a2bb28a3043d64cdc64ff50bfb627945c77bcfd57" + } + }, + { + "run_name": "p1_bugfix_gemma4-31b-q4_v5", + "verdict": "FAIL", + "grade_json": { + "bytes": 467, + "sha256": "3182de43effbec4148a35605a10fbcf8ec4479c3ece279d600fd34ba81243341" + } + }, + { + "run_name": "p1_bugfix_gemma4-31b-q4_v6", + "verdict": "PASS", + "grade_json": { + "bytes": 478, + "sha256": "d001be623e65b5b8df177bb5bed6ffcf720c20c1b09d9ce5410069011f837512" + } + }, + { + "run_name": "p1_bugfix_gemma4-31b-q4_v7", + "verdict": "FAIL", + "grade_json": { + "bytes": 467, + "sha256": "a98e38961cf492999b6fe0c7dffe92d4f2e9599e4324e27e32799fefe44ee66e" + } + }, + { + "run_name": "p1_bugfix_gemma4-31b-q4_v8", + "verdict": "FAIL", + "grade_json": { + "bytes": 469, + "sha256": "76c1029d08602ceb16361a7afaea843a018078ddc51151e7e969d2cdb5a263d9" + } + }, + { + "run_name": "p1_bugfix_gemma4-31b-q4_v9", + "verdict": "PASS", + "grade_json": { + "bytes": 480, + "sha256": "aea0e00c9cd364a33af96808e50aab16ed1a942f1b6c59f7bde8079586f7b452" + } + }, + { + "run_name": "p1_bugfix_gemma4-31b-q4_v10", + "verdict": "FAIL", + "grade_json": { + "bytes": 469, + "sha256": "cc4b412a05ac78893698a0ffe05fa4a3519ea42cea3051bdc679b29e472aa0c9" + } + }, + { + "run_name": "p1_testwrite_gemma4-31b-q4_v1", + "verdict": "FAIL", + "grade_json": { + "bytes": 1155, + "sha256": "73b7129ab31845b98ef25c4e1501e5367510c8b7fbda84357be4bb3294bbd2f3" + } + }, + { + "run_name": "p1_testwrite_gemma4-31b-q4_v2", + "verdict": "PASS", + "grade_json": { + "bytes": 341, + "sha256": "42c1da43e1c31be25775668d65010e502b8383356f35bfd846be7b583f7e153e" + } + }, + { + "run_name": "p1_testwrite_gemma4-31b-q4_v3", + "verdict": "PASS", + "grade_json": { + "bytes": 341, + "sha256": "b99e9530499f12935a79df989d2f182b0917530a4b3be44b95eaa26e7f2d2c33" + } + }, + { + "run_name": "p1_testwrite_gemma4-31b-q4_v4", + "verdict": "FAIL", + "grade_json": { + "bytes": 338, + "sha256": "a76068faab9d01220a3493092cadd4a8630aa720dc624df35bbb71b4f2772e05" + } + }, + { + "run_name": "p1_testwrite_gemma4-31b-q4_v5", + "verdict": "PASS", + "grade_json": { + "bytes": 341, + "sha256": "74243a504afb05edd995955710a388f8e221516f754407850844113c4c26e3db" + } + }, + { + "run_name": "p1_testwrite_gemma4-31b-q4_v6", + "verdict": "PASS", + "grade_json": { + "bytes": 341, + "sha256": "eeff0a8acc36edbcdc0b38e3fac798d8ec73d59bca53d8dc0664c4a8abf832ce" + } + }, + { + "run_name": "p1_testwrite_gemma4-31b-q4_v7", + "verdict": "PASS", + "grade_json": { + "bytes": 341, + "sha256": "0f83751c074427d12d916faba8ffd786e54d8feb1bd8d37ef8e6ada4128ea5a3" + } + }, + { + "run_name": "p1_testwrite_gemma4-31b-q4_v8", + "verdict": "PASS", + "grade_json": { + "bytes": 341, + "sha256": "cf4fb96563994dcd182edffa7fb117b91981086cef9907b87d1de3afa262b6dc" + } + }, + { + "run_name": "p1_testwrite_gemma4-31b-q4_v9", + "verdict": "PASS", + "grade_json": { + "bytes": 341, + "sha256": "bc0138be4f3c21f7687420af1fa496c244e735495742d7b09d1ba0d6818add51" + } + }, + { + "run_name": "p1_testwrite_gemma4-31b-q4_v10", + "verdict": "PASS", + "grade_json": { + "bytes": 341, + "sha256": "6ae706d907c18507866cea3b27f3f46bb025463558df00847cfcf905f4b50e77" + } + }, + { + "run_name": "p1_refactor_gemma4-31b-q4_v1", + "verdict": "PASS", + "grade_json": { + "bytes": 510, + "sha256": "7f69334378ac3187de489eb74ad81c8245267d1a5711612b94879f297e758bcf" + } + }, + { + "run_name": "p1_refactor_gemma4-31b-q4_v2", + "verdict": "PASS", + "grade_json": { + "bytes": 510, + "sha256": "7f69334378ac3187de489eb74ad81c8245267d1a5711612b94879f297e758bcf" + } + }, + { + "run_name": "p1_refactor_gemma4-31b-q4_v3", + "verdict": "PASS", + "grade_json": { + "bytes": 510, + "sha256": "7f69334378ac3187de489eb74ad81c8245267d1a5711612b94879f297e758bcf" + } + }, + { + "run_name": "p1_refactor_gemma4-31b-q4_v4", + "verdict": "PASS", + "grade_json": { + "bytes": 510, + "sha256": "7f69334378ac3187de489eb74ad81c8245267d1a5711612b94879f297e758bcf" + } + }, + { + "run_name": "p1_refactor_gemma4-31b-q4_v5", + "verdict": "PASS", + "grade_json": { + "bytes": 510, + "sha256": "7f69334378ac3187de489eb74ad81c8245267d1a5711612b94879f297e758bcf" + } + }, + { + "run_name": "p1_refactor_gemma4-31b-q4_v6", + "verdict": "FAIL", + "grade_json": { + "bytes": 529, + "sha256": "3b343e7e30b0fd413c183bcf25114cab48978d1928374717ad6eba1cc327e73d" + } + }, + { + "run_name": "p1_refactor_gemma4-31b-q4_v7", + "verdict": "FAIL", + "grade_json": { + "bytes": 529, + "sha256": "3b343e7e30b0fd413c183bcf25114cab48978d1928374717ad6eba1cc327e73d" + } + }, + { + "run_name": "p1_refactor_gemma4-31b-q4_v8", + "verdict": "PASS", + "grade_json": { + "bytes": 510, + "sha256": "7f69334378ac3187de489eb74ad81c8245267d1a5711612b94879f297e758bcf" + } + }, + { + "run_name": "p1_refactor_gemma4-31b-q4_v9", + "verdict": "FAIL", + "grade_json": { + "bytes": 411, + "sha256": "fca70add1343783f2d9a6e3bc82d5ae1e3c2ad48a00cce234501fd04b15cbe4f" + } + }, + { + "run_name": "p1_refactor_gemma4-31b-q4_v10", + "verdict": "PASS", + "grade_json": { + "bytes": 510, + "sha256": "7f69334378ac3187de489eb74ad81c8245267d1a5711612b94879f297e758bcf" + } + }, + { + "run_name": "p2_extract_gemma4-31b-q4_v1", + "verdict": "PASS", + "grade_json": { + "bytes": 3204, + "sha256": "b6bc97fc6e03c0b85c34d7e6b5aa1ffcfdf128e953247bc7569f9417170aa2b8" + } + }, + { + "run_name": "p2_extract_gemma4-31b-q4_v2", + "verdict": "PASS", + "grade_json": { + "bytes": 3204, + "sha256": "b6bc97fc6e03c0b85c34d7e6b5aa1ffcfdf128e953247bc7569f9417170aa2b8" + } + }, + { + "run_name": "p2_extract_gemma4-31b-q4_v3", + "verdict": "PASS", + "grade_json": { + "bytes": 3204, + "sha256": "b6bc97fc6e03c0b85c34d7e6b5aa1ffcfdf128e953247bc7569f9417170aa2b8" + } + }, + { + "run_name": "p2_extract_gemma4-31b-q4_v4", + "verdict": "PASS", + "grade_json": { + "bytes": 3204, + "sha256": "b6bc97fc6e03c0b85c34d7e6b5aa1ffcfdf128e953247bc7569f9417170aa2b8" + } + }, + { + "run_name": "p2_extract_gemma4-31b-q4_v5", + "verdict": "PASS", + "grade_json": { + "bytes": 3204, + "sha256": "b6bc97fc6e03c0b85c34d7e6b5aa1ffcfdf128e953247bc7569f9417170aa2b8" + } + }, + { + "run_name": "p2_extract_gemma4-31b-q4_v6", + "verdict": "PASS", + "grade_json": { + "bytes": 3204, + "sha256": "b6bc97fc6e03c0b85c34d7e6b5aa1ffcfdf128e953247bc7569f9417170aa2b8" + } + }, + { + "run_name": "p2_extract_gemma4-31b-q4_v7", + "verdict": "PASS", + "grade_json": { + "bytes": 3204, + "sha256": "b6bc97fc6e03c0b85c34d7e6b5aa1ffcfdf128e953247bc7569f9417170aa2b8" + } + }, + { + "run_name": "p2_extract_gemma4-31b-q4_v8", + "verdict": "PASS", + "grade_json": { + "bytes": 3204, + "sha256": "b6bc97fc6e03c0b85c34d7e6b5aa1ffcfdf128e953247bc7569f9417170aa2b8" + } + }, + { + "run_name": "p2_extract_gemma4-31b-q4_v9", + "verdict": "PASS", + "grade_json": { + "bytes": 3204, + "sha256": "b6bc97fc6e03c0b85c34d7e6b5aa1ffcfdf128e953247bc7569f9417170aa2b8" + } + }, + { + "run_name": "p2_extract_gemma4-31b-q4_v10", + "verdict": "PASS", + "grade_json": { + "bytes": 3204, + "sha256": "b6bc97fc6e03c0b85c34d7e6b5aa1ffcfdf128e953247bc7569f9417170aa2b8" + } + }, + { + "run_name": "p2_ci_gemma4-31b-q4_v1", + "verdict": "PASS", + "grade_json": { + "bytes": 482, + "sha256": "69e86f389fb98ab93705abc5fd07431d5a9730ba5b0214d67cccdc8eeb154ba6" + } + }, + { + "run_name": "p2_ci_gemma4-31b-q4_v2", + "verdict": "PASS", + "grade_json": { + "bytes": 482, + "sha256": "eae0f41e05a5ce0e1a370e295ccb33bfd10b26e892ff34e286253904640fd666" + } + }, + { + "run_name": "p2_ci_gemma4-31b-q4_v3", + "verdict": "PASS", + "grade_json": { + "bytes": 482, + "sha256": "1cdcddf815967c7395bceb0d8ac5b955294d82d5d932594f553e43f22ab4234d" + } + }, + { + "run_name": "p2_ci_gemma4-31b-q4_v4", + "verdict": "PASS", + "grade_json": { + "bytes": 482, + "sha256": "eae0f41e05a5ce0e1a370e295ccb33bfd10b26e892ff34e286253904640fd666" + } + }, + { + "run_name": "p2_ci_gemma4-31b-q4_v5", + "verdict": "PASS", + "grade_json": { + "bytes": 482, + "sha256": "1cdcddf815967c7395bceb0d8ac5b955294d82d5d932594f553e43f22ab4234d" + } + }, + { + "run_name": "p2_ci_gemma4-31b-q4_v6", + "verdict": "PASS", + "grade_json": { + "bytes": 482, + "sha256": "1cdcddf815967c7395bceb0d8ac5b955294d82d5d932594f553e43f22ab4234d" + } + }, + { + "run_name": "p2_ci_gemma4-31b-q4_v7", + "verdict": "PASS", + "grade_json": { + "bytes": 482, + "sha256": "eae0f41e05a5ce0e1a370e295ccb33bfd10b26e892ff34e286253904640fd666" + } + }, + { + "run_name": "p2_ci_gemma4-31b-q4_v8", + "verdict": "PASS", + "grade_json": { + "bytes": 482, + "sha256": "eae0f41e05a5ce0e1a370e295ccb33bfd10b26e892ff34e286253904640fd666" + } + }, + { + "run_name": "p2_ci_gemma4-31b-q4_v9", + "verdict": "PASS", + "grade_json": { + "bytes": 482, + "sha256": "1cdcddf815967c7395bceb0d8ac5b955294d82d5d932594f553e43f22ab4234d" + } + }, + { + "run_name": "p2_ci_gemma4-31b-q4_v10", + "verdict": "PASS", + "grade_json": { + "bytes": 482, + "sha256": "eae0f41e05a5ce0e1a370e295ccb33bfd10b26e892ff34e286253904640fd666" + } + }, + { + "run_name": "p2_hallucination_gemma4-31b-q4_v1", + "verdict": "PASS", + "grade_json": { + "bytes": 3586, + "sha256": "80910919c9e64a98ef035d0e5da409c4b26fc67de0ad65bcf63df6b4012393e6" + } + }, + { + "run_name": "p2_hallucination_gemma4-31b-q4_v2", + "verdict": "PASS", + "grade_json": { + "bytes": 3892, + "sha256": "2553e9edff94d4bf1cc8fda5c170ae25201bd4baa261fc4ab14d5754dbc11bed" + } + }, + { + "run_name": "p2_hallucination_gemma4-31b-q4_v3", + "verdict": "MISSING_OUTPUT", + "grade_json": { + "bytes": 153, + "sha256": "e4f8b9da1c7aba632e47e0a97771274c35662e1f1e2920b050928e498df8692a" + } + }, + { + "run_name": "p2_hallucination_gemma4-31b-q4_v4", + "verdict": "MISSING_OUTPUT", + "grade_json": { + "bytes": 153, + "sha256": "48183e26e443a49c4cc81ec4584bfa80e69da97f0759b05cf58e3c5b45bdc8f8" + } + }, + { + "run_name": "p2_hallucination_gemma4-31b-q4_v5", + "verdict": "MISSING_OUTPUT", + "grade_json": { + "bytes": 153, + "sha256": "019bbdcdc8ce03c38b61fabde3ebe89220203a5af63f6a565860128fd32d3a41" + } + }, + { + "run_name": "p2_hallucination_gemma4-31b-q4_v6", + "verdict": "PASS", + "grade_json": { + "bytes": 3775, + "sha256": "2e85eba6943daf1b7e00211c0a4339e3d91c4942516de7bdc6468b260d7665fe" + } + }, + { + "run_name": "p2_hallucination_gemma4-31b-q4_v7", + "verdict": "MISSING_OUTPUT", + "grade_json": { + "bytes": 153, + "sha256": "13f5071c95d5ba1db4673f578df59ec4f3441f0e16cce43008f8192563aebee9" + } + }, + { + "run_name": "p2_hallucination_gemma4-31b-q4_v8", + "verdict": "PASS", + "grade_json": { + "bytes": 3586, + "sha256": "80910919c9e64a98ef035d0e5da409c4b26fc67de0ad65bcf63df6b4012393e6" + } + }, + { + "run_name": "p2_hallucination_gemma4-31b-q4_v9", + "verdict": "PASS", + "grade_json": { + "bytes": 3836, + "sha256": "3c8f6c79a0738d920981fde9054c69e95edfb470cd987ad7b94a29a5fb160a1e" + } + }, + { + "run_name": "p2_hallucination_gemma4-31b-q4_v10", + "verdict": "PASS", + "grade_json": { + "bytes": 3826, + "sha256": "b934ae4e4b2b6163c0101cb6595b3d55ecc27ebcd931f263c0fb7572153204d5" + } + }, + { + "run_name": "p2_triage_gemma4-31b-q4_v1", + "verdict": "PASS", + "grade_json": { + "bytes": 2250, + "sha256": "dbe6f269c365da51b85a5053d621051e8877fc7a7642fcc441c97f6c7653b5f5" + } + }, + { + "run_name": "p2_triage_gemma4-31b-q4_v2", + "verdict": "PASS", + "grade_json": { + "bytes": 2159, + "sha256": "c14e218d16bc8ba2bfb56fd8577d0f8539d628d0366f9340f7d2f60f17c91290" + } + }, + { + "run_name": "p2_triage_gemma4-31b-q4_v3", + "verdict": "PASS", + "grade_json": { + "bytes": 1914, + "sha256": "8b145bdcbc912d2f546ce33b90d11fe348c69829b205abe91d6513b83ba9daef" + } + }, + { + "run_name": "p2_triage_gemma4-31b-q4_v4", + "verdict": "PASS", + "grade_json": { + "bytes": 1727, + "sha256": "0e5b1aa18e87a58f55f541ac729fc931d1e38dfa5205b3a047e2ad859cd06a04" + } + }, + { + "run_name": "p2_triage_gemma4-31b-q4_v5", + "verdict": "PASS", + "grade_json": { + "bytes": 2156, + "sha256": "0d1c9775790296f34db998a3c53224e4ad0516b0823bfa1cd074561a6b73d6c4" + } + }, + { + "run_name": "p2_triage_gemma4-31b-q4_v6", + "verdict": "PASS", + "grade_json": { + "bytes": 1823, + "sha256": "3d692968f27afc71beba131cd873661ed419e70ba9094a5b9bba5ede51d9d62d" + } + }, + { + "run_name": "p2_triage_gemma4-31b-q4_v7", + "verdict": "PASS", + "grade_json": { + "bytes": 2162, + "sha256": "e1f80a2b26dbf4d4a89166e3aeaa1d90b604d68fbb99286aa5e801231483922b" + } + }, + { + "run_name": "p2_triage_gemma4-31b-q4_v8", + "verdict": "FAIL", + "grade_json": { + "bytes": 2345, + "sha256": "ddd1baea329ab8b63b59b7f1fc949dce4d551b8f25b0816f865b1d3a803879e2" + } + }, + { + "run_name": "p2_triage_gemma4-31b-q4_v9", + "verdict": "FAIL", + "grade_json": { + "bytes": 2444, + "sha256": "8746c950be101673b46c2c22be3c3adfb6f214618b06b31788618933b8d74075" + } + }, + { + "run_name": "p2_triage_gemma4-31b-q4_v10", + "verdict": "PASS", + "grade_json": { + "bytes": 1541, + "sha256": "98c725368d22f6b29f262e70ca2cf1a55a7d70c0409f1ab0087a77a938b08a8d" + } + }, + { + "run_name": "p3_doc_gemma4-31b-q4_v1", + "verdict": "PASS", + "grade_json": { + "bytes": 2357, + "sha256": "addcd3cfc419649494adf40c60fd2385ec3b403f91e07f3423695aac6fd5c0dc" + } + }, + { + "run_name": "p3_doc_gemma4-31b-q4_v2", + "verdict": "PASS", + "grade_json": { + "bytes": 2357, + "sha256": "addcd3cfc419649494adf40c60fd2385ec3b403f91e07f3423695aac6fd5c0dc" + } + }, + { + "run_name": "p3_doc_gemma4-31b-q4_v3", + "verdict": "PASS", + "grade_json": { + "bytes": 2360, + "sha256": "a1e32ee1296ef6f8f122b96d91586e2fbd3df7e2ff1f5aea3b13c97c6cb2bdd8" + } + }, + { + "run_name": "p3_doc_gemma4-31b-q4_v4", + "verdict": "PASS", + "grade_json": { + "bytes": 2387, + "sha256": "10c1a63e62ecc97494797071010538305d197ff084808363feabdca104bfbbf1" + } + }, + { + "run_name": "p3_doc_gemma4-31b-q4_v5", + "verdict": "PASS", + "grade_json": { + "bytes": 2356, + "sha256": "2b471277132b5cb8888e33dfb28d6609a82f2df57ab635b45f369f157dff3d9d" + } + }, + { + "run_name": "p3_doc_gemma4-31b-q4_v6", + "verdict": "PASS", + "grade_json": { + "bytes": 2356, + "sha256": "2b471277132b5cb8888e33dfb28d6609a82f2df57ab635b45f369f157dff3d9d" + } + }, + { + "run_name": "p3_doc_gemma4-31b-q4_v7", + "verdict": "PASS", + "grade_json": { + "bytes": 2356, + "sha256": "2b471277132b5cb8888e33dfb28d6609a82f2df57ab635b45f369f157dff3d9d" + } + }, + { + "run_name": "p3_doc_gemma4-31b-q4_v8", + "verdict": "PASS", + "grade_json": { + "bytes": 2356, + "sha256": "2b471277132b5cb8888e33dfb28d6609a82f2df57ab635b45f369f157dff3d9d" + } + }, + { + "run_name": "p3_doc_gemma4-31b-q4_v9", + "verdict": "PASS", + "grade_json": { + "bytes": 2357, + "sha256": "addcd3cfc419649494adf40c60fd2385ec3b403f91e07f3423695aac6fd5c0dc" + } + }, + { + "run_name": "p3_doc_gemma4-31b-q4_v10", + "verdict": "PASS", + "grade_json": { + "bytes": 2357, + "sha256": "addcd3cfc419649494adf40c60fd2385ec3b403f91e07f3423695aac6fd5c0dc" + } + }, + { + "run_name": "p3_business_gemma4-31b-q4_v1", + "verdict": "FAIL", + "grade_json": { + "bytes": 3866, + "sha256": "3c85cb074672401082ba6adecc5f454e5124b81323f35fc312a8458bba0ab815" + } + }, + { + "run_name": "p3_business_gemma4-31b-q4_v2", + "verdict": "PASS", + "grade_json": { + "bytes": 3873, + "sha256": "13358399cce26683f59886f1cfc577330fd90233e4065a9af249f51edd05c4ab" + } + }, + { + "run_name": "p3_business_gemma4-31b-q4_v3", + "verdict": "PASS", + "grade_json": { + "bytes": 3873, + "sha256": "13358399cce26683f59886f1cfc577330fd90233e4065a9af249f51edd05c4ab" + } + }, + { + "run_name": "p3_business_gemma4-31b-q4_v4", + "verdict": "FAIL", + "grade_json": { + "bytes": 3866, + "sha256": "3c85cb074672401082ba6adecc5f454e5124b81323f35fc312a8458bba0ab815" + } + }, + { + "run_name": "p3_business_gemma4-31b-q4_v5", + "verdict": "FAIL", + "grade_json": { + "bytes": 3866, + "sha256": "3c85cb074672401082ba6adecc5f454e5124b81323f35fc312a8458bba0ab815" + } + }, + { + "run_name": "p3_business_gemma4-31b-q4_v6", + "verdict": "PASS", + "grade_json": { + "bytes": 3870, + "sha256": "8a18dd536897e3e16058a2b4c0f9f443faeeb0b1c71986a3a75245703cd484d1" + } + }, + { + "run_name": "p3_business_gemma4-31b-q4_v7", + "verdict": "PASS", + "grade_json": { + "bytes": 3870, + "sha256": "8a18dd536897e3e16058a2b4c0f9f443faeeb0b1c71986a3a75245703cd484d1" + } + }, + { + "run_name": "p3_business_gemma4-31b-q4_v8", + "verdict": "PASS", + "grade_json": { + "bytes": 3870, + "sha256": "8a18dd536897e3e16058a2b4c0f9f443faeeb0b1c71986a3a75245703cd484d1" + } + }, + { + "run_name": "p3_business_gemma4-31b-q4_v9", + "verdict": "FAIL", + "grade_json": { + "bytes": 3854, + "sha256": "8cd481bebf1c8c028e4ea2ea8a6157a3e21f758053374f05a45b73457036e209" + } + }, + { + "run_name": "p3_business_gemma4-31b-q4_v10", + "verdict": "PASS", + "grade_json": { + "bytes": 3870, + "sha256": "8a18dd536897e3e16058a2b4c0f9f443faeeb0b1c71986a3a75245703cd484d1" + } + }, + { + "run_name": "p3_market_gemma4-31b-q4_v1", + "verdict": "STRUCTURAL_PASS", + "grade_json": { + "bytes": 991, + "sha256": "ec7f817b56a4a6ab2128042d09466fdcdf5b77eaf246ea588cbb987a0865d5fa" + } + }, + { + "run_name": "p3_market_gemma4-31b-q4_v2", + "verdict": "STRUCTURAL_PASS", + "grade_json": { + "bytes": 991, + "sha256": "7a993d6b3110622d5558f7d78f2ebb57044bb1bdf2b0e9f52585b7fdcf81dddc" + } + }, + { + "run_name": "p3_market_gemma4-31b-q4_v3", + "verdict": "STRUCTURAL_PASS", + "grade_json": { + "bytes": 990, + "sha256": "5d1bcb2c532b0323a5acdb93607cd0a957b6254f0dc687a6751be84f10ba86b2" + } + }, + { + "run_name": "p3_market_gemma4-31b-q4_v4", + "verdict": "STRUCTURAL_PASS", + "grade_json": { + "bytes": 991, + "sha256": "3aa27fea1e0bb4aa05ff5805aa71f75268e0d1f32d9a65ac321b2fbbfed88ecd" + } + }, + { + "run_name": "p3_market_gemma4-31b-q4_v5", + "verdict": "STRUCTURAL_PASS", + "grade_json": { + "bytes": 991, + "sha256": "3234ae34a10c1abc8642ebfef870e65a48e9fe8740e174c611fb51428cb6b334" + } + }, + { + "run_name": "p3_market_gemma4-31b-q4_v6", + "verdict": "STRUCTURAL_PASS", + "grade_json": { + "bytes": 990, + "sha256": "bb5f60e04d85a7cf9d8aa866f0ef183cc643209b16f8d8cce14ec67c1c223b49" + } + }, + { + "run_name": "p3_market_gemma4-31b-q4_v7", + "verdict": "STRUCTURAL_PASS", + "grade_json": { + "bytes": 990, + "sha256": "5b462c433ddb337589a2f4efea0e65c99a966b6941f3f242231f69977831f002" + } + }, + { + "run_name": "p3_market_gemma4-31b-q4_v8", + "verdict": "STRUCTURAL_PASS", + "grade_json": { + "bytes": 991, + "sha256": "4e17cab2549fc8eefb34d1fbfc11c34b7bc33abd78a70a48955a403a19f02333" + } + }, + { + "run_name": "p3_market_gemma4-31b-q4_v9", + "verdict": "STRUCTURAL_PASS", + "grade_json": { + "bytes": 991, + "sha256": "ac1d5b61765d23524f6d2b292d9afb52d63e4086180ac1060c1c17b044ed23cf" + } + }, + { + "run_name": "p3_market_gemma4-31b-q4_v10", + "verdict": "STRUCTURAL_PASS", + "grade_json": { + "bytes": 991, + "sha256": "bae011370d19b11261427cb80a48bd422914bc2aabfd2df7a51be42616be2355" + } + }, + { + "run_name": "p3_writing_gemma4-31b-q4_v1", + "verdict": "PASS", + "grade_json": { + "bytes": 2879, + "sha256": "4d52cac7e70665c17a000630f5e703deb8c2dd7fa09b140e58cd916809855594" + } + }, + { + "run_name": "p3_writing_gemma4-31b-q4_v2", + "verdict": "PASS", + "grade_json": { + "bytes": 2879, + "sha256": "d5b96fa97f5173b2c03372286d3f1337cdad83a2d8433953285d8c1b96d080de" + } + }, + { + "run_name": "p3_writing_gemma4-31b-q4_v3", + "verdict": "PASS", + "grade_json": { + "bytes": 2876, + "sha256": "5554d7976a5268e573cd20176244ba7db50716c7dc7dad7df89386f9fc269564" + } + }, + { + "run_name": "p3_writing_gemma4-31b-q4_v4", + "verdict": "PASS", + "grade_json": { + "bytes": 2879, + "sha256": "b85303ff36d7da323a83154c84a296f2782ff66d51e5b0db1ef56a1379dc369b" + } + }, + { + "run_name": "p3_writing_gemma4-31b-q4_v5", + "verdict": "PASS", + "grade_json": { + "bytes": 2876, + "sha256": "5554d7976a5268e573cd20176244ba7db50716c7dc7dad7df89386f9fc269564" + } + }, + { + "run_name": "p3_writing_gemma4-31b-q4_v6", + "verdict": "PASS", + "grade_json": { + "bytes": 2879, + "sha256": "b85303ff36d7da323a83154c84a296f2782ff66d51e5b0db1ef56a1379dc369b" + } + }, + { + "run_name": "p3_writing_gemma4-31b-q4_v7", + "verdict": "PASS", + "grade_json": { + "bytes": 2876, + "sha256": "445eea72603fb268fb807b30c88e47ca7e8f6cae3b63812da43e4d4687dab137" + } + }, + { + "run_name": "p3_writing_gemma4-31b-q4_v8", + "verdict": "PASS", + "grade_json": { + "bytes": 2879, + "sha256": "a7f0d960afecef69f7e22bb44f38920d3b47486f61e8bd0d9c9d52c1fff00220" + } + }, + { + "run_name": "p3_writing_gemma4-31b-q4_v9", + "verdict": "PASS", + "grade_json": { + "bytes": 2879, + "sha256": "fe1a7369ee2c3314701f72704ca898889ee1d4952d825227c6a9d703d907aea3" + } + }, + { + "run_name": "p3_writing_gemma4-31b-q4_v10", + "verdict": "PASS", + "grade_json": { + "bytes": 2876, + "sha256": "5554d7976a5268e573cd20176244ba7db50716c7dc7dad7df89386f9fc269564" + } + }, + { + "run_name": "p3_pm_gemma4-31b-q4_v1", + "verdict": "FAIL", + "grade_json": { + "bytes": 2404, + "sha256": "eade149c12512150ebb28d5a446c6dc1037e0ded24a3ed5e83d65d87aba47bc4" + } + }, + { + "run_name": "p3_pm_gemma4-31b-q4_v2", + "verdict": "FAIL", + "grade_json": { + "bytes": 2425, + "sha256": "5e98c1230f111ddb4d1a15ae7063c80590fbee19b72661c4b56e9c94f0b12352" + } + }, + { + "run_name": "p3_pm_gemma4-31b-q4_v3", + "verdict": "FAIL", + "grade_json": { + "bytes": 2408, + "sha256": "be7951b5c0b3770b9ee61e8f3bd5042dfc0fa486d2be1086b97314e4a8976ca3" + } + }, + { + "run_name": "p3_pm_gemma4-31b-q4_v4", + "verdict": "FAIL", + "grade_json": { + "bytes": 2429, + "sha256": "013a8862905aa0eac7318b00edaa3a3e2f2ca542c4a329ef794b38486fd7a73b" + } + }, + { + "run_name": "p3_pm_gemma4-31b-q4_v5", + "verdict": "FAIL", + "grade_json": { + "bytes": 2429, + "sha256": "013a8862905aa0eac7318b00edaa3a3e2f2ca542c4a329ef794b38486fd7a73b" + } + }, + { + "run_name": "p3_pm_gemma4-31b-q4_v6", + "verdict": "FAIL", + "grade_json": { + "bytes": 2425, + "sha256": "5e98c1230f111ddb4d1a15ae7063c80590fbee19b72661c4b56e9c94f0b12352" + } + }, + { + "run_name": "p3_pm_gemma4-31b-q4_v7", + "verdict": "FAIL", + "grade_json": { + "bytes": 2425, + "sha256": "5e98c1230f111ddb4d1a15ae7063c80590fbee19b72661c4b56e9c94f0b12352" + } + }, + { + "run_name": "p3_pm_gemma4-31b-q4_v8", + "verdict": "FAIL", + "grade_json": { + "bytes": 2420, + "sha256": "5e5d3ca47674ec81eb8de446669c70129aff454157ba42e6bc018d6a13e64818" + } + }, + { + "run_name": "p3_pm_gemma4-31b-q4_v9", + "verdict": "FAIL", + "grade_json": { + "bytes": 2426, + "sha256": "8defbbde59eb9e22c6aa51b5bd69e93bcf5a548d56021d8f45b5fada421c1021" + } + }, + { + "run_name": "p3_pm_gemma4-31b-q4_v10", + "verdict": "FAIL", + "grade_json": { + "bytes": 2426, + "sha256": "8defbbde59eb9e22c6aa51b5bd69e93bcf5a548d56021d8f45b5fada421c1021" + } + } + ], + "correction_policy": "grade.json and label.json hashes are immutable raw evidence. Any correction must be a separate overlay naming the original hash, corrected grader hash, unchanged workspace archive hash, and reproducible defect." +} diff --git a/benchmarks/gemma4-31b-q4/gemma4-canonical-n10-project-mgmt-correction.json b/benchmarks/gemma4-31b-q4/gemma4-canonical-n10-project-mgmt-correction.json new file mode 100644 index 00000000..5defecf7 --- /dev/null +++ b/benchmarks/gemma4-31b-q4/gemma4-canonical-n10-project-mgmt-correction.json @@ -0,0 +1,1600 @@ +{ + "schema_version": 1, + "generated_at": "2026-08-02T05:43:15.002734+00:00", + "campaign": "gemma4-31b-q4-mmbt", + "target_n": 10, + "scope": "p3_pm legacy lexical false negatives only", + "policy": "immutable grade.json files remain raw evidence; this overlay changes no run artifact or raw verdict", + "correction_script": { + "path": "/home/michael/bench-gemma4-31b-q4/tooling/correct_gemma4_project_mgmt_grades.py", + "sha256": "86eeec68b6c7114fb0672566220938e35fa605cd1b89ac4b0863f7cfe305f581" + }, + "raw_grader": { + "path": "/home/michael/bench-gemma4-31b-q4/tooling/graders/phase3_project_mgmt_grade.py", + "sha256": "e1c3a9190e6d19cffd76734c8c922450b22d674bd5fd2db052fef31f91777bca" + }, + "rules": { + "R2": { + "description": "Maevia GA-to-private-beta pushback is expressed without one contiguous legacy keyword", + "patterns": [ + "\\bmaevia\\b.{0,240}\\b(?:push(?:ed)?[ -]?back|fallout|promis(?:e|ed)\\s+ga|private[ -]?beta)\\b" + ] + }, + "R3": { + "description": "legal/private-beta contract delay uses an equivalent non-contracted phrase", + "patterns": [ + "\\blegal\\b.{0,200}\\b(?:has\\s+not|not\\s+yet\\s+responded|sign[ -]?off|approval|unresponsive|silent|delay|contract)\\b" + ] + }, + "D3_mobile": { + "description": "web-responsive V1 and native V2 uses a hyphen the legacy keyword omits", + "patterns": [ + "\\bweb[ -]?responsive\\b.{0,160}\\bnative\\b.{0,100}\\bv2\\b" + ] + }, + "D4_option_b": { + "description": "private beta followed by the selected-customer count is Option B", + "patterns": [ + "\\bprivate[ -]?beta\\b.{0,120}\\b(?:3\\s*[-\u2013]\\s*5|selected\\s+customers?)\\b" + ] + } + }, + "aggregate": { + "cells": 10, + "raw_passes": 0, + "corrected_passes": 10, + "verdict_changes": 10 + }, + "cells": [ + { + "run_name": "p3_pm_gemma4-31b-q4_v1", + "raw_grade_sha256": "eade149c12512150ebb28d5a446c6dc1037e0ded24a3ed5e83d65d87aba47bc4", + "workspace_archive_sha256": "eb622a02fddb70c3aee43f652123e59640282815e55b230d70874175a61fcf08", + "status_report_source": "workspace_final.tar.gz:./status_report.md", + "status_report_sha256": "3c6f102bd297b6e8eaaeb15e0b1e67cadc359f09d01c3fd7b0f35005161ca48f", + "raw_verdict": "FAIL", + "corrected_verdict": "PASS", + "changes": [ + { + "item": "R2", + "category": "risks", + "reason": "Maevia GA-to-private-beta pushback is expressed without one contiguous legacy keyword", + "pattern": "\\bmaevia\\b.{0,240}\\b(?:push(?:ed)?[ -]?back|fallout|promis(?:e|ed)\\s+ga|private[ -]?beta)\\b" + }, + { + "item": "R3", + "category": "risks", + "reason": "legal/private-beta contract delay uses an equivalent non-contracted phrase", + "pattern": "\\blegal\\b.{0,200}\\b(?:has\\s+not|not\\s+yet\\s+responded|sign[ -]?off|approval|unresponsive|silent|delay|contract)\\b" + }, + { + "item": "D3_mobile", + "category": "decisions", + "reason": "web-responsive V1 and native V2 uses a hyphen the legacy keyword omits", + "pattern": "\\bweb[ -]?responsive\\b.{0,160}\\bnative\\b.{0,100}\\bv2\\b" + }, + { + "item": "D4_option_b", + "category": "decisions", + "reason": "private beta followed by the selected-customer count is Option B", + "pattern": "\\bprivate[ -]?beta\\b.{0,120}\\b(?:3\\s*[-\u2013]\\s*5|selected\\s+customers?)\\b" + } + ], + "corrected_grade": { + "task": "project_mgmt", + "verdict": "PASS", + "scores": { + "workstream_recall": "6/6", + "risk_recall": "4/6", + "decision_recall": "4/4", + "milestone_recall": "4/5", + "sections_present": [ + "workstream", + "risk", + "decision", + "milestone" + ], + "word_count": 321 + }, + "thresholds": { + "min_workstreams": 4, + "min_risks": 3, + "min_decisions": 3, + "min_milestones": 3, + "max_word_count": 700 + }, + "details": { + "workstreams": { + "WS1": { + "matched": true, + "keyword": "query layer" + }, + "WS2": { + "matched": true, + "keyword": "dashboard refresh" + }, + "WS3": { + "matched": true, + "keyword": "40-panel" + }, + "WS4": { + "matched": true, + "keyword": "access control" + }, + "WS5": { + "matched": true, + "keyword": "maevia" + }, + "WS6": { + "matched": true, + "keyword": "legal" + } + }, + "risks": { + "R1": { + "matched": false, + "keyword": null + }, + "R2": { + "matched": true, + "keyword": "correction-overlay-semantic-equivalent" + }, + "R3": { + "matched": true, + "keyword": "correction-overlay-semantic-equivalent" + }, + "R4": { + "matched": true, + "keyword": "40-panel" + }, + "R5": { + "matched": true, + "keyword": "custom branding" + }, + "R6": { + "matched": false, + "keyword": null + } + }, + "decisions": { + "D1_branding": { + "matched": true, + "keyword": "branding deferred" + }, + "D2_panel_limit": { + "matched": true, + "keyword": "40-panel" + }, + "D3_mobile": { + "matched": true, + "keyword": "correction-overlay-semantic-equivalent" + }, + "D4_option_b": { + "matched": true, + "keyword": "correction-overlay-semantic-equivalent" + } + }, + "milestones": { + "M1": { + "matched": true, + "keyword": "mid-may" + }, + "M2": { + "matched": true, + "keyword": "mid-may" + }, + "M3": { + "matched": false, + "keyword": null + }, + "M4": { + "matched": true, + "keyword": "v1.1" + }, + "M5": { + "matched": true, + "keyword": "deferred" + } + } + }, + "hand_rating_placeholders": { + "structure_quality_1to5": null, + "fabrication_count": null, + "owner_accuracy_0to6": null, + "rater_notes": null + } + } + }, + { + "run_name": "p3_pm_gemma4-31b-q4_v2", + "raw_grade_sha256": "5e98c1230f111ddb4d1a15ae7063c80590fbee19b72661c4b56e9c94f0b12352", + "workspace_archive_sha256": "6190888ab5900349d69d7d443477eb162cb5174acc340c36c5489a95b9a591c8", + "status_report_source": "workspace_final.tar.gz:./status_report.md", + "status_report_sha256": "57d5b455eb57566d5d616c0371fb5818e6fe643f1262799089be12f4b45e6588", + "raw_verdict": "FAIL", + "corrected_verdict": "PASS", + "changes": [ + { + "item": "R2", + "category": "risks", + "reason": "Maevia GA-to-private-beta pushback is expressed without one contiguous legacy keyword", + "pattern": "\\bmaevia\\b.{0,240}\\b(?:push(?:ed)?[ -]?back|fallout|promis(?:e|ed)\\s+ga|private[ -]?beta)\\b" + }, + { + "item": "R3", + "category": "risks", + "reason": "legal/private-beta contract delay uses an equivalent non-contracted phrase", + "pattern": "\\blegal\\b.{0,200}\\b(?:has\\s+not|not\\s+yet\\s+responded|sign[ -]?off|approval|unresponsive|silent|delay|contract)\\b" + }, + { + "item": "D4_option_b", + "category": "decisions", + "reason": "private beta followed by the selected-customer count is Option B", + "pattern": "\\bprivate[ -]?beta\\b.{0,120}\\b(?:3\\s*[-\u2013]\\s*5|selected\\s+customers?)\\b" + } + ], + "corrected_grade": { + "task": "project_mgmt", + "verdict": "PASS", + "scores": { + "workstream_recall": "6/6", + "risk_recall": "4/6", + "decision_recall": "4/4", + "milestone_recall": "5/5", + "sections_present": [ + "workstream", + "risk", + "decision", + "milestone" + ], + "word_count": 313 + }, + "thresholds": { + "min_workstreams": 4, + "min_risks": 3, + "min_decisions": 3, + "min_milestones": 3, + "max_word_count": 700 + }, + "details": { + "workstreams": { + "WS1": { + "matched": true, + "keyword": "query layer" + }, + "WS2": { + "matched": true, + "keyword": "mobile-responsive" + }, + "WS3": { + "matched": true, + "keyword": "40-panel" + }, + "WS4": { + "matched": true, + "keyword": "access-control" + }, + "WS5": { + "matched": true, + "keyword": "maevia" + }, + "WS6": { + "matched": true, + "keyword": "legal" + } + }, + "risks": { + "R1": { + "matched": false, + "keyword": null + }, + "R2": { + "matched": true, + "keyword": "correction-overlay-semantic-equivalent" + }, + "R3": { + "matched": true, + "keyword": "correction-overlay-semantic-equivalent" + }, + "R4": { + "matched": true, + "keyword": "40-panel" + }, + "R5": { + "matched": true, + "keyword": "custom branding" + }, + "R6": { + "matched": false, + "keyword": null + } + }, + "decisions": { + "D1_branding": { + "matched": true, + "keyword": "cut custom branding" + }, + "D2_panel_limit": { + "matched": true, + "keyword": "40-panel" + }, + "D3_mobile": { + "matched": true, + "keyword": "mobile responsive" + }, + "D4_option_b": { + "matched": true, + "keyword": "correction-overlay-semantic-equivalent" + } + }, + "milestones": { + "M1": { + "matched": true, + "keyword": "mid-may" + }, + "M2": { + "matched": true, + "keyword": "mid-may" + }, + "M3": { + "matched": true, + "keyword": "mid-july" + }, + "M4": { + "matched": true, + "keyword": "v1.1" + }, + "M5": { + "matched": true, + "keyword": "deferred" + } + } + }, + "hand_rating_placeholders": { + "structure_quality_1to5": null, + "fabrication_count": null, + "owner_accuracy_0to6": null, + "rater_notes": null + } + } + }, + { + "run_name": "p3_pm_gemma4-31b-q4_v3", + "raw_grade_sha256": "be7951b5c0b3770b9ee61e8f3bd5042dfc0fa486d2be1086b97314e4a8976ca3", + "workspace_archive_sha256": "fac5112297764b16bb742a38fcecc8617524ef13ebe869170e049842e63a4645", + "status_report_source": "workspace_final.tar.gz:./status_report.md", + "status_report_sha256": "1df31af7e0ef549e4a2db3fdae74f8aba184c79d530d810baae0c09215dac7b3", + "raw_verdict": "FAIL", + "corrected_verdict": "PASS", + "changes": [ + { + "item": "R2", + "category": "risks", + "reason": "Maevia GA-to-private-beta pushback is expressed without one contiguous legacy keyword", + "pattern": "\\bmaevia\\b.{0,240}\\b(?:push(?:ed)?[ -]?back|fallout|promis(?:e|ed)\\s+ga|private[ -]?beta)\\b" + }, + { + "item": "R3", + "category": "risks", + "reason": "legal/private-beta contract delay uses an equivalent non-contracted phrase", + "pattern": "\\blegal\\b.{0,200}\\b(?:has\\s+not|not\\s+yet\\s+responded|sign[ -]?off|approval|unresponsive|silent|delay|contract)\\b" + }, + { + "item": "D3_mobile", + "category": "decisions", + "reason": "web-responsive V1 and native V2 uses a hyphen the legacy keyword omits", + "pattern": "\\bweb[ -]?responsive\\b.{0,160}\\bnative\\b.{0,100}\\bv2\\b" + }, + { + "item": "D4_option_b", + "category": "decisions", + "reason": "private beta followed by the selected-customer count is Option B", + "pattern": "\\bprivate[ -]?beta\\b.{0,120}\\b(?:3\\s*[-\u2013]\\s*5|selected\\s+customers?)\\b" + } + ], + "corrected_grade": { + "task": "project_mgmt", + "verdict": "PASS", + "scores": { + "workstream_recall": "6/6", + "risk_recall": "4/6", + "decision_recall": "4/4", + "milestone_recall": "4/5", + "sections_present": [ + "workstream", + "risk", + "decision", + "milestone" + ], + "word_count": 303 + }, + "thresholds": { + "min_workstreams": 4, + "min_risks": 3, + "min_decisions": 3, + "min_milestones": 3, + "max_word_count": 700 + }, + "details": { + "workstreams": { + "WS1": { + "matched": true, + "keyword": "query-layer" + }, + "WS2": { + "matched": true, + "keyword": "dashboard refresh" + }, + "WS3": { + "matched": true, + "keyword": "architectural fix" + }, + "WS4": { + "matched": true, + "keyword": "access control" + }, + "WS5": { + "matched": true, + "keyword": "maevia" + }, + "WS6": { + "matched": true, + "keyword": "legal" + } + }, + "risks": { + "R1": { + "matched": false, + "keyword": null + }, + "R2": { + "matched": true, + "keyword": "correction-overlay-semantic-equivalent" + }, + "R3": { + "matched": true, + "keyword": "correction-overlay-semantic-equivalent" + }, + "R4": { + "matched": true, + "keyword": "40-panel" + }, + "R5": { + "matched": true, + "keyword": "custom-branding" + }, + "R6": { + "matched": false, + "keyword": null + } + }, + "decisions": { + "D1_branding": { + "matched": true, + "keyword": "branding cut" + }, + "D2_panel_limit": { + "matched": true, + "keyword": "40-panel" + }, + "D3_mobile": { + "matched": true, + "keyword": "correction-overlay-semantic-equivalent" + }, + "D4_option_b": { + "matched": true, + "keyword": "correction-overlay-semantic-equivalent" + } + }, + "milestones": { + "M1": { + "matched": true, + "keyword": "mid-may" + }, + "M2": { + "matched": true, + "keyword": "mid-may" + }, + "M3": { + "matched": false, + "keyword": null + }, + "M4": { + "matched": true, + "keyword": "v1.1" + }, + "M5": { + "matched": true, + "keyword": "deferred" + } + } + }, + "hand_rating_placeholders": { + "structure_quality_1to5": null, + "fabrication_count": null, + "owner_accuracy_0to6": null, + "rater_notes": null + } + } + }, + { + "run_name": "p3_pm_gemma4-31b-q4_v4", + "raw_grade_sha256": "013a8862905aa0eac7318b00edaa3a3e2f2ca542c4a329ef794b38486fd7a73b", + "workspace_archive_sha256": "2f376b7734e2927dc611ac896789acbb0fb16e066221f84604890fa138548b11", + "status_report_source": "workspace_final.tar.gz:./status_report.md", + "status_report_sha256": "02a5bf74ad575ac531f762cdd3803fc4ee30c72571b511136370a21d0e85d85b", + "raw_verdict": "FAIL", + "corrected_verdict": "PASS", + "changes": [ + { + "item": "R2", + "category": "risks", + "reason": "Maevia GA-to-private-beta pushback is expressed without one contiguous legacy keyword", + "pattern": "\\bmaevia\\b.{0,240}\\b(?:push(?:ed)?[ -]?back|fallout|promis(?:e|ed)\\s+ga|private[ -]?beta)\\b" + }, + { + "item": "R3", + "category": "risks", + "reason": "legal/private-beta contract delay uses an equivalent non-contracted phrase", + "pattern": "\\blegal\\b.{0,200}\\b(?:has\\s+not|not\\s+yet\\s+responded|sign[ -]?off|approval|unresponsive|silent|delay|contract)\\b" + }, + { + "item": "D4_option_b", + "category": "decisions", + "reason": "private beta followed by the selected-customer count is Option B", + "pattern": "\\bprivate[ -]?beta\\b.{0,120}\\b(?:3\\s*[-\u2013]\\s*5|selected\\s+customers?)\\b" + } + ], + "corrected_grade": { + "task": "project_mgmt", + "verdict": "PASS", + "scores": { + "workstream_recall": "6/6", + "risk_recall": "4/6", + "decision_recall": "4/4", + "milestone_recall": "4/5", + "sections_present": [ + "workstream", + "risk", + "decision", + "milestone" + ], + "word_count": 293 + }, + "thresholds": { + "min_workstreams": 4, + "min_risks": 3, + "min_decisions": 3, + "min_milestones": 3, + "max_word_count": 700 + }, + "details": { + "workstreams": { + "WS1": { + "matched": true, + "keyword": "query layer" + }, + "WS2": { + "matched": true, + "keyword": "dashboard refresh" + }, + "WS3": { + "matched": true, + "keyword": "architectural fix" + }, + "WS4": { + "matched": true, + "keyword": "access-control" + }, + "WS5": { + "matched": true, + "keyword": "maevia" + }, + "WS6": { + "matched": true, + "keyword": "legal" + } + }, + "risks": { + "R1": { + "matched": false, + "keyword": null + }, + "R2": { + "matched": true, + "keyword": "correction-overlay-semantic-equivalent" + }, + "R3": { + "matched": true, + "keyword": "correction-overlay-semantic-equivalent" + }, + "R4": { + "matched": true, + "keyword": "40-panel" + }, + "R5": { + "matched": true, + "keyword": "custom-branding" + }, + "R6": { + "matched": false, + "keyword": null + } + }, + "decisions": { + "D1_branding": { + "matched": true, + "keyword": "cut custom-branding" + }, + "D2_panel_limit": { + "matched": true, + "keyword": "40-panel" + }, + "D3_mobile": { + "matched": true, + "keyword": "mobile responsive" + }, + "D4_option_b": { + "matched": true, + "keyword": "correction-overlay-semantic-equivalent" + } + }, + "milestones": { + "M1": { + "matched": true, + "keyword": "mid-may" + }, + "M2": { + "matched": true, + "keyword": "mid-may" + }, + "M3": { + "matched": false, + "keyword": null + }, + "M4": { + "matched": true, + "keyword": "v1.1" + }, + "M5": { + "matched": true, + "keyword": "deferred" + } + } + }, + "hand_rating_placeholders": { + "structure_quality_1to5": null, + "fabrication_count": null, + "owner_accuracy_0to6": null, + "rater_notes": null + } + } + }, + { + "run_name": "p3_pm_gemma4-31b-q4_v5", + "raw_grade_sha256": "013a8862905aa0eac7318b00edaa3a3e2f2ca542c4a329ef794b38486fd7a73b", + "workspace_archive_sha256": "75cbdce99bebf45b3c96082686765bb2621fd862ea9dea6d76fc4505c008b127", + "status_report_source": "workspace_final.tar.gz:./status_report.md", + "status_report_sha256": "02a5bf74ad575ac531f762cdd3803fc4ee30c72571b511136370a21d0e85d85b", + "raw_verdict": "FAIL", + "corrected_verdict": "PASS", + "changes": [ + { + "item": "R2", + "category": "risks", + "reason": "Maevia GA-to-private-beta pushback is expressed without one contiguous legacy keyword", + "pattern": "\\bmaevia\\b.{0,240}\\b(?:push(?:ed)?[ -]?back|fallout|promis(?:e|ed)\\s+ga|private[ -]?beta)\\b" + }, + { + "item": "R3", + "category": "risks", + "reason": "legal/private-beta contract delay uses an equivalent non-contracted phrase", + "pattern": "\\blegal\\b.{0,200}\\b(?:has\\s+not|not\\s+yet\\s+responded|sign[ -]?off|approval|unresponsive|silent|delay|contract)\\b" + }, + { + "item": "D4_option_b", + "category": "decisions", + "reason": "private beta followed by the selected-customer count is Option B", + "pattern": "\\bprivate[ -]?beta\\b.{0,120}\\b(?:3\\s*[-\u2013]\\s*5|selected\\s+customers?)\\b" + } + ], + "corrected_grade": { + "task": "project_mgmt", + "verdict": "PASS", + "scores": { + "workstream_recall": "6/6", + "risk_recall": "4/6", + "decision_recall": "4/4", + "milestone_recall": "4/5", + "sections_present": [ + "workstream", + "risk", + "decision", + "milestone" + ], + "word_count": 293 + }, + "thresholds": { + "min_workstreams": 4, + "min_risks": 3, + "min_decisions": 3, + "min_milestones": 3, + "max_word_count": 700 + }, + "details": { + "workstreams": { + "WS1": { + "matched": true, + "keyword": "query layer" + }, + "WS2": { + "matched": true, + "keyword": "dashboard refresh" + }, + "WS3": { + "matched": true, + "keyword": "architectural fix" + }, + "WS4": { + "matched": true, + "keyword": "access-control" + }, + "WS5": { + "matched": true, + "keyword": "maevia" + }, + "WS6": { + "matched": true, + "keyword": "legal" + } + }, + "risks": { + "R1": { + "matched": false, + "keyword": null + }, + "R2": { + "matched": true, + "keyword": "correction-overlay-semantic-equivalent" + }, + "R3": { + "matched": true, + "keyword": "correction-overlay-semantic-equivalent" + }, + "R4": { + "matched": true, + "keyword": "40-panel" + }, + "R5": { + "matched": true, + "keyword": "custom-branding" + }, + "R6": { + "matched": false, + "keyword": null + } + }, + "decisions": { + "D1_branding": { + "matched": true, + "keyword": "cut custom-branding" + }, + "D2_panel_limit": { + "matched": true, + "keyword": "40-panel" + }, + "D3_mobile": { + "matched": true, + "keyword": "mobile responsive" + }, + "D4_option_b": { + "matched": true, + "keyword": "correction-overlay-semantic-equivalent" + } + }, + "milestones": { + "M1": { + "matched": true, + "keyword": "mid-may" + }, + "M2": { + "matched": true, + "keyword": "mid-may" + }, + "M3": { + "matched": false, + "keyword": null + }, + "M4": { + "matched": true, + "keyword": "v1.1" + }, + "M5": { + "matched": true, + "keyword": "deferred" + } + } + }, + "hand_rating_placeholders": { + "structure_quality_1to5": null, + "fabrication_count": null, + "owner_accuracy_0to6": null, + "rater_notes": null + } + } + }, + { + "run_name": "p3_pm_gemma4-31b-q4_v6", + "raw_grade_sha256": "5e98c1230f111ddb4d1a15ae7063c80590fbee19b72661c4b56e9c94f0b12352", + "workspace_archive_sha256": "42ebcf94466a424be2fd5176c6d1bb110004371a0c5462a3c059de3aaa900f54", + "status_report_source": "workspace_final.tar.gz:./status_report.md", + "status_report_sha256": "57d5b455eb57566d5d616c0371fb5818e6fe643f1262799089be12f4b45e6588", + "raw_verdict": "FAIL", + "corrected_verdict": "PASS", + "changes": [ + { + "item": "R2", + "category": "risks", + "reason": "Maevia GA-to-private-beta pushback is expressed without one contiguous legacy keyword", + "pattern": "\\bmaevia\\b.{0,240}\\b(?:push(?:ed)?[ -]?back|fallout|promis(?:e|ed)\\s+ga|private[ -]?beta)\\b" + }, + { + "item": "R3", + "category": "risks", + "reason": "legal/private-beta contract delay uses an equivalent non-contracted phrase", + "pattern": "\\blegal\\b.{0,200}\\b(?:has\\s+not|not\\s+yet\\s+responded|sign[ -]?off|approval|unresponsive|silent|delay|contract)\\b" + }, + { + "item": "D4_option_b", + "category": "decisions", + "reason": "private beta followed by the selected-customer count is Option B", + "pattern": "\\bprivate[ -]?beta\\b.{0,120}\\b(?:3\\s*[-\u2013]\\s*5|selected\\s+customers?)\\b" + } + ], + "corrected_grade": { + "task": "project_mgmt", + "verdict": "PASS", + "scores": { + "workstream_recall": "6/6", + "risk_recall": "4/6", + "decision_recall": "4/4", + "milestone_recall": "5/5", + "sections_present": [ + "workstream", + "risk", + "decision", + "milestone" + ], + "word_count": 313 + }, + "thresholds": { + "min_workstreams": 4, + "min_risks": 3, + "min_decisions": 3, + "min_milestones": 3, + "max_word_count": 700 + }, + "details": { + "workstreams": { + "WS1": { + "matched": true, + "keyword": "query layer" + }, + "WS2": { + "matched": true, + "keyword": "mobile-responsive" + }, + "WS3": { + "matched": true, + "keyword": "40-panel" + }, + "WS4": { + "matched": true, + "keyword": "access-control" + }, + "WS5": { + "matched": true, + "keyword": "maevia" + }, + "WS6": { + "matched": true, + "keyword": "legal" + } + }, + "risks": { + "R1": { + "matched": false, + "keyword": null + }, + "R2": { + "matched": true, + "keyword": "correction-overlay-semantic-equivalent" + }, + "R3": { + "matched": true, + "keyword": "correction-overlay-semantic-equivalent" + }, + "R4": { + "matched": true, + "keyword": "40-panel" + }, + "R5": { + "matched": true, + "keyword": "custom branding" + }, + "R6": { + "matched": false, + "keyword": null + } + }, + "decisions": { + "D1_branding": { + "matched": true, + "keyword": "cut custom branding" + }, + "D2_panel_limit": { + "matched": true, + "keyword": "40-panel" + }, + "D3_mobile": { + "matched": true, + "keyword": "mobile responsive" + }, + "D4_option_b": { + "matched": true, + "keyword": "correction-overlay-semantic-equivalent" + } + }, + "milestones": { + "M1": { + "matched": true, + "keyword": "mid-may" + }, + "M2": { + "matched": true, + "keyword": "mid-may" + }, + "M3": { + "matched": true, + "keyword": "mid-july" + }, + "M4": { + "matched": true, + "keyword": "v1.1" + }, + "M5": { + "matched": true, + "keyword": "deferred" + } + } + }, + "hand_rating_placeholders": { + "structure_quality_1to5": null, + "fabrication_count": null, + "owner_accuracy_0to6": null, + "rater_notes": null + } + } + }, + { + "run_name": "p3_pm_gemma4-31b-q4_v7", + "raw_grade_sha256": "5e98c1230f111ddb4d1a15ae7063c80590fbee19b72661c4b56e9c94f0b12352", + "workspace_archive_sha256": "191675310a0ca94b141e1b880425ba5476b1c71610050d238ce87fbe2b1ba7e4", + "status_report_source": "workspace_final.tar.gz:./status_report.md", + "status_report_sha256": "57d5b455eb57566d5d616c0371fb5818e6fe643f1262799089be12f4b45e6588", + "raw_verdict": "FAIL", + "corrected_verdict": "PASS", + "changes": [ + { + "item": "R2", + "category": "risks", + "reason": "Maevia GA-to-private-beta pushback is expressed without one contiguous legacy keyword", + "pattern": "\\bmaevia\\b.{0,240}\\b(?:push(?:ed)?[ -]?back|fallout|promis(?:e|ed)\\s+ga|private[ -]?beta)\\b" + }, + { + "item": "R3", + "category": "risks", + "reason": "legal/private-beta contract delay uses an equivalent non-contracted phrase", + "pattern": "\\blegal\\b.{0,200}\\b(?:has\\s+not|not\\s+yet\\s+responded|sign[ -]?off|approval|unresponsive|silent|delay|contract)\\b" + }, + { + "item": "D4_option_b", + "category": "decisions", + "reason": "private beta followed by the selected-customer count is Option B", + "pattern": "\\bprivate[ -]?beta\\b.{0,120}\\b(?:3\\s*[-\u2013]\\s*5|selected\\s+customers?)\\b" + } + ], + "corrected_grade": { + "task": "project_mgmt", + "verdict": "PASS", + "scores": { + "workstream_recall": "6/6", + "risk_recall": "4/6", + "decision_recall": "4/4", + "milestone_recall": "5/5", + "sections_present": [ + "workstream", + "risk", + "decision", + "milestone" + ], + "word_count": 313 + }, + "thresholds": { + "min_workstreams": 4, + "min_risks": 3, + "min_decisions": 3, + "min_milestones": 3, + "max_word_count": 700 + }, + "details": { + "workstreams": { + "WS1": { + "matched": true, + "keyword": "query layer" + }, + "WS2": { + "matched": true, + "keyword": "mobile-responsive" + }, + "WS3": { + "matched": true, + "keyword": "40-panel" + }, + "WS4": { + "matched": true, + "keyword": "access-control" + }, + "WS5": { + "matched": true, + "keyword": "maevia" + }, + "WS6": { + "matched": true, + "keyword": "legal" + } + }, + "risks": { + "R1": { + "matched": false, + "keyword": null + }, + "R2": { + "matched": true, + "keyword": "correction-overlay-semantic-equivalent" + }, + "R3": { + "matched": true, + "keyword": "correction-overlay-semantic-equivalent" + }, + "R4": { + "matched": true, + "keyword": "40-panel" + }, + "R5": { + "matched": true, + "keyword": "custom branding" + }, + "R6": { + "matched": false, + "keyword": null + } + }, + "decisions": { + "D1_branding": { + "matched": true, + "keyword": "cut custom branding" + }, + "D2_panel_limit": { + "matched": true, + "keyword": "40-panel" + }, + "D3_mobile": { + "matched": true, + "keyword": "mobile responsive" + }, + "D4_option_b": { + "matched": true, + "keyword": "correction-overlay-semantic-equivalent" + } + }, + "milestones": { + "M1": { + "matched": true, + "keyword": "mid-may" + }, + "M2": { + "matched": true, + "keyword": "mid-may" + }, + "M3": { + "matched": true, + "keyword": "mid-july" + }, + "M4": { + "matched": true, + "keyword": "v1.1" + }, + "M5": { + "matched": true, + "keyword": "deferred" + } + } + }, + "hand_rating_placeholders": { + "structure_quality_1to5": null, + "fabrication_count": null, + "owner_accuracy_0to6": null, + "rater_notes": null + } + } + }, + { + "run_name": "p3_pm_gemma4-31b-q4_v8", + "raw_grade_sha256": "5e5d3ca47674ec81eb8de446669c70129aff454157ba42e6bc018d6a13e64818", + "workspace_archive_sha256": "261171d6c3ef415e0486913dcf132b32162e4f84de90128b64cf2b344e41a4a5", + "status_report_source": "workspace_final.tar.gz:./status_report.md", + "status_report_sha256": "e42df5bb7e7cab0441618163adfac53f545af677fcd921524742946aaec3fb3f", + "raw_verdict": "FAIL", + "corrected_verdict": "PASS", + "changes": [ + { + "item": "R2", + "category": "risks", + "reason": "Maevia GA-to-private-beta pushback is expressed without one contiguous legacy keyword", + "pattern": "\\bmaevia\\b.{0,240}\\b(?:push(?:ed)?[ -]?back|fallout|promis(?:e|ed)\\s+ga|private[ -]?beta)\\b" + }, + { + "item": "R3", + "category": "risks", + "reason": "legal/private-beta contract delay uses an equivalent non-contracted phrase", + "pattern": "\\blegal\\b.{0,200}\\b(?:has\\s+not|not\\s+yet\\s+responded|sign[ -]?off|approval|unresponsive|silent|delay|contract)\\b" + }, + { + "item": "D3_mobile", + "category": "decisions", + "reason": "web-responsive V1 and native V2 uses a hyphen the legacy keyword omits", + "pattern": "\\bweb[ -]?responsive\\b.{0,160}\\bnative\\b.{0,100}\\bv2\\b" + }, + { + "item": "D4_option_b", + "category": "decisions", + "reason": "private beta followed by the selected-customer count is Option B", + "pattern": "\\bprivate[ -]?beta\\b.{0,120}\\b(?:3\\s*[-\u2013]\\s*5|selected\\s+customers?)\\b" + } + ], + "corrected_grade": { + "task": "project_mgmt", + "verdict": "PASS", + "scores": { + "workstream_recall": "6/6", + "risk_recall": "4/6", + "decision_recall": "4/4", + "milestone_recall": "5/5", + "sections_present": [ + "workstream", + "risk", + "decision", + "milestone" + ], + "word_count": 315 + }, + "thresholds": { + "min_workstreams": 4, + "min_risks": 3, + "min_decisions": 3, + "min_milestones": 3, + "max_word_count": 700 + }, + "details": { + "workstreams": { + "WS1": { + "matched": true, + "keyword": "query layer" + }, + "WS2": { + "matched": true, + "keyword": "mobile-responsive" + }, + "WS3": { + "matched": true, + "keyword": "architectural fix" + }, + "WS4": { + "matched": true, + "keyword": "access control" + }, + "WS5": { + "matched": true, + "keyword": "maevia" + }, + "WS6": { + "matched": true, + "keyword": "legal" + } + }, + "risks": { + "R1": { + "matched": false, + "keyword": null + }, + "R2": { + "matched": true, + "keyword": "correction-overlay-semantic-equivalent" + }, + "R3": { + "matched": true, + "keyword": "correction-overlay-semantic-equivalent" + }, + "R4": { + "matched": true, + "keyword": "40-panel" + }, + "R5": { + "matched": true, + "keyword": "custom-branding" + }, + "R6": { + "matched": false, + "keyword": null + } + }, + "decisions": { + "D1_branding": { + "matched": true, + "keyword": "cut custom-branding" + }, + "D2_panel_limit": { + "matched": true, + "keyword": "40-panel" + }, + "D3_mobile": { + "matched": true, + "keyword": "correction-overlay-semantic-equivalent" + }, + "D4_option_b": { + "matched": true, + "keyword": "correction-overlay-semantic-equivalent" + } + }, + "milestones": { + "M1": { + "matched": true, + "keyword": "mid-may" + }, + "M2": { + "matched": true, + "keyword": "mid-may" + }, + "M3": { + "matched": true, + "keyword": "mid-july" + }, + "M4": { + "matched": true, + "keyword": "v1.1" + }, + "M5": { + "matched": true, + "keyword": "deferred" + } + } + }, + "hand_rating_placeholders": { + "structure_quality_1to5": null, + "fabrication_count": null, + "owner_accuracy_0to6": null, + "rater_notes": null + } + } + }, + { + "run_name": "p3_pm_gemma4-31b-q4_v9", + "raw_grade_sha256": "8defbbde59eb9e22c6aa51b5bd69e93bcf5a548d56021d8f45b5fada421c1021", + "workspace_archive_sha256": "326c41dfdf2140156ad57262cc39fa5f85e034c0523db2ebd47834a129226762", + "status_report_source": "workspace_final.tar.gz:./status_report.md", + "status_report_sha256": "1ced2be945020ec4a7dbf92012f8651cb93a178d0b9c1b67c237d0f75a270a9c", + "raw_verdict": "FAIL", + "corrected_verdict": "PASS", + "changes": [ + { + "item": "R2", + "category": "risks", + "reason": "Maevia GA-to-private-beta pushback is expressed without one contiguous legacy keyword", + "pattern": "\\bmaevia\\b.{0,240}\\b(?:push(?:ed)?[ -]?back|fallout|promis(?:e|ed)\\s+ga|private[ -]?beta)\\b" + }, + { + "item": "R3", + "category": "risks", + "reason": "legal/private-beta contract delay uses an equivalent non-contracted phrase", + "pattern": "\\blegal\\b.{0,200}\\b(?:has\\s+not|not\\s+yet\\s+responded|sign[ -]?off|approval|unresponsive|silent|delay|contract)\\b" + }, + { + "item": "D4_option_b", + "category": "decisions", + "reason": "private beta followed by the selected-customer count is Option B", + "pattern": "\\bprivate[ -]?beta\\b.{0,120}\\b(?:3\\s*[-\u2013]\\s*5|selected\\s+customers?)\\b" + } + ], + "corrected_grade": { + "task": "project_mgmt", + "verdict": "PASS", + "scores": { + "workstream_recall": "6/6", + "risk_recall": "4/6", + "decision_recall": "4/4", + "milestone_recall": "5/5", + "sections_present": [ + "workstream", + "risk", + "decision", + "milestone" + ], + "word_count": 268 + }, + "thresholds": { + "min_workstreams": 4, + "min_risks": 3, + "min_decisions": 3, + "min_milestones": 3, + "max_word_count": 700 + }, + "details": { + "workstreams": { + "WS1": { + "matched": true, + "keyword": "query layer" + }, + "WS2": { + "matched": true, + "keyword": "mobile responsive" + }, + "WS3": { + "matched": true, + "keyword": "40-panel" + }, + "WS4": { + "matched": true, + "keyword": "access control" + }, + "WS5": { + "matched": true, + "keyword": "maevia" + }, + "WS6": { + "matched": true, + "keyword": "legal" + } + }, + "risks": { + "R1": { + "matched": false, + "keyword": null + }, + "R2": { + "matched": true, + "keyword": "correction-overlay-semantic-equivalent" + }, + "R3": { + "matched": true, + "keyword": "correction-overlay-semantic-equivalent" + }, + "R4": { + "matched": true, + "keyword": "40-panel" + }, + "R5": { + "matched": true, + "keyword": "custom-branding" + }, + "R6": { + "matched": false, + "keyword": null + } + }, + "decisions": { + "D1_branding": { + "matched": true, + "keyword": "cut custom-branding" + }, + "D2_panel_limit": { + "matched": true, + "keyword": "40-panel" + }, + "D3_mobile": { + "matched": true, + "keyword": "mobile responsive" + }, + "D4_option_b": { + "matched": true, + "keyword": "correction-overlay-semantic-equivalent" + } + }, + "milestones": { + "M1": { + "matched": true, + "keyword": "mid-may" + }, + "M2": { + "matched": true, + "keyword": "mid-may" + }, + "M3": { + "matched": true, + "keyword": "v1 launch" + }, + "M4": { + "matched": true, + "keyword": "v1.1" + }, + "M5": { + "matched": true, + "keyword": "deferred" + } + } + }, + "hand_rating_placeholders": { + "structure_quality_1to5": null, + "fabrication_count": null, + "owner_accuracy_0to6": null, + "rater_notes": null + } + } + }, + { + "run_name": "p3_pm_gemma4-31b-q4_v10", + "raw_grade_sha256": "8defbbde59eb9e22c6aa51b5bd69e93bcf5a548d56021d8f45b5fada421c1021", + "workspace_archive_sha256": "8b11367b2047223d189f3f7c731f6735558ae640d3d297d47d8e5a6f7bf31db5", + "status_report_source": "workspace_final.tar.gz:./status_report.md", + "status_report_sha256": "1ced2be945020ec4a7dbf92012f8651cb93a178d0b9c1b67c237d0f75a270a9c", + "raw_verdict": "FAIL", + "corrected_verdict": "PASS", + "changes": [ + { + "item": "R2", + "category": "risks", + "reason": "Maevia GA-to-private-beta pushback is expressed without one contiguous legacy keyword", + "pattern": "\\bmaevia\\b.{0,240}\\b(?:push(?:ed)?[ -]?back|fallout|promis(?:e|ed)\\s+ga|private[ -]?beta)\\b" + }, + { + "item": "R3", + "category": "risks", + "reason": "legal/private-beta contract delay uses an equivalent non-contracted phrase", + "pattern": "\\blegal\\b.{0,200}\\b(?:has\\s+not|not\\s+yet\\s+responded|sign[ -]?off|approval|unresponsive|silent|delay|contract)\\b" + }, + { + "item": "D4_option_b", + "category": "decisions", + "reason": "private beta followed by the selected-customer count is Option B", + "pattern": "\\bprivate[ -]?beta\\b.{0,120}\\b(?:3\\s*[-\u2013]\\s*5|selected\\s+customers?)\\b" + } + ], + "corrected_grade": { + "task": "project_mgmt", + "verdict": "PASS", + "scores": { + "workstream_recall": "6/6", + "risk_recall": "4/6", + "decision_recall": "4/4", + "milestone_recall": "5/5", + "sections_present": [ + "workstream", + "risk", + "decision", + "milestone" + ], + "word_count": 268 + }, + "thresholds": { + "min_workstreams": 4, + "min_risks": 3, + "min_decisions": 3, + "min_milestones": 3, + "max_word_count": 700 + }, + "details": { + "workstreams": { + "WS1": { + "matched": true, + "keyword": "query layer" + }, + "WS2": { + "matched": true, + "keyword": "mobile responsive" + }, + "WS3": { + "matched": true, + "keyword": "40-panel" + }, + "WS4": { + "matched": true, + "keyword": "access control" + }, + "WS5": { + "matched": true, + "keyword": "maevia" + }, + "WS6": { + "matched": true, + "keyword": "legal" + } + }, + "risks": { + "R1": { + "matched": false, + "keyword": null + }, + "R2": { + "matched": true, + "keyword": "correction-overlay-semantic-equivalent" + }, + "R3": { + "matched": true, + "keyword": "correction-overlay-semantic-equivalent" + }, + "R4": { + "matched": true, + "keyword": "40-panel" + }, + "R5": { + "matched": true, + "keyword": "custom-branding" + }, + "R6": { + "matched": false, + "keyword": null + } + }, + "decisions": { + "D1_branding": { + "matched": true, + "keyword": "cut custom-branding" + }, + "D2_panel_limit": { + "matched": true, + "keyword": "40-panel" + }, + "D3_mobile": { + "matched": true, + "keyword": "mobile responsive" + }, + "D4_option_b": { + "matched": true, + "keyword": "correction-overlay-semantic-equivalent" + } + }, + "milestones": { + "M1": { + "matched": true, + "keyword": "mid-may" + }, + "M2": { + "matched": true, + "keyword": "mid-may" + }, + "M3": { + "matched": true, + "keyword": "v1 launch" + }, + "M4": { + "matched": true, + "keyword": "v1.1" + }, + "M5": { + "matched": true, + "keyword": "deferred" + } + } + }, + "hand_rating_placeholders": { + "structure_quality_1to5": null, + "fabrication_count": null, + "owner_accuracy_0to6": null, + "rater_notes": null + } + } + } + ], + "errors": [], + "passed": true +} diff --git a/benchmarks/gemma4-31b-q4/gemma4-canonical-n10-scorecard.json b/benchmarks/gemma4-31b-q4/gemma4-canonical-n10-scorecard.json new file mode 100644 index 00000000..73fc6020 --- /dev/null +++ b/benchmarks/gemma4-31b-q4/gemma4-canonical-n10-scorecard.json @@ -0,0 +1,7563 @@ +{ + "schema_version": 1, + "generated_at": "2026-08-02T05:41:46.936883+00:00", + "label": "gemma4-31b-q4", + "target_n": 10, + "passed": true, + "errors": [], + "aggregate": { + "runs": 120, + "completed": 120, + "normal_completed": 120, + "terminal_outcomes": 0, + "graded": 120, + "scored_outcomes": 120, + "raw_passes": 89, + "raw_pass_rate": 0.741667, + "verdicts": { + "FAIL": 27, + "MISSING_OUTPUT": 4, + "PASS": 79, + "STRUCTURAL_PASS": 10 + }, + "quality_outcomes": { + "FAIL": 27, + "MISSING_OUTPUT": 4, + "PASS": 79, + "STRUCTURAL_PASS": 10 + }, + "finish_reasons": { + "done_signal": 116, + "model_stopped": 4 + }, + "wall_s": { + "sum": 22835.9, + "median": 113.05, + "max": 803.6 + }, + "completion_tokens": { + "sum": 978103, + "median": 5566.5, + "max": 30030 + }, + "model_call_completion_tps": { + "median": 55.85, + "min": 19.7, + "max": 61.9 + }, + "telemetry": { + "runs": 118, + "coverage_mean": 0.947183, + "active_gpu_mean_power_w_mean": 458.331, + "active_gpu_mean_sm_util_pct_mean": 86.054, + "max_temp_c": 87.0 + }, + "per_task": { + "p1_bugfix": { + "runs": 10, + "graded": 10, + "scored_outcomes": 10, + "raw_passes": 4, + "verdicts": { + "FAIL": 6, + "PASS": 4 + }, + "quality_outcomes": { + "FAIL": 6, + "PASS": 4 + }, + "finish_reasons": { + "done_signal": 10 + } + }, + "p1_testwrite": { + "runs": 10, + "graded": 10, + "scored_outcomes": 10, + "raw_passes": 8, + "verdicts": { + "FAIL": 2, + "PASS": 8 + }, + "quality_outcomes": { + "FAIL": 2, + "PASS": 8 + }, + "finish_reasons": { + "done_signal": 10 + } + }, + "p1_refactor": { + "runs": 10, + "graded": 10, + "scored_outcomes": 10, + "raw_passes": 7, + "verdicts": { + "FAIL": 3, + "PASS": 7 + }, + "quality_outcomes": { + "FAIL": 3, + "PASS": 7 + }, + "finish_reasons": { + "done_signal": 10 + } + }, + "p2_extract": { + "runs": 10, + "graded": 10, + "scored_outcomes": 10, + "raw_passes": 10, + "verdicts": { + "PASS": 10 + }, + "quality_outcomes": { + "PASS": 10 + }, + "finish_reasons": { + "done_signal": 10 + } + }, + "p2_ci": { + "runs": 10, + "graded": 10, + "scored_outcomes": 10, + "raw_passes": 10, + "verdicts": { + "PASS": 10 + }, + "quality_outcomes": { + "PASS": 10 + }, + "finish_reasons": { + "done_signal": 10 + } + }, + "p2_hallucination": { + "runs": 10, + "graded": 10, + "scored_outcomes": 10, + "raw_passes": 6, + "verdicts": { + "MISSING_OUTPUT": 4, + "PASS": 6 + }, + "quality_outcomes": { + "MISSING_OUTPUT": 4, + "PASS": 6 + }, + "finish_reasons": { + "done_signal": 6, + "model_stopped": 4 + } + }, + "p2_triage": { + "runs": 10, + "graded": 10, + "scored_outcomes": 10, + "raw_passes": 8, + "verdicts": { + "FAIL": 2, + "PASS": 8 + }, + "quality_outcomes": { + "FAIL": 2, + "PASS": 8 + }, + "finish_reasons": { + "done_signal": 10 + } + }, + "p3_doc": { + "runs": 10, + "graded": 10, + "scored_outcomes": 10, + "raw_passes": 10, + "verdicts": { + "PASS": 10 + }, + "quality_outcomes": { + "PASS": 10 + }, + "finish_reasons": { + "done_signal": 10 + } + }, + "p3_business": { + "runs": 10, + "graded": 10, + "scored_outcomes": 10, + "raw_passes": 6, + "verdicts": { + "FAIL": 4, + "PASS": 6 + }, + "quality_outcomes": { + "FAIL": 4, + "PASS": 6 + }, + "finish_reasons": { + "done_signal": 10 + } + }, + "p3_market": { + "runs": 10, + "graded": 10, + "scored_outcomes": 10, + "raw_passes": 10, + "verdicts": { + "STRUCTURAL_PASS": 10 + }, + "quality_outcomes": { + "STRUCTURAL_PASS": 10 + }, + "finish_reasons": { + "done_signal": 10 + } + }, + "p3_writing": { + "runs": 10, + "graded": 10, + "scored_outcomes": 10, + "raw_passes": 10, + "verdicts": { + "PASS": 10 + }, + "quality_outcomes": { + "PASS": 10 + }, + "finish_reasons": { + "done_signal": 10 + } + }, + "p3_pm": { + "runs": 10, + "graded": 10, + "scored_outcomes": 10, + "raw_passes": 0, + "verdicts": { + "FAIL": 10 + }, + "quality_outcomes": { + "FAIL": 10 + }, + "finish_reasons": { + "done_signal": 10 + } + } + } + }, + "runs": [ + { + "run_name": "p1_bugfix_gemma4-31b-q4_v1", + "task": "p1_bugfix", + "replicate": 1, + "summary": { + "started_at": "2026-08-02T00:21:07.782391+00:00", + "ended_at": "2026-08-02T00:32:56.494903+00:00", + "elapsed_s": 708.7, + "iterations": 80, + "total_completion_tokens": 18475, + "total_prompt_tokens": 1583932, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Stabilized the `logalyzer` codebase by fixing critical bugs and performance bottlenecks, removing a security vulnerability, and increasing test coverage to 91%.\n\nKey accomplishments:\n- **Performance**: Fixed an O(n\u00b2) log loading bottleneck, reducing 50MB log load time from >300s to ~3s.\n- **Security**: Removed the `eval()`-based `expression_filter` and corresponding CLI argument to prevent remote code execution.\n- **Stability**: Fixed a missing month in date parsing, corrected URL regex matching, and normalized CLI input types.\n- **Modernization**: Updated the codebase for Python 3.11+ and replaced manual JSON string building with the `json` module.\n- **Quality**: Removed dead code, added comprehensive tests for all modules, and achieved 91% line coverage.\n- **Documentation**: Created a full audit trail including a triage report, before-after metrics, ADRs, research notes, and a detailed CHANGELOG.\n\nThe final repository is located at `/workspace/` and is tagged as `v0.3.2`." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p1_bugfix_gemma4-31b-q4_v1", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 708.7, + "iters": 80, + "tokens": { + "completion_total": 18475, + "prompt_total": 1583932, + "all_total": 1602407 + }, + "throughput": { + "completion_tps_avg": 47.6, + "model_call_wall_s": 388.0, + "tool_call_wall_s": 314.2, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 404.7, + "first_commit_s": 0.1, + "first_tag_s": 689.9 + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 20.92, + "memory_used_mib_at_start": 65197.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0984, + "cost_usd_upper_bound": 0.0128, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": null + }, + { + "run_name": "p1_bugfix_gemma4-31b-q4_v2", + "task": "p1_bugfix", + "replicate": 2, + "summary": { + "started_at": "2026-08-02T00:21:07.781614+00:00", + "ended_at": "2026-08-02T00:31:47.326458+00:00", + "elapsed_s": 639.5, + "iterations": 58, + "total_completion_tokens": 14733, + "total_prompt_tokens": 984362, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "I have successfully adopted and stabilized the `logalyzer` codebase. \n\n**Key Accomplishments:**\n- **Fixed Critical Performance Issue**: Identified and fixed an O(n\u00b2) bottleneck in `logalyzer.io.load` where list concatenation was used in a loop. This resulted in a speedup from >300s to ~3.2s for a 50MB log file (approx. 90x improvement).\n- **Resolved Bugs**: \n - Fixed `parse_timestamp` missing 'Jul' in its lookup table, which caused crashes for July logs.\n - Fixed `status_filter` to correctly handle string inputs from the CLI.\n - Fixed `url_regex_filter` to perform substring matches using `re.search` instead of `re.match`.\n- **Modernized Codebase**: \n - Updated `collections.Iterable` to `collections.abc.Iterable` for Python 3.11 compatibility.\n - Optimized `ip_allowlist_filter` by converting the allowlist to a `set`, reducing lookup time from O(N) to O(1) per entry.\n- **Improved Quality & Coverage**:\n - Increased test coverage from 34% to 85% by adding comprehensive tests for `cli`, `io`, `output`, and `utils` modules.\n - Removed dead code (`logalyzer/experimental.py`).\n - Cleaned up unused imports and followed PEP8 standards using `ruff`.\n- **Documentation & Audit**:\n - Provided a full triage of issues in `/audit/triage.md`.\n - Documented hard metrics in `/audit/before-after.md`.\n - Created Architecture Decision Records (ADRs) and research logs in `/architecture/`, `/decisions/`, and `/research/`.\n - Updated `README.md` to remove false claims and accurately describe the tool's capabilities.\n - Created a user-facing `CHANGELOG.md`.\n\nThe final repository is tagged `v0.3.0` and is ready for hand-off." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p1_bugfix_gemma4-31b-q4_v2", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 639.5, + "iters": 58, + "tokens": { + "completion_total": 14733, + "prompt_total": 984362, + "all_total": 999095 + }, + "throughput": { + "completion_tps_avg": 45.7, + "model_call_wall_s": 322.4, + "tool_call_wall_s": 312.4, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 329.9, + "first_commit_s": 6.1, + "first_tag_s": 605.9 + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 20.92, + "memory_used_mib_at_start": 65197.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0888, + "cost_usd_upper_bound": 0.0115, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": null + }, + { + "run_name": "p1_bugfix_gemma4-31b-q4_v3", + "task": "p1_bugfix", + "replicate": 3, + "summary": { + "started_at": "2026-08-02T00:45:34.628923+00:00", + "ended_at": "2026-08-02T00:55:58.294596+00:00", + "elapsed_s": 623.7, + "iterations": 62, + "total_completion_tokens": 14968, + "total_prompt_tokens": 1161167, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Adopted and stabilized the `logalyzer` codebase. \n\nKey accomplishments:\n1. **Bug Fixes**: Fixed critical timestamp parsing error (missing July) and filter logic bugs (status type mismatch, URL regex match vs search).\n2. **Modernization**: Updated `collections` imports to `collections.abc` for Python 3.10+ compatibility.\n3. **Performance**: Optimized `logalyzer.io.load` from O(n\u00b2) to O(n) by replacing quadratic string and list concatenation with a streaming generator, reducing 50MB log processing time from >300s to ~3.3s.\n4. **Security**: Eliminated a critical Remote Code Execution (RCE) vulnerability by removing the `eval()`-based `expression_filter` and corresponding CLI flag.\n5. **Code Cleanup**: Removed dead code including an unused experimental module and several legacy utility functions.\n6. **Testing**: Expanded the test suite to cover CLI, IO, and Output modules, increasing total line coverage from 24% to 91% and achieving a 100% pass rate.\n7. **Documentation**: Updated the README (removed false claims), created a detailed CHANGELOG, and provided full audit and research trails including ADRs and decision records.\n\nThe final repository is tagged as `v0.3.0`." + }, + "verdict": "FAIL", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p1_bugfix_gemma4-31b-q4_v3", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 623.7, + "iters": 62, + "tokens": { + "completion_total": 14968, + "prompt_total": 1161167, + "all_total": 1176135 + }, + "throughput": { + "completion_tps_avg": 48.9, + "model_call_wall_s": 305.9, + "tool_call_wall_s": 312.7, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 337.5, + "first_commit_s": 3.3, + "first_tag_s": 597.7 + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 87.25, + "memory_used_mib_at_start": 65453.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0866, + "cost_usd_upper_bound": 0.0113, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9941, + "mean_power_w": 249.36, + "mean_sm_util_pct": 44.38, + "max_temp_c": 76.0 + } + }, + { + "run_name": "p1_bugfix_gemma4-31b-q4_v4", + "task": "p1_bugfix", + "replicate": 4, + "summary": { + "started_at": "2026-08-02T03:07:11.446114+00:00", + "ended_at": "2026-08-02T03:18:44.783977+00:00", + "elapsed_s": 693.3, + "iterations": 64, + "total_completion_tokens": 18708, + "total_prompt_tokens": 1257972, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Adopted and stabilized the logalyzer codebase. Fixed critical bugs in parsing and filtering, resolved a severe O(n\u00b2) performance bottleneck in log loading (reducing 50MB file processing from >300s to ~3s), removed dangerous `eval()`-based filtering, and improved test coverage from 0% to 93%. All changes are documented in ADRs, triage reports, and a detailed before-after metric table. Final release tagged as v0.3.2." + }, + "verdict": "FAIL", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p1_bugfix_gemma4-31b-q4_v4", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 693.3, + "iters": 64, + "tokens": { + "completion_total": 18708, + "prompt_total": 1257972, + "all_total": 1276680 + }, + "throughput": { + "completion_tps_avg": 49.9, + "model_call_wall_s": 374.9, + "tool_call_wall_s": 313.1, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 318.7, + "first_commit_s": 327.4, + "first_tag_s": 674.1 + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 19.75, + "memory_used_mib_at_start": 65325.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0963, + "cost_usd_upper_bound": 0.0125, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9952, + "mean_power_w": 259.64, + "mean_sm_util_pct": 46.02, + "max_temp_c": 78.0 + } + }, + { + "run_name": "p1_bugfix_gemma4-31b-q4_v5", + "task": "p1_bugfix", + "replicate": 5, + "summary": { + "started_at": "2026-08-02T03:07:11.446267+00:00", + "ended_at": "2026-08-02T03:17:22.187081+00:00", + "elapsed_s": 610.7, + "iterations": 67, + "total_completion_tokens": 13721, + "total_prompt_tokens": 1168165, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Adopted and stabilized the `logalyzer` codebase.\n\n### Key Accomplishments:\n- **Fixed Critical Bugs**:\n - Resolved `ImportError` caused by `collections.Iterable` deprecation in Python 3.10+.\n - Fixed a `ValueError` in `parse_timestamp` caused by a missing \"Jul\" in `MONTH_NAMES`.\n - Fixed `status_filter` and `url_regex_filter` to correctly handle CLI string input and substring matches.\n- **Performance Optimization**:\n - Rewrote `logalyzer.io.load` to eliminate O(n\u00b2) complexity.\n - **Result**: Benchmark time for a 50MB log reduced from >300s (timeout) to ~3.23s.\n- **Codebase Cleanup**:\n - Removed unused `experimental.py` module and legacy utility functions.\n - Updated README to remove false claims about \"tail mode\" and CIDR support.\n- **Quality Improvements**:\n - Increased test coverage from 0% (broken) to 90%.\n - Added comprehensive unit and integration tests for all core modules.\n - Eliminated all Ruff static analysis issues.\n- **Documentation**:\n - Created full audit trail: `triage.md`, `before-after.md`.\n - Documented architectural decisions and research (ADRs, Dead-ends, Questions).\n - Updated `CHANGELOG.md` and `README.md`.\n\nThe final repository is tagged as `v0.3.0` and is now stable, performant, and well-tested." + }, + "verdict": "FAIL", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p1_bugfix_gemma4-31b-q4_v5", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 610.7, + "iters": 67, + "tokens": { + "completion_total": 13721, + "prompt_total": 1168165, + "all_total": 1181886 + }, + "throughput": { + "completion_tps_avg": 47.9, + "model_call_wall_s": 286.4, + "tool_call_wall_s": 319.0, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 330.1, + "first_commit_s": 322.6, + "first_tag_s": 582.3 + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 19.75, + "memory_used_mib_at_start": 65325.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0848, + "cost_usd_upper_bound": 0.011, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9989, + "mean_power_w": 229.53, + "mean_sm_util_pct": 40.05, + "max_temp_c": 72.0 + } + }, + { + "run_name": "p1_bugfix_gemma4-31b-q4_v6", + "task": "p1_bugfix", + "replicate": 6, + "summary": { + "started_at": "2026-08-02T03:18:47.061801+00:00", + "ended_at": "2026-08-02T03:32:10.651413+00:00", + "elapsed_s": 803.6, + "iterations": 86, + "total_completion_tokens": 21150, + "total_prompt_tokens": 2088543, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "The logalyzer codebase has been adopted and stabilized. \n\nKey improvements include:\n- **Bug Fixes**: Resolved critical parsing error for July dates, and fixed logic errors in status and URL filters.\n- **Performance**: Eliminated a critical $O(n^2)$ bottleneck in `io.load`, reducing processing time for a 50MB log file from >300s (timeout) to ~3.25s. Optimized IP allowlist lookups to $O(1)$ using sets.\n- **Security**: Mitigated RCE vulnerability in `expression_filter` by implementing a whitelist-based AST validator. Fixed potential injection issues in JSON output by replacing manual string concatenation with the `json` module.\n- **Modernization**: Updated deprecated `collections` imports for Python 3.10+ compatibility.\n- **Test Suite**: Expanded test coverage from 38% to 90%, adding new tests for the CLI, IO, Output, and Utility modules.\n- **Codebase Cleanup**: Removed unused `experimental.py` module.\n- **Documentation**: Updated README and CHANGELOG to be truthful and detailed; provided full triage, research notes, and architecture decision records (ADRs).\n\nThe final repository is tagged as `v0.3.2` and is now in a state ready for hand-off to a new engineer.\n" + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p1_bugfix_gemma4-31b-q4_v6", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 803.6, + "iters": 86, + "tokens": { + "completion_total": 21150, + "prompt_total": 2088543, + "all_total": 2109693 + }, + "throughput": { + "completion_tps_avg": 43.9, + "model_call_wall_s": 481.2, + "tool_call_wall_s": 315.2, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 326.3, + "first_commit_s": 328.8, + "first_tag_s": 774.9 + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 25.48, + "memory_used_mib_at_start": 65325.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.1116, + "cost_usd_upper_bound": 0.0145, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9893, + "mean_power_w": 291.98, + "mean_sm_util_pct": 52.56, + "max_temp_c": 82.0 + } + }, + { + "run_name": "p1_bugfix_gemma4-31b-q4_v7", + "task": "p1_bugfix", + "replicate": 7, + "summary": { + "started_at": "2026-08-02T03:17:24.370595+00:00", + "ended_at": "2026-08-02T03:30:01.519631+00:00", + "elapsed_s": 757.1, + "iterations": 73, + "total_completion_tokens": 20875, + "total_prompt_tokens": 1635489, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "The logalyzer codebase has been adopted and stabilized. \n\nKey achievements:\n1. Fixed critical compatibility and parsing bugs (Python 3.11 ImportError, missing 'Jul' in month names).\n2. Resolved a major performance bottleneck in `io.load()`, reducing processing time for 50MB logs from >300s to 3.2s (~94x speedup).\n3. Hardened security by replacing `eval()` with a restricted AST evaluator and fixing manual JSON construction.\n4. Removed dead code (unused module `experimental.py` and utility functions).\n5. Expanded test coverage from effectively 0% to 88% by adding tests for all core modules.\n6. Modernized the codebase and documentation, including a detailed triage report, before-after metrics, and ADRs.\n\nThe repository is now at version v0.3.2 and is stable and well-documented." + }, + "verdict": "FAIL", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p1_bugfix_gemma4-31b-q4_v7", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 757.1, + "iters": 73, + "tokens": { + "completion_total": 20875, + "prompt_total": 1635489, + "all_total": 1656364 + }, + "throughput": { + "completion_tps_avg": 47.9, + "model_call_wall_s": 435.8, + "tool_call_wall_s": 315.4, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 322.3, + "first_commit_s": 353.1, + "first_tag_s": 728.9 + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 197.07, + "memory_used_mib_at_start": 65325.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.1052, + "cost_usd_upper_bound": 0.0137, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9906, + "mean_power_w": 290.56, + "mean_sm_util_pct": 51.36, + "max_temp_c": 73.0 + } + }, + { + "run_name": "p1_bugfix_gemma4-31b-q4_v8", + "task": "p1_bugfix", + "replicate": 8, + "summary": { + "started_at": "2026-08-02T03:32:12.959329+00:00", + "ended_at": "2026-08-02T03:44:19.151317+00:00", + "elapsed_s": 726.2, + "iterations": 83, + "total_completion_tokens": 19314, + "total_prompt_tokens": 1725963, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Adopted and stabilized the logalyzer codebase. \n\nKey accomplishments:\n1. **Fixed critical bugs:** Resolved Python 3.10+ import errors, fixed a missing month in the parser, and corrected filter logic for status codes and URLs.\n2. **Major performance boost:** Eliminated an O(n^2) bottleneck in the file loading logic, reducing processing time for a 50MB log from >300s to ~3.2s (approx 100x speedup).\n3. **Security hardening:** Replaced a dangerous `eval()` call with a restricted AST evaluator and replaced manual JSON string construction with the standard `json` module.\n4. **Code modernization:** Removed dead code (`experimental.py` and unused utils) and unified versioning.\n5. **Improved quality:** Increased test coverage from 0% to 92% and resolved all Ruff linting issues.\n6. **Documentation:** Updated the README to remove false claims and provided a detailed CHANGELOG, triage report, and architectural decisions.\n\nThe resulting repository is fully tested, performant, and secure.\n" + }, + "verdict": "FAIL", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p1_bugfix_gemma4-31b-q4_v8", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 726.2, + "iters": 83, + "tokens": { + "completion_total": 19314, + "prompt_total": 1725963, + "all_total": 1745277 + }, + "throughput": { + "completion_tps_avg": 47.4, + "model_call_wall_s": 407.4, + "tool_call_wall_s": 311.9, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 327.8, + "first_commit_s": 329.9, + "first_tag_s": 704.2 + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 124.83, + "memory_used_mib_at_start": 65581.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.1009, + "cost_usd_upper_bound": 0.0131, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9915, + "mean_power_w": 275.15, + "mean_sm_util_pct": 52.54, + "max_temp_c": 79.0 + } + }, + { + "run_name": "p1_bugfix_gemma4-31b-q4_v9", + "task": "p1_bugfix", + "replicate": 9, + "summary": { + "started_at": "2026-08-02T03:30:03.793776+00:00", + "ended_at": "2026-08-02T03:42:40.019013+00:00", + "elapsed_s": 756.2, + "iterations": 70, + "total_completion_tokens": 20369, + "total_prompt_tokens": 1602179, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Stabilized the logalyzer codebase by fixing critical bugs (ImportError, timestamp parsing), addressing a severe performance bottleneck (O(n^2) load time reduced from >300s to 3.3s), and eliminating a security vulnerability (replaced eval() with a safe AST evaluator). Improved test coverage from 34% to 84%, modernized deprecated Python patterns, and removed dead code. All changes are documented in ADRs, a triage report, and a before-after metrics table. The final repository is tagged as v0.3.0." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p1_bugfix_gemma4-31b-q4_v9", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 756.2, + "iters": 70, + "tokens": { + "completion_total": 20369, + "prompt_total": 1602179, + "all_total": 1622548 + }, + "throughput": { + "completion_tps_avg": 47.0, + "model_call_wall_s": 433.5, + "tool_call_wall_s": 316.9, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 323.3, + "first_commit_s": 331.5, + "first_tag_s": 698.3 + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 208.94, + "memory_used_mib_at_start": 65325.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.105, + "cost_usd_upper_bound": 0.0137, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9918, + "mean_power_w": 287.22, + "mean_sm_util_pct": 51.15, + "max_temp_c": 73.0 + } + }, + { + "run_name": "p1_bugfix_gemma4-31b-q4_v10", + "task": "p1_bugfix", + "replicate": 10, + "summary": { + "started_at": "2026-08-02T03:44:21.459804+00:00", + "ended_at": "2026-08-02T03:56:08.570275+00:00", + "elapsed_s": 707.1, + "iterations": 76, + "total_completion_tokens": 18580, + "total_prompt_tokens": 1567028, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "The logalyzer codebase has been stabilized and modernized. Key improvements include:\n- **Performance**: Fixed an O(n\u00b2) bottleneck in log loading, reducing 50MB log processing time from >300s to ~3.2s.\n- **Security**: Replaced `eval()` in the custom expression filter with a safe AST-based evaluator to prevent code injection.\n- **Bugs**: Fixed a critical date parsing bug (missing July), normalized CLI input for status filters, fixed URL regex behavior to support substring matches, and ensured time aggregations use UTC.\n- **Modernization**: Updated deprecated `collections.Iterable` to `collections.abc.Iterable` for Python 3.11 compatibility.\n- **Quality**: Increased test coverage from 39% to 91% by adding comprehensive tests for the CLI, IO, output, and utility modules.\n- **Maintenance**: Removed unused experimental code and legacy utilities.\n- **Documentation**: Provided a full triage report, an architecture ADR, a detailed changelog, and updated the README.\nThe final state is tagged as v0.3.2." + }, + "verdict": "FAIL", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p1_bugfix_gemma4-31b-q4_v10", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 707.1, + "iters": 76, + "tokens": { + "completion_total": 18580, + "prompt_total": 1567028, + "all_total": 1585608 + }, + "throughput": { + "completion_tps_avg": 48.0, + "model_call_wall_s": 387.2, + "tool_call_wall_s": 313.5, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 322.4, + "first_commit_s": 368.1, + "first_tag_s": 686.9 + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 133.25, + "memory_used_mib_at_start": 65581.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0982, + "cost_usd_upper_bound": 0.0128, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.997, + "mean_power_w": 256.16, + "mean_sm_util_pct": 43.77, + "max_temp_c": 80.0 + } + }, + { + "run_name": "p1_testwrite_gemma4-31b-q4_v1", + "task": "p1_testwrite", + "replicate": 1, + "summary": { + "started_at": "2026-08-02T00:45:34.629990+00:00", + "ended_at": "2026-08-02T00:54:18.228621+00:00", + "elapsed_s": 523.6, + "iterations": 49, + "total_completion_tokens": 25223, + "total_prompt_tokens": 1273638, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Increased test coverage of the logalyzer codebase from 25% to 72% by adding high-value tests for under-covered modules. \n\nKey accomplishments:\n- Added new test files: `tests/test_cli.py`, `tests/test_experimental.py`, `tests/test_io.py`, `tests/test_output.py`, `tests/test_utils.py`, and `tests/test_parser_edge.py`.\n- Achieved 100% coverage for `experimental.py`, `io.py`, `output.py`, and `utils.py`.\n- Significantly increased coverage for `cli.py` (98%), `filters.py` (97%), and `parser.py` (88%).\n- Implemented a monkeypatch in `tests/conftest.py` to resolve a Python 3.11 incompatibility in the production code (`collections.Iterable` removal) without modifying the production code itself.\n- Documented several baseline bugs (e.g., `status_filter` type mismatch, `re.match` instead of `re.search` in `url_regex_filter`, and a missing 'Jul' in `MONTH_NAMES`) in `research/dead-ends.md` as per the rules.\n- Maintained the original codebase's byte-identical status in `/logalyzer/`.\n- Followed the required repository structure and provided a detailed `CHANGELOG.md`.\n- Committed changes frequently with descriptive messages and tagged the final release.\n" + }, + "verdict": "FAIL", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p1_testwrite_gemma4-31b-q4_v1", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 523.6, + "iters": 49, + "tokens": { + "completion_total": 25223, + "prompt_total": 1273638, + "all_total": 1298861 + }, + "throughput": { + "completion_tps_avg": 49.1, + "model_call_wall_s": 514.0, + "tool_call_wall_s": 6.8, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 56.9, + "first_commit_s": 0.1, + "first_tag_s": 502.9 + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 87.25, + "memory_used_mib_at_start": 65453.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0727, + "cost_usd_upper_bound": 0.0095, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9836, + "mean_power_w": 481.16, + "mean_sm_util_pct": 91.31, + "max_temp_c": 81.0 + } + }, + { + "run_name": "p1_testwrite_gemma4-31b-q4_v2", + "task": "p1_testwrite", + "replicate": 2, + "summary": { + "started_at": "2026-08-02T00:56:00.566809+00:00", + "ended_at": "2026-08-02T01:02:29.405349+00:00", + "elapsed_s": 388.8, + "iterations": 48, + "total_completion_tokens": 18809, + "total_prompt_tokens": 1029851, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Increased test coverage of the logalyzer tool from 34% to 94% by adding high-value tests for previously uncovered modules (`experimental.py`, `io.py`, `output.py`, `utils.py`, and `cli.py`). All new tests pass, and baseline tests were preserved (with documented bugs in the starter codebase). Followed the requested directory structure and documented decisions in ADRs." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p1_testwrite_gemma4-31b-q4_v2", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 388.8, + "iters": 48, + "tokens": { + "completion_total": 18809, + "prompt_total": 1029851, + "all_total": 1048660 + }, + "throughput": { + "completion_tps_avg": 49.5, + "model_call_wall_s": 379.9, + "tool_call_wall_s": 6.1, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 24.6, + "first_commit_s": 6.5, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 500.03, + "memory_used_mib_at_start": 65453.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.054, + "cost_usd_upper_bound": 0.007, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9902, + "mean_power_w": 479.04, + "mean_sm_util_pct": 89.65, + "max_temp_c": 78.0 + } + }, + { + "run_name": "p1_testwrite_gemma4-31b-q4_v3", + "task": "p1_testwrite", + "replicate": 3, + "summary": { + "started_at": "2026-08-02T00:54:18.937486+00:00", + "ended_at": "2026-08-02T01:00:58.817170+00:00", + "elapsed_s": 399.9, + "iterations": 42, + "total_completion_tokens": 19900, + "total_prompt_tokens": 825275, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Increased test coverage of the logalyzer codebase from 34% to 98% by adding comprehensive test suites for previously uncovered modules. \n\nKey accomplishments:\n- Added `tests/test_experimental.py`, `tests/test_io.py`, `tests/test_output.py`, `tests/test_utils.py`, and `tests/test_cli.py`.\n- Improved `tests/test_parser.py` to cover error paths.\n- Implemented a monkeypatch in `tests/conftest.py` to handle `collections.Iterable` removal in Python 3.11, ensuring tests can run without modifying production code.\n- Documented three critical production bugs in `research/dead-ends.md` that caused baseline test failures (status filter type mismatch, URL regex matching behavior, and missing 'Jul' in MONTH_NAMES).\n- Maintained zero regressions to passing baseline tests.\n- Adhered strictly to the requirement of not modifying production code.\n" + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p1_testwrite_gemma4-31b-q4_v3", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 399.9, + "iters": 42, + "tokens": { + "completion_total": 19900, + "prompt_total": 825275, + "all_total": 845175 + }, + "throughput": { + "completion_tps_avg": 50.8, + "model_call_wall_s": 392.0, + "tool_call_wall_s": 5.4, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 24.0, + "first_commit_s": 4.7, + "first_tag_s": 380.1 + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 251.09, + "memory_used_mib_at_start": 65453.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0555, + "cost_usd_upper_bound": 0.0072, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9877, + "mean_power_w": 468.52, + "mean_sm_util_pct": 91.96, + "max_temp_c": 86.0 + } + }, + { + "run_name": "p1_testwrite_gemma4-31b-q4_v4", + "task": "p1_testwrite", + "replicate": 4, + "summary": { + "started_at": "2026-08-02T03:56:10.910803+00:00", + "ended_at": "2026-08-02T04:00:48.898736+00:00", + "elapsed_s": 278.0, + "iterations": 38, + "total_completion_tokens": 14087, + "total_prompt_tokens": 655459, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Increased test coverage of the `logalyzer` codebase from 34% to 88% by adding high-value tests for previously uncovered modules.\n\nKey accomplishments:\n1. **Comprehensive Test Suite**: Added four new test files targeting `experimental.py`, `utils.py`, `io.py`, and `cli.py`.\n2. **Targeted Coverage**:\n - `experimental.py`: 0% -> 100% (Anomaly detection and session windowing).\n - `utils.py`: 0% -> 100% (Date parsing, byte formatting, legacy support, and chunking).\n - `io.py`: 0% -> 100% (Log loading and iteration).\n - `cli.py`: 0% -> 85% (Integration tests for CLI arguments and output formats).\n3. **Zero Production Changes**: No modifications were made to the original `/logalyzer/` package.\n4. **Environment Compatibility**: Implemented a `sitecustomize.py` monkeypatch to ensure the codebase (which uses deprecated `collections.Iterable`) runs on Python 3.11 without modifying source code.\n5. **Regression Testing**: Verified that all baseline tests that previously passed still pass, and those that failed still fail (documenting these as existing bugs).\n6. **Documentation**: Maintained a detailed `CHANGELOG.md`, ADRs in `/decisions/`, and research notes in `/research/`.\n\nFinal coverage: 88%.\n" + }, + "verdict": "FAIL", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p1_testwrite_gemma4-31b-q4_v4", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 278.0, + "iters": 38, + "tokens": { + "completion_total": 14087, + "prompt_total": 655459, + "all_total": 669546 + }, + "throughput": { + "completion_tps_avg": 52.0, + "model_call_wall_s": 270.7, + "tool_call_wall_s": 5.1, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 37.6, + "first_commit_s": 5.4, + "first_tag_s": 255.2 + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 133.81, + "memory_used_mib_at_start": 65581.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0386, + "cost_usd_upper_bound": 0.005, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9892, + "mean_power_w": 471.63, + "mean_sm_util_pct": 88.3, + "max_temp_c": 82.0 + } + }, + { + "run_name": "p1_testwrite_gemma4-31b-q4_v5", + "task": "p1_testwrite", + "replicate": 5, + "summary": { + "started_at": "2026-08-02T03:42:42.285078+00:00", + "ended_at": "2026-08-02T03:48:15.550227+00:00", + "elapsed_s": 333.3, + "iterations": 37, + "total_completion_tokens": 16936, + "total_prompt_tokens": 648484, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Increased test coverage of logalyzer from 34% to 93% by adding high-value tests for under-covered modules. \n\nKey contributions:\n- Added `tests/test_experimental.py` providing 100% coverage for `experimental.py` (anomaly detection and session windowing).\n- Added `tests/test_utils.py` providing 100% coverage for `utils.py` (date parsing, byte formatting, legacy output, and chunking).\n- Added `tests/test_io.py` providing 100% coverage for `io.py` (file loading and iterative parsing).\n- Added `tests/test_output.py` providing 100% coverage for `output.py` (JSON, CSV, text formatters and aggregate rendering).\n- Added `tests/test_cli.py` providing 86% coverage for `cli.py` (integration tests for CLI arguments and end-to-end flow).\n- Implemented a runtime monkeypatch in `tests/conftest.py` to fix a Python 3.11 `ImportError` in the production code without modifying the production files.\n\nIdentified and documented several baseline bugs in the production code (e.g., missing month in `parser.py`, logic errors in `filters.py`) in ADRs and research notes, adhering to the constraint of not modifying production code.\n" + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p1_testwrite_gemma4-31b-q4_v5", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 333.3, + "iters": 37, + "tokens": { + "completion_total": 16936, + "prompt_total": 648484, + "all_total": 665420 + }, + "throughput": { + "completion_tps_avg": 51.9, + "model_call_wall_s": 326.4, + "tool_call_wall_s": 4.8, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 42.3, + "first_commit_s": 5.3, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 499.89, + "memory_used_mib_at_start": 65581.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0463, + "cost_usd_upper_bound": 0.006, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9751, + "mean_power_w": 480.35, + "mean_sm_util_pct": 90.21, + "max_temp_c": 76.0 + } + }, + { + "run_name": "p1_testwrite_gemma4-31b-q4_v6", + "task": "p1_testwrite", + "replicate": 6, + "summary": { + "started_at": "2026-08-02T04:00:49.672803+00:00", + "ended_at": "2026-08-02T04:08:34.370068+00:00", + "elapsed_s": 464.7, + "iterations": 37, + "total_completion_tokens": 22810, + "total_prompt_tokens": 818645, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Increased test coverage of the `logalyzer` codebase from 34% to 97% by adding high-value tests for the previously uncovered `cli.py`, `experimental.py`, `io.py`, `output.py`, and `utils.py` modules. Baseline tests were preserved, and bugs encountered in the production code were documented in `research/dead-ends.md` and `CHANGELOG.md` without modifying the original source code, as per the requirements. All new tests pass against the current behavior of the codebase." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p1_testwrite_gemma4-31b-q4_v6", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 464.7, + "iters": 37, + "tokens": { + "completion_total": 22810, + "prompt_total": 818645, + "all_total": 841455 + }, + "throughput": { + "completion_tps_avg": 49.9, + "model_call_wall_s": 457.2, + "tool_call_wall_s": 5.3, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 23.5, + "first_commit_s": 4.7, + "first_tag_s": 450.0 + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 234.92, + "memory_used_mib_at_start": 65581.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0645, + "cost_usd_upper_bound": 0.0084, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9899, + "mean_power_w": 481.71, + "mean_sm_util_pct": 92.33, + "max_temp_c": 82.0 + } + }, + { + "run_name": "p1_testwrite_gemma4-31b-q4_v7", + "task": "p1_testwrite", + "replicate": 7, + "summary": { + "started_at": "2026-08-02T03:48:16.350597+00:00", + "ended_at": "2026-08-02T03:55:13.320391+00:00", + "elapsed_s": 417.0, + "iterations": 44, + "total_completion_tokens": 20892, + "total_prompt_tokens": 856122, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Increased test coverage of the logalyzer codebase from 34% to 91%. Added high-value tests targeting under-covered modules (experimental, utils, output, io, and cli), focusing on edge cases, error paths, and integration. All new tests pass against the current codebase. Identified and documented three pre-existing bugs in the production code in ADR 0001, which were left unfixed to adhere to the no-production-change rule. Fixed a Python 3.11 compatibility issue in the test environment using a monkeypatch in conftest.py. Final repository structure adheres to the requested layout, including a detailed CHANGELOG.md." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p1_testwrite_gemma4-31b-q4_v7", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 417.0, + "iters": 44, + "tokens": { + "completion_total": 20892, + "prompt_total": 856122, + "all_total": 877014 + }, + "throughput": { + "completion_tps_avg": 51.0, + "model_call_wall_s": 409.6, + "tool_call_wall_s": 4.9, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 64.2, + "first_commit_s": 5.4, + "first_tag_s": 398.8 + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 16.47, + "memory_used_mib_at_start": 65581.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0579, + "cost_usd_upper_bound": 0.0075, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9952, + "mean_power_w": 479.92, + "mean_sm_util_pct": 90.1, + "max_temp_c": 74.0 + } + }, + { + "run_name": "p1_testwrite_gemma4-31b-q4_v8", + "task": "p1_testwrite", + "replicate": 8, + "summary": { + "started_at": "2026-08-02T04:08:35.234484+00:00", + "ended_at": "2026-08-02T04:16:19.199334+00:00", + "elapsed_s": 464.0, + "iterations": 50, + "total_completion_tokens": 22263, + "total_prompt_tokens": 1361395, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Increased test coverage of the logalyzer tool from 34% to 99% by adding high-value tests for the CLI, IO, Output, and Experimental modules, as well as edge cases for the parser and filters. Baseline test failures were identified as production bugs and documented in `research/dead-ends.md` without modifying the production code. A monkeypatch was added to `tests/conftest.py` to ensure Python 3.11 compatibility. All added tests pass against the current codebase behavior." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p1_testwrite_gemma4-31b-q4_v8", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 464.0, + "iters": 50, + "tokens": { + "completion_total": 22263, + "prompt_total": 1361395, + "all_total": 1383658 + }, + "throughput": { + "completion_tps_avg": 49.0, + "model_call_wall_s": 454.5, + "tool_call_wall_s": 6.6, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 33.2, + "first_commit_s": 4.6, + "first_tag_s": 442.2 + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 499.92, + "memory_used_mib_at_start": 65581.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0644, + "cost_usd_upper_bound": 0.0084, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9914, + "mean_power_w": 474.5, + "mean_sm_util_pct": 91.35, + "max_temp_c": 82.0 + } + }, + { + "run_name": "p1_testwrite_gemma4-31b-q4_v9", + "task": "p1_testwrite", + "replicate": 9, + "summary": { + "started_at": "2026-08-02T03:55:14.058907+00:00", + "ended_at": "2026-08-02T04:06:38.149680+00:00", + "elapsed_s": 684.1, + "iterations": 54, + "total_completion_tokens": 30030, + "total_prompt_tokens": 1728443, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Increased test coverage of `logalyzer` from 34% to 97% by adding high-value tests for `cli.py`, `experimental.py`, `io.py`, `output.py`, and `utils.py`. Handled Python 3.11 compatibility issues via `tests/conftest.py` to avoid production code changes. Documented baseline bugs and decisions in `/decisions/` and `/research/`." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p1_testwrite_gemma4-31b-q4_v9", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 684.1, + "iters": 54, + "tokens": { + "completion_total": 30030, + "prompt_total": 1728443, + "all_total": 1758473 + }, + "throughput": { + "completion_tps_avg": 44.6, + "model_call_wall_s": 673.0, + "tool_call_wall_s": 8.0, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 25.6, + "first_commit_s": 4.6, + "first_tag_s": 621.1 + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 499.95, + "memory_used_mib_at_start": 65581.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.095, + "cost_usd_upper_bound": 0.0123, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.994, + "mean_power_w": 487.82, + "mean_sm_util_pct": 93.58, + "max_temp_c": 75.0 + } + }, + { + "run_name": "p1_testwrite_gemma4-31b-q4_v10", + "task": "p1_testwrite", + "replicate": 10, + "summary": { + "started_at": "2026-08-02T04:16:20.083544+00:00", + "ended_at": "2026-08-02T04:22:40.621129+00:00", + "elapsed_s": 380.5, + "iterations": 53, + "total_completion_tokens": 17822, + "total_prompt_tokens": 1323034, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Increased test coverage of the logalyzer codebase from 34% to 99% by adding high-value tests for under-covered modules. Targeted `utils.py`, `experimental.py`, `io.py`, `output.py`, `cli.py`, and `parser.py`, focusing on edge cases, error paths, and integration. Maintained a strict policy of no changes to production code, documenting three baseline bugs in `research/dead-ends.md`. All added tests pass against the current codebase (with the exception of the documented baseline bugs). Implementation includes a Python 3.11 compatibility fix in `tests/conftest.py` to handle the removal of `collections.Iterable`. Full documentation of decisions and changes is provided in the repository." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p1_testwrite_gemma4-31b-q4_v10", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 380.5, + "iters": 53, + "tokens": { + "completion_total": 17822, + "prompt_total": 1323034, + "all_total": 1340856 + }, + "throughput": { + "completion_tps_avg": 48.2, + "model_call_wall_s": 369.7, + "tool_call_wall_s": 7.6, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 63.1, + "first_commit_s": 4.6, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 499.9, + "memory_used_mib_at_start": 65581.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0528, + "cost_usd_upper_bound": 0.0069, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9855, + "mean_power_w": 466.52, + "mean_sm_util_pct": 89.03, + "max_temp_c": 83.0 + } + }, + { + "run_name": "p1_refactor_gemma4-31b-q4_v1", + "task": "p1_refactor", + "replicate": 1, + "summary": { + "started_at": "2026-08-02T01:02:30.232102+00:00", + "ended_at": "2026-08-02T01:05:36.464935+00:00", + "elapsed_s": 186.2, + "iterations": 35, + "total_completion_tokens": 9779, + "total_prompt_tokens": 351967, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Refactored `logalyzer/output.py` into a focused package `logalyzer/output/` containing `json_renderer.py`, `csv_renderer.py`, `plain_renderer.py`, and `legacy.py`. Preserved the public API through `logalyzer/output/__init__.py`. Verified that existing call sites and tests remain unchanged in behavior, including the preservation of existing bugs in `filters.py` and `parser.py`. All documentation (ADRs, research notes, CHANGELOG) has been created in the required structure." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p1_refactor_gemma4-31b-q4_v1", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 186.2, + "iters": 35, + "tokens": { + "completion_total": 9779, + "prompt_total": 351967, + "all_total": 361746 + }, + "throughput": { + "completion_tps_avg": 53.9, + "model_call_wall_s": 181.5, + "tool_call_wall_s": 2.8, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 70.6, + "first_commit_s": 62.8, + "first_tag_s": 167.0 + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 498.71, + "memory_used_mib_at_start": 65453.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0259, + "cost_usd_upper_bound": 0.0034, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9667, + "mean_power_w": 465.36, + "mean_sm_util_pct": 85.35, + "max_temp_c": 78.0 + } + }, + { + "run_name": "p1_refactor_gemma4-31b-q4_v2", + "task": "p1_refactor", + "replicate": 2, + "summary": { + "started_at": "2026-08-02T01:00:59.565226+00:00", + "ended_at": "2026-08-02T01:04:31.084608+00:00", + "elapsed_s": 211.5, + "iterations": 41, + "total_completion_tokens": 11126, + "total_prompt_tokens": 508191, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Refactored `logalyzer/output.py` into a `logalyzer/output/` package with specialized modules (`json_renderer.py`, `csv_renderer.py`, `plain_renderer.py`, `legacy.py`) while maintaining full backward compatibility for the public API. Verified that the test suite (including existing bugs) remains unchanged and that other files are byte-identical. Documentation provided in `/decisions/`, `/research/`, and `CHANGELOG.md`." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p1_refactor_gemma4-31b-q4_v2", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 211.5, + "iters": 41, + "tokens": { + "completion_total": 11126, + "prompt_total": 508191, + "all_total": 519317 + }, + "throughput": { + "completion_tps_avg": 54.1, + "model_call_wall_s": 205.7, + "tool_call_wall_s": 3.5, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 101.4, + "first_commit_s": 98.0, + "first_tag_s": 196.3 + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 253.02, + "memory_used_mib_at_start": 65453.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0294, + "cost_usd_upper_bound": 0.0038, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9693, + "mean_power_w": 459.06, + "mean_sm_util_pct": 89.05, + "max_temp_c": 84.0 + } + }, + { + "run_name": "p1_refactor_gemma4-31b-q4_v3", + "task": "p1_refactor", + "replicate": 3, + "summary": { + "started_at": "2026-08-02T02:40:35.025584+00:00", + "ended_at": "2026-08-02T02:43:53.245183+00:00", + "elapsed_s": 198.2, + "iterations": 40, + "total_completion_tokens": 10797, + "total_prompt_tokens": 399232, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Refactored `logalyzer/output.py` into a modular subpackage `logalyzer/output/` containing `json_renderer.py`, `csv_renderer.py`, `plain_renderer.py`, and `legacy.py`. Preserved the public API surface and the exact behavior of the original codebase, including a pre-existing `ImportError` in tests caused by Python 3.11 compatibility issues in `logalyzer/filters.py` (which was left untouched per constraints). All verification steps passed except for the pre-existing test failure. Documentation (ADRs, research notes, and CHANGELOG) has been provided. Final release tagged as v0.2.1." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p1_refactor_gemma4-31b-q4_v3", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 198.2, + "iters": 40, + "tokens": { + "completion_total": 10797, + "prompt_total": 399232, + "all_total": 410029 + }, + "throughput": { + "completion_tps_avg": 56.0, + "model_call_wall_s": 192.8, + "tool_call_wall_s": 3.1, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 101.4, + "first_commit_s": 89.7, + "first_tag_s": 186.1 + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 17.56, + "memory_used_mib_at_start": 65197.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0275, + "cost_usd_upper_bound": 0.0036, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9586, + "mean_power_w": 471.53, + "mean_sm_util_pct": 87.9, + "max_temp_c": 71.0 + } + }, + { + "run_name": "p1_refactor_gemma4-31b-q4_v4", + "task": "p1_refactor", + "replicate": 4, + "summary": { + "started_at": "2026-08-02T04:22:41.544812+00:00", + "ended_at": "2026-08-02T04:25:29.279714+00:00", + "elapsed_s": 167.7, + "iterations": 33, + "total_completion_tokens": 8832, + "total_prompt_tokens": 322475, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Successfully refactored `logalyzer/output.py` into a focused package `logalyzer/output/` containing `json_renderer.py`, `csv_renderer.py`, `plain_renderer.py`, and `legacy.py`. The public API is preserved via `logalyzer/output/__init__.py`. All verification checks passed, including a check that non-output files and tests remained byte-identical. A pre-existing bug regarding `collections.Iterable` in Python 3.11 was identified and documented in `research/dead-ends.md` as per instructions. The final state is tagged as `v0.2.1`." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p1_refactor_gemma4-31b-q4_v4", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 167.7, + "iters": 33, + "tokens": { + "completion_total": 8832, + "prompt_total": 322475, + "all_total": 331307 + }, + "throughput": { + "completion_tps_avg": 54.2, + "model_call_wall_s": 163.1, + "tool_call_wall_s": 2.7, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 70.9, + "first_commit_s": 64.5, + "first_tag_s": 150.2 + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 499.93, + "memory_used_mib_at_start": 65581.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0233, + "cost_usd_upper_bound": 0.003, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9839, + "mean_power_w": 451.04, + "mean_sm_util_pct": 82.21, + "max_temp_c": 81.0 + } + }, + { + "run_name": "p1_refactor_gemma4-31b-q4_v5", + "task": "p1_refactor", + "replicate": 5, + "summary": { + "started_at": "2026-08-02T04:06:38.949743+00:00", + "ended_at": "2026-08-02T04:09:28.009494+00:00", + "elapsed_s": 169.1, + "iterations": 33, + "total_completion_tokens": 8879, + "total_prompt_tokens": 300783, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Refactored `logalyzer/output.py` into a modular subpackage `logalyzer/output/` containing `json_renderer.py`, `csv_renderer.py`, `plain_renderer.py`, and `legacy.py`. The public API was preserved via `logalyzer/output/__init__.py`. \n\nKey points:\n- `format_legacy` was found in `utils.py` instead of `output.py`. To satisfy both the \"don't touch `utils.py`\" and \"make `logalyzer.output.format_legacy` work\" constraints, the function was duplicated into `logalyzer/output/legacy.py`.\n- Verified that `from logalyzer.output import format_legacy` works.\n- Verified that no other files in the repository were modified.\n- Observed that the existing test suite fails on Python 3.11 due to a pre-existing bug (`collections.Iterable` import in `filters.py`), which was documented in `research/dead-ends.md` but not fixed as per the constraints.\n- All changes committed to git and tagged as `v1.0.0`.\n" + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p1_refactor_gemma4-31b-q4_v5", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 169.1, + "iters": 33, + "tokens": { + "completion_total": 8879, + "prompt_total": 300783, + "all_total": 309662 + }, + "throughput": { + "completion_tps_avg": 54.0, + "model_call_wall_s": 164.5, + "tool_call_wall_s": 2.6, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 73.6, + "first_commit_s": 68.4, + "first_tag_s": 147.6 + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 288.24, + "memory_used_mib_at_start": 65581.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0235, + "cost_usd_upper_bound": 0.0031, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9758, + "mean_power_w": 462.68, + "mean_sm_util_pct": 88.91, + "max_temp_c": 75.0 + } + }, + { + "run_name": "p1_refactor_gemma4-31b-q4_v6", + "task": "p1_refactor", + "replicate": 6, + "summary": { + "started_at": "2026-08-02T04:25:30.233598+00:00", + "ended_at": "2026-08-02T04:28:25.363299+00:00", + "elapsed_s": 175.1, + "iterations": 33, + "total_completion_tokens": 9583, + "total_prompt_tokens": 316151, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Successfully refactored `logalyzer/output.py` into a structured `logalyzer/output/` package. \n\nKey changes:\n- Created `logalyzer/output/` package with `json_renderer.py`, `csv_renderer.py`, `plain_renderer.py`, and `legacy.py`.\n- Implemented a public API in `logalyzer/output/__init__.py` to ensure backward compatibility.\n- Moved `format_legacy` from `logalyzer/utils.py` to `logalyzer/output/legacy.py` to satisfy the refactor requirements and verification checks.\n- Documented a known bug in the starter codebase (`collections.Iterable` import error in Python 3.11) in `research/dead-ends.md`, which was left unfixed per instructions.\n- Provided ADRs, a CHANGELOG, and verified all constraints.\n- Committed all changes to a new git repository at `/workspace/` and tagged the release as `v1.0.0`.\n" + }, + "verdict": "FAIL", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p1_refactor_gemma4-31b-q4_v6", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 175.1, + "iters": 33, + "tokens": { + "completion_total": 9583, + "prompt_total": 316151, + "all_total": 325734 + }, + "throughput": { + "completion_tps_avg": 56.2, + "model_call_wall_s": 170.6, + "tool_call_wall_s": 2.6, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 78.4, + "first_commit_s": 74.2, + "first_tag_s": 158.5 + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 165.75, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0243, + "cost_usd_upper_bound": 0.0032, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9709, + "mean_power_w": 453.36, + "mean_sm_util_pct": 83.6, + "max_temp_c": 79.0 + } + }, + { + "run_name": "p1_refactor_gemma4-31b-q4_v7", + "task": "p1_refactor", + "replicate": 7, + "summary": { + "started_at": "2026-08-02T04:09:28.890514+00:00", + "ended_at": "2026-08-02T04:12:55.318323+00:00", + "elapsed_s": 206.4, + "iterations": 42, + "total_completion_tokens": 10802, + "total_prompt_tokens": 514218, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "The `logalyzer/output.py` module has been refactored into a `logalyzer/output/` package, separating JSON, CSV, plain-text, and legacy rendering logic into their own modules while maintaining backward compatibility for the public API. Verification tests confirm that the behavior remains identical to the original codebase, including existing bugs and failures. Documentation (ADRs, CHANGELOG, research notes) has been provided in the repository." + }, + "verdict": "FAIL", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p1_refactor_gemma4-31b-q4_v7", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 206.4, + "iters": 42, + "tokens": { + "completion_total": 10802, + "prompt_total": 514218, + "all_total": 525020 + }, + "throughput": { + "completion_tps_avg": 54.0, + "model_call_wall_s": 200.1, + "tool_call_wall_s": 3.9, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 81.9, + "first_commit_s": 78.7, + "first_tag_s": 193.1 + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 255.48, + "memory_used_mib_at_start": 65581.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0287, + "cost_usd_upper_bound": 0.0037, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.969, + "mean_power_w": 473.97, + "mean_sm_util_pct": 84.59, + "max_temp_c": 77.0 + } + }, + { + "run_name": "p1_refactor_gemma4-31b-q4_v8", + "task": "p1_refactor", + "replicate": 8, + "summary": { + "started_at": "2026-08-02T04:28:26.383576+00:00", + "ended_at": "2026-08-02T04:32:25.090948+00:00", + "elapsed_s": 238.7, + "iterations": 35, + "total_completion_tokens": 12046, + "total_prompt_tokens": 404131, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Refactored `logalyzer/output.py` into a focused package `logalyzer/output/` with separate modules for JSON, CSV, plain-text rendering, and legacy compatibility. Maintained the public API via `__init__.py` and ensured all non-output files remained byte-identical to the original repository. Documented a pre-existing Python 3.11 compatibility bug in `dead-ends.md` and summarized the changes in `CHANGELOG.md`. Verified the refactor with the provided checks." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p1_refactor_gemma4-31b-q4_v8", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 238.7, + "iters": 35, + "tokens": { + "completion_total": 12046, + "prompt_total": 404131, + "all_total": 416177 + }, + "throughput": { + "completion_tps_avg": 51.5, + "model_call_wall_s": 233.8, + "tool_call_wall_s": 3.0, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 73.8, + "first_commit_s": 68.8, + "first_tag_s": 226.9 + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 161.77, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0332, + "cost_usd_upper_bound": 0.0043, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9845, + "mean_power_w": 460.08, + "mean_sm_util_pct": 88.42, + "max_temp_c": 83.0 + } + }, + { + "run_name": "p1_refactor_gemma4-31b-q4_v9", + "task": "p1_refactor", + "replicate": 9, + "summary": { + "started_at": "2026-08-02T04:12:56.151222+00:00", + "ended_at": "2026-08-02T04:16:28.275389+00:00", + "elapsed_s": 212.1, + "iterations": 39, + "total_completion_tokens": 11193, + "total_prompt_tokens": 483466, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "The `logalyzer/output.py` module was refactored into a `logalyzer/output/` package. The logic was split into `json_renderer.py`, `csv_renderer.py`, and `plain_renderer.py` modules. A `legacy.py` module was created as a shim to provide access to `format_legacy()` from `logalyzer.utils` while maintaining the public API surface at `logalyzer.output`. All public functions (`to_json`, `to_csv`, `to_text`, `render_aggregate`, and `format_legacy`) are re-exported from `logalyzer/output/__init__.py`. \n\nVerification results:\n- `pytest`: Pre-existing bugs in `filters.py` and `parser.py` were identified but not fixed as per constraints.\n- `from logalyzer.output import format_legacy`: Successfully imports.\n- `git diff` (excluding `__pycache__`): Confirmed that no files other than `logalyzer/output.py` were modified.\n\nDocumentation has been provided in `/decisions/`, `/research/`, and `CHANGELOG.md`.\n" + }, + "verdict": "FAIL", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p1_refactor_gemma4-31b-q4_v9", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 212.1, + "iters": 39, + "tokens": { + "completion_total": 11193, + "prompt_total": 483466, + "all_total": 494659 + }, + "throughput": { + "completion_tps_avg": 54.2, + "model_call_wall_s": 206.6, + "tool_call_wall_s": 3.4, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 109.3, + "first_commit_s": 97.7, + "first_tag_s": 195.7 + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 258.95, + "memory_used_mib_at_start": 65581.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0295, + "cost_usd_upper_bound": 0.0038, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9901, + "mean_power_w": 460.34, + "mean_sm_util_pct": 83.19, + "max_temp_c": 74.0 + } + }, + { + "run_name": "p1_refactor_gemma4-31b-q4_v10", + "task": "p1_refactor", + "replicate": 10, + "summary": { + "started_at": "2026-08-02T04:32:26.122957+00:00", + "ended_at": "2026-08-02T04:35:17.777451+00:00", + "elapsed_s": 171.7, + "iterations": 35, + "total_completion_tokens": 9197, + "total_prompt_tokens": 360546, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Successfully refactored `logalyzer/output.py` into a focused `logalyzer/output/` package. The split includes `json_renderer.py`, `csv_renderer.py`, `plain_renderer.py`, and `legacy.py`. The public API was preserved via `logalyzer/output/__init__.py`, ensuring no changes were needed in other modules. All verification checks passed, including byte-identical tests and unrelated files. Documentation (ADRs, research notes, and CHANGELOG) has been provided in the repository. The final state is tagged as `v0.2.1`." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p1_refactor_gemma4-31b-q4_v10", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 171.7, + "iters": 35, + "tokens": { + "completion_total": 9197, + "prompt_total": 360546, + "all_total": 369743 + }, + "throughput": { + "completion_tps_avg": 55.1, + "model_call_wall_s": 166.9, + "tool_call_wall_s": 2.8, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 82.7, + "first_commit_s": 64.4, + "first_tag_s": 157.6 + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 203.89, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0238, + "cost_usd_upper_bound": 0.0031, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.961, + "mean_power_w": 452.56, + "mean_sm_util_pct": 90.79, + "max_temp_c": 80.0 + } + }, + { + "run_name": "p2_extract_gemma4-31b-q4_v1", + "task": "p2_extract", + "replicate": 1, + "summary": { + "started_at": "2026-08-02T01:04:31.829002+00:00", + "ended_at": "2026-08-02T01:05:48.226058+00:00", + "elapsed_s": 76.4, + "iterations": 6, + "total_completion_tokens": 4449, + "total_prompt_tokens": 22299, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Extracted 20 financial and operational fields from the Veridyne Networks Q3 2026 press release. All fields matched the provided JSON schema in terms of type and units. Results are stored in /workspace/extraction_results.json." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p2_extract_gemma4-31b-q4_v1", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 76.4, + "iters": 6, + "tokens": { + "completion_total": 4449, + "prompt_total": 22299, + "all_total": 26748 + }, + "throughput": { + "completion_tps_avg": 58.8, + "model_call_wall_s": 75.7, + "tool_call_wall_s": 0.3, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 37.1, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 214.87, + "memory_used_mib_at_start": 65453.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0106, + "cost_usd_upper_bound": 0.0014, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9162, + "mean_power_w": 460.29, + "mean_sm_util_pct": 91.33, + "max_temp_c": 83.0 + } + }, + { + "run_name": "p2_extract_gemma4-31b-q4_v2", + "task": "p2_extract", + "replicate": 2, + "summary": { + "started_at": "2026-08-02T02:08:59.939496+00:00", + "ended_at": "2026-08-02T02:10:12.503118+00:00", + "elapsed_s": 72.6, + "iterations": 6, + "total_completion_tokens": 4449, + "total_prompt_tokens": 22299, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Extracted 20 financial and operational fields from the Veridyne Networks Q3 2026 press release. All fields matched the provided JSON schema in terms of type and units. Results are stored in /workspace/extraction_results.json." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p2_extract_gemma4-31b-q4_v2", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 72.6, + "iters": 6, + "tokens": { + "completion_total": 4449, + "prompt_total": 22299, + "all_total": 26748 + }, + "throughput": { + "completion_tps_avg": 61.9, + "model_call_wall_s": 71.9, + "tool_call_wall_s": 0.3, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 36.7, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 443.93, + "memory_used_mib_at_start": 66221.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0101, + "cost_usd_upper_bound": 0.0013, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.8953, + "mean_power_w": 493.67, + "mean_sm_util_pct": 90.86, + "max_temp_c": 75.0 + } + }, + { + "run_name": "p2_extract_gemma4-31b-q4_v3", + "task": "p2_extract", + "replicate": 3, + "summary": { + "started_at": "2026-08-02T01:05:49.001024+00:00", + "ended_at": "2026-08-02T01:07:02.880774+00:00", + "elapsed_s": 73.9, + "iterations": 6, + "total_completion_tokens": 4449, + "total_prompt_tokens": 22299, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Extracted 20 financial and operational fields from the Veridyne Networks Q3 2026 press release. All fields matched the provided JSON schema in terms of type and units. Results are stored in /workspace/extraction_results.json." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p2_extract_gemma4-31b-q4_v3", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 73.9, + "iters": 6, + "tokens": { + "completion_total": 4449, + "prompt_total": 22299, + "all_total": 26748 + }, + "throughput": { + "completion_tps_avg": 60.8, + "model_call_wall_s": 73.2, + "tool_call_wall_s": 0.3, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 37.1, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 260.34, + "memory_used_mib_at_start": 65453.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0103, + "cost_usd_upper_bound": 0.0013, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.8796, + "mean_power_w": 489.97, + "mean_sm_util_pct": 97.5, + "max_temp_c": 83.0 + } + }, + { + "run_name": "p2_extract_gemma4-31b-q4_v4", + "task": "p2_extract", + "replicate": 4, + "summary": { + "started_at": "2026-08-02T04:35:18.807160+00:00", + "ended_at": "2026-08-02T04:36:34.525299+00:00", + "elapsed_s": 75.7, + "iterations": 6, + "total_completion_tokens": 4449, + "total_prompt_tokens": 22299, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Extracted 20 financial and operational fields from the Veridyne Networks Q3 2026 press release. All fields matched the provided JSON schema in terms of type and units. Results are stored in /workspace/extraction_results.json." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p2_extract_gemma4-31b-q4_v4", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 75.7, + "iters": 6, + "tokens": { + "completion_total": 4449, + "prompt_total": 22299, + "all_total": 26748 + }, + "throughput": { + "completion_tps_avg": 59.3, + "model_call_wall_s": 75.0, + "tool_call_wall_s": 0.3, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 37.0, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 196.02, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0105, + "cost_usd_upper_bound": 0.0014, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9247, + "mean_power_w": 473.25, + "mean_sm_util_pct": 93.13, + "max_temp_c": 80.0 + } + }, + { + "run_name": "p2_extract_gemma4-31b-q4_v5", + "task": "p2_extract", + "replicate": 5, + "summary": { + "started_at": "2026-08-02T04:16:29.122426+00:00", + "ended_at": "2026-08-02T04:17:43.540204+00:00", + "elapsed_s": 74.4, + "iterations": 6, + "total_completion_tokens": 4449, + "total_prompt_tokens": 22299, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Extracted 20 financial and operational fields from the Veridyne Networks Q3 2026 press release. All fields matched the provided JSON schema in terms of type and units. Results are stored in /workspace/extraction_results.json." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p2_extract_gemma4-31b-q4_v5", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 74.4, + "iters": 6, + "tokens": { + "completion_total": 4449, + "prompt_total": 22299, + "all_total": 26748 + }, + "throughput": { + "completion_tps_avg": 60.4, + "model_call_wall_s": 73.7, + "tool_call_wall_s": 0.3, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 36.8, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 275.85, + "memory_used_mib_at_start": 65581.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0103, + "cost_usd_upper_bound": 0.0013, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9409, + "mean_power_w": 497.01, + "mean_sm_util_pct": 97.33, + "max_temp_c": 73.0 + } + }, + { + "run_name": "p2_extract_gemma4-31b-q4_v6", + "task": "p2_extract", + "replicate": 6, + "summary": { + "started_at": "2026-08-02T04:36:35.585240+00:00", + "ended_at": "2026-08-02T04:37:49.353628+00:00", + "elapsed_s": 73.8, + "iterations": 6, + "total_completion_tokens": 4449, + "total_prompt_tokens": 22299, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Extracted 20 financial and operational fields from the Veridyne Networks Q3 2026 press release. All fields matched the provided JSON schema in terms of type and units. Results are stored in /workspace/extraction_results.json." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p2_extract_gemma4-31b-q4_v6", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 73.8, + "iters": 6, + "tokens": { + "completion_total": 4449, + "prompt_total": 22299, + "all_total": 26748 + }, + "throughput": { + "completion_tps_avg": 60.9, + "model_call_wall_s": 73.0, + "tool_call_wall_s": 0.4, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 37.1, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 182.31, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0103, + "cost_usd_upper_bound": 0.0013, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9485, + "mean_power_w": 478.24, + "mean_sm_util_pct": 97.4, + "max_temp_c": 81.0 + } + }, + { + "run_name": "p2_extract_gemma4-31b-q4_v7", + "task": "p2_extract", + "replicate": 7, + "summary": { + "started_at": "2026-08-02T04:17:44.387764+00:00", + "ended_at": "2026-08-02T04:18:26.120360+00:00", + "elapsed_s": 41.7, + "iterations": 5, + "total_completion_tokens": 2449, + "total_prompt_tokens": 14081, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Extracted 20 fields from /input/repo/press_release.txt based on the schema in /input/repo/schema.json. All fields were found and correctly typed/scaled. Results are written to /workspace/extraction_results.json." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p2_extract_gemma4-31b-q4_v7", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 41.7, + "iters": 5, + "tokens": { + "completion_total": 2449, + "prompt_total": 14081, + "all_total": 16530 + }, + "throughput": { + "completion_tps_avg": 59.5, + "model_call_wall_s": 41.1, + "tool_call_wall_s": 0.3, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 29.1, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 268.65, + "memory_used_mib_at_start": 65581.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0058, + "cost_usd_upper_bound": 0.0008, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.8393, + "mean_power_w": 489.61, + "mean_sm_util_pct": 98.12, + "max_temp_c": 75.0 + } + }, + { + "run_name": "p2_extract_gemma4-31b-q4_v8", + "task": "p2_extract", + "replicate": 8, + "summary": { + "started_at": "2026-08-02T04:37:50.412655+00:00", + "ended_at": "2026-08-02T04:39:03.477328+00:00", + "elapsed_s": 73.1, + "iterations": 6, + "total_completion_tokens": 4449, + "total_prompt_tokens": 22299, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Extracted 20 financial and operational fields from the Veridyne Networks Q3 2026 press release. All fields matched the provided JSON schema in terms of type and units. Results are stored in /workspace/extraction_results.json." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p2_extract_gemma4-31b-q4_v8", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 73.1, + "iters": 6, + "tokens": { + "completion_total": 4449, + "prompt_total": 22299, + "all_total": 26748 + }, + "throughput": { + "completion_tps_avg": 61.5, + "model_call_wall_s": 72.3, + "tool_call_wall_s": 0.3, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 37.1, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 192.37, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0102, + "cost_usd_upper_bound": 0.0013, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9576, + "mean_power_w": 491.28, + "mean_sm_util_pct": 91.2, + "max_temp_c": 81.0 + } + }, + { + "run_name": "p2_extract_gemma4-31b-q4_v9", + "task": "p2_extract", + "replicate": 9, + "summary": { + "started_at": "2026-08-02T04:18:26.986836+00:00", + "ended_at": "2026-08-02T04:19:42.156470+00:00", + "elapsed_s": 75.2, + "iterations": 6, + "total_completion_tokens": 4449, + "total_prompt_tokens": 22299, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Extracted 20 financial and operational fields from the Veridyne Networks Q3 2026 press release. All fields matched the provided JSON schema in terms of type and units. Results are stored in /workspace/extraction_results.json." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p2_extract_gemma4-31b-q4_v9", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 75.2, + "iters": 6, + "tokens": { + "completion_total": 4449, + "prompt_total": 22299, + "all_total": 26748 + }, + "throughput": { + "completion_tps_avg": 59.8, + "model_call_wall_s": 74.5, + "tool_call_wall_s": 0.3, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 36.8, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 286.15, + "memory_used_mib_at_start": 65581.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0104, + "cost_usd_upper_bound": 0.0014, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9309, + "mean_power_w": 475.67, + "mean_sm_util_pct": 85.53, + "max_temp_c": 74.0 + } + }, + { + "run_name": "p2_extract_gemma4-31b-q4_v10", + "task": "p2_extract", + "replicate": 10, + "summary": { + "started_at": "2026-08-02T04:39:04.573687+00:00", + "ended_at": "2026-08-02T04:40:17.558183+00:00", + "elapsed_s": 73.0, + "iterations": 6, + "total_completion_tokens": 4449, + "total_prompt_tokens": 22299, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Extracted 20 financial and operational fields from the Veridyne Networks Q3 2026 press release. All fields matched the provided JSON schema in terms of type and units. Results are stored in /workspace/extraction_results.json." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p2_extract_gemma4-31b-q4_v10", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 73.0, + "iters": 6, + "tokens": { + "completion_total": 4449, + "prompt_total": 22299, + "all_total": 26748 + }, + "throughput": { + "completion_tps_avg": 61.6, + "model_call_wall_s": 72.2, + "tool_call_wall_s": 0.3, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 37.0, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 169.33, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0101, + "cost_usd_upper_bound": 0.0013, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.8904, + "mean_power_w": 485.19, + "mean_sm_util_pct": 95.64, + "max_temp_c": 82.0 + } + }, + { + "run_name": "p2_ci_gemma4-31b-q4_v1", + "task": "p2_ci", + "replicate": 1, + "summary": { + "started_at": "2026-08-02T02:10:13.425486+00:00", + "ended_at": "2026-08-02T02:11:41.932894+00:00", + "elapsed_s": 88.5, + "iterations": 29, + "total_completion_tokens": 4141, + "total_prompt_tokens": 234185, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Fixed CI failures in discountkit. \n1. Resolved ruff linting errors by removing an unused import and updating type annotations to Python 3.11 standards.\n2. Fixed a bug in `discount_amount` where the percent-off value was not divided by 100.\n3. Updated `test_cart.py` to align with the v0.3.0 breaking change regarding `Cart.line_count()`.\n4. Documented all changes in `CHANGELOG.md` and `diagnosis.md`, and added an ADR for the test update.\n5. Verified that `ruff check src/ tests/` and `pytest -q` both pass.\n6. Tagged the release as v0.3.2." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p2_ci_gemma4-31b-q4_v1", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 88.5, + "iters": 29, + "tokens": { + "completion_total": 4141, + "prompt_total": 234185, + "all_total": 238326 + }, + "throughput": { + "completion_tps_avg": 50.5, + "model_call_wall_s": 82.0, + "tool_call_wall_s": 4.7, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 34.1, + "first_commit_s": 0.1, + "first_tag_s": 77.9 + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 281.18, + "memory_used_mib_at_start": 66221.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0123, + "cost_usd_upper_bound": 0.0016, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.904, + "mean_power_w": 440.07, + "mean_sm_util_pct": 69.35, + "max_temp_c": 76.0 + } + }, + { + "run_name": "p2_ci_gemma4-31b-q4_v2", + "task": "p2_ci", + "replicate": 2, + "summary": { + "started_at": "2026-08-02T01:07:03.642776+00:00", + "ended_at": "2026-08-02T01:08:38.188117+00:00", + "elapsed_s": 94.5, + "iterations": 29, + "total_completion_tokens": 4551, + "total_prompt_tokens": 230536, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Fixed CI failures in discountkit v0.3.1. \n\n1. Fixed linting errors using `ruff check --fix` (unused imports, sort imports, deprecated type hints).\n2. Fixed `discount_amount` and `apply_discount` to correctly handle percentages (0-100) instead of fractions (0-1).\n3. Updated `test_line_count` to expect the number of lines instead of total quantity, as specified in the v0.3.0 changelog.\n4. Documented all changes in `CHANGELOG.md`, `diagnosis.md`, and an ADR in `decisions/`.\n5. Tagged the final state as `v0.3.2`.\n\nVerification:\n- `ruff check src/ tests/` -> Passed\n- `pytest -q` -> Passed (7 tests, 0 failed)\n" + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p2_ci_gemma4-31b-q4_v2", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 94.5, + "iters": 29, + "tokens": { + "completion_total": 4551, + "prompt_total": 230536, + "all_total": 235087 + }, + "throughput": { + "completion_tps_avg": 51.7, + "model_call_wall_s": 88.0, + "tool_call_wall_s": 4.9, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 36.3, + "first_commit_s": 0.1, + "first_tag_s": 84.1 + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 222.35, + "memory_used_mib_at_start": 65453.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0131, + "cost_usd_upper_bound": 0.0017, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.8995, + "mean_power_w": 427.5, + "mean_sm_util_pct": 81.5, + "max_temp_c": 84.0 + } + }, + { + "run_name": "p2_ci_gemma4-31b-q4_v3", + "task": "p2_ci", + "replicate": 3, + "summary": { + "started_at": "2026-08-02T02:11:42.872352+00:00", + "ended_at": "2026-08-02T02:13:03.154857+00:00", + "elapsed_s": 80.3, + "iterations": 21, + "total_completion_tokens": 3754, + "total_prompt_tokens": 162381, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Fixed CI failures in discountkit. \n- Corrected `discount_amount` to properly divide the percentage by 100.\n- Updated `test_line_count_matches_qty_sum` in `tests/test_cart.py` to expect the number of lines instead of total quantity, in accordance with the v0.3.0 breaking change documented in the CHANGELOG.\n- Fixed all `ruff` linting errors, including an unused import and outdated type annotations.\n- Verified that `ruff check src/ tests/` and `pytest -q` both pass.\n- Documented the fixes in `diagnosis.md` and updated `CHANGELOG.md` for the v0.3.2 release.\n\"\"\"\n" + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p2_ci_gemma4-31b-q4_v3", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 80.3, + "iters": 21, + "tokens": { + "completion_total": 3754, + "prompt_total": 162381, + "all_total": 166135 + }, + "throughput": { + "completion_tps_avg": 50.2, + "model_call_wall_s": 74.8, + "tool_call_wall_s": 4.2, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 31.6, + "first_commit_s": 0.1, + "first_tag_s": 66.1 + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 256.59, + "memory_used_mib_at_start": 66221.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0112, + "cost_usd_upper_bound": 0.0015, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.934, + "mean_power_w": 435.34, + "mean_sm_util_pct": 79.44, + "max_temp_c": 72.0 + } + }, + { + "run_name": "p2_ci_gemma4-31b-q4_v4", + "task": "p2_ci", + "replicate": 4, + "summary": { + "started_at": "2026-08-02T04:40:18.661538+00:00", + "ended_at": "2026-08-02T04:42:11.440790+00:00", + "elapsed_s": 112.8, + "iterations": 33, + "total_completion_tokens": 5568, + "total_prompt_tokens": 268477, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Fixed CI failures by:\n1. Removing an unused `json` import in `src/discountkit/cart.py` to satisfy `ruff`.\n2. Fixing a bug in `discount_amount` where it failed to divide the percentage by 100.\n3. Updating `test_line_count_matches_qty_sum` in `tests/test_cart.py` to expect the number of lines instead of total quantity, consistent with the v0.3.0 breaking change described in the changelog.\n4. Verified that `ruff check src/ tests/` and `pytest -q` now pass.\n5. Updated `CHANGELOG.md` and created `diagnosis.md`.\n6. Tagged the release as v0.3.2." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p2_ci_gemma4-31b-q4_v4", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 112.8, + "iters": 33, + "tokens": { + "completion_total": 5568, + "prompt_total": 268477, + "all_total": 274045 + }, + "throughput": { + "completion_tps_avg": 52.6, + "model_call_wall_s": 105.9, + "tool_call_wall_s": 5.0, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 22.2, + "first_commit_s": 0.1, + "first_tag_s": 90.5 + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 160.46, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0157, + "cost_usd_upper_bound": 0.002, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9309, + "mean_power_w": 413.59, + "mean_sm_util_pct": 71.55, + "max_temp_c": 80.0 + } + }, + { + "run_name": "p2_ci_gemma4-31b-q4_v5", + "task": "p2_ci", + "replicate": 5, + "summary": { + "started_at": "2026-08-02T04:19:43.050568+00:00", + "ended_at": "2026-08-02T04:21:12.212810+00:00", + "elapsed_s": 89.2, + "iterations": 26, + "total_completion_tokens": 4419, + "total_prompt_tokens": 178919, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Fixed CI failures by:\n1. Removing an unused `json` import in `src/discountkit/cart.py` to satisfy Ruff.\n2. Fixing the `discount_amount` implementation in `src/discountkit/discount.py` to correctly divide the percentage by 100, as required by the v0.3.0 API change.\n3. Updating `tests/test_cart.py` to expect `line_count()` to return the number of lines rather than the total quantity, consistent with the v0.3.0 breaking change.\n\nVerified that `ruff check src/ tests/` and `pytest -q` both pass.\nCreated `diagnosis.md` and updated `CHANGELOG.md`.\nTagged the release as v0.3.2." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p2_ci_gemma4-31b-q4_v5", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 89.2, + "iters": 26, + "tokens": { + "completion_total": 4419, + "prompt_total": 178919, + "all_total": 183338 + }, + "throughput": { + "completion_tps_avg": 53.0, + "model_call_wall_s": 83.4, + "tool_call_wall_s": 4.3, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 39.0, + "first_commit_s": 0.1, + "first_tag_s": 79.3 + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 270.34, + "memory_used_mib_at_start": 65581.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0124, + "cost_usd_upper_bound": 0.0016, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.8969, + "mean_power_w": 456.25, + "mean_sm_util_pct": 79.35, + "max_temp_c": 76.0 + } + }, + { + "run_name": "p2_ci_gemma4-31b-q4_v6", + "task": "p2_ci", + "replicate": 6, + "summary": { + "started_at": "2026-08-02T04:42:12.625736+00:00", + "ended_at": "2026-08-02T04:44:03.319889+00:00", + "elapsed_s": 110.7, + "iterations": 34, + "total_completion_tokens": 5323, + "total_prompt_tokens": 277652, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Fixed CI failures in discountkit v0.3.1. \n\n1. Fixed bug in `discount_amount` where percent was not divided by 100.\n2. Updated `test_line_count_matches_qty_sum` to `test_line_count` to expect the number of lines (2) instead of the total quantity (5), aligning with the v0.3.0 breaking change.\n3. Removed unused `json` import in `src/discountkit/cart.py`.\n4. Fixed all Ruff linting errors, including import sorting and updating type hints to Python 3.10+ syntax.\n5. Updated `pyproject.toml` and `CHANGELOG.md` for release v0.3.2.\n6. Provided detailed documentation in `diagnosis.md` and a decision record in `decisions/0001-update-line-count-test.md`.\n\nVerified:\n- `ruff check src/ tests/` exits 0.\n- `pytest -q` exits 0.\n" + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p2_ci_gemma4-31b-q4_v6", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 110.7, + "iters": 34, + "tokens": { + "completion_total": 5323, + "prompt_total": 277652, + "all_total": 282975 + }, + "throughput": { + "completion_tps_avg": 51.2, + "model_call_wall_s": 104.0, + "tool_call_wall_s": 4.8, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 27.8, + "first_commit_s": 0.1, + "first_tag_s": 93.1 + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 149.81, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0154, + "cost_usd_upper_bound": 0.002, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9485, + "mean_power_w": 427.74, + "mean_sm_util_pct": 85.82, + "max_temp_c": 78.0 + } + }, + { + "run_name": "p2_ci_gemma4-31b-q4_v7", + "task": "p2_ci", + "replicate": 7, + "summary": { + "started_at": "2026-08-02T04:21:13.158823+00:00", + "ended_at": "2026-08-02T04:22:50.837763+00:00", + "elapsed_s": 97.7, + "iterations": 31, + "total_completion_tokens": 4662, + "total_prompt_tokens": 225919, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Fixed CI failures in the discountkit project.\n- Resolved a `ruff` linting error by removing an unused import and modernizing type hints.\n- Fixed a logic bug in `discount_amount` where the discount percentage was not being divided by 100.\n- Updated an outdated test in `tests/test_cart.py` to match the v0.3.0 breaking change where `line_count()` now returns the number of lines instead of the total item quantity.\n- Verified that `ruff check src/ tests/` and `pytest -q` both pass.\n- Documented changes in `CHANGELOG.md` and `diagnosis.md`.\n- Tagged the final release as `v0.3.2`." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p2_ci_gemma4-31b-q4_v7", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 97.7, + "iters": 31, + "tokens": { + "completion_total": 4662, + "prompt_total": 225919, + "all_total": 230581 + }, + "throughput": { + "completion_tps_avg": 51.1, + "model_call_wall_s": 91.2, + "tool_call_wall_s": 4.7, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 29.5, + "first_commit_s": 0.1, + "first_tag_s": 84.9 + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 253.21, + "memory_used_mib_at_start": 65581.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0136, + "cost_usd_upper_bound": 0.0018, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9212, + "mean_power_w": 445.14, + "mean_sm_util_pct": 79.42, + "max_temp_c": 72.0 + } + }, + { + "run_name": "p2_ci_gemma4-31b-q4_v8", + "task": "p2_ci", + "replicate": 8, + "summary": { + "started_at": "2026-08-02T04:44:04.501555+00:00", + "ended_at": "2026-08-02T04:45:28.402160+00:00", + "elapsed_s": 83.9, + "iterations": 21, + "total_completion_tokens": 4121, + "total_prompt_tokens": 142741, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Fixed CI failures in the `discountkit` project. \n\nKey fixes:\n- Removed unused `json` import in `src/discountkit/cart.py` to satisfy Ruff.\n- Corrected `discount_amount` and `apply_discount` in `src/discountkit/discount.py` to divide the percentage by 100, aligning the implementation with the v0.3.0 API change.\n- Updated `test_line_count_matches_qty_sum` in `tests/test_cart.py` to expect the number of lines instead of the total quantity, as specified in the v0.3.0 breaking changes.\n\nVerified with:\n- `ruff check src/ tests/` -> 0 errors.\n- `pytest -q` -> 0 failed.\n\nFinal release tagged as v0.3.2." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p2_ci_gemma4-31b-q4_v8", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 83.9, + "iters": 21, + "tokens": { + "completion_total": 4121, + "prompt_total": 142741, + "all_total": 146862 + }, + "throughput": { + "completion_tps_avg": 52.5, + "model_call_wall_s": 78.6, + "tool_call_wall_s": 4.1, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 47.6, + "first_commit_s": 0.1, + "first_tag_s": 71.6 + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 135.73, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0117, + "cost_usd_upper_bound": 0.0015, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9535, + "mean_power_w": 439.8, + "mean_sm_util_pct": 71.47, + "max_temp_c": 77.0 + } + }, + { + "run_name": "p2_ci_gemma4-31b-q4_v9", + "task": "p2_ci", + "replicate": 9, + "summary": { + "started_at": "2026-08-02T04:22:51.768938+00:00", + "ended_at": "2026-08-02T04:24:38.727620+00:00", + "elapsed_s": 107.0, + "iterations": 24, + "total_completion_tokens": 4328, + "total_prompt_tokens": 158816, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Fixed CI failures by:\n1. Removing an unused `json` import in `src/discountkit/cart.py`.\n2. Fixing the `discount_amount` calculation in `src/discountkit/discount.py` to correctly handle percent-based discounts (dividing by 100).\n3. Updating `test_line_count_matches_qty_sum` in `tests/test_cart.py` to reflect the v0.3.0 breaking change where `line_count()` now returns the number of lines instead of total quantity.\n4. Applying automatic Ruff fixes for import sorting and modernization of type annotations.\n\nVerified that `ruff check src/ tests/` and `pytest -q` both pass.\nUpdated `CHANGELOG.md`, created `diagnosis.md`, and tagged the release as `v0.3.2`.\n" + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p2_ci_gemma4-31b-q4_v9", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 107.0, + "iters": 24, + "tokens": { + "completion_total": 4328, + "prompt_total": 158816, + "all_total": 163144 + }, + "throughput": { + "completion_tps_avg": 42.7, + "model_call_wall_s": 101.4, + "tool_call_wall_s": 4.1, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 23.1, + "first_commit_s": 0.1, + "first_tag_s": 95.5 + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 236.06, + "memory_used_mib_at_start": 65581.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0149, + "cost_usd_upper_bound": 0.0019, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9813, + "mean_power_w": 429.32, + "mean_sm_util_pct": 76.0, + "max_temp_c": 74.0 + } + }, + { + "run_name": "p2_ci_gemma4-31b-q4_v10", + "task": "p2_ci", + "replicate": 10, + "summary": { + "started_at": "2026-08-02T04:45:29.615869+00:00", + "ended_at": "2026-08-02T04:46:55.482798+00:00", + "elapsed_s": 85.9, + "iterations": 27, + "total_completion_tokens": 4115, + "total_prompt_tokens": 162685, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Fixed CI failure by:\n1. Removing an unused `json` import in `src/discountkit/cart.py` (fixed `ruff` failure).\n2. Correcting the calculation in `discount_amount` to divide the percentage by 100 (fixed `pytest` failures in `test_discount_amount` and `test_apply_discount`).\n3. Updating `test_line_count` in `tests/test_cart.py` to expect the number of lines instead of the sum of quantities, aligning with the breaking change introduced in v0.3.0 as documented in the changelog.\n\nAll verification steps passed:\n- `ruff check src/ tests/` exited with code 0.\n- `pytest -q` passed all tests with exit code 0.\n\nDeliverables created:\n- Fixed source code and tests.\n- Updated `CHANGELOG.md` with v0.3.2 release notes.\n- Created `diagnosis.md` detailing the fixes.\n- Created `decisions/adr-1.md` explaining the test update.\n- Tagged the final commit as `v0.3.2`." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p2_ci_gemma4-31b-q4_v10", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 85.9, + "iters": 27, + "tokens": { + "completion_total": 4115, + "prompt_total": 162685, + "all_total": 166800 + }, + "throughput": { + "completion_tps_avg": 51.5, + "model_call_wall_s": 79.8, + "tool_call_wall_s": 4.4, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 31.1, + "first_commit_s": 0.1, + "first_tag_s": 72.8 + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 141.15, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0119, + "cost_usd_upper_bound": 0.0015, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9313, + "mean_power_w": 444.87, + "mean_sm_util_pct": 84.88, + "max_temp_c": 78.0 + } + }, + { + "run_name": "p2_hallucination_gemma4-31b-q4_v1", + "task": "p2_hallucination", + "replicate": 1, + "summary": { + "started_at": "2026-08-02T01:08:38.973119+00:00", + "ended_at": "2026-08-02T01:10:34.526607+00:00", + "elapsed_s": 115.6, + "iterations": 16, + "total_completion_tokens": 6454, + "total_prompt_tokens": 121925, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Verified 15 issues from the code-review report against the logalyzer codebase. Found 6 real issues and 9 fabricated issues. The final results are recorded in /workspace/triage.json." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p2_hallucination_gemma4-31b-q4_v1", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 115.6, + "iters": 16, + "tokens": { + "completion_total": 6454, + "prompt_total": 121925, + "all_total": 128379 + }, + "throughput": { + "completion_tps_avg": 56.8, + "model_call_wall_s": 113.7, + "tool_call_wall_s": 0.9, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 105.4, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 220.81, + "memory_used_mib_at_start": 65453.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0161, + "cost_usd_upper_bound": 0.0021, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9516, + "mean_power_w": 488.6, + "mean_sm_util_pct": 91.83, + "max_temp_c": 83.0 + } + }, + { + "run_name": "p2_hallucination_gemma4-31b-q4_v2", + "task": "p2_hallucination", + "replicate": 2, + "summary": { + "started_at": "2026-08-02T02:13:04.107174+00:00", + "ended_at": "2026-08-02T02:14:58.207247+00:00", + "elapsed_s": 114.1, + "iterations": 17, + "total_completion_tokens": 6154, + "total_prompt_tokens": 128689, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Verified 15 issues from issue_report.md. 6 were found to be real and 9 were fabricated. Results are written to /workspace/triage.json." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p2_hallucination_gemma4-31b-q4_v2", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 114.1, + "iters": 17, + "tokens": { + "completion_total": 6154, + "prompt_total": 128689, + "all_total": 134843 + }, + "throughput": { + "completion_tps_avg": 54.9, + "model_call_wall_s": 112.1, + "tool_call_wall_s": 1.0, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 98.6, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 241.92, + "memory_used_mib_at_start": 66221.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0158, + "cost_usd_upper_bound": 0.0021, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9202, + "mean_power_w": 481.27, + "mean_sm_util_pct": 93.09, + "max_temp_c": 74.0 + } + }, + { + "run_name": "p2_hallucination_gemma4-31b-q4_v3", + "task": "p2_hallucination", + "replicate": 3, + "summary": { + "started_at": "2026-08-02T01:10:35.323382+00:00", + "ended_at": "2026-08-02T01:11:07.770827+00:00", + "elapsed_s": 32.4, + "iterations": 9, + "total_completion_tokens": 1494, + "total_prompt_tokens": 43015, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "model_stopped", + "done_summary": "(model stopped without explicit done())" + }, + "verdict": "MISSING_OUTPUT", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p2_hallucination_gemma4-31b-q4_v3", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 32.4, + "iters": 9, + "tokens": { + "completion_total": 1494, + "prompt_total": 43015, + "all_total": 44509 + }, + "throughput": { + "completion_tps_avg": 47.5, + "model_call_wall_s": 31.4, + "tool_call_wall_s": 0.5, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": null, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 225.87, + "memory_used_mib_at_start": 65453.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0045, + "cost_usd_upper_bound": 0.0006, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.7716, + "mean_power_w": 417.97, + "mean_sm_util_pct": 83.33, + "max_temp_c": 83.0 + } + }, + { + "run_name": "p2_hallucination_gemma4-31b-q4_v4", + "task": "p2_hallucination", + "replicate": 4, + "summary": { + "started_at": "2026-08-02T04:46:56.782076+00:00", + "ended_at": "2026-08-02T04:47:52.790784+00:00", + "elapsed_s": 56.0, + "iterations": 13, + "total_completion_tokens": 2921, + "total_prompt_tokens": 78794, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "model_stopped", + "done_summary": "(model stopped without explicit done())" + }, + "verdict": "MISSING_OUTPUT", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p2_hallucination_gemma4-31b-q4_v4", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 56.0, + "iters": 13, + "tokens": { + "completion_total": 2921, + "prompt_total": 78794, + "all_total": 81715 + }, + "throughput": { + "completion_tps_avg": 53.6, + "model_call_wall_s": 54.5, + "tool_call_wall_s": 0.8, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": null, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 134.74, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0078, + "cost_usd_upper_bound": 0.001, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.8929, + "mean_power_w": 413.1, + "mean_sm_util_pct": 74.64, + "max_temp_c": 81.0 + } + }, + { + "run_name": "p2_hallucination_gemma4-31b-q4_v5", + "task": "p2_hallucination", + "replicate": 5, + "summary": { + "started_at": "2026-08-02T04:24:39.643695+00:00", + "ended_at": "2026-08-02T04:25:35.406774+00:00", + "elapsed_s": 55.8, + "iterations": 13, + "total_completion_tokens": 2921, + "total_prompt_tokens": 78794, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "model_stopped", + "done_summary": "(model stopped without explicit done())" + }, + "verdict": "MISSING_OUTPUT", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p2_hallucination_gemma4-31b-q4_v5", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 55.8, + "iters": 13, + "tokens": { + "completion_total": 2921, + "prompt_total": 78794, + "all_total": 81715 + }, + "throughput": { + "completion_tps_avg": 53.8, + "model_call_wall_s": 54.3, + "tool_call_wall_s": 0.8, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": null, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 494.63, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0077, + "cost_usd_upper_bound": 0.001, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.8961, + "mean_power_w": 451.88, + "mean_sm_util_pct": 79.09, + "max_temp_c": 72.0 + } + }, + { + "run_name": "p2_hallucination_gemma4-31b-q4_v6", + "task": "p2_hallucination", + "replicate": 6, + "summary": { + "started_at": "2026-08-02T04:47:54.074833+00:00", + "ended_at": "2026-08-02T04:49:24.325525+00:00", + "elapsed_s": 90.3, + "iterations": 15, + "total_completion_tokens": 4894, + "total_prompt_tokens": 105469, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Verified 15 issues in issue_report.md. 6 were found to be real, and 9 were fabricated. The results have been written to /workspace/triage.json." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p2_hallucination_gemma4-31b-q4_v6", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 90.3, + "iters": 15, + "tokens": { + "completion_total": 4894, + "prompt_total": 105469, + "all_total": 110363 + }, + "throughput": { + "completion_tps_avg": 55.3, + "model_call_wall_s": 88.5, + "tool_call_wall_s": 0.9, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 79.7, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 136.45, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0125, + "cost_usd_upper_bound": 0.0016, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9413, + "mean_power_w": 479.87, + "mean_sm_util_pct": 92.06, + "max_temp_c": 79.0 + } + }, + { + "run_name": "p2_hallucination_gemma4-31b-q4_v7", + "task": "p2_hallucination", + "replicate": 7, + "summary": { + "started_at": "2026-08-02T04:25:36.349028+00:00", + "ended_at": "2026-08-02T04:26:09.911272+00:00", + "elapsed_s": 33.6, + "iterations": 9, + "total_completion_tokens": 1494, + "total_prompt_tokens": 43015, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "model_stopped", + "done_summary": "(model stopped without explicit done())" + }, + "verdict": "MISSING_OUTPUT", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p2_hallucination_gemma4-31b-q4_v7", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 33.6, + "iters": 9, + "tokens": { + "completion_total": 1494, + "prompt_total": 43015, + "all_total": 44509 + }, + "throughput": { + "completion_tps_avg": 45.9, + "model_call_wall_s": 32.6, + "tool_call_wall_s": 0.5, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": null, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 499.87, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0047, + "cost_usd_upper_bound": 0.0006, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.8929, + "mean_power_w": 411.38, + "mean_sm_util_pct": 57.86, + "max_temp_c": 71.0 + } + }, + { + "run_name": "p2_hallucination_gemma4-31b-q4_v8", + "task": "p2_hallucination", + "replicate": 8, + "summary": { + "started_at": "2026-08-02T04:49:25.557348+00:00", + "ended_at": "2026-08-02T04:51:20.887576+00:00", + "elapsed_s": 115.3, + "iterations": 16, + "total_completion_tokens": 6454, + "total_prompt_tokens": 121925, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Verified 15 issues from the code-review report against the logalyzer codebase. Found 6 real issues and 9 fabricated issues. The final results are recorded in /workspace/triage.json." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p2_hallucination_gemma4-31b-q4_v8", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 115.3, + "iters": 16, + "tokens": { + "completion_total": 6454, + "prompt_total": 121925, + "all_total": 128379 + }, + "throughput": { + "completion_tps_avg": 56.9, + "model_call_wall_s": 113.5, + "tool_call_wall_s": 0.9, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 105.4, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 181.93, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.016, + "cost_usd_upper_bound": 0.0021, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.954, + "mean_power_w": 478.96, + "mean_sm_util_pct": 91.7, + "max_temp_c": 80.0 + } + }, + { + "run_name": "p2_hallucination_gemma4-31b-q4_v9", + "task": "p2_hallucination", + "replicate": 9, + "summary": { + "started_at": "2026-08-02T04:26:10.881643+00:00", + "ended_at": "2026-08-02T04:28:01.532385+00:00", + "elapsed_s": 110.7, + "iterations": 17, + "total_completion_tokens": 6152, + "total_prompt_tokens": 133050, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Verified 15 issues from the preliminary report. 6 were found to be real and 9 were found to be fabricated. The results are stored in /workspace/triage.json." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p2_hallucination_gemma4-31b-q4_v9", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 110.7, + "iters": 17, + "tokens": { + "completion_total": 6152, + "prompt_total": 133050, + "all_total": 139202 + }, + "throughput": { + "completion_tps_avg": 56.6, + "model_call_wall_s": 108.6, + "tool_call_wall_s": 1.0, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 99.6, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 499.93, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0154, + "cost_usd_upper_bound": 0.002, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9485, + "mean_power_w": 480.52, + "mean_sm_util_pct": 92.18, + "max_temp_c": 73.0 + } + }, + { + "run_name": "p2_hallucination_gemma4-31b-q4_v10", + "task": "p2_hallucination", + "replicate": 10, + "summary": { + "started_at": "2026-08-02T04:51:22.150921+00:00", + "ended_at": "2026-08-02T04:53:25.952847+00:00", + "elapsed_s": 123.8, + "iterations": 19, + "total_completion_tokens": 6853, + "total_prompt_tokens": 159756, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Verified 15 issues from the preliminary report. 6 were found to be real and 9 were found to be fabricated. The results are detailed in /workspace/triage.json." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p2_hallucination_gemma4-31b-q4_v10", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 123.8, + "iters": 19, + "tokens": { + "completion_total": 6853, + "prompt_total": 159756, + "all_total": 166609 + }, + "throughput": { + "completion_tps_avg": 56.4, + "model_call_wall_s": 121.5, + "tool_call_wall_s": 1.2, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 105.7, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 138.01, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0172, + "cost_usd_upper_bound": 0.0022, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9289, + "mean_power_w": 476.4, + "mean_sm_util_pct": 91.5, + "max_temp_c": 81.0 + } + }, + { + "run_name": "p2_triage_gemma4-31b-q4_v1", + "task": "p2_triage", + "replicate": 1, + "summary": { + "started_at": "2026-08-02T02:14:59.166581+00:00", + "ended_at": "2026-08-02T02:16:47.170718+00:00", + "elapsed_s": 108.0, + "iterations": 8, + "total_completion_tokens": 6211, + "total_prompt_tokens": 57359, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Triaged 30 support tickets, classifying them into 12 categories with associated urgency levels. Identified two duplicate clusters and recorded non-obvious decisions in ADRs. The final output is a structured JSON file at /workspace/triage_results.json, accompanied by a README and decision logs." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p2_triage_gemma4-31b-q4_v1", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 108.0, + "iters": 8, + "tokens": { + "completion_total": 6211, + "prompt_total": 57359, + "all_total": 63570 + }, + "throughput": { + "completion_tps_avg": 58.1, + "model_call_wall_s": 107.0, + "tool_call_wall_s": 0.5, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 78.1, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 254.58, + "memory_used_mib_at_start": 66221.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.015, + "cost_usd_upper_bound": 0.0019, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9259, + "mean_power_w": 470.89, + "mean_sm_util_pct": 85.81, + "max_temp_c": 75.0 + } + }, + { + "run_name": "p2_triage_gemma4-31b-q4_v2", + "task": "p2_triage", + "replicate": 2, + "summary": { + "started_at": "2026-08-02T01:11:08.557629+00:00", + "ended_at": "2026-08-02T01:13:14.023931+00:00", + "elapsed_s": 125.5, + "iterations": 7, + "total_completion_tokens": 7489, + "total_prompt_tokens": 52650, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Triaged 30 support tickets from /input/repo/tickets.txt. Categorized each ticket, assigned urgency, and identified duplicate clusters. The results are stored in /workspace/triage_results.json, with a detailed approach in /workspace/README.md and a decision record for API limit classification in /workspace/decisions/ADR-001-API-Limit-Classification.md." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p2_triage_gemma4-31b-q4_v2", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 125.5, + "iters": 7, + "tokens": { + "completion_total": 7489, + "prompt_total": 52650, + "all_total": 60139 + }, + "throughput": { + "completion_tps_avg": 60.1, + "model_call_wall_s": 124.6, + "tool_call_wall_s": 0.4, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 99.2, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 259.2, + "memory_used_mib_at_start": 65453.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0174, + "cost_usd_upper_bound": 0.0023, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9562, + "mean_power_w": 499.84, + "mean_sm_util_pct": 97.36, + "max_temp_c": 84.0 + } + }, + { + "run_name": "p2_triage_gemma4-31b-q4_v3", + "task": "p2_triage", + "replicate": 3, + "summary": { + "started_at": "2026-08-02T02:16:48.127387+00:00", + "ended_at": "2026-08-02T02:18:56.402816+00:00", + "elapsed_s": 128.3, + "iterations": 7, + "total_completion_tokens": 7493, + "total_prompt_tokens": 52729, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Triaged 30 support tickets. Categorized each ticket by type and urgency, identified two duplicate clusters (dashboard permissions and API rate limits), and produced a structured JSON output at /workspace/triage_results.json. Additionally provided a README and an ADR explaining the decision to treat rate-limit blocks as active incidents." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p2_triage_gemma4-31b-q4_v3", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 128.3, + "iters": 7, + "tokens": { + "completion_total": 7493, + "prompt_total": 52729, + "all_total": 60222 + }, + "throughput": { + "completion_tps_avg": 58.8, + "model_call_wall_s": 127.4, + "tool_call_wall_s": 0.4, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 99.1, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 269.79, + "memory_used_mib_at_start": 66221.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0178, + "cost_usd_upper_bound": 0.0023, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9353, + "mean_power_w": 489.85, + "mean_sm_util_pct": 91.68, + "max_temp_c": 75.0 + } + }, + { + "run_name": "p2_triage_gemma4-31b-q4_v4", + "task": "p2_triage", + "replicate": 4, + "summary": { + "started_at": "2026-08-02T04:53:27.269927+00:00", + "ended_at": "2026-08-02T04:55:23.178375+00:00", + "elapsed_s": 115.9, + "iterations": 8, + "total_completion_tokens": 6758, + "total_prompt_tokens": 60330, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Triage of 30 support tickets completed. Results are stored in /workspace/triage_results.json. I identified 6 urgent tickets, 2 duplicate clusters, and classified all tickets into the required closed vocabulary. ADRs for specific decisions (rate limiting as an incident and extortion as noise) were created in /workspace/decisions/, and a summary is available in /workspace/README.md." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p2_triage_gemma4-31b-q4_v4", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 115.9, + "iters": 8, + "tokens": { + "completion_total": 6758, + "prompt_total": 60330, + "all_total": 67088 + }, + "throughput": { + "completion_tps_avg": 58.8, + "model_call_wall_s": 114.9, + "tool_call_wall_s": 0.5, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 87.1, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 193.15, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0161, + "cost_usd_upper_bound": 0.0021, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9491, + "mean_power_w": 495.74, + "mean_sm_util_pct": 93.52, + "max_temp_c": 81.0 + } + }, + { + "run_name": "p2_triage_gemma4-31b-q4_v5", + "task": "p2_triage", + "replicate": 5, + "summary": { + "started_at": "2026-08-02T04:28:02.535859+00:00", + "ended_at": "2026-08-02T04:30:13.704057+00:00", + "elapsed_s": 131.2, + "iterations": 4, + "total_completion_tokens": 7754, + "total_prompt_tokens": 29427, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Triaged 30 tickets from /input/repo/tickets.txt. Classified each into one of 12 categories and assigned urgency. Identified two duplicate clusters (sharing issues and API rate limits). Produced /workspace/triage_results.json and /workspace/README.md." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p2_triage_gemma4-31b-q4_v5", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 131.2, + "iters": 4, + "tokens": { + "completion_total": 7754, + "prompt_total": 29427, + "all_total": 37181 + }, + "throughput": { + "completion_tps_avg": 59.3, + "model_call_wall_s": 130.7, + "tool_call_wall_s": 0.2, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 110.5, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 308.81, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0182, + "cost_usd_upper_bound": 0.0024, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9527, + "mean_power_w": 494.15, + "mean_sm_util_pct": 96.46, + "max_temp_c": 76.0 + } + }, + { + "run_name": "p2_triage_gemma4-31b-q4_v6", + "task": "p2_triage", + "replicate": 6, + "summary": { + "started_at": "2026-08-02T04:55:24.499448+00:00", + "ended_at": "2026-08-02T04:57:00.403861+00:00", + "elapsed_s": 95.9, + "iterations": 9, + "total_completion_tokens": 5565, + "total_prompt_tokens": 63792, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Triage of 30 support tickets completed. Results are structured in `/workspace/triage_results.json`, with a summary of approach in `/workspace/README.md` and detailed reasoning for ambiguous cases in `/workspace/decisions/`. Identified 5 urgent issues, 2 duplicate clusters, and 4 noise/spam tickets." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p2_triage_gemma4-31b-q4_v6", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 95.9, + "iters": 9, + "tokens": { + "completion_total": 5565, + "prompt_total": 63792, + "all_total": 69357 + }, + "throughput": { + "completion_tps_avg": 58.7, + "model_call_wall_s": 94.8, + "tool_call_wall_s": 0.6, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 66.2, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 175.25, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0133, + "cost_usd_upper_bound": 0.0017, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9385, + "mean_power_w": 486.46, + "mean_sm_util_pct": 84.53, + "max_temp_c": 80.0 + } + }, + { + "run_name": "p2_triage_gemma4-31b-q4_v7", + "task": "p2_triage", + "replicate": 7, + "summary": { + "started_at": "2026-08-02T04:30:14.695255+00:00", + "ended_at": "2026-08-02T04:32:15.491709+00:00", + "elapsed_s": 120.8, + "iterations": 9, + "total_completion_tokens": 7105, + "total_prompt_tokens": 73842, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Triaged 30 support tickets. Created `/workspace/triage_results.json` containing classifications, urgency levels, and duplicate clusters. Provided a `/workspace/README.md` summary and detailed the reasoning for non-obvious classifications in `/workspace/decisions/` (ADRs)." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p2_triage_gemma4-31b-q4_v7", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 120.8, + "iters": 9, + "tokens": { + "completion_total": 7105, + "prompt_total": 73842, + "all_total": 80947 + }, + "throughput": { + "completion_tps_avg": 59.4, + "model_call_wall_s": 119.7, + "tool_call_wall_s": 0.6, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 92.3, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 499.83, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0168, + "cost_usd_upper_bound": 0.0022, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.952, + "mean_power_w": 494.35, + "mean_sm_util_pct": 96.04, + "max_temp_c": 74.0 + } + }, + { + "run_name": "p2_triage_gemma4-31b-q4_v8", + "task": "p2_triage", + "replicate": 8, + "summary": { + "started_at": "2026-08-02T04:57:01.757378+00:00", + "ended_at": "2026-08-02T04:58:46.442628+00:00", + "elapsed_s": 104.7, + "iterations": 5, + "total_completion_tokens": 6262, + "total_prompt_tokens": 28430, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Triaged 30 support tickets. Classified each into one of the 12 specified categories, assigned urgency levels, and identified two duplicate clusters. Produced a detailed JSON output at `/workspace/triage_results.json` and a summary README at `/workspace/README.md`." + }, + "verdict": "FAIL", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p2_triage_gemma4-31b-q4_v8", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 104.7, + "iters": 5, + "tokens": { + "completion_total": 6262, + "prompt_total": 28430, + "all_total": 34692 + }, + "throughput": { + "completion_tps_avg": 60.2, + "model_call_wall_s": 104.1, + "tool_call_wall_s": 0.3, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 85.4, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 146.31, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0145, + "cost_usd_upper_bound": 0.0019, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9551, + "mean_power_w": 482.51, + "mean_sm_util_pct": 92.76, + "max_temp_c": 80.0 + } + }, + { + "run_name": "p2_triage_gemma4-31b-q4_v9", + "task": "p2_triage", + "replicate": 9, + "summary": { + "started_at": "2026-08-02T04:32:16.501642+00:00", + "ended_at": "2026-08-02T04:33:41.263119+00:00", + "elapsed_s": 84.8, + "iterations": 5, + "total_completion_tokens": 4907, + "total_prompt_tokens": 24039, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Triaged 30 support tickets from `/input/repo/tickets.txt` into 12 predefined categories. Identified 3 urgent issues (production outage, API failure, and account compromise). Detected two duplicate clusters involving dashboard permissions and API rate limits. The results are saved in `/workspace/triage_results.json` and a detailed explanation of the methodology is provided in `/workspace/README.md`." + }, + "verdict": "FAIL", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p2_triage_gemma4-31b-q4_v9", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 84.8, + "iters": 5, + "tokens": { + "completion_total": 4907, + "prompt_total": 24039, + "all_total": 28946 + }, + "throughput": { + "completion_tps_avg": 58.3, + "model_call_wall_s": 84.1, + "tool_call_wall_s": 0.3, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 44.4, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 500.05, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0118, + "cost_usd_upper_bound": 0.0015, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9434, + "mean_power_w": 477.91, + "mean_sm_util_pct": 91.76, + "max_temp_c": 73.0 + } + }, + { + "run_name": "p2_triage_gemma4-31b-q4_v10", + "task": "p2_triage", + "replicate": 10, + "summary": { + "started_at": "2026-08-02T04:58:47.840205+00:00", + "ended_at": "2026-08-02T05:00:43.483473+00:00", + "elapsed_s": 115.6, + "iterations": 7, + "total_completion_tokens": 6794, + "total_prompt_tokens": 47178, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Triage of 30 support tickets completed. Results are stored in /workspace/triage_results.json. I categorized each ticket, assigned urgency, and identified duplicate clusters. Additionally, I provided a README.md summary and a decision log in /workspace/decisions/ADR-001.md explaining the classification of specific high-impact and noise tickets." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p2_triage_gemma4-31b-q4_v10", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 115.6, + "iters": 7, + "tokens": { + "completion_total": 6794, + "prompt_total": 47178, + "all_total": 53972 + }, + "throughput": { + "completion_tps_avg": 59.2, + "model_call_wall_s": 114.8, + "tool_call_wall_s": 0.4, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 77.6, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 193.83, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0161, + "cost_usd_upper_bound": 0.0021, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9516, + "mean_power_w": 490.21, + "mean_sm_util_pct": 94.61, + "max_temp_c": 82.0 + } + }, + { + "run_name": "p3_doc_gemma4-31b-q4_v1", + "task": "p3_doc", + "replicate": 1, + "summary": { + "started_at": "2026-08-02T01:13:14.831432+00:00", + "ended_at": "2026-08-02T01:15:16.584695+00:00", + "elapsed_s": 121.8, + "iterations": 8, + "total_completion_tokens": 7296, + "total_prompt_tokens": 51659, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Produced a 1-page executive brief recommending a 'Pass' on the follow-on investment for Nimbus Logistics. The brief synthesizes contradictions between the company's public press release and internal memos/leaked reports, highlighting valuation repricing, high customer concentration, and a shift from growth to survival. Supporting documentation including key facts, research notes, and decision records are provided." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_doc_gemma4-31b-q4_v1", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 121.8, + "iters": 8, + "tokens": { + "completion_total": 7296, + "prompt_total": 51659, + "all_total": 58955 + }, + "throughput": { + "completion_tps_avg": 60.4, + "model_call_wall_s": 120.7, + "tool_call_wall_s": 0.5, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 62.9, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 202.22, + "memory_used_mib_at_start": 65453.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0169, + "cost_usd_upper_bound": 0.0022, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9442, + "mean_power_w": 488.83, + "mean_sm_util_pct": 96.71, + "max_temp_c": 84.0 + } + }, + { + "run_name": "p3_doc_gemma4-31b-q4_v2", + "task": "p3_doc", + "replicate": 2, + "summary": { + "started_at": "2026-08-02T02:18:57.378753+00:00", + "ended_at": "2026-08-02T02:21:01.214442+00:00", + "elapsed_s": 123.8, + "iterations": 8, + "total_completion_tokens": 7296, + "total_prompt_tokens": 51659, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Produced a 1-page executive brief recommending a 'Pass' on the follow-on investment for Nimbus Logistics. The brief synthesizes contradictions between the company's public press release and internal memos/leaked reports, highlighting valuation repricing, high customer concentration, and a shift from growth to survival. Supporting documentation including key facts, research notes, and decision records are provided." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_doc_gemma4-31b-q4_v2", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 123.8, + "iters": 8, + "tokens": { + "completion_total": 7296, + "prompt_total": 51659, + "all_total": 58955 + }, + "throughput": { + "completion_tps_avg": 59.4, + "model_call_wall_s": 122.8, + "tool_call_wall_s": 0.5, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 62.4, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 230.18, + "memory_used_mib_at_start": 66221.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0172, + "cost_usd_upper_bound": 0.0022, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9693, + "mean_power_w": 480.4, + "mean_sm_util_pct": 92.72, + "max_temp_c": 78.0 + } + }, + { + "run_name": "p3_doc_gemma4-31b-q4_v3", + "task": "p3_doc", + "replicate": 3, + "summary": { + "started_at": "2026-08-02T01:15:17.434478+00:00", + "ended_at": "2026-08-02T01:17:10.700188+00:00", + "elapsed_s": 113.3, + "iterations": 10, + "total_completion_tokens": 6632, + "total_prompt_tokens": 67272, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Produced a 1-page executive brief recommending a 'Pass' on follow-on investment in Nimbus Logistics, Inc. The brief synthesizes data from five sources, highlighting critical tensions between the company's public growth narrative and internal austerity measures, including ARR inflation, customer concentration risk, and planned headcount reductions. Supporting documentation includes a key-facts list, research notes, and a decision record." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_doc_gemma4-31b-q4_v3", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 113.3, + "iters": 10, + "tokens": { + "completion_total": 6632, + "prompt_total": 67272, + "all_total": 73904 + }, + "throughput": { + "completion_tps_avg": 59.2, + "model_call_wall_s": 112.1, + "tool_call_wall_s": 0.6, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 50.8, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 215.59, + "memory_used_mib_at_start": 65453.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0157, + "cost_usd_upper_bound": 0.002, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9709, + "mean_power_w": 469.09, + "mean_sm_util_pct": 89.09, + "max_temp_c": 83.0 + } + }, + { + "run_name": "p3_doc_gemma4-31b-q4_v4", + "task": "p3_doc", + "replicate": 4, + "summary": { + "started_at": "2026-08-02T05:00:44.884278+00:00", + "ended_at": "2026-08-02T05:02:40.117753+00:00", + "elapsed_s": 115.2, + "iterations": 9, + "total_completion_tokens": 6893, + "total_prompt_tokens": 59862, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Produced a comprehensive investment evaluation for Nimbus Logistics, Inc. The primary deliverable is a 1-page executive brief (`brief.md`) recommending a 'Pass' on the follow-on investment due to management's lack of transparency and extreme customer concentration. Supporting documentation includes a list of key facts (`key-facts.md`), detailed research notes (`research/notes.md`), a decision record (`decisions/adr-01.md`), and a record of rejected perspectives (`research/dead-ends.md`)." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_doc_gemma4-31b-q4_v4", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 115.2, + "iters": 9, + "tokens": { + "completion_total": 6893, + "prompt_total": 59862, + "all_total": 66755 + }, + "throughput": { + "completion_tps_avg": 60.4, + "model_call_wall_s": 114.1, + "tool_call_wall_s": 0.6, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 45.6, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 173.6, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.016, + "cost_usd_upper_bound": 0.0021, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9549, + "mean_power_w": 497.4, + "mean_sm_util_pct": 97.57, + "max_temp_c": 78.0 + } + }, + { + "run_name": "p3_doc_gemma4-31b-q4_v5", + "task": "p3_doc", + "replicate": 5, + "summary": { + "started_at": "2026-08-02T04:33:42.316332+00:00", + "ended_at": "2026-08-02T04:35:11.926418+00:00", + "elapsed_s": 89.6, + "iterations": 9, + "total_completion_tokens": 5286, + "total_prompt_tokens": 51163, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Produced a 1-page executive brief recommending that the investor PASS on a follow-on investment in Nimbus Logistics. The recommendation is based on significant discrepancies between the company's public growth narrative and its internal survival-mode restructuring, including adjusted ARR figures, deferred international expansion, and undisclosed headcount reductions. Supporting files including key facts, research notes, and decision records are provided in the workspace." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_doc_gemma4-31b-q4_v5", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 89.6, + "iters": 9, + "tokens": { + "completion_total": 5286, + "prompt_total": 51163, + "all_total": 56449 + }, + "throughput": { + "completion_tps_avg": 59.7, + "model_call_wall_s": 88.5, + "tool_call_wall_s": 0.6, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 35.7, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 328.26, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0124, + "cost_usd_upper_bound": 0.0016, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.8929, + "mean_power_w": 487.26, + "mean_sm_util_pct": 93.88, + "max_temp_c": 73.0 + } + }, + { + "run_name": "p3_doc_gemma4-31b-q4_v6", + "task": "p3_doc", + "replicate": 6, + "summary": { + "started_at": "2026-08-02T05:02:41.589339+00:00", + "ended_at": "2026-08-02T05:04:11.430594+00:00", + "elapsed_s": 89.8, + "iterations": 9, + "total_completion_tokens": 5285, + "total_prompt_tokens": 51163, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Produced a 1-page executive brief recommending that the investor PASS on a follow-on investment in Nimbus Logistics. The recommendation is based on significant discrepancies between the company's public expansion narrative and its internal strategic retraction, as well as misleading ARR reporting and high customer concentration. Supporting documentation includes a detailed key-facts list, research notes, and a framing decision record." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_doc_gemma4-31b-q4_v6", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 89.8, + "iters": 9, + "tokens": { + "completion_total": 5285, + "prompt_total": 51163, + "all_total": 56448 + }, + "throughput": { + "completion_tps_avg": 59.6, + "model_call_wall_s": 88.7, + "tool_call_wall_s": 0.6, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 35.8, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 148.58, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0125, + "cost_usd_upper_bound": 0.0016, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9465, + "mean_power_w": 465.8, + "mean_sm_util_pct": 86.56, + "max_temp_c": 72.0 + } + }, + { + "run_name": "p3_doc_gemma4-31b-q4_v7", + "task": "p3_doc", + "replicate": 7, + "summary": { + "started_at": "2026-08-02T04:35:12.972841+00:00", + "ended_at": "2026-08-02T04:36:45.314414+00:00", + "elapsed_s": 92.3, + "iterations": 9, + "total_completion_tokens": 5286, + "total_prompt_tokens": 51163, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Produced a 1-page executive brief recommending that the investor PASS on a follow-on investment in Nimbus Logistics. The recommendation is based on significant discrepancies between the company's public growth narrative and its internal survival-mode restructuring, including adjusted ARR figures, deferred international expansion, and undisclosed headcount reductions. Supporting files including key facts, research notes, and decision records are provided in the workspace." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_doc_gemma4-31b-q4_v7", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 92.3, + "iters": 9, + "tokens": { + "completion_total": 5286, + "prompt_total": 51163, + "all_total": 56449 + }, + "throughput": { + "completion_tps_avg": 58.0, + "model_call_wall_s": 91.2, + "tool_call_wall_s": 0.6, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 37.0, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 379.72, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0128, + "cost_usd_upper_bound": 0.0017, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9209, + "mean_power_w": 479.29, + "mean_sm_util_pct": 88.06, + "max_temp_c": 73.0 + } + }, + { + "run_name": "p3_doc_gemma4-31b-q4_v8", + "task": "p3_doc", + "replicate": 8, + "summary": { + "started_at": "2026-08-02T05:04:12.879770+00:00", + "ended_at": "2026-08-02T05:05:43.807968+00:00", + "elapsed_s": 90.9, + "iterations": 9, + "total_completion_tokens": 5285, + "total_prompt_tokens": 51163, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Produced a 1-page executive brief recommending that the investor PASS on a follow-on investment in Nimbus Logistics. The recommendation is based on significant discrepancies between the company's public expansion narrative and its internal strategic retraction, as well as misleading ARR reporting and high customer concentration. Supporting documentation includes a detailed key-facts list, research notes, and a framing decision record." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_doc_gemma4-31b-q4_v8", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 90.9, + "iters": 9, + "tokens": { + "completion_total": 5285, + "prompt_total": 51163, + "all_total": 56448 + }, + "throughput": { + "completion_tps_avg": 58.9, + "model_call_wall_s": 89.8, + "tool_call_wall_s": 0.6, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 35.8, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 181.06, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0126, + "cost_usd_upper_bound": 0.0016, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9351, + "mean_power_w": 482.28, + "mean_sm_util_pct": 92.44, + "max_temp_c": 71.0 + } + }, + { + "run_name": "p3_doc_gemma4-31b-q4_v9", + "task": "p3_doc", + "replicate": 9, + "summary": { + "started_at": "2026-08-02T04:36:46.399014+00:00", + "ended_at": "2026-08-02T04:38:46.542822+00:00", + "elapsed_s": 120.1, + "iterations": 8, + "total_completion_tokens": 7296, + "total_prompt_tokens": 51659, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Produced a 1-page executive brief recommending a 'Pass' on the follow-on investment for Nimbus Logistics. The brief synthesizes contradictions between the company's public press release and internal memos/leaked reports, highlighting valuation repricing, high customer concentration, and a shift from growth to survival. Supporting documentation including key facts, research notes, and decision records are provided." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_doc_gemma4-31b-q4_v9", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 120.1, + "iters": 8, + "tokens": { + "completion_total": 7296, + "prompt_total": 51659, + "all_total": 58955 + }, + "throughput": { + "completion_tps_avg": 61.2, + "model_call_wall_s": 119.1, + "tool_call_wall_s": 0.5, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 62.5, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 505.18, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0167, + "cost_usd_upper_bound": 0.0022, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9575, + "mean_power_w": 488.73, + "mean_sm_util_pct": 92.38, + "max_temp_c": 74.0 + } + }, + { + "run_name": "p3_doc_gemma4-31b-q4_v10", + "task": "p3_doc", + "replicate": 10, + "summary": { + "started_at": "2026-08-02T05:05:45.319999+00:00", + "ended_at": "2026-08-02T05:07:45.361137+00:00", + "elapsed_s": 120.0, + "iterations": 8, + "total_completion_tokens": 7296, + "total_prompt_tokens": 51659, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Produced a 1-page executive brief recommending a 'Pass' on the follow-on investment for Nimbus Logistics. The brief synthesizes contradictions between the company's public press release and internal memos/leaked reports, highlighting valuation repricing, high customer concentration, and a shift from growth to survival. Supporting documentation including key facts, research notes, and decision records are provided." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_doc_gemma4-31b-q4_v10", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 120.0, + "iters": 8, + "tokens": { + "completion_total": 7296, + "prompt_total": 51659, + "all_total": 58955 + }, + "throughput": { + "completion_tps_avg": 61.3, + "model_call_wall_s": 119.0, + "tool_call_wall_s": 0.5, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 62.7, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 173.73, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0167, + "cost_usd_upper_bound": 0.0022, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9583, + "mean_power_w": 486.67, + "mean_sm_util_pct": 95.25, + "max_temp_c": 71.0 + } + }, + { + "run_name": "p3_business_gemma4-31b-q4_v1", + "task": "p3_business", + "replicate": 1, + "summary": { + "started_at": "2026-08-02T02:21:02.195182+00:00", + "ended_at": "2026-08-02T02:22:36.370720+00:00", + "elapsed_s": 94.2, + "iterations": 9, + "total_completion_tokens": 5487, + "total_prompt_tokens": 43691, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "I have completed the independent review of the Borealis Analytics acquisition diligence pack. I recommended a HOLD on the $84M LOI due to a fundamental contradiction between the stated strategic rationale (acquiring a differentiated product) and the integration plan (sunsetting the product in 12 months). Additional concerns include unexpectedly low gross margins (71%), high execution risk regarding the forced relocation of the engineering team, and thin customer diligence (only 5 of 280 customers contacted). I have provided a detailed memo, a full list of concerns, specific diligence asks for the next two weeks, and ADRs for my judgment calls." + }, + "verdict": "FAIL", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_business_gemma4-31b-q4_v1", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 94.2, + "iters": 9, + "tokens": { + "completion_total": 5487, + "prompt_total": 43691, + "all_total": 49178 + }, + "throughput": { + "completion_tps_avg": 59.0, + "model_call_wall_s": 93.1, + "tool_call_wall_s": 0.6, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 42.4, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 235.08, + "memory_used_mib_at_start": 66221.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0131, + "cost_usd_upper_bound": 0.0017, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9554, + "mean_power_w": 458.24, + "mean_sm_util_pct": 78.84, + "max_temp_c": 77.0 + } + }, + { + "run_name": "p3_business_gemma4-31b-q4_v2", + "task": "p3_business", + "replicate": 2, + "summary": { + "started_at": "2026-08-02T01:17:11.532735+00:00", + "ended_at": "2026-08-02T01:18:32.429360+00:00", + "elapsed_s": 80.9, + "iterations": 9, + "total_completion_tokens": 4742, + "total_prompt_tokens": 37120, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "I have completed the independent review of the Borealis Analytics acquisition diligence pack. My primary recommendation is to HOLD the LOI due to a critical discrepancy in the target's reported runway (implied 5.5 months vs claimed 22 months), a strategic contradiction in paying a growth multiple for a product slated for sunset, and overly optimistic staff retention assumptions. I have provided a detailed memo, a comprehensive list of concerns, a specific set of diligence asks for the next two weeks, and an ADR documenting the decision logic." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_business_gemma4-31b-q4_v2", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 80.9, + "iters": 9, + "tokens": { + "completion_total": 4742, + "prompt_total": 37120, + "all_total": 41862 + }, + "throughput": { + "completion_tps_avg": 59.4, + "model_call_wall_s": 79.8, + "tool_call_wall_s": 0.6, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 32.4, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 194.34, + "memory_used_mib_at_start": 65453.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0112, + "cost_usd_upper_bound": 0.0015, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9271, + "mean_power_w": 495.91, + "mean_sm_util_pct": 95.5, + "max_temp_c": 83.0 + } + }, + { + "run_name": "p3_business_gemma4-31b-q4_v3", + "task": "p3_business", + "replicate": 3, + "summary": { + "started_at": "2026-08-02T02:22:37.368829+00:00", + "ended_at": "2026-08-02T02:23:57.088651+00:00", + "elapsed_s": 79.7, + "iterations": 9, + "total_completion_tokens": 4742, + "total_prompt_tokens": 37120, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "I have completed the independent review of the Borealis Analytics acquisition diligence pack. My primary recommendation is to HOLD the LOI due to a critical discrepancy in the target's reported runway (implied 5.5 months vs claimed 22 months), a strategic contradiction in paying a growth multiple for a product slated for sunset, and overly optimistic staff retention assumptions. I have provided a detailed memo, a comprehensive list of concerns, a specific set of diligence asks for the next two weeks, and an ADR documenting the decision logic." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_business_gemma4-31b-q4_v3", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 79.7, + "iters": 9, + "tokens": { + "completion_total": 4742, + "prompt_total": 37120, + "all_total": 41862 + }, + "throughput": { + "completion_tps_avg": 60.3, + "model_call_wall_s": 78.6, + "tool_call_wall_s": 0.6, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 32.0, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 248.35, + "memory_used_mib_at_start": 66221.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0111, + "cost_usd_upper_bound": 0.0014, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.941, + "mean_power_w": 486.72, + "mean_sm_util_pct": 93.75, + "max_temp_c": 74.0 + } + }, + { + "run_name": "p3_business_gemma4-31b-q4_v4", + "task": "p3_business", + "replicate": 4, + "summary": { + "started_at": "2026-08-02T05:07:46.898565+00:00", + "ended_at": "2026-08-02T05:09:18.737029+00:00", + "elapsed_s": 91.8, + "iterations": 9, + "total_completion_tokens": 5487, + "total_prompt_tokens": 43691, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "I have completed the independent review of the Borealis Analytics acquisition diligence pack. I recommended a HOLD on the $84M LOI due to a fundamental contradiction between the stated strategic rationale (acquiring a differentiated product) and the integration plan (sunsetting the product in 12 months). Additional concerns include unexpectedly low gross margins (71%), high execution risk regarding the forced relocation of the engineering team, and thin customer diligence (only 5 of 280 customers contacted). I have provided a detailed memo, a full list of concerns, specific diligence asks for the next two weeks, and ADRs for my judgment calls." + }, + "verdict": "FAIL", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_business_gemma4-31b-q4_v4", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 91.8, + "iters": 9, + "tokens": { + "completion_total": 5487, + "prompt_total": 43691, + "all_total": 49178 + }, + "throughput": { + "completion_tps_avg": 60.5, + "model_call_wall_s": 90.7, + "tool_call_wall_s": 0.6, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 42.7, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 159.45, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0127, + "cost_usd_upper_bound": 0.0017, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9804, + "mean_power_w": 478.13, + "mean_sm_util_pct": 92.42, + "max_temp_c": 70.0 + } + }, + { + "run_name": "p3_business_gemma4-31b-q4_v5", + "task": "p3_business", + "replicate": 5, + "summary": { + "started_at": "2026-08-02T04:38:47.669072+00:00", + "ended_at": "2026-08-02T04:40:19.357782+00:00", + "elapsed_s": 91.7, + "iterations": 9, + "total_completion_tokens": 5487, + "total_prompt_tokens": 43691, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "I have completed the independent review of the Borealis Analytics acquisition diligence pack. I recommended a HOLD on the $84M LOI due to a fundamental contradiction between the stated strategic rationale (acquiring a differentiated product) and the integration plan (sunsetting the product in 12 months). Additional concerns include unexpectedly low gross margins (71%), high execution risk regarding the forced relocation of the engineering team, and thin customer diligence (only 5 of 280 customers contacted). I have provided a detailed memo, a full list of concerns, specific diligence asks for the next two weeks, and ADRs for my judgment calls." + }, + "verdict": "FAIL", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_business_gemma4-31b-q4_v5", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 91.7, + "iters": 9, + "tokens": { + "completion_total": 5487, + "prompt_total": 43691, + "all_total": 49178 + }, + "throughput": { + "completion_tps_avg": 60.6, + "model_call_wall_s": 90.5, + "tool_call_wall_s": 0.6, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 42.5, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 500.05, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0127, + "cost_usd_upper_bound": 0.0017, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9269, + "mean_power_w": 500.08, + "mean_sm_util_pct": 97.22, + "max_temp_c": 73.0 + } + }, + { + "run_name": "p3_business_gemma4-31b-q4_v6", + "task": "p3_business", + "replicate": 6, + "summary": { + "started_at": "2026-08-02T05:09:20.224774+00:00", + "ended_at": "2026-08-02T05:10:46.014436+00:00", + "elapsed_s": 85.8, + "iterations": 9, + "total_completion_tokens": 5087, + "total_prompt_tokens": 38550, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "I have completed the independent review of the Borealis Analytics acquisition deal pack. After analyzing the document, I identified critical financial discrepancies (specifically regarding runway and infrastructure synergies) and a major strategic contradiction regarding the planned sunsetting of the target's flagship product. Consequently, I have recommended a HOLD on the LOI. \n\nDeliverables include:\n- `memo.md`: Executive recommendation and high-level summary.\n- `concerns.md`: Detailed analysis of all red flags.\n- `asks.md`: Specific diligence requests for the next 2 weeks.\n- `decisions/ADR-001.md`: Documentation of the judgment call to hold the deal.\n- `README.md`: Overview of the output." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_business_gemma4-31b-q4_v6", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 85.8, + "iters": 9, + "tokens": { + "completion_total": 5087, + "prompt_total": 38550, + "all_total": 43637 + }, + "throughput": { + "completion_tps_avg": 60.1, + "model_call_wall_s": 84.6, + "tool_call_wall_s": 0.6, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 34.3, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 151.34, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0119, + "cost_usd_upper_bound": 0.0015, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9324, + "mean_power_w": 484.2, + "mean_sm_util_pct": 90.18, + "max_temp_c": 70.0 + } + }, + { + "run_name": "p3_business_gemma4-31b-q4_v7", + "task": "p3_business", + "replicate": 7, + "summary": { + "started_at": "2026-08-02T04:40:20.468921+00:00", + "ended_at": "2026-08-02T04:41:49.116045+00:00", + "elapsed_s": 88.6, + "iterations": 9, + "total_completion_tokens": 5087, + "total_prompt_tokens": 38550, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "I have completed the independent review of the Borealis Analytics acquisition deal pack. After analyzing the document, I identified critical financial discrepancies (specifically regarding runway and infrastructure synergies) and a major strategic contradiction regarding the planned sunsetting of the target's flagship product. Consequently, I have recommended a HOLD on the LOI. \n\nDeliverables include:\n- `memo.md`: Executive recommendation and high-level summary.\n- `concerns.md`: Detailed analysis of all red flags.\n- `asks.md`: Specific diligence requests for the next 2 weeks.\n- `decisions/ADR-001.md`: Documentation of the judgment call to hold the deal.\n- `README.md`: Overview of the output." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_business_gemma4-31b-q4_v7", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 88.6, + "iters": 9, + "tokens": { + "completion_total": 5087, + "prompt_total": 38550, + "all_total": 43637 + }, + "throughput": { + "completion_tps_avg": 58.1, + "model_call_wall_s": 87.5, + "tool_call_wall_s": 0.6, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 35.5, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 505.59, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0123, + "cost_usd_upper_bound": 0.0016, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9594, + "mean_power_w": 465.79, + "mean_sm_util_pct": 84.89, + "max_temp_c": 73.0 + } + }, + { + "run_name": "p3_business_gemma4-31b-q4_v8", + "task": "p3_business", + "replicate": 8, + "summary": { + "started_at": "2026-08-02T05:10:47.550941+00:00", + "ended_at": "2026-08-02T05:12:14.736858+00:00", + "elapsed_s": 87.2, + "iterations": 9, + "total_completion_tokens": 5087, + "total_prompt_tokens": 38550, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "I have completed the independent review of the Borealis Analytics acquisition deal pack. After analyzing the document, I identified critical financial discrepancies (specifically regarding runway and infrastructure synergies) and a major strategic contradiction regarding the planned sunsetting of the target's flagship product. Consequently, I have recommended a HOLD on the LOI. \n\nDeliverables include:\n- `memo.md`: Executive recommendation and high-level summary.\n- `concerns.md`: Detailed analysis of all red flags.\n- `asks.md`: Specific diligence requests for the next 2 weeks.\n- `decisions/ADR-001.md`: Documentation of the judgment call to hold the deal.\n- `README.md`: Overview of the output." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_business_gemma4-31b-q4_v8", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 87.2, + "iters": 9, + "tokens": { + "completion_total": 5087, + "prompt_total": 38550, + "all_total": 43637 + }, + "throughput": { + "completion_tps_avg": 59.1, + "model_call_wall_s": 86.1, + "tool_call_wall_s": 0.6, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 34.3, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 142.47, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0121, + "cost_usd_upper_bound": 0.0016, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9174, + "mean_power_w": 487.37, + "mean_sm_util_pct": 93.59, + "max_temp_c": 70.0 + } + }, + { + "run_name": "p3_business_gemma4-31b-q4_v9", + "task": "p3_business", + "replicate": 9, + "summary": { + "started_at": "2026-08-02T04:41:50.263950+00:00", + "ended_at": "2026-08-02T04:43:18.051807+00:00", + "elapsed_s": 87.8, + "iterations": 11, + "total_completion_tokens": 5171, + "total_prompt_tokens": 49232, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "I have completed the independent review of the Borealis Analytics acquisition diligence pack. I recommended a 'HOLD' on the LOI due to critical inconsistencies in the deal pack, including a synergy claim ($3.2M) that exceeds the target's total implied COGS ($2.64M) and a contradiction between the strategic goal of extending the product surface area and the plan to sunset the product within 12 months. I have provided a detailed memo, a comprehensive list of concerns, specific diligence asks for the next two weeks, and ADRs documenting my judgment calls." + }, + "verdict": "FAIL", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_business_gemma4-31b-q4_v9", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 87.8, + "iters": 11, + "tokens": { + "completion_total": 5171, + "prompt_total": 49232, + "all_total": 54403 + }, + "throughput": { + "completion_tps_avg": 59.8, + "model_call_wall_s": 86.4, + "tool_call_wall_s": 0.7, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 36.3, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 499.99, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0122, + "cost_usd_upper_bound": 0.0016, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9681, + "mean_power_w": 489.8, + "mean_sm_util_pct": 93.06, + "max_temp_c": 73.0 + } + }, + { + "run_name": "p3_business_gemma4-31b-q4_v10", + "task": "p3_business", + "replicate": 10, + "summary": { + "started_at": "2026-08-02T05:12:16.259672+00:00", + "ended_at": "2026-08-02T05:13:41.177482+00:00", + "elapsed_s": 84.9, + "iterations": 9, + "total_completion_tokens": 5087, + "total_prompt_tokens": 38550, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "I have completed the independent review of the Borealis Analytics acquisition deal pack. After analyzing the document, I identified critical financial discrepancies (specifically regarding runway and infrastructure synergies) and a major strategic contradiction regarding the planned sunsetting of the target's flagship product. Consequently, I have recommended a HOLD on the LOI. \n\nDeliverables include:\n- `memo.md`: Executive recommendation and high-level summary.\n- `concerns.md`: Detailed analysis of all red flags.\n- `asks.md`: Specific diligence requests for the next 2 weeks.\n- `decisions/ADR-001.md`: Documentation of the judgment call to hold the deal.\n- `README.md`: Overview of the output." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_business_gemma4-31b-q4_v10", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 84.9, + "iters": 9, + "tokens": { + "completion_total": 5087, + "prompt_total": 38550, + "all_total": 43637 + }, + "throughput": { + "completion_tps_avg": 60.7, + "model_call_wall_s": 83.8, + "tool_call_wall_s": 0.6, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 34.3, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 122.89, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0118, + "cost_usd_upper_bound": 0.0015, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9423, + "mean_power_w": 480.52, + "mean_sm_util_pct": 87.24, + "max_temp_c": 70.0 + } + }, + { + "run_name": "p3_market_gemma4-31b-q4_v1", + "task": "p3_market", + "replicate": 1, + "summary": { + "started_at": "2026-08-02T01:18:33.283109+00:00", + "ended_at": "2026-08-02T01:22:50.667125+00:00", + "elapsed_s": 257.4, + "iterations": 16, + "total_completion_tokens": 8071, + "total_prompt_tokens": 387754, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "I have researched and recommended a password manager for a 50-person engineering team. After evaluating 1Password, Bitwarden, Keeper, Dashlane, and NordPass, I recommended 1Password Business due to its superior developer experience and high adoption likelihood, which are critical for eliminating shadow-IT within 90 days. I provided a full feature/pricing comparison, detailed decision records (ADRs), and a comprehensive list of sources." + }, + "verdict": "STRUCTURAL_PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_market_gemma4-31b-q4_v1", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 257.4, + "iters": 16, + "tokens": { + "completion_total": 8071, + "prompt_total": 387754, + "all_total": 395825 + }, + "throughput": { + "completion_tps_avg": 32.0, + "model_call_wall_s": 251.9, + "tool_call_wall_s": 4.5, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 9.4, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 205.07, + "memory_used_mib_at_start": 65453.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0357, + "cost_usd_upper_bound": 0.0046, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9713, + "mean_power_w": 495.65, + "mean_sm_util_pct": 96.29, + "max_temp_c": 85.0 + } + }, + { + "run_name": "p3_market_gemma4-31b-q4_v2", + "task": "p3_market", + "replicate": 2, + "summary": { + "started_at": "2026-08-02T02:23:58.119775+00:00", + "ended_at": "2026-08-02T02:27:12.186064+00:00", + "elapsed_s": 194.1, + "iterations": 20, + "total_completion_tokens": 9028, + "total_prompt_tokens": 442437, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Researched and recommended a password manager for a 50-person engineering team. Evaluated 1Password, Bitwarden, Dashlane, Keeper, and NordPass. Recommended 1Password Business due to superior developer experience (CLI) and strong SOC2 compliance, with Bitwarden as a runner-up for open-source requirements. All findings are cited in sources.md and summarized in recommendation.md." + }, + "verdict": "STRUCTURAL_PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_market_gemma4-31b-q4_v2", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 194.1, + "iters": 20, + "tokens": { + "completion_total": 9028, + "prompt_total": 442437, + "all_total": 451465 + }, + "throughput": { + "completion_tps_avg": 48.0, + "model_call_wall_s": 188.1, + "tool_call_wall_s": 4.8, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 9.3, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 274.21, + "memory_used_mib_at_start": 66221.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.027, + "cost_usd_upper_bound": 0.0035, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9531, + "mean_power_w": 485.18, + "mean_sm_util_pct": 93.97, + "max_temp_c": 75.0 + } + }, + { + "run_name": "p3_market_gemma4-31b-q4_v3", + "task": "p3_market", + "replicate": 3, + "summary": { + "started_at": "2026-08-02T01:22:51.546562+00:00", + "ended_at": "2026-08-02T01:28:13.595380+00:00", + "elapsed_s": 322.0, + "iterations": 15, + "total_completion_tokens": 8059, + "total_prompt_tokens": 271715, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Successfully researched and recommended a password manager for a 50-person engineering team. Evaluated 5 products (1Password, Bitwarden, Dashlane, Keeper, NordPass) against a set of defined criteria including SSO, SOC2 compliance, and Developer Experience. Recommended 1Password Business for its superior CLI and UX, with Bitwarden Enterprise as a cost-effective, open-source runner-up. All factual claims are cited with public URLs in sources.md and the recommendation memo." + }, + "verdict": "STRUCTURAL_PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_market_gemma4-31b-q4_v3", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 322.0, + "iters": 15, + "tokens": { + "completion_total": 8059, + "prompt_total": 271715, + "all_total": 279774 + }, + "throughput": { + "completion_tps_avg": 25.3, + "model_call_wall_s": 318.3, + "tool_call_wall_s": 2.8, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 20.1, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 499.37, + "memory_used_mib_at_start": 65453.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0447, + "cost_usd_upper_bound": 0.0058, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9783, + "mean_power_w": 489.71, + "mean_sm_util_pct": 92.86, + "max_temp_c": 86.0 + } + }, + { + "run_name": "p3_market_gemma4-31b-q4_v4", + "task": "p3_market", + "replicate": 4, + "summary": { + "started_at": "2026-08-02T05:13:42.768241+00:00", + "ended_at": "2026-08-02T05:16:35.180147+00:00", + "elapsed_s": 172.4, + "iterations": 19, + "total_completion_tokens": 8745, + "total_prompt_tokens": 278772, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Researched and recommended a password manager for a 50-person engineering team. Evaluated 1Password, Bitwarden, Keeper, Dashlane, and NordPass. Recommended 1Password Business for its superior DX (CLI), security (Secret Key), and SSO capabilities, with Bitwarden Enterprise as the runner-up. All deliverables including the recommendation memo, comparison table, ADRs, and sources have been created." + }, + "verdict": "STRUCTURAL_PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_market_gemma4-31b-q4_v4", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 172.4, + "iters": 19, + "tokens": { + "completion_total": 8745, + "prompt_total": 278772, + "all_total": 287517 + }, + "throughput": { + "completion_tps_avg": 53.2, + "model_call_wall_s": 164.4, + "tool_call_wall_s": 6.8, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 7.9, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 123.59, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0239, + "cost_usd_upper_bound": 0.0031, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9571, + "mean_power_w": 476.49, + "mean_sm_util_pct": 90.62, + "max_temp_c": 70.0 + } + }, + { + "run_name": "p3_market_gemma4-31b-q4_v5", + "task": "p3_market", + "replicate": 5, + "summary": { + "started_at": "2026-08-02T04:43:19.259687+00:00", + "ended_at": "2026-08-02T04:45:53.837557+00:00", + "elapsed_s": 154.6, + "iterations": 14, + "total_completion_tokens": 7095, + "total_prompt_tokens": 302302, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "I have researched and recommended a password manager for a 50-person engineering team. I evaluated 1Password, Bitwarden, Dashlane, Keeper, and NordPass against specific criteria including SSO integration, developer experience (CLI), and SOC2 compliance. I recommended 1Password Business ($4,794/year) as the primary choice due to its superior adoption potential and developer tools, with Bitwarden Enterprise as the runner-up for open-source requirements. All claims are cited in sources.md." + }, + "verdict": "STRUCTURAL_PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_market_gemma4-31b-q4_v5", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 154.6, + "iters": 14, + "tokens": { + "completion_total": 7095, + "prompt_total": 302302, + "all_total": 309397 + }, + "throughput": { + "completion_tps_avg": 47.1, + "model_call_wall_s": 150.6, + "tool_call_wall_s": 3.1, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 10.7, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 338.76, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0215, + "cost_usd_upper_bound": 0.0028, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9702, + "mean_power_w": 482.93, + "mean_sm_util_pct": 89.84, + "max_temp_c": 75.0 + } + }, + { + "run_name": "p3_market_gemma4-31b-q4_v6", + "task": "p3_market", + "replicate": 6, + "summary": { + "started_at": "2026-08-02T05:16:36.745073+00:00", + "ended_at": "2026-08-02T05:20:05.199056+00:00", + "elapsed_s": 208.5, + "iterations": 26, + "total_completion_tokens": 8866, + "total_prompt_tokens": 772802, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "I have researched and recommended a password manager for a 50-person engineering team. I evaluated 1Password, Bitwarden, Dashlane, Keeper, and NordPass against criteria including SSO integration, SOC2 compliance, and developer experience. I recommended 1Password Business as the primary choice due to its superior DX and security architecture, with Bitwarden Enterprise as a runner-up for those prioritizing open source and cost. All deliverables including the recommendation memo, comparison matrix, and ADRs have been created in the workspace." + }, + "verdict": "STRUCTURAL_PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_market_gemma4-31b-q4_v6", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 208.5, + "iters": 26, + "tokens": { + "completion_total": 8866, + "prompt_total": 772802, + "all_total": 781668 + }, + "throughput": { + "completion_tps_avg": 44.0, + "model_call_wall_s": 201.4, + "tool_call_wall_s": 5.4, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 9.3, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 161.71, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.029, + "cost_usd_upper_bound": 0.0038, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9832, + "mean_power_w": 468.1, + "mean_sm_util_pct": 82.38, + "max_temp_c": 71.0 + } + }, + { + "run_name": "p3_market_gemma4-31b-q4_v7", + "task": "p3_market", + "replicate": 7, + "summary": { + "started_at": "2026-08-02T04:45:55.064417+00:00", + "ended_at": "2026-08-02T04:49:46.306011+00:00", + "elapsed_s": 231.2, + "iterations": 39, + "total_completion_tokens": 9745, + "total_prompt_tokens": 989064, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "I have completed the research and recommendation for a password manager for the 50-person engineering team. \n\nKey accomplishments:\n1. Evaluated 5 products: 1Password, Bitwarden, Dashlane, Keeper, and NordPass.\n2. Defined a rigorous set of decision criteria focusing on SSO, SOC2 compliance, and Developer Experience (CLI/API).\n3. Recommended **Bitwarden Enterprise** as the primary choice due to its transparent pricing ($6/user/month), open-source nature, and strong engineering tooling.\n4. Provided a detailed 90-day implementation plan to meet the CTO's deadline.\n5. Documented all findings in `recommendation.md`, `comparison.md`, and `sources.md`, with ADRs in `decisions/` for transparency.\n" + }, + "verdict": "STRUCTURAL_PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_market_gemma4-31b-q4_v7", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 231.2, + "iters": 39, + "tokens": { + "completion_total": 9745, + "prompt_total": 989064, + "all_total": 998809 + }, + "throughput": { + "completion_tps_avg": 44.7, + "model_call_wall_s": 218.0, + "tool_call_wall_s": 11.0, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 10.9, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 500.09, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0321, + "cost_usd_upper_bound": 0.0042, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9732, + "mean_power_w": 453.29, + "mean_sm_util_pct": 80.96, + "max_temp_c": 73.0 + } + }, + { + "run_name": "p3_market_gemma4-31b-q4_v8", + "task": "p3_market", + "replicate": 8, + "summary": { + "started_at": "2026-08-02T05:20:06.879311+00:00", + "ended_at": "2026-08-02T05:22:52.841778+00:00", + "elapsed_s": 166.0, + "iterations": 17, + "total_completion_tokens": 8104, + "total_prompt_tokens": 295610, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Researched and recommended Bitwarden Enterprise as the password manager for a 50-person engineering team. Evaluated 5 products (Bitwarden, 1Password, Keeper, Dashlane, NordPass) against criteria including SOC2 compliance, SSO integration, and Developer CLI support. Provided a detailed recommendation memo, a side-by-side comparison, and a full audit trail of sources and decisions." + }, + "verdict": "STRUCTURAL_PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_market_gemma4-31b-q4_v8", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 166.0, + "iters": 17, + "tokens": { + "completion_total": 8104, + "prompt_total": 295610, + "all_total": 303714 + }, + "throughput": { + "completion_tps_avg": 50.4, + "model_call_wall_s": 160.9, + "tool_call_wall_s": 4.1, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 7.9, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 167.97, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0231, + "cost_usd_upper_bound": 0.003, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9639, + "mean_power_w": 452.41, + "mean_sm_util_pct": 85.27, + "max_temp_c": 71.0 + } + }, + { + "run_name": "p3_market_gemma4-31b-q4_v9", + "task": "p3_market", + "replicate": 9, + "summary": { + "started_at": "2026-08-02T04:49:47.548784+00:00", + "ended_at": "2026-08-02T04:53:00.345777+00:00", + "elapsed_s": 192.8, + "iterations": 18, + "total_completion_tokens": 9097, + "total_prompt_tokens": 362891, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "I have completed the research and recommendation for the password manager standardization. \n\nKey outcomes:\n1. **Recommended Product**: Bitwarden Enterprise ($6/user/month, ~$3,600/year for 50 seats).\n2. **Runner-up**: 1Password Enterprise (Superior UX, but less transparent pricing).\n3. **Criteria**: Prioritized SSO integration, SOC2 compliance, and Developer Experience (CLI/API).\n4. **Deliverables**: \n - `/workspace/recommendation.md`: Final memo.\n - `/workspace/comparison.md`: Feature and pricing matrix.\n - `/workspace/sources.md`: Full list of cited URLs.\n - `/workspace/decisions/`: ADRs for product selection and evaluation criteria.\n - `/workspace/research/notes.md`: Working research notes.\n - `README.md`: Guide to the output.\n\nAll factual claims are cited in the recommendation and comparison documents." + }, + "verdict": "STRUCTURAL_PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_market_gemma4-31b-q4_v9", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 192.8, + "iters": 18, + "tokens": { + "completion_total": 9097, + "prompt_total": 362891, + "all_total": 371988 + }, + "throughput": { + "completion_tps_avg": 48.7, + "model_call_wall_s": 186.8, + "tool_call_wall_s": 4.9, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 10.9, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 378.53, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0268, + "cost_usd_upper_bound": 0.0035, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9595, + "mean_power_w": 484.11, + "mean_sm_util_pct": 93.21, + "max_temp_c": 76.0 + } + }, + { + "run_name": "p3_market_gemma4-31b-q4_v10", + "task": "p3_market", + "replicate": 10, + "summary": { + "started_at": "2026-08-02T05:22:54.437384+00:00", + "ended_at": "2026-08-02T05:26:22.529902+00:00", + "elapsed_s": 208.1, + "iterations": 18, + "total_completion_tokens": 10104, + "total_prompt_tokens": 328243, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Researched and recommended a password manager for a 50-person engineering team. Evaluated 1Password, Bitwarden, Dashlane, Keeper, and NordPass. Recommended 1Password Business ($4,794/year) based on its superior developer experience (CLI), SOC2 compliance, and high adoption potential. Provided a full comparison matrix, ADRs, and a comprehensive sources list." + }, + "verdict": "STRUCTURAL_PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_market_gemma4-31b-q4_v10", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 208.1, + "iters": 18, + "tokens": { + "completion_total": 10104, + "prompt_total": 328243, + "all_total": 338347 + }, + "throughput": { + "completion_tps_avg": 51.4, + "model_call_wall_s": 196.7, + "tool_call_wall_s": 10.3, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 9.3, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 122.44, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0289, + "cost_usd_upper_bound": 0.0038, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9611, + "mean_power_w": 476.83, + "mean_sm_util_pct": 89.93, + "max_temp_c": 71.0 + } + }, + { + "run_name": "p3_writing_gemma4-31b-q4_v1", + "task": "p3_writing", + "replicate": 1, + "summary": { + "started_at": "2026-08-02T02:27:13.236311+00:00", + "ended_at": "2026-08-02T02:28:46.227783+00:00", + "elapsed_s": 93.0, + "iterations": 13, + "total_completion_tokens": 5089, + "total_prompt_tokens": 72407, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Produced three audience-tailored rewrites of an incident memo: a CEO brief, a customer email, and a legal summary. Each rewrite adheres to specific constraints regarding tone, content, and word limits defined in `audience_briefs.json`. Documentation includes a README, research notes, and ADR-style decision records for transparency on content omissions and tone choices." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_writing_gemma4-31b-q4_v1", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 93.0, + "iters": 13, + "tokens": { + "completion_total": 5089, + "prompt_total": 72407, + "all_total": 77496 + }, + "throughput": { + "completion_tps_avg": 55.7, + "model_call_wall_s": 91.3, + "tool_call_wall_s": 0.8, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 24.1, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 226.75, + "memory_used_mib_at_start": 66221.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0129, + "cost_usd_upper_bound": 0.0017, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.914, + "mean_power_w": 460.49, + "mean_sm_util_pct": 86.0, + "max_temp_c": 73.0 + } + }, + { + "run_name": "p3_writing_gemma4-31b-q4_v2", + "task": "p3_writing", + "replicate": 2, + "summary": { + "started_at": "2026-08-02T01:28:14.508298+00:00", + "ended_at": "2026-08-02T01:31:26.342552+00:00", + "elapsed_s": 191.8, + "iterations": 11, + "total_completion_tokens": 4336, + "total_prompt_tokens": 58205, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Successfully produced three audience-specific rewrites of the internal outage memo. \n\n1. **CEO Brief**: Concise executive summary focusing on board-level signal (second outage in 90 days) and key remediations, without technical jargon.\n2. **Customer Email**: Accountable and warm communication explaining the outage in plain language and offering SLA credits, while omitting internal blame and board-level context.\n3. **Legal Summary**: Precise, dispassionate factual record for SLA credit determination, focusing on time windows, account impact, and contractual thresholds.\n\nAll outputs were verified against the `audience_briefs.json` constraints for content, tone, and word limits. Supporting documentation (decisions and research notes) has been provided." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_writing_gemma4-31b-q4_v2", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 191.8, + "iters": 11, + "tokens": { + "completion_total": 4336, + "prompt_total": 58205, + "all_total": 62541 + }, + "throughput": { + "completion_tps_avg": 22.8, + "model_call_wall_s": 190.4, + "tool_call_wall_s": 0.7, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 86.8, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 500.05, + "memory_used_mib_at_start": 65581.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0266, + "cost_usd_upper_bound": 0.0035, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9645, + "mean_power_w": 482.8, + "mean_sm_util_pct": 87.37, + "max_temp_c": 87.0 + } + }, + { + "run_name": "p3_writing_gemma4-31b-q4_v3", + "task": "p3_writing", + "replicate": 3, + "summary": { + "started_at": "2026-08-02T02:28:47.253711+00:00", + "ended_at": "2026-08-02T02:30:00.964816+00:00", + "elapsed_s": 73.7, + "iterations": 11, + "total_completion_tokens": 4250, + "total_prompt_tokens": 58290, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Produced three tailored rewrites of the outage memo (CEO brief, Customer email, and Legal summary) adhering to all constraints specified in audience_briefs.json. Included detailed decision records and research notes." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_writing_gemma4-31b-q4_v3", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 73.7, + "iters": 11, + "tokens": { + "completion_total": 4250, + "prompt_total": 58290, + "all_total": 62540 + }, + "throughput": { + "completion_tps_avg": 58.7, + "model_call_wall_s": 72.4, + "tool_call_wall_s": 0.7, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 38.9, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 259.41, + "memory_used_mib_at_start": 66221.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0102, + "cost_usd_upper_bound": 0.0013, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9498, + "mean_power_w": 441.34, + "mean_sm_util_pct": 79.0, + "max_temp_c": 77.0 + } + }, + { + "run_name": "p3_writing_gemma4-31b-q4_v4", + "task": "p3_writing", + "replicate": 4, + "summary": { + "started_at": "2026-08-02T05:26:24.123041+00:00", + "ended_at": "2026-08-02T05:27:53.468111+00:00", + "elapsed_s": 89.3, + "iterations": 11, + "total_completion_tokens": 5045, + "total_prompt_tokens": 56093, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Produced three tailored rewrites of the API outage incident: a concise CEO brief, a warm and accountable customer email, and a precise legal summary. All rewrites adhere strictly to the constraints provided in `audience_briefs.json`, including word limits, required content, and prohibited terms. Documentation of decisions and research notes are also provided." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_writing_gemma4-31b-q4_v4", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 89.3, + "iters": 11, + "tokens": { + "completion_total": 5045, + "prompt_total": 56093, + "all_total": 61138 + }, + "throughput": { + "completion_tps_avg": 57.4, + "model_call_wall_s": 87.9, + "tool_call_wall_s": 0.7, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 21.9, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 147.9, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0124, + "cost_usd_upper_bound": 0.0016, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9518, + "mean_power_w": 453.98, + "mean_sm_util_pct": 81.11, + "max_temp_c": 69.0 + } + }, + { + "run_name": "p3_writing_gemma4-31b-q4_v5", + "task": "p3_writing", + "replicate": 5, + "summary": { + "started_at": "2026-08-02T04:53:01.661441+00:00", + "ended_at": "2026-08-02T04:54:20.140087+00:00", + "elapsed_s": 78.5, + "iterations": 11, + "total_completion_tokens": 4250, + "total_prompt_tokens": 58290, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Produced three tailored rewrites of the outage memo (CEO brief, Customer email, and Legal summary) adhering to all constraints specified in audience_briefs.json. Included detailed decision records and research notes." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_writing_gemma4-31b-q4_v5", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 78.5, + "iters": 11, + "tokens": { + "completion_total": 4250, + "prompt_total": 58290, + "all_total": 62540 + }, + "throughput": { + "completion_tps_avg": 55.1, + "model_call_wall_s": 77.1, + "tool_call_wall_s": 0.7, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 40.6, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 499.94, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0109, + "cost_usd_upper_bound": 0.0014, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9554, + "mean_power_w": 430.88, + "mean_sm_util_pct": 66.25, + "max_temp_c": 73.0 + } + }, + { + "run_name": "p3_writing_gemma4-31b-q4_v6", + "task": "p3_writing", + "replicate": 6, + "summary": { + "started_at": "2026-08-02T05:27:55.090700+00:00", + "ended_at": "2026-08-02T05:29:22.048280+00:00", + "elapsed_s": 87.0, + "iterations": 11, + "total_completion_tokens": 5045, + "total_prompt_tokens": 56093, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Produced three tailored rewrites of the API outage incident: a concise CEO brief, a warm and accountable customer email, and a precise legal summary. All rewrites adhere strictly to the constraints provided in `audience_briefs.json`, including word limits, required content, and prohibited terms. Documentation of decisions and research notes are also provided." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_writing_gemma4-31b-q4_v6", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 87.0, + "iters": 11, + "tokens": { + "completion_total": 5045, + "prompt_total": 56093, + "all_total": 61138 + }, + "throughput": { + "completion_tps_avg": 59.0, + "model_call_wall_s": 85.5, + "tool_call_wall_s": 0.7, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 21.9, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 121.47, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0121, + "cost_usd_upper_bound": 0.0016, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9195, + "mean_power_w": 495.19, + "mean_sm_util_pct": 97.47, + "max_temp_c": 70.0 + } + }, + { + "run_name": "p3_writing_gemma4-31b-q4_v7", + "task": "p3_writing", + "replicate": 7, + "summary": { + "started_at": "2026-08-02T04:54:21.466351+00:00", + "ended_at": "2026-08-02T04:55:36.480824+00:00", + "elapsed_s": 75.0, + "iterations": 11, + "total_completion_tokens": 4318, + "total_prompt_tokens": 58254, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Produced three tailored rewrites of the outage memo (CEO brief, Customer email, Legal summary) based on specifications in audience_briefs.json. All rewrites adhere to word limits, tone requirements, and content constraints. Included ADR-style records of decisions and working notes." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_writing_gemma4-31b-q4_v7", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 75.0, + "iters": 11, + "tokens": { + "completion_total": 4318, + "prompt_total": 58254, + "all_total": 62572 + }, + "throughput": { + "completion_tps_avg": 58.7, + "model_call_wall_s": 73.6, + "tool_call_wall_s": 0.7, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 39.3, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 499.61, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0104, + "cost_usd_upper_bound": 0.0014, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9333, + "mean_power_w": 443.43, + "mean_sm_util_pct": 85.27, + "max_temp_c": 74.0 + } + }, + { + "run_name": "p3_writing_gemma4-31b-q4_v8", + "task": "p3_writing", + "replicate": 8, + "summary": { + "started_at": "2026-08-02T05:29:23.663196+00:00", + "ended_at": "2026-08-02T05:31:17.184264+00:00", + "elapsed_s": 113.5, + "iterations": 13, + "total_completion_tokens": 5311, + "total_prompt_tokens": 74398, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Successfully produced three audience-specific rewrites of the incident memo: a CEO brief, a customer email, and a legal summary. Each rewrite strictly adheres to the constraints provided in audience_briefs.json, including word limits, mandatory inclusions, and prohibited content. Supporting documentation, including decision records (ADRs), research notes, and a README, has been provided in the /workspace directory." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_writing_gemma4-31b-q4_v8", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 113.5, + "iters": 13, + "tokens": { + "completion_total": 5311, + "prompt_total": 74398, + "all_total": 79709 + }, + "throughput": { + "completion_tps_avg": 47.5, + "model_call_wall_s": 111.9, + "tool_call_wall_s": 0.9, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 23.3, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 154.17, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0158, + "cost_usd_upper_bound": 0.0021, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9251, + "mean_power_w": 483.94, + "mean_sm_util_pct": 90.91, + "max_temp_c": 72.0 + } + }, + { + "run_name": "p3_writing_gemma4-31b-q4_v9", + "task": "p3_writing", + "replicate": 9, + "summary": { + "started_at": "2026-08-02T04:55:37.835386+00:00", + "ended_at": "2026-08-02T04:57:07.101066+00:00", + "elapsed_s": 89.3, + "iterations": 13, + "total_completion_tokens": 4932, + "total_prompt_tokens": 73582, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Successfully produced three tailored rewrites of an incident memo: a CEO brief, a customer email, and a legal summary. Each document adheres to the specific constraints, word limits, and tones defined in `audience_briefs.json`. Included detailed decision records (ADRs), research notes, and a README for transparency and verification." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_writing_gemma4-31b-q4_v9", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 89.3, + "iters": 13, + "tokens": { + "completion_total": 4932, + "prompt_total": 73582, + "all_total": 78514 + }, + "throughput": { + "completion_tps_avg": 56.3, + "model_call_wall_s": 87.7, + "tool_call_wall_s": 0.8, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 24.5, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 499.94, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0124, + "cost_usd_upper_bound": 0.0016, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.8959, + "mean_power_w": 474.04, + "mean_sm_util_pct": 91.82, + "max_temp_c": 73.0 + } + }, + { + "run_name": "p3_writing_gemma4-31b-q4_v10", + "task": "p3_writing", + "replicate": 10, + "summary": { + "started_at": "2026-08-02T05:31:18.852345+00:00", + "ended_at": "2026-08-02T05:32:32.970755+00:00", + "elapsed_s": 74.1, + "iterations": 11, + "total_completion_tokens": 4250, + "total_prompt_tokens": 58290, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Produced three tailored rewrites of the outage memo (CEO brief, Customer email, and Legal summary) adhering to all constraints specified in audience_briefs.json. Included detailed decision records and research notes." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_writing_gemma4-31b-q4_v10", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 74.1, + "iters": 11, + "tokens": { + "completion_total": 4250, + "prompt_total": 58290, + "all_total": 62540 + }, + "throughput": { + "completion_tps_avg": 58.4, + "model_call_wall_s": 72.7, + "tool_call_wall_s": 0.7, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 39.0, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 164.51, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0103, + "cost_usd_upper_bound": 0.0013, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9447, + "mean_power_w": 495.55, + "mean_sm_util_pct": 97.73, + "max_temp_c": 69.0 + } + }, + { + "run_name": "p3_pm_gemma4-31b-q4_v1", + "task": "p3_pm", + "replicate": 1, + "summary": { + "started_at": "2026-08-02T01:31:27.254962+00:00", + "ended_at": "2026-08-02T01:34:31.145299+00:00", + "elapsed_s": 183.9, + "iterations": 7, + "total_completion_tokens": 3786, + "total_prompt_tokens": 28755, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Produced a structured project status report for Project Aurora, synthesizing six weeks of meeting notes. The final deliverables include `status_report.md` (the primary report), `research/notes.md` (working synthesis), and a `README.md`. The report highlights the shift of the embedded SDK to a private beta due to a discovered security gap and tracks milestones for mid-May and late-July." + }, + "verdict": "FAIL", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_pm_gemma4-31b-q4_v1", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 183.9, + "iters": 7, + "tokens": { + "completion_total": 3786, + "prompt_total": 28755, + "all_total": 32541 + }, + "throughput": { + "completion_tps_avg": 20.7, + "model_call_wall_s": 183.0, + "tool_call_wall_s": 0.4, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 115.9, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 499.9, + "memory_used_mib_at_start": 65581.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0255, + "cost_usd_upper_bound": 0.0033, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9788, + "mean_power_w": 484.17, + "mean_sm_util_pct": 90.46, + "max_temp_c": 86.0 + } + }, + { + "run_name": "p3_pm_gemma4-31b-q4_v2", + "task": "p3_pm", + "replicate": 2, + "summary": { + "started_at": "2026-08-02T02:30:01.998527+00:00", + "ended_at": "2026-08-02T02:31:14.189844+00:00", + "elapsed_s": 72.2, + "iterations": 7, + "total_completion_tokens": 4145, + "total_prompt_tokens": 29617, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "I have produced a structured project status report for Project Aurora based on the provided six weeks of meeting notes. The report includes a headline, a detailed workstream table, a risk register ordered by severity, a chronological decisions log, and a milestone list. All claims are cited by the meeting week, and the most recent status updates were used for all workstreams. All deliverables are located in the /workspace directory as requested." + }, + "verdict": "FAIL", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_pm_gemma4-31b-q4_v2", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 72.2, + "iters": 7, + "tokens": { + "completion_total": 4145, + "prompt_total": 29617, + "all_total": 33762 + }, + "throughput": { + "completion_tps_avg": 58.1, + "model_call_wall_s": 71.3, + "tool_call_wall_s": 0.4, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 44.1, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 249.69, + "memory_used_mib_at_start": 66221.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.01, + "cost_usd_upper_bound": 0.0013, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9695, + "mean_power_w": 472.88, + "mean_sm_util_pct": 90.27, + "max_temp_c": 76.0 + } + }, + { + "run_name": "p3_pm_gemma4-31b-q4_v3", + "task": "p3_pm", + "replicate": 3, + "summary": { + "started_at": "2026-08-02T01:34:32.050965+00:00", + "ended_at": "2026-08-02T01:37:24.333730+00:00", + "elapsed_s": 172.3, + "iterations": 6, + "total_completion_tokens": 3386, + "total_prompt_tokens": 25278, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Produced a structured project status report for Project Aurora. The final deliverable (`/workspace/status_report.md`) synthesizes six weeks of meeting notes into a high-level overview for the CEO, including a workstream roadmap, risk register, decisions log, and milestone list. All claims are cited by week and grounded in the provided source material. Synthesis notes and a README were also provided." + }, + "verdict": "FAIL", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_pm_gemma4-31b-q4_v3", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 172.3, + "iters": 6, + "tokens": { + "completion_total": 3386, + "prompt_total": 25278, + "all_total": 28664 + }, + "throughput": { + "completion_tps_avg": 19.7, + "model_call_wall_s": 171.5, + "tool_call_wall_s": 0.4, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 112.3, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 500.09, + "memory_used_mib_at_start": 65709.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0239, + "cost_usd_upper_bound": 0.0031, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9867, + "mean_power_w": 478.32, + "mean_sm_util_pct": 93.23, + "max_temp_c": 86.0 + } + }, + { + "run_name": "p3_pm_gemma4-31b-q4_v4", + "task": "p3_pm", + "replicate": 4, + "summary": { + "started_at": "2026-08-02T05:32:34.647427+00:00", + "ended_at": "2026-08-02T05:33:36.175291+00:00", + "elapsed_s": 61.5, + "iterations": 7, + "total_completion_tokens": 3600, + "total_prompt_tokens": 26878, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Produced a structured project status report for Project Aurora at `/workspace/status_report.md`. The report synthesizes six weeks of meeting notes into a concise executive summary including workstream status, a risk register, a decisions log, and key milestones, all cited by week. Additional working notes and a README are provided in `/workspace`." + }, + "verdict": "FAIL", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_pm_gemma4-31b-q4_v4", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 61.5, + "iters": 7, + "tokens": { + "completion_total": 3600, + "prompt_total": 26878, + "all_total": 30478 + }, + "throughput": { + "completion_tps_avg": 59.4, + "model_call_wall_s": 60.6, + "tool_call_wall_s": 0.4, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 34.5, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 121.06, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0085, + "cost_usd_upper_bound": 0.0011, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.8943, + "mean_power_w": 498.59, + "mean_sm_util_pct": 94.08, + "max_temp_c": 69.0 + } + }, + { + "run_name": "p3_pm_gemma4-31b-q4_v5", + "task": "p3_pm", + "replicate": 5, + "summary": { + "started_at": "2026-08-02T04:57:08.460008+00:00", + "ended_at": "2026-08-02T04:58:09.732291+00:00", + "elapsed_s": 61.3, + "iterations": 7, + "total_completion_tokens": 3600, + "total_prompt_tokens": 26878, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Produced a structured project status report for Project Aurora at `/workspace/status_report.md`. The report synthesizes six weeks of meeting notes into a concise executive summary including workstream status, a risk register, a decisions log, and key milestones, all cited by week. Additional working notes and a README are provided in `/workspace`." + }, + "verdict": "FAIL", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_pm_gemma4-31b-q4_v5", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 61.3, + "iters": 7, + "tokens": { + "completion_total": 3600, + "prompt_total": 26878, + "all_total": 30478 + }, + "throughput": { + "completion_tps_avg": 59.6, + "model_call_wall_s": 60.4, + "tool_call_wall_s": 0.4, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 34.3, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 451.24, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0085, + "cost_usd_upper_bound": 0.0011, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.8972, + "mean_power_w": 493.89, + "mean_sm_util_pct": 96.92, + "max_temp_c": 73.0 + } + }, + { + "run_name": "p3_pm_gemma4-31b-q4_v6", + "task": "p3_pm", + "replicate": 6, + "summary": { + "started_at": "2026-08-02T05:33:37.844408+00:00", + "ended_at": "2026-08-02T05:34:47.289362+00:00", + "elapsed_s": 69.4, + "iterations": 7, + "total_completion_tokens": 4145, + "total_prompt_tokens": 29617, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "I have produced a structured project status report for Project Aurora based on the provided six weeks of meeting notes. The report includes a headline, a detailed workstream table, a risk register ordered by severity, a chronological decisions log, and a milestone list. All claims are cited by the meeting week, and the most recent status updates were used for all workstreams. All deliverables are located in the /workspace directory as requested." + }, + "verdict": "FAIL", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_pm_gemma4-31b-q4_v6", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 69.4, + "iters": 7, + "tokens": { + "completion_total": 4145, + "prompt_total": 29617, + "all_total": 33762 + }, + "throughput": { + "completion_tps_avg": 60.4, + "model_call_wall_s": 68.6, + "tool_call_wall_s": 0.4, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 44.4, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 131.76, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0096, + "cost_usd_upper_bound": 0.0012, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.8646, + "mean_power_w": 494.58, + "mean_sm_util_pct": 97.69, + "max_temp_c": 70.0 + } + }, + { + "run_name": "p3_pm_gemma4-31b-q4_v7", + "task": "p3_pm", + "replicate": 7, + "summary": { + "started_at": "2026-08-02T04:58:11.118975+00:00", + "ended_at": "2026-08-02T04:59:23.605065+00:00", + "elapsed_s": 72.5, + "iterations": 7, + "total_completion_tokens": 4145, + "total_prompt_tokens": 29617, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "I have produced a structured project status report for Project Aurora based on the provided six weeks of meeting notes. The report includes a headline, a detailed workstream table, a risk register ordered by severity, a chronological decisions log, and a milestone list. All claims are cited by the meeting week, and the most recent status updates were used for all workstreams. All deliverables are located in the /workspace directory as requested." + }, + "verdict": "FAIL", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_pm_gemma4-31b-q4_v7", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 72.5, + "iters": 7, + "tokens": { + "completion_total": 4145, + "prompt_total": 29617, + "all_total": 33762 + }, + "throughput": { + "completion_tps_avg": 57.9, + "model_call_wall_s": 71.6, + "tool_call_wall_s": 0.4, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 45.6, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 499.83, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0101, + "cost_usd_upper_bound": 0.0013, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9655, + "mean_power_w": 469.05, + "mean_sm_util_pct": 88.0, + "max_temp_c": 72.0 + } + }, + { + "run_name": "p3_pm_gemma4-31b-q4_v8", + "task": "p3_pm", + "replicate": 8, + "summary": { + "started_at": "2026-08-02T05:34:48.955391+00:00", + "ended_at": "2026-08-02T05:35:59.243911+00:00", + "elapsed_s": 70.3, + "iterations": 7, + "total_completion_tokens": 4103, + "total_prompt_tokens": 28950, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Produced a structured project status report for Project Aurora based on six weeks of meeting notes. The deliverable includes a headline, workstream table, risk register, decisions log, milestones, and open asks, all cited by week. Supporting research notes and a README are also provided in /workspace." + }, + "verdict": "FAIL", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_pm_gemma4-31b-q4_v8", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 70.3, + "iters": 7, + "tokens": { + "completion_total": 4103, + "prompt_total": 28950, + "all_total": 33053 + }, + "throughput": { + "completion_tps_avg": 59.1, + "model_call_wall_s": 69.4, + "tool_call_wall_s": 0.4, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 41.2, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 167.61, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0098, + "cost_usd_upper_bound": 0.0013, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9246, + "mean_power_w": 483.78, + "mean_sm_util_pct": 97.71, + "max_temp_c": 69.0 + } + }, + { + "run_name": "p3_pm_gemma4-31b-q4_v9", + "task": "p3_pm", + "replicate": 9, + "summary": { + "started_at": "2026-08-02T04:59:25.001729+00:00", + "ended_at": "2026-08-02T05:00:22.776214+00:00", + "elapsed_s": 57.8, + "iterations": 8, + "total_completion_tokens": 3319, + "total_prompt_tokens": 33101, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Produced a structured project status report for Project Aurora based on 6 weeks of meeting notes. The report includes a headline, workstream tracking, risk register, decisions log, milestones, and open asks, all cited by week. Supporting research notes and a README were also provided." + }, + "verdict": "FAIL", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_pm_gemma4-31b-q4_v9", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 57.8, + "iters": 8, + "tokens": { + "completion_total": 3319, + "prompt_total": 33101, + "all_total": 36420 + }, + "throughput": { + "completion_tps_avg": 58.5, + "model_call_wall_s": 56.8, + "tool_call_wall_s": 0.5, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 33.3, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 499.98, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.008, + "cost_usd_upper_bound": 0.001, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.8651, + "mean_power_w": 487.69, + "mean_sm_util_pct": 88.64, + "max_temp_c": 73.0 + } + }, + { + "run_name": "p3_pm_gemma4-31b-q4_v10", + "task": "p3_pm", + "replicate": 10, + "summary": { + "started_at": "2026-08-02T05:36:01.014050+00:00", + "ended_at": "2026-08-02T05:36:58.747412+00:00", + "elapsed_s": 57.7, + "iterations": 8, + "total_completion_tokens": 3319, + "total_prompt_tokens": 33101, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Produced a structured project status report for Project Aurora based on 6 weeks of meeting notes. The report includes a headline, workstream tracking, risk register, decisions log, milestones, and open asks, all cited by week. Supporting research notes and a README were also provided." + }, + "verdict": "FAIL", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_pm_gemma4-31b-q4_v10", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 57.7, + "iters": 8, + "tokens": { + "completion_total": 3319, + "prompt_total": 33101, + "all_total": 36420 + }, + "throughput": { + "completion_tps_avg": 58.5, + "model_call_wall_s": 56.8, + "tool_call_wall_s": 0.5, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 33.4, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 184.98, + "memory_used_mib_at_start": 66571.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.008, + "cost_usd_upper_bound": 0.001, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9532, + "mean_power_w": 464.83, + "mean_sm_util_pct": 85.5, + "max_temp_c": 69.0 + } + } + ] +} diff --git a/benchmarks/gemma4-31b-q4/gemma4-canonical-n10-scorecard.md b/benchmarks/gemma4-31b-q4/gemma4-canonical-n10-scorecard.md new file mode 100644 index 00000000..ab98e584 --- /dev/null +++ b/benchmarks/gemma4-31b-q4/gemma4-canonical-n10-scorecard.md @@ -0,0 +1,31 @@ +# Gemma 4 31B Q4 raw canonical scorecard (N=10) + +> Raw grader verdicts only. Any reproducible correction is a separate overlay tied to unchanged archive and grader hashes. + +- Evidence-complete runs: 120/120 +- Normal completed workspaces: 120/120 +- Explicit terminal outcomes: 0/120 +- Graded runs: 120/120 +- Raw pass-equivalent outcomes: 89/120 +- Median model-call completion throughput: 55.85 tok/s +- Median cell wall time: 113.05 s +- Telemetry-complete runs: 118/120 + +| Task | Raw pass | Scored | Finish reasons | Quality outcomes | +|---|---:|---:|---|---| +| `p1_bugfix` | 4/10 | 10/10 | done_signal:10 | FAIL:6, PASS:4 | +| `p1_testwrite` | 8/10 | 10/10 | done_signal:10 | FAIL:2, PASS:8 | +| `p1_refactor` | 7/10 | 10/10 | done_signal:10 | FAIL:3, PASS:7 | +| `p2_extract` | 10/10 | 10/10 | done_signal:10 | PASS:10 | +| `p2_ci` | 10/10 | 10/10 | done_signal:10 | PASS:10 | +| `p2_hallucination` | 6/10 | 10/10 | done_signal:6, model_stopped:4 | MISSING_OUTPUT:4, PASS:6 | +| `p2_triage` | 8/10 | 10/10 | done_signal:10 | FAIL:2, PASS:8 | +| `p3_doc` | 10/10 | 10/10 | done_signal:10 | PASS:10 | +| `p3_business` | 6/10 | 10/10 | done_signal:10 | FAIL:4, PASS:6 | +| `p3_market` | 10/10 | 10/10 | done_signal:10 | STRUCTURAL_PASS:10 | +| `p3_writing` | 10/10 | 10/10 | done_signal:10 | PASS:10 | +| `p3_pm` | 0/10 | 10/10 | done_signal:10 | FAIL:10 | + +## Methodology boundary + +A `done_signal` is a finish behavior, not a pass. `PASS` and `STRUCTURAL_PASS` count only as raw pass-equivalent grader verdicts. A preserved terminal label is reported as a distinct non-pass quality outcome, never fabricated into a normal grader verdict. Model-call throughput excludes tool execution; wall time includes it. Telemetry is per attributed replica GPU, while CPU package power is shared host context and AC wall power is unavailable to software. diff --git a/benchmarks/gemma4-31b-q4/gemma4-canonical-n3-evidence-audit.json b/benchmarks/gemma4-31b-q4/gemma4-canonical-n3-evidence-audit.json new file mode 100644 index 00000000..6fa7132a --- /dev/null +++ b/benchmarks/gemma4-31b-q4/gemma4-canonical-n3-evidence-audit.json @@ -0,0 +1,2323 @@ +{ + "schema_version": 1, + "generated_at": "2026-08-02T03:05:53.534835+00:00", + "root": "/home/michael/bench-gemma4-31b-q4", + "label": "gemma4-31b-q4", + "target_n": 3, + "expected_runs": 36, + "audited_runs": 36, + "passed": true, + "errors": [], + "warnings": [ + "p1_bugfix_gemma4-31b-q4_v1: pre-telemetry valid attempt; supplemental telemetry required", + "p1_bugfix_gemma4-31b-q4_v2: pre-telemetry valid attempt; supplemental telemetry required", + "p2_hallucination_gemma4-31b-q4_v3: telemetry coverage below 80%: 0.7716" + ], + "raw_telemetry": { + "path": "/home/michael/gemma4-campaign-state/telemetry/snapshots/canonical-n3-plus-supplement.csv", + "bytes": 330168, + "lines": 3393, + "sha256_at_audit_time": "bd637547ff5ea9a2b6cb68ce8a203b13712f3b16bf708f2bb10badae962ef328" + }, + "runs": [ + { + "run_name": "p1_bugfix_gemma4-31b-q4_v1", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "7a1938178dbdde521151e1b1d392fec32d55434c", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 708.7, + "iterations": 80, + "completion_tokens": 18475, + "prompt_tokens_cumulative": 1583932, + "model_turns": 80, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": null, + "files": { + "receipt.json": { + "bytes": 8170, + "sha256": "b2a2b09a244b2461fa4f21616139b6d892c2f300e8a25b4155b5d0df03e318f7" + }, + "transcript.jsonl": { + "bytes": 77185, + "sha256": "72bb0d5cf9474eb1c2851c13001ec8f699ecb587a842962e6ad74328fad657b1" + }, + "summary.json": { + "bytes": 1313, + "sha256": "b97ae41ff92d952c9bbbef50b5534efa2498edb98b6d7d05de89271df613f1d5" + }, + "workspace_final.tar.gz": { + "bytes": 28932897, + "sha256": "4bc1a9558a52c9fb68a3fb2ad1e5748a30de4315f80d6dbcad109cdcfae24a76" + }, + "cost.json": { + "bytes": 1242, + "sha256": "aac1c08a5f940f83536e8dc0e66fd34662b89bd4ba17938de29a976d9d8a8748" + }, + "grade.json": { + "bytes": 467, + "sha256": "8ef51f6f2d9457eb209c41700017a65b6f13674d9251e4c8b02352a7ad89a4aa" + } + } + }, + { + "run_name": "p1_bugfix_gemma4-31b-q4_v2", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "7a1938178dbdde521151e1b1d392fec32d55434c", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 639.5, + "iterations": 58, + "completion_tokens": 14733, + "prompt_tokens_cumulative": 984362, + "model_turns": 58, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": null, + "files": { + "receipt.json": { + "bytes": 8170, + "sha256": "e1afd7f0706ff36f0395a9404de1df3c5875c2d95725d2da24a94f0f2ae0c1a6" + }, + "transcript.jsonl": { + "bytes": 52146, + "sha256": "f6731ff3cdd0af7fab869e689f3b6ded47de098838dc90a2859859e5c15018ba" + }, + "summary.json": { + "bytes": 1999, + "sha256": "576e3927a4e9cf68341324837d68a22d637570306554f829561e1c9298869609" + }, + "workspace_final.tar.gz": { + "bytes": 28949571, + "sha256": "841879c97462eda9cc1612eedc24ee8611eb3656efdda9ee880fea24497d5b57" + }, + "cost.json": { + "bytes": 1240, + "sha256": "cd746dc74c66ffa830be960b7be0d7ef66a7c70c6390f3350b8ca15e695c7423" + }, + "grade.json": { + "bytes": 468, + "sha256": "8c1caead3b6ffa35ac77aefd6150cc3bc5513596a7f125c9e024d34226cdfe3f" + } + } + }, + { + "run_name": "p1_bugfix_gemma4-31b-q4_v3", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 623.7, + "iterations": 62, + "completion_tokens": 14968, + "prompt_tokens_cumulative": 1161167, + "model_turns": 62, + "length_finishes": 0, + "grade_verdict": "FAIL", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9941, + "mean_power_w": 249.36, + "mean_sm_util_pct": 44.38, + "max_temp_c": 76.0, + "cpu_package_mean_power_w": 133.84 + }, + "files": { + "receipt.json": { + "bytes": 8953, + "sha256": "d3a7789767428adc34242dc59125578e4d4445a2ddc5a99b2e78d52718b127c3" + }, + "transcript.jsonl": { + "bytes": 55242, + "sha256": "a2f6e8cec02de3d745fc9894e4051ca328417a0b6edac6966ec9edfd1a8a8766" + }, + "summary.json": { + "bytes": 1514, + "sha256": "9850feb80175867ba25c84877acf3f08d16524b568e2119cfab241bf34c746e7" + }, + "workspace_final.tar.gz": { + "bytes": 28953180, + "sha256": "f07b17077a855af54375939b7c7c65fd0a613a18659456d2f493f4ec7a50fac3" + }, + "cost.json": { + "bytes": 1242, + "sha256": "6d408865f3a1d7aafbaf662eeec0e46bc32ef3baf889a20be018b84c4ec834a4" + }, + "gpu_telemetry.json": { + "bytes": 2450, + "sha256": "8db210ac88e35367989709f1bcdb813f856de6bfd241eca5b175bdfcd26a9b3d" + }, + "grade.json": { + "bytes": 468, + "sha256": "10243cc2a1b0f2c8df720128ec698608433a7b5176a9ffe716efd29205815667" + } + } + }, + { + "run_name": "p1_testwrite_gemma4-31b-q4_v1", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 523.6, + "iterations": 49, + "completion_tokens": 25223, + "prompt_tokens_cumulative": 1273638, + "model_turns": 49, + "length_finishes": 0, + "grade_verdict": "FAIL", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9836, + "mean_power_w": 481.16, + "mean_sm_util_pct": 91.31, + "max_temp_c": 81.0, + "cpu_package_mean_power_w": 133.28 + }, + "files": { + "receipt.json": { + "bytes": 8955, + "sha256": "982dcf5a0e4637072a83ce0a400bea7422122b943e97ce3e9f4f5d091fab143c" + }, + "transcript.jsonl": { + "bytes": 57036, + "sha256": "48fed858bc6c18f032a8e1f89230547b425815156e1afb29af37ddebe63a51bc" + }, + "summary.json": { + "bytes": 1479, + "sha256": "0a9dbf647408093d6503dd606d3725eb6778b66a04c69c51ab8531ac3fcba653" + }, + "workspace_final.tar.gz": { + "bytes": 149484, + "sha256": "a522588ef84bec7cf9edf8da1c090f858fbab32d75c6fa5bfa27c28107172a2a" + }, + "cost.json": { + "bytes": 1242, + "sha256": "034e404e054f1a8a705a4626c2df9372efd9503a5d7c3f25b98851efae7a13fe" + }, + "gpu_telemetry.json": { + "bytes": 2410, + "sha256": "e83b81526a96ab743abb1b3589532651c202f9d2686b44f3cbde702cdb2907b7" + }, + "grade.json": { + "bytes": 1155, + "sha256": "73b7129ab31845b98ef25c4e1501e5367510c8b7fbda84357be4bb3294bbd2f3" + } + } + }, + { + "run_name": "p1_testwrite_gemma4-31b-q4_v2", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 388.8, + "iterations": 48, + "completion_tokens": 18809, + "prompt_tokens_cumulative": 1029851, + "model_turns": 48, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9902, + "mean_power_w": 479.04, + "mean_sm_util_pct": 89.65, + "max_temp_c": 78.0, + "cpu_package_mean_power_w": 136.92 + }, + "files": { + "receipt.json": { + "bytes": 8958, + "sha256": "a88c0da0f29cdfa068f246f94cc94779032defc619e0042bf024aa5fddba632c" + }, + "transcript.jsonl": { + "bytes": 55995, + "sha256": "e29b8ec4f36990ccb224af8f215b357e02c5b9d29e4c544cef96d84717e69116" + }, + "summary.json": { + "bytes": 679, + "sha256": "3ea94ce4164c4c8bdbebf903d38495718e251b94140b9d18aeba137736609520" + }, + "workspace_final.tar.gz": { + "bytes": 138148, + "sha256": "ed2052b32ae52b5832aebc625413764854230bbe23611d931d735874219ab03d" + }, + "cost.json": { + "bytes": 1240, + "sha256": "881968bad40e6940fd25b639c3f4f10b759e11b440176941eea76146e05cf9b6" + }, + "gpu_telemetry.json": { + "bytes": 2449, + "sha256": "c4bce59813eea97fb8558046580fa2dc2526847c5f66a63dde285758cafdbad5" + }, + "grade.json": { + "bytes": 341, + "sha256": "42c1da43e1c31be25775668d65010e502b8383356f35bfd846be7b583f7e153e" + } + } + }, + { + "run_name": "p1_testwrite_gemma4-31b-q4_v3", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 399.9, + "iterations": 42, + "completion_tokens": 19900, + "prompt_tokens_cumulative": 825275, + "model_turns": 42, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9877, + "mean_power_w": 468.52, + "mean_sm_util_pct": 91.96, + "max_temp_c": 86.0, + "cpu_package_mean_power_w": 136.84 + }, + "files": { + "receipt.json": { + "bytes": 8958, + "sha256": "057d735baa76ae75cab8aae2d06c0c10d73a4adcd9dc49790be911cc83eeec94" + }, + "transcript.jsonl": { + "bytes": 50956, + "sha256": "fb8be270e1fb754d0585a92189259eb5a9e14426abe9983ac3c155621dbe5600" + }, + "summary.json": { + "bytes": 1158, + "sha256": "da9c949bb86558868d75aa50347057cf3f65510b908d1b5be4ac309e593d21b0" + }, + "workspace_final.tar.gz": { + "bytes": 140677, + "sha256": "f7e3e9cd2be319d6df53da86252e78481fecf0cc597e4b43efdafbf83d0ce07c" + }, + "cost.json": { + "bytes": 1241, + "sha256": "ba4db6fb45cc046dc3487fcc5ce45b00feae0406373ef7360e73951565a4bdc1" + }, + "gpu_telemetry.json": { + "bytes": 2449, + "sha256": "9281dd0ce4fd2d380117e959ce61d2baecbf1998709a91770e198e8bcda2e8fb" + }, + "grade.json": { + "bytes": 341, + "sha256": "b99e9530499f12935a79df989d2f182b0917530a4b3be44b95eaa26e7f2d2c33" + } + } + }, + { + "run_name": "p1_refactor_gemma4-31b-q4_v1", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 186.2, + "iterations": 35, + "completion_tokens": 9779, + "prompt_tokens_cumulative": 351967, + "model_turns": 35, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9667, + "mean_power_w": 465.36, + "mean_sm_util_pct": 85.35, + "max_temp_c": 78.0, + "cpu_package_mean_power_w": 137.0 + }, + "files": { + "receipt.json": { + "bytes": 8956, + "sha256": "c898eedb4993578fca466cb12e19f45091ab32aec824d1ba8bce1e2e72641454" + }, + "transcript.jsonl": { + "bytes": 25468, + "sha256": "945b7feb6c271ad743d914462ba601146de308ce04b3c22696a9938e35467918" + }, + "summary.json": { + "bytes": 783, + "sha256": "fd3a64637a800835a7f5d09dec5d6f1f2bb82aabfa540481a9af9727d745e1b1" + }, + "workspace_final.tar.gz": { + "bytes": 74100, + "sha256": "bd8d0ac1c21d887b949eb9d97fe36673933f5fe012fb8bad4657a59873722816" + }, + "cost.json": { + "bytes": 1240, + "sha256": "66f111f744443b50626b3050b4775796a6471a968811ca4d439f739bd2059384" + }, + "gpu_telemetry.json": { + "bytes": 2447, + "sha256": "4bae1b27dab4e67c47c0a07de21b37bf4bfc954ba9a966e0bf3cef89922cfa11" + }, + "grade.json": { + "bytes": 510, + "sha256": "7f69334378ac3187de489eb74ad81c8245267d1a5711612b94879f297e758bcf" + } + } + }, + { + "run_name": "p1_refactor_gemma4-31b-q4_v2", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 211.5, + "iterations": 41, + "completion_tokens": 11126, + "prompt_tokens_cumulative": 508191, + "model_turns": 41, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9693, + "mean_power_w": 459.06, + "mean_sm_util_pct": 89.05, + "max_temp_c": 84.0, + "cpu_package_mean_power_w": 137.01 + }, + "files": { + "receipt.json": { + "bytes": 8956, + "sha256": "a8c89f90890e9acacec48b8c73ede92286ffd047bdc5d52a7bace17a18d6930b" + }, + "transcript.jsonl": { + "bytes": 26507, + "sha256": "aa1d445c65f35cdba5254bc3c5de40eea9b75df89e5552bacc7bc9b94f8d7855" + }, + "summary.json": { + "bytes": 730, + "sha256": "6a11f94056186ebdb47360977f854063250f81ac92ea572868fd747efd5daf38" + }, + "workspace_final.tar.gz": { + "bytes": 75807, + "sha256": "49cee0476af144ba1185102887ffedd0141488e34acdff629800bef941cec037" + }, + "cost.json": { + "bytes": 1242, + "sha256": "eee7bd5e349aab1f394356abc6c67b8ec8b09d442367a7afd0ab1af554a802ef" + }, + "gpu_telemetry.json": { + "bytes": 2448, + "sha256": "8f05e2f6b887eacecae0187c3a54a053d0bb3acb6fe01cb67d080baac077a574" + }, + "grade.json": { + "bytes": 510, + "sha256": "7f69334378ac3187de489eb74ad81c8245267d1a5711612b94879f297e758bcf" + } + } + }, + { + "run_name": "p1_refactor_gemma4-31b-q4_v3", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "1a4a954715bf7b228a2d5496fa87547318be6ba8", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 198.2, + "iterations": 40, + "completion_tokens": 10797, + "prompt_tokens_cumulative": 399232, + "model_turns": 40, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9586, + "mean_power_w": 471.53, + "mean_sm_util_pct": 87.9, + "max_temp_c": 71.0, + "cpu_package_mean_power_w": 112.13 + }, + "files": { + "receipt.json": { + "bytes": 9512, + "sha256": "d8b8de86c6abb82b39c59fd2ed56571dc1c05e4c1cf2214fce1a54803bb20dca" + }, + "transcript.jsonl": { + "bytes": 26550, + "sha256": "5ac6ee6b8399781a2394c9e663e4ac0c0850655b6ccb4d80d8ad307e6b62d3ab" + }, + "summary.json": { + "bytes": 893, + "sha256": "e22036d8e57d95f93243ae7cd3f22cfc271686c8134c715dfc44b7459410fe61" + }, + "workspace_final.tar.gz": { + "bytes": 66503, + "sha256": "427d4a3a2eaef5eb3ed46b025996fdbc51fce966e234f63855e4542984c3bd1e" + }, + "cost.json": { + "bytes": 1241, + "sha256": "b296c66ff2312ad59757460488b509e23770bf1e6d79e3bee11d7a11110770cf" + }, + "gpu_telemetry.json": { + "bytes": 2400, + "sha256": "f6e6a834c58ed8e87854468ea3292129e98077bfc5b7155062a30fa570b2eddb" + }, + "grade.json": { + "bytes": 510, + "sha256": "7f69334378ac3187de489eb74ad81c8245267d1a5711612b94879f297e758bcf" + } + } + }, + { + "run_name": "p2_extract_gemma4-31b-q4_v1", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 76.4, + "iterations": 6, + "completion_tokens": 4449, + "prompt_tokens_cumulative": 22299, + "model_turns": 6, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9162, + "mean_power_w": 460.29, + "mean_sm_util_pct": 91.33, + "max_temp_c": 83.0, + "cpu_package_mean_power_w": 136.95 + }, + "files": { + "receipt.json": { + "bytes": 8954, + "sha256": "f4c6d7cec85a2b5e5b05cbfeeb2311e06b0c9392ec9ef3e3623963981c780b31" + }, + "transcript.jsonl": { + "bytes": 4462, + "sha256": "32ca9ee6cf929cd5d153e7f4d7e52d4dd120dc9e99a780b706f7ffeacc2de67c" + }, + "summary.json": { + "bytes": 530, + "sha256": "c5e97fda27f744cb1368b05e5a6faa3a166ee12264d47bc8e96759fcfb48e251" + }, + "workspace_final.tar.gz": { + "bytes": 11518, + "sha256": "65b92ede9b5f084ca706d4b201f5861fdd4c15d04e7211b5369607b5b9e4013a" + }, + "cost.json": { + "bytes": 1233, + "sha256": "e18fb448f9dcea19dadcac89dce493e73b58e678e779bfadf9fd77285b890fda" + }, + "gpu_telemetry.json": { + "bytes": 2445, + "sha256": "b8c84f438fc2203f7c8d77ec0ba90cadde556c01c18bea511ddfc930ac2fe357" + }, + "grade.json": { + "bytes": 3204, + "sha256": "b6bc97fc6e03c0b85c34d7e6b5aa1ffcfdf128e953247bc7569f9417170aa2b8" + } + } + }, + { + "run_name": "p2_extract_gemma4-31b-q4_v2", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 72.6, + "iterations": 6, + "completion_tokens": 4449, + "prompt_tokens_cumulative": 22299, + "model_turns": 6, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.8953, + "mean_power_w": 493.67, + "mean_sm_util_pct": 90.86, + "max_temp_c": 75.0, + "cpu_package_mean_power_w": 113.55 + }, + "files": { + "receipt.json": { + "bytes": 8952, + "sha256": "ce495248feeffe85113c7293821946f22107710355182b97cabc340902bf20e9" + }, + "transcript.jsonl": { + "bytes": 4461, + "sha256": "d3ecc25456ffc3638644f9ae4b60e1e7dcc53b77d3d30d85097d89bafc5c9bff" + }, + "summary.json": { + "bytes": 530, + "sha256": "06fbc403a3d9e561ea45f9bed30c58fd7f7c9dc8ba46c7c0e9fa798255f6a633" + }, + "workspace_final.tar.gz": { + "bytes": 11519, + "sha256": "f9eecea3446b9e0c0306a8fce7d0e79eb3fab3f5d33c27a44dfed5009b05fa2d" + }, + "cost.json": { + "bytes": 1233, + "sha256": "b7eaa3addf58279ba31c7eddbed435dd33430501fc64b3e745c2398b292bf976" + }, + "gpu_telemetry.json": { + "bytes": 2398, + "sha256": "ee106709da076e2ad1f97940ae43645bcc7f47fbd3a8df728565a08d35205e7e" + }, + "grade.json": { + "bytes": 3204, + "sha256": "b6bc97fc6e03c0b85c34d7e6b5aa1ffcfdf128e953247bc7569f9417170aa2b8" + } + } + }, + { + "run_name": "p2_extract_gemma4-31b-q4_v3", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 73.9, + "iterations": 6, + "completion_tokens": 4449, + "prompt_tokens_cumulative": 22299, + "model_turns": 6, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.8796, + "mean_power_w": 489.97, + "mean_sm_util_pct": 97.5, + "max_temp_c": 83.0, + "cpu_package_mean_power_w": 136.83 + }, + "files": { + "receipt.json": { + "bytes": 8954, + "sha256": "48090475a70a17e3285430ffb43d146eac293b9e056705af2ef36e9eb795d0e3" + }, + "transcript.jsonl": { + "bytes": 4461, + "sha256": "f4f8cd64f8bd2c10501cedad1365accd1846bc0a0e9c1f346d922c10c5664646" + }, + "summary.json": { + "bytes": 530, + "sha256": "05ae675a95053a282a2ec46decbd7dadc3433d5a663a298a82472d91daae96b5" + }, + "workspace_final.tar.gz": { + "bytes": 11512, + "sha256": "6f39c6c22c3cfa9114f079a566bf355f9c5d9123a6e33b836d24729c059bd962" + }, + "cost.json": { + "bytes": 1233, + "sha256": "0521bb3ebf24dcb019320cb7f8e968a73b0affbe8e9948cdd793a0cdc6e3f1be" + }, + "gpu_telemetry.json": { + "bytes": 2400, + "sha256": "d994fdb5e28fe41a221ee2acb88f3f6995f5f53cdcfd122ea6e9f345f2855bf8" + }, + "grade.json": { + "bytes": 3204, + "sha256": "b6bc97fc6e03c0b85c34d7e6b5aa1ffcfdf128e953247bc7569f9417170aa2b8" + } + } + }, + { + "run_name": "p2_ci_gemma4-31b-q4_v1", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 88.5, + "iterations": 29, + "completion_tokens": 4141, + "prompt_tokens_cumulative": 234185, + "model_turns": 29, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.904, + "mean_power_w": 440.07, + "mean_sm_util_pct": 69.35, + "max_temp_c": 76.0, + "cpu_package_mean_power_w": 113.04 + }, + "files": { + "receipt.json": { + "bytes": 8947, + "sha256": "1bfeb5c41370bb8e9c06a9cf1442b5475f7433baae1432b55df02397a601c3f3" + }, + "transcript.jsonl": { + "bytes": 18156, + "sha256": "c37e909e7d44dd910771a1a082a0dce0c783b13c2ddb90c6263078cf44397a15" + }, + "summary.json": { + "bytes": 854, + "sha256": "9e8f9cf9e8c965bb749bf17dcedcd8c6a9df7ded2bbb4d03b0644d36f103a361" + }, + "workspace_final.tar.gz": { + "bytes": 42534, + "sha256": "462800e2a5ba5abc66790efa44478c855dc72d785877714df198208cbd052176" + }, + "cost.json": { + "bytes": 1230, + "sha256": "60c46c1767609b23a08dceae3efa29b18b8aba6147e48b1e67fb42515c13cbc7" + }, + "gpu_telemetry.json": { + "bytes": 2391, + "sha256": "cde7c49411e541739642838f93aeabf561dcdf8d006aad272c1e8e924a889c0e" + }, + "grade.json": { + "bytes": 482, + "sha256": "69e86f389fb98ab93705abc5fd07431d5a9730ba5b0214d67cccdc8eeb154ba6" + } + } + }, + { + "run_name": "p2_ci_gemma4-31b-q4_v2", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 94.5, + "iterations": 29, + "completion_tokens": 4551, + "prompt_tokens_cumulative": 230536, + "model_turns": 29, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.8995, + "mean_power_w": 427.5, + "mean_sm_util_pct": 81.5, + "max_temp_c": 84.0, + "cpu_package_mean_power_w": 136.99 + }, + "files": { + "receipt.json": { + "bytes": 8949, + "sha256": "330300d0460fb67ce69d65b17ba1fabf6127921d0d430f14efbb628bd95e441f" + }, + "transcript.jsonl": { + "bytes": 18370, + "sha256": "3399f7f3cd0183e8e9f3be341cfdc9a8f90874139fce4bd9777c8ecac5bec873" + }, + "summary.json": { + "bytes": 929, + "sha256": "a4a1834961b6e83966b444c41f381cb804e7d4b50ecfd0362aece86fd902195e" + }, + "workspace_final.tar.gz": { + "bytes": 48425, + "sha256": "157aed82ecac3ad44166bc2b54c8d7cc12cda6cd7afa5663ee71c9fa71fac791" + }, + "cost.json": { + "bytes": 1230, + "sha256": "e9c8def56e3794c5f7e4506bdc9e1b09b1dba23f55c0f0836e72e3d38afd5317" + }, + "gpu_telemetry.json": { + "bytes": 2396, + "sha256": "1f20117fd246a72b61518b6fa137fe375c71d2855e5af3a42dac7db71a9c1b77" + }, + "grade.json": { + "bytes": 482, + "sha256": "eae0f41e05a5ce0e1a370e295ccb33bfd10b26e892ff34e286253904640fd666" + } + } + }, + { + "run_name": "p2_ci_gemma4-31b-q4_v3", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 80.3, + "iterations": 21, + "completion_tokens": 3754, + "prompt_tokens_cumulative": 162381, + "model_turns": 21, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.934, + "mean_power_w": 435.34, + "mean_sm_util_pct": 79.44, + "max_temp_c": 72.0, + "cpu_package_mean_power_w": 112.97 + }, + "files": { + "receipt.json": { + "bytes": 8947, + "sha256": "de8c9045350678be58207876eebbfa395ab86d5e5ad7b19676f11c12b711e3b1" + }, + "transcript.jsonl": { + "bytes": 14266, + "sha256": "1ad6dd78d4c1f09b610922e47b0d78686957cd2bad6071994ae6e2a484428350" + }, + "summary.json": { + "bytes": 881, + "sha256": "3ec5f398e260ad00d3ee87b7c5b6d40905369999b640150c306f2eb42b7a9612" + }, + "workspace_final.tar.gz": { + "bytes": 40255, + "sha256": "e2f34ad318240ff88eedc621842a08ba6715ff4eb33629f3e942637c0848dc86" + }, + "cost.json": { + "bytes": 1230, + "sha256": "ce7c5a844032037bb5de24b724dcaa6c97ce6caa2cbf5c0e5481602ed615241d" + }, + "gpu_telemetry.json": { + "bytes": 2391, + "sha256": "182a21c9e60bd0d47fcb6847c5d850647055d44642c2407b432bdd3477bfe845" + }, + "grade.json": { + "bytes": 482, + "sha256": "1cdcddf815967c7395bceb0d8ac5b955294d82d5d932594f553e43f22ab4234d" + } + } + }, + { + "run_name": "p2_hallucination_gemma4-31b-q4_v1", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 115.6, + "iterations": 16, + "completion_tokens": 6454, + "prompt_tokens_cumulative": 121925, + "model_turns": 16, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9516, + "mean_power_w": 488.6, + "mean_sm_util_pct": 91.83, + "max_temp_c": 83.0, + "cpu_package_mean_power_w": 136.58 + }, + "files": { + "receipt.json": { + "bytes": 8966, + "sha256": "000d8b774edf30dac9b4fe688ef6ed94bffddb1196edda77469269ecf806b602" + }, + "transcript.jsonl": { + "bytes": 9410, + "sha256": "769746a886657bd6a339d0b376a9405b5d877fcc770da5c1b28d8ffb09cbf142" + }, + "summary.json": { + "bytes": 489, + "sha256": "9922fd7d5dcb5739be2ca233f2ada76172b5beff1cf37779d874f87963e3d677" + }, + "workspace_final.tar.gz": { + "bytes": 11352, + "sha256": "b4de9834bc4281d8823e8aeadcb633e13e672640604304733e1b096aa58dab57" + }, + "cost.json": { + "bytes": 1245, + "sha256": "c13344093961e064a51b8a6550ab1464debb87e2a1989056e8b65b09cb9ae939" + }, + "gpu_telemetry.json": { + "bytes": 2409, + "sha256": "93b43cae80fc001d4ccb7f85cb106a365ef48f72c32736645c44e7a17bd38a5c" + }, + "grade.json": { + "bytes": 3586, + "sha256": "80910919c9e64a98ef035d0e5da409c4b26fc67de0ad65bcf63df6b4012393e6" + } + } + }, + { + "run_name": "p2_hallucination_gemma4-31b-q4_v2", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 114.1, + "iterations": 17, + "completion_tokens": 6154, + "prompt_tokens_cumulative": 128689, + "model_turns": 17, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9202, + "mean_power_w": 481.27, + "mean_sm_util_pct": 93.09, + "max_temp_c": 74.0, + "cpu_package_mean_power_w": 113.31 + }, + "files": { + "receipt.json": { + "bytes": 8964, + "sha256": "2cbada4e3d3fadf359becdf5da4b9b866f6d51d025453f642c23faadd5879e27" + }, + "transcript.jsonl": { + "bytes": 10566, + "sha256": "eafdf5cd5b4d9b307e1d51724bc1b042e2970ead94e6eed757d8e6de8edd4c66" + }, + "summary.json": { + "bytes": 442, + "sha256": "42545f2faeccbd3c3c5af74e4640c9da4e2454ed0a38ac1f3e81ba6392299d7e" + }, + "workspace_final.tar.gz": { + "bytes": 11814, + "sha256": "00ae973a8d6e9f1e567870e7403d1bf605bf3e82a70bdef609287d640946daf5" + }, + "cost.json": { + "bytes": 1244, + "sha256": "00fd00be3788b4948ddfc27c49d81d176eb0de734f3d872df545c37054021e58" + }, + "gpu_telemetry.json": { + "bytes": 2407, + "sha256": "be3f0c5d9aca73c4051fdc6b880a0dd9bb37bccc971301d20c01e666e0cc9e48" + }, + "grade.json": { + "bytes": 3892, + "sha256": "2553e9edff94d4bf1cc8fda5c170ae25201bd4baa261fc4ab14d5754dbc11bed" + } + } + }, + { + "run_name": "p2_hallucination_gemma4-31b-q4_v3", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "model_stopped", + "elapsed_s": 32.4, + "iterations": 9, + "completion_tokens": 1494, + "prompt_tokens_cumulative": 43015, + "model_turns": 9, + "length_finishes": 0, + "grade_verdict": "MISSING_OUTPUT", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.7716, + "mean_power_w": 417.97, + "mean_sm_util_pct": 83.33, + "max_temp_c": 83.0, + "cpu_package_mean_power_w": 137.12 + }, + "files": { + "receipt.json": { + "bytes": 8966, + "sha256": "c588242987936d1cc91253b8dfe727b4567ee3df212b10b0cc3de0097010adce" + }, + "transcript.jsonl": { + "bytes": 3643, + "sha256": "7096baadd99ee25917c930c0ef99ed2746e11fe4d85142197ebc2a05c12e8f03" + }, + "summary.json": { + "bytes": 346, + "sha256": "c3a15b3d9a9d02fd77292f80fc985d594c28eb5b848e1a57a496243816a89111" + }, + "workspace_final.tar.gz": { + "bytes": 10689, + "sha256": "cb4a3526503a052dac2aaaa769c353fca88417571ed16dfbefd9d3ab09e0d4aa" + }, + "cost.json": { + "bytes": 1239, + "sha256": "5954c2443675b7879dac302d9089518852403f8a613775af1d9ba61dd772b4f8" + }, + "gpu_telemetry.json": { + "bytes": 2400, + "sha256": "e4a260f5b5fe1973664034384e06173c381fb9b15bd6d022a27f76c98709e2ad" + }, + "grade.json": { + "bytes": 153, + "sha256": "e4f8b9da1c7aba632e47e0a97771274c35662e1f1e2920b050928e498df8692a" + } + } + }, + { + "run_name": "p2_triage_gemma4-31b-q4_v1", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 108.0, + "iterations": 8, + "completion_tokens": 6211, + "prompt_tokens_cumulative": 57359, + "model_turns": 8, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9259, + "mean_power_w": 470.89, + "mean_sm_util_pct": 85.81, + "max_temp_c": 75.0, + "cpu_package_mean_power_w": 113.07 + }, + "files": { + "receipt.json": { + "bytes": 8943, + "sha256": "667b801d0ba40236fb24011fe328c17b62ebee20c2ced924cb7751ed61e5ec29" + }, + "transcript.jsonl": { + "bytes": 10882, + "sha256": "98d8dfa0c742dadeb2ed12650b13f0e6b6c28aeb81de737f908ee1bd9c02ad15" + }, + "summary.json": { + "bytes": 600, + "sha256": "7a5d63bcd45effd15282edde7eef5d2319c72c0d0f89aa0306c3399913d34102" + }, + "workspace_final.tar.gz": { + "bytes": 12714, + "sha256": "9614fb670dc01897da49271da079012b43d8b729af22be864ffe61f18724b321" + }, + "cost.json": { + "bytes": 1233, + "sha256": "0d26ca75ca0e7cdf0d11ef36a3ef7aeed7a270c08a247a2a203b02b3c3a1db0a" + }, + "gpu_telemetry.json": { + "bytes": 2395, + "sha256": "a3fb3239f9459cb8e4f81d218623bd7e590961a509f2cffd5a18e1de1cf1f51c" + }, + "grade.json": { + "bytes": 2250, + "sha256": "dbe6f269c365da51b85a5053d621051e8877fc7a7642fcc441c97f6c7653b5f5" + } + } + }, + { + "run_name": "p2_triage_gemma4-31b-q4_v2", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 125.5, + "iterations": 7, + "completion_tokens": 7489, + "prompt_tokens_cumulative": 52650, + "model_turns": 7, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9562, + "mean_power_w": 499.84, + "mean_sm_util_pct": 97.36, + "max_temp_c": 84.0, + "cpu_package_mean_power_w": 136.63 + }, + "files": { + "receipt.json": { + "bytes": 8945, + "sha256": "0288e87079edfe2244a331754506643c233790a8cd98d5dc40707a11c9a19f68" + }, + "transcript.jsonl": { + "bytes": 10908, + "sha256": "bf08f200f4260bf44aff76dcea291177648d68dd8ebf33a8aa7d874ef862408f" + }, + "summary.json": { + "bytes": 659, + "sha256": "217354a083ead3542f13afd8380ca186fa28041cc8cc6726b8127a737c1c00bb" + }, + "workspace_final.tar.gz": { + "bytes": 12920, + "sha256": "3349fae04f9709e12ab95b4559a5d6568b7ba017c582733187b344025d7148d9" + }, + "cost.json": { + "bytes": 1233, + "sha256": "3e4975f2694ec3f49d6c712fa9ee2063143dacc684bc0a9d5d40b939e0a3f85e" + }, + "gpu_telemetry.json": { + "bytes": 2400, + "sha256": "7bbe3f407ab9b9d7c620ee5730059d3fb51a04e0228488c2495ea8ee714c5d46" + }, + "grade.json": { + "bytes": 2159, + "sha256": "c14e218d16bc8ba2bfb56fd8577d0f8539d628d0366f9340f7d2f60f17c91290" + } + } + }, + { + "run_name": "p2_triage_gemma4-31b-q4_v3", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 128.3, + "iterations": 7, + "completion_tokens": 7493, + "prompt_tokens_cumulative": 52729, + "model_turns": 7, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9353, + "mean_power_w": 489.85, + "mean_sm_util_pct": 91.68, + "max_temp_c": 75.0, + "cpu_package_mean_power_w": 113.07 + }, + "files": { + "receipt.json": { + "bytes": 8943, + "sha256": "cb906fe8c74989a3a0844c92838f901effd382d84db056815fa34d4cb808ca10" + }, + "transcript.jsonl": { + "bytes": 10654, + "sha256": "79b135973c3285e556806fcb26738fca7a775f02550338d4a61af085a06c9ade" + }, + "summary.json": { + "bytes": 644, + "sha256": "80f55c1f0a23af5a58a6d67552818f1af54ef24f223a46063e6c144ff0286d25" + }, + "workspace_final.tar.gz": { + "bytes": 12802, + "sha256": "92ab6a3283af89a091f8329811a9b05d781fd3dee3ee0985946fe5657048cf95" + }, + "cost.json": { + "bytes": 1234, + "sha256": "bd2fb5a02e7a3758b62c82d542de7c5cb4bf7a074b0c91381378c1b9eee5e63e" + }, + "gpu_telemetry.json": { + "bytes": 2399, + "sha256": "b5ce4aca228b743be558bb095669f32abb6d7bd9d47c4303464681ee93d3a63e" + }, + "grade.json": { + "bytes": 1914, + "sha256": "8b145bdcbc912d2f546ce33b90d11fe348c69829b205abe91d6513b83ba9daef" + } + } + }, + { + "run_name": "p3_doc_gemma4-31b-q4_v1", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 121.8, + "iterations": 8, + "completion_tokens": 7296, + "prompt_tokens_cumulative": 51659, + "model_turns": 8, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9442, + "mean_power_w": 488.83, + "mean_sm_util_pct": 96.71, + "max_temp_c": 84.0, + "cpu_package_mean_power_w": 136.91 + }, + "files": { + "receipt.json": { + "bytes": 8956, + "sha256": "e82000d2c622ccbd80272cb91a940f666af73f187dcfd791bcf0eaba5f2e1dd0" + }, + "transcript.jsonl": { + "bytes": 14483, + "sha256": "57e17b5eb071b15b2e3d1959f41f9823b94ae0166874adaf01c4a60962afaa39" + }, + "summary.json": { + "bytes": 723, + "sha256": "d2180a1ba15cba18e03d8f52b3595c30954f050f376df3b2106c2b99689937f5" + }, + "workspace_final.tar.gz": { + "bytes": 15602, + "sha256": "764bc3a872dd39c331ae67d35d1c5c9c26758f47668eac413222bfad50a7b355" + }, + "cost.json": { + "bytes": 1231, + "sha256": "b975a39a92cec1d99ccf68599a8edf7900d37e269eaa73d100a5ab38b9ff5eb9" + }, + "gpu_telemetry.json": { + "bytes": 2397, + "sha256": "f2747d5acd70ac8236f7e933300496baf220548e84141c142e1599ce409e5fb3" + }, + "grade.json": { + "bytes": 2357, + "sha256": "addcd3cfc419649494adf40c60fd2385ec3b403f91e07f3423695aac6fd5c0dc" + } + } + }, + { + "run_name": "p3_doc_gemma4-31b-q4_v2", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 123.8, + "iterations": 8, + "completion_tokens": 7296, + "prompt_tokens_cumulative": 51659, + "model_turns": 8, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9693, + "mean_power_w": 480.4, + "mean_sm_util_pct": 92.72, + "max_temp_c": 78.0, + "cpu_package_mean_power_w": 113.16 + }, + "files": { + "receipt.json": { + "bytes": 8954, + "sha256": "2d59a5d6e6bf429d767447461c521914274a8a45763a062dea250be8a7b5a5ec" + }, + "transcript.jsonl": { + "bytes": 14484, + "sha256": "d6a6a9795bbd69d388ab0f058bef562338055664d578af964c95f99a3eb4a667" + }, + "summary.json": { + "bytes": 723, + "sha256": "ce4d10ebe26c6b5612292822982dcb69116f6864744f166d2c388be997cc43b5" + }, + "workspace_final.tar.gz": { + "bytes": 15594, + "sha256": "cc19c24bdd130e1aaa7b2aec944779f0cff243f0eef1b465e0b647243c144d34" + }, + "cost.json": { + "bytes": 1231, + "sha256": "bbf0a3b4f8f96c6a97f15d4030e96f6b44d07d270d97b7610485e09d0f0be90e" + }, + "gpu_telemetry.json": { + "bytes": 2396, + "sha256": "ed567400b60c4e212b96157e552d2672320aafb3ff299950f8811c55fd56ebc7" + }, + "grade.json": { + "bytes": 2357, + "sha256": "addcd3cfc419649494adf40c60fd2385ec3b403f91e07f3423695aac6fd5c0dc" + } + } + }, + { + "run_name": "p3_doc_gemma4-31b-q4_v3", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 113.3, + "iterations": 10, + "completion_tokens": 6632, + "prompt_tokens_cumulative": 67272, + "model_turns": 10, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9709, + "mean_power_w": 469.09, + "mean_sm_util_pct": 89.09, + "max_temp_c": 83.0, + "cpu_package_mean_power_w": 137.02 + }, + "files": { + "receipt.json": { + "bytes": 8956, + "sha256": "4ecf36adec39b7279a65e69a7404cfc30a8da6a0e4318004718259f4af1364f3" + }, + "transcript.jsonl": { + "bytes": 16781, + "sha256": "a8a0b48977fadcb65aaa23a894d973153560173ccf590b3486f64f60fcfcffe2" + }, + "summary.json": { + "bytes": 747, + "sha256": "167d8934c09ec91e3e926ca9499379b34c0b1f6538f6c565eafc01641c42e073" + }, + "workspace_final.tar.gz": { + "bytes": 16341, + "sha256": "fdd57249f4f12c782da3dbb6af706d34983697efabbc32c500c8f0af27c0fad5" + }, + "cost.json": { + "bytes": 1231, + "sha256": "d41d93bd17e42e69cc0100bee0bafa79c063b61e7ce6e41d77b1caf2664e72e3" + }, + "gpu_telemetry.json": { + "bytes": 2401, + "sha256": "b81ad3d434787af21bca5223f71fbcbaaa96a3b4bf643b2f63b6978fe34babef" + }, + "grade.json": { + "bytes": 2360, + "sha256": "a1e32ee1296ef6f8f122b96d91586e2fbd3df7e2ff1f5aea3b13c97c6cb2bdd8" + } + } + }, + { + "run_name": "p3_business_gemma4-31b-q4_v1", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 94.2, + "iterations": 9, + "completion_tokens": 5487, + "prompt_tokens_cumulative": 43691, + "model_turns": 9, + "length_finishes": 0, + "grade_verdict": "FAIL", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9554, + "mean_power_w": 458.24, + "mean_sm_util_pct": 78.84, + "max_temp_c": 77.0, + "cpu_package_mean_power_w": 112.85 + }, + "files": { + "receipt.json": { + "bytes": 8959, + "sha256": "ffcb5af2fb2c4b8fa8cc26fc21c02c64cb905c1f267f576775bbfaee95408fc2" + }, + "transcript.jsonl": { + "bytes": 16635, + "sha256": "224abc79829aa387ea42f89a7e76e1349021dc086196fbcd7f7446794ae12cc8" + }, + "summary.json": { + "bytes": 956, + "sha256": "ea4c6d11864bda3ae075daffecc0989616185dacec6e981fa94c434e643aa988" + }, + "workspace_final.tar.gz": { + "bytes": 15908, + "sha256": "1217b662e553531a731c4cfe9596bae3d1ccbc80752dd7ac12c76da894d65d90" + }, + "cost.json": { + "bytes": 1234, + "sha256": "04e7b3f606e02679842522ae4eeb0ccd3d4b61f87f41e8eee7b38d1bf98e6bcc" + }, + "gpu_telemetry.json": { + "bytes": 2399, + "sha256": "cff9ffadb62632691ccbb74aa0b83ae059b51f205405f6ab30890d1589abcc6e" + }, + "grade.json": { + "bytes": 3866, + "sha256": "3c85cb074672401082ba6adecc5f454e5124b81323f35fc312a8458bba0ab815" + } + } + }, + { + "run_name": "p3_business_gemma4-31b-q4_v2", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 80.9, + "iterations": 9, + "completion_tokens": 4742, + "prompt_tokens_cumulative": 37120, + "model_turns": 9, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9271, + "mean_power_w": 495.91, + "mean_sm_util_pct": 95.5, + "max_temp_c": 83.0, + "cpu_package_mean_power_w": 137.15 + }, + "files": { + "receipt.json": { + "bytes": 8961, + "sha256": "6b54b92c8c0396ffbfbb5ed2f8907b802e3fe4d2e218cb9c888c9fe80f9b3f08" + }, + "transcript.jsonl": { + "bytes": 16134, + "sha256": "75a2f519c54888667c21ad3e3b021569ae835ad81fd233e100f56c1713549b45" + }, + "summary.json": { + "bytes": 853, + "sha256": "3142c3c654573d7bed6174124196fbce6facfb39fae3fbf95fda122e731f626c" + }, + "workspace_final.tar.gz": { + "bytes": 15801, + "sha256": "7a319197444df2e8d1b271a7ef4091fb663cf67241f120057cf084fddf1aaeb0" + }, + "cost.json": { + "bytes": 1234, + "sha256": "ca67bd0d05b32d34219d90fa1100d5963baf1dd028d1b618f9d460e27329bcfb" + }, + "gpu_telemetry.json": { + "bytes": 2399, + "sha256": "5c38965d3ed993c1c268eac6091a90a141477416d43088f5ae20d0e2057e2681" + }, + "grade.json": { + "bytes": 3873, + "sha256": "13358399cce26683f59886f1cfc577330fd90233e4065a9af249f51edd05c4ab" + } + } + }, + { + "run_name": "p3_business_gemma4-31b-q4_v3", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 79.7, + "iterations": 9, + "completion_tokens": 4742, + "prompt_tokens_cumulative": 37120, + "model_turns": 9, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.941, + "mean_power_w": 486.72, + "mean_sm_util_pct": 93.75, + "max_temp_c": 74.0, + "cpu_package_mean_power_w": 112.93 + }, + "files": { + "receipt.json": { + "bytes": 8959, + "sha256": "06fb3e91e1166b84c55cf7a70f66876735c67cf8e32d2e664bdd0bb45a26d035" + }, + "transcript.jsonl": { + "bytes": 16133, + "sha256": "e1cab5771101129766dc6ec65ae6d633df275d7a54929e1a3da8cf4d579e1571" + }, + "summary.json": { + "bytes": 853, + "sha256": "31fd9a8af67581a1a48bebb0450619fb80d03229cdf321ba55560a4421a54afd" + }, + "workspace_final.tar.gz": { + "bytes": 15805, + "sha256": "8abb097c9a0a7bfbd30f59b051be46e5907e13d1c85c9d2c5769a520dc532886" + }, + "cost.json": { + "bytes": 1234, + "sha256": "3cf0e3376dc28702234453ff8f4e5aeafebe87de47c33d5985e4bc4fabaf63a0" + }, + "gpu_telemetry.json": { + "bytes": 2396, + "sha256": "e177609ba7af6e1f34e365f2fc3f17ea6c34ecc9c94e80f6a865f0d515b898d9" + }, + "grade.json": { + "bytes": 3873, + "sha256": "13358399cce26683f59886f1cfc577330fd90233e4065a9af249f51edd05c4ab" + } + } + }, + { + "run_name": "p3_market_gemma4-31b-q4_v1", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 257.4, + "iterations": 16, + "completion_tokens": 8071, + "prompt_tokens_cumulative": 387754, + "model_turns": 16, + "length_finishes": 0, + "grade_verdict": "STRUCTURAL_PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9713, + "mean_power_w": 495.65, + "mean_sm_util_pct": 96.29, + "max_temp_c": 85.0, + "cpu_package_mean_power_w": 136.75 + }, + "files": { + "receipt.json": { + "bytes": 8928, + "sha256": "02bb435a5b4fdf5ce444c727c2aab9dc982c3820eb7c8ce22ae0aea7d06e9e67" + }, + "transcript.jsonl": { + "bytes": 22003, + "sha256": "f2f2cd2453d41f8c6c8dca98034f64ef3c90cd8d0ce1d6e0ff256a43f2b1fc4b" + }, + "summary.json": { + "bytes": 747, + "sha256": "903cf4fde70a72420d9f5a24d98a9bc7512675e5e9bd1c498b5f4066af5cff60" + }, + "workspace_final.tar.gz": { + "bytes": 264891, + "sha256": "961f9f5fed7d5aca0767339337210745112837366ca5c413405cb20dfdf7fd4f" + }, + "cost.json": { + "bytes": 1236, + "sha256": "640a00660e7718909fe2dc2eff09af048db6f6d20efae3b74af8a50565bc6fbe" + }, + "gpu_telemetry.json": { + "bytes": 2402, + "sha256": "4917b69f7d65423ca4014280e3bf2895fc7801c2df9c179cc4900bd65316f1c4" + }, + "grade.json": { + "bytes": 991, + "sha256": "ec7f817b56a4a6ab2128042d09466fdcdf5b77eaf246ea588cbb987a0865d5fa" + } + } + }, + { + "run_name": "p3_market_gemma4-31b-q4_v2", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 194.1, + "iterations": 20, + "completion_tokens": 9028, + "prompt_tokens_cumulative": 442437, + "model_turns": 20, + "length_finishes": 0, + "grade_verdict": "STRUCTURAL_PASS", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9531, + "mean_power_w": 485.18, + "mean_sm_util_pct": 93.97, + "max_temp_c": 75.0, + "cpu_package_mean_power_w": 112.49 + }, + "files": { + "receipt.json": { + "bytes": 8926, + "sha256": "723d063398588013f7b6052082026114bf29dec559f48dd1eb1e904a1f06ff2e" + }, + "transcript.jsonl": { + "bytes": 23827, + "sha256": "98c92be6cf1e6313be373078a28ab44174526e368a03f53ef5580c75917e6f4e" + }, + "summary.json": { + "bytes": 687, + "sha256": "5e5ccb999a51bb167d790eacd190e63c6384a465cb7bb2bdc7337fccde0858df" + }, + "workspace_final.tar.gz": { + "bytes": 264699, + "sha256": "18d375a961d52e800e0bc7a2445f3a61e29e08236b503731f971a08b6c179d15" + }, + "cost.json": { + "bytes": 1235, + "sha256": "9a0175ff5a729a9a190b9930f3690f1ac68398e2deb8ec6f64820172c4936d85" + }, + "gpu_telemetry.json": { + "bytes": 2398, + "sha256": "a189bbe833bd814fca08e0dec893993e892b63be121d4a029b2d45b0a7925ea1" + }, + "grade.json": { + "bytes": 991, + "sha256": "7a993d6b3110622d5558f7d78f2ebb57044bb1bdf2b0e9f52585b7fdcf81dddc" + } + } + }, + { + "run_name": "p3_market_gemma4-31b-q4_v3", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 322.0, + "iterations": 15, + "completion_tokens": 8059, + "prompt_tokens_cumulative": 271715, + "model_turns": 15, + "length_finishes": 0, + "grade_verdict": "STRUCTURAL_PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9783, + "mean_power_w": 489.71, + "mean_sm_util_pct": 92.86, + "max_temp_c": 86.0, + "cpu_package_mean_power_w": 137.66 + }, + "files": { + "receipt.json": { + "bytes": 8928, + "sha256": "46c0620b93bcdda525541bfd2c4dfb39ea289d92e44059663e740426c2278697" + }, + "transcript.jsonl": { + "bytes": 21639, + "sha256": "28283464e8685547b97fd7307796460573417109ae9ddaf10cf4cbe1ae6060d3" + }, + "summary.json": { + "bytes": 783, + "sha256": "7d159cccc3186adc8b392d9fcb3653164f567c12bb90844c0629604d66de0a1a" + }, + "workspace_final.tar.gz": { + "bytes": 277200, + "sha256": "576a9c8abf100141ecb229b88b7807074aa42e0bdf2dbe9974dfcc9cfea4e980" + }, + "cost.json": { + "bytes": 1237, + "sha256": "cbaa088ac34943052570e107ff00db849be27af9f0da3a2d16e6704231925ece" + }, + "gpu_telemetry.json": { + "bytes": 2404, + "sha256": "7407f92bc939fe9971ddab57be058867e4914c0d23bf36610fbe46f0872ab44d" + }, + "grade.json": { + "bytes": 990, + "sha256": "5d1bcb2c532b0323a5acdb93607cd0a957b6254f0dc687a6751be84f10ba86b2" + } + } + }, + { + "run_name": "p3_writing_gemma4-31b-q4_v1", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 93.0, + "iterations": 13, + "completion_tokens": 5089, + "prompt_tokens_cumulative": 72407, + "model_turns": 13, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.914, + "mean_power_w": 460.49, + "mean_sm_util_pct": 86.0, + "max_temp_c": 73.0, + "cpu_package_mean_power_w": 112.85 + }, + "files": { + "receipt.json": { + "bytes": 8962, + "sha256": "20d5e7d994199823fab97c326232441ca9bd5569ef11a972c402dc7d54589daa" + }, + "transcript.jsonl": { + "bytes": 15990, + "sha256": "404e7626e52a4928a4881eb7be832152ced40aaa01fdc3f1eaf8816a4071fa64" + }, + "summary.json": { + "bytes": 677, + "sha256": "15b672068f5fa802a679f4bea758c2a91c44cb3c664c2b3d3beb8b19b487e469" + }, + "workspace_final.tar.gz": { + "bytes": 15514, + "sha256": "f87a925d16003bb1d84da8c39b4491b168c6b8ba612ff6fd878ddf7e56cbd37c" + }, + "cost.json": { + "bytes": 1234, + "sha256": "3eae502e939914b0b6add91645d27f74ea3ff1db347bdddc9be820eff8b73841" + }, + "gpu_telemetry.json": { + "bytes": 2396, + "sha256": "4e5ad939c642f1ee1d6b0837d0561b856e3e8d3f90100a2b7ca575be1676ab70" + }, + "grade.json": { + "bytes": 2879, + "sha256": "4d52cac7e70665c17a000630f5e703deb8c2dd7fa09b140e58cd916809855594" + } + } + }, + { + "run_name": "p3_writing_gemma4-31b-q4_v2", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 191.8, + "iterations": 11, + "completion_tokens": 4336, + "prompt_tokens_cumulative": 58205, + "model_turns": 11, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9645, + "mean_power_w": 482.8, + "mean_sm_util_pct": 87.37, + "max_temp_c": 87.0, + "cpu_package_mean_power_w": 138.28 + }, + "files": { + "receipt.json": { + "bytes": 8964, + "sha256": "344295433cc7e8dfc413953057888c378ae8a2a8e95d408bdd6e1186fe503a42" + }, + "transcript.jsonl": { + "bytes": 11867, + "sha256": "9d5ca60adf16166a00054ac7e38ecadd7c220127d625ee7b7accb47537e7d600" + }, + "summary.json": { + "bytes": 1072, + "sha256": "25db8e039e0d70de15743604f8da5ca48a646708e5636b23777ff9d4d594e992" + }, + "workspace_final.tar.gz": { + "bytes": 13886, + "sha256": "e5da165c568e1e684254c5d7996c5654318e377afbead4859b05b741f945b5e1" + }, + "cost.json": { + "bytes": 1236, + "sha256": "ccfd5b41afd661d08ef93058d0a09e94569bdb2c31f991637462b4265748964c" + }, + "gpu_telemetry.json": { + "bytes": 2405, + "sha256": "2c020c94e033061b39df41bb589f113e699322cb58466ebf294b79f178d61ba3" + }, + "grade.json": { + "bytes": 2879, + "sha256": "d5b96fa97f5173b2c03372286d3f1337cdad83a2d8433953285d8c1b96d080de" + } + } + }, + { + "run_name": "p3_writing_gemma4-31b-q4_v3", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 73.7, + "iterations": 11, + "completion_tokens": 4250, + "prompt_tokens_cumulative": 58290, + "model_turns": 11, + "length_finishes": 0, + "grade_verdict": "PASS", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9498, + "mean_power_w": 441.34, + "mean_sm_util_pct": 79.0, + "max_temp_c": 77.0, + "cpu_package_mean_power_w": 113.07 + }, + "files": { + "receipt.json": { + "bytes": 8962, + "sha256": "0d34b93613273625c8093145f3bcdc4de4cf88758f864c1c758e10e1faf504e7" + }, + "transcript.jsonl": { + "bytes": 11561, + "sha256": "098a161c5fb99762e3df7f4cb8c490549ddc2435f33ffb7887eeed19aece747f" + }, + "summary.json": { + "bytes": 522, + "sha256": "5b1be637be95a148bb4ad11662deeddb74543a11a845653f7de3bfff000a72ac" + }, + "workspace_final.tar.gz": { + "bytes": 13991, + "sha256": "b612618510b43d4fb1793ecef2149ebd0da6c21685d10b742045429243d9bad5" + }, + "cost.json": { + "bytes": 1234, + "sha256": "dbd969a5548e2a9b59d0e081c28c181c18cab45e22d190de562d8ecbfdfa70a7" + }, + "gpu_telemetry.json": { + "bytes": 2395, + "sha256": "4a9fac128bb5cc8f62b10f059d590312317ebf7d4f8e1444bef9a9af28d1b25b" + }, + "grade.json": { + "bytes": 2876, + "sha256": "5554d7976a5268e573cd20176244ba7db50716c7dc7dad7df89386f9fc269564" + } + } + }, + { + "run_name": "p3_pm_gemma4-31b-q4_v1", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 183.9, + "iterations": 7, + "completion_tokens": 3786, + "prompt_tokens_cumulative": 28755, + "model_turns": 7, + "length_finishes": 0, + "grade_verdict": "FAIL", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9788, + "mean_power_w": 484.17, + "mean_sm_util_pct": 90.46, + "max_temp_c": 86.0, + "cpu_package_mean_power_w": 138.49 + }, + "files": { + "receipt.json": { + "bytes": 8953, + "sha256": "bf5b124d1fcf516f129d94fe8f22407461230499307ab01a24e0013f6de68136" + }, + "transcript.jsonl": { + "bytes": 7558, + "sha256": "d3291b6c0e20bef9bd143b9bc1d6e1c7f41e81c90ad1d867aad5a0cdc77a5e64" + }, + "summary.json": { + "bytes": 694, + "sha256": "83681b6cad2097056abf0f0bcd8bd63642670e1e9959c6c5cf1de1ff099c7369" + }, + "workspace_final.tar.gz": { + "bytes": 12841, + "sha256": "eb622a02fddb70c3aee43f652123e59640282815e55b230d70874175a61fcf08" + }, + "cost.json": { + "bytes": 1230, + "sha256": "4e6c26bf109501788f45951c8b086b0af788ec68ffbbc420b63a17b0167077db" + }, + "gpu_telemetry.json": { + "bytes": 2402, + "sha256": "6d8a6172614487d154a0260381351b5244d0076f298eca9a7f877097217a07ea" + }, + "grade.json": { + "bytes": 2404, + "sha256": "eade149c12512150ebb28d5a446c6dc1037e0ded24a3ed5e83d65d87aba47bc4" + } + } + }, + { + "run_name": "p3_pm_gemma4-31b-q4_v2", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 72.2, + "iterations": 7, + "completion_tokens": 4145, + "prompt_tokens_cumulative": 29617, + "model_turns": 7, + "length_finishes": 0, + "grade_verdict": "FAIL", + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9695, + "mean_power_w": 472.88, + "mean_sm_util_pct": 90.27, + "max_temp_c": 76.0, + "cpu_package_mean_power_w": 132.69 + }, + "files": { + "receipt.json": { + "bytes": 8951, + "sha256": "7975d3f2ebd6af42d6a203790f662b61c46f3d41a6c7497971305b7e830c8481" + }, + "transcript.jsonl": { + "bytes": 8573, + "sha256": "aaaa90cd1dd4cc1ccdb2241a17c9d721400672f226be72fdce168a7a7d2339de" + }, + "summary.json": { + "bytes": 754, + "sha256": "598cb8b9ee1c2b73a0f4e8728259609efe1b0f460331ab883920a660e7d1b7ff" + }, + "workspace_final.tar.gz": { + "bytes": 13139, + "sha256": "6190888ab5900349d69d7d443477eb162cb5174acc340c36c5489a95b9a591c8" + }, + "cost.json": { + "bytes": 1226, + "sha256": "95639c64f720d0ba360c8bf67a7e46bbc821a72c11105f96b653a0632742bb21" + }, + "gpu_telemetry.json": { + "bytes": 2395, + "sha256": "5126fc789f682ea5d3be582131d3b6d4ca5376ae177fb48827eb21c33833188f" + }, + "grade.json": { + "bytes": 2425, + "sha256": "5e98c1230f111ddb4d1a15ae7063c80590fbee19b72661c4b56e9c94f0b12352" + } + } + }, + { + "run_name": "p3_pm_gemma4-31b-q4_v3", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c4a5470607344412b1a3a9e0cdaec3902807d7c4", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 172.3, + "iterations": 6, + "completion_tokens": 3386, + "prompt_tokens_cumulative": 25278, + "model_turns": 6, + "length_finishes": 0, + "grade_verdict": "FAIL", + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9867, + "mean_power_w": 478.32, + "mean_sm_util_pct": 93.23, + "max_temp_c": 86.0, + "cpu_package_mean_power_w": 138.81 + }, + "files": { + "receipt.json": { + "bytes": 8953, + "sha256": "ced44d9abe5553e3d7b62f0bf8c0fb2137ec214ed8e567bb0e5b8a3854f891e5" + }, + "transcript.jsonl": { + "bytes": 7311, + "sha256": "b130ac807c1b2622b54e3e4eb30dd1d0fef1c25db3328b185ed18f360a17fe89" + }, + "summary.json": { + "bytes": 708, + "sha256": "20c8f932f6c38489911471d600d9236cc122deca72f731831bc09cbe2fb2b4e8" + }, + "workspace_final.tar.gz": { + "bytes": 12920, + "sha256": "fac5112297764b16bb742a38fcecc8617524ef13ebe869170e049842e63a4645" + }, + "cost.json": { + "bytes": 1231, + "sha256": "ed09867309f941998fbb60813dcee241c4b31cd8dc8960ccb41d02a20787ca1b" + }, + "gpu_telemetry.json": { + "bytes": 2403, + "sha256": "5ade9d5dba7dfa09feed864c81af1a0e92071e16eca52dee641f3660cd735c6a" + }, + "grade.json": { + "bytes": 2408, + "sha256": "be7951b5c0b3770b9ee61e8f3bd5042dfc0fa486d2be1086b97314e4a8976ca3" + } + } + } + ], + "pretelemetry_supplements": [ + { + "run_name": "p1_bugfix_gemma4-31b-q4-telemetry-supplement_v1", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 779.5, + "iterations": 68, + "completion_tokens": 22980, + "prompt_tokens_cumulative": 1360791, + "model_turns": 68, + "length_finishes": 0, + "grade_verdict": null, + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9942, + "mean_power_w": 297.76, + "mean_sm_util_pct": 51.88, + "max_temp_c": 74.0, + "cpu_package_mean_power_w": 128.3 + }, + "files": { + "receipt.json": { + "bytes": 9533, + "sha256": "160325870cc3b91505bfb372a0cf7f9950a3b51089e197b2b3696c5a6ede0da8" + }, + "transcript.jsonl": { + "bytes": 85245, + "sha256": "3de94b54b395a925cf7702c72e72331f86d3c412c662ab7ca925f436d8cca64d" + }, + "summary.json": { + "bytes": 2148, + "sha256": "ee4ec82f91cbade39cc9229f272ec83dee9389dba31fabde2fcddf33edf3fc23" + }, + "workspace_final.tar.gz": { + "bytes": 28950277, + "sha256": "ecb5307e3a8b2441b72e576298a3b84cfc304e1bf6e476ddbe1714fb48261f07" + }, + "cost.json": { + "bytes": 1265, + "sha256": "0b2b4c671701cc39ec343389a954b83904e95a3899ea18106f3798a8c5b4d9a0" + }, + "gpu_telemetry.json": { + "bytes": 2447, + "sha256": "5c16a3a10acd802272b831194a319a84ab346812c4aac63a2e055c378412398c" + } + }, + "supplements_run": "p1_bugfix_gemma4-31b-q4_v1" + }, + { + "run_name": "p1_bugfix_gemma4-31b-q4-telemetry-supplement_v2", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 1031.5, + "iterations": 78, + "completion_tokens": 19891, + "prompt_tokens_cumulative": 1730028, + "model_turns": 78, + "length_finishes": 0, + "grade_verdict": null, + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9937, + "mean_power_w": 194.22, + "mean_sm_util_pct": 34.09, + "max_temp_c": 74.0, + "cpu_package_mean_power_w": 124.95 + }, + "files": { + "receipt.json": { + "bytes": 9533, + "sha256": "fe67b13b061d11ba2b790bbcee3e3a3b7658a67a1ab1bf7167b9a695c9339ff0" + }, + "transcript.jsonl": { + "bytes": 77337, + "sha256": "809d769786b89a1edc29607dee67707f0308116ea12594690dd327c4823e4883" + }, + "summary.json": { + "bytes": 1570, + "sha256": "a55b9bd46dcefd0dd85a0b2c5fb12f27b48cbfbf37979707c8f20cc3fa86ab31" + }, + "workspace_final.tar.gz": { + "bytes": 28961325, + "sha256": "ab7b933a6113f17198c275ffc0300c81cc8795f2f7264d2665a2f17a90b0fb4d" + }, + "cost.json": { + "bytes": 1265, + "sha256": "4070249f078c72dfd93083d59d1511adb16ac2647e2dd50557594b071699b6e6" + }, + "gpu_telemetry.json": { + "bytes": 2492, + "sha256": "f58db836c68505993e74abba511ff238af517e35e06f3ce598b4c09f76e98f10" + } + }, + "supplements_run": "p1_bugfix_gemma4-31b-q4_v2" + } + ], + "invalid_attempt_classifications": { + "path": "/home/michael/bench-gemma4-31b-q4/tooling/deployments/gemma4-31b-q4-tower2/invalid-attempt-classifications.json", + "bytes": 6601, + "sha256": "09b45db2328ed1520f5491e37897a4cc031d66a9a1187b938977dfc872840176" + }, + "preserved_invalid_attempts": [ + { + "attempt": "p1_bugfix_gemma4-31b-q4_v3_retry-20260802T003822Z-dea60926", + "classification": { + "source_run": "p1_bugfix_gemma4-31b-q4_v3", + "classification": "infrastructure-invalid", + "reason_code": "operator-interrupted-companion-after-harness-dispatch-defect", + "incident_document": "canonical-harness-incident-20260802.json", + "classified_before_grade": true, + "affirmative_evidence": [ + "the companion lane raised the recorded KeyError in the shared defective harness", + "the supervisor was stopped immediately to prevent further evaluation under that harness" + ], + "expected_files": { + "receipt.json": "0a571fa5aaefacb15833eb56d63969d45b69f7e66b7622ec309ddf52ff1be4b5", + "transcript.jsonl": "7c9fa7c58e4ea36ab46938b690785b4a699b98d38c7f934d1aed083d87fd3163" + }, + "replacement": { + "required": true, + "canonical_run": "p1_bugfix_gemma4-31b-q4_v3", + "status": "completed" + } + }, + "files": { + "receipt.json": { + "bytes": 8174, + "sha256": "0a571fa5aaefacb15833eb56d63969d45b69f7e66b7622ec309ddf52ff1be4b5" + }, + "transcript.jsonl": { + "bytes": 1866, + "sha256": "7c9fa7c58e4ea36ab46938b690785b4a699b98d38c7f934d1aed083d87fd3163" + } + } + }, + { + "attempt": "p1_bugfix_gemma4-31b-q4_v3_retry-20260802T004533Z-64195329", + "classification": { + "source_run": "p1_bugfix_gemma4-31b-q4_v3", + "classification": "infrastructure-invalid", + "reason_code": "operator-interrupted-before-required-telemetry-gate", + "incident_document": "canonical-telemetry-gate-incident-20260802.json", + "classified_before_grade": true, + "affirmative_evidence": [ + "the retry began before per-run sidecar attribution was proven", + "the supervisor stopped before the attempt completed" + ], + "expected_files": { + "receipt.json": "c70a93f7255d8d4ef84311ed17187d84050c5e4faf4dbfd2768550e6773e760c", + "transcript.jsonl": "a07274f29bc3b36fe8c3259af0a8a4d53d1b7f8f6256e44dfc43c90b2b5ba7b2" + }, + "replacement": { + "required": true, + "canonical_run": "p1_bugfix_gemma4-31b-q4_v3", + "status": "completed" + } + }, + "files": { + "receipt.json": { + "bytes": 8170, + "sha256": "c70a93f7255d8d4ef84311ed17187d84050c5e4faf4dbfd2768550e6773e760c" + }, + "transcript.jsonl": { + "bytes": 1852, + "sha256": "a07274f29bc3b36fe8c3259af0a8a4d53d1b7f8f6256e44dfc43c90b2b5ba7b2" + } + } + }, + { + "attempt": "p1_refactor_gemma4-31b-q4_v3-server-timeout-20260802T020859Z", + "classification": { + "source_run": "p1_refactor_gemma4-31b-q4_v3", + "classification": "infrastructure-invalid", + "reason_code": "server-transport-timeout-below-native-envelope", + "incident_document": "canonical-refactor-timeout-incident-20260802.json", + "classified_before_grade": true, + "affirmative_evidence": [ + "server task 80497 was cancelled exactly 3600 seconds after launch", + "the slot had decoded 133606 tokens, reported truncated=false, and remained below its 262144-token context boundary" + ], + "expected_files": { + "cost.json": "c5d1062f477758436ba0c3baf32df97fbe4e30df09c831bde470ccf938fce7f7", + "gpu_telemetry.json": "4f717a0e9d7298401c41e1799f772186a03a25bdf21bee93b30425cab3b31494", + "gpu_telemetry.reanalyzed-fdcc2496.json": "51eff413ad7a1a01bc6116a52e31ac3c39a1f844bf0aa0f0e09d9a0731392801", + "grade.json": "7f69334378ac3187de489eb74ad81c8245267d1a5711612b94879f297e758bcf", + "receipt.json": "58b9ee02a9c9cd03829a91d8f2f4d6d5eedf064f1c96206bc6c4b38522c4bfd8", + "summary.json": "0445d7ff95e0fbfbbc06d21cb3fdcccf8a741249ea294dda3f136a3ea4ed2b56", + "transcript.jsonl": "e887af90ea26038d3d8aadf22998eaa443eeb78eaa59d204bc7fcc741124e00e", + "workspace_final.tar.gz": "90a1b1918640b2a496b66ed5fbaf8649eb5902aa2e4e36c1f6b4debc71c720bc" + }, + "replacement": { + "required": true, + "canonical_run": "p1_refactor_gemma4-31b-q4_v3", + "status": "completed", + "receipt_sha256": "d8b8de86c6abb82b39c59fd2ed56571dc1c05e4c1cf2214fce1a54803bb20dca", + "summary_sha256": "e22036d8e57d95f93243ae7cd3f22cfc271686c8134c715dfc44b7459410fe61", + "server_timeout_seconds": 14400, + "finish_reason": "done_signal" + } + }, + "files": { + "cost.json": { + "bytes": 1241, + "sha256": "c5d1062f477758436ba0c3baf32df97fbe4e30df09c831bde470ccf938fce7f7" + }, + "gpu_telemetry.json": { + "bytes": 2939, + "sha256": "4f717a0e9d7298401c41e1799f772186a03a25bdf21bee93b30425cab3b31494" + }, + "gpu_telemetry.reanalyzed-fdcc2496.json": { + "bytes": 2973, + "sha256": "51eff413ad7a1a01bc6116a52e31ac3c39a1f844bf0aa0f0e09d9a0731392801" + }, + "grade.json": { + "bytes": 510, + "sha256": "7f69334378ac3187de489eb74ad81c8245267d1a5711612b94879f297e758bcf" + }, + "receipt.json": { + "bytes": 8956, + "sha256": "58b9ee02a9c9cd03829a91d8f2f4d6d5eedf064f1c96206bc6c4b38522c4bfd8" + }, + "summary.json": { + "bytes": 321, + "sha256": "0445d7ff95e0fbfbbc06d21cb3fdcccf8a741249ea294dda3f136a3ea4ed2b56" + }, + "transcript.jsonl": { + "bytes": 30545, + "sha256": "e887af90ea26038d3d8aadf22998eaa443eeb78eaa59d204bc7fcc741124e00e" + }, + "workspace_final.tar.gz": { + "bytes": 67984, + "sha256": "90a1b1918640b2a496b66ed5fbaf8649eb5902aa2e4e36c1f6b4debc71c720bc" + } + } + }, + { + "attempt": "p1_testwrite_gemma4-31b-q4_v1_retry-20260802T003822Z-8061c2b0", + "classification": { + "source_run": "p1_testwrite_gemma4-31b-q4_v1", + "classification": "infrastructure-invalid", + "reason_code": "harness-dispatch-keyerror-required-tool-argument", + "incident_document": "canonical-harness-incident-20260802.json", + "classified_before_grade": true, + "affirmative_evidence": [ + "a syntactically valid tool call omitted path", + "the harness raised KeyError instead of returning a recoverable tool error" + ], + "expected_files": { + "receipt.json": "92313dc9d6070ad4fba78fb89a4d043281e17d0e5fb2203e0bdc660c551669e4", + "transcript.jsonl": "ca8b0f10985799cbd7511a4ce32d6b51939f6d6f32139719cb41b9c52d18aeb4" + }, + "replacement": { + "required": true, + "canonical_run": "p1_testwrite_gemma4-31b-q4_v1", + "status": "completed" + } + }, + "files": { + "receipt.json": { + "bytes": 8176, + "sha256": "92313dc9d6070ad4fba78fb89a4d043281e17d0e5fb2203e0bdc660c551669e4" + }, + "transcript.jsonl": { + "bytes": 9815, + "sha256": "ca8b0f10985799cbd7511a4ce32d6b51939f6d6f32139719cb41b9c52d18aeb4" + } + } + }, + { + "attempt": "p1_testwrite_gemma4-31b-q4_v1_retry-20260802T004533Z-09111a0e", + "classification": { + "source_run": "p1_testwrite_gemma4-31b-q4_v1", + "classification": "infrastructure-invalid", + "reason_code": "operator-interrupted-before-required-telemetry-gate", + "incident_document": "canonical-telemetry-gate-incident-20260802.json", + "classified_before_grade": true, + "affirmative_evidence": [ + "the retry began before per-run sidecar attribution was proven", + "the supervisor stopped before the attempt completed" + ], + "expected_files": { + "receipt.json": "e163856856b1a8323e4b881d00a43f0a4361ad00f8d5e800acb3a6d435d259a5", + "transcript.jsonl": "e0efc1aebe693ac92e2a78354316059e5332e761979bdbee220a42019e0367f3" + }, + "replacement": { + "required": true, + "canonical_run": "p1_testwrite_gemma4-31b-q4_v1", + "status": "completed" + } + }, + "files": { + "receipt.json": { + "bytes": 8172, + "sha256": "e163856856b1a8323e4b881d00a43f0a4361ad00f8d5e800acb3a6d435d259a5" + }, + "transcript.jsonl": { + "bytes": 14568, + "sha256": "e0efc1aebe693ac92e2a78354316059e5332e761979bdbee220a42019e0367f3" + } + } + }, + { + "attempt": "p1_testwrite_gemma4-31b-q4_v3_retry-20260802T005418Z-ff262c4a", + "classification": { + "source_run": "p1_testwrite_gemma4-31b-q4_v3", + "classification": "infrastructure-invalid", + "reason_code": "operator-interrupted-companion-after-harness-dispatch-defect", + "incident_document": "canonical-harness-incident-20260802.json", + "classified_before_grade": true, + "affirmative_evidence": [ + "the attempt was active under the same defective harness", + "the supervisor stopped before completion and the attempt was preserved without a grade" + ], + "expected_files": { + "receipt.json": "326934f8660018bdd081c3740222199dc6d5851d011489012ae93166c51440eb", + "transcript.jsonl": "63717381ad72db6d006a97dc24a1da14a557a92542826be3d92c264968556661" + }, + "replacement": { + "required": true, + "canonical_run": "p1_testwrite_gemma4-31b-q4_v3", + "status": "completed" + } + }, + "files": { + "receipt.json": { + "bytes": 8175, + "sha256": "326934f8660018bdd081c3740222199dc6d5851d011489012ae93166c51440eb" + }, + "transcript.jsonl": { + "bytes": 6114, + "sha256": "63717381ad72db6d006a97dc24a1da14a557a92542826be3d92c264968556661" + } + } + } + ] +} diff --git a/benchmarks/gemma4-31b-q4/gemma4-canonical-n3-grader-manifest.json b/benchmarks/gemma4-31b-q4/gemma4-canonical-n3-grader-manifest.json new file mode 100644 index 00000000..badf272e --- /dev/null +++ b/benchmarks/gemma4-31b-q4/gemma4-canonical-n3-grader-manifest.json @@ -0,0 +1,386 @@ +{ + "schema_version": 1, + "generated_at": "2026-08-02T03:06:02.921035+00:00", + "root": "/home/michael/bench-gemma4-31b-q4", + "label": "gemma4-31b-q4", + "target_n": 3, + "passed": true, + "errors": [], + "repository_commit": "c7d0ee9bbd287716a5ef11a077ec3d89652ce25d", + "python": "3.12.3 (main, Jun 19 2026, 12:46:00) [GCC 13.3.0]", + "docker_version": "29.2.1", + "bench_sandbox_image_id": "sha256:61fae3bdff98a8a307a3aff35d482cf6ad9bf31f83c151499c99ae4cee64a9cd", + "grader_files": { + "tooling/scripts/grade_microbench.sh": { + "bytes": 5789, + "sha256": "fb65e0be31abeb86ece4cb1eef6c894200266b57fe48a069952c53af84016c9d" + }, + "tooling/graders/phase1_grade.py": { + "bytes": 12036, + "sha256": "cccf4967bccb7e013d5ea29a21cba0b274a8c81037d6bbe9d91e3428389875cb" + }, + "tooling/graders/code_task_grader.py": { + "bytes": 9546, + "sha256": "82aeb55f87f123c7920458a5346412378cfd3f9db100d924f51ea2f4a91a1eb5" + }, + "tooling/graders/phase2_extraction_grade.py": { + "bytes": 4958, + "sha256": "9e390ab2531e72aab64613382b01ce83e60ca6242e9abef87a0deefb502726d6" + }, + "tooling/graders/phase2_ci_failure_grade.py": { + "bytes": 5242, + "sha256": "6ae717a937da8abc4aca74621f67445667ea34b06767d64a8310f0b5fae53ab4" + }, + "tooling/graders/phase2_hallucination_grade.py": { + "bytes": 4945, + "sha256": "24e64182b04810bb3938ecdd9d3886d994e440716f4889952edb69786c7bb237" + }, + "tooling/graders/phase2_triage_grade.py": { + "bytes": 4493, + "sha256": "ea3eebcf6a2fc085276e555f2cb1a54239e1dc9286bac3cb2300f01929b855a1" + }, + "tooling/graders/phase3_doc_synthesis_grade.py": { + "bytes": 4068, + "sha256": "8c117c0a3bfec94a91f94e49bb5b4281a458bcc4a650e9ae86e2450f0f5cac0a" + }, + "tooling/graders/phase3_business_memo_grade.py": { + "bytes": 5996, + "sha256": "691c5c6cec0f3bbf2fc86e80c47dd3fc2bde4873dc630297a0aea13edd548d7b" + }, + "tooling/graders/phase3_market_research_grade.py": { + "bytes": 4585, + "sha256": "e87c20a55bb74e9be883a9b2628fd954a99aaebf937f1d1bd4fbbf111b836de6" + }, + "tooling/graders/phase3_writing_editing_grade.py": { + "bytes": 5520, + "sha256": "a3461f31be9d24cc1c5132fff4bff96a27af77af2457dad839ef5498afba8f56" + }, + "tooling/graders/phase3_project_mgmt_grade.py": { + "bytes": 5706, + "sha256": "e1c3a9190e6d19cffd76734c8c922450b22d674bd5fd2db052fef31f91777bca" + }, + "tooling/graders/ground_truth/phase2_extraction.json": { + "bytes": 4012, + "sha256": "fead89b8051621404257afd0d5cbd5edb3b81cfcc4fe94fd58cf59e546a82d67" + }, + "tooling/graders/ground_truth/phase2_hallucination.json": { + "bytes": 4847, + "sha256": "eb2d03bbb161d1619306904a1850b6ac7106e35cbc49febff431e0f7280900c3" + }, + "tooling/graders/ground_truth/phase2_triage.json": { + "bytes": 5105, + "sha256": "6ca3763f32abfe3a5b8dfa56662c9083f91f81c2c3d06d3117bfb974367669d1" + }, + "tooling/graders/ground_truth/phase3_doc_synthesis.json": { + "bytes": 3569, + "sha256": "20e2baef1b99f4dc9339e87208105c9c5a57cf6b6266f09ce88518d5931ebe9d" + }, + "tooling/graders/ground_truth/phase3_business_memo.json": { + "bytes": 5197, + "sha256": "d422f262007c219a5b2da654e9c1282bfb94e19837eafd33ca70009e66b72313" + }, + "tooling/graders/ground_truth/phase3_market_research_rubric.json": { + "bytes": 3428, + "sha256": "d67aaf1664f2280d1e47855aaeaf543590f40b34e7d42aa194729e5e28fa90a5" + }, + "tooling/graders/ground_truth/phase3_project_mgmt.json": { + "bytes": 4382, + "sha256": "ff97a618bcb888a7f4d67eba55b2e5b1d0b459c694b90db0f6ec4072e0280220" + }, + "tooling/inputs/phase3_writing_editing/audience_briefs.json": { + "bytes": 4254, + "sha256": "0c22573bd97d973d6de73cbbf2e9b806b3099f90872a5f0ec8a4fa33948f94cd" + } + }, + "raw_grades": [ + { + "run_name": "p1_bugfix_gemma4-31b-q4_v1", + "verdict": "PASS", + "grade_json": { + "bytes": 467, + "sha256": "8ef51f6f2d9457eb209c41700017a65b6f13674d9251e4c8b02352a7ad89a4aa" + } + }, + { + "run_name": "p1_bugfix_gemma4-31b-q4_v2", + "verdict": "PASS", + "grade_json": { + "bytes": 468, + "sha256": "8c1caead3b6ffa35ac77aefd6150cc3bc5513596a7f125c9e024d34226cdfe3f" + } + }, + { + "run_name": "p1_bugfix_gemma4-31b-q4_v3", + "verdict": "FAIL", + "grade_json": { + "bytes": 468, + "sha256": "10243cc2a1b0f2c8df720128ec698608433a7b5176a9ffe716efd29205815667" + } + }, + { + "run_name": "p1_testwrite_gemma4-31b-q4_v1", + "verdict": "FAIL", + "grade_json": { + "bytes": 1155, + "sha256": "73b7129ab31845b98ef25c4e1501e5367510c8b7fbda84357be4bb3294bbd2f3" + } + }, + { + "run_name": "p1_testwrite_gemma4-31b-q4_v2", + "verdict": "PASS", + "grade_json": { + "bytes": 341, + "sha256": "42c1da43e1c31be25775668d65010e502b8383356f35bfd846be7b583f7e153e" + } + }, + { + "run_name": "p1_testwrite_gemma4-31b-q4_v3", + "verdict": "PASS", + "grade_json": { + "bytes": 341, + "sha256": "b99e9530499f12935a79df989d2f182b0917530a4b3be44b95eaa26e7f2d2c33" + } + }, + { + "run_name": "p1_refactor_gemma4-31b-q4_v1", + "verdict": "PASS", + "grade_json": { + "bytes": 510, + "sha256": "7f69334378ac3187de489eb74ad81c8245267d1a5711612b94879f297e758bcf" + } + }, + { + "run_name": "p1_refactor_gemma4-31b-q4_v2", + "verdict": "PASS", + "grade_json": { + "bytes": 510, + "sha256": "7f69334378ac3187de489eb74ad81c8245267d1a5711612b94879f297e758bcf" + } + }, + { + "run_name": "p1_refactor_gemma4-31b-q4_v3", + "verdict": "PASS", + "grade_json": { + "bytes": 510, + "sha256": "7f69334378ac3187de489eb74ad81c8245267d1a5711612b94879f297e758bcf" + } + }, + { + "run_name": "p2_extract_gemma4-31b-q4_v1", + "verdict": "PASS", + "grade_json": { + "bytes": 3204, + "sha256": "b6bc97fc6e03c0b85c34d7e6b5aa1ffcfdf128e953247bc7569f9417170aa2b8" + } + }, + { + "run_name": "p2_extract_gemma4-31b-q4_v2", + "verdict": "PASS", + "grade_json": { + "bytes": 3204, + "sha256": "b6bc97fc6e03c0b85c34d7e6b5aa1ffcfdf128e953247bc7569f9417170aa2b8" + } + }, + { + "run_name": "p2_extract_gemma4-31b-q4_v3", + "verdict": "PASS", + "grade_json": { + "bytes": 3204, + "sha256": "b6bc97fc6e03c0b85c34d7e6b5aa1ffcfdf128e953247bc7569f9417170aa2b8" + } + }, + { + "run_name": "p2_ci_gemma4-31b-q4_v1", + "verdict": "PASS", + "grade_json": { + "bytes": 482, + "sha256": "69e86f389fb98ab93705abc5fd07431d5a9730ba5b0214d67cccdc8eeb154ba6" + } + }, + { + "run_name": "p2_ci_gemma4-31b-q4_v2", + "verdict": "PASS", + "grade_json": { + "bytes": 482, + "sha256": "eae0f41e05a5ce0e1a370e295ccb33bfd10b26e892ff34e286253904640fd666" + } + }, + { + "run_name": "p2_ci_gemma4-31b-q4_v3", + "verdict": "PASS", + "grade_json": { + "bytes": 482, + "sha256": "1cdcddf815967c7395bceb0d8ac5b955294d82d5d932594f553e43f22ab4234d" + } + }, + { + "run_name": "p2_hallucination_gemma4-31b-q4_v1", + "verdict": "PASS", + "grade_json": { + "bytes": 3586, + "sha256": "80910919c9e64a98ef035d0e5da409c4b26fc67de0ad65bcf63df6b4012393e6" + } + }, + { + "run_name": "p2_hallucination_gemma4-31b-q4_v2", + "verdict": "PASS", + "grade_json": { + "bytes": 3892, + "sha256": "2553e9edff94d4bf1cc8fda5c170ae25201bd4baa261fc4ab14d5754dbc11bed" + } + }, + { + "run_name": "p2_hallucination_gemma4-31b-q4_v3", + "verdict": "MISSING_OUTPUT", + "grade_json": { + "bytes": 153, + "sha256": "e4f8b9da1c7aba632e47e0a97771274c35662e1f1e2920b050928e498df8692a" + } + }, + { + "run_name": "p2_triage_gemma4-31b-q4_v1", + "verdict": "PASS", + "grade_json": { + "bytes": 2250, + "sha256": "dbe6f269c365da51b85a5053d621051e8877fc7a7642fcc441c97f6c7653b5f5" + } + }, + { + "run_name": "p2_triage_gemma4-31b-q4_v2", + "verdict": "PASS", + "grade_json": { + "bytes": 2159, + "sha256": "c14e218d16bc8ba2bfb56fd8577d0f8539d628d0366f9340f7d2f60f17c91290" + } + }, + { + "run_name": "p2_triage_gemma4-31b-q4_v3", + "verdict": "PASS", + "grade_json": { + "bytes": 1914, + "sha256": "8b145bdcbc912d2f546ce33b90d11fe348c69829b205abe91d6513b83ba9daef" + } + }, + { + "run_name": "p3_doc_gemma4-31b-q4_v1", + "verdict": "PASS", + "grade_json": { + "bytes": 2357, + "sha256": "addcd3cfc419649494adf40c60fd2385ec3b403f91e07f3423695aac6fd5c0dc" + } + }, + { + "run_name": "p3_doc_gemma4-31b-q4_v2", + "verdict": "PASS", + "grade_json": { + "bytes": 2357, + "sha256": "addcd3cfc419649494adf40c60fd2385ec3b403f91e07f3423695aac6fd5c0dc" + } + }, + { + "run_name": "p3_doc_gemma4-31b-q4_v3", + "verdict": "PASS", + "grade_json": { + "bytes": 2360, + "sha256": "a1e32ee1296ef6f8f122b96d91586e2fbd3df7e2ff1f5aea3b13c97c6cb2bdd8" + } + }, + { + "run_name": "p3_business_gemma4-31b-q4_v1", + "verdict": "FAIL", + "grade_json": { + "bytes": 3866, + "sha256": "3c85cb074672401082ba6adecc5f454e5124b81323f35fc312a8458bba0ab815" + } + }, + { + "run_name": "p3_business_gemma4-31b-q4_v2", + "verdict": "PASS", + "grade_json": { + "bytes": 3873, + "sha256": "13358399cce26683f59886f1cfc577330fd90233e4065a9af249f51edd05c4ab" + } + }, + { + "run_name": "p3_business_gemma4-31b-q4_v3", + "verdict": "PASS", + "grade_json": { + "bytes": 3873, + "sha256": "13358399cce26683f59886f1cfc577330fd90233e4065a9af249f51edd05c4ab" + } + }, + { + "run_name": "p3_market_gemma4-31b-q4_v1", + "verdict": "STRUCTURAL_PASS", + "grade_json": { + "bytes": 991, + "sha256": "ec7f817b56a4a6ab2128042d09466fdcdf5b77eaf246ea588cbb987a0865d5fa" + } + }, + { + "run_name": "p3_market_gemma4-31b-q4_v2", + "verdict": "STRUCTURAL_PASS", + "grade_json": { + "bytes": 991, + "sha256": "7a993d6b3110622d5558f7d78f2ebb57044bb1bdf2b0e9f52585b7fdcf81dddc" + } + }, + { + "run_name": "p3_market_gemma4-31b-q4_v3", + "verdict": "STRUCTURAL_PASS", + "grade_json": { + "bytes": 990, + "sha256": "5d1bcb2c532b0323a5acdb93607cd0a957b6254f0dc687a6751be84f10ba86b2" + } + }, + { + "run_name": "p3_writing_gemma4-31b-q4_v1", + "verdict": "PASS", + "grade_json": { + "bytes": 2879, + "sha256": "4d52cac7e70665c17a000630f5e703deb8c2dd7fa09b140e58cd916809855594" + } + }, + { + "run_name": "p3_writing_gemma4-31b-q4_v2", + "verdict": "PASS", + "grade_json": { + "bytes": 2879, + "sha256": "d5b96fa97f5173b2c03372286d3f1337cdad83a2d8433953285d8c1b96d080de" + } + }, + { + "run_name": "p3_writing_gemma4-31b-q4_v3", + "verdict": "PASS", + "grade_json": { + "bytes": 2876, + "sha256": "5554d7976a5268e573cd20176244ba7db50716c7dc7dad7df89386f9fc269564" + } + }, + { + "run_name": "p3_pm_gemma4-31b-q4_v1", + "verdict": "FAIL", + "grade_json": { + "bytes": 2404, + "sha256": "eade149c12512150ebb28d5a446c6dc1037e0ded24a3ed5e83d65d87aba47bc4" + } + }, + { + "run_name": "p3_pm_gemma4-31b-q4_v2", + "verdict": "FAIL", + "grade_json": { + "bytes": 2425, + "sha256": "5e98c1230f111ddb4d1a15ae7063c80590fbee19b72661c4b56e9c94f0b12352" + } + }, + { + "run_name": "p3_pm_gemma4-31b-q4_v3", + "verdict": "FAIL", + "grade_json": { + "bytes": 2408, + "sha256": "be7951b5c0b3770b9ee61e8f3bd5042dfc0fa486d2be1086b97314e4a8976ca3" + } + } + ], + "correction_policy": "grade.json and label.json hashes are immutable raw evidence. Any correction must be a separate overlay naming the original hash, corrected grader hash, unchanged workspace archive hash, and reproducible defect." +} diff --git a/benchmarks/gemma4-31b-q4/gemma4-canonical-n3-project-mgmt-correction.json b/benchmarks/gemma4-31b-q4/gemma4-canonical-n3-project-mgmt-correction.json new file mode 100644 index 00000000..e13a276e --- /dev/null +++ b/benchmarks/gemma4-31b-q4/gemma4-canonical-n3-project-mgmt-correction.json @@ -0,0 +1,520 @@ +{ + "schema_version": 1, + "generated_at": "2026-08-02T03:06:03.021824+00:00", + "campaign": "gemma4-31b-q4-mmbt", + "target_n": 3, + "scope": "p3_pm legacy lexical false negatives only", + "policy": "immutable grade.json files remain raw evidence; this overlay changes no run artifact or raw verdict", + "correction_script": { + "path": "/home/michael/bench-gemma4-31b-q4/tooling/correct_gemma4_project_mgmt_grades.py", + "sha256": "84db7554665fbecae264c059839be16f9bc0cf3e5870c3d7b34c5c8c2f16be74" + }, + "raw_grader": { + "path": "/home/michael/bench-gemma4-31b-q4/tooling/graders/phase3_project_mgmt_grade.py", + "sha256": "e1c3a9190e6d19cffd76734c8c922450b22d674bd5fd2db052fef31f91777bca" + }, + "rules": { + "R2": { + "description": "Maevia GA-to-private-beta pushback is expressed without one contiguous legacy keyword", + "patterns": [ + "\\bmaevia\\b.{0,240}\\b(?:push(?:ed)?[ -]?back|fallout|promis(?:e|ed)\\s+ga|private[ -]?beta)\\b" + ] + }, + "R3": { + "description": "legal/private-beta contract delay uses an equivalent non-contracted phrase", + "patterns": [ + "\\blegal\\b.{0,200}\\b(?:has\\s+not|not\\s+yet\\s+responded|sign[ -]?off|approval|unresponsive|silent|delay|contract)\\b" + ] + }, + "D3_mobile": { + "description": "web-responsive V1 and native V2 uses a hyphen the legacy keyword omits", + "patterns": [ + "\\bweb[ -]?responsive\\b.{0,160}\\bnative\\b.{0,100}\\bv2\\b" + ] + }, + "D4_option_b": { + "description": "private beta followed by the selected-customer count is Option B", + "patterns": [ + "\\bprivate[ -]?beta\\b.{0,120}\\b(?:3\\s*[-\u2013]\\s*5|selected\\s+customers?)\\b" + ] + } + }, + "aggregate": { + "cells": 3, + "raw_passes": 0, + "corrected_passes": 3, + "verdict_changes": 3 + }, + "cells": [ + { + "run_name": "p3_pm_gemma4-31b-q4_v1", + "raw_grade_sha256": "eade149c12512150ebb28d5a446c6dc1037e0ded24a3ed5e83d65d87aba47bc4", + "workspace_archive_sha256": "eb622a02fddb70c3aee43f652123e59640282815e55b230d70874175a61fcf08", + "status_report_sha256": "3c6f102bd297b6e8eaaeb15e0b1e67cadc359f09d01c3fd7b0f35005161ca48f", + "raw_verdict": "FAIL", + "corrected_verdict": "PASS", + "changes": [ + { + "item": "R2", + "category": "risks", + "reason": "Maevia GA-to-private-beta pushback is expressed without one contiguous legacy keyword", + "pattern": "\\bmaevia\\b.{0,240}\\b(?:push(?:ed)?[ -]?back|fallout|promis(?:e|ed)\\s+ga|private[ -]?beta)\\b" + }, + { + "item": "R3", + "category": "risks", + "reason": "legal/private-beta contract delay uses an equivalent non-contracted phrase", + "pattern": "\\blegal\\b.{0,200}\\b(?:has\\s+not|not\\s+yet\\s+responded|sign[ -]?off|approval|unresponsive|silent|delay|contract)\\b" + }, + { + "item": "D3_mobile", + "category": "decisions", + "reason": "web-responsive V1 and native V2 uses a hyphen the legacy keyword omits", + "pattern": "\\bweb[ -]?responsive\\b.{0,160}\\bnative\\b.{0,100}\\bv2\\b" + }, + { + "item": "D4_option_b", + "category": "decisions", + "reason": "private beta followed by the selected-customer count is Option B", + "pattern": "\\bprivate[ -]?beta\\b.{0,120}\\b(?:3\\s*[-\u2013]\\s*5|selected\\s+customers?)\\b" + } + ], + "corrected_grade": { + "task": "project_mgmt", + "verdict": "PASS", + "scores": { + "workstream_recall": "6/6", + "risk_recall": "4/6", + "decision_recall": "4/4", + "milestone_recall": "4/5", + "sections_present": [ + "workstream", + "risk", + "decision", + "milestone" + ], + "word_count": 321 + }, + "thresholds": { + "min_workstreams": 4, + "min_risks": 3, + "min_decisions": 3, + "min_milestones": 3, + "max_word_count": 700 + }, + "details": { + "workstreams": { + "WS1": { + "matched": true, + "keyword": "query layer" + }, + "WS2": { + "matched": true, + "keyword": "dashboard refresh" + }, + "WS3": { + "matched": true, + "keyword": "40-panel" + }, + "WS4": { + "matched": true, + "keyword": "access control" + }, + "WS5": { + "matched": true, + "keyword": "maevia" + }, + "WS6": { + "matched": true, + "keyword": "legal" + } + }, + "risks": { + "R1": { + "matched": false, + "keyword": null + }, + "R2": { + "matched": true, + "keyword": "correction-overlay-semantic-equivalent" + }, + "R3": { + "matched": true, + "keyword": "correction-overlay-semantic-equivalent" + }, + "R4": { + "matched": true, + "keyword": "40-panel" + }, + "R5": { + "matched": true, + "keyword": "custom branding" + }, + "R6": { + "matched": false, + "keyword": null + } + }, + "decisions": { + "D1_branding": { + "matched": true, + "keyword": "branding deferred" + }, + "D2_panel_limit": { + "matched": true, + "keyword": "40-panel" + }, + "D3_mobile": { + "matched": true, + "keyword": "correction-overlay-semantic-equivalent" + }, + "D4_option_b": { + "matched": true, + "keyword": "correction-overlay-semantic-equivalent" + } + }, + "milestones": { + "M1": { + "matched": true, + "keyword": "mid-may" + }, + "M2": { + "matched": true, + "keyword": "mid-may" + }, + "M3": { + "matched": false, + "keyword": null + }, + "M4": { + "matched": true, + "keyword": "v1.1" + }, + "M5": { + "matched": true, + "keyword": "deferred" + } + } + }, + "hand_rating_placeholders": { + "structure_quality_1to5": null, + "fabrication_count": null, + "owner_accuracy_0to6": null, + "rater_notes": null + } + } + }, + { + "run_name": "p3_pm_gemma4-31b-q4_v2", + "raw_grade_sha256": "5e98c1230f111ddb4d1a15ae7063c80590fbee19b72661c4b56e9c94f0b12352", + "workspace_archive_sha256": "6190888ab5900349d69d7d443477eb162cb5174acc340c36c5489a95b9a591c8", + "status_report_sha256": "57d5b455eb57566d5d616c0371fb5818e6fe643f1262799089be12f4b45e6588", + "raw_verdict": "FAIL", + "corrected_verdict": "PASS", + "changes": [ + { + "item": "R2", + "category": "risks", + "reason": "Maevia GA-to-private-beta pushback is expressed without one contiguous legacy keyword", + "pattern": "\\bmaevia\\b.{0,240}\\b(?:push(?:ed)?[ -]?back|fallout|promis(?:e|ed)\\s+ga|private[ -]?beta)\\b" + }, + { + "item": "R3", + "category": "risks", + "reason": "legal/private-beta contract delay uses an equivalent non-contracted phrase", + "pattern": "\\blegal\\b.{0,200}\\b(?:has\\s+not|not\\s+yet\\s+responded|sign[ -]?off|approval|unresponsive|silent|delay|contract)\\b" + }, + { + "item": "D4_option_b", + "category": "decisions", + "reason": "private beta followed by the selected-customer count is Option B", + "pattern": "\\bprivate[ -]?beta\\b.{0,120}\\b(?:3\\s*[-\u2013]\\s*5|selected\\s+customers?)\\b" + } + ], + "corrected_grade": { + "task": "project_mgmt", + "verdict": "PASS", + "scores": { + "workstream_recall": "6/6", + "risk_recall": "4/6", + "decision_recall": "4/4", + "milestone_recall": "5/5", + "sections_present": [ + "workstream", + "risk", + "decision", + "milestone" + ], + "word_count": 313 + }, + "thresholds": { + "min_workstreams": 4, + "min_risks": 3, + "min_decisions": 3, + "min_milestones": 3, + "max_word_count": 700 + }, + "details": { + "workstreams": { + "WS1": { + "matched": true, + "keyword": "query layer" + }, + "WS2": { + "matched": true, + "keyword": "mobile-responsive" + }, + "WS3": { + "matched": true, + "keyword": "40-panel" + }, + "WS4": { + "matched": true, + "keyword": "access-control" + }, + "WS5": { + "matched": true, + "keyword": "maevia" + }, + "WS6": { + "matched": true, + "keyword": "legal" + } + }, + "risks": { + "R1": { + "matched": false, + "keyword": null + }, + "R2": { + "matched": true, + "keyword": "correction-overlay-semantic-equivalent" + }, + "R3": { + "matched": true, + "keyword": "correction-overlay-semantic-equivalent" + }, + "R4": { + "matched": true, + "keyword": "40-panel" + }, + "R5": { + "matched": true, + "keyword": "custom branding" + }, + "R6": { + "matched": false, + "keyword": null + } + }, + "decisions": { + "D1_branding": { + "matched": true, + "keyword": "cut custom branding" + }, + "D2_panel_limit": { + "matched": true, + "keyword": "40-panel" + }, + "D3_mobile": { + "matched": true, + "keyword": "mobile responsive" + }, + "D4_option_b": { + "matched": true, + "keyword": "correction-overlay-semantic-equivalent" + } + }, + "milestones": { + "M1": { + "matched": true, + "keyword": "mid-may" + }, + "M2": { + "matched": true, + "keyword": "mid-may" + }, + "M3": { + "matched": true, + "keyword": "mid-july" + }, + "M4": { + "matched": true, + "keyword": "v1.1" + }, + "M5": { + "matched": true, + "keyword": "deferred" + } + } + }, + "hand_rating_placeholders": { + "structure_quality_1to5": null, + "fabrication_count": null, + "owner_accuracy_0to6": null, + "rater_notes": null + } + } + }, + { + "run_name": "p3_pm_gemma4-31b-q4_v3", + "raw_grade_sha256": "be7951b5c0b3770b9ee61e8f3bd5042dfc0fa486d2be1086b97314e4a8976ca3", + "workspace_archive_sha256": "fac5112297764b16bb742a38fcecc8617524ef13ebe869170e049842e63a4645", + "status_report_sha256": "1df31af7e0ef549e4a2db3fdae74f8aba184c79d530d810baae0c09215dac7b3", + "raw_verdict": "FAIL", + "corrected_verdict": "PASS", + "changes": [ + { + "item": "R2", + "category": "risks", + "reason": "Maevia GA-to-private-beta pushback is expressed without one contiguous legacy keyword", + "pattern": "\\bmaevia\\b.{0,240}\\b(?:push(?:ed)?[ -]?back|fallout|promis(?:e|ed)\\s+ga|private[ -]?beta)\\b" + }, + { + "item": "R3", + "category": "risks", + "reason": "legal/private-beta contract delay uses an equivalent non-contracted phrase", + "pattern": "\\blegal\\b.{0,200}\\b(?:has\\s+not|not\\s+yet\\s+responded|sign[ -]?off|approval|unresponsive|silent|delay|contract)\\b" + }, + { + "item": "D3_mobile", + "category": "decisions", + "reason": "web-responsive V1 and native V2 uses a hyphen the legacy keyword omits", + "pattern": "\\bweb[ -]?responsive\\b.{0,160}\\bnative\\b.{0,100}\\bv2\\b" + }, + { + "item": "D4_option_b", + "category": "decisions", + "reason": "private beta followed by the selected-customer count is Option B", + "pattern": "\\bprivate[ -]?beta\\b.{0,120}\\b(?:3\\s*[-\u2013]\\s*5|selected\\s+customers?)\\b" + } + ], + "corrected_grade": { + "task": "project_mgmt", + "verdict": "PASS", + "scores": { + "workstream_recall": "6/6", + "risk_recall": "4/6", + "decision_recall": "4/4", + "milestone_recall": "4/5", + "sections_present": [ + "workstream", + "risk", + "decision", + "milestone" + ], + "word_count": 303 + }, + "thresholds": { + "min_workstreams": 4, + "min_risks": 3, + "min_decisions": 3, + "min_milestones": 3, + "max_word_count": 700 + }, + "details": { + "workstreams": { + "WS1": { + "matched": true, + "keyword": "query-layer" + }, + "WS2": { + "matched": true, + "keyword": "dashboard refresh" + }, + "WS3": { + "matched": true, + "keyword": "architectural fix" + }, + "WS4": { + "matched": true, + "keyword": "access control" + }, + "WS5": { + "matched": true, + "keyword": "maevia" + }, + "WS6": { + "matched": true, + "keyword": "legal" + } + }, + "risks": { + "R1": { + "matched": false, + "keyword": null + }, + "R2": { + "matched": true, + "keyword": "correction-overlay-semantic-equivalent" + }, + "R3": { + "matched": true, + "keyword": "correction-overlay-semantic-equivalent" + }, + "R4": { + "matched": true, + "keyword": "40-panel" + }, + "R5": { + "matched": true, + "keyword": "custom-branding" + }, + "R6": { + "matched": false, + "keyword": null + } + }, + "decisions": { + "D1_branding": { + "matched": true, + "keyword": "branding cut" + }, + "D2_panel_limit": { + "matched": true, + "keyword": "40-panel" + }, + "D3_mobile": { + "matched": true, + "keyword": "correction-overlay-semantic-equivalent" + }, + "D4_option_b": { + "matched": true, + "keyword": "correction-overlay-semantic-equivalent" + } + }, + "milestones": { + "M1": { + "matched": true, + "keyword": "mid-may" + }, + "M2": { + "matched": true, + "keyword": "mid-may" + }, + "M3": { + "matched": false, + "keyword": null + }, + "M4": { + "matched": true, + "keyword": "v1.1" + }, + "M5": { + "matched": true, + "keyword": "deferred" + } + } + }, + "hand_rating_placeholders": { + "structure_quality_1to5": null, + "fabrication_count": null, + "owner_accuracy_0to6": null, + "rater_notes": null + } + } + } + ], + "errors": [], + "passed": true +} diff --git a/benchmarks/gemma4-31b-q4/gemma4-canonical-n3-scorecard.json b/benchmarks/gemma4-31b-q4/gemma4-canonical-n3-scorecard.json new file mode 100644 index 00000000..81eb8e30 --- /dev/null +++ b/benchmarks/gemma4-31b-q4/gemma4-canonical-n3-scorecard.json @@ -0,0 +1,2435 @@ +{ + "schema_version": 1, + "generated_at": "2026-08-02T03:06:02.993594+00:00", + "label": "gemma4-31b-q4", + "target_n": 3, + "passed": true, + "errors": [], + "aggregate": { + "runs": 36, + "completed": 36, + "normal_completed": 36, + "terminal_outcomes": 0, + "graded": 36, + "scored_outcomes": 36, + "raw_passes": 29, + "raw_pass_rate": 0.805556, + "verdicts": { + "FAIL": 6, + "MISSING_OUTPUT": 1, + "PASS": 26, + "STRUCTURAL_PASS": 3 + }, + "quality_outcomes": { + "FAIL": 6, + "MISSING_OUTPUT": 1, + "PASS": 26, + "STRUCTURAL_PASS": 3 + }, + "finish_reasons": { + "done_signal": 35, + "model_stopped": 1 + }, + "wall_s": { + "sum": 7164.3, + "median": 122.8, + "max": 708.7 + }, + "completion_tokens": { + "sum": 291243, + "median": 6332.5, + "max": 25223 + }, + "model_call_completion_tps": { + "median": 54.5, + "min": 19.7, + "max": 61.9 + }, + "telemetry": { + "runs": 34, + "coverage_mean": 0.944515, + "active_gpu_mean_power_w_mean": 465.854, + "active_gpu_mean_sm_util_pct_mean": 87.904, + "max_temp_c": 87.0 + }, + "per_task": { + "p1_bugfix": { + "runs": 3, + "graded": 3, + "scored_outcomes": 3, + "raw_passes": 2, + "verdicts": { + "FAIL": 1, + "PASS": 2 + }, + "quality_outcomes": { + "FAIL": 1, + "PASS": 2 + }, + "finish_reasons": { + "done_signal": 3 + } + }, + "p1_testwrite": { + "runs": 3, + "graded": 3, + "scored_outcomes": 3, + "raw_passes": 2, + "verdicts": { + "FAIL": 1, + "PASS": 2 + }, + "quality_outcomes": { + "FAIL": 1, + "PASS": 2 + }, + "finish_reasons": { + "done_signal": 3 + } + }, + "p1_refactor": { + "runs": 3, + "graded": 3, + "scored_outcomes": 3, + "raw_passes": 3, + "verdicts": { + "PASS": 3 + }, + "quality_outcomes": { + "PASS": 3 + }, + "finish_reasons": { + "done_signal": 3 + } + }, + "p2_extract": { + "runs": 3, + "graded": 3, + "scored_outcomes": 3, + "raw_passes": 3, + "verdicts": { + "PASS": 3 + }, + "quality_outcomes": { + "PASS": 3 + }, + "finish_reasons": { + "done_signal": 3 + } + }, + "p2_ci": { + "runs": 3, + "graded": 3, + "scored_outcomes": 3, + "raw_passes": 3, + "verdicts": { + "PASS": 3 + }, + "quality_outcomes": { + "PASS": 3 + }, + "finish_reasons": { + "done_signal": 3 + } + }, + "p2_hallucination": { + "runs": 3, + "graded": 3, + "scored_outcomes": 3, + "raw_passes": 2, + "verdicts": { + "MISSING_OUTPUT": 1, + "PASS": 2 + }, + "quality_outcomes": { + "MISSING_OUTPUT": 1, + "PASS": 2 + }, + "finish_reasons": { + "done_signal": 2, + "model_stopped": 1 + } + }, + "p2_triage": { + "runs": 3, + "graded": 3, + "scored_outcomes": 3, + "raw_passes": 3, + "verdicts": { + "PASS": 3 + }, + "quality_outcomes": { + "PASS": 3 + }, + "finish_reasons": { + "done_signal": 3 + } + }, + "p3_doc": { + "runs": 3, + "graded": 3, + "scored_outcomes": 3, + "raw_passes": 3, + "verdicts": { + "PASS": 3 + }, + "quality_outcomes": { + "PASS": 3 + }, + "finish_reasons": { + "done_signal": 3 + } + }, + "p3_business": { + "runs": 3, + "graded": 3, + "scored_outcomes": 3, + "raw_passes": 2, + "verdicts": { + "FAIL": 1, + "PASS": 2 + }, + "quality_outcomes": { + "FAIL": 1, + "PASS": 2 + }, + "finish_reasons": { + "done_signal": 3 + } + }, + "p3_market": { + "runs": 3, + "graded": 3, + "scored_outcomes": 3, + "raw_passes": 3, + "verdicts": { + "STRUCTURAL_PASS": 3 + }, + "quality_outcomes": { + "STRUCTURAL_PASS": 3 + }, + "finish_reasons": { + "done_signal": 3 + } + }, + "p3_writing": { + "runs": 3, + "graded": 3, + "scored_outcomes": 3, + "raw_passes": 3, + "verdicts": { + "PASS": 3 + }, + "quality_outcomes": { + "PASS": 3 + }, + "finish_reasons": { + "done_signal": 3 + } + }, + "p3_pm": { + "runs": 3, + "graded": 3, + "scored_outcomes": 3, + "raw_passes": 0, + "verdicts": { + "FAIL": 3 + }, + "quality_outcomes": { + "FAIL": 3 + }, + "finish_reasons": { + "done_signal": 3 + } + } + } + }, + "runs": [ + { + "run_name": "p1_bugfix_gemma4-31b-q4_v1", + "task": "p1_bugfix", + "replicate": 1, + "summary": { + "started_at": "2026-08-02T00:21:07.782391+00:00", + "ended_at": "2026-08-02T00:32:56.494903+00:00", + "elapsed_s": 708.7, + "iterations": 80, + "total_completion_tokens": 18475, + "total_prompt_tokens": 1583932, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Stabilized the `logalyzer` codebase by fixing critical bugs and performance bottlenecks, removing a security vulnerability, and increasing test coverage to 91%.\n\nKey accomplishments:\n- **Performance**: Fixed an O(n\u00b2) log loading bottleneck, reducing 50MB log load time from >300s to ~3s.\n- **Security**: Removed the `eval()`-based `expression_filter` and corresponding CLI argument to prevent remote code execution.\n- **Stability**: Fixed a missing month in date parsing, corrected URL regex matching, and normalized CLI input types.\n- **Modernization**: Updated the codebase for Python 3.11+ and replaced manual JSON string building with the `json` module.\n- **Quality**: Removed dead code, added comprehensive tests for all modules, and achieved 91% line coverage.\n- **Documentation**: Created a full audit trail including a triage report, before-after metrics, ADRs, research notes, and a detailed CHANGELOG.\n\nThe final repository is located at `/workspace/` and is tagged as `v0.3.2`." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p1_bugfix_gemma4-31b-q4_v1", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 708.7, + "iters": 80, + "tokens": { + "completion_total": 18475, + "prompt_total": 1583932, + "all_total": 1602407 + }, + "throughput": { + "completion_tps_avg": 47.6, + "model_call_wall_s": 388.0, + "tool_call_wall_s": 314.2, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 404.7, + "first_commit_s": 0.1, + "first_tag_s": 689.9 + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 20.92, + "memory_used_mib_at_start": 65197.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0984, + "cost_usd_upper_bound": 0.0128, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": null + }, + { + "run_name": "p1_bugfix_gemma4-31b-q4_v2", + "task": "p1_bugfix", + "replicate": 2, + "summary": { + "started_at": "2026-08-02T00:21:07.781614+00:00", + "ended_at": "2026-08-02T00:31:47.326458+00:00", + "elapsed_s": 639.5, + "iterations": 58, + "total_completion_tokens": 14733, + "total_prompt_tokens": 984362, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "I have successfully adopted and stabilized the `logalyzer` codebase. \n\n**Key Accomplishments:**\n- **Fixed Critical Performance Issue**: Identified and fixed an O(n\u00b2) bottleneck in `logalyzer.io.load` where list concatenation was used in a loop. This resulted in a speedup from >300s to ~3.2s for a 50MB log file (approx. 90x improvement).\n- **Resolved Bugs**: \n - Fixed `parse_timestamp` missing 'Jul' in its lookup table, which caused crashes for July logs.\n - Fixed `status_filter` to correctly handle string inputs from the CLI.\n - Fixed `url_regex_filter` to perform substring matches using `re.search` instead of `re.match`.\n- **Modernized Codebase**: \n - Updated `collections.Iterable` to `collections.abc.Iterable` for Python 3.11 compatibility.\n - Optimized `ip_allowlist_filter` by converting the allowlist to a `set`, reducing lookup time from O(N) to O(1) per entry.\n- **Improved Quality & Coverage**:\n - Increased test coverage from 34% to 85% by adding comprehensive tests for `cli`, `io`, `output`, and `utils` modules.\n - Removed dead code (`logalyzer/experimental.py`).\n - Cleaned up unused imports and followed PEP8 standards using `ruff`.\n- **Documentation & Audit**:\n - Provided a full triage of issues in `/audit/triage.md`.\n - Documented hard metrics in `/audit/before-after.md`.\n - Created Architecture Decision Records (ADRs) and research logs in `/architecture/`, `/decisions/`, and `/research/`.\n - Updated `README.md` to remove false claims and accurately describe the tool's capabilities.\n - Created a user-facing `CHANGELOG.md`.\n\nThe final repository is tagged `v0.3.0` and is ready for hand-off." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p1_bugfix_gemma4-31b-q4_v2", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 639.5, + "iters": 58, + "tokens": { + "completion_total": 14733, + "prompt_total": 984362, + "all_total": 999095 + }, + "throughput": { + "completion_tps_avg": 45.7, + "model_call_wall_s": 322.4, + "tool_call_wall_s": 312.4, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 329.9, + "first_commit_s": 6.1, + "first_tag_s": 605.9 + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 20.92, + "memory_used_mib_at_start": 65197.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0888, + "cost_usd_upper_bound": 0.0115, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": null + }, + { + "run_name": "p1_bugfix_gemma4-31b-q4_v3", + "task": "p1_bugfix", + "replicate": 3, + "summary": { + "started_at": "2026-08-02T00:45:34.628923+00:00", + "ended_at": "2026-08-02T00:55:58.294596+00:00", + "elapsed_s": 623.7, + "iterations": 62, + "total_completion_tokens": 14968, + "total_prompt_tokens": 1161167, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Adopted and stabilized the `logalyzer` codebase. \n\nKey accomplishments:\n1. **Bug Fixes**: Fixed critical timestamp parsing error (missing July) and filter logic bugs (status type mismatch, URL regex match vs search).\n2. **Modernization**: Updated `collections` imports to `collections.abc` for Python 3.10+ compatibility.\n3. **Performance**: Optimized `logalyzer.io.load` from O(n\u00b2) to O(n) by replacing quadratic string and list concatenation with a streaming generator, reducing 50MB log processing time from >300s to ~3.3s.\n4. **Security**: Eliminated a critical Remote Code Execution (RCE) vulnerability by removing the `eval()`-based `expression_filter` and corresponding CLI flag.\n5. **Code Cleanup**: Removed dead code including an unused experimental module and several legacy utility functions.\n6. **Testing**: Expanded the test suite to cover CLI, IO, and Output modules, increasing total line coverage from 24% to 91% and achieving a 100% pass rate.\n7. **Documentation**: Updated the README (removed false claims), created a detailed CHANGELOG, and provided full audit and research trails including ADRs and decision records.\n\nThe final repository is tagged as `v0.3.0`." + }, + "verdict": "FAIL", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p1_bugfix_gemma4-31b-q4_v3", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 623.7, + "iters": 62, + "tokens": { + "completion_total": 14968, + "prompt_total": 1161167, + "all_total": 1176135 + }, + "throughput": { + "completion_tps_avg": 48.9, + "model_call_wall_s": 305.9, + "tool_call_wall_s": 312.7, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 337.5, + "first_commit_s": 3.3, + "first_tag_s": 597.7 + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 87.25, + "memory_used_mib_at_start": 65453.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0866, + "cost_usd_upper_bound": 0.0113, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9941, + "mean_power_w": 249.36, + "mean_sm_util_pct": 44.38, + "max_temp_c": 76.0 + } + }, + { + "run_name": "p1_testwrite_gemma4-31b-q4_v1", + "task": "p1_testwrite", + "replicate": 1, + "summary": { + "started_at": "2026-08-02T00:45:34.629990+00:00", + "ended_at": "2026-08-02T00:54:18.228621+00:00", + "elapsed_s": 523.6, + "iterations": 49, + "total_completion_tokens": 25223, + "total_prompt_tokens": 1273638, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Increased test coverage of the logalyzer codebase from 25% to 72% by adding high-value tests for under-covered modules. \n\nKey accomplishments:\n- Added new test files: `tests/test_cli.py`, `tests/test_experimental.py`, `tests/test_io.py`, `tests/test_output.py`, `tests/test_utils.py`, and `tests/test_parser_edge.py`.\n- Achieved 100% coverage for `experimental.py`, `io.py`, `output.py`, and `utils.py`.\n- Significantly increased coverage for `cli.py` (98%), `filters.py` (97%), and `parser.py` (88%).\n- Implemented a monkeypatch in `tests/conftest.py` to resolve a Python 3.11 incompatibility in the production code (`collections.Iterable` removal) without modifying the production code itself.\n- Documented several baseline bugs (e.g., `status_filter` type mismatch, `re.match` instead of `re.search` in `url_regex_filter`, and a missing 'Jul' in `MONTH_NAMES`) in `research/dead-ends.md` as per the rules.\n- Maintained the original codebase's byte-identical status in `/logalyzer/`.\n- Followed the required repository structure and provided a detailed `CHANGELOG.md`.\n- Committed changes frequently with descriptive messages and tagged the final release.\n" + }, + "verdict": "FAIL", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p1_testwrite_gemma4-31b-q4_v1", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 523.6, + "iters": 49, + "tokens": { + "completion_total": 25223, + "prompt_total": 1273638, + "all_total": 1298861 + }, + "throughput": { + "completion_tps_avg": 49.1, + "model_call_wall_s": 514.0, + "tool_call_wall_s": 6.8, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 56.9, + "first_commit_s": 0.1, + "first_tag_s": 502.9 + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 87.25, + "memory_used_mib_at_start": 65453.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0727, + "cost_usd_upper_bound": 0.0095, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9836, + "mean_power_w": 481.16, + "mean_sm_util_pct": 91.31, + "max_temp_c": 81.0 + } + }, + { + "run_name": "p1_testwrite_gemma4-31b-q4_v2", + "task": "p1_testwrite", + "replicate": 2, + "summary": { + "started_at": "2026-08-02T00:56:00.566809+00:00", + "ended_at": "2026-08-02T01:02:29.405349+00:00", + "elapsed_s": 388.8, + "iterations": 48, + "total_completion_tokens": 18809, + "total_prompt_tokens": 1029851, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Increased test coverage of the logalyzer tool from 34% to 94% by adding high-value tests for previously uncovered modules (`experimental.py`, `io.py`, `output.py`, `utils.py`, and `cli.py`). All new tests pass, and baseline tests were preserved (with documented bugs in the starter codebase). Followed the requested directory structure and documented decisions in ADRs." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p1_testwrite_gemma4-31b-q4_v2", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 388.8, + "iters": 48, + "tokens": { + "completion_total": 18809, + "prompt_total": 1029851, + "all_total": 1048660 + }, + "throughput": { + "completion_tps_avg": 49.5, + "model_call_wall_s": 379.9, + "tool_call_wall_s": 6.1, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 24.6, + "first_commit_s": 6.5, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 500.03, + "memory_used_mib_at_start": 65453.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.054, + "cost_usd_upper_bound": 0.007, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9902, + "mean_power_w": 479.04, + "mean_sm_util_pct": 89.65, + "max_temp_c": 78.0 + } + }, + { + "run_name": "p1_testwrite_gemma4-31b-q4_v3", + "task": "p1_testwrite", + "replicate": 3, + "summary": { + "started_at": "2026-08-02T00:54:18.937486+00:00", + "ended_at": "2026-08-02T01:00:58.817170+00:00", + "elapsed_s": 399.9, + "iterations": 42, + "total_completion_tokens": 19900, + "total_prompt_tokens": 825275, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Increased test coverage of the logalyzer codebase from 34% to 98% by adding comprehensive test suites for previously uncovered modules. \n\nKey accomplishments:\n- Added `tests/test_experimental.py`, `tests/test_io.py`, `tests/test_output.py`, `tests/test_utils.py`, and `tests/test_cli.py`.\n- Improved `tests/test_parser.py` to cover error paths.\n- Implemented a monkeypatch in `tests/conftest.py` to handle `collections.Iterable` removal in Python 3.11, ensuring tests can run without modifying production code.\n- Documented three critical production bugs in `research/dead-ends.md` that caused baseline test failures (status filter type mismatch, URL regex matching behavior, and missing 'Jul' in MONTH_NAMES).\n- Maintained zero regressions to passing baseline tests.\n- Adhered strictly to the requirement of not modifying production code.\n" + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p1_testwrite_gemma4-31b-q4_v3", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 399.9, + "iters": 42, + "tokens": { + "completion_total": 19900, + "prompt_total": 825275, + "all_total": 845175 + }, + "throughput": { + "completion_tps_avg": 50.8, + "model_call_wall_s": 392.0, + "tool_call_wall_s": 5.4, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 24.0, + "first_commit_s": 4.7, + "first_tag_s": 380.1 + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 251.09, + "memory_used_mib_at_start": 65453.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0555, + "cost_usd_upper_bound": 0.0072, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9877, + "mean_power_w": 468.52, + "mean_sm_util_pct": 91.96, + "max_temp_c": 86.0 + } + }, + { + "run_name": "p1_refactor_gemma4-31b-q4_v1", + "task": "p1_refactor", + "replicate": 1, + "summary": { + "started_at": "2026-08-02T01:02:30.232102+00:00", + "ended_at": "2026-08-02T01:05:36.464935+00:00", + "elapsed_s": 186.2, + "iterations": 35, + "total_completion_tokens": 9779, + "total_prompt_tokens": 351967, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Refactored `logalyzer/output.py` into a focused package `logalyzer/output/` containing `json_renderer.py`, `csv_renderer.py`, `plain_renderer.py`, and `legacy.py`. Preserved the public API through `logalyzer/output/__init__.py`. Verified that existing call sites and tests remain unchanged in behavior, including the preservation of existing bugs in `filters.py` and `parser.py`. All documentation (ADRs, research notes, CHANGELOG) has been created in the required structure." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p1_refactor_gemma4-31b-q4_v1", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 186.2, + "iters": 35, + "tokens": { + "completion_total": 9779, + "prompt_total": 351967, + "all_total": 361746 + }, + "throughput": { + "completion_tps_avg": 53.9, + "model_call_wall_s": 181.5, + "tool_call_wall_s": 2.8, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 70.6, + "first_commit_s": 62.8, + "first_tag_s": 167.0 + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 498.71, + "memory_used_mib_at_start": 65453.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0259, + "cost_usd_upper_bound": 0.0034, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9667, + "mean_power_w": 465.36, + "mean_sm_util_pct": 85.35, + "max_temp_c": 78.0 + } + }, + { + "run_name": "p1_refactor_gemma4-31b-q4_v2", + "task": "p1_refactor", + "replicate": 2, + "summary": { + "started_at": "2026-08-02T01:00:59.565226+00:00", + "ended_at": "2026-08-02T01:04:31.084608+00:00", + "elapsed_s": 211.5, + "iterations": 41, + "total_completion_tokens": 11126, + "total_prompt_tokens": 508191, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Refactored `logalyzer/output.py` into a `logalyzer/output/` package with specialized modules (`json_renderer.py`, `csv_renderer.py`, `plain_renderer.py`, `legacy.py`) while maintaining full backward compatibility for the public API. Verified that the test suite (including existing bugs) remains unchanged and that other files are byte-identical. Documentation provided in `/decisions/`, `/research/`, and `CHANGELOG.md`." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p1_refactor_gemma4-31b-q4_v2", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 211.5, + "iters": 41, + "tokens": { + "completion_total": 11126, + "prompt_total": 508191, + "all_total": 519317 + }, + "throughput": { + "completion_tps_avg": 54.1, + "model_call_wall_s": 205.7, + "tool_call_wall_s": 3.5, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 101.4, + "first_commit_s": 98.0, + "first_tag_s": 196.3 + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 253.02, + "memory_used_mib_at_start": 65453.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0294, + "cost_usd_upper_bound": 0.0038, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9693, + "mean_power_w": 459.06, + "mean_sm_util_pct": 89.05, + "max_temp_c": 84.0 + } + }, + { + "run_name": "p1_refactor_gemma4-31b-q4_v3", + "task": "p1_refactor", + "replicate": 3, + "summary": { + "started_at": "2026-08-02T02:40:35.025584+00:00", + "ended_at": "2026-08-02T02:43:53.245183+00:00", + "elapsed_s": 198.2, + "iterations": 40, + "total_completion_tokens": 10797, + "total_prompt_tokens": 399232, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Refactored `logalyzer/output.py` into a modular subpackage `logalyzer/output/` containing `json_renderer.py`, `csv_renderer.py`, `plain_renderer.py`, and `legacy.py`. Preserved the public API surface and the exact behavior of the original codebase, including a pre-existing `ImportError` in tests caused by Python 3.11 compatibility issues in `logalyzer/filters.py` (which was left untouched per constraints). All verification steps passed except for the pre-existing test failure. Documentation (ADRs, research notes, and CHANGELOG) has been provided. Final release tagged as v0.2.1." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p1_refactor_gemma4-31b-q4_v3", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 198.2, + "iters": 40, + "tokens": { + "completion_total": 10797, + "prompt_total": 399232, + "all_total": 410029 + }, + "throughput": { + "completion_tps_avg": 56.0, + "model_call_wall_s": 192.8, + "tool_call_wall_s": 3.1, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 101.4, + "first_commit_s": 89.7, + "first_tag_s": 186.1 + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 17.56, + "memory_used_mib_at_start": 65197.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0275, + "cost_usd_upper_bound": 0.0036, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9586, + "mean_power_w": 471.53, + "mean_sm_util_pct": 87.9, + "max_temp_c": 71.0 + } + }, + { + "run_name": "p2_extract_gemma4-31b-q4_v1", + "task": "p2_extract", + "replicate": 1, + "summary": { + "started_at": "2026-08-02T01:04:31.829002+00:00", + "ended_at": "2026-08-02T01:05:48.226058+00:00", + "elapsed_s": 76.4, + "iterations": 6, + "total_completion_tokens": 4449, + "total_prompt_tokens": 22299, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Extracted 20 financial and operational fields from the Veridyne Networks Q3 2026 press release. All fields matched the provided JSON schema in terms of type and units. Results are stored in /workspace/extraction_results.json." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p2_extract_gemma4-31b-q4_v1", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 76.4, + "iters": 6, + "tokens": { + "completion_total": 4449, + "prompt_total": 22299, + "all_total": 26748 + }, + "throughput": { + "completion_tps_avg": 58.8, + "model_call_wall_s": 75.7, + "tool_call_wall_s": 0.3, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 37.1, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 214.87, + "memory_used_mib_at_start": 65453.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0106, + "cost_usd_upper_bound": 0.0014, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9162, + "mean_power_w": 460.29, + "mean_sm_util_pct": 91.33, + "max_temp_c": 83.0 + } + }, + { + "run_name": "p2_extract_gemma4-31b-q4_v2", + "task": "p2_extract", + "replicate": 2, + "summary": { + "started_at": "2026-08-02T02:08:59.939496+00:00", + "ended_at": "2026-08-02T02:10:12.503118+00:00", + "elapsed_s": 72.6, + "iterations": 6, + "total_completion_tokens": 4449, + "total_prompt_tokens": 22299, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Extracted 20 financial and operational fields from the Veridyne Networks Q3 2026 press release. All fields matched the provided JSON schema in terms of type and units. Results are stored in /workspace/extraction_results.json." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p2_extract_gemma4-31b-q4_v2", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 72.6, + "iters": 6, + "tokens": { + "completion_total": 4449, + "prompt_total": 22299, + "all_total": 26748 + }, + "throughput": { + "completion_tps_avg": 61.9, + "model_call_wall_s": 71.9, + "tool_call_wall_s": 0.3, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 36.7, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 443.93, + "memory_used_mib_at_start": 66221.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0101, + "cost_usd_upper_bound": 0.0013, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.8953, + "mean_power_w": 493.67, + "mean_sm_util_pct": 90.86, + "max_temp_c": 75.0 + } + }, + { + "run_name": "p2_extract_gemma4-31b-q4_v3", + "task": "p2_extract", + "replicate": 3, + "summary": { + "started_at": "2026-08-02T01:05:49.001024+00:00", + "ended_at": "2026-08-02T01:07:02.880774+00:00", + "elapsed_s": 73.9, + "iterations": 6, + "total_completion_tokens": 4449, + "total_prompt_tokens": 22299, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Extracted 20 financial and operational fields from the Veridyne Networks Q3 2026 press release. All fields matched the provided JSON schema in terms of type and units. Results are stored in /workspace/extraction_results.json." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p2_extract_gemma4-31b-q4_v3", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 73.9, + "iters": 6, + "tokens": { + "completion_total": 4449, + "prompt_total": 22299, + "all_total": 26748 + }, + "throughput": { + "completion_tps_avg": 60.8, + "model_call_wall_s": 73.2, + "tool_call_wall_s": 0.3, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 37.1, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 260.34, + "memory_used_mib_at_start": 65453.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0103, + "cost_usd_upper_bound": 0.0013, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.8796, + "mean_power_w": 489.97, + "mean_sm_util_pct": 97.5, + "max_temp_c": 83.0 + } + }, + { + "run_name": "p2_ci_gemma4-31b-q4_v1", + "task": "p2_ci", + "replicate": 1, + "summary": { + "started_at": "2026-08-02T02:10:13.425486+00:00", + "ended_at": "2026-08-02T02:11:41.932894+00:00", + "elapsed_s": 88.5, + "iterations": 29, + "total_completion_tokens": 4141, + "total_prompt_tokens": 234185, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Fixed CI failures in discountkit. \n1. Resolved ruff linting errors by removing an unused import and updating type annotations to Python 3.11 standards.\n2. Fixed a bug in `discount_amount` where the percent-off value was not divided by 100.\n3. Updated `test_cart.py` to align with the v0.3.0 breaking change regarding `Cart.line_count()`.\n4. Documented all changes in `CHANGELOG.md` and `diagnosis.md`, and added an ADR for the test update.\n5. Verified that `ruff check src/ tests/` and `pytest -q` both pass.\n6. Tagged the release as v0.3.2." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p2_ci_gemma4-31b-q4_v1", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 88.5, + "iters": 29, + "tokens": { + "completion_total": 4141, + "prompt_total": 234185, + "all_total": 238326 + }, + "throughput": { + "completion_tps_avg": 50.5, + "model_call_wall_s": 82.0, + "tool_call_wall_s": 4.7, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 34.1, + "first_commit_s": 0.1, + "first_tag_s": 77.9 + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 281.18, + "memory_used_mib_at_start": 66221.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0123, + "cost_usd_upper_bound": 0.0016, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.904, + "mean_power_w": 440.07, + "mean_sm_util_pct": 69.35, + "max_temp_c": 76.0 + } + }, + { + "run_name": "p2_ci_gemma4-31b-q4_v2", + "task": "p2_ci", + "replicate": 2, + "summary": { + "started_at": "2026-08-02T01:07:03.642776+00:00", + "ended_at": "2026-08-02T01:08:38.188117+00:00", + "elapsed_s": 94.5, + "iterations": 29, + "total_completion_tokens": 4551, + "total_prompt_tokens": 230536, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Fixed CI failures in discountkit v0.3.1. \n\n1. Fixed linting errors using `ruff check --fix` (unused imports, sort imports, deprecated type hints).\n2. Fixed `discount_amount` and `apply_discount` to correctly handle percentages (0-100) instead of fractions (0-1).\n3. Updated `test_line_count` to expect the number of lines instead of total quantity, as specified in the v0.3.0 changelog.\n4. Documented all changes in `CHANGELOG.md`, `diagnosis.md`, and an ADR in `decisions/`.\n5. Tagged the final state as `v0.3.2`.\n\nVerification:\n- `ruff check src/ tests/` -> Passed\n- `pytest -q` -> Passed (7 tests, 0 failed)\n" + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p2_ci_gemma4-31b-q4_v2", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 94.5, + "iters": 29, + "tokens": { + "completion_total": 4551, + "prompt_total": 230536, + "all_total": 235087 + }, + "throughput": { + "completion_tps_avg": 51.7, + "model_call_wall_s": 88.0, + "tool_call_wall_s": 4.9, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 36.3, + "first_commit_s": 0.1, + "first_tag_s": 84.1 + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 222.35, + "memory_used_mib_at_start": 65453.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0131, + "cost_usd_upper_bound": 0.0017, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.8995, + "mean_power_w": 427.5, + "mean_sm_util_pct": 81.5, + "max_temp_c": 84.0 + } + }, + { + "run_name": "p2_ci_gemma4-31b-q4_v3", + "task": "p2_ci", + "replicate": 3, + "summary": { + "started_at": "2026-08-02T02:11:42.872352+00:00", + "ended_at": "2026-08-02T02:13:03.154857+00:00", + "elapsed_s": 80.3, + "iterations": 21, + "total_completion_tokens": 3754, + "total_prompt_tokens": 162381, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Fixed CI failures in discountkit. \n- Corrected `discount_amount` to properly divide the percentage by 100.\n- Updated `test_line_count_matches_qty_sum` in `tests/test_cart.py` to expect the number of lines instead of total quantity, in accordance with the v0.3.0 breaking change documented in the CHANGELOG.\n- Fixed all `ruff` linting errors, including an unused import and outdated type annotations.\n- Verified that `ruff check src/ tests/` and `pytest -q` both pass.\n- Documented the fixes in `diagnosis.md` and updated `CHANGELOG.md` for the v0.3.2 release.\n\"\"\"\n" + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p2_ci_gemma4-31b-q4_v3", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 80.3, + "iters": 21, + "tokens": { + "completion_total": 3754, + "prompt_total": 162381, + "all_total": 166135 + }, + "throughput": { + "completion_tps_avg": 50.2, + "model_call_wall_s": 74.8, + "tool_call_wall_s": 4.2, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 31.6, + "first_commit_s": 0.1, + "first_tag_s": 66.1 + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 256.59, + "memory_used_mib_at_start": 66221.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0112, + "cost_usd_upper_bound": 0.0015, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.934, + "mean_power_w": 435.34, + "mean_sm_util_pct": 79.44, + "max_temp_c": 72.0 + } + }, + { + "run_name": "p2_hallucination_gemma4-31b-q4_v1", + "task": "p2_hallucination", + "replicate": 1, + "summary": { + "started_at": "2026-08-02T01:08:38.973119+00:00", + "ended_at": "2026-08-02T01:10:34.526607+00:00", + "elapsed_s": 115.6, + "iterations": 16, + "total_completion_tokens": 6454, + "total_prompt_tokens": 121925, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Verified 15 issues from the code-review report against the logalyzer codebase. Found 6 real issues and 9 fabricated issues. The final results are recorded in /workspace/triage.json." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p2_hallucination_gemma4-31b-q4_v1", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 115.6, + "iters": 16, + "tokens": { + "completion_total": 6454, + "prompt_total": 121925, + "all_total": 128379 + }, + "throughput": { + "completion_tps_avg": 56.8, + "model_call_wall_s": 113.7, + "tool_call_wall_s": 0.9, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 105.4, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 220.81, + "memory_used_mib_at_start": 65453.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0161, + "cost_usd_upper_bound": 0.0021, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9516, + "mean_power_w": 488.6, + "mean_sm_util_pct": 91.83, + "max_temp_c": 83.0 + } + }, + { + "run_name": "p2_hallucination_gemma4-31b-q4_v2", + "task": "p2_hallucination", + "replicate": 2, + "summary": { + "started_at": "2026-08-02T02:13:04.107174+00:00", + "ended_at": "2026-08-02T02:14:58.207247+00:00", + "elapsed_s": 114.1, + "iterations": 17, + "total_completion_tokens": 6154, + "total_prompt_tokens": 128689, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Verified 15 issues from issue_report.md. 6 were found to be real and 9 were fabricated. Results are written to /workspace/triage.json." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p2_hallucination_gemma4-31b-q4_v2", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 114.1, + "iters": 17, + "tokens": { + "completion_total": 6154, + "prompt_total": 128689, + "all_total": 134843 + }, + "throughput": { + "completion_tps_avg": 54.9, + "model_call_wall_s": 112.1, + "tool_call_wall_s": 1.0, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 98.6, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 241.92, + "memory_used_mib_at_start": 66221.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0158, + "cost_usd_upper_bound": 0.0021, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9202, + "mean_power_w": 481.27, + "mean_sm_util_pct": 93.09, + "max_temp_c": 74.0 + } + }, + { + "run_name": "p2_hallucination_gemma4-31b-q4_v3", + "task": "p2_hallucination", + "replicate": 3, + "summary": { + "started_at": "2026-08-02T01:10:35.323382+00:00", + "ended_at": "2026-08-02T01:11:07.770827+00:00", + "elapsed_s": 32.4, + "iterations": 9, + "total_completion_tokens": 1494, + "total_prompt_tokens": 43015, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "model_stopped", + "done_summary": "(model stopped without explicit done())" + }, + "verdict": "MISSING_OUTPUT", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p2_hallucination_gemma4-31b-q4_v3", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 32.4, + "iters": 9, + "tokens": { + "completion_total": 1494, + "prompt_total": 43015, + "all_total": 44509 + }, + "throughput": { + "completion_tps_avg": 47.5, + "model_call_wall_s": 31.4, + "tool_call_wall_s": 0.5, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": null, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 225.87, + "memory_used_mib_at_start": 65453.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0045, + "cost_usd_upper_bound": 0.0006, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.7716, + "mean_power_w": 417.97, + "mean_sm_util_pct": 83.33, + "max_temp_c": 83.0 + } + }, + { + "run_name": "p2_triage_gemma4-31b-q4_v1", + "task": "p2_triage", + "replicate": 1, + "summary": { + "started_at": "2026-08-02T02:14:59.166581+00:00", + "ended_at": "2026-08-02T02:16:47.170718+00:00", + "elapsed_s": 108.0, + "iterations": 8, + "total_completion_tokens": 6211, + "total_prompt_tokens": 57359, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Triaged 30 support tickets, classifying them into 12 categories with associated urgency levels. Identified two duplicate clusters and recorded non-obvious decisions in ADRs. The final output is a structured JSON file at /workspace/triage_results.json, accompanied by a README and decision logs." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p2_triage_gemma4-31b-q4_v1", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 108.0, + "iters": 8, + "tokens": { + "completion_total": 6211, + "prompt_total": 57359, + "all_total": 63570 + }, + "throughput": { + "completion_tps_avg": 58.1, + "model_call_wall_s": 107.0, + "tool_call_wall_s": 0.5, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 78.1, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 254.58, + "memory_used_mib_at_start": 66221.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.015, + "cost_usd_upper_bound": 0.0019, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9259, + "mean_power_w": 470.89, + "mean_sm_util_pct": 85.81, + "max_temp_c": 75.0 + } + }, + { + "run_name": "p2_triage_gemma4-31b-q4_v2", + "task": "p2_triage", + "replicate": 2, + "summary": { + "started_at": "2026-08-02T01:11:08.557629+00:00", + "ended_at": "2026-08-02T01:13:14.023931+00:00", + "elapsed_s": 125.5, + "iterations": 7, + "total_completion_tokens": 7489, + "total_prompt_tokens": 52650, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Triaged 30 support tickets from /input/repo/tickets.txt. Categorized each ticket, assigned urgency, and identified duplicate clusters. The results are stored in /workspace/triage_results.json, with a detailed approach in /workspace/README.md and a decision record for API limit classification in /workspace/decisions/ADR-001-API-Limit-Classification.md." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p2_triage_gemma4-31b-q4_v2", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 125.5, + "iters": 7, + "tokens": { + "completion_total": 7489, + "prompt_total": 52650, + "all_total": 60139 + }, + "throughput": { + "completion_tps_avg": 60.1, + "model_call_wall_s": 124.6, + "tool_call_wall_s": 0.4, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 99.2, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 259.2, + "memory_used_mib_at_start": 65453.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0174, + "cost_usd_upper_bound": 0.0023, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9562, + "mean_power_w": 499.84, + "mean_sm_util_pct": 97.36, + "max_temp_c": 84.0 + } + }, + { + "run_name": "p2_triage_gemma4-31b-q4_v3", + "task": "p2_triage", + "replicate": 3, + "summary": { + "started_at": "2026-08-02T02:16:48.127387+00:00", + "ended_at": "2026-08-02T02:18:56.402816+00:00", + "elapsed_s": 128.3, + "iterations": 7, + "total_completion_tokens": 7493, + "total_prompt_tokens": 52729, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Triaged 30 support tickets. Categorized each ticket by type and urgency, identified two duplicate clusters (dashboard permissions and API rate limits), and produced a structured JSON output at /workspace/triage_results.json. Additionally provided a README and an ADR explaining the decision to treat rate-limit blocks as active incidents." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p2_triage_gemma4-31b-q4_v3", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 128.3, + "iters": 7, + "tokens": { + "completion_total": 7493, + "prompt_total": 52729, + "all_total": 60222 + }, + "throughput": { + "completion_tps_avg": 58.8, + "model_call_wall_s": 127.4, + "tool_call_wall_s": 0.4, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 99.1, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 269.79, + "memory_used_mib_at_start": 66221.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0178, + "cost_usd_upper_bound": 0.0023, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9353, + "mean_power_w": 489.85, + "mean_sm_util_pct": 91.68, + "max_temp_c": 75.0 + } + }, + { + "run_name": "p3_doc_gemma4-31b-q4_v1", + "task": "p3_doc", + "replicate": 1, + "summary": { + "started_at": "2026-08-02T01:13:14.831432+00:00", + "ended_at": "2026-08-02T01:15:16.584695+00:00", + "elapsed_s": 121.8, + "iterations": 8, + "total_completion_tokens": 7296, + "total_prompt_tokens": 51659, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Produced a 1-page executive brief recommending a 'Pass' on the follow-on investment for Nimbus Logistics. The brief synthesizes contradictions between the company's public press release and internal memos/leaked reports, highlighting valuation repricing, high customer concentration, and a shift from growth to survival. Supporting documentation including key facts, research notes, and decision records are provided." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_doc_gemma4-31b-q4_v1", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 121.8, + "iters": 8, + "tokens": { + "completion_total": 7296, + "prompt_total": 51659, + "all_total": 58955 + }, + "throughput": { + "completion_tps_avg": 60.4, + "model_call_wall_s": 120.7, + "tool_call_wall_s": 0.5, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 62.9, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 202.22, + "memory_used_mib_at_start": 65453.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0169, + "cost_usd_upper_bound": 0.0022, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9442, + "mean_power_w": 488.83, + "mean_sm_util_pct": 96.71, + "max_temp_c": 84.0 + } + }, + { + "run_name": "p3_doc_gemma4-31b-q4_v2", + "task": "p3_doc", + "replicate": 2, + "summary": { + "started_at": "2026-08-02T02:18:57.378753+00:00", + "ended_at": "2026-08-02T02:21:01.214442+00:00", + "elapsed_s": 123.8, + "iterations": 8, + "total_completion_tokens": 7296, + "total_prompt_tokens": 51659, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Produced a 1-page executive brief recommending a 'Pass' on the follow-on investment for Nimbus Logistics. The brief synthesizes contradictions between the company's public press release and internal memos/leaked reports, highlighting valuation repricing, high customer concentration, and a shift from growth to survival. Supporting documentation including key facts, research notes, and decision records are provided." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_doc_gemma4-31b-q4_v2", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 123.8, + "iters": 8, + "tokens": { + "completion_total": 7296, + "prompt_total": 51659, + "all_total": 58955 + }, + "throughput": { + "completion_tps_avg": 59.4, + "model_call_wall_s": 122.8, + "tool_call_wall_s": 0.5, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 62.4, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 230.18, + "memory_used_mib_at_start": 66221.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0172, + "cost_usd_upper_bound": 0.0022, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9693, + "mean_power_w": 480.4, + "mean_sm_util_pct": 92.72, + "max_temp_c": 78.0 + } + }, + { + "run_name": "p3_doc_gemma4-31b-q4_v3", + "task": "p3_doc", + "replicate": 3, + "summary": { + "started_at": "2026-08-02T01:15:17.434478+00:00", + "ended_at": "2026-08-02T01:17:10.700188+00:00", + "elapsed_s": 113.3, + "iterations": 10, + "total_completion_tokens": 6632, + "total_prompt_tokens": 67272, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Produced a 1-page executive brief recommending a 'Pass' on follow-on investment in Nimbus Logistics, Inc. The brief synthesizes data from five sources, highlighting critical tensions between the company's public growth narrative and internal austerity measures, including ARR inflation, customer concentration risk, and planned headcount reductions. Supporting documentation includes a key-facts list, research notes, and a decision record." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_doc_gemma4-31b-q4_v3", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 113.3, + "iters": 10, + "tokens": { + "completion_total": 6632, + "prompt_total": 67272, + "all_total": 73904 + }, + "throughput": { + "completion_tps_avg": 59.2, + "model_call_wall_s": 112.1, + "tool_call_wall_s": 0.6, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 50.8, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 215.59, + "memory_used_mib_at_start": 65453.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0157, + "cost_usd_upper_bound": 0.002, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9709, + "mean_power_w": 469.09, + "mean_sm_util_pct": 89.09, + "max_temp_c": 83.0 + } + }, + { + "run_name": "p3_business_gemma4-31b-q4_v1", + "task": "p3_business", + "replicate": 1, + "summary": { + "started_at": "2026-08-02T02:21:02.195182+00:00", + "ended_at": "2026-08-02T02:22:36.370720+00:00", + "elapsed_s": 94.2, + "iterations": 9, + "total_completion_tokens": 5487, + "total_prompt_tokens": 43691, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "I have completed the independent review of the Borealis Analytics acquisition diligence pack. I recommended a HOLD on the $84M LOI due to a fundamental contradiction between the stated strategic rationale (acquiring a differentiated product) and the integration plan (sunsetting the product in 12 months). Additional concerns include unexpectedly low gross margins (71%), high execution risk regarding the forced relocation of the engineering team, and thin customer diligence (only 5 of 280 customers contacted). I have provided a detailed memo, a full list of concerns, specific diligence asks for the next two weeks, and ADRs for my judgment calls." + }, + "verdict": "FAIL", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_business_gemma4-31b-q4_v1", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 94.2, + "iters": 9, + "tokens": { + "completion_total": 5487, + "prompt_total": 43691, + "all_total": 49178 + }, + "throughput": { + "completion_tps_avg": 59.0, + "model_call_wall_s": 93.1, + "tool_call_wall_s": 0.6, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 42.4, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 235.08, + "memory_used_mib_at_start": 66221.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0131, + "cost_usd_upper_bound": 0.0017, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9554, + "mean_power_w": 458.24, + "mean_sm_util_pct": 78.84, + "max_temp_c": 77.0 + } + }, + { + "run_name": "p3_business_gemma4-31b-q4_v2", + "task": "p3_business", + "replicate": 2, + "summary": { + "started_at": "2026-08-02T01:17:11.532735+00:00", + "ended_at": "2026-08-02T01:18:32.429360+00:00", + "elapsed_s": 80.9, + "iterations": 9, + "total_completion_tokens": 4742, + "total_prompt_tokens": 37120, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "I have completed the independent review of the Borealis Analytics acquisition diligence pack. My primary recommendation is to HOLD the LOI due to a critical discrepancy in the target's reported runway (implied 5.5 months vs claimed 22 months), a strategic contradiction in paying a growth multiple for a product slated for sunset, and overly optimistic staff retention assumptions. I have provided a detailed memo, a comprehensive list of concerns, a specific set of diligence asks for the next two weeks, and an ADR documenting the decision logic." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_business_gemma4-31b-q4_v2", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 80.9, + "iters": 9, + "tokens": { + "completion_total": 4742, + "prompt_total": 37120, + "all_total": 41862 + }, + "throughput": { + "completion_tps_avg": 59.4, + "model_call_wall_s": 79.8, + "tool_call_wall_s": 0.6, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 32.4, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 194.34, + "memory_used_mib_at_start": 65453.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0112, + "cost_usd_upper_bound": 0.0015, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9271, + "mean_power_w": 495.91, + "mean_sm_util_pct": 95.5, + "max_temp_c": 83.0 + } + }, + { + "run_name": "p3_business_gemma4-31b-q4_v3", + "task": "p3_business", + "replicate": 3, + "summary": { + "started_at": "2026-08-02T02:22:37.368829+00:00", + "ended_at": "2026-08-02T02:23:57.088651+00:00", + "elapsed_s": 79.7, + "iterations": 9, + "total_completion_tokens": 4742, + "total_prompt_tokens": 37120, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "I have completed the independent review of the Borealis Analytics acquisition diligence pack. My primary recommendation is to HOLD the LOI due to a critical discrepancy in the target's reported runway (implied 5.5 months vs claimed 22 months), a strategic contradiction in paying a growth multiple for a product slated for sunset, and overly optimistic staff retention assumptions. I have provided a detailed memo, a comprehensive list of concerns, a specific set of diligence asks for the next two weeks, and an ADR documenting the decision logic." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_business_gemma4-31b-q4_v3", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 79.7, + "iters": 9, + "tokens": { + "completion_total": 4742, + "prompt_total": 37120, + "all_total": 41862 + }, + "throughput": { + "completion_tps_avg": 60.3, + "model_call_wall_s": 78.6, + "tool_call_wall_s": 0.6, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 32.0, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 248.35, + "memory_used_mib_at_start": 66221.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0111, + "cost_usd_upper_bound": 0.0014, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.941, + "mean_power_w": 486.72, + "mean_sm_util_pct": 93.75, + "max_temp_c": 74.0 + } + }, + { + "run_name": "p3_market_gemma4-31b-q4_v1", + "task": "p3_market", + "replicate": 1, + "summary": { + "started_at": "2026-08-02T01:18:33.283109+00:00", + "ended_at": "2026-08-02T01:22:50.667125+00:00", + "elapsed_s": 257.4, + "iterations": 16, + "total_completion_tokens": 8071, + "total_prompt_tokens": 387754, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "I have researched and recommended a password manager for a 50-person engineering team. After evaluating 1Password, Bitwarden, Keeper, Dashlane, and NordPass, I recommended 1Password Business due to its superior developer experience and high adoption likelihood, which are critical for eliminating shadow-IT within 90 days. I provided a full feature/pricing comparison, detailed decision records (ADRs), and a comprehensive list of sources." + }, + "verdict": "STRUCTURAL_PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_market_gemma4-31b-q4_v1", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 257.4, + "iters": 16, + "tokens": { + "completion_total": 8071, + "prompt_total": 387754, + "all_total": 395825 + }, + "throughput": { + "completion_tps_avg": 32.0, + "model_call_wall_s": 251.9, + "tool_call_wall_s": 4.5, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 9.4, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 205.07, + "memory_used_mib_at_start": 65453.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0357, + "cost_usd_upper_bound": 0.0046, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9713, + "mean_power_w": 495.65, + "mean_sm_util_pct": 96.29, + "max_temp_c": 85.0 + } + }, + { + "run_name": "p3_market_gemma4-31b-q4_v2", + "task": "p3_market", + "replicate": 2, + "summary": { + "started_at": "2026-08-02T02:23:58.119775+00:00", + "ended_at": "2026-08-02T02:27:12.186064+00:00", + "elapsed_s": 194.1, + "iterations": 20, + "total_completion_tokens": 9028, + "total_prompt_tokens": 442437, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Researched and recommended a password manager for a 50-person engineering team. Evaluated 1Password, Bitwarden, Dashlane, Keeper, and NordPass. Recommended 1Password Business due to superior developer experience (CLI) and strong SOC2 compliance, with Bitwarden as a runner-up for open-source requirements. All findings are cited in sources.md and summarized in recommendation.md." + }, + "verdict": "STRUCTURAL_PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_market_gemma4-31b-q4_v2", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 194.1, + "iters": 20, + "tokens": { + "completion_total": 9028, + "prompt_total": 442437, + "all_total": 451465 + }, + "throughput": { + "completion_tps_avg": 48.0, + "model_call_wall_s": 188.1, + "tool_call_wall_s": 4.8, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 9.3, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 274.21, + "memory_used_mib_at_start": 66221.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.027, + "cost_usd_upper_bound": 0.0035, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9531, + "mean_power_w": 485.18, + "mean_sm_util_pct": 93.97, + "max_temp_c": 75.0 + } + }, + { + "run_name": "p3_market_gemma4-31b-q4_v3", + "task": "p3_market", + "replicate": 3, + "summary": { + "started_at": "2026-08-02T01:22:51.546562+00:00", + "ended_at": "2026-08-02T01:28:13.595380+00:00", + "elapsed_s": 322.0, + "iterations": 15, + "total_completion_tokens": 8059, + "total_prompt_tokens": 271715, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Successfully researched and recommended a password manager for a 50-person engineering team. Evaluated 5 products (1Password, Bitwarden, Dashlane, Keeper, NordPass) against a set of defined criteria including SSO, SOC2 compliance, and Developer Experience. Recommended 1Password Business for its superior CLI and UX, with Bitwarden Enterprise as a cost-effective, open-source runner-up. All factual claims are cited with public URLs in sources.md and the recommendation memo." + }, + "verdict": "STRUCTURAL_PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_market_gemma4-31b-q4_v3", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 322.0, + "iters": 15, + "tokens": { + "completion_total": 8059, + "prompt_total": 271715, + "all_total": 279774 + }, + "throughput": { + "completion_tps_avg": 25.3, + "model_call_wall_s": 318.3, + "tool_call_wall_s": 2.8, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 20.1, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 499.37, + "memory_used_mib_at_start": 65453.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0447, + "cost_usd_upper_bound": 0.0058, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9783, + "mean_power_w": 489.71, + "mean_sm_util_pct": 92.86, + "max_temp_c": 86.0 + } + }, + { + "run_name": "p3_writing_gemma4-31b-q4_v1", + "task": "p3_writing", + "replicate": 1, + "summary": { + "started_at": "2026-08-02T02:27:13.236311+00:00", + "ended_at": "2026-08-02T02:28:46.227783+00:00", + "elapsed_s": 93.0, + "iterations": 13, + "total_completion_tokens": 5089, + "total_prompt_tokens": 72407, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Produced three audience-tailored rewrites of an incident memo: a CEO brief, a customer email, and a legal summary. Each rewrite adheres to specific constraints regarding tone, content, and word limits defined in `audience_briefs.json`. Documentation includes a README, research notes, and ADR-style decision records for transparency on content omissions and tone choices." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_writing_gemma4-31b-q4_v1", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 93.0, + "iters": 13, + "tokens": { + "completion_total": 5089, + "prompt_total": 72407, + "all_total": 77496 + }, + "throughput": { + "completion_tps_avg": 55.7, + "model_call_wall_s": 91.3, + "tool_call_wall_s": 0.8, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 24.1, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 226.75, + "memory_used_mib_at_start": 66221.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0129, + "cost_usd_upper_bound": 0.0017, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.914, + "mean_power_w": 460.49, + "mean_sm_util_pct": 86.0, + "max_temp_c": 73.0 + } + }, + { + "run_name": "p3_writing_gemma4-31b-q4_v2", + "task": "p3_writing", + "replicate": 2, + "summary": { + "started_at": "2026-08-02T01:28:14.508298+00:00", + "ended_at": "2026-08-02T01:31:26.342552+00:00", + "elapsed_s": 191.8, + "iterations": 11, + "total_completion_tokens": 4336, + "total_prompt_tokens": 58205, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Successfully produced three audience-specific rewrites of the internal outage memo. \n\n1. **CEO Brief**: Concise executive summary focusing on board-level signal (second outage in 90 days) and key remediations, without technical jargon.\n2. **Customer Email**: Accountable and warm communication explaining the outage in plain language and offering SLA credits, while omitting internal blame and board-level context.\n3. **Legal Summary**: Precise, dispassionate factual record for SLA credit determination, focusing on time windows, account impact, and contractual thresholds.\n\nAll outputs were verified against the `audience_briefs.json` constraints for content, tone, and word limits. Supporting documentation (decisions and research notes) has been provided." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_writing_gemma4-31b-q4_v2", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 191.8, + "iters": 11, + "tokens": { + "completion_total": 4336, + "prompt_total": 58205, + "all_total": 62541 + }, + "throughput": { + "completion_tps_avg": 22.8, + "model_call_wall_s": 190.4, + "tool_call_wall_s": 0.7, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 86.8, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 500.05, + "memory_used_mib_at_start": 65581.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0266, + "cost_usd_upper_bound": 0.0035, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9645, + "mean_power_w": 482.8, + "mean_sm_util_pct": 87.37, + "max_temp_c": 87.0 + } + }, + { + "run_name": "p3_writing_gemma4-31b-q4_v3", + "task": "p3_writing", + "replicate": 3, + "summary": { + "started_at": "2026-08-02T02:28:47.253711+00:00", + "ended_at": "2026-08-02T02:30:00.964816+00:00", + "elapsed_s": 73.7, + "iterations": 11, + "total_completion_tokens": 4250, + "total_prompt_tokens": 58290, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Produced three tailored rewrites of the outage memo (CEO brief, Customer email, and Legal summary) adhering to all constraints specified in audience_briefs.json. Included detailed decision records and research notes." + }, + "verdict": "PASS", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_writing_gemma4-31b-q4_v3", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 73.7, + "iters": 11, + "tokens": { + "completion_total": 4250, + "prompt_total": 58290, + "all_total": 62540 + }, + "throughput": { + "completion_tps_avg": 58.7, + "model_call_wall_s": 72.4, + "tool_call_wall_s": 0.7, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 38.9, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 259.41, + "memory_used_mib_at_start": 66221.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0102, + "cost_usd_upper_bound": 0.0013, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9498, + "mean_power_w": 441.34, + "mean_sm_util_pct": 79.0, + "max_temp_c": 77.0 + } + }, + { + "run_name": "p3_pm_gemma4-31b-q4_v1", + "task": "p3_pm", + "replicate": 1, + "summary": { + "started_at": "2026-08-02T01:31:27.254962+00:00", + "ended_at": "2026-08-02T01:34:31.145299+00:00", + "elapsed_s": 183.9, + "iterations": 7, + "total_completion_tokens": 3786, + "total_prompt_tokens": 28755, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Produced a structured project status report for Project Aurora, synthesizing six weeks of meeting notes. The final deliverables include `status_report.md` (the primary report), `research/notes.md` (working synthesis), and a `README.md`. The report highlights the shift of the embedded SDK to a private beta due to a discovered security gap and tracks milestones for mid-May and late-July." + }, + "verdict": "FAIL", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_pm_gemma4-31b-q4_v1", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 183.9, + "iters": 7, + "tokens": { + "completion_total": 3786, + "prompt_total": 28755, + "all_total": 32541 + }, + "throughput": { + "completion_tps_avg": 20.7, + "model_call_wall_s": 183.0, + "tool_call_wall_s": 0.4, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 115.9, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 499.9, + "memory_used_mib_at_start": 65581.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0255, + "cost_usd_upper_bound": 0.0033, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9788, + "mean_power_w": 484.17, + "mean_sm_util_pct": 90.46, + "max_temp_c": 86.0 + } + }, + { + "run_name": "p3_pm_gemma4-31b-q4_v2", + "task": "p3_pm", + "replicate": 2, + "summary": { + "started_at": "2026-08-02T02:30:01.998527+00:00", + "ended_at": "2026-08-02T02:31:14.189844+00:00", + "elapsed_s": 72.2, + "iterations": 7, + "total_completion_tokens": 4145, + "total_prompt_tokens": 29617, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "I have produced a structured project status report for Project Aurora based on the provided six weeks of meeting notes. The report includes a headline, a detailed workstream table, a risk register ordered by severity, a chronological decisions log, and a milestone list. All claims are cited by the meeting week, and the most recent status updates were used for all workstreams. All deliverables are located in the /workspace directory as requested." + }, + "verdict": "FAIL", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_pm_gemma4-31b-q4_v2", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 72.2, + "iters": 7, + "tokens": { + "completion_total": 4145, + "prompt_total": 29617, + "all_total": 33762 + }, + "throughput": { + "completion_tps_avg": 58.1, + "model_call_wall_s": 71.3, + "tool_call_wall_s": 0.4, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 44.1, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 249.69, + "memory_used_mib_at_start": 66221.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.01, + "cost_usd_upper_bound": 0.0013, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9695, + "mean_power_w": 472.88, + "mean_sm_util_pct": 90.27, + "max_temp_c": 76.0 + } + }, + { + "run_name": "p3_pm_gemma4-31b-q4_v3", + "task": "p3_pm", + "replicate": 3, + "summary": { + "started_at": "2026-08-02T01:34:32.050965+00:00", + "ended_at": "2026-08-02T01:37:24.333730+00:00", + "elapsed_s": 172.3, + "iterations": 6, + "total_completion_tokens": 3386, + "total_prompt_tokens": 25278, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "finish_reason": "done_signal", + "done_summary": "Produced a structured project status report for Project Aurora. The final deliverable (`/workspace/status_report.md`) synthesizes six weeks of meeting notes into a high-level overview for the CEO, including a workstream roadmap, risk register, decisions log, and milestone list. All claims are cited by week and grounded in the provided source material. Synthesis notes and a README were also provided." + }, + "verdict": "FAIL", + "terminal_label": null, + "label": null, + "cost": { + "run_name": "p3_pm_gemma4-31b-q4_v3", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "wall_s": 172.3, + "iters": 6, + "tokens": { + "completion_total": 3386, + "prompt_total": 25278, + "all_total": 28664 + }, + "throughput": { + "completion_tps_avg": 19.7, + "model_call_wall_s": 171.5, + "tool_call_wall_s": 0.4, + "_note": "completion_tps is over model-API call time only, not wall (tool execution + sandbox not counted)" + }, + "time_to": { + "first_write_s": 112.3, + "first_commit_s": null, + "first_tag_s": null + }, + "gpu": { + "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "power_limit_w": 500.0, + "power_draw_w_at_start": 500.09, + "memory_used_mib_at_start": 65709.0, + "memory_total_mib": 97887.0, + "_note": "power and memory captured at run START only; real values vary during the run. For truthful power, future runs need a sidecar nvidia-smi sampler." + }, + "energy_estimate": { + "kwh_upper_bound": 0.0239, + "cost_usd_upper_bound": 0.0031, + "rate_usd_per_kwh": 0.13, + "_note": "Upper bound: assumes GPU drew at power.limit for the entire wall time. Real draw is lower (idle between calls, peaks during decode). Treat as ceiling, not point estimate." + } + }, + "telemetry": { + "coverage": 0.9867, + "mean_power_w": 478.32, + "mean_sm_util_pct": 93.23, + "max_temp_c": 86.0 + } + } + ] +} diff --git a/benchmarks/gemma4-31b-q4/gemma4-canonical-n3-scorecard.md b/benchmarks/gemma4-31b-q4/gemma4-canonical-n3-scorecard.md new file mode 100644 index 00000000..e94237c6 --- /dev/null +++ b/benchmarks/gemma4-31b-q4/gemma4-canonical-n3-scorecard.md @@ -0,0 +1,31 @@ +# Gemma 4 31B Q4 raw canonical scorecard (N=3) + +> Raw grader verdicts only. Any reproducible correction is a separate overlay tied to unchanged archive and grader hashes. + +- Evidence-complete runs: 36/36 +- Normal completed workspaces: 36/36 +- Explicit terminal outcomes: 0/36 +- Graded runs: 36/36 +- Raw pass-equivalent outcomes: 29/36 +- Median model-call completion throughput: 54.5 tok/s +- Median cell wall time: 122.8 s +- Telemetry-complete runs: 34/36 + +| Task | Raw pass | Scored | Finish reasons | Quality outcomes | +|---|---:|---:|---|---| +| `p1_bugfix` | 2/3 | 3/3 | done_signal:3 | FAIL:1, PASS:2 | +| `p1_testwrite` | 2/3 | 3/3 | done_signal:3 | FAIL:1, PASS:2 | +| `p1_refactor` | 3/3 | 3/3 | done_signal:3 | PASS:3 | +| `p2_extract` | 3/3 | 3/3 | done_signal:3 | PASS:3 | +| `p2_ci` | 3/3 | 3/3 | done_signal:3 | PASS:3 | +| `p2_hallucination` | 2/3 | 3/3 | done_signal:2, model_stopped:1 | MISSING_OUTPUT:1, PASS:2 | +| `p2_triage` | 3/3 | 3/3 | done_signal:3 | PASS:3 | +| `p3_doc` | 3/3 | 3/3 | done_signal:3 | PASS:3 | +| `p3_business` | 2/3 | 3/3 | done_signal:3 | FAIL:1, PASS:2 | +| `p3_market` | 3/3 | 3/3 | done_signal:3 | STRUCTURAL_PASS:3 | +| `p3_writing` | 3/3 | 3/3 | done_signal:3 | PASS:3 | +| `p3_pm` | 0/3 | 3/3 | done_signal:3 | FAIL:3 | + +## Methodology boundary + +A `done_signal` is a finish behavior, not a pass. `PASS` and `STRUCTURAL_PASS` count only as raw pass-equivalent grader verdicts. A preserved terminal label is reported as a distinct non-pass quality outcome, never fabricated into a normal grader verdict. Model-call throughput excludes tool execution; wall time includes it. Telemetry is per attributed replica GPU, while CPU package power is shared host context and AC wall power is unavailable to software. diff --git a/benchmarks/gemma4-31b-q4/gemma4-extended-evidence-audit.json b/benchmarks/gemma4-31b-q4/gemma4-extended-evidence-audit.json new file mode 100644 index 00000000..07d02f51 --- /dev/null +++ b/benchmarks/gemma4-31b-q4/gemma4-extended-evidence-audit.json @@ -0,0 +1,705 @@ +{ + "schema_version": 1, + "generated_at": "2026-08-02T06:44:07.041628+00:00", + "root": "/home/michael/bench-gemma4-31b-q4", + "matrix": { + "path": "/home/michael/bench-gemma4-31b-q4/tooling/gemma4-31b-q4-extended-matrix.json", + "sha256": "be3022cdf50872d8f5e9920a44eabfc9d4a51617302d49a4ea32d597e67c1359" + }, + "expected_runs": 12, + "audited_runs": 12, + "passed": false, + "errors": [ + "n1_gemma4-31b-q4_v1: artifact does not identify pinned subject ref 1678f19404c54f8588eceabfb21c3f4ce812b483", + "n1_gemma4-31b-q4_v1: artifact does not identify pinned subject ref 488f5b48f02946ff31ce3f342566d2a1a9687201", + "n1_gemma4-31b-q4_v1: artifact does not identify pinned subject ref ff20aadf067905b89c9c781cb5ee7ea0532551d5", + "n1_gemma4-31b-q4_v1: artifact does not identify pinned subject ref ab148dee598635644b759f87abb3f594918276b6", + "n1_gemma4-31b-q4_v2: artifact does not identify pinned subject ref 309e9cd0ad5d572ab9313eb577ec30a3e8524752", + "n1_gemma4-31b-q4_v2: artifact does not identify pinned subject ref e5ceb43ea0f9c3939a154fb42e6649b828eca878", + "n1_gemma4-31b-q4_v2: artifact does not identify pinned subject ref 1678f19404c54f8588eceabfb21c3f4ce812b483", + "n1_gemma4-31b-q4_v2: artifact does not identify pinned subject ref e4e83fedfee36c651afc6af7f3dba9851fa5dc96", + "n1_gemma4-31b-q4_v2: artifact does not identify pinned subject ref 488f5b48f02946ff31ce3f342566d2a1a9687201", + "n1_gemma4-31b-q4_v2: artifact does not identify pinned subject ref ff20aadf067905b89c9c781cb5ee7ea0532551d5", + "n1_gemma4-31b-q4_v2: artifact does not identify pinned subject ref ab148dee598635644b759f87abb3f594918276b6", + "n1_gemma4-31b-q4_v3: artifact does not identify pinned subject ref 1678f19404c54f8588eceabfb21c3f4ce812b483" + ], + "warnings": [], + "runs": [ + { + "run_name": "n1_gemma4-31b-q4_v1", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "2edfd7773c5243b460d7402468dbcb5c86efeb21", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 288.6, + "iterations": 32, + "completion_tokens": 11141, + "prompt_tokens_cumulative": 1095202, + "model_turns": 32, + "length_finishes": 0, + "grade_verdict": null, + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9875, + "mean_power_w": 458.59, + "mean_sm_util_pct": 85.19, + "max_temp_c": 73.0, + "cpu_package_mean_power_w": 132.56 + }, + "files": { + "receipt.json": { + "bytes": 9472, + "sha256": "4c3308a0e67ff0c66c121c0cd0e481c254d943ea11145e84165d8cf0c14042ec" + }, + "transcript.jsonl": { + "bytes": 33062, + "sha256": "4439ce94aacdf22011e95e438d2721551b5db657c4f85e70a7600b7338259bce" + }, + "summary.json": { + "bytes": 588, + "sha256": "ea1105b39de0245dda40eb4e94e6007bdddc8ec6f25c64a97de143892e0f0fc8" + }, + "workspace_final.tar.gz": { + "bytes": 44563623, + "sha256": "36a69bbc5635483742b3187dc8c26928a5abe5bdf9c640a4ece8f7e57d079af6" + }, + "cost.json": { + "bytes": 1232, + "sha256": "90139506ea41482a59f7d06c3baeacfb1b192016a36b98c61ef5ff54139ee478" + }, + "gpu_telemetry.json": { + "bytes": 2388, + "sha256": "5448acd781b2c8fd477755e821cad21717677aec9bf20a29f4bcb8bf3a0e81c6" + } + }, + "suite": "dreamserver-1-pr-audit", + "replicate": 1, + "ordinal": 0 + }, + { + "run_name": "n1_gemma4-31b-q4_v2", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "2edfd7773c5243b460d7402468dbcb5c86efeb21", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 290.8, + "iterations": 29, + "completion_tokens": 12262, + "prompt_tokens_cumulative": 979694, + "model_turns": 29, + "length_finishes": 0, + "grade_verdict": null, + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9801, + "mean_power_w": 471.42, + "mean_sm_util_pct": 92.14, + "max_temp_c": 80.0, + "cpu_package_mean_power_w": 132.56 + }, + "files": { + "receipt.json": { + "bytes": 9472, + "sha256": "8312185d2c7077f5b8a7290b53b011fd0071fc38cfb83060d99c6b026f12daa2" + }, + "transcript.jsonl": { + "bytes": 37260, + "sha256": "2b6bf9b02a8ef6ae6f0b178cfed1e16a9f65469a7dccb474fd7cbac591ba8338" + }, + "summary.json": { + "bytes": 727, + "sha256": "0a1cdb090b3f1cafd5926fc52b16ffe28e3f33fdc460e5c5e3c5b4afeb45ae83" + }, + "workspace_final.tar.gz": { + "bytes": 35319738, + "sha256": "675a82e99d568a7d67d0f27155e9e946e734c09f77caa675ef7dab9e7bc7e056" + }, + "cost.json": { + "bytes": 1229, + "sha256": "f6e9c7e7b9b0b8a2f69d055546e6b36eb878b142d78947489a8ba5b385f7caa0" + }, + "gpu_telemetry.json": { + "bytes": 2389, + "sha256": "2b744b53e9f1a88e60ed7ec8cb6519866d5caaaeadee2c271deb8e8c597cc91c" + } + }, + "suite": "dreamserver-1-pr-audit", + "replicate": 2, + "ordinal": 1 + }, + { + "run_name": "n1_gemma4-31b-q4_v3", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "2edfd7773c5243b460d7402468dbcb5c86efeb21", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 301.0, + "iterations": 37, + "completion_tokens": 12061, + "prompt_tokens_cumulative": 1199565, + "model_turns": 37, + "length_finishes": 0, + "grade_verdict": null, + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9801, + "mean_power_w": 463.7, + "mean_sm_util_pct": 86.8, + "max_temp_c": 74.0, + "cpu_package_mean_power_w": 134.74 + }, + "files": { + "receipt.json": { + "bytes": 9475, + "sha256": "ae07345f67f98fbf724d7042cf58023d7991bb87bf781a7620e0c65e32db1074" + }, + "transcript.jsonl": { + "bytes": 32696, + "sha256": "6543161fa9f6cc3053e264822d3bfe0622bcdbef0edb4cbbc26bec4cb33e9ec0" + }, + "summary.json": { + "bytes": 897, + "sha256": "a39102ba583247ab793f0dfde96238e811a60d64e0bd576aa7707b1987cce044" + }, + "workspace_final.tar.gz": { + "bytes": 35372453, + "sha256": "ecf822386e837254c8955986a05d042e814cd57488f73a5bc66912d68a263645" + }, + "cost.json": { + "bytes": 1235, + "sha256": "e00196d50b43ed63c16e4776a05b5050109a5bf09de5d77911b83d6e86149835" + }, + "gpu_telemetry.json": { + "bytes": 2395, + "sha256": "0811f37cc399b3f519fce10e1bfc66bfa83cdb76c84ab94db0862073653d8723" + } + }, + "suite": "dreamserver-1-pr-audit", + "replicate": 3, + "ordinal": 2 + }, + { + "run_name": "gemma4-31b-q4_invest_memo_v1", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "2edfd7773c5243b460d7402468dbcb5c86efeb21", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 499.1, + "iterations": 52, + "completion_tokens": 21513, + "prompt_tokens_cumulative": 1417966, + "model_turns": 52, + "length_finishes": 0, + "grade_verdict": null, + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9918, + "mean_power_w": 468.66, + "mean_sm_util_pct": 88.67, + "max_temp_c": 82.0, + "cpu_package_mean_power_w": 134.32 + }, + "files": { + "receipt.json": { + "bytes": 9488, + "sha256": "9825d68667206159264702789b0e99f8da8b889ee3f72299cb945f1f71882f09" + }, + "transcript.jsonl": { + "bytes": 63001, + "sha256": "bf1d7bdb0c9fe073e5a45a516e903b583fdd30de918eb4e19f99bfc662e52b88" + }, + "summary.json": { + "bytes": 1656, + "sha256": "05ad14fd22f58043b3e21e1ac8d5a9d36e98fe2ca31ef47b190d13631b722af9" + }, + "workspace_final.tar.gz": { + "bytes": 14871211, + "sha256": "2081780297fcdc6f6854ff89508602035b1e8a6dd7a8582aaf00addadff8a8db" + }, + "cost.json": { + "bytes": 1242, + "sha256": "0d8e1d5536413bb9b2e0e54f3d732ab6cc642ab1b8be954f46b78cd09bda3201" + }, + "gpu_telemetry.json": { + "bytes": 2481, + "sha256": "5766a571c30bd9a0ffa803bc1ad1d50c57456aaadcb85a813eeb9dfd9cbd1845" + } + }, + "suite": "wallstreet-investment-memo", + "replicate": 1, + "ordinal": 3 + }, + { + "run_name": "gemma4-31b-q4_invest_memo_v2", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "2edfd7773c5243b460d7402468dbcb5c86efeb21", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 527.9, + "iterations": 55, + "completion_tokens": 20235, + "prompt_tokens_cumulative": 2650157, + "model_turns": 55, + "length_finishes": 0, + "grade_verdict": null, + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9945, + "mean_power_w": 464.18, + "mean_sm_util_pct": 84.12, + "max_temp_c": 74.0, + "cpu_package_mean_power_w": 135.68 + }, + "files": { + "receipt.json": { + "bytes": 9487, + "sha256": "ec55215a3d6382df4dd49d106b5874d1ec0a0f499cb8c04cc92c01c0b10d1a3e" + }, + "transcript.jsonl": { + "bytes": 70598, + "sha256": "94302f3a30b52a995d03598229c5ca9c4676e830ca6e3d745f8037eea4c28547" + }, + "summary.json": { + "bytes": 1117, + "sha256": "1ea1bfa8e6b5cb6b84044b60dd1e32e4e703be9707bd49adcea7825fef49adbb" + }, + "workspace_final.tar.gz": { + "bytes": 2081783, + "sha256": "09b573d50aeda91dfc8eb2b8b35f9158b971b24ef0d45d027d6e3659d672dbdc" + }, + "cost.json": { + "bytes": 1242, + "sha256": "23a00ed6e081129360c1a0d2b3281e318bb10564d3faac30ad74ae295c759ac9" + }, + "gpu_telemetry.json": { + "bytes": 2492, + "sha256": "5a1720568453f58d9d437eee15460b1c7fc56202002dcd28994e1ba9fcc24bcf" + } + }, + "suite": "wallstreet-investment-memo", + "replicate": 2, + "ordinal": 4 + }, + { + "run_name": "gemma4-31b-q4_invest_memo_v3", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "2edfd7773c5243b460d7402468dbcb5c86efeb21", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 413.6, + "iterations": 53, + "completion_tokens": 15676, + "prompt_tokens_cumulative": 2347299, + "model_turns": 53, + "length_finishes": 0, + "grade_verdict": null, + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9913, + "mean_power_w": 457.58, + "mean_sm_util_pct": 83.6, + "max_temp_c": 81.0, + "cpu_package_mean_power_w": 135.76 + }, + "files": { + "receipt.json": { + "bytes": 9489, + "sha256": "d067e1d9291a7ed28a9702ac2ee75f37907bddd8b04849a1ee953fa0bbc12ddf" + }, + "transcript.jsonl": { + "bytes": 51513, + "sha256": "95a15eb5c1a08f07a0ca6864eea3301ad1ce8af08d883a2cd3e848e233d03687" + }, + "summary.json": { + "bytes": 1131, + "sha256": "fbd117d21e2bb9d2764cc275978dbb825471c39814f488c1537755c15d10f30b" + }, + "workspace_final.tar.gz": { + "bytes": 9286952, + "sha256": "a54068f11048fdab8529893a935856fb591491ce373545bbd734f9899d72dc40" + }, + "cost.json": { + "bytes": 1243, + "sha256": "4cfdc410bb505792a5efaac954e7f72b954b33e4a7a44014a22ea0cccf9864ff" + }, + "gpu_telemetry.json": { + "bytes": 2488, + "sha256": "d7b32114e33b188eb541af710919e7528bc458027dbde946f81547d702243f49" + } + }, + "suite": "wallstreet-investment-memo", + "replicate": 3, + "ordinal": 5 + }, + { + "run_name": "gemma4-31b-q4_board_pres_v1", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "2edfd7773c5243b460d7402468dbcb5c86efeb21", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 345.5, + "iterations": 55, + "completion_tokens": 16170, + "prompt_tokens_cumulative": 1335224, + "model_turns": 55, + "length_finishes": 0, + "grade_verdict": null, + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9841, + "mean_power_w": 475.35, + "mean_sm_util_pct": 88.68, + "max_temp_c": 75.0, + "cpu_package_mean_power_w": 136.49 + }, + "files": { + "receipt.json": { + "bytes": 9568, + "sha256": "9359047f762d4e921e232027fa399360c39c299b470e3d6067692d58d8bd91aa" + }, + "transcript.jsonl": { + "bytes": 61777, + "sha256": "2ad9abfdb05532a3abee4379115da3be5f4a66bf638c80a611e05b17cbe8a3e5" + }, + "summary.json": { + "bytes": 713, + "sha256": "ca5398982b2d89cda9e2096988b3e2be00e02c57328f2beb3b97f82b2c222dc9" + }, + "workspace_final.tar.gz": { + "bytes": 1013461, + "sha256": "c57c381efeafacf75b4ff11257b8040dfc5975a92c8dd6ae3fb4b898b60684bf" + }, + "cost.json": { + "bytes": 1240, + "sha256": "5968faaf42e4b8527fb61ad2e9ba683fcf962dcb78f30dbe554d804ee760b46c" + }, + "gpu_telemetry.json": { + "bytes": 2487, + "sha256": "2346cc1744a1c13129e199e3bfb4c831dac412668e9c6b552db47017bbe794ab" + } + }, + "suite": "wallstreet-board-presentation", + "replicate": 1, + "ordinal": 6 + }, + { + "run_name": "gemma4-31b-q4_board_pres_v2", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "2edfd7773c5243b460d7402468dbcb5c86efeb21", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 334.8, + "iterations": 39, + "completion_tokens": 17763, + "prompt_tokens_cumulative": 494272, + "model_turns": 39, + "length_finishes": 0, + "grade_verdict": null, + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9857, + "mean_power_w": 466.48, + "mean_sm_util_pct": 88.06, + "max_temp_c": 83.0, + "cpu_package_mean_power_w": 136.07 + }, + "files": { + "receipt.json": { + "bytes": 9569, + "sha256": "52067ac01815cef44f0b76f4b22079317cb6f4e4b4bd54066020dc189e073fb5" + }, + "transcript.jsonl": { + "bytes": 59495, + "sha256": "f88c6829668cb989122d0007f1f1793c6675c9287bbb693f0d3a1ec1333257c7" + }, + "summary.json": { + "bytes": 871, + "sha256": "5d0f896e9cd0acc1e06ccb20ceca8970d6d08b48595f8a7cfe5d6e28786a81e9" + }, + "workspace_final.tar.gz": { + "bytes": 662464, + "sha256": "270c8feebf15eb3e5a7d1413b04f2d1aa95126921bdd739f4986338392d1d4bf" + }, + "cost.json": { + "bytes": 1238, + "sha256": "6589e51cd7862d05ba9a7403bcbd4dd713edf9e834190357be66f54028037bd5" + }, + "gpu_telemetry.json": { + "bytes": 2484, + "sha256": "cef85b72bb935d97a038600f3e757ba82e7116241ad2ff6687e78d489d67c0e5" + } + }, + "suite": "wallstreet-board-presentation", + "replicate": 2, + "ordinal": 7 + }, + { + "run_name": "gemma4-31b-q4_board_pres_v3", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "2edfd7773c5243b460d7402468dbcb5c86efeb21", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 276.5, + "iterations": 39, + "completion_tokens": 13446, + "prompt_tokens_cumulative": 719464, + "model_turns": 39, + "length_finishes": 0, + "grade_verdict": null, + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9765, + "mean_power_w": 469.31, + "mean_sm_util_pct": 82.73, + "max_temp_c": 75.0, + "cpu_package_mean_power_w": 134.12 + }, + "files": { + "receipt.json": { + "bytes": 9568, + "sha256": "e4b677e582bcc80085bba12239dd2e6624b330a79ec1cbb85212685c235ec0bd" + }, + "transcript.jsonl": { + "bytes": 44658, + "sha256": "bea08e24361aaaaf39a32cf751e4e5e10502d1b25327c8bb61ece2207d206473" + }, + "summary.json": { + "bytes": 724, + "sha256": "d826023fac361cb2ea063e736cc846d9bad17275008e20882ac29650841b8d91" + }, + "workspace_final.tar.gz": { + "bytes": 889261, + "sha256": "7e13b973a52100048c866bada8fb4ef60cf6e5740ee6e569293d8f64d48e9613" + }, + "cost.json": { + "bytes": 1239, + "sha256": "752887cd2873372346c251cc4a329ccf4a20b11b401a426af9ed2759f56c3c55" + }, + "gpu_telemetry.json": { + "bytes": 2480, + "sha256": "40b7fdabae80d03ec04cb9f8ff14e8862910b219064edbc1ad3b31da498430d8" + } + }, + "suite": "wallstreet-board-presentation", + "replicate": 3, + "ordinal": 8 + }, + { + "run_name": "gemma4-31b-q4_75pr_v1", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "2edfd7773c5243b460d7402468dbcb5c86efeb21", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 496.2, + "iterations": 41, + "completion_tokens": 17996, + "prompt_tokens_cumulative": 2129829, + "model_turns": 41, + "length_finishes": 0, + "grade_verdict": null, + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9976, + "mean_power_w": 478.13, + "mean_sm_util_pct": 91.13, + "max_temp_c": 82.0, + "cpu_package_mean_power_w": 132.51 + }, + "files": { + "receipt.json": { + "bytes": 9556, + "sha256": "e391c2ae8b53e64597c7274da1c5a2407df958229b77d4618a3dbbd8746898ab" + }, + "transcript.jsonl": { + "bytes": 51607, + "sha256": "fedf67f0ef5f5faac2b2ee92057d9f4c0b67c506dfb2c08807190994af1dc6eb" + }, + "summary.json": { + "bytes": 903, + "sha256": "bcd3d259a03c267301caa15b59f8d1940887d8b39c4d686f2492a280195dd45f" + }, + "workspace_final.tar.gz": { + "bytes": 34921336, + "sha256": "2986ef9cc929b714eb48976c400957248f7f289376fbe5c2886c9d4f8f45141e" + }, + "cost.json": { + "bytes": 1233, + "sha256": "2647027537956b6ddfb84e29cd2b12b80c1e363dd7b8225a569dd62df841bbc4" + }, + "gpu_telemetry.json": { + "bytes": 2477, + "sha256": "512ebf212572dd92fa402e27ac903c3bf98fbfc939b6a6ecee11261d63348816" + } + }, + "suite": "dreamserver-75-pr-audit", + "replicate": 1, + "ordinal": 9 + }, + { + "run_name": "gemma4-31b-q4_75pr_v2", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "2edfd7773c5243b460d7402468dbcb5c86efeb21", + "endpoint": "http://127.0.0.1:8000/v1/chat/completions", + "finish_reason": "model_stopped", + "elapsed_s": 193.8, + "iterations": 32, + "completion_tokens": 7592, + "prompt_tokens_cumulative": 879716, + "model_turns": 32, + "length_finishes": 0, + "grade_verdict": null, + "telemetry": { + "gpu": [ + "0" + ], + "coverage": 0.9804, + "mean_power_w": 459.98, + "mean_sm_util_pct": 83.54, + "max_temp_c": 73.0, + "cpu_package_mean_power_w": 135.56 + }, + "files": { + "receipt.json": { + "bytes": 9556, + "sha256": "3f06c30bdae4813b39c0ccf74770892179550be0590463e672c0609f889d2b7c" + }, + "transcript.jsonl": { + "bytes": 23651, + "sha256": "01c6decba025a16b03929e68486de8aa8ce653f65e83dc4472195b36013a1c56" + }, + "summary.json": { + "bytes": 427, + "sha256": "84671b0076d1eed4faca9a26553b399a002e9bf90e732e37bbf32d477d02a0e4" + }, + "workspace_final.tar.gz": { + "bytes": 54035493, + "sha256": "f9a10fe0fa04125aa230dd9e588b5763544b3fa3ebcba9666e6d78bc9e427fd9" + }, + "cost.json": { + "bytes": 1231, + "sha256": "183017e80971ab13628c21da1aa56d27cd31c0b755f7caa351778f4259bd37e0" + }, + "gpu_telemetry.json": { + "bytes": 2394, + "sha256": "6f7cfc0572a40cea88c6d5d9abdfdf52f1e8449f3b26f89697fbf744dbbe7284" + } + }, + "suite": "dreamserver-75-pr-audit", + "replicate": 2, + "ordinal": 10 + }, + { + "run_name": "gemma4-31b-q4_75pr_v3", + "outcome_kind": "completed-workspace", + "terminal_label": null, + "harness_git_sha": "17e0128cb8c3268d4616a51147b1b754db8103b6", + "endpoint": "http://127.0.0.1:8001/v1/chat/completions", + "finish_reason": "done_signal", + "elapsed_s": 561.7, + "iterations": 41, + "completion_tokens": 25841, + "prompt_tokens_cumulative": 1193936, + "model_turns": 41, + "length_finishes": 0, + "grade_verdict": null, + "telemetry": { + "gpu": [ + "1" + ], + "coverage": 0.9881, + "mean_power_w": 475.31, + "mean_sm_util_pct": 90.95, + "max_temp_c": 71.0, + "cpu_package_mean_power_w": 113.63 + }, + "files": { + "receipt.json": { + "bytes": 9554, + "sha256": "ec88a1c45d985f0fb5e47b7b0a0f781e68087f4a7d01fb7d85bf796e3a752b18" + }, + "transcript.jsonl": { + "bytes": 84879, + "sha256": "3bdf12182ac973b2b7e647152b6a150aebf1f8cfdeb90172de95c3e81a4aae27" + }, + "summary.json": { + "bytes": 954, + "sha256": "ce48b91c08ea6759d22844de7533700fdecd0c4b7ce6a2c857cdba9f80bd95e4" + }, + "workspace_final.tar.gz": { + "bytes": 626792011, + "sha256": "3246d75825dc7c8b9f0abb76c3b4ddf2e1a716a4b18e679e10922a86756275cc" + }, + "cost.json": { + "bytes": 1233, + "sha256": "3286a47bec380fa22049ba1491353221c6760658d9181f8217af1a0e3747af5b" + }, + "gpu_telemetry.json": { + "bytes": 2399, + "sha256": "53cfcf980e4301302ad3a218ae8bef57c1375c5f9f96f7cb2b2875b8affc2411" + } + }, + "suite": "dreamserver-75-pr-audit", + "replicate": 3, + "ordinal": 11 + } + ], + "preserved_infrastructure_invalid_attempts": [ + { + "attempt": "gemma4-31b-q4_75pr_v3-attempt1-20260802T062523Z", + "files": { + "gpu_telemetry.json": { + "bytes": 2424, + "sha256": "09e3321f1d6730b6ea0d767f2b6becec50f7c07d529decdcc93592e50b362de9" + }, + "invalidation.json": { + "bytes": 1485, + "sha256": "bf2f4230faa203ccec907912c1133bf40f1141a0e1c562b8c6a52a2cf33fee74" + }, + "label.json": { + "bytes": 1351, + "sha256": "3aeee812852f6e4f71a8fe15899d7892b1a8e92f4806cc868f65fe1eac771d0d" + }, + "receipt.json": { + "bytes": 9556, + "sha256": "7b5d95e4779d530d1d7f3df84fea87d93906ea6b2cabe65a27972604e286df93" + }, + "transcript.jsonl": { + "bytes": 35844, + "sha256": "eda95eb4d40d46556d222b4ffff0d88ee9d56cead5a21b0c4685bc2890393d5c" + } + } + } + ], + "scope_note": "This audit proves identity, configuration, routing, telemetry, and artifact preservation. Suite-specific substantive grading and visual/workbook/code inspection are separate required overlays; this document is not a quality pass." +} diff --git a/benchmarks/gemma4-31b-q4/post-campaign-restore.json b/benchmarks/gemma4-31b-q4/post-campaign-restore.json new file mode 100644 index 00000000..e2bd5b96 --- /dev/null +++ b/benchmarks/gemma4-31b-q4/post-campaign-restore.json @@ -0,0 +1,45 @@ +{ + "schema_version": 1, + "validated_at": "2026-08-02T07:06:18Z", + "status": "passed", + "restored_stack": { + "model": "DeepSeek-V4-Flash-0731", + "max_model_len": 1048576, + "health_http_status": 200, + "image_digest": "sha256:48518e91cf87dd0c0483c76ff86e81dfc0f46de7e364b46f7a82c481ce08188f", + "launcher_sha256": "4fca6a4a478f0876a1e6d7241dd87a3c9574bd970f29943bb3696993d4bfe25d", + "openclaw_config_sha256": "79f2872821d68f76a2c42cb8f3972ae494b380a4d0d058fba405dbcc5e489c7f", + "gpu_power_limits_w": [500, 500] + }, + "production_checks": { + "gateway_active": true, + "sanctuary_portal_healthy": true, + "pixel_portal_healthy": true, + "deepseek_restart_count": 0, + "deepseek_oom_killed": false, + "sanctuary": { + "status": "ok", + "tool_command": "printf SANCTUARY_OK", + "tool_result": "SANCTUARY_OK", + "provider": "tower", + "model": "DeepSeek-V4-Flash-0731", + "context_tokens": 1048576, + "fallback_used": false + }, + "pixel": { + "status": "ok", + "tool_command": "printf PIXEL_OK", + "tool_result": "PIXEL_OK", + "provider": "tower", + "model": "DeepSeek-V4-Flash-0731", + "context_tokens": 1048576, + "fallback_used": false + } + }, + "cleanup": { + "gemma_serving_services_running": 0, + "gemma_serving_containers_running": 0, + "campaign_sandbox_containers_stopped_not_deleted": 96 + }, + "private_evidence_directory": "/home/michael/gemma4-campaign-state/production-restore/20260802T070054Z" +} diff --git a/benchmarks/gemma4-31b-q4/substantive-audit.json b/benchmarks/gemma4-31b-q4/substantive-audit.json new file mode 100644 index 00000000..4f3142e0 --- /dev/null +++ b/benchmarks/gemma4-31b-q4/substantive-audit.json @@ -0,0 +1,121 @@ +{ + "schema_version": 1, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "audited_at": "2026-08-02", + "evidence_audit": { + "path": "gemma4-extended-evidence-audit.json", + "sha256": "5cf402097ab50a900f021da44fdb3e5ac91862653980c673dc21fd53be1d6242", + "valid_runs": 12, + "preserved_infrastructure_invalid_attempts": 1, + "passed": false, + "failure_reason": "all three single-PR artifacts omit one or more pinned subject commits" + }, + "overall": { + "strict_substantive_passes": 0, + "runs": 12, + "note": "A normal harness completion or historical artifact gate is necessary but not sufficient for a substantive pass." + }, + "single_pr": [ + { + "replicate": 1, + "disposition": "MERGE", + "expected": "MERGE", + "classification": "COMMON_PROVENANCE_GATE_FAIL", + "findings": ["useful review", "76 tests preserved", "required pinned refs/history omitted", "final audit repo untagged"] + }, + { + "replicate": 2, + "disposition": "MERGE", + "expected": "MERGE", + "classification": "COMMON_PROVENANCE_GATE_FAIL", + "findings": ["correct disposition", "claimed 76 tests without preserved output", "required pinned refs/history omitted", "final audit repo untagged and nested audit workspace dirty"] + }, + { + "replicate": 3, + "disposition": "REJECT", + "expected": "MERGE", + "classification": "WRONG_SUBJECT_AND_VERDICT", + "findings": ["treated unrelated current-main extensions as PR contamination", "required squash/history refs omitted", "preserved 76/76 test claim", "clean tagged nested audit repo"] + } + ], + "investment_memo": [ + { + "replicate": 1, + "classification": "SHIPPED_WITH_MATERIAL_FINANCE_DEFECTS", + "historical_gate": "FAIL_MISSING_PDF", + "findings": ["tiny one-sheet workbook", "not a three-statement model", "no scenario/share-price checks", "stated $110 target contradicts workbook and independent bridge"] + }, + { + "replicate": 2, + "classification": "SHIPPED_WITH_MATERIAL_FINANCE_DEFECTS", + "historical_gate": "FAIL_MISSING_PDF", + "findings": ["one static historical sheet", "zero formulas", "no forecast statements or valuation", "unsupported $85 target"] + }, + { + "replicate": 3, + "classification": "SHIPPED_WITH_MATERIAL_FINANCE_DEFECTS", + "historical_gate": "FAIL_MISSING_PDF", + "findings": ["one static sheet", "zero formulas", "no merger combination, valuation, or share bridge", "unsupported $15 target"] + } + ], + "board_presentation": [ + { + "replicate": 1, + "classification": "MATERIAL_DECK_DEFECTS", + "historical_gate": "FAIL_MISSING_PDF", + "findings": ["15-slide PPTX", "severe overlap/clipping on slides 4, 5, 6, 11, and 13", "unresolved >X% placeholder", "unsupported $110 valuation"] + }, + { + "replicate": 2, + "classification": "MATERIAL_DECK_DEFECTS", + "historical_gate": "FAIL_MISSING_PDF", + "findings": ["16-slide PPTX", "sparse and generic", "clipped chart/subtitle", "synthetic unsupported positioning/distribution visuals", "unsupported $85 valuation"] + }, + { + "replicate": 3, + "classification": "SHIPPED_WITH_MATERIAL_DECK_DEFECTS", + "historical_gate": "PASS_PPTX_AND_PDF", + "findings": ["15-slide PPTX and 15-page PDF", "portrait layout with large whitespace and tiny visuals", "axis collision", "generic/synthetic visuals", "unsupported $15 valuation"] + } + ], + "frozen_75pr": [ + { + "replicate": 1, + "audit": "75pr-v1-audit.json", + "audit_sha256": "a8943774041bb7b419620c18fd4558e9481f7d755ae664713a0ce25b520a83d8", + "classification": "MODEL_TERMINAL_FAILURE", + "actual_pr_dirs": 6, + "missing_pr_dirs": 69, + "parsable_verdicts": 6, + "test_evidence": 6 + }, + { + "replicate": 2, + "audit": "75pr-v2-audit.json", + "audit_sha256": "889686929e81a49956a83aad3ed07692acdef47aa1f9f8e954c52631e5268cb8", + "classification": "MODEL_TERMINAL_FAILURE", + "actual_pr_dirs": 2, + "missing_pr_dirs": 73, + "parsable_verdicts": 1, + "test_evidence": 0 + }, + { + "replicate": 3, + "audit": "75pr-v3-audit.json", + "audit_sha256": "71af86f5da85e5be17eaa2b617b91ae0e441f100f9187faf173bcb5dd5ba0823", + "classification": "MODEL_TERMINAL_FAILURE", + "actual_pr_dirs": 75, + "missing_pr_dirs": 0, + "parsable_verdicts": 14, + "test_evidence": 4, + "reviews_under_800_bytes": 75, + "clean_final_repo": false + } + ], + "independent_reproduction": { + "single_pr_base": "67 passed", + "single_pr_head": "76 passed", + "single_pr_current": "261 passed", + "conclusion": "The pinned subject is a test-only rescue over runtime code already present on current main; MERGE is the expected disposition." + } +} diff --git a/claims.yaml b/claims.yaml index 324455ec..7060027a 100644 --- a/claims.yaml +++ b/claims.yaml @@ -19,7 +19,7 @@ # entry once it's been published. schema_version: 0.1 -last_updated: "2026-08-01" +last_updated: "2026-08-02" claims: @@ -509,7 +509,7 @@ claims: think, so the direction favors DeepSeek but is not a statistically matched global leaderboard claim. status: provisional - scope: published corrected local MMBT results as of 2026-08-01 + scope: published corrected local MMBT results as of 2026-08-02 evidence: "benchmarks/deepseek-v4-flash-0731/qwen397-corrected-score-overlay.json; benchmarks/deepseek-v4-flash-0731/DEEPSEEK_V4_FLASH_0731_VERIFIED_RESULTS.md" caveats: - "DeepSeek N=3 vs Qwen N=10; sampling, context, campaign date, and grader coverage differ." @@ -542,6 +542,62 @@ claims: # ─── Retracted ──────────────────────────────────────────────────────────────── + # Gemma 4 31B QAT Q4_0 cross-suite campaign + + - id: model.gemma4-31b-q4.tower2-serving + text: > + Gemma 4 31B QAT Q4_0 serves stably as two independent full-offload + llama.cpp replicas on Tower2, one per RTX PRO 6000 at 500 W. Each replica + exposes four Q8-KV slots with a hard native 262,144-token context per slot; + the pair reached 290.279 aggregate decode tok/s at eight total concurrent + requests and passed near-native-context recall plus chat/tool contracts. + status: provisional + scope: single Tower2 rig, pinned official GGUF and llama.cpp build + evidence: "tooling/deployments/gemma4-31b-q4-tower2/final-validation.json; benchmarks/gemma4-31b-q4/README.md" + caveats: + - "Single rig and runtime build; cross-engine and cross-day variance are not characterized." + - "Cross-GPU split candidates were corrupt or unsupported, so this is an independent-replica result." + + - id: model.gemma4-31b-q4.canonical-corrected + text: > + On the canonical 12-family MMBT, Gemma scores 29/36 raw and 32/36 + corrected at N=3, and 89/120 raw and 99/120 corrected at N=10. The + correction changes only project-management lexical false negatives and + is tied to unchanged raw grades, reports, archives, and grader hashes. + status: provisional + scope: Gemma 4 31B official QAT Q4_0, model-card sampling, native 256K context + evidence: "benchmarks/gemma4-31b-q4/gemma4-canonical-n10-scorecard.json; benchmarks/gemma4-31b-q4/gemma4-canonical-n10-project-mgmt-correction.json" + caveats: + - "The corrected total depends on a narrow post-run semantic correction; raw grades remain primary and published." + - "Market-research passes are structural and do not prove every citation." + + - id: model.gemma4-31b-q4.beats-qwen36-27b-bounded + text: > + Gemma is materially stronger than Qwen3.6-27B-AWQ thinking on the pinned + directly comparable N=3 quality matrix: 29/36 raw and 32/36 corrected + versus Qwen's 20/36 raw. Short-context single-stream speed is nearly tied, + while Qwen's vLLM stack is substantially stronger under dense batching. + status: provisional + scope: Tower2, historical N=3 MMBT comparison and separate 500 W serving measurements + evidence: "benchmarks/gemma4-31b-q4/comparison.json; tooling/gemma4-comparison-sources.json" + caveats: + - "Sampling, runtime engine, campaign date, and grader coverage differ." + - "Qwen no-think 113/118 is a done_signal rate, not a comparable quality score." + + - id: model.gemma4-31b-q4.extended-zero-of-twelve + text: > + Gemma scores 0/12 under the strict extended substantive audit. Every + single-PR artifact fails pinned-subject provenance, all three investment + workbooks are financially invalid and omit PDF, all three board decks + fail common or visual-quality gates, and all three frozen 75-PR attempts + are materially incomplete. + status: provisional + scope: four extended suites times three replicates on the pinned Tower2 deployment + evidence: "benchmarks/gemma4-31b-q4/substantive-audit.json; benchmarks/gemma4-31b-q4/gemma4-extended-evidence-audit.json" + caveats: + - "Small modality-specific N; older local entries have not all received the same strict overlay." + - "Per-PR factual accuracy is not exhaustively graded; the 75-PR result is artifact completeness and audit substance." + retracted: - id: hw.q8.27b.tower2-decode-narrows-at-long-ctx diff --git a/tooling/GEMMA4-EXTENDED-SUBSTANTIVE-AUDIT-PROTOCOL.md b/tooling/GEMMA4-EXTENDED-SUBSTANTIVE-AUDIT-PROTOCOL.md new file mode 100644 index 00000000..7566b39c --- /dev/null +++ b/tooling/GEMMA4-EXTENDED-SUBSTANTIVE-AUDIT-PROTOCOL.md @@ -0,0 +1,105 @@ +# Gemma 4 extended-suite substantive audit protocol + +This protocol is fixed before any Gemma extended-suite output exists. It keeps +the historical artifact/finish gates, adds the same strict overlays applied to +DeepSeek V4 Flash, and never converts “files exist” into a quality pass. + +## Common evidence gate + +Every run first passes `tooling/audit_gemma4_extended.py`: exact model/runtime, +task and subject hashes, model-card sampling, native context/output policy, +deterministic replica lane, 500 W limits, telemetry, dependency lineage, clean +tagged archive, and attempt preservation. Raw archives and grades remain +immutable. Substantive corrections are separate overlays tied to archive and +grader hashes; no model rerun can replace an unfavorable valid attempt. + +A missing deliverable, early stop, malformed tool use, length exhaustion at the +genuine dynamic context allowance, scroll loop, or wrong audited subject after +the pinned refs were available is a model outcome. Only affirmative endpoint, +fixture, network-before-subject-fetch, or harness evidence can make an attempt +infrastructure-invalid. + +## Single DreamServer PR #1057 + +- Audit the immutable subject in `gemma4-single-pr-subject-pin.json`: base + `309e9cd0`, head `e5ceb43e`, squash `1678f194`, and the original PR commits. + The stale word “open” and current `main` cannot redefine the subject. +- Require the complete requested repository structure, meaningful commit + history, final tag, explicit bounty/AMD/risk treatment, line-level traces, + tool log, questions, dead ends, decisions, and rerunnable evidence. +- Reproduce applicable behavior on the historical state and run documented + baseline/head/current tests. Claimed execution without command output is not + test evidence; a skip must be explicit and scoped. +- Reconcile the important history: runtime code arrived independently through + PR #1039, while #1057's surviving squash contribution is targeted tests. +- Independent expected disposition is **MERGE**. Report separately whether a + different disposition is defensible, over-strict, unsupported, or unsafe; + do not force binary agreement when the artifact contains materially useful + reasoning. + +## Investment memo + +The historical `SHIPPED` gate requires a clean tagged repository, PDF and +source memo, workbook, raw primary sources, extraction/analysis code, decisions, +questions, dead ends, source hashes, and tool log. The substantive overlay also +requires: + +- A US-listed $1B-$10B company at the recorded observation date. +- An eight-page-or-shorter readable memo leading with recommendation, 12-month + target, probability-weighted bear/base/bull cases, risks, differentiated + thesis or explicit efficient-pricing conclusion, confidence, and limits. +- A real three-statement workbook with formulas—not pasted outputs—whose + historical periods reconcile to primary filings and whose statements, + working capital, cash, debt, capex, interest, taxes, and valuation link + consistently. Formula caches are recalculated before inspection. +- Unit, sign, share-count, enterprise/equity-value, FCFF/FCFE, WACC, terminal + value, and scenario-probability checks. Five random memo numbers are traced + end-to-end from source bytes through extraction/model cells to prose. +- Every management quote and external claim is traceable to preserved source + content; citations or raw filenames without the claimed evidence are defects. + +Classify each run as `SUBSTANTIVE_FINANCE_PASS`, +`SHIPPED_WITH_MATERIAL_FINANCE_DEFECTS`, `SCAFFOLD_AND_STOP`, or +`MODEL_TERMINAL_FAILURE`. + +## Board presentation + +The board run must derive from the matching memo replicate; a terminal memo +produces a preserved dependency failure rather than a substituted input. + +- Require PPTX and PDF, 15-25 slides, source/build scripts, standalone chart + data and regeneration scripts, claim/number/quote traces, five-number + reconciliation, storyboard committed before the first deck, alternatives, + audience analysis, decisions, questions, dead ends, sources, and tool log. +- Render every slide and every PDF page. Inspect at full size for clipping, + overlap, tiny text, broken fonts, malformed axes, illegible contrast, + inconsistent scenario colors, missing images, and PPTX/PDF divergence. +- Check all required content, probability-weighted scenario visualization, + reasoning graph, rejected hypotheses, change-my-mind triggers, limitations, + and the self-audit slide. Re-run a sample of chart scripts and compare output. +- Trace five random displayed numbers and every management quote to the exact + matching memo-repository commit/file/line and preserved primary evidence. + +Classify each run as `SUBSTANTIVE_DECK_PASS`, +`SHIPPED_WITH_MATERIAL_DECK_DEFECTS`, `SCAFFOLD_AND_STOP`, +`DEPENDENCY_FAILURE`, or `MODEL_TERMINAL_FAILURE`. + +## Frozen 75-PR audit + +Apply `tooling/MMBT-75PR-AUDIT-PROTOCOL.md` unchanged against fixture baseline +`d5154c3` and canonical PR-set SHA-256 +`569b95b3384af0c4ae4b54a2c8c8f7c908b396124777927a37b5c8fa0211ecd1`. +Run the existing frozen-scope validator, then inspect individual review/test +substance, traceability, overlap/dependency reconciliation, bounty and AMD +coverage, risk matrix, executive reconciliation, boilerplate, and rerunnable +commands. Exactly three valid outcomes remain in the denominator; only proven +infrastructure-invalid attempts receive replacements. + +## Publication + +Publish the historical gate, evidence audit, and substantive classification as +separate axes. Compare Gemma to DeepSeek and historical Qwen artifacts only at +matching axes, with task, sampling, context, date, quantization, and N caveats. +All rendered images, workbook/deck audit JSON, compact hashes, grader versions, +and correction overlays are retained; oversized raw archives remain external +with paths and SHA-256 receipts. diff --git a/tooling/analyze_gemma4_queueing.py b/tooling/analyze_gemma4_queueing.py new file mode 100755 index 00000000..07b51b77 --- /dev/null +++ b/tooling/analyze_gemma4_queueing.py @@ -0,0 +1,123 @@ +#!/usr/bin/env python3 +"""Derive an explicitly labeled queue-wait estimate from preserved concurrency evidence.""" +from __future__ import annotations + +import argparse +import hashlib +import json +import statistics +from datetime import datetime, timezone +from pathlib import Path + + +def sha256(path: Path) -> str: + digest = hashlib.sha256() + with path.open("rb") as handle: + for chunk in iter(lambda: handle.read(1024 * 1024), b""): + digest.update(chunk) + return digest.hexdigest() + + +def median(values: list[float]) -> float: + return round(statistics.median(values), 6) + + +def analyze(source: Path, slots: int, concurrency: int) -> dict: + raw = json.loads(source.read_text()) + candidates = [ + row for row in raw.get("concurrency_results", []) + if row.get("concurrency") == concurrency + ] + if len(candidates) != 1: + raise ValueError(f"expected one concurrency={concurrency} result, found {len(candidates)}") + requests = candidates[0].get("requests") or [] + if len(requests) != concurrency: + raise ValueError(f"expected {concurrency} request results, found {len(requests)}") + if not 0 < slots < concurrency or concurrency % slots: + raise ValueError("this wave estimator requires concurrency to be a multiple of slots") + + derived = [] + for index, request in enumerate(requests): + timings = request.get("timings") or {} + wall = request.get("wall_seconds") + prompt_ms = timings.get("prompt_ms") + predicted_ms = timings.get("predicted_ms") + if not all(isinstance(value, (int, float)) for value in (wall, prompt_ms, predicted_ms)): + raise ValueError(f"request {index} lacks wall/prompt/predicted timing") + server_work = (prompt_ms + predicted_ms) / 1000.0 + derived.append({ + "request_index": request.get("request_index", index), + "wall_s": round(wall, 6), + "server_reported_prompt_plus_decode_s": round(server_work, 6), + "wall_minus_server_reported_work_s": round(wall - server_work, 6), + "tokens_evaluated": request.get("tokens_evaluated"), + "tokens_predicted": request.get("tokens_predicted"), + }) + ordered = sorted(derived, key=lambda row: row["wall_s"]) + first_wave = ordered[:slots] + queued_wave = ordered[slots:slots * 2] + first_wall = median([row["wall_s"] for row in first_wave]) + queued_wall = median([row["wall_s"] for row in queued_wave]) + first_overhead = median([ + row["wall_minus_server_reported_work_s"] for row in first_wave + ]) + queued_overhead = median([ + row["wall_minus_server_reported_work_s"] for row in queued_wave + ]) + return { + "schema_version": 1, + "generated_at": datetime.now(timezone.utc).isoformat(), + "source": { + "path": str(source.resolve()), + "bytes": source.stat().st_size, + "sha256": sha256(source), + }, + "operating_point": { + "parallel_slots": slots, + "simultaneous_requests": concurrency, + "waves": concurrency // slots, + }, + "first_wave": { + "requests": len(first_wave), + "median_wall_s": first_wall, + "median_wall_minus_server_work_s": first_overhead, + }, + "queued_second_wave": { + "requests": len(queued_wave), + "median_wall_s": queued_wall, + "median_wall_minus_server_work_s": queued_overhead, + }, + "derived": { + "second_wave_wall_penalty_s": round(queued_wall - first_wall, 6), + "estimated_queue_wait_delta_s": round(queued_overhead - first_overhead, 6), + "aggregate_decode_tokens_per_second_wall": candidates[0].get( + "aggregate_decode_tokens_per_second_wall" + ), + }, + "requests": derived, + "methodology": ( + "The client released all requests together. With four server slots and eight " + "requests, the four shortest walls are the first service wave and the four " + "longest are the queued wave. estimated_queue_wait_delta_s subtracts each " + "response's llama.cpp-reported prompt+decode work from client wall time, then " + "differences the wave medians. It includes scheduler/HTTP overhead and is an " + "estimate, not a direct server-side queue timestamp." + ), + } + + +def main() -> None: + parser = argparse.ArgumentParser() + parser.add_argument("--source", required=True, type=Path) + parser.add_argument("--slots", required=True, type=int) + parser.add_argument("--concurrency", required=True, type=int) + parser.add_argument("--output", required=True, type=Path) + args = parser.parse_args() + document = analyze(args.source.resolve(), args.slots, args.concurrency) + args.output.parent.mkdir(parents=True, exist_ok=True) + args.output.write_text(json.dumps(document, indent=2) + "\n") + print(json.dumps(document["derived"], sort_keys=True)) + + +if __name__ == "__main__": + main() diff --git a/tooling/audit_gemma4_campaign.py b/tooling/audit_gemma4_campaign.py new file mode 100755 index 00000000..83a82a75 --- /dev/null +++ b/tooling/audit_gemma4_campaign.py @@ -0,0 +1,444 @@ +#!/usr/bin/env python3 +"""Fail-closed audit of Gemma canonical evidence and preserved attempts.""" +from __future__ import annotations + +import argparse +import hashlib +import json +from datetime import datetime, timezone +from pathlib import Path + + +TASKS = [ + "p1_bugfix", "p1_testwrite", "p1_refactor", "p2_extract", "p2_ci", + "p2_hallucination", "p2_triage", "p3_doc", "p3_business", + "p3_market", "p3_writing", "p3_pm", +] +MODEL = "Gemma-4-31B-it-QAT-Q4_0" +MODEL_SHA = "179cfb99212709597eae5929112cfca677e1bbf566178b479ae1da0c4772874b" +SERVER_SHA = "200b403b5735418ff1f6da0cea1938e413e11869ae362e8044a12b0df04622fc" + + +def sha256(path: Path) -> str: + digest = hashlib.sha256() + with path.open("rb") as handle: + for chunk in iter(lambda: handle.read(1024 * 1024), b""): + digest.update(chunk) + return digest.hexdigest() + + +def read_json(path: Path): + try: + return json.loads(path.read_text()) + except (OSError, json.JSONDecodeError): + return None + + +def parse_supplement_mappings(values: list[str]) -> tuple[dict[str, str], list[str]]: + mappings = {} + errors = [] + for value in values: + original, separator, supplemental = value.partition("=") + if not separator or not original or not supplemental: + errors.append( + f"invalid pretelemetry supplement mapping {value!r}; " + "expected CANONICAL_RUN=SUPPLEMENTAL_RUN" + ) + continue + if original in mappings: + errors.append(f"duplicate pretelemetry supplement mapping for {original}") + continue + mappings[original] = supplemental + return mappings, errors + + +def default_root(script_path: Path) -> Path: + """Resolve the repository root for a script installed directly in tooling/.""" + return script_path.resolve().parents[1] + + +def audit_invalid_attempts( + invalid_root: Path, classification_path: Path, label: str, +) -> tuple[list[dict], dict | None, list[str]]: + """Inventory excluded attempts and prove every exclusion is classified.""" + errors: list[str] = [] + attempts = sorted( + attempt for attempt in invalid_root.glob(f"*{label}*") + if attempt.is_dir() + ) if invalid_root.is_dir() else [] + if not attempts: + return [], None, errors + + document = read_json(classification_path) + if not isinstance(document, dict): + return [], None, [f"invalid-attempt classification document missing or invalid: {classification_path}"] + entries = document.get("attempts") + if not isinstance(entries, dict): + return [], None, ["invalid-attempt classification document lacks attempts object"] + + records = [] + observed_names = {attempt.name for attempt in attempts} + extra_names = sorted(set(entries) - observed_names) + if extra_names: + errors.append(f"classified invalid attempts missing from evidence tree: {extra_names}") + + for attempt in attempts: + entry = entries.get(attempt.name) + evidence = {} + for path in sorted(path for path in attempt.rglob("*") if path.is_file()): + evidence[str(path.relative_to(attempt))] = { + "bytes": path.stat().st_size, + "sha256": sha256(path), + } + if not isinstance(entry, dict): + errors.append(f"{attempt.name}: missing invalid-attempt classification") + records.append({"attempt": attempt.name, "files": evidence}) + continue + + source_run = entry.get("source_run") + if not isinstance(source_run, str) or not attempt.name.startswith(source_run): + errors.append(f"{attempt.name}: classification source_run does not match attempt") + if entry.get("classification") != "infrastructure-invalid": + errors.append(f"{attempt.name}: exclusion is not classified infrastructure-invalid") + if entry.get("classified_before_grade") is not True: + errors.append(f"{attempt.name}: exclusion was not classified before grading") + if not entry.get("reason_code"): + errors.append(f"{attempt.name}: classification lacks reason_code") + affirmative = entry.get("affirmative_evidence") + if not isinstance(affirmative, list) or not affirmative: + errors.append(f"{attempt.name}: classification lacks affirmative evidence") + replacement = entry.get("replacement") or {} + if replacement.get("required") is not True or replacement.get("status") != "completed": + errors.append(f"{attempt.name}: exact canonical replacement is not completed") + incident = classification_path.parent / str(entry.get("incident_document") or "") + if not incident.is_file(): + errors.append(f"{attempt.name}: incident document is missing") + + expected_files = entry.get("expected_files") or {} + if not isinstance(expected_files, dict) or not expected_files: + errors.append(f"{attempt.name}: classification lacks expected file hashes") + else: + for relative, expected_hash in sorted(expected_files.items()): + observed = evidence.get(relative) + if observed is None: + errors.append(f"{attempt.name}: classified evidence missing {relative}") + elif observed["sha256"] != expected_hash: + errors.append(f"{attempt.name}: classified evidence hash mismatch for {relative}") + + records.append({ + "attempt": attempt.name, + "classification": entry, + "files": evidence, + }) + + classification_record = { + "path": str(classification_path.resolve()), + "bytes": classification_path.stat().st_size, + "sha256": sha256(classification_path), + } + return records, classification_record, errors + + +def audit_replacement_controls(logs: Path, invalid_records: list[dict]) -> list[str]: + """Prove defect-specific controls on exact canonical replacements.""" + errors: list[str] = [] + for record in invalid_records: + entry = record.get("classification") or {} + replacement = entry.get("replacement") or {} + run_name = replacement.get("canonical_run") + if not run_name: + continue + run_dir = logs / run_name + if not (run_dir / "summary.json").is_file() or not ( + run_dir / "workspace_final.tar.gz" + ).is_file(): + errors.append(f"{record['attempt']}: exact replacement evidence is incomplete") + continue + if entry.get("reason_code") != "server-transport-timeout-below-native-envelope": + continue + receipt = read_json(run_dir / "receipt.json") or {} + processes = (receipt.get("serving") or {}).get("host_processes") or [] + timeouts = [] + for process in processes: + argv = process.get("argv") or [] + for index, value in enumerate(argv[:-1]): + if value == "--timeout": + try: + timeouts.append(int(argv[index + 1])) + except (TypeError, ValueError): + pass + if not timeouts or min(timeouts) < 14400: + errors.append( + f"{record['attempt']}: replacement does not prove a >=14400-second server timeout" + ) + summary = read_json(run_dir / "summary.json") or {} + if "timed out" in str(summary.get("finish_reason") or "").lower(): + errors.append(f"{record['attempt']}: replacement also ended in a transport timeout") + return errors + + +def audit_run(run_dir: Path, allow_pretelemetry: bool, require_grades: bool) -> tuple[dict, list[str], list[str]]: + name = run_dir.name + errors: list[str] = [] + warnings: list[str] = [] + files = {} + for filename in ( + "receipt.json", "transcript.jsonl", "summary.json", "workspace_final.tar.gz", + "cost.json", "gpu_telemetry.json", "grade.json", "label.json", + ): + path = run_dir / filename + if path.is_file(): + files[filename] = {"bytes": path.stat().st_size, "sha256": sha256(path)} + + label = read_json(run_dir / "label.json") + summary = read_json(run_dir / "summary.json") + terminal_label = bool(label and label.get("primary")) + completed = bool(summary and (run_dir / "workspace_final.tar.gz").is_file()) + terminal_only = terminal_label and not completed + if not completed and not terminal_label: + errors.append("not a completed run or explicit terminal label") + return {"run_name": name, "files": files}, errors, warnings + if terminal_only: + primary = label.get("primary") + if primary == "dependency-failure": + if not label.get("source_run") or not label.get("source_primary"): + errors.append("dependency-failure label lacks source evidence") + return { + "run_name": name, "outcome_kind": "dependency-failure", + "terminal_label": primary, "files": files, + }, errors, warnings + + required_files = ["receipt.json", "transcript.jsonl", "cost.json"] + if completed: + required_files.extend(["summary.json", "workspace_final.tar.gz"]) + for required in required_files: + if required not in files: + prefix = "terminal outcome missing preserved " if terminal_only else "missing " + errors.append(f"{prefix}{required}") + if require_grades and not terminal_only and "grade.json" not in files: + errors.append("missing grade.json") + if "gpu_telemetry.json" not in files: + if allow_pretelemetry and not terminal_only: + warnings.append("pre-telemetry valid attempt; supplemental telemetry required") + else: + prefix = "terminal outcome missing preserved " if terminal_only else "missing " + errors.append(f"{prefix}gpu_telemetry.json") + if terminal_only and any( + required not in files + for required in ("receipt.json", "transcript.jsonl", "cost.json", "gpu_telemetry.json") + ): + return { + "run_name": name, "outcome_kind": "terminal-label", + "terminal_label": label.get("primary"), "files": files, + }, errors, warnings + + receipt = read_json(run_dir / "receipt.json") or {} + defaults = receipt.get("inference_request_defaults") or {} + if (receipt.get("harness") or {}).get("git_dirty") is not False: + errors.append("receipt reports dirty harness worktree") + if (receipt.get("vllm") or {}).get("served_model_name") != MODEL: + errors.append("wrong served model in receipt") + expected_defaults = { + "temperature": 1.0, + "top_p": 0.95, + "top_k": 64, + "max_model_len": 262144, + "max_output_tokens_cap": 262144, + } + for key, expected in expected_defaults.items(): + if defaults.get(key) != expected: + errors.append(f"wrong {key}: {defaults.get(key)!r}") + manifest = ((receipt.get("serving") or {}).get("manifest") or {}).get("payload") or {} + if ((manifest.get("artifact") or {}).get("sha256")) != MODEL_SHA: + errors.append("serving manifest model hash mismatch") + processes = (receipt.get("serving") or {}).get("host_processes") or [] + if not processes or any(process.get("exe_sha256") != SERVER_SHA for process in processes): + errors.append("host llama-server provenance mismatch") + models = (((receipt.get("serving") or {}).get("endpoint_models") or {}).get("payload") or {}).get("data") or [] + if [model.get("id") for model in models] != [MODEL]: + errors.append("live /v1/models identity mismatch") + hardware = (receipt.get("hardware") or {}).get("nvidia_smi") or [] + if len(hardware) != 2 or any("500.00 W" not in row for row in hardware): + errors.append("receipt does not prove two 500 W GPU limits") + if summary and summary.get("model") != MODEL: + errors.append("summary model mismatch") + + transcript_path = run_dir / "transcript.jsonl" + transcript = [] + try: + transcript = [json.loads(line) for line in transcript_path.read_text().splitlines() if line.strip()] + except (OSError, json.JSONDecodeError): + errors.append("transcript is not valid JSONL") + model_turns = [row for row in transcript if row.get("type") == "model"] + if not model_turns: + errors.append("transcript has no model turns") + + cost = read_json(run_dir / "cost.json") + if "cost.json" in files: + if not isinstance(cost, dict): + errors.append("cost.json is not valid JSON") + cost = {} + else: + if cost.get("run_name") != name: + errors.append("cost run name mismatch") + if cost.get("model") != MODEL: + errors.append("cost model mismatch") + if not isinstance(cost.get("wall_s"), (int, float)) or cost.get("wall_s") <= 0: + errors.append("cost wall_s is not positive") + + telemetry = read_json(run_dir / "gpu_telemetry.json") + telemetry_summary = None + if telemetry: + observed = (telemetry.get("attribution") or {}).get("active_gpu_ids_observed") or [] + coverage = (telemetry.get("sampling") or {}).get("coverage_fraction_of_wall") + cap = (telemetry.get("active_gpu") or {}).get("configured_cap_w") + if len(observed) != 1 or observed[0] not in ("0", "1"): + errors.append("telemetry does not identify exactly one replica GPU") + if cap != 500.0: + errors.append("telemetry cap is not 500 W") + if terminal_only and not (telemetry.get("sampling") or {}).get("window_source"): + errors.append("terminal telemetry lacks an explicit evidence window source") + if not isinstance(coverage, (int, float)) or coverage < 0.80: + warnings.append(f"telemetry coverage below 80%: {coverage!r}") + telemetry_summary = { + "gpu": observed, + "coverage": coverage, + "mean_power_w": (telemetry.get("active_gpu") or {}).get("mean_power_w"), + "mean_sm_util_pct": (telemetry.get("active_gpu") or {}).get("mean_sm_util_pct"), + "max_temp_c": (telemetry.get("active_gpu") or {}).get("max_temp_c"), + "cpu_package_mean_power_w": (telemetry.get("cpu_package_shared_context") or {}).get("mean_power_w"), + } + + grade = read_json(run_dir / "grade.json") + summary_doc = summary or {} + return { + "run_name": name, + "outcome_kind": "terminal-label" if terminal_only else "completed-workspace", + "terminal_label": label.get("primary") if terminal_only else None, + "harness_git_sha": (receipt.get("harness") or {}).get("git_sha"), + "endpoint": (receipt.get("vllm") or {}).get("api_url"), + "finish_reason": summary_doc.get("finish_reason") or ( + f"terminal:{label.get('primary')}" if terminal_only else None + ), + "elapsed_s": summary_doc.get("elapsed_s") or (cost or {}).get("wall_s"), + "iterations": summary_doc.get("iterations") or (cost or {}).get("iters"), + "completion_tokens": summary_doc.get("total_completion_tokens") or ( + ((cost or {}).get("tokens") or {}).get("completion_total") + ), + "prompt_tokens_cumulative": summary_doc.get("total_prompt_tokens") or ( + ((cost or {}).get("tokens") or {}).get("prompt_total") + ), + "model_turns": len(model_turns), + "length_finishes": sum(row.get("finish_reason") == "length" for row in model_turns), + "grade_verdict": grade.get("verdict") if grade else None, + "telemetry": telemetry_summary, + "files": files, + }, errors, warnings + + +def main() -> None: + parser = argparse.ArgumentParser() + parser.add_argument("--root", type=Path, default=default_root(Path(__file__))) + parser.add_argument("--label", default="gemma4-31b-q4") + parser.add_argument("--target-n", type=int, default=3) + parser.add_argument("--allow-pretelemetry-run", action="append", default=[]) + parser.add_argument( + "--pretelemetry-supplement", action="append", default=[], + metavar="CANONICAL_RUN=SUPPLEMENTAL_RUN", + ) + parser.add_argument("--require-grades", action="store_true") + parser.add_argument( + "--invalid-classifications", type=Path, + default=None, + help="classification ledger for preserved infrastructure-invalid attempts", + ) + parser.add_argument("--raw-telemetry", type=Path, default=Path("/home/michael/gemma4-campaign-state/telemetry/gemma4-31b-q4-gpu.csv")) + parser.add_argument("--output", required=True, type=Path) + args = parser.parse_args() + root = args.root.resolve() + logs = root / "logs" + allowed = set(args.allow_pretelemetry_run) + supplements, mapping_errors = parse_supplement_mappings(args.pretelemetry_supplement) + records = [] + errors = list(mapping_errors) + warnings = [] + for task in TASKS: + for rep in range(1, args.target_n + 1): + name = f"{task}_{args.label}_v{rep}" + record, run_errors, run_warnings = audit_run(logs / name, name in allowed, args.require_grades) + records.append(record) + errors.extend(f"{name}: {message}" for message in run_errors) + warnings.extend(f"{name}: {message}" for message in run_warnings) + + missing_mappings = sorted(allowed - set(supplements)) + extra_mappings = sorted(set(supplements) - allowed) + errors.extend( + f"{name}: allowed pre-telemetry run has no audited supplemental mapping" + for name in missing_mappings + ) + errors.extend( + f"{name}: supplemental mapping supplied without --allow-pretelemetry-run" + for name in extra_mappings + ) + supplemental_records = [] + for original, supplemental in sorted(supplements.items()): + record, run_errors, run_warnings = audit_run( + logs / supplemental, allow_pretelemetry=False, require_grades=False, + ) + record["supplements_run"] = original + supplemental_records.append(record) + errors.extend(f"{supplemental}: {message}" for message in run_errors) + warnings.extend(f"{supplemental}: {message}" for message in run_warnings) + + invalid_root = logs / "_invalid" + classification_path = args.invalid_classifications or ( + root / "tooling" / "deployments" / "gemma4-31b-q4-tower2" / + "invalid-attempt-classifications.json" + ) + invalid, invalid_classification_record, invalid_errors = audit_invalid_attempts( + invalid_root, classification_path, args.label, + ) + errors.extend(invalid_errors) + errors.extend(audit_replacement_controls(logs, invalid)) + + raw = None + if args.raw_telemetry.is_file(): + with args.raw_telemetry.open("rb") as handle: + lines = sum(1 for _ in handle) + raw = { + "path": str(args.raw_telemetry.resolve()), + "bytes": args.raw_telemetry.stat().st_size, + "lines": lines, + "sha256_at_audit_time": sha256(args.raw_telemetry), + } + else: + errors.append(f"raw telemetry missing: {args.raw_telemetry}") + + document = { + "schema_version": 1, + "generated_at": datetime.now(timezone.utc).isoformat(), + "root": str(root), + "label": args.label, + "target_n": args.target_n, + "expected_runs": len(TASKS) * args.target_n, + "audited_runs": len(records), + "passed": not errors, + "errors": errors, + "warnings": warnings, + "raw_telemetry": raw, + "runs": records, + "pretelemetry_supplements": supplemental_records, + "invalid_attempt_classifications": invalid_classification_record, + "preserved_invalid_attempts": invalid, + } + args.output.parent.mkdir(parents=True, exist_ok=True) + args.output.write_text(json.dumps(document, indent=2) + "\n") + print(json.dumps({ + "passed": document["passed"], "runs": len(records), + "errors": len(errors), "warnings": len(warnings), "invalid_attempts": len(invalid), + }, sort_keys=True)) + raise SystemExit(0 if document["passed"] else 1) + + +if __name__ == "__main__": + main() diff --git a/tooling/audit_gemma4_extended.py b/tooling/audit_gemma4_extended.py new file mode 100755 index 00000000..f192bdf9 --- /dev/null +++ b/tooling/audit_gemma4_extended.py @@ -0,0 +1,174 @@ +#!/usr/bin/env python3 +"""Fail-closed provenance and artifact audit for Gemma's extended MMBT suites.""" +from __future__ import annotations + +import argparse +import json +import tarfile +from datetime import datetime, timezone +from pathlib import Path + +from audit_gemma4_campaign import audit_run, read_json, sha256 + + +RUN_PREFIXES = { + "dreamserver-1-pr-audit": "n1_gemma4-31b-q4", + "wallstreet-investment-memo": "gemma4-31b-q4_invest_memo", + "wallstreet-board-presentation": "gemma4-31b-q4_board_pres", + "dreamserver-75-pr-audit": "gemma4-31b-q4_75pr", +} + + +def expected_run_name(suite_id: str, rep: int) -> str: + return f"{RUN_PREFIXES[suite_id]}_v{rep}" + + +def audit_subject_refs(root: Path, suite: dict, run_dir: Path) -> list[str]: + """Prove a one-PR artifact audited the pinned comparator subject.""" + errors = [] + pin_path = root / suite["subject_pin"] + if not pin_path.is_file(): + return [f"subject pin missing: {pin_path}"] + if sha256(pin_path) != suite["subject_pin_sha256"]: + errors.append("subject pin hash mismatch") + archive = run_dir / "workspace_final.tar.gz" + if not archive.is_file(): + return errors + text_parts = [] + try: + with tarfile.open(archive, "r:gz") as handle: + for member in handle.getmembers(): + if not member.isfile() or member.size > 5 * 1024 * 1024: + continue + if Path(member.name).suffix.lower() not in {".md", ".txt", ".json"}: + continue + extracted = handle.extractfile(member) + if extracted is not None: + text_parts.append(extracted.read().decode(errors="replace").lower()) + except (OSError, tarfile.TarError) as exc: + return errors + [f"cannot inspect subject refs in archive: {exc}"] + corpus = "\n".join(text_parts) + for required_sha in suite["required_subject_shas"]: + if required_sha.lower() not in corpus and required_sha[:8].lower() not in corpus: + errors.append(f"artifact does not identify pinned subject ref {required_sha}") + return errors + + +def audit_extended_run(root: Path, suite: dict, rep: int, ordinal: int, + lane_ports: list[int]) -> tuple[dict, list[str], list[str]]: + name = expected_run_name(suite["id"], rep) + run_dir = root / "logs" / name + record, errors, warnings = audit_run(run_dir, False, False) + record.update({"suite": suite["id"], "replicate": rep, "ordinal": ordinal}) + label = read_json(run_dir / "label.json") or {} + if record.get("terminal_label"): + if record["terminal_label"] == "dependency-failure": + expected_source = expected_run_name(suite["input_from"], rep) + if label.get("source_run") != expected_source: + errors.append( + f"dependency-failure source {label.get('source_run')!r} != {expected_source!r}" + ) + return record, errors, warnings + + receipt = read_json(run_dir / "receipt.json") or {} + task = receipt.get("task") or {} + if task.get("sha256") != suite["current_task_sha256"]: + errors.append("task hash differs from pinned extended matrix") + runtime = ((receipt.get("sandbox") or {}).get("runtime") or {}) + if runtime.get("require_git_tag") is not True: + errors.append("extended run did not require a git tag") + expected_port = lane_ports[ordinal % len(lane_ports)] + endpoint = ((receipt.get("vllm") or {}).get("api_url") or "") + if f":{expected_port}/" not in endpoint: + errors.append(f"run endpoint does not match deterministic lane port {expected_port}") + if suite.get("input_from"): + expected_source = expected_run_name(suite["input_from"], rep) + mounted = str(runtime.get("input_mount") or "") + if not mounted.endswith(f"tooling/workspace/{expected_source}"): + errors.append(f"board input mount does not derive from {expected_source}") + if suite.get("input_path"): + mounted = str(runtime.get("input_mount") or "") + if Path(mounted) != Path(suite["input_path"]): + errors.append("frozen-fixture input path mismatch") + if suite.get("subject_pin"): + errors.extend(audit_subject_refs(root, suite, run_dir)) + return record, errors, warnings + + +def main() -> None: + parser = argparse.ArgumentParser() + parser.add_argument("--root", type=Path, default=Path(__file__).resolve().parents[1]) + parser.add_argument( + "--matrix", type=Path, + default=Path(__file__).with_name("gemma4-31b-q4-extended-matrix.json"), + ) + parser.add_argument("--output", required=True, type=Path) + args = parser.parse_args() + root = args.root.resolve() + matrix_path = args.matrix.resolve() + matrix = json.loads(matrix_path.read_text()) + lane_ports = [int(port) for port in matrix["lane_ports"]] + records = [] + errors = [] + warnings = [] + ordinal = 0 + for suite in matrix["suites"]: + if suite["id"] not in RUN_PREFIXES: + errors.append(f"unknown suite id in matrix: {suite['id']}") + continue + for rep in range(1, int(matrix["replicates"]) + 1): + record, run_errors, run_warnings = audit_extended_run( + root, suite, rep, ordinal, lane_ports, + ) + records.append(record) + name = record["run_name"] + errors.extend(f"{name}: {message}" for message in run_errors) + warnings.extend(f"{name}: {message}" for message in run_warnings) + ordinal += 1 + + invalid = [] + invalid_root = root / "logs" / "_infra_invalid" + if invalid_root.is_dir(): + for attempt in sorted(path for path in invalid_root.iterdir() if path.is_dir()): + files = {} + for path in sorted(candidate for candidate in attempt.rglob("*") if candidate.is_file()): + files[str(path.relative_to(attempt))] = { + "bytes": path.stat().st_size, + "sha256": sha256(path), + } + invalid.append({"attempt": attempt.name, "files": files}) + + expected = len(matrix["suites"]) * int(matrix["replicates"]) + document = { + "schema_version": 1, + "generated_at": datetime.now(timezone.utc).isoformat(), + "root": str(root), + "matrix": { + "path": str(matrix_path), + "sha256": sha256(matrix_path), + }, + "expected_runs": expected, + "audited_runs": len(records), + "passed": len(records) == expected and not errors, + "errors": errors, + "warnings": warnings, + "runs": records, + "preserved_infrastructure_invalid_attempts": invalid, + "scope_note": ( + "This audit proves identity, configuration, routing, telemetry, and artifact " + "preservation. Suite-specific substantive grading and visual/workbook/code " + "inspection are separate required overlays; this document is not a quality pass." + ), + } + args.output.parent.mkdir(parents=True, exist_ok=True) + args.output.write_text(json.dumps(document, indent=2) + "\n") + print(json.dumps({ + "passed": document["passed"], "runs": len(records), + "errors": len(errors), "warnings": len(warnings), + "invalid_attempts": len(invalid), + }, sort_keys=True)) + raise SystemExit(0 if document["passed"] else 1) + + +if __name__ == "__main__": + main() diff --git a/tooling/bench_autopilot.py b/tooling/bench_autopilot.py old mode 100644 new mode 100755 index d1640017..83135b61 --- a/tooling/bench_autopilot.py +++ b/tooling/bench_autopilot.py @@ -47,7 +47,7 @@ HEARTBEAT_FRESH_SECS = 120 SUBSTANCE_CHECK_SECS = 300 -_last_substance_check = 0.0 +_last_substance_check: dict[str, float] = {} TASKS = ["p1_bugfix","p1_testwrite","p1_refactor","p2_extract","p2_ci", "p2_hallucination","p2_triage","p3_doc","p3_business","p3_market", @@ -198,33 +198,116 @@ def endpoint_up(port: int) -> bool: def container_running(name: str) -> bool: + if not name: + return False r = sh(["docker", "ps", "--format", "{{.Names}}"]) return name in r.stdout.split() +def configured_ports(cfg: dict) -> list[int]: + ports = cfg.get("lane_ports") or [cfg["port"]] + return list(dict.fromkeys(int(port) for port in ports)) + + +def active_harnesses_by_port() -> dict[int, dict]: + """Return exact live harness PID/run mappings for each endpoint lane.""" + active = {} + for proc_dir in Path("/proc").glob("[0-9]*"): + try: + argv = [ + part.decode(errors="replace") + for part in (proc_dir / "cmdline").read_bytes().split(b"\0") + if part + ] + harness_index = next( + index for index, arg in enumerate(argv) + if Path(arg).name == "harness.py" + ) + port_index = argv.index("--port") + run_name = argv[harness_index + 1] + port = int(argv[port_index + 1]) + except (OSError, StopIteration, ValueError, IndexError): + continue + active[port] = {"pid": int(proc_dir.name), "cell": run_name} + return active + + +def slot_progress_signature(payload) -> tuple: + """Extract active prompt/decode progress from llama.cpp's /slots payload.""" + if not isinstance(payload, list): + return () + processing = [] + for slot in payload: + if not isinstance(slot, dict) or not slot.get("is_processing"): + continue + next_token = slot.get("next_token") or [] + token_state = next_token[0] if next_token and isinstance(next_token[0], dict) else {} + processing.append(( + slot.get("id"), + slot.get("id_task"), + slot.get("n_prompt_tokens_processed"), + token_state.get("n_decoded"), + token_state.get("n_remain"), + )) + return tuple(sorted(processing, key=lambda row: (str(row[0]), str(row[1])))) + + +def endpoint_slot_progress(port: int) -> tuple: + result = sh(["curl", "-fsS", "--max-time", "5", f"http://127.0.0.1:{port}/slots"]) + if result.returncode != 0: + return () + try: + return slot_progress_signature(json.loads(result.stdout)) + except json.JSONDecodeError: + return () + + +def configured_services(cfg: dict) -> list[str]: + return [str(service) for service in (cfg.get("services") or [])] + + +def service_running(name: str) -> bool: + return sh(["systemctl", "--user", "is-active", "--quiet", name]).returncode == 0 + + +def runtime_running(cfg: dict) -> bool: + services = configured_services(cfg) + if services: + return all(service_running(service) for service in services) + return container_running(cfg.get("container")) + + def ensure_endpoint(cfg: dict) -> bool: engine = cfg.get("engine", "llamacpp") if engine == "external": - port, name = cfg["port"], cfg["container"] - if endpoint_up(port): + ports = configured_ports(cfg) + name = cfg.get("container") or ", ".join(configured_services(cfg)) or "external runtime" + if all(endpoint_up(port) for port in ports): return True launcher = cfg.get("launcher") if not launcher: raise ValueError("external engine config requires a launcher") - log(f"endpoint down — invoking external launcher {launcher}") - notify("bench: endpoint restart", f"{name} down — invoking {launcher}", 0) - sh(["bash", launcher]) - for _ in range(cfg["endpoint_grace_secs"] // 5): - if endpoint_up(port): - log("endpoint back up") + missing = [port for port in ports if not endpoint_up(port)] + log(f"endpoint lane(s) {missing} down — invoking external launcher {launcher}") + notify("bench: endpoint restart", f"{name} lane(s) {missing} down — invoking {launcher}", 0) + launched = sh(["bash", launcher]) + if launched.returncode != 0: + detail = (launched.stderr or launched.stdout).strip()[-800:] + log(f"ERROR: external launcher failed rc={launched.returncode}: {detail}") + notify("bench: endpoint FAILED", f"{name} launcher failed rc={launched.returncode}", 1) + return False + for _ in range(max(1, cfg["endpoint_grace_secs"] // 5)): + if all(endpoint_up(port) for port in ports): + log(f"all endpoint lanes back up: {ports}") return True - if not container_running(name): - log("ERROR: external endpoint container died during load") + if not runtime_running(cfg): + log("ERROR: external inference runtime died during load") notify("bench: endpoint FAILED", f"{name} died during model load", 1) return False time.sleep(5) - log("ERROR: external endpoint did not come up within grace period") - notify("bench: endpoint FAILED", f"{name} did not load within grace period", 1) + missing = [port for port in ports if not endpoint_up(port)] + log(f"ERROR: external endpoint lane(s) {missing} did not come up within grace period") + notify("bench: endpoint FAILED", f"{name} lane(s) {missing} did not load within grace period", 1) return False if engine != "llamacpp": note = cfg.get("_notimplemented", f"engine '{engine}' has no launcher") @@ -406,8 +489,39 @@ def kill_stuck(cell: str): sh(["docker", "rm", "-f", f"bench-sandbox-{cell}"]) +def active_substance_targets(cfg) -> list[dict]: + """Return every live lane's exact run/transcript/PID for substance checks.""" + targets = [] + active = active_harnesses_by_port() + for port in configured_ports(cfg): + harness = active.get(port) + if not harness: + continue + transcript = LOGS / harness["cell"] / "transcript.jsonl" + if transcript.exists(): + targets.append({ + "port": port, + "cell": harness["cell"], + "pid": harness["pid"], + "transcript": transcript, + }) + if targets: + return targets + # Preserve the historical single-lane fallback for non-Linux/dev contexts + # where procfs attribution is unavailable. + transcript, _ = newest_transcript(cfg) + if transcript and transcript.exists(): + return [{ + "port": int(cfg["port"]), + "cell": transcript.parent.name, + "pid": None, + "transcript": transcript, + }] + return [] + + def check_current_substance(cfg): - """Run the published five-minute substance monitor on the active cell. + """Run the published five-minute substance monitor on every active lane. Scroll loops are terminated by exact harness PID and receive the canonical explicit failure label. Runaway-generation signals are logged for endpoint @@ -415,40 +529,44 @@ def check_current_substance(cfg): """ global _last_substance_check now = time.time() - if now - _last_substance_check < SUBSTANCE_CHECK_SECS: - return - _last_substance_check = now - tp, _ = newest_transcript(cfg) - if not tp or not tp.exists(): - return - cell = tp.parent.name - result = sh(["python3", str(SCRIPTS / "check_substance.py"), str(tp)]) - with open(STATE_DIR / f"substance-{cell}.log", "a") as f: - f.write(f"\n[{time.strftime('%Y-%m-%dT%H:%M:%SZ', time.gmtime())}] rc={result.returncode}\n") - f.write(result.stdout or "") - f.write(result.stderr or "") - if result.returncode == 1: - pids = sh(["pgrep", "-f", f"harness.py {cell} "]).stdout.split() - if not pids: - log(f"SUBSTANCE: scroll-loop detected for {cell}, but exact harness PID was absent") - return - log(f"SUBSTANCE: scroll-loop detected for {cell}; SIGTERM exact PID(s) {','.join(pids)}") - for pid in pids: - sh(["kill", "-TERM", pid]) - label = { - "primary": "identical-call-loop", - "sub_labels": ["scroll-loop"], - "notes": ( - "Operator-SIGTERM per the documented >=30 identical digit-stripped " - "tool-command substance rule; see the preserved transcript and monitor log." - ), - "labeler": "operator-monitoring-supervisor", - "labeled_at": time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()), - } - (tp.parent / "label.json").write_text(json.dumps(label, indent=2) + "\n") - sh(["docker", "rm", "-f", f"bench-sandbox-{cell}"]) - elif result.returncode == 2: - log(f"SUBSTANCE: runaway-generation suspected for {cell}; endpoint_up={endpoint_up(cfg['port'])}") + for target in active_substance_targets(cfg): + cell = target["cell"] + if now - _last_substance_check.get(cell, 0.0) < SUBSTANCE_CHECK_SECS: + continue + _last_substance_check[cell] = now + transcript = target["transcript"] + result = sh(["python3", str(SCRIPTS / "check_substance.py"), str(transcript)]) + with open(STATE_DIR / f"substance-{cell}.log", "a") as f: + f.write(f"\n[{time.strftime('%Y-%m-%dT%H:%M:%SZ', time.gmtime())}] rc={result.returncode}\n") + f.write(result.stdout or "") + f.write(result.stderr or "") + if result.returncode == 1: + pid = target["pid"] + pids = [str(pid)] if pid is not None else \ + sh(["pgrep", "-f", f"harness.py {cell} "]).stdout.split() + if not pids: + log(f"SUBSTANCE: scroll-loop detected for {cell}, but exact harness PID was absent") + continue + log(f"SUBSTANCE: scroll-loop detected for {cell}; SIGTERM exact PID(s) {','.join(pids)}") + for exact_pid in pids: + sh(["kill", "-TERM", exact_pid]) + label = { + "primary": "identical-call-loop", + "sub_labels": ["scroll-loop"], + "notes": ( + "Operator-SIGTERM per the documented >=30 identical digit-stripped " + "tool-command substance rule; see the preserved transcript and monitor log." + ), + "labeler": "operator-monitoring-supervisor", + "labeled_at": time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()), + } + (transcript.parent / "label.json").write_text(json.dumps(label, indent=2) + "\n") + sh(["docker", "rm", "-f", f"bench-sandbox-{cell}"]) + elif result.returncode == 2: + log( + f"SUBSTANCE: runaway-generation suspected for {cell}; " + f"endpoint_up={endpoint_up(target['port'])}" + ) # --- recent cells + fails --------------------------------------------------- @@ -540,12 +658,15 @@ def write_status(cfg, target, phase, arms_done, started_at=None): eta_secs = int(statistics.median(all_walls) * rem) recent, fails = recent_cells_and_fails(cfg) now = time.time() + lane_endpoint_status = { + str(port): endpoint_up(port) for port in configured_ports(cfg) + } status = { # ---- original keys (unchanged) ---- "updated": time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()), "target_n": target, "phase": phase, - "endpoint_up": endpoint_up(cfg["port"]), - "container_up": container_running(cfg["container"]), + "endpoint_up": all(lane_endpoint_status.values()), + "container_up": runtime_running(cfg), "grand_done": grand_done, "grand_total": grand_total, "pct": round(100 * grand_done / grand_total, 1) if grand_total else 0, "current": cur, "arms": arms, "arms_done": arms_done, "tasks": TASKS, @@ -558,6 +679,8 @@ def write_status(cfg, target, phase, arms_done, started_at=None): "preset": cfg.get("preset"), # addition A: which preset is driving "model": cfg.get("model"), "engine": cfg.get("engine", "llamacpp"), + "lane_endpoints": lane_endpoint_status, + "runtime_up": runtime_running(cfg), } STATUS.write_text(json.dumps(status, indent=2)) return status @@ -704,36 +827,113 @@ def git(*args, check=False): log(f"publish: git step failed (non-fatal): {e}") +def benchmark_environment(cfg, base_env=None): + """Return the child environment carrying a model's pinned sampling point. + + ``run_microbench.sh`` intentionally keeps historical defaults, so model-card + campaigns must pass their explicit overrides through the supervisor. Keep + this in a testable helper: a config field that is only written to a receipt + or README but not sent to the harness is a methodology failure. + """ + run_env = dict(os.environ if base_env is None else base_env) + sampling_env = { + "benchmark_temperature": "BENCH_TEMP", + "benchmark_top_p": "BENCH_TOP_P", + "benchmark_top_k": "BENCH_TOP_K", + "benchmark_max_output_tokens_cap": "BENCH_MAX_OUTPUT_TOKENS_CAP", + "serving_manifest": "BENCH_SERVING_MANIFEST", + } + for config_key, env_key in sampling_env.items(): + value = cfg.get(config_key) + if value is not None: + run_env[env_key] = str(value) + else: + run_env.pop(env_key, None) + return run_env + + def run_arm_with_supervision(cfg, arm, target, started_at): - """Spawn run_microbench for one arm; watchdog endpoint + stuck cells until it exits.""" + """Shard one arm deterministically across configured replica endpoints.""" label, thinking = arm["label"], arm["thinking"] ids = sh("docker ps -aq --filter name=bench-sandbox-").stdout.split() if ids: sh(["docker", "rm", "-f", *ids]) sh(f"sudo rm -rf {TOOLING}/workspace/*{label}_v* 2>/dev/null") - log(f"RUN arm {label} (thinking={thinking}) target N={target}") - proc = subprocess.Popen( - ["bash", str(SCRIPTS / "run_microbench.sh"), cfg["model"], str(cfg["port"]), - label, str(target), "", thinking, str(cfg["max_model_len"])], - stdout=open(STATE_DIR / f"run-{label}.log", "a"), stderr=subprocess.STDOUT) - last_endpoint_ok = time.time() - while proc.poll() is None: - time.sleep(30) - write_heartbeat() - write_status(cfg, target, f"run:{label}", [], started_at) - check_harness_anomaly(cfg) # addition C - check_current_substance(cfg) - if not endpoint_up(cfg["port"]): - if time.time() - last_endpoint_ok > 90: - log("watchdog: endpoint down >90s — restarting") - ensure_endpoint(cfg) - last_endpoint_ok = time.time() - else: - last_endpoint_ok = time.time() - cur = current_cell_info(cfg) - if cur["cell"] and cur["frozen_secs"] and cur["frozen_secs"] > cfg["stuck_secs"]: - kill_stuck(cur["cell"]) - log(f"run_microbench {label} exited rc={proc.returncode}") + ports = configured_ports(cfg) + lane_count = len(ports) + log(f"RUN arm {label} (thinking={thinking}) target N={target} lanes={ports}") + procs = [] + handles = [] + try: + for lane_index, port in enumerate(ports): + run_env = benchmark_environment(cfg) + run_env["BENCH_LANE_INDEX"] = str(lane_index) + run_env["BENCH_LANE_COUNT"] = str(lane_count) + log_path = STATE_DIR / ( + f"run-{label}.log" if lane_count == 1 + else f"run-{label}-lane{lane_index}-port{port}.log" + ) + handle = open(log_path, "a") + handles.append(handle) + proc = subprocess.Popen( + ["bash", str(SCRIPTS / "run_microbench.sh"), cfg["model"], str(port), + label, str(target), "", thinking, str(cfg["max_model_len"])], + stdout=handle, stderr=subprocess.STDOUT, env=run_env) + procs.append((lane_index, port, proc)) + + last_endpoint_ok = {port: time.time() for port in ports} + lane_progress = { + port: {"cell": None, "signal": None, "last_progress": time.time()} + for port in ports + } + while any(proc.poll() is None for _, _, proc in procs): + time.sleep(30) + write_heartbeat() + write_status(cfg, target, f"run:{label}", [], started_at) + check_harness_anomaly(cfg) # addition C + check_current_substance(cfg) + for port in ports: + if endpoint_up(port): + last_endpoint_ok[port] = time.time() + elif time.time() - last_endpoint_ok[port] > 90: + log(f"watchdog: endpoint lane {port} down >90s — restarting runtime") + ensure_endpoint(cfg) + last_endpoint_ok[port] = time.time() + # Transcript mtimes do not advance while urllib waits for a + # non-streamed response. Watch each replica independently and count + # live llama.cpp slot-token movement as progress so native-context + # generations are not mistaken for hung harnesses. + now = time.time() + active = active_harnesses_by_port() + for port in ports: + harness = active.get(port) + watch = lane_progress[port] + if not harness: + watch.update(cell=None, signal=None, last_progress=now) + continue + cell = harness["cell"] + transcript = LOGS / cell / "transcript.jsonl" + try: + transcript_mtime_ns = transcript.stat().st_mtime_ns + except OSError: + transcript_mtime_ns = None + signal = (cell, transcript_mtime_ns, endpoint_slot_progress(port)) + if watch["cell"] != cell or watch["signal"] != signal: + watch.update(cell=cell, signal=signal, last_progress=now) + continue + frozen_secs = int(now - watch["last_progress"]) + if frozen_secs > cfg["stuck_secs"]: + log( + f"STUCK: lane port {port} cell {cell} showed no transcript " + f"or server-slot progress for {frozen_secs}s" + ) + kill_stuck(cell) + watch.update(cell=None, signal=None, last_progress=now) + lane_results = {port: proc.returncode for _, port, proc in procs} + log(f"run_microbench {label} lane exit codes={lane_results}") + finally: + for handle in handles: + handle.close() sh(f"sudo rm -rf /tmp/grade_*{label}_v* 2>/dev/null") g = sh(["bash", str(SCRIPTS / "grade_microbench.sh"), label]) (STATE_DIR / f"grade-{label}.log").write_text(g.stdout + g.stderr) diff --git a/tooling/capture_gemma4_grader_manifest.py b/tooling/capture_gemma4_grader_manifest.py new file mode 100755 index 00000000..0c4e88b4 --- /dev/null +++ b/tooling/capture_gemma4_grader_manifest.py @@ -0,0 +1,130 @@ +#!/usr/bin/env python3 +"""Fingerprint the exact grader stack and raw Gemma grades for publication.""" +from __future__ import annotations + +import argparse +import hashlib +import json +import subprocess +import sys +from datetime import datetime, timezone +from pathlib import Path + +from summarize_gemma4_campaign import TASKS + + +GRADER_FILES = [ + "tooling/scripts/grade_microbench.sh", + "tooling/graders/phase1_grade.py", + "tooling/graders/code_task_grader.py", + "tooling/graders/phase2_extraction_grade.py", + "tooling/graders/phase2_ci_failure_grade.py", + "tooling/graders/phase2_hallucination_grade.py", + "tooling/graders/phase2_triage_grade.py", + "tooling/graders/phase3_doc_synthesis_grade.py", + "tooling/graders/phase3_business_memo_grade.py", + "tooling/graders/phase3_market_research_grade.py", + "tooling/graders/phase3_writing_editing_grade.py", + "tooling/graders/phase3_project_mgmt_grade.py", + "tooling/graders/ground_truth/phase2_extraction.json", + "tooling/graders/ground_truth/phase2_hallucination.json", + "tooling/graders/ground_truth/phase2_triage.json", + "tooling/graders/ground_truth/phase3_doc_synthesis.json", + "tooling/graders/ground_truth/phase3_business_memo.json", + "tooling/graders/ground_truth/phase3_market_research_rubric.json", + "tooling/graders/ground_truth/phase3_project_mgmt.json", + "tooling/inputs/phase3_writing_editing/audience_briefs.json", +] + + +def sha256(path: Path) -> str: + digest = hashlib.sha256() + with path.open("rb") as handle: + for chunk in iter(lambda: handle.read(1024 * 1024), b""): + digest.update(chunk) + return digest.hexdigest() + + +def fingerprint_files(root: Path, relative_paths: list[str]) -> tuple[dict, list[str]]: + files = {} + errors = [] + for relative in relative_paths: + path = root / relative + if not path.is_file(): + errors.append(f"missing grader input: {relative}") + continue + files[relative] = {"bytes": path.stat().st_size, "sha256": sha256(path)} + return files, errors + + +def command_output(command: list[str]) -> str | None: + result = subprocess.run(command, capture_output=True, text=True, check=False) + return result.stdout.strip() if result.returncode == 0 else None + + +def main() -> None: + parser = argparse.ArgumentParser() + parser.add_argument("--root", type=Path, default=Path(__file__).resolve().parents[1]) + parser.add_argument("--label", default="gemma4-31b-q4") + parser.add_argument("--target-n", type=int, required=True) + parser.add_argument("--output", type=Path, required=True) + args = parser.parse_args() + root = args.root.resolve() + files, errors = fingerprint_files(root, GRADER_FILES) + grades = [] + for task in TASKS: + for rep in range(1, args.target_n + 1): + name = f"{task}_{args.label}_v{rep}" + run = root / "logs" / name + grade = run / "grade.json" + label = run / "label.json" + if grade.is_file(): + payload = json.loads(grade.read_text()) + grades.append({ + "run_name": name, "verdict": payload.get("verdict"), + "grade_json": {"bytes": grade.stat().st_size, "sha256": sha256(grade)}, + }) + elif label.is_file(): + payload = json.loads(label.read_text()) + grades.append({ + "run_name": name, "terminal_label": payload.get("primary"), + "label_json": {"bytes": label.stat().st_size, "sha256": sha256(label)}, + }) + else: + errors.append(f"{name}: neither raw grade nor terminal label exists") + git_status = command_output(["git", "-C", str(root), "status", "--porcelain"]) + if git_status: + errors.append("grader manifest captured from dirty worktree") + document = { + "schema_version": 1, + "generated_at": datetime.now(timezone.utc).isoformat(), + "root": str(root), + "label": args.label, + "target_n": args.target_n, + "passed": not errors, + "errors": errors, + "repository_commit": command_output(["git", "-C", str(root), "rev-parse", "HEAD"]), + "python": sys.version, + "docker_version": command_output(["docker", "version", "--format", "{{.Server.Version}}"]), + "bench_sandbox_image_id": command_output([ + "docker", "image", "inspect", "bench-sandbox:latest", "--format", "{{.Id}}", + ]), + "grader_files": files, + "raw_grades": grades, + "correction_policy": ( + "grade.json and label.json hashes are immutable raw evidence. Any correction " + "must be a separate overlay naming the original hash, corrected grader hash, " + "unchanged workspace archive hash, and reproducible defect." + ), + } + args.output.parent.mkdir(parents=True, exist_ok=True) + args.output.write_text(json.dumps(document, indent=2) + "\n") + print(json.dumps({ + "passed": document["passed"], "grader_files": len(files), + "raw_grades": len(grades), "errors": len(errors), + }, sort_keys=True)) + raise SystemExit(0 if document["passed"] else 1) + + +if __name__ == "__main__": + main() diff --git a/tooling/correct_gemma4_project_mgmt_grades.py b/tooling/correct_gemma4_project_mgmt_grades.py new file mode 100755 index 00000000..bdd23eaa --- /dev/null +++ b/tooling/correct_gemma4_project_mgmt_grades.py @@ -0,0 +1,203 @@ +#!/usr/bin/env python3 +"""Create a non-destructive overlay for narrow project-management grader misses.""" +from __future__ import annotations + +import argparse +import copy +import hashlib +import json +import re +import tarfile +from datetime import datetime, timezone +from pathlib import Path + + +RULES = { + "R2": { + "description": "Maevia GA-to-private-beta pushback is expressed without one contiguous legacy keyword", + "patterns": [ + r"\bmaevia\b.{0,240}\b(?:push(?:ed)?[ -]?back|fallout|promis(?:e|ed)\s+ga|private[ -]?beta)\b", + ], + }, + "R3": { + "description": "legal/private-beta contract delay uses an equivalent non-contracted phrase", + "patterns": [ + r"\blegal\b.{0,200}\b(?:has\s+not|not\s+yet\s+responded|sign[ -]?off|approval|unresponsive|silent|delay|contract)\b", + ], + }, + "D3_mobile": { + "description": "web-responsive V1 and native V2 uses a hyphen the legacy keyword omits", + "patterns": [ + r"\bweb[ -]?responsive\b.{0,160}\bnative\b.{0,100}\bv2\b", + ], + }, + "D4_option_b": { + "description": "private beta followed by the selected-customer count is Option B", + "patterns": [ + r"\bprivate[ -]?beta\b.{0,120}\b(?:3\s*[-–]\s*5|selected\s+customers?)\b", + ], + }, +} + + +def sha256(path: Path) -> str: + digest = hashlib.sha256() + with path.open("rb") as handle: + for chunk in iter(lambda: handle.read(1024 * 1024), b""): + digest.update(chunk) + return digest.hexdigest() + + +def sha256_bytes(data: bytes) -> str: + return hashlib.sha256(data).hexdigest() + + +def read_archived_report(archive: Path) -> tuple[str, str]: + """Read status_report.md from immutable run evidence without extracting it.""" + with tarfile.open(archive, "r:gz") as bundle: + matches = [ + member + for member in bundle.getmembers() + if member.isfile() + and member.name.removeprefix("./") == "status_report.md" + ] + if len(matches) != 1: + raise ValueError( + f"expected one ./status_report.md in {archive}, found {len(matches)}" + ) + handle = bundle.extractfile(matches[0]) + if handle is None: + raise ValueError(f"could not read ./status_report.md from {archive}") + data = handle.read() + return data.decode("utf-8"), sha256_bytes(data) + + +def apply_correction(raw: dict, report_text: str) -> tuple[dict, list[dict]]: + corrected = copy.deepcopy(raw) + details = corrected.get("details") or {} + changes = [] + normalized = re.sub(r"\s+", " ", report_text.lower()) + for item, rule in RULES.items(): + category = "risks" if item.startswith("R") else "decisions" + result = (details.get(category) or {}).get(item) + if not isinstance(result, dict) or result.get("matched"): + continue + matched_pattern = next( + (pattern for pattern in rule["patterns"] if re.search(pattern, normalized)), + None, + ) + if matched_pattern: + result["matched"] = True + result["keyword"] = "correction-overlay-semantic-equivalent" + changes.append({ + "item": item, + "category": category, + "reason": rule["description"], + "pattern": matched_pattern, + }) + + scores = corrected.get("scores") or {} + thresholds = corrected.get("thresholds") or {} + workstream_count = sum( + bool(value.get("matched")) for value in (details.get("workstreams") or {}).values() + ) + risk_count = sum( + bool(value.get("matched")) for value in (details.get("risks") or {}).values() + ) + decision_count = sum( + bool(value.get("matched")) for value in (details.get("decisions") or {}).values() + ) + milestone_count = sum( + bool(value.get("matched")) for value in (details.get("milestones") or {}).values() + ) + scores["workstream_recall"] = f"{workstream_count}/6" + scores["risk_recall"] = f"{risk_count}/6" + scores["decision_recall"] = f"{decision_count}/4" + scores["milestone_recall"] = f"{milestone_count}/5" + corrected["verdict"] = "PASS" if ( + workstream_count >= thresholds.get("min_workstreams", 4) + and risk_count >= thresholds.get("min_risks", 3) + and decision_count >= thresholds.get("min_decisions", 3) + and milestone_count >= thresholds.get("min_milestones", 3) + and scores.get("word_count", 10**9) <= thresholds.get("max_word_count", 700) + and len(scores.get("sections_present") or []) >= 3 + ) else "FAIL" + return corrected, changes + + +def build(root: Path, target_n: int, script_path: Path) -> tuple[dict, list[str]]: + raw_grader = root / "tooling" / "graders" / "phase3_project_mgmt_grade.py" + errors = [] + cells = [] + for replicate in range(1, target_n + 1): + name = f"p3_pm_gemma4-31b-q4_v{replicate}" + run = root / "logs" / name + raw_grade = run / "grade.json" + archive = run / "workspace_final.tar.gz" + missing = [str(path) for path in (raw_grade, archive) if not path.is_file()] + if missing: + errors.append(f"{name}: missing {missing}") + continue + try: + report_text, report_sha256 = read_archived_report(archive) + except (OSError, tarfile.TarError, UnicodeDecodeError, ValueError) as exc: + errors.append(f"{name}: archived status report error: {exc}") + continue + raw = json.loads(raw_grade.read_text()) + corrected, changes = apply_correction(raw, report_text) + cells.append({ + "run_name": name, + "raw_grade_sha256": sha256(raw_grade), + "workspace_archive_sha256": sha256(archive), + "status_report_source": "workspace_final.tar.gz:./status_report.md", + "status_report_sha256": report_sha256, + "raw_verdict": raw.get("verdict"), + "corrected_verdict": corrected.get("verdict"), + "changes": changes, + "corrected_grade": corrected, + }) + return { + "schema_version": 1, + "generated_at": datetime.now(timezone.utc).isoformat(), + "campaign": "gemma4-31b-q4-mmbt", + "target_n": target_n, + "scope": "p3_pm legacy lexical false negatives only", + "policy": "immutable grade.json files remain raw evidence; this overlay changes no run artifact or raw verdict", + "correction_script": { + "path": str(script_path.resolve()), + "sha256": sha256(script_path), + }, + "raw_grader": { + "path": str(raw_grader.resolve()), + "sha256": sha256(raw_grader), + }, + "rules": RULES, + "aggregate": { + "cells": len(cells), + "raw_passes": sum(cell["raw_verdict"] == "PASS" for cell in cells), + "corrected_passes": sum(cell["corrected_verdict"] == "PASS" for cell in cells), + "verdict_changes": sum( + cell["raw_verdict"] != cell["corrected_verdict"] for cell in cells + ), + }, + "cells": cells, + }, errors + + +def main() -> None: + parser = argparse.ArgumentParser() + parser.add_argument("--root", type=Path, default=Path(__file__).resolve().parents[1]) + parser.add_argument("--target-n", type=int, required=True) + parser.add_argument("--output", type=Path, required=True) + args = parser.parse_args() + document, errors = build(args.root.resolve(), args.target_n, Path(__file__)) + document["errors"] = errors + document["passed"] = not errors + args.output.parent.mkdir(parents=True, exist_ok=True) + args.output.write_text(json.dumps(document, indent=2) + "\n") + print(json.dumps({**document["aggregate"], "errors": len(errors)}, sort_keys=True)) + raise SystemExit(0 if not errors else 1) + + +if __name__ == "__main__": + main() diff --git a/tooling/deployments/gemma4-31b-q4-tower2/COMPLETION-CHECKLIST.md b/tooling/deployments/gemma4-31b-q4-tower2/COMPLETION-CHECKLIST.md new file mode 100644 index 00000000..1f2af210 --- /dev/null +++ b/tooling/deployments/gemma4-31b-q4-tower2/COMPLETION-CHECKLIST.md @@ -0,0 +1,107 @@ +# Gemma 4 31B Q4 campaign completion checklist + +This checklist maps each required campaign outcome to authoritative evidence. +It is an operational completion gate, not a substitute for the evidence and +does not relax `PREREGISTRATION.md`. + +## 1. Artifact and cohort pin + +- [x] Official QAT Q4_0 model, matching mmproj, Hugging Face revision, byte + sizes, and local SHA-256 values are recorded in `model-manifest.json`. +- [x] MMBT base, task fixtures, model-card sampling, native context, topology + candidates, validity policy, canonical N=3, N=10 expansion, and extended N=3 + suites were pinned before their applicable runs. +- [ ] Final validator output proves the pinned files and local model artifacts + still match at publication time. + +## 2. Tower2 serving choice + +- [x] One-GPU full offload, dual-layer split, dual-row split, and independent + replicas were measured under the 500 W limits. +- [x] Independent full-offload replicas won the preregistered quality-preserving + rule and passed chat, tool, concurrency, restart, and near-256K recall + gates; the decision and rejected candidates are documented in `README.md`, + `topology-matrix.json`, `MICROBENCH-INDEX.md`, and `final-validation.json`. +- [ ] Publication explicitly reports queued-request latency or queue wait at the + accepted four-slot operating point rather than inferring it from aggregate + concurrency alone. +- [ ] Publication artifacts include the exact accepted runtime/model hashes and + uncached-versus-cached labels for every quoted performance result. + +## 3. Sanctuary and Pixel + +- [x] Pre-campaign fallback and pre-canonical Gemma route checks are preserved. +- [x] Sanctuary and Pixel were placed on separate replica endpoints. +- [ ] Fresh post-campaign checks prove both agents can complete a real tool turn + before publication. +- [ ] After publication, the byte-for-byte pre-Gemma OpenClaw configuration is + restored, DeepSeek-V4-Flash-0731 is restarted, both agents use it without + fallback, and fresh end-to-end markers and tool turns pass. + +## 4. Instrumentation and validity + +- [x] Receipts capture model/runtime identity, live endpoint, sampling/context, + power caps, task hash, repository state, and archive provenance. +- [x] Per-replica GPU plus shared CPU-package telemetry, cost extraction, + deterministic two-lane scheduling, endpoint recovery, exact-PID substance + monitoring, and server-slot progress watchdogs have isolated tests. +- [x] Explicit terminal outcomes retain receipt, transcript, cost, telemetry, + label, and evidence-window provenance without fabricating normal grades. +- [x] The one-hour transport-timeout refactor attempt remains preserved as + infrastructure-invalid, and its exact replacement reaches a model-native + stop condition under the corrected 14,400-second server timeout. +- [ ] The committed boundary bundle passes syntax, unit, ShellCheck, + preregistration, and comparison-source validation from a clean worktree. + +## 5. Canonical matrix + +- [ ] All 36 immutable N=3 cells are completed or explicitly terminal, graded + or labeled, telemetry-audited, and summarized without cherry-picking. +- [ ] The two pre-telemetry bug-fix attempts retain their original quality + outcomes and map to separately labeled, idle-start supplemental observations. +- [ ] Raw N=3 grader manifest, scorecard, evidence audit, invalid-attempt + inventory, raw telemetry hash, and any reproducible correction overlay are + complete. +- [ ] All 120 N=10 cells, including immutable v1-v3, meet the same gates and + have a separate variance-aware scorecard and audit. + +## 6. Extended suites + +- [ ] Single-PR audit N=3 uses the exact pinned PR #1057 base/head/squash and + original commits; evidence and substantive code-review audits pass. +- [ ] Investment memo N=3 archives receive formula, statement, unit, + traceability, valuation, and material-finance review. +- [ ] Board presentation N=3 derives from the matching memo replicate; every + deck is rendered and receives structural, visual, and claim-trace review. +- [ ] Frozen historical 75-PR N=3 uses the pinned baseline and PR-set hash and + receives strict coverage, traceability, repository, test, and substance + review. +- [ ] Every attempt and dependency failure is preserved; only affirmatively + proven infrastructure-invalid attempts are excluded and replaced. + +## 7. Cross-model interpretation + +- [x] Comparator source documents and historical operating points were pinned + before Gemma grades in `gemma4-comparison-sources.json`. +- [ ] Gemma N=3 and N=10 raw/corrected results are compared with Qwen3.6-27B, + Qwen3.6-35B-A3B, Qwen3-Coder-Next, Qwen3.5-397B-A17B, and DeepSeek V4 Flash + only on supported axes, with sample-size, engine, quantization, context, + sampling, power, date, grader, and task-level caveats. +- [ ] The final report distinguishes bounded-task quality, performance, + artifact-modality quality, and marathon-agent reliability; it makes no + unsupported global-SOTA claim. + +## 8. Publish, merge, and restore + +- [ ] Compact reports, manifests, hashes, validators, deployment files, + scorecards, correction overlays, and external-archive inventories comply with + `REPO-SPACE.md`. +- [ ] Code, evidence, link, secret, size, synthetic-merge, and security audits + pass; the worktree is clean. +- [ ] A draft PR is opened against current MMBT main, independently audited, + updated if necessary, and merged only when clean. +- [ ] Gemma campaign and telemetry services are stopped after evidence capture; + the proven DeepSeek launcher/config is restored byte-for-byte and Sanctuary + and Pixel pass final DeepSeek health checks. +- [ ] Only after every box above has authoritative evidence is the persistent + campaign goal marked complete. diff --git a/tooling/deployments/gemma4-31b-q4-tower2/PREREGISTRATION.md b/tooling/deployments/gemma4-31b-q4-tower2/PREREGISTRATION.md new file mode 100644 index 00000000..38933ee3 --- /dev/null +++ b/tooling/deployments/gemma4-31b-q4-tower2/PREREGISTRATION.md @@ -0,0 +1,196 @@ +# Gemma 4 31B QAT Q4_0 MMBT preregistration + +Status: frozen before the first Gemma inference on Tower2 + +Campaign branch: `gemma4-31b-q4-mmbt` + +MMBT base: `dcd9431d82168a17f039de084dce1a46ce3cc01a` + +## Question + +What quality, reliability, long-context behavior, latency, and concurrent +throughput does the official Gemma 4 31B instruction-tuned QAT Q4_0 checkpoint +deliver on Tower2 when it is configured for the best capability-preserving +operation this hardware can sustain? How does that evidence compare with the +published Qwen3.6-27B, Qwen3.6-35B-A3B, Qwen3-Coder-Next, Qwen3.5-397B, and +DeepSeek V4 Flash entries? + +The campaign must separate bounded-task benchmark strength from complex +artifact quality, long-horizon agent control, and serving performance. No one +aggregate is allowed to stand in for those distinct questions. + +## Immutable model and operating point + +- Source: `google/gemma-4-31B-it-qat-q4_0-gguf` +- Revision: `59dde24573e7e61570dba08b18a2e1fe246955ed` +- Text GGUF SHA-256: `179cfb99212709597eae5929112cfca677e1bbf566178b479ae1da0c4772874b` +- Multimodal projector SHA-256: `6bd60bdb958548b4093196d38744b0f2290c12503a3fddd7486bffa9c5eb07a4` +- Native context: 262,144 tokens per sequence +- Sampling: temperature 1.0, top-p 0.95, top-k 64 +- Hardware: two RTX PRO 6000 Blackwell Workstation Edition GPUs +- Persistent power limit: 500 W per GPU + +The text benchmark uses the text GGUF. The projector is pinned and receives a +separate image smoke test, but multimodal tokens are not injected into the +historical text-only MMBT cells. + +The request output cap is 262,144 for every canonical and extended suite. The +harness subtracts actual prompt tokens and a 14,000-token safety reserve, so a +request can never exceed the native total context. A length termination at the +remaining safe ceiling is a valid model-control outcome, not infrastructure +invalidity. + +## Serving bakeoff before model benchmarking + +The model is only about 17.65 GB, so using both GPUs for each token is not +assumed to be optimal. The four preregistered candidates are: + +1. One fully offloaded instance on GPU 0. +2. One instance split 1:1 across both GPUs in llama.cpp layer mode. +3. One instance split 1:1 across both GPUs in llama.cpp row mode. +4. Two independent fully offloaded instances, one per GPU, behind a + health-aware local router. + +The exact machine-readable matrix and selection rule are in +`topology-matrix.json`. Every candidate must use the same model hash, chat +template, sampling, context, and benchmark prompts. Runtime builds are pinned by +commit or image digest. Current llama.cpp source is evaluated because the +already-cached b9641 image predates the validated QAT upload; the cached image is +a fallback candidate, not an assumed winner. + +Mandatory gates precede speed ranking: valid chat formatting, native tool-call +round trip, 250K prompt acceptance, start/middle/end recall near 256K, no OOM, +stable restart, and repeatability. A candidate failing any gate is ineligible. +Among passing candidates, concurrency-8 aggregate decode wins; candidates within +10% are broken by single-request end-to-end latency and uncached 128K prefill, +then power, operational simplicity, and failover behavior. + +For independent replicas, `-np` and the total context pool are capacity search +variables. A configuration may expose multiple slots only when each advertised +slot can still receive a full 262,144-token sequence. Queueing is preferable to +silently reducing per-request context. + +## Production safety gate + +DeepSeek remains the production model while non-conflicting preparation occurs. +Before releasing its GPU allocation, Sanctuary and Pixel must each pass a real +request and a fallback route must be proven. During the controlled cutover, +availability and route state are recorded. Both agents must pass again after +Gemma load, after a deliberate Gemma restart, and after the benchmark campaign. +Failure restores the prior known-good route before optimization continues. + +## Canonical cohort + +The immutable first cohort is the standard 12 task families at N=3: 36 model +runs. It is scored and reported independently for direct comparison with the +historical N=3 local entries and DeepSeek. + +After v1-v3 are complete, every family is extended through v10: 120 total model +runs. The first three are not replaced or reselected. Full N=10 is chosen rather +than expanding only favorable or differential cells; it gives Gemma a direct +variance-aware comparison with the Qwen3.5-397B N=10 campaign while retaining +the historical Qwen3.6-27B N=3 and differential-cell comparisons. + +All generated grades are preserved. The raw score is always published. + +## Extended suites + +Each suite runs at N=3 with the full safe 256K request envelope: + +- DreamServer single-PR audit +- Wall Street investment memo +- Wall Street board presentation, using the corresponding memo replicate +- Frozen historical DreamServer 75-PR audit at baseline `d5154c3` + +The frozen PR-number set SHA-256 is +`569b95b3384af0c4ae4b54a2c8c8f7c908b396124777927a37b5c8fa0211ecd1`. +Current and historical task hashes are recorded in +`tooling/gemma4-31b-q4-extended-matrix.json`. + +The single-PR prompt still calls PR #1057 “open,” but the PR merged before the +DeepSeek and Gemma campaigns. `tooling/gemma4-single-pr-subject-pin.json` +anchors the unchanged task to the exact subject audited by all three DeepSeek +replicates: base `309e9cd0`, head `e5ceb43e`, squash contribution `1678f194`, +and the four original PR commits. Each Gemma artifact must identify and +reconcile those refs. Current `main` is optional context, not a replacement +subject; auditing a different diff is a model failure, while a proven inability +to fetch an immutable ref is preserved as infrastructure-invalid. + +## Attempt preservation and validity + +Every attempt receives a unique run name and remains in the campaign ledger. +No completed or unfavorable run is overwritten. + +An attempt may be marked infrastructure-invalid only with affirmative evidence +that the intended model was never evaluated, such as wrong model identity, +fixture/hash mismatch caught before inference, endpoint unavailable for more +than 90 seconds with no model response, or a harness launch failure before the +first request. A server-visible API error is not automatically excluded. OOM, +length termination, looping, malformed tool calls, missing deliverables, bad +formatting, and failure to call tools are valid outcomes unless a separate +reproduction proves an infrastructure defect. + +Infrastructure-invalid attempts are preserved, classified, and replaced until +the preregistered valid N is reached. Replacement policy never depends on score. + +## Grading and artifact audit + +- Original `grade.json` outputs are immutable. +- Corrections require a reproducible grader defect against an unchanged archive. +- Raw grades are copied to `grade.raw.json`; corrected grades and a cell-by-cell + overlay record archive hash, old/new verdict, reason, and grader commit. +- Corrections may remove generated caches, host-runtime dependence, or a + contradictory rubric rule. They may not reinterpret a merely weak answer. +- Workbooks are inspected for formulas, hard-coded calculations, consistency, + missing claimed sheets, and material finance defects. +- Presentations are rendered and visually inspected for overflow, clipping, + unreadable text, charts, citations, and unsupported claims. +- Agent-audit repositories are checked for required structure, traceability, + executable evidence, git history/tag, and substantive completion. + +## Telemetry and reporting + +Record exact model/runtime revisions and commands; prompt and template controls; +sampling; context and output caps; seeds; request IDs; prompt, cached-prompt, +reasoning, and completion tokens; finish reasons; TTFT/prefill/decode where the +server exposes them; wall time; queue time; concurrency; GPU memory, utilization, +power, temperature, and clocks; host power where available; restart/OOM state; +artifact hashes; and grader commits. + +Performance probes use uncached prompts or explicitly label prefix-cache hits. +Power telemetry is sampled at five-second cadence and clipped to request/run +windows. Performance comparisons do not mix cached and uncached measurements. + +## Comparison and publication rules + +Matched claims use common task cells, cohort sizes, and raw scoring where +available. Context, engine, sampling, quantization, and date differences are +shown beside every cross-model aggregate. Gemma may be described as the highest +observed result only if the matched evidence supports it; no global-SOTA claim +is inferred from MMBT. + +Compact audits, manifests, hashes, validators, deployment files, scorecards, and +limitations are committed. Oversized raw archives remain external under +`REPO-SPACE.md`, with hashes and inventory in git. Publication requires syntax, +unit, link, evidence, secret, size, and synthetic-merge audits; the draft PR is +merged only after those pass and production agents are healthy. + +## Append-only serving amendment: six-slot measurement + +At 2026-08-01T23:44:00Z, before any canonical or extended MMBT quality run, the +independent-replica serving search was extended from slots `[1, 2, 4]` to +`[1, 2, 4, 6]`. The measured four-slot Q8 candidate occupied 64,789 MiB of +97,887 MiB while maintaining four hard 262,144-token slots. Its observed +incremental KV footprint supports measuring a six-slot, 1,572,864-token pool +with a material VRAM safety margin. This amendment uses only serving memory and +throughput evidence, responds to the explicit full-utilization objective, and +does not discard or replace any prior topology attempt. The exact machine- +readable amendment and anti-cherry-pick note are in `topology-matrix.json`. + +## Stop conditions + +The campaign is complete only when the selected serving topology is stable, all +valid N=3 and N=10 canonical cells and every extended N=3 suite are audited, +the results and deployment rationale are merged into MMBT, and Sanctuary and +Pixel pass the final production checks. Time or a favorable early score is not a +stop condition. diff --git a/tooling/deployments/gemma4-31b-q4-tower2/README.md b/tooling/deployments/gemma4-31b-q4-tower2/README.md new file mode 100644 index 00000000..ca0b3fad --- /dev/null +++ b/tooling/deployments/gemma4-31b-q4-tower2/README.md @@ -0,0 +1,233 @@ +# Gemma 4 31B QAT Q4_0 on Tower2 + +This directory is the reproducibility package for the Gemma 4 31B Q4 campaign. +It is created before serving optimization so runtime choices cannot be selected +after seeing benchmark quality. + +Read in this order: + +1. `PREREGISTRATION.md` — immutable questions, cohorts, validity policy, and + topology selection rule. +2. `model-manifest.json` — exact local artifacts, hashes, upstream revision, + hardware, and runtime candidates. +3. `topology-matrix.json` — serving candidates, workloads, and winner rule. +4. `final-validation.json` — populated only after a serving candidate passes. +5. `campaign/` — launchers, units, telemetry, and benchmark supervision added as + the preregistered phases execute. + +The source is Google's [official QAT Q4_0 GGUF](https://huggingface.co/google/gemma-4-31B-it-qat-q4_0-gguf). +Google documents [llama.cpp as an OpenAI-compatible Gemma serving route](https://ai.google.dev/gemma/docs/integrations/llamacpp), +and the model repository pins the 256K context and temperature 1.0 / top-p 0.95 / +top-k 64 operating point. + +This package deliberately tests one model per GPU as well as cross-GPU splitting. +The 17 GB text model fits on either 98 GB GPU; using both GPUs for every request +would add communication without necessarily improving latency or aggregate +throughput. + +The campaign runtime is temporary. `production-restore.json` fingerprints the +proven DeepSeek launcher, image, OpenClaw routing, and a permission-restricted +pre-campaign config backup. Campaign completion requires stopping Gemma, +restoring that DeepSeek state byte-for-byte, and passing fresh Sanctuary and +Pixel end-to-end checks. + +Serving files +------------- + +- `run-gemma4-server.sh` is the foreground server launcher. It refuses any + context pool that would leave a parallel slot with less than 262,144 tokens, + and refuses a server read/write timeout below 14,400 seconds. The four-hour + floor is long enough to reach the native output boundary at the slowest + long-generation decode rate observed during canonical N=3. +- `mmbt-gemma4@.service` and `topologies/*.env` define the preregistered serving + candidates. +- `gemma4-topology-control.sh` enforces mutually exclusive topology candidates + and refuses to start while the DeepSeek container still owns the GPUs. +- Split-GPU builds use the privately pinned NCCL runtime recorded in + `model-manifest.json`; no system CUDA or NCCL package is modified. + +Validated benchmark serving +--------------------------- + +The winning quality-preserving topology is two independent Q8-KV, four-slot +replicas: GPU 0 on port 8000 and GPU 1 on port 8001. Each request stays on one +GPU and retains a hard 262,144-token context; using both replicas concurrently +increases campaign throughput without introducing a cross-GPU dependency. + +- `benchmark-serving-manifest.json` is the compact immutable receipt anchor for + the artifact, runtime, topology, sampling point, context, output ceiling, and + 500 W power caps. +- `ensure-gemma4-winner.sh` starts and health-checks both systemd services; the + benchmark supervisor uses it for recovery rather than assuming a container. +- `tooling/gemma4-31b-q4-mmbt.json` sends temperature 1.0, top-p 0.95, top-k 64, + and the 262,144-token output ceiling to the harness. Work is assigned by + stable run ordinal modulo two, so the lanes are disjoint and reproducible. +- Every run receipt captures the deployment manifest, live `/v1/models` + identity, and the exact host `llama-server` path, arguments, and SHA-256. +- The supervisor watches each port independently and treats advancing + llama.cpp `/slots` prompt/decode counters as live progress. This matters at + the native 262K output envelope: a non-streamed HTTP call can decode for more + than the historical transcript-staleness timeout without appending a JSONL + row, and must not be misclassified as a hung run. The five-minute substance + monitor is likewise keyed to every live lane and exact harness PID, so work + on one replica cannot mask an identical-call loop on the other. +- `gemma4_gpu_telemetry.py` samples both GPUs and the CPU package every five + seconds. It attributes port 8000's active harness to GPU 0 and port 8001's to + GPU 1, so simultaneous cells are not mislabeled as one dual-GPU request. +- `analyze_replica_telemetry.py` clips samples to authoritative evidence + windows and records their exact source. Normally that is `summary.json`; an + explicit terminal outcome instead uses the preserved receipt/transcript + start and label/transcript end. It writes per-run power, utilization, memory, + temperature, clock, energy, CPU-package context, coverage, and concurrent-lane + evidence for both completed workspaces and terminal outcomes. AC wall draw is + explicitly unavailable to software and is not fabricated from component data. +- The telemetry sidecar derives `cost.json` and `gpu_telemetry.json` for an + explicit terminal label as well as a normal summary. A model-control failure + therefore cannot evade resource accounting merely because it never produced + a final workspace archive. +- `snapshot_gemma4_telemetry.sh` fail-closes cohort-boundary snapshots. It + refuses an active benchmark or overwrite, briefly stops the telemetry writer + and sidecar, validates that the copied CSV ends on a complete row with the + pinned schema, prints its SHA-256, and restores exactly the services that + were active. N=3 and N=10 audits point at immutable snapshots rather than + hashing a live CSV that later cohorts will append to. + +The two-lane optimization changes scheduling only. It does not change the +canonical run names, cohort membership, prompts, grading, or attempt-preservation +rules, and v1-v3 remain the immutable first cohort before expansion through v10. + +Canonical refactor v3 exposed that the initial llama.cpp default +`--timeout 3600` was shorter than the model's native envelope: the final API +call was cancelled at exactly one hour after 133,606 decoded tokens, while the +slot remained untruncated at 151,966 total tokens. The attempt is preserved in +`logs/_invalid/` and externally, and is excluded as infrastructure-invalid +independently of its score. Its prolonged post-work generation remains valid +operational reliability evidence. The exact cell is replaced under the same +prompt, model, sampling, context, and output controls after raising only the +server transport timeout to 14,400 seconds. See +`canonical-refactor-timeout-incident-20260802.json`. + +Canonical bug-fix v1/v2 completed before the telemetry sidecar was introduced. +After N=3 stops, `tooling/run_gemma4_supplemental_telemetry.sh` produces exactly +two separately labeled, one-per-GPU observations with identical operating +controls. They never replace or enter the quality cohort. The fail-closed audit +requires explicit canonical-to-supplement mappings and validates each +supplement's complete receipt, archive, cost, and telemetry evidence. The +supplement launcher also requires both replica slot pools to be idle at launch, +so an already-running Sanctuary or Pixel request cannot silently contaminate +the matched observation. + +At the clean N=3 boundary, the committed tools run in this order: + +```bash +bash tooling/run_gemma4_supplemental_telemetry.sh +bash tooling/snapshot_gemma4_telemetry.sh \ + /home/michael/gemma4-campaign-state/telemetry/snapshots/canonical-n3-plus-supplement.csv + +python3 tooling/audit_gemma4_campaign.py \ + --target-n 3 --require-grades \ + --allow-pretelemetry-run p1_bugfix_gemma4-31b-q4_v1 \ + --allow-pretelemetry-run p1_bugfix_gemma4-31b-q4_v2 \ + --pretelemetry-supplement \ + p1_bugfix_gemma4-31b-q4_v1=p1_bugfix_gemma4-31b-q4-telemetry-supplement_v1 \ + --pretelemetry-supplement \ + p1_bugfix_gemma4-31b-q4_v2=p1_bugfix_gemma4-31b-q4-telemetry-supplement_v2 \ + --raw-telemetry \ + /home/michael/gemma4-campaign-state/telemetry/snapshots/canonical-n3-plus-supplement.csv \ + --output logs/_campaign_audit/gemma4-canonical-n3-evidence-audit.json + +python3 tooling/capture_gemma4_grader_manifest.py \ + --target-n 3 \ + --output logs/_campaign_audit/gemma4-canonical-n3-grader-manifest.json +python3 tooling/summarize_gemma4_campaign.py \ + --target-n 3 \ + --allow-pretelemetry-run p1_bugfix_gemma4-31b-q4_v1 \ + --allow-pretelemetry-run p1_bugfix_gemma4-31b-q4_v2 \ + --output-json logs/_campaign_audit/gemma4-canonical-n3-scorecard.json \ + --output-markdown logs/_campaign_audit/gemma4-canonical-n3-scorecard.md + +python3 tooling/correct_gemma4_project_mgmt_grades.py \ + --target-n 3 \ + --output logs/_campaign_audit/gemma4-canonical-n3-project-mgmt-correction.json +``` + +N=10 repeats the snapshot, audit, grader-manifest, and scorecard commands with +`--target-n 10` and a distinct immutable telemetry snapshot. Extended suites +start only after that second boundary passes. + +`tooling/summarize_gemma4_campaign.py` generates the raw N=3 and N=10 JSON and +Markdown scorecards only when every expected completed workspace or explicit +terminal label has its required grade/label, cost, and telemetry evidence. It +keeps `done_signal`, raw PASS/STRUCTURAL_PASS, explicit terminal non-pass +outcomes, model-call throughput, total wall time, and per-replica telemetry as +separate axes; it never invents a normal grade for a terminal outcome, and +reproducible grader corrections remain separate hash-tied overlays. +`tooling/capture_gemma4_grader_manifest.py` fingerprints the grading driver, +all task graders and ground-truth inputs, the sandbox image, runtime versions, +repository commit, and every raw grade/terminal-label file before any overlay +is considered. + +The raw project-management grader searches for a few contiguous phrases and +misses semantically exact wording such as “Maevia … push back,” “Legal has not +yet responded,” hyphenated “web-responsive,” and “private beta (3-5 +customers).” `tooling/correct_gemma4_project_mgmt_grades.py` is a narrow, +non-destructive overlay fixed after N=3 and before N=10. It records the raw +grade, report, archive, raw-grader, and correction-script hashes; changes only +those four lexical false negatives; and never overwrites `grade.json`. Raw +scores remain primary and the corrected total is reported separately. No other +failure is reinterpreted without an independently reproducible grader defect. + +`invalid-attempt-classifications.json` is the exclusion ledger. The canonical +auditor requires every `logs/_invalid/` directory to have a matching +score-independent infrastructure classification, affirmative evidence, +hash-pinned preserved files, an incident document, and a completed exact +canonical replacement. An unknown, altered, or merely pending exclusion makes +the cohort audit fail. + +`tooling/gemma4-comparison-sources.json` freezes the pre-result evidence for +Qwen3.6-27B, Qwen3.6-35B-A3B, Qwen3-Coder-Next, Qwen3.5-397B-A17B, and +DeepSeek V4 Flash. `tooling/validate_gemma4_comparison_sources.py` verifies +every source hash and checks the extracted DeepSeek/Qwen397 raw and corrected +totals plus the historical serving controls. Qwen3.6-35B-A3B is deliberately +recorded as lacking a comparable canonical cohort rather than being assigned a +made-up rank. This pin happens before Gemma grades are known. + +`tooling/analyze_gemma4_queueing.py` derives the accepted four-slot queueing +result from the preserved simultaneous-eight-request probe. It separates the +four shortest first-wave requests from the four queued requests, subtracts +llama.cpp-reported prompt-plus-decode work from client wall time, fingerprints +the raw source, and labels the resulting queue-wait delta as an estimate rather +than a direct server timestamp. + +The committed `queueing-analysis.json` reports 112.6166 aggregate decode +tokens/s at eight simultaneous requests. With four slots, the second wave's +median client wall was 9.0752 seconds longer and the derived queue-wait delta +was 9.0921 seconds. The source is the hash-pinned uncached topology probe; the +queue delta remains explicitly an estimate because llama.cpp did not expose a +direct per-request queue-start timestamp. + +Extended-suite scheduling +------------------------- + +`tooling/run_gemma4_extended_suites.py` applies the same two-lane design to the +four DeepSeek-comparable extended suites. `mmbt-gemma4-extended@.service` pins +lane 0 to port 8000 and lane 1 to port 8001. Jobs are assigned by immutable +suite/replicate ordinal, and board-presentation jobs wait for their matching +investment-memo workspace even when the dependency is running on the other +lane. Any Python traceback stops that lane for operator review; endpoint outages +are the only automatically replaceable attempts. + +`tooling/audit_gemma4_extended.py` then fail-closes the 12-run matrix over +model/runtime identity, pinned task hashes, deterministic lane assignment, +sampling/context/output controls, 500 W receipts, telemetry, dependency +lineage, frozen-fixture mounting, archive hashes, and preserved invalid +attempts. The one-PR arm additionally requires every artifact to identify the +base, head, squash contribution, and original commits pinned in +`gemma4-single-pr-subject-pin.json`, matching the three DeepSeek comparator +archives despite the prompt's stale “open PR” wording. Its result is +deliberately an evidence audit, not a quality verdict: +the one-PR and 75-PR repositories, investment workbooks, and rendered board +decks still receive their separate substantive code, finance, and visual +overlays before publication. Those overlay rules are fixed before launch in +`tooling/GEMMA4-EXTENDED-SUBSTANTIVE-AUDIT-PROTOCOL.md`; its SHA-256 is pinned +by the extended matrix so completed artifacts cannot change the rubric. diff --git a/tooling/deployments/gemma4-31b-q4-tower2/analyze_replica_telemetry.py b/tooling/deployments/gemma4-31b-q4-tower2/analyze_replica_telemetry.py new file mode 100755 index 00000000..54cb1053 --- /dev/null +++ b/tooling/deployments/gemma4-31b-q4-tower2/analyze_replica_telemetry.py @@ -0,0 +1,297 @@ +#!/usr/bin/env python3 +"""Derive per-run telemetry for Gemma's independent one-GPU replicas.""" +from __future__ import annotations + +import argparse +import csv +import hashlib +import json +import statistics +import subprocess +from collections import Counter, defaultdict +from datetime import datetime, timezone +from pathlib import Path + + +def parse_time(value: str) -> float: + return datetime.fromisoformat(value.replace("Z", "+00:00")).timestamp() + + +def rounded(value, digits=4): + return None if value is None else round(value, digits) + + +def percentile(values: list[float], q: float): + if not values: + return None + ordered = sorted(values) + return ordered[min(len(ordered) - 1, int(q * (len(ordered) - 1)))] + + +def file_sha256(path: Path) -> str: + digest = hashlib.sha256() + with path.open("rb") as handle: + for chunk in iter(lambda: handle.read(65536), b""): + digest.update(chunk) + return digest.hexdigest() + + +def provenance() -> dict: + script = Path(__file__).resolve() + repo = script.parent.parent.parent.parent + git_sha = subprocess.run( + ["git", "-C", str(repo), "rev-parse", "HEAD"], + capture_output=True, text=True, check=False, + ).stdout.strip() or None + return {"path": str(script), "file_sha256": file_sha256(script), "git_sha": git_sha} + + +def load_rows(path: Path) -> tuple[list[dict], dict[float, dict[str, dict]]]: + rows = [] + by_time: dict[float, dict[str, dict]] = defaultdict(dict) + with path.open(newline="") as handle: + for raw in csv.DictReader(handle): + try: + row = { + "ts": parse_time(raw["ts"]), + "gpu": str(int(raw["gpu"])), + "port": int(raw["endpoint_port"]), + "power_limit_w": float(raw["power_limit_w"]), + "power_w": float(raw["power_w"]), + "util_sm": float(raw["util_sm"]), + "util_mem": float(raw["util_mem"]), + "mem_used_mib": float(raw["mem_used_mib"]), + "temp_c": float(raw["temp_c"]), + "sm_clk_mhz": float(raw["sm_clk_mhz"]), + "cell": raw.get("cell", "").strip(), + "harness_pid": int(raw["harness_pid"]) if raw.get("harness_pid") else None, + "cpu_package_power_w": float(raw["cpu_package_power_w"]) if raw.get("cpu_package_power_w") else None, + } + except (KeyError, TypeError, ValueError): + continue + rows.append(row) + by_time[row["ts"]][row["gpu"]] = row + return rows, by_time + + +def transcript_bounds(run_dir: Path) -> tuple[float, float] | None: + timestamps = [] + try: + for line in (run_dir / "transcript.jsonl").read_text().splitlines(): + if not line.strip(): + continue + row = json.loads(line) + if row.get("t"): + timestamps.append(parse_time(row["t"])) + except (OSError, TypeError, ValueError, json.JSONDecodeError): + return None + return (min(timestamps), max(timestamps)) if timestamps else None + + +def load_window(run_dir: Path) -> tuple[float, float, float, str] | None: + try: + summary = json.loads((run_dir / "summary.json").read_text()) + start = parse_time(summary["started_at"]) + end = parse_time(summary["ended_at"]) + wall = float(summary["elapsed_s"]) + return (start, end, wall, "summary.json") if end > start and wall > 0 else None + except (OSError, KeyError, TypeError, ValueError, json.JSONDecodeError): + pass + + label = None + receipt = None + try: + label = json.loads((run_dir / "label.json").read_text()) + except (OSError, TypeError, ValueError, json.JSONDecodeError): + pass + if not isinstance(label, dict) or not label.get("primary"): + return None + try: + receipt = json.loads((run_dir / "receipt.json").read_text()) + except (OSError, TypeError, ValueError, json.JSONDecodeError): + pass + bounds = transcript_bounds(run_dir) + + start = None + start_source = None + if isinstance(receipt, dict) and receipt.get("captured_at"): + try: + start = parse_time(receipt["captured_at"]) + start_source = "receipt.json:captured_at" + except (TypeError, ValueError): + pass + if start is None and bounds: + start = bounds[0] + start_source = "transcript.jsonl:first_t" + + end = None + end_source = None + if label.get("labeled_at"): + try: + end = parse_time(label["labeled_at"]) + end_source = "label.json:labeled_at" + except (TypeError, ValueError): + pass + if end is None and bounds: + end = bounds[1] + end_source = "transcript.jsonl:last_t" + + if start is None or end is None or end <= start: + return None + return start, end, end - start, f"{start_source}..{end_source}" + + +def integrate(samples: list[dict], key: str, end: float, max_gap: float) -> tuple[float, float]: + energy_ws = 0.0 + covered = 0.0 + for current, following in zip(samples, samples[1:]): + dt = min(following["ts"], end) - current["ts"] + value = current.get(key) + if value is not None and 0 < dt <= max_gap: + energy_ws += value * dt + covered += dt + return energy_ws, covered + + +def find_run(cell: str, logs_dirs: list[Path]) -> Path | None: + for logs_dir in logs_dirs: + candidate = logs_dir / cell + if candidate.is_dir(): + return candidate + return None + + +def analyze(csv_path: Path, logs_dirs: list[Path], cap_per_gpu: float, + rate: float, max_gap: float) -> dict[str, dict]: + rows, by_time = load_rows(csv_path) + by_cell: dict[str, list[dict]] = defaultdict(list) + for row in rows: + if row["cell"]: + by_cell[row["cell"]].append(row) + result = {} + analyzer = provenance() + for cell, all_samples in sorted(by_cell.items()): + run_dir = find_run(cell, logs_dirs) + window = load_window(run_dir) if run_dir else None + if not window: + continue + start, end, wall, window_source = window + samples = sorted((row for row in all_samples if start <= row["ts"] <= end), key=lambda row: row["ts"]) + if not samples: + continue + gpus = sorted({row["gpu"] for row in samples}) + active_energy_ws, covered = integrate(samples, "power_w", end, max_gap) + cpu_samples = [row for row in samples if row["cpu_package_power_w"] is not None] + cpu_energy_ws, cpu_covered = integrate(cpu_samples, "cpu_package_power_w", end, max_gap) + host_component_values = [] + other_cells = Counter() + simultaneous_decode = 0 + paired = 0 + for row in samples: + pair = by_time.get(row["ts"], {}) + if "0" in pair and "1" in pair: + paired += 1 + cpu = row["cpu_package_power_w"] or 0.0 + host_component_values.append(pair["0"]["power_w"] + pair["1"]["power_w"] + cpu) + other_gpu = "1" if row["gpu"] == "0" else "0" + other_cell = pair[other_gpu].get("cell") or "(idle/no benchmark harness)" + other_cells[other_cell] += 1 + if pair["0"]["util_sm"] > 20 and pair["1"]["util_sm"] > 20: + simultaneous_decode += 1 + active_power = [row["power_w"] for row in samples] + cpu_power = [row["cpu_package_power_w"] for row in cpu_samples] + active_kwh = active_energy_ws / 3_600_000.0 + cpu_kwh = cpu_energy_ws / 3_600_000.0 + result[cell] = { + "schema_version": 1, + "run_name": cell, + "source_csv": str(csv_path), + "generated_at": datetime.now(timezone.utc).isoformat(), + "analyzer": analyzer, + "attribution": { + "method": "live harness --port maps port 8000 to GPU 0 and port 8001 to GPU 1", + "active_gpu_ids_observed": gpus, + "other_gpu_work_is_reported_as_concurrency_not_charged_to_active_gpu_energy": True, + "cpu_package_power_is_shared_host_context_and_not_uniquely_attributable_when_runs_overlap": True + }, + "sampling": { + "samples": len(samples), + "paired_host_samples": paired, + "first_sample": datetime.fromtimestamp(samples[0]["ts"], timezone.utc).isoformat(), + "last_sample": datetime.fromtimestamp(samples[-1]["ts"], timezone.utc).isoformat(), + "window_source": window_source, + "window_started_at": datetime.fromtimestamp(start, timezone.utc).isoformat(), + "window_ended_at": datetime.fromtimestamp(end, timezone.utc).isoformat(), + "window_wall_s": rounded(wall, 1), + "integrated_coverage_s": rounded(covered, 1), + "coverage_fraction_of_wall": rounded(min(1.0, covered / wall), 4), + "max_accepted_gap_s": max_gap + }, + "active_gpu": { + "mean_power_w": rounded(statistics.mean(active_power), 2), + "time_weighted_power_w": rounded(active_energy_ws / covered if covered else None, 2), + "p90_power_w": rounded(percentile(active_power, 0.90), 2), + "max_power_w": rounded(max(active_power), 2), + "mean_sm_util_pct": rounded(statistics.mean(row["util_sm"] for row in samples), 2), + "p90_sm_util_pct": rounded(percentile([row["util_sm"] for row in samples], 0.90), 2), + "max_memory_used_mib": rounded(max(row["mem_used_mib"] for row in samples), 1), + "max_temp_c": rounded(max(row["temp_c"] for row in samples), 1), + "mean_sm_clock_mhz": rounded(statistics.mean(row["sm_clk_mhz"] for row in samples), 1), + "configured_cap_w": cap_per_gpu + }, + "cpu_package_shared_context": { + "samples": len(cpu_samples), + "mean_power_w": rounded(statistics.mean(cpu_power), 2) if cpu_power else None, + "max_power_w": rounded(max(cpu_power), 2) if cpu_power else None, + "integrated_coverage_s": rounded(cpu_covered, 1) + }, + "concurrency": { + "other_gpu_cell_sample_counts": dict(sorted(other_cells.items())), + "both_gpus_over_20pct_sm_fraction": rounded(simultaneous_decode / paired, 4) if paired else None, + "mean_observed_two_gpu_plus_cpu_package_w": rounded(statistics.mean(host_component_values), 2) if host_component_values else None, + "max_observed_two_gpu_plus_cpu_package_w": rounded(max(host_component_values), 2) if host_component_values else None, + "wall_power_note": "This is GPU plus CPU package telemetry, not AC wall draw; no software wall meter is available." + }, + "energy_and_cost": { + "active_gpu_sampled_kwh": rounded(active_kwh, 6), + "active_gpu_sampled_cost_usd": rounded(active_kwh * rate, 6), + "cpu_package_shared_sampled_kwh": rounded(cpu_kwh, 6), + "active_gpu_cap_upper_bound_kwh": rounded(cap_per_gpu * wall / 3_600_000.0, 6), + "rate_usd_per_kwh": rate + } + } + return result + + +def main() -> None: + parser = argparse.ArgumentParser() + parser.add_argument("--csv", required=True, type=Path) + parser.add_argument("--logs-dir", action="append", default=[], type=Path) + parser.add_argument("--cap-per-gpu", type=float, default=500.0) + parser.add_argument("--rate-usd-per-kwh", type=float, default=0.13) + parser.add_argument("--max-gap-s", type=float, default=15.0) + parser.add_argument("--write-run-artifacts", action="store_true") + parser.add_argument("--output", type=Path) + args = parser.parse_args() + logs_dirs = [path.resolve() for path in args.logs_dir] + report = analyze(args.csv.resolve(), logs_dirs, args.cap_per_gpu, + args.rate_usd_per_kwh, args.max_gap_s) + if args.write_run_artifacts: + for cell, document in report.items(): + run_dir = find_run(cell, logs_dirs) + if run_dir and ((run_dir / "summary.json").exists() or (run_dir / "label.json").exists()): + target = run_dir / "gpu_telemetry.json" + temporary = target.with_suffix(".tmp") + temporary.write_text(json.dumps(document, indent=2) + "\n") + temporary.replace(target) + rendered = json.dumps(report, indent=2) + "\n" + if args.output: + temporary = args.output.with_suffix(args.output.suffix + ".tmp") + temporary.write_text(rendered) + temporary.replace(args.output) + else: + print(rendered, end="") + + +if __name__ == "__main__": + main() diff --git a/tooling/deployments/gemma4-31b-q4-tower2/benchmark-serving-manifest.json b/tooling/deployments/gemma4-31b-q4-tower2/benchmark-serving-manifest.json new file mode 100644 index 00000000..f8b30bfd --- /dev/null +++ b/tooling/deployments/gemma4-31b-q4-tower2/benchmark-serving-manifest.json @@ -0,0 +1,92 @@ +{ + "schema_version": 1, + "campaign": "gemma4-31b-q4-mmbt", + "served_name": "Gemma-4-31B-it-QAT-Q4_0", + "artifact": { + "path": "/mnt/bulk/models/google-gemma-4-31B-it-QAT-Q4_0-GGUF/gemma-4-31B_q4_0-it.gguf", + "sha256": "179cfb99212709597eae5929112cfca677e1bbf566178b479ae1da0c4772874b", + "source_revision": "59dde24573e7e61570dba08b18a2e1fe246955ed" + }, + "multimodal_projector": { + "path": "/mnt/bulk/models/google-gemma-4-31B-it-QAT-Q4_0-GGUF/gemma-4-31B-it-mmproj.gguf", + "sha256": "6bd60bdb958548b4093196d38744b0f2290c12503a3fddd7486bffa9c5eb07a4" + }, + "runtime": { + "kind": "host-native llama.cpp", + "commit": "11924d4c17abc27383376a1ac6a24fa3e36c1c0c", + "version": 10223, + "llama_server_sha256": "200b403b5735418ff1f6da0cea1938e413e11869ae362e8044a12b0df04622fc", + "cuda": "13.1.115", + "cuda_architecture": "120a-real" + }, + "topology": { + "kind": "two independent full-offload replicas", + "lane_ports": [8000, 8001], + "services": [ + "mmbt-gemma4@replica-gpu0-q8-s4.service", + "mmbt-gemma4@replica-gpu1-q8-s4.service" + ], + "physical_gpu_per_lane": [0, 1], + "slots_per_replica": 4, + "context_tokens_per_slot": 262144, + "context_pool_tokens_per_replica": 1048576, + "cache_type_k": "q8_0", + "cache_type_v": "q8_0", + "kv_unified": false, + "flash_attention": true, + "gpu_power_limit_w_each": 500 + }, + "benchmark_operating_point": { + "temperature": 1.0, + "top_p": 0.95, + "top_k": 64, + "max_model_len": 262144, + "max_output_tokens_cap": 262144, + "server_read_write_timeout_seconds": 14400, + "prompt_growth_reserve_tokens": 12000, + "context_safety_tokens": 2048, + "minimum_estimated_prompt_tokens": 8000, + "minimum_request_max_tokens": 2048 + }, + "amendments": [ + { + "effective_at": "2026-08-02T02:12:03Z", + "field": "benchmark_operating_point.server_read_write_timeout_seconds", + "previous_value": 3600, + "current_value": 14400, + "reason": "the prior transport timeout cancelled a live native-envelope request below 262144 tokens", + "incident": "canonical-refactor-timeout-incident-20260802.json", + "quality_controls_changed": false + } + ], + "telemetry": { + "interval_seconds": 5, + "raw_csv": "/home/michael/gemma4-campaign-state/telemetry/gemma4-31b-q4-gpu.csv", + "gpu_fields": [ + "power_limit_w", + "power_w", + "util_sm", + "util_mem", + "mem_used_mib", + "temp_c", + "sm_clk_mhz" + ], + "cpu_field": "RAPL package power derived from energy_uj", + "attribution": "live harness port 8000 maps to GPU 0; port 8001 maps to GPU 1", + "wall_power": "unavailable to software; GPU plus CPU component telemetry is not labeled as AC wall draw", + "services": [ + "mmbt-gemma4-power-logger.service", + "mmbt-gemma4-telemetry-sidecar.service" + ] + }, + "evidence_references": [ + { + "path": "model-manifest.json", + "sha256": "9208a912a473c94c0d016700025fd3576777c1b864ae381111f4ceb6695e1796" + }, + { + "path": "final-validation.json", + "sha256": "295945748ca82fd55fb768368aa1f5b75827411bf33dfe21c7c7adfd257883a8" + } + ] +} diff --git a/tooling/deployments/gemma4-31b-q4-tower2/canonical-harness-incident-20260802.json b/tooling/deployments/gemma4-31b-q4-tower2/canonical-harness-incident-20260802.json new file mode 100644 index 00000000..c6a83179 --- /dev/null +++ b/tooling/deployments/gemma4-31b-q4-tower2/canonical-harness-incident-20260802.json @@ -0,0 +1,36 @@ +{ + "schema_version": 1, + "detected_at": "2026-08-02T00:33:40Z", + "campaign": "gemma4-31b-q4-mmbt", + "classification": "harness-invalid attempt requiring preserved replacement", + "trigger_attempt": "p1_testwrite_gemma4-31b-q4_v1", + "trigger_evidence": { + "harness_commit": "7a1938178dbdde521151e1b1d392fec32d55434c", + "exception": "KeyError: 'path'", + "condition": "a syntactically valid tool call omitted the required path argument", + "effect": "the harness exited nonzero instead of returning a tool error and allowing model recovery", + "receipt_sha256": "92313dc9d6070ad4fba78fb89a4d043281e17d0e5fb2203e0bdc660c551669e4", + "transcript_sha256": "ca8b0f10985799cbd7511a4ce32d6b51939f6d6f32139719cb41b9c52d18aeb4", + "lane_log_sha256": "7e2176880f46874af26b1308ba5b5ac0d3d34b9c6adde5bd819111d211c3497b", + "supervisor_log_sha256": "e4db2a98b5cd9a14cb0c37d64a1144fb878114d7d5ae95439714db8b06adb51e" + }, + "operator_action": { + "action": "stopped the two-lane canonical supervisor immediately after detection", + "reason": "prevent additional cells from being evaluated under the defective error path", + "completed_valid_cells_retained": [ + "p1_bugfix_gemma4-31b-q4_v1", + "p1_bugfix_gemma4-31b-q4_v2" + ], + "incomplete_attempts_to_preserve_and_replace": [ + "p1_bugfix_gemma4-31b-q4_v3", + "p1_testwrite_gemma4-31b-q4_v1", + "p1_testwrite_gemma4-31b-q4_v3" + ] + }, + "correction": { + "scope": "validate required tool arguments and malformed types before dispatch", + "new_behavior": "return deterministic TOOL_ERROR text to the model; continue the same attempt", + "completed_cell_equivalence": "retained completed cells never encountered this error path, so their prompts, inference, tools, termination, and artifacts are unaffected", + "anti_cherry_pick": "replacement is determined solely by completeness and the recorded harness exception, before any affected cell was graded" + } +} diff --git a/tooling/deployments/gemma4-31b-q4-tower2/canonical-refactor-timeout-incident-20260802.json b/tooling/deployments/gemma4-31b-q4-tower2/canonical-refactor-timeout-incident-20260802.json new file mode 100644 index 00000000..b5615ce0 --- /dev/null +++ b/tooling/deployments/gemma4-31b-q4-tower2/canonical-refactor-timeout-incident-20260802.json @@ -0,0 +1,91 @@ +{ + "schema_version": 1, + "campaign": "gemma4-31b-q4-mmbt", + "classified_at": "2026-08-02T02:12:03Z", + "classification": "infrastructure-invalid canonical attempt with valid operational model-control evidence", + "run_name": "p1_refactor_gemma4-31b-q4_v3", + "score_independence": "classified and preserved before campaign grading because the server cancelled at its configured transport timeout", + "attempt": { + "started_at": "2026-08-02T01:05:37.203845Z", + "ended_at": "2026-08-02T02:08:59.029658Z", + "elapsed_s": 3801.8, + "iterations": 46, + "summary_finish_reason": "api_error: timed out", + "summary_completion_tokens_excluding_cancelled_call": 10283, + "summary_prompt_tokens": 489088, + "final_server_task_id": 80497, + "final_call_started_at": "2026-08-02T01:08:59Z", + "final_call_cancelled_at": "2026-08-02T02:08:59Z", + "configured_server_timeout_s": 3600, + "server_decoded_tokens_at_cancel": 133606, + "server_slot_total_tokens_at_cancel": 151966, + "server_truncated": false + }, + "validity_reason": { + "infrastructure_defect": "llama.cpp cancelled the live request at exactly the configured 3600-second read/write timeout while the slot was still decoding below the 262144-token native context/output boundary", + "quality_treatment": "the timeout-truncated attempt is not a canonical model-quality outcome and must be replaced without consulting its grade", + "operational_treatment": "the model's continued generation after completing the requested repository changes remains evidence of weak termination discipline" + }, + "preservation": { + "git_visible_path": "logs/_invalid/p1_refactor_gemma4-31b-q4_v3-server-timeout-20260802T020859Z", + "external_root": "/home/michael/gemma4-campaign-state/canonical-invalid/p1_refactor_gemma4-31b-q4_v3-server-timeout-20260802T020859Z", + "files": { + "cost.json": "c5d1062f477758436ba0c3baf32df97fbe4e30df09c831bde470ccf938fce7f7", + "gpu_telemetry.json": "4f717a0e9d7298401c41e1799f772186a03a25bdf21bee93b30425cab3b31494", + "gpu_telemetry.reanalyzed-fdcc2496.json": "51eff413ad7a1a01bc6116a52e31ac3c39a1f844bf0aa0f0e09d9a0731392801", + "grade.json": "7f69334378ac3187de489eb74ad81c8245267d1a5711612b94879f297e758bcf", + "receipt.json": "58b9ee02a9c9cd03829a91d8f2f4d6d5eedf064f1c96206bc6c4b38522c4bfd8", + "summary.json": "0445d7ff95e0fbfbbc06d21cb3fdcccf8a741249ea294dda3f136a3ea4ed2b56", + "transcript.jsonl": "e887af90ea26038d3d8aadf22998eaa443eeb78eaa59d204bc7fcc741124e00e", + "workspace_final.tar.gz": "90a1b1918640b2a496b66ed5fbaf8649eb5902aa2e4e36c1f6b4debc71c720bc", + "server-task-80497.log": "ce70ea512597a14e8262e1300989243f1ca6bdfeee04a000a6a2cd69da4d8a45" + } + }, + "correction": { + "server_read_write_timeout_s": 14400, + "changed_control": "transport timeout only", + "unchanged_controls": [ + "model and projector bytes", + "llama.cpp binary", + "one-GPU full-offload topology", + "Q8 KV cache and four slots", + "262144-token per-slot context", + "262144-token benchmark output ceiling", + "temperature 1.0, top_p 0.95, top_k 64", + "500 W GPU cap", + "task prompt and fixture", + "canonical run name" + ], + "replacement_policy": "preserve the invalid attempt, remove only its active canonical copy after the N=3 supervisor stops, and rerun the exact cell before grading/audit acceptance" + }, + "post_classification_grade": { + "verdict": "PASS", + "treatment": "preserved with the invalid attempt but excluded; the infrastructure classification and replacement decision were recorded before this grade existed" + }, + "post_boundary_reanalysis": { + "git_sha": "fdcc24964bdde6db7c602bacb3a53b8780229985", + "treatment": "the original telemetry is immutable; a separately named reanalysis adds explicit evidence-window provenance without replacing it" + }, + "replacement_result": { + "started_at": "2026-08-02T02:40:35.025584Z", + "ended_at": "2026-08-02T02:43:53.245183Z", + "elapsed_s": 198.2, + "iterations": 40, + "completion_tokens": 10797, + "prompt_tokens_cumulative": 399232, + "finish_reason": "done_signal", + "server_read_write_timeout_s": 14400, + "harness_git_sha": "1a4a954715bf7b228a2d5496fa87547318be6ba8", + "harness_git_dirty": false, + "grade_verdict": "PASS", + "files": { + "receipt.json": "d8b8de86c6abb82b39c59fd2ed56571dc1c05e4c1cf2214fce1a54803bb20dca", + "summary.json": "e22036d8e57d95f93243ae7cd3f22cfc271686c8134c715dfc44b7459410fe61", + "transcript.jsonl": "5ac6ee6b8399781a2394c9e663e4ac0c0850655b6ccb4d80d8ad307e6b62d3ab", + "workspace_final.tar.gz": "427d4a3a2eaef5eb3ed46b025996fdbc51fce966e234f63855e4542984c3bd1e", + "cost.json": "b296c66ff2312ad59757460488b509e23770bf1e6d79e3bee11d7a11110770cf", + "gpu_telemetry.json": "0c735a89e20aa5047ba35be526f88aeda1add9f5b234af221f4034ed3d706a7e", + "grade.json": "7f69334378ac3187de489eb74ad81c8245267d1a5711612b94879f297e758bcf" + } + } +} diff --git a/tooling/deployments/gemma4-31b-q4-tower2/canonical-telemetry-gate-incident-20260802.json b/tooling/deployments/gemma4-31b-q4-tower2/canonical-telemetry-gate-incident-20260802.json new file mode 100644 index 00000000..4e82efe9 --- /dev/null +++ b/tooling/deployments/gemma4-31b-q4-tower2/canonical-telemetry-gate-incident-20260802.json @@ -0,0 +1,32 @@ +{ + "schema_version": 1, + "campaign": "gemma4-31b-q4-mmbt", + "detected_at": "2026-08-02T00:45:33Z", + "classification": "infrastructure-invalid attempts interrupted to enforce the preregistered telemetry gate", + "condition": "the restarted canonical lanes began before per-run sidecar attribution was proven for every new cell", + "operator_action": { + "action": "stopped the supervisor before either active retry completed", + "score_independence": "the attempts were interrupted and classified before grading", + "preserved_attempts": [ + "p1_bugfix_gemma4-31b-q4_v3_retry-20260802T004533Z-64195329", + "p1_testwrite_gemma4-31b-q4_v1_retry-20260802T004533Z-09111a0e" + ], + "replacement_policy": "rerun both exact canonical cells only after the telemetry writer and sidecar gate passed" + }, + "completed_pretelemetry_cells": { + "runs": [ + "p1_bugfix_gemma4-31b-q4_v1", + "p1_bugfix_gemma4-31b-q4_v2" + ], + "treatment": "retain their original quality outcomes and map them to separately labeled, idle-start supplemental telemetry observations" + }, + "external_state": { + "root": "/home/michael/gemma4-campaign-state/autopilot-state-telemetry-gate-20260802T004533Z", + "files": { + "autopilot.log": "b7f30ab45440710a2b372dd042b1643c4c9ae55b047881fd3a3ee34f202855c4", + "run-gemma4-31b-q4-lane0-port8000.log": "94b8d3c640618a53b38b7d606dd5be03a413afebec6eac04d07d6898ae5c0b92", + "run-gemma4-31b-q4-lane1-port8001.log": "9ea8dba0c7b30bea97dce427f4828cc22355745b33bd761115eea151f75073cf", + "status.json": "c88dca145c48cb28dee1c2a68574022d75a5d35a2d3395b2c02502cefc7b1156" + } + } +} diff --git a/tooling/deployments/gemma4-31b-q4-tower2/configure-openclaw-dual-replicas.sh b/tooling/deployments/gemma4-31b-q4-tower2/configure-openclaw-dual-replicas.sh new file mode 100755 index 00000000..8a8fdf1f --- /dev/null +++ b/tooling/deployments/gemma4-31b-q4-tower2/configure-openclaw-dual-replicas.sh @@ -0,0 +1,65 @@ +#!/usr/bin/env bash +set -euo pipefail + +CONFIG=/home/michael/.openclaw/openclaw.json +STATE_DIR=/home/michael/gemma4-campaign-state +BACKUP=$STATE_DIR/openclaw.gemma-single-provider.json +MODEL=Gemma-4-31B-it-QAT-Q4_0 +OPENCLAW=/home/michael/.npm-global/bin/openclaw + +install -d -m 700 "$STATE_DIR" +if [[ ! -e "$BACKUP" ]]; then + install -m 600 "$CONFIG" "$BACKUP" +fi + +patch_file="$(mktemp "$STATE_DIR/openclaw-dual-replica-patch.XXXXXX.json")" +dry_run_file="$(mktemp "$STATE_DIR/openclaw-dual-replica-dry-run.XXXXXX.json")" +chmod 0600 "$patch_file" "$dry_run_file" +cleanup() { + rm -f -- "$patch_file" "$dry_run_file" +} +trap cleanup EXIT + +jq --arg model "$MODEL" ' + { + models: { + providers: { + tower1: (.models.providers.tower | .baseUrl = "http://127.0.0.1:8001/v1") + } + }, + agents: { + defaults: { + models: { + ("tower1/" + $model): {alias: "Pixel"} + } + }, + list: ( + .agents.list + | map( + if .id == "main" then + . + {model: {primary: ("tower/" + $model)}} + elif .id == "pixel" then + . + {model: {primary: ("tower1/" + $model)}} + else + . + end + ) + ) + } + } +' "$CONFIG" >"$patch_file" + +"$OPENCLAW" config patch --file "$patch_file" --dry-run --json >"$dry_run_file" +jq -e '.ok == true' "$dry_run_file" >/dev/null +"$OPENCLAW" config patch --file "$patch_file" >/dev/null +"$OPENCLAW" config validate --json | jq -e '.valid == true' + +jq -r --arg model "$MODEL" ' + [ + .models.providers.tower.baseUrl, + .models.providers.tower1.baseUrl, + (.agents.list[] | select(.id == "main") | .model.primary), + (.agents.list[] | select(.id == "pixel") | .model.primary) + ] | @tsv +' "$CONFIG" +sha256sum "$BACKUP" "$CONFIG" diff --git a/tooling/deployments/gemma4-31b-q4-tower2/dual_lane_probe.py b/tooling/deployments/gemma4-31b-q4-tower2/dual_lane_probe.py new file mode 100755 index 00000000..368efe71 --- /dev/null +++ b/tooling/deployments/gemma4-31b-q4-tower2/dual_lane_probe.py @@ -0,0 +1,143 @@ +#!/usr/bin/env python3 +"""Launch matched server microbenches on both independent GPU replicas.""" + +from __future__ import annotations + +import argparse +import datetime as dt +import hashlib +import json +import pathlib +import subprocess +import time +from typing import Any + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("--endpoint0", default="http://127.0.0.1:8000") + parser.add_argument("--endpoint1", default="http://127.0.0.1:8001") + parser.add_argument("--label", required=True) + parser.add_argument("--output-root", required=True, type=pathlib.Path) + parser.add_argument( + "--microbench", + default=( + "/home/michael/bench-gemma4-31b-q4/tooling/deployments/" + "gemma4-31b-q4-tower2/server_microbench.py" + ), + ) + parser.add_argument("--lane-concurrency", type=int, default=4) + parser.add_argument("--prompt-tokens", type=int, default=1024) + parser.add_argument("--predict-tokens", type=int, default=256) + parser.add_argument("--timeout", type=float, default=900) + args = parser.parse_args() + + timestamp = dt.datetime.now(dt.timezone.utc).strftime("%Y%m%dT%H%M%SZ") + out_dir = args.output_root / f"{args.label}-{timestamp}" + out_dir.mkdir(parents=True, exist_ok=False) + + telemetry_handle = (out_dir / "nvidia-smi.csv").open("wb") + query = ( + "timestamp,index,uuid,memory.used,utilization.gpu,utilization.memory," + "power.draw,power.limit,temperature.gpu,clocks.sm,clocks.mem" + ) + sampler = subprocess.Popen( + [ + "nvidia-smi", + f"--query-gpu={query}", + "--format=csv,noheader,nounits", + "--loop-ms=200", + ], + stdout=telemetry_handle, + stderr=subprocess.STDOUT, + ) + + processes: list[subprocess.Popen[str]] = [] + started = time.perf_counter() + try: + for lane, endpoint in enumerate((args.endpoint0, args.endpoint1)): + command = [ + args.microbench, + "--endpoint", + endpoint, + "--label", + f"{args.label}-lane{lane}", + "--output-root", + str(args.output_root), + "--prompt-tokens", + "128", + "--stream-predict", + "64", + "--concurrency", + str(args.lane_concurrency), + "--concurrency-prompt-tokens", + str(args.prompt_tokens), + "--concurrency-predict", + str(args.predict_tokens), + "--timeout", + str(args.timeout), + ] + processes.append( + subprocess.Popen(command, stdout=subprocess.PIPE, stderr=subprocess.PIPE, text=True) + ) + + lane_results: list[dict[str, Any]] = [] + return_codes: list[int] = [] + for lane, process in enumerate(processes): + stdout, stderr = process.communicate(timeout=args.timeout) + return_codes.append(process.returncode) + (out_dir / f"lane{lane}.stdout.json").write_text(stdout, encoding="utf-8") + (out_dir / f"lane{lane}.stderr.log").write_text(stderr, encoding="utf-8") + lane_results.append(json.loads(stdout) if stdout.strip() else {}) + finally: + for process in processes: + if process.poll() is None: + process.terminate() + try: + process.wait(timeout=5) + except subprocess.TimeoutExpired: + process.kill() + process.wait(timeout=5) + wall_seconds = time.perf_counter() - started + sampler.terminate() + try: + sampler.wait(timeout=5) + except subprocess.TimeoutExpired: + sampler.kill() + sampler.wait(timeout=5) + telemetry_handle.close() + + lane_tps = [ + float(result["concurrency"][0]["aggregate_decode_tokens_per_second_wall"]) + for result in lane_results + ] + summary = { + "schema_version": 1, + "timestamp": timestamp, + "label": args.label, + "topology": "two independent GPU replicas", + "lane_concurrency": args.lane_concurrency, + "total_concurrency": args.lane_concurrency * 2, + "prompt_tokens_per_request": args.prompt_tokens, + "predict_tokens_per_request": args.predict_tokens, + "wall_seconds_including_lane_warmups": wall_seconds, + "lane_return_codes": return_codes, + "lane_results": lane_results, + "lane_aggregate_decode_tps": lane_tps, + "combined_aggregate_decode_tps": sum(lane_tps), + "passed": return_codes == [0, 0], + } + (out_dir / "summary.json").write_text( + json.dumps(summary, indent=2, sort_keys=True) + "\n", encoding="utf-8" + ) + sums = [] + for path in sorted(out_dir.iterdir()): + if path.is_file() and path.name != "SHA256SUMS": + sums.append(f"{hashlib.sha256(path.read_bytes()).hexdigest()} {path.name}") + (out_dir / "SHA256SUMS").write_text("\n".join(sums) + "\n", encoding="utf-8") + print(json.dumps(summary, indent=2)) + return 0 if summary["passed"] else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/tooling/deployments/gemma4-31b-q4-tower2/ensure-gemma4-winner.sh b/tooling/deployments/gemma4-31b-q4-tower2/ensure-gemma4-winner.sh new file mode 100755 index 00000000..9728dbba --- /dev/null +++ b/tooling/deployments/gemma4-31b-q4-tower2/ensure-gemma4-winner.sh @@ -0,0 +1,38 @@ +#!/usr/bin/env bash +# Recover the preregistered two-replica Gemma winner and prove both API lanes. +set -euo pipefail + +SERVICES=( + mmbt-gemma4@replica-gpu0-q8-s4.service + mmbt-gemma4@replica-gpu1-q8-s4.service +) +PORTS=(8000 8001) +GRACE_SECS="${GEMMA_RECOVERY_GRACE_SECS:-300}" + +for service in "${SERVICES[@]}"; do + systemctl --user start "$service" +done + +deadline=$(( $(date +%s) + GRACE_SECS )) +while (( $(date +%s) < deadline )); do + ready=1 + for i in "${!SERVICES[@]}"; do + if ! systemctl --user is-active --quiet "${SERVICES[$i]}" || \ + ! curl -fsS --max-time 5 "http://127.0.0.1:${PORTS[$i]}/v1/models" >/dev/null; then + ready=0 + fi + done + if (( ready )); then + printf 'Gemma winner healthy: %s\n' "${PORTS[*]}" + exit 0 + fi + sleep 5 +done + +for i in "${!SERVICES[@]}"; do + printf '%s active=%s endpoint=%s\n' \ + "${SERVICES[$i]}" \ + "$(systemctl --user is-active "${SERVICES[$i]}" 2>/dev/null || true)" \ + "$(curl -fsS --max-time 5 "http://127.0.0.1:${PORTS[$i]}/v1/models" >/dev/null && echo up || echo down)" >&2 +done +exit 1 diff --git a/tooling/deployments/gemma4-31b-q4-tower2/extended-lane0.env b/tooling/deployments/gemma4-31b-q4-tower2/extended-lane0.env new file mode 100644 index 00000000..caceda37 --- /dev/null +++ b/tooling/deployments/gemma4-31b-q4-tower2/extended-lane0.env @@ -0,0 +1,3 @@ +GEMMA_EXTENDED_LANE_INDEX=0 +GEMMA_EXTENDED_LANE_COUNT=2 +GEMMA_EXTENDED_PORT=8000 diff --git a/tooling/deployments/gemma4-31b-q4-tower2/extended-lane1.env b/tooling/deployments/gemma4-31b-q4-tower2/extended-lane1.env new file mode 100644 index 00000000..6d19fa10 --- /dev/null +++ b/tooling/deployments/gemma4-31b-q4-tower2/extended-lane1.env @@ -0,0 +1,3 @@ +GEMMA_EXTENDED_LANE_INDEX=1 +GEMMA_EXTENDED_LANE_COUNT=2 +GEMMA_EXTENDED_PORT=8001 diff --git a/tooling/deployments/gemma4-31b-q4-tower2/final-validation.json b/tooling/deployments/gemma4-31b-q4-tower2/final-validation.json new file mode 100644 index 00000000..ed48dca8 --- /dev/null +++ b/tooling/deployments/gemma4-31b-q4-tower2/final-validation.json @@ -0,0 +1,251 @@ +{ + "schema_version": 1, + "captured_at": "2026-08-02T00:05:39Z", + "scope": "serving topology and production integration; this is not the final MMBT campaign result", + "status": "serving_validated", + "winner": { + "topology": "dual-independent-replicas", + "selection_reason": "Best quality-preserving aggregate throughput with independent failure domains, native context per slot, stable restart behavior, and no cross-GPU dependency.", + "instances": [ + { + "physical_gpu": 0, + "endpoint": "http://127.0.0.1:8000/v1", + "service": "mmbt-gemma4@replica-gpu0-q8-s4.service", + "production_agent": "Sanctuary" + }, + { + "physical_gpu": 1, + "endpoint": "http://127.0.0.1:8001/v1", + "service": "mmbt-gemma4@replica-gpu1-q8-s4.service", + "production_agent": "Pixel" + } + ], + "per_instance": { + "model": "Gemma-4-31B-it-QAT-Q4_0", + "slots": 4, + "context_pool_tokens": 1048576, + "context_tokens_per_slot": 262144, + "kv_unified": false, + "cache_type_k": "q8_0", + "cache_type_v": "q8_0", + "gpu_layers": "all", + "split_mode": "none", + "flash_attention": true, + "multimodal_projector": true, + "steady_vram_mib": 64789, + "vram_headroom_mib": 33098, + "gpu_power_limit_w": 500, + "enabled_at_boot": true + } + }, + "runtime": { + "kind": "llama.cpp-current-source", + "commit": "11924d4c17abc27383376a1ac6a24fa3e36c1c0c", + "version": 10223, + "llama_server_sha256": "200b403b5735418ff1f6da0cea1938e413e11869ae362e8044a12b0df04622fc", + "cuda": "13.1.115", + "cuda_architecture": "120a-real", + "nccl": "2.29.7-private-runtime", + "runtime_build_has_embedded_nccl_runpath": true, + "cache_reuse_note": "llama.cpp disables cache_reuse when the multimodal projector is loaded; this is logged and not hidden." + }, + "topology_results": { + "single_gpu_f16_one_slot": { + "steady_vram_mib": 40337, + "prefill_tokens_per_second": { + "1024": 1565.4088924388461, + "32768": 2830.1683910895604, + "131072": 1483.8109873875555, + "250000": 945.7353381742734 + }, + "decode_tokens_per_second": { + "1024": 70.28992948492308, + "32768": 62.066820265883585, + "131072": 47.26540986218625, + "250000": 36.69319756850233 + }, + "concurrency_8_aggregate_decode_tps": 48.04099537091844 + }, + "single_gpu_f16_two_slots": { + "steady_vram_mib": 62101, + "hard_context_tokens_per_slot": 262144, + "concurrency_8_aggregate_decode_tps": 68.03549794679058 + }, + "single_gpu_q8_four_slots": { + "steady_vram_mib": 64789, + "hard_context_tokens_per_slot": 262144, + "concurrency_4_aggregate_decode_tps": 128.4227371181164, + "concurrency_8_replicates_tps": [ + 112.61659474507437, + 129.7601716359008, + 123.99151098153261 + ], + "concurrency_8_mean_tps": 122.12275912083594, + "concurrency_8_median_tps": 123.99151098153261 + }, + "single_gpu_q8_six_slots_append_only_extension": { + "steady_vram_mib": 87833, + "hard_context_tokens_per_slot": 262144, + "concurrency_6_aggregate_decode_tps": 136.91836422803846, + "concurrency_8_replicates_tps": [ + 103.49956096453603, + 122.87712597047079, + 124.37803366674967 + ], + "concurrency_8_mean_tps": 116.9182402005855, + "concurrency_8_median_tps": 122.87712597047079, + "selection_note": "Within the 10 percent tie band but slower at the preregistered concurrency-8 point, with 23044 MiB less headroom than four slots." + }, + "two_gpu_independent_four_slot_replicas": { + "total_concurrency": 8, + "lane0_aggregate_decode_tps": 143.72858299108827, + "lane1_aggregate_decode_tps": 146.5506168778769, + "combined_aggregate_decode_tps": 290.2791998689652, + "passed": true + } + }, + "quality_and_operational_gates": { + "q8_near_native_recall": { + "prompt_tokens": 245016, + "completion_tokens": 331, + "total_tokens": 245347, + "start_marker": true, + "middle_marker": true, + "end_marker": true, + "finish_reason": "stop", + "passed": true + }, + "postrestart_contract_gpu0": { + "chat": true, + "tool_call": true, + "tool_followup": true, + "passed": true + }, + "postrestart_contract_gpu1": { + "chat": true, + "tool_call": true, + "tool_followup": true, + "passed": true + }, + "cold_service_restart": { + "gpu0_endpoint": true, + "gpu1_endpoint": true, + "passed": true + }, + "production_routing": { + "sanctuary_provider": "tower", + "sanctuary_endpoint": "http://127.0.0.1:8000/v1", + "pixel_provider": "tower1", + "pixel_endpoint": "http://127.0.0.1:8001/v1", + "sanctuary_marker_pass": true, + "pixel_marker_pass": true, + "live_openclaw_config_sha256": "e81e856ec5c9a7bb20532ca479c24c7fa78caa6ce7222982df7ebd95929e6c25", + "passed": true + } + }, + "rejected_candidates": { + "dual_layer_1to1": { + "classification": "valid topology failure; not quality preserving", + "evidence": [ + "Deterministic ordinary chat returned corrupted unused-token markers instead of the requested marker.", + "Deterministic required tool call returned HTTP 500: output did not match peg-gemma4 format.", + "A full-envelope OpenClaw retry produced more than 4754 tokens for a one-marker response and was cancelled as operationally runaway." + ], + "passed": false + }, + "dual_row_1to1": { + "classification": "objectively pre-inference invalid/unsupported on this runtime and PCIe topology", + "error": "device CUDA0 does not support split buffers", + "model_was_evaluated": false, + "passed": false + }, + "health_probe_128_token_attempt": { + "classification": "infrastructure-invalid probe retained in evidence", + "reason": "The probe imposed max_tokens=128 and the server stopped exactly at 128 with OpenClaw stopReason=length. The harness was corrected to a 262144-token ceiling and unique sessions before valid health checks." + } + }, + "external_evidence": [ + {"path": "/home/michael/gemma4-campaign-state/topology/single-gpu1-f16-s1-fullctx-20260801T233216Z/summary.json", "sha256": "66bbe34b6a3172a5c8f889821136f4061774ee5f47afcce91e77e445a5fe8fd4"}, + {"path": "/home/michael/gemma4-campaign-state/topology/replica-gpu1-f16-s2-concurrency-20260801T234153Z/summary.json", "sha256": "6393eb21966ee00b909d9e1fd96276040b40fa2c2bf01ae165cddb5cd3e158f1"}, + {"path": "/home/michael/gemma4-campaign-state/topology/replica-gpu1-q8-s4-concurrency-20260801T234354Z/summary.json", "sha256": "98f977520b227f03a1abd720cfc21f6178f6c9b3a2ef4279ae9b85e2b77f9906"}, + {"path": "/home/michael/gemma4-campaign-state/topology/replica-gpu1-q8-s4-c8-r2-20260801T235333Z/summary.json", "sha256": "a4f77f8c36752ac83aeb7186dc516ad00615204b550b4ee399aa1b1b20dd28eb"}, + {"path": "/home/michael/gemma4-campaign-state/topology/replica-gpu1-q8-s4-c8-r3-20260801T235350Z/summary.json", "sha256": "94391b12860e28f49c4b466144626a486c9eaed401e8c4d96e1b4b28f1b738ca"}, + {"path": "/home/michael/gemma4-campaign-state/topology/replica-gpu1-q8-s6-concurrency-20260801T235025Z/summary.json", "sha256": "7edd9664f979f726d92ec31bf943a5ed1629710d4d2e7a507b4ffd9c5ccdf111"}, + {"path": "/home/michael/gemma4-campaign-state/topology/replica-gpu1-q8-s6-c8-r2-20260801T235214Z/summary.json", "sha256": "e92485614733d79e52b6a721aba2c3ba8662d01d40f2fe46f99653a24d0f06a4"}, + {"path": "/home/michael/gemma4-campaign-state/topology/replica-gpu1-q8-s6-c8-r3-20260801T235236Z/summary.json", "sha256": "282314c14edaf21b8ecc6e30e4c9806393036168ce7cd9489c06912cacd81637"}, + {"path": "/home/michael/gemma4-campaign-state/topology/dual-replica-q8-s4-c4-20260802T000125Z/summary.json", "sha256": "d2b81569b27872996cb3918e5585beaa9cd7ff9ec74c6b5731a1b2bd23f2079a"}, + {"path": "/home/michael/gemma4-campaign-state/recall/replica-gpu1-q8-s4-20260801T234503Z/summary.json", "sha256": "f03fc416496c743bd2901a29a23ffc57e14c15b535da15bd19110c17deab797f"}, + {"path": "/home/michael/gemma4-campaign-state/contracts/dual-layer-1to1-seed424242-20260801T235856Z/summary.json", "sha256": "4534ceaeb22e79b35c344594674ed46e606e477d718089bdc5e8ed9c8a7ef842"}, + {"path": "/home/michael/gemma4-campaign-state/contracts/winner-gpu0-postrestart-20260802T000218Z/summary.json", "sha256": "ea1eb74e318830092804f2d8432e20af641f2260d398cfa8c8f3a743d32b01cc"}, + {"path": "/home/michael/gemma4-campaign-state/contracts/winner-gpu1-postrestart-20260802T000218Z/summary.json", "sha256": "bf49d5b619ec0d8e981f2327ad10efc97dc0f79dc3edc4ece0c6c90644f7455d"}, + {"path": "/home/michael/gemma4-campaign-state/health/gemma-dual-provider-routing-20260802T000514Z/SHA256SUMS", "sha256": "ea5595dac02256abffb81fc0bc53e5407ff694d5dddb7d88f5916ce1d968e0e4"}, + {"path": "/home/michael/gemma4-campaign-state/health/gemma-winner-postrestart-20260802T000229Z/SHA256SUMS", "sha256": "1c8f697e5757a037822cdb563b8f6f73a1e5bed41d9b524dd48dc19c86d1cc98"}, + {"path": "/home/michael/gemma4-campaign-state/serving-final-evidence/dual-layer.journal.log", "sha256": "d3e166bf2de796c58b187b3651b2f289cccd10faf2f3913c16f09f1a36076743"}, + {"path": "/home/michael/gemma4-campaign-state/serving-final-evidence/dual-row.journal.log", "sha256": "02f20ac33f1a0bebc27a16d45306494d32985a0486e60e64b3b98b304d097297"}, + {"path": "/home/michael/gemma4-campaign-state/serving-final-evidence/dual-provider-routing.log", "sha256": "5e755a232f49388aaf086841c8a12f986f4ec9eb42a23dbf2cb9f69f972f69cf"} + ], + "production_restore": { + "required_after_campaign": true, + "manifest": "production-restore.json", + "deepseek_launcher_sha256": "4fca6a4a478f0876a1e6d7241dd87a3c9574bd970f29943bb3696993d4bfe25d", + "pre_campaign_openclaw_sha256": "79f2872821d68f76a2c42cb8f3972ae494b380a4d0d058fba405dbcc5e489c7f", + "status": "completed_and_validated", + "validated_at": "2026-08-02T07:06:18Z", + "evidence_directory": "/home/michael/gemma4-campaign-state/production-restore/20260802T070054Z", + "deepseek": { + "served_model": "DeepSeek-V4-Flash-0731", + "max_model_len": 1048576, + "health_http_status": 200, + "image_digest": "sha256:48518e91cf87dd0c0483c76ff86e81dfc0f46de7e364b46f7a82c481ce08188f", + "container_restart_policy": "unless-stopped", + "container_restart_count": 0, + "container_oom_killed": false, + "gpu_power_limits_w": [500, 500] + }, + "openclaw": { + "restored_config_sha256": "79f2872821d68f76a2c42cb8f3972ae494b380a4d0d058fba405dbcc5e489c7f", + "restored_config_mode": "600", + "gateway_active": true, + "sanctuary_portal_healthy": true, + "pixel_portal_healthy": true + }, + "fresh_agent_tool_checks": { + "sanctuary": { + "status": "ok", + "provider": "tower", + "model": "DeepSeek-V4-Flash-0731", + "context_tokens": 1048576, + "fallback_used": false, + "tool": "exec", + "tool_command": "printf SANCTUARY_OK", + "tool_result": "SANCTUARY_OK", + "final_text": "SANCTUARY_OK", + "result_sha256": "e52e459db9165e488fddf799e32304f4066a3392810f87ea885604d206e5f8a9", + "session_trace_sha256": "e8e6af4f8dfab6ce19acdd0df1e95658415b4f94a6c39639b738d750df1491e9" + }, + "pixel": { + "status": "ok", + "provider": "tower", + "model": "DeepSeek-V4-Flash-0731", + "context_tokens": 1048576, + "fallback_used": false, + "tool": "exec", + "tool_command": "printf PIXEL_OK", + "tool_result": "PIXEL_OK", + "final_text": "PIXEL_OK", + "result_sha256": "4099f3faf1effcad60c8a2afae04636b9930afe4eac6c1e380f0491229183647", + "session_trace_sha256": "e55d69c95581e9c7ef4b679a1005d3c69f56235207edea9e4c9f6af0922fec8a" + } + }, + "campaign_cleanup": { + "gemma_serving_services_running": 0, + "gemma_serving_containers_running": 0, + "campaign_sandbox_containers_stopped_not_deleted": 96 + }, + "invalid_validation_attempts_retained": { + "count": 2, + "classification": "harness_invalid_before_model_invocation", + "reason": "SSH retokenized a multiword --message argument; the replacement checks used --message-file and fresh session IDs." + } + } +} diff --git a/tooling/deployments/gemma4-31b-q4-tower2/gemma4-telemetry-sidecar.sh b/tooling/deployments/gemma4-31b-q4-tower2/gemma4-telemetry-sidecar.sh new file mode 100755 index 00000000..42b3008b --- /dev/null +++ b/tooling/deployments/gemma4-31b-q4-tower2/gemma4-telemetry-sidecar.sh @@ -0,0 +1,30 @@ +#!/usr/bin/env bash +set -uo pipefail + +ROOT=/home/michael/bench-gemma4-31b-q4 +LOGS="$ROOT/logs" +DEPLOY="$ROOT/tooling/deployments/gemma4-31b-q4-tower2" +CSV=/home/michael/gemma4-campaign-state/telemetry/gemma4-31b-q4-gpu.csv +REPORT=/home/michael/gemma4-campaign-state/telemetry/gemma4-31b-q4-report.json +SIDELOG=/home/michael/gemma4-campaign-state/telemetry/sidecar.log + +while true; do + if [[ -s "$CSV" ]]; then + python3 "$DEPLOY/analyze_replica_telemetry.py" \ + --csv "$CSV" --logs-dir "$LOGS" --cap-per-gpu 500 \ + --write-run-artifacts --output "$REPORT" >>"$SIDELOG" 2>&1 || true + fi + + while IFS= read -r run_dir; do + [[ -f "$run_dir/receipt.json" && -f "$run_dir/transcript.jsonl" ]] || continue + if [[ -f "$run_dir/summary.json" && ! -f "$run_dir/workspace_final.tar.gz" ]]; then + continue + fi + [[ -f "$run_dir/cost.json" ]] && continue + python3 "$ROOT/tooling/scripts/extract_cost.py" "$run_dir" >>"$SIDELOG" 2>&1 || true + done < <(find "$LOGS" -mindepth 2 -maxdepth 2 -type f \ + \( -name summary.json -o -name label.json \) \ + -path '*gemma4-31b-q4*' -printf '%h\n' | sort -u) + + sleep 60 +done diff --git a/tooling/deployments/gemma4-31b-q4-tower2/gemma4-topology-control.sh b/tooling/deployments/gemma4-31b-q4-tower2/gemma4-topology-control.sh new file mode 100755 index 00000000..060f06cf --- /dev/null +++ b/tooling/deployments/gemma4-31b-q4-tower2/gemma4-topology-control.sh @@ -0,0 +1,67 @@ +#!/usr/bin/env bash +set -euo pipefail + +UNIT_SOURCE=/home/michael/bench-gemma4-31b-q4/tooling/deployments/gemma4-31b-q4-tower2/mmbt-gemma4@.service +UNIT_NAME=mmbt-gemma4@.service +KNOWN_INSTANCES=( + single-gpu0 single-gpu1 dual-layer-1to1 dual-row-1to1 + replica-gpu0 replica-gpu1 + replica-gpu0-f16-s2 replica-gpu1-f16-s2 + replica-gpu0-q8-s4 replica-gpu1-q8-s4 + replica-gpu0-q8-s6 replica-gpu1-q8-s6 +) + +usage() { + printf 'Usage: %s install|start |stop|status\n' "$0" >&2 + exit 64 +} + +install_unit() { + test -r "$UNIT_SOURCE" + systemctl --user link "$UNIT_SOURCE" >/dev/null + systemctl --user daemon-reload +} + +stop_all() { + local instance + for instance in "${KNOWN_INSTANCES[@]}"; do + systemctl --user stop "mmbt-gemma4@${instance}.service" 2>/dev/null || true + done +} + +deepseek_is_running() { + [[ "$(docker inspect -f '{{.State.Running}}' deepseek-v4-flash-0731 2>/dev/null || true)" == true ]] +} + +start_candidate() { + local candidate="$1" + if deepseek_is_running; then + printf 'Refusing to start Gemma while deepseek-v4-flash-0731 owns the GPUs.\n' >&2 + exit 69 + fi + install_unit + stop_all + case "$candidate" in + single-gpu0|single-gpu1|dual-layer-1to1|dual-row-1to1) + systemctl --user start "mmbt-gemma4@${candidate}.service" + ;; + dual-independent-replicas) + systemctl --user start mmbt-gemma4@replica-gpu0.service mmbt-gemma4@replica-gpu1.service + ;; + *) usage ;; + esac +} + +show_status() { + systemctl --user --no-pager --full list-units 'mmbt-gemma4@*.service' + nvidia-smi --query-gpu=index,power.limit,memory.total,memory.used,utilization.gpu,power.draw,temperature.gpu \ + --format=csv,noheader,nounits +} + +case "${1:-}" in + install) install_unit ;; + start) [[ $# -eq 2 ]] || usage; start_candidate "$2" ;; + stop) stop_all ;; + status) show_status ;; + *) usage ;; +esac diff --git a/tooling/deployments/gemma4-31b-q4-tower2/gemma4_gpu_telemetry.py b/tooling/deployments/gemma4-31b-q4-tower2/gemma4_gpu_telemetry.py new file mode 100755 index 00000000..bc4ed173 --- /dev/null +++ b/tooling/deployments/gemma4-31b-q4-tower2/gemma4_gpu_telemetry.py @@ -0,0 +1,171 @@ +#!/usr/bin/env python3 +"""Five-second Tower2 telemetry with deterministic per-replica run attribution.""" +from __future__ import annotations + +import argparse +import csv +import os +import subprocess +import time +from datetime import datetime, timezone +from pathlib import Path + + +FIELDS = [ + "ts", "gpu", "endpoint_port", "power_limit_w", "power_w", "util_sm", + "util_mem", "mem_used_mib", "temp_c", "sm_clk_mhz", "cell", + "harness_pid", "cpu_package_power_w", +] +PORT_BY_GPU = {0: 8000, 1: 8001} +RAPL_ENERGY = Path("/sys/devices/virtual/powercap/intel-rapl/intel-rapl:0/energy_uj") +RAPL_MAX = Path("/sys/devices/virtual/powercap/intel-rapl/intel-rapl:0/max_energy_range_uj") + + +def read_proc_argv(pid_dir: Path) -> list[str]: + try: + return [ + part.decode("utf-8", "replace") + for part in (pid_dir / "cmdline").read_bytes().split(b"\0") + if part + ] + except (OSError, PermissionError): + return [] + + +def active_harnesses(proc_root: Path = Path("/proc")) -> dict[int, tuple[str, int]]: + """Return endpoint port -> (run name, PID) for live Gemma harnesses.""" + found: dict[int, tuple[str, int]] = {} + if not proc_root.exists(): + return found + for pid_dir in proc_root.iterdir(): + if not pid_dir.name.isdigit(): + continue + argv = read_proc_argv(pid_dir) + harness_index = next( + (i for i, arg in enumerate(argv) if arg.endswith("/tooling/harness.py")), + None, + ) + if harness_index is None or harness_index + 1 >= len(argv): + continue + if "bench-gemma4-31b-q4" not in argv[harness_index]: + continue + try: + port_index = argv.index("--port") + port = int(argv[port_index + 1]) + except (ValueError, IndexError): + continue + if port not in PORT_BY_GPU.values(): + continue + candidate = (argv[harness_index + 1], int(pid_dir.name)) + # Multiple benchmark harnesses on one inference lane are forbidden. If + # encountered, retain the lowest PID deterministically and make the raw + # process list available through ordinary host evidence for diagnosis. + if port not in found or candidate[1] < found[port][1]: + found[port] = candidate + return found + + +def read_rapl_energy_uj() -> int | None: + if not RAPL_ENERGY.exists(): + return None + result = subprocess.run( + ["sudo", "-n", "cat", str(RAPL_ENERGY)], + capture_output=True, text=True, check=False, + ) + try: + return int(result.stdout.strip()) if result.returncode == 0 else None + except ValueError: + return None + + +def read_rapl_max_uj() -> int | None: + try: + return int(RAPL_MAX.read_text().strip()) + except (OSError, ValueError): + return None + + +def rapl_power(previous: tuple[int, float] | None, current_uj: int | None, + current_t: float, max_uj: int | None) -> tuple[float | None, tuple[int, float] | None]: + if current_uj is None: + return None, previous + if previous is None: + return None, (current_uj, current_t) + previous_uj, previous_t = previous + delta_uj = current_uj - previous_uj + if delta_uj < 0 and max_uj: + delta_uj += max_uj + delta_t = current_t - previous_t + power = delta_uj / 1_000_000.0 / delta_t if delta_uj >= 0 and delta_t > 0 else None + return power, (current_uj, current_t) + + +def nvidia_rows() -> list[list[str]]: + query = ( + "index,power.limit,power.draw,utilization.gpu,utilization.memory," + "memory.used,temperature.gpu,clocks.current.sm" + ) + result = subprocess.run( + ["nvidia-smi", f"--query-gpu={query}", "--format=csv,noheader,nounits"], + capture_output=True, text=True, check=False, + ) + if result.returncode != 0: + return [] + return [[part.strip() for part in line.split(",")] for line in result.stdout.splitlines()] + + +def prepare_output(path: Path) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + if path.exists() and path.stat().st_size: + with path.open(newline="") as handle: + header = next(csv.reader(handle), []) + if header != FIELDS: + raise RuntimeError(f"refusing incompatible telemetry CSV: {path}") + return + with path.open("w", newline="") as handle: + csv.writer(handle).writerow(FIELDS) + + +def run(output: Path, interval: float) -> None: + prepare_output(output) + rapl_state = None + rapl_max = read_rapl_max_uj() + while True: + sample_t = time.time() + ts = datetime.now(timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ") + current_energy = read_rapl_energy_uj() + cpu_power, rapl_state = rapl_power(rapl_state, current_energy, sample_t, rapl_max) + harnesses = active_harnesses() + rows = [] + for values in nvidia_rows(): + if len(values) != 8: + continue + gpu = int(values[0]) + port = PORT_BY_GPU.get(gpu) + cell, pid = harnesses.get(port, ("", "")) + rows.append([ + ts, gpu, port or "", *values[1:], cell, pid, + "" if cpu_power is None else f"{cpu_power:.4f}", + ]) + if rows: + with output.open("a", newline="") as handle: + writer = csv.writer(handle) + writer.writerows(rows) + handle.flush() + os.fsync(handle.fileno()) + elapsed = time.time() - sample_t + time.sleep(max(0.1, interval - elapsed)) + + +def main() -> None: + parser = argparse.ArgumentParser() + parser.add_argument("--output", required=True, type=Path) + parser.add_argument("--interval", type=float, default=5.0) + args = parser.parse_args() + if args.interval <= 0: + parser.error("--interval must be positive") + run(args.output.resolve(), args.interval) + + +if __name__ == "__main__": + main() diff --git a/tooling/deployments/gemma4-31b-q4-tower2/invalid-attempt-classifications.json b/tooling/deployments/gemma4-31b-q4-tower2/invalid-attempt-classifications.json new file mode 100644 index 00000000..0104e3f9 --- /dev/null +++ b/tooling/deployments/gemma4-31b-q4-tower2/invalid-attempt-classifications.json @@ -0,0 +1,117 @@ +{ + "schema_version": 1, + "campaign": "gemma4-31b-q4-mmbt", + "policy": "Every excluded attempt must have affirmative infrastructure evidence, be classified without consulting a score, remain hash-preserved, and receive an exact canonical replacement.", + "attempts": { + "p1_bugfix_gemma4-31b-q4_v3_retry-20260802T003822Z-dea60926": { + "source_run": "p1_bugfix_gemma4-31b-q4_v3", + "classification": "infrastructure-invalid", + "reason_code": "operator-interrupted-companion-after-harness-dispatch-defect", + "incident_document": "canonical-harness-incident-20260802.json", + "classified_before_grade": true, + "affirmative_evidence": [ + "the companion lane raised the recorded KeyError in the shared defective harness", + "the supervisor was stopped immediately to prevent further evaluation under that harness" + ], + "expected_files": { + "receipt.json": "0a571fa5aaefacb15833eb56d63969d45b69f7e66b7622ec309ddf52ff1be4b5", + "transcript.jsonl": "7c9fa7c58e4ea36ab46938b690785b4a699b98d38c7f934d1aed083d87fd3163" + }, + "replacement": {"required": true, "canonical_run": "p1_bugfix_gemma4-31b-q4_v3", "status": "completed"} + }, + "p1_testwrite_gemma4-31b-q4_v1_retry-20260802T003822Z-8061c2b0": { + "source_run": "p1_testwrite_gemma4-31b-q4_v1", + "classification": "infrastructure-invalid", + "reason_code": "harness-dispatch-keyerror-required-tool-argument", + "incident_document": "canonical-harness-incident-20260802.json", + "classified_before_grade": true, + "affirmative_evidence": [ + "a syntactically valid tool call omitted path", + "the harness raised KeyError instead of returning a recoverable tool error" + ], + "expected_files": { + "receipt.json": "92313dc9d6070ad4fba78fb89a4d043281e17d0e5fb2203e0bdc660c551669e4", + "transcript.jsonl": "ca8b0f10985799cbd7511a4ce32d6b51939f6d6f32139719cb41b9c52d18aeb4" + }, + "replacement": {"required": true, "canonical_run": "p1_testwrite_gemma4-31b-q4_v1", "status": "completed"} + }, + "p1_testwrite_gemma4-31b-q4_v3_retry-20260802T005418Z-ff262c4a": { + "source_run": "p1_testwrite_gemma4-31b-q4_v3", + "classification": "infrastructure-invalid", + "reason_code": "operator-interrupted-companion-after-harness-dispatch-defect", + "incident_document": "canonical-harness-incident-20260802.json", + "classified_before_grade": true, + "affirmative_evidence": [ + "the attempt was active under the same defective harness", + "the supervisor stopped before completion and the attempt was preserved without a grade" + ], + "expected_files": { + "receipt.json": "326934f8660018bdd081c3740222199dc6d5851d011489012ae93166c51440eb", + "transcript.jsonl": "63717381ad72db6d006a97dc24a1da14a557a92542826be3d92c264968556661" + }, + "replacement": {"required": true, "canonical_run": "p1_testwrite_gemma4-31b-q4_v3", "status": "completed"} + }, + "p1_bugfix_gemma4-31b-q4_v3_retry-20260802T004533Z-64195329": { + "source_run": "p1_bugfix_gemma4-31b-q4_v3", + "classification": "infrastructure-invalid", + "reason_code": "operator-interrupted-before-required-telemetry-gate", + "incident_document": "canonical-telemetry-gate-incident-20260802.json", + "classified_before_grade": true, + "affirmative_evidence": [ + "the retry began before per-run sidecar attribution was proven", + "the supervisor stopped before the attempt completed" + ], + "expected_files": { + "receipt.json": "c70a93f7255d8d4ef84311ed17187d84050c5e4faf4dbfd2768550e6773e760c", + "transcript.jsonl": "a07274f29bc3b36fe8c3259af0a8a4d53d1b7f8f6256e44dfc43c90b2b5ba7b2" + }, + "replacement": {"required": true, "canonical_run": "p1_bugfix_gemma4-31b-q4_v3", "status": "completed"} + }, + "p1_testwrite_gemma4-31b-q4_v1_retry-20260802T004533Z-09111a0e": { + "source_run": "p1_testwrite_gemma4-31b-q4_v1", + "classification": "infrastructure-invalid", + "reason_code": "operator-interrupted-before-required-telemetry-gate", + "incident_document": "canonical-telemetry-gate-incident-20260802.json", + "classified_before_grade": true, + "affirmative_evidence": [ + "the retry began before per-run sidecar attribution was proven", + "the supervisor stopped before the attempt completed" + ], + "expected_files": { + "receipt.json": "e163856856b1a8323e4b881d00a43f0a4361ad00f8d5e800acb3a6d435d259a5", + "transcript.jsonl": "e0efc1aebe693ac92e2a78354316059e5332e761979bdbee220a42019e0367f3" + }, + "replacement": {"required": true, "canonical_run": "p1_testwrite_gemma4-31b-q4_v1", "status": "completed"} + }, + "p1_refactor_gemma4-31b-q4_v3-server-timeout-20260802T020859Z": { + "source_run": "p1_refactor_gemma4-31b-q4_v3", + "classification": "infrastructure-invalid", + "reason_code": "server-transport-timeout-below-native-envelope", + "incident_document": "canonical-refactor-timeout-incident-20260802.json", + "classified_before_grade": true, + "affirmative_evidence": [ + "server task 80497 was cancelled exactly 3600 seconds after launch", + "the slot had decoded 133606 tokens, reported truncated=false, and remained below its 262144-token context boundary" + ], + "expected_files": { + "cost.json": "c5d1062f477758436ba0c3baf32df97fbe4e30df09c831bde470ccf938fce7f7", + "gpu_telemetry.json": "4f717a0e9d7298401c41e1799f772186a03a25bdf21bee93b30425cab3b31494", + "gpu_telemetry.reanalyzed-fdcc2496.json": "51eff413ad7a1a01bc6116a52e31ac3c39a1f844bf0aa0f0e09d9a0731392801", + "grade.json": "7f69334378ac3187de489eb74ad81c8245267d1a5711612b94879f297e758bcf", + "receipt.json": "58b9ee02a9c9cd03829a91d8f2f4d6d5eedf064f1c96206bc6c4b38522c4bfd8", + "summary.json": "0445d7ff95e0fbfbbc06d21cb3fdcccf8a741249ea294dda3f136a3ea4ed2b56", + "transcript.jsonl": "e887af90ea26038d3d8aadf22998eaa443eeb78eaa59d204bc7fcc741124e00e", + "workspace_final.tar.gz": "90a1b1918640b2a496b66ed5fbaf8649eb5902aa2e4e36c1f6b4debc71c720bc" + }, + "replacement": { + "required": true, + "canonical_run": "p1_refactor_gemma4-31b-q4_v3", + "status": "completed", + "receipt_sha256": "d8b8de86c6abb82b39c59fd2ed56571dc1c05e4c1cf2214fce1a54803bb20dca", + "summary_sha256": "e22036d8e57d95f93243ae7cd3f22cfc271686c8134c715dfc44b7459410fe61", + "server_timeout_seconds": 14400, + "finish_reason": "done_signal" + } + } + } +} diff --git a/tooling/deployments/gemma4-31b-q4-tower2/long_context_recall.py b/tooling/deployments/gemma4-31b-q4-tower2/long_context_recall.py new file mode 100755 index 00000000..e10049b1 --- /dev/null +++ b/tooling/deployments/gemma4-31b-q4-tower2/long_context_recall.py @@ -0,0 +1,179 @@ +#!/usr/bin/env python3 +"""Near-native-context start/middle/end recall gate for a served topology.""" + +from __future__ import annotations + +import argparse +import datetime as dt +import gzip +import hashlib +import json +import pathlib +import subprocess +import time +import urllib.error +import urllib.request +import uuid +from typing import Any + + +SAMPLING = {"temperature": 1.0, "top_p": 0.95, "top_k": 64} + + +def request_json( + base_url: str, path: str, payload: dict[str, Any], timeout: float +) -> tuple[int, bytes]: + request = urllib.request.Request( + base_url.rstrip("/") + path, + data=json.dumps(payload, separators=(",", ":")).encode(), + headers={"Content-Type": "application/json"}, + method="POST", + ) + try: + with urllib.request.urlopen(request, timeout=timeout) as response: + return response.status, response.read() + except urllib.error.HTTPError as exc: + return exc.code, exc.read() + + +def tokenize(base_url: str, content: str, timeout: float) -> int: + status, body = request_json( + base_url, "/tokenize", {"content": content, "add_special": False}, timeout + ) + if status != 200: + raise RuntimeError(f"Tokenize failed with HTTP {status}: {body[:1000]!r}") + return len(json.loads(body)["tokens"]) + + +def make_content( + base_url: str, target_tokens: int, markers: dict[str, str], timeout: float +) -> tuple[str, int]: + fixed = ( + "This is a long-context retrieval test. Memorize all three markers.\n" + f"START_MARKER={markers['start']}\n" + "{FIRST_FILLER}\n" + f"MIDDLE_MARKER={markers['middle']}\n" + "{SECOND_FILLER}\n" + f"END_MARKER={markers['end']}\n" + "Return one compact JSON object with keys start, middle, and end whose values are the exact markers." + ) + units = target_tokens + content = "" + actual = 0 + for _ in range(4): + first = " x" * (units // 2) + second = " y" * (units - units // 2) + content = fixed.replace("{FIRST_FILLER}", first).replace("{SECOND_FILLER}", second) + actual = tokenize(base_url, content, timeout) + if abs(actual - target_tokens) <= 2: + break + units = max(1, round(units * target_tokens / actual)) + return content, actual + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("--endpoint", required=True) + parser.add_argument("--model", default="Gemma-4-31B-it-QAT-Q4_0") + parser.add_argument("--label", required=True) + parser.add_argument("--output-root", required=True, type=pathlib.Path) + parser.add_argument("--target-content-tokens", type=int, default=245000) + parser.add_argument("--max-output-tokens", type=int, default=2048) + parser.add_argument("--timeout", type=float, default=3600) + args = parser.parse_args() + + timestamp = dt.datetime.now(dt.timezone.utc).strftime("%Y%m%dT%H%M%SZ") + out_dir = args.output_root / f"{args.label}-{timestamp}" + out_dir.mkdir(parents=True, exist_ok=False) + run_id = uuid.uuid4().hex + markers = { + "start": f"S_{run_id}_7C19", + "middle": f"M_{run_id}_4A62", + "end": f"E_{run_id}_9F35", + } + content, actual_content_tokens = make_content( + args.endpoint, args.target_content_tokens, markers, args.timeout + ) + payload = { + "model": args.model, + "messages": [{"role": "user", "content": content}], + **SAMPLING, + "max_tokens": args.max_output_tokens, + "stream": False, + } + payload_bytes = json.dumps(payload, separators=(",", ":")).encode() + with gzip.open(out_dir / "request.json.gz", "wb", compresslevel=9) as handle: + handle.write(payload_bytes) + + telemetry_path = out_dir / "nvidia-smi.csv" + telemetry_handle = telemetry_path.open("wb") + query = ( + "timestamp,index,uuid,memory.used,utilization.gpu,utilization.memory," + "power.draw,power.limit,temperature.gpu,clocks.sm,clocks.mem" + ) + sampler = subprocess.Popen( + [ + "nvidia-smi", + f"--query-gpu={query}", + "--format=csv,noheader,nounits", + "--loop-ms=200", + ], + stdout=telemetry_handle, + stderr=subprocess.STDOUT, + ) + started = time.perf_counter() + try: + status, response_body = request_json( + args.endpoint, "/v1/chat/completions", payload, args.timeout + ) + finally: + wall_seconds = time.perf_counter() - started + sampler.terminate() + try: + sampler.wait(timeout=5) + except subprocess.TimeoutExpired: + sampler.kill() + sampler.wait(timeout=5) + telemetry_handle.close() + + (out_dir / "response.json").write_bytes(response_body) + try: + response = json.loads(response_body) + except json.JSONDecodeError as exc: + raise RuntimeError(f"Non-JSON response, HTTP {status}: {response_body[:1000]!r}") from exc + content_out = response.get("choices", [{}])[0].get("message", {}).get("content", "") or "" + marker_results = {key: value in content_out for key, value in markers.items()} + summary = { + "schema_version": 1, + "timestamp": timestamp, + "label": args.label, + "endpoint": args.endpoint, + "model": args.model, + "sampling": SAMPLING, + "target_content_tokens": args.target_content_tokens, + "actual_content_tokens": actual_content_tokens, + "max_output_tokens": args.max_output_tokens, + "request_sha256": hashlib.sha256(payload_bytes).hexdigest(), + "http_status": status, + "wall_seconds": wall_seconds, + "finish_reason": response.get("choices", [{}])[0].get("finish_reason"), + "usage": response.get("usage"), + "markers": markers, + "marker_results": marker_results, + "error": response.get("error"), + "passed": status == 200 and all(marker_results.values()) and not response.get("error"), + } + (out_dir / "summary.json").write_text( + json.dumps(summary, indent=2, sort_keys=True) + "\n", encoding="utf-8" + ) + sums = [] + for path in sorted(out_dir.iterdir()): + if path.is_file() and path.name != "SHA256SUMS": + sums.append(f"{hashlib.sha256(path.read_bytes()).hexdigest()} {path.name}") + (out_dir / "SHA256SUMS").write_text("\n".join(sums) + "\n", encoding="utf-8") + print(json.dumps(summary, indent=2)) + return 0 if summary["passed"] else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/tooling/deployments/gemma4-31b-q4-tower2/mmbt-gemma4-extended@.service b/tooling/deployments/gemma4-31b-q4-tower2/mmbt-gemma4-extended@.service new file mode 100644 index 00000000..bafbe5c3 --- /dev/null +++ b/tooling/deployments/gemma4-31b-q4-tower2/mmbt-gemma4-extended@.service @@ -0,0 +1,15 @@ +[Unit] +Description=Gemma 4 31B Q4 extended MMBT lane %i +After=mmbt-gemma4-power-logger.service +Wants=mmbt-gemma4-power-logger.service + +[Service] +Type=simple +WorkingDirectory=/home/michael/bench-gemma4-31b-q4 +EnvironmentFile=/home/michael/bench-gemma4-31b-q4/tooling/deployments/gemma4-31b-q4-tower2/extended-lane%i.env +ExecStart=/usr/bin/python3 /home/michael/bench-gemma4-31b-q4/tooling/run_gemma4_extended_suites.py +Restart=on-failure +RestartSec=30s + +[Install] +WantedBy=default.target diff --git a/tooling/deployments/gemma4-31b-q4-tower2/mmbt-gemma4-power-logger.service b/tooling/deployments/gemma4-31b-q4-tower2/mmbt-gemma4-power-logger.service new file mode 100644 index 00000000..8088dbec --- /dev/null +++ b/tooling/deployments/gemma4-31b-q4-tower2/mmbt-gemma4-power-logger.service @@ -0,0 +1,12 @@ +[Unit] +Description=Gemma 4 31B Q4 five-second per-replica MMBT telemetry +After=network-online.target + +[Service] +Type=simple +ExecStart=/usr/bin/python3 /home/michael/bench-gemma4-31b-q4/tooling/deployments/gemma4-31b-q4-tower2/gemma4_gpu_telemetry.py --output /home/michael/gemma4-campaign-state/telemetry/gemma4-31b-q4-gpu.csv --interval 5 +Restart=always +RestartSec=10s + +[Install] +WantedBy=default.target diff --git a/tooling/deployments/gemma4-31b-q4-tower2/mmbt-gemma4-telemetry-sidecar.service b/tooling/deployments/gemma4-31b-q4-tower2/mmbt-gemma4-telemetry-sidecar.service new file mode 100644 index 00000000..a7e5ba86 --- /dev/null +++ b/tooling/deployments/gemma4-31b-q4-tower2/mmbt-gemma4-telemetry-sidecar.service @@ -0,0 +1,13 @@ +[Unit] +Description=Gemma 4 31B Q4 per-run telemetry and cost artifacts +After=mmbt-gemma4-power-logger.service +Wants=mmbt-gemma4-power-logger.service + +[Service] +Type=simple +ExecStart=/usr/bin/bash /home/michael/bench-gemma4-31b-q4/tooling/deployments/gemma4-31b-q4-tower2/gemma4-telemetry-sidecar.sh +Restart=always +RestartSec=30s + +[Install] +WantedBy=default.target diff --git a/tooling/deployments/gemma4-31b-q4-tower2/mmbt-gemma4@.service b/tooling/deployments/gemma4-31b-q4-tower2/mmbt-gemma4@.service new file mode 100644 index 00000000..6f9aa7d2 --- /dev/null +++ b/tooling/deployments/gemma4-31b-q4-tower2/mmbt-gemma4@.service @@ -0,0 +1,20 @@ +[Unit] +Description=MMBT Gemma 4 31B Q4 server (%i) +After=network-online.target +Wants=network-online.target + +[Service] +Type=simple +WorkingDirectory=/home/michael/bench-gemma4-31b-q4 +EnvironmentFile=/home/michael/bench-gemma4-31b-q4/tooling/deployments/gemma4-31b-q4-tower2/topologies/%i.env +ExecStart=/home/michael/bench-gemma4-31b-q4/tooling/deployments/gemma4-31b-q4-tower2/run-gemma4-server.sh +Restart=on-failure +RestartSec=5 +TimeoutStartSec=900 +TimeoutStopSec=180 +KillSignal=SIGINT +LimitMEMLOCK=infinity +LimitNOFILE=1048576 + +[Install] +WantedBy=default.target diff --git a/tooling/deployments/gemma4-31b-q4-tower2/model-manifest.json b/tooling/deployments/gemma4-31b-q4-tower2/model-manifest.json new file mode 100644 index 00000000..87d39ddf --- /dev/null +++ b/tooling/deployments/gemma4-31b-q4-tower2/model-manifest.json @@ -0,0 +1,75 @@ +{ + "schema_version": 1, + "campaign": "gemma4-31b-q4-mmbt", + "captured_at": "2026-08-01", + "source_repository": "google/gemma-4-31B-it-qat-q4_0-gguf", + "source_revision": "59dde24573e7e61570dba08b18a2e1fe246955ed", + "base_model": "google/gemma-4-31B-it-qat-q4_0-unquantized", + "served_name": "Gemma-4-31B-it-QAT-Q4_0", + "model": { + "path": "/mnt/bulk/models/google-gemma-4-31B-it-QAT-Q4_0-GGUF/gemma-4-31B_q4_0-it.gguf", + "bytes": 17651001568, + "sha256": "179cfb99212709597eae5929112cfca677e1bbf566178b479ae1da0c4772874b" + }, + "multimodal_projector": { + "path": "/mnt/bulk/models/google-gemma-4-31B-it-QAT-Q4_0-GGUF/gemma-4-31B-it-mmproj.gguf", + "bytes": 1200726368, + "sha256": "6bd60bdb958548b4093196d38744b0f2290c12503a3fddd7486bffa9c5eb07a4" + }, + "model_card": { + "parameters": 30700000000, + "architecture": "dense", + "quantization": "QAT Q4_0 GGUF", + "native_context_tokens": 262144, + "temperature": 1.0, + "top_p": 0.95, + "top_k": 64, + "modalities": ["text", "image"] + }, + "tower2": { + "gpus": "2x NVIDIA RTX PRO 6000 Blackwell Workstation Edition", + "gpu_memory_mib_each": 97887, + "gpu_power_limit_w_each": 500, + "mmbt_base_commit": "dcd9431d82168a17f039de084dce1a46ce3cc01a" + }, + "runtime_candidates": [ + { + "kind": "llama.cpp-current-source", + "commit": "11924d4c17abc27383376a1ac6a24fa3e36c1c0c", + "version": 10223, + "status": "built_pending_model_validation", + "source_path": "/home/michael/llama.cpp-gemma4-11924d4", + "build_path": "/home/michael/llama.cpp-gemma4-11924d4/build-cuda-tower2", + "toolchain": { + "host_compiler": "GNU 13.3.0", + "cuda": "13.1.115", + "cuda_architectures": "120a-real", + "build_type": "Release", + "cuda_flash_attention": true, + "cuda_graphs": true, + "nccl_requested": true, + "nccl_found": true, + "nccl_version": "2.29.7", + "nccl_root": "/home/michael/llama.cpp-gemma4-runtime-deps/nccl", + "nccl_header_sha256": "fc9720fe6d8b45a62775a2f9d13c43ed6e88b0e01fd771081d6f665e6df94bfb", + "nccl_library_sha256": "aa957cdfb91b516eae0d54a28e9ee5db52730d02e0ab45580efc3c19a68327a4" + }, + "binaries": { + "llama-server": "200b403b5735418ff1f6da0cea1938e413e11869ae362e8044a12b0df04622fc", + "llama-cli": "53186d087d962fd22d5c575351e8c95d40b8ab01aeb4a3c351f93897c8970e3d", + "llama-bench": "0cdacab3ebecec7e0d9811b66b81401a9b0ffd7c5027b6df8cc43179133e1a47" + }, + "notes": [ + "Built natively for Blackwell sm_120a from the pinned source commit.", + "NCCL 2.29.7 was copied from the pinned production DeepSeek image into an isolated runtime-dependency directory and linked through an embedded RUNPATH; no system package was changed.", + "Independent single-GPU replicas do not depend on NCCL, but the private NCCL runtime removes a known handicap from split-GPU candidates." + ] + }, + { + "kind": "llama.cpp-container-fallback", + "image": "ghcr.io/ggml-org/llama.cpp@sha256:9b17daa3579c05543a0e7c87f3af92f05434b9f3d1608b45ae9ec855130eb8ac", + "build": 9641, + "status": "predates_qat_upload_revalidate_before_use" + } + ] +} diff --git a/tooling/deployments/gemma4-31b-q4-tower2/openai_contract_probe.py b/tooling/deployments/gemma4-31b-q4-tower2/openai_contract_probe.py new file mode 100755 index 00000000..60775999 --- /dev/null +++ b/tooling/deployments/gemma4-31b-q4-tower2/openai_contract_probe.py @@ -0,0 +1,203 @@ +#!/usr/bin/env python3 +"""Validate OpenAI-compatible chat and native tool-call behavior for Gemma.""" + +from __future__ import annotations + +import argparse +import datetime as dt +import hashlib +import json +import pathlib +import urllib.error +import urllib.request +import uuid +from typing import Any + + +SAMPLING = {"temperature": 1.0, "top_p": 0.95, "top_k": 64} + + +def post(base_url: str, payload: dict[str, Any], timeout: float) -> tuple[int, bytes]: + request = urllib.request.Request( + base_url.rstrip("/") + "/v1/chat/completions", + data=json.dumps(payload, separators=(",", ":")).encode(), + headers={"Content-Type": "application/json"}, + method="POST", + ) + try: + with urllib.request.urlopen(request, timeout=timeout) as response: + return response.status, response.read() + except urllib.error.HTTPError as exc: + return exc.code, exc.read() + + +def write_json(path: pathlib.Path, value: Any) -> None: + path.write_text(json.dumps(value, indent=2, sort_keys=True) + "\n", encoding="utf-8") + + +def decode(status: int, body: bytes, label: str) -> dict[str, Any]: + try: + parsed = json.loads(body) + except json.JSONDecodeError as exc: + raise RuntimeError(f"{label}: non-JSON HTTP {status}: {body[:1000]!r}") from exc + return parsed + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("--endpoint", required=True) + parser.add_argument("--model", default="Gemma-4-31B-it-QAT-Q4_0") + parser.add_argument("--label", required=True) + parser.add_argument("--output-root", required=True, type=pathlib.Path) + parser.add_argument("--timeout", type=float, default=900) + parser.add_argument("--seed", type=int, default=424242) + args = parser.parse_args() + + timestamp = dt.datetime.now(dt.timezone.utc).strftime("%Y%m%dT%H%M%SZ") + out_dir = args.output_root / f"{args.label}-{timestamp}" + out_dir.mkdir(parents=True, exist_ok=False) + marker = f"MMBT_CONTRACT_{uuid.uuid4().hex}" + + base = { + "model": args.model, + **SAMPLING, + "seed": args.seed, + "max_tokens": 1024, + "stream": False, + } + chat_request = { + **base, + "messages": [ + {"role": "system", "content": "Follow the user's exact response-format request."}, + {"role": "user", "content": f"Reply with exactly {marker} and nothing else."}, + ], + } + chat_status, chat_body = post(args.endpoint, chat_request, args.timeout) + (out_dir / "chat.response.json").write_bytes(chat_body) + chat = decode(chat_status, chat_body, "chat") + chat_content = chat.get("choices", [{}])[0].get("message", {}).get("content", "") + chat_passed = chat_status == 200 and marker in (chat_content or "") + + tool_name = "return_benchmark_marker" + tool_request = { + **base, + "messages": [ + { + "role": "user", + "content": f"Call {tool_name} once with marker set exactly to {marker}. Do not answer in prose.", + } + ], + "tools": [ + { + "type": "function", + "function": { + "name": tool_name, + "description": "Return the exact benchmark marker supplied by the user.", + "parameters": { + "type": "object", + "properties": {"marker": {"type": "string"}}, + "required": ["marker"], + "additionalProperties": False, + }, + }, + } + ], + "tool_choice": "required", + } + tool_status, tool_body = post(args.endpoint, tool_request, args.timeout) + (out_dir / "tool-call.response.json").write_bytes(tool_body) + tool_response = decode(tool_status, tool_body, "tool-call") + assistant_message = tool_response.get("choices", [{}])[0].get("message", {}) + tool_calls = assistant_message.get("tool_calls") or [] + parsed_calls: list[dict[str, Any]] = [] + for call in tool_calls: + function = call.get("function") or {} + arguments = function.get("arguments") or "{}" + try: + parsed_arguments = json.loads(arguments) if isinstance(arguments, str) else arguments + except json.JSONDecodeError: + parsed_arguments = {"_unparseable": arguments} + parsed_calls.append( + {"id": call.get("id"), "name": function.get("name"), "arguments": parsed_arguments} + ) + tool_passed = ( + tool_status == 200 + and len(parsed_calls) == 1 + and parsed_calls[0]["name"] == tool_name + and parsed_calls[0]["arguments"].get("marker") == marker + ) + + followup_passed = False + followup: dict[str, Any] | None = None + followup_status: int | None = None + if tool_passed: + call_id = parsed_calls[0]["id"] + followup_request = { + **base, + "messages": tool_request["messages"] + + [ + assistant_message, + { + "role": "tool", + "tool_call_id": call_id, + "content": json.dumps({"marker": marker}), + }, + { + "role": "user", + "content": f"Now reply with exactly {marker} and nothing else.", + }, + ], + "tools": tool_request["tools"], + "tool_choice": "auto", + } + followup_status, followup_body = post(args.endpoint, followup_request, args.timeout) + (out_dir / "tool-followup.response.json").write_bytes(followup_body) + followup = decode(followup_status, followup_body, "tool-followup") + followup_content = followup.get("choices", [{}])[0].get("message", {}).get("content", "") + followup_passed = followup_status == 200 and marker in (followup_content or "") + + summary = { + "schema_version": 1, + "timestamp": timestamp, + "label": args.label, + "endpoint": args.endpoint, + "model": args.model, + "sampling": SAMPLING, + "seed": args.seed, + "marker": marker, + "chat": { + "http_status": chat_status, + "finish_reason": chat.get("choices", [{}])[0].get("finish_reason"), + "error": chat.get("error"), + "passed": chat_passed, + }, + "tool_call": { + "http_status": tool_status, + "finish_reason": tool_response.get("choices", [{}])[0].get("finish_reason"), + "error": tool_response.get("error"), + "parsed_calls": parsed_calls, + "passed": tool_passed, + }, + "tool_followup": { + "attempted": followup is not None, + "http_status": followup_status, + "finish_reason": None + if followup is None + else followup.get("choices", [{}])[0].get("finish_reason"), + "error": None if followup is None else followup.get("error"), + "passed": followup_passed, + }, + "passed": chat_passed and tool_passed and followup_passed, + } + write_json(out_dir / "summary.json", summary) + sums = [] + for path in sorted(out_dir.iterdir()): + if path.is_file() and path.name != "SHA256SUMS": + sums.append(f"{hashlib.sha256(path.read_bytes()).hexdigest()} {path.name}") + (out_dir / "SHA256SUMS").write_text("\n".join(sums) + "\n", encoding="utf-8") + print(json.dumps(summary, indent=2)) + return 0 if summary["passed"] else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/tooling/deployments/gemma4-31b-q4-tower2/pixel-production-runaway-20260802.json b/tooling/deployments/gemma4-31b-q4-tower2/pixel-production-runaway-20260802.json new file mode 100644 index 00000000..b37c6e45 --- /dev/null +++ b/tooling/deployments/gemma4-31b-q4-tower2/pixel-production-runaway-20260802.json @@ -0,0 +1,67 @@ +{ + "schema_version": 1, + "campaign": "gemma4-31b-q4-mmbt", + "detected_at": "2026-08-02T02:01:00Z", + "classification": "production-runaway-generation", + "scope": "extended operational reliability evidence; not part of the canonical MMBT quality score", + "agent": "pixel", + "source": { + "job_id": "e761638a-e72e-4234-a324-93107adcb83b", + "job_description": "Initiative check-in (6h)", + "task_id": "887c3aca-c065-4515-8818-d0e89517af7a", + "session_id": "887d0e23-96d9-4c0b-a14c-392e91ed0155", + "server_task_id": 144620, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "endpoint_port": 8001 + }, + "observation": { + "task_started_at": "2026-08-02T01:20:00.020Z", + "model_call_started_at": "2026-08-02T01:21:01Z", + "cancelled_at": "2026-08-02T02:01:38.705Z", + "task_duration_s": 2498.685, + "approximate_model_call_duration_s": 2437, + "decoded_tokens_at_cancel": 81753, + "slot_total_tokens_at_cancel": 101091, + "configured_max_output_tokens": 238108, + "average_decode_tokens_per_s": 33.58, + "persisted_assistant_turn": false, + "condition": "the production request streamed continuously for approximately 40 minutes and remained far below the configured native-context output ceiling", + "attribution": "real production request with model-control failure; not benchmark residue, queue contamination, or an infrastructure timeout" + }, + "operator_action": { + "command": "/home/michael/.npm-global/bin/openclaw tasks cancel 887c3aca-c065-4515-8818-d0e89517af7a", + "result": "cancelled", + "ended_at": "2026-08-02T02:01:38.705Z", + "error": "Cancelled by operator", + "effect": "the GPU 1 slot became idle immediately and board power fell from approximately 500 W to approximately 110 W" + }, + "recovery": { + "checked_at": "2026-08-02T02:02:14Z", + "session_id": "7136b04a-1932-42e8-a50b-cb7c6dac80bf", + "marker": "MMBT_PIXEL_RUNAWAY_RECOVERY_20260802T020214Z", + "status": "ok", + "summary": "completed", + "wall_s": 16.1, + "model_fetch_s": 4.062, + "finish_reason": "stop", + "output_tokens": 105, + "context_window": 262144, + "aborted_last_run": false + }, + "external_evidence": { + "root": "/home/michael/gemma4-campaign-state/production-incidents/pixel-initiative-runaway-20260802T012101Z", + "privacy": "metadata-only bundle; user content redacted", + "files": { + "gateway-session.log": "5f40f1a9e79876ddc5f8f477fe7a6d065d8312e5bb738618939955647e0bca4b", + "recovery-session.json": "229c779edc8d226f3f344f7d2005bced89a65265163c1340811328183f60175f", + "server-task.log": "245e52ccb50e72055908f3679ddfa97cf5a7d3b88663355c8b636e507bff4551", + "session-metadata.json": "b857df09849c241daef64e6c192ffba4c5567501d23690b858e92f7d79293a42", + "task.json": "138629757d4b14686f8258004db76342e098ae49d726cd709dbad17f2dbf6c57" + } + }, + "interpretation": { + "quality_scoring": "do not count this event as a canonical task failure", + "operational_finding": "native-context serving needs explicit production stop governance because a coherent request can consume tens of thousands of tokens without reaching a natural stop", + "campaign_follow_up": "retain the full safe model envelope for benchmark fairness while documenting and testing production-specific output and cancellation controls" + } +} diff --git a/tooling/deployments/gemma4-31b-q4-tower2/production-restore.json b/tooling/deployments/gemma4-31b-q4-tower2/production-restore.json new file mode 100644 index 00000000..d38caa93 --- /dev/null +++ b/tooling/deployments/gemma4-31b-q4-tower2/production-restore.json @@ -0,0 +1,45 @@ +{ + "schema_version": 1, + "captured_at": "2026-08-01", + "required_final_state": "Restore the proven DeepSeek production deployment after the Gemma campaign, then prove Sanctuary and Pixel health before campaign completion.", + "deepseek": { + "container_name": "deepseek-v4-flash-0731", + "served_model": "DeepSeek-V4-Flash-0731", + "endpoint": "http://127.0.0.1:8000/v1", + "launcher": "/home/michael/start-deepseek-v4-flash-0731.sh", + "launcher_sha256": "4fca6a4a478f0876a1e6d7241dd87a3c9574bd970f29943bb3696993d4bfe25d", + "image": "voipmonitor/vllm@sha256:48518e91cf87dd0c0483c76ff86e81dfc0f46de7e364b46f7a82c481ce08188f", + "restart_policy": "unless-stopped", + "gpu_power_limit_w_each": 500 + }, + "openclaw": { + "live_config": "/home/michael/.openclaw/openclaw.json", + "pre_campaign_backup": "/home/michael/gemma4-campaign-state/openclaw.pre-gemma.json", + "pre_campaign_sha256": "79f2872821d68f76a2c42cb8f3972ae494b380a4d0d058fba405dbcc5e489c7f", + "backup_mode": "0600", + "service": "openclaw-gateway.service", + "provider": "tower", + "provider_api": "openai-completions", + "provider_base_url": "http://127.0.0.1:8000/v1", + "default_primary": "tower/DeepSeek-V4-Flash-0731" + }, + "portals": { + "sanctuary": { + "container": "sanctuary-portal", + "openai_base_url": "http://127.0.0.1:18789/v1", + "default_model": "openclaw" + }, + "pixel": { + "container": "pixel-portal", + "openai_base_url": "http://127.0.0.1:18789/v1", + "default_model": "openclaw/pixel" + } + }, + "completion_gates": [ + "Stop and disable every Gemma campaign service.", + "Verify the live OpenClaw config has been restored byte-for-byte to the recorded pre-campaign SHA256.", + "Start DeepSeek only with the recorded launcher and verify the image digest, 500 W limits, model ID, and endpoint health.", + "Restart OpenClaw and prove real end-to-end Sanctuary and Pixel requests complete under DeepSeek.", + "Record post-restore hashes, timestamps, model responses, container/service health, and GPU telemetry in final-validation.json." + ] +} diff --git a/tooling/deployments/gemma4-31b-q4-tower2/prove-agent-health.sh b/tooling/deployments/gemma4-31b-q4-tower2/prove-agent-health.sh new file mode 100755 index 00000000..5c7dc15e --- /dev/null +++ b/tooling/deployments/gemma4-31b-q4-tower2/prove-agent-health.sh @@ -0,0 +1,66 @@ +#!/usr/bin/env bash +set -euo pipefail + +PHASE="${1:-}" +OUT_ROOT="${2:-/home/michael/gemma4-campaign-state/health}" +if [[ -z "$PHASE" ]]; then + printf 'Usage: %s PHASE [OUTPUT_ROOT]\n' "$0" >&2 + exit 64 +fi + +timestamp="$(date -u +%Y%m%dT%H%M%SZ)" +out_dir="$OUT_ROOT/$PHASE-$timestamp" +install -d -m 700 "$out_dir" +health_max_tokens="${HEALTH_MAX_TOKENS:-262144}" +health_timeout_seconds="${HEALTH_TIMEOUT_SECONDS:-180}" + +probe() { + local agent="$1" + local portal_container="$2" + local requested_model="$3" + local marker="MMBT_${PHASE}_${agent}_${timestamp}" + local api_key response http_code body_file summary_file + + api_key="$(docker inspect "$portal_container" | jq -r '.[0].Config.Env[]' | sed -n 's/^OPENAI_API_KEY=//p' | head -n 1)" + if [[ -z "$api_key" ]]; then + printf 'No portal API key found for %s\n' "$portal_container" >&2 + return 1 + fi + + body_file="$out_dir/${agent}.response.json" + summary_file="$out_dir/${agent}.summary.json" + response="$(curl --silent --show-error --max-time "$health_timeout_seconds" \ + --write-out $'\n%{http_code}' \ + --config <(printf 'header = "Authorization: Bearer %s"\nheader = "Content-Type: application/json"\n' "$api_key") \ + --data "$(jq -cn --arg model "$requested_model" --arg marker "$marker" --argjson max_tokens "$health_max_tokens" \ + '{model:$model,user:$marker,messages:[{role:"user",content:("Health check. Reply with exactly this marker and nothing else: " + $marker)}],temperature:0,max_tokens:$max_tokens,stream:false}')" \ + http://127.0.0.1:18789/v1/chat/completions)" + http_code="${response##*$'\n'}" + printf '%s' "${response%$'\n'*}" >"$body_file" + + jq -n \ + --arg phase "$PHASE" \ + --arg timestamp "$timestamp" \ + --arg agent "$agent" \ + --arg portal "$portal_container" \ + --arg requested_model "$requested_model" \ + --argjson max_tokens "$health_max_tokens" \ + --arg marker "$marker" \ + --arg http_code "$http_code" \ + --arg response_id "$(jq -r '.id // ""' "$body_file")" \ + --arg finish_reason "$(jq -r '.choices[0].finish_reason // ""' "$body_file")" \ + --arg content "$(jq -r '.choices[0].message.content // ""' "$body_file")" \ + --arg error "$(jq -r '.error.message // .detail // ""' "$body_file")" \ + --arg raw_sha256 "$(sha256sum "$body_file" | cut -d' ' -f1)" \ + '{phase:$phase,timestamp:$timestamp,agent:$agent,portal:$portal,requested_model:$requested_model,max_tokens:$max_tokens,http_code:($http_code|tonumber),response_id:$response_id,finish_reason:$finish_reason,content:$content,error:$error,raw_sha256:$raw_sha256,passed:(($http_code=="200") and ($content|contains($marker)) and ($error==""))}' \ + >"$summary_file" + + jq -c . "$summary_file" + jq -e '.passed == true' "$summary_file" >/dev/null +} + +probe sanctuary sanctuary-portal openclaw +probe pixel pixel-portal openclaw/pixel + +find "$out_dir" -maxdepth 1 -type f -print0 | sort -z | xargs -0 sha256sum >"$out_dir/SHA256SUMS" +printf '%s\n' "$out_dir" diff --git a/tooling/deployments/gemma4-31b-q4-tower2/queueing-analysis.json b/tooling/deployments/gemma4-31b-q4-tower2/queueing-analysis.json new file mode 100644 index 00000000..7df27b05 --- /dev/null +++ b/tooling/deployments/gemma4-31b-q4-tower2/queueing-analysis.json @@ -0,0 +1,96 @@ +{ + "schema_version": 1, + "generated_at": "2026-08-02T02:17:59.511743+00:00", + "source": { + "path": "/home/michael/gemma4-campaign-state/topology/replica-gpu1-q8-s4-concurrency-20260801T234354Z/summary.json", + "bytes": 35730, + "sha256": "98f977520b227f03a1abd720cfc21f6178f6c9b3a2ef4279ae9b85e2b77f9906" + }, + "operating_point": { + "parallel_slots": 4, + "simultaneous_requests": 8, + "waves": 2 + }, + "first_wave": { + "requests": 4, + "median_wall_s": 9.108222, + "median_wall_minus_server_work_s": 2.117381 + }, + "queued_second_wave": { + "requests": 4, + "median_wall_s": 18.183445, + "median_wall_minus_server_work_s": 11.209493 + }, + "derived": { + "second_wave_wall_penalty_s": 9.075223, + "estimated_queue_wait_delta_s": 9.092112, + "aggregate_decode_tokens_per_second_wall": 112.61659474507437 + }, + "requests": [ + { + "request_index": 0, + "wall_s": 9.108004, + "server_reported_prompt_plus_decode_s": 7.009431, + "wall_minus_server_reported_work_s": 2.098573, + "tokens_evaluated": 1025, + "tokens_predicted": 256 + }, + { + "request_index": 1, + "wall_s": 9.107791, + "server_reported_prompt_plus_decode_s": 7.040805, + "wall_minus_server_reported_work_s": 2.066986, + "tokens_evaluated": 1025, + "tokens_predicted": 256 + }, + { + "request_index": 2, + "wall_s": 18.183528, + "server_reported_prompt_plus_decode_s": 6.990353, + "wall_minus_server_reported_work_s": 11.193175, + "tokens_evaluated": 1025, + "tokens_predicted": 256 + }, + { + "request_index": 3, + "wall_s": 9.108829, + "server_reported_prompt_plus_decode_s": 6.972641, + "wall_minus_server_reported_work_s": 2.136188, + "tokens_evaluated": 1025, + "tokens_predicted": 256 + }, + { + "request_index": 4, + "wall_s": 9.10844, + "server_reported_prompt_plus_decode_s": 6.931966, + "wall_minus_server_reported_work_s": 2.176474, + "tokens_evaluated": 1025, + "tokens_predicted": 256 + }, + { + "request_index": 5, + "wall_s": 18.182898, + "server_reported_prompt_plus_decode_s": 7.026039, + "wall_minus_server_reported_work_s": 11.156859, + "tokens_evaluated": 1025, + "tokens_predicted": 256 + }, + { + "request_index": 6, + "wall_s": 18.183917, + "server_reported_prompt_plus_decode_s": 6.958106, + "wall_minus_server_reported_work_s": 11.225811, + "tokens_evaluated": 1025, + "tokens_predicted": 256 + }, + { + "request_index": 7, + "wall_s": 18.183362, + "server_reported_prompt_plus_decode_s": 6.921437, + "wall_minus_server_reported_work_s": 11.261925, + "tokens_evaluated": 1025, + "tokens_predicted": 256 + } + ], + "methodology": "The client released all requests together. With four server slots and eight requests, the four shortest walls are the first service wave and the four longest are the queued wave. estimated_queue_wait_delta_s subtracts each response's llama.cpp-reported prompt+decode work from client wall time, then differences the wave medians. It includes scheduler/HTTP overhead and is an estimate, not a direct server-side queue timestamp." +} diff --git a/tooling/deployments/gemma4-31b-q4-tower2/run-gemma4-server.sh b/tooling/deployments/gemma4-31b-q4-tower2/run-gemma4-server.sh new file mode 100755 index 00000000..bffd653a --- /dev/null +++ b/tooling/deployments/gemma4-31b-q4-tower2/run-gemma4-server.sh @@ -0,0 +1,98 @@ +#!/usr/bin/env bash +set -euo pipefail + +# Reproducible foreground launcher for one Gemma server process. A service +# manager or the topology harness owns backgrounding, logs, and restart policy. + +RUNTIME_ROOT="${GEMMA_RUNTIME_ROOT:-/home/michael/llama.cpp-gemma4-11924d4/build-cuda-tower2}" +SERVER_BIN="${GEMMA_SERVER_BIN:-$RUNTIME_ROOT/bin/llama-server}" +MODEL="${GEMMA_MODEL:-/mnt/bulk/models/google-gemma-4-31B-it-QAT-Q4_0-GGUF/gemma-4-31B_q4_0-it.gguf}" +MMPROJ="${GEMMA_MMPROJ:-/mnt/bulk/models/google-gemma-4-31B-it-QAT-Q4_0-GGUF/gemma-4-31B-it-mmproj.gguf}" + +export CUDA_DEVICE_ORDER=PCI_BUS_ID +export CUDA_VISIBLE_DEVICES="${GEMMA_CUDA_VISIBLE_DEVICES:-0}" +export GGML_CUDA_ENABLE_UNIFIED_MEMORY=0 + +HOST="${GEMMA_HOST:-127.0.0.1}" +PORT="${GEMMA_PORT:-8000}" +# The DeepSeek compatibility alias was used only during the safe first cutover. +# Once OpenClaw records Gemma's actual identity, expose only that identity. +ALIASES="${GEMMA_ALIASES:-Gemma-4-31B-it-QAT-Q4_0}" +CTX_SIZE="${GEMMA_CTX_SIZE:-262144}" +PARALLEL="${GEMMA_PARALLEL:-1}" +SPLIT_MODE="${GEMMA_SPLIT_MODE:-none}" +MAIN_GPU="${GEMMA_MAIN_GPU:-0}" +TENSOR_SPLIT="${GEMMA_TENSOR_SPLIT:-}" +CACHE_TYPE_K="${GEMMA_CACHE_TYPE_K:-f16}" +CACHE_TYPE_V="${GEMMA_CACHE_TYPE_V:-f16}" +BATCH_SIZE="${GEMMA_BATCH_SIZE:-2048}" +UBATCH_SIZE="${GEMMA_UBATCH_SIZE:-512}" +THREADS="${GEMMA_THREADS:-16}" +THREADS_BATCH="${GEMMA_THREADS_BATCH:-24}" +THREADS_HTTP="${GEMMA_THREADS_HTTP:-16}" +CACHE_REUSE="${GEMMA_CACHE_REUSE:-256}" +KV_UNIFIED="${GEMMA_KV_UNIFIED:-0}" +HTTP_TIMEOUT="${GEMMA_HTTP_TIMEOUT:-14400}" + +NATIVE_CONTEXT=262144 +required_ctx=$((PARALLEL * NATIVE_CONTEXT)) +if (( CTX_SIZE < required_ctx )); then + printf 'Refusing GEMMA_CTX_SIZE=%s with GEMMA_PARALLEL=%s: each slot must retain the full %s-token native context.\n' \ + "$CTX_SIZE" "$PARALLEL" "$NATIVE_CONTEXT" >&2 + exit 64 +fi +if ! [[ "$HTTP_TIMEOUT" =~ ^[1-9][0-9]*$ ]] || (( HTTP_TIMEOUT < 14400 )); then + printf 'Refusing GEMMA_HTTP_TIMEOUT=%s: native-context benchmark serving requires at least 14400 seconds.\n' \ + "$HTTP_TIMEOUT" >&2 + exit 64 +fi + +test -x "$SERVER_BIN" +test -r "$MODEL" +test -r "$MMPROJ" + +cmd=( + "$SERVER_BIN" + --model "$MODEL" + --mmproj "$MMPROJ" + --alias "$ALIASES" + --host "$HOST" + --port "$PORT" + --ctx-size "$CTX_SIZE" + --parallel "$PARALLEL" + --n-predict -1 + --n-gpu-layers all + --split-mode "$SPLIT_MODE" + --main-gpu "$MAIN_GPU" + --flash-attn on + --cache-type-k "$CACHE_TYPE_K" + --cache-type-v "$CACHE_TYPE_V" + --batch-size "$BATCH_SIZE" + --ubatch-size "$UBATCH_SIZE" + --threads "$THREADS" + --threads-batch "$THREADS_BATCH" + --threads-http "$THREADS_HTTP" + --timeout "$HTTP_TIMEOUT" + --cont-batching + --cache-prompt + --cache-reuse "$CACHE_REUSE" + --metrics + --slots + --jinja + --no-webui +) + +if [[ "$KV_UNIFIED" == 1 ]]; then + cmd+=(--kv-unified) +else + cmd+=(--no-kv-unified) +fi + +if [[ -n "$TENSOR_SPLIT" ]]; then + cmd+=(--tensor-split "$TENSOR_SPLIT") +fi + +printf 'Launching:' >&2 +printf ' %q' "${cmd[@]}" >&2 +printf '\n' >&2 +exec "${cmd[@]}" diff --git a/tooling/deployments/gemma4-31b-q4-tower2/server_microbench.py b/tooling/deployments/gemma4-31b-q4-tower2/server_microbench.py new file mode 100755 index 00000000..d4f146cc --- /dev/null +++ b/tooling/deployments/gemma4-31b-q4-tower2/server_microbench.py @@ -0,0 +1,373 @@ +#!/usr/bin/env python3 +"""Controlled llama.cpp HTTP performance probe with raw, attributable evidence.""" + +from __future__ import annotations + +import argparse +import concurrent.futures +import datetime as dt +import hashlib +import json +import os +import pathlib +import subprocess +import threading +import time +import urllib.error +import urllib.request +from typing import Any + + +SAMPLING = {"temperature": 1.0, "top_p": 0.95, "top_k": 64} + + +def utc_now() -> str: + return dt.datetime.now(dt.timezone.utc).isoformat() + + +def http_get(base_url: str, path: str, timeout: float = 30) -> tuple[int, bytes, dict[str, str]]: + req = urllib.request.Request(base_url.rstrip("/") + path) + try: + with urllib.request.urlopen(req, timeout=timeout) as response: + return response.status, response.read(), dict(response.headers.items()) + except urllib.error.HTTPError as exc: + return exc.code, exc.read(), dict(exc.headers.items()) + + +def http_post_json( + base_url: str, path: str, payload: dict[str, Any], timeout: float = 3600 +) -> tuple[int, bytes, dict[str, str]]: + body = json.dumps(payload, separators=(",", ":")).encode() + req = urllib.request.Request( + base_url.rstrip("/") + path, + data=body, + headers={"Content-Type": "application/json"}, + method="POST", + ) + try: + with urllib.request.urlopen(req, timeout=timeout) as response: + return response.status, response.read(), dict(response.headers.items()) + except urllib.error.HTTPError as exc: + return exc.code, exc.read(), dict(exc.headers.items()) + + +def sha256_bytes(data: bytes) -> str: + return hashlib.sha256(data).hexdigest() + + +def write_bytes(path: pathlib.Path, data: bytes) -> str: + path.write_bytes(data) + return sha256_bytes(data) + + +def write_json(path: pathlib.Path, value: Any) -> str: + data = (json.dumps(value, indent=2, sort_keys=True) + "\n").encode() + return write_bytes(path, data) + + +def get_json(base_url: str, path: str, timeout: float = 30) -> Any: + status, body, _ = http_get(base_url, path, timeout) + if status != 200: + raise RuntimeError(f"GET {path} returned HTTP {status}: {body[:500]!r}") + return json.loads(body) + + +def post_json(base_url: str, path: str, payload: dict[str, Any], timeout: float = 3600) -> Any: + status, body, _ = http_post_json(base_url, path, payload, timeout) + if status != 200: + raise RuntimeError(f"POST {path} returned HTTP {status}: {body[:1000]!r}") + return json.loads(body) + + +def build_prompt(base_url: str, requested_tokens: int, nonce: str) -> tuple[str, int]: + # The repeated unit is cheap to construct and deterministic. Calibrate by + # asking the served tokenizer; the benchmark records the resulting count. + units = max(1, requested_tokens) + prefix = f"MMBT synthetic prefill {nonce}. Preserve every token.\n" + for _ in range(3): + prompt = prefix + (" x" * units) + tokenized = post_json(base_url, "/tokenize", {"content": prompt, "add_special": False}) + count = len(tokenized["tokens"]) + if count == requested_tokens or count <= 0: + return prompt, count + units = max(1, round(units * requested_tokens / count)) + prompt = prefix + (" x" * units) + tokenized = post_json(base_url, "/tokenize", {"content": prompt, "add_special": False}) + return prompt, len(tokenized["tokens"]) + + +class NvidiaSampler: + def __init__(self, output_path: pathlib.Path, interval_ms: int = 200) -> None: + self.output_path = output_path + self.interval_ms = interval_ms + self.process: subprocess.Popen[bytes] | None = None + self.handle: Any = None + + def __enter__(self) -> "NvidiaSampler": + self.handle = self.output_path.open("wb") + query = ( + "timestamp,index,uuid,memory.used,utilization.gpu,utilization.memory," + "power.draw,power.limit,temperature.gpu,clocks.sm,clocks.mem" + ) + self.process = subprocess.Popen( + [ + "nvidia-smi", + f"--query-gpu={query}", + "--format=csv,noheader,nounits", + f"--loop-ms={self.interval_ms}", + ], + stdout=self.handle, + stderr=subprocess.STDOUT, + ) + return self + + def __exit__(self, exc_type: Any, exc: Any, traceback: Any) -> None: + if self.process is not None: + self.process.terminate() + try: + self.process.wait(timeout=5) + except subprocess.TimeoutExpired: + self.process.kill() + self.process.wait(timeout=5) + if self.handle is not None: + self.handle.close() + + +def stream_completion( + base_url: str, + payload: dict[str, Any], + raw_path: pathlib.Path, + timeout: float, +) -> dict[str, Any]: + request_payload = dict(payload) + request_payload["stream"] = True + body = json.dumps(request_payload, separators=(",", ":")).encode() + req = urllib.request.Request( + base_url.rstrip("/") + "/completion", + data=body, + headers={"Content-Type": "application/json"}, + method="POST", + ) + started = time.perf_counter() + first_content_at: float | None = None + events: list[dict[str, Any]] = [] + with urllib.request.urlopen(req, timeout=timeout) as response: + status = response.status + for raw_line in response: + line = raw_line.decode("utf-8", errors="replace").strip() + if not line.startswith("data: "): + continue + encoded = line[6:] + if encoded == "[DONE]": + continue + event = json.loads(encoded) + events.append(event) + if first_content_at is None and event.get("content"): + first_content_at = time.perf_counter() + ended = time.perf_counter() + write_json(raw_path, events) + final_event = events[-1] if events else {} + return { + "http_status": status, + "wall_seconds": ended - started, + "ttft_seconds": None if first_content_at is None else first_content_at - started, + "event_count": len(events), + "content_chars": sum(len(str(event.get("content", ""))) for event in events), + "timings": final_event.get("timings"), + "tokens_evaluated": final_event.get("tokens_evaluated"), + "tokens_predicted": final_event.get("tokens_predicted"), + "stop_type": final_event.get("stop_type"), + "stopped_eos": final_event.get("stopped_eos"), + "stopped_limit": final_event.get("stopped_limit"), + } + + +def nonstream_completion(base_url: str, payload: dict[str, Any], timeout: float) -> dict[str, Any]: + started = time.perf_counter() + status, body, headers = http_post_json(base_url, "/completion", payload, timeout) + ended = time.perf_counter() + parsed = json.loads(body) + return { + "http_status": status, + "wall_seconds": ended - started, + "headers": headers, + "body": parsed, + "body_sha256": sha256_bytes(body), + } + + +def completion_payload(prompt: str, n_predict: int, seed: int) -> dict[str, Any]: + return { + "prompt": prompt, + "n_predict": n_predict, + "ignore_eos": True, + "cache_prompt": False, + "seed": seed, + **SAMPLING, + } + + +def run_concurrency( + base_url: str, + output_dir: pathlib.Path, + concurrency: int, + prompt_tokens: int, + n_predict: int, + seed: int, + timeout: float, +) -> dict[str, Any]: + start_event = threading.Event() + work: list[tuple[dict[str, Any], int, int]] = [] + for index in range(concurrency): + prompt, actual = build_prompt(base_url, prompt_tokens, f"c{concurrency}-r{index}") + work.append((completion_payload(prompt, n_predict, seed + index), actual, index)) + + def worker(item: tuple[dict[str, Any], int, int]) -> dict[str, Any]: + payload, actual, index = item + start_event.wait() + result = nonstream_completion(base_url, payload, timeout) + body = result.pop("body") + write_json(output_dir / f"concurrency-{concurrency}-request-{index}.response.json", body) + result.update( + { + "request_index": index, + "request_payload_sha256": sha256_bytes( + json.dumps(payload, separators=(",", ":")).encode() + ), + "actual_prompt_tokens": actual, + "tokens_evaluated": body.get("tokens_evaluated"), + "tokens_predicted": body.get("tokens_predicted"), + "timings": body.get("timings"), + "stop_type": body.get("stop_type"), + "stopped_limit": body.get("stopped_limit"), + "error": body.get("error"), + } + ) + return result + + batch_started = time.perf_counter() + with concurrent.futures.ThreadPoolExecutor(max_workers=concurrency) as executor: + futures = [executor.submit(worker, item) for item in work] + start_event.set() + requests = [future.result() for future in futures] + batch_wall = time.perf_counter() - batch_started + total_predicted = sum(int(item.get("tokens_predicted") or 0) for item in requests) + return { + "concurrency": concurrency, + "batch_wall_seconds": batch_wall, + "aggregate_predicted_tokens": total_predicted, + "aggregate_decode_tokens_per_second_wall": total_predicted / batch_wall, + "requests": requests, + } + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("--endpoint", required=True) + parser.add_argument("--label", required=True) + parser.add_argument("--output-root", required=True, type=pathlib.Path) + parser.add_argument("--prompt-tokens", default="1024,32768,131072,250000") + parser.add_argument("--stream-predict", type=int, default=256) + parser.add_argument("--concurrency", default="1,2,4,8") + parser.add_argument("--concurrency-prompt-tokens", type=int, default=1024) + parser.add_argument("--concurrency-predict", type=int, default=256) + parser.add_argument("--seed", type=int, default=424242) + parser.add_argument("--timeout", type=float, default=3600) + args = parser.parse_args() + + timestamp = dt.datetime.now(dt.timezone.utc).strftime("%Y%m%dT%H%M%SZ") + output_dir = args.output_root / f"{args.label}-{timestamp}" + output_dir.mkdir(parents=True, exist_ok=False) + + health_status, health_body, _ = http_get(args.endpoint, "/health", 10) + if health_status != 200: + raise RuntimeError(f"Endpoint is not healthy: HTTP {health_status} {health_body!r}") + + metadata: dict[str, Any] = { + "schema_version": 1, + "label": args.label, + "endpoint": args.endpoint, + "started_at": utc_now(), + "sampling": SAMPLING, + "seed": args.seed, + "model_list": get_json(args.endpoint, "/v1/models"), + "slots_before": get_json(args.endpoint, "/slots"), + "arguments": vars(args) | {"output_root": str(args.output_root)}, + } + metrics_status, metrics_before, _ = http_get(args.endpoint, "/metrics") + if metrics_status == 200: + write_bytes(output_dir / "metrics-before.prom", metrics_before) + + prompt_results: list[dict[str, Any]] = [] + concurrency_results: list[dict[str, Any]] = [] + with NvidiaSampler(output_dir / "nvidia-smi.csv"): + for index, requested in enumerate(int(value) for value in args.prompt_tokens.split(",") if value): + prompt, actual = build_prompt(args.endpoint, requested, f"p{requested}") + payload = completion_payload(prompt, args.stream_predict, args.seed + index) + raw_path = output_dir / f"prefill-{requested}-stream-events.json" + result = stream_completion(args.endpoint, payload, raw_path, args.timeout) + result.update({"requested_prompt_tokens": requested, "actual_prompt_tokens": actual}) + prompt_results.append(result) + write_json(output_dir / f"prefill-{requested}-summary.json", result) + + for concurrency in (int(value) for value in args.concurrency.split(",") if value): + result = run_concurrency( + args.endpoint, + output_dir, + concurrency, + args.concurrency_prompt_tokens, + args.concurrency_predict, + args.seed + 1000 + concurrency * 10, + args.timeout, + ) + concurrency_results.append(result) + write_json(output_dir / f"concurrency-{concurrency}.json", result) + + metrics_status, metrics_after, _ = http_get(args.endpoint, "/metrics") + if metrics_status == 200: + write_bytes(output_dir / "metrics-after.prom", metrics_after) + metadata.update( + { + "finished_at": utc_now(), + "prompt_results": prompt_results, + "concurrency_results": concurrency_results, + "slots_after": get_json(args.endpoint, "/slots"), + } + ) + write_json(output_dir / "summary.json", metadata) + + sums: list[str] = [] + for path in sorted(output_dir.iterdir()): + if path.is_file() and path.name != "SHA256SUMS": + sums.append(f"{hashlib.sha256(path.read_bytes()).hexdigest()} {path.name}") + (output_dir / "SHA256SUMS").write_text("\n".join(sums) + "\n", encoding="utf-8") + + compact = { + "output_dir": str(output_dir), + "label": args.label, + "prompt_results": [ + { + "prompt_tokens": item["actual_prompt_tokens"], + "ttft_seconds": item["ttft_seconds"], + "wall_seconds": item["wall_seconds"], + "timings": item["timings"], + } + for item in prompt_results + ], + "concurrency": [ + { + "concurrency": item["concurrency"], + "batch_wall_seconds": item["batch_wall_seconds"], + "aggregate_decode_tokens_per_second_wall": item[ + "aggregate_decode_tokens_per_second_wall" + ], + } + for item in concurrency_results + ], + } + print(json.dumps(compact, indent=2)) + return 0 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/tooling/deployments/gemma4-31b-q4-tower2/test_gemma4_telemetry.py b/tooling/deployments/gemma4-31b-q4-tower2/test_gemma4_telemetry.py new file mode 100644 index 00000000..91dbbc26 --- /dev/null +++ b/tooling/deployments/gemma4-31b-q4-tower2/test_gemma4_telemetry.py @@ -0,0 +1,88 @@ +#!/usr/bin/env python3 +import csv +import importlib.util +import json +from pathlib import Path + + +DEPLOY = Path(__file__).parent + + +def load(name, filename): + spec = importlib.util.spec_from_file_location(name, DEPLOY / filename) + module = importlib.util.module_from_spec(spec) + assert spec.loader is not None + spec.loader.exec_module(module) + return module + + +LOGGER = load("gemma4_gpu_telemetry", "gemma4_gpu_telemetry.py") +ANALYZER = load("analyze_replica_telemetry", "analyze_replica_telemetry.py") + + +def test_rapl_power_handles_counter_wrap(): + power, state = LOGGER.rapl_power((990, 10.0), 20, 12.0, 1000) + assert power == 30 / 1_000_000 / 2 + assert state == (20, 12.0) + + +def test_replica_analyzer_attributes_only_the_active_lane(tmp_path): + logs = tmp_path / "logs" + run = logs / "p2_extract_gemma4-31b-q4_v1" + run.mkdir(parents=True) + (run / "summary.json").write_text(json.dumps({ + "started_at": "2026-08-02T00:00:00+00:00", + "ended_at": "2026-08-02T00:00:10+00:00", + "elapsed_s": 10, + })) + csv_path = tmp_path / "gpu.csv" + with csv_path.open("w", newline="") as handle: + writer = csv.writer(handle) + writer.writerow(LOGGER.FIELDS) + for second in (0, 5, 10): + ts = f"2026-08-02T00:00:{second:02d}Z" + writer.writerow([ts, 0, 8000, 500, 400, 95, 20, 65000, 70, 2700, + run.name, 123, 100]) + writer.writerow([ts, 1, 8001, 500, 300, 80, 15, 65000, 65, 2600, + "other_cell", 456, 100]) + report = ANALYZER.analyze(csv_path, [logs], 500, 0.13, 15) + doc = report[run.name] + assert doc["attribution"]["active_gpu_ids_observed"] == ["0"] + assert doc["active_gpu"]["mean_power_w"] == 400 + assert doc["sampling"]["coverage_fraction_of_wall"] == 1.0 + assert doc["concurrency"]["other_gpu_cell_sample_counts"] == {"other_cell": 3} + assert doc["concurrency"]["mean_observed_two_gpu_plus_cpu_package_w"] == 800 + assert doc["sampling"]["window_source"] == "summary.json" + + +def test_replica_analyzer_uses_preserved_terminal_outcome_window(tmp_path): + logs = tmp_path / "logs" + run = logs / "p3_market_gemma4-31b-q4_v1" + run.mkdir(parents=True) + (run / "receipt.json").write_text(json.dumps({ + "captured_at": "2026-08-02T00:00:00+00:00", + })) + (run / "label.json").write_text(json.dumps({ + "primary": "identical-call-loop", + "labeled_at": "2026-08-02T00:00:10+00:00", + })) + (run / "transcript.jsonl").write_text("\n".join([ + json.dumps({"t": "2026-08-02T00:00:02+00:00", "type": "model"}), + json.dumps({"t": "2026-08-02T00:00:08+00:00", "type": "tool"}), + ]) + "\n") + csv_path = tmp_path / "gpu.csv" + with csv_path.open("w", newline="") as handle: + writer = csv.writer(handle) + writer.writerow(LOGGER.FIELDS) + for second in (0, 5, 10): + ts = f"2026-08-02T00:00:{second:02d}Z" + writer.writerow([ts, 1, 8001, 500, 450, 96, 20, 65000, 75, 2700, + run.name, 456, 110]) + report = ANALYZER.analyze(csv_path, [logs], 500, 0.13, 15) + doc = report[run.name] + assert doc["sampling"]["window_source"] == ( + "receipt.json:captured_at..label.json:labeled_at" + ) + assert doc["sampling"]["window_wall_s"] == 10.0 + assert doc["sampling"]["coverage_fraction_of_wall"] == 1.0 + assert doc["attribution"]["active_gpu_ids_observed"] == ["1"] diff --git a/tooling/deployments/gemma4-31b-q4-tower2/test_gemma4_telemetry_sidecar.py b/tooling/deployments/gemma4-31b-q4-tower2/test_gemma4_telemetry_sidecar.py new file mode 100644 index 00000000..aa46404b --- /dev/null +++ b/tooling/deployments/gemma4-31b-q4-tower2/test_gemma4_telemetry_sidecar.py @@ -0,0 +1,17 @@ +#!/usr/bin/env python3 +from pathlib import Path + + +SCRIPT = Path(__file__).with_name("gemma4-telemetry-sidecar.sh").read_text() + + +def test_sidecar_derives_cost_for_completed_and_terminal_outcomes(): + assert "-name summary.json -o -name label.json" in SCRIPT + assert '[[ -f "$run_dir/receipt.json" && -f "$run_dir/transcript.jsonl" ]]' in SCRIPT + assert "-printf '%h\\n' | sort -u" in SCRIPT + assert 'extract_cost.py" "$run_dir"' in SCRIPT + + +def test_sidecar_writes_per_run_telemetry_artifacts(): + assert "--write-run-artifacts" in SCRIPT + assert "--cap-per-gpu 500" in SCRIPT diff --git a/tooling/deployments/gemma4-31b-q4-tower2/topologies/dual-layer-1to1.env b/tooling/deployments/gemma4-31b-q4-tower2/topologies/dual-layer-1to1.env new file mode 100644 index 00000000..47efc787 --- /dev/null +++ b/tooling/deployments/gemma4-31b-q4-tower2/topologies/dual-layer-1to1.env @@ -0,0 +1,9 @@ +GEMMA_CUDA_VISIBLE_DEVICES=0,1 +GEMMA_PORT=8000 +GEMMA_CTX_SIZE=262144 +GEMMA_PARALLEL=1 +GEMMA_SPLIT_MODE=layer +GEMMA_MAIN_GPU=0 +GEMMA_TENSOR_SPLIT=1,1 +GEMMA_CACHE_TYPE_K=f16 +GEMMA_CACHE_TYPE_V=f16 diff --git a/tooling/deployments/gemma4-31b-q4-tower2/topologies/dual-row-1to1.env b/tooling/deployments/gemma4-31b-q4-tower2/topologies/dual-row-1to1.env new file mode 100644 index 00000000..d909f880 --- /dev/null +++ b/tooling/deployments/gemma4-31b-q4-tower2/topologies/dual-row-1to1.env @@ -0,0 +1,9 @@ +GEMMA_CUDA_VISIBLE_DEVICES=0,1 +GEMMA_PORT=8000 +GEMMA_CTX_SIZE=262144 +GEMMA_PARALLEL=1 +GEMMA_SPLIT_MODE=row +GEMMA_MAIN_GPU=0 +GEMMA_TENSOR_SPLIT=1,1 +GEMMA_CACHE_TYPE_K=f16 +GEMMA_CACHE_TYPE_V=f16 diff --git a/tooling/deployments/gemma4-31b-q4-tower2/topologies/replica-gpu0-f16-s2.env b/tooling/deployments/gemma4-31b-q4-tower2/topologies/replica-gpu0-f16-s2.env new file mode 100644 index 00000000..5c877e28 --- /dev/null +++ b/tooling/deployments/gemma4-31b-q4-tower2/topologies/replica-gpu0-f16-s2.env @@ -0,0 +1,9 @@ +GEMMA_CUDA_VISIBLE_DEVICES=0 +GEMMA_PORT=8000 +GEMMA_CTX_SIZE=524288 +GEMMA_PARALLEL=2 +GEMMA_SPLIT_MODE=none +GEMMA_MAIN_GPU=0 +GEMMA_CACHE_TYPE_K=f16 +GEMMA_CACHE_TYPE_V=f16 +GEMMA_KV_UNIFIED=0 diff --git a/tooling/deployments/gemma4-31b-q4-tower2/topologies/replica-gpu0-q8-s4.env b/tooling/deployments/gemma4-31b-q4-tower2/topologies/replica-gpu0-q8-s4.env new file mode 100644 index 00000000..97118cb4 --- /dev/null +++ b/tooling/deployments/gemma4-31b-q4-tower2/topologies/replica-gpu0-q8-s4.env @@ -0,0 +1,9 @@ +GEMMA_CUDA_VISIBLE_DEVICES=0 +GEMMA_PORT=8000 +GEMMA_CTX_SIZE=1048576 +GEMMA_PARALLEL=4 +GEMMA_SPLIT_MODE=none +GEMMA_MAIN_GPU=0 +GEMMA_CACHE_TYPE_K=q8_0 +GEMMA_CACHE_TYPE_V=q8_0 +GEMMA_KV_UNIFIED=0 diff --git a/tooling/deployments/gemma4-31b-q4-tower2/topologies/replica-gpu0-q8-s6.env b/tooling/deployments/gemma4-31b-q4-tower2/topologies/replica-gpu0-q8-s6.env new file mode 100644 index 00000000..bf8a3c31 --- /dev/null +++ b/tooling/deployments/gemma4-31b-q4-tower2/topologies/replica-gpu0-q8-s6.env @@ -0,0 +1,9 @@ +GEMMA_CUDA_VISIBLE_DEVICES=0 +GEMMA_PORT=8000 +GEMMA_CTX_SIZE=1572864 +GEMMA_PARALLEL=6 +GEMMA_SPLIT_MODE=none +GEMMA_MAIN_GPU=0 +GEMMA_CACHE_TYPE_K=q8_0 +GEMMA_CACHE_TYPE_V=q8_0 +GEMMA_KV_UNIFIED=0 diff --git a/tooling/deployments/gemma4-31b-q4-tower2/topologies/replica-gpu0.env b/tooling/deployments/gemma4-31b-q4-tower2/topologies/replica-gpu0.env new file mode 100644 index 00000000..986f9271 --- /dev/null +++ b/tooling/deployments/gemma4-31b-q4-tower2/topologies/replica-gpu0.env @@ -0,0 +1,8 @@ +GEMMA_CUDA_VISIBLE_DEVICES=0 +GEMMA_PORT=8000 +GEMMA_CTX_SIZE=262144 +GEMMA_PARALLEL=1 +GEMMA_SPLIT_MODE=none +GEMMA_MAIN_GPU=0 +GEMMA_CACHE_TYPE_K=f16 +GEMMA_CACHE_TYPE_V=f16 diff --git a/tooling/deployments/gemma4-31b-q4-tower2/topologies/replica-gpu1-f16-s2.env b/tooling/deployments/gemma4-31b-q4-tower2/topologies/replica-gpu1-f16-s2.env new file mode 100644 index 00000000..9d93d5c5 --- /dev/null +++ b/tooling/deployments/gemma4-31b-q4-tower2/topologies/replica-gpu1-f16-s2.env @@ -0,0 +1,9 @@ +GEMMA_CUDA_VISIBLE_DEVICES=1 +GEMMA_PORT=8001 +GEMMA_CTX_SIZE=524288 +GEMMA_PARALLEL=2 +GEMMA_SPLIT_MODE=none +GEMMA_MAIN_GPU=0 +GEMMA_CACHE_TYPE_K=f16 +GEMMA_CACHE_TYPE_V=f16 +GEMMA_KV_UNIFIED=0 diff --git a/tooling/deployments/gemma4-31b-q4-tower2/topologies/replica-gpu1-q8-s4.env b/tooling/deployments/gemma4-31b-q4-tower2/topologies/replica-gpu1-q8-s4.env new file mode 100644 index 00000000..25eebe8f --- /dev/null +++ b/tooling/deployments/gemma4-31b-q4-tower2/topologies/replica-gpu1-q8-s4.env @@ -0,0 +1,9 @@ +GEMMA_CUDA_VISIBLE_DEVICES=1 +GEMMA_PORT=8001 +GEMMA_CTX_SIZE=1048576 +GEMMA_PARALLEL=4 +GEMMA_SPLIT_MODE=none +GEMMA_MAIN_GPU=0 +GEMMA_CACHE_TYPE_K=q8_0 +GEMMA_CACHE_TYPE_V=q8_0 +GEMMA_KV_UNIFIED=0 diff --git a/tooling/deployments/gemma4-31b-q4-tower2/topologies/replica-gpu1-q8-s6.env b/tooling/deployments/gemma4-31b-q4-tower2/topologies/replica-gpu1-q8-s6.env new file mode 100644 index 00000000..33756ca4 --- /dev/null +++ b/tooling/deployments/gemma4-31b-q4-tower2/topologies/replica-gpu1-q8-s6.env @@ -0,0 +1,9 @@ +GEMMA_CUDA_VISIBLE_DEVICES=1 +GEMMA_PORT=8001 +GEMMA_CTX_SIZE=1572864 +GEMMA_PARALLEL=6 +GEMMA_SPLIT_MODE=none +GEMMA_MAIN_GPU=0 +GEMMA_CACHE_TYPE_K=q8_0 +GEMMA_CACHE_TYPE_V=q8_0 +GEMMA_KV_UNIFIED=0 diff --git a/tooling/deployments/gemma4-31b-q4-tower2/topologies/replica-gpu1.env b/tooling/deployments/gemma4-31b-q4-tower2/topologies/replica-gpu1.env new file mode 100644 index 00000000..3cc9daed --- /dev/null +++ b/tooling/deployments/gemma4-31b-q4-tower2/topologies/replica-gpu1.env @@ -0,0 +1,8 @@ +GEMMA_CUDA_VISIBLE_DEVICES=1 +GEMMA_PORT=8001 +GEMMA_CTX_SIZE=262144 +GEMMA_PARALLEL=1 +GEMMA_SPLIT_MODE=none +GEMMA_MAIN_GPU=0 +GEMMA_CACHE_TYPE_K=f16 +GEMMA_CACHE_TYPE_V=f16 diff --git a/tooling/deployments/gemma4-31b-q4-tower2/topologies/single-gpu0.env b/tooling/deployments/gemma4-31b-q4-tower2/topologies/single-gpu0.env new file mode 100644 index 00000000..986f9271 --- /dev/null +++ b/tooling/deployments/gemma4-31b-q4-tower2/topologies/single-gpu0.env @@ -0,0 +1,8 @@ +GEMMA_CUDA_VISIBLE_DEVICES=0 +GEMMA_PORT=8000 +GEMMA_CTX_SIZE=262144 +GEMMA_PARALLEL=1 +GEMMA_SPLIT_MODE=none +GEMMA_MAIN_GPU=0 +GEMMA_CACHE_TYPE_K=f16 +GEMMA_CACHE_TYPE_V=f16 diff --git a/tooling/deployments/gemma4-31b-q4-tower2/topologies/single-gpu1.env b/tooling/deployments/gemma4-31b-q4-tower2/topologies/single-gpu1.env new file mode 100644 index 00000000..e7f390d9 --- /dev/null +++ b/tooling/deployments/gemma4-31b-q4-tower2/topologies/single-gpu1.env @@ -0,0 +1,8 @@ +GEMMA_CUDA_VISIBLE_DEVICES=1 +GEMMA_PORT=8000 +GEMMA_CTX_SIZE=262144 +GEMMA_PARALLEL=1 +GEMMA_SPLIT_MODE=none +GEMMA_MAIN_GPU=0 +GEMMA_CACHE_TYPE_K=f16 +GEMMA_CACHE_TYPE_V=f16 diff --git a/tooling/deployments/gemma4-31b-q4-tower2/topology-matrix.json b/tooling/deployments/gemma4-31b-q4-tower2/topology-matrix.json new file mode 100644 index 00000000..80774199 --- /dev/null +++ b/tooling/deployments/gemma4-31b-q4-tower2/topology-matrix.json @@ -0,0 +1,73 @@ +{ + "schema_version": 1, + "campaign": "gemma4-31b-q4-tower2-topology-bakeoff", + "mandatory_per_sequence_context_tokens": 262144, + "gpu_power_limit_w_each": 500, + "sampling": {"temperature": 1.0, "top_p": 0.95, "top_k": 64}, + "candidates": [ + { + "id": "single-gpu0", + "devices": [0], + "instances": 1, + "split_mode": "none", + "initial_parallel_slots": 1 + }, + { + "id": "dual-layer-1to1", + "devices": [0, 1], + "instances": 1, + "split_mode": "layer", + "tensor_split": [1, 1], + "initial_parallel_slots": 1 + }, + { + "id": "dual-row-1to1", + "devices": [0, 1], + "instances": 1, + "split_mode": "row", + "tensor_split": [1, 1], + "initial_parallel_slots": 1 + }, + { + "id": "dual-independent-replicas", + "devices": [0, 1], + "instances": 2, + "split_mode": "none", + "initial_parallel_slots_per_instance": 1, + "parallel_slots_search_per_instance": [1, 2, 4, 6], + "context_pool_tokens_search_per_instance": [262144, 524288, 1048576, 1572864] + } + ], + "workloads": [ + "chat_format_and_native_tool_call_contract", + "uncached_prefill_1k_32k_128k_250k", + "decode_256_and_2048_tokens", + "aggregate_concurrency_1_2_4_8", + "queue_and_backpressure", + "prefix_cache_reuse", + "near_256k_start_middle_end_recall", + "restart_recovery_and_postrestart_repeat", + "sanctuary_and_pixel_pre_post_cutover" + ], + "execution_policy": { + "topology_bakeoff": "Run one candidate at a time in an isolated measurement window so GPU, power, latency, and throughput results remain comparable.", + "post_selection_campaign": "When the selected topology exposes two independent replicas, partition run IDs deterministically across GPU0 and GPU1 and execute both lanes concurrently.", + "lane_artifacts": "Keep endpoint, GPU index, process identity, timings, telemetry, outputs, grades, and retry history attributable to each run and lane.", + "shared_state": "Use disjoint run directories and atomic result publication; do not allow simultaneous workers to claim the same run ID." + }, + "selection_rule": [ + "Reject any candidate that changes the pinned model/template/sampling, cannot expose a full 262144-token sequence, fails tool-call or recall gates, OOMs, or is unstable across restart.", + "Among passing candidates, maximize aggregate concurrency-8 decode throughput.", + "If candidates are within 10 percent at concurrency 8, prefer lower median single-request end-to-end latency and faster 128K uncached prefill.", + "If still tied, prefer lower combined power, fewer cross-GPU dependencies, and simpler restart/failover behavior." + ], + "append_only_amendments": [ + { + "timestamp": "2026-08-01T23:44:00Z", + "phase": "serving optimization before any canonical or extended MMBT quality run", + "change": "Extend independent-replica slot search from [1,2,4] to [1,2,4,6] and add a 1572864-token context pool.", + "rationale": "The measured four-slot Q8 candidate used 64789 MiB of 97887 MiB while preserving four hard 262144-token slots. The measured incremental Q8 KV cost indicates six slots should remain within VRAM with a material safety margin and better satisfy the user's full-utilization objective.", + "anti_cherry_pick_note": "This amendment responds only to serving memory and throughput evidence. No canonical or extended MMBT model-quality outcome had been run or observed. All earlier topology attempts remain preserved." + } + ] +} diff --git a/tooling/deployments/gemma4-31b-q4-tower2/validate_preregistration.py b/tooling/deployments/gemma4-31b-q4-tower2/validate_preregistration.py new file mode 100755 index 00000000..ec1421ac --- /dev/null +++ b/tooling/deployments/gemma4-31b-q4-tower2/validate_preregistration.py @@ -0,0 +1,128 @@ +#!/usr/bin/env python3 +"""Fail closed when the Gemma campaign no longer matches its preregistration.""" + +from __future__ import annotations + +import hashlib +import json +import subprocess +from pathlib import Path + + +HERE = Path(__file__).resolve().parent +ROOT = HERE.parents[2] +MODEL_MANIFEST = HERE / "model-manifest.json" +TOPOLOGY = HERE / "topology-matrix.json" +MICROBENCH = ROOT / "tooling/gemma4-31b-q4-mmbt.json" +EXTENDED = ROOT / "tooling/gemma4-31b-q4-extended-matrix.json" + + +def require(condition: bool, message: str) -> None: + if not condition: + raise RuntimeError(message) + + +def load(path: Path) -> dict: + return json.loads(path.read_text(encoding="utf-8")) + + +def sha256(path: Path) -> str: + digest = hashlib.sha256() + with path.open("rb") as stream: + for chunk in iter(lambda: stream.read(8 * 1024 * 1024), b""): + digest.update(chunk) + return digest.hexdigest() + + +manifest = load(MODEL_MANIFEST) +topology = load(TOPOLOGY) +microbench = load(MICROBENCH) +extended = load(EXTENDED) + +require(manifest["source_revision"] == "59dde24573e7e61570dba08b18a2e1fe246955ed", "source revision drift") +for key in ("model", "multimodal_projector"): + artifact = manifest[key] + path = Path(artifact["path"]) + require(path.is_file(), f"missing {key}: {path}") + require(path.stat().st_size == artifact["bytes"], f"size drift for {key}") + require(sha256(path) == artifact["sha256"], f"SHA-256 drift for {key}") + metadata = path.parent / ".cache/huggingface/download" / f"{path.name}.metadata" + require(metadata.is_file(), f"missing Hugging Face metadata for {key}") + require(metadata.read_text().splitlines()[0] == manifest["source_revision"], f"metadata revision drift for {key}") + +context = manifest["model_card"]["native_context_tokens"] +sampling = { + "temperature": manifest["model_card"]["temperature"], + "top_p": manifest["model_card"]["top_p"], + "top_k": manifest["model_card"]["top_k"], +} +require(context == microbench["max_model_len"] == extended["served_context_tokens"], "context mismatch across configs") +require(context == topology["mandatory_per_sequence_context_tokens"], "topology context mismatch") +require(microbench["benchmark_temperature"] == sampling["temperature"], "microbench temperature drift") +require(microbench["benchmark_top_p"] == sampling["top_p"], "microbench top-p drift") +require(microbench["benchmark_top_k"] == sampling["top_k"], "microbench top-k drift") +require(all(extended[key] == value for key, value in sampling.items()), "extended sampling drift") +require(topology["sampling"] == sampling, "topology sampling drift") +require(microbench["canonical_n"] == extended["replicates"] == 3, "canonical N mismatch") +require(microbench["variance_expansion_n"] == 10, "variance expansion N drift") + +extended_protocol = ROOT / extended["substantive_audit_protocol"] +require(extended_protocol.is_file(), "missing extended substantive audit protocol") +require( + sha256(extended_protocol) == extended["substantive_audit_protocol_sha256"], + "extended substantive audit protocol hash drift", +) + +for suite in extended["suites"]: + task = ROOT / suite["task"] + require(task.is_file(), f"missing task: {task}") + require(sha256(task) == suite["current_task_sha256"], f"task hash drift: {suite['id']}") + require(suite["max_output_tokens_cap"] == context, f"artificial output cap: {suite['id']}") + require(all(suite[key] == value for key, value in sampling.items()), f"suite sampling drift: {suite['id']}") + if suite.get("subject_pin"): + subject_pin = ROOT / suite["subject_pin"] + require(subject_pin.is_file(), f"missing subject pin: {suite['id']}") + require( + sha256(subject_pin) == suite["subject_pin_sha256"], + f"subject pin hash drift: {suite['id']}", + ) + pinned = load(subject_pin) + pinned_shas = { + pinned["base_sha"], pinned["head_sha"], pinned["squash_merge_sha"], + *pinned["pr_commit_shas"], + } + require( + set(suite["required_subject_shas"]).issubset(pinned_shas), + f"required subject refs drift: {suite['id']}", + ) + +fixture = next(suite for suite in extended["suites"] if suite["id"] == "dreamserver-75-pr-audit") +pr_set = Path(fixture["input_path"]) / "canonical-prs.txt" +require(sha256(pr_set) == "569b95b3384af0c4ae4b54a2c8c8f7c908b396124777927a37b5c8fa0211ecd1", "frozen PR set drift") + +base = manifest["tower2"]["mmbt_base_commit"] +ancestor = subprocess.run( + ["git", "-C", str(ROOT), "merge-base", "--is-ancestor", base, "HEAD"], + check=False, +).returncode +require(ancestor == 0, "campaign branch no longer descends from preregistered MMBT base") + +power_rows = subprocess.check_output( + ["nvidia-smi", "--query-gpu=power.limit", "--format=csv,noheader,nounits"], + text=True, +).splitlines() +power_limits = [round(float(row.strip()), 2) for row in power_rows] +require(power_limits == [500.0, 500.0], f"GPU power limits are not pinned: {power_limits}") + +print(json.dumps({ + "status": "valid", + "source_revision": manifest["source_revision"], + "model_sha256": manifest["model"]["sha256"], + "mmproj_sha256": manifest["multimodal_projector"]["sha256"], + "context_tokens": context, + "sampling": sampling, + "canonical_n": microbench["canonical_n"], + "variance_expansion_n": microbench["variance_expansion_n"], + "extended_suites": [suite["id"] for suite in extended["suites"]], + "power_limits_w": power_limits, +}, indent=2)) diff --git a/tooling/gemma4-31b-q4-extended-matrix.json b/tooling/gemma4-31b-q4-extended-matrix.json new file mode 100644 index 00000000..3d22b58d --- /dev/null +++ b/tooling/gemma4-31b-q4-extended-matrix.json @@ -0,0 +1,97 @@ +{ + "schema_version": 1, + "model": "Gemma-4-31B-it-QAT-Q4_0", + "endpoints": [ + "http://127.0.0.1:8000/v1", + "http://127.0.0.1:8001/v1" + ], + "lane_ports": [8000, 8001], + "served_context_tokens": 262144, + "temperature": 1.0, + "top_p": 0.95, + "top_k": 64, + "replicates": 3, + "substantive_audit_protocol": "tooling/GEMMA4-EXTENDED-SUBSTANTIVE-AUDIT-PROTOCOL.md", + "substantive_audit_protocol_sha256": "cb167238f97a6b23d3e1d1bd64b71349c6b9846df9de5b34b07dcb3bcb664699", + "historical_comparison_arms": [ + "Qwen3.6-27B-AWQ", + "Qwen3.6-35B-A3B-AWQ", + "Qwen3-Coder-Next-AWQ", + "Qwen3.5-397B-A17B-Q3", + "DeepSeek-V4-Flash-0731" + ], + "suites": [ + { + "id": "dreamserver-1-pr-audit", + "task": "tooling/tasks/task_pr_audit_n1.md", + "historical_task_sha256": "2e8770b3dbd210679e291f043f75b0d8aa5ec030c38b956926b8ba41c2cfb55c", + "current_task_sha256": "18c722e23dc9b38ce2be015bd2f3ab912c678bda7b44e5089d20af8a2f3e61c1", + "subject_pin": "tooling/gemma4-single-pr-subject-pin.json", + "subject_pin_sha256": "3e22c76de5eff21b17881bb14ee301226c6eb0e3fd6f01065a8e074e98227a1e", + "required_subject_shas": [ + "309e9cd0ad5d572ab9313eb577ec30a3e8524752", + "e5ceb43ea0f9c3939a154fb42e6649b828eca878", + "1678f19404c54f8588eceabfb21c3f4ce812b483", + "e4e83fedfee36c651afc6af7f3dba9851fa5dc96", + "488f5b48f02946ff31ce3f342566d2a1a9687201", + "ff20aadf067905b89c9c781cb5ee7ea0532551d5", + "ab148dee598635644b759f87abb3f594918276b6" + ], + "temperature": 1.0, + "top_p": 0.95, + "top_k": 64, + "max_output_tokens_cap": 262144, + "stuck_threshold": 500, + "require_git_tag": true + }, + { + "id": "wallstreet-investment-memo", + "task": "tooling/tasks/task_investment_memo.md", + "historical_task_sha256": "02f48c90f5c951cff424ad0646dab2d00e7225194c9a64aefee1c6aa736c61a3", + "current_task_sha256": "02f48c90f5c951cff424ad0646dab2d00e7225194c9a64aefee1c6aa736c61a3", + "temperature": 1.0, + "top_p": 0.95, + "top_k": 64, + "max_output_tokens_cap": 262144, + "stuck_threshold": 500, + "require_git_tag": true + }, + { + "id": "wallstreet-board-presentation", + "task": "tooling/tasks/task_board_presentation.md", + "historical_task_sha256": null, + "current_task_sha256": "b9f5e451d04bf7cad85eeb5fa98051b44b61028847b984048930437aadbbf4be", + "temperature": 1.0, + "top_p": 0.95, + "top_k": 64, + "max_output_tokens_cap": 262144, + "stuck_threshold": 500, + "require_git_tag": true, + "input_from": "wallstreet-investment-memo" + }, + { + "id": "dreamserver-75-pr-audit", + "task": "tooling/tasks/task_dreamserver_pr_audit_frozen75.md", + "historical_task_sha256": "5f3ee8854907bbd78d9a6a6aaaad91e632821b0c77e09ac3fbcf42e0c5b33fde", + "current_task_sha256": "0d426e15303b3f886a75db81165cdbe9d91534a474217f91f4e2baf41c0056a6", + "temperature": 1.0, + "top_p": 0.95, + "top_k": 64, + "max_output_tokens_cap": 262144, + "stuck_threshold": 500, + "require_git_tag": true, + "input_path": "/mnt/bulk/benchmark-fixtures/dreamserver-75-pr-audit-2026-04-27" + } + ], + "comparability_notes": [ + "Gemma is served at its model-card 262144-token native limit. Every suite uses that value as the request cap; the harness estimates prompt growth with a 12000-token reserve, adds 2048 context-safety tokens, and limits each request to the remaining context.", + "Gemma is evaluated at its published sampling point: temperature=1.0, top_p=0.95, top_k=64. Historical arms with different sampling are compared with an explicit operating-point caveat.", + "The canonical first cohort is N=3 for direct historical comparison. The same immutable v1-v3 cohort is then extended through v10 for every one of the 12 families; v1-v3 are never replaced by later runs.", + "Both RTX PRO 6000 GPUs remain capped at 500 W. The selected serving topology is chosen by a preregistered bakeoff and may use independent per-GPU replicas if that beats cross-GPU splitting.", + "The current public 1-PR and 75-PR task files differ bytewise from historical private-receipt hashes after extraction and sanitization; both historical and current hashes are pinned.", + "The one-PR prompt's word 'open' is stale because PR #1057 merged on 2026-05-10. The unchanged task is anchored to the same base 309e9cd0, head e5ceb43e, squash 1678f194, and original PR commits evidenced in all three DeepSeek comparator archives; current main cannot substitute for that subject.", + "The 75-PR suite uses the frozen historical fixture at baseline d5154c3. Its canonical PR-number set SHA-256 is 569b95b3384af0c4ae4b54a2c8c8f7c908b396124777927a37b5c8fa0211ecd1.", + "Every attempt is preserved. Only a proven pre-response infrastructure failure may be excluded; length termination, looping, malformed tool use, artifact omissions, server-visible API failures without a demonstrated endpoint outage, and other model-control failures remain valid outcomes.", + "Raw grades are immutable. Any grader correction is stored as a non-destructive overlay tied to unchanged archive hashes and grader commits." + ] +} diff --git a/tooling/gemma4-31b-q4-mmbt.json b/tooling/gemma4-31b-q4-mmbt.json new file mode 100644 index 00000000..a3615c9b --- /dev/null +++ b/tooling/gemma4-31b-q4-mmbt.json @@ -0,0 +1,30 @@ +{ + "schema_version": 1, + "preset": "gemma4-31b-q4", + "engine": "external", + "model": "Gemma-4-31B-it-QAT-Q4_0", + "port": 8000, + "lane_ports": [8000, 8001], + "max_model_len": 262144, + "benchmark_max_output_tokens_cap": 262144, + "benchmark_temperature": 1.0, + "benchmark_top_p": 0.95, + "benchmark_top_k": 64, + "services": [ + "mmbt-gemma4@replica-gpu0-q8-s4.service", + "mmbt-gemma4@replica-gpu1-q8-s4.service" + ], + "launcher": "/home/michael/bench-gemma4-31b-q4/tooling/deployments/gemma4-31b-q4-tower2/ensure-gemma4-winner.sh", + "serving_manifest": "/home/michael/bench-gemma4-31b-q4/tooling/deployments/gemma4-31b-q4-tower2/benchmark-serving-manifest.json", + "arms": [ + { + "label": "gemma4-31b-q4", + "thinking": "" + } + ], + "canonical_n": 3, + "variance_expansion_n": 10, + "variance_expansion_scope": "all_12_families", + "stuck_secs": 1200, + "endpoint_grace_secs": 300 +} diff --git a/tooling/gemma4-comparison-sources.json b/tooling/gemma4-comparison-sources.json new file mode 100644 index 00000000..a333bf75 --- /dev/null +++ b/tooling/gemma4-comparison-sources.json @@ -0,0 +1,146 @@ +{ + "schema_version": 1, + "source_repository_commit": "dc8f57d443aeccba652a33852fda32d37cd701f3", + "purpose": "Pin the exact published evidence used for the post-campaign Gemma comparison before Gemma grades are known.", + "source_documents": [ + { + "path": "SCORECARD.md", + "sha256": "50d4a8b23addb1ceb32ff7d36c91be99e4a470a495d9a5d1f4bb18337526be30", + "verify_at_source_commit": true + }, + { + "path": "benchmarks/microbench-2026-04-28/bug-fixing/Qwen3.6-27B-AWQ/receipt.json", + "sha256": "99e0e1b780f11452ccf7294f1a510a0154840699077a1c415cee36d79c8cb91c" + }, + { + "path": "benchmarks/microbench-2026-04-28/bug-fixing/Qwen3-Coder-Next-AWQ/receipt.json", + "sha256": "fe36925f249b6162e494dceb7848e4cf124609e95233e3ddb5afc281dc4a9540" + }, + { + "path": "benchmarks/dreamserver-1-pr-audit/Qwen3.6-35B-A3B-AWQ/receipt.json", + "sha256": "b488efbd3859ef73275fd1cfd97bbb4130e1f7bd34d2ad6c95843c498cfba0ec" + }, + { + "path": "benchmarks/dreamserver-1-pr-audit/Qwen3.6-35B-A3B-AWQ/README.md", + "sha256": "7ba4c41b603ddec3858eb4663ea8c0ea989b21aeeaac72eefb08c3a3f511c325" + }, + { + "path": "benchmarks/wallstreet-intern-test/Qwen3.6-35B-A3B-AWQ/README.md", + "sha256": "dae02dff30ef07ce47a045e7503deacc05f1459cd012867ecf70e4053565d153" + }, + { + "path": "benchmarks/deepseek-v4-flash-0731/canonical-regrade-audit.json", + "sha256": "ea9330029d725458c70d7efac2f298d6bb0f31198a759880d03b9e079de842b4" + }, + { + "path": "benchmarks/deepseek-v4-flash-0731/DEEPSEEK_V4_FLASH_0731_VERIFIED_RESULTS.md", + "sha256": "efb686161b6f6a4e481d0bca864a74e37b38e0e89700867de9346de6920aa517" + }, + { + "path": "benchmarks/deepseek-v4-flash-0731/qwen397-corrected-score-overlay.json", + "sha256": "ae0f65248ed74f8fcbd5d01093720cb4e65c8d2da2c278d7460870e55a309e0f" + }, + { + "path": "hardware-tests/qwen3.5-397b-vs-step3.7-flash-2026-05-29/findings-n10.md", + "sha256": "58a98ffd94d6401bbb38c05cbe895ca9ac6974f2ea6af2dc23633b4136d697f2" + }, + { + "path": "hardware-tests/qwen3.5-397b-vs-step3.7-flash-2026-05-29/manifest.json", + "sha256": "eb9ecbb63304868facba58a2eea5b146a807afe077881836ce393c152d36a772" + }, + { + "path": "benchmarks/microbench-phase-b-2026-05-02/findings.md", + "sha256": "69ba6c4b5dd6a6f961f053fb7d6a1aea64e464d20442acab66828b8f13c73e42" + } + ], + "comparators": [ + { + "id": "qwen3.6-27b-awq", + "display_name": "Qwen3.6-27B-AWQ", + "canonical": { + "cohort": "12 families x N=3", + "raw_passes": 20, + "total": 36 + }, + "operating_point": { + "engine": "vLLM", + "quantization": "AWQ 4-bit", + "context_tokens": 262144, + "temperature": 0.3, + "historical_per_response_ceiling": 180000 + } + }, + { + "id": "qwen3-coder-next-awq", + "display_name": "Qwen3-Coder-Next-AWQ", + "canonical": { + "cohort": "12 families x N=3", + "raw_passes": 20, + "total": 36 + }, + "operating_point": { + "engine": "vLLM", + "quantization": "AWQ 4-bit", + "context_tokens": 262144, + "temperature": 0.3, + "historical_per_response_ceiling": 180000 + } + }, + { + "id": "qwen3.6-35b-a3b-awq", + "display_name": "Qwen3.6-35B-A3B-AWQ", + "canonical": null, + "available_evidence": "No complete comparable 12-family canonical cohort. The published single-PR run produced 0/13 required artifacts and the Wall Street arm produced no usable deliverable across three attempts. Report this as sparse extended-suite evidence, not a canonical pass-rate rank.", + "operating_point": { + "engine": "vLLM", + "quantization": "AWQ 4-bit", + "context_tokens": 262144, + "temperature": 0.3, + "historical_per_response_ceiling": 180000 + } + }, + { + "id": "qwen3.5-397b-a17b-q3-nothink", + "display_name": "Qwen3.5-397B-A17B UD-Q3_K_XL (no-think)", + "canonical": { + "cohort": "12 families x N=10", + "raw_passes": 82, + "corrected_passes": 92, + "total": 120 + }, + "operating_point": { + "engine": "llama.cpp", + "quantization": "UD-Q3_K_XL", + "context_tokens": 131072, + "temperature": 0.3, + "reasoning_mode": "no-think", + "gpu_power_limit_w": 600 + } + }, + { + "id": "deepseek-v4-flash-0731", + "display_name": "DeepSeek V4 Flash 0731", + "canonical": { + "cohort": "12 families x N=3", + "raw_passes": 23, + "corrected_passes": 35, + "total": 36 + }, + "operating_point": { + "engine": "vLLM", + "quantization": "FP8", + "context_tokens": 1048576, + "temperature": 1.0, + "top_p": 0.95, + "gpu_power_limit_w": 500 + } + } + ], + "comparison_rules": [ + "Raw and corrected results are separate columns; never replace historical raw scores silently.", + "Use Gemma N=3 for directional comparison with historical N=3 cohorts and Gemma N=10 for variance-aware comparison with Qwen3.5-397B N=10.", + "Do not rank Qwen3.6-35B-A3B on the canonical matrix because no complete comparable cohort exists.", + "Show engine, quantization, sampling, context, output-ceiling, power-cap, date, and grader differences next to aggregate rates.", + "A cross-model rate is not a global SOTA claim and does not erase task-level or extended-suite failure modes." + ] +} diff --git a/tooling/gemma4-single-pr-subject-pin.json b/tooling/gemma4-single-pr-subject-pin.json new file mode 100644 index 00000000..f7f0102e --- /dev/null +++ b/tooling/gemma4-single-pr-subject-pin.json @@ -0,0 +1,48 @@ +{ + "schema_version": 1, + "captured_at": "2026-08-02T01:26:00Z", + "repository_at_capture": "Osmantic/ODS", + "legacy_repository_url_in_task": "https://github.com/Light-Heart-Labs/DreamServer", + "pull_request": 1057, + "task_text_claims_open": true, + "observed_state": "closed", + "merged": true, + "merged_at": "2026-05-10T15:25:34Z", + "title": "fix(host-agent): runtime hygiene — narrow pull, surface failures, normalize bind volumes", + "author": "yasinBursali", + "base_ref": "main", + "base_sha": "309e9cd0ad5d572ab9313eb577ec30a3e8524752", + "head_ref": "fix/host-agent-runtime-hygiene", + "head_sha": "e5ceb43ea0f9c3939a154fb42e6649b828eca878", + "squash_merge_sha": "1678f19404c54f8588eceabfb21c3f4ce812b483", + "pr_commit_shas": [ + "e4e83fedfee36c651afc6af7f3dba9851fa5dc96", + "488f5b48f02946ff31ce3f342566d2a1a9687201", + "ff20aadf067905b89c9c781cb5ee7ea0532551d5", + "ab148dee598635644b759f87abb3f594918276b6", + "e5ceb43ea0f9c3939a154fb42e6649b828eca878" + ], + "api_diff_sha256_at_capture": "a9738afb716ee746313da36d973df74e6ed1ebd58c8508a40a02beffc0e2d676", + "api_diff_at_capture_note": "The post-merge API diff contains only the surviving 168-line test contribution. The full review subject includes the pinned PR commits and merge-history reconciliation above.", + "deepseek_comparator_archives": [ + { + "run": "n1_deepseek-v4-flash-0731_v1", + "workspace_archive_sha256": "622d2cb56c583d20c7a8a39ff7f02825098bc6c1eff6ae35fdbcd34d44557506" + }, + { + "run": "n1_deepseek-v4-flash-0731_v2", + "workspace_archive_sha256": "4a1aa42dbcab8d69487a801a366c91631257766ce2d8662d5e9ab18d73caebbc" + }, + { + "run": "n1_deepseek-v4-flash-0731_v3", + "workspace_archive_sha256": "2c3bc2a85065eae416f852f6f3132de95bda1b688bf3856cead3c62b463eaf92" + } + ], + "comparability_policy": [ + "Keep the original one-PR task text unchanged.", + "Treat the stale word 'open' as task metadata drift, not as permission to substitute another PR.", + "Every Gemma audit must identify and reconcile the pinned head, base, squash merge, and original PR commits.", + "Current main may be used as additional context but may not replace the immutable PR subject.", + "A run that cannot obtain the pinned refs because of a proven upstream or network failure is infrastructure-invalid and preserved; a run that obtains them but audits the wrong subject is a model-quality failure." + ] +} diff --git a/tooling/harness.py b/tooling/harness.py index 2dcb7e76..064a6467 100644 --- a/tooling/harness.py +++ b/tooling/harness.py @@ -7,6 +7,7 @@ from datetime import datetime, timezone from pathlib import Path import urllib.request, urllib.error +from urllib.parse import urlparse SANDBOX = "bench-sandbox-run" # default; overridden per-run in main() so parallel runs can coexist IMAGE = "bench-sandbox:latest" @@ -63,11 +64,84 @@ def docker_inspect(name, fmt=None): return p.stdout.strip() if p.returncode == 0 else None +def _redact_argv(argv): + """Return process arguments without accidentally persisting credentials.""" + redacted = [] + hide_next = False + secret_flags = {"--api-key", "--api-key-file", "--token", "--auth-token"} + for arg in argv: + if hide_next: + redacted.append("") + hide_next = False + continue + lowered = arg.lower() + if lowered in secret_flags: + redacted.append(arg) + hide_next = True + elif any(lowered.startswith(f"{flag}=") for flag in secret_flags): + redacted.append(arg.split("=", 1)[0] + "=") + else: + redacted.append(arg) + return redacted + + +def _host_server_processes(api_url): + """Capture host-native llama-server provenance for the request endpoint.""" + port = urlparse(api_url).port + processes = [] + proc_root = Path("/proc") + if not proc_root.exists(): + return processes + for entry in proc_root.iterdir(): + if not entry.name.isdigit(): + continue + try: + argv = (entry / "cmdline").read_bytes().split(b"\0") + argv = [part.decode("utf-8", "replace") for part in argv if part] + except (OSError, PermissionError): + continue + if not argv or not any("llama-server" in part for part in argv): + continue + port_matches = any( + arg == str(port) and i > 0 and argv[i - 1] in ("--port", "-p") + for i, arg in enumerate(argv) + ) or any(arg == f"--port={port}" for arg in argv) + if not port_matches: + continue + exe_path = None + exe_sha256 = None + try: + exe_path = str((entry / "exe").resolve(strict=True)) + exe_sha256 = file_sha256(exe_path) + except (OSError, PermissionError): + pass + processes.append({ + "pid": int(entry.name), + "exe": exe_path, + "exe_sha256": exe_sha256, + "argv": _redact_argv(argv), + }) + return processes + + +def _endpoint_models(api_url): + models_url = api_url.rsplit("/chat/completions", 1)[0] + "/models" + try: + with urllib.request.urlopen(models_url, timeout=10) as response: + return { + "url": models_url, + "http_status": response.status, + "payload": json.loads(response.read().decode("utf-8")), + } + except Exception as exc: + return {"url": models_url, "error": f"{type(exc).__name__}: {exc}"} + + def record_environment(run_name, model, api_url, task_file, log_dir, *, sandbox_runtime=None, temperature=0.0, stuck_threshold=30, max_iters=10000, reasoning_effort=None, enable_thinking=None, max_model_len=262144, max_output_tokens_cap=180000, - top_p=None, top_k=None): + top_p=None, top_k=None, serving_manifest=None): """Capture everything needed to reproduce the run. Written before the loop starts. sandbox_runtime: dict of per-run sandbox flags (gh_token_set, docker_socket, @@ -112,6 +186,27 @@ def record_environment(run_name, model, api_url, task_file, log_dir, *, }, } + # The benchmark may use a host-native systemd service rather than a Docker + # inference container. Record the immutable deployment manifest, live model + # identity, and exact host process/binary serving this endpoint. + receipt["serving"] = { + "manifest": None, + "endpoint_models": _endpoint_models(api_url), + "host_processes": _host_server_processes(api_url), + } + if serving_manifest: + manifest_path = Path(serving_manifest).resolve() + manifest = { + "path": str(manifest_path), + "sha256": file_sha256(manifest_path), + "byte_size": manifest_path.stat().st_size, + } + try: + manifest["payload"] = json.loads(manifest_path.read_text()) + except (OSError, json.JSONDecodeError) as exc: + manifest["read_error"] = f"{type(exc).__name__}: {exc}" + receipt["serving"]["manifest"] = manifest + # Best-effort: identify running vLLM containers either by the historical # ``vllm-*`` naming convention or by the exact served model name in the # container name/arguments. Production deployments often use host @@ -212,9 +307,13 @@ def record_environment(run_name, model, api_url, task_file, log_dir, *, "top_p": top_p, "top_k": top_k, "max_tokens_strategy": ( - f"min({max_output_tokens_cap}, max_model_len - " - "last_prompt_tokens - 14000), floor 2048" + f"min({max_output_tokens_cap}, max(2048, max_model_len - " + "max(last_prompt_tokens + 12000, 8000) - 2048))" ), + "prompt_growth_reserve_tokens": 12000, + "minimum_estimated_prompt_tokens": 8000, + "context_safety_tokens": 2048, + "minimum_request_max_tokens": 2048, "max_model_len": max_model_len, "max_output_tokens_cap": max_output_tokens_cap, "stream": False, @@ -479,7 +578,9 @@ def validate_done(require_files, require_git_tag): File requirements are matched as bare filenames against `find /workspace -maxdepth 2 -name ` so the agent's choice of audit-repo location (e.g. /workspace/ vs /workspace/audit-repo/ vs /workspace/audit-pr-1057/) - doesn't matter. Same for the git-tag check.""" + doesn't matter. The git check accepts only a clean candidate repository + whose HEAD has an annotated tag, so a pre-existing tag in an input clone + cannot satisfy the completion gate.""" missing = [] for fname in (require_files or []): # Strip leading slashes so we always match by basename pattern; the @@ -489,17 +590,24 @@ def validate_done(require_files, require_git_tag): if not r['stdout'].strip(): missing.append(fname) if require_git_tag: - # Find any git repo under /workspace with at least one annotated tag. + # Find a clean git repo under /workspace with an annotated tag at HEAD. # /workspace itself, /workspace/*/, and /workspace/*/*/ — covers nested # audit repos like /workspace/dreamserver-audit/. cmd = ( "for d in /workspace /workspace/*/ /workspace/*/*/; do " - " [ -d \"$d/.git\" ] && (cd \"$d\" && git tag -l 2>/dev/null | grep -q . && echo TAG_FOUND && break); " + " [ -d \"$d/.git\" ] || continue; " + " git -C \"$d\" diff --quiet 2>/dev/null || continue; " + " git -C \"$d\" diff --cached --quiet 2>/dev/null || continue; " + " [ -z \"$(git -C \"$d\" ls-files --others --exclude-standard 2>/dev/null)\" ] || continue; " + " for t in $(git -C \"$d\" tag --points-at HEAD 2>/dev/null); do " + " [ \"$(git -C \"$d\" cat-file -t \"refs/tags/$t\" 2>/dev/null)\" = tag ] " + " && echo TAG_FOUND && break 2; " + " done; " "done" ) r = docker_exec(cmd, timeout=15) if "TAG_FOUND" not in r['stdout']: - missing.append("(no annotated git tag in any workspace repo)") + missing.append("(no clean workspace repo with an annotated tag at HEAD)") if missing: return ( "DONE_REJECTED: Required artifacts missing — task spec demands these before completion: " @@ -510,10 +618,34 @@ def validate_done(require_files, require_git_tag): def execute_tool(name, args, log_dir, require_files=None, require_git_tag=False): + # Tool schemas declare these fields required, but a model can still emit a + # syntactically valid call that omits one. That must become an observable + # tool error the model can repair, never a harness exception that destroys + # the benchmark attempt. + if not isinstance(args, dict): + return f"TOOL_ERROR: {name} arguments must be a JSON object" + required = { + "bash": ("command",), + "write_file": ("path", "content"), + "read_file": ("path",), + } + missing = [key for key in required.get(name, ()) if key not in args] + if missing: + return f"TOOL_ERROR: {name} missing required argument(s): {', '.join(missing)}" + if name == "bash": - cmd = args.get("command", "") + cmd = args["command"] + if not isinstance(cmd, str) or not cmd: + return "TOOL_ERROR: bash argument 'command' must be a non-empty string" workdir = args.get("workdir") or "/workspace" - timeout = int(args.get("timeout_s") or 300) + if not isinstance(workdir, str): + return "TOOL_ERROR: bash argument 'workdir' must be a string" + try: + timeout = int(args.get("timeout_s") or 300) + except (TypeError, ValueError): + return "TOOL_ERROR: bash argument 'timeout_s' must be a positive integer" + if timeout <= 0: + return "TOOL_ERROR: bash argument 'timeout_s' must be a positive integer" r = docker_exec(cmd, workdir=workdir, timeout=timeout) body = f"rc={r['rc']} duration={r['duration_s']}s\n--- stdout ---\n{r['stdout']}" if r['stderr']: @@ -523,9 +655,13 @@ def execute_tool(name, args, log_dir, require_files=None, require_git_tag=False) return body elif name == "write_file": path = args["path"] + content = args["content"] + if not isinstance(path, str) or not path: + return "TOOL_ERROR: write_file argument 'path' must be a non-empty string" + if not isinstance(content, str): + return "TOOL_ERROR: write_file argument 'content' must be a string" if not path.startswith("/"): path = "/workspace/" + path - content = args["content"] # Stage via tempfile on host, docker cp into sandbox (handles binary/special chars cleanly) tmp = Path(log_dir) / f".write_{uuid.uuid4().hex}.tmp" tmp.write_text(content, encoding="utf-8") @@ -540,6 +676,8 @@ def execute_tool(name, args, log_dir, require_files=None, require_git_tag=False) return f"wrote {size} bytes to {path}" elif name == "read_file": path = args["path"] + if not isinstance(path, str) or not path: + return "TOOL_ERROR: read_file argument 'path' must be a non-empty string" if not path.startswith("/"): path = "/workspace/" + path r = docker_exec(f"head -c 200000 {path!r}", timeout=10) @@ -794,7 +932,7 @@ def main(): "every request and recorded in the receipt. Leave unset for other models.") ap.add_argument("--max-model-len", type=int, default=262144, help="The endpoint's context window, used to size each request's max_tokens " - "(max_tokens = min(180000, max_model_len - prompt - safety)). Default 262144 " + "(max_tokens = min(output cap, remaining estimated context)). Default 262144 " "matches the vLLM models benched so far. Set to the served --ctx-size for " "models hosted with a smaller window (e.g. 131072 for the 397B GGUF on llama.cpp) " "so requests don't exceed the context and 400.") @@ -802,6 +940,9 @@ def main(): help="Hard cap for each request's max_tokens. Default 180000 matches the " "current PR-audit and microbench harness. Use 64000 to reproduce the " "historical investment-memo/board harness operating point.") + ap.add_argument("--serving-manifest", default=None, + help="Path to the immutable deployment manifest for the live inference " + "endpoint. Its path, hash, size, and JSON payload are captured in receipt.json.") ap.add_argument("--temperature", type=float, default=0.0, help="Sampling temperature sent on every request. Default 0.0 (deterministic). " "At temp=0 with seed=42, models can fall into fixed-point loops on long-horizon " @@ -960,6 +1101,7 @@ def main(): max_output_tokens_cap=args.max_output_tokens_cap, top_p=args.top_p, top_k=args.top_k, + serving_manifest=args.serving_manifest, ) print(f"receipt -> {log_dir / 'receipt.json'} (vllm containers logged: {len(receipt['vllm']['containers'])})") diff --git a/tooling/run_gemma4_extended_suites.py b/tooling/run_gemma4_extended_suites.py new file mode 100755 index 00000000..dd1d95cf --- /dev/null +++ b/tooling/run_gemma4_extended_suites.py @@ -0,0 +1,447 @@ +#!/usr/bin/env python3 +"""Persistent sharded Gemma 4 runner for the non-microbench MMBT suites.""" +from __future__ import annotations + +import hashlib +import json +import os +import shutil +import subprocess +import sys +import time +from pathlib import Path + +ROOT = Path(__file__).resolve().parent.parent +TOOLING = ROOT / "tooling" +LOGS = ROOT / "logs" +WORKSPACES = TOOLING / "workspace" +MATRIX_PATH = TOOLING / "gemma4-31b-q4-extended-matrix.json" +LANE_INDEX = int(os.environ.get("GEMMA_EXTENDED_LANE_INDEX", "0")) +LANE_COUNT = int(os.environ.get("GEMMA_EXTENDED_LANE_COUNT", "1")) +PORT = int(os.environ.get("GEMMA_EXTENDED_PORT", "8000")) +STATE_ROOT = Path("/tmp/mmbt-gemma4-31b-q4-extended") +STATE = STATE_ROOT / f"lane{LANE_INDEX}" +STATUS = STATE / "status.json" +EVENTS = STATE / "events.jsonl" +MAIN_STATUS = Path("/tmp/bench-autopilot/status.json") +MODEL = "Gemma-4-31B-it-QAT-Q4_0" +LAUNCHER = ROOT / "tooling/deployments/gemma4-31b-q4-tower2/ensure-gemma4-winner.sh" +SERVING_MANIFEST = ROOT / "tooling/deployments/gemma4-31b-q4-tower2/benchmark-serving-manifest.json" +SUBSTANCE = TOOLING / "scripts" / "check_substance.py" +STATE.mkdir(parents=True, exist_ok=True) + + +def event(kind: str, **data) -> None: + row = {"ts": time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()), "kind": kind, **data} + with EVENTS.open("a") as f: + f.write(json.dumps(row, sort_keys=True) + "\n") + print(json.dumps(row, sort_keys=True), flush=True) + + +def write_status(**data) -> None: + data["updated"] = time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()) + tmp = STATUS.with_suffix(".tmp") + tmp.write_text(json.dumps(data, indent=2, sort_keys=True) + "\n") + tmp.replace(STATUS) + + +def sha256(path: Path) -> str: + h = hashlib.sha256() + with path.open("rb") as f: + for chunk in iter(lambda: f.read(1024 * 1024), b""): + h.update(chunk) + return h.hexdigest() + + +def endpoint_up() -> bool: + return subprocess.run( + ["curl", "-fsS", "--max-time", "5", f"http://127.0.0.1:{PORT}/v1/models"], + stdout=subprocess.DEVNULL, + stderr=subprocess.DEVNULL, + ).returncode == 0 + + +def ensure_endpoint() -> bool: + if endpoint_up(): + return True + event("endpoint_restart", launcher=str(LAUNCHER)) + subprocess.run(["bash", str(LAUNCHER)], check=False) + for _ in range(60): + if endpoint_up(): + event("endpoint_recovered") + return True + time.sleep(5) + event("endpoint_recovery_failed") + return False + + +def git_clean() -> bool: + r = subprocess.run( + ["git", "-C", str(ROOT), "status", "--porcelain"], + capture_output=True, + text=True, + check=False, + ) + if r.stdout.strip(): + event("dirty_worktree", details=r.stdout.strip()) + return False + return True + + +def main_microbench_complete() -> bool: + try: + j = json.loads(MAIN_STATUS.read_text()) + except Exception: + return False + return j.get("phase") == "COMPLETE" and j.get("grand_done") == j.get("grand_total") == 120 + + +def wait_for_microbench() -> None: + while not main_microbench_complete(): + write_status(phase="WAITING_FOR_MICROBENCH", lane=LANE_INDEX, port=PORT) + time.sleep(60) + event("microbench_gate_passed") + + +def run_name(suite_id: str, rep: int) -> str: + names = { + "dreamserver-1-pr-audit": "n1_gemma4-31b-q4", + "wallstreet-investment-memo": "gemma4-31b-q4_invest_memo", + "wallstreet-board-presentation": "gemma4-31b-q4_board_pres", + "dreamserver-75-pr-audit": "gemma4-31b-q4_75pr", + } + return f"{names[suite_id]}_v{rep}" + + +def terminal_pathology(log_dir: Path) -> bool: + try: + label = json.loads((log_dir / "label.json").read_text()) + return label.get("primary") == "identical-call-loop" + except Exception: + return False + + +def terminal_labeled_outcome(log_dir: Path) -> bool: + """Return whether a run has an explicit terminal benchmark label.""" + try: + label = json.loads((log_dir / "label.json").read_text()) + except Exception: + return False + return bool(label.get("primary")) + + +def completed(log_dir: Path) -> bool: + return ((log_dir / "summary.json").exists() and (log_dir / "workspace_final.tar.gz").exists()) \ + or terminal_labeled_outcome(log_dir) + + +def infra_invalid(log_dir: Path) -> bool: + """Return whether a completed run is invalid infrastructure evidence. + + MMBT's published failure taxonomy treats both ``api_error: timed out`` + and other ``api_error: ...`` finish reasons as benchmark outcomes. Do not + discard or retry those here: they can reflect model context, parser, OOM, + or single-call latency failures. The live supervisor separately detects + an endpoint outage lasting more than 90 seconds and sets + ``killed_for_infra`` for the genuinely retryable case. + """ + try: + reason = str(json.loads((log_dir / "summary.json").read_text()).get("finish_reason") or "") + except Exception: + # A missing summary is not evidence of an endpoint outage. It can be + # a model/tool-schema failure or a harness crash and must not inflate + # shipped rate through an automatic retry. + return False + return reason.startswith("endpoint_") + + +def archive_invalid(name: str, attempt: int) -> None: + stamp = time.strftime("%Y%m%dT%H%M%SZ", time.gmtime()) + invalid_root = LOGS / "_infra_invalid" + invalid_root.mkdir(parents=True, exist_ok=True) + src = LOGS / name + if src.exists(): + shutil.move(str(src), str(invalid_root / f"{name}-attempt{attempt}-{stamp}")) + ws = WORKSPACES / name + if ws.exists(): + ws_invalid = WORKSPACES / "_infra_invalid" + ws_invalid.mkdir(parents=True, exist_ok=True) + shutil.move(str(ws), str(ws_invalid / f"{name}-attempt{attempt}-{stamp}")) + + +def label_scroll_loop(log_dir: Path, check_output: str) -> None: + label = { + "primary": "identical-call-loop", + "sub_labels": ["scroll-loop"], + "notes": "Exact-PID SIGTERM after the published >=30 identical digit-stripped command rule.", + "labeler": "extended-suite-supervisor", + "labeled_at": time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()), + "substance_check": check_output[-4000:], + } + (log_dir / "label.json").write_text(json.dumps(label, indent=2) + "\n") + + +def label_missing_artifacts(log_dir: Path, stdout_path: Path, rc: int) -> str: + """Record a non-endpoint terminal outcome instead of silently retrying it.""" + log_dir.mkdir(parents=True, exist_ok=True) + try: + output = stdout_path.read_text(errors="replace") + except Exception: + output = "" + if "Traceback (most recent call last)" in output: + primary = "harness-crash" + sub_labels = ["operator-review-required", "not-auto-scored"] + else: + primary = "missing-artifacts" + sub_labels = ["model-terminal-failure"] + label = { + "schema_version": 1, + "primary": primary, + "sub_labels": sub_labels, + "notes": ( + "Harness exited without the required summary/archive and no sustained " + "endpoint outage was observed. A Python traceback stops the lane for " + "operator review; otherwise this is a preserved terminal model outcome." + ), + "labeler": "extended-suite-supervisor", + "labeled_at": time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()), + "return_code": rc, + "harness_output_tail": output[-8000:], + } + (log_dir / "label.json").write_text(json.dumps(label, indent=2) + "\n") + return primary + + +def harness_command(suite: dict, rep: int, max_model_len: int, + top_p: float, top_k: int) -> list[str]: + name = run_name(suite["id"], rep) + cmd = [ + "python3", str(TOOLING / "harness.py"), name, str(ROOT / suite["task"]), + "--model", MODEL, "--port", str(PORT), + "--temperature", str(suite["temperature"]), + "--top-p", str(top_p), + "--top-k", str(top_k), + "--stuck-threshold", str(suite["stuck_threshold"]), + "--max-model-len", str(max_model_len), + "--max-output-tokens-cap", str(suite["max_output_tokens_cap"]), + "--serving-manifest", str(SERVING_MANIFEST), + "--docker-socket", "--gpus", "all", + ] + if suite.get("require_git_tag"): + cmd.append("--require-git-tag") + if suite.get("input_from"): + source = WORKSPACES / run_name(suite["input_from"], rep) + cmd += ["--input-mount", str(source)] + if suite.get("input_path"): + cmd += ["--input-mount", str(Path(suite["input_path"]).resolve())] + return cmd + + +def wait_for_dependency(source_name: str) -> bool: + """Wait until the corresponding memo ships or gets a terminal label.""" + source_log = LOGS / source_name + while True: + if terminal_labeled_outcome(source_log): + return False + if completed(source_log): + return True + write_status( + phase="WAITING_FOR_DEPENDENCY", source_run=source_name, + lane=LANE_INDEX, port=PORT, + ) + time.sleep(30) + + +def supervise_one(suite: dict, rep: int, max_model_len: int, + top_p: float, top_k: int) -> None: + name = run_name(suite["id"], rep) + log_dir = LOGS / name + if suite.get("input_from"): + source_name = run_name(suite["input_from"], rep) + source_log = LOGS / source_name + if not wait_for_dependency(source_name): + log_dir.mkdir(parents=True, exist_ok=True) + source_label = json.loads((source_log / "label.json").read_text()) + label = { + "schema_version": 1, + "primary": "dependency-failure", + "sub_labels": ["input-run-did-not-ship"], + "notes": ( + f"Not launched because required input run {source_name} ended " + f"with terminal label {source_label.get('primary')}." + ), + "labeler": "extended-suite-supervisor", + "labeled_at": time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()), + "source_run": source_name, + "source_primary": source_label.get("primary"), + } + (log_dir / "label.json").write_text(json.dumps(label, indent=2) + "\n") + event("run_dependency_failure", run=name, source_run=source_name, + source_primary=source_label.get("primary")) + return + if completed(log_dir) and not infra_invalid(log_dir): + event("run_skip_complete", run=name) + return + for attempt in range(1, 4): + if not ensure_endpoint(): + time.sleep(60) + continue + if not git_clean(): + raise RuntimeError("extended suite worktree became dirty") + cmd = harness_command(suite, rep, max_model_len, top_p, top_k) + stdout_path = STATE / f"{name}-attempt{attempt}.log" + event("run_start", run=name, suite=suite["id"], rep=rep, attempt=attempt, command=cmd) + # Publish the run identity before spawning the harness; the telemetry + # logger independently attributes the live harness PID/port to its GPU. + write_status(phase="RUNNING", suite=suite["id"], run=name, + rep=rep, attempt=attempt, harness_pid=None, + lane=LANE_INDEX, port=PORT) + with stdout_path.open("a") as out: + proc = subprocess.Popen(cmd, cwd=ROOT, stdout=out, stderr=subprocess.STDOUT) + write_status(phase="RUNNING", suite=suite["id"], run=name, + rep=rep, attempt=attempt, harness_pid=proc.pid, + lane=LANE_INDEX, port=PORT) + last_substance = 0.0 + endpoint_down_since = None + killed_for_infra = False + while proc.poll() is None: + time.sleep(30) + transcript = log_dir / "transcript.jsonl" + now = time.time() + if transcript.exists() and now - last_substance >= 300: + last_substance = now + check = subprocess.run( + ["python3", str(SUBSTANCE), str(transcript)], + capture_output=True, text=True, check=False, + ) + event("substance", run=name, rc=check.returncode) + if check.returncode == 1: + proc.terminate() + try: + proc.wait(timeout=20) + except subprocess.TimeoutExpired: + proc.kill() + label_scroll_loop(log_dir, (check.stdout or "") + (check.stderr or "")) + subprocess.run(["docker", "rm", "-f", f"bench-sandbox-{name}"], + stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL) + event("run_terminal_pathology", run=name, pathology="scroll-loop") + return + if endpoint_up(): + endpoint_down_since = None + elif endpoint_down_since is None: + endpoint_down_since = now + elif now - endpoint_down_since > 90: + proc.terminate() + try: + proc.wait(timeout=20) + except subprocess.TimeoutExpired: + proc.kill() + killed_for_infra = True + event("run_killed_for_endpoint", run=name) + break + write_status(phase="RUNNING", suite=suite["id"], run=name, + rep=rep, attempt=attempt, harness_pid=proc.pid, + lane=LANE_INDEX, port=PORT) + rc = proc.wait() + if killed_for_infra or infra_invalid(log_dir): + event("run_infra_invalid", run=name, attempt=attempt, rc=rc) + archive_invalid(name, attempt) + ensure_endpoint() + continue + if completed(log_dir): + subprocess.run( + ["python3", str(TOOLING / "scripts" / "extract_cost.py"), str(log_dir)], + cwd=ROOT, check=False, + ) + event("run_complete", run=name, attempt=attempt, rc=rc) + return + primary = label_missing_artifacts(log_dir, stdout_path, rc) + event("run_terminal_failure", run=name, attempt=attempt, rc=rc, + primary=primary) + if primary == "harness-crash": + raise RuntimeError(f"{name} hit a harness traceback; operator review required") + return + raise RuntimeError(f"{name} exhausted three infrastructure retries") + + +def validate_matrix(matrix: dict) -> None: + if LANE_COUNT < 1 or not 0 <= LANE_INDEX < LANE_COUNT: + raise RuntimeError(f"invalid lane {LANE_INDEX}/{LANE_COUNT}") + lane_ports = [int(port) for port in matrix.get("lane_ports", [8000])] + if len(lane_ports) != LANE_COUNT or lane_ports[LANE_INDEX] != PORT: + raise RuntimeError( + f"lane wiring mismatch: lane={LANE_INDEX}/{LANE_COUNT} port={PORT} " + f"matrix_ports={lane_ports}" + ) + if not SERVING_MANIFEST.is_file(): + raise RuntimeError(f"serving manifest missing: {SERVING_MANIFEST}") + protocol = ROOT / matrix["substantive_audit_protocol"] + if not protocol.is_file(): + raise RuntimeError(f"substantive audit protocol missing: {protocol}") + protocol_hash = sha256(protocol) + if protocol_hash != matrix["substantive_audit_protocol_sha256"]: + raise RuntimeError(f"substantive audit protocol hash mismatch: {protocol_hash}") + for suite in matrix["suites"]: + task = ROOT / suite["task"] + actual = sha256(task) + if actual != suite["current_task_sha256"]: + raise RuntimeError(f"task hash mismatch for {suite['id']}: {actual}") + if suite.get("subject_pin"): + subject_pin = ROOT / suite["subject_pin"] + if not subject_pin.is_file(): + raise RuntimeError(f"subject pin missing for {suite['id']}: {subject_pin}") + actual_pin = sha256(subject_pin) + if actual_pin != suite.get("subject_pin_sha256"): + raise RuntimeError( + f"subject pin hash mismatch for {suite['id']}: {actual_pin}" + ) + if suite.get("input_from") and suite.get("input_path"): + raise RuntimeError(f"suite {suite['id']} cannot set both input_from and input_path") + if suite.get("input_path") and not Path(suite["input_path"]).is_dir(): + raise RuntimeError( + f"input fixture missing for {suite['id']}: {suite['input_path']}" + ) + + +def main() -> None: + matrix = json.loads(MATRIX_PATH.read_text()) + validate_matrix(matrix) + max_model_len = int(matrix["served_context_tokens"]) + if max_model_len <= 0: + raise RuntimeError("served_context_tokens must be positive") + top_p = float(matrix["top_p"]) + if not 0.0 < top_p <= 1.0: + raise RuntimeError("top_p must be in (0, 1]") + top_k = int(matrix["top_k"]) + if top_k <= 0: + raise RuntimeError("top_k must be positive") + if "--validate-only" in sys.argv[1:]: + print( + f"VALID: {len(matrix['suites'])} suites x N={matrix['replicates']} " + f"lane={LANE_INDEX}/{LANE_COUNT} port={PORT}" + ) + return + wait_for_microbench() + if not git_clean(): + raise SystemExit("refusing extended suites from dirty worktree") + all_jobs = [ + (suite, rep) + for suite in matrix["suites"] + for rep in range(1, matrix["replicates"] + 1) + ] + jobs = [job for ordinal, job in enumerate(all_jobs) if ordinal % LANE_COUNT == LANE_INDEX] + total = len(jobs) + done = 0 + for suite, rep in jobs: + supervise_one(suite, rep, max_model_len, top_p, top_k) + done += 1 + write_status(phase="RUNNING", done=done, total=total, + suite=suite["id"], rep=rep, lane=LANE_INDEX, port=PORT) + write_status(phase="COMPLETE", done=done, total=total, + lane=LANE_INDEX, port=PORT) + event("extended_lane_complete", done=done, total=total, + lane=LANE_INDEX, port=PORT) + + +if __name__ == "__main__": + main() diff --git a/tooling/run_gemma4_supplemental_telemetry.sh b/tooling/run_gemma4_supplemental_telemetry.sh new file mode 100755 index 00000000..ee154160 --- /dev/null +++ b/tooling/run_gemma4_supplemental_telemetry.sh @@ -0,0 +1,110 @@ +#!/usr/bin/env bash +# Produce telemetry-matched supplemental observations for the two canonical +# bug-fix runs that completed before the telemetry sidecar existed. These runs +# use a distinct label and never replace or enter the canonical quality cohort. +set -euo pipefail + +ROOT="$(cd "$(dirname "$0")/.." && pwd)" +LABEL="gemma4-31b-q4-telemetry-supplement" +MODEL="Gemma-4-31B-it-QAT-Q4_0" +MANIFEST="$ROOT/tooling/deployments/gemma4-31b-q4-tower2/benchmark-serving-manifest.json" +RUNNER="$ROOT/tooling/scripts/run_microbench.sh" +EXPECTED=( + "p1_bugfix_${LABEL}_v1" + "p1_bugfix_${LABEL}_v2" +) + +cd "$ROOT" +if [ -n "$(git status --porcelain)" ]; then + echo "ERROR: refusing supplemental telemetry from a dirty worktree" >&2 + exit 2 +fi +if systemctl --user is-active --quiet mmbt-gemma4-canonical-n3-r3.service; then + echo "ERROR: canonical N=3 service is still active" >&2 + exit 2 +fi +if pgrep -af 'bench_autopilot.py|run_microbench.sh|tooling/harness.py' | grep -v run_gemma4_supplemental_telemetry >/dev/null; then + echo "ERROR: another benchmark harness is active" >&2 + exit 2 +fi +for port in 8000 8001; do + observed="$(curl -fsS "http://127.0.0.1:${port}/v1/models" | jq -r '.data[0].id')" + if [ "$observed" != "$MODEL" ]; then + echo "ERROR: port $port serves $observed, expected $MODEL" >&2 + exit 2 + fi + if ! curl -fsS "http://127.0.0.1:${port}/slots" \ + | jq -e 'all(.[]; .is_processing == false)' >/dev/null; then + echo "ERROR: port $port has an active request; retry at an idle boundary" >&2 + exit 2 + fi +done +mapfile -t limits < <(nvidia-smi --query-gpu=power.limit --format=csv,noheader,nounits) +if [ "${limits[*]}" != "500.00 500.00" ]; then + echo "ERROR: GPU power limits are not both 500 W: ${limits[*]}" >&2 + exit 2 +fi + +for run in "${EXPECTED[@]}"; do + if [ -e "logs/$run" ]; then + echo "ERROR: supplemental run already exists and will not be overwritten: $run" >&2 + exit 2 + fi +done + +common_env=( + BENCH_LANE_COUNT=24 + BENCH_TEMP=1.0 + BENCH_TOP_P=0.95 + BENCH_TOP_K=64 + BENCH_MAX_OUTPUT_TOKENS_CAP=262144 + BENCH_SERVING_MANIFEST="$MANIFEST" +) + +# N=2 has 24 total ordinals. Lane-count 24 assigns ordinal 0 (bugfix v1) +# exclusively to index 0 and ordinal 1 (bugfix v2) exclusively to index 1. +env "${common_env[@]}" BENCH_LANE_INDEX=0 \ + bash "$RUNNER" "$MODEL" 8000 "$LABEL" 2 "" "" 262144 & +pid0=$! +env "${common_env[@]}" BENCH_LANE_INDEX=1 \ + bash "$RUNNER" "$MODEL" 8001 "$LABEL" 2 "" "" 262144 & +pid1=$! + +rc=0 +wait "$pid0" || rc=1 +wait "$pid1" || rc=1 +if [ "$rc" -ne 0 ]; then + echo "ERROR: a supplemental lane failed; preserve and audit both attempts" >&2 + exit 1 +fi + +for run in "${EXPECTED[@]}"; do + for required in receipt.json transcript.jsonl summary.json workspace_final.tar.gz; do + if [ ! -s "logs/$run/$required" ]; then + echo "ERROR: $run missing $required" >&2 + exit 1 + fi + done + python3 tooling/scripts/extract_cost.py "logs/$run" +done + +# The telemetry sidecar attributes and clips completed runs once per minute. +for _ in $(seq 1 36); do + ready=1 + for run in "${EXPECTED[@]}"; do + [ -s "logs/$run/cost.json" ] && [ -s "logs/$run/gpu_telemetry.json" ] || ready=0 + done + [ "$ready" -eq 1 ] && break + sleep 5 +done +for run in "${EXPECTED[@]}"; do + if [ ! -s "logs/$run/gpu_telemetry.json" ]; then + echo "ERROR: telemetry sidecar did not materialize $run/gpu_telemetry.json" >&2 + exit 1 + fi +done + +printf '%s\n' \ + "SUPPLEMENTAL_TELEMETRY_COMPLETE" \ + "canonical outcomes preserved: p1_bugfix_gemma4-31b-q4_v1/v2" \ + "supplements: ${EXPECTED[*]}" diff --git a/tooling/scripts/audit_frozen_75pr_run.py b/tooling/scripts/audit_frozen_75pr_run.py index 0eb3f7e2..085551c0 100644 --- a/tooling/scripts/audit_frozen_75pr_run.py +++ b/tooling/scripts/audit_frozen_75pr_run.py @@ -41,6 +41,14 @@ def command(*args: str) -> str: return subprocess.check_output(args, text=True).strip() +def read_text_or_empty(path: Path) -> str: + """Treat a missing required artifact as empty so the audit fails closed.""" + try: + return path.read_text(errors="replace") + except FileNotFoundError: + return "" + + def main() -> int: parser = argparse.ArgumentParser() parser.add_argument("log_dir", type=Path) @@ -78,10 +86,10 @@ def main() -> int: structural = json.loads(validation.stdout) prs = sorted((workspace / "prs").glob("pr-*"), key=lambda path: int(path.name[3:])) - reviews = {int(path.name[3:]): (path / "review.md").read_text(errors="replace") for path in prs} - traces = {int(path.name[3:]): (path / "trace.md").read_text(errors="replace") for path in prs} - diffs = {int(path.name[3:]): (path / "diff-analysis.md").read_text(errors="replace") for path in prs} - verdicts = {int(path.name[3:]): (path / "verdict.md").read_text(errors="replace") for path in prs} + reviews = {int(path.name[3:]): read_text_or_empty(path / "review.md") for path in prs} + traces = {int(path.name[3:]): read_text_or_empty(path / "trace.md") for path in prs} + diffs = {int(path.name[3:]): read_text_or_empty(path / "diff-analysis.md") for path in prs} + verdicts = {int(path.name[3:]): read_text_or_empty(path / "verdict.md") for path in prs} test_prs = sorted( { int(path.parents[1].name[3:]) @@ -97,7 +105,7 @@ def main() -> int: except Exception: continue tool_events = sum(row.get("type") == "tool" for row in transcript_rows) - tool_log = (workspace / "tool-log.md").read_text(errors="replace") + tool_log = read_text_or_empty(workspace / "tool-log.md") numbered_tool_entries = sum( bool(re.match(r"^\s*\d+[.)]\s+", line)) for line in tool_log.splitlines() ) diff --git a/tooling/scripts/check_substance.py b/tooling/scripts/check_substance.py index 354a405a..f726c853 100755 --- a/tooling/scripts/check_substance.py +++ b/tooling/scripts/check_substance.py @@ -45,6 +45,26 @@ def digit_strip(s: str) -> str: return re.sub(r"\d+", "#", s) +def command_template(entry: dict) -> str: + """Return a stable template for valid and malformed tool commands. + + Tool arguments are model output and therefore untrusted. Preserve a + malformed value as a type-qualified template so repeated malformed calls + can still trip the loop detector without crashing the supervisor. + """ + args = entry.get("args") or {} + if not isinstance(args, dict): + return f"" + command = args.get("command", "") or "" + if not isinstance(command, str): + try: + rendered = json.dumps(command, sort_keys=True, separators=(",", ":")) + except (TypeError, ValueError): + rendered = repr(command) + return f""[:300] + return digit_strip(command)[:300] + + def main() -> int: ap = argparse.ArgumentParser() ap.add_argument("transcript", help="Path to transcript.jsonl") @@ -91,8 +111,7 @@ def main() -> int: return 0 window = tool_calls[-args.window:] - templates = [digit_strip((e.get("args") or {}).get("command", "") or "")[:300] - for e in window] + templates = [command_template(entry) for entry in window] unique = len(set(templates)) streak = 1 diff --git a/tooling/scripts/run_microbench.sh b/tooling/scripts/run_microbench.sh index 359e1a99..3eaccddb 100755 --- a/tooling/scripts/run_microbench.sh +++ b/tooling/scripts/run_microbench.sh @@ -64,6 +64,25 @@ THINKING_FLAG="" MAXLEN="${7:-}" # optional: served context window (e.g. 131072 for the 397B GGUF on llama.cpp) MAXLEN_FLAG="" [ -n "$MAXLEN" ] && MAXLEN_FLAG="--max-model-len $MAXLEN" +MAX_OUTPUT_TOKENS_CAP="${BENCH_MAX_OUTPUT_TOKENS_CAP:-180000}" +if ! [[ "$MAX_OUTPUT_TOKENS_CAP" =~ ^[1-9][0-9]*$ ]]; then + echo "ERROR: invalid BENCH_MAX_OUTPUT_TOKENS_CAP=$MAX_OUTPUT_TOKENS_CAP" >&2 + exit 2 +fi +MAX_OUTPUT_TOKENS_CAP_FLAG="--max-output-tokens-cap $MAX_OUTPUT_TOKENS_CAP" +SERVING_MANIFEST_FLAG="" +[ -n "${BENCH_SERVING_MANIFEST:-}" ] && SERVING_MANIFEST_FLAG="--serving-manifest $BENCH_SERVING_MANIFEST" + +# Deterministic run sharding allows one supervisor to drive independent GPU +# replicas without duplicate claims. The default remains the historical single +# lane. A run's zero-based ordinal modulo BENCH_LANE_COUNT owns the run. +LANE_INDEX="${BENCH_LANE_INDEX:-0}" +LANE_COUNT="${BENCH_LANE_COUNT:-1}" +if ! [[ "$LANE_INDEX" =~ ^[0-9]+$ && "$LANE_COUNT" =~ ^[1-9][0-9]*$ ]] || \ + (( LANE_INDEX >= LANE_COUNT )); then + echo "ERROR: invalid BENCH_LANE_INDEX=$LANE_INDEX BENCH_LANE_COUNT=$LANE_COUNT" >&2 + exit 2 +fi # Sampling overrides via env (default keeps the cross-model temp=0.3 protocol). Some models # specify a required operating point and loop under low temp — e.g. MiniMax-M2 card mandates @@ -130,9 +149,15 @@ TASKS=( ) TOTAL_RUNS=$(( ${#TASKS[@]} * N )) +if (( LANE_INDEX < TOTAL_RUNS )); then + ASSIGNED_RUNS=$(( (TOTAL_RUNS - 1 - LANE_INDEX) / LANE_COUNT + 1 )) +else + ASSIGNED_RUNS=0 +fi START_T=$(date +%s) echo "==> Microbench chain: $TOTAL_RUNS runs (${#TASKS[@]} task families × N=$N)" +echo " lane: $LANE_INDEX/$LANE_COUNT ($ASSIGNED_RUNS assigned runs)" echo " model: $MODEL (label: $LABEL)" echo " port: $PORT" echo " started: $(date -u +%Y-%m-%dT%H:%M:%SZ)" @@ -141,10 +166,16 @@ echo "" DONE=0 SKIPPED=0 FAILED=0 +RUN_ORDINAL=0 for entry in "${TASKS[@]}"; do IFS='|' read -r task_short task_file input_dir <<< "$entry" for v in $(seq 1 "$N"); do + ordinal="$RUN_ORDINAL" + RUN_ORDINAL=$((RUN_ORDINAL + 1)) + if (( ordinal % LANE_COUNT != LANE_INDEX )); then + continue + fi run_name="${task_short}_${LABEL}_v${v}" DONE=$((DONE + 1)) @@ -184,6 +215,8 @@ for entry in "${TASKS[@]}"; do $REASONING_FLAG \ $THINKING_FLAG \ $MAXLEN_FLAG \ + $MAX_OUTPUT_TOKENS_CAP_FLAG \ + $SERVING_MANIFEST_FLAG \ $INPUT_FLAG \ --docker-socket \ --gpus all 2>&1 | tail -3 @@ -202,7 +235,7 @@ M=$(( (ELAPSED % 3600) / 60 )) echo "" echo "==> Microbench chain complete" -echo " total: $TOTAL_RUNS runs" +echo " total: $ASSIGNED_RUNS assigned runs (of $TOTAL_RUNS campaign runs)" echo " skipped: $SKIPPED (already complete from prior invocations)" echo " failed: $FAILED (see logs//transcript.jsonl)" echo " elapsed: ${H}h${M}m" diff --git a/tooling/scripts/test_audit_frozen_75pr_run.py b/tooling/scripts/test_audit_frozen_75pr_run.py new file mode 100644 index 00000000..e1fbc406 --- /dev/null +++ b/tooling/scripts/test_audit_frozen_75pr_run.py @@ -0,0 +1,13 @@ +from pathlib import Path + +import audit_frozen_75pr_run as audit + + +def test_read_text_or_empty_reads_existing_file(tmp_path: Path) -> None: + path = tmp_path / "review.md" + path.write_text("review") + assert audit.read_text_or_empty(path) == "review" + + +def test_read_text_or_empty_fails_closed_for_missing_file(tmp_path: Path) -> None: + assert audit.read_text_or_empty(tmp_path / "missing.md") == "" diff --git a/tooling/snapshot_gemma4_telemetry.sh b/tooling/snapshot_gemma4_telemetry.sh new file mode 100755 index 00000000..baf7d5a4 --- /dev/null +++ b/tooling/snapshot_gemma4_telemetry.sh @@ -0,0 +1,73 @@ +#!/usr/bin/env bash +# Take an immutable, newline-complete telemetry snapshot at a clean cohort boundary. +set -euo pipefail + +CSV=/home/michael/gemma4-campaign-state/telemetry/gemma4-31b-q4-gpu.csv +SNAPSHOT_ROOT=/home/michael/gemma4-campaign-state/telemetry/snapshots +LOGGER=mmbt-gemma4-power-logger.service +SIDECAR=mmbt-gemma4-telemetry-sidecar.service + +if [ "$#" -ne 1 ]; then + echo "usage: $0 $SNAPSHOT_ROOT/.csv" >&2 + exit 2 +fi +destination=$1 +case "$destination" in + "$SNAPSHOT_ROOT"/*.csv) ;; + *) echo "ERROR: snapshot destination must be an absolute CSV under $SNAPSHOT_ROOT" >&2; exit 2 ;; +esac +if [ -e "$destination" ] || [ -e "$destination.tmp" ]; then + echo "ERROR: refusing to overwrite telemetry snapshot: $destination" >&2 + exit 2 +fi +if pgrep -af 'bench_autopilot.py|run_microbench.sh|tooling/harness.py|run_gemma4_extended_suites.py' \ + | grep -v snapshot_gemma4_telemetry >/dev/null; then + echo "ERROR: benchmark work is active; telemetry snapshot requires a clean boundary" >&2 + exit 2 +fi +if [ ! -s "$CSV" ]; then + echo "ERROR: live telemetry CSV is missing or empty: $CSV" >&2 + exit 2 +fi + +logger_was_active=0 +sidecar_was_active=0 +systemctl --user is-active --quiet "$LOGGER" && logger_was_active=1 +systemctl --user is-active --quiet "$SIDECAR" && sidecar_was_active=1 +restart_services() { + [ "$logger_was_active" -eq 0 ] || systemctl --user start "$LOGGER" + [ "$sidecar_was_active" -eq 0 ] || systemctl --user start "$SIDECAR" +} +trap restart_services EXIT + +# Stop the writer before copying so the snapshot cannot end with a partial row. +systemctl --user stop "$SIDECAR" "$LOGGER" +mkdir -p "$SNAPSHOT_ROOT" +cp --reflink=auto --preserve=mode,timestamps "$CSV" "$destination.tmp" +python3 - "$destination.tmp" <<'PY' +import csv +import sys +from pathlib import Path + +path = Path(sys.argv[1]) +raw = path.read_bytes() +if not raw.endswith(b"\n"): + raise SystemExit("telemetry snapshot does not end at a complete line") +with path.open(newline="") as handle: + rows = list(csv.reader(handle)) +expected = [ + "ts", "gpu", "endpoint_port", "power_limit_w", "power_w", "util_sm", + "util_mem", "mem_used_mib", "temp_c", "sm_clk_mhz", "cell", + "harness_pid", "cpu_package_power_w", +] +if not rows or rows[0] != expected: + raise SystemExit("telemetry snapshot header drift") +if len(rows) < 3: + raise SystemExit("telemetry snapshot has no paired GPU samples") +PY +mv "$destination.tmp" "$destination" +sha256sum "$destination" + +restart_services +trap - EXIT +printf '%s\n' TELEMETRY_SNAPSHOT_COMPLETE diff --git a/tooling/summarize_gemma4_campaign.py b/tooling/summarize_gemma4_campaign.py new file mode 100755 index 00000000..725b6f7e --- /dev/null +++ b/tooling/summarize_gemma4_campaign.py @@ -0,0 +1,233 @@ +#!/usr/bin/env python3 +"""Generate compact raw N=3/N=10 Gemma scorecards from immutable run evidence.""" +from __future__ import annotations + +import argparse +import json +import statistics +from collections import Counter +from datetime import datetime, timezone +from pathlib import Path + + +TASKS = [ + "p1_bugfix", "p1_testwrite", "p1_refactor", "p2_extract", "p2_ci", + "p2_hallucination", "p2_triage", "p3_doc", "p3_business", + "p3_market", "p3_writing", "p3_pm", +] +PASS_VERDICTS = {"PASS", "STRUCTURAL_PASS"} + + +def read_json(path: Path): + try: + return json.loads(path.read_text()) + except (OSError, json.JSONDecodeError): + return None + + +def median(values): + return round(statistics.median(values), 3) if values else None + + +def finish_reason(record: dict) -> str: + summary = record.get("summary") or {} + if summary.get("finish_reason"): + return summary["finish_reason"] + if record.get("terminal_label"): + return f"terminal:{record['terminal_label']}" + return "MISSING" + + +def quality_outcome(record: dict) -> str: + if record.get("verdict"): + return record["verdict"] + if record.get("terminal_label"): + return f"TERMINAL:{record['terminal_label']}" + return "MISSING" + + +def aggregate_records(records: list[dict]) -> dict: + completed = [ + record for record in records + if record.get("summary") or record.get("terminal_label") + ] + normal_completed = [record for record in records if record.get("summary")] + terminal = [record for record in records if record.get("terminal_label")] + graded = [record for record in records if record.get("verdict")] + scored = [record for record in records if record.get("verdict") or record.get("terminal_label")] + telemetry = [record["telemetry"] for record in records if record.get("telemetry")] + per_task = {} + for task in TASKS: + task_records = [record for record in records if record["task"] == task] + verdicts = Counter(record.get("verdict") or "MISSING" for record in task_records) + outcomes = Counter(quality_outcome(record) for record in task_records) + finishes = Counter(finish_reason(record) for record in task_records) + per_task[task] = { + "runs": len(task_records), + "graded": sum(record.get("verdict") is not None for record in task_records), + "scored_outcomes": sum( + record.get("verdict") is not None or record.get("terminal_label") is not None + for record in task_records + ), + "raw_passes": sum(record.get("verdict") in PASS_VERDICTS for record in task_records), + "verdicts": dict(sorted(verdicts.items())), + "quality_outcomes": dict(sorted(outcomes.items())), + "finish_reasons": dict(sorted(finishes.items())), + } + model_tps = [ + record["cost"]["throughput"]["completion_tps_avg"] + for record in records + if isinstance((((record.get("cost") or {}).get("throughput") or {}).get("completion_tps_avg")), (int, float)) + ] + wall_s = [] + completion_tokens = [] + for record in completed: + cost = record.get("cost") or {} + summary = record.get("summary") or {} + wall = cost.get("wall_s", summary.get("elapsed_s")) + tokens = (cost.get("tokens") or {}).get( + "completion_total", summary.get("total_completion_tokens") + ) + if isinstance(wall, (int, float)): + wall_s.append(wall) + if isinstance(tokens, (int, float)): + completion_tokens.append(tokens) + return { + "runs": len(records), + "completed": len(completed), + "normal_completed": len(normal_completed), + "terminal_outcomes": len(terminal), + "graded": len(graded), + "scored_outcomes": len(scored), + "raw_passes": sum(record.get("verdict") in PASS_VERDICTS for record in records), + "raw_pass_rate": ( + round(sum(record.get("verdict") in PASS_VERDICTS for record in records) / len(scored), 6) + if scored else None + ), + "verdicts": dict(sorted(Counter(record.get("verdict") or "MISSING" for record in records).items())), + "quality_outcomes": dict(sorted(Counter(quality_outcome(record) for record in records).items())), + "finish_reasons": dict(sorted(Counter(finish_reason(record) for record in records).items())), + "wall_s": {"sum": round(sum(wall_s), 3), "median": median(wall_s), "max": max(wall_s) if wall_s else None}, + "completion_tokens": { + "sum": sum(completion_tokens), "median": median(completion_tokens), + "max": max(completion_tokens) if completion_tokens else None, + }, + "model_call_completion_tps": {"median": median(model_tps), "min": min(model_tps) if model_tps else None, + "max": max(model_tps) if model_tps else None}, + "telemetry": { + "runs": len(telemetry), + "coverage_mean": round(sum(row["coverage"] for row in telemetry) / len(telemetry), 6) if telemetry else None, + "active_gpu_mean_power_w_mean": round(sum(row["mean_power_w"] for row in telemetry) / len(telemetry), 3) if telemetry else None, + "active_gpu_mean_sm_util_pct_mean": round(sum(row["mean_sm_util_pct"] for row in telemetry) / len(telemetry), 3) if telemetry else None, + "max_temp_c": max((row["max_temp_c"] for row in telemetry), default=None), + }, + "per_task": per_task, + } + + +def markdown(document: dict) -> str: + aggregate = document["aggregate"] + target_n = document["target_n"] + lines = [ + f"# Gemma 4 31B Q4 raw canonical scorecard (N={target_n})", + "", + "> Raw grader verdicts only. Any reproducible correction is a separate overlay tied to unchanged archive and grader hashes.", + "", + f"- Evidence-complete runs: {aggregate['completed']}/{aggregate['runs']}", + f"- Normal completed workspaces: {aggregate['normal_completed']}/{aggregate['runs']}", + f"- Explicit terminal outcomes: {aggregate['terminal_outcomes']}/{aggregate['runs']}", + f"- Graded runs: {aggregate['graded']}/{aggregate['runs']}", + f"- Raw pass-equivalent outcomes: {aggregate['raw_passes']}/{aggregate['scored_outcomes']}", + f"- Median model-call completion throughput: {aggregate['model_call_completion_tps']['median']} tok/s", + f"- Median cell wall time: {aggregate['wall_s']['median']} s", + f"- Telemetry-complete runs: {aggregate['telemetry']['runs']}/{aggregate['runs']}", + "", + "| Task | Raw pass | Scored | Finish reasons | Quality outcomes |", + "|---|---:|---:|---|---|", + ] + for task in TASKS: + row = aggregate["per_task"][task] + finishes = ", ".join(f"{key}:{value}" for key, value in row["finish_reasons"].items()) + outcomes = ", ".join(f"{key}:{value}" for key, value in row["quality_outcomes"].items()) + lines.append(f"| `{task}` | {row['raw_passes']}/{row['scored_outcomes']} | {row['scored_outcomes']}/{row['runs']} | {finishes} | {outcomes} |") + lines += [ + "", + "## Methodology boundary", + "", + "A `done_signal` is a finish behavior, not a pass. `PASS` and `STRUCTURAL_PASS` count only as raw pass-equivalent grader verdicts. A preserved terminal label is reported as a distinct non-pass quality outcome, never fabricated into a normal grader verdict. Model-call throughput excludes tool execution; wall time includes it. Telemetry is per attributed replica GPU, while CPU package power is shared host context and AC wall power is unavailable to software.", + "", + ] + return "\n".join(lines) + + +def main() -> None: + parser = argparse.ArgumentParser() + parser.add_argument("--root", type=Path, default=Path(__file__).resolve().parents[1]) + parser.add_argument("--label", default="gemma4-31b-q4") + parser.add_argument("--target-n", type=int, required=True) + parser.add_argument("--allow-pretelemetry-run", action="append", default=[]) + parser.add_argument("--output-json", required=True, type=Path) + parser.add_argument("--output-markdown", required=True, type=Path) + args = parser.parse_args() + root = args.root.resolve() + allowed = set(args.allow_pretelemetry_run) + records = [] + errors = [] + for task in TASKS: + for rep in range(1, args.target_n + 1): + name = f"{task}_{args.label}_v{rep}" + run = root / "logs" / name + summary = read_json(run / "summary.json") + grade = read_json(run / "grade.json") + cost = read_json(run / "cost.json") + label_doc = read_json(run / "label.json") + terminal_label = ( + label_doc.get("primary") + if summary is None and isinstance(label_doc, dict) else None + ) + telemetry_raw = read_json(run / "gpu_telemetry.json") + if summary is None and not terminal_label: + errors.append(f"{name}: missing summary.json or explicit terminal label") + if grade is None and not terminal_label: + errors.append(f"{name}: missing grade.json") + if cost is None: + errors.append(f"{name}: missing cost.json") + if telemetry_raw is None and (name not in allowed or terminal_label): + errors.append(f"{name}: missing gpu_telemetry.json") + telemetry = None + if telemetry_raw: + telemetry = { + "coverage": telemetry_raw["sampling"]["coverage_fraction_of_wall"], + "mean_power_w": telemetry_raw["active_gpu"]["mean_power_w"], + "mean_sm_util_pct": telemetry_raw["active_gpu"]["mean_sm_util_pct"], + "max_temp_c": telemetry_raw["active_gpu"]["max_temp_c"], + } + records.append({ + "run_name": name, "task": task, "replicate": rep, + "summary": summary, "verdict": grade.get("verdict") if grade else None, + "terminal_label": terminal_label, "label": label_doc, + "cost": cost, "telemetry": telemetry, + }) + document = { + "schema_version": 1, + "generated_at": datetime.now(timezone.utc).isoformat(), + "label": args.label, + "target_n": args.target_n, + "passed": not errors, + "errors": errors, + "aggregate": aggregate_records(records), + "runs": records, + } + args.output_json.parent.mkdir(parents=True, exist_ok=True) + args.output_markdown.parent.mkdir(parents=True, exist_ok=True) + args.output_json.write_text(json.dumps(document, indent=2) + "\n") + args.output_markdown.write_text(markdown(document)) + print(json.dumps({ + "passed": document["passed"], "runs": len(records), + "raw_passes": document["aggregate"]["raw_passes"], "errors": len(errors), + }, sort_keys=True)) + raise SystemExit(0 if document["passed"] else 1) + + +if __name__ == "__main__": + main() diff --git a/tooling/test_analyze_gemma4_queueing.py b/tooling/test_analyze_gemma4_queueing.py new file mode 100644 index 00000000..82cf34ac --- /dev/null +++ b/tooling/test_analyze_gemma4_queueing.py @@ -0,0 +1,53 @@ +#!/usr/bin/env python3 +import importlib.util +import json +from pathlib import Path + + +SCRIPT = Path(__file__).with_name("analyze_gemma4_queueing.py") +SPEC = importlib.util.spec_from_file_location("queueing", SCRIPT) +MODULE = importlib.util.module_from_spec(SPEC) +assert SPEC.loader is not None +SPEC.loader.exec_module(MODULE) + + +def request(wall, prompt_ms=1000, predicted_ms=4000): + return { + "wall_seconds": wall, + "timings": {"prompt_ms": prompt_ms, "predicted_ms": predicted_ms}, + "tokens_evaluated": 100, + "tokens_predicted": 200, + } + + +def test_queueing_analyzer_separates_service_waves_and_labels_estimate(tmp_path): + source = tmp_path / "summary.json" + source.write_text(json.dumps({ + "concurrency_results": [{ + "concurrency": 8, + "aggregate_decode_tokens_per_second_wall": 100.0, + "requests": [ + request(7.0), request(7.2), request(7.1), request(7.3), + request(14.0), request(14.2), request(14.1), request(14.3), + ], + }], + })) + result = MODULE.analyze(source, slots=4, concurrency=8) + assert result["first_wave"]["median_wall_s"] == 7.15 + assert result["queued_second_wave"]["median_wall_s"] == 14.15 + assert result["derived"]["second_wave_wall_penalty_s"] == 7.0 + assert result["derived"]["estimated_queue_wait_delta_s"] == 7.0 + assert "estimate, not a direct server-side queue timestamp" in result["methodology"] + + +def test_queueing_analyzer_rejects_incomplete_request_inventory(tmp_path): + source = tmp_path / "summary.json" + source.write_text(json.dumps({ + "concurrency_results": [{"concurrency": 8, "requests": [request(7.0)]}], + })) + try: + MODULE.analyze(source, slots=4, concurrency=8) + except ValueError as exc: + assert "expected 8 request results" in str(exc) + else: + raise AssertionError("incomplete concurrency evidence was accepted") diff --git a/tooling/test_audit_gemma4_campaign.py b/tooling/test_audit_gemma4_campaign.py new file mode 100644 index 00000000..1291e978 --- /dev/null +++ b/tooling/test_audit_gemma4_campaign.py @@ -0,0 +1,209 @@ +#!/usr/bin/env python3 +import importlib.util +import json +from pathlib import Path + + +SCRIPT = Path(__file__).with_name("audit_gemma4_campaign.py") +SPEC = importlib.util.spec_from_file_location("audit", SCRIPT) +AUDIT = importlib.util.module_from_spec(SPEC) +assert SPEC.loader is not None +SPEC.loader.exec_module(AUDIT) + + +def write_json(path, value): + path.write_text(json.dumps(value)) + + +def make_run(tmp_path, telemetry=True): + run = tmp_path / "p2_extract_gemma4-31b-q4_v1" + run.mkdir() + write_json(run / "receipt.json", { + "harness": {"git_dirty": False, "git_sha": "abc"}, + "vllm": {"served_model_name": AUDIT.MODEL, "api_url": "http://127.0.0.1:8000/v1/chat/completions"}, + "inference_request_defaults": { + "temperature": 1.0, "top_p": 0.95, "top_k": 64, + "max_model_len": 262144, "max_output_tokens_cap": 262144, + }, + "serving": { + "manifest": {"payload": {"artifact": {"sha256": AUDIT.MODEL_SHA}}}, + "host_processes": [{"exe_sha256": AUDIT.SERVER_SHA}], + "endpoint_models": {"payload": {"data": [{"id": AUDIT.MODEL}]}}, + }, + "hardware": {"nvidia_smi": ["0, GPU, driver, 500.00 W", "1, GPU, driver, 500.00 W"]}, + }) + write_json(run / "summary.json", { + "model": AUDIT.MODEL, "finish_reason": "done_signal", "elapsed_s": 10, + "iterations": 1, "total_completion_tokens": 20, "total_prompt_tokens": 30, + }) + (run / "transcript.jsonl").write_text(json.dumps({ + "type": "model", "finish_reason": "stop", "completion_tokens": 20, + }) + "\n") + (run / "workspace_final.tar.gz").write_bytes(b"archive") + write_json(run / "cost.json", { + "run_name": run.name, "model": AUDIT.MODEL, "wall_s": 10, + "iters": 1, "tokens": {"completion_total": 20, "prompt_total": 30}, + }) + if telemetry: + write_json(run / "gpu_telemetry.json", { + "attribution": {"active_gpu_ids_observed": ["0"]}, + "sampling": {"coverage_fraction_of_wall": 0.95}, + "active_gpu": {"configured_cap_w": 500.0, "mean_power_w": 400, + "mean_sm_util_pct": 90, "max_temp_c": 70}, + "cpu_package_shared_context": {"mean_power_w": 100}, + }) + return run + + +def test_complete_run_passes_strict_evidence_audit(tmp_path): + record, errors, warnings = AUDIT.audit_run(make_run(tmp_path), False, False) + assert errors == [] + assert warnings == [] + assert record["telemetry"]["gpu"] == ["0"] + + +def test_deployed_default_root_is_repository_parent_of_tooling(): + deployed = Path("/home/michael/bench-gemma4-31b-q4/tooling/audit_gemma4_campaign.py") + assert AUDIT.default_root(deployed) == Path("/home/michael/bench-gemma4-31b-q4") + + +def test_pretelemetry_exception_is_explicit_and_warned(tmp_path): + record, errors, warnings = AUDIT.audit_run(make_run(tmp_path, telemetry=False), True, False) + assert errors == [] + assert warnings == ["pre-telemetry valid attempt; supplemental telemetry required"] + assert record["telemetry"] is None + + +def test_wrong_sampling_fails(tmp_path): + run = make_run(tmp_path) + receipt = json.loads((run / "receipt.json").read_text()) + receipt["inference_request_defaults"]["temperature"] = 0.3 + write_json(run / "receipt.json", receipt) + _, errors, _ = AUDIT.audit_run(run, False, False) + assert "wrong temperature: 0.3" in errors + + +def test_supplement_mapping_parser_is_fail_closed(): + mappings, errors = AUDIT.parse_supplement_mappings([ + "canonical_v1=supplement_v1", + "canonical_v1=duplicate", + "malformed", + ]) + assert mappings == {"canonical_v1": "supplement_v1"} + assert errors == [ + "duplicate pretelemetry supplement mapping for canonical_v1", + "invalid pretelemetry supplement mapping 'malformed'; expected " + "CANONICAL_RUN=SUPPLEMENTAL_RUN", + ] + + +def test_terminal_outcome_requires_preserved_receipt_and_transcript(tmp_path): + run = tmp_path / "p3_market_gemma4-31b-q4_v1" + run.mkdir() + write_json(run / "label.json", {"primary": "identical-call-loop"}) + _, errors, _ = AUDIT.audit_run(run, False, False) + assert errors == [ + "terminal outcome missing preserved receipt.json", + "terminal outcome missing preserved transcript.jsonl", + "terminal outcome missing preserved cost.json", + "terminal outcome missing preserved gpu_telemetry.json", + ] + + +def test_terminal_outcome_passes_only_with_full_identity_cost_and_telemetry(tmp_path): + run = make_run(tmp_path) + (run / "summary.json").unlink() + (run / "workspace_final.tar.gz").unlink() + write_json(run / "label.json", { + "primary": "identical-call-loop", + "labeled_at": "2026-08-02T00:00:10+00:00", + }) + telemetry = json.loads((run / "gpu_telemetry.json").read_text()) + telemetry["sampling"]["window_source"] = ( + "receipt.json:captured_at..label.json:labeled_at" + ) + write_json(run / "gpu_telemetry.json", telemetry) + record, errors, warnings = AUDIT.audit_run(run, False, True) + assert errors == [] + assert warnings == [] + assert record["outcome_kind"] == "terminal-label" + assert record["finish_reason"] == "terminal:identical-call-loop" + assert record["grade_verdict"] is None + + +def test_dependency_failure_requires_source_evidence(tmp_path): + run = tmp_path / "gemma4-31b-q4_board_pres_v1" + run.mkdir() + write_json(run / "label.json", {"primary": "dependency-failure"}) + _, errors, _ = AUDIT.audit_run(run, False, False) + assert errors == ["dependency-failure label lacks source evidence"] + + +def test_invalid_attempt_requires_hash_tied_classification_and_completed_replacement(tmp_path): + invalid_root = tmp_path / "logs" / "_invalid" + attempt = invalid_root / "p1_refactor_gemma4-31b-q4_v3-timeout" + attempt.mkdir(parents=True) + (attempt / "receipt.json").write_bytes(b"receipt") + incident = tmp_path / "incident.json" + write_json(incident, {"classification": "test"}) + classifications = tmp_path / "classifications.json" + write_json(classifications, { + "attempts": { + attempt.name: { + "source_run": "p1_refactor_gemma4-31b-q4_v3", + "classification": "infrastructure-invalid", + "reason_code": "server-timeout", + "incident_document": incident.name, + "classified_before_grade": True, + "affirmative_evidence": ["exact timeout"], + "expected_files": { + "receipt.json": AUDIT.sha256(attempt / "receipt.json"), + }, + "replacement": { + "required": True, + "canonical_run": "p1_refactor_gemma4-31b-q4_v3", + "status": "completed", + }, + }, + }, + }) + records, source, errors = AUDIT.audit_invalid_attempts( + invalid_root, classifications, "gemma4-31b-q4", + ) + assert errors == [] + assert records[0]["files"]["receipt.json"]["sha256"] == AUDIT.sha256( + attempt / "receipt.json" + ) + assert source["sha256"] == AUDIT.sha256(classifications) + + document = json.loads(classifications.read_text()) + document["attempts"][attempt.name]["replacement"]["status"] = "pending" + write_json(classifications, document) + _, _, errors = AUDIT.audit_invalid_attempts( + invalid_root, classifications, "gemma4-31b-q4", + ) + assert errors == [f"{attempt.name}: exact canonical replacement is not completed"] + + +def test_timeout_replacement_requires_corrected_server_control(tmp_path): + logs = tmp_path / "logs" + logs.mkdir() + run = make_run(logs) + receipt = json.loads((run / "receipt.json").read_text()) + receipt["serving"]["host_processes"][0]["argv"] = [ + "llama-server", "--timeout", "14400", + ] + write_json(run / "receipt.json", receipt) + record = { + "attempt": "p2_extract_gemma4-31b-q4_v1-timeout", + "classification": { + "reason_code": "server-transport-timeout-below-native-envelope", + "replacement": {"canonical_run": run.name}, + }, + } + assert AUDIT.audit_replacement_controls(logs, [record]) == [] + receipt["serving"]["host_processes"][0]["argv"][-1] = "3600" + write_json(run / "receipt.json", receipt) + assert AUDIT.audit_replacement_controls(logs, [record]) == [ + f"{record['attempt']}: replacement does not prove a >=14400-second server timeout" + ] diff --git a/tooling/test_audit_gemma4_extended.py b/tooling/test_audit_gemma4_extended.py new file mode 100644 index 00000000..3f906fe4 --- /dev/null +++ b/tooling/test_audit_gemma4_extended.py @@ -0,0 +1,75 @@ +#!/usr/bin/env python3 +import importlib.util +import io +import json +import tarfile +from pathlib import Path + + +SCRIPT = Path(__file__).with_name("audit_gemma4_extended.py") +SPEC = importlib.util.spec_from_file_location("audit_extended", SCRIPT) +MODULE = importlib.util.module_from_spec(SPEC) +assert SPEC.loader is not None +SPEC.loader.exec_module(MODULE) + + +def test_expected_run_names_are_stable(): + assert MODULE.expected_run_name("dreamserver-1-pr-audit", 2) == "n1_gemma4-31b-q4_v2" + assert MODULE.expected_run_name("wallstreet-investment-memo", 3) == \ + "gemma4-31b-q4_invest_memo_v3" + assert MODULE.expected_run_name("wallstreet-board-presentation", 1) == \ + "gemma4-31b-q4_board_pres_v1" + assert MODULE.expected_run_name("dreamserver-75-pr-audit", 3) == \ + "gemma4-31b-q4_75pr_v3" + + +def test_dependency_terminal_label_must_match_same_replicate(tmp_path, monkeypatch): + logs = tmp_path / "logs" / "gemma4-31b-q4_board_pres_v2" + logs.mkdir(parents=True) + (logs / "label.json").write_text(json.dumps({ + "primary": "dependency-failure", + "source_run": "gemma4-31b-q4_invest_memo_v1", + "source_primary": "model-terminal-failure", + })) + record, errors, _ = MODULE.audit_extended_run( + tmp_path, + { + "id": "wallstreet-board-presentation", + "input_from": "wallstreet-investment-memo", + "current_task_sha256": "unused", + }, + rep=2, + ordinal=5, + lane_ports=[8000, 8001], + ) + assert record["terminal_label"] == "dependency-failure" + assert errors == [ + "dependency-failure source 'gemma4-31b-q4_invest_memo_v1' != " + "'gemma4-31b-q4_invest_memo_v2'" + ] + + +def test_subject_ref_audit_requires_every_pinned_ref(tmp_path): + pin = tmp_path / "subject.json" + pin.write_text("{}") + run = tmp_path / "logs" / "n1_gemma4-31b-q4_v1" + run.mkdir(parents=True) + archive = run / "workspace_final.tar.gz" + payload = b"base abcdef12 head 12345678\n" + with tarfile.open(archive, "w:gz") as handle: + info = tarfile.TarInfo("audit/trace.md") + info.size = len(payload) + handle.addfile(info, io.BytesIO(payload)) + suite = { + "subject_pin": "subject.json", + "subject_pin_sha256": MODULE.sha256(pin), + "required_subject_shas": [ + "abcdef12aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa", + "12345678bbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbb", + "deadbeefcccccccccccccccccccccccccccccccc", + ], + } + assert MODULE.audit_subject_refs(tmp_path, suite, run) == [ + "artifact does not identify pinned subject ref " + "deadbeefcccccccccccccccccccccccccccccccc" + ] diff --git a/tooling/test_bench_autopilot_sampling.py b/tooling/test_bench_autopilot_sampling.py new file mode 100644 index 00000000..90ac3af0 --- /dev/null +++ b/tooling/test_bench_autopilot_sampling.py @@ -0,0 +1,99 @@ +#!/usr/bin/env python3 +import importlib.util +from pathlib import Path + + +SCRIPT = Path(__file__).with_name("bench_autopilot.py") +SPEC = importlib.util.spec_from_file_location("bench_autopilot", SCRIPT) +MODULE = importlib.util.module_from_spec(SPEC) +assert SPEC.loader is not None +SPEC.loader.exec_module(MODULE) + + +def test_model_card_sampling_reaches_child_environment(): + env = MODULE.benchmark_environment( + { + "benchmark_temperature": 1.0, + "benchmark_top_p": 0.95, + "benchmark_top_k": 64, + "benchmark_max_output_tokens_cap": 262144, + "serving_manifest": "/tmp/serving.json", + }, + {"PRESERVED": "yes", "BENCH_TOP_K": "stale"}, + ) + assert env["PRESERVED"] == "yes" + assert env["BENCH_TEMP"] == "1.0" + assert env["BENCH_TOP_P"] == "0.95" + assert env["BENCH_TOP_K"] == "64" + assert env["BENCH_MAX_OUTPUT_TOKENS_CAP"] == "262144" + assert env["BENCH_SERVING_MANIFEST"] == "/tmp/serving.json" + + +def test_unset_sampling_does_not_leak_parent_overrides(): + env = MODULE.benchmark_environment( + {}, + { + "PRESERVED": "yes", + "BENCH_TEMP": "9", + "BENCH_TOP_P": "0.1", + "BENCH_TOP_K": "1", + "BENCH_MAX_OUTPUT_TOKENS_CAP": "123", + "BENCH_SERVING_MANIFEST": "/tmp/stale.json", + }, + ) + assert env == {"PRESERVED": "yes"} + + +def test_configured_ports_are_ordered_and_deduplicated(): + assert MODULE.configured_ports({"port": 8000}) == [8000] + assert MODULE.configured_ports({"port": 8000, "lane_ports": [8000, "8001", 8000]}) == [8000, 8001] + + +def test_slot_progress_signature_tracks_only_live_decode_progress(): + payload = [ + { + "id": 0, + "id_task": 10, + "is_processing": False, + "n_prompt_tokens_processed": 100, + "next_token": [{"n_decoded": 20, "n_remain": 1000}], + }, + { + "id": 3, + "id_task": 99, + "is_processing": True, + "n_prompt_tokens_processed": 67, + "next_token": [{"n_decoded": 16296, "n_remain": 213539}], + }, + ] + assert MODULE.slot_progress_signature(payload) == ((3, 99, 67, 16296, 213539),) + payload[1]["next_token"][0]["n_decoded"] += 1 + assert MODULE.slot_progress_signature(payload) == ((3, 99, 67, 16297, 213539),) + + +def test_slot_progress_signature_is_empty_for_idle_or_bad_payload(): + assert MODULE.slot_progress_signature([]) == () + assert MODULE.slot_progress_signature({"not": "a slot list"}) == () + assert MODULE.slot_progress_signature([{"id": 0, "is_processing": False}]) == () + + +def test_substance_targets_include_both_live_lanes(tmp_path): + old_logs = MODULE.LOGS + old_active = MODULE.active_harnesses_by_port + try: + MODULE.LOGS = tmp_path + for cell in ("p1_refactor_model_v3", "p3_market_model_v1"): + (tmp_path / cell).mkdir() + (tmp_path / cell / "transcript.jsonl").write_text("{}\n") + MODULE.active_harnesses_by_port = lambda: { + 8000: {"pid": 101, "cell": "p1_refactor_model_v3"}, + 8001: {"pid": 202, "cell": "p3_market_model_v1"}, + } + targets = MODULE.active_substance_targets({"port": 8000, "lane_ports": [8000, 8001]}) + assert [(target["port"], target["pid"], target["cell"]) for target in targets] == [ + (8000, 101, "p1_refactor_model_v3"), + (8001, 202, "p3_market_model_v1"), + ] + finally: + MODULE.LOGS = old_logs + MODULE.active_harnesses_by_port = old_active diff --git a/tooling/test_capture_gemma4_grader_manifest.py b/tooling/test_capture_gemma4_grader_manifest.py new file mode 100644 index 00000000..83880931 --- /dev/null +++ b/tooling/test_capture_gemma4_grader_manifest.py @@ -0,0 +1,29 @@ +#!/usr/bin/env python3 +import importlib.util +from pathlib import Path + + +SCRIPT = Path(__file__).with_name("capture_gemma4_grader_manifest.py") +SPEC = importlib.util.spec_from_file_location("grader_manifest", SCRIPT) +MODULE = importlib.util.module_from_spec(SPEC) +assert SPEC.loader is not None +SPEC.loader.exec_module(MODULE) + + +def test_fingerprint_files_hashes_and_fails_closed(tmp_path): + present = tmp_path / "present.py" + present.write_bytes(b"print('fixed grader')\n") + files, errors = MODULE.fingerprint_files(tmp_path, ["present.py", "missing.json"]) + assert files["present.py"]["bytes"] == len(present.read_bytes()) + assert files["present.py"]["sha256"] == MODULE.sha256(present) + assert errors == ["missing grader input: missing.json"] + + +def test_manifest_covers_all_canonical_task_graders(): + joined = "\n".join(MODULE.GRADER_FILES) + for task in MODULE.TASKS: + family = task.split("_", 1)[0] + assert family in {"p1", "p2", "p3"} + assert "phase1_grade.py" in joined + assert "phase2_hallucination_grade.py" in joined + assert "phase3_project_mgmt_grade.py" in joined diff --git a/tooling/test_check_substance.py b/tooling/test_check_substance.py new file mode 100644 index 00000000..756b90a1 --- /dev/null +++ b/tooling/test_check_substance.py @@ -0,0 +1,25 @@ +#!/usr/bin/env python3 +import importlib.util +from pathlib import Path + + +SCRIPT = Path(__file__).parent / "scripts" / "check_substance.py" +SPEC = importlib.util.spec_from_file_location("check_substance", SCRIPT) +MODULE = importlib.util.module_from_spec(SPEC) +assert SPEC.loader is not None +SPEC.loader.exec_module(MODULE) + + +def test_command_template_digit_strips_valid_commands(): + entry = {"args": {"command": "sed -n '100,160p' report.md"}} + assert MODULE.command_template(entry) == "sed -n '#,#p' report.md" + + +def test_command_template_preserves_malformed_boolean_without_crashing(): + assert MODULE.command_template({"args": {"command": True}}) == \ + "" + + +def test_command_template_handles_non_mapping_args_without_crashing(): + assert MODULE.command_template({"args": ["unexpected"]}) == \ + "" diff --git a/tooling/test_correct_gemma4_project_mgmt_grades.py b/tooling/test_correct_gemma4_project_mgmt_grades.py new file mode 100644 index 00000000..3d0bc3ac --- /dev/null +++ b/tooling/test_correct_gemma4_project_mgmt_grades.py @@ -0,0 +1,92 @@ +#!/usr/bin/env python3 +import importlib.util +import io +import tarfile +from pathlib import Path + + +SCRIPT = Path(__file__).with_name("correct_gemma4_project_mgmt_grades.py") +SPEC = importlib.util.spec_from_file_location("pm_correction", SCRIPT) +MODULE = importlib.util.module_from_spec(SPEC) +assert SPEC.loader is not None +SPEC.loader.exec_module(MODULE) + + +def raw_grade(): + return { + "verdict": "FAIL", + "scores": { + "sections_present": ["workstream", "risk", "decision", "milestone"], + "word_count": 300, + }, + "thresholds": { + "min_workstreams": 4, + "min_risks": 3, + "min_decisions": 3, + "min_milestones": 3, + "max_word_count": 700, + }, + "details": { + "workstreams": {f"WS{i}": {"matched": True} for i in range(1, 7)}, + "risks": { + "R1": {"matched": False}, "R2": {"matched": False}, + "R3": {"matched": False}, "R4": {"matched": True}, + "R5": {"matched": True}, "R6": {"matched": False}, + }, + "decisions": { + "D1_branding": {"matched": True}, + "D2_panel_limit": {"matched": True}, + "D3_mobile": {"matched": False}, + "D4_option_b": {"matched": False}, + }, + "milestones": {f"M{i}": {"matched": True} for i in range(1, 6)}, + }, + } + + +def test_semantic_equivalents_correct_only_the_known_lexical_misses(): + report = """ + Risks: Maevia was promised GA and may push back on private beta. + Legal has not yet responded to the private-beta contract draft. + Decisions: web-responsive in V1; native in V2. + SDK release changed to private beta (3-5 customers). + """ + corrected, changes = MODULE.apply_correction(raw_grade(), report) + assert corrected["verdict"] == "PASS" + assert corrected["scores"]["risk_recall"] == "4/6" + assert corrected["scores"]["decision_recall"] == "4/4" + assert {change["item"] for change in changes} == { + "R2", "R3", "D3_mobile", "D4_option_b", + } + + +def test_unrelated_text_does_not_change_a_failure(): + corrected, changes = MODULE.apply_correction(raw_grade(), "No additional evidence.") + assert corrected["verdict"] == "FAIL" + assert changes == [] + + +def test_archived_report_is_authoritative_when_live_workspace_is_gone(tmp_path): + report = b"Risks: Maevia may push back on private beta.\n" + archive = tmp_path / "workspace_final.tar.gz" + with tarfile.open(archive, "w:gz") as bundle: + member = tarfile.TarInfo("./status_report.md") + member.size = len(report) + bundle.addfile(member, io.BytesIO(report)) + + text, digest = MODULE.read_archived_report(archive) + assert text == report.decode("utf-8") + assert digest == MODULE.sha256_bytes(report) + + +def test_archived_report_rejects_missing_member(tmp_path): + archive = tmp_path / "workspace_final.tar.gz" + with tarfile.open(archive, "w:gz"): + pass + + try: + MODULE.read_archived_report(archive) + except ValueError as exc: + assert "expected one ./status_report.md" in str(exc) + else: + raise AssertionError("missing status_report.md should fail closed") diff --git a/tooling/test_gemma4_extended_runner.py b/tooling/test_gemma4_extended_runner.py new file mode 100644 index 00000000..84c17591 --- /dev/null +++ b/tooling/test_gemma4_extended_runner.py @@ -0,0 +1,38 @@ +#!/usr/bin/env python3 +import importlib.util +import os +from pathlib import Path + + +os.environ["GEMMA_EXTENDED_LANE_INDEX"] = "0" +os.environ["GEMMA_EXTENDED_LANE_COUNT"] = "2" +os.environ["GEMMA_EXTENDED_PORT"] = "8000" +SCRIPT = Path(__file__).with_name("run_gemma4_extended_suites.py") +SPEC = importlib.util.spec_from_file_location("runner", SCRIPT) +RUNNER = importlib.util.module_from_spec(SPEC) +assert SPEC.loader is not None +SPEC.loader.exec_module(RUNNER) + + +def test_names_are_campaign_specific_and_replicated(): + assert RUNNER.run_name("dreamserver-1-pr-audit", 2) == "n1_gemma4-31b-q4_v2" + assert RUNNER.run_name("wallstreet-board-presentation", 3) == "gemma4-31b-q4_board_pres_v3" + + +def test_harness_command_pins_lane_and_full_operating_point(): + suite = { + "id": "dreamserver-1-pr-audit", + "task": "tooling/tasks/task_pr_audit_n1.md", + "temperature": 1.0, + "stuck_threshold": 500, + "max_output_tokens_cap": 262144, + "require_git_tag": True, + } + command = RUNNER.harness_command(suite, 1, 262144, 0.95, 64) + joined = " ".join(command) + assert "--port 8000" in joined + assert "--max-model-len 262144" in joined + assert "--max-output-tokens-cap 262144" in joined + assert "--temperature 1.0 --top-p 0.95 --top-k 64" in joined + assert "--serving-manifest" in command + assert "--require-git-tag" in command diff --git a/tooling/test_gemma4_server_launcher.py b/tooling/test_gemma4_server_launcher.py new file mode 100644 index 00000000..5693b5d8 --- /dev/null +++ b/tooling/test_gemma4_server_launcher.py @@ -0,0 +1,17 @@ +from pathlib import Path + + +ROOT = Path(__file__).resolve().parent + + +def test_native_envelope_transport_timeout_is_enforced(): + candidates = [ + ROOT / "deployments" / "gemma4-31b-q4-tower2" / "run-gemma4-server.sh", + ROOT / "run-gemma4-server.sh", + ] + launcher_path = next(path for path in candidates if path.is_file()) + launcher = launcher_path.read_text() + assert 'HTTP_TIMEOUT="${GEMMA_HTTP_TIMEOUT:-14400}"' in launcher + assert "HTTP_TIMEOUT < 14400" in launcher + assert '--timeout "$HTTP_TIMEOUT"' in launcher + assert "--timeout 3600" not in launcher diff --git a/tooling/test_gemma4_supplemental_telemetry.py b/tooling/test_gemma4_supplemental_telemetry.py new file mode 100644 index 00000000..ec3952ed --- /dev/null +++ b/tooling/test_gemma4_supplemental_telemetry.py @@ -0,0 +1,31 @@ +#!/usr/bin/env python3 +from pathlib import Path + + +SCRIPT = Path(__file__).with_name("run_gemma4_supplemental_telemetry.sh").read_text() + + +def test_supplements_use_distinct_noncanonical_label_and_exact_shards(): + assert 'LABEL="gemma4-31b-q4-telemetry-supplement"' in SCRIPT + assert "BENCH_LANE_COUNT=24" in SCRIPT + assert "BENCH_LANE_INDEX=0" in SCRIPT + assert "BENCH_LANE_INDEX=1" in SCRIPT + assert '"$LABEL" 2' in SCRIPT + + +def test_supplements_fail_closed_on_campaign_overlap_and_power_drift(): + assert "mmbt-gemma4-canonical-n3-r3.service" in SCRIPT + assert "another benchmark harness is active" in SCRIPT + assert '"500.00 500.00"' in SCRIPT + assert "will not be overwritten" in SCRIPT + + +def test_supplements_require_telemetry_before_success(): + assert "gpu_telemetry.json" in SCRIPT + assert "SUPPLEMENTAL_TELEMETRY_COMPLETE" in SCRIPT + + +def test_supplements_require_idle_replica_slots_at_launch(): + assert 'http://127.0.0.1:${port}/slots' in SCRIPT + assert "all(.[]; .is_processing == false)" in SCRIPT + assert "retry at an idle boundary" in SCRIPT diff --git a/tooling/test_harness_tool_validation.py b/tooling/test_harness_tool_validation.py new file mode 100644 index 00000000..6a081f8f --- /dev/null +++ b/tooling/test_harness_tool_validation.py @@ -0,0 +1,65 @@ +#!/usr/bin/env python3 +import importlib.util +from pathlib import Path + + +SCRIPT = Path(__file__).with_name("harness.py") +SPEC = importlib.util.spec_from_file_location("harness", SCRIPT) +MODULE = importlib.util.module_from_spec(SPEC) +assert SPEC.loader is not None +SPEC.loader.exec_module(MODULE) + + +def execute(name, args): + return MODULE.execute_tool(name, args, Path("/tmp/not-used")) + + +def test_missing_required_arguments_are_recoverable_tool_errors(): + assert execute("read_file", {}) == "TOOL_ERROR: read_file missing required argument(s): path" + assert execute("write_file", {"path": "x"}) == ( + "TOOL_ERROR: write_file missing required argument(s): content" + ) + assert execute("bash", {}) == "TOOL_ERROR: bash missing required argument(s): command" + + +def test_malformed_argument_types_are_recoverable_tool_errors(): + assert execute("read_file", {"path": None}) == ( + "TOOL_ERROR: read_file argument 'path' must be a non-empty string" + ) + assert execute("write_file", {"path": "x", "content": 1}) == ( + "TOOL_ERROR: write_file argument 'content' must be a string" + ) + assert execute("bash", {"command": "true", "timeout_s": "later"}) == ( + "TOOL_ERROR: bash argument 'timeout_s' must be a positive integer" + ) + + +def test_non_object_arguments_are_recoverable_tool_errors(): + assert execute("read_file", []) == "TOOL_ERROR: read_file arguments must be a JSON object" + + +def test_git_tag_gate_requires_clean_annotated_tag_at_head(monkeypatch): + calls = [] + + def fake_docker_exec(command, timeout): + calls.append(command) + return {"stdout": "TAG_FOUND\n", "stderr": "", "returncode": 0} + + monkeypatch.setattr(MODULE, "docker_exec", fake_docker_exec) + assert MODULE.validate_done([], True) is None + command = calls[-1] + assert "diff --quiet" in command + assert "diff --cached --quiet" in command + assert "ls-files --others --exclude-standard" in command + assert "tag --points-at HEAD" in command + assert "cat-file -t" in command + + +def test_git_tag_gate_fails_closed_without_qualifying_repo(monkeypatch): + monkeypatch.setattr( + MODULE, + "docker_exec", + lambda command, timeout: {"stdout": "", "stderr": "", "returncode": 0}, + ) + error = MODULE.validate_done([], True) + assert "no clean workspace repo with an annotated tag at HEAD" in error diff --git a/tooling/test_snapshot_gemma4_telemetry.py b/tooling/test_snapshot_gemma4_telemetry.py new file mode 100644 index 00000000..21cb11a1 --- /dev/null +++ b/tooling/test_snapshot_gemma4_telemetry.py @@ -0,0 +1,19 @@ +#!/usr/bin/env python3 +from pathlib import Path + + +SCRIPT = Path(__file__).with_name("snapshot_gemma4_telemetry.sh").read_text() + + +def test_snapshot_requires_a_clean_boundary_and_refuses_overwrite(): + assert "benchmark work is active" in SCRIPT + assert "refusing to overwrite telemetry snapshot" in SCRIPT + assert "snapshot destination must be an absolute CSV" in SCRIPT + + +def test_snapshot_pauses_writer_validates_csv_and_restores_services(): + assert 'systemctl --user stop "$SIDECAR" "$LOGGER"' in SCRIPT + assert "telemetry snapshot does not end at a complete line" in SCRIPT + assert "telemetry snapshot header drift" in SCRIPT + assert "trap restart_services EXIT" in SCRIPT + assert "sha256sum" in SCRIPT diff --git a/tooling/test_summarize_gemma4_campaign.py b/tooling/test_summarize_gemma4_campaign.py new file mode 100644 index 00000000..489eaf7f --- /dev/null +++ b/tooling/test_summarize_gemma4_campaign.py @@ -0,0 +1,74 @@ +#!/usr/bin/env python3 +import importlib.util +from pathlib import Path + + +SCRIPT = Path(__file__).with_name("summarize_gemma4_campaign.py") +SPEC = importlib.util.spec_from_file_location("summary", SCRIPT) +MODULE = importlib.util.module_from_spec(SPEC) +assert SPEC.loader is not None +SPEC.loader.exec_module(MODULE) + + +def record(name, task, verdict, finish, wall, tokens, tps, telemetry=None): + return { + "run_name": name, + "task": task, + "verdict": verdict, + "summary": {"finish_reason": finish, "elapsed_s": wall, "total_completion_tokens": tokens}, + "cost": {"throughput": {"completion_tps_avg": tps}}, + "telemetry": telemetry, + } + + +def test_aggregate_keeps_finish_and_quality_axes_separate(): + rows = [ + record("a", "p1_bugfix", "PASS", "done_signal", 10, 100, 20, + {"coverage": 0.9, "mean_power_w": 400, "mean_sm_util_pct": 90, "max_temp_c": 70}), + record("b", "p1_bugfix", "FAIL", "done_signal", 20, 200, 30), + record("c", "p2_extract", "STRUCTURAL_PASS", "model_stopped", 30, 300, 40), + ] + result = MODULE.aggregate_records(rows) + assert result["raw_passes"] == 2 + assert result["finish_reasons"] == {"done_signal": 2, "model_stopped": 1} + assert result["model_call_completion_tps"]["median"] == 30 + assert result["per_task"]["p1_bugfix"]["raw_passes"] == 1 + assert result["telemetry"]["runs"] == 1 + + +def test_markdown_explicitly_rejects_done_equals_pass(): + document = {"target_n": 1, "aggregate": MODULE.aggregate_records([])} + rendered = MODULE.markdown(document) + assert "A `done_signal` is a finish behavior, not a pass" in rendered + assert "Raw grader verdicts only" in rendered + + +def test_terminal_label_is_scored_as_distinct_nonpass_without_fake_grade(): + terminal = { + "run_name": "p3_market_gemma4-31b-q4_v1", + "task": "p3_market", + "verdict": None, + "summary": None, + "terminal_label": "identical-call-loop", + "cost": { + "wall_s": 45, + "tokens": {"completion_total": 1234}, + "throughput": {"completion_tps_avg": 27.4}, + }, + "telemetry": { + "coverage": 0.92, "mean_power_w": 450, + "mean_sm_util_pct": 95, "max_temp_c": 76, + }, + } + result = MODULE.aggregate_records([terminal]) + assert result["completed"] == 1 + assert result["normal_completed"] == 0 + assert result["terminal_outcomes"] == 1 + assert result["graded"] == 0 + assert result["scored_outcomes"] == 1 + assert result["raw_passes"] == 0 + assert result["raw_pass_rate"] == 0.0 + assert result["finish_reasons"] == {"terminal:identical-call-loop": 1} + assert result["quality_outcomes"] == {"TERMINAL:identical-call-loop": 1} + assert result["wall_s"]["sum"] == 45 + assert result["completion_tokens"]["sum"] == 1234 diff --git a/tooling/test_validate_gemma4_comparison_sources.py b/tooling/test_validate_gemma4_comparison_sources.py new file mode 100644 index 00000000..9a6cf9a6 --- /dev/null +++ b/tooling/test_validate_gemma4_comparison_sources.py @@ -0,0 +1,67 @@ +#!/usr/bin/env python3 +import importlib.util +import hashlib +import json +from pathlib import Path +from types import SimpleNamespace + + +SCRIPT = Path(__file__).with_name("validate_gemma4_comparison_sources.py") +SPEC = importlib.util.spec_from_file_location("comparison_sources", SCRIPT) +MODULE = importlib.util.module_from_spec(SPEC) +assert SPEC.loader is not None +SPEC.loader.exec_module(MODULE) + + +def test_expected_comparator_set_is_explicit_and_complete(): + assert MODULE.EXPECTED_IDS == { + "qwen3.6-27b-awq", + "qwen3-coder-next-awq", + "qwen3.6-35b-a3b-awq", + "qwen3.5-397b-a17b-q3-nothink", + "deepseek-v4-flash-0731", + } + + +def test_manifest_keeps_raw_corrected_and_missing_cohort_distinct(): + manifest = json.loads( + Path(__file__).with_name("gemma4-comparison-sources.json").read_text() + ) + rows = {row["id"]: row for row in manifest["comparators"]} + assert rows["deepseek-v4-flash-0731"]["canonical"] == { + "cohort": "12 families x N=3", + "raw_passes": 23, + "corrected_passes": 35, + "total": 36, + } + assert rows["qwen3.5-397b-a17b-q3-nothink"]["canonical"]["total"] == 120 + assert rows["qwen3.6-35b-a3b-awq"]["canonical"] is None + assert any("not a global SOTA claim" in rule for rule in manifest["comparison_rules"]) + + +def test_mutable_scorecard_is_verified_at_the_pinned_source_commit(): + manifest = json.loads( + Path(__file__).with_name("gemma4-comparison-sources.json").read_text() + ) + assert manifest["source_repository_commit"] == ( + "dc8f57d443aeccba652a33852fda32d37cd701f3" + ) + scorecard = next( + row for row in manifest["source_documents"] if row["path"] == "SCORECARD.md" + ) + assert scorecard["verify_at_source_commit"] is True + + +def test_git_blob_hash_uses_exact_git_show_bytes(monkeypatch, tmp_path): + payload = b"pre-campaign scorecard\n" + + def fake_run(command, capture_output, check): + assert command[-1] == "abc123:SCORECARD.md" + assert capture_output is True + assert check is False + return SimpleNamespace(returncode=0, stdout=payload) + + monkeypatch.setattr(MODULE.subprocess, "run", fake_run) + assert MODULE.git_blob_sha256(tmp_path, "abc123", "SCORECARD.md") == hashlib.sha256( + payload + ).hexdigest() diff --git a/tooling/validate_gemma4_comparison_sources.py b/tooling/validate_gemma4_comparison_sources.py new file mode 100755 index 00000000..638c593b --- /dev/null +++ b/tooling/validate_gemma4_comparison_sources.py @@ -0,0 +1,154 @@ +#!/usr/bin/env python3 +"""Fail closed if a pinned Gemma comparison source or extracted claim drifts.""" +from __future__ import annotations + +import argparse +import hashlib +import json +import subprocess +from pathlib import Path + + +EXPECTED_IDS = { + "qwen3.6-27b-awq", + "qwen3-coder-next-awq", + "qwen3.6-35b-a3b-awq", + "qwen3.5-397b-a17b-q3-nothink", + "deepseek-v4-flash-0731", +} +RECEIPTS = { + "qwen3.6-27b-awq": "benchmarks/microbench-2026-04-28/bug-fixing/Qwen3.6-27B-AWQ/receipt.json", + "qwen3-coder-next-awq": "benchmarks/microbench-2026-04-28/bug-fixing/Qwen3-Coder-Next-AWQ/receipt.json", + "qwen3.6-35b-a3b-awq": "benchmarks/dreamserver-1-pr-audit/Qwen3.6-35B-A3B-AWQ/receipt.json", +} + + +def sha256(path: Path) -> str: + digest = hashlib.sha256() + with path.open("rb") as handle: + for chunk in iter(lambda: handle.read(1024 * 1024), b""): + digest.update(chunk) + return digest.hexdigest() + + +def read_json(path: Path) -> dict: + return json.loads(path.read_text()) + + +def git_blob_sha256(root: Path, commit: str, path: str) -> str | None: + result = subprocess.run( + ["git", "-C", str(root), "show", f"{commit}:{path}"], + capture_output=True, + check=False, + ) + if result.returncode != 0: + return None + return hashlib.sha256(result.stdout).hexdigest() + + +def validate(root: Path, manifest_path: Path) -> list[str]: + errors = [] + manifest = read_json(manifest_path) + source_commit = manifest.get("source_repository_commit") + for source in manifest.get("source_documents", []): + path = root / source["path"] + if source.get("verify_at_source_commit"): + observed = git_blob_sha256(root, source_commit, source["path"]) + if observed is None: + errors.append(f"missing source document at pinned commit: {source['path']}") + elif observed != source["sha256"]: + errors.append(f"source hash drift at pinned commit: {source['path']}") + elif not path.is_file(): + errors.append(f"missing source document: {source['path']}") + elif sha256(path) != source["sha256"]: + errors.append(f"source hash drift: {source['path']}") + + ancestor = subprocess.run( + ["git", "-C", str(root), "merge-base", "--is-ancestor", source_commit, "HEAD"], + capture_output=True, + check=False, + ) + if ancestor.returncode != 0: + errors.append("pinned comparison source commit is not an ancestor of HEAD") + + comparators = {row.get("id"): row for row in manifest.get("comparators", [])} + if set(comparators) != EXPECTED_IDS: + errors.append( + f"comparator ids differ: observed={sorted(comparators)} expected={sorted(EXPECTED_IDS)}" + ) + for model_id, row in comparators.items(): + canonical = row.get("canonical") + if canonical is None: + continue + total = canonical.get("total") + for field in ("raw_passes", "corrected_passes"): + value = canonical.get(field) + if value is not None and ( + not isinstance(value, int) or not isinstance(total, int) or not 0 <= value <= total + ): + errors.append(f"invalid {model_id} {field}/{total}") + + deepseek = read_json(root / "benchmarks/deepseek-v4-flash-0731/canonical-regrade-audit.json") + pinned = (comparators.get("deepseek-v4-flash-0731") or {}).get("canonical") or {} + for field, source_field in ( + ("raw_passes", "raw_reported_passes"), + ("corrected_passes", "corrected_passes"), + ("total", "total_cells"), + ): + if pinned.get(field) != deepseek.get(source_field): + errors.append(f"DeepSeek extracted {field} differs from its audit") + + qwen397 = read_json( + root / "benchmarks/deepseek-v4-flash-0731/qwen397-corrected-score-overlay.json" + ) + pinned = (comparators.get("qwen3.5-397b-a17b-q3-nothink") or {}).get("canonical") or {} + qwen_raw = (qwen397.get("raw_published") or {}).get("nothink") or {} + qwen_corrected = (qwen397.get("corrected_overlay") or {}).get("nothink") or {} + if pinned.get("raw_passes") != qwen_raw.get("pass"): + errors.append("Qwen3.5-397B extracted raw_passes differs from its overlay") + if pinned.get("corrected_passes") != qwen_corrected.get("pass"): + errors.append("Qwen3.5-397B extracted corrected_passes differs from its overlay") + if pinned.get("total") != qwen_corrected.get("total"): + errors.append("Qwen3.5-397B extracted total differs from its overlay") + + for model_id, receipt_rel in RECEIPTS.items(): + row = comparators.get(model_id) or {} + op = row.get("operating_point") or {} + receipt = read_json(root / receipt_rel) + defaults = receipt.get("inference_request_defaults") or {} + if defaults.get("temperature") != op.get("temperature"): + errors.append(f"{model_id} temperature differs from pinned receipt") + if defaults.get("max_model_len") != op.get("context_tokens"): + errors.append(f"{model_id} context differs from pinned receipt") + ceiling = op.get("historical_per_response_ceiling") + if ceiling is not None and str(ceiling) not in str(defaults.get("max_tokens_strategy")): + errors.append(f"{model_id} output ceiling differs from pinned receipt") + + scorecard = (root / "SCORECARD.md").read_text() + for required in ( + "Qwen3.6-27B-AWQ", + "Qwen3-Coder-Next-AWQ", + "Qwen3.6-35B-A3B-AWQ", + "DeepSeek V4 Flash 0731", + "397B-A17B", + ): + if required not in scorecard: + errors.append(f"scorecard no longer identifies comparator {required}") + return errors + + +def main() -> None: + parser = argparse.ArgumentParser() + parser.add_argument("--root", type=Path, default=Path(__file__).resolve().parents[1]) + parser.add_argument( + "--manifest", type=Path, + default=Path(__file__).with_name("gemma4-comparison-sources.json"), + ) + args = parser.parse_args() + errors = validate(args.root.resolve(), args.manifest.resolve()) + print(json.dumps({"passed": not errors, "errors": errors}, indent=2)) + raise SystemExit(0 if not errors else 1) + + +if __name__ == "__main__": + main()