Skip to content

docs(run): §7 report — clean-room root-caused, 11 PRs unstranded - #3310

Merged
noahgift merged 4 commits into
mainfrom
PMAT-3305-run-report
Sep 15, 2026
Merged

noahgift merged 4 commits into
mainfrom
PMAT-3305-run-report

Conversation

@noahgift

@noahgift noahgift commented Sep 15, 2026

Copy link
Copy Markdown
Contributor

APR-RELEASE-001 §7 interval report for the 2026-09-15 autonomous run.

Two mechanisms this interval, both of the declared fix that cannot fire class:

Also records what is not claimed: the A1 half of #3189 — "the clean-room gate cannot pass on a release commit at all" — stays unreproduced, with the command that would settle it named. #3305 explains all 11 observed errors and none is release-shaped.

Docs-only; no code, no roadmap change.

ont-delta: none this interval produced no ontology row — the two findings are merge-path and publish-manifest mechanisms, both already tracked as #3305/#3306/#3308

🤖 Generated with Claude Code

no-close: this is the §7 interval report. It cites #3189, #3305, #3306 and #3308 as the findings it records; each is closed by its own PR, and a docs report closes nothing.

APR-RELEASE-001 §7. The interval's two mechanisms, both of the
declared-fix-that-cannot-fire class:

  * cargo publish strips path-only dev-deps, and aprender-compute's src/ named
    three of them -- 8 consecutive red clean-room runs, two releases shipped
    over the hard gate.
  * a .gitattributes merge driver is read from the side being merged INTO, so
    the union declaration from #3256 never fired on any older branch.

Records what is NOT claimed: the A1 half of #3189 stays unreproduced, and the
command that would settle it is named.

Pmat-Ticket: PMAT-3305

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@noahgift
noahgift enabled auto-merge September 15, 2026 13:06
noahgift and others added 2 commits September 15, 2026 15:27
PR-1's premise held and widened: 355 tests dark, and the make target that would
have run them errored, which is why nobody noticed. PR-2 named 14 anonymous
obligations and repointed 4 proofs that referenced nothing -- then the scan
showed 824 of 876 contracts carry the same defect, so #3091's "0/17 bound" was
never a Qwen finding.

Records the judgement call on KANI-QHF-001 as a decision request rather than
burying it, and the staging-contracts hazard that was deliberately not actioned.

Pmat-Ticket: PMAT-3305

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…gations named

Leads with the scope line: GDN 0/5 discharged (5 obligations, 5 tests, all
ignored with empty bodies), #3303 not started, andon clock stated.

Records that pv proof-status reports L4 for a contract whose every test is an
empty stub, and that two of my own changes were caught by the repo's guards
rather than by me.

Pmat-Ticket: PMAT-3305

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@github-actions

github-actions Bot commented Sep 15, 2026

Copy link
Copy Markdown

§13.11 rung 1 — quorum shadow verdict

S13-SHADOW pr=3310 head=9ecec9084fe28a4586ecbb8a72e427172a6032cc verdict=REFUSE class=Q1 arm_rc=1

Shadow mode: this records a verdict and merges nothing. A refusal
to arm is not a block (§13 adds zero rows to §7) — the pull request is
exactly as green as it was.

…3303 blocker named

The scope line now carries the #3303 blocker with its number, as ruled.
Records the two P0·Instrument rules (sampling window >= p95 gate; BEHIND=0
without group batching pays the burst all at once), and the census gate
catching my own naming widening duplicate-stem drift.

Pmat-Ticket: PMAT-3305

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@noahgift
noahgift added this pull request to the merge queue Sep 15, 2026
@noahgift

Copy link
Copy Markdown
Contributor Author

§7 interval — 18:04Z (posted as a comment: this PR is in the merge queue, which locks its branch)

scope: GDN 5/5 discharged, mutation-verified (#3322, queue pos 2)
       QHF 6/7 discharged, 15 mutations all red (local branch PMAT-3091-qhf-bodies, off #3114)
           QHF-BND-005 NOT discharged: contract defect, formal "for finite M" is vacuous
       E2E 0/7 — worker triaging now
       #3303: reference EXISTS, raw float32 logits, 78 positions, n=5 byte-identical,
              falsified against llama-perplexity (0 of 9,436,312 words differ)
       #3091: PARITY MEASURED — apr #3114 vs llama.cpp d1d3c3396, same host:
              min cosine 0.995905 @pos4, mean 0.998989, 0 of 78 below 0.98,
              4 argmax mismatches, all rank-2 on near-ties (ref top-2 gap as low as 0.013)
              re-run independently by the orchestrator: matches
              remaining to merge: E2E triage, QHF-BND-005 amendment, contracts pushed into
              #3114, #3114 onto main, CI green, then ready-for-review
       andon: 66h14m to §8 at 2026-09-18T12:19:39Z
queue: depth 8, UNCHANGED since 16:49Z; only #3309 merged since 11:09Z; head under diagnosis
pins:  llama.cpp d1d3c3396 HELD. `-no-cnv` is registered only for LLAMA_EXAMPLE_COMPLETION, so the
       parity test must call llama-completion, not llama-cli. But llama-completion with the test's
       exact flags exits rc=0 with 0 stdout bytes; diagnosing before the test switches binaries

§8 decisions, logged not asked

  1. Pin bump by measurement. 39173bcac (2026-01-15) predates qwen3.5 support (fc0fe4004, 2026-02-10). Target d1d3c3396 loads general.architecture = qwen35 on intel CPU, rc=0.
  2. Raw-logit producer, not a longer corpus. llama-perplexity keeps 38 of 78 positions as uint16 log-softmax. Lengthening the corpus would move the 0.66 GPU basis, and the ruling says a moved baseline is a ticket, never a threshold edit.
  3. QHF-BND-005 gets amended, not faked. ∀l: ||h_l||_inf <= M for finite M only says "finite". Amend the formal to explicit finiteness, which is what FALSIFY-QHF-005 already tests. Add a quantitative bound only after measuring the growth ratio on the reference.
  4. The shared-registry page after apr run panics on GX10 (GB10): CUDA succeeds but wgpu fallback triggers anyway #612 is correct. The reclaim needs per-job CARGO_HOME plus a drain window, not a narrower live check. The infra#509 crate was resolved 10 minutes before it was read, and nothing held the path. Filed in paiml/infra.

infra#612 — implemented, verifying before PR

Split by tree shape, budget cap 60 GB, 24 consecutive skips → exit 3, ledger at /var/lib/ci-reaper/ledger.jsonl. Harness 45 ok / 0 FAIL, 9 mutations red, bashrs 0 errors (5 pre-existing SEC011 fixed), forjar validate rc 0 on gx10 and yoga. gh is absent on gx10, so attribution uses live paths only. Second, pre-existing defect: the old per-PR rule matched <pr>/*/src (depth 3), but the layout is <pr>/src (depth 2). It would have matched nothing even with no live job.

Coordination

session holds agreement
infra-0b infra#607 (gx10/yoga forjar fold, STOPPED) we hold infra#612. It edits only ci-reaper-* / ci-disk-watch-*. No forjar apply on gx10 with the other's unmerged branch, and we message before applying.
infra-66 PMAT-615 / infra#615 (scripts/pvl/lib.sh) we stay off that file. We do not own aprender#3281 (we only merged main into it).
paiml-implement-a5 PMAT-207 (paiml-implement repo only) nothing in aprender

Filed this interval: aprender#3329 (pin bump), aprender#3330 (row-0a instrument refuses every c=1 lane at both pins).

@noahgift

Copy link
Copy Markdown
Contributor Author

§7 interval — 18:30Z

scope: GDN 5/5 — MERGED to main (#3322, 17:49Z)
       QHF 6/7 — in #3114 (pushed, CI running); QHF-BND-005 formal amended to finiteness, body in progress
       E2E 3/7 — in #3114; ORD-003 amended (n%256 precondition); INV-001 equation fixed (+untied lm_head);
                 INV-001 / BND-002 / BND-005 / MON-004 need N/A — schema has NO N/A form; adding AppliesTo::NotApplicable
       #3303: reference built + falsified; structure record MERGED (#3323, 17:49Z)
       #3091: parity min cosine 0.995905 / 78 positions, 0 below 0.98; #3114 now carries contracts + tests + main
              remaining: #3114 CI green, BND-005 body, N/A schema, apply N/A, un-draft
       andon: 65h49m to §8 at 2026-09-18T12:19:39Z
queue: #3322 + #3323 merged 17:49Z; head advancing; #3295 had never been armed (explicit --squash is refused under the merge queue)
pins:  llama.cpp #3331 open + armed

§8 decisions, logged

  1. The pin PR was never blocked by the pin. Given the parity test's flags, the old pin's llama-cli looped printing "please use llama-completion instead" until timeout (rc=124, 532 MB of stdout). -no-cnv has only ever been registered for llama-completion. The test is broken at every pin → separate ticket, with the verified fix: llama-completion without --log-disable, which at d1d3c3396 empties stdout.
  2. Contract N/A needs a schema variant, not a string. AppliesTo has #[serde(other)], so applies_to: N/A would parse as an algorithm target named "N/A". Every N/A precedent in contracts/ is Lean-only. Adding AppliesTo::NotApplicable with a required typed reason and owner, counted as K N/A and never as passed.
  3. feat(qwen35): Qwen3.5 / Qwen3.8 hybrid GGUFs run on the CPU — apr run + apr chat (#3091) #3114 leaves draft once its CI is green and QHF-BND-005 has a body. The draft ruling's stated reason (zero contracts) no longer holds.

Shipped / opened this interval

#3331 llama.cpp pin bump, open + armed; -no-cnv record corrected
infra#620 ci-reaper per-PR registry fix, open + armed; harness 45/0 re-verified by the orchestrator
infra#619 shared-registry reclaim needs per-job CARGO_HOME + a drain window
#3114 contract tests (QHF 6, E2E 3) + main, pushed
#3297 README CONTRACT_COUNT regenerated (it adds a contract and never re-ran make readme-sync)

Merged via the queue into main with commit 9ab550d Sep 15, 2026
18 of 19 checks passed
@noahgift
noahgift deleted the PMAT-3305-run-report branch September 15, 2026 18:47
@noahgift

Copy link
Copy Markdown
Contributor Author

§7 interval — 19:18Z

scope: GDN 5/5 — on main (#3322)
       QHF 7/7 — on #3114 (QHF-BND-005 body pushed: 48 layers, 9 adversarial-but-finite inputs,
                 2 eps mutations red; the softplus clamp proven unreachable, arg capped at 3.31 by RMSNorm)
       E2E 3/7 — on #3114; the other 4 are declared-N/A candidates, waiting on the schema form
       contract amendments #3333 (QHF-BND-005 finiteness, QE2E-ORD-003 n%256 precondition,
                 QE2E-INV-001 untied lm_head term) — open, armed
       #3303: reference built + falsified (0/9,436,312 words differ); structure record on main (#3323)
       #3091: parity min cosine 0.995905 / 78 positions, 0 below 0.98, 4 argmax flips all rank-2 on near-ties
              remaining: #3114 CI green → onto main → ready; N/A schema (#3327 → schema-na) → apply N/A
       andon: 65h01m to §8 at 2026-09-18T12:19:39Z
queue: #3307 MERGED (the clean-room fix is on main) · #3316, #3319, #3326, #3312 merged; infra#620 MERGED
pins:  llama.cpp #3331 open (guard-cargo red, diagnosing)

§8 decisions, logged

  1. ci.yml:570 → fragments, not a file. The worker's merge proof: a single file with two EOF appends still conflicts; only inserts at different positions merge. The ruling's premise was "same lock-elimination as roadmap fragments", so one shared file only moves the lock. The target is ci/explicit-test-commands.d/NNN-<slug>.cmd, one command per file, gapped ordinals, read in sort order, with a guard refusing duplicate commands, shared ordinals and empty dirs. merge=union is refused: GitHub server-side honouring is unverified, it is inert for older branches, and it keeps both sides of a same-line edit silently.
  2. The N/A schema collision → migrate, keep the rule strict. silu-kernel-v1 and tokenizer-v1 already carried an untyped na_reason (meaning "no Lean proof") that serde dropped. Those 5 moved into the typed lean.notes, which explain_render.rs:51 actually reads. pv proof-status is byte-identical before and after. SCHEMA-021..023 stay Errors; a warning is the decoration class.
  3. cascade-publish must refuse to start unless clean-room is green on the tag's sha — two releases shipped over a red hard gate #3318 lands, and it adds a step to the train. clean-room clones aprender main at 23:00 UTC, so it almost never tests a tag commit (retro: none of the 7 runs after v0.66.0 or the 3 after v0.67.0 tested the tag). Until infra takes a ref input and records the full sha in structured output (filed in paiml/infra), every release must dispatch clean-room while main == tag, or the cascade refuses. Correct fail-closed behaviour; it must be in place before 0.68's T-4.

#3041 — split plan (#3334)

96 files: 42 already identical on main (#3004 registry, #3026 parity), 20 differ only because R-0b is behind, and 1 is a contract main deleted on purpose. Recommendation: close with a 33-file residual, cut fresh from main; the mechanical and contract slices come out empty after a trial merge.

Coordination

infra-0b was messaged before the gx10 reaper apply, as agreed: forjar plan is read-only first, and only ci-reaper-* / ci-disk-watch-* from infra origin/main get applied, never eph-* or #607's branch.

@noahgift

Copy link
Copy Markdown
Contributor Author

§7 interval — 19:58Z

scope: GDN 5/5 — on main (#3322)
       QHF 7/7 — on #3114 (draft). Red: a claim-literal guard (inference_result.rs:1213) + the silicon-coverage NO-GO (below)
       E2E 3/7 — on #3114; the other 4 are declared-N/A candidates waiting on the typed N/A schema (stacked on #3327)
       #3303: reference built + a per-token producer mode.
              llama.cpp batched vs per-token: min cos 0.998197, and it flips argmax against ITSELF at 2, 21, 36, 77 (orchestrator re-ran)
       #3091: parity — per-token confound RESOLVED (orchestrator re-ran all three comparisons):
                flips 2 and 36 were llama batch-vs-per-token near-ties; 28 and 73 persist against both llama modes (gaps 0.315 / 0.245)
              prompt variation, 5 prompts / 468 positions (worker receipt; orchestrator re-verification in flight):
                threads: apr byte-identical at 1 / 8 / 32 threads (pin proven by /proc Threads 2 / 9 / 33) → confound ELIMINATED
                12 persistent flips, ALL below the pooled median llama top-2 gap (1.148); 0 at a gap ≥ 1.0 → no defect candidate
                28 / 73 are ordinary members of that set (gap ranks 6 and 9 of 12)
                outlier: chat markup tokenized as plain text — min cos 0.961 @pos1, 4 positions < 0.98
                → layer-wise hidden-state diff dispatched (Lane B, now)
              estimate: code on #3114; left = 2 red guards on #3114 → main, the N/A chain (#3327 → schema-na), and the markup root cause [U]
       andon: 64h20m to §8 at 2026-09-18T12:19:39Z
pins:  llama.cpp #3331 red on 2 guards (DET002 at lane.sh:13; a pipe into grep -q) — fixing now

Landed / armed since the last interval

Red, with the cause read from the log (not the step name)

PR check cause action
#3327 guard-tree G2.1 cli.rs moved; surface_audit.csv not touched rows re-pointed through a line diff, byte-identical cited lines only; pushing
#3331 guard-cargo bashrs DET002 at evidence/parity/pin-bump-d1d3c3396/lane.sh:13 fixing
#3331 guard-tree check_no_pipe_into_grep_q.sh fixing
#3320, #3114 guard-tree check_silicon_coverage.sh NO-GO: "schedule=0" in CI, while the same script on the same tree run locally reads 60 scheduled runs and GO root-causing; not rerun-until-green

§8 decision, logged

gx10 / yoga reaper deploy: a blanket make provision is refused. Planned from a fresh worktree, forjar reads 76 creates, including enP7s7-dhcp, ollama-binary and the pip installs. That is because the state dir is gitignored and the worktree has none, not because the host drifted. The apply will be forjar apply -r over only the 9 ci-reaper-* / ci-disk-watch-* resources, against each machine's live state dir, and only after a scoped plan shows nothing outside them. infra-0b was notified before any apply.

@noahgift

Copy link
Copy Markdown
Contributor Author

§7 interval — 21:36Z

scope: GDN 5/5 — on main (#3322)
       QHF 7/7 — on #3114 (draft); its claim-literal red was line-keyed baseline drift → unmeasured tok/s deleted, pushed
       E2E 3/7 — on #3114; 4 N/A candidates wait on the typed N/A schema (stacked on #3327)
       #3303: reference built; per-token + tensor-dump producer modes; llama batched vs per-token flips against itself at 4/78
       #3091: layer observer (local measurement branch; logits byte-identical with it on — f6f79264 / 92f1b54d):
                embedding byte-identical apr = llama = gguf-py
                largest step at EVERY position measured: layer-0 DeltaNet mixer, across ONE matmul (Q5_K ssm_out)
                there apr = float64 dequant(W)·x to 1e-6; llama departs 0.034–0.042 — ggml quantizes the activation to Q8_K (worker receipt; source line + hashes being re-verified)
              → running now: apr with ggml's Q8_K activation quantization emulated (port proven bit-exact vs ggml C first).
                If apr then converges on llama, the parity gap is the reference's quantization, not an apr defect.
              estimate: #3114 code-complete; left = #3114 guards green → main, the N/A chain (#3327 → schema-na), the emulation verdict
       andon: 62h42m to §8 at 2026-09-18T12:19:39Z
pins:  llama.cpp #3331 red — DET002 (lane.sh:13 `date`, and its started_utc is cited by LEDGER + contract) + one pipe into grep -q; fix pending the precedent check

Moved since the last interval

§8 decisions, logged

  1. pv strictness splits. SCHEMA-024 (id required), SCHEMA-025 (unique) and R4 (N contracts evaluated; exit 2 only when nothing was evaluated) land first. SCHEMA-026 (kani citations resolve) stays an Error, never a warning, and lands after corpus repair. It measured 511 errors on the named corpus: 101 id-shaped citations dangle because contracts: 3,612 obligations had no id, so nothing could cite them #3320's generator numbered by position instead of reusing the ids kani already cites, plus 7 cross-contract prefix collisions and 389 prose citations.
  2. That generator bug is fixed in contracts: 3,612 obligations had no id, so nothing could cite them #3320 before it lands (ids must be stable), with selftest cases, per the ruling that generator bugs get a case. The mapping rule is adopted only if it scores 100% on the ground-truth contracts that already had ids.
  3. The observer stays on a local measurement branch, not in feat(qwen35): Qwen3.5 / Qwen3.8 hybrid GGUFs run on the CPU — apr run + apr chat (#3091) #3114. Whether to upstream it is decided once the emulation verdict shows its value.
  4. silicon-coverage NO-GO on contracts: 3,612 obligations had no id, so nothing could cite them #3320/feat(qwen35): Qwen3.5 / Qwen3.8 hybrid GGUFs run on the CPU — apr run + apr chat (#3091) #3114: CI read schedule=0 in the event-filtered listing, while the same script on the same tree reads 60 scheduled runs locally. Root cause in progress; it will not be rerun until green.

@noahgift

Copy link
Copy Markdown
Contributor Author

§7 interval — 2026-09-15T22:24Z

scope: GDN 5/5 — on main (#3322)
       QHF 7/7 — on #3114 (draft); claim-literal guard green (2cab0906b); red only on the silicon-coverage guard defect (issue + worker, below)
       E2E 3/7 — on #3114; 4 N/A candidates wait on the typed N/A schema (stacked on #3327)
       #3303: reference built; producer gains per-token, tensor-dump, and (running now) KV-type / flash-attn modes
       #3091: ggml vec_dot EMULATION on apr (a port bit-exact vs ggml's C on 10 fixtures; switch-OFF shas unchanged):
                mechanism CONFIRMED — apr equals llama to ≤2e-6 through layer 2 with it ON
                magnitude KILLED — it removes only 0.8–32% of the logits gap (orig 15.7%, p4 32.3%)
                the residual enters at layer 3, the first full-attention mixer, with matching input;
                at pos 0, f16 rounding reproduces llama's output 2044/2048 bit-equal (llama: f16 KV + flash-attn)
              → running now: the llama reference at f32 KV with FA off, and apr OFF/ON against A/B/C
              estimate: #3114 code-complete; left = the silicon guard fix → #3114 green → main, the N/A chain, the KV-config verdict
       andon: 61h55m to §8 at 2026-09-18T12:19:39Z
pins:  llama.cpp #3331 — DET002 fix verified in the CI image's bashrs 7.4.1, pushed (orchestrator gate: positive control on the old file, 0 codes on the new)

Landed / deployed

pv chain (re-verified by the orchestrator)

Correction to the previous interval

Decision 2 there ("the generator bug is fixed in #3320 before it lands") is withdrawn on measurement. On the 34 citations that already had ids, no binding rule scores 100% (position 33/34, type+ordinal 1/34). Naming also reduced dangling citations (576 on main → 497), so #3320 lands unchanged. Ids are names and stay stable; the citations are what gets repaired.

Red, with cause

PR cause action
#3320, #3114 check_silicon_coverage.sh: a global RUN_CAP=60 dropped the run carrying ada-yoga* (UNCOVERED), and an empty schedule event page (NO-GO) issue filed; a worker is making it read per axis workflow, with gh-shim rows RED on today's script
#3336 ci / lint, ci / coverage, guard-cargo diagnosing from the logs

@noahgift

Copy link
Copy Markdown
Contributor Author

§7 interval — 2026-09-16T05:35Z

scope: GDN 5/5 — on main (#3322)
       QHF 7/7 — on #3114 (draft)
       E2E 3/7 — the typed N/A schema is open as #3340 (one red: the pv surface-gate table needs the 3 new rules)
       #3303: reference, per-token, tensor-dump and KV-config producer modes — all evidence committed
       #3091: RESULT. apr's Qwen3.5 CPU forward is BIT-IDENTICAL to ggml's own arithmetic.
              Against a SCALAR-built llama.cpp d1d3c3396 (no AVX/FMA/repack) in config C (f32 KV, flash-attn off),
              with ggml's arithmetic emulated in apr: 1500/1500 + 750/750 dumped points bit-equal (max rel L2 0.0),
              the per-token logit STREAMS byte-identical (cmp rc=0), 0 argmax mismatches over 78+82 positions.
              Six iterations, five departures, EVERY one closed by reproducing ggml's float arithmetic —
              SSE2 exp vs libm, and double vs f32 accumulation — and NOT ONE by changing what is computed.
              There is no algorithmic difference between apr and ggml on this path.
              The shipped apr path never moved: switch-OFF logits stayed f6f79264… / 92f1b54d… at every iteration.
              Amplification floor measured: 1 ulp in, 1.5e-7 out, and only through the element that sets the
              Q8_K block scale; for 7 of 9 probed elements the output is bit-identical. The earlier 1224x was
              amplification of an already-1e-6 difference, not a floor.
              left: #3114 green → main, the N/A declarations (#3340), and the evidence PR series (being prepared)
       andon: 54h44m to §8 at 2026-09-18T12:19:39Z

0.68 scope (operator ask, done this interval)

The repo has no per-release label — milestone 0.68.0 (ms#5, due 2026-09-15, already overdue) is the release scope. Qwen3.5 was already on it (#3091, #3114, #3303); the rest of this train's work had fallen off it, so #3320, #3331, #3340, #3297, #3313, #3318, #3321, #3324, #3325, #3328, #3329, #3330, #3332, #3334, #3335, #3336, #3337, #3338 are now on 0.68.0, and new PRs are filed onto it as they open. This is release scoping, not the CI label edits the standing rule forbids.

Landed / open

Red, with cause read from the log

  • contracts: 3,612 obligations had no id, so nothing could cite them #3320 'Chapter Examples Compile' is ENV, not code: the intel clean-room runner could not execute …/rustup/toolchains/1.93.0-x86_64-unknown-linux-gnu/bin/rustcNo such file or directory — and lost target/debug/deps/*.d mid-build. That runner root now holds only 1.89.0 and 1.91. Tracing which sweeper removed a toolchain under a live build; intel still runs the OLD reaper (96ec2124 vs repo 8304c8c8), since only gx10 was applied.

@noahgift

Copy link
Copy Markdown
Contributor Author

§7 interval — 2026-09-16T06:15Z

scope: GDN 5/5 — on main (#3322)
       QHF 7/7 — on #3114
       E2E 3/7 + 4 N/A — the schema is #3340 (now green on its own tests: SCHEMA-021/022/023 got REAL rows in the
                 pv-surface decision table, each mutation-proven, not an exemption). The four QE2E declarations
                 are being written now, each with a grounded reason and owner.
       #3303: reference + per-token + tensor-dump + KV-config producer modes, evidence committed
       #3091: RESULT STANDS — apr's Qwen3.5 CPU forward is BIT-IDENTICAL to ggml's own arithmetic
                 (scalar llama.cpp d1d3c3396, config C): 2250/2250 dumped points bit-equal, logit streams
                 byte-identical, 0 argmax mismatches, no algorithmic departure in six iterations.
              The evidence stack is rebased onto the pin branch and has a 3-PR landing plan; 25 of 25 in-repo
              sha256 citations re-hashed OK (2 stale citations found and fixed), 8 broken *.log citations
              replaced with committed transcripts. It opens as soon as #3331 is green.
       andon: 54h04m to §8 at 2026-09-18T12:19:39Z

The train is blocked on ONE thing, and it is on main

check_baseline_ratchets.sh and check_complexity_ratchet.sh fail on a pristine origin/main worktree:

FAIL tool_version: scripts/cb200_baseline.txt was recorded under pmat 3.40.1,
     runner has pmat 3.40.2 — verdicts would compare two instruments

That is not a PR defect — every open PR inherits it through guard-tree. It is also the guard being right: two instruments cannot be compared. #3291 already records that the fleet declares 3.40.1 while aprender's pins say 3.40.0 and /opt/ci-tools lags both, so the version disagrees in three places.

A worker is measuring pmat on every host and in the CI image first, then restamping by re-measuring, never by editing a number: any baseline value that moves is a ticket, not a restamp. Converge, prove, then move the assertion.

Opened / fixed this interval

0.68 scope (operator ask) — done

Milestone 0.68.0 is the release scope; the repo has no per-release label. 19 items of this train were off it and are now on it, all 13 unlabeled 0.68 issues now carry a type label, and every PR opened since is filed onto 0.68.0 at creation. Qwen3.5 is on it: #3091, #3114, #3303, #3320, #3331, #3340.

@noahgift

Copy link
Copy Markdown
Contributor Author

§7 correction — two claims from the last interval are WITHDRAWN

Both were falsified by workers who checked the premise instead of executing it. Recording them here because both were published above.

1. "The train is blocked on main" — WRONG. Main is green; this workstation is drifted.

I read check_baseline_ratchets / check_complexity_ratchet failing on a pristine origin/main worktree and reported main as red. Measured properly:

where pmat bashrs
intel, gx10, yoga, mini (the 28 registered runners) 3.40.1 7.4.1
repo pins + both baselines + .pmat-gates.toml 3.40.1 7.4.1
infra forjar, every machine including this one 3.40.1 7.4.1
lambda-labs (this box, NOT a runner) 3.40.2 7.3.0

guard-tree is success on main at a6066555f and the two runs before it. Main already carries the 3.40.1 restamp (#3301). The red reproduces only here, and restamping to 3.40.2 would have asserted a version zero runners have — converge, prove, then move the assertion, in that order. No repo change was made.

Two consequences worth recording. ~/.cargo/bin/pmat on this box was rewritten at 2026-09-16 01:13, after the fleet settled on 3.40.1 — the un-owned ALWAYS-LATEST mechanism in #3291 is still running here. And it cuts both ways: my earlier local bashrs checks ran 7.3.0 against a 7.4.1 fleet, which is exactly why lane.sh's DET002 would not reproduce locally and had to be verified on intel.

2. "E2E 3/7 + 4 N/A pending the schema" — WRONG. None of the four is a device claim.

The contract says so, and the contract wins:

obligation what it actually is
QE2E-INV-001 a deterministic sum over parameter shapes
QE2E-BND-002 arithmetic over architecture constants
QE2E-MON-004 monotonicity of min(bw/x, c)tok_s is defined by this contract's own equation, not measured on a device
QE2E-BND-005 a corpus claim, and currently FALSE: bindings_implemented is 0 of 6

All four are CPU-decidable. Declaring them N/A would have moved obligations out of the unproved column without proving anything — the exact laundering the N/A rules were written to prevent, applied to the ticket that wrote them. Zero N/A declarations were made.

The real root cause is that this contract binds 0 of 6 obligations to a test. Lane B is re-scoped to that: write the four tests, tighten QE2E-BND-002's O(...) into a concrete bound (it cannot fail as written), and bind them. #3340 stands on its own merits — the schema plus the migration of 5 untyped na_reason justifications serde was silently dropping.

The scope line should read E2E 3/7, 4 provable on CPU and unbound — not "4 N/A".

@noahgift

Copy link
Copy Markdown
Contributor Author

§7 interval — 2026-09-16T07:17Z

scope: GDN 5/5 — on main (#3322)
       QHF 7/7 — on #3114
       E2E 6/7 bodies — three QE2E tests written and pushed to #3114 (INV-001 parameter sum, BND-002 FLOPs bound,
                 MON-004 roofline monotonicity), each with a mutation that turns it RED. NOT declared N/A: the
                 contract shows all four are CPU-decidable, so N/A would have been laundering.
                 QE2E-BND-005 stays UNPROVED and false on purpose — bindings are 0 of 6 (#3347)
       #3303: reference + per-token + tensor-dump + KV-config producer modes
       #3091: result stands — apr's Qwen3.5 CPU forward is BIT-IDENTICAL to ggml's arithmetic
                 (2250/2250 points, byte-identical logit streams, no algorithmic departure in six iterations)
              the 3-PR evidence series is rebased and waiting on #3331
       andon: 53h02m to §8 at 2026-09-18T12:19:39Z

Every red on the board was read from its log, and each had a different cause

PR cause fix
#3342, #3344 check_dogfood_coverage G2.1 — the slices move code the ledger cites by path:line 55 and 56 rows re-pointed through a line-level diff, accepting only rows whose cited TEXT is byte-identical; a normalise-and-cmp proves every other byte of the ledger unchanged
#3344 ont-delta: restores is outside the vocabulary {type|shape|reason|resolves|none} body rewritten to resolves REG-OB-004 …, proved RED→GREEN offline with the guard's own --body/--changed mode
#3331 check_hardcoded_paths — the evidence lane wrote into ONE agent session's scratchpad, so nobody else could re-run it the path becomes ${LANE_OUT:?}
#3114 no ont-delta: line (§11.1) added, and the body was edited BEFORE the push so the guard sees a fresh payload

Corrections — three claims of mine withdrawn

  1. "The train is blocked on main." No: guard-tree is green on main; the ratchet red was this workstation's instrument skew.
  2. "E2E 3/7 + 4 N/A." No: all four are CPU-decidable. Three now have bodies; the fourth is false and stays failing.
  3. "An un-owned ALWAYS-LATEST mechanism rewrote pmat at 01:13." No — timezone. stat prints local time; TZ=UTC stat shows 23:13 UTC, which is exactly infra-0b's authorized cargo install pmat 3.40.2 after they published it. There is nothing un-owned to hunt.
    That one had a consequence: my scoped forjar apply -r stack-tool-pmat converged this box back to the declared 3.40.1, undoing a version the operator had asked for. I have reinstalled 3.40.2 with the same command, and I will not re-run that convergence while the pin question is open. The bashrs half (7.3.0 → the declared 7.4.1) was real drift and stays — it immediately paid off: a DET002 finding that would not reproduce under 7.3.0 now reproduces locally.
    Standing rule from this: TZ=UTC stat, never bare stat, before drawing a conclusion from a file's mtime.

Filed

0.68

Milestone 0.68.0 carries this train: 19 items added, all 13 unlabeled issues typed, every new PR filed onto it at creation. Qwen3.5 is on it.

@noahgift

Copy link
Copy Markdown
Contributor Author

§7 interval — 2026-09-16T08:13Z

scope: GDN 5/5 — on main (#3322)
       QHF 7/7 — on #3114 (draft, per the standing decision)
       E2E 6/7 bodies on #3114; QE2E-BND-005 stays FALSE and failing (bindings 5/6, #3347/#3348)
       #3303: reference + per-token + tensor-dump + KV-config producer modes
       #3091: apr's Qwen3.5 CPU forward is BIT-IDENTICAL to ggml's arithmetic — 2250/2250 points,
              byte-identical logit streams, no algorithmic departure. 3-PR evidence series waits on #3331.
       NEW, and it is the strongest number of the night:
              apr's parameter arithmetic now reproduces a REAL Qwen3.5 GGUF EXACTLY.
              ~/models/Qwen3.5-0.8B-Q4_K_M.gguf, header parsed directly: 320 tensors, 752,393,024 params.
              model_parameter_count, fed the descriptor: 752,393,024. Delta ZERO, both layer kinds
              matching tensor for tensor. Dropping the attn_gate term turns it RED by exactly d*inner_size.
       andon: 52h06m to §8 at 2026-09-18T12:19:39Z

What that measurement found on the way

ModelConstraints carried none of the five shape keys contracts/model-families/qwen3_5.yaml declares (inner_size, state_size, conv_kernel, group_count, full_attention_interval), so 18 of every 24 Qwen3.5 layers were counted as if their conv, gates, state norm and mixer projections did not exist. Two shapes in the real file also contradict dense accounting and are now modelled from the tensors rather than assumed: attn_q is [1024, 4096] = 2·n_h·d_k (the q projection emits the output gate), and the file is tied — no output.weight.

QE2E-INV-001 is still not asserted, and the range was not widened. P(9B) moves 8.209B → 8.345B, still 0.655B below [9.0B, 9.2B], and the 9b descriptor contradicts itself: inner_size: 2048 at four times the hidden dim, and group_count: 8 failing group_count·state_size == inner_size where the measured 0.8B satisfies it at 16·128. A real Qwen3.5-9B GGUF settles it in minutes; none is on the fleet, so the obligation stays unproved and the number is pinned by a test.

Open PRs, all armed, all filed on 0.68.0

PR what
#3331 the llama.cpp pin bump (the evidence series' base)
#3342 / #3344 / #3349 #3041 slices S3a / S3b / S3c — library, caller switch + static guard, effective-config + runtime refusal
#3348 5 of 6 qwen35-e2e equations bound to real implementations
#3350 the DeltaNet shape keys (above)
#3320 obligation naming — armed this interval

Honest counters, not flattering ones

  • The 6th binding, verification_ladder, is left unbound and NOT marked pendingcount_binding_coverage excludes Pending from the denominator, so "pending" would have displayed a clean 5/5 and hidden the gap. It reads 5/6.
  • Lane B is now on the reason it cannot be defined: pv's L2 column computes idx < falsification_tests.len(), an index comparison, so all 7 obligations showed ✓ before any test existed. A column that ticks on a count cannot report coverage. The corpus-wide before/after will be reported as measured, and a large drop is the point.

@noahgift

Copy link
Copy Markdown
Contributor Author

§7 interval — 2026-09-16T09:13Z

scope: CRITICAL PATH (operator ruling): #3331 -> parity evidence series (3 PRs) -> #3114 out of draft -> #3091 -> cut
       #3331  GREEN, armed, 0 red, 0 pending — the last red (check_hardcoded_paths, +4) was paid down route A:
              the 4 /props captures stay as recorded; 7 OTHER portability findings fixed; delta +4 -> -3, PASS
       parity series, all 6 branches rebased onto the current pin head 920385c9c, clean, pin self-test rc 0 each:
              PR1 (reference + raw-logits)      0 paydown, 1 cycle   — all five guards green now
              PR2 (measurement + per-token + variation)  29 paydown  — ONE authored file, classify_spec.json
              PR3 (layerwise: emulation/KV/scalar)        8 paydown  — 7x SEC011 rm -rf $VAR, 1x SEC001 in a COMMENT
              both paydowns dispatched; they are visible locally, so each is 1 cycle if cleared before opening
       estimate vs andon: ci.yml PR runs p50 26.7 min / p90 75.8 min (n=16, measured; guard-tree p50 4.1, guard-cargo p50 8.3)
              stacked PRs are CI-dark, so the three are SERIAL: 3 x (PR run + post-merge main run)
              machine time 3.5 h p50 / 9.2 h p90; series finishes 2026-09-16T12:30Z (p50) .. 2026-09-17T06:00Z (p90)
              margin to andon: ~30-48 h. NOT past the andon on CI time.
              [U] review latency and merge-queue wait — not measured; that, not CI, is what could still move it
       andon: 51h06m to §8 at 2026-09-18T12:19:39Z
pack:  (from `make -C machines/intel verify-fleet-bin` + /proc; no runners API)
       intel  listeners(online)=16  workers(busy)=15  load 44.6  free 401G   fleet-bin: 16/16 live listeners converged, 0 stale/foreign
       gx10   listeners(online)=6   workers(busy)=2   load 9.9   free 61G
       yoga   listeners(online)=5   workers(busy)=3   load 12.0  free 354G
       mini-m4: idle BY DESIGN (guide §3) and NOT a P0 — measured: 0 workflows in .github/workflows name a
                macos/mini-m4/apple runner label, so no Apple-irreducible job is queued or queueable today
ms#5:  0.68.0 open = 2 issue(s) + 2 PR(s)  (was 296 open items)

Ruling 1 — ms#5 retargeted, mechanically, no discretion

296 open items were on 0.68.0. Kept: #3091, #3208, #3114, #3331 (#3303 and #3335 are closed/merged, and PR3 joins when it opens). Everything else moved by label: P0/P1 → 0.69.0 (61), everything else → 0.70.0 (231), 0 failures. 0.68.0 now reads as the Qwen3.5 CPU floor, so "overdue by a day" measures the real critical path.

Ruling 3 — clean-room-on-tag is now a gate, not a runbook line

paiml/infra#622 is open (closes infra#621), and it is the durable fix:

  • a ref input (sha or tag), and a first-step assertion that HEAD equals it or the job fails immediately, printing both;
  • the tested commit recorded as structured outputtested_sha (full 40 hex) in results.csv and tested-sha: in the step summary — with the Makefile's commit: <abbrev> log line byte-unchanged, so aprender's existing gate keeps working and can move off the log parse next release.

Three measurements decided it: git clone --depth 1 --branch <40-hex> exits 128, so --branch cannot express the one form a release gate needs; after git fetch origin <tag>, FETCH_HEAD is the tag object while checkout leaves HEAD on the peeled commit, so every ref resolves through ^{commit}; and a 4-char prefix resolves cleanly in a fresh shallow clone, so abbreviations are refused by policy. Falsifier: 35 rows, and disabling only the refusal flips 4 rows RED.

It also caught the consumer that would have broken at T-4: sync-readme.sh folds extra commas into duration_sec, so a new column would have killed fmt_duration.

Reaper: applied to all three, per the ruling

ci-reaper.sh is now 8304c8c8 (= repo) on intel, gx10 and yoga; timers active; NeedDaemonReload=no everywhere. gx10's ledger is writing hourly (latest: reclaimed to 76 GB free from 2 GB). yoga applied per-resource via forjar, plan-gated; intel via its own deploy-systemd-units.

Two findings from doing it, filed in infra: make verify-systemd-units dispatches two targets that do not exist (verify-ci-hardware, verify-metric-liveness), so the aggregate has never been able to return 0 — a verification nobody could act on. And intel's /var/lib/ci-reaper was missing after the make deploy and had to be converged with forjar apply -r ci-reaper-state-dir: the make path and the forjar path disagree about who owns the directory the reaper needs for its counters and ledger.

@noahgift

Copy link
Copy Markdown
Contributor Author

§7 interval — 2026-09-16T09:32Z

scope: CRITICAL PATH — #3331 -> parity series (3 PRs) -> #3114 out of draft -> #3091 -> cut
       #3331  UNSTABLE red=[present] pending=1 ; merge queue: #3331@1
       parity series OPENED, all three at once so their CI runs in PARALLEL, not serially:
              PR1 #3354  reference + raw-logits        ARMED   (0 red, CI running)
              PR2 #3355  measurement + per-token + variation   unarmed until PR1 merges (order enforced by arming)
              PR3 #3356  layerwise (emulation/KV/scalar)       unarmed until PR2 merges
              both blocking paydowns were cleared BEFORE opening, so the 1-cycle estimate holds:
                PR2's +29 machine paths were ONE authored spec -> ${HOME}-rooted, helper taught to expand,
                  and flips.json re-run on intel reproduces BYTE-IDENTICAL (cmp rc 0) before the citations moved
                PR3's 8 bashrs findings: 7x rm -rf through an unvalidated var (now "${VAR:?}/literal"), 1x the
                  word "eval" in a COMMENT -> reworded, no suppression anywhere
       estimate vs andon: ci.yml PR runs p50 26.7 / p90 75.8 min (n=16, measured). Opening in parallel removes
              two serial CI waits from the earlier estimate: series finishes ~2026-09-16T11:00Z (p50)
              .. 2026-09-16T18:00Z (p90) if each merges on its first green. Margin ~44-51 h. NOT past the andon.
              [U] review latency and merge-queue wait — still the only thing that could move it; not measured
       cut gate: PR3-of-the-queue-architecture is open as #3352 (a roadmap edit without its fragment is refused)
       andon: 50h47m to §8 at 2026-09-18T12:19:39Z
pack:  intel:listeners=16,busy=15,load=52.95,free=708G gx10:listeners=5,busy=5,load=15.06,free=82G yoga:listeners=5,busy=4,load=16.06,free=377G 
       mini-m4: idle BY DESIGN and not a P0 — 0 workflows name a macos/mini-m4/apple label, so no Apple job is queueable
ms#5:  2 issue(s) + 6 PR(s) open (was 296 items; kept = the Qwen3.5 CPU floor + the cut gates)

Reaper: applied to all three, and the apply found the defect that mattered

ci-reaper.sh is 8304c8c8 (= repo) on intel, gx10 and yoga; timers active; state dirs present. gx10's ledger writes hourly. intel's has never been written, and the apply is what exposed it:

Result=timeout  ExecMainStatus=15/2   Consumed 6min 40.560s CPU time
09:12:03  ci-reaper.service: start operation timed out. Terminating.

It was not hung — it was deleting a per-PR registry copy every ~6.5 s until SIGTERM. 183 copies wait on intel (12 on gx10, 7 on yoga), so one rule needs ~20 min against a 10-min unit budget: the run is killed inside it, every rule after it has never run there, and the escalation is blind on the host that needs it most.

paiml/infra#627 is open (closes infra#626): the sweep is budgeted and resumable rather than given a bigger timeout — a bigger timeout was rejected because the script holds a flock, so a genuinely stuck reaper would silently swallow every later hourly fire instead of failing. reaper_finish is now always reached; the ledger gains status: partial, remaining, deferred; TimeoutStartSec keeps its meaning as a hang detector. Falsifier: 3 new legs forced by a stepping clock in a file — 6 rows RED against a pristine origin/main copy, 0 failures in the 11 pre-existing legs.

Also filed from the same apply: infra#625make verify-systemd-units dispatches two targets that do not exist, so the aggregate the operator asked me to verify with has never been able to return 0.

Also filed

#3353resolve_base's git rev-list --first-parent | grep -qx SIGPIPEs the producer: rc 0 without pipefail, 141 with it, and every caller sets it. So the "HEAD is on the first-parent line" arm is unreachable under CI's own shell, and diff-scoped guards silently compare against a different base. check_roadmap_diff_additive.sh inherits it.

@noahgift

Copy link
Copy Markdown
Contributor Author

§7 interval — 2026-09-16T09:48Z

scope: CRITICAL PATH — #3331 -> parity series -> #3114 out of draft -> #3091 -> cut
       #3331  merge queue POSITION 1, AWAITING_CHECKS, 1 check outstanding, 0 red that is required
              (`present` is the review backlog and is not a required check)
       parity series, all three open and running CI IN PARALLEL rather than serially:
              PR1 #3354 reference + raw-logits         ARMED, 15 checks running, 0 required-red
              PR2 #3355 measurement + per-token + variation   17 running — arms when PR1 merges
              PR3 #3356 layerwise (emulation/KV/scalar)       12 running — arms when PR2 merges
       cut gates: #3352 (roadmap fragment required) 1 check left; #3357 (cascade reads the structured tested-sha) 14 running
       #3114  BEHIND, draft, 5 running. Its body now says `Closes #3091`, and GraphQL confirms the link
              (closes #3091 and #3303). The body guards read the FROZEN payload, so that takes effect on its
              next push — which is the post-series update-branch, not a silent fix.
       estimate vs andon: unchanged and measured — ci.yml PR runs p50 26.7 / p90 75.8 min (n=16).
              Series finishes ~2026-09-16T11:00Z (p50) .. ~18:00Z (p90) if each merges on its first green.
              [U] review latency and merge-queue wait — still the only thing that could move it.
       andon: 50h31m to §8 at 2026-09-18T12:19:39Z — NOT threatened by CI time
pack:  intel  listeners=16 busy=15 load 72.9  free 693G
       gx10   listeners=5  busy=4  load 15.3  free  26G   <-- watch: see below
       yoga   listeners=4  busy=2  load 38.0  free 358G
       mini-m4: idle BY DESIGN, not a P0 — 0 workflows name a macos/mini-m4/apple label, so no Apple job is queueable
       lambda (not a runner): load 7.9, / 358G, raid 2.7T
ms#5:  2 issue(s) + 7 PR(s) open — the Qwen3.5 CPU floor plus the cut gates

Reaper ledger: 2 of 3, and the third is diagnosed, not unknown

The operator asked for ledger records on all three. Measured just now:

host script last run ledger
gx10 8304c8c8 = repo 09:42 success 86 lines, hourly
yoga 8304c8c8 = repo 09:02 success 1 line — its first fire reclaimed 113 GB (61 sweeps)
intel 8304c8c8 = repo 09:02 timeout ABSENT

intel is the one that cannot write it, for the reason already measured: 183 per-PR registry copies at ~6.5 s each is ~20 min of work against a 10-min unit budget, so the run is SIGTERM'd inside one rule and reaper_finish is never reached. paiml/infra#627 is open and armed with the fix — the sweep is budgeted and resumable, the ledger is always written, and TimeoutStartSec reverts to being a hang detector. intel's ledger will exist on its first run after that merges and is deployed; until then I am not claiming 3 of 3.

Fleet risk worth naming: gx10 is reclaiming slower than CI consumes

gx10's own ledger line from 09:46 reads free_gb_before: 50 → free_gb_after: 23, having reclaimed 2.2 GB — i.e. free space fell while its sweep was running, and it is now at 26 GB against a declared REAPER_CRITICAL_GB=90. The reaper is working; the intake is simply faster. That matters here because gx10 is running jobs for the parity series on the critical path, and a disk-full runner is how #3336 lost its ci / lint and ci / coverage earlier tonight ("because the disk was full"). Watching it per interval; if it keeps falling I will say so rather than discover it as a red PR.

Since the last interval

  • paiml/infra#622 MERGED — ruling 3's durable fix is live: clean-room.yml takes a ref, asserts HEAD == tag in its first step or fails, and records tested-sha as structured output. The interim T-3 is now executable end to end: tag → gh workflow run clean-room.yml -f ref=<tag> → job asserts → cascade refuses without that run id.
  • feat(release): the cascade gate reads the structured tested-sha, and still refuses everything else #3357 opened — the aprender half: the cascade gate prefers the structured record, falls back to the strict log line, and requires the assert step to have passed. Case table 25 → 35 rows; two mutations each turning named rows red; the v0.66.0 and v0.67.0 retro still REFUSE, now with the reason each gives.
  • PR3: a roadmap edit without its fragment is refused — the contention #3297 removed cannot come back #3352's only red is cleared — the new guard's own case table piped into grep -q, which check_no_pipe_into_grep_q refuses and is right to: under pipefail the pipeline reports the producer's SIGPIPE, not grep's verdict, so a row could have passed on a death rather than on a match.

@noahgift

Copy link
Copy Markdown
Contributor Author

§7 interval — 2026-09-16T10:17Z

scope: CRITICAL PATH — #3331 -> parity series -> #3114 out of draft -> #3091 -> cut
       #3331  OPEN UNSTABLE pending=0
       PR1 #3354 armed, ENOSPC-killed jobs re-run; PR2 #3355 and PR3 #3356 STOOD DOWN (8 runs cancelled)
       estimate vs andon: unchanged on CI time (p50 26.7 / p90 75.8 min, n=16). Serialising the three PRs
              costs the parallel wall-clock I should not have taken; it does not threaten the andon.
       andon: 50h02m to §8 at 2026-09-18T12:19:39Z
pack:  gx10 97% used, 32G free (see below) ; ms#5: 2 issue(s) + 7 PR(s)

Correction: my gx10 diagnosis was wrong, and the intake was mine

I reported the reaper as "losing a race" — reclaiming 2.2 GB per run while intake outpaced it. infra-0b measured it properly and that framing does not survive:

  • The reaper's 09:52Z pass logged free 2GB -> 67GB, swept=0. 65 GB came back because jobs ended, not because the sweep worked. It was never racing.
  • The fill was three live build trees, not many small stale ones: gh-readonly-queue 67 G, 3355 47 G, 3354 46 G — all written within the hour.

3354 and 3355 are mine. I opened all three parity evidence PRs at once specifically to get their CI cycles in parallel rather than serially. That optimisation put three ~47 GB target trees on a 916 GB host at the same time, and it is what took gx10 to 0 bytes. The serial order was already enforced by arming, so the parallelism bought nothing the queue would not have given me anyway.

This is admission control, not a reaper defect. paiml/infra#627 correctly does not address it.

What I did about it, in my own lane

  1. Cancelled 8 runs (evidence(parity): apr's Qwen3.5 CPU forward vs the d1d3c3396 reference — 78 positions, per-token prefill, 5 prompts #3355 and evidence(parity): with ggml's arithmetic emulated, apr's Qwen3.5 CPU forward is BIT-IDENTICAL to scalar llama.cpp #3356, in-progress and queued), leaving PR1 and the merge-queue build. PR2's CI runs when PR1 merges, PR3's after PR2 — one build tree at a time.
  2. Reclaimed 44.5 GB of aprender's own stale cargo output: ~/data/targets/{aprender,realizar,trueno}CACHEDIR.TAG present, newest write 2026-04-03, no holder, and cited nowhere in the tree. gx10 went 0 → 37 GB free.
  3. Deliberately did NOT delete 108 GB more, because the citation check came back positive: ~/runs (33 G) is cited by two specs and an evidence/.../findings.json, and the distill-7b-* runs by three contracts. Deleting cited evidence to free disk would fabricate a cleaner history. ~/.apr/cache/hf (62 G) is regenerable but the honest fix there is the knob, not my hand.

All three PR reds were the disk, not the code

#3354 workspace-test — "Failed to create … because the disk was full". #3355 gpu-quick — "rustc-LLVM ERROR: IO failure on output stream: No space left on device". #3356 guard-tree — "cannot create directory at 'evidence/dogfood/0.67.0': No space left on device". PR1's have been re-run now that space exists.

Filed / open from this

paiml/infra#628 carries three real measurement defects found on the way, none of which is today's cause and the body says so: REAPER_APR_HF_BUDGET_GB=80 sits above the 62 GB occupancy so a correct rule logs ok: 62GB <= budget 80GB hourly and never fires; ~/data/sccache is a symlink and du -sb/find do not follow a start point, so that budget reads 27 bytes and 0 of 13,941 files and has never swept a byte; and ~/data/targets matched no rule at all.

@noahgift

Copy link
Copy Markdown
Contributor Author

§7 interval — 2026-09-16T10:45Z

scope: CRITICAL PATH — #3331 -> parity series -> #3114 out of draft -> #3091 -> cut
       #3331  in the merge queue, position 1, building. Required check (gate) SUCCESS; only 'present' red, not required.
       PR1 #3354 armed, rebuilding after its ENOSPC kill. PR2/PR3 stood down, serial by design.
       estimate vs andon: CI time unchanged (p50 26.7 / p90 75.8 min, n=16). Serialisation costs wall-clock I
              should not have taken in the first place; the andon is not threatened by CI.
       andon: 49h34m to §8 at 2026-09-18T12:19:39Z
pack:  gx10 83% used, 152G free (0 free at 10:16Z) ; intel 16 listeners ; yoga 4 ; mini-m4 idle by design, 0 Apple jobs queueable
ms#5:  2 issue(s) + 7 PR(s)

gx10: 0 → 153 GB free, and the fix was mine to make

The host hit 0 bytes at 10:16Z and was killing critical-path jobs. Recovery, in order:

  1. Cancelled 8 runs of my own PR2/PR3 — I had opened all three parity PRs in parallel for CI speed, which put three ~47 GB build trees on a 916 GB host simultaneously. That was the intake.
  2. Reclaimed 44.5 GB of aprender's stale cargo output (~/data/targets/*, CACHEDIR.TAG present, 5 months idle, unheld, cited nowhere).
  3. Reclaimed 47 GB more — PR2's build tree, stale by construction once I cancelled its runs.
  4. Left 108 GB alone: ~/runs and the distill-7b-* runs are cited by two specs, an evidence/.../findings.json and three contracts; ~/.apr/cache/hf (62 GB) is regenerable but the honest fix there is the knob, not my hand.

infra-0b's measurement corrected mine and is worth recording: the reaper was never losing a race — its 09:52Z pass logged free 2GB → 67GB, swept=0, i.e. the space came back because jobs ended. This was admission control, not a reaper defect.

The reaper is deployed to all three — and could not take effect

ci-reaper.sh is now cc38a5604e (= repo) on intel, gx10 and yoga, via forjar per-resource on gx10/yoga and the make target on intel. But every host still read needreload=yes, REAPER_RUN_BUDGET_SEC=NONE: the file converged and systemd kept serving the old unit.

paiml/infra#629 is open with the fix, and the sweep found far more than the one resource I filed: 35 activation checks across 6 configs were blind to unit content — they asked only loaded/enabled/active, all of which stay true when the unit body changes. Three intel tasks had no daemon-reload in their command at all. And make deploy-systemd-units reloads before the forjar apply lines that write the units, so its own reload could never cover its own writes.

That is the third and fourth instance in 24 hours of a number that cannot go red — after the fail-open roadmap sync and pv's L2 index comparison. Falsifier: 35 violations and 6/6 rule-B failures against the pre-fix tree, 0 and 6/6 after.

Consequence for the operator's "ledger records for all three": gx10 95 lines, yoga 2, intel still ABSENT — and it will stay absent until #629 lands, because intel's fix lives in the unit env that never loaded. I am not hand-running systemctl daemon-reload on three shared hosts to close a status line.

Infra board

#622 MERGED (clean-room ref + HEAD==tag assert — ruling 3's durable fix is live). #627 MERGED (reaper budget/resumability). #628 rebased, MERGEABLE, armed (the sccache symlink reading 27 bytes, the apr-hf budget above occupancy, the unwalked 45 GB root). #629 open and armed.

@noahgift

Copy link
Copy Markdown
Contributor Author

§7 interval — 2026-09-16T11:19Z

scope: CRITICAL PATH — #3331 MERGED (11:0xZ). main is now d83592af8.
       parity series rebased --onto main from the pin head (squash, so a plain rebase would have replayed
              the pin's commits against their own squashed copy). All six branches clean:
              PR1 #3354 now 2 commits / 9 files / +739-0 — the pin's diff is gone from it
              PR2 #3355 and PR3 #3356 REBASED BUT NOT PUSHED, on purpose: pushing them restarts their CI and
              puts three ~47GB build trees back on gx10, which is what I did this morning
       #3114  conflicted on #3331's squash; both conflicts are the SAME portability fix landed twice
              (gguf_real.rs, tests_encode_decode.rs) — main's side taken, nothing of #3114's own work dropped
       estimate vs andon: CI time unchanged; the live risk is not CI, it is the wedge below
       andon: 48h59m to §8 at 2026-09-18T12:19:39Z
ms#5:  2 issue(s) + 6 PR(s)

Ledger records for all three — SATISFIED, and the evidence is better than the status line

intel wrote its first ledger line at 11:07:19Z. Not because #632 landed (it has not), but because #627's budgets carry script-level defaultsREAPER_RUN_BUDGET_SEC="${REAPER_RUN_BUDGET_SEC:-420}" at line 249 of the deployed script — so the unit env was never what gated it. My earlier "intel is blocked on #632" was wrong on that point.

The line itself is #627's design working end to end:

{"ts":"2026-09-16T11:07:19Z","host":"mac-server","free_gb_before":870,"free_gb_after":1005,
 "swept":30,"reclaimed_bytes":56731312299,"reason":"partial: 105 unit(s) and 0 step(s) deferred"}

It stopped cleanly at its budget, recorded what it left, and reclaimed 56.7 GB — where the same host was previously SIGTERM'd mid-sweep and recorded nothing. Current: intel 1, gx10 101, yoga 3.

#632 is still correct and still worth landing — 35 activation checks blind to unit content is a defect either way — but it is no longer what gates this line.

PR1 is wedged on a run that never dispatched a job

CI run 35088329165 for #3354: created 11:04:08, updated_at 11:04:09, status pending, jobs: 0, ten minutes later — while gx10 had 6 listeners / 4 busy and yoga 5 / 3. That is not starvation. It is the concurrency-group wedge: ci.yml groups on ci-${{ pull_request.number }}, and the force-push that rebased this branch replaced an in-flight run in group ci-3354.

Consequence: both required contexts (gate, workspace-test) are ABSENT from #3354 — not failing, absent — so it reads BLOCKED with one check (present, not required) and auto-merge cannot enqueue it. #3352 and #3357, untouched by a force-push, enqueued normally at positions 1 and 2.

Cancel + rerun did not re-dispatch, so I am minting a fresh event. Recording the shape because it is a trap with a force-push in any rebase-onto-main flow, which is exactly what a squash-merged stack requires.

@noahgift noahgift mentioned this pull request Sep 16, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant