Skip to content

Latest commit

 

History

History
413 lines (310 loc) · 125 KB

File metadata and controls

413 lines (310 loc) · 125 KB
type project-state
title PPT-LAT — PROJECT STATE LEDGER (the proven record)
description This is the authoritative record of what is PROVEN.
tags
project-state
timestamp 2026-06-18 05:48:11 UTC
resource shannon-prime-lattice/papers/PPT-LAT-STATE.md
sp_status GREEN
sp_gate none
sp_commit TBD
sp_repro none

PPT-LAT — PROJECT STATE LEDGER (the proven record)

This is the authoritative record of what is PROVEN. Read it first. Trust it. Build on it. Do NOT re-derive it.

To the next session/agent: this is the 20th rewrite. Six months of validation live behind these lines. The operator remembers; you don't. Your default is TRUST. Each PROVEN line cites its evidence (commit / tag / gate / closure doc). If — and only if — you have a concrete reason to doubt a specific line, the reverify command is given; running it is the exception, not the warm-up. Re-proving the whole stack from scratch is the failure mode this project has hit 20 times. Don't.

Status vocabulary: [PROVEN] evidence cited · [WIRED] built + in-tree, gated · [DESIGN] spec'd, unbuilt · [TARGET] a number to measure, not yet measured · [SPECULATIVE] idea. Promotion requires a gate + an entry here.

READ FIRST, before any work: papers/PPT-LAT-Theory.md — the canonical theory (the 13-step PPT substitution, O_K/Q(√−163), CRT primes, the frozen Spinor + KSTE formats, theorems T1–T8, production status). Skipping it caused real drift this session (a fresh agent inverted the PPT/Lattice hierarchy AND measured the wrong Spinor primitive — both because the theory wasn't read). It IS in the repo; read it. Companion docs: PPT-LAT-Theory.md (the math/why — FIRST) · RFC-001 (north-star preamble) · PPT-LAT-Systems-v1.md (the canonical systems narrative — supersedes v0 Systems + the two standalone v0 specs, which are now its Appendices A/B) · the C1–C6 contracts (forward work) · per-cell SESSION-CLOSED-*.md (closure detail) · the roadmap (sequence). This ledger is the backward record; the contracts are the forward plan; Systems v1 is the current synthesis.

Last updated: 2026-07-01 (FAITHFULNESS AXIS CLOSED end-to-end. Recall selector graduated to L5-cosine (G-L5-RECALL-LIVE 86.89% paraphrase; the fact signal is layer-localized in global layer 5). The heavy generative judge is PARKED via the hard-foreign kill-test (0 benefit vs L5-direct+τ; PASSed 15/18). The zero-prior/private-data hole (SNE crucible: 80% confabulation / 5% secret-leak on novel entities) is CLOSED by a deterministic attribute-grounding gate + query-token guard + ZERO-INFERENCE symbolic decline (G-SNE-ATTRGATE-ZEROINF: confab→0, leak→0, recall 100%, paraphrase untouched; the decline streams a fixed string with NO gemma4 forward). Governing law = ADR-002 (Decide→Execute spine); empirical lever/constant map = papers/PPT-LAT-FINDINGS-LEDGER.md. SWARM/DHT (SP-SWARM memory mesh) re-elevated to a PRIMARY forward axis alongside ADR-002 — and now BUILT END-TO-END (L0–L4 GREEN, cross-lang byte-parity, sp-daemon-integrated default-off; see the SP-SWARM PROVEN block below + PPT-LAT-MESH-API.md). Engine fc2e8467daf2fa; lattice 930bfcd→(this commit).).

[PROVEN/WIRED — L0–L4 GREEN, cross-lang byte-parity, integrated default-off] SP-SWARM private memory mesh (2026-07-01; engine 7daf2fa, sp_swarm crate). Replication + discovery over the content-addressed MEM-OKF store — not a new store. Five gated layers, every one reusing a proven asset: L1 content addressing (addr=sha256(norm(body))[:16], plus the C2-SimHash address class for episodes), L2 have/want replication with verify-on-arrival (G-SWARM-REPLICATE-CONVERGE), L3 Ed25519 sign-on-write / verify-vs-roster provenance (audited ed25519-dalek ↔ pynacl parity; G-SWARM-PROVENANCE-ED25519), L0 QUIC transport (quinn/rustls 1.3, reused from network::quic_shard) + Ed25519 mutual roster handshake (G-SWARM-TRANSPORT-QUIC), L4 C2-SimHash discovery gossip (SIM shortlist → exact-fetch verify; G-SWARM-GOSSIP-DISCOVERY). Rust↔Python byte-parity proven (G-SWARM-RUST-PARITY, incl. the CRLF-normalize interop fix). Integrated into sp-daemon behind an optional, default-off swarm feature (SP_SWARM=1 gate; G-SWARM-NODE/G-SWARM-DAEMON-WIRE), plus a standalone sp-swarm-node binary. Honest-negative (G-SWARM-C2-SEMANTIC): C2-256 is a shortlist (recall@5 0.885 == L5-cosine top-1), NOT a top-1 oracle (recall@1 0.607) — used as a hint, confirmed by exact-fetch; bit-count is the lever. Governing law honored: the mesh moves curated records, never raw latent (ADR-002 / TELE-12). Remaining = multi-host deployment only. Call surface + flags: papers/PPT-LAT-MESH-API.md; design + rejected mechanics: papers/PPT-LAT-DESIGN-SWARM-MEMORY-MESH.md.

Last updated: 2026-06-30 (O(1) CONVERSATION KV LIVE — G-PERSIST-KV GREEN, SP_PERSIST_KV default-ON (6-turn byte-identical, TTFT off 7.47× vs on flat); + LATENT INTERCEPTOR + TELEPATHY CAPSTONE CLOSED + PARKED (TELE-14 sovereign native delegate). See the two sections at the foot of this ledger.). Prior: 2026-06-24 (WHOLE-MACHINE DIFFUSION-JUDGE PERF ~2x stacked byte-exact + COLA north-star doc + a self-cond-OOB recall regression IN FLIGHT).

[PROVEN] SP_DG_SCRATCHREUSE diffusion-judge speedup, default-on (engine e31c70d): hoist the per-expert synchronizing cudaMalloc/cudaFree into a reused device pool; reversed 2x2 A/B OFF 281/285s vs ON 193/194s (order-independent) = ~1.46x; byte-identical by construction. [WIRED -- byte-exact, default-OFF, promotion HELD] SP_DG_ASYNC pinned double-buffer prefetch of spillover experts (engine 2a1c830): overlap upload(N+1) with compute(N) via dg_ustream + 4 double-buffered slots + up_ev/cons_ev + host-W-A-R guard; SP_DG_MOECHK determinism oracle 240/240 single-item + tonight's 6-diverse-item concurrency stress 1440/1440 byte-exact OFF==ON; ~2x stacked on scratch-reuse. Root cause of async parity = a PRE-EXISTING dg_self_cond OOB (harness sized self-cond to CL=16 vs forward C=256 -> a dg_k_softmax_rows/V=262144 vocab-softmax over-read of uninitialised memory), compute-sanitizer memcheck-pinpointed + fixed (zero-init dev alloc + full-canvas sizing). [CORRECTED 2026-06-24-later -- the "regression" was a misread, WITHDRAWN] the killed native multi-step recall run scored 5/60 = 8.3%. This is NOT an OOB-fix regression: the self-conditioning is CORRECTLY wired (verified test_diffjudge_denoise.c:437 have_prev gate -> step-0 plain forward, steps 1+ feed prior logits); the 95.6% is the EXTERNAL llama.cpp oracle, our NATIVE judge was always ~25% single-forward (f8f76a5). Honest finding: the "iterative multi-step denoise rescues the native judge" hypothesis is REFUTED (8.3% multi-step is no better than ~25% single-forward). The dg_self_cond OOB fix remains a genuine correctness fix; SP_DG_ASYNC is byte-exact + default-off, NO regression blocks it. [PREFIX-KV OVERTURNED via Cola E1, verified from _diffgemma_reference/diffusion-gemma.cpp:43-54 + ARCH-NOTES.md:40-52]: the reference mask is ASYMMETRIC -- prompt queries are causal-over-prompt and NEVER attend the canvas; only canvas queries are bidirectional. So prompt K/V is canvas-invariant BY CONSTRUCTION and the reference SHIPS a prefix-KV decode variant (llm_graph_input_attn_diffusion_decode, rectangular [P+C,C], cache prompt K/V, forward only canvas). Our 6.9e-4/NaN refutation was FALSE (fp-noise + the now-fixed OOB). prefix-KV is VALID on the current model -- NOT a train-time property, NOT a Cola finetune. prefix-KV RECLAIMED -- answer-lossless (G-DG-PREFIXKV-PARITY GREEN): proof re-run shows the K/V byte-delta is fp non-associativity (mask verified asymmetric cuda_forward.cu:5477-5482), NOT coupling -> byte-delta is the wrong gate; the ANSWER-parity gate is GREEN (SP_DG_PREFIXKV 0 vs 1 = bit-identical picks + ans_tok over 3 items, fast quicker). The fast path already exists behind SP_DG_PREFIXKV (the N6 port). PRODUCTION CONFIRMED (CANVAS=256 STEPS=4): parity HOLDS (base==fast bit-identical picks+ans_tok 3/3) + SPEEDUP ~1.5-1.6x (33-38% faster). prefix-KV = SHIP-IT (lossless + ~1.6x). FULL GATE DONE (G-DG-PREFIXKV-FULL, 140 items STEPS=12): LEG A==LEG B on AGGREGATE (recall 44/90=48.9%, reject 49/50=98.0%) + 1.621x faster, BUT 11/140 per-item picks DIFFER (net-zero; fp-jitter compounds over 12 steps) => prefix-KV is "accuracy-neutral, NOT byte-exact at depth" (weaker than async byte-exactness; shippable for the judge since recall/reject preserved, but default-on is a VALUES call). Native judge full-config = 48.9% recall / 98.0% reject (reject BEATS oracle 96.0%). DEPTH (STEPS=48 partial ~51% ~= STEPS=12 48.9%): depth SATURATES at ~12 steps => the 46pt gap to oracle 95.6% is NOT denoise depth (divergence-ladder #1 REFUTED-weak). LEAD = #2 self-cond divergence (harness feeds MASKED answer-row logits to SC, test_diffjudge_denoise.c:481-489; reference feeds RAW canvas logits). T33 DONE (SC A/B, G-DG-SC-AB): masked-vs-raw self-cond REFUTED (recall flat 55->50%, reject 90->100%, N=30; the masked answer-row SC was benign). BOTH depth (#1) AND self-cond (#2) now refuted -> native diffusion judge PLATEAUS ~50% recall / ~95% reject with these levers. Likely real gap = METHODOLOGY (the oracle REASONS via <|channel>thought + _TAGPOOL before the tag; our native harness hard-constrains the answer row from step 0 = a blind classifier, no reasoning). NEXT = bake-off (T32): the resident 12B GENERATIVE judge is a proven native 85.7% (Phase 4) and sidesteps the streaming/sampler/quant/methodology rabbit hole. [DESIGN] COLA (DESIGN-COLA-DLM-MAPPING.md section 2, corrected): block-causal does NOT map to prefix-KV (the model is already prompt-causal); Cola residual value = latent diffusion / avoid the V=262144 vocab softmax (the exact bug class hit tonight).

Last updated: 2026-06-21 (NIGHTSHIFT OFFLINE CURATOR GREEN + MEM-OKF anti-rebuild store ACTIVE).

[PROVEN/WIRED] Recall-organism architecture, fixed + recorded (2026-06-21): the causal ablation oracle (TAU=−8) is the official ADMISSION gate (teacher-forced knockout: load-bearing novel memory collapses ≪ −8, parametric ≈ 0) — and is the NIGHTSHIFT curator's admit step; the learned latent W_c head (SP_B3_WC, G-CHAT-B3-WC-DEPLOY, engine edc8079) is the live RECALL selector on the served chat (an in-distribution memorizer: 99.7% trained / 28.3% OOD@K8); the native Diffusion Judge WON the OOD kill-test (G-DIFFJUDGE-OOD-H2H, lattice dda7ffa: 94.4% recall / 98.0% reject vs W_c 28.3% / 96.9% @K8, +66pp, clears the pre-registered +10pp/≥96.9% criterion) → it is the zero-shot Stage-2 adjudicator over the W_c/LSH top-K, and N5b (resident reservoir) is JUSTIFIED (94.4% is the external 26B oracle proxy; N5b makes our native judge fast enough to match it). (The first-pass "run INVALID" call was an analyst error — grepped the harness 40-char truncated result-line, not the full reply; corrected + receipted.) [WIRED, gated-GREEN on SYNTHETIC] NIGHTSHIFT offline curator (run_kairos_curator, engine 9ad7ede9ee46686107f3e, default-off SP_NIGHTSHIFT_OFFLINE): 12B model-call ep.secret extractor → ablation admit → conformant MEM-OKF emit. G-NIGHTSHIFT-CURATOR criteria 1-4 GREEN (novel "8-FALCON-7729" collapse −33.59 ACCEPT / parametric "Paris" 0.00 REJECT, ~33-nat sep; emit rc=0, addr-join verified) — criterion 5 (B4 distributional/provenance fix) CLOSED live (engine 3ccba61, receipt G-CHAT-B4-NIGHTSHIFT-provenance.log): the documented live-0.084-vs-curated-9.858 collapse is GONE — a 2-token BOS+trailing-newline tokenization mismatch (+ a process-static write-once guard) fixed → the live episode is byte-compatible with the curated ep.k → identical text scores 9.858 == 9.858, a novel fact scores in-band (6.295) with clean foreign-reject (−15.498). End-to-end machinery correct (capture → byte-exact provenance → in-band scoring → foreign-reject). Residual (pre-scoped, NOT a miss): a novel instance doesn't out-rank the 90 curated needles under the closed-set W_c head — now superseded on the hot path by the general L5-cosine selector; the one follow-on = gate L5 recall on a nightshift-captured episode. Record papers/CONTRACT-NIGHTSHIFT-CURATOR.md §7. [ACTIVE] MEM-OKF = the content-addressed LUT→summary→full anti-rebuild store (tools/okf_mem.py + memory-okf/, spec MEMORY-OKF-PROFILE.md); the lookup-before-you-build pre-flight is now binding in prompt.md §0/§8 + CLAUDE.md.

Prior: 2026-06-20 (B3-WC AUTONOMOUS RECALL CAMPAIGN RESOLVED — a learned W_c head does LIVE instance-level episodic recall on the served gemma-4-12B chat, with clean foreign-reject; engine edc8079). XBAR UNIFIED onto the exact-integer O_K substrate + the BYTE-EXACT FORWARD CLOSED GREEN on the real gemma-4-12B (engine 0019b86→d2d7ceb, byte-exact 69c0588 + submodule d9d96f3).


0. The frame (so the record is read correctly)

PPT-ARM is the primary, load-bearing product (the 13-step transformer-forward replacement + the Spinor KV / two-ring memory architecture). The Lattice fell out of it. The value is the envelope — inline KV compression → unlimited context, Ring-2 offload + residual/CRT bandwidth bypass → multi-device, speed on integer pipes — with bit-exact as the invariant floor, not the headline. See RFC-001.

Crucial gating rule (operator, 2026-06-02): the system does not work in isolation. A stage will be slow / miss a system-level number that it only hits once the rest of the envelope is in place (e.g. tok/s is not achievable until the Spinor cache + Ring-2 + island sharding are wired). Therefore a stage is gated on ITS OWN correctness/metric (bit-exact output; the kernel's own throughput; the compressor's own ratio), NEVER on end-system tok/s. Do not declare a stage failed because the assembled-system number isn't there yet — that number is a system gate, measured only when the envelope is assembled. Penalizing a stage for a system number it structurally cannot hit alone is a category error.


1. PROVEN — PPT forward (the bolt-on works on real models)

The PPT discrete forward reproduces stock models argmax bit-exact to llama.cpp. This is the precondition that licences compression/context/speed on top.

Model Evidence Reverify Status
Qwen3-0.6B core E_CPU_2 / forward gates ctest -R E_CPU_2 [PROVEN]
Qwen2.5 qwen25_forward gates core suite [PROVEN]
Gemma3-1B M_GEMMA3_CPU + T_FRO_4 PPL ctest -R T_FRO_4 [PROVEN]
Gemma4-E2B M_GEMMA4 PPL gate, top-1 bit-exact oracle ctest -R M_GEMMA4 (engine) [PROVEN] this session
Qwen3.6-35B-A3B (qwen35moe) M_QWEN36 top-1 bit-exact (3/3 5444 8 198), 218s ctest -R M_QWEN36 (core) [PROVEN] this session

qwen35moe is a Gated DeltaNet (Qwen3-Next) + 256-expert MoE + IMRoPE hybrid, NOT Mamba2 (the old GGUF-INVEST doc mislabeled it — superseded). Full per-block validation: GDN recurrence, MoE router/experts/shared, gated full-attn — all matched oracle fingerprints. Closure: SESSION-CLOSED-lat-3-moe-forward.md. Commits: core 568b678→d8e614f.


2. PROVEN — math-core primitives (the substrate)

Primitive Evidence Status
Barrett mod-mul (+ nvcc paired-register fix) NTT/mod_q tests; HVX K.beta.2.5b [PROVEN]
Frobenius lift Q4/Q8 + arena packed weights (zero-inflation) E_CPU_*, arena tests [PROVEN]
Q4_K + Q6_K k-quant dequant (ggml-exact) weight_dtype; qwen35moe conv fingerprint matched core 25809d8
NTT-CRT host, dual-prime, byte-exact (negacyclic, N≤512) ntt_crt tests [PROVEN]
NTT N≤512 frozen-prime cap (2N|q−1) + Bluestein arbitrary-N≤512 NTT.0–5 closures [PROVEN constraint]
KSTE encoder + Friedman sieve E_CPU_6, sieve tests [PROVEN — re-gated fresh 2026-07-01, system tests/fixtures/G-KSTE-SIEVE-REVERIFY-2026-07-01.log @ 15698b0: T_KSTE 22/22 + T_SIEVE/T_POUW 37/37, gcc -O2. core primitives. The archived Paper III §11.6 application numbers were then re-derived fresh (system tests/fixtures/G-SIEVE-MEASURE-2026-07-01.log, harness tests/sieve_measure.c): termination + speed CONFIRMED (Dickson wqo saturates; raw dominance 16.77 ns/pair; per-cand p99 0.063 µs; 50 µs gate clear ~800×) — but discrimination is an HONEST-NEGATIVE on the in-tree v1 encoder (intra/inter dedup ratio only 1.0–1.3× vs the paper's 17×; plateau 2 slots vs ~307). The rich discrimination/plateau belong to the fuller 60-node encoder in the anti-contaminate-gated OLD repos, NOT this fixed-13-node "Phase 1" core]
KSTE v2 — magnitude-as-depth encoder + Dickson σ0⊕σ1 dominance (core/kste_md/) T_KMD 11/11; G-KSTE-MD-2026-07-01.log [PROVEN 2026-07-01 — the discrimination win, re-derived fresh from Paper III/IV spec, anti-contamination-clean] Built as a NEW module (frozen v1 untouched). Head-to-head vs v1 on the cluster probe: at near-duplicate cos≈1.0 (σ=0.005) intra/inter dedup ratio 37.6× (vs v1 1.2×, vs paper ~17×), 8.4× @ cos 0.9998, degrading gracefully with noise; rich frontier plateau 76 (vs v1's degenerate 2); Dickson check 16.66 ns/pair. The count-based signature (sign-split nB/nC + chain-shape MBB/MCC) moves independently → real semantic discrimination where v1's magnitude-correlated order-stats collapsed. Scope: CPU-core encoder, NOT yet wired to the served daemon / PoUW sieve (frozen v1 still there); encode is qsort-bound (9.1 µs/vec). INPUT-GATED (honest negative, G-KSTE-MD-REALDATA-2026-07-01.log): on REAL last-token global-Q the 37.6× collapses to 1.00× / recall@1 = random floor — because the data itself barely separates (same-fact vs diff-fact cosine gap 0.038; even dense cosine only 11.5%). It amplifies whatever separability the input HAS; last-token global-Q has ~none. So it does NOT go on the global-Q memory/recall/dedup path; its only real home is content-bearing vectors where cosine already separates (pooled embedding, TELE-1) — as a cheap exact-integer discrete dedup key, not a magic discriminator.
★ L5 RECALL FINDING — cracks the LN-1 paraphrase wall system tests/fixtures/G-REP-LAYER-L5-2026-07-01.log (harnesses tests/rep_sweep.py,rep_layer.py,rep_hybrid.py,kste_md_confirm.c) [PROVEN 2026-07-01 — CORRECTS the LN-1 "last-token global-Q is content-poor" conclusion] The fact signal is NOT absent — it is layer-localized in global layer 5, and averaging all 8 global layers destroyed it. Per-layer exact→paraphrase recall@1: L5 85.2% / L6 78.7% / all-layer-avg 11.5% (greedy subset picks L5 alone). Deployable: L5 (mean-heads, 512-d) query-key recall via cosine = 100% exact / 88.5% paraphrase on the 61 fact-conflicts, vs deployed Jaccard 100%/8.2% — the paraphrase wall that ended the faithfulness arc. Combine with Jaccard by GATE/max (naive z-sum hurts: 49%). Data is ALREADY captured (read_global_q emits all 8 global layers; selector indexes L5). Sub-findings: KSTE-MD is the WRONG tool here (signal is directional/angular; kste_md is magnitude-shape-blind, 3.3% vs cosine 85%); Q-vs-K cosine is floor (different subspaces → needs W_c-fed-L5 if matching stored K). Honest: n=61 small, needs scale confirm; verify live layer-index ordering == offline L5 before wiring. WIRED + GATED LIVE 2026-07-01 (GREEN-LIVE): SP_RECALL_L5 on the served gemma-4-12B (branch feat/l5-recall, release build 0 errors, ep.l5 query-keys per episode) → G-L5-RECALL-LIVE-2026-07-01.log: paraphrase OBEY = 53/61 = 86.89% (vs deployed Jaccard ~8.2%; matches offline 88.5% within noise); exact-query alignment perfect (cos=1.000). Default-off = byte-identical null floor. The 7 leaks are hard parametric conflicts (Sun-is-a-star etc.). Combiner (Jaccard+L5+KSTE-MD) does NOT beat L5 alone (redundant/dead); its value would be foreign-reject, untested.
Spinor block: KV encode/decode + 64-byte receipt ABI vht2 / spinor tests; silicon-confirmed [PROVEN]
Garner 2-prime recombination constants ntt_crt.c [PROVEN]
OK_Q4 reducing codecSP_DT_OK_Q4=11, transcoder add_q4/use_q4, Frobenius Q4 pack/unpack, arena Q4 path (row_prec[] 8/4 + q4_unpack) gate E_PARITY_2 [PROVEN/WIRED]
C1 — qwen35moe .sp-model reducing + output-lossless production path — transcode OK_Q4 -> sp_model_load -> sp_model_to_qwen36 (the loader/SWIVEL) -> qwen36_forward. Rank-3 experts via arena (build_packed_q4 rank-3 fix); F32 router via sp_as_f32; arena-aware expert_mm. Measured: 16.33 GB vs 19.7 GB Q4_K_M source (~17% reduction); round-trip top-1 = 5444 == oracle (OUTPUT-LOSSLESS). core 66ccab9; qwen36_spmodel_top1 [PROVEN] this session — first real envelope number

3. PROVEN-PER-RECORD — Hexagon / backend / daemon track

(Reported via the operator's session logs + memory; in sibling repos — engine, sp_compute_skel, sp_daemon. Trust the tags/closures; do not re-run unless touching that code.)

Item Evidence Status
Hexagon HVX: Barrett, mod_q matmul, NTT.0–5c, Bluestein tags lat-phase-2-... ; closures [PROVEN-per-record]
HX.3b: vrmpy int8 forward = 1.04× faster than ARM fp32, byte-equal tag lat-phase-2-hx-3b-hvx-vectorized; CLOSURE-HX-3b.md [PROVEN-per-record] — the integer-substrate-matches-fp32 proof point
Cross-backend determinism (ARM ↔ cDSP scalar) WIRE-HEX-FINISH [PROVEN-per-record]
QNN HTP NPU dispatch in Unsigned PD (1.329 ms execute, 64/64 byte-exact) K.2-spike closure [PROVEN-per-record]
Mode-D FastRPC bridge (S22U), concurrent dispatch (Arc) Sprint K closures [PROVEN-per-record]
L3 daemon: chat + dialogue + ledger + QUIC peer wire daemon closures [PROVEN-per-record]
M.4 PoUW ledger + mesh canonical Garner order M.4 closure [PROVEN-per-record]
Trick #1 substrate (dual HVX vector contexts 1.935×) Sprint K v0.alpha [PROVEN-per-record]

Honest negatives also proven (do not re-litigate):

  • NTT-attention is slower than fp32 dot at HD ≤ 256 (~0.15–0.72×). The substrate win is over HD (poly length), not ctx. Speed comes from compression + bandwidth-bypass + integer pipes + multi-device, NOT from NTT. [PROVEN-per-record: NTT.6]
  • HX.3b chat-shape inner loop is memory-bandwidth-bound, not ALU-bound → v2 wsum-precompute only got 1.065× (gate was 1.20×, FAILED honestly). Attack bandwidth (prefetch/VTCM/2-row), not ALU. [PROVEN-per-record]

4. The ARM memory architecture — regime split (System-1 / System-2)

Prior proven design (old SP, to be re-established here): the KV/memory path is regime-adaptive, not one-size:

  • System-1 (small context): a fast simple path — keep latency low where compression overhead wouldn't pay. (Old SP ran small-ctx differently to hold speed.)
  • System-2 (large context): the Spinor-compressed + Ring-2-offload path — where the 120× and unlimited-context envelope lives.
  • A crossover oracle predicts when to switch System-1 → System-2 (by ctx length / bandwidth pressure / cache occupancy).

[DESIGN here / PROVEN-pattern in old SP]. This is why §0's gating rule matters: System-2 looks slow at small ctx in isolation — it's not meant to run there. Belongs in contract C2 (ARM memory) with the oracle spec'd explicitly.

[MEASURED 2026-06-02, harness tests/c2_sparse_recall.c, contract C2.0.4/C2.0.5] Sparse-recall fidelity vs full attention (needle-in-haystack, N=4096): ORACLE top-B reproduces full attention at B=64 (cosine 1.0, 8/8 needles) — so Ring-2 storage does become usable context IF the recall router is good. Query-agnostic SWA/φ capture the recent cluster but miss distant needles (0–1/8); pure Fibonacci-φ is the worst router (φ is for eviction coverage, not peaked-mass recall). KSTE-signature recall looked best (4–6/8) but ADVERSARIAL test (tests/c2_kste_router_adv.c) FALSIFIES it: permuted decoys (same histogram, ~0 dot product) get KSTE tier-0 distance-to-q 937 vs needles' 985 — INDISTINGUISHABLE; the router pulls 21/32 zero-score decoys. KSTE order-statistics are a histogram (permutation-invariant), dot-product is directional → KSTE is structurally NOT a recall router (the 6/8 was an artifact of needles being the only histogram-distinctive vectors). Recall router MUST be a cheap directional score (low-rank/projected dot-product or NTT coarse pre-score); KSTE ruled out (stays valid for dedup/dominance only). Caught by adversarial verification, not theory. System-1/2 oracle DERIVED: switch at Ncrit = min(RAMbudget/pt_f32, ~(W+B)) ≈ 1–4 k tokens (forced by RAM at tens of k); System-2 quality is bounded by router fidelity → oracle exposes a quality floor (widen B / denser fallback). The recall-router fidelity gap is the open C2 item. → SOLVED 2026-06-02 (C2.0.6, tests/c2_router_proj.c, core 186aadb): a ±1 Rademacher random projection (rank-16 = 32 B/token, SMALLER than KSTE 64B) is ORACLE-PERFECT — 8/8 needles, cosine 1.0000, 0 decoys at B=64, where KSTE got 0/8. JL preserves the dot; ±1 keeps it INTEGER/Z_q-native (Lattice-pure, not "dirty float"). Router=±1 projection sidecar, Compressor=Spinor/Tail-Slayer, cleanly separated. Path B (NTT low-freq) unnecessary. Unblocks System-2 usable context.


5. TARGET — the envelope (the value; measure, don't assume)

These justify the project and are not yet measured here. They are the point. Measuring them is the next phase's job (contracts C1/C2).

Target Where measured Status
KV compression — TWO overlays (per PPT-LAT-Theory.md §3–4, T2, T8.2) C2 [MEASURED 2026-06-02, gate C2_KV_RATIO, harness tests/c2_kv_measure.c] (a) sp_spinor_encode_vec (faithful, dimension-preserving, 63 B/block, NBLK=⌈HD/55⌉): asymptotes ~3.5×/f32 but only ~1.0–1.7×/f16, at cosine ≥ 0.99996 (top-1-safe pending real-model confirm). NOT 120× — it is 1 int8/elem + per-block scale (candidate "anchor-basis reconstructs HD≫55" FALSIFIED by linear NBLK scaling). (b) sp_kste_encode (lossy ⪯_d signature, 64 B regardless of HD): ratio = HD/16 → the ~130× headline = KSTE at HD≈2048; discrimination ~1.0 even at scale 65536 (the M.5 i16-clamp does NOT collapse continuous-K discrimination — only token-ID). It is a dedup/routing signature, NOT reconstructable attention KV. Decisive: neither overlay alone is "120× reconstructable KV" → the headline = faithful ~3.5× × Ring-2 effective-context multiplier. Ring-2 recall (C2_RING2_RECALL) is now the load-bearing unmeasured piece, not the per-vector block. REAL-MODEL CONFIRM (C2_KV_DECODE_DETERMINISM, live E_CPU_8 test_kv_spinor Qwen3-0.6B scalar, 2026-06-02): Spinor-KV vs f32-KV argmax 29/31, KL mean 2.300e-02 (gate ≤2.0e-1) — PASS as a BOUNDED-DIVERGENCE overlay. HONEST: Spinor-KV is LOSSY (~6.5% argmax flips over 28 layers), NOT bit-exact; the per-vector cosine 0.99996 did NOT carry to 31/31. Bit-exact floor = weight path + gate-OFF, not the Spinor-KV overlay. Gate-OFF == f32 forward bit-identical still holds.
.sp-model converter REDUCTION ratio (≤ source; sub-Q4) C1 [TARGET]
tok/s vs the bar: beat llama.cpp + old SP hier-KV @ 40 tok/s, Qwen3.6 P1 SPEED contract; system gate [TARGET] — REFERENCE MEASURED 2026-06-02: llama.cpp Qwen3-0.6B-f16 CPU greedy = 210 t/s prompt / 28.2 t/s gen (dev host i9-11900KB). f16 decode already bandwidth-bound at 28 t/s → confirms the SP lever is reduced weight-read traffic (packed Q8/Q4), not ALU. SP-side MEASURED 2026-06-02 (sp_toks, after segfault fix 0fb39ab): f16 0.84 → Q8 arena 1.58 (1.88×, bandwidth lever) → Q8+threaded matmul 10.53 (8975753) → +threaded attn/per-head 12.55 (d7735a4) → +AVX2 int8×f32 dot 39.52 (5e443c9, 3.15×) = 47× over the 0.84 f16 baseline. FAIR quant-matched scoreboard: SP-Q8 39.52 vs llama.cpp-Q8_0 52.8 → SP ~0.75× (llama.cpp ~1.34× faster); SP-Q8 beats llama.cpp-f16 (28.2) but that's not apples-to-apples. From ~33× behind to ~1.34× behind on the fair fight — competitive, not yet winning. VNNI int8×int8 TESTED (gated SP_VNNI=1, engine a2ad1dc) — DOCUMENTED NEGATIVE: only +9% (43.95 vs 40.38) AND top-1 gate FAILS (divergent tokens). Falsifies the "ALU gap" hypothesis: Q8 decode is BANDWIDTH-bound (VNNI reads the same int8 weight bytes as AVX2; 4× ALU barely helps), and naive per-vector int8 act-quant is too lossy (needs per-channel/SmoothQuant). AVX2-f32 dot (40 t/s, accurate, parity-safe) stays the production CPU kernel. The ~1.34× gap to llama.cpp-Q8 is memory layout/bandwidth (Q8_0 32-elem blocks / fewer passes), NOT ALU — that's the real follow-up. Threading + AVX2 parity-safe (oracle=SP_CPU_SCALAR). The speed thesis is validated: SP within ~2.7× of llama.cpp on 0.6B from packed-Q8 + threading, no SIMD yet. See CONTRACT-SPEED SPEED_WIRE_CPU ladder.
Ring-2 disk offload + recall cost + effective-context multiplier C2/C3 [MEASURED 2026-06-02, gate C2_RING2_RECALL PASS, harness tests/c2_ring2_measure.c] spill→recall byte-identical (2000/2000); recall ~10 µs/token + ~2 µs/block decode (page-cached) ≪ recompute (no weights/matmul). Effective-context multiplier = (RAM_window+Optane)/RAM_window: ~400× @16 GB, ~794× @32 GB, ~1190× @48 GB at a 512-token window → the "~120×/unlimited context" headline is CONSERVATIVE and lives HERE (candidate #2), not in the per-vector codec. Scope limit: solves the memory wall, NOT the compute wall — usable context needs a sparse/recalled-attention pattern (SWA / Fibonacci φ sub-sampling / retrieval); re-measure on real Optane.
Dual-GPU / multi-device residue-sharing (ship residues, not tensors) C3 + Trick #1 [RESOLVED — HONEST NEGATIVE, do NOT re-offer. engine tests/perf/SESSION-PERF-SYNTHESIS-2026-06-23.md §1.2/§2, engine 1d0e414] The "ship residues, not tensors" pitch is FALSIFIED on physics: dual-prime CRT shards NUMBERS, not experts — each prime channel computes the ENTIRE network and needs ALL the weights, so assigning primes to devices cuts per-device memory by zero. Real device-sharding needs full activation exchange (not tiny residue arrays). CRT's genuine gift to a heterogeneous split is BIT-EXACTNESS (no cross-device float drift) — the enabler of a clean split, not the split mechanism. The iGPU here (Intel UHD, Xe-LP, 32 EU, no XMX, ~0.75 TFLOPS, shares DRAM) is a phase-3-maybe expert-offload target at best, never a CRT island. Keep CRT for what it's for: exact arithmetic / auditability.
int-end-to-end (no per-matmul fp dequant; only logits) WIRE-* + C2 [DESIGN]

5.05 C2.1 COMPLETE (2026-06-03) — two-ring recall wired live, all three walls down

C2.1 wired the C2 measurements into the live qwen3_generate_kv decode path and drove the three walls down, each gated on an N=512 NIAH parity gauntlet (GEN_KV bit-parity + real 837492-needle HIT). Engine commits 67f4997f8ea920; full record in CONTRACT-C2 §C2.1.

  • Router (Step 1, 67f4997): ±1 Rademacher projection sidecar, recall-set = sinks ∪ top-(B−W−sink) ∪ recent-W; parity-exact when off / B≥ctx.
  • G1 NIAH (7055964, tests/niah.c): decode-path needle gate. r=32 holds 2×/4×/ at N=2k; depth 10/50/90 all HIT (no recency bias). Budget B is absolute → achievable ratio grows with context.
  • G2 PPL (d56c1a7+e916365): autoregressive decode-path PPL. v1 FAILED (4× +40%, 8× +104% — dropped softmax tail + attention sinks). Fix = Möbius-pinned sinks (SP_RECALL_SINK=4). v2 N=2k: 2× −0.71%, 4× −0.92%, 8× +0.69% — all <2%. Intelligence wall solved @8×.
  • Step 2b Optane (2707f60/fdc0f07/e895ef4, ring2_disk.c): NO_BUFFERING + IOCP async. Latency v0 48.7 → v1a 18.9 (dedupe) → v1b 7.57 µs/read (≈ media floor). NIAH HIT off F: Optane.
  • Compute wall (b7a1f92): O(B·N) max-extract → O(N) quickselect (the ~10-h-at-32k bottleneck removed).
  • Memory wall (f8ea920): Ring-1 kc/vc(sink+W) ring buffer when offloading. N=512 15× shrink; 32k = 910× (7.5 GB → 8.3 MB).

Honest RAM floor @32k: Ring-1 = 8.3 MB but the projk ±1 router index stays full-P ≈ 940 MB → net 7.5 GB → ~950 MB (~8×), projk-dominated; int8/int4 router-index quant is the next optimization (owned, not hidden).

5.055 C2.1 prefill/decode modes + fusion + release prep (2026-06-03)

Three decode modes now exist, all parity-exact when off, additive/gated (proven paths untouched):

  • Streaming (default): recall during prefill, always-low-RAM (Ring-1 = sink+W throughout). O(B·N) prefill. This is the 32k headline path.
  • Decode-only (SP_RECALL_DECODE_ONLY, engine a5e9b86): dense-exact prefill in RAM, recall engages only at decode (pos ≥ n_prompt). Trades peak-RAM-during-ingest for fast exact ingest.
  • Fusion / compact-and-spill (SP_RECALL_FUSE, engine 7896bc4): dense-exact prefill in a full-P RAM buffer, then ONE bulk spill of the cold tail to Optane at the prefill→decode boundary + copy sinks/window into the (sink+W) cache + free the prefill buffer → window-sized decode RAM with exact ingest. Verified N=512 (boundary fired, needle off disk) and timed N=8192 (51.4 min wall, 1.88 GB buffer freed, HIT 837492, 10.62 µs/read). Upgrades paper-01 §3.7 future-work → result.

Honest cost note: fusion prefill is exact O(N²) attention (recall off during ingest), so 32k dense-exact is ~18 h on one f16 core — the stock cost of exact attention, not a fusion defect, consistent with the ~1.34× throughput gap. The 32k headline therefore runs on the streaming path (always-low-RAM, O(B·N)); fusion's receipt is the 512 + 8k runs. R9 (streaming 32k) in flight — drop its retrieval/read-count/latency/wall-clock into paper-01 §4 + abstract + EXPECTED.md + landing hero on completion.

Release prep (publishing track): paper-02 repro green — 6/6 E_FMT gates, L1 reducing (Qwen3-0.6B-f16 1,439.4 → 719.6 MB, 50.0%), L4 bit-faithful forward on gemma-3 + qwen3, captured in EXPECTED.md. License = MIT (wired through LICENSE / CITATION.cff / README / site). Release staging assembled in comms/release/ (papers 01–02, site, ledger); shannon-prime-papers repo set up for the papers series. Public remote URL + first release tag deferred (operator deciding).

5.06 C2.2 COMPLETE (2026-06-04) — canonical two-ring + NTT fusion + the network tier

The day the discrete object closed at every scale. Full record in CONTRACT-C2 §C2.2; math-core 9c26475→54ee28b, engine 005473d→57c9a53, suite 21/21.

  • [PROVEN] ARM in math-core. core/arm/ + the abstract Ring-2 backend in the L1 ABI (sp_arm_ring2_register; registered = borrowed). The full two-ring decode lives in core/forward/decode.cthe only decode in the tree (engine duplicate ~430 lines incinerated; engine resolves the canonical decode at engine speed via the cpu_overlay.c dispatch seam: 22.62 tok/s). Run-gates T_ARM_GENKV (11 gates incl. counting-backend registration proof). Reverify: ctest -R T_ARM.
  • [PROVEN] Dual-prime NTT keystore fusion. SP_NTT_KV: K cached write-once as the dual-prime residue block; score = residue dot + Garner = exact ⟨q,k⟩, no inverse butterflies. Bluestein keystore (empirical coefficient-0 weights folded into the key) covers every pow-2 HD ≤ 256 — the HD=8 fixture runs fusion natively in-tree. sp_pr_resdot deferred reduction (15 u64 products per mod) + engine AVX2 Barrett-SIMD override: fusion 15.68 → 18.35 tok/s (84% of the 22.6 f32 baseline), sequences bit-identical. Residual −16% = per-token q-transforms (448 fwd NTT pairs + 224 K-encodes) — the named next lever.
  • [PROVEN] Optane tier, dual block size. Two NO_BUFFERING+IOCP stores (8192 B K-residue / 4096 B V-f32) registered through the L1 hook. Live F: run: 16.25 tok/s, sequence identical — inherits the C2.1 7.57 µs/read queue-depth floor.
  • [PROVEN] Network tier — Trick #8 closed. QUIC peer as sp_arm_ring2_backend (M_NET_RING2), then the two-process showpiece (sp_ring2_showpiece + SP_RING2_SERVE): 20,160 KB of raw untranslated u32 residue payload (840 writes + 2520 reads, 504 batched flights) over 127.0.0.1 mid-decode, zero serialization/fp translation, SEQUENCE IDENTICAL to baseline (1.46 vs 1.48 tok/s — transport absorbed by compute). Caveat: loopback; cross-host pending.
  • [PROVEN, honest-negative corrected] MTP T8 (CONTRACT-C4-C5-C6): KV-reuse verify machinery bit-identical; the 1.76× was a degenerate-prompt artifact — real prose/code prompts 0.87× with prompt-lookup drafting; needs a real draft source. The rollback substrate stands.

Composed claim, now run-gated end-to-end: compute operand = cache line (Ring-1, 910× @32k) = disk block (Optane, 7.57 µs floor) = wire packet (QUIC loopback, 20.16 MB) — one dual-prime residue object, byte-exact at every boundary.

  • [PROVEN, regime-bounded] Bit-packed popcount router (C2.3, same day). SP_RECALL_BITS: projk sidecar → one u64 of projection signs per (pos,kvh) (~940 MB r=32-float → ~59 MB @32k, 16×; 32× vs the r=64-float equivalent — an earlier revision said "~29 MB/32×", corrected: 940/16 = 58.7 MB; the decode banner prints actuals), scoring = popcount XOR. NIAH gate: r=32 passes 2× all depths (d50 answer identical to f32) but MISSES 4×/d90 where f32 HITs — honest SimHash resolution loss; r=64 (same 8 bytes) restores full fidelity 6/6. Production: SP_RECALL_BITS=1 SP_RECALL_R=64. Named remaining gate before default-on: PPL deflection at bits-r=64. CONTRACT-C2 §C2.3; math-core 92c07fe, engine 3d2d2c3.
  • [PROVEN] q-transform SoA head-batch (same day). sp_ntt_fwd_batch seam + AVX2 lanes=heads override: fusion 18.35 → 22.3 tok/s (gap to f32 16% → ~7%), fusion×rings×backend 16.08 → 20.14, sequences identical; T_PR_BATCH bit-exact. math-core f7b9b6d, engine 144d445. The NTT compute-optimization arc is CLOSED — residual ~7% buys the whole discrete envelope.

5.07 THE STAGE TAXONOMY (operator-minted 2026-06-04) — the heterogeneous deployment ladder

Canonical names for the physical deployment tiers. Each stage is a strict optimization target for the compiler/router/orchestrator; a stage is claimed only when its composed run-gate exists (the Alpha discipline).

Stage Substrate Status
Alpha CPU / RAM / Optane (Beast Canyon) PROVEN-in-parts; composed 32k finale COMPLETED 2026-06-06 — NIAH verdict MISS (infrastructure proven at 16.3 h scale; retrieval quality at 64× budget not — see §5.11 + CONTRACT-C2 §C2.4-CLOSURE) — context machinery cache-adjacent (~82 MB @32k), Optane = active memory tier, weights remain the DDR bandwidth budget
Beta RTX 2060 12GB (pure VRAM) NEXT after Alpha closeout. SoA lanes→warps; model+context <6% of VRAM; sm_75 constraints pinned: no cp.async/ldmatrix/mbarrier, single INT32 ALU port (see reference-cuda-sm-feature-tiers), and mul.wide/mad.wide.u32 are BANNED (nvcc paired-register miscompile, reference-nvcc-paired-register-bug — decompose to mul.lo/mul.hi + add.cc, anchor ptx_ntt.cuh)
Gamma RTX 2060 + Optane beyond-VRAM contexts/models: cold tail DMA'd across PCIe; Optane stays load-bearing for MUCH larger models even with the GPU present
Delta CPU + RAM + Optane + RTX 2060 as ONE engine asymmetric split: CPU owns the O(N) popcount routing scan + Optane I/O; GPU owns the O(B) residue inner products; pinned-memory (cudaHostAlloc) DMA keeps the QUIC zero-serialization packet intact across PCIe
Epsilon Snapdragon DSP/SVM/ARM/ISP/NPU/UFS (S22U) UMA frontier — groundwork live (Mode-D FastRPC, V69 HVX, QNN NPU Unsigned PD, Tricks #1-#10)
Zeta Alpha–Delta + Epsilon as one system the QUIC mesh as nervous system: the byte on the Optane platter == the byte scored in VTCM, no translation anywhere
Eta Gemma 4 — the Native Sensory Lattice (encoder-free multimodal ingest) RUNWAY CLEAR: gemma-4-12b-it-Q4_K_M.gguf on disk; ingest = ONE OK_Q4 matmul per modality (48×48 patches / 640-float 40ms audio frames), so pixels and sound enter the discrete pipeline at the sensor boundary; same 3-G4 port spine (oracle → bridge → transcode → top-1/PPL); MTP drafters = standalone checkpoints (the T8 draft source); embedding-layer audio round-trip (overcomplete 640→E injective projection ⇒ exact-then-pseudoinverse recovery) gets an SNR gate — "the cache is the audio file," claimed only at the embedding layer, never the K layer
Omicron ο Intel GNA 2.0 — the small-o coprocessor (NUC11 on-die) PROBE QUEUED: the always-on milliwatt integer affine engine (int16/int8 MAC, int32 accum, EXACT in range, 64-byte-aligned DMA — already our ABI). Named for little-o: the lower-order term that never dominates but never sleeps, and Ptolemy's ο-as-zero — the placeholder that holds the cell while the system idles. Ladder of ambition (each gated by Stage-0 read of the archived LGPL kernel source + die probe, NEVER by "designed for" copy): (1) wake-gate VAD, (2) the entire Gemma-4 audio embedder as a native affine layer, (3) the ±1 Rademacher router projection (SimHash minted off-core), (speculative) small-prime CRT-NTT in int16 lanes
Kairos καιρός the sp-kernel — escape from turn-based execution (the time/agency axis) [DESIGN — registered 2026-06-10; opens AFTER P2.b/P3 close] hierarchical tick (GNA→router→Exec), latent interrupts (the X-R1 mechanism as delivery path), eos→yield + gated NIGHTSHIFT idle loop, drivers (ears/HA/TTS), registry+permissions, receipted flywheel. Own doc set: ROADMAP-KAIROS.md + CONTRACT-KAIROS-*.md (kept separate to not pollute the live campaign). Reference corpus: CosySim/NEXUS/Project X (KAI-0 extraction first-pass done)
Holon ⬢⃝ the bonded whole the Universal Discrete Architecture, one Garner formula from L1 cache line to QUIC packet (Trick #8 closed the wire; #9 the ABI; #10 the receipts). The space/distribution axis — composes with (does not absorb) Kairos

5.08 STAGE ALPHA CLOSED + STAGE BETA OPENED ON THE GPU (2026-06-05/06)

STAGE ALPHA (CPU/RAM/Optane) — the C2 envelope is closed-in-parts; the composed 32k gate is deliberately deferred behind the amplification fix. Full detail in CONTRACT-C2 §C2.2/C2.3/C2.4 + CONTRACT-SPEED. The day's arc:

  • C2.2 canonical two-ring in math-core + dual-prime NTT fusion (18.35 tok/s, 84% of f32, bit-identical) + Optane dual-size + QUIC peer + two-process showpiece (20.16 MB raw residues over loopback, sequence identical). math-core 9c26475→54ee28b, engine 005473d→57c9a53.
  • q-transform SoA head-batch (math-core f7b9b6d, engine 144d445): fusion 18.35→22.3 tok/s, gap to f32 16%→7%; T_PR_BATCH bit-exact. NTT compute-optimization arc CLOSED.
  • C2.3 bit-packed popcount router (math-core 92c07fe, engine 3d2d2c3): projk → 1 u64/(pos,kvh) (16× @r32). Both named gates GREEN: NIAH 6/6 + PPL transparent (−0.97%/−0.12%) at ≤4×; FAILS 8× (+6.08% vs f32 +0.69%). Production: SP_RECALL_BITS=1 R=64 ≤4×.
  • Amplification bundle (2441e0b/9cd502f + 200c0ec/2484650/16e15e3): KVSEL group-centroid kv-head selection (NIAH 3/3 @4× incl d90; PPL −0.92%, beats per-Q-head), split-device Optane (K→F: CPU-slot / V→E: PCH), read_batch2 device-overlap ABI (the serialization-tax fix: serial split 4u/S WORSE than single-device 3u/S; overlapped 2u/S), bounded LRU temporal staging cache (T_CACHE_EXACT bit-identical).
  • THE TEMPORAL-LOCALITY FINDING (novel): adjacent decode steps' recall sets DRIFT slowly — measured ~9.46 TB Optane reads to serve ~3 GB unique blocks at 32k; the 2 GB LRU cache absorbs 86%→42% as the reuse-window union outgrows it with depth (full curve mapped). The router's working set glides; the cache surfs it. CONTRACT-SPEED.
  • C2.4 composed 32k finale: TERMINATED at 44.7%, gate PENDING — NOT claimed. Four versions chased four real walls (v2 scan, v3 split, v4 overlap, v5 cache); each shipped+gated, but the single composed log was never banked. DECISION (operator, economic): v5 is terminal; the 4 GB-slab scaling is documented-not-run; no more finale relaunches. Partial run proved indestructibility (8.6 h saturated, RAM flat, zero leak). The composed gate closes when re-run cheaply post-amplification — not load-bearing for the rest of the project. (Outcome: v5 completed 2026-06-06 — verdict MISS; see §5.11.)

STAGE BETA (RTX 2060 12GB, Turing sm_75) — OPENED + GPU GENERATION LIVE (2026-06-06). Full detail in Roadmap §21 + SESSION-CLOSED-stage-beta-s0.md.

  • Stage 0 verified ON THE CARD: CUDA 13.2 still targets compute_75; build-cuda clean (48/48). CUDA_SMOKE + E_CU_5 NTT-attention (KL 2.4e-10) + E_CU_6 KSTE all PASS. The discrete dual-prime poly-ring attention reproduces math-core scalar to fp noise on Turing.
  • Prefill forward gated: M_GEMMA3_CUDA PASS; M_QWEN3_CUDA functionally PASS on f32 + Q8 (argmax 31/31, KL ~1e-11 — the ship precisions); fp16 sub-gate fails the f32-parity threshold (precision floor, GPU twin of E_CPU_8 — decision owed).
  • GPU autoregressive KV-cache DECODE built + gated (engine 3b6831c): qwen3_decode_cuda (k_rope_at position-aware RoPE + k_attn_decode single-query + k_argmax device reduction, KV resident in VRAM, zero per-step host sync). M_QWEN3_DECODE_CUDA = GPU decode == GPU prefill teacher-forced, 5/5.
  • Speed pass 1 (engine 1af7c9a): f32 6.93 → Q8 11.97 tok/s (1.7×). Honest wall = kernel-launch overhead (~250 launches/token at 0.6B); device-argmax barely moved f32 (sync wasn't the wall). NEXT = CUDA graphs → fused kernels → discrete router on GPU (shared-mem-staged, NO L2 pin on Turing) → llama.cpp head-to-head. Then Stage Gamma (pinned-mem Optane→VRAM, consumer Turing has no GPUDirect Storage).

5.09 STAGE BETA SPEED CLOSED-IN-PARTS + STAGE ETA OPENED (2026-06-06)

The Speed-pass-1 numbers above (6.93/11.97) were COLD-START artifacts — corrected in this session. Full detail in CONTRACT-SPEED §BETA.2/3a/v3/v4 + SESSION-CLOSED-stage-beta-speed.md. The arc, all gated bit-exact / top-1-lossless on the actual RTX 2060:

  • BETA.2 — CUDA graphs. Position-indirect decode kernels (device-scalar int *dpos) make the per-token launch sequence capturable; capture once, replay/token. First commit claimed 7.24→91.55, 12.65×wrong (per-step ran cold, graph ran warm). Anchored (warm + n_gen=256 + both clocks pinned): graphs are ~1.06×. Launch overhead was never the wall — cold-start was (CUDA lazy module load + cuBLAS JIT ≈ 13× first-call; a persistent warm daemon captures it).
  • BETA.3 — the INT8/Q4 dp4a bandwidth ladder. Fused dp4a GEMV reads packed Q8/Q4 arena codes straight from VRAM (no f32 scratch); warp-per-row + 128-bit int4 loads + shuffle reduction; per-tensor precision dispatch (DevTensor.prec) handles K-quant mixes (Q8 head + Q4 body). Isolated GEMV sweep (tests/bench_gemv_int8.cu, both clocks pinned): f32 1× (~290 GB/s = 86% of the 2060's 336 GB/s peak, bus-saturated) → int8 ~3.8× → Q4 ~7.06× at 12B-scale dims, hugging the byte ratio (4:1 / 8:1). Q4 correctness vs host ref: 1.34e-7. At 0.6B/full-clock the decode is overhead-bound (~91 tok/s, all precisions converge); the win binds only at large-model scale.
  • Production wiring + gate. Q4-dp4a wired into qwen3_decode_cuda (the K-quant-mix bug — Q8 head misread as Q4 → 0/256 — caught by the production gate, fixed via per-tensor precision). M_QWEN3_DECODE_CUDA = 28/28: f32/Q8/Q4/.sp-model all 256/256 top-1 lossless.
  • .sp-model adapter fix (engine 2138f89): sp_model_to_qwen3/qwen25 now honor the arch_struct growth discipline (min-copy of min(arch_struct_size, sizeof), zero-fill the appended tail) per PPT-LAT-SP-MODEL-v0 §3 — older artifacts (e.g. arch_struct_size=56 = base+FP16) now load instead of being hard-rejected.
  • METHODOLOGY (now standing discipline): no GPU tok/s number without warmup + long window + both clocks pinned (-lgc locks SM only; a weight-GEMV is memory-bound → GDDR6 clock must be at full speed); confirm the kernel sits on the binding bottleneck (Amdahl); isolated benches validate kernel MATH, production gates validate the DATA-STRUCTURE handoff. See feedback-gpu-microbench-methodology.

STAGE ETA OPENED (2026-06-06, branch stage-eta-gemma4-cuda). Gemma4 (MatFormer/Gemma-3n E-series: AltUp + shared-KV + per-layer geometry + softcap) CUDA forward+decode, so the 6.6 GB Gemma-4-12B-Q4_K_M runs on the 2060 with the ~7× Q4 win. CPU core/forward/gemma4.c is the bit-exact oracle; gate target = gemma-4-E4B. Reference read + 6-stage gated plan (ETA.1–5) banked. Do not rush the variable-geometry port; resume at ETA.1 (adapter + weightless V-norm). Detail in Roadmap §21/§19 + memory project-stage-eta-gemma4-cuda.

5.10 STAGE ETA PHASE 1 CLOSED — THE GEMMA4 CUDA ENGINE (2026-06-06)

The full Gemma 4 (MatFormer E-series) architecture runs on the RTX 2060 — forward AND autoregressive decode — gated 38/38 against the CPU oracle, both live runs green FIRST TRY. Full receipt in SESSION-CLOSED-stage-eta-phase1.md; engine merged to main at 559435c; Roadmap §19.

  • gemma4_forward_cuda: 35 layers of per-layer GLOBAL/SWA geometry + shared-KV (15 owners/20 sharers) + proportional rope_freqs + weightless V-norm + elastic FFN + AltUp + out_scale + tied head + softcap → argmax 12/12, max KL 2.663e-10 vs gemma4_forward. Distributional identity at machine-noise level.
  • gemma4_decode_cuda: autoregressive greedy over a JAGGED shared-KV cache (per-owner [P×kvd_L]; sharers allocate nothing), per-step AltUp, windowed single-query attention → the oracle teacher-forced-predicts every generated token.
  • Why first-try: the bisection bulkheads. Weight-ingest gate (8/8, incl. the cross-seam link — the fork tax collapsed to ONE as_f32→sp_as_f32 shim), L0 math lock (+ the inline-Frobenius-lift finding: gemm_w_lift enforces the oracle's exact-integer-accumulate arithmetic on cuBLAS; + the ×25 norm-amplification analysis), L4 geometry-shift breach (rope_freqs handoff at the floor), L15 sharer-seam proof (cross-layer VRAM addressing exact at 1.1e-5). The monolith had nowhere left to fail.
  • Numerical findings banked: the RMSNorm re-condenses amplified noise at each layer entry (self-healing over depth, no explosion through 16 layers); sharer attention is CLEANER than native (inherited normalized K/V + softmax squashing); ABS error is the gate currency at norm boundaries (rel inflates on near-zeros).
  • NEXT — ETA.5b, the velocity pass (pure physics, gated top-1): device-side PLE gather (sever the per-step host sync) → CUDA-graph capture of the jagged topology → Q4-dp4a routing (the proven ~7× byte diet) → 12B-Q4_K_M transcode + load + tok/s vs llama.cpp. The architecture is proven; what remains is the speedometer.

5.11 C2.4 CLOSED ON AN HONEST NEGATIVE — THE v5 FINALE VERDICT IS MISS (2026-06-06)

The composed 32k Optane finale COMPLETED (FINALEDONE, 58,674 s = 16.3 h, zero errors) and the needle was not retrieved (837492 absent; the answer parroted the needle's context word "Optane" but generic digits). Full disposition: CONTRACT-C2 §C2.4-CLOSURE.

  • Infrastructure: PROVEN at full scale. 1.353B device reads/stream (K 11.1 TB, V 5.5 TB) at 19.55/19.58 µs/read at queue depth; 2 GB LRU absorbed 67.0% on both streams; Ring-1 482×; NTT fusion exact for the whole run; clean teardown. The storage thesis stands.
  • Config regression exposed (no spin): v5 ran the f32 r=16 routerSP_RECALL_BITS/R=64 were dropped from the runner at v4 and never restored, and SP_RECALL_KVSEL was never set in ANY finale script. The script's "KVSEL + bits-r64" header was prose, not config; niah's knob-echo (r=16 B=512) caught it. Also why 16.3 h, not ~2–3 h. Lesson: banners must echo getenv, never aspiration.
  • Regime honesty: B=512 @ 32k = 64× selection — all quality gates were at 2×–8× (N=2048). And there is no full-attention 32k control for Qwen3-0.6B, so router-dilution vs model-ceiling is not yet separable.
  • Disposition: closed AS A MISS per the standing operator decision; NO Optane relaunch. A RAM-only NIAH ladder (N=2k/4k/8k/16k, B=512, mock RAM Ring-2 — minutes/rung) localizes the break point on the budget axis. Paper 01 releases on the 512-position-proven claims; the public README's pre-claimed 32k HIT was corrected same-day.

5.13 THE GEMMA-4 CAMPAIGN CLOSED — SOVEREIGN PIPELINE + CITABLE 06-R10 (2026-06-08)

The ETA.5b anchor (§5.12 below) detonated, and the campaign that followed closed the whole arc in two days. Full record: CONTRACT-SPEED (GOLD INSTRUMENT addendum → RESOLUTION → Q4B SPEC + decision matrix → CLOSED GREEN); receipts lattice tests/gemma4_gold/; public LEDGER 06-R8/R9/R10 + papers 04/05/06 (written).

  • THE GOLD INSTRUMENT: a from-scratch reference forward off the official safetensors measured gemma-4-12B's TRUE wikitext PPL at 4.6776 — llama.cpp's 397–506 was the ARTIFACTS, not the engine (same arithmetic over GGUF-dequantized tensors reproduces the breakage: pre-fix 271–364, post-June-5 rebuilt still 192.9). Forensics: no permutation, in-place period-6 damage, layer_output_scale class independently defective (scalar swap 364→97). GGUF lane DEAD for this model.
  • SAFETENSORS DIRECT (the sovereign pipeline): sp_transcode --st takes weight values from the checkpoint (GGUF = verified-clean metadata/tokenizer only; mapped-but-missing = hard error). OK_Q8 artifact: 4.7396 (+1.33%).
  • OK_Q4B (arena layout v2, formal migration): per-32-block f16 scales, store-then-derive; recipe B1 (Q4B gate/up + Q8 rest, chosen from a 6-recipe SIMULATION matrix — gemma4 is PTQ-hostile, all-sym-32 = +45%) → 9.4 GB artifact at 5.1259 = the simulation to four decimals; GPU kernel k_gemv_q4b_dp4a_v2 (one weight block per 128-bit chunk) landed 5.1160 — triple-instrument agreement. Core 85aadd3, engine bea361e.
  • SHOOTOUT-2 (CITABLE, 06-R10): 26.1 tok/s at PPL 5.12 on the 2060-12GB (graph EXACT 256/256, dp4a top-1 256/256, 24/24). llama.cpp: 31.29 tok/s at PPL 192–506. Engine bandwidth 245 vs 207 GB/s (+18%). §5.12's 34.2 is formally RETIRED with its quality-failed per-row artifact — the series' own anchor rule caught it. Community deliverables: public GEMMA4-QUANT-FIX.md + issue post.
  • Open from the campaign: gemma4 tokenizer dispatch (SPEC written), B2 asym upgrade, in-engine CPU 12B PPL gate behind harness fixes (progress prints + score-only-positions; the serial oracle was killed undetermined at 331 min).

5.14 XBAR — THE AUDITABLE LATENT CROSSBAR (RFC v1.1; P1 CITABLE · P2.b CLOSED · P3 READ-PATH COMPLETE (P3.1/b-1/b-2 bit-exact) · P3.2 WRITE-PATH: spill BYTE-EXACT + paged-read BIT-EXACT + SWA RING SHRINK (40/48 layers) · §P3.2-b-2b GLOBAL SPARSE RECALL CLOSED GREEN — Learned-LSH wins 8× at +0.47% PPL on a 512×32 matrix; KV cache DECOUPLED from context · Phase C alloc-shrink in progress (C-a device-select + C-b.1 sidecar GREEN, C-b.2 VRAM cut next)) (2026-06-09; refreshed 2026-06-13)

The post-gemma4 campaign: a token-free inter-model memory architecture — Exec + a small Memo curator sharing the cyclotomic rings, every write receipted/gated/rewindable. Docs: RFC-XBAR-auditable-latent-crossbar.md (v1.1) + CONTRACT-XBAR-P1/P2/P2b/C1-lite. [PROVEN/CITABLE] and [WIRED] items:

  • [PROVEN — ledger X-R1, public] P1 Inception Probe: a 12B's generation is steered by direct KV-cache transplant, no tokens — 15/15 (5×3 matrix) lexical incorporation, 15/15 selectivity (double dissociation), max 3.69 orders rank pull, dose-response (1 row ~4% attn mass bends ranks; 6 contiguous rows bend words), G0 null bit-identical 7/7, dual-metric coherence (gold-instrument PPL 1.70–4.10). Engine: SP_XBAR_* knobs in cuda_forward.cu (capture/splice/emb/rank), tests/test_xbar_p1_cuda.c. Ledger row X-R1 in Position_Is_Arithmetic.
  • [PROVEN — measured, internal] P2.b Phase 0 (cloud inversion, RunPod A6000): k=2 pseudo-tokens recover a 6-token span on the real bf16 12B — existence proven in two regimes; kl-dropped-weighted Pareto: Arm F (free) ~94–96% gap-closed but off-manifold (~17–25), Arm H (hull) ~64–73% on-manifold (~0.8). Operating point = a P2.b training-time λ-selection (recall-invariant primary), NOT a parity test — the free-generation parity gate was convicted as unusable (greedy decode loops on this it-model in every regime; honest negative banked).
  • [WIRED — gated] C1-lite curator (the local immune system, qwen3 CPU two-ring): C1L.1 transactional core (tools/curator/curator_core.c: clone/gate/atomic-promote/rewind/receipts, G-C1L-1 null PASS) + C1L.0a episode persistence + router re-projection determinism (tools/curator/curator_replay.c links real arm.c/arm_scan.c; projk recovered bit-identically from the persisted K store — no projk serialization needed). Episode format {k.bin, manifest} = NIGHTSHIFT's standard.
  • [DESIGN] Ring 3 (RFC §3.1): a four-tier hierarchy — Ring1 working / Ring2 verbatim episodic (hippocampus) / Ring2′ transient shadow / Ring3 adapter-compressed consolidated (neocortex); NIGHTSHIFT transfer-and-transform under the irreversible-aware G-R3-LOSS gate; resolves the C2.4 64× recall ceiling. RFC §6.1 (verified external work — AI Harness Engineering / RHO, with the PPL-delta-over-self-preference sharpening) + §6.2 (latent-vs-lexical threat-landscape note).
  • [PROVEN — 2026-06-09, tag xbar-c1-lite-complete] C1-lite COMPLETE (supersedes the C1L.0b "NEXT" that stood here): C1L.0a re-projection + C1L.0b replay (SP_REPLAY seam in decode.c, off-path bit-exact; T_GENKV_REPLAY_NULL 34/34) + C1L.1 transaction + C1L.2 cold-evict (T_GENKV_COLD_EVICT 45/45 — lossless cold-evict PROMOTES, hot-evict diverges and REWINDS). The curator's transactional control flow + episode replay proven on the uniform-geometry qwen3 ring. P3 pre-flight audit (CONTRACT-C1-lite §3b, line-cited against gemma4.c): TWO real port gaps — G-P3-GEOM (per-layer-class NKV/HD; decode.c:317 allocates projk on a single uniform HD) + G-P3-SHARED (shared-KV owner-indirect spill/recall); episode byte layout + SP_REPLAY seam transfer as-is; the V-less alarm was a false read (SP projects V independently).
  • [MEASURED — 2026-06-09] P2.b Fork-2: the recall invariant WORKS, 3-seed reproducible. Adding the through-model readback-CE term (+λ_read·L_readback) flipped held-out span recall from chance (§3g: 58/100) to 80–84/100 final on all 3 seeds with recovery held (~0.18) — clean causal attribution; the loss must be computed at the END of the forward. λ_read sweep: usable band [0.25, 0.5], knee 0.5 (n=1, unconfirmed vs the 3-seed 0.25 anchor). Fork-1 k-sweep DONE (2026-06-09): verdict ADAPTER-LIMITED — the k=6 no-compression control did NOT lift recovery (0.155 best / 0.191 final ≈ the k≥2 plateau; both pre-stated predictions hit, incl. the ~60/100 recall endpoint); dilution monotone over the full curve (recall 88→84→72→70→60 for k=1→6). k=2 = the operating knee (dominates all k≥3 on both axes; the only 3-seed point); the recovery lever is adapter capacity/data, NOT k (CONTRACT-P2b §3i verdict).
  • [WIRED — SW-emu proven; HW bring-up unblocked] GNA 2.0 audio lane (RFC-XBAR §3.2): Stage 0–2 probes on real linked libGNA pin the envelope — FiLM = native ElementWiseAffine (accepted to N=32768; ~65k element cap), device memory ≈224 MB usable, 1D conv is the 2.0-native primitive (i16 conv weights, batch=1, single in-channel), i8/i16 + i32 accum (OK_Q4/Q8 maps as a storage codec). Bring-up kit staged locally (archive/notes_and_stuff/GNA/): Windows + Linux drivers, wsj_dnn5b / rm_lstm4f / librispeech_s5 (incl. OpenVINO IR) reference models, aclnet int8 ONNX audio-CNN exemplar (rm_cnn4a_smbr absent — aclnet is the local conv-layout candidate).
  • [MEASURED — 2026-06-10] CAPACITY ARM: verdict NOT-CAPACITY. 12-receipt grid (4 configs ×3 seeds): all recovery medians inside the baseline noise band (0.145–0.168 vs 0.148±0.034); 4.4× params bought zero recovery AND degraded recall (84→68 at 49.9M; overfit signature exactly as pre-named). The ~0.18 plateau is data/objective-limited. λ-leg resolved too: λ=0.5's edge was seed luck (one final-epoch collapse); OPERATING POINT PINNED: k=2, λ_read=0.25, d512/L2 11.3M — the smallest config is Pareto-optimal on both axes (CONTRACT-P2b §3j verdict). NIGHTSHIFT inherits a high-selectivity (80–84/100), bounded-loss substrate; G-R3-LOSS governs.
  • [CLOSED 2026-06-12] P2.b CAMPAIGN + P3.1/P3.1b GREEN. Diagnostic arms all converged: not-capacity (§3j) · not-k (k-sweep) · grok HARD-UNDERFIT (§3k) · context channel-limited (§3o Fork-4) · horizon ASYMPTOTE ~0.28 (§3m) · KV-prefix injection net-harmful (§3p Fork-5, incl. RoPE'd kv2 best −0.249) — GENERATION dead at k=2; the wall is the objective↔task mismatch, NOT channel width. Pivot to RECOGNITION (§3q Fork-6 contrastive native-attention addressing) = real-but-sub-usable: 32-way top-1 0.462 < 0.50 PASS ⇒ REST (pre-registered, no goalpost move; 15× chance, beats native-key 3.3×; top-5 0.77 = shortlister-not-sniper → the two-stage retrieve-verify door). P2.b lane rests; memory cells get a heuristic/two-stage addresser, not a learned sole-top-1. P3 ring-on-Exec is now the active lane: P3.0 manifest CLOSED GREEN (system 9a2b0a9) → P3.1 decode-wiring G-P3-1 BIT-EXACT GREEN on the real 12B (off[L] episode-store indirection in the gemma4 CUDA decode; recall seq == legacy, token-identical; engine cuda_forward.cu) → P3.1b-1 serialized-store G-P3-1b BIT-EXACT GREEN (xbar_episode.c linked into sp_engine_cuda; serialize→disk→deserialize→mount→decode == legacy) → P3.1b-2 recall-as-history G-P3-1b-2 BIT-EXACT GREEN (mount episode into the live-cache FRONT [0,H), decode prompt at [H,..); no offset threading since pos is already absolute; fix = seed dpos=H because the loop-skip bypassed k_incr_pos; continuation == monolithic, diffs=0). THE XBAR READ-PATH IS COMPLETE — bit-exact on the real 12B at every rung (off[L] mirror → serialized disk episode → prepended history). P3.2-a WRITE-PATH STARTED — shadow spill G-P3-R2.a BYTE-EXACT GREEN (the inverse: SP_XBAR_SPILL=dir, per-step owner K/V spilled through the sp_arm_ring2_backend stdio ABI at off[L]+pos·kvd·4, read back byte-identical to the live cache — diffs=0 at the DEC gate (P=16, 5.2 MiB) AND the 259-position velocity decode (85.3 MiB), 48 owners / 0 sharer blocks in store; store length = store_bytes − last-owner-unwritten-slot confirms the [0,P-1) byte law; one-staging-buffer/one-sync batching, per-step-sync perf tax deferred to a P3.2 overlap follow-on). Built on the VS2022/VS18 host (VS2019 can't build the CUDA tree — <stdatomic.h>). P3.2-b-1 PAGED-READ G-P3-R2.b-1 BIT-EXACT GREEN — the closed loop runs live: per step spill pos → POISON [0,pos] (zero the live cache) → page [0,pos) back off Ring-2 (read_block(off[L]) → H2D) before attention. Paged decode token-identical to legacy full-cache (diffs[4..16)=0, both 2 10 100 1000 497 564 …); the poison proves the bytes came off disk, not a stale live copy. The model's entire history lived on disk and fed attention bit-exactly — write-path ∘ read-path as one loop. SCOPE (honest): proves the recall READ; does NOT shrink VRAM (globals attend all positions → need sparse recall, the router; only SWA owners shrink on the substrate alone). P3.2-b-2a SWA RING SHRINK G-P3-R2.b-2a BIT-EXACT GREEN — the FIRST REAL VRAM WIN. Refinement (caught): gemma SWA is a pure sliding window, NO sinks → the window is always live, nothing to page → the shrink is a W-slot RING (not two-source-with-paging, which is the globals' job). The 40 SWA owners (carrying the dominant kvd=2048) shrink from P to W slots: write pos→slot pos%W, k_attn_decode_ring reads in POSITION order (s0+j)%W so the fp reduction is byte-identical. Gate: ring-of-4 == full-cache window-4 decode, diffs[4..16)=0, P=16 wraps 3×. Globals untouched (full P; b-2b's job). Effect: the dominant context-linear cache term (SWA, ~21 GB @ 32k) becomes CONSTANT (~0.67 GB @ W=1024); only the 8 small globals still scale (~1 GB @ 32k) — the last linear term, which b-2b's router collapses. Bit-exact shrink on 40/48 layers, substrate-only, no router, no disk. Full record CONTRACT-XBAR-P2b §3j–§3q + CONTRACT-XBAR-P3 §P3.1 + §P3.2-a (G-P3-R2.a) + §P3.2-b-1 (G-P3-R2.b-1) + §P3.2-b-2a (G-P3-R2.b-2a). P3.2-b-2b GLOBAL SHRINK — MECHANISM CLOSED GREEN 2026-06-13. Frozen sp_arm_select_geom ±1 projection router on the 8 globals, built null-first: Phase 0/1 bit-exact (shadow-select + projk oracle-parity mism=0, live output unchanged) → G2 PPL-deflection 4× −0.31% / 8× −3.21% inside the locked < 2.0% (OK_Q4B -b1 baseline 4.6665 == gold bf16 4.68; negative = C2.1 denoise sign on a noisy n_ctx=84 window) → G1 served-off-disk bit-exact (NaN-poison live globals + sparse page off Ring-2; gather-from-disk == gather-from-live diffs=0, the needle survives the poison). The router SELECTS + SERVES off disk + PRESERVES quality. FOLLOW-ONS (mechanism ≠ deployment): larger-N G2, the literal alloc-shrink (B not P) + device select, and the M_GEMMA4 mis-registration (plain path = coarse QAT variant 7.4M PPL → repoint to -b1). Engine a0b8d42. (spec + locked params: CONTRACT-XBAR-P3 §P3.2-b-2b.) The 8 global owners attend only a budget-B recalled subset via the frozen sp_arm_select_geom ±1 projection router (v0), paged from Ring-2 (composes b-1 page-in ∘ b-2a). First non-bit-exact stage — gate G-P3-R2.b-2b REPLACES diffs=0 with a bounded-degradation pair: G1 needle-HIT ≥ full-attention (NaN-poison rigor) + G2 PPL-deflection < 2.0%, both within the ≤ 8× / N ≤ 2k band (32k/64× explicitly ungated, the C2.4-cliff lesson), top-1 retention reported non-gating; §3q learned two-stage shortlist→verify held as the fallback lever (only un-rested on a v0 breach). (v0 cut + mechanism closed below.)
  • [CLOSED GREEN 2026-06-13] §P3.2-b-2b GLOBAL SPARSE RECALL — LEARNED-LSH WINS 8× = +0.47% PPL. The full arc, measured on the real 12B (N=2048×3 = 3072 scored, vs FULL 5.1551): (1) larger-N G2 (G-P3-R2.b-2b-N) — the small-N (42-pos) negatives were an ILLUSION; on the full wikitext-2 corpus the frozen ±1 router gives 4× +1.65% GREEN / 8× +4.17% RED; a W-probe {4,64,128} floors at +3.74% = the frozen-router 8× Pareto frontier. (2) oracle ceiling (G-P3-R2.b-2b-ORACLE)SP_ARM_ORACLE exact top-B by q·K: 8× = −0.08%, 4× = −0.01% → 8× is NOT information-bounded; the frozen +4.17% is 100% router quality (the diffuse 7.7% dropped mass is noise — C2.1 denoise; offline diag: gemma globals keep only 92.3% mass even at the oracle = DIFFUSE not concentrated). (3) Learned-LSH (G-P3-R2.b-2b-LSH) — a shared 512×32 projection R trained by forward-KL distillation of the true attention distribution (tools/xbar_lsh/train_lsh.py, GPU 0.8s/ep on the 2060; CUDA torch 2.6.0+cu124 installed); deployed via SP_ARM_LSH=M.bin (M=R·Rᵀ, select = top-B by (Mq)·K, reuses k_qk_scores+k_apply_M, zero new hot-path kernels, cost independent of r). G2 8×: LSH r=32 = 5.1791 = +0.47% GREEN (16,384 params, identical inference cost to the v0 frozen router, 0.55pt off the oracle). Weight tests/fixtures/lsh/lsh_M_r32.bin. Engine 08d4d79 (dump) + dab7d36 (oracle) + 222463a (LSH). §P3.2-b-2b CLOSED end-to-end; the KV cache DECOUPLES from context length — SWA ring (b-2a) caps the dominant term at W, the 8 globals cap at the GQA union nh·B (corrected from "B" by C-b.2 below: with n_kv=1, n_h=16 the 16 query heads pick near-orthogonal top-B sets on gemma's diffuse globals, so the per-step union → nh·B = 16·256 = 4096 — constant in P but 16× the per-head B). Both terms O(1) in context (the XBAR thesis), the global constant being nh·B not B; decoupling is visible only for N > nh·B (≤4096 the union is context-bound). ("measure before you mutate" earned its keep twice — the concede-4× bet was flipped to train-8× by the oracle; the mass-proxy was flipped by the on-engine PPL.)
  • [CLOSED GREEN 2026-06-14] Phase C ALLOC-SHRINK (turn the proven selection into realized VRAM, then prove retention): C-a DONE (engine 7195100)SP_ARM_DEVSEL device-side top-B (k_topb_dev), selection-invariant (5.1791==5.1791), severs the host round-trip. C-b.1 DONE (engine 7cd7482)SP_ARM_LSH_R projected-key sidecar (resident r=32 RᵀK, 16× smaller than full K; the architecture finding: you can't rank evicted keys, so the compact slab requires this resident router state), gate +0.24% deflection-invariant (the sidecar is the truer r-dim form). C-b.2 MECHANICS DONE / VRAM-CONSTANT CORRECTED (engine 725058c)SP_ARM_SLAB: full global K/V in host-RAM Ring-2, per-step union paged into compact slots [0,m), gather remapped absolute→compact-slot. Output-invariant GREEN: SP PPL=5.1676 == C-b.1 (G-P3-R2.b-2c). The measure-first catch: per-step union = 1500/1511/1442 of 2048, clip=0 → the compact slab is bounded by the GQA union nh·B = 4096, NOT B+sink = 258 (a B-sized slab would have clipped ~1240 valid keys/step). Corrected thesis above. NEXT (the actual VRAM demonstration): run at N > nh·B (8k/16k/32k, SP_ARM_BSLAB≈4400) so the slab (nh·B, constant) sits below the context — nvidia-smi flat-line is only visible there; at N=2048 (< 4096) the context bounds the union so no shrink shows. C-b.2 O(1) LADDER GREEN (engine 33ac632): N=8192 union=2058 clip=0 PPL 5.0549 VRAM ~11440-11476 MiB | N=16384 union=2287 clip=0 PPL 5.1371 VRAM ~11477-11524 MiB ⇒ ~50 MiB delta across 2× context = O(1) cache CONFIRMED (a full O(N) cache adds ~5.4 GiB 8k→16k); cache alloc byte-identical by construction (globals nh·B, SWA W). Fixed a latent N>Bslab crash first (the post-loop G-P3-GEOM.a oracle-parity diagnostic re-projected dKc[L] over [0,P) but the slab holds only Bslab<P slots → OOB cudaMemcpy → "invalid argument"; guarded !arm_slab). C-c NIAH CLOSED GREEN (engine 8e35877/3218d73, contract 5242955): the needle SURVIVES the O(1) compaction at every depth, only with the learned router. test_gemma4_cuda SP_G4_NIAH mode (token-space, SWA-isolation asserted, slab active Bslab=4400, full K/V in host Ring-2, ranked by resident r=32 sidecar): depth 10%@16k (gap 14729) HIT · 50%@8k (gap 4093) HIT · 90%@16k (gap 1649, SWA lip) HIT — all exact 837492; NEG CONTROL 50% slab+FROZEN±1 = MISS (5/6 digits then corrupted → loop). Full-attn baseline @16k physically impossible on the 2060 (ctx-softmax shared-mem >64KB + cache OOM) = the motivation. §P3.2-b-2b / Phase C CLOSED end-to-end: SELECT (LSH 8× +0.47%) → REALIZE (slab O(1), 8k↔16k flat) → RETAIN (NIAH all depths). SCOPE: the ~11.4 GiB floor = the 9.4 GiB resident model (backend-direct gemma4_decode_cuda harness bypasses the arena's zero-copy streaming); the KV term XBAR controls (~0.8 GiB) is what's O(1). Public LEDGER stays DARK pending operator wording (claim = KV O(1) + NIAH retention, NOT "12B@16k on 12GB"). NEXT = P3.3 SP_REPLAY → P3.4, then KAIROS harness. Full record: CONTRACT-XBAR-P3 §P3.2-b-2b + G-P3-R2.b-2b-N / -ORACLE / -LSH / -2c / -2c-8k / -NIAH run-records.
  • [CLOSED GREEN — 2026-06-17] P3.3 SP_REPLAY (replay-write) + P3.4 G-P3-PPL (recall quality) — XBAR P3 CLOSED END-TO-END. P3.3 wires the inverse of the read path: SP_REPLAY injects a stored episode's owner-K/V over the prefill rows [0,NPOS) at the CUDA cache-store boundary, before attention. G-P3-SHARED 3-leg PASS on BOTH 12B (gemma4-12b-b1, 48 owners) AND E2B (gemma4-e2b, 15 owners / 20 sharers = owner-indirection exercised): an intact replayed episode is bit-identical to baseline (diffs=0), a zeroed episode diverges 12/12 (collapse to a degenerate loop), SP_REPLAY unset = floor — intact-equals-baseline proves the seam is well-formed, zeroed-diverges proves the payload is load-bearing not inert. The inject is placed at both prefill stores in gemma4_decode_cuda (graph-capture ~L2516 AND velocity ~L2825) — the velocity path is the one the gate runs (use_graph false under recall). Harness SP_G4_REPLAY_GATE; runners _run_p33.bat/_run_p33_e2b.bat; receipts tests/fixtures/xbar_p3_replay/G-P3-SHARED_{12B,E2B}_GREEN.log. P3.4 is the recall-quality gate: the PPL scorer IS gemma4_decode_cuda in SP_G4_SCORE mode, so SP_REPLAY composed with it with ZERO new engine code (primitives snap together at the boundary). G-P3-PPL: wiki.tiny n_ctx=84, recall-OFF baseline SP PPL 4.6665 → recall-ON (proven episode, NPOS=4) 4.7311 = +1.38% deflection < 2.0% gate → PASS (foreign episode over the 4 earliest of 84 positions; the model holds focus). Complements the §P3.2-b-2b sparse-recall deflection (learned-LSH 8× = +0.47%). CAVEAT (on the record): n_scored=42, a single chunk — deterministic (replay, not router sampling, so NOT a noise-flippable small-N illusion) but a larger-N multi-chunk run is the named hardening lever before any public headline. Receipt tests/fixtures/xbar_p3_replay/G-P3-PPL_run.log; runner _run_p34_ppl.bat. XBAR P3 IS NOW CLOSED END-TO-END: P3.0 manifest → P3.1 read → P3.2 spill/page/SWA-ring → P3.2-b-2b learned-LSH select → Phase C O(1) VRAM → C-c NIAH → P3.3 replay-write → P3.4 recall-quality. The crossbar reads, writes, compresses to O(1), retrieves under poison, replays bit-exactly, and recalls without breaking PPL.
  • [CLOSED GREEN — 2026-06-17] THE ORCHESTRATION TIER ABOVE P3 — C2 Memo curator + #222 O(1) rewind + Ring-3 Path A + the EAR→Ring-2 organism bridge. The whole XBAR memory stack now stands.
    • C2 Memo curator CLOSED (autonomous Ring-2 recall loop; engine tools/curator/, receipts tests/fixtures/xbar_c2/, contract CONTRACT-XBAR-C2-memo-curator-loop.md). Step 1 registry + centroid-sig writer (G-MEMO-CUE offline). Step 2 resolver RE-ORIENTED to a discrete bit-collision gate (Shannon-Prime course-correct off a float threshold): 256-bit LSH hash, XOR+popcount, integer Hamming radius TAU_BITS=168, r=256 — an r-sweep proved sign-binarize collapses @r=32 (gap −1) and recovers @r≥128; the integer gate is reduction-order-immune (a float cosine near τ can flip across fp reduction orders). Verified-not-adopted: Gemini's "the ARM dot IS a Hamming distance" was FALSE as-built (real centroids ≠ ±1). Step 3.0 G-MEMO-NULL GREEN — orchestrator inert when off (PPL 4.6665 == 4.6665 bit-identical, cue-extraction seam fired, empty-registry→NULL; gemma4_decode_cuda byte-untouched). Step 3.1 G-MEMO-LOOP GREEN on the 12B — ACCEPT matched recall +0.000% deflection / REJECT corrupted recall +40106% (the safety valve flags + discards). SELECT is order-immune ⇒ the offline G-MEMO-CUE(discrete) verdict transfers online by construction.
    • #222 CLOSEDSP_REPLAY ported into the persistent gemma4_kv_* ABI as gemma4_kv_replay + O(1) bit-exact rewind (cuda_forward.cu). G-222 GREEN E2B (15 owners) + 12B (48 owners): replay-inject load-bearing, gemma4_kv_rewind resets the pre-injection prefix [0,anchor) byte-identical (layer-diffs=0). G-222-WRAP GREEN — SWA-ring (KAI-1c journal): replay into an active sliding-window session, each clobbered ring slot journaled before overwrite, journal-backed rewind diffs=0. Local KV airtight in BOTH regimes (full-cache + SWA-ring). The curator speculates a recall in the resident cache and undoes a rejected one in O(1) byte-exact — the §4-trap guarantee made mechanical. Receipts tests/fixtures/xbar_c2/G-222*.log. Engine b4b037a / 24071bc.
    • Ring-3 gist consolidation Path A CLOSED end-to-end (VSA/HRR, parameter-free, zero training budget; engine tools/ring3/, receipts tests/fixtures/xbar_r3/, contract CONTRACT-XBAR-R3-consolidation.md). Architecture = retrieve-and-verify (honoring the P2.b verdict: generative gist-fill dead, top-5 shortlisting is the door); consolidation-time only (§4 recall-time upsampling forbidden). R3.1 G-R3-BINDM = Σ(addr⊛id), circular conv = the NTT algebra; addr seeded by the episode's real C2 256-bit sig, id = clean ±1 label; recall@1=1.0 to N=32 @ D=1024 (margins +0.586/+0.568), ±1 substrate carrier ≈ ideal unitary (metric-bug caught+fixed: SNR ratio→margin/z-score at N=2). R3.2 G-R3-LOSS — the consolidation loss is a STEP FUNCTION: hit lossless +0.000%, miss +8.04% caught by the 2% gate (degrade-safe, never silent corruption); promotion budget ≤32 episodes/vector; latency shear = 71µs unbind + 1 Optane read. R3.3 G-R3-DUALROUTE — the continuous pipe (cue→VSA unbind→top-K shortlist→#222 verify scan→land): clean-hit + decoy-scan (reject foreign rank-1 +8.04%/rewind → accept correct rank-2) + null parity. R3.4 G-R3-NIGHTSHIFT — idle-loop consolidation state machine (bind→shadow-gate the whole bound set→promote+evict-to-Optane→saturate&seal): D=1024 seals at CAP=32 (349.8 MB resident KV → 16.3 KB Ring-3 index); D=128 gate-driven seal proves the cap is the capacity math, not a constant. Deferred (named): the Z_q/NTT engine port of the host-numpy VSA; the G-R3-PROV provenance tag; Path B (the trained adapter) stays budget-gated, untouched. Engine 23539b7a64a916.
    • G-XBAR-ORGANISM step 1 GREEN — the EAR→Ring-2 write seam (engine run_kai3_write / SP_G4_KAI3_WRITE, receipt tests/fixtures/xbar_organism/, engine 6600cf4). A real audio packet (the KAI-3 gemma4_kv_inject_seq path that pivots the 12B 7/8) → conditioned cache npos=114ep_audio serialized in the canonical uniform-512 episode format (layout-bug caught via a size sanity-check: 12B cache is jagged global 1×512 / SWA 8×256=2048; the episode clamps to the global 512, ep.k = 48×114×512×4 = 11,206,656 B = the _c2_ep_wiki format). Signature gate: the audio-derived 256-bit sig separates cleanly (self 211/256, margin +79; distinct from ep_wiki/ep_toy). Round-trip: SP_REPLAY=ep_audio loads + injects clean (RT_EXIT=0); the +1989% deflection is FOREIGN-BY-DESIGN (audio episode vs a wikitext score context — ~0% is matched-context only, the high deflection IS the reject signal). The auditory front-end and the episodic memory core are physically connected.
  • NEXT (queue): (1) the period-6 sig re-base — the C2/Ring-3 256-bit sig pipeline uses PERIOD=8 (L%8==7) as a consistent content-hash layer subset; the 12B true SWA period is 6 (L%6==5, confirmed by the organism diag). Separation is robust to the choice so all prior C2/R3 gates STAND — this is a correctness tidy-up, not a result fix. (2) the full G-XBAR-ORGANISM loop — drive a raw audio cue → Ring-3 shortlist → #222 verify scan (reject mismatched text blocks) → autonomously land ep_audio into the resident cache. Deferred/named: the Z_q/NTT engine port of the VSA bind/unbind (deployment); the P3.4 larger-N multi-chunk hardening run; Gemma-4 MTP draft head (PPL-vs-4.68-gold gated first). (historic: P3 end-to-end, C-b.2 O(1) VRAM, C-c NIAH — all CLOSED above.)
  • Cloud infra proven + documented: SSH-free HF-mediated self-terminating RunPod pattern (A6000 $0.33/hr); see RUNBOOK-cloud-compute.md + memory reference-cloud-compute-runpod-hf.

5.12 ETA.5b CLOSED — THE 12B SHOOTOUT WON (2026-06-07) (SUPERSEDED by §5.13: the 34.2 is RETIRED — its artifact failed the PPL gate)

SP 34.2 tok/s vs llama.cpp-CUDA 31.29 ± 0.20 (+9.3%) — Gemma-4-12B, RTX 2060, tg256, SM pinned (-lmc unsupported on GeForce; memory free-ran for BOTH engines). Engine af738f9, core e8708f7. Full record CONTRACT-SPEED §ETA.5b; receipt _12b_shootout.log.

  • ANCHOR: not citable until the PPL gate closes — the SP artifact squeezes Q6_K source tensors to Q4 (5.56 GB vs the 6.62 GB GGUF; fewer bytes = part of the win, more weight-quant error). Named release-blocking gate for paper 06: wikitext PPL, both engines.
  • E2B ladder (44/44): lift 10.3 → graph 10.6 → dp4a 62.3 (6.05×) → graph+dp4a 75.7 tok/s (7.35×) — Amdahl-clean composition (device PLE gather + packed tied head + jagged graph capture).
  • The dense 12B ≠ E-series: PL=0 but out_scale + rope_freqs present (now presence-keyed); no KV sharing; per-layer kv-head ARRAY (8 SWA/1 global); V-less globals (V = raw K projection) — landed across transcode/bridge/oracle/CUDA, E2B regression held throughout.
  • THE L11 KILL: per-VECTOR int8 activation quant collapsed on the 12B's outlier-heavy activations (L11, trained out_scale 0.005) → oracle-rank 205596. Operator-directed bisection (provenance → embed 0.000e+00 → norms smooth → layer bisect → LIFT discriminator: structure at 1.5e-4 floors everywhere) pinned it to the quant. Fix = per-16-BLOCK scales aligned to the 128-bit loads (zero extra bus). Verdict: rank 2 at gap 0.31 — a measured top-2 near-tie (gates now print oracle-rank on any flip). 12B 24/24, E2B 44/44, qwen3 green.

5.1 FORWARD PRIORITY (re-ordered 2026-06-02 — differentiators ahead of context)

C2's measurement phase is done and re-ranked the work (KV ~3.5× lossy; Ring-2 context ~hundreds× but largely disk-tiering). The unmeasured load-bearing differentiators now lead. Full rationale in RFC-001 §11:

  1. P1 — SPEED / WIRE gap → tok/s vs llama.cpp (the north-star; integer pipes still scalar-f32 off-Hexagon; HX.3b 1.04× bandwidth-bound is the warning).
  2. P2 — C4 MTP (T8 exact O(1) rollback).
  3. P3 — C3 multi-device CRT residues + Garner service (2-node CRT-shard byte-exact vertical slice = the proof). [RESOLVED — do NOT re-offer, see §4 table row + SESSION-PERF-SYNTHESIS-2026-06-23.md] CRT shards numbers not experts (per-device memory saved = 0); it is a bit-exactness enabler for a heterogeneous split, not the split mechanism, and this box has no viable second island (iGPU 0.75 TFLOPS, shares DRAM). The only residual sliver is a literal 2-discrete-GPU bit-identity diff for auditability — out-of-band, low priority, NOT a perf/VRAM lever.
  4. P4 — remaining C2 (fp16 swivel, qwen36 Spinor-KV wiring, a directional recall router — KSTE ruled out) — DEMOTED, secondary context axis.
  5. P5 — C5 eMeMo, C6 cyclotomic paper.

6. Open blockers (honest)

  • BUILD (the recurring root cause, diagnosed 2026-06-02). The engine is GCC-authored and was never made MSVC-clean; the MinGW build dir SEGFAULTS at runtime (test_gen_kv 0xC0000005, known-good code). So the CPU build does not cleanly build+run on EITHER toolchain. Env pin FIXED + committed (engine 33c6a27): scripts/env/env-common.bat now pins SP_PIN_VS_BUILDTOOLS=D:\Program Files (x86)\Microsoft Visual Studio\18\BuildTools (MSVC v14.50, cl 19.50) — the prior ...\2019\BuildTools pin was the phantom P0.1 flagged. CORRECTION (per docs/BUILD-ENV.md, the authoritative build doc): the canonical CPU backend is MinGW gcc 15.2 (build/ dir), operator-approved — MSVC CANNOT build CPU (known, Tier-3-deferred); build-cpu/=CUDA-host only. A prior step this session wrongly repinned the CUDA-host VS to VS18 + chased an MSVC CPU build; reverted (engine 6fc0832), VS18 saved as the separate Tier-3 SP_PIN_VS2022_BUILDTOOLS. SP-side tok/s baseline ALREADY EXISTS: WIRE-CPU daemon ~1.206 t/s, WIRE-CUDA 1.526 t/s on Qwen3-0.6B (engine ea0d0ac/a299ed0; scalar hot path — the integer-pipe wiring WIRE-CPU-V2 is the P1 gap; llama.cpp ref 28.2 t/s → ~23× to close). The de-GCC work below is Tier-3 MSVC-parity progress (kept, GCC-safe), NOT the CPU build. Tier-3 MSVC-parity: engine lib + sp_toks.exe now COMPILE + LINK UNDER VS18 (de-GCC committed). Commits: engine db84bf3 (SP_TARGET macro / ternlog alignas / persist tail) + core submodule 777a10e (sp_channel /experimental:c11atomics) + 33c6a27 (persist atomics shim). Still GCC-only but NON-BLOCKING (micro-benches off the forward path): tests/test_avx512_persist.c (__ATOMIC_*), tests/bench_avx_spinor_sweep.c (stream_nt) — de-GCC later for the full suite; cmake --build build-cpu --target sp_toks builds the forward path now. RUNTIME SEGFAULT — FIXED 2026-06-02 (engine 0fb39ab). Root cause = the fork-tax struct divergence (cf. project-arch-struct-divergence): the engine's include/sp_engine/model.h qwen3_config/qwen3_layer/qwen3_model were STALE — missing the gemma4 (g4_*) + qwen36 (q36_*) fields the core added in the d8e614f bump. Since cfg is embedded by value, sizeof(qwen3_config) differed → token_embd (every field after cfg) sat at the wrong offset → core's qwen3_load wrote it where the engine read NULL → embed_row segfaulted. Localized by instrumentation (cfg read fine, token_embd=NULL in embed_row), NOT the toolchain/de-GCC/harness. Fix: synced the engine's three structs to core byte-for-byte (+ do-not-diverge note). sp_toks now RUNS. SPEED_BASELINE MEASURED: Qwen3-0.6B-f16 CPU = 0.84 tok/s (sp_toks, as-is f16 path) vs llama.cpp 28.2 t/s → ~33× gap = the WIRE-CPU-V2 integer-pipe work (P1). (Consistent with the WIRE-CPU daemon's ~1.2 t/s.) See reference-cpu-build-toolchain memory.

  • WIRE gap: CPU/CUDA/Vulkan forward shells still call scalar f32 (Hexagon done via HX.3b). The envelope isn't realized until shells call the integer + Spinor-KV primitives.

  • qwen35moe .sp-model: needs the reducing OK_Q4 artifact (OK_Q8 was backwards → 35 GB) + sp_model_to_qwen36 bridge + arena-aware expert path. Forward already gated GGUF-direct (M_QWEN36).

  • engine↔core fork tax: duplicated forwards / dequant / row_bytes / arch-id enums. "One object" presupposes de-duplication.


6.1 Disk / storage layout (Knack's host, 2026-06-02)

Drive Kind Role
C: OS SSD (~50 GB free) build trees, scratch
D: working SSD (~50 GB free) repos (D:\F\shannon-prime-repos), source GGUFs, active .sp-model (transcode source+artifact side-by-side, ~40 GB peak — fits)
E: / F: Intel Optane (16 / 32 GB) reserved for Ring-2 KV offload spill/recall — near-RAM latency + byte-addressable = the tier that keeps "context beyond RAM" fast. A C2/C3 design input, not just storage.
G: Google Drive 5 TB (streaming) cold archive only (backups, superseded artifacts). Never mmap'd live.
H: external SD 1 TB bulk cold model storage + transcode overflow

qwen35moe transcode: source (D:, 19.7 GB) → OK_Q4 .sp-model (D:, ~20 GB). No longer blocked.

7. Discipline that is working (keep it)

Clean rewrite · bounded crates + frozen seams · contract system (RFC + C1–C6) · per-cell closure docs · oracle-fingerprint validation · honest PROVEN/TARGET tagging · surface-upstream-never-silently-revise-a-gate · separate worktrees for parallel agents · this STATE ledger updated every session. This is the structure that finally works. Maintain it. Update this file at the end of every session.

KAIROS — the time/agency axis (KAI-1 + KAI-1b + KAI-1c CLOSED 2026-06-14; 6h soak GREEN 2026-06-16; KAI-2 CLOSED-BOUNDED + KAI-3 audio-port CLOSED GREEN 2026-06-16; GNA "EAR" line CLOSED on PHYSICAL SILICON 2026-06-17; C2 curator + Ring-3 Path A + #222 + organism step-1 CLOSED 2026-06-17; XBAR UNIFIED onto the exact-integer O_K substrate — 10 receipts, boundary thesis — 2026-06-18; BYTE-EXACT FORWARD CLOSED GREEN on the 12B — 2026-06-18; CHAT-FULLSTACK LIVE — coherent + byte-exact + O(1) + single-entry 12B chat — 2026-06-19; B3-WC AUTONOMOUS RECALL RESOLVED — learned-head instance recall LIVE on the 12B chat — 2026-06-20)

Framing (load-bearing — keep these two programs distinct): KAIROS is the latent-interrupt / agency-time axis and is the BASIS OF THE XBAR latent-space memory (the token-free receipted crossbar). Its lineage is KAI-1/1b/1c (heartbeat, O(1) bit-exact rewind, journaled SWA ring) + KAI-2 (latent interrupt). The GNA "EAR" line is a separate-but-related sibling program: real AUDIO in/out via the Intel NUC "Beast Canyon" GNA 2.0 always-on hardware (an always-on "ear" giving the model real-world audio). KAI-3 (the audio-port frame projector) is the BRIDGE into the GNA line — it shares KAIROS's frozen gemma4_kv_inject residual-entry seam, but it is not a replacement for KAIROS latent memory. The audio/GNA work is a deliberate near-term pivot; the project pivots back to XBAR (KAIROS latent memory) afterward.

  • [CLOSED GREEN] KAI-1 heartbeat null — control-plane mechanism proven on qwen3-0.6B (sp_daemon kairos: cold-evict prune + SALIENCE≥0.5 policy + O(Δ) flat); production cognition+stability proven on gemma4-12B (perfect 24-tick crucible: 21/21 idle→NO_OP, 3/3 salient→coherent ACTION, 0 false/0 missed/0 malformed; tick-5 post-action reversion defeats the 0.6B corruption attractor). Public ledger KAIROS-01 (Position_Is_Arithmetic). Contract §4.

  • [CLOSED GREEN] KAI-1b metal eviction — cold-evict at the XBAR tensor layer: persistent-KV gemma4_kv_* + O(1) rewind(Δ) in cuda_forward.cu (one-shot gemma4_decode_cuda byte-untouched). G-1b-REWIND-NULL bit-exact (16.5 MB / 48 layers / diffs=0 + EQUIV gen-reproduce); O(actions)→O(1) latency receipt (metal slope 0.0073 vs prefix-grow 0.924 s/action, 127× shallower; 16.7× @ A=16). Engine 0bb94f1, contract §5.5. The crossbar (X-R2 O(1) memory) and the heartbeat (KAIROS-01 agency) are now one system: memory that doesn't grow with context + agency that doesn't spend compute on silence + eviction that is an O(1) coordinate shear.

  • [CLOSED GREEN] KAI-1c wrap-aware journaled ring — unites KAI-1b's O(1)-time rewind with X-R2's O(1)-space SWA ring. The ring aliasing hazard (an idle tick's wrapped writes overwrite still-live window slots) is defeated by an undo-journal (save-before-overwrite, restore-in-reverse on rewind, cleared by commit); journal bound = min(k,W)/owner/tick = constant ⇒ O(1) in both axes. G-1b-WRAP-NULL byte-exact across a forced wrap (non-vacuous: 40/40 SWA owners clobbered, post-rewind diffs=0 + EQUIV); journaled-ring O(1) telemetry ring slope 0.00365 ≈ full-cache 0.00371 s/action (the journal adds no asymptotic cost — exact per-tick tax is below the 2060's wall-clock floor since its memory clock can't be pinned, filed for cudaEvent timing). run_kairos_metal semantic crucible (commit-on-action / rewind-on-idle on the journaled ring): 24-tick tape 0 false / 0 missed / 0 malformed / 0 pos-violations, 3 salient → coherent ACTIONs (start/clean/renew), every post-action idle tick reverts to NO_OP. Engine through b0d2bf6, contract §5.6-5.8. Wrap-correctness and semantic-correctness are proven on orthogonal axes (G-1b-WRAP-NULL in isolation; the crucible on the faithful W=1024 window) — combining them would corrupt one of the proofs.

  • [GREEN — 2026-06-16] G-KAIROS-1 6h endurance soak (run_kairos_soak, _run_kairos_soak.bat 6, DEDICATED local RTX 2060, ring_W=1024 Jmax=160, clocks pinned 1680): SOAK_EXIT=0; 351 loops / ~8,400 ticks / 6h01m; 0 false / 0 missed / 0 malformed / 0 pos-violation — salient→ACTION, idle→NO_OP throughout, clocks reset on exit. The journaled-ring metal ran a multi-hour reflex loop unattended on consumer silicon with zero drift/leak — the strongest endurance receipt to date. The dedicated GPU gave the uninterrupted run the shared desktop kept false-aborting (prior best 6.5h = contention-aborted on the global-free tripwire, a harness/contention issue NOT a substrate failure). The formal ≥24h gate is un-pursued by operator choice (NOT failed). Logs engine/results/{soak_console.log, kairos_soak_detail.log}. Contract §5.9.

  • [CLOSED — BOUNDED — 2026-06-16] KAI-2 latent interrupt (engine c5628e4, contract §6.6 2675c79). Two findings, both on the record: (1) Phase-1 latent-delivery seam gemma4_kv_inject = GREEN / frozen verified asset — the EMB control passed 2/2 on the 12B OK_Q4B / RTX 2060 (a real-token embedding SEQUENCE pivots salient→ACTION, idle→NO_OP); this seam is the load-bearing primitive KAI-3 + the GNA EAR line build on. (2) Phase-2 learned compressed single-event codec KAI2Codec = BOUNDED — the maximally-constrained t10 packet (k=16, on-manifold cos 0.9913, sharp τ=0.2, held-out val_KL plateau 0.9157) MISSED the salient pivot (PACKET 1/2). The wall is SEQUENCE-POSITIONAL (a fixed-width static packet compresses out the per-position directional variance attention routes on; NOT manifold-distance, NOT capacity). No more codec-compression cycles. This honest negative is what motivated KAI-3 (the inverse — a sequence, not a packet).

  • [CLOSED GREEN — 2026-06-16] KAI-3 audio-port frame projector (engine e35a227, contract §7.3 e826950) — the inverse of KAI-2 + the BRIDGE into the GNA "EAR" line (a separate-but-related sibling of KAIROS, sharing the same inject seam, NOT a replacement for KAIROS latent memory). Inject a SEQUENCE of N projected frames (1:1 with positions, no compression). New engine ABI gemma4_kv_inject_seq (strict loop over the frozen inject+prefill primitives; G-KAIROS-3-NULL 2/2 byte-identical to the inline EMB loop). Projector tools/audio_port/{gen_synth_frames,frame_projector,emit_corpus}.py = per-position MLP 640→V_sub + on-manifold binder softmax(logits/τ)·W_sub (W_sub = real embed rows×√H), trained with DENSE PER-POSITION cross-entropy (the fix for the t10 sparse-gradient plateau; the pivot is a consequence, never the train signal). Done LOCAL / NO CLOUD — the engine owns the gemma tokenizer (new SP_G4_TOK_DUMP mode); a cloud G4 for a tiny MLP would be over-provisioning. Synthetic ladder noise_rel=0.1 (2.5× noise:signal) → held-out per-position top-1 1.000, manifold cos 0.9998 (binder noise-independent); real-token train V_sub=60 → top-1 0.931, cos 0.9937. G-KAIROS-3 metal gate (SP_G4_KAI3 manifest): 8/8 SEMANTIC pivots on the resident 12B (salient → event-specific ACTION like "Restart the build process" / "Check disk status and run SMART"; idle → NO_OP), KAI3_GATE_EXIT=0. Receipts _xbar/p2b/kai3_gate.log, tools/audio_port/KAI3-LADDER-RESULTS.md (engine repo).

  • [CLOSED — 2026-06-17] GNA "EAR" line — REAL AUDIO front-end PHYSICALLY REALIZED on the Intel GNA 2.0 silicon (contract §7.4/7.5/7.6). The synthetic anchor is replaced by a real audio path, lowered onto the GNA hardware and verified end-to-end. (1) Real speech → 12B pivot 7/8 — real TTS speech → log-mel → a GNA-conservative Conv1d encoder + CTC head → gemma4_kv_inject_seq → the resident 12B pivots 7/8 (up from 3/7 first run); held-out CTC token recovery 0.44 → 0.868 under a multi-voice bake (924 samples / 2 voices / 400 ep) — the 3/7 ceiling was data-starvation, not architecture (exactly as predicted). All 4 NO_OPs correct; 3/4 ACTIONs correct + coherent; 1 miss = conservative ACTION→NO_OP. (2) Quant ladder (OpenVINO 2023.3, GNA 2.0) — ONNX→OV-IR FP32 bit-exact (CPU 0.877 == torch 0.877); GNA default i16 naive PTQ shears 0.877→0.667 (−0.211, scale-invariant ⇒ real int16 quant loss = the predicted CTC-head shear); NNCF calibrated INT8 on CPU recovers the head (0.860) but its FakeQuantize won't compile on GNA; POT DefaultQuantization, GNA-native i16 = 0.877 FULL RECOVERY (== FP32). Two GNA conv constraints fixed at zero cost: encoder conv padding 1→0 (GNA = VALID only) and the CTC head 33→36 out-channels (GNA filters must be a multiple of 4; dummy channels sliced). (3) GNA_HW on the physical accelerator = 0.877 — run on the real Intel GNA 2.0 in the NUC "Beast Canyon" (i9-11900KB, driver gna_03.05.00.2116, BIOS-enabled): 0.877 == GNA_SW_EXACT emulation == FP32. The EAR front-end is PHYSICALLY REALIZED (native-Windows OpenVINO 2023.3 — WSL2 has no GNA MMIO passthrough). Engine tooling tools/audio_port/{ov_gna_score,ov_score_ir,pot_gna_quantize}.py + run_gna_hw.bat + GNA_HW_BRINGUP.md; receipts _xbar/p2b/kai3/G-KAIROS-3-{AUDIO_7of8,GNA-i16_quant_gate,GNA-HW}.log. The near-term audio/GNA pivot is DONE; the project has pivoted BACK to XBAR (KAIROS latent memory).

  • [CLOSED GREEN — 2026-06-17] C2 MEMO CURATOR + #222 + G-XBAR-ORGANISM step 1. The autonomous Ring-2 recall loop and its local-KV deployment tier are CLOSED end-to-end. Full record: CONTRACT-XBAR-C2-memo-curator-loop.md. Summary:

    • Step 1 — registry + centroid-sig writer (tools/curator/build_registry.py): two real 12B episodes written; separation margin ep_toy +0.2158, ep_wiki +0.1534 (held-out cues, self-score is row-max, positives clear background). G-MEMO-CUE (offline) GREEN. Fixed: npos bounds the sig to the true filled prefix (garbage-tail fix).
    • Step 2 — offline resolver RE-ORIENTED to a discrete bit-collision gate (the Shannon-Prime call against a float threshold): 256-bit LSH hash (sig_bits), XOR+popcount, integer Hamming radius TAU_BITS=168. r-sweep (tools/curator/rsweep.py) proved sign-binarize COLLAPSES at r=32 (bit-gap −1; ep_wiki margin lives in magnitude, not sign), recovers at r≥128 (+6), ships at r=256 (+19). Why discrete: reduction-order-immune and hardware-independent (a float cosine near τ can flip across reduction orders; an integer popcount over a fixed bit-hash cannot). G-MEMO-CUE (discrete r=256) GREEN: held-out cues resolve to own id at 177/178, all 8 unrelated queries → NULL ≤140. tools/curator/discrete_resolve.py. Engine 6dd87b9.
    • Step 3.0 — G-MEMO-NULL GREEN (engine 3ea0587): curator host state machine (tools/curator/curator_loop.py) wired over proven seams; gemma4_decode_cuda byte-untouched (null floor). LEG A baseline PPL 4.6665 == LEG B cue-extraction-ON PPL 4.6665 bit-identical; cue observer fired (23,133,792 B dumped); empty-registry → NULL → inert. Option A = one-shot loop first (reject = discard-and-rerun, O(context)); the O(1)-rewind port is #222.
    • Step 3.1 — G-MEMO-LOOP GREEN on 12B (engine 627dfad): ACCEPT matched recall +0.000% deflection → PROMOTE; REJECT zeroed recall +40106.6% → FLAG + DISCARD. SELECT is the discrete cue (proven offline; transfer confirmed). Corrections applied: negative control = fresh negatives + G-MEMO-NULL (not ep_toy which is a true positive for itself); deflection is the safety valve, not the selector; REJECT leg explicitly exercised.
    • #222gemma4_kv_replay into the persistent KV ABI (engine b4b037a): injects a stored episode's owner-K/V directly at [dpos, dpos+npos), advancing dpos. G-222 GREEN on E2B (15 owners) + 12B (48 owners): replay load-bearing (zeroed slots read back all-zero: 0/36864 E2B, 0/688128 12B); rewind [0,anchor) byte-identical (diffs=0). Receipt G-222-REWIND-NULL.log. #222 WRAP (engine 24071bc): journal-backed SWA-ring replay (each clobbered ring slot checkpointed before overwrite); G-222-WRAP GREEN E2B+12B. Receipt G-222-WRAP.log. The local KV substrate is airtight in both regimes — full-cache and SWA-ring.
    • G-XBAR-ORGANISM step 1 GREEN (engine 6600cf4): EAR→Ring-2 write seam — real audio packet (KAI-3 inject_seq path) → conditioned cache npos=114 → ep_audio serialized in the canonical uniform-512 format [48,114,512] (clamp fix: 12B SWA jagged global 512 / SWA 2048 → clamp to global 512 = _c2_ep_wiki format). Signature separates (self 211/256, margin +79, distinct from ep_wiki/ep_toy). SP_REPLAY=ep_audio loads+injects clean (RT_EXIT=0); +1989% deflection is foreign-by-design (audio≠wiki context; ~0% is matched-context only). Artifacts at _xbar/p2b/kai3 + HF bucket D:\Files\Models. NOTE: C2/R3 pipeline uses PERIOD=8 (L%8==7) as the content-hash layer subset; the 12B's true SWA period is 6 (L%6==5). Separation is robust to this choice and ALL prior C2/R3 gates STAND — a period-6 correctness tidy-up is the named open follow-on.
  • [CLOSED GREEN — 2026-06-17] Ring-3 Path A CLOSED end-to-end — parameter-free, zero training budget. Full record: CONTRACT-XBAR-R3-consolidation.md. VSA/HRR superposition on the existing NTT/CRT substrate; the neocortical gist tier above closed Ring-2.

    • R3.1 G-R3-BIND (engine 23539b7): store M = Σ_i (addr_i ⊛ id_i) (circular conv = NTT-over-Z_q algebra), addr_i seeded by episode's real C2 256-bit signature (content-derived). N=2 recall@1=1.0, margins +0.586/+0.568; capacity (D=1024) recall@5≥0.90 to N=64; ±1 Rademacher carrier ≈ ideal unitary. Caught+fixed: metric bug SNR ratio → margin/z-score at N=2.
    • R3.2 G-R3-LOSS (engine aae3131): consolidation loss is a STEP FUNCTION — hit fidelity 0.000% (verbatim Ring-2 verify); capacity miss +8.04% (>> 2% gate → caught + O(1)-rewound, never silent); promotion budget ≤32 episodes/vector @ D=1024; unbind latency 71µs + one Optane ReadFile (negligible).
    • R3.3 G-R3-DUALROUTE (engine 69638cf): continuous retrieve-and-verify pipe — cue→VSA unbind→top-K shortlist→#222 verify scan→land. Three pipes: (a) clean hit: +0.000% ACCEPT; (b) decoy scan: rank-1 foreign +8.04% REJECT+rewind → rank-2 correct +0.000% ACCEPT; (c) null parity: empty Ring-3 → NULL → baseline byte-exact (== no module). Survives wrong candidates; degrade-safe throughout.
    • R3.4 G-R3-NIGHTSHIFT (engine a64a916): idle-loop consolidation state machine — SELECT→BIND(shadow copy)→SHADOW-GATE(re-verify every bound episode, crosstalk-safe)→PROMOTE+EVICT(verbatim stays on Optane, tier-demotion not delete)→SATURATE&SEAL(gate-driven, CAP=32 safety cap). D=1024/CAP=32: 40 episodes → 349.8 MB resident KV demoted to Optane, Ring-3 resident index 16.3 KB. D=128 small run: gate fires before cap (seal at [10,6,15,8,1] max 15 < 32), proving seal is the math, not a hardcoded value. ⇒ Ring-3 Path A CLOSED end-to-end, parameter-free, zero training budget. Deferred: Z_q/NTT engine port (deployment), G-R3-PROV provenance tag, Path B adapter (only if shortlist insufficient, behind operator budget green).
  • [CLOSED GREEN — 2026-06-18] XBAR UNIFIED ONTO THE EXACT-INTEGER O_K SUBSTRATE. The whole XBAR memory architecture was re-carried from the generic float carriers it had been running on onto the exact-integer O_K substrate (Q(√−163), the dual-prime negacyclic CRT-NTT, core/ntt_crt + core/poly_ring — the algebra the project is built on). Ten receipts, all GREEN or honest-negative. Engine origin/main 0019b86→d2d7ceb (all pushed). Receipts: engine tests/fixtures/xbar_r3/{G-R3-BIND-on-OK.log,-legB,G-R2-FROB-PARITY,G-R2-FROB-AB,G-R2-FROB-ENTROPY,G-R3-MOBIUS,G-T2-WEIGHTS,G-PERIOD6-REBASE} + tests/fixtures/xbar_organism/G-XBAR-ORGANISM-FULL.log.

    • (1) G-R3-BIND-on-O_K Leg A — GREEN (engine 0019b86, tools/ring3/g_r3_bind_ok.py). The Ring-3 VSA bind moved off host float-FFT onto the engine-native exact-integer dual-prime negacyclic CRT-NTT (frozen primes q1=1073738753, q2=1073732609, M=1152908312643096577 — already linked into sp_engine_cuda, zero new linkage). (a) C-ENGINE PARITY 256/256 bit-identical: numpy-int negacyclic == native sp_pr_mul / ntt_forward∘pointwise∘inverse / sp_pr_inner / sp_pr_score_kstore (encode) — the header EXACTNESS CONTRACT holds. (b) MARGIN PARITY: ±1 carrier int == float recall (lossless encode); recall@1=1.0 to N=16, recall@5=1.0 to N=32 @ deg-512. (c) REDUCTION-ORDER IMMUNITY: integer superposition M byte-identical across 8 summation permutations; the float M diverges 4.44e-15 (non-associative). Seam survey: the gemma4_kv_* resident cache is pure f32 (the INT8/dp4a path is the WEIGHT gemv, not the cache) ⇒ the memory tier was float-decoupled and poly_ring/ntt_crt are reachable from the XBAR host lane directly.
    • (2) G-R3-BIND-on-O_K Leg B — HONEST NEGATIVE (engine d7d96fe). Split-prime O_K Dirichlet-character carriers (Kronecker χ_d; Heegner ladder d=−67 streak-16 vs d=−163 streak-40, matched deg N=512) DO drop native mutual coherence and Heegner-order it (mean@N=64 random 0.0355 > OK(−67) 0.0153 > OK(−163) 0.0086, the Weil bound) — but it is operationally inert: Ring-3 recall got WORSE (small-period character = spiky spectrum = poor self-unbind), and C2 SimHash Hamming was unchanged (random projection washes out native coherence). Random ±1 stays the carrier.
    • (3) G-R3-ORGANISM-NATIVE — GREEN (engine 1f0f6be). Host float-FFT RIPPED OUT of the live loop: tools/ring3/ok_bind.py routes bind/unbind through native sp_pr_mul via ctypes; D=1024 tiled as a DIRECT SUM of two 512-blocks so the CAP=32 capacity gate did NOT regress. g_r3_dualroute + g_r3_nightshift both GREEN native (dualroute clean-hit/decoy/null; nightshift seals at CAP=32 @ D=1024, recall@1 all, 349.8 MB demoted to Optane / 16.3 KB index; D=128 gate-fires-before-cap).
    • (4) G-R2-FROB — GREEN (engine dbe4103, rank-2 d076797, tools/curator/frob_episode.py). The Frobenius π^k INTEGER Ring-2 episode store (Theorem-T4 storage form): per-(layer,channel) Frobenius-scaled integer coords; a 2-step rank-2 O_K lattice (coarse coord a + error-feedback residual coord b, x=a·s_a+b·s_b, REAL scales — NOT the literal complex ω, which would inject an imaginary part). Schemes a16(16b)/a8b4(12b)/a16b8(24b). Fidelity: a16 relL2 3e-5 (effectively lossless, 2.0× store); a16b8 24b relL2 1.2e-7 SUB-ULP (18% byte-exact, 0.76× store); a16b16 32b 8e-11 (98.9% byte-exact). T4 confirmed operationally: the per-tensor π^k scale is FREE (no propagation), the integer episode replays clean on 12B. HONEST SCOPE: the n_scored=42 wiki.tiny SP_REPLAY PPL gate is BLIND below ~1% (small-N tie-flip: 24b@1e-7 swings like 8b@8e-3, non-monotonic). Only the float-exact replay reads a clean +0.000% (== baseline 4.6665). "Lossless" is established by reconstruction fidelity, not by an n=42 PPL — no fake +0.000% was manufactured. The compression lever is BIT-WIDTH (12b 2.86× / 16b 2.0× / 24b 1.36× vs float32).
    • (5) G-R2-FROB-ENTROPY — NEGATIVE (engine e6d17bb). Lossless entropy coding (lzma/zlib) on the Frobenius codes is dead weight (a16b8+lzma 1.02×; the int8 residual is incompressible high-entropy; transpose/delta layouts don't help). The earlier 2.56× on dense M was a red-herring (leading-zero int32 bytes).
    • (6) G-R3-MOBIUS — NEGATIVE (engine 1e70763). Möbius square-free compression of the dense Ring-3 M FAILS (M is 99.6% dense, no multiplicative redundancy; divisor-reconstruction error 1.35× the signal; masking sheds memories, recall 1.000→0.969@N=32). The 6/π² density is real but does not transfer to a holographic vector.
    • (7) G-XBAR-ORGANISM-FULL — GREEN (engine 15e7051, tools/ring3/g_xbar_organism_full.py). The FULL loop end-to-end on REAL episodes — ep_audio (EAR) + text decoys ep_wiki/ep_toy. C2 256-bit sig separation (audio self 256, decoys 147/129 << TAU_BITS=168); nightshift native integer bind shadow-gate recall@1 all; dualroute audio cue → ep_audio TOP-1, C2 Hamming verify ACCEPT audio / REJECT text (cross-modal verify = SIGNATURE, not text-PPL, since audio is foreign to any text scoring context); LAND = Frobenius a16b8 integer store decoded → float sub-ULP; METAL = SP_REPLAY of the integer-decoded ep_audio into the 12B resident cache, checks=5 fails=0 (clean inject; PPL 88.89 = foreign-by-design). Continuous audio → discrete integer memory → continuous KV back out, autonomous.
    • (8) G-T2-WEIGHTS — NEGATIVE on T2's own object (engine ac76c8e). T2 Möbius tested on the real gemma-4-12b embed_tokens (V=262144 E=3840): the Möbius transform puts 43.6% energy on non-square-free (claim ~0%); composite-row reconstruction cos 0.032 vs random 0.039 (worse than random). Trained embeddings have NO multiplicative index structure (token IDs are BPE merge ranks). For the record: T2 (Möbius) was a DESIGN PROPOSAL in the theory paper, never empirically validated, unlike T4 (Frobenius cancellation, validated 6-sig-fig on Gemma3-1B).
    • (9) G-PERIOD6-REBASE — GREEN (engine d2d7ceb). C2/Ring-3 content-hash period 8→6 to match the TRUE gemma4 SWA geometry (cuda_forward.cu: globals = Li%g4_period==g4_period-1, g4_period default 6 → {5,11,17,23,29,35,41,47}). The old PERIOD=8 hashed mostly NON-global layers. Rebased 9 sig-pipeline files; re-gated GREEN, separation cleaner (decoy 154→129). The gates were robust to the layer subset as predicted — the period-6 rebase named in the prior session's NEXT is now CLOSED.
    • THE BOUNDARY THESIS (the session keystone): the O_K substrate's value is exact arithmetic — the indestructible algebraic CONTAINER (bind, Frobenius integer store, reduction-order immunity). Every attempt to impose number-theoretic STRUCTURE onto the high-entropy CONTENT was measured-inert (Leg B carriers, Möbius-on-M, entropy-on-codes, T2-on-weights). Edge-of-chaos: unstructured chaos bound inside rigid algebraic order — wins on the container, never on the content. The host-numpy VSA → native Z_q/NTT port (the long-deferred follow-on) is now DONE (Leg A + ORGANISM-NATIVE).
  • [CLOSED GREEN — 2026-06-18] THE BYTE-EXACT FORWARD — the gemma-4-12B forward path now runs EXACT-INTEGER end-to-end (auditability, NOT compression). The same exact-integer O_K container that holds the memory tier now holds the forward pass itself: the four nonlinear fp32 "islands" (RMSNorm / softmax / GELU-tanh / RoPE) and the attention (Q·K / p·V) convert to exact-integer arithmetic, so logits become bit-identical across reduction-order and machine. Full design + gate record: CONTRACT-BYTEEXACT-forward.md §5.1/§5.2/§7/§8. Engine 69c0588, math-core submodule d9d96f3, lattice contract 2751407. Receipts: engine tests/fixtures/xbar_r3/{G-ISLANDS-Q-REF,G-BYTEEXACT-ISLANDS-CUDA,G-BYTEEXACT-FORWARD-12B,G-WIRE-CUDA-GEMMA4,G-WIRE-CUDA-DECODE-GEMMA4}.log. Byte-exact = EXACT ARITHMETIC / cross-machine determinism (the auditability mission), explicitly NOT a compression lever (the compression axis is CONVICTED — CONTRACT-BYTEEXACT §0/§1).

    • The frozen container. Dual-prime negacyclic CRT-NTT, the SAME primes/Garner as the memory ring: q1=1073738753, q2=1073732609, Garner inv 894602413, M=q1·q2≈2⁶⁰ — M fits u64 ⇒ no __int128 anywhere (device __umul64hi for wide products, a 64-bit-split isqrt for the RMS numerator, device CORDIC for RoPE angles). One substrate carries both the XBAR bind and the forward's exact attention.
    • The four integer islands (the genuinely-new nonlinear piece). The byte-exact linear algebra (dual-prime Barrett / mod-q matmul / Garner CRT / NTT ladder) was already built + bit-exact-gated in the universal Rust crate engine tools/sp_dsp_smoke (L2 orchestrator + scalar reference; sp_barrett_oracle.rs / sp_matmul_q_ref.rs Q1_INV_MOD_Q2=894602413 / sp_ntt_0..5b); the islands were the missing piece, now added as sp_islands_q_ref.rs (+ math-core core/exact_islands/). G-ISLANDS-Q-REF GREEN (host x86, no DSP): RMS 5.8e-6 / softmax 1.3e-6 / GELU 2.8e-6 / RoPE 9.2e-6, all order-immune; RoPE via fixed-point CORDIC. (Lesson banked: grep the crate before re-deriving byte-exact arithmetic — the ATTN-NTT/bx_* prototypes re-derived proven code. See lessons.md + [[reference-byteexact-already-in-rust-crate]].)
    • On-12B gates (real gemma-4-12B, gemma4-12b-b1.sp-model, RTX 2060). G-BYTEEXACT-ISLANDS-CUDA GREEN (engine b93f157): the crate *_q_ref references agree with the live CUDA float-island outputs on REAL layer-24 activations — RMS 3.84e-5 / GELU 8.18e-7 / RoPE 9.62e-6 (softmax gated offline 1.3e-6), all < 1e-4. G-BYTEEXACT-FORWARD-12B GREEN (the wire-in, behind default-off SP_BYTEEXACT): LEG A (off) = PPL 4.6665 == baseline BYTE-IDENTICAL (the null floor — the one-shot gemma4_decode_cuda stays byte-untouched), LEG B (on) = 4.6569 PPL parity (−0.21% on n=42; the integer islands + integer attention cost nothing measurable), and run-to-run BIT-IDENTICAL (the integer reductions are order-immune — the cross-machine bit-identity proxy).
    • Driven by the universal daemon (the L1 ABI grew a verb). The universal sp_daemon drives the real 12B prefill through the existing forward-backend hook — G-WIRE-CUDA-GEMMA4 (engine eee3aac: cuda_forward_count 0→1, wire_cuda_active:true) — and now token-by-token DECODES it through a new persistent-KV L1 verb sp_session_register_kvdecode_backend (open/prefill/decode_step/rewind/position/close) — G-WIRE-CUDA-DECODE-GEMMA4 (submodule d9d96f3 → engine 6b9a786): 32/32 tokens bit-identical to the gemma4_kv_decode oracle, VRAM flat (O(1) cache). The verb is additive/append-only (the prefill-only sp_session_register_forward_backend is unchanged; ABI grew, didn't break) — documented as PPT-LAT-L1-ABI-v0 §6b. .sp-model OK_Q4B loader reconciled (decision B, engine e9fb9b0): the crate consumes the engine's resident weights, OK_Q4B is decoded engine-side (PPT-LAT-SP-MODEL-v0 §6.5).
    • HONEST NEGATIVES (kept attached): the ONLY remaining item is the external 2-physical-GPU bit-identical check (the reduction-order immunity is the on-one-machine proxy; a true two-GPU diff is out-of-band); the LEG-B PPL parity is n=42 single-chunk (deterministic, but small-N — [[feedback_small_n_deflection_illusion]]); and the boundary thesis stands (byte-exact buys auditability, not compression). Public: Papers 19/20/21 + LEDGER.md in Position_Is_Arithmetic (19f3d8c); PPT-ARM-System §8 islands CLOSED.
  • [CLOSED GREEN — 2026-06-20] B3-WC — THE AUTONOMOUS-RECALL CAMPAIGN RESOLVED: a learned W_c head does LIVE instance-level episodic recall on the served gemma-4-12B chat. The chat (CONTRACT-CHAT-FULLSTACK, already coherent/byte-exact/O(1)/single-entry GREEN) now has an autonomous librarian: the operator states a fact, the curator stores it as an episode, and on a later query the head selects the right episode — or refuses if none is relevant. Engine edc8079 (pushed); receipts engine tests/fixtures/chat_fullstack/G-CHAT-B3-WC-{DEPLOY,DIV2,DIV}.log; full run-record CONTRACT-CHAT-FULLSTACK.md.

    • The whole arc is a boundary-thesis win, honest negatives intact. ALL hand-designed relevance signals FAIL open-world — 6 verifiers + 4 Disposer signals (Yes/No bridge, Δ-continuation, multi-token ΣΔLL, consensus) + cosine-qK + ΔLL-polarity, each dominated by per-episode/shared-attractor bias at N=3 wiki (best = raw q·K + argmax multitok-ΔLL = 2/3 rank; the two "pristine" normalizations made it worse). Commits 4dba6c8/2b623ab/acd7b3a. ROOT CAUSE (the ablation oracle): wiki facts are PARAMETRIC (the 12B regenerates them from weights) → episodic dependency is unmeasurable → the corpus was the wall (B3-v10 27038d3, N-invariant 485b5b1).
    • THE UNLOCK — teacher-forced ablation knockout on a NOVEL needle. SP_B3_SECRET teacher-forces the known secret tokens + cudaMemset-ablates exactly their source KV rows, scores ΣΔLL with/without. B3-v12 (15738c1) novel −33.56 vs parametric −0.15 (~220×); B3-v13 (b6470cc) 3-archetype matrix code −33.56 / contradiction −18.58 / relational −16.10 vs parametric band [−0.15,+1.45], TAU pinned −8.0, ~16-nat gap. The gate 9 prior signals couldn't build = a SHIPPABLE oracle, and a PERFECT LABELER (perfect-diagonal separable matrix, 7556d04).
    • CORPUS + HEAD. Split-entropy minter (mint_corpus.py 52ec468; secret stated ONCE) + single-pass ep.secret admission oracle (adfcb06) scaled 200 needles + 60 foreign null-class (200/200 accept, control reject, f4166c7). Learned W_c (HD=512→r=32; relevance = logsumexp-over-positions then mean-over-heads; InfoNCE over [E episodes + NULL/s0] + reject hinge). Templated → 34% instance top-1 (carrier collision); DIVERSITY FIX (mint_corpus_v2 72d9c77, f62e6ef): 90 unique-subject needles → instance top-1 34%→100%. Reduction discovery: max=12/361, top-8-mean=16/361, logsumexp-mean=361/361, int16==f32 for all three (the int16 "degradation" was a wrong-reduction bug, NOT quantization).
    • G-CHAT-B3-WC-DIV2 (87044d8, deploy gate): (E+1)-argmax over [episodes, NULL=s0] = 360/361 instance recall + 50/50 foreign reject, int16-exact, s0=+0.102. G-CHAT-B3-WC-DEPLOY (edc8079, LIVE): recall.rs WcHead/wc_score (numerically-stable LSE) + routes.rs SP_B3_WC branch (score all episodes → NULL-argmax → replay winner @SP_REPLAY_MTARGET=42 or clean prompt; default-off = null floor). Live on the 90-needle registry: matched → RECALL ep_n_div_000 (9.858); foreign "capital of France" → whole population negative → NULL → clean "Paris." Blob _b3_wc/wc_deploy.bin (WCB1 hd512 r32 s0=+0.1021 sscale=0.17678) via tools/xbar_lsh/export_wc_deploy.py; launcher run_console_recall.bat (eb6332d, PMAX 4096 82d77c8); console stop-button (fc25b07).
    • BOUNDARY-THESIS extension: autonomous recall is won by a LEARNED head on a DIVERSE corpus — NOT hand-designed number-theoretic/geometric signals (every one a measured negative). Corpus diversity (not machinery, not normalization, not more-memory) was the binding constraint at every granularity; novel-vs-parametric is the only measurable episodic-dependency axis. B4 NIGHTSHIFT (between-turn consolidation: turn→curator-mint→ablation-admit→hot-append; W_c needs no retrain to score a new episode) = DEFERRED, pre-scoped (SESSION-HANDOFF §0d).
  • [CLOSED — HONEST-NEGATIVE — 2026-07-01] G-T4-WEIGHTS: T4 Frobenius π^k of the model WEIGHTS is REDUNDANT vs OK_Q4B (as a compression lever). The named open frontier is now resolved by a pre-registered kill-test (papers/PPT-LAT-T4-WEIGHTS-SCOPE.md; probe engine tools/t4_weights_probe.py on the bf16 source; receipt engine tests/fixtures/t4_weights/G-T4-WEIGHTS.log). On real 12B tensors (3/3: L0 q_proj, L0 down_proj, L23 o_proj), OK_Q4B (per-32-block int4+f16, 4.5 eff bits/w) gives relL2 ≈ 0.10–0.12; the Frobenius "free" per-tensor π^k scale at 4 b/w gives relL2 ≈ 0.40 (~3.3–4× worse for 0.5 bits saved), per-row ≈ 0.21–0.28, and matching OK_Q4B fidelity requires the scale to shrink back to per-32-block (where the codec is identical to OK_Q4B). The rank-2 residual lowers error only by spending MORE bits. The free-scale property that holds on Ring-2 EPISODES (G-R2-FROB, sub-ULP at 24b) does NOT transfer to trained WEIGHT tensors (per-block dynamic range / outliers the block scale already captures). Resolves the roadmap-vs-CLAUDE.md tension in favour of "CONVICTED redundant vs OK_Q4B." No overclaim: this refutes T4 as a weight-compression lever, NOT the T4 exact-cancellation property (real, needs Q8, buys auditability). Boundary thesis extended to the weights (like T2/Möbius in G-T2-WEIGHTS). (Prior framing, now superseded: "T4 of the weights is the validated untouched NEXT lever.")

  • NEXT trajectory (post-T4-weights): with T4-of-weights closed negative, the live open edges are (per Roadmap §2 + the swarm gap): SP-SWARM multi-host deploy (graduate the L0–L4 mesh to GREEN-LIVE), Persistent-KV P2/P3 (~40K sessions → unbounded context via Ring-2 eviction), NIGHTSHIFT criterion-5 (curator on live turns), and the external 2-physical-GPU byte-exact check. Native-C XBAR port stands. Full human+agent synthesis: CURRENT-STATE-OF-PROJECT.md (repo root). Then B4 NIGHTSHIFT (between-turn memory consolidation, pre-scoped) and KAIROS post-organism state. Standing hygiene: the external 2-GPU byte-exact check, cudaEvent journal-tax, gemma4_kv_decode first-token boundary reconcile, compact-slab globals wrap-rewind, the P3.4 larger-N multi-chunk hardening run. Full human+agent synthesis: CURRENT-STATE-OF-PROJECT.md (repo root).

LATENT INTERCEPTOR + TELEPATHY (LatentBridge) — CAPSTONE CLOSED + PARKED (2026-06-30)

The finetuned gemma4-draft body is now a latent-native router, and a tokenizer-free cross-family bridge runs natively in the engine. Built, characterized, gated, and deliberately parked at a safe architectural end-state. Spec: papers/PPT-LAT-TELEPATHY-LatentBridge-spec.md. Four pillars:

  • [PROVEN] (1) The Decision Suite — four hardened, non-hallucinating heads. Tool / Action / Memory / Route, all near-miss-hardened so they never fire on idle chatter: false-fire 0.000 on isolated cross-distribution OOD across the suite; KEEP recall lifted 0.429→1.000. The Route head (LOCAL vs TELEPATHY) is anti-laziness-trained (local-doable code + tool tasks + model-mentions all labelled LOCAL → isolated-OOD 1.000) and governs decide_route in-engine. Gates G-TH-HARD (tool OOD 1.000), G-ACT-HARD (action OOD 0.979), G-ROUTE-WIRE (in-engine route == Python; headless ⇒ Local null floor).
  • [PROVEN] (2) The Geometry — cross-family alignment + ordered positional bandwidth. gemma-3n-E2B ↔ qwen2.5-coder-0.5b via a ridge affine adapter + adapter registry: representation alignment (retr@1 1.000 / roundtrip 0.891), foreign-reject (AUC 0.999), causal generation-steering (steer-acc 1.000, dLL_self +0.414 / dLL_cross −0.496), and a multi-vector prefix read positionally (ORDER_gain corr−shuf = +1.45 nats, corr>shuf 100% — order matters beyond aggregate mass). Gates G-TELEPATHY-ROUNDTRIP / -REJECT / -GEN-TRIGGER / -READABLE-PREFIX.
  • [PROVEN] (3) The Boundary — pure latent fusion degrades; the two-stage architecture is correct. Telepathy is a gist/intent channel, NOT a precise-symbolic one (TELE-11b): fused latent+text FAILS (hybrid 0.000), and the channel loses precise operands while the text delegate is itself capable (text arithmetic 0.806). The cemented architecture is therefore sequential decide_route(latent) → delegate_execute(CLEAN TEXT), never-fuse contract enforced. Gate G-TELEPATHY-TWOSTAGE (TELE-12/13).
  • [WIRED] (4) The Sovereign Engine — native, zero-Python, memory-isolated, fail-closed LatentBridge. tools/sp_daemon/src/telepathy.rs (SP_TELEPATHY, default-off): the bridge object + adapter-load + pure-Rust affine transfer (in-engine == Python, max|Δ|=6.7e-6) + the routing primitive + the fail-closed license gate (license unset ⇒ inert). Gate G-TELEPATHY-WIRE. TELE-14 — the standalone SOVEREIGN native delegate (engine 2f57520): SP_TELEPATHY_NATIVEeagle_accept::run_telepathy_native loads the qwen2.5-coder in its own L1 session and answers a clean-text task entirely in-engine via prefill_chunkdecode_step argmax (the sp_memo_* host-CPU pattern) — coherent in-engine Python answer, ~0.8 tok/s CPU. Gate G-TELEPATHY-NATIVE GREEN.
    • The free co-residency win + the gotchas (banked, MEM-OKF 61b9acd). The working decode is CPU L1, which touches no CUDA g_w — so this also solves live co-residency with the resident 12B with zero contention and zero g_w refactor (the risky path we intentionally avoided). On drop the SpModel/SpSession free with zero residual. Gotchas that cost real time: CUDA qwen3_decode_cuda/qwen3_forward_cuda are strict SP_ARCH_QWEN3-only and reject the QWEN25 coder (no CUDA QWEN25 decode); qwen3_generate_kv on the wire_cuda session segfaults (host weights freed post-upload). HF parity inherited from the top-1-lossless transcode gate (identical-token argmax == oracle) — re-running the Python oracle was deemed redundant and skipped by operator choice. A QWEN3-arch transcode (to unlock the GPU decode) is declined on purpose — it would re-collide with g_w.
    • Boundary-thesis extension: the routing/decision wins were taken by learned heads + clean-text delegation, never by forcing precise content through the latent channel — consistent with the project-wide "win on the container, not the content" thesis.

PARKED. The LatentBridge framework is feature-complete for this milestone: sovereign, native, co-resident, Python-free, fail-closed. Licensing enforcement remains SPEC (fail-closed self-disabling only; no host-external effects). Receipts: engine 2f57520; lattice spec/STATE b82f447+; MEM-OKF 61b9acd.

[PROVEN] O(1) CONVERSATION KV — G-PERSIST-KV GREEN + DEFAULT-ON (2026-06-30)

The served 12B no longer re-prefills the whole conversation each turn. SP_PERSIST_KV is now DEFAULT-ON (engine d211fd2): a follow-up turn reuses the resident cache via longest-common-prefix match of the committed token sequence and prefills only the new suffix (bounded-tail rewind ≤32 inside the SWA journal); SP_PERSIST_KV=0 forces the O(n) re-prefill null floor; the cache-mutating paths (replay / inject / agency-writers SP_DECIDE/SP_FORGET/SP_B4_NIGHTSHIFT / speculative recall SP_B3_JUDGE/SP_B3_DISPOSER/SP_INT2) still reset every turn. The append was found already implemented but dormant during the scope pass (anti-rebuild win — no new C); P1 was gate + flip.

  • G-PERSIST-KV GREEN (tests/perf/_g_persist_kv_gate.py + _persist_gate_compare.py, on the live 12B, temp 0, byteexact, SWA ring W=2048, Pmax=4096): a 6-turn scripted chat is byte-identical across ALL turns with persist off (O(n) re-prefill) vs on (O(1) suffix-append) — matching sha16 + text + char counts. The default-config run (no env) == the off baseline byte-identical, so the flip is safe. The kill-test (persisted KV must equal a fresh re-prefill) holds.
  • O(1) telemetry: TTFT off = [1351,2294,4034,5522,7748,10096] ms (grows 7.47× over 6 turns); on = [1539,984,1074,1142,1418,1347] ms (flat, 0.88×). Mechanism confirmed in the daemon log: reuse 172/172 committed (drop 0); prefill suffix 20 (full would be 192).
  • Scope/limits (honest): one active conversation at a time (KV_COMMITTED is a single global sequence; a chat_id switch LCP-misses → reset); drop_n>32 → reset fallback; both correct (reset = the O(n) null floor). Receipts tests/perf/_persist_gate_{off,on,def}.json; commit d211fd2; MEM-OKF bb2f5a13. Scope doc papers/PPT-LAT-OKV-Persistent-KV-SCOPE.md.
  • Open (P2/P3): larger Pmax + ring tuning for ~40K-token sessions; evict the 8 GLOBAL layers (the O(n) floor) to the existing XBAR Ring-2/Optane demotion for unbounded context.