| type | findings-ledger | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| title | Findings Ledger — the measured constants, layers, levers, boundaries, and test outcomes | ||||||||
| description | The single lookup table for everything the campaign has MEASURED on the metal: which layer carries the fact signal, what every threshold/tau/K/alpha is and where it came from, the architectural boundaries (what works vs what is measured-inert), the substrate constants, and a per-test ledger (corpus + config + outcome + receipt + commit). Read this when you need 'which layer / what tau / what K / where the lever is' without re-deriving. Receipts-first: every number names a gate log + commit. | ||||||||
| tags |
|
||||||||
| timestamp | 2026-07-01 00:00:00 UTC | ||||||||
| resource | shannon-prime-lattice/papers/PPT-LAT-FINDINGS-LEDGER.md | ||||||||
| sp_status | GREEN | ||||||||
| sp_gate | each row names its gate | ||||||||
| sp_commit | engine fc2e846 (attr-gate zero-inference HEAD of this arc); see per-row commits | ||||||||
| sp_repro | each finding cites the launcher + harness + fixture that reproduces it |
Gist. The fact signal is layer-localized (global layer 5). Recall = decide in latent, execute in clean text, never fuse. The cheap levers win: L5 cosine (τ=0.30) selects at 86.89%; a Jaccard/attribute check grounds delivery; the model's own robustness rejects general-knowledge foreign; a deterministic attribute-gate closes the zero-prior hole at zero inference. The heavy generative judge is parked (it earned nothing the cheap levers didn't). This doc is the flat lookup table of the constants and boundaries behind those sentences.
L5 global-layer-5 → §1. tau τ=0.30 Jaccard-0.6 attr_tau-0.5 K=2 TAU=-8 M_target=42 EOT_BIAS=4.0 → §2. never-fuse deciders-dont-execute native-robustness-boundary judge-needs-2 query-token-guard → §3. dual-prime period-6 HD=512 Ring-3 C2-256 → §4. test-ledger receipts → §5. honest-negatives measured-inert → §6.
| what | value | how measured | receipt / commit |
|---|---|---|---|
| The fact signal lives in global layer 5 | L5 exact→paraphrase recall@1 = 85.2%; all-8-global-layer average = 11.5% (averaging dilutes the signal ~7×) | per-layer sweep of the global-Q representation vs a Jaccard oracle | G-REP-LAYER-L5 |
| L5-cosine query-key recall (offline) | 100% exact / 88.5% paraphrase (vs Jaccard 100% / 8.2%) | cosine of L5 query-embed vs stored L5 key | G-REP-LAYER-L5 |
| L5-cosine recall LIVE (served 12B) | G-L5-OBEY-REPRO)systemecho delivery). 54/61 == selector top-1: every correctly-selected episode obeyed; misses = selection cross-picks. Delivery solved; obey ceiling = selector. Law: obey receipts carry build SHA + canary + full env dump |
live /v1/chat, SP_RECALL_L5 + SP_RECALL_L5_PROMPT=systemecho |
G-DELIVERY-SWEEP + G-ONECONFIG-LIVE, engine 8ae343b |
SP_RECALL_L5_PROMPT |
systemecho (fact as SYSTEM authority + verbatim-echo priming; multi-turn preserved) |
delivery wording is a MEASURED lever: plain 40.98 < scaled 63.93 < sandwich < factecho/system < systemecho 88.52 (0 leak) | G-DELIVERY-SWEEP |
SP_RECALL_QONLY |
1 (canonical config) | non-interrogative turns skip the L5 stage (in-registry cos background ≥0.9 otherwise injects an irrelevant fact) | G-RECALL-QONLY-LEXICAL 188/188 + G-ONECONFIG-LIVE Q-phase |
| gemma4 global layers | period 6 → global layers {5, 11, 17, 23, 29, 35, 41, 47}; n_global = NL/PERIOD = 8 |
arch (content-hash period rebased 8→6 to the true global layers at G-PERIOD6-REBASE) | d2d7ceb |
| L5 query-embed construction | mean over G_NH=16 heads of global layer 5 (last token), L2-normalized → 512-d (HD=512) |
recall::l5_query_embed; L5_LAYER=5 |
engine recall.rs |
| observed L5 cosine range (private entities) | exact-match query ≈ 1.000; same-entity different-attribute query ≈ 0.93–0.95 (the shared entity token dominates the embedding) | SNE serve log | G-SNE-CRUCIBLE-L5DIRECT |
Takeaway: select on L5 alone; never average the global layers. The entity/subject dominates the L5 embedding — which is why a same-entity mismatched-attribute query still scores ~0.95 (see §3 boundary).
| lever (env / const) | value | role | receipt |
|---|---|---|---|
SP_RECALL_L5_TAU |
0.30 | L5-cosine recall gate; below-τ = no recall (silently rejects genuine out-of-corpus foreign) | G-L5-RECALL-LIVE, G-HARDFOREIGN-L5DIRECT |
recall::OVERLAP_THR (Jaccard) |
0.6 | the production faithfulness verifier — token-overlap of the cited span vs the fact | G-FAITHFUL-RECALL-JACCARD (15/15) |
SP_RECALL_ATTR_TAU |
0.5 | attribute-gate: if ≥50% of the query's salient words are ABSENT from the fact → decline | G-SNE-ATTRGATE |
query_has_entity_token guard |
token len≥4 AND contains a digit |
fires the attribute-gate ONLY for private-entity queries (paraphrase-safe) | G-ATTRGATE-GUARD-PARA |
SP_B3_JUDGE_K |
2 (minimum working) | judge candidate count. K=1 breaks the reject (single candidate → always-PICK, NULLs=0); K=2 = 83% clean-reject (== K=8, cheaper) | G-JUDGE-KSWEEP-K2 |
causal-ablation admit oracle TAU |
−8 | W_c admission: novel needle ΣΔLL ≈ −33.56 vs parametric ≈ −0.15 → clean separation | B3-v13 15738c1 |
SP_REPLAY_MTARGET |
42 | KV-replay injection-mass clamp (recall replay winner) | G-CHAT-B3-WC-DEPLOY |
SP_EOT_BIAS |
4.0 | end-of-turn logit bias (clean turn-stop on the served chat) | chat_fullstack |
| Diffusion Judge (OOD proxy) | recall 94.4% / reject 98.0% | won the OOD kill-test as a proxy; NOT the production path (native judge plateaus ~50%) | G-DIFFJUDGE-OOD-H2H |
| C2 SimHash bit-count (similarity overlay) | 256 (frozen C2) → recall@1 0.607 / recall@5 0.885; 512/1024/2048 → @1 0.72/0.82/0.87 | wider sigs recover top-1 toward L5-cosine (0.885); 256 = the shortlist-hint default | G-SWARM-C2-SEMANTIC |
The governing law (ADR-002): DECIDE in latent, EXECUTE in clean symbol/text, NEVER fuse latent content into generation, and a decider must not execute. The measured boundaries:
| boundary | finding | evidence |
|---|---|---|
| Never fuse | fused latent+text transmit = 0.000; sequential decide→execute is the win | TELE-12/13 |
| Deciders don't execute | when the judge delivered its own pick it lost context-authority (right pick, parametric answer); routing delivery through the judge session costs recall (50% vs 86.89%) | G-JUDGE-KSWEEP-K2, authority fix 4ff75d1 |
| Judge needs ≥2 candidates | at K=1 the generative judge degenerates to always-PICK (no reject) | G-JUDGE-KSWEEP-K2 |
| Native robustness protects GENERAL knowledge, not PRIVATE data | general-knowledge hard-foreign (Athens/silver/…) = 0/18 spurious even when a mismatched fact is delivered (the model has priors); zero-prior novel entities = 80% confabulation + 5% leak (no priors to fall back on) | G-HARDFOREIGN-L5DIRECT, G-SNE-CRUCIBLE-L5DIRECT |
| tau + robustness ≥ the generative judge on reject | judge = 0 benefit over L5-direct+τ (0/18 == 0/18) AND PASSed 15/18 hard-foreign (failed its own reject job) → PARKED | G-HARDFOREIGN-JUDGE |
| Attribute grounding is a query property, not a query∩fact property | gate on whether the QUERY carries a private-entity token (a wrong retrieval must NOT lower the shield); a shared-token guard broke SNE decline to 2/6 when L5 delivered a wrong same-structure entity | G-SNE-ATTRGATE-GUARD |
| Zero-inference decline | on attribute-absence the reject streams a fixed string with NO gemma4 forward → confabulation/leak mathematically impossible + microsecond latency | G-SNE-ATTRGATE-ZEROINF (16/16 "no gemma4 decode") |
| Byte-exact ENVELOPE GAP on the served path (2026-07-02) | served chat already runs byte-exact islands+decode-attention per request (byteexact default-true → gemma4_kv_byteexact_set); SP_BYTEEXACT env = no-op there (61/61 byte-identical A/B). The residual float surface = prefill cuBLAS GEMMs — the un-pinned surface behind the para-obey receipt divergence. Name the envelope in every byte-exact claim. |
G-BX-OBEY-AB |
| Blunt closed-book prompt over-declines | SP_RECALL_STRICT made the model refuse even valid matches (0/4) — dead lever |
G-SNE-STRICT-OVERDECLINE |
| A speed ladder lives in COMPILE FLAGS, not just code (2026-07-02) | the engine's standard build-cpu math-core libs compile WITHOUT /openmp /arch:AVX2 → any binary linking them silently loses the brick-4/5/6 qwen36 ladder (served 35B ran 3×-slow; CPU-only A/B = 0.17 tok/s = exactly the pre-OMP rung). Perf-critical exes must link build-cpu-perf (+ LLVM libomp.lib — MSVC 14.29's lacks __kmpc_dispatch_deinit). Two enums trap in the same campaign: L1 wire arch_id (QWEN36=8) ≠ internal sp_arch_t (QWEN36=4). |
G-QWEN36-SERVE |
| CUDA kernel launches fail SILENTLY — and Pmax-derived shared memory is the trap (2026-07-03) | at SP_DAEMON_KVDECODE_PMAX=20000 the float decode attention launches with Pmax*4B = 78KB dynamic shared > Turing's 64KB max → cudaErrorInvalidConfiguration on EVERY launch, invisible at stream-sync → stale attention out = the "float path garbage" (the S1-era fragility narrative is CORRECTED for this config: the float attention never ran). One-shot loud launch-failure telemetry now permanent at both attention call sites. |
G-12B-SERVE-ROOTCAUSE |
| VRAM oversubscription on WDDM taxes EVERY kernel launch, body-independent (2026-07-03) | Pmax-sized KV caches + 8GB weights ≈ full 12GB card → the driver's residency management adds ~0.6 ms per attention launch even for an immediately-returning kernel (probe8 == probe16 == 9.3s; kernel arithmetic was ~1s of an 18s prefill). Effective serve collapsed to 0.02–2 tok/s. PMAX 20000→4096 (resident) = coherent BOTH decode paths, 22.6 tok/s short-ctx, 13 @340. Law: "GPU busy but slow" + flat per-launch cost ⇒ check residency BEFORE kernel math; decode-only tok/s receipts do not describe the operator experience (min of decode/prefill/recall-chain does). |
G-12B-SERVE-ROOTCAUSE |
| The recall stack's turn cost multiplies prefills (OPEN, pre-scoped) | per recall turn: the dead B3-v2 q·K scan (TAU=inf ⇒ can never fire) runs over all 81 episodes × full-conversation positions, TWICE (base + rebuilt delivery prompt), and systemecho delivery full-re-prefills (system-fact-first breaks prefix reuse) → turn-4 prefill 86.9s even at PMAX=4096. Fix brick: skip the dead scan; don't re-run the recall pipeline on the delivery prompt. | G-12B-SERVE-ROOTCAUSE §C |
| constant | value |
|---|---|
| dual-prime CRT-NTT primes | q1 = 1073738753, q2 = 1073732609 |
| CRT modulus | M = 1152908312643096577 (≈2⁶⁰, fits u64 ⇒ no __int128 in the ring) |
| O_K | Z[(1+√−163)/2] (class number 1) |
| content-hash period (C2/Ring-3 ↔ global layers) | 6 |
| C2 signature | 256-bit (Rademacher ±1 projection router) |
| Ring-3 VSA dimension | D = 1024 (two 512-blocks); superposition CAP = 32 |
| byte-exact forward | SP_BYTEEXACT (default-off = byte-identical null floor); 4 nonlinear islands (RMSNorm/softmax/GELU/RoPE) on the same CRT-NTT |
| MEM-OKF address classes | content-addressed = sha256(norm(body))[:16] (agent facts, tamper-evident by re-hash); C2-addressed = 256-bit C2 SimHash via --addr (episodes, not body-re-hashable; L3 Ed25519 for cross-node authenticity). Store split: 87 content / 26 c2 |
| CRLF cross-platform interop | Python-on-Windows WRITES MEM-OKF as CRLF but computes addr over the LF-normalized body (text-mode read translates CRLF→LF). A Rust/other node MUST normalize line endings on read (parse_fm normalizes first) or every cross-node address mismatches. |
All receipts under shannon-prime-system-engine/tests/fixtures/chat_fullstack/ unless noted; served 12B (OK_Q4B) on the RTX 2060, wire_cuda_backend.
| gate / test | corpus + config | outcome | commit |
|---|---|---|---|
G-L5-RECALL-LIVE |
61 faithful facts, paraphrase queries, SP_RECALL_L5 τ=0.30 |
86.89% paraphrase obey (default-off) | d9099cd |
G-JUDGE-KSWEEP-K2 |
12 in-mem para + 12 foreign-v2, SP_B3_JUDGE_L5 |
K=1 reject broken; K=2 = 50% recall / 17% spurious (83% clean-reject); recall-thru-judge is the cost, not K | 4ff75d1, 5d336c7 |
G-JUDGE-REORDER-K2 |
same 12/12, reorder (judge veto → L5-#1 deliver) | 50% recall / 17% spurious — preserves L5 recall, doesn't beat L5-direct+τ | 61160e9 |
G-L5-DIRECT-SAME12 |
same hard-12 baseline | 50% recall / 8% spurious (the hard-12 is a hard slice of the 86.89% full-61) | 61160e9 |
G-HARDFOREIGN-L5DIRECT |
18 same-domain high-cosine UNANSWERABLE (general knowledge) | 0/18 spurious — native robustness absolute | 8dbbcdc |
G-HARDFOREIGN-JUDGE |
same 18, judge@K=2 | 0/18 spurious BUT judge PASSed 15/18 → judge PARKED | 8dbbcdc |
G-SNE-CRUCIBLE-L5DIRECT |
20 synthetic novel entities (uuid ids, high-entropy override codes, 0 parametric prior; audited) | MATCH recall 20/20=100%; MISMATCH confab 80% / leak 5% / decline 15% — the vulnerability | 45149a1 |
G-SNE-STRICT-OVERDECLINE |
SNE + SP_RECALL_STRICT |
0/4 match (over-declines) — STRICT rejected | f161a27 |
G-SNE-ATTRGATE |
SNE + SP_RECALL_ATTR_GATE τ=0.5 |
100% recall + 100% decline, confab 80→0, leak 5→0 | f161a27 |
G-ATTRGATE-PARA-REGRESSION |
fct paraphrase + attr-gate (no guard) | 0/12 — lexical gate over-declines paraphrase (needs guard) | f161a27 |
G-SNE-ATTRGATE-GUARD / G-ATTRGATE-GUARD-PARA |
SNE + fct, attr-gate + query-token guard | SNE 100%/decline preserved; paraphrase 6/12 = baseline, 0 over-decline (globally safe) | a4ebe3d |
G-SNE-ATTRGATE-ZEROINF |
SNE + zero-inference symbolic decline | 16/16 match+decline; all declines "no gemma4 decode" (hallucination-immune reject) | fc2e846 |
G-KSTE-MD / G-KSTE-MD-REALDATA |
magnitude-depth encoder + Dickson σ0⊕σ1 | 37.6× synthetic discrimination but INPUT-GATED (directional signal → magnitude-shape-blind on real global-Q) → honest negative | 104-109 (this session) |
G-SWARM-REPLICATE-CONVERGE |
2 divergent store-dir "nodes" over the real 113-object MEM-OKF | content-address round-trip + have/want convergence (113/113 byte-identical) + verify-on-arrival + idempotence + tamper-reject; SP-SWARM L1+L2 core, transport-agnostic | lattice (this session) |
G-SWARM-PROVENANCE-ED25519 |
Ed25519 (libsodium/PyNaCl) sign-on-write + verify-on-pull vs invite-only roster, over real MEM-OKF objects | signed content+episode commit; tampered-episode→sig-invalid, stripped→unsigned, forged→sig-invalid, unrostered→untrusted-signer, tampered-content→integrity-fail (all rejected pre-commit); C2 episodes now tamper-evident cross-node; SP-SWARM L3 | lattice (this session) |
G-SWARM-RUST-PARITY |
Rust tools/sp_swarm (sha2 + ed25519-dalek) vs pynacl-signed fixture + real store |
6/6: Rust reproduces Python addresses (89 content), ed25519-dalek verifies the pynacl signature over addr‖body, tamper+roster reject; cross-lang byte-parity | engine (this session) |
G-SWARM-TRANSPORT-QUIC |
2-node localhost over quinn/rustls; Ed25519 mutual roster auth | A↔B bidirectional convergence (pull 3 / pull 2), tampered object rejected on arrival (integrity-fail), off-roster peer dropped (0 objects); SP-SWARM L0 (reused engine QUIC, not rust-libp2p) | engine (this session) |
G-SWARM-NODE |
2-node localhost, run_node autonomous periodic sync |
converge 5/5 both directions; persistent identity stable across reloads; roster file parsed; SP-SWARM integration orchestration | engine (this session) |
G-SWARM-DAEMON-WIRE |
cargo build --features wire_cuda_backend,swarm |
sp-daemon builds+links with the mesh wired (19.33s); SP_SWARM=1 spawns run_node, unset=no-op null floor; SP-SWARM daemon integration |
engine (this session) |
G-SWARM-C2-INDEX |
synthetic near/far 256-bit sigs | C2Index find_similar top-k Hamming: near>far, exact=0, monotone, hex round-trip; L4 index mechanics | engine (this session) |
G-SWARM-C2-SEMANTIC |
61 paraphrase L5 embeds → nearest episode; SimHash vs L5-cosine | C2-256 recall@1 0.607 vs cosine 0.885 (retains 69% — weak top-1) BUT recall@5 0.885 = cosine's top-1 (strong shortlist); bit-count lever (512/1024/2048 → 0.72/0.82/0.87 @1). L4 = hint/shortlist, not top-1 | engine (this session) |
G-SWARM-GOSSIP-DISCOVERY |
2-node localhost QUIC; SIM shortlist gossip + exact-fetch verify | A discovers a B-only object via C2 shortlist → exact-fetch (accept L1+L2+L3) → converge; decoys A holds skipped; off-roster rejected; k≥5. SP-SWARM L4 network discovery | engine (this session) |
-
KSTE-MD / Friedman magnitude-depth router: 37.6× synthetic discrimination collapses on real global-Q (input-gated, directional-blind). Parked as a dedup/eviction primitive, OFF the recall path.
-
C2-256 SimHash as a TOP-1 retriever: 0.607 recall@1 (retains only 69% of L5-cosine 0.885) — honest-negative for standalone discovery. VIABLE only as a top-k shortlist→exact-fetch hint (recall@5 0.885, its designed §6 role). Boundary thesis again: the quantized structure-on-content signal is a hint, not the answer.
-
Query-side foreign reject (L5 / Jaccard / margin thresholds): ~90% false-accept on clean-v2 foreign — cheap query-side signals cannot reject; the model's robustness + τ do.
-
Generative judge (SP_B3_JUDGE): 0 benefit over L5-direct+τ on hard-foreign, PASSes 15/18 → PARKED.
-
STRICT closed-book prompt: over-declines valid matches (0/4) — dead lever.
-
Shared-token attribute guard (query∩fact): broke SNE decline (2/6) on wrong-entity delivery → replaced by the query-token guard.
-
Fused latent+text (Telepathy precise channel): 0.000. CRT multi-device: loopback-only (resolved-negative pending 2-GPU). Möbius/entropy-coding on M / T2-Möbius embedding: measured-inert (see STATE boundary thesis).
-
T4 Frobenius π^k of the model WEIGHTS as a compression lever (
G-T4-WEIGHTS, 2026-07-01): REDUNDANT vs OK_Q4B. On 3/3 real 12B tensors, OK_Q4B (per-32-block int4+f16, 4.5 eff bits/w) = relL2 0.10–0.12; the Frobenius "free" per-tensor scale at 4 b/w = 0.40 (~3.3–4× worse to save 0.5 bits), per-row 0.21–0.28; matching fidelity needs the scale back at per-block (== OK_Q4B). The free-scale property that holds on Ring-2 EPISODES (G-R2-FROB, sub-ULP@24b) does not transfer to trained weight tensors (per-block outliers). Refutes T4-as-weight-compression, NOT the T4 exact-cancellation property (needs Q8, buys auditability). Scope+receipt:PPT-LAT-T4-WEIGHTS-SCOPE.md, enginetests/fixtures/t4_weights/G-T4-WEIGHTS.log. Boundary thesis extended to the weights (as T2 was). -
GEODESIC pre-flight (2026-07-03, ADR-003 v2): the faithfulness steering field is CONSTANT — Tier A. F3 pairs (61 fct + 20 SNE,
G-F3-CAPTUREengine6c03996): x0 = clean-parametric vs x1 = systemecho-delivered post-output_norm state, last-prompt-token frame.G-FLOW-STRAIGHTNESS(engine1b3c234, pins pre-registered same day): topPC(unc) 0.731 / cos-to-mean 0.839 (pins 0.70/0.70), centered topPC 0.218 ⇒ the structure IS the mean vector; Tier-B legs also pass (kNN-conflict 0.796, ridge 1-step rel-err 0.204) so a 1–2-step FM head stays viable in reserve; robust on fct-only/sne-only/no-echo slices; frame-1 (first-answer-token) NOT constant (0.293/0.478) — steer at the DECIDE state. Δ-norm mean 173.6 (min 118.4), never degenerate. Honest negative: directional coherence does NOT separate fct/sne (0.846 vs 0.816) — no free zero-prior detector; the attr-gate stays the shield. Consequence: the Faithfulness Head's first form is ONE steering vector (mean v), gateG-FM-STEER-OBEYvs systemecho 88.52%/0-leak.* Receipts: enginetests/fixtures/chat_fullstack/G-F3-CAPTURE.log+G-FLOW-STRAIGHTNESS.log; data_faithful_corpus/f3/{A,B}(tracked). -
Pre-head steering with the faithfulness v̄ is an HONEST NEGATIVE — the final-norm surface expresses the obey decision, it does not make it (2026-07-03,
G-FM-STEER-OBEY, enginedd9dbbb). New layer-agnostic rail:gemma4_kv_steer(persistent per-step axpy α·v onto the pre-head hidden;SP_STEER_VEC/SP_STEER_ALPHA; default-off null floor). With the Tier-A v̄ under plain delivery: positive α monotonically degrades obey (slice-16: baseline 9/16·6-leak → α0.5 6/16·9 → α1.0 4/16·10), α−0.5 == baseline exactly. Mechanism (three receipts converge): Tier-A constancy at the final norm + frame-1 non-constancy + TELE-2's ~16–22 seam ⇒ the obey choice is made upstream in attention over the fact text; the final-norm delta is the phrasing/expression of the already-made choice. Final-norm steering CLOSED; escalation = per-layer tap at the TELE-2 seam (rung 2), per-layer sweep (rung 3, the L5 methodology). Receipttests/fixtures/chat_fullstack/G-FM-STEER-OBEY.log. -
Obey/leak is NOT linearly readable from the final-norm F3 states — the Tier-A ray is the TREATMENT SIGNATURE, not an outcome axis (2026-07-03,
G-OBEY-PROBE-OFFLINE, enginec9d3834). New run P (plain delivery, capture rail on: 26/61 obey — balanced labels vs A's 54/61). Pre-registered pin (pooled AUC ≥0.80 AND cross-mode ≥0.70) FAILS on both frames: within-mode ≈ chance (frame-0: A 0.468, P-logistic 0.491; frame-1 P-only 0.205 = n=61×3840d overfit); the pooled 0.79/0.76 is the pre-registered mode confound (separates prompt condition, not outcome); cross-mode 0.302/0.563. Completes the surface's conviction: the final-norm state neither causes (G-FM-STEER-OBEY) nor linearly predicts (this gate) obedience — the decision lives upstream in attention over the fact text. Deployed detectors remain B3-JUDGE lexical grounding + the attr-gate. Hint banked, not claimed: P frame-1 v̄-proj means 85.5 (obey) vs 28.6 (not), variance-swamped — nonlinear probe or ~10× data to revisit. Receipttests/fixtures/chat_fullstack/G-OBEY-PROBE-OFFLINE.log; data_faithful_corpus/f3/P(tracked). -
THE SYSTEM SEALS: live memory grows from EMPTY, is recalled by paraphrase, and survives restart (2026-07-03,
G-B4-GROW-RECALL-L5GREEN, engine934f853). Two one-line-class holes had kept B4 growth invisible to the deployed selector: (1) live captures leftl5keyempty (the "follow-up" comment at routes.rs:735) — fixed bymint_live_ep_l5(standalone position-0 scratch forward →read_global_q→l5_query_embed;ep.l5sidecar written for restart; both the inline B4 and LAYER-3 merge writers); (2) cold-start bug: an EMPTY registry file loaded asNone⇒auto_recalldisabled for the entire serve ⇒ a production memory could never bootstrap — diagnosed live (run-1 grew 5 episodes, recalled 0/5, zeroRECALL-L5log lines), fixed (empty+env-set arms the chain; L5 scans registry ∪ nightshift). Gate on wholly novel facts (Biscuit / 4471 / Hobart / Volvo / lapsang): R 4/5 · F clean · persist PASS. New canonical daily driver:run_console_system.bat(production registry_memory_live);run_console_faithful.batdemoted to GATE-ONLY (RUNBOOK §6). Receipttests/fixtures/chat_fullstack/G-B4-GROW-RECALL-L5.log. -
SPECTEST (draft→test→veto→clean-execute) is LIVE and PARTIAL — and it convicts lexical draft-testing at the value-substitution boundary (2026-07-03,
G-SPECTEST-V1, engine5a2cf1c).SP_SPECTEST=1(default-off null floor): on delivery turns the full draft decodes with the stream HELD; grounding-tested before any byte reaches the client; PASS releases verbatim, VETO executes the clean symbolic answer from the record (leak-impossible on vetoed turns). Measured on plain delivery full-61: 26/61 → 41/61 obey (+24 pts), veto 23/pass 38 — but the 14 PASS-leaks are invariant across v1 (any-salient) and v1.1 (fact-minus-query): subject-grounded value-substitution drafts ("The capital of Canada is Ottawa.") speak in record tokens while swapping the value, and no lexical rule can identify "the value". Convergence law: the v2 semantic test-head IS the nonlinear obey/leak probe (the two operator ideas are one build); its training data = the 6-mode × 61 capture (_f3_modes_capture.bat→f3/M_*, mode-held-out design kills the prompt-condition confound). Telepathy low-quant external tester PARKED (redundancy law: a weaker external judge must beat deployed guards on signal — the test the 12B judge already failed). Receipttests/fixtures/chat_fullstack/G-SPECTEST-V1.log. -
THE TEST-HEAD IS REAL: obey/leak is READABLE from the final-norm DECIDE state at 0.89–0.90 held-out-mode AUC (2026-07-03,
G-TESTHEAD-OFFLINE, engine285a949) — and the signal is LINEAR. 6-mode × 61 capture (366 states, all VERIFY PASS,f3/M_*tracked); pins pre-registered mid-capture (0.75 mean / 0.65 min): frame-1 0.901/0.791, frame-0 (PRE-first-token) 0.887/0.717 — both REAL. obey-vs-LEAK 0.78–0.98: the value-substitution class that convicted lexical testing (G-SPECTEST-V1) is latent-separable. Mechanism: linear == MLP (0.890 vs 0.887) ⇒ theG-OBEY-PROBE-OFFLINEchance result is SUPERSEDED-IN-SCOPE: it was a power+confound failure (n=61, 1–2 modes), not absence — the banked "10× data" revisit condition fired exactly. The final-norm story is now symmetric and complete: the surface EXPRESSES the upcoming decision (readable, 0.90) but cannot be PUSHED (G-FM-STEER-OBEYconviction stands). NEXT: retrain-on-366 export →SP_SPECTEST_HEADat the held-stream seam →G-SPECTEST-V2(close the 14-leak class; systemecho-class faithfulness without the scaffold). Receipttests/fixtures/chat_fullstack/G-TESTHEAD-OFFLINE.log. -
THE HALLUCINATION VETO IS LIVE:
G-SPECTEST-V2GREEN — plain delivery + the linear test-head = 52/61 obey / 2 leak, systemecho-class faithfulness WITHOUT the scaffold (2026-07-03, engine9adda35). The ladder, all same-day full-61: plain 26/61·14+leaks → +lexical veto 41/61·14 → +linear head 52/61·2 → systemecho 54/61·0 (reference). Head = theG-TESTHEAD-OFFLINEpipeline collapsed to one dot product (SPH1 blob 15KB,score = x·v + con the frame-1 state via the one-shotcapture_feattap); veto = lexical ∪ head(τ=0.6271, 95% train-leak-catch @ 3.2% false-veto); head-vetoes 30 / lexical 8 / passes 23. Rigor: deployed head trained on 5 modes EXCLUDING plain ⇒ the live run is mode-held-out. Limits on record: 2 residual head-false-passes; ITEM-level generalization unproven (61 facts appear in training under other modes) — V3 = fresh-fact corpus gate before any canonical promotion; selection ceiling untouched. FlagsSP_SPECTEST+SP_SPECTEST_HEAD, default-off null floor. Receipttests/fixtures/chat_fullstack/G-SPECTEST-V2.log. -
The veto's safety property GENERALIZES; promotion correctly blocked by its own pin (2026-07-03,
G-SPECTEST-V3PARTIAL, enginece094c1). The hardest available configuration — 30 novel varied-template counterfacts (pre-registered + committed before any serve), grown as LIVE episodes via the B4 seal, gated plain±head, item-held-out × grown-registry × mode-held-out at once: baseline 7/30 obey · 20 leaks (strongest-prior corpus yet) → head run 10/30 · 1 leak. Pin 1 (LEAK ≤ 2): PASS — 20→1 on wholly unseen items. Pin 2 (OBEY ≥ 15): FAIL, cause isolated OUTSIDE the head — every miss is "From the record: 〈WRONG fact〉": SELECTION cross-picks on assertion-minted L5 keys (statement-space; cos 0.74–0.81 background) across same-register encyclopedic facts, where the curated corpus mints keys from exact-question global-Q. Failure mode shifted confidently-wrong → faithful-to-the-wrong-page (auditable). Named lever: question-spaceep.l5minting at B4 capture (question-form of the assertion / multi-key episodes), then re-run this gate unchanged. Receipttests/fixtures/chat_fullstack/G-SPECTEST-V3.log.
- Law + architecture: PPT-LAT-ADR-002-DECIDE-EXECUTE-SPINE.md
- What is built + open: VERIFIED-SCOREBOARD.md
- Proven ledger (full detail): PPT-LAT-STATE.md
- Swarm/DHT primary axis: PPT-LAT-DESIGN-SWARM-MEMORY-MESH.md
- MEM-OKF banked findings (hashed, gisted):
memory-okf/(python tools/okf_mem.py lookup --root memory-okf "<keyword>")
Banking discipline: durable findings here are also banked to MEM-OKF (content-addressed, SHA-hashed, gist+full tiers) so they are anti-rebuild-searchable. This doc is the human-readable flat index; MEM-OKF is the hashed store.