| type | project-state | ||||
|---|---|---|---|---|---|
| title | Verified Scoreboard — what is built, what is open (receipts-checked) | ||||
| description | The proven-capability scoreboard, re-verified 2026-07-01 against actual commits + gate fixtures by the Phase 0 ground-truth fleet (not transcribed from prior docs); NIGHTSHIFT graduated PARTIAL→VERIFIED 2026-07-02 (criterion-5 provenance fix closed). 10 VERIFIED, 0 PARTIAL, 0 false-greens. Plus the open frontier, the honest negatives kept attached, and the contradictions reconciled. | ||||
| tags |
|
||||
| timestamp | 2026-07-01 00:00:00 UTC | ||||
| resource | shannon-prime-lattice/papers/VERIFIED-SCOREBOARD.md | ||||
| sp_status | GREEN | ||||
| sp_gate | each row names its gate | ||||
| sp_commit | see rows | ||||
| sp_repro | Phase 0 fleet audit 2026-07-01: every commit confirmed an ancestor of engine HEAD; every gate fixture read, not just stat'd |
Method: a read-only fleet checked each claim against (a) a commit that resolves and is an ancestor of engine HEAD, and (b) a gate fixture/receipt file that exists and contains a GREEN verdict. A claim with no commit+gate was to be downgraded. Result at the 2026-07-01 audit: 9 VERIFIED · 1 PARTIAL · 0 UNVERIFIED · 0 false-greens; current (2026-07-04, LIVING-MEMORY + SELF-IMPROVEMENT LOOP + PERSONALITY FRAMEWORK graduated): 13 VERIFIED · 0 PARTIAL · 0 false-greens (LM bricks composed + faithfulness-neutral; prior re-spot-check 2026-07-02, AUDIT-2026-07-02.md).
| Capability | Status | Commit | Gate / receipt | Verdict |
|---|---|---|---|---|
Byte-exact 12B forward (SP_BYTEEXACT) |
gated-GREEN, default-off | 69c0588 |
tests/fixtures/xbar_r3/G-BYTEEXACT-FORWARD-12B.log |
VERIFIED — off 4.6665 null-floor / on 4.6569 / run-to-run bit-identical |
O(1) persistent conversation KV (SP_PERSIST_KV) |
PROVEN, default-ON | d211fd2 |
tests/perf/G-PERSIST-KV-PARITY.log |
VERIFIED — 6-turn SHA-256 ON==OFF byte-identical |
| SWA ring (sliding-window KV shrink) | GREEN | in 0019b86→d2d7ceb |
tests/fixtures/xbar_p3_replay/ |
VERIFIED |
| Global-eviction slab + learned-LSH + NIAH | CLOSED GREEN | 222463a,33ac632,8e35877,3218d73 |
tests/fixtures/lsh/lsh_R_r32_raw.bin+lsh_M_r32.bin; NIAH/ladder logs |
VERIFIED — needle HIT at depth 10/50/90%; FROZEN±1 = MISS (neg control) |
| B3-WC autonomous episodic recall | CLOSED GREEN-LIVE | edc8079 |
tests/fixtures/chat_fullstack/G-CHAT-B3-WC-{DEPLOY,DIV2}.log |
VERIFIED — live on the 12B; foreign-reject clean |
| Diffusion Judge OOD | PROVEN | lattice dda7ffa |
tests/fixtures/chat_fullstack/G-DIFFJUDGE-OOD-H2H.log |
VERIFIED — recall 94.4% / reject 98.0% |
| CHAT-FULLSTACK (served 12B chat) | GREEN-LIVE | chat_fullstack arc | CONTRACT-CHAT-FULLSTACK.md + G-CHAT-*.log |
VERIFIED — coherent + byte-exact + O(1) + single-entry |
| Memory agency: forget / decide / merge | GREEN | chat arc | G-FORGET.log,G-DECIDE.log,G-MERGE.log |
VERIFIED — all three GREEN, null-floor when off |
| Telepathy TELE-1..15 (bridge + route + two-stage + native delegate + LIVE delegate) | PROVEN (1-3) / WIRED (4) / native CPU delegate / LIVE two-stage delegate | 2f57520 (TELE chain) + c3c4b22,ef6c282 (TELE-15) |
tests/fixtures/telepathy/G-TELEPATHY-LIVE.log + receipts in tools/telepathy/* |
VERIFIED — TELE-15 G-TELEPATHY-LIVE GREEN: decide_route(latent) → delegate_execute runs the qwen coder on CLEAN TEXT (CPU L1, ~0.8 tok/s), coherent answer; never-fuse honored. Honest scope = gist/intent routing + clean-text execution, NOT latent verbatim |
| MEM-OKF anti-rebuild store | ACTIVE | tooling | tools/okf_mem.py + memory-okf/ (verify GREEN, 83 objects) |
VERIFIED |
L5-direct recall (SP_RECALL_L5 τ=0.30 + systemecho delivery) |
RECOVERED GREEN-LIVE 2026-07-02 (regressed then re-won same day) | d9099cd → downgrade ad27b1b → recovery 8ae343b |
G-L5-OBEY-REPRO + G-DELIVERY-SWEEP + G-ONECONFIG-LIVE |
The 07-01 86.89% receipt proved NOT reproducible (plain recite 40.98% on the identical stack — bounded-unidentified env divergence, G-CUBLAS-PIN-CANARY); the systemecho delivery (fact as SYSTEM authority + echo priming) re-won it at 88.52% OBEY / 0 LEAK — 54/61 == selector top-1, i.e. every correctly-selected episode obeyed; delivery solved, ceiling = selector. Obey receipts now carry build SHA + canary + full env dump (law) |
ONE-CONFIG composition (run_console_faithful.bat: Tier0 + L5+attr-gate+QONLY+systemecho, merged 81-ep registry) |
RE-GATED GREEN-LIVE 2026-07-03 @ PMAX=4096 (first 8ae343b @20000) |
engine d9ee34b |
G-ONECONFIG-LIVE (v2, re-gate) |
VERIFIED — RE-RUN on the SPEED-FIXED stack (PMAX 20000→4096 per G-12B-SERVE-ROOTCAUSE + dead-scan skip + bx chunked-fold): P 54/61 IDENTICAL to the 20000 gate · SNE 3/3 zero-inference declines, 0 spurious · hard-foreign 2/2 clean · multi-turn coherent · QONLY 2/2 · byte-identical rerun. The whole 12B speed campaign is faithfulness-neutral; numbers re-earned on the config the operator runs (no silent gate revision). NOTE: run_console_faithful.bat = the GATE launcher (loads the 61-fact TEST corpus); chat daily-driver = run_console_chat.bat (no test registry, recall off). Spec+history RUNBOOK-ONE-CONFIG.md |
SP-SWARM L4 gossip discovery (SIM + discover_similar) |
GREEN (2-node localhost) | engine (this session) | tests/fixtures/swarm/G-SWARM-GOSSIP-DISCOVERY.log |
VERIFIED — C2 shortlist gossip over QUIC → exact-fetch verify → converge; A discovers a B-only object, decoys skipped, off-roster rejected; k≥5 shortlist regime |
SP-SWARM L4 C2 similarity overlay (sp_swarm::similarity) |
mechanics GREEN; semantic HINT-only | engine (this session) | G-SWARM-C2-INDEX (mechanics) + G-SWARM-C2-SEMANTIC (recall) |
VERIFIED — index top-k Hamming correct; C2-256 = weak top-1 (0.607 vs cosine 0.885) but strong shortlist (recall@5 0.885); expose as top-k hint→exact-fetch (its designed role), NOT top-1; bit-count is the lever (2048≈cosine) |
SP-SWARM daemon integration (swarm feature + run_node) |
GREEN (default-off) | engine (this session) | G-SWARM-NODE.log + G-SWARM-DAEMON-WIRE.log |
VERIFIED — run_node autonomous sync converges (5/5 both ways); sp-daemon builds+links with --features swarm; SP_SWARM=1 spawns the mesh on its own port w/ persistent identity, unset=no-op |
| SP-SWARM L0 network transport (QUIC + Ed25519 roster) | GREEN (2-node localhost) | engine (this session) | tests/fixtures/swarm/G-SWARM-TRANSPORT-QUIC.log |
VERIFIED — reuses engine QUIC (quinn/rustls); Ed25519 mutual roster auth; A↔B bidirectional convergence, tampered rejected on arrival, off-roster peer dropped; behind transport feature (not rust-libp2p — anti-rebuild) |
SP-SWARM Rust crate tools/sp_swarm (L1/L2/L3 port) |
GREEN (local, cross-lang parity) | engine (this session) | tests/fixtures/swarm/G-SWARM-RUST-PARITY.log |
VERIFIED — 6/6: Rust sha2 reproduces Python addresses (89 content), ed25519-dalek verifies pynacl sig, tamper/roster reject; CRLF cross-platform interop bug found+fixed; L0 libp2p transport behind transport feature (next) |
SP-SWARM L3 Ed25519 provenance (swarm_provenance.py) |
GREEN (local, libsodium/PyNaCl) | lattice (this session) | tests/fixtures/swarm/G-SWARM-PROVENANCE-ED25519.log |
VERIFIED — sign-on-write + verify-on-pull against invite-only roster; tampered-episode/stripped/forged/unrostered/tampered-content ALL rejected pre-commit; C2 episodes tamper-evident cross-node; L0 libp2p transport still DESIGN |
SP-SWARM L1+L2 replication core (swarm_sync.py) |
GREEN (local, transport-agnostic) | lattice (this session) | tests/fixtures/swarm/G-SWARM-REPLICATE-CONVERGE.log |
VERIFIED — content-address round-trip + have/want convergence (113/113 byte-identical) + verify-on-arrival + idempotence + tamper-reject over the real MEM-OKF store; L0 libp2p transport still DESIGN |
Attribute-grounding gate (SP_RECALL_ATTR_GATE + query-token guard, zero-inference decline) |
GREEN, default-off | fc2e846 |
G-SNE-ATTRGATE-ZEROINF / G-ATTRGATE-GUARD-PARA |
VERIFIED — SNE confab 80→0, leak 5→0, recall 100%; paraphrase 6/12 baseline, 0 over-decline; decline runs NO gemma4 forward |
| NIGHTSHIFT offline curator | GREEN (default-off) | 6107f3e,9ad7ede,9ee4668,3ccba61 |
G-NIGHTSHIFT-CURATOR.log + G-CHAT-B4-NIGHTSHIFT-provenance.log |
VERIFIED — synthetic gate GREEN (criteria 1-4); criterion 5 (B4 distributional/provenance fix) CLOSED live: live episode == curated (9.858 byte-perfect; the 0.084 collapse gone), novel fact in-band (6.295) + clean foreign-reject (−15.498). Residual (pre-scoped, not a miss): novel-instance out-ranking under the closed-set W_c head → superseded on the hot path by the shipped L5-cosine selector; the one follow-on = gate L5 recall on a nightshift-captured episode |
Decide→Execute SPINE (ADR-002 §8.2 realized, spine.rs) |
GREEN-LIVE 2026-07-03 | engine 579552d |
G-SPINE-ONECONFIG (tests/fixtures/chat_fullstack/G-SPINE-ONECONFIG.log) |
VERIFIED — the unified latent-decision framework: LatentView (immutable) → priority-folded Deciders [L5Recall, AttrGate] → discrete LatentDecision → Executor, compiler-enforced boundary (deciders can't touch the stream, executor can't read a tensor). SP_SPINE=1 reproduces the one-config stack byte-for-byte: P 54/61 · S 3/3 zero-inference · F 2/2 · C · Q 2/2 · X identical to inline (behavior-preserving; the ~1500-line env-branch ladder collapsed to a fold). Default-off = inline untouched. Heads/telepathy/judge scaffolded as extension-point Deciders. Design: PPT-LAT-SPINE-FRAMEWORK.md |
SPEED_NORTHSTAR: qwen36 35B-A3B SERVED chat (run_console_qwen36.bat, port 3001) |
GREEN-LIVE | engine c12d1ea+c0ec86b; submodule 5d1fdaa |
chat_fullstack/G-QWEN36-SERVE.log (+G-MOE-* ladder, 0.018→6.073 tok/s = 337×) |
VERIFIED — /v1/chat coherent + multi-turn + greedy-deterministic at 5.33–5.55 tok/s on the 2060 (sessionless daemon on arch_id 8; GPU boot sp_q36gpu_boot: dense 852.5MB + 25/40 expert layers resident + pinned streaming rump). KEY BOUNDARY BANKED: the standard build-cpu math-core libs compile WITHOUT /openmp /arch:AVX2 → any exe linking them loses the brick-4/5/6 ladder (3× serve hole, CPU A/B hit the pre-OMP 0.17 rung); the qwen36 launcher uses the target-wirecuda-perf exe (build-cpu-perf libs + LLVM libomp); the 12B one-config exe untouched |
| LIVING MEMORY stack (ADR-005: hot-reload · decision-telemetry · turn-telemetry · idle refine) | GREEN, default-off | engine de8506e→6431e55→e272f4a→70e4992 |
chat_fullstack/G-LM-{RECONCILE,TELEMETRY,REFINE,TURNTELEM,COMPOSE}.log |
VERIFIED — 4 bricks, each isolated-gated, then COMPOSED (G-LM-COMPOSE): all flags on at once serve 3 policies (counterfact/private-secret/persona) while the idle thread BOTH reconciles a concept written mid-run (+1 hot-loaded → served "Marlowe City", no restart) AND model-refines a mis-classed secret (counterfact→private-secret on idle), every turn logs a decision + turn record (4+4), private-secret redacted (secret 0 hits). B1 SP_MEM_RECONCILE idle reconciler; B2 SP_TELEMETRY class-redacted decision log; B2b turn-outcome (output+tok_s+obeyed, secret output redacted = finetune data); B3 SP_MEM_CLASSIFY_REFINE idle model-refine (safety-monotone). Faithfulness-neutral: none of these flags are set by run_console_faithful.bat, and none touches L5 selection scoring → the 54/61 one-config stack is byte-identical. Design: PPT-LAT-ADR-005-LIVING-MEMORY.md + CONTRACT-LIVING-MEMORY.md |
| SELF-IMPROVEMENT LOOP (data-gen + finetune: telemetry → train → promote → deploy, autonomous) | GREEN, default-off | harness b3fb755→bfaa4eb; engine 6d418a4 |
harness tests/G-DF-{CONVERT,SEED,TRAIN-CLOUD,EVAL,DEPLOY,PARITY,AUTOTRAIN}.log + engine chat_fullstack/G-DF-LIVE.log |
VERIFIED — the LM-B2 telemetry flywheel closed into a live self-improvement loop (CONTRACT-DATAGEN-FINETUNE, every brick lifts a named CosySim module, anti-rebuild). DF-B1 telemetry→Alpaca JSONL (privacy choke: redacted skipped); DF-B2 synthetic mem_class seed (6 balanced, incl synthetic private-secret); DF-B3 QLoRA on Colab T4 (Qwen2.5-0.5B, HF-mediated, adapter→KnackAU/sp-mem-class-adapter); DF-B4 distinct-held-out A/B (20%→83.3% vs base) + registry gate_and_promote (MUST_IMPROVE); DF-B6 harness curator deploys the 0.5B on CPU (safety-monotone) replacing the engine's 12B model_classify; DF-B5 auto-train trigger (accrued telemetry→fire, idempotent, DRY unless SP_AUTOTRAIN_LIVE). G-DF-LIVE: curator corrects the served store live → engine reconcile-on-edit serves the correction, no restart. G-DF-PARITY: the deployed 0.5B BEATS the 12B it replaces (0.83 vs 0.33 vs ground truth — the 12B model_classify was the weak classifier). Whole loop default-off; the 12B refine stays as fallback. |
| PERSONALITY FRAMEWORK (self-modifiable + system-curatable persona/self-model) | GREEN, default-off (SP_PERSONALITY) |
harness a1c59ea→e35cfdf |
harness tests/G-PF-{OWNERSHIP,PERSONA,TAGS,DECORATORS,CURATE}.log |
VERIFIED — CONTRACT-PERSONALITY PF-B1..B5, every brick EXTENDS a named existing seam (anti-rebuild). PF-B1 mem_owner axis (self|user) orthogonal to mem_class, owner-tagged OKF concepts (memory-okf-self/), no classifier; PF-B2 structured persona.md ## Personality state block folded into load_agent_system + self-model injection; PF-B3 PersonalityStateInterceptor persists the [MOOD]/[VOICE]/[TRAIT±] tags the model already emits; PF-B4 @personality decorators (=@skill pack) — the model CALLS adjust_mood/set_voice/set_trait/remember_self via run_with_tools to durably self-modify; PF-B5 consolidate_personality (mirrors consolidate_conversation) extracts transcript shifts + prunes stale traits + snapshots to content-addressed memory-okf-personality/ tier, wired into agency.py gated SP_PERSONALITY. Personality self-modifiable (model) AND system-curatable (NIGHTSHIFT). PF-B6 (engine-native personality head) DEFERRED. ADR-002: personality = DECISION → clean tag/decorator/label → EXECUTE. |
PRODUCT KEYSTONE-2 T1: prefill "stall" fix (gemma4_kv_prefill + daemon.rs CRT bridge) |
GREEN-LIVE 2026-07-07 | engine (this session) | tests/perf/G-PK2-PREFILL.log |
VERIFIED — the >~1000-tok "stall" was the SWA ring never arming in served daemons (Rust set_var invisible to the CUDA C getenv on Windows) → ring-off full-cache @PMAX=20000 → VRAM oversubscription/WDDM thrash. Fix: _putenv_s CRT bridge + chunked-sync/fail-fast/telemetry. SP_WORDS=1300 → got_DONE=True 98.2s (was forever); n=299 regression 12.0s; ring-ON chunk times flatten past dpos=1024 (O(1)-context). LESSON: Rust→C env needs _putenv_s |
| PRODUCT KEYSTONE-2 T2/T3/T4: harness product hardening (coding tools · task loop · MEM-OKF v2 · UI surfaces) | GREEN offline; T2 live = honest capability boundary | harness (this session) | harness tests/G-PK2-{TOOLROBUST,MEMOKF-V2,UI-ENDPOINTS}.log + G-PK2-TRANCHE-SUMMARY.md |
VERIFIED (offline) — T2 G-PK2-TOOLROBUST 10/10: edit_file/run_tests/git tools, resumable run_task + work queue, malformed-recovery + no-progress break + verify-before-accept (closes the 12B "DONE"-confabulation gap found live); T3 G-PK2-MEMOKF-V2 6/6: provenance lane (remember(source)+provenance()), near-dup guard, registry verify/compact; T4 G-PK2-UI-ENDPOINTS 5/5: /v1/memory·/v1/tasks·/v1/persona gateway surfaces + operator.html panel + self-knowledge refresh. Live task-loop: the 12B autonomously editing code to green is at/beyond this model's reliable capability (honest negative); the harness now refuses to confabulate a pass. Contract: CONTRACT-PRODUCT-KEYSTONE-2.md |
| ADR-006 Verified Agency (verify-before-accept law · SSE v2 typed events · agency-tick hygiene) | GREEN offline, default-off | harness (wave 2) | harness tests/G-PK2-{SSE-V2,UI-ENDPOINTS}.log |
VERIFIED (offline) — PPT-LAT-ADR-006-VERIFIED-AGENCY.md. SSE v2 G-PK2-SSE-V2 5/5: gateway stream carries typed {persona}/{tool}/{delta} events, typed_events:false = pure-delta backward-compat. Hygiene on KAIROS tick + /v1/persona/state chip: G-PK2-UI-ENDPOINTS 7/7. Verify-before-accept named as the law unifying task-loop verify + NIGHTSHIFT admission + DECIDE |
PK2 batched dp4a prefill GEMM (SP_KV_PREFILL_DP4A) |
HONEST NEGATIVE — arithmetic correct, not a speed win; default-off | engine (wave 2) | tests/perf/G-PK2-PREFILL-DP4A.log |
VERIFIED-negative — packed OK_Q4B int4 dp4a batched GEMM (no f32 materialization): coherent "Paris" through the full 48-layer forward (arithmetic CORRECT) but naive + register-tiled both memory/occupancy-bound on the 2060 → NOT faster than cublas-lift. Ships default-off (unset = gemm_w_lift byte-identical null floor). Real levers = cublasGemmEx int8 / SMEM tiling / batched-under-ring (follow-on); kernels committed dormant for a future A/B |
| ADR-007 Harness Spine (decide→execute→VERIFY fold; persona/hygiene/recall deciders; SSE persona-change + heartbeat) | REALIZED, offline-GREEN | harness (wave 3) | harness tests/G-PK2-SPINE.log (9/9) + G-PK2-SSE-V2.log (7/7) |
VERIFIED (offline) — PPT-LAT-ADR-007-HARNESS-SPINE.md. control/spine.py: TurnView→Deciders→Executors→Verifiers→SpineReceipts (~150 lines). persona_tags decider reuses PF-B3 apply_personality_tags, verified by re-reading persona.md; tick hygiene verified-clean; recall decider (ranked search_memories_ranked) abstains on foreign, wiring opt-in; VERIFY_FAIL honesty case: a lying executor is flagged, never trusted. Gateway post-turn + agency tick route through the spine; verified mid-conversation persona shift → {"persona",changed:true} SSE + persona.md persistence; {"hb"} keep-alive; console renders tool cards + live persona chip. Memory: search_memories/memory_stats (hot set stays ≤6, extras on the OKFS tier) |
ADR-008 Adaptive Turn (toolset decider · recall wired · receipt ring /v1/spine) |
REALIZED — offline 12/12 + recall LIVE-GREEN 4/4 | harness (wave 4) | harness tests/G-PK2-SPINE-2.log (12/12) + G-PK2-RECALL-LIVE.log (4/4, live 12B via gateway) |
VERIFIED — PPT-LAT-ADR-008-ADAPTIVE-TURN.md. Toolset decider picks the RIGHT ≤6 tools per turn (coding/memory/core, deterministic, null-floor on chat; SP_SPINE_TOOLSET). Recall decider WIRED (SP_SPINE_RECALL): matched facts → system note + {"recall"} SSE; live faithfulness on the real 12B: matched → "Your favorite color is teal." (80s), foreign → abstains + clean "Paris" (78s, no hijack). Receipt ring: every decide→execute→verify verdict at /v1/spine + operator pane (VERIFY_FAIL renders red). One recall authority at a time — see the wave-5 composition row |
Flywheel persistence of spine receipts (telemetry-okf kind:"spine") |
GREEN offline, default-on (cheap+safe) | harness (wave 5) | harness tests/G-PK2-FLYWHEEL.log (6/6) |
VERIFIED — spine receipts flush into the durable telemetry-okf tier via the EXISTING TelemetrySink (content-addressed, idempotent, watermarked incremental; anti-rebuild). Flushed each agency tick + gateway turn; the ADR-005 flywheel corpus now includes what the harness DECIDED + whether verify held. Private-secret never routes through spine payloads by construction |
| recall∘L5 composition (the two recall authorities, both armed) | FREE composition REFUTED → one-authority guard ENFORCED, LIVE-GREEN 6/6 | harness (wave 5) | G-PK2-RECALL-L5-COMPOSE-FREE.log (the honest negative, 4/6) + G-PK2-RECALL-L5-COMPOSE.log (enforced, 6/6 live) |
VERIFIED — Phase A (both armed): L5's systemecho CROSS-PICKED color-adjacent queries from the counterfact corpus and overrode the harness note ("favorite color?" → "Human blood is green"; "sky?" → "Green") — L5 selection cross-picks are a known daemon residual; composition surfaces them user-visibly. Phase B: the gateway auto-disarms spine recall whenever the request arms L5 (auto_recall=true), with an {"authority":"L5"} receipt event; guard held, L5 authority intact ("Lyon"), harness authority faithful when L5 off ("Teal."). Plus auto_recall body passthrough (default false). The one-authority rule is now STRUCTURAL, not operator discipline |
ADR-009 batched prefill under the SWA ring (SP_KV_PREFILL_BATCH, the ring-off precondition lifted) |
LIVE-GREEN 3/3 — bounded ~7× win + graceful fallback; default-off | engine (wave 6) | tests/perf/G-PK2-BATCHRING.log (3/3) |
VERIFIED — PPT-LAT-ADR-009-PREFILL-SPEED.md. The batched cold prefill was BLOCKED under the ring (so the ring config never used it); wave 6 lifts it (ring-layout k_ring_sink for SWA owners + retained contiguous shared-owner K/V for sharers). Live @ PMAX=4096: n≈837 → 6.7s coherent (~7× vs ~39s per-token); n≈1765 → VRAM guard DECLINES → graceful per-token fallback 99s coherent (null floor, live-proven). Two guards: VRAM (SP_KV_BATCH_VRAM_MARGIN_MB) + persist (declines under SP_PERSIST_KV, since a batched-ring cache isn't a valid persist-continuation base). Verdict across 3 waves: batching is a bounded cold-prefill win, not the general lever — the f32 activation scratch caps it; cublasGemmEx-int8 would hit the same ceiling (not pursued). Default-off; production per-token path unaffected |
| ADR-010 whole-machine balance (measured VRAM breakdown + KV auto-fit; the CPU-offload verdict) | ANALYSIS (measured) + auto-fit realized default-off | engine (wave 7) | tests/perf/G-PK2-AUTOFIT.log |
VERIFIED (measurement) — PPT-LAT-ADR-010-WHOLE-MACHINE-BALANCE.md. Measured live: model ~10.9 GB resident FILLS the 12 GB card; KV cache only ~0.78 GB (shared-KV — NOT the 2.5 GB I first assumed); ~0.57 GB free. ★HONEST CORRECTION: the original ">1000-tok stall" root was co-resident LM Studio (~6.4 GB oversubscription), and the ADR-009 batch thrash was the transient scratch eating the thin headroom — the KV was never the movable mass, THE MODEL IS. Every weight-offload refuted by PCIe x8 (6.2 GB/s forbids per-token weight streaming); the ONE lever that frees real VRAM = CPU-resident layer offload (weights in 40 GB/s DRAM, compute on AVX2, exchange only activations; CRT keeps CPU+GPU legs bit-identical) — the ADR-011 candidate. SP_G4_KV_AUTOFIT (default-off) clamps Pmax to free VRAM for the co-tenant case; dedicated card serves coherent ("Paris" 12.4s), null floor intact |
| ADR-011 CPU-resident layer offload (FFN-tail hybrid — REALIZED + perf-gated) | STAGE-2 REALIZED + PERF-GATED — coherent, frees VRAM, usable at ~6 tok/s (K=4); steep sync-bound slope | math-core+engine (waves 8–9) | tests/perf/G-ADR11-{MEASURE,HYBRID,HYBRID-PERF}.log |
VERIFIED — PPT-LAT-ADR-011-CPU-LAYER-OFFLOAD.md. gemma4_ffn_block_cpu (math-core, sp_matmul on OK_Q4B) + g4_kv_step routes last SP_G4_CPU_TAIL FFNs to CPU + build_weights skips those uploads. ★DESIGN PIVOT: shared-KV → offload the FFN only (~90% of layer weight, stateless). PERF re-gate on the AVX2/OpenMP daemon (build-cpu-perf+libomp, target-wirecuda-perf), PMAX=4096: K=0 23.84 tok/s → K=4 5.99 tok/s (freed ~0.46 GB) → K=8 3.25 tok/s (freed ~0.96 GB), coherent all K. Perf libs made it usable (K=8 0.5→3.25 tok/s ~6.5×). HONEST slope: ~33 ms/offloaded-FFN/token = the per-FFN GPU↔CPU round-trip (D2H+2×sync+H2D, ~2K syncs/token) — sync-bound, not compute-bound, far steeper than the ~8.4× memory-bound ideal. VRAM-for-latency lever; usable where VRAM is the hard constraint. Default-off = 23.84 tok/s null floor. Flatten-the-curve (next) = cross-token decode pipeline |
Grouped by the 4 roadmap axes:
Axis 1 — Sovereign Telepathy. v1 DONE (TELE-15 G-TELEPATHY-LIVE + TELE-16 G-TELEPATHY-CHAT-LIVE, both GREEN): the cemented two-stage delegate runs live AND is wired into the served /v1/chat SSE path (SP_TELEPATHY_CHAT=1, default-off null floor) — a routed turn streams the qwen coder's clean-text answer into the live {delta} stream with ⟦delegate⟧ markers, never fusing (TELE-12). Engine c3c4b22/50200eb; tag telepathy-v1. Remaining: (a) autonomous feat-route on the served path — currently route is SP_ROUTE_FORCE-driven; the TELE-7 head needs a NON-COMMITTING capture_feat dispatch verb (the existing gemma4_kv_capture_feat is async-armed and would commit the cache) + d_model plumbing; (b) GPU-speed transmit, blocked by the arch gap (CUDA qwen3_decode_cuda is SP_ARCH_QWEN3-only; qwen3_generate_kv segfaults on the CUDA session) → coder runs CPU-only (~0.8 tok/s) by design. NOTE: a latent-prefix inputs_embeds seam is deliberately NOT built — fusing latent+text is the TELE-12 0.000 negative. Licensing/attestation = SPEC (fail-closed).
Axis 2 — CRT residue split / multi-device. Garner 2-prime constants PROVEN; residue-exchange multi-device is [DESIGN]; QUIC residue transport proven on loopback only. The 2-physical-GPU byte-exact bit-identical check is the one remaining external byte-exact item.
**Axis 3 — Absolute faithfulness. SOLVED (2026-07-01, 184994b); reopened-in-part 2026-07-02 (obedience receipt unreproducible) and RE-CLOSED the same day via the systemecho delivery: 88.52% obey / 0 leak on the full 61, delivery-obedience effectively 100% on correct selections, whole-stack composition gated GREEN (G-ONECONFIG-LIVE). Residual obey ceiling = the selector (54/61 top-1). Recall-path fact obedience = 100% (G-FAITHFUL-RECALL-JACCARD 15/15) on a strong-prior conflict set. The chain: in-context obedience was already 100% (F1); pure-KV replay recall = 0% (F1b.1 — W_c mis-selects natural facts + attenuated K/V can't override a prior); the fix (F2b, SP_RECALL_JACCARD, default-off) = select the episode by token-overlap/Jaccard (recall::token_overlap, not the geometric W_c) + deliver text-in-context with the faithfulness system prompt. Rule: natural facts → Jaccard+text; novel high-entropy needles → keep W_c+replay. Receipts tests/fixtures/faithful/. Extended 2026-07-01 (this session): the recall SELECTOR graduated from Jaccard to L5-cosine (the fact signal is layer-localized in global layer 5; G-L5-RECALL-LIVE = 86.89% paraphrase, default-off); the heavy generative judge was PARKED (hard-foreign kill-test: 0 benefit over L5-direct+τ, and it PASSed 15/18 — G-HARDFOREIGN-L5DIRECT/-JUDGE); and the one regime native robustness does NOT cover — zero-prior/private data (the SNE crucible: 80% confabulation + 5% secret-leak on novel entities) — is closed by a deterministic attribute-grounding gate (SP_RECALL_ATTR_GATE + query-token guard) with a ZERO-INFERENCE symbolic decline (G-SNE-ATTRGATE-ZEROINF: confab→0, leak→0, recall 100%, paraphrase untouched, the decline streams a fixed string with no gemma4 forward). Faithfulness is now solved on BOTH general-knowledge conflict AND private zero-prior data. Full lever/constant map: PPT-LAT-FINDINGS-LEDGER.md; law: PPT-LAT-ADR-002-DECIDE-EXECUTE-SPINE.md.
Axis 4 — Native consolidation. Port host-Python XBAR tooling (C2 signatures, Frobenius episode codec) to C/Rust; T4 Frobenius of model weights (validated lever, unbuilt); single-binary deployment. Wire the proven P3 eviction slab from the test-bin into the live daemon (PORT, not rebuild — lsh_R_r32_raw.bin is already trained).
Also open: WIRE-CPU integer-pipe speed (~23× behind llama.cpp on 0.6B CPU — the RFC north-star P1); the one NIGHTSHIFT follow-on = gate L5-cosine recall on a nightshift-captured live episode (criterion-5's provenance/distributional fix is CLOSED); P3.4 larger-N multi-chunk PPL hardening; P3 global-eviction live win-gate (needs GPU headroom past the 2060's ~20K resident ceiling).
- P2.b generation channel at k=2: latent generation refuted; recognition sub-usable (32-way top-1 0.462). Adopted fallback = two-stage retrieve-then-verify.
- Telepathy precise-symbolic channel: latent carries gist/intent only; fused latent+text = 0.000 (corrupts downstream). Architecture is strictly two-stage: decide via latent, execute via clean text. Never fuse.
- C2.4 32k/64× Optane finale: needle not retrieved at 64× selection — closed as a miss.
- KSTE as a recall router: falsified (histogram, permutation-invariant); router is the ±1 Rademacher projection.
- ETA.5b 34.2 tok/s: retired (quant artifact failed PPL gate); real number 26.1 tok/s @ PPL 5.12.
- Generative recall judge (
SP_B3_JUDGE): PARKED (2026-07-01) — hard-foreign kill-test = 0 benefit over L5-direct+τ (0/18 == 0/18) and it PASSed 15/18 (failed its own reject job). Code kept default-off as an honest negative. - STRICT closed-book prompt (
SP_RECALL_STRICT): over-declines even valid matches (0/4) — dead lever, replaced by the deterministic attribute-gate. - Shared-token attribute guard (query∩fact): broke SNE decline (2/6) on wrong-entity delivery → replaced by the query-token guard (gate on the QUERY carrying a private-entity token).
- T4 Frobenius π^k of the model WEIGHTS as a compression lever (
G-T4-WEIGHTS, 2026-07-01): REDUNDANT vs OK_Q4B — pre-registered kill-test on 3/3 real 12B tensors. OK_Q4B (per-32-block int4+f16, 4.5 eff bits/w) relL2 0.10–0.12; the Frobenius "free" per-tensor scale at 4 b/w is ~3.3–4× worse for 0.5 bits saved; matching fidelity needs the scale back at per-block (== OK_Q4B). The free-scale property that holds on Ring-2 episodes (G-R2-FROB) does NOT transfer to trained weights. Refutes T4-as-weight-compression, NOT the exact-cancellation property (needs Q8, buys auditability). Resolves the roadmap/CLAUDE.md tension → "convicted redundant." ScopePPT-LAT-T4-WEIGHTS-SCOPE.md; receipt enginetests/fixtures/t4_weights/G-T4-WEIGHTS.log.
These were drifting across dated docs; the authoritative resolution:
- Diffusion judge — the 26B cascade is retired; production recall gate is the deterministic Jaccard verifier @0.6. The native diffusion judge plateaus ~50% recall (won the OOD kill-test as a proxy, but is not the production path). The 06-21 docs claiming "N5b justified / native judge is production" are superseded (
STATUS-MAP-2026-06-21.md,RFC-ORGANISM-unified.md→ marked stale on this point). - MTP / spec-decode —
G-EAGLE-ACCEPTlive probe shipped (5689e3f); the EAGLE/MTP draft pipeline is wired. STATE's older "parked, needs draft source" framing is stale; FRAMEWORK-INDEX §G/§S is current. - Gate registry — single source is
PPT-LAT-FRAMEWORK-INDEX.md §G;gate-receipts.mdis archived (superseded). - API reference — single source is
PPT-LAT-FRAMEWORK-API.md;PPT-LAT-KEYSTONE-API.mdis superseded. - Systems narrative —
PPT-LAT-Systems-v1.mdis current;PPT-LAT-Systems.md(v0) archived.