Skip to content

Commit 69e4723

Browse files
committed
PK2 wave7: ADR-010 whole-machine balance -- what fills the 12GB (model dominates, KV small, honest correction), levers ranked, CPU layer-offload = the real lever (ADR-011 candidate)
1 parent 318563d commit 69e4723

2 files changed

Lines changed: 120 additions & 0 deletions

File tree

Lines changed: 119 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,119 @@
1+
---
2+
type: design
3+
title: "ADR-010 — Whole-machine balance: what fills the 12 GB, and which offloads actually pay on THIS box"
4+
description: "The systems answer to 'stop fighting VRAM inside the 2060, use the whole NUC11.' Grounded in MEASURED hardware truth (project-perf-wholemachine) + this session's live VRAM numbers. What fills the 12 GB: model weights 8.8 GB (fixed, must stay resident) + KV caches ~2.4 GB (the ONLY movable mass) + buffers. The binding constraint is PCIe gen3 x8 at ~6.2 GB/s, which forbids per-token weight streaming and kills most naive offloads. Each of the operator's ideas tested against physics: CPU/AVX2 matmul (bounded — frees VRAM at ~8× slowdown, only the KV-attention math is a candidate, not weights), iGPU Xe-LP (0.75 TFLOPS, dp4a only — cold-tail only), Optane M10 (block ~1GB/s, NOT PMEM — cold episodes only, never weights/hot-KV), the 24 MB L3 (activation locality, not a store). The ONE lever that pays: the global-KV cache is the entire Pmax-scaling term (~128 KB/token across the 8 full-attention layers), and it is what oversubscribes — bound/spill it (the XBAR global-eviction slab, already GREEN but NOT wired into the served decode; a resident-window cap is the low-risk first step) to free the VRAM that thrashes. The split-crate design (math-core exact substrate / engine CUDA / host-CPU-perf libs / the dormant XBAR slab) is exactly the set of levers — but only the KV-offload one clears the PCIe wall."
5+
tags: [design, adr, whole-machine, vram, offload, pcie, cpu, igpu, optane, kv-cache, xbar, balance, honest-negative]
6+
timestamp: 2026-07-08T00:00:00Z
7+
resource: shannon-prime-lattice/papers/PPT-LAT-ADR-010-WHOLE-MACHINE-BALANCE.md
8+
sp_status: "ANALYSIS (measured, with an honest mid-flight correction: the MODEL fills the card, the KV is small) + auto-fit realized default-off. The real lever = CPU-resident layer offload (ADR-011 candidate)."
9+
sp_gate: "G-PK2-AUTOFIT (measured VRAM breakdown + auto-fit behavior) — tests/perf/G-PK2-AUTOFIT.log"
10+
sp_commit: "builds on project-perf-wholemachine (hardware truth) + ADR-009 (prefill VRAM ceiling)"
11+
sp_repro: "nvidia-smi breakdown @ Pmax=4096 vs 20000; SP_G4_KV_GLOBAL_W cap"
12+
---
13+
14+
# ADR-010 — Whole-machine balance
15+
16+
**Status: DESIGN + one lever realized.** The systems answer to "don't fight VRAM inside the 2060 —
17+
balance the whole NUC11." Grounded in the MEASURED hardware truth ([[project-perf-wholemachine]])
18+
and this session's live VRAM numbers, not architecture-diagram optimism.
19+
20+
## 1. What actually fills the 12 GB (measured — and the correction that matters)
21+
22+
Live, RTX 2060 12 GB, served 12B (`.sp-model` = 8.79 GB on disk). Measured by reading free VRAM
23+
*inside* `gemma4_kv_open` (after weights resident, before the KV/buffer allocs) and again at idle:
24+
25+
| Component | VRAM | Movable? | Why |
26+
|---|---|---|---|
27+
| **Model RESIDENT** (OK_Q4B codes + PLE + packed embd + arena/scratch pools) | **~10.9 GB** | only by moving LAYERS off-GPU (§3) | measured: 10 937 MiB used before any KV alloc |
28+
| **KV cache** (SHARED-KV: only a few owner layers × Pmax) | **~0.78 GB @ Pmax20000** | yes, but small | shared-KV → the global cache is ~2 owner layers, NOT 8; ~0.65 GB |
29+
| **Working buffers** (dq/k/v, dscr, dlog[262144], dseq[Pmax]) | **~0.1 GB** | no | small |
30+
| **Free (dedicated card)** | **~0.57 GB** @ Pmax20000 (11 719 used) || the thin headroom |
31+
| **Batched-prefill scratch** (transient) | **~0.4–1 GB** | ADR-009 | O(n) f32 activations; eats the thin headroom → the batch ceiling |
32+
33+
**★ THE CORRECTION (honest):** my first pass assumed the globals were ~2.5 GB (8 full-attention
34+
layers × Pmax). Wrong — gemma4-12b uses **shared-KV**, so only a handful of owner layers hold a
35+
cache; the KV is **~0.78 GB, not 2.5 GB**. The dominant term is the **MODEL itself at ~10.9 GB
36+
resident** — the 12B *fills the 12 GB card by design*, leaving ~0.57 GB free. Two consequences:
37+
(a) the **original "stall" was co-resident LM Studio** (~6.4 GB) → model+LMStudio = ~15 GB ≫ 12 GB
38+
oversubscription, NOT the KV; (b) the **ADR-009 batch thrash** was the transient batch scratch
39+
(~0.4–1 GB) eating the thin ~0.57 GB headroom. The KV cache was never the movable mass — **the
40+
model is.**
41+
42+
## 2. The binding constraint (why most offloads don't pay)
43+
44+
**PCIe gen3 x8 ≈ 6.2 GB/s real** (measured `_pcie_bw.cu`; NOT x16, NOT Gen5). To run a layer
45+
somewhere other than the 2060 you must move either its **weights** or its **activations** across
46+
that bus every token. Weights: the 8.8 GB model at 6.2 GB/s = **~1.4 s/token** just to stream —
47+
catastrophic. So **weights must stay resident**; no CPU/iGPU/Optane weight offload pays. Only data
48+
that is (a) small per token, or (b) accessed with locality / infrequently is a candidate: the KV
49+
cache and cold episodes. Everything below follows from this one number.
50+
51+
## 3. The operator's ideas, each against the physics (re-judged on the corrected breakdown)
52+
53+
- **"Run the matmul on the CPU (24 MB SVM) / system RAM."****This is the relevant lever —
54+
because the MODEL, not the KV, fills the card.** The i9-11900KB has **24 MB L3** (that's the
55+
"SVM"), 8 cores, ~40 GB/s DDR4. To free real VRAM you must evict *weights*, and the only weights
56+
that can move are whole LAYERS: keep a subset of layers' OK_Q4B weights in **host RAM** (31.5 GB
57+
free) and **compute those layers on the CPU** (AVX2/OpenMP — the `build-cpu-perf` libs already
58+
ship, the qwen36 lane uses them), exchanging only the E-sized activation (~15 KB/token/layer)
59+
over PCIe. Cost: those layers run ~8–10× slower (40 GB/s DRAM vs the 2060's ~336 GB/s), but each
60+
offloaded layer frees ~0.2 GB VRAM. **Verdict: a genuine VRAM-for-latency trade** — offload k
61+
layers → free ~0.2k GB → slower by ~(k/48)×8×. Worth it to (a) create headroom that stops the
62+
batch/co-tenant thrash, or (b) fit a *second* small model / longer everything. NOT a speedup;
63+
a rebalance. The `_int8` weight streaming path (per-token weight fetch over PCIe) stays refuted
64+
(~1.4 s/token) — the layer must COMPUTE where its weights live (CPU), not stream them to the GPU.
65+
- **"Offload the ring to RAM / Optane."** Re-judged: the KV is only **~0.78 GB**, so this frees
66+
little. The XBAR host-spill + LSH slab (GREEN, un-wired) is still architecturally right for
67+
*very long contexts* (where even shared-KV grows), but it is NOT the VRAM lever on today's
68+
persona-sized chats — the model is. Deferred to when context length, not the model, is the term.
69+
- **"Offload weights to Optane."** Optane M10 = **block NVMe ~1 GB/s, NOT PMEM** — slower than the
70+
PCIe it feeds. Refuted for weights. Right role = the **cold XBAR episode store**. (Note: for the
71+
CPU-layer-offload above, the offloaded weights live in **DRAM**, not Optane.)
72+
- **"Shard experts across devices via CRT."** CRT splits *numbers*; a GEMM runs in FULL in each
73+
prime channel → every device needs ALL weights → per-device cut = zero. Refuted as a sharding
74+
mechanism. CRT's gift = *bit-exactness across devices* (the enabler of a clean split), not the split.
75+
- **"iGPU (Xe-LP)."** 32 EU, no XMX, ~0.75 TFLOPS, dp4a only — 10–25× weaker; feeding it costs the
76+
same activation exchange as the CPU but with less RAM and no AVX2 maturity. The CPU is the better
77+
offload target on this box. iGPU = cold-tail only.
78+
79+
## 4. The verdict, ranked (corrected)
80+
81+
1. **CPU-resident layer offload** (weights in DRAM, compute on CPU AVX2, exchange only activations).
82+
The one lever that frees *real* VRAM, because the model is what fills the card. VRAM-for-latency;
83+
the `build-cpu-perf` machinery exists. The big-but-correct project.
84+
2. **Auto-fit the KV to free VRAM** (`SP_G4_KV_AUTOFIT`, realized §5) — the co-tenant safety clamp:
85+
size Pmax to what's actually free so a shared card (the original LM-Studio stall) never
86+
oversubscribes. Small VRAM effect (KV is small) but directly fixes the *original stall's cause*.
87+
3. **Chunked batch scratch** (ADR-009 follow-on) — cap the transient O(n) f32 scratch so the batch
88+
win fires without eating the thin headroom.
89+
4. **XBAR host-spill KV slab** — for *long-context* regimes (deferred; not today's term).
90+
5. **Optane = cold episodes, iGPU = cold-tail, CRT = exactness** — already in role.
91+
92+
## 5. The realized first step — auto-fit the KV to free VRAM (G-PK2-AUTOFIT)
93+
94+
`SP_G4_KV_AUTOFIT=1` (default-off = the passed Pmax verbatim = **null floor**): at
95+
`gemma4_kv_open`, read free VRAM (weights already resident) and clamp Pmax so the KV cache + a
96+
margin fit — so a **shared card never oversubscribes** (the original stall = co-resident LM Studio;
97+
autofit would have clamped Pmax to fit the remaining VRAM instead of thrashing). Small effect on a
98+
dedicated card (the 12B already fits at Pmax=20000 with ~0.57 GB free), decisive on a shared one.
99+
Honest limitation found in the gate: on a nearly-full card the free-read at open under-estimates
100+
(the arena/scratch pools allocate lazily), so the clamp is conservative — margin-tunable via
101+
`SP_G4_KV_AUTOFIT_MARGIN_MB`. It never RAISES Pmax and never breaks the dedicated-card path.
102+
103+
## 6. The real conclusion (the operator's thesis, confirmed — and it points at the CPU)
104+
105+
The split multi-crate design **is** the set of levers. But the corrected measurement flips which
106+
one matters: **the 12B model at ~10.9 GB resident fills the 12 GB card — the KV cache is a rounding
107+
error (~0.78 GB), and the ~0.57 GB headroom is the whole game.** So the physics verdict is:
108+
- **You cannot free meaningful VRAM by moving DATA (KV/episodes) — there isn't much.** You free it
109+
by moving **WEIGHTS**, and PCIe x8 forbids streaming them per token — so the weights must
110+
**compute where they live**. That is the **CPU-resident layer offload**: a few layers' weights in
111+
40 GB/s DRAM, computed on 8 AVX2 cores, exchanging only the tiny activation. The operator's
112+
instinct ("run the matmul on the CPU… balance of the entire system") is **correct** — it is the
113+
one lever that rebalances the dominant mass, at a bounded latency cost.
114+
- The split design enables it cleanly: the math-core forward is already CPU-capable (the reference
115+
path), the `build-cpu-perf` AVX2/OpenMP libs exist, and CRT guarantees the CPU and GPU legs stay
116+
**bit-identical** — so a hybrid CPU+GPU forward is drift-free and auditable, which is exactly the
117+
property that makes a split trustworthy. That is the next real project (ADR-011 candidate):
118+
a per-layer placement policy (GPU for the hot majority, CPU-resident for the cold tail that buys
119+
back headroom), gated on end-to-end tok/s and VRAM freed.

papers/VERIFIED-SCOREBOARD.md

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -54,6 +54,7 @@ Method: a read-only fleet checked each claim against (a) a commit that resolves
5454
| **Flywheel persistence of spine receipts** (telemetry-okf `kind:"spine"`) | **GREEN offline, default-on (cheap+safe)** | harness (wave 5) | harness `tests/G-PK2-FLYWHEEL.log` (6/6) | VERIFIED — spine receipts flush into the durable telemetry-okf tier via the EXISTING `TelemetrySink` (content-addressed, idempotent, watermarked incremental; anti-rebuild). Flushed each agency tick + gateway turn; the ADR-005 flywheel corpus now includes what the harness DECIDED + whether verify held. Private-secret never routes through spine payloads by construction |
5555
| **recall∘L5 composition** (the two recall authorities, both armed) | **FREE composition REFUTED → one-authority guard ENFORCED, LIVE-GREEN 6/6** | harness (wave 5) | `G-PK2-RECALL-L5-COMPOSE-FREE.log` (the honest negative, 4/6) + **`G-PK2-RECALL-L5-COMPOSE.log` (enforced, 6/6 live)** | VERIFIED — Phase A (both armed): L5's systemecho CROSS-PICKED color-adjacent queries from the counterfact corpus and overrode the harness note ("favorite color?" → "Human blood is green"; "sky?" → "Green") — L5 selection cross-picks are a known daemon residual; composition surfaces them user-visibly. Phase B: the gateway auto-disarms spine recall whenever the request arms L5 (`auto_recall=true`), with an `{"authority":"L5"}` receipt event; guard held, L5 authority intact ("Lyon"), harness authority faithful when L5 off ("Teal."). Plus `auto_recall` body passthrough (default false). The one-authority rule is now STRUCTURAL, not operator discipline |
5656
| **ADR-009 batched prefill under the SWA ring** (`SP_KV_PREFILL_BATCH`, the ring-off precondition lifted) | **LIVE-GREEN 3/3 — bounded ~7× win + graceful fallback; default-off** | engine (wave 6) | `tests/perf/G-PK2-BATCHRING.log` (3/3) | VERIFIED — `PPT-LAT-ADR-009-PREFILL-SPEED.md`. The batched cold prefill was BLOCKED under the ring (so the ring config never used it); wave 6 lifts it (ring-layout `k_ring_sink` for SWA owners + retained contiguous shared-owner K/V for sharers). Live @ PMAX=4096: **n≈837 → 6.7s coherent (~7× vs ~39s per-token)**; n≈1765 → VRAM guard DECLINES → graceful per-token fallback 99s coherent (null floor, live-proven). Two guards: VRAM (`SP_KV_BATCH_VRAM_MARGIN_MB`) + persist (declines under `SP_PERSIST_KV`, since a batched-ring cache isn't a valid persist-continuation base). Verdict across 3 waves: **batching is a bounded cold-prefill win, not the general lever** — the f32 activation scratch caps it; cublasGemmEx-int8 would hit the same ceiling (not pursued). Default-off; production per-token path unaffected |
57+
| **ADR-010 whole-machine balance** (measured VRAM breakdown + KV auto-fit; the CPU-offload verdict) | **ANALYSIS (measured) + auto-fit realized default-off** | engine (wave 7) | `tests/perf/G-PK2-AUTOFIT.log` | VERIFIED (measurement) — `PPT-LAT-ADR-010-WHOLE-MACHINE-BALANCE.md`. Measured live: **model ~10.9 GB resident FILLS the 12 GB card; KV cache only ~0.78 GB (shared-KV — NOT the 2.5 GB I first assumed); ~0.57 GB free.** ★HONEST CORRECTION: the original ">1000-tok stall" root was co-resident LM Studio (~6.4 GB oversubscription), and the ADR-009 batch thrash was the transient scratch eating the thin headroom — the KV was never the movable mass, THE MODEL IS. Every weight-offload refuted by PCIe x8 (6.2 GB/s forbids per-token weight streaming); the ONE lever that frees real VRAM = **CPU-resident layer offload** (weights in 40 GB/s DRAM, compute on AVX2, exchange only activations; CRT keeps CPU+GPU legs bit-identical) — the ADR-011 candidate. `SP_G4_KV_AUTOFIT` (default-off) clamps Pmax to free VRAM for the co-tenant case; dedicated card serves coherent ("Paris" 12.4s), null floor intact |
5758

5859
## The open frontier (what we actually skipped)
5960

0 commit comments

Comments
 (0)