lobes runs the fleet (four co-resident vLLM backends + a gateway, by
default — a mesh-brain deployment shape can drop one of them to a peer box;
see docs/deployment-shapes.md) with knob values tuned to the hardware it
lands on. A machine profile is the per-card tuning declaration: which
models serve each role, their
GPU memory budget, context length, attention backend, and other vLLM knobs
the compose template substitutes. The profile schema now covers six core
roles (cortex / senses / muse / worker / embedder / reranker) —
muse being the opt-in-hosted seventh Colleague role and worker the
opt-in-hosted eighth (thor-worker-lobe plan), both of whose full declarations
live in a deployment shape, not a card profile (see "muse: shape-declared,
not card-declared" below, which now covers worker too). This document
walks the detection flow,
how a profile is chosen, the knobs' meanings and provenance, and how to
write custom profiles for new hardware.
See lobes explain profiles or lobes explain tuning for the brief version.
That command reads from lobes/explain/catalog.py; this file is the deep
reference.
The card-detection flow below runs only during lobes init — either
auto-detected from the card, or forced with an explicit --profile <name> —
to resolve the per-role Profile this document describes and render it into
.env. lobes serve and lobes status run no detection at all and
accept neither --profile nor --machine: they only read the deployment's
already-rendered .env (including the persisted LOBES_PROFILE) — profile
selection happens once, at init time, not on every serve/status call.
lobes switch has its own, separate --machine flag — a different,
legacy single-model machine-profile knob (lobes/profiles/__init__.py's
MachineProfile / resolve_machine(), tuning the single-model template's
GPU/context defaults), not the fleet Profile system covered here; left at
its default (auto), it re-detects the GPU name via nvidia-smi -L on its
own code path, not lobes/runtime/_detect.py::detect_card().
The following steps resolve the machine name at lobes init time:
-
Gather raw facts (from
lobes/runtime/_detect.py::detect_card()):- Device name: from
nvidia-smi --query-gpu=name,compute_cap - Compute capability: from the same nvidia-smi query (e.g.,
"11.0"→"sm_110") - Total memory: from
/proc/meminfo MemTotal(honest system total on unified-memory Jetson boards; nvidia-smi memory fields report[N/A]on these boards, so we never parse them) - Hostname: from
socket.gethostname()(e.g.,"thor") - Device-tree model (Jetson boards only): from
/proc/device-tree/model, NUL-terminated (e.g.,"NVIDIA Jetson AGX Thor")
Every probe is independent and degrades gracefully: a missing binary, a timeout, a missing file, or unparsable output never raises. The corresponding fact is simply
None, and resolution continues. - Device name: from
-
Resolve to a card name (from
lobes/machines/_registry.py::detect()):- Pass the gathered facts to the live
lobes.machinesregistry (a dict ofCardStrategyinstances, one per supported chip, in detection-precedence order). - The first card's
DetectionSignaturethat matches wins. Signatures are matched byname_markers(device name or hostname substrings) orcompute_capability(e.g.,"sm_110"). - If no card matches, the result is
UNKNOWN.
- Pass the gathered facts to the live
-
Never guess a fallback silently. An
UNKNOWNresult is first-class and honest — it is not replaced by a "closest" card. A deployment on unrecognized hardware gets an explicit warning (inlobes init/lobes statusoutput), and the conservativebaseprofile is used to avoid OOM crashes on first boot.
Once detection (or an explicit --profile flag on lobes init) resolves the
card name, a profile is looked up via lobes/profiles/loader.py::resolve_profile():
-
Explicit always wins: if
--profile <name>was given (or aLOBES_PROFILEenv var), that profile is used, even if it diverges from the auto-detected machine name. A warning is printed if forced. -
Operator override next: if a file
<deployment-dir>/profiles/<name>.tomlexists (where<name>is the machine name or explicit--profile), it wins over the built-in. -
Built-in fallback: the packaged profile in
lobes/profiles/builtin/, resolved by name.
A same-named operator file completely overrides the built-in — they are
never merged field-by-field. The operator profile is a self-contained
declaration; any role or knob it stays silent on falls back to "no opinion",
meaning the compose template's own ${VAR:-default} applies.
Every knob is optional (None = "profile takes no position, template default
applies"). Only knobs the profile diverges on appear in the TOML or env; the
rest are inherited from the template.
The six roles and seven knobs map to env vars via lobes/profiles/render.py:
| role | env prefix | feasible | model | gpu_mem_util | max_model_len | quantization | kv_cache_dtype | attention_backend | enforce_eager | max_num_seqs |
|---|---|---|---|---|---|---|---|---|---|---|
cortex |
PRIMARY_ |
FEASIBLE |
MODEL |
GPU_MEM_UTIL |
MAX_MODEL_LEN |
QUANTIZATION |
KV_CACHE_DTYPE |
ATTENTION_BACKEND |
ENFORCE_EAGER |
MAX_NUM_SEQS |
senses |
MULTIMODAL_ |
FEASIBLE |
MODEL |
GPU_MEM_UTIL |
MAX_MODEL_LEN |
QUANTIZATION |
KV_CACHE_DTYPE |
ATTENTION_BACKEND |
(not used) | MAX_NUM_SEQS |
muse |
MUSE_ |
FEASIBLE |
MODEL |
GPU_MEM_UTIL |
MAX_MODEL_LEN |
QUANTIZATION |
(not used) | ATTENTION_BACKEND |
(not used) | (not used) |
worker |
WORKER_ |
FEASIBLE |
MODEL |
GPU_MEM_UTIL |
MAX_MODEL_LEN |
QUANTIZATION |
(not used) | ATTENTION_BACKEND |
(not used) | (not used) |
embedder |
EMBED_ |
FEASIBLE |
MODEL |
GPU_MEM_UTIL |
MAX_MODEL_LEN |
(not used) | (not used) | ATTENTION_BACKEND |
(not used) | (not used) |
reranker |
RERANK_ |
FEASIBLE |
MODEL |
GPU_MEM_UTIL |
MAX_MODEL_LEN |
(not used) | (not used) | ATTENTION_BACKEND |
ENFORCE_EAGER |
(not used) |
muse and worker are the two opt-in core roles
(lobes/profiles/shapes.py's OPT_IN_CORE_ROLES = ("muse", "worker")): each
carries the full per-machine knob set above, but the built-in card
profiles stay silent on both — spark.toml and thor.toml declare
nothing for [roles.muse] or [roles.worker], because no card hosts either
heavy lobe alongside the default cortex+senses duo. The one exception is
base.toml, which vetoes both ([roles.muse] feasible = false /
[roles.worker] feasible = false) under its conservative unknown-card rule,
exactly as it already disables senses. The full muse/worker declaration
(model, budget, quantization, attention backend) lives in the
muse-/worker-hosting deployment shape's own overrides — thor-muse for
muse, thor-worker (forthcoming, thor-worker-lobe plan task t7) for
worker — see docs/deployment-shapes.md — while
every non-hosting shape renders nothing for either role at all, passing the
card's own declaration through verbatim: base.toml's veto marker survives,
and a card that stays silent on muse/worker renders no MUSE_*/WORKER_*
line. thor-muse's budget values are measured (2026-07-17 live boot on
the physical Thor: util 0.55 at the full 262144 window), though the shape
itself stays UNVALIDATED, and is now additionally DORMANT — the box that
measured it moved to hosting worker instead; see
docs/gemma-4-31b-nvfp4.md. thor-worker's budget
values are not yet measured — no number is declared here; see
docs/qwen3.6-35b-a3b-nvfp4.md.
What it does: If false, the role is marked infeasible on this box.
lobes/profiles/render.py renders it as <PREFIX>_FEASIBLE=false in the
.env, and (once wired up in t6) the gateway omits it from GET /capabilities
and returns 404 role_infeasible on POST requests. If true (the default),
no <PREFIX>_FEASIBLE key is emitted — "feasible" is the assumed default.
When set:
sparkprofile (GB10): all four default rolesfeasible=true(all load-tested here); silent onmuseandworker(both shape-declared — see above).thorprofile (Jetson AGX Thor): all four default rolesfeasible=true(validated live 2026-07-13); silent onmuseandworker(both shape-declared — see above).baseprofile (unknown card):cortex=true,senses=false,muse=false,worker=false,embedder=true,reranker=true— the multimodal gear, the 31B muse lobe, and the 35B-A3B worker lobe are all disabled to save memory on unknown hardware.
What it does: The HuggingFace model id (e.g.,
sakamakismile/Qwen3.6-27B-Text-NVFP4-MTP). Rendered to two env vars:
<PREFIX>_MODEL and <PREFIX>_SERVED_NAME (both set to the same value). The
compose template passes the served name to vLLM's --served-model-name
separately from the model id it downloads; the two must agree for the gateway
to route correctly.
When set:
sparkcortex:sakamakismile/Qwen3.6-27B-Text-NVFP4-MTP(the 256K-native MTP primary, re-exported with its draft head restored; load-tested 2026-05-31 on GB10, ~2.4x single-stream decode over the archived 27B baseline).sparksenses:coolthor/gemma-4-12B-it-NVFP4A16(12B multimodal, load-tested on GB10).sparkembedder:Qwen/Qwen3-Embedding-0.6B(0.6B dense embedder, 1024-dim, 32K context).sparkreranker:Qwen/Qwen3-Reranker-0.6B(0.6B cross-encoder).thorcortex/senses/embedder/reranker: same as spark (both are 128 GB unified Blackwell-class boards; Thor diverges only on vLLM knobs, not model ids).basecortex:Qwen/Qwen3.5-4B(small model for unknown hardware — seedocs/qwen3.5-4b-minor.md, 256K native, bf16).basesenses: omitted (infeasible; seefeasible=falseabove).baseembedder/reranker: same as spark.
What it does: The fraction of total GPU memory lobes tells vLLM the
generate/pooling lane may use. vLLM uses this to cap --max-num-seqs (batch
size) and prefill throughput so OOM never happens. On unified-memory boards
(Spark, Thor) the "GPU memory" is actually a carved-out slice of the unified
pool; on dedicated-VRAM boards (RTX PRO 6000) it is true discrete VRAM.
When set and why:
sparkcortex:0.30— the primary serves at this utilization budget (0.30 + 0.14 + 0.06 + 0.06 = 0.56total across all four roles on 128 GB), chosen to leave headroom for other mesh services on the shared GB10 (load-tested 2026-05-31).sparksenses:0.14— the 12B multimodal gear, smaller budget than cortex.sparkembedder/reranker:0.06each — the two ~0.6B pooling gears.thorcortex/senses/embedder/reranker: same as spark (Thor is also 128 GB unified; the budget is hardware-independent, not a divergence).basecortex:0.30— safe default for unknown hardware (small model, so conservative util).baseembedder/reranker:0.06each.
What it does: The maximum sequence length (in tokens) vLLM will accept for this role. Determines GPU memory overhead; longer sequences = larger KV cache. Tuned per card to balance headroom and capability.
When set and why:
sparkcortex:131072(128K) — the primary's load-tested context on the GB10 (the checkpoint has 256K native, but at util 0.30 on a shared board, 128K is the validated cap; issue #107 tracks broader context migration).sparksenses:32768(32K) — the multimodal gear's context.sparkembedder/reranker:8192(8K) — the pooling gears' context.thorcortex:131072(128K) — same as spark (both boards are 128 GB, same fleet budget; cortex is util-bound, not context-bound at these numbers).thorsenses/embedder/reranker: same as spark.basecortex:32768(32K) — the small 4B model's cap on unknown hardware, conservative.baseembedder/reranker:8192(8K) — same as spark.
What it does: The quantization format. Passed to vLLM's --quantization
flag. Per-model semantics (defined in lobes/catalog.py); some models have
multiple quantization options.
When set and why:
sparkcortex:"modelopt"— the MTP primary is an nvidia/ ModelOpt FP4 export; requires this quantization.sparksenses:"compressed-tensors"— the Gemma 4 checkpoint uses compressed-tensors format.sparkembedder/reranker: omitted (these models do not require explicit quantization on vLLM nv26.04).thorcortex/senses: same as spark.thorembedder/reranker: same as spark.basecortex: omitted (the 4B model is bf16/unquantized, not quantized; seedocs/qwen3.5-4b-minor.md).baseembedder/reranker: omitted (same as spark).
What it does: The dtype of the KV cache buffers in vLLM. Traded against
accuracy for memory savings. Passed to vLLM's --kv-cache-dtype flag.
When set and why:
sparkcortex:"fp8"— FP8 KV cache saves memory and is validated on this checkpoint/board pairing (load-tested 2026-05-31).sparksenses: omitted (the compose template's default, typically"float16"or None, applies).sparkembedder/reranker: omitted.thorcortex:"auto"— the hand-validated value the live Thor fleet runs with. The checkpoint ships no calibrated KV scales (kv_cache_quant_algo: null); an earlier session recorded anassert layer.k_scale > 0.0crash underfp8on Thor, but on the currently pinned nightly that assert did not reproduce (2026-07-13):fp8now boots with "Using KV cache scaling factor 1.0" / "uncalibrated q_scale" warnings — an accuracy risk, not a crash.autoserves the KV cache in the model dtype and avoids the uncalibrated-fp8 question entirely. Provenance: the hand-edited deployment that ran this box for weeks, plus the 2026-07-13 boot logs recorded in issue #109 (which also tracks the GB10 side of the question).thorsenses/embedder/reranker: omitted (same as spark).basecortex/senses/embedder/reranker: omitted.
What it does: The attention kernel backend. vLLM chooses at inference time based on the attention shape and the backend's capabilities; the flag provides an override or force. Affects both speed and stability.
When set and why:
sparkcortex: omitted (vLLM auto-selects; FlashInfer on the GB10).sparkandthorsenses:"TRITON_ATTN"— not a Thor divergence. This mirrors the fleet template's cross-machine default for the Gemma gear: Gemma 4's heterogeneous head sizes (sliding 256 / full-attention 512) need the Triton backend. Same value on every machine; its env plumbing (MULTIMODAL_ATTENTION_BACKEND) is unchanged pending the GB10 check in issue #109.sparkembedder/reranker: omitted (vLLM auto-selects on the GB10).thorcortex: omitted (the generate lane runs stably with the auto-selected backend on Thor).thorembedder:"TRITON_ATTN"— a real sm_110 divergence: the auto-picked FLASH_ATTN pooling path hangs the forward pass on sm_110 (requests accepted, never answered —/healthstays green, which is why the correctness probes exist). Inherited from the sharedSM_110trait inlobes/machines/_traits.py, overlaid bylobes/profiles/loader.py.thorreranker:"TRITON_ATTN"— the same sm_110 pooling bug, surfaced as NaN relevance scores (wrong orderings, issue #105) rather than a hang. Inherited from the SM_110 trait.basecortex/senses/embedder/reranker: omitted (template defaults apply; the conservative small model on unknown hardware avoids tuning for any backend).
What it does: If true, vLLM disables CUDA graph capture and runs every
request in eager mode. Slower but more stable. Passed to vLLM as
--enforce-eager (or --no-enforce-eager if false). Note: the env var
holds the token, not a bare boolean (see lobes/profiles/render.py for
the translation).
When set and why:
sparkcortex/senses/embedder: omitted (CUDA graphs are stable on GB10).sparkreranker: omitted (template default; CUDA graphs work here).thorreranker:true— validated live on Thor 2026-07-13. With Triton attention, the reranker's CUDA-graph capture is unstable on sm_110: requests fail withcudaErrorLaunchFailurewhen CUDA graphs are enabled. Enforce eager mode (no graphs) and the reranker runs stably and correctly. Provenance: live reranker correctness probe (ordering test) crashed with graph capture enabled but passed withenforce_eager=true; see issue #105 comment for the crash log. Inherited from the SM_110 trait.thorcortex/senses/embedder: omitted (CUDA graphs are stable elsewhere on Thor).basecortex/senses/embedder/reranker: omitted.
What it does: The maximum number of sequences (batch size) vLLM will
schedule concurrently in a single forward pass. Tuned per workload
(--purpose), not per machine. Affects both throughput (more batches) and
latency (higher tail latency per request).
When set and why:
sparkcortex:2— the fleet's cortex defaults to small batches for low latency (appropriate for the agentic workload — agents mostly read one response at a time, not many concurrent requests).sparksenses/embedder/reranker: omitted (template defaults apply, typically 4–8 for pooling layers).thorcortex:2— same as spark (machine-independent; this knob is workload-tuned, not machine-tuned, so thor and spark match).thorsenses/embedder/reranker: omitted (same as spark).basecortex/senses/embedder/reranker: omitted.
Operator-defined profiles go in <deployment-dir>/profiles/<name>.toml (the
file's stem is the profile name). They are discovered automatically by
lobes init (the only verb that resolves a Profile — see "How detection
works" above) and override any built-in with the same name.
A profile is a self-contained TOML declaration. Example:
name = "my-custom-box"
summary = "Custom RTX 6000 tuning with larger context"
[roles.cortex]
feasible = true
model = "sakamakismile/Qwen3.6-27B-Text-NVFP4-MTP"
gpu_mem_util = 0.75
max_model_len = 180000
quantization = "modelopt"
kv_cache_dtype = "fp8"
attention_backend = "flashinfer"
[roles.senses]
feasible = true
model = "coolthor/gemma-4-12B-it-NVFP4A16"
gpu_mem_util = 0.15
max_model_len = 32768
# attention_backend inherited from template default
[roles.embedder]
feasible = true
model = "Qwen/Qwen3-Embedding-0.6B"
gpu_mem_util = 0.05
max_model_len = 8192
[roles.reranker]
feasible = true
model = "Qwen/Qwen3-Reranker-0.6B"
gpu_mem_util = 0.05
max_model_len = 8192Rules:
- File name is the profile name.
profiles/my-custom-box.toml→ profile name ismy-custom-box. Use it aslobes init --profile my-custom-box --apply(the flag is--profileoninit,--machineonswitch). - Inline
namefield must match the file stem. If the file declaresname = "something-else", profile loading fails with an error (this prevents confusion from stale copies). - Roles are optional. A role omitted from the TOML means "take no position
— use the template's defaults for this role." You don't have to restate all
six roles; a minimal profile might only touch
cortex. (Staying silent onmuse/workeris the norm — the built-in card profiles do; a muse-/worker-hosting shape carries each role's own declaration.) - Knobs within a role are optional. A knob omitted means "no opinion" (the
template's
${VAR:-default}applies). This lets you override only the knobs you care about and inherit sensible defaults for the rest. - Unknown knob names are errors. A typo in a knob name (e.g.,
gpu_memory_utilinstead ofgpu_mem_util) causes profile loading to fail with remediation (list of known knobs). This is strict-by-design to catch operator mistakes early. - Profile override is complete, not merged. If you define a custom
my-custom-boxprofile and a built-inmy-custom-boxexists, the custom one replaces it entirely; the two are never merged field-by-field.
To use your profile:
# Explicit flag (--profile on init, --machine on switch) or LOBES_PROFILE env var:
lobes init --profile my-custom-box --apply
lobes switch my-model --machine my-custom-box --apply
# Automatic if the detected machine name matches:
# (detection resolves "my-custom-box" → looks for profile "my-custom-box" →
# finds it in <deploy-dir>/profiles/my-custom-box.toml)Auto-detection only works if the detected machine name (from
lobes/machines/_registry.py) matches the profile name. If you want a custom
profile to be auto-selected on a new card, add a new CardStrategy module to
lobes/machines/ (following the spark.py / thor.py pattern).
For a complete worked example — a live-validated operator profile for the
Jetson AGX Orin 64GB (Gemma senses at its full 128K context on Ampere sm_87,
with measured knob values and the Jetson divergences found on the way) — see
orin-profiles.md.
CI guards against cross-machine breakage via golden .env files — one per
built-in profile. Every PR that edits profile or machine-strategy code must
regenerate them:
uv run python tests/goldens/regen.pyThis reads every built-in profile, renders it to .env format, and byte-diffs
against tests/goldens/<profile>.env (e.g., tests/goldens/spark.env).
The invariant: editing the thor profile or the SM_110 trait must not
change the spark golden (i.e., spark.env stays identical). This ensures that
work on one machine's profile does not accidentally break another's. If you edit
machine strategy or profile code and the golden diff fails CI, re-run the regen
script (it updates the goldens to the new render output) and commit the change.
The template-defaults.env golden verifies that the compose template's own
built-in defaults (when no profile is loaded) render as expected.
lobes is deployed and validated on the following hardware, with the stated
caveat profile as the truthful, tested configuration. Aspirational entries are
marked clearly as unvalidated. This table's scope is BUILT-IN
profiles/shapes (lobes/profiles/builtin/*.toml,
lobes/profiles/builtin_shapes/*.toml) — the #108 rule is that a built-in
earns "validated" status only via a repo-recorded boot of that exact
built-in, so a validated operator-defined profile (<deploy-dir>/profiles/ <name>.toml) does not by itself validate a built-in for the same card. The
Orin row below carries both facts side by side so they read as complementary,
not contradictory.
| card | machine | profile | status | validation |
|---|---|---|---|---|
| DGX Spark (Grace Blackwell, 128 GB unified) | spark |
spark.toml |
load-tested | 2026-06-03 — lobes benchmark ran at ~7.8–8.0 tok/s decode (27B primary, util 0.30, single-stream) on the fleet duo (cortex 128K, senses 32K) with FlashInfer attention. The correctness probes postdate that run and have not been executed on the GB10 — rerank ordering there is explicitly unverified (issue #106). See docs/tuning-profiles.md and the Spark run on docs/qwen3.6-27b-text-nvfp4-mtp.md. |
| Jetson AGX Thor (Blackwell-class sm_110, 128 GB unified) | thor |
thor.toml |
load-tested | 2026-07-13 — the three correctness probes (cortex known-answer, embed paraphrase-ranking, rerank ordering) all pass with the profile's four divergences (cortex kv_cache_dtype=auto, embedder TRITON_ATTN, reranker TRITON_ATTN + enforce_eager=true); senses' health was not confirmed in that run (it boots last and the concurrent-boot race, caveat 4 below, can catch it). |
| unknown card (UNKNOWN) | generic |
base.toml |
conservative fallback | — — no hardware beyond Spark/Thor is load-tested; UNKNOWN cards are served a small 4B model, no 27B, and no multimodal (senses disabled). Avoids OOM crashes on first boot. See issue #107 for broader "small default on every card" work. |
| Jetson AGX Orin / Orin Nano Super | (not yet) |
(not yet) | no built-in profile — unvalidated at that scope; operator-validated separately | Built-ins: no built-in orin profile exists, and the built-in orin-small shape remains DECLARED/UNVALIDATED (#108 — an unbooted built-in stays unvalidated regardless of operator activity elsewhere). Do not claim built-in Orin support from this row. Operator deployment: a physical Jetson AGX Orin 64GB WAS operator-validated 2026-07-16/17, using a hand-written operator profile (<deploy-dir>/profiles/orin.toml — senses/embedder/reranker, senses at its full native 131072 context, measured gpu_mem_util=0.45) composed with the built-in thor-lobe shape (not orin-small). See orin-profiles.md for the measured knobs and the Jetson/sm_87 divergences found live. |
This section documents Thor's validated divergences from Spark's baseline, with honest provenance. Thor is load-tested and ready for production, but with these specific knob values; they exist because of real hardware constraints.
The problem: The checkpoint ships no calibrated KV scales
(kv_cache_quant_algo: null). An earlier session on this box recorded an
assert layer.k_scale > 0.0 crash under the Spark default fp8; on the
currently pinned nightly that assert did not reproduce (2026-07-13) —
fp8 boots with Using KV cache scaling factor 1.0 for fp8_e4m3 /
uncalibrated q_scale warnings instead, i.e. an accuracy risk rather
than a crash. Whether the GB10 has the same exposure is the open half of
issue #109.
The fix: Set cortex kv_cache_dtype=auto — the KV cache is kept in the
model dtype, sidestepping uncalibrated-fp8 entirely.
Validation: the hand-tuned deployment that has served this box runs
auto; on the clean-boot validation (2026-07-13) the cortex known-answer
probe passed under auto. No probe was run under fp8 — the fp8 evidence
is boot-log warnings, not a measured accuracy delta.
Provenance: Measured on Thor (sm_110, 128 GB unified, Jetson AGX).
Carried by lobes/machines/thor.py and overlaid by
lobes/profiles/loader.py from the live registry (not hardcoded in TOML).
The problem: On sm_110 (Thor), the auto-picked FLASH_ATTN pooling path is
broken: on the embedder it hangs the forward pass (requests accepted,
never answered — /health stays green, which is exactly why the correctness
probes exist); on the reranker the same bug surfaces as NaN relevance
scores, i.e. wrong orderings (issue #105).
The fix: Force attention_backend=TRITON_ATTN for the two pooling roles.
(senses also runs TRITON_ATTN, but that is the cross-machine template
default for Gemma 4's heterogeneous head sizes — identical on Spark, not a
Thor divergence.)
Validation: Live correctness probes on the clean-boot validation
(2026-07-13): the embed paraphrase probe passed (cos(paraphrase) 0.89 >
cos(unrelated) 0.23, no hang) and the rerank ordering probe ranked the
relevant document first, with no cudaErrorLaunchFailure (see caveat 3).
Provenance: The SM_110 trait (a compute-capability characteristic, not
Thor-only — a future sm_110 board inherits it) in
lobes/machines/_traits.py defines the knobs for both pooling roles;
lobes/machines/thor.py composes the trait. Issue #105 documents the
symptoms; issue #106 tracks checking rerank ordering on the GB10 (still
open).
The trait does NOT cover the opt-in
embed-deepgear — set it by hand on sm_110.
SM_110keys its knobs by profile role name ("embedder","reranker"—lobes/machines/_traits.py:22,28).embed-deepis a gear, not a role: it is deliberately outsidelobes.profiles.schema.ROLES, so it has noRoleProfileand no trait can reach it. ItsEMBED_DEEP_ATTENTION_BACKENDtherefore stays at the template defaultautoon every machine, including sm_110 — whereautoresolves to the broken FLASH_ATTN pooling path.A third pooling lane is subject to the identical bug, and it fails the same insidious way: the forward pass hangs while
/healthstays green. An operator hostingembed-deepon Thor (or any future sm_110 board) MUST set it explicitly in.env:EMBED_DEEP_ATTENTION_BACKEND=TRITON_ATTNThis is the known, accepted cost of
embed-deepbeing a gear rather than an eighth role — every opt-in gear (minor,middle,multimodal-coder) shares it;embed-deepis the first where the consequence is a silent hang rather than a suboptimal knob. UNVALIDATED on sm_110 either way: no Thor has bootedembed-deep, so the TRITON_ATTN workaround is inferred from the embedder's measured behaviour, not measured for this gear (#108).
The problem: On sm_110, even with Triton attention, the reranker's CUDA
graph capture is unstable. When graph capture is enabled (the default), some
requests fail with cudaErrorLaunchFailure.
The fix: Set reranker enforce_eager=true. This disables CUDA graph
capture and runs every request in eager mode, avoiding the error.
Validation: The cudaErrorLaunchFailure under CUDA graphs was recorded in
issue #105 (earlier session on this box). On the clean-boot validation
(2026-07-13) with enforce_eager=true the rerank ordering probe passed with
no engine crash.
Provenance: Measured on Thor. The SM_110 trait in lobes/machines/_traits.py
defines reranker enforce_eager=true as a sm_110 characteristic. It is
inherited by the Thor CardStrategy (composes the trait).
The problem (measured 2026-07-13, 4/4 concurrent-boot attempts failed):
when all gears start at once, each vLLM engine's memory-profiling window sees
the other gears' weight loads — on Jetson unified memory, weight files read
through the page cache count as occupied — so the cortex fails either the
free-memory gate (Free memory on device … is less than desired GPU memory utilization) or KV allocation (No available memory for the cache blocks),
regardless of KV dtype or profile. The same values boot cleanly when the
primary starts alone: this is a bring-up race, not a steady-state
misconfiguration. Bring-up ordering is not expressible as a per-gear env knob
— tracked as follow-up work (plan risk r7).
Workaround (validated):
# after any teardown, reclaim page cache before rebooting the fleet:
sudo sh -c 'sync; echo 3 > /proc/sys/vm/drop_caches'
# bring the primary up first, then the rest:
docker compose up -d vllm-primary # wait until healthy
docker compose up -dWith that sequence the thor profile boots clean and all three correctness
probes pass. senses (the Gemma 12B) boots last and is the gear most
often caught by the race — its health was not confirmed in the 2026-07-13
clean-boot run; the steady-state values (32K / util 0.14) are the ones the
long-running hand-tuned deployment serves with.
- Issue #106 tracks re-verifying the Triton-ATTN knobs on GB10 (Spark) to confirm they do not cause regressions there. They are expected to work (Spark can run FlashInfer fine, so Triton is not needed, but forcing it anyway should not break correctness).
- Issue #109 tracks two GB10 verifications: whether uncalibrated-fp8 KV is
an accuracy risk there too (the Thor crash originally attributed to it did
not reproduce on the pinned nightly — fp8 now warns instead of asserting),
and whether the legacy
VLLM_ATTENTION_BACKENDenv is truly dead on the GB10's pinned image before the template drops it. - Issue #107 tracks broader work to tuned-small-model defaults for every card, not just UNKNOWN. Spark and Thor are mature; this is future work.
lobes explain tuning/lobes explain profiles— brief in-CLI referencelobes explain machine— machine profile flags and detectiondocs/tuning-profiles.md— workload (purpose) profiles and shahizat's benchmark baselinedocs/colleague-stack.md— the fleet contract (roles, context, memory budget)docs/deployment-shapes.md— the orthogonal deployment-shape axis (which roles a box hosts at all, composed over this per-machine tuning)lobes/machines/— CardStrategy modules (one per chip; the source of truth for detection signatures and machine knobs)lobes/profiles/— profile schema, loader, renderer, and built-in TOMLlobes/runtime/_detect.py— the detection fact-gatherer- Issue #105 — reranker
cudaErrorLaunchFailureon Thor (fixed by enforce_eager + TRITON_ATTN) - Issue #106 — GB10 re-verification of TRITON_ATTN knobs (pending)
- Issue #107 — broader tuned-small-model defaults for every card (future work)
- Issue #109 — GB10 verifications: uncalibrated-fp8 KV accuracy exposure, and
whether
VLLM_ATTENTION_BACKENDis dead on the pinned image