All notable changes to this project are documented here. The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.
- The
workerrole — an eighth first-class Colleague role servingunsloth/Qwen3.6-35B-A3B-NVFP4, a multimodal (image+video, no audio) ~3B-active MoE with a self-hosted MTP draft at 262K native.workeris the fast ground-work DOER — the first role besidescortexpermittedrepo_action(it may act on the repo under cortex's direction; forbidden onlyfinal_decision/security_decision), served multimodal (no--language-model-only) — a "seeing doer" distinct fromsenses, which perceives but must not act. Wired end-to-end: catalog gear + capability tier (minor < multimodal < worker < muse < main), the eight-role registry, the opt-in-core gateway backend (WORKER_BASE_URL; unwired ⇒ infeasible, never a silent fallback),OPT_IN_CORE_ROLES+base.tomlunknown-card veto, a profile-gatedvllm-workercompose service on the Qwen nightly lane (compressed-tensors,qwen3_coder+qwen3parser pair, self-draft MTP), thelobes up workerverb, thethor-workerdeployment shape, and full docs. VALIDATED live on the physical Jetson AGX Thor (sm_110), 2026-07-31 (docs/evidence/2026-07-31-accept-worker-thor.txt): boots and serves at measured util 0.45 / full 262144 window (KV pool 41.78 GiB = 14.07× concurrency, weights 24.81 GiB); MTP self-draft loads and accepts at 89.1%; decode 50.8 tok/s (with thinking) / 73.5 tok/s sustained, TTFT ~2.1 s; image AND video intake correct (image: red/blue with a negative control; video: a real Spark-webcam clip described accurately — subject/scene/motion, not a #101-style drop); thinking + tool + reasoning parsers work;model=workerroutes through the gateway andmodel=muse404srole_infeasiblewith no referral. Key finding — the vllm-worker lane must NOT force--moe-backendon sm_110: every forced NVFP4 backend was refused (flashinfer_* lack sm_110 kernels; marlin/triton reject the mixed quantized-main/unquantized-MTP experts), so the compose omits the flag and vLLM auto-selects per path. Spec/plan underdocs/specs/anddocs/plans/2026-07-31-thor-worker-lobe-qwen3-6-35b-a3b.md. - Muse goes dormant/unhosted mesh-wide: Thor moves off the Gemma 4 31B
museto hostworkerinstead.model=musenow 404srole_infeasiblewith nohosted_byreferral; themuserole, its catalog entry, and thethor-museshape all stay in-tree (cite-don't-delete), and the tier vocabulary still ranksworker < muse.
- CLAUDE.md / docs: the stale "pressure policy degrades … to
minor" wording now matches the shipped code — the degrade-to-minorsubstitution was removed; under pressure the full tiers (cortex/senses/worker/muse) are shed with HTTP 429 +Retry-After, andminoris the servable floor. - SonarCloud cleanup (folded in from #163): cleared the 5 genuinely-open
issues — S8997 ×4 (
test_chatterbox_pcm16.pyandtest_readiness_peer_probe.pynow use themonkeypatchfixture instead of manual save/restore) and S3776 (tts_client.py::_synthesize_singlerefactored, cognitive complexity 18 → 6, behavior-preserving). No suppressions were added — the false-positive candidates were already resolved.
lobes capabilitiesrenders a thirdloadedstate,by-proxy, for a role this box serves by forwarding to a peer — derived from the payload's existingfeasible+proxiedpair and shown alongside the hosting peer. Two roles in the same proxied state previously printed different values (yesvsno) purely because of leftover local<PREFIX>_BASE_URLwiring, while both forwarded happily and both answered 200.
docs/colleague-stack.mdnow states thatloadedis a LOCAL wiring fact that does not indicate whether a proxied role works, and records whysensesreportsloaded: trueon every fleet deployment:multimodalis the only optional generate backend whose compose default is non-empty (MULTIMODAL_BASE_URL=${MULTIMODAL_BASE_URL:-http://vllm-multimodal:8000}), and${VAR:-default}ignores an explicit empty value, so it cannot be unwired from.env.
- The stale comment in
lobes/cli/_commands/capabilities.pythat admitted an infeasible role's row "still shows loaded=yes" now describes the shipped three-state behaviour.
lobes.realtime.tts_client— the TTS retry sentinel is now a dedicated_Retrytype rather than a bareobject(), so_handle_tts_response()annotates asbytes | _Retryand_synthesize_single()narrows withisinstanceinstead of suppressing the mismatch with# type: ignore[return-value]. The runtime contract is unchanged — callers still only ever seebytes— but that invariant is now statically checked rather than asserted, so a future edit cannot leak a non-bytes sentinel past the retry check unnoticed (Qodo review, PR #157).
- tests/test_tts_pause_and_truncation.py — pause-mapping, truncation-detection and a ReDoS regression test for the S8786 rewrite (skipped offline; httpx is a realtime-extra dependency).
- Sonar S8786: replaced the quadratic
!{3,}$regex in trailing_pause_ms with a linear rstrip count — a 40k-character input took ~2.6s to scan and now takes microseconds. - Sonar S3776: extracted
_handle_tts_response/_is_truncated/_min_plausible_durationfrom_synthesize_single, dropping its cognitive complexity from 30 (limit 15). - Sonar S8513 (x2): collapsed chained endswith calls into single tuple-argument calls.
- Sonar S8572: the catch-all TTS retry arm now uses logging.exception so the traceback survives.
- Sonar S8410 (x3): FastAPI /v1/audio/transcriptions parameters now use Annotated type hints;
modelis correctly typedstr | None. - Sonar S5778 (x2): hoisted non-asserting setup calls out of pytest.raises blocks so the tests assert what they claim.
- Spec and plan for issue #155 — STT readiness truth: device-agnostic readiness, per-lane audio gating, a completed realtime composite, and cause-naming gateway refusals.
lobes.realtime._conversation.delivery_pause_ms-- milliseconds to wait before the next audio chunk, tracking playback rather than socket drain. Runs at mostDELIVERY_LEAD_MS(400 ms) ahead of the playhead: enough that a client buffer never runs dry, short enough that a barge-in onset lands inside a turn the server can still cancel. Returns 0 whenever delivery is already at or behind the playhead, so it only ever slows a run-ahead and never adds latency to audio the client is waiting on.
secureContextHint(Astro harness) derives a copy-pasteable remedy from the page's own origin -- the exactssh -Lcommand and URL -- instead of describing the secure-context rule and leaving the reader to apply it. A page already on localhost is told the origin is not the problem rather than sent chasing it.
- Realtime voice-out is now paced to the playhead instead of drained as fast as the socket accepts it.
_drive_responsepumped every delta in 2-4 ms for 7.5-8.5 s of audio (measured,docs/evidence/2026-07-22-accept-realtime-voice-to-voice-spark.txt) and then leftSPEAKING, so a user talking over a still-playing reply was talking while the floor had already returned toLISTENING: a new turn opened andresponse.interruptedwas never emitted. Every barge-in guarantee inlobes/realtime/_floor.pywas correct and completely inert. It also made session history dishonest -- the floor trims an interrupted reply to the prefix plausibly heard, but with instant delivery nothing is ever undelivered, so the machine recorded the whole reply as spoken. Barge-in itself remains UNVALIDATED live (#108); this is the mechanism the next acceptance run must exercise, not a claim that it passed. - The Astro harness's
no-devicemessage now names the PipeWire/PulseAudio output-only-profile trap, where a card shows a working device toarecordand nothing at all to the browser, and points atpactl list sources shortplus the duplex-profile fix. - The
unsupportedstate no longer blames the origin unconditionally (Qodo #154-1). It is reached viaisSupported(), which checks only API presence — so a browser missingAudioWorkletNodewas being sent to set up SSH forwarding that would change nothing.unsupportedMessage()now readsisSecureContextand asserts a cause only when one is actually known: insecure origin → the forwarding remedy, secure origin → a missing-Web-Audio-APIs message, unreadable → both named and neither asserted. The pre-existing test that asserted the conflated copy is corrected and now pins both branches explicitly instead of depending on an ambientisSecureContext. CLAUDE.mdno longer overclaims the #151 realtime work on two counts. It said "no live acceptance transcript for this work has landed underdocs/evidence/yet", which went stale in the very squash merge that added the 2026-07-22 Spark transcript; and it stated "Speaking during playback interrupts" as fact, which that transcript disproves. Both are corrected to match the evidence: the block is now PARTIALLY validated, naming what the run proved (voice-to-voice end-to-end) and what it did not (barge-in), and flagging that the transcript predates the 0.54.1 pacing fix so it is not evidence for barge-in either way.secureContextHint()no longer invents port 4321 whenlocation.portis empty (Qodo #154-2). A page served on a default port reported""and got a forwarding command for a service it was never served from; the real port is now derived from the protocol (80/443), and a privileged remote port maps to an unprivileged local one, sossh -Lnever emits a command that needs root on the pasting side.
- Voice-to-voice on /v1/realtime (#151): an opt-in conversation surface — a committed turn can trigger a server-side generate + Chatterbox reply streamed back over the SAME WebSocket as response.audio.delta events, interruptible by barge-in. Ears-only stays the default: a session that never sends response.create emits exactly the #149 transcription-only sequence.
- Barge-in with cancel-both semantics: speech during playback cancels the in-flight generate AND TTS, never sends the undelivered audio remainder, and emits response.interrupted — only the plausibly-heard prefix enters history. Arms the previously-dormant BARGE_IN_WINDOW_MS.
- New stdlib modules, all offline-tested:
_wire.py(base64 event codec),_floor.py(floor/turn state machine with per-stage deadlines),_turn.py(generate payload shaping),_conversation.py(the wiring layer).app.pystays apragma: no covershell. - Per-session conversation history + system prompt, with an operator default via DEFAULT_SYSTEM_PROMPT and a per-session connect-config override.
- site/ — a local-only Astro harness driving the realtime surface from a browser: mic capture with echoCancellation, live event stream with per-error-code distinction and VAD at_ms timing, audio-out playback, a user mute/mic-off control, and a conversation toggle. Reached through a local credential-injecting WebSocket proxy via ssh -L; never deployed. New site-build CI job.
- New env knobs threaded end-to-end (settings -> compose -> env.audio.example -> doctor --fix): BARGE_IN_WINDOW_MS, BARGE_IN_MODEL, TTS_VOICE_CONCURRENCY, DEFAULT_SYSTEM_PROMPT.
- Named error codes generate_failed, tts_failed, response_timeout and invalid_wire_event; boundary events now carry at_ms and reason, so VAD tuning is observable and a max_turn force-commit is finally distinguishable from a silence-confirmed stop.
- BREAKING (#151, coordinated): /v1/realtime audio is now OpenAI-shaped base64 JSON events in BOTH directions — input_audio_buffer.append inbound, response.audio.delta outbound. Raw binary audio frames are gone and now yield invalid_wire_event. The deployed reachy-mini-cli speaks the old wire and cannot stream until it adapts — tracked in reachy-mini-cli#115. In-repo clients (realtime-smoke.py, realtime-voice-loop.py) are migrated.
- Per-lane TTS concurrency: a spoken reply no longer queues behind unrelated batch /v1/audio/speech work. The batch lane's observable behaviour is unchanged (verified byte-for-byte).
- Muting is a narrowed ban, not a lifted one (deviation d1): AUTOMATIC mute-during-playback stays forbidden — it is the AEC substitute that makes barge-in impossible — while user-initiated mute/mic-off is allowed, because AEC is owned at the client edge (Reachy firmware, browser echoCancellation).
- The segmenter's at_ms and reason were computed and then dropped before reaching the wire (pre-existing since #149), so VAD boundary timing was unobservable by any client and a max_turn force-commit looked identical to a silence-confirmed stop.
- A stock Python-template .gitignore rule (/site, meant for mkdocs output) silently swallowed the new Astro harness: git add reported nothing, with no error.
- Three env knobs the settings module read but the deployment never passed — BARGE_IN_WINDOW_MS, BARGE_IN_MODEL and TTS_CONCURRENCY — silently pinned to their in-container defaults. A new AST-based coverage test fails CI if any future Settings field is added unwired.
- Six defects in the voice loop and its tests, all found by review (Qodo) and all accepted: the event reader swallowed EVERY exception from
read_frameand retried, so an EOF after a disconnect became an endless spin and the main loop waited out its idle timeout — a dead session reading as 'nobody spoke' (timeouts now continue, anything else ends the session and says why);speak()ran the audio player with no timeout, which is not hypothetical — paplay was OBSERVED hanging on a sink whose ALSA device another process held, and with the mic muted for the duration that hang deafens the session permanently (now a bounded 60s per backend, falling through to the next); a mid-session mic EOF broke the feeder loop without stopping the session, faking silence again;arecord's stderr was never piped, so the 'no audio' failure message could not actually quote the ALSA error it promised to; and two test weaknesses — a singlerecv()asserting an exact byte count where a stream socket may legitimately return less (now drains to the expected total, verified stable over 12 consecutive runs), and ajoin(timeout=...)that never asserted the thread finished, which is precisely where a deadlock should be reported.
scripts/realtime-voice-loop.py— talk to the machine, voice to voice, by composing the three endpoints the fleet already serves:/v1/realtimefor ears, a generate lane for the brain,/v1/audio/speechfor the mouth. It is also the richest live test of the realtime surface, because it holds a LONG-LIVED DUPLEX session — reading while writing, for minutes — which the one-shotrealtime-smoke.pycannot exercise. It defaults to the Gemma 4 12B lane (--model multimodal, ~1s to a short reply on the DGX Spark) rather than a thinking model, since in a spoken turn latency is dead air; takes the API key fromLOBES_API_KEYbecause argv is world-readable via/proc; supports stereo-only capture devices (--channels 2, downmixed); and routes playback with--sinkfor boxes where another process owns the audio device.
docs/realtime-pipeline.mddocuments the voice loop and three behaviours that a client of/v1/realtimemust get right, each learned on live hardware: answerPINGwithPONG(uvicorn pings every ~20s and closes a peer that never pongs — a duplex client that ignores them dies after tens of seconds while a short smoke run never notices); half-duplex turn-taking (the mic is muted for the whole synthesize-and-play window, because without AEC the session transcribes the machine talking to itself — real barge-in needs AEC, tracked in #151); and picking a FAST generate lane for spoken replies.
- The WebSocket client in
scripts/realtime-smoke.pybounded its reads withsettimeout, which is a property of the SOCKET rather than of one call — so a reader handed its deadline to any thread writing on the same socket, and a write that timed out part-sent left a torn frame that desynchronised the peer's parser. Reads now wait withselect, which observes readability without mutating socket state. Invisible to the one-shot smoke run (which streams, then reads); found by building a client that listens while it talks.
- The
/v1/realtimeroute is split into_open_session,_arm_segmenter,_to_pcm16k,_emit_turn_events,_transcribe_turn, and_pump_session(SonarCloud S3776: cognitive complexity 30 against a limit of 15). Behaviour is unchanged and, since the route carries no unit coverage by design, the refactor was gated on a fresh live run at both wire rates rather than on the offline suite alone.
- The
/v1/realtimehandshake forwarded the caller'sAuthorization(andCookie) to the realtime bridge. The credential is spent the moment the gateway's own inbound gate validates it and the bridge has no auth of its own, so forwarding it only widened a gateway key's blast radius to the bridge's logs and telemetry. Both headers are now dropped from the forwarded handshake (Qodo). - A dead realtime bridge could strand a gateway handler thread. When the upstream pump ended it half-closed the client (
SHUT_WR), which does NOT wake arecvblocked on that same socket — so an idle client left the thread parked until it happened to speak or hang up, leaking one thread per open session on every bridge restart. Each pump now shuts its peer down in both directions on exit. The unit test missed it because its fake socket returned EOF the moment its script ran dry; a fake that genuinely blocks is now the regression guard (Qodo). VAD_MAX_TURN_MSwas read unvalidated, so0or a negative value made the segmenter force-commit on the first chunk of every turn and keep committing — an event storm plus one STT forward per 32 ms chunk from a single typo in.env. Clamped to a 1000 ms floor, matching the existingtts_speed/tts_concurrencytreatment (Qodo).
/v1/realtimeis VALIDATED live on the DGX Spark GB10 (2026-07-21,docs/evidence/2026-07-21-accept-realtime-spark.txt): a full session ran through the gateway tunnel against the real Silero model and the real Parakeet/Chatterbox sidecars —session.createdthrough transcription on one connection, at 24000 Hz and the 16000 Hz passthrough, plus the 401 on an unauthenticated handshake and the 426 on a plain GET. The #149 motivating case (a five-word question that a client-side energy threshold shattered into "Ready, she") now arrives as one whole utterance with a single speech boundary pair. Four things stay UNVALIDATED and are documented as such: a real microphone (the live runs used synthesized audio), the VAD-unavailable path, concurrent sessions, and the max-turn force-commit.
- The gateway tunnel sent the bridge's FIRST FRAME back upstream instead of to the client.
read_headreturns any bytes the upstream packed into the same TCP segment as its 101, andrun_tunnelwrote them toupstream— sosession.creatednever reached the caller, and an unmasked server frame arrived at the bridge, which RFC 6455 §5.1 requires it to close on. Every session died the instant it opened. Caught by the first live run on a DGX Spark GB10, NOT by the unit suite, whose test had asserted the wrong direction as correct — the test is now inverted and joined by a regression test naming the failure.
- Pre-merge review of the realtime work caught three defects, all reproduced before fixing: the gateway's
/v1/realtimehandshake hung on every real connection (BufferedReader.read(n)waits to fill its buffer on a blocking socket, and a bridge that just accepted a session sends nothing more — nowread1(), guarded by real-socketpair tests instead of a fake that returned early); Silero inference and scipy resampling ran on the bridge's single asyncio event loop, so one talking session starved every other session and every batch/v1/audio/*request (both now offloaded viaanyio.to_thread.run_sync, matching the Chatterbox sidecar's convention, with the 16 kHz no-op passthrough left inline); andscripts/realtime-smoke.pydesynced its frame parser if a read timed out between a frame's header and its payload.
/v1/realtime— the server_vad WebSocket session the realtime bridge has promised since it shipped (issue #149). One connection replaces a WS-plus-batch-POST dance: stream PCM16 mono little-endian in (24000 Hz default, 16000 Hz accepted; the server resamples to 16 kHz itself) and receivesession.created/input_audio_buffer.speech_started/...speech_stopped/conversation.item.input_audio_transcription.completed/errorevents back on the SAME connection, with committed turns transcribed by Parakeet. This redeems two in-tree IOUs —app.py's own "PR2 adds the /v1/realtime WebSocket route" docstring andrealtime-pipeline.md's "planned for a later release" boundary claim — against the #149 baseline probe, where the deployed facade served four batch routes and no WebSocket, forcing reachy-mini-cli to endpoint from a client-side energy threshold that shattered sentences at inter-word dips.lobes/realtime/_segmenter.py— the server_vad turn segmenter as a pure state machine over 512-sample / 32 ms chunks, with Silero injected as a callable so the offline suite tests segmentation with a fake VAD (no torch, no GPU). A never-silent turn force-commits atVAD_MAX_TURN_MS(default 30 s) withreason="max_turn"— a normal boundary event, never an error — so a stuck stream cannot grow bridge memory without bound.lobes/realtime/_session.py— the session event schema, config parsing, teardown bookkeeping, and session-id-scoped logging (stdlib-only, so it is unit-tested without the[realtime]extra). A singleerrorevent type discriminated byErrorCodeis what makes VAD-down distinguishable from silence by event type alone; credential-shaped config fields are redacted before any log line.lobes/gateway/_realtime.py— WebSocket passthrough in the stdlib gateway: a 101-upgrade and bidirectional byte relay fronting the bridge, so the session is reached through the same origin and the same opt-inGATEWAY_API_KEYbearer gate as every other/v1/*route. The handshake is relayed verbatim rather than reimplemented, both pump directions unwind on either side's close, and the session legs drop the HTTP read timeout that would otherwise kill a listening session mid-silence.- The
sttrole advertisesrealtime_vad_sessiononlobes capabilitiesand gatewayGET /capabilities, so a client discovers the session surface instead of probing for it. Advertised only when the audio overlay is wired AND the lane is feasible — a text-only fleet or aSTT_FEASIBLE=falsedeployment withholds the claim. scripts/realtime-smoke.py— a live end-to-end session check (synthesize a known phrase, stream it, assert the boundary and transcript events arrive on one connection) written against a hand-rolled RFC 6455 client with no torch, no OpenAI SDK, and no WebSocket dependency, because keeping those out of a robot CLI's dependency tree is the point of server-side VAD.docs/evidence/README-realtime-acceptance.mddocuments the acceptance-evidence procedure.
/v1/realtimeis DECLARED/UNVALIDATED live (#108): the offline suite proves the session, VAD, and config logic with a scripted fake VAD, but nothing has run against real Silero, real Parakeet, or hardware yet — nodocs/evidence/transcript exists for issue #149. A plain GET to the route answers 426 Upgrade Required, and a declared-offsttlane answers 404role_infeasiblenaming its peer; a session is never proxied cross-box, since the #129 proxy-lobes forwarder is POST-only.
- The realtime container never received its own VAD knobs:
_settings.pyreadVAD_THRESHOLD,VAD_SILENCE_MS,VAD_PREFIX_PADDING_MS,DEFAULT_TURN_DETECTION, andDEFAULT_AEC_MODE, but neitherenv.audio.examplenor the composeenvironment:block passed any of them, so every value silently pinned to its code default and no operator could tune them. All six (plus the newVAD_MAX_TURN_MS) are now wired with compose defaults identical to the code's, anddoctor --fixheals a pre-existing deployment append-only, never rewriting an operator-customised line.
SupportedModel.default_gpu_mem_util— a per-model pooling budget override. The shared embed/score default (0.06) is sized for the ~0.6B gears and is SMALLER than the 4B embedder's own weights (7.56 GiB measured, vs a 0.06 x 121.69 = 7.30 GiB budget), solobes switch Qwen/Qwen3-Embedding-4Bpreviously wrote a budget the model could not load in. The 4B declares 0.11; every other model keeps the shared default.- Served-name collision guard in the gateway: two wired backends claiming one
served_namemakeresolve_model/order_backendsownership order-dependent, which on the embed lane means answering from the WRONG VECTOR SPACE. The gateway still starts (a name clash must not take the fleet down) but now warns loudly on stderr, naming the colliding backends.
- The
/recalland/rememberwrappers forcedEIDETIC_EMBED_URL=http://localhost:8002/v1— a port nothing listens on — so every semantic query silently ran on eidetic's 128-dim lexical-hash fallback while the docs claimed a live embedder. Both now default to the lobes gateway (http://localhost:8001/v1), matching eidetic >= 0.12's own default; verifiedonline=Trueat 1024 dim. The previous docs-only correction fixed the prose and left the scripts broken. - Sonar
python:S1192: extracted_CONTEXT_32K_NATIVEfor the"32K native"literal, now shared by three catalog entries.
embed-deep— an opt-in second embedding gear (Qwen/Qwen3-Embedding-4B, 2560-dim Matryoshka, MTEB multilingual 69.45 vs the 0.6B's ~64.3) beside the always-on hot-path embedder, addressed through the gateway asmodel=embed-deep. Opt-in on both axes — thevllm-embed-deepservice isCOMPOSE_PROFILES=embed-deep-gated and the gateway route is wired only whenEMBED_DEEP_BASE_URLis set — so every existing deployment renders byte-identically (no shape golden changed;vllm-embed's service hash is unchanged). Structurally it follows themultimodal-coderprecedent: an opt-in backend plus a wired-only alias whose name is its backend name.EMBED_DEEP_*env block (MODEL,SERVED_NAME,BASE_URL,MAX_MODEL_LEN,GPU_MEM_UTIL,ATTENTION_BACKEND) anddocs/qwen3-embedding-4b.md.- GB10 acceptance transcript (
docs/evidence/2026-07-20-accept-embed-deep-gb10.txt): booted live onspark-f8a9alongside the running spark-lobe fleet (zero fleet mutation) and serves 2560 dim — previously declared fromconfig.jsononly. Matryoshka honoured at all 6 probed ladder points; paraphrase probe 0.7362 vs 0.2818 unrelated; boots atgpu_mem_util=0.11(weights 7.56 GiB, KV 11.34 GiB / 82,592 tokens, CUDA graph pool 0.84 GiB); 42.4 ms median vs the 0.6B's 11.5 ms. - Pooling-lane attention-backend documentation (
docs/tuning-profiles.md,docs/machine-profiles.md,docs/qwen3-embedding-4b.md): theSM_110trait keys its knobs by PROFILE ROLE name, so it cannot reach theembed-deepGEAR — a Thor operator must setEMBED_DEEP_ATTENTION_BACKEND=TRITON_ATTNby hand or the forward pass hangs while/healthstays green (#105). This is the first place the gear-vs-role tradeoff costs something sharp.
- The gateway's opt-in alias wiring now covers both
multimodal-coderandembed-deep(same wired-only contract).embed-deepdeliberately gets no fallback to the 0.6B: the two embedders occupy different vector spaces, so a silent downgrade would return meaningless similarity instead of an honest unknown-model failure.tier_aliasesfalls back upward, which is right for generation and wrong for embeddings. Qwen/Qwen3-Embedding-4Bcatalog statusconfigured->load-tested(GB10 only; sm_110 remains UNVALIDATED per #108).EMBED_DEEP_GPU_MEM_UTIL=0.11is documented as MEASURED — with the caveat that vLLM's actual allocation (19.74 GiB) does not reconcile withutil x total(13.39 GiB) on this unified-memory card, so the knob is empirical, as it was for spark-lobe's 0.44.- The catalog's
test_exactly_one_embed_and_one_score_modelinvariant split: the score lane still pins exactly one model, while the embed lane now pins the property that actually matters — exactly one entry carriesrole_hint="embedding", so theembedderrole's reported model stays unambiguous under_catalog_by_role_hint's first-match lookup. - Corrected the
recall/rememberskill docs: the embed endpoint is the lobes gateway onlocalhost:8001/v1(not:8002), eidetic >= 0.12 sendsAuthorizationfromEIDETIC_EMBED_API_KEYor a borrowedCOLLEAGUE_API_KEY/CULTURE_VLLM_API_KEY, and records live in TWO stores — public in the COMMITTED<repo-root>/.eidetic/memory, private in$HOME— whichrecall.sh's own header already documented correctly whileSKILL.mdcontradicted it. Filed the silent hash-fallback bug upstream asagentculture/eidetic-cli#34.
lobes switchon an embed-task model always named thevllm-embedservice, so switching to the 4B told operators to replace the hot-path 0.6B in place — silently invalidating any index built with it, precisely the hazard this design exists to prevent._pooling_noticehardcoded one service for the whole embed task; it now resolves per model viarole_hint. Found by an independent colleague review.
toolson the role contract — every role inlobes capabilities --jsonand gatewayGET /capabilitiesnow reports whether its endpoint accepts OpenAItoolson a request. Derived from the catalog'stool_parser(the same field the served--tool-call-parserflag is built from, so it cannot drift from reality withouttests/test_catalog.py's pairing guard failing first):truefor cortex/senses/muse,falsefor the pooling roles (embedder/reranker serve no chat lane) and stt/tts. Previously NO role advertised tool support anywhere, so a Colleague could not discover it. Deliberately a bool, not a parser name — the served parser can diverge from the catalog's (PRIMARY_TOOL_CALL_PARSER, theqwen3_coder_thinkingplugin), so naming one would be a claimlobes.rolescannot honestly make.tool_useinmuse's declared responsibilities. Not a widening of its authority:final_decision/repo_action/security_decisionstay forbidden, so muse calls tools to RESEARCH a proposal, never to enact one —cortexremains the only lobe that acts.
- Gemma 4 lanes now serve the
gemma4parser PAIR (senses, the opt-in coder candidate, andmuse):--tool-call-parser=gemma4(Gemma4EngineToolParser, replacing the genericpythonic) plus--reasoning-parser=gemma4(Gemma4ParserReasoningAdapter, previously absent entirely). This mirrors the cortex lane's long-standing--reasoning-parser=qwen3+qwen3_coderpairing. See Fixed, below, for why each half is load-bearing. Operators on an existing scaffold must re-runlobes init(or edit the deployeddocker-compose.yml) to pick this up; a running container keeps its old flags until recreated.
- Gemma 4 tool calling was silently broken on every Gemma lane. Gemma 4 does
not emit Python-style calls — it emits
<|tool_call>call:name{...}<tool_call|>, whose delimiters are special tokens. Thepythonicparser is served withskip_special_tokens=True, so those delimiters were stripped before it ran; it then matched nothing and vLLM relayed the model's perfectly well-formed call as ordinary assistant content, withtool_calls: nullandfinish_reason: "stop". A caller passingtoolsgot prose shaped like a tool call and no callable one — no error, no warning.pythonicwas never evidence-backed:runtime/_parser.pycarried its own "risk r2, pending live validation" caveat from the start, and that check had never run. It ran on 2026-07-17 against the live 31B on a physical Jetson AGX Thor and disproved the guess. Validated on the 31Bmuselane only (docs/evidence/2026-07-17-accept-muse-tool-calling-thor.txt); the 12B lanes inherit the family rule and remain UNVALIDATED (#108) — a strictly better default than a parser proven wrong for the family, not a measured claim. - Gemma 4 channel markers leaked into
content. The correct tool parser forcesskip_special_tokens=False(that is how it sees<|tool_call>), which also exposes Gemma's<|channel>thoughtmarkers — which a tool parser has no business stripping. A plain answer came back as"<|channel>thought\n<channel|>The weather in Paris is...". vLLM ships the matching half (--reasoning-parser=gemma4) and lobes wired it on no Gemma lane; it is now paired on all three. Enable both or neither: the tool parser alone trades a broken tool call for dirty content. lobes capabilitiesno longer misreports an older gateway as unreachable. Its gateway sanity-check required every currentRoleInfofield, so a NEWER CLI probing an OLDER gateway (routine on a mixed-version mesh) judged a perfectly good response malformed, fell back to.envguesses, and printed "gateway unreachable" — false, since the gateway had answered, and the exact #92 dishonesty inverted. Fields added after the original #81 contract shape (tools,feasible) are now tolerated when absent; the stable core is already conclusive for that check's real job (telling a lobes gateway from a stray daemon on a guessed port)._render_tablealready.get-ed both with safe defaults, so an older payload renders without fabricating either.
- First-class stt/tts (#129): the audio roles joined the feasibility + peer referral/proxy channels —
STT_/TTS_FEASIBLE(absent = feasible, the sleeping-lobe default; every pre-#129 deployment renders byte-identically),STT_/TTS_PEER_ORIGIN/_PEER_PROXY/_PEER_API_KEYwith the same three-condition arming,hosted_by/proxiedcapabilities annotations via the one shared annotator, and a capabilities-based peer readiness probe (audio roles never appear on a peer/v1/models). - Per-endpoint audio routing (#129):
/v1/audio/speech(tts) and/v1/audio/transcriptions(stt) route independently — a declared-off lane proxies to its peer through the same data-plane machinery as core roles (callerAuthorizationstripped + pairwise key injected, body forwarded VERBATIM,X-Lobes-Proxiedsingle-hop guard with 508proxy_loop,X-Lobes-Proxied-Byattribution) or 404srole_infeasiblewith the honest referral;AUDIO_URLstays the local-bridge lane. The live trigger: Chatterbox on the Thor, Parakeet local on the Spark. DECLARED/UNVALIDATED until the live acceptance transcript lands (#108).
- Auth stragglers (#129 items 1-2): the stt/tts measure probes (
roles_measure.py) and the minor client (minor/_client.py, used bybenchmark --all-lobesandlobes route) now merge the same contextvar-scoped gateway auth header every assess-backed verb attaches;lobes routeresolves and installs the deployment key like its sibling verbs. With no key configured every request is byte-identical.
lobes doctorgains two fleet checks (#119):scaffold_files(every expected scaffold file on disk — the 2026-07-17 Spark partial-audio incident served for hours with/healthgreen and two Dockerfiles absent) andprofile_staleness(the deployed.envcarries the knobs the resolved machine profile requires — the 2026-07-14 Thor incident: a pre-#110.envmissing the SM_110 divergences hung its rerank lane silently). A key still carrying the template default where the profile requires a divergence is named; a genuine operator override only downgrades to info; shape-dropped roles are never demanded.lobes doctor --fix— the missing-only heal lane (#119):--fixprints the plan (still read-only),--fix --applywrites only ABSENT scaffold files and appends only ABSENT.envkeys, never rewriting an existing line (composeenv_filelast-duplicate-wins would let an appended default clobber an operator value). The safe path betweenlobes initrefusing (any file exists) and--forceclobbering the whole template set,.envincluded.
- Doctor remediation strings for stale/partial scaffolds name
doctor --fix, neverlobes init --apply --force(which would wipe the gateway key, peer/proxy config, and shape reclaim values).
accept-shape.shandvalidate-tiers.shnow consume the compose-fchain fromlobes fleet filesinstead of re-implementing it in bash (#138):--restoreno longer boots the very lobe a mesh-shape backup dropped (and no longer re-introduces the gatewaydepends_onedge the shape resets), and validate-tiers gateway recreates keep the shape overlay so results describe the topology the operator actually runs. The restore path fails loudly if the chain cannot be resolved; the best-effort_compose_downdegrades to a baredown --remove-orphans. A drift-proofing test pins both scripts to the CLI authority.
lobes fleet files— read-only verb printing the resolved docker compose-fchain (one argv token per line; empty for a plain deployment), so scripts consume the chain from the CLI instead of re-implementing it (#137)
- One compose
-fchain authority (#137):compose_file_argsinlobes/runtime/_compose.pybuilds every-flist;_compose_filesdelegates to it,up.pydelegates its targeted-services semantics as a parameter, andcompose_up_detachednow resolves the full chain — fixingswitchtearing down with the full chain but bringing up with none, andlobes serve --applybooting a shape-dropped lobe on a shape deployment - Deleted the dead
shape_render.py::shape_compose_filesandShapeRender.compose_files(never consumed by init; a latent duplicate chain builder)
lobes fleet up/downandlobes up <role>now honour an operator-authoreddocker-compose.override.yml.docker composeauto-discovers that file only when it resolves the project itself; ANY explicit-fsuppresses it._compose_filesreturns[]for a plain fleet (so the override applied), but scaffolding an unrelated overlay — the--audiooverlay or a deployment-shape override — switched it to an explicit-fchain that silently STOPPED applying the operator's file. Behaviour thus flipped on unrelated state, with no warning. Found live on the DGX Spark GB10, where the operator override publishes the Parakeet STT container on127.0.0.1:9002: alobes fleet up --applywould have recreated the container without that publish and brokenreachy-mini-cli's defaultREACHY_STT_URL. The override is now named explicitly in the-fchain, LAST — after even the shape overlay — because that is what an override file means to compose: last wins. Deployments without the file are byte-identical to before.
muse— the seventh first-class Colleague role: the creative/ideation lobe, servingnvidia/Gemma-4-31B-IT-NVFP4(Gemma 4 31B IT, NVIDIA's official modelopt NVFP4 export; 256K native; declares vision+audio configs (plain-gemma4 line, NOT the Unified family); MTP DECLARED viagoogle/gemma-4-31B-it-assistant, unmeasured). Addressable asmodel=muse; capability order is nowminor < multimodal < muse < primary. Responsibilities: creative_generation / long_form_writing / ideation / style_variation / divergent_second_opinion; forbidden: final_decision / repo_action / security_decision (muse proposes, cortex decides).- Opt-in core roles (
lobes.profiles.shapes.OPT_IN_CORE_ROLES): muse carries the full per-machine Profile knob set (MUSE_*prefix, schema is now five core roles) butmachine-as-brainnever hosts it — a 31B cannot co-reside with the cortex+senses duo on a 128 GB box. The machine-as-brain identity set is nowDEFAULT_HOSTED_ROLES(the six), and the machine-as-brain-equals-bare-card byte-identity invariant is preserved exactly (a non-hosted opt-in core role renders nothing). thor-musebuilt-in deployment shape (DECLARED — budget measured live; UNVALIDATED pending the acceptance transcript, #108 rule): hosts muse + embedder + reranker + audio; drops BOTH heavy default lobes (cortex and senses) to peer boxes. Carries the full muse declaration in its[overrides.muse](model,gpu_mem_util=0.55— measured live on the physical Thor 2026-07-17 (26.47 GiB KV pool, 611,415 tokens, 2.33x concurrency; the 0.40 hypothesis was refused with 0.6 GiB KV),max_model_len=262144— the full 256K native window,quantization=modelopt,attention_backend=TRITON_ATTN); hosting muse renders its activation env (COMPOSE_PROFILES=muse+MUSE_BASE_URL).base.tomlvetoes muse on unrecognised cards.vllm-musecompose service — profile-gated behind themuseDocker Compose profile (never started by a plaindocker compose up), same custom Gemma 4 image asvllm-multimodal(MUSE_IMAGEoverrides the tag).scripts/accept-shape.sh: drops reclaimable page caches beforefleet upwhen passwordless sudo is available (the documented Thor first-boot ritual, now automated in the acceptance flow), and the referral phase is proxy-aware — a dropped role with<PREFIX>_PEER_PROXYarmed is checked for a proxied answer carryingX-Lobes-Proxied-By: <origin>instead of the referral 404.- Gateway: muse joins all four peer/feasibility env channels
(
MUSE_FEASIBLE/MUSE_PEER_ORIGIN/MUSE_PEER_PROXY/MUSE_PEER_API_KEY) — referral and proxy-lobes work for muse exactly like every core role.lobes up muse,lobes measure(llm family), pressure shedding (muse degrades tominor), andlobes fleet status(container included when activated) all cover the new role.lobes up colleague-stackdeliberately stays the six default roles.
OPT_IN_BACKENDS(gateway): an unwired, unflagged muse backend is infeasible by default, somodel=museon every pre-muse/stale.env404srole_infeasible(honest, referable, proxyable) instead of silently upward-falling-back to cortex./capabilities(gateway + CLI) now reports SEVEN roles; docs and the in-CLI explain catalog updated throughout.tests/test_live_capabilities.py: generate-role discovery now covers muse and the third lobe state — a proxied dropped role must answer with theX-Lobes-Proxied-Bymarker naminghosted_by(a relayed non-2xx counts as an honest, marked relay; referral-only drops still demand the 404).
vllm-museboots behinddepends_on: service_healthyonvllm-embed/vllm-rerank— a concurrent cold boot at 31B scale crashed CUDA-graph capture (CUBLAS_STATUS_EXECUTION_FAILED) and every restart then failed vLLM's free-at-boot check against the dirtied page cache (measured live on the physical Thor, 2026-07-17).
docs/orin-profiles.md— Jetson AGX Orin 64GB live-validation evidence (2026-07-16/17, issue #127 mesh work): the operator profile serving Gemmasensesat its full 128K context on Ampere sm_87 (measuredgpu_mem_util0.45; 0.30 refused with 2.25 GiB KV vs the 3.08 GiB that 131072 needs; KV pool 802,644 tokens / 6.12x concurrency), embedder/reranker probe results, whycortex(modelopt NVFP4 W4A4) is architecturally infeasible on Ampere, three Jetson/sm_87 divergences found live (csv-mode GPU access needsruntime: nvidia; the Parakeet base image is Spark-only — no sm_87 kernels; unified-memory use far exceeds the util sum), the validated #127 cortex proxy wiring, and the gateway→gateway audio-chaining limitation (readiness probes/v1/health/ready, which a peer gateway 401s/404s — first-class audio referral knobs are the phase-2 candidate).
docs/machine-profiles.mdcross-links the Orin worked example from the custom-profile section.
docs/machine-profiles.mddocumentedlobes init --machinefor selecting a profile;init's actual flag is--profile(--machinebelongs toswitch). The wrong flag made the documented custom-profile flow error out.
- CI hardening: pin
astral-sh/setup-uv'sversionto0.11.29(waslatest) and turn onenable-cache: trueacross both workflows (tests.yml,publish.yml, 6 usages total), so a transient GitHub release-CDN outage can no longer take down every job by failing uv's "resolve latest" step; two identical CI failures on PR #132 (2026-07-16 22:39Z/22:43Z, GitHub 503 HTML from the setup-uv download) prompted the change. The action's SHA pin, tokens, and cache-dependency-glob are unchanged
- t5 live-verification evidence (colleague#320, spark-lobe go-live): docs/evidence/2026-07-14-strict-tools-spark-lobe-spark.txt — acceptance gates PASS at util 0.44/262144, lobes assess --strict-tools 3/3 legs PASS (strict+thinking was HTTP 500), colleague captured-bytes replay returns clean read_file with thinking intact via the armed gateway knob, MTP 100% draft acceptance under the constrained grammar, and the end-to-end colleague work repro delivers changed files in 4 clean steps (was 13 steps / 0 files). The deployment under test ran lobes-cli 0.44.0 (the release carrying the fix); 0.44.1 is docs-only — it records that verification, it is not itself what was verified
- Strict, grammar-constrained tool calls with thinking enabled (colleague#320): new lobes.vllm_plugins package ships the qwen3_coder_thinking vLLM tool-parser plugin (overrides get_structural_tag to derive reasoning from the request's effective enable_thinking, fixing the strict+thinking HTTP 500 caused by the served build's hardcoded reasoning=False); the fleet template mounts it into vllm-primary via --tool-parser-plugin and flips PRIMARY_TOOL_CALL_PARSER default to qwen3_coder_thinking; lobes init materialises the plugin file into the deployment dir; the gateway gains an opt-in GATEWAY_FORCE_STRICT_TOOLS knob injecting function.strict=true on cortex-lane tools requests with a retry-once-without-strict fallback on schema-compile failures
- spec + plan (devague /think + /spec-to-plan): docs/specs/2026-07-14-lobes-serves-strict-grammar-constrained-tool-calls.md and docs/plans/2026-07-14-… — converged frame with the proven root cause (server-side parser-salvage mangle of off-template emissions, deterministic at temp 0) and the user decisions (arm strict BOTH ways; retry-without-strict on schema-compile failure); live evidence recorded in the in-repo eidetic store
- Mesh-brain end-state implementation (#112, t2–t6):
orin-smallbuilt-in shape — the small-model reference shape for the Jetson AGX Orin 64GB (minor + pooling gears, BOTH heavies dropped), shipped as declared-but-UNVALIDATED data per the #108 rule, with theminoropt-in role added to the shapehostsvocabulary (OPT_IN_ROLES) rather than dishonestly reusing the cortex slot - Honest cross-box referral (#112 t3, the confirmed direct+referral decision): opt-in
PRIMARY/MULTIMODAL/EMBED/RERANK_PEER_ORIGINenv vars (operator-declared full origins, never derived — #92); with a peer declared,lobes capabilities, gatewayGET /capabilities, and the 404role_infeasiblebody name the hosting peer (hosted_by); annotation only — the gateway never forwards a request to a peer, and zero peer config renders byte-identical pre-referral responses (pinned byte-for-byte in tests) - Contract-test matrix (#112 t4): data-driven per-(built-in shape, dropped role) honesty tests — capabilities flag/omit, /v1/models omission, per-alias 404 with referral, no-outbound-connection tripwire — plus pinned t1 budget regressions (spark-lobe 262144 / thor-lobe 131072 cannot be silently lowered by a golden regen)
- Acceptance evidence (#112 t5):
scripts/accept-shape.shgains the orin-small arm and an opt-in referral phase; live Thor transcriptdocs/evidence/2026-07-14-accept-referral-thor.txt(referral 404s with a real declared Spark origin, cross-box cortex reachability, byte-for-byte shape restore) - Mesh-brain end-state docs (#112 t6): the four recorded decisions, the measured co-residency tax table (cortex 131072→262144, senses 32768→131072), and evidence citations in
docs/deployment-shapes.md,lobes explain shapes, and CLAUDE.md
- Deployment shapes (#113 implementation): Shape schema + three built-in shapes as TOML data (machine-as-brain, spark-lobe, thor-lobe) over the #110 Profile machinery; shape-aware budget re-derivation as declared overrides with provenance (measured live: spark-lobe cortex 0.44/262144, thor-lobe senses 0.30/131072); pure shape×card render composition with per-(shape,card) goldens; lobes init --shape behind the dry-run/--apply contract (bare init byte-identical, --single conflict); gateway dev lane: PIP_EXTRA_INDEX_URL build-arg passthrough so from-source boxes can deploy a TestPyPI .devN build without hand-edits
- Dropped-lobe honesty: a request for an unwired dropped role (e.g. model=senses on a spark-lobe box) now returns 404 role_infeasible on every alias instead of silently rerouting to the primary model (#92 invariant, caught by the t5 contract tests)
- Spec + plan for the mesh-brain end-state (#112, devague /think + /spec-to-plan): one heavy lobe per box as the far end of a backward-compatible, mixable shape axis — full-native budgets (spark-lobe cortex 262144, thor-lobe senses 131072), a declared-but-unvalidated orin-small shape, and cross-box direct + honest referral (opt-in peer config; no data-plane proxying); cheap gears co-reside everywhere
- Spec + plan: deployment shapes — machine-as-brain (default) vs per-box mesh-brain lobe profiles (spark-lobe drops Gemma senses, thor-lobe drops the Qwen cortex), flag-first selection on lobes init, per-box honesty; end-state tracked as #112 (devague frame + 8-task plan)
- Per-machine hardware profiles (spec + plan shipped in 0.40.2/0.40.3; this
is the implementation — 13 tasks, 4 waves):
lobes/machines/— per-chip strategy registry (oneCardStrategymodule per chip: spark, thor, blackwell, generic) with a sharedSM_110trait; the legacyMachineProfile/MACHINE_PROFILES/detect_machine()API is derived from the registry, every pre-existing test unchanged. Adding a chip = one file + one registration line.lobes/profiles/— profile schema (per-rolefeasible/model+ seven machine knobs), TOML built-ins (spark,thor,base), loader (operator profile in<deploy-dir>/profiles/<name>.tomloverrides built-ins), and the profile→env renderer. Thor's four divergent knobs stay single-sourced in the machines registry and overlay at load time.lobes/runtime/_detect.py— host card detection (device name + compute capability + total memory from/proc/meminfo; never nvidia-smi memory fields — they are[N/A]on Thor; UNKNOWN is first-class).lobes initdetects the card and applies the resolved profile (--profileoverrides with a warning); an UNKNOWN card warns and serves the conservativebaseprofile (4B generate model + the two 0.6B pooling gears, senses disabled, no 27B) instead of refusing (#107's unknown-card slice).lobes doctor/statusreport the detected card (device, compute capability, memory) and the chosen profile, warning on forced/unvalidated combinations; init persists the choice asLOBES_PROFILE.- Hardware feasibility honoured end-to-end:
<PREFIX>_FEASIBLE=falseremoves the role fromlobes capabilities/GET /capabilitiesand the gateway answers 404role_infeasibleinstead of silently rerouting (extends the #92 invariant). - Per-role correctness probes (
lobes assess --probes [--role r] [--timeout s]): cortex known-answer, embed paraphrase-beats-unrelated, rerank relevant-doc-first; timeout counts as FAIL (catches the sm_110 FLASH_ATTN hang that/healthmisses). - Golden rendered artifacts per shipped profile + the template-default
surface (
tests/goldens/, byte-diffed, GPU-less) — a change for one machine cannot silently alter another's rendering (the cross-machine no-breakage guard). - Upgrade-compat proof: a main-scaffolded deployment keeps working with
the new CLI (zero bytes changed by upgrade, env-name tripwire, re-init is
diffed and
--force --apply-gated). docs/machine-profiles.md+lobes explain profiles+ honest support tables in README/CLAUDE.md. Thor validated live 2026-07-13: 3/3 correctness probes pass on a clean boot (rerank correct and stable under TRITON_ATTN + eager); senses unconfirmed in that run. Orin / Orin Nano Super named but unvalidated.- Fleet compose knobs are env-parameterised (per-gear kv-cache dtype,
--attention-config, enforce-eager, models per role); defaults reproduce the shipped GB10 behavior byte-for-byte.MULTIMODAL_ATTENTION_BACKENDdeliberately kept pending the GB10 check (#109). - Live-validation findings, recorded as plan risk r7 and in the #109 thread:
the fp8
k_scaleassert did not reproduce on the pinned nightly — uncalibrated fp8-KV now boots with scale-1.0 warnings (an accuracy risk, not a crash) — and concurrent fleet first-boot on Thor fails a memory race (each engine's profiling window sees co-resident weight loads via page cache) regardless of profile; sequential bring-up (primary first) plusdrop_cachesafter teardown boots clean. Boot ordering is not expressible as a per-gear env knob — follow-up work.
GATEWAY_DEFAULT_MODELdefault is now empty = follow the primary gear's served name (was: hardcoded 27B id). Identical behavior on spark/thor; correct on thebaseprofile where the primary serves a 4B.
- The cortex correctness probe disables thinking per-request
(
chat_template_kwargs: {"enable_thinking": false}, thelobes routeidiom) — on the thinking-mode cortex the 16-token budget was consumed inside the<think>trace and a correct model failed the probe.
-
Plan (
/spec-to-plan, converged): per-machine hardware profiles —docs/plans/2026-07-13-lobes-fits-the-machine-it-lands-on-one-command-det.md. Thirteen tasks over four file-disjoint dependency waves (fifteen drafted, two rejected during convergence): the per-chip strategy pattern + profile schema + spark/thor profiles, card detection, template parameterisation and per-role correctness probes (wave 1);initapplies the profile and role-feasibility reaches capabilities/gateway (wave 2);doctor/statusreport the profile, and the Thor profile is validated on the physical board (wave 3); docs + an honest support table (wave 4).Per-chip knowledge goes behind a strategy pattern — one module per chip (
lobes/machines/<chip>.py) owning its own detection signature, per-role knobs and provenance, plus a small shared registry. Adding a chip is one new file and one registration line; it must not mean editing shared tables, and a change for one chip must not be able to break another. No existing code is deleted:MachineProfile,MACHINE_PROFILESanddetect_machine()keep working, rebuilt from the registry rather than duplicated.On Thor the reranker stays served and advertised — it runs, it is simply not yet correct; its ordering probe is recorded as a known failure pointing at #105 / #106, rather than the role being hidden.
-
Spec + plan rework (same frame, re-converged): per-machine knowledge is keyed by causal capability trait, not board name — of the four Thor hand-edits, three trace to
sm_110(the FLASH_ATTN pooling hang, the CUDA-graph classify fault) and one to the checkpoint's missing KV scales; none traces to "Thor the board". A machine profile becomes a named validated bundle (detection signature + memory budget + model-per-role + applicable traits), so an unrecognised board sharing a trait inherits its fix. Three new plan tasks: golden rendered compose/.env per shipped profile byte-diffed in CI, so a change for one machine cannot silently alter another's rendering (t13); an unknown card warns and serves a conservative small base — no 27B — instead of refusing, folding the unknown-card slice of #107 in (t14); and upgrading lobes-cli never breaks an existing scaffold — zero bytes changed in the deployment dir, old env names honoured, re-init always diffed +--apply(t15). Two GB10 verifications are parked as tracked risks: whether the fp8-KV crash is checkpoint-driven (shared fix) rather than Thor-specific, and whether theVLLM_ATTENTION_BACKENDenv is truly dead on the GB10's pinned image before t3 deletes it. Packaging per #107: profiles ship in the wheel, no pip extras.
- Spec correction: the first cut of this spec claimed lobes had no machine-profile
concept at all. It does —
lobes/profiles.pyalready shipsMachineProfile/MACHINE_PROFILES(spark/thor/blackwell/generic) anddetect_machine(), wired toVLLM_MACHINEand used byswitch/benchmark. The real gaps, now stated honestly in the spec's before-state: it is one knob-set per machine, not per role; it lacks the knobs that actually mattered on Thor (KV-cache dtype, enforce-eager, model-per-role, role feasibility); the fleet compose (the default path) ignores it entirely and hardcodes the Spark values; itsthorrow is an unvalidated guess (status="configured": flashinfer / 32768 / util 0.6) that live Thor testing contradicts; anddetect_machine()silently falls back togenericinstead of admitting it does not know the card.
-
Spec (
/think, converged): per-machine hardware profiles —docs/specs/2026-07-13-lobes-fits-the-machine-it-lands-on-one-command-det.md. lobes detects the host card and applies a profile tuned for that box (feasible roles + model per role + util / context / quantization / KV dtype / attention backend / enforce-eager), instead of the fleet compose's hardcoded GB10 values. Ships Spark (default) + Thor as supported; Orin / Orin Nano Super are named but unvalidated; an unrecognised card refuses-or-warns rather than silently applying the Spark profile. Every role gains a correctness probe, not just/health— a role that is healthy but semantically wrong must fail.Motivated by bringing the fleet up on a Jetson AGX Thor (sm_110): the Spark-tuned template scored 1 of 4 roles correct on first boot (senses clean; cortex crash-looped on an fp8-KV assert; embedder accepted requests and never answered; reranker killed its engine with
cudaErrorLaunchFailure), and reaching 3 of 4 took four hand-edits to the generated compose. Rerank is still wrong there — tracked in #105 / #106.
- Real-socket streaming regression test: a dribbling chunked-SSE upstream through the real
open_upstreamasserts the first relayed frame arrives while the upstream is still mid-stream — the fake-upstream tests (whoseread()already had read1 semantics) could never catch this bug class. - Spec + plan for the fix under
docs/specs/anddocs/plans/(devague framelobes-gateway-relays-sse-streams-frame-by-frame-th).
- Gateway SSE streaming:
_Upstream.readnow usesHTTPResponse.read1, sodata:frames relay to the client as the backend generates them. PreviouslyHTTPResponse.read(65536)blocked until 64 KiB or EOF, releasing a whole streamed turn in one terminal burst (frames=21 first=last=3.06s through the gateway) — full-turn latency before the first visible token for anystream: trueclient (issue #103; reported by colleague, agentculture/colleague#318).
- The
Hostrequest header is validated against a strict host-authority allowlist before it is echoed as an advertised origin inGET /capabilities. Previously the c29 origin-resolution change reflected an unsanitized, client-controlledHostheader into every role'sendpoint(e.g.http://evil.example/../<script>or auser@hostauthority), so a client scraping the contract could be handed an attacker's origin to dial. An invalid host is now treated as "no origin supplied" (empty endpoint), exactly as an absent header is. The trusted operator overrideGATEWAY_PUBLIC_URLis unaffected. (SonarCloudpythonsecurity:S5131.)
scripts/live-check.sh+tests/test_live_capabilities.py— a LOCAL, single-trigger, unattended pre-PR gate that dials every advertised role endpoint+path, every id in/v1/models, checks CLI/gateway agreement, reproduces Colleague's role-discovery path, and fails on deployed-gateway version skew. It FAILS rather than skips when armed. A 429 (pressure shed) and a 503 carryingRetry-Aftercount as reachable; only a 404, a connection failure, aRetry-After-less 503, or a bare 5xx are faults. that dials every advertised role endpoint+path, every id in/v1/models, checks CLI/gateway agreement, reproduces Colleague's role-discovery path, and fails on deployed-gateway version skew. It FAILS rather than skips when armed. A 429 (pressure shed) and a 503 carryingRetry-Aftercount as reachable; only a 404, a connection failure, aRetry-After-less 503, or a bare 5xx are faults.lobes/gateway/_readiness.py— a bounded background probe of each backend's/health, mirroringPressureCache. Tri-state (Truehealthy /Falsereached-but-unhealthy /Noneunreachable), daemon thread, socket-free.current(), and a probe that degrades toNoneonOSError,http.client.HTTPExceptionandValueError.GET /healthnow reports{"version": ...}, andlobes doctorgains agateway_version_matchcheck that fails on skew between the deployed gateway and the CLI wheel (issue #99).lobes/gateway/_routing.py::is_unknown_model— a pure predicate separating "unknown model id" from "unspecified model".- Ground-truth perception probes for the
sensesrole: a stdlib-generated solid-colour PNG whose colour the model must name, and a Chatterbox-synthesized word the model must transcribe. Both carry negative controls; the old placeholder-media tests are relabelled as wire checks.
- No cross-backend failover.
order_backendsnow returns at most one backend. A request naming the cortex model can never be answered by the Gemma backend, which protects thefinal_authorityrole contract from #81. GET /v1/modelsandGET /capabilities.readyare backed by the live readiness cache rather than by configuration. A wired-but-dead backend is no longer advertised.RoleInfo.readyis no longer an alias ofloadedforcortex/senses/embedder/reranker.build_role_registryself-enforces the invariant: a suppliedbackend_readymap is authoritative, and a presentNone, a presentFalse, and a missing key all mean not-ready.lobes capabilities/lobes endpointnow render the running gateway'sGET /capabilitieswhen it is reachable, falling back to an offline.env-derived view withready=falseon every role. The CLI and the gateway can no longer disagree, because there is now one derivation instead of two. The--jsonpayload is keyed strictly by role name in both modes (byte-identical to the gateway's payload in live mode); the live-vs-offline signal travels on stderr (asource:notice) rather than as an extra top-level key, so a strictset(keys) == ROLESconsumer never trips (Qodo).- A backend is wired only when its
*_BASE_URLis set; theor *_SERVED_NAMEclause is gone (issue #97). sensesis documented as vision-only intake. The checkpoint declares audio support but vLLM'sgemma4_unifiedpath does not serve it (issue #101);sttremains the supported speech path.docs/gemma4-mtp-draft.mdcarries a superseded banner: DSpark does not load on vLLM 0.23 (#75). Issue #69's disabled-experiment-entry criterion is closed answered-negative.- README's quickstart no longer claims
lobes initscaffolds the single-model deployment; the fleet duo has been the default since #69.
- #91 — a dead cortex backend no longer surfaces as a terminal
404 model does not exist.handle_postrewrote the model id once, before the failover loop, then retried the same body against a backend serving a different model, which correctly 404'd; the4xx = client error, no failoverrule relayed that as terminal and killed multi-step agent loops. A dead, unreachable or warming owner now yields 503 +Retry-Afterwithtype: backend_unavailable. - #92 / #95 — the gateway never advertises an origin built from its internal listen port. Precedence is
GATEWAY_PUBLIC_URL(an operator override for a tunnel or Host-rewriting proxy) > the requestHostheader > an empty endpoint.GATEWAY_PUBLIC_URLis deliberately NOT defaulted: a defaultedpublic_urloutranksHostand would advertise loopback to every remote client. - #96 —
AUDIO_URLnow reaches the gateway from the base fleet compose, sostt/ttsstop advertisingready=trueon a path that 404s when the audio overlay is not composed in. - #97 —
GET /v1/modelsno longer advertises phantom backends wired from a*_SERVED_NAMEalone against adefault_urlnaming a container that need not exist. ReadinessCacheseeds its initial per-backend snapshot withdict.fromkeysinstead of a dict comprehension (SonarCloudpython:S7519).ReadinessCache.stop()no longer clears its thread reference while the refresh thread is still alive, so a subsequentstart()cannot spawn a second concurrent refresh loop; the join bound now covers a full sequential refresh pass andstart()cleanly replaces a genuinely-dead thread (Qodo reliability).- An unknown model id returns
404 model_not_foundinstead of being silently served by the default backend under a different model's weights. Unknown-ness is decided against the routing table, never against the readiness-filtered/v1/modelslist, so a wired-but-dead backend still yields 503 rather than 404. test_live_main_text_returns_nonempty_contentno longer fails on a thinking model:max_tokens=16was consumed entirely by the reasoning trace, leavingcontent=Noneandfinish_reason=length(issue #93).
preserve_thinkingon the cortex/main vLLM lane — both compose templates now launch the primary/cortex service with--default-chat-template-kwargs '{"preserve_thinking": true}'next to--reasoning-parser=qwen3, so the served Qwen3.6 chat template retains all historical<think>blocks across a multi-turn conversation (default keeps only the reasoning after the last user turn). Default-on but per-request overridable; scoped to the cortex/main lane (embed/rerank/senses untouched,lobes routestill forcesenable_thinking=false). Closes #93.- Reasoning-aware
lobes.minorclient —assistant_turn_from_response()builds an assistant history message preserving thereasoning/reasoning_contenttrace, and a newhistory=parameter onchat_completion/chat_textround-trips it back on subsequent turns (single-turn behaviour unchanged). Part of #93. lobes assess --preserve-thinking— a read-only two-turn token-delta diagnostic that proves the reasoning round-trip is live (prompt-token count rises when the assistant<think>history is preserved vs content-only). Part of #93.
GET /capabilitiesnow advertises a client-reachable origin for every role'sendpoint— derived from the requestHostheader, overridable with the new optionalGATEWAY_PUBLIC_URLenv (for tunnels / Host-rewriting proxies) — so a consumer can dial any role'sendpointdirectly instead of a non-routable internal host (vllm-primary:8000,realtime:8080). This covers all six roles, includingstt/tts. Closes #87.- The realtime bridge exposes
GET /v1/health/ready, aggregating Chatterbox (TTS) + Parakeet (STT) readiness into one signal the gateway live-probes.
lobes statusis now fleet-aware: on a fleet deployment it reports per-gear container states and points atlobes fleet status/lobes capabilities, instead of the contradictory single-modelstate: model-gear-vllm — not createdline printed next tohealth: ok. A single-model deployment's output is byte-for-byte unchanged. Closes #84.GET /capabilitiesreportsstt/ttsreadyfrom a live probe of the audio backend (not merelyAUDIO_URLbeing set), so an advertised-ready audio role is genuinely consumable;lobes capabilities --jsonkeeps the configured signal since the host CLI can't reach the internal backends. Closes #89 (readiness).
- The gateway returns a clear 503 (
Retry-After) for/v1/audio/transcriptionsand/v1/audio/speechwhen the audio backend is reachable but still warming (Chatterbox/Parakeet loading, or a poisoned-CUDA context now surfaced honestly by Chatterbox's/v1/health/ready), distinct from the 502 for a genuinely unreachable backend — closing the "advertised ready but 502s / not client-consumable" gap. Closes #89. - The gateway's audio-readiness probe (
probe_audio_ready) now degrades a malformedAUDIO_URL(a non-numeric port makesurlsplit(...).portraiseValueError) or a broken HTTP exchange (http.client.HTTPException) to "readiness unknown" instead of letting the exception crash theGET /capabilities/POST /v1/audio/*handler — mirroringopen_upstream's guard. (Qodo review, #90.) build_role_registryclamps the stt/ttsreadysignal on the audio overlay being configured, so an unconfigured overlay can never reportready=Truewith an empty endpoint even if a caller passesaudio_ready=True— the public builder now enforces the "unconfigured ⇒ not ready" invariant its docstring already promised. (Qodo review, #90.)cmd_statusis split into_cmd_status_fleet/_cmd_status_singlerender helpers (mirroring the existing_cmd_status_pressure), leaving the verb as pure dispatch — dropping its Cognitive Complexity from 16 to under the gate's 15. Each helper uses a single-returnif/else(like_cmd_status_pressure) rather than an early return, so it renders its one exit path without trippingpython:S3516("always returns the same value"). Behavior-preserving; the fleet and single-model outputs are byte-for-byte unchanged. (SonarCloudpython:S3776+python:S3516, #90.)
- Under swap/iowait pressure the gateway now sheds a
main/cortexormultimodal/sensesrequest with HTTP 429 +Retry-After(busy — retry shortly) instead of silently degrading it onto a different model. The degrade-to-minor path is removed outright (noLOBES_PRESSURE_POLICYtoggle); an explicitminorrequest is still served as the floor. Busy is disclosed viaX-Lobes-Tier-Reason: busyand an OpenAI-shapedserver_busyerror body, distinct from the hard 502upstream_unavailable; callers (the acpvllm-localprovider, colleague, generic OpenAI SDKs) must retry with backoff.lobes status --pressureand the gatewayGET /statusnow report the busy-policy state (mode/shed/retry_after). The trigger reuses the existing swap/iowait signal; the tunable thresholds (#86) and the/procsampler are untouched.
- Pressure no longer crosses capabilities: on a default fleet with
minorunwired, a pressuredcortex/mainrequest previously fell through the tier upward-fallback onto the Gemma multimodal gear (a different capability), silently answering an authoritative-reasoning request with a perception model. Removing the degrade path makes that cross-capability substitution structurally impossible. Closes #85.
- Fleet gateway now receives the pressure-policy thresholds (
LOBES_SWAP_DEGRADED_THRESHOLD/LOBES_IOWAIT_DEGRADED_THRESHOLD) via its composeenvironment, so operators can tune the swap/iowait degrade triggers per box through.env. Previously those vars were only read from the gateway container's env, which the compose never populated, so the knobs silently stayed on the code defaults (75 / 50). Notably lets a box with an unreliable/proc/pressure/io(phantom high iowait on an idle disk, e.g. the DGX Spark GB10) raiseLOBES_IOWAIT_DEGRADED_THRESHOLDto 100 instead of permanently degrading the generate lane. Documented inenv.example+ regression test added.
- Gateway GET /capabilities (and
lobes capabilitieswhen read via the gateway) reported catalog-native context instead of the served--max-model-len, because the gateway container's environment lackedPRIMARY_/MULTIMODAL_/EMBED_/RERANK_MAX_MODEL_LEN. The fleet compose now passes those into the gateway service so the #81 served-context overlay resolves (cortex 131072, senses 32768); regression test added.
- cortex/senses role-based Colleague contract: six first-class roles (cortex, senses, embedder, reranker, stt, tts) with responsibilities/forbidden_responsibilities (#81)
lobes capabilities [--json]andlobes endpoint <role>— read-only role discovery over the role registry- Gateway GET /capabilities — machine-readable {role: {endpoint, model, context, ready, responsibilities, ...}} contract for Colleague
lobes up <role>andlobes up colleague-stack— role-based serving (dry-run by default, --apply to run; colleague-stack bundles the audio overlay)- lobes measure [--role] [--json] — per-role runtime-only metrics (TTFT/decode-tps/prefill, docs-per-sec, RTF)
- lobes benchmark --profile {cortex-only,cortex+senses,senses-direct,qwen-nvfp4-vs-bf16,all} — comparison profiles (runtime-only)
- lobes explain roles (aliases: colleague, colleague-stack, capabilities) and docs/colleague-stack.md
- Fleet context rebalance: cortex (Qwen 27B MTP) served at 128K, senses (Gemma 4 12B) at 32K (util 0.14, provisional); default fleet budget 0.30+0.14+0.06+0.06 = 0.56
- cortex->primary and senses->multimodal added to catalog.TIER_ROLE as the primary contract; main|multimodal|hard|normal|cheap|minor kept as back-compat aliases; brain forbidden
- lobes fleet status now reports the always-on Gemma (senses) container (FLEET_MULTIMODAL added to FLEET_CONTAINERS)
- Latent circular import between lobes.roles and lobes.gateway surfaced when lobes.roles was imported first
- catalog: coolthor/gemma-4-12B-it-NVFP4A16 (Gemma 4 12B NVFP4 base it-model) as the new default multimodal gear, native MTP wired ON by default ({"method": "mtp", "model": "google/gemma-4-12B-it-assistant", "num_speculative_tokens": 1}) -- measured 28.6 tok/s decode @ 57.9% draft acceptance, the fastest Gemma config on the DGX Spark (docs/vllm-nightly-migration.md §7)
- fleet compose: opt-in vllm-multimodal-coder service (profiles: [multimodal-coder]) so the demoted coder gear stays reachable; gateway wires an opt-in MULTIMODAL_CODER_BASE_URL/MULTIMODAL_CODER_SERVED_NAME backend + a multimodal-coder alias, added only once wired
- catalog: sakamakismile/gemma-4-12B-coder-fable5-composer2.5-MTP-NVFP4 demoted from role_hint=multimodal to role_hint=candidate (kept, cite-don't-delete) -- native MTP measured only 30.8% draft acceptance on the coder fine-tune, not worth wiring by default
- gateway: _DEFAULT_MULTIMODAL now points at coolthor/gemma-4-12B-it-NVFP4A16; the multimodal/normal tier aliases resolve to it
- docs/gemma-4-12b-nvfp4.md, README.md, docs/gateway-fleet.md, docs/qwen3-14b-nvfp4.md updated to describe both Gemma gears (default base + opt-in coder) and cite docs/vllm-nightly-migration.md §7 for the benchmark evidence
- docs/gemma-4-12b-nvfp4.md — first throughput/prefill benchmark for the Gemma 4 12B multimodal gear on the DGX Spark GB10: ~23 tok/s single-stream decode (23.0 sustained over 1,500 tok), prefill ~2,650 tok/s (847 tok) / ~1,954 tok/s (6,682 tok), on vLLM 0.23.1rc1.dev672 native gemma4_unified
- README acknowledgement of Mieszko Syty (FutureProofHomes; Jetson AI Lab) alongside shahizat
- docs/gemma-4-12b-nvfp4.md notes a config-drift follow-up: on the current :nightly-audio (dev672) image the default lane util 0.12 + VLLM_MEMORY_PROFILER_ESTIMATE_CUDAGRAPHS=0 no longer boots 8192 co-resident (cudagraph accounting changed) — benchmarked at util 0.15 with a trimmed cudagraph capture set
- docs/gemma-4-12b-nvfp4.md — made the #75 speculative-decoding section internally consistent: it now reads as CLOSED (route resolved, wire/measure/verdict not implemented) throughout, matching the "Resolved" bullet, instead of framing #75 as active work (Qodo); also corrected the stale claim that the 12B lane decodes slower than the primary — the benchmark shows it out-decodes the primary single-stream (~23 vs ~18–19 tok/s)
- Gemma 4 12B multimodal gear now SERVES (text + image + audio) — live-validated on the DGX Spark GB10 via vLLM nightly's native gemma4_unified class (#71/#73); catalog status promoted configured → load-tested
- Dockerfile.vllm-gemma4 rebased FROM vllm/vllm-openai nightly (pinned by digest) + the vllm[audio] extra (librosa==0.11.0 soundfile==0.14.0 av==17.1.0 soxr==1.1.0, pinned to the live-validated set) (was NGC 26.06 / vLLM 0.22.1 + a transformers overlay); vllm-multimodal compose/env now set VLLM_MEMORY_PROFILER_ESTIMATE_CUDAGRAPHS=0 and default MULTIMODAL_MAX_MODEL_LEN 8192 (co-resident util-0.12 KV holds ~24K tokens, not the 128K native)
- Gemma 4 12B serve blocker root-caused and fixed: gemma4_unified has heterogeneous per-layer head sizes (40 sliding@256 + 8 full@512) that only vLLM's native class handles — released vLLM ≤0.22.1 fell back to the transformers backend and crashed the full-attention o_proj (4096≠8192); a TRITON_ATTN backend flag does not fix it
- Issue #75 spec + plan (Gemma 4 12B gear speculative decoding) converged via devague /think + /spec-to-plan
- docs/gemma4-mtp-draft.md — resolved spec-decode draft route (DSpark draft_model first; native google/gemma-4-12B-it-assistant recorded as escalation candidate)
- docs/gemma-4-12b-nvfp4.md — added the Speculative decoding (#75) before-state / gap / scope-split subsection
- Custom vLLM image Dockerfile.vllm-gemma4 (FROM nvcr.io/nvidia/vllm:26.06-py3, vLLM 0.22.1, + a pinned from-source Transformers 181beb3) so the Gemma 4 12B gemma4_unified multimodal gear loads (#71)
- MULTIMODAL_IMAGE override for the vllm-multimodal service (local build by default; optional ghcr.io/local-registry tag)
- MULTIMODAL_ATTENTION_BACKEND env (TRITON_ATTN) for Gemma 4 non-square attention
- Multimodal gear quantization corrected to compressed-tensors (was modelopt_fp4) after live validation (#71)
- Removed the invalid gemma4_mtp speculative_config from the multimodal gear (vLLM Gemma4 MTP needs a separate gemma4_assistant draft model)
- docs/compose comments now name the correct base image and MULTIMODAL_IMAGE registry semantics
- docs/specs: Gemma 4 12B multimodal-duo spec (issue #69) — /think frame for default-serving the Qwen3.6-27B-MTP + Gemma4-12B duo as main/minor/multimodal tiers (vision+audio), native-MTP on by default, DSpark draft as a disabled experiment, 14B demoted to a legacy candidate
- docs/plans: buildable plan for the Gemma duo (9 tasks across 5 dependency waves, 6 accepted-risk objects) via /spec-to-plan — covers all 26 spec targets; resolves the main/minor/multimodal pressure-ladder seam as a first-class task
- Reworded the three Gemma-risk markers in
catalog.py/runtime/_parser.pyfrom bareTODO(risk …)comments toRisk … (pending #71)— the deferred live-validation work is tracked in issue #71 (gemma4_unified won't load on released vLLM images), so the comments now cite the tracking issue instead of an untracked TODO (clears SonarCloudpython:S1135).
- Gateway: re-wire the legacy 14B
middlegenerate backend fromMIDDLE_BASE_URL/MIDDLE_SERVED_NAMEinbuild_config(). The #69 14B demotion dropped the wiring but the compose template still ships thevllm-middleprofile + those env vars, so enabling the profile silently fell back to the primary; the 14B is again reachable by its explicit served name (and, as intended, gets no tier alias). (Qodo) - Gateway: a
GATEWAY_ALIASESoperator override keyed by a legacy tier alias (hard/cheap/normal) is now honoured on the pressure-aware tier path. Tier requests normalize to the new vocabulary (hard→main) before the alias lookup, which bypassed a legacy-keyed override;build_config()now mirrors a tier-keyed override onto its vocabulary synonyms (explicit keys still win). (Qodo)
- Third capability tier: opt-in
vllm-middle14B-NVFP4 generate gear (COMPOSE_PROFILES=middle, GPU mem-util 0.12), inference-only (not a LoRA base). - Gateway capability-tier aliases — callers send
model=cheap|normal|hardand the gateway resolves to the 4B/14B/27B generate gears (same-task alias on top of task-family routing) with upward fallback when a tier is absent. - Read-only host memory-pressure sampler (
swap_used_percent/iowait_percentfrom /proc) and a swap/iowait pressure policy with a degraded-mode state machine (env-overridable thresholds). - Pressure-aware tier downgrade at the gateway with an
X-Lobes-Overridebypass header; the served tier and reason cross the OpenAI boundary viaX-Lobes-Tier/X-Lobes-Tier-Reasonresponse headers. lobes status --pressure— read-only snapshot of the current tier ceiling, mode, reason, and live swap/iowait.scripts/validate-tiers.sh+docs/validate-tiers.md— operator-run live validation harness for the three-tier fleet on the Spark.
- 27B primary default served context trimmed 256K→128K (
PRIMARY_MAX_MODEL_LEN=131072) andPRIMARY_GPU_MEM_UTILlowered to 0.45 so the co-resident 14B middle gear fits within the 128GB unified-memory budget (0.45 + 0.12 + 0.10 + 0.06 + 0.06 = 0.79).
- Mutation-safety prose in
lobes learnand CLAUDE.md now lists thefleet up/fleet downwrite verbs (was only in the--jsonpayload). - CLAUDE.md documents the turn-on/turn-off lifecycle explicitly (
serve/stopandfleet up/down) instead of leaving it implicit in the verb names.
lobes benchmark --all-lobes --concurrency auto: per-lobe (minor + primary) performance benchmark routed through the gateway — single-stream decode tok/s, prefill TTFT, concurrent throughput with auto-ramp to the throughput knee (req/s + p50/p95 latency + ms/token), plus the logprobs cat soft-score, rendered as one combined minor-vs-primary report with per-metric deltas.lobes eval cat --score logprobs --mode open|closed: read-only 'Where is the cat?' temporal-reasoning probe, scored by logprobs (softmax over candidate-location full-sequence echo logprobs as the headline, with a chat first-token-mass cross-check and graceful fallback when echo is unavailable).lobes.benchpackage:cat_probe(deterministic, seeded timestamped-narrative generator with exactly one unambiguous current location; open + closed modes),cat_score(echo-softmax headline scorer + first-token cross-check + fallback), andreport(per-lobe markdown report renderer with minor-vs-primary deltas).lobes.minorlogprobs plumbing:chat_completionnow forwardslogprobs/top_logprobs; newcompletions_echo(full-sequence/v1/completionsecho scoring) andgateway_supports_echocapability probe (never raises; lets callers fall back).lobes.assessper-lobe perf engine:measure_prefill_ttft,run_concurrent(requests/sec + p50/p95 latency + ms/token), andauto_ramp_concurrency(1→2→4→… ramp with plateau/knee detection).
- The
minorlobe — a cheap, warm co-resident Qwen3.5-4B small-brain (issue #64). A new switchable catalog gearQwen/Qwen3.5-4B(role_hint="minor", served bf16 — the first unsloth-LoRA fine-tune target; multimodal, served text-only via--language-model-only), reachable both as a switchable gear and as an opt-in warm co-resident backend alongside the 27B primary.- New read-only verbs:
lobes run minor "<prompt>"(call the minor model),lobes route "<text>"(classify a task across catalog gears with an escalate flag + confidence), andlobes eval minor --suite <path>(run a JSONL eval suite). All three default--base-urlto the gateway (http://localhost:8000/v1) and reuse a new stdlib-only urllib client (lobes.minor) — no new runtime dependencies. - Governance + escalation (
lobes.minor.governance): the minor role may prepare/classify/format/validate/suggest/summarize/route, and escalates on forbidden actions (approve/finalize/delete/deploy/architectural) or any of five escalation conditions. Role-keyed, not model-keyed. - Warm co-residency: an opt-in
vllm-minorfleet service (compose profileminor+MINOR_BASE_URL/MINOR_SERVED_NAMEgateway env gate); the gateway routes the minor model id to it with failover to the primary. Default fleet behavior is unchanged.
- New read-only verbs:
runtime/_parser.pyrecognizes the Qwen3.5 family →qwen3_coder(it emits the XML function-call format, not Hermes JSON).- Catalog supports an unquantized bf16 generate gear via a
quantization="none"sentinel thatlobes switchnormalizes to "omit--quantization" and surfaces as a required compose edit.
- Memory-discipline "Conventions and workflow" section in
CLAUDE.md— a per-task recall-before / remember-after convention (scope localized to this repo's nick) so the vendoredremember/recallskills are actually used, not just present:/recallbefore non-trivial work to build on prior decisions instead of re-deriving them, and/rememberwhen a non-obvious decision, constraint, fix-and-why, or hard-won gotcha surfaces. The section documents this repo's memory as in-repo and public — records resolve to<repo-root>/.eidetic/memory(committed, team- and mesh-shared). Inserted idempotently (skipped if already present), slotted under an existing "Conventions and workflow" heading when one exists, else appended.
- Refreshed the
remember+recallwrappers from eidetic-cli 0.10.0 (cite-don't-import) — picks up eidetic's project-local store default: the files backend now resolves per record by visibility — PUBLIC records inside a git repo go to<repo-root>/.eidetic/memory(committed, team-shared), PRIVATE records (or any record outside a repo) go to$HOME/.eidetic/memory(never committed), an explicitEIDETIC_DATA_DIRstill wins, and recall reads both stores and merges. Also carries the 0.9.3 hardening (interactive-stdin guard,helpas a search term, SIGPIPE-safe suffix parsing). Recipe policy override (the wrappers here are NOT byte-verbatim): the injected default visibility is flipped from eidetic'sprivatetopublic, so a plain/rememberlands the note in./.eidetic/memoryin this repo, kept as part of the repo — pass--visibility privateto route a record to$HOMEinstead.rememberdriveseidetic remember(idempotent upsert of one JSON record or an NDJSON batch on stdin);recalldriveseidetic recallwith four search modes (exact / approximate / keyword / hybrid). EachSKILL.mdis localized only in the illustrative--scope <nick>examples (Provenance keeps "First-party to eidetic-cli"). Runtime dep: theeideticCLI on PATH (else a local eidetic-cli checkout withuv) —eidetic >= 0.10.0for the in-repo routing; on an older CLI the public records still work but are stored in$HOME/.eidetic/memoryinstead of in-repo. Propagated by rollout-cli'seidetic-memoryrecipe.
docs/tensorrt-llm-investigation.md— a dated desk investigation (no live run) of serving the MTP 27B primary (sakamakismile/Qwen3.6-27B-Text-NVFP4-MTP) with TensorRT-LLM (trtllm-serve) on the DGX Spark (GB10/SM121) instead of vLLM. Verdict: not yet — TRT-LLM MTP spec-decode is DeepSeek-only in stable releases and the Qwen3.6 hybrid GDN/DeltaNet kernels are RC-only (both land in 1.3.0 RC builds); serving on a stable TRT-LLM today would forfeit the ~2.4× decode win the checkpoint exists for. Records the engine-integration seam (the request path — gateway routing +lobes assess/benchmark— is already engine-agnostic, while the gateway/statusvllm:*metrics path,catalog.py,switch.py, templates, andVLLM_*env vars are vLLM-specific), a feasibility table by dimension with confidence levels, a comparison against the recorded vLLM baseline, a minimal spike recipe, an explicit revisit trigger (TRT-LLM 1.3.0 stable), and 11 cited sources. Linked from the README per-model notes.
- Vendored the
remember+recallmemory skills from eidetic-cli (cite-don't-import) — the write/read halves of eidetic's shared~/.eidetic/memorysurface, so this agent (Claude and its colleague backend) can persist facts across sessions and recall them later, sharing one store.rememberdriveseidetic remember(idempotent upsert of one JSON record or an NDJSON batch on stdin, dedup by id + content hash);recalldriveseidetic recallwith four search modes — exact / approximate / keyword / hybrid — each hit carrying text, full provenance metadata, a relevance score, and a freshness signal. The.shwrappers are byte-verbatim from eidetic-cli (their first-party origin); eachSKILL.mdis localized only in the illustrative--scope <nick>examples (Provenance keeps "First-party to eidetic-cli"). Both default to this agent's PRIVATE scope, reading the suffix fromculture.yaml. Runtime dep: theeideticCLI on PATH (else a local eidetic-cli checkout withuv). Propagated by rollout-cli'seidetic-memoryrecipe.
- Renamed the tool from
model-gear/modeltolobes/lobes-cli. The binary is nowlobes(lobes switch,lobes serve,lobes assess, …), the import package islobes, and the PyPI distribution islobes-cli. The deployed Culture agent is renamedmodel-gear→lobes(culture.yaml,AGENTS.md). - Deployment dir is now
~/.lobes(env$LOBES_DIR). The legacy$MODEL_GEAR_DIRand~/.model-gearare still resolved as fallbacks, so a pre-rename deployment keeps working with the renamed CLI without redeploying. - Deployment-internal names are intentionally kept as
model-gearso a live fleet isn't disrupted: Dockercontainer_names (model-gear-vllm,model-gear-gateway, …),mg-logwrap.sh,MODEL_GEAR_LOG_DIR, and the served-model id are unchanged.
modelis kept as a deprecated alias command forlobes(same entry point);--version/help reflect whichever name was invoked.model-gearis published on PyPI as a deprecated alias oflobes-cli: a metadata-only shim package (packaging/model-gear/) that depends onlobes-cli==<same version>, plus apublish-aliasjob inpublish.ymlthat builds and publishes it after the main release.
- Relicensed the project from MIT to Apache 2.0 — full Apache 2.0 LICENSE text, pyproject
license/classifier metadata, and a new README License section. Aligns with sibling AgentCulture repos (e.g. colleague, data-refinery-cli).
- Doc consistency:
docs/mistral-small-3.2-24b-nvfp4.mdanddocs/qwen3.6-35b-a3b-nvfp4.mdstill framed Mistral as the default fleet fallback the gateway pairs with the primary. The fleet has run one generate backend by default since the single-backend default (#42); the warm fallback is opt-in. Reframed both docs to match — closing the drift with themodel explaincatalog corrected in 0.26.2. - Corrected the Mistral doc's "How it runs in the fleet" section, which described
a fallback wiring that no longer exists:
FALLBACK_MODEL/FALLBACK_MAX_MODEL_LEN/FALLBACK_GPU_MEM_UTIL/….envkeys "scaffolded bymodel init --fleet" and a shippedmodel-gear-vllm-fallbackservice. The current templates ship no fallback service and the gateway reads onlyFALLBACK_URL+FALLBACK_SERVED_NAME(set after you manually add avllm-fallbackservice). Following the old text produced a non-working config; rewrote it to the actual two-step opt-in, matchingdocs/gateway-fleet.md→ "Adding a fallback".
- Doc alignment pass across the audio + fleet surfaces (no behavior change):
docs/chatterbox-tts.md: healthcheck showspython3.12(not the stalepython3) and the compose snippet includescontainer_name: model-gear-chatterbox— matching the 0.26.1 fixes.docs/realtime-pipeline.md: the TTS service ischatterbox(wastts); the overlay scaffolds a Chatterbox Dockerfile too.docs/openai-api.md: corrected the auth caveat — the gateway is a pass-through and is not auth-aware for any proxied endpoint (the previous wording implied it gated/v1/chat/completions).docs/gateway-fleet.md: endpoint list now includes/v1/audio/*; added an "Auth (known limitation)" note.model explain gateway/model explain tunnel(explain/catalog.py): added the gateway-not-auth-aware caveat; fixed the Mistral entry (opt-in fallback candidate, not the active default pairing); listedtunnel/fleetas write verbs.model learn(learn.py): added an "Auth / exposure" section + anauth_exposureJSON field, andmodel explain tunnel/gatewaypointers.
model init --fleet --audionow scaffoldsDockerfile.chatterbox. The Chatterbox sidecar landed in 0.25 (the composechatterboxservice builds fromDockerfile.chatterbox), but the build file was never added toAUDIO_TEMPLATES, so the scaffold omitted it anddocker compose build chatterboxfailed with "Dockerfile.chatterbox: no such file". Added it to the audio template set (twin of theDockerfile.realtime/Dockerfile.parakeetwiring) so the audio overlay can actually build and serve TTS.model fleet statusnow reports the TTS gear.FLEET_TTSstill pointed at the oldmodel-gear-ttscontainer name, but the Chatterbox sidecar renamed the container tomodel-gear-chatterbox— so status listed the live TTS gear as "not created". PinnedFLEET_TTStomodel-gear-chatterboxand added a test that asserts everyFLEET_AUDIO_CONTAINERSname matches acontainer_name:in the packaged audio compose (catches future rename drift).- Chatterbox container now reports healthy. Its
Dockerfile.chatterboxinstalls the interpreter aspython3.12(nopython3symlink), but the compose healthcheck called barepython3— which exec-failed every interval, pinning the working container at "starting"/"unhealthy". Switched the healthcheck topython3.12and added a test tying the healthcheck interpreter to the one the Dockerfile provides.
Documentation pass for the realtime audio overlay and the OpenAI API front: a
feature doc per audio backend, a consolidated endpoint reference, and the same
information surfaced through model learn, model explain, the README, and
CLAUDE.md.
docs/parakeet-stt.md: per-model feature doc for the Parakeet STT backend (nvidia/parakeet-tdt-0.6b-v2, NeMo ASR) — the only audio model that lacked one. Covers the HTTP contract, the real (model-loaded + CUDA-live) readiness probe, fleet integration, the stale-CUDA-context runbook, and why it is not a switchable catalog gear.docs/openai-api.md: consolidated OpenAI-compatible API surface reference — every endpoint (/v1/chat/completions,/v1/completions,/v1/embeddings,/v1/rerank,/v1/score,/v1/audio/transcriptions,/v1/audio/speech,/v1/models,/v1/models/supported,/health), routing semantics (name / default / failover / SSE / audio fan-out), per-endpointcurlexamples, the loaded-vs-supported split, and auth/exposure.model explaintopics:realtime/audio(the/v1/audio/*overlay),transcribe/stt/parakeet(STT),speak/tts/chatterbox(TTS), andapi/openai(the endpoint surface); linked from the explain root.model learnnow documents the realtime audio overlay and the OpenAI API surface (text +--jsonrealtime_audio/api_surfacefields).- README sections for Realtime audio (STT + TTS) and The OpenAI-compatible API surface, plus the two audio backends added to the per-model notes.
CLAUDE.md: documents the realtime audio overlay alongside the fleet; the CLI package tree now lists thegateway/,realtime/,explain/, andcatalog.pysurfaces.
docs/realtime-pipeline.md: removed the staleNGC_API_KEYbring-up step (a Magpie leftover — Chatterbox needs no NGC key) and documented the TTS → STT round-trip inscripts/audio-smoke.py.
- Chatterbox TTS sidecar (
model_gear/realtime/chatterbox_server.py): a FastAPI HTTP server (GET /v1/health/ready,POST /v1/audio/synthesize) that wraps Resemble AI's Chatterbox model and returns raw PCM16 mono 24 kHz audio. Supports zero-shot voice cloning via a.wavreference path. Runs as thechatterboxfleet service built by the newDockerfile.chatterbox(arm64 cu128 recipe). [chatterbox]optional-deps group (fastapi,uvicorn) inpyproject.toml.docs/chatterbox-tts.md: bake-off numbers, arm64 install recipe, sidecar HTTP contract, and integration notes.model_gear/templates/fleet/Dockerfile.chatterbox: arm64 CUDA build — rebased onnvidia/cuda:12.8.0-cudnn-runtime-ubuntu24.04(no preinstalled torch) with pinnedtorch==2.11.0+cu128+torchaudio==2.11.0+cu128+ Perth, fixing the NGC pytorch ABI conflict with Perth observed at runtime.- numpy fast path in
float_tensor_to_pcm16(stdlib fallback kept for offline CI); parity test intests/test_chatterbox_pcm16.py. resolve_voice().wavcheck is now case-insensitive (.WAV/.Wavwork).- Dead SSML code removed from
tts_client.py(_insert_ssml_breaks); stale Magpie references updated to Chatterbox across app, client, and tests; speed-ignored warning emitted when a non-default speed is passed tosynthesize().
- Replaced Magpie TTS with Chatterbox across the realtime stack:
protocol.py(TTS_SAMPLE_RATE22050→24000,resolve_voicerewritten for Chatterbox —.wavpath for cloning,""for default),tts_client.py(plain JSON POST, no SSML/prosody wrapping),_settings.py(tts_urldefault →http://chatterbox:9000,default_voicedefault →""),docker-compose.audio.yml(newchatterbox:service replacestts:),env.audio.example(Magpie/NGC vars removed,CHATTERBOX_PORTadded).
model overview --live— a live fleet dashboard.overviewwas a static description;--livenow probes the running deployment and reports the five "what is it doing right now" views: online (per-backend health), offered (served + candidate models, task families, the endpoint list), busy (in-flight / queued requests), usage (cumulative prompt/generation tokens and finished requests by reason), and endpoints. It is read-only and HTTP-only — it works against a local deployment or amodel tunnelhostname alike, and degrades gracefully when a backend or its metrics is unreachable.- Gateway
GET /status— a model-gear-native JSON aggregate. The fleet's backends are internal-only, so the gateway fans out to each one's/health+/metricsand returns{object: "model-gear.fleet_status", default_model, busy: {running, waiting}, backends: [...], endpoints: [...]}. This is the sourcemodel overview --livereads in the fleet (a bare single-model server is read directly from its/metrics+/health). model_gear._metrics— a small stdlib-only helper that parses vLLM's Prometheus/metrics(running/waiting, prompt/generation tokens,request_success_totalby finish reason, KV-cache usage) and best-effort HTTP probes that never raise.
- Durable vLLM logs that survive restart/recreate (#50). When a vLLM
container restarted, its
docker logs— and any EngineCore crash trace — were lost, which blocked root-causing #50 for lack of data.model initnow scaffoldsmg-logwrap.sh, bind-mounted as each vLLM service's entrypoint: it tees stdout+stderr to a per-boot file<service>-<boot>.logunder a host-mounted log dir (${MODEL_GEAR_LOG_DIR:-<deploy>/logs}→/logs/model-gear), thenexecs the real command so vLLM stays the signal target (graceful shutdown) and the exit code (andrestart:policy) are unchanged. Teeing at the process-I/O level captures both Python tracebacks and native CUDA/C++ aborts; if logging can't be set up it falls back to a plainexecand never blocks serving. The crash boot is preserved as its own file. Wired into the single-model and fleet (primary/embed/rerank) compose templates. Seedocs/durable-logs.md. model logs— new read-only verb to list/tail the durable logs, reading the host files directly so it works even after the crashed container is gone:model logs(list boots),model logs <service>(tail latest), andmodel logs <service> --previous(tail the boot that crashed, after a restart).
model init/model serve/model fleet uppre-create the host log dir (user-owned) before compose bind-mounts it, so logs are never root-owned.
model fleet statusnow reports the embedding + reranker gears.FLEET_CONTAINERSlisted onlyvllm-primary+gateway, somodel fleet statussilently omitted thevllm-embed/vllm-rerankcontainers the default fleet (#44/#47) actually runs. AddedFLEET_EMBED/FLEET_RERANKto the default container set — status now lists all four (the opt-in generate fallback stays excluded, as it is not in the default compose).
- Aligned the agent/human-facing prose with the co-resident gears (#44/#47).
model learn,model overview,model explain(root + fleet),model init --fleethelp, thefleetdocstring, the scaffoldedenv.example/docker-compose.ymlcomments,README.md,CLAUDE.md, anddocs/gateway-fleet.mdstill described the fleet as a "2-model" / "two-container" / "single-backend" deployment. They now describe the default fleet as the generate primary plus co-resident embedding + reranker gears behind one gateway, routed by task family (generate / embed / score / rerank), with the generate fallback as the only opt-in backend. Added a "Task families & gears" section +explain embeddings/explain rerankpointers tomodel learn.
- Embedding + reranker gears (closes #44). model-gear now serves two pooling
gears alongside the chat primary, reachable through the same OpenAI-compatible
gateway and routed by the request's
modelfield:Qwen/Qwen3-Embedding-0.6B—POST /v1/embeddings(vLLM--runner pooling --convert embed), native 1024-dim, MRL-truncatable via thedimensionsparam (Matryoshka--hf-overrides).Qwen/Qwen3-Reranker-0.6B—POST /v1/rerank+/v1/score(vLLM--runner pooling --convert classify, served via theQwen3ForSequenceClassification--hf-overrides).- Catalog:
SupportedModelgainstask(generate/embed/score),dimension, andhf_overrides; both gears surface inmodel overview --listandGET /v1/models/supported. - Fleet:
vllm-embed+vllm-rerankservices in the fleet compose (always-warm, small--max-model-len/--gpu-memory-utilizationso they co-reside with the 27B on a single GB10), wired as gateway backends. - Gateway: task-aware failover — an embed/score request never fails over to a generate backend (and vice versa); chat primary↔fallback failover preserved.
- CLI:
model switch --task {generate,embed,score}for solo serving;model explain embeddings/rerank/scoredocument the call shapes; per-model docs underdocs/. - Boundary: model-gear serves the gears only — no vector store, index, chunker, or retrieval lands here (guarded by a test); storage + retrieval are the consumer's half (eidetic-cli).
- markdownlint: exempt skill prompt templates (
.claude/skills/**/prompts/**) from markdownlint. These are model-facing prompts fed verbatim to a backend (first line is$ARGUMENTSor a prose instruction), so MD041 (first-line H1) and MD032 are inapplicable — a heading would be injected into the prompt.SKILL.mdis still linted; onlyprompts/is exempt. Unblocks thelintCI job after theask-colleagueskill was vendored in.
- The fleet is now single-backend by default (Qwen primary only); the Mistral
fallback is removed. Live validation showed two ~30B NVFP4 models don't co-fit
a shared GB10, so the warm dense Mistral-Small-3.2-24B fallback has been dropped
from the default fleet and the primary restored to its load-tested solo
headroom:
PRIMARY_GPU_MEM_UTIL0.40 → 0.6andPRIMARY_MAX_MODEL_LEN32768 → 262144(full 256K). Thevllm-fallbackservice is gone fromfleet/docker-compose.yml, andFLEET_CONTAINERSno longer includes it. - The gateway makes the fallback optional.
build_confignow adds a second backend only whenFALLBACK_URLorFALLBACK_SERVED_NAMEis set in env — so the default gateway serves the primary alone (no failover target), and a two-backend fleet still works for anyone who wires one up. Routing/failover primitives are unchanged;order_backendsreturns just the primary when solo. - Mistral stays a selectable catalog candidate (
model overview --list) and the documented opt-in fallback — only its role as the default fleet fallback is removed. README,docs/gateway-fleet.md, and themodel explain fleet/gateway/model init --helptext are updated to the single-backend default (with an "Adding a fallback" guide).
docs/gateway-fleet.mduses$HOME/.model-gearinstead of the non-portable~/.model-gear.
Qodo review of #41:
model init --fleet --audionow scaffolds_readiness.py— addedfleet/_readiness.py → _readiness.pyto_compose.AUDIO_TEMPLATES. The ParakeetDockerfile.parakeetCOPY _readiness.pyrequires it at the deployment-dir root, so a clean audio init previously produced a tree wheredocker compose build sttwould fail. Covered bytest_init.py.- Parakeet readiness drift guard + simplification — removed the third
(inline) copy of the readiness decision from
listen_server.py(the scaffold now guarantees the vendored_readiness.pyis present), and added a test asserting the vendored twin stays behaviourally identical to the canonicalmodel_gear/realtime/_readiness.py. - CUDA readiness probe failures are now logged —
listen_server.health()emits alogger.warningwith the exception type/message before returning503, so operators can distinguish driver-down / OOM / stale-context. scripts/audio-smoke.pynow exercises/v1/audio/speech(it previously claimed both routes but only tested transcriptions) and wires the formerly unused--stt-urlto a direct-Parakeet transcription check.docs/realtime-pipeline.mduses$HOME/.model-gearinstead of the non-portable~/.model-gear.
docs/realtime-pipeline.md— the previously-missing runbook for the audio surface: that model-gear owns the live:8080realtime facade, themodel init --fleet --audio/model fleet upbring-up, the topology (gateway path-routes/v1/audio/*→ realtime → Parakeet/Magpie), the drift it fixed (#39/#40), the cheap readiness probe, and the stale-Parakeet-CUDA restart runbook. Resolves a doc referenced frompyproject.toml, the audio overlay, and the realtime app docstring but never written.scripts/audio-smoke.py— a stdlib-only live smoke test for the audio routes: assertsGET :8080/openapi.jsonlists both/v1/audio/transcriptionsand/v1/audio/speech, then POSTs an in-memory 16 kHz WAV and asserts200 {text: …}. Reproduces issue #39's repro to confirm the 500→200 fix. Requires a running GPU box (not a CI unit test).model_gear/realtime/_readiness.py— a stdlib-onlyevaluate_readiness()helper backing the Parakeet/v1/health/readycheap probe; unit-tested in CI without torch/nemo/GPU.
- Parakeet STT healthcheck now reflects real model readiness (#39). The
vendored
templates/fleet/listen_server.py/v1/health/readyreturned{"status": "ready"}unconditionally — process liveness only — so a container whose CUDA context had gone stale (CUDA error: unknown error, every transcription 500ing) still reported Docker "healthy". The probe now reports ready only when the NeMo model is loaded and a trivial CUDA tensor op succeeds, returning503otherwise (a cheap probe, not a full transcription each interval). The pure decision is vendored into the Parakeet build context andCOPY'd into the image so it resolves without the wheel.
scripts/gen-api-key.py— generate or rotate the bearer key (CULTURE_VLLM_API_KEY) that gates the served API. The secret is created with the stdlibsecretsmodule and never hardcoded, so the script is safe in the open-source repo; the key only ever lands in the gitignored deployment.env(written0o600, best-effort). Hidden by default (no echo into logs/scrollback);--showprints it,--forcerotates an existing key, and--bytes(min 16) is validated. Resolves the deployment dir like themodelCLI (--dir→$MODEL_GEAR_DIR→$HOME/.model-gear), degrades gracefully on an unreadable or non-regular.env, and runs from a wheel install (nomodel_gearimport). Referenced from the README "Expose the API" section.
model tunnel— expose the local OpenAI-compatible API from anywhere via a Cloudflare Tunnel (#35). Dry-run by default (prints thecloudflaredcommand and the publichttps://<host>/v1URL);--applystarts a standalonecloudflared tunnel runin the background (logging tocloudflared.login the deployment dir), and--stop --applytears it down. The public hostname resolves--hostname→$CULTURE_VLLM_PUBLIC_HOSTNAME→CULTURE_VLLM_PUBLIC_HOSTNAMEin a gitignored.cf-tunnel.env; the run-token comes fromCULTURE_CF_TUNNEL_TOKEN_SHUSHU(a shushu-sealed secret name, preferred) orCULTURE_CF_TUNNEL_TOKEN(plaintext fallback). The token is never placed on the process argv (so it can't leak viapsor the log) — cloudflared reads it from theTUNNEL_TOKENenvironment variable, whichshushuinjects (sealed mode) or the launcher sets directly (fallback). The resolved hostname and sealed-secret name are validated against a conservative charset before they reach the argv (an argument-injection guard).--applypreflights thatcloudflared(andshushu) is on PATH, that no tunnel is already running for the deployment, and that the local server answers/health;--stopsignals the recorded process group and confirms exit (SIGTERM → SIGKILL) before clearing a PID-reuse-safe pidfile (the recorded pid is identity-checked against/procso a reused pid can't be killed). No hostname, token, or backend checkpoint id is committed. The Cloudflare side (tunnel + ingress + DNS) is provisioned once bycultureflare remote-login --no-access.- Optional bearer auth on the served API via
CULTURE_VLLM_API_KEY, wired into the single-modeldocker-compose.ymlasVLLM_API_KEY=${CULTURE_VLLM_API_KEY:-}. Empty (default) leaves local dev open; set it and vLLM requiresAuthorization: Bearer— the gate for any public exposure. Documented inenv.examplealongside a note thatVLLM_SERVED_NAMEcan be a generic alias to keep the checkpoint name out of the public/v1/models. cf-tunnel.env.examplescaffolded bymodel init(single + fleet), a placeholder-only template the owner copies to the gitignored.cf-tunnel.env.- README "Expose the API from anywhere (Cloudflare Tunnel)" section and a
model explain tunnelcatalog entry.
- Served context raised 128K → full 256K (native) for the MTP primary on DGX
Spark. The
sparkmachine profile'smax_model_lendefault is now262144(was131072), with matching changes to the single-modelenv.example/docker-compose.ymldefaults and themodel switch --help/model explaintext. Load-tested 2026-06-03 on the shared GB10 (util 0.6,--max-num-seqs 2, KV-FP8, MTP n=3): boots clean (CUDA-graph capture, PIECEWISE, 0.71 GiB in 2 s — no OOM), 17.8 tok/s decode, 74.0 % MTP draft acceptance, bothmodel assessprobesfinish=stop, tool-calling probe passes, and 71,601 MiB (~70 GiB) resident — the same footprint as 32K/128K, because--gpu-memory-utilizationfixes the KV-pool reservation (only the addressable context grows). vLLM reports 5.29× max concurrency at a full 256K request, well above the--max-num-seqs 2decode cap, so there is no practical concurrency cost versus the 128K default.model switch --max-model-len <N>still overrides per deployment, and util stays a conservative0.6(shared box). Seedocs/qwen3.6-27b-text-nvfp4-mtp.md(new 256K benchmark) anddocs/tuning-profiles.md. - Catalog
contextstring updated. The MTP primary now reads"256K native (served at full 256K on the shared GB10)". - Scope — deliberately left at the old contexts: fleet templates stay at 32K
(co-residence with the 24B fallback is a different, still-unvalidated memory
regime;
fleet/env.examplenotes this), and thethor/genericmachine profiles stay at 32K (unmeasured estimates) withblackwellat 64K. Themodel switchnative-ceiling clamp (added in 0.16.0) still pins 32K-native candidates (nvidia/Qwen3-32B-NVFP4,mmangkad/Qwen3.6-35B-A3B-NVFP4) down to their own ceilings under the new 256K spark default.
model switchwarns when an uncatalogued model would inherit an unclamped machine context default. The native-ceiling clamp only protects catalogued models; an uncatalogued model ID (whichswitchsupports) inherits the machine default (now spark's 262144) and would boot-fail if the checkpoint's native context is smaller.switchnow emits a clear warning pointing at--max-model-len/ cataloguing, rather than silently applying the high default (no silent clamp — an uncatalogued ceiling is unknown, so guessing one is wrong both ways). Addresses a Qodo reliability finding on #34.
- Served context raised 32K → 128K for the MTP primary on DGX Spark. The
sparkmachine profile'smax_model_lendefault is now131072(was32768), with matching changes in the single-modelenv.example/docker-compose.ymldefaults. Load-tested 2026-06-03 on the shared GB10 (util 0.6,--max-num-seqs 2, KV-FP8, MTP n=3): boots clean (no CUDA-graph-capture OOM), 18.3 tok/s decode, 73.3 % MTP draft acceptance, bothmodel assessprobesfinish=stop, and 71,963 MiB (~70 GiB) resident — the same footprint as 32K, because--gpu-memory-utilizationfixes the KV-pool reservation (the pool holds 9.6× a full 128K request).model switch --max-model-len <N>still overrides per deployment, and util stays a conservative0.6(the box is shared). Seedocs/qwen3.6-27b-text-nvfp4-mtp.md(new 128K benchmark) anddocs/tuning-profiles.md. - Catalog
contextstrings clarified. The MTP primary now reads"256K native (served at 128K on the shared GB10)"; the non-served candidate / fallback entries (mmangkad/Qwen3.6-27B-NVFP4, the Mistral fallback) drop the stale per-model "capped to 32K" note and state native context only. - Scope — deliberately left at the old contexts: fleet templates stay at 32K
(the fleet runs the primary co-resident with a 24B fallback at lower util — a
different memory regime the single-model 128K test does not validate;
fleet/env.examplenotes this), and thethor/genericmachine profiles stay at 32K (unmeasured estimates) withblackwellat 64K.
model switchclamps the machine context default to a model's native ceiling. Raising spark'smax_model_lendefault to131072made it apply to every model switched to on spark — including the 32K-native catalog candidates (nvidia/Qwen3-32B-NVFP4,mmangkad/Qwen3.6-35B-A3B-NVFP4), where vLLM refuses a--max-model-lenabove the checkpoint's native limit (no YaRN) and the container fails to boot.SupportedModelnow carries a numericnative_max_model_len, andmodel switchclamps the resolved context down to it when no explicit--max-model-lenis given (an explicit value still wins, for opted-in YaRN configs). Fixes a Qodo correctness finding on #33.
- Fleet default primary →
sakamakismile/Qwen3.6-27B-Text-NVFP4-MTP(the MTP build), replacingmmangkad/Qwen3.6-27B-NVFP4(issue #26 follow-up). The tool-calling gate that kept it a candidate is now closed: served through the production compose it emits a validqwen3_codertool call, completes a full tool round-trip, keeps its reasoning trace, and runs MTP spec-decode at 78.6% draft acceptance with tool calling on — ~2.4× single-stream decode (8 → ~19 tok/s), ~71 GB footprint, bothmodel assessprobesfinish=stop. Promoted across the catalog (role_hint), the gateway default (_DEFAULT_PRIMARY),whoami, both templateenv.example/docker-compose.ymlfiles, andculture.yaml. - The MTP serve flags are now baked into the compose templates (single-model +
fleet
vllm-primary):--speculative-config,--trust-remote-code,--language-model-only, the--tokenizer=mmangkad/Qwen3.6-27B-NVFP4override, and--max-num-seqs=2. A freshmodel init && model serveof the default now works out of the box. Quantization default ismodelopt. model switchnotices inverted. Because the template ships the MTP primary's flags, switching to a non-MTP model now prints "REMOVE these 4command:lines" (was "add" for the MTP candidate); the MoE--moe-backendadd-notice is unchanged. Switching to the MTP primary force-caps--max-num-seqsto 2.mmangkad/Qwen3.6-27B-NVFP4archived to a candidate — retained as the MTP primary's tokenizer source and the only vision-capable 27B in the catalog.
model switch --applyno longer takes a healthy deployment down when a manual compose edit is required (Qodo review). Switching to a non-MTP model (the template ships the MTP primary's incompatible flags) now writes.envand stops before the restart, printing the lines to remove;--forceoverrides to recreate the container anyway.- MTP compose flags are a single source of truth (
catalog.mtp_compose_command_items()) — consumed by bothmodel switch's removal notice and guarded against drift from the packaged templates by a new test (Qodo review). - Security guidance for the now-default
--trust-remote-codeadded to both compose templates andenv.example: HF_TOKEN is only needed for gated repos (defaults are public) — leave it empty or use a minimal-scope read-only token, and pin trusted revisions (Qodo review). Tracking the upstream tokenizer fix that would let us drop the override in #29.
- MTP (Multi-Token Prediction) candidate for the 27B (issue #26). New catalog
entry
sakamakismile/Qwen3.6-27B-Text-NVFP4-MTP— a text-only re-export of the 27B primary with its MTP draft head restored in bf16 so vLLM speculative decoding actually works. The lesson from the 35B MoE applied: the baseline NVFP4 export drops the MTP head (~0 % draft acceptance), and a newer vLLM isn't installable on the aarch64 GB10 — so the fix is a checkpoint that ships the MTP weights, not a newer engine. Carries a catalogspeculative_config({"method":"qwen3_5_mtp","num_speculative_tokens":3}); quantization ismodelopt. Load-tested on the DGX Spark (GB10) 2026-05-31: 19.1 tok/s decode (~2.4× the baseline 27B's ~8 tok/s) at 72 % MTP draft acceptance on vLLM 0.19.0+nv26.04 — the open risk (does the stock image acceptqwen3_5_mtp?) is cleared (it resolves theQwen3_5MTPdraft head). One tokenizer override is required: the checkpoint declares the newerTokenizersBackendclass (absent from nv26.04), so serve with--tokenizer=mmangkad/Qwen3.6-27B-NVFP4(the cached sibling, same vocab);model switchprints it.- New per-model doc
docs/qwen3.6-27b-text-nvfp4-mtp.mdwith the serve recipe, the live benchmark table (decode tok/s + acceptance vs the baseline), and the caveats (--max-num-seqs 2or it silently OOMs; the tokenizer override).
- New per-model doc
model switchsurfaces MTP serve-extras, not just MoE._moe_notice→_serve_notices(now a list): a model with a catalogspeculative_configprints the exact--speculative-config/--trust-remote-code/--language-model-onlycompose edits (+ theVLLM_MAX_NUM_SEQS=2reminder), the same hand-edit pattern as--moe-backend. The--jsondry-run replaces themoe_noticekey with acompose_editslist.env.example+ theexplaincatalog prose updated to match.
- Workload
purpose+ machine tuning profiles.model switchnow resolves the serve config from three layers — a machine profile (--machine, default auto-detected fromnvidia-smi+ hostname: GPU-memory fraction, context, attention backend), a workload profile (--purpose, defaultbalanced: the batching knobs and the shapemodel benchmarkexercises), and the model's catalog entry — with explicit--max-model-len/--gpu-mem-utilflags overriding the machine defaults.- New
model_gear/profiles.py(pure data module, likecatalog.py):WorkloadProfile(balanced≈1K/1K,prompt-heavy≈8K/1K,decode-heavy≈1K/8K) andMachineProfile(sparkload-tested,thor/blackwell/genericconfigured), guarded bytests/test_profiles.py. - Richer single-model template — the serve command now passes
--attention-backend,--max-num-seqs,--max-num-batched-tokens(env-driven), plus static--enable-chunked-prefill/--async-scheduling. New.envkeys:VLLM_PURPOSE,VLLM_MACHINE,VLLM_ATTENTION_BACKEND,VLLM_MAX_NUM_SEQS,VLLM_MAX_NUM_BATCHED_TOKENS. - Per-model MoE serve extras — the catalog gains
moe_backend/speculative_config(set only on theQwen3.6-35B-A3BMoE candidate).model switchto the MoE prints them as a documented compose edit (they break the dense/hybrid models and can't be defaulted in the shared template). model benchmarkis tied to the config — its workload shape defaults to the configuredVLLM_PURPOSE(overridable with--purpose/--input-len/--output-len).model whoami/model overviewsurface the activegear(purpose/machine);model explain tuningdocuments the layering;docs/tuning-profiles.mdis new.- Credit: the serve tuning and the three workload shapes follow shahizat's
cross-machine NVFP4 benchmark (NVIDIA Developer Forums) — see the README
Acknowledgements and
docs/tuning-profiles.md. - Live-replicated on the shared DGX Spark (2026-05-31) rather than trusting
the post: with the new flags the 35B MoE candidate loads solo (util 0.70,
marlin) and runs single-stream decode ~35 tok/s vs the 27B's ~7.8 — ~4.6×
faster (the MoE's ~3B-active advantage). Numbers + method in
docs/tuning-profiles.mdanddocs/qwen3.6-35b-a3b-nvfp4.md.
- New
model switch--max-model-len/--gpu-mem-utilnow default to the machine profile (was a fixed 32768 / 0.6); pass them explicitly to override.model benchmarkreplaces--decode-tokenswith purpose-driven--input-len/--output-len.- Catalog: dropped the MTP
speculative_configfrom themmangkad/Qwen3.6-35B-A3B-NVFP4entry (kept--moe-backend=marlin). Live testing showed shahizat's MTP draft fails to load on themmangkad/copy (qwen3_5_mtp.pyweight-shape mismatch on vLLM nv26.04) — it is tied to hisnvidia/checkpoint.model switchno longer prints a recipe that wouldn't load.
- Audio I/O behind the gateway (STT + TTS) — issue #18, part 1 of 3. model-gear
now serves OpenAI-compatible
POST /v1/audio/transcriptionsandPOST /v1/audio/speechon the same host port as the text API, fronted by the same stdlib gateway. The audio backends are the same models the standalone realtime-api stack ran — NVIDIA Parakeet STT + Magpie TTS NIM — consolidated into the fleet (no separate compose project; the realtime bridge's LLM is the fleet gateway itself, so there is no extra vLLM container).- New
[realtime]extra +model_gear.realtimepackage (vendored from therealtime-apisibling, cite-don't-import): a FastAPI bridge that exposes the OpenAI audio surface (/v1/audio/speechadapts Magpie's proprietary/v1/audio/synthesize;/v1/audio/transcriptionsforwards to Parakeet). The base wheel and the gateway stay stdlib-only — torch/fastapi never leak into them. - Gateway audio routing —
/v1/audio/*is path-routed to the audio backend (AUDIO_URL) with no model rewrite and no failover; binary responses relayed streamed (chunked) so a large TTS body never buffers whole in the gateway. UnsetAUDIO_URL(a text-only fleet) → those paths 404, unchanged. model init --fleet --audioscaffolds the audio overlay (docker-compose.audio.yml+Dockerfile.realtime+ a vendoredDockerfile.parakeet/listen_server.py) and appends the audio keys to.env.model fleet up/down/statusauto-include the overlay when present.- Co-residence caveat: the audio services share the GPU with the LLM fleet — the overlay is opt-in so text-only boxes keep their GPU budget. See the per-model docs (PR3) for live numbers.
- The realtime WebSocket (
/v1/realtime) and themodel overview/doctor/explainsurface land in the follow-up PRs (parts 2 and 3).
- New
- Audio review hardening (PR #24 review).
- Gateway no longer buffers whole audio bodies —
/v1/audio/*responses are relayed chunked instead ofread_all()'d into memory, so one large TTS WAV can't OOM the fleet's single front door. TTS_CONCURRENCY/TTS_SPEEDclamped to ≥ 1 —TTS_CONCURRENCY=0previously seeded anasyncio.Semaphore(0)that hung every TTS request; a 0/negative speed emitted nonsensicalrate="0%"SSML./v1/audio/speechspeedclamped to OpenAI's 0.25–4.0 range before the Magpie percentage conversion, so out-of-range values no longer reach the backend asrate="{huge|negative}%"and 502.- SonarCloud config — coverage exclusions now mirror
coverage.runomit(the[realtime]-extra modules can't be unit-imported offline), and the deployment scaffolds undermodel_gear/templates/**are excluded from analysis (container Dockerfiles + the vendored Parakeet server aren't package runtime). Added unit tests forrealtime.protocol, the settings clamps, the speed clamp, and the streamed audio relay.
- Gateway no longer buffers whole audio bodies —
model learn --jsonnow includes amodelsobject (supported_catalog/loaded_now) — a machine-readable version of the catalog-vs-loaded explainer for agent consumers. (Additive field; the only observable behavior change in this release.)
- Documented "supported catalog vs. loaded now" consistently across the README,
docs/gateway-fleet.md(new "Supported catalog vs. warm backends" subsection), the per-model docs, and the CLI teaching surfaces (model learn,model explain models/overview/status/whoami/root, and theoverview/status/whoami/fleet statushelp strings). The distinction:model overview --list/GET /v1/models/supported= the gears you can switch to (taggedload-tested/configured, static); the liveGET /v1/models(whichmodel fleet statusqueries) = what's actually loaded now.model status/model whoamireport the configured served model (from.env) + health — not a live/v1/modelsquery. Docs + help text (no serving/runtime behavior change).
RedHatAI/Mistral-Small-3.2-24B-Instruct-2506-NVFP4support — added to the supported-model catalog (model overview --list,GET /v1/models/supported) with a per-model doc,docs/mistral-small-3.2-24b-nvfp4.md. Load-tested on the DGX Spark (GB10): ~15 GiB weights, ~14.9 tok/s decode, prefill 2,009 tok in 1.49 s, tool calling ✅.mistraltool-call parser inference —model_gear.runtime._parsernow maps Mistral-family ids (incl. themistralai/org) to themistralparser;model switchauto-selects it.model switch --quantization— the served--quantizationis now set per model (read from the catalog for a known model, e.g.compressed-tensorsfor the RedHatAI NVFP4 Mistral vsmodelopt_fp4for the nvidia/mmangkad checkpoints);--quantizationoverrides it. The single-model compose readsVLLM_QUANTIZATION.
- Fleet default fallback is now the dense Mistral-Small-3.2-24B, replacing the
mmangkad/Qwen3.6-35B-A3B-NVFP4MoE, which never loaded on the GB10 (OOM co-resident, stall solo — no benchmark obtained). Mistral is dense, loads reliably, and is smaller (~15 GiB weights). The fleet compose serves it with the mistral tokenizer + images limited to 0 (required for tool-call parsing on the nv26.04 build; the HF tokenizer leaks[TOOL_CALLS]markup, and the mistral tokenizer alone crashes the Pixtral profiler) and no--reasoning-parser(instruct model). The 35B MoE is demoted to a catalogue candidate. model_gear.gateway._config._DEFAULT_FALLBACK, the fleetdocker-compose.yml/env.exampleFALLBACK_*defaults,docs/gateway-fleet.md, andREADME.mdupdated for the new fallback.
- Fleet default GPU-mem utilisations rebalanced
0.55/0.30→0.40/0.35. Live validation on a DGX Spark (GB10) showed0.55/0.30OOM-crash-loops the fallback: the 27B primary alone takes ~75 GiB at util 0.6, and--gpu-memory-utilizationis fraction-of-total per process (the two backends don't coordinate). The new values are a dedicated-box estimate; the templates and docs now state plainly that co-residence of two ~30B models needs a dedicated box.
- Docs corrected against live findings (2026-05-30):
docs/gateway-fleet.mdgains a "Live validation findings" section (27B warm-up ~7 min, ~75 GiB footprint, 8.0 tok/s decode; co-residence not viable on a shared GB10).docs/qwen3.6-35b-a3b-nvfp4.mdupdated from "not yet load-tested" to the actual result — the MoE fallback does not load reliably on this box (OOM co-resident; crash/stall even solo).docs/qwen3.6-27b-nvfp4.mdreframed as the fleet default primary (was "candidate") with the warm-up measurement and a corrected recommendation.
GET /v1/models/supportedgateway endpoint — the "change gears" catalog. Alongside the OpenAI-standard/v1/models(which lists only the two loaded backends), the gateway now serves the full catalog of supported models a client can change gears to, each flaggedloaded(a backend serves it now) anddefault(the gateway routes unknown/missing names there). Non-OpenAI shape ("object": "model-gear.supported_models") so/v1/modelsstays standard for existing clients. Puresupported_models_payload()ingateway/_routing.py.- New packaged catalog
model_gear/catalog.py— a dependency-freeSUPPORTED_MODELStuple (the 27B primary, the 32B dense candidate, the 35B-A3B MoE fallback) that is the single source of truth for both the gateway (which runs from a wheel and can't readdocs/) and the CLI.model overview --listis now catalog-backed, so it is populated even in a wheel install.
- Fleet (and single-model) default primary →
mmangkad/Qwen3.6-27B-NVFP4. The scaffolded default served model is now the Qwen3.6 27B (hybrid Mamba/linear-attn + ViT, 256K native context) with--tool-call-parser=qwen3_coder— matching what runs on the DGX Spark and convertible's parent model. The densenvidia/Qwen3-32B-NVFP4remains a supported candidate (PRIMARY_MODEL/model switch). Recomputed co-resident GPU memory:PRIMARY_GPU_MEM_UTIL=0.55andFALLBACK_GPU_MEM_UTIL=0.30(the 27B is heavier than the 32B). Updated the fleet + single-model templates,gateway/_config.py,whoamidefault,culture.yaml/AGENTS.md/CLAUDE.md(served-model coherence chain), and the per-model + gateway-fleet docs.
- Fallback model + single front OpenAI gateway ("fleet"). A new
scaffold-based deployment runs two always-warm vLLM backends behind one
stdlib gateway that model-gear manages as three containers
(
model-gear-gateway,model-gear-vllm-primary,model-gear-vllm-fallback). The gateway routes each request by itsmodelfield, defaults an unknown/missing name to the primary, and fails over to the other backend when the chosen one refuses the connection or returns a 5xx before the response body (4xx is returned verbatim; no mid-stream retry). SSE streams are relayed chunk-by-chunk. Default fallback: the MoEmmangkad/Qwen3.6-35B-A3B-NVFP4. - New gateway package
model_gear/gateway/— a pure-stdlib (http.server+http.client, no runtime deps) reverse proxy:_routing.py(pure name/alias/default routing + failover ordering),_config.py(env → routing table + server config),server.py(thehandle_postfailover seam, upstream client, andThreadingHTTPServerhandler), run aspython -m model_gear.gateway. model init --fleetscaffolds the fleet templates (docker-compose.yml+.env+Dockerfile.gateway) and pinsMODEL_GEAR_VERSIONto the running release;model fleet up | down | statusdrives the deployment (up/downdry-run by default,--applyto commit;statusis read-only and reports all three containers + the gateway/health+/v1/models).- Docs:
docs/gateway-fleet.md(topology, routing/failover, memory, verbs),docs/qwen3.6-35b-a3b-nvfp4.md(the MoE fallback), a README "fleet" section, andmodel explain fleet/model explain gatewayentries.
model_gear/runtime/_compose.pygained a template registry (SINGLE_TEMPLATES/FLEET_TEMPLATES), atemplates=argument onscaffold_plan/write_scaffold(single-model stays the default — existing callers unchanged), acompose_up_buildhelper, andFLEET_CONTAINERS.- The fleet
.envmirrorsVLLM_MODEL/VLLM_SERVED_NAME/VLLM_TOOL_CALL_PARSER(= the primary) so the read-only single-model verbs (status/whoami/doctor) stay coherent on a fleet deployment.model switchremains single-model only.
- SonarCloud cleanup (no behavior change). Split
cmd_switchinto_select_parser/_emit_dry_run/_apply_switchhelpers to bring its cognitive complexity under the gate, and hoisted the repeated"(unset)"literal inmodel statusinto a_UNSETconstant.
- Per-model tool-call parser auto-selection. New
model_gear/runtime/_parser.pyinfer_parser()maps a model name to its parser (qwen3_coderfor Qwen3-Coder / Qwen3.6,hermesfor Qwen3 dense, unknown → leave untouched).model switchnow picks the right parser automatically so tool calling keeps working across a switch without the caller remembering it;--tool-call-parserstill overrides (issue #13). - Post-switch / post-start tool-calling probe.
model switch --applyandmodel serve --applynow probetool_choice:"auto"once the container is healthy and report PASS/FAIL (with the called tool names) — reusing the existingassessprobe.--no-probeskips it; the probe never aborts the command (unreachable / HTTP 400 degrade to a FAIL result). model statusreports the activetool_call_parser(VLLM_TOOL_CALL_PARSER), so "which gear am I in" is complete withoutdocker inspect.
lepenseuris retired; the deployed agent is nowmodel-gear. The tool and the deployed agent share one identity. Updatedculture.yaml(suffix: model-gear), theAGENTS.mdsystem prompt,model whoami/learn/explainoutput, the posting nick (.claude/skills.local.yaml.example), the compose/.envtemplates,README.md, andCLAUDE.md(the former "two identities" section now describes one).
- OpenAI tool/function calling on the served vLLM model. The packaged compose
template (
model_gear/templates/docker-compose.yml) now serves with--enable-auto-tool-choiceand--tool-call-parser=${VLLM_TOOL_CALL_PARSER:-hermes}, sotool_choice:"auto"requests return atool_callsarray instead of HTTP 400. Additive — plain chat/reasoning is unaffected, no extra GPU/memory cost. Unblocks coder-agent harnesses that drive the model entirely through tool calls (issue #9). VLLM_TOOL_CALL_PARSERenv var (defaulthermes) +model switch --tool-call-parser— the parser is per-model:hermesfits Qwen3 dense (e.g.Qwen3-32B), while Qwen3-Coder / Qwen3.6 checkpoints emit the XML function format and needqwen3_coder.switchwrites the var only when the flag is given, so retuning a model never clobbers its parser.model assess --tools— an opt-in tool-calling probe that verifies atool_choice:"auto"request returns atool_callsarray naming afinishfunction. Degrades gracefully (a FAIL row, no abort) against a server that lacks the flags.
- devague workflow trio vendored under
.claude/skills/(cite-don't-import):think(idea→spec),spec-to-plan(spec→plan), andassign-to-workforce(plan→parallel implementation) — the operator chain for the deterministicdevagueCLI. Authored inagentculture/devague, vendored via guildmaster; each carriestype: command(load-bearing on the culture/agex backend, where aSKILL.mdwithouttype:is silently skipped). They drive thedevagueCLI at runtime (uv tool install devague), resolved portably by the wrappers. docs/skill-sources.md— provenance ledger recording the citation path and authoring origin of every vendored skill (the trio plus the six steward-sourced skills).
Redesigned the repo around running, assessing, and switching the local vLLM
model. The model-ops logic that lived in the model-runner skill is now a
first-class CLI. lepenseur is still the deployed agent that consumes the served
model; model-gear is the tool that runs it.
- Model-ops verbs on the
modelCLI:switch <model>,serve(aliasstart) /stop,status,assess(correctness probes),benchmark(decode throughput + prefill), andinit(scaffold a deployment dir). Write verbs (switch/serve/stop/init) are dry-run by default and require--apply(mutation-safety rule). - Scaffold-based deployment.
docker-compose.yml+env.exampleship as packaged templates undermodel_gear/templates/;model initmaterialises them into~/.model-gear(default), aTARGET, or the local folder. Every model-ops verb resolves the deployment dir via--compose-dir→$MODEL_GEAR_DIR→~/.model-gear. - Ported runtime modules (
model_gear/runtime/+model_gear/assess.py), stdlib-only (urllib, fixed-argvsubprocess), with full unit tests. model overviewnow folds in the currently-served model and the candidate-model list, filterable with--current/--list.
- PyPI distribution renamed
lepenseur→model-gear; binarylepenseur→model; Python packagelepenseur→model_gear. Error classLepenseurError→ModelGearError. Thelepenseurconsole script is removed. - Agent-first verbs reframed for the tool:
whoamireports tool/machine/served model/container health/agent;learnteaches the model-ops surface;explaincatalog rewritten (switch/assess/backend/models/…). doctoris now real — checks docker availability, deployment scaffold,.env↔culture.yamlcoherence, and/healthreachability (a down model is a warning, not a failure).- The
model-runnerskill is now a thin shim thatexecsmodel; its_assess.pywas removed (the logic lives inmodel_gear/assess.py). AGENTS.md/culture.yamlclarified: they describe the deployedlepenseuragent, not the repo. README + CLAUDE.md reoriented around model-gear.
- BREAKING: the vLLM container is renamed
lepenseur-vllm→model-gear-vllm. A box running the old container mustdocker compose downunder the old name, thenmodel init --apply+model serve --apply.
model-runnerskill (local, not vendored):switchthe local vLLM runtime model andassess/benchmark it (stdlib_assess.pyfor correctness + throughput, host-side facts via the wrapper). Drives this repo's compose +.env; documented in CLAUDE.md and README. Mutating verbs (switch,down) are dry-run by default and require--apply(CLAUDE.md mutation-safety rule);--portdefaults to.env'sVLLM_PORT(then 8000).
docs/qwen3.6-27b-nvfp4.md: filled with the live load-test (DGX Spark/GB10, 2026-05-27).mmangkad/Qwen3.6-27B-NVFP4loads and serves under our vLLM image (no--trust-remote-code); ~7.9–8.0 tok/s decode, ~70 GB reserved, 29 GB weights. It is a hybrid Mamba/linear-attention vision-language model and is slower on decode than the 32B here — recommendation: keep the 32B. All pre-flight caveats (SGLang-only, multimodal, ModelOpt rc) validated/resolved.
docs/qwen3-32b-nvfp4.md: per-model doc for the current runtime model, with a live test on DGX Spark (GB10) —nvcr.io/nvidia/vllm:26.04-py3(engine0.19.0+...nv26.04), ~9.7 tok/s decode (batch=1), ~2,800 tok/s prefill, ~72 GB reserved atgpu-memory-utilization=0.6, correctness verified.docs/qwen3.6-27b-nvfp4.md: per-model doc for candidatemmangkad/Qwen3.6-27B-NVFP4. ItsQwen3_5ForConditionalGenerationarch is registered in the current vLLM image (so the same compose can serve it); live load-test/benchmark tracked by issue #6.- README "Per-model notes" linking both docs.
docker-compose.yml: corrected the--reasoning-parser=qwen3comment — on the nv26.04 build the<think>trace is returned in thereasoningfield, notreasoning_content.
docker-compose.yml+.env.example: a local vLLM server (NGCnvcr.io/nvidia/vllmimage) that serves the runtime model as an OpenAI-compatible API on:8000for theacpbackend, tuned for DGX Spark (GB10 Blackwell, 128 GB unified memory).- README "Running the model locally (vLLM)" section.
- Switched lepenseur's runtime model from
nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4tonvidia/Qwen3-32B-NVFP4acrossculture.yaml,AGENTS.md,lepenseur/explain/catalog.py,README.md, andCLAUDE.md(32B dense NVFP4 reasoning model with a thinking mode).
- Initial CLI/PyPI sibling scaffold (copied and adapted from the
lecodeurtwin): top-levellepenseurpackage with thelepenseurconsole script. - Read-only verbs:
whoami,learn,explain,overview, and aclinoun withcli overview. doctorverb shipped as a rubric-shaped stub; real self-diagnosis semantics for a thinking ("non-doer") agent are deferred to a follow-up.- Runtime identity files:
AGENTS.mdandculture.yaml(acp backend,vllm-local/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4). - CI:
tests.yml(test + lint +afi cli doctor . --strictgate + version-check) andpublish.yml(PyPI/TestPyPI via Trusted Publishing). - Six vendored skills under
.claude/skills/(cicd, communicate, version-bump, run-tests, sonarclaude, doc-test-alignment), provenance: steward.