All notable changes to this project are documented here. The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.
model overview --live— a live fleet dashboard.overviewwas a static description;--livenow probes the running deployment and reports the five "what is it doing right now" views: online (per-backend health), offered (served + candidate models, task families, the endpoint list), busy (in-flight / queued requests), usage (cumulative prompt/generation tokens and finished requests by reason), and endpoints. It is read-only and HTTP-only — it works against a local deployment or amodel tunnelhostname alike, and degrades gracefully when a backend or its metrics is unreachable.- Gateway
GET /status— a model-gear-native JSON aggregate. The fleet's backends are internal-only, so the gateway fans out to each one's/health+/metricsand returns{object: "model-gear.fleet_status", default_model, busy: {running, waiting}, backends: [...], endpoints: [...]}. This is the sourcemodel overview --livereads in the fleet (a bare single-model server is read directly from its/metrics+/health). model_gear._metrics— a small stdlib-only helper that parses vLLM's Prometheus/metrics(running/waiting, prompt/generation tokens,request_success_totalby finish reason, KV-cache usage) and best-effort HTTP probes that never raise.
- Durable vLLM logs that survive restart/recreate (#50). When a vLLM
container restarted, its
docker logs— and any EngineCore crash trace — were lost, which blocked root-causing #50 for lack of data.model initnow scaffoldsmg-logwrap.sh, bind-mounted as each vLLM service's entrypoint: it tees stdout+stderr to a per-boot file<service>-<boot>.logunder a host-mounted log dir (${MODEL_GEAR_LOG_DIR:-<deploy>/logs}→/logs/model-gear), thenexecs the real command so vLLM stays the signal target (graceful shutdown) and the exit code (andrestart:policy) are unchanged. Teeing at the process-I/O level captures both Python tracebacks and native CUDA/C++ aborts; if logging can't be set up it falls back to a plainexecand never blocks serving. The crash boot is preserved as its own file. Wired into the single-model and fleet (primary/embed/rerank) compose templates. Seedocs/durable-logs.md. model logs— new read-only verb to list/tail the durable logs, reading the host files directly so it works even after the crashed container is gone:model logs(list boots),model logs <service>(tail latest), andmodel logs <service> --previous(tail the boot that crashed, after a restart).
model init/model serve/model fleet uppre-create the host log dir (user-owned) before compose bind-mounts it, so logs are never root-owned.
model fleet statusnow reports the embedding + reranker gears.FLEET_CONTAINERSlisted onlyvllm-primary+gateway, somodel fleet statussilently omitted thevllm-embed/vllm-rerankcontainers the default fleet (#44/#47) actually runs. AddedFLEET_EMBED/FLEET_RERANKto the default container set — status now lists all four (the opt-in generate fallback stays excluded, as it is not in the default compose).
- Aligned the agent/human-facing prose with the co-resident gears (#44/#47).
model learn,model overview,model explain(root + fleet),model init --fleethelp, thefleetdocstring, the scaffoldedenv.example/docker-compose.ymlcomments,README.md,CLAUDE.md, anddocs/gateway-fleet.mdstill described the fleet as a "2-model" / "two-container" / "single-backend" deployment. They now describe the default fleet as the generate primary plus co-resident embedding + reranker gears behind one gateway, routed by task family (generate / embed / score / rerank), with the generate fallback as the only opt-in backend. Added a "Task families & gears" section +explain embeddings/explain rerankpointers tomodel learn.
- Embedding + reranker gears (closes #44). model-gear now serves two pooling
gears alongside the chat primary, reachable through the same OpenAI-compatible
gateway and routed by the request's
modelfield:Qwen/Qwen3-Embedding-0.6B—POST /v1/embeddings(vLLM--runner pooling --convert embed), native 1024-dim, MRL-truncatable via thedimensionsparam (Matryoshka--hf-overrides).Qwen/Qwen3-Reranker-0.6B—POST /v1/rerank+/v1/score(vLLM--runner pooling --convert classify, served via theQwen3ForSequenceClassification--hf-overrides).- Catalog:
SupportedModelgainstask(generate/embed/score),dimension, andhf_overrides; both gears surface inmodel overview --listandGET /v1/models/supported. - Fleet:
vllm-embed+vllm-rerankservices in the fleet compose (always-warm, small--max-model-len/--gpu-memory-utilizationso they co-reside with the 27B on a single GB10), wired as gateway backends. - Gateway: task-aware failover — an embed/score request never fails over to a generate backend (and vice versa); chat primary↔fallback failover preserved.
- CLI:
model switch --task {generate,embed,score}for solo serving;model explain embeddings/rerank/scoredocument the call shapes; per-model docs underdocs/. - Boundary: model-gear serves the gears only — no vector store, index, chunker, or retrieval lands here (guarded by a test); storage + retrieval are the consumer's half (eidetic-cli).
- markdownlint: exempt skill prompt templates (
.claude/skills/**/prompts/**) from markdownlint. These are model-facing prompts fed verbatim to a backend (first line is$ARGUMENTSor a prose instruction), so MD041 (first-line H1) and MD032 are inapplicable — a heading would be injected into the prompt.SKILL.mdis still linted; onlyprompts/is exempt. Unblocks thelintCI job after theask-colleagueskill was vendored in.
- The fleet is now single-backend by default (Qwen primary only); the Mistral
fallback is removed. Live validation showed two ~30B NVFP4 models don't co-fit
a shared GB10, so the warm dense Mistral-Small-3.2-24B fallback has been dropped
from the default fleet and the primary restored to its load-tested solo
headroom:
PRIMARY_GPU_MEM_UTIL0.40 → 0.6andPRIMARY_MAX_MODEL_LEN32768 → 262144(full 256K). Thevllm-fallbackservice is gone fromfleet/docker-compose.yml, andFLEET_CONTAINERSno longer includes it. - The gateway makes the fallback optional.
build_confignow adds a second backend only whenFALLBACK_URLorFALLBACK_SERVED_NAMEis set in env — so the default gateway serves the primary alone (no failover target), and a two-backend fleet still works for anyone who wires one up. Routing/failover primitives are unchanged;order_backendsreturns just the primary when solo. - Mistral stays a selectable catalog candidate (
model overview --list) and the documented opt-in fallback — only its role as the default fleet fallback is removed. README,docs/gateway-fleet.md, and themodel explain fleet/gateway/model init --helptext are updated to the single-backend default (with an "Adding a fallback" guide).
docs/gateway-fleet.mduses$HOME/.model-gearinstead of the non-portable~/.model-gear.
Qodo review of #41:
model init --fleet --audionow scaffolds_readiness.py— addedfleet/_readiness.py → _readiness.pyto_compose.AUDIO_TEMPLATES. The ParakeetDockerfile.parakeetCOPY _readiness.pyrequires it at the deployment-dir root, so a clean audio init previously produced a tree wheredocker compose build sttwould fail. Covered bytest_init.py.- Parakeet readiness drift guard + simplification — removed the third
(inline) copy of the readiness decision from
listen_server.py(the scaffold now guarantees the vendored_readiness.pyis present), and added a test asserting the vendored twin stays behaviourally identical to the canonicalmodel_gear/realtime/_readiness.py. - CUDA readiness probe failures are now logged —
listen_server.health()emits alogger.warningwith the exception type/message before returning503, so operators can distinguish driver-down / OOM / stale-context. scripts/audio-smoke.pynow exercises/v1/audio/speech(it previously claimed both routes but only tested transcriptions) and wires the formerly unused--stt-urlto a direct-Parakeet transcription check.docs/realtime-pipeline.mduses$HOME/.model-gearinstead of the non-portable~/.model-gear.
docs/realtime-pipeline.md— the previously-missing runbook for the audio surface: that model-gear owns the live:8080realtime facade, themodel init --fleet --audio/model fleet upbring-up, the topology (gateway path-routes/v1/audio/*→ realtime → Parakeet/Magpie), the drift it fixed (#39/#40), the cheap readiness probe, and the stale-Parakeet-CUDA restart runbook. Resolves a doc referenced frompyproject.toml, the audio overlay, and the realtime app docstring but never written.scripts/audio-smoke.py— a stdlib-only live smoke test for the audio routes: assertsGET :8080/openapi.jsonlists both/v1/audio/transcriptionsand/v1/audio/speech, then POSTs an in-memory 16 kHz WAV and asserts200 {text: …}. Reproduces issue #39's repro to confirm the 500→200 fix. Requires a running GPU box (not a CI unit test).model_gear/realtime/_readiness.py— a stdlib-onlyevaluate_readiness()helper backing the Parakeet/v1/health/readycheap probe; unit-tested in CI without torch/nemo/GPU.
- Parakeet STT healthcheck now reflects real model readiness (#39). The
vendored
templates/fleet/listen_server.py/v1/health/readyreturned{"status": "ready"}unconditionally — process liveness only — so a container whose CUDA context had gone stale (CUDA error: unknown error, every transcription 500ing) still reported Docker "healthy". The probe now reports ready only when the NeMo model is loaded and a trivial CUDA tensor op succeeds, returning503otherwise (a cheap probe, not a full transcription each interval). The pure decision is vendored into the Parakeet build context andCOPY'd into the image so it resolves without the wheel.
scripts/gen-api-key.py— generate or rotate the bearer key (CULTURE_VLLM_API_KEY) that gates the served API. The secret is created with the stdlibsecretsmodule and never hardcoded, so the script is safe in the open-source repo; the key only ever lands in the gitignored deployment.env(written0o600, best-effort). Hidden by default (no echo into logs/scrollback);--showprints it,--forcerotates an existing key, and--bytes(min 16) is validated. Resolves the deployment dir like themodelCLI (--dir→$MODEL_GEAR_DIR→$HOME/.model-gear), degrades gracefully on an unreadable or non-regular.env, and runs from a wheel install (nomodel_gearimport). Referenced from the README "Expose the API" section.
model tunnel— expose the local OpenAI-compatible API from anywhere via a Cloudflare Tunnel (#35). Dry-run by default (prints thecloudflaredcommand and the publichttps://<host>/v1URL);--applystarts a standalonecloudflared tunnel runin the background (logging tocloudflared.login the deployment dir), and--stop --applytears it down. The public hostname resolves--hostname→$CULTURE_VLLM_PUBLIC_HOSTNAME→CULTURE_VLLM_PUBLIC_HOSTNAMEin a gitignored.cf-tunnel.env; the run-token comes fromCULTURE_CF_TUNNEL_TOKEN_SHUSHU(a shushu-sealed secret name, preferred) orCULTURE_CF_TUNNEL_TOKEN(plaintext fallback). The token is never placed on the process argv (so it can't leak viapsor the log) — cloudflared reads it from theTUNNEL_TOKENenvironment variable, whichshushuinjects (sealed mode) or the launcher sets directly (fallback). The resolved hostname and sealed-secret name are validated against a conservative charset before they reach the argv (an argument-injection guard).--applypreflights thatcloudflared(andshushu) is on PATH, that no tunnel is already running for the deployment, and that the local server answers/health;--stopsignals the recorded process group and confirms exit (SIGTERM → SIGKILL) before clearing a PID-reuse-safe pidfile (the recorded pid is identity-checked against/procso a reused pid can't be killed). No hostname, token, or backend checkpoint id is committed. The Cloudflare side (tunnel + ingress + DNS) is provisioned once bycultureflare remote-login --no-access.- Optional bearer auth on the served API via
CULTURE_VLLM_API_KEY, wired into the single-modeldocker-compose.ymlasVLLM_API_KEY=${CULTURE_VLLM_API_KEY:-}. Empty (default) leaves local dev open; set it and vLLM requiresAuthorization: Bearer— the gate for any public exposure. Documented inenv.examplealongside a note thatVLLM_SERVED_NAMEcan be a generic alias to keep the checkpoint name out of the public/v1/models. cf-tunnel.env.examplescaffolded bymodel init(single + fleet), a placeholder-only template the owner copies to the gitignored.cf-tunnel.env.- README "Expose the API from anywhere (Cloudflare Tunnel)" section and a
model explain tunnelcatalog entry.
- Served context raised 128K → full 256K (native) for the MTP primary on DGX
Spark. The
sparkmachine profile'smax_model_lendefault is now262144(was131072), with matching changes to the single-modelenv.example/docker-compose.ymldefaults and themodel switch --help/model explaintext. Load-tested 2026-06-03 on the shared GB10 (util 0.6,--max-num-seqs 2, KV-FP8, MTP n=3): boots clean (CUDA-graph capture, PIECEWISE, 0.71 GiB in 2 s — no OOM), 17.8 tok/s decode, 74.0 % MTP draft acceptance, bothmodel assessprobesfinish=stop, tool-calling probe passes, and 71,601 MiB (~70 GiB) resident — the same footprint as 32K/128K, because--gpu-memory-utilizationfixes the KV-pool reservation (only the addressable context grows). vLLM reports 5.29× max concurrency at a full 256K request, well above the--max-num-seqs 2decode cap, so there is no practical concurrency cost versus the 128K default.model switch --max-model-len <N>still overrides per deployment, and util stays a conservative0.6(shared box). Seedocs/qwen3.6-27b-text-nvfp4-mtp.md(new 256K benchmark) anddocs/tuning-profiles.md. - Catalog
contextstring updated. The MTP primary now reads"256K native (served at full 256K on the shared GB10)". - Scope — deliberately left at the old contexts: fleet templates stay at 32K
(co-residence with the 24B fallback is a different, still-unvalidated memory
regime;
fleet/env.examplenotes this), and thethor/genericmachine profiles stay at 32K (unmeasured estimates) withblackwellat 64K. Themodel switchnative-ceiling clamp (added in 0.16.0) still pins 32K-native candidates (nvidia/Qwen3-32B-NVFP4,mmangkad/Qwen3.6-35B-A3B-NVFP4) down to their own ceilings under the new 256K spark default.
model switchwarns when an uncatalogued model would inherit an unclamped machine context default. The native-ceiling clamp only protects catalogued models; an uncatalogued model ID (whichswitchsupports) inherits the machine default (now spark's 262144) and would boot-fail if the checkpoint's native context is smaller.switchnow emits a clear warning pointing at--max-model-len/ cataloguing, rather than silently applying the high default (no silent clamp — an uncatalogued ceiling is unknown, so guessing one is wrong both ways). Addresses a Qodo reliability finding on #34.
- Served context raised 32K → 128K for the MTP primary on DGX Spark. The
sparkmachine profile'smax_model_lendefault is now131072(was32768), with matching changes in the single-modelenv.example/docker-compose.ymldefaults. Load-tested 2026-06-03 on the shared GB10 (util 0.6,--max-num-seqs 2, KV-FP8, MTP n=3): boots clean (no CUDA-graph-capture OOM), 18.3 tok/s decode, 73.3 % MTP draft acceptance, bothmodel assessprobesfinish=stop, and 71,963 MiB (~70 GiB) resident — the same footprint as 32K, because--gpu-memory-utilizationfixes the KV-pool reservation (the pool holds 9.6× a full 128K request).model switch --max-model-len <N>still overrides per deployment, and util stays a conservative0.6(the box is shared). Seedocs/qwen3.6-27b-text-nvfp4-mtp.md(new 128K benchmark) anddocs/tuning-profiles.md. - Catalog
contextstrings clarified. The MTP primary now reads"256K native (served at 128K on the shared GB10)"; the non-served candidate / fallback entries (mmangkad/Qwen3.6-27B-NVFP4, the Mistral fallback) drop the stale per-model "capped to 32K" note and state native context only. - Scope — deliberately left at the old contexts: fleet templates stay at 32K
(the fleet runs the primary co-resident with a 24B fallback at lower util — a
different memory regime the single-model 128K test does not validate;
fleet/env.examplenotes this), and thethor/genericmachine profiles stay at 32K (unmeasured estimates) withblackwellat 64K.
model switchclamps the machine context default to a model's native ceiling. Raising spark'smax_model_lendefault to131072made it apply to every model switched to on spark — including the 32K-native catalog candidates (nvidia/Qwen3-32B-NVFP4,mmangkad/Qwen3.6-35B-A3B-NVFP4), where vLLM refuses a--max-model-lenabove the checkpoint's native limit (no YaRN) and the container fails to boot.SupportedModelnow carries a numericnative_max_model_len, andmodel switchclamps the resolved context down to it when no explicit--max-model-lenis given (an explicit value still wins, for opted-in YaRN configs). Fixes a Qodo correctness finding on #33.
- Fleet default primary →
sakamakismile/Qwen3.6-27B-Text-NVFP4-MTP(the MTP build), replacingmmangkad/Qwen3.6-27B-NVFP4(issue #26 follow-up). The tool-calling gate that kept it a candidate is now closed: served through the production compose it emits a validqwen3_codertool call, completes a full tool round-trip, keeps its reasoning trace, and runs MTP spec-decode at 78.6% draft acceptance with tool calling on — ~2.4× single-stream decode (8 → ~19 tok/s), ~71 GB footprint, bothmodel assessprobesfinish=stop. Promoted across the catalog (role_hint), the gateway default (_DEFAULT_PRIMARY),whoami, both templateenv.example/docker-compose.ymlfiles, andculture.yaml. - The MTP serve flags are now baked into the compose templates (single-model +
fleet
vllm-primary):--speculative-config,--trust-remote-code,--language-model-only, the--tokenizer=mmangkad/Qwen3.6-27B-NVFP4override, and--max-num-seqs=2. A freshmodel init && model serveof the default now works out of the box. Quantization default ismodelopt. model switchnotices inverted. Because the template ships the MTP primary's flags, switching to a non-MTP model now prints "REMOVE these 4command:lines" (was "add" for the MTP candidate); the MoE--moe-backendadd-notice is unchanged. Switching to the MTP primary force-caps--max-num-seqsto 2.mmangkad/Qwen3.6-27B-NVFP4archived to a candidate — retained as the MTP primary's tokenizer source and the only vision-capable 27B in the catalog.
model switch --applyno longer takes a healthy deployment down when a manual compose edit is required (Qodo review). Switching to a non-MTP model (the template ships the MTP primary's incompatible flags) now writes.envand stops before the restart, printing the lines to remove;--forceoverrides to recreate the container anyway.- MTP compose flags are a single source of truth (
catalog.mtp_compose_command_items()) — consumed by bothmodel switch's removal notice and guarded against drift from the packaged templates by a new test (Qodo review). - Security guidance for the now-default
--trust-remote-codeadded to both compose templates andenv.example: HF_TOKEN is only needed for gated repos (defaults are public) — leave it empty or use a minimal-scope read-only token, and pin trusted revisions (Qodo review). Tracking the upstream tokenizer fix that would let us drop the override in #29.
- MTP (Multi-Token Prediction) candidate for the 27B (issue #26). New catalog
entry
sakamakismile/Qwen3.6-27B-Text-NVFP4-MTP— a text-only re-export of the 27B primary with its MTP draft head restored in bf16 so vLLM speculative decoding actually works. The lesson from the 35B MoE applied: the baseline NVFP4 export drops the MTP head (~0 % draft acceptance), and a newer vLLM isn't installable on the aarch64 GB10 — so the fix is a checkpoint that ships the MTP weights, not a newer engine. Carries a catalogspeculative_config({"method":"qwen3_5_mtp","num_speculative_tokens":3}); quantization ismodelopt. Load-tested on the DGX Spark (GB10) 2026-05-31: 19.1 tok/s decode (~2.4× the baseline 27B's ~8 tok/s) at 72 % MTP draft acceptance on vLLM 0.19.0+nv26.04 — the open risk (does the stock image acceptqwen3_5_mtp?) is cleared (it resolves theQwen3_5MTPdraft head). One tokenizer override is required: the checkpoint declares the newerTokenizersBackendclass (absent from nv26.04), so serve with--tokenizer=mmangkad/Qwen3.6-27B-NVFP4(the cached sibling, same vocab);model switchprints it.- New per-model doc
docs/qwen3.6-27b-text-nvfp4-mtp.mdwith the serve recipe, the live benchmark table (decode tok/s + acceptance vs the baseline), and the caveats (--max-num-seqs 2or it silently OOMs; the tokenizer override).
- New per-model doc
model switchsurfaces MTP serve-extras, not just MoE._moe_notice→_serve_notices(now a list): a model with a catalogspeculative_configprints the exact--speculative-config/--trust-remote-code/--language-model-onlycompose edits (+ theVLLM_MAX_NUM_SEQS=2reminder), the same hand-edit pattern as--moe-backend. The--jsondry-run replaces themoe_noticekey with acompose_editslist.env.example+ theexplaincatalog prose updated to match.
- Workload
purpose+ machine tuning profiles.model switchnow resolves the serve config from three layers — a machine profile (--machine, default auto-detected fromnvidia-smi+ hostname: GPU-memory fraction, context, attention backend), a workload profile (--purpose, defaultbalanced: the batching knobs and the shapemodel benchmarkexercises), and the model's catalog entry — with explicit--max-model-len/--gpu-mem-utilflags overriding the machine defaults.- New
model_gear/profiles.py(pure data module, likecatalog.py):WorkloadProfile(balanced≈1K/1K,prompt-heavy≈8K/1K,decode-heavy≈1K/8K) andMachineProfile(sparkload-tested,thor/blackwell/genericconfigured), guarded bytests/test_profiles.py. - Richer single-model template — the serve command now passes
--attention-backend,--max-num-seqs,--max-num-batched-tokens(env-driven), plus static--enable-chunked-prefill/--async-scheduling. New.envkeys:VLLM_PURPOSE,VLLM_MACHINE,VLLM_ATTENTION_BACKEND,VLLM_MAX_NUM_SEQS,VLLM_MAX_NUM_BATCHED_TOKENS. - Per-model MoE serve extras — the catalog gains
moe_backend/speculative_config(set only on theQwen3.6-35B-A3BMoE candidate).model switchto the MoE prints them as a documented compose edit (they break the dense/hybrid models and can't be defaulted in the shared template). model benchmarkis tied to the config — its workload shape defaults to the configuredVLLM_PURPOSE(overridable with--purpose/--input-len/--output-len).model whoami/model overviewsurface the activegear(purpose/machine);model explain tuningdocuments the layering;docs/tuning-profiles.mdis new.- Credit: the serve tuning and the three workload shapes follow shahizat's
cross-machine NVFP4 benchmark (NVIDIA Developer Forums) — see the README
Acknowledgements and
docs/tuning-profiles.md. - Live-replicated on the shared DGX Spark (2026-05-31) rather than trusting
the post: with the new flags the 35B MoE candidate loads solo (util 0.70,
marlin) and runs single-stream decode ~35 tok/s vs the 27B's ~7.8 — ~4.6×
faster (the MoE's ~3B-active advantage). Numbers + method in
docs/tuning-profiles.mdanddocs/qwen3.6-35b-a3b-nvfp4.md.
- New
model switch--max-model-len/--gpu-mem-utilnow default to the machine profile (was a fixed 32768 / 0.6); pass them explicitly to override.model benchmarkreplaces--decode-tokenswith purpose-driven--input-len/--output-len.- Catalog: dropped the MTP
speculative_configfrom themmangkad/Qwen3.6-35B-A3B-NVFP4entry (kept--moe-backend=marlin). Live testing showed shahizat's MTP draft fails to load on themmangkad/copy (qwen3_5_mtp.pyweight-shape mismatch on vLLM nv26.04) — it is tied to hisnvidia/checkpoint.model switchno longer prints a recipe that wouldn't load.
- Audio I/O behind the gateway (STT + TTS) — issue #18, part 1 of 3. model-gear
now serves OpenAI-compatible
POST /v1/audio/transcriptionsandPOST /v1/audio/speechon the same host port as the text API, fronted by the same stdlib gateway. The audio backends are the same models the standalone realtime-api stack ran — NVIDIA Parakeet STT + Magpie TTS NIM — consolidated into the fleet (no separate compose project; the realtime bridge's LLM is the fleet gateway itself, so there is no extra vLLM container).- New
[realtime]extra +model_gear.realtimepackage (vendored from therealtime-apisibling, cite-don't-import): a FastAPI bridge that exposes the OpenAI audio surface (/v1/audio/speechadapts Magpie's proprietary/v1/audio/synthesize;/v1/audio/transcriptionsforwards to Parakeet). The base wheel and the gateway stay stdlib-only — torch/fastapi never leak into them. - Gateway audio routing —
/v1/audio/*is path-routed to the audio backend (AUDIO_URL) with no model rewrite and no failover; binary responses relayed streamed (chunked) so a large TTS body never buffers whole in the gateway. UnsetAUDIO_URL(a text-only fleet) → those paths 404, unchanged. model init --fleet --audioscaffolds the audio overlay (docker-compose.audio.yml+Dockerfile.realtime+ a vendoredDockerfile.parakeet/listen_server.py) and appends the audio keys to.env.model fleet up/down/statusauto-include the overlay when present.- Co-residence caveat: the audio services share the GPU with the LLM fleet — the overlay is opt-in so text-only boxes keep their GPU budget. See the per-model docs (PR3) for live numbers.
- The realtime WebSocket (
/v1/realtime) and themodel overview/doctor/explainsurface land in the follow-up PRs (parts 2 and 3).
- New
- Audio review hardening (PR #24 review).
- Gateway no longer buffers whole audio bodies —
/v1/audio/*responses are relayed chunked instead ofread_all()'d into memory, so one large TTS WAV can't OOM the fleet's single front door. TTS_CONCURRENCY/TTS_SPEEDclamped to ≥ 1 —TTS_CONCURRENCY=0previously seeded anasyncio.Semaphore(0)that hung every TTS request; a 0/negative speed emitted nonsensicalrate="0%"SSML./v1/audio/speechspeedclamped to OpenAI's 0.25–4.0 range before the Magpie percentage conversion, so out-of-range values no longer reach the backend asrate="{huge|negative}%"and 502.- SonarCloud config — coverage exclusions now mirror
coverage.runomit(the[realtime]-extra modules can't be unit-imported offline), and the deployment scaffolds undermodel_gear/templates/**are excluded from analysis (container Dockerfiles + the vendored Parakeet server aren't package runtime). Added unit tests forrealtime.protocol, the settings clamps, the speed clamp, and the streamed audio relay.
- Gateway no longer buffers whole audio bodies —
model learn --jsonnow includes amodelsobject (supported_catalog/loaded_now) — a machine-readable version of the catalog-vs-loaded explainer for agent consumers. (Additive field; the only observable behavior change in this release.)
- Documented "supported catalog vs. loaded now" consistently across the README,
docs/gateway-fleet.md(new "Supported catalog vs. warm backends" subsection), the per-model docs, and the CLI teaching surfaces (model learn,model explain models/overview/status/whoami/root, and theoverview/status/whoami/fleet statushelp strings). The distinction:model overview --list/GET /v1/models/supported= the gears you can switch to (taggedload-tested/configured, static); the liveGET /v1/models(whichmodel fleet statusqueries) = what's actually loaded now.model status/model whoamireport the configured served model (from.env) + health — not a live/v1/modelsquery. Docs + help text (no serving/runtime behavior change).
RedHatAI/Mistral-Small-3.2-24B-Instruct-2506-NVFP4support — added to the supported-model catalog (model overview --list,GET /v1/models/supported) with a per-model doc,docs/mistral-small-3.2-24b-nvfp4.md. Load-tested on the DGX Spark (GB10): ~15 GiB weights, ~14.9 tok/s decode, prefill 2,009 tok in 1.49 s, tool calling ✅.mistraltool-call parser inference —model_gear.runtime._parsernow maps Mistral-family ids (incl. themistralai/org) to themistralparser;model switchauto-selects it.model switch --quantization— the served--quantizationis now set per model (read from the catalog for a known model, e.g.compressed-tensorsfor the RedHatAI NVFP4 Mistral vsmodelopt_fp4for the nvidia/mmangkad checkpoints);--quantizationoverrides it. The single-model compose readsVLLM_QUANTIZATION.
- Fleet default fallback is now the dense Mistral-Small-3.2-24B, replacing the
mmangkad/Qwen3.6-35B-A3B-NVFP4MoE, which never loaded on the GB10 (OOM co-resident, stall solo — no benchmark obtained). Mistral is dense, loads reliably, and is smaller (~15 GiB weights). The fleet compose serves it with the mistral tokenizer + images limited to 0 (required for tool-call parsing on the nv26.04 build; the HF tokenizer leaks[TOOL_CALLS]markup, and the mistral tokenizer alone crashes the Pixtral profiler) and no--reasoning-parser(instruct model). The 35B MoE is demoted to a catalogue candidate. model_gear.gateway._config._DEFAULT_FALLBACK, the fleetdocker-compose.yml/env.exampleFALLBACK_*defaults,docs/gateway-fleet.md, andREADME.mdupdated for the new fallback.
- Fleet default GPU-mem utilisations rebalanced
0.55/0.30→0.40/0.35. Live validation on a DGX Spark (GB10) showed0.55/0.30OOM-crash-loops the fallback: the 27B primary alone takes ~75 GiB at util 0.6, and--gpu-memory-utilizationis fraction-of-total per process (the two backends don't coordinate). The new values are a dedicated-box estimate; the templates and docs now state plainly that co-residence of two ~30B models needs a dedicated box.
- Docs corrected against live findings (2026-05-30):
docs/gateway-fleet.mdgains a "Live validation findings" section (27B warm-up ~7 min, ~75 GiB footprint, 8.0 tok/s decode; co-residence not viable on a shared GB10).docs/qwen3.6-35b-a3b-nvfp4.mdupdated from "not yet load-tested" to the actual result — the MoE fallback does not load reliably on this box (OOM co-resident; crash/stall even solo).docs/qwen3.6-27b-nvfp4.mdreframed as the fleet default primary (was "candidate") with the warm-up measurement and a corrected recommendation.
GET /v1/models/supportedgateway endpoint — the "change gears" catalog. Alongside the OpenAI-standard/v1/models(which lists only the two loaded backends), the gateway now serves the full catalog of supported models a client can change gears to, each flaggedloaded(a backend serves it now) anddefault(the gateway routes unknown/missing names there). Non-OpenAI shape ("object": "model-gear.supported_models") so/v1/modelsstays standard for existing clients. Puresupported_models_payload()ingateway/_routing.py.- New packaged catalog
model_gear/catalog.py— a dependency-freeSUPPORTED_MODELStuple (the 27B primary, the 32B dense candidate, the 35B-A3B MoE fallback) that is the single source of truth for both the gateway (which runs from a wheel and can't readdocs/) and the CLI.model overview --listis now catalog-backed, so it is populated even in a wheel install.
- Fleet (and single-model) default primary →
mmangkad/Qwen3.6-27B-NVFP4. The scaffolded default served model is now the Qwen3.6 27B (hybrid Mamba/linear-attn + ViT, 256K native context) with--tool-call-parser=qwen3_coder— matching what runs on the DGX Spark and convertible's parent model. The densenvidia/Qwen3-32B-NVFP4remains a supported candidate (PRIMARY_MODEL/model switch). Recomputed co-resident GPU memory:PRIMARY_GPU_MEM_UTIL=0.55andFALLBACK_GPU_MEM_UTIL=0.30(the 27B is heavier than the 32B). Updated the fleet + single-model templates,gateway/_config.py,whoamidefault,culture.yaml/AGENTS.md/CLAUDE.md(served-model coherence chain), and the per-model + gateway-fleet docs.
- Fallback model + single front OpenAI gateway ("fleet"). A new
scaffold-based deployment runs two always-warm vLLM backends behind one
stdlib gateway that model-gear manages as three containers
(
model-gear-gateway,model-gear-vllm-primary,model-gear-vllm-fallback). The gateway routes each request by itsmodelfield, defaults an unknown/missing name to the primary, and fails over to the other backend when the chosen one refuses the connection or returns a 5xx before the response body (4xx is returned verbatim; no mid-stream retry). SSE streams are relayed chunk-by-chunk. Default fallback: the MoEmmangkad/Qwen3.6-35B-A3B-NVFP4. - New gateway package
model_gear/gateway/— a pure-stdlib (http.server+http.client, no runtime deps) reverse proxy:_routing.py(pure name/alias/default routing + failover ordering),_config.py(env → routing table + server config),server.py(thehandle_postfailover seam, upstream client, andThreadingHTTPServerhandler), run aspython -m model_gear.gateway. model init --fleetscaffolds the fleet templates (docker-compose.yml+.env+Dockerfile.gateway) and pinsMODEL_GEAR_VERSIONto the running release;model fleet up | down | statusdrives the deployment (up/downdry-run by default,--applyto commit;statusis read-only and reports all three containers + the gateway/health+/v1/models).- Docs:
docs/gateway-fleet.md(topology, routing/failover, memory, verbs),docs/qwen3.6-35b-a3b-nvfp4.md(the MoE fallback), a README "fleet" section, andmodel explain fleet/model explain gatewayentries.
model_gear/runtime/_compose.pygained a template registry (SINGLE_TEMPLATES/FLEET_TEMPLATES), atemplates=argument onscaffold_plan/write_scaffold(single-model stays the default — existing callers unchanged), acompose_up_buildhelper, andFLEET_CONTAINERS.- The fleet
.envmirrorsVLLM_MODEL/VLLM_SERVED_NAME/VLLM_TOOL_CALL_PARSER(= the primary) so the read-only single-model verbs (status/whoami/doctor) stay coherent on a fleet deployment.model switchremains single-model only.
- SonarCloud cleanup (no behavior change). Split
cmd_switchinto_select_parser/_emit_dry_run/_apply_switchhelpers to bring its cognitive complexity under the gate, and hoisted the repeated"(unset)"literal inmodel statusinto a_UNSETconstant.
- Per-model tool-call parser auto-selection. New
model_gear/runtime/_parser.pyinfer_parser()maps a model name to its parser (qwen3_coderfor Qwen3-Coder / Qwen3.6,hermesfor Qwen3 dense, unknown → leave untouched).model switchnow picks the right parser automatically so tool calling keeps working across a switch without the caller remembering it;--tool-call-parserstill overrides (issue #13). - Post-switch / post-start tool-calling probe.
model switch --applyandmodel serve --applynow probetool_choice:"auto"once the container is healthy and report PASS/FAIL (with the called tool names) — reusing the existingassessprobe.--no-probeskips it; the probe never aborts the command (unreachable / HTTP 400 degrade to a FAIL result). model statusreports the activetool_call_parser(VLLM_TOOL_CALL_PARSER), so "which gear am I in" is complete withoutdocker inspect.
lepenseuris retired; the deployed agent is nowmodel-gear. The tool and the deployed agent share one identity. Updatedculture.yaml(suffix: model-gear), theAGENTS.mdsystem prompt,model whoami/learn/explainoutput, the posting nick (.claude/skills.local.yaml.example), the compose/.envtemplates,README.md, andCLAUDE.md(the former "two identities" section now describes one).
- OpenAI tool/function calling on the served vLLM model. The packaged compose
template (
model_gear/templates/docker-compose.yml) now serves with--enable-auto-tool-choiceand--tool-call-parser=${VLLM_TOOL_CALL_PARSER:-hermes}, sotool_choice:"auto"requests return atool_callsarray instead of HTTP 400. Additive — plain chat/reasoning is unaffected, no extra GPU/memory cost. Unblocks coder-agent harnesses that drive the model entirely through tool calls (issue #9). VLLM_TOOL_CALL_PARSERenv var (defaulthermes) +model switch --tool-call-parser— the parser is per-model:hermesfits Qwen3 dense (e.g.Qwen3-32B), while Qwen3-Coder / Qwen3.6 checkpoints emit the XML function format and needqwen3_coder.switchwrites the var only when the flag is given, so retuning a model never clobbers its parser.model assess --tools— an opt-in tool-calling probe that verifies atool_choice:"auto"request returns atool_callsarray naming afinishfunction. Degrades gracefully (a FAIL row, no abort) against a server that lacks the flags.
- devague workflow trio vendored under
.claude/skills/(cite-don't-import):think(idea→spec),spec-to-plan(spec→plan), andassign-to-workforce(plan→parallel implementation) — the operator chain for the deterministicdevagueCLI. Authored inagentculture/devague, vendored via guildmaster; each carriestype: command(load-bearing on the culture/agex backend, where aSKILL.mdwithouttype:is silently skipped). They drive thedevagueCLI at runtime (uv tool install devague), resolved portably by the wrappers. docs/skill-sources.md— provenance ledger recording the citation path and authoring origin of every vendored skill (the trio plus the six steward-sourced skills).
Redesigned the repo around running, assessing, and switching the local vLLM
model. The model-ops logic that lived in the model-runner skill is now a
first-class CLI. lepenseur is still the deployed agent that consumes the served
model; model-gear is the tool that runs it.
- Model-ops verbs on the
modelCLI:switch <model>,serve(aliasstart) /stop,status,assess(correctness probes),benchmark(decode throughput + prefill), andinit(scaffold a deployment dir). Write verbs (switch/serve/stop/init) are dry-run by default and require--apply(mutation-safety rule). - Scaffold-based deployment.
docker-compose.yml+env.exampleship as packaged templates undermodel_gear/templates/;model initmaterialises them into~/.model-gear(default), aTARGET, or the local folder. Every model-ops verb resolves the deployment dir via--compose-dir→$MODEL_GEAR_DIR→~/.model-gear. - Ported runtime modules (
model_gear/runtime/+model_gear/assess.py), stdlib-only (urllib, fixed-argvsubprocess), with full unit tests. model overviewnow folds in the currently-served model and the candidate-model list, filterable with--current/--list.
- PyPI distribution renamed
lepenseur→model-gear; binarylepenseur→model; Python packagelepenseur→model_gear. Error classLepenseurError→ModelGearError. Thelepenseurconsole script is removed. - Agent-first verbs reframed for the tool:
whoamireports tool/machine/served model/container health/agent;learnteaches the model-ops surface;explaincatalog rewritten (switch/assess/backend/models/…). doctoris now real — checks docker availability, deployment scaffold,.env↔culture.yamlcoherence, and/healthreachability (a down model is a warning, not a failure).- The
model-runnerskill is now a thin shim thatexecsmodel; its_assess.pywas removed (the logic lives inmodel_gear/assess.py). AGENTS.md/culture.yamlclarified: they describe the deployedlepenseuragent, not the repo. README + CLAUDE.md reoriented around model-gear.
- BREAKING: the vLLM container is renamed
lepenseur-vllm→model-gear-vllm. A box running the old container mustdocker compose downunder the old name, thenmodel init --apply+model serve --apply.
model-runnerskill (local, not vendored):switchthe local vLLM runtime model andassess/benchmark it (stdlib_assess.pyfor correctness + throughput, host-side facts via the wrapper). Drives this repo's compose +.env; documented in CLAUDE.md and README. Mutating verbs (switch,down) are dry-run by default and require--apply(CLAUDE.md mutation-safety rule);--portdefaults to.env'sVLLM_PORT(then 8000).
docs/qwen3.6-27b-nvfp4.md: filled with the live load-test (DGX Spark/GB10, 2026-05-27).mmangkad/Qwen3.6-27B-NVFP4loads and serves under our vLLM image (no--trust-remote-code); ~7.9–8.0 tok/s decode, ~70 GB reserved, 29 GB weights. It is a hybrid Mamba/linear-attention vision-language model and is slower on decode than the 32B here — recommendation: keep the 32B. All pre-flight caveats (SGLang-only, multimodal, ModelOpt rc) validated/resolved.
docs/qwen3-32b-nvfp4.md: per-model doc for the current runtime model, with a live test on DGX Spark (GB10) —nvcr.io/nvidia/vllm:26.04-py3(engine0.19.0+...nv26.04), ~9.7 tok/s decode (batch=1), ~2,800 tok/s prefill, ~72 GB reserved atgpu-memory-utilization=0.6, correctness verified.docs/qwen3.6-27b-nvfp4.md: per-model doc for candidatemmangkad/Qwen3.6-27B-NVFP4. ItsQwen3_5ForConditionalGenerationarch is registered in the current vLLM image (so the same compose can serve it); live load-test/benchmark tracked by issue #6.- README "Per-model notes" linking both docs.
docker-compose.yml: corrected the--reasoning-parser=qwen3comment — on the nv26.04 build the<think>trace is returned in thereasoningfield, notreasoning_content.
docker-compose.yml+.env.example: a local vLLM server (NGCnvcr.io/nvidia/vllmimage) that serves the runtime model as an OpenAI-compatible API on:8000for theacpbackend, tuned for DGX Spark (GB10 Blackwell, 128 GB unified memory).- README "Running the model locally (vLLM)" section.
- Switched lepenseur's runtime model from
nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4tonvidia/Qwen3-32B-NVFP4acrossculture.yaml,AGENTS.md,lepenseur/explain/catalog.py,README.md, andCLAUDE.md(32B dense NVFP4 reasoning model with a thinking mode).
- Initial CLI/PyPI sibling scaffold (copied and adapted from the
lecodeurtwin): top-levellepenseurpackage with thelepenseurconsole script. - Read-only verbs:
whoami,learn,explain,overview, and aclinoun withcli overview. doctorverb shipped as a rubric-shaped stub; real self-diagnosis semantics for a thinking ("non-doer") agent are deferred to a follow-up.- Runtime identity files:
AGENTS.mdandculture.yaml(acp backend,vllm-local/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4). - CI:
tests.yml(test + lint +afi cli doctor . --strictgate + version-check) andpublish.yml(PyPI/TestPyPI via Trusted Publishing). - Six vendored skills under
.claude/skills/(cicd, communicate, version-bump, run-tests, sonarclaude, doc-test-alignment), provenance: steward.