Every OpenAI-compatible endpoint lobes serves, and how the gateway routes it.
All endpoints are served on a single port (default :8000, set by VLLM_PORT
in the deployment .env). In single-model mode that port is the raw vLLM container
itself; in fleet mode it is the model-gear-gateway container, which routes by the
request's model field. Clients point at the same URL either way.
| Endpoint | Method | Backend | Notes |
|---|---|---|---|
/v1/chat/completions |
POST | generate primary (opt-in fallback) | routed by model field; unknown/missing → default primary. SSE when "stream": true. |
/v1/completions |
POST | generate primary (opt-in fallback) | same routing as chat/completions |
/v1/embeddings |
POST | Qwen/Qwen3-Embedding-0.6B (warm embed gear) |
1024-dim, Matryoshka-truncatable via "dimensions" |
/v1/rerank |
POST | Qwen/Qwen3-Reranker-0.6B (warm rerank gear) |
Jina/Cohere shape, sorted best-first |
/v1/score |
POST | Qwen/Qwen3-Reranker-0.6B (same backend as rerank) |
raw cross-encoder scores, input order |
/v1/audio/transcriptions |
POST | Parakeet STT via the realtime bridge | multipart file upload → {"text": ...} |
/v1/audio/speech |
POST | Chatterbox TTS via the realtime bridge | text → audio bytes (24 kHz) |
/v1/realtime |
GET (WebSocket upgrade) | realtime bridge, tunneled through the gateway | server_vad session; base64 PCM16 mono LE JSON events both ways, 24000 Hz default / 16000 Hz accepted. Transcription-only by default; opt in with response.create for a spoken, interruptible reply on the same socket (issues #149, #151) — see below |
/v1/models |
GET | gateway | OpenAI-standard list of loaded backends (what is hot now) |
/v1/models/supported |
GET | gateway | full supported-model catalog (every gear you can switch to; each flagged loaded/default) |
/capabilities |
GET | gateway | the eight-role Colleague contract (cortex/senses/muse/worker/embedder/reranker/stt/tts) resolved to live endpoint + metadata — non-OpenAI, lobes-native |
/health |
GET | gateway | liveness |
Embeddings, rerank, score, and audio (including the /v1/realtime WebSocket
session) require the fleet overlay — they are not available when running
the raw vLLM container in single-model mode. The generate endpoints
(/v1/chat/completions, /v1/completions) and /v1/models work in both
modes.
The gateway inspects the request's model field and forwards the request to the
backend that declares that served name (or any alias listed in GATEWAY_ALIASES).
The forwarded body's model is rewritten to the backend's actual
--served-model-name so the backend accepts aliased or defaulted routes without
complaint.
A missing or unknown model value routes to GATEWAY_DEFAULT_MODEL (the primary),
so existing single-model clients — the acp vllm-local provider, plain curl
calls — keep working with no changes.
When an opt-in warm fallback is configured, a chosen generate backend that refuses the connection or returns a 5xx before any response body is retried against the other generate backend. A 4xx (client error) is returned verbatim — no failover. Once a 2xx body has started streaming, there is no retry; the client already has bytes. By default the fleet runs one generate backend (the primary), so there is no failover peer for generate; the embed/rerank gears are separate task families and are never failover targets for each other.
When the host is under swap/iowait pressure, a full-tier generate request
(main/cortex, multimodal/senses, worker, or muse) is shed with 429 Too Many Requests rather than silently degraded onto a different model — under
pressure the gateway never substitutes a cheaper or different-capability model
(issue #85). An explicit model=minor request is the servable floor and is
always served. The 429 carries:
| Field | Value |
|---|---|
Retry-After |
seconds to wait before retrying (5 by default) |
X-Lobes-Tier-Reason |
busy |
| Body | {"error": {"type": "server_busy", "code": "busy", "message": "…"}} |
This is distinct from a 502 (type: upstream_unavailable, every backend
down — do not retry): a 429 means the model is up but the box is pressured.
Clients MUST treat 429 + Retry-After as a retryable transient and back off
— the acp vllm-local provider, colleague, and generic OpenAI SDKs all do.
Send X-Lobes-Override: true to force the requested tier and be served instead of
shed (the manual escape hatch). See
docs/gateway-fleet.md
for the full policy and thresholds.
Set GATEWAY_FORCE_STRICT_TOOLS=1 in the deployment .env (unset/false by
default) to have the gateway inject "strict": true into every
tools[i].function on a /v1/chat/completions request that both (a) routes
to the cortex/primary lane and (b) carries a non-empty tools array —
before forwarding it on. strict: true arms vLLM's xgrammar structural-tag
constrained decoding, so a tool call's arguments are grammar-constrained
against the tool's own JSON schema instead of merely parsed after the model
emits them — see
docs/qwen3.6-27b-text-nvfp4-mtp.md for why
this matters for a thinking model (colleague#320: unconstrained generation
can drift off the tool-call template and get "salvaged" into a mangled call;
the served build's structural-tag call site also needed a request-aware
reasoning flag to stay compatible with thinking mode — the paired
qwen3_coder_thinking tool-parser plugin, not this knob, supplies that).
- Caller wins. A tool that already carries any
strictkey (trueORfalse) is left untouched — the gateway fills in only an absent key. Nothing else about the request changes. - Scope. Only
/v1/chat/completions, and only when the resolved backend is the primary (cortex) lane./v1/completions, embeddings, rerank, and audio are never touched; amultimodal/sensesrequest is unaffected.museis excluded too, even though it serves tool calls and declarestool_use— that omission is deliberate, not an oversight: on the muse lane the knob is inert. Measured live on the 31B (2026-07-17),strict: truenever engages xgrammar there — a schema with a regex xgrammar cannot compile is accepted with HTTP 200 rather than failing, no grammar log line is emitted, and the output matchesstrict: false. Injectingstrictwould advertise a grammar-constrained lane that isn't one. Seedocs/gemma-4-31b-nvfp4.md#tool-calling, which also records the two disproven rationales an earlier draft gave (thesupports_required_and_namedflag, which cortex's own parser shares; and an EngineCore-crash risk that did not reproduce).workeris excluded by the same primary-only scope, even though it also declarestool_use(and, uniquely among non-cortexroles,repo_action) — whetherstrictengages xgrammar on the worker lane is UNMEASURED (no live boot has happened yet, seedocs/qwen3.6-35b-a3b-nvfp4.md), so this scope stays the conservative default rather than a claim either way. The lane set islobes.gateway.server._STRICT_TOOL_LANES; widen it only with a live transcript showing strict decoding actually constrains decoding there. - Retry-without-strict fallback. If the injected request comes back with
a 4xx/5xx whose body matches a schema/grammar-compile-failure signature
(a heuristic substring list —
structural_tag,xgrammar,grammar,json_schema— pending live discovery of the real error text), the gateway retries exactly once with the original, un-injected body and relays that retry's response verbatim (success or failure) — never a second retry. The offending tool schema name(s) and an upstream error snippet are logged server-side. A failure that does not match the signature (the caller's own error, or an unrelated owner problem) is relayed as-is, with no retry. - Default off is byte-identical passthrough. With the knob unset (or on
any non-eligible request — no tools, non-primary lane, every tool already
declaring its own
strict), none of this logic runs: the request takes the exact code path it took before the knob existed.
See docs/gateway-fleet.md for where this sits in the
gateway's fleet-fronting behavior, and
docs/qwen3.6-27b-text-nvfp4-mtp.md for the
tool-parser plugin this pairs with.
"stream": true requests are relayed chunk-by-chunk with per-chunk flushing.
Normal (non-streaming) responses are buffered with Content-Length.
/v1/audio/* requests are forwarded by the gateway to the realtime bridge
(model-gear-realtime, default http://realtime:8080, configured via AUDIO_URL).
The bridge proxies:
POST /v1/audio/transcriptions→ Parakeet STT (model-gear-stt, port 9002)POST /v1/audio/speech→ Chatterbox TTS (model-gear-chatterbox, port 9000)
The audio overlay is enabled with lobes init --fleet --audio --apply. See
docs/realtime-pipeline.md for full bring-up instructions.
/v1/realtime reaches the same bridge a different way: not an HTTP forward,
but a 101-upgrade + bidirectional byte tunnel
(lobes/gateway/_realtime.py). The gateway relays the client's WebSocket
handshake to the local realtime bridge verbatim, relays the bridge's 101
back, and then pumps opaque bytes both directions until either side closes —
it never parses the WebSocket protocol itself. A plain GET with no
Upgrade: websocket header gets 426; a declared-off stt lane gets the
same 404 role_infeasible (naming hosted_by) the batch STT route
gets. Unlike the batch audio routes, this tunnel is never proxied
cross-box even when STT_PEER_PROXY is armed — the #129 proxy-lobes
forwarder is POST-only. See Realtime session
below for the session contract.
Because the tunnel is opaque and already pumps both directions in parallel,
the spoken replies #151 adds — server-to-client response.audio.delta frames
— ride it with zero gateway change: the gateway neither knows nor cares
that the session now talks back.
Auth is opt-in. Set GATEWAY_API_KEY (fallback CULTURE_VLLM_API_KEY) in
the deployment .env to require Authorization: Bearer <key> on every
data-plane request, /v1/audio/* included — unset (the default) leaves the
gateway exactly as before, no header inspected. See
Auth and exposure below for the full gated-vs-keyless
route policy and the 401 shape.
Routes to the generate primary (sakamakismile/Qwen3.6-27B-Text-NVFP4-MTP by
default) or the opt-in fallback. Supply the served model name in model, or omit
it to hit the primary.
curl -s http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "sakamakismile/Qwen3.6-27B-Text-NVFP4-MTP",
"messages": [{"role": "user", "content": "What is 17 * 23?"}]
}'Streaming:
curl -s http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "sakamakismile/Qwen3.6-27B-Text-NVFP4-MTP",
"messages": [{"role": "user", "content": "Count to five."}],
"stream": true
}'The vllm-local provider in culture.yaml points at this endpoint; an unknown
model defaults to the primary, so model: default also works when you set
VLLM_SERVED_NAME=default in .env.
POST /v1/embeddings — served by the warm Qwen/Qwen3-Embedding-0.6B gear (fleet
only). Native dimension is 1024; pass "dimensions" to Matryoshka-truncate.
Request:
{
"model": "Qwen/Qwen3-Embedding-0.6B",
"input": ["text a", "text b"]
}input accepts a string or a list of strings. "dimensions": 512 truncates to any
supported Matryoshka sub-dimension (32 / 64 / 128 / 256 / 512 / 768 / 1024); omit
to get the native 1024-dim output.
Response:
{
"object": "list",
"data": [
{"object": "embedding", "index": 0, "embedding": [/* 1024 floats */]},
{"object": "embedding", "index": 1, "embedding": [/* 1024 floats */]}
],
"model": "Qwen/Qwen3-Embedding-0.6B",
"usage": {"prompt_tokens": 4, "total_tokens": 4}
}curl -s http://localhost:8000/v1/embeddings \
-H "Content-Type: application/json" \
-d '{"model":"Qwen/Qwen3-Embedding-0.6B","input":["Hello world"]}'POST /v1/rerank — Jina/Cohere-compatible; served by the warm
Qwen/Qwen3-Reranker-0.6B gear. Results are sorted best-first; index is the
position in the original documents list.
Request:
{
"model": "Qwen/Qwen3-Reranker-0.6B",
"query": "What is the capital of France?",
"documents": ["Paris is the capital.", "Berlin is the capital.", "Rome is the capital."]
}Response:
{
"results": [
{"index": 0, "relevance_score": 0.91},
{"index": 2, "relevance_score": 0.18},
{"index": 1, "relevance_score": 0.07}
]
}curl -s http://localhost:8000/v1/rerank \
-H "Content-Type: application/json" \
-d '{
"model": "Qwen/Qwen3-Reranker-0.6B",
"query": "capital of France",
"documents": ["Paris is the capital.","Berlin is the capital."]
}'POST /v1/score — raw cross-encoder scores from the same
Qwen/Qwen3-Reranker-0.6B backend as /v1/rerank. Results are in input order
(not sorted); use /v1/rerank when you need sorted output.
Request:
{
"model": "Qwen/Qwen3-Reranker-0.6B",
"text_1": "What is the capital of France?",
"text_2": ["Paris is the capital.", "Berlin is the capital."]
}text_1 is the query string; text_2 is a string or list of strings.
Response:
{
"object": "list",
"data": [
{"index": 0, "score": 0.91},
{"index": 1, "score": 0.07}
]
}curl -s http://localhost:8000/v1/score \
-H "Content-Type: application/json" \
-d '{
"model": "Qwen/Qwen3-Reranker-0.6B",
"text_1": "capital of France",
"text_2": ["Paris is the capital.","Berlin is the capital."]
}'POST /v1/audio/transcriptions — multipart upload; served by Parakeet NeMo ASR
via the realtime bridge. Requires the --audio fleet overlay.
curl -s http://localhost:8000/v1/audio/transcriptions \
-F "file=@clip.wav" \
-F "language=en"The backend model is fixed (nvidia/parakeet-tdt-0.6b-v2), so no model field is
needed; language is accepted for forward-compatibility (Parakeet is English-only).
Response:
{"text": "the transcribed words here"}POST /v1/audio/speech — served by Chatterbox TTS via the realtime bridge.
Returns audio bytes (24 kHz). Requires the --audio fleet overlay.
curl -s http://localhost:8000/v1/audio/speech \
-H "Content-Type: application/json" \
-d '{"model":"chatterbox","input":"Hello from lobes.","voice":""}' \
-o speech.wavThe response body is raw audio — pipe to a file or an audio player. Leave voice
empty (or null) for Chatterbox's built-in default voice; set it to a .wav path on
the sidecar for zero-shot voice cloning (passed through as audio_prompt_path). See
docs/chatterbox-tts.md for the voice-cloning contract.
Audio endpoints are gated by the same opt-in bearer check as every other POST route — see Auth and exposure below.
GET /v1/realtime with an Upgrade: websocket handshake — a persistent
server_vad session that replaces separate STT batch calls with one
connection: stream audio in, receive VAD boundary and transcription events
back on the same socket, and (opt-in) a spoken reply too. Requires the
--audio fleet overlay (issues #149, #151).
Wire format: base64 PCM16 mono little-endian inside OpenAI-shaped JSON
events, both directions. Audio in is input_audio_buffer.append; audio out
is response.audio.delta. input_sample_rate defaults to 24000 Hz;
16000 Hz is also accepted — the server resamples 24 kHz down to 16 kHz
itself before feeding Silero VAD and Parakeet (both native at 16 kHz), so a
client never needs to know or match that rate. Any other rate is rejected as
an invalid session config before any audio is accepted.
Breaking change (issue #151). #149 shipped raw binary frames as the input wire. They are no longer accepted: a binary frame now yields the named
error/invalid_wire_eventevent instead of being read as audio. Migration is mechanical — where you calledws.send_binary(pcm), send{"type": "input_audio_buffer.append", "audio": base64(pcm)}as a text frame. The deployed reachy-mini-cli speaks the old wire and cannot stream until it adapts (tracked in reachy-mini-cli#115); that break is a recorded, operator-accepted decision, not a regression.
Session config is set via connect-URL query parameters, not a first message:
ws://localhost:8000/v1/realtime?input_sample_rate=16000
| Param | Default | Accepted |
|---|---|---|
input_audio_format |
pcm16 |
pcm16 only |
input_sample_rate |
24000 |
24000 or 16000 |
input_channels |
1 |
1 (mono) only |
turn_detection |
server_vad |
server_vad only |
aec_mode |
none |
none or aec |
system_prompt |
env DEFAULT_SYSTEM_PROMPT, else the built-in |
any string |
import base64, json, websocket # pip install websocket-client
ws = websocket.create_connection("ws://localhost:8000/v1/realtime?input_sample_rate=16000")
print(ws.recv()) # {"type": "session.created", ...} — confirms negotiated config
ws.send(json.dumps({ # any chunk size; the server reassembles
"type": "input_audio_buffer.append",
"audio": base64.b64encode(pcm_chunk).decode(),
}))
print(ws.recv()) # speech_started / speech_stopped / transcription.completed / errorEvents come back as JSON text frames:
session.created— sent immediately after the handshake; confirms the negotiated config (read effective defaults off it if you sent none).input_audio_buffer.speech_started/...speech_stopped— server-side Silero VAD boundaries, each carryingat_ms(32 ms-quantised audio-stream time, a different clock from the event'stimestamp_ms). A stopped turn carriesreason: "silence"(normal) orreason: "max_turn"— the max-turn cap (VAD_MAX_TURN_MS, default 30000 ms) force-committing an unusually long, never-silent turn.max_turnis not an error — noerrorevent fires on this path. Both fields reached the wire in #151; before that a client could not observe either the boundary timing or the force-commit.conversation.item.input_audio_transcription.completed— the committed turn's Parakeet transcript, on the same connection.error— a namedcode:invalid_session_config(rejected before any session exists),vad_unavailable(Silero down — distinct from ordinary silence, which emits nothing),invalid_wire_event(a malformed client frame — bad JSON, a bad base64audiofield, or a raw binary frame; the session stays open),stt_forward_failed(the Parakeet forward failed; a turn is never silently dropped), plus the conversation-onlygenerate_failed,tts_failedandresponse_timeoutbelow.
Conversation is opt-in; transcription-only is the default. A session
answers nothing until the client sends a response.create event (idempotent,
session-level). One that never sends it emits exactly the sequence above —
the transcription-only contract reachy-mini-cli depends on, which #151 leaves
unchanged in behaviour. Once armed, each committed turn is answered on the
same connection:
response.created→response.text.done→response.audio.delta× N →response.done. The reply is generated through the gateway's own/v1/chat/completions(OPENAI_BASE_URL), synthesized by Chatterbox, and streamed back as base64 PCM16 at 24 kHz — the same rate the client sends, so audio-out never resamples.- The voice lane defaults to
multimodal, notcortex(~1 s to a short reply, versus a thinking lane spending that on a trace nobody hears); override withOPENAI_MODEL. A lane this box does not host surfaces asgenerate_failedcarrying thehosted_bypeer hint — never a silent fallback. - Speaking during playback interrupts: generate and TTS are both
cancelled, the undelivered audio is never sent, and
response.interrupted(truncated: true) goes out.BARGE_IN_WINDOW_MS(default 750) guards the moment just after the machine takes the floor.BARGE_IN_MODELis threaded but unconsumed — window-only barge-in ships. - Every stage is deadlined (60 s each). On expiry the floor returns to the
caller with
response_timeout, the stage named in the message text — never a session wedged mid-response. - Per-session conversation history and system prompt live on the server, in memory only, and die with the session.
Sessions are ephemeral. There is no resume: a disconnect for any reason (idle, mid-speech, mid-transcription, mid-response) tears the session down completely, and a reconnecting client gets a brand-new session id — reconnect-and-restart is the client contract, not a bug to work around.
Reached the same way as the batch audio routes — through the gateway, tunneled to the realtime bridge (101 upgrade + byte relay, see Realtime WebSocket tunnel above) — gated by the same opt-in bearer check (a missing/wrong key is rejected before any tunnel is opened); see Auth and exposure below.
Live status. The #149 session mechanism is VALIDATED on the DGX Spark
GB10, 2026-07-21 — a full session ran through the gateway tunnel against real
Silero and real Parakeet at both wire rates (transcript:
docs/evidence/2026-07-21-accept-realtime-spark.txt).
That run predates the #151 wire change, so it proves the tunnel, the VAD and
the STT forward — not the wire the server speaks today. Everything #151
adds — the base64 event wire, the conversation surface, barge-in, the
per-stage deadlines — is DECLARED and offline-proven, not validated: per
the #108 rule no transcript for it has landed under docs/evidence/ yet.
Still unvalidated from #149 and not retired by this work: a real microphone
(the live runs used synthesized audio), the VAD-unavailable path, concurrent
sessions, and the max-turn cap. See
docs/realtime-pipeline.md
for the full contract, the wire migration in detail, and the #149 baseline
this redeems.
GET /v1/models — OpenAI-standard list of backends currently loaded in GPU memory.
In single-model mode this is one entry; in fleet mode it includes the generate
primary plus the embedding and reranker gears.
curl -s http://localhost:8000/v1/modelsGET /v1/models/supported — the full lobes supported catalog: every gear you
can switch to, each flagged loaded (in GPU memory now) and default (the primary
the gateway defaults to). This is the HTTP equivalent of lobes overview --list.
curl -s http://localhost:8000/v1/models/supportedGET /capabilities — the EIGHT first-class, Colleague-facing roles (cortex,
senses, muse, worker, embedder, reranker, stt, tts — issue #81;
worker joined as the eighth, thor-worker-lobe plan), each
resolved to
live metadata: role, model, runtime, endpoint, path, context,
quant, mtp, responsibilities, forbidden_responsibilities, ready, and
loaded. This is the discovery contract a Colleague client uses to drive the
fleet by capability instead of a hardcoded model id — lobes capabilities --json returns the identical shape over the CLI. Non-OpenAI (lobes-native), a
sibling to /status.
Generate requests are addressed by capability-tier alias — model=main|minor| multimodal|worker|muse (back-compat hard|cheap|normal), or the
Colleague-role names model=cortex|senses|muse|worker. muse (the opt-in
creative/ideation lobe, Gemma 4 31B NVFP4) and worker (the opt-in fast
ground-work DOER, Qwen3.6 35B-A3B) are each hosted only by their own
opt-in-hosting deployment shape: on a deployment that doesn't host one,
model=muse / model=worker gets an honest 404 role_infeasible (never a
silent fallback to the primary — the inverted feasibility default; see
docs/gateway-fleet.md). As of
this writing muse is DORMANT mesh-wide (no box declares its hosting shape)
and worker's hosting shape is forthcoming — both currently 404
role_infeasible everywhere.
curl -s http://localhost:8000/capabilities{
"cortex": {
"role": "cortex", "model": "sakamakismile/Qwen3.6-27B-Text-NVFP4-MTP",
"runtime": "vllm", "endpoint": "http://localhost:8000",
"path": "/v1/chat/completions", "context": 131072, "quant": "modelopt",
"mtp": true, "responsibilities": ["reasoning", "deciding", "..."],
"forbidden_responsibilities": [], "ready": true, "loaded": true
},
"senses": { "...": "..." },
"muse": { "...": "..." },
"worker": { "...": "..." },
"embedder": { "...": "..." },
"reranker": { "...": "..." },
"stt": { "...": "..." },
"tts": { "...": "..." }
}All eight roles report this one client-reachable gateway endpoint —
including stt/tts — because routing is by the model field / OpenAI path,
not by distinct URLs (issue #87). The gateway advertises the origin you dialed
(the request Host header; override with GATEWAY_PUBLIC_URL for a tunnel), so
endpoint is reachable as-is and internal hosts are never leaked. For stt/tts,
GET /capabilities reports a live readiness probe (issue #89) — ready: true
only when an audio round-trip would truly succeed. See
docs/colleague-stack.md for the full contract
(responsibilities per role, the cortex/senses↔primary/multimodal mapping,
lobes up <role>, lobes measure, and the client-flow / rename-safety proof).
Two questions that look alike but are not:
| Question | CLI | HTTP |
|---|---|---|
| What can I run? (catalog) | lobes overview --list |
GET /v1/models/supported |
| What's loaded right now? | lobes fleet status |
GET /v1/models |
| What's the deployment set to serve? | lobes status / lobes whoami |
— |
Mnemonic: the catalog is what's on the menu (and which dishes have been
cooked); GET /v1/models is what's hot now.
lobes status / lobes whoami report the model the deployment is configured to
serve (from .env) plus container health — normally the same model, but it is
configuration, not a live query. For runtime truth, query /v1/models.
See docs/gateway-fleet.md
for the full discussion.
Set CULTURE_VLLM_API_KEY in $HOME/.lobes/.env before serving. vLLM then
requires Authorization: Bearer $CULTURE_VLLM_API_KEY on every request. An empty
key leaves the API open — only safe for local development.
curl -s http://localhost:8000/v1/chat/completions \
-H "Authorization: Bearer $CULTURE_VLLM_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"default","messages":[{"role":"user","content":"hi"}]}'Generate the key with python3 scripts/gen-api-key.py (writes to the deployment
.env; --show to print, --force to rotate), then lobes serve --apply to
enforce it.
Set GATEWAY_API_KEY in the fleet .env (fallback CULTURE_VLLM_API_KEY —
the first non-blank of the two wins) to require Authorization: Bearer <key>
on the gateway's own data-plane routes, /v1/audio/* included. Unset — the
default — leaves the gateway exactly as before: no header is ever inspected.
Gated: every POST (chat/completions, completions, embeddings, rerank,
score, audio) and GET /v1/models + GET /v1/models/supported. Keyless by
design, regardless of the key: GET /health (container/peer probes must
reach it before any key is distributed), GET /capabilities (the
control-plane discovery surface peers read to learn what a box hosts, before
they hold a key), and GET /status (operator observability, no inference).
A missing/wrong/malformed key gets 401 with an OpenAI-shaped
invalid_api_key body and a WWW-Authenticate: Bearer header:
curl -s http://localhost:8000/v1/chat/completions \
-H "Authorization: Bearer $GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"cortex","messages":[{"role":"user","content":"hi"}]}'{"error": {"message": "Invalid API key. Pass this gateway's configured key as 'Authorization: Bearer <key>'.", "type": "invalid_api_key", "code": "invalid_api_key"}}The message never echoes what was sent nor any part of the expected key; the
comparison is timing-safe (hmac.compare_digest). Keeping the gateway port
off the public internet — the Cloudflare Tunnel (lobes tunnel), Cloudflare
Access, an IP allowlist, or binding to localhost and exposing only via the
tunnel — remains good practice, but with GATEWAY_API_KEY set it is now
defense-in-depth on top of the bearer gate, not the only protection.
A proxied response carries an extra header. When this gateway forwards a
request to a peer box for a role it dropped (proxy-lobes, opt-in — see
docs/gateway-fleet.md#proxy-lobes-the-third-lobe-state-opt-in),
the response carries X-Lobes-Proxied-By: <peer origin>; a locally-served
response never does. GET /v1/models only lists a proxied role's id while a
live probe of the peer's own /v1/models confirms it actually serves that
id — the same "advertised implies reachable" rule (issue #92) a local backend
is already held to, extended across the box boundary.
lobes tunnel --apply publishes the local API at an owner-chosen hostname through
a Cloudflare Tunnel — no inbound ports, no static IP required. Set
CULTURE_VLLM_API_KEY before tunnelling.
lobes tunnel # dry-run: prints the cloudflared command + public URL
lobes tunnel --apply # start the tunnel in the background
lobes tunnel --stop --apply # tear it downSee the README "Expose the API" section and lobes explain tunnel for the full
two-step provisioning flow (cultureflare + lobes tunnel).
lobes explain gateway— routing semantics (name / default / failover / SSE)lobes explain fleet— the multi-container fleet topologylobes explain roles— the eight-role Colleague contract (GET /capabilities)lobes explain embeddings—/v1/embeddingsrequest/response detaillobes explain rerank—/v1/rerankrequest/response detaillobes explain score—/v1/scorerequest/response detaillobes explain tunnel— Cloudflare Tunnel bring-uplobes explain realtime— the/v1/realtimesession surface, in-CLIdocs/gateway-fleet.md— full fleet topology, memory guidance, live validation findingsdocs/colleague-stack.md— the eight-role Colleague contract,GET /capabilitiesJSON shape,lobes up/measure/benchmark --profiledocs/realtime-pipeline.md— audio overlay bring-up (STT + TTS), the/v1/realtimesession contract, health/readiness, runbooksdocs/chatterbox-tts.md— Chatterbox TTS details, voice prompting