Attach third-party model providers to Grok Build (xAI's grok CLI) — GLM/Z.ai, DeepSeek,
opencode Zen, Alibaba, any OpenAI-compatible endpoint — without hand-editing TOML, and without
copying a single API key into a config file.
English · Tiếng Việt · 中文
$ grok-connect list
provider key protocol status base_url
deepseek api chat_completions attached https://api.deepseek.com
opencode-go api chat_completions attached https://opencode.ai/zen/go/v1
zai-coding-plan api chat_completions ready https://api.z.ai/api/coding/paas/v4
$ grok-connect add zai-coding-plan glm-5.2 glm-4.7
attached 'zai-coding-plan' to /Users/me/.grok/config.toml
Calling the endpoint:
ok glm-5_2 HTTP 200 works
ok glm-4_7 HTTP 200 works
All working. Use: grok -m <model> · or /model inside the TUI
Type /connect and the agent runs grok-connect list for you, reads the table back, and asks
which provider to attach — it never picks for you. (The skill ships in English, Vietnamese and
Chinese; the screenshot shows the Vietnamese build.)
Note what it says at the bottom: only a successful HTTP status counts as proof. Appearing in
the registry or in /model does not mean the endpoint answers. New models show up from the next
grok session on.
Grok Build is already a multi-provider client — its sampler speaks three protocols
(chat_completions, responses, Anthropic messages) and [model.*] blocks let you point at
any endpoint. What it does not have is a way to use that:
grok loginonly authenticates with xAI (OAuth · device code · OIDC · an external auth binary).- There is no
/connect, no provider picker, no "paste your API key" screen. - Adding a provider means writing three TOML blocks by hand and getting every field right.
grok-connect is that missing surface. It reads the models.dev registry for
base URLs and model catalogues, resolves your key, writes the config, and calls the endpoint to
prove the thing actually works.
| Piece | What it does |
|---|---|
grok-connect |
CLI: list · models · add · test · remove. Writes/removes the config blocks. |
grok-cred |
Credential helper wired into grok's [auth_provider.*]. Resolves keys at call time so no secret is ever written into config.toml. |
/connect skill |
The CLI, surfaced as a slash command inside the grok TUI. Ships in English, Vietnamese and Chinese. |
grok-skills-prune |
Trims grok's skill catalogue down to what it actually uses — see the section below. |
chatgpt-responses-shim |
Optional, not installed by default. Bridges grok to a ChatGPT subscription through an unofficial endpoint — see the section below. |
- Grok Build ≥ 1.0.5 (
grok --version) - Python 3.9+
- Keys from one of:
- opencode's credential store (
~/.local/share/opencode/auth.json) — used automatically, nothing to configure - environment variables — the names models.dev declares (
DEEPSEEK_API_KEY,ZHIPU_API_KEY,OPENCODE_API_KEY, …) orGROK_CRED_<PROVIDER>
- opencode's credential store (
git clone <this-repo> && cd grok-connect
./install.sh # scripts into ~/.local/bin, /connect skill into every ~/.grok*
SKILL_LANG=vi ./install.sh # Vietnamese skill (also: zh)install.sh finds every grok home on the machine (~/.grok, ~/.grok-100, …) and installs the
skill into each. Set GROK_HOME to target just one.
grok-connect list # who has a key, what's attached
grok-connect models zai-coding-plan # model catalogue for a provider
grok-connect add zai-coding-plan glm-5.2 # attach (no model = first three)
grok-connect test zai-coding-plan # re-verify any time
grok-connect remove zai-coding-plan # detachOr inside grok, type /connect and let the agent drive it.
New models appear in the next session — then /model <name> or grok -m <name>.
Three blocks per provider. Note there is no key anywhere:
[auth_provider.zai-coding-plan]
command = "grok-cred zai-coding-plan"
[model_providers.zai-coding-plan]
base_url = "https://api.z.ai/api/coding/paas/v4"
api_backend = "chat_completions"
auth_provider = "zai-coding-plan"
[model.glm-5_2]
model = "glm-5.2"
model_provider = "zai-coding-plan"
name = "glm-5.2 (zai-coding-plan)"
context_window = 1000000Grok runs grok-cred before any turn that needs a token, caches the result in memory, and re-runs
it on expiry or rejection. Tokens never touch disk. A model backed by an auth provider is strictly
BYOK: your xAI session token is never sent to a third-party endpoint.
grok models only reads the config file. A model with a broken base_url still shows up in that
list; you find out at chat time, with a 404. add and test make a real HTTP call per model:
| HTTP | Meaning |
|---|---|
| 200 | works |
| 401 / 403 | key wrong or expired · key not entitled to this model |
| 402 | out of balance — config is right, top up and it runs |
| 404 | wrong base_url or wrong model name |
| 429 | out of quota / rate limited — config is right, wait for reset |
That split matters: 402 and 429 are account problems, 404 is your mistake.
- Don't append
/v1blindly. models.dev already records the full base. Z.ai is…/api/coding/paas/v4; adding/v1yields a 404. Only a bare host (https://api.deepseek.com) needs it. This tool gets it right — but if you hand-edit, this is the trap. - A duplicate table kills the whole config — silently. TOML forbids declaring
[skills]or[plugins]twice, and grok responds to an invalid config by discarding the entire file — your model providers included. Nothing is printed; onlygrok inspect --jsonshowsconfigSourceswith"note": "parse error".grok-connectrefuses to write rather than produce that, but if you hand-edit, merge into the existing table instead of appending a second one. - Grok rewrites
config.toml. Switching models in the TUI makes grok normalise the file and strip every comment. It has also been seen rewriting[models] defaultto the last model used. Re-check that line after playing with the model picker. - A probe without a real
User-Agentgets 403. The WAF in front ofopencode.airejectsPython-urllib, which reads as "key not entitled" when the key is fine.grok-connectsends a UA; remember it if you roll your own check. - Anthropic-protocol providers are not supported yet. MiniMax's coding plan and friends
authenticate with
x-api-key, andgrok-credonly mintsAuthorization: Bearer. - Providers with their own SDK can't be attached (Google, Anthropic direct, xAI): models.dev lists no REST base URL for them, so grok has nothing to call.
Expand for the setup and the two protocol mismatches it works around.
You can point grok at a ChatGPT Plus/Pro subscription instead of an OpenAI API key. It works,
tool calling included, the same way opencode and other CLIs offer ChatGPT sign-in. It routes
through chatgpt.com/backend-api/codex.
That is an internal endpoint, not a public API. It carries no stability guarantee: the request or response shape can change without notice and this shim will break when it does. Whether third-party use fits your agreement with OpenAI is your call, not this README's. For production work an OpenAI Platform key (
base_url = "https://api.openai.com/v1",api_backend = "responses") is the path that is documented, versioned, and supported.
./install.sh --with-chatgpt
codex login # or sign in through opencode; the shim reuses that token
grok-connect add chatgptTwo protocol mismatches had to be bridged, and both are worth knowing if you ever wire a client to a Responses-style backend:
-
400 System messages are not allowed.Grok serializes the system prompt as an input item withrole: "system". The Codex backend only accepts it in the top-levelinstructionsfield. The shim hoists it. -
One question, four identical requests. With
store: falsethe backend returnsresponse.completedcarryingoutput: []— the text exists only in the streamed deltas. Grok builds its assistant message from that final object, sees an empty response, classifies it as a retryable server error, and re-sends the same payload until the retry budget runs out. The shim collects items fromresponse.output_item.doneand splices them back intooutput. -
serialization error: unknown variant 'keepalive'. During long thinking the backend emits an SSEkeepaliveheartbeat. Grok parses the stream with a closed enum, so an unknown event is a non-retryable error that kills the turn outright. The shim forwards onlyresponse.*events (pluserror) and drops transport noise.
The model name must be one the ChatGPT plan allows (gpt-5.6-sol works; gpt-5 is rejected).
grok-cred starts the shim on demand, so there is nothing to run by hand. Expect one extra
upstream request per session: grok tries HTTP/2 first and falls back to HTTP/1.1.
Questions, a provider that won't attach, or a setup worth sharing — bring it to OpenCode Vietnam — a Facebook group for people wiring AI coding CLIs (opencode, Grok Build, Codex, Claude Code) into real work. Vietnamese-first, English welcome.
Bug reports and pull requests go in the issue tracker; the group is for the messier "has anyone got X talking to Y" conversations.
Not about providers — but it ships here because it fixes the other way a grok session dies.
Grok collects skills from ~/.grok*/skills, ~/.agents/skills, ~/.claude/skills and plugins
(Claude Code compatibility, no off switch), lists all of them in one <system-reminder>, and
re-injects that list on every model call — keeping every copy in the transcript.
Measured on a machine with a large skill library:
460 skills = 145 KB = ~37,000 tokens per model call
72 copies in one session = 10.4 MB = 98% of the transcript
three sessions passed 3M tokens and could not recover
Auto-compact cannot rescue that: the compaction request carries the same history, so it is rejected too. And a third-party provider reports the overflow inside the SSE stream with no HTTP status, so grok reads it as retryable and re-sends the identical oversized payload 15 times — about nine minutes of looking frozen before the turn fails.
grok-skills-prune # show what would be kept/hidden, write nothing
grok-skills-prune --apply # write [skills] ignore into every grok homeIt keeps the skills grok has actually opened (<skill>/SKILL.md reads in session history — grok
has no skill tool, reading the file is how it loads one) and hides the rest. No skill file is
moved or edited; only [skills] ignore is written, so your skill store stays the source of truth.
Optional usage registry. Grok's own history is narrow if grok is new to your setup. Point the tool at a table that already tracks skill usage across your tools:
// ~/.config/grok-connect/registry.json
{"base_token": "…", "table_id": "tbl…", "keep_status": "Frequently used",
"name_field": "Skill", "status_field": "Status"}Needs lark-cli. Missing file, missing CLI, or an unreachable Base — the tool silently falls back
to grok's history. Re-run after adding skills.
MIT. Not affiliated with xAI, OpenAI, Z.ai, DeepSeek, or opencode.

