OpenCode
sst/opencode is an
OpenAI-compatible TUI that talks
/v1/chat/completions. It uses the AI SDK's
openai-compatible provider, so you point it at
rapid-mlx's base URL and any served alias just works.
Quick start
Needs: OpenCode (brew install sst/tap/opencode) and a running server (rapid-mlx serve qwen3-coder-30b-4bit).
$ rapid-mlx agents opencode --setup $ opencode run "say hello" # a reply = it works
A reply means OpenCode is using your local model. rapid-mlx agents opencode --test runs the full integration check, including a real headless task.
/v1/chat/completions ·
Setup: rapid-mlx agents opencode --setup ·
Matrix cell:
✅ ✅ XFAIL(arch) ✅
(DeepSeek R1-Distill tool-emission gap — see
XFAIL arch).
Install
$ brew install sst/tap/opencode # or $ npm install -g opencode-ai
Config
One-shot, with a server running: rapid-mlx agents opencode --setup
writes ~/.config/opencode/opencode.json, filling in the
model the server reports. Setup checks the server first and refuses
if it isn't reachable (--no-check writes the config
offline). An existing file is merged: your other keys are kept, no
backup is made, and --dry-run previews without writing.
By hand, the template is:
// ~/.config/opencode/opencode.json { "$schema": "https://opencode.ai/config.json", "provider": { "rapid-mlx": { "npm": "@ai-sdk/openai-compatible", "name": "Rapid-MLX", "options": { "baseURL": "http://localhost:8000/v1", "apiKey": "not-needed" }, "models": { "default": {} } } }, "model": "rapid-mlx/default" }
Auth. A bare rapid-mlx serve needs no key, so
apiKey is a placeholder. For a keyed server (--api-key,
RAPID_MLX_API_KEY, or the desktop app), export
RAPID_MLX_API_KEY before running setup: it writes
"apiKey": "{env:RAPID_MLX_API_KEY}", which OpenCode reads
from the environment, so the key is not stored in the file.
Run
$ rapid-mlx serve qwen3-coder-30b-4bit $ opencode # TUI $ opencode run "add a --verbose flag to cli.py" # headless one-shot $ opencode run --continue "now document it" # follow-up
Re-verified 2026-10-02 — opencode 1.18.34
passed all six tasks of our agent harness plus a follow-up turn with
opencode run and this config, against
qwen3.6-35b-4bit on rapid-mlx 0.15.4 (M4 Pro, 48 GB).
Recommended aliases (per rapid-mlx agents opencode):
qwen3-coder-30b-4bit, qwen3.5-9b-4bit,
qwen3.6-35b-4bit, qwen3.5-4b-4bit.
Gotchas
-
Config schema drifts across releases. If the template
doesn't load, run
opencode --helpand check~/.config/opencode/opencode.jsonagainst opencode.ai docs. -
Headless mode.
opencode run "<task>"runs one task without the TUI, andopencode run --continue "<task>"continues the last session in the current directory.rapid-mlx agents opencode --testuses it to run a real end-to-end task. - First-run Anthropic key prompt. Pick "skip" — rapid-mlx is providing the model through the openai-compatible provider, not Anthropic.
- DeepSeek R1-Distill XFAILs on tool calling. R1 was reasoning-only post-trained; the distilled Qwen 32B doesn't emit function calls. Text-only OpenCode chat works; agent workflows need a tool-trained family (Qwen 3.5/3.6, Gemma 4, gpt-oss).
Empirical
The opencode row of the integration matrix
is ✅ on Qwen 3.6, Gemma 4, and gpt-oss; DeepSeek R1-Distill
XFAIL(arch) per the shared
R1-Distill tool-emission
gap. Source of truth for this page:
TestOpenCode
+ rapid-mlx agents opencode.