Qwen Code
QwenLM/qwen-code
is Alibaba's coding CLI, tuned for the Qwen 3 tool format. It
speaks the plain OpenAI wire through an openai entry in
the modelProviders block of its settings file.
Quick start
Needs: Qwen Code (npm install -g @qwen-code/qwen-code@latest) and a running server (rapid-mlx serve qwen3-coder-30b-4bit).
$ rapid-mlx agents qwen-code --setup $ qwen -p "say hello" # a reply = it works
A reply means Qwen Code is using your local model. Setup keeps your other OpenAI-compatible providers in ~/.qwen/settings.json and only adds or updates the rapid-mlx entry.
/v1/chat/completions ·
Setup: rapid-mlx agents qwen-code --setup ·
Matrix cell:
✅ ✅ XFAIL(arch) ✅
(DeepSeek R1-Distill — see
XFAIL arch).
Install
$ npm install -g @qwen-code/qwen-code@latest
Config
One-shot, with a server running: rapid-mlx agents qwen-code --setup
writes ~/.qwen/settings.json, filling in the model and
context window the server reports. It shows the exact diff, asks before
writing (--yes in scripts) and backs up an existing file.
Your other keys are kept, and in the modelProviders.openai
list only the entry with rapid-mlx's model id is added or updated;
other providers stay as they are, including their keys, which the
preview redacts. Setup sets env.RAPID_MLX_API_KEY in the
file to the template's placeholder. A file that is not valid
JSON is refused, not overwritten. --dry-run previews without writing. The
template:
// ~/.qwen/settings.json { "modelProviders": { "openai": [ { "id": "qwen3.6-35b-4bit", "name": "qwen3.6-35b-4bit (Rapid-MLX)", "envKey": "RAPID_MLX_API_KEY", "baseUrl": "http://localhost:8000/v1", "generationConfig": { "timeout": 300000, "maxRetries": 1, "contextWindowSize": 262144, "samplingParams": {"max_tokens": 8192} } } ] }, "env": {"RAPID_MLX_API_KEY": "not-needed"}, "security": {"auth": {"selectedType": "openai"}}, "model": {"name": "qwen3.6-35b-4bit"} }
Run
$ rapid-mlx serve qwen3-coder-30b-4bit $ qwen # interactive $ qwen -p "explain what main.py does" # headless one-shot
Recommended aliases (per rapid-mlx agents qwen-code):
qwen3-coder-30b-4bit, qwen3.6-35b-4bit,
qwen3.6-27b-4bit, qwen3.5-9b-4bit.
Gotchas
-
/v1suffix is hardcoded. qwen-code appends no prefix — setbaseUrltohttp://localhost:8000/v1, not the bare root. -
Older releases (< 1.2) stripped
chat_template_kwargs. Upgrade if you see<think>traces leaking intocontenton Qwen 3.6.
Empirical
The qwen-code row of the integration matrix
is ✅ on Qwen 3.6, Gemma 4, and gpt-oss; DeepSeek R1-Distill
XFAIL(arch). Source of truth for this page:
TestQwenCode
+ rapid-mlx agents qwen-code.