Agent guide · rapid-mlx 0.15.7 · ← Back to README

Qwen Code

QwenLM/qwen-code is Alibaba's coding CLI, tuned for the Qwen 3 tool format. It speaks the plain OpenAI wire through an openai entry in the modelProviders block of its settings file.

Quick start

Needs: Qwen Code (npm install -g @qwen-code/qwen-code@latest) and a running server (rapid-mlx serve qwen3-coder-30b-4bit).

$ rapid-mlx agents qwen-code --setup
$ qwen -p "say hello"               # a reply = it works

A reply means Qwen Code is using your local model. Setup keeps your other OpenAI-compatible providers in ~/.qwen/settings.json and only adds or updates the rapid-mlx entry.

Wire: /v1/chat/completions · Setup: rapid-mlx agents qwen-code --setup · Matrix cell: ✅ ✅ XFAIL(arch) ✅ (DeepSeek R1-Distill — see XFAIL arch).

Install

$ npm install -g @qwen-code/qwen-code@latest

Config

One-shot, with a server running: rapid-mlx agents qwen-code --setup writes ~/.qwen/settings.json, filling in the model and context window the server reports. It shows the exact diff, asks before writing (--yes in scripts) and backs up an existing file. Your other keys are kept, and in the modelProviders.openai list only the entry with rapid-mlx's model id is added or updated; other providers stay as they are, including their keys, which the preview redacts. Setup sets env.RAPID_MLX_API_KEY in the file to the template's placeholder. A file that is not valid JSON is refused, not overwritten. --dry-run previews without writing. The template:

// ~/.qwen/settings.json
{
  "modelProviders": {
    "openai": [
      {
        "id": "qwen3.6-35b-4bit",
        "name": "qwen3.6-35b-4bit (Rapid-MLX)",
        "envKey": "RAPID_MLX_API_KEY",
        "baseUrl": "http://localhost:8000/v1",
        "generationConfig": {
          "timeout": 300000,
          "maxRetries": 1,
          "contextWindowSize": 262144,
          "samplingParams": {"max_tokens": 8192}
        }
      }
    ]
  },
  "env": {"RAPID_MLX_API_KEY": "not-needed"},
  "security": {"auth": {"selectedType": "openai"}},
  "model": {"name": "qwen3.6-35b-4bit"}
}

Run

$ rapid-mlx serve qwen3-coder-30b-4bit
$ qwen                                   # interactive
$ qwen -p "explain what main.py does"       # headless one-shot

Recommended aliases (per rapid-mlx agents qwen-code): qwen3-coder-30b-4bit, qwen3.6-35b-4bit, qwen3.6-27b-4bit, qwen3.5-9b-4bit.

Gotchas

Empirical

The qwen-code row of the integration matrix is ✅ on Qwen 3.6, Gemma 4, and gpt-oss; DeepSeek R1-Distill XFAIL(arch). Source of truth for this page: TestQwenCode + rapid-mlx agents qwen-code.

See also