Integration matrix
Which agents and frameworks work with which model families, tested end
to end. To connect one, use its agent guide
— every supported agent has a one-line setup command
(rapid-mlx agents <name> --setup).
rapid-mlx serve, made a real tool call (or a real file
edit for the CLI agents) and checked the outcome, not just that the
API answered. The cells below were last run on rapid-mlx 0.11.3
(Mac Studio M3 Ultra, 2026-07-28) with an independent re-run on 0.11.5
(Mac mini M2 Pro 32 GB, 2026-08-01). Results for the current
release with real agent binaries are on the
agent guides page.
Agent tiers
TIER-1
claude-code (/v1/messages) ·
codex-cli (/v1/responses) ·
hermes · aider · deepseek-harness
(/v1/chat/completions). Tier-1 agents are exercised end to
end on real weights before each release, have a first-class
--setup and a maintained guide.
TIER-2 cursor (IDE chat only; its agent path runs on Cursor's backend) · opencode · qwen-code · openhands · kilo-code · copilot · droid · kimi-code · gpt-oss (raw SDK) · and the three frameworks. These are covered by wire-level smoke tests and on-demand runs.
Results at a glance
60 cells run → 49 PASS · 11 XFAIL · 0 FAIL on the M3 Ultra run.
Every XFAIL has a documented root cause (10 architectural, 1
upstream format gap). The M2 Pro re-run on 0.11.5 gave 43 PASS · 12
skipped · 1 XFAIL · 0 FAIL: the skips were OpenHands without a Docker
daemon (4) and the DeepSeek R1-Distill model declining to emit tool
calls (8), and it used the 4- and 8-bit aliases a 32 GB machine
can hold (qwen3.5-9b-4bit, gemma-4-12b-4bit,
gpt-oss-20b-mxfp4-q8, deepseek-r1-32b-4bit).
The suite also has Hunyuan 3 (Hy3) and Muse Glimmer columns, not shown here: Hy3 needs a 192 GB+ Mac and runs only in the weekly large-Mac job, and the Muse cells were not part of the runs above.
Cheap-alias policy
Each family column boots one representative alias — the smallest that still exercises the family's tool-call and reasoning parsers:
| Family | Boot alias | Why this alias |
|---|---|---|
| Qwen 3.6 | qwen3.6-35b-8bit (MoE, 3 B active) |
The Qwen3.6 flagship MoE. |
| Gemma 4 | gemma-4-31b-4bit (dense) |
Dense Gemma 4, text-only path. |
| DeepSeek | deepseek-r1-32b-4bit (R1-distilled Qwen 32B) |
About 16 GB; exercises the DeepSeek reasoning parser. Every V4-Flash build needs a 192–256 GB Mac, so it is not in the matrix. |
| gpt-oss | gpt-oss-120b-mxfp4-q8 (MoE) |
Harmony tool and reasoning parsers. |
Agent × Family matrix
11 agents × 4 families = 44 cells. ✅ = passed; XFAIL = a known, root-caused failure (see below).
| Agent | Qwen 3.6 | Gemma 4 | DeepSeek | gpt-oss |
|---|---|---|---|---|
codex-cli T1 /v1/responses |
✅ | ✅ | ✅ | ✅ |
claude-code T1 /v1/messages · 5-min guide |
✅ | ✅ | ✅ | ✅ |
| opencode | ✅ | ✅ | XFAIL (arch) | ✅ |
| qwen-code | ✅ | ✅ | XFAIL (arch) | ✅ |
| openhands Docker E2E | ✅ | ✅ | ✅ | XFAIL (format) |
| hermes-agent T1 | ✅ | ✅ | XFAIL (arch) | ✅ |
| aider T1 real bash-CLI harness | ✅ | ✅ | ✅ | ✅ |
| kilo-code | ✅ | ✅ | XFAIL (arch) | ✅ |
| deepseek-harness T1 | ✅ | ✅ | XFAIL (arch) | ✅ |
| copilot wire smoke | ✅ | ✅ | XFAIL (arch) | ✅ |
| droid wire smoke | ✅ | ✅ | XFAIL (arch) | ✅ |
| kimi-code wire smoke | ✅ | ✅ | XFAIL (arch) | ✅ |
Framework × Family matrix
3 frameworks × 4 families = 12 cells.
| Framework | Qwen 3.6 | Gemma 4 | DeepSeek | gpt-oss |
|---|---|---|---|---|
| LangChain + LangGraph | ✅ | ✅ | XFAIL (arch) | ✅ |
| PydanticAI | ✅ | ✅ | XFAIL (arch) | ✅ |
| smolagents | ✅ | ✅ | ✅ | ✅ |
Why the XFAIL cells fail
XFAIL (arch) — DeepSeek R1-Distill does not emit tool calls
The DeepSeek column boots deepseek-r1-32b-4bit, a
reasoning distill of Qwen 32B. R1's post-training was reasoning-only
(arXiv 2501.12948 §2.3.3), and the distill lost the base model's
tool-calling: it answers "I cannot provide the current weather in Tokyo
as I cannot access the get_weather tool." deterministically at 4-bit
and 8-bit. It is a model limit, not a parser bug — the cells that do
not need OpenAI tool_calls (Codex CLI, Claude Code,
smolagents' code execution, and Aider / OpenHands, which parse text
actions) pass on the same server.
XFAIL (format) — gpt-oss with OpenHands
gpt-oss's native Harmony output never emits the
<execute_bash> / <execute_ipython>
text-action tags that OpenHands' CodeActAgent parses, so this one cell
fails on the OpenHands side — tracked upstream at
All-Hands-AI/OpenHands #15167.
Reproduce locally
From an engine checkout, boot one family's alias, then run its shard:
$ rapid-mlx serve qwen3.6-35b-8bit # strict, one family per booted server (the intended workflow) $ RAPID_MLX_MATRIX_STRICT=1 RAPID_MLX_AGENT_MATRIX_FAMILY=qwen36 \ pytest tests/integrations/test_agents_matrix.py
Or test one agent against your own running server:
rapid-mlx agents <name> --test.
Environment overrides
| Variable | Default | Purpose |
|---|---|---|
RAPID_MLX_BASE_URL | http://localhost:8000/v1 | Where matrix clients point. |
RAPID_MLX_AGENT_MATRIX_FAMILY | (all) | Restrict to qwen36 / gemma4 / deepseek / gptoss / hy3 / muse. |
RAPID_MLX_MATRIX_STRICT | 0 | If 1, a missing server fails instead of skipping. |