Run Qwen3 Coder 30B (A3B MoE) on a Mac
The coding specialist: a 30 B mixture-of-experts trained for code that decodes at 53 tok/s on an M2 Pro. This is the model to serve behind a local coding agent if the 35B generalist is too big.
MoE · 30 B total / 3 B active · 4-bit · code coding agents · code completion · 32 GB Macs
One command
$ rapid-mlx serve qwen3-coder-30b-4bit
Weights (16.0 GB) download on first run; you get an OpenAI-compatible
endpoint at http://localhost:8000/v1. No rapid-mlx yet? It's one line —
curl -fsSL https://rapidmlx.com/install.sh | bash — or take the
desktop app.
Measured on real hardware
| Machine | Decode | First token | Peak memory | Cold boot | Weights |
|---|---|---|---|---|---|
| Mac mini M2 Pro · 32 GB | 53.0 tok/s | 0.54 s | 10.0 GB | 11.1 s | 16.0 GB |
Measured on rapid-mlx 0.12.10, 2026-08-11. Decode: median of 3 runs, temperature 0, 256-token saturating generation, engine-reported token counts, unique salt per request (no prefix-cache hits). TTFT: median of 3, short prompt. Boot: process spawn to first completed token. Peak RSS: 0.5 s sampling across the run.
Will it fit your Mac?
Peak resident memory measured 10.0 GB during a 256-token generation. The KV cache grows with context length, so treat that as a floor, not a ceiling. For the conservative install-default placement see the hardware tiers table; to compare against every model your RAM can hold, use the live picker.
In this 256-token run, peak resident memory came in below the on-disk weight size — with mixture-of-experts weights the runtime doesn't have to touch every expert right away. Don't budget by that number: for sustained use, follow our memory guide and plan for the full weight size plus context headroom. The measured RSS here is a floor, not a plan.
Variants & alternatives
qwen3.6-35b— Qwen3.6 35B — the generalist that also codes well
FAQ
How much memory does Qwen3 Coder 30B (A3B MoE) need on a Mac?
Measured peak resident memory was 10.0 GB on rapid-mlx 0.12.10 during a 256-token generation (M2 Pro, 32 GB). The weights are 16.0 GB on disk. Longer contexts grow the KV cache beyond this, so leave headroom.
How fast is Qwen3 Coder 30B (A3B MoE) on Apple Silicon?
We measured 53.0 tokens/sec sustained decode and 0.54 s time-to-first-token on a Mac mini M2 Pro (32 GB), median of 3 runs at temperature 0.
How do I run Qwen3 Coder 30B (A3B MoE) locally?
Install rapid-mlx (curl -fsSL https://rapidmlx.com/install.sh | bash, or brew install rapid-mlx), then: rapid-mlx serve qwen3-coder-30b-4bit — the weights download on first run and you get an OpenAI-compatible endpoint at localhost:8000/v1 that works with Cursor, Claude Code, Aider, and any OpenAI client.