Model pages · rapid-mlx 0.15.3

qwen3.6-35b

Run Qwen3.6 35B (A3B MoE) on a Mac

The flagship local agent model: 35 B of capability decoding at 60 tok/s on an M2 Pro because only ~3 B parameters are active per token. This is the model our release gate drives real coding agents with.

MoE · 35 B total / 3 B active · 4-bitcoding agents · the 32 GB Mac flagship · Tier-1 family

Alibaba 256K context Base model intelligence 32.1 AA

Measured decode
59.6tok/s
on Mac mini M2 Pro · 32 GB
Weights
19.1 GB
Peak memory
9.5 GB

One command

rapid-mlx serve qwen3.6-35b

Weights (19.1 GB) download on first run; you get an OpenAI-compatible endpoint at http://localhost:8000/v1. No rapid-mlx yet? It's one line — curl -fsSL https://rapidmlx.com/install.sh | bash — or take the desktop app.

Measured on real hardware

Headline machine · Mac mini M2 Pro · 32 GB

Decode
59.6tok/s
First token
0.55s
Peak memory
9.5GB
Cold boot
14.2s

All measured machines

MachineDecodeFirst tokenPeak memoryCold bootWeights
Mac mini M2 Pro · 32 GB59.6 tok/s0.55 s9.5 GB14.2 s19.1 GB

Measured on rapid-mlx 0.12.10, 2026-08-11. Decode: median of 3 runs, temperature 0, 256-token saturating generation, engine-reported token counts, unique salt per request (no prefix-cache hits). TTFT: median of 3, short prompt. Boot: process spawn to first completed token. Peak RSS: 0.5 s sampling across the run.

Will it fit your Mac?

Measured peak memory
9.5 GB
Weights on disk
19.1 GB

Peak resident memory measured 9.5 GB during a 256-token generation. The KV cache grows with context length, so treat that as a floor, not a ceiling. For the conservative install-default placement see the hardware tiers table; to compare against every model your RAM can hold, use the live picker.

In this 256-token run, peak resident memory came in below the on-disk weight size — with mixture-of-experts weights the runtime doesn't have to touch every expert right away. Don't budget by that number: for sustained use, follow our memory guide and plan for the full weight size plus context headroom. The measured RSS here is a floor, not a plan.

Agents: This exact 4-bit build drove Codex CLI's apply_patch edit format first-try in our end-to-end tests, and the 8-bit build is what our Tier-1 agent release gate runs Claude Code, Codex, Hermes and Aider against on every release.

Variants & alternatives

  • qwen3.6-35b-mxfp4 — The MXFP4 variant
  • qwen3-coder-30b-4bit — Qwen3 Coder 30B — the code specialist

FAQ

How much memory does Qwen3.6 35B (A3B MoE) need on a Mac?

Measured peak resident memory was 9.5 GB on rapid-mlx 0.12.10 during a 256-token generation (M2 Pro, 32 GB). The weights are 19.1 GB on disk. Longer contexts grow the KV cache beyond this, so leave headroom.

How fast is Qwen3.6 35B (A3B MoE) on Apple Silicon?

We measured 59.6 tokens/sec sustained decode and 0.55 s time-to-first-token on a Mac mini M2 Pro (32 GB), median of 3 runs at temperature 0.

How do I run Qwen3.6 35B (A3B MoE) locally?

Install rapid-mlx (curl -fsSL https://rapidmlx.com/install.sh | bash, or brew install rapid-mlx), then: rapid-mlx serve qwen3.6-35b — the weights download on first run and you get an OpenAI-compatible endpoint at localhost:8000/v1 that works with Cursor, Claude Code, Aider, and any OpenAI client.

Where next