qwen3.8-27b-4bit
Run Qwen3.8 27B on a Mac
Qwen's 27B dense reasoner and the highest-scoring open-weights model the engine serves (Artificial Analysis index 52 at full precision) — which is why the recommendation policy makes it the smart pick for every RAM tier from 32 GB up. Hybrid architecture with a native MTP head, vision input, and qwen3 thinking separation built in.
27 B dense · 4-bit · hybrid · native MTP · vision32 GB+ Macs · the engine's smart pick · reasoning · agents
One command
rapid-mlx serve qwen3.8-27b-4bit
Weights (16.3 GB) download on first run; you get an OpenAI-compatible
endpoint at http://localhost:8000/v1. No rapid-mlx yet? It's one line —
curl -fsSL https://rapidmlx.com/install.sh | bash — or take the
desktop app.
Point a client at it. OpenAI SDK: base_url="http://localhost:8000/v1",
model="qwen3.8-27b-4bit". Claude Code:
ANTHROPIC_BASE_URL=http://localhost:8000 (no /v1 suffix) —
see the Claude Code guide.
Measured on real hardware
Headline machine · Mac mini M4 Pro · 48 GB
All measured machines
| Machine | Decode | First token | Peak memory | Cold boot | Weights |
|---|---|---|---|---|---|
| Mac mini M4 Pro · 48 GB rapid-mlx 0.15.4 · 2026-10-02 raw data | 24.1 tok/s | 18.8 s | 25.3 GB | — | 16.3 GB |
| Mac Studio M3 Ultra · 256 GB rapid-mlx 0.13.4 · 2026-09-02 engine benchmark doc | 43.4 tok/s | 24.7 s | 26.7 / 27.1 GB | — | 16.3 GB |
Each row cites its own run; numbers are never copied across machines. Mac mini M4 Pro · 48 GB: rapid-mlx 0.15.4, 2026-10-02 — raw data + method — 4 concurrent streams: 24.0 tok/s aggregate · Mac Studio M3 Ultra · 256 GB: rapid-mlx 0.13.4, 2026-09-02 — raw data + method.
Will it fit your Mac?
32 GB Mac (engine tier floor) — measured 25.3 GB peak on a 48 GB M4 Pro through a 2K-prompt run; the M3 Ultra qualification measured ≈20 GB for the complete process tree through an 8K workload.
For the conservative install-default placement see the hardware tiers table; to compare against every model your RAM can hold, use the live picker.
Variants & profiles
qwen3.8-27b-tensorfold— Experimental accelerated profile — one request at a time, no tools/media/grammarqwen3.5-4b-4bit— Qwen3.5 4B — the fast pick for the same tiersqwen3.5-9b-4bit— Qwen3.5 9B — the 18–23 GB tier's smart pick
Tensorfold: An experimental opt-in profile (qwen3.8-27b-tensorfold) pairs this checkpoint with its DFlash2 drafter on 0.15.4. It admits one text request at a time and rejects tools, media and grammar; the ordinary qwen3.8-27b-4bit path stays the feature-complete default. The release notes make no fixed speed claim for it — neither do we.
FAQ
How much memory does Qwen3.8 27B need on a Mac?
32 GB Mac (engine tier floor) — measured 25.3 GB peak on a 48 GB M4 Pro through a 2K-prompt run; the M3 Ultra qualification measured ≈20 GB for the complete process tree through an 8K workload.
How fast is Qwen3.8 27B on Apple Silicon?
We measured 24.1 on a Mac mini M4 Pro · 48 GB, rapid-mlx 0.15.4, 2026-10-02; peak memory 25.3 GB on that run. Every row on this page keeps its own chip, version and date — no number is copied across machines.
How do I run Qwen3.8 27B locally?
rapid-mlx serve qwen3.8-27b-4bit — the weights download on first run and you get an OpenAI-compatible endpoint at localhost:8000/v1 that works with Cursor, Claude Code, Aider, and any OpenAI client.
What is the tensorfold profile for Qwen3.8 27B?
An experimental opt-in profile (qwen3.8-27b-tensorfold) pairs this checkpoint with its DFlash2 drafter on 0.15.4. It admits one text request at a time and rejects tools, media and grammar; the ordinary qwen3.8-27b-4bit path stays the feature-complete default. The release notes make no fixed speed claim for it — neither do we.