Model pages · hot models · rapid-mlx 0.15.4

qwen3.8-27b-4bit

Run Qwen3.8 27B on a Mac

Qwen's 27B dense reasoner and the highest-scoring open-weights model the engine serves (Artificial Analysis index 52 at full precision) — which is why the recommendation policy makes it the smart pick for every RAM tier from 32 GB up. Hybrid architecture with a native MTP head, vision input, and qwen3 thinking separation built in.

27 B dense · 4-bit · hybrid · native MTP · vision32 GB+ Macs · the engine's smart pick · reasoning · agents

Measured decode
24.1tok/s
on Mac mini M4 Pro · 48 GB · rapid-mlx 0.15.4 · 2026-10-02
Weights
16.3 GB
Peak memory · Metal
25.3 GB

One command

rapid-mlx serve qwen3.8-27b-4bit

Weights (16.3 GB) download on first run; you get an OpenAI-compatible endpoint at http://localhost:8000/v1. No rapid-mlx yet? It's one line — curl -fsSL https://rapidmlx.com/install.sh | bash — or take the desktop app.

Point a client at it. OpenAI SDK: base_url="http://localhost:8000/v1", model="qwen3.8-27b-4bit". Claude Code: ANTHROPIC_BASE_URL=http://localhost:8000 (no /v1 suffix) — see the Claude Code guide.

Measured on real hardware

Headline machine · Mac mini M4 Pro · 48 GB

Decode
24.1tok/s
First token · 2K prompt
18.8 s
Peak memory · Metal
25.3 GB
Weights on disk
16.3 GB

All measured machines

MachineDecodeFirst tokenPeak memoryCold bootWeights
Mac mini M4 Pro · 48 GB rapid-mlx 0.15.4 · 2026-10-02 raw data24.1 tok/s18.8 s25.3 GB—16.3 GB
Mac Studio M3 Ultra · 256 GB rapid-mlx 0.13.4 · 2026-09-02 engine benchmark doc43.4 tok/s24.7 s26.7 / 27.1 GB—16.3 GB

Each row cites its own run; numbers are never copied across machines. Mac mini M4 Pro · 48 GB: rapid-mlx 0.15.4, 2026-10-02 — raw data + method — 4 concurrent streams: 24.0 tok/s aggregate · Mac Studio M3 Ultra · 256 GB: rapid-mlx 0.13.4, 2026-09-02 — raw data + method.

Will it fit your Mac?

Peak memory · Metal
25.3 GB
Weights on disk
16.3 GB

32 GB Mac (engine tier floor) — measured 25.3 GB peak on a 48 GB M4 Pro through a 2K-prompt run; the M3 Ultra qualification measured ≈20 GB for the complete process tree through an 8K workload.

For the conservative install-default placement see the hardware tiers table; to compare against every model your RAM can hold, use the live picker.

Variants & profiles

  • qwen3.8-27b-tensorfold — Experimental accelerated profile — one request at a time, no tools/media/grammar
  • qwen3.5-4b-4bit — Qwen3.5 4B — the fast pick for the same tiers
  • qwen3.5-9b-4bit — Qwen3.5 9B — the 18–23 GB tier's smart pick

Tensorfold: An experimental opt-in profile (qwen3.8-27b-tensorfold) pairs this checkpoint with its DFlash2 drafter on 0.15.4. It admits one text request at a time and rejects tools, media and grammar; the ordinary qwen3.8-27b-4bit path stays the feature-complete default. The release notes make no fixed speed claim for it — neither do we.

FAQ

How much memory does Qwen3.8 27B need on a Mac?

32 GB Mac (engine tier floor) — measured 25.3 GB peak on a 48 GB M4 Pro through a 2K-prompt run; the M3 Ultra qualification measured ≈20 GB for the complete process tree through an 8K workload.

How fast is Qwen3.8 27B on Apple Silicon?

We measured 24.1 on a Mac mini M4 Pro · 48 GB, rapid-mlx 0.15.4, 2026-10-02; peak memory 25.3 GB on that run. Every row on this page keeps its own chip, version and date — no number is copied across machines.

How do I run Qwen3.8 27B locally?

rapid-mlx serve qwen3.8-27b-4bit — the weights download on first run and you get an OpenAI-compatible endpoint at localhost:8000/v1 that works with Cursor, Claude Code, Aider, and any OpenAI client.

What is the tensorfold profile for Qwen3.8 27B?

An experimental opt-in profile (qwen3.8-27b-tensorfold) pairs this checkpoint with its DFlash2 drafter on 0.15.4. It admits one text request at a time and rejects tools, media and grammar; the ordinary qwen3.8-27b-4bit path stays the feature-complete default. The release notes make no fixed speed claim for it — neither do we.

Where next