deepseek-v41-flash-reap-2bit
Run DeepSeek V4.1 Flash (REAP) on a Mac
DeepSeek's V4.1 Flash compressed into a 2-bit REAP build with a narrow DSpark speculative sidecar — the 256 GB lane. Deliberately constrained: greedy generation only, requests serialize, no tools or images. #3 trending on Hugging Face when it landed (2026-09-10).
MoE · REAP 2-bit + DSpark K4 sidecar256 GB Mac Studio · frontier DeepSeek on one Mac
One command
rapid-mlx serve deepseek-v41-flash-reap-2bit
The sidecar is a separate pinned download: rapid-mlx pull deepseek-v41-flash-reap-2bit first.
Weights (212.9 GB) download on first run; you get an OpenAI-compatible
endpoint at http://localhost:8000/v1. No rapid-mlx yet? It's one line —
curl -fsSL https://rapidmlx.com/install.sh | bash — or take the
desktop app.
Point a client at it. OpenAI SDK: base_url="http://localhost:8000/v1",
model="deepseek-v41-flash-reap-2bit". Claude Code:
ANTHROPIC_BASE_URL=http://localhost:8000 (no /v1 suffix) —
see the Claude Code guide.
Measured on real hardware
Headline machine · Mac Studio M3 Ultra · 256 GB
All measured machines
| Machine | Decode | First token | Peak memory | Cold boot | Weights |
|---|---|---|---|---|---|
| Mac Studio M3 Ultra · 256 GB rapid-mlx 0.15.0 · 2026-09 engine model reference | 19.4 tok/s with the K4 sidecar (9.6 autoregressive) | — | 218.2 GB | — | 212.9 GB |
Each row cites its own run; numbers are never copied across machines. Mac Studio M3 Ultra · 256 GB: rapid-mlx 0.15.0, 2026-09 — raw data + method — four-workload 128-token greedy suite, deterministic across repeats — approximately 20 tok/s for the measured configuration, not a per-prompt guarantee.
Will it fit your Mac?
224 GB measured memory floor — startup fails closed below it.
For the conservative install-default placement see the hardware tiers table; to compare against every model your RAM can hold, use the live picker.
Variants & profiles
deepseek-v4-flash-2bit— DeepSeek V4-Flash 2-bit — the V4 generation, continuous batchingdeepseek-v4-flash-4bit— DeepSeek V4-Flash 4-bit — 1M context
FAQ
How much memory does DeepSeek V4.1 Flash (REAP) need on a Mac?
224 GB measured memory floor — startup fails closed below it.
How fast is DeepSeek V4.1 Flash (REAP) on Apple Silicon?
We measured 19.4 tok/s with the K4 sidecar (9.6 autoregressive) on a Mac Studio M3 Ultra · 256 GB, rapid-mlx 0.15.0, 2026-09; peak memory 218.2 GB on that run. Every row on this page keeps its own chip, version and date — no number is copied across machines.
How do I run DeepSeek V4.1 Flash (REAP) locally?
rapid-mlx pull deepseek-v41-flash-reap-2bit, then rapid-mlx serve deepseek-v41-flash-reap-2bit — the weights download on first run and you get an OpenAI-compatible endpoint at localhost:8000/v1 that works with Cursor, Claude Code, Aider, and any OpenAI client.
What exactly did the 0.15.0 measurement cover?
Four-workload 128-token greedy suite, deterministic across repeats — approximately 20 tok/s for the measured configuration, not a per-prompt guarantee.