Model pages · hot models · rapid-mlx 0.15.4

deepseek-v41-flash-reap-2bit

Run DeepSeek V4.1 Flash (REAP) on a Mac

DeepSeek's V4.1 Flash compressed into a 2-bit REAP build with a narrow DSpark speculative sidecar — the 256 GB lane. Deliberately constrained: greedy generation only, requests serialize, no tools or images. #3 trending on Hugging Face when it landed (2026-09-10).

MoE · REAP 2-bit + DSpark K4 sidecar256 GB Mac Studio · frontier DeepSeek on one Mac

Measured decode
19.4tok/s
on Mac Studio M3 Ultra · 256 GB · rapid-mlx 0.15.0 · 2026-09
Weights
212.9 GB
Peak MLX memory
218.2 GB

One command

rapid-mlx serve deepseek-v41-flash-reap-2bit

The sidecar is a separate pinned download: rapid-mlx pull deepseek-v41-flash-reap-2bit first.

Weights (212.9 GB) download on first run; you get an OpenAI-compatible endpoint at http://localhost:8000/v1. No rapid-mlx yet? It's one line — curl -fsSL https://rapidmlx.com/install.sh | bash — or take the desktop app.

Point a client at it. OpenAI SDK: base_url="http://localhost:8000/v1", model="deepseek-v41-flash-reap-2bit". Claude Code: ANTHROPIC_BASE_URL=http://localhost:8000 (no /v1 suffix) — see the Claude Code guide.

Measured on real hardware

Headline machine · Mac Studio M3 Ultra · 256 GB

Decode
19.4tok/s
First token
—
Peak MLX memory
218.2 GB
Weights on disk
212.9 GB

All measured machines

MachineDecodeFirst tokenPeak memoryCold bootWeights
Mac Studio M3 Ultra · 256 GB rapid-mlx 0.15.0 · 2026-09 engine model reference19.4 tok/s with the K4 sidecar (9.6 autoregressive)—218.2 GB—212.9 GB

Each row cites its own run; numbers are never copied across machines. Mac Studio M3 Ultra · 256 GB: rapid-mlx 0.15.0, 2026-09 — raw data + method — four-workload 128-token greedy suite, deterministic across repeats — approximately 20 tok/s for the measured configuration, not a per-prompt guarantee.

Will it fit your Mac?

Peak MLX memory
218.2 GB
Weights on disk
212.9 GB

224 GB measured memory floor — startup fails closed below it.

For the conservative install-default placement see the hardware tiers table; to compare against every model your RAM can hold, use the live picker.

Variants & profiles

  • deepseek-v4-flash-2bit — DeepSeek V4-Flash 2-bit — the V4 generation, continuous batching
  • deepseek-v4-flash-4bit — DeepSeek V4-Flash 4-bit — 1M context

FAQ

How much memory does DeepSeek V4.1 Flash (REAP) need on a Mac?

224 GB measured memory floor — startup fails closed below it.

How fast is DeepSeek V4.1 Flash (REAP) on Apple Silicon?

We measured 19.4 tok/s with the K4 sidecar (9.6 autoregressive) on a Mac Studio M3 Ultra · 256 GB, rapid-mlx 0.15.0, 2026-09; peak memory 218.2 GB on that run. Every row on this page keeps its own chip, version and date — no number is copied across machines.

How do I run DeepSeek V4.1 Flash (REAP) locally?

rapid-mlx pull deepseek-v41-flash-reap-2bit, then rapid-mlx serve deepseek-v41-flash-reap-2bit — the weights download on first run and you get an OpenAI-compatible endpoint at localhost:8000/v1 that works with Cursor, Claude Code, Aider, and any OpenAI client.

What exactly did the 0.15.0 measurement cover?

Four-workload 128-token greedy suite, deterministic across repeats — approximately 20 tok/s for the measured configuration, not a per-prompt guarantee.

Where next