DeepSeek V4-Flash on Apple Silicon
DeepSeek's V4 family — Channel-Sparse Attention (CSA) plus
Hierarchical-Cluster Attention (HCA) — on a single 256 GB Mac
Studio, with the deepseek tool parser and the
deepseek_v4 reasoning parser selected for you.
Pick a quant and run it
V4-Flash needs a high-memory Mac Studio. Start with the 2-bit DQ build:
$ rapid-mlx serve deepseek-v4-flash-2bit
- 192 GB Mac (256 GB comfortable):
deepseek-v4-flash-2bit— ~97 GB of weights - 256 GB Mac Studio:
deepseek-v4-flash-4bit(~152 GB) ordeepseek-v4-flash-8bit(~155 GB) - 192 GB+ with DSpark:
deepseek-v4-flash-0731-mxfp4— the 0731 checkpoint with its built-in drafter (see below)
The endpoint is OpenAI-compatible at
http://localhost:8000/v1; reasoning comes back as a
structured field. For the V4.1 Flash REAP build see its
measured model page.
Download
Every alias on this page downloads with one command — the pull buttons in the tables below copy it. All 3 aliases on this page are mirrored on the rapid-mlx CDN — with automatic mid-pull fallback to Hugging Face if a mirror file slows down. Weights land in the standard Hugging Face cache, and rapid-mlx serve pulls automatically on first use. Live mirror status →
Why V4-Flash matters
The V4 family is DeepSeek's first generation built around CSA + HCA sparse attention — instead of full O(n²) attention, each query head attends only to a learned subset of channels (CSA) grouped into hierarchical clusters (HCA). The result is a model that scales context length without the usual quadratic KV blow-up and serves at meaningfully higher throughput than a dense V3 of comparable parameter count. V4-Flash is the smaller member of the family, sized to fit the unified-memory envelope of a workstation Mac rather than a datacenter rack.
What Rapid-MLX adds
-
Parsers. The V4 aliases use the
deepseektool-call parser and thedeepseek_v4reasoning parser, which keeps V4's internal scratch reasoning out of the user-visiblecontent. The 0731 checkpoint uses its owndeepseek_v4_0731tool parser. - Prefix-cache reuse. A repeated or extended prompt reuses the longest usable cached prefix, so warm turns skip re-prefilling the shared part of a long agent prompt.
-
DSpark speculative decoding. The 0731 checkpoint ships its
own drafter — three MTP stages plus low-rank Markov heads — and
rapid-mlx drafts and verifies with it
(
--speculative-config '{"method":"dspark","num_speculative_tokens":5}'). Detection is fail-closed against the checkpoint'sinference/config.jsonand safetensors index, so it needs the local checkpoint path rather than the alias — see perf flags. - MoE + MXFP4 guardrail. MoE models in MXFP4 can corrupt when sharded across several MLX devices (mlx#3402); rapid-mlx warns at load time instead of serving garbled tokens.
Compared with plain mlx-lm
mlx-lm loads V4
(mlx-lm #1192),
but serving it over an OpenAI-compatible API also needs the
tokenizer, tool parser and reasoning parser wired up. The rapid-mlx
aliases carry all of that.
All quants
| Alias | Bits | Approx. disk | Unified-memory floor | HF repo | get it |
|---|---|---|---|---|---|
deepseek-v4-flash-2bit |
2 (DQ) | ~97 GB | 192 GB (256 GB comfortable) | DeepSeek-V4-Flash-2bit-DQ | CDN |
deepseek-v4-flash-4bit |
4 | ~152 GB | 256 GB Mac Studio | DeepSeek-V4-Flash-4bit | CDN |
deepseek-v4-flash-8bit |
8 | ~155 GB | 256 GB Mac Studio | DeepSeek-V4-Flash-8bit | CDN |
Disk sizes are the Hugging Face repository totals. The 2-bit DQ build is the smallest quant we recommend for V4-Flash; on a 256 GB Mac Studio, start there.
Recommended settings
V4-Flash inherits the DeepSeek-R1 reasoning lineage. The alias
selects the deepseek tool parser and the
deepseek_v4 reasoning parser, so
/v1/chat/completions returns a separate
reasoning_content field and /v1/responses
returns reasoning items. The 2/4/8-bit aliases set no sampling
defaults; send the model card's values per request, or set them
for every request at startup:
rapid-mlx serve deepseek-v4-flash-2bit \
--default-temperature 0.6 --default-top-p 0.95
Tutorial — chat + tool call
curl http://127.0.0.1:8000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "deepseek-v4-flash-2bit",
"messages": [{"role": "user", "content": "What is the weather in Tokyo?"}],
"tools": [
{"type": "function", "function": {
"name": "get_weather",
"description": "Look up current weather",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"]
}
}}
]
}'
Python via the OpenAI client, using the reasoning surface:
from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="none")
resp = client.responses.create(
model="deepseek-v4-flash-2bit",
input="Plan a three-step approach to optimizing a slow SQL query.",
reasoning={"effort": "high"},
)
print(resp.output_text)
for item in resp.output:
if item.type == "reasoning":
for chunk in item.summary:
print("[reasoning]", chunk.text)
Known limitations
- 2-bit DQ is the right starting point on a 256 GB Mac Studio, but the model is still big — leave room for system memory or you'll page. Don't run another LLM next to it.
- Sparse attention pays off most at long context. At 1-2K prompt + short decode you won't see the throughput edge over a comparably-sized dense model.
- The 4-bit and 8-bit builds leave little headroom even on a 256 GB Mac — keep contexts moderate and other apps closed.