Hero model · rapid-mlx 0.15.7

DeepSeek V4-Flash on Apple Silicon

DeepSeek's V4 family — Channel-Sparse Attention (CSA) plus Hierarchical-Cluster Attention (HCA) — on a single 256 GB Mac Studio, with the deepseek tool parser and the deepseek_v4 reasoning parser selected for you.

Pick a quant and run it

V4-Flash needs a high-memory Mac Studio. Start with the 2-bit DQ build:

$ rapid-mlx serve deepseek-v4-flash-2bit

The endpoint is OpenAI-compatible at http://localhost:8000/v1; reasoning comes back as a structured field. For the V4.1 Flash REAP build see its measured model page.

Download

Every alias on this page downloads with one command — the pull buttons in the tables below copy it. All 3 aliases on this page are mirrored on the rapid-mlx CDN — with automatic mid-pull fallback to Hugging Face if a mirror file slows down. Weights land in the standard Hugging Face cache, and rapid-mlx serve pulls automatically on first use. Live mirror status →

Why V4-Flash matters

The V4 family is DeepSeek's first generation built around CSA + HCA sparse attention — instead of full O(n²) attention, each query head attends only to a learned subset of channels (CSA) grouped into hierarchical clusters (HCA). The result is a model that scales context length without the usual quadratic KV blow-up and serves at meaningfully higher throughput than a dense V3 of comparable parameter count. V4-Flash is the smaller member of the family, sized to fit the unified-memory envelope of a workstation Mac rather than a datacenter rack.

What Rapid-MLX adds

Compared with plain mlx-lm

mlx-lm loads V4 (mlx-lm #1192), but serving it over an OpenAI-compatible API also needs the tokenizer, tool parser and reasoning parser wired up. The rapid-mlx aliases carry all of that.

All quants

Alias Bits Approx. disk Unified-memory floor HF repo get it
deepseek-v4-flash-2bit 2 (DQ) ~97 GB 192 GB (256 GB comfortable) DeepSeek-V4-Flash-2bit-DQ CDN
deepseek-v4-flash-4bit 4 ~152 GB 256 GB Mac Studio DeepSeek-V4-Flash-4bit CDN
deepseek-v4-flash-8bit 8 ~155 GB 256 GB Mac Studio DeepSeek-V4-Flash-8bit CDN

Disk sizes are the Hugging Face repository totals. The 2-bit DQ build is the smallest quant we recommend for V4-Flash; on a 256 GB Mac Studio, start there.

Recommended settings

V4-Flash inherits the DeepSeek-R1 reasoning lineage. The alias selects the deepseek tool parser and the deepseek_v4 reasoning parser, so /v1/chat/completions returns a separate reasoning_content field and /v1/responses returns reasoning items. The 2/4/8-bit aliases set no sampling defaults; send the model card's values per request, or set them for every request at startup:

rapid-mlx serve deepseek-v4-flash-2bit \
  --default-temperature 0.6 --default-top-p 0.95

Tutorial — chat + tool call

curl http://127.0.0.1:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "deepseek-v4-flash-2bit",
    "messages": [{"role": "user", "content": "What is the weather in Tokyo?"}],
    "tools": [
      {"type": "function", "function": {
        "name": "get_weather",
        "description": "Look up current weather",
        "parameters": {
          "type": "object",
          "properties": {"city": {"type": "string"}},
          "required": ["city"]
        }
      }}
    ]
  }'

Python via the OpenAI client, using the reasoning surface:

from openai import OpenAI

client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="none")

resp = client.responses.create(
    model="deepseek-v4-flash-2bit",
    input="Plan a three-step approach to optimizing a slow SQL query.",
    reasoning={"effort": "high"},
)

print(resp.output_text)
for item in resp.output:
    if item.type == "reasoning":
        for chunk in item.summary:
            print("[reasoning]", chunk.text)

Known limitations

Related