mimo-v2.6-flash-4bit
Run MiMo-V2.6 Flash on a Mac
Xiaomi's frontier MoE reasoner, new 2026-09-21: 309 B total, 15 B active per token, tool calling and JSON output through verified parsers. Text-only — the integration claims no vision, audio or speculative execution for it.
309 B total / 15 B active MoE · 4-bit · MTP checkpoint192 GB+ Mac Studio · frontier reasoning per active parameter
One command
rapid-mlx serve mimo-v2.6-flash-4bit
Weights (171.7 GB) download on first run; you get an OpenAI-compatible
endpoint at http://localhost:8000/v1. No rapid-mlx yet? It's one line —
curl -fsSL https://rapidmlx.com/install.sh | bash — or take the
desktop app.
Point a client at it. OpenAI SDK: base_url="http://localhost:8000/v1",
model="mimo-v2.6-flash-4bit". Claude Code:
ANTHROPIC_BASE_URL=http://localhost:8000 (no /v1 suffix) —
see the Claude Code guide.
Measured on real hardware
Headline machine · Mac Studio M3 Ultra · 256 GB
All measured machines
| Machine | Decode | First token | Peak memory | Cold boot | Weights |
|---|---|---|---|---|---|
| Mac Studio M3 Ultra · 256 GB rapid-mlx 0.15.0 · 2026-09 0.15.0 release notes | 59 → 49 tok/s across the 4K→30K context sweep | — | 164.3 GB | — | 171.7 GB |
Each row cites its own run; numbers are never copied across machines. Mac Studio M3 Ultra · 256 GB: rapid-mlx 0.15.0, 2026-09 — raw data + method — an exact cached repeat of a 15K document fell from 34.2 s cold to 0.95 s.
Will it fit your Mac?
192 GB (alias floor) — the qualified run used 164.3 GB resident and peaked at 174 GB at 30K context.
For the conservative install-default placement see the hardware tiers table; to compare against every model your RAM can hold, use the live picker.
Variants & profiles
glm5.3-flash-4bit— GLM-5.3 Flash — the other 192 GB+ frontier MoEqwen3.8-flash-next-4bit— Qwen3.8 Flash-Next — 180B total / 6B active
FAQ
How much memory does MiMo-V2.6 Flash need on a Mac?
192 GB (alias floor) — the qualified run used 164.3 GB resident and peaked at 174 GB at 30K context.
How fast is MiMo-V2.6 Flash on Apple Silicon?
We measured 59 → 49 tok/s across the 4K→30K context sweep on a Mac Studio M3 Ultra · 256 GB, rapid-mlx 0.15.0, 2026-09; peak memory 164.3 GB on that run. Every row on this page keeps its own chip, version and date — no number is copied across machines.
How do I run MiMo-V2.6 Flash locally?
rapid-mlx serve mimo-v2.6-flash-4bit — the weights download on first run and you get an OpenAI-compatible endpoint at localhost:8000/v1 that works with Cursor, Claude Code, Aider, and any OpenAI client.
What exactly did the 0.15.0 measurement cover?
An exact cached repeat of a 15k document fell from 34.2 s cold to 0.95 s.