Model pages · hot models · rapid-mlx 0.15.4

mimo-v2.6-flash-4bit

Run MiMo-V2.6 Flash on a Mac

Xiaomi's frontier MoE reasoner, new 2026-09-21: 309 B total, 15 B active per token, tool calling and JSON output through verified parsers. Text-only — the integration claims no vision, audio or speculative execution for it.

309 B total / 15 B active MoE · 4-bit · MTP checkpoint192 GB+ Mac Studio · frontier reasoning per active parameter

Measured decode
59 → 49tok/s
on Mac Studio M3 Ultra · 256 GB · rapid-mlx 0.15.0 · 2026-09
Weights
171.7 GB
Resident memory (174 GB peak @30K)
164.3 GB

One command

rapid-mlx serve mimo-v2.6-flash-4bit

Weights (171.7 GB) download on first run; you get an OpenAI-compatible endpoint at http://localhost:8000/v1. No rapid-mlx yet? It's one line — curl -fsSL https://rapidmlx.com/install.sh | bash — or take the desktop app.

Point a client at it. OpenAI SDK: base_url="http://localhost:8000/v1", model="mimo-v2.6-flash-4bit". Claude Code: ANTHROPIC_BASE_URL=http://localhost:8000 (no /v1 suffix) — see the Claude Code guide.

Measured on real hardware

Headline machine · Mac Studio M3 Ultra · 256 GB

Decode
59 → 49tok/s
First token
—
Resident memory (174 GB peak @30K)
164.3 GB
Weights on disk
171.7 GB

All measured machines

MachineDecodeFirst tokenPeak memoryCold bootWeights
Mac Studio M3 Ultra · 256 GB rapid-mlx 0.15.0 · 2026-09 0.15.0 release notes59 → 49 tok/s across the 4K→30K context sweep—164.3 GB—171.7 GB

Each row cites its own run; numbers are never copied across machines. Mac Studio M3 Ultra · 256 GB: rapid-mlx 0.15.0, 2026-09 — raw data + method — an exact cached repeat of a 15K document fell from 34.2 s cold to 0.95 s.

Will it fit your Mac?

Resident memory (174 GB peak @30K)
164.3 GB
Weights on disk
171.7 GB

192 GB (alias floor) — the qualified run used 164.3 GB resident and peaked at 174 GB at 30K context.

For the conservative install-default placement see the hardware tiers table; to compare against every model your RAM can hold, use the live picker.

Variants & profiles

  • glm5.3-flash-4bit — GLM-5.3 Flash — the other 192 GB+ frontier MoE
  • qwen3.8-flash-next-4bit — Qwen3.8 Flash-Next — 180B total / 6B active

FAQ

How much memory does MiMo-V2.6 Flash need on a Mac?

192 GB (alias floor) — the qualified run used 164.3 GB resident and peaked at 174 GB at 30K context.

How fast is MiMo-V2.6 Flash on Apple Silicon?

We measured 59 → 49 tok/s across the 4K→30K context sweep on a Mac Studio M3 Ultra · 256 GB, rapid-mlx 0.15.0, 2026-09; peak memory 164.3 GB on that run. Every row on this page keeps its own chip, version and date — no number is copied across machines.

How do I run MiMo-V2.6 Flash locally?

rapid-mlx serve mimo-v2.6-flash-4bit — the weights download on first run and you get an OpenAI-compatible endpoint at localhost:8000/v1 that works with Cursor, Claude Code, Aider, and any OpenAI client.

What exactly did the 0.15.0 measurement cover?

An exact cached repeat of a 15k document fell from 34.2 s cold to 0.95 s.

Where next