Changelog / release

0.11.6 — DeepSeek V4 gets its own drafter, and a guard against runaway output

Released 2026-08-02 · full changelog · GitHub releases

Upgrade: pip install -U rapid-mlx  ·  brew upgrade rapid-mlx  ·  or grab the desktop app.

Four DeepSeek V4 changes that compound: the checkpoint's native speculative-decoding heads are finally used, its reasoning is parsed by a parser that understands it, a repeated prompt stops re-prefilling from scratch — and a model that falls into a loop is stopped instead of running until the client gives up.

DeepSeek V4 Flash 0731 · MXFP4 · 8,131-token prompt, 124 completion tokensTTFTDecode
Cold25.830 s25.81–26.11 tok/s
Warm — before~24.6 s
Warm — after0.318 s (81.2× vs cold)25.81–26.11 tok/s