Changelog / release

0.13.2 — Flash-Next decodes faster, the app gets 43% smaller

Released 2026-08-30 · full changelog · GitHub releases

Upgrade: pip install -U rapid-mlx  ·  brew upgrade rapid-mlx  ·  or grab the desktop app.

Long sessions get quicker on both ends: Qwen3.8-Flash-Next can opt into its own multi-token prediction head, long prompts reach the first token roughly a third sooner, and a repeated prompt reuses its cache instead of paying for it twice. The Mac app ships as a download less than half the size it was, Kokoro is ready before you go offline, and attachments, credentials and model switching got a safety pass.