Changelog / release

0.13.3 — GLM-5.3-Flash, and vision chats that stop recomputing

Released 2026-09-01 · full changelog · GitHub releases

Upgrade: pip install -U rapid-mlx  ·  brew upgrade rapid-mlx  ·  or grab the desktop app.

A 320B-parameter model now answers on a desktop: glm5.3-flash-4bit runs through the CLI, the OpenAI-compatible server and the Mac app. Vision models stop recomputing the text they have already seen, finished video jobs survive a restart, and the app gains a command palette.