Changelog / release

0.15.5 — Less disruptive Computer Use, earlier model checks, clearer first starts

Released 2026-10-03 · full changelog · GitHub releases

Upgrade: pip install -U rapid-mlx  ·  brew upgrade rapid-mlx  ·  or grab the desktop app.

0.15.5 makes experimental Computer Use less disruptive, reduces host overhead in qualified speculative decoding, and checks bring-your-own models before downloading them. First-start failures now name the failing stage and the fix.

  • Computer Use works in the background for eligible actions. Eligible clicks, plain text entry, and non-Command keystrokes can target the exact app window approved for the task without bringing it forward; window identity is checked immediately before input. Finder, Command shortcuts, and the automatic fallback when background routing is unavailable still use the foreground, and forced background mode refuses that fallback. The feature remains experimental and still needs Screen Recording and Accessibility permission.
  • Less host overhead during accepted MTP drafts. Qualified continuous self-MTP decoding on the Qwen and GLM speculative paths now performs one host synchronization per accepted draft cycle instead of one per proposed token. Output and profile eligibility are unchanged; the gain depends on draft acceptance and workload, and the release notes quote no fixed speedup.
  • Bring your own model, checked before download. Models outside the catalog are checked for format, architecture, and fit before anything downloads; when one cannot run, Rapid-MLX suggests a runnable alternative and can file a support request, which stays opt-in. The new rapid-mlx import command converts a compatible bf16/fp16 safetensors model to quantized MLX (--quantize 2, 3, 4, 6 or 8 bits), and a cancelled or failed import leaves the cache unchanged.
  • First starts that fail say how to fix it. Failures identify the stage that failed and give a specific remedy: draft-only MTP checkpoints are refused before download and point to the base model that uses them automatically, missing optional extras print the exact install command, and context-length rejections say how to fit the request. Text-capable vision checkpoints now fall back to the text lane on a base install instead of failing, and a new --context-length flag sets the per-request window up to the model's declared limit.
  • SillyTavern-compatible sampling. DRY sampling and repetition_penalty_range are now honoured, and samplers Rapid-MLX does not implement (such as XTC or mirostat) are refused with an HTTP 400 naming the field instead of being silently dropped.
  • More reliable local coding agents. A new pi profile (rapid-mlx agents pi --setup), dsh 0.2 setup, and headless opencode join the agent integrations. Claude Code prefix caching, integer tool arguments, and undeclared tool markup are handled more reliably.
  • Benchmarks and the leaderboard. Install, recipe, and benchmark-share output now link to the matching leaderboard views. bench --submit no longer submits anything and points to benchmark run + benchmark share; Desktop benchmark sharing follows the same workflow and refuses empty or invalid runs.
  • Desktop. Model unload is a labelled action beside the active model (Unload all for multi-model pools), with the existing guarded behaviour while work is in progress.