Changelog / release
0.15.6 — A guided first run and lossless prompts by default
Released 2026-10-04 · full changelog · GitHub releases
0.15.6 makes the first local-model session easier to start and long-running servers more predictable. Bare rapid-mlx opens an interactive start screen, PFlash prompt compression becomes opt-in for every model, and model setup, cache reuse, memory recovery and bring-your-own-model checks are more reliable.
- A clearer first step. Running bare
rapid-mlxin a terminal opens a compact start screen: it shows the active or last-used model when there is one, separates the quick-start model from the best model for this Mac, and waits for a keypress before chatting, starting a server, connecting a detected coding agent or opening the model picker.rapid-mlx --helpis grouped by task, and non-interactive use prints deterministic copy-and-paste commands. See the CLI reference. - Lossless prompts by default. PFlash prompt compression is off for every alias unless you pass
--pflash autoor--pflash always. When an opted-in request is compressed, supported APIs report the original and retained token counts (metrics.prompt_compressionor theX-Rapid-MLX-Prompt-Compressedheader). See PFlash. - Agent setup ready for the first task. Generated Codex, OpenCode, Claude and Pi configurations preserve explicit user settings, use the live server's model and context metadata, and work with authenticated local servers without storing secrets in the generated files.
- More reliable bring-your-own-model flows. External revision probes are bounded and respect warm-cache and offline behaviour, and a requested server log level takes effect before model-serving modules load. With telemetry enabled, BYOM preflight, suggestion, support-request and import outcomes are reported as closed categories; local paths and custom model names are never sent (see Telemetry).
- Safer memory and shared-cache reuse. Completed single-request KV state is released promptly, and an idle server re-measures reclaimed memory before refusing new work at the Metal memory limit. Valid deduplicated shared Hugging Face blob layouts pass the model-integrity checks, and a stray non-model folder named like a catalog alias no longer hides that alias.
- MTP under concurrency. A single request keeps the qualified MTP path; when requests overlap, an eligible running request yields at a token boundary and they decode together through ordinary batching.
- Clef decision models in System One. The System One service can run Clef and Clef-Flash models for text decisions, ranking, bounded images and multi-frame video, with explicit errors for unsupported media and oversized requests; the optional runtime stays out of the base install.
- Also fixed. Community Qwen3.8 checkpoints get the correct reasoning and nested tool-call parsers automatically, and failed video jobs remove partial files before reporting failure.