Changelog
The 0.8 through 0.12 lines, newest first. Every release also has a permalink page of its own; for one-shot release notes per tag, check the GitHub releases page — this list is a curated summary of the most load-bearing changes.
Community contributors through v0.11.9
Twenty-six people contributed 36 merged pull requests. Thank you for making rapid-mlx faster, safer, easier to integrate, and better documented.
- @66Ton99 — Responses SSE heartbeats and Codex long-context handling (#1061, #1141)
- @ImL1s — no-thinking Anthropic streams for Claude Code (#213)
- @MaXoS-Agent — Qwen 3.6 community benchmark on Apple M2 Max (#1065)
- @MetricorTechnologies — prefix-cache hit accounting in usage blocks (#478)
- @MishkinBerteig — MLX-step worker request routing (#182)
- @Oxygen56 — safer rate-limit identities and the AI-client compatibility guide (#560, #561)
- @ShiroKSH — TurboQuant Metal source packaging (#1086)
- @Shreyas-Gowda26 — configurable CORS support (#48)
- @Sumu004 — benchmarking tooling (#263)
- @XiaoPengMei — server log-level control (#57)
- @angusgastle — Claude Code integration and mixed-temperature concurrent batching (#557, #619, #665)
- @dashitongzhi — documentation URL and tooling-reference cleanup (#256)
- @jedisct1 — Swival client setup documentation (#199)
- @jpcarranza94 — cross-format streaming tool-call fallback (#426)
- @kumosan2 — bounded trim-free prefix reuse for hybrid models (#1111)
- @loriz-art — README logo and banner design (#548)
- @marsmike — Anthropic system-message normalization (#976)
- @masonjames — OutputRouter migrations for Qwen, DeepSeek, GPT-OSS and streaming, plus authenticated health management (#227, #228, #229, #259, #260)
- @michaelasper — Ollama comparison benchmarking and Anthropic SDK authentication (#176, #262)
- @parth-951 — preserving HTTPException responses (#491)
- @pierre427 — visibility into pflash compression collapse (#1106)
- @romanbsd — nested Gemma 4 tool arguments (#1102)
- @samuelfaj — MLX generation-stream thread ownership (#161)
- @unsaltedbutter-ai — streaming interval accumulation and Gemma 4 parser coverage (#210, #307)
- @wparuch — Qwen 3.6 and GPT-OSS community benchmarks on Apple M2 Max (#1068)
- @wuwangzhang1216 — Codex namespace tool groups on the Responses API (#993)
0.15.7 — 2026-10-07 · EmbeddingGemma 2, two more accelerated profiles, steadier agent setup
0.15.7 adds native text and code embeddings with EmbeddingGemma 2, two experimental TensorFold profiles, and agent setup that keeps more of your existing client configuration. Servers gain switches for disk writes and log output, and responses can report per-request timing.
- EmbeddingGemma 2 text and code embeddings.
embeddinggemma-2-4bitandembeddinggemma-2-bf16serve normalized 768-dimensional vectors through/v1/embeddings, with optional smaller dimensions and an explicit 8192-token limit. Text and code inputs only; images, audio and video are refused. - Two more experimental accelerated profiles.
nemotron-3.5-lightning-tensorfold(48 GB minimum) andqwen3.8-flash-next-tensorfold(192 GB minimum) use the MTP head published with each checkpoint, serve text chat without tools, images or grammar, and leave the ordinary aliases unchanged. All four TensorFold profiles now share one pinned TensorFold runtime revision, so an existing TensorFold install needs reinstalling at that revision. No fixed speedup is claimed. - Agent setup that keeps what you have. Continue and Cline get their configuration where their clients read it, and a keyed server's credential reaches the Continue entry. Qwen Code setup keeps your other OpenAI-compatible providers, previews the exact change with their keys redacted, and backs up the file. Codex setup ends with a report of the saved defaults read back from disk. See agent setup.
- Disk writes and logs under your control.
--disable-disk-cachesstops optional cache writes (prefix-cache snapshot, KV checkpoints, vision cache disk tier) while keeping in-memory reuse,--log-filesends server output to a file or/dev/null, and the global--disable-version-checkskips release checks. Prefix-cache persistence leaves a free-disk reserve on the volume. See Disk writes. - Per-request timing. Chat Completions, Completions and Responses can report experimental
metrics.time_to_first_token_msandmetrics.mean_itl_msfor the request (see API). - Warmer coding-agent turns. A warm turn reuses the tokenized prompt instead of re-tokenizing it whole, and greedy MTP output is reproducible from run to run. Gains depend on the model, hardware and workload.
- LTX-2.5 video extension.
POST /v1/videos/extendappends frames to an existing MP4 through the usual video job API. - Model catalog.
qwopus-27b-8bitis retired because its upstream repository is gone; useqwopus-27b-4bit.qwen3.8-27b-abliterated-4bitfollows the publisher's relocated oQ4e build. - Desktop. Broader Simplified Chinese coverage, clearer timeout handling and a verified saved Codex model.
- Also fixed. An unreadable model cache on an external volume now says how to grant access instead of looking like a missing model (see Troubleshooting), and pinned-revision image, video and speculative checkpoints download through the mirror. Computer Use remains experimental.
0.15.6 — 2026-10-04 · A guided first run and lossless prompts by default
0.15.6 makes the first local-model session easier to start and long-running servers more predictable. Bare rapid-mlx opens an interactive start screen, PFlash prompt compression becomes opt-in for every model, and model setup, cache reuse, memory recovery and bring-your-own-model checks are more reliable.
- A clearer first step. Running bare
rapid-mlxin a terminal opens a compact start screen: it shows the active or last-used model when there is one, separates the quick-start model from the best model for this Mac, and waits for a keypress before chatting, starting a server, connecting a detected coding agent or opening the model picker.rapid-mlx --helpis grouped by task, and non-interactive use prints deterministic copy-and-paste commands. See the CLI reference. - Lossless prompts by default. PFlash prompt compression is off for every alias unless you pass
--pflash autoor--pflash always. When an opted-in request is compressed, supported APIs report the original and retained token counts (metrics.prompt_compressionor theX-Rapid-MLX-Prompt-Compressedheader). See PFlash. - Agent setup ready for the first task. Generated Codex, OpenCode, Claude and Pi configurations preserve explicit user settings, use the live server's model and context metadata, and work with authenticated local servers without storing secrets in the generated files.
- More reliable bring-your-own-model flows. External revision probes are bounded and respect warm-cache and offline behaviour, and a requested server log level takes effect before model-serving modules load. With telemetry enabled, BYOM preflight, suggestion, support-request and import outcomes are reported as closed categories; local paths and custom model names are never sent (see Telemetry).
- Safer memory and shared-cache reuse. Completed single-request KV state is released promptly, and an idle server re-measures reclaimed memory before refusing new work at the Metal memory limit. Valid deduplicated shared Hugging Face blob layouts pass the model-integrity checks, and a stray non-model folder named like a catalog alias no longer hides that alias.
- MTP under concurrency. A single request keeps the qualified MTP path; when requests overlap, an eligible running request yields at a token boundary and they decode together through ordinary batching.
- Clef decision models in System One. The System One service can run Clef and Clef-Flash models for text decisions, ranking, bounded images and multi-frame video, with explicit errors for unsupported media and oversized requests; the optional runtime stays out of the base install.
- Also fixed. Community Qwen3.8 checkpoints get the correct reasoning and nested tool-call parsers automatically, and failed video jobs remove partial files before reporting failure.
0.15.5 — 2026-10-03 · Less disruptive Computer Use, earlier model checks, clearer first starts
0.15.5 makes experimental Computer Use less disruptive, reduces host overhead in qualified speculative decoding, and checks bring-your-own models before downloading them. First-start failures now name the failing stage and the fix.
- Computer Use works in the background for eligible actions. Eligible clicks, plain text entry, and non-Command keystrokes can target the exact app window approved for the task without bringing it forward; window identity is checked immediately before input. Finder, Command shortcuts, and the automatic fallback when background routing is unavailable still use the foreground, and forced background mode refuses that fallback. The feature remains experimental and still needs Screen Recording and Accessibility permission.
- Less host overhead during accepted MTP drafts. Qualified continuous self-MTP decoding on the Qwen and GLM speculative paths now performs one host synchronization per accepted draft cycle instead of one per proposed token. Output and profile eligibility are unchanged; the gain depends on draft acceptance and workload, and the release notes quote no fixed speedup.
- Bring your own model, checked before download. Models outside the catalog are checked for format, architecture, and fit before anything downloads; when one cannot run, Rapid-MLX suggests a runnable alternative and can file a support request, which stays opt-in. The new
rapid-mlx importcommand converts a compatible bf16/fp16 safetensors model to quantized MLX (--quantize2, 3, 4, 6 or 8 bits), and a cancelled or failed import leaves the cache unchanged. - First starts that fail say how to fix it. Failures identify the stage that failed and give a specific remedy: draft-only MTP checkpoints are refused before download and point to the base model that uses them automatically, missing optional extras print the exact install command, and context-length rejections say how to fit the request. Text-capable vision checkpoints now fall back to the text lane on a base install instead of failing, and a new
--context-lengthflag sets the per-request window up to the model's declared limit. - SillyTavern-compatible sampling. DRY sampling and
repetition_penalty_rangeare now honoured, and samplers Rapid-MLX does not implement (such as XTC or mirostat) are refused with an HTTP 400 naming the field instead of being silently dropped. - More reliable local coding agents. A new
piprofile (rapid-mlx agents pi --setup), dsh 0.2 setup, and headless opencode join the agent integrations. Claude Code prefix caching, integer tool arguments, and undeclared tool markup are handled more reliably. - Benchmarks and the leaderboard. Install, recipe, and benchmark-share output now link to the matching leaderboard views.
bench --submitno longer submits anything and points tobenchmark run+benchmark share; Desktop benchmark sharing follows the same workflow and refuses empty or invalid runs. - Desktop. Model unload is a labelled action beside the active model (
Unload allfor multi-model pools), with the existing guarded behaviour while work is in progress.
0.15.4 — 2026-10-02 · Experimental acceleration profiles and Share Compute in Desktop
0.15.4 adds two narrowly qualified experimental acceleration paths for Qwen3.8 27B and GLM-5.3 Flash, plus a complete Share Compute experience in Desktop. The accelerated profiles pin their model and runtime inputs and fail closed when a request or machine is outside the measured contract.
- Experimental Qwen3.8 27B acceleration. The opt-in
qwen3.8-27b-tensorfoldprofile pairs one qualified 4-bit checkpoint with its DFlash2 drafter and exposes the active backend and limitations through the server status APIs and Desktop settings. It admits one text request at a time and rejects tools, media, grammar constraints, and unqualified model or runtime revisions. The ordinaryqwen3.8-27b-4bitpath remains available as the feature-complete fallback. This release makes no fixed speed claim for the experimental profile. - Experimental GLM-5.3 Flash acceleration on 256 GB Macs. The dedicated
glm5.3-flash-tensorfoldalias uses the checkpoint's embedded MTP head and selects its qualified accelerated backend by default on eligible systems. Startup validates the exact runtime and checkpoint revisions; unsupported tools, media, grammar, and general batching fail explicitly. Desktop labels it Experimental with an opt-out that restarts on the ordinary GLM serving path, and Desktop chat works with the profile out of the box. On the tracked qualification fixture, a 460-token coding prompt with a 256-token thinking budget and 768-token reply cap measured 57.5 tok/s with MTP drafting against 47.5 tok/s on the same loaded engine with drafting disabled, a 1.21× ratio; these bounded samples do not establish a fixed multiplier or quality gain. - Share Compute arrives in Desktop. The new Experimental workspace guides users through choosing an eligible local model, reviewing a connection, contributing to the live pool, stopping a session, and restoring their previous local model. A separate read-only key stored in Keychain can show settled API credit across nodes on the same account, while local activity remains scoped to this Mac. Provider credentials are sent only to the local provider process and are not saved in app settings.
- Faster high-concurrency dense inference. An opt-in row-invariant lane matrix-multiplication path accelerates qualified dense 4-bit workloads at eight or more concurrent requests. Measured aggregate decode gains ranged from 50–81% on an M5 Max and 52–59% on an M3 Pro at 8–16 requests. It provides no measured gain at four or fewer rows and skips mixture-of-experts models.
- Serving is more predictable and better verified. Qwen acceleration caps admission at one active request, and both accelerated profiles verify installed runtime provenance before loading. GLM streaming preserves code whitespace and reasoning boundaries across streaming and non-streaming responses, Stable Audio 3 weight downloads are pinned to a reviewed immutable revision, and rapid server restart loops suppress duplicate anonymous startup telemetry within a bounded window.
0.15.3 — 2026-09-30 · Supervised Computer Use and more resilient local serving
0.15.3 adds opt-in, supervised Computer Use for bounded local Mac tasks, along with a documented API for clients that provide their own interface. It also improves lower-memory image generation, first-run setup, model serving, and recovery from common startup errors.
- Computer Use stays bounded and approval-first. Desktop can resolve a task to a selected macOS app or browser window, keep that scope visible, and pause for approval before consequential actions. Credentials, payment details, and commerce actions remain blocked, and runs stop when their target or approval context changes.
- A documented API supports custom Computer Use clients. The authenticated
/v1/cuaAPI covers capabilities, permissions, app and window discovery, bounded observations, run creation, event polling, approval decisions, cancellation, and a model-free--cua-onlyserver mode. Screenshots require opt-in from both server and request, and raw click and typing operations are not exposed. - Qwen Image 2.1 can fit on lower-memory Macs. The image path now supports a quantized mflux configuration with an 8 GB minimum-memory gate, while the original BF16 checkpoint remains available explicitly as
qwen-image-2.1-bf16. Actual fit still depends on the model and memory available to the process. - First runs and startup failures are easier to recover from. Running bare
rapid-mlxshows first-run guidance. Missing optional runtimes provide commands for the active install method, while missing local paths and curated model IDs receive clearer errors. - Desktop counts its first-run funnel without identifying an install. The anonymous milestones run once per new install, contain only the app version and a closed milestone name, and honour both telemetry and update-check settings. The privacy disclosure documents the complete boundary.
- Agent requests behave more predictably. Strict structured output works with streaming and tool calls, completion budgets stay within the model context, and QuickSilver prompt admission and hybrid checkpoint caching gain additional bounds for long sessions. Structured-output and context-length rejections can include a closed
reject_reasonwithout exposing request text. - Agent sessions keep their cached prefix on 16–32 GB Macs. A minimum cache budget keeps the agent-session prefix resident, and shutdown saves it for the next start, so repeat turns reuse the cached prefix instead of prefilling it again. See the measured comparisons.
- Blank PDF pages no longer stall OCR. Blank pages bypass text recognition, allowing scanned imports to finish or report that no readable content was found. Inference failures use closed telemetry categories, and private error text and Computer Use observations are not sent as telemetry.
0.15.2 — 2026-09-25 · Easier connections, safer starts, and useful failures
0.15.2 makes local model serving easier to connect, harder to misconfigure, and much more useful when startup fails.
- Connect an agent without hunting for server details. Desktop now puts the local endpoint and copy-ready connection information first, followed by guided integration choices. The server’s purpose is visible as soon as it is running.
- More capable local model paths. GLM-5.3 Flash can participate in the QuickSilver pool and gains an explicit default reasoning-effort control. Qwen4 can load validated PLE rows from a bounded sidecar. Laya and CLM System One are supported, and LFM2.5-VL gains a qualified DSpark companion.
- The default port no longer turns a healthy launch into a dead end. If port 8000 is occupied and you did not explicitly choose a port,
rapid-mlx serveselects the next free port from a bounded range and prints the choice. An explicit--portremains strict, and inherited listeners remain validated. - Missing optional support is actionable. Vision, image, video, and audio startup failures identify the exact extra and can offer to install it. Use
--yesfor unattended setup. - Startup failures keep their real cause. Hugging Face access failures, invalid configs, tokenizer failures, incompatible weights, quantization mismatches, and insufficient memory retain stable categories across CLI and Desktop instead of collapsing into a generic error.
- Crash diagnostics survive the process. Fatal server tracebacks are persisted, and the next launch can report that the previous startup ended before reaching a terminal state instead of leaving a silent gap.
Telemetry remains restricted to the documented closed event schema. Crash details stay local on the Mac and are not uploaded as telemetry.
0.15.1 — 2026-09-24 · Failed first starts tell you what went wrong
0.15.1 makes a failed first start diagnosable — for you and for us.
- A model that needs an optional extra now tells you which one. Vision, video, image, and audio serve paths now converge on the same actionable failure when their optional runtime is unavailable. The CLI and
python -m rapid_mlx.serverprint the exactpip install 'rapid-mlx[<extra>]'command, while the Desktop startup-failure panel names the closed reason and extra and links to the Startup Log for the installation details. Ready:means ready. The server prints its ready banner only after the port is actually bound, so a script that waits for the line can connect immediately.- Every serving lane reports success the same way. Image, video, audio, embedding, and specialised text servers now emit the same
model_servedevent as the default lane. - Telemetry: two additions, still anonymous, still closed enums.
server_start_staterecordsattemptedfollowed byreadyorfailed, with a failed stage ofresolve,download,preflight,prepare,engine_start, orbind. A missing optional runtime recordserror_class=missing_extraonmodel_serve_failed, withextralimited tovision,video,audio, orimage. There is no new free text or identifier. Full disclosure: rapidmlx.com/docs/telemetry.
Caveat: the bundled privacy policy names the new optional-extra field, but the Settings → Privacy summary does not yet name the new startup-state fields.
0.14.3 — 2026-09-18 · Repeat image questions stop re-reading the image, and Desktop can finish small jobs
Asking a second question about the same picture no longer re-encodes it or re-reads the conversation. Personal Intelligence moves from answering to doing a bounded piece of work on your Mac — with the exact path, content and command shown before anything happens. The Community Benchmark says where it is while it runs. And the downloaded DMG opens correctly again on current macOS.
- A follow-up question about the same image skips the work it already did. On a qualified Qwen3.6 35B media workload the second turn reached its first token 54.8% sooner and finished 22.4% faster, and the singleton lane generated 35.4% more tokens per second. The media-aware prefix cache keeps the exact prior-turn media boundary so the image is not encoded again and the unchanged conversation is not prefilled again; its one-time storage cost measured 10–25 ms on short eligible first turns and nothing on long prompts. A separate fast path stops merging and repacking a cache whose shape is already right — across a 21-case suite the response hashes and peak memory were identical. Both are default-on, fail closed outside their qualified boundary, and keep an explicit rollback switch. Multimodal sampling also reaches parity with text:
seed,top_kandmin_pare no longer ignored by the media scheduler. - Personal Intelligence can finish a small job, inside a fence you can see. A qualified local model can search and read local text, write a file you reviewed, move one file to recoverable Trash, and compile or run a bounded development command — enough to find a note in a folder, save a draft, or write and run a small program. It is not a shell. Desktop shows the exact path, content, command and arguments first; search and read can be approved for the session, while writes, execution and Trash need a fresh approval every time. Paths stay under your home directory, and hidden files, symlinks, credential stores, package descendants and oversized or non-text inputs are refused. Commands run with no shell in a macOS sandbox bounded on time, output, filesystem, network and processes, with compiler plugins and indirect argument files unavailable. Work lands in
~/Rapid Workspacewhen you do not name a destination. - The Community Benchmark tells you where it is, and what it measured stays what it measured. Desktop shows live stages, pass counts, an ETA and the latest measurement while a run is going, then offers a direct route to run another model or repeat the same one. Contributor identity survives an app restart. My Results and Community views distinguish an exact total from bounded recent evidence rather than inventing a zero or a comparison. A completed result keeps its model, protocol, execution, Mac and contributor scope even if the on-screen selection changes afterwards, and packaged builds carry provenance that is stamped and verified before anything can be published — a modified source build still measures locally but fails closed at publication. DeepSeek V4.1 can now run the registered benchmark through its qualified serial runtime: on a 256 GB M3 Ultra the pinned 2-bit path measured about 31.09 tok/s short-case and 29.42 tok/s long-case.
- The downloaded DMG opens again. The 0.14.0–0.14.2 disk images stored two Finder view records in a nested binary-plist shape that could close the install window on current macOS, and crash and relaunch Finder on Sonoma. The layout is now generated deterministically using the flat dictionary records Finder writes itself, the release verifier rejects the old encoding, and the branded background and Applications drag target render correctly.
- A 32K prompt on GLM-5.3 Flash no longer has to fit in one forward pass. Native MTP prefills in bounded 1,024-token chunks instead of materializing the whole prompt at once; two consecutive real 32K HTTP requests completed on a 256 GB M3 Ultra without Metal running out of memory. This does not claim the checkpoint's advertised 1M context is usable. Alongside it: stopping
rapid-mlx serveno longer hits the macOS 15 Python-finalization race that could hang or crash the process at shutdown, a cancelled DNS-pinned Desktop request cannot resume the same continuation twice, and a sandboxed command returns at its hard deadline even when macOS is slow to tear down a denied child. - One more text alias, and the package has its real name.
bonsai2-27b-2bitbrings the catalog to 193 text aliases, 247 total; Bonsai 2 Hadamard MLX packs now load through a Rapid-owned loader for their packed Hadamard projections and inverse embeddings, qualified across text, streaming and image HTTP requests. The Python distribution also moves to the canonicalrapid_mlxpackage name, keeping a compatibility shim so existingimport vllm_mlxcode goes on working.
0.14.2 — 2026-09-15 · Measured decode gains across four model families, and an approval-first Agent mode
Five qualified paths got faster on named hardware, and every one of them falls back to the stock route outside the shape it was tested on. Desktop gains an opt-in Agent mode that pauses before a consequential action and waits for you to approve the exact pending call. K2 Horizon 7B arrives on a Rapid-owned native runtime, and a set of prompt-cache and tool-loop failures found in real Desktop use are closed.
- Five measured speed-ups on the product path. DeepSeek V4.1 Flash on mixed DSpark K4 9.58 → 19.39 tok/s (2.02×); Qwen3.6 35B on the native MTP fixed suite 83.32 → 130.93 tok/s (+57.1%); GLM-5.3 Flash across six real tasks 26.58 → 35.54 tok/s (+33.7%); Qwen3.6 35B compiled no-MTP HTTP decode 101.7 → 121.6 tok/s (+19.6%); and its fused GDN HTTP path 62.3 → 79.1 tok/s (+27.0%). GLM-5.3 preserved all 12 reasoning traces and final answers byte-for-byte in its paired qualification. The Qwen3.6 kernels and the compiled replay fail closed to the stock route when the request shape, cache mode, concurrency or runtime is outside the tested set.
- Agent mode in Desktop, as an experiment. An opt-in mode backed by the same bounded runtime as Server: it can drive a local model through configured MCP tools, pause before a consequential action, show a bounded and credential-redacted summary, and continue only after you approve the exact pending call. The first version is deliberately conservative — runs are authenticated and process-local, completed state expires, tool sets are capped, unknown tools are treated as side effects, and an MCP reload cannot redirect an already approved call to a replacement tool. A denial records that the action did not execute rather than asking a compact model to reinterpret the refusal (#3480, #3439).
- K2 Horizon 7B on a native runtime, as an experiment. Server and Desktop run it through a Rapid-owned implementation including its reasoning and tool-call formats. On an Apple M4 Pro with 48 GB of unified memory the shipped paths measured 49–51 tok/s, 0.693 s median TTFT and 5.52 GB peak MLX memory — roughly Qwen3.5 9B decode speed at slightly lower peak memory, and slower and heavier than Qwen3.5 4B. The Rapid adapter and the checkpoint-bundled implementation produced the same output-token hash with median throughput 0.0089% apart, so no speculative performance patch was added. It stays experimental and is not a default recommendation (#3486, #3483).
- Long Desktop conversations stop paying to re-read themselves. A web-search or tool turn no longer moves temporary guidance to the front of the prompt, so a long conversation keeps its prefix cache instead of spending roughly 20 seconds re-reading itself on both the tool turn and its follow-up. Hybrid-model conversations resume from recurrent-state checkpoints, and exact prompt preparation is reused rather than repeated (#3454, #3446, #3435).
- Tool calls fail legibly instead of corrupting the edit. Malformed tool arguments now preserve the answer text already generated and give the model a precise correction target instead of surfacing raw protocol debris. An XML parameter value keeps the indentation of its first line, so a coding agent's replacement block still compiles — the parser was stripping all surrounding whitespace where exactly one newline per side is markup, which matches what vLLM and SGLang do for the same wire format (#3419, #3369). Document search, outlines, numbered sections and malformed retrieval cursors fail honestly rather than returning plausible but incorrect results.
- Code-editing turns reuse what is already on screen. Copy-draft now reaches sampled and cache-reusing requests: measured Qwen3.5 and Qwen3.8 turns accepted 71–99% of copied tokens and decoded 1.6–2× faster when the answer reused text already present in the conversation. Requests with nothing useful to copy are unchanged (#3417, #3398).
- Two more text aliases.
k2-horizon-7b-4bitanddeepseek-v41-flash-reap-2bitbring the catalog to 192 text aliases, 246 total. The DeepSeek REAP 2-bit checkpoint measured in 0.14.1 and held back now ships as an explicitly experimental alias behind a 224 GB minimum-memory gate — reachable by name on a machine that can hold it, still not a recommendation for ordinary Macs and still not a managed download.
0.14.1 — 2026-09-11 · Large and scanned PDFs, and a repetition loop that stops safely
Desktop can read a 300-page PDF or a stack of scans without forcing the whole file into one prompt. A multimodal request that fell into a token loop now returns what it produced instead of exhausting Metal and timing out. And an expensive 256 GiB checkpoint was qualified, measured, and deliberately not added to the catalog — the evidence for that decision ships with the release.
- Large and scanned PDFs. Selectable PDFs, image-only scans and documents that mix the two are all analysable. Extraction runs behind a bounded cache with hard read deadlines, and a follow-up question can retrieve a specific outline, page or section rather than re-sending the document. The dogfood covered both ends: a 300-page selectable Chinese PDF cached 125,833 characters with all 300 pages completed and 30 chapter rows resolved to real page offsets, and a 6-page image-only PDF reached the model through OCR with multi-turn synthesis decoding at 45 tok/s. Several failure boundaries closed with it: an OCR or render failure can no longer masquerade as complete extraction, a progress heartbeat cannot extend a read forever, and a model that keeps asking for tools is forced into synthesis.
- A runaway repetition loop now ends safely. An exact token loop in a tool-free multimodal request used to continue until Metal ran out of resource handles, leaving the client waiting on an empty HTTP 500. The scheduler now checks the existing conservative detector every eight generated tokens, retires the stopped batch row immediately, and returns the valid partial response with a normal length finish. The reported 61-token cycle is caught at token 184.
- Sampled MTP has a stronger correctness net. New distribution-level tests cover independent speculative acceptance draws and rejection-residual sampling, with real-weight checks forcing the K=2 and K=3 paths. They earned their place immediately: the prior focused suite passed all 246 tests, while four of the five new distribution checks failed and exposed total-variation error up to 0.14. The real-weight smoke recorded 118 K=2 and 147 K=3 speculative attempts, so the deeper arithmetic is exercised rather than inferred from a nonzero acceptance rate.
- A qualification we published for saying no. An experimental DeepSeek V4.1 Flash REAP 2-bit checkpoint was measured on a 256 GiB M3 Ultra: a 199 GiB artifact, 239.44 s to load, 213.51 GB peak, and 7.31 tok/s conservative decode against a 12 tok/s product floor — short by 34%. It preserved 99.9949% mean robust routing mass and passed tiny-configuration parity within 1.4e-6, so the quality was there and the speed was not. The tooling and evidence ship in the repository and the checkpoint is archived for expert reproduction, but it stays outside the catalog: no Server entry, no Desktop entry, not a managed download.
0.14.0 — 2026-09-10 · Measured decode gains, two experimental workspaces, and a much larger catalog
Four measured speed-ups land on the product path, not in a micro-benchmark: qualified verification, a qualified 8-bit pairing, a fused kernel and the scheduler each got faster on a named Mac. Computer Use and Share Compute arrive as explicit opt-ins, the Community Benchmark says what it is doing while it runs, serving starts from one command, and the catalog grows to 190 text aliases plus seven more image models.
- Four measured improvements on the product path. On an M3 Ultra: Qwen3.8 27B MTP GDN verification 45.87 → 52.28 tok/s (+14.0% decode, verification sync 4.954 → 4.644 s); Muse-Glimmer 30B 8-bit DFlash 22.1–22.6 → 41.3–46.3 tok/s (2.00× code median, 1.83× chat); Qwen3.8 Flash-Next 4-bit fused GDN +6.35% end-to-end (+26.28% on the isolated layer); and a short request queued behind three long prompts reached its first token in 1.206 s instead of 3.067 s (2.54× faster). The MTP result used one checkpoint, prompt, seed and K=3 draft depth across three alternating runs with matching output hashes. Fused Flash-Next GDN and shortest-validated-tail scheduling stay explicit opt-ins while their hardware coverage grows.
- Computer Use, as an experiment. Enable it under Experimental Features for bounded workflows that plan and execute locally across supported Mac apps — supervised Downloads cleanup, and drafting content before posting it through Safari or Chrome. Permissions and app targets are scoped, retries are bounded, the visual recovery loop is tied to the active server session, and consequential actions stay reviewable. This is an early preview, not unattended general-purpose automation.
- Share Compute, as an experiment. A new workspace lets you offer compatible local capacity to a configured compute pool. Server and Desktop validate model fit and configuration before accepting work, keep service bindings on safe local interfaces, expose live health and lifecycle state, and propagate cancellation back to the local server. It is off by default and does not participate until you opt in.
- The Community Benchmark tells you where it is. Desktop and CLI now show the selected model, the current phase, elapsed time and an ETA while a run is in progress. Rows share one definition of median decode throughput and TTFT, record run conditions and cached model metadata, distinguish local results from shared ones, and preview contributor identity before upload. Results stay local unless you explicitly share them.
- Ten more text aliases.
g9v3-39a5b-4bitis a 21.3 GiB mixed-precision MoE for 32 GB+ Macs, its reasoning probe decoding at 85 tok/s on an M3 Ultra. Granite 4.2 arrives as six aliases (30B/8B/3B × 4-bit/8-bit), all through the same 38-check server journey, measured from 20.8 tok/s (30B 8-bit) to 175.1 tok/s (3B 4-bit).minicpm5-2b-4bitis a compact 1.3 GiB option that scored 24/31 in 7.17 s against 26/31 in 32.54 s for the 4B comparison — faster, with the quality tradeoff documented.neohorse-9b-4bitmeasured 44.5 tok/s but did not beat the default across the full quality suite, so it ships available rather than selected. - Seven more image models. FLUX.1 schnell, Stable Diffusion 3.5 Large, SDXL Base, Bonsai Image 4B 2-bit, HiDream O1 Dev and Qwen Image Edit join the catalog. From the release qualification matrix: FLUX.1 schnell at 1024×1024 in 23.36 s with 9.46 GiB peak, Bonsai 4B 2-bit at 512×512 in 5.39 s with 6.52 GiB peak, SDXL Base at 1024×1024 in 16.46 s after construction. These are single-machine measurements; the readiness UI shows the memory tier before a costly load.
- Serving starts from one command.
rapid-mlx startpicks a fitting cached model, starts the canonical server, waits for readiness and configures a supported agent profile —--dry-runpreviews every mutation.rapid-mlx serviceinstalls and operates a least-privileged headless macOS service with status, logs, restart, uninstall and rollback.rapid-mlx doctoradds lifecycle, import, cache, configuration and service diagnostics with verified repairs. Desktop can unload an idle resident model in one click while keeping the selection, and it refuses to interrupt active chat, image, audio, video or dictation work. - An optimization we turned off. Qwen3.5 4B MTP is now off by default: release qualification measured single-stream decode 25–37% slower on M2 Pro and M3 Ultra with no concurrency gain. It remains available explicitly. Telemetry now honours
RAPID_MLX_TELEMETRY=0,DO_NOT_TRACK=1and common CI markers in both the engine and the Desktop app.
0.13.4 — 2026-09-03 · Qualified Qwen models go faster on their own, and benchmarks stay on your Mac
Four qualified Qwen artifacts now pick their validated speculative-decoding preset without being asked, so concurrent work finishes sooner with no flag to remember. Video generation gets a real workspace in the Mac app, the Community Benchmark keeps every result local until you explicitly share one, and a model that will not fit now says so before it tries to load.
- Qualified Qwen models select their own MTP preset. Qwen3.5 4B/9B, Qwen3.6 27B and Qwen3.8 27B (the exact qualified 4-bit artifacts) now choose their validated multi-token-prediction preset and the continuous scheduler when no speculative-decoding option is given. Across mixed four-request cohorts, aggregate throughput improved 14.1%–30.8%. For Qwen3.8 27B the single-request path also scales with context: 1.43× the 0.13.3 ordinary path at 128 tokens and 2.34× at 32K.
--no-spec-decoderestores ordinary decoding; Desktop has the same switch. - Video generation in the Mac app. Enable it under Experimental Features for a local workspace with queued jobs, progress, cancellation and restart-safe results. The signed app carries the runtime and encoder for LTX 2.5 q8 (text and image to video), Wan 2.1 1.3B bf16 and CogVideoX-Fun 5B q4 — nothing to provision after install. Wan 2.1 needs a Mac with at least 40 GB of unified memory.
- Community Benchmark is private by default.
rapid-mlx benchmarkand the new Desktop page plan, run, validate and archive text, image and video measurements locally. Nothing leaves the Mac unless you choose to send it, and the consent step shows the exact destination and payload digest before you do. - A model that will not fit tells you first. CLI, server and Desktop now read one atomic model registry, per-model Metal limits are derived from the machine's available unified memory, and a rejected launch returns fit guidance instead of failing deep inside model loading.
- Headless deployment is documented and qualified. A LaunchDaemon guide covers dedicated service accounts, loopback-first networking, KeepAlive, reboot recovery and FileVault — see Run Rapid-MLX at boot on a headless Mac.
- Desktop quality of life. Opt-in persistent memory across conversations (editable and removable in Settings), Mermaid blocks rendered as isolated network-blocked previews, dictation that overlaps model page-in with your speech, and an update card that says whether the installed app is current, behind, or ahead of the public release.
0.13.3 — 2026-09-01 · GLM-5.3-Flash, and vision chats that stop recomputing
A 320B-parameter model now answers on a desktop: glm5.3-flash-4bit runs through the CLI, the OpenAI-compatible server and the Mac app. Vision models stop recomputing the text they have already seen, finished video jobs survive a restart, and the app gains a command palette.
- GLM-5.3-Flash is served.
rapid-mlx serve glm5.3-flash-4bitruns Zhipu's 320B MoE (18B active per token) with the processor and runtime compatibility the checkpoint needs. Measured on a 256 GB M3 Ultra: a sustained 512-token response decoded at a median 29.2 tok/s using 165.4 GB of active MLX memory, so the alias asks for the 192 GB tier. Speculative decoding is deliberately off for it — its qualification run showed no throughput gain, and we would rather ship the honest default than a flag that looks impressive. - Vision chats reuse the text they have already processed. A hybrid multimodal conversation whose current turn is text-only can now reuse its text prefix instead of recomputing it, with cache identity bound to the rendered request and model state. Requests that actually carry media still go through the validated vision path — a picture never scores a text-only cache hit.
- Finished video jobs survive a restart. Completed generations persist across server restarts, and a client can ask what a model is capable of before paying to load it.
- A command palette in the Mac app. Common actions are reachable from the keyboard, a diagnostics shortcut makes failures easier to report, model pickers stay scoped to the task you are in, and streaming Markdown keeps rendering when a window comes back from a disconnected display.
- Packaged builds can repair themselves. A Desktop install that fails after the DMG step can now recover its engine instead of leaving a half-installed app.
0.13.2 — 2026-08-30 · Flash-Next decodes faster, the app gets 43% smaller
Long sessions get quicker on both ends: Qwen3.8-Flash-Next can opt into its own multi-token prediction head, long prompts reach the first token roughly a third sooner, and a repeated prompt reuses its cache instead of paying for it twice. The Mac app ships as a download less than half the size it was, Kokoro is ready before you go offline, and attachments, credentials and model switching got a safety pass.
- Native MTP for Qwen3.8-Flash-Next. Opt in and the checkpoint's own one-layer prediction head proposes tokens while the full model verifies every one of them, so the output is the model's — not an approximation. Measured on an M3 Ultra: 25.2 → 34.9 tok/s at a short prompt and 21.2 → 28.8 tok/s at 32K (+36–42%), with 76.41% of proposals accepted. It costs up to 6.6 GB more active memory, and seeded requests or stateful grammar, tool and reasoning processors deliberately stay on ordinary decoding rather than quietly weakening their contract. 192 GB remains the recommended tier for this experimental checkpoint (#2572, #2655).
- Long prompts reach the first token ~a third sooner. QSA index-cache work is now batched across eligible prefills: 2K 3.35 → 2.26 s, 8K 13.69 → 9.24 s, 32K 62.85 → 44.66 s (−29–32%), with decode speed and cache precision unchanged (#2574, #2596).
- Repeated prompts get their cache back. Prefix snapshots now follow the exact rendered prompt and survive batching and persistence, so a warm 5,288-token request reused 5,273 tokens and finished in 0.539 s — and a completed 32K request can no longer poison the next one in the same process (#2588, #2644).
- Kokoro is ready before you go offline.
rapid-mlx pullnow fetches the voice assets and prepares the English G2P requirement as part of the pull — validated by generating speech with networking disabled after a clean cache. A missing runtime requirement returns an actionable readiness error instead of installing software mid-request (#2648, #2664). - The Mac app download is 43% smaller. LZMA packaging and dependency-proven pruning took the signed, notarized DMG from 187.5 MB to 106.3 MB. The honest tradeoff: the one-time Finder copy is slower on the measured host, 5.94 s → 17.80 s. Releases now publish the exact approved and notarized bytes rather than rebuilding after the tag (#2668, #2775).
- Attachments, credentials and model switching got a safety pass. A message takes up to four images and 6 MiB of encoded image data and says which limit rejected the rest; deleting a generated image uses an app-owned sheet where Return cannot delete; switching away from a busy model asks first and Cancel keeps the in-flight response; and Tools settings resolve a search key's real Keychain state without ever exposing it (#2541, #2578, #2543, #2514).
- Tool calls and explicit flags mean what they say. Qwen3.8-27B required and named tool calls use the checkpoint-native format and return OpenAI-compatible JSON arguments, an invalid required argument now fails as a client error instead of looking executable, explicit
--mllmwins over automatic fallback, and an explicit variant pull persists its subfolder so a laterserveresolves the checkpoint you actually pulled (#2660, #2643, #2558).
0.13.1 — 2026-08-27 · Qwen3.8-Flash-Next, day-0
Qwen's newest flagship runs on your Mac the day it exists: qwen3.8-flash-next-4bit serves our own 4-bit MLX conversion of the 125B Qwen3.8-Flash-Next through a vendored qwen4_exp backbone. Also in this release: pull exactly one variant of a multi-variant repo, subfolder-quant repos download from the mirror again, and switching models mid-session got a real lifecycle.
- Qwen3.8-Flash-Next is served.
rapid-mlx serve qwen3.8-flash-next-4bitruns our 4-bit conversion of Qwen's 125B MoE (~6B active parameters per token) through a vendored implementation of itsqwen4_exparchitecture, verified for parity against the reference implementation. Experimental, text lane, hermes tool calling; plan for a 128 GB+ Mac. The weights pull from Hugging Face — at 4-bit × 125B they are large enough that we deliberately serve them from the source rather than the mirror. - Pull one variant, not the whole repo.
rapid-mlx pull <repo> --bits 4/--format mlxselects a single variant of a multi-variant repository instead of downloading every quantization it ships (#2338). - Subfolder-quant repos download from the mirror again. Repositories that keep each quantization in its own subfolder (LFM2.5 among them) were hard-declining the CDN mirror and always falling back to Hugging Face; they now pull the exact subfolder from the mirror (#2279).
- Switching models mid-session got a real lifecycle. Replacing the resident model now quiesces admitted requests first, hands off before the old engine retires, and surfaces an honest abort reason to any stream it interrupts — a dictation session no longer dies because chat swapped assistants underneath it.
- The 8 GB Quickstart is safe again. The Mac app's Quickstart pick on the smallest Macs is back to a model that actually fits (#2432).
0.13.0 — 2026-08-26 · Ornith-1.5, and prefill that tunes itself
0.13.0 brings the Ornith-1.5 family (9B dense + 35B-A3B MoE) and Nemotron-Labs-Diffusion 3B into the catalog, adopts MLX 0.32.1, auto-tunes recurrent prefill on verified profiles, and hardens the tool loop. The desktop app redesigns audio into two lanes, polishes the DMG install, and wires the first-run wizard onto the RAM-recommendation SSOT.
- New models. Ornith-1.5 — an open, MIT-licensed family — lands as first-class aliases:
ornith-1.5-9b-bf16(dense) andornith-1.5-35b-a3b-bf16(MoE A3B), both on the Qwen3.5-class GatedDeltaNet architecture with hermes tool calling (#2197). Nemotron-Labs-Diffusion 3B is served in AR mode asnemotron-labs-diffusion-3b-4bit(#2194). Qwen3-Next models can stream weights from disk (#2192). - MLX 0.32.1. Adopted after Mac model benchmarks (#2199), with a guard that prevents quantized-KV corruption on the new runtime (#2234) and persisted KV isolated by model revision (#2202).
- Prefill tunes itself. Recurrent-model prompt prefill picks its chunk size from measured profiles instead of one fixed default, scoped to verified model profiles (#2210, #2211, #2213), and multimodal serving separates the MLLM prefill budget from the vision budget (#2215).
- A sturdier tool loop. Malformed tool calls are recovered through the tool loop instead of dropped (#2278), chat templates resolve from model profiles at load time (#2288), Aider stays pinned to the served local model, and Hermes gets cross-version toolsets.
- Desktop: audio in two lanes. Dictation and voice notes get separate readiness UX with the download cliff surfaced up front (#2188); the DMG install flow is polished (#2214); the first-run wizard picks models from the same RAM-recommendation SSOT as the CLI (#2232).
- Desktop: agents before the engine. “Connect your agents” is useful before an engine starts (#2308), agent health is surfaced with a re-run setup path (#2230), and the tray menu gains “Copy API endpoint” (#2246).
0.12.18 — 2026-08-21 · Branch your chat, run more models
Regenerating an answer no longer throws the old one away — every alternative stays a switchable branch. This release also widens the catalog (LTX-2.5 video, Qwen Image, Ternary Bonsai 8B, North-Mini-Code, GPT-OSS Puzzle), makes dictation start hot, speeds up MoE decode, and hardens multimodal serving.
- Regenerating an answer keeps the old one. Regenerate, Retry, and prompt edits now preserve every alternative as a switchable branch — step between them with the
‹ 2/3 ›control under the bubble. Deleting a turn tells you exactly how many turns it removes, each fork remembers where you left off, and old conversations load unchanged (#2147). - Dictation starts hot. The catalog lookup and lazy weight load that used to land inside the first dictation of a session now run at enable / model-pick time — a 30-second clip lands in well under half a second once warm, and the Dictation tab attributes slow runs (“model 1.2 s · asr 0.3 s”) (#2152).
- Faster MoE decode. Expert gate+up projections are fused into a single
gather_qmmlaunch (#2151), and GDN prompt prefill gets a blocked-seq Metal kernel (#2144); GatedDeltaNet input projections fuse into one quantized matmul (#2159). - New models. LTX-2.5 video generation (#2158), Qwen Image through mflux (#2157), Ternary Bonsai 8B (#2162), North-Mini-Code (Cohere2-MoE, with a dedicated reasoning parser so chain-of-thought stops leaking into content) (#1223, #2171), and GPT-OSS Puzzle checkpoints (#1224).
- Speculative decoding stays under your control. Tested model families expose experimental MTP presets in the desktop (#2173), and advanced users can opt into structurally compatible target/drafter pairs even when Rapid does not recommend them by default (#2177) — capability and recommendation stay separate, so an experimental path never silently becomes the default.
- Safer multimodal serving. Remote and local media inputs are validated against SSRF and arbitrary-file-read paths before the VLM loader sees them (#2167). Gemma 4 checkpoints with stale embedded chat templates are upgraded to the shipped compatible template at load time.
- SVG code blocks render a preview — offscreen, never on the network (#2148).
- Mac & CLI polish. Background update progress is visible again (#2129), WON’T FIT rows open a read-only Review (#2131), a fresh install can quit before choosing telemetry (#2178);
agentssizes its name column to the widest alias (#2139), and a flaky update-check response no longer crashesrapid-mlx upgrade(#2168).
0.12.17 — 2026-08-19 · Dictation fixed — update now
The 0.12.16 release build shipped without the macOS microphone entitlement, so dictation could never be enabled. 0.12.17 fixes it and gates every future release on the sealed signature carrying the key.
- Dictation works now — please update. Speech to Text’s “Allow…” button did nothing in 0.12.16: the notarized build lacked the microphone entitlement, so macOS refused the request silently — no prompt, no error. The entitlement now ships in the signed app, and the build pipeline refuses to sign a release without it (#2134, #2135).
- Agents see the context window your Mac can actually serve: memory-aware
max_model_lenon/v1/models, andrapid-mlx models --jsonfor scripts (#2122). - Codex can dispatch MCP tools again over the Responses API (#2130); out-of-range
--portfails with an actionable message (#2127); the unknown-model hint drops its drifting alias count (#2128).
0.12.16 — 2026-08-19 · Dictate into any app
Tap right Option and speak — the words land at your cursor in any app, transcribed on your Mac. Also in this release: web search that works out of the box, model picks that follow the Artificial Analysis index, chat that typesets math, sliding-window prompt-cache reuse, and a page-by-page docs audit against the shipped code.
- Dictate into any app. Tap right Option, say a sentence, and it lands wherever your cursor is — any app, even with Rapid’s window closed. Audio goes to your local server’s transcription endpoint and nowhere else (#2049).
- Web search now works out of the box. Desktop chat’s web search defaults to Keenable’s keyless service — no account, no API key, no more DuckDuckGo rate-limit dead ends. A free Parallel key in Settings → Tools buys the best measured quality (tops the Artificial Analysis Search Index; ~1,000 free searches a month) (#2040–#2044).
- Your Mac is offered the smartest model it can actually run. Recommendations follow the Artificial Analysis Intelligence Index: every Mac from 32 GB up gets Qwen3.8-27B — GPT-5.6-class intelligence at ~40 tok/s with multi-token prediction on (#2055).
- Chat renders like it should. Inline and block LaTeX are typeset the way models actually write them (#2107); code blocks stop flickering while an answer streams (#2106); markdown code blocks and tables are visible again (#2056).
- Conversations get folders and Markdown export, and getting a model is two explicit steps — Download, then Start (#2053). Onboarding fills the window instead of floating as a small card (#2063). Voice notes in Apple formats (M4A, CAF) transcribe directly (#2101).
- Faster in long chats. Sliding-window models keep reusing their prompt cache across turns instead of re-prefilling (#2064); the prefix cache releases its memory after sitting idle (#2038) and chat can opt out of it for sensitive prompts (#2048).
- The docs were audited page by page against the shipped code (~20 fix PRs): all 94
serveflags, every operator environment variable and the full endpoint surface are documented; six drifted defaults fixed and pinned by tests; PRIVACY/SECURITY tell the whole story (#2065–#2113). And theinstall.shyou curl is now published only at release tags — byte-identical to the tagged source (#2092). - CLI polish.
chat --port Nasks the server what it serves (#2046);launch continueandcontinue-devresolve to the same client and theagentsfooter counts honestly (#2108);recipechecks free disk before printing aservecommand that cannot fit (#2119).
- @xiaoxiunique — math typesetting (#2107), streaming-fence stability (#2106), jump-to-latest, window floor and duplicate-attach fixes (#2021, #2022, #2023)
- @osdodo — markdown visibility (#2056) and the image progress sweep (#2028)
- @ryo1sato — release the prefix cache after idle (#2038)
- @loriz-art — the full-window onboarding design (#2063)
- @guo — NOTICE and upstream attribution (#1968)
0.12.15 — 2026-08-18 · Tool calls stop truncating what your agent writes
If a coding agent wrote files through a local Qwen model on 0.12.5–0.12.14, upgrade first — outputs could be silently shortened. Also in this release: Qwen3.8-27B up to 23% faster with one flag (identical output, checksummed), downloads that reroute instead of stalling, voice transcription that stops inventing words, and safer model loading.
- Upgrade first if you run coding agents. On 0.12.5–0.12.14, an agent (Claude Code, Aider, and friends) writing files through a Qwen 3.5/3.6 model could have its output silently shortened: a file meant to be 710 bytes arrived as 11 bytes, and nothing looked wrong — no error, valid response. Even short values could corrupt (
Tokyo→Toyo). The cause was our constrained decoding fighting the model’s own output format; 0.12.15 fixes it. If your agent edited anything important on those versions, it’s worth re-checking those files. Gemma-4 and JSON-wire models were never affected (#1996, #1997). - Qwen3.8-27B gets faster — for free. Add one flag —
--speculative-config '{"method":"mtp"}'— and generation goes 38.9 → 44.1 tok/s on an M3 Ultra, up to 23% faster when the model writes code, JSON or lists. The output is exactly the same: same sha256 with it on and off. Works on the aliases that ship an MTP head (#2004); in the desktop app it’s a switch under Settings → Performance (#1987). - Downloads stay fast. Pulling a model now watches its own transfer speed, file by file; if a download slows to a crawl, that file automatically finishes from Hugging Face instead — mid-pull, no restart. In our fresh-install test a 9B model was downloaded and chatting in about two minutes at 90 MB/s.
RAPID_MLX_MIRROR_MIN_MBPS=0restores the old behaviour if you ever need it (#2015). - DeepSeek Harness works out of the box.
rapid-mlx agents dsh --setupwires thedshcoding agent to your local server — it joins Claude Code, Codex, Hermes and Aider as agents we run against every release before shipping (#1982). And when the served model can’t do step-by-step reasoning,dshno longer shows an effort selector that does nothing (#1984). - Voice transcription stops inventing words. Pauses and silence used to come back as imagined speech — Whisper was transcribing its own padding. A silent clip now returns an empty transcript, like it should (#1964).
- Safer to try models you just found. A model repository can no longer run code on your machine simply by being loaded; downloaded Python artifacts are integrity-checked (SHA256); the audio loader refuses pickled weights unless you explicitly allow them;
rapid-mlx sharewarns before a key leaves your machine; and--trusted-hostscan pin which HTTP Host headers your server answers (#2011). - Fixes.
install.shrequires a native arm64 Python and repairs a half-built venv (#1994); hybrid vision models get the right cache instead of answering from stale state (#1963); media-only models no longer appear as launchable chat models (#1965);rapid-mlx psandmodelsstop colliding columns andconnectno longer announces a server that is not running (#2006); the unusedjlenssubcommand is removed (#1992). - Desktop. Settings gains a Developer section in development builds only — rehearse first-run without touching real state, with every erasure opt-in and named before it happens (#1993); the status footer sheds readouts instead of squeezing six chips into four’s width (#1991); an MTP sidecar whose quantization differs from the base now warns instead of silently degrading (#1989).
- @xiaoxiunique — Settings → Developer with re-onboarding (#1993) and the status-footer fix (#1991)
- @rinaldofesta — warn when an MTP sidecar’s quantization differs from the base (#1989)
- @YUHAO-corn — strip multi-line release-note comments (#1990)
- @moinulmoin — reported the x86_64-Python installer trap (#1994)
0.12.14 — 2026-08-15 · Qwen3.8-27B in our own mixed-precision build
A Rapid-MLX mixed-precision quantization of Qwen3.8-27B that spends its bits where they matter, plus user-owned model aliases, the DeepSeek Harness agent, and the desktop app handing all updates to Sparkle.
- Qwen3.8-27B, mixed-precision.
rapid-mlx chat qwen3.8-27b-mixed-3.5bpwserves our own quantization of rapid-mlx/Qwen3.8-27B-mixed-3.5bpw-MLX — 13.0 GB of weights versus the 4-bit build’s 15.0 GB, spending bits where they matter instead of flattening every tensor to 4-bit. Measured throughrapid-mlx serveon an M3 Ultra (8K prompt, peak across the process tree): 20 GB peak, 323 tok/s prefill, 40 tok/s decode, zero swap — plan for a 48 GB Mac or larger. Strong at code (10/10) and tool-calling (25/30, including parallel and sequential multi-tool), but 5/10 on the MATH subset with thinking off — which is why it is not a built-in RAM-tier default; for arithmetic-heavy workqwen3.6-35b-4bitorgemma-4-26b-4bitstay the safer picks (#1954). 172 text aliases, 221 total. - Name your own models.
rapid-mlx alias set fast qwen3.5-4b-4bitgives any built-in model or Hugging Face repo id a name you choose, stored in your own config — no repo edit, no fork. Aliases can’t shadow a built-in name or chain to another user alias, and a malformed store fails closed rather than resolving to the wrong weights (#1942). - DeepSeek Harness joins the agents.
rapid-mlx agents dsh --setupwires DeepSeek Harness to your local server like the Tier-1 agents, andrapid-mlx launchlists it. The agent test sweep now runs each agent inside a throwaway home, so it can no longer read or rewrite your real agent config (#1952). - Vision pixel bounds.
--vision-min-pixels/--vision-max-pixelscap what a multimodal model’s image preprocessor produces, so a large photo can’t balloon the KV cache on a memory-tight Mac. Both default to off (#1943). - Desktop: Sparkle owns updates, end to end. The in-app DMG installer is gone; the app checks, downloads, verifies the signature, and installs through Sparkle. Builds from 0.12.12 onward keep themselves up to date; older builds are told to download once and never hit a broken in-app install path again (#1947). Qwen 3.8 models are recognized as their own family — correct picker name and tool-calling shown as a capability (#1959).
- Fixes.
/v1/modelsreports a hybrid model’s live metadata instead of a stale snapshot (#1955); a client disconnecting mid-stream no longer logs a false “missing finish” warning (#1951); rejected CORS preflights return the correct content length (#1949).
0.12.13 — 2026-08-14 · Qwen3.8-27B, text and vision
Qwen’s newest 27B lands day-0 — a hybrid GatedDeltaNet model serving as a strong general-purpose and tool-using text model with full batched throughput, plus a vision-language mode — alongside a batch of desktop app fixes, including chat attachments.
- Qwen3.8-27B is one command away.
rapid-mlx chat qwen3.8-27b-4bitserves the hybrid (GatedDeltaNet) checkpoint on the default lane; start with--mllmto answer questions about images through the serialized single-stream vision lane — verified describing a real photo end-to-end. The stale routing notice claiming--mllm“will error” for hybrid backbones was corrected (#1939). Catalog: 171 text aliases, 220 total. - Model list. An empty cache no longer shows a phantom “No” entry parsed out of the CLI’s “no models cached yet” notice (#1920).
- Onboarding. Completing macOS onboarding now requires explicit confirmation, and model inspection is deferred until consent (#1917, #1926).
- Desktop. Chat accepts PDF, CSV and TXT attachments (#1932); the Launch page’s Codex/Hermes rows actually launch (#1933); resident image downloads show activity instead of a dead button (#1922); audio models are downloaded and verified before serving (#1928, #1923); long answers no longer degrade from O(n²) link cursor rects (#1930); hover no longer masks button labels (#1934); the composer placeholder stays clear of IME text (#1935); zero-token stopped turns are dropped from wire history (#1914).
- @loriz-art — onboarding requires explicit confirmation; model selection kept inside Step 2 (#1917, #1931)
- @osdodo — chat attachments (PDF/CSV/TXT) and audio-model verification before serving (#1932, #1928)
- @xiaoxiunique — Launch rows fixed, hover-label fix, O(n²) cursor-rect fix (#1933, #1934, #1930)
- @Jevin-F — composer placeholder clear of IME text (#1935)
0.12.12 — 2026-08-13 · Vision gets its first token faster
A vision-lane performance release plus resilience: streaming vision requests stop rebuilding the tokenizer vocab on every call, model loads can no longer hang on a cold network, and the desktop app gains signed background updates.
- Vision first-token latency drops sharply. The output-router detection that runs on every streaming request was rebuilding the full tokenizer vocab (~262k entries on Gemma) on each call; it is now detected once and memoized, and the MLLM scheduler wakes on an event instead of a 10 ms idle poll. gemma-3 image requests that previously returned an empty completion (or crashed) now answer correctly (#1909).
- Model loads no longer hang on a cold network. Warm starts could stall on unbounded Hugging Face metadata calls; the loader now hands a local snapshot path to the model, so an offline or poisoned-DNS start can’t wedge the server (#1908, #1888).
- Streaming errors tell the truth. MLLM streaming preflight errors return HTTP 400 instead of a silent 200 with empty content (#1897).
- Scheduler & caches. Batch-adaptive recurrent-barrier interval for single-stream decode (#1895); hybrid sliding/full KV caches keep their per-layer cache class when trimmed (#1863); image/video models are single-slot residents so they stop accumulating in memory (#1889).
- Catalog. Ling-3.0-tiny gains block-FP8 checkpoint support (#1910).
- Desktop. Signed background updates via Sparkle (#1907); streaming aligned with native chat (#1906); conversation search in the sidebar (#1879); global and per-conversation custom instructions (#1883); Launch integrations derived from the engine registry (#1894).
- @guo — release-pipeline hardening: PF-2 secrets gate, and the GitHub Release created last, after the updater pointer (#1859, #1891)
- @osdodo — desktop conversation search, custom instructions, and Sparkle signed background updates (#1879, #1883, #1907)
- @xiaoxiunique — desktop streaming aligned with native chat (#1906)
- @loriz-art — core workspace unified with the Direction D visual system (#1905)
- @Jevin-F — model loads no longer hang on unbounded Hugging Face metadata calls (#1908)
- @aftersnow — Ling-3.0-tiny block-FP8 support (#1910)
0.12.11 — 2026-08-12 · Faster under load, faster alone
A bench-driven performance release: concurrent throughput climbs 8–41% depending on the model, a lone request gets its own KV-cache fast path, and forced tool calls always arrive as valid JSON objects.
- Concurrent requests prefill as one wave. The hybrid admission throttle is retired — requests that arrive together are admitted together instead of being staggered 200 ms apart, which kept batched attention off its fast path for the batch’s whole lifetime (#1866). Measured on an M3 Ultra: Qwen3.6-35B-A3B-8bit at 8 concurrent streams goes 190 → 268 aggregate tok/s (+41%); Qwen3.5-4B +20%, 9B +13%, gpt-oss-20b +12%, 27B +8%.
- Deep batches coalesce scheduler steps. With four or more running requests the engine runs up to 4 scheduler steps per dispatch, cutting per-step executor round-trips (~10 ms/step at high aggregate rates). Finishes, pending admissions, and the memory-pressure check cadence are preserved (#1878).
- A lone request keeps a singleton KV cache. Single-stream decode skips batched-form bookkeeping — plain causal mask straight onto the native SDPA fast path, and the dense-sampler fast path now engages at batch size 1, promoting to batched caches only when a second request joins (#1874). Qwen3.5-4B single-stream: 174.8 tok/s at short context, 147.5 at 16k.
- Forced tool calls always carry object arguments. Under
tool_choice, a weak continuation could produce a bare scalar ("arguments": "1") — shipped as-is on the non-stream path and silently dropped as an empty turn when streaming. Both paths now repair to the OpenAI wire contract, with schema-required violations still surfacing explicitly, and tool logits processors are no longer built for plain-chat requests (#1880, #1878). - Reliability. Prefix-cache entries can no longer be corrupted by the new singleton lane (loaned state is copied on admission); coalesced-step errors deliver already-produced tokens before surfacing; admission-wave timing survives malformed tuning env vars (#1874, #1866, #1878).
- Catalog. New alias
nemotron-3.5-lightning-30b-4bitjoins the Nemotron line (170 text aliases, 219 total).
0.12.10 — 2026-08-11 · A 131K-context reasoner on an 8 GB Mac
inclusionAI’s Ling-3.0-tiny gets the first MLX conversion anywhere and ships served natively — a 7.9B mixture-of-experts reasoner with 1.3B active parameters that fits an 8 GB Mac. Underneath: MoE experts can stream from disk, the scheduler fails loudly instead of hanging clients, and auth fails closed.
- Ling-3.0-tiny is served natively.
rapid-mlx serve ling-3.0-tiny-4bitruns inclusionAI’s 7.9B-total / 1.3B-active sparse-MoE reasoner (128 experts, top-8 plus one shared, KDA + MLA hybrid attention, 131K context, MIT) through a vendoredbailing_hybridbackbone verified against the official modeling code on identical weights (#1817). Thinking streams asreasoning_content— toggled byenable_thinkingor the model’s detailed thinking on/off switch — and tool calls parse natively in both stream modes. The 4-bit conversion (4.2 GB) is the first MLX conversion of this architecture anywhere, published under the rapid-mlx org on Hugging Face and mirrored on the model CDN (#1819, #1826). - MoE experts can stream from disk.
--disk-streamstreams routed-expert weights from disk instead of holding every expert resident (#1803). - Scheduler errors fail loudly. An engine error mid-batch now fails the in-flight requests instead of leaving their clients hanging (#1810).
- Auth fails closed. An empty companion key or a duplicated
x-api-keyheader is rejected instead of slipping through (#1811). - The desktop app gets a unified Settings surface. Settings UI and menu-bar branding are reworked into one coherent surface (#1828), and the CLI’s RAM-tier model recommendations are now shared with the Mac app so both surfaces suggest the same models for your machine (#1832).
- Desktop paper cuts. Downloads report honestly and web tools are hardened (#1813); cached-model defaults are chosen quality-aware (#1805); onboarding is clearer and first-launch catalog probes are gated (#1804, #1816); third-party license texts ship inside the .app bundle (#1812); Markdown tables read correctly under VoiceOver (#1822, #1824); and native XCUITest pixel coverage joins the desktop test stack (#1820).
- Release infrastructure. Installer checksums are version-stamped so cosign signing is idempotent (#1821), the mlx-compat import-order contract is restored in
bailing_hybridwith its guard test green (#1823), and two order-dependent full-suite flakes are gone (#1830).
- @loriz-art — unified macOS Settings UI and menu-bar branding (#1828)
- @MIt9 —
--disk-streamfor MoE routed-expert weights (#1803)
0.12.9 — 2026-08-10 · Muse Glimmer, natively
Meta released Muse Glimmer 30B and rapid-mlx serves it the same week — through its own vendored implementation of the architecture, with no external runtime dependency. The desktop app learns to hear and speak, and MCP tools arrive in desktop chat.
- Muse Glimmer 30B is served natively.
rapid-mlx serve muse-glimmer-30b-4bitruns Meta’s 29.6B dense reasoner on the standard text lane — sliding-window attention, 131K context — through a vendored backbone verified against the reference implementation to a 1.4e-5 max deviation on identical weights (#1802). The model thinks on a private channel and answers on another; rapid-mlx demultiplexes that wire intoreasoning_contentandcontent, and parses its ATEM tool-call envelope natively — streaming and non-streaming, including multi-turn tool results (#1791). Muse joins the release gate as the sixth agent-verified family (#1808), and the 4-bit weights (19.4 GB) are mirrored on the model CDN. The checkpoint’s vision tower ships in the weights but image input stays off until the multimodal path lands upstream. - The desktop app can hear and speak. A new Audio tab turns recordings into text and text into speech, in a voice you pick — all on your Mac (#1786).
- MCP tools reach desktop chat. The engine’s MCP support is wired through to the desktop app, so chat can call your configured MCP servers (#1787).
- Chat and Images keep more than one model warm. Budgeted multi-model residency lets the desktop hold a chat model and an image model at once instead of evicting on every tab switch (#1788).
- Models you already downloaded elsewhere are found. Weights another MLX runtime pulled are discovered and reused instead of downloaded again (#1785).
- Hybrid vision models stop crashing the batching engine. They are served through a serialized lane instead (#1798), and MLLM serves now expose their status metrics (#1781).
- Desktop paper cuts. An isolated first run uses cached models instead of re-downloading (#1801); overlapping memory confirmations no longer stack (#1800); math renders correctly in the shipped app (#1797); closing the main window behaves (#1795); the updater and image cache state were repaired (#1794); and a grounded answer that denies real-time access re-synthesizes once instead of shipping the refusal (#1784).
- @ryo1sato — MLLM status metrics (#1781)
- @xiaoxiunique — discover models other MLX runtimes downloaded (#1785)
- @osdodo — desktop Audio tab (#1786) and multi-model residency (#1788)
- @Jevin-F — MCP in the desktop app (#1787)
0.12.8 — 2026-08-10 · The desktop app makes pictures
A new Images tab renders locally, Chat reads the images you attach, and a wide band of engine correctness fixes lands underneath — including the packaging bug that would have shipped the Images tab unable to render anything in a downloaded copy.
- The desktop app generates images. The Images tab renders through the same one-model-per-process server the chat uses: pick an image model (
z-image-turboorflux2-klein-4b), load it through the usual readiness gate, prompt, and refine. Renders land in a filmstrip you can step back through, and selecting an older one restores the prompt that produced it (#1705). Two launch-blockers were caught before shipping: the bundled sidecar was missingmflux, so a downloaded copy could never actually generate (#1768), and the disk-space gate mis-counted already-cached component-layout weights as an impending download and refused to start (#1773). - Chat reads the images you attach. With a vision model loaded, attach an image to a message and ask about it; models that cannot read images show the attach button disabled with an explanation instead of failing later (#1723).
- The Qwen tool-call parser handles awkward arguments correctly. A series of fixes to
qwen3_coder_xmladdresses legacy raw string arguments whose own content contains XML-like closing tags — where the parser cannot tell an argument’s text from the wrapper around it. Some variants produced wrong arguments, others leaked wrapper framing into the answer or dropped the text after a call (#1730).AutoToolParser’s balanced-JSON scan was fixed alongside (#1726), and a replayed terminal chunk undertool_choice: autono longer duplicates content into the answer (#1711). tool_choice="none"is honored everywhere. A tool-trained model could emit a call — with a name and arguments that were never even declared — on the exact turn a client had turned tools off; the call is now dropped and the prose kept, across every parser and both stream modes (#1761). Literal<tool_call>or<think>tags in ordinary prose survive instead of being eaten (#1766, #1779), mid-conversationsystemmessages reach Harmony / gpt-oss models instead of being silently dropped (#1769), and a truncated chain-of-thought is marked as incomplete rather than shipped as the answer (#1770).- The first follow-up message no longer re-reads the opening context. The opening turn never saved a reusable cache boundary, so the second message paid to re-prefill everything; reuse already worked from the third message on. Measured on
qwen3.6-27b-4bitwith a ~9.9K-token document: the follow-up prefills 34 tokens instead of 9,941 — 1.45 s instead of 32.4 s (#1732). LFM2.5-2.6B gains bounded prefix reuse too: its alias saidis_hybrid: falsewhile the runtime found hybrid layers, so every turn re-prefilled the whole context (#1764). - The server survives sustained load and rude clients. The scheduler reclaims paged full-KV and free-block memory instead of wedging on a
D-METAL-CAP503 (#1646), a streaming client disconnecting mid-generation no longer leaks a running slot until the server wedges (#1782), andGET /healthreturns 200 for image-gen and video-gen serves instead of 500 for the life of the serve (#1783). - Reasoning-plus-tools turns get a real token budget. The desktop’s floor for those turns was set to exactly the default budget, so the
max()meant to lift it never lifted anyone; it is now 16,384 while Max Tokens sits at its default. A short budget does not fail loudly — it returns a cut-off answer that reads as a model that “could not do it” (#1722). - Desktop paper cuts. “Browse all models” opens the catalogue instead of closing the wizard (#1662); a stale last-served alias is validated before restore, so a removed model no longer fails the launch (#1729); returning to Chat after loading an Images model switches the server back instead of leaving the load button inert (#1739); the chat tab reports one throughput number instead of three, because prefill is no longer charged to tokens/second (#1728); web-page approvals gain “Always allow” with private and local addresses still blocked (#1695); and Claude Code gets an agent profile (#1720).
- The release process itself got trustworthy. App and engine are cut in one event instead of two that could drift (#1649); a release gate that had not run for eleven releases was found dead and repaired (#1671); the Codex review step fails closed on backend, auth, timeout or execution failure instead of passing silently (#1700); and four AX-only GUI golden flows run on every desktop PR — driving the app through the accessibility API with no screen recording, so they work unattended in CI (#1721, #1731, #1708).
0.12.7 — 2026-08-07 · One version number
From this release the engine and the Rapid-MLX Desktop app carry one version number, enforced in CI. 0.12.6 is skipped deliberately — the desktop app 0.12.6 that users already have was built before the work below landed, so reusing that number would have made it mean two different things.
- The two version numbers can no longer drift. They were maintained by hand in different files —
pyproject.tomlfor the engine, the app’sInfo.plistfor the desktop build — with nothing comparing them, so drift was the default rather than a risk. On 2026-08-07 it reached users: the engine was 0.12.5 while the app was 0.12.6, and both were correctly called “the latest release”. A check now fails any change where the two disagree, closing the chain end to end: git tag =Info.plist=pyproject.toml. - The privacy consent switch works. In Settings → Privacy, pressing Send anonymous usage data wrote the preference but left the switch showing its old value — a privacy control that appeared to refuse your choice. Measured by driving accessibility against real builds: 0 → 0 before, 0 → 1 after.
- Model recommendations are grounded in measurement. Every RAM tier now offers exactly two picks — faster and smarter — with measured ~8K peak memory and throughput from an M2 Pro 32 GB Mac mini. Two exceptions are stated rather than hidden: the smarter primaries at 64 GB and 96 GB+ exceed the benchmark host and are carried forward unmeasured; their faster alternatives are measured.
- “Untested” no longer looks like a score of zero. A model with no compatible benchmark rendered an empty dashed track, so a missing measurement and a poor result were indistinguishable. The sweep is reproducible now too — persisted prefix-cache entries were inflating 8K prefill from ~313 to ~20,701 tok/s between reruns.
agents --setupcan no longer eat the operator’s config. The release gate runs it on machines that are also someone’s daily driver. It used to back up and restore the real file, which failed on SIGKILL and — worse — faithfully restored a file that had already been damaged, leaving a developer’s codex pointed at a local server for weeks. It now honoursCODEX_HOME/HERMES_HOME, and the gate refuses to run against the real config directory.- A GUI golden-flow suite that finds controls by accessibility identifier, not screen position — six journeys plus three invariants, with committed structural snapshots, so a control that vanishes, is renamed, moves in the hierarchy, or silently disables becomes a reviewable diff. The snapshots discard coordinates by design, so they say nothing about visual layout.
- Model tabs consolidated into a single Model Management surface; syntax highlighting and markdown tables in chat; the tray’s Check for updates… reports a result instead of nothing; a machine that fails the live memory guard is offered an honestly-labelled smaller model rather than a chooser whose smallest option is the one that just failed.
Desktop app 0.12.6 — 2026-08-07 · Rapid can look things up
Bundles the 0.12.5 engine. The headline is a set of built-in tools the model can reach mid-answer — with you deciding what it is allowed to touch.
- Weather, web search, and reading a page. Ask about something recent and the model can go and find out. Each lookup shows a card you can open to see exactly what it asked for and what came back. Fetching a page asks permission first and names the site — and “Don’t allow” is a normal answer, not an error. Search works out of the box; the backend is switchable in Settings → Tools, and private or local addresses are refused.
- Conversations can be pinned, renamed and archived. Right-click a chat in the sidebar or use the
···button on hover. Pinned chats get their own group at the top; archived ones collapse into a group you can reopen. Renaming a chat stops Rapid from re-titling it afterwards. - Deleting a conversation asks first, names the chat, and says plainly that it cannot be undone.
- Every message has its own actions — copy any message, edit a question you already sent, regenerate an answer you did not like.
- Models that can never chat are no longer offered. The picker was listing entries a conversation could never use.
0.12.5 — 2026-08-07 · Tool calling tells the truth
Ten defects across the parsers, the streaming path and the prompt builder, all the same shape: what the model meant to call was not what got called, or not what you saw.
- Every streamed string argument could come back wrapped in quotes. On the streaming path only, formatting whitespace arriving before a JSON value’s opening quote made the Qwen3-Coder XML parser read the wrapper quotes as argument bytes.
browsegot"https://example.com"— quotes included — and rejected it as not a URL; a file read got"/path/to/file"and was told it does not exist. Non-streaming was unaffected, which is exactly what made it read as a weak model rather than a parser bug: same model, same prompt, onlystreamdiffering — 4/4 correct without it, 5/5 corrupt with it. Found from the other end, as five stable failures in the Hermes agent suite (PR #1600). agents codex --setupno longer deletes your Codex config. It rewrote~/.codex/config.tomlfrom a template instead of merging, so anyone who had customised Codex lost it the first time they pointed it at Rapid-MLX. It now deep-merges, backs the original up first, and copies metadata through the destination descriptor rather than by pathname — so the backup cannot be diverted by a symlink swapped in mid-write (PR #1539).- Tool-call arguments survive the round trip. Seven defects, one shape — a value the model emitted did not reach the tool: values mangled converting between wire formats (#1518), XML delimiters stripped out of string arguments (#1552), undeclared parameters accepted ambiguously (#1551), Qwen truncating its own arguments at the length cutoff (#1574), streamed arguments unvalidated on some parsers (#1559), Nemotron dropping the declared-tool gate while streaming (#1540), and forced-stream tool wire leaking into the response (#1546).
- LFM models no longer print their tool calls at you. On the streaming path LFM2.5 emitted raw markup into the visible answer while the call itself parsed fine, so it looked like a display glitch. It was not: for parsers without native tool-format support the engine serialises prior calls into the prompt as
[Calling tool: name({args})], the model imitates its own transcript, and the parser only knew the pythonic dialect. 10/10 before, 0/10 after (PR #1592). - Concurrent chat and tool calls no longer kill the engine loop. With both in flight,
PromptProcessingBatch.extendwroteNoneinto mlx-lm’s per-slot logits-processor list and took the scheduler down — reproducibly, in 3 of 5 runs (PR #1526). rapid-mlx lswas reporting cached sizes at roughly double reality. It walked the Hugging Face cache following symlinks, counting each blob once underblobs/and again through everysnapshots/link. Every “reclaim space” figure inherited the error, including the ones the Mac app renders — Ternary Bonsai 27B read 15.87 GiB against a true 7.9 GiB (PR #1584).- Ministral 3 is withdrawn. It could return no reply at all on some Macs; the alias is gone rather than left in the picker. Model count is now 215.
Desktop app 0.12.5 — 2026-08-05 · An ordinary maths question could take the app down
A desktop-only release — the engine stays at 0.12.4, which this build bundles. Two of these shipped in front of every new user: a reply containing a formula quit the app outright, and the starter model everyone landed on fell apart on the kind of question people try first.
- Maths no longer crashes the app. When a reply contained a formula the app quit with a macOS crash dialog, and a plain
2 + 2 = 4was enough to trigger it. Formulas now render as their plain-text source rather than taking the app down; proper typesetting is still to come. - The starter model was unusable and has been replaced. The pick that shipped with 0.12.1 came apart on ordinary multi-step questions — doubling words together, then looping until it ran out of room. Measured on the same Mac against the same question: 0/4 correct before, 16/16 after, with the first answer arriving in about a second. Anyone still on the retired starter is walked back through first-run setup rather than left on it (PR #1484).
- A silent server no longer hangs the chat. Sending a message used to sit there indefinitely if the model stopped responding. It now fails after a bounded wait and the message stays retryable.
- Loading a model that will not fit is blocked instead of freezing the Mac. The check reads free memory at that moment and warns, rather than letting the load proceed into a freeze or a kernel panic.
- Your API key is no longer shown in the clear in the copyable setup snippets on the Launch page.
- Cursor no longer gets a configuration that cannot work. Cursor routes requests through its own servers, so a
localhostaddress was never reachable from it. The app now says so and points at Claude Code, Cline, or Continue for a genuinely local connection. - Two models that could fail to answer are hidden for now. Ministral 3 and Gemma 4 E2B looked like ordinary small chat models in the picker, but on some Macs they returned no reply at all. They stay hidden until that is fixed.
- Quitting no longer stalls or crashes, and launching the app no longer risks shutting down a
rapid-mlx serveyou started yourself in a terminal. Conversation history keeps its order, the compose box grows with what you type, and the transcript stops yanking you to the bottom while you read further up.
0.12.4 — 2026-08-04 · Caches that were quietly not caching
Four of these are the same shape: a cache that reported success while storing nothing, or storing the wrong thing. None of them raised an error — they just made every turn pay full price. Plus an 8 GB Mac finally gets an answer instead of a rejection.
- An 8 GB Mac gets a recommendation instead of a rejection. Every previous pick was refused by auto-start on the smallest Macs, so the tier existed only to say no. It now recommends
lfm2.5-2.6b-4bit— about 2.0 GB resident, fast, and explicitly not a coding model. Serving it required teaching the alias resolver about Hugging Face repos that ship every quantization in a subdirectory of one repo:LiquidAI/LFM2.5-2.6B-MLXholds eight, so a naive pull fetched ~20 GB to use 1.6 GB. Thecurl | bashbanner now recommends exactly what the desktop app recommends, with one test that fails if the two ever drift apart (PRs #1442, #1443, #1447). --enable-prefix-cachewas a silent no-op on the dense Qwen 3.5 / 3.6 aliases and Ternary Bonsai. These models carry non-trimmable recurrent layers whose state is reusable only through the bounded snapshot path — but the auto-default that enables that path keyed on a hybrid flag these models are deliberately pinned off, so it skipped exactly the models that needed it. Every store was dropped and a multi-turn tool agent re-prefilled its whole accumulated context every turn: turn-2 TTFT 22.3 s onqwen3.5-9b-4bit. Routing is unchanged; only the reuse path is switched on (PR #1445).- A three-token drift at a message boundary could discard a 40K+ reusable prefix. Tokenization is not compositional: appending the next assistant or tool segment can change the previous prompt's final few byte-BPE tokens. A non-trimmable cache cannot trim to absorb that, so the whole snapshot was thrown away. An 8-token replay window covers the drift observed on Codex turns at negligible re-prefill cost (PR #1449).
- Shutdown persisted the shallowest prefix first. The save has a short deadline and may commit only one entry; LRU order could start with a one-token bootstrap entry, spend the single guaranteed slot on it, and skip the real 100K+ frontier. Deadline-aware saves now go longest-first (PR #1451).
- The repetition guard answered a runaway loop with 503 and discarded the partial output. Worse, 503 is the same code as a genuine Metal-runtime abort — where “retry smaller” is correct advice — so agent frameworks read it as a transient outage and re-sent the identical prompt straight back into the loop. The guard stop is now distinguishable from a runtime abort and returns the tokens already generated. Measured on
qwen3.5-4b-4bit, a 30-turn tool-agent loop hit this twice; larger models, not at all (PR #1450). - DeepSeek V4 long decodes could exhaust Metal's resource count while memory looked healthy. The V4 cache updates functionally, so without realizing each forward's leaf arrays the decode retained one live lazy-graph chain per layer per token — hitting the fixed
499000resource limit with plenty of bytes free (PR #1448). - DeepSeek V4 could reopen thinking after it had been turned off, re-entering the reasoning parser mid-turn and leaking reasoning into content. Prompt shaping cannot enforce this; it is now suppressed at decode time. V4 also now accepts the
r:-prefixed DSML tool aliases it actually samples, instead of forwarding them as malformed calls (PRs #1452, #1446). - The desktop app's memory column was two different metrics wearing one hat. Some rows were a bare-
mlx_lmallocation high-water mark, others were whatrapid-mlx serveactually uses — andservequantizes the KV cache to int4 by default, worth roughly 2 GB on a 27B. Re-measured throughserveon one machine, the tier table, the hardware-tiers doc and the blog now agree (PR #1453). - The first text-to-speech request could 500 when Kokoro's spaCy G2P model had not been resolved yet; it is now pre-resolved at the route gate (PR #1254).
- Benchmark artifacts carry a schema version and a methodology hash, so results produced under different methodologies can no longer be silently aggregated together (PR #1455).
0.12.3 — 2026-08-04 · A picker that answers “will this run on my Mac?”
- Gemma 4 checkpoints stopped serving at all. mlx-lm 0.31.x dropped
mlx_lm/chat_templates/gemma4.py, whilemlx-community/gemma-4-26b-a4b-it-4bitstill declares"chat_template_type": "gemma4"in its tokenizer config. mlx-lm imports that module by name with no guard, so the import raisedModuleNotFoundErrorandrapid-mlx servedied at startup with no fallback to the checkpoint’s ownchat_template.jinja— a regression against 0.6.71, which still shipped the module. Weights load before the tokenizer inmlx_lm.load, so a catch-and-retry would re-read multiple gigabytes; the offending field is neutralized up front instead (issue #1420, PR #1439). - The desktop app recommends a model by your Mac’s memory. The old five-role matrix (Coding / Chat / Vision / …) asked you to classify yourself before running anything. It is replaced by one RAM-tier table: per tier a smart pick — the most capable model that fits — and, where a genuinely faster model is worth a second card, a fast alternative. About 1,500 lines lighter (PR #1437).
- The bundled app moves to 0.12.1, a signed release carrying engine 0.12.1 (PR #1433).
0.12.1 — 2026-08-03 · A desktop app, and a Gemma 4 page that was quoting numbers nobody measured
- A menu-bar Mac app, in the open-source repo.
apps/rapid-macis an Ollama-style local-LLM app under Apache 2.0 — the same engine with a one-click interface, download-only, with therapid-mlxname reserved for the engine itself (PRs #1406, #1413, #1427). - The Gemma 4 26B page was quoting fabricated speculative-decoding numbers. They were replaced with measured ones, alongside a corrected MoE flag and a serve guide sized for a 32 GB Mac (PR #1429). A companion recipe covers Gemma 4 12B on an 18 GB Mac (PR #1423).
- SuffixDecoding is no longer a tax on traffic it does not suit. Its floor on free-form generation went from −32% to −6.7% (gemma-4-12b-4bit, M3 Pro) with the code-edit win intact, and M2 Pro’s high-overlap case moved from exactly 1.00× to +13%. It still does not clear the bar for defaulting on, so it stays opt-in — the point is that turning it on is no longer a gamble (PR #1419).
- Gemma 4 gets live KV quantization, a working bench path, and an exposed KV-projection override (PR #1408).
- Long-context prefill adapts to memory pressure instead of pushing through it (PR #1410).
pip install rapid-mlx[all]now installs an audio stack that actually works — the extra was pulling a combination that had never been validated together (PR #1421).- mlx / mlx-lm / mlx-vlm version bounds are capped, and moving one is gated on a full-family output-coherence sweep. An upstream heuristic change had previously shipped garbage generations (issue #1248, PR #1400).
- Community benchmarks submit over HTTP rather than by opening a GitHub PR (PR #1403).
- Codex and DeepSeek V4 engineering turns are steadier — bounded action priming, evidence-aware retries, recovery when a test command is unavailable, and a hardened DSML tool protocol (PRs #1402, #1407, #1417, #1418, #1424, #1428).
0.11.9 — 2026-08-02 · Requests that hang, and requests that lie
Every fix here is a case where the server failed to give an honest answer: it either never came back, or it reported success on output it had already judged bad. Nothing new is added; seven ways to be misled are removed.
- A streaming tool call could wedge a connection indefinitely. Once the tool parser sees an opener (
<tool_call>,<function=), the post-processor withholds every delta until the block closes — and that suppression had no bound. A block that never closed held all content forever, so the SSE generator emitted nothing but: keepalivecomments and no timeout at any layer could fire: the reported request generated 2,207 tokens, put 1 chunk on the wire, and ran 518 s — outliving its own 300 s client timeout. From the client it was indistinguishable from a server that never answered. Withheld bytes are now capped at 64 KB, after which the text is released as content and tool suppression latches off for the rest of the turn. The cap only applies before any tool call has reached the wire, so genuine incrementally-streamed tool calls are never touched (issue #1359, PR #1391). - A forced tool call no longer fabricates arguments it knows are invalid. Under
tool_choice: "required"or a named function, when the parser surfaced no call the route synthesised one to honour the tool-call-guaranteed contract — witharguments: "{}". If the target tool's schema declared required properties, that was a call the server already knew was schema-invalid, handed to the client underfinish_reason: "tool_calls"for it to parse and execute. Unmet required properties now return 422 on chat non-stream and on both Responses sites; streaming cannot 422 after headers are sent, so it finishesstoprather than fabricating. A tool with no required fields still synthesises{}exactly as before (issue #1256, PR #1394). - The repetition guard reports a failure instead of a success. Running Codex against a local DeepSeek V4 build, the model emitted
ambiguover and over. rapid-mlx correctly stopped it at 304 completion tokens — then loggedfinished normally, so Codex emittedturn.completedand exited successfully without changing a single file. Guard termination is now an aborted generation carrying an explicit error (PR #1396). - Short degenerate loops are caught before they exhaust Metal. A 46,475-token Responses request fell into a repeated short CJK loop, generated 10,844 completion tokens, and hit
[metal::malloc] Resource limit (499000) exceededaround scheduler step 12,032 — and the recovery path emitted a terminal output with noerror, so the request was reported as a success. Sustained 1–5 token loops are now detected, with a conservative 256-repeat floor, reading only the trailing 768 tokens regardless of context length, and still restricted to tool-bearing requests. Scheduler-local Metal failures now propagate throughRequestOutput.error(PR #1393). - Embeddings no longer leak Metal buffers.
/v1/embeddingsnever calledmx.clear_cache(), unlike every LLM path in the engine — and becausepadding=Truemakes each batch a different sequence length, MLX's size-keyed allocator pool almost never had a reusable block. It only grew: roughly 70 MB retained per input text, 2.3 GB → 24 GB over 320 texts, and about 50 GB in a week-old production process. Buffers are now released after each batch on both the string and pre-tokenized paths (issue #1380, PR #1390). - A warm start is no longer an outage. The persisted prefix cache was loaded synchronously between engine start and readiness, so
/health/readyand/v1/modelsreturned 503 for the entire multi-second read — an orchestrator gating traffic on readiness saw a warm start as downtime. Readiness now flips first and the cache warms in the background; shutdown cancels a load still in flight. This is safe because each entry installs as a single atomic swap under its own lock, so an early request either misses and recomputes (always correct) or hits a fully-installed cache — never a half-populated one (issue #1350, PR #1392). - ffmpeg is found when it is not on
PATH. Video remux and crop now resolve the binary fromFFMPEG_BINARY, then the processPATH, then the usual Homebrew and system locations — so a GUI-launched process with a minimal environment completes instead of failing at the last step (issue #1352, PR #1397).
0.11.8 — 2026-08-02 · Embeddings stop truncating silently at 512 tokens
/v1/embeddings hardcoded the tokenizer at max_length=512. Anything longer came back HTTP 200, correctly shaped, and quietly missing its tail — which degrades a vector index with no signal anywhere. Reported by a user indexing code chunks with Qwen3-Embedding-4B, whose architecture supports 32K (issue #1381).
- The limit is now derived from the model.
autoreadsconfig.max_position_embeddings, thentokenizer.model_max_length, guarding Hugging Face's large "unset" sentinel. It falls back to 512 only when the model declares nothing at all — and logs that it did. --embedding-max-length auto|<int>— an operator ceiling for a lower memory or service limit. An explicit value is validated against and clamped to the model maximum.--embedding-overflow-policy truncate|error—truncate(the default) still discards the tail, but logs a warning and incrementsrapid_mlx_embedding_truncations_totalon/metrics, so it is no longer silent.errorreturns a structured 400 withcode: "input_too_long"carrying the observed and allowed token counts.- Both apply to string and pre-tokenized inputs alike, and the usage block no longer over-reports the pre-truncation token count (PR #1386).
0.11.7 — 2026-08-02 · Test-only
One change, and no behaviour change: proof that the RAPID_MLX_BASE_URL release guard (G7) fails loudly rather than passing vacuously — driven from bash across the boundary, with the exit code shadowed so a silent skip cannot read as a pass (PR #1384). Released so the published tag matches the tree the gate ran against.
0.11.6 — 2026-08-02 · DeepSeek V4 gets its own drafter, and a guard against runaway output
Four DeepSeek V4 changes that compound: the checkpoint's native speculative-decoding heads are finally used, its reasoning is parsed by a parser that understands it, a repeated prompt stops re-prefilling from scratch — and a model that falls into a loop is stopped instead of running until the client gives up.
- DSpark — DeepSeek V4's checkpoint-native speculative decoding. DeepSeek ships three MTP stages plus low-rank Markov heads inside the V4 Flash 0731 checkpoint; rapid-mlx now drafts and verifies with them instead of ignoring them. Turn it on with
--speculative-config '{"method":"dspark","num_speculative_tokens":5}'. Detection is fail-closed: the block geometry is read from the checkpoint's owninference/config.json, and every required tensor is checked against the safetensors index — so a stripped conversion refuses to start rather than decoding wrong.num_speculative_tokensmust equal the checkpoint'sdspark_block_size(5 on the 0731 checkpoint); a partial block is rejected. On real Codex/v1/responsestool rounds, 2.6–2.9 tokens were accepted per round. One limitation worth knowing before you try it: DSpark reads the checkpoint directory directly, so today it needs a local path —rapid-mlx serve /path/to/DeepSeek-V4-Flash-0731-MXFP4-MLX— not thedeepseek-v4-flash-0731-mxfp4alias, which resolves to a Hugging Face id. See perf flags (PR #1379). - A DeepSeek V4 reasoning parser. V4 was being parsed by the generic DeepSeek-R1 parser, which does not understand V4's protocol — so part of the model's internal scratch reasoning was emitted as user-visible content.
reasoning_parser: deepseek_v4now backs all four V4 aliases, wired request-aware through Chat, Anthropic and Responses. An explicit-thinking stream keeps its reasoning in the reasoning channel with no<think>leakage into raw SSE, reasoning or content (PR #1383). - Warm TTFT on a repeated prompt: ~24.6 s → 0.318 s. An exact-repeat 8,131-token prompt was still paying almost the full cold cost, because an unusable full-length cache snapshot masked a usable 8,113-token message-boundary one. An exact non-trimmable hit now falls back to the longest usable strict prefix instead of silently full-prefilling. Measured on the real MXFP4 checkpoint, with byte-identical output:
| DeepSeek V4 Flash 0731 · MXFP4 · 8,131-token prompt, 124 completion tokens | TTFT | Decode |
|---|---|---|
| Cold | 25.830 s | 25.81–26.11 tok/s |
| Warm — before | ~24.6 s | — |
| Warm — after | 0.318 s (81.2× vs cold) | 25.81–26.11 tok/s |
- The prefix cache survives a restart. Persisted V4 prefixes are reloaded and reused across server restarts, so the warm path above is warm on the first request after a restart rather than the second.
- A runaway agent no longer generates until the client gives up. A real Codex + DeepSeek V4 Flash session repeated the same investigation sentence for 4,922 tokens over 234 seconds, until the client disconnected — the existing post-completion coherence telemetry can only notice that after the fact, never stop it mid-flight. A token-level guard now checks a bounded suffix every eight decode tokens and stops exact periodic output, exposing
rapid_mlx_repetition_loop_stops_total. It is deliberately conservative — at least 72 repeated tokens, and restricted to tool-bearing agent requests, so plain chat is untouched (PR #1377). A follow-up made the copy threshold adaptive to loop length: a long periodic block stops after three copies while a short phrase still needs proportionally more, because a real 61-token paragraph loop was letting 424–464 tokens reach the UI before the guard fired (PR #1378). In 0.11.9 this guard stops reporting success — see above.
0.11.5 — 2026-08-01 · DeepSeek V4 Flash 0731, and a canary for bad output
- DeepSeek V4 Flash 0731 is served. The 0731 checkpoint gets a first-class alias,
deepseek-v4-flash-0731-mxfp4— the most capable open-weights model in the registry, and the cheapest 50-point model on the Artificial Analysis index (PR #1364). Serving it took three follow-ups: mixed-length batches (PR #1365), a stabilization pass (PR #1369), and making prefix reuse actually engage on V4 (PR #1371), where it had been silently doing nothing. - DeepSeek DSML markup no longer leaks into Responses streams (PR #1373).
- An output-quality canary. Request telemetry can carry
output_degenerate, a boolean computed locally that flags runaway repetition — so a bad quantization surfaces as a version-scoped spike instead of scattered bug reports. The check runs on your machine and only its yes/no answer is sent; no prompt, no completion (PR #1250, #1266). Documented on the telemetry page. - Small-model GPU smoke on free CI runners — coherence plus tool-calling on Qwen and Llama, catching the cheapest class of regression without waiting for the Studio.
0.11.4 — 2026-07-31 · Controls for the new multimodal lanes
A follow-up to 0.11.3 that makes the audio and video surfaces controllable, discoverable and reproducible — no new models, just the knobs the new lanes were missing. Every change here was exercised end-to-end on the Studio (M3 Ultra) against the published wheel.
- Video motion controls.
POST /v1/videosnow acceptsguidance_scale,negative_prompt, explicitfps/frames, and (LTX image-to-video)conditioning_strength— each validated against the served backend, so an out-of-range value returns a clear 400 instead of a bad clip (PR #1340). - Discoverable limits.
GET /v1/videos/capabilitiesreports the served backend's real bounds — the size range and multiple-of-64 rule, the8n+1(LTX) and4n+1(Wan) frame shapes, native fps, the pixel-frame workload ceiling and the reference-image caps — so a client can size a request without guessing (PR #1343). - Audio output format.
/v1/audio/speechand/v1/audio/musictakesample_rate(8k–96k) andchannels(mono or stereo); the returned WAV is resampled and up- or down-mixed to match, and echoes the actual rate and channels inX-Audio-*response headers (PR #1346). - Reproducible designed voices. Qwen3-TTS VoiceDesign accepts a
voice_seed, so a voice you described in natural language comes back byte-for-byte identical on the next call — the missing piece for keeping one narrator across a project (PR #1347). - LTX videos come back video-only. The LTX checkpoint is an audio-video one, but its audio track was silent; that empty track is now stripped so the MP4 is a single video stream instead of reading to downstream tools as "this clip has sound" (PR #1349).
- Docs caught up. The audio guide had five whole model families missing (PR #1345), and the video guide now states plainly that these backends return no audio (PR #1353).
0.11.3 — 2026-07-30 · Audio grows up, and video ships
The largest single jump in the model registry so far: 191 → 214 aliases. Eleven new TTS/STT aliases, eight video-generation aliases, and a whole new modality behind an asynchronous jobs API. Version 0.11.2 was cut but never released: its Tier-1 agent gate failed. 0.11.3 is 0.11.2 plus a first attempt at the fix (PR #1341) — and its gate failed too. The defects turned out to be in the gate harness itself, not in the release payload: the harness ran the Hermes setup without --base-url, leaving its context at the 32K fallback below the 64K minimum Hermes requires to start at all, and its per-agent time budget was too tight for the 35B hybrid gate model's cold shader compile. The payload was then verified by hand on the Studio (M3 Ultra) — Claude Code, Codex and Aider pass, and Hermes passes once its context is set correctly — and 0.11.3 shipped through the documented emergency-release path, with the bypass and its reason stamped into the GitHub release notes. The harness has since been fixed.
- Video generation, a new lane.
POST /v1/videosfollows OpenAI's job-based Videos API — submit, poll, download the MP4. Three backends ship: Wan 2.1 / 2.2 with four converted checkpoints, native frame-rate handling,4n+1temporal-shape enforcement, per-checkpoint pixel-area ceilings and LoRA support (PR #1322); CogVideoX-Fun with its runtime bundled so no source checkout is needed (PR #1313, #1329); and MLX-native LTX-2.3 (PR #1298). Jobs are serialized on purpose — two diffusion pipelines resident at once will exhaust unified memory. Needs Python 3.11+ and ffmpeg; core text and audio stay on 3.10 (PR #1325). See the video family page. - Voice cloning, four ways. Qwen3-TTS Base clones from a reference clip plus its transcript, so a channel can keep one branded narrator (PR #1305). IndexTTS 1.5 clones from the clip alone — no transcript — which makes it the least fiddly path (PR #1318). F5-TTS (pure MLX, no torch) closes the Chinese expressive-cloning gap that Qwen3-TTS reads flat on and Chatterbox can't reach, being English-only (PR #1297, #1304). Chatterbox gained zero-shot cloning and an exaggeration control (PR #1303).
- Qwen3-TTS VoiceDesign — describe a voice instead of picking one. Timbre, gender, age, accent, emotion and prosody all come from a natural-language
instructionsstring; there are no named speakers at all. That is a strictly richer surface than CustomVoice'sinstructions, which only modulates a fixed speaker (PR #1309). - Forced alignment — timings without recognition error. Give
Qwen3-ForcedAligneraudio plus the transcript you already have and it returns per-character start/end times. Because it never guesses at words, it cannot mis-hear them — which is exactly what karaoke captions and beat-synced editing need, and it was unoccupied on MLX (PR #1301). The lane was then hardened: blocking work moved off the event loop, the aligner given its own model cache so it stops evicting the ASR model, and corrupted uploads no longer misreported as bad requests (PR #1317, #1327). - SenseVoice — fast Asian-language ASR. FunAudioLLM SenseVoice Small (~234M, non-autoregressive CTC), strongest in the registry on Chinese, Cantonese, Japanese and Korean, and it emits per-segment emotion and audio-event tags alongside the transcript (PR #1308).
- Word-level timestamps on transcription.
/v1/audio/transcriptionsnow returns per-word timings (PR #1299). - Local text-to-music. A vendored MLX Stable Audio 3 engine behind
POST /v1/audio/music(PR #1307, #1316). - Kokoro stops taking the worker down. A broken espeak G2P setup used to crash the audio worker or return 500s and crowd out other candidates; it now degrades and reports instead (PR #1312, #1314). VoxCPM's advertised Chinese support was withdrawn — it did not work, and claiming it was worse than not having it (PR #1302).
- Release gating got stricter and better calibrated. A KV-quant differential quality gate with a chip-tier classifier (PR #1289), DeepSeek-R1 excluded from the coherence sweep where it was a known false positive (PR #1324), the Hermes gauntlet profile stabilized (PR #1330), and the G12 random-coverage gate recalibrated (PR #1332).
- Dead code removed. The unused vLLM platform prototype (PR #1288), unused cloud routing (PR #1290) and legacy branding on public surfaces (PR #1292) are gone.
rapid-mlx modelsnow shows download size per model (PR #1293).
0.11.1 — 2026-07-28 · Correctness & first-run polish across the 0.11 line
Community contributors:
@pierre427 — DeepSeek-V4 explicit YaRN attention_factor
(#1225) and containing local model_file imports to the model root
(#1226).
- Live KV-cache quantization now engages on vision models too. The
--kv-cache-dtype int8/int4path probes the nested language-modelhead_dim, so multimodal servers get the same steady-state KV savings as text models; a fail-safe keeps bf16 wherever no supported group size fits (PR #1208, #1231). - Qwen3.6-35B-A3B serves coherently again. An upstream mlx-lm change had applied a spurious +1.0 RMSNorm shift that garbled this checkpoint; it is now undone (PR #1234).
- MTP no longer returns intermittent empty responses. The Qwen3.6 MTP sidecar path could occasionally emit nothing; sidecar extraction and eligibility checks were hardened (PR #1201, #1214, #1215).
- A quieter, clearer first run. Cold-model downloads show a TTY-aware single-line progress bar with the correct starter size; the unauthenticated-download advisory and leaked Hugging Face progress bars are silenced; and cold-start dead air is gone (PR #1237, #1259, #1260, #1263).
- MCP config accepts the standard
mcpServerskey. Following the standard examples previously resolved to zero servers; the standard key is now honored (PR #1243). - Output-coherence release gate. Six blocking golden checks plus an advisory garbage detector now guard every release across the model fleet (PR #1247, #1262).
0.11.0 — 2026-07-24 · Quantized live KV cache — int4/int8 on the continuous-batching cache
Community contributor: @66Ton99 — Codex Responses long-context handling (#1141).
- The live continuous-batching KV cache is now quantized.
--kv-cache-dtype int4(the default) andint8previously only shrank the retained prefix cache — the live decode cache stayed bf16, so a long-context or high-concurrency batch still grew its KV memory at full precision (and--disable-prefix-cachesaved nothing). The continuous-batching cache is now quantized with dequant-on-read, cutting steady-state KV memory on long-context and multi-request serving. Hybrid, sliding-window and MLA-latent caches stay bf16 where no supported group size fits, so the path is safe to leave on by default (PR #1197, #1199). - Forced tool-call arguments are now grammar-constrained on reasoning models too. A forced or named tool call (
tool_choice="required"or a specific function) on a reasoning model previously opted out of the decode-time grammar while it was inside its thinking budget, leaving the arguments unconstrained. They are now hard-constrained like every other tool call — closing the last gap in the #558 tool-call-integrity work: forced or free, thinking or not, a tool call is guaranteed parseable (PR #1192, #558 line①).
0.10.18 — 2026-07-24 · Bounded prefill memory on long prompts
- Long text-only prompts no longer spike memory. The multimodal batch generator was pushing the entire prompt through the language model in a single forward and projecting logits over every position, peaking around 35 GB on a ~20k-token prompt — enough to max out a 48 GB Mac. Text-only prefill is now chunked to bound peak memory (PR #1193, fixes #1187).
0.10.17 — 2026-07-22 · Base-wheel VLM serve + install-hint fix
- Hybrid-VLM checkpoints serve cleanly from the base wheel. Serving a hybrid vision-language checkpoint without the
[vision]extra now skips the vision path instead of erroring, and the Gemma-4 install hint points at the right extra (PR #1179).
0.10.16 — 2026-07-22 · Grammar-constrained Gemma-4 native tool calls
- Gemma 4 native tool calls are grammar-constrained. Gemma 4's native tool-call format is now constrained by a grammar at decode time, extending guaranteed-parseable tool-calling to the Gemma-4 family (PR #1171, #558 E4).
0.10.15 — 2026-07-21 · Grammar-constrained tool calling, out of the box
llguidancepromoted to a core dependency. The default-on grammar-constrained tool-calling from 0.10.14 now works with zero extra setup — the grammar engine ships with the base install instead of being an opt-in extra, so guaranteed-parseable tool calls are on out of the box (PR #1146, part of #558).
0.10.14 — 2026-07-21 · Default-on grammar-constrained tool calling
- Grammar-constrained tool calling, on by default. Tool-call generation is now constrained by a grammar at decode time, so the model cannot emit a malformed tool call — the structured output is guaranteed parseable rather than parsed-and-hoped. Includes an automatic path that engages it without any configuration (PR #1143, part of #558).
0.10.12 — 2026-07-17 · Response cache + trim-free prefix reuse
Community contributors: @kumosan2 — hybrid prefix reuse (#1111); @romanbsd — nested Gemma 4 tool arguments (#1102); @pierre427 — pflash compression visibility (#1106).
- Prompt-deterministic response cache — opt-in exact-match short-circuit. New
--response-cache-entries Nretains up to N fully-computed greedy (temperature 0/top_k 1) chat completions; a completely repeated request returns the stored completion verbatim with zero GPU decode — the second and later identical calls are effectively free. The cache key spans every output-affecting field (model id, the canonical messages + chat kwargs, resolved sampling kwargs,response_format/logprobs), so any change is a clean miss and recompute; sampled requests are never short-circuited. It is distinct from and complementary to the prefix / KV cache — the prefix cache reuses prefill state and still decodes, this returns the entire stored completion and decodes nothing. New Prometheus countersrapid_mlx_response_cache_hits_total/_misses_total. Ideal for agent retry loops, eval harnesses, and dashboard / health probes that re-issue an identical request. Default0= fully disabled, request path byte-for-byte unchanged (PR #1123). - Trim-free prefix reuse for hybrid & sliding-window models — opt-in. New
--hybrid-cache-entries Nretains up to N non-trimmable prefix-cache entries so a stable prefix + a new suffix each turn reuses prior prefill instead of recomputing it. Covers both hybrid recurrent-state (GatedDeltaNet / Mamba) and sliding-window (Gemma 4, GPT-OSS) models — best for stable-system-prompt / long-context agent workloads. An exact re-request of a rotated sliding-window prompt safely falls back to a full prefill (byte-equal to cold). Default0= disabled (PR #1111, #1124). - Gemma 4 nested tool arguments. Correctly parse Gemma 4 tool calls whose arguments contain nested objects (PR #1102).
- pflash compression visibility. Surface the previously-silent endpoints-only compression collapse so a degraded pflash configuration is no longer hidden (PR #1106).
- Dependency hygiene. Exclude the broken
mlx-vlm 0.6.4from all extras (PR #1119).
0.10.10 — 2026-07-15 · Ternary Bonsai 27B + Qwen3-Coder-Next 80B
- Ternary Bonsai 27B — new flagship small-footprint model.
bonsai-27b-2bitis a 2-bit ternary Qwen 3.5-class 27B that packs into 7.9 GB and runs on a 16 GB Mac (~46 tok/s on M3 Ultra). The quant is stock MLX 2-bit affine — no custom kernel — and it loads through the mlx-lm text path (served text-only via theis_text_onlyprofile flag to sidestep an mlx-vlm SSM bug, even though the config declares a vision tower). Strong on the mainstream: code, math, reasoning, EN/ZH general writing, and agent/framework tool-calling (verified end-to-end with the OpenAI SDK, LangChain, pydantic-ai, smolagents CodeAgent, and Aider). Known limitation: strict Chinese classical regulated verse (五言绝句 with fixed rhyme) can trip a repetition loop — a niche edge case. Weights are on the R2 mirror. - Qwen3-Coder-Next 80B — new coder family SKUs. Added
qwen3-coder-next-80b-4bit(44.8 GB) andqwen3-coder-next-80b-8bit(84.7 GB), the 80B-A3B Coder-Next MoE at 4-bit and 8-bit MLX with theqwen3_coder_xmltool-call parser. HF-pull (not yet on the R2 mirror). See the alias reference. - KV-cache export / import. Serialize a warmed prompt-prefix KV cache to disk and re-hydrate it on a later serve, so a long shared system prompt / document context is paid for once rather than re-prefilled every session.
- Gemma 4 hardening. Correctness and stability fixes to the vendored Gemma 4 text path.
- Release-infra hardening. Toughened the auto-release / version-bump pipeline so version markers and the PyPI + Homebrew publish steps stay in lockstep.
0.10.9 — 2026-07-12 · share serve-flag passthrough + MTP K=3 default
- Systematic serve-flag passthrough over
share.rapid-mlx share <model> -- <serve flags>now forwards everything after a literal--verbatim to therapid-mlx servethat share spawns (the git / cargo / kubectl end-of-options idiom), so every serve flag works over the tunnel without share re-declaring each one. A prefix-aware denylist keeps share-owned flags (--host/--api-key/--port/--listen-fd/--log-level) from being forwarded, preventing accidental LAN re-exposure or API-key override (PR #1096). - Spec-decode over share is opt-in, MTP K=3 by default. Speculative decoding stays off by default across a share; opt in per-share with
-- --force-spec-decode --speculative-config '{"method":"mtp"}'. When--force-spec-decodeis set, MTP defaults to K=3 unless you pass an explicitnum_speculative_tokens(--mtp-max-kand--speculative-configare mutually exclusive).
0.10.8 — 2026-07-11 · HY3 native MTP (opt-in self-speculative decoding)
- HY3 native MTP — opt-in self-speculative decoding. Hunyuan 3's built-in 3.8B MTP head (a DeepSeek-V3-style prediction layer that the 4-bit conversion had stripped) is re-extracted into a sidecar and wired into the vendored spec-decode installer, giving HY3 factory-trained self-speculative drafting with no separate draft model (PR #1094). It is off by default — enable per serve with
--force-spec-decode --speculative-config '{"method":"mtp"}'. On M3 Ultra the head currently measures roughly break-even-to-slightly-slower on decode throughput (draft accept ~52–74% at K=3 doesn't overcome the per-step draft + verify cost), so plain autoregressive decode stays the default and recommended path.
0.10.7 — 2026-07-10 · Long-run OOM fix, Hunyuan 3 + Liquid
Community contributors: @66Ton99 — idle Responses SSE heartbeats (#1061); @MaXoS-Agent and @wparuch — Apple M2 Max community benchmarks (#1065, #1068); @ShiroKSH — TurboQuant Metal packaging (#1086).
- Long-run Metal OOM fixed. The scheduler no longer caches non-trimmable GatedDeltaNet recurrent state in the reuse cache, so long-lived / high-concurrency sessions stop leaking Metal memory and eventually hitting an OOM. The load-bearing stability win of the release (PR #1075, fixes #1025 and the #1058 leak).
- Hunyuan 3 (HY3) — new vendor family (Ultra-only preview). Day-one support for Tencent's Hunyuan 3, a 295B-total / 21B-active MoE with a 3.8B MTP head. Ultra-only: peak resident memory is ~156 GB, so it requires an M3 Ultra with 256 GB unified memory and will not fit smaller Macs. Alias
hy3-preview-4bit(4-bit MLX) with a dedicatedhy_v3tool-call + reasoning parser (PR #1070). - Tool-call parser coverage expansion. New parser families so tool-calling works across more of the catalog: LiquidAI LFM2.x (
lfm), Mistral / Devstral / Ministral (mistral), DeepSeek-Coder-V2-Lite (deepseek_v3), and NVIDIA Nemotron (fail-open). Ported from the vLLM / SGLang reference parsers rather than hand-rolled.
0.10.5 — 2026-07-08 · Fix batch
- Spec-decode unified interface migration. Legacy
_install_mtpinstaller removed (−475 LOC dead code) — every MTP path now flows through the vendored installer that was already the live runtime.mtp_optimisticis hard-rejected at all three entry points (SchedulerConfig.__post_init__,server.load_model, CLI normalizer) so misuse fails loud instead of silently ignoring the flag.spec_decode='suffix'is preserved as a canonical value;mtp_max_kconflict detection is now idempotent for the disabled default. Matched-vendored bench on Qwen3.6-27B-4bit + MTP sidecar (90 requests × 3 runs) confirmed behavior parity: tok/s median 43.83 → 46.03 (+5.0% noise), MTP rounds 17704 → 17619 (0.5% delta), finish reasons 59/31 → 60/30 (PR #1050). - README trim. Slimmed from 1142 → ~280 lines. Each trimmed section carries a → link to the specific docs page anchor instead of losing content silently (PR #1054).
- install.sh Tier-1 refresh.
RECOMMENDED_MODELnow maps to the 0.10 Tier-1 families (gpt-oss-120b · qwen3.6-35b-a3b · gpt-oss-20b · qwen3.5-4b fallback). Canonical URLs migrated torapidmlx.comacross the script (PR #1053). - HF → R2 mirror tool. New
scripts/mirror_to_r2.pyhandles streaming multipart mirrors of gated / large model weights (used to seed gpt-oss-120b MXFP4-Q8 and Qwen3.6-{27,35}B-MTP-4bit into R2). Bench / maintenance tooling only, not a runtime dependency (PR #1056). - OpenHands integration harness. Corrected the gpt-oss XFAIL reason (CodeActAgent parses text-action XML tags; gpt-oss emits harmony analysis + final channels — the format mismatch, not a stop-scoping bug). Docker digest handling in the harness now uses a two-ref pattern that avoids the OpenHands 0.9.0
base_image.split(':')bug on 3-way digest tags (PR #1055, tracked upstream at All-Hands-AI/OpenHands#15167).
0.10.3 — 2026-07-07 · Integrations + parser fixes
- Harmony parser final-channel stop scoping. User-supplied
stop=sequences now only match against the harmony final channel, so gpt-oss on CodeActAgent no longer prematurely truncates when a stop token appears inside the analysis channel. Unblocks gpt-oss under OpenHands / Aider (PR #1051, closes #1049). - Real Aider bash-CLI integration harness. Four Tier-1 family × Aider cells were un-xfailed after wiring a live
aider --model openai/<alias>harness that actually issues a code-edit round-trip (PR #1047). - Real OpenHands Docker E2E harness. CodeActAgent runs against
ghcr.io/all-hands-ai/openhands:0.9.0with a digest-pinned runtime image; three of four Tier-1 cells un-xfailed (gpt-oss cell still XFAIL — see the 0.10.5 note above) (PR #1048). - R2 mirror bypass fixes.
jlensandbench/bench --submitnow honor the R2 mirror before falling back to Hugging Face — matching whatserveandpullalready did (PRs #1045, #1046). - Spec-decode unified interface (staging). MTP moved onto the vLLM-style
speculative_configplumbing that#1050then completes in 0.10.5 (PR #1044, operator lane).
0.10.2 — 2026-07-07 · Interpretability
rapid-mlx jlens— read a model's internal draft. A read-only Jacobian-lens command that decodes what a model is disposed to say at every layer, reports where the answer crystallizes (early-exit headroom) and how far ahead of the logit lens it reads. Runs locally on quantized MLX models via a Jacobian–vector product — the fulld×dJacobian is never materialized and it differentiates through 4/8-bit weights. Follows Anthropic's Global Workspace work. Supports dense decoders (Qwen3, Llama, Phi); linear-attention / hybrid and VLM vision towers are flagged as unsupported rather than crashing.--verboseadds per-layer ranked readouts + the answer's rank trajectory;--jsonfor machine output.
0.10.1 — 2026-07-05
- Version hygiene. Reverted an accidental early 0.10.1 bump from #1010 and re-tagged cleanly (PRs #1020, #1021). No functional change.
0.10.0 — 2026-07-04
Community contributor: @wuwangzhang1216 — Codex namespace tool groups on the Responses API (#993).
- Codex tool groups on
/v1/responses. The Responses endpoint now accepts Codex namespace tool groups, fixing tool-call handling for Codex-style clients (PR #993).
0.9.14 — 2026-07-04
- OpenAI
reasoning_efforttranslation. Thereasoning_effortrequest field is now translated at the route layer, so OpenAI-shaped clients that set it map cleanly onto the engine's reasoning controls (PR #1009, closes #448).
0.9.13 — 2026-07-04
- Gemma-4 MTP sidecar CLI. New
--mtp-sidecarflag plus an eligibility gate for Gemma-4 speculative decoding — the CLI-facing half of the 0.9.13 MTP work (PR #1000).
0.9.12 — 2026-07-02
- Gemma-4 MTP drafter, stabilized. The 0.9.11 assistant-drafter inject was reverted (PR #997) and re-landed in a polished form (PR #998), so Gemma-4 speculative decoding runs without the first-cut rough edges.
0.9.11 — 2026-07-02
- Gemma-4 speculative decoding (MTP). Google's official assistant-drafter inject plus a Gemma-4 entry in the MTP allowlist (PRs #990 + #988) — faster Gemma-4 generation with byte-identical output.
- Scheduler prefix-cache fix. On an exact prompt-cache hit the cache is trimmed by one token to prevent a duplicate KV entry (PR #994).
- CLI audio aliases.
rapid-mlx pull/rmnow resolve audio aliases (PR #992).
0.9.10 — 2026-07-01
- Whisper silence-hallucination guard. A Silero VAD pre-trim strips the silent lead-in so STT stops inventing text over quiet audio (PR #980).
- Tool-call argument normalization. Assistant
tool_call.argumentsis coerced to a dict at the chat-template boundary, fixing downstream shape mismatches (PR #981). - Qwen3-Coder XML streaming. The streaming parser now anchors on
<function=…>instead of the<tool_call>wrapper (PR #979). - Anthropic adapter. Non-leading
systemmessages are merged before the chat template is applied (PR #976). - Server flags.
python -m vllm_mlx.servernow threads TurboQuant flags intoSchedulerConfig(PR #983). - README audited and rewritten for clarity and value-prop (PR #601).
0.9.9 — 2026-06-30 · Multi-user & long-context release
- Compressed KV cache by default on 9 hero MoE aliases — Qwen3.5-9B / 27B and the Qwen3.6-35B-A3B family (4 / 6 / 8-bit + DWQ). Roughly 2× more concurrent users at the same RAM budget, with byte-identical output vs the uncompressed baseline.
- Radix-tree prefix cache is default-on. Concurrent requests that share a system prompt (IDE assistants, Cursor / Claude Code / Aider fleets, multi-user chat) reuse the KV automatically. Verified 13× aggregate throughput at 10 concurrent clients.
- 128k context on Qwen3.5-9B out of the box. No flag flips — the alias resolves to the model's native 256k position-embedding cap.
- Speculative decoding (MTP + DFlash) is lossless when enabled. Byte-identical output verified vs the non-spec baseline. MTP works with the
qwen3.5-9b-mtp-4bitsidecar; DFlash works withqwen3.5-27b-8bit+ the z-lab drafter. - DFlash is now flag-gated as experimental. Install with
pip install 'rapid-mlx[dflash]', enable with--enable-dflash. Measured 1.4× pooled speedup on Qwen3.5-27B-8bit — the previously advertised 3.5× number was from a broken bench and has been corrected here. - Long-running session stability verified. 8+ hour agent loops show flat memory (±1% RSS drift over 8 h, 2,878 turns).
- GLM-5.2 removed from the roadmap. Both offline GGUF → MLX conversion attempts (Q2_K and Q3_K_M sources) landed a working load path but the model produced incoherent output on this REAP-50-pruned checkpoint.
- Config. New CLI flag
--kv-cache-turboquant noneopts out of the compressed KV cache on a per-server basis. Default is per-alias auto.
0.8.19 — 2026-06-25
- Tmax-27B hybrid-cache fix.
is_hybrid: trueon all Tmax-27B variants — was incorrectlyfalsein 0.8.18, which corrupted the RNN state on prefix-cache hits (PR #902). - Responses item shape defaults to
type: "message"when role+content are present (PR #904). - thinking accounting.
enable_thinkingnow threads into prompt accounting on the chat + responses paths so reasoning models report token usage correctly (PR #906, follow-up to #891). - Auto-config routing. V4 / V5 model-type detection aligned with the alias classifier (PR #903).
- CLI:
rapid-mlx servenow exits non-zero on a port-already-in-use error (PR #905). - CI: pr_validate full_unit flakes on macOS uv-managed Python fixed (PR #907).
0.8.18 — 2026-06-24
- Tmax-9B + Tmax-27B aliases added. First-mover MLX support — see the Tmax docs page (PR #899).
- tool_call promotion in reasoning ported from upstream into the refactored think_parser (closes #344, PRs #896 + #898).
- Auto-disable thinking on casual chat completions (matches the M-2 / #891 pattern, PR #895).
- Auto-disable thinking when tools are provided — strict-JSON pattern carried over to reasoning models (PR #891).
- Audio:
_serve_audio_modenow honours--embedding-modeland--served-model-name(closes R11-K #258, PR #894). - Tool parser: prevent
deepseek_v3parser from binding to non-V3 Qwen2 distills + warn on misbind (PR #887). - Tempfile leak fix:
managed_tempfilehelper applied across leak sites (closes #719, PR #888). - Strict JSON schema retry on context-length re-check before the repair retry (closes #267b, PR #886).
0.8.16 — 2026-06-22
/v1/modelssurfaces effective parsers. Thetool_call_parserandreasoning_parserthe alias resolved to are now visible per model (V-1, S-2, PR #878).- reasoning_content sanitization on the
tool_choice=requiredpath so the reasoning trace is not leaked into the tool-call output (V-2, PR #881).
0.8.15 — 2026-06-21
- Strict JSON schema mode enforced via post-generate validation + repair retry (closes #423, PR #873).
- Reasoning tail rescue: tail re-routed to
contentwhenfinish_reason=lengthhits mid-reasoning (closes 8-round D-carry #259, PR #875). - CLI:
rapid-mlx launch <client>— one-shot IDE bootstrap for Cursor, Claude Code, Aider, etc. (closes #566, PR #870). - deepseek_v3 variant added for the R1-0528 family (PR #874).
- Chat output always emits the
contentkey on the assistant message — fixes clients that throw on its absence (PR #872).
0.8.14 — 2026-06-20
- Responses streaming fix: exclude the reasoning-cutoff sentinel from
downstream_output_seen— regression from PR #860 (PR #869). - README documents the 0.8.13 audio support (TTS + STT, 26 aliases) (PR #868).
- Length-stop rescue stub for reasoning models restored after refactor (closes #858, PR #860).
0.8.13 — 2026-06-19 · Audio launch
- Audio support shipped. 26 aliases — 13 TTS (Kokoro, Chatterbox, VibeVoice, VoxCPM, Dia) and 13 STT (Whisper, Parakeet) — via
/v1/audio/speechand/v1/audio/transcriptions. Install withpip install 'rapid-mlx[audio]'. - Audio CLI:
--audio-mode+ comprehensive alias registry (R10-A, PR #854). - Responses: emit reasoning item on
max_output_tokenscutoff (R11-M-F1, PR #866). - Embeddings:
[embeddings]extra + 503 on missing model +/v1/modelsvisibility (H-08 + H-09 + H-13, PR #861). - VLM penalties: frequency / presence / repetition penalties now pass through to the VLM sampler (closes #512, PR #864).
- Tool-choice=required: finalize invariant pinned in streaming (R11-V1 + R11-V2, PR #859).
- UI-TARS: accumulator-anchor + honour
enable_thinking=false(R10-F, PR #850).
Older
Full pre-0.8.13 release notes are in the GitHub releases history. Notable highlights from the 0.7 → 0.8 era:
- 0.8.x: Reasoning-parser refactor, multi-turn prompt cache with KV trimming for sub-100 ms TTFT, RNN state snapshots for hybrid architectures.
- 0.7.41: Cross-thread
Stream(gpu, 4)crash fix in logprobs path (hotfix). - 0.7.40: 5 systematic integration fixes — tool_choice / stop / template-escape / logprobs / think-detection (PR #716).
- 0.7.x: VibeThinker integration, alias bundle for ≤5B models, diffusion-lane fix for mlx-vlm eos_token_ids alias.
Free, open source, runs entirely on your Mac.
curl -fsSL https://rapidmlx.com/install.sh | bash