Changelog / release

0.15.4 — Experimental acceleration profiles and Share Compute in Desktop

Released 2026-10-02 · full changelog · GitHub releases

Upgrade: pip install -U rapid-mlx  ·  brew upgrade rapid-mlx  ·  or grab the desktop app.

0.15.4 adds two narrowly qualified experimental acceleration paths for Qwen3.8 27B and GLM-5.3 Flash, plus a complete Share Compute experience in Desktop. The accelerated profiles pin their model and runtime inputs and fail closed when a request or machine is outside the measured contract.

  • Experimental Qwen3.8 27B acceleration. The opt-in qwen3.8-27b-tensorfold profile pairs one qualified 4-bit checkpoint with its DFlash2 drafter and exposes the active backend and limitations through the server status APIs and Desktop settings. It admits one text request at a time and rejects tools, media, grammar constraints, and unqualified model or runtime revisions. The ordinary qwen3.8-27b-4bit path remains available as the feature-complete fallback. On the tracked qualification fixture, a 460-token coding prompt with a 256-token thinking budget and 768-token reply cap measured 57.5 tok/s with drafting against 47.5 tok/s without, a 1.21× ratio; these bounded samples do not establish a fixed multiplier or quality gain.
  • Experimental GLM-5.3 Flash acceleration on 256 GB Macs. The dedicated glm5.3-flash-tensorfold alias uses the checkpoint's embedded MTP head and selects its qualified accelerated backend by default on eligible systems. Startup validates the exact runtime and checkpoint revisions; unsupported tools, media, grammar, and general batching fail explicitly. Desktop labels it Experimental with an opt-out that restarts on the ordinary GLM serving path, and Desktop chat works with the profile out of the box.
  • Share Compute arrives in Desktop. The new Experimental workspace guides users through choosing an eligible local model, reviewing a connection, contributing to the live pool, stopping a session, and restoring their previous local model. A separate read-only key stored in Keychain can show settled API credit across nodes on the same account, while local activity remains scoped to this Mac. Provider credentials are sent only to the local provider process and are not saved in app settings.
  • Faster high-concurrency dense inference. An opt-in row-invariant lane matrix-multiplication path accelerates qualified dense 4-bit workloads at eight or more concurrent requests. Measured aggregate decode gains ranged from 50–81% on an M5 Max and 52–59% on an M3 Pro at 8–16 requests. It provides no measured gain at four or fewer rows and skips mixture-of-experts models.
  • Serving is more predictable and better verified. Qwen acceleration caps admission at one active request, and both accelerated profiles verify installed runtime provenance before loading. GLM streaming preserves code whitespace and reasoning boundaries across streaming and non-streaming responses, Stable Audio 3 weight downloads are pinned to a reviewed immutable revision, and rapid server restart loops suppress duplicate anonymous startup telemetry within a bounded window.