Changelog / release

0.15.7 — EmbeddingGemma 2, two more accelerated profiles, steadier agent setup

Released 2026-10-07 · full changelog · GitHub releases

Upgrade: pip install -U rapid-mlx  ·  brew upgrade rapid-mlx  ·  or grab the desktop app.

0.15.7 adds native text and code embeddings with EmbeddingGemma 2, two experimental TensorFold profiles, and agent setup that keeps more of your existing client configuration. Servers gain switches for disk writes and log output, and responses can report per-request timing.

  • EmbeddingGemma 2 text and code embeddings. embeddinggemma-2-4bit and embeddinggemma-2-bf16 serve normalized 768-dimensional vectors through /v1/embeddings, with optional smaller dimensions and an explicit 8192-token limit. Text and code inputs only; images, audio and video are refused.
  • Two more experimental accelerated profiles. nemotron-3.5-lightning-tensorfold (48 GB minimum) and qwen3.8-flash-next-tensorfold (192 GB minimum) use the MTP head published with each checkpoint, serve text chat without tools, images or grammar, and leave the ordinary aliases unchanged. All four TensorFold profiles now share one pinned TensorFold runtime revision, so an existing TensorFold install needs reinstalling at that revision. No fixed speedup is claimed.
  • Agent setup that keeps what you have. Continue and Cline get their configuration where their clients read it, and a keyed server's credential reaches the Continue entry. Qwen Code setup keeps your other OpenAI-compatible providers, previews the exact change with their keys redacted, and backs up the file. Codex setup ends with a report of the saved defaults read back from disk. See agent setup.
  • Disk writes and logs under your control. --disable-disk-caches stops optional cache writes (prefix-cache snapshot, KV checkpoints, vision cache disk tier) while keeping in-memory reuse, --log-file sends server output to a file or /dev/null, and the global --disable-version-check skips release checks. Prefix-cache persistence leaves a free-disk reserve on the volume. See Disk writes.
  • Per-request timing. Chat Completions, Completions and Responses can report experimental metrics.time_to_first_token_ms and metrics.mean_itl_ms for the request (see API).
  • Warmer coding-agent turns. A warm turn reuses the tokenized prompt instead of re-tokenizing it whole, and greedy MTP output is reproducible from run to run. Gains depend on the model, hardware and workload.
  • LTX-2.5 video extension. POST /v1/videos/extend appends frames to an existing MP4 through the usual video job API.
  • Model catalog. qwopus-27b-8bit is retired because its upstream repository is gone; use qwopus-27b-4bit. qwen3.8-27b-abliterated-4bit follows the publisher's relocated oQ4e build.
  • Desktop. Broader Simplified Chinese coverage, clearer timeout handling and a verified saved Codex model.
  • Also fixed. An unreadable model cache on an external volume now says how to grant access instead of looking like a missing model (see Troubleshooting), and pinned-revision image, video and speculative checkpoints download through the mirror. Computer Use remains experimental.