Changelog / release

0.15.3 — Supervised Computer Use and more resilient local serving

Released 2026-09-30 · full changelog · GitHub releases

Upgrade: pip install -U rapid-mlx  ·  brew upgrade rapid-mlx  ·  or grab the desktop app.

0.15.3 adds opt-in, supervised Computer Use for bounded local Mac tasks, along with a documented API for clients that provide their own interface. It also improves lower-memory image generation, first-run setup, model serving, and recovery from common startup errors.

  • Computer Use stays bounded and approval-first. Desktop can resolve a task to a selected macOS app or browser window, keep that scope visible, and pause for approval before consequential actions. Credentials, payment details, and commerce actions remain blocked, and runs stop when their target or approval context changes.
  • A documented API supports custom Computer Use clients. The authenticated /v1/cua API covers capabilities, permissions, app and window discovery, bounded observations, run creation, event polling, approval decisions, cancellation, and a model-free --cua-only server mode. Screenshots require opt-in from both server and request, and raw click and typing operations are not exposed.
  • Qwen Image 2.1 can fit on lower-memory Macs. The image path now supports a quantized mflux configuration with an 8 GB minimum-memory gate, while the original BF16 checkpoint remains available explicitly as qwen-image-2.1-bf16. Actual fit still depends on the model and memory available to the process.
  • First runs and startup failures are easier to recover from. Running bare rapid-mlx shows first-run guidance. Missing optional runtimes provide commands for the active install method, while missing local paths and curated model IDs receive clearer errors.
  • Desktop counts its first-run funnel without identifying an install. The anonymous milestones run once per new install, contain only the app version and a closed milestone name, and honour both telemetry and update-check settings. The privacy disclosure documents the complete boundary.
  • Agent requests behave more predictably. Strict structured output works with streaming and tool calls, completion budgets stay within the model context, and QuickSilver prompt admission and hybrid checkpoint caching gain additional bounds for long sessions. Structured-output and context-length rejections can include a closed reject_reason without exposing request text.
  • Agent sessions keep their cached prefix on 16–32 GB Macs. A minimum cache budget keeps the agent-session prefix resident, and shutdown saves it for the next start, so repeat turns reuse the cached prefix instead of prefilling it again. See the measured comparisons.
  • Blank PDF pages no longer stall OCR. Blank pages bypass text recognition, allowing scanned imports to finish or report that no readable content was found. Inference failures use closed telemetry categories, and private error text and Computer Use observations are not sent as telemetry.