Roadmap · rapid-mlx 0.15.7

Roadmap

Rapid-MLX ships small releases often rather than working to a fixed long-range plan. The best view of what is coming is the open work on GitHub; the best view of what is settled is everything outside the experimental list below.

Where to follow what's next

Experimental in 0.15.7

These work today but are marked experimental: their behaviour, flags or hardware requirements may still change.

AreaWhat it is
Large-Mac model lanesglm5.3-flash-4bit, mimo-v2.6-flash-4bit, qwen3.8-flash-next-4bit and deepseek-v41-flash-reap-2bit — frontier MoE models for 192–256 GB Mac Studios. rapid-mlx info <alias> shows each one's limits.
Accelerated profilesqwen3.8-27b-tensorfold, glm5.3-flash-tensorfold, nemotron-3.5-lightning-tensorfold (48 GB) and qwen3.8-flash-next-tensorfold (192 GB) — opt-in speculative profiles that need the separately installed TensorFold runtime. They serve text chat without tools, media or grammar.
Other experimental aliasesk2-horizon-7b-4bit, neohorse-9b-4bit, qwen3.8-27b-abliterated-4bit.
Decision modelsrapid-mlx system-one — a separate typed-decision server (Laya, CLM, and Clef / Clef-Flash through the [clef] extra) for /v1/systemone and /v1/rank.
Computer userapid-mlx cua — a native macOS accessibility agent with a configurable planner.
Speculative decodingDDTree, DFlash pairs below 8-bit, and SuffixDecoding on aliases not yet benchmarked for it (see performance flags).
Memory techniques--use-paged-cache, --kv-cache-turboquant v4, --disk-stream for MoE experts, and --kv-disk-checkpoint-interval (snapshots are written but not yet reloaded).

Contribute

Issues labelled for contributors and the contribution guide are on GitHub. Benchmarks from your own Mac help too: rapid-mlx benchmark run <alias> (any alias from rapid-mlx benchmark catalog), then rapid-mlx benchmark share <run-id> adds them to the leaderboard.