Roadmap · rapid-mlx 0.15.7
Roadmap
Rapid-MLX ships small releases often rather than working to a fixed long-range plan. The best view of what is coming is the open work on GitHub; the best view of what is settled is everything outside the experimental list below.
Where to follow what's next
- GitHub issues — planned work, model requests and bugs. When
pullorserverefuses a Hugging Face model, adding--requestfiles a support request for it. - Pull requests — what is being built right now.
- Changelog — what each release changed.
- Discord — tell us what you want (
rapid-mlx feedbackopens it).
Experimental in 0.15.7
These work today but are marked experimental: their behaviour, flags or hardware requirements may still change.
| Area | What it is |
|---|---|
| Large-Mac model lanes | glm5.3-flash-4bit, mimo-v2.6-flash-4bit, qwen3.8-flash-next-4bit and deepseek-v41-flash-reap-2bit — frontier MoE models for 192–256 GB Mac Studios. rapid-mlx info <alias> shows each one's limits. |
| Accelerated profiles | qwen3.8-27b-tensorfold, glm5.3-flash-tensorfold, nemotron-3.5-lightning-tensorfold (48 GB) and qwen3.8-flash-next-tensorfold (192 GB) — opt-in speculative profiles that need the separately installed TensorFold runtime. They serve text chat without tools, media or grammar. |
| Other experimental aliases | k2-horizon-7b-4bit, neohorse-9b-4bit, qwen3.8-27b-abliterated-4bit. |
| Decision models | rapid-mlx system-one — a separate typed-decision server (Laya, CLM, and Clef / Clef-Flash through the [clef] extra) for /v1/systemone and /v1/rank. |
| Computer use | rapid-mlx cua — a native macOS accessibility agent with a configurable planner. |
| Speculative decoding | DDTree, DFlash pairs below 8-bit, and SuffixDecoding on aliases not yet benchmarked for it (see performance flags). |
| Memory techniques | --use-paged-cache, --kv-cache-turboquant v4, --disk-stream for MoE experts, and --kv-disk-checkpoint-interval (snapshots are written but not yet reloaded). |
Contribute
Issues labelled for contributors and the contribution guide are
on GitHub.
Benchmarks from your own Mac help too:
rapid-mlx benchmark run <alias> (any alias from rapid-mlx benchmark catalog), then
rapid-mlx benchmark share <run-id> adds them to the
leaderboard.