Model families on Rapid-MLX
Every model rapid-mlx serves, by family. If you just want to know what to run on your Mac, start with the first section.
Pick a model for your Mac
Ask rapid-mlx — it reads your Mac's memory and prints two picks with the exact command:
$ rapid-mlx recipe
Recommended for this 18.0 GB Mac (18 GB tier)
1. Smart — qwen3.5-9b-4bit
rapid-mlx serve qwen3.5-9b-4bit
2. Fast — qwen3.5-4b-4bit
rapid-mlx serve qwen3.5-4b-4bit
Then run the line it prints — rapid-mlx serve <alias>
downloads the weights on first use and starts an OpenAI-compatible
server at http://localhost:8000/v1. The alias selects the
right tool-call parser, reasoning parser and cache settings for you.
The catalog holds 197 text, vision and reasoning aliases, 44 audio
aliases and 10 video-generation aliases, plus image-generation and
embedding models. rapid-mlx models lists them all in
your terminal.
Hero deep dives
Six models get their own page because serving them cleanly on MLX takes model-specific work — a custom parser, a separate engine, or special cache handling. Each page starts with which variant fits your Mac and the command to run it.
qwen3_xml tool calling, hybrid cache path.ui_tars tool parser.holo3.1-35b-a3b (4-bit / 8-bit) — reported running well in production by a peer Apple-Silicon agent.All families
Each family page groups its version lines under one heading each,
with every alias, its parsers and capability notes. Pass any alias
to rapid-mlx serve <alias>; the parsers and cache
settings are wired up for you. The registry itself is
rapid_mlx/aliases.json (audio:
rapid_mlx/audio/aliases.json).
llama tool envelope.muse ATEM tool envelope · text serving.hy_v3 tool + reasoning parser./v1/audio/*./v1/videos job API · needs Python 3.11+ and ffmpeg./v1/images · needs the [image] extra.Notable small models
Headline-family pages are organized by the big lines (Qwen, Gemma, Llama). But the most interesting work in 2026 is often happening at the small-and-weird end — focused-purpose models, single-shop research checkpoints, frontier-style tricks compressed into 4 GB. The list below is editorial: small models we find worth trying, or that another team has reported working well in production.
holo3.1-35b-a3b (4-bit / 8-bit) — reported running well in production by a peer Apple-Silicon agent.qwen3_xml tool calling and the hybrid cache path./v1/embeddings for RAG and code search without an API call — 1.13 GB at 4-bit.