Docs
Rapid-MLX runs open models on your Mac's GPU and serves them through an OpenAI- and Anthropic-compatible API. Three commands get you from nothing to a running model:
$ curl -fsSL https://rapidmlx.com/install.sh | bash # or: brew install rapid-mlx $ rapid-mlx serve qwen3.5-4b-4bit # any alias, or an org/repo from Hugging Face $ curl http://localhost:8000/v1/chat/completions -H "Content-Type: application/json" \ -d '{"model":"qwen3.5-4b-4bit","messages":[{"role":"user","content":"hi"}]}'
Or chat in the terminal with rapid-mlx chat — it starts and stops the server for you.
Run rapid-mlx on its own for a one-screen start menu: chat, best model for this Mac, connect a coding agent, start a server, or choose a model.
Full walkthrough: Install & first request.
Next, pick what you came for
rapid-mlx recipe prints the smart and fast pick for your RAM; the tiers page explains each pick and links measured speeds on Macs like yours.http://localhost:8000/v1, port 8000, the endpoints, and curl, Python and Node examples.rapid-mlx service install starts the server at boot on a headless Mac, with health-gated upgrades.rapid-mlx doctor, then the common symptoms and fixes.Agents & integrations
Rapid-MLX serves an OpenAI-compatible endpoint at
localhost:8000/v1, so most tools work by just pointing them
there. These guides give the exact per-tool config, tested against the
engine.
launch claude-code setup and endpoint config.apply_patch tool path.Hero models
These are the models where Rapid-MLX does engineering work beyond the upstream MLX-LM stack to serve them cleanly: our own MLX quants, alias configuration, and dedicated tool-call and reasoning parsers.
ui_tars tool parser.reasoning_content./v1/embeddings endpoint, next to your chat model — for RAG and semantic or code search. 1.13 GB at 4-bit.Looking for everything else?
The full set of supported models — 197 text/vision/reasoning aliases (250 total incl. 44 audio and 10 video generation) across 15 vendor families including Qwen (3.8 / 3.6 / 3.5 / Coder / VL / legacy / 2.5 / Qwopus), Gemma (4 / 4-mobile / 3 / EmbeddingGemma), Llama, DeepSeek, GLM, Mistral, Phi, GPT-OSS, MiniMax, Hermes, Hunyuan, Liquid, Granite and a curated small-models bucket — is browseable as a family directory with one page per vendor, or as a flat alias catalog for ctrl-F. For "what fits my Mac" use the live RAM picker at models.rapidmlx.com.
Notable small models
Editorial short-list of small-but-interesting models worth watching outside the headline families — including Holo3.1-35B-A3B, VibeThinker, Tmax-9B, Granite 4 H-Micro, and EmbeddingGemma 2. The full callout lives on the model families page.
What is on by default
rapid-mlx serve <alias> applies the alias's tuned
profile with no extra flags: the right tool-call and reasoning parsers,
the prefix cache, compressed KV cache on the verified Qwen3.5 / 3.6
aliases, and MTP speculative decoding on the aliases qualified for it.
PFlash long-prompt compression is opt-in (--pflash auto). The
performance flags reference lists each
technique, which aliases get it, and how to turn it off. Release
history lives in the changelog.