Rapid-MLX docs · v0.15.7

Docs

Rapid-MLX runs open models on your Mac's GPU and serves them through an OpenAI- and Anthropic-compatible API. Three commands get you from nothing to a running model:

$ curl -fsSL https://rapidmlx.com/install.sh | bash     # or: brew install rapid-mlx
$ rapid-mlx serve qwen3.5-4b-4bit                       # any alias, or an org/repo from Hugging Face
$ curl http://localhost:8000/v1/chat/completions -H "Content-Type: application/json" \
    -d '{"model":"qwen3.5-4b-4bit","messages":[{"role":"user","content":"hi"}]}'

Or chat in the terminal with rapid-mlx chat — it starts and stops the server for you. Run rapid-mlx on its own for a one-screen start menu: chat, best model for this Mac, connect a coding agent, start a server, or choose a model. Full walkthrough: Install & first request.

Next, pick what you came for

Agents & integrations

Rapid-MLX serves an OpenAI-compatible endpoint at localhost:8000/v1, so most tools work by just pointing them there. These guides give the exact per-tool config, tested against the engine.

Hero models

These are the models where Rapid-MLX does engineering work beyond the upstream MLX-LM stack to serve them cleanly: our own MLX quants, alias configuration, and dedicated tool-call and reasoning parsers.

Looking for everything else?

The full set of supported models — 197 text/vision/reasoning aliases (250 total incl. 44 audio and 10 video generation) across 15 vendor families including Qwen (3.8 / 3.6 / 3.5 / Coder / VL / legacy / 2.5 / Qwopus), Gemma (4 / 4-mobile / 3 / EmbeddingGemma), Llama, DeepSeek, GLM, Mistral, Phi, GPT-OSS, MiniMax, Hermes, Hunyuan, Liquid, Granite and a curated small-models bucket — is browseable as a family directory with one page per vendor, or as a flat alias catalog for ctrl-F. For "what fits my Mac" use the live RAM picker at models.rapidmlx.com.

Notable small models

Editorial short-list of small-but-interesting models worth watching outside the headline families — including Holo3.1-35B-A3B, VibeThinker, Tmax-9B, Granite 4 H-Micro, and EmbeddingGemma 2. The full callout lives on the model families page.

What is on by default

rapid-mlx serve <alias> applies the alias's tuned profile with no extra flags: the right tool-call and reasoning parsers, the prefix cache, compressed KV cache on the verified Qwen3.5 / 3.6 aliases, and MTP speculative decoding on the aliases qualified for it. PFlash long-prompt compression is opt-in (--pflash auto). The performance flags reference lists each technique, which aliases get it, and how to turn it off. Release history lives in the changelog.