Install & first request
Install Rapid-MLX, start a server, and make your first OpenAI-compatible chat completion.
$ curl -fsSL https://rapidmlx.com/install.sh | bash $ rapid-mlx serve qwen3.5-4b-4bit $ curl http://localhost:8000/v1/chat/completions -H "Content-Type: application/json" \ -d '{"model":"qwen3.5-4b-4bit","messages":[{"role":"user","content":"ping"}]}'
That is the whole path. The sections below explain each step and the options you have.
Requirements
- Apple Silicon Mac — M1 or newer.
- macOS 14 (Sonoma) or later.
- Python 3.10 or newer (the installer auto-fetches one if you do not have it).
Intel Macs are not supported — Rapid-MLX is built on Apple's MLX framework which targets the unified-memory Apple Silicon GPU.
Install
One curl — no sudo, nothing to sign up for.
$ curl -fsSL https://rapidmlx.com/install.sh | bash
Prefer Homebrew? It's in homebrew/core — no tap, no trust:
$ brew install rapid-mlx
The installer looks for Python 3.10+ (installing Python 3.12 through
Homebrew, or a standalone build, if none is found), creates a venv at
~/.rapid-mlx, symlinks the rapid-mlx
command into ~/.local/bin, and suggests a starter model
for your Mac's RAM. You can also install from pip into a
venv you manage yourself:
$ python3 -m venv .venv && source .venv/bin/activate $ pip install rapid-mlx
Not sure where to start? Run rapid-mlx with no command. It
shows one screen with your Mac's chip and RAM and the model it will
chat with, then waits for a key: Enter chats (the quick-start
qwen3.5-4b-4bit on a cold cache, or a recommended model
you already have), b picks the best model for this Mac,
c connects a detected coding agent, s starts a server and
m chooses another model. Each key prints the exact command it
runs; the CLI reference has the details.
Start a server
Pick any alias from the catalog
and pass it to rapid-mlx serve. Aliases are short
memorable names that resolve to the canonical Hugging Face repo
and apply the right tool-call parser, reasoning parser and performance
defaults. A Hugging Face org/name works too (see
Bring your own model):
$ rapid-mlx serve qwen3.5-4b-4bit # downloads the MLX weights on first run, then: INFO: Uvicorn running on http://127.0.0.1:8000 (Press CTRL+C to quit)
The server listens on 127.0.0.1 only. Without
--port it takes the first free port from 8000 to 8009
(8000 unless something already uses it); an explicit
--port 9000 never falls back. Set the served name with
--served-model-name my-model, or pick a quant explicitly
with the full alias suffix (e.g. qwen3.5-9b-8bit).
--host 0.0.0.0 exposes the server to your network — set
--api-key first.
Make your first request
The endpoints are drop-in OpenAI-compatible — point any OpenAI
client at http://localhost:8000/v1 with any non-empty
API key.
cURL
$ curl http://localhost:8000/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{ "model": "qwen3.5-4b-4bit", "messages": [{"role": "user", "content": "ping"}] }'
Python (openai SDK)
from openai import OpenAI client = OpenAI( base_url="http://localhost:8000/v1", api_key="not-used", ) resp = client.chat.completions.create( model="qwen3.5-4b-4bit", messages=[{"role": "user", "content": "ping"}], ) print(resp.choices[0].message.content)
What is on by default
Each alias carries a tuned profile, so these apply without any extra flag on the aliases that qualify for them:
- Prefix cache — when requests share a system prompt or tool schema, the server reuses the already-computed KV instead of prefilling it again. On for every model.
-
Compressed KV cache (K8V4) — on 9 verified Qwen3.5 / 3.6
aliases (Qwen3.5-9B / 27B / 35B and Qwen3.6-35B-A3B in 4 / 6 / 8-bit
and DWQ), so more concurrent sessions fit in the same RAM. Turn off
with
--kv-cache-turboquant none. -
MTP speculative decoding — on for
qwen3.5-9b-4bit,qwen3.6-27b-4bit,qwen3.6-35b-4bit,qwen3.8-27b-4bitandglm5.3-flash-4bit. Turn off with--no-spec-decode.
PFlash long-prompt compression is off for every model; opt in with
--pflash auto or --pflash always.
rapid-mlx info <alias> shows exactly what a given
alias turns on. The performance flags
reference covers every technique and its opt-in variants.
Wire it into a tool
Anything that speaks the OpenAI API works without code changes —
point its base URL at http://localhost:8000/v1. For coding
agents there is a setup command per agent, and a one-step start:
$ rapid-mlx agents # list supported agents $ rapid-mlx agents codex --setup # point Codex CLI at the running server $ rapid-mlx start claude-code # start a server and configure Claude Code
start picks a recommended model that fits your Mac (and is
already downloaded, if one is), starts the server and configures the
agent; add --dry-run to preview it. The
agent guides list what each setup writes.
Next steps
rapid-mlx recipe prints the smart and fast pick for your RAM; the tiers page explains them.org/name, or import and quantize one.