OpenAI-compatible local LLM server on Mac (MLX)
rapid-mlx serve <model> runs a model on your Mac's GPU
with MLX and serves it at http://localhost:8000/v1: port
8000, the OpenAI paths under /v1. Point any OpenAI
client at that base URL and use the model alias as model.
$ brew install rapid-mlx # or: curl -fsSL https://rapidmlx.com/install.sh | bash $ rapid-mlx serve qwen3.5-4b-4bit $ curl http://localhost:8000/v1/chat/completions -H "Content-Type: application/json" \ -d '{"model": "qwen3.5-4b-4bit", "messages": [{"role": "user", "content": "Hello"}]}'
| Setting | Value |
|---|---|
| OpenAI base URL | http://localhost:8000/v1 |
| Anthropic base URL | http://localhost:8000 (no /v1; the SDK adds /v1/messages) |
| Port | 8000. If it's taken, the next free port up to 8009, and the startup banner says which. --port N pins it. |
| Host | 127.0.0.1, this Mac only. --host 0.0.0.0 opens it to your network. |
| API key | Not checked by default; SDKs need any non-empty string. Set --api-key or RAPID_MLX_API_KEY to require one. |
| Model name | The alias you served (qwen3.5-4b-4bit) or its Hugging Face repo id. GET /v1/models lists both. |
Start the server
Install with Homebrew (brew install rapid-mlx) or the
one-line installer, then serve a model by alias or by Hugging Face
repo id. The first run downloads the weights. When the server is up it
prints the URLs to use:
$ rapid-mlx serve qwen3.5-4b-4bit … Ready: http://127.0.0.1:8000 OpenAI: http://127.0.0.1:8000/v1 Anthropic: http://127.0.0.1:8000 Model: qwen3.5-4b-4bit
localhost and 127.0.0.1 are the same address
here. Not sure which model fits your Mac? rapid-mlx recipe
prints the pick for your memory, and
Pick a model for my Mac explains each
tier. In another terminal, rapid-mlx connect prints the same
URLs for a server that is already running.
Port and base URL
- Default port 8000. Without
--port,servetakes the first free port from 8000 to 8009 and says so:Port 8000 is in use; using 8001 instead (pass --port to choose).The base URL is thenhttp://localhost:8001/v1. - A fixed port.
rapid-mlx serve qwen3.5-4b-4bit --port 8080. An explicit port never falls back: if it's busy,serveexits with an error. - Include
/v1in the base URL for OpenAI clients (they append/chat/completions,/responsesand so on). Leave it off for Anthropic clients such as Claude Code. - Other machines on your network.
--host 0.0.0.0listens on every interface. Start it with an API key first (below).
Endpoints
Paths are relative to http://localhost:8000.
| Endpoint | API | Notes |
|---|---|---|
| POST /v1/chat/completions | OpenAI Chat Completions | Streaming, tool calling, vision inputs, reasoning_content. |
| POST /v1/responses | OpenAI Responses | What Codex CLI speaks; streaming events, reasoning and tool-call items. |
| POST /v1/messages | Anthropic Messages | What Claude Code speaks; tools, streaming, thinking blocks. Plus /v1/messages/count_tokens. |
| POST /v1/completions | OpenAI Completions | Raw text completion, no chat template. |
| GET /v1/models | OpenAI Models | The served model, with its context window and capabilities. |
| POST /v1/embeddings | OpenAI Embeddings | Attach an embedding model with --embedding-model (EmbeddingGemma 2 needs the [vision] extra); without one it returns an error that says so. See EmbeddingGemma 2. |
| POST /v1/audio/speech, /v1/audio/transcriptions | OpenAI Audio | Text-to-speech and speech-to-text with an audio model; needs the [audio] extra. |
| POST /v1/images/generations | OpenAI Images | Text-to-image with an image model; needs the [image] extra. |
| GET /health | Probe | Health check; also /healthz, /readyz, /livez. |
| GET /docs, /openapi.json | Schema | Swagger UI and the OpenAPI schema of the running server. |
Video, audio translation and music, image edits, metrics and model management are in the full API reference, with every place the server differs from OpenAI.
Call it from curl, Python and Node
curl
$ curl http://localhost:8000/v1/models $ curl http://localhost:8000/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{"model": "qwen3.5-4b-4bit", "messages": [{"role": "user", "content": "Hello"}]}'
Python (openai)
from openai import OpenAI client = OpenAI(base_url="http://localhost:8000/v1", api_key="not-needed") r = client.chat.completions.create( model="qwen3.5-4b-4bit", messages=[{"role": "user", "content": "Hello"}], ) print(r.choices[0].message.content) # the Responses API works on the same client print(client.responses.create(model="qwen3.5-4b-4bit", input="Hello").output_text)
Node / TypeScript (openai)
import OpenAI from "openai"; const client = new OpenAI({ baseURL: "http://localhost:8000/v1", apiKey: "not-needed" }); const r = await client.chat.completions.create({ model: "qwen3.5-4b-4bit", messages: [{ role: "user", content: "Hello" }], }); console.log(r.choices[0].message.content);
Without changing code
The OpenAI SDKs read the base URL and key from the environment:
$ export OPENAI_BASE_URL=http://localhost:8000/v1 $ export OPENAI_API_KEY=not-needed $ python your_script.py
qwen3.5-4b-4bit on a Mac mini (M2 Pro, 32 GB): the
startup banner and port fallback above, the curl calls, the Python
snippet (openai 3.26.0) and the Node snippet (openai 7.30.0) all
returned replies.
API key and network access
A plain rapid-mlx serve doesn't check keys, and it only
listens on 127.0.0.1. To require a key, set it when you
start the server:
$ export RAPID_MLX_API_KEY="$(openssl rand -hex 24)" $ rapid-mlx serve qwen3.5-4b-4bit # or: --api-key <key>
Clients then send Authorization: Bearer <key>
(what the OpenAI SDKs do with api_key); on
/v1/messages, Anthropic's x-api-key header
works too. A missing or wrong key gets 401; the health
probes stay open. Set a key before you use --host 0.0.0.0
or put the server behind a tunnel.
How it differs from mlx_lm.server
mlx_lm.server, the server that ships with Apple's mlx-lm,
is also OpenAI-compatible. Rapid-MLX runs the same MLX weights; the
differences are in what the server exposes (mlx-lm 0.32.0 checked):
| Rapid-MLX | mlx_lm.server | |
|---|---|---|
| Start | rapid-mlx serve qwen3.5-4b-4bit | mlx_lm.server --model <repo> |
| Default port | 8000 (next free up to 8009) | 8080 |
| OpenAI routes | Chat Completions, Completions, Models, Responses, Embeddings, Audio, Images | Chat Completions, Completions, Models |
| Anthropic Messages | /v1/messages | No |
| API key | --api-key / RAPID_MLX_API_KEY | No option; its own help says it is not recommended for production |
| Coding agents | One setup command per agent (guides) | Configure by hand |
Moving a client over means changing the base URL from
http://localhost:8080/v1 to
http://localhost:8000/v1 and the model name to an alias
(or keep the repo id). On speed, Rapid-MLX 0.15.6 measured up to 4×
faster than mlx-lm 0.32.0 and 1.5× on a typical task, with the same
weights: Qwen3.5-9B 4-bit on a Mac mini M4 Pro (48 GB), best case
whole-file code edits, long-prompt agent turns close to a tie. Every
task and the raw data are on Rapid-MLX vs
mlx-lm.
Frequently asked questions
What port does the MLX OpenAI-compatible server use?
Rapid-MLX listens on port 8000 by default, so the OpenAI base URL is http://localhost:8000/v1. If 8000 is taken it uses the next free port up to 8009 and prints it; --port sets it explicitly. mlx-lm's own mlx_lm.server defaults to port 8080.
What is the base URL for a local OpenAI-compatible server on Mac?
http://localhost:8000/v1 for OpenAI clients (Python, Node, LangChain, Codex CLI) and http://localhost:8000 for Anthropic clients such as Claude Code. Use the alias you served as the model name.
Does it need an OpenAI API key?
No. The server doesn't check keys unless you start it with --api-key or RAPID_MLX_API_KEY. The OpenAI SDKs refuse an empty key, so pass any placeholder string such as not-needed.
Which OpenAI endpoints are supported?
Chat Completions, Completions, Models, Responses, Embeddings (with an embedding model attached), Audio speech and transcription, and Images, plus Anthropic Messages. The API reference lists every route and where it differs from OpenAI.
Can I use it from another computer on my network?
Yes. Start it with --host 0.0.0.0 and an API key, then use http://<your-mac's-ip>:8000/v1 from the other machine. By default it only accepts connections from the Mac itself.
Is Rapid-MLX a drop-in replacement for mlx_lm.server?
For OpenAI clients, yes: change the base URL from port 8080 to 8000 and set the model name. It runs the same MLX weights and adds the Responses and Anthropic Messages APIs, API keys, and embeddings, audio and image routes.
Next
- Connect a coding agent: Claude Code, Codex CLI, OpenCode and more, one setup command each.
- OpenAI SDK guide: the base URL switch, streaming and tool calls.
- API reference: every route and each place it differs from OpenAI.
- Run it as a service: start the server at boot.