Reference · Local server

OpenAI-compatible local LLM server on Mac (MLX)

rapid-mlx serve <model> runs a model on your Mac's GPU with MLX and serves it at http://localhost:8000/v1: port 8000, the OpenAI paths under /v1. Point any OpenAI client at that base URL and use the model alias as model.

$ brew install rapid-mlx                # or: curl -fsSL https://rapidmlx.com/install.sh | bash
$ rapid-mlx serve qwen3.5-4b-4bit
$ curl http://localhost:8000/v1/chat/completions -H "Content-Type: application/json" \
    -d '{"model": "qwen3.5-4b-4bit", "messages": [{"role": "user", "content": "Hello"}]}'
SettingValue
OpenAI base URLhttp://localhost:8000/v1
Anthropic base URLhttp://localhost:8000 (no /v1; the SDK adds /v1/messages)
Port8000. If it's taken, the next free port up to 8009, and the startup banner says which. --port N pins it.
Host127.0.0.1, this Mac only. --host 0.0.0.0 opens it to your network.
API keyNot checked by default; SDKs need any non-empty string. Set --api-key or RAPID_MLX_API_KEY to require one.
Model nameThe alias you served (qwen3.5-4b-4bit) or its Hugging Face repo id. GET /v1/models lists both.

Start the server

Install with Homebrew (brew install rapid-mlx) or the one-line installer, then serve a model by alias or by Hugging Face repo id. The first run downloads the weights. When the server is up it prints the URLs to use:

$ rapid-mlx serve qwen3.5-4b-4bit
…
  Ready: http://127.0.0.1:8000

  OpenAI:    http://127.0.0.1:8000/v1
  Anthropic: http://127.0.0.1:8000
  Model:     qwen3.5-4b-4bit

localhost and 127.0.0.1 are the same address here. Not sure which model fits your Mac? rapid-mlx recipe prints the pick for your memory, and Pick a model for my Mac explains each tier. In another terminal, rapid-mlx connect prints the same URLs for a server that is already running.

Port and base URL

Endpoints

Paths are relative to http://localhost:8000.

EndpointAPINotes
POST /v1/chat/completionsOpenAI Chat CompletionsStreaming, tool calling, vision inputs, reasoning_content.
POST /v1/responsesOpenAI ResponsesWhat Codex CLI speaks; streaming events, reasoning and tool-call items.
POST /v1/messagesAnthropic MessagesWhat Claude Code speaks; tools, streaming, thinking blocks. Plus /v1/messages/count_tokens.
POST /v1/completionsOpenAI CompletionsRaw text completion, no chat template.
GET /v1/modelsOpenAI ModelsThe served model, with its context window and capabilities.
POST /v1/embeddingsOpenAI EmbeddingsAttach an embedding model with --embedding-model (EmbeddingGemma 2 needs the [vision] extra); without one it returns an error that says so. See EmbeddingGemma 2.
POST /v1/audio/speech, /v1/audio/transcriptionsOpenAI AudioText-to-speech and speech-to-text with an audio model; needs the [audio] extra.
POST /v1/images/generationsOpenAI ImagesText-to-image with an image model; needs the [image] extra.
GET /healthProbeHealth check; also /healthz, /readyz, /livez.
GET /docs, /openapi.jsonSchemaSwagger UI and the OpenAPI schema of the running server.

Video, audio translation and music, image edits, metrics and model management are in the full API reference, with every place the server differs from OpenAI.

Call it from curl, Python and Node

curl

$ curl http://localhost:8000/v1/models
$ curl http://localhost:8000/v1/chat/completions \
    -H "Content-Type: application/json" \
    -d '{"model": "qwen3.5-4b-4bit", "messages": [{"role": "user", "content": "Hello"}]}'

Python (openai)

from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="not-needed")

r = client.chat.completions.create(
    model="qwen3.5-4b-4bit",
    messages=[{"role": "user", "content": "Hello"}],
)
print(r.choices[0].message.content)

# the Responses API works on the same client
print(client.responses.create(model="qwen3.5-4b-4bit", input="Hello").output_text)

Node / TypeScript (openai)

import OpenAI from "openai";

const client = new OpenAI({ baseURL: "http://localhost:8000/v1", apiKey: "not-needed" });
const r = await client.chat.completions.create({
  model: "qwen3.5-4b-4bit",
  messages: [{ role: "user", content: "Hello" }],
});
console.log(r.choices[0].message.content);

Without changing code

The OpenAI SDKs read the base URL and key from the environment:

$ export OPENAI_BASE_URL=http://localhost:8000/v1
$ export OPENAI_API_KEY=not-needed
$ python your_script.py
Tested 2026-10-07 with rapid-mlx 0.15.7 and qwen3.5-4b-4bit on a Mac mini (M2 Pro, 32 GB): the startup banner and port fallback above, the curl calls, the Python snippet (openai 3.26.0) and the Node snippet (openai 7.30.0) all returned replies.

API key and network access

A plain rapid-mlx serve doesn't check keys, and it only listens on 127.0.0.1. To require a key, set it when you start the server:

$ export RAPID_MLX_API_KEY="$(openssl rand -hex 24)"
$ rapid-mlx serve qwen3.5-4b-4bit            # or: --api-key <key>

Clients then send Authorization: Bearer <key> (what the OpenAI SDKs do with api_key); on /v1/messages, Anthropic's x-api-key header works too. A missing or wrong key gets 401; the health probes stay open. Set a key before you use --host 0.0.0.0 or put the server behind a tunnel.

How it differs from mlx_lm.server

mlx_lm.server, the server that ships with Apple's mlx-lm, is also OpenAI-compatible. Rapid-MLX runs the same MLX weights; the differences are in what the server exposes (mlx-lm 0.32.0 checked):

Rapid-MLXmlx_lm.server
Startrapid-mlx serve qwen3.5-4b-4bitmlx_lm.server --model <repo>
Default port8000 (next free up to 8009)8080
OpenAI routesChat Completions, Completions, Models, Responses, Embeddings, Audio, ImagesChat Completions, Completions, Models
Anthropic Messages/v1/messagesNo
API key--api-key / RAPID_MLX_API_KEYNo option; its own help says it is not recommended for production
Coding agentsOne setup command per agent (guides)Configure by hand

Moving a client over means changing the base URL from http://localhost:8080/v1 to http://localhost:8000/v1 and the model name to an alias (or keep the repo id). On speed, Rapid-MLX 0.15.6 measured up to 4× faster than mlx-lm 0.32.0 and 1.5× on a typical task, with the same weights: Qwen3.5-9B 4-bit on a Mac mini M4 Pro (48 GB), best case whole-file code edits, long-prompt agent turns close to a tie. Every task and the raw data are on Rapid-MLX vs mlx-lm.

Frequently asked questions

What port does the MLX OpenAI-compatible server use?

Rapid-MLX listens on port 8000 by default, so the OpenAI base URL is http://localhost:8000/v1. If 8000 is taken it uses the next free port up to 8009 and prints it; --port sets it explicitly. mlx-lm's own mlx_lm.server defaults to port 8080.

What is the base URL for a local OpenAI-compatible server on Mac?

http://localhost:8000/v1 for OpenAI clients (Python, Node, LangChain, Codex CLI) and http://localhost:8000 for Anthropic clients such as Claude Code. Use the alias you served as the model name.

Does it need an OpenAI API key?

No. The server doesn't check keys unless you start it with --api-key or RAPID_MLX_API_KEY. The OpenAI SDKs refuse an empty key, so pass any placeholder string such as not-needed.

Which OpenAI endpoints are supported?

Chat Completions, Completions, Models, Responses, Embeddings (with an embedding model attached), Audio speech and transcription, and Images, plus Anthropic Messages. The API reference lists every route and where it differs from OpenAI.

Can I use it from another computer on my network?

Yes. Start it with --host 0.0.0.0 and an API key, then use http://<your-mac's-ip>:8000/v1 from the other machine. By default it only accepts connections from the Mac itself.

Is Rapid-MLX a drop-in replacement for mlx_lm.server?

For OpenAI clients, yes: change the base URL from port 8080 to 8000 and set the model name. It runs the same MLX weights and adds the Responses and Anthropic Messages APIs, API keys, and embeddings, audio and image routes.

Next