API surface
Point any OpenAI client at http://localhost:8000/v1 and use
the alias you served as the model name. Any non-empty API key works
unless you start the server with --api-key.
Starting the server, the port and the base URL are covered in
OpenAI-compatible server.
from openai import OpenAI client = OpenAI(base_url="http://localhost:8000/v1", api_key="not-used") stream = client.chat.completions.create( model="qwen3.5-4b-4bit", messages=[{"role": "user", "content": "Hello"}], stream=True, ) for chunk in stream: print(chunk.choices[0].delta.content or "", end="")
The three things most clients need:
- Chat:
POST /v1/chat/completions— tools, vision, streaming andreasoning_content. - Agents:
POST /v1/responses(Codex CLI) andPOST /v1/messages(Anthropic Messages, used by Claude Code). - Discovery:
GET /v1/modelslists what the server serves, with its parsers and context window.
GET /openapi.json and an
interactive Swagger UI at GET /docs. The list below
is the human-readable summary; the JSON schema is the source of truth.
Endpoints
| Endpoint | What it does |
|---|---|
| POST /v1/chat/completions | Chat completions — tools, vision, streaming and reasoning_content. |
| POST /v1/responses | OpenAI Responses API — streaming with item-by-item events, reasoning items, tool calls. |
| POST /v1/messages | Anthropic Messages API — tools, streaming and thinking blocks. |
| POST /v1/messages/count_tokens | Anthropic token count for a Messages request. |
| POST /v1/completions | Text completion without a chat template. |
| GET /v1/models | The served models with their parsers, modality, capabilities and context window. |
| POST /v1/embeddings | Embeddings — needs an embedding model attached with --embedding-model (e.g. embeddinggemma-2-4bit, which uses the [vision] runtime). |
| POST /v1/audio/speech | Text-to-speech (Kokoro, Qwen3-TTS, IndexTTS, Chatterbox, VibeVoice, VoxCPM, F5-TTS, Dia) — needs the [audio] extra. Cloning models take a ref_audio field; Qwen3-TTS Base and F5-TTS also need ref_text. |
| POST /v1/audio/transcriptions | Speech-to-text (Whisper, Parakeet, SenseVoice, Qwen3-ASR) with word-level timestamps — needs the [audio] extra. A text field switches to forced alignment against that transcript. |
| POST /v1/audio/translations | Speech-to-English translation — needs the [audio] extra. |
| POST /v1/audio/music | Text-to-music and sound effects — needs the [audio] extra. |
| POST /v1/images/generations | Text-to-image (FLUX, Qwen-Image, Z-Image and more) — needs the [image] extra and Python 3.11+. Returns b64_json. |
| POST /v1/images/edits | Image editing on models that support it (e.g. qwen-image-edit). |
| POST /v1/videos | Create a video job (LTX-2.5 / 2.3, Wan, CogVideoX-Fun) — needs the [video] extra, ffmpeg and Python 3.11+. Returns a job id. |
| GET /v1/videos/{id} | Poll a video job's status. |
| GET /v1/videos/{id}/content | Download the finished MP4. |
| POST /v1/videos/extend | Extend an LTX-2.5 MP4 (multipart upload): returns one MP4 with the original and the appended frames, through the same job and download endpoints. Source: 24 fps, 256–1920 px in multiples of 32, an 8n+1 frame count of at least 9, up to 20 MiB; extend_frames 8–48 in steps of 8. The result may hold at most 97 frames and 24 million pixel-frames. Allow the upload with --max-request-bytes 23068672. |
| GET /v1/videos/capabilities | What the loaded video model accepts — modes, native fps, size and frame-count rules, and the extension limits. |
| POST /v1/requests/{id}/cancel | Cancel an in-flight request. |
| GET /v1/status | Running and waiting requests, cache and Metal memory. |
| GET /health | Health probe (also /healthz, /readyz, /livez, /health/ready). |
| GET /metrics | Prometheus counters and histograms. |
| GET /openapi.json | OpenAPI 3 schema for every route. |
| GET /docs | Swagger UI. |
MCP tool routes (/v1/mcp/*) appear when the server starts
with --mcp-config; model management, cache and agent-run
routes are listed in /openapi.json.
What's the same as OpenAI
- Request and response shapes for
/chat/completions,/completions,/embeddings,/responses. - SSE streaming via
stream: truewith adata: [DONE]terminator. - Function / tool calling via
toolsandtool_choice, returned as an OpenAItool_callsarray. - Vision inputs as
image_urlcontent parts on multimodal models (Gemma 4, Qwen3-VL and others). - Authorization: without
--api-key(orRAPID_MLX_API_KEY) the header is accepted but not checked; with it, a missing or wrong key returns401.
Where it differs
-
Reasoning is off unless you ask for it on chat completions.
For thinking models, a
/v1/chat/completionsrequest that sets no thinking preference — plain chat, or a request withtools— gets a direct answer. Ask for reasoning withchat_template_kwargs: {"enable_thinking": true},reasoning_effortorreasoning_max_tokens. -
reasoning_contentis a sibling ofcontent(the DeepSeek-R1 field name), not OpenAI'sreasoningblocks. -
reasoning_max_tokenscaps the thinking part of a response; the model then answers. It does not bound the answer — pair it withmax_tokensfor total length. -
Each alias uses its own tool-call format internally (Hermes, Qwen3
XML, GLM, Harmony, Gemma 4, UI-TARS, Llama, DeepSeek, MiniMax…); the
client always receives OpenAI's
tool_callsarray. -
/v1/modelsadds Rapid-MLX fields —tool_call_parser,reasoning_parser,is_hybrid,is_moe,modality,capabilities,context_window,recommended_sampling— that OpenAI-only clients ignore. -
A finished response can carry a Rapid-MLX
metricsobject:speculative_decodingwhen the request ran MTP, andprompt_compression(original_tokens,kept_tokens) when an opted-in PFlash shortened the prompt. Chat Completions, single-prompt Completions and Responses also report experimental request timing:time_to_first_token_ms(scheduler arrival to first output token, so queueing and prefill are included) and, from two tokens on,mean_itl_ms(mean interval between output tokens;1000 / mean_itl_msis that request's decode tokens/second). Streaming sends them once on the final event, whether or notstream_options.include_usageis set. Failed, cancelled and response-cache replays omit them, and/v1/messagesdoes not carry them. Themetricsobject is omitted when none of these apply. Non-streaming responses on every text API, including/v1/messages, also announce compression with theX-Rapid-MLX-Prompt-Compressed: <kept>/<original>header. -
When the server can't take a request it returns
503withRetry-After. The message says why:Server is busy (max concurrent requests reached)above--max-concurrent-requests, orServer is at its Metal memory limitwith how much memory the request needs (see Troubleshooting). -
Streaming SSE for
/v1/responsesemitsresponse.output_item.added,response.output_item.done, reasoning items and tool-call items in OpenAI's order; a diffusion model emits a block of tokens per delta instead of one.
Examples
List models
$ curl -s http://localhost:8000/v1/models | jq '.data[] | select(.id=="qwen3.5-4b-4bit")' { "id": "qwen3.5-4b-4bit", "object": "model", "owned_by": "rapid-mlx", "is_hybrid": false, "is_moe": false, "tool_call_parser": "hermes", "reasoning_parser": "qwen3", "modality": "text", "context_window": 262144, "capabilities": ["text", "tools"], … }
The list also carries the underlying Hugging Face repo id as its own
entry, so either name works as model.
Chat completion with cURL
$ curl http://localhost:8000/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{"model": "qwen3.5-4b-4bit", "messages": [{"role": "user", "content": "What is 17*23?"}], "chat_template_kwargs": {"enable_thinking": true}}'
With thinking on, the reply carries both reasoning_content
and the final content.
Embeddings
# first install: pip install 'rapid-mlx[vision]' $ rapid-mlx serve qwen3.5-4b-4bit \ --embedding-model embeddinggemma-2-4bit --port 8000 & $ curl http://localhost:8000/v1/embeddings \ -H "Content-Type: application/json" \ -d '{"model": "embeddinggemma-2-4bit", "input": "ping"}'
An embedding model is attached to a chat server — on its own,
rapid-mlx serve embeddinggemma-2-4bit exits with a hint to
use --embedding-model. Setup, task prefixes,
dimensions and a worked search example are on the
EmbeddingGemma 2 page.
Long inputs. The input limit comes from the model
(config.max_position_embeddings, else
tokenizer.model_max_length), so a 32K-context embedder uses
its full window. Inputs above the limit are never cut silently: by
default the tail is truncated with a logged warning and a
rapid_mlx_embedding_truncations_total increment on
/metrics; with --embedding-overflow-policy error
the request is rejected with a 400:
$ rapid-mlx serve qwen3.5-4b-4bit --port 8000 \ --embedding-model embeddinggemma-2-4bit \ --embedding-max-length 2048 --embedding-overflow-policy error # a 3,000-token input then returns: {"error": {"code": "input_too_long", "message": "…observed vs allowed tokens…"}}
Audio (TTS)
# first install: pip install 'rapid-mlx[audio]' $ rapid-mlx serve kokoro --port 8002 & $ curl http://localhost:8002/v1/audio/speech \ -H "Content-Type: application/json" \ -d '{"model": "kokoro", "voice": "af_heart", "input": "hello"}' \ -o speech.wav
Where to next
For every request and response field, open the Swagger UI at
http://localhost:8000/docs on your own server — it always
matches the running build. The CLI reference
covers server flags; the performance flags
reference covers what is on by default and how to change it.