Reference · HTTP API

API surface

Point any OpenAI client at http://localhost:8000/v1 and use the alias you served as the model name. Any non-empty API key works unless you start the server with --api-key. Starting the server, the port and the base URL are covered in OpenAI-compatible server.

from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="not-used")
stream = client.chat.completions.create(
    model="qwen3.5-4b-4bit",
    messages=[{"role": "user", "content": "Hello"}],
    stream=True,
)
for chunk in stream:
    print(chunk.choices[0].delta.content or "", end="")

The three things most clients need:

Live OpenAPI schema. Every running server exposes its own machine-readable schema at GET /openapi.json and an interactive Swagger UI at GET /docs. The list below is the human-readable summary; the JSON schema is the source of truth.

Endpoints

EndpointWhat it does
POST /v1/chat/completionsChat completions — tools, vision, streaming and reasoning_content.
POST /v1/responsesOpenAI Responses API — streaming with item-by-item events, reasoning items, tool calls.
POST /v1/messagesAnthropic Messages API — tools, streaming and thinking blocks.
POST /v1/messages/count_tokensAnthropic token count for a Messages request.
POST /v1/completionsText completion without a chat template.
GET /v1/modelsThe served models with their parsers, modality, capabilities and context window.
POST /v1/embeddingsEmbeddings — needs an embedding model attached with --embedding-model (e.g. embeddinggemma-2-4bit, which uses the [vision] runtime).
POST /v1/audio/speechText-to-speech (Kokoro, Qwen3-TTS, IndexTTS, Chatterbox, VibeVoice, VoxCPM, F5-TTS, Dia) — needs the [audio] extra. Cloning models take a ref_audio field; Qwen3-TTS Base and F5-TTS also need ref_text.
POST /v1/audio/transcriptionsSpeech-to-text (Whisper, Parakeet, SenseVoice, Qwen3-ASR) with word-level timestamps — needs the [audio] extra. A text field switches to forced alignment against that transcript.
POST /v1/audio/translationsSpeech-to-English translation — needs the [audio] extra.
POST /v1/audio/musicText-to-music and sound effects — needs the [audio] extra.
POST /v1/images/generationsText-to-image (FLUX, Qwen-Image, Z-Image and more) — needs the [image] extra and Python 3.11+. Returns b64_json.
POST /v1/images/editsImage editing on models that support it (e.g. qwen-image-edit).
POST /v1/videosCreate a video job (LTX-2.5 / 2.3, Wan, CogVideoX-Fun) — needs the [video] extra, ffmpeg and Python 3.11+. Returns a job id.
GET /v1/videos/{id}Poll a video job's status.
GET /v1/videos/{id}/contentDownload the finished MP4.
POST /v1/videos/extendExtend an LTX-2.5 MP4 (multipart upload): returns one MP4 with the original and the appended frames, through the same job and download endpoints. Source: 24 fps, 256–1920 px in multiples of 32, an 8n+1 frame count of at least 9, up to 20 MiB; extend_frames 8–48 in steps of 8. The result may hold at most 97 frames and 24 million pixel-frames. Allow the upload with --max-request-bytes 23068672.
GET /v1/videos/capabilitiesWhat the loaded video model accepts — modes, native fps, size and frame-count rules, and the extension limits.
POST /v1/requests/{id}/cancelCancel an in-flight request.
GET /v1/statusRunning and waiting requests, cache and Metal memory.
GET /healthHealth probe (also /healthz, /readyz, /livez, /health/ready).
GET /metricsPrometheus counters and histograms.
GET /openapi.jsonOpenAPI 3 schema for every route.
GET /docsSwagger UI.

MCP tool routes (/v1/mcp/*) appear when the server starts with --mcp-config; model management, cache and agent-run routes are listed in /openapi.json.

What's the same as OpenAI

Where it differs

Examples

List models

$ curl -s http://localhost:8000/v1/models | jq '.data[] | select(.id=="qwen3.5-4b-4bit")'
{
  "id": "qwen3.5-4b-4bit",
  "object": "model",
  "owned_by": "rapid-mlx",
  "is_hybrid": false,
  "is_moe": false,
  "tool_call_parser": "hermes",
  "reasoning_parser": "qwen3",
  "modality": "text",
  "context_window": 262144,
  "capabilities": ["text", "tools"],
  …
}

The list also carries the underlying Hugging Face repo id as its own entry, so either name works as model.

Chat completion with cURL

$ curl http://localhost:8000/v1/chat/completions \
    -H "Content-Type: application/json" \
    -d '{"model": "qwen3.5-4b-4bit", "messages": [{"role": "user", "content": "What is 17*23?"}],
         "chat_template_kwargs": {"enable_thinking": true}}'

With thinking on, the reply carries both reasoning_content and the final content.

Embeddings

# first install: pip install 'rapid-mlx[vision]'
$ rapid-mlx serve qwen3.5-4b-4bit \
    --embedding-model embeddinggemma-2-4bit --port 8000 &
$ curl http://localhost:8000/v1/embeddings \
    -H "Content-Type: application/json" \
    -d '{"model": "embeddinggemma-2-4bit", "input": "ping"}'

An embedding model is attached to a chat server — on its own, rapid-mlx serve embeddinggemma-2-4bit exits with a hint to use --embedding-model. Setup, task prefixes, dimensions and a worked search example are on the EmbeddingGemma 2 page.

Long inputs. The input limit comes from the model (config.max_position_embeddings, else tokenizer.model_max_length), so a 32K-context embedder uses its full window. Inputs above the limit are never cut silently: by default the tail is truncated with a logged warning and a rapid_mlx_embedding_truncations_total increment on /metrics; with --embedding-overflow-policy error the request is rejected with a 400:

$ rapid-mlx serve qwen3.5-4b-4bit --port 8000 \
    --embedding-model embeddinggemma-2-4bit \
    --embedding-max-length 2048 --embedding-overflow-policy error

# a 3,000-token input then returns:
{"error": {"code": "input_too_long", "message": "…observed vs allowed tokens…"}}

Audio (TTS)

# first install: pip install 'rapid-mlx[audio]'
$ rapid-mlx serve kokoro --port 8002 &
$ curl http://localhost:8002/v1/audio/speech \
    -H "Content-Type: application/json" \
    -d '{"model": "kokoro", "voice": "af_heart", "input": "hello"}' \
    -o speech.wav

Where to next

For every request and response field, open the Swagger UI at http://localhost:8000/docs on your own server — it always matches the running build. The CLI reference covers server flags; the performance flags reference covers what is on by default and how to change it.