Get started

Install & first request

Install Rapid-MLX, start a server, and make your first OpenAI-compatible chat completion.

$ curl -fsSL https://rapidmlx.com/install.sh | bash
$ rapid-mlx serve qwen3.5-4b-4bit
$ curl http://localhost:8000/v1/chat/completions -H "Content-Type: application/json" \
    -d '{"model":"qwen3.5-4b-4bit","messages":[{"role":"user","content":"ping"}]}'

That is the whole path. The sections below explain each step and the options you have.

Requirements

Intel Macs are not supported — Rapid-MLX is built on Apple's MLX framework which targets the unified-memory Apple Silicon GPU.

Install

One curl — no sudo, nothing to sign up for.

$ curl -fsSL https://rapidmlx.com/install.sh | bash

Prefer Homebrew? It's in homebrew/core — no tap, no trust:

$ brew install rapid-mlx

The installer looks for Python 3.10+ (installing Python 3.12 through Homebrew, or a standalone build, if none is found), creates a venv at ~/.rapid-mlx, symlinks the rapid-mlx command into ~/.local/bin, and suggests a starter model for your Mac's RAM. You can also install from pip into a venv you manage yourself:

$ python3 -m venv .venv && source .venv/bin/activate
$ pip install rapid-mlx

Not sure where to start? Run rapid-mlx with no command. It shows one screen with your Mac's chip and RAM and the model it will chat with, then waits for a key: Enter chats (the quick-start qwen3.5-4b-4bit on a cold cache, or a recommended model you already have), b picks the best model for this Mac, c connects a detected coding agent, s starts a server and m chooses another model. Each key prints the exact command it runs; the CLI reference has the details.

Start a server

Pick any alias from the catalog and pass it to rapid-mlx serve. Aliases are short memorable names that resolve to the canonical Hugging Face repo and apply the right tool-call parser, reasoning parser and performance defaults. A Hugging Face org/name works too (see Bring your own model):

$ rapid-mlx serve qwen3.5-4b-4bit
# downloads the MLX weights on first run, then:
INFO:     Uvicorn running on http://127.0.0.1:8000 (Press CTRL+C to quit)

The server listens on 127.0.0.1 only. Without --port it takes the first free port from 8000 to 8009 (8000 unless something already uses it); an explicit --port 9000 never falls back. Set the served name with --served-model-name my-model, or pick a quant explicitly with the full alias suffix (e.g. qwen3.5-9b-8bit). --host 0.0.0.0 exposes the server to your network — set --api-key first.

Make your first request

The endpoints are drop-in OpenAI-compatible — point any OpenAI client at http://localhost:8000/v1 with any non-empty API key.

cURL

$ curl http://localhost:8000/v1/chat/completions \
    -H "Content-Type: application/json" \
    -d '{
      "model": "qwen3.5-4b-4bit",
      "messages": [{"role": "user", "content": "ping"}]
    }'

Python (openai SDK)

from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8000/v1",
    api_key="not-used",
)
resp = client.chat.completions.create(
    model="qwen3.5-4b-4bit",
    messages=[{"role": "user", "content": "ping"}],
)
print(resp.choices[0].message.content)

What is on by default

Each alias carries a tuned profile, so these apply without any extra flag on the aliases that qualify for them:

PFlash long-prompt compression is off for every model; opt in with --pflash auto or --pflash always.

rapid-mlx info <alias> shows exactly what a given alias turns on. The performance flags reference covers every technique and its opt-in variants.

Wire it into a tool

Anything that speaks the OpenAI API works without code changes — point its base URL at http://localhost:8000/v1. For coding agents there is a setup command per agent, and a one-step start:

$ rapid-mlx agents                      # list supported agents
$ rapid-mlx agents codex --setup        # point Codex CLI at the running server
$ rapid-mlx start claude-code           # start a server and configure Claude Code

start picks a recommended model that fits your Mac (and is already downloaded, if one is), starts the server and configures the agent; add --dry-run to preview it. The agent guides list what each setup writes.

Next steps