CLI reference
Every subcommand shipped in rapid-mlx, grouped by
role. Captured verbatim from rapid-mlx <subcmd> --help
on a fresh pip install rapid-mlx venv — so
what is on this page is exactly what the binary reports.
The five commands you'll use
$ rapid-mlx serve qwen3.5-9b-4bit # start the OpenAI/Anthropic-compatible server $ rapid-mlx chat qwen3.5-9b-4bit # chat in the terminal $ rapid-mlx pull qwen3.5-9b-4bit # download only $ rapid-mlx import Qwen/Qwen3-0.6B --quantize 4 # convert a bf16 model to 4-bit MLX $ rapid-mlx agents claude-code --setup # point an agent at the server
Details: serve ·
chat ·
pull ·
import ·
agents. Everything else is below,
grouped by role.
Running rapid-mlx with no command
In a terminal, bare rapid-mlx opens a one-screen start
menu: your Mac's chip and RAM, the model Enter will chat with, and
single-key actions. Nothing downloads or starts until you press a key,
and each action prints the exact command it runs.
| Key | Action |
|---|---|
| Enter | Chat with the ready model: a running server's model, your last-used model, or a downloaded recommended model; on a cold cache, the quick-start qwen3.5-4b-4bit (about 3 GB download). |
| b | Chat with the best model for this Mac (the rapid-mlx recipe pick), when it differs. |
| c | Connect a detected coding agent, starting a server for it if none is running (shown only when an agent is detected). |
| s | Start the server for your own tools. |
| m | Choose another model. |
| q | Quit. |
Without a terminal (a pipe, CI, TERM=dumb) it prints the
recommended models and copy-paste commands instead and exits 1.
rapid-mlx --help lists the commands grouped by task: get
started, coding agents, manage, and advanced / experimental.
Every subcommand accepts -h / --help. The
top-level command also honours --version, a global
--no-telemetry escape hatch (equivalent to
RAPID_MLX_TELEMETRY=0 for the current run), and
--disable-version-check (equivalent to
RAPID_MLX_DISABLE_VERSION_CHECK=1; servers started by
chat, start and share inherit
it). Global options go before the subcommand:
rapid-mlx --disable-version-check serve qwen3.5-4b-4bit.
Server
serve
Start the OpenAI-compatible HTTP server. This is the workhorse — every chat / responses / embeddings / audio route lives behind it.
$ rapid-mlx serve qwen3.5-4b-4bit $ rapid-mlx serve qwen3.6-35b-8bit --port 9000 --host 0.0.0.0 $ rapid-mlx serve gpt-oss-20b-mxfp4-q8 --enable-auto-tool-choice \ --tool-call-parser harmony
serve has a large surface (100+ flags across host binding,
KV cache, speculative decode, tool / reasoning parsing, sampler
defaults, PFlash long-prompt compression, and MCP wiring). Highlights below; the complete flag list is one
rapid-mlx serve --help away.
| Category | Key flags |
|---|---|
| Binding | --host (default 127.0.0.1), --port, --listen-fd FD (for launchd/systemd socket activation), --served-model-name, --api-key, --cors-origins, --rate-limit, --max-request-bytes, --timeout, --log-level (falls back to RAPID_MLX_LOG_LEVEL), --log-file TARGET (all server output to a file, appended; - for stdout, /dev/null to discard; falls back to RAPID_MLX_LOG_FILE) |
| Batching & concurrency | --max-num-seqs, --max-concurrent-requests (default 256; above it, HTTP 503 "Server is busy" with Retry-After), --prefill-batch-size, --completion-batch-size, --stream-interval (continuous batching is always on) |
| KV cache | --kv-cache-dtype {bf16,int8,int4} (default bf16), --kv-cache-turboquant [v4|k8v4|none] (v4 default, 4.6× compression on k8v4), --reasoning (pins int8 for AIME-class math), --enable-prefix-cache (on by default), --prefix-cache-index {radix,hash}, --cache-memory-mb, --disable-disk-caches (see Disk writes) |
| Speculative decode | --speculative-config '{...}' — one vLLM-style JSON knob for every method: {"method":"mtp"}, {"method":"dflash"}, {"method":"ddtree"}, {"method":"suffix", "num_speculative_tokens":8}, {"method":"dspark", "num_speculative_tokens":5} (DeepSeek V4, checkpoint-native — needs a local checkpoint path). Plus --force-spec-decode / --no-spec-decode to override the per-model eligibility profile. |
| Parsers | --enable-auto-tool-choice, --tool-call-parser (auto, hermes, qwen3, qwen3_coder, harmony/gpt_oss, gemma4, deepseek_v31, kimi_k2, minimax, ui_tars, …), --reasoning-parser NAME (NAME: gemma4, qwen3, hy3, deepseek_r1, deepseek_v4, vibethinker, glm4, gpt_oss, harmony, minimax, ui_tars), --no-thinking, --no-tool-call-parser, --no-reasoning-parser |
| PFlash long prompt | --pflash {off,auto,always} (default off for every model; auto / always opt in), --pflash-threshold (default 32k), --pflash-keep-ratio (default 0.20), --pflash-sink-tokens, --pflash-tail-tokens, --pflash-include-tools |
| Embeddings | --embedding-model, --embedding-max-length {auto,<int>} (default auto — derived from config.max_position_embeddings, else tokenizer.model_max_length; an explicit value is clamped to the model maximum), --embedding-overflow-policy {truncate,error} (default truncate, which logs and increments rapid_mlx_embedding_truncations_total; error returns HTTP 400 input_too_long) |
| MCP + tools | --mcp-config <path>, --enable-tool-logits-bias |
| Escape hatches | --force-hybrid / --no-hybrid, --force-openai-harmony-streaming / --no-openai-harmony-streaming, --mllm / --no-mllm, --force-disk-check, --watchdog-ppid PID |
Models outside the catalog. For a Hugging Face repo that isn't
a catalog alias and isn't cached yet, serve first runs a
metadata-only check: format, architecture and memory fit. Before
downloading, it refuses a model it can prove won't run on this Mac.
--no-preflight skips the check. --request
files an opt-in support request without asking when a public model is
refused for its architecture or format. It sends the repo id,
architecture, format, failure class and Rapid-MLX version.
--disk-stream serves MoE experts from disk, so the memory
refusal doesn't apply. See Bring your own model.
Disk writes
Beyond the model download, serve can write these to the SSD:
| What is written | When | How to turn it off |
|---|---|---|
Prefix-cache snapshot in ~/.cache/rapid-mlx/prefix_cache/ (rewritten whole each time) | Each server shutdown, including the end of a chat session, and each --idle-unload-seconds unload | --disable-disk-caches, or --disable-prefix-cache (also turns off the in-memory cache) |
KV checkpoints in ~/.cache/rapid-mlx/kv_checkpoints/ | Only with --kv-disk-checkpoint-interval N > 0 | Off by default; --disable-disk-caches overrules an interval |
Vision prefix-cache disk tier in ~/.cache/mlx-vlm/apc/ | Only when APC_DISK_ENABLED=1 is set | Off by default; --disable-disk-caches overrules it |
| Server log output | Continuously | --log-file /dev/null, or a path on a RAM disk |
--disable-disk-caches (or RAPID_MLX_DISABLE_DISK_CACHES=1)
also stops a saved prefix-cache snapshot from loading at startup; the
in-memory prefix cache keeps working. The server logs which explicit
settings it overrules. When it does persist the prefix cache, it skips
entries that would leave less than 5 GiB free on the volume
(RAPID_MLX_PREFIX_CACHE_MIN_FREE_DISK_BYTES; 0
turns the check off).
share
Start rapid-mlx serve and open a public Cloudflare-fronted
URL on rapidmlx.com so a friend on a different device can hit
the model. Uses the built-in websockets reverse tunnel — no
port-forwarding.
$ rapid-mlx share qwen3.5-4b-4bit # prints URL + share key + one-click chat link, Ctrl-C stops
Flags: --port (default 8765), --thinking / --no-thinking
(default off, so chat UIs see content immediately),
--cors-origins (defaults to the rapidmlx.com chat-frontend allowlist),
--rate-limit RPM (default 120),
--chat-frontend URL (default https://rapid-pro.quicksilverpro.io;
pass empty to suppress), --disable-disk-caches and
--log-file (passed to the server it starts).
launch
One-shot bootstrap: patch a detected IDE / agent client (Cline, Claude
Code, Continue, Cursor) so it routes at the local rapid-mlx server.
rapid-mlx launch list prints the detection matrix.
$ rapid-mlx launch cursor $ rapid-mlx launch claude-code --model qwen3.6-35b-8bit $ rapid-mlx launch --all --start-server --port 8000
Flags: client (or list), --all,
--model (default $RAPID_MLX_DEFAULT_MODEL or
qwen3.5-4b-4bit),
--server-url (default http://127.0.0.1:8000),
--port, --start-server, --dry-run,
--json (with list). With
RAPID_MLX_API_KEY exported, the key is used for the
client config and passed to a server started with
--start-server. For Claude Code, launch also
records the model's context window when the server (or the local cache)
reports it.
start
Start a server for an agent and configure the agent in one command: picks the agent's first recommended model that fits this Mac and is already cached (or previews the download and asks), runs the server in the foreground, and writes the agent's config once the server is ready.
$ rapid-mlx start claude-code $ rapid-mlx start pi --dry-run # model, port and config changes; starts nothing $ rapid-mlx start # generic OpenAI-compatible endpoint
Flags: profile (agent name; omit for a generic endpoint),
--model, --port (default 8000),
--host (default 127.0.0.1), --no-download
(fail unless a recommended model is cached), --dry-run,
--yes / -y, --no-setup (only
print the agent's instructions), --ready-timeout (default
600 s), --disable-disk-caches and --log-file
(passed to the server it starts; see serve).
connect
Print the running server's connection details (OpenAI and Anthropic base URLs, model), or the setup steps for one tool.
$ rapid-mlx connect $ rapid-mlx connect claude-code $ rapid-mlx connect --json
Flags: target (claude-code,
continue or openai-python), --json,
--host, --port, --model,
--base-url (the live server's OpenAI-style URL, e.g.
http://localhost:8123/v1).
Chat
chat alias: run
Interactive REPL against a spawned or already-running server. Defaults
to qwen3.5-4b-4bit when the model arg is omitted. The run
alias is Ollama-parity. Answers render as proper terminal Markdown
(headings, tables, syntax-highlighted code) with a live token counter
while streaming; pipes and NO_COLOR keep plain text.
$ rapid-mlx chat $ rapid-mlx chat qwen3.5-9b-4bit --think $ rapid-mlx chat --port 8000 # connect to existing server $ rapid-mlx chat --mcp-config ~/mcp.json # give the REPL MCP tools $ rapid-mlx run gemma-4-12b-4bit # Ollama-style alias
Flags: --system, --think / --no-think
(default off in the REPL to avoid reasoning-model CoT leak),
--mcp-config <path> (connect MCP servers so the REPL can
call their tools — parallel across servers, with live tool-activity),
--max-tokens (default 2048; 4096 with --think),
--temperature (default 0.7),
--mcp-max-rounds (tool-call rounds per turn, default 8),
--context-length (per-request window for the server chat
starts), --disable-prefix-cache (no prefix cache in the
server chat starts, in memory or on disk),
--disable-disk-caches (keep the in-memory cache, write no
optional caches to disk), --log-file (where the spawned
server's output goes; without it chat keeps that output in a temporary
log file, and - prints it into the session),
--port, --base-url,
--ready-timeout / --response-timeout (default 600 s each).
Model management
models
List every alias registered in rapid_mlx/aliases.json
with per-alias tool/reasoning parser, spec-decode
eligibility, suffix-decoding tier, and DFlash / DDTree readiness.
$ rapid-mlx models # all aliases $ rapid-mlx models --cached # same as `rapid-mlx ls` $ rapid-mlx models --search qwen $ rapid-mlx models --modality audio
Flags: --cached, --json (machine-readable;
pairs with --cached), --search TERM
(alias substring), --modality {text,video-gen,image-gen,audio}.
ls
List every model in the local HuggingFace cache — alias (or
(unmapped)), HF repo, on-disk size, last modified. Equivalent
to rapid-mlx models --cached.
pull
Download a model into the HuggingFace cache — no server.
$ rapid-mlx pull qwen3.5-9b-4bit $ rapid-mlx pull mlx-community/gemma-4-12b-it-4bit $ rapid-mlx pull Qwen/Qwen3-0.6B --no-preflight
Flags: model (alias or org/name),
--bits N (pull only the <N>bit/ variant of
a multi-variant repo), --format name (only the named format
variant, e.g. mxfp4; gguf is refused because
Rapid-MLX can't run GGUF files), --request,
--no-preflight. Like serve, pull
checks models outside the catalog before downloading (format and
architecture). It reports memory fit but doesn't refuse on it, since you
may be pulling for another Mac. The check is skipped with
--bits / --format.
import
Convert a bf16/fp16 safetensors model to quantized MLX: explicit and
cancel-safe. serve and pull never convert.
$ rapid-mlx import Qwen/Qwen3-0.6B --quantize 4 $ rapid-mlx import ~/models/my-finetune --quantize 4 --name my-ft-4bit $ rapid-mlx serve my-ft-4bit
Flags: source (Hugging Face org/name or a
local model directory), --quantize BITS (2, 3, 4, 6 or 8;
default 4), --name NAME (default
<source-name>-<bits>bit, with -local
appended if that is already an alias), --force (replace an
existing import of the same name from another source).
Before downloading or converting, it checks free disk and memory. The
conversion and a one-token smoke test run in a temporary directory, and
only a passing model is published to
~/.rapid-mlx/imports/<name>. Ctrl-C leaves the cache
unchanged, and re-running an identical import is a no-op. Imports show
up in rapid-mlx models --cached and are served and removed
by name. Full walkthrough: Bring your own model.
rm
Delete a cached model, or an import by its name. Confirms first; -y skips the prompt.
$ rapid-mlx rm qwen3.5-9b-4bit -y
info
Print the per-alias profile — HF repo, tool/reasoning parsers, hybrid flag, MoE flag, spec-decode support, PFlash tier — for one alias or raw HF repo.
$ rapid-mlx info qwen3.6-35b-8bit $ rapid-mlx info mlx-community/SmolLM3-3B-4bit
ps
List every running rapid-mlx serve process on this
machine — PID, port, loaded model, uptime.
recipe
Recommend the smart and the fast model for this Mac's memory. See hardware tiers for every tier.
$ rapid-mlx recipe $ rapid-mlx recipe --max-ram 32 --json # plan for another Mac
alias
Give a model your own short name. User aliases live in
~/.config/rapid-mlx/user-aliases.json and point at a
catalog alias or a Hugging Face repo id.
$ rapid-mlx alias set my-coder qwen3-coder-30b-4bit $ rapid-mlx alias list $ rapid-mlx alias remove my-coder
Benchmark
benchmark run + benchmark share is the
local-first Community Benchmark flow and the only one
that feeds the leaderboard: a shared run lands
in your Mac's row, next to everyone else with the same chip and memory,
and on your contributor page. Leaderboard cells the current benchmark has
not measured show legacy bench --submit runs, badged with
their rapid-mlx version and never mixed into current medians.
bench freeform + validation tiers · --submit removed
Run a benchmark against a model. Freeform mode by default (any
--num-prompts / --max-tokens / batching combo),
with validation tiers. --submit is removed: it exits with an
error and sends nothing. Use benchmark run + share
to put a run on the leaderboard.
$ rapid-mlx bench qwen3.5-4b-4bit $ rapid-mlx bench gpt-oss-20b-mxfp4-q8 --tier speed $ rapid-mlx bench gpt-oss-20b-mxfp4-q8 --tier all # smoke → speed → harness
Flags: freeform batching / KV / prefix-cache flags mirror
serve (they configure the throwaway server bench spawns);
--submit is removed: it prints the replacement
(rapid-mlx benchmark run <model>, then
rapid-mlx benchmark share <run-id>) and exits 2, and
nothing is run or uploaded. The flags that only shaped a submission
(--spec-decode, --run-group, --sampled,
--notes, --repo-root) are still listed in
--help;
--tier {smoke,speed,harness,all} runs a validation ladder
(harness drives 5 first-class agent flows: codex/opencode/qwen-code/hermes/aider);
--base-url reuses an already-running server for --tier;
--long-prompt-tokens + --pflash replicate the long-prompt
TTFT profile from PR #649.
benchmark Community Benchmark · local-first → the leaderboard + your contributor page
Run or inspect reproducible local benchmarks. Every run is saved on your
Mac first (~/.rapid-mlx/benchmarks/, or
$RAPID_MLX_BENCHMARK_HOME); nothing is uploaded until you
run share on a specific run id, and share
shows you the exact payload and asks before sending. Six subcommands,
each with --json for scripts.
$ rapid-mlx benchmark catalog # models with a protocol + fit for this Mac $ rapid-mlx benchmark plan qwen3.5-9b-4bit # exact workload; runs nothing $ rapid-mlx benchmark run qwen3.5-9b-4bit # run, save locally, print run id $ rapid-mlx benchmark results --limit 5 # local runs + share receipts $ rapid-mlx benchmark inspect RUN_ID # one saved run, in full $ rapid-mlx benchmark share --preview RUN_ID # exact payload, nothing sent $ rapid-mlx benchmark share RUN_ID # show payload, ask [y/N], upload
catalog [--all] [--memory-gib N] [--json] — lists
the recommended models for this Mac (the ★ focus models
that fit), then a one-line count of everything else; --all
prints every model with a registered protocol (text, image and video
workloads). The fit column is computed against this Mac's unified
memory, or against --memory-gib when you are planning for
another machine.
plan <model> [--json] — prints the model,
task, protocol id and version, and every case the run will measure (for
text: pp512-tg128 and pp2048-tg512, one warmup
round then five measured rounds each). Runs nothing.
run <model> [--json] — runs the registered
protocol and saves the result locally. While it runs it reports the
model load, an estimated time remaining after the warmup round, and
each round's tok/s; at the end it prints
Saved local result <run_id>, the measurements, and
Nothing was uploaded. (--json keeps stdout to
the result document; progress goes to stderr.)
results [--limit N] [--json] — lists locally saved
runs, newest first. The JSON form also carries a receipts
map, so you can see which runs have been shared.
inspect <run_id> [--json] — prints one saved
result in full.
share <run_id> [--preview] [--yes] [--json] —
explicitly uploads one saved run. --preview prints the
exact payload (plus its digests) and sends nothing. Without
--yes the payload is printed and you are asked
Upload this result? [y/N]; --yes is for
callers that already showed their own consent dialog, and
--json requires --yes or
--preview.
share sends — and never sends.
Published: chip, RAM, core counts, OS and package versions, the model
(its Hugging Face repo id), execution settings, the timings, and a
random resettable install id. Never sent: your name, hostname, serial
number, hardware UUID, prompts, outputs, or file paths — the payload
has no IP-address field. Like any HTTPS service, the endpoint observes
the source IP; it uses it transiently for abuse limits and never
stores it in the benchmark record. share --preview shows
you the exact bytes before you decide.
Receipts and idempotency. An accepted upload returns a receipt
(submission_id, accepted_at,
run_digest, already_exists, and your
pseudonymous contributor name + tag). The CLI checks that
the receipt names the run and payload it just sent, then saves it under
~/.rapid-mlx/benchmarks/receipts/ so
results can show it. Sharing the same run again is safe: a
run id with identical content is accepted again with
already_exists: true ("already uploaded") and does not
create a second row; the same run id with different content is
rejected (409), and a run that was retracted stays gone (410). Each
install also has a daily acceptance cap.
Your board name. The pseudonym on your contributor page
(/leaderboard/contributors/<name>-<tag>) is
derived from a random per-install id stored at
~/.rapid-mlx/bench-install-id, created the first time a
share is confirmed. It is not a serial number and identifies nothing
but that file. Delete the file and the next share is issued a new id and
a new name; runs shared under the old name stay under the old name.
Agent / IDE integration
agents
List, configure, and integration-test agent integrations: claude-code,
codex, pi, opencode, hermes, deepseek-harness (dsh),
qwen-code, kilo-code, aider, openhands, continue (or
continue-dev), plus the langchain, pydanticai and smolagents
framework profiles. Per-agent pages: agent guides.
$ rapid-mlx agents # list $ rapid-mlx agents codex # setup guide $ rapid-mlx agents pi --setup --dry-run # preview, write nothing $ rapid-mlx agents pi --setup # auto-configure $ rapid-mlx agents hermes --test # run integration flow $ rapid-mlx agents aider --agent-version 0.90.0 --setup
Flags: agent_name (omit to list),
--setup, --dry-run (preview setup without
writing), --yes / -y (apply without the
confirmation prompt), --no-check (skip the server health
and model check, to write a config while no server is running),
--test,
--model (auto-detected from the running server),
--base-url (default http://localhost:8000/v1),
--agent-version.
Diagnostics
doctor
One-shot environment check — Apple Silicon detection, macOS + Darwin versions, free disk, HF cache size, Python version, mlx / mlx-lm / transformers / fastapi / uvicorn / rapid-mlx versions, optional-extras status, network probes — and can apply verified repairs.
$ rapid-mlx doctor $ rapid-mlx doctor --fix --dry-run # show the repair plan $ rapid-mlx doctor --fix # apply verified repairs only
Flags: --verbose / -v (probe detail),
--json (versioned machine-readable report),
--summary (one line), --deep (dependency,
DNS and route probes, up to 30 s), --fix with
--dry-run / --yes, --only SECTION
/ --skip SECTION (repeatable).
telemetry
Manage anonymous usage telemetry, which is on by default after a one-time notice. See the telemetry page for the full event schema.
$ rapid-mlx telemetry status # enabled/disabled + why $ rapid-mlx telemetry off # or: disable $ rapid-mlx telemetry on # or: enable $ rapid-mlx telemetry preview # a sample payload, exactly as sent $ rapid-mlx telemetry reset-id # rotate the client id, keep the choice $ rapid-mlx telemetry reset # delete the stored choice + rotate the id
For one run: rapid-mlx --no-telemetry <command>,
RAPID_MLX_TELEMETRY=0 or DO_NOT_TRACK=1.
Maintenance
upgrade
Detect this install's method (brew, pip, install.sh) and run the right
upgrade command. --dry-run just prints what it would do;
-y skips the prompt. rapid-mlx update is the
same command.
$ rapid-mlx upgrade $ rapid-mlx upgrade --dry-run $ rapid-mlx upgrade -y
version
Print the version. Same as the top-level --version / -V.
help
Print help for a subcommand. rapid-mlx help serve is equivalent to rapid-mlx serve --help.
feedback
Open the community Discord invite in your browser; --no-open only prints the link.
Advanced and experimental
| Command | What it does |
|---|---|
service | Run Rapid-MLX as a headless macOS LaunchDaemon (install, configure, apply, status, logs, restart, upgrade, uninstall). See Headless macOS service. |
system-one | Serve a typed decision model (Laya, CLM, or Clef with clef / clef-flash) on POST /v1/systemone and POST /v1/rank. A separate process from serve. CLM is built in; Laya needs the [system-one] extra and Clef the [clef] extra. |
cua | Experimental native-accessibility computer-use agent: rapid-mlx cua run --planner local-9b. Needs the [computer-use] extra. |
Deprecated aliases
The pre-rename binaries vllm-mlx, vllm-mlx-chat,
and vllm-mlx-bench still ship as entry points and forward
to the same code — kept so muscle memory and old scripts keep working.
New usage should always be rapid-mlx, or its short form
rmlx.