Reference · rapid-mlx 0.15.7 · ← Back to README

CLI reference

Every subcommand shipped in rapid-mlx, grouped by role. Captured verbatim from rapid-mlx <subcmd> --help on a fresh pip install rapid-mlx venv — so what is on this page is exactly what the binary reports.

The five commands you'll use

$ rapid-mlx serve qwen3.5-9b-4bit                # start the OpenAI/Anthropic-compatible server
$ rapid-mlx chat qwen3.5-9b-4bit                 # chat in the terminal
$ rapid-mlx pull qwen3.5-9b-4bit                 # download only
$ rapid-mlx import Qwen/Qwen3-0.6B --quantize 4  # convert a bf16 model to 4-bit MLX
$ rapid-mlx agents claude-code --setup           # point an agent at the server

Details: serve · chat · pull · import · agents. Everything else is below, grouped by role.

Running rapid-mlx with no command

In a terminal, bare rapid-mlx opens a one-screen start menu: your Mac's chip and RAM, the model Enter will chat with, and single-key actions. Nothing downloads or starts until you press a key, and each action prints the exact command it runs.

KeyAction
EnterChat with the ready model: a running server's model, your last-used model, or a downloaded recommended model; on a cold cache, the quick-start qwen3.5-4b-4bit (about 3 GB download).
bChat with the best model for this Mac (the rapid-mlx recipe pick), when it differs.
cConnect a detected coding agent, starting a server for it if none is running (shown only when an agent is detected).
sStart the server for your own tools.
mChoose another model.
qQuit.

Without a terminal (a pipe, CI, TERM=dumb) it prints the recommended models and copy-paste commands instead and exits 1. rapid-mlx --help lists the commands grouped by task: get started, coding agents, manage, and advanced / experimental.

Every subcommand accepts -h / --help. The top-level command also honours --version, a global --no-telemetry escape hatch (equivalent to RAPID_MLX_TELEMETRY=0 for the current run), and --disable-version-check (equivalent to RAPID_MLX_DISABLE_VERSION_CHECK=1; servers started by chat, start and share inherit it). Global options go before the subcommand: rapid-mlx --disable-version-check serve qwen3.5-4b-4bit.

Server

serve

Start the OpenAI-compatible HTTP server. This is the workhorse — every chat / responses / embeddings / audio route lives behind it.

$ rapid-mlx serve qwen3.5-4b-4bit
$ rapid-mlx serve qwen3.6-35b-8bit --port 9000 --host 0.0.0.0
$ rapid-mlx serve gpt-oss-20b-mxfp4-q8 --enable-auto-tool-choice \
    --tool-call-parser harmony

serve has a large surface (100+ flags across host binding, KV cache, speculative decode, tool / reasoning parsing, sampler defaults, PFlash long-prompt compression, and MCP wiring). Highlights below; the complete flag list is one rapid-mlx serve --help away.

CategoryKey flags
Binding --host (default 127.0.0.1), --port, --listen-fd FD (for launchd/systemd socket activation), --served-model-name, --api-key, --cors-origins, --rate-limit, --max-request-bytes, --timeout, --log-level (falls back to RAPID_MLX_LOG_LEVEL), --log-file TARGET (all server output to a file, appended; - for stdout, /dev/null to discard; falls back to RAPID_MLX_LOG_FILE)
Batching & concurrency --max-num-seqs, --max-concurrent-requests (default 256; above it, HTTP 503 "Server is busy" with Retry-After), --prefill-batch-size, --completion-batch-size, --stream-interval (continuous batching is always on)
KV cache --kv-cache-dtype {bf16,int8,int4} (default bf16), --kv-cache-turboquant [v4|k8v4|none] (v4 default, 4.6× compression on k8v4), --reasoning (pins int8 for AIME-class math), --enable-prefix-cache (on by default), --prefix-cache-index {radix,hash}, --cache-memory-mb, --disable-disk-caches (see Disk writes)
Speculative decode --speculative-config '{...}' — one vLLM-style JSON knob for every method: {"method":"mtp"}, {"method":"dflash"}, {"method":"ddtree"}, {"method":"suffix", "num_speculative_tokens":8}, {"method":"dspark", "num_speculative_tokens":5} (DeepSeek V4, checkpoint-native — needs a local checkpoint path). Plus --force-spec-decode / --no-spec-decode to override the per-model eligibility profile.
Parsers --enable-auto-tool-choice, --tool-call-parser (auto, hermes, qwen3, qwen3_coder, harmony/gpt_oss, gemma4, deepseek_v31, kimi_k2, minimax, ui_tars, …), --reasoning-parser NAME (NAME: gemma4, qwen3, hy3, deepseek_r1, deepseek_v4, vibethinker, glm4, gpt_oss, harmony, minimax, ui_tars), --no-thinking, --no-tool-call-parser, --no-reasoning-parser
PFlash long prompt --pflash {off,auto,always} (default off for every model; auto / always opt in), --pflash-threshold (default 32k), --pflash-keep-ratio (default 0.20), --pflash-sink-tokens, --pflash-tail-tokens, --pflash-include-tools
Embeddings --embedding-model, --embedding-max-length {auto,<int>} (default auto — derived from config.max_position_embeddings, else tokenizer.model_max_length; an explicit value is clamped to the model maximum), --embedding-overflow-policy {truncate,error} (default truncate, which logs and increments rapid_mlx_embedding_truncations_total; error returns HTTP 400 input_too_long)
MCP + tools --mcp-config <path>, --enable-tool-logits-bias
Escape hatches --force-hybrid / --no-hybrid, --force-openai-harmony-streaming / --no-openai-harmony-streaming, --mllm / --no-mllm, --force-disk-check, --watchdog-ppid PID

Models outside the catalog. For a Hugging Face repo that isn't a catalog alias and isn't cached yet, serve first runs a metadata-only check: format, architecture and memory fit. Before downloading, it refuses a model it can prove won't run on this Mac. --no-preflight skips the check. --request files an opt-in support request without asking when a public model is refused for its architecture or format. It sends the repo id, architecture, format, failure class and Rapid-MLX version. --disk-stream serves MoE experts from disk, so the memory refusal doesn't apply. See Bring your own model.

Disk writes

Beyond the model download, serve can write these to the SSD:

What is writtenWhenHow to turn it off
Prefix-cache snapshot in ~/.cache/rapid-mlx/prefix_cache/ (rewritten whole each time)Each server shutdown, including the end of a chat session, and each --idle-unload-seconds unload--disable-disk-caches, or --disable-prefix-cache (also turns off the in-memory cache)
KV checkpoints in ~/.cache/rapid-mlx/kv_checkpoints/Only with --kv-disk-checkpoint-interval N > 0Off by default; --disable-disk-caches overrules an interval
Vision prefix-cache disk tier in ~/.cache/mlx-vlm/apc/Only when APC_DISK_ENABLED=1 is setOff by default; --disable-disk-caches overrules it
Server log outputContinuously--log-file /dev/null, or a path on a RAM disk

--disable-disk-caches (or RAPID_MLX_DISABLE_DISK_CACHES=1) also stops a saved prefix-cache snapshot from loading at startup; the in-memory prefix cache keeps working. The server logs which explicit settings it overrules. When it does persist the prefix cache, it skips entries that would leave less than 5 GiB free on the volume (RAPID_MLX_PREFIX_CACHE_MIN_FREE_DISK_BYTES; 0 turns the check off).

share

Start rapid-mlx serve and open a public Cloudflare-fronted URL on rapidmlx.com so a friend on a different device can hit the model. Uses the built-in websockets reverse tunnel — no port-forwarding.

$ rapid-mlx share qwen3.5-4b-4bit
# prints URL + share key + one-click chat link, Ctrl-C stops

Flags: --port (default 8765), --thinking / --no-thinking (default off, so chat UIs see content immediately), --cors-origins (defaults to the rapidmlx.com chat-frontend allowlist), --rate-limit RPM (default 120), --chat-frontend URL (default https://rapid-pro.quicksilverpro.io; pass empty to suppress), --disable-disk-caches and --log-file (passed to the server it starts).

launch

One-shot bootstrap: patch a detected IDE / agent client (Cline, Claude Code, Continue, Cursor) so it routes at the local rapid-mlx server. rapid-mlx launch list prints the detection matrix.

$ rapid-mlx launch cursor
$ rapid-mlx launch claude-code --model qwen3.6-35b-8bit
$ rapid-mlx launch --all --start-server --port 8000

Flags: client (or list), --all, --model (default $RAPID_MLX_DEFAULT_MODEL or qwen3.5-4b-4bit), --server-url (default http://127.0.0.1:8000), --port, --start-server, --dry-run, --json (with list). With RAPID_MLX_API_KEY exported, the key is used for the client config and passed to a server started with --start-server. For Claude Code, launch also records the model's context window when the server (or the local cache) reports it.

start

Start a server for an agent and configure the agent in one command: picks the agent's first recommended model that fits this Mac and is already cached (or previews the download and asks), runs the server in the foreground, and writes the agent's config once the server is ready.

$ rapid-mlx start claude-code
$ rapid-mlx start pi --dry-run       # model, port and config changes; starts nothing
$ rapid-mlx start                    # generic OpenAI-compatible endpoint

Flags: profile (agent name; omit for a generic endpoint), --model, --port (default 8000), --host (default 127.0.0.1), --no-download (fail unless a recommended model is cached), --dry-run, --yes / -y, --no-setup (only print the agent's instructions), --ready-timeout (default 600 s), --disable-disk-caches and --log-file (passed to the server it starts; see serve).

connect

Print the running server's connection details (OpenAI and Anthropic base URLs, model), or the setup steps for one tool.

$ rapid-mlx connect
$ rapid-mlx connect claude-code
$ rapid-mlx connect --json

Flags: target (claude-code, continue or openai-python), --json, --host, --port, --model, --base-url (the live server's OpenAI-style URL, e.g. http://localhost:8123/v1).

Chat

chat  alias: run

Interactive REPL against a spawned or already-running server. Defaults to qwen3.5-4b-4bit when the model arg is omitted. The run alias is Ollama-parity. Answers render as proper terminal Markdown (headings, tables, syntax-highlighted code) with a live token counter while streaming; pipes and NO_COLOR keep plain text.

$ rapid-mlx chat
$ rapid-mlx chat qwen3.5-9b-4bit --think
$ rapid-mlx chat --port 8000     # connect to existing server
$ rapid-mlx chat --mcp-config ~/mcp.json  # give the REPL MCP tools
$ rapid-mlx run gemma-4-12b-4bit # Ollama-style alias

Flags: --system, --think / --no-think (default off in the REPL to avoid reasoning-model CoT leak), --mcp-config <path> (connect MCP servers so the REPL can call their tools — parallel across servers, with live tool-activity), --max-tokens (default 2048; 4096 with --think), --temperature (default 0.7), --mcp-max-rounds (tool-call rounds per turn, default 8), --context-length (per-request window for the server chat starts), --disable-prefix-cache (no prefix cache in the server chat starts, in memory or on disk), --disable-disk-caches (keep the in-memory cache, write no optional caches to disk), --log-file (where the spawned server's output goes; without it chat keeps that output in a temporary log file, and - prints it into the session), --port, --base-url, --ready-timeout / --response-timeout (default 600 s each).

Model management

models

List every alias registered in rapid_mlx/aliases.json with per-alias tool/reasoning parser, spec-decode eligibility, suffix-decoding tier, and DFlash / DDTree readiness.

$ rapid-mlx models              # all aliases
$ rapid-mlx models --cached     # same as `rapid-mlx ls`
$ rapid-mlx models --search qwen
$ rapid-mlx models --modality audio

Flags: --cached, --json (machine-readable; pairs with --cached), --search TERM (alias substring), --modality {text,video-gen,image-gen,audio}.

ls

List every model in the local HuggingFace cache — alias (or (unmapped)), HF repo, on-disk size, last modified. Equivalent to rapid-mlx models --cached.

pull

Download a model into the HuggingFace cache — no server.

$ rapid-mlx pull qwen3.5-9b-4bit
$ rapid-mlx pull mlx-community/gemma-4-12b-it-4bit
$ rapid-mlx pull Qwen/Qwen3-0.6B --no-preflight

Flags: model (alias or org/name), --bits N (pull only the <N>bit/ variant of a multi-variant repo), --format name (only the named format variant, e.g. mxfp4; gguf is refused because Rapid-MLX can't run GGUF files), --request, --no-preflight. Like serve, pull checks models outside the catalog before downloading (format and architecture). It reports memory fit but doesn't refuse on it, since you may be pulling for another Mac. The check is skipped with --bits / --format.

import

Convert a bf16/fp16 safetensors model to quantized MLX: explicit and cancel-safe. serve and pull never convert.

$ rapid-mlx import Qwen/Qwen3-0.6B --quantize 4
$ rapid-mlx import ~/models/my-finetune --quantize 4 --name my-ft-4bit
$ rapid-mlx serve my-ft-4bit

Flags: source (Hugging Face org/name or a local model directory), --quantize BITS (2, 3, 4, 6 or 8; default 4), --name NAME (default <source-name>-<bits>bit, with -local appended if that is already an alias), --force (replace an existing import of the same name from another source).

Before downloading or converting, it checks free disk and memory. The conversion and a one-token smoke test run in a temporary directory, and only a passing model is published to ~/.rapid-mlx/imports/<name>. Ctrl-C leaves the cache unchanged, and re-running an identical import is a no-op. Imports show up in rapid-mlx models --cached and are served and removed by name. Full walkthrough: Bring your own model.

rm

Delete a cached model, or an import by its name. Confirms first; -y skips the prompt.

$ rapid-mlx rm qwen3.5-9b-4bit -y

info

Print the per-alias profile — HF repo, tool/reasoning parsers, hybrid flag, MoE flag, spec-decode support, PFlash tier — for one alias or raw HF repo.

$ rapid-mlx info qwen3.6-35b-8bit
$ rapid-mlx info mlx-community/SmolLM3-3B-4bit

ps

List every running rapid-mlx serve process on this machine — PID, port, loaded model, uptime.

recipe

Recommend the smart and the fast model for this Mac's memory. See hardware tiers for every tier.

$ rapid-mlx recipe
$ rapid-mlx recipe --max-ram 32 --json   # plan for another Mac

alias

Give a model your own short name. User aliases live in ~/.config/rapid-mlx/user-aliases.json and point at a catalog alias or a Hugging Face repo id.

$ rapid-mlx alias set my-coder qwen3-coder-30b-4bit
$ rapid-mlx alias list
$ rapid-mlx alias remove my-coder

Benchmark

benchmark run + benchmark share is the local-first Community Benchmark flow and the only one that feeds the leaderboard: a shared run lands in your Mac's row, next to everyone else with the same chip and memory, and on your contributor page. Leaderboard cells the current benchmark has not measured show legacy bench --submit runs, badged with their rapid-mlx version and never mixed into current medians.

bench  freeform + validation tiers · --submit removed

Run a benchmark against a model. Freeform mode by default (any --num-prompts / --max-tokens / batching combo), with validation tiers. --submit is removed: it exits with an error and sends nothing. Use benchmark run + share to put a run on the leaderboard.

$ rapid-mlx bench qwen3.5-4b-4bit
$ rapid-mlx bench gpt-oss-20b-mxfp4-q8 --tier speed
$ rapid-mlx bench gpt-oss-20b-mxfp4-q8 --tier all          # smoke → speed → harness

Flags: freeform batching / KV / prefix-cache flags mirror serve (they configure the throwaway server bench spawns); --submit is removed: it prints the replacement (rapid-mlx benchmark run <model>, then rapid-mlx benchmark share <run-id>) and exits 2, and nothing is run or uploaded. The flags that only shaped a submission (--spec-decode, --run-group, --sampled, --notes, --repo-root) are still listed in --help; --tier {smoke,speed,harness,all} runs a validation ladder (harness drives 5 first-class agent flows: codex/opencode/qwen-code/hermes/aider); --base-url reuses an already-running server for --tier; --long-prompt-tokens + --pflash replicate the long-prompt TTFT profile from PR #649.

benchmark  Community Benchmark · local-first → the leaderboard + your contributor page

Run or inspect reproducible local benchmarks. Every run is saved on your Mac first (~/.rapid-mlx/benchmarks/, or $RAPID_MLX_BENCHMARK_HOME); nothing is uploaded until you run share on a specific run id, and share shows you the exact payload and asks before sending. Six subcommands, each with --json for scripts.

$ rapid-mlx benchmark catalog                     # models with a protocol + fit for this Mac
$ rapid-mlx benchmark plan qwen3.5-9b-4bit        # exact workload; runs nothing
$ rapid-mlx benchmark run qwen3.5-9b-4bit         # run, save locally, print run id
$ rapid-mlx benchmark results --limit 5           # local runs + share receipts
$ rapid-mlx benchmark inspect RUN_ID              # one saved run, in full
$ rapid-mlx benchmark share --preview RUN_ID      # exact payload, nothing sent
$ rapid-mlx benchmark share RUN_ID                # show payload, ask [y/N], upload

catalog [--all] [--memory-gib N] [--json] — lists the recommended models for this Mac (the ★ focus models that fit), then a one-line count of everything else; --all prints every model with a registered protocol (text, image and video workloads). The fit column is computed against this Mac's unified memory, or against --memory-gib when you are planning for another machine.
plan <model> [--json] — prints the model, task, protocol id and version, and every case the run will measure (for text: pp512-tg128 and pp2048-tg512, one warmup round then five measured rounds each). Runs nothing.
run <model> [--json] — runs the registered protocol and saves the result locally. While it runs it reports the model load, an estimated time remaining after the warmup round, and each round's tok/s; at the end it prints Saved local result <run_id>, the measurements, and Nothing was uploaded. (--json keeps stdout to the result document; progress goes to stderr.)
results [--limit N] [--json] — lists locally saved runs, newest first. The JSON form also carries a receipts map, so you can see which runs have been shared.
inspect <run_id> [--json] — prints one saved result in full.
share <run_id> [--preview] [--yes] [--json] — explicitly uploads one saved run. --preview prints the exact payload (plus its digests) and sends nothing. Without --yes the payload is printed and you are asked Upload this result? [y/N]; --yes is for callers that already showed their own consent dialog, and --json requires --yes or --preview.

What share sends — and never sends. Published: chip, RAM, core counts, OS and package versions, the model (its Hugging Face repo id), execution settings, the timings, and a random resettable install id. Never sent: your name, hostname, serial number, hardware UUID, prompts, outputs, or file paths — the payload has no IP-address field. Like any HTTPS service, the endpoint observes the source IP; it uses it transiently for abuse limits and never stores it in the benchmark record. share --preview shows you the exact bytes before you decide.

Receipts and idempotency. An accepted upload returns a receipt (submission_id, accepted_at, run_digest, already_exists, and your pseudonymous contributor name + tag). The CLI checks that the receipt names the run and payload it just sent, then saves it under ~/.rapid-mlx/benchmarks/receipts/ so results can show it. Sharing the same run again is safe: a run id with identical content is accepted again with already_exists: true ("already uploaded") and does not create a second row; the same run id with different content is rejected (409), and a run that was retracted stays gone (410). Each install also has a daily acceptance cap.

Your board name. The pseudonym on your contributor page (/leaderboard/contributors/<name>-<tag>) is derived from a random per-install id stored at ~/.rapid-mlx/bench-install-id, created the first time a share is confirmed. It is not a serial number and identifies nothing but that file. Delete the file and the next share is issued a new id and a new name; runs shared under the old name stay under the old name.

Agent / IDE integration

agents

List, configure, and integration-test agent integrations: claude-code, codex, pi, opencode, hermes, deepseek-harness (dsh), qwen-code, kilo-code, aider, openhands, continue (or continue-dev), plus the langchain, pydanticai and smolagents framework profiles. Per-agent pages: agent guides.

$ rapid-mlx agents                          # list
$ rapid-mlx agents codex                    # setup guide
$ rapid-mlx agents pi --setup --dry-run     # preview, write nothing
$ rapid-mlx agents pi --setup               # auto-configure
$ rapid-mlx agents hermes --test            # run integration flow
$ rapid-mlx agents aider --agent-version 0.90.0 --setup

Flags: agent_name (omit to list), --setup, --dry-run (preview setup without writing), --yes / -y (apply without the confirmation prompt), --no-check (skip the server health and model check, to write a config while no server is running), --test, --model (auto-detected from the running server), --base-url (default http://localhost:8000/v1), --agent-version.

Diagnostics

doctor

One-shot environment check — Apple Silicon detection, macOS + Darwin versions, free disk, HF cache size, Python version, mlx / mlx-lm / transformers / fastapi / uvicorn / rapid-mlx versions, optional-extras status, network probes — and can apply verified repairs.

$ rapid-mlx doctor
$ rapid-mlx doctor --fix --dry-run   # show the repair plan
$ rapid-mlx doctor --fix             # apply verified repairs only

Flags: --verbose / -v (probe detail), --json (versioned machine-readable report), --summary (one line), --deep (dependency, DNS and route probes, up to 30 s), --fix with --dry-run / --yes, --only SECTION / --skip SECTION (repeatable).

telemetry

Manage anonymous usage telemetry, which is on by default after a one-time notice. See the telemetry page for the full event schema.

$ rapid-mlx telemetry status    # enabled/disabled + why
$ rapid-mlx telemetry off       # or: disable
$ rapid-mlx telemetry on        # or: enable
$ rapid-mlx telemetry preview   # a sample payload, exactly as sent
$ rapid-mlx telemetry reset-id  # rotate the client id, keep the choice
$ rapid-mlx telemetry reset     # delete the stored choice + rotate the id

For one run: rapid-mlx --no-telemetry <command>, RAPID_MLX_TELEMETRY=0 or DO_NOT_TRACK=1.

Maintenance

upgrade

Detect this install's method (brew, pip, install.sh) and run the right upgrade command. --dry-run just prints what it would do; -y skips the prompt. rapid-mlx update is the same command.

$ rapid-mlx upgrade
$ rapid-mlx upgrade --dry-run
$ rapid-mlx upgrade -y

version

Print the version. Same as the top-level --version / -V.

help

Print help for a subcommand. rapid-mlx help serve is equivalent to rapid-mlx serve --help.

feedback

Open the community Discord invite in your browser; --no-open only prints the link.

Advanced and experimental

CommandWhat it does
serviceRun Rapid-MLX as a headless macOS LaunchDaemon (install, configure, apply, status, logs, restart, upgrade, uninstall). See Headless macOS service.
system-oneServe a typed decision model (Laya, CLM, or Clef with clef / clef-flash) on POST /v1/systemone and POST /v1/rank. A separate process from serve. CLM is built in; Laya needs the [system-one] extra and Clef the [clef] extra.
cuaExperimental native-accessibility computer-use agent: rapid-mlx cua run --planner local-9b. Needs the [computer-use] extra.

Deprecated aliases

The pre-rename binaries vllm-mlx, vllm-mlx-chat, and vllm-mlx-bench still ship as entry points and forward to the same code — kept so muscle memory and old scripts keep working. New usage should always be rapid-mlx, or its short form rmlx.

Next steps