Rapid-MLX vs oMLX, measured (and when to pick oMLX)
Two Apache-2.0 MLX servers with OpenAI and Anthropic APIs. Same model, two Macs, default settings apart from the few listed. oMLX wins throughput under load and the first agent turn after a restart; Rapid-MLX wins single-stream decode.
Disclosure: we build Rapid-MLX. oMLX is a good project and beats us on several of the tests below; the tables show which and by how much. The Rapid-MLX numbers are Rapid-MLX 0.15.3 (measured at commit 1394f16, which 0.15.3 includes) (the previous release, 0.15.2, is covered on /compare). The raw data, the script and the summary are published.
What each one is
- oMLX is an LLM inference server for Apple Silicon with continuous batching and a two-tier KV cache: a hot tier in memory and a cold tier on SSD that survives restarts. It runs from a macOS menu-bar app with a web admin panel, and it serves OpenAI and Anthropic (
/v1/messages) APIs. Apache-2.0. - Rapid-MLX is an OpenAI- and Anthropic-compatible inference server for Apple Silicon, built on MLX, with continuous batching, a prefix cache, per-model tool-call and reasoning parsers, a curated model catalog, a Desktop app, and
brew install rapid-mlxfrom homebrew/core. Apache-2.0.
Head-to-head
Qwen3.5-9B at 4-bit, each server at its defaults except: oMLX with its documented --paged-ssd-cache-dir, Rapid-MLX with PFlash off (see /compare). Medians of 3 runs (2 for the 4-stream row). oMLX was measured in the earlier run on each machine with the same harness; its version has not changed. On the 18 GB Mac, oMLX refused one of its three ~16k-token requests with its prefill memory guard, so that cell is the median of two.
| Rapid-MLX 0.15.3 (@1394f16), 48 GB | oMLX 0.7.0rc1, 48 GB | Rapid-MLX 0.15.3 (@1394f16), 18 GB | oMLX 0.7.0rc1, 18 GB | |
|---|---|---|---|---|
| Decode, 1 stream | 66.1 tok/s | 51.2 tok/s | 43.6 tok/s | 27.9 tok/s |
| TTFT, ~2k-token prompt (cold) | 5.30 s | 5.40 s | 5.50 s | 5.73 s |
| TTFT, ~16k-token prompt (cold) | 42.8 s | 41.7 s | 44.7 s | 43.9 s |
| 4 streams, aggregate | 68.3 tok/s | 97.1 tok/s | 62.3 tok/s | 85.3 tok/s |
| Tool calls passed (30 prompts) | 30/30 | 30/30 | 30/30 | 30/30 |
| Agent session: turn 1, cold | 65.1 s | 63.1 s | 68.5 s | 67.0 s |
| Agent session: turns 2–10, median | 0.49 s | 0.49 s | 0.69 s | 0.66 s |
| After a restart: turn 1 | 12.3 s | 0.27 s | 13.0 s | 0.56 s |
| After a restart: turns 2–10, median | 0.49 s | 0.27 s | 0.67 s | 0.44 s |
The agent session is a Claude-Code-shaped replay: a 22,819-token first request (a long system prompt plus 30 tool schemas), then 10 turns from a fixed transcript, then the same 10 turns after a server restart. Full method on /compare.
Where oMLX wins
- After a restart. oMLX's SSD tier brought the first turn back in 0.27 s (48 GB) and 0.56 s (18 GB). Rapid-MLX saves its prefix cache to disk at shutdown and reloads it at start, which restored 18,944 of the 22,819 prompt tokens; the first turn still took 12.3 s and 13.0 s. This is oMLX's "5 seconds, not 90" claim, and the data backs it.
- Four concurrent streams: 97.1 tok/s against 68.3 tok/s on 48 GB, 85.3 tok/s against 62.3 tok/s on 18 GB.
- Cold prefill at ~16k tokens, narrowly on both machines: 41.7 s against 42.8 s on 48 GB, 43.9 s against 44.7 s on 18 GB. oMLX also had the shorter cold first agent turn on both machines.
- Follow-up turns on 18 GB, slightly: 0.66 s against 0.69 s (level on 48 GB), and turns 2–10 after a restart on both machines.
Where Rapid-MLX wins
- Single-stream decode: 66.1 tok/s against 51.2 tok/s on 48 GB, 43.6 tok/s against 27.9 tok/s on 18 GB. Rapid-MLX runs this model with MTP speculative decoding on by default.
- Install and packaging:
brew install rapid-mlxfrom homebrew/core, pip, a one-line script, or the Desktop app; per-model aliases pick the tool-call and reasoning parsers for you.
Both passed all 30 tool-call prompts. The previous release, Rapid-MLX 0.15.2, re-prefilled the whole prompt every turn on the 18 GB Mac (70.0 s); that is fixed in 0.15.3.
So which one?
- Pick oMLX if you restart the server or switch models often during long Claude Code, OpenClaw or Cursor sessions, or run several agents at once. Pick it too if you like a menu-bar app with a web admin panel.
- Pick Rapid-MLX if single-stream speed matters most, or you want homebrew/core packaging, a curated catalog with per-model parsers, and a Desktop app.
Looking at other oMLX alternatives? The full comparison also covers Ollama, LM Studio and mlx-lm.
Frequently asked questions
Is Rapid-MLX faster than oMLX?
For single-stream decode, yes in our runs: 66.1 tok/s against 51.2 tok/s on an M4 Pro Mac mini with 48 GB (Qwen3.5-9B at 4-bit), with Rapid-MLX 0.15.3 (measured at commit 1394f16, which 0.15.3 includes). With four concurrent streams, no: oMLX was ahead (97.1 tok/s against 68.3 tok/s). In a long coding-agent session follow-up turns were level on 48 GB (0.49 s and 0.49 s), but after a server restart oMLX answered the first turn in 0.27 s against 12.3 s.
What are the alternatives to oMLX on a Mac?
Rapid-MLX (MLX, OpenAI and Anthropic APIs, Apache-2.0), Ollama (GGUF, cross-platform, MIT), LM Studio (desktop app, GGUF and MLX), and Apple's mlx-lm server. The comparison page has a feature table and measured numbers for all of them except LM Studio.
Do both work with Claude Code?
Yes. Both serve the Anthropic Messages API at /v1/messages, so Claude Code can point at either with ANTHROPIC_BASE_URL. See Claude Code on a local model.