Rapid-MLX vs oMLX, measured (and when to pick oMLX)

Two Apache-2.0 MLX servers with OpenAI and Anthropic APIs. Same model, two Macs, default settings apart from the few listed. oMLX wins throughput under load and the first agent turn after a restart; Rapid-MLX wins single-stream decode.

Disclosure: we build Rapid-MLX. oMLX is a good project and beats us on several of the tests below; the tables show which and by how much. The Rapid-MLX numbers are Rapid-MLX 0.15.3 (measured at commit 1394f16, which 0.15.3 includes) (the previous release, 0.15.2, is covered on /compare). The raw data, the script and the summary are published.

What each one is

Head-to-head

Qwen3.5-9B at 4-bit, each server at its defaults except: oMLX with its documented --paged-ssd-cache-dir, Rapid-MLX with PFlash off (see /compare). Medians of 3 runs (2 for the 4-stream row). oMLX was measured in the earlier run on each machine with the same harness; its version has not changed. On the 18 GB Mac, oMLX refused one of its three ~16k-token requests with its prefill memory guard, so that cell is the median of two.

Rapid-MLX 0.15.3 (@1394f16), 48 GB oMLX 0.7.0rc1, 48 GB Rapid-MLX 0.15.3 (@1394f16), 18 GB oMLX 0.7.0rc1, 18 GB
Decode, 1 stream 66.1 tok/s 51.2 tok/s 43.6 tok/s 27.9 tok/s
TTFT, ~2k-token prompt (cold) 5.30 s 5.40 s 5.50 s 5.73 s
TTFT, ~16k-token prompt (cold) 42.8 s 41.7 s 44.7 s 43.9 s
4 streams, aggregate 68.3 tok/s 97.1 tok/s 62.3 tok/s 85.3 tok/s
Tool calls passed (30 prompts) 30/30 30/30 30/30 30/30
Agent session: turn 1, cold 65.1 s 63.1 s 68.5 s 67.0 s
Agent session: turns 2–10, median 0.49 s 0.49 s 0.69 s 0.66 s
After a restart: turn 1 12.3 s 0.27 s 13.0 s 0.56 s
After a restart: turns 2–10, median 0.49 s 0.27 s 0.67 s 0.44 s

The agent session is a Claude-Code-shaped replay: a 22,819-token first request (a long system prompt plus 30 tool schemas), then 10 turns from a fixed transcript, then the same 10 turns after a server restart. Full method on /compare.

Where oMLX wins

Where Rapid-MLX wins

Both passed all 30 tool-call prompts. The previous release, Rapid-MLX 0.15.2, re-prefilled the whole prompt every turn on the 18 GB Mac (70.0 s); that is fixed in 0.15.3.

So which one?

Looking at other oMLX alternatives? The full comparison also covers Ollama, LM Studio and mlx-lm.

Frequently asked questions

Is Rapid-MLX faster than oMLX?

For single-stream decode, yes in our runs: 66.1 tok/s against 51.2 tok/s on an M4 Pro Mac mini with 48 GB (Qwen3.5-9B at 4-bit), with Rapid-MLX 0.15.3 (measured at commit 1394f16, which 0.15.3 includes). With four concurrent streams, no: oMLX was ahead (97.1 tok/s against 68.3 tok/s). In a long coding-agent session follow-up turns were level on 48 GB (0.49 s and 0.49 s), but after a server restart oMLX answered the first turn in 0.27 s against 12.3 s.

What are the alternatives to oMLX on a Mac?

Rapid-MLX (MLX, OpenAI and Anthropic APIs, Apache-2.0), Ollama (GGUF, cross-platform, MIT), LM Studio (desktop app, GGUF and MLX), and Apple's mlx-lm server. The comparison page has a feature table and measured numbers for all of them except LM Studio.

Do both work with Claude Code?

Yes. Both serve the Anthropic Messages API at /v1/messages, so Claude Code can point at either with ANTHROPIC_BASE_URL. See Claude Code on a local model.