An Ollama alternative for Mac, measured against Ollama
Same Qwen3.5-9B model, two Macs, default settings apart from the few listed. Where an MLX-native server helps, where Ollama is still the better choice, and how to try both side by side.
Disclosure: we build Rapid-MLX. The Rapid-MLX numbers below are Rapid-MLX 0.15.3 (measured at commit 1394f16, which 0.15.3 includes); the raw data and the script are published. For the longer August study across three models, including llama.cpp and LM Studio, see Rapid-MLX vs Ollama vs LM Studio, measured. To actually switch, the migration guide maps every Ollama command.
Rapid-MLX vs Ollama, measured
Rapid-MLX loads the MLX 4-bit build of Qwen3.5-9B; Ollama loads its own qwen3.5:9b (GGUF Q4_K_M), the tag ollama pull qwen3.5:9b gives you. The weights are the same model at a similar bit width, not the same file. Ollama ran with a 32k context so no prompt was truncated, and with OLLAMA_NUM_PARALLEL=4, which Ollama 0.34.3 ignores for this model (its log: "model architecture does not currently support parallel requests"), so its four streams ran one after another. Rapid-MLX ran with PFlash off (see /compare). Ollama was measured in the earlier run on each machine with the same harness; its version has not changed. Medians of 3 runs (2 for the 4-stream row).
| Rapid-MLX 0.15.3 (@1394f16), 48 GB | Ollama 0.34.3, 48 GB | Rapid-MLX 0.15.3 (@1394f16), 18 GB | Ollama 0.34.3, 18 GB | |
|---|---|---|---|---|
| Decode, 1 stream | 66.1 tok/s | 38.4 tok/s | 43.6 tok/s | 22.8 tok/s |
| TTFT, ~2k-token prompt (cold) | 5.30 s | 6.35 s | 5.50 s | 6.24 s |
| TTFT, ~16k-token prompt (cold) | 42.8 s | 49.4 s | 44.7 s | 49.9 s |
| 4 streams, aggregate | 68.3 tok/s | 37.4 tok/s | 62.3 tok/s | 22.5 tok/s |
| Tool calls passed (30 prompts) | 30/30 | 30/30 | 30/30 | 30/30 |
| Agent session: turns 2–10, median | 0.49 s | 0.41 s | 0.69 s | 0.44 s |
| After a restart (Ollama: model reload): turn 1 | 12.3 s | 73.1 s | 13.0 s | 76.9 s |
How to read it: Rapid-MLX decoded faster on one stream on both Macs. The 4-stream gap is mostly Ollama serving this model one request at a time, not a like-for-like batching comparison. Cold prefill was close, with Ollama 10–15% slower. Both passed all 30 tool-call prompts.
Where Ollama won: follow-up turns in the coding-agent replay (0.41 s against 0.49 s on 48 GB, 0.44 s against 0.69 s on 18 GB, where the gap is larger). The previous release, Rapid-MLX 0.15.2, lost that test badly on the 18 GB Mac (70.0 s, re-prefilling the whole prompt every turn); that is fixed in 0.15.3. After a restart, Rapid-MLX reloaded most of the prefix from its disk save and Ollama prefilled from scratch. The oMLX page covers the session in more detail.
When to stay on Ollama
- You need Windows or Linux, or one tool across several machines.
- You depend on a model from the Ollama library that is not packaged for MLX.
- You are happy with its speed on your Mac. Ollama also ships an MLX engine in preview for Macs with 32 GB or more (announcement); we did not run it, so it is not in the table.
When Rapid-MLX is the better fit
- You are on an Apple Silicon Mac and want an MLX-native server with both OpenAI and Anthropic (
/v1/messages) APIs, per-model tool-call and reasoning parsers, and Apache-2.0 licensing. - You want
brew install rapid-mlxfrom homebrew/core, or the Desktop app.
You do not have to choose on day one: the two servers listen on different ports (Ollama on 11434, Rapid-MLX on 8000), so they can run side by side while you compare them on your own prompts.
Other options, including oMLX, LM Studio and mlx-lm, are on the full comparison page.
Frequently asked questions
Is Rapid-MLX faster than Ollama?
For decode, yes in our runs: on an M4 Pro Mac mini with 48 GB (Qwen3.5-9B at 4-bit), Rapid-MLX 0.15.3 (measured at commit 1394f16, which 0.15.3 includes) decoded 66.1 tok/s on one stream against Ollama's 38.4 tok/s, and 68.3 tok/s against 37.4 tok/s across four streams (Ollama serves this model one request at a time). In a long coding-agent session Ollama answered follow-up turns faster: 0.41 s against 0.49 s on 48 GB, and 0.44 s against 0.69 s on an 18 GB Mac. The previous release, 0.15.2, was far slower there (70.0 s); that is fixed in 0.15.3.
What is a good Ollama alternative for a Mac?
On Apple Silicon, the MLX-native servers: Rapid-MLX, oMLX and Apple's mlx-lm server. LM Studio is the pick if you want a desktop app first. The comparison page puts them side by side with measured numbers.
Can I keep using my Ollama tools?
Anything that speaks the OpenAI API only needs a new base URL (http://localhost:8000/v1). The migration guide has the command-by-command mapping.