An Ollama alternative for Mac, measured against Ollama

Same Qwen3.5-9B model, two Macs, default settings apart from the few listed. Where an MLX-native server helps, where Ollama is still the better choice, and how to try both side by side.

Disclosure: we build Rapid-MLX. The Rapid-MLX numbers below are Rapid-MLX 0.15.3 (measured at commit 1394f16, which 0.15.3 includes); the raw data and the script are published. For the longer August study across three models, including llama.cpp and LM Studio, see Rapid-MLX vs Ollama vs LM Studio, measured. To actually switch, the migration guide maps every Ollama command.

Rapid-MLX vs Ollama, measured

Rapid-MLX loads the MLX 4-bit build of Qwen3.5-9B; Ollama loads its own qwen3.5:9b (GGUF Q4_K_M), the tag ollama pull qwen3.5:9b gives you. The weights are the same model at a similar bit width, not the same file. Ollama ran with a 32k context so no prompt was truncated, and with OLLAMA_NUM_PARALLEL=4, which Ollama 0.34.3 ignores for this model (its log: "model architecture does not currently support parallel requests"), so its four streams ran one after another. Rapid-MLX ran with PFlash off (see /compare). Ollama was measured in the earlier run on each machine with the same harness; its version has not changed. Medians of 3 runs (2 for the 4-stream row).

Rapid-MLX 0.15.3 (@1394f16), 48 GB Ollama 0.34.3, 48 GB Rapid-MLX 0.15.3 (@1394f16), 18 GB Ollama 0.34.3, 18 GB
Decode, 1 stream 66.1 tok/s 38.4 tok/s 43.6 tok/s 22.8 tok/s
TTFT, ~2k-token prompt (cold) 5.30 s 6.35 s 5.50 s 6.24 s
TTFT, ~16k-token prompt (cold) 42.8 s 49.4 s 44.7 s 49.9 s
4 streams, aggregate 68.3 tok/s 37.4 tok/s 62.3 tok/s 22.5 tok/s
Tool calls passed (30 prompts) 30/30 30/30 30/30 30/30
Agent session: turns 2–10, median 0.49 s 0.41 s 0.69 s 0.44 s
After a restart (Ollama: model reload): turn 1 12.3 s 73.1 s 13.0 s 76.9 s

How to read it: Rapid-MLX decoded faster on one stream on both Macs. The 4-stream gap is mostly Ollama serving this model one request at a time, not a like-for-like batching comparison. Cold prefill was close, with Ollama 10–15% slower. Both passed all 30 tool-call prompts.

Where Ollama won: follow-up turns in the coding-agent replay (0.41 s against 0.49 s on 48 GB, 0.44 s against 0.69 s on 18 GB, where the gap is larger). The previous release, Rapid-MLX 0.15.2, lost that test badly on the 18 GB Mac (70.0 s, re-prefilling the whole prompt every turn); that is fixed in 0.15.3. After a restart, Rapid-MLX reloaded most of the prefix from its disk save and Ollama prefilled from scratch. The oMLX page covers the session in more detail.

When to stay on Ollama

When Rapid-MLX is the better fit

You do not have to choose on day one: the two servers listen on different ports (Ollama on 11434, Rapid-MLX on 8000), so they can run side by side while you compare them on your own prompts.

Other options, including oMLX, LM Studio and mlx-lm, are on the full comparison page.

Frequently asked questions

Is Rapid-MLX faster than Ollama?

For decode, yes in our runs: on an M4 Pro Mac mini with 48 GB (Qwen3.5-9B at 4-bit), Rapid-MLX 0.15.3 (measured at commit 1394f16, which 0.15.3 includes) decoded 66.1 tok/s on one stream against Ollama's 38.4 tok/s, and 68.3 tok/s against 37.4 tok/s across four streams (Ollama serves this model one request at a time). In a long coding-agent session Ollama answered follow-up turns faster: 0.41 s against 0.49 s on 48 GB, and 0.44 s against 0.69 s on an 18 GB Mac. The previous release, 0.15.2, was far slower there (70.0 s); that is fixed in 0.15.3.

What is a good Ollama alternative for a Mac?

On Apple Silicon, the MLX-native servers: Rapid-MLX, oMLX and Apple's mlx-lm server. LM Studio is the pick if you want a desktop app first. The comparison page puts them side by side with measured numbers.

Can I keep using my Ollama tools?

Anything that speaks the OpenAI API only needs a new base URL (http://localhost:8000/v1). The migration guide has the command-by-command mapping.