Rapid-MLX vs oMLX vs Ollama vs LM Studio vs mlx-lm, features and measurements
Five ways to serve a local model on a Mac. A feature table, one reproducible measurement on the same model, the raw data, and an honest "choose X if" for each.
Disclosure: we build Rapid-MLX. Expect a page like this to flatter us. That is why every number below comes from scripted runs with the script, the driver, the raw per-request JSON and the derived summary published next to it. Every number on this page is generated from those files. Where another engine wins, the table says so.
The short version
On an M4 Pro Mac mini with 48 GB (Qwen3.5-9B at 4-bit), with each engine at its defaults except the settings listed under the table. The Rapid-MLX rows are Rapid-MLX 0.15.3 (measured at commit 1394f16, which 0.15.3 includes); the previous release, 0.15.2, is compared in What changed in Rapid-MLX.
- Single-stream decode: Rapid-MLX was fastest at 66.1 tok/s (oMLX 51.2 tok/s, mlx-lm 47.7 tok/s, Ollama 38.4 tok/s). Rapid-MLX runs this model with MTP speculative decoding on by default.
- Four concurrent streams: Rapid-MLX loses. oMLX (97.1 tok/s) and mlx-lm (96.6 tok/s) led; Rapid-MLX was third (68.3 tok/s). Ollama (37.4 tok/s) serves this model one request at a time.
- Long coding-agent session, follow-up turns: every engine reused the ~23k-token prefix and answered in under a second; Ollama was fastest: Ollama 0.41 s, mlx-lm 0.47 s, oMLX 0.49 s, Rapid-MLX 0.49 s (medians of turns 2–10).
- Same session after a server restart: oMLX wins. Its SSD tier answered the first turn in 0.27 s. Rapid-MLX restored most of the prefix from its disk save and took 12.3 s; mlx-lm (65.7 s) and Ollama (73.1 s) prefilled from scratch.
- Cold prefill was close for everyone: 41.7 s (oMLX, fastest) to 49.4 s at ~16k tokens.
- Tool calls: every engine passed all 30 prompts, so this set does not separate them.
- On an 18 GB Mac the order is the same, except that mlx-lm's server ran out of GPU memory in the agent session; the 18 GB table has all numbers.
Feature comparison
What each project ships today, from its own docs and our installs. "Yes" means documented and working in the version we ran, not a roadmap.
| Rapid-MLX | oMLX | Ollama | LM Studio | mlx-lm (mlx_lm.server) |
|
|---|---|---|---|---|---|
| What it is | OpenAI- and Anthropic-compatible inference server for Apple Silicon, built on MLX | Inference server for Apple Silicon, built on MLX, run from a macOS menu-bar app | Cross-platform model runner and server | Desktop app for downloading and chatting with models, with a local server | Apple's reference MLX library; its server is a minimal example |
| OpenAI API | Yes | Yes | Yes (plus its own /api) |
Yes (plus its own REST API) | Yes, basic |
Anthropic /v1/messages |
Yes | Yes | Yes (since v0.14) | Yes (since 0.4.1) | No |
| GUI | Rapid-MLX Desktop (macOS app); CLI | Menu-bar app plus a web admin panel | Desktop app with chat; CLI | Full desktop app; lms CLI |
None (Python CLI) |
| License | Apache-2.0 (engine and Desktop) | Apache-2.0 | MIT | Free to use, closed-source app | MIT |
| Platforms | macOS on Apple Silicon | macOS on Apple Silicon | macOS, Windows, Linux | macOS, Windows, Linux | macOS on Apple Silicon (MLX also has Linux backends) |
| Model format / backend | MLX | MLX (mlx-lm) | GGUF (llama.cpp); MLX engine in preview on Macs with 32 GB+ | GGUF (llama.cpp) and MLX | MLX |
| Concurrent requests | Continuous batching | Continuous batching | Parallel slots (OLLAMA_NUM_PARALLEL), not for every architecture |
Parallel predictions (--parallel) |
Batching (off with a quantized KV cache) |
| KV / prompt cache | In-memory prefix cache; written to disk at shutdown if it fits a short time budget, reloaded at start | In-memory hot tier plus an SSD cold tier that survives restarts | In-memory, per loaded model | In-memory | In-memory prompt cache |
| Tool-call parsing | Dedicated parser per model family, picked by the model alias | Auto-detected per model family (JSON <tool_call>, Mistral, GLM, MiniMax, Kimi and others) |
Per-model templates | Per-model templates | From the tokenizer's chat template |
| Model catalog | Curated aliases with per-model profiles, or any MLX repo on Hugging Face | Any mlx-lm model; downloader in the admin panel | ollama.com library | Hugging Face search in the app | Any MLX repo on Hugging Face |
| Install | brew install rapid-mlx, pip, one-line script, or the Desktop DMG |
DMG, Homebrew tap, or pip from a release wheel | DMG, Homebrew, or script | DMG | pip install mlx-lm |
Sources: each project's README and docs, checked 2026-09-26 (oMLX, Ollama Anthropic compatibility, Ollama MLX preview, LM Studio Anthropic endpoint, mlx-lm server). Rapid-MLX flags are from rapid-mlx serve --help at 0.15.2.
Measured on an M4 Pro Mac mini (48 GB)
| Engine (config) | Decode, 1 stream | TTFT, ~2k-token prompt | TTFT, ~16k-token prompt | 4 streams, aggregate | Tool calls (30 prompts) |
|---|---|---|---|---|---|
| Rapid-MLX 0.15.3 (@1394f16) (PFlash off) | 66.1 tok/s | 5.30 s | 42.8 s | 68.3 tok/s | 30/30 |
| Rapid-MLX 0.15.3 (@1394f16) (defaults, PFlash on) | not run | 5.31 s | 8.23 s | not run | not run |
| oMLX 0.7.0rc1 (SSD cache on) | 51.2 tok/s | 5.40 s | 41.7 s | 97.1 tok/s | 30/30 |
| Ollama 0.34.3 (qwen3.5:9b, Q4_K_M) | 38.4 tok/s | 6.35 s | 49.4 s | 37.4 tok/s | 30/30 |
| mlx-lm 0.31.3 (mlx_lm.server) | 47.7 tok/s | 5.51 s | 43.6 s | 96.6 tok/s | 30/30 |
Conditions. Mac mini, Apple M4 Pro, 48 GB unified memory, macOS 26.5.1, on AC power, idle before each engine (GPU 0% after a 60 s rest + 10 idle seconds). Model: Qwen3.5-9B at 4-bit (MLX engines: mlx-community/Qwen3.5-9B-4bit @ 8b2b98c, MTP sidecar mlx-community/Qwen3.5-9B-MTP-4bit @ 222dfd2; Ollama: qwen3.5:9b, GGUF Q4_K_M, id 6488c96fa5fa), same as run 2. One engine at a time, loopback only, temperature 0, streaming. Before each engine the driver rested 60 s and waited for the GPU to read idle, and gave the engine a fresh home directory (no cache or settings carried over). The competitors were measured in run 3; Rapid-MLX 0.15.3 (measured at commit 1394f16, which 0.15.3 includes) was measured a few hours later the same night with the same harness and protocol in run 4 (the competitors' versions had not changed, so they were not re-run). Decode and TTFT cells are the median of 3 runs, the 4-stream cell the median of 2. Decode tok/s is (completion tokens − 1) / (last token − first token), using each engine's own token count. TTFT is request sent to the first streamed token of any kind, thinking included; each prompt starts with a random nonce so that no prefix cache applies. Engines group tokens into stream chunks differently (oMLX sends several tokens per chunk), which can add a fraction of a second to a TTFT. Every request asked for thinking explicitly (chat_template_kwargs.enable_thinking=true), because Rapid-MLX turns thinking off by default for casual requests while the others leave it on; the raw files store the generated text and reasoning, so you can check. Engine settings: Rapid-MLX with --pflash off (see below), everything else default, which for this model includes MTP speculative decoding; oMLX with its documented --paged-ssd-cache-dir; Ollama with a 32k context so no prompt was truncated and OLLAMA_NUM_PARALLEL=4 (Ollama 0.34.3 ignores it for this model and logs "model architecture does not currently support parallel requests"; its four streams' first tokens arrived about 7 s apart, one after another); mlx-lm's server at its defaults. Ollama loads its own GGUF Q4_K_M build of the model, the others load the same MLX 4-bit files.
Why the main Rapid-MLX row has PFlash off. For Qwen3.5 and Qwen3.6 models Rapid-MLX turns on PFlash by default: long-prompt prefill compression that keeps a fraction of the prompt's tokens (20% by default) before prefilling. It makes long prompts much faster (8.23 s instead of 42.8 s at ~16k tokens; it did not engage on the agent session, whose requests carry tools), but the model then does not see every prompt token, so it is not a like-for-like comparison with engines that prefill all of them. We did not measure answer quality with PFlash on, and ran only the prefill and agent-session tests in that configuration.
On an 18 GB Mac
MacBook Pro, Apple M3 Pro, 18 GB unified memory, macOS 15.6.1, on power. Competitors from run 2, Rapid-MLX 0.15.3 (measured at commit 1394f16, which 0.15.3 includes) from run 4, same protocol. On this machine oMLX refused one of its three ~16k-token requests with its prefill memory guard ("Prefill would require ~9.08 GB peak … dynamic ceiling is 9.00 GB"), so its ~16k cell is the median of the two that ran, and Ollama's four streams started about 11 s apart.
| Engine (config) | Decode, 1 stream | TTFT, ~2k-token prompt | TTFT, ~16k-token prompt | 4 streams, aggregate | Tool calls (30 prompts) |
|---|---|---|---|---|---|
| Rapid-MLX 0.15.3 (@1394f16) (PFlash off) | 43.6 tok/s | 5.50 s | 44.7 s | 62.3 tok/s | 30/30 |
| Rapid-MLX 0.15.3 (@1394f16) (defaults, PFlash on) | not run | 5.50 s | 8.58 s | not run | not run |
| oMLX 0.7.0rc1 (SSD cache on) | 27.9 tok/s | 5.73 s | 43.9 s | 85.3 tok/s | 30/30 |
| Ollama 0.34.3 (qwen3.5:9b, Q4_K_M) | 22.8 tok/s | 6.24 s | 49.9 s | 22.5 tok/s | 30/30 |
| mlx-lm 0.31.3 (mlx_lm.server) | 25.6 tok/s | 5.87 s | 45.6 s | 80.7 tok/s | 30/30 |
The coding-agent session (oMLX's headline case)
oMLX's pitch is that Claude Code "responds in 5 seconds, not 90" (omlx.ai): coding agents resend a long system prompt and tool list on every turn, and oMLX keeps that prefix's KV cache on SSD so it survives a restart. We replayed a Claude-Code-shaped session to measure exactly that: a 22,819-token first request (a long system prompt plus 30 tool schemas), then 10 user turns from a fixed transcript (the assistant replies are canned, so every engine sees identical prompts). Then we restarted the server, warmed it with an unrelated prompt, and replayed the same 10 turns. For Ollama, "restart" means unloading the model (ollama stop), which drops its KV cache the same way. Numbers are time to first token per turn; when a server sends nothing until its 16-token reply is complete, the number is time to the end of the stream instead, and the raw records flag those turns (ttft_fallback_end_of_stream).
M4 Pro, 48 GB:
| Engine (config) | Turn 1 (cold) | Turns 2–10, median | 10 turns, total | After restart: turn 1 | After restart: turns 2–10, median |
|---|---|---|---|---|---|
| Rapid-MLX 0.15.3 (@1394f16) (PFlash off) | 65.1 s | 0.49 s | 69.6 s | 12.3 s | 0.49 s |
| Rapid-MLX 0.15.3 (@1394f16) (defaults, PFlash on) | 65.3 s | 0.49 s | 69.8 s | 12.3 s | 0.49 s |
| oMLX 0.7.0rc1 (SSD cache on) | 63.1 s | 0.49 s | 67.6 s | 0.27 s | 0.27 s |
| Ollama 0.34.3 (qwen3.5:9b, Q4_K_M) | 73.4 s | 0.41 s | 76.9 s | 73.1 s | 0.41 s |
| mlx-lm 0.31.3 (mlx_lm.server) | 65.7 s | 0.47 s | 69.7 s | 65.7 s | 0.47 s |
M3 Pro, 18 GB:
| Engine (config) | Turn 1 (cold) | Turns 2–10, median | 10 turns, total | After restart: turn 1 | After restart: turns 2–10, median |
|---|---|---|---|---|---|
| Rapid-MLX 0.15.3 (@1394f16) (PFlash off) | 68.5 s | 0.69 s | 74.9 s | 13.0 s | 0.67 s |
| Rapid-MLX 0.15.3 (@1394f16) (defaults, PFlash on) | 68.1 s | 0.69 s | 74.5 s | 13.0 s | 0.65 s |
| oMLX 0.7.0rc1 (SSD cache on) | 67.0 s | 0.66 s | 73.0 s | 0.56 s | 0.44 s |
| Ollama 0.34.3 (qwen3.5:9b, Q4_K_M) | 77.6 s | 0.44 s | 81.4 s | 76.9 s | 0.44 s |
| mlx-lm 0.31.3 (mlx_lm.server) | 71.4 s | 0.85 s, 0.77 s; server failed on turn 4 | n/a | 68.9 s | 0.86 s; server failed on turn 3 |
oMLX's claim holds: after one cold prefill it answered every turn in under a second, and its SSD tier made the first turn after a restart nearly free (0.27 s on 48 GB, 0.56 s on 18 GB). Ollama and mlx-lm reused their caches within the session just as well, and lost them on restart. On 18 GB, mlx-lm's server died with a Metal out-of-memory error on turn 4, and after the restart on turn 3 (server log).
Rapid-MLX 0.15.3 (measured at commit 1394f16, which 0.15.3 includes) kept the session prefix resident on both machines: turns 2–10 prefilled only the 42–66 new tokens each, bringing it under a second like the others (0.49 s on 48 GB, 0.69 s on 18 GB), though still behind Ollama on both machines and slightly behind oMLX on 18 GB. At shutdown it saved the session to disk and reloaded it at start, so the first turn after the restart restored 18,944 of the 22,819 prompt tokens and prefilled the remaining 3,875: 12.3 s on 48 GB. That is far better than a full prefill but still well behind oMLX's 0.27 s. The save is not guaranteed: on 18 GB, one of the four shutdowns (the last one, after the restart replay, so no number above depends on it) skipped the entry because its predicted write time (3.1 s) did not fit the 3.5 s shutdown budget minus its 0.4 s commit headroom. Cache budgets and save results for every start are in each run's versions.json and server logs. The prompt token counts differ slightly between engines (about 150 tokens) because each renders the tool schemas through its own chat template.
What changed in Rapid-MLX
The first published version of this page measured the previous release, Rapid-MLX 0.15.2, and it lost the agent session badly on the 18 GB Mac: its prefix cache sized itself to 0.86 GB there, each ~0.96 GB session entry was "too large" to store, and every turn paid the full prefill; what it saved at shutdown did not include the session, so the first turn after a restart was a full prefill too. Commit 1394f16 (#3794) gives agent sessions a minimum cache budget on small-RAM Macs and makes the shutdown save fit in most cases (see above). Same harness, same protocol, PFlash off:
| Rapid-MLX, agent session | 0.15.2, 18 GB | 0.15.3 (@1394f16), 18 GB | 0.15.2, 48 GB | 0.15.3 (@1394f16), 48 GB |
|---|---|---|---|---|
| Turns 2–10, median | 70.0 s | 0.69 s | 0.50 s | 0.49 s |
| Turn 1 after a restart | 69.1 s | 13.0 s | 65.1 s | 12.3 s |
Install 0.15.3 with pip install rapid-mlx==0.15.3, brew upgrade rapid-mlx, or Rapid-MLX Desktop 0.15.3. The benchmark was measured at commit 1394f16, which 0.15.3 includes; it was not re-run on the release tag. The 0.15.2 numbers are in summary-run2.json and summary-run3-m4pro-48gb.json.
Tool calling
Thirty fixed prompts, each with the same eight tools (weather, calculator, web search, email, calendar, read file, shell, currency). A prompt passes when the first tool call names the right function and its JSON arguments contain the expected values. Non-streaming, temperature 0. The prompts, expected calls and every raw response are in the raw data.
Rapid-MLX 0.15.3 (@1394f16) (PFlash off): 30/30 · oMLX 0.7.0rc1 (SSD cache on): 30/30 · Ollama 0.34.3 (qwen3.5:9b, Q4_K_M): 30/30 · mlx-lm 0.31.3 (mlx_lm.server): 30/30. Every engine passed every prompt. These are single-call, clearly worded requests, so the result says the plumbing works on all four servers for this model; it does not rank them. Harder multi-step tool use is where parsers differ, and this test does not cover it.
What we did not measure, and why
- LM Studio. Its headless install puts a background service under the user's home directory, which our clean-scratch-directory rule for the benchmark machines excluded, and on the one Mac where it was already installed (our Mac Studio M3 Ultra) the GPU was busy with an unrelated fine-tuning job for the whole window. LM Studio is in the feature table only.
- M3 Ultra numbers, for the same reason.
- Ollama's MLX engine (
qwen3.5:9b-mlx), which needs 32 GB or more and was not in our run-3 plan; it is a gap in the 48 GB table. - Other models. One model, one quantization. Treat this as two data points, and run the script on your own Mac and models.
When to choose which
- Choose oMLX if you restart the server or switch models often during long coding-agent sessions: its SSD cache made the first turn after a restart nearly free, and it led 4-stream throughput. Also if you want a menu-bar app with a web admin panel.
- Choose Ollama if you need Windows or Linux, a single tool across machines, or a model from its library that is not packaged for MLX.
- Choose LM Studio if you want the most polished desktop app for browsing, downloading and chatting with models, on any OS.
- Choose mlx-lm if you are writing Python against MLX directly or need a minimal reference server.
- Choose Rapid-MLX if you want an Apache-2.0 MLX server with both OpenAI and Anthropic APIs, the fastest single-stream decode in our tests, a curated catalog with per-model tool-call and reasoning parsers, a Desktop app, and
brew installfrom homebrew/core. For long agent sessions on a small-RAM Mac, use 0.15.3, which includes commit 1394f16; 0.15.2 re-prefills every turn there.
Focused pages: Rapid-MLX vs oMLX · Ollama alternative for Mac · our August Ollama/LM Studio/llama.cpp benchmark.
How these runs came about
We publish every run. Run 1 (18 GB) was superseded after a review found three flaws that favored Rapid-MLX or muddied the result: Rapid-MLX had thinking off by default while the other engines were thinking; the engines ran back to back in a fixed order with Rapid-MLX first, on the coolest machine; and one Rapid-MLX start overlapped the shutdown of the previous server. The same review caught two errors on the first version of this page: it said Ollama's four streams ran in parallel (they did not), and the server logs it cited were missing from the published data. Run 2 (18 GB) fixed all of that; run 3 repeated it on a 48 GB M4 Pro; run 4 re-measured only Rapid-MLX at commit 1394f16, which 0.15.3 includes, on both machines. The data README maps each table to its run.
Reproduce it
curl -O https://rapidmlx.com/blog/assets/compare-2026-09/bench_compare.py # start any OpenAI-compatible server on loopback, then, per scenario a b c d e: python3 bench_compare.py --engine NAME --base-url http://127.0.0.1:PORT/v1 --model MODEL --thinking on --scenario a --out a.json
The client is standard-library Python. run_all.sh is the driver we used: it starts and stops each server, rests between engines, and restarts the server for the agent session. It expects the engines installed in a scratch directory (W, see its header) and your normal Hugging Face cache in HF_HOME.
Frequently asked questions
Is Rapid-MLX faster than Ollama?
For decode, yes in our runs: on an M4 Pro Mac mini with 48 GB (Qwen3.5-9B at 4-bit), Rapid-MLX 0.15.3 (measured at commit 1394f16, which 0.15.3 includes) decoded 66.1 tok/s on one stream against Ollama's 38.4 tok/s, and 68.3 tok/s against 37.4 tok/s across four streams (Ollama serves this model one request at a time). In a long coding-agent session Ollama answered follow-up turns faster: a median of 0.41 s against 0.49 s for Rapid-MLX, and 0.44 s against 0.69 s on an 18 GB Mac. The previous release, 0.15.2, was far slower there (70.0 s); that is fixed in 0.15.3.
Rapid-MLX vs oMLX: which should I use?
Pick by workload. On an M4 Pro Mac mini with 48 GB, oMLX led four concurrent streams (97.1 tok/s against 68.3 tok/s) and the first agent turn after a restart (0.27 s against 12.3 s, thanks to its SSD cache). Rapid-MLX 0.15.3 (measured at commit 1394f16, which 0.15.3 includes) led single-stream decode (66.1 tok/s against 51.2 tok/s). Follow-up agent turns were level on 48 GB (0.49 s and 0.49 s) with oMLX slightly ahead on 18 GB (0.66 s against 0.69 s), and both passed all 30 tool-call prompts. Details on the Rapid-MLX vs oMLX page.
Does Rapid-MLX work with Claude Code?
Yes. Rapid-MLX serves the Anthropic Messages API at /v1/messages, so Claude Code connects with ANTHROPIC_BASE_URL (see Claude Code on a local model). oMLX, Ollama (v0.14+) and LM Studio (0.4.1+) also serve that API. In our replay of a 23k-token Claude Code session, Rapid-MLX 0.15.3 (measured at commit 1394f16, which 0.15.3 includes) answered follow-up turns in a median of 0.49 s on a 48 GB Mac and 0.69 s on an 18 GB Mac. The previous release, 0.15.2, re-prefilled the whole prompt every turn on the 18 GB Mac (70.0 s); that is fixed in 0.15.3.
Where is the raw data?
Every request's timings, the engine-reported token usage and every tool-call response are listed in raw/manifest.json, with the script (bench_compare.py), the driver (run_all.sh), the engine versions and flags (each run's versions.json) and the numbers used on this page (summary.json).