Rapid-MLX vs oMLX vs Ollama vs LM Studio vs mlx-lm, features and measurements

Five ways to serve a local model on a Mac. A feature table, one reproducible measurement on the same model, the raw data, and an honest "choose X if" for each.

Disclosure: we build Rapid-MLX. Expect a page like this to flatter us. That is why every number below comes from scripted runs with the script, the driver, the raw per-request JSON and the derived summary published next to it. Every number on this page is generated from those files. Where another engine wins, the table says so.

The short version

On an M4 Pro Mac mini with 48 GB (Qwen3.5-9B at 4-bit), with each engine at its defaults except the settings listed under the table. The Rapid-MLX rows are Rapid-MLX 0.15.3 (measured at commit 1394f16, which 0.15.3 includes); the previous release, 0.15.2, is compared in What changed in Rapid-MLX.

Feature comparison

What each project ships today, from its own docs and our installs. "Yes" means documented and working in the version we ran, not a roadmap.

Rapid-MLX oMLX Ollama LM Studio mlx-lm (mlx_lm.server)
What it is OpenAI- and Anthropic-compatible inference server for Apple Silicon, built on MLX Inference server for Apple Silicon, built on MLX, run from a macOS menu-bar app Cross-platform model runner and server Desktop app for downloading and chatting with models, with a local server Apple's reference MLX library; its server is a minimal example
OpenAI API Yes Yes Yes (plus its own /api) Yes (plus its own REST API) Yes, basic
Anthropic /v1/messages Yes Yes Yes (since v0.14) Yes (since 0.4.1) No
GUI Rapid-MLX Desktop (macOS app); CLI Menu-bar app plus a web admin panel Desktop app with chat; CLI Full desktop app; lms CLI None (Python CLI)
License Apache-2.0 (engine and Desktop) Apache-2.0 MIT Free to use, closed-source app MIT
Platforms macOS on Apple Silicon macOS on Apple Silicon macOS, Windows, Linux macOS, Windows, Linux macOS on Apple Silicon (MLX also has Linux backends)
Model format / backend MLX MLX (mlx-lm) GGUF (llama.cpp); MLX engine in preview on Macs with 32 GB+ GGUF (llama.cpp) and MLX MLX
Concurrent requests Continuous batching Continuous batching Parallel slots (OLLAMA_NUM_PARALLEL), not for every architecture Parallel predictions (--parallel) Batching (off with a quantized KV cache)
KV / prompt cache In-memory prefix cache; written to disk at shutdown if it fits a short time budget, reloaded at start In-memory hot tier plus an SSD cold tier that survives restarts In-memory, per loaded model In-memory In-memory prompt cache
Tool-call parsing Dedicated parser per model family, picked by the model alias Auto-detected per model family (JSON <tool_call>, Mistral, GLM, MiniMax, Kimi and others) Per-model templates Per-model templates From the tokenizer's chat template
Model catalog Curated aliases with per-model profiles, or any MLX repo on Hugging Face Any mlx-lm model; downloader in the admin panel ollama.com library Hugging Face search in the app Any MLX repo on Hugging Face
Install brew install rapid-mlx, pip, one-line script, or the Desktop DMG DMG, Homebrew tap, or pip from a release wheel DMG, Homebrew, or script DMG pip install mlx-lm

Sources: each project's README and docs, checked 2026-09-26 (oMLX, Ollama Anthropic compatibility, Ollama MLX preview, LM Studio Anthropic endpoint, mlx-lm server). Rapid-MLX flags are from rapid-mlx serve --help at 0.15.2.

Measured on an M4 Pro Mac mini (48 GB)

Engine (config) Decode, 1 stream TTFT, ~2k-token prompt TTFT, ~16k-token prompt 4 streams, aggregate Tool calls (30 prompts)
Rapid-MLX 0.15.3 (@1394f16) (PFlash off) 66.1 tok/s 5.30 s 42.8 s 68.3 tok/s 30/30
Rapid-MLX 0.15.3 (@1394f16) (defaults, PFlash on) not run 5.31 s 8.23 s not run not run
oMLX 0.7.0rc1 (SSD cache on) 51.2 tok/s 5.40 s 41.7 s 97.1 tok/s 30/30
Ollama 0.34.3 (qwen3.5:9b, Q4_K_M) 38.4 tok/s 6.35 s 49.4 s 37.4 tok/s 30/30
mlx-lm 0.31.3 (mlx_lm.server) 47.7 tok/s 5.51 s 43.6 s 96.6 tok/s 30/30

Conditions. Mac mini, Apple M4 Pro, 48 GB unified memory, macOS 26.5.1, on AC power, idle before each engine (GPU 0% after a 60 s rest + 10 idle seconds). Model: Qwen3.5-9B at 4-bit (MLX engines: mlx-community/Qwen3.5-9B-4bit @ 8b2b98c, MTP sidecar mlx-community/Qwen3.5-9B-MTP-4bit @ 222dfd2; Ollama: qwen3.5:9b, GGUF Q4_K_M, id 6488c96fa5fa), same as run 2. One engine at a time, loopback only, temperature 0, streaming. Before each engine the driver rested 60 s and waited for the GPU to read idle, and gave the engine a fresh home directory (no cache or settings carried over). The competitors were measured in run 3; Rapid-MLX 0.15.3 (measured at commit 1394f16, which 0.15.3 includes) was measured a few hours later the same night with the same harness and protocol in run 4 (the competitors' versions had not changed, so they were not re-run). Decode and TTFT cells are the median of 3 runs, the 4-stream cell the median of 2. Decode tok/s is (completion tokens − 1) / (last token − first token), using each engine's own token count. TTFT is request sent to the first streamed token of any kind, thinking included; each prompt starts with a random nonce so that no prefix cache applies. Engines group tokens into stream chunks differently (oMLX sends several tokens per chunk), which can add a fraction of a second to a TTFT. Every request asked for thinking explicitly (chat_template_kwargs.enable_thinking=true), because Rapid-MLX turns thinking off by default for casual requests while the others leave it on; the raw files store the generated text and reasoning, so you can check. Engine settings: Rapid-MLX with --pflash off (see below), everything else default, which for this model includes MTP speculative decoding; oMLX with its documented --paged-ssd-cache-dir; Ollama with a 32k context so no prompt was truncated and OLLAMA_NUM_PARALLEL=4 (Ollama 0.34.3 ignores it for this model and logs "model architecture does not currently support parallel requests"; its four streams' first tokens arrived about 7 s apart, one after another); mlx-lm's server at its defaults. Ollama loads its own GGUF Q4_K_M build of the model, the others load the same MLX 4-bit files.

Why the main Rapid-MLX row has PFlash off. For Qwen3.5 and Qwen3.6 models Rapid-MLX turns on PFlash by default: long-prompt prefill compression that keeps a fraction of the prompt's tokens (20% by default) before prefilling. It makes long prompts much faster (8.23 s instead of 42.8 s at ~16k tokens; it did not engage on the agent session, whose requests carry tools), but the model then does not see every prompt token, so it is not a like-for-like comparison with engines that prefill all of them. We did not measure answer quality with PFlash on, and ran only the prefill and agent-session tests in that configuration.

On an 18 GB Mac

MacBook Pro, Apple M3 Pro, 18 GB unified memory, macOS 15.6.1, on power. Competitors from run 2, Rapid-MLX 0.15.3 (measured at commit 1394f16, which 0.15.3 includes) from run 4, same protocol. On this machine oMLX refused one of its three ~16k-token requests with its prefill memory guard ("Prefill would require ~9.08 GB peak … dynamic ceiling is 9.00 GB"), so its ~16k cell is the median of the two that ran, and Ollama's four streams started about 11 s apart.

Engine (config) Decode, 1 stream TTFT, ~2k-token prompt TTFT, ~16k-token prompt 4 streams, aggregate Tool calls (30 prompts)
Rapid-MLX 0.15.3 (@1394f16) (PFlash off) 43.6 tok/s 5.50 s 44.7 s 62.3 tok/s 30/30
Rapid-MLX 0.15.3 (@1394f16) (defaults, PFlash on) not run 5.50 s 8.58 s not run not run
oMLX 0.7.0rc1 (SSD cache on) 27.9 tok/s 5.73 s 43.9 s 85.3 tok/s 30/30
Ollama 0.34.3 (qwen3.5:9b, Q4_K_M) 22.8 tok/s 6.24 s 49.9 s 22.5 tok/s 30/30
mlx-lm 0.31.3 (mlx_lm.server) 25.6 tok/s 5.87 s 45.6 s 80.7 tok/s 30/30

The coding-agent session (oMLX's headline case)

oMLX's pitch is that Claude Code "responds in 5 seconds, not 90" (omlx.ai): coding agents resend a long system prompt and tool list on every turn, and oMLX keeps that prefix's KV cache on SSD so it survives a restart. We replayed a Claude-Code-shaped session to measure exactly that: a 22,819-token first request (a long system prompt plus 30 tool schemas), then 10 user turns from a fixed transcript (the assistant replies are canned, so every engine sees identical prompts). Then we restarted the server, warmed it with an unrelated prompt, and replayed the same 10 turns. For Ollama, "restart" means unloading the model (ollama stop), which drops its KV cache the same way. Numbers are time to first token per turn; when a server sends nothing until its 16-token reply is complete, the number is time to the end of the stream instead, and the raw records flag those turns (ttft_fallback_end_of_stream).

M4 Pro, 48 GB:

Engine (config) Turn 1 (cold) Turns 2–10, median 10 turns, total After restart: turn 1 After restart: turns 2–10, median
Rapid-MLX 0.15.3 (@1394f16) (PFlash off) 65.1 s 0.49 s 69.6 s 12.3 s 0.49 s
Rapid-MLX 0.15.3 (@1394f16) (defaults, PFlash on) 65.3 s 0.49 s 69.8 s 12.3 s 0.49 s
oMLX 0.7.0rc1 (SSD cache on) 63.1 s 0.49 s 67.6 s 0.27 s 0.27 s
Ollama 0.34.3 (qwen3.5:9b, Q4_K_M) 73.4 s 0.41 s 76.9 s 73.1 s 0.41 s
mlx-lm 0.31.3 (mlx_lm.server) 65.7 s 0.47 s 69.7 s 65.7 s 0.47 s

M3 Pro, 18 GB:

Engine (config) Turn 1 (cold) Turns 2–10, median 10 turns, total After restart: turn 1 After restart: turns 2–10, median
Rapid-MLX 0.15.3 (@1394f16) (PFlash off) 68.5 s 0.69 s 74.9 s 13.0 s 0.67 s
Rapid-MLX 0.15.3 (@1394f16) (defaults, PFlash on) 68.1 s 0.69 s 74.5 s 13.0 s 0.65 s
oMLX 0.7.0rc1 (SSD cache on) 67.0 s 0.66 s 73.0 s 0.56 s 0.44 s
Ollama 0.34.3 (qwen3.5:9b, Q4_K_M) 77.6 s 0.44 s 81.4 s 76.9 s 0.44 s
mlx-lm 0.31.3 (mlx_lm.server) 71.4 s 0.85 s, 0.77 s; server failed on turn 4 n/a 68.9 s 0.86 s; server failed on turn 3

oMLX's claim holds: after one cold prefill it answered every turn in under a second, and its SSD tier made the first turn after a restart nearly free (0.27 s on 48 GB, 0.56 s on 18 GB). Ollama and mlx-lm reused their caches within the session just as well, and lost them on restart. On 18 GB, mlx-lm's server died with a Metal out-of-memory error on turn 4, and after the restart on turn 3 (server log).

Rapid-MLX 0.15.3 (measured at commit 1394f16, which 0.15.3 includes) kept the session prefix resident on both machines: turns 2–10 prefilled only the 42–66 new tokens each, bringing it under a second like the others (0.49 s on 48 GB, 0.69 s on 18 GB), though still behind Ollama on both machines and slightly behind oMLX on 18 GB. At shutdown it saved the session to disk and reloaded it at start, so the first turn after the restart restored 18,944 of the 22,819 prompt tokens and prefilled the remaining 3,875: 12.3 s on 48 GB. That is far better than a full prefill but still well behind oMLX's 0.27 s. The save is not guaranteed: on 18 GB, one of the four shutdowns (the last one, after the restart replay, so no number above depends on it) skipped the entry because its predicted write time (3.1 s) did not fit the 3.5 s shutdown budget minus its 0.4 s commit headroom. Cache budgets and save results for every start are in each run's versions.json and server logs. The prompt token counts differ slightly between engines (about 150 tokens) because each renders the tool schemas through its own chat template.

What changed in Rapid-MLX

The first published version of this page measured the previous release, Rapid-MLX 0.15.2, and it lost the agent session badly on the 18 GB Mac: its prefix cache sized itself to 0.86 GB there, each ~0.96 GB session entry was "too large" to store, and every turn paid the full prefill; what it saved at shutdown did not include the session, so the first turn after a restart was a full prefill too. Commit 1394f16 (#3794) gives agent sessions a minimum cache budget on small-RAM Macs and makes the shutdown save fit in most cases (see above). Same harness, same protocol, PFlash off:

Rapid-MLX, agent session 0.15.2, 18 GB 0.15.3 (@1394f16), 18 GB 0.15.2, 48 GB 0.15.3 (@1394f16), 48 GB
Turns 2–10, median 70.0 s 0.69 s 0.50 s 0.49 s
Turn 1 after a restart 69.1 s 13.0 s 65.1 s 12.3 s

Install 0.15.3 with pip install rapid-mlx==0.15.3, brew upgrade rapid-mlx, or Rapid-MLX Desktop 0.15.3. The benchmark was measured at commit 1394f16, which 0.15.3 includes; it was not re-run on the release tag. The 0.15.2 numbers are in summary-run2.json and summary-run3-m4pro-48gb.json.

Tool calling

Thirty fixed prompts, each with the same eight tools (weather, calculator, web search, email, calendar, read file, shell, currency). A prompt passes when the first tool call names the right function and its JSON arguments contain the expected values. Non-streaming, temperature 0. The prompts, expected calls and every raw response are in the raw data.

Rapid-MLX 0.15.3 (@1394f16) (PFlash off): 30/30 · oMLX 0.7.0rc1 (SSD cache on): 30/30 · Ollama 0.34.3 (qwen3.5:9b, Q4_K_M): 30/30 · mlx-lm 0.31.3 (mlx_lm.server): 30/30. Every engine passed every prompt. These are single-call, clearly worded requests, so the result says the plumbing works on all four servers for this model; it does not rank them. Harder multi-step tool use is where parsers differ, and this test does not cover it.

What we did not measure, and why

When to choose which

Focused pages: Rapid-MLX vs oMLX · Ollama alternative for Mac · our August Ollama/LM Studio/llama.cpp benchmark.

How these runs came about

We publish every run. Run 1 (18 GB) was superseded after a review found three flaws that favored Rapid-MLX or muddied the result: Rapid-MLX had thinking off by default while the other engines were thinking; the engines ran back to back in a fixed order with Rapid-MLX first, on the coolest machine; and one Rapid-MLX start overlapped the shutdown of the previous server. The same review caught two errors on the first version of this page: it said Ollama's four streams ran in parallel (they did not), and the server logs it cited were missing from the published data. Run 2 (18 GB) fixed all of that; run 3 repeated it on a 48 GB M4 Pro; run 4 re-measured only Rapid-MLX at commit 1394f16, which 0.15.3 includes, on both machines. The data README maps each table to its run.

Reproduce it

curl -O https://rapidmlx.com/blog/assets/compare-2026-09/bench_compare.py
# start any OpenAI-compatible server on loopback, then, per scenario a b c d e:
python3 bench_compare.py --engine NAME --base-url http://127.0.0.1:PORT/v1 --model MODEL --thinking on --scenario a --out a.json

The client is standard-library Python. run_all.sh is the driver we used: it starts and stops each server, rests between engines, and restarts the server for the agent session. It expects the engines installed in a scratch directory (W, see its header) and your normal Hugging Face cache in HF_HOME.

Frequently asked questions

Is Rapid-MLX faster than Ollama?

For decode, yes in our runs: on an M4 Pro Mac mini with 48 GB (Qwen3.5-9B at 4-bit), Rapid-MLX 0.15.3 (measured at commit 1394f16, which 0.15.3 includes) decoded 66.1 tok/s on one stream against Ollama's 38.4 tok/s, and 68.3 tok/s against 37.4 tok/s across four streams (Ollama serves this model one request at a time). In a long coding-agent session Ollama answered follow-up turns faster: a median of 0.41 s against 0.49 s for Rapid-MLX, and 0.44 s against 0.69 s on an 18 GB Mac. The previous release, 0.15.2, was far slower there (70.0 s); that is fixed in 0.15.3.

Rapid-MLX vs oMLX: which should I use?

Pick by workload. On an M4 Pro Mac mini with 48 GB, oMLX led four concurrent streams (97.1 tok/s against 68.3 tok/s) and the first agent turn after a restart (0.27 s against 12.3 s, thanks to its SSD cache). Rapid-MLX 0.15.3 (measured at commit 1394f16, which 0.15.3 includes) led single-stream decode (66.1 tok/s against 51.2 tok/s). Follow-up agent turns were level on 48 GB (0.49 s and 0.49 s) with oMLX slightly ahead on 18 GB (0.66 s against 0.69 s), and both passed all 30 tool-call prompts. Details on the Rapid-MLX vs oMLX page.

Does Rapid-MLX work with Claude Code?

Yes. Rapid-MLX serves the Anthropic Messages API at /v1/messages, so Claude Code connects with ANTHROPIC_BASE_URL (see Claude Code on a local model). oMLX, Ollama (v0.14+) and LM Studio (0.4.1+) also serve that API. In our replay of a 23k-token Claude Code session, Rapid-MLX 0.15.3 (measured at commit 1394f16, which 0.15.3 includes) answered follow-up turns in a median of 0.49 s on a 48 GB Mac and 0.69 s on an 18 GB Mac. The previous release, 0.15.2, re-prefilled the whole prompt every turn on the 18 GB Mac (70.0 s); that is fixed in 0.15.3.

Where is the raw data?

Every request's timings, the engine-reported token usage and every tool-call response are listed in raw/manifest.json, with the script (bench_compare.py), the driver (run_all.sh), the engine versions and flags (each run's versions.json) and the numbers used on this page (summary.json).