Rapid-MLX vs mlx-lm, up to 4× faster than Apple's MLX, measured task by task
Same Mac, same model files, greedy decoding, both servers at their defaults. Every task, every run and the scripts are published, including the tasks where Rapid-MLX is barely ahead.
Rapid-MLX 0.15.6 is up to 4× faster than mlx-lm 0.32.0, Apple's MLX library and server, and 1.5× faster on a typical task. Measured on 2026-10-04 on a Mac mini M4 Pro (48 GB) with the same Hugging Face weights in both servers, greedy decoding and both servers at their defaults: Qwen3.5-9B 4-bit decoded whole-file code edits at up to 198.4 tok/s on Rapid-MLX against 46.4 tok/s on mlx-lm, with a median decode speedup of 1.50× across the 18 everyday tasks that stream text (1.37× end to end over all 20); Qwen3.6-35B-A3B 4-bit (MoE) was up to 3.16× faster. The speedup comes from speculative decoding, which Rapid-MLX turns on by default for these models. Long-prompt agent turns are close to a tie.
Disclosure: we build Rapid-MLX. Rapid-MLX is built on MLX, Apple's machine-learning framework, and on parts of mlx-lm itself. The baseline here is mlx-lm, the reference library and server for running language models on MLX, published by Apple's ml-explore team. When we say "Apple's MLX" we mean mlx-lm's server, mlx_lm.server. Every number on this page is generated from the raw records into summary.json; the client, the driver and the tasks are published next to them.
The short version
- Up to 4× faster. Qwen3.5-9B 4-bit, task "Code edit: rename across a file": mlx-lm decoded 46.4 tok/s, Rapid-MLX 198.4 tok/s (4.27×). mlx-lm without its server (
mlx_lm.generate) is a little faster (49.6 tok/s); against that, the same task is 4.00×. That is why we say 4× and not more. - A typical task: 1.5×. Across the 18 tasks that stream text, the median decode speedup on Qwen3.5-9B 4-bit was 1.50× (lowest 1.28×). Measured end to end, from sending the request to the last token, the median over all 20 tasks was 1.37× and the best 3.13×.
- The MoE model too. Qwen3.6-35B-A3B 4-bit (MoE): up to 3.16×, median 1.57×.
- Where it is close. A coding-agent turn with an ~8k-token prompt (1.05× end to end on 9B), the first token of a fresh prompt, and Qwen3.5-4B at its defaults (1.12×), where speculative decoding is off. Details below.
- Same model, same weights, not always the same words. Output was byte-identical on 11 of 20 tasks on 9B; on the rest the two engines picked different but closely ranked tokens at some point. Why.
How it is faster
A language model writes one token at a time, and every step reads all of the model's weights from memory. On a Mac that memory traffic, not arithmetic, is what limits speed. Speculative decoding guesses several tokens cheaply, then has the full model check the whole guess in a single step, which costs about the same as writing one token. Every guessed token the model agrees with is a token you did not wait for; at the first disagreement the model's own token is kept and the rest of the guess is thrown away.
Rapid-MLX turns this on by default for these Qwen models, with two guessers:
- MTP (multi-token prediction). The Qwen3.5 and Qwen3.6 models ship a small extra head trained to predict the next couple of tokens. It drafts every step, and it is right often enough on structured text (JSON, SQL, tables, translations) to give about 1.5×, less on free prose.
- Prompt lookup. When the last few tokens the model wrote also appear in your prompt, Rapid-MLX proposes whatever followed them in the prompt, a long run at a time. Ask for a whole file back with a rename or a bug fix, and most of the answer is a copy of the file you sent, so most of each guess is accepted. That is where the largest gains come from.
mlx-lm can also do speculative decoding if you give it a separate draft model (--draft-model); it is off by default. This page compares both servers as they come.
Per-task results
Each row is one task: the speedup (Rapid-MLX tokens per second divided by mlx-lm's), and underneath it both engines' decode speed and total time. The two tool-call tasks return their answer in one chunk on both servers, so they show the end-to-end ratio instead. Each cell is the median of 3 runs, each run on a freshly started server.
Qwen3.5-9B, 4-bit
Rapid-MLX default: MTP and prompt lookup on. Up to 4.27×, median 1.50×.
- Code edit: rename across a file 4.27×decode 46.4 → 198.4 tok/s · total 34.3 s → 11.0 s · same output
- Code edit: fix a bug, return whole file 4.22×decode 45.9 → 193.7 tok/s · total 46.3 s → 14.8 s · output differs
- Code edit: apply review comments 3.85×decode 45.2 → 174.1 tok/s · total 56.9 s → 19.2 s · same output
- Code edit: produce a unified diff 1.37×decode 46.3 → 63.2 tok/s · total 11.4 s → 9.37 s · same output
- JSON transform (rename keys, convert types) 1.64×decode 46.0 → 75.3 tok/s · total 35.1 s → 23.0 s · same output
- Format conversion: JSON to CSV 1.53×decode 46.3 → 70.7 tok/s · total 15.8 s → 11.8 s · same output
- Structured output: fill a JSON schema 1.58×decode 47.0 → 74.2 tok/s · total 8.03 s → 5.96 s · same output
- Extract fields from a document 1.58×decode 47.0 → 74.3 tok/s · total 9.12 s → 6.33 s · same output
- Tool call (OpenAI tools) 1.25× e2eend to end 5.14 s → 4.12 s (the reply arrives in one chunk, so no decode rate) · same output
- Coding-agent turn, ~8k-token context with tool results 1.05× e2eend to end 28.5 s → 27.0 s (the reply arrives in one chunk, so no decode rate) · same output
- Summarize an article 1.36×decode 47.0 → 63.9 tok/s · total 8.38 s → 7.01 s · output differs
- Translate (English to Spanish) 1.47×decode 47.5 → 69.5 tok/s · total 4.51 s → 3.18 s · output differs
- Proofread a document, return full text 2.18×decode 47.1 → 102.8 tok/s · total 10.7 s → 5.63 s · same output
- Write an email 1.38×decode 47.7 → 65.6 tok/s · total 3.43 s → 2.53 s · output differs
- Free-form prose (short story) 1.28×decode 47.1 → 60.1 tok/s · total 15.9 s → 13.5 s · output differs
- Explain code 1.40×decode 46.2 → 64.7 tok/s · total 19.9 s → 15.5 s · output differs
- SQL 1.49×decode 47.0 → 69.9 tok/s · total 10.2 s → 7.00 s · output differs
- Regex 1.39×decode 47.1 → 65.5 tok/s · total 9.89 s → 7.19 s · output differs
- Markdown table from notes 1.52×decode 47.2 → 71.6 tok/s · total 7.17 s → 4.99 s · same output
- Long technical answer (~800 tokens) 1.36×decode 46.8 → 63.4 tok/s · total 23.7 s → 17.7 s · output differs
Qwen3.6-35B-A3B, 4-bit (MoE)
Rapid-MLX default: MTP and prompt lookup on. Up to 3.16×, median 1.57×.
- Code edit: rename across a file 3.16×decode 71.9 → 227.6 tok/s · total 21.4 s → 8.07 s · same output
- Code edit: fix a bug, return whole file 3.09×decode 71.3 → 220.6 tok/s · total 28.6 s → 10.8 s · output differs
- Code edit: apply review comments 2.79×decode 70.6 → 196.7 tok/s · total 35.2 s → 14.4 s · same output
- Code edit: produce a unified diff 1.47×decode 72.3 → 106.1 tok/s · total 5.41 s → 4.13 s · output differs
- JSON transform (rename keys, convert types) 1.60×decode 71.6 → 114.7 tok/s · total 21.5 s → 14.0 s · same output
- Format conversion: JSON to CSV 1.56×decode 72.4 → 113.1 tok/s · total 9.36 s → 6.59 s · same output
- Structured output: fill a JSON schema 1.57×decode 74.0 → 116.0 tok/s · total 4.80 s → 3.39 s · same output
- Extract fields from a document 1.61×decode 74.1 → 119.7 tok/s · total 5.73 s → 3.82 s · same output
- Tool call (OpenAI tools) 1.23× e2eend to end 3.08 s → 2.51 s (the reply arrives in one chunk, so no decode rate) · same output
- Coding-agent turn, ~8k-token context with tool results 1.10× e2eend to end 15.2 s → 13.8 s (the reply arrives in one chunk, so no decode rate) · same output
- Summarize an article 1.36×decode 73.8 → 100.6 tok/s · total 5.36 s → 4.28 s · output differs
- Translate (English to Spanish) 1.56×decode 74.8 → 116.4 tok/s · total 2.91 s → 2.06 s · output differs
- Proofread a document, return full text 1.96×decode 74.2 → 145.4 tok/s · total 6.61 s → 3.75 s · same output
- Write an email 1.43×decode 75.1 → 107.4 tok/s · total 2.28 s → 1.63 s · output differs
- Free-form prose (short story) 1.19×decode 74.3 → 88.4 tok/s · total 8.43 s → 7.52 s · output differs
- Explain code 1.29×decode 71.7 → 92.6 tok/s · total 12.0 s → 9.76 s · output differs
- SQL 1.58×decode 74.5 → 117.4 tok/s · total 4.51 s → 3.22 s · output differs
- Regex 1.57×decode 74.4 → 116.7 tok/s · total 6.32 s → 4.13 s · output differs
- Markdown table from notes 1.59×decode 74.5 → 118.4 tok/s · total 4.38 s → 2.95 s · output differs
- Long technical answer (~800 tokens) 1.48×decode 74.1 → 109.7 tok/s · total 13.8 s → 9.42 s · output differs
Qwen3.5-4B, 4-bit, at defaults
Rapid-MLX turns MTP off by default for this model, because earlier measurements on M2 Pro and M3 Ultra Macs found it slower there. Without speculative decoding the gain is small: up to 1.14×, median 1.12×.
- Code edit: rename across a file 1.13×decode 73.1 → 82.4 tok/s · total 21.4 s → 19.2 s · same output
- Code edit: fix a bug, return whole file 1.13×decode 72.3 → 82.0 tok/s · total 29.5 s → 26.0 s · output differs
- Code edit: apply review comments 1.14×decode 71.1 → 81.2 tok/s · total 35.5 s → 31.3 s · same output
- Code edit: produce a unified diff 1.12×decode 72.1 → 81.0 tok/s · total 15.5 s → 13.9 s · same output
- JSON transform (rename keys, convert types) 1.12×decode 72.4 → 80.9 tok/s · total 22.0 s → 19.9 s · same output
- Format conversion: JSON to CSV 1.12×decode 73.0 → 81.6 tok/s · total 9.81 s → 8.96 s · same output
- Structured output: fill a JSON schema 1.11×decode 74.5 → 83.1 tok/s · total 4.90 s → 4.51 s · same output
- Extract fields from a document 1.12×decode 74.7 → 83.3 tok/s · total 5.87 s → 5.34 s · same output
- Tool call (OpenAI tools) 1.24× e2eend to end 3.36 s → 2.71 s (the reply arrives in one chunk, so no decode rate) · output differs
- Coding-agent turn, ~8k-token context with tool results 0.99× e2eend to end 17.0 s → 17.1 s (the reply arrives in one chunk, so no decode rate) · output differs
- Summarize an article 1.11×decode 74.7 → 83.3 tok/s · total 4.54 s → 3.98 s · output differs
- Translate (English to Spanish) 1.11×decode 75.8 → 84.2 tok/s · total 2.76 s → 2.53 s · same output
- Proofread a document, return full text 1.12×decode 74.8 → 83.5 tok/s · total 6.63 s → 6.03 s · same output
- Write an email 1.11×decode 76.2 → 84.8 tok/s · total 1.77 s → 1.71 s · output differs
- Free-form prose (short story) 1.12×decode 75.0 → 83.8 tok/s · total 8.69 s → 8.29 s · output differs
- Explain code 1.12×decode 73.0 → 81.7 tok/s · total 12.2 s → 11.1 s · output differs
- SQL 1.12×decode 75.4 → 84.1 tok/s · total 4.98 s → 4.41 s · output differs
- Regex 1.12×decode 75.2 → 84.0 tok/s · total 6.17 s → 5.55 s · output differs
- Markdown table from notes 1.12×decode 75.5 → 84.2 tok/s · total 4.47 s → 4.07 s · output differs
- Long technical answer (~800 tokens) 1.12×decode 74.7 → 83.6 tok/s · total 14.9 s → 13.4 s · output differs
Qwen3.5-4B with MTP turned on
The same model with --speculative-config '{"method":"mtp"}'. On this M4 Pro it was faster on every task: up to 4.13×, median 1.52×. It is not the default, and we have not re-checked it on other chips, so the headline does not use it.
- Code edit: rename across a file 4.13×decode 73.1 → 302.0 tok/s · total 21.4 s → 6.79 s · same output
- Code edit: fix a bug, return whole file 4.05×decode 72.3 → 292.5 tok/s · total 29.5 s → 9.21 s · output differs
- Code edit: apply review comments 3.72×decode 71.1 → 264.7 tok/s · total 35.5 s → 12.0 s · same output
- Code edit: produce a unified diff 1.49×decode 72.1 → 107.3 tok/s · total 15.5 s → 11.0 s · same output
- JSON transform (rename keys, convert types) 1.60×decode 72.4 → 116.1 tok/s · total 22.0 s → 14.4 s · same output
- Format conversion: JSON to CSV 1.53×decode 73.0 → 111.8 tok/s · total 9.81 s → 7.04 s · same output
- Structured output: fill a JSON schema 1.58×decode 74.5 → 117.9 tok/s · total 4.90 s → 3.48 s · same output
- Extract fields from a document 1.60×decode 74.7 → 119.6 tok/s · total 5.87 s → 3.94 s · same output
- Tool call (OpenAI tools) 1.42× e2eend to end 3.36 s → 2.36 s (the reply arrives in one chunk, so no decode rate) · output differs
- Coding-agent turn, ~8k-token context with tool results 1.05× e2eend to end 17.0 s → 16.2 s (the reply arrives in one chunk, so no decode rate) · output differs
- Summarize an article 1.37×decode 74.7 → 102.6 tok/s · total 4.54 s → 3.67 s · output differs
- Translate (English to Spanish) 1.52×decode 75.8 → 115.1 tok/s · total 2.76 s → 1.94 s · same output
- Proofread a document, return full text 2.08×decode 74.8 → 155.5 tok/s · total 6.63 s → 3.60 s · same output
- Write an email 1.28×decode 76.2 → 97.3 tok/s · total 1.77 s → 1.47 s · output differs
- Free-form prose (short story) 1.30×decode 75.0 → 97.7 tok/s · total 8.69 s → 7.56 s · output differs
- Explain code 1.41×decode 73.0 → 103.4 tok/s · total 12.2 s → 9.36 s · output differs
- SQL 1.51×decode 75.4 → 114.3 tok/s · total 4.98 s → 3.34 s · output differs
- Regex 1.44×decode 75.2 → 108.6 tok/s · total 6.17 s → 4.34 s · output differs
- Markdown table from notes 1.56×decode 75.5 → 118.0 tok/s · total 4.47 s → 3.01 s · same output
- Long technical answer (~800 tokens) 1.35×decode 74.7 → 100.8 tok/s · total 14.9 s → 10.2 s · output differs
Where it is close
- Long-prompt agent turns. The coding-agent task sends an ~8k-token prompt (a repository conversation with tool results) and gets a short tool call back. Almost all of the time goes into reading the prompt, which both servers do with the same MLX kernels: time to first token 28.4 s on mlx-lm and 27.0 s on Rapid-MLX (9B). End to end that is 1.05× on 9B, 1.10× on 35B and 0.99× on 4B.
- The first token of a fresh prompt. For a prompt neither server has seen, time to first token is level, for example 3.95 s against 3.88 s on the 9B rename task.
- Qwen3.5-4B at its defaults: 1.12× at the median, because speculative decoding is off for that model (see above).
- Free prose. A short story is the smallest win on every model with speculative decoding: 1.28× on 9B, 1.19× on 35B. There is nothing to copy from the prompt, and MTP guesses free prose less well.
- Four concurrent streams. With the harness from our server comparison (9B, thinking on), Rapid-MLX aggregated 104.2 tok/s and mlx-lm 100.2 tok/s (1.04×): parity, not a win.
- Follow-up turns of a long agent session, mlx-lm is faster. In the same harness's 22,675-token session, both servers reuse the cached prefix, and mlx-lm answered turns 2–10 in a median of 0.36 s against 0.47 s for Rapid-MLX.
Method
- Machine. Mac mini, Apple M4 Pro, 48 GB unified memory, macOS 26.5.1, on AC power, idle before every server start (30 s rest, then 10 seconds of idle GPU; no other model server running).
- Baseline. mlx-lm 0.32.0, the latest release on PyPI when we measured, started with
python -m mlx_lm server --model <snapshot>and nothing else. - Rapid-MLX. 0.15.6 from PyPI, started with
rapid-mlx serve <alias>and nothing else (the MTP row on 4B adds the one flag shown). - Same weights. Both servers load the same Hugging Face snapshots: mlx-community/Qwen3.5-9B-4bit, mlx-community/Qwen3.6-35B-A3B-4bit and mlx-community/Qwen3.5-4B-MLX-4bit, at the revisions listed in versions.json. Where MTP is on, Rapid-MLX also loads the model's small MTP head from a companion repository.
- Requests. Streaming
/v1/chat/completions, temperature 0 (greedy), a fixed token limit per task, Qwen thinking switched off, usage reported by each engine. - Tasks. 20 tasks of the kind people send a local model: whole-file code edits, a unified diff, JSON transforms and extraction, a tool call, a coding-agent turn, summary, translation, proofreading, an email, a story, code explanation, SQL, a regex, a Markdown table and a long technical answer. The code in the prompts is unmodified Python standard-library source. All prompts are in tasks.json.
- Runs. 3 runs per model and server. Every run starts a fresh server with a fresh home directory, so no cache carries over, sends two warm-up requests, then each task once. The order of the servers rotates between runs, and before every start the driver waited for the GPU to be idle.
- Measures. Decode speed is (completion tokens − 1) divided by the time from the first to the last streamed token, using each engine's own token count. End to end is request sent to last token. Cells are medians over runs.
- Measured on 2026-10-04.
Are the answers the same?
Not always, and this page does not claim they are. With greedy decoding, the output was byte-identical on 11 of 20 tasks on 9B, 9 of 20 on 35B and 9 of 20 on 4B. Most code edits and the JSON, CSV and extraction tasks matched; prose usually did not.
Where we tested it, speculative decoding is not the cause. On 9B, Rapid-MLX with speculative decoding switched off (--no-spec-decode) differed from mlx-lm on every re-run task as well, and Qwen3.5-4B at its defaults uses no speculative decoding at all and still differs on prose. We did not run that control on 35B. The two servers run different code paths over the same weights (Rapid-MLX 0.15.6 uses its own scheduler on top of an earlier mlx-lm), so tiny numerical differences can flip a choice between two nearly equally likely tokens, after which the texts go separate ways. We looked at the divergences with a plain mlx-lm forward pass; where the two chosen tokens can be matched to the model's top candidates, they were close alternatives near the top of its ranking (tie-check data). With MTP on, Rapid-MLX's greedy output on some prose tasks also varied slightly between runs; code and JSON output did not.
What this does not cover
- One Mac, an M4 Pro. Speculative decoding pays off differently on other chips; that is why it is off by default for the 4B model.
- Thinking was off. With thinking on, most of the output is reasoning prose, which gains less than code edits. We did not measure that here.
- Defaults only. We did not run mlx-lm with a draft model.
Reproduce it
B=https://rapidmlx.com/blog/assets/vs-mlx-2026-10 curl -O $B/bench_tasks.py -O $B/tasks.json # start any OpenAI-compatible server on loopback, then: python3 bench_tasks.py --base-url http://127.0.0.1:PORT/v1 --model MODEL --label NAME --tasks tasks.json --out NAME.json
The client is standard-library Python. run_tasks.sh and lib.sh are the driver we used (fresh server and home directory per run, rotating order, idle-GPU wait), analyze.py computes the medians and the identical-output check, and the data README lists every file.
Frequently asked questions
Is Rapid-MLX faster than MLX?
Yes, against mlx-lm, Apple's MLX library and server, on the models we measured. On a Mac mini M4 Pro (48 GB), Rapid-MLX 0.15.6 decoded up to 4× faster than mlx-lm 0.32.0 (Qwen3.5-9B 4-bit, whole-file code edits) and 1.5× faster on a typical task, with the same weights, greedy decoding and both servers at their defaults (measured 2026-10-04). Rapid-MLX itself runs on MLX; the difference is speculative decoding, on by default for these models.
How much faster is Rapid-MLX than mlx-lm?
It depends on the task. On Qwen3.5-9B 4-bit: 4.27× at best (a whole file returned with a rename, 46.4 tok/s against 198.4 tok/s), 1.50× at the median of the 18 tasks that stream text (1.37× end to end over all 20), 1.28× at the least (Free-form prose (short story)), and 1.05× end to end on a long-prompt agent turn. Qwen3.6-35B-A3B 4-bit (MoE): up to 3.16×, median 1.57×. Qwen3.5-4B at its defaults, without speculative decoding: median 1.12×. The headline says 4× rather than 4.27× because mlx-lm without its server (mlx_lm.generate) is slightly faster, and against it the best task is 4.00×.
Does Rapid-MLX change model outputs?
It loads the same weights, and speculative decoding only keeps tokens the full model agrees with. The text is still not always byte-identical to mlx-lm's: with greedy decoding it matched on 11 of 20 tasks on Qwen3.5-9B 4-bit. Where it differed, the two engines typically picked between nearly equally likely tokens, and on the models we re-ran with speculative decoding switched off the outputs differed too, which points to small numerical differences between the two code paths rather than the speedup.
Which models get the speedup?
Models where Rapid-MLX turns speculative decoding on by default, such as Qwen3.5-9B and Qwen3.6-35B-A3B. Qwen3.5-4B has it off by default and was 1.12× faster at the median; with MTP turned on it reached 4.13× on this Mac.