Rapid-MLX vs mlx-lm, up to 4× faster than Apple's MLX, measured task by task

Same Mac, same model files, greedy decoding, both servers at their defaults. Every task, every run and the scripts are published, including the tasks where Rapid-MLX is barely ahead.

Rapid-MLX 0.15.6 is up to 4× faster than mlx-lm 0.32.0, Apple's MLX library and server, and 1.5× faster on a typical task. Measured on 2026-10-04 on a Mac mini M4 Pro (48 GB) with the same Hugging Face weights in both servers, greedy decoding and both servers at their defaults: Qwen3.5-9B 4-bit decoded whole-file code edits at up to 198.4 tok/s on Rapid-MLX against 46.4 tok/s on mlx-lm, with a median decode speedup of 1.50× across the 18 everyday tasks that stream text (1.37× end to end over all 20); Qwen3.6-35B-A3B 4-bit (MoE) was up to 3.16× faster. The speedup comes from speculative decoding, which Rapid-MLX turns on by default for these models. Long-prompt agent turns are close to a tie.

Disclosure: we build Rapid-MLX. Rapid-MLX is built on MLX, Apple's machine-learning framework, and on parts of mlx-lm itself. The baseline here is mlx-lm, the reference library and server for running language models on MLX, published by Apple's ml-explore team. When we say "Apple's MLX" we mean mlx-lm's server, mlx_lm.server. Every number on this page is generated from the raw records into summary.json; the client, the driver and the tasks are published next to them.

The short version

How it is faster

A language model writes one token at a time, and every step reads all of the model's weights from memory. On a Mac that memory traffic, not arithmetic, is what limits speed. Speculative decoding guesses several tokens cheaply, then has the full model check the whole guess in a single step, which costs about the same as writing one token. Every guessed token the model agrees with is a token you did not wait for; at the first disagreement the model's own token is kept and the rest of the guess is thrown away.

Rapid-MLX turns this on by default for these Qwen models, with two guessers:

mlx-lm can also do speculative decoding if you give it a separate draft model (--draft-model); it is off by default. This page compares both servers as they come.

Per-task results

Each row is one task: the speedup (Rapid-MLX tokens per second divided by mlx-lm's), and underneath it both engines' decode speed and total time. The two tool-call tasks return their answer in one chunk on both servers, so they show the end-to-end ratio instead. Each cell is the median of 3 runs, each run on a freshly started server.

Qwen3.5-9B, 4-bit

Rapid-MLX default: MTP and prompt lookup on. Up to 4.27×, median 1.50×.

Qwen3.5-9B 4-bit: Rapid-MLX speedup over mlx-lm per task (decode tokens/s unless marked e2e). Median of 3 runs per cell.
  1. Code edit: rename across a file 4.27×
    decode 46.4 → 198.4 tok/s · total 34.3 s → 11.0 s · same output
  2. Code edit: fix a bug, return whole file 4.22×
    decode 45.9 → 193.7 tok/s · total 46.3 s → 14.8 s · output differs
  3. Code edit: apply review comments 3.85×
    decode 45.2 → 174.1 tok/s · total 56.9 s → 19.2 s · same output
  4. Code edit: produce a unified diff 1.37×
    decode 46.3 → 63.2 tok/s · total 11.4 s → 9.37 s · same output
  5. JSON transform (rename keys, convert types) 1.64×
    decode 46.0 → 75.3 tok/s · total 35.1 s → 23.0 s · same output
  6. Format conversion: JSON to CSV 1.53×
    decode 46.3 → 70.7 tok/s · total 15.8 s → 11.8 s · same output
  7. Structured output: fill a JSON schema 1.58×
    decode 47.0 → 74.2 tok/s · total 8.03 s → 5.96 s · same output
  8. Extract fields from a document 1.58×
    decode 47.0 → 74.3 tok/s · total 9.12 s → 6.33 s · same output
  9. Tool call (OpenAI tools) 1.25× e2e
    end to end 5.14 s → 4.12 s (the reply arrives in one chunk, so no decode rate) · same output
  10. Coding-agent turn, ~8k-token context with tool results 1.05× e2e
    end to end 28.5 s → 27.0 s (the reply arrives in one chunk, so no decode rate) · same output
  11. Summarize an article 1.36×
    decode 47.0 → 63.9 tok/s · total 8.38 s → 7.01 s · output differs
  12. Translate (English to Spanish) 1.47×
    decode 47.5 → 69.5 tok/s · total 4.51 s → 3.18 s · output differs
  13. Proofread a document, return full text 2.18×
    decode 47.1 → 102.8 tok/s · total 10.7 s → 5.63 s · same output
  14. Write an email 1.38×
    decode 47.7 → 65.6 tok/s · total 3.43 s → 2.53 s · output differs
  15. Free-form prose (short story) 1.28×
    decode 47.1 → 60.1 tok/s · total 15.9 s → 13.5 s · output differs
  16. Explain code 1.40×
    decode 46.2 → 64.7 tok/s · total 19.9 s → 15.5 s · output differs
  17. SQL 1.49×
    decode 47.0 → 69.9 tok/s · total 10.2 s → 7.00 s · output differs
  18. Regex 1.39×
    decode 47.1 → 65.5 tok/s · total 9.89 s → 7.19 s · output differs
  19. Markdown table from notes 1.52×
    decode 47.2 → 71.6 tok/s · total 7.17 s → 4.99 s · same output
  20. Long technical answer (~800 tokens) 1.36×
    decode 46.8 → 63.4 tok/s · total 23.7 s → 17.7 s · output differs

Qwen3.6-35B-A3B, 4-bit (MoE)

Rapid-MLX default: MTP and prompt lookup on. Up to 3.16×, median 1.57×.

Qwen3.6-35B-A3B 4-bit (MoE): Rapid-MLX speedup over mlx-lm per task (decode tokens/s unless marked e2e). Median of 3 runs per cell.
  1. Code edit: rename across a file 3.16×
    decode 71.9 → 227.6 tok/s · total 21.4 s → 8.07 s · same output
  2. Code edit: fix a bug, return whole file 3.09×
    decode 71.3 → 220.6 tok/s · total 28.6 s → 10.8 s · output differs
  3. Code edit: apply review comments 2.79×
    decode 70.6 → 196.7 tok/s · total 35.2 s → 14.4 s · same output
  4. Code edit: produce a unified diff 1.47×
    decode 72.3 → 106.1 tok/s · total 5.41 s → 4.13 s · output differs
  5. JSON transform (rename keys, convert types) 1.60×
    decode 71.6 → 114.7 tok/s · total 21.5 s → 14.0 s · same output
  6. Format conversion: JSON to CSV 1.56×
    decode 72.4 → 113.1 tok/s · total 9.36 s → 6.59 s · same output
  7. Structured output: fill a JSON schema 1.57×
    decode 74.0 → 116.0 tok/s · total 4.80 s → 3.39 s · same output
  8. Extract fields from a document 1.61×
    decode 74.1 → 119.7 tok/s · total 5.73 s → 3.82 s · same output
  9. Tool call (OpenAI tools) 1.23× e2e
    end to end 3.08 s → 2.51 s (the reply arrives in one chunk, so no decode rate) · same output
  10. Coding-agent turn, ~8k-token context with tool results 1.10× e2e
    end to end 15.2 s → 13.8 s (the reply arrives in one chunk, so no decode rate) · same output
  11. Summarize an article 1.36×
    decode 73.8 → 100.6 tok/s · total 5.36 s → 4.28 s · output differs
  12. Translate (English to Spanish) 1.56×
    decode 74.8 → 116.4 tok/s · total 2.91 s → 2.06 s · output differs
  13. Proofread a document, return full text 1.96×
    decode 74.2 → 145.4 tok/s · total 6.61 s → 3.75 s · same output
  14. Write an email 1.43×
    decode 75.1 → 107.4 tok/s · total 2.28 s → 1.63 s · output differs
  15. Free-form prose (short story) 1.19×
    decode 74.3 → 88.4 tok/s · total 8.43 s → 7.52 s · output differs
  16. Explain code 1.29×
    decode 71.7 → 92.6 tok/s · total 12.0 s → 9.76 s · output differs
  17. SQL 1.58×
    decode 74.5 → 117.4 tok/s · total 4.51 s → 3.22 s · output differs
  18. Regex 1.57×
    decode 74.4 → 116.7 tok/s · total 6.32 s → 4.13 s · output differs
  19. Markdown table from notes 1.59×
    decode 74.5 → 118.4 tok/s · total 4.38 s → 2.95 s · output differs
  20. Long technical answer (~800 tokens) 1.48×
    decode 74.1 → 109.7 tok/s · total 13.8 s → 9.42 s · output differs

Qwen3.5-4B, 4-bit, at defaults

Rapid-MLX turns MTP off by default for this model, because earlier measurements on M2 Pro and M3 Ultra Macs found it slower there. Without speculative decoding the gain is small: up to 1.14×, median 1.12×.

Qwen3.5-4B 4-bit: Rapid-MLX speedup over mlx-lm per task (decode tokens/s unless marked e2e). Median of 3 runs per cell.
  1. Code edit: rename across a file 1.13×
    decode 73.1 → 82.4 tok/s · total 21.4 s → 19.2 s · same output
  2. Code edit: fix a bug, return whole file 1.13×
    decode 72.3 → 82.0 tok/s · total 29.5 s → 26.0 s · output differs
  3. Code edit: apply review comments 1.14×
    decode 71.1 → 81.2 tok/s · total 35.5 s → 31.3 s · same output
  4. Code edit: produce a unified diff 1.12×
    decode 72.1 → 81.0 tok/s · total 15.5 s → 13.9 s · same output
  5. JSON transform (rename keys, convert types) 1.12×
    decode 72.4 → 80.9 tok/s · total 22.0 s → 19.9 s · same output
  6. Format conversion: JSON to CSV 1.12×
    decode 73.0 → 81.6 tok/s · total 9.81 s → 8.96 s · same output
  7. Structured output: fill a JSON schema 1.11×
    decode 74.5 → 83.1 tok/s · total 4.90 s → 4.51 s · same output
  8. Extract fields from a document 1.12×
    decode 74.7 → 83.3 tok/s · total 5.87 s → 5.34 s · same output
  9. Tool call (OpenAI tools) 1.24× e2e
    end to end 3.36 s → 2.71 s (the reply arrives in one chunk, so no decode rate) · output differs
  10. Coding-agent turn, ~8k-token context with tool results 0.99× e2e
    end to end 17.0 s → 17.1 s (the reply arrives in one chunk, so no decode rate) · output differs
  11. Summarize an article 1.11×
    decode 74.7 → 83.3 tok/s · total 4.54 s → 3.98 s · output differs
  12. Translate (English to Spanish) 1.11×
    decode 75.8 → 84.2 tok/s · total 2.76 s → 2.53 s · same output
  13. Proofread a document, return full text 1.12×
    decode 74.8 → 83.5 tok/s · total 6.63 s → 6.03 s · same output
  14. Write an email 1.11×
    decode 76.2 → 84.8 tok/s · total 1.77 s → 1.71 s · output differs
  15. Free-form prose (short story) 1.12×
    decode 75.0 → 83.8 tok/s · total 8.69 s → 8.29 s · output differs
  16. Explain code 1.12×
    decode 73.0 → 81.7 tok/s · total 12.2 s → 11.1 s · output differs
  17. SQL 1.12×
    decode 75.4 → 84.1 tok/s · total 4.98 s → 4.41 s · output differs
  18. Regex 1.12×
    decode 75.2 → 84.0 tok/s · total 6.17 s → 5.55 s · output differs
  19. Markdown table from notes 1.12×
    decode 75.5 → 84.2 tok/s · total 4.47 s → 4.07 s · output differs
  20. Long technical answer (~800 tokens) 1.12×
    decode 74.7 → 83.6 tok/s · total 14.9 s → 13.4 s · output differs

Qwen3.5-4B with MTP turned on

The same model with --speculative-config '{"method":"mtp"}'. On this M4 Pro it was faster on every task: up to 4.13×, median 1.52×. It is not the default, and we have not re-checked it on other chips, so the headline does not use it.

Qwen3.5-4B 4-bit: Rapid-MLX speedup over mlx-lm per task (decode tokens/s unless marked e2e). Median of 3 runs per cell.
  1. Code edit: rename across a file 4.13×
    decode 73.1 → 302.0 tok/s · total 21.4 s → 6.79 s · same output
  2. Code edit: fix a bug, return whole file 4.05×
    decode 72.3 → 292.5 tok/s · total 29.5 s → 9.21 s · output differs
  3. Code edit: apply review comments 3.72×
    decode 71.1 → 264.7 tok/s · total 35.5 s → 12.0 s · same output
  4. Code edit: produce a unified diff 1.49×
    decode 72.1 → 107.3 tok/s · total 15.5 s → 11.0 s · same output
  5. JSON transform (rename keys, convert types) 1.60×
    decode 72.4 → 116.1 tok/s · total 22.0 s → 14.4 s · same output
  6. Format conversion: JSON to CSV 1.53×
    decode 73.0 → 111.8 tok/s · total 9.81 s → 7.04 s · same output
  7. Structured output: fill a JSON schema 1.58×
    decode 74.5 → 117.9 tok/s · total 4.90 s → 3.48 s · same output
  8. Extract fields from a document 1.60×
    decode 74.7 → 119.6 tok/s · total 5.87 s → 3.94 s · same output
  9. Tool call (OpenAI tools) 1.42× e2e
    end to end 3.36 s → 2.36 s (the reply arrives in one chunk, so no decode rate) · output differs
  10. Coding-agent turn, ~8k-token context with tool results 1.05× e2e
    end to end 17.0 s → 16.2 s (the reply arrives in one chunk, so no decode rate) · output differs
  11. Summarize an article 1.37×
    decode 74.7 → 102.6 tok/s · total 4.54 s → 3.67 s · output differs
  12. Translate (English to Spanish) 1.52×
    decode 75.8 → 115.1 tok/s · total 2.76 s → 1.94 s · same output
  13. Proofread a document, return full text 2.08×
    decode 74.8 → 155.5 tok/s · total 6.63 s → 3.60 s · same output
  14. Write an email 1.28×
    decode 76.2 → 97.3 tok/s · total 1.77 s → 1.47 s · output differs
  15. Free-form prose (short story) 1.30×
    decode 75.0 → 97.7 tok/s · total 8.69 s → 7.56 s · output differs
  16. Explain code 1.41×
    decode 73.0 → 103.4 tok/s · total 12.2 s → 9.36 s · output differs
  17. SQL 1.51×
    decode 75.4 → 114.3 tok/s · total 4.98 s → 3.34 s · output differs
  18. Regex 1.44×
    decode 75.2 → 108.6 tok/s · total 6.17 s → 4.34 s · output differs
  19. Markdown table from notes 1.56×
    decode 75.5 → 118.0 tok/s · total 4.47 s → 3.01 s · same output
  20. Long technical answer (~800 tokens) 1.35×
    decode 74.7 → 100.8 tok/s · total 14.9 s → 10.2 s · output differs

Where it is close

Method

Are the answers the same?

Not always, and this page does not claim they are. With greedy decoding, the output was byte-identical on 11 of 20 tasks on 9B, 9 of 20 on 35B and 9 of 20 on 4B. Most code edits and the JSON, CSV and extraction tasks matched; prose usually did not.

Where we tested it, speculative decoding is not the cause. On 9B, Rapid-MLX with speculative decoding switched off (--no-spec-decode) differed from mlx-lm on every re-run task as well, and Qwen3.5-4B at its defaults uses no speculative decoding at all and still differs on prose. We did not run that control on 35B. The two servers run different code paths over the same weights (Rapid-MLX 0.15.6 uses its own scheduler on top of an earlier mlx-lm), so tiny numerical differences can flip a choice between two nearly equally likely tokens, after which the texts go separate ways. We looked at the divergences with a plain mlx-lm forward pass; where the two chosen tokens can be matched to the model's top candidates, they were close alternatives near the top of its ranking (tie-check data). With MTP on, Rapid-MLX's greedy output on some prose tasks also varied slightly between runs; code and JSON output did not.

What this does not cover

Reproduce it

B=https://rapidmlx.com/blog/assets/vs-mlx-2026-10
curl -O $B/bench_tasks.py -O $B/tasks.json
# start any OpenAI-compatible server on loopback, then:
python3 bench_tasks.py --base-url http://127.0.0.1:PORT/v1 --model MODEL --label NAME --tasks tasks.json --out NAME.json

The client is standard-library Python. run_tasks.sh and lib.sh are the driver we used (fresh server and home directory per run, rotating order, idle-GPU wait), analyze.py computes the medians and the identical-output check, and the data README lists every file.

Frequently asked questions

Is Rapid-MLX faster than MLX?

Yes, against mlx-lm, Apple's MLX library and server, on the models we measured. On a Mac mini M4 Pro (48 GB), Rapid-MLX 0.15.6 decoded up to 4× faster than mlx-lm 0.32.0 (Qwen3.5-9B 4-bit, whole-file code edits) and 1.5× faster on a typical task, with the same weights, greedy decoding and both servers at their defaults (measured 2026-10-04). Rapid-MLX itself runs on MLX; the difference is speculative decoding, on by default for these models.

How much faster is Rapid-MLX than mlx-lm?

It depends on the task. On Qwen3.5-9B 4-bit: 4.27× at best (a whole file returned with a rename, 46.4 tok/s against 198.4 tok/s), 1.50× at the median of the 18 tasks that stream text (1.37× end to end over all 20), 1.28× at the least (Free-form prose (short story)), and 1.05× end to end on a long-prompt agent turn. Qwen3.6-35B-A3B 4-bit (MoE): up to 3.16×, median 1.57×. Qwen3.5-4B at its defaults, without speculative decoding: median 1.12×. The headline says 4× rather than 4.27× because mlx-lm without its server (mlx_lm.generate) is slightly faster, and against it the best task is 4.00×.

Does Rapid-MLX change model outputs?

It loads the same weights, and speculative decoding only keeps tokens the full model agrees with. The text is still not always byte-identical to mlx-lm's: with greedy decoding it matched on 11 of 20 tasks on Qwen3.5-9B 4-bit. Where it differed, the two engines typically picked between nearly equally likely tokens, and on the models we re-ran with speculative decoding switched off the outputs differed too, which points to small numerical differences between the two code paths rather than the speedup.

Which models get the speedup?

Models where Rapid-MLX turns speculative decoding on by default, such as Qwen3.5-9B and Qwen3.6-35B-A3B. Qwen3.5-4B has it off by default and was 1.12× faster at the median; with MTP turned on it reached 4.13× on this Mac.