Best local LLMs for an M2 Pro Mac with 32 GB
34 models have been measured on an Apple M2 Pro with 32 GB of unified memory. Every number below was measured on that chip and memory size, by the Rapid-MLX community or by us. None is scaled from another Mac.
What Rapid-MLX picks for 32 GB
These two come from the engine's recommendation policy for its 32 GB tier
(rapid/default/text-generation/ram-v1): the same picks rapid-mlx recipe, the
installer and the desktop app show. They are the policy's choice, not a ranking of
the measurements below. How the tiers are chosen.
Measured on this Mac
Decode speed is how fast the model writes once it has read your prompt; first-token time is how long it reads before it starts. Each section has its own scale and method, so compare rows within a section.
Community runs
Current community benchmark: a 512-token prompt, then 128 generated tokens. Median across runs; each run on the contributor's own Mac.
- qwen3-0.6b258tok/s
- lfm2.5-1b-4bit210tok/s
- bonsai-1.7b-2bit190tok/s
- vibethinker-1.5b-4bit131tok/s
- lfm2.5-8b-a1b-4bit124tok/s
- lfm2.5-2.6b-4bit93.2tok/s
- llama3-3b-4bit85.6tok/s
- nemotron-3.5-lightning-30b-4bit64.7tok/s
- gemma3-4b-qat-4bit64.3tok/s
- qwen3-4b-thinking-2507-4bit63.2tok/s
- qwen3.6-35b-mxfp463.0tok/s
- qwen3.5-4b-4bit62.4tok/s
- qwen3.6-35b-nvfp461.9tok/s
- qwen3.6-35b61.3tok/s
- qwen3-4b-instruct-2507-4bit59.9tok/s
- bonsai-8b-2bit59.8tok/s
- gemma-4-e4b-4bit52.2tok/s
- gpt-oss-20b51.9tok/s
- qwen3.6-35b-dwq46.6tok/s
- qwen3-8b-4bit37.1tok/s
- deepseek-r1-8b-4bit37.0tok/s
- qwen3.5-9b-4bit36.0tok/s
- k2-horizon-7b-4bit35.8tok/s
- qwen3.5-4b-8bit35.5tok/s
- qwen3.5-9b-6bit24.8tok/s
- gemma-4-12b-4bit23.0tok/s
- qwen3.5-9b-8bit20.1tok/s
- bonsai-27b-2bit19.3tok/s
- devstral-v2-24b-4bit12.6tok/s
- qwen3.6-27b11.4tok/s
- qwen3.5-27b-4bit11.2tok/s
- ornith-1.5-9b-bf1610.9tok/s
- qwen3.8-27b-4bit9.9tok/s
- qwen3.5-27b-6bit7.7tok/s
Older community benchmark
The previous community suite (short prompt). Its protocol differs from the current one, so compare these rows with each other, not with the sections above.
- qwen3.5-4b-4bit61.3tok/s
- gpt-oss-20b48.2tok/s
- qwen3.5-9b-4bit36.0tok/s
Run one
rapid-mlx recipe detects your Mac's memory and prints the pick with its
command. Any model above runs with rapid-mlx serve <alias> and serves
an OpenAI- and Anthropic-compatible API on http://localhost:8000.
$ curl -fsSL https://rapidmlx.com/install.sh | bash $ rapid-mlx recipe $ rapid-mlx serve qwen3.8-27b-4bit
Add your Mac's numbers
Have an M2 Pro Mac with 32 GB? Run the benchmark and share it. It lands on the live leaderboard view for this Mac and on this page at its next build.
$ rapid-mlx benchmark run qwen3.8-27b-4bit $ rapid-mlx benchmark share <run-id>