Best local LLMs for an M2 Pro Mac with 32 GB

34 models have been measured on an Apple M2 Pro with 32 GB of unified memory. Every number below was measured on that chip and memory size, by the Rapid-MLX community or by us. None is scaled from another Mac.

34
models measured
0.11.9–0.14.3
rapid-mlx versions
2026-09-19
latest run

What Rapid-MLX picks for 32 GB

These two come from the engine's recommendation policy for its 32 GB tier (rapid/default/text-generation/ram-v1): the same picks rapid-mlx recipe, the installer and the desktop app show. They are the policy's choice, not a ranking of the measurements below. How the tiers are chosen.

Most capable that fits
qwen3.8-27b-4bit
Measured here: 9.9 tok/s (community runs, rapid-mlx 0.14.3).
rapid-mlx serve qwen3.8-27b-4bit
Fastest pick
Measured here: 62.4 tok/s (community runs, rapid-mlx 0.14.2).
rapid-mlx serve qwen3.5-4b-4bit

Measured on this Mac

Decode speed is how fast the model writes once it has read your prompt; first-token time is how long it reads before it starts. Each section has its own scale and method, so compare rows within a section.

Community runs

Current community benchmark: a 512-token prompt, then 128 generated tokens. Median across runs; each run on the contributor's own Mac.

Older community benchmark

The previous community suite (short prompt). Its protocol differs from the current one, so compare these rows with each other, not with the sections above.

Run one

rapid-mlx recipe detects your Mac's memory and prints the pick with its command. Any model above runs with rapid-mlx serve <alias> and serves an OpenAI- and Anthropic-compatible API on http://localhost:8000.

$ curl -fsSL https://rapidmlx.com/install.sh | bash
$ rapid-mlx recipe
$ rapid-mlx serve qwen3.8-27b-4bit

Add your Mac's numbers

Have an M2 Pro Mac with 32 GB? Run the benchmark and share it. It lands on the live leaderboard view for this Mac and on this page at its next build.

$ rapid-mlx benchmark run qwen3.8-27b-4bit
$ rapid-mlx benchmark share <run-id>

More