Best local LLMs for an M4 Pro Mac with 48 GB

23 models have been measured on an Apple M4 Pro with 48 GB of unified memory. Every number below was measured on that chip and memory size, by the Rapid-MLX community or by us. None is scaled from another Mac.

23
models measured
0.13.1–0.15.5
rapid-mlx versions
2026-10-03
latest run

What Rapid-MLX picks for 48 GB

These two come from the engine's recommendation policy for its 48 GB tier (rapid/default/text-generation/ram-v1): the same picks rapid-mlx recipe, the installer and the desktop app show. They are the policy's choice, not a ranking of the measurements below. How the tiers are chosen.

Most capable that fits
qwen3.8-27b-4bit
Measured here: 24.1 tok/s (rapid-mlx team runs, rapid-mlx 0.15.4).
rapid-mlx serve qwen3.8-27b-4bit
Fastest pick
qwen3.6-35b-4bit
Measured here: 104 tok/s (rapid-mlx team runs, rapid-mlx 0.15.4).
rapid-mlx serve qwen3.6-35b-4bit

Measured on this Mac

Decode speed is how fast the model writes once it has read your prompt; first-token time is how long it reads before it starts. Each section has its own scale and method, so compare rows within a section.

Community runs

Current community benchmark: a 512-token prompt, then 128 generated tokens. Median across runs; each run on the contributor's own Mac.

Rapid-MLX team runs

Our own runs on a Mac mini with this chip and memory, each model on its default rapid-mlx serve settings (which turn on MTP speculative decoding for the models that ship it): decode speed and first-token time at a ~2,000-token prompt, server-reported Metal peak memory.

Older community benchmark

The previous community suite (short prompt). Its protocol differs from the current one, so compare these rows with each other, not with the sections above.

Run one

rapid-mlx recipe detects your Mac's memory and prints the pick with its command. Any model above runs with rapid-mlx serve <alias> and serves an OpenAI- and Anthropic-compatible API on http://localhost:8000.

$ curl -fsSL https://rapidmlx.com/install.sh | bash
$ rapid-mlx recipe
$ rapid-mlx serve qwen3.8-27b-4bit

Add your Mac's numbers

Have an M4 Pro Mac with 48 GB? Run the benchmark and share it. It lands on the live leaderboard view for this Mac and on this page at its next build.

$ rapid-mlx benchmark run qwen3.8-27b-4bit
$ rapid-mlx benchmark share <run-id>

More