Best local LLMs for an M4 Pro Mac with 48 GB
23 models have been measured on an Apple M4 Pro with 48 GB of unified memory. Every number below was measured on that chip and memory size, by the Rapid-MLX community or by us. None is scaled from another Mac.
What Rapid-MLX picks for 48 GB
These two come from the engine's recommendation policy for its 48 GB tier
(rapid/default/text-generation/ram-v1): the same picks rapid-mlx recipe, the
installer and the desktop app show. They are the policy's choice, not a ranking of
the measurements below. How the tiers are chosen.
Measured on this Mac
Decode speed is how fast the model writes once it has read your prompt; first-token time is how long it reads before it starts. Each section has its own scale and method, so compare rows within a section.
Community runs
Current community benchmark: a 512-token prompt, then 128 generated tokens. Median across runs; each run on the contributor's own Mac.
- qwen3.6-35b85.4tok/s
- qwen3-vl-30b-4bit84.9tok/s
- qwen3.5-4b-4bit83.3tok/s
- gemma-4-26b-4bit75.5tok/s
- gpt-oss-20b73.2tok/s
- qwen3.5-9b-4bit49.7tok/s
- gemma-4-12b-4bit31.8tok/s
- qwen3.6-27b15.5tok/s
- qwen3.8-27b-4bit15.4tok/s
Rapid-MLX team runs
Our own runs on a Mac mini with this chip and memory, each model on its default rapid-mlx serve settings (which turn on MTP speculative decoding for the models that ship it): decode speed and first-token time at a ~2,000-token prompt, server-reported Metal peak memory.
- qwen3-0.6b384tok/s
- lfm2.5-1b-4bit314tok/s
- bonsai-1.7b-2bit236tok/s
- llama3-3b-4bit121tok/s
- qwen3.6-35b104tok/s
- qwen3-4b-instruct-2507-4bit93.0tok/s
- qwen3-4b-thinking-2507-4bit93.0tok/s
- qwen3-coder-30b-4bit91.5tok/s
- gemma3-4b-qat-4bit91.3tok/s
- qwen3.5-4b-4bit85.8tok/s
- gemma-4-26b-4bit78.5tok/s
- gpt-oss-20b77.2tok/s
- qwen3.5-9b-4bit63.0tok/s
- qwen3-8b-4bit54.0tok/s
- gemma-4-12b-4bit33.5tok/s
- qwen3.8-27b-4bit24.1tok/s
- qwen3.6-27b23.4tok/s
- bonsai-27b-2bit22.4tok/s
- bonsai2-27b-2bit17.7tok/s
Older community benchmark
The previous community suite (short prompt). Its protocol differs from the current one, so compare these rows with each other, not with the sections above.
- qwen3.6-35b87.9tok/s
- qwen3.5-4b-4bit83.8tok/s
- gemma-4-26b-4bit77.1tok/s
- gemma-4-26b-qat-4bit70.8tok/s
- qwen3.6-35b-optiq-4bit66.2tok/s
- qwen3-8b-4bit52.6tok/s
- qwen3.5-9b-4bit50.1tok/s
- qwen3.8-27b-mixed-3.5bpw17.1tok/s
- qwen3.8-27b-4bit15.5tok/s
Run one
rapid-mlx recipe detects your Mac's memory and prints the pick with its
command. Any model above runs with rapid-mlx serve <alias> and serves
an OpenAI- and Anthropic-compatible API on http://localhost:8000.
$ curl -fsSL https://rapidmlx.com/install.sh | bash $ rapid-mlx recipe $ rapid-mlx serve qwen3.8-27b-4bit
Add your Mac's numbers
Have an M4 Pro Mac with 48 GB? Run the benchmark and share it. It lands on the live leaderboard view for this Mac and on this page at its next build.
$ rapid-mlx benchmark run qwen3.8-27b-4bit $ rapid-mlx benchmark share <run-id>