Best local LLMs for an M1 Max Mac with 64 GB
7 models have been measured on an Apple M1 Max with 64 GB of unified memory. Every number below was measured on that chip and memory size, by the Rapid-MLX community or by us. None is scaled from another Mac.
What Rapid-MLX picks for 64 GB
These two come from the engine's recommendation policy for its 64 GB tier
(rapid/default/text-generation/ram-v1): the same picks rapid-mlx recipe, the
installer and the desktop app show. They are the policy's choice, not a ranking of
the measurements below. How the tiers are chosen.
Measured on this Mac
Decode speed is how fast the model writes once it has read your prompt; first-token time is how long it reads before it starts. Each section has its own scale and method, so compare rows within a section.
Community runs
Current community benchmark: a 512-token prompt, then 128 generated tokens. Median across runs; each run on the contributor's own Mac.
- bonsai-1.7b-2bit217tok/s
- gpt-oss-20b69.0tok/s
- gemma-4-e4b-4bit67.7tok/s
- qwen3.6-35b62.5tok/s
- qwen3.5-9b-4bit45.7tok/s
- qwen3.8-27b-4bit12.1tok/s
Older community benchmark
The previous community suite (short prompt). Its protocol differs from the current one, so compare these rows with each other, not with the sections above.
- qwen3.5-4b-4bit91.9tok/s
Run one
rapid-mlx recipe detects your Mac's memory and prints the pick with its
command. Any model above runs with rapid-mlx serve <alias> and serves
an OpenAI- and Anthropic-compatible API on http://localhost:8000.
$ curl -fsSL https://rapidmlx.com/install.sh | bash $ rapid-mlx recipe $ rapid-mlx serve qwen3.8-27b-4bit
Add your Mac's numbers
Have an M1 Max Mac with 64 GB? Run the benchmark and share it. It lands on the live leaderboard view for this Mac and on this page at its next build.
$ rapid-mlx benchmark run qwen3.8-27b-4bit $ rapid-mlx benchmark share <run-id>