Best local LLM by Mac · measured
Best local LLMs for your Mac, by chip and memory
One page per Mac configuration that has been measured on enough models: decode speed, first-token time and memory for each, measured on that exact chip and memory size. Nothing is scaled from another Mac.
Pick your Mac
M2 Pro, 32 GB
34 models measured · latest run 2026-09-19
M4 Pro, 48 GB
23 models measured · latest run 2026-10-03
M1 Max, 64 GB
7 models measured · latest run 2026-10-04
M1 Pro, 32 GB
7 models measured · latest run 2026-10-03
M3 Ultra, 256 GB
5 models measured · latest run 2026-09-22 · see MLX models benchmarked on M3 Ultra
A configuration gets a page once 5 or more models have been measured on it, including at least one run on rapid-mlx 0.14 or newer.
Almost there
These have 3 or more models measured but no page yet: fewer than 5 models, or no run on a current rapid-mlx. If you have one, a run adds to it:
$ rapid-mlx benchmark run qwen3.5-9b-4bit $ rapid-mlx benchmark share <run-id>
- M5 Max, 36 GB10 models measured
- M4 Max, 64 GB5 models measured
- M2, 24 GB4 models measured
- M1 Ultra, 128 GB3 models measured
- M3 Pro, 18 GB3 models measured
- M3 Pro, 36 GB3 models measured
- M5 Max, 64 GB3 models measured
- M5 Pro, 64 GB3 models measured
My Mac isn't here
Pick by memory instead: rapid-mlx recipe prints the engine's choice for
your Mac, and hardware tiers lists every tier.
The best local LLM for your Mac, by
memory walks through the same choice from an 8 GB Air up.