Best local LLM by Mac · measured

Best local LLMs for your Mac, by chip and memory

One page per Mac configuration that has been measured on enough models: decode speed, first-token time and memory for each, measured on that exact chip and memory size. Nothing is scaled from another Mac.

Pick your Mac

A configuration gets a page once 5 or more models have been measured on it, including at least one run on rapid-mlx 0.14 or newer.

Almost there

These have 3 or more models measured but no page yet: fewer than 5 models, or no run on a current rapid-mlx. If you have one, a run adds to it:

$ rapid-mlx benchmark run qwen3.5-9b-4bit
$ rapid-mlx benchmark share <run-id>

My Mac isn't here

Pick by memory instead: rapid-mlx recipe prints the engine's choice for your Mac, and hardware tiers lists every tier. The best local LLM for your Mac, by memory walks through the same choice from an 8 GB Air up.