Apple Silicon LLM benchmarks
Measured local-inference performance, contributed by
rapid-mlx users: 35 model-and-chip combinations from 58
submissions across 11 machines (updated 2026-09-27). Explore and
filter on the interactive leaderboard, or jump
straight to a measurement below. Every number is reproducible:
rapid-mlx bench <model> --tier speed --submit.
/api/benchmarks
combinations
Submissions are anonymous and retractable — see the leaderboard for details.
Every number here is reproducible
rapid-mlx bench <model> --tier speed --submit
Bar length is proportional to tok/s on one shared scale across every chip group — the longest bar is the corpus maximum, 329 tok/s on Apple M5 Max.
No model or chip matches
35 model-and-chip combinations exist in this corpus. Clear the search to bring every group back.
Add your machine
Install the engine, run the standardized speed tier, and your Mac joins the corpus above.
1 · Install the engine
curl -fsSL https://rapidmlx.com/install.sh | bash
2 · Run the speed tier and submit
rapid-mlx bench qwen3.5-9b-4bit --tier speed --submit
Raw corpus: /api/benchmarks · Submissions are anonymous and retractable — see the leaderboard for details.