Model pages · rapid-mlx 0.12.11
Local models, measured
The most-pulled models on rapid-mlx, each with its own page: one command to serve it, and real numbers — decode speed, time to first token, peak memory, cold boot — all measured on the same Mac mini M2 Pro (32 GB), the kind of machine most people actually have. Sorted by decode speed.
| Model | Architecture | Decode | First token | Peak memory | Weights |
|---|---|---|---|---|---|
| Qwen3 0.6B | 0.6 B dense · 4-bit | 225.9 tok/s | 0.17 s | 1.2 GB | 0.3 GB |
| LFM2.5 1.2B | 1.2 B hybrid conv+attention · 4-bit | 209.8 tok/s | 0.2 s | 1.1 GB | 0.6 GB |
| Ternary Bonsai 1.7B | 1.7 B ternary · 2-bit | 169.0 tok/s | 0.23 s | 1.3 GB | 0.5 GB |
| LFM2.5 2.6B | 2.6 B hybrid conv+attention · 4-bit | 93.5 tok/s | 0.33 s | 2.2 GB | 1.5 GB |
| Llama 3.2 3B | 3 B dense · 4-bit | 82.2 tok/s | 0.25 s | 2.4 GB | 1.7 GB |
| Gemma 3 4B (QAT) | 4 B dense · QAT 4-bit · vision | 66.7 tok/s | 0.45 s | 4.4 GB | 2.8 GB |
| Qwen3 4B Thinking (2507) | 4 B dense · 4-bit · reasoning | 61.3 tok/s | 0.41 s | 3.0 GB | 2.1 GB |
| Qwen3 4B Instruct (2507) | 4 B dense · 4-bit | 61.2 tok/s | 0.4 s | 2.9 GB | 2.1 GB |
| Qwen3.5 4B | 4 B dense · 4-bit | 60.7 tok/s | 0.56 s | 3.4 GB | 2.9 GB |
| Qwen3.6 35B (A3B MoE) | MoE · 35 B total / 3 B active · 4-bit | 59.6 tok/s | 0.55 s | 9.5 GB | 19.1 GB |
| Qwen3 Coder 30B (A3B MoE) | MoE · 30 B total / 3 B active · 4-bit · code | 53.0 tok/s | 0.54 s | 10.0 GB | 16.0 GB |
| Gemma 4 26B (A4B MoE) | MoE · 26 B total / 4 B active · 4-bit | 50.4 tok/s | 0.68 s | 6.0 GB | 14.3 GB |
| GPT-OSS 20B | MoE · 20 B total · MXFP4-Q8 · reasoning | 47.9 tok/s | 0.57 s | 9.4 GB | 11.3 GB |
| Qwen3 8B | 8 B dense · 4-bit | 37.1 tok/s | 0.62 s | 5.1 GB | 4.3 GB |
| Qwen3.5 9B | 9 B dense · 4-bit | 36.4 tok/s | 0.84 s | 5.8 GB | 5.6 GB |
| Gemma 4 12B | 12 B dense · 4-bit · vision | 22.6 tok/s | 0.95 s | 7.8 GB | 6.3 GB |
| Ternary Bonsai 27B | 27 B ternary · 2-bit | 17.4 tok/s | 1.48 s | 8.4 GB | 7.9 GB |
| Qwen3.6 27B | 27 B dense · 4-bit | 11.1 tok/s | 2.41 s | 8.7 GB | 15.0 GB |
rapid-mlx 0.12.10, measured 2026-08-11. Median of 3 runs, temperature 0, engine-reported token counts, prefix cache defeated with a unique salt per request. Every page carries the full method line.