Model pages · rapid-mlx 0.12.11

Local models, measured

The most-pulled models on rapid-mlx, each with its own page: one command to serve it, and real numbers — decode speed, time to first token, peak memory, cold boot — all measured on the same Mac mini M2 Pro (32 GB), the kind of machine most people actually have. Sorted by decode speed.

ModelArchitectureDecodeFirst tokenPeak memoryWeights
Qwen3 0.6B0.6 B dense · 4-bit225.9 tok/s0.17 s1.2 GB0.3 GB
LFM2.5 1.2B1.2 B hybrid conv+attention · 4-bit209.8 tok/s0.2 s1.1 GB0.6 GB
Ternary Bonsai 1.7B1.7 B ternary · 2-bit169.0 tok/s0.23 s1.3 GB0.5 GB
LFM2.5 2.6B2.6 B hybrid conv+attention · 4-bit93.5 tok/s0.33 s2.2 GB1.5 GB
Llama 3.2 3B3 B dense · 4-bit82.2 tok/s0.25 s2.4 GB1.7 GB
Gemma 3 4B (QAT)4 B dense · QAT 4-bit · vision66.7 tok/s0.45 s4.4 GB2.8 GB
Qwen3 4B Thinking (2507)4 B dense · 4-bit · reasoning61.3 tok/s0.41 s3.0 GB2.1 GB
Qwen3 4B Instruct (2507)4 B dense · 4-bit61.2 tok/s0.4 s2.9 GB2.1 GB
Qwen3.5 4B4 B dense · 4-bit60.7 tok/s0.56 s3.4 GB2.9 GB
Qwen3.6 35B (A3B MoE)MoE · 35 B total / 3 B active · 4-bit59.6 tok/s0.55 s9.5 GB19.1 GB
Qwen3 Coder 30B (A3B MoE)MoE · 30 B total / 3 B active · 4-bit · code53.0 tok/s0.54 s10.0 GB16.0 GB
Gemma 4 26B (A4B MoE)MoE · 26 B total / 4 B active · 4-bit50.4 tok/s0.68 s6.0 GB14.3 GB
GPT-OSS 20BMoE · 20 B total · MXFP4-Q8 · reasoning47.9 tok/s0.57 s9.4 GB11.3 GB
Qwen3 8B8 B dense · 4-bit37.1 tok/s0.62 s5.1 GB4.3 GB
Qwen3.5 9B9 B dense · 4-bit36.4 tok/s0.84 s5.8 GB5.6 GB
Gemma 4 12B12 B dense · 4-bit · vision22.6 tok/s0.95 s7.8 GB6.3 GB
Ternary Bonsai 27B27 B ternary · 2-bit17.4 tok/s1.48 s8.4 GB7.9 GB
Qwen3.6 27B27 B dense · 4-bit11.1 tok/s2.41 s8.7 GB15.0 GB

rapid-mlx 0.12.10, measured 2026-08-11. Median of 3 runs, temperature 0, engine-reported token counts, prefix cache defeated with a unique salt per request. Every page carries the full method line.

Where next