Run Gemma 3 4B (QAT) on a Mac
Google's quantisation-aware-trained Gemma 3 4B — the QAT checkpoint was trained to be quantised, so it loses less than a post-hoc 4-bit conversion. Takes image inputs too.
4 B dense · QAT 4-bit · vision vision + text on small Macs · Google-ecosystem prompts
One command
$ rapid-mlx serve gemma3-4b-qat-4bit
Weights (2.8 GB) download on first run; you get an OpenAI-compatible
endpoint at http://localhost:8000/v1. No rapid-mlx yet? It's one line —
curl -fsSL https://rapidmlx.com/install.sh | bash — or take the
desktop app.
Measured on real hardware
| Machine | Decode | First token | Peak memory | Cold boot | Weights |
|---|---|---|---|---|---|
| Mac mini M2 Pro · 32 GB | 66.7 tok/s | 0.45 s | 4.4 GB | 10.3 s | 2.8 GB |
Measured on rapid-mlx 0.12.10, 2026-08-11. Decode: median of 3 runs, temperature 0, 256-token saturating generation, engine-reported token counts, unique salt per request (no prefix-cache hits). TTFT: median of 3, short prompt. Boot: process spawn to first completed token. Peak RSS: 0.5 s sampling across the run.
Will it fit your Mac?
Peak resident memory measured 4.4 GB during a 256-token generation. The KV cache grows with context length, so treat that as a floor, not a ceiling. For the conservative install-default placement see the hardware tiers table; to compare against every model your RAM can hold, use the live picker.
Variants & alternatives
gemma-4-12b-4bit— Gemma 4 12B — the current generation, dense
FAQ
How much memory does Gemma 3 4B (QAT) need on a Mac?
Measured peak resident memory was 4.4 GB on rapid-mlx 0.12.10 during a 256-token generation (M2 Pro, 32 GB). The weights are 2.8 GB on disk. Longer contexts grow the KV cache beyond this, so leave headroom.
How fast is Gemma 3 4B (QAT) on Apple Silicon?
We measured 66.7 tokens/sec sustained decode and 0.45 s time-to-first-token on a Mac mini M2 Pro (32 GB), median of 3 runs at temperature 0.
How do I run Gemma 3 4B (QAT) locally?
Install rapid-mlx (curl -fsSL https://rapidmlx.com/install.sh | bash, or brew install rapid-mlx), then: rapid-mlx serve gemma3-4b-qat-4bit — the weights download on first run and you get an OpenAI-compatible endpoint at localhost:8000/v1 that works with Cursor, Claude Code, Aider, and any OpenAI client.