Gemma
35 MLX aliases · Google's open-weights family — Gemma 4 / 4-mobile (3n) / 3 / EmbeddingGemma.
Pick one
One command per line of this family, smallest download first. Your Mac needs the download size in free memory plus room for macOS and the context window; the hardware tiers page has the engine's picks for every RAM size. The first run downloads the weights and starts an OpenAI-compatible server on http://localhost:8000/v1.
rapid-mlx serve gemma3-1b-4bitrapid-mlx serve gemma-4-e2b-4bitrapid-mlx serve gemma-4-12b-4bit
The Gemma family on rapid-mlx covers Google's open-weights line end-to-end: the frontier Gemma 4 (12B / 26B / 31B, plus QAT variants for sharper low-bit quants), the mobile-tier Gemma 4 E-series and Gemma 3n, the older but still useful Gemma 3 predecessor, and EmbeddingGemma for local /v1/embeddings. All chat variants use the gemma4 tool-call parser (Gemma 3 uses hermes); EmbeddingGemma routes to the embeddings endpoint, not chat.
- family
- Gemma
- aliases
- 37
- lines
- 4
- install
- rapid-mlx serve <alias>
- OpenAI base URL
- http://localhost:8000/v1
Download
Every alias on this page downloads with one command — the pull buttons in the tables below copy it. All 37 aliases on this page are mirrored on the rapid-mlx CDN — with automatic mid-pull fallback to Hugging Face if a mirror file slows down. Weights land in the standard Hugging Face cache, and rapid-mlx serve pulls automatically on first use. Live mirror status →
Lines in this family
- Gemma 4 (16 aliases)
- Gemma 4 (mobile / 3n) (11 aliases)
- Gemma 3 (6 aliases)
- EmbeddingGemma (4 aliases)
Gemma 4 · 16 aliases
Google's frontier open-weights family — text + vision. rapid-mlx ships the full quant ladder including Google's QAT (quantization-aware training) and assistant-tuned variants for sharper 4/8-bit quality. All variants use the gemma4 tool-call parser and the gemma4 reasoning parser.
parser: gemma4
| alias | hf repo | tool parser | reasoning | flags | context | AA index | get it |
|---|---|---|---|---|---|---|---|
| gemma-4-12b-4bit | mlx-community/gemma-4-12B-it-4bit | gemma4 | gemma4 | spec | 256K | 22.2 | CDN |
| gemma-4-12b-8bit | mlx-community/gemma-4-12B-it-8bit | gemma4 | gemma4 | spec | 256K | 22.2 | CDN |
| gemma-4-12b-assistant | mlx-community/gemma-4-12B-it-assistant-bf16 | gemma4 | gemma4 | — | 128K | 22.2 | CDN |
| gemma-4-12b-optiq-4bit | mlx-community/gemma-4-12B-it-OptiQ-4bit | gemma4 | gemma4 | spec | 128K | 22.2 | CDN |
| gemma-4-12b-qat-4bit | mlx-community/gemma-4-12B-it-qat-4bit | gemma4 | gemma4 | spec | 256K | 22.2 | CDN |
| gemma-4-12b-qat-8bit | mlx-community/gemma-4-12B-it-qat-8bit | gemma4 | gemma4 | spec | 256K | 22.2 | CDN |
| gemma-4-12b-qat-assistant-4bit | mlx-community/gemma-4-12B-it-qat-assistant-4bit | gemma4 | gemma4 | — | 256K | 22.2 | CDN |
| gemma-4-26b-4bit | mlx-community/gemma-4-26b-a4b-it-4bit | gemma4 | gemma4 | moe · spec | 256K | 26.1 | CDN |
| gemma-4-26b-8bit | mlx-community/gemma-4-26B-A4B-it-8bit | gemma4 | gemma4 | moe · spec | 256K | 26.1 | CDN |
| gemma-4-26b-assistant | mlx-community/gemma-4-26B-A4B-it-assistant-bf16 | gemma4 | gemma4 | — | 256K | 26.1 | CDN |
| gemma-4-26b-qat-4bit | mlx-community/gemma-4-26B-A4B-it-qat-4bit | gemma4 | gemma4 | moe · spec | 256K | 26.1 | CDN |
| gemma-4-31b-4bit | mlx-community/gemma-4-31b-it-4bit | gemma4 | gemma4 | spec | 256K | 29.7 | CDN |
| gemma-4-31b-8bit | mlx-community/gemma-4-31b-it-8bit | gemma4 | gemma4 | spec | 256K | 29.7 | CDN |
| gemma-4-31b-assistant | mlx-community/gemma-4-31B-it-assistant-bf16 | gemma4 | gemma4 | — | 256K | 29.7 | CDN |
| gemma-4-31b-qat-4bit | mlx-community/gemma-4-31B-it-qat-4bit | gemma4 | gemma4 | spec | 256K | 29.7 | CDN |
| gemma-4-31b-qat-8bit | mlx-community/gemma-4-31B-it-qat-8bit | gemma4 | gemma4 | spec | 256K | 29.7 | CDN |
Notes & caveats
- QAT variants are noticeably sharper at low bit-widths — prefer
-qat-4bitover plain-4bitwhen both exist. - 31B is the recommended ceiling for a 64 GB Mac; 12B fits comfortably in 16 GB.
Gemma 4 (mobile / 3n) · 11 aliases
The mobile-tier Gemma line — Gemma 4 E-series (effective 2B / 4B with weight-sharing) and the older Gemma 3n family. These are the smallest production-grade Gemmas, designed to fit phones and base-RAM Macs.
parser: gemma4
| alias | hf repo | tool parser | reasoning | flags | context | AA index | get it |
|---|---|---|---|---|---|---|---|
| gemma-3n-e2b-4bit | mlx-community/gemma-3n-E2B-it-4bit | — | — | spec | 32K | 1 | CDN |
| gemma-3n-e4b-4bit | lmstudio-community/gemma-3n-E4B-it-MLX-4bit | — | — | spec | 32K | 1 | CDN |
| gemma-4-e2b-4bit | mlx-community/gemma-4-e2b-it-4bit | gemma4 | gemma4 | spec | 128K | 9.5 | CDN |
| gemma-4-e2b-6bit | mlx-community/gemma-4-e2b-it-6bit | gemma4 | gemma4 | spec | 128K | 9.5 | CDN |
| gemma-4-e2b-8bit | mlx-community/gemma-4-e2b-it-8bit | gemma4 | gemma4 | spec | 128K | 9.5 | CDN |
| gemma-4-e2b-assistant | mlx-community/gemma-4-e2b-it-assistant-bf16 | gemma4 | gemma4 | — | 128K | 9.5 | CDN |
| gemma-4-e4b-4bit | mlx-community/gemma-4-e4b-it-4bit | gemma4 | gemma4 | spec | 128K | 12.2 | CDN |
| gemma-4-e4b-6bit | mlx-community/gemma-4-e4b-it-6bit | gemma4 | gemma4 | spec | 128K | 12.2 | CDN |
| gemma-4-e4b-8bit | mlx-community/gemma-4-e4b-it-8bit | gemma4 | gemma4 | spec | 128K | 12.2 | CDN |
| gemma-4-e4b-assistant | mlx-community/gemma-4-e4b-it-assistant-bf16 | gemma4 | gemma4 | — | 128K | 12.2 | CDN |
| gemma-4-e4b-optiq-4bit | mlx-community/gemma-4-e4b-it-OptiQ-4bit | gemma4 | gemma4 | spec | 128K | 12.2 | CDN |
Notes & caveats
- Gemma 3n entries have no tool-call parser configured — they are text-only.
Gemma 3 · 6 aliases
Kept in the registry for users who want the older, more battle-tested generation. All variants share the hermes tool envelope; reasoning is plain (no <think> stream).
parser: hermes
| alias | hf repo | tool parser | reasoning | flags | context | AA index | get it |
|---|---|---|---|---|---|---|---|
| gemma3-12b-4bit | mlx-community/gemma-3-12b-it-qat-4bit | hermes | — | spec | 128K | 5.5 | CDN |
| gemma3-1b-4bit | mlx-community/gemma-3-1b-it-4bit | hermes | — | spec | 32K | 1 | CDN |
| gemma3-1b-qat-4bit | mlx-community/gemma-3-1b-it-qat-4bit | hermes | — | spec | 32K | 1 | CDN |
| gemma3-27b-4bit | mlx-community/gemma-3-27b-it-4bit | hermes | — | spec | — | 7.4 | CDN |
| gemma3-27b-qat-4bit | mlx-community/gemma-3-27b-it-qat-4bit | hermes | — | spec | 128K | 7.4 | CDN |
| gemma3-4b-qat-4bit | mlx-community/gemma-3-4b-it-qat-4bit | hermes | — | spec | 128K | 1 | CDN |
Notes & caveats
- QAT variants ship for the 1B / 4B / 12B / 27B tiers — same quant-quality benefit as Gemma 4.
- If you can pick Gemma 4 instead, you should — it's strictly better on the standard benchmarks.
EmbeddingGemma · 4 aliases
Google's embedding models, served on a local /v1/embeddings endpoint next to your chat model — RAG, semantic search and code search without an API call. Start with EmbeddingGemma 2 (embeddinggemma-2-4bit / -bf16, text and code, 8,192-token input); the older 300M model stays for existing indexes. Setup, task prefixes and a worked search example: EmbeddingGemma 2 hero page.
parser: —
| alias | hf repo | tool parser | reasoning | flags | context | AA index | get it |
|---|---|---|---|---|---|---|---|
| embeddinggemma-2-4bit | mlx-community/embeddinggemma-2-4bit | — | — | embedding | 8K | — | CDN |
| embeddinggemma-2-bf16 | mlx-community/embeddinggemma-2-bf16 | — | — | embedding | 8K | — | CDN |
| embeddinggemma-300m-6bit | mlx-community/embeddinggemma-300m-6bit | — | — | embedding | 2K | — | CDN |
| embeddinggemma-300m-8bit | mlx-community/embeddinggemma-300m-8bit | — | — | embedding | 2K | — | CDN |
Notes & caveats
- Attach one to a chat server with
rapid-mlx serve <chat-alias> --embedding-model embeddinggemma-2-4bit; it answers on/v1/embeddings, not/v1/chat/completions. - EmbeddingGemma 2 needs the
[vision]runtime; the 300M aliases need the[embeddings]extra. Vectors from the two models do not mix — re-embed when you switch.
Context is read from the config.json of the exact build each alias pulls; for an embedding model it is the most input tokens the engine embeds, which can be less than the config declares. AA index is the Artificial Analysis Intelligence Index for the base model at full precision with reasoning on — a property of the model, not a score for our quantised build.
Notes & caveats
- Gemma 4 + 4-mobile use the
gemma4tool-call parser; Gemma 3 useshermes; EmbeddingGemma has no tool parser (it's an embeddings model). - QAT (quantization-aware trained) variants exist for Gemma 4 and Gemma 3 — prefer them at 4/8-bit when both exist.