DiffusionGemma 26B-A4B-it
Google's discrete-diffusion text generation model — parallel block
decoding instead of one token at a time — served by Rapid-MLX behind
the same OpenAI-compatible API as every autoregressive model, even
though mlx-lm has no diffusion_gemma model
type.
Pick a quant and run it
- 32 GB Mac:
rapid-mlx serve diffusion-gemma-26b-4bit(~15 GB of weights — the recommended quant) - 64 GB Mac or larger:
rapid-mlx serve diffusion-gemma-26b-8bit(~28 GB, highest quality)
The diffusion engine runs on mlx-vlm, so install the
[vision] extra first:
pip install "rapid-mlx[vision]==0.15.7". You get an
OpenAI-compatible endpoint at http://localhost:8000/v1
backed by a non-autoregressive decoder that emits whole blocks of
tokens per refinement step.
Download
Every alias on this page downloads with one command — the pull buttons in the tables below copy it. All 2 aliases on this page are mirrored on the rapid-mlx CDN — with automatic mid-pull fallback to Hugging Face if a mirror file slows down. Weights land in the standard Hugging Face cache, and rapid-mlx serve pulls automatically on first use. Live mirror status →
What "discrete diffusion text gen" means
A standard autoregressive LLM emits text one token at a time — sample a token, append it to the context, run the model again. A discrete-diffusion model instead starts from a fully masked block, runs a fixed number of "denoising" passes, and at each pass replaces the most-confident masked positions with real tokens. Because every position in the block can be updated in the same forward pass, the decoder is naturally parallel.
The practical effect on a Mac: each forward pass produces many tokens of output rather than one, which on Apple Silicon's bandwidth-bound regime amortises memory traffic across the whole block.
Why Rapid-MLX ships it
mlx-lm rejects this checkpoint with
"Model type diffusion_gemma not supported"
(mlx-lm#1391).
Rapid-MLX routes the text-diffusion aliases to a
dedicated diffusion engine built on mlx-vlm's
diffusion generator, and adapts it to the OpenAI
/v1/chat/completions contract — from the client's point
of view it behaves like a standard AR model, except that streaming
delta chunks arrive in larger groups instead of one
token each.
Quants we publish
| alias | hf repo | cache footprint | recommended for | get it |
|---|---|---|---|---|
| diffusion-gemma-26b-4bit | mlx-community/diffusiongemma-26B-A4B-it-4bit | ~15 GB | 32 GB Macs — the recommended quant | CDN |
| diffusion-gemma-26b-8bit | mlx-community/diffusiongemma-26B-A4B-it-8bit | ~28 GB | 64 GB Macs — highest quality | CDN |
Run it
$ pip install "rapid-mlx[vision]==0.15.7" # provides mlx-vlm, which the diffusion engine runs on $ rapid-mlx serve diffusion-gemma-26b-4bit
Streaming chat completion
The deltas come in larger batches than an AR model — every iteration of the diffusion loop reveals a block of tokens — but the wire format is identical SSE chunks.
from openai import OpenAI client = OpenAI(base_url="http://localhost:8000/v1", api_key="not-used") stream = client.chat.completions.create( model="diffusion-gemma-26b-4bit", messages=[ {"role": "user", "content": "Write a haiku about Apple Silicon."} ], stream=True, ) for chunk in stream: delta = chunk.choices[0].delta.content if delta: print(delta, end="", flush=True)
Engineering notes
-
Diffusion engine. The aliases carry
modality: text-diffusion, which routes them to a separate engine wrappingmlx-vlm'sstream_diffusion_generateinstead of the autoregressive scheduler. Requests run one at a time on the GPU; speculative decoding does not apply. -
Sampling knobs. Only
temperatureis taken from the request. The denoising step count (48 per block by default) and the entropy-bound sampler come from the model's own defaults and are not exposed as OpenAI request fields. -
Tool calling. The aliases use the
gemma4tool-call parser, and parsedtool_callsare returned on non-streaming requests. Withstream: truethe tool declarations still reach the model but calls are not parsed out — for streaming agentic workloads use an autoregressive alias (Tmax, Qwen3.5, Gemma 4).
References
- Quant family on Hugging Face: mlx-community/diffusiongemma-26B-A4B-it-4bit
- Upstream mlx-lm support issue: ml-explore/mlx-lm#1391
- Alias config:
rapid_mlx/aliases.json