Hero model · rapid-mlx 0.15.7

DiffusionGemma 26B-A4B-it

Google's discrete-diffusion text generation model — parallel block decoding instead of one token at a time — served by Rapid-MLX behind the same OpenAI-compatible API as every autoregressive model, even though mlx-lm has no diffusion_gemma model type.

Pick a quant and run it

The diffusion engine runs on mlx-vlm, so install the [vision] extra first: pip install "rapid-mlx[vision]==0.15.7". You get an OpenAI-compatible endpoint at http://localhost:8000/v1 backed by a non-autoregressive decoder that emits whole blocks of tokens per refinement step.

Download

Every alias on this page downloads with one command — the pull buttons in the tables below copy it. All 2 aliases on this page are mirrored on the rapid-mlx CDN — with automatic mid-pull fallback to Hugging Face if a mirror file slows down. Weights land in the standard Hugging Face cache, and rapid-mlx serve pulls automatically on first use. Live mirror status →

What "discrete diffusion text gen" means

A standard autoregressive LLM emits text one token at a time — sample a token, append it to the context, run the model again. A discrete-diffusion model instead starts from a fully masked block, runs a fixed number of "denoising" passes, and at each pass replaces the most-confident masked positions with real tokens. Because every position in the block can be updated in the same forward pass, the decoder is naturally parallel.

The practical effect on a Mac: each forward pass produces many tokens of output rather than one, which on Apple Silicon's bandwidth-bound regime amortises memory traffic across the whole block.

Why Rapid-MLX ships it

mlx-lm rejects this checkpoint with "Model type diffusion_gemma not supported" (mlx-lm#1391). Rapid-MLX routes the text-diffusion aliases to a dedicated diffusion engine built on mlx-vlm's diffusion generator, and adapts it to the OpenAI /v1/chat/completions contract — from the client's point of view it behaves like a standard AR model, except that streaming delta chunks arrive in larger groups instead of one token each.

Honest caveat. We have not benchmarked this against every other MLX-flavoured serving stack head-to-head, so we make no speed claim. What we commit to: OpenAI-shape responses and an alias that just works. If you have a reproducible comparison against another local runtime, please send it our way for the community performance page.

Quants we publish

aliashf repocache footprintrecommended forget it
diffusion-gemma-26b-4bit mlx-community/diffusiongemma-26B-A4B-it-4bit ~15 GB 32 GB Macs — the recommended quant CDN
diffusion-gemma-26b-8bit mlx-community/diffusiongemma-26B-A4B-it-8bit ~28 GB 64 GB Macs — highest quality CDN

Run it

$ pip install "rapid-mlx[vision]==0.15.7"   # provides mlx-vlm, which the diffusion engine runs on
$ rapid-mlx serve diffusion-gemma-26b-4bit

Streaming chat completion

The deltas come in larger batches than an AR model — every iteration of the diffusion loop reveals a block of tokens — but the wire format is identical SSE chunks.

from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="not-used")

stream = client.chat.completions.create(
    model="diffusion-gemma-26b-4bit",
    messages=[
        {"role": "user", "content": "Write a haiku about Apple Silicon."}
    ],
    stream=True,
)
for chunk in stream:
    delta = chunk.choices[0].delta.content
    if delta:
        print(delta, end="", flush=True)

Engineering notes

References