Local embeddings on a Mac with EmbeddingGemma 2
EmbeddingGemma 2 is Google's new embedding model for text and code.
Rapid-MLX runs it on your Mac as an OpenAI-compatible
/v1/embeddings endpoint, on the same server as your chat
model — for RAG, semantic search and code search, with nothing leaving
the machine.
Pick a variant and run it
- Most Macs:
embeddinggemma-2-4bit— a 1.13 GB download - When embedding fidelity matters most:
embeddinggemma-2-bf16— 1.52 GB, not quantized
An embedding model has no chat surface, so you attach it to a chat
server with --embedding-model. Three commands:
# EmbeddingGemma 2 runs on the [vision] runtime $ pip install 'rapid-mlx[vision]' $ rapid-mlx pull embeddinggemma-2-4bit $ rapid-mlx serve qwen3.5-4b-4bit \ --embedding-model embeddinggemma-2-4bit --port 8000
Chat stays on /v1/chat/completions; embeddings are on
/v1/embeddings at the same address. If the
[vision] runtime is missing, serve stops
before loading anything and prints the exact install command for the
Python it runs under. The older 300M aliases use the
[embeddings] extra instead — see
below.
Download
Every alias on this page downloads with one command — the pull buttons in the tables below copy it. Both aliases on this page are mirrored on the rapid-mlx CDN — with automatic mid-pull fallback to Hugging Face if a mirror file slows down. Weights land in the standard Hugging Face cache, and rapid-mlx serve pulls automatically on first use. Live mirror status →
What it is for
- RAG. Embed your documents once, embed each question, and hand the closest passages to the chat model running on the same server.
- Semantic search. Find notes, tickets or docs by meaning rather than by exact words.
- Code search. Use the code-retrieval prefix for the query and the file name as the document title (see task prefixes).
Every vector comes back unit-length with 768 dimensions by default, so a plain dot product is the cosine similarity — any vector store works.
What Rapid-MLX adds
-
A native encoder. 0.15.7 ships its own EmbeddingGemma 2
encoder for text and code, loaded from the MLX weights on the
[vision]runtime —mlx-embeddingsis not needed. -
One server, two models. The embedding model sits next to the
chat model, and
/v1/modelslists it with"capabilities": ["embedding"], so clients that discover models through the OpenAI API find it. -
The OpenAI request shape.
inputtakes one string, a list of strings (a batch) or pre-tokenized ids;dimensionsshortens the vector and re-normalizes it;encoding_formatisfloatorbase64. -
Long inputs are never cut silently. The input window is 8,192
tokens, prefix included. Longer inputs are truncated with a logged
warning and a
/metricscounter, or rejected with a 400 when you start the server with--embedding-overflow-policy error.
Sources: rapid_mlx/embedding.py,
rapid_mlx/models/embedding_gemma2/
and the embeddings guide.
All quants
| Alias | Weights | Download | Recommended for | HF repo | get it |
|---|---|---|---|---|---|
embeddinggemma-2-4bit |
4-bit affine | 1.13 GB | Default — the smaller download | embeddinggemma-2-4bit | CDN |
embeddinggemma-2-bf16 |
BF16 | 1.52 GB | Closest to the original model | embeddinggemma-2-bf16 | CDN |
4-bit or BF16?
The difference is size against fidelity. The conversion checks published on each model card compare the MLX weights with the original checkpoint: BF16 text embeddings match it to a cosine of at least 0.9999; 4-bit to at least 0.98, which the card calls measurable drift. Those are numerical checks on a handful of inputs, not a retrieval benchmark — if ranking quality matters, compare both on your own data.
Pick one per index: vectors from different variants (or from the older 300M model) do not mix, so re-embed your documents when you switch.
Task prefixes and dimensions
The endpoint does not guess whether an input is a query or a document — you add the prefix:
- Search query:
task: search result | query: … - Code query:
task: code retrieval | query: … - Document:
title: none | text: …— put a real title (or the file name for code) in place ofnonewhen you have one
dimensions defaults to 768. 512, 256 and 128 are the
recommended shorter sizes; queries and documents must use the same
one. Asking for more than 768 returns a 400.
Tutorial — search with cosine similarity
With the server from the top of the page running, one request embeds a query and a document:
$ curl http://localhost:8000/v1/embeddings \ -H 'Content-Type: application/json' \ -d '{"model": "embeddinggemma-2-4bit", "input": ["task: search result | query: Which planet is the Red Planet?", "title: none | text: Mars is known as the Red Planet."]}'
The response has one data entry per input, each with a
768-number embedding, plus token usage.
The same thing from Python with the OpenAI client, ranking three
documents for a question and for a code query (save as
search.py, run python search.py):
from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="not-needed")
MODEL = "embeddinggemma-2-4bit"
docs = [
"title: none | text: Mars is known as the Red Planet.",
"title: none | text: Venus is often called Earth's twin.",
"title: utils.py | text: def slugify(s): return s.lower().replace(' ', '-')",
]
queries = [
"task: search result | query: Which planet is the Red Planet?",
"task: code retrieval | query: turn a title into a URL slug",
]
def embed(texts):
resp = client.embeddings.create(model=MODEL, input=texts, dimensions=256)
return [d.embedding for d in resp.data]
# Vectors come back unit-length, so a dot product is the cosine similarity.
def cosine(a, b):
return sum(x * y for x, y in zip(a, b))
doc_vecs = embed(docs)
for query, q in zip(queries, embed(queries)):
print(query)
for score, doc in sorted(((cosine(q, v), d) for d, v in zip(docs, doc_vecs)), reverse=True):
print(f" {score:.3f} {doc}")
Output on an M2 Pro Mac mini with rapid-mlx 0.15.7 — the right document wins each query:
task: search result | query: Which planet is the Red Planet?
0.867 title: none | text: Mars is known as the Red Planet.
0.668 title: none | text: Venus is often called Earth's twin.
0.567 title: utils.py | text: def slugify(s): return s.lower().replace(' ', '-')
task: code retrieval | query: turn a title into a URL slug
0.799 title: utils.py | text: def slugify(s): return s.lower().replace(' ', '-')
0.586 title: none | text: Mars is known as the Red Planet.
0.563 title: none | text: Venus is often called Earth's twin.
Run it beside a chat model
One server (simplest). The serve command at the
top already does this: the chat model and the embedding model share
one process and one port.
A second port. To leave a running chat server untouched, start a second server for embeddings on port 8001. It still needs a chat model as its main model, so give it the smallest one:
$ rapid-mlx serve qwen3-0.6b-4bit \
--embedding-model embeddinggemma-2-4bit --port 8001
Then point your embedding client at
http://localhost:8001/v1. The extra host model is a
0.35 GB download.
EmbeddingGemma 2 and the older 300M model
embeddinggemma-300m-6bit and
-8bit are still in the catalog. They are a different,
smaller model (2,048-token window) and run on the
[embeddings] extra
(pip install 'rapid-mlx[embeddings]') instead of
[vision]. For a new index, start with EmbeddingGemma 2;
an existing 300M index keeps working, but its vectors cannot be
compared with EmbeddingGemma 2 vectors.
Known limitations
- Text and code only. The upstream model also embeds images, audio and video; this release does not serve those inputs, and a request with media gets a 400.
-
Not a main model.
rapid-mlx serve embeddinggemma-2-4biton its own exits with a hint to use--embedding-model. - One embedding model per server. A request naming another embedding model gets a 400 that names the loaded one; switching means restarting the server.