Hero model · rapid-mlx 0.15.7

Local embeddings on a Mac with EmbeddingGemma 2

EmbeddingGemma 2 is Google's new embedding model for text and code. Rapid-MLX runs it on your Mac as an OpenAI-compatible /v1/embeddings endpoint, on the same server as your chat model — for RAG, semantic search and code search, with nothing leaving the machine.

Pick a variant and run it

An embedding model has no chat surface, so you attach it to a chat server with --embedding-model. Three commands:

# EmbeddingGemma 2 runs on the [vision] runtime
$ pip install 'rapid-mlx[vision]'
$ rapid-mlx pull embeddinggemma-2-4bit
$ rapid-mlx serve qwen3.5-4b-4bit \
    --embedding-model embeddinggemma-2-4bit --port 8000

Chat stays on /v1/chat/completions; embeddings are on /v1/embeddings at the same address. If the [vision] runtime is missing, serve stops before loading anything and prints the exact install command for the Python it runs under. The older 300M aliases use the [embeddings] extra instead — see below.

Download

Every alias on this page downloads with one command — the pull buttons in the tables below copy it. Both aliases on this page are mirrored on the rapid-mlx CDN — with automatic mid-pull fallback to Hugging Face if a mirror file slows down. Weights land in the standard Hugging Face cache, and rapid-mlx serve pulls automatically on first use. Live mirror status →

What it is for

Every vector comes back unit-length with 768 dimensions by default, so a plain dot product is the cosine similarity — any vector store works.

What Rapid-MLX adds

Sources: rapid_mlx/embedding.py, rapid_mlx/models/embedding_gemma2/ and the embeddings guide.

All quants

Alias Weights Download Recommended for HF repo get it
embeddinggemma-2-4bit 4-bit affine 1.13 GB Default — the smaller download embeddinggemma-2-4bit CDN
embeddinggemma-2-bf16 BF16 1.52 GB Closest to the original model embeddinggemma-2-bf16 CDN

4-bit or BF16?

The difference is size against fidelity. The conversion checks published on each model card compare the MLX weights with the original checkpoint: BF16 text embeddings match it to a cosine of at least 0.9999; 4-bit to at least 0.98, which the card calls measurable drift. Those are numerical checks on a handful of inputs, not a retrieval benchmark — if ranking quality matters, compare both on your own data.

Pick one per index: vectors from different variants (or from the older 300M model) do not mix, so re-embed your documents when you switch.

Task prefixes and dimensions

The endpoint does not guess whether an input is a query or a document — you add the prefix:

dimensions defaults to 768. 512, 256 and 128 are the recommended shorter sizes; queries and documents must use the same one. Asking for more than 768 returns a 400.

Tutorial — search with cosine similarity

With the server from the top of the page running, one request embeds a query and a document:

$ curl http://localhost:8000/v1/embeddings \
    -H 'Content-Type: application/json' \
    -d '{"model": "embeddinggemma-2-4bit",
         "input": ["task: search result | query: Which planet is the Red Planet?",
                   "title: none | text: Mars is known as the Red Planet."]}'

The response has one data entry per input, each with a 768-number embedding, plus token usage. The same thing from Python with the OpenAI client, ranking three documents for a question and for a code query (save as search.py, run python search.py):

from openai import OpenAI

client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="not-needed")
MODEL = "embeddinggemma-2-4bit"

docs = [
    "title: none | text: Mars is known as the Red Planet.",
    "title: none | text: Venus is often called Earth's twin.",
    "title: utils.py | text: def slugify(s): return s.lower().replace(' ', '-')",
]
queries = [
    "task: search result | query: Which planet is the Red Planet?",
    "task: code retrieval | query: turn a title into a URL slug",
]

def embed(texts):
    resp = client.embeddings.create(model=MODEL, input=texts, dimensions=256)
    return [d.embedding for d in resp.data]

# Vectors come back unit-length, so a dot product is the cosine similarity.
def cosine(a, b):
    return sum(x * y for x, y in zip(a, b))

doc_vecs = embed(docs)
for query, q in zip(queries, embed(queries)):
    print(query)
    for score, doc in sorted(((cosine(q, v), d) for d, v in zip(docs, doc_vecs)), reverse=True):
        print(f"  {score:.3f}  {doc}")

Output on an M2 Pro Mac mini with rapid-mlx 0.15.7 — the right document wins each query:

task: search result | query: Which planet is the Red Planet?
  0.867  title: none | text: Mars is known as the Red Planet.
  0.668  title: none | text: Venus is often called Earth's twin.
  0.567  title: utils.py | text: def slugify(s): return s.lower().replace(' ', '-')
task: code retrieval | query: turn a title into a URL slug
  0.799  title: utils.py | text: def slugify(s): return s.lower().replace(' ', '-')
  0.586  title: none | text: Mars is known as the Red Planet.
  0.563  title: none | text: Venus is often called Earth's twin.

Run it beside a chat model

One server (simplest). The serve command at the top already does this: the chat model and the embedding model share one process and one port.

A second port. To leave a running chat server untouched, start a second server for embeddings on port 8001. It still needs a chat model as its main model, so give it the smallest one:

$ rapid-mlx serve qwen3-0.6b-4bit \
    --embedding-model embeddinggemma-2-4bit --port 8001

Then point your embedding client at http://localhost:8001/v1. The extra host model is a 0.35 GB download.

EmbeddingGemma 2 and the older 300M model

embeddinggemma-300m-6bit and -8bit are still in the catalog. They are a different, smaller model (2,048-token window) and run on the [embeddings] extra (pip install 'rapid-mlx[embeddings]') instead of [vision]. For a new index, start with EmbeddingGemma 2; an existing 300M index keeps working, but its vectors cannot be compared with EmbeddingGemma 2 vectors.

Known limitations

Related