Model aliases
An alias is a short, stable model name — qwen3.5-4b-4bit —
that resolves to a Hugging Face MLX repo and applies that model's
tool-call parser, reasoning parser and performance defaults. Rapid-MLX
ships 197 text aliases plus image, video, embedding and audio
registries. Find one and serve it:
$ rapid-mlx models --search qwen # filter the catalog $ rapid-mlx recipe # the two picks for this Mac $ rapid-mlx info qwen3.5-9b-4bit # what an alias turns on $ rapid-mlx serve qwen3.5-9b-4bit
Not in the catalog? Serve the Hugging Face org/name
directly (Bring your own model), or give it your
own short name with rapid-mlx alias set <name> <org/name>
(alias list / alias remove manage them).
rapid_mlx/aliases.json
in the engine repo (text, image, video and embedding aliases; audio is in
rapid_mlx/audio/aliases.json).
rapid-mlx models prints the same data.
Naming convention
Text aliases follow one template:
<family>-<version>-<params>-<modality?>-<technique?>-<quant>
| Segment | Meaning | Examples |
|---|---|---|
| family | Model family | qwen, gemma, llama, mistral, deepseek, phi, hermes |
| version | Major version | 3.5, 3.6, -4, -r1, -v4-flash |
| params | Parameter count (MoE includes active) | 4b, 12b, 27b, 35b (Qwen3.6 35B-A3B MoE) |
| modality (optional) | Non-text variants | -vl (vision), -coder (code) |
| technique (optional) | Training-time modifier | -qat (QAT), -distill, -thinking, -assistant (MTP assistant drafter) |
| quant | Quantization tier | -4bit, -6bit, -8bit, -mxfp4, -mxfp4-q8, -qat-4bit |
Most aliases end in their quant — qwen3.5-4b-4bit — so the
precision is visible in the name. A few short names
(qwen3.6-35b, gpt-oss-20b,
qwen3-0.6b, tmax-9b) point at the standard
build of that model; rapid-mlx info shows which repo each
resolves to. -qat is a technique, not a quant —
it stacks before the quant, as in gemma-4-12b-qat-4bit.
Registry schema
Each entry in aliases.json is either a bare repo id or an
object with a closed set of keys: an unknown key fails alias loading, so
copy the field set from the closest existing alias when adding one. The
keys most entries use:
| Key | Type | Meaning |
|---|---|---|
hf_path | string | Hugging Face repo id in org/name form. The only mandatory field. |
tool_call_parser | string | Tool-call parser selected at serve time (hermes, qwen3_coder_xml, gemma4, harmony, deepseek, …). |
reasoning_parser | string | Reasoning parser selected at serve time (qwen3, gemma4, deepseek_r1, gpt_oss, glm5, minimax, …). |
is_hybrid | bool | Hybrid architecture (linear attention / Mamba / GatedDeltaNet). Rules out SuffixDecoding and generic speculative paths. |
is_moe | bool | Mixture-of-experts model. |
modality | string | text (default), text-diffusion, embedding, image-gen or video-gen. |
supports_spec_decode | bool | Whether speculative-decoding routes are eligible. |
mtp_draft_model, mtp_default_enabled | string, bool | The MTP sidecar checkpoint, and whether serve turns MTP on by default. |
pflash_tier, turboquant_tier | string | Whether PFlash and K8V4 KV compression are verified for this alias. Verified K8V4 is on by default; PFlash is always opt-in. |
min_memory_gb, experimental | number, bool | Memory floor and experimental status, shown by info and the pickers. |
recommended_sampling | object | Sampling defaults applied when a request doesn't set them. |
The full key list — DFlash / DDTree drafters, suffix tier, vision memory
floor, chat-template id and more — is the
_ALLOWED_PROFILE_KEYS set in
rapid_mlx/model_aliases.py.
What rapid-mlx models prints
The main table lists text and embedding aliases; image, video and audio
models follow in their own tables. --modality picks one
table, --search filters by name, --json gives
stable machine-readable output, and --cached (or
rapid-mlx ls) shows only what is downloaded.
| Column | Meaning |
|---|---|
| Alias | The name you pass to serve / chat / pull. |
| Size | Download size. |
| Tools | Tool-call parser the alias selects (or —). |
| Reasoning | Reasoning parser the alias selects (or —). |
| Spec-Decode | ✓ eligible; ✗ hybrid ruled out by a hybrid architecture; ✗ not eligible. |
| Suffix Tier | SuffixDecoding recommendation: verified, neutral, avoid, unknown, n/a. |
| DFlash / DDTree | verified = curated pair, exp = explicit opt-in, ✗ = incompatible. |
| Preset | The alias's speculative preset: MTP@<drafter>@<k>, MTP@native@<k>, Suffix, or —. |
The model picker shows the same catalog with a RAM filter.
Text families
197 text aliases across these families (image, video, audio and embedding aliases are counted in their own registries):
| Family | # aliases | Notes |
|---|---|---|
| Qwen 3 (3.5 · 3.6 · 3.8 · Coder · Coder-Next 80B · VL · legacy) | 59 | Anchor family — includes MoE 35B-A3B / 122B, the Coder-Next 80B (qwen3-coder-next-80b-4bit / -8bit, HF-pull), hero coder alias, VL variants, plus the 3.8 line — the 27B set (qwen3.8-27b-4bit, -4bit-fp16, -mixed-3.5bpw, -abliterated-4bit, plus the experimental -tensorfold accelerated profile) and the experimental qwen3.8-flash-next-4bit (128 GB minimum, 192 GB recommended), whose experimental qwen3.8-flash-next-tensorfold accelerated profile needs 192 GB. |
| Gemma 4 (dense · e2b / e4b · qat · assistant · optiq · DiffusionGemma) | 29 | Includes the DiffusionGemma aliases. EmbeddingGemma is an embedding alias, listed below. |
| GPT-OSS (20B · 120B · MXFP4-Q4/Q8 · safeguard) | 10 | Harmony-native tool + reasoning parsers. |
| UI-TARS | 9 | ByteDance vision-grounded UI automation agent. Custom ui_tars tool parser. |
| Granite (4.2 30B / 8B / 3B · 4-bit + 8-bit · Granite 4 H-Micro / Tiny) | 8 | Granite 4.2 in 30B / 8B / 3B, 4-bit and 8-bit, plus Granite 4 H-Micro and H-Tiny. |
| Tmax (9B / 27B, 4/6/8-bit + bf16) | 7 | Hybrid Gated-DeltaNet agent model (Ai2). |
| DeepSeek (R1-Distill · V4-Flash 2/4/8bit + 0731 MXFP4 · V4.1-Flash REAP · Coder-V2-lite) | 8 | V4-Flash needs a 192–256 GB Mac; R1-Distill covers the 32B slot. deepseek-v41-flash-reap-2bit is experimental and refuses to start below 224 GB. |
| Mistral (24B · Small-4 119B · Devstral) | 6 | Devstral is the code variant; Small-4 119B is the large MoE SKU. |
| Gemma 3 (1B–27B · qat) | 6 | Kept for continuity; new work targets Gemma 4. |
| Phi 4 (mini + mini-flash 4bit variants) | 4 | |
| Llama (3.1 · 3.2) | 4 | |
| GLM (4.5-Air · 4.7 · 5.2 REAP · 5.3-Flash) | 5 | glm5.3-flash-4bit is a 320B MoE and asks for the 192 GB tier. glm5.3-flash-tensorfold is its experimental 256 GB accelerated profile. |
| Liquid (LFM2 24B-A2B · LFM2.5 8B-A1B / 2.6B / 1.2B) | 4 | |
| Qwopus | 2 | |
| Bonsai (Ternary 1.7B · 8B · 27B · Bonsai 2 27B) | 4 | The ternary image model bonsai-image-4b-2bit is image-generation and counts in that registry, not here. bonsai2-27b-2bit takes image input and loads through Rapid-MLX's own loader for its packed Hadamard projections. |
| MiniCPM (MiniCPM5 1B · 2B · optiq) | 3 | minicpm5-2b-4bit is a compact 1.4 GB option. |
| Muse (Glimmer 30B, 4-bit / 8-bit / bf16) | 3 | The 8-bit pairing is the one qualified for DFlash acceleration. |
| MiniMax (M2.5 / M2.7 MXFP4) | 2 | |
| VibeThinker (1.5B / 3B) | 2 | |
| Holo3 | 2 | |
| Hermes 3 · 4 | 2 | |
| Ornith 1.5 (9B dense · 35B-A3B MoE) | 2 | |
| Hunyuan 3 (Hy3) | 1 | |
| Ling (3.0 Tiny) | 1 | |
| Others (G9v3-39A5B · NeoHorse 1 · Kimi K2.6 · K2 Horizon 7B · MiMo V2.6 Flash · Qwen 2.5 · Nemotron 3 / 3.5-Lightning / Labs-Diffusion · North Mini Code · Nanbeige 4.1 · SmolLM3) | 14 | One or two aliases each; nemotron-labs-diffusion-3b-4bit is a text-diffusion model. k2-horizon-7b-4bit runs on Rapid-MLX's own native runtime and is experimental. nemotron-3.5-lightning-tensorfold is the experimental 48 GB accelerated profile of nemotron-3.5-lightning-30b-4bit. |
Audio registry
Audio aliases live in a separate registry — start them the same way:
rapid-mlx serve whisper-large-v3. The 44 audio aliases:
| Kind | Aliases |
|---|---|
| TTS · kokoro | kokoro, kokoro-82m, kokoro-82m-bf16, kokoro-82m-4bit, kokoro-82m-8bit, kokoro-4bit, kokoro-8bit |
| TTS · vibevoice / voxcpm / chatterbox / dia | vibevoice, vibevoice-realtime, voxcpm, chatterbox, chatterbox-4bit, dia |
| TTS · qwen3-tts | qwen3-tts, qwen3-tts-4bit, qwen3-tts-6bit, qwen3-tts-clone, qwen3-tts-customvoice, qwen3-tts-voicedesign, qwen3-tts-voicedesign-4bit, qwen3-tts-voicedesign-8bit |
| TTS · cloning (indextts / f5) | f5-tts-zh, indextts, indextts-1.5 |
| STT · whisper | whisper, whisper-1, whisper-large-v3, whisper-large-v3-turbo, whisper-medium, whisper-small, whisper-base, whisper-tiny |
| STT · parakeet | parakeet, parakeet-tdt-0.6b, parakeet-tdt-0.6b-v2, parakeet-v3, parakeet-tdt-0.6b-v3 |
| STT · sensevoice | sensevoice, sensevoice-small |
| STT · qwen3-asr | qwen3-asr, qwen3-asr-0.6b, qwen3-asr-1.7b |
| Forced alignment · qwen3-aligner | qwen3-aligner, qwen3-forced-aligner |
Audio requires the [audio] extra —
pip install "rapid-mlx[audio]==0.15.7". See the
extras page
for the full dep list.
Video generation registry
Video aliases live in rapid_mlx/aliases.json with
modality: video-gen and serve through the asynchronous
/v1/videos job API, not /v1/chat/completions.
The 10 video aliases:
| Line | Aliases |
|---|---|
| Wan 2.1 | wan2.1-t2v-1.3b-bf16 |
| Wan 2.2 | wan2.2-i2v-a14b-q8, wan2.2-t2v-a14b-bf16, wan2.2-ti2v-5b-bf16, wan2.2-ti2v-5b-q8 |
| CogVideoX-Fun | cogvideox-fun-5b-bf16, cogvideox-fun-5b-q4, cogvideox-fun-5b-q8 |
| LTX-2.3 | ltx-2.3-mlx-q4 |
| LTX-2.5 | ltx-2.5-mlx-q8 |
Video requires the [video] extra plus ffmpeg and Python 3.11+ —
pip install "rapid-mlx[video]==0.15.7". See the
video family page.
Image generation registry
Image aliases (modality: image-gen) serve through
/v1/images/generations (and /v1/images/edits
where the model edits). They need the [image] extra and
Python 3.11+ — pip install "rapid-mlx[image]==0.15.7".
| Line | Aliases |
|---|---|
| FLUX | flux2-klein-4b, flux2-klein-4b-bf16, flux-schnell |
| Qwen-Image | qwen-image, qwen-image-2.1, qwen-image-2.1-bf16, qwen-image-edit |
| Z-Image | z-image-turbo |
| Others | bonsai-image-4b-2bit, hidream-o1-dev, sdxl-base, sd35-large-4bit |
Embedding aliases
Embedding aliases attach to a chat server with
--embedding-model and answer on /v1/embeddings.
embeddinggemma-2-4bit and embeddinggemma-2-bf16
(EmbeddingGemma 2, text and code) need the [vision] runtime
— start there for a new index; see the
EmbeddingGemma 2 page.
The older embeddinggemma-300m-6bit and
embeddinggemma-300m-8bit need the [embeddings]
extra.
The (unmapped) column
rapid-mlx ls (or models --cached) shows every
Hugging Face repo already on disk. A repo with no matching alias shows
as (unmapped) in the Alias column — serve it by its
full org/name (e.g.
rapid-mlx serve mlx-community/Holo3-35B-A3B-8bit), or name
it with rapid-mlx alias set.