Reference · rapid-mlx 0.15.7 · ← Back to README

Model aliases

An alias is a short, stable model name — qwen3.5-4b-4bit — that resolves to a Hugging Face MLX repo and applies that model's tool-call parser, reasoning parser and performance defaults. Rapid-MLX ships 197 text aliases plus image, video, embedding and audio registries. Find one and serve it:

$ rapid-mlx models --search qwen         # filter the catalog
$ rapid-mlx recipe                       # the two picks for this Mac
$ rapid-mlx info qwen3.5-9b-4bit         # what an alias turns on
$ rapid-mlx serve qwen3.5-9b-4bit

Not in the catalog? Serve the Hugging Face org/name directly (Bring your own model), or give it your own short name with rapid-mlx alias set <name> <org/name> (alias list / alias remove manage them).

Source of truth. The registry lives at rapid_mlx/aliases.json in the engine repo (text, image, video and embedding aliases; audio is in rapid_mlx/audio/aliases.json). rapid-mlx models prints the same data.

Naming convention

Text aliases follow one template:

<family>-<version>-<params>-<modality?>-<technique?>-<quant>
SegmentMeaningExamples
familyModel familyqwen, gemma, llama, mistral, deepseek, phi, hermes
versionMajor version3.5, 3.6, -4, -r1, -v4-flash
paramsParameter count (MoE includes active)4b, 12b, 27b, 35b (Qwen3.6 35B-A3B MoE)
modality (optional)Non-text variants-vl (vision), -coder (code)
technique (optional)Training-time modifier-qat (QAT), -distill, -thinking, -assistant (MTP assistant drafter)
quantQuantization tier-4bit, -6bit, -8bit, -mxfp4, -mxfp4-q8, -qat-4bit

Most aliases end in their quant — qwen3.5-4b-4bit — so the precision is visible in the name. A few short names (qwen3.6-35b, gpt-oss-20b, qwen3-0.6b, tmax-9b) point at the standard build of that model; rapid-mlx info shows which repo each resolves to. -qat is a technique, not a quant — it stacks before the quant, as in gemma-4-12b-qat-4bit.

Registry schema

Each entry in aliases.json is either a bare repo id or an object with a closed set of keys: an unknown key fails alias loading, so copy the field set from the closest existing alias when adding one. The keys most entries use:

KeyTypeMeaning
hf_pathstringHugging Face repo id in org/name form. The only mandatory field.
tool_call_parserstringTool-call parser selected at serve time (hermes, qwen3_coder_xml, gemma4, harmony, deepseek, …).
reasoning_parserstringReasoning parser selected at serve time (qwen3, gemma4, deepseek_r1, gpt_oss, glm5, minimax, …).
is_hybridboolHybrid architecture (linear attention / Mamba / GatedDeltaNet). Rules out SuffixDecoding and generic speculative paths.
is_moeboolMixture-of-experts model.
modalitystringtext (default), text-diffusion, embedding, image-gen or video-gen.
supports_spec_decodeboolWhether speculative-decoding routes are eligible.
mtp_draft_model, mtp_default_enabledstring, boolThe MTP sidecar checkpoint, and whether serve turns MTP on by default.
pflash_tier, turboquant_tierstringWhether PFlash and K8V4 KV compression are verified for this alias. Verified K8V4 is on by default; PFlash is always opt-in.
min_memory_gb, experimentalnumber, boolMemory floor and experimental status, shown by info and the pickers.
recommended_samplingobjectSampling defaults applied when a request doesn't set them.

The full key list — DFlash / DDTree drafters, suffix tier, vision memory floor, chat-template id and more — is the _ALLOWED_PROFILE_KEYS set in rapid_mlx/model_aliases.py.

What rapid-mlx models prints

The main table lists text and embedding aliases; image, video and audio models follow in their own tables. --modality picks one table, --search filters by name, --json gives stable machine-readable output, and --cached (or rapid-mlx ls) shows only what is downloaded.

ColumnMeaning
AliasThe name you pass to serve / chat / pull.
SizeDownload size.
ToolsTool-call parser the alias selects (or —).
ReasoningReasoning parser the alias selects (or —).
Spec-Decode✓ eligible; ✗ hybrid ruled out by a hybrid architecture; ✗ not eligible.
Suffix TierSuffixDecoding recommendation: verified, neutral, avoid, unknown, n/a.
DFlash / DDTreeverified = curated pair, exp = explicit opt-in, ✗ = incompatible.
PresetThe alias's speculative preset: MTP@<drafter>@<k>, MTP@native@<k>, Suffix, or —.

The model picker shows the same catalog with a RAM filter.

Text families

197 text aliases across these families (image, video, audio and embedding aliases are counted in their own registries):

Family# aliasesNotes
Qwen 3 (3.5 · 3.6 · 3.8 · Coder · Coder-Next 80B · VL · legacy)59Anchor family — includes MoE 35B-A3B / 122B, the Coder-Next 80B (qwen3-coder-next-80b-4bit / -8bit, HF-pull), hero coder alias, VL variants, plus the 3.8 line — the 27B set (qwen3.8-27b-4bit, -4bit-fp16, -mixed-3.5bpw, -abliterated-4bit, plus the experimental -tensorfold accelerated profile) and the experimental qwen3.8-flash-next-4bit (128 GB minimum, 192 GB recommended), whose experimental qwen3.8-flash-next-tensorfold accelerated profile needs 192 GB.
Gemma 4 (dense · e2b / e4b · qat · assistant · optiq · DiffusionGemma)29Includes the DiffusionGemma aliases. EmbeddingGemma is an embedding alias, listed below.
GPT-OSS (20B · 120B · MXFP4-Q4/Q8 · safeguard)10Harmony-native tool + reasoning parsers.
UI-TARS9ByteDance vision-grounded UI automation agent. Custom ui_tars tool parser.
Granite (4.2 30B / 8B / 3B · 4-bit + 8-bit · Granite 4 H-Micro / Tiny)8Granite 4.2 in 30B / 8B / 3B, 4-bit and 8-bit, plus Granite 4 H-Micro and H-Tiny.
Tmax (9B / 27B, 4/6/8-bit + bf16)7Hybrid Gated-DeltaNet agent model (Ai2).
DeepSeek (R1-Distill · V4-Flash 2/4/8bit + 0731 MXFP4 · V4.1-Flash REAP · Coder-V2-lite)8V4-Flash needs a 192–256 GB Mac; R1-Distill covers the 32B slot. deepseek-v41-flash-reap-2bit is experimental and refuses to start below 224 GB.
Mistral (24B · Small-4 119B · Devstral)6Devstral is the code variant; Small-4 119B is the large MoE SKU.
Gemma 3 (1B–27B · qat)6Kept for continuity; new work targets Gemma 4.
Phi 4 (mini + mini-flash 4bit variants)4
Llama (3.1 · 3.2)4
GLM (4.5-Air · 4.7 · 5.2 REAP · 5.3-Flash)5glm5.3-flash-4bit is a 320B MoE and asks for the 192 GB tier. glm5.3-flash-tensorfold is its experimental 256 GB accelerated profile.
Liquid (LFM2 24B-A2B · LFM2.5 8B-A1B / 2.6B / 1.2B)4
Qwopus2
Bonsai (Ternary 1.7B · 8B · 27B · Bonsai 2 27B)4The ternary image model bonsai-image-4b-2bit is image-generation and counts in that registry, not here. bonsai2-27b-2bit takes image input and loads through Rapid-MLX's own loader for its packed Hadamard projections.
MiniCPM (MiniCPM5 1B · 2B · optiq)3minicpm5-2b-4bit is a compact 1.4 GB option.
Muse (Glimmer 30B, 4-bit / 8-bit / bf16)3The 8-bit pairing is the one qualified for DFlash acceleration.
MiniMax (M2.5 / M2.7 MXFP4)2
VibeThinker (1.5B / 3B)2
Holo32
Hermes 3 · 42
Ornith 1.5 (9B dense · 35B-A3B MoE)2
Hunyuan 3 (Hy3)1
Ling (3.0 Tiny)1
Others (G9v3-39A5B · NeoHorse 1 · Kimi K2.6 · K2 Horizon 7B · MiMo V2.6 Flash · Qwen 2.5 · Nemotron 3 / 3.5-Lightning / Labs-Diffusion · North Mini Code · Nanbeige 4.1 · SmolLM3)14One or two aliases each; nemotron-labs-diffusion-3b-4bit is a text-diffusion model. k2-horizon-7b-4bit runs on Rapid-MLX's own native runtime and is experimental. nemotron-3.5-lightning-tensorfold is the experimental 48 GB accelerated profile of nemotron-3.5-lightning-30b-4bit.

Audio registry

Audio aliases live in a separate registry — start them the same way: rapid-mlx serve whisper-large-v3. The 44 audio aliases:

KindAliases
TTS · kokorokokoro, kokoro-82m, kokoro-82m-bf16, kokoro-82m-4bit, kokoro-82m-8bit, kokoro-4bit, kokoro-8bit
TTS · vibevoice / voxcpm / chatterbox / diavibevoice, vibevoice-realtime, voxcpm, chatterbox, chatterbox-4bit, dia
TTS · qwen3-ttsqwen3-tts, qwen3-tts-4bit, qwen3-tts-6bit, qwen3-tts-clone, qwen3-tts-customvoice, qwen3-tts-voicedesign, qwen3-tts-voicedesign-4bit, qwen3-tts-voicedesign-8bit
TTS · cloning (indextts / f5)f5-tts-zh, indextts, indextts-1.5
STT · whisperwhisper, whisper-1, whisper-large-v3, whisper-large-v3-turbo, whisper-medium, whisper-small, whisper-base, whisper-tiny
STT · parakeetparakeet, parakeet-tdt-0.6b, parakeet-tdt-0.6b-v2, parakeet-v3, parakeet-tdt-0.6b-v3
STT · sensevoicesensevoice, sensevoice-small
STT · qwen3-asrqwen3-asr, qwen3-asr-0.6b, qwen3-asr-1.7b
Forced alignment · qwen3-alignerqwen3-aligner, qwen3-forced-aligner

Audio requires the [audio] extra — pip install "rapid-mlx[audio]==0.15.7". See the extras page for the full dep list.

Video generation registry

Video aliases live in rapid_mlx/aliases.json with modality: video-gen and serve through the asynchronous /v1/videos job API, not /v1/chat/completions. The 10 video aliases:

LineAliases
Wan 2.1wan2.1-t2v-1.3b-bf16
Wan 2.2wan2.2-i2v-a14b-q8, wan2.2-t2v-a14b-bf16, wan2.2-ti2v-5b-bf16, wan2.2-ti2v-5b-q8
CogVideoX-Funcogvideox-fun-5b-bf16, cogvideox-fun-5b-q4, cogvideox-fun-5b-q8
LTX-2.3ltx-2.3-mlx-q4
LTX-2.5ltx-2.5-mlx-q8

Video requires the [video] extra plus ffmpeg and Python 3.11+ — pip install "rapid-mlx[video]==0.15.7". See the video family page.

Image generation registry

Image aliases (modality: image-gen) serve through /v1/images/generations (and /v1/images/edits where the model edits). They need the [image] extra and Python 3.11+ — pip install "rapid-mlx[image]==0.15.7".

LineAliases
FLUXflux2-klein-4b, flux2-klein-4b-bf16, flux-schnell
Qwen-Imageqwen-image, qwen-image-2.1, qwen-image-2.1-bf16, qwen-image-edit
Z-Imagez-image-turbo
Othersbonsai-image-4b-2bit, hidream-o1-dev, sdxl-base, sd35-large-4bit

Embedding aliases

Embedding aliases attach to a chat server with --embedding-model and answer on /v1/embeddings. embeddinggemma-2-4bit and embeddinggemma-2-bf16 (EmbeddingGemma 2, text and code) need the [vision] runtime — start there for a new index; see the EmbeddingGemma 2 page. The older embeddinggemma-300m-6bit and embeddinggemma-300m-8bit need the [embeddings] extra.

The (unmapped) column

rapid-mlx ls (or models --cached) shows every Hugging Face repo already on disk. A repo with no matching alias shows as (unmapped) in the Alias column — serve it by its full org/name (e.g. rapid-mlx serve mlx-community/Holo3-35B-A3B-8bit), or name it with rapid-mlx alias set.

Next steps