Models

Model families on Rapid-MLX

Every model rapid-mlx serves, by family. If you just want to know what to run on your Mac, start with the first section.

Pick a model for your Mac

Ask rapid-mlx — it reads your Mac's memory and prints two picks with the exact command:

$ rapid-mlx recipe
Recommended for this 18.0 GB Mac (18 GB tier)

1. Smart — qwen3.5-9b-4bit
   rapid-mlx serve qwen3.5-9b-4bit

2. Fast — qwen3.5-4b-4bit
   rapid-mlx serve qwen3.5-4b-4bit

Then run the line it prints — rapid-mlx serve <alias> downloads the weights on first use and starts an OpenAI-compatible server at http://localhost:8000/v1. The alias selects the right tool-call parser, reasoning parser and cache settings for you.

The catalog holds 197 text, vision and reasoning aliases, 44 audio aliases and 10 video-generation aliases, plus image-generation and embedding models. rapid-mlx models lists them all in your terminal.

Hero deep dives

Six models get their own page because serving them cleanly on MLX takes model-specific work — a custom parser, a separate engine, or special cache handling. Each page starts with which variant fits your Mac and the command to run it.

All families

Each family page groups its version lines under one heading each, with every alias, its parsers and capability notes. Pass any alias to rapid-mlx serve <alias>; the parsers and cache settings are wired up for you. The registry itself is rapid_mlx/aliases.json (audio: rapid_mlx/audio/aliases.json).

Qwen workhorse
Qwen3.8 / 3.6 / 3.5 / Coder / VL / legacy / 2.5 / Qwopus · the default picks from 16 GB up.
Gemma QAT
Gemma 4 / 4-mobile (3n) / 3 / EmbeddingGemma · QAT variants for sharper low-bit quants.
Llama
Llama 3.1 / 3.2 · 1B / 3B / 8B · llama tool envelope.
Muse Glimmer
Meta's 30B dense reasoner · 131K context · native muse ATEM tool envelope · text serving.
Ling
inclusionAI's 7.9B MoE reasoner, 1.3B active · 131K context on an 8 GB Mac.
DeepSeek
V4-Flash / V4.1-Flash + R1 reasoning-distilled + Coder MoE · V4-Flash also has its own page.
GLM custom parser
GLM-5.3-Flash + GLM-5.2 (REAP) + GLM-4.7 + GLM-4.5-Air · custom GLM tool envelopes.
Mistral
Ministral 3B + Mistral 24B + Mistral-Small-4 119B + Devstral V1 / V2 24B.
Phi
Phi-3.5-mini + Phi-4 mini / 14B + mini-reasoning.
Granite hybrid SSM
Granite 4 H-Micro / Tiny (hybrid state-space + attention) + Granite 4.2 3B / 8B / 30B.
GPT-OSS Harmony
20B / 120B · MXFP4 · Harmony-native tools and reasoning streams.
MiniMax
M2.5 / M2.7 · custom tool envelope + reasoning format.
Hermes
Hermes 3 8B + Hermes 4 70B · Nous Research instruct line.
Hunyuan Ultra-only
Hy3 295B MoE preview · 256 GB Mac Studio · hy_v3 tool + reasoning parser.
Liquid on-device
LiquidAI LFM2 / LFM2.5 (24B-A2B / 8B-A1B / 2.6B / 1.2B) · the 8 GB tier picks.
Audio TTS + STT
TTS (Kokoro / Qwen3-TTS / IndexTTS / Chatterbox / VibeVoice / VoxCPM / F5-TTS / Dia) + STT (Whisper / Parakeet / SenseVoice / Qwen3-ASR / Qwen3-ForcedAligner) · OpenAI-compat /v1/audio/*.
Video generation
LTX-2.5 / 2.3 · Wan 2.1 / 2.2 · CogVideoX-Fun · async /v1/videos job API · needs Python 3.11+ and ffmpeg.
Image generation
Z-Image-Turbo · FLUX · Qwen-Image · OpenAI-compatible /v1/images · needs the [image] extra.
Ornith
Ornith 1.5 · 9B dense and 35B-A3B MoE.
Small / curated research
Bonsai / SmolLM3 / Nanbeige / Nemotron — sub-10B research checkpoints we keep on the registry.

Notable small models

Headline-family pages are organized by the big lines (Qwen, Gemma, Llama). But the most interesting work in 2026 is often happening at the small-and-weird end — focused-purpose models, single-shop research checkpoints, frontier-style tricks compressed into 4 GB. The list below is editorial: small models we find worth trying, or that another team has reported working well in production.