Models · family

Qwen

62 MLX aliases · Alibaba's full open-weights line — Qwen3.8 / 3.6 / 3.5 / Coder / VL / legacy / 2.5 / Qwopus.

Pick one

One command per line of this family, smallest download first. Your Mac needs the download size in free memory plus room for macOS and the context window; the hardware tiers page has the engine's picks for every RAM size. The first run downloads the weights and starts an OpenAI-compatible server on http://localhost:8000/v1.

The Qwen family on rapid-mlx covers every Alibaba open-weights line we support — from the workhorse Qwen3.5 (4B → 122B-A10B MoE) and the post-3.5 refresh Qwen3.6, through the code-tuned Qwen3 Coder and the vision-language Qwen3-VL, down to the smaller legacy Qwen3 / Qwen2.5 baselines and the community Opus-aligned Qwopus. Each alias selects its own tool-call parser (hermes on Qwen3.5, qwen3_coder_xml on Qwen3.6 / 3.8 and Coder) and, on the thinking models, the qwen3 reasoning parser — the tables below list each one — so structured tool calls and chain-of-thought stream out of /v1/chat/completions and /v1/responses with no extra config.

family
Qwen
aliases
62
lines
8
install
rapid-mlx serve <alias>
OpenAI base URL
http://localhost:8000/v1

Download

Every alias on this page downloads with one command — the pull buttons in the tables below copy it. 60 of the 62 aliases on this page are mirrored on the rapid-mlx CDN; the rest pull from Hugging Face directly — with automatic mid-pull fallback to Hugging Face if a mirror file slows down. Weights land in the standard Hugging Face cache, and rapid-mlx serve pulls automatically on first use. Live mirror status →

Lines in this family

Qwen3.8 · 7 aliases

The newest line. The 27B is a hybrid GatedDeltaNet checkpoint that runs text on the default lane and answers questions about images via --mllm. Alongside the standard 4-bit build we publish our own mixed-precision 3.5-bpw quantization (rapid-mlx/Qwen3.8-27B-mixed-3.5bpw-MLX) that spends its bits where they matter — 13.0 GB of weights against the 4-bit build's 15.0 GB.

parser: hermes

aliashf repotool parserreasoningflagscontextAA indexget it
qwen3.8-27b-4bitrapid-mlx/Qwen3.8-27B-4bit-MTP-MLXqwen3_coder_xmlqwen3hybrid256K52CDN
qwen3.8-27b-4bit-fp16rapid-mlx/Qwen3.8-27B-4bit-MTP-fp16-MLXqwen3_coder_xmlqwen3hybrid256K52CDN
qwen3.8-27b-abliterated-4bitwindowsxp811203/Qwen3.8-27B-Abliterated-MLX-oQ4e-mtpqwen3_coder_xmlqwen3hybrid256K—CDN
qwen3.8-27b-mixed-3.5bpwrapid-mlx/Qwen3.8-27B-mixed-3.5bpw-MLXqwen3_coder_xmlqwen3hybrid256K—CDN
qwen3.8-27b-tensorfoldVontra/Qwen3.8-27B-MLX-4bit—qwen3hybrid256K—CDN
qwen3.8-flash-next-4bitrapid-mlx/Qwen3.8-Flash-Next-4bithermesqwen3hybrid · moe256K55.8CDN
qwen3.8-flash-next-tensorfoldTensorFold/Qwen3.8-Flash-Next-MLX-4bit-MTP—qwen3hybrid · moe——HF

Notes & caveats

Qwen3.6 · 19 aliases

Post-3.5 refresh. Same hybrid-cache MoE architecture as 3.5-35B but a tighter post-training mix and an XML-style tool-call envelope, parsed with our qwen3_coder_xml parser. We ship DWQ (data-aware quant) and Unsloth's UD-MLX quants alongside the standard 4/6/8-bit set across both 27B and 35B.

parser: qwen3_coder_xml

aliashf repotool parserreasoningflagscontextAA indexget it
qwen3.6-27bmlx-community/Qwen3.6-27B-4bitqwen3_coder_xmlqwen3—256K37.7CDN
qwen3.6-27b-4bitmlx-community/Qwen3.6-27B-4bitqwen3_coder_xmlqwen3—256K37.7CDN
qwen3.6-27b-6bitlmstudio-community/Qwen3.6-27B-MLX-6bitqwen3_coder_xmlqwen3—256K37.7CDN
qwen3.6-27b-8bitunsloth/Qwen3.6-27B-MLX-8bitqwen3_coder_xmlqwen3—256K37.7CDN
qwen3.6-27b-mtp-4bitmlx-community/Qwen3.6-27B-MTP-4bitqwen3_coder_xmlqwen3—256K37.7CDN
qwen3.6-27b-optiq-4bitmlx-community/Qwen3.6-27B-OptiQ-4bitqwen3_coder_xmlqwen3—256K37.7CDN
qwen3.6-27b-udunsloth/Qwen3.6-27B-UD-MLX-4bitqwen3_coder_xmlqwen3—256K37.7CDN
qwen3.6-27b-ud-3bitunsloth/Qwen3.6-27B-UD-MLX-3bitqwen3_coder_xmlqwen3—256K37.7CDN
qwen3.6-27b-ud-6bitunsloth/Qwen3.6-27B-UD-MLX-6bitqwen3_coder_xmlqwen3—256K37.7CDN
qwen3.6-35bmlx-community/Qwen3.6-35B-A3B-4bitqwen3_coder_xmlqwen3hybrid · moe256K32.1CDN
qwen3.6-35b-4bitmlx-community/Qwen3.6-35B-A3B-4bitqwen3_coder_xmlqwen3hybrid · moe256K32.1CDN
qwen3.6-35b-6bitmlx-community/Qwen3.6-35B-A3B-6bitqwen3_coder_xmlqwen3hybrid · moe256K32.1CDN
qwen3.6-35b-8bitmlx-community/Qwen3.6-35B-A3B-8bitqwen3_coder_xmlqwen3hybrid · moe256K32.1CDN
qwen3.6-35b-dwqmlx-community/Qwen3.6-35B-A3B-4bit-DWQqwen3_coder_xmlqwen3hybrid · moe256K32.1CDN
qwen3.6-35b-mtp-4bitmlx-community/Qwen3.6-35B-A3B-MTP-4bitqwen3_coder_xmlqwen3hybrid · moe256K32.1CDN
qwen3.6-35b-mxfp4mlx-community/Qwen3.6-35B-A3B-mxfp4qwen3_coder_xmlqwen3hybrid · moe256K32.1CDN
qwen3.6-35b-nvfp4mlx-community/Qwen3.6-35B-A3B-nvfp4qwen3_coder_xmlqwen3hybrid · moe256K32.1CDN
qwen3.6-35b-optiq-4bitmlx-community/Qwen3.6-35B-A3B-OptiQ-4bitqwen3_coder_xmlqwen3hybrid · moe256K32.1CDN
qwen3.6-35b-udunsloth/Qwen3.6-35B-A3B-UD-MLX-4bitqwen3_coder_xmlqwen3hybrid · moe256K32.1CDN

Notes & caveats

Qwen3.5 · 15 aliases

The workhorse line on rapid-mlx. MLX-native quants from 4 GB-RAM-friendly 4B up to the 122B-A10B MoE that needs a Mac Studio with 192 GB+ unified memory. Default tool-call parser is hermes; reasoning parser is qwen3.

parser: hermes

aliashf repotool parserreasoningflagscontextAA indexget it
qwen3.5-122b-6bitmlx-community/Qwen3.5-122B-A10B-6bithermesqwen3hybrid · moe—32.8HF
qwen3.5-122b-8bitmlx-community/Qwen3.5-122B-A10B-8bithermesqwen3hybrid · moe256K32.8CDN
qwen3.5-122b-mxfp4nightmedia/Qwen3.5-122B-A10B-Text-mxfp4-mlxhermesqwen3hybrid · moe256K32.8CDN
qwen3.5-27b-4bitmlx-community/Qwen3.5-27B-4bithermesqwen3—256K34.6CDN
qwen3.5-27b-6bitmlx-community/Qwen3.5-27B-6bithermesqwen3—256K34.6CDN
qwen3.5-27b-8bitmlx-community/Qwen3.5-27B-8bithermesqwen3—256K34.6CDN
qwen3.5-35b-4bitmlx-community/Qwen3.5-35B-A3B-4bithermesqwen3hybrid · moe256K29.9CDN
qwen3.5-35b-6bitmlx-community/Qwen3.5-35B-A3B-6bithermesqwen3hybrid · moe256K29.9CDN
qwen3.5-35b-8bitmlx-community/Qwen3.5-35B-A3B-8bithermesqwen3hybrid · moe256K29.9CDN
qwen3.5-4b-4bitmlx-community/Qwen3.5-4B-MLX-4bithermesqwen3—256K20.4CDN
qwen3.5-4b-6bitmlx-community/Qwen3.5-4B-6bithermesqwen3—256K20.4CDN
qwen3.5-4b-8bitmlx-community/Qwen3.5-4B-8bithermesqwen3—256K20.4CDN
qwen3.5-9b-4bitmlx-community/Qwen3.5-9B-4bithermesqwen3—256K21.8CDN
qwen3.5-9b-6bitmlx-community/Qwen3.5-9B-6bithermesqwen3—256K21.8CDN
qwen3.5-9b-8bitmlx-community/Qwen3.5-9B-8bithermesqwen3—256K21.8CDN

Notes & caveats

Qwen3 Coder · 4 aliases

Alibaba's code-specialised line — instruction-tuned on code corpora and aligned for IDE-style completion + agentic edit workflows. Both aliases use the hermes tool envelope so Cursor / Claude Code / Aider / Continue all work without per-model config.

parser: hermes

aliashf repotool parserreasoningflagscontextAA indexget it
qwen3-coder-30b-4bitmlx-community/Qwen3-Coder-30B-A3B-Instruct-4bithermes—moe · spec256K13.6CDN
qwen3-coder-4bitlmstudio-community/Qwen3-Coder-Next-MLX-4bitqwen3_coder_xml—hybrid · moe256K21.3CDN
qwen3-coder-next-80b-4bitmlx-community/Qwen3-Coder-Next-4bitqwen3_coder_xml—hybrid · moe256K21.3CDN
qwen3-coder-next-80b-8bitlmstudio-community/Qwen3-Coder-Next-MLX-8bitqwen3_coder_xml—hybrid · moe256K21.3CDN

Notes & caveats

Qwen3-VL · 4 aliases

Vision-language extension of Qwen3. rapid-mlx serves the OpenAI-compatible /v1/chat/completions vision shape — pass image URLs or base64 in the standard message-content array. The 2B / 4B / 8B variants fit a 16 GB Mac; the 30B-A3B MoE wants 32 GB+.

parser: hermes

aliashf repotool parserreasoningflagscontextAA indexget it
qwen3-vl-2b-4bitmlx-community/Qwen3-VL-2B-Instruct-4bithermes—spec256K—CDN
qwen3-vl-30b-4bitmlx-community/Qwen3-VL-30B-A3B-Instruct-4bithermesqwen3moe · spec256K9.9CDN
qwen3-vl-4b-4bitmlx-community/Qwen3-VL-4B-Instruct-4bithermesqwen3spec256K3.7CDN
qwen3-vl-8b-4bitmlx-community/Qwen3-VL-8B-Instruct-4bithermesqwen3spec256K8.2CDN

Notes & caveats

Qwopus · 2 aliases

Community-trained "Qwen meets Opus" line — Qwen3 base weights post-trained on an Opus-style alignment mix, with hybrid cache enabled. Three MLX quants: 9B-4bit, 27B-4bit, 27B-8bit.

parser: hermes

aliashf repotool parserreasoningflagscontextAA indexget it
qwopus-27b-4bitJackrong/MLX-Qwopus3.5-27B-v3-4bithermesqwen3hybrid256K—CDN
qwopus-9b-4bitJackrong/MLX-Qwopus3.5-9B-v3-4bithermesqwen3hybrid256K—CDN

Notes & caveats

Qwen3 (legacy) · 10 aliases

The original Qwen3 line — kept for users who want a small, well-understood baseline. Includes the 2507 instruct and thinking checkpoints, which give you the same family on two different alignment recipes.

parser: hermes

aliashf repotool parserreasoningflagscontextAA indexget it
qwen3-0.6bmlx-community/Qwen3-0.6B-4bithermesqwen3spec40K1CDN
qwen3-0.6b-4bitmlx-community/Qwen3-0.6B-4bithermesqwen3spec40K1CDN
qwen3-0.6b-8bitmlx-community/Qwen3-0.6B-8bithermesqwen3spec40K1CDN
qwen3-1.7bmlx-community/Qwen3-1.7B-4bithermesqwen3spec40K2.2CDN
qwen3-1.7b-4bitmlx-community/Qwen3-1.7B-4bithermesqwen3spec40K2.2CDN
qwen3-4b-8bitmlx-community/Qwen3-4B-8bithermesqwen3spec40K8.2CDN
qwen3-4b-instruct-2507-4bitmlx-community/Qwen3-4B-Instruct-2507-4bithermes—spec256K6.9CDN
qwen3-4b-thinking-2507-4bitmlx-community/Qwen3-4B-Thinking-2507-4bithermesqwen3spec256K11.9CDN
qwen3-8b-4bitmlx-community/Qwen3-8B-4bithermesqwen3spec40K8.3CDN
qwen3-8b-8bitmlx-community/Qwen3-8B-8bithermesqwen3spec40K8.3CDN

Notes & caveats

Qwen2.5 · 1 alias

One Qwen2.5 alias kept in the registry — the 14B Instruct 4-bit — as a reference baseline for anyone benchmarking the Qwen3 line against its predecessor. Same hermes tool envelope as the rest of the Qwen family.

parser: hermes

aliashf repotool parserreasoningflagscontextAA indexget it
qwen2.5-14b-4bitmlx-community/Qwen2.5-14B-Instruct-4bithermes—spec32K—CDN

Notes & caveats

Context is read from the config.json of the exact build each alias pulls; for an embedding model it is the most input tokens the engine embeds, which can be less than the config declares. AA index is the Artificial Analysis Intelligence Index for the base model at full precision with reasoning on — a property of the model, not a score for our quantised build.

Notes & caveats

External reading

Where next