Qwen
62 MLX aliases · Alibaba's full open-weights line — Qwen3.8 / 3.6 / 3.5 / Coder / VL / legacy / 2.5 / Qwopus.
Pick one
One command per line of this family, smallest download first. Your Mac needs the download size in free memory plus room for macOS and the context window; the hardware tiers page has the engine's picks for every RAM size. The first run downloads the weights and starts an OpenAI-compatible server on http://localhost:8000/v1.
rapid-mlx serve qwen3-0.6b-4bitrapid-mlx serve qwen3-vl-2b-4bitrapid-mlx serve qwen3.5-4b-4bitrapid-mlx serve qwopus-9b-4bitrapid-mlx serve qwen3.5-9b-4bitrapid-mlx serve qwen2.5-14b-4bitrapid-mlx serve qwen3-coder-30b-4bitrapid-mlx serve qwen3.8-27b-4bitrapid-mlx serve qwen3.6-35b-4bit
The Qwen family on rapid-mlx covers every Alibaba open-weights line we support — from the workhorse Qwen3.5 (4B → 122B-A10B MoE) and the post-3.5 refresh Qwen3.6, through the code-tuned Qwen3 Coder and the vision-language Qwen3-VL, down to the smaller legacy Qwen3 / Qwen2.5 baselines and the community Opus-aligned Qwopus. Each alias selects its own tool-call parser (hermes on Qwen3.5, qwen3_coder_xml on Qwen3.6 / 3.8 and Coder) and, on the thinking models, the qwen3 reasoning parser — the tables below list each one — so structured tool calls and chain-of-thought stream out of /v1/chat/completions and /v1/responses with no extra config.
- family
- Qwen
- aliases
- 62
- lines
- 8
- install
- rapid-mlx serve <alias>
- OpenAI base URL
- http://localhost:8000/v1
Download
Every alias on this page downloads with one command — the pull buttons in the tables below copy it. 60 of the 62 aliases on this page are mirrored on the rapid-mlx CDN; the rest pull from Hugging Face directly — with automatic mid-pull fallback to Hugging Face if a mirror file slows down. Weights land in the standard Hugging Face cache, and rapid-mlx serve pulls automatically on first use. Live mirror status →
Lines in this family
- Qwen3.8 (7 aliases)
- Qwen3.6 (19 aliases)
- Qwen3.5 (15 aliases)
- Qwen3 Coder (4 aliases)
- Qwen3-VL (4 aliases)
- Qwopus (2 aliases)
- Qwen3 (legacy) (10 aliases)
- Qwen2.5 (1 alias)
Qwen3.8 · 7 aliases
The newest line. The 27B is a hybrid GatedDeltaNet checkpoint that runs text on the default lane and answers questions about images via --mllm. Alongside the standard 4-bit build we publish our own mixed-precision 3.5-bpw quantization (rapid-mlx/Qwen3.8-27B-mixed-3.5bpw-MLX) that spends its bits where they matter — 13.0 GB of weights against the 4-bit build's 15.0 GB.
parser: hermes
| alias | hf repo | tool parser | reasoning | flags | context | AA index | get it |
|---|---|---|---|---|---|---|---|
| qwen3.8-27b-4bit | rapid-mlx/Qwen3.8-27B-4bit-MTP-MLX | qwen3_coder_xml | qwen3 | hybrid | 256K | 52 | CDN |
| qwen3.8-27b-4bit-fp16 | rapid-mlx/Qwen3.8-27B-4bit-MTP-fp16-MLX | qwen3_coder_xml | qwen3 | hybrid | 256K | 52 | CDN |
| qwen3.8-27b-abliterated-4bit | windowsxp811203/Qwen3.8-27B-Abliterated-MLX-oQ4e-mtp | qwen3_coder_xml | qwen3 | hybrid | 256K | — | CDN |
| qwen3.8-27b-mixed-3.5bpw | rapid-mlx/Qwen3.8-27B-mixed-3.5bpw-MLX | qwen3_coder_xml | qwen3 | hybrid | 256K | — | CDN |
| qwen3.8-27b-tensorfold | Vontra/Qwen3.8-27B-MLX-4bit | — | qwen3 | hybrid | 256K | — | CDN |
| qwen3.8-flash-next-4bit | rapid-mlx/Qwen3.8-Flash-Next-4bit | hermes | qwen3 | hybrid · moe | 256K | 55.8 | CDN |
| qwen3.8-flash-next-tensorfold | TensorFold/Qwen3.8-Flash-Next-MLX-4bit-MTP | — | qwen3 | hybrid · moe | — | — | HF |
Notes & caveats
- Multi-token prediction is on by default for
qwen3.8-27b-4bitandqwen3.8-27b-4bit-fp16(--no-spec-decodeturns it off); other builds opt in with--speculative-config '{"method":"mtp"}'.rapid-mlx info <alias>shows which. - The mixed-3.5bpw build fits a 48 GB Mac with headroom; it is strong at code and tool calling but not a RAM-tier default.
Qwen3.6 · 19 aliases
Post-3.5 refresh. Same hybrid-cache MoE architecture as 3.5-35B but a tighter post-training mix and an XML-style tool-call envelope, parsed with our qwen3_coder_xml parser. We ship DWQ (data-aware quant) and Unsloth's UD-MLX quants alongside the standard 4/6/8-bit set across both 27B and 35B.
parser: qwen3_coder_xml
| alias | hf repo | tool parser | reasoning | flags | context | AA index | get it |
|---|---|---|---|---|---|---|---|
| qwen3.6-27b | mlx-community/Qwen3.6-27B-4bit | qwen3_coder_xml | qwen3 | — | 256K | 37.7 | CDN |
| qwen3.6-27b-4bit | mlx-community/Qwen3.6-27B-4bit | qwen3_coder_xml | qwen3 | — | 256K | 37.7 | CDN |
| qwen3.6-27b-6bit | lmstudio-community/Qwen3.6-27B-MLX-6bit | qwen3_coder_xml | qwen3 | — | 256K | 37.7 | CDN |
| qwen3.6-27b-8bit | unsloth/Qwen3.6-27B-MLX-8bit | qwen3_coder_xml | qwen3 | — | 256K | 37.7 | CDN |
| qwen3.6-27b-mtp-4bit | mlx-community/Qwen3.6-27B-MTP-4bit | qwen3_coder_xml | qwen3 | — | 256K | 37.7 | CDN |
| qwen3.6-27b-optiq-4bit | mlx-community/Qwen3.6-27B-OptiQ-4bit | qwen3_coder_xml | qwen3 | — | 256K | 37.7 | CDN |
| qwen3.6-27b-ud | unsloth/Qwen3.6-27B-UD-MLX-4bit | qwen3_coder_xml | qwen3 | — | 256K | 37.7 | CDN |
| qwen3.6-27b-ud-3bit | unsloth/Qwen3.6-27B-UD-MLX-3bit | qwen3_coder_xml | qwen3 | — | 256K | 37.7 | CDN |
| qwen3.6-27b-ud-6bit | unsloth/Qwen3.6-27B-UD-MLX-6bit | qwen3_coder_xml | qwen3 | — | 256K | 37.7 | CDN |
| qwen3.6-35b | mlx-community/Qwen3.6-35B-A3B-4bit | qwen3_coder_xml | qwen3 | hybrid · moe | 256K | 32.1 | CDN |
| qwen3.6-35b-4bit | mlx-community/Qwen3.6-35B-A3B-4bit | qwen3_coder_xml | qwen3 | hybrid · moe | 256K | 32.1 | CDN |
| qwen3.6-35b-6bit | mlx-community/Qwen3.6-35B-A3B-6bit | qwen3_coder_xml | qwen3 | hybrid · moe | 256K | 32.1 | CDN |
| qwen3.6-35b-8bit | mlx-community/Qwen3.6-35B-A3B-8bit | qwen3_coder_xml | qwen3 | hybrid · moe | 256K | 32.1 | CDN |
| qwen3.6-35b-dwq | mlx-community/Qwen3.6-35B-A3B-4bit-DWQ | qwen3_coder_xml | qwen3 | hybrid · moe | 256K | 32.1 | CDN |
| qwen3.6-35b-mtp-4bit | mlx-community/Qwen3.6-35B-A3B-MTP-4bit | qwen3_coder_xml | qwen3 | hybrid · moe | 256K | 32.1 | CDN |
| qwen3.6-35b-mxfp4 | mlx-community/Qwen3.6-35B-A3B-mxfp4 | qwen3_coder_xml | qwen3 | hybrid · moe | 256K | 32.1 | CDN |
| qwen3.6-35b-nvfp4 | mlx-community/Qwen3.6-35B-A3B-nvfp4 | qwen3_coder_xml | qwen3 | hybrid · moe | 256K | 32.1 | CDN |
| qwen3.6-35b-optiq-4bit | mlx-community/Qwen3.6-35B-A3B-OptiQ-4bit | qwen3_coder_xml | qwen3 | hybrid · moe | 256K | 32.1 | CDN |
| qwen3.6-35b-ud | unsloth/Qwen3.6-35B-A3B-UD-MLX-4bit | qwen3_coder_xml | qwen3 | hybrid · moe | 256K | 32.1 | CDN |
Notes & caveats
- Both 27B and 35B variants ship — pick 27B for a 32 GB Mac, 35B-A3B for 64 GB+.
- DWQ + UD-MLX quants are higher-quality 4-bit than the stock release.
- Hybrid cache on all 35B variants — long context with bounded KV growth.
Qwen3.5 · 15 aliases
The workhorse line on rapid-mlx. MLX-native quants from 4 GB-RAM-friendly 4B up to the 122B-A10B MoE that needs a Mac Studio with 192 GB+ unified memory. Default tool-call parser is hermes; reasoning parser is qwen3.
parser: hermes
| alias | hf repo | tool parser | reasoning | flags | context | AA index | get it |
|---|---|---|---|---|---|---|---|
| qwen3.5-122b-6bit | mlx-community/Qwen3.5-122B-A10B-6bit | hermes | qwen3 | hybrid · moe | — | 32.8 | HF |
| qwen3.5-122b-8bit | mlx-community/Qwen3.5-122B-A10B-8bit | hermes | qwen3 | hybrid · moe | 256K | 32.8 | CDN |
| qwen3.5-122b-mxfp4 | nightmedia/Qwen3.5-122B-A10B-Text-mxfp4-mlx | hermes | qwen3 | hybrid · moe | 256K | 32.8 | CDN |
| qwen3.5-27b-4bit | mlx-community/Qwen3.5-27B-4bit | hermes | qwen3 | — | 256K | 34.6 | CDN |
| qwen3.5-27b-6bit | mlx-community/Qwen3.5-27B-6bit | hermes | qwen3 | — | 256K | 34.6 | CDN |
| qwen3.5-27b-8bit | mlx-community/Qwen3.5-27B-8bit | hermes | qwen3 | — | 256K | 34.6 | CDN |
| qwen3.5-35b-4bit | mlx-community/Qwen3.5-35B-A3B-4bit | hermes | qwen3 | hybrid · moe | 256K | 29.9 | CDN |
| qwen3.5-35b-6bit | mlx-community/Qwen3.5-35B-A3B-6bit | hermes | qwen3 | hybrid · moe | 256K | 29.9 | CDN |
| qwen3.5-35b-8bit | mlx-community/Qwen3.5-35B-A3B-8bit | hermes | qwen3 | hybrid · moe | 256K | 29.9 | CDN |
| qwen3.5-4b-4bit | mlx-community/Qwen3.5-4B-MLX-4bit | hermes | qwen3 | — | 256K | 20.4 | CDN |
| qwen3.5-4b-6bit | mlx-community/Qwen3.5-4B-6bit | hermes | qwen3 | — | 256K | 20.4 | CDN |
| qwen3.5-4b-8bit | mlx-community/Qwen3.5-4B-8bit | hermes | qwen3 | — | 256K | 20.4 | CDN |
| qwen3.5-9b-4bit | mlx-community/Qwen3.5-9B-4bit | hermes | qwen3 | — | 256K | 21.8 | CDN |
| qwen3.5-9b-6bit | mlx-community/Qwen3.5-9B-6bit | hermes | qwen3 | — | 256K | 21.8 | CDN |
| qwen3.5-9b-8bit | mlx-community/Qwen3.5-9B-8bit | hermes | qwen3 | — | 256K | 21.8 | CDN |
Notes & caveats
- Default workhorse — start with
qwen3.5-9b-4biton a 16 GB Mac,qwen3.5-27b-4biton 32 GB+. - The 35B and 122B variants are MoE (A3B / A10B) — heavy weights, light active params.
- Hybrid-cache enabled on the 35B and 122B quants for long-context throughput.
Qwen3 Coder · 4 aliases
Alibaba's code-specialised line — instruction-tuned on code corpora and aligned for IDE-style completion + agentic edit workflows. Both aliases use the hermes tool envelope so Cursor / Claude Code / Aider / Continue all work without per-model config.
parser: hermes
| alias | hf repo | tool parser | reasoning | flags | context | AA index | get it |
|---|---|---|---|---|---|---|---|
| qwen3-coder-30b-4bit | mlx-community/Qwen3-Coder-30B-A3B-Instruct-4bit | hermes | — | moe · spec | 256K | 13.6 | CDN |
| qwen3-coder-4bit | lmstudio-community/Qwen3-Coder-Next-MLX-4bit | qwen3_coder_xml | — | hybrid · moe | 256K | 21.3 | CDN |
| qwen3-coder-next-80b-4bit | mlx-community/Qwen3-Coder-Next-4bit | qwen3_coder_xml | — | hybrid · moe | 256K | 21.3 | CDN |
| qwen3-coder-next-80b-8bit | lmstudio-community/Qwen3-Coder-Next-MLX-8bit | qwen3_coder_xml | — | hybrid · moe | 256K | 21.3 | CDN |
Notes & caveats
- 30B-A3B-Instruct is the recommended pick for general code work on a 32 GB+ Mac.
- The MLX-Next 4-bit quant runs hybrid cache for long-file context.
Qwen3-VL · 4 aliases
Vision-language extension of Qwen3. rapid-mlx serves the OpenAI-compatible /v1/chat/completions vision shape — pass image URLs or base64 in the standard message-content array. The 2B / 4B / 8B variants fit a 16 GB Mac; the 30B-A3B MoE wants 32 GB+.
parser: hermes
| alias | hf repo | tool parser | reasoning | flags | context | AA index | get it |
|---|---|---|---|---|---|---|---|
| qwen3-vl-2b-4bit | mlx-community/Qwen3-VL-2B-Instruct-4bit | hermes | — | spec | 256K | — | CDN |
| qwen3-vl-30b-4bit | mlx-community/Qwen3-VL-30B-A3B-Instruct-4bit | hermes | qwen3 | moe · spec | 256K | 9.9 | CDN |
| qwen3-vl-4b-4bit | mlx-community/Qwen3-VL-4B-Instruct-4bit | hermes | qwen3 | spec | 256K | 3.7 | CDN |
| qwen3-vl-8b-4bit | mlx-community/Qwen3-VL-8B-Instruct-4bit | hermes | qwen3 | spec | 256K | 8.2 | CDN |
Notes & caveats
- Use the standard OpenAI
image_urlmessage-content shape — no custom transport. - 30B-A3B variant has the
qwen3reasoning parser wired so you also getreasoning_contenton/v1/responses.
Qwopus · 2 aliases
Community-trained "Qwen meets Opus" line — Qwen3 base weights post-trained on an Opus-style alignment mix, with hybrid cache enabled. Three MLX quants: 9B-4bit, 27B-4bit, 27B-8bit.
parser: hermes
| alias | hf repo | tool parser | reasoning | flags | context | AA index | get it |
|---|---|---|---|---|---|---|---|
| qwopus-27b-4bit | Jackrong/MLX-Qwopus3.5-27B-v3-4bit | hermes | qwen3 | hybrid | 256K | — | CDN |
| qwopus-9b-4bit | Jackrong/MLX-Qwopus3.5-9B-v3-4bit | hermes | qwen3 | hybrid | 256K | — | CDN |
Notes & caveats
- Hybrid cache on all variants — long-context cost bounded.
- Reasoning parser is
qwen3, so<think>-style CoT streams cleanly.
Qwen3 (legacy) · 10 aliases
The original Qwen3 line — kept for users who want a small, well-understood baseline. Includes the 2507 instruct and thinking checkpoints, which give you the same family on two different alignment recipes.
parser: hermes
| alias | hf repo | tool parser | reasoning | flags | context | AA index | get it |
|---|---|---|---|---|---|---|---|
| qwen3-0.6b | mlx-community/Qwen3-0.6B-4bit | hermes | qwen3 | spec | 40K | 1 | CDN |
| qwen3-0.6b-4bit | mlx-community/Qwen3-0.6B-4bit | hermes | qwen3 | spec | 40K | 1 | CDN |
| qwen3-0.6b-8bit | mlx-community/Qwen3-0.6B-8bit | hermes | qwen3 | spec | 40K | 1 | CDN |
| qwen3-1.7b | mlx-community/Qwen3-1.7B-4bit | hermes | qwen3 | spec | 40K | 2.2 | CDN |
| qwen3-1.7b-4bit | mlx-community/Qwen3-1.7B-4bit | hermes | qwen3 | spec | 40K | 2.2 | CDN |
| qwen3-4b-8bit | mlx-community/Qwen3-4B-8bit | hermes | qwen3 | spec | 40K | 8.2 | CDN |
| qwen3-4b-instruct-2507-4bit | mlx-community/Qwen3-4B-Instruct-2507-4bit | hermes | — | spec | 256K | 6.9 | CDN |
| qwen3-4b-thinking-2507-4bit | mlx-community/Qwen3-4B-Thinking-2507-4bit | hermes | qwen3 | spec | 256K | 11.9 | CDN |
| qwen3-8b-4bit | mlx-community/Qwen3-8B-4bit | hermes | qwen3 | spec | 40K | 8.3 | CDN |
| qwen3-8b-8bit | mlx-community/Qwen3-8B-8bit | hermes | qwen3 | spec | 40K | 8.3 | CDN |
Notes & caveats
qwen3-0.6b-4bitis the smallest practical chat-quality model in the registry — useful for laptop-class CI smoke tests.- The
-thinking-2507-variant streamsreasoning_content; the-instruct-2507-variant does not.
Qwen2.5 · 1 alias
One Qwen2.5 alias kept in the registry — the 14B Instruct 4-bit — as a reference baseline for anyone benchmarking the Qwen3 line against its predecessor. Same hermes tool envelope as the rest of the Qwen family.
parser: hermes
| alias | hf repo | tool parser | reasoning | flags | context | AA index | get it |
|---|---|---|---|---|---|---|---|
| qwen2.5-14b-4bit | mlx-community/Qwen2.5-14B-Instruct-4bit | hermes | — | spec | 32K | — | CDN |
Notes & caveats
- Use Qwen3.5-9B or Qwen3.6-27B instead unless you have a specific reason to pin Qwen2.5.
Context is read from the config.json of the exact build each alias pulls; for an embedding model it is the most input tokens the engine embeds, which can be less than the config declares. AA index is the Artificial Analysis Intelligence Index for the base model at full precision with reasoning on — a property of the model, not a score for our quantised build.
Notes & caveats
- Tool-call parser is
hermesacross the family except for Qwen3.6, which usesqwen3_coder_xml. - Reasoning parser is
qwen3across the family —<think>-style CoT streams toreasoning_contenton/v1/responses. - For the larger MoE variants (35B-A3B, 122B-A10B) hybrid cache is on by default — long-context cost stays bounded.