Public MLX model mirror

Choose the model your Mac can carry.

Pick your unified memory, compare verified model families, and copy one command. Public weights, no account, built for Apple Silicon.

Mirror status Operational · R2 edge
rapid-mlx pull qwen3.5-4b-4bit
Catalog
239 aliases
Mirror status
221 mirrored
Access
anonymous · public
Optimized for
Apple Silicon

Pick by Mac RAM

Pick your Mac's unified-memory tier and see the two models rapid-mlx recommend surfaces for it — smart first, fast second. Same picks as the CLI recommender. The full catalog stays available below.

8 GB · SELECTED Two curated picks — smart first, fast second
Smart

LFM2.5 2.6B 4bit

  • 64% capability
  • 3.0 GB footprint
  • 93.5 tok/s

Not for coding

rapid-mlx pull lfm2.5-2.6b-4bit
Fast

LFM2.5 1B 4bit

  • 47% capability
  • 1.9 GB footprint
  • 208.4 tok/s

Basic chat

rapid-mlx pull lfm2.5-1b-4bit
Smart and fast model picks for each Mac memory tier, from the engine recommender
RAMSmart pickFast pick
8 GBLFM2.5 2.6B 4bitLFM2.5 1B 4bit
16 GBQwen3.5 4B 4bitLFM2.5 1B 4bit
18 GBQwen3.5 9B 4bitQwen3.5 4B 4bit
24 GBTernary Bonsai 27B 2bitQwen3.5 4B 4bit
32 GBQwen3.8 27B 4bitQwen3.5 4B 4bit
48 GBQwen3.8 27B 4bitQwen3.6 35B-A3B 4bit
64 GBQwen3.8 27B 4bitQwen3.6 35B-A3B 4bit
96 GBQwen3.8 27B 4bitQwen3.6 35B-A3B 4bit

Picks come straight from the engine recommender (rapid-mlx recommend / hardware tiers): smart = highest capability that fits, fast = highest throughput. Footprint is measured resident memory; context length, KV cache and concurrent requests consume more.

Verified first-class support

Curated model families

Verified aliases, shard sizes, RAM tiers and copy-ready commands. Hunyuan 3 is the one family fetched from Hugging Face rather than the R2 mirror.

Qwen 3.6

MoE workhorse · 8–95 GB

The 35B-A3B MoE anchors the 48-95 GB tier (3B active, 35B total). Native MTP head baked into every checkpoint — pair with a matching MTP drafter (below) for speculative decode.

qwen3.6-35b-8bitBrowse · mlx-community/Qwen3.6-35B-A3B-8bit
35.2 GB48-95 GB

35B-A3B MoE at 8-bit MLX. Anchor pick of the 48-95 GB tier.

rapid-mlx pull qwen3.6-35b-8bit
qwen3.6-35b-4bitBrowse · mlx-community/Qwen3.6-35B-A3B-4bit
19.0 GB24-47 GB

35B-A3B MoE at 4-bit MLX. Fits at 32-48 GB with headroom.

rapid-mlx pull qwen3.6-35b-4bit
qwen3.6-27b-4bitBrowse · mlx-community/Qwen3.6-27B-4bit
15.0 GB24-47 GB

Dense 27B at 4-bit MLX. Hybrid attention layout.

rapid-mlx pull qwen3.6-27b-4bit
qwen3.5-4b-4bitBrowse · mlx-community/Qwen3.5-4B-MLX-4bit
2.9 GB8-23 GB

Qwen 3.5 4B 4-bit — the small-tier stand-in until Qwen 3.6 ships a <8B SKU.

rapid-mlx pull qwen3.5-4b-4bit

Model catalog

Every rapid-mlx alias joined against the live R2 mirror.

loading catalog…

Context is read from the config.json of the exact build each alias pulls, so two quantisations of one model can differ. AA index is the Artificial Analysis Intelligence Index for the base model at full precision with reasoning on — a property of the model, not a score for our quantised build. Sort by it to find the most capable model your Mac can actually hold.

Same Mac, same method

Measured on a Mac mini M2 Pro

Eighteen of the most-pulled models, each with its own page: decode speed, first token, peak memory and weights.

ModelArchitectureDecodeFirst tokenPeak memoryWeights
Qwen3 0.6B0.6 B dense · 4-bit225.9 tok/s0.17 s1.2 GB0.3 GB
LFM2.5 1.2B1.2 B hybrid conv+attention · 4-bit209.8 tok/s0.2 s1.1 GB0.6 GB
Ternary Bonsai 1.7B1.7 B ternary · 2-bit169.0 tok/s0.23 s1.3 GB0.5 GB
LFM2.5 2.6B2.6 B hybrid conv+attention · 4-bit93.5 tok/s0.33 s2.2 GB1.5 GB
Llama 3.2 3B3 B dense · 4-bit82.2 tok/s0.25 s2.4 GB1.7 GB
Gemma 3 4B (QAT)4 B dense · QAT 4-bit · vision66.7 tok/s0.45 s4.4 GB2.8 GB
Qwen3 4B Thinking (2507)4 B dense · 4-bit · reasoning61.3 tok/s0.41 s3.0 GB2.1 GB
Qwen3 4B Instruct (2507)4 B dense · 4-bit61.2 tok/s0.4 s2.9 GB2.1 GB
Qwen3.5 4B4 B dense · 4-bit60.7 tok/s0.56 s3.4 GB2.9 GB
Qwen3.6 35B (A3B MoE)MoE · 35 B total / 3 B active · 4-bit59.6 tok/s0.55 s9.5 GB19.1 GB
Qwen3 Coder 30B (A3B MoE)MoE · 30 B total / 3 B active · 4-bit · code53.0 tok/s0.54 s10.0 GB16.0 GB
Gemma 4 26B (A4B MoE)MoE · 26 B total / 4 B active · 4-bit50.4 tok/s0.68 s6.0 GB14.3 GB
GPT-OSS 20BMoE · 20 B total · MXFP4-Q8 · reasoning47.9 tok/s0.57 s9.4 GB11.3 GB
Qwen3 8B8 B dense · 4-bit37.1 tok/s0.62 s5.1 GB4.3 GB
Qwen3.5 9B9 B dense · 4-bit36.4 tok/s0.84 s5.8 GB5.6 GB
Gemma 4 12B12 B dense · 4-bit · vision22.6 tok/s0.95 s7.8 GB6.3 GB
Ternary Bonsai 27B27 B ternary · 2-bit17.4 tok/s1.48 s8.4 GB7.9 GB
Qwen3.6 27B27 B dense · 4-bit11.1 tok/s2.41 s8.7 GB15.0 GB

rapid-mlx 0.12.10, measured 2026-08-11. Median of 3 runs, temperature 0, engine-reported token counts, prefix cache defeated with a unique salt per request. Every page carries the full method line.

No rapid-mlx required

Download directly from the mirror.

Every file is a public, CORS-enabled Cloudflare R2 URL. Grab one file, a whole repository, or point any MLX runtime at the result — the files are the exact upstream MLX weights, usable with mlx-lm, LM Studio, mlx-vlm, or your own runtime.

One file
curl -LO https://models.rapidmlx.com/mlx-community/Qwen3.5-9B-6bit/model.safetensors
A whole repository — file list comes from the mirror, no Hugging Face needed
REPO=mlx-community/Qwen3.5-9B-6bit curl -s "https://models.rapidmlx.com/$REPO/" | grep -o '"name":"[^"]*"' | cut -d'"' -f4 | xargs -P4 -I{} curl -fL --create-dirs -o "$REPO/{}" "https://models.rapidmlx.com/$REPO/{}"
Run with any MLX tool
mlx_lm.generate --model ./$REPO --prompt "Hello"