Small / curated models
Curated research checkpoints — Bonsai / MiniCPM5 / SmolLM3 / Nanbeige / Nemotron — plus one-off big checkpoints (Kimi K2.6, North-Mini-Code, MiMo-V2.6 Flash) and experimental community builds.
Pick one
One command per line of this family, smallest download first. Your Mac needs the download size in free memory plus room for macOS and the context window; the hardware tiers page has the engine's picks for every RAM size. The first run downloads the weights and starts an OpenAI-compatible server on http://localhost:8000/v1.
rapid-mlx serve minicpm5-1b-4bitrapid-mlx serve smollm3-3b-4bitrapid-mlx serve nanbeige4.1-3b-4bitrapid-mlx serve nemotron-labs-diffusion-3b-4bitrapid-mlx serve bonsai-27b-2bitrapid-mlx serve north-mini-code-4bitrapid-mlx serve g9v3-39a5b-4bit
A grab-bag of small research models that don't fit a larger vendor family but earn their place in the registry by being unusually good at something specific — ternary-weight agent reasoning (Bonsai), 1B on-device chat (MiniCPM5), long-context at 3B (SmolLM3), or pruning-recipe research (Nemotron). See also our notable small models callout for the curated short-list.
- family
- Small / curated models
- aliases
- 21
- lines
- 7
- install
- rapid-mlx serve <alias>
- OpenAI base URL
- http://localhost:8000/v1
Download
Every alias on this page downloads with one command — the pull buttons in the tables below copy it. 18 of the 21 aliases on this page are mirrored on the rapid-mlx CDN; the rest pull from Hugging Face directly — with automatic mid-pull fallback to Hugging Face if a mirror file slows down. Weights land in the standard Hugging Face cache, and rapid-mlx serve pulls automatically on first use. Live mirror status →
Lines in this family
- Bonsai (ternary) (5 aliases)
- MiniCPM5 (3 aliases)
- SmolLM3 (1 alias)
- Nanbeige (1 alias)
- One-off big checkpoints (4 aliases)
- Experimental community builds (3 aliases)
- Nemotron (4 aliases)
Bonsai (ternary) · 5 aliases
Ternary-weight research line — 2-bit packed 1.7B and 27B checkpoints, tool-call capable via the hermes envelope. Tiny footprint (~2.5 GB at 1.7B), useful for agent and draft-model experiments.
parser: hermes
| alias | hf repo | tool parser | reasoning | flags | context | AA index | get it |
|---|---|---|---|---|---|---|---|
| bonsai-1.7b-2bit | prism-ml/Ternary-Bonsai-1.7B-mlx-2bit | hermes | — | — | 32K | — | CDN |
| bonsai-27b-2bit | prism-ml/Ternary-Bonsai-27B-mlx-2bit | hermes | qwen3 | — | 256K | — | CDN |
| bonsai-8b-2bit | prism-ml/Ternary-Bonsai-8B-mlx-2bit | hermes | qwen3 | — | 64K | — | CDN |
| bonsai-image-4b-2bit | prism-ml/bonsai-image-ternary-4B-mlx-2bit | — | — | image-gen | — | — | CDN |
| bonsai2-27b-2bit | prism-ml/Ternary-Bonsai-2-27B-mlx-2bit | qwen3_coder_xml | qwen3 | hybrid | 256K | — | CDN |
Notes & caveats
- Ternary 2-bit weights — the 1.7B lands around 2.5 GB on disk yet still passes tool-call formatting.
- The 27B ternary variant adds a
qwen3reasoning parser for<think>-style traces.
MiniCPM5 · 3 aliases
OpenBMB's on-device 1B — strong sub-1 GB chat/agent model with a minicpm tool-call envelope and qwen3 reasoning. Ships a stock 4-bit and an OptiQ 4-bit quant.
parser: minicpm
| alias | hf repo | tool parser | reasoning | flags | context | AA index | get it |
|---|---|---|---|---|---|---|---|
| minicpm5-1b-4bit | openbmb/MiniCPM5-1B-MLX | minicpm | qwen3 | — | 128K | 11.9 | CDN |
| minicpm5-1b-optiq-4bit | mlx-community/MiniCPM5-1B-OptiQ-4bit | minicpm | qwen3 | — | 128K | 11.9 | CDN |
| minicpm5-2b-4bit | openbmb/MiniCPM5-2B-MLX | minicpm | qwen3 | — | 128K | — | CDN |
Notes & caveats
- 1B footprint — the smallest tool-call-capable text model in the registry.
- OptiQ 4-bit is a higher-quality quant than the stock 4-bit at the same size.
SmolLM3 · 1 alias
Hugging Face's small-but-mighty 3B — strong sub-4 GB chat model with long-context training.
parser: hermes
| alias | hf repo | tool parser | reasoning | flags | context | AA index | get it |
|---|---|---|---|---|---|---|---|
| smollm3-3b-4bit | mlx-community/SmolLM3-3B-4bit | hermes | qwen3 | spec | 64K | — | CDN |
Notes & caveats
- SmolLM3 3B 4-bit is a strong sub-4 GB chat model — see the notable-small-models callout for context.
Nanbeige · 1 alias
Small instruction-tuned checkpoint kept for research diversity.
parser: hermes
| alias | hf repo | tool parser | reasoning | flags | context | AA index | get it |
|---|---|---|---|---|---|---|---|
| nanbeige4.1-3b-4bit | mlx-community/Nanbeige4.1-3B-4bit | hermes | deepseek_r1 | spec | 256K | 11 | CDN |
One-off big checkpoints · 4 aliases
Checkpoints that earn a place in the registry without a vendor family around them. Kimi K2.6 (470 GB, DQ3_K_M-q8) is a frontier MoE for the 512 GB Mac Studio, served with its own kimi tool envelope and deepseek_r1 reasoning parser. North-Mini-Code is a Cohere2-MoE code model (18.5 GB 4-bit / 61 GB bf16) with a dedicated cohere_command4 reasoning parser so chain-of-thought stays out of the answer. MiMo-V2.6 Flash is Xiaomi's 309B / 15B-active MoE reasoner (experimental, 192 GB+ Macs, text-only) — measured on its model page.
parser: kimi
| alias | hf repo | tool parser | reasoning | flags | context | AA index | get it |
|---|---|---|---|---|---|---|---|
| kimi-k2.6 | mlx-community/Kimi-K2.6-mlx-DQ3_K_M-q8 | kimi | deepseek_r1 | moe | — | 45.1 | HF |
| mimo-v2.6-flash-4bit | Vontra/MiMo-V2.6-Flash-RL-MLX-4bit-MTP | qwen3_coder_xml | qwen3 | moe | — | — | HF |
| north-mini-code-4bit | mlx-community/North-Mini-Code-1.0-4bit | — | cohere_command4 | moe | 500,000 | 20.2 | CDN |
| north-mini-code-bf16 | mlx-community/North-Mini-Code-1.0-bf16 | — | cohere_command4 | moe | — | 20.2 | HF |
Notes & caveats
- kimi-k2.6 is the largest checkpoint in the registry — Ultra-class hardware only.
- north-mini-code-4bit is the practical pick; bf16 exists for quantization research.
Experimental community builds · 3 aliases
Community checkpoints served behind the experimental flag and a minimum-memory floor: G9v3-39A5B (MoE, 32 GB+), K2 Horizon 7B (own k2_horizon tool and reasoning parsers, 16 GB+) and NeoHorse 1 9B (18 GB+).
parser: —
| alias | hf repo | tool parser | reasoning | flags | context | AA index | get it |
|---|---|---|---|---|---|---|---|
| g9v3-39a5b-4bit | rapid-mlx/G9v3-39A5B-MLX-4bit | minicpm | qwen3 | moe | 128K | — | CDN |
| k2-horizon-7b-4bit | abenzerps/K2-Horizon-7B-MLX-4bit | k2_horizon | k2_horizon | text | 512K | — | CDN |
| neohorse-9b-4bit | rapid-mlx/NeoHorse-1-9B-MLX-4bit | hermes | qwen3 | text | 256K | — | CDN |
Nemotron · 4 aliases
Pruning-research checkpoint with hybrid cache — useful for studying compression-vs-quality trade-offs.
parser: hermes
| alias | hf repo | tool parser | reasoning | flags | context | AA index | get it |
|---|---|---|---|---|---|---|---|
| nemotron-3.5-lightning-30b-4bit | mlx-community/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-4bit | nemotron | qwen3 | hybrid · moe | 256K | — | CDN |
| nemotron-3.5-lightning-tensorfold | TensorFold/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-MLX-4bit | — | qwen3 | hybrid · moe | 256K | — | CDN |
| nemotron-30b-4bit | lmstudio-community/NVIDIA-Nemotron-3-Nano-30B-A3B-MLX-4bit | hermes | qwen3 | hybrid · moe | 256K | 14.5 | CDN |
| nemotron-labs-diffusion-3b-4bit | mlx-community/Nemotron-Labs-Diffusion-3B-4bit | — | — | spec | 256K | — | CDN |
Notes & caveats
- Nemotron 30B is a pruning-research checkpoint with hybrid cache.
Context is read from the config.json of the exact build each alias pulls; for an embedding model it is the most input tokens the engine embeds, which can be less than the config declares. AA index is the Artificial Analysis Intelligence Index for the base model at full precision with reasoning on — a property of the model, not a score for our quantised build.