bonsai2-27b-2bit
Run Ternary Bonsai 2 27B on a Mac
Prism ML's second-generation ternary 27B — a full-size model in 8.6 GB on disk, and it takes image inputs. The most-excited-about text model on Hugging Face the week it landed (trending #1, GGUF builds, 2026-10-02); this is the MLX build, measured.
27 B ternary · 2-bit · vision24–32 GB Macs · maximum model per gigabyte · vision + text
One command
rapid-mlx serve bonsai2-27b-2bit
Weights (8.6 GB) download on first run; you get an OpenAI-compatible
endpoint at http://localhost:8000/v1. No rapid-mlx yet? It's one line —
curl -fsSL https://rapidmlx.com/install.sh | bash — or take the
desktop app.
Point a client at it. OpenAI SDK: base_url="http://localhost:8000/v1",
model="bonsai2-27b-2bit". Claude Code:
ANTHROPIC_BASE_URL=http://localhost:8000 (no /v1 suffix) —
see the Claude Code guide.
Measured on real hardware
Headline machine · Mac mini M4 Pro · 48 GB
All measured machines
| Machine | Decode | First token | Peak memory | Cold boot | Weights |
|---|---|---|---|---|---|
| Mac mini M4 Pro · 48 GB rapid-mlx 0.15.4 · 2026-10-02 raw data | 17.7 tok/s | 23.4 s | 24.4 GB | — | 8.6 GB |
Each row cites its own run; numbers are never copied across machines. Mac mini M4 Pro · 48 GB: rapid-mlx 0.15.4, 2026-10-02 — raw data + method — 4 concurrent streams: 16.8 tok/s aggregate.
Will it fit your Mac?
Not tiered by the installer's policy — measured 24.4 GB peak on a 48 GB M4 Pro; treat 32 GB as the practical floor.
For the conservative install-default placement see the hardware tiers table; to compare against every model your RAM can hold, use the live picker.
Variants & profiles
bonsai-27b-2bit— Ternary Bonsai 27B (1st gen) — 13 GB measured peak, the 24 GB tier's smart pickqwen3.8-27b-4bit— Qwen3.8 27B — 4-bit dense, about 2.7× the decode at 2× the memory
FAQ
How much memory does Ternary Bonsai 2 27B need on a Mac?
Not tiered by the installer's policy — measured 24.4 GB peak on a 48 GB M4 Pro; treat 32 GB as the practical floor.
How fast is Ternary Bonsai 2 27B on Apple Silicon?
We measured 17.7 on a Mac mini M4 Pro · 48 GB, rapid-mlx 0.15.4, 2026-10-02; peak memory 24.4 GB on that run. Every row on this page keeps its own chip, version and date — no number is copied across machines.
How do I run Ternary Bonsai 2 27B locally?
rapid-mlx serve bonsai2-27b-2bit — the weights download on first run and you get an OpenAI-compatible endpoint at localhost:8000/v1 that works with Cursor, Claude Code, Aider, and any OpenAI client.