Model pages · hot models · rapid-mlx 0.15.4

bonsai2-27b-2bit

Run Ternary Bonsai 2 27B on a Mac

Prism ML's second-generation ternary 27B — a full-size model in 8.6 GB on disk, and it takes image inputs. The most-excited-about text model on Hugging Face the week it landed (trending #1, GGUF builds, 2026-10-02); this is the MLX build, measured.

27 B ternary · 2-bit · vision24–32 GB Macs · maximum model per gigabyte · vision + text

Measured decode
17.7tok/s
on Mac mini M4 Pro · 48 GB · rapid-mlx 0.15.4 · 2026-10-02
Weights
8.6 GB
Peak memory · Metal
24.4 GB

One command

rapid-mlx serve bonsai2-27b-2bit

Weights (8.6 GB) download on first run; you get an OpenAI-compatible endpoint at http://localhost:8000/v1. No rapid-mlx yet? It's one line — curl -fsSL https://rapidmlx.com/install.sh | bash — or take the desktop app.

Point a client at it. OpenAI SDK: base_url="http://localhost:8000/v1", model="bonsai2-27b-2bit". Claude Code: ANTHROPIC_BASE_URL=http://localhost:8000 (no /v1 suffix) — see the Claude Code guide.

Measured on real hardware

Headline machine · Mac mini M4 Pro · 48 GB

Decode
17.7tok/s
First token · 2K prompt
23.4 s
Peak memory · Metal
24.4 GB
Weights on disk
8.6 GB

All measured machines

MachineDecodeFirst tokenPeak memoryCold bootWeights
Mac mini M4 Pro · 48 GB rapid-mlx 0.15.4 · 2026-10-02 raw data17.7 tok/s23.4 s24.4 GB—8.6 GB

Each row cites its own run; numbers are never copied across machines. Mac mini M4 Pro · 48 GB: rapid-mlx 0.15.4, 2026-10-02 — raw data + method — 4 concurrent streams: 16.8 tok/s aggregate.

Will it fit your Mac?

Peak memory · Metal
24.4 GB
Weights on disk
8.6 GB

Not tiered by the installer's policy — measured 24.4 GB peak on a 48 GB M4 Pro; treat 32 GB as the practical floor.

For the conservative install-default placement see the hardware tiers table; to compare against every model your RAM can hold, use the live picker.

Variants & profiles

  • bonsai-27b-2bit — Ternary Bonsai 27B (1st gen) — 13 GB measured peak, the 24 GB tier's smart pick
  • qwen3.8-27b-4bit — Qwen3.8 27B — 4-bit dense, about 2.7× the decode at 2× the memory

FAQ

How much memory does Ternary Bonsai 2 27B need on a Mac?

Not tiered by the installer's policy — measured 24.4 GB peak on a 48 GB M4 Pro; treat 32 GB as the practical floor.

How fast is Ternary Bonsai 2 27B on Apple Silicon?

We measured 17.7 on a Mac mini M4 Pro · 48 GB, rapid-mlx 0.15.4, 2026-10-02; peak memory 24.4 GB on that run. Every row on this page keeps its own chip, version and date — no number is copied across machines.

How do I run Ternary Bonsai 2 27B locally?

rapid-mlx serve bonsai2-27b-2bit — the weights download on first run and you get an OpenAI-compatible endpoint at localhost:8000/v1 that works with Cursor, Claude Code, Aider, and any OpenAI client.

Where next