Model pages · rapid-mlx 0.15.3

bonsai-27b-2bit

Run Ternary Bonsai 27B on a Mac

A 27 B model in 8.4 GB of resident memory — ternary 2-bit quantisation makes a 27 B fit where a 4-bit 12B sits. Decode is slower per token than MoE peers, but the quality-per-gigabyte is the point.

27 B ternary · 2-bitbiggest model that fits on a 16 GB Mac

256K context

Measured decode
17.4tok/s
on Mac mini M2 Pro · 32 GB
Weights
7.9 GB
Peak memory
8.4 GB

One command

rapid-mlx serve bonsai-27b-2bit

Weights (7.9 GB) download on first run; you get an OpenAI-compatible endpoint at http://localhost:8000/v1. No rapid-mlx yet? It's one line — curl -fsSL https://rapidmlx.com/install.sh | bash — or take the desktop app.

Measured on real hardware

Headline machine · Mac mini M2 Pro · 32 GB

Decode
17.4tok/s
First token
1.48s
Peak memory
8.4GB
Cold boot
9.3s

All measured machines

MachineDecodeFirst tokenPeak memoryCold bootWeights
Mac mini M2 Pro · 32 GB17.4 tok/s1.48 s8.4 GB9.3 s7.9 GB

Measured on rapid-mlx 0.12.10, 2026-08-11. Decode: median of 3 runs, temperature 0, 256-token saturating generation, engine-reported token counts, unique salt per request (no prefix-cache hits). TTFT: median of 3, short prompt. Boot: process spawn to first completed token. Peak RSS: 0.5 s sampling across the run.

Will it fit your Mac?

Measured peak memory
8.4 GB
Weights on disk
7.9 GB

Peak resident memory measured 8.4 GB during a 256-token generation. The KV cache grows with context length, so treat that as a floor, not a ceiling. For the conservative install-default placement see the hardware tiers table; to compare against every model your RAM can hold, use the live picker.

Variants & alternatives

  • gemma-4-26b-4bit — Gemma 4 26B MoE — similar class, 3× the decode speed, more memory

FAQ

How much memory does Ternary Bonsai 27B need on a Mac?

Measured peak resident memory was 8.4 GB on rapid-mlx 0.12.10 during a 256-token generation (M2 Pro, 32 GB). The weights are 7.9 GB on disk. Longer contexts grow the KV cache beyond this, so leave headroom.

How fast is Ternary Bonsai 27B on Apple Silicon?

We measured 17.4 tokens/sec sustained decode and 1.48 s time-to-first-token on a Mac mini M2 Pro (32 GB), median of 3 runs at temperature 0.

How do I run Ternary Bonsai 27B locally?

Install rapid-mlx (curl -fsSL https://rapidmlx.com/install.sh | bash, or brew install rapid-mlx), then: rapid-mlx serve bonsai-27b-2bit — the weights download on first run and you get an OpenAI-compatible endpoint at localhost:8000/v1 that works with Cursor, Claude Code, Aider, and any OpenAI client.

Where next