Run Ternary Bonsai 1.7B on a Mac
A ternary (2-bit) model that fits in half a gigabyte on disk. The quality ceiling is what you'd expect at this size — the interesting part is how much of it survives 2-bit quantisation.
1.7 B ternary · 2-bit tiny-footprint experiments · embedded-style use
One command
$ rapid-mlx serve bonsai-1.7b-2bit
Weights (0.5 GB) download on first run; you get an OpenAI-compatible
endpoint at http://localhost:8000/v1. No rapid-mlx yet? It's one line —
curl -fsSL https://rapidmlx.com/install.sh | bash — or take the
desktop app.
Measured on real hardware
| Machine | Decode | First token | Peak memory | Cold boot | Weights |
|---|---|---|---|---|---|
| Mac mini M2 Pro · 32 GB | 169.0 tok/s | 0.23 s | 1.3 GB | 4.3 s | 0.5 GB |
Measured on rapid-mlx 0.12.10, 2026-08-11. Decode: median of 3 runs, temperature 0, 256-token saturating generation, engine-reported token counts, unique salt per request (no prefix-cache hits). TTFT: median of 3, short prompt. Boot: process spawn to first completed token. Peak RSS: 0.5 s sampling across the run.
Will it fit your Mac?
Peak resident memory measured 1.3 GB during a 256-token generation. The KV cache grows with context length, so treat that as a floor, not a ceiling. For the conservative install-default placement see the hardware tiers table; to compare against every model your RAM can hold, use the live picker.
Variants & alternatives
bonsai-27b-2bit— Bonsai 27B — the 27 B ternary that runs in 8.4 GB
FAQ
How much memory does Ternary Bonsai 1.7B need on a Mac?
Measured peak resident memory was 1.3 GB on rapid-mlx 0.12.10 during a 256-token generation (M2 Pro, 32 GB). The weights are 0.5 GB on disk. Longer contexts grow the KV cache beyond this, so leave headroom.
How fast is Ternary Bonsai 1.7B on Apple Silicon?
We measured 169.0 tokens/sec sustained decode and 0.23 s time-to-first-token on a Mac mini M2 Pro (32 GB), median of 3 runs at temperature 0.
How do I run Ternary Bonsai 1.7B locally?
Install rapid-mlx (curl -fsSL https://rapidmlx.com/install.sh | bash, or brew install rapid-mlx), then: rapid-mlx serve bonsai-1.7b-2bit — the weights download on first run and you get an OpenAI-compatible endpoint at localhost:8000/v1 that works with Cursor, Claude Code, Aider, and any OpenAI client.