Pi Coding Agent with a Local Model on Mac
Pi is a small, fast terminal coding agent. Point it at a model running on your own Mac and it edits code with no cloud account and no per-token bill.
Pi is a minimal terminal coding agent: a short system prompt, a handful of tools, and any model you give it. Custom providers live in one file, ~/.pi/agent/models.json, and Rapid-MLX can write that file for you.
The whole thing
- Install Rapid-MLX and serve a model.
rapid-mlx agents pi --setupadds arapid-mlxprovider to pi.pi --provider rapid-mlx --model <alias>.
1 · Install Rapid-MLX and pi
curl -fsSL https://rapidmlx.com/install.sh | bash npm install -g @earendil-works/pi-coding-agent
2 · Serve a model
rapid-mlx serve qwen3.5-4b-4bit
That's the engine's pick for a 16 GB Mac and the model we verified with. rapid-mlx recipe prints the pick for your memory; Best local LLMs by Mac shows measured speeds per chip. Keep the server running.
3 · Point pi at it
rapid-mlx agents pi --setup
You get a diff preview and a confirmation prompt, and any existing models.json is backed up first. The provider it adds uses pi's openai-completions API with http://localhost:8000/v1 as the base URL and lists the served model with its context window.
4 · Run it
Start pi interactively with the local provider, or give it one task:
# inside your project folder pi --provider rapid-mlx --model qwen3.5-4b-4bit \ -p "Add a function add(a, b) to hello.py that returns a + b. Edit the file."
From our run:
I've added the `add(a, b)` function to `hello.py` that returns `a + b`.
and hello.py gained the function.
Verified on 2026-10-06 with rapid-mlx 0.15.6, pi 1.0.4 and qwen3.5-4b-4bit on a Mac mini (M2 Pro, 32 GB): after rapid-mlx agents pi --setup, the pi -p task edited the file through pi's tools in about 8 seconds.
Frequently asked questions
Can pi use a local model?
Yes. Pi reads custom providers from ~/.pi/agent/models.json. rapid-mlx agents pi --setup adds one that points at your Rapid-MLX server, then pi --provider rapid-mlx uses it.
Does pi need an API key for a local model?
No. The provider entry uses a placeholder key, and a plain rapid-mlx serve doesn't check it. If you start the server with RAPID_MLX_API_KEY set, export the same variable before running setup: the entry then reads the key from that variable instead.
Which model works best with pi?
Pi's prompt is small, so even small models respond quickly. For real coding work, use the biggest model rapid-mlx recipe suggests for your Mac.
Where to go next
- Pi reference: what setup writes and how to roll it back.
- Best local LLMs by Mac: measured speed for your chip and memory.
- Rapid-MLX vs mlx-lm: the same weights, measured task by task.
- All agent guides: Claude Code, Codex CLI, OpenCode and more.
rapid-mlx serve <alias> serves any model on
localhost:8000/v1.
curl -fsSL https://rapidmlx.com/install.sh | bash
New models and speedups, in your inbox
A short note whenever Rapid-MLX gets faster or adds models worth running on your Mac. No spam — unsubscribe anytime.