Pi Coding Agent with a Local Model on Mac

Pi is a small, fast terminal coding agent. Point it at a model running on your own Mac and it edits code with no cloud account and no per-token bill.

Pi is a minimal terminal coding agent: a short system prompt, a handful of tools, and any model you give it. Custom providers live in one file, ~/.pi/agent/models.json, and Rapid-MLX can write that file for you.

The whole thing

  1. Install Rapid-MLX and serve a model.
  2. rapid-mlx agents pi --setup adds a rapid-mlx provider to pi.
  3. pi --provider rapid-mlx --model <alias>.

1 · Install Rapid-MLX and pi

curl -fsSL https://rapidmlx.com/install.sh | bash
npm install -g @earendil-works/pi-coding-agent

2 · Serve a model

rapid-mlx serve qwen3.5-4b-4bit

That's the engine's pick for a 16 GB Mac and the model we verified with. rapid-mlx recipe prints the pick for your memory; Best local LLMs by Mac shows measured speeds per chip. Keep the server running.

3 · Point pi at it

rapid-mlx agents pi --setup

You get a diff preview and a confirmation prompt, and any existing models.json is backed up first. The provider it adds uses pi's openai-completions API with http://localhost:8000/v1 as the base URL and lists the served model with its context window.

4 · Run it

Start pi interactively with the local provider, or give it one task:

# inside your project folder
pi --provider rapid-mlx --model qwen3.5-4b-4bit \
  -p "Add a function add(a, b) to hello.py that returns a + b. Edit the file."

From our run:

I've added the `add(a, b)` function to `hello.py` that returns `a + b`.

and hello.py gained the function.

Verified on 2026-10-06 with rapid-mlx 0.15.6, pi 1.0.4 and qwen3.5-4b-4bit on a Mac mini (M2 Pro, 32 GB): after rapid-mlx agents pi --setup, the pi -p task edited the file through pi's tools in about 8 seconds.

Frequently asked questions

Can pi use a local model?

Yes. Pi reads custom providers from ~/.pi/agent/models.json. rapid-mlx agents pi --setup adds one that points at your Rapid-MLX server, then pi --provider rapid-mlx uses it.

Does pi need an API key for a local model?

No. The provider entry uses a placeholder key, and a plain rapid-mlx serve doesn't check it. If you start the server with RAPID_MLX_API_KEY set, export the same variable before running setup: the entry then reads the key from that variable instead.

Which model works best with pi?

Pi's prompt is small, so even small models respond quickly. For real coding work, use the biggest model rapid-mlx recipe suggests for your Mac.

Where to go next


Run this yourself. Rapid-MLX is an open-source, OpenAI- and Anthropic-compatible inference server for Apple Silicon. One command installs it, then rapid-mlx serve <alias> serves any model on localhost:8000/v1.
curl -fsSL https://rapidmlx.com/install.sh | bash

New models and speedups, in your inbox

A short note whenever Rapid-MLX gets faster or adds models worth running on your Mac. No spam — unsubscribe anytime.