Pi Coding Agent
Pi (npm
@earendil-works/pi-coding-agent) is a small, fast terminal
coding agent. It reads custom OpenAI-compatible providers from
~/.pi/agent/models.json, so pointing it at rapid-mlx is
one provider block.
Quick start
Needs: Pi (npm install -g @earendil-works/pi-coding-agent) and a running server (rapid-mlx serve qwen3.6-35b-4bit).
$ rapid-mlx agents pi --setup $ pi --print --no-session "say hello" # a reply = it works
A reply means Pi is using your local model (if pi has credentials for other providers too, add --provider rapid-mlx). rapid-mlx agents pi --test runs the full integration check.
/v1/chat/completions (pi's
openai-completions API) ·
Setup: rapid-mlx agents pi --setup
Re-verified 2026-10-02 — pi 1.0.0 passed all six tasks of our agent harness plus a follow-up turn with this exact config, against
qwen3.6-35b-4bit on
rapid-mlx 0.15.4 (M4 Pro, 48 GB). Its prompt is small (four
tools), so later turns reuse the prefix cache.
Install
$ npm install -g @earendil-works/pi-coding-agent
Setup
$ rapid-mlx serve qwen3.6-35b-4bit # in another terminal $ rapid-mlx agents pi --setup --dry-run # preview only $ rapid-mlx agents pi --setup # preview, confirm, write
--setup reads the model id and context window from the
running server. It then:
- prints an exact diff of
~/.pi/agent/models.json(or$PI_CODING_AGENT_DIR/models.json) - asks
Apply this configuration? [y/N]. Pass--yesin scripts, because a non-interactive run without it writes nothing. - copies an existing file to
models.json.bak.<unix-time> - writes atomically, owner-only (
0600). Ifmodels.jsonis a symlink, it updates the file the link points to. - refuses to write if the file changed between preview and confirmation
- checks the server's
/healthand/v1/models(skip this with--no-check)
Your other providers are kept. Under the rapid-mlx
provider, models are merged by id, so models you added there
yourself are kept too. The block it writes:
// ~/.pi/agent/models.json { "providers": { "rapid-mlx": { "baseUrl": "http://localhost:8000/v1", "api": "openai-completions", "apiKey": "not-needed", "models": [ { "id": "qwen3.6-35b-4bit", "name": "qwen3.6-35b-4bit (Rapid-MLX)", "contextWindow": 262144, "maxTokens": 8192, "input": ["text"], "reasoning": false } ] } } }
id and contextWindow come from the running
server. The values above are for qwen3.6-35b-4bit.
Run
$ pi # interactive $ pi --print --no-session "summarize this repo" # one-shot $ pi --print --continue "now add a test for it" # follow-up in the same session
If rapid-mlx is the only provider pi has credentials for, no flags are
needed. Otherwise see the first gotcha below. Recommended aliases (per
rapid-mlx agents pi): qwen3.6-35b-4bit,
qwen3.5-9b-4bit, qwen3-coder-30b-4bit.
rapid-mlx start pi starts a recommended model and then
runs the same setup.
Gotchas
-
Other providers can win. If pi has other providers authenticated
(its
auth.jsonor provider env keys), it may default to one of them. Pick the local model with/model(Ctrl+S saves it as the default), or pass--provider rapid-mlx --model <id>. -
"reasoning": falseis the setting we verified. pi then sends no reasoning preference, and rapid-mlx turns thinking off on tool turns."reasoning": true(so pi's/thinkingdrivesreasoning_effort) has not been verified. -
Rapid-MLX desktop app. The app starts its server with a
per-launch API key. Set
"apiKey": "$RAPID_MLX_API_KEY"(pi expands$NAMEinmodels.json) and export the key. -
pi's own telemetry.
PI_TELEMETRY=0turns off pi's install/update telemetry and provider attribution headers.PI_OFFLINE=1stops its catalog refreshes.