Agent guide · rapid-mlx 0.15.5 · ← All agents

Pi Coding Agent

Pi (npm @earendil-works/pi-coding-agent) is a small, fast terminal coding agent. It reads custom OpenAI-compatible providers from ~/.pi/agent/models.json, so pointing it at rapid-mlx is one provider block.

Quick start

Needs: Pi (npm install -g @earendil-works/pi-coding-agent) and a running server (rapid-mlx serve qwen3.6-35b-4bit).

$ rapid-mlx agents pi --setup
$ pi --print --no-session "say hello"   # a reply = it works

A reply means Pi is using your local model (if pi has credentials for other providers too, add --provider rapid-mlx). rapid-mlx agents pi --test runs the full integration check.

Wire: /v1/chat/completions (pi's openai-completions API) · Setup: rapid-mlx agents pi --setup
Re-verified 2026-10-02 — pi 1.0.0 passed all six tasks of our agent harness plus a follow-up turn with this exact config, against qwen3.6-35b-4bit on rapid-mlx 0.15.4 (M4 Pro, 48 GB). Its prompt is small (four tools), so later turns reuse the prefix cache.

Install

$ npm install -g @earendil-works/pi-coding-agent

Setup

$ rapid-mlx serve qwen3.6-35b-4bit

# in another terminal
$ rapid-mlx agents pi --setup --dry-run   # preview only
$ rapid-mlx agents pi --setup             # preview, confirm, write

--setup reads the model id and context window from the running server. It then:

Your other providers are kept. Under the rapid-mlx provider, models are merged by id, so models you added there yourself are kept too. The block it writes:

// ~/.pi/agent/models.json
{
  "providers": {
    "rapid-mlx": {
      "baseUrl": "http://localhost:8000/v1",
      "api": "openai-completions",
      "apiKey": "not-needed",
      "models": [
        {
          "id": "qwen3.6-35b-4bit",
          "name": "qwen3.6-35b-4bit (Rapid-MLX)",
          "contextWindow": 262144,
          "maxTokens": 8192,
          "input": ["text"],
          "reasoning": false
        }
      ]
    }
  }
}

id and contextWindow come from the running server. The values above are for qwen3.6-35b-4bit.

Run

$ pi                                                      # interactive
$ pi --print --no-session "summarize this repo"             # one-shot
$ pi --print --continue "now add a test for it"             # follow-up in the same session

If rapid-mlx is the only provider pi has credentials for, no flags are needed. Otherwise see the first gotcha below. Recommended aliases (per rapid-mlx agents pi): qwen3.6-35b-4bit, qwen3.5-9b-4bit, qwen3-coder-30b-4bit. rapid-mlx start pi starts a recommended model and then runs the same setup.

Gotchas

See also