Agent guide · rapid-mlx 0.15.7 · ← Back to README

Claude Code

Claude Code is Anthropic's official CLI. It speaks the Anthropic SDK wire (/v1/messages), not OpenAI's. rapid-mlx exposes /v1/messages as a first-class route so a one-liner env var flip is all you need.

Quick start

Needs: Claude Code (npm install -g @anthropic-ai/claude-code) and a running server (rapid-mlx serve qwen3.6-35b-4bit).

$ rapid-mlx agents claude-code --setup
$ claude -p "say hello"            # a reply = it works

A reply from claude -p means Claude Code is talking to your local model. rapid-mlx agents claude-code --test runs the full integration check.

TIER-1 AGENT   Wire: /v1/messages (Anthropic SDK) · Setup: rapid-mlx agents claude-code --setup · Matrix cell: ✅ ✅ ✅ ✅ (Qwen 3.6 / Gemma 4 / DeepSeek / gpt-oss on 2026-07-07 pilot).
Re-verified 2026-08-01 — a real end-to-end bug fix: Claude Code pointed at the local server via ANTHROPIC_BASE_URL and asked to find and fix a failing factorial, which it did (agent_smoke, claude 2.1.218). Plus the Anthropic SDK deep flow — plain, system prompt, multi-turn, streaming, tool use — 5/5. On rapid-mlx 0.11.5: the fix against gpt-oss-20b-mxfp4-q8, the SDK flow against qwen3.5-9b-4bit, both on an M2 Pro Mac mini (32 GB).

New to this? The 5-minute walkthrough on the blog is the fastest path from zero to a working local agent — this page is the complete reference behind it.

Install

Claude Code is upstream Anthropic — install their vendor CLI:

$ npm install -g @anthropic-ai/claude-code
# or download the binary from anthropic.com/claude-code

Config

One-shot, with a server running:

$ rapid-mlx agents claude-code --setup --dry-run   # preview the diff
$ rapid-mlx agents claude-code --setup             # confirm and write
# or, same file:
$ rapid-mlx launch claude-code --model qwen3.6-35b-4bit

Both write the env block of ~/.claude/settings.json. Every other key (permissions, MCP servers, …) is kept, and the old file is copied to settings.json.bak.<unix-time> first. agents --setup also shows the exact diff, asks before writing (--yes in scripts), and checks the server afterwards.

// ~/.claude/settings.json
{
  "env": {
    "ANTHROPIC_BASE_URL": "http://localhost:8000",
    "ANTHROPIC_API_KEY": "sk-noop",
    "ANTHROPIC_AUTH_TOKEN": "",
    "ANTHROPIC_MODEL": "qwen3.6-35b-4bit",
    "CLAUDE_CODE_MAX_CONTEXT_TOKENS": "262144"
  }
}

ANTHROPIC_MODEL and CLAUDE_CODE_MAX_CONTEXT_TOKENS come from the running server (the values above are for qwen3.6-35b-4bit), so Claude Code uses the model's real context window instead of guessing one for a model id it doesn't know. ANTHROPIC_AUTH_TOKEN is blanked so that a token inherited from a proxy or account switcher can't conflict with the key.

Auth. A bare rapid-mlx serve needs no key, so ANTHROPIC_API_KEY is a placeholder. For a keyed server (--api-key, RAPID_MLX_API_KEY, or the desktop app), export RAPID_MLX_API_KEY before running setup: its value is written as ANTHROPIC_API_KEY, because Claude Code reads the key from this file. The setup diff shows credential values as <redacted>.

To try it in one shell without touching the settings file: Anthropic's SDK honors the ANTHROPIC_BASE_URL env var. Set it to the root of your rapid-mlx server (no /v1 suffix — the SDK appends the path itself):

$ export ANTHROPIC_BASE_URL=http://localhost:8000
$ export ANTHROPIC_API_KEY=not-needed
$ claude
# or inline:
$ ANTHROPIC_BASE_URL=http://localhost:8000 ANTHROPIC_API_KEY=not-needed claude

Anthropic SDK (Python)

If you're driving the same wire from a script, the Python SDK snippet is (verified by tests/integrations/test_anthropic_sdk.py):

$ pip install anthropic
# Note the base_url is the ROOT, not /v1 — the SDK appends /v1/messages itself.
# The test file strips /v1 with ``_BASE.rstrip('/').removesuffix('/v1')``.
from anthropic import Anthropic

client = Anthropic(
    base_url="http://localhost:8000",
    api_key="not-needed",
)

message = client.messages.create(
    model="default",
    max_tokens=256,
    messages=[{"role": "user", "content": "Say hello"}],
)
# Reasoning models emit a ThinkingBlock first — walk content and find the
# first block with type == "text" (see _first_text in the test file).
for block in message.content:
    if getattr(block, "type", None) == "text":
        print(block.text)
        break

Run

$ rapid-mlx serve qwen3.6-35b-4bit
# in another shell:
$ ANTHROPIC_BASE_URL=http://localhost:8000 ANTHROPIC_API_KEY=not-needed claude
# headless one-shot, and a follow-up in the same session:
$ claude -p "find and fix the failing test"
$ claude -p --continue "now add a regression test"

Prefix cache across turns

Claude Code (2.1.287 and later) adds a small system message to every request that holds only a running token counter (<total_tokens>… tokens left</total_tokens>). The server drops that counter-only message, but only in the position Claude Code puts it; it carries no instruction. Its changing value therefore never lands at the front of the prompt, and later turns of a session reuse the prefix cache instead of re-processing Claude Code's whole system prompt and tool list. No server flag is needed.

No tool flags needed. serve auto-configures the tool-call parser from the model's alias profile and turns on grammar-constrained tool calling by default, so tool calls come back guaranteed-parseable. Verified end-to-end: a bare serve + an Anthropic /v1/messages request with a tool returns a tool_use block.

Gotchas

Empirical

The claude-code row of the integration matrix is ✅ across all four Tier-1 families (Qwen 3.6, Gemma 4, DeepSeek R1-Distill, gpt-oss) on the 2026-07-07 pilot. The dedicated Anthropic SDK deep flow at test_anthropic_sdk.py covers plain, system prompt, multi-turn, streaming, and tool use.

See also