Claude Code
Claude Code is
Anthropic's official CLI. It speaks the Anthropic SDK wire
(/v1/messages), not OpenAI's. rapid-mlx exposes
/v1/messages as a first-class route so a one-liner
env var flip is all you need.
Quick start
Needs: Claude Code (npm install -g @anthropic-ai/claude-code) and a running server (rapid-mlx serve qwen3.6-35b-4bit).
$ rapid-mlx agents claude-code --setup $ claude -p "say hello" # a reply = it works
A reply from claude -p means Claude Code is talking to your local model. rapid-mlx agents claude-code --test runs the full integration check.
/v1/messages (Anthropic SDK) ·
Setup: rapid-mlx agents claude-code --setup ·
Matrix cell:
✅ ✅ ✅ ✅
(Qwen 3.6 / Gemma 4 / DeepSeek / gpt-oss on 2026-07-07 pilot).
Re-verified 2026-08-01 — a real end-to-end bug fix: Claude Code pointed at the local server via
ANTHROPIC_BASE_URL and asked to find and fix a failing factorial, which it did (agent_smoke, claude 2.1.218). Plus the Anthropic SDK deep flow — plain, system prompt, multi-turn, streaming, tool use — 5/5. On rapid-mlx 0.11.5: the fix against gpt-oss-20b-mxfp4-q8, the SDK flow against qwen3.5-9b-4bit, both on an M2 Pro Mac mini (32 GB).
New to this? The 5-minute walkthrough on the blog is the fastest path from zero to a working local agent — this page is the complete reference behind it.
Install
Claude Code is upstream Anthropic — install their vendor CLI:
$ npm install -g @anthropic-ai/claude-code # or download the binary from anthropic.com/claude-code
Config
One-shot, with a server running:
$ rapid-mlx agents claude-code --setup --dry-run # preview the diff $ rapid-mlx agents claude-code --setup # confirm and write # or, same file: $ rapid-mlx launch claude-code --model qwen3.6-35b-4bit
Both write the env block of
~/.claude/settings.json. Every other key
(permissions, MCP servers, …) is kept, and the old file is
copied to settings.json.bak.<unix-time> first.
agents --setup also shows the exact diff, asks before
writing (--yes in scripts), and checks the server afterwards.
// ~/.claude/settings.json { "env": { "ANTHROPIC_BASE_URL": "http://localhost:8000", "ANTHROPIC_API_KEY": "sk-noop", "ANTHROPIC_AUTH_TOKEN": "", "ANTHROPIC_MODEL": "qwen3.6-35b-4bit", "CLAUDE_CODE_MAX_CONTEXT_TOKENS": "262144" } }
ANTHROPIC_MODEL and CLAUDE_CODE_MAX_CONTEXT_TOKENS
come from the running server (the values above are for
qwen3.6-35b-4bit), so Claude Code uses the model's real
context window instead of guessing one for a model id it doesn't know.
ANTHROPIC_AUTH_TOKEN is blanked so that a token inherited
from a proxy or account switcher can't conflict with the key.
Auth. A bare rapid-mlx serve needs no key, so
ANTHROPIC_API_KEY is a placeholder. For a keyed server
(--api-key, RAPID_MLX_API_KEY, or the desktop
app), export RAPID_MLX_API_KEY before running setup: its
value is written as ANTHROPIC_API_KEY, because Claude Code
reads the key from this file. The setup diff shows credential values as
<redacted>.
To try it in one shell without touching the settings file:
Anthropic's SDK honors the ANTHROPIC_BASE_URL env var.
Set it to the root of your rapid-mlx server (no /v1
suffix — the SDK appends the path itself):
$ export ANTHROPIC_BASE_URL=http://localhost:8000 $ export ANTHROPIC_API_KEY=not-needed $ claude # or inline: $ ANTHROPIC_BASE_URL=http://localhost:8000 ANTHROPIC_API_KEY=not-needed claude
Anthropic SDK (Python)
If you're driving the same wire from a script, the Python SDK
snippet is (verified by
tests/integrations/test_anthropic_sdk.py):
$ pip install anthropic
# Note the base_url is the ROOT, not /v1 — the SDK appends /v1/messages itself. # The test file strips /v1 with ``_BASE.rstrip('/').removesuffix('/v1')``. from anthropic import Anthropic client = Anthropic( base_url="http://localhost:8000", api_key="not-needed", ) message = client.messages.create( model="default", max_tokens=256, messages=[{"role": "user", "content": "Say hello"}], ) # Reasoning models emit a ThinkingBlock first — walk content and find the # first block with type == "text" (see _first_text in the test file). for block in message.content: if getattr(block, "type", None) == "text": print(block.text) break
Run
$ rapid-mlx serve qwen3.6-35b-4bit # in another shell: $ ANTHROPIC_BASE_URL=http://localhost:8000 ANTHROPIC_API_KEY=not-needed claude # headless one-shot, and a follow-up in the same session: $ claude -p "find and fix the failing test" $ claude -p --continue "now add a regression test"
Prefix cache across turns
Claude Code (2.1.287 and later) adds a small system message to every
request that holds only a running token counter
(<total_tokens>… tokens left</total_tokens>).
The server drops that counter-only message, but only in the position
Claude Code puts it; it carries no instruction. Its changing value
therefore never lands at the front of the prompt, and later turns of a
session reuse the prefix cache instead of re-processing Claude Code's
whole system prompt and tool list. No server flag is needed.
serve auto-configures the
tool-call parser from the model's alias profile and turns on
grammar-constrained tool calling by default, so tool calls come back
guaranteed-parseable. Verified end-to-end: a bare serve +
an Anthropic /v1/messages request with a tool returns a
tool_use block.
Gotchas
-
Base URL is root, not
/v1. The Anthropic SDK appends/v1/messagesitself.ANTHROPIC_BASE_URL=http://localhost:8000/v1would make it POST to/v1/v1/messagesand 404. -
Reasoning models emit a
ThinkingBlockfirst. The legacyresponse.content[0].textlookup raisesAttributeError. Walkresponse.contentand pick the first block whosetype == "text". -
Tool calls use Anthropic's
input_schema, not OpenAI'sparameters. Thetest_anthropic_sdk.pysmoke posts aninput_schemawith aget_weather(city: str)function and asserts atool_useblock comes back. -
Env vars are per shell. Exported variables only last for the
current shell.
--setup/launch claude-codeput them in~/.claude/settings.jsonso they persist. -
Model-catalog warning. Without
CLAUDE_CODE_MAX_CONTEXT_TOKENS(for example when setup ran while no server was reachable), Claude Code may print that the model "isn't described by this version's model catalog" and guess a context window. Re-runrapid-mlx agents claude-code --setupwith the server up to fill it in.
Empirical
The claude-code row of the integration matrix
is ✅ across all four Tier-1 families (Qwen 3.6, Gemma 4,
DeepSeek R1-Distill, gpt-oss) on the 2026-07-07 pilot. The dedicated
Anthropic SDK deep flow at
test_anthropic_sdk.py
covers plain, system prompt, multi-turn, streaming, and tool use.