Agent guide · rapid-mlx 0.15.7 · ← Back to README

Codex CLI

OpenAI Codex CLI (Rust, 0.136.0+) is the official OpenAI agent CLI. It speaks /v1/responses — the "responses" API, not the chat-completions one. rapid-mlx exposes /v1/responses as a first-class route so Codex talks to your local server the same way it talks to api.openai.com.

Quick start

Needs: Codex CLI (brew install codex) and a running server (rapid-mlx serve qwen3.6-35b-4bit).

$ rapid-mlx agents codex --setup
$ codex exec --skip-git-repo-check "say hello"   # a reply = it works

A reply means Codex is using your local model. rapid-mlx agents codex --test runs the full integration check.

TIER-1 AGENT   Wire: /v1/responses · Setup: rapid-mlx agents codex --setup · Matrix cell: ✅ ✅ ✅ ✅ (Qwen 3.6 / Gemma 4 / DeepSeek / gpt-oss on 2026-07-07 pilot).
Re-verified 2026-08-01 — a /v1/responses wire smoke, on rapid-mlx 0.11.5, Qwen 3.5-9B-4bit on an M2 Pro Mac mini (32 GB) — Qwen family only; the other three families were not re-run in this pass.

Install

Homebrew or npm — Codex CLI is upstream OpenAI, install the vendor binary:

$ brew install codex
# or
$ npm install -g @openai/codex

Config

With a server running, rapid-mlx agents codex --setup writes ~/.codex/config.toml (or $CODEX_HOME/config.toml). It also writes rapid-mlx-model-catalog.json next to it, a local model catalog. This stops Codex from using its flagship fallback metadata on the first turn, before its remote catalog request completes. An existing config.toml is merged: your other keys (approval_policy, [mcp_servers.*], [projects.*] …) are kept, but TOML comments are not, and no backup is made. Setup checks the server before it writes and refuses if the server isn't reachable (--no-check writes the config offline); a config that is already current is left untouched. --dry-run previews without writing.

Setup then prints a Saved Codex defaults report, read back from the files it wrote: model, provider, endpoint and catalog. The catalog must contain the saved model's exact id. A passed connection check means the server answered its health and model-list checks, not that a coding task has run. Project settings, profiles and codex --model / --config can override these defaults; if Codex reports an unknown model or generic model metadata, compare its model id with the one setup printed. A full Hugging Face repo name and its short alias are different catalog ids even for the same weights, so to use another id rerun setup with --model set to it.

# ~/.codex/config.toml
model = "qwen3.6-35b-4bit"
model_provider = "rapid-mlx"
model_context_window = 262144
model_catalog_json = "/Users/you/.codex/rapid-mlx-model-catalog.json"
sandbox_mode = "workspace-write"

[model_providers.rapid-mlx]
name = "Rapid-MLX (local)"
base_url = "http://localhost:8000/v1"

model and model_context_window come from the running server. The values above are for qwen3.6-35b-4bit. base_url must include /v1. sandbox_mode = "workspace-write" lets codex exec edit files in the workspace; if your config already sets sandbox_mode, setup keeps your value.

Auth. A bare rapid-mlx serve needs no key. If the server has one (--api-key, RAPID_MLX_API_KEY, or the Rapid-MLX desktop app, which mints a key on every launch), export RAPID_MLX_API_KEY before running setup: setup then adds env_key = "RAPID_MLX_API_KEY" to the provider block (the key itself is not written to the file). Keep it exported in the shell you run Codex from — Codex errors out if env_key names a variable that isn't set. Re-running setup without the variable against an unkeyed server removes that line again.

Run

$ rapid-mlx serve qwen3.6-35b-4bit
# in another shell:
$ codex
# headless one-shot (outside a trusted git repo, add --skip-git-repo-check):
$ codex exec --skip-git-repo-check "read main.py and summarize what it does"

Recommended aliases (per rapid-mlx agents codex): qwen3.6-35b-4bit, qwen3.5-9b-4bit, qwen3-coder-30b-4bit.

Gotchas

Empirical

The codex-cli row of the integration matrix is ✅ across all four Tier-1 families (Qwen 3.6, Gemma 4, DeepSeek R1-Distill, gpt-oss) on the 2026-07-07 pilot. The matrix cell posts a minimal /v1/responses envelope, verifies the SSE-shaped output body is non-empty, and asserts no <think> or <|channel|>analysis leak in the emitted text. Source of truth for this page: tests/integrations/test_agents_matrix.py::TestCodexCLI and the CLI output of rapid-mlx agents codex.

See also