Codex CLI
OpenAI Codex CLI
(Rust, 0.136.0+) is the official OpenAI agent CLI.
It speaks /v1/responses — the "responses" API, not the
chat-completions one. rapid-mlx exposes /v1/responses
as a first-class route so Codex talks to your local server the same
way it talks to api.openai.com.
Quick start
Needs: Codex CLI (brew install codex) and a running server (rapid-mlx serve qwen3.6-35b-4bit).
$ rapid-mlx agents codex --setup $ codex exec --skip-git-repo-check "say hello" # a reply = it works
A reply means Codex is using your local model. rapid-mlx agents codex --test runs the full integration check.
/v1/responses ·
Setup: rapid-mlx agents codex --setup ·
Matrix cell:
✅ ✅ ✅ ✅
(Qwen 3.6 / Gemma 4 / DeepSeek / gpt-oss on 2026-07-07 pilot).
Re-verified 2026-08-01 — a
/v1/responses wire smoke, on rapid-mlx 0.11.5, Qwen 3.5-9B-4bit on an M2 Pro Mac mini (32 GB) — Qwen family only; the other three families were not re-run in this pass.
Install
Homebrew or npm — Codex CLI is upstream OpenAI, install the vendor binary:
$ brew install codex # or $ npm install -g @openai/codex
Config
With a server running, rapid-mlx agents codex --setup writes
~/.codex/config.toml (or $CODEX_HOME/config.toml).
It also writes rapid-mlx-model-catalog.json next to it, a
local model catalog. This stops Codex from using its flagship fallback
metadata on the first turn, before its remote catalog request
completes. An existing config.toml is merged: your other
keys (approval_policy, [mcp_servers.*],
[projects.*] …) are kept, but TOML comments are not, and no
backup is made. Setup checks the server before it writes and refuses if
the server isn't reachable (--no-check writes the config
offline); a config that is already current is left untouched.
--dry-run previews without writing.
Setup then prints a Saved Codex defaults report, read back from
the files it wrote: model, provider, endpoint and catalog. The catalog
must contain the saved model's exact id. A passed connection check
means the server answered its health and model-list checks, not that a
coding task has run. Project settings, profiles and
codex --model / --config can override these
defaults; if Codex reports an unknown model or generic model metadata,
compare its model id with the one setup printed. A full Hugging Face
repo name and its short alias are different catalog ids even for the
same weights, so to use another id rerun setup with
--model set to it.
# ~/.codex/config.toml model = "qwen3.6-35b-4bit" model_provider = "rapid-mlx" model_context_window = 262144 model_catalog_json = "/Users/you/.codex/rapid-mlx-model-catalog.json" sandbox_mode = "workspace-write" [model_providers.rapid-mlx] name = "Rapid-MLX (local)" base_url = "http://localhost:8000/v1"
model and model_context_window come from the
running server. The values above are for qwen3.6-35b-4bit.
base_url must include /v1.
sandbox_mode = "workspace-write" lets codex exec
edit files in the workspace; if your config already sets
sandbox_mode, setup keeps your value.
Auth. A bare rapid-mlx serve needs no key. If the
server has one (--api-key, RAPID_MLX_API_KEY,
or the Rapid-MLX desktop app, which mints a key on every launch),
export RAPID_MLX_API_KEY before running setup: setup then
adds env_key = "RAPID_MLX_API_KEY" to the provider block
(the key itself is not written to the file). Keep it exported in the
shell you run Codex from — Codex errors out if env_key
names a variable that isn't set. Re-running setup without the variable
against an unkeyed server removes that line again.
Run
$ rapid-mlx serve qwen3.6-35b-4bit # in another shell: $ codex # headless one-shot (outside a trusted git repo, add --skip-git-repo-check): $ codex exec --skip-git-repo-check "read main.py and summarize what it does"
Recommended aliases (per rapid-mlx agents codex):
qwen3.6-35b-4bit, qwen3.5-9b-4bit,
qwen3-coder-30b-4bit.
Gotchas
-
Codex hardcodes
stream: true. rapid-mlx's/v1/responsesemits the 7 SSE events Codex parses (response.created,response.output_text.delta,response.function_call_arguments.delta,response.completed, …). No override needed. -
Stateless shim.
previous_response_idis not implemented (upstream codex#3841); Codex re-sends the full conversation ininput[]each turn, which is what makes the local backend work. -
Integer tool arguments. With the Qwen3-Coder XML tool format,
an integer passed to a
type: numberparameter arrives as an integer (10000), which Codex's integer-typed fields require; values like1.5stay floats. -
Reasoning effort. Codex's
reasoning.effortis passed to the model, e.g.codex exec -c model_reasoning_effort=high …. Accepted values arenone,minimal,low,medium,highandxhigh. Anything else gets a 400. -
Gemma 4 12B hangs Codex ~60% of tool-use prompts due to a
model-side
thought\n…degenerate loop (upstream Gemma issue, downstream tracked as#686, closed). Use Qwen 3.5 / 3.6 for Codex agent workflows. -
Sandbox. Codex CLI runs commands in its own sandbox
(Seatbelt on macOS). The
sandbox_modesetup writes allows edits inside the workspace; change it inconfig.tomlif you want Codex read-only.
Empirical
The codex-cli row of the integration matrix
is ✅ across all four Tier-1 families (Qwen 3.6, Gemma 4,
DeepSeek R1-Distill, gpt-oss) on the 2026-07-07 pilot. The matrix cell
posts a minimal /v1/responses envelope, verifies the
SSE-shaped output body is non-empty, and asserts no
<think> or <|channel|>analysis
leak in the emitted text. Source of truth for this page:
tests/integrations/test_agents_matrix.py::TestCodexCLI
and the CLI output of
rapid-mlx agents codex.