Agent guide · rapid-mlx 0.15.7

Goose with a local model on Mac

Run Goose against a model served on your Mac by Rapid-MLX. Goose uses a custom OpenAI-compatible provider with the local Chat Completions endpoint.

Quick start

Install Rapid-MLX, then keep this server running. Install Goose if you do not already have it.

$ pip install 'rapid-mlx==0.15.7'
$ rapid-mlx serve qwen3.5-4b-4bit

Create ~/.config/goose/custom_providers/ if needed, then save this as ~/.config/goose/custom_providers/custom_rapid_mlx.json:

{
  "name": "custom_rapid_mlx",
  "engine": "openai",
  "display_name": "Rapid-MLX",
  "description": "Local OpenAI-compatible server for Apple Silicon",
  "api_key_env": "",
  "base_url": "http://127.0.0.1:8000/v1/chat/completions",
  "models": [{ "name": "qwen3.5-4b-4bit" }],
  "supports_streaming": true,
  "requires_auth": false
}

In another terminal, select the custom provider for a new Goose session:

$ goose session start --provider custom_rapid_mlx

In Goose Desktop, open Settings → Models → Configure providers and select Rapid-MLX for a new session.

How the provider connects

The models name must match the alias passed to rapid-mlx serve. Goose's custom provider accepts the full /v1/chat/completions endpoint in base_url. Other OpenAI clients generally use http://127.0.0.1:8000/v1; see the server URL guide. If port 8000 is in use, use the port printed by the server in the JSON.

A default local server needs no key, so this provider sends no authorization header. If you start Rapid-MLX with --api-key, set api_key_env to RAPID_MLX_API_KEY, set requires_auth to true, and export that environment variable with the same key before launching Goose.

Verification and next steps

This configuration was checked against Goose custom-provider docs, its declarative provider schema and OpenAI endpoint handling, plus the Rapid-MLX 0.15.7 source. No model inference was run for this guide.