Goose with a local model on Mac
Run Goose against a model served on your Mac by Rapid-MLX. Goose uses a custom OpenAI-compatible provider with the local Chat Completions endpoint.
Quick start
Install Rapid-MLX, then keep this server running. Install Goose if you do not already have it.
$ pip install 'rapid-mlx==0.15.7' $ rapid-mlx serve qwen3.5-4b-4bit
Create ~/.config/goose/custom_providers/ if needed, then save this as ~/.config/goose/custom_providers/custom_rapid_mlx.json:
{
"name": "custom_rapid_mlx",
"engine": "openai",
"display_name": "Rapid-MLX",
"description": "Local OpenAI-compatible server for Apple Silicon",
"api_key_env": "",
"base_url": "http://127.0.0.1:8000/v1/chat/completions",
"models": [{ "name": "qwen3.5-4b-4bit" }],
"supports_streaming": true,
"requires_auth": false
}
In another terminal, select the custom provider for a new Goose session:
$ goose session start --provider custom_rapid_mlx
In Goose Desktop, open Settings → Models → Configure providers and select Rapid-MLX for a new session.
How the provider connects
The models name must match the alias passed to rapid-mlx serve. Goose's custom provider accepts the full /v1/chat/completions endpoint in base_url. Other OpenAI clients generally use http://127.0.0.1:8000/v1; see the server URL guide. If port 8000 is in use, use the port printed by the server in the JSON.
A default local server needs no key, so this provider sends no authorization header. If you start Rapid-MLX with --api-key, set api_key_env to RAPID_MLX_API_KEY, set requires_auth to true, and export that environment variable with the same key before launching Goose.
Verification and next steps
This configuration was checked against Goose custom-provider docs, its declarative provider schema and OpenAI endpoint handling, plus the Rapid-MLX 0.15.7 source. No model inference was run for this guide.