Run Goose with a Local Model on Your Mac
Goose is Block's open-source agent for the terminal and desktop. Its OpenAI provider accepts a custom host, so it can run on a model served from your own Mac.
Goose is an open-source agent from Block that edits files, runs commands and uses extensions. It's built around tool calls, and its OpenAI provider lets you set the host, so a Rapid-MLX server on your Mac works as its model.
The whole thing
- Install Rapid-MLX and serve a model.
- Set Goose's provider to
openaiwith your local host. goose session, orgoose run -t "<task>".
1 · Install Rapid-MLX and Goose
curl -fsSL https://rapidmlx.com/install.sh | bash curl -fsSL https://github.com/block/goose/releases/download/stable/download_cli.sh | CONFIGURE=false bash
The second command installs the Goose CLI to ~/.local/bin without starting its interactive setup. The Goose desktop app is a separate download from the Goose site.
2 · Serve a model
rapid-mlx serve qwen3.5-4b-4bit
That's the engine's pick for a 16 GB Mac and the model we verified with. rapid-mlx recipe prints the pick for your memory; Best local LLMs by Mac has measured speeds per chip. Keep the server running.
3 · Point Goose at it
Goose reads ~/.config/goose/config.yaml. Set the OpenAI provider's host to your server and turn on the built-in developer extension, which gives Goose its file and shell tools:
GOOSE_PROVIDER: openai
GOOSE_MODEL: qwen3.5-4b-4bit
OPENAI_HOST: http://localhost:8000
OPENAI_BASE_PATH: v1/chat/completions
extensions:
developer:
enabled: true
type: builtin
name: developer
timeout: 300
Goose also wants an OpenAI key to exist. A plain rapid-mlx serve doesn't check it, so any value works:
export OPENAI_API_KEY=not-needed
4 · Run it
# inside your project folder goose run --no-session -t "In hello.py, keep the existing greet function and add a new function add(a, b) that returns a + b."
From our run, trimmed:
● new session · openai qwen3.5-4b-4bit
▸ edit
path hello.py
Edited hello.py (2 lines -> 5 lines)
Done. I've kept the existing `greet` function and added a new `add(a, b)` function that returns `a + b`.
Verified on 2026-10-06 with rapid-mlx 0.15.6, Goose 1.53.0 and qwen3.5-4b-4bit on a Mac mini (M2 Pro, 32 GB): with the config above, goose run edited the file through the developer extension's tools in about 20 seconds (over SSH, so we also set GOOSE_DISABLE_KEYRING=1; Goose couldn't reach the login keychain there). In an earlier run with a vaguer prompt, the 4B model rewrote the whole file instead of adding to it. Small models need precise instructions; a bigger model is more forgiving.
Frequently asked questions
Can Goose use a local LLM on a Mac?
Yes. Set GOOSE_PROVIDER: openai with OPENAI_HOST: http://localhost:8000 in ~/.config/goose/config.yaml, serve a model with Rapid-MLX, and Goose's tool calls run against your Mac.
Do I need an OpenAI account to use Goose locally?
No. Goose needs OPENAI_API_KEY to be set, but a plain rapid-mlx serve doesn't check its value, and no request goes to OpenAI.
Which local model works best with Goose?
Goose leans on tool calls, so use the largest model rapid-mlx recipe suggests for your Mac, such as qwen3.8-27b-4bit on 32 GB or more.
Where to go next
- Best local LLMs by Mac: measured speed for your chip and memory.
- Rapid-MLX vs mlx-lm: the same weights, measured task by task.
- Agent guides: the agents
rapid-mlx agents --setupconfigures for you. - API reference: the endpoints Goose and other tools call.
rapid-mlx serve <alias> serves any model on
localhost:8000/v1.
curl -fsSL https://rapidmlx.com/install.sh | bash
New models and speedups, in your inbox
A short note whenever Rapid-MLX gets faster or adds models worth running on your Mac. No spam — unsubscribe anytime.