Run Goose with a Local Model on Your Mac

Goose is Block's open-source agent for the terminal and desktop. Its OpenAI provider accepts a custom host, so it can run on a model served from your own Mac.

Goose is an open-source agent from Block that edits files, runs commands and uses extensions. It's built around tool calls, and its OpenAI provider lets you set the host, so a Rapid-MLX server on your Mac works as its model.

The whole thing

  1. Install Rapid-MLX and serve a model.
  2. Set Goose's provider to openai with your local host.
  3. goose session, or goose run -t "<task>".

1 · Install Rapid-MLX and Goose

curl -fsSL https://rapidmlx.com/install.sh | bash
curl -fsSL https://github.com/block/goose/releases/download/stable/download_cli.sh | CONFIGURE=false bash

The second command installs the Goose CLI to ~/.local/bin without starting its interactive setup. The Goose desktop app is a separate download from the Goose site.

2 · Serve a model

rapid-mlx serve qwen3.5-4b-4bit

That's the engine's pick for a 16 GB Mac and the model we verified with. rapid-mlx recipe prints the pick for your memory; Best local LLMs by Mac has measured speeds per chip. Keep the server running.

3 · Point Goose at it

Goose reads ~/.config/goose/config.yaml. Set the OpenAI provider's host to your server and turn on the built-in developer extension, which gives Goose its file and shell tools:

GOOSE_PROVIDER: openai
GOOSE_MODEL: qwen3.5-4b-4bit
OPENAI_HOST: http://localhost:8000
OPENAI_BASE_PATH: v1/chat/completions
extensions:
  developer:
    enabled: true
    type: builtin
    name: developer
    timeout: 300

Goose also wants an OpenAI key to exist. A plain rapid-mlx serve doesn't check it, so any value works:

export OPENAI_API_KEY=not-needed

4 · Run it

# inside your project folder
goose run --no-session -t "In hello.py, keep the existing greet function and add a new function add(a, b) that returns a + b."

From our run, trimmed:

● new session · openai qwen3.5-4b-4bit
▸ edit
    path hello.py
Edited hello.py (2 lines -> 5 lines)
Done. I've kept the existing `greet` function and added a new `add(a, b)` function that returns `a + b`.

Verified on 2026-10-06 with rapid-mlx 0.15.6, Goose 1.53.0 and qwen3.5-4b-4bit on a Mac mini (M2 Pro, 32 GB): with the config above, goose run edited the file through the developer extension's tools in about 20 seconds (over SSH, so we also set GOOSE_DISABLE_KEYRING=1; Goose couldn't reach the login keychain there). In an earlier run with a vaguer prompt, the 4B model rewrote the whole file instead of adding to it. Small models need precise instructions; a bigger model is more forgiving.

Frequently asked questions

Can Goose use a local LLM on a Mac?

Yes. Set GOOSE_PROVIDER: openai with OPENAI_HOST: http://localhost:8000 in ~/.config/goose/config.yaml, serve a model with Rapid-MLX, and Goose's tool calls run against your Mac.

Do I need an OpenAI account to use Goose locally?

No. Goose needs OPENAI_API_KEY to be set, but a plain rapid-mlx serve doesn't check its value, and no request goes to OpenAI.

Which local model works best with Goose?

Goose leans on tool calls, so use the largest model rapid-mlx recipe suggests for your Mac, such as qwen3.8-27b-4bit on 32 GB or more.

Where to go next


Run this yourself. Rapid-MLX is an open-source, OpenAI- and Anthropic-compatible inference server for Apple Silicon. One command installs it, then rapid-mlx serve <alias> serves any model on localhost:8000/v1.
curl -fsSL https://rapidmlx.com/install.sh | bash

New models and speedups, in your inbox

A short note whenever Rapid-MLX gets faster or adds models worth running on your Mac. No spam — unsubscribe anytime.