Run Continue with a Local Model on Your Mac
Continue is the open-source AI assistant for VS Code, JetBrains and the terminal. One model entry in its config points it at a model served on your own Mac.
Continue puts chat, inline edits and an agent into VS Code, JetBrains and a terminal CLI (cn). Every model it uses is described in one config file, and a model can be any OpenAI-compatible server, including Rapid-MLX on your Mac.
With a local model, the code Continue reads and the prompts you write stay on your machine, and there's nothing to pay per token.
The whole thing
- Install Rapid-MLX and serve a model.
- Add a Rapid-MLX model to
~/.continue/config.yaml. - Pick it in Continue.
1 · Install Rapid-MLX
curl -fsSL https://rapidmlx.com/install.sh | bash
Prefer Homebrew? brew install rapid-mlx. Continue itself comes from your editor's extension marketplace, or npm install -g @continuedev/cli for the terminal.
2 · Serve a model
rapid-mlx serve qwen3.5-4b-4bit
That's the engine's pick for a 16 GB Mac and the model we verified with. rapid-mlx recipe prints the pick for your Mac's memory, and Best local LLMs by Mac has measured speeds per chip. Keep this terminal open.
3 · Add the model to Continue
Continue reads ~/.continue/config.yaml. Add a model with the openai provider and your local server as its apiBase:
name: Rapid-MLX local
version: 1.0.0
schema: v1
models:
- name: Rapid-MLX
provider: openai
model: qwen3.5-4b-4bit
apiBase: http://localhost:8000/v1
apiKey: not-needed
roles:
- chat
- edit
- apply
The model value is the alias you served. If the file already lists other models, add this entry to the same models: list.
rapid-mlx agents continue --setup and rapid-mlx launch continue-dev currently write Continue's older config.json. The Continue CLI we tested didn't use that file, so this guide uses config.yaml.
4 · Use it
In VS Code or JetBrains, pick Rapid-MLX in Continue's model dropdown and chat. In a terminal, cn picks up the same file:
# inside your project folder cn -p "Add a function add(a, b) to hello.py that returns a + b. Edit the file."
From our run:
I've added the `add(a, b)` function to `hello.py`. The file now contains: - `greet(name)` - returns "Hello, " + name - `add(a, b)` - returns a + b
Verified on 2026-10-06 with rapid-mlx 0.15.6, Continue CLI 1.5.47 and qwen3.5-4b-4bit on a Mac mini (M2 Pro, 32 GB): with the config.yaml above, cn -p edited the file through Continue's tools in under 10 seconds. We didn't drive the VS Code or JetBrains extension; they read the same config file.
Frequently asked questions
Can Continue use a local LLM on a Mac?
Yes. Add a model with provider: openai and apiBase: http://localhost:8000/v1 to ~/.continue/config.yaml, serve it with Rapid-MLX, and pick it in Continue. Chat, edit and the agent then run on your Mac.
Is Continue with a local model private?
Requests go from Continue to localhost, so your prompts and code stay on the Mac. Continue's own optional features, such as signing in to Continue Hub, follow Continue's policies; the local model path doesn't need them. Rapid-MLX itself sends anonymous usage counts by default, never prompts or code; rapid-mlx telemetry off turns that off (details).
Which model should I use with Continue?
rapid-mlx recipe prints the engine's pick for your memory. For the agent and multi-file edits, use the largest that fits, such as qwen3.8-27b-4bit on 32 GB or more.
Where to go next
- Continue reference: the
launchandagentscommands in detail. - Best local LLMs by Mac: measured speed for your chip and memory.
- Rapid-MLX vs mlx-lm: the same weights, measured task by task.
- All agent guides: every agent Rapid-MLX configures.
rapid-mlx serve <alias> serves any model on
localhost:8000/v1.
curl -fsSL https://rapidmlx.com/install.sh | bash
New models and speedups, in your inbox
A short note whenever Rapid-MLX gets faster or adds models worth running on your Mac. No spam — unsubscribe anytime.