Run Cline with a Local Model on Your Mac
Cline is an open-source coding agent for VS Code and the terminal. Its OpenAI Compatible provider takes a base URL, so the model behind it can run on your own Mac.
Cline plans and makes changes across your project, asking before it runs commands or edits files. It works with any OpenAI-compatible endpoint, including Rapid-MLX on your Mac, in both its VS Code extension and its CLI.
The whole thing
- Install Rapid-MLX and serve a model.
cline auth -p openai -b http://localhost:8000/v1 …(CLI), or set the same values in the extension.- Give Cline a task.
1 · Install Rapid-MLX and Cline
curl -fsSL https://rapidmlx.com/install.sh | bash npm install -g cline
For the editor, install Cline from the VS Code marketplace instead of, or as well as, the CLI.
2 · Serve a model
rapid-mlx serve qwen3.5-4b-4bit
That's the engine's pick for a 16 GB Mac and the model we verified with. rapid-mlx recipe prints the pick for your memory; Best local LLMs by Mac has measured speeds per chip. Keep the server running.
3 · Point Cline at it
CLI. One command stores the provider, base URL, key and model:
cline auth -p openai -b http://localhost:8000/v1 -k not-needed -m qwen3.5-4b-4bit
It replies Provider configured: openai-compatible (qwen3.5-4b-4bit). A plain rapid-mlx serve doesn't check the key, so any value works.
VS Code extension. Open Cline's settings and choose OpenAI Compatible as the API provider, then enter the same three values: base URL http://localhost:8000/v1, any API key, and the model alias you served as the model ID.
4 · Run it
# inside your project folder cline "In hello.py, keep the existing greet function and add a new function add(a, b) that returns a + b."
The CLI runs in act mode with tool approval on by default, so it edits straight away. From our run, trimmed:
- `greet(name)` - returns a greeting message - `add(a, b)` - returns the sum of two numbers The task is complete. The hello.py file now contains both the existing `greet` function and the new `add(a, b)` function.
Verified on 2026-10-06 with rapid-mlx 0.15.6, Cline CLI 3.0.68 and qwen3.5-4b-4bit on a Mac mini (M2 Pro, 32 GB): after cline auth as above, the task edited the file through Cline's tools in under 30 seconds. We didn't drive the VS Code extension; it uses the same OpenAI Compatible provider settings.
Frequently asked questions
Can Cline use a local model on a Mac?
Yes. Cline's OpenAI Compatible provider takes any base URL. Serve a model with Rapid-MLX and set the base URL to http://localhost:8000/v1, with cline auth in the terminal or in the extension's settings.
Does Cline with a local model send my code anywhere?
Model requests go to localhost, so prompts and code stay on your Mac. Cline's own optional account features follow Cline's policies; the local provider doesn't need them. Rapid-MLX itself sends anonymous usage counts by default, never prompts or code; rapid-mlx telemetry off turns that off (details).
Which model should I use with Cline?
Cline's prompts are long and it relies on tool calls, so use the largest model rapid-mlx recipe suggests, such as qwen3.8-27b-4bit on 32 GB or more.
Where to go next
- Best local LLMs by Mac: measured speed for your chip and memory.
- Rapid-MLX vs mlx-lm: the same weights, measured task by task.
- Kilo Code reference: the Cline fork, configured by
rapid-mlx agents kilo-code --setup. - All agent guides: every agent Rapid-MLX configures.
rapid-mlx serve <alias> serves any model on
localhost:8000/v1.
curl -fsSL https://rapidmlx.com/install.sh | bash
New models and speedups, in your inbox
A short note whenever Rapid-MLX gets faster or adds models worth running on your Mac. No spam — unsubscribe anytime.