Kilo Code
Kilo Code is a
Cline fork with independent maintenance. It ships as a VS Code
extension and a CLI; both honor an openai API provider
block that points at any OpenAI-compat base URL.
Quick start
Needs: The Kilo CLI (npm install -g @kilocode/cli) and a running server (rapid-mlx serve qwen3-coder-30b-4bit).
$ rapid-mlx agents kilo-code --setup $ kilo run "say hello" # a reply = it works
A reply means Kilo is using your local model. rapid-mlx agents kilo-code --test runs the full integration check.
/v1/chat/completions ·
Setup: rapid-mlx agents kilo-code --setup ·
Matrix cell:
✅ ✅ XFAIL(arch) ✅
(DeepSeek R1-Distill — see
XFAIL arch).
Install
$ npm install -g @kilocode/cli
Config
One-shot, with a server running: rapid-mlx agents kilo-code --setup
writes the Kilo CLI's ~/.config/kilo/kilo.json, filling in
the model and context window the server reports. An existing file is merged: your other keys are kept, no backup is made,
and --dry-run previews without writing. The template:
// ~/.config/kilo/kilo.json { "$schema": "https://app.kilo.ai/config.json", "model": "rapid-mlx/qwen3-coder-30b-4bit", "provider": { "rapid-mlx": { "npm": "@ai-sdk/openai-compatible", "name": "Rapid-MLX", "options": { "apiKey": "not-needed", "baseURL": "http://localhost:8000/v1" }, "models": { "qwen3-coder-30b-4bit": { "name": "qwen3-coder-30b-4bit (Rapid-MLX)", "tool_call": true, "reasoning": true, "limit": {"context": 262144, "output": 8192} } } } } }
For the VS Code extension: Settings → Kilo Code → API Provider =
openai; Base URL = http://localhost:8000/v1.
Run
$ rapid-mlx serve qwen3-coder-30b-4bit $ kilo
Recommended aliases (per rapid-mlx agents kilo-code):
qwen3-coder-30b-4bit, qwen3.6-35b-4bit,
qwen3.5-9b-4bit, gpt-oss-20b-mxfp4-q8.
Gotchas
- Do not point Kilo at a Cline config. Kilo is a Cline fork; the schema drifted after the split.
-
First-run key prompt. Choose "OpenAI Compatible" and
paste
not-neededas the API key. - Continue.dev migrants: Kilo's JSON config is stricter. Quote all API base URLs.
Empirical
The kilo-code row of the integration matrix
is ✅ on Qwen 3.6, Gemma 4, and gpt-oss; DeepSeek R1-Distill
XFAIL(arch). Source of truth for this page:
TestKiloCode
+ rapid-mlx agents kilo-code.