Agent guide · rapid-mlx 0.15.7 · ← Back to README

DeepSeek Harness

DeepSeek Harness (dsh) is DeepSeek's own coding agent — a plugin-composed harness with filesystem, bash, web and sub-agent tools, available as a web UI, a TUI, and a one-shot headless mode. Pointed at rapid-mlx it runs entirely on your Mac: your code never leaves the machine, and there is no DeepSeek account, API key, or token bill involved.

Quick start

Needs: Node 22.15 or newer, dsh 0.2.0-rc.2 or newer (npm install -g @deepseek-ai/dsh) and a running server (rapid-mlx serve qwen3.6-35b-8bit).

$ rapid-mlx agents dsh --setup
$ rapid-mlx agents dsh --test        # runs a real headless task

dsh exits 0 even when a task fails, so check the work, not the exit status. --test grades the result for you.

TIER-1 AGENT   Wire: generic openai-completions provider → /v1/chat/completions · Setup: rapid-mlx agents dsh --setup · Matrix cell: ✅ ✅ XFAIL (arch) ✅ (the DeepSeek column XFAILs on the R1-Distill checkpoint, which cannot emit OpenAI-shape tool_calls — DeepSeek's own client gets no exemption from its own checkpoint's gap).
Promoted to Tier-1 2026-08-17 — real end-to-end multi-step bug fix (ran the failing test, diagnosed an off-by-one, edited the file, re-ran to confirm) on dsh against qwen3.6-35b-8bit, 64 s cold / 17 s warm on an M3 Ultra.

Install

DSH needs Node's Zstd stream API. Node 22.15 or newer works; the npm manifest does not declare that minimum, so an older Node installs cleanly and then fails at plugin boot with an opaque ESM stack.

$ node --version   # must be >= 22.15
$ npm install -g @deepseek-ai/dsh
$ dsh --version    # 0.2.0-rc.2 or newer

Setup supports dsh 0.2.0-rc.2 and newer (npm latest). dsh 0.1.x reads providers from settings.yaml instead, and 0.1.0-rc.7's --profile headless crashes with a Cordis HMR error (a dsh bug), so upgrade before running setup.

Config

Start a server, then let rapid-mlx write the provider block. It previews an exact diff first, so you can see the change before it happens:

$ rapid-mlx serve qwen3.6-35b-8bit

# in another terminal
$ rapid-mlx agents dsh --setup --dry-run   # preview
$ rapid-mlx agents dsh --setup             # apply

--setup asks the server which model is running, how big its context window is, and whether it has a reasoning parser. It then writes dsh's home-level patch layer, $DSH_HOME/cordis.patch.yml (default ~/.dsh/cordis.patch.yml). dsh 0.2 builds every profile from Cordis patch layers and loads this file automatically, so a plain dsh --profile headless picks the provider up with no extra flags. The file is a top-level YAML list of {id, config} layers:

# ~/.dsh/cordis.patch.yml
- id: llm-pi-ai
  config:
    providers:
      rapid-mlx:
        displayName: Rapid-MLX Local
        apiKeyEnv: RAPID_MLX_API_KEY
        api: openai-completions
        baseURL: http://localhost:8000/v1
        defaultContextWindow: 262144
        defaultMaxTokens: 8192
        compat:
          supportsReasoningEffort: true
        models:
        - id: qwen3.6-35b-4bit
          name: qwen3.6-35b-4bit (Rapid-MLX)
          contextWindow: 262144
          maxTokens: 8192
          reasoningEfforts:
            'off': none
            low: low
            medium: medium
            high: high
- id: agent-default-model
  config:
    provider: rapid-mlx
    model: qwen3.6-35b-4bit

Harness's generic OpenAI transport insists on resolving a credential even for an unauthenticated loopback server, so rapid-mlx also adds the non-secret sentinel RAPID_MLX_API_KEY: not-needed to Harness's owner-only .credentials.yaml when that key is absent. An existing value is preserved, and credential values are redacted from the preview — if you keep real remote-provider keys in that file, they are never echoed to your terminal.

One-off alternative. Save the same list anywhere and pass it explicitly. This is the form our 2026-10-02 harness run used:

$ RAPID_MLX_API_KEY=local dsh --profile headless --patch rapid-mlx.yml "summarize this workspace"

Run

$ dsh web                                    # browser UI
$ dsh --profile tui                          # terminal UI
$ dsh --profile headless "summarize this workspace"   # one-shot

Why the generic provider, not deepseek-official

Harness ships a native deepseek-official adapter, and rapid-mlx deliberately does not use it. Pointing it at a loopback server would mean inventing an API key and asserting that your Mac carries the DeepSeek cloud service's model identities, capacity, and rate contract. The generic openai-completions route says exactly what is true: a local OpenAI-compatible endpoint serving whatever model you booted. Reasoning-effort levels are still wired through, so the off/low/medium/high control in the UI reaches reasoning_effort on the wire.

What Tier-1 means here

Tier-1 is a release gate, not a badge. Every version bump runs dsh — the real binary, not a mock — through a real multi-step bug-fix task against a booted 35B model, and grades it by re-running the repository's own test suite. If dsh regresses, the release cannot tag or publish.

The grading deliberately ignores dsh's exit status, because it is not a pass signal: measured on rc.7, a Harness pointed at an unregistered provider prints NO_ADAPTER and still exits 0. Only the work actually landing on disk counts.

You can run the same integration suite yourself:

$ rapid-mlx agents dsh --test

It uses an isolated DSH_HOME and workspace, so it never loads or writes your real Harness sessions or credentials.

Gotchas

See also