OpenHands
OpenHands
runs a sandboxed agent inside Docker. The CodeActAgent parses
text-action tags (<execute_bash>,
<execute_ipython>) out of plain-text LLM output —
NOT OpenAI function calls — and applies file edits through its
Docker runtime. rapid-mlx is wired as an OpenAI-compatible provider
through LiteLLM.
Quick start
Needs: OpenHands (pip install openhands), Docker running, and a server (rapid-mlx serve qwen3.6-35b-4bit).
$ rapid-mlx agents openhands --setup # prints three export lines # paste them into your shell, then: $ openhands
If LiteLLM can't route the model, prefix it: LLM_MODEL=openai/qwen3.6-35b-4bit. rapid-mlx agents openhands --test runs the integration check.
/v1/chat/completions ·
Setup: env vars LLM_BASE_URL /
LLM_API_KEY / LLM_MODEL ·
Matrix cell:
✅ ✅ ✅ XFAIL
(Qwen3.6 / Gemma 4 / DeepSeek pass; gpt-oss is
XFAIL (format) — its native harmony output never emits the
<execute_bash> text-action tags CodeActAgent
parses, an upstream OpenHands gap
#15167).
Install
$ pip install openhands # Docker Desktop / dockerd must be running — OpenHands ships a # sandbox-runtime container image and a docker-in-docker sock passthrough.
Config
LiteLLM env vars. rapid-mlx agents openhands --setup
prints them for the model the server reports and writes no files:
$ export LLM_BASE_URL=http://localhost:8000/v1 $ export LLM_API_KEY=not-needed $ export LLM_MODEL=qwen3.6-35b-4bit # Setup prints the bare model id. If LiteLLM can't route it, prefix it with # ``openai/`` (LLM_MODEL=openai/qwen3.6-35b-4bit). The prefix is a LiteLLM # routing hint; without it LiteLLM guesses the provider from the model name, # which fails on non-canonical rapid-mlx names.
Run
$ rapid-mlx serve qwen3.6-35b-4bit $ LLM_BASE_URL=http://localhost:8000/v1 \ LLM_API_KEY=not-needed \ LLM_MODEL=openai/qwen3.6-35b-4bit \ openhands
Recommended aliases (per rapid-mlx agents openhands):
qwen3.5-9b-4bit, qwen3.6-35b-4bit,
qwen3.5-4b-4bit.
Docker E2E harness
The integration test
tests/integrations/test_openhands.sh
drives the pinned OpenHands 0.9.0 app + runtime images against a
running rapid-mlx serve:
# pinned by manifest-list digest so a moved tag can't silently swap the image OPENHANDS_IMAGE="ghcr.io/all-hands-ai/openhands:0.9.0@sha256:d4b028…" OPENHANDS_RUNTIME_IMAGE="ghcr.io/all-hands-ai/runtime:od_v0.9.0_…" $ docker run --rm \ -e SANDBOX_CONTAINER_IMAGE="$OPENHANDS_RUNTIME_IMAGE" \ -e LLM_BASE_URL="http://host.docker.internal:8000/v1" \ -e LLM_MODEL="openai/$MODEL" \ -e LLM_API_KEY="rapidmlx" \ -v /var/run/docker.sock:/var/run/docker.sock \ -v "$WORKDIR:/opt/workspace_base" \ --add-host host.docker.internal:host-gateway \ "$OPENHANDS_IMAGE" \ python -m openhands.core.main \ -i 10 -d /opt/workspace_base \ -t "Fix the bug in add.py — this function should add, not subtract."
When the LLM is running on the same Mac as Docker Desktop,
localhost inside the container refers to the container
itself. The harness rewrites the URL host to
host.docker.internal so the container can reach the
rapid-mlx server. Any non-loopback host (remote-serve node, RFC1918
IP, DNS name) is preserved as-is.
Gotchas
- Docker daemon required. OpenHands' sandbox runtime is a docker-in-docker sock passthrough — non-Docker CI skips the harness cleanly.
-
Text-action, not function calling. OpenHands parses
<execute_bash>/<execute_ipython>blocks out of plain text. This is why R1-Distill (which can't emit function calls) still passes OpenHands cells — it can emit text-action tags. -
gpt-oss — XFAIL (upstream format gap). On the wire, rapid-mlx
applies user-supplied
stop=['</execute_ipython>', ...]only to the harmony final channel, so analysis-channel reasoning does not trigger a premature stop. The cell still XFAILs: gpt-oss's native harmony output never emits the<execute_bash>/<execute_ipython>text-action tags CodeActAgent parses — an upstream OpenHands parser gap tracked at OpenHands #15167, not a rapid-mlx bug. - Small local models struggle with the complex action format. 7B-and-below can loop; 12B-and-up (Gemma 4, Qwen 3.5 9B, Qwen 3.6 35B-A3B) hit the AST whitelist reliably.
Empirical
The 2026-07-07 pilot ran the harness across all four Tier-1
families with the pinned image pair. Family × wall-time (single
cell, cached-image path):
Qwen 3.5-4B-4bit 32.14 s (2 CodeAct steps: read →
edit_file_by_replace → finish),
Gemma-4-31B-4bit 47.87 s,
DeepSeek R1-Distill-32B-4bit 72.08 s (long analysis-channel CoT
before the edit action, still one-shot),
gpt-oss-20B-MXFP4-Q8 XFAIL (harmony output emits no
<execute_bash> text-action tag for CodeActAgent to
parse — upstream OpenHands
#15167,
see the gotcha above). The three tool-emitting families satisfied the
AST whitelist on add.py; gpt-oss stays strict-xfail.