UI-TARS on Apple Silicon
ByteDance's vision-grounded UI automation agent. UI-TARS reads screen
pixels and emits click / type / scroll actions — the operator side
of a computer-use loop. rapid-mlx serves nine MLX quants across the
1.5-7B, 7B-SFT, 7B-DPO, and 72B-DPO variants, with the custom
ui_tars tool parser wired into the alias config.
Pick a variant and run it
- 16 GB Mac (default):
rapid-mlx serve ui-tars-1.5-7b-4bit— the 1.5 checkpoint at 4-bit, ~4 GB - 24 GB Mac or more:
ui-tars-1.5-7b-6bitor-8bitfor higher precision - 96 GB Mac or more:
rapid-mlx serve ui-tars-72b-dpo-4bit(~40 GB) for the best accuracy
Screenshots go through the vision path, so install the extra first:
pip install "rapid-mlx[vision]==0.15.7". The
ui_tars tool and reasoning parsers are selected
automatically, so /v1/chat/completions returns clean
tool_calls for every screen action. Request examples are
below; every quant is in the
table.
Download
Every alias on this page downloads with one command — the pull buttons in the tables below copy it. 8 of the 9 aliases on this page are mirrored on the rapid-mlx CDN; the rest pull from Hugging Face directly — with automatic mid-pull fallback to Hugging Face if a mirror file slows down. Weights land in the standard Hugging Face cache, and rapid-mlx serve pulls automatically on first use. Live mirror status →
Why UI-TARS matters
Computer-use agents are the niche where a small open VLM serving on a laptop is dramatically more useful than a cloud API: latency, privacy of the screen, and the long tail of "this app, on this OS, with these accessibility settings" all bias toward local execution. UI-TARS is the strongest open competitor in that niche — better grounding than generic VLMs because it was post-trained on a large screen-action dataset. The 1.5-7B is small enough to live next to whatever it is operating; the 72B-DPO is for the rare case where a single workstation drives a fleet of remote desktops.
What Rapid-MLX adds
-
Tool-coupled system prompt. UI-TARS only produces its
structured
Thought: / Action: …trace with the right system prompt, so the server prepends it whenever a request carriestools. Your own system message is kept and added after it. -
Streaming-safe actions. The streaming parser holds back a
partial
Action:until its coordinates are complete, so clients never execute a half-parsed click. -
Chat and Responses parity.
/v1/chat/completionsand/v1/responsesreturn the same actions, andenable_thinking=falsereturns the action without the reasoning prose.
Sources:
rapid_mlx/tool_parsers/ui_tars_tool_parser.py
(actions) and
rapid_mlx/reasoning/ui_tars_parser.py
(streaming reasoning).
Compared with plain mlx-vlm
mlx-vlm on its own can load the UI-TARS weights but does
not understand the action grammar — it emits the raw
Thought: … Action: click(x,y) string as
chat content. rapid-mlx converts that into
tool_calls with a strongly-typed JSON arguments
payload, which is what downstream operator frameworks (browser
controllers, OSWorld harnesses) actually consume.
All quants
All nine quants serve through rapid-mlx serve <alias> with no extra flags.
| Alias | Bits | Approx. disk | Recommended for | HF repo | get it |
|---|---|---|---|---|---|
ui-tars-1.5-7b-4bit | 4 | ~4.0 GB | Default — latest 1.5 SFT, fits 16 GB Mac | UI-TARS-1.5-7B-4bit | CDN |
ui-tars-1.5-7b-6bit | 6 | ~5.5 GB | Quality bump on 24 GB+ | UI-TARS-1.5-7B-6bit | CDN |
ui-tars-1.5-7b-8bit | 8 | ~7.5 GB | Lossless tier for eval runs | UI-TARS-1.5-7B-8bit | CDN |
ui-tars-7b-sft-4bit | 4 | ~4.0 GB | Original 7B SFT — broadest task coverage | UI-TARS-7B-SFT-4bit | CDN |
ui-tars-7b-sft-8bit | 8 | ~7.5 GB | Original 7B SFT, lossless tier | UI-TARS-7B-SFT-8bit | CDN |
ui-tars-7b-dpo-4bit | 4 | ~4.0 GB | DPO-aligned 7B — better instruction following | UI-TARS-7B-DPO-4bit | CDN |
ui-tars-7b-dpo-6bit | 6 | ~5.5 GB | DPO 7B quality bump | UI-TARS-7B-DPO-6bit | CDN |
ui-tars-7b-dpo-8bit | 8 | ~7.5 GB | DPO 7B lossless tier | UI-TARS-7B-DPO-8bit | CDN |
ui-tars-72b-dpo-4bit | 4 | ~40 GB | M3 Ultra workstation — best accuracy | UI-TARS-72B-DPO-4bit | HF |
Recommended settings
UI-TARS is sensitive to sampling. Send temperature: 0
(greedy) — operator agents that take real actions on the user's
machine should be reproducible — or set it once for every request
with --default-temperature 0. Use sampling only for
exploration or data-collection workloads.
rapid-mlx serve ui-tars-1.5-7b-4bit --default-temperature 0
The UI-TARS system prompt is prepended when the request includes
tools=. If you supply your own sysprompt, the parser
will still produce tool_calls as long as the
Thought: / Action: grammar is present
in the model's output.
Tutorial — cURL with a screen-action tool
curl http://127.0.0.1:8000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "ui-tars-1.5-7b-4bit",
"messages": [
{"role": "user", "content": [
{"type": "text", "text": "Click the Submit button."},
{"type": "image_url", "image_url": {"url": "data:image/png;base64,<screenshot>"}}
]}
],
"tools": [
{"type": "function", "function": {
"name": "click",
"description": "Click at screen coordinates",
"parameters": {
"type": "object",
"properties": {"x": {"type": "number"}, "y": {"type": "number"}},
"required": ["x", "y"]
}
}}
]
}'
Python via the OpenAI client:
from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="none")
resp = client.chat.completions.create(
model="ui-tars-1.5-7b-4bit",
messages=[
{"role": "user", "content": [
{"type": "text", "text": "Click the Submit button."},
{"type": "image_url", "image_url": {"url": f"data:image/png;base64,{screenshot_b64}"}},
]},
],
tools=[{
"type": "function",
"function": {
"name": "click",
"description": "Click at screen coordinates",
"parameters": {
"type": "object",
"properties": {"x": {"type": "number"}, "y": {"type": "number"}},
"required": ["x", "y"],
},
},
}],
)
for call in resp.choices[0].message.tool_calls:
print(call.function.name, call.function.arguments)
Known limitations
- UI-TARS expects a single screenshot per turn. Multi-image context (before/after) works but the grounding accuracy drops meaningfully — feed the current frame only.
-
enable_thinking=trueon the 7B-SFT variant emits a verboseThought:trace that consumes the context budget. Setenable_thinking=falsefor headless loops where you only need the action. - The 72B-DPO-4bit is borderline on a 64 GB Mac — it works but leaves no headroom for a long context or another process. 96 GB unified memory is the comfortable floor.