Hero model · rapid-mlx 0.15.7

UI-TARS on Apple Silicon

ByteDance's vision-grounded UI automation agent. UI-TARS reads screen pixels and emits click / type / scroll actions — the operator side of a computer-use loop. rapid-mlx serves nine MLX quants across the 1.5-7B, 7B-SFT, 7B-DPO, and 72B-DPO variants, with the custom ui_tars tool parser wired into the alias config.

Pick a variant and run it

Screenshots go through the vision path, so install the extra first: pip install "rapid-mlx[vision]==0.15.7". The ui_tars tool and reasoning parsers are selected automatically, so /v1/chat/completions returns clean tool_calls for every screen action. Request examples are below; every quant is in the table.

Download

Every alias on this page downloads with one command — the pull buttons in the tables below copy it. 8 of the 9 aliases on this page are mirrored on the rapid-mlx CDN; the rest pull from Hugging Face directly — with automatic mid-pull fallback to Hugging Face if a mirror file slows down. Weights land in the standard Hugging Face cache, and rapid-mlx serve pulls automatically on first use. Live mirror status →

Why UI-TARS matters

Computer-use agents are the niche where a small open VLM serving on a laptop is dramatically more useful than a cloud API: latency, privacy of the screen, and the long tail of "this app, on this OS, with these accessibility settings" all bias toward local execution. UI-TARS is the strongest open competitor in that niche — better grounding than generic VLMs because it was post-trained on a large screen-action dataset. The 1.5-7B is small enough to live next to whatever it is operating; the 72B-DPO is for the rare case where a single workstation drives a fleet of remote desktops.

What Rapid-MLX adds

Sources: rapid_mlx/tool_parsers/ui_tars_tool_parser.py (actions) and rapid_mlx/reasoning/ui_tars_parser.py (streaming reasoning).

Compared with plain mlx-vlm

mlx-vlm on its own can load the UI-TARS weights but does not understand the action grammar — it emits the raw Thought: … Action: click(x,y) string as chat content. rapid-mlx converts that into tool_calls with a strongly-typed JSON arguments payload, which is what downstream operator frameworks (browser controllers, OSWorld harnesses) actually consume.

All quants

All nine quants serve through rapid-mlx serve <alias> with no extra flags.

Alias Bits Approx. disk Recommended for HF repo get it
ui-tars-1.5-7b-4bit4~4.0 GBDefault — latest 1.5 SFT, fits 16 GB MacUI-TARS-1.5-7B-4bitCDN
ui-tars-1.5-7b-6bit6~5.5 GBQuality bump on 24 GB+UI-TARS-1.5-7B-6bitCDN
ui-tars-1.5-7b-8bit8~7.5 GBLossless tier for eval runsUI-TARS-1.5-7B-8bitCDN
ui-tars-7b-sft-4bit4~4.0 GBOriginal 7B SFT — broadest task coverageUI-TARS-7B-SFT-4bitCDN
ui-tars-7b-sft-8bit8~7.5 GBOriginal 7B SFT, lossless tierUI-TARS-7B-SFT-8bitCDN
ui-tars-7b-dpo-4bit4~4.0 GBDPO-aligned 7B — better instruction followingUI-TARS-7B-DPO-4bitCDN
ui-tars-7b-dpo-6bit6~5.5 GBDPO 7B quality bumpUI-TARS-7B-DPO-6bitCDN
ui-tars-7b-dpo-8bit8~7.5 GBDPO 7B lossless tierUI-TARS-7B-DPO-8bitCDN
ui-tars-72b-dpo-4bit4~40 GBM3 Ultra workstation — best accuracyUI-TARS-72B-DPO-4bitHF

Recommended settings

UI-TARS is sensitive to sampling. Send temperature: 0 (greedy) — operator agents that take real actions on the user's machine should be reproducible — or set it once for every request with --default-temperature 0. Use sampling only for exploration or data-collection workloads.

rapid-mlx serve ui-tars-1.5-7b-4bit --default-temperature 0

The UI-TARS system prompt is prepended when the request includes tools=. If you supply your own sysprompt, the parser will still produce tool_calls as long as the Thought: / Action: grammar is present in the model's output.

Tutorial — cURL with a screen-action tool

curl http://127.0.0.1:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "ui-tars-1.5-7b-4bit",
    "messages": [
      {"role": "user", "content": [
        {"type": "text", "text": "Click the Submit button."},
        {"type": "image_url", "image_url": {"url": "data:image/png;base64,<screenshot>"}}
      ]}
    ],
    "tools": [
      {"type": "function", "function": {
        "name": "click",
        "description": "Click at screen coordinates",
        "parameters": {
          "type": "object",
          "properties": {"x": {"type": "number"}, "y": {"type": "number"}},
          "required": ["x", "y"]
        }
      }}
    ]
  }'

Python via the OpenAI client:

from openai import OpenAI

client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="none")

resp = client.chat.completions.create(
    model="ui-tars-1.5-7b-4bit",
    messages=[
        {"role": "user", "content": [
            {"type": "text", "text": "Click the Submit button."},
            {"type": "image_url", "image_url": {"url": f"data:image/png;base64,{screenshot_b64}"}},
        ]},
    ],
    tools=[{
        "type": "function",
        "function": {
            "name": "click",
            "description": "Click at screen coordinates",
            "parameters": {
                "type": "object",
                "properties": {"x": {"type": "number"}, "y": {"type": "number"}},
                "required": ["x", "y"],
            },
        },
    }],
)
for call in resp.choices[0].message.tool_calls:
    print(call.function.name, call.function.arguments)

Known limitations

Related