Hero model · rapid-mlx 0.15.7

VibeThinker on Apple Silicon

A 1.5B and 3B reasoning model that punches well above its weight class — long chain-of-thought emitted as a structured reasoning_content field, suitable for distillation pipelines or low-RAM agents. rapid-mlx serves two MLX quants with the custom vibethinker reasoning parser.

Pick a size and run it

The vibethinker reasoning parser and the model card's sampling defaults (temperature 1.0, top_p 0.95) are applied automatically. The reasoning trace comes back as a structured field rather than inline text — see the /v1/responses example.

Download

Every alias on this page downloads with one command — the pull buttons in the tables below copy it. All 2 aliases on this page are mirrored on the rapid-mlx CDN — with automatic mid-pull fallback to Hugging Face if a mirror file slows down. Weights land in the standard Hugging Face cache, and rapid-mlx serve pulls automatically on first use. Live mirror status →

Why VibeThinker matters

The "reasoning at the low end" niche has been the gap between DeepSeek R1 (huge, server-only) and Qwen2.5-7B (small, but not reasoning-tuned). VibeThinker fills it — a 1.5B / 3B distilled reasoner that can show its work on math, code, and logic, while fitting comfortably in the RAM of an entry-level Apple Silicon Mac. If you need a chain-of-thought trace and you don't have 4 GB to spare for it, this is the model.

What Rapid-MLX adds

Sources: the reasoning parsers in rapid_mlx/reasoning/ and the alias profile in rapid_mlx/aliases.json.

Compared with plain mlx-lm

mlx-lm on its own can load VibeThinker weights but does not have a reasoning-parser surface — the <think>...</think> trace appears inline in the chat content. rapid-mlx splits the trace into a structured reasoning_content field on /v1/responses and on the OpenAI-style streaming delta, so clients don't have to string-match.

All quants

Alias Bits Approx. disk Recommended for HF repo get it
vibethinker-1.5b-4bit 4 ~1.0 GB Default — the smallest footprint; runs on any Apple Silicon Mac VibeThinker-1.5B-mlx-4bit CDN
vibethinker-3b-8bit 8 ~3.0 GB Quality bump — 8-bit 3B for distillation / long-CoT runs VibeThinker-3B-8bit CDN

Position vs the big reasoners

DeepSeek R1, V4-Flash, GLM-Z1 — these are stronger reasoners, but they need a high-end workstation or a cloud GPU. Qwen3-Thinking and GLM-4.5-Air are middle-ground, requiring 24-32 GB of unified memory. VibeThinker is one of the smallest reasoning models in the catalog — a chain-of-thought trace from about a gigabyte of weights. On a 256 GB Mac Studio, DeepSeek V4-Flash is meaningfully stronger.

Recommended settings

The alias applies temperature=1.0, top_p=0.95 from the model card to requests that do not set their own — VibeThinker was trained with high-temperature sampling for diversity of reasoning paths, and lower temperatures collapse the trace into shallow restatements. Tool calling uses the hermes parser.

rapid-mlx serve vibethinker-1.5b-4bit

Tutorial — reasoning chain via /v1/responses

The structured reasoning trace shows up as response.reasoning — the message content is just the final answer, so you can render the trace separately or strip it for downstream evals.

curl http://127.0.0.1:8000/v1/responses \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "vibethinker-1.5b-4bit",
    "input": "If a train leaves Chicago at 60 mph and another leaves NYC at 80 mph, when do they meet (790 miles apart)?",
    "reasoning": {"effort": "medium"}
  }'

Python with the OpenAI client:

from openai import OpenAI

client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="none")

resp = client.responses.create(
    model="vibethinker-1.5b-4bit",
    input="Prove that the sum of the first n odd numbers is n^2.",
    reasoning={"effort": "high"},
)

# Final answer
print(resp.output_text)

# Structured reasoning trace
for item in resp.output:
    if item.type == "reasoning":
        for chunk in item.summary:
            print("[reasoning]", chunk.text)

Known limitations

Related