VibeThinker on Apple Silicon
A 1.5B and 3B reasoning model that punches well above its weight
class — long chain-of-thought emitted as a structured
reasoning_content field, suitable for distillation
pipelines or low-RAM agents. rapid-mlx serves two MLX quants with
the custom vibethinker reasoning parser.
Pick a size and run it
- Any Apple Silicon Mac (8 GB and up):
rapid-mlx serve vibethinker-1.5b-4bit— ~1 GB of weights - 16 GB Mac or more:
rapid-mlx serve vibethinker-3b-8bit— ~3 GB, the stronger reasoner
The vibethinker reasoning parser and the model card's
sampling defaults (temperature 1.0,
top_p 0.95) are applied automatically. The reasoning
trace comes back as a structured field rather than inline text —
see the /v1/responses example.
Download
Every alias on this page downloads with one command — the pull buttons in the tables below copy it. All 2 aliases on this page are mirrored on the rapid-mlx CDN — with automatic mid-pull fallback to Hugging Face if a mirror file slows down. Weights land in the standard Hugging Face cache, and rapid-mlx serve pulls automatically on first use. Live mirror status →
Why VibeThinker matters
The "reasoning at the low end" niche has been the gap between DeepSeek R1 (huge, server-only) and Qwen2.5-7B (small, but not reasoning-tuned). VibeThinker fills it — a 1.5B / 3B distilled reasoner that can show its work on math, code, and logic, while fitting comfortably in the RAM of an entry-level Apple Silicon Mac. If you need a chain-of-thought trace and you don't have 4 GB to spare for it, this is the model.
What Rapid-MLX adds
-
Alias profile. Both aliases select the
vibethinkerreasoning parser, thehermestool-call parser and the model card's sampling defaults (temperature 1.0,top_p 0.95), applied to any request that does not set its own. -
Streaming-safe reasoning split. The parser keeps the
preamble intact when
<think>is split across stream chunks, and does not emit a partial tool call when a<think>block appears mid-call.
Sources: the reasoning parsers in
rapid_mlx/reasoning/
and the alias profile in
rapid_mlx/aliases.json.
Compared with plain mlx-lm
mlx-lm on its own can load VibeThinker weights but does
not have a reasoning-parser surface — the
<think>...</think> trace appears inline in
the chat content. rapid-mlx splits the trace into a structured
reasoning_content field on
/v1/responses and on the OpenAI-style streaming
delta, so clients don't have to string-match.
All quants
| Alias | Bits | Approx. disk | Recommended for | HF repo | get it |
|---|---|---|---|---|---|
vibethinker-1.5b-4bit |
4 | ~1.0 GB | Default — the smallest footprint; runs on any Apple Silicon Mac | VibeThinker-1.5B-mlx-4bit | CDN |
vibethinker-3b-8bit |
8 | ~3.0 GB | Quality bump — 8-bit 3B for distillation / long-CoT runs | VibeThinker-3B-8bit | CDN |
Position vs the big reasoners
DeepSeek R1, V4-Flash, GLM-Z1 — these are stronger reasoners, but they need a high-end workstation or a cloud GPU. Qwen3-Thinking and GLM-4.5-Air are middle-ground, requiring 24-32 GB of unified memory. VibeThinker is one of the smallest reasoning models in the catalog — a chain-of-thought trace from about a gigabyte of weights. On a 256 GB Mac Studio, DeepSeek V4-Flash is meaningfully stronger.
Recommended settings
The alias applies temperature=1.0,
top_p=0.95 from the model card to requests that do not
set their own — VibeThinker was
trained with high-temperature sampling for diversity of reasoning
paths, and lower temperatures collapse the trace into shallow
restatements. Tool calling uses the
hermes parser.
rapid-mlx serve vibethinker-1.5b-4bit
Tutorial — reasoning chain via /v1/responses
The structured reasoning trace shows up as
response.reasoning — the message content is just the
final answer, so you can render the trace separately or strip it
for downstream evals.
curl http://127.0.0.1:8000/v1/responses \
-H 'Content-Type: application/json' \
-d '{
"model": "vibethinker-1.5b-4bit",
"input": "If a train leaves Chicago at 60 mph and another leaves NYC at 80 mph, when do they meet (790 miles apart)?",
"reasoning": {"effort": "medium"}
}'
Python with the OpenAI client:
from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="none")
resp = client.responses.create(
model="vibethinker-1.5b-4bit",
input="Prove that the sum of the first n odd numbers is n^2.",
reasoning={"effort": "high"},
)
# Final answer
print(resp.output_text)
# Structured reasoning trace
for item in resp.output:
if item.type == "reasoning":
for chunk in item.summary:
print("[reasoning]", chunk.text)
Known limitations
-
At
temperature=0.0the model collapses to short non-reasoning answers — the alias defaults to 1.0 for a reason. Don't override unless you understand what you're losing. - Reasoning traces can run to several thousand tokens on hard problems. If you're streaming to a UI, throttle the render or the user will see a wall of "thinking" before the answer.
-
Tool calling works (via the
hermesparser) but is meaningfully weaker than on Qwen3-Thinking; pick the model for reasoning, not for agentic tool use.