glm5.3-flash-4bit
Run GLM-5.3 Flash on a Mac
Z AI's frontier MoE reasoner: 320 B of parameters with 18 B active per token, served on a 256 GB Mac Studio with glm5 thinking separation. The embedded MTP head ships in the checkpoint — the big-Mac flagship the engine qualifies down to the runtime revision.
320 B total / 18 B active MoE · 4-bit · MTP256 GB Mac Studio · frontier reasoning · long context
One command
rapid-mlx serve glm5.3-flash-4bit
Weights (181.7 GB) download on first run; you get an OpenAI-compatible
endpoint at http://localhost:8000/v1. No rapid-mlx yet? It's one line —
curl -fsSL https://rapidmlx.com/install.sh | bash — or take the
desktop app.
Point a client at it. OpenAI SDK: base_url="http://localhost:8000/v1",
model="glm5.3-flash-4bit". Claude Code:
ANTHROPIC_BASE_URL=http://localhost:8000 (no /v1 suffix) —
see the Claude Code guide.
Measured on real hardware
Headline machine · Mac Studio M3 Ultra · 256 GB
All measured machines
| Machine | Decode | First token | Peak memory | Cold boot | Weights |
|---|---|---|---|---|---|
| Mac Studio M3 Ultra · 256 GB rapid-mlx 0.13.4 · 2026-09-02 engine benchmark doc | 27.8 tok/s | 22.8 s | 180.6 / 195.6 GB | — | 181.7 GB |
Each row cites its own run; numbers are never copied across machines. Mac Studio M3 Ultra · 256 GB: rapid-mlx 0.13.4, 2026-09-02 — raw data + method.
Will it fit your Mac?
192 GB catalog floor — the measured 32K workload peaks at 195.6 GB MLX and wants a 256 GB Mac.
For the conservative install-default placement see the hardware tiers table; to compare against every model your RAM can hold, use the live picker.
Variants & profiles
glm5.3-flash-tensorfold— Experimental accelerated profile on 256 GB Macs — Desktop opt-outglm4.5-air-4bit— GLM-4.5-Air — the older, much smaller GLM MoE
Tensorfold: An experimental accelerated profile (glm5.3-flash-tensorfold) ships alongside the ordinary serving path on 256 GB Macs: same checkpoint, embedded MTP head, fail-closed validation, Desktop opt-out. No speed ratio is claimed for it here.
FAQ
How much memory does GLM-5.3 Flash need on a Mac?
192 GB catalog floor — the measured 32K workload peaks at 195.6 GB MLX and wants a 256 GB Mac.
How fast is GLM-5.3 Flash on Apple Silicon?
We measured 27.8 on a Mac Studio M3 Ultra · 256 GB, rapid-mlx 0.13.4, 2026-09-02; peak memory 180.6 / 195.6 GB on that run. Every row on this page keeps its own chip, version and date — no number is copied across machines.
How do I run GLM-5.3 Flash locally?
rapid-mlx serve glm5.3-flash-4bit — the weights download on first run and you get an OpenAI-compatible endpoint at localhost:8000/v1 that works with Cursor, Claude Code, Aider, and any OpenAI client.
What is the tensorfold profile for GLM-5.3 Flash?
An experimental accelerated profile (glm5.3-flash-tensorfold) ships alongside the ordinary serving path on 256 GB Macs: same checkpoint, embedded MTP head, fail-closed validation, Desktop opt-out. No speed ratio is claimed for it here.