Model pages · hot models · rapid-mlx 0.15.4

glm5.3-flash-4bit

Run GLM-5.3 Flash on a Mac

Z AI's frontier MoE reasoner: 320 B of parameters with 18 B active per token, served on a 256 GB Mac Studio with glm5 thinking separation. The embedded MTP head ships in the checkpoint — the big-Mac flagship the engine qualifies down to the runtime revision.

320 B total / 18 B active MoE · 4-bit · MTP256 GB Mac Studio · frontier reasoning · long context

Measured decode
27.8tok/s
on Mac Studio M3 Ultra · 256 GB · rapid-mlx 0.13.4 · 2026-09-02
Weights
181.7 GB
MLX active / peak (32K)
180.6 / 195.6 GB

One command

rapid-mlx serve glm5.3-flash-4bit

Weights (181.7 GB) download on first run; you get an OpenAI-compatible endpoint at http://localhost:8000/v1. No rapid-mlx yet? It's one line — curl -fsSL https://rapidmlx.com/install.sh | bash — or take the desktop app.

Point a client at it. OpenAI SDK: base_url="http://localhost:8000/v1", model="glm5.3-flash-4bit". Claude Code: ANTHROPIC_BASE_URL=http://localhost:8000 (no /v1 suffix) — see the Claude Code guide.

Measured on real hardware

Headline machine · Mac Studio M3 Ultra · 256 GB

Decode
27.8tok/s
First token · 8K prompt
22.8 s
MLX active / peak (32K)
180.6 / 195.6 GB
Weights on disk
181.7 GB

All measured machines

MachineDecodeFirst tokenPeak memoryCold bootWeights
Mac Studio M3 Ultra · 256 GB rapid-mlx 0.13.4 · 2026-09-02 engine benchmark doc27.8 tok/s22.8 s180.6 / 195.6 GB—181.7 GB

Each row cites its own run; numbers are never copied across machines. Mac Studio M3 Ultra · 256 GB: rapid-mlx 0.13.4, 2026-09-02 — raw data + method.

Will it fit your Mac?

MLX active / peak (32K)
180.6 / 195.6 GB
Weights on disk
181.7 GB

192 GB catalog floor — the measured 32K workload peaks at 195.6 GB MLX and wants a 256 GB Mac.

For the conservative install-default placement see the hardware tiers table; to compare against every model your RAM can hold, use the live picker.

Variants & profiles

  • glm5.3-flash-tensorfold — Experimental accelerated profile on 256 GB Macs — Desktop opt-out
  • glm4.5-air-4bit — GLM-4.5-Air — the older, much smaller GLM MoE

Tensorfold: An experimental accelerated profile (glm5.3-flash-tensorfold) ships alongside the ordinary serving path on 256 GB Macs: same checkpoint, embedded MTP head, fail-closed validation, Desktop opt-out. No speed ratio is claimed for it here.

FAQ

How much memory does GLM-5.3 Flash need on a Mac?

192 GB catalog floor — the measured 32K workload peaks at 195.6 GB MLX and wants a 256 GB Mac.

How fast is GLM-5.3 Flash on Apple Silicon?

We measured 27.8 on a Mac Studio M3 Ultra · 256 GB, rapid-mlx 0.13.4, 2026-09-02; peak memory 180.6 / 195.6 GB on that run. Every row on this page keeps its own chip, version and date — no number is copied across machines.

How do I run GLM-5.3 Flash locally?

rapid-mlx serve glm5.3-flash-4bit — the weights download on first run and you get an OpenAI-compatible endpoint at localhost:8000/v1 that works with Cursor, Claude Code, Aider, and any OpenAI client.

What is the tensorfold profile for GLM-5.3 Flash?

An experimental accelerated profile (glm5.3-flash-tensorfold) ships alongside the ordinary serving path on 256 GB Macs: same checkpoint, embedded MTP head, fail-closed validation, Desktop opt-out. No speed ratio is claimed for it here.

Where next