Models · agent models · rapid-mlx 0.15.7

Holo3 — small-MoE agent model

Holo3.1-35B-A3B is a 35B-parameter / 3B-active mixture-of-experts agent model from Holo Labs, served by rapid-mlx as the holo3.1-35b-a3b (4-bit) and holo3.1-35b-a3b-8bit (8-bit) aliases. A peer Apple-Silicon agent team reported it running well in production.

Pick a quant and run it

rapid-mlx serve holo3.1-35b-a3b        # 4-bit — fits a 32 GB Mac with headroom
rapid-mlx serve holo3.1-35b-a3b-8bit   # 8-bit — 64 GB

Weights pull from Hugging Face (pipenetwork/) on first use. The endpoint is the standard OpenAI-compatible one at http://localhost:8000/v1, with a Hermes-shape tool-call envelope and the qwen3 reasoning parser selected for you, so existing agent stacks work unchanged.

Why it's notable

The bar for a rapid-mlx mirror is narrow: a model has to fit a real Apple-Silicon Mac (not a Mac Studio cluster), do something specific better than its size class, and have at least one credible production user. Holo3 hits all three.

Architecture sketch

family
Holo3 (Holo Labs)
variant
Holo3.1-35B-A3B
shape
MoE · 35B total · 3B active
focus
Agent / tool-using workflows
upstream HF
huggingface.co/HoloLabs
rapid-mlx alias
holo3.1-35b-a3b · holo3.1-35b-a3b-8bit
Hugging Face repos
pipenetwork/Holo-3.1-35B-A3B-MLX-4bit · -8bit
tool parser
hermes · reasoning: qwen3

Download

Both quants pull from Hugging Face directly. One command each:

rapid-mlx pull holo3.1-35b-a3b
rapid-mlx pull holo3.1-35b-a3b-8bit

Weights land in the standard Hugging Face cache, and rapid-mlx serve pulls automatically on first use. Live mirror status →

Where next