Models · family

Muse Glimmer

3 MLX aliases · Meta's Muse Glimmer 30B · dense reasoner, 131K context.

Pick one

One command per line of this family, smallest download first. Your Mac needs the download size in free memory plus room for macOS and the context window; the hardware tiers page has the engine's picks for every RAM size. The first run downloads the weights and starts an OpenAI-compatible server on http://localhost:8000/v1.

Meta's Muse Glimmer — a 29.6B dense reasoning model with a sliding-window attention stack and recipient-routed channels: the model thinks on a private channel and answers on another, and rapid-mlx demultiplexes that wire natively into reasoning_content and content. Tool calls use the ATEM function-calls envelope, parsed by the dedicated muse parser pair. rapid-mlx serves the text backbone through its own vendored implementation — no external runtime dependency. The checkpoint ships a vision tower, but image input is not served; text, reasoning and tool calls are fully supported.

family
Muse Glimmer
aliases
3
lines
1
install
rapid-mlx serve <alias>
OpenAI base URL
http://localhost:8000/v1

Download

Every alias on this page downloads with one command — the pull buttons in the tables below copy it. All 3 aliases on this page are mirrored on the rapid-mlx CDN — with automatic mid-pull fallback to Hugging Face if a mirror file slows down. Weights land in the standard Hugging Face cache, and rapid-mlx serve pulls automatically on first use. Live mirror status →

Muse Glimmer 30B · 3 aliases

Dense 29.6B with [sliding ×3, global] attention, 131K context and channel-routed reasoning. Uses the muse tool-call parser (ATEM envelope) and the muse reasoning parser — <think>-free chain-of-thought streams out of reasoning_content with no extra config.

parser: muse

aliashf repotool parserreasoningflagscontextAA indexget it
muse-glimmer-30b-4bitmlx-community/Muse-Glimmer-30B-4bitmusemuse—128K35.1CDN
muse-glimmer-30b-8bitmlx-community/Muse-Glimmer-30B-8bitmusemuse—128K35.1CDN
muse-glimmer-30b-bf16mlx-community/Muse-Glimmer-30B-bf16musemuse—128K35.1CDN

Notes & caveats

Context is read from the config.json of the exact build each alias pulls; for an embedding model it is the most input tokens the engine embeds, which can be less than the config declares. AA index is the Artificial Analysis Intelligence Index for the base model at full precision with reasoning on — a property of the model, not a score for our quantised build.

Where next