Ling
1 MLX alias · inclusionAI's Ling-3.0-tiny · 7.9B MoE with 1.3B active, KDA+MLA hybrid, MIT.
Pick one
One command per line of this family, smallest download first. Your Mac needs the download size in free memory plus room for macOS and the context window; the hardware tiers page has the engine's picks for every RAM size. The first run downloads the weights and starts an OpenAI-compatible server on http://localhost:8000/v1.
rapid-mlx serve ling-3.0-tiny-4bit
inclusionAI's Ling 3.0 line — sparse-MoE reasoners on a KDA + MLA hybrid attention stack (three linear-attention layers per full-attention layer), MIT licensed. Ling-3.0-tiny is 7.9B parameters total with only 1.3B active per token (128 experts, top-8 plus one shared), which is how a 131K-context reasoner fits and runs on an 8 GB Mac. rapid-mlx serves it through its own vendored bailing_hybrid implementation — verified against the official modeling code — with thinking streamed as reasoning_content and native tool calling.
- family
- Ling
- aliases
- 1
- lines
- 1
- install
- rapid-mlx serve <alias>
- OpenAI base URL
- http://localhost:8000/v1
Download
Every alias on this page downloads with one command — the pull buttons in the tables below copy it. All 1 aliases on this page are mirrored on the rapid-mlx CDN — with automatic mid-pull fallback to Hugging Face if a mirror file slows down. Weights land in the standard Hugging Face cache, and rapid-mlx serve pulls automatically on first use. Live mirror status →
Ling 3.0 tiny · 1 alias
7.9B-total / 1.3B-active MoE with the KDA + MLA hybrid stack and 131K context. Uses the glm47 tool-call parser and the qwen3 reasoning parser; thinking toggles via chat_template_kwargs: {"enable_thinking": true} or the model's detailed thinking on/off system-prompt switch.
parser: glm47
| alias | hf repo | tool parser | reasoning | flags | context | AA index | get it |
|---|---|---|---|---|---|---|---|
| ling-3.0-tiny-4bit | rapid-mlx/Ling-3.0-tiny-MLX-4bit | glm47 | qwen3 | hybrid · moe | 128K | 24.5 | CDN |
Notes & caveats
- ling-3.0-tiny-4bit (4.2 GB) is the pick — the strongest reasoner that fits an 8 GB Mac, mirrored on our edge CDN.
- First MLX conversion of this family — published under the rapid-mlx org on Hugging Face.
Context is read from the config.json of the exact build each alias pulls; for an embedding model it is the most input tokens the engine embeds, which can be less than the config declares. AA index is the Artificial Analysis Intelligence Index for the base model at full precision with reasoning on — a property of the model, not a score for our quantised build.