Models · family

GPT-OSS

10 MLX aliases · OpenAI's GPT-OSS 20B & 120B — MXFP4 / 4 / 8-bit · Harmony-native tool calling.

Pick one

One command per line of this family, smallest download first. Your Mac needs the download size in free memory plus room for macOS and the context window; the hardware tiers page has the engine's picks for every RAM size. The first run downloads the weights and starts an OpenAI-compatible server on http://localhost:8000/v1.

OpenAI's GPT-OSS line in MLX form — both the 20B and the 120B in MXFP4 / 4-bit / 8-bit quants, plus the safeguard-tuned 20B. Harmony-native tool calling (use the harmony tool-call parser) and Harmony reasoning streams. mxfp4-q8 is the recommended-by-OpenAI low-bit format — the 20B runs comfortably on a 32 GB Mac; the 120B wants a Mac Studio with 96 GB+.

family
GPT-OSS
aliases
10
lines
1
install
rapid-mlx serve <alias>
OpenAI base URL
http://localhost:8000/v1

Download

Every alias on this page downloads with one command — the pull buttons in the tables below copy it. 9 of the 10 aliases on this page are mirrored on the rapid-mlx CDN; the rest pull from Hugging Face directly — with automatic mid-pull fallback to Hugging Face if a mirror file slows down. Weights land in the standard Hugging Face cache, and rapid-mlx serve pulls automatically on first use. Live mirror status →

GPT-OSS 20B & 120B · 10 aliases

20B and 120B in MXFP4 / 4-bit / 8-bit, plus the safeguard-tuned 20B. mxfp4-q8 is the recommended-by-OpenAI low-bit format. Harmony-native tool calling + Harmony reasoning streams.

parser: harmony

aliashf repotool parserreasoningflagscontextAA indexget it
gpt-oss-120bmlx-community/gpt-oss-120b-MXFP4-Q8harmonyharmonymoe · spec128K24.1CDN
gpt-oss-120b-4bitmlx-community/gpt-oss-120b-4bitharmonyharmonymoe · spec—24.1HF
gpt-oss-120b-mxfp4-q4mlx-community/gpt-oss-120b-MXFP4-Q4harmonyharmonymoe · spec128K24.1CDN
gpt-oss-120b-mxfp4-q8mlx-community/gpt-oss-120b-MXFP4-Q8harmonyharmonymoe · spec128K24.1CDN
gpt-oss-20bmlx-community/gpt-oss-20b-MXFP4-Q8harmonyharmonymoe · spec128K15.2CDN
gpt-oss-20b-4bitmlx-community/gpt-oss-20b-OptiQ-4bitharmonyharmonymoe · spec128K15.2CDN
gpt-oss-20b-8bitlmstudio-community/gpt-oss-20b-MLX-8bitharmonyharmonymoe · spec128K15.2CDN
gpt-oss-20b-mxfp4-q4mlx-community/gpt-oss-20b-MXFP4-Q4harmonyharmonymoe · spec128K15.2CDN
gpt-oss-20b-mxfp4-q8mlx-community/gpt-oss-20b-MXFP4-Q8harmonyharmonymoe · spec128K15.2CDN
gpt-oss-safeguard-20blmstudio-community/gpt-oss-safeguard-20b-MLX-MXFP4harmonyharmonymoe · spec128K—CDN

Notes & caveats

Context is read from the config.json of the exact build each alias pulls; for an embedding model it is the most input tokens the engine embeds, which can be less than the config declares. AA index is the Artificial Analysis Intelligence Index for the base model at full precision with reasoning on — a property of the model, not a score for our quantised build.

Where next