Bring your own model
Found a model on Hugging Face? Hand rapid-mlx serve its repo
id. For a model outside the catalog that isn't downloaded yet, it first
checks that your Mac can run it, then downloads and serves it.
Run it
$ rapid-mlx serve Qwen/Qwen3-0.6B # any org/repo, or a local folder ✓ Checked before download Format safetensors · bf16 · 1.4 GB Architecture qwen3 (supported) Chat template found Fits your Mac yes · ~2.7 GB of 18 GB …
The check reads only the repo's metadata. Then the download runs, and
once the model has loaded the server prints Ready: with its
address (http://127.0.0.1:8000 by default). Talk to it from
another terminal:
$ curl -s http://localhost:8000/v1/chat/completions -H 'Content-Type: application/json' \
-d '{"model":"Qwen/Qwen3-0.6B","messages":[{"role":"user","content":"Say hello in five words."}]}' \
| jq -r '.choices[0].message.content'
Hello!
…). Your sizes and suggestions will differ.
If it doesn't run
A refused model exits with status 1 and nothing is downloaded. Find the message you got:
| You see | What it means | Next |
|---|---|---|
only has GGUF files |
Rapid-MLX runs MLX and safetensors weights, not GGUF. The message lists MLX builds of the same model when it can verify one (third-party; read the model card), or links a Hugging Face search. | rapid-mlx serve <mlx-build>, e.g. mlx-community/Qwen3-0.6B-4bit |
only has PyTorch .bin weights |
Look for a safetensors or MLX upload of the same model. | rapid-mlx serve <that-repo> |
Architecture … is not supported yet |
No loader in this install can build that model type. rapid-mlx models lists what runs here. |
Ask us to support it: rapid-mlx serve <org/repo> --request (what it sends) |
needs ~N GB of memory; this Mac has M GB |
The weights don't fit this Mac's memory. The message names a catalog model that does. | The suggested model, a more quantized build, or rapid-mlx import <org/repo> --quantize 4 for an unquantized one |
| The model runs but is an unquantized bf16 / fp16 checkpoint | It uses several times the memory of a 4-bit build. | rapid-mlx import <org/repo> --quantize 4, then serve the name it prints (details) |
access to '<org/repo>' is gated |
Accept the licence or request access on the model's Hugging Face page first. | huggingface-cli login or export HF_TOKEN=…, then run again |
| The model is already on disk | Pass the folder instead of a repo id. The same checks run on its files. | rapid-mlx serve ~/models/my-model |
Think this is wrong? |
Every refusal ends with this line. The check can be skipped for one run. | Re-run with --no-preflight |
Import and quantize: rapid-mlx import
serve and pull never convert anything.
Converting an unquantized (bf16 / fp16) safetensors checkpoint to
quantized MLX is a separate, explicit command:
$ rapid-mlx import Qwen/Qwen3-0.6B --quantize 4 Source bf16 · 1.4 GB · qwen3 (supported) Needs ~2.0 GB free disk (you have 99 GB) Output ~0.4 GB, fits in 18 GB ✓ ✓ Smoke test passed rapid-mlx serve qwen3-0.6b-4bit-local (stored at ~/.rapid-mlx/imports/qwen3-0.6b-4bit-local) $ rapid-mlx serve qwen3-0.6b-4bit-local
| Flag | Meaning |
|---|---|
source | A Hugging Face repo id (org/name) or a local model directory. |
--quantize BITS | 2, 3, 4, 6 or 8 (default 4). Affine MLX quantization, group size 64. |
--name NAME | The name to serve it by. The default is <source-name>-<bits>bit; when that's already a catalog or user alias, -local is appended (as above), so an import never shadows an alias. |
--force | Replace an existing import of the same name that came from another source or recipe. |
What it checks first, from metadata, before any download or conversion:
- The source has
model*.safetensorsweights at its root. GGUF /.bin-only repos are refused. - The source isn't already quantized. Otherwise:
is already quantized; serve it directly: rapid-mlx serve <repo>. - Its
model_typehas an mlx-lm converter in this install (text models). Repos that ship their own model code (model_file/auto_map) are refused, because import never runs repository code. - Disk: the output plus 10%, plus the source download plus 10% when the Hugging Face cache is on the same disk and the source isn't cached yet.
- Memory: the quantized output plus
max(4 GB, 20% of RAM) must fit this Mac. Otherwise it's refused
with
Try fewer bits (--quantize 3 or 2).
Cancel-safe and cached.
- Conversion and a one-token smoke test run in child processes, inside a temporary directory next to the import cache. Only a converted, smoke-tested model is renamed into place.
- Ctrl-C, a failed conversion or a failed smoke test leaves the cache
untouched:
Cancelled. The import cache is unchanged.If a--forcereplacement is killed between its two renames, the next import of that name restores the previous copy. - A per-name lock keeps two imports of the same name apart.
- Hub sources download into the normal Hugging Face cache, so an interrupted download resumes. The source is never deleted.
- An import is keyed by source, revision, mlx-lm version and recipe.
Re-running the same import is a no-op (
✓ Already imported). The same name with a different key is refused unless you pass--nameor--force.
Where it lives. ~/.rapid-mlx/imports/<name>
(or $RAPID_MLX_HOME/imports).
$ rapid-mlx models --cached # or: rapid-mlx ls … ── Imported with `rapid-mlx import` (serve or rm by name) ── • qwen3-0.6b-4bit-local 330.9 MiB from Qwen/Qwen3-0.6B · 4-bit $ rapid-mlx serve qwen3-0.6b-4bit-local $ rapid-mlx rm qwen3-0.6b-4bit-local # asks first; -y skips the prompt
Reference
What the check looks at
rapid-mlx serve <org/repo> and
rapid-mlx pull <org/repo> run the check for models
outside the Rapid-MLX catalog. It reads metadata only: one Hugging Face
model_info call, plus config.json fetched into
a temporary directory when an architecture verdict needs confirming. No
weights are fetched.
| Check | What it looks at | Refuses when |
|---|---|---|
| Format | The repo's file list. MLX (quantized) and Hugging Face safetensors both run. | The repo has no safetensors at all and only GGUF files, or only PyTorch pytorch_model*.bin weights. Repos with weights in per-quantization subfolders, or .npz audio weights, are not refused. |
| Architecture | model_type in config.json, compared with every loader this install has: mlx-lm, mlx-vlm, mlx-audio (when the [audio] extra is installed) and the architectures Rapid-MLX registers itself. |
The config proves it is a text-generating model that no installed loader can build. The check confirms this against the full config.json before refusing. Repos that ship their own model code are never refused on architecture. |
| Chat template | chat_template.jinja / .json, or the template in the tokenizer config. |
Never. It's informational. Chat template none — likely a base model; prefer /v1/completions means use text completion, not chat. |
| Memory fit | The root model*.safetensors shards, plus an allowance for KV cache and runtime (weights × 1.2 + 1 GB), against this Mac's physical memory. yes up to 75% of memory, tight above that, no when the weights alone are larger than memory. |
Only serve, and only for no. pull reports the fit but still downloads (you may be pulling for another Mac), and serve --disk-stream skips this refusal. |
The check is conservative. Anything it can't establish gives no verdict,
and serve / pull continue without a verdict.
That covers network trouble, a Hub answer slower than 10 seconds, a
gated repo you can't read yet, and an unusual repo layout. It doesn't
run at all for catalog aliases, repos already in your Hugging Face
cache, offline mode (HF_HUB_OFFLINE), or
pull --bits / --format.
--no-preflight skips the check for one run, on both
serve and pull.
What refusals look like
$ rapid-mlx pull Qwen/Qwen3-0.6B-GGUF
! Qwen/Qwen3-0.6B-GGUF only has GGUF files (Q8_0).
Rapid-MLX runs MLX and safetensors weights, not GGUF. Nothing was downloaded.
MLX builds of the same base model (third-party, not reviewed by Rapid-MLX;
matched by base model, architecture and size):
mlx-community/Qwen3-0.6B-4bit · 4-bit · ~0.3 GB · apache-2.0
lmstudio-community/Qwen3-0.6B-MLX-4bit · 4-bit · ~0.3 GB · apache-2.0
To try one, check its model card first, then: rapid-mlx pull mlx-community/Qwen3-0.6B-4bit
(Think this is wrong? Re-run with --no-preflight to skip this check.)
rapid-mlx pull <repo> --format gguf is refused outright:
Rapid-MLX cannot run GGUF files, so `--format gguf` is not available.
$ rapid-mlx serve Qwen/Qwen3-235B-A22B
✗ Qwen/Qwen3-235B-A22B needs ~526 GB of memory; this Mac has 18 GB.
Its weights alone are 438 GB. Nothing was downloaded.
Pick a smaller or more quantized build: rapid-mlx models
A catalog model that runs well on your Mac:
rapid-mlx serve qwen3.5-9b-4bit
(Think this is wrong? Re-run with --no-preflight to skip this check.)
Suggested alternatives (never switched automatically)
A refusal can name up to two things that will run. Neither is ever substituted for the model you asked for. To use one, type the command it prints.
-
An MLX build of the same model (GGUF-only Hub repos). The
refused repo must declare itself a quantization of exactly one base
repo (finetunes and merges count as different models). A candidate is
shown only when all of these hold:
- its
model_typematches the base - its parameter count is within 5%
- it uses an MLX quantization (not AWQ / GPTQ)
- this install supports its architecture
- it fits this Mac comfortably
- its
- A catalog model of similar size, picked by the same recommendation policy the installer and Desktop use, limited to what this Mac's memory supports. When the refusal was "too big", it suggests this Mac's own recommendation instead of the nearest size.
All Hub lookups for suggestions share one short time budget. A slow Hub means fewer suggestions, never a slower refusal.
“Ask us to support it?” (opt-in)
When the check refuses a public Hugging Face model for its
architecture or its weight format (GGUF / .bin only), an
interactive terminal asks:
Ask us to support it? This sends the repo id, architecture, format and your Rapid-MLX version. Nothing else. [y/N]
The request also carries the refusal's failure class, as the payload
below shows. Nothing is sent unless you answer y. In scripts and other
non-interactive shells there's no prompt. Re-run with
--request (on serve or pull) to
send it without asking. When you do, exactly five fields go to
rapidmlx.com/api/model-request:
{
"repo": "Qwen/Qwen3-0.6B-GGUF",
"model_type": null,
"format": "gguf",
"failure": "unsupported_format",
"version": "0.15.5"
}
- Never offered for gated or private repos, local paths, or "too big for this Mac" refusals.
- The site groups requests by failure class + architecture/format into one GitHub issue per key. Yours opens it or adds a vote, and the CLI prints the issue URL so you can follow it.
- If rapidmlx.com can't be reached, the CLI prints a prefilled GitHub issue link so you can file the request yourself.
- The HTTP
User-Agentis the fixed stringrapid-mlx-cli. Separately, if anonymous usage telemetry is on, the refusal is recorded like any failed start: its failure class, plus a model identity that is a repo id only when the repo is proven public, and otherwise<custom>or<local>(telemetry details).
Gated and private repos
Until you have access, the check can't read a gated repo's metadata, so it gives no verdict and the download stops with:
Error: access to '<org/repo>' is gated on Hugging Face. Accept the licence or request access at https://huggingface.co/<org/repo> Then run `huggingface-cli login` or set `HF_TOKEN`, and try again.
Once you're signed in, the check runs normally. Support requests are
never offered for gated or private repos. rapid-mlx import
of a gated repo needs the same sign-in.
Local directories
The format, architecture and memory checks run on the folder's own
files and config.json. There's nothing to download, and
support requests aren't offered for local paths.
$ rapid-mlx serve ~/models/my-finetune $ rapid-mlx import ~/models/my-finetune --quantize 4 --name my-ft-4bit $ rapid-mlx serve my-ft-4bit
A local import is keyed by a fingerprint of the folder (file names,
sizes and modification times, plus the contents of small files like
configs and tokenizers). Re-running on an unchanged folder is a no-op.
After you edit the folder, the stale import isn't reused: re-run with
--force to replace it, or pick a new --name.
Frequently asked questions
Can Rapid-MLX run a GGUF model?
No. Rapid-MLX runs MLX and safetensors weights. A GGUF-only repo is refused before anything is downloaded and, when the check can verify one, it names an MLX build of the same base model that fits your Mac.
Does the check download the model?
No. It reads Hugging Face metadata only: one model_info call, plus config.json into a temporary directory when it needs to confirm an architecture verdict. If it can't read the metadata, it gives no verdict and the normal download proceeds.
What does a support request send?
It's sent only after you answer y or pass --request, and it contains the repo id, the architecture (model_type), the weight format, the failure class and your Rapid-MLX version. It's never offered for gated, private or local models, or for models that are simply too big for your Mac.
How do I quantize a bf16 model for my Mac?
rapid-mlx import <org/repo> --quantize 4 (2, 3, 4, 6 or 8 bits). It checks disk and memory first, converts in a temporary directory, runs a one-token smoke test, and only then publishes the model under a name you can serve, list and remove.