Model guide · rapid-mlx 0.15.5

Bring your own model

Found a model on Hugging Face? Hand rapid-mlx serve its repo id. For a model outside the catalog that isn't downloaded yet, it first checks that your Mac can run it, then downloads and serves it.

Run it

$ rapid-mlx serve Qwen/Qwen3-0.6B       # any org/repo, or a local folder

  ✓ Checked before download
    Format        safetensors · bf16 · 1.4 GB
    Architecture  qwen3 (supported)
    Chat template found
    Fits your Mac yes · ~2.7 GB of 18 GB
  …

The check reads only the repo's metadata. Then the download runs, and once the model has loaded the server prints Ready: with its address (http://127.0.0.1:8000 by default). Talk to it from another terminal:

$ curl -s http://localhost:8000/v1/chat/completions -H 'Content-Type: application/json' \
    -d '{"model":"Qwen/Qwen3-0.6B","messages":[{"role":"user","content":"Say hello in five words."}]}' \
    | jq -r '.choices[0].message.content'
Hello!
Output captured from a rapid-mlx 0.15.5 install on an 18 GB Mac (abridged at …). Your sizes and suggestions will differ.

If it doesn't run

A refused model exits with status 1 and nothing is downloaded. Find the message you got:

You seeWhat it meansNext
only has GGUF files Rapid-MLX runs MLX and safetensors weights, not GGUF. The message lists MLX builds of the same model when it can verify one (third-party; read the model card), or links a Hugging Face search. rapid-mlx serve <mlx-build>, e.g. mlx-community/Qwen3-0.6B-4bit
only has PyTorch .bin weights Look for a safetensors or MLX upload of the same model. rapid-mlx serve <that-repo>
Architecture … is not supported yet No loader in this install can build that model type. rapid-mlx models lists what runs here. Ask us to support it: rapid-mlx serve <org/repo> --request (what it sends)
needs ~N GB of memory; this Mac has M GB The weights don't fit this Mac's memory. The message names a catalog model that does. The suggested model, a more quantized build, or rapid-mlx import <org/repo> --quantize 4 for an unquantized one
The model runs but is an unquantized bf16 / fp16 checkpoint It uses several times the memory of a 4-bit build. rapid-mlx import <org/repo> --quantize 4, then serve the name it prints (details)
access to '<org/repo>' is gated Accept the licence or request access on the model's Hugging Face page first. huggingface-cli login or export HF_TOKEN=…, then run again
The model is already on disk Pass the folder instead of a repo id. The same checks run on its files. rapid-mlx serve ~/models/my-model
Think this is wrong? Every refusal ends with this line. The check can be skipped for one run. Re-run with --no-preflight

Import and quantize: rapid-mlx import

serve and pull never convert anything. Converting an unquantized (bf16 / fp16) safetensors checkpoint to quantized MLX is a separate, explicit command:

$ rapid-mlx import Qwen/Qwen3-0.6B --quantize 4

  Source   bf16 · 1.4 GB · qwen3 (supported)
  Needs    ~2.0 GB free disk  (you have 99 GB)
  Output   ~0.4 GB, fits in 18 GB ✓

  ✓ Smoke test passed
  rapid-mlx serve qwen3-0.6b-4bit-local
  (stored at ~/.rapid-mlx/imports/qwen3-0.6b-4bit-local)

$ rapid-mlx serve qwen3-0.6b-4bit-local
FlagMeaning
sourceA Hugging Face repo id (org/name) or a local model directory.
--quantize BITS2, 3, 4, 6 or 8 (default 4). Affine MLX quantization, group size 64.
--name NAMEThe name to serve it by. The default is <source-name>-<bits>bit; when that's already a catalog or user alias, -local is appended (as above), so an import never shadows an alias.
--forceReplace an existing import of the same name that came from another source or recipe.

What it checks first, from metadata, before any download or conversion:

Cancel-safe and cached.

Where it lives. ~/.rapid-mlx/imports/<name> (or $RAPID_MLX_HOME/imports).

$ rapid-mlx models --cached      # or: rapid-mlx ls
  …
  ── Imported with `rapid-mlx import` (serve or rm by name) ──
  • qwen3-0.6b-4bit-local  330.9 MiB  from Qwen/Qwen3-0.6B · 4-bit

$ rapid-mlx serve qwen3-0.6b-4bit-local
$ rapid-mlx rm qwen3-0.6b-4bit-local       # asks first; -y skips the prompt

Reference

What the check looks at

rapid-mlx serve <org/repo> and rapid-mlx pull <org/repo> run the check for models outside the Rapid-MLX catalog. It reads metadata only: one Hugging Face model_info call, plus config.json fetched into a temporary directory when an architecture verdict needs confirming. No weights are fetched.

CheckWhat it looks atRefuses when
Format The repo's file list. MLX (quantized) and Hugging Face safetensors both run. The repo has no safetensors at all and only GGUF files, or only PyTorch pytorch_model*.bin weights. Repos with weights in per-quantization subfolders, or .npz audio weights, are not refused.
Architecture model_type in config.json, compared with every loader this install has: mlx-lm, mlx-vlm, mlx-audio (when the [audio] extra is installed) and the architectures Rapid-MLX registers itself. The config proves it is a text-generating model that no installed loader can build. The check confirms this against the full config.json before refusing. Repos that ship their own model code are never refused on architecture.
Chat template chat_template.jinja / .json, or the template in the tokenizer config. Never. It's informational. Chat template none — likely a base model; prefer /v1/completions means use text completion, not chat.
Memory fit The root model*.safetensors shards, plus an allowance for KV cache and runtime (weights × 1.2 + 1 GB), against this Mac's physical memory. yes up to 75% of memory, tight above that, no when the weights alone are larger than memory. Only serve, and only for no. pull reports the fit but still downloads (you may be pulling for another Mac), and serve --disk-stream skips this refusal.

The check is conservative. Anything it can't establish gives no verdict, and serve / pull continue without a verdict. That covers network trouble, a Hub answer slower than 10 seconds, a gated repo you can't read yet, and an unusual repo layout. It doesn't run at all for catalog aliases, repos already in your Hugging Face cache, offline mode (HF_HUB_OFFLINE), or pull --bits / --format.

--no-preflight skips the check for one run, on both serve and pull.

What refusals look like

$ rapid-mlx pull Qwen/Qwen3-0.6B-GGUF

  ! Qwen/Qwen3-0.6B-GGUF only has GGUF files (Q8_0).
    Rapid-MLX runs MLX and safetensors weights, not GGUF. Nothing was downloaded.
    MLX builds of the same base model (third-party, not reviewed by Rapid-MLX;
    matched by base model, architecture and size):
      mlx-community/Qwen3-0.6B-4bit · 4-bit · ~0.3 GB · apache-2.0
      lmstudio-community/Qwen3-0.6B-MLX-4bit · 4-bit · ~0.3 GB · apache-2.0
    To try one, check its model card first, then: rapid-mlx pull mlx-community/Qwen3-0.6B-4bit
    (Think this is wrong? Re-run with --no-preflight to skip this check.)

rapid-mlx pull <repo> --format gguf is refused outright: Rapid-MLX cannot run GGUF files, so `--format gguf` is not available.

$ rapid-mlx serve Qwen/Qwen3-235B-A22B

  ✗ Qwen/Qwen3-235B-A22B needs ~526 GB of memory; this Mac has 18 GB.
    Its weights alone are 438 GB. Nothing was downloaded.
    Pick a smaller or more quantized build: rapid-mlx models
    A catalog model that runs well on your Mac:
      rapid-mlx serve qwen3.5-9b-4bit
    (Think this is wrong? Re-run with --no-preflight to skip this check.)

Suggested alternatives (never switched automatically)

A refusal can name up to two things that will run. Neither is ever substituted for the model you asked for. To use one, type the command it prints.

All Hub lookups for suggestions share one short time budget. A slow Hub means fewer suggestions, never a slower refusal.

“Ask us to support it?” (opt-in)

When the check refuses a public Hugging Face model for its architecture or its weight format (GGUF / .bin only), an interactive terminal asks:

  Ask us to support it? This sends the repo id, architecture,
  format and your Rapid-MLX version. Nothing else. [y/N]

The request also carries the refusal's failure class, as the payload below shows. Nothing is sent unless you answer y. In scripts and other non-interactive shells there's no prompt. Re-run with --request (on serve or pull) to send it without asking. When you do, exactly five fields go to rapidmlx.com/api/model-request:

{
  "repo": "Qwen/Qwen3-0.6B-GGUF",
  "model_type": null,
  "format": "gguf",
  "failure": "unsupported_format",
  "version": "0.15.5"
}

Gated and private repos

Until you have access, the check can't read a gated repo's metadata, so it gives no verdict and the download stops with:

  Error: access to '<org/repo>' is gated on Hugging Face.
  Accept the licence or request access at https://huggingface.co/<org/repo>
  Then run `huggingface-cli login` or set `HF_TOKEN`, and try again.

Once you're signed in, the check runs normally. Support requests are never offered for gated or private repos. rapid-mlx import of a gated repo needs the same sign-in.

Local directories

The format, architecture and memory checks run on the folder's own files and config.json. There's nothing to download, and support requests aren't offered for local paths.

$ rapid-mlx serve ~/models/my-finetune
$ rapid-mlx import ~/models/my-finetune --quantize 4 --name my-ft-4bit
$ rapid-mlx serve my-ft-4bit

A local import is keyed by a fingerprint of the folder (file names, sizes and modification times, plus the contents of small files like configs and tokenizers). Re-running on an unchanged folder is a no-op. After you edit the folder, the stale import isn't reused: re-run with --force to replace it, or pick a new --name.

Frequently asked questions

Can Rapid-MLX run a GGUF model?

No. Rapid-MLX runs MLX and safetensors weights. A GGUF-only repo is refused before anything is downloaded and, when the check can verify one, it names an MLX build of the same base model that fits your Mac.

Does the check download the model?

No. It reads Hugging Face metadata only: one model_info call, plus config.json into a temporary directory when it needs to confirm an architecture verdict. If it can't read the metadata, it gives no verdict and the normal download proceeds.

What does a support request send?

It's sent only after you answer y or pass --request, and it contains the repo id, the architecture (model_type), the weight format, the failure class and your Rapid-MLX version. It's never offered for gated, private or local models, or for models that are simply too big for your Mac.

How do I quantize a bf16 model for my Mac?

rapid-mlx import <org/repo> --quantize 4 (2, 3, 4, 6 or 8 bits). It checks disk and memory first, converts in a temporary directory, runs a one-token smoke test, and only then publishes the model under a name you can serve, list and remove.

Next: connect an agent or a chat app to your model.