Optional install extras
The base install serves any text model. Add an extra only for the feature you need:
| You want | Install |
|---|---|
| Images as input (Gemma 4, Qwen-VL, UI-TARS…) | rapid-mlx[vision] |
| Speech: text-to-speech, transcription, music | rapid-mlx[audio] |
| Image generation and editing | rapid-mlx[image] (Python 3.11+) |
| Video generation | rapid-mlx[video] (Python 3.11+, ffmpeg) |
/v1/embeddings | rapid-mlx[embeddings] |
Typed decision models (rapid-mlx system-one) | Laya: rapid-mlx[system-one]; Clef: rapid-mlx[clef] (CLM needs neither) |
| DFlash or opt-in MTP speculative decoding | rapid-mlx[dflash] / rapid-mlx[mtp] |
| Most of the above at once | rapid-mlx[all] |
$ pip install "rapid-mlx[vision]==0.15.7" $ pip install "rapid-mlx[vision,audio]==0.15.7" # combine; quote for zsh
When a model needs an extra you don't have, serve says
which one and offers to install it (--yes accepts without
asking). The lists below are transcribed from
[project.optional-dependencies] in the engine's
pyproject.toml
for 0.15.6.
User-facing extras
[vision] — image input for multimodal models
Required for any model that takes images; text-only models work without it. EmbeddingGemma 2 (embeddinggemma-2-4bit / -bf16) also runs on this runtime. The mlx-vlm pin matches the Desktop app.
mlx-vlm==0.7.2opencv-python>=4.8.0torch>=2.3.0,torchvision>=0.18.0pillow>=10.0.0
$ pip install "rapid-mlx[vision]==0.15.7"
[audio] — speech and music
Every audio alias (see the audio registry) needs this extra: Kokoro, Qwen3-TTS, IndexTTS, F5-TTS, Chatterbox, VibeVoice, VoxCPM, Dia, Whisper, Parakeet, SenseVoice, Qwen3-ASR and the forced aligner.
mlx-audio>=0.5.3,<0.6f5-tts-mlx==0.2.6— EN+ZH voice cloningsounddevice>=0.4.0,soundfile>=0.12.0,scipy>=1.10.0,numba>=0.57.0tiktoken>=0.5.0,misaki[zh,ja]>=0.5.0,spacy>=3.7.0,num2words>=0.5.0,loguru>=0.7.0espeakng-loader>=0.2.0,phonemizer-fork>=3.3.0(the fork keeps the API Kokoro's G2P calls)cn2an>=0.5.0— Chinese number conversion
$ pip install "rapid-mlx[audio]==0.15.7"
[image] — image generation and editing
Text-to-image and image editing (FLUX.2 Klein, FLUX.1-schnell, Qwen-Image, Z-Image and the other image aliases). Apple Silicon and Python 3.11+ only; on Python 3.10 the image preflight explains the upgrade.
mflux>=0.20.0,<0.21(Python 3.11+)mlx-vlm==0.7.2(Python 3.11+)pillow>=10.0.0,sentencepiece>=0.2.0
$ pip install "rapid-mlx[image]==0.15.7"
[video] — video generation
Text-to-video and image-to-video (LTX-2.5 / 2.3 with synchronized audio, Wan, CogVideoX-Fun) through the asynchronous /v1/videos API. Apple Silicon and Python 3.11+ only, plus ffmpeg.
mlx-video-with-audio==0.1.36mlx-arsenal>=0.10.1imageio[ffmpeg]>=2.34.0pillow>=10.0.0,sentencepiece>=0.2.0
$ pip install "rapid-mlx[video]==0.15.7"
[embeddings] — /v1/embeddings
Needed for embedding models served through mlx-embeddings, such as embeddinggemma-300m-6bit, attached to a chat server with --embedding-model. EmbeddingGemma 2 does not use this extra — it needs [vision]; see the EmbeddingGemma 2 page.
mlx-embeddings>=0.1.0
$ pip install "rapid-mlx[embeddings]==0.15.7"
[dflash] — DFlash speculative decoding
For the verified DFlash pairs (e.g. qwen3.5-27b-8bit, muse-glimmer-30b-8bit) with --speculative-config '{"method":"dflash",…}'. Text-only — it does not pull torch or OpenCV. Drafter weights download on first use.
mlx-vlm==0.7.2
$ pip install "rapid-mlx[dflash]==0.15.7"
[mtp] — opt-in MTP backends
For the explicit native-MTP backend and Gemma 4 assistant sidecars set through --speculative-config. The default MTP on qualified Qwen aliases does not need it. Text-only.
mlx-vlm==0.7.2pillow>=10.0.0
$ pip install "rapid-mlx[mtp]==0.15.7"
[chat] — Gradio chat UI
Adds the rapid-mlx-chat web UI (separate from the terminal REPL, rapid-mlx chat).
gradio>=4.0.0pytz>=2024.1
$ pip install "rapid-mlx[chat]==0.15.7"
[computer-use] — native macOS computer use
The accessibility bindings behind the experimental rapid-mlx cua agent.
pyobjc-framework-ApplicationServices==12.2.2pyobjc-framework-Quartz==12.2.2
$ pip install "rapid-mlx[computer-use]==0.15.7"
[system-one] — typed decision models
The Laya backend for rapid-mlx system-one, which serves typed decision models (the CLM backend is built in). The extra installs on Python 3.10, but its laya-mlx backend needs 3.11+.
laya-mlx>=0.2,<0.3(Python 3.11+)
$ pip install "rapid-mlx[system-one]==0.15.7"
[clef] — Clef decision models for System One
Runtime for the clef and clef-flash decision models in rapid-mlx system-one (--backend clef). Clef's joint head runs on PyTorch with Apple Metal (MPS), so this extra is kept separate from [system-one] and is not part of [all].
torch>=2.11.0,torchvision>=0.26.0(macOS)transformers>=5.10.2,!=5.13.0,<5.16,accelerate>=1.4.0packaging>=24.0,pillow>=10.0.0
$ pip install "rapid-mlx[clef]==0.15.7" $ rapid-mlx system-one clef-flash --port 8700
[audio-desktop] — the Desktop app's audio subset
The trimmed audio stack inside the Desktop app: transcription and Qwen3 preset-voice speech, without F5 voice cloning or Kokoro's language toolchain. On the CLI, install [audio] instead.
mlx-audio>=0.5.3,<0.6soundfile>=0.12.0
$ pip install "rapid-mlx[audio-desktop]==0.15.7"
[guided] — kept for old install commands
Structured output runs on llguidance, which is a core dependency. This extra only keeps pip install "rapid-mlx[guided]" working.
llguidance>=1.7.6
$ pip install "rapid-mlx[guided]==0.15.7"
[all] — everything a user would install
[vision], [chat], [embeddings],
[audio], [computer-use] and
[system-one] in one install.
$ pip install "rapid-mlx[all]==0.15.7"
Its packages also cover [dflash], [mtp] and
[guided]. It does not include [image]
or [video] (Python 3.11+ only), [clef], or the contributor extras
below — add those explicitly.
Contributor extras
[test] — what pytest tests/ needs
Enough to import and collect every test; kept narrow so the PR validator can rebuild a venv without linters.
pytest>=7.0.0,pytest-asyncio>=0.21.0(asyncio_mode = auto)referencing>=0.28.4,prometheus_client>=0.16.0,aiohttp>=3.9.0,pillow>=10.0.0tomli>=2.0.1(Python < 3.11)- Darwin only:
mlx-vlm==0.7.2,mlx-audio>=0.5.3,<0.6 hypothesis>=6.100.0,pytest-cov>=4.0.0,diff-cover>=8.0.0,<9.0.0
$ pip install "rapid-mlx[test]==0.15.7"
[dev] — tests plus linters
Everything in [test] plus Ruff and mypy.
- Everything in
[test] ruff>=0.1.0,mypy>=1.0.0
$ pip install "rapid-mlx[dev]==0.15.7"
[ci-linux] and [ci-apple] are the engine CI's
own environments, and [mirror] (boto3) is
maintainer tooling for the model mirror; users never need them.
Check what's installed
$ rapid-mlx doctor --only packages.optional
The doctor lists each optional package with its version, or the exact install command when it is missing.