Reference · rapid-mlx 0.15.7 · ← Back to README

Optional install extras

The base install serves any text model. Add an extra only for the feature you need:

You wantInstall
Images as input (Gemma 4, Qwen-VL, UI-TARS…)rapid-mlx[vision]
Speech: text-to-speech, transcription, musicrapid-mlx[audio]
Image generation and editingrapid-mlx[image] (Python 3.11+)
Video generationrapid-mlx[video] (Python 3.11+, ffmpeg)
/v1/embeddingsrapid-mlx[embeddings]
Typed decision models (rapid-mlx system-one)Laya: rapid-mlx[system-one]; Clef: rapid-mlx[clef] (CLM needs neither)
DFlash or opt-in MTP speculative decodingrapid-mlx[dflash] / rapid-mlx[mtp]
Most of the above at oncerapid-mlx[all]
$ pip install "rapid-mlx[vision]==0.15.7"
$ pip install "rapid-mlx[vision,audio]==0.15.7"    # combine; quote for zsh

When a model needs an extra you don't have, serve says which one and offers to install it (--yes accepts without asking). The lists below are transcribed from [project.optional-dependencies] in the engine's pyproject.toml for 0.15.6.

User-facing extras

[vision] — image input for multimodal models

Required for any model that takes images; text-only models work without it. EmbeddingGemma 2 (embeddinggemma-2-4bit / -bf16) also runs on this runtime. The mlx-vlm pin matches the Desktop app.

$ pip install "rapid-mlx[vision]==0.15.7"

[audio] — speech and music

Every audio alias (see the audio registry) needs this extra: Kokoro, Qwen3-TTS, IndexTTS, F5-TTS, Chatterbox, VibeVoice, VoxCPM, Dia, Whisper, Parakeet, SenseVoice, Qwen3-ASR and the forced aligner.

$ pip install "rapid-mlx[audio]==0.15.7"

[image] — image generation and editing

Text-to-image and image editing (FLUX.2 Klein, FLUX.1-schnell, Qwen-Image, Z-Image and the other image aliases). Apple Silicon and Python 3.11+ only; on Python 3.10 the image preflight explains the upgrade.

$ pip install "rapid-mlx[image]==0.15.7"

[video] — video generation

Text-to-video and image-to-video (LTX-2.5 / 2.3 with synchronized audio, Wan, CogVideoX-Fun) through the asynchronous /v1/videos API. Apple Silicon and Python 3.11+ only, plus ffmpeg.

$ pip install "rapid-mlx[video]==0.15.7"

[embeddings] — /v1/embeddings

Needed for embedding models served through mlx-embeddings, such as embeddinggemma-300m-6bit, attached to a chat server with --embedding-model. EmbeddingGemma 2 does not use this extra — it needs [vision]; see the EmbeddingGemma 2 page.

$ pip install "rapid-mlx[embeddings]==0.15.7"

[dflash] — DFlash speculative decoding

For the verified DFlash pairs (e.g. qwen3.5-27b-8bit, muse-glimmer-30b-8bit) with --speculative-config '{"method":"dflash",…}'. Text-only — it does not pull torch or OpenCV. Drafter weights download on first use.

$ pip install "rapid-mlx[dflash]==0.15.7"

[mtp] — opt-in MTP backends

For the explicit native-MTP backend and Gemma 4 assistant sidecars set through --speculative-config. The default MTP on qualified Qwen aliases does not need it. Text-only.

$ pip install "rapid-mlx[mtp]==0.15.7"

[chat] — Gradio chat UI

Adds the rapid-mlx-chat web UI (separate from the terminal REPL, rapid-mlx chat).

$ pip install "rapid-mlx[chat]==0.15.7"

[computer-use] — native macOS computer use

The accessibility bindings behind the experimental rapid-mlx cua agent.

$ pip install "rapid-mlx[computer-use]==0.15.7"

[system-one] — typed decision models

The Laya backend for rapid-mlx system-one, which serves typed decision models (the CLM backend is built in). The extra installs on Python 3.10, but its laya-mlx backend needs 3.11+.

$ pip install "rapid-mlx[system-one]==0.15.7"

[clef] — Clef decision models for System One

Runtime for the clef and clef-flash decision models in rapid-mlx system-one (--backend clef). Clef's joint head runs on PyTorch with Apple Metal (MPS), so this extra is kept separate from [system-one] and is not part of [all].

$ pip install "rapid-mlx[clef]==0.15.7"
$ rapid-mlx system-one clef-flash --port 8700

[audio-desktop] — the Desktop app's audio subset

The trimmed audio stack inside the Desktop app: transcription and Qwen3 preset-voice speech, without F5 voice cloning or Kokoro's language toolchain. On the CLI, install [audio] instead.

$ pip install "rapid-mlx[audio-desktop]==0.15.7"

[guided] — kept for old install commands

Structured output runs on llguidance, which is a core dependency. This extra only keeps pip install "rapid-mlx[guided]" working.

$ pip install "rapid-mlx[guided]==0.15.7"

[all] — everything a user would install

[vision], [chat], [embeddings], [audio], [computer-use] and [system-one] in one install.

$ pip install "rapid-mlx[all]==0.15.7"

Its packages also cover [dflash], [mtp] and [guided]. It does not include [image] or [video] (Python 3.11+ only), [clef], or the contributor extras below — add those explicitly.

Contributor extras

[test] — what pytest tests/ needs

Enough to import and collect every test; kept narrow so the PR validator can rebuild a venv without linters.

$ pip install "rapid-mlx[test]==0.15.7"

[dev] — tests plus linters

Everything in [test] plus Ruff and mypy.

$ pip install "rapid-mlx[dev]==0.15.7"

[ci-linux] and [ci-apple] are the engine CI's own environments, and [mirror] (boto3) is maintainer tooling for the model mirror; users never need them.

Check what's installed

$ rapid-mlx doctor --only packages.optional

The doctor lists each optional package with its version, or the exact install command when it is missing.

Next steps