Models · family

Audio

44 audio aliases · TTS (Kokoro / Qwen3-TTS / Chatterbox / VibeVoice / F5 / IndexTTS / VoxCPM / Dia) + STT (Whisper / Parakeet / SenseVoice / Qwen3-ASR / ForcedAligner).

rapid-mlx serves OpenAI-compatible /v1/audio/transcriptions and /v1/audio/speech endpoints — drop-in for any OpenAI SDK client. 44 aliases ship today across two modalities: text-to-speech (Kokoro, Qwen3-TTS, Chatterbox, VibeVoice, F5, IndexTTS, VoxCPM, Dia) and speech-to-text (Whisper, Parakeet, SenseVoice, Qwen3-ASR, ForcedAligner). All aliases route through the same mlx_audio-backed runtime so swapping models is one CLI flag. rapid-mlx pull fetches a voice model's runtime assets as part of the pull, so a machine that pulled while online can synthesise while offline.

family
Audio
aliases
44
lines
13
install
rapid-mlx serve <alias>
OpenAI base URL
http://localhost:8000/v1

Download

Every alias on this page downloads with one command — the pull buttons in the tables below copy it. All 44 aliases on this page are mirrored on the rapid-mlx CDN — with automatic mid-pull fallback to Hugging Face if a mirror file slows down. Weights land in the standard Hugging Face cache, and rapid-mlx serve pulls automatically on first use. Live mirror status →

Usage

The audio aliases route through OpenAI-compatible endpoints — any OpenAI SDK works without per-model config.

# Text-to-speech
curl -X POST http://localhost:8000/v1/audio/speech \
  -H "Content-Type: application/json" \
  -d '{"model": "kokoro", "input": "hello from rapid-mlx"}' \
  --output hello.wav

# Speech-to-text (multipart upload)
curl -X POST http://localhost:8000/v1/audio/transcriptions \
  -H "Content-Type: multipart/form-data" \
  -F file=@hello.wav \
  -F model=whisper-large-v3-turbo

System-wide local dictation

Rapid-MLX Desktop turns the STT aliases on this page into a system-wide Mac dictation tool. Tap Right Command or Right Option in any app, speak, then tap again: Rapid records only during that session, transcribes through the local /v1/audio/transcriptions endpoint, and pastes the result at the cursor. Raw audio is not archived unless the user explicitly enables audio archiving.

The Desktop model picker explains the practical tradeoffs: Whisper Large v3 Turbo is the balanced multilingual default; Qwen3-ASR targets Chinese/English code-switching and vocabulary hints; SenseVoice is the fast Chinese, Cantonese, Japanese and Korean option; Parakeet v2 is English-only, while v3 supports 25 European languages but not Chinese. See the complete local dictation guide or the 中文本地听写指南.

Lines in this family

Text-to-speech (TTS) · 24 aliases

OpenAI-compat endpoint: POST /v1/audio/speech. Pass model (any alias from the tables below), input (the text to synthesise), and optionally voice (per-family voice id — falls back to the alias's default voice when omitted).

Kokoro · 7 aliases

Lightweight 82M open-weights TTS — the default rapid-mlx TTS pick. Three quants (bf16 / 8-bit / 4-bit) with short- and long-form aliases for each.

aliashf repodefault voicenotesget it
kokoromlx-community/Kokoro-82M-bf16af_heartDefault Kokoro 82M (bf16). Drop-in OpenAI-style TTS.CDN
kokoro-4bitmlx-community/Kokoro-82M-4bitaf_heartShort-form 4-bit alias.CDN
kokoro-82mmlx-community/Kokoro-82M-bf16af_heartLong-form alias for ``kokoro`` — same HF id.CDN
kokoro-82m-4bitmlx-community/Kokoro-82M-4bitaf_heart4-bit quant of Kokoro 82M — smaller download, mild quality drop.CDN
kokoro-82m-8bitmlx-community/Kokoro-82M-8bitaf_heart8-bit quant of Kokoro 82M.CDN
kokoro-82m-bf16mlx-community/Kokoro-82M-bf16af_heartExplicit bf16 quant of Kokoro 82M.CDN
kokoro-8bitmlx-community/Kokoro-82M-8bitaf_heartShort-form 8-bit alias.CDN

Chatterbox · 2 aliases

Multilingual expressive TTS with voice cloning via reference audio. Both fp16 and 4-bit ship.

aliashf repodefault voicenotesget it
chatterboxmlx-community/chatterbox-turbo-fp16defaultMultilingual expressive TTS. Voice cloning via reference audio.CDN
chatterbox-4bitmlx-community/chatterbox-turbo-4bitdefault4-bit quant of Chatterbox Turbo.CDN

VibeVoice · 2 aliases

Realtime low-latency 0.5B TTS. Per-language voice caches; en-Grace_woman is the canonical English-female default.

aliashf repodefault voicenotesget it
vibevoicemlx-community/VibeVoice-Realtime-0.5B-4biten-Grace_womanRealtime low-latency TTS, 0.5B 4bit. Per-language voice caches; en-Grace_woman is the canonical English-female default. Full set: en-{Grace_woman,Mike_man,Carter_man,Davis_man,Frank_man,Emma_woman} + de/fr/it/jp/kr/nl/pl/pt/sp Spk0/Spk1 pairs + in-Samuel_man.CDN
vibevoice-realtimemlx-community/VibeVoice-Realtime-0.5B-4biten-Grace_womanLong-form alias for ``vibevoice``.CDN

VoxCPM · 1 alias

Chinese/English high-quality TTS, CPM-based. Single alias.

aliashf repodefault voicenotesget it
voxcpmmlx-community/VoxCPM1.5defaultEnglish (experimental) TTS (CPM-based). WARNING: Chinese is currently BROKEN in this MLX port — Chinese input produces Thai-script gibberish plus long runaway generation, so do NOT use voxcpm for Chinese. For Chinese TTS use `f5-tts-zh` (EN+ZH, cloneable) or `qwen3-tts` (named Chinese speakers) instead. English-only is usable but best-effort; report upstream if synth fails.CDN

Dia · 1 alias

Multi-speaker dialogue TTS in 4-bit. Best-effort — covered by the probe but mlx_audio support is experimental.

aliashf repodefault voicenotesget it
diamlx-community/Dia-1.6B-4bitdefaultDia 1.6B multi-speaker dialogue TTS (4-bit). Best-effort — covered by the audio probe but mlx_audio support is experimental; report bugs upstream if synth fails.CDN

Qwen3-TTS · 8 aliases

Alibaba's TTS line: Base clones a voice from a reference clip plus its transcript; CustomVoice modulates fixed speakers with an instructions string; VoiceDesign builds a voice from a natural-language description alone — and voice_seed brings the exact same narrator back on the next call.

aliashf repodefault voicenotesget it
qwen3-ttsmlx-community/Qwen3-TTS-12Hz-1.7B-CustomVoice-bf16SerenaAlibaba Qwen3-TTS 1.7B CustomVoice (bf16). Multilingual predefined speakers (Chinese: Vivian, Serena, Uncle_Fu, Dylan, Eric; English: Ryan, Aiden; Japanese: Ono_Anna; Korean: Sohee) matched case-insensitively, with emotion/style control via the OpenAI `instructions` field. The CustomVoice variant is used (not 0.6B-Base, which is voice-cloning-only and needs a reference clip).CDN
qwen3-tts-4bitmlx-community/Qwen3-TTS-12Hz-1.7B-CustomVoice-4bitSerena4-bit quant of Qwen3-TTS 1.7B CustomVoice — smallest download.CDN
qwen3-tts-6bitmlx-community/Qwen3-TTS-12Hz-1.7B-CustomVoice-6bitSerena6-bit quant of Qwen3-TTS 1.7B CustomVoice — smaller download, mild quality drop.CDN
qwen3-tts-clonemlx-community/Qwen3-TTS-12Hz-1.7B-Base-bf16cloneAlibaba Qwen3-TTS 1.7B **Base** — zero-shot Chinese/multilingual voice CLONING. Requires ref_audio (a clean 5-10s reference clip at the model's native sample rate) plus ref_text (its exact transcript); ignores predefined speakers. Lets a channel reuse ONE consistent branded narrator voice. Distinct from the CustomVoice variant (named speakers + `instruct` emotion, no cloning).CDN
qwen3-tts-customvoicemlx-community/Qwen3-TTS-12Hz-1.7B-CustomVoice-bf16SerenaLong-form alias for ``qwen3-tts`` — same HF id (1.7B CustomVoice bf16).CDN
qwen3-tts-voicedesignmlx-community/Qwen3-TTS-12Hz-1.7B-VoiceDesign-bf16describeQwen3-TTS 1.7B VoiceDesign (bf16). The VOICE-DESIGN sibling of CustomVoice: instead of picking a named speaker, the ENTIRE voice — timbre, gender, age, accent, emotion, prosody — is described in natural language via the OpenAI `instructions` field (mapped to mlx_audio's mandatory `instruct` arg; the weights self-dispatch to `generate_voice_design`). Richer style control than CustomVoice's `instruct` (which only modulates a fixed speaker). Has NO named speakers, so `voice` is ignored — the voice surface is the single `describe` sentinel (like F5's `clone`) and `default_voice` is that sentinel so the voice-omitted path validates. Supply a description, e.g. `instructions="a warm, low female narrator, calm and measured"`.CDN
qwen3-tts-voicedesign-4bitmlx-community/Qwen3-TTS-12Hz-1.7B-VoiceDesign-4bitdescribe4-bit quant of Qwen3-TTS 1.7B VoiceDesign — smallest download.CDN
qwen3-tts-voicedesign-8bitmlx-community/Qwen3-TTS-12Hz-1.7B-VoiceDesign-8bitdescribe8-bit quant of Qwen3-TTS 1.7B VoiceDesign — smaller download, mild quality drop.CDN

F5-TTS · 1 alias

Pure-MLX expressive voice cloning (no torch) — closes the Chinese expressive-cloning gap that Qwen3-TTS reads flat on.

aliashf repodefault voicenotesget it
f5-tts-zhlucasnewman/f5-tts-mlxcloneF5-TTS (pure-MLX, no torch): EN+ZH multilingual, zero-shot voice CLONING via ref_audio (24kHz clip) + ref_text. Fills the Chinese expressive/cloneable gap (Qwen3-TTS is flat, Chatterbox English-only). With no ref it uses a packaged default voice; supply a Chinese ref for a natural Chinese narrator.CDN

IndexTTS · 2 aliases

Clones a voice from the reference clip alone — no transcript needed, the least fiddly cloning path.

aliashf repodefault voicenotesget it
indexttsmlx-community/IndexTTS-1.5cloneIndexTTS 1.5 zero-shot voice cloning. Requires ref_audio; the reference transcript is not required. No predefined speakers.CDN
indextts-1.5mlx-community/IndexTTS-1.5cloneLong-form alias for indextts — same IndexTTS 1.5 MLX checkpoint.CDN

Speech-to-text (STT) · 20 aliases

OpenAI-compat endpoint: POST /v1/audio/transcriptions. Send multipart/form-data with file (the audio bytes) and model (any alias from the tables below). Optional language hint accepted by Whisper aliases.

Whisper · 8 aliases

OpenAI's multilingual STT family — tiny → large-v3 + large-v3-turbo. whisper-1 is wired as an OpenAI-compat placeholder pointing at large-v3.

aliashf repolanguagesnotesget it
whispermlx-community/whisper-large-v3-mlxmultilingualBare ``whisper`` -> largest-v3. Matches OpenAI's ``whisper-1`` placeholder.CDN
whisper-1mlx-community/whisper-large-v3-mlxmultilingualOpenAI-spec id from the legacy Whisper API. Same destination as ``whisper-large-v3``.CDN
whisper-basemlx-community/whisper-base-mlxmultilingualWhisper base.CDN
whisper-large-v3mlx-community/whisper-large-v3-mlxmultilingualWhisper large-v3 (multilingual, 99+ languages).CDN
whisper-large-v3-turbomlx-community/whisper-large-v3-turbomultilingualWhisper large-v3 turbo (faster decode, slight quality drop).CDN
whisper-mediummlx-community/whisper-medium-mlxmultilingualWhisper medium.CDN
whisper-smallmlx-community/whisper-small-mlxmultilingualWhisper small.CDN
whisper-tinymlx-community/whisper-tiny-mlxmultilingualWhisper tiny (smallest, fastest, lowest accuracy).CDN

Parakeet · 5 aliases

NVIDIA Parakeet TDT 0.6B — v2 is English-only; v3 supports 25 European languages with automatic language detection. Both checkpoints ship.

aliashf repolanguagesnotesget it
parakeetmlx-community/parakeet-tdt-0.6b-v2enNVIDIA Parakeet TDT 0.6B v2 (English, fastest STT).CDN
parakeet-tdt-0.6bmlx-community/parakeet-tdt-0.6b-v2enLong-form alias for ``parakeet``.CDN
parakeet-tdt-0.6b-v2mlx-community/parakeet-tdt-0.6b-v2enParakeet TDT 0.6B v2.CDN
parakeet-tdt-0.6b-v3mlx-community/parakeet-tdt-0.6b-v3en,es,fr,de,bg,hr,cs,da,nl,et,fi,el,hu,it,lv,lt,mt,pl,pt,ro,sk,sl,sv,ru,ukParakeet TDT 0.6B v3 with automatic language detection. Chinese is not supported.CDN
parakeet-v3mlx-community/parakeet-tdt-0.6b-v3en,es,fr,de,bg,hr,cs,da,nl,et,fi,el,hu,it,lv,lt,mt,pl,pt,ro,sk,sl,sv,ru,ukParakeet TDT 0.6B v3 with automatic language detection. Chinese is not supported.CDN

Qwen3-ASR · 3 aliases

Alibaba's Qwen3-ASR — an audio LLM rather than an encoder-decoder ASR model. Keeps Chinese/English code-switching straight without a decoding hint and takes hotwords as a system_prompt. Bare qwen3-asr resolves to the 1.7B; a 0.6B ships too.

aliashf repolanguagesnotesget it
qwen3-asrmlx-community/Qwen3-ASR-1.7B-5bitmultilingualQwen3-ASR 1.7B (5-bit). An audio LLM rather than an encoder-decoder ASR model: keeps Chinese/English code-switching straight without a decoding hint, and emits normal punctuation plus spacing between scripts. Takes hotwords as ``system_prompt`` (STTEngine maps the neutral ``context`` argument onto it). Bare ``qwen3-asr`` -> the 1.7B.CDN
qwen3-asr-0.6bmlx-community/Qwen3-ASR-0.6B-5bitmultilingualSmaller Qwen3-ASR. Faster than the 1.7B with a small accuracy cost.CDN
qwen3-asr-1.7bmlx-community/Qwen3-ASR-1.7B-5bitmultilingualExplicit long-form alias for ``qwen3-asr`` — same HF id.CDN

SenseVoice · 2 aliases

FunAudioLLM SenseVoice Small (~234M, non-autoregressive CTC) — the strongest pick for Chinese, Cantonese, Japanese and Korean, and it emits per-segment emotion and audio-event tags alongside the transcript.

aliashf repolanguagesnotesget it
sensevoicemlx-community/SenseVoiceSmallzh,yue,ja,ko,enFunAudioLLM SenseVoice Small (~234M, non-autoregressive CTC). Dedicated fast Asian-language ASR (strong on Chinese/Cantonese/Japanese/Korean). Also emits per-segment emotion + audio-event tags. Bare ``sensevoice`` -> the small model.CDN
sensevoice-smallmlx-community/SenseVoiceSmallzh,yue,ja,ko,enExplicit long-form alias for ``sensevoice`` — same HF id.CDN

Qwen3-ForcedAligner · 2 aliases

Give it audio plus the transcript you already have and it returns per-character start/end times. It never guesses at words, so it cannot mis-hear them — exactly what karaoke captions and beat-synced editing need.

aliashf repolanguagesnotesget it
qwen3-alignermlx-community/Qwen3-ForcedAligner-0.6B-8bitmultilingualForced ALIGNMENT (not ASR): given audio + the KNOWN transcript, returns per-character start/end times with zero recognition error. Use STTEngine.align(audio, text, language) or /v1/audio/transcriptions with a `text` field. Powers karaoke captions + beat-synced editing; unoccupied on MLX.CDN
qwen3-forced-alignermlx-community/Qwen3-ForcedAligner-0.6B-8bitmultilingualLong-form alias for ``qwen3-aligner`` — same HF id.CDN

Notes & caveats

External reading

Where next