Audio
44 audio aliases · TTS (Kokoro / Qwen3-TTS / Chatterbox / VibeVoice / F5 / IndexTTS / VoxCPM / Dia) + STT (Whisper / Parakeet / SenseVoice / Qwen3-ASR / ForcedAligner).
rapid-mlx serves OpenAI-compatible /v1/audio/transcriptions and /v1/audio/speech endpoints — drop-in for any OpenAI SDK client. 44 aliases ship today across two modalities: text-to-speech (Kokoro, Qwen3-TTS, Chatterbox, VibeVoice, F5, IndexTTS, VoxCPM, Dia) and speech-to-text (Whisper, Parakeet, SenseVoice, Qwen3-ASR, ForcedAligner). All aliases route through the same mlx_audio-backed runtime so swapping models is one CLI flag. rapid-mlx pull fetches a voice model's runtime assets as part of the pull, so a machine that pulled while online can synthesise while offline.
- family
- Audio
- aliases
- 44
- lines
- 13
- install
- rapid-mlx serve <alias>
- OpenAI base URL
- http://localhost:8000/v1
Download
Every alias on this page downloads with one command — the pull buttons in the tables below copy it. All 44 aliases on this page are mirrored on the rapid-mlx CDN — with automatic mid-pull fallback to Hugging Face if a mirror file slows down. Weights land in the standard Hugging Face cache, and rapid-mlx serve pulls automatically on first use. Live mirror status →
Usage
The audio aliases route through OpenAI-compatible endpoints — any OpenAI SDK works without per-model config.
# Text-to-speech
curl -X POST http://localhost:8000/v1/audio/speech \
-H "Content-Type: application/json" \
-d '{"model": "kokoro", "input": "hello from rapid-mlx"}' \
--output hello.wav
# Speech-to-text (multipart upload)
curl -X POST http://localhost:8000/v1/audio/transcriptions \
-H "Content-Type: multipart/form-data" \
-F file=@hello.wav \
-F model=whisper-large-v3-turbo
System-wide local dictation
Rapid-MLX Desktop turns the STT aliases on this page into a system-wide Mac dictation tool. Tap Right Command or Right Option in any app, speak, then tap again: Rapid records only during that session, transcribes through the local /v1/audio/transcriptions endpoint, and pastes the result at the cursor. Raw audio is not archived unless the user explicitly enables audio archiving.
The Desktop model picker explains the practical tradeoffs: Whisper Large v3 Turbo is the balanced multilingual default; Qwen3-ASR targets Chinese/English code-switching and vocabulary hints; SenseVoice is the fast Chinese, Cantonese, Japanese and Korean option; Parakeet v2 is English-only, while v3 supports 25 European languages but not Chinese. See the complete local dictation guide or the 中文本地听写指南.
Lines in this family
- Text-to-speech (TTS) · Kokoro · Chatterbox · VibeVoice · VoxCPM · Dia · Qwen3-TTS · F5-TTS · IndexTTS
- Speech-to-text (STT) · Whisper · Parakeet · Qwen3-ASR · SenseVoice · Qwen3-ForcedAligner
Text-to-speech (TTS) · 24 aliases
OpenAI-compat endpoint: POST /v1/audio/speech. Pass model (any alias from the tables below), input (the text to synthesise), and optionally voice (per-family voice id — falls back to the alias's default voice when omitted).
Kokoro · 7 aliases
Lightweight 82M open-weights TTS — the default rapid-mlx TTS pick. Three quants (bf16 / 8-bit / 4-bit) with short- and long-form aliases for each.
| alias | hf repo | default voice | notes | get it |
|---|---|---|---|---|
| kokoro | mlx-community/Kokoro-82M-bf16 | af_heart | Default Kokoro 82M (bf16). Drop-in OpenAI-style TTS. | CDN |
| kokoro-4bit | mlx-community/Kokoro-82M-4bit | af_heart | Short-form 4-bit alias. | CDN |
| kokoro-82m | mlx-community/Kokoro-82M-bf16 | af_heart | Long-form alias for ``kokoro`` — same HF id. | CDN |
| kokoro-82m-4bit | mlx-community/Kokoro-82M-4bit | af_heart | 4-bit quant of Kokoro 82M — smaller download, mild quality drop. | CDN |
| kokoro-82m-8bit | mlx-community/Kokoro-82M-8bit | af_heart | 8-bit quant of Kokoro 82M. | CDN |
| kokoro-82m-bf16 | mlx-community/Kokoro-82M-bf16 | af_heart | Explicit bf16 quant of Kokoro 82M. | CDN |
| kokoro-8bit | mlx-community/Kokoro-82M-8bit | af_heart | Short-form 8-bit alias. | CDN |
Chatterbox · 2 aliases
Multilingual expressive TTS with voice cloning via reference audio. Both fp16 and 4-bit ship.
| alias | hf repo | default voice | notes | get it |
|---|---|---|---|---|
| chatterbox | mlx-community/chatterbox-turbo-fp16 | default | Multilingual expressive TTS. Voice cloning via reference audio. | CDN |
| chatterbox-4bit | mlx-community/chatterbox-turbo-4bit | default | 4-bit quant of Chatterbox Turbo. | CDN |
VibeVoice · 2 aliases
Realtime low-latency 0.5B TTS. Per-language voice caches; en-Grace_woman is the canonical English-female default.
| alias | hf repo | default voice | notes | get it |
|---|---|---|---|---|
| vibevoice | mlx-community/VibeVoice-Realtime-0.5B-4bit | en-Grace_woman | Realtime low-latency TTS, 0.5B 4bit. Per-language voice caches; en-Grace_woman is the canonical English-female default. Full set: en-{Grace_woman,Mike_man,Carter_man,Davis_man,Frank_man,Emma_woman} + de/fr/it/jp/kr/nl/pl/pt/sp Spk0/Spk1 pairs + in-Samuel_man. | CDN |
| vibevoice-realtime | mlx-community/VibeVoice-Realtime-0.5B-4bit | en-Grace_woman | Long-form alias for ``vibevoice``. | CDN |
VoxCPM · 1 alias
Chinese/English high-quality TTS, CPM-based. Single alias.
| alias | hf repo | default voice | notes | get it |
|---|---|---|---|---|
| voxcpm | mlx-community/VoxCPM1.5 | default | English (experimental) TTS (CPM-based). WARNING: Chinese is currently BROKEN in this MLX port — Chinese input produces Thai-script gibberish plus long runaway generation, so do NOT use voxcpm for Chinese. For Chinese TTS use `f5-tts-zh` (EN+ZH, cloneable) or `qwen3-tts` (named Chinese speakers) instead. English-only is usable but best-effort; report upstream if synth fails. | CDN |
Dia · 1 alias
Multi-speaker dialogue TTS in 4-bit. Best-effort — covered by the probe but mlx_audio support is experimental.
| alias | hf repo | default voice | notes | get it |
|---|---|---|---|---|
| dia | mlx-community/Dia-1.6B-4bit | default | Dia 1.6B multi-speaker dialogue TTS (4-bit). Best-effort — covered by the audio probe but mlx_audio support is experimental; report bugs upstream if synth fails. | CDN |
Qwen3-TTS · 8 aliases
Alibaba's TTS line: Base clones a voice from a reference clip plus its transcript; CustomVoice modulates fixed speakers with an instructions string; VoiceDesign builds a voice from a natural-language description alone — and voice_seed brings the exact same narrator back on the next call.
| alias | hf repo | default voice | notes | get it |
|---|---|---|---|---|
| qwen3-tts | mlx-community/Qwen3-TTS-12Hz-1.7B-CustomVoice-bf16 | Serena | Alibaba Qwen3-TTS 1.7B CustomVoice (bf16). Multilingual predefined speakers (Chinese: Vivian, Serena, Uncle_Fu, Dylan, Eric; English: Ryan, Aiden; Japanese: Ono_Anna; Korean: Sohee) matched case-insensitively, with emotion/style control via the OpenAI `instructions` field. The CustomVoice variant is used (not 0.6B-Base, which is voice-cloning-only and needs a reference clip). | CDN |
| qwen3-tts-4bit | mlx-community/Qwen3-TTS-12Hz-1.7B-CustomVoice-4bit | Serena | 4-bit quant of Qwen3-TTS 1.7B CustomVoice — smallest download. | CDN |
| qwen3-tts-6bit | mlx-community/Qwen3-TTS-12Hz-1.7B-CustomVoice-6bit | Serena | 6-bit quant of Qwen3-TTS 1.7B CustomVoice — smaller download, mild quality drop. | CDN |
| qwen3-tts-clone | mlx-community/Qwen3-TTS-12Hz-1.7B-Base-bf16 | clone | Alibaba Qwen3-TTS 1.7B **Base** — zero-shot Chinese/multilingual voice CLONING. Requires ref_audio (a clean 5-10s reference clip at the model's native sample rate) plus ref_text (its exact transcript); ignores predefined speakers. Lets a channel reuse ONE consistent branded narrator voice. Distinct from the CustomVoice variant (named speakers + `instruct` emotion, no cloning). | CDN |
| qwen3-tts-customvoice | mlx-community/Qwen3-TTS-12Hz-1.7B-CustomVoice-bf16 | Serena | Long-form alias for ``qwen3-tts`` — same HF id (1.7B CustomVoice bf16). | CDN |
| qwen3-tts-voicedesign | mlx-community/Qwen3-TTS-12Hz-1.7B-VoiceDesign-bf16 | describe | Qwen3-TTS 1.7B VoiceDesign (bf16). The VOICE-DESIGN sibling of CustomVoice: instead of picking a named speaker, the ENTIRE voice — timbre, gender, age, accent, emotion, prosody — is described in natural language via the OpenAI `instructions` field (mapped to mlx_audio's mandatory `instruct` arg; the weights self-dispatch to `generate_voice_design`). Richer style control than CustomVoice's `instruct` (which only modulates a fixed speaker). Has NO named speakers, so `voice` is ignored — the voice surface is the single `describe` sentinel (like F5's `clone`) and `default_voice` is that sentinel so the voice-omitted path validates. Supply a description, e.g. `instructions="a warm, low female narrator, calm and measured"`. | CDN |
| qwen3-tts-voicedesign-4bit | mlx-community/Qwen3-TTS-12Hz-1.7B-VoiceDesign-4bit | describe | 4-bit quant of Qwen3-TTS 1.7B VoiceDesign — smallest download. | CDN |
| qwen3-tts-voicedesign-8bit | mlx-community/Qwen3-TTS-12Hz-1.7B-VoiceDesign-8bit | describe | 8-bit quant of Qwen3-TTS 1.7B VoiceDesign — smaller download, mild quality drop. | CDN |
F5-TTS · 1 alias
Pure-MLX expressive voice cloning (no torch) — closes the Chinese expressive-cloning gap that Qwen3-TTS reads flat on.
| alias | hf repo | default voice | notes | get it |
|---|---|---|---|---|
| f5-tts-zh | lucasnewman/f5-tts-mlx | clone | F5-TTS (pure-MLX, no torch): EN+ZH multilingual, zero-shot voice CLONING via ref_audio (24kHz clip) + ref_text. Fills the Chinese expressive/cloneable gap (Qwen3-TTS is flat, Chatterbox English-only). With no ref it uses a packaged default voice; supply a Chinese ref for a natural Chinese narrator. | CDN |
IndexTTS · 2 aliases
Clones a voice from the reference clip alone — no transcript needed, the least fiddly cloning path.
| alias | hf repo | default voice | notes | get it |
|---|---|---|---|---|
| indextts | mlx-community/IndexTTS-1.5 | clone | IndexTTS 1.5 zero-shot voice cloning. Requires ref_audio; the reference transcript is not required. No predefined speakers. | CDN |
| indextts-1.5 | mlx-community/IndexTTS-1.5 | clone | Long-form alias for indextts — same IndexTTS 1.5 MLX checkpoint. | CDN |
Speech-to-text (STT) · 20 aliases
OpenAI-compat endpoint: POST /v1/audio/transcriptions. Send multipart/form-data with file (the audio bytes) and model (any alias from the tables below). Optional language hint accepted by Whisper aliases.
Whisper · 8 aliases
OpenAI's multilingual STT family — tiny → large-v3 + large-v3-turbo. whisper-1 is wired as an OpenAI-compat placeholder pointing at large-v3.
| alias | hf repo | languages | notes | get it |
|---|---|---|---|---|
| whisper | mlx-community/whisper-large-v3-mlx | multilingual | Bare ``whisper`` -> largest-v3. Matches OpenAI's ``whisper-1`` placeholder. | CDN |
| whisper-1 | mlx-community/whisper-large-v3-mlx | multilingual | OpenAI-spec id from the legacy Whisper API. Same destination as ``whisper-large-v3``. | CDN |
| whisper-base | mlx-community/whisper-base-mlx | multilingual | Whisper base. | CDN |
| whisper-large-v3 | mlx-community/whisper-large-v3-mlx | multilingual | Whisper large-v3 (multilingual, 99+ languages). | CDN |
| whisper-large-v3-turbo | mlx-community/whisper-large-v3-turbo | multilingual | Whisper large-v3 turbo (faster decode, slight quality drop). | CDN |
| whisper-medium | mlx-community/whisper-medium-mlx | multilingual | Whisper medium. | CDN |
| whisper-small | mlx-community/whisper-small-mlx | multilingual | Whisper small. | CDN |
| whisper-tiny | mlx-community/whisper-tiny-mlx | multilingual | Whisper tiny (smallest, fastest, lowest accuracy). | CDN |
Parakeet · 5 aliases
NVIDIA Parakeet TDT 0.6B — v2 is English-only; v3 supports 25 European languages with automatic language detection. Both checkpoints ship.
| alias | hf repo | languages | notes | get it |
|---|---|---|---|---|
| parakeet | mlx-community/parakeet-tdt-0.6b-v2 | en | NVIDIA Parakeet TDT 0.6B v2 (English, fastest STT). | CDN |
| parakeet-tdt-0.6b | mlx-community/parakeet-tdt-0.6b-v2 | en | Long-form alias for ``parakeet``. | CDN |
| parakeet-tdt-0.6b-v2 | mlx-community/parakeet-tdt-0.6b-v2 | en | Parakeet TDT 0.6B v2. | CDN |
| parakeet-tdt-0.6b-v3 | mlx-community/parakeet-tdt-0.6b-v3 | en,es,fr,de,bg,hr,cs,da,nl,et,fi,el,hu,it,lv,lt,mt,pl,pt,ro,sk,sl,sv,ru,uk | Parakeet TDT 0.6B v3 with automatic language detection. Chinese is not supported. | CDN |
| parakeet-v3 | mlx-community/parakeet-tdt-0.6b-v3 | en,es,fr,de,bg,hr,cs,da,nl,et,fi,el,hu,it,lv,lt,mt,pl,pt,ro,sk,sl,sv,ru,uk | Parakeet TDT 0.6B v3 with automatic language detection. Chinese is not supported. | CDN |
Qwen3-ASR · 3 aliases
Alibaba's Qwen3-ASR — an audio LLM rather than an encoder-decoder ASR model. Keeps Chinese/English code-switching straight without a decoding hint and takes hotwords as a system_prompt. Bare qwen3-asr resolves to the 1.7B; a 0.6B ships too.
| alias | hf repo | languages | notes | get it |
|---|---|---|---|---|
| qwen3-asr | mlx-community/Qwen3-ASR-1.7B-5bit | multilingual | Qwen3-ASR 1.7B (5-bit). An audio LLM rather than an encoder-decoder ASR model: keeps Chinese/English code-switching straight without a decoding hint, and emits normal punctuation plus spacing between scripts. Takes hotwords as ``system_prompt`` (STTEngine maps the neutral ``context`` argument onto it). Bare ``qwen3-asr`` -> the 1.7B. | CDN |
| qwen3-asr-0.6b | mlx-community/Qwen3-ASR-0.6B-5bit | multilingual | Smaller Qwen3-ASR. Faster than the 1.7B with a small accuracy cost. | CDN |
| qwen3-asr-1.7b | mlx-community/Qwen3-ASR-1.7B-5bit | multilingual | Explicit long-form alias for ``qwen3-asr`` — same HF id. | CDN |
SenseVoice · 2 aliases
FunAudioLLM SenseVoice Small (~234M, non-autoregressive CTC) — the strongest pick for Chinese, Cantonese, Japanese and Korean, and it emits per-segment emotion and audio-event tags alongside the transcript.
| alias | hf repo | languages | notes | get it |
|---|---|---|---|---|
| sensevoice | mlx-community/SenseVoiceSmall | zh,yue,ja,ko,en | FunAudioLLM SenseVoice Small (~234M, non-autoregressive CTC). Dedicated fast Asian-language ASR (strong on Chinese/Cantonese/Japanese/Korean). Also emits per-segment emotion + audio-event tags. Bare ``sensevoice`` -> the small model. | CDN |
| sensevoice-small | mlx-community/SenseVoiceSmall | zh,yue,ja,ko,en | Explicit long-form alias for ``sensevoice`` — same HF id. | CDN |
Qwen3-ForcedAligner · 2 aliases
Give it audio plus the transcript you already have and it returns per-character start/end times. It never guesses at words, so it cannot mis-hear them — exactly what karaoke captions and beat-synced editing need.
| alias | hf repo | languages | notes | get it |
|---|---|---|---|---|
| qwen3-aligner | mlx-community/Qwen3-ForcedAligner-0.6B-8bit | multilingual | Forced ALIGNMENT (not ASR): given audio + the KNOWN transcript, returns per-character start/end times with zero recognition error. Use STTEngine.align(audio, text, language) or /v1/audio/transcriptions with a `text` field. Powers karaoke captions + beat-synced editing; unoccupied on MLX. | CDN |
| qwen3-forced-aligner | mlx-community/Qwen3-ForcedAligner-0.6B-8bit | multilingual | Long-form alias for ``qwen3-aligner`` — same HF id. | CDN |
Notes & caveats
- TTS aliases default to a sensible per-family voice but accept any voice supported by the underlying model — pass
voicein the JSON body of/v1/audio/speech. - STT aliases accept both
multipart/form-dataaudio uploads and the OpenAI-stylefilefield on/v1/audio/transcriptions. whisper-1is wired as an OpenAI-compat placeholder — it resolves towhisper-large-v3so legacy OpenAI client code works unchanged.- Dia 1.6B is best-effort — covered by the audio probe but
mlx_audiosupport is experimental upstream; report bugs upstream if synth fails.
External reading
- Kokoro 82M on HF
- Whisper large-v3 (MLX) on HF
- Parakeet TDT 0.6B v2 (MLX) on HF
- Parakeet TDT 0.6B v3 upstream model card