Chat apps
Any chat app that speaks the OpenAI API can use a model you serve with Rapid-MLX. Start the server, then give the app its address.
$ rapid-mlx serve qwen3.5-9b-4bit # or any alias, org/repo or import name
Open WebUI
In Open WebUI, open Settings → Admin → Connections, find Manage OpenAI API Connections and click + Add Connection. Paste this URL:
http://localhost:8000/v1
- API key: any non-empty value, or your
--api-key - Model IDs: leave empty. Open WebUI reads them from
/v1/models.
If Open WebUI runs in Docker, the container's localhost is
not your Mac. Use http://host.docker.internal:8000/v1
instead, as Open WebUI's docs describe. If the container still can't
connect, start the server with --host 0.0.0.0, which also
exposes it to your network, so set --api-key too.
SillyTavern
In SillyTavern's API connections panel, choose Chat Completion →
Custom (OpenAI-compatible) or Text Completion, and fill in the
connection settings (base URL
http://localhost:8000/v1). SillyTavern sends its whole
sampler payload on every request; a stock preset works as-is, and the
sampler contract below says what happens to each
setting.
Connection settings
Every app asks for the same few things:
| Base URL | http://localhost:8000/v1. The server binds to loopback by default; see --host and --port in the CLI reference. |
| API key | Anything, unless the server was started with --api-key. The Rapid-MLX desktop app always sets one. |
| Model | The alias or import name you served. Apps that read /v1/models list it for you. |
| Routes | Chat: /v1/chat/completions. Text completion: /v1/completions. Both apply the samplers below. |
A model with no chat template (the pre-download
check prints Chat template none) is usually a base
model. Use the app's text-completion mode for it.
Samplers
Rapid-MLX applies the samplers it implements, accepts the rest only at
their "off" value, and refuses any you turn on, so a setting is never
silently ignored. This holds on both /v1/chat/completions
and /v1/completions, whichever app sends the request.
| Sampler | Notes |
|---|---|
temperature, top_p, top_k, min_p | Standard. |
repetition_penalty + repetition_penalty_range | The range is the penalty window in tokens. 0 means the whole context; the default window is 20 tokens. KoboldCpp's rep_pen_range spelling is accepted. |
presence_penalty, frequency_penalty | −2.0 to 2.0. |
DRY: dry_multiplier, dry_base, dry_allowed_length, dry_penalty_last_n, dry_sequence_breakers | dry_multiplier 0 means off. dry_sequence_breakers takes a list of strings, or the JSON-encoded list SillyTavern sends (up to 64 non-empty strings of up to 32 characters each). Some special serving lanes (certain speculative-decoding modes) refuse DRY with an error instead of ignoring it. |
Accepted only when off: typical_p, tfs,
top_a, mirostat, XTC
(xtc_probability), smoothing_factor, dynamic
temperature, epsilon_cutoff / eta_cutoff,
nsigma, top_n_sigma and a few more. Turn one
of these on and the request gets HTTP 400 naming the field:
sampler setting 'xtc_probability' is not supported by Rapid-MLX; set it to 0.0 or remove it. Supported samplers: temperature, top_p, top_k, min_p, repetition_penalty (+ repetition_penalty_range), presence_penalty, frequency_penalty, DRY (dry_multiplier, dry_base, dry_allowed_length, dry_penalty_last_n, dry_sequence_breakers).