Chat app guide · rapid-mlx 0.15.5

Chat apps

Any chat app that speaks the OpenAI API can use a model you serve with Rapid-MLX. Start the server, then give the app its address.

$ rapid-mlx serve qwen3.5-9b-4bit     # or any alias, org/repo or import name

Open WebUI

In Open WebUI, open Settings → Admin → Connections, find Manage OpenAI API Connections and click + Add Connection. Paste this URL:

http://localhost:8000/v1

If Open WebUI runs in Docker, the container's localhost is not your Mac. Use http://host.docker.internal:8000/v1 instead, as Open WebUI's docs describe. If the container still can't connect, start the server with --host 0.0.0.0, which also exposes it to your network, so set --api-key too.

SillyTavern

In SillyTavern's API connections panel, choose Chat Completion → Custom (OpenAI-compatible) or Text Completion, and fill in the connection settings (base URL http://localhost:8000/v1). SillyTavern sends its whole sampler payload on every request; a stock preset works as-is, and the sampler contract below says what happens to each setting.

Connection settings

Every app asks for the same few things:

Base URLhttp://localhost:8000/v1. The server binds to loopback by default; see --host and --port in the CLI reference.
API keyAnything, unless the server was started with --api-key. The Rapid-MLX desktop app always sets one.
ModelThe alias or import name you served. Apps that read /v1/models list it for you.
RoutesChat: /v1/chat/completions. Text completion: /v1/completions. Both apply the samplers below.

A model with no chat template (the pre-download check prints Chat template none) is usually a base model. Use the app's text-completion mode for it.

Samplers

Rapid-MLX applies the samplers it implements, accepts the rest only at their "off" value, and refuses any you turn on, so a setting is never silently ignored. This holds on both /v1/chat/completions and /v1/completions, whichever app sends the request.

SamplerNotes
temperature, top_p, top_k, min_pStandard.
repetition_penalty + repetition_penalty_rangeThe range is the penalty window in tokens. 0 means the whole context; the default window is 20 tokens. KoboldCpp's rep_pen_range spelling is accepted.
presence_penalty, frequency_penalty−2.0 to 2.0.
DRY: dry_multiplier, dry_base, dry_allowed_length, dry_penalty_last_n, dry_sequence_breakersdry_multiplier 0 means off. dry_sequence_breakers takes a list of strings, or the JSON-encoded list SillyTavern sends (up to 64 non-empty strings of up to 32 characters each). Some special serving lanes (certain speculative-decoding modes) refuse DRY with an error instead of ignoring it.

Accepted only when off: typical_p, tfs, top_a, mirostat, XTC (xtc_probability), smoothing_factor, dynamic temperature, epsilon_cutoff / eta_cutoff, nsigma, top_n_sigma and a few more. Turn one of these on and the request gets HTTP 400 naming the field:

sampler setting 'xtc_probability' is not supported by Rapid-MLX; set it to 0.0 or
remove it. Supported samplers: temperature, top_p, top_k, min_p, repetition_penalty
(+ repetition_penalty_range), presence_penalty, frequency_penalty, DRY (dry_multiplier,
dry_base, dry_allowed_length, dry_penalty_last_n, dry_sequence_breakers).

Next steps