Performance flags reference
Most users need no flags: rapid-mlx serve <alias>
already turns on the speed-ups that alias is qualified for. This page
shows what is on, how to check a specific alias, how to turn things
off, and the opt-in techniques for advanced use.
What is on by default
| Technique | On for | Turn off |
|---|---|---|
| Prefix cache | Every model. Repeated prompt prefixes (system prompt, tool schema, earlier turns) are reused instead of prefilled again. | --disable-prefix-cache |
| MTP speculative decoding | qwen3.5-9b-4bit, qwen3.6-27b-4bit, qwen3.6-35b-4bit, qwen3.8-27b-4bit (and the aliases that resolve to the same checkpoints) and glm5.3-flash-4bit. | --no-spec-decode |
| Compressed KV cache (K8V4) | 9 verified Qwen3.5 / 3.6 aliases: Qwen3.5-9B and 27B (4 / 8-bit), Qwen3.5-35B 6-bit, Qwen3.6-35B-A3B (4 / 6 / 8-bit and DWQ). | --kv-cache-turboquant none |
PFlash long-prompt compression is lossy, so it is
off for every model until you opt in with --pflash auto
or --pflash always.
Check what applies to an alias
rapid-mlx info <alias> is the source of truth: it
shows the parsers, whether MTP is on by default or opt-in, and whether
DFlash, DDTree and SuffixDecoding apply — with the exact opt-in command
when one exists. No download needed.
$ rapid-mlx info qwen3.5-27b-8bit
Alias: qwen3.5-27b-8bit → mlx-community/Qwen3.5-27B-8bit
┌──────────────────────────────────────────────────────────────┐
│ Model: mlx-community/Qwen3.5-27B-8bit │
│ ──────────────────────────────────────────────────────────── │
│ Tool format : hermes │
│ Reasoning parser : qwen3 │
│ Architecture : dense GatedDeltaNet (standard scheduler) │
│ Spec decode : ✗ try --speculative-config {"method":"dflash"} │
│ MTP path : disabled │
│ KV-share : no │
│ Throttle : ✗ not needed │
│ Suffix tier : n/a (no MTP/drafter — spec decode off) │
└──────────────────────────────────────────────────────────────┘
┌──────────────────────────────────────────────────────────────┐
│ DFlash eligibility: ⚠ verified pair; runtime unavailable │
│ ──────────────────────────────────────────────────────────── │
│ Declared support : ✓ yes (supports_dflash=true) │
│ Not MoE : ✓ yes (dense) │
│ Target precision : ✓ 8-bit or higher │
│ Drafter algorithm : dflash │
│ Drafter declared : ✓ z-lab/Qwen3.5-27B-DFlash │
│ mlx-vlm 0.5.0+ : ✗ missing (need rapid-mlx[dflash]) │
└──────────────────────────────────────────────────────────────┘
Experimental opt-in: rapid-mlx serve qwen3.5-27b-8bit --speculative-config '{"method":"dflash","model":"z-lab/Qwen3.5-27B-DFlash"}'
"Runtime unavailable" here means the [dflash] extra is not
installed on this machine; the pair itself is verified.
One flag for speculative decoding
Every speculative-decoding method is configured through a single JSON
flag, --speculative-config. Two flags override the
alias's eligibility: --no-spec-decode turns all of it off
(including the default MTP), and --force-spec-decode
enables it on an alias whose profile says it does not apply — use that
only when you have verified the profile is wrong.
MTP · Multi-Token Prediction
A small draft head proposes the next few tokens and the full model verifies every proposal in one pass, so the tokens you get are the model's own. Qwen3.5 / 3.6 / 3.8 aliases use an MTP sidecar checkpoint; GLM-5.3 Flash and Qwen3.8 Flash-Next use the head built into their checkpoint. On the default-on aliases above nothing needs to be set; on the other MTP-capable aliases it is opt-in:
$ rapid-mlx serve qwen3.6-35b-8bit --speculative-config '{"method":"mtp"}'
| Flag | Default | Purpose |
|---|---|---|
--speculative-config '{"method":"mtp",…}' | alias-driven | Optional fields: num_speculative_tokens (draft tokens per step), disable_auto_k, continuous_batching and allow_dynamic_membership. Continuous self-MTP runs only with "continuous_batching": true; dynamic membership is off unless the alias is qualified for it. |
--no-spec-decode | off | Turns off MTP and every other speculative method, including an alias's default MTP. |
By default MTP runs on one request at a time. When another request arrives, an eligible running request yields at a token boundary and the requests decode together on the ordinary batched path; a request that is alone again goes back to MTP. Nothing needs to be set for this.
On the Qwen3.5 / 3.6 aliases, default-on MTP also proposes runs copied from the prompt (prompt lookup), which is what speeds up answers that repeat most of their input, such as a whole file returned with a small edit. The Rapid-MLX vs mlx-lm benchmark measures both against Apple's MLX server, task by task, with the raw data published.
Gemma 4 assistant sidecar (opt-in). Qualified Gemma 4 targets can
be paired with an assistant drafter, e.g.
gemma-4-12b-assistant, by naming it as the
"model" of an MTP config. It is never on by default because
the gain depends on the workload: on an M4 Pro eight-prompt matrix
(rapid-mlx 0.15.0), fixed K=3 raised pooled decode from 57.88 to
67.12 tok/s (1.16×) at 64.64% draft acceptance.
Qwen3.8 Flash-Next (opt-in). Its native MTP head needs
--speculative-config; 192 GB is the recommended tier.
Measured on a 256 GB M3 Ultra (rapid-mlx 0.13.2, August 2026) at
fixed K=1, with 76.41% of proposals accepted and up to 6.6 GB more
active memory:
| Prompt | Serial decode | Native MTP | Change |
|---|---|---|---|
| 128 | 25.17 tok/s | 34.85 tok/s | +38.5% |
| 2K | 23.64 tok/s | 33.53 tok/s | +41.8% |
| 8K | 22.82 tok/s | 32.20 tok/s | +41.1% |
| 32K | 21.16 tok/s | 28.82 tok/s | +36.2% |
$ rapid-mlx serve qwen3.8-flash-next-4bit \ --speculative-config '{"method":"mtp","num_speculative_tokens":1}'
DFlash · block-diffusion drafter
A separate block-diffusion drafter proposes a block of tokens per pass;
the target verifies them. Single-user, serial mode. Needs the
[dflash] extra and a verified target/drafter pair —
qwen3.5-27b-8bit and muse-glimmer-30b-8bit are
verified pairs; rapid-mlx info prints the exact command for
each alias and marks other pairs as experimental.
$ pip install "rapid-mlx[dflash]" $ rapid-mlx serve qwen3.5-27b-8bit \ --speculative-config '{"method":"dflash","model":"z-lab/Qwen3.5-27B-DFlash"}'
The eligibility gates rapid-mlx info checks: declared
support, a dense (non-MoE) target, 8-bit or higher target precision
(4-bit pairs are experimental), a declared drafter, and the
[dflash] runtime. Mixture-of-experts aliases are
incompatible.
DSpark · DeepSeek V4's built-in drafter
DeepSeek V4 Flash checkpoints carry their drafter inside the
checkpoint (MTP stages plus low-rank Markov heads), so there is
nothing extra to download. Serve the downloaded checkpoint
directory, not the alias: DSpark reads its block geometry from
the checkpoint's inference/config.json, and
num_speculative_tokens must equal its
dspark_block_size (5 on the 0731 checkpoint).
$ rapid-mlx serve /path/to/DeepSeek-V4-Flash-0731-MXFP4-MLX \ --speculative-config '{"method":"dspark","num_speculative_tokens":5}'
Detection fails closed: the geometry is read from the checkpoint and
every required tensor is checked against
model.safetensors.index.json, so a conversion without the
drafter weights refuses to start instead of decoding without it. The
experimental deepseek-v41-flash-reap-2bit alias needs no
flag — it always decodes with its pinned DSpark K4 sidecar.
LFM2.5-VL companion. DSpark also has one qualified external
drafter: serving the BF16 LiquidAI/LFM2.5-VL-3B with
--speculative-config '{"method":"dspark"}' attaches
LiquidAI/LFM2.5-VL-3B-DSpark with 7 proposals per step. It
is serial and greedy-only, accepts text and image prompts, and rejects
sampling, logprobs, tools and seeds with HTTP 400.
DDTree · draft-tree verification (experimental)
Verifies a tree of DFlash drafts per step. Single-user, serial mode;
needs a DDTree-declared alias and the external dtree-mlx
runtime, and fails at startup if either is missing.
$ rapid-mlx serve <alias> --speculative-config \ '{"method":"ddtree","model":"<drafter>","num_speculative_tokens":16,"tree_budget":24}'
SuffixDecoding · drafter-free drafts
Model-free speculative decoding from suffix trees built over past
outputs. Explicit opt-in, for long high-overlap workloads (copying from
the prompt, code edits, repeated tool XML) on aliases whose suffix tier
is not avoid. It does not run on hybrid models.
$ rapid-mlx serve gemma-4-12b-4bit \ --speculative-config '{"method":"suffix","num_speculative_tokens":8}'
num_speculative_tokens is the maximum draft per verify
step; the verify cost grows with it. rapid-mlx info and
rapid-mlx models show each alias's suffix tier
(verified, neutral, avoid,
unknown, or n/a). "Unknown" means the alias
has not been benchmarked for it yet.
PFlash · long-prompt prefill compression
For long prompts, PFlash scores the middle of the prompt and prefills only the start, the recent tail and the most query-relevant middle blocks. It is off by default for every model; opt in per server:
$ rapid-mlx serve qwen3.6-35b-4bit --pflash auto # prompts of at least --pflash-threshold tokens $ rapid-mlx serve qwen3.6-35b-4bit --pflash always # every prompt long enough to trim
With always and the default sizes it compresses prompts
longer than about 11.5K tokens (shorter ones have no middle left to
trim); auto also waits for --pflash-threshold.
The bench-validated profile on the verified Qwen3.5 / 3.6 aliases
measured 3.87×–8.5× faster first tokens and 5/5 needle-in-a-haystack
recall. Five clean trials only loosely bound the failure rate, and
PFlash drops about 80% of the middle tokens — if exact recall from the
middle of a long document matters, benchmark your workload first.
Prompts that carry tool definitions are skipped unless you pass
--pflash-include-tools.
A compressed request says so. Non-streaming responses set the
X-Rapid-MLX-Prompt-Compressed: <kept>/<original>
header, and the OpenAI-style responses (and the final chunk of a
streamed chat completion) carry metrics.prompt_compression
with original_tokens and kept_tokens.
usage.prompt_tokens still counts the prompt you sent.
| Flag | Default | Purpose |
|---|---|---|
--pflash {off,auto,always} | off | Master switch. always compresses any prompt long enough to have a middle to trim; auto only prompts above the threshold. |
--pflash-threshold N | 32768 | Minimum prompt tokens before auto compresses. |
--pflash-keep-ratio F | 0.20 (or the alias's own value) | Fraction of prompt tokens to keep. |
--pflash-min-keep-tokens N | 2048 | Minimum tokens kept. |
--pflash-sink-tokens N | 256 | Leading prompt tokens always kept. |
--pflash-tail-tokens N | 2048 | Trailing prompt tokens always kept. |
--pflash-block-size N | 128 | Middle-token scoring block size. |
--pflash-query-window N | 512 | Trailing window used to score middle blocks. |
--pflash-stride-blocks N | 8 | Keep every Nth middle block as an anchor (0 disables anchors). |
--pflash-include-tools | off | Also compress prompts that carry tool definitions. |
KV-cache compression
TurboQuant compresses the KV cache so long contexts and concurrent
sessions fit in less memory. k8v4 keeps K at 8-bit
(Walsh-Hadamard) and V at 4-bit (Lloyd-Max) — about 4.6× smaller on
dense models — and is on by default for the 9 verified aliases listed
above. v4 compresses V only (3–4 bit) and keeps K in FP16.
Separately, --kv-cache-dtype int8 / int4
shrink the live KV cache 2× / 4× at a decode-speed cost on long
contexts; sliding-window (Gemma 3 / 4, GPT-OSS) and MLA (DeepSeek V3+,
Kimi K2.5) models reject an explicit int8 /
int4 at startup.
| Flag | Default | Purpose |
|---|---|---|
--kv-cache-turboquant {v4,k8v4,none} | alias-driven (k8v4 on 9 verified aliases) | none turns it off; a bare flag means v4. |
--kv-cache-turboquant-bits {3,4} | auto by head_dim | V-side bit width. Ignored with k8v4. |
--kv-cache-turboquant-group-size N | 32 | Group size for V-side quantization. |
--kv-cache-dtype {bf16,int8,int4} | bf16 | Live KV-cache precision. |
--reasoning | off | Pins the KV cache to int8 — use for hard math, where 4-bit KV loses accuracy. |
Prompt and response caching
The prefix cache is always on and makes a repeated prompt cheap; how much it saves grows with the length of the shared prefix. Two more caches are off until you size them:
| Flag | Default | Purpose |
|---|---|---|
--response-cache-entries N | 0 (off) | Keep up to N finished greedy (temperature 0 / top_k 1) chat completions; an identical repeat returns the stored answer with no decode. Sampled requests are never served from it. |
--hybrid-cache-entries N | 0 (off) | Keep up to N prefix-cache entries for hybrid (GatedDeltaNet / Mamba) and sliding-window (Gemma 4, GPT-OSS) models, so a stable prefix plus a new suffix each turn skips re-prefilling the prefix. Best for agent sessions with a fixed system prompt. |
--pin-system-prompt | off | Keep the system prompt in the prefix cache under memory pressure. |
--cache-memory-mb N | about 20% of RAM | Memory budget for the prefix cache. |