Reference · rapid-mlx 0.15.7 · ← Back to README

Performance flags reference

Most users need no flags: rapid-mlx serve <alias> already turns on the speed-ups that alias is qualified for. This page shows what is on, how to check a specific alias, how to turn things off, and the opt-in techniques for advanced use.

What is on by default

TechniqueOn forTurn off
Prefix cacheEvery model. Repeated prompt prefixes (system prompt, tool schema, earlier turns) are reused instead of prefilled again.--disable-prefix-cache
MTP speculative decodingqwen3.5-9b-4bit, qwen3.6-27b-4bit, qwen3.6-35b-4bit, qwen3.8-27b-4bit (and the aliases that resolve to the same checkpoints) and glm5.3-flash-4bit.--no-spec-decode
Compressed KV cache (K8V4)9 verified Qwen3.5 / 3.6 aliases: Qwen3.5-9B and 27B (4 / 8-bit), Qwen3.5-35B 6-bit, Qwen3.6-35B-A3B (4 / 6 / 8-bit and DWQ).--kv-cache-turboquant none

PFlash long-prompt compression is lossy, so it is off for every model until you opt in with --pflash auto or --pflash always.

Check what applies to an alias

rapid-mlx info <alias> is the source of truth: it shows the parsers, whether MTP is on by default or opt-in, and whether DFlash, DDTree and SuffixDecoding apply — with the exact opt-in command when one exists. No download needed.

$ rapid-mlx info qwen3.5-27b-8bit
  Alias: qwen3.5-27b-8bit → mlx-community/Qwen3.5-27B-8bit

┌──────────────────────────────────────────────────────────────┐
│ Model: mlx-community/Qwen3.5-27B-8bit                        │
│ ──────────────────────────────────────────────────────────── │
│ Tool format      : hermes                                    │
│ Reasoning parser : qwen3                                     │
│ Architecture     : dense GatedDeltaNet (standard scheduler)  │
│ Spec decode      : ✗ try --speculative-config {"method":"dflash"} │
│ MTP path         : disabled                                  │
│ KV-share         : no                                        │
│ Throttle         : ✗ not needed                              │
│ Suffix tier      : n/a (no MTP/drafter — spec decode off)    │
└──────────────────────────────────────────────────────────────┘

┌──────────────────────────────────────────────────────────────┐
│ DFlash eligibility: ⚠ verified pair; runtime unavailable     │
│ ──────────────────────────────────────────────────────────── │
│ Declared support  : ✓ yes (supports_dflash=true)             │
│ Not MoE           : ✓ yes (dense)                            │
│ Target precision  : ✓ 8-bit or higher                        │
│ Drafter algorithm : dflash                                   │
│ Drafter declared  : ✓ z-lab/Qwen3.5-27B-DFlash               │
│ mlx-vlm 0.5.0+    : ✗ missing (need rapid-mlx[dflash])       │
└──────────────────────────────────────────────────────────────┘

  Experimental opt-in: rapid-mlx serve qwen3.5-27b-8bit --speculative-config '{"method":"dflash","model":"z-lab/Qwen3.5-27B-DFlash"}'

"Runtime unavailable" here means the [dflash] extra is not installed on this machine; the pair itself is verified.

One flag for speculative decoding

Every speculative-decoding method is configured through a single JSON flag, --speculative-config. Two flags override the alias's eligibility: --no-spec-decode turns all of it off (including the default MTP), and --force-spec-decode enables it on an alias whose profile says it does not apply — use that only when you have verified the profile is wrong.

MTP · Multi-Token Prediction

A small draft head proposes the next few tokens and the full model verifies every proposal in one pass, so the tokens you get are the model's own. Qwen3.5 / 3.6 / 3.8 aliases use an MTP sidecar checkpoint; GLM-5.3 Flash and Qwen3.8 Flash-Next use the head built into their checkpoint. On the default-on aliases above nothing needs to be set; on the other MTP-capable aliases it is opt-in:

$ rapid-mlx serve qwen3.6-35b-8bit --speculative-config '{"method":"mtp"}'
FlagDefaultPurpose
--speculative-config '{"method":"mtp",…}'alias-drivenOptional fields: num_speculative_tokens (draft tokens per step), disable_auto_k, continuous_batching and allow_dynamic_membership. Continuous self-MTP runs only with "continuous_batching": true; dynamic membership is off unless the alias is qualified for it.
--no-spec-decodeoffTurns off MTP and every other speculative method, including an alias's default MTP.

By default MTP runs on one request at a time. When another request arrives, an eligible running request yields at a token boundary and the requests decode together on the ordinary batched path; a request that is alone again goes back to MTP. Nothing needs to be set for this.

On the Qwen3.5 / 3.6 aliases, default-on MTP also proposes runs copied from the prompt (prompt lookup), which is what speeds up answers that repeat most of their input, such as a whole file returned with a small edit. The Rapid-MLX vs mlx-lm benchmark measures both against Apple's MLX server, task by task, with the raw data published.

Gemma 4 assistant sidecar (opt-in). Qualified Gemma 4 targets can be paired with an assistant drafter, e.g. gemma-4-12b-assistant, by naming it as the "model" of an MTP config. It is never on by default because the gain depends on the workload: on an M4 Pro eight-prompt matrix (rapid-mlx 0.15.0), fixed K=3 raised pooled decode from 57.88 to 67.12 tok/s (1.16×) at 64.64% draft acceptance.

Qwen3.8 Flash-Next (opt-in). Its native MTP head needs --speculative-config; 192 GB is the recommended tier. Measured on a 256 GB M3 Ultra (rapid-mlx 0.13.2, August 2026) at fixed K=1, with 76.41% of proposals accepted and up to 6.6 GB more active memory:

PromptSerial decodeNative MTPChange
12825.17 tok/s34.85 tok/s+38.5%
2K23.64 tok/s33.53 tok/s+41.8%
8K22.82 tok/s32.20 tok/s+41.1%
32K21.16 tok/s28.82 tok/s+36.2%
$ rapid-mlx serve qwen3.8-flash-next-4bit \
    --speculative-config '{"method":"mtp","num_speculative_tokens":1}'

DFlash · block-diffusion drafter

A separate block-diffusion drafter proposes a block of tokens per pass; the target verifies them. Single-user, serial mode. Needs the [dflash] extra and a verified target/drafter pair — qwen3.5-27b-8bit and muse-glimmer-30b-8bit are verified pairs; rapid-mlx info prints the exact command for each alias and marks other pairs as experimental.

$ pip install "rapid-mlx[dflash]"
$ rapid-mlx serve qwen3.5-27b-8bit \
    --speculative-config '{"method":"dflash","model":"z-lab/Qwen3.5-27B-DFlash"}'

The eligibility gates rapid-mlx info checks: declared support, a dense (non-MoE) target, 8-bit or higher target precision (4-bit pairs are experimental), a declared drafter, and the [dflash] runtime. Mixture-of-experts aliases are incompatible.

DSpark · DeepSeek V4's built-in drafter

DeepSeek V4 Flash checkpoints carry their drafter inside the checkpoint (MTP stages plus low-rank Markov heads), so there is nothing extra to download. Serve the downloaded checkpoint directory, not the alias: DSpark reads its block geometry from the checkpoint's inference/config.json, and num_speculative_tokens must equal its dspark_block_size (5 on the 0731 checkpoint).

$ rapid-mlx serve /path/to/DeepSeek-V4-Flash-0731-MXFP4-MLX \
    --speculative-config '{"method":"dspark","num_speculative_tokens":5}'

Detection fails closed: the geometry is read from the checkpoint and every required tensor is checked against model.safetensors.index.json, so a conversion without the drafter weights refuses to start instead of decoding without it. The experimental deepseek-v41-flash-reap-2bit alias needs no flag — it always decodes with its pinned DSpark K4 sidecar.

LFM2.5-VL companion. DSpark also has one qualified external drafter: serving the BF16 LiquidAI/LFM2.5-VL-3B with --speculative-config '{"method":"dspark"}' attaches LiquidAI/LFM2.5-VL-3B-DSpark with 7 proposals per step. It is serial and greedy-only, accepts text and image prompts, and rejects sampling, logprobs, tools and seeds with HTTP 400.

DDTree · draft-tree verification (experimental)

Verifies a tree of DFlash drafts per step. Single-user, serial mode; needs a DDTree-declared alias and the external dtree-mlx runtime, and fails at startup if either is missing.

$ rapid-mlx serve <alias> --speculative-config \
    '{"method":"ddtree","model":"<drafter>","num_speculative_tokens":16,"tree_budget":24}'

SuffixDecoding · drafter-free drafts

Model-free speculative decoding from suffix trees built over past outputs. Explicit opt-in, for long high-overlap workloads (copying from the prompt, code edits, repeated tool XML) on aliases whose suffix tier is not avoid. It does not run on hybrid models.

$ rapid-mlx serve gemma-4-12b-4bit \
    --speculative-config '{"method":"suffix","num_speculative_tokens":8}'

num_speculative_tokens is the maximum draft per verify step; the verify cost grows with it. rapid-mlx info and rapid-mlx models show each alias's suffix tier (verified, neutral, avoid, unknown, or n/a). "Unknown" means the alias has not been benchmarked for it yet.

PFlash · long-prompt prefill compression

For long prompts, PFlash scores the middle of the prompt and prefills only the start, the recent tail and the most query-relevant middle blocks. It is off by default for every model; opt in per server:

$ rapid-mlx serve qwen3.6-35b-4bit --pflash auto     # prompts of at least --pflash-threshold tokens
$ rapid-mlx serve qwen3.6-35b-4bit --pflash always   # every prompt long enough to trim

With always and the default sizes it compresses prompts longer than about 11.5K tokens (shorter ones have no middle left to trim); auto also waits for --pflash-threshold. The bench-validated profile on the verified Qwen3.5 / 3.6 aliases measured 3.87×–8.5× faster first tokens and 5/5 needle-in-a-haystack recall. Five clean trials only loosely bound the failure rate, and PFlash drops about 80% of the middle tokens — if exact recall from the middle of a long document matters, benchmark your workload first. Prompts that carry tool definitions are skipped unless you pass --pflash-include-tools.

A compressed request says so. Non-streaming responses set the X-Rapid-MLX-Prompt-Compressed: <kept>/<original> header, and the OpenAI-style responses (and the final chunk of a streamed chat completion) carry metrics.prompt_compression with original_tokens and kept_tokens. usage.prompt_tokens still counts the prompt you sent.

FlagDefaultPurpose
--pflash {off,auto,always}offMaster switch. always compresses any prompt long enough to have a middle to trim; auto only prompts above the threshold.
--pflash-threshold N32768Minimum prompt tokens before auto compresses.
--pflash-keep-ratio F0.20 (or the alias's own value)Fraction of prompt tokens to keep.
--pflash-min-keep-tokens N2048Minimum tokens kept.
--pflash-sink-tokens N256Leading prompt tokens always kept.
--pflash-tail-tokens N2048Trailing prompt tokens always kept.
--pflash-block-size N128Middle-token scoring block size.
--pflash-query-window N512Trailing window used to score middle blocks.
--pflash-stride-blocks N8Keep every Nth middle block as an anchor (0 disables anchors).
--pflash-include-toolsoffAlso compress prompts that carry tool definitions.

KV-cache compression

TurboQuant compresses the KV cache so long contexts and concurrent sessions fit in less memory. k8v4 keeps K at 8-bit (Walsh-Hadamard) and V at 4-bit (Lloyd-Max) — about 4.6× smaller on dense models — and is on by default for the 9 verified aliases listed above. v4 compresses V only (3–4 bit) and keeps K in FP16. Separately, --kv-cache-dtype int8 / int4 shrink the live KV cache 2× / 4× at a decode-speed cost on long contexts; sliding-window (Gemma 3 / 4, GPT-OSS) and MLA (DeepSeek V3+, Kimi K2.5) models reject an explicit int8 / int4 at startup.

FlagDefaultPurpose
--kv-cache-turboquant {v4,k8v4,none}alias-driven (k8v4 on 9 verified aliases)none turns it off; a bare flag means v4.
--kv-cache-turboquant-bits {3,4}auto by head_dimV-side bit width. Ignored with k8v4.
--kv-cache-turboquant-group-size N32Group size for V-side quantization.
--kv-cache-dtype {bf16,int8,int4}bf16Live KV-cache precision.
--reasoningoffPins the KV cache to int8 — use for hard math, where 4-bit KV loses accuracy.

Prompt and response caching

The prefix cache is always on and makes a repeated prompt cheap; how much it saves grows with the length of the shared prefix. Two more caches are off until you size them:

FlagDefaultPurpose
--response-cache-entries N0 (off)Keep up to N finished greedy (temperature 0 / top_k 1) chat completions; an identical repeat returns the stored answer with no decode. Sampled requests are never served from it.
--hybrid-cache-entries N0 (off)Keep up to N prefix-cache entries for hybrid (GatedDeltaNet / Mamba) and sliding-window (Gemma 4, GPT-OSS) models, so a stable prefix plus a new suffix each turn skips re-prefilling the prefix. Best for agent sessions with a fixed system prompt.
--pin-system-promptoffKeep the system prompt in the prefix cache under memory pressure.
--cache-memory-mb Nabout 20% of RAMMemory budget for the prefix cache.

See also