# Rapid-MLX 0.15.6 vs mlx-lm 0.32.0 — per-task benchmark (2026-10-04)

Data behind https://rapidmlx.com/compare/mlx-lm. Every number on that page,
the homepage headline and the /compare four-stream note is computed by
`landing/scripts/vs_mlx.py` from the files here into `summary.json`.

- `versions.json` — machine, engine versions, commands, request settings, protocol.
- `tasks.json` — the 20 tasks (built by `build_tasks.py`; code in prompts is unmodified CPython stdlib source).
- `bench_tasks.py` — the client (standard-library Python). `run_tasks.sh` + `lib.sh` — the driver
  (fresh server and fresh HOME per config per rep, rotating order). `all.sh` — what was run.
- `raw/tasks/<model>/<config>/rep<N>.json` — every request: timings, engine-reported usage, generated text.
  Configs: `mlxlm` (mlx_lm.server defaults), `rapid` (Rapid-MLX defaults), `rapid_mtp` (4B with
  `--speculative-config '{"method":"mtp"}'`), `rapid_nospec` (`--no-spec-decode`, 1 rep, attribution only).
  `progress.txt` logs each server start with the load average.
- `analyze.py` — the original analysis script (medians, speedups, identical-output check).
- `raw/post/generate2.txt` — `mlx_lm.generate` (no server) on two tasks per model (`gen_check.sh`).
- `raw/post/*.tiecheck.json`, `tiecheck.log` — for each task whose output differs, the reference model's
  top-5 next-token log-probs at the first differing position (`tie_check.py`, run by `post.sh`).
- `raw/extras-9b/` — four concurrent streams and the agent session with the /compare harness (`extras.sh`).

Absolute paths under the benchmark host's home directory are replaced with `~`.
