Notes from the engine room

Benchmarks, internals,
and practical Mac AI.

Field notes from the people building Rapid-MLX: measured comparisons, model guides, and repeatable local workflows.

All articles

28 stories · newest first

python Use a Local Model from Python on Your MacCall a local LLM from Python on Apple Silicon with the OpenAI SDK or LangChain: chat, streaming and tool calls against Rapid-MLX, verified end to end. cline Run Cline with a Local Model on Your MacUse Cline with a model running on your Mac via Rapid-MLX: one cline auth command, or the OpenAI Compatible provider in VS Code. Verified on Apple Silicon. continue Run Continue with a Local Model on Your MacUse Continue's chat, edit and agent with a model running on your Mac via Rapid-MLX. One config.yaml entry; verified with the Continue CLI on Apple Silicon. goose Run Goose with a Local Model on Your MacPoint Block's Goose agent at a model running on your Mac with Rapid-MLX. Goose's OpenAI provider plus a local host; verified with tool calls end to end. opencode Run OpenCode with a Local Model on Your MacPoint the OpenCode terminal agent at a model running on your Mac with Rapid-MLX. One setup command writes the provider config; model requests stay local. pi Pi Coding Agent with a Local Model on MacRun the pi coding agent on a model served from your Mac with Rapid-MLX. One setup command writes pi's models.json; verified end to end on Apple Silicon. deepseek DeepSeek V4.1 on a Mac, and Why Your LLM API Router Is Reading Your PromptsWe shipped a native MLX runtime for DeepSeek V4.1 Flash on day zero and published the 2-bit weights. Then a study of 428 LLM API routers explained why running it locally is the point. local-dictation How to Dictate into Any Mac App with a Fully Local Speech ModelUse Rapid-MLX Desktop for private, system-wide Mac dictation with Whisper, Qwen3-ASR, SenseVoice, or Parakeet. Audio stays on your Mac. deepseek Run DeepSeek Locally on Your MacBook ProHow to run DeepSeek locally on a Mac with MLX — open reasoning (R1) and coding models from a 4 GB 8B on a MacBook Air to the V4-Flash flagship on a Studio, one command to serve them, and how to wire them into Claude Code or Cursor. gemma Run Gemma 4 Locally on Your MacBook ProHow to run Google's Gemma locally on a Mac with MLX — which size fits your RAM (from a 1 GB nano on an 8 GB Air to the 31B dense on a Studio), one command to serve it, and how to wire it into Claude Code or Cursor. gpt-oss Run GPT-OSS Locally on Your MacBook ProHow to run OpenAI's GPT-OSS locally on a Mac with MLX — the 20B fits a 16 GB MacBook, the 120B runs on a Studio, both mixture-of-experts with native reasoning and tool calling. One command to serve it, and how to wire it into Claude Code or Cursor. llama Run Llama Locally on Your MacBook ProHow to run Meta's Llama locally on a Mac with MLX — which size fits your RAM (from a 0.5 GB 1B on a base Air to the 8B on 16 GB), one command to serve it, and how to wire it into Claude Code or Cursor. mistral Run Mistral Locally on Your MacBook ProHow to run Mistral locally on a Mac with MLX — Mistral Small 24B on a 24 GB MacBook, Devstral for coding, and the Mistral Small 4 119B flagship on a Studio. One command to serve it, and how to wire it into Claude Code or Cursor. mlx MLX vs llama.cpp on Apple Silicon: Measured PerformanceMLX vs llama.cpp on the same Mac and models: MLX wins decode (1.5× on MoE), llama.cpp wins prefill (1.7–2×). Which is faster depends on your workload. benchmarks MLX vs Ollama vs LM Studio: Rapid-MLX Speed BenchmarkMLX vs Ollama, llama.cpp and LM Studio on one M2 Pro Mac mini, raw data published: Rapid-MLX wins MoE decode by 1.5×, llama.cpp wins cold prefill. qwen Run Qwen 3.5 & Qwen 3.6 Locally on Your MacBook ProHow to run Qwen 3.5 and Qwen 3.6 locally on a Mac with MLX — which size fits your RAM (from 8 GB Airs to 128 GB Studios), one command to serve it, and how to wire it into Claude Code or Cursor. apple-silicon Best Local LLM for an 8, 16, 24 or 32 GB Mac, Measured8 GB: lfm2.5-2.6b. 16 GB: qwen3.5-4b. 24 GB: bonsai-27b-2bit. 32 GB and up: qwen3.8-27b. The engine's pick for each Mac's memory, measured on real Macs. apple-silicon How Much RAM Do You Need to Run a Local LLM on a Mac?8 GB fits 23 of 82 models, Qwen3.5 4B included. 24 GB fits 60, up to Qwen3.6 27B. The most capable one needs 256 GB. Estimated tiers, per model. coding-agents Local Coding Agents That Actually Work on Apple SiliconWe re-verified Claude Code, Codex CLI, Aider, and Hermes end-to-end on local models on Apple Silicon (July 2026) — real edits, real tool calls, current binaries. What works, why, and the one-line setup for each. ollama Ollama vs LM Studio vs Rapid-MLX (MLX): Which to UseOllama, LM Studio or an MLX server like Rapid-MLX? Where each fits on a Mac (platforms, API, GUI, license) and a 5-minute benchmark to run yourself. release Rapid-MLX 0.11.0 — 25.6× faster to first token, meet your Mac's new agentFifteen releases in three weeks. Prefix-cache reuse takes TTFT from 13.1s to 0.51s, rapid-mlx chat becomes a real MCP agent, five new model families land day one, and tool calls are now schema-valid by construction — all measured on real hardware. ollama Moving from Ollama to Rapid-MLX on Apple SiliconA practical migration guide. The CLI commands map almost 1:1 (run, pull, rm, ps), you point your OpenAI-compatible tools at a new port, and you can run both servers side by side while you switch. Includes a model-name cheat sheet. aider Run Aider with a Local Model on Your MacPair-program with Aider on a model running on your Mac via Rapid-MLX. Three environment variables, no API bill, code stays local. Verified on a Mac. claude-code Run Claude Code with a Local Model on a MacPoint Claude Code at a model on your Mac with Rapid-MLX: one setup command or two environment variables. Model requests stay local, no per-token bill. codex How to Run Codex CLI with Local Models on a MacRun Codex CLI with a local model on your Mac: serve it with Rapid-MLX, run rapid-mlx agents codex --setup (base URL http://localhost:8000/v1), then codex. cursor Cursor with a Local LLM: Use a Local Model on Your MacCursor can use a local model for chat and composer, but not at localhost. Serve it with an API key, open an HTTPS tunnel, run rapid-mlx launch cursor. benchmarks Best Local LLMs for a MacBook Pro (2026)32 GB+: Gemma-4-26B, best average at about 75 tok/s on an M4 Pro. Coding: Qwen3.6-27B. Speed: GPT-OSS-20B. Five LLMs scored and timed on M1 to M5 chips. benchmarks Best MLX Models for Coding: 17 Benchmarked on M3 UltraWhich MLX model to run for coding: Qwen3-Coder-Next for agents, Qwen3.5-35B-A3B on a smaller Mac. 17 models benchmarked for code, tools, speed and RAM.
28 articles
Per page