Use a Local Model from Python on Your Mac
Your Python code already knows how to talk to a model: the OpenAI SDK. Change the base URL and the same code runs against a model on your own Mac, streaming and tool calls included.
If your code uses the openai package, or a framework built on it such as LangChain, you can run it against a local model without rewriting anything. Rapid-MLX serves the OpenAI API on your Mac, so the only change is where the client points.
That's useful for development (no API bill while you iterate), for private data, and for tests that shouldn't need a network.
The whole thing
- Install Rapid-MLX and serve a model.
pip install openai.OpenAI(base_url="http://localhost:8000/v1", api_key="not-needed").
1 · Install Rapid-MLX
curl -fsSL https://rapidmlx.com/install.sh | bash
Prefer Homebrew? brew install rapid-mlx.
2 · Serve a model
rapid-mlx serve qwen3.5-4b-4bit
That's the engine's pick for a 16 GB Mac and the model we verified with. rapid-mlx recipe prints the pick for your memory, and Best local LLMs by Mac shows measured speeds per chip. Keep the server running.
3 · Call it with the OpenAI SDK
pip install openai
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="not-needed")
resp = client.chat.completions.create(
model="default",
messages=[{"role": "user", "content": "Name three Apple Silicon chips. One line."}],
)
print(resp.choices[0].message.content)
model="default" means whatever the server is serving, so the script doesn't change when you switch models. Our run printed:
Three examples of Apple Silicon chips are the M1, M2, and M3.
Streaming works the same way as against the cloud API: pass stream=True and read chunk.choices[0].delta.content.
4 · Call tools
Tool calling uses the standard tools parameter:
tools = [{
"type": "function",
"function": {
"name": "get_weather",
"description": "Current weather for a city",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"],
},
},
}]
resp = client.chat.completions.create(
model="default",
messages=[{"role": "user", "content": "What is the weather in Berlin?"}],
tools=tools,
)
call = resp.choices[0].message.tool_calls[0]
print(call.function.name, call.function.arguments)
Our run printed get_weather {"city": "Berlin"}.
5 · Use it from LangChain
pip install langchain-openai
from langchain_core.tools import tool
from langchain_openai import ChatOpenAI
llm = ChatOpenAI(base_url="http://localhost:8000/v1", api_key="not-needed", model="default")
print(llm.invoke("In one sentence, what is unified memory?").content)
@tool
def get_weather(city: str) -> str:
"""Current weather for a city."""
return f"Sunny in {city}"
msg = llm.bind_tools([get_weather]).invoke("What is the weather in Berlin?")
print(msg.tool_calls)
Our run returned a one-sentence answer and [{'name': 'get_weather', 'args': {'city': 'Berlin'}, ...}].
Verified on 2026-10-06 with rapid-mlx 0.15.6, openai 3.24.0, langchain-openai 1.6.7 and qwen3.5-4b-4bit on a Mac mini (M2 Pro, 32 GB): a chat completion, a streamed completion, a tool call through the OpenAI SDK, and invoke plus bind_tools through LangChain. Each script finished in about 3 seconds with the model loaded.
Frequently asked questions
Can I use the OpenAI Python SDK with a local model?
Yes. Set base_url="http://localhost:8000/v1" on the client and serve a model with Rapid-MLX. Chat completions, streaming and tool calls work as they do against OpenAI.
What do I pass as the API key for a local model?
Any string. A plain rapid-mlx serve doesn't check it. If you start the server with RAPID_MLX_API_KEY set, pass that value instead.
Does LangChain work with a local model on a Mac?
Yes. ChatOpenAI from langchain-openai takes the same base_url, and bind_tools produces tool calls from the local model.
Where to go next
- OpenAI SDK reference and LangChain reference: more patterns and options.
- API reference: every endpoint Rapid-MLX serves.
- Best local LLMs by Mac: measured speed for your chip and memory.
- Rapid-MLX vs mlx-lm: the same weights, measured task by task.
rapid-mlx serve <alias> serves any model on
localhost:8000/v1.
curl -fsSL https://rapidmlx.com/install.sh | bash
New models and speedups, in your inbox
A short note whenever Rapid-MLX gets faster or adds models worth running on your Mac. No spam — unsubscribe anytime.