Use a Local Model from Python on Your Mac

Your Python code already knows how to talk to a model: the OpenAI SDK. Change the base URL and the same code runs against a model on your own Mac, streaming and tool calls included.

If your code uses the openai package, or a framework built on it such as LangChain, you can run it against a local model without rewriting anything. Rapid-MLX serves the OpenAI API on your Mac, so the only change is where the client points.

That's useful for development (no API bill while you iterate), for private data, and for tests that shouldn't need a network.

The whole thing

  1. Install Rapid-MLX and serve a model.
  2. pip install openai.
  3. OpenAI(base_url="http://localhost:8000/v1", api_key="not-needed").

1 · Install Rapid-MLX

curl -fsSL https://rapidmlx.com/install.sh | bash

Prefer Homebrew? brew install rapid-mlx.

2 · Serve a model

rapid-mlx serve qwen3.5-4b-4bit

That's the engine's pick for a 16 GB Mac and the model we verified with. rapid-mlx recipe prints the pick for your memory, and Best local LLMs by Mac shows measured speeds per chip. Keep the server running.

3 · Call it with the OpenAI SDK

pip install openai
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="not-needed")

resp = client.chat.completions.create(
    model="default",
    messages=[{"role": "user", "content": "Name three Apple Silicon chips. One line."}],
)
print(resp.choices[0].message.content)

model="default" means whatever the server is serving, so the script doesn't change when you switch models. Our run printed:

Three examples of Apple Silicon chips are the M1, M2, and M3.

Streaming works the same way as against the cloud API: pass stream=True and read chunk.choices[0].delta.content.

4 · Call tools

Tool calling uses the standard tools parameter:

tools = [{
    "type": "function",
    "function": {
        "name": "get_weather",
        "description": "Current weather for a city",
        "parameters": {
            "type": "object",
            "properties": {"city": {"type": "string"}},
            "required": ["city"],
        },
    },
}]
resp = client.chat.completions.create(
    model="default",
    messages=[{"role": "user", "content": "What is the weather in Berlin?"}],
    tools=tools,
)
call = resp.choices[0].message.tool_calls[0]
print(call.function.name, call.function.arguments)

Our run printed get_weather {"city": "Berlin"}.

5 · Use it from LangChain

pip install langchain-openai
from langchain_core.tools import tool
from langchain_openai import ChatOpenAI

llm = ChatOpenAI(base_url="http://localhost:8000/v1", api_key="not-needed", model="default")
print(llm.invoke("In one sentence, what is unified memory?").content)

@tool
def get_weather(city: str) -> str:
    """Current weather for a city."""
    return f"Sunny in {city}"

msg = llm.bind_tools([get_weather]).invoke("What is the weather in Berlin?")
print(msg.tool_calls)

Our run returned a one-sentence answer and [{'name': 'get_weather', 'args': {'city': 'Berlin'}, ...}].

Verified on 2026-10-06 with rapid-mlx 0.15.6, openai 3.24.0, langchain-openai 1.6.7 and qwen3.5-4b-4bit on a Mac mini (M2 Pro, 32 GB): a chat completion, a streamed completion, a tool call through the OpenAI SDK, and invoke plus bind_tools through LangChain. Each script finished in about 3 seconds with the model loaded.

Frequently asked questions

Can I use the OpenAI Python SDK with a local model?

Yes. Set base_url="http://localhost:8000/v1" on the client and serve a model with Rapid-MLX. Chat completions, streaming and tool calls work as they do against OpenAI.

What do I pass as the API key for a local model?

Any string. A plain rapid-mlx serve doesn't check it. If you start the server with RAPID_MLX_API_KEY set, pass that value instead.

Does LangChain work with a local model on a Mac?

Yes. ChatOpenAI from langchain-openai takes the same base_url, and bind_tools produces tool calls from the local model.

Where to go next


Run this yourself. Rapid-MLX is an open-source, OpenAI- and Anthropic-compatible inference server for Apple Silicon. One command installs it, then rapid-mlx serve <alias> serves any model on localhost:8000/v1.
curl -fsSL https://rapidmlx.com/install.sh | bash

New models and speedups, in your inbox

A short note whenever Rapid-MLX gets faster or adds models worth running on your Mac. No spam — unsubscribe anytime.