Ollama, the tool for running large language models (LLMs) locally, is now compatible with Anthropic's Messages API[1]. Available in version 0.14.0 and later, it lets you use Anthropic's terminal coding tool, Claude Code, together with open-source models[1]. You can run models on your own machine, or connect to models through Ollama's cloud[1].
What "Anthropic API Compatibility" Changes
Claude Code is Anthropic's agentic coding tool that runs in the terminal[1]. It normally assumes Anthropic's own models, but now that Ollama is compatible with Anthropic's Messages API, any model running on Ollama can be called from Claude Code[1].
Using a local model means your code and instructions stay on your own machine instead of being sent to an external server. When you need a larger model, you can instead connect to a cloud model through ollama.com[1].
Setup: Installing Claude Code and Connecting Ollama
First, install Claude Code itself. The installation commands for each OS are as follows[1].
# macOS / Linux / WSL
curl -fsSL https://claude.ai/install.sh | bash
# Windows PowerShell
irm https://claude.ai/install.ps1 | iex
Next, point Claude Code at Ollama. You only need to set two environment variables[1].
export ANTHROPIC_AUTH_TOKEN=ollama
export ANTHROPIC_BASE_URL=http://localhost:11434
Then start Claude Code with the model you want to use[1].
claude --model gpt-oss:20b
Cloud models on Ollama can be specified the same way[1].
claude --model glm-4.7:cloud
No API key or additional configuration file is required; you can switch the destination just by setting environment variables.
Recommended Models and Context Length
The models Ollama recommends for coding are as follows[1]. For local models, "gpt-oss:20b" and "qwen3-coder" are suggested, while for cloud models, "glm-4.7:cloud" and "minimax-m2.1:cloud" are listed[1].
For work that handles long context, such as coding, a context length (the number of tokens a model can handle at once) of at least 32K (32,000 tokens) is recommended[1]. For local models, you change the context length in Ollama's settings[1]. Ollama's cloud models always run at their maximum context length[1].
It Works with the Anthropic SDK, Too
Beyond Claude Code, existing apps that use the Anthropic SDK (software development kit) can also be pointed at Ollama simply by changing the base URL[1]. In Python, set base_url to Ollama's address when initializing the client, and set api_key to "ollama." The key is required, but its value itself is not used[1].
import anthropic
client = anthropic.Anthropic(
base_url='http://localhost:11434',
api_key='ollama', # required, but the value is ignored
)
message = client.messages.create(
model='qwen3-coder',
messages=[
{'role': 'user', 'content': 'Write a function to check if a number is prime'}
]
)
print(message.content[0].text)
A wide range of features is supported: multi-turn conversations, streaming (incremental output), system prompts, tool calling (function calling), extended thinking, and vision for image input[1]. This widens the possibility of running many of the tools and apps built around Anthropic's models on open-source models instead.
Summary
With version 0.14.0 and later, Ollama is now compatible with Anthropic's Messages API, making it possible to run the terminal coding tool Claude Code and the Anthropic SDK on open-source models. Setup is simple — install Claude Code and point it at Ollama with environment variables — and you can choose either local or cloud models. A context length of 32K or more is recommended for coding, and features such as tool calling, extended thinking, and image input are supported. For anyone who wants to try a development environment built around Anthropic's models on their own hardware or on open models, this is a welcome step that broadens the options.
Source: https://ollama.com/blog/claude
