Skip to content

Ollama provides compatibility with the Anthropic Messages API to help connect existing applications to Ollama, including tools like Claude Code.

Usage ​

Environment variables ​

To use Ollama with tools that expect the Anthropic API (like Claude Code), set these environment variables:

shell
export ANTHROPIC_AUTH_TOKEN=ollama  # required but ignored
export ANTHROPIC_BASE_URL=http://localhost:11434

Simple /v1/messages example ​

Streaming example ​

Tool calling example ​

Using with Claude Code ​

Claude Code can be configured to use Ollama as its backend.

For coding use cases, models like glm-4.7, minimax-m2.1, and qwen3-coder are recommended.

Download a model before use:

shell
ollama pull qwen3-coder

Note: Qwen 3 coder is a 30B parameter model requiring at least 24GB of VRAM to run smoothly. More is required for longer context lengths.

shell
ollama pull glm-4.7:cloud

Quick setup ​

shell
ollama launch claude

This will prompt you to select a model, configure Claude Code automatically, and launch it. To configure without launching:

shell
ollama launch claude --config

Manual setup ​

Set the environment variables and run Claude Code:

shell
ANTHROPIC_AUTH_TOKEN=ollama ANTHROPIC_BASE_URL=http://localhost:11434 claude --model qwen3-coder

Or set the environment variables in your shell profile:

shell
export ANTHROPIC_AUTH_TOKEN=ollama
export ANTHROPIC_BASE_URL=http://localhost:11434

Then run Claude Code with any Ollama model:

shell
claude --model qwen3-coder

Endpoints ​

/v1/messages ​

Supported features ​

  • [x] Messages
  • [x] Streaming
  • [x] System prompts
  • [x] Multi-turn conversations
  • [x] Vision (images)
  • [x] Tools (function calling)
  • [x] Tool results
  • [x] Thinking/extended thinking

Supported request fields ​

  • [x] model
  • [x] max_tokens
  • [x] messages
    • [x] Text content
    • [x] Image content (base64)
    • [x] Array of content blocks
    • [x] tool_use blocks
    • [x] tool_result blocks
    • [x] thinking blocks
  • [x] system (string or array)
  • [x] stream
  • [x] temperature
  • [x] top_p
  • [x] top_k
  • [x] stop_sequences
  • [x] tools
  • [x] thinking
  • [ ] tool_choice
  • [ ] metadata

Supported response fields ​

  • [x] id
  • [x] type
  • [x] role
  • [x] model
  • [x] content (text, tool_use, thinking blocks)
  • [x] stop_reason (end_turn, max_tokens, tool_use)
  • [x] usage (input_tokens, output_tokens)

Streaming events ​

  • [x] message_start
  • [x] content_block_start
  • [x] content_block_delta (text_delta, input_json_delta, thinking_delta)
  • [x] content_block_stop
  • [x] message_delta
  • [x] message_stop
  • [x] ping
  • [x] error

Models ​

Ollama supports both local and cloud models.

Local models ​

Pull a local model before use:

shell
ollama pull qwen3-coder

Recommended local models:

  • qwen3-coder - Excellent for coding tasks
  • gpt-oss:20b - Strong general-purpose model

Cloud models ​

Cloud models are available immediately without pulling:

  • glm-4.7:cloud - High-performance cloud model
  • minimax-m2.1:cloud - Fast cloud model

Default model names ​

For tooling that relies on default Anthropic model names such as claude-3-5-sonnet, use ollama cp to copy an existing model name:

shell
ollama cp qwen3-coder claude-3-5-sonnet

Afterwards, this new model name can be specified in the model field:

shell
curl http://localhost:11434/v1/messages \
    -H "Content-Type: application/json" \
    -d '{
        "model": "claude-3-5-sonnet",
        "max_tokens": 1024,
        "messages": [
            {
                "role": "user",
                "content": "Hello!"
            }
        ]
    }'

Differences from the Anthropic API ​

Behavior differences ​

  • API key is accepted but not validated
  • anthropic-version header is accepted but not used
  • Token counts are approximations based on the underlying model's tokenizer

Not supported ​

The following Anthropic API features are not currently supported:

FeatureDescription
/v1/messages/count_tokensToken counting endpoint
tool_choiceForcing specific tool use or disabling tools
metadataRequest metadata (user_id)
Prompt cachingcache_control blocks for caching prefixes
Batches API/v1/messages/batches for async batch processing
Citationscitations content blocks
PDF supportdocument content blocks with PDF files
Server-sent errorserror events during streaming (errors return HTTP status)

Partial support ​

FeatureStatus
Image contentBase64 images supported; URL images not supported
Extended thinkingBasic support; budget_tokens accepted but not enforced