diff --git a/.SYNC_INFO.md b/.SYNC_INFO.md index f949a35..b7a03ec 100644 --- a/.SYNC_INFO.md +++ b/.SYNC_INFO.md @@ -4,8 +4,8 @@ This is a mirror of the Ollama repository. **Synced from:** https://github.com/ollama/ollama.git **Branch:** main -**Commit:** 9f7822851c1f080d7d2a1dbe0e4d51233e5a28bc -**Sync Date:** 2025-12-12 +**Commit:** 361d6c16c23b60f84eb1d2134914610c0ae5836b +**Sync Date:** 2026-01-12 **Content:** Paths: docs --- diff --git a/docs/README.md b/docs/README.md index 74544a3..4483eb5 100644 --- a/docs/README.md +++ b/docs/README.md @@ -14,6 +14,7 @@ * [API Reference](https://docs.ollama.com/api) * [Modelfile Reference](https://docs.ollama.com/modelfile) * [OpenAI Compatibility](https://docs.ollama.com/api/openai-compatibility) +* [Anthropic Compatibility](./api/anthropic-compatibility.mdx) ### Resources diff --git a/docs/api.md b/docs/api.md index 03f6dbe..7c32c95 100644 --- a/docs/api.md +++ b/docs/api.md @@ -895,11 +895,11 @@ curl http://localhost:11434/api/chat -d '{ "tool_calls": [ { "function": { - "name": "get_temperature", + "name": "get_weather", "arguments": { "city": "Toronto" } - }, + } } ] }, @@ -907,7 +907,7 @@ curl http://localhost:11434/api/chat -d '{ { "role": "tool", "content": "11 degrees celsius", - "tool_name": "get_temperature", + "tool_name": "get_weather" } ], "stream": false, diff --git a/docs/api/anthropic-compatibility.mdx b/docs/api/anthropic-compatibility.mdx new file mode 100644 index 0000000..a0f2cd7 --- /dev/null +++ b/docs/api/anthropic-compatibility.mdx @@ -0,0 +1,406 @@ +--- +title: Anthropic compatibility +--- + +Ollama provides compatibility with the [Anthropic Messages API](https://docs.anthropic.com/en/api/messages) to help connect existing applications to Ollama, including tools like Claude Code. + +## Recommended models + +For coding use cases, models like `glm-4.7:cloud`, `minimax-m2.1:cloud`, and `qwen3-coder` are recommended. + +Pull a model before use: +```shell +ollama pull qwen3-coder +ollama pull glm-4.7:cloud +``` + +## Usage + +### Environment variables + +To use Ollama with tools that expect the Anthropic API (like Claude Code), set these environment variables: + +```shell +export ANTHROPIC_BASE_URL=http://localhost:11434 +export ANTHROPIC_API_KEY=ollama # required but ignored +``` + +### Simple `/v1/messages` example + + + +```python basic.py +import anthropic + +client = anthropic.Anthropic( + base_url='http://localhost:11434', + api_key='ollama', # required but ignored +) + +message = client.messages.create( + model='qwen3-coder', + max_tokens=1024, + messages=[ + {'role': 'user', 'content': 'Hello, how are you?'} + ] +) +print(message.content[0].text) +``` + +```javascript basic.js +import Anthropic from "@anthropic-ai/sdk"; + +const anthropic = new Anthropic({ + baseURL: "http://localhost:11434", + apiKey: "ollama", // required but ignored +}); + +const message = await anthropic.messages.create({ + model: "qwen3-coder", + max_tokens: 1024, + messages: [{ role: "user", content: "Hello, how are you?" }], +}); + +console.log(message.content[0].text); +``` + +```shell basic.sh +curl -X POST http://localhost:11434/v1/messages \ +-H "Content-Type: application/json" \ +-H "x-api-key: ollama" \ +-H "anthropic-version: 2023-06-01" \ +-d '{ + "model": "qwen3-coder", + "max_tokens": 1024, + "messages": [{ "role": "user", "content": "Hello, how are you?" }] +}' +``` + + + +### Streaming example + + + +```python streaming.py +import anthropic + +client = anthropic.Anthropic( + base_url='http://localhost:11434', + api_key='ollama', +) + +with client.messages.stream( + model='qwen3-coder', + max_tokens=1024, + messages=[{'role': 'user', 'content': 'Count from 1 to 10'}] +) as stream: + for text in stream.text_stream: + print(text, end='', flush=True) +``` + +```javascript streaming.js +import Anthropic from "@anthropic-ai/sdk"; + +const anthropic = new Anthropic({ + baseURL: "http://localhost:11434", + apiKey: "ollama", +}); + +const stream = await anthropic.messages.stream({ + model: "qwen3-coder", + max_tokens: 1024, + messages: [{ role: "user", content: "Count from 1 to 10" }], +}); + +for await (const event of stream) { + if ( + event.type === "content_block_delta" && + event.delta.type === "text_delta" + ) { + process.stdout.write(event.delta.text); + } +} +``` + +```shell streaming.sh +curl -X POST http://localhost:11434/v1/messages \ +-H "Content-Type: application/json" \ +-d '{ + "model": "qwen3-coder", + "max_tokens": 1024, + "stream": true, + "messages": [{ "role": "user", "content": "Count from 1 to 10" }] +}' +``` + + + +### Tool calling example + + + +```python tools.py +import anthropic + +client = anthropic.Anthropic( + base_url='http://localhost:11434', + api_key='ollama', +) + +message = client.messages.create( + model='qwen3-coder', + max_tokens=1024, + tools=[ + { + 'name': 'get_weather', + 'description': 'Get the current weather in a location', + 'input_schema': { + 'type': 'object', + 'properties': { + 'location': { + 'type': 'string', + 'description': 'The city and state, e.g. San Francisco, CA' + } + }, + 'required': ['location'] + } + } + ], + messages=[{'role': 'user', 'content': "What's the weather in San Francisco?"}] +) + +for block in message.content: + if block.type == 'tool_use': + print(f'Tool: {block.name}') + print(f'Input: {block.input}') +``` + +```javascript tools.js +import Anthropic from "@anthropic-ai/sdk"; + +const anthropic = new Anthropic({ + baseURL: "http://localhost:11434", + apiKey: "ollama", +}); + +const message = await anthropic.messages.create({ + model: "qwen3-coder", + max_tokens: 1024, + tools: [ + { + name: "get_weather", + description: "Get the current weather in a location", + input_schema: { + type: "object", + properties: { + location: { + type: "string", + description: "The city and state, e.g. San Francisco, CA", + }, + }, + required: ["location"], + }, + }, + ], + messages: [{ role: "user", content: "What's the weather in San Francisco?" }], +}); + +for (const block of message.content) { + if (block.type === "tool_use") { + console.log("Tool:", block.name); + console.log("Input:", block.input); + } +} +``` + +```shell tools.sh +curl -X POST http://localhost:11434/v1/messages \ +-H "Content-Type: application/json" \ +-d '{ + "model": "qwen3-coder", + "max_tokens": 1024, + "tools": [ + { + "name": "get_weather", + "description": "Get the current weather in a location", + "input_schema": { + "type": "object", + "properties": { + "location": { + "type": "string", + "description": "The city and state" + } + }, + "required": ["location"] + } + } + ], + "messages": [{ "role": "user", "content": "What is the weather in San Francisco?" }] +}' +``` + + + +## Using with Claude Code + +[Claude Code](https://code.claude.com/docs/en/overview) can be configured to use Ollama as its backend: + +```shell +ANTHROPIC_BASE_URL=http://localhost:11434 ANTHROPIC_API_KEY=ollama claude --model qwen3-coder +``` + +Or set the environment variables in your shell profile: + +```shell +export ANTHROPIC_BASE_URL=http://localhost:11434 +export ANTHROPIC_API_KEY=ollama +``` + +Then run Claude Code with any Ollama model: + +```shell +# Local models +claude --model qwen3-coder +claude --model gpt-oss:20b + +# Cloud models +claude --model glm-4.7:cloud +claude --model minimax-m2.1:cloud +``` + +## Endpoints + +### `/v1/messages` + +#### Supported features + +- [x] Messages +- [x] Streaming +- [x] System prompts +- [x] Multi-turn conversations +- [x] Vision (images) +- [x] Tools (function calling) +- [x] Tool results +- [x] Thinking/extended thinking + +#### Supported request fields + +- [x] `model` +- [x] `max_tokens` +- [x] `messages` + - [x] Text `content` + - [x] Image `content` (base64) + - [x] Array of content blocks + - [x] `tool_use` blocks + - [x] `tool_result` blocks + - [x] `thinking` blocks +- [x] `system` (string or array) +- [x] `stream` +- [x] `temperature` +- [x] `top_p` +- [x] `top_k` +- [x] `stop_sequences` +- [x] `tools` +- [x] `thinking` +- [ ] `tool_choice` +- [ ] `metadata` + +#### Supported response fields + +- [x] `id` +- [x] `type` +- [x] `role` +- [x] `model` +- [x] `content` (text, tool_use, thinking blocks) +- [x] `stop_reason` (end_turn, max_tokens, tool_use) +- [x] `usage` (input_tokens, output_tokens) + +#### Streaming events + +- [x] `message_start` +- [x] `content_block_start` +- [x] `content_block_delta` (text_delta, input_json_delta, thinking_delta) +- [x] `content_block_stop` +- [x] `message_delta` +- [x] `message_stop` +- [x] `ping` +- [x] `error` + +## Models + +Ollama supports both local and cloud models. + +### Local models + +Pull a local model before use: + +```shell +ollama pull qwen3-coder +``` + +Recommended local models: +- `qwen3-coder` - Excellent for coding tasks +- `gpt-oss:20b` - Strong general-purpose model + +### Cloud models + +Cloud models are available immediately without pulling: + +- `glm-4.7:cloud` - High-performance cloud model +- `minimax-m2.1:cloud` - Fast cloud model + +### Default model names + +For tooling that relies on default Anthropic model names such as `claude-3-5-sonnet`, use `ollama cp` to copy an existing model name: + +```shell +ollama cp qwen3-coder claude-3-5-sonnet +``` + +Afterwards, this new model name can be specified in the `model` field: + +```shell +curl http://localhost:11434/v1/messages \ + -H "Content-Type: application/json" \ + -d '{ + "model": "claude-3-5-sonnet", + "max_tokens": 1024, + "messages": [ + { + "role": "user", + "content": "Hello!" + } + ] + }' +``` + +## Differences from the Anthropic API + +### Behavior differences + +- API key is accepted but not validated +- `anthropic-version` header is accepted but not used +- Token counts are approximations based on the underlying model's tokenizer + +### Not supported + +The following Anthropic API features are not currently supported: + +| Feature | Description | +|---------|-------------| +| `/v1/messages/count_tokens` | Token counting endpoint | +| `tool_choice` | Forcing specific tool use or disabling tools | +| `metadata` | Request metadata (user_id) | +| Prompt caching | `cache_control` blocks for caching prefixes | +| Batches API | `/v1/messages/batches` for async batch processing | +| Citations | `citations` content blocks | +| PDF support | `document` content blocks with PDF files | +| Server-sent errors | `error` events during streaming (errors return HTTP status) | + +### Partial support + +| Feature | Status | +|---------|--------| +| Image content | Base64 images supported; URL images not supported | +| Extended thinking | Basic support; `budget_tokens` accepted but not enforced | diff --git a/docs/api/openai-compatibility.mdx b/docs/api/openai-compatibility.mdx index 94febc3..a088205 100644 --- a/docs/api/openai-compatibility.mdx +++ b/docs/api/openai-compatibility.mdx @@ -277,6 +277,8 @@ curl -X POST http://localhost:11434/v1/chat/completions \ ### `/v1/responses` +> Note: Added in Ollama v0.13.3 + Ollama supports the [OpenAI Responses API](https://platform.openai.com/docs/api-reference/responses). Only the non-stateful flavor is supported (i.e., there is no `previous_response_id` or `conversation` support). #### Supported features diff --git a/docs/capabilities/vision.mdx b/docs/capabilities/vision.mdx index 3342eae..81fa260 100644 --- a/docs/capabilities/vision.mdx +++ b/docs/capabilities/vision.mdx @@ -36,7 +36,6 @@ Provide an `images` array. SDKs accept file paths, URLs or raw bytes while the R }], "stream": false }' - " ``` diff --git a/docs/docs.json b/docs/docs.json index 71a6f17..810e947 100644 --- a/docs/docs.json +++ b/docs/docs.json @@ -32,7 +32,9 @@ "codeblocks": "system" }, "contextual": { - "options": ["copy"] + "options": [ + "copy" + ] }, "navbar": { "links": [ @@ -52,7 +54,9 @@ "display": "simple" }, "examples": { - "languages": ["curl"] + "languages": [ + "curl" + ] } }, "redirects": [ @@ -97,6 +101,7 @@ { "group": "Integrations", "pages": [ + "/integrations/claude-code", "/integrations/vscode", "/integrations/jetbrains", "/integrations/codex", @@ -139,7 +144,8 @@ "/api/streaming", "/api/usage", "/api/errors", - "/api/openai-compatibility" + "/api/openai-compatibility", + "/api/anthropic-compatibility" ] }, { diff --git a/docs/faq.mdx b/docs/faq.mdx index 13ef8a2..4237da4 100644 --- a/docs/faq.mdx +++ b/docs/faq.mdx @@ -14,11 +14,11 @@ curl -fsSL https://ollama.com/install.sh | sh ## How can I view the logs? -Review the [Troubleshooting](./troubleshooting.md) docs for more about using logs. +Review the [Troubleshooting](./troubleshooting) docs for more about using logs. ## Is my GPU compatible with Ollama? -Please refer to the [GPU docs](./gpu.md). +Please refer to the [GPU docs](./gpu). ## How can I specify the context window size? diff --git a/docs/gpu.mdx b/docs/gpu.mdx index 36bfd3d..9cb2d3a 100644 --- a/docs/gpu.mdx +++ b/docs/gpu.mdx @@ -33,7 +33,7 @@ Check your compute compatibility to see if your card is supported: | 5.0 | GeForce GTX | `GTX 750 Ti` `GTX 750` `NVS 810` | | | Quadro | `K2200` `K1200` `K620` `M1200` `M520` `M5000M` `M4000M` `M3000M` `M2000M` `M1000M` `K620M` `M600M` `M500M` | -For building locally to support older GPUs, see [developer.md](./development.md#linux-cuda-nvidia) +For building locally to support older GPUs, see [developer](./development#linux-cuda-nvidia) ### GPU Selection @@ -54,7 +54,7 @@ sudo modprobe nvidia_uvm` Ollama supports the following AMD GPUs via the ROCm library: -> [!NOTE] +> **NOTE:** > Additional AMD GPU support is provided by the Vulkan Library - see below. @@ -132,9 +132,9 @@ Ollama supports GPU acceleration on Apple devices via the Metal API. ## Vulkan GPU Support -> [!NOTE] +> **NOTE:** > Vulkan is currently an Experimental feature. To enable, you must set OLLAMA_VULKAN=1 for the Ollama server as -described in the [FAQ](faq.md#how-do-i-configure-ollama-server) +described in the [FAQ](faq#how-do-i-configure-ollama-server) Additional GPU support on Windows and Linux is provided via [Vulkan](https://www.vulkan.org/). On Windows most GPU vendors drivers come @@ -161,6 +161,6 @@ sudo setcap cap_perfmon+ep /usr/local/bin/ollama To select specific Vulkan GPU(s), you can set the environment variable `GGML_VK_VISIBLE_DEVICES` to one or more numeric IDs on the Ollama server as -described in the [FAQ](faq.md#how-do-i-configure-ollama-server). If you +described in the [FAQ](faq#how-do-i-configure-ollama-server). If you encounter any problems with Vulkan based GPUs, you can disable all Vulkan GPUs by setting `GGML_VK_VISIBLE_DEVICES=-1` \ No newline at end of file diff --git a/docs/integrations/claude-code.mdx b/docs/integrations/claude-code.mdx new file mode 100644 index 0000000..6d1d832 --- /dev/null +++ b/docs/integrations/claude-code.mdx @@ -0,0 +1,69 @@ +--- +title: Claude Code +--- + +## Install + +Install [Claude Code](https://code.claude.com/docs/en/overview): + + + +```shell macOS / Linux +curl -fsSL https://claude.ai/install.sh | bash +``` + +```powershell Windows +irm https://claude.ai/install.ps1 | iex +``` + + + +## Usage with Ollama + +Claude Code connects to Ollama using the Anthropic-compatible API. + +1. Set the environment variables: + +```shell +export ANTHROPIC_BASE_URL=http://localhost:11434 +export ANTHROPIC_API_KEY=ollama +``` + +2. Run Claude Code with an Ollama model: + +```shell +claude --model qwen3-coder +``` + +Or run with environment variables inline: + +```shell +ANTHROPIC_BASE_URL=http://localhost:11434 ANTHROPIC_API_KEY=ollama claude --model qwen3-coder +``` + +## Connecting to ollama.com + +1. Create an [API key](https://ollama.com/settings/keys) on ollama.com +2. Set the environment variables: + +```shell +export ANTHROPIC_BASE_URL=https://ollama.com +export ANTHROPIC_API_KEY= +``` + +3. Run Claude Code with a cloud model: + +```shell +claude --model glm-4.7:cloud +``` + +## Recommended Models + +### Cloud models +- `glm-4.7:cloud` - High-performance cloud model +- `minimax-m2.1:cloud` - Fast cloud model +- `qwen3-coder:480b` - Large coding model + +### Local models +- `qwen3-coder` - Excellent for coding tasks +- `gpt-oss:20b` - Strong general-purpose model diff --git a/docs/linux.mdx b/docs/linux.mdx index c40ab05..cf3ccc6 100644 --- a/docs/linux.mdx +++ b/docs/linux.mdx @@ -1,5 +1,5 @@ --- -title: Linux +title: "Linux" --- ## Install @@ -13,8 +13,7 @@ curl -fsSL https://ollama.com/install.sh | sh ## Manual install - If you are upgrading from a prior version, you should remove the old libraries - with `sudo rm -rf /usr/lib/ollama` first. + If you are upgrading from a prior version, you should remove the old libraries with `sudo rm -rf /usr/lib/ollama` first. Download and extract the package: @@ -113,11 +112,7 @@ sudo systemctl status ollama ``` - While AMD has contributed the `amdgpu` driver upstream to the official linux - kernel source, the version is older and may not support all ROCm features. We - recommend you install the latest driver from - https://www.amd.com/en/support/linux-drivers for best support of your Radeon - GPU. + While AMD has contributed the `amdgpu` driver upstream to the official linux kernel source, the version is older and may not support all ROCm features. We recommend you install the latest driver from https://www.amd.com/en/support/linux-drivers for best support of your Radeon GPU. ## Customizing @@ -196,4 +191,4 @@ Remove the downloaded models and Ollama service user and group: sudo userdel ollama sudo groupdel ollama sudo rm -r /usr/share/ollama -``` +``` \ No newline at end of file diff --git a/docs/modelfile.mdx b/docs/modelfile.mdx index a3eca20..ce91bbf 100644 --- a/docs/modelfile.mdx +++ b/docs/modelfile.mdx @@ -41,6 +41,7 @@ INSTRUCTION arguments | [`ADAPTER`](#adapter) | Defines the (Q)LoRA adapters to apply to the model. | | [`LICENSE`](#license) | Specifies the legal license. | | [`MESSAGE`](#message) | Specify message history. | +| [`REQUIRES`](#requires) | Specify the minimum version of Ollama required by the model. | ## Examples @@ -248,6 +249,16 @@ MESSAGE user Is Ontario in Canada? MESSAGE assistant yes ``` +### REQUIRES + +The `REQUIRES` instruction allows you to specify the minimum version of Ollama required by the model. + +``` +REQUIRES +``` + +The version should be a valid Ollama version (e.g. 0.14.0). + ## Notes - the **`Modelfile` is not case sensitive**. In the examples, uppercase instructions are used to make it easier to distinguish it from arguments. diff --git a/docs/troubleshooting.md b/docs/troubleshooting.md deleted file mode 100644 index c141bf4..0000000 --- a/docs/troubleshooting.md +++ /dev/null @@ -1,3 +0,0 @@ -# Troubleshooting - -For troubleshooting, see [https://docs.ollama.com/troubleshooting](https://docs.ollama.com/troubleshooting) diff --git a/docs/troubleshooting.mdx b/docs/troubleshooting.mdx index ec66257..9083dd8 100644 --- a/docs/troubleshooting.mdx +++ b/docs/troubleshooting.mdx @@ -87,7 +87,7 @@ When Ollama starts up, it takes inventory of the GPUs present in the system to d ### Linux NVIDIA Troubleshooting -If you are using a container to run Ollama, make sure you've set up the container runtime first as described in [docker.md](./docker.md) +If you are using a container to run Ollama, make sure you've set up the container runtime first as described in [docker](./docker) Sometimes the Ollama can have difficulties initializing the GPU. When you check the server logs, this can show up as various error codes, such as "3" (not initialized), "46" (device unavailable), "100" (no device), "999" (unknown), or others. The following troubleshooting techniques may help resolve the problem