Sync ollama docs from 82f905cd on 2026-07-12
This commit is contained in:
+1
-1
@@ -343,7 +343,7 @@ When loading a new model, Ollama evaluates the required VRAM for the model again
|
||||
|
||||
## How can I enable Flash Attention?
|
||||
|
||||
Flash Attention is a feature of most modern models that can significantly reduce memory usage as the context size grows. To enable Flash Attention, set the `OLLAMA_FLASH_ATTENTION` environment variable to `1` when starting the Ollama server.
|
||||
Flash Attention is a feature of most modern models that can significantly reduce memory usage as the context size grows. Ollama uses Flash Attention automatically when the selected backend and devices support it. To force Flash Attention on, set `OLLAMA_FLASH_ATTENTION=1` when starting the Ollama server. To disable it, set `OLLAMA_FLASH_ATTENTION=0`.
|
||||
|
||||
## How can I set the quantization type for the K/V cache?
|
||||
|
||||
|
||||
Reference in New Issue
Block a user