refactor: consolidate Ollama model configuration
Build and Push / build (release) Successful in 27s

- Add OLLAMA_EMBEDDING_MODEL for embeddings (nomic-embed-text)
- OLLAMA_MODEL now used for all LLM operations (mistral-nemo-large:latest)
- Remove separate reranker_model setting
- Update WikiPageWriter to use settings instead of hardcoded model
- Improves VRAM efficiency by keeping one model hot

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
This commit is contained in:
2025-12-22 11:15:27 +01:00
co-authored by Claude Opus 4.5
parent 5be31a5a00
commit a1832e3245
9 changed files with 27 additions and 14 deletions
+1 -1
View File
@@ -113,7 +113,7 @@ def get_ollama_client() -> OllamaClient:
settings = get_settings()
client = OllamaClient(
base_url=settings.ollama_url,
model=settings.ollama_model
model=settings.ollama_embedding_model
)
logger.debug("Created Ollama client instance")
return client