Build and Push / build (release) Successful in 27s
- Add OLLAMA_EMBEDDING_MODEL for embeddings (nomic-embed-text) - OLLAMA_MODEL now used for all LLM operations (mistral-nemo-large:latest) - Remove separate reranker_model setting - Update WikiPageWriter to use settings instead of hardcoded model - Improves VRAM efficiency by keeping one model hot 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
155 lines
5.5 KiB
Markdown
155 lines
5.5 KiB
Markdown
# Changelog
|
|
|
|
All notable changes to Library Desk will be documented in this file.
|
|
|
|
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
|
|
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
|
|
|
|
## [1.3.2] - 2025-12-22
|
|
|
|
### Changed
|
|
|
|
- **Consolidated Ollama model configuration** - All LLM operations now use single `OLLAMA_MODEL` environment variable
|
|
- Removed separate `reranker_model` setting
|
|
- HybridRAG re-ranking, consolidation analysis, and wiki page writing all use the same model
|
|
- Improves VRAM efficiency by keeping one model hot
|
|
- Added `OLLAMA_EMBEDDING_MODEL` environment variable for embedding model (previously overloaded `OLLAMA_MODEL`)
|
|
- Updated WikiPageWriter to accept settings instead of hardcoded model name
|
|
|
|
## [1.3.1] - 2025-12-16
|
|
|
|
### Fixed
|
|
|
|
- Smart create endpoint missing `content_extractor` dependency causing 500 errors on `POST /wiki/pages/smart-create`
|
|
|
|
## [1.3.0] - 2025-12-15
|
|
|
|
### Changed
|
|
|
|
- **Two-Stage RRF Architecture** - Major refactor to level the playing field between wiki and web results
|
|
- Stage 1: Vector and graph results merged into single "wiki" ranking using mini-RRF
|
|
- Stage 2: Final RRF between wiki (single source) and web (single source)
|
|
- Wiki pages no longer get 2x advantage from appearing in both vector and graph searches
|
|
- Multi-source confirmation still determines wiki internal ranking
|
|
|
|
- **Skip synonyms in graph search** - LLM-generated synonyms (e.g., "author") no longer match unrelated graph entities (e.g., "author2000")
|
|
- Vector search still uses synonyms for semantic similarity
|
|
- Graph search uses only core keywords for exact entity matching
|
|
|
|
### Added
|
|
|
|
- `VECTOR_SIMILARITY_THRESHOLD` config setting (default: 0.7) to filter weak vector matches
|
|
- Deduplication in graph search to prevent same document appearing multiple times
|
|
|
|
### Fixed
|
|
|
|
- Graph search duplicate entity bug where same document could appear twice if entity linked multiple times
|
|
|
|
## [1.2.1] - 2025-12-15
|
|
|
|
### Fixed
|
|
|
|
- HybridRAG router missing `content_extractor` dependency causing 500 errors on `/query/hybrid` endpoint
|
|
|
|
## [1.2.0] - 2025-12-15
|
|
|
|
### Added
|
|
|
|
- **RAG Search Endpoint** (`POST /rag/search`)
|
|
- Web, news, and image search via SearXNG
|
|
- Full content extraction using Trafilatura (F1 score 0.958)
|
|
- Redis caching with configurable TTL
|
|
- Markdown sources summary for LLM consumption
|
|
- Returns both extracted content and original snippets
|
|
|
|
- **Content Extraction Endpoints** (`/content/*`)
|
|
- `POST /content/extract` - Extract content from a single URL
|
|
- `POST /content/extract/batch` - Batch extraction (up to 20 URLs)
|
|
- Reusable ContentExtractor client for use across the codebase
|
|
|
|
- **HybridRAG Content Extraction Enhancement**
|
|
- Web search results now include full extracted content via Trafilatura
|
|
- Falls back to original snippets if extraction fails
|
|
- Improves context quality for LLM re-ranking and consumption
|
|
|
|
### Changed
|
|
|
|
- Added new configuration options:
|
|
- `SEARCH_CACHE_TTL` - Search cache TTL in seconds (default: 300)
|
|
- `SEARCH_TIMEOUT` - SearXNG timeout (default: 10s)
|
|
- `CONTENT_EXTRACTION_TIMEOUT` - Per-URL extraction timeout (default: 5s)
|
|
- `CONTENT_MAX_LENGTH` - Max extracted content length (default: 2000)
|
|
- `SEARCH_DEFAULT_LIMIT` - Default search results (default: 10)
|
|
|
|
### Dependencies
|
|
|
|
- Added `trafilatura~=1.12.0` for content extraction
|
|
|
|
## [1.1.3] - 2025-12-14
|
|
|
|
### Added
|
|
|
|
- Watchtower update trigger in Gitea workflow after successful build
|
|
|
|
## [1.1.2] - 2025-12-14
|
|
|
|
### Fixed
|
|
|
|
- Updated registry login URL in Gitea workflow (git.schweitz.net → git.schweitz.internal)
|
|
|
|
## [1.1.1] - 2025-12-14
|
|
|
|
### Fixed
|
|
|
|
- Updated container registry tag URLs in Gitea workflow (git.schweitz.net → git.schweitz.internal)
|
|
|
|
### Added
|
|
|
|
- Tests for Smart Page Creation feature (`test_smart_create.py`)
|
|
- Model validation tests for WikiSmartCreateRequest/Response
|
|
- WikiService.smart_create_page method tests
|
|
- Bidirectional entity linking utility tests
|
|
- Endpoint validation tests
|
|
|
|
## [1.1.0] - 2025-12-11
|
|
|
|
### Added
|
|
|
|
- **Smart Page Creation Endpoint** (`POST /wiki/pages/smart-create`)
|
|
- Combines HybridRAG research with LLM content generation
|
|
- Searches existing wiki, knowledge graph, and web for topic context
|
|
- Uses WikiPageWriter to synthesize findings into structured wiki content
|
|
- Auto-generates page path from topic if not provided
|
|
- Returns research summary with source counts
|
|
|
|
- **Bidirectional Entity Linking**
|
|
- New shared utility (`entity_linking_utils.py`) for reusable entity linking
|
|
- Forward links: Links entities mentioned in new pages to existing entity pages
|
|
- Backward links: Updates existing pages that mention the new entity
|
|
- Runs automatically in background after smart page creation
|
|
|
|
- **Version Management**
|
|
- Added `pyproject.toml` with project metadata and version
|
|
- Version is now read from `pyproject.toml` (single source of truth)
|
|
- Health check endpoint returns current version
|
|
- FastAPI docs show current version
|
|
|
|
### Changed
|
|
|
|
- Updated `config.py` to read version from `pyproject.toml`
|
|
- Updated `main.py` to use centralized version
|
|
|
|
## [1.0.0] - 2025-12-10
|
|
|
|
### Added
|
|
|
|
- Initial release extracted from portainer-core
|
|
- Wiki page management (`/wiki/pages` CRUD endpoints)
|
|
- HybridRAG search (`/query/hybrid`) with vector, graph, and web search
|
|
- Knowledge graph operations (`/graph/*`)
|
|
- Vector search operations (`/vector/*`)
|
|
- Knowledge consolidation from search results (`/consolidate/knowledge`)
|
|
- Entity linking and extraction
|
|
- Wiki.js change listener for auto-processing user edits
|
|
- Multi-tenant architecture with user namespace isolation
|