Build and Push / build (release) Successful in 1m10s
🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
383 lines
15 KiB
Markdown
383 lines
15 KiB
Markdown
# Changelog
|
|
|
|
All notable changes to Library Desk will be documented in this file.
|
|
|
|
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
|
|
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
|
|
|
|
## [1.5.0] - 2025-12-26
|
|
|
|
### Added
|
|
|
|
- **Central Settings Database** - Tatlock-wide configuration via PostgreSQL
|
|
- `SettingsClient` for async access to `system_settings` database
|
|
- User-scoped settings with global fallback
|
|
- API config storage with `enabled` toggle and per-source category filters
|
|
- JSON Schema support for future UI rendering
|
|
|
|
- **External API Providers** - Modular `src/apis/` package with swappable implementations
|
|
- `OpenMeteoProvider` - Weather with geocoding (free, no API key)
|
|
- `NOSProvider` - Dutch news RSS (16 categories including sports)
|
|
- `BBCProvider` - English news RSS (21 categories including sports)
|
|
- `AggregatedNewsProvider` - Merges sources chronologically with category filtering
|
|
- `AlphaVantageProvider` - Stock/crypto quotes (API key from settings DB)
|
|
- Abstract base classes for provider interoperability
|
|
|
|
- **Provider Dependency Injection**
|
|
- `WeatherProviderDep`, `NewsProviderDep`, `AlphaVantageProviderDep` type aliases
|
|
- Async initialization with settings database integration
|
|
- Lifecycle management in `shutdown_clients()`
|
|
|
|
- **Development Dependencies** - `requirements-dev.txt`
|
|
- `pip-audit` for security vulnerability scanning
|
|
- `ruff` for code quality
|
|
- Testing packages moved from main requirements
|
|
|
|
### Changed
|
|
|
|
- News sources configurable via `news.sources` setting
|
|
- Per-source category filtering via `api.{source}.categories`
|
|
- Categories default to all if not specified
|
|
|
|
## [1.4.8] - 2025-12-25
|
|
|
|
### Added
|
|
|
|
- **Paperless Orphan Cleanup** - `POST /maintenance/cleanup/paperless` endpoint
|
|
- Detects documents deleted from Paperless but still indexed in Library Desk
|
|
- Removes orphaned vectors and graph nodes
|
|
- Supports `dry_run=true` for preview mode
|
|
|
|
## [1.4.7] - 2025-12-25
|
|
|
|
### Fixed
|
|
|
|
- **Paperless Custom Field Update** - Fixed 400 error when marking documents as indexed
|
|
- Paperless API requires field ID (integer) not field name (string)
|
|
- Now looks up `library_indexed` field ID before updating
|
|
- Webhook params format: `doc_url` and `title` from Jinja templates
|
|
|
|
### Added
|
|
|
|
- **Webhook Debug Endpoint** - `POST /documents/webhook-capture` for development testing
|
|
|
|
## [1.4.6] - 2025-12-25
|
|
|
|
### Fixed
|
|
|
|
- **Paperless Webhook Payload Format** - Updated model to match Paperless `include_document=true` format
|
|
- Paperless sends `id` instead of `document_id`
|
|
- Paperless sends full document data including `content`, `title`, `tags`, etc.
|
|
- Webhook now uses content from payload, skipping extra Paperless API call
|
|
- Added `extra = "ignore"` to handle additional Paperless fields
|
|
|
|
## [1.4.5] - 2025-12-25
|
|
|
|
### Added
|
|
|
|
- **Document Storage Integration** - Paperless-ngx integration for PDFs, images, and documents
|
|
- Event-driven architecture via Paperless webhooks
|
|
- `POST /documents/webhook` - Receive document events from Paperless workflows
|
|
- `POST /documents/upload` - Upload files directly to Paperless
|
|
- `POST /documents/upload-url` - Download and upload documents from URL
|
|
- `POST /documents/search` - Semantic search across indexed documents
|
|
- `GET /documents/health` - Paperless connectivity health check
|
|
- **DocumentSyncService** - Indexes Paperless documents into vectors and graph
|
|
- Fetches document content via Paperless API
|
|
- Chunks text and generates embeddings for Qdrant
|
|
- Creates Document nodes in Neo4j knowledge graph
|
|
- Supports multi-tenancy via user parameter in webhook URL
|
|
- **PaperlessClient** - REST API client for Paperless-ngx
|
|
- Document retrieval, upload, and update operations
|
|
- Health check support
|
|
- **Paperless Workflow Configuration**
|
|
- Production workflow: Document Added (NOT tagged llm-test) → webhook to Library Desk
|
|
- Test workflow: Document Added (tagged llm-test) → webhook with test user
|
|
|
|
### Changed
|
|
|
|
- Updated `src/config.py` with Paperless configuration settings
|
|
- Added `PaperlessDep` dependency injection for document endpoints
|
|
|
|
## [1.4.4] - 2025-12-24
|
|
|
|
### Added
|
|
|
|
- **Test Data Cleanup Endpoint** - `POST /maintenance/cleanup/test-data`
|
|
- Purges LLM test data from wiki, graph, and vectors
|
|
- Security-restricted to test user namespace only (`users/llm-tester/*`, `users/llm_tester/*`)
|
|
- Supports `dry_run=true` (default) to preview before deleting
|
|
- Scheduler task configured for weekly cleanup (Sunday 3:00 AM)
|
|
|
|
## [1.4.3] - 2025-12-24
|
|
|
|
### Changed
|
|
|
|
- **Volatile Cache System Refactored to Vector Storage**
|
|
- Backend migrated from Redis to Qdrant for semantic search capability
|
|
- Data converted to natural language for embedding and semantic retrieval
|
|
- Collection naming: `volatile_{user}` for per-user isolation
|
|
- TTL implemented via `ttl_expiry` timestamp in vector payload
|
|
- Simplified endpoints:
|
|
- `GET /volatile/search?q=...` - Semantic search across volatile data
|
|
- `POST /volatile/store?namespace=...&key=...` - Store with query params
|
|
- `GET /volatile/{namespace}/{key}` - Get specific record
|
|
- `DELETE /volatile/{namespace}/{key}` - Delete record
|
|
- Removed namespace-specific URL patterns (simpler API for LLM tool use)
|
|
|
|
### Added
|
|
|
|
- **HybridRAG Volatile Integration** - Volatile cache now included in multi-source search
|
|
- Volatile results get priority boost in RRF fusion (current data ranks higher)
|
|
- New config options: `enable_volatile`, `volatile_limit` (default 1), `volatile_threshold`
|
|
- Timing breakdown includes `volatile_ms`
|
|
- **Volatile Cleanup Endpoint** - `POST /maintenance/cleanup/volatile`
|
|
- Purges expired records across all `volatile_*` collections
|
|
- Scheduler task for every 10 minutes recommended
|
|
- Returns per-collection cleanup counts
|
|
- **Natural Language Conversion** - Structured data converted for embedding
|
|
- Template-based conversion for each namespace (weather, news, financial, etc.)
|
|
- Fallback for custom namespaces
|
|
|
|
## [1.4.2] - 2025-12-24
|
|
|
|
### Added
|
|
|
|
- **Volatile Cache System** - Ephemeral data storage with TTL
|
|
- `GET /volatile/{namespace}/{key}` - Retrieve cached record
|
|
- `POST /volatile/{namespace}/{key}` - Store/update record with TTL
|
|
- `DELETE /volatile/{namespace}/{key}` - Remove record
|
|
- `GET /volatile/{namespace}` - List keys in namespace
|
|
- `DELETE /volatile/{namespace}` - Clear all records in namespace
|
|
- `GET /volatile/stats` - Cache statistics by namespace
|
|
- `GET /volatile/scheduled` - Records needing refresh (for scheduler)
|
|
- `GET /volatile/namespaces` - List available namespaces with default TTLs
|
|
- **Volatile Namespaces** - Predefined categories with appropriate TTLs:
|
|
- `weather` (30min) - Weather conditions and forecasts
|
|
- `news` (1hr) - Headlines and breaking news
|
|
- `financial` (5min) - Stock prices, exchange rates
|
|
- `transit` (5min) - Train/bus schedules, delays
|
|
- `traffic` (10min) - Commute times, road conditions
|
|
- `air_quality` (1hr) - Pollution, pollen counts
|
|
- `sports` (1min) - Live scores, matches
|
|
- `social` (10min) - Social notifications
|
|
- `system` (1min) - Service health status
|
|
- `context` (1hr) - Session state
|
|
- `custom` (1hr) - User-defined data
|
|
- **Refresh Schedule Support** - Optional cron expressions for scheduler integration
|
|
|
|
## [1.4.1] - 2025-12-24
|
|
|
|
### Fixed
|
|
|
|
- Wiki.js API token now optional - GraphQL API works without authentication
|
|
- Container startup failure when `WIKI_GRAPHQL_API` env var not set
|
|
|
|
## [1.4.0] - 2025-12-24
|
|
|
|
### Added
|
|
|
|
- **Maintenance Router** - New `/maintenance` endpoints for system health and cleanup
|
|
- `GET /maintenance/health` - Lightweight health check (detailed mode available)
|
|
- `POST /maintenance/cleanup/all` - Full orphan cleanup (vectors + graph)
|
|
- `POST /maintenance/cleanup/vectors` - Purge orphan vector chunks
|
|
- `POST /maintenance/cleanup/graph` - Purge orphan graph nodes
|
|
- `POST /maintenance/reconcile-index` - Combined cleanup + reindex missing pages
|
|
- **Bidirectional Orphan Detection** - Cross-validate vectors and graph nodes
|
|
- `find_documents_without_vectors()` - Graph nodes missing vector chunks
|
|
- `find_chunks_without_graph_nodes()` - Vector chunks missing graph nodes
|
|
- **Qdrant Client Methods** - Bulk operations for maintenance
|
|
- `scroll_all_points()` - Iterate all points with pagination
|
|
- `delete_by_ids()` - Batch delete by point IDs
|
|
- **Graph Service Cleanup** - Node deletion methods
|
|
- `delete_document_node()` - Remove document and relationships
|
|
- `delete_collection_node()` - Remove collection and contained documents
|
|
- `get_all_document_references()` - Get all document references for validation
|
|
- **Redis Timestamp Tracking** - `last_cleanup` timestamp for scheduler integration
|
|
- **Memory System Plan** - Documented three-tier architecture (volatile/documents/knowledge)
|
|
|
|
### Changed
|
|
|
|
- **Wiki.js Authentication** - Switched from username/password to API token
|
|
- New `WIKI_GRAPHQL_API` environment variable for JWT token
|
|
- Deprecated `WIKIJS_USERNAME` and `WIKIJS_PASSWORD` (kept for backwards compatibility)
|
|
- **Service Dependencies** - Added `VectorServiceDep` and `GraphServiceDep` type aliases
|
|
|
|
### Fixed
|
|
|
|
- Wiki.js client now properly handles API token auth without login flow
|
|
|
|
## [1.3.3] - 2025-12-23
|
|
|
|
### Added
|
|
|
|
- Temperature parameter to `OllamaClient.generate_text()` for controlling output determinism
|
|
- `TODO.md` tracking remaining stub endpoints to implement
|
|
- Wired `/query/semantic` endpoint to VectorService
|
|
- Wired `/query/graph` endpoint to GraphService
|
|
|
|
### Changed
|
|
|
|
- **Improved LLM prompts** based on llm-findings.md recommendations:
|
|
- Keyword extraction: temperature 0.0, negative constraints
|
|
- LLM re-ranking: temperature 0.0, explicit rules
|
|
- Conflict detection: temperature 0.0, analysis steps (CoT)
|
|
- Wiki page creation: temperature 0.3, anti-hallucination constraints
|
|
- Page reconstruction: temperature 0.2, preservation constraints
|
|
- Web results analysis: temperature 0.0, conservative approach
|
|
- Test fixtures now use configurable host (TEST_HOST) instead of Docker hostnames
|
|
|
|
### Removed
|
|
|
|
- Dead code: unused `get_default_user()` function
|
|
- Unused imports from routers (wiki.py, graph.py, hybrid_rag.py)
|
|
- Stub endpoints shadowed by real implementations (/stats, /ingest/document, /ingest/batch)
|
|
|
|
## [1.3.2] - 2025-12-22
|
|
|
|
### Changed
|
|
|
|
- **Consolidated Ollama model configuration** - All LLM operations now use single `OLLAMA_MODEL` environment variable
|
|
- Removed separate `reranker_model` setting
|
|
- HybridRAG re-ranking, consolidation analysis, and wiki page writing all use the same model
|
|
- Improves VRAM efficiency by keeping one model hot
|
|
- Added `OLLAMA_EMBEDDING_MODEL` environment variable for embedding model (previously overloaded `OLLAMA_MODEL`)
|
|
- Updated WikiPageWriter to accept settings instead of hardcoded model name
|
|
|
|
## [1.3.1] - 2025-12-16
|
|
|
|
### Fixed
|
|
|
|
- Smart create endpoint missing `content_extractor` dependency causing 500 errors on `POST /wiki/pages/smart-create`
|
|
|
|
## [1.3.0] - 2025-12-15
|
|
|
|
### Changed
|
|
|
|
- **Two-Stage RRF Architecture** - Major refactor to level the playing field between wiki and web results
|
|
- Stage 1: Vector and graph results merged into single "wiki" ranking using mini-RRF
|
|
- Stage 2: Final RRF between wiki (single source) and web (single source)
|
|
- Wiki pages no longer get 2x advantage from appearing in both vector and graph searches
|
|
- Multi-source confirmation still determines wiki internal ranking
|
|
|
|
- **Skip synonyms in graph search** - LLM-generated synonyms (e.g., "author") no longer match unrelated graph entities (e.g., "author2000")
|
|
- Vector search still uses synonyms for semantic similarity
|
|
- Graph search uses only core keywords for exact entity matching
|
|
|
|
### Added
|
|
|
|
- `VECTOR_SIMILARITY_THRESHOLD` config setting (default: 0.7) to filter weak vector matches
|
|
- Deduplication in graph search to prevent same document appearing multiple times
|
|
|
|
### Fixed
|
|
|
|
- Graph search duplicate entity bug where same document could appear twice if entity linked multiple times
|
|
|
|
## [1.2.1] - 2025-12-15
|
|
|
|
### Fixed
|
|
|
|
- HybridRAG router missing `content_extractor` dependency causing 500 errors on `/query/hybrid` endpoint
|
|
|
|
## [1.2.0] - 2025-12-15
|
|
|
|
### Added
|
|
|
|
- **RAG Search Endpoint** (`POST /rag/search`)
|
|
- Web, news, and image search via SearXNG
|
|
- Full content extraction using Trafilatura (F1 score 0.958)
|
|
- Redis caching with configurable TTL
|
|
- Markdown sources summary for LLM consumption
|
|
- Returns both extracted content and original snippets
|
|
|
|
- **Content Extraction Endpoints** (`/content/*`)
|
|
- `POST /content/extract` - Extract content from a single URL
|
|
- `POST /content/extract/batch` - Batch extraction (up to 20 URLs)
|
|
- Reusable ContentExtractor client for use across the codebase
|
|
|
|
- **HybridRAG Content Extraction Enhancement**
|
|
- Web search results now include full extracted content via Trafilatura
|
|
- Falls back to original snippets if extraction fails
|
|
- Improves context quality for LLM re-ranking and consumption
|
|
|
|
### Changed
|
|
|
|
- Added new configuration options:
|
|
- `SEARCH_CACHE_TTL` - Search cache TTL in seconds (default: 300)
|
|
- `SEARCH_TIMEOUT` - SearXNG timeout (default: 10s)
|
|
- `CONTENT_EXTRACTION_TIMEOUT` - Per-URL extraction timeout (default: 5s)
|
|
- `CONTENT_MAX_LENGTH` - Max extracted content length (default: 2000)
|
|
- `SEARCH_DEFAULT_LIMIT` - Default search results (default: 10)
|
|
|
|
### Dependencies
|
|
|
|
- Added `trafilatura~=1.12.0` for content extraction
|
|
|
|
## [1.1.3] - 2025-12-14
|
|
|
|
### Added
|
|
|
|
- Watchtower update trigger in Gitea workflow after successful build
|
|
|
|
## [1.1.2] - 2025-12-14
|
|
|
|
### Fixed
|
|
|
|
- Updated registry login URL in Gitea workflow (git.schweitz.net → git.schweitz.internal)
|
|
|
|
## [1.1.1] - 2025-12-14
|
|
|
|
### Fixed
|
|
|
|
- Updated container registry tag URLs in Gitea workflow (git.schweitz.net → git.schweitz.internal)
|
|
|
|
### Added
|
|
|
|
- Tests for Smart Page Creation feature (`test_smart_create.py`)
|
|
- Model validation tests for WikiSmartCreateRequest/Response
|
|
- WikiService.smart_create_page method tests
|
|
- Bidirectional entity linking utility tests
|
|
- Endpoint validation tests
|
|
|
|
## [1.1.0] - 2025-12-11
|
|
|
|
### Added
|
|
|
|
- **Smart Page Creation Endpoint** (`POST /wiki/pages/smart-create`)
|
|
- Combines HybridRAG research with LLM content generation
|
|
- Searches existing wiki, knowledge graph, and web for topic context
|
|
- Uses WikiPageWriter to synthesize findings into structured wiki content
|
|
- Auto-generates page path from topic if not provided
|
|
- Returns research summary with source counts
|
|
|
|
- **Bidirectional Entity Linking**
|
|
- New shared utility (`entity_linking_utils.py`) for reusable entity linking
|
|
- Forward links: Links entities mentioned in new pages to existing entity pages
|
|
- Backward links: Updates existing pages that mention the new entity
|
|
- Runs automatically in background after smart page creation
|
|
|
|
- **Version Management**
|
|
- Added `pyproject.toml` with project metadata and version
|
|
- Version is now read from `pyproject.toml` (single source of truth)
|
|
- Health check endpoint returns current version
|
|
- FastAPI docs show current version
|
|
|
|
### Changed
|
|
|
|
- Updated `config.py` to read version from `pyproject.toml`
|
|
- Updated `main.py` to use centralized version
|
|
|
|
## [1.0.0] - 2025-12-10
|
|
|
|
### Added
|
|
|
|
- Initial release extracted from portainer-core
|
|
- Wiki page management (`/wiki/pages` CRUD endpoints)
|
|
- HybridRAG search (`/query/hybrid`) with vector, graph, and web search
|
|
- Knowledge graph operations (`/graph/*`)
|
|
- Vector search operations (`/vector/*`)
|
|
- Knowledge consolidation from search results (`/consolidate/knowledge`)
|
|
- Entity linking and extraction
|
|
- Wiki.js change listener for auto-processing user edits
|
|
- Multi-tenant architecture with user namespace isolation
|