# Changelog All notable changes to Library Desk will be documented in this file. The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html). ## [1.5.0] - 2025-12-26 ### Added - **Central Settings Database** - Tatlock-wide configuration via PostgreSQL - `SettingsClient` for async access to `system_settings` database - User-scoped settings with global fallback - API config storage with `enabled` toggle and per-source category filters - JSON Schema support for future UI rendering - **External API Providers** - Modular `src/apis/` package with swappable implementations - `OpenMeteoProvider` - Weather with geocoding (free, no API key) - `NOSProvider` - Dutch news RSS (16 categories including sports) - `BBCProvider` - English news RSS (21 categories including sports) - `AggregatedNewsProvider` - Merges sources chronologically with category filtering - `AlphaVantageProvider` - Stock/crypto quotes (API key from settings DB) - Abstract base classes for provider interoperability - **Provider Dependency Injection** - `WeatherProviderDep`, `NewsProviderDep`, `AlphaVantageProviderDep` type aliases - Async initialization with settings database integration - Lifecycle management in `shutdown_clients()` - **Development Dependencies** - `requirements-dev.txt` - `pip-audit` for security vulnerability scanning - `ruff` for code quality - Testing packages moved from main requirements ### Changed - News sources configurable via `news.sources` setting - Per-source category filtering via `api.{source}.categories` - Categories default to all if not specified ## [1.4.8] - 2025-12-25 ### Added - **Paperless Orphan Cleanup** - `POST /maintenance/cleanup/paperless` endpoint - Detects documents deleted from Paperless but still indexed in Library Desk - Removes orphaned vectors and graph nodes - Supports `dry_run=true` for preview mode ## [1.4.7] - 2025-12-25 ### Fixed - **Paperless Custom Field Update** - Fixed 400 error when marking documents as indexed - Paperless API requires field ID (integer) not field name (string) - Now looks up `library_indexed` field ID before updating - Webhook params format: `doc_url` and `title` from Jinja templates ### Added - **Webhook Debug Endpoint** - `POST /documents/webhook-capture` for development testing ## [1.4.6] - 2025-12-25 ### Fixed - **Paperless Webhook Payload Format** - Updated model to match Paperless `include_document=true` format - Paperless sends `id` instead of `document_id` - Paperless sends full document data including `content`, `title`, `tags`, etc. - Webhook now uses content from payload, skipping extra Paperless API call - Added `extra = "ignore"` to handle additional Paperless fields ## [1.4.5] - 2025-12-25 ### Added - **Document Storage Integration** - Paperless-ngx integration for PDFs, images, and documents - Event-driven architecture via Paperless webhooks - `POST /documents/webhook` - Receive document events from Paperless workflows - `POST /documents/upload` - Upload files directly to Paperless - `POST /documents/upload-url` - Download and upload documents from URL - `POST /documents/search` - Semantic search across indexed documents - `GET /documents/health` - Paperless connectivity health check - **DocumentSyncService** - Indexes Paperless documents into vectors and graph - Fetches document content via Paperless API - Chunks text and generates embeddings for Qdrant - Creates Document nodes in Neo4j knowledge graph - Supports multi-tenancy via user parameter in webhook URL - **PaperlessClient** - REST API client for Paperless-ngx - Document retrieval, upload, and update operations - Health check support - **Paperless Workflow Configuration** - Production workflow: Document Added (NOT tagged llm-test) → webhook to Library Desk - Test workflow: Document Added (tagged llm-test) → webhook with test user ### Changed - Updated `src/config.py` with Paperless configuration settings - Added `PaperlessDep` dependency injection for document endpoints ## [1.4.4] - 2025-12-24 ### Added - **Test Data Cleanup Endpoint** - `POST /maintenance/cleanup/test-data` - Purges LLM test data from wiki, graph, and vectors - Security-restricted to test user namespace only (`users/llm-tester/*`, `users/llm_tester/*`) - Supports `dry_run=true` (default) to preview before deleting - Scheduler task configured for weekly cleanup (Sunday 3:00 AM) ## [1.4.3] - 2025-12-24 ### Changed - **Volatile Cache System Refactored to Vector Storage** - Backend migrated from Redis to Qdrant for semantic search capability - Data converted to natural language for embedding and semantic retrieval - Collection naming: `volatile_{user}` for per-user isolation - TTL implemented via `ttl_expiry` timestamp in vector payload - Simplified endpoints: - `GET /volatile/search?q=...` - Semantic search across volatile data - `POST /volatile/store?namespace=...&key=...` - Store with query params - `GET /volatile/{namespace}/{key}` - Get specific record - `DELETE /volatile/{namespace}/{key}` - Delete record - Removed namespace-specific URL patterns (simpler API for LLM tool use) ### Added - **HybridRAG Volatile Integration** - Volatile cache now included in multi-source search - Volatile results get priority boost in RRF fusion (current data ranks higher) - New config options: `enable_volatile`, `volatile_limit` (default 1), `volatile_threshold` - Timing breakdown includes `volatile_ms` - **Volatile Cleanup Endpoint** - `POST /maintenance/cleanup/volatile` - Purges expired records across all `volatile_*` collections - Scheduler task for every 10 minutes recommended - Returns per-collection cleanup counts - **Natural Language Conversion** - Structured data converted for embedding - Template-based conversion for each namespace (weather, news, financial, etc.) - Fallback for custom namespaces ## [1.4.2] - 2025-12-24 ### Added - **Volatile Cache System** - Ephemeral data storage with TTL - `GET /volatile/{namespace}/{key}` - Retrieve cached record - `POST /volatile/{namespace}/{key}` - Store/update record with TTL - `DELETE /volatile/{namespace}/{key}` - Remove record - `GET /volatile/{namespace}` - List keys in namespace - `DELETE /volatile/{namespace}` - Clear all records in namespace - `GET /volatile/stats` - Cache statistics by namespace - `GET /volatile/scheduled` - Records needing refresh (for scheduler) - `GET /volatile/namespaces` - List available namespaces with default TTLs - **Volatile Namespaces** - Predefined categories with appropriate TTLs: - `weather` (30min) - Weather conditions and forecasts - `news` (1hr) - Headlines and breaking news - `financial` (5min) - Stock prices, exchange rates - `transit` (5min) - Train/bus schedules, delays - `traffic` (10min) - Commute times, road conditions - `air_quality` (1hr) - Pollution, pollen counts - `sports` (1min) - Live scores, matches - `social` (10min) - Social notifications - `system` (1min) - Service health status - `context` (1hr) - Session state - `custom` (1hr) - User-defined data - **Refresh Schedule Support** - Optional cron expressions for scheduler integration ## [1.4.1] - 2025-12-24 ### Fixed - Wiki.js API token now optional - GraphQL API works without authentication - Container startup failure when `WIKI_GRAPHQL_API` env var not set ## [1.4.0] - 2025-12-24 ### Added - **Maintenance Router** - New `/maintenance` endpoints for system health and cleanup - `GET /maintenance/health` - Lightweight health check (detailed mode available) - `POST /maintenance/cleanup/all` - Full orphan cleanup (vectors + graph) - `POST /maintenance/cleanup/vectors` - Purge orphan vector chunks - `POST /maintenance/cleanup/graph` - Purge orphan graph nodes - `POST /maintenance/reconcile-index` - Combined cleanup + reindex missing pages - **Bidirectional Orphan Detection** - Cross-validate vectors and graph nodes - `find_documents_without_vectors()` - Graph nodes missing vector chunks - `find_chunks_without_graph_nodes()` - Vector chunks missing graph nodes - **Qdrant Client Methods** - Bulk operations for maintenance - `scroll_all_points()` - Iterate all points with pagination - `delete_by_ids()` - Batch delete by point IDs - **Graph Service Cleanup** - Node deletion methods - `delete_document_node()` - Remove document and relationships - `delete_collection_node()` - Remove collection and contained documents - `get_all_document_references()` - Get all document references for validation - **Redis Timestamp Tracking** - `last_cleanup` timestamp for scheduler integration - **Memory System Plan** - Documented three-tier architecture (volatile/documents/knowledge) ### Changed - **Wiki.js Authentication** - Switched from username/password to API token - New `WIKI_GRAPHQL_API` environment variable for JWT token - Deprecated `WIKIJS_USERNAME` and `WIKIJS_PASSWORD` (kept for backwards compatibility) - **Service Dependencies** - Added `VectorServiceDep` and `GraphServiceDep` type aliases ### Fixed - Wiki.js client now properly handles API token auth without login flow ## [1.3.3] - 2025-12-23 ### Added - Temperature parameter to `OllamaClient.generate_text()` for controlling output determinism - `TODO.md` tracking remaining stub endpoints to implement - Wired `/query/semantic` endpoint to VectorService - Wired `/query/graph` endpoint to GraphService ### Changed - **Improved LLM prompts** based on llm-findings.md recommendations: - Keyword extraction: temperature 0.0, negative constraints - LLM re-ranking: temperature 0.0, explicit rules - Conflict detection: temperature 0.0, analysis steps (CoT) - Wiki page creation: temperature 0.3, anti-hallucination constraints - Page reconstruction: temperature 0.2, preservation constraints - Web results analysis: temperature 0.0, conservative approach - Test fixtures now use configurable host (TEST_HOST) instead of Docker hostnames ### Removed - Dead code: unused `get_default_user()` function - Unused imports from routers (wiki.py, graph.py, hybrid_rag.py) - Stub endpoints shadowed by real implementations (/stats, /ingest/document, /ingest/batch) ## [1.3.2] - 2025-12-22 ### Changed - **Consolidated Ollama model configuration** - All LLM operations now use single `OLLAMA_MODEL` environment variable - Removed separate `reranker_model` setting - HybridRAG re-ranking, consolidation analysis, and wiki page writing all use the same model - Improves VRAM efficiency by keeping one model hot - Added `OLLAMA_EMBEDDING_MODEL` environment variable for embedding model (previously overloaded `OLLAMA_MODEL`) - Updated WikiPageWriter to accept settings instead of hardcoded model name ## [1.3.1] - 2025-12-16 ### Fixed - Smart create endpoint missing `content_extractor` dependency causing 500 errors on `POST /wiki/pages/smart-create` ## [1.3.0] - 2025-12-15 ### Changed - **Two-Stage RRF Architecture** - Major refactor to level the playing field between wiki and web results - Stage 1: Vector and graph results merged into single "wiki" ranking using mini-RRF - Stage 2: Final RRF between wiki (single source) and web (single source) - Wiki pages no longer get 2x advantage from appearing in both vector and graph searches - Multi-source confirmation still determines wiki internal ranking - **Skip synonyms in graph search** - LLM-generated synonyms (e.g., "author") no longer match unrelated graph entities (e.g., "author2000") - Vector search still uses synonyms for semantic similarity - Graph search uses only core keywords for exact entity matching ### Added - `VECTOR_SIMILARITY_THRESHOLD` config setting (default: 0.7) to filter weak vector matches - Deduplication in graph search to prevent same document appearing multiple times ### Fixed - Graph search duplicate entity bug where same document could appear twice if entity linked multiple times ## [1.2.1] - 2025-12-15 ### Fixed - HybridRAG router missing `content_extractor` dependency causing 500 errors on `/query/hybrid` endpoint ## [1.2.0] - 2025-12-15 ### Added - **RAG Search Endpoint** (`POST /rag/search`) - Web, news, and image search via SearXNG - Full content extraction using Trafilatura (F1 score 0.958) - Redis caching with configurable TTL - Markdown sources summary for LLM consumption - Returns both extracted content and original snippets - **Content Extraction Endpoints** (`/content/*`) - `POST /content/extract` - Extract content from a single URL - `POST /content/extract/batch` - Batch extraction (up to 20 URLs) - Reusable ContentExtractor client for use across the codebase - **HybridRAG Content Extraction Enhancement** - Web search results now include full extracted content via Trafilatura - Falls back to original snippets if extraction fails - Improves context quality for LLM re-ranking and consumption ### Changed - Added new configuration options: - `SEARCH_CACHE_TTL` - Search cache TTL in seconds (default: 300) - `SEARCH_TIMEOUT` - SearXNG timeout (default: 10s) - `CONTENT_EXTRACTION_TIMEOUT` - Per-URL extraction timeout (default: 5s) - `CONTENT_MAX_LENGTH` - Max extracted content length (default: 2000) - `SEARCH_DEFAULT_LIMIT` - Default search results (default: 10) ### Dependencies - Added `trafilatura~=1.12.0` for content extraction ## [1.1.3] - 2025-12-14 ### Added - Watchtower update trigger in Gitea workflow after successful build ## [1.1.2] - 2025-12-14 ### Fixed - Updated registry login URL in Gitea workflow (git.schweitz.net → git.schweitz.internal) ## [1.1.1] - 2025-12-14 ### Fixed - Updated container registry tag URLs in Gitea workflow (git.schweitz.net → git.schweitz.internal) ### Added - Tests for Smart Page Creation feature (`test_smart_create.py`) - Model validation tests for WikiSmartCreateRequest/Response - WikiService.smart_create_page method tests - Bidirectional entity linking utility tests - Endpoint validation tests ## [1.1.0] - 2025-12-11 ### Added - **Smart Page Creation Endpoint** (`POST /wiki/pages/smart-create`) - Combines HybridRAG research with LLM content generation - Searches existing wiki, knowledge graph, and web for topic context - Uses WikiPageWriter to synthesize findings into structured wiki content - Auto-generates page path from topic if not provided - Returns research summary with source counts - **Bidirectional Entity Linking** - New shared utility (`entity_linking_utils.py`) for reusable entity linking - Forward links: Links entities mentioned in new pages to existing entity pages - Backward links: Updates existing pages that mention the new entity - Runs automatically in background after smart page creation - **Version Management** - Added `pyproject.toml` with project metadata and version - Version is now read from `pyproject.toml` (single source of truth) - Health check endpoint returns current version - FastAPI docs show current version ### Changed - Updated `config.py` to read version from `pyproject.toml` - Updated `main.py` to use centralized version ## [1.0.0] - 2025-12-10 ### Added - Initial release extracted from portainer-core - Wiki page management (`/wiki/pages` CRUD endpoints) - HybridRAG search (`/query/hybrid`) with vector, graph, and web search - Knowledge graph operations (`/graph/*`) - Vector search operations (`/vector/*`) - Knowledge consolidation from search results (`/consolidate/knowledge`) - Entity linking and extraction - Wiki.js change listener for auto-processing user edits - Multi-tenant architecture with user namespace isolation