Files
library-desk/CHANGELOG.md
T
jpmschweitzerandClaude Opus 4.5 2552bfd1f9
Build and Push / build (release) Successful in 33s
fix: make Wiki.js API token optional for open GraphQL endpoints
The Wiki.js GraphQL API is accessible without authentication.
Make WIKI_GRAPHQL_API env var optional with empty default to fix
container startup failures.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2025-12-24 16:59:41 +01:00

8.5 KiB

Changelog

All notable changes to Library Desk will be documented in this file.

The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.

[1.4.1] - 2025-12-24

Fixed

  • Wiki.js API token now optional - GraphQL API works without authentication
  • Container startup failure when WIKI_GRAPHQL_API env var not set

[1.4.0] - 2025-12-24

Added

  • Maintenance Router - New /maintenance endpoints for system health and cleanup
    • GET /maintenance/health - Lightweight health check (detailed mode available)
    • POST /maintenance/cleanup/all - Full orphan cleanup (vectors + graph)
    • POST /maintenance/cleanup/vectors - Purge orphan vector chunks
    • POST /maintenance/cleanup/graph - Purge orphan graph nodes
    • POST /maintenance/reconcile-index - Combined cleanup + reindex missing pages
  • Bidirectional Orphan Detection - Cross-validate vectors and graph nodes
    • find_documents_without_vectors() - Graph nodes missing vector chunks
    • find_chunks_without_graph_nodes() - Vector chunks missing graph nodes
  • Qdrant Client Methods - Bulk operations for maintenance
    • scroll_all_points() - Iterate all points with pagination
    • delete_by_ids() - Batch delete by point IDs
  • Graph Service Cleanup - Node deletion methods
    • delete_document_node() - Remove document and relationships
    • delete_collection_node() - Remove collection and contained documents
    • get_all_document_references() - Get all document references for validation
  • Redis Timestamp Tracking - last_cleanup timestamp for scheduler integration
  • Memory System Plan - Documented three-tier architecture (volatile/documents/knowledge)

Changed

  • Wiki.js Authentication - Switched from username/password to API token
    • New WIKI_GRAPHQL_API environment variable for JWT token
    • Deprecated WIKIJS_USERNAME and WIKIJS_PASSWORD (kept for backwards compatibility)
  • Service Dependencies - Added VectorServiceDep and GraphServiceDep type aliases

Fixed

  • Wiki.js client now properly handles API token auth without login flow

[1.3.3] - 2025-12-23

Added

  • Temperature parameter to OllamaClient.generate_text() for controlling output determinism
  • TODO.md tracking remaining stub endpoints to implement
  • Wired /query/semantic endpoint to VectorService
  • Wired /query/graph endpoint to GraphService

Changed

  • Improved LLM prompts based on llm-findings.md recommendations:
    • Keyword extraction: temperature 0.0, negative constraints
    • LLM re-ranking: temperature 0.0, explicit rules
    • Conflict detection: temperature 0.0, analysis steps (CoT)
    • Wiki page creation: temperature 0.3, anti-hallucination constraints
    • Page reconstruction: temperature 0.2, preservation constraints
    • Web results analysis: temperature 0.0, conservative approach
  • Test fixtures now use configurable host (TEST_HOST) instead of Docker hostnames

Removed

  • Dead code: unused get_default_user() function
  • Unused imports from routers (wiki.py, graph.py, hybrid_rag.py)
  • Stub endpoints shadowed by real implementations (/stats, /ingest/document, /ingest/batch)

[1.3.2] - 2025-12-22

Changed

  • Consolidated Ollama model configuration - All LLM operations now use single OLLAMA_MODEL environment variable
    • Removed separate reranker_model setting
    • HybridRAG re-ranking, consolidation analysis, and wiki page writing all use the same model
    • Improves VRAM efficiency by keeping one model hot
  • Added OLLAMA_EMBEDDING_MODEL environment variable for embedding model (previously overloaded OLLAMA_MODEL)
  • Updated WikiPageWriter to accept settings instead of hardcoded model name

[1.3.1] - 2025-12-16

Fixed

  • Smart create endpoint missing content_extractor dependency causing 500 errors on POST /wiki/pages/smart-create

[1.3.0] - 2025-12-15

Changed

  • Two-Stage RRF Architecture - Major refactor to level the playing field between wiki and web results

    • Stage 1: Vector and graph results merged into single "wiki" ranking using mini-RRF
    • Stage 2: Final RRF between wiki (single source) and web (single source)
    • Wiki pages no longer get 2x advantage from appearing in both vector and graph searches
    • Multi-source confirmation still determines wiki internal ranking
  • Skip synonyms in graph search - LLM-generated synonyms (e.g., "author") no longer match unrelated graph entities (e.g., "author2000")

    • Vector search still uses synonyms for semantic similarity
    • Graph search uses only core keywords for exact entity matching

Added

  • VECTOR_SIMILARITY_THRESHOLD config setting (default: 0.7) to filter weak vector matches
  • Deduplication in graph search to prevent same document appearing multiple times

Fixed

  • Graph search duplicate entity bug where same document could appear twice if entity linked multiple times

[1.2.1] - 2025-12-15

Fixed

  • HybridRAG router missing content_extractor dependency causing 500 errors on /query/hybrid endpoint

[1.2.0] - 2025-12-15

Added

  • RAG Search Endpoint (POST /rag/search)

    • Web, news, and image search via SearXNG
    • Full content extraction using Trafilatura (F1 score 0.958)
    • Redis caching with configurable TTL
    • Markdown sources summary for LLM consumption
    • Returns both extracted content and original snippets
  • Content Extraction Endpoints (/content/*)

    • POST /content/extract - Extract content from a single URL
    • POST /content/extract/batch - Batch extraction (up to 20 URLs)
    • Reusable ContentExtractor client for use across the codebase
  • HybridRAG Content Extraction Enhancement

    • Web search results now include full extracted content via Trafilatura
    • Falls back to original snippets if extraction fails
    • Improves context quality for LLM re-ranking and consumption

Changed

  • Added new configuration options:
    • SEARCH_CACHE_TTL - Search cache TTL in seconds (default: 300)
    • SEARCH_TIMEOUT - SearXNG timeout (default: 10s)
    • CONTENT_EXTRACTION_TIMEOUT - Per-URL extraction timeout (default: 5s)
    • CONTENT_MAX_LENGTH - Max extracted content length (default: 2000)
    • SEARCH_DEFAULT_LIMIT - Default search results (default: 10)

Dependencies

  • Added trafilatura~=1.12.0 for content extraction

[1.1.3] - 2025-12-14

Added

  • Watchtower update trigger in Gitea workflow after successful build

[1.1.2] - 2025-12-14

Fixed

  • Updated registry login URL in Gitea workflow (git.schweitz.net → git.schweitz.internal)

[1.1.1] - 2025-12-14

Fixed

  • Updated container registry tag URLs in Gitea workflow (git.schweitz.net → git.schweitz.internal)

Added

  • Tests for Smart Page Creation feature (test_smart_create.py)
    • Model validation tests for WikiSmartCreateRequest/Response
    • WikiService.smart_create_page method tests
    • Bidirectional entity linking utility tests
    • Endpoint validation tests

[1.1.0] - 2025-12-11

Added

  • Smart Page Creation Endpoint (POST /wiki/pages/smart-create)

    • Combines HybridRAG research with LLM content generation
    • Searches existing wiki, knowledge graph, and web for topic context
    • Uses WikiPageWriter to synthesize findings into structured wiki content
    • Auto-generates page path from topic if not provided
    • Returns research summary with source counts
  • Bidirectional Entity Linking

    • New shared utility (entity_linking_utils.py) for reusable entity linking
    • Forward links: Links entities mentioned in new pages to existing entity pages
    • Backward links: Updates existing pages that mention the new entity
    • Runs automatically in background after smart page creation
  • Version Management

    • Added pyproject.toml with project metadata and version
    • Version is now read from pyproject.toml (single source of truth)
    • Health check endpoint returns current version
    • FastAPI docs show current version

Changed

  • Updated config.py to read version from pyproject.toml
  • Updated main.py to use centralized version

[1.0.0] - 2025-12-10

Added

  • Initial release extracted from portainer-core
  • Wiki page management (/wiki/pages CRUD endpoints)
  • HybridRAG search (/query/hybrid) with vector, graph, and web search
  • Knowledge graph operations (/graph/*)
  • Vector search operations (/vector/*)
  • Knowledge consolidation from search results (/consolidate/knowledge)
  • Entity linking and extraction
  • Wiki.js change listener for auto-processing user edits
  • Multi-tenant architecture with user namespace isolation