Compare commits

...
7 Commits
Author SHA1 Message Date
jpmschweitzerandClaude Opus 4.5 2ac4595b37 fix: update registry login URL in Gitea workflow
🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2025-12-14 11:21:29 +01:00
jpmschweitzerandClaude Opus 4.5 37552b926f fix: update registry URL and add smart create tests
Build and Push / build (release) Failing after 12s
- Updated container registry URL in Gitea workflow (git.schweitz.net → git.schweitz.internal)
- Added comprehensive tests for smart page creation feature

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2025-12-14 11:16:44 +01:00
jpmschweitzer 23ccd5ca5b added AGENTS.md llm instruction file 2025-12-11 21:16:49 +01:00
jpmschweitzerandClaude Opus 4.5 2bfb1e29d7 fix: include pyproject.toml in Docker image for version info
The version was showing as 0.0.0 because pyproject.toml wasn't
being copied into the Docker image.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2025-12-11 21:08:53 +01:00
jpmschweitzerandClaude Opus 4.5 0f224b460e docs: add AGENTS.md for Claude Code context
Build and Push / build (release) Failing after 16s
🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2025-12-11 20:29:04 +01:00
jpmschweitzerandClaude Opus 4.5 3deed7cbcb chore: add version management and changelog
- Add pyproject.toml with project metadata and version (1.1.0)
- Update config.py to read version from pyproject.toml
- Update main.py to use centralized version in FastAPI app
- Add CHANGELOG.md documenting v1.0.0 and v1.1.0 changes

Version is now the single source of truth in pyproject.toml and is
displayed in the health check endpoint and API docs.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2025-12-11 20:23:56 +01:00
jpmschweitzerandClaude Opus 4.5 e05e7aeae3 feat: add smart page creation endpoint with HybridRAG research
Add POST /wiki/pages/smart-create endpoint that combines research with
content generation for the librarian agent:

- Run HybridRAG search on topic (wiki + graph + web)
- Use LLM (WikiPageWriter) to synthesize findings into wiki content
- Create page with proper attribution and sources
- Schedule background tasks for vector/graph indexing
- Apply bidirectional entity linking (forward + backward links)

New files:
- src/services/entity_linking_utils.py - shared entity linking helper

Modified:
- src/models/wiki.py - WikiSmartCreateRequest/Response models
- src/services/wiki_service.py - smart_create_page() method
- src/routers/wiki.py - /pages/smart-create endpoint

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2025-12-11 20:23:46 +01:00
12 changed files with 1208 additions and 9 deletions
+3 -3
View File
@@ -13,7 +13,7 @@ jobs:
- name: Login to Gitea Registry
uses: docker/login-action@v3
with:
registry: git.schweitz.net
registry: git.schweitz.internal
username: ${{ secrets.REGISTRY_USER }}
password: ${{ secrets.REGISTRY_PASSWORD }}
@@ -23,5 +23,5 @@ jobs:
context: .
push: true
tags: |
git.schweitz.net/jpmschweitzer/library-desk:latest
git.schweitz.net/jpmschweitzer/library-desk:${{ github.ref_name }}
git.schweitz.internal/jpmschweitzer/library-desk:latest
git.schweitz.internal/jpmschweitzer/library-desk:${{ github.ref_name }}
+46
View File
@@ -0,0 +1,46 @@
# AGENTS.md
> **Start every session by reading this file.**
> This file outlines the operational protocols, coding standards, and architectural decisions for this FastAPI project.
## 1. Agent Operational Protocols
### 🧠 Work Patterns (Plan-Act-Reflect)
* **Plan:** Before writing code, briefly outline your plan. Identify which files you will touch and what the side effects might be.
* **Act:** Execute the changes in small, atomic steps.
* **Reflect:** After coding, verify your work. Did you break existing tests? Did you add new tests?
### 🛡️ Git Discipline
* **NEVER commit to `main` or `master` directly.** Always create a feature branch: `feature/your-feature-name` or `fix/issue-description`.
* **Commit Messages:** Use the [Conventional Commits](https://www.conventionalcommits.org/) format.
* `feat: add user login endpoint`
* `fix: resolve database connection timeout`
* `refactor: split monolith dependency file`
* **Atomic Commits:** Keep commits small. One logical change = one commit.
### 📝 Changelog Maintenance
* **Update `CHANGELOG.md`** with every user-facing change.
* Format: `## [Unreleased] - YYYY-MM-DD` followed by `### Added`, `### Changed`, or `### Fixed`.
---
## 2. FastAPI Architecture & Best Practices
*Reference: [FastAPI Best Practices](https://github.com/zhanymkanov/fastapi-best-practices)*
### 📂 Project Structure (Directory-based, NOT File-type based)
Do **not** group files by type (e.g., one huge `routers` folder). Group by **domain/module** inside a `src/` directory.
**Correct Structure:**
```text
src/
├── auth/
│ ├── router.py # Endpoints
│ ├── schemas.py # Pydantic models
│ ├── service.py # Business logic (CRUD, etc.)
│ ├── dependencies.py# Module-specific dependencies
│ └── config.py # Module-specific settings
├── posts/
│ ├── router.py
│ └── ...
└── main.py # App entry point
+68
View File
@@ -0,0 +1,68 @@
# Changelog
All notable changes to Library Desk will be documented in this file.
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
## [1.1.2] - 2025-12-14
### Fixed
- Updated registry login URL in Gitea workflow (git.schweitz.net → git.schweitz.internal)
## [1.1.1] - 2025-12-14
### Fixed
- Updated container registry tag URLs in Gitea workflow (git.schweitz.net → git.schweitz.internal)
### Added
- Tests for Smart Page Creation feature (`test_smart_create.py`)
- Model validation tests for WikiSmartCreateRequest/Response
- WikiService.smart_create_page method tests
- Bidirectional entity linking utility tests
- Endpoint validation tests
## [1.1.0] - 2025-12-11
### Added
- **Smart Page Creation Endpoint** (`POST /wiki/pages/smart-create`)
- Combines HybridRAG research with LLM content generation
- Searches existing wiki, knowledge graph, and web for topic context
- Uses WikiPageWriter to synthesize findings into structured wiki content
- Auto-generates page path from topic if not provided
- Returns research summary with source counts
- **Bidirectional Entity Linking**
- New shared utility (`entity_linking_utils.py`) for reusable entity linking
- Forward links: Links entities mentioned in new pages to existing entity pages
- Backward links: Updates existing pages that mention the new entity
- Runs automatically in background after smart page creation
- **Version Management**
- Added `pyproject.toml` with project metadata and version
- Version is now read from `pyproject.toml` (single source of truth)
- Health check endpoint returns current version
- FastAPI docs show current version
### Changed
- Updated `config.py` to read version from `pyproject.toml`
- Updated `main.py` to use centralized version
## [1.0.0] - 2025-12-10
### Added
- Initial release extracted from portainer-core
- Wiki page management (`/wiki/pages` CRUD endpoints)
- HybridRAG search (`/query/hybrid`) with vector, graph, and web search
- Knowledge graph operations (`/graph/*`)
- Vector search operations (`/vector/*`)
- Knowledge consolidation from search results (`/consolidate/knowledge`)
- Entity linking and extraction
- Wiki.js change listener for auto-processing user edits
- Multi-tenant architecture with user namespace isolation
+1
View File
@@ -12,6 +12,7 @@ COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
# Copy application
COPY pyproject.toml .
COPY src/ ./src/
COPY static/ ./static/
+31
View File
@@ -0,0 +1,31 @@
[project]
name = "library-desk"
version = "1.1.2"
description = "Coordination service for The Library system - HybridRAG queries, document ingestion, entity extraction, and knowledge consolidation"
readme = "README.md"
requires-python = ">=3.12"
license = {text = "MIT"}
authors = [
{name = "JP Schweitzer"}
]
keywords = ["rag", "knowledge-graph", "wiki", "semantic-search", "neo4j", "qdrant"]
classifiers = [
"Development Status :: 4 - Beta",
"Framework :: FastAPI",
"Intended Audience :: Developers",
"License :: OSI Approved :: MIT License",
"Programming Language :: Python :: 3",
"Programming Language :: Python :: 3.12",
]
[project.urls]
Homepage = "https://github.com/jpmschweitzer/library-desk"
Documentation = "https://github.com/jpmschweitzer/library-desk#readme"
[build-system]
requires = ["setuptools>=61.0"]
build-backend = "setuptools.build_meta"
[tool.setuptools.packages.find]
where = ["."]
include = ["src*"]
+12 -1
View File
@@ -4,9 +4,20 @@ Following best practices: modular settings, environment-based config.
"""
from functools import lru_cache
from pathlib import Path
from pydantic import Field
from pydantic_settings import BaseSettings, SettingsConfigDict
# Read version from pyproject.toml
try:
import tomllib
_pyproject_path = Path(__file__).parent.parent / "pyproject.toml"
with open(_pyproject_path, "rb") as f:
_pyproject = tomllib.load(f)
__version__ = _pyproject["project"]["version"]
except Exception:
__version__ = "0.0.0" # Fallback if pyproject.toml not found
class Settings(BaseSettings):
"""Application settings loaded from environment variables."""
@@ -75,7 +86,7 @@ class Settings(BaseSettings):
# Application
app_name: str = Field(default="Library Desk", description="Application name")
app_version: str = Field(default="1.0.0", description="Application version")
app_version: str = Field(default=__version__, description="Application version")
debug: bool = Field(default=False, description="Debug mode")
@property
+2 -2
View File
@@ -16,7 +16,7 @@ from typing import Dict, Any
import logging
from pathlib import Path
from src.config import Settings, get_settings
from src.config import Settings, get_settings, __version__
from src.core.dependencies import verify_api_key
# Configure logging
@@ -30,7 +30,7 @@ logger = logging.getLogger(__name__)
app = FastAPI(
title="Library Desk API",
description="Coordination service for The Library system - HybridRAG queries, document ingestion, entity extraction, and mind map generation",
version="1.0.0",
version=__version__,
docs_url="/docs",
redoc_url="/redoc",
)
+42 -1
View File
@@ -8,7 +8,7 @@ Models for:
"""
from pydantic import BaseModel, Field, field_validator
from typing import Optional, List
from typing import Optional, List, Dict, Any
from datetime import datetime
@@ -190,3 +190,44 @@ class DossierOperationResponse(BaseModel):
dossier_name: str = Field(..., description="Dossier name")
index_page_id: Optional[int] = Field(None, description="Index page ID (if created)")
index_page_path: Optional[str] = Field(None, description="Index page path (if created)")
# Smart create models (HybridRAG-powered page creation)
class WikiSmartCreateRequest(BaseModel):
"""Request model for smart page creation with research."""
topic: str = Field(..., min_length=1, max_length=500, description="Topic to research and create page about")
path: Optional[str] = Field(None, description="Page path (auto-generated from topic if not provided)")
tags: List[str] = Field(default_factory=list, description="Tags for the page")
user: Optional[str] = Field(None, description="User identifier")
include_web_research: bool = Field(default=True, description="Include web search results")
include_wiki_search: bool = Field(default=True, description="Include existing wiki knowledge")
@field_validator("tags")
@classmethod
def validate_tags(cls, v: List[str]) -> List[str]:
"""Validate and clean tags."""
cleaned = [tag.strip() for tag in v if tag.strip()]
return list(set(cleaned))
@field_validator("path")
@classmethod
def validate_path(cls, v: Optional[str]) -> Optional[str]:
"""Validate page path if provided."""
if v is None:
return None
# Ensure path starts with /
if not v.startswith("/"):
v = f"/{v}"
# Remove trailing slash
if v.endswith("/") and v != "/":
v = v.rstrip("/")
return v
class WikiSmartCreateResponse(BaseModel):
"""Response model for smart page creation."""
page: WikiPage = Field(..., description="Created wiki page")
research_summary: Dict[str, Any] = Field(..., description="Summary of research used")
sources_used: int = Field(..., description="Number of sources incorporated")
search_id: Optional[str] = Field(None, description="HybridRAG search ID for reference")
entity_linking: Dict[str, int] = Field(default_factory=dict, description="Entity linking statistics")
+127 -2
View File
@@ -13,7 +13,8 @@ import logging
from src.models.wiki import (
WikiPage, WikiPageList, WikiPageCreate, WikiPageUpdate, WikiPageMove,
WikiOperationResponse, WikiSearchResponse,
DossierList, WikiSearchResult
DossierList, WikiSearchResult,
WikiSmartCreateRequest, WikiSmartCreateResponse
)
from src.services.wiki_service import WikiService
from src.services.graph_service import GraphService
@@ -22,8 +23,15 @@ from src.clients.wikijs_client import WikiJSClient
from src.clients.neo4j_client import Neo4jClient
from src.clients.qdrant_client import QdrantClientWrapper
from src.clients.ollama_client import OllamaClient
from src.core.dependencies import WikiJSDep, Neo4jDep, QdrantDep, OllamaDep, verify_api_key
from src.core.dependencies import (
WikiJSDep, Neo4jDep, QdrantDep, OllamaDep, SearXNGDep,
verify_api_key, get_settings, get_hybrid_rag_service, get_ingestion_service
)
from src.core.multi_tenancy import DEFAULT_USER
from src.services.hybrid_rag_service import HybridRAGService
from src.services.wiki_page_writer import WikiPageWriter
from src.services.entity_linking_utils import apply_bidirectional_entity_linking
from src.config import Settings
logger = logging.getLogger(__name__)
@@ -163,6 +171,123 @@ async def create_page(
raise HTTPException(status_code=500, detail="Internal server error")
@router.post("/pages/smart-create", response_model=WikiSmartCreateResponse, status_code=201)
async def smart_create_page(
request: WikiSmartCreateRequest,
background_tasks: BackgroundTasks,
wiki_client: WikiJSDep,
neo4j_client: Neo4jDep,
qdrant_client: QdrantDep,
ollama_client: OllamaDep,
searxng_client: SearXNGDep,
settings: Settings = Depends(get_settings),
api_key: str = Depends(verify_api_key)
):
"""
Create wiki page with intelligent research.
Combines HybridRAG search with LLM content generation to create
rich, well-researched wiki pages in a single API call.
**Process:**
1. Runs HybridRAG search on the topic (wiki + graph + web)
2. Uses LLM to synthesize findings into structured wiki content
3. Creates the page with proper attribution/sources
4. Indexes into vectors + knowledge graph (background)
5. Applies bidirectional entity linking (background)
**Example Request:**
```json
{
"topic": "Docker orchestration patterns",
"path": "/technology/containers/docker-orchestration",
"tags": ["technology", "devops", "containers"],
"user": "jpmschweitzer",
"include_web_research": true,
"include_wiki_search": true
}
```
**Returns:**
- Created page with ID, path, content
- Research summary (wiki/web/graph result counts)
- Entity linking statistics (forward/backward links)
"""
try:
user = request.user or DEFAULT_USER
# Build services
wiki_service = WikiService(wiki_client)
vector_service = VectorService(qdrant_client, wiki_client, ollama_client)
graph_service = GraphService(neo4j_client, wiki_client)
hybrid_rag_service = HybridRAGService(
vector_service=vector_service,
graph_service=graph_service,
searxng_client=searxng_client,
ollama_client=ollama_client,
settings=settings
)
wiki_page_writer = WikiPageWriter(ollama_client=ollama_client)
# Step 1-5: Research + Generate + Create page
page, research_data = await wiki_service.smart_create_page(
topic=request.topic,
user=user,
path=request.path,
tags=request.tags,
hybrid_rag_service=hybrid_rag_service,
wiki_page_writer=wiki_page_writer,
include_web=request.include_web_research,
include_wiki=request.include_wiki_search
)
# Schedule graph and vector updates in background
background_tasks.add_task(
graph_service.update_from_page,
page_id=page.id,
user=user
)
background_tasks.add_task(
vector_service.update_from_page,
page_id=page.id,
user=user
)
# Schedule bidirectional entity linking in background
async def run_entity_linking():
ingestion_service = get_ingestion_service()
return await apply_bidirectional_entity_linking(
page_id=page.id,
page_title=page.title,
user=user,
neo4j_client=neo4j_client,
wiki_service=wiki_service,
ingestion_service=ingestion_service
)
background_tasks.add_task(run_entity_linking)
logger.info(
f"Smart page created: id={page.id}, path={page.path}, "
f"sources={research_data['sources_used']}"
)
return WikiSmartCreateResponse(
page=page,
research_summary=research_data["research_summary"],
sources_used=research_data["sources_used"],
search_id=research_data["search_id"],
entity_linking={"forward_links": 0, "backward_links": 0, "pages_updated": 0}
# Note: entity_linking stats are 0 here as it runs in background
)
except ValueError as e:
raise HTTPException(status_code=400, detail=str(e))
except Exception as e:
logger.error(f"Failed to smart create page: {e}", exc_info=True)
raise HTTPException(status_code=500, detail="Internal server error")
@router.put("/pages/{page_id}", response_model=WikiPage)
async def update_page(
page_id: int,
+160
View File
@@ -0,0 +1,160 @@
"""
Shared entity linking utilities for Library Desk.
Provides bidirectional entity linking functionality that can be used by:
- Consolidation service (knowledge consolidation)
- Wiki router (smart page creation)
- Any other service that creates wiki pages
"""
import logging
from typing import Dict, Any, Optional
from src.core.multi_tenancy import get_neo4j_user_base_label
logger = logging.getLogger(__name__)
async def apply_bidirectional_entity_linking(
page_id: int,
page_title: str,
user: str,
neo4j_client: "Neo4jClient",
wiki_service: "WikiService",
ingestion_service: Optional["IngestionService"] = None
) -> Dict[str, int]:
"""
Apply bidirectional entity linking after page creation/update.
This runs AFTER ingestion so entities are extracted and in the graph.
Steps:
1. Link entities in the new page (forward links to existing entities)
2. Find pages that mention the new entity (reverse references)
3. Link entities in those pages (backward links to the new entity)
Args:
page_id: Wiki page ID
page_title: Page title (used to find reverse references)
user: User identifier
neo4j_client: Neo4j client for graph queries
wiki_service: Wiki service for page operations
ingestion_service: Optional ingestion service for re-indexing
Returns:
Dict with link counts: {
"forward_links": int, # Links added to the new page
"backward_links": int, # Links added to other pages pointing to new page
"pages_updated": int # Number of other pages updated
}
"""
from src.routers.entity_linking import (
link_entities_in_page,
EntityLinkingRequest,
get_entities_with_paths,
add_entity_links_to_content
)
from src.core.dependencies import get_graph_service, get_wiki_service, get_ingestion_service
from src.models.wiki import WikiPageUpdate
forward_links = 0
backward_links = 0
pages_updated = 0
try:
graph_service = get_graph_service()
# Use provided services or get defaults
wiki_svc = wiki_service
ingestion_svc = ingestion_service or get_ingestion_service()
# STEP 1: Forward linking - link entities in the new page
logger.info(f"Step 1/3: Linking entities in page {page_id} ('{page_title}')")
try:
forward_result = await link_entities_in_page(
request=EntityLinkingRequest(
user=user,
page_id=page_id,
create_relationships=True,
re_index_if_changed=False # Already indexed, no need to re-index
),
wiki_service=wiki_svc,
graph_service=graph_service,
ingestion_service=ingestion_svc,
api_key="" # Internal call, no auth needed
)
forward_links = forward_result.content_links_added
logger.info(f"Added {forward_links} forward links in page {page_id}")
except Exception as e:
logger.error(f"Failed to add forward links: {e}")
# STEP 2: Find reverse references - which pages mention this new entity?
logger.info(f"Step 2/3: Finding pages that mention '{page_title}'")
user_base_label = get_neo4j_user_base_label(user)
# Query to find documents that mention entities with this page's title
reverse_query = f"""
// Find entities with the same name as the page title
MATCH (e:{user_base_label})
WHERE toLower(e.name) = toLower($title)
AND NOT e:Document
// Find documents that mention those entities
MATCH (d:Document)-[r:MENTIONS]->(e)
WHERE d.page_id <> $page_id // Exclude the page itself
RETURN DISTINCT d.page_id as page_id, d.title as title
LIMIT 50
"""
try:
reverse_refs = await neo4j_client.execute_query(
reverse_query,
{"title": page_title, "page_id": page_id}
)
logger.info(f"Found {len(reverse_refs)} pages that mention '{page_title}'")
except Exception as e:
logger.error(f"Failed to find reverse references: {e}")
reverse_refs = []
# STEP 3: Backward linking - add links in those pages to the new entity
if reverse_refs:
logger.info(f"Step 3/3: Adding backward links in {len(reverse_refs)} pages")
for ref in reverse_refs:
try:
backward_result = await link_entities_in_page(
request=EntityLinkingRequest(
user=user,
page_id=ref['page_id'],
create_relationships=False, # Relationships already exist
re_index_if_changed=False # Don't re-index for link updates
),
wiki_service=wiki_svc,
graph_service=graph_service,
ingestion_service=ingestion_svc,
api_key=""
)
if backward_result.content_links_added > 0:
backward_links += backward_result.content_links_added
pages_updated += 1
logger.info(
f"Added {backward_result.content_links_added} links "
f"in page {ref['page_id']} ('{ref['title']}')"
)
except Exception as e:
logger.error(f"Failed to add backward links in page {ref['page_id']}: {e}")
else:
logger.info("Step 3/3: No reverse references found, skipping backward linking")
return {
"forward_links": forward_links,
"backward_links": backward_links,
"pages_updated": pages_updated
}
except Exception as e:
logger.error(f"Bidirectional entity linking failed: {e}", exc_info=True)
return {
"forward_links": 0,
"backward_links": 0,
"pages_updated": 0
}
+163
View File
@@ -431,3 +431,166 @@ class WikiService:
WikiPageList filtered by dossier tag
"""
return await self.list_pages(user, tag=dossier_name, limit=limit)
async def smart_create_page(
self,
topic: str,
user: str,
path: Optional[str],
tags: List[str],
hybrid_rag_service: "HybridRAGService",
wiki_page_writer: "WikiPageWriter",
include_web: bool = True,
include_wiki: bool = True
) -> tuple["WikiPage", Dict[str, Any]]:
"""
Create wiki page with research from HybridRAG.
This method combines research + content generation + page creation:
1. Run HybridRAG search on topic
2. Format results for WikiPageWriter
3. Generate page content with LLM
4. Create page in Wiki.js
5. Return page + research summary
Args:
topic: Topic to research and create page about
user: User identifier
path: Optional page path (auto-generated from topic if not provided)
tags: Tags for the page
hybrid_rag_service: HybridRAG service for multi-source search
wiki_page_writer: WikiPageWriter for LLM content generation
include_web: Include web search results
include_wiki: Include existing wiki knowledge
Returns:
Tuple of (created WikiPage, research summary dict)
"""
from src.models.hybrid_rag import HybridRAGConfig
logger.info(f"Smart create page: topic='{topic}', user='{user}'")
# Step 1: Run HybridRAG search on the topic
config = HybridRAGConfig(
enable_vector=include_wiki,
enable_graph=include_wiki,
enable_web=include_web,
enable_reranking=True,
enable_enrichment=True,
final_result_count=15 # Get more results for rich content
)
search_response = await hybrid_rag_service.search(
query=topic,
user=user,
config=config
)
logger.info(
f"HybridRAG search completed: {search_response.total_results} results, "
f"search_id={search_response.search_id}"
)
# Step 2: Format results for WikiPageWriter
source_information = []
wiki_results_count = 0
web_results_count = 0
graph_entities_count = 0
for result in search_response.results:
source_type = result.source_type
if "web" in source_type:
web_results_count += 1
source_information.append({
"title": result.title,
"url": result.url or "",
"content": result.content[:500] if result.content else ""
})
elif "vector" in source_type or "graph" in source_type:
wiki_results_count += 1
# For wiki results, use page path as URL
source_information.append({
"title": result.title,
"url": f"/{result.page_path}" if result.page_path else "",
"content": result.content[:500] if result.content else ""
})
# Count entities from related dossiers
if result.related_dossiers:
graph_entities_count += len(result.related_dossiers)
# Step 3: Generate page content with LLM
# Use topic as summary and let WikiPageWriter create structured content
topic_summary = f"Research findings about: {topic}"
if search_response.keywords:
topic_summary += f"\n\nKey concepts: {', '.join(search_response.keywords.core_keywords)}"
# Extract entities from search results for knowledge graph linking
entities = []
if search_response.keywords and search_response.keywords.core_keywords:
entities = search_response.keywords.core_keywords[:10]
# Get related documents for cross-linking
related_docs = []
for result in search_response.results[:5]:
if result.page_path:
related_docs.append(f"[{result.title}](/{result.page_path})")
content = await wiki_page_writer.create_page(
title=topic,
topic_summary=topic_summary,
source_information=source_information[:10], # Limit sources
entities=entities,
related_docs=related_docs
)
logger.info(f"Generated page content: {len(content)} characters")
# Step 4: Auto-generate path from topic if not provided
if not path:
# Convert topic to kebab-case path
import re
path_slug = topic.lower()
path_slug = re.sub(r'[^\w\s-]', '', path_slug) # Remove special chars
path_slug = re.sub(r'\s+', '-', path_slug) # Spaces to hyphens
path_slug = re.sub(r'-+', '-', path_slug) # Multiple hyphens to single
path_slug = path_slug.strip('-')
# Infer category from tags or use reference
category = "reference"
if tags:
category = tags[0].lower()
path = f"/{category}/{path_slug}"
# Step 5: Create page using existing create_page method
from src.models.wiki import WikiPageCreate
page_data = WikiPageCreate(
title=topic,
path=path,
content=content,
description=f"Research summary about {topic}",
tags=tags,
user=user
)
page = await self.create_page(page_data)
logger.info(f"Created page: id={page.id}, path={page.path}")
# Build research summary
research_summary = {
"wiki_results": wiki_results_count,
"web_results": web_results_count,
"graph_entities": graph_entities_count,
"keywords_extracted": len(search_response.keywords.core_keywords) if search_response.keywords else 0,
"timing_ms": search_response.timing.total_ms if search_response.timing else 0
}
return page, {
"research_summary": research_summary,
"sources_used": len(source_information),
"search_id": search_response.search_id
}
+553
View File
@@ -0,0 +1,553 @@
"""
Tests for Smart Page Creation functionality.
Tests the new smart-create feature including:
- WikiSmartCreateRequest/Response models
- smart_create_page() method in WikiService
- Bidirectional entity linking utilities
- POST /wiki/pages/smart-create endpoint
"""
import pytest
from unittest.mock import AsyncMock, MagicMock, patch
from src.models.wiki import (
WikiSmartCreateRequest,
WikiSmartCreateResponse,
WikiPage
)
# =============================================================================
# Model Tests
# =============================================================================
class TestWikiSmartCreateRequest:
"""Tests for WikiSmartCreateRequest model validation."""
def test_minimal_request(self):
"""Test request with only required field."""
request = WikiSmartCreateRequest(topic="Docker containers")
assert request.topic == "Docker containers"
assert request.path is None
assert request.tags == []
assert request.user is None
assert request.include_web_research is True
assert request.include_wiki_search is True
def test_full_request(self):
"""Test request with all fields."""
request = WikiSmartCreateRequest(
topic="Kubernetes orchestration",
path="/technology/kubernetes",
tags=["devops", "containers"],
user="testuser",
include_web_research=False,
include_wiki_search=True
)
assert request.topic == "Kubernetes orchestration"
assert request.path == "/technology/kubernetes"
# Tags are deduplicated via set, so order is not guaranteed
assert set(request.tags) == {"devops", "containers"}
assert request.user == "testuser"
assert request.include_web_research is False
assert request.include_wiki_search is True
def test_topic_min_length(self):
"""Test that topic requires at least 1 character."""
with pytest.raises(ValueError):
WikiSmartCreateRequest(topic="")
def test_topic_max_length(self):
"""Test that topic is limited to 500 characters."""
long_topic = "x" * 501
with pytest.raises(ValueError):
WikiSmartCreateRequest(topic=long_topic)
def test_path_validation_adds_leading_slash(self):
"""Test that path without leading slash gets one added."""
request = WikiSmartCreateRequest(
topic="Test",
path="technology/test"
)
assert request.path == "/technology/test"
def test_path_validation_removes_trailing_slash(self):
"""Test that trailing slash is removed."""
request = WikiSmartCreateRequest(
topic="Test",
path="/technology/test/"
)
assert request.path == "/technology/test"
def test_tags_deduplication(self):
"""Test that duplicate tags are removed."""
request = WikiSmartCreateRequest(
topic="Test",
tags=["devops", "devops", "containers", "devops"]
)
assert len(request.tags) == 2
assert "devops" in request.tags
assert "containers" in request.tags
def test_tags_whitespace_cleanup(self):
"""Test that tag whitespace is cleaned."""
request = WikiSmartCreateRequest(
topic="Test",
tags=[" devops ", "containers", " ", ""]
)
assert "devops" in request.tags
assert "containers" in request.tags
assert "" not in request.tags
assert " " not in request.tags
class TestWikiSmartCreateResponse:
"""Tests for WikiSmartCreateResponse model."""
def test_response_structure(self):
"""Test response model with all fields."""
page = WikiPage(
id=123,
path="/users/test/technology/docker",
title="Docker",
content="# Docker\n\nContent here",
tags=["technology"],
is_published=True,
created_at="2024-01-15T10:00:00Z",
updated_at="2024-01-15T10:00:00Z"
)
response = WikiSmartCreateResponse(
page=page,
research_summary={
"wiki_results": 3,
"web_results": 5,
"graph_entities": 2
},
sources_used=8,
search_id="test-uuid-123",
entity_linking={
"forward_links": 4,
"backward_links": 2,
"pages_updated": 1
}
)
assert response.page.id == 123
assert response.sources_used == 8
assert response.research_summary["wiki_results"] == 3
assert response.entity_linking["forward_links"] == 4
def test_response_default_entity_linking(self):
"""Test that entity_linking defaults to empty dict."""
page = WikiPage(
id=1,
path="/test",
title="Test",
content="Content",
tags=[],
is_published=True,
created_at="2024-01-15T10:00:00Z",
updated_at="2024-01-15T10:00:00Z"
)
response = WikiSmartCreateResponse(
page=page,
research_summary={},
sources_used=0
)
assert response.entity_linking == {}
assert response.search_id is None
# =============================================================================
# WikiService.smart_create_page Tests
# =============================================================================
class TestSmartCreatePage:
"""Tests for WikiService.smart_create_page method."""
@pytest.fixture
def mock_hybrid_rag_service(self):
"""Mock HybridRAG service."""
service = AsyncMock()
# Create mock response
mock_response = MagicMock()
mock_response.total_results = 5
mock_response.search_id = "search-123"
mock_response.results = [
MagicMock(
source_type="vector",
title="Existing Docker Page",
url=None,
page_path="users/testuser/docker-basics",
content="Docker is a containerization platform...",
related_dossiers=["containers"]
),
MagicMock(
source_type="web",
title="Docker Documentation",
url="https://docs.docker.com",
page_path=None,
content="Official Docker documentation...",
related_dossiers=None
)
]
mock_response.keywords = MagicMock()
mock_response.keywords.core_keywords = ["docker", "containers", "virtualization"]
mock_response.timing = MagicMock()
mock_response.timing.total_ms = 1500
service.search = AsyncMock(return_value=mock_response)
return service
@pytest.fixture
def mock_wiki_page_writer(self):
"""Mock WikiPageWriter."""
writer = AsyncMock()
writer.create_page = AsyncMock(return_value="# Docker Containers\n\n## Overview\n\nGenerated content about Docker...")
return writer
@pytest.mark.asyncio
async def test_smart_create_basic(
self,
mock_hybrid_rag_service,
mock_wiki_page_writer
):
"""Test basic smart page creation flow."""
from src.services.wiki_service import WikiService
mock_wiki_client = AsyncMock()
wiki_service = WikiService(mock_wiki_client)
# Mock the create_page method on the service itself
mock_page = WikiPage(
id=42,
path="/users/testuser/technology/docker",
title="Docker containers",
content="# Docker\n\nGenerated content",
tags=["technology"],
is_published=True,
created_at="2024-01-15T10:00:00Z",
updated_at="2024-01-15T10:00:00Z"
)
with patch.object(wiki_service, 'create_page', new_callable=AsyncMock) as mock_create:
mock_create.return_value = mock_page
page, research_data = await wiki_service.smart_create_page(
topic="Docker containers",
user="testuser",
path="/technology/docker",
tags=["technology"],
hybrid_rag_service=mock_hybrid_rag_service,
wiki_page_writer=mock_wiki_page_writer,
include_web=True,
include_wiki=True
)
# Verify HybridRAG was called
mock_hybrid_rag_service.search.assert_called_once()
# Verify WikiPageWriter was called
mock_wiki_page_writer.create_page.assert_called_once()
# Verify page was created
mock_create.assert_called_once()
# Verify research data
assert "research_summary" in research_data
assert "sources_used" in research_data
assert "search_id" in research_data
assert research_data["search_id"] == "search-123"
@pytest.mark.asyncio
async def test_smart_create_auto_generates_path(
self,
mock_hybrid_rag_service,
mock_wiki_page_writer
):
"""Test that path is auto-generated from topic when not provided."""
from src.services.wiki_service import WikiService
mock_wiki_client = AsyncMock()
wiki_service = WikiService(mock_wiki_client)
mock_page = WikiPage(
id=42,
path="/users/testuser/tutorials/docker-compose-tutorial",
title="Docker Compose Tutorial",
content="# Docker Compose\n\nContent",
tags=["tutorials"],
is_published=True,
created_at="2024-01-15T10:00:00Z",
updated_at="2024-01-15T10:00:00Z"
)
with patch.object(wiki_service, 'create_page', new_callable=AsyncMock) as mock_create:
mock_create.return_value = mock_page
await wiki_service.smart_create_page(
topic="Docker Compose Tutorial",
user="testuser",
path=None, # No path provided
tags=["tutorials"],
hybrid_rag_service=mock_hybrid_rag_service,
wiki_page_writer=mock_wiki_page_writer
)
# Check that create_page was called
mock_create.assert_called_once()
# The WikiPageCreate passed should have auto-generated path
call_args = mock_create.call_args[0][0] # First positional arg
assert "docker-compose-tutorial" in call_args.path.lower()
@pytest.mark.asyncio
async def test_smart_create_respects_web_flag(
self,
mock_hybrid_rag_service,
mock_wiki_page_writer
):
"""Test that include_web flag is passed to HybridRAG."""
from src.services.wiki_service import WikiService
mock_wiki_client = AsyncMock()
wiki_service = WikiService(mock_wiki_client)
mock_page = WikiPage(
id=1,
path="/test",
title="Test",
content="Content",
tags=[],
is_published=True,
created_at="2024-01-15T10:00:00Z",
updated_at="2024-01-15T10:00:00Z"
)
with patch.object(wiki_service, 'create_page', new_callable=AsyncMock) as mock_create:
mock_create.return_value = mock_page
await wiki_service.smart_create_page(
topic="Test",
user="testuser",
path="/test",
tags=[],
hybrid_rag_service=mock_hybrid_rag_service,
wiki_page_writer=mock_wiki_page_writer,
include_web=False,
include_wiki=True
)
# Check HybridRAG config
call_args = mock_hybrid_rag_service.search.call_args
config = call_args[1]["config"]
assert config.enable_web is False
assert config.enable_vector is True
@pytest.mark.asyncio
async def test_smart_create_counts_sources(
self,
mock_hybrid_rag_service,
mock_wiki_page_writer
):
"""Test that sources are counted correctly."""
from src.services.wiki_service import WikiService
mock_wiki_client = AsyncMock()
wiki_service = WikiService(mock_wiki_client)
mock_page = WikiPage(
id=1,
path="/test",
title="Test",
content="Content",
tags=[],
is_published=True,
created_at="2024-01-15T10:00:00Z",
updated_at="2024-01-15T10:00:00Z"
)
with patch.object(wiki_service, 'create_page', new_callable=AsyncMock) as mock_create:
mock_create.return_value = mock_page
page, research_data = await wiki_service.smart_create_page(
topic="Test",
user="testuser",
path="/test",
tags=[],
hybrid_rag_service=mock_hybrid_rag_service,
wiki_page_writer=mock_wiki_page_writer
)
# Should have 2 sources (1 wiki + 1 web from mock)
assert research_data["sources_used"] == 2
assert research_data["research_summary"]["wiki_results"] == 1
assert research_data["research_summary"]["web_results"] == 1
# =============================================================================
# Entity Linking Utils Tests
# =============================================================================
class TestBidirectionalEntityLinking:
"""Tests for entity_linking_utils.apply_bidirectional_entity_linking."""
@pytest.fixture
def mock_neo4j_client(self):
"""Mock Neo4j client."""
client = AsyncMock()
client.execute_query = AsyncMock(return_value=[])
return client
@pytest.fixture
def mock_wiki_service(self):
"""Mock WikiService."""
service = AsyncMock()
return service
@pytest.fixture
def mock_ingestion_service(self):
"""Mock IngestionService."""
service = AsyncMock()
return service
@pytest.mark.asyncio
async def test_returns_link_counts(
self,
mock_neo4j_client,
mock_wiki_service,
mock_ingestion_service
):
"""Test that function returns proper link count structure."""
from src.services.entity_linking_utils import apply_bidirectional_entity_linking
# Patch at the import location within the module
with patch('src.routers.entity_linking.link_entities_in_page') as mock_link:
mock_result = MagicMock()
mock_result.content_links_added = 3
mock_link.return_value = mock_result
with patch('src.core.dependencies.get_graph_service'):
with patch('src.core.dependencies.get_ingestion_service', return_value=mock_ingestion_service):
result = await apply_bidirectional_entity_linking(
page_id=42,
page_title="Docker",
user="testuser",
neo4j_client=mock_neo4j_client,
wiki_service=mock_wiki_service,
ingestion_service=mock_ingestion_service
)
assert "forward_links" in result
assert "backward_links" in result
assert "pages_updated" in result
@pytest.mark.asyncio
async def test_handles_no_reverse_references(
self,
mock_neo4j_client,
mock_wiki_service,
mock_ingestion_service
):
"""Test graceful handling when no reverse references found."""
from src.services.entity_linking_utils import apply_bidirectional_entity_linking
# No reverse references
mock_neo4j_client.execute_query = AsyncMock(return_value=[])
with patch('src.routers.entity_linking.link_entities_in_page') as mock_link:
mock_result = MagicMock()
mock_result.content_links_added = 2
mock_link.return_value = mock_result
with patch('src.core.dependencies.get_graph_service'):
with patch('src.core.dependencies.get_ingestion_service', return_value=mock_ingestion_service):
result = await apply_bidirectional_entity_linking(
page_id=42,
page_title="NewEntity",
user="testuser",
neo4j_client=mock_neo4j_client,
wiki_service=mock_wiki_service,
ingestion_service=mock_ingestion_service
)
assert result["backward_links"] == 0
assert result["pages_updated"] == 0
@pytest.mark.asyncio
async def test_handles_errors_gracefully(
self,
mock_neo4j_client,
mock_wiki_service,
mock_ingestion_service
):
"""Test that errors don't crash the function."""
from src.services.entity_linking_utils import apply_bidirectional_entity_linking
with patch('src.routers.entity_linking.link_entities_in_page') as mock_link:
mock_link.side_effect = Exception("Test error")
with patch('src.core.dependencies.get_graph_service'):
with patch('src.core.dependencies.get_ingestion_service', return_value=mock_ingestion_service):
result = await apply_bidirectional_entity_linking(
page_id=42,
page_title="Test",
user="testuser",
neo4j_client=mock_neo4j_client,
wiki_service=mock_wiki_service,
ingestion_service=mock_ingestion_service
)
# Should return zeros, not raise
assert result["forward_links"] == 0
assert result["backward_links"] == 0
assert result["pages_updated"] == 0
# =============================================================================
# Endpoint Tests
# =============================================================================
class TestSmartCreateEndpoint:
"""Tests for POST /wiki/pages/smart-create endpoint."""
@pytest.fixture
def mock_clients(self):
"""Create all mock clients needed for the endpoint."""
return {
"wiki_client": AsyncMock(),
"neo4j_client": AsyncMock(),
"qdrant_client": MagicMock(),
"ollama_client": AsyncMock(),
"searxng_client": AsyncMock()
}
@pytest.mark.asyncio
async def test_endpoint_returns_201(self, mock_clients):
"""Test that successful creation returns 201 status."""
from fastapi.testclient import TestClient
from unittest.mock import patch
# This test would require more setup with FastAPI TestClient
# For now, we test the model validation
request = WikiSmartCreateRequest(
topic="Test Topic",
tags=["test"]
)
assert request.topic == "Test Topic"
def test_request_validation_rejects_empty_topic(self):
"""Test that empty topic is rejected."""
with pytest.raises(ValueError):
WikiSmartCreateRequest(topic="")
def test_request_accepts_minimal_input(self):
"""Test that only topic is required."""
request = WikiSmartCreateRequest(topic="Minimal test")
assert request.topic == "Minimal test"
assert request.include_web_research is True # default
assert request.include_wiki_search is True # default