refactor: extract library-desk to standalone repository
- Move library-desk service to git.schweitz.net/jpmschweitzer/library-desk - Update stacks/library-desk.yml to use container image - Add library-desk to External Repositories in CONTAINERS.md - Remove source code (now in external repo) 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
This commit is contained in:
@@ -51,6 +51,7 @@ Services marked "(external)" have source code in separate Gitea repositories. Co
|
||||
|---------|------------|-------|
|
||||
| **Scheduler** | [scheduler](https://git.schweitz.net/jpmschweitzer/scheduler) | `git.schweitz.net/jpmschweitzer/scheduler:latest` |
|
||||
| **Core API** | [core-api](https://git.schweitz.net/jpmschweitzer/core-api) | `git.schweitz.net/jpmschweitzer/core-api:latest` |
|
||||
| **Library Desk** | [library-desk](https://git.schweitz.net/jpmschweitzer/library-desk) | `git.schweitz.net/jpmschweitzer/library-desk:latest` |
|
||||
|
||||
**Development workflow:** Clone repo → make changes → create Gitea release → Watchtower auto-updates container.
|
||||
|
||||
|
||||
@@ -1,473 +0,0 @@
|
||||
# Librarian Integration Guide
|
||||
|
||||
**How The Scheduler (Librarian) interacts with Library Desk for documentation management**
|
||||
|
||||
## Overview
|
||||
|
||||
The Scheduler's documentation mirroring tasks feed content into The Library system via Library Desk API. This creates a knowledge graph and vector index of all documentation for semantic search and relationship discovery.
|
||||
|
||||
## Architecture Flow
|
||||
|
||||
```
|
||||
┌─────────────────┐
|
||||
│ The Scheduler │ (The Librarian)
|
||||
│ (scheduler) │
|
||||
└────────┬────────┘
|
||||
│ 1. Mirror docs from sources
|
||||
│ (GitHub, Gitea, etc.)
|
||||
↓
|
||||
┌─────────────────┐
|
||||
│ Gitea Repo │
|
||||
│ docs-mirror/* │
|
||||
└────────┬────────┘
|
||||
│ 2. Ingest to Library
|
||||
↓
|
||||
┌─────────────────┐
|
||||
│ Library Desk │ (Coordination API)
|
||||
│ (library-desk) │
|
||||
└────────┬────────┘
|
||||
│ 3. Process & Index
|
||||
├──→ Neo4j (relationships)
|
||||
├──→ Qdrant (embeddings)
|
||||
└──→ Wiki.js (dossiers)
|
||||
```
|
||||
|
||||
## Required Endpoints
|
||||
|
||||
### 1. Document Ingestion
|
||||
|
||||
**POST /ingest/document**
|
||||
```json
|
||||
{
|
||||
"source": "github",
|
||||
"repository": "anthropics/anthropic-cookbook",
|
||||
"path": "skills/citation/guide.md",
|
||||
"content": "# Citation Guide\n...",
|
||||
"metadata": {
|
||||
"commit_sha": "abc123",
|
||||
"author": "Anthropic",
|
||||
"updated_at": "2025-12-08T10:30:00Z",
|
||||
"gitea_mirror_path": "docs-mirror/anthropic-cookbook/skills/citation/guide.md",
|
||||
"language": "markdown",
|
||||
"tags": ["skills", "citation", "prompting"]
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**Response:**
|
||||
```json
|
||||
{
|
||||
"document_id": "doc_abc123xyz",
|
||||
"status": "processing",
|
||||
"operations": {
|
||||
"chunking": "pending",
|
||||
"embedding": "pending",
|
||||
"entity_extraction": "pending",
|
||||
"graph_indexing": "pending"
|
||||
},
|
||||
"estimated_completion": "2025-12-08T10:30:15Z"
|
||||
}
|
||||
```
|
||||
|
||||
### 2. Batch Ingestion
|
||||
|
||||
**POST /ingest/batch**
|
||||
```json
|
||||
{
|
||||
"source": "github",
|
||||
"repository": "anthropics/anthropic-cookbook",
|
||||
"documents": [
|
||||
{
|
||||
"path": "skills/citation/guide.md",
|
||||
"content": "...",
|
||||
"metadata": {...}
|
||||
},
|
||||
{
|
||||
"path": "skills/summarization/techniques.md",
|
||||
"content": "...",
|
||||
"metadata": {...}
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
**Response:**
|
||||
```json
|
||||
{
|
||||
"batch_id": "batch_xyz789",
|
||||
"total_documents": 2,
|
||||
"status": "processing",
|
||||
"documents": [
|
||||
{"document_id": "doc_1", "status": "queued"},
|
||||
{"document_id": "doc_2", "status": "queued"}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
### 3. Document Status Check
|
||||
|
||||
**GET /ingest/status/{document_id}**
|
||||
|
||||
**Response:**
|
||||
```json
|
||||
{
|
||||
"document_id": "doc_abc123xyz",
|
||||
"status": "completed",
|
||||
"operations": {
|
||||
"chunking": "completed",
|
||||
"embedding": "completed",
|
||||
"entity_extraction": "completed",
|
||||
"graph_indexing": "completed"
|
||||
},
|
||||
"results": {
|
||||
"chunks_created": 12,
|
||||
"vectors_indexed": 12,
|
||||
"entities_extracted": 8,
|
||||
"relationships_created": 15
|
||||
},
|
||||
"completed_at": "2025-12-08T10:30:14Z"
|
||||
}
|
||||
```
|
||||
|
||||
### 4. Content Update Detection
|
||||
|
||||
**POST /ingest/check-updates**
|
||||
```json
|
||||
{
|
||||
"documents": [
|
||||
{
|
||||
"path": "docs-mirror/anthropic-cookbook/skills/citation/guide.md",
|
||||
"content_hash": "sha256:abc123...",
|
||||
"updated_at": "2025-12-08T10:30:00Z"
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
**Response:**
|
||||
```json
|
||||
{
|
||||
"updates_needed": [
|
||||
{
|
||||
"path": "docs-mirror/anthropic-cookbook/skills/citation/guide.md",
|
||||
"reason": "content_changed",
|
||||
"last_indexed": "2025-12-07T10:30:00Z",
|
||||
"action": "re-index"
|
||||
}
|
||||
],
|
||||
"up_to_date": [],
|
||||
"new_documents": []
|
||||
}
|
||||
```
|
||||
|
||||
### 5. Deduplication Check
|
||||
|
||||
**POST /deduplicate/check**
|
||||
```json
|
||||
{
|
||||
"document_id": "doc_abc123xyz",
|
||||
"similarity_threshold": 0.85
|
||||
}
|
||||
```
|
||||
|
||||
**Response:**
|
||||
```json
|
||||
{
|
||||
"duplicates": [
|
||||
{
|
||||
"document_id": "doc_def456",
|
||||
"similarity": 0.92,
|
||||
"path": "docs-mirror/claude-docs/citation-best-practices.md",
|
||||
"overlap_summary": "Both documents cover citation formatting"
|
||||
}
|
||||
],
|
||||
"suggestions": {
|
||||
"action": "merge_or_cross_reference",
|
||||
"confidence": 0.88
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### 6. Repository Sync Status
|
||||
|
||||
**GET /ingest/repo-status/{repository_name}**
|
||||
|
||||
**Response:**
|
||||
```json
|
||||
{
|
||||
"repository": "anthropic-cookbook",
|
||||
"total_documents": 156,
|
||||
"indexed_documents": 156,
|
||||
"pending_updates": 0,
|
||||
"failed_documents": 0,
|
||||
"last_sync": "2025-12-08T03:00:00Z",
|
||||
"next_scheduled_sync": "2025-12-09T03:00:00Z"
|
||||
}
|
||||
```
|
||||
|
||||
## Scheduler Integration Workflow
|
||||
|
||||
### Phase 1: Mirror Documentation (Current)
|
||||
```python
|
||||
# This already exists in the Scheduler
|
||||
async def mirror_documentation():
|
||||
"""Mirror docs from external sources to Gitea"""
|
||||
repos = [
|
||||
"anthropics/anthropic-cookbook",
|
||||
"anthropics/prompt-eng-interactive-tutorial",
|
||||
# ... etc
|
||||
]
|
||||
|
||||
for repo in repos:
|
||||
# Clone/pull to /docs-mirror/{repo-name}
|
||||
await git_sync(repo, f"/docs-mirror/{repo}")
|
||||
```
|
||||
|
||||
### Phase 2: Index to Library (New)
|
||||
```python
|
||||
async def index_to_library():
|
||||
"""Send mirrored docs to Library Desk for indexing"""
|
||||
|
||||
# Get list of documents in docs-mirror
|
||||
docs_path = Path("/docs-mirror")
|
||||
|
||||
for repo_dir in docs_path.iterdir():
|
||||
if not repo_dir.is_dir():
|
||||
continue
|
||||
|
||||
# Find all markdown files
|
||||
md_files = list(repo_dir.rglob("*.md"))
|
||||
|
||||
# Check what needs updating
|
||||
update_check = await check_library_updates(md_files)
|
||||
|
||||
if update_check["updates_needed"]:
|
||||
# Batch ingest updated documents
|
||||
await batch_ingest_documents(
|
||||
repository=repo_dir.name,
|
||||
documents=update_check["updates_needed"]
|
||||
)
|
||||
|
||||
# Wait for processing to complete
|
||||
await wait_for_batch_completion(batch_id)
|
||||
|
||||
# Check for duplicates
|
||||
await check_and_resolve_duplicates(repo_dir.name)
|
||||
```
|
||||
|
||||
### Phase 3: Monitor & Maintain
|
||||
```python
|
||||
async def maintain_library_index():
|
||||
"""Periodic maintenance of Library index"""
|
||||
|
||||
# Check for orphaned entries (deleted from source)
|
||||
await cleanup_orphaned_documents()
|
||||
|
||||
# Update embeddings if model changed
|
||||
await refresh_embeddings_if_needed()
|
||||
|
||||
# Generate relationship maps for new content
|
||||
await discover_document_relationships()
|
||||
```
|
||||
|
||||
## Scheduler Task Definition
|
||||
|
||||
**New Task: `library_sync`**
|
||||
|
||||
```yaml
|
||||
Task Name: library_sync
|
||||
Description: Sync mirrored documentation to Library for indexing and search
|
||||
Schedule: Daily at 03:30 (after doc mirroring at 03:00)
|
||||
Priority: 15 (user maintenance)
|
||||
Service: library
|
||||
Executor: scheduler.tasks.library_tasks.sync_library_index
|
||||
Dependencies:
|
||||
- docs_mirror (must complete first)
|
||||
Configuration:
|
||||
- LIBRARY_DESK_URL: http://library-desk:8089
|
||||
- LIBRARY_API_KEY: ${LIBRARY_API_KEY}
|
||||
- BATCH_SIZE: 50
|
||||
- CHECK_UPDATES_ONLY: true
|
||||
Outputs:
|
||||
- Documents indexed
|
||||
- Duplicates found
|
||||
- Relationships created
|
||||
```
|
||||
|
||||
## API Client Example
|
||||
|
||||
```python
|
||||
# scheduler/src/clients/library_desk.py
|
||||
|
||||
import httpx
|
||||
from typing import List, Dict, Any
|
||||
from pathlib import Path
|
||||
|
||||
class LibraryDeskClient:
|
||||
def __init__(self, base_url: str, api_key: str):
|
||||
self.base_url = base_url
|
||||
self.headers = {
|
||||
"Authorization": f"Bearer {api_key}",
|
||||
"Content-Type": "application/json"
|
||||
}
|
||||
|
||||
async def ingest_document(
|
||||
self,
|
||||
source: str,
|
||||
repository: str,
|
||||
path: str,
|
||||
content: str,
|
||||
metadata: Dict[str, Any]
|
||||
) -> Dict[str, Any]:
|
||||
"""Ingest a single document"""
|
||||
async with httpx.AsyncClient() as client:
|
||||
response = await client.post(
|
||||
f"{self.base_url}/ingest/document",
|
||||
headers=self.headers,
|
||||
json={
|
||||
"source": source,
|
||||
"repository": repository,
|
||||
"path": path,
|
||||
"content": content,
|
||||
"metadata": metadata
|
||||
}
|
||||
)
|
||||
response.raise_for_status()
|
||||
return response.json()
|
||||
|
||||
async def batch_ingest(
|
||||
self,
|
||||
source: str,
|
||||
repository: str,
|
||||
documents: List[Dict[str, Any]]
|
||||
) -> Dict[str, Any]:
|
||||
"""Ingest multiple documents"""
|
||||
async with httpx.AsyncClient() as client:
|
||||
response = await client.post(
|
||||
f"{self.base_url}/ingest/batch",
|
||||
headers=self.headers,
|
||||
json={
|
||||
"source": source,
|
||||
"repository": repository,
|
||||
"documents": documents
|
||||
},
|
||||
timeout=300.0 # 5 minutes for large batches
|
||||
)
|
||||
response.raise_for_status()
|
||||
return response.json()
|
||||
|
||||
async def check_updates(
|
||||
self,
|
||||
documents: List[Dict[str, str]]
|
||||
) -> Dict[str, Any]:
|
||||
"""Check which documents need updating"""
|
||||
async with httpx.AsyncClient() as client:
|
||||
response = await client.post(
|
||||
f"{self.base_url}/ingest/check-updates",
|
||||
headers=self.headers,
|
||||
json={"documents": documents}
|
||||
)
|
||||
response.raise_for_status()
|
||||
return response.json()
|
||||
|
||||
async def get_repo_status(self, repository: str) -> Dict[str, Any]:
|
||||
"""Get indexing status for a repository"""
|
||||
async with httpx.AsyncClient() as client:
|
||||
response = await client.get(
|
||||
f"{self.base_url}/ingest/repo-status/{repository}",
|
||||
headers=self.headers
|
||||
)
|
||||
response.raise_for_status()
|
||||
return response.json()
|
||||
```
|
||||
|
||||
## Typical Workflow Sequence
|
||||
|
||||
### Daily Documentation Sync (03:00-03:45)
|
||||
|
||||
1. **03:00** - Scheduler runs `docs_mirror` task
|
||||
- Pulls latest from all configured repos
|
||||
- Writes to `/docs-mirror/*`
|
||||
|
||||
2. **03:30** - Scheduler runs `library_sync` task
|
||||
- Scans `/docs-mirror/` for changes
|
||||
- Calls `POST /ingest/check-updates` with file hashes
|
||||
- Gets list of updated/new documents
|
||||
|
||||
3. **03:31-03:40** - Batch ingestion
|
||||
- Groups documents by repo
|
||||
- Calls `POST /ingest/batch` for each repo
|
||||
- Monitors `GET /ingest/status/{batch_id}`
|
||||
|
||||
4. **03:41-03:44** - Post-processing
|
||||
- Calls `POST /deduplicate/check` for new docs
|
||||
- Reviews duplicate suggestions
|
||||
- Logs statistics to Scheduler database
|
||||
|
||||
5. **03:45** - Completion
|
||||
- Scheduler marks task complete
|
||||
- Sends summary to logs
|
||||
- Updates next run time
|
||||
|
||||
## Error Handling
|
||||
|
||||
### Retry Strategy
|
||||
```python
|
||||
async def ingest_with_retry(document: Dict, max_retries: int = 3):
|
||||
"""Ingest with exponential backoff"""
|
||||
for attempt in range(max_retries):
|
||||
try:
|
||||
result = await library_client.ingest_document(**document)
|
||||
return result
|
||||
except httpx.TimeoutException:
|
||||
if attempt < max_retries - 1:
|
||||
await asyncio.sleep(2 ** attempt) # 1s, 2s, 4s
|
||||
else:
|
||||
# Log failure and continue
|
||||
logger.error(f"Failed to ingest {document['path']} after {max_retries} attempts")
|
||||
return None
|
||||
```
|
||||
|
||||
### Graceful Degradation
|
||||
- If Library Desk is down, queue documents for later ingestion
|
||||
- Store failed ingestions in Scheduler database
|
||||
- Retry failed ingestions on next run
|
||||
|
||||
## Metrics to Track
|
||||
|
||||
The Scheduler should track:
|
||||
- Documents mirrored vs. documents indexed
|
||||
- Average ingestion time per document
|
||||
- Deduplication rate
|
||||
- Failed ingestions
|
||||
- Library Desk response times
|
||||
|
||||
These can be displayed in the Scheduler UI dashboard.
|
||||
|
||||
## Environment Variables
|
||||
|
||||
Add to Scheduler's environment:
|
||||
```bash
|
||||
# Library Integration
|
||||
LIBRARY_DESK_URL=http://library-desk:8089
|
||||
LIBRARY_API_KEY=${LIBRARY_API_KEY}
|
||||
LIBRARY_BATCH_SIZE=50
|
||||
LIBRARY_SYNC_ENABLED=true
|
||||
```
|
||||
|
||||
## Next Steps
|
||||
|
||||
1. Implement ingestion endpoints in Library Desk
|
||||
2. Add LibraryDeskClient to Scheduler
|
||||
3. Create `library_sync` task in Scheduler
|
||||
4. Test with small batch of documents
|
||||
5. Monitor and tune performance
|
||||
6. Expand to full documentation corpus
|
||||
|
||||
## Benefits
|
||||
|
||||
- **Automatic Knowledge Base**: All mirrored docs automatically indexed
|
||||
- **Semantic Search**: Find docs by meaning, not just keywords
|
||||
- **Relationship Discovery**: Understand connections between docs
|
||||
- **Deduplication**: Identify overlapping content across repos
|
||||
- **HybridRAG Ready**: Knowledge graph + vectors enable advanced AI queries
|
||||
@@ -1,236 +0,0 @@
|
||||
# Library Desk - API Coordination Service
|
||||
|
||||
**FastAPI service that coordinates all Library operations**
|
||||
|
||||
## Overview
|
||||
|
||||
Library Desk is the central coordination layer for The Library system, providing a unified API for:
|
||||
|
||||
- **HybridRAG Queries** - Combines Neo4j (structure) + Qdrant (semantics) + SearXNG (web)
|
||||
- **Document Ingestion** - Parse, chunk, embed, and index documents
|
||||
- **Entity Extraction** - NLP to identify classes, functions, concepts
|
||||
- **Relationship Mapping** - Link entities in Neo4j knowledge graph
|
||||
- **Mind Map Generation** - Query Neo4j graph → Render D3.js visualizations
|
||||
- **Wiki.js Proxy** - CRUD operations for dossiers
|
||||
- **Deduplication** - Vector similarity + graph analysis
|
||||
|
||||
## Architecture
|
||||
|
||||
```
|
||||
Library Desk API (FastAPI)
|
||||
├── Neo4j (knowledge graph)
|
||||
├── Qdrant (vector search)
|
||||
├── Wiki.js (wiki operations)
|
||||
├── SearXNG (web search)
|
||||
├── Ollama (embeddings)
|
||||
└── Redis (caching)
|
||||
```
|
||||
|
||||
## Requirements
|
||||
|
||||
- **Python**: 3.12+
|
||||
- **Dependencies**: See `requirements.txt`
|
||||
|
||||
## Configuration
|
||||
|
||||
Environment variables (set in Portainer stack):
|
||||
|
||||
```bash
|
||||
# Required
|
||||
LIBRARY_API_KEY=<generate-with-openssl-rand-hex-32>
|
||||
NEO4J_PASSWORD=<neo4j-password>
|
||||
WIKIJS_API_KEY=<from-wiki-admin-panel>
|
||||
|
||||
# Optional (defaults provided)
|
||||
NEO4J_URI=bolt://neo4j:7687
|
||||
NEO4J_USER=neo4j
|
||||
QDRANT_HOST=qdrant
|
||||
QDRANT_PORT=6333
|
||||
WIKIJS_URL=http://wiki:3000
|
||||
SEARXNG_URL=http://searxng:8080
|
||||
OLLAMA_URL=http://ollama:11434
|
||||
OLLAMA_MODEL=nomic-embed-text
|
||||
REDIS_HOST=redis-shared
|
||||
REDIS_PORT=6379
|
||||
REDIS_DB=2
|
||||
```
|
||||
|
||||
## API Endpoints
|
||||
|
||||
### System
|
||||
|
||||
- `GET /` - Root endpoint
|
||||
- `GET /health` - Health check (public)
|
||||
- `GET /stats` - System statistics (authenticated)
|
||||
|
||||
### Query (All require API key)
|
||||
|
||||
- `POST /query/hybrid` - HybridRAG (graph + vector + web)
|
||||
- `POST /query/semantic` - Vector search only
|
||||
- `POST /query/graph` - Graph traversal only
|
||||
- `GET /query/related/{id}` - Find related content
|
||||
|
||||
### Content Management (Future)
|
||||
|
||||
- `POST /ingest/document` - Index new document
|
||||
- `POST /ingest/wiki-page` - Sync Wiki.js page
|
||||
- `POST /wiki/dossier` - Create dossier (proxies to Wiki.js)
|
||||
- `PUT /wiki/dossier/{id}` - Update dossier
|
||||
- `DELETE /wiki/dossier/{id}` - Delete dossier
|
||||
|
||||
### Graph Operations (Future)
|
||||
|
||||
- `GET /graph/entities` - List entities
|
||||
- `GET /graph/mindmap/{id}` - Generate mind map
|
||||
- `POST /graph/query` - Execute Cypher query
|
||||
|
||||
### Deduplication (Future)
|
||||
|
||||
- `POST /deduplicate/find` - Find duplicates
|
||||
- `POST /deduplicate/merge` - Merge duplicates
|
||||
|
||||
## Development
|
||||
|
||||
### Local Setup
|
||||
|
||||
```bash
|
||||
# Create virtual environment
|
||||
python3 -m venv venv
|
||||
source venv/bin/activate
|
||||
|
||||
# Install dependencies
|
||||
pip install -r requirements.txt
|
||||
|
||||
# Run development server
|
||||
uvicorn src.main:app --reload --host 0.0.0.0 --port 8089
|
||||
```
|
||||
|
||||
### Project Structure
|
||||
|
||||
Following [FastAPI Best Practices](https://github.com/zhanymkanov/fastapi-best-practices):
|
||||
|
||||
```
|
||||
library-desk/
|
||||
├── src/
|
||||
│ ├── __init__.py # Package initialization
|
||||
│ ├── main.py # FastAPI application
|
||||
│ ├── config.py # Pydantic settings
|
||||
│ └── [future modules] # Domain-specific modules
|
||||
├── requirements.txt # Python dependencies
|
||||
└── README.md # This file
|
||||
```
|
||||
|
||||
Future structure (as features are added):
|
||||
|
||||
```
|
||||
library-desk/
|
||||
├── src/
|
||||
│ ├── query/ # Query domain
|
||||
│ │ ├── router.py
|
||||
│ │ ├── schemas.py
|
||||
│ │ ├── service.py
|
||||
│ │ └── dependencies.py
|
||||
│ ├── graph/ # Graph domain
|
||||
│ ├── wiki/ # Wiki domain
|
||||
│ └── ingest/ # Ingestion domain
|
||||
```
|
||||
|
||||
## Best Practices Implemented
|
||||
|
||||
✅ **Async routes** for I/O operations
|
||||
✅ **Dependency injection** for configuration and auth
|
||||
✅ **Pydantic models** for request/response validation
|
||||
✅ **Modular settings** using Pydantic Settings
|
||||
✅ **API key authentication** with Bearer tokens
|
||||
✅ **OpenAPI documentation** auto-generated
|
||||
✅ **Proper logging** with structured format
|
||||
✅ **CORS middleware** configured
|
||||
✅ **Health checks** for monitoring
|
||||
✅ **Minor version locking** (`~=`) in requirements
|
||||
✅ **CVE-checked dependencies** (Dec 2025)
|
||||
|
||||
## API Documentation
|
||||
|
||||
Once running, access:
|
||||
|
||||
- **Interactive docs**: http://localhost:8089/docs
|
||||
- **ReDoc**: http://localhost:8089/redoc
|
||||
- **OpenAPI spec**: http://localhost:8089/openapi.json
|
||||
|
||||
## Authentication
|
||||
|
||||
All protected endpoints require a Bearer token:
|
||||
|
||||
```bash
|
||||
curl -H "Authorization: Bearer ${LIBRARY_API_KEY}" \
|
||||
http://localhost:8089/stats
|
||||
```
|
||||
|
||||
## Testing
|
||||
|
||||
```bash
|
||||
# Run tests (when implemented)
|
||||
pytest
|
||||
|
||||
# With coverage
|
||||
pytest --cov=src --cov-report=term
|
||||
```
|
||||
|
||||
## Deployment
|
||||
|
||||
Deployed via Portainer stack: `/stacks/library-desk.yml`
|
||||
|
||||
The container:
|
||||
- Runs on port 8089
|
||||
- Auto-creates venv on startup
|
||||
- Installs dependencies from requirements.txt
|
||||
- Starts uvicorn with 2 workers
|
||||
- Mounts source code for live editing
|
||||
|
||||
## Monitoring
|
||||
|
||||
- **Uptime Kuma**: Monitor `/health` endpoint
|
||||
- **Logs**: `docker logs library-desk`
|
||||
- **Stats**: `GET /stats` (requires API key)
|
||||
|
||||
## Scheduler Integration
|
||||
|
||||
See [LIBRARIAN_INTEGRATION.md](./LIBRARIAN_INTEGRATION.md) for details on how The Scheduler (Librarian) integrates with Library Desk for automated documentation indexing.
|
||||
|
||||
**Key Workflow:**
|
||||
1. Scheduler mirrors docs to Gitea (daily 03:00)
|
||||
2. Scheduler syncs to Library Desk (daily 03:30)
|
||||
3. Library Desk ingests, chunks, embeds, and indexes
|
||||
4. Content becomes searchable via HybridRAG
|
||||
|
||||
## Future Enhancements
|
||||
|
||||
- [ ] Implement HybridRAG query logic
|
||||
- [ ] Add Neo4j connection pooling
|
||||
- [ ] Add Qdrant client initialization
|
||||
- [ ] Implement Wiki.js API proxy
|
||||
- [ ] Add entity extraction (spaCy/NLP)
|
||||
- [ ] Implement mind map generation
|
||||
- [ ] Add deduplication logic
|
||||
- [ ] **Implement Scheduler integration endpoints** (see LIBRARIAN_INTEGRATION.md)
|
||||
- [ ] Add comprehensive tests
|
||||
- [ ] Add rate limiting
|
||||
- [ ] Add request tracing
|
||||
|
||||
## References
|
||||
|
||||
- [FastAPI Best Practices](https://github.com/zhanymkanov/fastapi-best-practices)
|
||||
- [FastAPI Documentation](https://fastapi.tiangolo.com/)
|
||||
- [Pydantic Documentation](https://docs.pydantic.dev/)
|
||||
- [Neo4j Python Driver](https://neo4j.com/docs/python-manual/current/)
|
||||
- [Qdrant Python Client](https://python-client.qdrant.tech/)
|
||||
|
||||
## License
|
||||
|
||||
Part of Portainer Core infrastructure.
|
||||
|
||||
## Support
|
||||
|
||||
- Check logs: `docker logs library-desk`
|
||||
- Health check: `curl http://localhost:8089/health`
|
||||
- API docs: http://localhost:8089/docs
|
||||
@@ -1,507 +0,0 @@
|
||||
# Wiki.js Change Detection - Complete Setup Guide
|
||||
|
||||
Complete setup instructions for real-time Wiki.js page change detection with loop prevention.
|
||||
|
||||
---
|
||||
|
||||
## Overview
|
||||
|
||||
This system automatically processes user-edited Wiki.js pages using PostgreSQL database triggers and NOTIFY/LISTEN.
|
||||
|
||||
**Features:**
|
||||
- ✅ Real-time change detection via database triggers
|
||||
- ✅ Read-only database user (security best practice)
|
||||
- ✅ Loop prevention (automated edits don't trigger re-processing)
|
||||
- ✅ Debouncing (prevents duplicate processing)
|
||||
- ✅ Graceful error handling
|
||||
|
||||
---
|
||||
|
||||
## Architecture
|
||||
|
||||
```
|
||||
┌────────────────────────────────────────────────┐
|
||||
│ User edits page in Wiki.js │
|
||||
└──────────────────┬─────────────────────────────┘
|
||||
│
|
||||
↓
|
||||
┌────────────────────────────────────────────────┐
|
||||
│ PostgreSQL Trigger fires │
|
||||
│ pg_notify('wiki_page_changes', 'UPDATE:123:...')│
|
||||
└──────────────────┬─────────────────────────────┘
|
||||
│
|
||||
↓
|
||||
┌────────────────────────────────────────────────┐
|
||||
│ WikiChangeListener (library-desk) │
|
||||
│ ├─ LOOP CHECK #1: Is author automated user? │ ← Prevents loops
|
||||
│ ├─ LOOP CHECK #2: Recently processed? │ ← Debouncing
|
||||
│ └─ Process page (ingest + entity linking) │
|
||||
└──────────────────┬─────────────────────────────┘
|
||||
│
|
||||
↓
|
||||
┌────────────────────────────────────────────────┐
|
||||
│ Entity linking writes back to Wiki.js │
|
||||
│ (via API as wikijs_username) │
|
||||
└──────────────────┬─────────────────────────────┘
|
||||
│
|
||||
↓
|
||||
┌────────────────────────────────────────────────┐
|
||||
│ Trigger fires again BUT... │
|
||||
│ LOOP CHECK #1 catches it: author is │
|
||||
│ automated user → SKIP processing │ ← Loop prevented!
|
||||
└────────────────────────────────────────────────┘
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Prerequisites
|
||||
|
||||
- Wiki.js running with PostgreSQL database
|
||||
- Library-desk has network access to PostgreSQL
|
||||
- PostgreSQL admin access (for initial setup)
|
||||
|
||||
---
|
||||
|
||||
## Step 1: Create Read-Only Database User
|
||||
|
||||
**Why:** Library-desk only needs to LISTEN for notifications, not modify data. Using a read-only user is a security best practice.
|
||||
|
||||
### Run SQL as PostgreSQL admin:
|
||||
|
||||
```bash
|
||||
docker exec -it postgres psql -U postgres -d wiki
|
||||
```
|
||||
|
||||
Then execute:
|
||||
|
||||
```sql
|
||||
-- Create read-only user for change listener
|
||||
CREATE USER library_desk_listener WITH PASSWORD 'secure_password_here';
|
||||
|
||||
-- Grant minimal permissions
|
||||
GRANT CONNECT ON DATABASE wiki TO library_desk_listener;
|
||||
GRANT USAGE ON SCHEMA public TO library_desk_listener;
|
||||
GRANT SELECT ON TABLE pages TO library_desk_listener;
|
||||
|
||||
-- Verify permissions
|
||||
\du library_desk_listener
|
||||
-- Should show: No roles, just basic login
|
||||
|
||||
-- Test LISTEN capability (important!)
|
||||
\c wiki library_desk_listener
|
||||
LISTEN wiki_page_changes;
|
||||
-- Should succeed even with read-only permissions
|
||||
\q
|
||||
```
|
||||
|
||||
**Save the password** - you'll need it for `.env` configuration.
|
||||
|
||||
---
|
||||
|
||||
## Step 2: Install Database Triggers
|
||||
|
||||
Triggers emit NOTIFY events when pages are created/updated/deleted.
|
||||
|
||||
### Run the setup script:
|
||||
|
||||
```bash
|
||||
docker exec -it library-desk bash
|
||||
cd /app
|
||||
python3 setup_wiki_triggers.py
|
||||
```
|
||||
|
||||
**Expected output:**
|
||||
```
|
||||
======================================================================
|
||||
Wiki.js Database Trigger Setup
|
||||
======================================================================
|
||||
|
||||
Connecting to Wiki.js database at postgres:5432
|
||||
Database: wiki
|
||||
User: wikijs
|
||||
|
||||
✓ Connected to Wiki.js database
|
||||
|
||||
Installing triggers...
|
||||
|
||||
✓ Triggers installed successfully:
|
||||
|
||||
- wiki_page_delete_trigger (DELETE)
|
||||
- wiki_page_insert_trigger (INSERT)
|
||||
- wiki_page_update_trigger (UPDATE)
|
||||
|
||||
======================================================================
|
||||
Setup complete!
|
||||
======================================================================
|
||||
```
|
||||
|
||||
### Manual verification:
|
||||
|
||||
```sql
|
||||
SELECT trigger_name, event_manipulation
|
||||
FROM information_schema.triggers
|
||||
WHERE event_object_table = 'pages'
|
||||
ORDER BY trigger_name;
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Step 3: Configure Environment Variables
|
||||
|
||||
Add to your `services/library-desk/.env`:
|
||||
|
||||
```bash
|
||||
# Wiki.js API User (used for entity linking)
|
||||
# IMPORTANT: This username will be filtered out to prevent loops
|
||||
WIKIJS_USERNAME=your_wikijs_api_username
|
||||
|
||||
# Wiki.js Database Configuration (for change listener)
|
||||
WIKIJS_DB_HOST=postgres
|
||||
WIKIJS_DB_PORT=5432
|
||||
WIKIJS_DB_NAME=wiki
|
||||
WIKIJS_DB_USER=library_desk_listener # Read-only user created in Step 1
|
||||
WIKIJS_DB_PASSWORD=secure_password_here # From Step 1
|
||||
|
||||
# Loop Prevention: Debounce duration (optional, default: 5 seconds)
|
||||
# Prevents processing duplicate notifications for the same page
|
||||
WIKIJS_CHANGE_LISTENER_DEBOUNCE_SECONDS=5
|
||||
```
|
||||
|
||||
**Critical:** The `WIKIJS_USERNAME` must match the Wiki.js user that library-desk uses for API calls. This prevents infinite loops when entity linking updates pages.
|
||||
|
||||
---
|
||||
|
||||
## Step 4: Restart Library-Desk
|
||||
|
||||
```bash
|
||||
docker restart library-desk
|
||||
```
|
||||
|
||||
### Check logs for successful startup:
|
||||
|
||||
```bash
|
||||
docker logs library-desk --tail 50 | grep -E "(Wiki.js change listener|wiki_page_changes)"
|
||||
```
|
||||
|
||||
**Expected log output:**
|
||||
```
|
||||
INFO - Starting Wiki.js database change listener
|
||||
INFO - Listening for Wiki.js page changes via PostgreSQL NOTIFY
|
||||
INFO - Wiki.js change listener started successfully
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Step 5: Test the System
|
||||
|
||||
### Test 1: Manual Page Edit
|
||||
|
||||
1. **Edit a page in Wiki.js** (add some text)
|
||||
2. **Save the page**
|
||||
3. **Watch library-desk logs:**
|
||||
|
||||
```bash
|
||||
docker logs library-desk --follow | grep -E "(Received|notification|Processing page|Entity linking)"
|
||||
```
|
||||
|
||||
**Expected log output:**
|
||||
```
|
||||
INFO - Received UPDATE notification for page 123 by user@example.com
|
||||
INFO - Processing page.update for page 123 ('Test Page')
|
||||
INFO - Ingesting page 123 into knowledge base
|
||||
INFO - Ingestion complete for page 123
|
||||
INFO - Applying bidirectional entity linking for page 123
|
||||
INFO - Entity linking complete for page 123: 5 forward links, 3 backward links
|
||||
INFO - Successfully processed page.update for page 123
|
||||
```
|
||||
|
||||
### Test 2: Loop Prevention
|
||||
|
||||
After the entity linking completes (which updates the Wiki.js page), you should see:
|
||||
|
||||
```
|
||||
DEBUG - Skipping notification for page 123 - automated edit by wikijs_username (likely entity linking)
|
||||
```
|
||||
|
||||
**This means loop prevention is working!** The automated edit triggered a notification, but the listener correctly filtered it out.
|
||||
|
||||
### Test 3: Debouncing
|
||||
|
||||
If you rapidly save the same page multiple times (within 5 seconds), you should see:
|
||||
|
||||
```
|
||||
INFO - Received UPDATE notification for page 123 by user@example.com
|
||||
INFO - Processing page.update for page 123...
|
||||
DEBUG - Skipping notification for page 123 - processed within last 5s (debouncing)
|
||||
DEBUG - Skipping notification for page 123 - processed within last 5s (debouncing)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Loop Prevention Details
|
||||
|
||||
The system has **two layers of loop prevention**:
|
||||
|
||||
### Layer 1: Automated User Filtering
|
||||
|
||||
**Problem:** Entity linking updates Wiki.js pages via API, which triggers database changes, which triggers notifications again.
|
||||
|
||||
**Solution:** Filter out notifications where the author matches `WIKIJS_USERNAME`:
|
||||
|
||||
```python
|
||||
# In wiki_change_listener.py:125-141
|
||||
def _is_automated_user(self, email: str) -> bool:
|
||||
automated_users = [
|
||||
self.settings.wikijs_username, # Your Wiki.js API user
|
||||
"library-desk@system",
|
||||
"automation@system"
|
||||
]
|
||||
return email.lower() in [u.lower() for u in automated_users]
|
||||
```
|
||||
|
||||
**Customize:** If you use different automation users, add them to the list.
|
||||
|
||||
### Layer 2: Debouncing
|
||||
|
||||
**Problem:** Multiple rapid notifications for the same page (network hiccups, Wiki.js behavior).
|
||||
|
||||
**Solution:** Track recently processed pages and ignore duplicates within N seconds:
|
||||
|
||||
```python
|
||||
# In wiki_change_listener.py:143-164
|
||||
def _is_recently_processed(self, page_id: int) -> bool:
|
||||
if page_id not in self._recent_notifications:
|
||||
return False
|
||||
|
||||
last_processed = self._recent_notifications[page_id]
|
||||
elapsed = (datetime.now() - last_processed).total_seconds()
|
||||
|
||||
return elapsed < self._debounce_seconds # Default: 5 seconds
|
||||
```
|
||||
|
||||
**Configure:** Set `WIKIJS_CHANGE_LISTENER_DEBOUNCE_SECONDS` in `.env` (1-60 seconds).
|
||||
|
||||
---
|
||||
|
||||
## Monitoring
|
||||
|
||||
### Check Listener Status
|
||||
|
||||
```bash
|
||||
# Is listener running?
|
||||
docker logs library-desk --tail 50 | grep "Wiki.js change listener"
|
||||
|
||||
# Expected: "Wiki.js change listener started successfully"
|
||||
```
|
||||
|
||||
### Monitor Real-Time Changes
|
||||
|
||||
```bash
|
||||
# Watch for page change notifications
|
||||
docker logs library-desk --follow | grep -E "(Received|Processing page|Entity linking|Skipping)"
|
||||
```
|
||||
|
||||
### Check Loop Prevention Stats
|
||||
|
||||
```bash
|
||||
# Count automated user filters (should be non-zero after entity linking runs)
|
||||
docker logs library-desk | grep "automated edit" | wc -l
|
||||
|
||||
# Count debouncing skips
|
||||
docker logs library-desk | grep "debouncing" | wc -l
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
### Listener Not Starting
|
||||
|
||||
**Error:** `Failed to start Wiki.js change listener`
|
||||
|
||||
**Solutions:**
|
||||
|
||||
1. **Check database connection:**
|
||||
```bash
|
||||
docker exec library-desk ping postgres
|
||||
```
|
||||
|
||||
2. **Test database credentials:**
|
||||
```bash
|
||||
docker exec library-desk psql -h postgres -U library_desk_listener -d wiki -c "\dt"
|
||||
```
|
||||
|
||||
3. **Verify user exists:**
|
||||
```bash
|
||||
docker exec postgres psql -U postgres -d wiki -c "\du library_desk_listener"
|
||||
```
|
||||
|
||||
### Changes Not Being Processed
|
||||
|
||||
**Problem:** Page edits but no logs appear
|
||||
|
||||
**Solutions:**
|
||||
|
||||
1. **Verify triggers installed:**
|
||||
```bash
|
||||
docker exec postgres psql -U postgres -d wiki -c "
|
||||
SELECT trigger_name FROM information_schema.triggers
|
||||
WHERE event_object_table = 'pages';
|
||||
"
|
||||
```
|
||||
Should list 3 triggers (INSERT, UPDATE, DELETE).
|
||||
|
||||
2. **Test notification manually:**
|
||||
```bash
|
||||
docker exec postgres psql -U postgres -d wiki -c "
|
||||
NOTIFY wiki_page_changes, 'TEST:999:test@example.com';
|
||||
"
|
||||
```
|
||||
Check library-desk logs for "Received TEST notification".
|
||||
|
||||
3. **Check authorEmail field exists:**
|
||||
```bash
|
||||
docker exec postgres psql -U postgres -d wiki -c "
|
||||
SELECT column_name FROM information_schema.columns
|
||||
WHERE table_name = 'pages' AND column_name = 'authorEmail';
|
||||
"
|
||||
```
|
||||
|
||||
### Infinite Loop Detected
|
||||
|
||||
**Symptoms:** Continuous processing logs for the same page
|
||||
|
||||
**Solutions:**
|
||||
|
||||
1. **Verify WIKIJS_USERNAME is correct:**
|
||||
```bash
|
||||
docker exec library-desk env | grep WIKIJS_USERNAME
|
||||
```
|
||||
Must match the Wiki.js username library-desk uses for API calls.
|
||||
|
||||
2. **Check automated user filtering:**
|
||||
```bash
|
||||
docker logs library-desk | grep "automated edit"
|
||||
```
|
||||
If you see no "Skipping" logs, the filter isn't working.
|
||||
|
||||
3. **Manually stop processing:**
|
||||
```bash
|
||||
docker restart library-desk
|
||||
```
|
||||
Then fix the `WIKIJS_USERNAME` configuration.
|
||||
|
||||
### Permission Denied
|
||||
|
||||
**Error:** `permission denied to create trigger`
|
||||
|
||||
**Solutions:**
|
||||
|
||||
1. **Run trigger setup as postgres admin:**
|
||||
```bash
|
||||
docker exec -it postgres psql -U postgres -d wiki
|
||||
# Then run setup_wiki_triggers.py with admin user
|
||||
```
|
||||
|
||||
2. **Grant trigger creation permission:**
|
||||
```sql
|
||||
GRANT CREATE ON DATABASE wiki TO wikijs;
|
||||
ALTER TABLE pages OWNER TO wikijs;
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Configuration Reference
|
||||
|
||||
### Environment Variables
|
||||
|
||||
| Variable | Default | Description |
|
||||
|----------|---------|-------------|
|
||||
| `WIKIJS_USERNAME` | *required* | Wiki.js API user (for loop prevention) |
|
||||
| `WIKIJS_DB_HOST` | `postgres` | PostgreSQL host |
|
||||
| `WIKIJS_DB_PORT` | `5432` | PostgreSQL port |
|
||||
| `WIKIJS_DB_NAME` | `wiki` | Wiki.js database name |
|
||||
| `WIKIJS_DB_USER` | `library_desk_listener` | Read-only database user |
|
||||
| `WIKIJS_DB_PASSWORD` | *required* | Database password |
|
||||
| `WIKIJS_CHANGE_LISTENER_DEBOUNCE_SECONDS` | `5` | Debounce duration (1-60s) |
|
||||
|
||||
### Automated Users List
|
||||
|
||||
Edits by these users are filtered out to prevent loops:
|
||||
|
||||
```python
|
||||
automated_users = [
|
||||
self.settings.wikijs_username, # From WIKIJS_USERNAME
|
||||
"library-desk@system",
|
||||
"automation@system",
|
||||
"bot@system"
|
||||
]
|
||||
```
|
||||
|
||||
**Customize:** Edit `wiki_change_listener.py:134-139` to add your automation users.
|
||||
|
||||
---
|
||||
|
||||
## Security Best Practices
|
||||
|
||||
✅ **Read-only database user** - `library_desk_listener` cannot modify data
|
||||
✅ **No direct database writes** - All Wiki.js updates go through API
|
||||
✅ **Loop prevention** - Automated edits don't trigger re-processing
|
||||
✅ **Error isolation** - Listener failures don't affect Wiki.js
|
||||
✅ **Connection resilience** - asyncpg handles reconnection automatically
|
||||
|
||||
---
|
||||
|
||||
## Uninstalling
|
||||
|
||||
To remove the change detection system:
|
||||
|
||||
### 1. Stop the listener:
|
||||
|
||||
Edit `src/main.py` and comment out lines 354-364 (listener startup).
|
||||
|
||||
### 2. Remove database triggers:
|
||||
|
||||
```bash
|
||||
docker exec -it postgres psql -U postgres -d wiki -c "
|
||||
DROP TRIGGER IF EXISTS wiki_page_insert_trigger ON pages;
|
||||
DROP TRIGGER IF EXISTS wiki_page_update_trigger ON pages;
|
||||
DROP TRIGGER IF EXISTS wiki_page_delete_trigger ON pages;
|
||||
DROP FUNCTION IF EXISTS notify_page_change();
|
||||
"
|
||||
```
|
||||
|
||||
### 3. Remove read-only user:
|
||||
|
||||
```bash
|
||||
docker exec -it postgres psql -U postgres -d wiki -c "
|
||||
DROP USER IF EXISTS library_desk_listener;
|
||||
"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Related Files
|
||||
|
||||
- `src/services/wiki_change_listener.py` - LISTEN service implementation
|
||||
- `src/routers/webhooks.py` - Processing logic (shared)
|
||||
- `src/services/consolidation_service.py` - Bidirectional entity linking
|
||||
- `setup_wiki_triggers.py` - Trigger installation script
|
||||
- `setup_wiki_readonly_user.sql` - Read-only user creation SQL
|
||||
- `src/main.py:354-383` - Listener startup/shutdown hooks
|
||||
- `src/config.py:38-49` - Configuration settings
|
||||
|
||||
---
|
||||
|
||||
## Summary
|
||||
|
||||
With this system enabled:
|
||||
|
||||
| User Action | Processing | Loop Prevention |
|
||||
|-------------|-----------|-----------------|
|
||||
| Manual page edit | ✅ Processed | N/A |
|
||||
| AI agent creates page | ✅ Processed | N/A |
|
||||
| Entity linking updates page | ❌ Skipped | ✅ Automated user filter |
|
||||
| Rapid duplicate edits | ⚡ Debounced | ✅ Timestamp check |
|
||||
|
||||
**Result:** User-edited and AI-generated pages are indistinguishable in the knowledge base, with zero risk of infinite loops!
|
||||
@@ -1,344 +0,0 @@
|
||||
# Wiki.js Webhook Implementation - Complete Summary
|
||||
|
||||
## What We Built
|
||||
|
||||
A comprehensive webhook system that makes **user-edited wiki pages indistinguishable from AI-generated content** in your knowledge base.
|
||||
|
||||
### Files Created/Modified
|
||||
|
||||
| File | Purpose | Lines |
|
||||
|------|---------|-------|
|
||||
| `src/routers/webhooks.py` | Webhook handler for all CRUD events | ~450 |
|
||||
| `src/main.py` | Registered webhook router | +1 |
|
||||
| `WEBHOOK_SETUP.md` | Configuration guide | Documentation |
|
||||
| `WEBHOOK_IMPLEMENTATION_SUMMARY.md` | This file | Documentation |
|
||||
|
||||
---
|
||||
|
||||
## Features Implemented
|
||||
|
||||
### ✅ 1. **Page Creation** (`page.create`)
|
||||
- Triggers full ingestion pipeline
|
||||
- Generates vector embeddings (Qdrant)
|
||||
- Extracts entities (Neo4j)
|
||||
- **Bidirectional entity linking**:
|
||||
- New page links to existing entities
|
||||
- Existing entity pages link back to new page
|
||||
|
||||
### ✅ 2. **Page Updates** (`page.update`)
|
||||
- Re-ingests page with force refresh
|
||||
- Updates vector embeddings
|
||||
- Refreshes entity extraction
|
||||
- **Re-applies bidirectional linking**
|
||||
|
||||
### ✅ 3. **Page Deletion** (`page.delete`)
|
||||
**Comprehensive cleanup:**
|
||||
- ✅ Removes vector embeddings from Qdrant
|
||||
- ✅ Identifies orphaned entities (only mentioned in deleted page)
|
||||
- ✅ Deletes Document node from Neo4j
|
||||
- ✅ Deletes orphaned entities
|
||||
- ✅ Cleans up broken `MENTIONS` relationships
|
||||
- ✅ Cleans up broken `SearchQuery` relationships
|
||||
|
||||
### ✅ 4. **Page Rename/Move** (`page.rename`)
|
||||
**Intelligent handling:**
|
||||
- ✅ Updates Document node path in Neo4j
|
||||
- ✅ Updates title in Neo4j
|
||||
- **Path-only change** (move): Quick path update, no re-processing
|
||||
- **Title change** (rename): Full re-processing + entity relinking
|
||||
|
||||
---
|
||||
|
||||
## How It Works
|
||||
|
||||
### Architecture
|
||||
|
||||
```
|
||||
┌─────────────┐
|
||||
│ Wiki.js │ User edits page
|
||||
│ (Frontend) │
|
||||
└──────┬──────┘
|
||||
│ HTTP POST (webhook)
|
||||
↓
|
||||
┌─────────────────────────────────────┐
|
||||
│ library-desk:8089/webhooks/wikijs │
|
||||
│ (Webhook Handler) │
|
||||
└──────┬──────────────────────────────┘
|
||||
│ Background task
|
||||
↓
|
||||
┌──────────────────────────────────────────┐
|
||||
│ Processing Pipeline │
|
||||
│ ├─ Ingestion (vectors + graph) │
|
||||
│ ├─ Entity extraction │
|
||||
│ ├─ Forward linking (page → entities) │
|
||||
│ └─ Backward linking (entities → page) │
|
||||
└──────────────────────────────────────────┘
|
||||
│
|
||||
↓
|
||||
┌───────────────────────────────────┐
|
||||
│ Knowledge Base Updated │
|
||||
│ ├─ Qdrant (vectors) │
|
||||
│ ├─ Neo4j (graph + entities) │
|
||||
│ └─ Wiki.js (entity links in MD) │
|
||||
└───────────────────────────────────┘
|
||||
```
|
||||
|
||||
### Event Processing Flow
|
||||
|
||||
**1. Page Create/Update**
|
||||
```
|
||||
Webhook received → Extract user from email → Background task:
|
||||
├─ Ingest page (embeddings + entities)
|
||||
├─ Apply forward entity linking
|
||||
├─ Find reverse references
|
||||
└─ Apply backward entity linking
|
||||
```
|
||||
|
||||
**2. Page Delete**
|
||||
```
|
||||
Webhook received → Background task:
|
||||
├─ Remove vectors from Qdrant
|
||||
├─ Find orphaned entities
|
||||
├─ Delete Document node (cascades MENTIONS)
|
||||
├─ Delete orphaned entities
|
||||
└─ Clean broken SearchQuery relationships
|
||||
```
|
||||
|
||||
**3. Page Rename/Move**
|
||||
```
|
||||
Webhook received → Check what changed:
|
||||
├─ Path only? → Update Document.path
|
||||
└─ Title changed? → Re-ingest + re-link entities
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Configuration Required
|
||||
|
||||
### Step 1: Wiki.js Webhook Setup
|
||||
|
||||
Navigate to **Administration** → **Webhooks** → **Add Webhook**
|
||||
|
||||
| Setting | Value |
|
||||
|---------|-------|
|
||||
| Endpoint URL | `http://library-desk:8089/webhooks/wikijs` |
|
||||
| Events | ☑ page.create, page.update, page.delete, page.rename |
|
||||
| Authorization | `Bearer af88ed8f44bed81bdb20d0534f1c4547340b29e2aba4963f61a71b993d7eb6e5` |
|
||||
|
||||
### Step 2: Deploy Updated Code
|
||||
|
||||
```bash
|
||||
cd services/library-desk
|
||||
docker-compose up -d --build library-desk
|
||||
```
|
||||
|
||||
### Step 3: Verify Webhook Endpoint
|
||||
|
||||
```bash
|
||||
curl http://192.168.86.149:8089/webhooks/health
|
||||
|
||||
# Expected response:
|
||||
# {"status": "ok", "service": "webhooks"}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Testing the Webhook
|
||||
|
||||
### Test 1: Create Page
|
||||
1. Create new page in Wiki.js: "Test Entity Linking"
|
||||
2. Add content: "This page discusses Docker and Kubernetes orchestration"
|
||||
3. Save page
|
||||
|
||||
**Expected Result:**
|
||||
- Vectors created in Qdrant
|
||||
- Document node created in Neo4j with entities: Docker, Kubernetes
|
||||
- "Docker" and "Kubernetes" become hyperlinks (if entities exist)
|
||||
- Pages about Docker/Kubernetes now link back to this page
|
||||
|
||||
### Test 2: Update Page
|
||||
1. Edit the page
|
||||
2. Add: "We also use Terraform for infrastructure"
|
||||
3. Save
|
||||
|
||||
**Expected Result:**
|
||||
- Vectors updated
|
||||
- "Terraform" entity extracted
|
||||
- "Terraform" becomes hyperlink
|
||||
- Terraform entity page links back to this page
|
||||
|
||||
### Test 3: Delete Page
|
||||
1. Delete the page from Wiki.js
|
||||
|
||||
**Expected Result:**
|
||||
- Vectors removed from Qdrant
|
||||
- Document node removed from Neo4j
|
||||
- If "Test Entity Linking" was the only page mentioning an entity, that entity is deleted (orphan cleanup)
|
||||
|
||||
### Test 4: Move Page
|
||||
1. Create page at `/tech/docker-guide`
|
||||
2. Move to `/infrastructure/docker-guide`
|
||||
|
||||
**Expected Result:**
|
||||
- Document.path updated to `/infrastructure/docker-guide`
|
||||
- No re-processing (vectors/entities unchanged)
|
||||
- Links remain valid
|
||||
|
||||
### Test 5: Rename Page
|
||||
1. Create page "Docker Guide"
|
||||
2. Rename to "Docker Best Practices"
|
||||
|
||||
**Expected Result:**
|
||||
- Document.title updated
|
||||
- Full re-processing (title is entity-relevant)
|
||||
- Entity linking refreshed
|
||||
|
||||
---
|
||||
|
||||
## Performance Characteristics
|
||||
|
||||
### Response Times
|
||||
- **Webhook acknowledgment**: <100ms (non-blocking)
|
||||
- **Background processing**: 2-10s depending on page size
|
||||
- **No impact on Wiki.js editing experience**
|
||||
|
||||
### Resource Usage
|
||||
Per page edit:
|
||||
- **Ollama**: 1-2 embedding calls (~500ms each)
|
||||
- **Neo4j**: 3-5 queries (~100ms total)
|
||||
- **Qdrant**: 1 upsert operation (~50ms)
|
||||
- **Wiki.js**: 1-N page updates for backward links
|
||||
|
||||
### Scaling Considerations
|
||||
- Background tasks run asynchronously (non-blocking)
|
||||
- For bulk operations (>100 pages), consider:
|
||||
1. Temporarily disable webhook
|
||||
2. Use batch ingestion endpoint
|
||||
3. Re-enable webhook after bulk import
|
||||
|
||||
---
|
||||
|
||||
## Monitoring & Debugging
|
||||
|
||||
### Check Webhook Activity
|
||||
|
||||
```bash
|
||||
# Watch webhook logs
|
||||
docker logs library-desk --follow | grep webhook
|
||||
|
||||
# Look for:
|
||||
# - "Received Wiki.js webhook: page.create"
|
||||
# - "Processing page.update for page 123"
|
||||
# - "Entity linking complete: 5 forward links, 3 backward links"
|
||||
# - "Cleanup complete for deleted page 456"
|
||||
```
|
||||
|
||||
### Verify Processing Complete
|
||||
|
||||
```bash
|
||||
# Check if page was indexed
|
||||
curl -X POST http://192.168.86.149:8089/query/hybrid \
|
||||
-H "Authorization: Bearer $API_KEY" \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{"query": "your page content", "user": "jpmschweitzer"}'
|
||||
|
||||
# Check graph entities
|
||||
# In Neo4j Browser:
|
||||
MATCH (d:Document {page_id: 123})-[:MENTIONS]->(e)
|
||||
RETURN d.title, collect(e.name) as entities
|
||||
```
|
||||
|
||||
### Common Issues & Solutions
|
||||
|
||||
**Problem:** Webhook not triggering
|
||||
- ✅ Check Wiki.js webhook configuration
|
||||
- ✅ Verify library-desk is running
|
||||
- ✅ Check authorization header
|
||||
|
||||
**Problem:** Processing fails silently
|
||||
- ✅ Check library-desk logs: `docker logs library-desk`
|
||||
- ✅ Verify Neo4j/Qdrant/Ollama are accessible
|
||||
- ✅ Check user email mapping in `extract_user_from_email()`
|
||||
|
||||
**Problem:** Orphaned entities remain after deletion
|
||||
- ✅ Orphan detection only removes entities with ZERO `MENTIONS` relationships
|
||||
- ✅ If entity is mentioned in other pages, it's preserved (correct behavior)
|
||||
|
||||
---
|
||||
|
||||
## Advanced Customization
|
||||
|
||||
### 1. Custom User Mapping
|
||||
|
||||
Edit `src/routers/webhooks.py`:
|
||||
|
||||
```python
|
||||
def extract_user_from_email(email: str) -> str:
|
||||
"""Map Wiki.js user emails to library-desk users"""
|
||||
user_map = {
|
||||
"admin@example.com": "admin",
|
||||
"john@example.com": "jpmschweitzer"
|
||||
}
|
||||
return user_map.get(email, email.split("@")[0])
|
||||
```
|
||||
|
||||
### 2. Selective Processing
|
||||
|
||||
Process only specific paths:
|
||||
|
||||
```python
|
||||
# In handle_wikijs_webhook():
|
||||
page_path = payload.page.get("path", "")
|
||||
if not page_path.startswith("/tech/"):
|
||||
logger.info(f"Skipping page outside /tech/")
|
||||
return
|
||||
```
|
||||
|
||||
### 3. Disable Backward Linking
|
||||
|
||||
To save processing time, disable backward linking:
|
||||
|
||||
```python
|
||||
# In process_wiki_page_change():
|
||||
# Comment out:
|
||||
# link_stats = await consolidation_service._apply_bidirectional_entity_linking(...)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Success Criteria
|
||||
|
||||
✅ **User-edited pages** have vectors in Qdrant
|
||||
✅ **User-edited pages** have entities in Neo4j
|
||||
✅ **User-edited pages** have bidirectional entity links
|
||||
✅ **Deleted pages** are cleaned from vectors + graph
|
||||
✅ **Renamed/moved pages** update correctly
|
||||
✅ **No difference** between manual and AI-generated pages
|
||||
|
||||
---
|
||||
|
||||
## Next Steps
|
||||
|
||||
1. **Deploy the code** (restart library-desk)
|
||||
2. **Configure Wiki.js webhook** (Administration → Webhooks)
|
||||
3. **Test with a sample page** (create, edit, delete)
|
||||
4. **Monitor logs** (verify processing completes)
|
||||
5. **Verify results** (check vectors, graph, entity links)
|
||||
|
||||
---
|
||||
|
||||
## Summary
|
||||
|
||||
**Before webhooks:**
|
||||
- Only AI-generated pages fully integrated
|
||||
- Manual edits not indexed/linked
|
||||
- Inconsistent knowledge base
|
||||
|
||||
**After webhooks:**
|
||||
- **All pages processed identically**
|
||||
- **Automatic entity extraction**
|
||||
- **Bidirectional linking**
|
||||
- **Clean deletion handling**
|
||||
- **Rename/move support**
|
||||
|
||||
**Result:** A unified, consistent knowledge base regardless of content source! 🎉
|
||||
@@ -1,302 +0,0 @@
|
||||
# Wiki.js Webhook Setup Guide
|
||||
|
||||
Complete guide to enable automatic processing of user-edited wiki pages.
|
||||
|
||||
## Overview
|
||||
|
||||
When you manually edit pages in Wiki.js, the webhook system ensures they're processed identically to AI-generated content:
|
||||
|
||||
- ✅ **Vectors updated** - Qdrant embeddings refreshed
|
||||
- ✅ **Graph updated** - Entities extracted, relationships created
|
||||
- ✅ **Bidirectional links** - Automatic entity linking both ways
|
||||
- ✅ **Deletions handled** - Clean removal from vectors + graph
|
||||
- ✅ **Moves processed** - Path updates propagated to graph
|
||||
|
||||
## Step 1: Configure Wiki.js Webhook
|
||||
|
||||
### Access Wiki.js Administration
|
||||
|
||||
1. Log into Wiki.js as administrator
|
||||
2. Navigate to **Administration** → **Webhooks**
|
||||
3. Click **Add New Webhook**
|
||||
|
||||
### Webhook Configuration
|
||||
|
||||
| Field | Value |
|
||||
|-------|-------|
|
||||
| **Name** | Library Desk Integration |
|
||||
| **Endpoint URL** | `http://library-desk:8089/webhooks/wikijs` |
|
||||
| **Content Type** | `application/json` |
|
||||
| **Events** | ☑ page.create<br>☑ page.update<br>☑ page.delete<br>☑ page.rename |
|
||||
| **Status** | ✅ Active |
|
||||
|
||||
### Authentication Header
|
||||
|
||||
Add custom header for API authentication:
|
||||
|
||||
```
|
||||
Authorization: Bearer af88ed8f44bed81bdb20d0534f1c4547340b29e2aba4963f61a71b993d7eb6e5
|
||||
```
|
||||
|
||||
*(Use your actual Library Desk API key from `.env`)*
|
||||
|
||||
## Step 2: Verify Webhook Endpoint
|
||||
|
||||
Test the webhook health endpoint:
|
||||
|
||||
```bash
|
||||
curl http://192.168.86.149:8089/webhooks/health
|
||||
```
|
||||
|
||||
Expected response:
|
||||
```json
|
||||
{
|
||||
"status": "ok",
|
||||
"service": "webhooks"
|
||||
}
|
||||
```
|
||||
|
||||
## Step 3: Test Webhook Integration
|
||||
|
||||
### Test 1: Page Creation
|
||||
|
||||
1. Create a new page in Wiki.js
|
||||
2. Add content mentioning existing entities (e.g., "Docker", "Kubernetes")
|
||||
3. Save the page
|
||||
|
||||
**Expected Result:**
|
||||
- Vectors created in Qdrant
|
||||
- Document node created in Neo4j
|
||||
- Entities extracted and linked
|
||||
- Bidirectional links added (new page links to entities, entities' pages link back)
|
||||
|
||||
### Test 2: Page Update
|
||||
|
||||
1. Edit an existing page
|
||||
2. Add new content with entity mentions
|
||||
3. Save changes
|
||||
|
||||
**Expected Result:**
|
||||
- Vectors updated in Qdrant (old removed, new added)
|
||||
- Graph updated with new entities
|
||||
- Entity links refreshed
|
||||
|
||||
### Test 3: Page Deletion
|
||||
|
||||
1. Delete a page from Wiki.js
|
||||
|
||||
**Expected Result:**
|
||||
- Vectors removed from Qdrant
|
||||
- Document node removed from Neo4j
|
||||
- Orphaned entities cleaned up (entities only mentioned in this page)
|
||||
- Broken relationships removed
|
||||
|
||||
### Test 4: Page Move/Rename
|
||||
|
||||
1. Move a page to different location (e.g., `/tech/docker` → `/infrastructure/docker`)
|
||||
2. Or rename the page
|
||||
|
||||
**Expected Result:**
|
||||
- Document node path updated in Neo4j
|
||||
- Entity links updated if title changed
|
||||
- Vectors remain valid (no re-embedding needed unless content changed)
|
||||
|
||||
## Event Processing Details
|
||||
|
||||
### page.create
|
||||
|
||||
```
|
||||
User creates page → Webhook → Library Desk
|
||||
├─ Ingest page (vector + graph)
|
||||
├─ Extract entities
|
||||
├─ Forward linking (page → entities)
|
||||
└─ Backward linking (entities → page)
|
||||
```
|
||||
|
||||
### page.update
|
||||
|
||||
```
|
||||
User edits page → Webhook → Library Desk
|
||||
├─ Re-ingest page (force refresh)
|
||||
├─ Update entities
|
||||
├─ Refresh forward links
|
||||
└─ Refresh backward links
|
||||
```
|
||||
|
||||
### page.delete
|
||||
|
||||
```
|
||||
User deletes page → Webhook → Library Desk
|
||||
├─ Remove vectors from Qdrant
|
||||
├─ Find orphaned entities
|
||||
├─ Delete Document node
|
||||
├─ Delete orphaned entities
|
||||
└─ Clean broken relationships
|
||||
```
|
||||
|
||||
### page.rename
|
||||
|
||||
```
|
||||
User moves/renames page → Webhook → Library Desk
|
||||
├─ Update Document node path
|
||||
├─ Update title if changed
|
||||
├─ Re-link entities if title changed
|
||||
└─ Update references
|
||||
```
|
||||
|
||||
## Monitoring Webhook Activity
|
||||
|
||||
### Check Logs
|
||||
|
||||
```bash
|
||||
# Library Desk logs
|
||||
docker logs library-desk --follow | grep webhook
|
||||
|
||||
# Look for:
|
||||
# - "Received Wiki.js webhook: page.create"
|
||||
# - "Processing page.update for page 123"
|
||||
# - "Cleanup complete for deleted page 456"
|
||||
```
|
||||
|
||||
### Verify Processing
|
||||
|
||||
After making changes in Wiki.js, verify:
|
||||
|
||||
**1. Vectors Updated:**
|
||||
```bash
|
||||
curl -X POST http://192.168.86.149:8089/query/hybrid \
|
||||
-H "Authorization: Bearer $API_KEY" \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{"query": "your page content", "user": "jpmschweitzer"}'
|
||||
```
|
||||
|
||||
**2. Graph Updated:**
|
||||
```cypher
|
||||
// In Neo4j Browser
|
||||
MATCH (d:Document {page_id: 123})
|
||||
OPTIONAL MATCH (d)-[:MENTIONS]->(e)
|
||||
RETURN d, collect(e.name) as entities
|
||||
```
|
||||
|
||||
**3. Entity Links Added:**
|
||||
Visit the page in Wiki.js and verify entity names are hyperlinked.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
### Webhook Not Triggering
|
||||
|
||||
**Problem:** Changes in Wiki.js don't trigger processing
|
||||
|
||||
**Solutions:**
|
||||
1. Check webhook configuration in Wiki.js Admin
|
||||
2. Verify endpoint URL is correct: `http://library-desk:8089/webhooks/wikijs`
|
||||
3. Check authorization header is set
|
||||
4. Verify library-desk is running: `docker ps | grep library-desk`
|
||||
5. Check library-desk logs: `docker logs library-desk --tail 50`
|
||||
|
||||
### Processing Failures
|
||||
|
||||
**Problem:** Webhook triggers but processing fails
|
||||
|
||||
**Solutions:**
|
||||
1. Check library-desk logs for errors
|
||||
2. Verify Neo4j is accessible
|
||||
3. Verify Qdrant is accessible
|
||||
4. Check Ollama is running (for embeddings)
|
||||
5. Verify user mapping in `extract_user_from_email()`
|
||||
|
||||
### Orphaned Entities
|
||||
|
||||
**Problem:** Deleted pages leave entities behind
|
||||
|
||||
**Solution:**
|
||||
Manually clean orphaned entities:
|
||||
```cypher
|
||||
// Find orphaned entities (no MENTIONS relationships)
|
||||
MATCH (e:User_Jpmschweitzer)
|
||||
WHERE NOT e:Document
|
||||
AND NOT EXISTS {(d:Document)-[:MENTIONS]->(e)}
|
||||
DETACH DELETE e
|
||||
RETURN count(e) as deleted
|
||||
```
|
||||
|
||||
### Broken Links
|
||||
|
||||
**Problem:** Entity links broken after page moves
|
||||
|
||||
**Solution:**
|
||||
Re-run entity linking on all pages:
|
||||
```bash
|
||||
curl -X POST http://192.168.86.149:8089/entity-linking/link-all \
|
||||
-H "Authorization: Bearer $API_KEY" \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{"user": "jpmschweitzer"}'
|
||||
```
|
||||
|
||||
## Performance Considerations
|
||||
|
||||
### Background Processing
|
||||
|
||||
Webhook handler processes events in the background to avoid blocking Wiki.js:
|
||||
- Webhook response: <100ms (immediate acknowledgment)
|
||||
- Processing: 2-10s (runs asynchronously)
|
||||
- No impact on Wiki.js editing experience
|
||||
|
||||
### Rate Limiting
|
||||
|
||||
For bulk operations (importing many pages):
|
||||
- Consider temporarily disabling webhook
|
||||
- Use batch ingestion endpoint instead: `/ingest/batch`
|
||||
- Re-enable webhook after bulk import
|
||||
|
||||
## Advanced Configuration
|
||||
|
||||
### Custom User Mapping
|
||||
|
||||
By default, user is extracted from email (`user@domain` → `user`).
|
||||
|
||||
Customize in `/services/library-desk/src/routers/webhooks.py`:
|
||||
|
||||
```python
|
||||
def extract_user_from_email(email: str) -> str:
|
||||
# Option 1: Map specific emails to users
|
||||
user_map = {
|
||||
"admin@example.com": "admin",
|
||||
"john.doe@example.com": "jpmschweitzer"
|
||||
}
|
||||
return user_map.get(email, email.split("@")[0])
|
||||
```
|
||||
|
||||
### Event Filtering
|
||||
|
||||
To process only certain pages, add filtering in webhook handler:
|
||||
|
||||
```python
|
||||
# Only process pages in /tech/ path
|
||||
if not payload.page.get("path", "").startswith("/tech/"):
|
||||
logger.info(f"Skipping page outside /tech/: {payload.page.get('path')}")
|
||||
return
|
||||
```
|
||||
|
||||
### Disable Bidirectional Linking
|
||||
|
||||
To only do forward linking (page → entities) without backward linking:
|
||||
|
||||
Edit `process_wiki_page_change()`:
|
||||
```python
|
||||
# Comment out the bidirectional linking call
|
||||
# link_stats = await consolidation_service._apply_bidirectional_entity_linking(...)
|
||||
```
|
||||
|
||||
## Summary
|
||||
|
||||
With webhooks configured:
|
||||
|
||||
| User Action | Automatic Processing |
|
||||
|-------------|---------------------|
|
||||
| Create page | Ingest + entity link (bidirectional) |
|
||||
| Edit page | Re-ingest + refresh links |
|
||||
| Delete page | Clean vectors + graph + orphaned entities |
|
||||
| Move page | Update paths + refresh links |
|
||||
|
||||
**Result:** User-edited and AI-generated pages are **indistinguishable** in the knowledge base!
|
||||
@@ -1,441 +0,0 @@
|
||||
# Wiki.js Page Change Detection
|
||||
|
||||
Automatic processing system for user-edited Wiki.js pages using PostgreSQL database triggers.
|
||||
|
||||
## Overview
|
||||
|
||||
**Problem:** Wiki.js doesn't have built-in webhooks (feature is [planned but not released](https://requarks.canny.io/wiki/p/webhooks))
|
||||
|
||||
**Solution:** PostgreSQL NOTIFY/LISTEN triggers to detect page changes in real-time
|
||||
|
||||
When you manually edit pages in Wiki.js, the change detection system ensures they're processed identically to AI-generated content:
|
||||
|
||||
- ✅ **Vectors updated** - Qdrant embeddings refreshed
|
||||
- ✅ **Graph updated** - Entities extracted, relationships created
|
||||
- ✅ **Bidirectional links** - Automatic entity linking both ways
|
||||
- ✅ **Deletions handled** - Clean removal from vectors + graph
|
||||
|
||||
---
|
||||
|
||||
## How It Works
|
||||
|
||||
```
|
||||
┌─────────────────────────────────────────────┐
|
||||
│ Wiki.js (User edits page) │
|
||||
└──────────────────┬──────────────────────────┘
|
||||
│
|
||||
↓
|
||||
┌─────────────────────────────────────────────┐
|
||||
│ PostgreSQL "pages" table │
|
||||
│ INSERT/UPDATE/DELETE triggers fire │
|
||||
└──────────────────┬──────────────────────────┘
|
||||
│ pg_notify('wiki_page_changes', ...)
|
||||
↓
|
||||
┌─────────────────────────────────────────────┐
|
||||
│ library-desk WikiChangeListener │
|
||||
│ Receives notification via LISTEN │
|
||||
└──────────────────┬──────────────────────────┘
|
||||
│
|
||||
↓
|
||||
┌─────────────────────────────────────────────┐
|
||||
│ Processing Pipeline (same as webhooks) │
|
||||
│ ├─ Ingest page (vectors + graph) │
|
||||
│ ├─ Extract entities │
|
||||
│ ├─ Forward linking (page → entities) │
|
||||
│ └─ Backward linking (entities → page) │
|
||||
└─────────────────────────────────────────────┘
|
||||
```
|
||||
|
||||
**Key Components:**
|
||||
|
||||
1. **PostgreSQL Triggers** - Installed in Wiki.js database, emit NOTIFY events on page changes
|
||||
2. **WikiChangeListener** - Python service that LISTENs for notifications and processes changes
|
||||
3. **Processing Pipeline** - Same logic as webhook handler (uses `process_wiki_page_change()`)
|
||||
|
||||
---
|
||||
|
||||
## Setup Instructions
|
||||
|
||||
### Prerequisites
|
||||
|
||||
- Library-desk has network access to Wiki.js PostgreSQL database
|
||||
- Database user has `CREATE FUNCTION` and `CREATE TRIGGER` permissions
|
||||
|
||||
### Step 1: Configure Environment Variables
|
||||
|
||||
Add these to your `.env` file:
|
||||
|
||||
```bash
|
||||
# Wiki.js Database Configuration (for change detection)
|
||||
WIKIJS_DB_HOST=postgres # PostgreSQL host (usually same as Wiki.js)
|
||||
WIKIJS_DB_PORT=5432 # PostgreSQL port
|
||||
WIKIJS_DB_NAME=wiki # Wiki.js database name
|
||||
WIKIJS_DB_USER=wikijs # Database user (needs trigger permissions)
|
||||
WIKIJS_DB_PASSWORD=your_password # Database password
|
||||
```
|
||||
|
||||
### Step 2: Install Database Triggers
|
||||
|
||||
Run the setup script from within the library-desk container:
|
||||
|
||||
```bash
|
||||
# Enter the container
|
||||
docker exec -it library-desk bash
|
||||
|
||||
# Run trigger setup script
|
||||
cd /app
|
||||
python3 setup_wiki_triggers.py
|
||||
```
|
||||
|
||||
**Expected output:**
|
||||
```
|
||||
======================================================================
|
||||
Wiki.js Database Trigger Setup
|
||||
======================================================================
|
||||
|
||||
Connecting to Wiki.js database at postgres:5432
|
||||
Database: wiki
|
||||
User: wikijs
|
||||
|
||||
✓ Connected to Wiki.js database
|
||||
|
||||
Installing triggers...
|
||||
|
||||
✓ Triggers installed successfully:
|
||||
|
||||
- wiki_page_delete_trigger (DELETE)
|
||||
- wiki_page_insert_trigger (INSERT)
|
||||
- wiki_page_update_trigger (UPDATE)
|
||||
|
||||
======================================================================
|
||||
Setup complete!
|
||||
======================================================================
|
||||
|
||||
Next steps:
|
||||
1. Restart library-desk service to activate the listener
|
||||
2. Edit a page in Wiki.js to test
|
||||
3. Check library-desk logs for processing messages
|
||||
```
|
||||
|
||||
### Step 3: Restart Library-Desk
|
||||
|
||||
```bash
|
||||
docker restart library-desk
|
||||
```
|
||||
|
||||
**Check logs for successful startup:**
|
||||
```bash
|
||||
docker logs library-desk --tail 50 | grep -E "(Wiki.js change listener|wiki_page_changes)"
|
||||
```
|
||||
|
||||
You should see:
|
||||
```
|
||||
INFO - Starting Wiki.js database change listener
|
||||
INFO - Listening for Wiki.js page changes via PostgreSQL NOTIFY
|
||||
INFO - Wiki.js change listener started successfully
|
||||
```
|
||||
|
||||
### Step 4: Test the System
|
||||
|
||||
1. **Create a new page in Wiki.js**
|
||||
- Add content mentioning existing entities (e.g., "Docker", "Kubernetes")
|
||||
- Save the page
|
||||
|
||||
2. **Check library-desk logs:**
|
||||
```bash
|
||||
docker logs library-desk --follow | grep -E "(Received|Processing|Entity linking)"
|
||||
```
|
||||
|
||||
**Expected log output:**
|
||||
```
|
||||
INFO - Received INSERT notification for page 123 by user@example.com
|
||||
INFO - Processing page.create for page 123 ('Test Page')
|
||||
INFO - Ingesting page 123 into knowledge base
|
||||
INFO - Ingestion complete for page 123
|
||||
INFO - Applying bidirectional entity linking for page 123
|
||||
INFO - Entity linking complete for page 123: 5 forward links, 3 backward links (2 pages updated)
|
||||
INFO - Successfully processed page.create for page 123
|
||||
```
|
||||
|
||||
3. **Verify results:**
|
||||
- Check vectors in Qdrant
|
||||
- Check Neo4j for Document node and entities
|
||||
- Visit the page in Wiki.js - entities should be hyperlinked
|
||||
|
||||
---
|
||||
|
||||
## Event Processing Details
|
||||
|
||||
### page.create (INSERT trigger)
|
||||
|
||||
```
|
||||
User creates page → PostgreSQL INSERT trigger fires
|
||||
├─ pg_notify('wiki_page_changes', 'INSERT:123:user@example.com')
|
||||
├─ WikiChangeListener receives notification
|
||||
├─ Ingest page (vector + graph)
|
||||
├─ Extract entities
|
||||
├─ Forward linking (page → entities)
|
||||
└─ Backward linking (entities → page)
|
||||
```
|
||||
|
||||
### page.update (UPDATE trigger)
|
||||
|
||||
```
|
||||
User edits page → PostgreSQL UPDATE trigger fires
|
||||
├─ pg_notify('wiki_page_changes', 'UPDATE:123:user@example.com')
|
||||
├─ WikiChangeListener receives notification
|
||||
├─ Re-ingest page (force refresh)
|
||||
├─ Update entities
|
||||
├─ Refresh forward links
|
||||
└─ Refresh backward links
|
||||
```
|
||||
|
||||
### page.delete (DELETE trigger)
|
||||
|
||||
```
|
||||
User deletes page → PostgreSQL DELETE trigger fires
|
||||
├─ pg_notify('wiki_page_changes', 'DELETE:123:user@example.com')
|
||||
├─ WikiChangeListener receives notification
|
||||
├─ Remove vectors from Qdrant
|
||||
├─ Find orphaned entities
|
||||
├─ Delete Document node
|
||||
├─ Delete orphaned entities
|
||||
└─ Clean broken relationships
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Monitoring
|
||||
|
||||
### Check Listener Status
|
||||
|
||||
```bash
|
||||
# Check if listener is running
|
||||
docker logs library-desk --tail 50 | grep "Wiki.js change listener"
|
||||
|
||||
# Should show:
|
||||
# INFO - Wiki.js change listener started successfully
|
||||
# INFO - Listening for Wiki.js page changes via PostgreSQL NOTIFY
|
||||
```
|
||||
|
||||
### Monitor Page Changes
|
||||
|
||||
```bash
|
||||
# Watch for page change notifications
|
||||
docker logs library-desk --follow | grep -E "(Received.*notification|Processing page)"
|
||||
```
|
||||
|
||||
### Verify Trigger Installation
|
||||
|
||||
Connect to Wiki.js database and check:
|
||||
|
||||
```sql
|
||||
-- List installed triggers
|
||||
SELECT trigger_name, event_manipulation, event_object_table
|
||||
FROM information_schema.triggers
|
||||
WHERE event_object_table = 'pages'
|
||||
ORDER BY trigger_name;
|
||||
|
||||
-- Expected result:
|
||||
-- wiki_page_delete_trigger | DELETE | pages
|
||||
-- wiki_page_insert_trigger | INSERT | pages
|
||||
-- wiki_page_update_trigger | UPDATE | pages
|
||||
```
|
||||
|
||||
### Test Notification Manually
|
||||
|
||||
You can manually test the NOTIFY/LISTEN system:
|
||||
|
||||
```sql
|
||||
-- In one session, listen for notifications:
|
||||
LISTEN wiki_page_changes;
|
||||
|
||||
-- In another session, manually trigger:
|
||||
NOTIFY wiki_page_changes, 'TEST:999:test@example.com';
|
||||
|
||||
-- The first session should receive the notification
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
### Change Listener Not Starting
|
||||
|
||||
**Problem:** Logs show "Failed to start Wiki.js change listener"
|
||||
|
||||
**Solutions:**
|
||||
1. Check database connection settings in `.env`
|
||||
2. Verify library-desk can reach Wiki.js PostgreSQL:
|
||||
```bash
|
||||
docker exec library-desk ping postgres
|
||||
```
|
||||
3. Test database credentials:
|
||||
```bash
|
||||
docker exec library-desk psql -h postgres -U wikijs -d wiki -c "\dt"
|
||||
```
|
||||
|
||||
### Changes Not Being Processed
|
||||
|
||||
**Problem:** Pages edited but no processing logs appear
|
||||
|
||||
**Solutions:**
|
||||
|
||||
1. **Check triggers are installed:**
|
||||
```bash
|
||||
docker exec -it postgres psql -U wikijs -d wiki -c "
|
||||
SELECT trigger_name FROM information_schema.triggers
|
||||
WHERE event_object_table = 'pages';
|
||||
"
|
||||
```
|
||||
|
||||
2. **Verify listener is running:**
|
||||
```bash
|
||||
docker logs library-desk | grep "Listening for Wiki.js page changes"
|
||||
```
|
||||
|
||||
3. **Test notification manually:**
|
||||
```bash
|
||||
docker exec -it postgres psql -U wikijs -d wiki -c "
|
||||
NOTIFY wiki_page_changes, 'INSERT:123:test@example.com';
|
||||
"
|
||||
```
|
||||
Check library-desk logs for "Received INSERT notification"
|
||||
|
||||
4. **Check authorEmail field exists:**
|
||||
The triggers use `authorEmail` field from Wiki.js pages table. Verify:
|
||||
```sql
|
||||
SELECT column_name
|
||||
FROM information_schema.columns
|
||||
WHERE table_name = 'pages' AND column_name = 'authorEmail';
|
||||
```
|
||||
|
||||
### Trigger Permission Denied
|
||||
|
||||
**Problem:** Setup script fails with "permission denied to create trigger"
|
||||
|
||||
**Solutions:**
|
||||
|
||||
1. Grant necessary permissions to database user:
|
||||
```sql
|
||||
GRANT CREATE ON DATABASE wiki TO wikijs;
|
||||
ALTER TABLE pages OWNER TO wikijs;
|
||||
```
|
||||
|
||||
2. Or run setup script with database admin user:
|
||||
```bash
|
||||
# Modify WIKIJS_DB_USER temporarily to postgres
|
||||
WIKIJS_DB_USER=postgres python3 setup_wiki_triggers.py
|
||||
```
|
||||
|
||||
### Processing Errors
|
||||
|
||||
**Problem:** Notification received but processing fails
|
||||
|
||||
**Solutions:**
|
||||
|
||||
1. Check library-desk logs for specific error:
|
||||
```bash
|
||||
docker logs library-desk --tail 100 | grep -A 10 "Failed to handle notification"
|
||||
```
|
||||
|
||||
2. Verify user mapping works:
|
||||
- Triggers extract user from `authorEmail` field
|
||||
- Check `extract_user_from_email()` in `wiki_change_listener.py`
|
||||
- Default: `user@domain.com` → `user`
|
||||
|
||||
3. Ensure page_id exists and is accessible
|
||||
|
||||
---
|
||||
|
||||
## User Mapping
|
||||
|
||||
By default, the user is extracted from the Wiki.js `authorEmail` field:
|
||||
|
||||
```python
|
||||
# user@example.com → user
|
||||
user = user_email.split('@')[0]
|
||||
```
|
||||
|
||||
**Custom mapping:** Edit `src/services/wiki_change_listener.py`:
|
||||
|
||||
```python
|
||||
async def _handle_notification(self, connection, pid, channel, payload):
|
||||
# ... existing code ...
|
||||
|
||||
# Custom user mapping
|
||||
user_map = {
|
||||
"admin@example.com": "admin",
|
||||
"john.doe@example.com": "jpmschweitzer"
|
||||
}
|
||||
user = user_map.get(user_email, user_email.split('@')[0])
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Comparison: Webhooks vs Database Triggers
|
||||
|
||||
| Feature | Webhooks (Not Available) | PostgreSQL Triggers |
|
||||
|---------|-------------------------|---------------------|
|
||||
| **Availability** | Planned, not released | Available now |
|
||||
| **Setup** | Simple (UI configuration) | Requires database access |
|
||||
| **Reliability** | HTTP-based, can fail | Database-native, highly reliable |
|
||||
| **Performance** | Network overhead | Minimal overhead |
|
||||
| **Maintenance** | None | Must survive database migrations |
|
||||
| **Isolation** | Loose coupling | Tight coupling to database |
|
||||
|
||||
**Recommendation:** Use database triggers until Wiki.js webhooks are officially released. When webhooks become available, migrate to that approach for better isolation.
|
||||
|
||||
---
|
||||
|
||||
## Uninstalling
|
||||
|
||||
To remove the triggers (if switching to webhooks later):
|
||||
|
||||
```bash
|
||||
docker exec -it postgres psql -U wikijs -d wiki -c "
|
||||
DROP TRIGGER IF EXISTS wiki_page_insert_trigger ON pages;
|
||||
DROP TRIGGER IF EXISTS wiki_page_update_trigger ON pages;
|
||||
DROP TRIGGER IF EXISTS wiki_page_delete_trigger ON pages;
|
||||
DROP FUNCTION IF EXISTS notify_page_change();
|
||||
"
|
||||
```
|
||||
|
||||
Then remove the listener startup from `src/main.py` (lines 354-364).
|
||||
|
||||
---
|
||||
|
||||
## Summary
|
||||
|
||||
With database triggers configured:
|
||||
|
||||
| User Action | Automatic Processing |
|
||||
|-------------|---------------------:|
|
||||
| Create page | Ingest + entity link (bidirectional) |
|
||||
| Edit page | Re-ingest + refresh links |
|
||||
| Delete page | Clean vectors + graph + orphaned entities |
|
||||
|
||||
**Result:** User-edited and AI-generated pages are **indistinguishable** in the knowledge base!
|
||||
|
||||
---
|
||||
|
||||
## Related Files
|
||||
|
||||
- `src/services/wiki_change_listener.py` - LISTEN service implementation
|
||||
- `src/routers/webhooks.py` - Processing logic (shared with future webhook approach)
|
||||
- `src/services/consolidation_service.py` - Bidirectional entity linking
|
||||
- `setup_wiki_triggers.py` - Trigger installation script
|
||||
- `src/main.py` - Application startup/shutdown hooks
|
||||
|
||||
---
|
||||
|
||||
## Future: Migration to Webhooks
|
||||
|
||||
When Wiki.js webhooks are released:
|
||||
|
||||
1. Configure webhook in Wiki.js admin panel
|
||||
2. Point to `http://library-desk:8089/webhooks/wikijs`
|
||||
3. Remove database triggers (see Uninstalling section)
|
||||
4. Disable listener startup in `src/main.py`
|
||||
|
||||
The webhook handler already exists and uses the same processing logic!
|
||||
@@ -1,216 +0,0 @@
|
||||
#!/usr/bin/env python3
|
||||
"""
|
||||
Cleanup script to remove all test data from Neo4j graph database.
|
||||
This prepares the system for production by removing:
|
||||
- SearchQuery nodes (HybridRAG search history)
|
||||
- WebResult nodes (external web search results)
|
||||
- Any orphaned nodes without relationships
|
||||
"""
|
||||
import requests
|
||||
import sys
|
||||
|
||||
# Configuration
|
||||
BASE_URL = "http://192.168.86.149:8089"
|
||||
API_KEY = "af88ed8f44bed81bdb20d0534f1c4547340b29e2aba4963f61a71b993d7eb6e5"
|
||||
USER = "jpmschweitzer"
|
||||
|
||||
headers = {
|
||||
"Authorization": f"Bearer {API_KEY}",
|
||||
"Content-Type": "application/json"
|
||||
}
|
||||
|
||||
def execute_query(query_text, description):
|
||||
"""Execute a Cypher query via API."""
|
||||
url = f"{BASE_URL}/graph/query"
|
||||
payload = {
|
||||
"user": USER,
|
||||
"query": query_text
|
||||
}
|
||||
|
||||
print(f"🔄 {description}...")
|
||||
response = requests.post(url, headers=headers, json=payload)
|
||||
|
||||
if response.status_code != 200:
|
||||
print(f" ❌ Failed: {response.status_code}")
|
||||
print(f" {response.text}")
|
||||
return None
|
||||
|
||||
data = response.json()
|
||||
return data
|
||||
|
||||
def count_nodes(node_type):
|
||||
"""Count nodes of a specific type."""
|
||||
if node_type == "SearchQuery":
|
||||
query = f'MATCH (n:SearchQuery) WHERE n.user = "{USER}" RETURN count(n) as count'
|
||||
elif node_type == "WebResult":
|
||||
query = f'MATCH (n:User_{USER}_WebResult:WebResult) RETURN count(n) as count'
|
||||
elif node_type == "Document":
|
||||
query = f'MATCH (n:User_{USER}_Document:Document) RETURN count(n) as count'
|
||||
else:
|
||||
query = f'MATCH (n:User_{USER}) WHERE NOT n:Document AND NOT n:SearchQuery AND NOT n:WebResult RETURN count(n) as count'
|
||||
|
||||
result = execute_query(query, f"Counting {node_type} nodes")
|
||||
if result and result.get("results"):
|
||||
return result["results"][0]["count"]
|
||||
return 0
|
||||
|
||||
def delete_search_queries():
|
||||
"""Delete all SearchQuery nodes and their relationships."""
|
||||
query = f'''
|
||||
MATCH (sq:SearchQuery)
|
||||
WHERE sq.user = "{USER}"
|
||||
DETACH DELETE sq
|
||||
RETURN count(*) as deleted
|
||||
'''
|
||||
|
||||
result = execute_query(query, "Deleting SearchQuery nodes")
|
||||
if result and result.get("results"):
|
||||
return result["results"][0].get("deleted", 0)
|
||||
return 0
|
||||
|
||||
def delete_web_results():
|
||||
"""Delete all WebResult nodes."""
|
||||
query = f'''
|
||||
MATCH (wr:User_{USER}_WebResult:WebResult)
|
||||
DETACH DELETE wr
|
||||
RETURN count(*) as deleted
|
||||
'''
|
||||
|
||||
result = execute_query(query, "Deleting WebResult nodes")
|
||||
if result and result.get("results"):
|
||||
return result["results"][0].get("deleted", 0)
|
||||
return 0
|
||||
|
||||
def delete_orphaned_nodes():
|
||||
"""Delete any orphaned entity nodes without relationships."""
|
||||
query = f'''
|
||||
MATCH (n:User_{USER})
|
||||
WHERE NOT n:Document
|
||||
AND NOT n:SearchQuery
|
||||
AND NOT n:WebResult
|
||||
AND NOT (n)--()
|
||||
DETACH DELETE n
|
||||
RETURN count(*) as deleted
|
||||
'''
|
||||
|
||||
result = execute_query(query, "Deleting orphaned nodes")
|
||||
if result and result.get("results"):
|
||||
return result["results"][0].get("deleted", 0)
|
||||
return 0
|
||||
|
||||
def get_all_node_stats():
|
||||
"""Get statistics on all node types."""
|
||||
query = f'''
|
||||
MATCH (n:User_{USER})
|
||||
WITH labels(n) as labels, count(n) as count
|
||||
RETURN labels, count
|
||||
ORDER BY count DESC
|
||||
'''
|
||||
|
||||
result = execute_query(query, "Getting node statistics")
|
||||
if result and result.get("results"):
|
||||
return result["results"]
|
||||
return []
|
||||
|
||||
def main():
|
||||
print("=" * 70)
|
||||
print("🧹 Neo4j Graph Database Cleanup Script")
|
||||
print("=" * 70)
|
||||
print()
|
||||
|
||||
# Show current state
|
||||
print("📊 Current Database State:")
|
||||
print("-" * 70)
|
||||
|
||||
search_count = count_nodes("SearchQuery")
|
||||
web_count = count_nodes("WebResult")
|
||||
doc_count = count_nodes("Document")
|
||||
other_count = count_nodes("Other")
|
||||
|
||||
print(f" SearchQuery nodes: {search_count}")
|
||||
print(f" WebResult nodes: {web_count}")
|
||||
print(f" Document nodes: {doc_count}")
|
||||
print(f" Other entity nodes: {other_count}")
|
||||
print()
|
||||
|
||||
if search_count == 0 and web_count == 0 and doc_count == 0 and other_count == 0:
|
||||
print("✨ Database is already clean!")
|
||||
return
|
||||
|
||||
total_to_delete = search_count + web_count + other_count
|
||||
if total_to_delete == 0 and doc_count > 0:
|
||||
print("⚠️ Only Document nodes found (legitimate data)")
|
||||
print(" No cleanup needed!")
|
||||
return
|
||||
|
||||
# Ask for confirmation
|
||||
print("⚠️ WARNING: About to delete:")
|
||||
if search_count > 0:
|
||||
print(f" - {search_count} SearchQuery nodes (search history)")
|
||||
if web_count > 0:
|
||||
print(f" - {web_count} WebResult nodes (cached web results)")
|
||||
if other_count > 0:
|
||||
print(f" - {other_count} orphaned entity nodes")
|
||||
|
||||
print()
|
||||
print(" Document nodes will NOT be deleted (they represent wiki pages)")
|
||||
print()
|
||||
|
||||
response = input("Continue with cleanup? (yes/no): ").strip().lower()
|
||||
if response != "yes":
|
||||
print("❌ Cancelled")
|
||||
return
|
||||
|
||||
print()
|
||||
print("🚀 Starting cleanup...")
|
||||
print()
|
||||
|
||||
# Delete SearchQuery nodes
|
||||
if search_count > 0:
|
||||
deleted_sq = delete_search_queries()
|
||||
print(f" ✅ Deleted {deleted_sq} SearchQuery nodes")
|
||||
|
||||
# Delete WebResult nodes
|
||||
if web_count > 0:
|
||||
deleted_wr = delete_web_results()
|
||||
print(f" ✅ Deleted {deleted_wr} WebResult nodes")
|
||||
|
||||
# Delete orphaned nodes
|
||||
if other_count > 0:
|
||||
deleted_orphans = delete_orphaned_nodes()
|
||||
print(f" ✅ Deleted {deleted_orphans} orphaned nodes")
|
||||
|
||||
print()
|
||||
print("=" * 70)
|
||||
print("📊 Final Database State")
|
||||
print("=" * 70)
|
||||
|
||||
# Show final state
|
||||
search_count_final = count_nodes("SearchQuery")
|
||||
web_count_final = count_nodes("WebResult")
|
||||
doc_count_final = count_nodes("Document")
|
||||
other_count_final = count_nodes("Other")
|
||||
|
||||
print(f" SearchQuery nodes: {search_count_final}")
|
||||
print(f" WebResult nodes: {web_count_final}")
|
||||
print(f" Document nodes: {doc_count_final}")
|
||||
print(f" Other entity nodes: {other_count_final}")
|
||||
print()
|
||||
|
||||
if search_count_final == 0 and web_count_final == 0 and other_count_final == 0:
|
||||
print("✨ Graph database is now production-ready!")
|
||||
print(" Only legitimate wiki-derived data remains.")
|
||||
else:
|
||||
print("⚠️ Some nodes remain - this might be expected if you have active wiki pages")
|
||||
|
||||
if __name__ == "__main__":
|
||||
try:
|
||||
main()
|
||||
except KeyboardInterrupt:
|
||||
print("\n\n❌ Interrupted by user")
|
||||
sys.exit(1)
|
||||
except Exception as e:
|
||||
print(f"\n\n❌ Error: {e}")
|
||||
import traceback
|
||||
traceback.print_exc()
|
||||
sys.exit(1)
|
||||
@@ -1,126 +0,0 @@
|
||||
#!/usr/bin/env python3
|
||||
"""
|
||||
Cleanup script to delete all wiki pages via REST API.
|
||||
This will trigger automatic vector and graph cleanup for each page.
|
||||
"""
|
||||
import requests
|
||||
import time
|
||||
import sys
|
||||
|
||||
# Configuration
|
||||
BASE_URL = "http://192.168.86.149:8089"
|
||||
API_KEY = "af88ed8f44bed81bdb20d0534f1c4547340b29e2aba4963f61a71b993d7eb6e5"
|
||||
USER = "jpmschweitzer"
|
||||
|
||||
headers = {
|
||||
"Authorization": f"Bearer {API_KEY}",
|
||||
"Content-Type": "application/json"
|
||||
}
|
||||
|
||||
def list_all_pages():
|
||||
"""List all wiki pages."""
|
||||
url = f"{BASE_URL}/wiki/pages"
|
||||
params = {"user": USER, "limit": 200}
|
||||
|
||||
print(f"🔍 Listing all pages for user: {USER}")
|
||||
response = requests.get(url, headers=headers, params=params)
|
||||
|
||||
if response.status_code != 200:
|
||||
print(f"❌ Failed to list pages: {response.status_code}")
|
||||
print(response.text)
|
||||
return []
|
||||
|
||||
data = response.json()
|
||||
pages = data.get("pages", [])
|
||||
print(f"📚 Found {len(pages)} pages")
|
||||
return pages
|
||||
|
||||
def delete_page(page_id, title):
|
||||
"""Delete a single wiki page."""
|
||||
url = f"{BASE_URL}/wiki/pages/{page_id}"
|
||||
params = {"user": USER}
|
||||
|
||||
print(f"🗑️ Deleting page {page_id}: {title}")
|
||||
response = requests.delete(url, headers=headers, params=params)
|
||||
|
||||
if response.status_code == 200:
|
||||
print(f" ✅ Deleted successfully")
|
||||
return True
|
||||
else:
|
||||
print(f" ❌ Failed: {response.status_code}")
|
||||
print(f" {response.text}")
|
||||
return False
|
||||
|
||||
def main():
|
||||
print("=" * 60)
|
||||
print("🧹 Wiki Cleanup Script")
|
||||
print("=" * 60)
|
||||
print()
|
||||
|
||||
# List all pages
|
||||
pages = list_all_pages()
|
||||
|
||||
if not pages:
|
||||
print("✨ No pages found. Wiki is already clean!")
|
||||
return
|
||||
|
||||
print()
|
||||
print(f"⚠️ WARNING: About to delete {len(pages)} pages!")
|
||||
print(" This will also delete:")
|
||||
print(" - Vector embeddings in Qdrant")
|
||||
print(" - Document nodes in Neo4j")
|
||||
print()
|
||||
|
||||
# Ask for confirmation
|
||||
response = input("Continue? (yes/no): ").strip().lower()
|
||||
if response != "yes":
|
||||
print("❌ Cancelled")
|
||||
return
|
||||
|
||||
print()
|
||||
print("🚀 Starting deletion...")
|
||||
print()
|
||||
|
||||
# Delete each page
|
||||
success_count = 0
|
||||
failed_count = 0
|
||||
|
||||
for i, page in enumerate(pages, 1):
|
||||
page_id = page.get("id")
|
||||
title = page.get("title", "Untitled")
|
||||
|
||||
print(f"[{i}/{len(pages)}] ", end="")
|
||||
|
||||
if delete_page(page_id, title):
|
||||
success_count += 1
|
||||
else:
|
||||
failed_count += 1
|
||||
|
||||
# Small delay to avoid overwhelming the API
|
||||
time.sleep(0.2)
|
||||
|
||||
print()
|
||||
print("=" * 60)
|
||||
print("📊 Cleanup Summary")
|
||||
print("=" * 60)
|
||||
print(f"✅ Successfully deleted: {success_count} pages")
|
||||
print(f"❌ Failed to delete: {failed_count} pages")
|
||||
print(f"📝 Total pages: {len(pages)}")
|
||||
print()
|
||||
|
||||
if success_count > 0:
|
||||
print("🧹 Background cleanup tasks are running to remove:")
|
||||
print(" - Vector embeddings from Qdrant")
|
||||
print(" - Document nodes from Neo4j")
|
||||
print()
|
||||
print("✨ Wiki cleanup complete!")
|
||||
|
||||
if __name__ == "__main__":
|
||||
try:
|
||||
main()
|
||||
except KeyboardInterrupt:
|
||||
print("\n\n❌ Interrupted by user")
|
||||
sys.exit(1)
|
||||
except Exception as e:
|
||||
print(f"\n\n❌ Error: {e}")
|
||||
sys.exit(1)
|
||||
@@ -1,469 +0,0 @@
|
||||
# Knowledge Consolidation (Librarian Task)
|
||||
|
||||
Automated processing of HybridRAG search results to consolidate knowledge into wiki pages.
|
||||
|
||||
## Overview
|
||||
|
||||
The **Knowledge Consolidation** system is a Librarian task that:
|
||||
1. Processes unprocessed `SearchQuery` nodes from Neo4j (created by HybridRAG searches)
|
||||
2. Analyzes web results with Ollama (mistral-nemo) for novel information
|
||||
3. Creates new wiki pages for new concepts/technologies
|
||||
4. Updates existing pages with new facts and citations
|
||||
5. Adds new entities to the knowledge graph
|
||||
6. Marks `SearchQuery` nodes as processed
|
||||
|
||||
**Architecture:**
|
||||
- **Library Desk**: Exposes `/consolidate/knowledge` REST endpoint
|
||||
- **Scheduler**: Calls endpoint periodically via generic `rest_api_executor`
|
||||
- **Separation of concerns**: Scheduler triggers, Library Desk executes
|
||||
|
||||
---
|
||||
|
||||
## Endpoint: POST /consolidate/knowledge
|
||||
|
||||
**URL:** `http://library-desk:8089/consolidate/knowledge`
|
||||
|
||||
**Authentication:** Bearer token (LIBRARY_DESK_API_KEY)
|
||||
|
||||
**Request Body:**
|
||||
```json
|
||||
{
|
||||
"process_limit": 10, // Max searches to process per run (1-100)
|
||||
"lookback_days": 7, // Only process searches from last N days (1-90)
|
||||
"min_web_results": 2, // Minimum web results required (1-20)
|
||||
"dry_run": false // If true, analyze but don't create pages
|
||||
}
|
||||
```
|
||||
|
||||
**Response:**
|
||||
```json
|
||||
{
|
||||
"total_found": 5,
|
||||
"processed_count": 4,
|
||||
"pages_created": 2,
|
||||
"pages_updated": 3,
|
||||
"entities_added": 7,
|
||||
"errors": ["Search abc123: Failed to parse response"],
|
||||
"results": [
|
||||
{
|
||||
"search_id": "uuid-1",
|
||||
"query": "docker orchestration kubernetes",
|
||||
"pages_created": 1,
|
||||
"pages_updated": 1,
|
||||
"entities_added": 3,
|
||||
"error": null
|
||||
}
|
||||
],
|
||||
"dry_run": false
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Manual Testing
|
||||
|
||||
### Test the endpoint directly:
|
||||
|
||||
```bash
|
||||
# Set API key
|
||||
export LIBRARY_DESK_API_KEY="your-api-key-here"
|
||||
|
||||
# Test with dry run (no changes)
|
||||
curl -X POST http://192.168.86.149:8089/consolidate/knowledge \
|
||||
-H "Authorization: Bearer $LIBRARY_DESK_API_KEY" \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"process_limit": 5,
|
||||
"lookback_days": 7,
|
||||
"min_web_results": 2,
|
||||
"dry_run": true
|
||||
}' | python3 -m json.tool
|
||||
|
||||
# Real run (creates pages)
|
||||
curl -X POST http://192.168.86.149:8089/consolidate/knowledge \
|
||||
-H "Authorization: Bearer $LIBRARY_DESK_API_KEY" \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"process_limit": 10,
|
||||
"lookback_days": 7,
|
||||
"min_web_results": 2,
|
||||
"dry_run": false
|
||||
}' | python3 -m json.tool
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Scheduler Integration
|
||||
|
||||
### Step 1: Set Environment Variable
|
||||
|
||||
The Scheduler needs the Library Desk API key to authenticate:
|
||||
|
||||
```bash
|
||||
# Add to Scheduler's .env or docker-compose
|
||||
LIBRARY_DESK_API_KEY=your-library-desk-api-key-here
|
||||
```
|
||||
|
||||
### Step 2: Register Task via Scheduler API
|
||||
|
||||
Create the scheduled task using the Scheduler's REST API:
|
||||
|
||||
```bash
|
||||
# Set Scheduler API key
|
||||
export SCHEDULER_API_KEY="your-scheduler-api-key"
|
||||
|
||||
# Create knowledge consolidation task (runs hourly)
|
||||
curl -X POST http://192.168.86.149:8090/tasks \
|
||||
-H "Authorization: Bearer $SCHEDULER_API_KEY" \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"task_name": "knowledge_consolidation",
|
||||
"service": "library-desk",
|
||||
"executor": "rest_api_executor",
|
||||
"priority": 50,
|
||||
"minute": 0,
|
||||
"hour": -1,
|
||||
"day_of_month": -1,
|
||||
"month": -1,
|
||||
"day_of_week": -1,
|
||||
"enabled": true,
|
||||
"description": "Hourly Librarian knowledge consolidation from HybridRAG search results",
|
||||
"config": {
|
||||
"url": "http://library-desk:8089/consolidate/knowledge",
|
||||
"method": "POST",
|
||||
"payload": {
|
||||
"process_limit": 10,
|
||||
"lookback_days": 7,
|
||||
"min_web_results": 2,
|
||||
"dry_run": false
|
||||
},
|
||||
"auth": {
|
||||
"type": "bearer",
|
||||
"token": "${LIBRARY_DESK_API_KEY}"
|
||||
},
|
||||
"timeout": 300,
|
||||
"response_path": "processed_count"
|
||||
}
|
||||
}' | python3 -m json.tool
|
||||
```
|
||||
|
||||
**Schedule Patterns:**
|
||||
|
||||
```bash
|
||||
# Every hour at minute 0
|
||||
"minute": 0, "hour": -1, ...
|
||||
|
||||
# Every 4 hours at minute 15
|
||||
"minute": 15, "hour": [0, 4, 8, 12, 16, 20], ...
|
||||
|
||||
# Daily at 3:00 AM
|
||||
"minute": 0, "hour": 3, "day_of_month": -1, ...
|
||||
|
||||
# Every Monday at 9:00 AM
|
||||
"minute": 0, "hour": 9, "day_of_week": 0, ...
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Generic REST API Executor
|
||||
|
||||
The Scheduler's `rest_api_executor` is a universal executor for calling any REST API across the system.
|
||||
|
||||
### Features:
|
||||
- **HTTP Methods**: GET, POST, PUT, DELETE, PATCH
|
||||
- **Authentication**: Bearer token, Basic auth, API key
|
||||
- **Environment Variables**: Use `${VAR_NAME}` for secrets
|
||||
- **Configurable timeouts and SSL verification**
|
||||
- **JSON payload support**
|
||||
- **Response extraction via JSONPath**
|
||||
|
||||
### Example Configurations:
|
||||
|
||||
**1. Library Desk Knowledge Consolidation:**
|
||||
```json
|
||||
{
|
||||
"executor": "rest_api_executor",
|
||||
"config": {
|
||||
"url": "http://library-desk:8089/consolidate/knowledge",
|
||||
"method": "POST",
|
||||
"payload": {"process_limit": 10, "lookback_days": 7},
|
||||
"auth": {"type": "bearer", "token": "${LIBRARY_DESK_API_KEY}"}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**2. Core API Container Restart:**
|
||||
```json
|
||||
{
|
||||
"executor": "rest_api_executor",
|
||||
"config": {
|
||||
"url": "http://core-api:8088/v1/infrastructure/containers/nginx/restart",
|
||||
"method": "POST",
|
||||
"auth": {"type": "bearer", "token": "${CORE_API_KEY}"}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**3. External Webhook (Slack notification):**
|
||||
```json
|
||||
{
|
||||
"executor": "rest_api_executor",
|
||||
"config": {
|
||||
"url": "https://hooks.slack.com/services/YOUR/WEBHOOK/URL",
|
||||
"method": "POST",
|
||||
"payload": {"text": "Daily backup completed"},
|
||||
"verify_ssl": true
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**4. Basic Auth Example:**
|
||||
```json
|
||||
{
|
||||
"executor": "rest_api_executor",
|
||||
"config": {
|
||||
"url": "http://internal-service:8080/api/sync",
|
||||
"method": "GET",
|
||||
"auth": {
|
||||
"type": "basic",
|
||||
"username": "${SERVICE_USER}",
|
||||
"password": "${SERVICE_PASSWORD}"
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Managing the Task
|
||||
|
||||
### View Task Status:
|
||||
```bash
|
||||
curl -H "Authorization: Bearer $SCHEDULER_API_KEY" \
|
||||
http://192.168.86.149:8090/tasks/knowledge_consolidation | python3 -m json.tool
|
||||
```
|
||||
|
||||
### View Execution History:
|
||||
```bash
|
||||
curl -H "Authorization: Bearer $SCHEDULER_API_KEY" \
|
||||
"http://192.168.86.149:8090/executions?task_name=knowledge_consolidation&limit=10" \
|
||||
| python3 -m json.tool
|
||||
```
|
||||
|
||||
### Manually Trigger Task:
|
||||
```bash
|
||||
curl -X POST \
|
||||
-H "Authorization: Bearer $SCHEDULER_API_KEY" \
|
||||
http://192.168.86.149:8090/tasks/knowledge_consolidation/trigger
|
||||
```
|
||||
|
||||
### Update Task Configuration:
|
||||
```bash
|
||||
curl -X PUT \
|
||||
-H "Authorization: Bearer $SCHEDULER_API_KEY" \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"enabled": false,
|
||||
"config": {
|
||||
"payload": {
|
||||
"process_limit": 20,
|
||||
"dry_run": true
|
||||
}
|
||||
}
|
||||
}' \
|
||||
http://192.168.86.149:8090/tasks/knowledge_consolidation
|
||||
```
|
||||
|
||||
### Delete Task:
|
||||
```bash
|
||||
curl -X DELETE \
|
||||
-H "Authorization: Bearer $SCHEDULER_API_KEY" \
|
||||
http://192.168.86.149:8090/tasks/knowledge_consolidation
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## How It Works
|
||||
|
||||
### 1. HybridRAG Search Creates SearchQuery Nodes
|
||||
|
||||
When a user performs a HybridRAG search (e.g., "docker orchestration kubernetes"):
|
||||
- Library Desk stores search in Neo4j as `SearchQuery` node
|
||||
- Web results stored as `WebResult` nodes with `FOUND` relationships
|
||||
- `processed: false` flag indicates it needs consolidation
|
||||
|
||||
### 2. Scheduler Triggers Consolidation (Hourly)
|
||||
|
||||
The Scheduler runs the task at configured intervals:
|
||||
- Calls Library Desk `/consolidate/knowledge` endpoint via HTTP
|
||||
- Passes configuration (process_limit, lookback_days, etc.)
|
||||
- Authenticates with Bearer token from environment
|
||||
|
||||
### 3. Library Desk Processes Searches
|
||||
|
||||
For each unprocessed `SearchQuery`:
|
||||
1. **Query Neo4j** for searches with `processed: false` from last N days
|
||||
2. **Retrieve web results** via `FOUND` relationships
|
||||
3. **Analyze with Ollama** (mistral-nemo) for novel information:
|
||||
- Extract new concepts/technologies worth documenting
|
||||
- Identify facts to add to existing pages
|
||||
- Find new entities for knowledge graph
|
||||
4. **Create/update wiki pages** (TODO: implement wiki API)
|
||||
5. **Update knowledge graph** with new entities
|
||||
6. **Mark SearchQuery as processed** (`processed: true, processed_at: timestamp`)
|
||||
|
||||
### 4. Response Returned to Scheduler
|
||||
|
||||
Library Desk returns summary:
|
||||
- Total searches found and processed
|
||||
- Pages created/updated
|
||||
- Entities added
|
||||
- Errors encountered
|
||||
|
||||
Scheduler logs the result and marks execution as complete.
|
||||
|
||||
---
|
||||
|
||||
## Future Enhancements
|
||||
|
||||
### 1. Wiki Page Creation
|
||||
Currently logs what would be created. Need to implement:
|
||||
- Wiki.js API integration for page creation
|
||||
- Template-based page generation
|
||||
- Citation formatting from web sources
|
||||
|
||||
### 2. Enhanced Entity Extraction
|
||||
- Use spaCy or similar NLP for entity recognition
|
||||
- Link entities between documents
|
||||
- Build entity relationship graph
|
||||
|
||||
### 3. Smart Consolidation
|
||||
- Detect duplicate/similar pages
|
||||
- Merge related content
|
||||
- Suggest tag improvements
|
||||
|
||||
### 4. User Feedback Loop
|
||||
- Flag low-confidence consolidations for review
|
||||
- Allow users to approve/reject suggestions
|
||||
- Learn from user feedback
|
||||
|
||||
### 5. Metrics and Monitoring
|
||||
- Track consolidation success rates
|
||||
- Measure wiki growth over time
|
||||
- Identify knowledge gaps
|
||||
|
||||
---
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
### Task not running?
|
||||
|
||||
1. **Check task is enabled:**
|
||||
```bash
|
||||
curl -H "Authorization: Bearer $SCHEDULER_API_KEY" \
|
||||
http://192.168.86.149:8090/tasks/knowledge_consolidation
|
||||
```
|
||||
|
||||
2. **Check environment variable:**
|
||||
```bash
|
||||
docker exec scheduler env | grep LIBRARY_DESK_API_KEY
|
||||
```
|
||||
|
||||
3. **Check Scheduler logs:**
|
||||
```bash
|
||||
docker logs scheduler --tail 50
|
||||
```
|
||||
|
||||
### Task failing?
|
||||
|
||||
1. **Check execution history:**
|
||||
```bash
|
||||
curl -H "Authorization: Bearer $SCHEDULER_API_KEY" \
|
||||
"http://192.168.86.149:8090/executions?task_name=knowledge_consolidation&limit=5"
|
||||
```
|
||||
|
||||
2. **Test endpoint directly:**
|
||||
```bash
|
||||
curl -X POST http://192.168.86.149:8089/consolidate/knowledge \
|
||||
-H "Authorization: Bearer $LIBRARY_DESK_API_KEY" \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{"process_limit": 1, "dry_run": true}'
|
||||
```
|
||||
|
||||
3. **Check Library Desk logs:**
|
||||
```bash
|
||||
docker logs library-desk --tail 100 | grep -i consolidation
|
||||
```
|
||||
|
||||
### No searches being processed?
|
||||
|
||||
Searches may already be processed or too old:
|
||||
- Increase `lookback_days` to process older searches
|
||||
- Lower `min_web_results` threshold
|
||||
- Check Neo4j for unprocessed searches:
|
||||
```cypher
|
||||
MATCH (sq:SearchQuery {processed: false})
|
||||
RETURN sq.query, sq.timestamp, sq.web_count
|
||||
ORDER BY sq.timestamp DESC
|
||||
LIMIT 10
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Example Workflow
|
||||
|
||||
**1. User searches:** "What is Docker Swarm orchestration?"
|
||||
|
||||
**2. HybridRAG creates SearchQuery:**
|
||||
```cypher
|
||||
CREATE (sq:SearchQuery {
|
||||
id: "uuid-123",
|
||||
query: "What is Docker Swarm orchestration?",
|
||||
user: "jpmschweitzer",
|
||||
timestamp: datetime(),
|
||||
processed: false,
|
||||
web_count: 5
|
||||
})
|
||||
// + WebResult nodes for each web search result
|
||||
```
|
||||
|
||||
**3. Scheduler triggers consolidation (hourly):**
|
||||
```
|
||||
[2025-01-09 14:00:00] Scheduler: Running knowledge_consolidation task
|
||||
[2025-01-09 14:00:01] REST API Executor: POST http://library-desk:8089/consolidate/knowledge
|
||||
```
|
||||
|
||||
**4. Library Desk processes:**
|
||||
```
|
||||
[2025-01-09 14:00:02] Found 3 unprocessed searches
|
||||
[2025-01-09 14:00:03] Processing: "What is Docker Swarm orchestration?"
|
||||
[2025-01-09 14:00:04] Retrieved 5 web results
|
||||
[2025-01-09 14:00:08] LLM analysis: Novel info found
|
||||
[2025-01-09 14:00:08] Would create page: "Docker Swarm Orchestration"
|
||||
[2025-01-09 14:00:08] Marked search uuid-123 as processed
|
||||
[2025-01-09 14:00:10] Consolidation complete: 3/3 searches, 2 pages created
|
||||
```
|
||||
|
||||
**5. Scheduler logs result:**
|
||||
```
|
||||
[2025-01-09 14:00:10] Task knowledge_consolidation succeeded: Processed 3/3 searches. Created 2 pages, updated 1 pages, added 5 entities.
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Summary
|
||||
|
||||
The Knowledge Consolidation system provides:
|
||||
- ✅ **Automated knowledge capture** from HybridRAG searches
|
||||
- ✅ **Separation of concerns** via REST API
|
||||
- ✅ **Generic REST executor** reusable across system
|
||||
- ✅ **Flexible scheduling** via Scheduler
|
||||
- ✅ **Comprehensive logging** and error handling
|
||||
- ✅ **Dry run mode** for testing
|
||||
- ⏳ **Wiki integration** (TODO)
|
||||
- ⏳ **Advanced entity extraction** (TODO)
|
||||
|
||||
For more information:
|
||||
- Scheduler: `/services/scheduler/README.md`
|
||||
- Library Desk: `/services/library-desk/README.md`
|
||||
- HybridRAG: `/services/library-desk/docs/HYBRID_RAG.md`
|
||||
@@ -1,107 +0,0 @@
|
||||
# [Topic Title]
|
||||
|
||||
> **Last Updated:** YYYY-MM-DD | **Status:** Draft/Active/Archived
|
||||
> **Tags:** #tag1 #tag2 #tag3
|
||||
|
||||
## Executive Summary
|
||||
|
||||
A concise 2-3 sentence overview of the topic. This should capture the essence and primary purpose for both AI and human readers.
|
||||
|
||||
**Key Facts:**
|
||||
- Most important fact #1
|
||||
- Most important fact #2
|
||||
- Most important fact #3
|
||||
|
||||
---
|
||||
|
||||
## Overview
|
||||
|
||||
Detailed introduction to the topic. Explain what it is, why it matters, and its context within the broader ecosystem.
|
||||
|
||||
## Core Concepts
|
||||
|
||||
### Concept 1: [Name]
|
||||
|
||||
Explanation of the first core concept.
|
||||
|
||||
### Concept 2: [Name]
|
||||
|
||||
Explanation of the second core concept.
|
||||
|
||||
## Technical Details
|
||||
|
||||
### Architecture
|
||||
|
||||
Description of how the system/concept is structured.
|
||||
|
||||
```
|
||||
[Diagram or code block if applicable]
|
||||
```
|
||||
|
||||
### Specifications
|
||||
|
||||
| Property | Value | Notes |
|
||||
|----------|-------|-------|
|
||||
| Property 1 | Value 1 | Additional context |
|
||||
| Property 2 | Value 2 | Additional context |
|
||||
|
||||
## Use Cases
|
||||
|
||||
### Primary Use Case
|
||||
|
||||
Description and example.
|
||||
|
||||
### Secondary Use Cases
|
||||
|
||||
- Use case 1
|
||||
- Use case 2
|
||||
|
||||
## Best Practices
|
||||
|
||||
1. **Practice 1**: Description
|
||||
2. **Practice 2**: Description
|
||||
|
||||
## Common Issues & Solutions
|
||||
|
||||
| Issue | Solution | Reference |
|
||||
|-------|----------|-----------|
|
||||
| Problem description | How to resolve | [Link] |
|
||||
|
||||
## Related Topics
|
||||
|
||||
- [Related Topic 1](link) - Brief description
|
||||
- [Related Topic 2](link) - Brief description
|
||||
|
||||
## Changes & Updates
|
||||
|
||||
### 2025-01-09
|
||||
- Updated section X with new information from [Source]
|
||||
- Corrected fact Y (previously stated Z)
|
||||
|
||||
---
|
||||
|
||||
## Sources
|
||||
|
||||
1. [Source Title](URL) - Publication Name, Date
|
||||
2. [Source Title](URL) - Publication Name, Date
|
||||
3. [Source Title](URL) - Publication Name, Date
|
||||
|
||||
## Knowledge Graph
|
||||
|
||||
**Related Entities:**
|
||||
- [Entity 1](graph-link) - Relationship description
|
||||
- [Entity 2](graph-link) - Relationship description
|
||||
|
||||
**Mentioned In:**
|
||||
- [Document 1](wiki-link) - Context
|
||||
- [Document 2](wiki-link) - Context
|
||||
|
||||
## Mind Map
|
||||
|
||||
View this topic in context: [Mind Map Link](mindmap-url)
|
||||
|
||||
---
|
||||
|
||||
*This page follows the Library Desk Wiki Standard v1.0*
|
||||
*Generated/Updated by: [Human/Librarian Agent]*
|
||||
*Quality Score: [0-100] | Completeness: [0-100%]*
|
||||
@@ -1,30 +0,0 @@
|
||||
[pytest]
|
||||
# Pytest configuration for Library Desk
|
||||
|
||||
# Test discovery
|
||||
python_files = test_*.py
|
||||
python_classes = Test*
|
||||
python_functions = test_*
|
||||
|
||||
# Asyncio settings
|
||||
asyncio_mode = auto
|
||||
asyncio_default_fixture_loop_scope = function
|
||||
|
||||
# Output options
|
||||
addopts =
|
||||
-v
|
||||
--strict-markers
|
||||
--tb=short
|
||||
--disable-warnings
|
||||
|
||||
# Markers for test categorization
|
||||
markers =
|
||||
unit: Unit tests (no external dependencies)
|
||||
integration: Integration tests (require services)
|
||||
slow: Slow tests that take more than 1 second
|
||||
|
||||
# Test paths
|
||||
testpaths = tests
|
||||
|
||||
# Minimum Python version
|
||||
minversion = 3.12
|
||||
@@ -1,30 +0,0 @@
|
||||
# Requires Python 3.12+
|
||||
|
||||
# FastAPI Framework (latest Dec 2024)
|
||||
fastapi~=0.115.0
|
||||
uvicorn[standard]~=0.32.0
|
||||
pydantic~=2.10.0
|
||||
pydantic-settings~=2.6.0
|
||||
|
||||
# HTTP Client (no known CVEs)
|
||||
httpx~=0.27.0
|
||||
|
||||
# Database & Vector Store
|
||||
neo4j~=6.0.3
|
||||
qdrant-client~=1.16.1
|
||||
asyncpg~=0.29.0 # PostgreSQL async driver for Wiki.js change detection
|
||||
|
||||
# Redis (Python 3.12 compatible)
|
||||
redis~=7.1.0
|
||||
|
||||
# Security
|
||||
python-jose[cryptography]~=3.3.0
|
||||
passlib[bcrypt]~=1.7.4
|
||||
python-multipart~=0.0.20
|
||||
|
||||
# Utilities
|
||||
python-dateutil~=2.9.0
|
||||
|
||||
# Testing
|
||||
pytest~=8.3.0
|
||||
pytest-asyncio~=0.24.0
|
||||
@@ -1,91 +0,0 @@
|
||||
#!/usr/bin/env python3
|
||||
"""
|
||||
Setup PostgreSQL triggers in Wiki.js database for change detection.
|
||||
|
||||
This script creates database triggers that emit NOTIFY events when
|
||||
pages are created, updated, or deleted in Wiki.js.
|
||||
|
||||
Run this once to enable automatic processing of user-edited pages.
|
||||
"""
|
||||
import asyncio
|
||||
import asyncpg
|
||||
import sys
|
||||
from src.config import get_settings
|
||||
|
||||
# Import the corrected trigger SQL from the listener module
|
||||
from src.services.wiki_change_listener import SETUP_TRIGGERS_SQL
|
||||
|
||||
|
||||
async def setup_triggers():
|
||||
"""Install database triggers in Wiki.js PostgreSQL database."""
|
||||
settings = get_settings()
|
||||
|
||||
print("=" * 70)
|
||||
print("Wiki.js Database Trigger Setup")
|
||||
print("=" * 70)
|
||||
print()
|
||||
print(f"Connecting to Wiki.js database at {settings.wikijs_db_host}:{settings.wikijs_db_port}")
|
||||
print(f"Database: {settings.wikijs_db_name}")
|
||||
print(f"User: {settings.wikijs_db_user}")
|
||||
print()
|
||||
|
||||
try:
|
||||
# Connect to Wiki.js database
|
||||
connection = await asyncpg.connect(
|
||||
host=settings.wikijs_db_host,
|
||||
port=settings.wikijs_db_port,
|
||||
user=settings.wikijs_db_user,
|
||||
password=settings.wikijs_db_password,
|
||||
database=settings.wikijs_db_name
|
||||
)
|
||||
|
||||
print("✓ Connected to Wiki.js database")
|
||||
print()
|
||||
print("Installing triggers...")
|
||||
print()
|
||||
|
||||
# Execute setup SQL
|
||||
await connection.execute(SETUP_TRIGGERS_SQL)
|
||||
|
||||
# Verify triggers were created
|
||||
triggers = await connection.fetch("""
|
||||
SELECT trigger_name, event_manipulation
|
||||
FROM information_schema.triggers
|
||||
WHERE event_object_table = 'pages'
|
||||
ORDER BY trigger_name
|
||||
""")
|
||||
|
||||
if triggers:
|
||||
print("✓ Triggers installed successfully:")
|
||||
print()
|
||||
for trigger in triggers:
|
||||
print(f" - {trigger['trigger_name']} ({trigger['event_manipulation']})")
|
||||
print()
|
||||
print("=" * 70)
|
||||
print("Setup complete!")
|
||||
print("=" * 70)
|
||||
print()
|
||||
print("Next steps:")
|
||||
print(" 1. Restart library-desk service to activate the listener")
|
||||
print(" 2. Edit a page in Wiki.js to test")
|
||||
print(" 3. Check library-desk logs for processing messages")
|
||||
print()
|
||||
else:
|
||||
print("✗ No triggers found after installation")
|
||||
sys.exit(1)
|
||||
|
||||
await connection.close()
|
||||
|
||||
except Exception as e:
|
||||
print(f"✗ Error: {e}")
|
||||
print()
|
||||
print("Common issues:")
|
||||
print(" - Check database connection settings in .env")
|
||||
print(" - Ensure database user has CREATE FUNCTION and CREATE TRIGGER permissions")
|
||||
print(" - Verify Wiki.js database is accessible from library-desk service")
|
||||
print()
|
||||
sys.exit(1)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
asyncio.run(setup_triggers())
|
||||
@@ -1,12 +0,0 @@
|
||||
"""
|
||||
Library Desk - FastAPI Coordination Service
|
||||
|
||||
The coordination layer for The Library system, providing:
|
||||
- HybridRAG queries (Neo4j + Qdrant + SearXNG)
|
||||
- Document ingestion and indexing
|
||||
- Entity extraction and relationship mapping
|
||||
- Mind map generation
|
||||
- Wiki.js API proxy
|
||||
"""
|
||||
|
||||
__version__ = "1.0.0"
|
||||
@@ -1,413 +0,0 @@
|
||||
"""
|
||||
Neo4j async client for Library Desk.
|
||||
|
||||
Provides async Neo4j operations with:
|
||||
- Connection pooling via AsyncGraphDatabase
|
||||
- Session management with context managers
|
||||
- Multi-tenancy support via user labels
|
||||
- Automatic retry on transient failures
|
||||
"""
|
||||
|
||||
from neo4j import AsyncGraphDatabase, AsyncDriver, AsyncSession
|
||||
from typing import Optional, List, Dict, Any
|
||||
import logging
|
||||
|
||||
from src.core.multi_tenancy import get_neo4j_user_label
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
|
||||
class Neo4jClient:
|
||||
"""
|
||||
Async Neo4j client with connection pooling.
|
||||
|
||||
Features:
|
||||
- Singleton driver instance (expensive to create)
|
||||
- Session-per-request pattern (lightweight)
|
||||
- Automatic transaction retry
|
||||
- Multi-tenancy via user-specific labels
|
||||
"""
|
||||
|
||||
def __init__(self, uri: str, user: str, password: str):
|
||||
"""
|
||||
Initialize Neo4j client.
|
||||
|
||||
Args:
|
||||
uri: Neo4j Bolt URI (e.g., "bolt://neo4j:7687")
|
||||
user: Neo4j username
|
||||
password: Neo4j password
|
||||
"""
|
||||
self.uri = uri
|
||||
self._driver: Optional[AsyncDriver] = None
|
||||
self._auth = (user, password)
|
||||
|
||||
async def connect(self):
|
||||
"""
|
||||
Initialize connection pool.
|
||||
|
||||
Should be called once at app startup.
|
||||
Driver handles connection pooling internally.
|
||||
"""
|
||||
if not self._driver:
|
||||
self._driver = AsyncGraphDatabase.driver(
|
||||
self.uri,
|
||||
auth=self._auth,
|
||||
max_connection_pool_size=50,
|
||||
connection_timeout=30.0,
|
||||
max_transaction_retry_time=30.0
|
||||
)
|
||||
# Verify connectivity
|
||||
await self._driver.verify_connectivity()
|
||||
logger.info(f"Connected to Neo4j at {self.uri}")
|
||||
|
||||
async def close(self):
|
||||
"""
|
||||
Close connection pool.
|
||||
|
||||
Should be called once at app shutdown.
|
||||
"""
|
||||
if self._driver:
|
||||
await self._driver.close()
|
||||
self._driver = None
|
||||
logger.info("Closed Neo4j connection")
|
||||
|
||||
async def execute_query(
|
||||
self,
|
||||
cypher: str,
|
||||
parameters: Optional[Dict[str, Any]] = None
|
||||
) -> List[Dict[str, Any]]:
|
||||
"""
|
||||
Execute Cypher query and return results.
|
||||
|
||||
Args:
|
||||
cypher: Cypher query string
|
||||
parameters: Query parameters
|
||||
|
||||
Returns:
|
||||
List of result records as dictionaries
|
||||
|
||||
Raises:
|
||||
Exception: If driver not initialized or query fails
|
||||
"""
|
||||
if not self._driver:
|
||||
await self.connect()
|
||||
|
||||
async with self._driver.session() as session:
|
||||
result = await session.run(cypher, parameters or {})
|
||||
records = await result.data()
|
||||
return records
|
||||
|
||||
async def execute_write(
|
||||
self,
|
||||
cypher: str,
|
||||
parameters: Optional[Dict[str, Any]] = None
|
||||
) -> List[Dict[str, Any]]:
|
||||
"""
|
||||
Execute write transaction with automatic retry.
|
||||
|
||||
Args:
|
||||
cypher: Cypher query string
|
||||
parameters: Query parameters
|
||||
|
||||
Returns:
|
||||
List of result records as dictionaries
|
||||
"""
|
||||
if not self._driver:
|
||||
await self.connect()
|
||||
|
||||
async def write_tx(tx):
|
||||
result = await tx.run(cypher, parameters or {})
|
||||
return await result.data()
|
||||
|
||||
async with self._driver.session() as session:
|
||||
return await session.execute_write(write_tx)
|
||||
|
||||
# Multi-tenancy helpers
|
||||
|
||||
def get_user_label(self, user: str) -> str:
|
||||
"""
|
||||
Get Neo4j label for user's documents.
|
||||
|
||||
Args:
|
||||
user: User identifier
|
||||
|
||||
Returns:
|
||||
Neo4j label string (e.g., "User_Jpmschweitzer_Document")
|
||||
"""
|
||||
return get_neo4j_user_label(user)
|
||||
|
||||
# Document node operations
|
||||
|
||||
async def create_document_node(
|
||||
self,
|
||||
user: str,
|
||||
doc_id: str,
|
||||
properties: Dict[str, Any]
|
||||
) -> Optional[Dict[str, Any]]:
|
||||
"""
|
||||
Create document node with user label.
|
||||
|
||||
Node structure:
|
||||
(doc:Document:User_{user}_Document {
|
||||
id: "doc_123",
|
||||
source: "github",
|
||||
repository: "anthropic-cookbook",
|
||||
path: "skills/citation/guide.md",
|
||||
title: "Citation Guide",
|
||||
created_at: timestamp(),
|
||||
updated_at: timestamp(),
|
||||
content_hash: "sha256:..."
|
||||
})
|
||||
|
||||
Args:
|
||||
user: User identifier
|
||||
doc_id: Unique document ID
|
||||
properties: Document properties
|
||||
|
||||
Returns:
|
||||
Created node properties or None on failure
|
||||
"""
|
||||
user_label = self.get_user_label(user)
|
||||
|
||||
# Ensure required properties
|
||||
properties["id"] = doc_id
|
||||
if "created_at" not in properties:
|
||||
properties["created_at"] = "timestamp()"
|
||||
|
||||
cypher = f"""
|
||||
CREATE (doc:Document:{user_label})
|
||||
SET doc = $properties
|
||||
SET doc.created_at = timestamp()
|
||||
SET doc.updated_at = timestamp()
|
||||
RETURN doc
|
||||
"""
|
||||
|
||||
try:
|
||||
result = await self.execute_write(cypher, {"properties": properties})
|
||||
return result[0]["doc"] if result else None
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to create document node: {e}", exc_info=True)
|
||||
return None
|
||||
|
||||
async def get_document_node(
|
||||
self,
|
||||
user: str,
|
||||
doc_id: str
|
||||
) -> Optional[Dict[str, Any]]:
|
||||
"""
|
||||
Get document node by ID.
|
||||
|
||||
Args:
|
||||
user: User identifier
|
||||
doc_id: Document ID
|
||||
|
||||
Returns:
|
||||
Document node properties or None if not found
|
||||
"""
|
||||
user_label = self.get_user_label(user)
|
||||
|
||||
cypher = f"""
|
||||
MATCH (doc:Document:{user_label} {{id: $doc_id}})
|
||||
RETURN doc
|
||||
"""
|
||||
|
||||
try:
|
||||
result = await self.execute_query(cypher, {"doc_id": doc_id})
|
||||
return result[0]["doc"] if result else None
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to get document node: {e}", exc_info=True)
|
||||
return None
|
||||
|
||||
async def delete_document_node(
|
||||
self,
|
||||
user: str,
|
||||
doc_id: str
|
||||
) -> bool:
|
||||
"""
|
||||
Delete document node and all its relationships.
|
||||
|
||||
Args:
|
||||
user: User identifier
|
||||
doc_id: Document ID
|
||||
|
||||
Returns:
|
||||
True if deleted, False otherwise
|
||||
"""
|
||||
user_label = self.get_user_label(user)
|
||||
|
||||
cypher = f"""
|
||||
MATCH (doc:Document:{user_label} {{id: $doc_id}})
|
||||
DETACH DELETE doc
|
||||
RETURN count(doc) as deleted
|
||||
"""
|
||||
|
||||
try:
|
||||
result = await self.execute_write(cypher, {"doc_id": doc_id})
|
||||
return result[0]["deleted"] > 0 if result else False
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to delete document node: {e}", exc_info=True)
|
||||
return False
|
||||
|
||||
async def find_similar_documents(
|
||||
self,
|
||||
user: str,
|
||||
doc_ids: List[str],
|
||||
max_depth: int = 2,
|
||||
limit: int = 20
|
||||
) -> List[Dict[str, Any]]:
|
||||
"""
|
||||
Find documents similar to given docs via graph traversal.
|
||||
|
||||
Uses: Shared concepts, shared entities, citation links.
|
||||
|
||||
Args:
|
||||
user: User identifier
|
||||
doc_ids: List of source document IDs
|
||||
max_depth: Maximum traversal depth
|
||||
limit: Maximum results to return
|
||||
|
||||
Returns:
|
||||
List of similar documents with connection strength
|
||||
"""
|
||||
user_label = self.get_user_label(user)
|
||||
|
||||
cypher = f"""
|
||||
MATCH (source:Document:{user_label})
|
||||
WHERE source.id IN $doc_ids
|
||||
MATCH (source)-[*1..{max_depth}]-(related:Document:{user_label})
|
||||
WHERE related.id <> source.id AND NOT related.id IN $doc_ids
|
||||
WITH related, count(*) as connection_strength
|
||||
ORDER BY connection_strength DESC
|
||||
LIMIT $limit
|
||||
RETURN related, connection_strength
|
||||
"""
|
||||
|
||||
try:
|
||||
result = await self.execute_query(
|
||||
cypher,
|
||||
{"doc_ids": doc_ids, "limit": limit}
|
||||
)
|
||||
return result
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to find similar documents: {e}", exc_info=True)
|
||||
return []
|
||||
|
||||
async def list_user_documents(
|
||||
self,
|
||||
user: str,
|
||||
limit: int = 100,
|
||||
offset: int = 0
|
||||
) -> List[Dict[str, Any]]:
|
||||
"""
|
||||
List all documents for a user.
|
||||
|
||||
Args:
|
||||
user: User identifier
|
||||
limit: Maximum results to return
|
||||
offset: Number of results to skip
|
||||
|
||||
Returns:
|
||||
List of document nodes
|
||||
"""
|
||||
user_label = self.get_user_label(user)
|
||||
|
||||
cypher = f"""
|
||||
MATCH (doc:Document:{user_label})
|
||||
RETURN doc
|
||||
ORDER BY doc.created_at DESC
|
||||
SKIP $offset
|
||||
LIMIT $limit
|
||||
"""
|
||||
|
||||
try:
|
||||
result = await self.execute_query(
|
||||
cypher,
|
||||
{"offset": offset, "limit": limit}
|
||||
)
|
||||
return [r["doc"] for r in result]
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to list documents: {e}", exc_info=True)
|
||||
return []
|
||||
|
||||
# Concept/entity operations
|
||||
|
||||
async def create_concept_node(
|
||||
self,
|
||||
concept_name: str,
|
||||
concept_type: str,
|
||||
properties: Optional[Dict[str, Any]] = None
|
||||
) -> Optional[Dict[str, Any]]:
|
||||
"""
|
||||
Create or update concept node.
|
||||
|
||||
Args:
|
||||
concept_name: Concept name
|
||||
concept_type: Concept type (Technique, Tool, Pattern, etc.)
|
||||
properties: Additional properties
|
||||
|
||||
Returns:
|
||||
Concept node properties
|
||||
"""
|
||||
cypher = """
|
||||
MERGE (concept:Concept {name: $name})
|
||||
ON CREATE SET
|
||||
concept.type = $type,
|
||||
concept.first_seen = timestamp(),
|
||||
concept.mention_count = 1
|
||||
ON MATCH SET
|
||||
concept.mention_count = concept.mention_count + 1
|
||||
SET concept += $properties
|
||||
RETURN concept
|
||||
"""
|
||||
|
||||
try:
|
||||
result = await self.execute_write(
|
||||
cypher,
|
||||
{
|
||||
"name": concept_name,
|
||||
"type": concept_type,
|
||||
"properties": properties or {}
|
||||
}
|
||||
)
|
||||
return result[0]["concept"] if result else None
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to create concept node: {e}", exc_info=True)
|
||||
return None
|
||||
|
||||
async def link_document_to_concept(
|
||||
self,
|
||||
user: str,
|
||||
doc_id: str,
|
||||
concept_name: str
|
||||
) -> bool:
|
||||
"""
|
||||
Create MENTIONS relationship between document and concept.
|
||||
|
||||
Args:
|
||||
user: User identifier
|
||||
doc_id: Document ID
|
||||
concept_name: Concept name
|
||||
|
||||
Returns:
|
||||
True if link created
|
||||
"""
|
||||
user_label = self.get_user_label(user)
|
||||
|
||||
cypher = f"""
|
||||
MATCH (doc:Document:{user_label} {{id: $doc_id}})
|
||||
MATCH (concept:Concept {{name: $concept_name}})
|
||||
MERGE (doc)-[r:MENTIONS]->(concept)
|
||||
ON CREATE SET r.count = 1
|
||||
ON MATCH SET r.count = r.count + 1
|
||||
RETURN r
|
||||
"""
|
||||
|
||||
try:
|
||||
result = await self.execute_write(
|
||||
cypher,
|
||||
{"doc_id": doc_id, "concept_name": concept_name}
|
||||
)
|
||||
return len(result) > 0
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to link document to concept: {e}", exc_info=True)
|
||||
return False
|
||||
@@ -1,386 +0,0 @@
|
||||
"""
|
||||
Ollama client for embeddings generation.
|
||||
|
||||
Provides async embedding operations via Ollama API:
|
||||
- Text embedding generation
|
||||
- Batch embedding support
|
||||
- Model management
|
||||
"""
|
||||
|
||||
import httpx
|
||||
from typing import List, Dict, Any, Optional
|
||||
import logging
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
|
||||
class OllamaClient:
|
||||
"""
|
||||
Ollama API client for embeddings.
|
||||
|
||||
Documentation: https://github.com/ollama/ollama/blob/main/docs/api.md
|
||||
Default model: nomic-embed-text (768-dimensional embeddings)
|
||||
"""
|
||||
|
||||
def __init__(self, base_url: str, model: str = "nomic-embed-text"):
|
||||
"""
|
||||
Initialize Ollama client.
|
||||
|
||||
Args:
|
||||
base_url: Ollama server URL (e.g., "http://ollama:11434")
|
||||
model: Embedding model name (default: "nomic-embed-text")
|
||||
"""
|
||||
self.base_url = base_url.rstrip("/")
|
||||
self.model = model
|
||||
self.embeddings_url = f"{self.base_url}/api/embeddings"
|
||||
self.generate_url = f"{self.base_url}/api/generate"
|
||||
self.tags_url = f"{self.base_url}/api/tags"
|
||||
self.client = httpx.AsyncClient(timeout=120.0) # Embeddings can be slow
|
||||
logger.info(f"Initialized Ollama client: {base_url}, model: {model}")
|
||||
|
||||
async def close(self):
|
||||
"""Close HTTP client"""
|
||||
await self.client.aclose()
|
||||
|
||||
async def embed(self, text: str) -> Optional[List[float]]:
|
||||
"""
|
||||
Generate embedding for single text.
|
||||
|
||||
Args:
|
||||
text: Text to embed
|
||||
|
||||
Returns:
|
||||
Embedding vector (768-dimensional for nomic-embed-text) or None on failure
|
||||
|
||||
Example:
|
||||
>>> embedding = await client.embed("Hello world")
|
||||
>>> len(embedding)
|
||||
768
|
||||
"""
|
||||
try:
|
||||
payload = {
|
||||
"model": self.model,
|
||||
"prompt": text
|
||||
}
|
||||
|
||||
response = await self.client.post(
|
||||
self.embeddings_url,
|
||||
json=payload
|
||||
)
|
||||
response.raise_for_status()
|
||||
data = response.json()
|
||||
|
||||
embedding = data.get("embedding")
|
||||
if not embedding:
|
||||
logger.error(f"No embedding in response: {data}")
|
||||
return None
|
||||
|
||||
return embedding
|
||||
|
||||
except httpx.HTTPStatusError as e:
|
||||
logger.error(f"HTTP error: {e.response.status_code} - {e.response.text}")
|
||||
return None
|
||||
except Exception as e:
|
||||
logger.error(f"Embedding failed: {e}", exc_info=True)
|
||||
return None
|
||||
|
||||
async def embed_batch(
|
||||
self,
|
||||
texts: List[str],
|
||||
show_progress: bool = False
|
||||
) -> List[Optional[List[float]]]:
|
||||
"""
|
||||
Generate embeddings for multiple texts.
|
||||
|
||||
Args:
|
||||
texts: List of texts to embed
|
||||
show_progress: Log progress for large batches
|
||||
|
||||
Returns:
|
||||
List of embedding vectors (same order as input)
|
||||
None entries for texts that failed to embed
|
||||
|
||||
Example:
|
||||
>>> texts = ["Hello", "World", "Test"]
|
||||
>>> embeddings = await client.embed_batch(texts)
|
||||
>>> len(embeddings)
|
||||
3
|
||||
"""
|
||||
embeddings = []
|
||||
|
||||
for i, text in enumerate(texts):
|
||||
if show_progress and i % 10 == 0:
|
||||
logger.info(f"Embedding progress: {i}/{len(texts)}")
|
||||
|
||||
embedding = await self.embed(text)
|
||||
embeddings.append(embedding)
|
||||
|
||||
if show_progress:
|
||||
logger.info(f"Embedding complete: {len(embeddings)}/{len(texts)}")
|
||||
|
||||
return embeddings
|
||||
|
||||
async def embed_batch_filtered(
|
||||
self,
|
||||
texts: List[str],
|
||||
show_progress: bool = False
|
||||
) -> List[List[float]]:
|
||||
"""
|
||||
Generate embeddings for multiple texts, filtering out failures.
|
||||
|
||||
Args:
|
||||
texts: List of texts to embed
|
||||
show_progress: Log progress for large batches
|
||||
|
||||
Returns:
|
||||
List of successful embedding vectors (may be shorter than input)
|
||||
|
||||
Example:
|
||||
>>> texts = ["Hello", "World", "Test"]
|
||||
>>> embeddings = await client.embed_batch_filtered(texts)
|
||||
>>> all(e is not None for e in embeddings)
|
||||
True
|
||||
"""
|
||||
all_embeddings = await self.embed_batch(texts, show_progress)
|
||||
return [e for e in all_embeddings if e is not None]
|
||||
|
||||
async def embed_documents(
|
||||
self,
|
||||
documents: List[Dict[str, Any]],
|
||||
content_field: str = "content",
|
||||
show_progress: bool = False
|
||||
) -> List[Dict[str, Any]]:
|
||||
"""
|
||||
Embed documents with metadata preservation.
|
||||
|
||||
Args:
|
||||
documents: List of document dictionaries
|
||||
content_field: Field name containing text to embed
|
||||
show_progress: Log progress for large batches
|
||||
|
||||
Returns:
|
||||
List of documents with added "embedding" field
|
||||
|
||||
Example:
|
||||
>>> docs = [
|
||||
... {"content": "Hello world", "id": 1},
|
||||
... {"content": "Test doc", "id": 2}
|
||||
... ]
|
||||
>>> embedded = await client.embed_documents(docs)
|
||||
>>> "embedding" in embedded[0]
|
||||
True
|
||||
"""
|
||||
texts = [doc.get(content_field, "") for doc in documents]
|
||||
embeddings = await self.embed_batch(texts, show_progress)
|
||||
|
||||
results = []
|
||||
for doc, embedding in zip(documents, embeddings):
|
||||
doc_copy = doc.copy()
|
||||
doc_copy["embedding"] = embedding
|
||||
results.append(doc_copy)
|
||||
|
||||
return results
|
||||
|
||||
async def get_embedding_dimension(self) -> Optional[int]:
|
||||
"""
|
||||
Get embedding dimension for current model.
|
||||
|
||||
Returns:
|
||||
Embedding dimension (e.g., 768 for nomic-embed-text) or None on failure
|
||||
|
||||
Example:
|
||||
>>> dim = await client.get_embedding_dimension()
|
||||
>>> dim
|
||||
768
|
||||
"""
|
||||
# Generate a test embedding to determine dimension
|
||||
test_embedding = await self.embed("test")
|
||||
if test_embedding:
|
||||
return len(test_embedding)
|
||||
return None
|
||||
|
||||
async def list_models(self) -> List[Dict[str, Any]]:
|
||||
"""
|
||||
List available Ollama models.
|
||||
|
||||
Returns:
|
||||
List of model information dictionaries
|
||||
|
||||
Example:
|
||||
>>> models = await client.list_models()
|
||||
>>> any(m["name"] == "nomic-embed-text" for m in models)
|
||||
True
|
||||
"""
|
||||
try:
|
||||
response = await self.client.get(self.tags_url)
|
||||
response.raise_for_status()
|
||||
data = response.json()
|
||||
return data.get("models", [])
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to list models: {e}")
|
||||
return []
|
||||
|
||||
async def check_model_available(self, model_name: Optional[str] = None) -> bool:
|
||||
"""
|
||||
Check if a model is available.
|
||||
|
||||
Args:
|
||||
model_name: Model name to check (defaults to self.model)
|
||||
|
||||
Returns:
|
||||
True if model is available, False otherwise
|
||||
"""
|
||||
check_model = model_name or self.model
|
||||
models = await self.list_models()
|
||||
|
||||
# Check for exact match or match with :latest tag
|
||||
for m in models:
|
||||
name = m.get("name", "")
|
||||
# Exact match
|
||||
if name == check_model:
|
||||
return True
|
||||
# Match without tag (e.g., "nomic-embed-text" matches "nomic-embed-text:latest")
|
||||
if name.startswith(f"{check_model}:"):
|
||||
return True
|
||||
|
||||
return False
|
||||
|
||||
async def generate_text(
|
||||
self,
|
||||
prompt: str,
|
||||
model: Optional[str] = None,
|
||||
stream: bool = False
|
||||
) -> Optional[str]:
|
||||
"""
|
||||
Generate text completion (for non-embedding use cases).
|
||||
|
||||
Args:
|
||||
prompt: Input prompt
|
||||
model: Model name (defaults to self.model)
|
||||
stream: Enable streaming response
|
||||
|
||||
Returns:
|
||||
Generated text or None on failure
|
||||
|
||||
Note: This is primarily for debugging/testing. Use specialized
|
||||
LLM services for production text generation.
|
||||
"""
|
||||
try:
|
||||
payload = {
|
||||
"model": model or self.model,
|
||||
"prompt": prompt,
|
||||
"stream": stream
|
||||
}
|
||||
|
||||
response = await self.client.post(
|
||||
self.generate_url,
|
||||
json=payload
|
||||
)
|
||||
response.raise_for_status()
|
||||
|
||||
if stream:
|
||||
# For streaming, return first chunk
|
||||
# Full streaming implementation would need async generator
|
||||
return response.text
|
||||
else:
|
||||
data = response.json()
|
||||
return data.get("response")
|
||||
|
||||
except Exception as e:
|
||||
logger.error(f"Text generation failed: {e}")
|
||||
return None
|
||||
|
||||
async def health_check(self) -> bool:
|
||||
"""
|
||||
Check if Ollama server is reachable and model is available.
|
||||
|
||||
Returns:
|
||||
True if healthy, False otherwise
|
||||
"""
|
||||
try:
|
||||
# Check server is up and get models in one call
|
||||
response = await self.client.get(self.tags_url, timeout=5.0)
|
||||
response.raise_for_status()
|
||||
data = response.json()
|
||||
models = data.get("models", [])
|
||||
|
||||
# Check model is available
|
||||
check_model = self.model
|
||||
model_found = False
|
||||
for m in models:
|
||||
name = m.get("name", "")
|
||||
if name == check_model or name.startswith(f"{check_model}:"):
|
||||
model_found = True
|
||||
break
|
||||
|
||||
if not model_found:
|
||||
logger.warning(f"Model '{self.model}' not found in Ollama")
|
||||
return False
|
||||
|
||||
return True
|
||||
|
||||
except Exception as e:
|
||||
logger.error(f"Health check failed: {e}")
|
||||
return False
|
||||
|
||||
def estimate_tokens(self, text: str) -> int:
|
||||
"""
|
||||
Rough estimate of token count for text.
|
||||
|
||||
Uses simple heuristic: ~4 characters per token.
|
||||
|
||||
Args:
|
||||
text: Input text
|
||||
|
||||
Returns:
|
||||
Estimated token count
|
||||
"""
|
||||
return len(text) // 4
|
||||
|
||||
def chunk_text_for_embedding(
|
||||
self,
|
||||
text: str,
|
||||
max_tokens: int = 512,
|
||||
overlap: int = 50
|
||||
) -> List[str]:
|
||||
"""
|
||||
Chunk text into segments suitable for embedding.
|
||||
|
||||
Args:
|
||||
text: Input text
|
||||
max_tokens: Maximum tokens per chunk
|
||||
overlap: Token overlap between chunks
|
||||
|
||||
Returns:
|
||||
List of text chunks
|
||||
|
||||
Example:
|
||||
>>> chunks = client.chunk_text_for_embedding(long_text, max_tokens=512)
|
||||
>>> all(client.estimate_tokens(c) <= 512 for c in chunks)
|
||||
True
|
||||
"""
|
||||
# Convert tokens to approximate character count
|
||||
max_chars = max_tokens * 4
|
||||
overlap_chars = overlap * 4
|
||||
|
||||
if len(text) <= max_chars:
|
||||
return [text]
|
||||
|
||||
chunks = []
|
||||
start = 0
|
||||
|
||||
while start < len(text):
|
||||
end = start + max_chars
|
||||
chunk = text[start:end]
|
||||
|
||||
# Try to break at sentence boundary
|
||||
if end < len(text):
|
||||
last_period = chunk.rfind(". ")
|
||||
if last_period > max_chars * 0.5: # Only break if > 50% through chunk
|
||||
end = start + last_period + 1
|
||||
chunk = text[start:end]
|
||||
|
||||
chunks.append(chunk.strip())
|
||||
start = end - overlap_chars
|
||||
|
||||
return chunks
|
||||
@@ -1,588 +0,0 @@
|
||||
"""
|
||||
Qdrant client wrapper for Library Desk.
|
||||
|
||||
Provides async vector operations with:
|
||||
- Collection-per-user multi-tenancy
|
||||
- Document chunk storage with embeddings
|
||||
- Semantic search
|
||||
- Similarity queries
|
||||
"""
|
||||
|
||||
from qdrant_client import QdrantClient
|
||||
from qdrant_client.models import (
|
||||
Distance, VectorParams, PointStruct,
|
||||
Filter, FieldCondition, MatchValue
|
||||
)
|
||||
from typing import List, Dict, Any, Optional
|
||||
import uuid
|
||||
import logging
|
||||
|
||||
from src.core.multi_tenancy import get_qdrant_collection_name
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
|
||||
class QdrantClientWrapper:
|
||||
"""
|
||||
Qdrant client wrapper with multi-tenancy support.
|
||||
|
||||
Pattern: Collection per user (from qdrant_memory.py)
|
||||
Each user has isolated vector collection for their documents.
|
||||
"""
|
||||
|
||||
def __init__(self, url: str, embedding_dim: int = 768):
|
||||
"""
|
||||
Initialize Qdrant client.
|
||||
|
||||
Args:
|
||||
url: Qdrant server URL (e.g., "http://qdrant:6333")
|
||||
embedding_dim: Vector embedding dimension (default 768 for nomic-embed-text)
|
||||
"""
|
||||
self.client = QdrantClient(url=url)
|
||||
self.embedding_dim = embedding_dim
|
||||
logger.info(f"Initialized Qdrant client: {url}")
|
||||
|
||||
def get_collection_name(self, user: str) -> str:
|
||||
"""
|
||||
Get collection name for user.
|
||||
|
||||
Args:
|
||||
user: User identifier
|
||||
|
||||
Returns:
|
||||
Collection name (e.g., "library_desk_jpmschweitzer")
|
||||
"""
|
||||
return get_qdrant_collection_name(user)
|
||||
|
||||
async def ensure_collection(self, collection_name: str):
|
||||
"""
|
||||
Create collection if it doesn't exist.
|
||||
|
||||
Args:
|
||||
collection_name: Collection name
|
||||
"""
|
||||
try:
|
||||
collections = self.client.get_collections()
|
||||
existing = [c.name for c in collections.collections]
|
||||
|
||||
if collection_name not in existing:
|
||||
logger.info(f"Creating Qdrant collection: {collection_name}")
|
||||
self.client.create_collection(
|
||||
collection_name=collection_name,
|
||||
vectors_config=VectorParams(
|
||||
size=self.embedding_dim,
|
||||
distance=Distance.COSINE
|
||||
)
|
||||
)
|
||||
logger.info(f"Created collection: {collection_name}")
|
||||
except Exception as e:
|
||||
logger.error(f"Error ensuring collection: {e}", exc_info=True)
|
||||
raise
|
||||
|
||||
async def collection_exists(self, collection_name: str) -> bool:
|
||||
"""
|
||||
Check if collection exists.
|
||||
|
||||
Args:
|
||||
collection_name: Collection name
|
||||
|
||||
Returns:
|
||||
True if collection exists
|
||||
"""
|
||||
try:
|
||||
collections = self.client.get_collections()
|
||||
existing = [c.name for c in collections.collections]
|
||||
return collection_name in existing
|
||||
except Exception as e:
|
||||
logger.error(f"Error checking collection: {e}", exc_info=True)
|
||||
return False
|
||||
|
||||
async def upsert_document_chunks(
|
||||
self,
|
||||
user: str,
|
||||
doc_id: str,
|
||||
chunks: List[Dict[str, Any]],
|
||||
embeddings: List[List[float]]
|
||||
) -> int:
|
||||
"""
|
||||
Upsert document chunks with embeddings.
|
||||
|
||||
Point structure:
|
||||
{
|
||||
id: "doc_123_chunk_0",
|
||||
vector: [...],
|
||||
payload: {
|
||||
doc_id: "doc_123",
|
||||
chunk_index: 0,
|
||||
content: "text content",
|
||||
metadata: {...}
|
||||
}
|
||||
}
|
||||
|
||||
Args:
|
||||
user: User identifier
|
||||
doc_id: Document ID
|
||||
chunks: List of chunk dictionaries with content
|
||||
embeddings: List of embedding vectors
|
||||
|
||||
Returns:
|
||||
Number of chunks upserted
|
||||
|
||||
Raises:
|
||||
ValueError: If chunks and embeddings length mismatch
|
||||
"""
|
||||
if len(chunks) != len(embeddings):
|
||||
raise ValueError(
|
||||
f"Chunks ({len(chunks)}) and embeddings ({len(embeddings)}) length mismatch"
|
||||
)
|
||||
|
||||
collection_name = self.get_collection_name(user)
|
||||
await self.ensure_collection(user)
|
||||
|
||||
points = []
|
||||
for i, (chunk, embedding) in enumerate(zip(chunks, embeddings)):
|
||||
# Generate deterministic point ID
|
||||
point_id = str(uuid.uuid5(
|
||||
uuid.NAMESPACE_DNS,
|
||||
f"{doc_id}_chunk_{i}"
|
||||
))
|
||||
|
||||
points.append(PointStruct(
|
||||
id=point_id,
|
||||
vector=embedding,
|
||||
payload={
|
||||
"doc_id": doc_id,
|
||||
"chunk_index": i,
|
||||
"content": chunk.get("content", ""),
|
||||
"metadata": chunk.get("metadata", {})
|
||||
}
|
||||
))
|
||||
|
||||
try:
|
||||
self.client.upsert(
|
||||
collection_name=collection_name,
|
||||
points=points
|
||||
)
|
||||
logger.info(f"Upserted {len(points)} chunks for {doc_id}")
|
||||
return len(points)
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to upsert chunks: {e}", exc_info=True)
|
||||
raise
|
||||
|
||||
async def search(
|
||||
self,
|
||||
user: str,
|
||||
query_vector: List[float],
|
||||
limit: int = 10,
|
||||
score_threshold: float = 0.7,
|
||||
filter_dict: Optional[Dict[str, Any]] = None
|
||||
) -> List[Dict[str, Any]]:
|
||||
"""
|
||||
Semantic search in user's collection.
|
||||
|
||||
Args:
|
||||
user: User identifier
|
||||
query_vector: Query embedding vector
|
||||
limit: Maximum results to return
|
||||
score_threshold: Minimum similarity score (0.0-1.0)
|
||||
filter_dict: Optional payload filters
|
||||
|
||||
Returns:
|
||||
List of matches with scores and payloads
|
||||
|
||||
Example:
|
||||
results = await client.search(
|
||||
user="jpmschweitzer",
|
||||
query_vector=[0.1, 0.2, ...],
|
||||
limit=5,
|
||||
filter_dict={"doc_id": "doc_123"}
|
||||
)
|
||||
"""
|
||||
collection_name = self.get_collection_name(user)
|
||||
|
||||
# Build filter if provided
|
||||
query_filter = None
|
||||
if filter_dict:
|
||||
conditions = []
|
||||
for key, value in filter_dict.items():
|
||||
conditions.append(
|
||||
FieldCondition(key=key, match=MatchValue(value=value))
|
||||
)
|
||||
query_filter = Filter(must=conditions)
|
||||
|
||||
try:
|
||||
response = self.client.query_points(
|
||||
collection_name=collection_name,
|
||||
query=query_vector,
|
||||
limit=limit,
|
||||
score_threshold=score_threshold,
|
||||
query_filter=query_filter,
|
||||
with_payload=True
|
||||
)
|
||||
|
||||
return [
|
||||
{
|
||||
"id": str(point.id),
|
||||
"score": point.score,
|
||||
"doc_id": point.payload["doc_id"],
|
||||
"chunk_index": point.payload["chunk_index"],
|
||||
"content": point.payload["content"],
|
||||
"metadata": point.payload.get("metadata", {})
|
||||
}
|
||||
for point in response.points
|
||||
]
|
||||
except Exception as e:
|
||||
logger.error(f"Search failed: {e}", exc_info=True)
|
||||
return []
|
||||
|
||||
async def delete_document(self, user: str, doc_id: str) -> bool:
|
||||
"""
|
||||
Delete all chunks for a document.
|
||||
|
||||
Args:
|
||||
user: User identifier
|
||||
doc_id: Document ID
|
||||
|
||||
Returns:
|
||||
True if deleted successfully
|
||||
"""
|
||||
collection_name = self.get_collection_name(user)
|
||||
|
||||
try:
|
||||
self.client.delete(
|
||||
collection_name=collection_name,
|
||||
points_selector=Filter(
|
||||
must=[
|
||||
FieldCondition(
|
||||
key="doc_id",
|
||||
match=MatchValue(value=doc_id)
|
||||
)
|
||||
]
|
||||
)
|
||||
)
|
||||
logger.info(f"Deleted chunks for {doc_id}")
|
||||
return True
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to delete document: {e}", exc_info=True)
|
||||
return False
|
||||
|
||||
async def get_document_chunks(
|
||||
self,
|
||||
user: str,
|
||||
doc_id: str
|
||||
) -> List[Dict[str, Any]]:
|
||||
"""
|
||||
Get all chunks for a document.
|
||||
|
||||
Args:
|
||||
user: User identifier
|
||||
doc_id: Document ID
|
||||
|
||||
Returns:
|
||||
List of chunks with content and metadata
|
||||
"""
|
||||
collection_name = self.get_collection_name(user)
|
||||
|
||||
try:
|
||||
# Scroll through points with doc_id filter
|
||||
points, _ = self.client.scroll(
|
||||
collection_name=collection_name,
|
||||
scroll_filter=Filter(
|
||||
must=[
|
||||
FieldCondition(
|
||||
key="doc_id",
|
||||
match=MatchValue(value=doc_id)
|
||||
)
|
||||
]
|
||||
),
|
||||
limit=1000,
|
||||
with_payload=True,
|
||||
with_vectors=False
|
||||
)
|
||||
|
||||
return [
|
||||
{
|
||||
"id": str(point.id),
|
||||
"chunk_index": point.payload["chunk_index"],
|
||||
"content": point.payload["content"],
|
||||
"metadata": point.payload.get("metadata", {})
|
||||
}
|
||||
for point in points
|
||||
]
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to get document chunks: {e}", exc_info=True)
|
||||
return []
|
||||
|
||||
async def count_documents(self, user: str) -> int:
|
||||
"""
|
||||
Count total number of unique documents in user's collection.
|
||||
|
||||
Args:
|
||||
user: User identifier
|
||||
|
||||
Returns:
|
||||
Number of unique documents
|
||||
"""
|
||||
collection_name = self.get_collection_name(user)
|
||||
|
||||
try:
|
||||
# Get collection info
|
||||
collection_info = self.client.get_collection(collection_name)
|
||||
# This gives total points, not unique docs
|
||||
# For unique docs, would need to aggregate by doc_id
|
||||
return collection_info.points_count
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to count documents: {e}", exc_info=True)
|
||||
return 0
|
||||
|
||||
async def delete_collection(self, user: str) -> bool:
|
||||
"""
|
||||
Delete user's entire collection.
|
||||
|
||||
Warning: This removes all data for the user!
|
||||
|
||||
Args:
|
||||
user: User identifier
|
||||
|
||||
Returns:
|
||||
True if deleted successfully
|
||||
"""
|
||||
collection_name = self.get_collection_name(user)
|
||||
|
||||
try:
|
||||
self.client.delete_collection(collection_name)
|
||||
logger.warning(f"Deleted collection: {collection_name}")
|
||||
return True
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to delete collection: {e}", exc_info=True)
|
||||
return False
|
||||
|
||||
async def find_similar_chunks(
|
||||
self,
|
||||
user: str,
|
||||
doc_id: str,
|
||||
limit: int = 10
|
||||
) -> List[Dict[str, Any]]:
|
||||
"""
|
||||
Find chunks similar to those in a given document.
|
||||
|
||||
Strategy: Get all chunks from doc, use their vectors to find similar chunks.
|
||||
|
||||
Args:
|
||||
user: User identifier
|
||||
doc_id: Source document ID
|
||||
limit: Maximum results per chunk
|
||||
|
||||
Returns:
|
||||
List of similar chunks from other documents
|
||||
"""
|
||||
collection_name = self.get_collection_name(user)
|
||||
|
||||
try:
|
||||
# Get source document chunks with vectors
|
||||
source_points, _ = self.client.scroll(
|
||||
collection_name=collection_name,
|
||||
scroll_filter=Filter(
|
||||
must=[
|
||||
FieldCondition(
|
||||
key="doc_id",
|
||||
match=MatchValue(value=doc_id)
|
||||
)
|
||||
]
|
||||
),
|
||||
limit=10, # Sample first 10 chunks
|
||||
with_payload=True,
|
||||
with_vectors=True
|
||||
)
|
||||
|
||||
if not source_points:
|
||||
return []
|
||||
|
||||
# Search using first chunk's vector
|
||||
# (could aggregate multiple chunks for better results)
|
||||
first_vector = source_points[0].vector
|
||||
|
||||
response = self.client.query_points(
|
||||
collection_name=collection_name,
|
||||
query=first_vector,
|
||||
limit=limit * 2, # Get more to filter out same doc
|
||||
with_payload=True
|
||||
)
|
||||
|
||||
# Filter out chunks from same document
|
||||
similar = [
|
||||
{
|
||||
"id": str(point.id),
|
||||
"score": point.score,
|
||||
"doc_id": point.payload["doc_id"],
|
||||
"content": point.payload["content"]
|
||||
}
|
||||
for point in response.points
|
||||
if point.payload["doc_id"] != doc_id
|
||||
]
|
||||
|
||||
return similar[:limit]
|
||||
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to find similar chunks: {e}", exc_info=True)
|
||||
return []
|
||||
|
||||
async def upsert_vector(
|
||||
self,
|
||||
collection_name: str,
|
||||
vector_id: str,
|
||||
vector: List[float],
|
||||
payload: Dict[str, Any]
|
||||
) -> bool:
|
||||
"""
|
||||
Upsert a single vector point.
|
||||
|
||||
Args:
|
||||
collection_name: Collection name
|
||||
vector_id: Point ID
|
||||
vector: Embedding vector
|
||||
payload: Point payload/metadata
|
||||
|
||||
Returns:
|
||||
True if successful
|
||||
"""
|
||||
try:
|
||||
self.client.upsert(
|
||||
collection_name=collection_name,
|
||||
points=[PointStruct(
|
||||
id=vector_id,
|
||||
vector=vector,
|
||||
payload=payload
|
||||
)]
|
||||
)
|
||||
return True
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to upsert vector: {e}", exc_info=True)
|
||||
return False
|
||||
|
||||
async def delete_by_filter(
|
||||
self,
|
||||
collection_name: str,
|
||||
filter_conditions: Dict[str, Any]
|
||||
) -> int:
|
||||
"""
|
||||
Delete points matching filter conditions.
|
||||
|
||||
Args:
|
||||
collection_name: Collection name
|
||||
filter_conditions: Filter conditions (e.g., {"page_id": 5})
|
||||
|
||||
Returns:
|
||||
Number of points deleted (approximation)
|
||||
"""
|
||||
try:
|
||||
# Build filter
|
||||
conditions = []
|
||||
for key, value in filter_conditions.items():
|
||||
conditions.append(
|
||||
FieldCondition(key=key, match=MatchValue(value=value))
|
||||
)
|
||||
query_filter = Filter(must=conditions)
|
||||
|
||||
# Delete points
|
||||
result = self.client.delete(
|
||||
collection_name=collection_name,
|
||||
points_selector=query_filter
|
||||
)
|
||||
|
||||
# Return operation status (Qdrant doesn't return count directly)
|
||||
logger.info(f"Deleted points with filter {filter_conditions}")
|
||||
return 1 # Placeholder, actual count not available from API
|
||||
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to delete by filter: {e}", exc_info=True)
|
||||
return 0
|
||||
|
||||
async def search_vectors(
|
||||
self,
|
||||
collection_name: str,
|
||||
query_vector: List[float],
|
||||
limit: int = 10,
|
||||
score_threshold: float = 0.5,
|
||||
filter_conditions: Optional[Dict[str, Any]] = None
|
||||
) -> List[Dict[str, Any]]:
|
||||
"""
|
||||
Search vectors in collection.
|
||||
|
||||
Args:
|
||||
collection_name: Collection name
|
||||
query_vector: Query embedding vector
|
||||
limit: Maximum results
|
||||
score_threshold: Minimum similarity score
|
||||
filter_conditions: Optional filter conditions
|
||||
|
||||
Returns:
|
||||
List of search results with scores and payloads
|
||||
"""
|
||||
# Build filter if provided
|
||||
query_filter = None
|
||||
if filter_conditions:
|
||||
conditions = []
|
||||
for key, value in filter_conditions.items():
|
||||
conditions.append(
|
||||
FieldCondition(key=key, match=MatchValue(value=value))
|
||||
)
|
||||
query_filter = Filter(must=conditions)
|
||||
|
||||
try:
|
||||
response = self.client.query_points(
|
||||
collection_name=collection_name,
|
||||
query=query_vector,
|
||||
limit=limit,
|
||||
score_threshold=score_threshold,
|
||||
query_filter=query_filter,
|
||||
with_payload=True
|
||||
)
|
||||
|
||||
return [
|
||||
{
|
||||
"id": str(point.id),
|
||||
"score": point.score,
|
||||
"payload": dict(point.payload)
|
||||
}
|
||||
for point in response.points
|
||||
]
|
||||
except Exception as e:
|
||||
logger.error(f"Search failed: {e}", exc_info=True)
|
||||
return []
|
||||
|
||||
async def list_collections(self) -> List[Dict[str, Any]]:
|
||||
"""
|
||||
List all collections with stats.
|
||||
|
||||
Returns:
|
||||
List of collection info dictionaries
|
||||
"""
|
||||
try:
|
||||
collections = self.client.get_collections()
|
||||
result = []
|
||||
|
||||
for coll in collections.collections:
|
||||
# Get detailed collection info
|
||||
try:
|
||||
info = self.client.get_collection(coll.name)
|
||||
result.append({
|
||||
"name": coll.name,
|
||||
"vectors_count": info.vectors_count or 0,
|
||||
"points_count": info.points_count or 0,
|
||||
"segments_count": info.segments_count or 0
|
||||
})
|
||||
except Exception as e:
|
||||
logger.warning(f"Failed to get info for collection {coll.name}: {e}")
|
||||
result.append({
|
||||
"name": coll.name,
|
||||
"vectors_count": 0,
|
||||
"points_count": 0,
|
||||
"segments_count": 0
|
||||
})
|
||||
|
||||
return result
|
||||
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to list collections: {e}", exc_info=True)
|
||||
return []
|
||||
@@ -1,348 +0,0 @@
|
||||
"""
|
||||
SearXNG search client for Library Desk.
|
||||
|
||||
Provides async web search via SearXNG instance:
|
||||
- General web search
|
||||
- Category-specific search
|
||||
- Language filtering
|
||||
- Result formatting
|
||||
"""
|
||||
|
||||
import httpx
|
||||
from typing import Optional, List, Dict, Any
|
||||
import logging
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
|
||||
class SearXNGClient:
|
||||
"""
|
||||
SearXNG search API client.
|
||||
|
||||
Documentation: https://docs.searxng.org/dev/search_api.html
|
||||
API Format: JSON (?format=json)
|
||||
"""
|
||||
|
||||
def __init__(self, base_url: str):
|
||||
"""
|
||||
Initialize SearXNG client.
|
||||
|
||||
Args:
|
||||
base_url: SearXNG base URL (e.g., "http://searxng:8080")
|
||||
"""
|
||||
self.base_url = base_url.rstrip("/")
|
||||
self.search_url = f"{self.base_url}/search"
|
||||
self.client = httpx.AsyncClient(timeout=30.0)
|
||||
logger.info(f"Initialized SearXNG client: {base_url}")
|
||||
|
||||
async def close(self):
|
||||
"""Close HTTP client"""
|
||||
await self.client.aclose()
|
||||
|
||||
async def search(
|
||||
self,
|
||||
query: str,
|
||||
categories: Optional[List[str]] = None,
|
||||
language: str = "en",
|
||||
time_range: Optional[str] = None,
|
||||
safesearch: int = 0,
|
||||
pageno: int = 1,
|
||||
limit: int = 10
|
||||
) -> Dict[str, Any]:
|
||||
"""
|
||||
Search via SearXNG.
|
||||
|
||||
Args:
|
||||
query: Search query string
|
||||
categories: List of categories (e.g., ["general", "images", "news"])
|
||||
language: Language code (e.g., "en", "fr", "auto")
|
||||
time_range: Time range filter ("day", "week", "month", "year", None for all)
|
||||
safesearch: Safe search level (0=off, 1=moderate, 2=strict)
|
||||
pageno: Page number (starts at 1)
|
||||
limit: Maximum results to return
|
||||
|
||||
Returns:
|
||||
Search results dictionary with:
|
||||
- query: Original query
|
||||
- results: List of result items
|
||||
- number_of_results: Total results found
|
||||
- suggestions: Query suggestions
|
||||
|
||||
Raises:
|
||||
Exception: If search fails
|
||||
"""
|
||||
params = {
|
||||
"q": query,
|
||||
"format": "json",
|
||||
"language": language,
|
||||
"safesearch": safesearch,
|
||||
"pageno": pageno
|
||||
}
|
||||
|
||||
# Add optional parameters
|
||||
if categories:
|
||||
params["categories"] = ",".join(categories)
|
||||
if time_range:
|
||||
params["time_range"] = time_range
|
||||
|
||||
try:
|
||||
response = await self.client.get(self.search_url, params=params)
|
||||
response.raise_for_status()
|
||||
data = response.json()
|
||||
|
||||
# Limit results if requested
|
||||
results = data.get("results", [])
|
||||
if limit:
|
||||
results = results[:limit]
|
||||
|
||||
return {
|
||||
"query": data.get("query", query),
|
||||
"results": results,
|
||||
"number_of_results": data.get("number_of_results", 0),
|
||||
"suggestions": data.get("suggestions", []),
|
||||
"answers": data.get("answers", []),
|
||||
"infoboxes": data.get("infoboxes", [])
|
||||
}
|
||||
|
||||
except httpx.HTTPStatusError as e:
|
||||
logger.error(f"HTTP error: {e.response.status_code} - {e.response.text}")
|
||||
raise Exception(f"SearXNG search failed: {e.response.status_code}")
|
||||
except Exception as e:
|
||||
logger.error(f"Search failed: {e}", exc_info=True)
|
||||
raise
|
||||
|
||||
async def search_general(
|
||||
self,
|
||||
query: str,
|
||||
language: str = "en",
|
||||
limit: int = 10
|
||||
) -> List[Dict[str, Any]]:
|
||||
"""
|
||||
Simplified general web search.
|
||||
|
||||
Args:
|
||||
query: Search query
|
||||
language: Language code
|
||||
limit: Maximum results
|
||||
|
||||
Returns:
|
||||
List of search results with:
|
||||
- title: Page title
|
||||
- url: Page URL
|
||||
- content: Page description/snippet
|
||||
- engine: Search engine that provided result
|
||||
- score: Relevance score (if available)
|
||||
"""
|
||||
try:
|
||||
result = await self.search(
|
||||
query=query,
|
||||
categories=["general"],
|
||||
language=language,
|
||||
limit=limit
|
||||
)
|
||||
return result.get("results", [])
|
||||
except Exception as e:
|
||||
logger.error(f"General search failed: {e}")
|
||||
return []
|
||||
|
||||
async def search_with_context(
|
||||
self,
|
||||
query: str,
|
||||
context: Optional[str] = None,
|
||||
language: str = "en",
|
||||
limit: int = 10
|
||||
) -> List[Dict[str, Any]]:
|
||||
"""
|
||||
Search with additional context appended to query.
|
||||
|
||||
Useful for RAG queries where you want to enhance search with context.
|
||||
|
||||
Args:
|
||||
query: Primary search query
|
||||
context: Additional context to append
|
||||
language: Language code
|
||||
limit: Maximum results
|
||||
|
||||
Returns:
|
||||
List of search results
|
||||
"""
|
||||
enhanced_query = f"{query} {context}" if context else query
|
||||
return await self.search_general(enhanced_query, language, limit)
|
||||
|
||||
async def search_documentation(
|
||||
self,
|
||||
query: str,
|
||||
limit: int = 10
|
||||
) -> List[Dict[str, Any]]:
|
||||
"""
|
||||
Search for technical documentation.
|
||||
|
||||
Enhances query with "documentation" and filters for technical content.
|
||||
|
||||
Args:
|
||||
query: Search query (e.g., "FastAPI")
|
||||
limit: Maximum results
|
||||
|
||||
Returns:
|
||||
List of documentation search results
|
||||
"""
|
||||
enhanced_query = f"{query} documentation"
|
||||
return await self.search_general(enhanced_query, limit=limit)
|
||||
|
||||
async def search_recent(
|
||||
self,
|
||||
query: str,
|
||||
time_range: str = "month",
|
||||
limit: int = 10
|
||||
) -> List[Dict[str, Any]]:
|
||||
"""
|
||||
Search for recent content only.
|
||||
|
||||
Args:
|
||||
query: Search query
|
||||
time_range: Time range ("day", "week", "month", "year")
|
||||
limit: Maximum results
|
||||
|
||||
Returns:
|
||||
List of recent search results
|
||||
"""
|
||||
try:
|
||||
result = await self.search(
|
||||
query=query,
|
||||
categories=["general"],
|
||||
time_range=time_range,
|
||||
limit=limit
|
||||
)
|
||||
return result.get("results", [])
|
||||
except Exception as e:
|
||||
logger.error(f"Recent search failed: {e}")
|
||||
return []
|
||||
|
||||
async def get_suggestions(self, query: str) -> List[str]:
|
||||
"""
|
||||
Get search query suggestions.
|
||||
|
||||
Args:
|
||||
query: Partial query
|
||||
|
||||
Returns:
|
||||
List of suggested queries
|
||||
"""
|
||||
try:
|
||||
result = await self.search(query=query, limit=1)
|
||||
return result.get("suggestions", [])
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to get suggestions: {e}")
|
||||
return []
|
||||
|
||||
async def search_images(
|
||||
self,
|
||||
query: str,
|
||||
safesearch: int = 1,
|
||||
limit: int = 10
|
||||
) -> List[Dict[str, Any]]:
|
||||
"""
|
||||
Search for images.
|
||||
|
||||
Args:
|
||||
query: Search query
|
||||
safesearch: Safe search level (0=off, 1=moderate, 2=strict)
|
||||
limit: Maximum results
|
||||
|
||||
Returns:
|
||||
List of image results with:
|
||||
- title: Image title
|
||||
- url: Image URL
|
||||
- thumbnail_src: Thumbnail URL
|
||||
- img_src: Full image URL
|
||||
- content: Description
|
||||
"""
|
||||
try:
|
||||
result = await self.search(
|
||||
query=query,
|
||||
categories=["images"],
|
||||
safesearch=safesearch,
|
||||
limit=limit
|
||||
)
|
||||
return result.get("results", [])
|
||||
except Exception as e:
|
||||
logger.error(f"Image search failed: {e}")
|
||||
return []
|
||||
|
||||
async def search_news(
|
||||
self,
|
||||
query: str,
|
||||
time_range: str = "week",
|
||||
limit: int = 10
|
||||
) -> List[Dict[str, Any]]:
|
||||
"""
|
||||
Search for news articles.
|
||||
|
||||
Args:
|
||||
query: Search query
|
||||
time_range: Time range ("day", "week", "month")
|
||||
limit: Maximum results
|
||||
|
||||
Returns:
|
||||
List of news results
|
||||
"""
|
||||
try:
|
||||
result = await self.search(
|
||||
query=query,
|
||||
categories=["news"],
|
||||
time_range=time_range,
|
||||
limit=limit
|
||||
)
|
||||
return result.get("results", [])
|
||||
except Exception as e:
|
||||
logger.error(f"News search failed: {e}")
|
||||
return []
|
||||
|
||||
async def format_results_for_rag(
|
||||
self,
|
||||
results: List[Dict[str, Any]],
|
||||
max_snippet_length: int = 500
|
||||
) -> List[Dict[str, str]]:
|
||||
"""
|
||||
Format search results for RAG context.
|
||||
|
||||
Extracts relevant fields and truncates content.
|
||||
|
||||
Args:
|
||||
results: Raw search results
|
||||
max_snippet_length: Maximum length of content snippet
|
||||
|
||||
Returns:
|
||||
Formatted results with title, url, snippet
|
||||
"""
|
||||
formatted = []
|
||||
for result in results:
|
||||
content = result.get("content", "")
|
||||
if len(content) > max_snippet_length:
|
||||
content = content[:max_snippet_length] + "..."
|
||||
|
||||
formatted.append({
|
||||
"title": result.get("title", ""),
|
||||
"url": result.get("url", ""),
|
||||
"snippet": content,
|
||||
"engine": result.get("engine", ""),
|
||||
"score": result.get("score", 0.0)
|
||||
})
|
||||
|
||||
return formatted
|
||||
|
||||
async def health_check(self) -> bool:
|
||||
"""
|
||||
Simple health check - verify SearXNG is responding.
|
||||
|
||||
Returns:
|
||||
True if service is reachable, False otherwise
|
||||
"""
|
||||
try:
|
||||
# Just hit the base URL to check if service is up
|
||||
response = await self.client.get(self.base_url, timeout=5.0)
|
||||
# Accept any 2xx or 3xx status (redirects are ok)
|
||||
return response.status_code < 400
|
||||
except Exception as e:
|
||||
logger.error(f"SearXNG health check failed: {e}")
|
||||
return False
|
||||
@@ -1,697 +0,0 @@
|
||||
"""
|
||||
Wiki.js GraphQL client for Library Desk.
|
||||
|
||||
Provides async Wiki.js operations via GraphQL API:
|
||||
- Page CRUD (create, read, update, delete)
|
||||
- Search and listing
|
||||
- Tag management
|
||||
- Multi-tenancy via path namespaces
|
||||
"""
|
||||
|
||||
import httpx
|
||||
from typing import Optional, List, Dict, Any
|
||||
import logging
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
|
||||
class WikiJSClient:
|
||||
"""
|
||||
Wiki.js GraphQL API client.
|
||||
|
||||
Documentation: https://docs.requarks.io/dev/api
|
||||
Authentication: Username/password login to get user-specific JWT token
|
||||
"""
|
||||
|
||||
def __init__(self, base_url: str, username: str, password: str):
|
||||
"""
|
||||
Initialize Wiki.js client.
|
||||
|
||||
Args:
|
||||
base_url: Wiki.js base URL (e.g., "http://wiki:3000")
|
||||
username: Wiki.js username (e.g., "librarian@schweitz.net")
|
||||
password: Wiki.js password
|
||||
"""
|
||||
self.base_url = base_url.rstrip("/")
|
||||
self.graphql_url = f"{self.base_url}/graphql"
|
||||
self.username = username
|
||||
self.password = password
|
||||
self.jwt_token: Optional[str] = None
|
||||
self.client = httpx.AsyncClient(timeout=30.0)
|
||||
logger.info(f"Initialized Wiki.js client: {base_url} (user: {username})")
|
||||
|
||||
async def close(self):
|
||||
"""Close HTTP client"""
|
||||
await self.client.aclose()
|
||||
|
||||
async def login(self) -> bool:
|
||||
"""
|
||||
Authenticate with Wiki.js using username/password.
|
||||
|
||||
Returns:
|
||||
True if login successful, False otherwise
|
||||
"""
|
||||
login_mutation = """
|
||||
mutation Login($username: String!, $password: String!, $strategy: String!) {
|
||||
authentication {
|
||||
login(username: $username, password: $password, strategy: $strategy) {
|
||||
responseResult {
|
||||
succeeded
|
||||
errorCode
|
||||
message
|
||||
}
|
||||
jwt
|
||||
}
|
||||
}
|
||||
}
|
||||
"""
|
||||
|
||||
variables = {
|
||||
"username": self.username,
|
||||
"password": self.password,
|
||||
"strategy": "local"
|
||||
}
|
||||
|
||||
try:
|
||||
response = await self.client.post(
|
||||
self.graphql_url,
|
||||
headers={"Content-Type": "application/json"},
|
||||
json={"query": login_mutation, "variables": variables}
|
||||
)
|
||||
response.raise_for_status()
|
||||
result = response.json()
|
||||
|
||||
if "errors" in result:
|
||||
logger.error(f"Login failed: {result['errors']}")
|
||||
return False
|
||||
|
||||
login_result = result.get("data", {}).get("authentication", {}).get("login", {})
|
||||
response_result = login_result.get("responseResult", {})
|
||||
|
||||
if not response_result.get("succeeded"):
|
||||
logger.error(f"Login failed: {response_result.get('message')}")
|
||||
return False
|
||||
|
||||
self.jwt_token = login_result.get("jwt")
|
||||
if not self.jwt_token:
|
||||
logger.error("Login succeeded but no JWT token received")
|
||||
return False
|
||||
|
||||
logger.info(f"Successfully authenticated as {self.username}")
|
||||
return True
|
||||
|
||||
except Exception as e:
|
||||
logger.error(f"Login failed: {e}", exc_info=True)
|
||||
return False
|
||||
|
||||
async def _ensure_authenticated(self):
|
||||
"""Ensure we have a valid JWT token, login if needed."""
|
||||
if not self.jwt_token:
|
||||
success = await self.login()
|
||||
if not success:
|
||||
raise Exception("Failed to authenticate with Wiki.js")
|
||||
|
||||
async def _execute_query(
|
||||
self,
|
||||
query: str,
|
||||
variables: Optional[Dict] = None
|
||||
) -> Dict[str, Any]:
|
||||
"""
|
||||
Execute GraphQL query.
|
||||
|
||||
Args:
|
||||
query: GraphQL query string
|
||||
variables: Query variables
|
||||
|
||||
Returns:
|
||||
Query result data
|
||||
|
||||
Raises:
|
||||
Exception: If query fails or returns errors
|
||||
"""
|
||||
# Ensure we're authenticated before making requests
|
||||
await self._ensure_authenticated()
|
||||
|
||||
payload = {
|
||||
"query": query,
|
||||
"variables": variables or {}
|
||||
}
|
||||
|
||||
headers = {
|
||||
"Authorization": f"Bearer {self.jwt_token}",
|
||||
"Content-Type": "application/json"
|
||||
}
|
||||
|
||||
try:
|
||||
response = await self.client.post(
|
||||
self.graphql_url,
|
||||
headers=headers,
|
||||
json=payload
|
||||
)
|
||||
response.raise_for_status()
|
||||
result = response.json()
|
||||
|
||||
if "errors" in result:
|
||||
logger.error(f"GraphQL errors: {result['errors']}")
|
||||
raise Exception(f"GraphQL errors: {result['errors']}")
|
||||
|
||||
return result.get("data", {})
|
||||
|
||||
except httpx.HTTPStatusError as e:
|
||||
logger.error(f"HTTP error: {e.response.status_code} - {e.response.text}")
|
||||
raise
|
||||
except Exception as e:
|
||||
logger.error(f"GraphQL query failed: {e}", exc_info=True)
|
||||
raise
|
||||
|
||||
async def list_pages(
|
||||
self,
|
||||
path_prefix: str = "",
|
||||
tags: Optional[List[str]] = None,
|
||||
limit: int = 50
|
||||
) -> List[Dict[str, Any]]:
|
||||
"""
|
||||
List pages with optional filtering.
|
||||
|
||||
Multi-tenancy: Use path_prefix to filter by user namespace.
|
||||
|
||||
Args:
|
||||
path_prefix: Filter by path prefix (e.g., "/users/jpmschweitzer")
|
||||
tags: Filter by tags (e.g., ["projects"])
|
||||
limit: Maximum results
|
||||
|
||||
Returns:
|
||||
List of page objects
|
||||
"""
|
||||
query = """
|
||||
query ListPages($limit: Int, $orderBy: PageOrderBy) {
|
||||
pages {
|
||||
list(limit: $limit, orderBy: $orderBy) {
|
||||
id
|
||||
path
|
||||
title
|
||||
description
|
||||
tags
|
||||
createdAt
|
||||
updatedAt
|
||||
isPublished
|
||||
}
|
||||
}
|
||||
}
|
||||
"""
|
||||
|
||||
variables = {
|
||||
"limit": limit,
|
||||
"orderBy": "TITLE"
|
||||
}
|
||||
|
||||
data = await self._execute_query(query, variables)
|
||||
pages = data.get("pages", {}).get("list", [])
|
||||
|
||||
# Ensure tags is always a list
|
||||
for page in pages:
|
||||
if "tags" not in page or page["tags"] is None:
|
||||
page["tags"] = []
|
||||
|
||||
# Filter by path prefix (client-side if API doesn't support)
|
||||
if path_prefix:
|
||||
# Normalize paths to have leading slash for consistent comparison
|
||||
normalized_prefix = "/" + path_prefix.lstrip("/")
|
||||
pages = [
|
||||
p for p in pages
|
||||
if ("/" + p["path"].lstrip("/")).startswith(normalized_prefix)
|
||||
]
|
||||
|
||||
# Filter by tags (dossiers)
|
||||
if tags:
|
||||
pages = [
|
||||
p for p in pages
|
||||
if any(tag in (p.get("tags") or []) for tag in tags)
|
||||
]
|
||||
|
||||
return pages
|
||||
|
||||
async def list_all_pages(
|
||||
self,
|
||||
path_prefix: str = "",
|
||||
tags: Optional[List[str]] = None,
|
||||
batch_size: int = 100
|
||||
) -> List[Dict[str, Any]]:
|
||||
"""
|
||||
List ALL pages with pagination support.
|
||||
|
||||
Fetches pages in batches until all are retrieved.
|
||||
|
||||
Args:
|
||||
path_prefix: Filter by path prefix (e.g., "users/jpmschweitzer")
|
||||
tags: Filter by tags
|
||||
batch_size: Number of pages per batch (max 100)
|
||||
|
||||
Returns:
|
||||
Complete list of page objects
|
||||
"""
|
||||
all_pages = []
|
||||
offset = 0
|
||||
|
||||
while True:
|
||||
# Wiki.js list doesn't support offset, but limit is enough
|
||||
# since we filter client-side by path_prefix
|
||||
# Just fetch a large batch
|
||||
pages = await self.list_pages(
|
||||
path_prefix=path_prefix,
|
||||
tags=tags,
|
||||
limit=1000 # Fetch up to 1000 at once
|
||||
)
|
||||
|
||||
if not pages:
|
||||
break
|
||||
|
||||
all_pages = pages
|
||||
break # Wiki.js list doesn't paginate, so one call is enough
|
||||
|
||||
logger.info(f"list_all_pages: found {len(all_pages)} pages (prefix: {path_prefix or 'all'})")
|
||||
return all_pages
|
||||
|
||||
async def get_taxonomy_structure(
|
||||
self,
|
||||
user_namespace: str,
|
||||
limit: int = 500
|
||||
) -> Dict[str, List[str]]:
|
||||
"""
|
||||
Get the existing taxonomy structure for a user namespace.
|
||||
|
||||
Returns a dict mapping top-level categories to their subcategories.
|
||||
This helps the LLM choose existing paths rather than creating new ones.
|
||||
|
||||
Args:
|
||||
user_namespace: User namespace (e.g., "users/jpmschweitzer")
|
||||
limit: Maximum pages to fetch
|
||||
|
||||
Returns:
|
||||
Dict like {"reference": ["political-entities", "tech"], "places": ["the-netherlands"]}
|
||||
"""
|
||||
from collections import defaultdict
|
||||
|
||||
pages = await self.list_pages(path_prefix=user_namespace, limit=limit)
|
||||
|
||||
# Extract structure: category -> set of subcategories
|
||||
structure: Dict[str, set] = defaultdict(set)
|
||||
prefix = user_namespace.strip("/") + "/"
|
||||
|
||||
for page in pages:
|
||||
path = page.get("path", "")
|
||||
if not path.startswith(prefix):
|
||||
continue
|
||||
|
||||
# Get relative path within user namespace
|
||||
rel_path = path[len(prefix):]
|
||||
parts = rel_path.split("/")
|
||||
|
||||
if len(parts) >= 1:
|
||||
category = parts[0]
|
||||
if len(parts) >= 2:
|
||||
# Has subcategory (e.g., reference/political-entities/nato)
|
||||
structure[category].add(parts[1])
|
||||
else:
|
||||
# Just ensure category exists in structure
|
||||
structure[category] # Access to ensure key exists
|
||||
|
||||
# Convert sets to sorted lists
|
||||
return {cat: sorted(list(subs)) for cat, subs in sorted(structure.items())}
|
||||
|
||||
async def get_page(self, page_id: int) -> Optional[Dict[str, Any]]:
|
||||
"""
|
||||
Get single page by ID.
|
||||
|
||||
Args:
|
||||
page_id: Page ID
|
||||
|
||||
Returns:
|
||||
Page object or None if not found
|
||||
"""
|
||||
query = """
|
||||
query GetPage($id: Int!) {
|
||||
pages {
|
||||
single(id: $id) {
|
||||
id
|
||||
path
|
||||
title
|
||||
description
|
||||
content
|
||||
tags {
|
||||
tag
|
||||
}
|
||||
createdAt
|
||||
updatedAt
|
||||
isPublished
|
||||
editor
|
||||
}
|
||||
}
|
||||
}
|
||||
"""
|
||||
|
||||
try:
|
||||
# Ensure page_id is int (GraphQL requires Int, not String)
|
||||
data = await self._execute_query(query, {"id": int(page_id)})
|
||||
page = data.get("pages", {}).get("single")
|
||||
|
||||
# Extract tag strings from tag objects
|
||||
if page and "tags" in page and page["tags"]:
|
||||
page["tags"] = [t["tag"] for t in page["tags"]]
|
||||
elif page:
|
||||
page["tags"] = []
|
||||
|
||||
return page
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to get page {page_id}: {e}")
|
||||
return None
|
||||
|
||||
async def create_page(
|
||||
self,
|
||||
path: str,
|
||||
title: str,
|
||||
content: str,
|
||||
description: str = "",
|
||||
tags: Optional[List[str]] = None,
|
||||
is_published: bool = True,
|
||||
is_private: bool = False,
|
||||
editor: str = "markdown"
|
||||
) -> Dict[str, Any]:
|
||||
"""
|
||||
Create new page.
|
||||
|
||||
Multi-tenancy: Ensure path starts with user namespace.
|
||||
|
||||
Args:
|
||||
path: Page path (e.g., "/users/jpmschweitzer/projects/library-desk")
|
||||
title: Page title
|
||||
content: Page content (markdown)
|
||||
description: Short description
|
||||
tags: List of tags (for dossier organization)
|
||||
is_published: Whether page is published
|
||||
is_private: Whether page is private
|
||||
editor: Editor type (markdown, wysiwyg, etc.)
|
||||
|
||||
Returns:
|
||||
Created page object
|
||||
|
||||
Raises:
|
||||
Exception: If creation fails
|
||||
"""
|
||||
mutation = """
|
||||
mutation CreatePage(
|
||||
$content: String!,
|
||||
$description: String!,
|
||||
$editor: String!,
|
||||
$isPublished: Boolean!,
|
||||
$isPrivate: Boolean!,
|
||||
$locale: String!,
|
||||
$path: String!,
|
||||
$tags: [String]!,
|
||||
$title: String!
|
||||
) {
|
||||
pages {
|
||||
create(
|
||||
content: $content,
|
||||
description: $description,
|
||||
editor: $editor,
|
||||
isPublished: $isPublished,
|
||||
isPrivate: $isPrivate,
|
||||
locale: $locale,
|
||||
path: $path,
|
||||
tags: $tags,
|
||||
title: $title
|
||||
) {
|
||||
responseResult {
|
||||
succeeded
|
||||
errorCode
|
||||
message
|
||||
}
|
||||
page {
|
||||
id
|
||||
path
|
||||
title
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
"""
|
||||
|
||||
variables = {
|
||||
"path": path,
|
||||
"title": title,
|
||||
"content": content,
|
||||
"description": description,
|
||||
"tags": tags or [],
|
||||
"isPublished": is_published,
|
||||
"isPrivate": is_private,
|
||||
"editor": editor,
|
||||
"locale": "en"
|
||||
}
|
||||
|
||||
data = await self._execute_query(mutation, variables)
|
||||
result = data.get("pages", {}).get("create", {})
|
||||
|
||||
if not result.get("responseResult", {}).get("succeeded"):
|
||||
error = result.get("responseResult", {})
|
||||
raise Exception(f"Failed to create page: {error}")
|
||||
|
||||
logger.info(f"Created page: {path}")
|
||||
return result.get("page")
|
||||
|
||||
async def update_page(
|
||||
self,
|
||||
page_id: int,
|
||||
content: Optional[str] = None,
|
||||
title: Optional[str] = None,
|
||||
description: Optional[str] = None,
|
||||
tags: Optional[List[str]] = None,
|
||||
is_published: Optional[bool] = None
|
||||
) -> Dict[str, Any]:
|
||||
"""
|
||||
Update existing page.
|
||||
|
||||
Args:
|
||||
page_id: Page ID
|
||||
content: New content (optional)
|
||||
title: New title (optional)
|
||||
description: New description (optional)
|
||||
tags: New tags (optional)
|
||||
is_published: Published status (optional)
|
||||
|
||||
Returns:
|
||||
Updated page object
|
||||
|
||||
Raises:
|
||||
Exception: If update fails
|
||||
"""
|
||||
mutation = """
|
||||
mutation UpdatePage(
|
||||
$id: Int!,
|
||||
$content: String,
|
||||
$title: String,
|
||||
$description: String,
|
||||
$tags: [String],
|
||||
$isPublished: Boolean
|
||||
) {
|
||||
pages {
|
||||
update(
|
||||
id: $id,
|
||||
content: $content,
|
||||
title: $title,
|
||||
description: $description,
|
||||
tags: $tags,
|
||||
isPublished: $isPublished
|
||||
) {
|
||||
responseResult {
|
||||
succeeded
|
||||
errorCode
|
||||
message
|
||||
}
|
||||
page {
|
||||
id
|
||||
updatedAt
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
"""
|
||||
|
||||
variables = {"id": page_id}
|
||||
if content is not None:
|
||||
variables["content"] = content
|
||||
if title is not None:
|
||||
variables["title"] = title
|
||||
if description is not None:
|
||||
variables["description"] = description
|
||||
if tags is not None:
|
||||
variables["tags"] = tags
|
||||
if is_published is not None:
|
||||
variables["isPublished"] = is_published
|
||||
|
||||
data = await self._execute_query(mutation, variables)
|
||||
result = data.get("pages", {}).get("update", {})
|
||||
|
||||
if not result.get("responseResult", {}).get("succeeded"):
|
||||
error = result.get("responseResult", {})
|
||||
raise Exception(f"Failed to update page: {error}")
|
||||
|
||||
logger.info(f"Updated page: {page_id}")
|
||||
return result.get("page")
|
||||
|
||||
async def delete_page(self, page_id: int):
|
||||
"""
|
||||
Delete page.
|
||||
|
||||
Args:
|
||||
page_id: Page ID
|
||||
|
||||
Raises:
|
||||
Exception: If deletion fails
|
||||
"""
|
||||
mutation = """
|
||||
mutation DeletePage($id: Int!) {
|
||||
pages {
|
||||
delete(id: $id) {
|
||||
responseResult {
|
||||
succeeded
|
||||
errorCode
|
||||
message
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
"""
|
||||
|
||||
data = await self._execute_query(mutation, {"id": page_id})
|
||||
result = data.get("pages", {}).get("delete", {})
|
||||
|
||||
if not result.get("responseResult", {}).get("succeeded"):
|
||||
error = result.get("responseResult", {})
|
||||
raise Exception(f"Failed to delete page: {error}")
|
||||
|
||||
logger.info(f"Deleted page: {page_id}")
|
||||
|
||||
async def search_pages(
|
||||
self,
|
||||
query: str,
|
||||
path_prefix: Optional[str] = None
|
||||
) -> List[Dict[str, Any]]:
|
||||
"""
|
||||
Search pages by content.
|
||||
|
||||
Args:
|
||||
query: Search query
|
||||
path_prefix: Optional path prefix filter
|
||||
|
||||
Returns:
|
||||
List of matching pages
|
||||
"""
|
||||
gql_query = """
|
||||
query SearchPages($query: String!) {
|
||||
pages {
|
||||
search(query: $query) {
|
||||
results {
|
||||
id
|
||||
path
|
||||
title
|
||||
description
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
"""
|
||||
|
||||
data = await self._execute_query(gql_query, {"query": query})
|
||||
results = data.get("pages", {}).get("search", {}).get("results", [])
|
||||
|
||||
# Filter by path prefix if provided
|
||||
if path_prefix:
|
||||
results = [r for r in results if r["path"].startswith(path_prefix)]
|
||||
|
||||
return results
|
||||
|
||||
async def get_page_tree(self, path: str = "/") -> List[Dict[str, Any]]:
|
||||
"""
|
||||
Get page tree structure.
|
||||
|
||||
Args:
|
||||
path: Root path
|
||||
|
||||
Returns:
|
||||
Tree structure of pages
|
||||
"""
|
||||
query = """
|
||||
query GetPageTree($parent: Int, $mode: String!) {
|
||||
pages {
|
||||
tree(parent: $parent, mode: $mode) {
|
||||
id
|
||||
path
|
||||
title
|
||||
isFolder
|
||||
pageId
|
||||
}
|
||||
}
|
||||
}
|
||||
"""
|
||||
|
||||
try:
|
||||
data = await self._execute_query(
|
||||
query,
|
||||
{"parent": 0, "mode": "all"}
|
||||
)
|
||||
return data.get("pages", {}).get("tree", [])
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to get page tree: {e}")
|
||||
return []
|
||||
|
||||
async def move_page(
|
||||
self,
|
||||
page_id: int,
|
||||
new_path: str,
|
||||
locale: str = "en"
|
||||
) -> bool:
|
||||
"""
|
||||
Move/rename page.
|
||||
|
||||
Args:
|
||||
page_id: Page ID
|
||||
new_path: New page path
|
||||
locale: Page locale
|
||||
|
||||
Returns:
|
||||
True if successful
|
||||
"""
|
||||
mutation = """
|
||||
mutation MovePage($id: Int!, $destinationPath: String!, $destinationLocale: String!) {
|
||||
pages {
|
||||
move(id: $id, destinationPath: $destinationPath, destinationLocale: $destinationLocale) {
|
||||
responseResult {
|
||||
succeeded
|
||||
errorCode
|
||||
message
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
"""
|
||||
|
||||
try:
|
||||
data = await self._execute_query(
|
||||
mutation,
|
||||
{
|
||||
"id": page_id,
|
||||
"destinationPath": new_path,
|
||||
"destinationLocale": locale
|
||||
}
|
||||
)
|
||||
result = data.get("pages", {}).get("move", {})
|
||||
success = result.get("responseResult", {}).get("succeeded", False)
|
||||
|
||||
if success:
|
||||
logger.info(f"Moved page {page_id} to {new_path}")
|
||||
return success
|
||||
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to move page: {e}")
|
||||
return False
|
||||
@@ -1,98 +0,0 @@
|
||||
"""
|
||||
Application configuration using Pydantic Settings.
|
||||
Following best practices: modular settings, environment-based config.
|
||||
"""
|
||||
|
||||
from functools import lru_cache
|
||||
from pydantic import Field
|
||||
from pydantic_settings import BaseSettings, SettingsConfigDict
|
||||
|
||||
|
||||
class Settings(BaseSettings):
|
||||
"""Application settings loaded from environment variables."""
|
||||
|
||||
model_config = SettingsConfigDict(
|
||||
env_file=".env",
|
||||
env_file_encoding="utf-8",
|
||||
case_sensitive=False,
|
||||
extra="ignore",
|
||||
)
|
||||
|
||||
# API Configuration
|
||||
library_api_key: str = Field(..., description="API key for authentication")
|
||||
|
||||
# Neo4j Configuration
|
||||
neo4j_uri: str = Field(default="bolt://neo4j:7687", description="Neo4j Bolt URI")
|
||||
neo4j_user: str = Field(default="neo4j", description="Neo4j username")
|
||||
neo4j_password: str = Field(..., description="Neo4j password")
|
||||
|
||||
# Qdrant Configuration
|
||||
qdrant_host: str = Field(default="qdrant", description="Qdrant host")
|
||||
qdrant_port: int = Field(default=6333, description="Qdrant port")
|
||||
|
||||
# Wiki.js Configuration
|
||||
wikijs_url: str = Field(default="http://wiki:3000", description="Wiki.js URL")
|
||||
wikijs_username: str = Field(..., description="Wiki.js username")
|
||||
wikijs_password: str = Field(..., description="Wiki.js password")
|
||||
|
||||
# Wiki.js Database Configuration (for change listener)
|
||||
wikijs_db_host: str = Field(default="postgres-shared", description="Wiki.js PostgreSQL host")
|
||||
wikijs_db_port: int = Field(default=5432, description="Wiki.js PostgreSQL port")
|
||||
wikijs_db_name: str = Field(default="library", description="Wiki.js database name")
|
||||
wikijs_db_user: str = Field(default="library_desk_listener", description="Wiki.js database user (read-only)")
|
||||
wikijs_db_password: str = Field(..., description="Wiki.js database password")
|
||||
wikijs_change_listener_debounce_seconds: int = Field(
|
||||
default=5,
|
||||
ge=1,
|
||||
le=60,
|
||||
description="Debounce duration to prevent processing duplicate notifications"
|
||||
)
|
||||
|
||||
# SearXNG Configuration
|
||||
searxng_url: str = Field(default="http://searxng:8080", description="SearXNG URL")
|
||||
|
||||
# Ollama Configuration (for embeddings)
|
||||
ollama_url: str = Field(default="http://ollama:11434", description="Ollama URL")
|
||||
ollama_model: str = Field(default="nomic-embed-text", description="Ollama embedding model")
|
||||
|
||||
# HybridRAG Configuration
|
||||
reranker_model: str = Field(default="mistral-nemo", description="Model for LLM re-ranking")
|
||||
reranker_enabled: bool = Field(default=True, description="Enable LLM re-ranking")
|
||||
hybrid_rag_vector_limit: int = Field(default=10, ge=1, le=50, description="Vector search limit")
|
||||
hybrid_rag_graph_limit: int = Field(default=10, ge=1, le=50, description="Graph search limit")
|
||||
hybrid_rag_web_limit: int = Field(default=5, ge=1, le=20, description="Web search limit")
|
||||
|
||||
# Entity Linking Fuzzy Matching Configuration
|
||||
entity_linking_min_confidence: float = Field(default=0.70, ge=0.0, le=1.0, description="Minimum confidence for entity-document matching")
|
||||
entity_linking_min_entity_length: int = Field(default=5, ge=1, le=50, description="Minimum entity name length for matching")
|
||||
entity_linking_min_containment_ratio: float = Field(default=0.30, ge=0.0, le=1.0, description="Minimum containment ratio for substring matching")
|
||||
entity_linking_min_token_overlap: float = Field(default=0.60, ge=0.0, le=1.0, description="Minimum token overlap ratio for matching")
|
||||
|
||||
# Redis Configuration (for job tracking - separate DB from wiki)
|
||||
redis_host: str = Field(default="redis-shared", description="Redis host")
|
||||
redis_port: int = Field(default=6379, description="Redis port")
|
||||
redis_db: int = Field(default=4, description="Redis database number (4 for library-desk jobs)")
|
||||
|
||||
# Application
|
||||
app_name: str = Field(default="Library Desk", description="Application name")
|
||||
app_version: str = Field(default="1.0.0", description="Application version")
|
||||
debug: bool = Field(default=False, description="Debug mode")
|
||||
|
||||
@property
|
||||
def qdrant_url(self) -> str:
|
||||
"""Computed Qdrant URL."""
|
||||
return f"http://{self.qdrant_host}:{self.qdrant_port}"
|
||||
|
||||
@property
|
||||
def redis_url(self) -> str:
|
||||
"""Computed Redis URL."""
|
||||
return f"redis://{self.redis_host}:{self.redis_port}/{self.redis_db}"
|
||||
|
||||
|
||||
@lru_cache
|
||||
def get_settings() -> Settings:
|
||||
"""
|
||||
Get cached settings instance.
|
||||
Uses lru_cache to ensure single instance across app.
|
||||
"""
|
||||
return Settings()
|
||||
@@ -1,384 +0,0 @@
|
||||
"""
|
||||
Dependency injection for Library Desk.
|
||||
|
||||
Provides FastAPI dependencies for service clients with:
|
||||
- Singleton pattern via @lru_cache
|
||||
- Lazy initialization
|
||||
- Proper lifecycle management
|
||||
- Type aliases for clean endpoint signatures
|
||||
"""
|
||||
|
||||
from functools import lru_cache
|
||||
from typing import Annotated
|
||||
from fastapi import Depends
|
||||
import logging
|
||||
|
||||
from src.config import Settings, get_settings
|
||||
from src.clients.neo4j_client import Neo4jClient
|
||||
from src.clients.qdrant_client import QdrantClientWrapper
|
||||
from src.clients.wikijs_client import WikiJSClient
|
||||
from src.clients.searxng_client import SearXNGClient
|
||||
from src.clients.ollama_client import OllamaClient
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
|
||||
# Settings dependency
|
||||
SettingsDep = Annotated[Settings, Depends(get_settings)]
|
||||
|
||||
|
||||
# Client factory functions with @lru_cache for singletons
|
||||
@lru_cache
|
||||
def get_neo4j_client() -> Neo4jClient:
|
||||
"""
|
||||
Get Neo4j client singleton.
|
||||
|
||||
Returns:
|
||||
Initialized Neo4j client (not yet connected)
|
||||
|
||||
Note: Call client.connect() during app startup
|
||||
"""
|
||||
settings = get_settings()
|
||||
client = Neo4jClient(
|
||||
uri=settings.neo4j_uri,
|
||||
user=settings.neo4j_user,
|
||||
password=settings.neo4j_password
|
||||
)
|
||||
logger.debug("Created Neo4j client instance")
|
||||
return client
|
||||
|
||||
|
||||
@lru_cache
|
||||
def get_qdrant_client() -> QdrantClientWrapper:
|
||||
"""
|
||||
Get Qdrant client singleton.
|
||||
|
||||
Returns:
|
||||
Initialized Qdrant client
|
||||
|
||||
Note: Collections are created lazily per-user
|
||||
"""
|
||||
settings = get_settings()
|
||||
client = QdrantClientWrapper(
|
||||
url=settings.qdrant_url,
|
||||
embedding_dim=768 # nomic-embed-text default
|
||||
)
|
||||
logger.debug("Created Qdrant client instance")
|
||||
return client
|
||||
|
||||
|
||||
@lru_cache
|
||||
def get_wikijs_client() -> WikiJSClient:
|
||||
"""
|
||||
Get Wiki.js client singleton.
|
||||
|
||||
Returns:
|
||||
Initialized Wiki.js GraphQL client with username/password auth
|
||||
"""
|
||||
settings = get_settings()
|
||||
client = WikiJSClient(
|
||||
base_url=settings.wikijs_url,
|
||||
username=settings.wikijs_username,
|
||||
password=settings.wikijs_password
|
||||
)
|
||||
logger.debug("Created Wiki.js client instance")
|
||||
return client
|
||||
|
||||
|
||||
@lru_cache
|
||||
def get_searxng_client() -> SearXNGClient:
|
||||
"""
|
||||
Get SearXNG client singleton.
|
||||
|
||||
Returns:
|
||||
Initialized SearXNG search client
|
||||
"""
|
||||
settings = get_settings()
|
||||
client = SearXNGClient(base_url=settings.searxng_url)
|
||||
logger.debug("Created SearXNG client instance")
|
||||
return client
|
||||
|
||||
|
||||
@lru_cache
|
||||
def get_ollama_client() -> OllamaClient:
|
||||
"""
|
||||
Get Ollama client singleton.
|
||||
|
||||
Returns:
|
||||
Initialized Ollama embeddings client
|
||||
"""
|
||||
settings = get_settings()
|
||||
client = OllamaClient(
|
||||
base_url=settings.ollama_url,
|
||||
model=settings.ollama_model
|
||||
)
|
||||
logger.debug("Created Ollama client instance")
|
||||
return client
|
||||
|
||||
|
||||
# Type aliases for FastAPI endpoint dependencies
|
||||
# Usage: def my_endpoint(neo4j: Neo4jDep):
|
||||
Neo4jDep = Annotated[Neo4jClient, Depends(get_neo4j_client)]
|
||||
QdrantDep = Annotated[QdrantClientWrapper, Depends(get_qdrant_client)]
|
||||
WikiJSDep = Annotated[WikiJSClient, Depends(get_wikijs_client)]
|
||||
SearXNGDep = Annotated[SearXNGClient, Depends(get_searxng_client)]
|
||||
OllamaDep = Annotated[OllamaClient, Depends(get_ollama_client)]
|
||||
|
||||
|
||||
# Lifecycle management functions
|
||||
async def startup_clients():
|
||||
"""
|
||||
Initialize all service clients at application startup.
|
||||
|
||||
Should be called in FastAPI lifespan or startup event.
|
||||
|
||||
Performs:
|
||||
- Neo4j connection pool initialization
|
||||
- Neo4j connectivity verification
|
||||
- Ollama model availability check
|
||||
"""
|
||||
logger.info("Starting up service clients...")
|
||||
|
||||
# Initialize Neo4j connection pool
|
||||
neo4j = get_neo4j_client()
|
||||
try:
|
||||
await neo4j.connect()
|
||||
logger.info("✓ Neo4j connected")
|
||||
except Exception as e:
|
||||
logger.error(f"✗ Neo4j connection failed: {e}")
|
||||
# Don't fail startup - allow degraded operation
|
||||
pass
|
||||
|
||||
# Check Ollama availability
|
||||
ollama = get_ollama_client()
|
||||
try:
|
||||
is_healthy = await ollama.health_check()
|
||||
if is_healthy:
|
||||
logger.info(f"✓ Ollama ready (model: {ollama.model})")
|
||||
else:
|
||||
logger.warning(f"✗ Ollama model '{ollama.model}' not available")
|
||||
except Exception as e:
|
||||
logger.error(f"✗ Ollama health check failed: {e}")
|
||||
pass
|
||||
|
||||
# Qdrant, Wiki.js, SearXNG are lazy-initialized
|
||||
logger.info("Service clients startup complete")
|
||||
|
||||
|
||||
async def shutdown_clients():
|
||||
"""
|
||||
Cleanup all service clients at application shutdown.
|
||||
|
||||
Should be called in FastAPI lifespan or shutdown event.
|
||||
|
||||
Performs:
|
||||
- Close Neo4j connection pool
|
||||
- Close HTTP clients
|
||||
"""
|
||||
logger.info("Shutting down service clients...")
|
||||
|
||||
# Close Neo4j driver
|
||||
neo4j = get_neo4j_client()
|
||||
try:
|
||||
await neo4j.close()
|
||||
logger.info("✓ Neo4j closed")
|
||||
except Exception as e:
|
||||
logger.error(f"Error closing Neo4j: {e}")
|
||||
|
||||
# Close HTTP clients
|
||||
clients_to_close = [
|
||||
("Wiki.js", get_wikijs_client()),
|
||||
("SearXNG", get_searxng_client()),
|
||||
("Ollama", get_ollama_client())
|
||||
]
|
||||
|
||||
for name, client in clients_to_close:
|
||||
try:
|
||||
await client.close()
|
||||
logger.info(f"✓ {name} client closed")
|
||||
except Exception as e:
|
||||
logger.error(f"Error closing {name} client: {e}")
|
||||
|
||||
logger.info("Service clients shutdown complete")
|
||||
|
||||
|
||||
async def check_service_health() -> dict:
|
||||
"""
|
||||
Check health of all service clients.
|
||||
|
||||
Returns:
|
||||
Dictionary with health status of each service:
|
||||
{
|
||||
"neo4j": bool,
|
||||
"qdrant": bool,
|
||||
"wikijs": bool,
|
||||
"searxng": bool,
|
||||
"ollama": bool
|
||||
}
|
||||
|
||||
Usage:
|
||||
>>> health = await check_service_health()
|
||||
>>> health["neo4j"]
|
||||
True
|
||||
"""
|
||||
health = {}
|
||||
|
||||
# Neo4j
|
||||
try:
|
||||
neo4j = get_neo4j_client()
|
||||
# Simple query to check connectivity
|
||||
await neo4j.execute_query("RETURN 1 as test", {})
|
||||
health["neo4j"] = True
|
||||
except Exception as e:
|
||||
logger.error(f"Neo4j health check failed: {e}")
|
||||
health["neo4j"] = False
|
||||
|
||||
# Qdrant
|
||||
try:
|
||||
qdrant = get_qdrant_client()
|
||||
# Check if we can list collections
|
||||
collections = qdrant.client.get_collections()
|
||||
health["qdrant"] = True
|
||||
except Exception as e:
|
||||
logger.error(f"Qdrant health check failed: {e}")
|
||||
health["qdrant"] = False
|
||||
|
||||
# Wiki.js
|
||||
try:
|
||||
wikijs = get_wikijs_client()
|
||||
# Try a simple query (list pages with limit 1)
|
||||
await wikijs.list_pages(limit=1)
|
||||
health["wikijs"] = True
|
||||
except Exception as e:
|
||||
logger.error(f"Wiki.js health check failed: {e}")
|
||||
health["wikijs"] = False
|
||||
|
||||
# SearXNG
|
||||
try:
|
||||
searxng = get_searxng_client()
|
||||
# Just check if service is up (no actual search)
|
||||
health["searxng"] = await searxng.health_check()
|
||||
except Exception as e:
|
||||
logger.error(f"SearXNG health check failed: {e}")
|
||||
health["searxng"] = False
|
||||
|
||||
# Ollama
|
||||
try:
|
||||
ollama = get_ollama_client()
|
||||
is_healthy = await ollama.health_check()
|
||||
health["ollama"] = is_healthy
|
||||
except Exception as e:
|
||||
logger.error(f"Ollama health check failed: {e}")
|
||||
health["ollama"] = False
|
||||
|
||||
return health
|
||||
|
||||
|
||||
# Service factory functions
|
||||
@lru_cache
|
||||
def get_vector_service() -> "VectorService":
|
||||
"""Get VectorService singleton."""
|
||||
from src.services.vector_service import VectorService
|
||||
return VectorService(
|
||||
qdrant_client=get_qdrant_client(),
|
||||
wikijs_client=get_wikijs_client(),
|
||||
ollama_client=get_ollama_client()
|
||||
)
|
||||
|
||||
|
||||
@lru_cache
|
||||
def get_graph_service() -> "GraphService":
|
||||
"""Get GraphService singleton."""
|
||||
from src.services.graph_service import GraphService
|
||||
return GraphService(
|
||||
neo4j_client=get_neo4j_client(),
|
||||
wikijs_client=get_wikijs_client()
|
||||
)
|
||||
|
||||
|
||||
@lru_cache
|
||||
def get_wiki_service() -> "WikiService":
|
||||
"""Get WikiService singleton."""
|
||||
from src.services.wiki_service import WikiService
|
||||
return WikiService(wiki_client=get_wikijs_client())
|
||||
|
||||
|
||||
@lru_cache
|
||||
def get_consolidation_service() -> "ConsolidationService":
|
||||
"""Get ConsolidationService singleton."""
|
||||
from src.services.consolidation_service import ConsolidationService
|
||||
return ConsolidationService(
|
||||
neo4j=get_neo4j_client(),
|
||||
ollama=get_ollama_client(),
|
||||
wiki=get_wikijs_client(),
|
||||
settings=get_settings(),
|
||||
ingestion_service=get_ingestion_service()
|
||||
)
|
||||
|
||||
|
||||
@lru_cache
|
||||
def get_ingestion_service() -> "IngestionService":
|
||||
"""Get IngestionService singleton."""
|
||||
from src.services.ingestion_service import IngestionService
|
||||
return IngestionService(
|
||||
vector_service=get_vector_service(),
|
||||
graph_service=get_graph_service(),
|
||||
wiki_client=get_wikijs_client()
|
||||
)
|
||||
|
||||
|
||||
@lru_cache
|
||||
def get_hybrid_rag_service() -> "HybridRAGService":
|
||||
"""Get HybridRAGService singleton."""
|
||||
from src.services.hybrid_rag_service import HybridRAGService
|
||||
return HybridRAGService(
|
||||
vector_service=get_vector_service(),
|
||||
graph_service=get_graph_service(),
|
||||
searxng_client=get_searxng_client(),
|
||||
ollama_client=get_ollama_client(),
|
||||
settings=get_settings()
|
||||
)
|
||||
|
||||
|
||||
# Utility: Get default user from settings or multi_tenancy
|
||||
def get_default_user() -> str:
|
||||
"""
|
||||
Get default user for operations.
|
||||
|
||||
Returns:
|
||||
Default user identifier
|
||||
"""
|
||||
from src.core.multi_tenancy import DEFAULT_USER
|
||||
return DEFAULT_USER
|
||||
|
||||
|
||||
# Authentication
|
||||
from fastapi import Security, HTTPException
|
||||
from fastapi.security import HTTPBearer
|
||||
|
||||
security = HTTPBearer()
|
||||
|
||||
|
||||
async def verify_api_key(
|
||||
credentials: Annotated[HTTPBearer, Security(security)],
|
||||
settings: SettingsDep
|
||||
) -> str:
|
||||
"""
|
||||
Verify API key from Bearer token.
|
||||
|
||||
Args:
|
||||
credentials: HTTP Bearer credentials
|
||||
settings: Application settings
|
||||
|
||||
Returns:
|
||||
API key if valid
|
||||
|
||||
Raises:
|
||||
HTTPException: If API key is invalid
|
||||
"""
|
||||
if credentials.credentials != settings.library_api_key:
|
||||
raise HTTPException(
|
||||
status_code=403,
|
||||
detail="Invalid API key"
|
||||
)
|
||||
return credentials.credentials
|
||||
@@ -1,209 +0,0 @@
|
||||
"""
|
||||
Multi-tenancy helpers for Library Desk.
|
||||
|
||||
Provides utilities for user namespace management across:
|
||||
- Wiki.js (path-based namespaces)
|
||||
- Neo4j (user-specific labels)
|
||||
- Qdrant (collection per user)
|
||||
"""
|
||||
|
||||
import re
|
||||
|
||||
# Default user for all operations
|
||||
DEFAULT_USER = "jpmschweitzer"
|
||||
|
||||
|
||||
def sanitize_user_id(user_id: str) -> str:
|
||||
"""
|
||||
Sanitize user ID for use in collection names, labels, and paths.
|
||||
|
||||
Converts special characters to underscores and ensures alphanumeric safety.
|
||||
|
||||
Args:
|
||||
user_id: Raw user identifier (email, username, etc.)
|
||||
|
||||
Returns:
|
||||
Sanitized user ID safe for use in identifiers
|
||||
|
||||
Examples:
|
||||
>>> sanitize_user_id("john@example.com")
|
||||
'john_at_example_com'
|
||||
>>> sanitize_user_id("user.name")
|
||||
'user_name'
|
||||
>>> sanitize_user_id("User Name")
|
||||
'user_name'
|
||||
"""
|
||||
sanitized = user_id.lower()
|
||||
|
||||
# Convert @ to _at_
|
||||
sanitized = sanitized.replace("@", "_at_")
|
||||
|
||||
# Convert dots to underscores
|
||||
sanitized = sanitized.replace(".", "_")
|
||||
|
||||
# Replace any non-alphanumeric characters with underscores
|
||||
sanitized = re.sub(r'[^a-z0-9_]', '_', sanitized)
|
||||
|
||||
# Remove consecutive underscores
|
||||
sanitized = re.sub(r'_+', '_', sanitized)
|
||||
|
||||
# Remove leading/trailing underscores
|
||||
sanitized = sanitized.strip('_')
|
||||
|
||||
return sanitized
|
||||
|
||||
|
||||
def get_qdrant_collection_name(user_id: str) -> str:
|
||||
"""
|
||||
Get Qdrant collection name for user.
|
||||
|
||||
Pattern: library_desk_{sanitized_user_id}
|
||||
|
||||
Args:
|
||||
user_id: User identifier
|
||||
|
||||
Returns:
|
||||
Qdrant collection name
|
||||
|
||||
Examples:
|
||||
>>> get_qdrant_collection_name("jpmschweitzer")
|
||||
'library_desk_jpmschweitzer'
|
||||
>>> get_qdrant_collection_name("john@example.com")
|
||||
'library_desk_john_at_example_com'
|
||||
"""
|
||||
sanitized = sanitize_user_id(user_id)
|
||||
return f"library_desk_{sanitized}"
|
||||
|
||||
|
||||
def get_wikijs_namespace(user_id: str) -> str:
|
||||
"""
|
||||
Get Wiki.js namespace (path prefix) for user.
|
||||
|
||||
Pattern: /users/{sanitized_user_id}
|
||||
|
||||
All wiki pages for a user will be under this namespace.
|
||||
|
||||
Args:
|
||||
user_id: User identifier
|
||||
|
||||
Returns:
|
||||
Wiki.js path prefix
|
||||
|
||||
Examples:
|
||||
>>> get_wikijs_namespace("jpmschweitzer")
|
||||
'/users/jpmschweitzer'
|
||||
>>> get_wikijs_namespace("john@example.com")
|
||||
'/users/john_at_example_com'
|
||||
"""
|
||||
sanitized = sanitize_user_id(user_id)
|
||||
return f"/users/{sanitized}"
|
||||
|
||||
|
||||
def get_neo4j_user_base_label(user_id: str) -> str:
|
||||
"""
|
||||
Get Neo4j base label for user's nodes (entities, etc).
|
||||
|
||||
Pattern: User_{Sanitized}
|
||||
|
||||
Uses title case for Neo4j label convention.
|
||||
|
||||
Args:
|
||||
user_id: User identifier
|
||||
|
||||
Returns:
|
||||
Neo4j base label for user's nodes
|
||||
|
||||
Examples:
|
||||
>>> get_neo4j_user_base_label("jpmschweitzer")
|
||||
'User_Jpmschweitzer'
|
||||
>>> get_neo4j_user_base_label("john@example.com")
|
||||
'User_John_At_Example_Com'
|
||||
"""
|
||||
sanitized = sanitize_user_id(user_id)
|
||||
# Title case each segment for Neo4j label convention
|
||||
parts = sanitized.split('_')
|
||||
titled = '_'.join(part.capitalize() for part in parts if part)
|
||||
return f"User_{titled}"
|
||||
|
||||
|
||||
def get_neo4j_user_label(user_id: str) -> str:
|
||||
"""
|
||||
Get Neo4j label for user's documents.
|
||||
|
||||
Pattern: User_{Sanitized}_Document
|
||||
|
||||
Uses title case for Neo4j label convention.
|
||||
|
||||
Args:
|
||||
user_id: User identifier
|
||||
|
||||
Returns:
|
||||
Neo4j label for user's document nodes
|
||||
|
||||
Examples:
|
||||
>>> get_neo4j_user_label("jpmschweitzer")
|
||||
'User_Jpmschweitzer_Document'
|
||||
>>> get_neo4j_user_label("john@example.com")
|
||||
'User_John_At_Example_Com_Document'
|
||||
"""
|
||||
sanitized = sanitize_user_id(user_id)
|
||||
# Title case each segment for Neo4j label convention
|
||||
parts = sanitized.split('_')
|
||||
titled = '_'.join(part.capitalize() for part in parts if part)
|
||||
return f"User_{titled}_Document"
|
||||
|
||||
|
||||
def validate_user_id(user_id: str) -> bool:
|
||||
"""
|
||||
Validate that a user ID is acceptable.
|
||||
|
||||
Checks:
|
||||
- Not empty
|
||||
- Not too long (max 100 chars)
|
||||
- Contains some alphanumeric characters
|
||||
|
||||
Args:
|
||||
user_id: User identifier to validate
|
||||
|
||||
Returns:
|
||||
True if valid, False otherwise
|
||||
|
||||
Examples:
|
||||
>>> validate_user_id("jpmschweitzer")
|
||||
True
|
||||
>>> validate_user_id("")
|
||||
False
|
||||
>>> validate_user_id("a" * 101)
|
||||
False
|
||||
"""
|
||||
if not user_id or len(user_id) > 100:
|
||||
return False
|
||||
|
||||
# Must contain at least one alphanumeric character
|
||||
if not re.search(r'[a-zA-Z0-9]', user_id):
|
||||
return False
|
||||
|
||||
return True
|
||||
|
||||
|
||||
def is_path_in_user_namespace(path: str, user_id: str) -> bool:
|
||||
"""
|
||||
Check if a Wiki.js path belongs to user's namespace.
|
||||
|
||||
Args:
|
||||
path: Wiki.js page path
|
||||
user_id: User identifier
|
||||
|
||||
Returns:
|
||||
True if path is in user's namespace
|
||||
|
||||
Examples:
|
||||
>>> is_path_in_user_namespace("/users/jpmschweitzer/projects", "jpmschweitzer")
|
||||
True
|
||||
>>> is_path_in_user_namespace("/users/other/projects", "jpmschweitzer")
|
||||
False
|
||||
>>> is_path_in_user_namespace("/public/docs", "jpmschweitzer")
|
||||
False
|
||||
"""
|
||||
namespace = get_wikijs_namespace(user_id)
|
||||
return path.startswith(namespace)
|
||||
@@ -1,426 +0,0 @@
|
||||
"""
|
||||
Redis-based job tracking for Library Desk.
|
||||
|
||||
Provides background job management with:
|
||||
- Job creation and status tracking
|
||||
- Progress updates
|
||||
- Result storage
|
||||
- Auto-expiration after 24 hours
|
||||
- User-scoped job queries
|
||||
"""
|
||||
|
||||
import redis.asyncio as redis
|
||||
import json
|
||||
import uuid
|
||||
from datetime import datetime, timezone
|
||||
from typing import Optional, Dict, Any, List
|
||||
from enum import Enum
|
||||
import logging
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
|
||||
class JobStatus(str, Enum):
|
||||
"""Job status enumeration."""
|
||||
QUEUED = "queued"
|
||||
PROCESSING = "processing"
|
||||
COMPLETED = "completed"
|
||||
FAILED = "failed"
|
||||
CANCELLED = "cancelled"
|
||||
|
||||
|
||||
class JobType(str, Enum):
|
||||
"""Job type enumeration."""
|
||||
DOCUMENT_INGESTION = "document_ingestion"
|
||||
BATCH_INGESTION = "batch_ingestion"
|
||||
GRAPH_UPDATE = "graph_update"
|
||||
VECTOR_UPDATE = "vector_update"
|
||||
EMBEDDINGS_GENERATION = "embeddings_generation"
|
||||
|
||||
|
||||
class JobManager:
|
||||
"""
|
||||
Redis-based job tracking manager.
|
||||
|
||||
Key structure:
|
||||
- library:job:{job_id} - Job data (JSON, TTL 24h)
|
||||
- library:user_jobs:{user} - Set of job IDs for user
|
||||
- library:active_jobs - Set of active (non-completed) job IDs
|
||||
|
||||
Job data format:
|
||||
{
|
||||
"job_id": "uuid",
|
||||
"job_type": "document_ingestion",
|
||||
"status": "processing",
|
||||
"user": "jpmschweitzer",
|
||||
"parameters": {...},
|
||||
"result": {...},
|
||||
"error": "error message",
|
||||
"progress": 50,
|
||||
"created_at": "2024-01-15T10:30:00Z",
|
||||
"started_at": "2024-01-15T10:30:05Z",
|
||||
"completed_at": "2024-01-15T10:35:00Z"
|
||||
}
|
||||
"""
|
||||
|
||||
def __init__(self, redis_url: str):
|
||||
"""
|
||||
Initialize job manager.
|
||||
|
||||
Args:
|
||||
redis_url: Redis connection URL (e.g., "redis://redis-shared:6379/4")
|
||||
"""
|
||||
self.redis_url = redis_url
|
||||
self._redis: Optional[redis.Redis] = None
|
||||
self.ttl = 86400 # 24 hours
|
||||
logger.info(f"Initialized JobManager: {redis_url}")
|
||||
|
||||
async def connect(self):
|
||||
"""Initialize Redis connection."""
|
||||
if not self._redis:
|
||||
self._redis = await redis.from_url(
|
||||
self.redis_url,
|
||||
encoding="utf-8",
|
||||
decode_responses=True
|
||||
)
|
||||
logger.info("Connected to Redis for job tracking")
|
||||
|
||||
async def close(self):
|
||||
"""Close Redis connection."""
|
||||
if self._redis:
|
||||
await self._redis.close()
|
||||
self._redis = None
|
||||
logger.info("Closed Redis connection")
|
||||
|
||||
def _job_key(self, job_id: str) -> str:
|
||||
"""Get Redis key for job."""
|
||||
return f"library:job:{job_id}"
|
||||
|
||||
def _user_jobs_key(self, user: str) -> str:
|
||||
"""Get Redis key for user's jobs."""
|
||||
return f"library:user_jobs:{user}"
|
||||
|
||||
def _active_jobs_key(self) -> str:
|
||||
"""Get Redis key for active jobs."""
|
||||
return "library:active_jobs"
|
||||
|
||||
async def create_job(
|
||||
self,
|
||||
job_type: JobType,
|
||||
user: str,
|
||||
parameters: Dict[str, Any]
|
||||
) -> str:
|
||||
"""
|
||||
Create new job.
|
||||
|
||||
Args:
|
||||
job_type: Type of job
|
||||
user: User identifier
|
||||
parameters: Job parameters
|
||||
|
||||
Returns:
|
||||
Job ID (UUID)
|
||||
|
||||
Example:
|
||||
>>> job_id = await manager.create_job(
|
||||
... JobType.DOCUMENT_INGESTION,
|
||||
... "jpmschweitzer",
|
||||
... {"doc_url": "https://example.com/doc.pdf"}
|
||||
... )
|
||||
"""
|
||||
if not self._redis:
|
||||
await self.connect()
|
||||
|
||||
job_id = str(uuid.uuid4())
|
||||
job_data = {
|
||||
"job_id": job_id,
|
||||
"job_type": job_type.value,
|
||||
"status": JobStatus.QUEUED.value,
|
||||
"user": user,
|
||||
"parameters": parameters,
|
||||
"result": None,
|
||||
"error": None,
|
||||
"progress": 0,
|
||||
"created_at": datetime.now(timezone.utc).isoformat()
|
||||
}
|
||||
|
||||
# Store job data
|
||||
await self._redis.setex(
|
||||
self._job_key(job_id),
|
||||
self.ttl,
|
||||
json.dumps(job_data)
|
||||
)
|
||||
|
||||
# Add to user's job set
|
||||
await self._redis.sadd(self._user_jobs_key(user), job_id)
|
||||
await self._redis.expire(self._user_jobs_key(user), self.ttl)
|
||||
|
||||
# Add to active jobs set
|
||||
await self._redis.sadd(self._active_jobs_key(), job_id)
|
||||
|
||||
logger.info(f"Created job {job_id}: {job_type.value} for user {user}")
|
||||
return job_id
|
||||
|
||||
async def get_job(self, job_id: str) -> Optional[Dict[str, Any]]:
|
||||
"""
|
||||
Get job by ID.
|
||||
|
||||
Args:
|
||||
job_id: Job ID
|
||||
|
||||
Returns:
|
||||
Job data dictionary or None if not found
|
||||
|
||||
Example:
|
||||
>>> job = await manager.get_job(job_id)
|
||||
>>> job["status"]
|
||||
'processing'
|
||||
"""
|
||||
if not self._redis:
|
||||
await self.connect()
|
||||
|
||||
job_json = await self._redis.get(self._job_key(job_id))
|
||||
if not job_json:
|
||||
return None
|
||||
|
||||
return json.loads(job_json)
|
||||
|
||||
async def update_job_status(
|
||||
self,
|
||||
job_id: str,
|
||||
status: JobStatus,
|
||||
progress: Optional[int] = None,
|
||||
result: Optional[Dict[str, Any]] = None,
|
||||
error: Optional[str] = None
|
||||
):
|
||||
"""
|
||||
Update job status.
|
||||
|
||||
Args:
|
||||
job_id: Job ID
|
||||
status: New status
|
||||
progress: Progress percentage (0-100)
|
||||
result: Result data (for completed jobs)
|
||||
error: Error message (for failed jobs)
|
||||
|
||||
Example:
|
||||
>>> await manager.update_job_status(
|
||||
... job_id,
|
||||
... JobStatus.PROCESSING,
|
||||
... progress=50
|
||||
... )
|
||||
"""
|
||||
if not self._redis:
|
||||
await self.connect()
|
||||
|
||||
job = await self.get_job(job_id)
|
||||
if not job:
|
||||
logger.error(f"Job {job_id} not found")
|
||||
return
|
||||
|
||||
# Update fields
|
||||
job["status"] = status.value
|
||||
if progress is not None:
|
||||
job["progress"] = progress
|
||||
if result is not None:
|
||||
job["result"] = result
|
||||
if error is not None:
|
||||
job["error"] = error
|
||||
|
||||
# Update timestamps
|
||||
now = datetime.now(timezone.utc).isoformat()
|
||||
if status == JobStatus.PROCESSING and not job.get("started_at"):
|
||||
job["started_at"] = now
|
||||
if status in (JobStatus.COMPLETED, JobStatus.FAILED, JobStatus.CANCELLED):
|
||||
job["completed_at"] = now
|
||||
|
||||
# Save updated job
|
||||
await self._redis.setex(
|
||||
self._job_key(job_id),
|
||||
self.ttl,
|
||||
json.dumps(job)
|
||||
)
|
||||
|
||||
# Remove from active jobs if completed
|
||||
if status in (JobStatus.COMPLETED, JobStatus.FAILED, JobStatus.CANCELLED):
|
||||
await self._redis.srem(self._active_jobs_key(), job_id)
|
||||
|
||||
logger.info(f"Updated job {job_id}: status={status.value}, progress={progress}")
|
||||
|
||||
async def get_user_jobs(
|
||||
self,
|
||||
user: str,
|
||||
limit: int = 50
|
||||
) -> List[Dict[str, Any]]:
|
||||
"""
|
||||
Get all jobs for a user.
|
||||
|
||||
Args:
|
||||
user: User identifier
|
||||
limit: Maximum jobs to return
|
||||
|
||||
Returns:
|
||||
List of job dictionaries (most recent first)
|
||||
|
||||
Example:
|
||||
>>> jobs = await manager.get_user_jobs("jpmschweitzer", limit=10)
|
||||
>>> len(jobs)
|
||||
10
|
||||
"""
|
||||
if not self._redis:
|
||||
await self.connect()
|
||||
|
||||
# Get job IDs for user
|
||||
job_ids = await self._redis.smembers(self._user_jobs_key(user))
|
||||
if not job_ids:
|
||||
return []
|
||||
|
||||
# Fetch job data
|
||||
jobs = []
|
||||
for job_id in job_ids:
|
||||
job = await self.get_job(job_id)
|
||||
if job:
|
||||
jobs.append(job)
|
||||
|
||||
# Sort by created_at descending
|
||||
jobs.sort(key=lambda j: j.get("created_at", ""), reverse=True)
|
||||
|
||||
return jobs[:limit]
|
||||
|
||||
async def get_active_jobs(self, limit: int = 100) -> List[Dict[str, Any]]:
|
||||
"""
|
||||
Get all active (non-completed) jobs.
|
||||
|
||||
Args:
|
||||
limit: Maximum jobs to return
|
||||
|
||||
Returns:
|
||||
List of active job dictionaries
|
||||
|
||||
Example:
|
||||
>>> active = await manager.get_active_jobs()
|
||||
>>> all(j["status"] in ["queued", "processing"] for j in active)
|
||||
True
|
||||
"""
|
||||
if not self._redis:
|
||||
await self.connect()
|
||||
|
||||
job_ids = await self._redis.smembers(self._active_jobs_key())
|
||||
if not job_ids:
|
||||
return []
|
||||
|
||||
jobs = []
|
||||
for job_id in job_ids:
|
||||
job = await self.get_job(job_id)
|
||||
if job:
|
||||
jobs.append(job)
|
||||
|
||||
return jobs[:limit]
|
||||
|
||||
async def cancel_job(self, job_id: str):
|
||||
"""
|
||||
Cancel a job.
|
||||
|
||||
Args:
|
||||
job_id: Job ID
|
||||
|
||||
Note: This only marks the job as cancelled. The actual job
|
||||
worker must check status and stop processing.
|
||||
"""
|
||||
await self.update_job_status(
|
||||
job_id,
|
||||
JobStatus.CANCELLED
|
||||
)
|
||||
logger.info(f"Cancelled job {job_id}")
|
||||
|
||||
async def delete_job(self, job_id: str, user: str):
|
||||
"""
|
||||
Delete a job.
|
||||
|
||||
Args:
|
||||
job_id: Job ID
|
||||
user: User identifier (for authorization)
|
||||
|
||||
Returns:
|
||||
True if deleted, False if not found or unauthorized
|
||||
"""
|
||||
if not self._redis:
|
||||
await self.connect()
|
||||
|
||||
job = await self.get_job(job_id)
|
||||
if not job:
|
||||
return False
|
||||
|
||||
# Check user authorization
|
||||
if job.get("user") != user:
|
||||
logger.warning(f"User {user} attempted to delete job {job_id} owned by {job.get('user')}")
|
||||
return False
|
||||
|
||||
# Delete job data
|
||||
await self._redis.delete(self._job_key(job_id))
|
||||
|
||||
# Remove from user's jobs
|
||||
await self._redis.srem(self._user_jobs_key(user), job_id)
|
||||
|
||||
# Remove from active jobs
|
||||
await self._redis.srem(self._active_jobs_key(), job_id)
|
||||
|
||||
logger.info(f"Deleted job {job_id}")
|
||||
return True
|
||||
|
||||
async def cleanup_expired_jobs(self):
|
||||
"""
|
||||
Clean up expired jobs from sets.
|
||||
|
||||
Redis will auto-expire job data, but set memberships need manual cleanup.
|
||||
Should be called periodically (e.g., hourly).
|
||||
"""
|
||||
if not self._redis:
|
||||
await self.connect()
|
||||
|
||||
# Cleanup active jobs set
|
||||
job_ids = await self._redis.smembers(self._active_jobs_key())
|
||||
for job_id in job_ids:
|
||||
exists = await self._redis.exists(self._job_key(job_id))
|
||||
if not exists:
|
||||
await self._redis.srem(self._active_jobs_key(), job_id)
|
||||
|
||||
logger.info("Cleaned up expired jobs")
|
||||
|
||||
async def get_job_stats(self, user: Optional[str] = None) -> Dict[str, int]:
|
||||
"""
|
||||
Get job statistics.
|
||||
|
||||
Args:
|
||||
user: Optional user filter
|
||||
|
||||
Returns:
|
||||
Statistics dictionary:
|
||||
{
|
||||
"total": 100,
|
||||
"queued": 5,
|
||||
"processing": 10,
|
||||
"completed": 80,
|
||||
"failed": 5
|
||||
}
|
||||
"""
|
||||
if user:
|
||||
jobs = await self.get_user_jobs(user, limit=1000)
|
||||
else:
|
||||
jobs = await self.get_active_jobs(limit=1000)
|
||||
|
||||
stats = {
|
||||
"total": len(jobs),
|
||||
"queued": 0,
|
||||
"processing": 0,
|
||||
"completed": 0,
|
||||
"failed": 0,
|
||||
"cancelled": 0
|
||||
}
|
||||
|
||||
for job in jobs:
|
||||
status = job.get("status", "")
|
||||
if status in stats:
|
||||
stats[status] += 1
|
||||
|
||||
return stats
|
||||
@@ -1,383 +0,0 @@
|
||||
"""
|
||||
Library Desk - Main FastAPI Application
|
||||
|
||||
Following best practices:
|
||||
- Async routes for I/O operations
|
||||
- Dependency injection for configuration
|
||||
- Proper error handling
|
||||
- OpenAPI documentation
|
||||
"""
|
||||
|
||||
from fastapi import FastAPI, HTTPException, Depends
|
||||
from fastapi.middleware.cors import CORSMiddleware
|
||||
from fastapi.staticfiles import StaticFiles
|
||||
from pydantic import BaseModel
|
||||
from typing import Dict, Any
|
||||
import logging
|
||||
from pathlib import Path
|
||||
|
||||
from src.config import Settings, get_settings
|
||||
from src.core.dependencies import verify_api_key
|
||||
|
||||
# Configure logging
|
||||
logging.basicConfig(
|
||||
level=logging.INFO,
|
||||
format="%(asctime)s - %(name)s - %(levelname)s - %(message)s"
|
||||
)
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
# Initialize FastAPI app
|
||||
app = FastAPI(
|
||||
title="Library Desk API",
|
||||
description="Coordination service for The Library system - HybridRAG queries, document ingestion, entity extraction, and mind map generation",
|
||||
version="1.0.0",
|
||||
docs_url="/docs",
|
||||
redoc_url="/redoc",
|
||||
)
|
||||
|
||||
# CORS middleware
|
||||
app.add_middleware(
|
||||
CORSMiddleware,
|
||||
allow_origins=["*"], # Configure appropriately for production
|
||||
allow_credentials=True,
|
||||
allow_methods=["*"],
|
||||
allow_headers=["*"],
|
||||
)
|
||||
|
||||
# Register routers
|
||||
from src.routers import wiki, tools, graph, vector, hybrid_rag, consolidation, ingestion, entity_linking, webhooks
|
||||
|
||||
app.include_router(wiki.router)
|
||||
app.include_router(tools.router)
|
||||
app.include_router(graph.router)
|
||||
app.include_router(vector.router)
|
||||
app.include_router(hybrid_rag.router)
|
||||
app.include_router(consolidation.router)
|
||||
app.include_router(ingestion.router)
|
||||
app.include_router(entity_linking.router)
|
||||
app.include_router(webhooks.router)
|
||||
|
||||
# Mount static files directory for Wiki.js integration scripts
|
||||
static_dir = Path(__file__).parent.parent / "static"
|
||||
if static_dir.exists():
|
||||
app.mount("/static", StaticFiles(directory=str(static_dir)), name="static")
|
||||
logger.info(f"Mounted static files from {static_dir}")
|
||||
|
||||
|
||||
# Response Models
|
||||
class HealthResponse(BaseModel):
|
||||
"""Health check response model."""
|
||||
status: str
|
||||
app_name: str
|
||||
version: str
|
||||
services: Dict[str, Any]
|
||||
|
||||
|
||||
class StatsResponse(BaseModel):
|
||||
"""Statistics response model."""
|
||||
wiki_pages: int
|
||||
neo4j_nodes: int
|
||||
qdrant_vectors: int
|
||||
|
||||
|
||||
# Routes
|
||||
@app.get("/", tags=["Root"])
|
||||
async def root() -> Dict[str, str]:
|
||||
"""Root endpoint."""
|
||||
return {
|
||||
"message": "Library Desk API",
|
||||
"docs": "/docs",
|
||||
"health": "/health"
|
||||
}
|
||||
|
||||
|
||||
@app.get("/health", response_model=HealthResponse, tags=["System"])
|
||||
async def health(settings: Settings = Depends(get_settings)) -> HealthResponse:
|
||||
"""
|
||||
Health check endpoint.
|
||||
Returns status of all connected services.
|
||||
"""
|
||||
from src.core.dependencies import check_service_health
|
||||
|
||||
# Check service connectivity
|
||||
service_health = await check_service_health()
|
||||
|
||||
# Overall status is healthy if at least Neo4j and Qdrant are up
|
||||
all_healthy = service_health.get("neo4j", False) and service_health.get("qdrant", False)
|
||||
overall_status = "healthy" if all_healthy else "degraded"
|
||||
|
||||
return HealthResponse(
|
||||
status=overall_status,
|
||||
app_name=settings.app_name,
|
||||
version=settings.app_version,
|
||||
services={
|
||||
"neo4j": {
|
||||
"url": settings.neo4j_uri,
|
||||
"healthy": service_health.get("neo4j", False)
|
||||
},
|
||||
"qdrant": {
|
||||
"url": settings.qdrant_url,
|
||||
"healthy": service_health.get("qdrant", False)
|
||||
},
|
||||
"wikijs": {
|
||||
"url": settings.wikijs_url,
|
||||
"healthy": service_health.get("wikijs", False)
|
||||
},
|
||||
"searxng": {
|
||||
"url": settings.searxng_url,
|
||||
"healthy": service_health.get("searxng", False)
|
||||
},
|
||||
"ollama": {
|
||||
"url": settings.ollama_url,
|
||||
"model": settings.ollama_model,
|
||||
"healthy": service_health.get("ollama", False)
|
||||
}
|
||||
}
|
||||
)
|
||||
|
||||
|
||||
@app.get("/stats", response_model=StatsResponse, tags=["System"])
|
||||
async def stats(
|
||||
api_key: str = Depends(verify_api_key)
|
||||
) -> StatsResponse:
|
||||
"""
|
||||
Get system statistics.
|
||||
Protected endpoint - requires API key.
|
||||
|
||||
TODO: Implement actual stats gathering from:
|
||||
- Neo4j (node count)
|
||||
- Qdrant (vector count)
|
||||
- Wiki.js (page count)
|
||||
"""
|
||||
return StatsResponse(
|
||||
wiki_pages=0,
|
||||
neo4j_nodes=0,
|
||||
qdrant_vectors=0
|
||||
)
|
||||
|
||||
|
||||
# Ingestion endpoints (for Scheduler integration)
|
||||
@app.post("/ingest/document", tags=["Ingestion"])
|
||||
async def ingest_document(
|
||||
document: Dict[str, Any],
|
||||
api_key: str = Depends(verify_api_key)
|
||||
) -> Dict[str, Any]:
|
||||
"""
|
||||
Ingest a single document for indexing.
|
||||
Used by The Scheduler to add mirrored documentation to the knowledge base.
|
||||
|
||||
Expected fields:
|
||||
- source: str (e.g., "github", "gitea")
|
||||
- repository: str (e.g., "anthropic-cookbook")
|
||||
- path: str (file path)
|
||||
- content: str (document content)
|
||||
- metadata: dict (commit, author, tags, etc.)
|
||||
|
||||
TODO: Implement document ingestion pipeline:
|
||||
1. Chunk content
|
||||
2. Generate embeddings (Ollama)
|
||||
3. Extract entities (NLP)
|
||||
4. Index in Qdrant
|
||||
5. Create graph nodes/relationships in Neo4j
|
||||
"""
|
||||
return {
|
||||
"message": "Document ingestion not yet implemented",
|
||||
"document_id": f"doc_{document.get('path', 'unknown')}",
|
||||
"status": "stub"
|
||||
}
|
||||
|
||||
|
||||
@app.post("/ingest/batch", tags=["Ingestion"])
|
||||
async def batch_ingest(
|
||||
batch: Dict[str, Any],
|
||||
api_key: str = Depends(verify_api_key)
|
||||
) -> Dict[str, Any]:
|
||||
"""
|
||||
Ingest multiple documents in a batch.
|
||||
More efficient than individual ingestion for large syncs.
|
||||
|
||||
TODO: Implement batch processing with task queue
|
||||
"""
|
||||
document_count = len(batch.get("documents", []))
|
||||
return {
|
||||
"message": "Batch ingestion not yet implemented",
|
||||
"batch_id": "batch_stub",
|
||||
"total_documents": document_count,
|
||||
"status": "stub"
|
||||
}
|
||||
|
||||
|
||||
@app.post("/ingest/check-updates", tags=["Ingestion"])
|
||||
async def check_updates(
|
||||
documents: Dict[str, Any],
|
||||
api_key: str = Depends(verify_api_key)
|
||||
) -> Dict[str, Any]:
|
||||
"""
|
||||
Check which documents need updating based on content hashes.
|
||||
Used by Scheduler to determine what changed since last sync.
|
||||
|
||||
TODO: Implement update detection:
|
||||
1. Query existing documents by path
|
||||
2. Compare content hashes
|
||||
3. Return list of updates needed
|
||||
"""
|
||||
return {
|
||||
"message": "Update checking not yet implemented",
|
||||
"updates_needed": [],
|
||||
"up_to_date": [],
|
||||
"new_documents": []
|
||||
}
|
||||
|
||||
|
||||
@app.get("/ingest/status/{document_id}", tags=["Ingestion"])
|
||||
async def get_ingestion_status(
|
||||
document_id: str,
|
||||
api_key: str = Depends(verify_api_key)
|
||||
) -> Dict[str, Any]:
|
||||
"""
|
||||
Get processing status for a document.
|
||||
|
||||
TODO: Implement status tracking
|
||||
"""
|
||||
return {
|
||||
"message": "Status tracking not yet implemented",
|
||||
"document_id": document_id,
|
||||
"status": "unknown"
|
||||
}
|
||||
|
||||
|
||||
@app.get("/ingest/repo-status/{repository}", tags=["Ingestion"])
|
||||
async def get_repo_status(
|
||||
repository: str,
|
||||
api_key: str = Depends(verify_api_key)
|
||||
) -> Dict[str, Any]:
|
||||
"""
|
||||
Get indexing status for an entire repository.
|
||||
|
||||
TODO: Implement repository-level statistics
|
||||
"""
|
||||
return {
|
||||
"message": "Repository status not yet implemented",
|
||||
"repository": repository,
|
||||
"total_documents": 0,
|
||||
"indexed_documents": 0
|
||||
}
|
||||
|
||||
|
||||
# Query endpoints (stubs for future implementation)
|
||||
# NOTE: /query/hybrid is now implemented in routers/hybrid_rag.py
|
||||
|
||||
@app.post("/query/semantic", tags=["Query"])
|
||||
async def semantic_query(
|
||||
query: Dict[str, Any],
|
||||
api_key: str = Depends(verify_api_key)
|
||||
) -> Dict[str, Any]:
|
||||
"""
|
||||
Semantic search via Qdrant.
|
||||
Pure vector similarity search.
|
||||
|
||||
TODO: Implement semantic search
|
||||
"""
|
||||
return {
|
||||
"message": "Semantic search not yet implemented",
|
||||
"query": query
|
||||
}
|
||||
|
||||
|
||||
@app.post("/query/graph", tags=["Query"])
|
||||
async def graph_query(
|
||||
query: Dict[str, Any],
|
||||
api_key: str = Depends(verify_api_key)
|
||||
) -> Dict[str, Any]:
|
||||
"""
|
||||
Graph traversal via Neo4j.
|
||||
Execute Cypher queries.
|
||||
|
||||
TODO: Implement graph queries
|
||||
"""
|
||||
return {
|
||||
"message": "Graph query not yet implemented",
|
||||
"query": query
|
||||
}
|
||||
|
||||
|
||||
# Deduplication endpoints
|
||||
@app.post("/deduplicate/check", tags=["Deduplication"])
|
||||
async def check_duplicates(
|
||||
request: Dict[str, Any],
|
||||
api_key: str = Depends(verify_api_key)
|
||||
) -> Dict[str, Any]:
|
||||
"""
|
||||
Check for duplicate or highly similar documents.
|
||||
Uses vector similarity and graph analysis.
|
||||
|
||||
Expected fields:
|
||||
- document_id: str
|
||||
- similarity_threshold: float (default 0.85)
|
||||
|
||||
TODO: Implement deduplication:
|
||||
1. Get document embedding from Qdrant
|
||||
2. Find similar vectors above threshold
|
||||
3. Check graph relationships
|
||||
4. Return candidates with similarity scores
|
||||
"""
|
||||
document_id = request.get("document_id")
|
||||
threshold = request.get("similarity_threshold", 0.85)
|
||||
|
||||
return {
|
||||
"message": "Deduplication not yet implemented",
|
||||
"document_id": document_id,
|
||||
"threshold": threshold,
|
||||
"duplicates": [],
|
||||
"suggestions": None
|
||||
}
|
||||
|
||||
|
||||
# Application lifecycle
|
||||
@app.on_event("startup")
|
||||
async def startup_event():
|
||||
"""Initialize connections and resources on startup."""
|
||||
from src.core.dependencies import startup_clients
|
||||
from src.services.wiki_change_listener import WikiChangeListener
|
||||
|
||||
settings = get_settings()
|
||||
logger.info(f"Starting {settings.app_name} v{settings.app_version}")
|
||||
logger.info(f"Neo4j: {settings.neo4j_uri}")
|
||||
logger.info(f"Qdrant: {settings.qdrant_url}")
|
||||
logger.info(f"Wiki.js: {settings.wikijs_url}")
|
||||
logger.info(f"SearXNG: {settings.searxng_url}")
|
||||
logger.info(f"Ollama: {settings.ollama_url}")
|
||||
|
||||
# Initialize all service clients
|
||||
await startup_clients()
|
||||
|
||||
# Start Wiki.js change listener (PostgreSQL NOTIFY/LISTEN)
|
||||
# This enables automatic processing of user-edited pages
|
||||
try:
|
||||
wiki_listener = WikiChangeListener()
|
||||
await wiki_listener.start()
|
||||
# Store reference for shutdown
|
||||
app.state.wiki_listener = wiki_listener
|
||||
logger.info("Wiki.js change listener started successfully")
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to start Wiki.js change listener: {e}", exc_info=True)
|
||||
logger.warning("Continuing without change listener - manual page updates will not be auto-processed")
|
||||
|
||||
|
||||
@app.on_event("shutdown")
|
||||
async def shutdown_event():
|
||||
"""Clean up resources on shutdown."""
|
||||
from src.core.dependencies import shutdown_clients
|
||||
|
||||
logger.info("Shutting down Library Desk API")
|
||||
|
||||
# Stop Wiki.js change listener if running
|
||||
if hasattr(app.state, "wiki_listener"):
|
||||
try:
|
||||
await app.state.wiki_listener.stop()
|
||||
logger.info("Wiki.js change listener stopped")
|
||||
except Exception as e:
|
||||
logger.error(f"Error stopping Wiki.js change listener: {e}")
|
||||
|
||||
# Close all service clients
|
||||
await shutdown_clients()
|
||||
@@ -1,49 +0,0 @@
|
||||
"""
|
||||
Knowledge Consolidation models for Librarian processing.
|
||||
|
||||
Used by the consolidation endpoint to process SearchQuery nodes
|
||||
and consolidate knowledge into wiki pages.
|
||||
"""
|
||||
from pydantic import BaseModel, Field
|
||||
from typing import List, Optional, Dict, Any
|
||||
|
||||
|
||||
class ConsolidationRequest(BaseModel):
|
||||
"""Request for knowledge consolidation from search results."""
|
||||
process_limit: int = Field(default=10, ge=1, le=100, description="Max searches to process")
|
||||
lookback_days: int = Field(default=7, ge=1, le=90, description="Process searches from last N days")
|
||||
min_web_results: int = Field(default=2, ge=1, le=20, description="Minimum web results needed")
|
||||
dry_run: bool = Field(default=False, description="If true, analyze but don't create pages")
|
||||
|
||||
|
||||
class SearchQueryInfo(BaseModel):
|
||||
"""Information about a search query to process."""
|
||||
id: str
|
||||
query: str
|
||||
user: str
|
||||
timestamp: str
|
||||
total_results: int
|
||||
web_count: int
|
||||
keywords: List[str] = []
|
||||
|
||||
|
||||
class ConsolidationResult(BaseModel):
|
||||
"""Result of processing a single search query."""
|
||||
search_id: str
|
||||
query: str
|
||||
pages_created: int = 0
|
||||
pages_updated: int = 0
|
||||
entities_added: int = 0
|
||||
error: Optional[str] = None
|
||||
|
||||
|
||||
class ConsolidationResponse(BaseModel):
|
||||
"""Response from knowledge consolidation."""
|
||||
total_found: int = Field(description="Total unprocessed searches found")
|
||||
processed_count: int = Field(description="Successfully processed searches")
|
||||
pages_created: int = Field(description="New wiki pages created")
|
||||
pages_updated: int = Field(description="Existing pages updated")
|
||||
entities_added: int = Field(description="New entities added to graph")
|
||||
errors: List[str] = Field(default=[], description="Error messages")
|
||||
results: List[ConsolidationResult] = Field(description="Per-search results")
|
||||
dry_run: bool = Field(description="Whether this was a dry run")
|
||||
@@ -1,126 +0,0 @@
|
||||
"""
|
||||
Graph models for Library Desk Neo4j operations.
|
||||
|
||||
Provides models for knowledge graph nodes, relationships, and queries.
|
||||
"""
|
||||
|
||||
from pydantic import BaseModel, Field
|
||||
from typing import List, Dict, Any, Optional
|
||||
from datetime import datetime
|
||||
|
||||
|
||||
class GraphNode(BaseModel):
|
||||
"""Graph node representation."""
|
||||
id: str = Field(..., description="Node ID")
|
||||
labels: List[str] = Field(..., description="Node labels")
|
||||
properties: Dict[str, Any] = Field(default_factory=dict, description="Node properties")
|
||||
|
||||
|
||||
class GraphRelationship(BaseModel):
|
||||
"""Graph relationship representation."""
|
||||
id: str = Field(..., description="Relationship ID")
|
||||
type: str = Field(..., description="Relationship type")
|
||||
start_node: str = Field(..., description="Start node ID")
|
||||
end_node: str = Field(..., description="End node ID")
|
||||
properties: Dict[str, Any] = Field(default_factory=dict, description="Relationship properties")
|
||||
|
||||
|
||||
class GraphNodeDetail(BaseModel):
|
||||
"""Detailed node with relationships."""
|
||||
node: GraphNode = Field(..., description="Node data")
|
||||
relationships: List[GraphRelationship] = Field(
|
||||
default_factory=list,
|
||||
description="Connected relationships"
|
||||
)
|
||||
related_nodes: List[GraphNode] = Field(
|
||||
default_factory=list,
|
||||
description="Connected nodes"
|
||||
)
|
||||
|
||||
|
||||
class CypherQueryRequest(BaseModel):
|
||||
"""Request to execute a Cypher query."""
|
||||
query: str = Field(..., description="Cypher query to execute")
|
||||
parameters: Dict[str, Any] = Field(
|
||||
default_factory=dict,
|
||||
description="Query parameters"
|
||||
)
|
||||
user: str = Field(
|
||||
default="jpmschweitzer",
|
||||
description="User for filtering (automatically scopes query)"
|
||||
)
|
||||
|
||||
|
||||
class CypherQueryResponse(BaseModel):
|
||||
"""Response from Cypher query execution."""
|
||||
results: List[Dict[str, Any]] = Field(..., description="Query results")
|
||||
count: int = Field(..., description="Number of results")
|
||||
query_time_ms: float = Field(..., description="Query execution time in milliseconds")
|
||||
|
||||
|
||||
class UpdateFromPageRequest(BaseModel):
|
||||
"""Request to update graph from a wiki page."""
|
||||
page_id: int = Field(..., description="Wiki page ID to process")
|
||||
user: str = Field(
|
||||
default="jpmschweitzer",
|
||||
description="User identifier for namespace scoping"
|
||||
)
|
||||
force_refresh: bool = Field(
|
||||
default=False,
|
||||
description="Force re-extraction even if page hasn't changed"
|
||||
)
|
||||
|
||||
|
||||
class EntityMention(BaseModel):
|
||||
"""Extracted entity mention."""
|
||||
text: str = Field(..., description="Entity text")
|
||||
type: str = Field(..., description="Entity type (Person, Project, Concept, etc.)")
|
||||
confidence: float = Field(default=1.0, description="Extraction confidence (0-1)")
|
||||
|
||||
|
||||
class GraphUpdateSummary(BaseModel):
|
||||
"""Summary of graph update operation."""
|
||||
page_id: int = Field(..., description="Page ID processed")
|
||||
page_title: str = Field(..., description="Page title")
|
||||
nodes_created: int = Field(default=0, description="New nodes created")
|
||||
nodes_updated: int = Field(default=0, description="Existing nodes updated")
|
||||
relationships_created: int = Field(default=0, description="New relationships created")
|
||||
entities_extracted: List[EntityMention] = Field(
|
||||
default_factory=list,
|
||||
description="Entities extracted from page"
|
||||
)
|
||||
processing_time_ms: float = Field(..., description="Processing time in milliseconds")
|
||||
success: bool = Field(default=True, description="Whether update succeeded")
|
||||
error_message: Optional[str] = Field(default=None, description="Error message if failed")
|
||||
|
||||
|
||||
class NodeListResponse(BaseModel):
|
||||
"""Response for node listing."""
|
||||
nodes: List[GraphNode] = Field(..., description="List of nodes")
|
||||
total: int = Field(..., description="Total number of nodes")
|
||||
user: str = Field(..., description="User filter applied")
|
||||
|
||||
|
||||
class MindMapNode(BaseModel):
|
||||
"""Mind map node for visualization."""
|
||||
id: str = Field(..., description="Node ID")
|
||||
label: str = Field(..., description="Node label/name")
|
||||
type: str = Field(..., description="Node type")
|
||||
size: int = Field(default=10, description="Visual size")
|
||||
color: Optional[str] = Field(default=None, description="Node color")
|
||||
|
||||
|
||||
class MindMapLink(BaseModel):
|
||||
"""Mind map link for visualization."""
|
||||
source: str = Field(..., description="Source node ID")
|
||||
target: str = Field(..., description="Target node ID")
|
||||
type: str = Field(..., description="Relationship type")
|
||||
strength: float = Field(default=1.0, description="Link strength")
|
||||
|
||||
|
||||
class MindMapResponse(BaseModel):
|
||||
"""Mind map data for D3.js or similar visualization."""
|
||||
nodes: List[MindMapNode] = Field(..., description="Graph nodes")
|
||||
links: List[MindMapLink] = Field(..., description="Graph edges")
|
||||
center_node: str = Field(..., description="Central node ID")
|
||||
depth: int = Field(..., description="Traversal depth")
|
||||
@@ -1,87 +0,0 @@
|
||||
"""
|
||||
HybridRAG models for multi-source search with RRF fusion.
|
||||
|
||||
Combines vector search (Qdrant), knowledge graph (Neo4j), and web search (SearXNG)
|
||||
with Reciprocal Rank Fusion and LLM re-ranking.
|
||||
"""
|
||||
|
||||
from pydantic import BaseModel, Field
|
||||
from typing import List, Optional, Dict, Any
|
||||
|
||||
|
||||
class HybridRAGConfig(BaseModel):
|
||||
"""Configuration for HybridRAG query."""
|
||||
vector_limit: int = Field(default=10, ge=1, le=50, description="Max vector results")
|
||||
graph_limit: int = Field(default=10, ge=1, le=50, description="Max graph results")
|
||||
web_limit: int = Field(default=5, ge=1, le=20, description="Max web results")
|
||||
enable_vector: bool = Field(default=True, description="Enable vector search")
|
||||
enable_graph: bool = Field(default=True, description="Enable graph search")
|
||||
enable_web: bool = Field(default=True, description="Enable web search")
|
||||
enable_reranking: bool = Field(default=True, description="Enable LLM re-ranking")
|
||||
enable_enrichment: bool = Field(default=True, description="Enable graph enrichment")
|
||||
final_result_count: int = Field(default=10, ge=1, le=50, description="Final results to return")
|
||||
rrf_k: int = Field(default=60, ge=1, le=100, description="RRF constant")
|
||||
|
||||
|
||||
class RelatedDossier(BaseModel):
|
||||
"""Related document metadata from graph enrichment."""
|
||||
page_id: int
|
||||
title: str
|
||||
path: str
|
||||
tag: str
|
||||
shared_entities: int
|
||||
|
||||
|
||||
class HybridRAGResult(BaseModel):
|
||||
"""Single result from HybridRAG query."""
|
||||
source_type: str = Field(..., description="Source: 'vector', 'graph', 'web'")
|
||||
title: str
|
||||
content: str
|
||||
url: Optional[str] = Field(None, description="URL for web results")
|
||||
page_id: Optional[int] = Field(None, description="Page ID for wiki results")
|
||||
page_path: Optional[str] = Field(None, description="Wiki page path")
|
||||
rrf_score: float = Field(..., description="Reciprocal Rank Fusion score")
|
||||
final_rank: int = Field(..., description="Final rank after re-ranking")
|
||||
sources: List[str] = Field(..., description="Which sources included this result")
|
||||
related_dossiers: List[RelatedDossier] = Field(default=[], description="Related documents via shared entities")
|
||||
metadata: Dict[str, Any] = Field(default={}, description="Additional metadata")
|
||||
|
||||
|
||||
class TimingBreakdown(BaseModel):
|
||||
"""Performance timing breakdown for each phase."""
|
||||
query_enhancement_ms: float = Field(..., description="Phase 0: Keyword/synonym extraction")
|
||||
vector_ms: float = Field(..., description="Phase 1: Vector search")
|
||||
graph_ms: float = Field(..., description="Phase 1: Graph search")
|
||||
web_ms: float = Field(..., description="Phase 1: Web search")
|
||||
fusion_ms: float = Field(..., description="Phase 2: RRF fusion")
|
||||
enrichment_ms: float = Field(..., description="Phase 3: Graph enrichment")
|
||||
reranking_ms: float = Field(..., description="Phase 4: LLM re-ranking")
|
||||
persistence_ms: float = Field(..., description="Phase 6: Search persistence")
|
||||
total_ms: float = Field(..., description="Total end-to-end time")
|
||||
|
||||
|
||||
class KeywordExtraction(BaseModel):
|
||||
"""Extracted keywords and synonyms from query enhancement."""
|
||||
core_keywords: List[str] = Field(default=[], description="Primary keywords")
|
||||
entities: List[str] = Field(default=[], description="Named entities")
|
||||
synonyms: Dict[str, List[str]] = Field(default={}, description="Synonyms map")
|
||||
expansions: Dict[str, List[str]] = Field(default={}, description="Abbreviation expansions")
|
||||
|
||||
|
||||
class HybridRAGResponse(BaseModel):
|
||||
"""Response from HybridRAG query."""
|
||||
query: str = Field(..., description="Original search query")
|
||||
keywords: KeywordExtraction = Field(..., description="Extracted keywords/synonyms")
|
||||
results: List[HybridRAGResult] = Field(..., description="Ranked search results")
|
||||
context: str = Field(..., description="Formatted context for LLM consumption")
|
||||
source_counts: Dict[str, int] = Field(..., description="Result counts by source")
|
||||
total_results: int = Field(..., description="Total number of results")
|
||||
timing: TimingBreakdown = Field(..., description="Performance breakdown")
|
||||
config_used: HybridRAGConfig = Field(..., description="Configuration used")
|
||||
search_id: Optional[str] = Field(None, description="Search ID for Librarian tracking")
|
||||
|
||||
|
||||
class HybridRAGRequest(BaseModel):
|
||||
"""Request for HybridRAG query."""
|
||||
query: str = Field(..., min_length=1, max_length=500, description="Search query")
|
||||
config: Optional[HybridRAGConfig] = Field(None, description="Custom configuration")
|
||||
@@ -1,67 +0,0 @@
|
||||
"""
|
||||
Pydantic models for Document Ingestion system.
|
||||
"""
|
||||
from pydantic import BaseModel, Field
|
||||
from typing import Optional, List, Dict, Any
|
||||
from datetime import datetime
|
||||
|
||||
|
||||
class IngestionRequest(BaseModel):
|
||||
"""Request to ingest a wiki page."""
|
||||
page_id: int = Field(..., description="Wiki page ID to ingest")
|
||||
user: str = Field(default="jpmschweitzer", description="User identifier")
|
||||
force_refresh: bool = Field(
|
||||
default=False,
|
||||
description="Force re-ingestion even if page hasn't changed"
|
||||
)
|
||||
skip_vectors: bool = Field(default=False, description="Skip vector embedding generation")
|
||||
skip_graph: bool = Field(default=False, description="Skip graph entity extraction")
|
||||
|
||||
|
||||
class BatchIngestionRequest(BaseModel):
|
||||
"""Request to ingest multiple wiki pages."""
|
||||
page_ids: List[int] = Field(..., description="List of wiki page IDs to ingest")
|
||||
user: str = Field(default="jpmschweitzer", description="User identifier")
|
||||
force_refresh: bool = Field(default=False)
|
||||
skip_vectors: bool = Field(default=False)
|
||||
skip_graph: bool = Field(default=False)
|
||||
max_concurrent: int = Field(
|
||||
default=3,
|
||||
ge=1,
|
||||
le=10,
|
||||
description="Maximum concurrent ingestion tasks"
|
||||
)
|
||||
|
||||
|
||||
class IngestionResult(BaseModel):
|
||||
"""Result of a single page ingestion."""
|
||||
page_id: int
|
||||
page_title: str
|
||||
page_path: Optional[str] = None
|
||||
success: bool
|
||||
error: Optional[str] = None
|
||||
vector_chunks_created: int = 0
|
||||
graph_entities_extracted: int = 0
|
||||
graph_relationships_created: int = 0
|
||||
processing_time_ms: float
|
||||
|
||||
|
||||
class BatchIngestionResult(BaseModel):
|
||||
"""Result of batch ingestion."""
|
||||
total_pages: int
|
||||
successful: int
|
||||
failed: int
|
||||
results: List[IngestionResult]
|
||||
total_processing_time_ms: float
|
||||
|
||||
|
||||
class IngestionStatus(BaseModel):
|
||||
"""Status of an ingestion job."""
|
||||
job_id: str
|
||||
status: str # "queued", "processing", "completed", "failed"
|
||||
progress: int # 0-100
|
||||
page_id: Optional[int] = None
|
||||
result: Optional[IngestionResult] = None
|
||||
created_at: datetime
|
||||
started_at: Optional[datetime] = None
|
||||
completed_at: Optional[datetime] = None
|
||||
@@ -1,61 +0,0 @@
|
||||
"""
|
||||
Tool catalog models for Library Desk.
|
||||
|
||||
Provides simplified tool definitions optimized for AI agent consumption.
|
||||
"""
|
||||
|
||||
from pydantic import BaseModel, Field
|
||||
from typing import List, Dict, Any, Optional
|
||||
from enum import Enum
|
||||
|
||||
|
||||
class ParameterType(str, Enum):
|
||||
"""Parameter data types."""
|
||||
STRING = "string"
|
||||
INTEGER = "integer"
|
||||
BOOLEAN = "boolean"
|
||||
ARRAY = "array"
|
||||
OBJECT = "object"
|
||||
|
||||
|
||||
class ToolParameter(BaseModel):
|
||||
"""Tool parameter definition."""
|
||||
name: str = Field(..., description="Parameter name")
|
||||
type: ParameterType = Field(..., description="Parameter type")
|
||||
description: str = Field(..., description="Parameter description")
|
||||
required: bool = Field(default=False, description="Whether parameter is required")
|
||||
default: Optional[Any] = Field(default=None, description="Default value if not required")
|
||||
example: Optional[Any] = Field(default=None, description="Example value")
|
||||
|
||||
|
||||
class ToolDefinition(BaseModel):
|
||||
"""Individual tool definition."""
|
||||
name: str = Field(..., description="Tool identifier (e.g., 'wiki_create_page')")
|
||||
category: str = Field(..., description="Tool category (e.g., 'wiki', 'graph')")
|
||||
description: str = Field(..., description="What this tool does")
|
||||
method: str = Field(..., description="HTTP method (GET, POST, PUT, DELETE)")
|
||||
endpoint: str = Field(..., description="API endpoint path")
|
||||
parameters: List[ToolParameter] = Field(default_factory=list, description="Tool parameters")
|
||||
returns: str = Field(..., description="What the tool returns")
|
||||
example: Optional[Dict[str, Any]] = Field(default=None, description="Example request")
|
||||
fast: bool = Field(default=True, description="Whether operation completes quickly (<5s)")
|
||||
|
||||
|
||||
class CategoryInfo(BaseModel):
|
||||
"""Tool category information."""
|
||||
name: str = Field(..., description="Category name")
|
||||
description: str = Field(..., description="Category description")
|
||||
tool_count: int = Field(..., description="Number of tools in category")
|
||||
|
||||
|
||||
class ToolCatalog(BaseModel):
|
||||
"""Complete tool catalog response."""
|
||||
service: str = Field(default="library-desk", description="Service name")
|
||||
version: str = Field(default="1.0.0", description="API version")
|
||||
base_url: str = Field(..., description="Base URL for API")
|
||||
categories: List[CategoryInfo] = Field(..., description="Available categories")
|
||||
tools: List[ToolDefinition] = Field(..., description="All available tools")
|
||||
authentication: str = Field(
|
||||
default="Bearer token via Authorization header",
|
||||
description="Authentication method"
|
||||
)
|
||||
@@ -1,100 +0,0 @@
|
||||
"""
|
||||
Vector models for Library Desk Qdrant operations.
|
||||
|
||||
Provides models for semantic search, document chunks, and embeddings.
|
||||
"""
|
||||
|
||||
from pydantic import BaseModel, Field
|
||||
from typing import List, Dict, Any, Optional
|
||||
|
||||
|
||||
class DocumentChunk(BaseModel):
|
||||
"""Document chunk with embedding."""
|
||||
chunk_id: str = Field(..., description="Unique chunk ID (page_id:chunk_index)")
|
||||
page_id: int = Field(..., description="Wiki page ID")
|
||||
chunk_index: int = Field(..., description="Chunk index within document")
|
||||
content: str = Field(..., description="Chunk text content")
|
||||
metadata: Dict[str, Any] = Field(default_factory=dict, description="Additional metadata")
|
||||
|
||||
|
||||
class SearchResult(BaseModel):
|
||||
"""Semantic search result."""
|
||||
chunk_id: str = Field(..., description="Chunk ID")
|
||||
page_id: int = Field(..., description="Wiki page ID")
|
||||
page_title: Optional[str] = Field(None, description="Page title")
|
||||
page_path: Optional[str] = Field(None, description="Page path")
|
||||
chunk_index: int = Field(..., description="Chunk index")
|
||||
content: str = Field(..., description="Chunk content")
|
||||
score: float = Field(..., description="Similarity score (0-1)")
|
||||
metadata: Dict[str, Any] = Field(default_factory=dict, description="Additional metadata")
|
||||
|
||||
|
||||
class SearchRequest(BaseModel):
|
||||
"""Semantic search request."""
|
||||
query: str = Field(..., min_length=1, description="Search query")
|
||||
user: str = Field(default="jpmschweitzer", description="User identifier")
|
||||
limit: int = Field(default=10, ge=1, le=100, description="Maximum results")
|
||||
score_threshold: float = Field(default=0.5, ge=0.0, le=1.0, description="Minimum similarity score")
|
||||
|
||||
|
||||
class SearchResponse(BaseModel):
|
||||
"""Semantic search response."""
|
||||
query: str = Field(..., description="Search query")
|
||||
results: List[SearchResult] = Field(..., description="Search results")
|
||||
total: int = Field(..., description="Number of results")
|
||||
user: str = Field(..., description="User filter applied")
|
||||
|
||||
|
||||
class VectorUpdateRequest(BaseModel):
|
||||
"""Request to update vectors from a wiki page."""
|
||||
page_id: int = Field(..., description="Wiki page ID to process")
|
||||
user: str = Field(
|
||||
default="jpmschweitzer",
|
||||
description="User identifier for namespace scoping"
|
||||
)
|
||||
force_refresh: bool = Field(
|
||||
default=False,
|
||||
description="Force re-embedding even if page hasn't changed"
|
||||
)
|
||||
|
||||
|
||||
class VectorUpdateSummary(BaseModel):
|
||||
"""Summary of vector update operation."""
|
||||
page_id: int = Field(..., description="Page ID processed")
|
||||
page_title: str = Field(..., description="Page title")
|
||||
chunks_created: int = Field(default=0, description="New chunks created")
|
||||
chunks_updated: int = Field(default=0, description="Existing chunks updated")
|
||||
chunks_deleted: int = Field(default=0, description="Old chunks deleted")
|
||||
total_chunks: int = Field(default=0, description="Total chunks for this page")
|
||||
embedding_dim: int = Field(default=768, description="Embedding dimensionality")
|
||||
processing_time_ms: float = Field(..., description="Processing time in milliseconds")
|
||||
success: bool = Field(default=True, description="Whether update succeeded")
|
||||
error_message: Optional[str] = Field(default=None, description="Error message if failed")
|
||||
|
||||
|
||||
class CollectionInfo(BaseModel):
|
||||
"""Qdrant collection information."""
|
||||
name: str = Field(..., description="Collection name")
|
||||
vectors_count: int = Field(..., description="Number of vectors")
|
||||
points_count: int = Field(..., description="Number of points")
|
||||
segments_count: int = Field(..., description="Number of segments")
|
||||
|
||||
|
||||
class CollectionListResponse(BaseModel):
|
||||
"""List of Qdrant collections."""
|
||||
collections: List[CollectionInfo] = Field(..., description="List of collections")
|
||||
total: int = Field(..., description="Total number of collections")
|
||||
|
||||
|
||||
class DeletePageChunksRequest(BaseModel):
|
||||
"""Request to delete all chunks for a page."""
|
||||
page_id: int = Field(..., description="Wiki page ID")
|
||||
user: str = Field(default="jpmschweitzer", description="User identifier")
|
||||
|
||||
|
||||
class DeletePageChunksResponse(BaseModel):
|
||||
"""Response from deleting page chunks."""
|
||||
page_id: int = Field(..., description="Page ID")
|
||||
chunks_deleted: int = Field(..., description="Number of chunks deleted")
|
||||
success: bool = Field(..., description="Whether deletion succeeded")
|
||||
message: str = Field(..., description="Result message")
|
||||
@@ -1,192 +0,0 @@
|
||||
"""
|
||||
Pydantic models for Wiki.js operations.
|
||||
|
||||
Models for:
|
||||
- Wiki pages (CRUD operations)
|
||||
- Dossiers (tag-based collections)
|
||||
- Search results
|
||||
"""
|
||||
|
||||
from pydantic import BaseModel, Field, field_validator
|
||||
from typing import Optional, List
|
||||
from datetime import datetime
|
||||
|
||||
|
||||
# Base models
|
||||
class WikiPageBase(BaseModel):
|
||||
"""Base wiki page fields."""
|
||||
title: str = Field(..., min_length=1, max_length=500, description="Page title")
|
||||
description: Optional[str] = Field(None, max_length=1000, description="Page description")
|
||||
tags: List[str] = Field(default_factory=list, description="Tags (for dossier organization)")
|
||||
is_published: bool = Field(default=True, description="Whether page is published")
|
||||
|
||||
@field_validator("tags")
|
||||
@classmethod
|
||||
def validate_tags(cls, v: List[str]) -> List[str]:
|
||||
"""Validate and clean tags."""
|
||||
# Remove empty tags and strip whitespace
|
||||
cleaned = [tag.strip() for tag in v if tag.strip()]
|
||||
# Ensure uniqueness
|
||||
return list(set(cleaned))
|
||||
|
||||
|
||||
class WikiPageCreate(WikiPageBase):
|
||||
"""Request model for creating a wiki page."""
|
||||
content: str = Field(..., description="Page content (markdown)")
|
||||
path: str = Field(..., min_length=1, max_length=500, description="Page path (e.g., '/projects/library-desk')")
|
||||
editor: str = Field(default="markdown", description="Editor type")
|
||||
user: Optional[str] = Field(None, description="User identifier (defaults to configured user)")
|
||||
|
||||
@field_validator("path")
|
||||
@classmethod
|
||||
def validate_path(cls, v: str) -> str:
|
||||
"""Validate page path."""
|
||||
# Ensure path starts with /
|
||||
if not v.startswith("/"):
|
||||
v = f"/{v}"
|
||||
# Remove trailing slash
|
||||
if v.endswith("/") and v != "/":
|
||||
v = v.rstrip("/")
|
||||
return v
|
||||
|
||||
|
||||
class WikiPageUpdate(BaseModel):
|
||||
"""Request model for updating a wiki page."""
|
||||
content: Optional[str] = Field(None, description="Updated content")
|
||||
title: Optional[str] = Field(None, min_length=1, max_length=500, description="Updated title")
|
||||
description: Optional[str] = Field(None, max_length=1000, description="Updated description")
|
||||
tags: Optional[List[str]] = Field(None, description="Updated tags")
|
||||
|
||||
@field_validator("tags")
|
||||
@classmethod
|
||||
def validate_tags(cls, v: Optional[List[str]]) -> Optional[List[str]]:
|
||||
"""Validate and clean tags."""
|
||||
if v is None:
|
||||
return None
|
||||
cleaned = [tag.strip() for tag in v if tag.strip()]
|
||||
return list(set(cleaned))
|
||||
|
||||
|
||||
class WikiPage(WikiPageBase):
|
||||
"""Response model for a wiki page."""
|
||||
id: int = Field(..., description="Page ID")
|
||||
path: str = Field(..., description="Page path")
|
||||
content: Optional[str] = Field(None, description="Page content")
|
||||
created_at: Optional[str] = Field(None, description="Creation timestamp")
|
||||
updated_at: Optional[str] = Field(None, description="Last update timestamp")
|
||||
editor: Optional[str] = Field(None, description="Editor type")
|
||||
|
||||
class Config:
|
||||
from_attributes = True
|
||||
|
||||
|
||||
class WikiPageSummary(BaseModel):
|
||||
"""Summarized wiki page (for list responses)."""
|
||||
id: int = Field(..., description="Page ID")
|
||||
path: str = Field(..., description="Page path")
|
||||
title: str = Field(..., description="Page title")
|
||||
description: Optional[str] = Field(None, description="Page description")
|
||||
tags: List[str] = Field(default_factory=list, description="Page tags")
|
||||
updated_at: Optional[str] = Field(None, description="Last update timestamp")
|
||||
is_published: bool = Field(..., description="Publication status")
|
||||
|
||||
|
||||
class WikiPageList(BaseModel):
|
||||
"""Response model for list of pages."""
|
||||
pages: List[WikiPageSummary] = Field(..., description="List of pages")
|
||||
total: int = Field(..., description="Total number of pages")
|
||||
filtered_by_tag: Optional[str] = Field(None, description="Tag filter applied")
|
||||
user: str = Field(..., description="User namespace")
|
||||
|
||||
|
||||
# Dossier models
|
||||
class DossierCreate(BaseModel):
|
||||
"""Request model for creating a dossier."""
|
||||
name: str = Field(..., min_length=1, max_length=100, description="Dossier name (becomes a tag)")
|
||||
title: str = Field(..., min_length=1, max_length=200, description="Human-readable title")
|
||||
description: str = Field(..., min_length=1, description="Dossier description")
|
||||
create_index_page: bool = Field(default=True, description="Create an index page for the dossier")
|
||||
user: Optional[str] = Field(None, description="User identifier")
|
||||
|
||||
@field_validator("name")
|
||||
@classmethod
|
||||
def validate_name(cls, v: str) -> str:
|
||||
"""Validate dossier name (will be used as tag)."""
|
||||
# Convert to lowercase, replace spaces with hyphens
|
||||
name = v.lower().strip()
|
||||
name = name.replace(" ", "-")
|
||||
# Remove special characters except hyphens and underscores
|
||||
name = "".join(c for c in name if c.isalnum() or c in "-_")
|
||||
if not name:
|
||||
raise ValueError("Dossier name must contain alphanumeric characters")
|
||||
return name
|
||||
|
||||
|
||||
class DossierInfo(BaseModel):
|
||||
"""Response model for dossier information."""
|
||||
name: str = Field(..., description="Dossier name (tag)")
|
||||
title: str = Field(..., description="Dossier title")
|
||||
description: str = Field(..., description="Dossier description")
|
||||
page_count: int = Field(..., description="Number of pages in dossier")
|
||||
index_page_id: Optional[int] = Field(None, description="ID of index page")
|
||||
index_page_path: Optional[str] = Field(None, description="Path to index page")
|
||||
created_at: Optional[str] = Field(None, description="Creation timestamp")
|
||||
|
||||
|
||||
class DossierList(BaseModel):
|
||||
"""Response model for list of dossiers."""
|
||||
dossiers: List[DossierInfo] = Field(..., description="List of dossiers")
|
||||
total: int = Field(..., description="Total number of dossiers")
|
||||
user: str = Field(..., description="User namespace")
|
||||
|
||||
|
||||
# Search models
|
||||
class WikiSearchResult(BaseModel):
|
||||
"""Search result item."""
|
||||
id: int = Field(..., description="Page ID")
|
||||
path: str = Field(..., description="Page path")
|
||||
title: str = Field(..., description="Page title")
|
||||
description: Optional[str] = Field(None, description="Page description")
|
||||
relevance: Optional[float] = Field(None, description="Search relevance score")
|
||||
|
||||
|
||||
class WikiSearchResponse(BaseModel):
|
||||
"""Response model for search results."""
|
||||
results: List[WikiSearchResult] = Field(..., description="Search results")
|
||||
query: str = Field(..., description="Search query")
|
||||
total: int = Field(..., description="Total results found")
|
||||
|
||||
|
||||
# Move/rename models
|
||||
class WikiPageMove(BaseModel):
|
||||
"""Request model for moving/renaming a page."""
|
||||
new_path: str = Field(..., min_length=1, description="New page path")
|
||||
locale: str = Field(default="en", description="Page locale")
|
||||
|
||||
@field_validator("new_path")
|
||||
@classmethod
|
||||
def validate_new_path(cls, v: str) -> str:
|
||||
"""Validate new path."""
|
||||
if not v.startswith("/"):
|
||||
v = f"/{v}"
|
||||
if v.endswith("/") and v != "/":
|
||||
v = v.rstrip("/")
|
||||
return v
|
||||
|
||||
|
||||
# Response models for operations
|
||||
class WikiOperationResponse(BaseModel):
|
||||
"""Generic response for wiki operations."""
|
||||
success: bool = Field(..., description="Whether operation succeeded")
|
||||
message: str = Field(..., description="Operation message")
|
||||
page_id: Optional[int] = Field(None, description="Page ID (if applicable)")
|
||||
page_path: Optional[str] = Field(None, description="Page path (if applicable)")
|
||||
|
||||
|
||||
class DossierOperationResponse(BaseModel):
|
||||
"""Response for dossier operations."""
|
||||
success: bool = Field(..., description="Whether operation succeeded")
|
||||
message: str = Field(..., description="Operation message")
|
||||
dossier_name: str = Field(..., description="Dossier name")
|
||||
index_page_id: Optional[int] = Field(None, description="Index page ID (if created)")
|
||||
index_page_path: Optional[str] = Field(None, description="Index page path (if created)")
|
||||
@@ -1,160 +0,0 @@
|
||||
"""
|
||||
Knowledge Consolidation router for Librarian processing.
|
||||
|
||||
Provides endpoints for the Scheduler to trigger knowledge consolidation
|
||||
from HybridRAG search results into wiki pages.
|
||||
"""
|
||||
|
||||
from fastapi import APIRouter, HTTPException, Depends
|
||||
import logging
|
||||
|
||||
from src.models.consolidation import ConsolidationRequest, ConsolidationResponse
|
||||
from src.services.consolidation_service import ConsolidationService
|
||||
from src.core.dependencies import (
|
||||
Neo4jDep, OllamaDep, WikiJSDep,
|
||||
verify_api_key, get_settings, get_ingestion_service
|
||||
)
|
||||
from src.config import Settings
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
router = APIRouter(prefix="/consolidate", tags=["Consolidation"])
|
||||
|
||||
|
||||
# Dependency to get Consolidation service
|
||||
def get_consolidation_service(
|
||||
neo4j_client: Neo4jDep,
|
||||
ollama_client: OllamaDep,
|
||||
wiki_client: WikiJSDep,
|
||||
settings: Settings = Depends(get_settings)
|
||||
) -> ConsolidationService:
|
||||
"""Get ConsolidationService instance with all dependencies."""
|
||||
return ConsolidationService(
|
||||
neo4j=neo4j_client,
|
||||
ollama=ollama_client,
|
||||
wiki=wiki_client,
|
||||
settings=settings,
|
||||
ingestion_service=get_ingestion_service()
|
||||
)
|
||||
|
||||
|
||||
@router.post("/knowledge", response_model=ConsolidationResponse)
|
||||
async def consolidate_knowledge(
|
||||
request: ConsolidationRequest,
|
||||
consolidation_service: ConsolidationService = Depends(get_consolidation_service),
|
||||
api_key: str = Depends(verify_api_key)
|
||||
):
|
||||
"""
|
||||
Consolidate knowledge from HybridRAG search results into wiki pages.
|
||||
|
||||
**Librarian Task** - Processes unprocessed SearchQuery nodes to:
|
||||
1. Find searches with web results from last N days
|
||||
2. Analyze web content with LLM for novel information
|
||||
3. Create new wiki pages for new concepts/technologies
|
||||
4. Update existing pages with new facts and citations
|
||||
5. Add new entities to knowledge graph
|
||||
6. Mark SearchQuery nodes as processed
|
||||
|
||||
**Typically called by The Scheduler** on a periodic basis (e.g., hourly).
|
||||
|
||||
**Parameters:**
|
||||
- `process_limit`: Maximum searches to process per run (default: 10)
|
||||
- `lookback_days`: Only process searches from last N days (default: 7)
|
||||
- `min_web_results`: Minimum web results required to consolidate (default: 2)
|
||||
- `dry_run`: If true, analyze but don't create pages (default: false)
|
||||
|
||||
**Returns:**
|
||||
- `total_found`: Number of unprocessed searches found
|
||||
- `processed_count`: Successfully processed searches
|
||||
- `pages_created`: New wiki pages created
|
||||
- `pages_updated`: Existing pages updated with new facts
|
||||
- `entities_added`: New entities added to knowledge graph
|
||||
- `errors`: List of error messages if any failed
|
||||
- `results`: Per-search processing results
|
||||
|
||||
**Example Request:**
|
||||
```json
|
||||
{
|
||||
"process_limit": 10,
|
||||
"lookback_days": 7,
|
||||
"min_web_results": 2,
|
||||
"dry_run": false
|
||||
}
|
||||
```
|
||||
|
||||
**Example Response:**
|
||||
```json
|
||||
{
|
||||
"total_found": 5,
|
||||
"processed_count": 4,
|
||||
"pages_created": 2,
|
||||
"pages_updated": 3,
|
||||
"entities_added": 7,
|
||||
"errors": ["Search abc123: Failed to parse response"],
|
||||
"results": [
|
||||
{
|
||||
"search_id": "uuid-1",
|
||||
"query": "docker orchestration kubernetes",
|
||||
"pages_created": 1,
|
||||
"pages_updated": 1,
|
||||
"entities_added": 3
|
||||
}
|
||||
],
|
||||
"dry_run": false
|
||||
}
|
||||
```
|
||||
|
||||
**Scheduler Task Configuration:**
|
||||
```json
|
||||
{
|
||||
"task_name": "knowledge_consolidation",
|
||||
"service": "library-desk",
|
||||
"executor": "rest_api_executor",
|
||||
"priority": 50,
|
||||
"minute": 0,
|
||||
"hour": -1,
|
||||
"description": "Hourly knowledge consolidation from search results",
|
||||
"config": {
|
||||
"url": "http://library-desk:8089/consolidate/knowledge",
|
||||
"method": "POST",
|
||||
"payload": {
|
||||
"process_limit": 10,
|
||||
"lookback_days": 7,
|
||||
"min_web_results": 2,
|
||||
"dry_run": false
|
||||
},
|
||||
"auth": {
|
||||
"type": "bearer",
|
||||
"token": "${LIBRARY_DESK_API_KEY}"
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
"""
|
||||
try:
|
||||
logger.info(
|
||||
f"Knowledge consolidation requested: "
|
||||
f"limit={request.process_limit}, lookback={request.lookback_days}d, "
|
||||
f"dry_run={request.dry_run}"
|
||||
)
|
||||
|
||||
response = await consolidation_service.consolidate_knowledge(
|
||||
process_limit=request.process_limit,
|
||||
lookback_days=request.lookback_days,
|
||||
min_web_results=request.min_web_results,
|
||||
dry_run=request.dry_run
|
||||
)
|
||||
|
||||
logger.info(
|
||||
f"Consolidation completed: {response.processed_count}/{response.total_found} searches, "
|
||||
f"{response.pages_created} pages created, {response.pages_updated} updated"
|
||||
)
|
||||
|
||||
return response
|
||||
|
||||
except ValueError as e:
|
||||
logger.error(f"Invalid request: {e}")
|
||||
raise HTTPException(status_code=400, detail=str(e))
|
||||
except Exception as e:
|
||||
logger.error(f"Knowledge consolidation failed: {e}", exc_info=True)
|
||||
raise HTTPException(status_code=500, detail="Consolidation failed")
|
||||
@@ -1,474 +0,0 @@
|
||||
"""
|
||||
Entity Linking Router
|
||||
|
||||
Finds and links mentions of existing entities in wiki pages to the knowledge graph.
|
||||
Creates both:
|
||||
1. Graph relationships (MENTIONS) in Neo4j
|
||||
2. Hyperlinks in wiki page content (markdown links)
|
||||
"""
|
||||
import logging
|
||||
from fastapi import APIRouter, Depends, HTTPException
|
||||
from pydantic import BaseModel
|
||||
from typing import List, Dict, Any, Optional, Tuple
|
||||
import re
|
||||
|
||||
from src.config import get_settings
|
||||
|
||||
from src.core.dependencies import (
|
||||
get_wiki_service,
|
||||
get_graph_service,
|
||||
get_ingestion_service,
|
||||
verify_api_key
|
||||
)
|
||||
from src.services.wiki_service import WikiService
|
||||
from src.services.graph_service import GraphService
|
||||
from src.services.ingestion_service import IngestionService
|
||||
from src.models.wiki import WikiPageUpdate
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
router = APIRouter(prefix="/entity-linking", tags=["Entity Linking"])
|
||||
|
||||
|
||||
class EntityLinkingResult(BaseModel):
|
||||
"""Result of entity linking operation."""
|
||||
page_id: int
|
||||
page_title: str
|
||||
entities_found: List[Dict[str, Any]]
|
||||
new_graph_links_created: int # New MENTIONS relationships in Neo4j
|
||||
total_graph_links: int
|
||||
content_links_added: int # New hyperlinks added to wiki page content
|
||||
content_updated: bool # Whether wiki page content was modified
|
||||
re_indexed: bool
|
||||
processing_time_ms: float
|
||||
|
||||
|
||||
class EntityLinkingRequest(BaseModel):
|
||||
"""Request for entity linking."""
|
||||
user: str
|
||||
page_id: int
|
||||
create_relationships: bool = True
|
||||
re_index_if_changed: bool = True
|
||||
|
||||
|
||||
@router.post("/link-page", response_model=EntityLinkingResult)
|
||||
async def link_entities_in_page(
|
||||
request: EntityLinkingRequest,
|
||||
wiki_service: WikiService = Depends(get_wiki_service),
|
||||
graph_service: GraphService = Depends(get_graph_service),
|
||||
ingestion_service: IngestionService = Depends(get_ingestion_service),
|
||||
api_key: str = Depends(verify_api_key)
|
||||
) -> EntityLinkingResult:
|
||||
"""
|
||||
Find and link entities mentioned in a wiki page.
|
||||
|
||||
This endpoint:
|
||||
1. Retrieves the page content from Wiki.js
|
||||
2. Fetches all existing entities from the knowledge graph
|
||||
3. Finds mentions of those entities in the page text
|
||||
4. Creates MENTIONS relationships in Neo4j
|
||||
5. Optionally re-indexes the page if new links were created
|
||||
|
||||
Args:
|
||||
request: EntityLinkingRequest with page_id and user
|
||||
|
||||
Returns:
|
||||
EntityLinkingResult with statistics about linked entities
|
||||
"""
|
||||
import time
|
||||
start_time = time.time()
|
||||
|
||||
try:
|
||||
logger.info(f"Entity linking for page {request.page_id} (user: {request.user})")
|
||||
|
||||
# Step 1: Get page content
|
||||
page = await wiki_service.get_page(request.page_id, request.user)
|
||||
if not page:
|
||||
raise HTTPException(status_code=404, detail=f"Page {request.page_id} not found")
|
||||
|
||||
page_title = page.title
|
||||
page_content = page.content
|
||||
original_content = page_content # Keep original for comparison
|
||||
|
||||
logger.info(f"Processing page: {page_title}")
|
||||
|
||||
# Step 2: Get all entities from knowledge graph with their wiki page paths
|
||||
entities = await get_entities_with_paths(graph_service, request.user)
|
||||
logger.info(f"Found {len(entities)} existing entities in graph")
|
||||
|
||||
# Step 3: Find entity mentions in page content
|
||||
found_entities = find_entity_mentions(page_content, entities)
|
||||
logger.info(f"Found {len(found_entities)} entity mentions in page")
|
||||
|
||||
# Filter out self-referential links (entities linking to the current page)
|
||||
found_entities = [
|
||||
e for e in found_entities
|
||||
if e.get("doc_page_id") != request.page_id
|
||||
]
|
||||
logger.info(f"After filtering self-references: {len(found_entities)} entities to link")
|
||||
|
||||
# Step 4: Add hyperlinks to wiki content for entities with pages
|
||||
content_links_added = 0
|
||||
content_updated = False
|
||||
|
||||
if found_entities:
|
||||
updated_content, content_links_added = add_entity_links_to_content(
|
||||
page_content,
|
||||
found_entities
|
||||
)
|
||||
|
||||
if updated_content != original_content:
|
||||
content_updated = True
|
||||
logger.info(f"Added {content_links_added} hyperlinks to page content")
|
||||
|
||||
# Update the wiki page (preserve existing title, description, tags)
|
||||
try:
|
||||
update_data = WikiPageUpdate(
|
||||
content=updated_content,
|
||||
title=page.title,
|
||||
description=page.description if hasattr(page, 'description') else None,
|
||||
tags=page.tags if hasattr(page, 'tags') else None
|
||||
)
|
||||
await wiki_service.update_page(
|
||||
page_id=request.page_id,
|
||||
page_data=update_data,
|
||||
user=request.user
|
||||
)
|
||||
logger.info(f"Updated wiki page {request.page_id} with entity links")
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to update wiki page content: {e}")
|
||||
# Continue anyway - graph links can still be created
|
||||
|
||||
# Step 5: Create graph relationships if requested
|
||||
new_graph_links = 0
|
||||
total_graph_links = 0
|
||||
|
||||
if request.create_relationships and found_entities:
|
||||
new_graph_links = await graph_service.create_entity_mentions(
|
||||
page_id=request.page_id,
|
||||
user=request.user,
|
||||
entity_names=found_entities
|
||||
)
|
||||
total_graph_links = len(found_entities)
|
||||
logger.info(f"Created {new_graph_links} new MENTIONS relationships")
|
||||
|
||||
# Step 6: Re-index if changes were made (content or graph)
|
||||
re_indexed = False
|
||||
if request.re_index_if_changed and (content_updated or new_graph_links > 0):
|
||||
logger.info(f"Re-indexing page {request.page_id} due to entity linking changes")
|
||||
try:
|
||||
await ingestion_service.ingest_page(request.page_id, request.user)
|
||||
re_indexed = True
|
||||
except Exception as e:
|
||||
logger.error(f"Re-indexing failed: {e}")
|
||||
# Don't fail the whole operation if re-indexing fails
|
||||
|
||||
processing_time = (time.time() - start_time) * 1000
|
||||
|
||||
return EntityLinkingResult(
|
||||
page_id=request.page_id,
|
||||
page_title=page_title,
|
||||
entities_found=found_entities,
|
||||
new_graph_links_created=new_graph_links,
|
||||
total_graph_links=total_graph_links,
|
||||
content_links_added=content_links_added,
|
||||
content_updated=content_updated,
|
||||
re_indexed=re_indexed,
|
||||
processing_time_ms=processing_time
|
||||
)
|
||||
|
||||
except HTTPException:
|
||||
raise
|
||||
except Exception as e:
|
||||
logger.error(f"Entity linking failed: {e}", exc_info=True)
|
||||
raise HTTPException(status_code=500, detail=f"Entity linking failed: {str(e)}")
|
||||
|
||||
|
||||
def find_entity_mentions(content: str, entities: List[Dict[str, Any]]) -> List[Dict[str, Any]]:
|
||||
"""
|
||||
Find mentions of entities in the page content.
|
||||
|
||||
Uses case-insensitive regex matching to find entity names in the text.
|
||||
|
||||
Args:
|
||||
content: Page content to search
|
||||
entities: List of entities with 'name', 'type', and optionally 'path' fields
|
||||
|
||||
Returns:
|
||||
List of found entities with additional 'mentions' field
|
||||
"""
|
||||
found = []
|
||||
content_lower = content.lower()
|
||||
|
||||
for entity in entities:
|
||||
entity_name = entity.get("name", "")
|
||||
if not entity_name or len(entity_name) < 3:
|
||||
continue
|
||||
|
||||
# Create regex pattern for whole word matching
|
||||
# This avoids matching "John" in "Johnson"
|
||||
pattern = r'\b' + re.escape(entity_name.lower()) + r'\b'
|
||||
|
||||
# Find all matches
|
||||
matches = list(re.finditer(pattern, content_lower))
|
||||
|
||||
if matches:
|
||||
found.append({
|
||||
"name": entity_name,
|
||||
"type": entity.get("type", "unknown"),
|
||||
"mentions": len(matches),
|
||||
"entity_id": entity.get("id"),
|
||||
"path": entity.get("path"), # Include path if available
|
||||
"doc_page_id": entity.get("doc_page_id") # Include for self-reference filtering
|
||||
})
|
||||
|
||||
# Sort by number of mentions (descending)
|
||||
found.sort(key=lambda x: x["mentions"], reverse=True)
|
||||
|
||||
return found
|
||||
|
||||
|
||||
def fuzzy_match_entity_to_document(
|
||||
entity_name: str,
|
||||
doc_title: str,
|
||||
settings
|
||||
) -> Tuple[bool, float]:
|
||||
"""
|
||||
Match entity name to document title with tolerance for variations.
|
||||
|
||||
Uses multiple strategies with confidence scoring:
|
||||
1. Exact match (confidence: 1.0)
|
||||
2. Containment match (confidence: 0.85-0.80)
|
||||
3. Token overlap match (confidence: 0.60-0.80)
|
||||
|
||||
Args:
|
||||
entity_name: Entity name to match
|
||||
doc_title: Document title to match against
|
||||
settings: Application settings with matching thresholds
|
||||
|
||||
Returns:
|
||||
Tuple of (is_match, confidence_score)
|
||||
"""
|
||||
e_lower = entity_name.lower().strip()
|
||||
d_lower = doc_title.lower().strip()
|
||||
|
||||
# Strategy 1: Exact match (confidence: 1.0)
|
||||
if e_lower == d_lower:
|
||||
return True, 1.0
|
||||
|
||||
# Strategy 2: Containment match (confidence: 0.85-0.80)
|
||||
# Require minimum length to avoid false positives
|
||||
if len(entity_name) >= settings.entity_linking_min_entity_length:
|
||||
# Entity is substring of title: "lingecollege" in "RSG Lingecollege"
|
||||
if e_lower in d_lower:
|
||||
confidence = len(entity_name) / len(doc_title)
|
||||
if confidence >= settings.entity_linking_min_containment_ratio:
|
||||
return True, 0.85
|
||||
|
||||
# Title is substring of entity: "Google" in "Google Cloud Platform"
|
||||
if d_lower in e_lower:
|
||||
confidence = len(doc_title) / len(entity_name)
|
||||
if confidence >= settings.entity_linking_min_containment_ratio:
|
||||
return True, 0.80
|
||||
|
||||
# Strategy 3: Word-level overlap (confidence: 0.60-0.80)
|
||||
entity_tokens = set(e_lower.split())
|
||||
title_tokens = set(d_lower.split())
|
||||
|
||||
# Remove common stop words to reduce false positives
|
||||
stop_words = {'the', 'a', 'an', 'of', 'for', 'and', 'or', 'in', 'on', 'at', 'to'}
|
||||
entity_tokens -= stop_words
|
||||
title_tokens -= stop_words
|
||||
|
||||
if entity_tokens and title_tokens:
|
||||
overlap = len(entity_tokens & title_tokens)
|
||||
total = len(entity_tokens | title_tokens)
|
||||
overlap_ratio = overlap / total
|
||||
|
||||
# Require significant overlap to avoid weak matches
|
||||
if overlap_ratio >= settings.entity_linking_min_token_overlap:
|
||||
confidence = 0.75 * overlap_ratio # Scale: 0.45-0.75
|
||||
return True, confidence
|
||||
|
||||
return False, 0.0
|
||||
|
||||
|
||||
async def get_entities_with_paths(graph_service: GraphService, user: str) -> List[Dict[str, Any]]:
|
||||
"""
|
||||
Get all entities and check which ones have corresponding wiki pages.
|
||||
|
||||
Uses fuzzy matching to handle name variations, typos, and partial matches
|
||||
while minimizing false positives through confidence thresholds.
|
||||
|
||||
Returns entities with their names, types, paths, and match confidence.
|
||||
"""
|
||||
from src.core.multi_tenancy import get_neo4j_user_base_label
|
||||
|
||||
settings = get_settings()
|
||||
user_base_label = get_neo4j_user_base_label(user)
|
||||
|
||||
# Step 1: Get all entities
|
||||
entities_query = f"""
|
||||
MATCH (e:{user_base_label})
|
||||
WHERE NOT e:Document
|
||||
RETURN e.name as name,
|
||||
e.type as type,
|
||||
e.id as id
|
||||
"""
|
||||
|
||||
# Step 2: Get all documents
|
||||
documents_query = """
|
||||
MATCH (d:Document)
|
||||
RETURN d.title as title,
|
||||
d.path as path,
|
||||
d.page_id as page_id
|
||||
"""
|
||||
|
||||
try:
|
||||
entities = await graph_service.neo4j.execute_query(entities_query, {})
|
||||
documents = await graph_service.neo4j.execute_query(documents_query, {})
|
||||
|
||||
# Step 3: Fuzzy match entities to documents
|
||||
results = []
|
||||
for entity in entities:
|
||||
entity_name = entity["name"]
|
||||
if not entity_name:
|
||||
continue
|
||||
|
||||
best_match = None
|
||||
best_confidence = 0.0
|
||||
|
||||
# Try to match this entity to any document
|
||||
for doc in documents:
|
||||
doc_title = doc["title"]
|
||||
if not doc_title:
|
||||
continue
|
||||
|
||||
is_match, confidence = fuzzy_match_entity_to_document(
|
||||
entity_name,
|
||||
doc_title,
|
||||
settings
|
||||
)
|
||||
|
||||
if is_match and confidence > best_confidence:
|
||||
best_match = doc
|
||||
best_confidence = confidence
|
||||
|
||||
# Only include matches above minimum confidence threshold
|
||||
if best_match and best_confidence >= settings.entity_linking_min_confidence:
|
||||
results.append({
|
||||
"name": entity_name,
|
||||
"type": entity.get("type", "unknown"),
|
||||
"id": entity.get("id"),
|
||||
"path": best_match["path"],
|
||||
"doc_page_id": best_match["page_id"],
|
||||
"match_confidence": best_confidence # NEW: track confidence
|
||||
})
|
||||
else:
|
||||
# No matching document found (or below threshold)
|
||||
results.append({
|
||||
"name": entity_name,
|
||||
"type": entity.get("type", "unknown"),
|
||||
"id": entity.get("id"),
|
||||
"path": None,
|
||||
"doc_page_id": None,
|
||||
"match_confidence": 0.0
|
||||
})
|
||||
|
||||
logger.info(f"Matched {len([r for r in results if r['path']])} entities to documents (threshold: {settings.entity_linking_min_confidence})")
|
||||
return results
|
||||
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to get entities with paths: {e}")
|
||||
return []
|
||||
|
||||
|
||||
def add_entity_links_to_content(
|
||||
content: str,
|
||||
entities_with_paths: List[Dict[str, Any]]
|
||||
) -> Tuple[str, int]:
|
||||
"""
|
||||
Add markdown hyperlinks for entities in the content.
|
||||
|
||||
Only links entities that:
|
||||
1. Have a wiki page (path is not None)
|
||||
2. Are not already inside existing links
|
||||
3. Are not already linked in the content
|
||||
|
||||
Returns:
|
||||
Tuple of (updated_content, links_added_count)
|
||||
"""
|
||||
if not entities_with_paths:
|
||||
return content, 0
|
||||
|
||||
# Filter to only entities with paths
|
||||
linkable_entities = [e for e in entities_with_paths if e.get("path")]
|
||||
if not linkable_entities:
|
||||
return content, 0
|
||||
|
||||
# Sort by length (longest first) to avoid partial replacements
|
||||
# e.g., "Machine Learning" before "Machine"
|
||||
linkable_entities.sort(key=lambda x: len(x["name"]), reverse=True)
|
||||
|
||||
updated_content = content
|
||||
links_added = 0
|
||||
|
||||
for entity in linkable_entities:
|
||||
entity_name = entity["name"]
|
||||
entity_path = entity["path"]
|
||||
|
||||
# Skip short entity names to avoid false positives
|
||||
if len(entity_name) < 3:
|
||||
continue
|
||||
|
||||
# Create markdown link with full path (including user namespace)
|
||||
# Wiki.js expects full paths like /users/jpmschweitzer/...
|
||||
markdown_link = f"[{entity_name}](/{entity_path})"
|
||||
|
||||
# Find all existing markdown links to protect them (recompute each iteration)
|
||||
link_pattern = r'\[([^\]]+)\]\([^\)]+\)'
|
||||
existing_links = list(re.finditer(link_pattern, updated_content))
|
||||
|
||||
# Create a list of protected ranges (start, end) for ENTIRE links (text + URL)
|
||||
# This prevents linking entity names inside existing link URLs
|
||||
protected_ranges = [(m.start(), m.end()) for m in existing_links]
|
||||
|
||||
# Find all potential matches
|
||||
pattern = r'\b(' + re.escape(entity_name) + r')\b'
|
||||
matches = list(re.finditer(pattern, updated_content, flags=re.IGNORECASE))
|
||||
|
||||
# Filter out matches that are inside existing links
|
||||
valid_matches = []
|
||||
for match in matches:
|
||||
match_start = match.start()
|
||||
match_end = match.end()
|
||||
|
||||
# Check if this match is inside any protected range
|
||||
inside_link = False
|
||||
for prot_start, prot_end in protected_ranges:
|
||||
if prot_start <= match_start < prot_end or prot_start < match_end <= prot_end:
|
||||
inside_link = True
|
||||
break
|
||||
|
||||
# Also check if already a link (pattern like [entity_name](...))
|
||||
if match_end < len(updated_content) - 1:
|
||||
next_chars = updated_content[match_end:match_end+2]
|
||||
if next_chars == '](':
|
||||
inside_link = True
|
||||
|
||||
if not inside_link:
|
||||
valid_matches.append(match)
|
||||
|
||||
if not valid_matches:
|
||||
continue
|
||||
|
||||
# Replace valid matches in reverse order (to preserve positions)
|
||||
for match in reversed(valid_matches):
|
||||
updated_content = (
|
||||
updated_content[:match.start()] +
|
||||
markdown_link +
|
||||
updated_content[match.end():]
|
||||
)
|
||||
links_added += 1
|
||||
|
||||
return updated_content, links_added
|
||||
@@ -1,253 +0,0 @@
|
||||
"""
|
||||
Graph router for Library Desk API.
|
||||
|
||||
Endpoints for Neo4j knowledge graph operations.
|
||||
"""
|
||||
|
||||
from fastapi import APIRouter, HTTPException, Depends, Query
|
||||
from typing import Optional, List
|
||||
import logging
|
||||
|
||||
from src.models.graph import (
|
||||
CypherQueryRequest, CypherQueryResponse,
|
||||
UpdateFromPageRequest, GraphUpdateSummary,
|
||||
NodeListResponse, GraphNodeDetail,
|
||||
MindMapResponse
|
||||
)
|
||||
from src.services.graph_service import GraphService
|
||||
from src.clients.neo4j_client import Neo4jClient
|
||||
from src.clients.wikijs_client import WikiJSClient
|
||||
from src.core.dependencies import Neo4jDep, WikiJSDep, verify_api_key
|
||||
from src.core.multi_tenancy import DEFAULT_USER
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
router = APIRouter(prefix="/graph", tags=["Graph"])
|
||||
|
||||
|
||||
# Dependency to get graph service
|
||||
def get_graph_service(
|
||||
neo4j_client: Neo4jDep,
|
||||
wiki_client: WikiJSDep
|
||||
) -> GraphService:
|
||||
"""Get graph service instance."""
|
||||
return GraphService(neo4j_client, wiki_client)
|
||||
|
||||
|
||||
@router.post("/query", response_model=CypherQueryResponse)
|
||||
async def execute_cypher_query(
|
||||
request: CypherQueryRequest,
|
||||
graph_service: GraphService = Depends(get_graph_service),
|
||||
api_key: str = Depends(verify_api_key)
|
||||
):
|
||||
"""
|
||||
Execute a user-scoped Cypher query.
|
||||
|
||||
The query is automatically scoped to the user's data for security.
|
||||
This prevents users from accessing other users' graph data.
|
||||
|
||||
**Example Request:**
|
||||
```json
|
||||
{
|
||||
"query": "MATCH (d:Document) RETURN d LIMIT 10",
|
||||
"parameters": {},
|
||||
"user": "jpmschweitzer"
|
||||
}
|
||||
```
|
||||
|
||||
**Security:** Query is automatically scoped with user label.
|
||||
"""
|
||||
try:
|
||||
return await graph_service.execute_query(
|
||||
query=request.query,
|
||||
parameters=request.parameters,
|
||||
user=request.user
|
||||
)
|
||||
except ValueError as e:
|
||||
raise HTTPException(status_code=400, detail=str(e))
|
||||
except Exception as e:
|
||||
logger.error(f"Cypher query failed: {e}", exc_info=True)
|
||||
raise HTTPException(status_code=500, detail="Query execution failed")
|
||||
|
||||
|
||||
@router.get("/nodes", response_model=NodeListResponse)
|
||||
async def list_nodes(
|
||||
user: str = Query(default=DEFAULT_USER, description="User identifier"),
|
||||
node_type: Optional[str] = Query(default=None, description="Node type filter"),
|
||||
limit: int = Query(default=100, ge=1, le=500, description="Maximum nodes"),
|
||||
graph_service: GraphService = Depends(get_graph_service),
|
||||
api_key: str = Depends(verify_api_key)
|
||||
):
|
||||
"""
|
||||
List graph nodes for a user.
|
||||
|
||||
Optionally filter by node type (Document, Person, Project, Concept, etc.).
|
||||
|
||||
**Example:** `/graph/nodes?user=jpmschweitzer&node_type=Document&limit=50`
|
||||
"""
|
||||
try:
|
||||
return await graph_service.list_nodes(
|
||||
user=user,
|
||||
node_type=node_type,
|
||||
limit=limit
|
||||
)
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to list nodes: {e}", exc_info=True)
|
||||
raise HTTPException(status_code=500, detail="Failed to list nodes")
|
||||
|
||||
|
||||
@router.get("/nodes/{node_id}", response_model=GraphNodeDetail)
|
||||
async def get_node(
|
||||
node_id: str,
|
||||
user: str = Query(default=DEFAULT_USER, description="User identifier"),
|
||||
graph_service: GraphService = Depends(get_graph_service),
|
||||
api_key: str = Depends(verify_api_key)
|
||||
):
|
||||
"""
|
||||
Get detailed information about a graph node.
|
||||
|
||||
Returns the node, its relationships, and connected nodes.
|
||||
|
||||
**Example:** `/graph/nodes/4:abc123def:0?user=jpmschweitzer`
|
||||
"""
|
||||
try:
|
||||
node = await graph_service.get_node(node_id, user)
|
||||
if not node:
|
||||
raise HTTPException(status_code=404, detail=f"Node {node_id} not found")
|
||||
return node
|
||||
except HTTPException:
|
||||
raise
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to get node {node_id}: {e}", exc_info=True)
|
||||
raise HTTPException(status_code=500, detail="Failed to get node")
|
||||
|
||||
|
||||
@router.post("/update-from-page/{page_id}", response_model=GraphUpdateSummary)
|
||||
async def update_graph_from_page(
|
||||
page_id: int,
|
||||
user: str = Query(default=DEFAULT_USER, description="User identifier"),
|
||||
force_refresh: bool = Query(default=False, description="Force re-extraction"),
|
||||
graph_service: GraphService = Depends(get_graph_service),
|
||||
api_key: str = Depends(verify_api_key)
|
||||
):
|
||||
"""
|
||||
Update knowledge graph from a wiki page.
|
||||
|
||||
This endpoint:
|
||||
1. Fetches the page from Wiki.js
|
||||
2. Extracts entities (people, projects, concepts, technologies)
|
||||
3. Creates/updates Document node
|
||||
4. Creates entity nodes and MENTIONS relationships
|
||||
5. Returns summary of what was updated
|
||||
|
||||
**Use Cases:**
|
||||
- Called automatically after page creation/update (via BackgroundTasks)
|
||||
- Called manually by user/Librarian to refresh graph
|
||||
- Called by Scheduler for batch processing
|
||||
|
||||
**Example:** `POST /graph/update-from-page/4?user=jpmschweitzer`
|
||||
|
||||
**Returns:** Summary with nodes/relationships created and entities extracted
|
||||
"""
|
||||
try:
|
||||
summary = await graph_service.update_from_page(
|
||||
page_id=page_id,
|
||||
user=user,
|
||||
force_refresh=force_refresh
|
||||
)
|
||||
|
||||
if not summary.success:
|
||||
raise HTTPException(
|
||||
status_code=500,
|
||||
detail=f"Graph update failed: {summary.error_message}"
|
||||
)
|
||||
|
||||
return summary
|
||||
|
||||
except HTTPException:
|
||||
raise
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to update graph from page {page_id}: {e}", exc_info=True)
|
||||
raise HTTPException(status_code=500, detail="Graph update failed")
|
||||
|
||||
|
||||
@router.post("/mindmap", response_model=MindMapResponse)
|
||||
async def generate_mindmap(
|
||||
center_node_id: str = Query(..., description="Central node ID"),
|
||||
user: str = Query(default=DEFAULT_USER, description="User identifier"),
|
||||
depth: int = Query(default=2, ge=1, le=5, description="Traversal depth"),
|
||||
graph_service: GraphService = Depends(get_graph_service),
|
||||
api_key: str = Depends(verify_api_key)
|
||||
):
|
||||
"""
|
||||
Generate mind map data for visualization.
|
||||
|
||||
Traverses the graph from a center node and returns nodes/links
|
||||
in a format suitable for D3.js or similar visualization libraries.
|
||||
|
||||
**Parameters:**
|
||||
- `center_node_id`: The node to center the mind map on
|
||||
- `depth`: How many hops away from center to include (1-5)
|
||||
- `user`: User identifier for scoping
|
||||
|
||||
**Example:** `POST /graph/mindmap?center_node_id=4:abc:0&depth=2`
|
||||
|
||||
**Returns:** Nodes and links for visualization
|
||||
"""
|
||||
try:
|
||||
return await graph_service.generate_mindmap(
|
||||
center_node_id=center_node_id,
|
||||
user=user,
|
||||
depth=depth
|
||||
)
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to generate mindmap: {e}", exc_info=True)
|
||||
raise HTTPException(status_code=500, detail="Mindmap generation failed")
|
||||
|
||||
|
||||
@router.post("/generate-entity-pages")
|
||||
async def generate_entity_pages(
|
||||
user: str = Query(default=DEFAULT_USER, description="User identifier"),
|
||||
min_mentions: int = Query(default=5, ge=1, le=100, description="Minimum mentions threshold"),
|
||||
entity_types: Optional[List[str]] = Query(default=None, description="Entity types to process"),
|
||||
graph_service: GraphService = Depends(get_graph_service),
|
||||
api_key: str = Depends(verify_api_key)
|
||||
):
|
||||
"""
|
||||
Generate wiki stub pages for graph entities.
|
||||
|
||||
Creates pages in `/entities/{type}/{name}` namespace for entities
|
||||
that have been mentioned in multiple documents. This creates a
|
||||
bidirectional knowledge graph ↔ wiki synchronization.
|
||||
|
||||
**Threshold:** Entities must be mentioned in at least `min_mentions` documents (default: 5)
|
||||
|
||||
**Auto-stub flow:**
|
||||
1. Find entities with >= min_mentions
|
||||
2. Check if entity already has a wiki page
|
||||
3. If not, create stub page with:
|
||||
- List of mentioning documents
|
||||
- Related entities (co-occurring)
|
||||
- Auto-generated tag to prevent feedback loops
|
||||
|
||||
**Feedback loop protection:** Pages tagged with `entity-stub` skip entity extraction
|
||||
|
||||
**Parameters:**
|
||||
- `min_mentions`: Minimum number of document mentions required (default: 5)
|
||||
- `entity_types`: List of types to process (default: Person, Technology, Concept, Project)
|
||||
- `user`: User identifier for scoping
|
||||
|
||||
**Example:** `POST /graph/generate-entity-pages?min_mentions=3&user=jpmschweitzer`
|
||||
|
||||
**Returns:** Summary with list of pages created and skipped
|
||||
"""
|
||||
try:
|
||||
result = await graph_service.generate_entity_stubs(
|
||||
user=user,
|
||||
min_mentions=min_mentions,
|
||||
entity_types=entity_types
|
||||
)
|
||||
return result
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to generate entity pages: {e}", exc_info=True)
|
||||
raise HTTPException(status_code=500, detail="Entity page generation failed")
|
||||
@@ -1,117 +0,0 @@
|
||||
"""
|
||||
HybridRAG router for multi-source search API.
|
||||
|
||||
Provides endpoint for combining vector, graph, and web search
|
||||
with RRF fusion and LLM re-ranking.
|
||||
"""
|
||||
|
||||
from fastapi import APIRouter, HTTPException, Depends, Query
|
||||
import logging
|
||||
|
||||
from src.models.hybrid_rag import HybridRAGRequest, HybridRAGResponse
|
||||
from src.services.hybrid_rag_service import HybridRAGService
|
||||
from src.services.vector_service import VectorService
|
||||
from src.services.graph_service import GraphService
|
||||
from src.clients.searxng_client import SearXNGClient
|
||||
from src.clients.ollama_client import OllamaClient
|
||||
from src.core.dependencies import (
|
||||
Neo4jDep, WikiJSDep, QdrantDep, OllamaDep,
|
||||
SearXNGDep, verify_api_key, get_settings
|
||||
)
|
||||
from src.config import Settings
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
router = APIRouter(prefix="/query", tags=["HybridRAG"])
|
||||
|
||||
|
||||
# Dependency to get HybridRAG service
|
||||
def get_hybrid_rag_service(
|
||||
neo4j_client: Neo4jDep,
|
||||
wiki_client: WikiJSDep,
|
||||
qdrant_client: QdrantDep,
|
||||
ollama_client: OllamaDep,
|
||||
searxng_client: SearXNGDep,
|
||||
settings: Settings = Depends(get_settings)
|
||||
) -> HybridRAGService:
|
||||
"""Get HybridRAG service instance with all dependencies."""
|
||||
from src.services.vector_service import VectorService
|
||||
from src.services.graph_service import GraphService
|
||||
|
||||
# Create component services
|
||||
vector_service = VectorService(qdrant_client, wiki_client, ollama_client)
|
||||
graph_service = GraphService(neo4j_client, wiki_client)
|
||||
|
||||
# Create HybridRAG service
|
||||
return HybridRAGService(
|
||||
vector_service=vector_service,
|
||||
graph_service=graph_service,
|
||||
searxng_client=searxng_client,
|
||||
ollama_client=ollama_client,
|
||||
settings=settings
|
||||
)
|
||||
|
||||
|
||||
@router.post("/hybrid", response_model=HybridRAGResponse)
|
||||
async def hybrid_search(
|
||||
request: HybridRAGRequest,
|
||||
user: str = Query(default="jpmschweitzer", description="User identifier for multi-tenancy"),
|
||||
hybrid_rag_service: HybridRAGService = Depends(get_hybrid_rag_service),
|
||||
api_key: str = Depends(verify_api_key)
|
||||
):
|
||||
"""
|
||||
Execute HybridRAG query combining vector, graph, and web search.
|
||||
|
||||
**6-Phase Pipeline:**
|
||||
1. **Query Enhancement**: Extract keywords/synonyms with LLM
|
||||
2. **Parallel Retrieval**: Search vector (Qdrant), graph (Neo4j), web (SearXNG)
|
||||
3. **RRF Fusion**: Merge results with Reciprocal Rank Fusion
|
||||
4. **Enrichment**: Add related documents via shared entities
|
||||
5. **LLM Re-ranking**: Re-rank with mistral-nemo for relevance
|
||||
6. **Context Formatting**: Format for LLM consumption
|
||||
7. **Persistence**: Store for Librarian knowledge consolidation
|
||||
|
||||
**Example Request:**
|
||||
```json
|
||||
{
|
||||
"query": "How does Docker orchestration work with Kubernetes?",
|
||||
"user": "jpmschweitzer",
|
||||
"config": {
|
||||
"vector_limit": 10,
|
||||
"graph_limit": 10,
|
||||
"web_limit": 5,
|
||||
"enable_reranking": true,
|
||||
"final_result_count": 10
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**Returns:**
|
||||
- Ranked results from all sources
|
||||
- Extracted keywords/synonyms
|
||||
- Related dossiers (via graph)
|
||||
- Formatted context for LLM
|
||||
- Performance timing breakdown
|
||||
- Search ID for Librarian tracking
|
||||
"""
|
||||
try:
|
||||
logger.info(f"HybridRAG request: '{request.query}' for user '{user}'")
|
||||
|
||||
response = await hybrid_rag_service.search(
|
||||
query=request.query,
|
||||
user=user,
|
||||
config=request.config
|
||||
)
|
||||
|
||||
logger.info(
|
||||
f"HybridRAG completed: {response.total_results} results in {response.timing.total_ms:.0f}ms"
|
||||
)
|
||||
|
||||
return response
|
||||
|
||||
except ValueError as e:
|
||||
logger.error(f"Invalid request: {e}")
|
||||
raise HTTPException(status_code=400, detail=str(e))
|
||||
except Exception as e:
|
||||
logger.error(f"HybridRAG search failed: {e}", exc_info=True)
|
||||
raise HTTPException(status_code=500, detail="Search failed")
|
||||
@@ -1,169 +0,0 @@
|
||||
"""
|
||||
Document Ingestion API Router
|
||||
|
||||
Endpoints for ingesting wiki pages into the knowledge base (vectors + graph).
|
||||
"""
|
||||
from fastapi import APIRouter, Depends, HTTPException, Query
|
||||
from typing import Optional
|
||||
|
||||
from src.services.ingestion_service import IngestionService
|
||||
from src.models.ingestion import (
|
||||
IngestionRequest,
|
||||
IngestionResult,
|
||||
BatchIngestionRequest,
|
||||
BatchIngestionResult
|
||||
)
|
||||
from src.core.dependencies import get_ingestion_service, verify_api_key
|
||||
|
||||
router = APIRouter(prefix="/ingest", tags=["Document Ingestion"])
|
||||
|
||||
|
||||
@router.post("/page", response_model=IngestionResult)
|
||||
async def ingest_page(
|
||||
request: IngestionRequest,
|
||||
ingestion: IngestionService = Depends(get_ingestion_service),
|
||||
api_key: str = Depends(verify_api_key)
|
||||
):
|
||||
"""
|
||||
Ingest a single wiki page into the knowledge base.
|
||||
|
||||
This endpoint:
|
||||
1. Fetches page content from Wiki.js
|
||||
2. Chunks content and generates embeddings (Qdrant)
|
||||
3. Extracts entities and updates knowledge graph (Neo4j)
|
||||
|
||||
## Use Cases
|
||||
|
||||
- **After page creation**: Automatically called by consolidation service
|
||||
- **Manual re-indexing**: Force refresh a page after manual edits
|
||||
- **Selective ingestion**: Skip vectors or graph if only one is needed
|
||||
|
||||
## Performance
|
||||
|
||||
- Typical page: 1-3 seconds
|
||||
- Large page (>5000 words): 5-10 seconds
|
||||
- Vector and graph ingestion run in parallel
|
||||
|
||||
## Example
|
||||
|
||||
```bash
|
||||
curl -X POST http://192.168.86.149:8089/ingest/page \
|
||||
-H "Authorization: Bearer $API_KEY" \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"page_id": 19,
|
||||
"user": "jpmschweitzer",
|
||||
"force_refresh": false
|
||||
}'
|
||||
```
|
||||
"""
|
||||
result = await ingestion.ingest_page(
|
||||
page_id=request.page_id,
|
||||
user=request.user,
|
||||
force_refresh=request.force_refresh,
|
||||
skip_vectors=request.skip_vectors,
|
||||
skip_graph=request.skip_graph
|
||||
)
|
||||
|
||||
if not result.success:
|
||||
raise HTTPException(
|
||||
status_code=500,
|
||||
detail=f"Ingestion failed: {result.error}"
|
||||
)
|
||||
|
||||
return result
|
||||
|
||||
|
||||
@router.post("/batch", response_model=BatchIngestionResult)
|
||||
async def ingest_batch(
|
||||
request: BatchIngestionRequest,
|
||||
ingestion: IngestionService = Depends(get_ingestion_service),
|
||||
api_key: str = Depends(verify_api_key)
|
||||
):
|
||||
"""
|
||||
Ingest multiple wiki pages concurrently.
|
||||
|
||||
## Concurrency Control
|
||||
|
||||
The `max_concurrent` parameter controls how many pages are processed simultaneously:
|
||||
- **Low (1-2)**: Safer for resource-constrained systems
|
||||
- **Medium (3-5)**: Good balance of speed and stability
|
||||
- **High (6-10)**: Maximum speed, requires good resources
|
||||
|
||||
## Batch Size Recommendations
|
||||
|
||||
- **Small batches (<10 pages)**: Use max_concurrent=3-5
|
||||
- **Medium batches (10-50 pages)**: Use max_concurrent=3
|
||||
- **Large batches (>50 pages)**: Use max_concurrent=2, consider splitting
|
||||
|
||||
## Example
|
||||
|
||||
```bash
|
||||
curl -X POST http://192.168.86.149:8089/ingest/batch \
|
||||
-H "Authorization: Bearer $API_KEY" \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"page_ids": [19, 20, 21, 22],
|
||||
"user": "jpmschweitzer",
|
||||
"max_concurrent": 3
|
||||
}'
|
||||
```
|
||||
"""
|
||||
result = await ingestion.ingest_batch(
|
||||
page_ids=request.page_ids,
|
||||
user=request.user,
|
||||
force_refresh=request.force_refresh,
|
||||
skip_vectors=request.skip_vectors,
|
||||
skip_graph=request.skip_graph,
|
||||
max_concurrent=request.max_concurrent
|
||||
)
|
||||
|
||||
return result
|
||||
|
||||
|
||||
@router.post("/all", response_model=BatchIngestionResult)
|
||||
async def ingest_all_pages(
|
||||
user: str = Query(default="jpmschweitzer", description="User identifier"),
|
||||
path_prefix: Optional[str] = Query(None, description="Path prefix filter (e.g., 'users/jpmschweitzer/tech')"),
|
||||
force_refresh: bool = Query(False, description="Force re-ingestion of all pages"),
|
||||
max_concurrent: int = Query(3, ge=1, le=10, description="Maximum concurrent ingestion tasks"),
|
||||
ingestion: IngestionService = Depends(get_ingestion_service),
|
||||
api_key: str = Depends(verify_api_key)
|
||||
):
|
||||
"""
|
||||
Ingest all wiki pages for a user (bulk re-indexing).
|
||||
|
||||
## Use Cases
|
||||
|
||||
- **Initial setup**: Index all existing pages
|
||||
- **Full re-index**: After major schema changes
|
||||
- **Path-specific**: Re-index a specific section (e.g., all tech docs)
|
||||
|
||||
## Performance
|
||||
|
||||
- **Small wiki (<50 pages)**: 2-5 minutes
|
||||
- **Medium wiki (50-200 pages)**: 5-20 minutes
|
||||
- **Large wiki (>200 pages)**: 20+ minutes
|
||||
|
||||
**Recommendation**: Run as background job for large wikis
|
||||
|
||||
## Example
|
||||
|
||||
```bash
|
||||
# Ingest all pages for user
|
||||
curl -X POST "http://192.168.86.149:8089/ingest/all?user=jpmschweitzer" \
|
||||
-H "Authorization: Bearer $API_KEY"
|
||||
|
||||
# Ingest only tech docs
|
||||
curl -X POST "http://192.168.86.149:8089/ingest/all?user=jpmschweitzer&path_prefix=users/jpmschweitzer/tech" \
|
||||
-H "Authorization: Bearer $API_KEY"
|
||||
```
|
||||
"""
|
||||
result = await ingestion.ingest_all_pages(
|
||||
user=user,
|
||||
path_prefix=path_prefix,
|
||||
force_refresh=force_refresh,
|
||||
max_concurrent=max_concurrent
|
||||
)
|
||||
|
||||
return result
|
||||
@@ -1,334 +0,0 @@
|
||||
"""
|
||||
Tools router for Library Desk API.
|
||||
|
||||
Provides tool discovery endpoint for AI agents.
|
||||
"""
|
||||
|
||||
from fastapi import APIRouter
|
||||
from src.models.tools import (
|
||||
ToolCatalog, ToolDefinition, ToolParameter, CategoryInfo, ParameterType
|
||||
)
|
||||
|
||||
router = APIRouter(prefix="/tools", tags=["Tools"])
|
||||
|
||||
|
||||
def get_wiki_tools() -> list[ToolDefinition]:
|
||||
"""Get wiki tool definitions."""
|
||||
return [
|
||||
ToolDefinition(
|
||||
name="wiki_list_pages",
|
||||
category="wiki",
|
||||
description="List wiki pages for a user with optional tag filtering",
|
||||
method="GET",
|
||||
endpoint="/wiki/pages",
|
||||
parameters=[
|
||||
ToolParameter(
|
||||
name="user",
|
||||
type=ParameterType.STRING,
|
||||
description="User identifier",
|
||||
required=False,
|
||||
default="jpmschweitzer",
|
||||
example="jpmschweitzer"
|
||||
),
|
||||
ToolParameter(
|
||||
name="tag",
|
||||
type=ParameterType.STRING,
|
||||
description="Filter by tag (dossier)",
|
||||
required=False,
|
||||
example="projects"
|
||||
),
|
||||
ToolParameter(
|
||||
name="limit",
|
||||
type=ParameterType.INTEGER,
|
||||
description="Maximum pages to return",
|
||||
required=False,
|
||||
default=50,
|
||||
example=20
|
||||
),
|
||||
],
|
||||
returns="List of pages with summaries",
|
||||
fast=True
|
||||
),
|
||||
ToolDefinition(
|
||||
name="wiki_get_page",
|
||||
category="wiki",
|
||||
description="Get a single wiki page by ID with full content",
|
||||
method="GET",
|
||||
endpoint="/wiki/pages/{page_id}",
|
||||
parameters=[
|
||||
ToolParameter(
|
||||
name="page_id",
|
||||
type=ParameterType.INTEGER,
|
||||
description="Page ID to retrieve",
|
||||
required=True,
|
||||
example=123
|
||||
),
|
||||
ToolParameter(
|
||||
name="user",
|
||||
type=ParameterType.STRING,
|
||||
description="User identifier for access control",
|
||||
required=False,
|
||||
default="jpmschweitzer"
|
||||
),
|
||||
],
|
||||
returns="Complete page object with content",
|
||||
fast=True
|
||||
),
|
||||
ToolDefinition(
|
||||
name="wiki_create_page",
|
||||
category="wiki",
|
||||
description="Create a new wiki page in user's namespace",
|
||||
method="POST",
|
||||
endpoint="/wiki/pages",
|
||||
parameters=[
|
||||
ToolParameter(
|
||||
name="title",
|
||||
type=ParameterType.STRING,
|
||||
description="Page title",
|
||||
required=True,
|
||||
example="Project Documentation"
|
||||
),
|
||||
ToolParameter(
|
||||
name="path",
|
||||
type=ParameterType.STRING,
|
||||
description="Page path (will be prefixed with user namespace)",
|
||||
required=True,
|
||||
example="/projects/my-project"
|
||||
),
|
||||
ToolParameter(
|
||||
name="content",
|
||||
type=ParameterType.STRING,
|
||||
description="Page content in Markdown format",
|
||||
required=True,
|
||||
example="# Overview\n\nThis is the content."
|
||||
),
|
||||
ToolParameter(
|
||||
name="description",
|
||||
type=ParameterType.STRING,
|
||||
description="Short page description",
|
||||
required=False,
|
||||
example="Documentation for my project"
|
||||
),
|
||||
ToolParameter(
|
||||
name="tags",
|
||||
type=ParameterType.ARRAY,
|
||||
description="Tags for categorization (dossiers)",
|
||||
required=False,
|
||||
example=["projects", "documentation"]
|
||||
),
|
||||
ToolParameter(
|
||||
name="user",
|
||||
type=ParameterType.STRING,
|
||||
description="User identifier",
|
||||
required=False,
|
||||
default="jpmschweitzer"
|
||||
),
|
||||
],
|
||||
returns="Created page object",
|
||||
example={
|
||||
"title": "My Project",
|
||||
"path": "/projects/my-project",
|
||||
"content": "# My Project\n\nProject description here.",
|
||||
"tags": ["projects"],
|
||||
"user": "jpmschweitzer"
|
||||
},
|
||||
fast=True
|
||||
),
|
||||
ToolDefinition(
|
||||
name="wiki_update_page",
|
||||
category="wiki",
|
||||
description="Update an existing wiki page",
|
||||
method="PUT",
|
||||
endpoint="/wiki/pages/{page_id}",
|
||||
parameters=[
|
||||
ToolParameter(
|
||||
name="page_id",
|
||||
type=ParameterType.INTEGER,
|
||||
description="Page ID to update",
|
||||
required=True,
|
||||
example=123
|
||||
),
|
||||
ToolParameter(
|
||||
name="title",
|
||||
type=ParameterType.STRING,
|
||||
description="New page title",
|
||||
required=False
|
||||
),
|
||||
ToolParameter(
|
||||
name="content",
|
||||
type=ParameterType.STRING,
|
||||
description="New page content",
|
||||
required=False
|
||||
),
|
||||
ToolParameter(
|
||||
name="description",
|
||||
type=ParameterType.STRING,
|
||||
description="New description",
|
||||
required=False
|
||||
),
|
||||
ToolParameter(
|
||||
name="tags",
|
||||
type=ParameterType.ARRAY,
|
||||
description="New tags",
|
||||
required=False
|
||||
),
|
||||
ToolParameter(
|
||||
name="user",
|
||||
type=ParameterType.STRING,
|
||||
description="User identifier",
|
||||
required=False,
|
||||
default="jpmschweitzer"
|
||||
),
|
||||
],
|
||||
returns="Updated page object",
|
||||
fast=True
|
||||
),
|
||||
ToolDefinition(
|
||||
name="wiki_delete_page",
|
||||
category="wiki",
|
||||
description="Delete a wiki page",
|
||||
method="DELETE",
|
||||
endpoint="/wiki/pages/{page_id}",
|
||||
parameters=[
|
||||
ToolParameter(
|
||||
name="page_id",
|
||||
type=ParameterType.INTEGER,
|
||||
description="Page ID to delete",
|
||||
required=True,
|
||||
example=123
|
||||
),
|
||||
ToolParameter(
|
||||
name="user",
|
||||
type=ParameterType.STRING,
|
||||
description="User identifier",
|
||||
required=False,
|
||||
default="jpmschweitzer"
|
||||
),
|
||||
],
|
||||
returns="Success confirmation",
|
||||
fast=True
|
||||
),
|
||||
ToolDefinition(
|
||||
name="wiki_search_pages",
|
||||
category="wiki",
|
||||
description="Search wiki pages by content within user's namespace",
|
||||
method="GET",
|
||||
endpoint="/wiki/search",
|
||||
parameters=[
|
||||
ToolParameter(
|
||||
name="q",
|
||||
type=ParameterType.STRING,
|
||||
description="Search query",
|
||||
required=True,
|
||||
example="docker configuration"
|
||||
),
|
||||
ToolParameter(
|
||||
name="user",
|
||||
type=ParameterType.STRING,
|
||||
description="User identifier",
|
||||
required=False,
|
||||
default="jpmschweitzer"
|
||||
),
|
||||
ToolParameter(
|
||||
name="limit",
|
||||
type=ParameterType.INTEGER,
|
||||
description="Maximum results",
|
||||
required=False,
|
||||
default=20
|
||||
),
|
||||
],
|
||||
returns="List of matching pages",
|
||||
fast=True
|
||||
),
|
||||
ToolDefinition(
|
||||
name="wiki_list_dossiers",
|
||||
category="wiki",
|
||||
description="List all dossiers (unique tags) for a user with page counts",
|
||||
method="GET",
|
||||
endpoint="/wiki/dossiers",
|
||||
parameters=[
|
||||
ToolParameter(
|
||||
name="user",
|
||||
type=ParameterType.STRING,
|
||||
description="User identifier",
|
||||
required=False,
|
||||
default="jpmschweitzer"
|
||||
),
|
||||
],
|
||||
returns="List of dossiers with page counts",
|
||||
fast=True
|
||||
),
|
||||
ToolDefinition(
|
||||
name="wiki_get_dossier_pages",
|
||||
category="wiki",
|
||||
description="Get all pages in a specific dossier",
|
||||
method="GET",
|
||||
endpoint="/wiki/dossiers/{dossier_name}/pages",
|
||||
parameters=[
|
||||
ToolParameter(
|
||||
name="dossier_name",
|
||||
type=ParameterType.STRING,
|
||||
description="Dossier name (tag)",
|
||||
required=True,
|
||||
example="projects"
|
||||
),
|
||||
ToolParameter(
|
||||
name="user",
|
||||
type=ParameterType.STRING,
|
||||
description="User identifier",
|
||||
required=False,
|
||||
default="jpmschweitzer"
|
||||
),
|
||||
ToolParameter(
|
||||
name="limit",
|
||||
type=ParameterType.INTEGER,
|
||||
description="Maximum pages",
|
||||
required=False,
|
||||
default=100
|
||||
),
|
||||
],
|
||||
returns="List of pages in dossier",
|
||||
fast=True
|
||||
),
|
||||
]
|
||||
|
||||
|
||||
@router.get("", response_model=ToolCatalog)
|
||||
async def get_tool_catalog() -> ToolCatalog:
|
||||
"""
|
||||
Get simplified tool catalog for AI agent consumption.
|
||||
|
||||
This endpoint provides a machine-readable catalog of all Library Desk tools,
|
||||
optimized for discovery and use by AI agents like The Librarian.
|
||||
|
||||
Returns tool definitions with:
|
||||
- Clear descriptions
|
||||
- Parameter specifications
|
||||
- Usage examples
|
||||
- Performance characteristics
|
||||
"""
|
||||
wiki_tools = get_wiki_tools()
|
||||
|
||||
# Calculate category stats
|
||||
categories = {}
|
||||
for tool in wiki_tools:
|
||||
if tool.category not in categories:
|
||||
categories[tool.category] = 0
|
||||
categories[tool.category] += 1
|
||||
|
||||
category_info = [
|
||||
CategoryInfo(
|
||||
name="wiki",
|
||||
description="Wiki.js page and dossier management operations",
|
||||
tool_count=categories.get("wiki", 0)
|
||||
)
|
||||
]
|
||||
|
||||
return ToolCatalog(
|
||||
service="library-desk",
|
||||
version="1.0.0",
|
||||
base_url="http://library-desk:8089",
|
||||
categories=category_info,
|
||||
tools=wiki_tools,
|
||||
authentication="Bearer token via Authorization header (LIBRARY_API_KEY)"
|
||||
)
|
||||
@@ -1,173 +0,0 @@
|
||||
"""
|
||||
Vector router for Library Desk API.
|
||||
|
||||
Endpoints for semantic search and vector operations.
|
||||
"""
|
||||
|
||||
from fastapi import APIRouter, HTTPException, Depends, Query
|
||||
from typing import Optional
|
||||
import logging
|
||||
|
||||
from src.models.vector import (
|
||||
SearchRequest, SearchResponse,
|
||||
VectorUpdateRequest, VectorUpdateSummary,
|
||||
CollectionListResponse,
|
||||
DeletePageChunksRequest, DeletePageChunksResponse
|
||||
)
|
||||
from src.services.vector_service import VectorService
|
||||
from src.clients.qdrant_client import QdrantClientWrapper
|
||||
from src.clients.wikijs_client import WikiJSClient
|
||||
from src.clients.ollama_client import OllamaClient
|
||||
from src.core.dependencies import QdrantDep, WikiJSDep, OllamaDep, verify_api_key
|
||||
from src.core.multi_tenancy import DEFAULT_USER
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
router = APIRouter(prefix="/vector", tags=["Vector"])
|
||||
|
||||
|
||||
# Dependency to get vector service
|
||||
def get_vector_service(
|
||||
qdrant_client: QdrantDep,
|
||||
wiki_client: WikiJSDep,
|
||||
ollama_client: OllamaDep
|
||||
) -> VectorService:
|
||||
"""Get vector service instance."""
|
||||
return VectorService(qdrant_client, wiki_client, ollama_client)
|
||||
|
||||
|
||||
@router.post("/search", response_model=SearchResponse)
|
||||
async def semantic_search(
|
||||
request: SearchRequest,
|
||||
vector_service: VectorService = Depends(get_vector_service),
|
||||
api_key: str = Depends(verify_api_key)
|
||||
):
|
||||
"""
|
||||
Perform semantic search across user's documents.
|
||||
|
||||
Uses Ollama to generate query embedding, then searches Qdrant
|
||||
for similar document chunks.
|
||||
|
||||
**Example Request:**
|
||||
```json
|
||||
{
|
||||
"query": "how to configure docker",
|
||||
"user": "jpmschweitzer",
|
||||
"limit": 10,
|
||||
"score_threshold": 0.5
|
||||
}
|
||||
```
|
||||
|
||||
**Returns:** List of matching chunks with similarity scores
|
||||
"""
|
||||
try:
|
||||
return await vector_service.search(
|
||||
query=request.query,
|
||||
user=request.user,
|
||||
limit=request.limit,
|
||||
score_threshold=request.score_threshold
|
||||
)
|
||||
except ValueError as e:
|
||||
raise HTTPException(status_code=400, detail=str(e))
|
||||
except Exception as e:
|
||||
logger.error(f"Semantic search failed: {e}", exc_info=True)
|
||||
raise HTTPException(status_code=500, detail="Search failed")
|
||||
|
||||
|
||||
@router.post("/update-from-page/{page_id}", response_model=VectorUpdateSummary)
|
||||
async def update_vectors_from_page(
|
||||
page_id: int,
|
||||
user: str = Query(default=DEFAULT_USER, description="User identifier"),
|
||||
force_refresh: bool = Query(default=False, description="Force re-embedding"),
|
||||
vector_service: VectorService = Depends(get_vector_service),
|
||||
api_key: str = Depends(verify_api_key)
|
||||
):
|
||||
"""
|
||||
Update vector embeddings from a wiki page.
|
||||
|
||||
This endpoint:
|
||||
1. Fetches the page from Wiki.js
|
||||
2. Chunks the content (500 tokens with 50 token overlap)
|
||||
3. Generates embeddings via Ollama
|
||||
4. Upserts chunks to Qdrant with metadata
|
||||
|
||||
**Use Cases:**
|
||||
- Called automatically after page creation/update (via BackgroundTasks)
|
||||
- Called manually by user/Librarian to refresh vectors
|
||||
- Called by Scheduler for batch processing
|
||||
|
||||
**Example:** `POST /vector/update-from-page/5?user=jpmschweitzer`
|
||||
|
||||
**Returns:** Summary with chunks created and processing time
|
||||
"""
|
||||
try:
|
||||
summary = await vector_service.update_from_page(
|
||||
page_id=page_id,
|
||||
user=user,
|
||||
force_refresh=force_refresh
|
||||
)
|
||||
|
||||
if not summary.success:
|
||||
raise HTTPException(
|
||||
status_code=500,
|
||||
detail=f"Vector update failed: {summary.error_message}"
|
||||
)
|
||||
|
||||
return summary
|
||||
|
||||
except HTTPException:
|
||||
raise
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to update vectors from page {page_id}: {e}", exc_info=True)
|
||||
raise HTTPException(status_code=500, detail="Vector update failed")
|
||||
|
||||
|
||||
@router.delete("/pages/{page_id}", response_model=DeletePageChunksResponse)
|
||||
async def delete_page_chunks(
|
||||
page_id: int,
|
||||
user: str = Query(default=DEFAULT_USER, description="User identifier"),
|
||||
vector_service: VectorService = Depends(get_vector_service),
|
||||
api_key: str = Depends(verify_api_key)
|
||||
):
|
||||
"""
|
||||
Delete all vector chunks for a wiki page.
|
||||
|
||||
This is automatically called when a page is deleted from the wiki.
|
||||
|
||||
**Example:** `DELETE /vector/pages/5?user=jpmschweitzer`
|
||||
"""
|
||||
try:
|
||||
deleted_count = await vector_service.delete_page_chunks(
|
||||
page_id=page_id,
|
||||
user=user
|
||||
)
|
||||
|
||||
return DeletePageChunksResponse(
|
||||
page_id=page_id,
|
||||
chunks_deleted=deleted_count,
|
||||
success=deleted_count > 0,
|
||||
message=f"Deleted {deleted_count} chunks for page {page_id}"
|
||||
)
|
||||
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to delete chunks for page {page_id}: {e}", exc_info=True)
|
||||
raise HTTPException(status_code=500, detail="Failed to delete chunks")
|
||||
|
||||
|
||||
@router.get("/collections", response_model=CollectionListResponse)
|
||||
async def list_collections(
|
||||
vector_service: VectorService = Depends(get_vector_service),
|
||||
api_key: str = Depends(verify_api_key)
|
||||
):
|
||||
"""
|
||||
List all Qdrant collections with statistics.
|
||||
|
||||
Returns collection names, vector counts, and point counts.
|
||||
|
||||
**Example:** `GET /vector/collections`
|
||||
"""
|
||||
try:
|
||||
return await vector_service.list_collections()
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to list collections: {e}", exc_info=True)
|
||||
raise HTTPException(status_code=500, detail="Failed to list collections")
|
||||
@@ -1,498 +0,0 @@
|
||||
"""
|
||||
Wiki.js Webhook Handler
|
||||
|
||||
Receives webhook events from Wiki.js for page CRUD operations
|
||||
and processes them identically to AI-generated content.
|
||||
"""
|
||||
import logging
|
||||
from fastapi import APIRouter, Depends, HTTPException, BackgroundTasks
|
||||
from pydantic import BaseModel
|
||||
from typing import Optional, Literal
|
||||
|
||||
from src.core.dependencies import (
|
||||
get_ingestion_service,
|
||||
get_wiki_service,
|
||||
get_graph_service,
|
||||
verify_api_key
|
||||
)
|
||||
from src.services.ingestion_service import IngestionService
|
||||
from src.services.consolidation_service import ConsolidationService
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
router = APIRouter(prefix="/webhooks", tags=["Webhooks"])
|
||||
|
||||
|
||||
class WikiJSWebhookPayload(BaseModel):
|
||||
"""
|
||||
Wiki.js webhook payload structure.
|
||||
|
||||
See: https://docs.requarks.io/webhooks
|
||||
"""
|
||||
event: Literal["page.create", "page.update", "page.delete", "page.rename"]
|
||||
page: dict # Contains: id, title, path, content, etc.
|
||||
user: dict # Contains: id, name, email
|
||||
timestamp: str
|
||||
|
||||
|
||||
class WebhookProcessingResult(BaseModel):
|
||||
"""Result of webhook processing."""
|
||||
success: bool
|
||||
page_id: int
|
||||
page_title: str
|
||||
event: str
|
||||
ingested: bool
|
||||
entity_linking: Optional[dict] = None
|
||||
error: Optional[str] = None
|
||||
processing_time_ms: float
|
||||
|
||||
|
||||
@router.post("/wikijs", response_model=WebhookProcessingResult)
|
||||
async def handle_wikijs_webhook(
|
||||
payload: WikiJSWebhookPayload,
|
||||
background_tasks: BackgroundTasks,
|
||||
ingestion_service: IngestionService = Depends(get_ingestion_service),
|
||||
api_key: str = Depends(verify_api_key)
|
||||
):
|
||||
"""
|
||||
Handle Wiki.js webhook events.
|
||||
|
||||
Processes page changes identically to AI-generated content:
|
||||
1. Ingest page → Vector embeddings + Graph entity extraction
|
||||
2. Apply bidirectional entity linking
|
||||
|
||||
This ensures user-edited pages have the same structure and
|
||||
connectivity as AI-generated pages.
|
||||
|
||||
Configuration in Wiki.js:
|
||||
- Administration → Webhooks
|
||||
- Add new webhook:
|
||||
- URL: http://library-desk:8089/webhooks/wikijs
|
||||
- Events: page.create, page.update
|
||||
- Headers: Authorization: Bearer <API_KEY>
|
||||
"""
|
||||
import time
|
||||
start_time = time.time()
|
||||
|
||||
page_id = payload.page.get("id")
|
||||
page_title = payload.page.get("title")
|
||||
user_email = payload.user.get("email", "unknown")
|
||||
|
||||
# Extract user identifier from email (assumes format: user@domain)
|
||||
# Adjust this based on your user mapping strategy
|
||||
user = extract_user_from_email(user_email)
|
||||
|
||||
logger.info(
|
||||
f"Received Wiki.js webhook: {payload.event} "
|
||||
f"for page {page_id} ('{page_title}') by {user_email}"
|
||||
)
|
||||
|
||||
try:
|
||||
# Handle different events
|
||||
if payload.event == "page.delete":
|
||||
# For deletions, clean up vectors and graph
|
||||
logger.info(f"Page {page_id} deleted, cleaning up knowledge base")
|
||||
background_tasks.add_task(
|
||||
cleanup_deleted_page,
|
||||
page_id=page_id,
|
||||
page_title=page_title,
|
||||
user=user,
|
||||
ingestion_service=ingestion_service
|
||||
)
|
||||
processing_time = (time.time() - start_time) * 1000
|
||||
return WebhookProcessingResult(
|
||||
success=True,
|
||||
page_id=page_id,
|
||||
page_title=page_title,
|
||||
event=payload.event,
|
||||
ingested=False,
|
||||
processing_time_ms=processing_time
|
||||
)
|
||||
|
||||
# For create/update events, process the page
|
||||
if payload.event in ["page.create", "page.update"]:
|
||||
# Run ingestion and entity linking in background
|
||||
# to avoid blocking the webhook response
|
||||
background_tasks.add_task(
|
||||
process_wiki_page_change,
|
||||
page_id=page_id,
|
||||
page_title=page_title,
|
||||
user=user,
|
||||
event=payload.event,
|
||||
ingestion_service=ingestion_service
|
||||
)
|
||||
|
||||
processing_time = (time.time() - start_time) * 1000
|
||||
|
||||
return WebhookProcessingResult(
|
||||
success=True,
|
||||
page_id=page_id,
|
||||
page_title=page_title,
|
||||
event=payload.event,
|
||||
ingested=True, # Will be processed in background
|
||||
processing_time_ms=processing_time
|
||||
)
|
||||
|
||||
# Handle page rename/move events
|
||||
if payload.event == "page.rename":
|
||||
# Update path and re-link if title changed
|
||||
new_path = payload.page.get("path")
|
||||
new_title = payload.page.get("title")
|
||||
old_path = payload.page.get("oldPath", new_path)
|
||||
old_title = payload.page.get("oldTitle", new_title)
|
||||
|
||||
background_tasks.add_task(
|
||||
process_page_rename,
|
||||
page_id=page_id,
|
||||
old_path=old_path,
|
||||
new_path=new_path,
|
||||
old_title=old_title,
|
||||
new_title=new_title,
|
||||
user=user,
|
||||
ingestion_service=ingestion_service
|
||||
)
|
||||
|
||||
processing_time = (time.time() - start_time) * 1000
|
||||
|
||||
return WebhookProcessingResult(
|
||||
success=True,
|
||||
page_id=page_id,
|
||||
page_title=page_title,
|
||||
event=payload.event,
|
||||
ingested=False, # Path update only, no re-embedding needed
|
||||
processing_time_ms=processing_time
|
||||
)
|
||||
|
||||
except Exception as e:
|
||||
logger.error(f"Webhook processing failed: {e}", exc_info=True)
|
||||
processing_time = (time.time() - start_time) * 1000
|
||||
|
||||
return WebhookProcessingResult(
|
||||
success=False,
|
||||
page_id=page_id,
|
||||
page_title=page_title,
|
||||
event=payload.event,
|
||||
ingested=False,
|
||||
error=str(e),
|
||||
processing_time_ms=processing_time
|
||||
)
|
||||
|
||||
|
||||
async def process_wiki_page_change(
|
||||
page_id: int,
|
||||
page_title: str,
|
||||
user: str,
|
||||
event: str,
|
||||
ingestion_service: IngestionService
|
||||
):
|
||||
"""
|
||||
Process wiki page change identically to AI-generated content.
|
||||
|
||||
This ensures consistency between manual and AI workflows.
|
||||
|
||||
Steps:
|
||||
1. Ingest page → Vector embeddings + Graph entities
|
||||
2. Apply bidirectional entity linking
|
||||
"""
|
||||
from src.core.dependencies import get_neo4j_client, get_ollama_client, get_wikijs_client
|
||||
from src.config import get_settings
|
||||
|
||||
try:
|
||||
logger.info(f"Processing {event} for page {page_id} ('{page_title}')")
|
||||
|
||||
# STEP 1: Ingest page (vector + graph)
|
||||
logger.info(f"Ingesting page {page_id} into knowledge base")
|
||||
await ingestion_service.ingest_page(
|
||||
page_id=page_id,
|
||||
user=user,
|
||||
force_refresh=(event == "page.update") # Force refresh on updates
|
||||
)
|
||||
logger.info(f"Ingestion complete for page {page_id}")
|
||||
|
||||
# STEP 2: Apply bidirectional entity linking
|
||||
logger.info(f"Applying bidirectional entity linking for page {page_id}")
|
||||
|
||||
# Import consolidation service to use the entity linking method
|
||||
settings = get_settings()
|
||||
consolidation_service = ConsolidationService(
|
||||
neo4j=get_neo4j_client(),
|
||||
ollama=get_ollama_client(),
|
||||
wiki=get_wikijs_client(),
|
||||
settings=settings,
|
||||
ingestion_service=ingestion_service
|
||||
)
|
||||
|
||||
link_stats = await consolidation_service._apply_bidirectional_entity_linking(
|
||||
page_id=page_id,
|
||||
page_title=page_title,
|
||||
user=user
|
||||
)
|
||||
|
||||
logger.info(
|
||||
f"Entity linking complete for page {page_id}: "
|
||||
f"{link_stats['forward_links']} forward links, "
|
||||
f"{link_stats['backward_links']} backward links "
|
||||
f"({link_stats['pages_updated']} pages updated)"
|
||||
)
|
||||
|
||||
logger.info(f"Successfully processed {event} for page {page_id}")
|
||||
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to process page {page_id}: {e}", exc_info=True)
|
||||
|
||||
|
||||
def extract_user_from_email(email: str) -> str:
|
||||
"""
|
||||
Extract user identifier from email.
|
||||
|
||||
Customize this based on your user mapping strategy:
|
||||
- Option 1: Use email prefix (user@domain → user)
|
||||
- Option 2: Map email to Wiki.js username
|
||||
- Option 3: Use email directly
|
||||
|
||||
Args:
|
||||
email: User email from Wiki.js webhook
|
||||
|
||||
Returns:
|
||||
User identifier for multi-tenancy
|
||||
"""
|
||||
# Option 1: Extract prefix from email
|
||||
if "@" in email:
|
||||
return email.split("@")[0]
|
||||
|
||||
# Fallback: use email as-is
|
||||
return email
|
||||
|
||||
|
||||
async def process_page_rename(
|
||||
page_id: int,
|
||||
old_path: str,
|
||||
new_path: str,
|
||||
old_title: str,
|
||||
new_title: str,
|
||||
user: str,
|
||||
ingestion_service: IngestionService
|
||||
):
|
||||
"""
|
||||
Process page rename/move events.
|
||||
|
||||
Handles two scenarios:
|
||||
1. Path change only (move to different location) - update path in graph
|
||||
2. Title change (rename) - update title, re-link entities, re-process
|
||||
|
||||
Args:
|
||||
page_id: Wiki page ID
|
||||
old_path: Previous page path
|
||||
new_path: New page path
|
||||
old_title: Previous page title
|
||||
new_title: New page title
|
||||
user: User identifier
|
||||
ingestion_service: Ingestion service instance
|
||||
"""
|
||||
from src.core.dependencies import get_neo4j_client
|
||||
from src.core.multi_tenancy import get_neo4j_user_label
|
||||
|
||||
try:
|
||||
logger.info(
|
||||
f"Processing rename for page {page_id}: "
|
||||
f"'{old_title}' → '{new_title}', "
|
||||
f"'{old_path}' → '{new_path}'"
|
||||
)
|
||||
|
||||
neo4j = get_neo4j_client()
|
||||
user_doc_label = get_neo4j_user_label(user)
|
||||
|
||||
# Check if title changed (rename) or just path changed (move)
|
||||
title_changed = old_title != new_title
|
||||
path_changed = old_path != new_path
|
||||
|
||||
if not title_changed and not path_changed:
|
||||
logger.info("No changes detected, skipping processing")
|
||||
return
|
||||
|
||||
# Update Document node in graph
|
||||
update_query = f"""
|
||||
MATCH (d:{user_doc_label}:Document {{page_id: $page_id}})
|
||||
SET d.path = $new_path,
|
||||
d.title = $new_title,
|
||||
d.updated_at = datetime()
|
||||
RETURN d
|
||||
"""
|
||||
|
||||
try:
|
||||
await neo4j.execute_query(update_query, {
|
||||
"page_id": page_id,
|
||||
"new_path": new_path,
|
||||
"new_title": new_title
|
||||
})
|
||||
logger.info(f"Updated Document node for page {page_id}")
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to update Document node: {e}")
|
||||
|
||||
# If title changed, need to re-process for entity linking
|
||||
if title_changed:
|
||||
logger.info(f"Title changed, re-processing page {page_id}")
|
||||
|
||||
# Re-ingest to update entities (title might be an entity)
|
||||
try:
|
||||
await ingestion_service.ingest_page(
|
||||
page_id=page_id,
|
||||
user=user,
|
||||
force_refresh=True
|
||||
)
|
||||
logger.info(f"Re-ingested page {page_id} after title change")
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to re-ingest page {page_id}: {e}")
|
||||
|
||||
# Re-apply bidirectional entity linking
|
||||
from src.config import get_settings
|
||||
from src.core.dependencies import get_ollama_client, get_wikijs_client
|
||||
from src.services.consolidation_service import ConsolidationService
|
||||
|
||||
settings = get_settings()
|
||||
consolidation_service = ConsolidationService(
|
||||
neo4j=neo4j,
|
||||
ollama=get_ollama_client(),
|
||||
wiki=get_wikijs_client(),
|
||||
settings=settings,
|
||||
ingestion_service=ingestion_service
|
||||
)
|
||||
|
||||
try:
|
||||
link_stats = await consolidation_service._apply_bidirectional_entity_linking(
|
||||
page_id=page_id,
|
||||
page_title=new_title,
|
||||
user=user
|
||||
)
|
||||
logger.info(
|
||||
f"Entity linking complete for renamed page {page_id}: "
|
||||
f"{link_stats['forward_links']} forward links, "
|
||||
f"{link_stats['backward_links']} backward links"
|
||||
)
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to apply entity linking: {e}")
|
||||
|
||||
elif path_changed:
|
||||
logger.info(f"Path changed only (move), no re-processing needed")
|
||||
|
||||
logger.info(f"Rename processing complete for page {page_id}")
|
||||
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to process rename for page {page_id}: {e}", exc_info=True)
|
||||
|
||||
|
||||
async def cleanup_deleted_page(
|
||||
page_id: int,
|
||||
page_title: str,
|
||||
user: str,
|
||||
ingestion_service: IngestionService
|
||||
):
|
||||
"""
|
||||
Clean up vectors and graph when a page is deleted.
|
||||
|
||||
Steps:
|
||||
1. Remove vector embeddings from Qdrant
|
||||
2. Remove Document node from Neo4j
|
||||
3. Clean up orphaned entities (entities only connected to this document)
|
||||
4. Remove broken MENTIONS relationships
|
||||
"""
|
||||
from src.core.dependencies import get_neo4j_client, get_vector_service
|
||||
from src.core.multi_tenancy import get_neo4j_user_base_label, get_neo4j_user_label
|
||||
|
||||
try:
|
||||
logger.info(f"Cleaning up deleted page {page_id} ('{page_title}')")
|
||||
|
||||
vector_service = get_vector_service()
|
||||
neo4j = get_neo4j_client()
|
||||
|
||||
user_base_label = get_neo4j_user_base_label(user)
|
||||
user_doc_label = get_neo4j_user_label(user)
|
||||
|
||||
# STEP 1: Remove vectors from Qdrant
|
||||
logger.info(f"Removing vectors for page {page_id}")
|
||||
try:
|
||||
await vector_service.delete_page_chunks(page_id, user)
|
||||
logger.info(f"Removed vectors for page {page_id}")
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to remove vectors for page {page_id}: {e}")
|
||||
|
||||
# STEP 2: Find and store orphaned entities before deletion
|
||||
# (entities that only have this document mentioning them)
|
||||
orphaned_entities_query = f"""
|
||||
MATCH (d:{user_doc_label}:Document {{page_id: $page_id}})-[:MENTIONS]->(e:{user_base_label})
|
||||
WHERE NOT e:Document
|
||||
WITH e, count{{(d2:Document)-[:MENTIONS]->(e)}} as mention_count
|
||||
WHERE mention_count = 1
|
||||
RETURN e.id as entity_id, e.name as entity_name, labels(e) as labels
|
||||
"""
|
||||
|
||||
try:
|
||||
orphaned = await neo4j.execute_query(orphaned_entities_query, {"page_id": page_id})
|
||||
logger.info(f"Found {len(orphaned)} orphaned entities for page {page_id}")
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to find orphaned entities: {e}")
|
||||
orphaned = []
|
||||
|
||||
# STEP 3: Delete Document node (this will cascade delete MENTIONS relationships)
|
||||
delete_doc_query = f"""
|
||||
MATCH (d:{user_doc_label}:Document {{page_id: $page_id}})
|
||||
DETACH DELETE d
|
||||
RETURN count(d) as deleted_count
|
||||
"""
|
||||
|
||||
try:
|
||||
result = await neo4j.execute_query(delete_doc_query, {"page_id": page_id})
|
||||
deleted_count = result[0]["deleted_count"] if result else 0
|
||||
logger.info(f"Deleted {deleted_count} Document node(s) for page {page_id}")
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to delete Document node: {e}")
|
||||
|
||||
# STEP 4: Delete orphaned entities
|
||||
if orphaned:
|
||||
for entity in orphaned:
|
||||
entity_id = entity["entity_id"]
|
||||
entity_name = entity["entity_name"]
|
||||
|
||||
delete_entity_query = f"""
|
||||
MATCH (e:{user_base_label} {{id: $entity_id}})
|
||||
WHERE NOT e:Document
|
||||
AND NOT EXISTS {{(d:Document)-[:MENTIONS]->(e)}}
|
||||
DETACH DELETE e
|
||||
RETURN count(e) as deleted_count
|
||||
"""
|
||||
|
||||
try:
|
||||
result = await neo4j.execute_query(delete_entity_query, {"entity_id": entity_id})
|
||||
deleted = result[0]["deleted_count"] if result else 0
|
||||
if deleted > 0:
|
||||
logger.info(f"Deleted orphaned entity: {entity_name}")
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to delete orphaned entity {entity_name}: {e}")
|
||||
|
||||
# STEP 5: Clean up broken SearchQuery relationships
|
||||
cleanup_search_query = f"""
|
||||
MATCH (sq:SearchQuery)-[r:FOUND]->(d:Document)
|
||||
WHERE NOT EXISTS {{(d)}}
|
||||
DELETE r
|
||||
RETURN count(r) as cleaned_count
|
||||
"""
|
||||
|
||||
try:
|
||||
result = await neo4j.execute_query(cleanup_search_query, {})
|
||||
cleaned = result[0]["cleaned_count"] if result else 0
|
||||
if cleaned > 0:
|
||||
logger.info(f"Cleaned up {cleaned} broken SearchQuery relationships")
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to clean SearchQuery relationships: {e}")
|
||||
|
||||
logger.info(f"Cleanup complete for deleted page {page_id}")
|
||||
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to cleanup deleted page {page_id}: {e}", exc_info=True)
|
||||
|
||||
|
||||
# Health check endpoint
|
||||
@router.get("/health")
|
||||
async def webhook_health():
|
||||
"""Health check for webhook endpoint."""
|
||||
return {"status": "ok", "service": "webhooks"}
|
||||
@@ -1,409 +0,0 @@
|
||||
"""
|
||||
Wiki router for Library Desk API.
|
||||
|
||||
Endpoints for wiki page and dossier management.
|
||||
All operations are scoped to user namespaces for multi-tenancy.
|
||||
"""
|
||||
|
||||
from fastapi import APIRouter, HTTPException, Depends, Query, Security, BackgroundTasks
|
||||
from fastapi.security import HTTPAuthorizationCredentials
|
||||
from typing import Optional
|
||||
import logging
|
||||
|
||||
from src.models.wiki import (
|
||||
WikiPage, WikiPageList, WikiPageCreate, WikiPageUpdate, WikiPageMove,
|
||||
WikiOperationResponse, WikiSearchResponse,
|
||||
DossierList, WikiSearchResult
|
||||
)
|
||||
from src.services.wiki_service import WikiService
|
||||
from src.services.graph_service import GraphService
|
||||
from src.services.vector_service import VectorService
|
||||
from src.clients.wikijs_client import WikiJSClient
|
||||
from src.clients.neo4j_client import Neo4jClient
|
||||
from src.clients.qdrant_client import QdrantClientWrapper
|
||||
from src.clients.ollama_client import OllamaClient
|
||||
from src.core.dependencies import WikiJSDep, Neo4jDep, QdrantDep, OllamaDep, verify_api_key
|
||||
from src.core.multi_tenancy import DEFAULT_USER
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
router = APIRouter(prefix="/wiki", tags=["Wiki"])
|
||||
|
||||
|
||||
# Dependency to get wiki service
|
||||
def get_wiki_service(wiki_client: WikiJSDep) -> WikiService:
|
||||
"""Get wiki service instance."""
|
||||
return WikiService(wiki_client)
|
||||
|
||||
|
||||
# Dependency to get graph service
|
||||
def get_graph_service(neo4j_client: Neo4jDep, wiki_client: WikiJSDep) -> GraphService:
|
||||
"""Get graph service instance."""
|
||||
return GraphService(neo4j_client, wiki_client)
|
||||
|
||||
|
||||
# Dependency to get vector service
|
||||
def get_vector_service(
|
||||
qdrant_client: QdrantDep,
|
||||
wiki_client: WikiJSDep,
|
||||
ollama_client: OllamaDep
|
||||
) -> VectorService:
|
||||
"""Get vector service instance."""
|
||||
return VectorService(qdrant_client, wiki_client, ollama_client)
|
||||
|
||||
|
||||
# Page operations
|
||||
@router.get("/pages", response_model=WikiPageList)
|
||||
async def list_pages(
|
||||
user: str = Query(default=DEFAULT_USER, description="User identifier"),
|
||||
tag: Optional[str] = Query(default=None, description="Filter by tag (dossier)"),
|
||||
limit: int = Query(default=50, ge=1, le=200, description="Maximum pages to return"),
|
||||
wiki_service: WikiService = Depends(get_wiki_service),
|
||||
api_key: str = Depends(verify_api_key)
|
||||
):
|
||||
"""
|
||||
List wiki pages for a user.
|
||||
|
||||
Optionally filter by tag (dossier). Pages are scoped to user's namespace.
|
||||
|
||||
**Example:** `/wiki/pages?user=jpmschweitzer&tag=projects&limit=20`
|
||||
"""
|
||||
try:
|
||||
return await wiki_service.list_pages(user=user, tag=tag, limit=limit)
|
||||
except ValueError as e:
|
||||
raise HTTPException(status_code=400, detail=str(e))
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to list pages: {e}", exc_info=True)
|
||||
raise HTTPException(status_code=500, detail="Internal server error")
|
||||
|
||||
|
||||
@router.get("/pages/{page_id}", response_model=WikiPage)
|
||||
async def get_page(
|
||||
page_id: int,
|
||||
user: str = Query(default=DEFAULT_USER, description="User identifier"),
|
||||
wiki_service: WikiService = Depends(get_wiki_service),
|
||||
api_key: str = Depends(verify_api_key)
|
||||
):
|
||||
"""
|
||||
Get a single wiki page by ID.
|
||||
|
||||
Access is restricted to pages within the user's namespace.
|
||||
|
||||
**Example:** `/wiki/pages/123?user=jpmschweitzer`
|
||||
"""
|
||||
try:
|
||||
page = await wiki_service.get_page(page_id=page_id, user=user)
|
||||
if not page:
|
||||
raise HTTPException(status_code=404, detail=f"Page {page_id} not found or access denied")
|
||||
return page
|
||||
except HTTPException:
|
||||
raise
|
||||
except ValueError as e:
|
||||
raise HTTPException(status_code=400, detail=str(e))
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to get page {page_id}: {e}", exc_info=True)
|
||||
raise HTTPException(status_code=500, detail="Internal server error")
|
||||
|
||||
|
||||
@router.post("/pages", response_model=WikiPage, status_code=201)
|
||||
async def create_page(
|
||||
page_data: WikiPageCreate,
|
||||
background_tasks: BackgroundTasks,
|
||||
wiki_service: WikiService = Depends(get_wiki_service),
|
||||
graph_service: GraphService = Depends(get_graph_service),
|
||||
vector_service: VectorService = Depends(get_vector_service),
|
||||
api_key: str = Depends(verify_api_key)
|
||||
):
|
||||
"""
|
||||
Create a new wiki page.
|
||||
|
||||
The page will be created in the user's namespace. If path doesn't start
|
||||
with namespace, it will be automatically prefixed.
|
||||
|
||||
**Auto-updates knowledge graph and vector embeddings**: After creating
|
||||
the page, the graph and vectors are automatically updated in the
|
||||
background to extract entities/relationships and generate semantic embeddings.
|
||||
|
||||
**Example Request:**
|
||||
```json
|
||||
{
|
||||
"title": "Library Desk Architecture",
|
||||
"path": "/projects/library-desk/architecture",
|
||||
"content": "# Architecture\\n\\nThis describes...",
|
||||
"description": "Architecture documentation",
|
||||
"tags": ["projects", "architecture"],
|
||||
"user": "jpmschweitzer"
|
||||
}
|
||||
```
|
||||
"""
|
||||
try:
|
||||
page = await wiki_service.create_page(page_data)
|
||||
|
||||
user = page_data.user or DEFAULT_USER
|
||||
|
||||
# Schedule BOTH graph and vector updates in background (non-blocking)
|
||||
background_tasks.add_task(
|
||||
graph_service.update_from_page,
|
||||
page_id=page.id,
|
||||
user=user
|
||||
)
|
||||
background_tasks.add_task(
|
||||
vector_service.update_from_page,
|
||||
page_id=page.id,
|
||||
user=user
|
||||
)
|
||||
|
||||
logger.info(f"Page {page.id} created, graph and vector updates scheduled")
|
||||
return page
|
||||
|
||||
except ValueError as e:
|
||||
raise HTTPException(status_code=400, detail=str(e))
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to create page: {e}", exc_info=True)
|
||||
raise HTTPException(status_code=500, detail="Internal server error")
|
||||
|
||||
|
||||
@router.put("/pages/{page_id}", response_model=WikiPage)
|
||||
async def update_page(
|
||||
page_id: int,
|
||||
page_data: WikiPageUpdate,
|
||||
background_tasks: BackgroundTasks,
|
||||
user: str = Query(default=DEFAULT_USER, description="User identifier"),
|
||||
wiki_service: WikiService = Depends(get_wiki_service),
|
||||
graph_service: GraphService = Depends(get_graph_service),
|
||||
vector_service: VectorService = Depends(get_vector_service),
|
||||
api_key: str = Depends(verify_api_key)
|
||||
):
|
||||
"""
|
||||
Update an existing wiki page.
|
||||
|
||||
Only pages within the user's namespace can be updated.
|
||||
Partial updates are supported - only provided fields will be updated.
|
||||
|
||||
**Auto-updates knowledge graph and vector embeddings**: After updating
|
||||
the page, the graph and vectors are automatically refreshed in the
|
||||
background to reflect the changes.
|
||||
|
||||
**Example Request:**
|
||||
```json
|
||||
{
|
||||
"title": "Updated Title",
|
||||
"tags": ["projects", "updated"]
|
||||
}
|
||||
```
|
||||
"""
|
||||
try:
|
||||
page = await wiki_service.update_page(
|
||||
page_id=page_id,
|
||||
page_data=page_data,
|
||||
user=user
|
||||
)
|
||||
|
||||
# Schedule BOTH graph and vector updates in background (non-blocking)
|
||||
background_tasks.add_task(
|
||||
graph_service.update_from_page,
|
||||
page_id=page_id,
|
||||
user=user,
|
||||
force_refresh=True # Force refresh on updates
|
||||
)
|
||||
background_tasks.add_task(
|
||||
vector_service.update_from_page,
|
||||
page_id=page_id,
|
||||
user=user,
|
||||
force_refresh=True # Force refresh on updates
|
||||
)
|
||||
|
||||
logger.info(f"Page {page_id} updated, graph and vector refresh scheduled")
|
||||
return page
|
||||
|
||||
except ValueError as e:
|
||||
raise HTTPException(status_code=400, detail=str(e))
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to update page {page_id}: {e}", exc_info=True)
|
||||
raise HTTPException(status_code=500, detail="Internal server error")
|
||||
|
||||
|
||||
@router.delete("/pages/{page_id}", response_model=WikiOperationResponse)
|
||||
async def delete_page(
|
||||
page_id: int,
|
||||
background_tasks: BackgroundTasks,
|
||||
user: str = Query(default=DEFAULT_USER, description="User identifier"),
|
||||
wiki_service: WikiService = Depends(get_wiki_service),
|
||||
vector_service: VectorService = Depends(get_vector_service),
|
||||
graph_service: GraphService = Depends(get_graph_service),
|
||||
api_key: str = Depends(verify_api_key)
|
||||
):
|
||||
"""
|
||||
Delete a wiki page.
|
||||
|
||||
Only pages within the user's namespace can be deleted.
|
||||
This operation cannot be undone.
|
||||
|
||||
**Auto-cleanup**: Vector chunks and graph nodes for this page are
|
||||
automatically deleted in the background.
|
||||
"""
|
||||
try:
|
||||
success = await wiki_service.delete_page(page_id=page_id, user=user)
|
||||
|
||||
# Schedule vector cleanup in background
|
||||
background_tasks.add_task(
|
||||
vector_service.delete_page_chunks,
|
||||
page_id=page_id,
|
||||
user=user
|
||||
)
|
||||
|
||||
# Schedule graph cleanup in background
|
||||
background_tasks.add_task(
|
||||
graph_service.delete_page,
|
||||
page_id=page_id,
|
||||
user=user
|
||||
)
|
||||
|
||||
logger.info(f"Page {page_id} deleted, vector and graph cleanup scheduled")
|
||||
return WikiOperationResponse(
|
||||
success=success,
|
||||
message=f"Page {page_id} deleted successfully",
|
||||
page_id=page_id
|
||||
)
|
||||
except ValueError as e:
|
||||
raise HTTPException(status_code=404, detail=str(e))
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to delete page {page_id}: {e}", exc_info=True)
|
||||
raise HTTPException(status_code=500, detail="Internal server error")
|
||||
|
||||
|
||||
@router.post("/pages/{page_id}/move", response_model=WikiOperationResponse)
|
||||
async def move_page(
|
||||
page_id: int,
|
||||
move_data: WikiPageMove,
|
||||
user: str = Query(default=DEFAULT_USER, description="User identifier"),
|
||||
wiki_service: WikiService = Depends(get_wiki_service),
|
||||
api_key: str = Depends(verify_api_key)
|
||||
):
|
||||
"""
|
||||
Move or rename a wiki page.
|
||||
|
||||
The new path must be within the user's namespace.
|
||||
|
||||
**Example Request:**
|
||||
```json
|
||||
{
|
||||
"new_path": "/projects/library-desk/docs/architecture"
|
||||
}
|
||||
```
|
||||
"""
|
||||
try:
|
||||
success = await wiki_service.move_page(
|
||||
page_id=page_id,
|
||||
new_path=move_data.new_path,
|
||||
user=user
|
||||
)
|
||||
if not success:
|
||||
raise HTTPException(status_code=500, detail="Failed to move page")
|
||||
|
||||
return WikiOperationResponse(
|
||||
success=True,
|
||||
message=f"Page {page_id} moved successfully",
|
||||
page_id=page_id,
|
||||
page_path=move_data.new_path
|
||||
)
|
||||
except ValueError as e:
|
||||
raise HTTPException(status_code=400, detail=str(e))
|
||||
except HTTPException:
|
||||
raise
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to move page {page_id}: {e}", exc_info=True)
|
||||
raise HTTPException(status_code=500, detail="Internal server error")
|
||||
|
||||
|
||||
# Search operations
|
||||
@router.get("/search", response_model=WikiSearchResponse)
|
||||
async def search_pages(
|
||||
q: str = Query(..., min_length=1, description="Search query"),
|
||||
user: str = Query(default=DEFAULT_USER, description="User identifier"),
|
||||
limit: int = Query(default=20, ge=1, le=100, description="Maximum results"),
|
||||
wiki_service: WikiService = Depends(get_wiki_service),
|
||||
api_key: str = Depends(verify_api_key)
|
||||
):
|
||||
"""
|
||||
Search wiki pages within user's namespace.
|
||||
|
||||
**Example:** `/wiki/search?q=architecture&user=jpmschweitzer&limit=10`
|
||||
"""
|
||||
try:
|
||||
results = await wiki_service.search_pages(
|
||||
query=q,
|
||||
user=user,
|
||||
limit=limit
|
||||
)
|
||||
|
||||
return WikiSearchResponse(
|
||||
results=[
|
||||
WikiSearchResult(
|
||||
id=r.id,
|
||||
path=r.path,
|
||||
title=r.title,
|
||||
description=r.description,
|
||||
relevance=None
|
||||
)
|
||||
for r in results
|
||||
],
|
||||
query=q,
|
||||
total=len(results)
|
||||
)
|
||||
except ValueError as e:
|
||||
raise HTTPException(status_code=400, detail=str(e))
|
||||
except Exception as e:
|
||||
logger.error(f"Search failed: {e}", exc_info=True)
|
||||
raise HTTPException(status_code=500, detail="Internal server error")
|
||||
|
||||
|
||||
# Dossier operations
|
||||
@router.get("/dossiers", response_model=DossierList)
|
||||
async def list_dossiers(
|
||||
user: str = Query(default=DEFAULT_USER, description="User identifier"),
|
||||
wiki_service: WikiService = Depends(get_wiki_service),
|
||||
api_key: str = Depends(verify_api_key)
|
||||
):
|
||||
"""
|
||||
List all dossiers (unique tags) for a user.
|
||||
|
||||
Dossiers are tag-based collections of pages.
|
||||
|
||||
**Example:** `/wiki/dossiers?user=jpmschweitzer`
|
||||
"""
|
||||
try:
|
||||
return await wiki_service.list_dossiers(user=user)
|
||||
except ValueError as e:
|
||||
raise HTTPException(status_code=400, detail=str(e))
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to list dossiers: {e}", exc_info=True)
|
||||
raise HTTPException(status_code=500, detail="Internal server error")
|
||||
|
||||
|
||||
@router.get("/dossiers/{dossier_name}/pages", response_model=WikiPageList)
|
||||
async def get_dossier_pages(
|
||||
dossier_name: str,
|
||||
user: str = Query(default=DEFAULT_USER, description="User identifier"),
|
||||
limit: int = Query(default=100, ge=1, le=500, description="Maximum pages"),
|
||||
wiki_service: WikiService = Depends(get_wiki_service),
|
||||
api_key: str = Depends(verify_api_key)
|
||||
):
|
||||
"""
|
||||
Get all pages in a dossier.
|
||||
|
||||
Returns pages tagged with the dossier name.
|
||||
|
||||
**Example:** `/wiki/dossiers/projects/pages?user=jpmschweitzer`
|
||||
"""
|
||||
try:
|
||||
return await wiki_service.get_dossier_pages(
|
||||
dossier_name=dossier_name,
|
||||
user=user,
|
||||
limit=limit
|
||||
)
|
||||
except ValueError as e:
|
||||
raise HTTPException(status_code=400, detail=str(e))
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to get dossier pages: {e}", exc_info=True)
|
||||
raise HTTPException(status_code=500, detail="Internal server error")
|
||||
@@ -1,932 +0,0 @@
|
||||
"""
|
||||
Knowledge Consolidation Service (Librarian Logic)
|
||||
|
||||
Processes unprocessed SearchQuery nodes from HybridRAG searches
|
||||
to consolidate new knowledge into wiki pages.
|
||||
|
||||
This service:
|
||||
1. Queries Neo4j for unprocessed SearchQuery nodes
|
||||
2. Analyzes web results with Ollama for novel information
|
||||
3. Creates/updates wiki pages with new facts
|
||||
4. Updates knowledge graph with new entities
|
||||
5. Marks SearchQuery nodes as processed
|
||||
"""
|
||||
import logging
|
||||
import json
|
||||
from datetime import datetime, timedelta
|
||||
from typing import List, Dict, Any, Optional
|
||||
|
||||
from src.clients.neo4j_client import Neo4jClient
|
||||
from src.clients.ollama_client import OllamaClient
|
||||
from src.clients.wikijs_client import WikiJSClient
|
||||
from src.services.wiki_page_writer import WikiPageWriter
|
||||
from src.models.consolidation import (
|
||||
SearchQueryInfo,
|
||||
ConsolidationResult,
|
||||
ConsolidationResponse
|
||||
)
|
||||
from src.config import Settings
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
|
||||
class ConsolidationService:
|
||||
"""
|
||||
Service for consolidating knowledge from search results.
|
||||
"""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
neo4j: Neo4jClient,
|
||||
ollama: OllamaClient,
|
||||
wiki: WikiJSClient,
|
||||
settings: Settings,
|
||||
ingestion_service: Optional["IngestionService"] = None
|
||||
):
|
||||
self.neo4j = neo4j
|
||||
self.ollama = ollama
|
||||
self.wiki = wiki
|
||||
self.settings = settings
|
||||
self.wiki_page_writer = WikiPageWriter(ollama_client=ollama)
|
||||
self.ingestion_service = ingestion_service # Optional to avoid circular dependency
|
||||
|
||||
async def consolidate_knowledge(
|
||||
self,
|
||||
process_limit: int = 10,
|
||||
lookback_days: int = 7,
|
||||
min_web_results: int = 2,
|
||||
dry_run: bool = False
|
||||
) -> ConsolidationResponse:
|
||||
"""
|
||||
Process unprocessed search queries and consolidate knowledge.
|
||||
|
||||
Args:
|
||||
process_limit: Maximum searches to process
|
||||
lookback_days: Only process searches from last N days
|
||||
min_web_results: Minimum web results required to consolidate
|
||||
dry_run: If True, analyze but don't create pages
|
||||
|
||||
Returns:
|
||||
ConsolidationResponse with processing results
|
||||
"""
|
||||
logger.info(f"Starting knowledge consolidation")
|
||||
logger.info(f"Limits: process={process_limit}, lookback={lookback_days}d, min_web={min_web_results}")
|
||||
if dry_run:
|
||||
logger.warning("DRY RUN MODE - will not create wiki pages")
|
||||
|
||||
# Find unprocessed searches
|
||||
unprocessed = await self._find_unprocessed_searches(lookback_days, process_limit)
|
||||
|
||||
if not unprocessed:
|
||||
logger.info("No unprocessed searches found")
|
||||
return ConsolidationResponse(
|
||||
total_found=0,
|
||||
processed_count=0,
|
||||
pages_created=0,
|
||||
pages_updated=0,
|
||||
entities_added=0,
|
||||
errors=[],
|
||||
results=[],
|
||||
dry_run=dry_run
|
||||
)
|
||||
|
||||
logger.info(f"Found {len(unprocessed)} unprocessed searches")
|
||||
|
||||
# Process each search
|
||||
results: List[ConsolidationResult] = []
|
||||
total_pages_created = 0
|
||||
total_pages_updated = 0
|
||||
total_entities_added = 0
|
||||
errors: List[str] = []
|
||||
|
||||
for search in unprocessed:
|
||||
try:
|
||||
result = await self._process_search(
|
||||
search=search,
|
||||
min_web_results=min_web_results,
|
||||
dry_run=dry_run
|
||||
)
|
||||
|
||||
if result:
|
||||
results.append(result)
|
||||
total_pages_created += result.pages_created
|
||||
total_pages_updated += result.pages_updated
|
||||
total_entities_added += result.entities_added
|
||||
|
||||
# Mark as processed if not dry run (even if skipped)
|
||||
# This prevents searches from accumulating when they don't meet criteria
|
||||
if not dry_run:
|
||||
await self._mark_search_processed(search['id'])
|
||||
|
||||
except Exception as e:
|
||||
error_msg = f"Search {search['id'][:8]}: {str(e)}"
|
||||
logger.error(f"Failed to process search: {error_msg}", exc_info=True)
|
||||
errors.append(error_msg)
|
||||
results.append(ConsolidationResult(
|
||||
search_id=search['id'],
|
||||
query=search['query'],
|
||||
error=str(e)
|
||||
))
|
||||
|
||||
# Mark as processed even on error (to avoid retrying failed searches forever)
|
||||
if not dry_run:
|
||||
await self._mark_search_processed(search['id'])
|
||||
|
||||
# Build response
|
||||
processed_count = len([r for r in results if not r.error])
|
||||
|
||||
response = ConsolidationResponse(
|
||||
total_found=len(unprocessed),
|
||||
processed_count=processed_count,
|
||||
pages_created=total_pages_created,
|
||||
pages_updated=total_pages_updated,
|
||||
entities_added=total_entities_added,
|
||||
errors=errors,
|
||||
results=results,
|
||||
dry_run=dry_run
|
||||
)
|
||||
|
||||
logger.info(
|
||||
f"Consolidation complete: {processed_count}/{len(unprocessed)} searches, "
|
||||
f"{total_pages_created} pages created, {total_pages_updated} updated, "
|
||||
f"{total_entities_added} entities added"
|
||||
)
|
||||
|
||||
return response
|
||||
|
||||
async def _find_unprocessed_searches(
|
||||
self,
|
||||
lookback_days: int,
|
||||
limit: int
|
||||
) -> List[Dict[str, Any]]:
|
||||
"""
|
||||
Find unprocessed SearchQuery nodes from Neo4j.
|
||||
"""
|
||||
lookback_date = datetime.now() - timedelta(days=lookback_days)
|
||||
|
||||
query = """
|
||||
MATCH (sq:SearchQuery {processed: false})
|
||||
WHERE sq.timestamp > datetime($lookback_date)
|
||||
RETURN sq.id as id,
|
||||
sq.query as query,
|
||||
sq.user as user,
|
||||
sq.timestamp as timestamp,
|
||||
sq.total_results as total_results,
|
||||
sq.web_count as web_count,
|
||||
sq.keywords as keywords
|
||||
ORDER BY sq.timestamp DESC
|
||||
LIMIT $limit
|
||||
"""
|
||||
|
||||
try:
|
||||
results = await self.neo4j.execute_query(
|
||||
query,
|
||||
{
|
||||
"lookback_date": lookback_date.isoformat(),
|
||||
"limit": limit
|
||||
}
|
||||
)
|
||||
|
||||
searches = []
|
||||
for record in results:
|
||||
searches.append({
|
||||
'id': record['id'],
|
||||
'query': record['query'],
|
||||
'user': record['user'],
|
||||
'timestamp': record['timestamp'],
|
||||
'total_results': record.get('total_results', 0),
|
||||
'web_count': record.get('web_count', 0),
|
||||
'keywords': record.get('keywords', [])
|
||||
})
|
||||
|
||||
return searches
|
||||
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to find unprocessed searches: {e}")
|
||||
return []
|
||||
|
||||
async def _process_search(
|
||||
self,
|
||||
search: Dict[str, Any],
|
||||
min_web_results: int,
|
||||
dry_run: bool
|
||||
) -> Optional[ConsolidationResult]:
|
||||
"""
|
||||
Process a single search query for knowledge consolidation.
|
||||
"""
|
||||
search_id = search['id']
|
||||
query = search['query']
|
||||
user = search['user']
|
||||
web_count = search.get('web_count', 0)
|
||||
|
||||
logger.info(f"Processing: '{query}' (user: {user}, web: {web_count})")
|
||||
|
||||
# Skip if insufficient web results
|
||||
if web_count < min_web_results:
|
||||
logger.info(f"Skipping - insufficient web results ({web_count} < {min_web_results})")
|
||||
return None
|
||||
|
||||
# Get web results from SearchQuery
|
||||
web_results = await self._get_web_results(search_id)
|
||||
if not web_results:
|
||||
logger.info("No web results found in database")
|
||||
return None
|
||||
|
||||
logger.info(f"Retrieved {len(web_results)} web results")
|
||||
|
||||
# Analyze web results with Ollama for novel information
|
||||
analysis = await self._analyze_web_results(
|
||||
query=query,
|
||||
web_results=web_results,
|
||||
keywords=search.get('keywords', []),
|
||||
user=user
|
||||
)
|
||||
|
||||
if not analysis or not analysis.get('has_novel_info'):
|
||||
logger.info("No novel information found")
|
||||
return ConsolidationResult(
|
||||
search_id=search_id,
|
||||
query=query
|
||||
)
|
||||
|
||||
# Extract consolidation actions
|
||||
pages_to_create = analysis.get('new_pages', [])
|
||||
pages_to_update = analysis.get('update_pages', [])
|
||||
new_entities = analysis.get('new_entities', [])
|
||||
|
||||
logger.info(
|
||||
f"Analysis: {len(pages_to_create)} new pages, "
|
||||
f"{len(pages_to_update)} updates, {len(new_entities)} entities"
|
||||
)
|
||||
|
||||
if dry_run:
|
||||
logger.info("[DRY RUN] Would create/update pages and entities")
|
||||
return ConsolidationResult(
|
||||
search_id=search_id,
|
||||
query=query,
|
||||
pages_created=len(pages_to_create),
|
||||
pages_updated=len(pages_to_update),
|
||||
entities_added=len(new_entities)
|
||||
)
|
||||
|
||||
# Create/update wiki pages
|
||||
pages_created = 0
|
||||
pages_updated = 0
|
||||
entities_added = 0
|
||||
|
||||
# Create new pages
|
||||
for page_data in pages_to_create:
|
||||
try:
|
||||
await self._create_or_consolidate_page(
|
||||
user=user,
|
||||
title=page_data.get('title'),
|
||||
path=page_data.get('path'),
|
||||
summary=page_data.get('summary'),
|
||||
source_query=query,
|
||||
web_results=web_results
|
||||
)
|
||||
pages_created += 1
|
||||
logger.info(f"Created page: {page_data.get('title')}")
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to create page {page_data.get('title')}: {e}")
|
||||
|
||||
# Update existing pages
|
||||
for page_data in pages_to_update:
|
||||
try:
|
||||
await self._update_page_with_facts(
|
||||
title=page_data.get('title'),
|
||||
new_facts=page_data.get('new_facts', []),
|
||||
source_url=page_data.get('source_url'),
|
||||
user=user
|
||||
)
|
||||
pages_updated += 1
|
||||
logger.info(f"Updated page: {page_data.get('title')}")
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to update page {page_data.get('title')}: {e}")
|
||||
|
||||
# Add new entities to graph
|
||||
for entity_data in new_entities:
|
||||
try:
|
||||
await self._add_entity_to_graph(
|
||||
user=user,
|
||||
entity_name=entity_data.get('name'),
|
||||
entity_type=entity_data.get('type'),
|
||||
description=entity_data.get('description'),
|
||||
source_search_id=search_id
|
||||
)
|
||||
entities_added += 1
|
||||
logger.info(f"Added entity: {entity_data.get('name')}")
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to add entity {entity_data.get('name')}: {e}")
|
||||
|
||||
return ConsolidationResult(
|
||||
search_id=search_id,
|
||||
query=query,
|
||||
pages_created=pages_created,
|
||||
pages_updated=pages_updated,
|
||||
entities_added=entities_added
|
||||
)
|
||||
|
||||
async def _get_web_results(self, search_id: str) -> List[Dict[str, Any]]:
|
||||
"""Get web results for a search from Neo4j."""
|
||||
query = """
|
||||
MATCH (sq:SearchQuery {id: $search_id})-[f:FOUND]->(wr:WebResult)
|
||||
RETURN wr.url as url,
|
||||
wr.title as title,
|
||||
wr.content as content,
|
||||
f.rank as rank,
|
||||
f.rrf_score as rrf_score
|
||||
ORDER BY f.rank
|
||||
LIMIT 20
|
||||
"""
|
||||
|
||||
try:
|
||||
results = await self.neo4j.execute_query(query, {"search_id": search_id})
|
||||
|
||||
web_results = []
|
||||
for record in results:
|
||||
web_results.append({
|
||||
'url': record['url'],
|
||||
'title': record['title'],
|
||||
'content': record['content'],
|
||||
'rank': record['rank'],
|
||||
'rrf_score': record['rrf_score']
|
||||
})
|
||||
|
||||
return web_results
|
||||
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to get web results: {e}")
|
||||
return []
|
||||
|
||||
async def _analyze_web_results(
|
||||
self,
|
||||
query: str,
|
||||
web_results: List[Dict[str, Any]],
|
||||
keywords: List[str],
|
||||
user: str = "jpmschweitzer"
|
||||
) -> Optional[Dict[str, Any]]:
|
||||
"""
|
||||
Analyze web results with Ollama for novel information.
|
||||
|
||||
Returns analysis with has_novel_info, new_pages, update_pages, new_entities.
|
||||
"""
|
||||
# Fetch existing taxonomy structure for this user
|
||||
try:
|
||||
taxonomy_structure = await self.wiki.get_taxonomy_structure(f"users/{user}")
|
||||
existing_paths_info = self._format_taxonomy_for_prompt(taxonomy_structure)
|
||||
logger.info(f"Fetched taxonomy with {len(taxonomy_structure)} categories for user {user}")
|
||||
except Exception as e:
|
||||
logger.warning(f"Failed to fetch taxonomy structure: {e}")
|
||||
existing_paths_info = ""
|
||||
|
||||
# Build analysis prompt
|
||||
web_summary = "\n\n".join([
|
||||
f"[{i+1}] {r['title']}\n{r['url']}\n{r['content'][:300]}..."
|
||||
for i, r in enumerate(web_results[:5])
|
||||
])
|
||||
|
||||
prompt = f"""You are a Librarian helping build a personal knowledge base and extended memory system.
|
||||
|
||||
Analyze these web search results for information worth documenting in our personal wiki.
|
||||
|
||||
Query: "{query}"
|
||||
Keywords: {', '.join(keywords) if keywords else 'none'}
|
||||
|
||||
Web Results:
|
||||
{web_summary}
|
||||
|
||||
This is a PERSONAL knowledge base using Schema.org-aligned taxonomy that captures:
|
||||
- People: Family members, friends, colleagues, public figures (Schema.org: Person)
|
||||
- Companies: Businesses, organizations, institutions (Schema.org: Organization)
|
||||
- Places: Locations, restaurants, travel destinations (Schema.org: Place)
|
||||
- Entertainment: Books, movies, TV, music, games (Schema.org: CreativeWork)
|
||||
- Recipes: Food, cooking techniques, ingredients (Schema.org: CreativeWork/Recipe)
|
||||
- Products: Purchased items, gear, tools, equipment (Schema.org: Product)
|
||||
- Technology: Software, applications, infrastructure (Schema.org: SoftwareApplication)
|
||||
- Health: Medical info, fitness, wellness (Schema.org: MedicalEntity)
|
||||
- Events: Concerts, travel, appointments, important dates (Schema.org: Event)
|
||||
- Hobbies: Personal interests, activities, pastimes (Custom extension)
|
||||
- Projects: Work projects, personal projects (Schema.org: Project)
|
||||
- Reference: General knowledge, how-tos (Custom extension)
|
||||
|
||||
Identify information worth documenting:
|
||||
1. New topics/people/things that deserve their own wiki page
|
||||
2. Facts that could enhance existing pages
|
||||
3. Entities (people, places, things, concepts) for the knowledge graph
|
||||
|
||||
Be INCLUSIVE - if someone searched for it, it's likely worth documenting.
|
||||
Personal information is just as valuable as technical information.
|
||||
|
||||
**CRITICAL: Use ONLY these Schema.org-aligned path prefixes (case-sensitive):**
|
||||
|
||||
- People: `people/<name>` (Schema.org: Person)
|
||||
- Companies: `companies/<company-name>` (Schema.org: Organization)
|
||||
- Places: `places/<location>` (Schema.org: Place)
|
||||
- Entertainment (Schema.org: CreativeWork):
|
||||
- Books: `entertainment/books/<title>`
|
||||
- Movies: `entertainment/movies/<title>`
|
||||
- TV: `entertainment/tv/<title>`
|
||||
- Music: `entertainment/music/<artist-or-album>`
|
||||
- Games: `entertainment/games/<title>`
|
||||
- Recipes: `recipes/<cuisine-or-category>/<dish>` (Schema.org: Recipe)
|
||||
- Products: `products/<category>/<product-name>` (Schema.org: Product)
|
||||
- Technology: `technology/<category>/<topic>` (Schema.org: SoftwareApplication)
|
||||
- Health: `health/<category>/<topic>` (Schema.org: MedicalEntity)
|
||||
- Events: `events/<event-type>/<event-name>` (Schema.org: Event)
|
||||
- Hobbies: `hobbies/<hobby-name>` (Custom extension)
|
||||
- Projects: `projects/<project-name>` (Schema.org: Project)
|
||||
- Reference: `reference/<category>/<topic>` (Custom extension)
|
||||
|
||||
**Path Rules:**
|
||||
- Use lowercase with hyphens (kebab-case): "machine-learning" not "Machine_Learning"
|
||||
- Keep paths 2-3 levels deep maximum
|
||||
- Be consistent with existing paths when possible
|
||||
|
||||
{existing_paths_info}
|
||||
|
||||
Return ONLY valid JSON:
|
||||
{{
|
||||
"has_novel_info": true,
|
||||
"new_pages": [
|
||||
{{"title": "Page Title", "path": "companies/example-company", "summary": "What information to include"}}
|
||||
],
|
||||
"update_pages": [
|
||||
{{"title": "Existing Page", "new_facts": ["fact 1"], "source_url": "url"}}
|
||||
],
|
||||
"new_entities": [
|
||||
{{"name": "Entity Name", "type": "person/place/thing/concept/recipe/media", "description": "Brief description"}}
|
||||
]
|
||||
}}
|
||||
|
||||
JSON:"""
|
||||
|
||||
try:
|
||||
# Call Ollama for analysis
|
||||
response = await self.ollama.generate_text(
|
||||
prompt=prompt,
|
||||
model=self.settings.reranker_model, # Use mistral-nemo
|
||||
stream=False
|
||||
)
|
||||
|
||||
if not response:
|
||||
logger.warning("Empty response from Ollama")
|
||||
return None
|
||||
|
||||
# Extract JSON from response
|
||||
response_clean = response.strip()
|
||||
if '{' in response_clean:
|
||||
json_start = response_clean.find('{')
|
||||
json_end = response_clean.rfind('}') + 1
|
||||
response_clean = response_clean[json_start:json_end]
|
||||
|
||||
analysis = json.loads(response_clean)
|
||||
return analysis
|
||||
|
||||
except json.JSONDecodeError as e:
|
||||
logger.error(f"Failed to parse Ollama response as JSON: {e}")
|
||||
logger.debug(f"Response was: {response[:500]}")
|
||||
return None
|
||||
except Exception as e:
|
||||
logger.error(f"Analysis failed: {e}", exc_info=True)
|
||||
return None
|
||||
|
||||
def _format_taxonomy_for_prompt(self, taxonomy: Dict[str, List[str]]) -> str:
|
||||
"""
|
||||
Format taxonomy structure for inclusion in LLM prompt.
|
||||
|
||||
Args:
|
||||
taxonomy: Dict mapping categories to subcategories
|
||||
|
||||
Returns:
|
||||
Formatted string showing existing paths
|
||||
"""
|
||||
if not taxonomy:
|
||||
return ""
|
||||
|
||||
lines = ["**Existing paths in your wiki (PREFER these over creating new ones):**"]
|
||||
for category, subcategories in taxonomy.items():
|
||||
if subcategories:
|
||||
lines.append(f"- {category}/")
|
||||
for sub in subcategories:
|
||||
lines.append(f" - {category}/{sub}/")
|
||||
else:
|
||||
lines.append(f"- {category}/")
|
||||
|
||||
lines.append("")
|
||||
lines.append("**IMPORTANT:** If a suitable existing path exists, use it instead of creating a new category.")
|
||||
lines.append("Example: NATO should go in `reference/political-entities/` not a new `reference/military-alliances/`")
|
||||
|
||||
return "\n".join(lines)
|
||||
|
||||
async def _mark_search_processed(self, search_id: str):
|
||||
"""Mark SearchQuery node as processed."""
|
||||
query = """
|
||||
MATCH (sq:SearchQuery {id: $search_id})
|
||||
SET sq.processed = true,
|
||||
sq.processed_at = datetime()
|
||||
RETURN sq.id
|
||||
"""
|
||||
|
||||
try:
|
||||
await self.neo4j.execute_query(query, {"search_id": search_id})
|
||||
logger.debug(f"Marked search {search_id} as processed")
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to mark search as processed: {e}")
|
||||
|
||||
async def _apply_bidirectional_entity_linking(
|
||||
self,
|
||||
page_id: int,
|
||||
page_title: str,
|
||||
user: str
|
||||
) -> Dict[str, int]:
|
||||
"""
|
||||
Apply bidirectional entity linking after page creation/update.
|
||||
|
||||
This runs AFTER ingestion so entities are extracted and in the graph.
|
||||
|
||||
Steps:
|
||||
1. Link entities in the new page (forward links to existing entities)
|
||||
2. Find pages that mention the new entity (reverse references)
|
||||
3. Link entities in those pages (backward links to the new entity)
|
||||
|
||||
Args:
|
||||
page_id: Wiki page ID
|
||||
page_title: Page title (used to find reverse references)
|
||||
user: User identifier
|
||||
|
||||
Returns:
|
||||
Dict with link counts: {
|
||||
"forward_links": int, # Links added to the new page
|
||||
"backward_links": int, # Links added to other pages pointing to new page
|
||||
"pages_updated": int # Number of other pages updated
|
||||
}
|
||||
"""
|
||||
from src.core.multi_tenancy import get_neo4j_user_base_label
|
||||
|
||||
forward_links = 0
|
||||
backward_links = 0
|
||||
pages_updated = 0
|
||||
|
||||
try:
|
||||
# Import here to avoid circular dependency
|
||||
from src.routers.entity_linking import link_entities_in_page, EntityLinkingRequest
|
||||
from src.core.dependencies import get_wiki_service, get_graph_service
|
||||
|
||||
wiki_service = get_wiki_service()
|
||||
graph_service = get_graph_service()
|
||||
|
||||
# STEP 1: Forward linking - link entities in the new page
|
||||
logger.info(f"Step 1/3: Linking entities in page {page_id} ('{page_title}')")
|
||||
try:
|
||||
forward_result = await link_entities_in_page(
|
||||
request=EntityLinkingRequest(
|
||||
user=user,
|
||||
page_id=page_id,
|
||||
create_relationships=True,
|
||||
re_index_if_changed=False # Already indexed, no need to re-index
|
||||
),
|
||||
wiki_service=wiki_service,
|
||||
graph_service=graph_service,
|
||||
ingestion_service=self.ingestion_service,
|
||||
api_key="" # Internal call, no auth needed
|
||||
)
|
||||
forward_links = forward_result.content_links_added
|
||||
logger.info(f"Added {forward_links} forward links in page {page_id}")
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to add forward links: {e}")
|
||||
|
||||
# STEP 2: Find reverse references - which pages mention this new entity?
|
||||
logger.info(f"Step 2/3: Finding pages that mention '{page_title}'")
|
||||
user_base_label = get_neo4j_user_base_label(user)
|
||||
|
||||
# Query to find documents that mention entities with this page's title
|
||||
reverse_query = f"""
|
||||
// Find entities with the same name as the page title
|
||||
MATCH (e:{user_base_label})
|
||||
WHERE toLower(e.name) = toLower($title)
|
||||
AND NOT e:Document
|
||||
|
||||
// Find documents that mention those entities
|
||||
MATCH (d:Document)-[r:MENTIONS]->(e)
|
||||
WHERE d.page_id <> $page_id // Exclude the page itself
|
||||
|
||||
RETURN DISTINCT d.page_id as page_id, d.title as title
|
||||
LIMIT 50
|
||||
"""
|
||||
|
||||
try:
|
||||
reverse_refs = await self.neo4j.execute_query(
|
||||
reverse_query,
|
||||
{"title": page_title, "page_id": page_id}
|
||||
)
|
||||
logger.info(f"Found {len(reverse_refs)} pages that mention '{page_title}'")
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to find reverse references: {e}")
|
||||
reverse_refs = []
|
||||
|
||||
# STEP 3: Backward linking - add links in those pages to the new entity
|
||||
if reverse_refs:
|
||||
logger.info(f"Step 3/3: Adding backward links in {len(reverse_refs)} pages")
|
||||
for ref in reverse_refs:
|
||||
try:
|
||||
backward_result = await link_entities_in_page(
|
||||
request=EntityLinkingRequest(
|
||||
user=user,
|
||||
page_id=ref['page_id'],
|
||||
create_relationships=False, # Relationships already exist
|
||||
re_index_if_changed=False # Don't re-index for link updates
|
||||
),
|
||||
wiki_service=wiki_service,
|
||||
graph_service=graph_service,
|
||||
ingestion_service=self.ingestion_service,
|
||||
api_key=""
|
||||
)
|
||||
if backward_result.content_links_added > 0:
|
||||
backward_links += backward_result.content_links_added
|
||||
pages_updated += 1
|
||||
logger.info(
|
||||
f"Added {backward_result.content_links_added} links "
|
||||
f"in page {ref['page_id']} ('{ref['title']}')"
|
||||
)
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to add backward links in page {ref['page_id']}: {e}")
|
||||
else:
|
||||
logger.info("Step 3/3: No reverse references found, skipping backward linking")
|
||||
|
||||
return {
|
||||
"forward_links": forward_links,
|
||||
"backward_links": backward_links,
|
||||
"pages_updated": pages_updated
|
||||
}
|
||||
|
||||
except Exception as e:
|
||||
logger.error(f"Bidirectional entity linking failed: {e}", exc_info=True)
|
||||
return {
|
||||
"forward_links": 0,
|
||||
"backward_links": 0,
|
||||
"pages_updated": 0
|
||||
}
|
||||
|
||||
async def _create_or_consolidate_page(
|
||||
self,
|
||||
user: str,
|
||||
title: str,
|
||||
path: str,
|
||||
summary: str,
|
||||
source_query: str,
|
||||
web_results: List[Dict[str, Any]]
|
||||
):
|
||||
"""
|
||||
Create wiki page or consolidate with existing synonym page.
|
||||
|
||||
Uses WikiPageWriter for intelligent LLM-based content generation:
|
||||
- For new pages: Holistic structured content creation
|
||||
- For existing pages: Zero-loss reconstruction with conflict detection
|
||||
"""
|
||||
# Normalize path to user namespace
|
||||
if not path.startswith(f"users/{user}"):
|
||||
path = f"users/{user}/{path.lstrip('/')}"
|
||||
|
||||
# Format web results as source information
|
||||
source_information = [
|
||||
{
|
||||
'title': r['title'],
|
||||
'url': r['url'],
|
||||
'content': r['content']
|
||||
}
|
||||
for r in web_results[:5] # Top 5 web results
|
||||
]
|
||||
|
||||
# Search for existing pages with similar titles (synonym consolidation)
|
||||
existing_pages = await self.wiki.search_pages(title, path_prefix=f"users/{user}")
|
||||
|
||||
if existing_pages:
|
||||
# Page exists - reconstruct with new information using LLM
|
||||
logger.info(f"Found existing page for '{title}', will reconstruct with new info")
|
||||
page_id = existing_pages[0]['id']
|
||||
|
||||
# Get current content
|
||||
existing_page = await self.wiki.get_page(page_id)
|
||||
if existing_page:
|
||||
# Build new information text from summary and web results
|
||||
new_information = f"{summary}\n\n"
|
||||
for r in web_results[:3]:
|
||||
new_information += f"- {r['title']}: {r['content'][:200]}...\n"
|
||||
|
||||
# Use WikiPageWriter to reconstruct with LLM
|
||||
reconstructed_content, conflicts = await self.wiki_page_writer.reconstruct_page(
|
||||
title=title,
|
||||
existing_content=existing_page['content'],
|
||||
new_information=new_information,
|
||||
new_sources=source_information,
|
||||
detect_conflicts=True
|
||||
)
|
||||
|
||||
if conflicts:
|
||||
logger.warning(
|
||||
f"Detected {len(conflicts)} conflicts when updating '{title}' - "
|
||||
"LLM chose most authoritative sources"
|
||||
)
|
||||
|
||||
await self.wiki.update_page(
|
||||
page_id=page_id,
|
||||
content=reconstructed_content
|
||||
)
|
||||
logger.info(f"Reconstructed existing page: {title}")
|
||||
|
||||
# Trigger ingestion to update vectors and graph
|
||||
if self.ingestion_service:
|
||||
try:
|
||||
await self.ingestion_service.ingest_page(
|
||||
page_id=page_id,
|
||||
user=user,
|
||||
force_refresh=True
|
||||
)
|
||||
logger.info(f"Ingested updated page {page_id} into knowledge base")
|
||||
|
||||
# Apply bidirectional entity linking after ingestion
|
||||
link_stats = await self._apply_bidirectional_entity_linking(
|
||||
page_id=page_id,
|
||||
page_title=title,
|
||||
user=user
|
||||
)
|
||||
logger.info(
|
||||
f"Entity linking complete: {link_stats['forward_links']} forward links, "
|
||||
f"{link_stats['backward_links']} backward links "
|
||||
f"({link_stats['pages_updated']} pages updated)"
|
||||
)
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to ingest updated page {page_id}: {e}")
|
||||
|
||||
return
|
||||
|
||||
# Create new page with LLM-generated structured content
|
||||
logger.info(f"Creating new page: {title}")
|
||||
|
||||
# Use WikiPageWriter to create structured content
|
||||
content = await self.wiki_page_writer.create_page(
|
||||
title=title,
|
||||
topic_summary=summary,
|
||||
source_information=source_information,
|
||||
entities=None, # Could extract from keywords if available
|
||||
related_docs=None
|
||||
)
|
||||
|
||||
# Extract tags from path for dossier organization
|
||||
path_parts = path.split('/')
|
||||
tags = [part for part in path_parts if part and part not in ['users', user]]
|
||||
|
||||
created_page = await self.wiki.create_page(
|
||||
path=path,
|
||||
title=title,
|
||||
content=content,
|
||||
description=f"Consolidated from search: {source_query}",
|
||||
tags=tags[:3], # Limit to 3 tags
|
||||
is_published=True
|
||||
)
|
||||
|
||||
page_id = created_page.get("id") if created_page else None
|
||||
logger.info(f"Created new page: {path} (page_id: {page_id})")
|
||||
|
||||
# Trigger ingestion to update vectors and graph
|
||||
if self.ingestion_service and page_id:
|
||||
try:
|
||||
await self.ingestion_service.ingest_page(
|
||||
page_id=page_id,
|
||||
user=user,
|
||||
force_refresh=False # New page, no need to force
|
||||
)
|
||||
logger.info(f"Ingested new page {page_id} into knowledge base")
|
||||
|
||||
# Apply bidirectional entity linking after ingestion
|
||||
link_stats = await self._apply_bidirectional_entity_linking(
|
||||
page_id=page_id,
|
||||
page_title=title,
|
||||
user=user
|
||||
)
|
||||
logger.info(
|
||||
f"Entity linking complete: {link_stats['forward_links']} forward links, "
|
||||
f"{link_stats['backward_links']} backward links "
|
||||
f"({link_stats['pages_updated']} pages updated)"
|
||||
)
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to ingest new page {page_id}: {e}")
|
||||
|
||||
async def _update_page_with_facts(
|
||||
self,
|
||||
title: str,
|
||||
new_facts: List[str],
|
||||
source_url: str,
|
||||
user: str
|
||||
):
|
||||
"""
|
||||
Update existing page with new facts using LLM reconstruction.
|
||||
|
||||
Uses WikiPageWriter to intelligently merge facts with zero loss.
|
||||
"""
|
||||
# Search for page
|
||||
pages = await self.wiki.search_pages(title, path_prefix=f"users/{user}")
|
||||
|
||||
if not pages:
|
||||
logger.warning(f"Page '{title}' not found for update")
|
||||
return
|
||||
|
||||
page_id = pages[0]['id']
|
||||
existing_page = await self.wiki.get_page(page_id)
|
||||
|
||||
if not existing_page:
|
||||
return
|
||||
|
||||
# Build new information from facts
|
||||
new_information = "\n".join([f"- {fact}" for fact in new_facts])
|
||||
|
||||
# Format source
|
||||
source_information = [{
|
||||
'title': source_url,
|
||||
'url': source_url,
|
||||
'content': new_information
|
||||
}]
|
||||
|
||||
# Use WikiPageWriter to reconstruct with LLM
|
||||
reconstructed_content, conflicts = await self.wiki_page_writer.reconstruct_page(
|
||||
title=title,
|
||||
existing_content=existing_page['content'],
|
||||
new_information=new_information,
|
||||
new_sources=source_information,
|
||||
detect_conflicts=True
|
||||
)
|
||||
|
||||
if conflicts:
|
||||
logger.warning(
|
||||
f"Detected {len(conflicts)} conflicts when updating '{title}' with new facts"
|
||||
)
|
||||
|
||||
await self.wiki.update_page(
|
||||
page_id=page_id,
|
||||
content=reconstructed_content
|
||||
)
|
||||
|
||||
# Trigger ingestion to update vectors and graph
|
||||
logger.debug(f"ingestion_service available: {self.ingestion_service is not None}")
|
||||
if self.ingestion_service:
|
||||
try:
|
||||
logger.info(f"Starting ingestion for updated page {page_id}")
|
||||
await self.ingestion_service.ingest_page(
|
||||
page_id=page_id,
|
||||
user=user,
|
||||
force_refresh=True
|
||||
)
|
||||
logger.info(f"Ingested updated page {page_id} into knowledge base")
|
||||
|
||||
# Apply bidirectional entity linking after ingestion
|
||||
link_stats = await self._apply_bidirectional_entity_linking(
|
||||
page_id=page_id,
|
||||
page_title=title,
|
||||
user=user
|
||||
)
|
||||
logger.info(
|
||||
f"Entity linking complete: {link_stats['forward_links']} forward links, "
|
||||
f"{link_stats['backward_links']} backward links "
|
||||
f"({link_stats['pages_updated']} pages updated)"
|
||||
)
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to ingest updated page {page_id}: {e}")
|
||||
|
||||
async def _add_entity_to_graph(
|
||||
self,
|
||||
user: str,
|
||||
entity_name: str,
|
||||
entity_type: str,
|
||||
description: str,
|
||||
source_search_id: str
|
||||
):
|
||||
"""Add new entity to knowledge graph."""
|
||||
from src.core.multi_tenancy import get_neo4j_user_base_label
|
||||
|
||||
user_base_label = get_neo4j_user_base_label(user)
|
||||
|
||||
# Create entity node with appropriate type label
|
||||
type_label = entity_type.capitalize() if entity_type else "Entity"
|
||||
|
||||
query = f"""
|
||||
MERGE (e:{user_base_label}:{type_label} {{name: $name}})
|
||||
ON CREATE SET
|
||||
e.description = $description,
|
||||
e.created_at = datetime(),
|
||||
e.source = 'librarian_consolidation',
|
||||
e.source_search_id = $search_id
|
||||
ON MATCH SET
|
||||
e.updated_at = datetime()
|
||||
RETURN e
|
||||
"""
|
||||
|
||||
try:
|
||||
await self.neo4j.execute_query(query, {
|
||||
"name": entity_name,
|
||||
"description": description,
|
||||
"search_id": source_search_id
|
||||
})
|
||||
logger.debug(f"Added entity to graph: {entity_name} ({entity_type})")
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to add entity to graph: {e}")
|
||||
File diff suppressed because it is too large
Load Diff
@@ -1,739 +0,0 @@
|
||||
"""
|
||||
HybridRAG service combining vector, graph, and web search.
|
||||
|
||||
6-Phase Pipeline:
|
||||
0. Query Enhancement - Extract keywords/synonyms with LLM
|
||||
1. Parallel Retrieval - Vector + Graph + Web search
|
||||
2. RRF Fusion - Merge results with Reciprocal Rank Fusion
|
||||
3. Enrichment - Add related dossiers via graph
|
||||
4. LLM Re-ranking - Re-rank with mistral-nemo
|
||||
5. Context Formatting - Format for LLM consumption
|
||||
6. Persistence - Store for Librarian processing
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
import time
|
||||
import json
|
||||
import uuid
|
||||
from typing import List, Dict, Any, Optional
|
||||
import logging
|
||||
|
||||
from src.services.vector_service import VectorService
|
||||
from src.services.graph_service import GraphService
|
||||
from src.clients.searxng_client import SearXNGClient
|
||||
from src.clients.ollama_client import OllamaClient
|
||||
from src.config import Settings
|
||||
from src.models.hybrid_rag import (
|
||||
HybridRAGConfig, HybridRAGRequest, HybridRAGResponse,
|
||||
HybridRAGResult, TimingBreakdown, KeywordExtraction,
|
||||
RelatedDossier
|
||||
)
|
||||
from src.core.multi_tenancy import get_neo4j_user_base_label, get_neo4j_user_label
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
|
||||
class HybridRAGService:
|
||||
"""
|
||||
Service for HybridRAG multi-source search with fusion and re-ranking.
|
||||
"""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
vector_service: VectorService,
|
||||
graph_service: GraphService,
|
||||
searxng_client: SearXNGClient,
|
||||
ollama_client: OllamaClient,
|
||||
settings: Settings
|
||||
):
|
||||
"""
|
||||
Initialize HybridRAG service.
|
||||
|
||||
Args:
|
||||
vector_service: Service for Qdrant vector search
|
||||
graph_service: Service for Neo4j graph search
|
||||
searxng_client: Client for web search
|
||||
ollama_client: Client for LLM (keyword extraction, re-ranking)
|
||||
settings: Application settings
|
||||
"""
|
||||
self.vector = vector_service
|
||||
self.graph = graph_service
|
||||
self.searxng = searxng_client
|
||||
self.ollama = ollama_client
|
||||
self.settings = settings
|
||||
self.reranker_model = settings.reranker_model
|
||||
|
||||
async def search(
|
||||
self,
|
||||
query: str,
|
||||
user: str,
|
||||
config: Optional[HybridRAGConfig] = None
|
||||
) -> HybridRAGResponse:
|
||||
"""
|
||||
Execute HybridRAG search across all sources.
|
||||
|
||||
Args:
|
||||
query: Search query
|
||||
user: User identifier
|
||||
config: Optional configuration override
|
||||
|
||||
Returns:
|
||||
Complete search response with ranked results and timing
|
||||
"""
|
||||
start_time = time.time()
|
||||
timing = {}
|
||||
|
||||
# Use default config if not provided
|
||||
if not config:
|
||||
config = HybridRAGConfig()
|
||||
|
||||
logger.info(f"HybridRAG search: '{query}' for user '{user}'")
|
||||
|
||||
# Phase 0: Query Enhancement
|
||||
phase0_start = time.time()
|
||||
keywords_data = await self._extract_keywords_and_synonyms(query)
|
||||
timing["query_enhancement_ms"] = (time.time() - phase0_start) * 1000
|
||||
|
||||
# Phase 1: Parallel Retrieval
|
||||
phase1_start = time.time()
|
||||
raw_results = await self._retrieve_parallel(query, user, config, keywords_data)
|
||||
timing["vector_ms"] = raw_results.get("timing", {}).get("vector_ms", 0)
|
||||
timing["graph_ms"] = raw_results.get("timing", {}).get("graph_ms", 0)
|
||||
timing["web_ms"] = raw_results.get("timing", {}).get("web_ms", 0)
|
||||
|
||||
# Phase 2: RRF Fusion
|
||||
phase2_start = time.time()
|
||||
fused_results = self._reciprocal_rank_fusion(
|
||||
results_by_source={
|
||||
"vector": raw_results.get("vector", []),
|
||||
"graph": raw_results.get("graph", []),
|
||||
"web": raw_results.get("web", [])
|
||||
},
|
||||
k=config.rrf_k
|
||||
)
|
||||
timing["fusion_ms"] = (time.time() - phase2_start) * 1000
|
||||
|
||||
# Phase 3: Enrichment
|
||||
phase3_start = time.time()
|
||||
if config.enable_enrichment:
|
||||
enriched_results = await self._enrich_with_related_dossiers(fused_results, user)
|
||||
else:
|
||||
enriched_results = fused_results
|
||||
timing["enrichment_ms"] = (time.time() - phase3_start) * 1000
|
||||
|
||||
# Phase 4: LLM Re-ranking
|
||||
phase4_start = time.time()
|
||||
if config.enable_reranking and len(enriched_results) > 1:
|
||||
reranked_results = await self._rerank_with_llm(enriched_results[:20], query)
|
||||
else:
|
||||
reranked_results = enriched_results
|
||||
timing["reranking_ms"] = (time.time() - phase4_start) * 1000
|
||||
|
||||
# Limit to final result count
|
||||
final_results = reranked_results[:config.final_result_count]
|
||||
|
||||
# Update final ranks
|
||||
for i, result in enumerate(final_results, start=1):
|
||||
result["final_rank"] = i
|
||||
|
||||
# Convert to HybridRAGResult models
|
||||
result_models = self._convert_to_result_models(final_results)
|
||||
|
||||
# Phase 5: Context Formatting
|
||||
context = self._format_context_for_llm(result_models)
|
||||
|
||||
# Calculate source counts
|
||||
source_counts = {}
|
||||
for result in result_models:
|
||||
for source in result.sources:
|
||||
source_counts[source] = source_counts.get(source, 0) + 1
|
||||
|
||||
timing["total_ms"] = (time.time() - start_time) * 1000
|
||||
|
||||
# Phase 6: Persistence (async, non-blocking)
|
||||
phase6_start = time.time()
|
||||
search_id = await self._persist_search_for_librarian(
|
||||
query=query,
|
||||
user=user,
|
||||
keywords_data=keywords_data,
|
||||
raw_results=raw_results,
|
||||
final_results=final_results,
|
||||
timing=timing
|
||||
)
|
||||
timing["persistence_ms"] = (time.time() - phase6_start) * 1000
|
||||
|
||||
# Build response
|
||||
return HybridRAGResponse(
|
||||
query=query,
|
||||
keywords=KeywordExtraction(**keywords_data),
|
||||
results=result_models,
|
||||
context=context,
|
||||
source_counts=source_counts,
|
||||
total_results=len(result_models),
|
||||
timing=TimingBreakdown(**timing),
|
||||
config_used=config,
|
||||
search_id=search_id
|
||||
)
|
||||
|
||||
async def _extract_keywords_and_synonyms(self, query: str) -> Dict[str, Any]:
|
||||
"""
|
||||
Phase 0: Extract keywords, entities, and synonyms using LLM.
|
||||
|
||||
Args:
|
||||
query: Search query
|
||||
|
||||
Returns:
|
||||
Dictionary with keywords, entities, synonyms, expansions
|
||||
"""
|
||||
prompt = f"""Extract search terms from this query. For each important word, provide synonyms and expansions.
|
||||
|
||||
Query: "{query}"
|
||||
|
||||
Return ONLY valid JSON:
|
||||
{{
|
||||
"core_keywords": ["key", "words", "from", "query"],
|
||||
"synonyms": {{
|
||||
"word": ["alternative", "terms"]
|
||||
}}
|
||||
}}
|
||||
|
||||
Example for "Docker container hosting":
|
||||
{{
|
||||
"core_keywords": ["docker", "container", "hosting"],
|
||||
"synonyms": {{
|
||||
"docker": ["containerization", "container runtime"],
|
||||
"hosting": ["server", "infrastructure"]
|
||||
}}
|
||||
}}
|
||||
|
||||
JSON:"""
|
||||
|
||||
try:
|
||||
response = await self.ollama.generate_text(
|
||||
prompt=prompt,
|
||||
model=self.reranker_model
|
||||
)
|
||||
|
||||
# Parse JSON response (handle potential extra text)
|
||||
response_clean = response.strip()
|
||||
# Try to extract JSON if wrapped in text
|
||||
if '{' in response_clean:
|
||||
json_start = response_clean.find('{')
|
||||
json_end = response_clean.rfind('}') + 1
|
||||
response_clean = response_clean[json_start:json_end]
|
||||
|
||||
keywords_data = json.loads(response_clean)
|
||||
|
||||
# Ensure all required fields exist
|
||||
result = {
|
||||
"core_keywords": keywords_data.get("core_keywords", []),
|
||||
"entities": keywords_data.get("entities", []),
|
||||
"synonyms": keywords_data.get("synonyms", {}),
|
||||
"expansions": keywords_data.get("expansions", {})
|
||||
}
|
||||
|
||||
logger.info(f"Extracted keywords: {result['core_keywords'][:5]}, synonyms: {len(result['synonyms'])} terms")
|
||||
return result
|
||||
|
||||
except json.JSONDecodeError as e:
|
||||
logger.warning(f"Failed to parse LLM keyword extraction: {e}, using fallback")
|
||||
# Fallback to simple extraction
|
||||
words = query.split()
|
||||
return {
|
||||
"core_keywords": words,
|
||||
"entities": [],
|
||||
"synonyms": {},
|
||||
"expansions": {}
|
||||
}
|
||||
except Exception as e:
|
||||
logger.error(f"Keyword extraction failed: {e}", exc_info=True)
|
||||
return {
|
||||
"core_keywords": query.split(),
|
||||
"entities": [],
|
||||
"synonyms": {},
|
||||
"expansions": {}
|
||||
}
|
||||
|
||||
async def _retrieve_parallel(
|
||||
self,
|
||||
query: str,
|
||||
user: str,
|
||||
config: HybridRAGConfig,
|
||||
keywords_data: Dict[str, Any]
|
||||
) -> Dict[str, List]:
|
||||
"""
|
||||
Phase 1: Retrieve results from all sources in parallel.
|
||||
|
||||
Args:
|
||||
query: Search query
|
||||
user: User identifier
|
||||
config: Search configuration
|
||||
keywords_data: Extracted keywords/synonyms
|
||||
|
||||
Returns:
|
||||
Dictionary with results from each source and timing
|
||||
"""
|
||||
tasks = {}
|
||||
timing = {}
|
||||
|
||||
# Vector search
|
||||
if config.enable_vector:
|
||||
async def vector_search():
|
||||
start = time.time()
|
||||
try:
|
||||
response = await self.vector.search(
|
||||
query=query,
|
||||
user=user,
|
||||
limit=config.vector_limit
|
||||
)
|
||||
results = [
|
||||
{
|
||||
"page_id": r.page_id,
|
||||
"title": r.page_title,
|
||||
"content": r.content,
|
||||
"path": r.page_path,
|
||||
"score": r.score,
|
||||
"source": "vector"
|
||||
}
|
||||
for r in response.results
|
||||
]
|
||||
return results, (time.time() - start) * 1000
|
||||
except Exception as e:
|
||||
logger.error(f"Vector search failed: {e}", exc_info=True)
|
||||
return [], (time.time() - start) * 1000
|
||||
|
||||
tasks["vector"] = vector_search()
|
||||
|
||||
# Graph search
|
||||
if config.enable_graph:
|
||||
async def graph_search():
|
||||
start = time.time()
|
||||
try:
|
||||
results = await self.graph.search_documents(
|
||||
query=query,
|
||||
user=user,
|
||||
limit=config.graph_limit,
|
||||
keywords_data=keywords_data
|
||||
)
|
||||
formatted = [
|
||||
{
|
||||
"page_id": r["page_id"],
|
||||
"title": r["title"],
|
||||
"content": "", # Graph doesn't return content
|
||||
"path": r["path"],
|
||||
"entity_matches": r.get("entity_matches", 0),
|
||||
"matched_entities": r.get("matched_entities", []),
|
||||
"source": "graph"
|
||||
}
|
||||
for r in results
|
||||
]
|
||||
return formatted, (time.time() - start) * 1000
|
||||
except Exception as e:
|
||||
logger.error(f"Graph search failed: {e}", exc_info=True)
|
||||
return [], (time.time() - start) * 1000
|
||||
|
||||
tasks["graph"] = graph_search()
|
||||
|
||||
# Web search
|
||||
if config.enable_web:
|
||||
async def web_search():
|
||||
start = time.time()
|
||||
try:
|
||||
results = await self.searxng.search_general(
|
||||
query=query,
|
||||
limit=config.web_limit
|
||||
)
|
||||
formatted = [
|
||||
{
|
||||
"url": r.get("url"),
|
||||
"title": r.get("title", ""),
|
||||
"content": r.get("content", ""),
|
||||
"engine": r.get("engine", ""),
|
||||
"source": "web"
|
||||
}
|
||||
for r in results
|
||||
]
|
||||
return formatted, (time.time() - start) * 1000
|
||||
except Exception as e:
|
||||
logger.error(f"Web search failed: {e}", exc_info=True)
|
||||
return [], (time.time() - start) * 1000
|
||||
|
||||
tasks["web"] = web_search()
|
||||
|
||||
# Execute all searches in parallel
|
||||
results_dict = await asyncio.gather(*tasks.values())
|
||||
|
||||
# Combine results with timing
|
||||
output = {"timing": {}}
|
||||
for i, source in enumerate(tasks.keys()):
|
||||
results, source_timing = results_dict[i]
|
||||
output[source] = results
|
||||
output["timing"][f"{source}_ms"] = source_timing
|
||||
|
||||
logger.info(
|
||||
f"Parallel retrieval: vector={len(output.get('vector', []))}, "
|
||||
f"graph={len(output.get('graph', []))}, web={len(output.get('web', []))}"
|
||||
)
|
||||
|
||||
return output
|
||||
|
||||
def _reciprocal_rank_fusion(
|
||||
self,
|
||||
results_by_source: Dict[str, List],
|
||||
k: int = 60
|
||||
) -> List[Dict[str, Any]]:
|
||||
"""
|
||||
Phase 2: Merge results using Reciprocal Rank Fusion.
|
||||
|
||||
RRF formula: score = sum(1 / (k + rank)) for each source
|
||||
|
||||
Args:
|
||||
results_by_source: Results from each source
|
||||
k: RRF constant (default 60)
|
||||
|
||||
Returns:
|
||||
Merged and sorted results
|
||||
"""
|
||||
rrf_scores = {}
|
||||
|
||||
for source, results in results_by_source.items():
|
||||
for rank, result in enumerate(results, start=1):
|
||||
# Use page_id for wiki results, url hash for web results
|
||||
if result.get("page_id"):
|
||||
result_id = f"page_{result['page_id']}"
|
||||
elif result.get("url"):
|
||||
result_id = f"url_{hash(result['url'])}"
|
||||
else:
|
||||
continue # Skip results without ID
|
||||
|
||||
if result_id not in rrf_scores:
|
||||
rrf_scores[result_id] = {
|
||||
"result": result,
|
||||
"rrf_score": 0.0,
|
||||
"sources": [],
|
||||
"source_type": source
|
||||
}
|
||||
|
||||
# RRF formula: sum of 1/(k + rank) across sources
|
||||
rrf_scores[result_id]["rrf_score"] += 1 / (k + rank)
|
||||
rrf_scores[result_id]["sources"].append(source)
|
||||
|
||||
# If result appears in multiple sources, update source_type
|
||||
if len(rrf_scores[result_id]["sources"]) > 1:
|
||||
rrf_scores[result_id]["source_type"] = "+".join(
|
||||
sorted(set(rrf_scores[result_id]["sources"]))
|
||||
)
|
||||
|
||||
# Sort by RRF score descending
|
||||
sorted_results = sorted(
|
||||
rrf_scores.values(),
|
||||
key=lambda x: x["rrf_score"],
|
||||
reverse=True
|
||||
)
|
||||
|
||||
logger.info(f"RRF fusion: {len(sorted_results)} unique results from {len(results_by_source)} sources")
|
||||
|
||||
return sorted_results
|
||||
|
||||
async def _enrich_with_related_dossiers(
|
||||
self,
|
||||
results: List[Dict[str, Any]],
|
||||
user: str
|
||||
) -> List[Dict[str, Any]]:
|
||||
"""
|
||||
Phase 3: Enrich results with related documents via shared entities.
|
||||
|
||||
Args:
|
||||
results: Fused results
|
||||
user: User identifier
|
||||
|
||||
Returns:
|
||||
Results with related_dossiers added
|
||||
"""
|
||||
for result in results:
|
||||
result_data = result.get("result", {})
|
||||
page_id = result_data.get("page_id")
|
||||
|
||||
if page_id:
|
||||
try:
|
||||
related_docs = await self.graph.get_related_documents(
|
||||
page_id=page_id,
|
||||
user=user,
|
||||
limit=5
|
||||
)
|
||||
|
||||
# Convert to RelatedDossier format
|
||||
related_dossiers = []
|
||||
for doc in related_docs:
|
||||
for tag in doc.get("tags", [])[:3]: # Max 3 tags per doc
|
||||
related_dossiers.append({
|
||||
"page_id": doc["page_id"],
|
||||
"title": doc["title"],
|
||||
"path": doc["path"],
|
||||
"tag": tag,
|
||||
"shared_entities": doc["shared_entities"]
|
||||
})
|
||||
|
||||
result["related_dossiers"] = related_dossiers[:5] # Limit to 5 total
|
||||
|
||||
except Exception as e:
|
||||
logger.warning(f"Failed to get related docs for page {page_id}: {e}")
|
||||
result["related_dossiers"] = []
|
||||
else:
|
||||
result["related_dossiers"] = []
|
||||
|
||||
return results
|
||||
|
||||
async def _rerank_with_llm(
|
||||
self,
|
||||
results: List[Dict[str, Any]],
|
||||
query: str
|
||||
) -> List[Dict[str, Any]]:
|
||||
"""
|
||||
Phase 4: Re-rank results using LLM for better relevance.
|
||||
|
||||
Args:
|
||||
results: Results to re-rank (top 20)
|
||||
query: Original search query
|
||||
|
||||
Returns:
|
||||
Re-ranked results
|
||||
"""
|
||||
if len(results) <= 1:
|
||||
return results
|
||||
|
||||
try:
|
||||
# Build prompt with numbered results
|
||||
docs_text = "\n".join([
|
||||
f"{i+1}. {r['result'].get('title', 'Untitled')} - {r['result'].get('content', '')[:200]}..."
|
||||
for i, r in enumerate(results)
|
||||
])
|
||||
|
||||
prompt = f"""Given this search query and documents, rank them by relevance.
|
||||
|
||||
Query: {query}
|
||||
|
||||
Documents:
|
||||
{docs_text}
|
||||
|
||||
Return only the numbers in order of relevance (most relevant first).
|
||||
Example: 3,1,5,2,4
|
||||
|
||||
Ranking:"""
|
||||
|
||||
response = await self.ollama.generate_text(
|
||||
prompt=prompt,
|
||||
model=self.reranker_model
|
||||
)
|
||||
|
||||
# Parse response: "3,1,5,2,4" → [2, 0, 4, 1, 3] (0-indexed)
|
||||
indices_str = response.strip().split('\n')[0] # Take first line
|
||||
indices = [int(x.strip()) - 1 for x in indices_str.split(",") if x.strip().isdigit()]
|
||||
|
||||
# Reorder results according to LLM ranking
|
||||
reranked = []
|
||||
for idx in indices:
|
||||
if 0 <= idx < len(results):
|
||||
reranked.append(results[idx])
|
||||
|
||||
# Add any results that weren't in the LLM response
|
||||
for i, result in enumerate(results):
|
||||
if i not in indices and result not in reranked:
|
||||
reranked.append(result)
|
||||
|
||||
logger.info(f"LLM re-ranking: reordered {len(reranked)} results")
|
||||
return reranked
|
||||
|
||||
except Exception as e:
|
||||
logger.warning(f"LLM re-ranking failed: {e}, using RRF order")
|
||||
return results # Fallback to RRF order
|
||||
|
||||
def _format_context_for_llm(self, results: List[HybridRAGResult]) -> str:
|
||||
"""
|
||||
Phase 5: Format results into context for LLM consumption.
|
||||
|
||||
Args:
|
||||
results: Ranked results
|
||||
|
||||
Returns:
|
||||
Formatted context string
|
||||
"""
|
||||
context_parts = []
|
||||
|
||||
for i, result in enumerate(results[:10], start=1):
|
||||
# Source indicator
|
||||
source_tag = f"[{result.source_type.upper()}]"
|
||||
|
||||
# Related dossiers if available
|
||||
related = ""
|
||||
if result.related_dossiers:
|
||||
tags = ", ".join([d.tag for d in result.related_dossiers[:3]])
|
||||
related = f"\n Related research: {tags}"
|
||||
|
||||
# Build context entry
|
||||
content_preview = result.content[:300] if result.content else "(no content)"
|
||||
context_parts.append(
|
||||
f"{i}. {source_tag} {result.title}\n"
|
||||
f" {content_preview}...{related}"
|
||||
)
|
||||
|
||||
return "\n\n".join(context_parts)
|
||||
|
||||
async def _persist_search_for_librarian(
|
||||
self,
|
||||
query: str,
|
||||
user: str,
|
||||
keywords_data: Dict[str, Any],
|
||||
raw_results: Dict[str, List],
|
||||
final_results: List[Dict[str, Any]],
|
||||
timing: Dict[str, float]
|
||||
) -> Optional[str]:
|
||||
"""
|
||||
Phase 6: Store search query and results for Librarian processing.
|
||||
|
||||
Creates SearchQuery node in Neo4j with relationships to found documents
|
||||
and web results for offline knowledge consolidation.
|
||||
|
||||
Args:
|
||||
query: Search query
|
||||
user: User identifier
|
||||
keywords_data: Extracted keywords/synonyms
|
||||
raw_results: Results from each source
|
||||
final_results: Final ranked results
|
||||
timing: Performance timing
|
||||
|
||||
Returns:
|
||||
Search ID for tracking
|
||||
"""
|
||||
try:
|
||||
user_base_label = get_neo4j_user_base_label(user)
|
||||
search_id = str(uuid.uuid4())
|
||||
|
||||
# Create SearchQuery node
|
||||
create_query = f"""
|
||||
CREATE (sq:{user_base_label}_SearchQuery:SearchQuery {{
|
||||
id: $search_id,
|
||||
query: $query,
|
||||
user: $user,
|
||||
timestamp: datetime(),
|
||||
processed: false,
|
||||
total_results: $total_results,
|
||||
vector_count: $vector_count,
|
||||
graph_count: $graph_count,
|
||||
web_count: $web_count,
|
||||
keywords: $keywords,
|
||||
synonyms: $synonyms,
|
||||
timing_ms: $timing_ms
|
||||
}})
|
||||
RETURN sq.id as id
|
||||
"""
|
||||
|
||||
result = await self.graph.neo4j.execute_query(create_query, {
|
||||
"search_id": search_id,
|
||||
"query": query,
|
||||
"user": user,
|
||||
"total_results": len(final_results),
|
||||
"vector_count": len(raw_results.get("vector", [])),
|
||||
"graph_count": len(raw_results.get("graph", [])),
|
||||
"web_count": len(raw_results.get("web", [])),
|
||||
"keywords": keywords_data.get("core_keywords", []),
|
||||
"synonyms": json.dumps(keywords_data.get("synonyms", {})),
|
||||
"timing_ms": timing.get("total_ms", 0)
|
||||
})
|
||||
|
||||
# Link to found wiki documents (top 20)
|
||||
for rank, result_data in enumerate(final_results[:20], start=1):
|
||||
result = result_data.get("result", {})
|
||||
page_id = result.get("page_id")
|
||||
|
||||
if page_id:
|
||||
link_doc_query = f"""
|
||||
MATCH (sq:{user_base_label}_SearchQuery:SearchQuery {{id: $search_id}})
|
||||
MATCH (d:Document {{page_id: $page_id}})
|
||||
MERGE (sq)-[f:FOUND]->(d)
|
||||
SET f.source = $source,
|
||||
f.rank = $rank,
|
||||
f.rrf_score = $rrf_score,
|
||||
f.final_rank = $final_rank
|
||||
"""
|
||||
|
||||
await self.graph.neo4j.execute_query(link_doc_query, {
|
||||
"search_id": search_id,
|
||||
"page_id": page_id,
|
||||
"source": result_data.get("source_type", "unknown"),
|
||||
"rank": rank,
|
||||
"rrf_score": result_data.get("rrf_score", 0),
|
||||
"final_rank": result_data.get("final_rank", rank)
|
||||
})
|
||||
|
||||
# Store web results as WebResult nodes (top 10)
|
||||
web_results = [r for r in final_results[:10] if r.get("result", {}).get("url")]
|
||||
for rank, result_data in enumerate(web_results, start=1):
|
||||
result = result_data.get("result", {})
|
||||
create_web_query = f"""
|
||||
MATCH (sq:{user_base_label}_SearchQuery:SearchQuery {{id: $search_id}})
|
||||
CREATE (wr:{user_base_label}_WebResult:WebResult {{
|
||||
url: $url,
|
||||
title: $title,
|
||||
content: $content,
|
||||
search_id: $search_id,
|
||||
timestamp: datetime()
|
||||
}})
|
||||
CREATE (sq)-[:FOUND {{
|
||||
source: "web",
|
||||
rank: $rank,
|
||||
rrf_score: $rrf_score
|
||||
}}]->(wr)
|
||||
"""
|
||||
|
||||
await self.graph.neo4j.execute_query(create_web_query, {
|
||||
"search_id": search_id,
|
||||
"url": result.get("url"),
|
||||
"title": result.get("title", ""),
|
||||
"content": result.get("content", "")[:1000], # Truncate
|
||||
"rank": rank,
|
||||
"rrf_score": result_data.get("rrf_score", 0)
|
||||
})
|
||||
|
||||
logger.info(f"Persisted search {search_id} for Librarian processing")
|
||||
return search_id
|
||||
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to persist search for Librarian: {e}", exc_info=True)
|
||||
return None
|
||||
|
||||
def _convert_to_result_models(self, results: List[Dict[str, Any]]) -> List[HybridRAGResult]:
|
||||
"""
|
||||
Convert internal result format to HybridRAGResult models.
|
||||
|
||||
Args:
|
||||
results: Internal result dictionaries
|
||||
|
||||
Returns:
|
||||
List of HybridRAGResult models
|
||||
"""
|
||||
models = []
|
||||
|
||||
for result_data in results:
|
||||
result = result_data.get("result", {})
|
||||
related_dossiers = result_data.get("related_dossiers", [])
|
||||
|
||||
models.append(HybridRAGResult(
|
||||
source_type=result_data.get("source_type", "unknown"),
|
||||
title=result.get("title", "Untitled"),
|
||||
content=result.get("content", ""),
|
||||
url=result.get("url"),
|
||||
page_id=result.get("page_id"),
|
||||
page_path=result.get("path"),
|
||||
rrf_score=result_data.get("rrf_score", 0),
|
||||
final_rank=result_data.get("final_rank", 0),
|
||||
sources=result_data.get("sources", []),
|
||||
related_dossiers=[RelatedDossier(**d) for d in related_dossiers],
|
||||
metadata={
|
||||
"entity_matches": result.get("entity_matches"),
|
||||
"matched_entities": result.get("matched_entities"),
|
||||
"engine": result.get("engine")
|
||||
}
|
||||
))
|
||||
|
||||
return models
|
||||
@@ -1,414 +0,0 @@
|
||||
"""
|
||||
Document Ingestion Service
|
||||
|
||||
Orchestrates the ingestion of wiki pages into the knowledge base:
|
||||
1. Fetches page content from Wiki.js
|
||||
2. Generates vector embeddings (Qdrant)
|
||||
3. Extracts entities and updates knowledge graph (Neo4j)
|
||||
|
||||
This service is called by:
|
||||
- Consolidation service (after creating/updating pages)
|
||||
- Manual ingestion endpoints
|
||||
- Batch ingestion jobs
|
||||
"""
|
||||
import logging
|
||||
import asyncio
|
||||
from typing import List, Optional
|
||||
from datetime import datetime
|
||||
import time
|
||||
|
||||
from src.services.vector_service import VectorService
|
||||
from src.services.graph_service import GraphService
|
||||
from src.clients.wikijs_client import WikiJSClient
|
||||
from src.models.ingestion import (
|
||||
IngestionRequest,
|
||||
IngestionResult,
|
||||
BatchIngestionRequest,
|
||||
BatchIngestionResult
|
||||
)
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
|
||||
class IngestionService:
|
||||
"""
|
||||
Service for ingesting wiki pages into the knowledge base.
|
||||
"""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
vector_service: VectorService,
|
||||
graph_service: GraphService,
|
||||
wiki_client: WikiJSClient
|
||||
):
|
||||
self.vector = vector_service
|
||||
self.graph = graph_service
|
||||
self.wiki = wiki_client
|
||||
|
||||
async def ingest_page(
|
||||
self,
|
||||
page_id: int,
|
||||
user: str,
|
||||
force_refresh: bool = False,
|
||||
skip_vectors: bool = False,
|
||||
skip_graph: bool = False,
|
||||
skip_entity_linking: bool = False
|
||||
) -> IngestionResult:
|
||||
"""
|
||||
Ingest a single wiki page into the knowledge base.
|
||||
|
||||
Args:
|
||||
page_id: Wiki page ID
|
||||
user: User identifier
|
||||
force_refresh: Force re-ingestion even if unchanged
|
||||
skip_vectors: Skip vector embedding generation
|
||||
skip_graph: Skip graph entity extraction
|
||||
skip_entity_linking: Skip automatic entity linking
|
||||
|
||||
Returns:
|
||||
IngestionResult with operation details
|
||||
"""
|
||||
start_time = time.time()
|
||||
|
||||
logger.info(f"Starting ingestion for page {page_id} (user: {user})")
|
||||
|
||||
try:
|
||||
# Fetch page to get metadata
|
||||
page = await self.wiki.get_page(page_id)
|
||||
if not page:
|
||||
return IngestionResult(
|
||||
page_id=page_id,
|
||||
page_title=f"Page {page_id}",
|
||||
success=False,
|
||||
error="Page not found in Wiki.js",
|
||||
processing_time_ms=(time.time() - start_time) * 1000
|
||||
)
|
||||
|
||||
page_title = page.get("title", f"Page {page_id}")
|
||||
page_path = page.get("path", "")
|
||||
|
||||
# Ingest vectors and graph in parallel
|
||||
tasks = []
|
||||
|
||||
if not skip_vectors:
|
||||
tasks.append(self._ingest_vectors(page_id, user, force_refresh))
|
||||
else:
|
||||
tasks.append(asyncio.create_task(asyncio.sleep(0))) # Dummy task
|
||||
|
||||
if not skip_graph:
|
||||
tasks.append(self._ingest_graph(page_id, user, force_refresh))
|
||||
else:
|
||||
tasks.append(asyncio.create_task(asyncio.sleep(0))) # Dummy task
|
||||
|
||||
# Execute in parallel
|
||||
vector_result, graph_result = await asyncio.gather(*tasks, return_exceptions=True)
|
||||
|
||||
# Handle errors
|
||||
vector_chunks = 0
|
||||
graph_entities = 0
|
||||
graph_relationships = 0
|
||||
errors = []
|
||||
|
||||
if not skip_vectors:
|
||||
if isinstance(vector_result, Exception):
|
||||
errors.append(f"Vector ingestion failed: {str(vector_result)}")
|
||||
logger.error(f"Vector ingestion failed for page {page_id}: {vector_result}")
|
||||
else:
|
||||
vector_chunks = vector_result.get("chunks_created", 0)
|
||||
|
||||
if not skip_graph:
|
||||
if isinstance(graph_result, Exception):
|
||||
errors.append(f"Graph ingestion failed: {str(graph_result)}")
|
||||
logger.error(f"Graph ingestion failed for page {page_id}: {graph_result}")
|
||||
else:
|
||||
# entities_extracted is a list, get its length
|
||||
entities_list = graph_result.get("entities_extracted", [])
|
||||
graph_entities = len(entities_list) if isinstance(entities_list, list) else 0
|
||||
graph_relationships = graph_result.get("relationships_created", 0)
|
||||
|
||||
# Step 3: Link existing entities in the page content (after graph extraction)
|
||||
entity_links_created = 0
|
||||
if not skip_entity_linking and not skip_graph and not isinstance(graph_result, Exception):
|
||||
try:
|
||||
entity_links_created = await self._link_existing_entities(page_id, user, page)
|
||||
logger.info(f"Created {entity_links_created} entity mention links for page {page_id}")
|
||||
except Exception as e:
|
||||
logger.warning(f"Entity linking failed for page {page_id}: {e}")
|
||||
# Don't fail the whole ingestion if entity linking fails
|
||||
|
||||
processing_time_ms = (time.time() - start_time) * 1000
|
||||
|
||||
result = IngestionResult(
|
||||
page_id=page_id,
|
||||
page_title=page_title,
|
||||
page_path=page_path,
|
||||
success=len(errors) == 0,
|
||||
error="; ".join(errors) if errors else None,
|
||||
vector_chunks_created=vector_chunks,
|
||||
graph_entities_extracted=graph_entities,
|
||||
graph_relationships_created=graph_relationships,
|
||||
processing_time_ms=processing_time_ms
|
||||
)
|
||||
|
||||
if result.success:
|
||||
logger.info(
|
||||
f"Successfully ingested page {page_id}: "
|
||||
f"{vector_chunks} chunks, {graph_entities} entities, "
|
||||
f"{graph_relationships} relationships, {entity_links_created} entity links "
|
||||
f"in {processing_time_ms:.0f}ms"
|
||||
)
|
||||
else:
|
||||
logger.warning(f"Partial ingestion failure for page {page_id}: {result.error}")
|
||||
|
||||
return result
|
||||
|
||||
except Exception as e:
|
||||
logger.error(f"Ingestion failed for page {page_id}: {e}", exc_info=True)
|
||||
return IngestionResult(
|
||||
page_id=page_id,
|
||||
page_title=f"Page {page_id}",
|
||||
success=False,
|
||||
error=str(e),
|
||||
processing_time_ms=(time.time() - start_time) * 1000
|
||||
)
|
||||
|
||||
async def _ingest_vectors(
|
||||
self,
|
||||
page_id: int,
|
||||
user: str,
|
||||
force_refresh: bool
|
||||
) -> dict:
|
||||
"""
|
||||
Ingest page into vector database.
|
||||
|
||||
Returns:
|
||||
Dict with chunks_created count
|
||||
"""
|
||||
try:
|
||||
summary = await self.vector.update_from_page(
|
||||
page_id=page_id,
|
||||
user=user,
|
||||
force_refresh=force_refresh
|
||||
)
|
||||
|
||||
return {
|
||||
"chunks_created": summary.chunks_created,
|
||||
"chunks_deleted": summary.chunks_deleted
|
||||
}
|
||||
|
||||
except Exception as e:
|
||||
logger.error(f"Vector ingestion failed for page {page_id}: {e}")
|
||||
raise
|
||||
|
||||
async def _ingest_graph(
|
||||
self,
|
||||
page_id: int,
|
||||
user: str,
|
||||
force_refresh: bool
|
||||
) -> dict:
|
||||
"""
|
||||
Ingest page into knowledge graph.
|
||||
|
||||
Returns:
|
||||
Dict with entities_extracted and relationships_created counts
|
||||
"""
|
||||
try:
|
||||
summary = await self.graph.update_from_page(
|
||||
page_id=page_id,
|
||||
user=user,
|
||||
force_refresh=force_refresh
|
||||
)
|
||||
|
||||
return {
|
||||
"entities_extracted": summary.entities_extracted,
|
||||
"relationships_created": summary.relationships_created
|
||||
}
|
||||
|
||||
except Exception as e:
|
||||
logger.error(f"Graph ingestion failed for page {page_id}: {e}")
|
||||
raise
|
||||
|
||||
async def _link_existing_entities(
|
||||
self,
|
||||
page_id: int,
|
||||
user: str,
|
||||
page: dict
|
||||
) -> int:
|
||||
"""
|
||||
Find and link mentions of existing entities in the page content.
|
||||
|
||||
This runs automatically after graph extraction to create MENTIONS relationships
|
||||
for entities that already exist in the knowledge graph but were mentioned in
|
||||
this page.
|
||||
|
||||
Args:
|
||||
page_id: Wiki page ID
|
||||
user: User identifier
|
||||
page: Page dict with content (from WikiJSClient)
|
||||
|
||||
Returns:
|
||||
Number of new entity mention links created
|
||||
"""
|
||||
import re
|
||||
|
||||
try:
|
||||
page_content = page.get("content", "")
|
||||
if not page_content or len(page_content) < 10:
|
||||
return 0
|
||||
|
||||
# Get all existing entities from the knowledge graph
|
||||
entities = await self.graph.get_all_entities(user)
|
||||
if not entities:
|
||||
logger.debug(f"No existing entities found for user {user}, skipping entity linking")
|
||||
return 0
|
||||
|
||||
# Find entity mentions in page content
|
||||
found_entities = []
|
||||
content_lower = page_content.lower()
|
||||
|
||||
for entity in entities:
|
||||
entity_name = entity.get("name", "")
|
||||
if not entity_name or len(entity_name) < 3:
|
||||
continue
|
||||
|
||||
# Create regex pattern for whole word matching
|
||||
# This avoids matching "John" in "Johnson"
|
||||
pattern = r'\b' + re.escape(entity_name.lower()) + r'\b'
|
||||
|
||||
# Find all matches
|
||||
matches = list(re.finditer(pattern, content_lower))
|
||||
|
||||
if matches:
|
||||
found_entities.append({
|
||||
"name": entity_name,
|
||||
"type": entity.get("type", "unknown"),
|
||||
"mentions": len(matches),
|
||||
"entity_id": entity.get("id")
|
||||
})
|
||||
|
||||
if not found_entities:
|
||||
logger.debug(f"No entity mentions found in page {page_id}")
|
||||
return 0
|
||||
|
||||
# Create MENTIONS relationships
|
||||
new_links_created = await self.graph.create_entity_mentions(
|
||||
page_id=page_id,
|
||||
user=user,
|
||||
entity_names=found_entities
|
||||
)
|
||||
|
||||
return new_links_created
|
||||
|
||||
except Exception as e:
|
||||
logger.error(f"Entity linking failed for page {page_id}: {e}")
|
||||
raise
|
||||
|
||||
async def ingest_batch(
|
||||
self,
|
||||
page_ids: List[int],
|
||||
user: str,
|
||||
force_refresh: bool = False,
|
||||
skip_vectors: bool = False,
|
||||
skip_graph: bool = False,
|
||||
max_concurrent: int = 3
|
||||
) -> BatchIngestionResult:
|
||||
"""
|
||||
Ingest multiple wiki pages concurrently.
|
||||
|
||||
Args:
|
||||
page_ids: List of wiki page IDs to ingest
|
||||
user: User identifier
|
||||
force_refresh: Force re-ingestion
|
||||
skip_vectors: Skip vector embedding generation
|
||||
skip_graph: Skip graph entity extraction
|
||||
max_concurrent: Maximum concurrent ingestion tasks
|
||||
|
||||
Returns:
|
||||
BatchIngestionResult with per-page results
|
||||
"""
|
||||
start_time = time.time()
|
||||
|
||||
logger.info(f"Starting batch ingestion of {len(page_ids)} pages (user: {user})")
|
||||
|
||||
results = []
|
||||
semaphore = asyncio.Semaphore(max_concurrent)
|
||||
|
||||
async def ingest_with_semaphore(page_id: int):
|
||||
async with semaphore:
|
||||
return await self.ingest_page(
|
||||
page_id=page_id,
|
||||
user=user,
|
||||
force_refresh=force_refresh,
|
||||
skip_vectors=skip_vectors,
|
||||
skip_graph=skip_graph
|
||||
)
|
||||
|
||||
# Execute all ingestions with concurrency control
|
||||
tasks = [ingest_with_semaphore(page_id) for page_id in page_ids]
|
||||
results = await asyncio.gather(*tasks)
|
||||
|
||||
# Calculate summary
|
||||
successful = sum(1 for r in results if r.success)
|
||||
failed = len(results) - successful
|
||||
total_processing_time_ms = (time.time() - start_time) * 1000
|
||||
|
||||
batch_result = BatchIngestionResult(
|
||||
total_pages=len(page_ids),
|
||||
successful=successful,
|
||||
failed=failed,
|
||||
results=results,
|
||||
total_processing_time_ms=total_processing_time_ms
|
||||
)
|
||||
|
||||
logger.info(
|
||||
f"Batch ingestion complete: {successful}/{len(page_ids)} successful "
|
||||
f"in {total_processing_time_ms:.0f}ms"
|
||||
)
|
||||
|
||||
return batch_result
|
||||
|
||||
async def ingest_all_pages(
|
||||
self,
|
||||
user: str,
|
||||
path_prefix: Optional[str] = None,
|
||||
force_refresh: bool = False,
|
||||
max_concurrent: int = 3
|
||||
) -> BatchIngestionResult:
|
||||
"""
|
||||
Ingest all wiki pages for a user.
|
||||
|
||||
Args:
|
||||
user: User identifier
|
||||
path_prefix: Optional path prefix filter (e.g., "users/jpmschweitzer")
|
||||
force_refresh: Force re-ingestion
|
||||
max_concurrent: Maximum concurrent ingestion tasks
|
||||
|
||||
Returns:
|
||||
BatchIngestionResult
|
||||
"""
|
||||
logger.info(f"Finding all pages for user {user} (prefix: {path_prefix or 'all'})")
|
||||
|
||||
# List all pages (not search - search requires a query and may have stale index)
|
||||
pages = await self.wiki.list_all_pages(
|
||||
path_prefix=path_prefix or f"users/{user}"
|
||||
)
|
||||
|
||||
if not pages:
|
||||
logger.warning(f"No pages found for user {user}")
|
||||
return BatchIngestionResult(
|
||||
total_pages=0,
|
||||
successful=0,
|
||||
failed=0,
|
||||
results=[],
|
||||
total_processing_time_ms=0
|
||||
)
|
||||
|
||||
page_ids = [p['id'] for p in pages]
|
||||
logger.info(f"Found {len(page_ids)} pages to ingest")
|
||||
|
||||
return await self.ingest_batch(
|
||||
page_ids=page_ids,
|
||||
user=user,
|
||||
force_refresh=force_refresh,
|
||||
max_concurrent=max_concurrent
|
||||
)
|
||||
@@ -1,358 +0,0 @@
|
||||
"""
|
||||
Vector service for Library Desk Qdrant operations.
|
||||
|
||||
Handles semantic search, document chunking, and embeddings.
|
||||
"""
|
||||
|
||||
import re
|
||||
import time
|
||||
import hashlib
|
||||
import uuid
|
||||
from typing import List, Dict, Any, Optional
|
||||
import logging
|
||||
|
||||
from src.clients.qdrant_client import QdrantClientWrapper
|
||||
from src.clients.wikijs_client import WikiJSClient
|
||||
from src.clients.ollama_client import OllamaClient
|
||||
from src.core.multi_tenancy import get_qdrant_collection_name
|
||||
from src.models.vector import (
|
||||
SearchResult, SearchResponse, VectorUpdateSummary,
|
||||
DocumentChunk, CollectionInfo, CollectionListResponse
|
||||
)
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
|
||||
class VectorService:
|
||||
"""
|
||||
Service for Qdrant vector operations.
|
||||
|
||||
Responsibilities:
|
||||
- Document chunking
|
||||
- Embedding generation
|
||||
- Semantic search
|
||||
- Vector CRUD operations
|
||||
"""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
qdrant_client: QdrantClientWrapper,
|
||||
wikijs_client: WikiJSClient,
|
||||
ollama_client: OllamaClient,
|
||||
chunk_size: int = 500,
|
||||
chunk_overlap: int = 50
|
||||
):
|
||||
"""
|
||||
Initialize vector service.
|
||||
|
||||
Args:
|
||||
qdrant_client: Qdrant database client
|
||||
wikijs_client: Wiki.js client for fetching pages
|
||||
ollama_client: Ollama client for embeddings
|
||||
chunk_size: Target chunk size in tokens (approximate)
|
||||
chunk_overlap: Overlap between chunks in tokens
|
||||
"""
|
||||
self.qdrant = qdrant_client
|
||||
self.wiki = wikijs_client
|
||||
self.ollama = ollama_client
|
||||
self.chunk_size = chunk_size
|
||||
self.chunk_overlap = chunk_overlap
|
||||
|
||||
def _chunk_text(self, text: str) -> List[str]:
|
||||
"""
|
||||
Chunk text into overlapping segments.
|
||||
|
||||
Simple word-based chunking for now.
|
||||
TODO: Use tiktoken or similar for token-accurate chunking.
|
||||
|
||||
Args:
|
||||
text: Text to chunk
|
||||
|
||||
Returns:
|
||||
List of text chunks
|
||||
"""
|
||||
# Remove extra whitespace
|
||||
text = re.sub(r'\s+', ' ', text).strip()
|
||||
|
||||
# Split into words (approximates tokens)
|
||||
words = text.split()
|
||||
|
||||
if len(words) <= self.chunk_size:
|
||||
return [text]
|
||||
|
||||
chunks = []
|
||||
start = 0
|
||||
|
||||
while start < len(words):
|
||||
end = start + self.chunk_size
|
||||
chunk_words = words[start:end]
|
||||
chunks.append(' '.join(chunk_words))
|
||||
|
||||
# Move start forward with overlap
|
||||
start = end - self.chunk_overlap
|
||||
|
||||
return chunks
|
||||
|
||||
async def update_from_page(
|
||||
self,
|
||||
page_id: int,
|
||||
user: str,
|
||||
force_refresh: bool = False
|
||||
) -> VectorUpdateSummary:
|
||||
"""
|
||||
Update vector embeddings from a wiki page.
|
||||
|
||||
Chunks the page content, generates embeddings, and upserts to Qdrant.
|
||||
|
||||
Args:
|
||||
page_id: Wiki page ID
|
||||
user: User identifier
|
||||
force_refresh: Force re-embedding even if unchanged
|
||||
|
||||
Returns:
|
||||
Summary of update operation
|
||||
"""
|
||||
start_time = time.time()
|
||||
|
||||
try:
|
||||
# Fetch page from Wiki.js
|
||||
page = await self.wiki.get_page(page_id)
|
||||
if not page:
|
||||
raise ValueError(f"Page {page_id} not found")
|
||||
|
||||
# Get collection name for user
|
||||
collection_name = get_qdrant_collection_name(user)
|
||||
|
||||
# Ensure collection exists
|
||||
await self.qdrant.ensure_collection(collection_name)
|
||||
|
||||
# Extract content
|
||||
content = page.get("content", "")
|
||||
title = page.get("title", "")
|
||||
path = page.get("path", "")
|
||||
|
||||
if not content:
|
||||
logger.warning(f"Page {page_id} has no content, skipping vector update")
|
||||
return VectorUpdateSummary(
|
||||
page_id=page_id,
|
||||
page_title=title,
|
||||
processing_time_ms=(time.time() - start_time) * 1000,
|
||||
success=True
|
||||
)
|
||||
|
||||
# Chunk the content
|
||||
chunks = self._chunk_text(content)
|
||||
logger.info(f"Split page {page_id} into {len(chunks)} chunks")
|
||||
|
||||
# Delete existing chunks for this page
|
||||
deleted_count = await self.qdrant.delete_by_filter(
|
||||
collection_name=collection_name,
|
||||
filter_conditions={"page_id": page_id}
|
||||
)
|
||||
|
||||
# Generate embeddings and upsert chunks
|
||||
chunks_created = 0
|
||||
for idx, chunk_text in enumerate(chunks):
|
||||
# Generate deterministic UUID from page_id and chunk_index
|
||||
chunk_id = str(uuid.uuid5(uuid.NAMESPACE_DNS, f"page_{page_id}_chunk_{idx}"))
|
||||
|
||||
# Generate embedding
|
||||
embedding = await self.ollama.embed(chunk_text)
|
||||
if not embedding:
|
||||
logger.error(f"Failed to generate embedding for chunk {chunk_id}")
|
||||
continue
|
||||
|
||||
# Prepare metadata
|
||||
metadata = {
|
||||
"page_id": page_id,
|
||||
"page_title": title,
|
||||
"page_path": path,
|
||||
"chunk_index": idx,
|
||||
"chunk_text": chunk_text,
|
||||
"user": user
|
||||
}
|
||||
|
||||
# Upsert to Qdrant
|
||||
success = await self.qdrant.upsert_vector(
|
||||
collection_name=collection_name,
|
||||
vector_id=chunk_id,
|
||||
vector=embedding,
|
||||
payload=metadata
|
||||
)
|
||||
|
||||
if success:
|
||||
chunks_created += 1
|
||||
|
||||
processing_time_ms = (time.time() - start_time) * 1000
|
||||
|
||||
logger.info(
|
||||
f"Updated vectors for page {page_id}: "
|
||||
f"{chunks_created} chunks created, {deleted_count} old chunks deleted"
|
||||
)
|
||||
|
||||
return VectorUpdateSummary(
|
||||
page_id=page_id,
|
||||
page_title=title,
|
||||
chunks_created=chunks_created,
|
||||
chunks_deleted=deleted_count,
|
||||
total_chunks=chunks_created,
|
||||
embedding_dim=len(embedding) if embedding else 768,
|
||||
processing_time_ms=processing_time_ms,
|
||||
success=True
|
||||
)
|
||||
|
||||
except Exception as e:
|
||||
processing_time_ms = (time.time() - start_time) * 1000
|
||||
logger.error(f"Failed to update vectors for page {page_id}: {e}", exc_info=True)
|
||||
|
||||
return VectorUpdateSummary(
|
||||
page_id=page_id,
|
||||
page_title="Unknown",
|
||||
processing_time_ms=processing_time_ms,
|
||||
success=False,
|
||||
error_message=str(e)
|
||||
)
|
||||
|
||||
async def search(
|
||||
self,
|
||||
query: str,
|
||||
user: str,
|
||||
limit: int = 10,
|
||||
score_threshold: float = 0.5
|
||||
) -> SearchResponse:
|
||||
"""
|
||||
Perform semantic search across user's documents.
|
||||
|
||||
Args:
|
||||
query: Search query text
|
||||
user: User identifier
|
||||
limit: Maximum results to return
|
||||
score_threshold: Minimum similarity score (0-1)
|
||||
|
||||
Returns:
|
||||
Search results with similarity scores
|
||||
"""
|
||||
start_time = time.time()
|
||||
|
||||
try:
|
||||
# Get collection name
|
||||
collection_name = get_qdrant_collection_name(user)
|
||||
|
||||
# Check if collection exists
|
||||
exists = await self.qdrant.collection_exists(collection_name)
|
||||
if not exists:
|
||||
logger.info(f"Collection {collection_name} doesn't exist, returning empty results")
|
||||
return SearchResponse(
|
||||
query=query,
|
||||
results=[],
|
||||
total=0,
|
||||
user=user
|
||||
)
|
||||
|
||||
# Generate query embedding
|
||||
query_embedding = await self.ollama.embed(query)
|
||||
if not query_embedding:
|
||||
raise ValueError("Failed to generate query embedding")
|
||||
|
||||
# Search in Qdrant
|
||||
search_results = await self.qdrant.search_vectors(
|
||||
collection_name=collection_name,
|
||||
query_vector=query_embedding,
|
||||
limit=limit,
|
||||
score_threshold=score_threshold
|
||||
)
|
||||
|
||||
# Convert to SearchResult models
|
||||
results = []
|
||||
for result in search_results:
|
||||
payload = result.get("payload", {})
|
||||
results.append(SearchResult(
|
||||
chunk_id=result["id"],
|
||||
page_id=payload.get("page_id", 0),
|
||||
page_title=payload.get("page_title"),
|
||||
page_path=payload.get("page_path"),
|
||||
chunk_index=payload.get("chunk_index", 0),
|
||||
content=payload.get("chunk_text", ""),
|
||||
score=result["score"],
|
||||
metadata=payload
|
||||
))
|
||||
|
||||
query_time_ms = (time.time() - start_time) * 1000
|
||||
logger.info(f"Semantic search completed in {query_time_ms:.2f}ms: {len(results)} results")
|
||||
|
||||
return SearchResponse(
|
||||
query=query,
|
||||
results=results,
|
||||
total=len(results),
|
||||
user=user
|
||||
)
|
||||
|
||||
except Exception as e:
|
||||
logger.error(f"Semantic search failed: {e}", exc_info=True)
|
||||
return SearchResponse(
|
||||
query=query,
|
||||
results=[],
|
||||
total=0,
|
||||
user=user
|
||||
)
|
||||
|
||||
async def delete_page_chunks(
|
||||
self,
|
||||
page_id: int,
|
||||
user: str
|
||||
) -> int:
|
||||
"""
|
||||
Delete all chunks for a wiki page.
|
||||
|
||||
Args:
|
||||
page_id: Wiki page ID
|
||||
user: User identifier
|
||||
|
||||
Returns:
|
||||
Number of chunks deleted
|
||||
"""
|
||||
collection_name = get_qdrant_collection_name(user)
|
||||
|
||||
try:
|
||||
deleted_count = await self.qdrant.delete_by_filter(
|
||||
collection_name=collection_name,
|
||||
filter_conditions={"page_id": page_id}
|
||||
)
|
||||
|
||||
logger.info(f"Deleted {deleted_count} chunks for page {page_id}")
|
||||
return deleted_count
|
||||
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to delete chunks for page {page_id}: {e}", exc_info=True)
|
||||
return 0
|
||||
|
||||
async def list_collections(self) -> CollectionListResponse:
|
||||
"""
|
||||
List all Qdrant collections.
|
||||
|
||||
Returns:
|
||||
List of collections with stats
|
||||
"""
|
||||
try:
|
||||
collections_data = await self.qdrant.list_collections()
|
||||
|
||||
collections = []
|
||||
for coll in collections_data:
|
||||
collections.append(CollectionInfo(
|
||||
name=coll["name"],
|
||||
vectors_count=coll.get("vectors_count", 0),
|
||||
points_count=coll.get("points_count", 0),
|
||||
segments_count=coll.get("segments_count", 0)
|
||||
))
|
||||
|
||||
return CollectionListResponse(
|
||||
collections=collections,
|
||||
total=len(collections)
|
||||
)
|
||||
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to list collections: {e}", exc_info=True)
|
||||
return CollectionListResponse(
|
||||
collections=[],
|
||||
total=0
|
||||
)
|
||||
@@ -1,279 +0,0 @@
|
||||
"""
|
||||
Wiki.js Database Change Listener
|
||||
|
||||
Listens to PostgreSQL NOTIFY events for page changes in Wiki.js
|
||||
and triggers the same processing as webhooks would.
|
||||
|
||||
This is an alternative to Wiki.js webhooks (which don't exist in open-source version).
|
||||
"""
|
||||
import logging
|
||||
import asyncio
|
||||
import asyncpg
|
||||
from typing import Optional
|
||||
from datetime import datetime
|
||||
|
||||
from src.config import get_settings
|
||||
from src.core.dependencies import get_ingestion_service
|
||||
from src.services.consolidation_service import ConsolidationService
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
|
||||
class WikiChangeListener:
|
||||
"""
|
||||
Listens to PostgreSQL NOTIFY events from Wiki.js database.
|
||||
|
||||
This requires setting up triggers in the Wiki.js database to emit
|
||||
NOTIFY events on INSERT/UPDATE/DELETE to the pages table.
|
||||
"""
|
||||
|
||||
def __init__(self):
|
||||
self.settings = get_settings()
|
||||
self.connection: Optional[asyncpg.Connection] = None
|
||||
self.running = False
|
||||
|
||||
# Loop prevention: Track recently processed pages
|
||||
# Key: page_id, Value: timestamp of last processing
|
||||
self._recent_notifications = {}
|
||||
self._debounce_seconds = self.settings.wikijs_change_listener_debounce_seconds
|
||||
|
||||
async def start(self):
|
||||
"""Start listening to database changes."""
|
||||
logger.info("Starting Wiki.js database change listener")
|
||||
|
||||
# Connect to Wiki.js PostgreSQL database
|
||||
self.connection = await asyncpg.connect(
|
||||
host=self.settings.wikijs_db_host,
|
||||
port=self.settings.wikijs_db_port,
|
||||
user=self.settings.wikijs_db_user,
|
||||
password=self.settings.wikijs_db_password,
|
||||
database=self.settings.wikijs_db_name
|
||||
)
|
||||
|
||||
# Listen to the wiki_page_changes channel
|
||||
await self.connection.add_listener('wiki_page_changes', self._handle_notification)
|
||||
|
||||
self.running = True
|
||||
logger.info("Listening for Wiki.js page changes via PostgreSQL NOTIFY")
|
||||
|
||||
async def stop(self):
|
||||
"""Stop listening and close connection."""
|
||||
if self.connection:
|
||||
await self.connection.remove_listener('wiki_page_changes', self._handle_notification)
|
||||
await self.connection.close()
|
||||
self.running = False
|
||||
logger.info("Stopped Wiki.js change listener")
|
||||
|
||||
async def _handle_notification(self, connection, pid, channel, payload):
|
||||
"""Handle NOTIFY event from database."""
|
||||
try:
|
||||
# Payload format: "operation:page_id:user_email"
|
||||
# e.g., "INSERT:123:user@example.com"
|
||||
parts = payload.split(':')
|
||||
if len(parts) < 3:
|
||||
logger.warning(f"Invalid notification payload: {payload}")
|
||||
return
|
||||
|
||||
operation = parts[0] # INSERT, UPDATE, DELETE
|
||||
page_id = int(parts[1])
|
||||
user_email = parts[2]
|
||||
|
||||
logger.info(f"Received {operation} notification for page {page_id} by {user_email}")
|
||||
|
||||
# LOOP PREVENTION: Debouncing - ignore rapid duplicate notifications
|
||||
# Note: We rely solely on debouncing for loop prevention because:
|
||||
# - The user_email in notifications is the page creator, not the editor
|
||||
# - Creator != namespace owner (e.g., 'librarian' creates page in 'users/jpmschweitzer/')
|
||||
# - Filtering by creator breaks legitimate page ingestion
|
||||
if self._is_recently_processed(page_id):
|
||||
logger.debug(
|
||||
f"Skipping notification for page {page_id} - "
|
||||
f"processed within last {self._debounce_seconds}s (debouncing)"
|
||||
)
|
||||
return
|
||||
|
||||
# Mark as recently processed
|
||||
self._mark_as_processed(page_id)
|
||||
|
||||
# Map operation to webhook-style event
|
||||
event_map = {
|
||||
'INSERT': 'page.create',
|
||||
'UPDATE': 'page.update',
|
||||
'DELETE': 'page.delete'
|
||||
}
|
||||
event = event_map.get(operation, 'page.update')
|
||||
|
||||
# Extract user from email
|
||||
user = user_email.split('@')[0] if '@' in user_email else 'jpmschweitzer'
|
||||
|
||||
# Process the change
|
||||
await self._process_page_change(
|
||||
page_id=page_id,
|
||||
event=event,
|
||||
user=user
|
||||
)
|
||||
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to handle notification: {e}", exc_info=True)
|
||||
|
||||
def _is_automated_user(self, email: str) -> bool:
|
||||
"""
|
||||
Check if email belongs to an automated system user.
|
||||
|
||||
These are edits made by library-desk via Wiki.js API (entity linking).
|
||||
We skip processing these to prevent loops.
|
||||
|
||||
Customize this list based on your Wiki.js username for library-desk.
|
||||
"""
|
||||
automated_users = [
|
||||
self.settings.wikijs_username, # Library-desk's Wiki.js API user
|
||||
"library-desk@system",
|
||||
"automation@system",
|
||||
"bot@system"
|
||||
]
|
||||
|
||||
return email.lower() in [u.lower() for u in automated_users]
|
||||
|
||||
def _is_recently_processed(self, page_id: int) -> bool:
|
||||
"""Check if page was processed recently (debouncing)."""
|
||||
if page_id not in self._recent_notifications:
|
||||
return False
|
||||
|
||||
last_processed = self._recent_notifications[page_id]
|
||||
elapsed = (datetime.now() - last_processed).total_seconds()
|
||||
|
||||
return elapsed < self._debounce_seconds
|
||||
|
||||
def _mark_as_processed(self, page_id: int):
|
||||
"""Mark page as recently processed."""
|
||||
self._recent_notifications[page_id] = datetime.now()
|
||||
|
||||
# Clean up old entries (keep last 100 pages)
|
||||
if len(self._recent_notifications) > 100:
|
||||
# Remove oldest entries
|
||||
sorted_items = sorted(
|
||||
self._recent_notifications.items(),
|
||||
key=lambda x: x[1]
|
||||
)
|
||||
self._recent_notifications = dict(sorted_items[-100:])
|
||||
|
||||
async def _process_page_change(self, page_id: int, event: str, user: str):
|
||||
"""Process page change identically to webhook handler."""
|
||||
from src.routers.webhooks import process_wiki_page_change, cleanup_deleted_page
|
||||
|
||||
ingestion_service = get_ingestion_service()
|
||||
|
||||
if event == 'page.delete':
|
||||
# For deletions, need to handle cleanup
|
||||
# Note: We don't have page_title at this point, use page_id
|
||||
await cleanup_deleted_page(
|
||||
page_id=page_id,
|
||||
page_title=f"Page {page_id}",
|
||||
user=user,
|
||||
ingestion_service=ingestion_service
|
||||
)
|
||||
else:
|
||||
# For create/update, get page details and process
|
||||
from src.core.dependencies import get_wiki_service
|
||||
wiki_service = get_wiki_service()
|
||||
|
||||
try:
|
||||
page = await wiki_service.get_page(page_id, user)
|
||||
|
||||
# If page access failed (wrong user), try to extract correct user from page path
|
||||
if not page:
|
||||
# Try to get page metadata without user validation to find correct namespace
|
||||
try:
|
||||
# Query Wiki.js directly for page path
|
||||
page_info = await wiki_service.wiki_client.get_page(page_id)
|
||||
if page_info and page_info.get('path'):
|
||||
# Extract user from path: users/{user}/...
|
||||
path_parts = page_info['path'].split('/')
|
||||
if len(path_parts) >= 2 and path_parts[0] == 'users':
|
||||
correct_user = path_parts[1]
|
||||
logger.debug(f"Retrying page {page_id} with correct user: {correct_user}")
|
||||
page = await wiki_service.get_page(page_id, correct_user)
|
||||
user = correct_user
|
||||
except Exception as e:
|
||||
logger.debug(f"Could not extract user from page {page_id} path: {e}")
|
||||
|
||||
if page:
|
||||
await process_wiki_page_change(
|
||||
page_id=page_id,
|
||||
page_title=page.title,
|
||||
user=user,
|
||||
event=event,
|
||||
ingestion_service=ingestion_service
|
||||
)
|
||||
else:
|
||||
logger.warning(f"Could not retrieve page {page_id} for processing")
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to process page {page_id}: {e}")
|
||||
|
||||
|
||||
# SQL to set up triggers in Wiki.js database
|
||||
SETUP_TRIGGERS_SQL = """
|
||||
-- Create function to notify on page changes
|
||||
-- Note: Wiki.js pages table has authorId (FK to users.id), not authorEmail
|
||||
-- We look up the email from the users table
|
||||
CREATE OR REPLACE FUNCTION notify_page_change()
|
||||
RETURNS TRIGGER AS $$
|
||||
DECLARE
|
||||
author_email TEXT;
|
||||
BEGIN
|
||||
IF TG_OP = 'DELETE' THEN
|
||||
-- Look up email from users table using OLD.authorId
|
||||
SELECT email INTO author_email FROM users WHERE id = OLD."authorId";
|
||||
IF author_email IS NULL THEN
|
||||
author_email := 'unknown@system';
|
||||
END IF;
|
||||
|
||||
PERFORM pg_notify(
|
||||
'wiki_page_changes',
|
||||
TG_OP || ':' || OLD.id || ':' || author_email
|
||||
);
|
||||
RETURN OLD;
|
||||
ELSE
|
||||
-- Look up email from users table using NEW.authorId
|
||||
SELECT email INTO author_email FROM users WHERE id = NEW."authorId";
|
||||
IF author_email IS NULL THEN
|
||||
author_email := 'unknown@system';
|
||||
END IF;
|
||||
|
||||
PERFORM pg_notify(
|
||||
'wiki_page_changes',
|
||||
TG_OP || ':' || NEW.id || ':' || author_email
|
||||
);
|
||||
RETURN NEW;
|
||||
END IF;
|
||||
END;
|
||||
$$ LANGUAGE plpgsql;
|
||||
|
||||
-- Create triggers on pages table
|
||||
DROP TRIGGER IF EXISTS wiki_page_insert_trigger ON pages;
|
||||
CREATE TRIGGER wiki_page_insert_trigger
|
||||
AFTER INSERT ON pages
|
||||
FOR EACH ROW
|
||||
EXECUTE FUNCTION notify_page_change();
|
||||
|
||||
DROP TRIGGER IF EXISTS wiki_page_update_trigger ON pages;
|
||||
CREATE TRIGGER wiki_page_update_trigger
|
||||
AFTER UPDATE ON pages
|
||||
FOR EACH ROW
|
||||
EXECUTE FUNCTION notify_page_change();
|
||||
|
||||
DROP TRIGGER IF EXISTS wiki_page_delete_trigger ON pages;
|
||||
CREATE TRIGGER wiki_page_delete_trigger
|
||||
AFTER DELETE ON pages
|
||||
FOR EACH ROW
|
||||
EXECUTE FUNCTION notify_page_change();
|
||||
|
||||
-- Verify triggers are created
|
||||
SELECT
|
||||
trigger_name,
|
||||
event_manipulation,
|
||||
event_object_table
|
||||
FROM information_schema.triggers
|
||||
WHERE event_object_table = 'pages'
|
||||
ORDER BY trigger_name;
|
||||
"""
|
||||
@@ -1,493 +0,0 @@
|
||||
"""
|
||||
Intelligent Wiki Page Writer Service
|
||||
|
||||
Uses LLM (mistral-nemo) to create and reconstruct wiki pages with:
|
||||
- Holistic content restructuring
|
||||
- Zero fact loss (unless superseded)
|
||||
- Conflict detection and flagging
|
||||
- Standard formatting with template adherence
|
||||
- Professional organization (summary, tables, chapters)
|
||||
|
||||
This service is used by:
|
||||
- Consolidation service (Librarian knowledge consolidation)
|
||||
- Any other service that needs to create/update wiki pages
|
||||
"""
|
||||
import logging
|
||||
import json
|
||||
from typing import Dict, Any, List, Optional, Tuple
|
||||
from datetime import datetime
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
|
||||
class WikiPageWriter:
|
||||
"""
|
||||
Intelligent wiki page writer using LLM for content generation and restructuring.
|
||||
"""
|
||||
|
||||
def __init__(self, ollama_client):
|
||||
"""
|
||||
Initialize wiki page writer.
|
||||
|
||||
Args:
|
||||
ollama_client: OllamaClient for LLM operations
|
||||
"""
|
||||
self.ollama = ollama_client
|
||||
self.model = "mistral-nemo" # Default model for writing
|
||||
|
||||
async def create_page(
|
||||
self,
|
||||
title: str,
|
||||
topic_summary: str,
|
||||
source_information: List[Dict[str, str]],
|
||||
entities: Optional[List[str]] = None,
|
||||
related_docs: Optional[List[str]] = None
|
||||
) -> str:
|
||||
"""
|
||||
Create new wiki page with structured content.
|
||||
|
||||
Args:
|
||||
title: Page title
|
||||
topic_summary: Brief summary of the topic
|
||||
source_information: List of {title, url, content} dicts
|
||||
entities: Related entities from knowledge graph
|
||||
related_docs: Related documents/pages
|
||||
|
||||
Returns:
|
||||
Formatted markdown content
|
||||
"""
|
||||
logger.info(f"Creating wiki page: {title}")
|
||||
|
||||
# Build source context
|
||||
sources_text = self._format_sources_for_llm(source_information)
|
||||
|
||||
# Create page using LLM
|
||||
prompt = self._build_create_prompt(
|
||||
title=title,
|
||||
summary=topic_summary,
|
||||
sources=sources_text,
|
||||
entities=entities or [],
|
||||
related_docs=related_docs or []
|
||||
)
|
||||
|
||||
content = await self._call_llm(prompt)
|
||||
|
||||
# Post-process to ensure template compliance
|
||||
content = self._ensure_standard_sections(
|
||||
content=content,
|
||||
title=title,
|
||||
sources=source_information,
|
||||
entities=entities or [],
|
||||
related_docs=related_docs or []
|
||||
)
|
||||
|
||||
return content
|
||||
|
||||
async def reconstruct_page(
|
||||
self,
|
||||
title: str,
|
||||
existing_content: str,
|
||||
new_information: str,
|
||||
new_sources: List[Dict[str, str]],
|
||||
detect_conflicts: bool = True
|
||||
) -> Tuple[str, Optional[List[Dict[str, Any]]]]:
|
||||
"""
|
||||
Reconstruct existing page with new information.
|
||||
|
||||
Intelligently merges new content with existing, restructures for clarity,
|
||||
and detects factual conflicts.
|
||||
|
||||
Args:
|
||||
title: Page title
|
||||
existing_content: Current page content
|
||||
new_information: New information to integrate
|
||||
new_sources: Sources for new information
|
||||
detect_conflicts: Whether to detect and flag conflicts
|
||||
|
||||
Returns:
|
||||
Tuple of (reconstructed_content, conflicts)
|
||||
conflicts: List of detected conflicts or None
|
||||
"""
|
||||
logger.info(f"Reconstructing wiki page: {title}")
|
||||
|
||||
# Detect conflicts first
|
||||
conflicts = None
|
||||
if detect_conflicts:
|
||||
conflicts = await self._detect_conflicts(
|
||||
existing_content=existing_content,
|
||||
new_information=new_information
|
||||
)
|
||||
|
||||
if conflicts:
|
||||
logger.warning(f"Detected {len(conflicts)} potential conflicts in {title}")
|
||||
|
||||
# Build reconstruction prompt
|
||||
prompt = self._build_reconstruct_prompt(
|
||||
title=title,
|
||||
existing_content=existing_content,
|
||||
new_information=new_information,
|
||||
new_sources=self._format_sources_for_llm(new_sources),
|
||||
conflicts=conflicts
|
||||
)
|
||||
|
||||
# Reconstruct with LLM
|
||||
reconstructed = await self._call_llm(prompt)
|
||||
|
||||
# Ensure standard sections are present
|
||||
reconstructed = self._ensure_standard_sections(
|
||||
content=reconstructed,
|
||||
title=title,
|
||||
sources=new_sources,
|
||||
is_update=True
|
||||
)
|
||||
|
||||
return reconstructed, conflicts
|
||||
|
||||
async def _detect_conflicts(
|
||||
self,
|
||||
existing_content: str,
|
||||
new_information: str
|
||||
) -> Optional[List[Dict[str, Any]]]:
|
||||
"""
|
||||
Detect factual conflicts between existing and new content.
|
||||
|
||||
Returns:
|
||||
List of conflicts with: {fact_a, fact_b, confidence, context}
|
||||
"""
|
||||
prompt = f"""Analyze these two pieces of content for factual conflicts.
|
||||
|
||||
EXISTING CONTENT:
|
||||
{existing_content[:2000]}
|
||||
|
||||
NEW INFORMATION:
|
||||
{new_information[:2000]}
|
||||
|
||||
Identify any facts that contradict each other. For each conflict, provide:
|
||||
1. The fact from existing content
|
||||
2. The contradicting fact from new information
|
||||
3. Confidence level (low/medium/high)
|
||||
4. Context/explanation
|
||||
|
||||
Return ONLY valid JSON:
|
||||
{{
|
||||
"conflicts": [
|
||||
{{
|
||||
"existing_fact": "fact from old content",
|
||||
"new_fact": "contradicting fact",
|
||||
"confidence": "medium",
|
||||
"context": "explanation of why these conflict"
|
||||
}}
|
||||
]
|
||||
}}
|
||||
|
||||
If no conflicts, return: {{"conflicts": []}}
|
||||
|
||||
JSON:"""
|
||||
|
||||
try:
|
||||
response = await self.ollama.generate_text(
|
||||
prompt=prompt,
|
||||
model=self.model,
|
||||
stream=False
|
||||
)
|
||||
|
||||
# Extract JSON
|
||||
response_clean = response.strip()
|
||||
if '{' in response_clean:
|
||||
json_start = response_clean.find('{')
|
||||
json_end = response_clean.rfind('}') + 1
|
||||
response_clean = response_clean[json_start:json_end]
|
||||
|
||||
result = json.loads(response_clean)
|
||||
conflicts = result.get('conflicts', [])
|
||||
|
||||
return conflicts if conflicts else None
|
||||
|
||||
except Exception as e:
|
||||
logger.error(f"Conflict detection failed: {e}")
|
||||
return None
|
||||
|
||||
def _build_create_prompt(
|
||||
self,
|
||||
title: str,
|
||||
summary: str,
|
||||
sources: str,
|
||||
entities: List[str],
|
||||
related_docs: List[str]
|
||||
) -> str:
|
||||
"""Build LLM prompt for creating new page."""
|
||||
return f"""You are a Librarian creating a dossier for a personal knowledge base and extended memory system.
|
||||
|
||||
Create a comprehensive, well-structured wiki page with appropriate sections for the content type.
|
||||
|
||||
TOPIC: {title}
|
||||
|
||||
SUMMARY: {summary}
|
||||
|
||||
SOURCE INFORMATION:
|
||||
{sources}
|
||||
|
||||
RELATED ENTITIES: {', '.join(entities) if entities else 'None'}
|
||||
|
||||
RELATED DOCUMENTS: {', '.join(related_docs) if related_docs else 'None'}
|
||||
|
||||
CONTENT TYPE GUIDELINES (Schema.org-aligned):
|
||||
|
||||
For PEOPLE (family, friends, colleagues, public figures) (Schema.org: Person):
|
||||
- Executive Summary (who they are, key facts)
|
||||
- Background & Biography
|
||||
- Relationships & Connections
|
||||
- Professional Info / Career
|
||||
- Interests & Preferences
|
||||
- Important Dates & Events
|
||||
- Notes & Observations
|
||||
|
||||
For COMPANIES (businesses, organizations, startups) (Schema.org: Organization):
|
||||
- Executive Summary (what they do, industry, key facts)
|
||||
- Overview & Mission
|
||||
- Products & Services
|
||||
- History & Milestones
|
||||
- Leadership & Team
|
||||
- Personal Connection / Experience
|
||||
- Notable Projects or Achievements
|
||||
|
||||
For PLACES (locations, restaurants, destinations) (Schema.org: Place):
|
||||
- Executive Summary (what/where, key details)
|
||||
- Location & How to Get There
|
||||
- Description & Atmosphere
|
||||
- Features & Amenities
|
||||
- Personal Experiences / Visits
|
||||
- Recommendations & Tips
|
||||
|
||||
For ENTERTAINMENT (books, movies, TV, music, games) (Schema.org: CreativeWork):
|
||||
- Executive Summary (title, creator, key facts)
|
||||
- Synopsis / Overview
|
||||
- Key Characters / Themes
|
||||
- Personal Thoughts & Ratings
|
||||
- Memorable Moments / Quotes
|
||||
- Related Works
|
||||
|
||||
For RECIPES & FOOD (Schema.org: Recipe):
|
||||
- Executive Summary (dish name, cuisine type)
|
||||
- Ingredients (formatted as table or list)
|
||||
- Instructions (step-by-step)
|
||||
- Cooking Tips & Variations
|
||||
- Personal Notes & Modifications
|
||||
- Source / Origin
|
||||
|
||||
For PRODUCTS (gear, tools, purchases) (Schema.org: Product):
|
||||
- Executive Summary (what it is, brand/model, key specs)
|
||||
- Overview & Purpose
|
||||
- Specifications (formatted as table)
|
||||
- Purchase Information (where, when, price)
|
||||
- Personal Experience / Review
|
||||
- Maintenance & Care
|
||||
- Related Products / Alternatives
|
||||
|
||||
For TECHNOLOGY (software, applications, infrastructure) (Schema.org: SoftwareApplication):
|
||||
- Executive Summary (what it is, key facts)
|
||||
- Overview & Purpose
|
||||
- Technical Details (tables for specs)
|
||||
- Setup & Configuration
|
||||
- Use Cases & Applications
|
||||
- Best Practices
|
||||
- Common Issues & Solutions
|
||||
|
||||
For EVENTS (concerts, travel, appointments) (Schema.org: Event):
|
||||
- Executive Summary (what, when, where)
|
||||
- Event Details (date, time, location, venue)
|
||||
- Participants / Attendees
|
||||
- Planning & Preparation
|
||||
- Experience / Highlights
|
||||
- Photos / Media
|
||||
- Notes & Reflections
|
||||
|
||||
For HEALTH (medical, fitness, wellness) (Schema.org: MedicalEntity):
|
||||
- Executive Summary (condition/topic, key facts)
|
||||
- Overview & Background
|
||||
- Symptoms / Signs / Characteristics
|
||||
- Treatments / Approaches / Recommendations
|
||||
- Personal Experience / Progress
|
||||
- Resources & References
|
||||
- Important Dates (appointments, changes)
|
||||
|
||||
For HOBBIES (activities, interests, pastimes) (Custom extension):
|
||||
- Executive Summary (what it is, why interesting)
|
||||
- Getting Started / Basics
|
||||
- Equipment & Materials
|
||||
- Techniques & Skills
|
||||
- Personal Progress / Achievements
|
||||
- Resources & Communities
|
||||
- Goals & Future Plans
|
||||
|
||||
For PROJECTS (work projects, personal projects) (Schema.org: Project):
|
||||
- Executive Summary (what, why, status)
|
||||
- Goals & Objectives
|
||||
- Timeline & Milestones
|
||||
- Team / Collaborators
|
||||
- Technical Details / Architecture
|
||||
- Current Status & Next Steps
|
||||
- Lessons Learned / Reflections
|
||||
|
||||
For REFERENCE (general knowledge, how-tos) (Custom extension):
|
||||
- Executive Summary
|
||||
- Overview & Context
|
||||
- Key Concepts & Definitions
|
||||
- Step-by-Step Guide (if applicable)
|
||||
- Examples & Use Cases
|
||||
- Tips & Best Practices
|
||||
- Related Topics & Further Reading
|
||||
|
||||
FORMATTING RULES:
|
||||
- Use markdown headers (##, ###)
|
||||
- Create tables for structured data (ingredients, specs, comparisons)
|
||||
- Use bullet points for lists
|
||||
- Include code blocks with ``` where applicable
|
||||
- Bold important terms
|
||||
- Keep sections focused and scannable
|
||||
- Adapt structure to content - not all sections apply to all topics
|
||||
|
||||
Generate ONLY the markdown content (do not include Sources, Knowledge Graph, or Mind Map sections - those are added automatically).
|
||||
|
||||
MARKDOWN:"""
|
||||
|
||||
def _build_reconstruct_prompt(
|
||||
self,
|
||||
title: str,
|
||||
existing_content: str,
|
||||
new_information: str,
|
||||
new_sources: str,
|
||||
conflicts: Optional[List[Dict[str, Any]]]
|
||||
) -> str:
|
||||
"""Build LLM prompt for reconstructing page."""
|
||||
conflicts_note = ""
|
||||
if conflicts:
|
||||
conflicts_note = "\n\nDETECTED CONFLICTS:\n"
|
||||
for i, c in enumerate(conflicts, 1):
|
||||
conflicts_note += f"{i}. Existing: '{c['existing_fact']}'\n"
|
||||
conflicts_note += f" New: '{c['new_fact']}'\n"
|
||||
conflicts_note += f" Confidence: {c['confidence']}\n"
|
||||
conflicts_note += f" Note: {c['context']}\n\n"
|
||||
conflicts_note += "IMPORTANT: For conflicts, prefer the most recent/authoritative source. Add a note in 'Changes & Updates' section when facts are superseded.\n"
|
||||
|
||||
return f"""Reconstruct this wiki page by intelligently merging new information with existing content.
|
||||
|
||||
TITLE: {title}
|
||||
|
||||
EXISTING CONTENT:
|
||||
{existing_content}
|
||||
|
||||
NEW INFORMATION TO INTEGRATE:
|
||||
{new_information}
|
||||
|
||||
NEW SOURCES:
|
||||
{new_sources}
|
||||
{conflicts_note}
|
||||
|
||||
RECONSTRUCTION REQUIREMENTS:
|
||||
1. **Zero Fact Loss**: Preserve ALL facts from existing content unless superseded
|
||||
2. **Holistic Restructuring**: Reorganize for better flow and clarity
|
||||
3. **Conflict Resolution**: When facts conflict, choose most authoritative/recent
|
||||
4. **Professional Structure**:
|
||||
- Update Executive Summary with key facts
|
||||
- Organize into clear chapters
|
||||
- Use tables for specifications/comparisons
|
||||
- Maintain consistent formatting
|
||||
5. **Update Tracking**: Add entry to "Changes & Updates" section with today's date
|
||||
|
||||
FORMATTING RULES:
|
||||
- Maintain markdown structure
|
||||
- Use tables for data (| col1 | col2 |)
|
||||
- Keep existing good structure, improve where needed
|
||||
- Bold important terms
|
||||
- Add subsections (###) where it improves clarity
|
||||
|
||||
OUTPUT INSTRUCTIONS:
|
||||
- Return complete page content (do not include Sources, Knowledge Graph, Mind Map - those are added automatically)
|
||||
- Include updated "Changes & Updates" section noting what was changed today
|
||||
- If facts were superseded, note it clearly
|
||||
|
||||
RECONSTRUCTED MARKDOWN:"""
|
||||
|
||||
async def _call_llm(self, prompt: str) -> str:
|
||||
"""Call LLM with prompt and return response."""
|
||||
try:
|
||||
response = await self.ollama.generate_text(
|
||||
prompt=prompt,
|
||||
model=self.model,
|
||||
stream=False
|
||||
)
|
||||
|
||||
if not response:
|
||||
raise Exception("Empty response from LLM")
|
||||
|
||||
return response.strip()
|
||||
|
||||
except Exception as e:
|
||||
logger.error(f"LLM call failed: {e}")
|
||||
raise
|
||||
|
||||
def _format_sources_for_llm(self, sources: List[Dict[str, str]]) -> str:
|
||||
"""Format source information for LLM prompt."""
|
||||
formatted = []
|
||||
for i, source in enumerate(sources, 1):
|
||||
formatted.append(f"[{i}] {source.get('title', 'Untitled')}")
|
||||
formatted.append(f" URL: {source.get('url', 'N/A')}")
|
||||
content = source.get('content', '')[:500] # Limit content length
|
||||
formatted.append(f" Content: {content}...\n")
|
||||
|
||||
return "\n".join(formatted)
|
||||
|
||||
def _ensure_standard_sections(
|
||||
self,
|
||||
content: str,
|
||||
title: str,
|
||||
sources: List[Dict[str, str]],
|
||||
entities: Optional[List[str]] = None,
|
||||
related_docs: Optional[List[str]] = None,
|
||||
is_update: bool = False
|
||||
) -> str:
|
||||
"""
|
||||
Ensure page has standard footer sections (Sources, Knowledge Graph, Mind Map).
|
||||
|
||||
These sections are standardized and appended automatically.
|
||||
"""
|
||||
# Remove any existing standard sections
|
||||
for section in ["## Sources", "## Knowledge Graph", "## Mind Map"]:
|
||||
if section in content:
|
||||
content = content.split(section)[0]
|
||||
|
||||
# Add horizontal rule before footer
|
||||
content = content.rstrip() + "\n\n---\n\n"
|
||||
|
||||
# Add Sources section
|
||||
content += "## Sources\n\n"
|
||||
if sources:
|
||||
for i, source in enumerate(sources, 1):
|
||||
content += f"{i}. [{source.get('title', 'Source')}]({source.get('url', '#')})\n"
|
||||
else:
|
||||
content += "*No sources listed*\n"
|
||||
|
||||
# Add Knowledge Graph section
|
||||
content += "\n## Knowledge Graph\n\n"
|
||||
if entities:
|
||||
content += "**Related Entities:**\n"
|
||||
for entity in entities[:10]: # Limit to 10
|
||||
content += f"- {entity}\n"
|
||||
else:
|
||||
content += "*No entities linked yet*\n"
|
||||
|
||||
content += "\n**View in Neo4j:** [Explore Graph](/graph)\n"
|
||||
|
||||
# Add Mind Map section
|
||||
content += "\n## Mind Map\n\n"
|
||||
content += f"**Interactive Mind Map:** [View Topic Map](/mindmap?topic={title.replace(' ', '+')})\n"
|
||||
|
||||
# Add footer metadata
|
||||
content += "\n---\n\n"
|
||||
timestamp = datetime.now().strftime('%Y-%m-%d %H:%M')
|
||||
action = "Updated" if is_update else "Created"
|
||||
content += f"*{action}: {timestamp} | Generated by: Librarian Agent* \n"
|
||||
content += "*Template: Library Desk Wiki Standard v1.0*\n"
|
||||
|
||||
return content
|
||||
@@ -1,433 +0,0 @@
|
||||
"""
|
||||
Wiki service layer for Library Desk.
|
||||
|
||||
Handles business logic for wiki operations with:
|
||||
- Multi-tenant path scoping
|
||||
- Dossier management (tag-based)
|
||||
- Page CRUD operations
|
||||
- Search functionality
|
||||
"""
|
||||
|
||||
from typing import List, Optional, Dict, Any
|
||||
import logging
|
||||
|
||||
from src.clients.wikijs_client import WikiJSClient
|
||||
from src.core.multi_tenancy import get_wikijs_namespace, validate_user_id, DEFAULT_USER
|
||||
from src.models.wiki import (
|
||||
WikiPage, WikiPageSummary, WikiPageList,
|
||||
WikiPageCreate, WikiPageUpdate,
|
||||
DossierInfo, DossierList
|
||||
)
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
|
||||
class WikiService:
|
||||
"""
|
||||
Service layer for wiki operations.
|
||||
|
||||
Responsibilities:
|
||||
- Enforce multi-tenant path scoping
|
||||
- Convert between client and API models
|
||||
- Handle dossier (tag) operations
|
||||
- Provide business logic layer
|
||||
"""
|
||||
|
||||
def __init__(self, wiki_client: WikiJSClient):
|
||||
"""
|
||||
Initialize wiki service.
|
||||
|
||||
Args:
|
||||
wiki_client: Initialized Wiki.js client
|
||||
"""
|
||||
self.wiki_client = wiki_client
|
||||
|
||||
def _get_user_namespace(self, user: str) -> str:
|
||||
"""
|
||||
Get user's wiki namespace with validation.
|
||||
|
||||
Args:
|
||||
user: User identifier
|
||||
|
||||
Returns:
|
||||
Wiki.js namespace path
|
||||
|
||||
Raises:
|
||||
ValueError: If user ID is invalid
|
||||
"""
|
||||
if not validate_user_id(user):
|
||||
raise ValueError(f"Invalid user ID: {user}")
|
||||
return get_wikijs_namespace(user)
|
||||
|
||||
def _ensure_user_path(self, path: str, user: str) -> str:
|
||||
"""
|
||||
Ensure path is within user's namespace.
|
||||
|
||||
Args:
|
||||
path: Requested page path
|
||||
user: User identifier
|
||||
|
||||
Returns:
|
||||
Full path within user namespace
|
||||
|
||||
Example:
|
||||
>>> self._ensure_user_path("/projects/foo", "jpmschweitzer")
|
||||
'/users/jpmschweitzer/projects/foo'
|
||||
"""
|
||||
namespace = self._get_user_namespace(user)
|
||||
|
||||
# If path already starts with namespace, return as-is
|
||||
if path.startswith(namespace):
|
||||
return path
|
||||
|
||||
# Remove leading slash from path if present
|
||||
path = path.lstrip("/")
|
||||
|
||||
# Combine namespace and path
|
||||
return f"{namespace}/{path}"
|
||||
|
||||
async def list_pages(
|
||||
self,
|
||||
user: str,
|
||||
tag: Optional[str] = None,
|
||||
limit: int = 50
|
||||
) -> WikiPageList:
|
||||
"""
|
||||
List pages for a user, optionally filtered by tag.
|
||||
|
||||
Args:
|
||||
user: User identifier
|
||||
tag: Optional tag filter (dossier)
|
||||
limit: Maximum pages to return
|
||||
|
||||
Returns:
|
||||
WikiPageList with pages and metadata
|
||||
"""
|
||||
namespace = self._get_user_namespace(user)
|
||||
|
||||
# Get pages with filtering
|
||||
pages = await self.wiki_client.list_pages(
|
||||
path_prefix=namespace,
|
||||
tags=[tag] if tag else None,
|
||||
limit=limit
|
||||
)
|
||||
|
||||
# Convert to summary format
|
||||
summaries = [
|
||||
WikiPageSummary(
|
||||
id=p["id"],
|
||||
path=p["path"],
|
||||
title=p["title"],
|
||||
description=p.get("description"),
|
||||
tags=p.get("tags", []),
|
||||
updated_at=p.get("updatedAt"),
|
||||
is_published=p.get("isPublished", True)
|
||||
)
|
||||
for p in pages
|
||||
]
|
||||
|
||||
return WikiPageList(
|
||||
pages=summaries,
|
||||
total=len(summaries),
|
||||
filtered_by_tag=tag,
|
||||
user=user
|
||||
)
|
||||
|
||||
async def get_page(self, page_id: int, user: str) -> Optional[WikiPage]:
|
||||
"""
|
||||
Get a single page by ID.
|
||||
|
||||
Args:
|
||||
page_id: Page ID
|
||||
user: User identifier (for validation)
|
||||
|
||||
Returns:
|
||||
WikiPage or None if not found or access denied
|
||||
|
||||
Note: Validates that page belongs to user's namespace
|
||||
"""
|
||||
page = await self.wiki_client.get_page(page_id)
|
||||
|
||||
if not page:
|
||||
return None
|
||||
|
||||
# Validate page is in user's namespace
|
||||
namespace = self._get_user_namespace(user)
|
||||
page_path = "/" + page["path"].lstrip("/") # Normalize path with leading slash
|
||||
if not page_path.startswith(namespace):
|
||||
logger.warning(f"User {user} attempted to access page outside namespace: {page['path']}")
|
||||
return None
|
||||
|
||||
return WikiPage(
|
||||
id=page["id"],
|
||||
path=page["path"],
|
||||
title=page["title"],
|
||||
description=page.get("description"),
|
||||
content=page.get("content"),
|
||||
tags=page.get("tags", []),
|
||||
created_at=page.get("createdAt"),
|
||||
updated_at=page.get("updatedAt"),
|
||||
is_published=page.get("isPublished", True),
|
||||
editor=page.get("editor")
|
||||
)
|
||||
|
||||
async def create_page(self, page_data: WikiPageCreate) -> WikiPage:
|
||||
"""
|
||||
Create a new wiki page.
|
||||
|
||||
Args:
|
||||
page_data: Page creation data
|
||||
|
||||
Returns:
|
||||
Created WikiPage
|
||||
|
||||
Raises:
|
||||
ValueError: If creation fails
|
||||
"""
|
||||
user = page_data.user or DEFAULT_USER
|
||||
|
||||
# Ensure path is in user's namespace
|
||||
full_path = self._ensure_user_path(page_data.path, user)
|
||||
|
||||
try:
|
||||
created = await self.wiki_client.create_page(
|
||||
path=full_path,
|
||||
title=page_data.title,
|
||||
content=page_data.content,
|
||||
description=page_data.description or "",
|
||||
tags=page_data.tags,
|
||||
is_published=page_data.is_published,
|
||||
editor=page_data.editor
|
||||
)
|
||||
|
||||
# Fetch full page details
|
||||
page = await self.wiki_client.get_page(created["id"])
|
||||
if not page:
|
||||
raise ValueError("Page created but could not be retrieved")
|
||||
|
||||
return WikiPage(
|
||||
id=page["id"],
|
||||
path=page["path"],
|
||||
title=page["title"],
|
||||
description=page.get("description"),
|
||||
content=page.get("content"),
|
||||
tags=page.get("tags", []),
|
||||
created_at=page.get("createdAt"),
|
||||
updated_at=page.get("updatedAt"),
|
||||
is_published=page.get("isPublished", True),
|
||||
editor=page.get("editor")
|
||||
)
|
||||
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to create page: {e}", exc_info=True)
|
||||
raise ValueError(f"Failed to create page: {str(e)}")
|
||||
|
||||
async def update_page(
|
||||
self,
|
||||
page_id: int,
|
||||
page_data: WikiPageUpdate,
|
||||
user: str
|
||||
) -> WikiPage:
|
||||
"""
|
||||
Update an existing page.
|
||||
|
||||
Args:
|
||||
page_id: Page ID to update
|
||||
page_data: Update data
|
||||
user: User identifier (for validation)
|
||||
|
||||
Returns:
|
||||
Updated WikiPage
|
||||
|
||||
Raises:
|
||||
ValueError: If page not found or update fails
|
||||
"""
|
||||
# Verify page exists and belongs to user
|
||||
existing = await self.get_page(page_id, user)
|
||||
if not existing:
|
||||
raise ValueError(f"Page {page_id} not found or access denied")
|
||||
|
||||
try:
|
||||
await self.wiki_client.update_page(
|
||||
page_id=page_id,
|
||||
content=page_data.content,
|
||||
title=page_data.title,
|
||||
description=page_data.description,
|
||||
tags=page_data.tags,
|
||||
is_published=True # Always keep pages published for internal wiki
|
||||
)
|
||||
|
||||
# Fetch updated page
|
||||
updated = await self.get_page(page_id, user)
|
||||
if not updated:
|
||||
raise ValueError("Page updated but could not be retrieved")
|
||||
|
||||
return updated
|
||||
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to update page {page_id}: {e}", exc_info=True)
|
||||
raise ValueError(f"Failed to update page: {str(e)}")
|
||||
|
||||
async def delete_page(self, page_id: int, user: str) -> bool:
|
||||
"""
|
||||
Delete a page.
|
||||
|
||||
Args:
|
||||
page_id: Page ID to delete
|
||||
user: User identifier (for validation)
|
||||
|
||||
Returns:
|
||||
True if deleted successfully
|
||||
|
||||
Raises:
|
||||
ValueError: If page not found or deletion fails
|
||||
"""
|
||||
# Verify page exists and belongs to user
|
||||
existing = await self.get_page(page_id, user)
|
||||
if not existing:
|
||||
raise ValueError(f"Page {page_id} not found or access denied")
|
||||
|
||||
try:
|
||||
await self.wiki_client.delete_page(page_id)
|
||||
logger.info(f"Deleted page {page_id} for user {user}")
|
||||
return True
|
||||
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to delete page {page_id}: {e}", exc_info=True)
|
||||
raise ValueError(f"Failed to delete page: {str(e)}")
|
||||
|
||||
async def search_pages(
|
||||
self,
|
||||
query: str,
|
||||
user: str,
|
||||
limit: int = 20
|
||||
) -> List[WikiPageSummary]:
|
||||
"""
|
||||
Search pages in user's namespace.
|
||||
|
||||
Args:
|
||||
query: Search query
|
||||
user: User identifier
|
||||
limit: Maximum results
|
||||
|
||||
Returns:
|
||||
List of matching pages
|
||||
"""
|
||||
namespace = self._get_user_namespace(user)
|
||||
|
||||
results = await self.wiki_client.search_pages(
|
||||
query=query,
|
||||
path_prefix=namespace
|
||||
)
|
||||
|
||||
# Convert to summaries (limit results)
|
||||
return [
|
||||
WikiPageSummary(
|
||||
id=r["id"],
|
||||
path=r["path"],
|
||||
title=r["title"],
|
||||
description=r.get("description"),
|
||||
tags=[], # Search results don't include tags
|
||||
updated_at=None,
|
||||
is_published=True
|
||||
)
|
||||
for r in results[:limit]
|
||||
]
|
||||
|
||||
async def move_page(
|
||||
self,
|
||||
page_id: int,
|
||||
new_path: str,
|
||||
user: str
|
||||
) -> bool:
|
||||
"""
|
||||
Move/rename a page.
|
||||
|
||||
Args:
|
||||
page_id: Page ID to move
|
||||
new_path: New path (within user namespace)
|
||||
user: User identifier
|
||||
|
||||
Returns:
|
||||
True if moved successfully
|
||||
|
||||
Raises:
|
||||
ValueError: If operation fails
|
||||
"""
|
||||
# Verify page exists and belongs to user
|
||||
existing = await self.get_page(page_id, user)
|
||||
if not existing:
|
||||
raise ValueError(f"Page {page_id} not found or access denied")
|
||||
|
||||
# Ensure new path is in user's namespace
|
||||
full_new_path = self._ensure_user_path(new_path, user)
|
||||
|
||||
try:
|
||||
success = await self.wiki_client.move_page(page_id, full_new_path)
|
||||
if success:
|
||||
logger.info(f"Moved page {page_id} to {full_new_path}")
|
||||
return success
|
||||
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to move page {page_id}: {e}", exc_info=True)
|
||||
raise ValueError(f"Failed to move page: {str(e)}")
|
||||
|
||||
# Dossier operations (tag-based)
|
||||
|
||||
async def list_dossiers(self, user: str) -> DossierList:
|
||||
"""
|
||||
List all dossiers (unique tags) for a user.
|
||||
|
||||
Args:
|
||||
user: User identifier
|
||||
|
||||
Returns:
|
||||
DossierList with all dossiers
|
||||
"""
|
||||
# Get all pages for user
|
||||
pages = await self.list_pages(user, limit=1000)
|
||||
|
||||
# Collect unique tags
|
||||
tag_counts: Dict[str, int] = {}
|
||||
for page in pages.pages:
|
||||
for tag in page.tags:
|
||||
tag_counts[tag] = tag_counts.get(tag, 0) + 1
|
||||
|
||||
# Create dossier info for each tag
|
||||
dossiers = [
|
||||
DossierInfo(
|
||||
name=tag,
|
||||
title=tag.replace("-", " ").title(),
|
||||
description=f"Dossier for {tag}",
|
||||
page_count=count,
|
||||
index_page_id=None,
|
||||
index_page_path=None,
|
||||
created_at=None
|
||||
)
|
||||
for tag, count in tag_counts.items()
|
||||
]
|
||||
|
||||
return DossierList(
|
||||
dossiers=sorted(dossiers, key=lambda d: d.page_count, reverse=True),
|
||||
total=len(dossiers),
|
||||
user=user
|
||||
)
|
||||
|
||||
async def get_dossier_pages(
|
||||
self,
|
||||
dossier_name: str,
|
||||
user: str,
|
||||
limit: int = 100
|
||||
) -> WikiPageList:
|
||||
"""
|
||||
Get all pages in a dossier (by tag).
|
||||
|
||||
Args:
|
||||
dossier_name: Dossier name (tag)
|
||||
user: User identifier
|
||||
limit: Maximum pages
|
||||
|
||||
Returns:
|
||||
WikiPageList filtered by dossier tag
|
||||
"""
|
||||
return await self.list_pages(user, tag=dossier_name, limit=limit)
|
||||
@@ -1,569 +0,0 @@
|
||||
/**
|
||||
* Library Desk Integration for Wiki.js
|
||||
* Combined re-index and entity linking buttons
|
||||
*
|
||||
* Usage: Add to Wiki.js Code Injection:
|
||||
* <script src="http://192.168.86.149:8089/static/wikijs-integration.js"></script>
|
||||
*/
|
||||
(function() {
|
||||
'use strict';
|
||||
|
||||
// Auto-detect Library Desk URL
|
||||
const scriptTag = document.currentScript;
|
||||
const scriptUrl = scriptTag ? scriptTag.src : '';
|
||||
const libraryDeskUrl = scriptUrl ? scriptUrl.split('/static/')[0] : 'http://192.168.86.149:8089';
|
||||
|
||||
// Shared configuration
|
||||
const CONFIG = window.LIBRARY_DESK_CONFIG || {
|
||||
libraryDeskUrl: libraryDeskUrl,
|
||||
apiKey: 'af88ed8f44bed81bdb20d0534f1c4547340b29e2aba4963f61a71b993d7eb6e5',
|
||||
user: 'jpmschweitzer',
|
||||
buttonPosition: 'toolbar', // 'toolbar' or 'floating'
|
||||
debug: true
|
||||
};
|
||||
|
||||
function log() {
|
||||
if (CONFIG.debug) {
|
||||
console.log('[Library Desk]', ...arguments);
|
||||
}
|
||||
}
|
||||
|
||||
// Store page ID globally once fetched
|
||||
let CACHED_PAGE_ID = null;
|
||||
let CACHED_PAGE_PATH = null;
|
||||
|
||||
// Initialize when DOM is ready
|
||||
if (document.readyState === 'loading') {
|
||||
document.addEventListener('DOMContentLoaded', init);
|
||||
} else {
|
||||
init();
|
||||
}
|
||||
|
||||
function init() {
|
||||
log('Initializing Library Desk integration...');
|
||||
log('Library Desk URL:', CONFIG.libraryDeskUrl);
|
||||
|
||||
// Wait for page to be ready, then add buttons
|
||||
setTimeout(addButtons, 1500);
|
||||
}
|
||||
|
||||
async function addButtons() {
|
||||
const pagePath = getPagePath();
|
||||
if (!pagePath) {
|
||||
log('Skipping buttons - not on a content page');
|
||||
return;
|
||||
}
|
||||
|
||||
log('Page path:', pagePath);
|
||||
|
||||
// Check if buttons already exist
|
||||
if (document.getElementById('library-desk-buttons')) {
|
||||
log('Buttons already exist');
|
||||
return;
|
||||
}
|
||||
|
||||
// Fetch and cache page ID during initialization
|
||||
CACHED_PAGE_PATH = pagePath;
|
||||
CACHED_PAGE_ID = await getPageIdFromPath(pagePath);
|
||||
log('Cached page ID for this session:', CACHED_PAGE_ID);
|
||||
|
||||
if (CONFIG.buttonPosition === 'floating') {
|
||||
addFloatingButtons(pagePath);
|
||||
} else {
|
||||
addToolbarButtons(pagePath);
|
||||
}
|
||||
}
|
||||
|
||||
function addToolbarButtons(pagePath) {
|
||||
const toolbar = findToolbar();
|
||||
if (!toolbar) {
|
||||
log('Toolbar not found, falling back to floating buttons');
|
||||
addFloatingButtons(pagePath);
|
||||
return;
|
||||
}
|
||||
|
||||
// Container for both buttons
|
||||
const btnContainer = document.createElement('div');
|
||||
btnContainer.id = 'library-desk-buttons';
|
||||
btnContainer.style.cssText = 'display: inline-flex; align-items: center; gap: 4px; margin-left: 8px;';
|
||||
|
||||
// Re-index button
|
||||
const reindexBtn = createToolbarButton(
|
||||
'library-desk-reindex-btn',
|
||||
'Re-index this page in Library Desk (vectors + knowledge graph)',
|
||||
'mdi-database-sync',
|
||||
() => reindexPage(pagePath, reindexBtn)
|
||||
);
|
||||
|
||||
// Entity link button
|
||||
const entityLinkBtn = createToolbarButton(
|
||||
'library-desk-entitylink-btn',
|
||||
'Link entities mentioned in this page to the knowledge graph',
|
||||
'mdi-graph-outline',
|
||||
() => linkEntities(pagePath, entityLinkBtn)
|
||||
);
|
||||
|
||||
btnContainer.appendChild(reindexBtn);
|
||||
btnContainer.appendChild(entityLinkBtn);
|
||||
toolbar.appendChild(btnContainer);
|
||||
log('Toolbar buttons added');
|
||||
}
|
||||
|
||||
function createToolbarButton(id, title, icon, onClick) {
|
||||
const button = document.createElement('button');
|
||||
button.id = id;
|
||||
button.className = 'v-btn v-btn--icon v-btn--round theme--dark v-size--default';
|
||||
button.type = 'button';
|
||||
button.title = title;
|
||||
button.setAttribute('aria-label', title);
|
||||
|
||||
button.innerHTML = `<span class="v-btn__content"><i class="v-icon notranslate mdi ${icon} theme--dark" style="font-size: 20px;"></i></span>`;
|
||||
|
||||
button.addEventListener('mouseenter', function() {
|
||||
this.style.backgroundColor = 'rgba(255, 255, 255, 0.08)';
|
||||
});
|
||||
button.addEventListener('mouseleave', function() {
|
||||
this.style.backgroundColor = '';
|
||||
});
|
||||
|
||||
button.onclick = function(e) {
|
||||
e.preventDefault();
|
||||
e.stopPropagation();
|
||||
onClick();
|
||||
};
|
||||
|
||||
return button;
|
||||
}
|
||||
|
||||
function addFloatingButtons(pagePath) {
|
||||
const container = document.createElement('div');
|
||||
container.id = 'library-desk-buttons';
|
||||
container.style.cssText = `
|
||||
position: fixed;
|
||||
bottom: 20px;
|
||||
right: 20px;
|
||||
display: flex;
|
||||
flex-direction: column;
|
||||
gap: 10px;
|
||||
z-index: 9999;
|
||||
`;
|
||||
|
||||
// Re-index button
|
||||
const reindexBtn = createFloatingButton(
|
||||
'library-desk-reindex-btn',
|
||||
'🔄 Re-index',
|
||||
'Re-index this page in Library Desk',
|
||||
'#1976d2',
|
||||
'#1565c0',
|
||||
() => reindexPage(pagePath, reindexBtn)
|
||||
);
|
||||
|
||||
// Entity link button
|
||||
const entityLinkBtn = createFloatingButton(
|
||||
'library-desk-entitylink-btn',
|
||||
'🔗 Link Entities',
|
||||
'Link entities in this page',
|
||||
'#43a047',
|
||||
'#388e3c',
|
||||
() => linkEntities(pagePath, entityLinkBtn)
|
||||
);
|
||||
|
||||
container.appendChild(reindexBtn);
|
||||
container.appendChild(entityLinkBtn);
|
||||
document.body.appendChild(container);
|
||||
log('Floating buttons added');
|
||||
}
|
||||
|
||||
function createFloatingButton(id, text, title, bgColor, hoverColor, onClick) {
|
||||
const button = document.createElement('button');
|
||||
button.id = id;
|
||||
button.innerHTML = text;
|
||||
button.title = title;
|
||||
button.style.cssText = `
|
||||
padding: 10px 15px;
|
||||
background: ${bgColor};
|
||||
color: white;
|
||||
border: none;
|
||||
border-radius: 4px;
|
||||
cursor: pointer;
|
||||
font-size: 14px;
|
||||
font-weight: 500;
|
||||
box-shadow: 0 2px 8px rgba(0,0,0,0.3);
|
||||
transition: all 0.3s;
|
||||
`;
|
||||
|
||||
button.onmouseover = function() { this.style.background = hoverColor; };
|
||||
button.onmouseout = function() { this.style.background = bgColor; };
|
||||
button.onclick = onClick;
|
||||
|
||||
return button;
|
||||
}
|
||||
|
||||
// ============================================================================
|
||||
// RE-INDEX FUNCTIONALITY
|
||||
// ============================================================================
|
||||
|
||||
async function reindexPage(pagePath, button) {
|
||||
const isFloating = button.id === 'library-desk-reindex-btn' && button.innerHTML.includes('Re-index');
|
||||
const icon = isFloating ? null : button.querySelector('.v-icon');
|
||||
const originalContent = isFloating ? button.innerHTML : null;
|
||||
const originalIcon = icon ? icon.className : null;
|
||||
|
||||
button.disabled = true;
|
||||
|
||||
if (isFloating) {
|
||||
button.innerHTML = '⏳ Loading...';
|
||||
button.style.background = '#757575';
|
||||
} else {
|
||||
icon.className = 'v-icon notranslate mdi mdi-loading mdi-spin theme--dark';
|
||||
}
|
||||
|
||||
try {
|
||||
// Use cached page ID from initialization
|
||||
if (!CACHED_PAGE_ID) {
|
||||
throw new Error('Page ID not available - was not fetched during initialization');
|
||||
}
|
||||
|
||||
log('Re-indexing page', CACHED_PAGE_ID, '...');
|
||||
|
||||
// Re-index directly
|
||||
const response = await fetch(CONFIG.libraryDeskUrl + '/ingest/page', {
|
||||
method: 'POST',
|
||||
headers: {
|
||||
'Authorization': 'Bearer ' + CONFIG.apiKey,
|
||||
'Content-Type': 'application/json'
|
||||
},
|
||||
body: JSON.stringify({
|
||||
user: CONFIG.user,
|
||||
page_id: CACHED_PAGE_ID,
|
||||
force_refresh: true
|
||||
})
|
||||
});
|
||||
|
||||
const result = await response.json();
|
||||
|
||||
if (response.ok && result.success) {
|
||||
if (isFloating) {
|
||||
button.innerHTML = '✓ Done!';
|
||||
button.style.background = '#4caf50';
|
||||
} else {
|
||||
icon.className = 'v-icon notranslate mdi mdi-check-circle theme--dark';
|
||||
icon.style.color = '#4caf50';
|
||||
}
|
||||
|
||||
const message = 'Re-indexed: ' + result.page_title + '\n\n' +
|
||||
'✓ Vectors: ' + result.vector_chunks_created + ' chunks\n' +
|
||||
'✓ Entities: ' + result.graph_entities_extracted + '\n' +
|
||||
'✓ Relationships: ' + result.graph_relationships_created + '\n' +
|
||||
'⏱ Time: ' + Math.round(result.processing_time_ms) + 'ms';
|
||||
|
||||
log('Success:', message);
|
||||
|
||||
// Show notification
|
||||
showNotification('success', 'Page Re-indexed', message);
|
||||
|
||||
// Reset button
|
||||
setTimeout(function() {
|
||||
if (isFloating) {
|
||||
button.innerHTML = originalContent;
|
||||
button.style.background = '#1976d2';
|
||||
} else {
|
||||
icon.className = originalIcon;
|
||||
icon.style.color = '';
|
||||
}
|
||||
button.disabled = false;
|
||||
}, 3000);
|
||||
} else {
|
||||
throw new Error(result.detail || result.error || 'Re-indexing failed');
|
||||
}
|
||||
} catch (error) {
|
||||
console.error('[Library Desk] Error:', error);
|
||||
|
||||
if (isFloating) {
|
||||
button.innerHTML = '✗ Failed';
|
||||
button.style.background = '#f44336';
|
||||
} else {
|
||||
icon.className = 'v-icon notranslate mdi mdi-alert-circle theme--dark';
|
||||
icon.style.color = '#f44336';
|
||||
}
|
||||
|
||||
showNotification('error', 'Re-indexing Failed', error.message + '\n\nCheck browser console for details.');
|
||||
|
||||
// Reset button
|
||||
setTimeout(function() {
|
||||
if (isFloating) {
|
||||
button.innerHTML = originalContent;
|
||||
button.style.background = '#1976d2';
|
||||
} else {
|
||||
icon.className = originalIcon;
|
||||
icon.style.color = '';
|
||||
}
|
||||
button.disabled = false;
|
||||
}, 4000);
|
||||
}
|
||||
}
|
||||
|
||||
// ============================================================================
|
||||
// ENTITY LINKING FUNCTIONALITY
|
||||
// ============================================================================
|
||||
|
||||
async function linkEntities(pagePath, button) {
|
||||
const isFloating = button.id === 'library-desk-entitylink-btn' && button.innerHTML.includes('Link');
|
||||
const icon = isFloating ? null : button.querySelector('.v-icon');
|
||||
const originalContent = isFloating ? button.innerHTML : null;
|
||||
const originalIcon = icon ? icon.className : null;
|
||||
|
||||
button.disabled = true;
|
||||
|
||||
if (isFloating) {
|
||||
button.innerHTML = '⏳ Finding...';
|
||||
button.style.background = '#757575';
|
||||
} else {
|
||||
icon.className = 'v-icon notranslate mdi mdi-loading mdi-spin theme--dark';
|
||||
}
|
||||
|
||||
try {
|
||||
// Get page ID from Wiki.js GraphQL API
|
||||
const pageId = await getPageIdFromPath(pagePath);
|
||||
if (!pageId) {
|
||||
throw new Error('Could not get page ID from Wiki.js');
|
||||
}
|
||||
|
||||
log('Linking entities for page', pageId, '...');
|
||||
|
||||
// Call entity linking endpoint
|
||||
const response = await fetch(CONFIG.libraryDeskUrl + '/entity-linking/link-page', {
|
||||
method: 'POST',
|
||||
headers: {
|
||||
'Authorization': 'Bearer ' + CONFIG.apiKey,
|
||||
'Content-Type': 'application/json'
|
||||
},
|
||||
body: JSON.stringify({
|
||||
user: CONFIG.user,
|
||||
page_id: pageId,
|
||||
create_relationships: true,
|
||||
re_index_if_changed: true
|
||||
})
|
||||
});
|
||||
|
||||
const result = await response.json();
|
||||
|
||||
if (response.ok) {
|
||||
if (isFloating) {
|
||||
button.innerHTML = '✓ Linked!';
|
||||
button.style.background = '#4caf50';
|
||||
} else {
|
||||
icon.className = 'v-icon notranslate mdi mdi-check-circle theme--dark';
|
||||
icon.style.color = '#4caf50';
|
||||
}
|
||||
|
||||
const entityList = result.entities_found
|
||||
.slice(0, 5)
|
||||
.map(e => `- ${e.name} (${e.mentions} mentions)${e.path ? ' ✓ linked' : ''}`)
|
||||
.join('\n');
|
||||
|
||||
const message = 'Entity Linking Complete!\n\n' +
|
||||
'Found: ' + result.entities_found.length + ' entities\n' +
|
||||
'Wiki links added: ' + result.content_links_added + '\n' +
|
||||
'Graph links added: ' + result.new_graph_links_created + '\n' +
|
||||
'Content updated: ' + (result.content_updated ? 'Yes' : 'No') + '\n' +
|
||||
'Re-indexed: ' + (result.re_indexed ? 'Yes' : 'No') + '\n\n' +
|
||||
'Top entities:\n' + (entityList || 'None');
|
||||
|
||||
log('Success:', result);
|
||||
alert(message);
|
||||
|
||||
// Reload page if content was updated
|
||||
if (result.content_updated) {
|
||||
log('Reloading page to show updated content...');
|
||||
setTimeout(function() {
|
||||
window.location.reload();
|
||||
}, 1500);
|
||||
} else {
|
||||
setTimeout(function() {
|
||||
if (isFloating) {
|
||||
button.innerHTML = originalContent;
|
||||
button.style.background = '#43a047';
|
||||
} else {
|
||||
icon.className = originalIcon;
|
||||
icon.style.color = '';
|
||||
}
|
||||
button.disabled = false;
|
||||
}, 3000);
|
||||
}
|
||||
} else {
|
||||
throw new Error(result.detail || 'Entity linking failed');
|
||||
}
|
||||
} catch (error) {
|
||||
console.error('[Library Desk] Error:', error);
|
||||
|
||||
if (isFloating) {
|
||||
button.innerHTML = '✗ Failed';
|
||||
button.style.background = '#f44336';
|
||||
} else {
|
||||
icon.className = 'v-icon notranslate mdi mdi-alert-circle theme--dark';
|
||||
icon.style.color = '#f44336';
|
||||
}
|
||||
|
||||
alert('Entity linking failed:\n\n' + error.message);
|
||||
|
||||
setTimeout(function() {
|
||||
if (isFloating) {
|
||||
button.innerHTML = originalContent;
|
||||
button.style.background = '#43a047';
|
||||
} else {
|
||||
icon.className = originalIcon;
|
||||
icon.style.color = '';
|
||||
}
|
||||
button.disabled = false;
|
||||
}, 4000);
|
||||
}
|
||||
}
|
||||
|
||||
// ============================================================================
|
||||
// SHARED UTILITIES
|
||||
// ============================================================================
|
||||
|
||||
async function getPageIdFromPath(pagePath) {
|
||||
/**
|
||||
* Query Wiki.js GraphQL API to get page ID from path.
|
||||
* Uses singleByPath query which requires path and locale.
|
||||
*/
|
||||
try {
|
||||
// Get locale from URL or default to 'en'
|
||||
const urlLocaleMatch = window.location.pathname.match(/^\/([a-z]{2})\//);
|
||||
const locale = urlLocaleMatch ? urlLocaleMatch[1] : 'en';
|
||||
|
||||
log('Querying Wiki.js GraphQL for page:', pagePath, 'locale:', locale);
|
||||
|
||||
const query = `
|
||||
query ($path: String!, $locale: String!) {
|
||||
pages {
|
||||
singleByPath(path: $path, locale: $locale) {
|
||||
id
|
||||
}
|
||||
}
|
||||
}
|
||||
`;
|
||||
|
||||
const response = await fetch('/graphql', {
|
||||
method: 'POST',
|
||||
headers: { 'Content-Type': 'application/json' },
|
||||
credentials: 'same-origin',
|
||||
body: JSON.stringify({
|
||||
query: query,
|
||||
variables: { path: pagePath, locale: locale }
|
||||
})
|
||||
});
|
||||
|
||||
const result = await response.json();
|
||||
log('GraphQL response:', result);
|
||||
|
||||
if (result.data && result.data.pages && result.data.pages.singleByPath) {
|
||||
const pageId = result.data.pages.singleByPath.id;
|
||||
log('Found page ID from GraphQL:', pageId);
|
||||
return parseInt(pageId, 10);
|
||||
}
|
||||
|
||||
if (result.errors) {
|
||||
log('GraphQL errors:', result.errors.map(e => e.message).join(', '));
|
||||
}
|
||||
|
||||
log('GraphQL query returned no page');
|
||||
return null;
|
||||
} catch (error) {
|
||||
log('GraphQL query failed:', error.message);
|
||||
return null;
|
||||
}
|
||||
}
|
||||
|
||||
function getPageId() {
|
||||
/**
|
||||
* Synchronous wrapper - returns null and triggers async fetch.
|
||||
* Actual buttons will need to call getPageIdFromPath() directly.
|
||||
*/
|
||||
log('getPageId() called - page ID must be fetched asynchronously via getPageIdFromPath()');
|
||||
return null;
|
||||
}
|
||||
|
||||
function waitForPageElement(callback, maxAttempts = 10, interval = 500) {
|
||||
/**
|
||||
* Wait for Vue store to be available, then call callback.
|
||||
*/
|
||||
let attempts = 0;
|
||||
|
||||
function check() {
|
||||
attempts++;
|
||||
|
||||
// Check if Vue store or page data is available
|
||||
const hasVueStore = window.$store && (window.$store.state || window.$store.get);
|
||||
const hasPageData = window.$page || window.pageId;
|
||||
|
||||
if (hasVueStore || hasPageData) {
|
||||
log('Vue/page data available after', attempts, 'attempts');
|
||||
callback();
|
||||
} else if (attempts < maxAttempts) {
|
||||
log('Waiting for Vue/page data... attempt', attempts);
|
||||
setTimeout(check, interval);
|
||||
} else {
|
||||
log('Vue/page data never appeared after', maxAttempts, 'attempts - adding buttons anyway');
|
||||
callback(); // Add buttons anyway, they just won't work
|
||||
}
|
||||
}
|
||||
|
||||
check();
|
||||
}
|
||||
|
||||
function findToolbar() {
|
||||
const selectors = [
|
||||
'header.v-toolbar .v-toolbar__content',
|
||||
'.v-app-bar .v-toolbar__content',
|
||||
'nav.v-toolbar .v-toolbar__content',
|
||||
'header .v-toolbar__content',
|
||||
'.v-app-bar__content'
|
||||
];
|
||||
|
||||
for (var i = 0; i < selectors.length; i++) {
|
||||
var elem = document.querySelector(selectors[i]);
|
||||
if (elem) return elem;
|
||||
}
|
||||
|
||||
return null;
|
||||
}
|
||||
|
||||
function getPagePath() {
|
||||
var path = window.location.pathname;
|
||||
path = path.replace(/^\//, '').replace(/\/$/, '');
|
||||
path = path.replace(/^[a-z]{2}\//, ''); // Remove language code
|
||||
|
||||
// Don't add buttons to home page or special pages
|
||||
if (path === '' || path === 'home' || path.startsWith('_') || path.startsWith('a/')) {
|
||||
return null;
|
||||
}
|
||||
|
||||
return path;
|
||||
}
|
||||
|
||||
function showNotification(type, title, message) {
|
||||
// Try Wiki.js notifications if available
|
||||
if (window.$store && typeof window.$store.commit === 'function') {
|
||||
try {
|
||||
window.$store.commit('showNotification', {
|
||||
message: title + ': ' + message,
|
||||
style: type,
|
||||
icon: type === 'success' ? 'check' : 'alert'
|
||||
});
|
||||
return;
|
||||
} catch (e) {
|
||||
// Fall through to alert
|
||||
}
|
||||
}
|
||||
|
||||
// Fallback to browser alert
|
||||
alert(title + '\n\n' + message);
|
||||
}
|
||||
|
||||
log('Library Desk integration loaded successfully');
|
||||
log('Version: 2.0.0 - Combined re-index and entity linking');
|
||||
|
||||
})();
|
||||
@@ -1 +0,0 @@
|
||||
"""Tests for Library Desk service."""
|
||||
@@ -1,109 +0,0 @@
|
||||
"""Pytest configuration and shared fixtures for Library Desk tests."""
|
||||
|
||||
import pytest
|
||||
import pytest_asyncio
|
||||
from typing import AsyncGenerator
|
||||
|
||||
# Test configuration
|
||||
pytest_plugins = ("pytest_asyncio",)
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def test_user() -> str:
|
||||
"""Default test user."""
|
||||
return "test_user"
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def neo4j_test_uri() -> str:
|
||||
"""Test Neo4j URI."""
|
||||
return "bolt://neo4j:7687"
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def neo4j_test_auth() -> tuple:
|
||||
"""Test Neo4j authentication."""
|
||||
return ("neo4j", "test_password")
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def qdrant_test_url() -> str:
|
||||
"""Test Qdrant URL."""
|
||||
return "http://qdrant:6333"
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def wikijs_test_config() -> dict:
|
||||
"""Test Wiki.js configuration."""
|
||||
return {
|
||||
"base_url": "http://wiki:3000",
|
||||
"api_key": "test_api_key"
|
||||
}
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def searxng_test_url() -> str:
|
||||
"""Test SearXNG URL."""
|
||||
return "http://searxng:8080"
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def ollama_test_config() -> dict:
|
||||
"""Test Ollama configuration."""
|
||||
return {
|
||||
"base_url": "http://ollama:11434",
|
||||
"model": "nomic-embed-text"
|
||||
}
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def redis_test_url() -> str:
|
||||
"""Test Redis URL."""
|
||||
return "redis://redis-shared:6379/4"
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def sample_document() -> dict:
|
||||
"""Sample document for testing."""
|
||||
return {
|
||||
"id": "test_doc_1",
|
||||
"title": "Test Document",
|
||||
"content": "This is a test document for unit testing.",
|
||||
"metadata": {
|
||||
"source": "test",
|
||||
"author": "test_user"
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def sample_chunks() -> list:
|
||||
"""Sample document chunks for testing."""
|
||||
return [
|
||||
{
|
||||
"content": "First chunk of text.",
|
||||
"metadata": {"chunk_index": 0}
|
||||
},
|
||||
{
|
||||
"content": "Second chunk of text.",
|
||||
"metadata": {"chunk_index": 1}
|
||||
},
|
||||
{
|
||||
"content": "Third chunk of text.",
|
||||
"metadata": {"chunk_index": 2}
|
||||
}
|
||||
]
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def sample_embeddings() -> list:
|
||||
"""Sample embeddings for testing (768-dimensional for nomic-embed-text)."""
|
||||
import random
|
||||
random.seed(42) # Reproducible embeddings
|
||||
|
||||
# Generate 3 sample 768-dimensional embeddings
|
||||
return [
|
||||
[random.random() for _ in range(768)],
|
||||
[random.random() for _ in range(768)],
|
||||
[random.random() for _ in range(768)]
|
||||
]
|
||||
@@ -1,754 +0,0 @@
|
||||
"""
|
||||
Comprehensive tests for Knowledge Consolidation system.
|
||||
|
||||
Tests cover:
|
||||
- ConsolidationService (unit tests with mocks)
|
||||
- Consolidation API endpoint (integration tests)
|
||||
- Model validation
|
||||
- Error handling
|
||||
- Dry run mode
|
||||
|
||||
Run with: pytest tests/test_consolidation.py -v -s
|
||||
"""
|
||||
|
||||
import pytest
|
||||
import pytest_asyncio
|
||||
from unittest.mock import AsyncMock, MagicMock, patch
|
||||
from typing import AsyncGenerator
|
||||
from datetime import datetime
|
||||
import json
|
||||
|
||||
from src.services.consolidation_service import ConsolidationService
|
||||
from src.models.consolidation import (
|
||||
ConsolidationRequest,
|
||||
ConsolidationResponse,
|
||||
ConsolidationResult,
|
||||
SearchQueryInfo
|
||||
)
|
||||
|
||||
# Test constants
|
||||
TEST_USER = "consolidation-tester"
|
||||
TEST_SEARCH_ID = "test-search-123"
|
||||
|
||||
|
||||
# Fixtures
|
||||
|
||||
@pytest.fixture
|
||||
def settings():
|
||||
"""Get mocked application settings for testing."""
|
||||
mock_settings = MagicMock()
|
||||
mock_settings.reranker_model = "mistral-nemo"
|
||||
mock_settings.ollama_model = "mistral-nemo"
|
||||
return mock_settings
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def mock_neo4j():
|
||||
"""Mock Neo4j client."""
|
||||
mock = AsyncMock()
|
||||
mock.execute_query = AsyncMock()
|
||||
return mock
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def mock_ollama():
|
||||
"""Mock Ollama client."""
|
||||
mock = AsyncMock()
|
||||
mock.generate_text = AsyncMock()
|
||||
return mock
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def mock_wiki():
|
||||
"""Mock Wiki.js client."""
|
||||
mock = AsyncMock()
|
||||
# Default mock for taxonomy structure
|
||||
mock.get_taxonomy_structure = AsyncMock(return_value={
|
||||
"companies": [],
|
||||
"people": [],
|
||||
"places": ["the-netherlands"],
|
||||
"reference": ["political-entities", "tech"],
|
||||
"technology": ["tools", "services"]
|
||||
})
|
||||
return mock
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def mock_ingestion():
|
||||
"""Mock Ingestion service."""
|
||||
mock = AsyncMock()
|
||||
mock.ingest_page = AsyncMock(return_value=MagicMock(success=True))
|
||||
return mock
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def consolidation_service(mock_neo4j, mock_ollama, mock_wiki, mock_ingestion, settings):
|
||||
"""Get ConsolidationService with mocked dependencies."""
|
||||
return ConsolidationService(
|
||||
neo4j=mock_neo4j,
|
||||
ollama=mock_ollama,
|
||||
wiki=mock_wiki,
|
||||
settings=settings,
|
||||
ingestion_service=mock_ingestion
|
||||
)
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def consolidation_service_no_ingestion(mock_neo4j, mock_ollama, mock_wiki, settings):
|
||||
"""Get ConsolidationService without ingestion service (legacy behavior)."""
|
||||
return ConsolidationService(
|
||||
neo4j=mock_neo4j,
|
||||
ollama=mock_ollama,
|
||||
wiki=mock_wiki,
|
||||
settings=settings,
|
||||
ingestion_service=None
|
||||
)
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def sample_unprocessed_searches():
|
||||
"""Sample unprocessed search queries."""
|
||||
return [
|
||||
{
|
||||
'id': 'search-1',
|
||||
'query': 'docker orchestration kubernetes',
|
||||
'user': TEST_USER,
|
||||
'timestamp': datetime.now().isoformat(),
|
||||
'total_results': 10,
|
||||
'web_count': 5,
|
||||
'keywords': ['docker', 'orchestration', 'kubernetes']
|
||||
},
|
||||
{
|
||||
'id': 'search-2',
|
||||
'query': 'python async programming',
|
||||
'user': TEST_USER,
|
||||
'timestamp': datetime.now().isoformat(),
|
||||
'total_results': 8,
|
||||
'web_count': 3,
|
||||
'keywords': ['python', 'async', 'programming']
|
||||
}
|
||||
]
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def sample_web_results():
|
||||
"""Sample web search results."""
|
||||
return [
|
||||
{
|
||||
'url': 'https://kubernetes.io/docs',
|
||||
'title': 'Kubernetes Documentation',
|
||||
'content': 'Kubernetes is an orchestration platform for containers...',
|
||||
'rank': 1,
|
||||
'rrf_score': 0.05
|
||||
},
|
||||
{
|
||||
'url': 'https://docs.docker.com/swarm',
|
||||
'title': 'Docker Swarm Documentation',
|
||||
'content': 'Docker Swarm is a container orchestration tool...',
|
||||
'rank': 2,
|
||||
'rrf_score': 0.04
|
||||
},
|
||||
{
|
||||
'url': 'https://example.com/k8s-tutorial',
|
||||
'title': 'Kubernetes Tutorial',
|
||||
'content': 'Learn how to use Kubernetes for container orchestration...',
|
||||
'rank': 3,
|
||||
'rrf_score': 0.03
|
||||
}
|
||||
]
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def sample_llm_analysis():
|
||||
"""Sample LLM analysis response."""
|
||||
return {
|
||||
"has_novel_info": True,
|
||||
"new_pages": [
|
||||
{
|
||||
"title": "Kubernetes Container Orchestration",
|
||||
"path": "infrastructure/kubernetes",
|
||||
"summary": "Overview of Kubernetes orchestration capabilities"
|
||||
}
|
||||
],
|
||||
"update_pages": [
|
||||
{
|
||||
"title": "Docker Infrastructure",
|
||||
"new_facts": [
|
||||
"Kubernetes provides automatic bin packing",
|
||||
"Self-healing capabilities with automatic restarts"
|
||||
],
|
||||
"source_url": "https://kubernetes.io/docs"
|
||||
}
|
||||
],
|
||||
"new_entities": [
|
||||
{
|
||||
"name": "Kubernetes",
|
||||
"type": "technology",
|
||||
"description": "Container orchestration platform"
|
||||
},
|
||||
{
|
||||
"name": "Docker Swarm",
|
||||
"type": "technology",
|
||||
"description": "Docker's native orchestration tool"
|
||||
}
|
||||
]
|
||||
}
|
||||
|
||||
|
||||
# Model Tests
|
||||
|
||||
def test_consolidation_request_validation():
|
||||
"""Test ConsolidationRequest model validation."""
|
||||
# Valid request
|
||||
request = ConsolidationRequest(
|
||||
process_limit=10,
|
||||
lookback_days=7,
|
||||
min_web_results=2,
|
||||
dry_run=False
|
||||
)
|
||||
assert request.process_limit == 10
|
||||
assert request.lookback_days == 7
|
||||
assert request.min_web_results == 2
|
||||
assert request.dry_run is False
|
||||
|
||||
# Default values
|
||||
request = ConsolidationRequest()
|
||||
assert request.process_limit == 10
|
||||
assert request.lookback_days == 7
|
||||
assert request.min_web_results == 2
|
||||
assert request.dry_run is False
|
||||
|
||||
# Validate limits
|
||||
with pytest.raises(Exception):
|
||||
ConsolidationRequest(process_limit=0) # Too low
|
||||
|
||||
with pytest.raises(Exception):
|
||||
ConsolidationRequest(process_limit=101) # Too high
|
||||
|
||||
|
||||
def test_consolidation_response_model():
|
||||
"""Test ConsolidationResponse model."""
|
||||
response = ConsolidationResponse(
|
||||
total_found=5,
|
||||
processed_count=4,
|
||||
pages_created=2,
|
||||
pages_updated=3,
|
||||
entities_added=5,
|
||||
errors=["Error 1"],
|
||||
results=[],
|
||||
dry_run=False
|
||||
)
|
||||
|
||||
assert response.total_found == 5
|
||||
assert response.processed_count == 4
|
||||
assert response.pages_created == 2
|
||||
assert len(response.errors) == 1
|
||||
|
||||
|
||||
def test_consolidation_result_model():
|
||||
"""Test ConsolidationResult model."""
|
||||
result = ConsolidationResult(
|
||||
search_id="test-123",
|
||||
query="test query",
|
||||
pages_created=1,
|
||||
pages_updated=2,
|
||||
entities_added=3,
|
||||
error=None
|
||||
)
|
||||
|
||||
assert result.search_id == "test-123"
|
||||
assert result.query == "test query"
|
||||
assert result.pages_created == 1
|
||||
assert result.error is None
|
||||
|
||||
|
||||
# Service Unit Tests
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_find_unprocessed_searches_empty(consolidation_service, mock_neo4j):
|
||||
"""Test finding unprocessed searches when none exist."""
|
||||
# Mock empty result
|
||||
mock_neo4j.execute_query.return_value = []
|
||||
|
||||
searches = await consolidation_service._find_unprocessed_searches(
|
||||
lookback_days=7,
|
||||
limit=10
|
||||
)
|
||||
|
||||
assert len(searches) == 0
|
||||
mock_neo4j.execute_query.assert_called_once()
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_find_unprocessed_searches_with_results(
|
||||
consolidation_service,
|
||||
mock_neo4j,
|
||||
sample_unprocessed_searches
|
||||
):
|
||||
"""Test finding unprocessed searches with results."""
|
||||
# Mock Neo4j response
|
||||
mock_neo4j.execute_query.return_value = sample_unprocessed_searches
|
||||
|
||||
searches = await consolidation_service._find_unprocessed_searches(
|
||||
lookback_days=7,
|
||||
limit=10
|
||||
)
|
||||
|
||||
assert len(searches) == 2
|
||||
assert searches[0]['query'] == 'docker orchestration kubernetes'
|
||||
assert searches[1]['query'] == 'python async programming'
|
||||
mock_neo4j.execute_query.assert_called_once()
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_get_web_results(consolidation_service, mock_neo4j, sample_web_results):
|
||||
"""Test retrieving web results for a search."""
|
||||
# Mock Neo4j response
|
||||
mock_neo4j.execute_query.return_value = sample_web_results
|
||||
|
||||
results = await consolidation_service._get_web_results(TEST_SEARCH_ID)
|
||||
|
||||
assert len(results) == 3
|
||||
assert results[0]['title'] == 'Kubernetes Documentation'
|
||||
assert results[1]['url'] == 'https://docs.docker.com/swarm'
|
||||
mock_neo4j.execute_query.assert_called_once()
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_get_web_results_empty(consolidation_service, mock_neo4j):
|
||||
"""Test retrieving web results when none exist."""
|
||||
mock_neo4j.execute_query.return_value = []
|
||||
|
||||
results = await consolidation_service._get_web_results(TEST_SEARCH_ID)
|
||||
|
||||
assert len(results) == 0
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_analyze_web_results_with_novel_info(
|
||||
consolidation_service,
|
||||
mock_ollama,
|
||||
mock_wiki,
|
||||
sample_web_results,
|
||||
sample_llm_analysis
|
||||
):
|
||||
"""Test analyzing web results with Ollama - novel info found."""
|
||||
# Mock Ollama response
|
||||
mock_ollama.generate_text.return_value = json.dumps(sample_llm_analysis)
|
||||
|
||||
analysis = await consolidation_service._analyze_web_results(
|
||||
query="docker orchestration",
|
||||
web_results=sample_web_results,
|
||||
keywords=["docker", "orchestration"],
|
||||
user=TEST_USER
|
||||
)
|
||||
|
||||
assert analysis is not None
|
||||
assert analysis['has_novel_info'] is True
|
||||
assert len(analysis['new_pages']) == 1
|
||||
assert len(analysis['update_pages']) == 1
|
||||
assert len(analysis['new_entities']) == 2
|
||||
mock_ollama.generate_text.assert_called_once()
|
||||
# Verify taxonomy was fetched
|
||||
mock_wiki.get_taxonomy_structure.assert_called_once_with(f"users/{TEST_USER}")
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_analyze_web_results_no_novel_info(
|
||||
consolidation_service,
|
||||
mock_ollama,
|
||||
sample_web_results
|
||||
):
|
||||
"""Test analyzing web results - no novel info."""
|
||||
# Mock Ollama response with no novel info
|
||||
analysis_no_novel = {
|
||||
"has_novel_info": False,
|
||||
"new_pages": [],
|
||||
"update_pages": [],
|
||||
"new_entities": []
|
||||
}
|
||||
mock_ollama.generate_text.return_value = json.dumps(analysis_no_novel)
|
||||
|
||||
analysis = await consolidation_service._analyze_web_results(
|
||||
query="common topic",
|
||||
web_results=sample_web_results,
|
||||
keywords=[],
|
||||
user=TEST_USER
|
||||
)
|
||||
|
||||
assert analysis is not None
|
||||
assert analysis['has_novel_info'] is False
|
||||
assert len(analysis['new_pages']) == 0
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_analyze_web_results_invalid_json(
|
||||
consolidation_service,
|
||||
mock_ollama,
|
||||
sample_web_results
|
||||
):
|
||||
"""Test analyzing web results with invalid JSON response."""
|
||||
# Mock Ollama response with invalid JSON
|
||||
mock_ollama.generate_text.return_value = "This is not JSON"
|
||||
|
||||
analysis = await consolidation_service._analyze_web_results(
|
||||
query="test query",
|
||||
web_results=sample_web_results,
|
||||
keywords=[],
|
||||
user=TEST_USER
|
||||
)
|
||||
|
||||
assert analysis is None
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_analyze_web_results_json_in_markdown(
|
||||
consolidation_service,
|
||||
mock_ollama,
|
||||
sample_web_results,
|
||||
sample_llm_analysis
|
||||
):
|
||||
"""Test extracting JSON from markdown-wrapped response."""
|
||||
# Mock Ollama response with JSON wrapped in markdown
|
||||
wrapped_response = f"""Here's the analysis:
|
||||
|
||||
```json
|
||||
{json.dumps(sample_llm_analysis)}
|
||||
```
|
||||
|
||||
Hope this helps!"""
|
||||
mock_ollama.generate_text.return_value = wrapped_response
|
||||
|
||||
analysis = await consolidation_service._analyze_web_results(
|
||||
query="test",
|
||||
web_results=sample_web_results,
|
||||
keywords=[],
|
||||
user=TEST_USER
|
||||
)
|
||||
|
||||
assert analysis is not None
|
||||
assert analysis['has_novel_info'] is True
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_analyze_web_results_taxonomy_failure(
|
||||
consolidation_service,
|
||||
mock_ollama,
|
||||
mock_wiki,
|
||||
sample_web_results,
|
||||
sample_llm_analysis
|
||||
):
|
||||
"""Test that analysis continues even if taxonomy fetch fails."""
|
||||
# Mock taxonomy fetch failure
|
||||
mock_wiki.get_taxonomy_structure.side_effect = Exception("Connection error")
|
||||
mock_ollama.generate_text.return_value = json.dumps(sample_llm_analysis)
|
||||
|
||||
analysis = await consolidation_service._analyze_web_results(
|
||||
query="test",
|
||||
web_results=sample_web_results,
|
||||
keywords=[],
|
||||
user=TEST_USER
|
||||
)
|
||||
|
||||
# Should still succeed, just without taxonomy info in prompt
|
||||
assert analysis is not None
|
||||
assert analysis['has_novel_info'] is True
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_format_taxonomy_for_prompt(consolidation_service):
|
||||
"""Test taxonomy formatting for LLM prompt."""
|
||||
taxonomy = {
|
||||
"companies": [],
|
||||
"places": ["the-netherlands", "rotterdam"],
|
||||
"reference": ["political-entities"]
|
||||
}
|
||||
|
||||
formatted = consolidation_service._format_taxonomy_for_prompt(taxonomy)
|
||||
|
||||
assert "Existing paths" in formatted
|
||||
assert "companies/" in formatted
|
||||
assert "places/the-netherlands/" in formatted
|
||||
assert "places/rotterdam/" in formatted
|
||||
assert "reference/political-entities/" in formatted
|
||||
assert "IMPORTANT" in formatted
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_format_taxonomy_empty(consolidation_service):
|
||||
"""Test taxonomy formatting with empty taxonomy."""
|
||||
formatted = consolidation_service._format_taxonomy_for_prompt({})
|
||||
assert formatted == ""
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_mark_search_processed(consolidation_service, mock_neo4j):
|
||||
"""Test marking search as processed."""
|
||||
await consolidation_service._mark_search_processed(TEST_SEARCH_ID)
|
||||
|
||||
mock_neo4j.execute_query.assert_called_once()
|
||||
call_args = mock_neo4j.execute_query.call_args
|
||||
assert TEST_SEARCH_ID in str(call_args)
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_process_search_insufficient_web_results(
|
||||
consolidation_service,
|
||||
sample_unprocessed_searches
|
||||
):
|
||||
"""Test processing search with insufficient web results."""
|
||||
search = sample_unprocessed_searches[1].copy()
|
||||
search['web_count'] = 1 # Below minimum
|
||||
|
||||
result = await consolidation_service._process_search(
|
||||
search=search,
|
||||
min_web_results=2,
|
||||
dry_run=False
|
||||
)
|
||||
|
||||
assert result is None
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_process_search_no_web_results_in_db(
|
||||
consolidation_service,
|
||||
mock_neo4j,
|
||||
sample_unprocessed_searches
|
||||
):
|
||||
"""Test processing search when web results not found in DB."""
|
||||
mock_neo4j.execute_query.return_value = []
|
||||
|
||||
result = await consolidation_service._process_search(
|
||||
search=sample_unprocessed_searches[0],
|
||||
min_web_results=2,
|
||||
dry_run=False
|
||||
)
|
||||
|
||||
assert result is None
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_process_search_dry_run(
|
||||
consolidation_service,
|
||||
mock_neo4j,
|
||||
mock_ollama,
|
||||
sample_unprocessed_searches,
|
||||
sample_web_results,
|
||||
sample_llm_analysis
|
||||
):
|
||||
"""Test processing search in dry run mode."""
|
||||
# Mock responses
|
||||
mock_neo4j.execute_query.return_value = sample_web_results
|
||||
mock_ollama.generate_text.return_value = json.dumps(sample_llm_analysis)
|
||||
|
||||
result = await consolidation_service._process_search(
|
||||
search=sample_unprocessed_searches[0],
|
||||
min_web_results=2,
|
||||
dry_run=True
|
||||
)
|
||||
|
||||
assert result is not None
|
||||
assert result.search_id == 'search-1'
|
||||
assert result.pages_created == 1
|
||||
assert result.pages_updated == 1
|
||||
assert result.entities_added == 2
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_consolidate_knowledge_no_searches(
|
||||
consolidation_service,
|
||||
mock_neo4j
|
||||
):
|
||||
"""Test consolidation when no unprocessed searches found."""
|
||||
mock_neo4j.execute_query.return_value = []
|
||||
|
||||
response = await consolidation_service.consolidate_knowledge(
|
||||
process_limit=10,
|
||||
lookback_days=7,
|
||||
min_web_results=2,
|
||||
dry_run=False
|
||||
)
|
||||
|
||||
assert response.total_found == 0
|
||||
assert response.processed_count == 0
|
||||
assert response.pages_created == 0
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_consolidate_knowledge_success(
|
||||
consolidation_service,
|
||||
mock_neo4j,
|
||||
mock_ollama,
|
||||
mock_wiki,
|
||||
sample_unprocessed_searches,
|
||||
sample_web_results,
|
||||
sample_llm_analysis
|
||||
):
|
||||
"""Test successful knowledge consolidation."""
|
||||
# Mock finding searches and entity creation
|
||||
# Each search processes: get web results, add 2 entities, mark processed
|
||||
mock_neo4j.execute_query.side_effect = [
|
||||
sample_unprocessed_searches, # Find searches
|
||||
sample_web_results, # Get web results for search 1
|
||||
None, # Add entity 1 (Kubernetes)
|
||||
None, # Add entity 2 (Docker Swarm)
|
||||
None, # Mark search 1 processed
|
||||
sample_web_results, # Get web results for search 2
|
||||
None, # Add entity 1 (Kubernetes)
|
||||
None, # Add entity 2 (Docker Swarm)
|
||||
None, # Mark search 2 processed
|
||||
]
|
||||
|
||||
# Mock wiki operations
|
||||
mock_wiki.search_pages.return_value = [] # No existing pages
|
||||
mock_wiki.create_page.return_value = None
|
||||
mock_wiki.update_page.return_value = None
|
||||
mock_wiki.get_page.return_value = None
|
||||
|
||||
# Mock LLM analysis and WikiPageWriter LLM calls
|
||||
mock_ollama.generate_text.return_value = json.dumps(sample_llm_analysis)
|
||||
|
||||
response = await consolidation_service.consolidate_knowledge(
|
||||
process_limit=10,
|
||||
lookback_days=7,
|
||||
min_web_results=2,
|
||||
dry_run=False
|
||||
)
|
||||
|
||||
assert response.total_found == 2
|
||||
assert response.processed_count == 2
|
||||
assert response.dry_run is False
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_consolidate_knowledge_with_errors(
|
||||
consolidation_service,
|
||||
mock_neo4j,
|
||||
mock_ollama,
|
||||
sample_unprocessed_searches
|
||||
):
|
||||
"""Test consolidation with some searches failing."""
|
||||
# Mock finding searches - return empty for web results to trigger internal error handling
|
||||
mock_neo4j.execute_query.side_effect = [
|
||||
sample_unprocessed_searches, # Find searches
|
||||
[], # Empty web results for search 1 (causes skip, not error)
|
||||
[], # Empty web results for search 2 (causes skip, not error)
|
||||
]
|
||||
|
||||
response = await consolidation_service.consolidate_knowledge(
|
||||
process_limit=10,
|
||||
lookback_days=7,
|
||||
min_web_results=2,
|
||||
dry_run=False
|
||||
)
|
||||
|
||||
assert response.total_found == 2
|
||||
# Both searches skipped due to no web results (not errors)
|
||||
assert response.processed_count == 0
|
||||
|
||||
|
||||
# Integration Tests (API Endpoint)
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_consolidation_endpoint_minimal_request(consolidation_service):
|
||||
"""Test consolidation endpoint with minimal request."""
|
||||
pytest.importorskip("fastapi") # Skip if fastapi not available
|
||||
|
||||
# This would require proper test client setup
|
||||
# Placeholder for integration test structure
|
||||
request = ConsolidationRequest()
|
||||
assert request.process_limit == 10
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_consolidation_endpoint_custom_config(consolidation_service):
|
||||
"""Test consolidation endpoint with custom configuration."""
|
||||
request = ConsolidationRequest(
|
||||
process_limit=5,
|
||||
lookback_days=14,
|
||||
min_web_results=3,
|
||||
dry_run=True
|
||||
)
|
||||
|
||||
assert request.process_limit == 5
|
||||
assert request.lookback_days == 14
|
||||
assert request.min_web_results == 3
|
||||
assert request.dry_run is True
|
||||
|
||||
|
||||
# Edge Cases
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_consolidate_with_max_limits(consolidation_service, mock_neo4j):
|
||||
"""Test consolidation with maximum limits."""
|
||||
mock_neo4j.execute_query.return_value = []
|
||||
|
||||
response = await consolidation_service.consolidate_knowledge(
|
||||
process_limit=100, # Max
|
||||
lookback_days=90, # Max
|
||||
min_web_results=20, # Max
|
||||
dry_run=True
|
||||
)
|
||||
|
||||
assert response.total_found == 0
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_analyze_empty_web_results(consolidation_service, mock_ollama):
|
||||
"""Test analyzing with empty web results list."""
|
||||
mock_ollama.generate_text.return_value = json.dumps({
|
||||
"has_novel_info": False,
|
||||
"new_pages": [],
|
||||
"update_pages": [],
|
||||
"new_entities": []
|
||||
})
|
||||
|
||||
analysis = await consolidation_service._analyze_web_results(
|
||||
query="test",
|
||||
web_results=[],
|
||||
keywords=[],
|
||||
user=TEST_USER
|
||||
)
|
||||
|
||||
# Should still call LLM but return no novel info
|
||||
assert analysis is not None
|
||||
|
||||
|
||||
# Performance/Load Tests (optional)
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_process_many_searches_dry_run(
|
||||
consolidation_service,
|
||||
mock_neo4j,
|
||||
mock_ollama
|
||||
):
|
||||
"""Test processing many searches in dry run mode."""
|
||||
# Generate many test searches
|
||||
many_searches = [
|
||||
{
|
||||
'id': f'search-{i}',
|
||||
'query': f'test query {i}',
|
||||
'user': TEST_USER,
|
||||
'timestamp': datetime.now().isoformat(),
|
||||
'total_results': 5,
|
||||
'web_count': 3,
|
||||
'keywords': ['test']
|
||||
}
|
||||
for i in range(50)
|
||||
]
|
||||
|
||||
mock_neo4j.execute_query.return_value = many_searches[:10] # Limit by config
|
||||
|
||||
response = await consolidation_service.consolidate_knowledge(
|
||||
process_limit=10,
|
||||
lookback_days=7,
|
||||
min_web_results=2,
|
||||
dry_run=True
|
||||
)
|
||||
|
||||
# Should only process up to limit
|
||||
assert response.total_found == 10
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
pytest.main([__file__, "-v", "-s"])
|
||||
@@ -1,484 +0,0 @@
|
||||
"""
|
||||
Comprehensive tests for Entity Linking system.
|
||||
|
||||
Tests cover:
|
||||
- Finding entity mentions in pages
|
||||
- Creating MENTIONS relationships in Neo4j
|
||||
- Adding hyperlinks to wiki content
|
||||
- Idempotency (safe to run multiple times)
|
||||
- Protection of existing links (no nesting)
|
||||
- Multi-tenancy isolation
|
||||
|
||||
Run with: pytest tests/test_entity_linking.py -v -s
|
||||
"""
|
||||
|
||||
import pytest
|
||||
import pytest_asyncio
|
||||
from typing import AsyncGenerator
|
||||
|
||||
from src.clients.neo4j_client import Neo4jClient
|
||||
from src.clients.wikijs_client import WikiJSClient
|
||||
from src.services.graph_service import GraphService
|
||||
from src.services.wiki_service import WikiService
|
||||
from src.routers.entity_linking import (
|
||||
find_entity_mentions,
|
||||
add_entity_links_to_content,
|
||||
get_entities_with_paths
|
||||
)
|
||||
from src.config import get_settings
|
||||
|
||||
# Test user to isolate test data
|
||||
TEST_USER = "entity-link-tester"
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def settings():
|
||||
"""Get application settings."""
|
||||
return get_settings()
|
||||
|
||||
|
||||
@pytest_asyncio.fixture
|
||||
async def neo4j_client(settings) -> AsyncGenerator[Neo4jClient, None]:
|
||||
"""Get connected Neo4j client."""
|
||||
client = Neo4jClient(
|
||||
uri=settings.neo4j_uri,
|
||||
user=settings.neo4j_user,
|
||||
password=settings.neo4j_password
|
||||
)
|
||||
await client.connect()
|
||||
yield client
|
||||
await client.close()
|
||||
|
||||
|
||||
@pytest_asyncio.fixture
|
||||
async def wiki_client(settings) -> AsyncGenerator[WikiJSClient, None]:
|
||||
"""Get Wiki.js client."""
|
||||
client = WikiJSClient(
|
||||
base_url=settings.wikijs_url,
|
||||
username=settings.wikijs_username,
|
||||
password=settings.wikijs_password
|
||||
)
|
||||
yield client
|
||||
|
||||
|
||||
@pytest_asyncio.fixture
|
||||
async def graph_service(neo4j_client, wiki_client):
|
||||
"""Get GraphService instance."""
|
||||
return GraphService(neo4j_client, wiki_client)
|
||||
|
||||
|
||||
@pytest_asyncio.fixture
|
||||
async def wiki_service(wiki_client):
|
||||
"""Get WikiService instance."""
|
||||
return WikiService(wiki_client)
|
||||
|
||||
|
||||
@pytest_asyncio.fixture
|
||||
async def test_entities(graph_service):
|
||||
"""Create test entities in graph."""
|
||||
from src.core.multi_tenancy import get_neo4j_user_base_label
|
||||
|
||||
user_base_label = get_neo4j_user_base_label(TEST_USER)
|
||||
|
||||
# Clean up any existing test entities
|
||||
cleanup_query = f"""
|
||||
MATCH (n:{user_base_label})
|
||||
WHERE n.name IN ['Docker', 'Kubernetes', 'PostgreSQL']
|
||||
DETACH DELETE n
|
||||
"""
|
||||
await graph_service.neo4j.execute_query(cleanup_query)
|
||||
|
||||
# Create test entities
|
||||
create_query = f"""
|
||||
CREATE (d:{user_base_label}:Technology {{name: 'Docker', type: 'technology'}})
|
||||
CREATE (k:{user_base_label}:Technology {{name: 'Kubernetes', type: 'technology'}})
|
||||
CREATE (p:{user_base_label}:Technology {{name: 'PostgreSQL', type: 'technology'}})
|
||||
RETURN d.name, k.name, p.name
|
||||
"""
|
||||
await graph_service.neo4j.execute_query(create_query)
|
||||
|
||||
yield ["Docker", "Kubernetes", "PostgreSQL"]
|
||||
|
||||
# Cleanup after test
|
||||
await graph_service.neo4j.execute_query(cleanup_query)
|
||||
|
||||
|
||||
# ============================================================================
|
||||
# Unit Tests - Entity Mention Detection
|
||||
# ============================================================================
|
||||
|
||||
class TestFindEntityMentions:
|
||||
"""Test finding entity mentions in content."""
|
||||
|
||||
def test_find_single_mention(self):
|
||||
"""Test finding a single entity mention."""
|
||||
content = "Docker is a containerization platform."
|
||||
entities = [
|
||||
{"name": "Docker", "type": "technology"}
|
||||
]
|
||||
|
||||
found = find_entity_mentions(content, entities)
|
||||
|
||||
assert len(found) == 1
|
||||
assert found[0]["name"] == "Docker"
|
||||
assert found[0]["mentions"] == 1
|
||||
|
||||
def test_find_multiple_mentions(self):
|
||||
"""Test finding multiple mentions of same entity."""
|
||||
content = "Docker containers run on Docker Engine. Docker is great!"
|
||||
entities = [
|
||||
{"name": "Docker", "type": "technology"}
|
||||
]
|
||||
|
||||
found = find_entity_mentions(content, entities)
|
||||
|
||||
assert len(found) == 1
|
||||
assert found[0]["name"] == "Docker"
|
||||
assert found[0]["mentions"] == 3
|
||||
|
||||
def test_case_insensitive_matching(self):
|
||||
"""Test case-insensitive entity matching."""
|
||||
content = "docker and DOCKER and Docker are the same"
|
||||
entities = [
|
||||
{"name": "Docker", "type": "technology"}
|
||||
]
|
||||
|
||||
found = find_entity_mentions(content, entities)
|
||||
|
||||
assert len(found) == 1
|
||||
assert found[0]["mentions"] == 3
|
||||
|
||||
def test_whole_word_matching(self):
|
||||
"""Test that partial word matches are excluded."""
|
||||
content = "Kubernetes and Kubernetes-based and MyKubernetes"
|
||||
entities = [
|
||||
{"name": "Kubernetes", "type": "technology"}
|
||||
]
|
||||
|
||||
found = find_entity_mentions(content, entities)
|
||||
|
||||
assert len(found) == 1
|
||||
# Regex \b matches at hyphens, so "Kubernetes-based" contains "Kubernetes"
|
||||
# Only "MyKubernetes" is excluded (no word boundary)
|
||||
assert found[0]["mentions"] == 2 # "Kubernetes" and "Kubernetes-based"
|
||||
|
||||
def test_ignore_short_names(self):
|
||||
"""Test that entities with names <3 chars are ignored."""
|
||||
content = "Go is a programming language by Google"
|
||||
entities = [
|
||||
{"name": "Go", "type": "language"}, # Too short
|
||||
{"name": "Google", "type": "organization"}
|
||||
]
|
||||
|
||||
found = find_entity_mentions(content, entities)
|
||||
|
||||
assert len(found) == 1
|
||||
assert found[0]["name"] == "Google"
|
||||
|
||||
def test_sort_by_mention_count(self):
|
||||
"""Test results are sorted by mention count."""
|
||||
content = "Docker Docker Docker. Kubernetes Kubernetes. PostgreSQL."
|
||||
entities = [
|
||||
{"name": "PostgreSQL", "type": "database"},
|
||||
{"name": "Docker", "type": "technology"},
|
||||
{"name": "Kubernetes", "type": "technology"}
|
||||
]
|
||||
|
||||
found = find_entity_mentions(content, entities)
|
||||
|
||||
assert len(found) == 3
|
||||
assert found[0]["name"] == "Docker" # Most mentions
|
||||
assert found[0]["mentions"] == 3
|
||||
assert found[1]["name"] == "Kubernetes"
|
||||
assert found[1]["mentions"] == 2
|
||||
assert found[2]["name"] == "PostgreSQL"
|
||||
assert found[2]["mentions"] == 1
|
||||
|
||||
|
||||
# ============================================================================
|
||||
# Unit Tests - Content Link Addition
|
||||
# ============================================================================
|
||||
|
||||
class TestAddEntityLinksToContent:
|
||||
"""Test adding hyperlinks to content."""
|
||||
|
||||
def test_add_single_link(self):
|
||||
"""Test adding a single entity link."""
|
||||
content = "Docker is a containerization platform."
|
||||
entities = [
|
||||
{"name": "Docker", "path": "users/test/docker"}
|
||||
]
|
||||
|
||||
updated, count = add_entity_links_to_content(content, entities)
|
||||
|
||||
assert count == 1
|
||||
assert "[Docker](/docker)" in updated
|
||||
|
||||
def test_add_multiple_instances(self):
|
||||
"""Test linking all instances of an entity."""
|
||||
content = "Docker containers run on Docker Engine."
|
||||
entities = [
|
||||
{"name": "Docker", "path": "users/test/docker"}
|
||||
]
|
||||
|
||||
updated, count = add_entity_links_to_content(content, entities)
|
||||
|
||||
assert count == 2 # Both instances linked
|
||||
assert updated.count("[Docker](/docker)") == 2
|
||||
|
||||
def test_skip_entities_without_path(self):
|
||||
"""Test that entities without wiki pages are not linked."""
|
||||
content = "Docker and Kubernetes are used together."
|
||||
entities = [
|
||||
{"name": "Docker", "path": "users/test/docker"},
|
||||
{"name": "Kubernetes", "path": None} # No page
|
||||
]
|
||||
|
||||
updated, count = add_entity_links_to_content(content, entities)
|
||||
|
||||
assert count == 1 # Only Docker
|
||||
assert "[Docker](/docker)" in updated
|
||||
assert "[Kubernetes]" not in updated
|
||||
|
||||
def test_protect_existing_links(self):
|
||||
"""Test that existing markdown links are not modified."""
|
||||
content = "See [Docker](https://docker.com) for more info. Docker is great!"
|
||||
entities = [
|
||||
{"name": "Docker", "path": "users/test/docker"}
|
||||
]
|
||||
|
||||
updated, count = add_entity_links_to_content(content, entities)
|
||||
|
||||
# Should link the second "Docker" but not the one already linked
|
||||
assert count == 1
|
||||
assert "[Docker](https://docker.com)" in updated # Preserved
|
||||
assert updated.count("[Docker](/docker)") == 1
|
||||
|
||||
def test_no_nested_links(self):
|
||||
"""Test that entity names in URLs are not linked."""
|
||||
content = "Check [Docker Hub](/docker/hub) for images."
|
||||
entities = [
|
||||
{"name": "Docker", "path": "users/test/docker"}
|
||||
]
|
||||
|
||||
updated, count = add_entity_links_to_content(content, entities)
|
||||
|
||||
# "Docker" in the URL path should not be linked
|
||||
assert count == 0
|
||||
assert "[Docker Hub](/docker/hub)" in updated # Unchanged
|
||||
|
||||
def test_longest_first_matching(self):
|
||||
"""Test that longer entity names are matched first."""
|
||||
content = "Machine Learning and Machine are different."
|
||||
entities = [
|
||||
{"name": "Machine Learning", "path": "users/test/ml"},
|
||||
{"name": "Machine", "path": "users/test/machine"}
|
||||
]
|
||||
|
||||
updated, count = add_entity_links_to_content(content, entities)
|
||||
|
||||
# Should link "Machine Learning" first, leaving "Machine" alone
|
||||
assert "[Machine Learning](/ml)" in updated
|
||||
assert count >= 1
|
||||
|
||||
|
||||
# ============================================================================
|
||||
# Integration Tests - Full Entity Linking Flow
|
||||
# ============================================================================
|
||||
|
||||
class TestEntityLinkingIntegration:
|
||||
"""Test full entity linking flow."""
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_get_entities_with_paths(self, graph_service, test_entities):
|
||||
"""Test retrieving entities and their wiki page paths."""
|
||||
entities = await get_entities_with_paths(graph_service, TEST_USER)
|
||||
|
||||
# Should find our test entities
|
||||
entity_names = [e["name"] for e in entities]
|
||||
assert "Docker" in entity_names
|
||||
assert "Kubernetes" in entity_names
|
||||
assert "PostgreSQL" in entity_names
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_create_mentions_relationships(self, graph_service, test_entities):
|
||||
"""Test creating MENTIONS relationships."""
|
||||
from src.core.multi_tenancy import get_neo4j_user_label
|
||||
|
||||
user_doc_label = get_neo4j_user_label(TEST_USER)
|
||||
|
||||
# Create a test document node
|
||||
doc_query = f"""
|
||||
CREATE (d:{user_doc_label}:Document {{
|
||||
page_id: 9999,
|
||||
title: 'Test Doc',
|
||||
path: 'users/test/doc'
|
||||
}})
|
||||
RETURN d
|
||||
"""
|
||||
await graph_service.neo4j.execute_query(doc_query)
|
||||
|
||||
# Create MENTIONS relationships
|
||||
found_entities = [
|
||||
{"name": "Docker"},
|
||||
{"name": "Kubernetes"}
|
||||
]
|
||||
|
||||
new_links = await graph_service.create_entity_mentions(
|
||||
page_id=9999,
|
||||
user=TEST_USER,
|
||||
entity_names=found_entities
|
||||
)
|
||||
|
||||
assert new_links == 2
|
||||
|
||||
# Verify relationships exist
|
||||
verify_query = f"""
|
||||
MATCH (d:{user_doc_label}:Document {{page_id: 9999}})-[r:MENTIONS]->(e)
|
||||
RETURN count(r) as mention_count
|
||||
"""
|
||||
result = await graph_service.neo4j.execute_query(verify_query)
|
||||
assert result[0]["mention_count"] == 2
|
||||
|
||||
# Cleanup
|
||||
cleanup_query = f"""
|
||||
MATCH (d:{user_doc_label}:Document {{page_id: 9999}})
|
||||
DETACH DELETE d
|
||||
"""
|
||||
await graph_service.neo4j.execute_query(cleanup_query)
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_idempotency(self, graph_service):
|
||||
"""Test that entity linking is idempotent."""
|
||||
from src.core.multi_tenancy import get_neo4j_user_label, get_neo4j_user_base_label
|
||||
|
||||
user_doc_label = get_neo4j_user_label(TEST_USER)
|
||||
user_base_label = get_neo4j_user_base_label(TEST_USER)
|
||||
|
||||
# Aggressively clean up ALL test data first (fresh start)
|
||||
cleanup_all = f"""
|
||||
MATCH (n)
|
||||
WHERE (n:{user_base_label} OR n:{user_doc_label})
|
||||
AND (n.page_id = 9998 OR n.name = 'TestDockerEntity')
|
||||
DETACH DELETE n
|
||||
"""
|
||||
await graph_service.neo4j.execute_query(cleanup_all)
|
||||
|
||||
# Create a unique test entity
|
||||
entity_query = f"""
|
||||
CREATE (e:{user_base_label}:Technology {{name: 'TestDockerEntity', type: 'technology'}})
|
||||
RETURN e
|
||||
"""
|
||||
await graph_service.neo4j.execute_query(entity_query)
|
||||
|
||||
# Create test document
|
||||
doc_query = f"""
|
||||
CREATE (d:{user_doc_label}:Document {{
|
||||
page_id: 9998,
|
||||
title: 'Test Doc 2',
|
||||
path: 'users/test/doc2'
|
||||
}})
|
||||
RETURN d
|
||||
"""
|
||||
await graph_service.neo4j.execute_query(doc_query)
|
||||
|
||||
found_entities = [{"name": "TestDockerEntity"}]
|
||||
|
||||
# Link once
|
||||
first_run = await graph_service.create_entity_mentions(
|
||||
page_id=9998,
|
||||
user=TEST_USER,
|
||||
entity_names=found_entities
|
||||
)
|
||||
assert first_run == 1
|
||||
|
||||
# Link again - should not create duplicates
|
||||
second_run = await graph_service.create_entity_mentions(
|
||||
page_id=9998,
|
||||
user=TEST_USER,
|
||||
entity_names=found_entities
|
||||
)
|
||||
assert second_run == 0 # No new links
|
||||
|
||||
# Verify only one relationship exists
|
||||
verify_query = f"""
|
||||
MATCH (d:{user_doc_label}:Document {{page_id: 9998}})-[r:MENTIONS]->()
|
||||
RETURN count(r) as mention_count
|
||||
"""
|
||||
result = await graph_service.neo4j.execute_query(verify_query)
|
||||
assert result[0]["mention_count"] == 1
|
||||
|
||||
# Cleanup
|
||||
cleanup_query = f"""
|
||||
MATCH (n)
|
||||
WHERE (n:{user_base_label} OR n:{user_doc_label})
|
||||
AND (n.page_id = 9998 OR n.name = 'TestDockerEntity')
|
||||
DETACH DELETE n
|
||||
"""
|
||||
await graph_service.neo4j.execute_query(cleanup_query)
|
||||
|
||||
|
||||
# ============================================================================
|
||||
# Multi-Tenancy Tests
|
||||
# ============================================================================
|
||||
|
||||
class TestEntityLinkingMultiTenancy:
|
||||
"""Test multi-tenancy isolation in entity linking."""
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_user_isolation(self, graph_service):
|
||||
"""Test that entities are isolated by user."""
|
||||
from src.core.multi_tenancy import get_neo4j_user_base_label
|
||||
|
||||
user1_label = get_neo4j_user_base_label("user1")
|
||||
user2_label = get_neo4j_user_base_label("user2")
|
||||
|
||||
# Create entity for user1
|
||||
create_user1 = f"""
|
||||
CREATE (e:{user1_label}:Technology {{name: 'Docker', type: 'technology'}})
|
||||
RETURN e
|
||||
"""
|
||||
await graph_service.neo4j.execute_query(create_user1)
|
||||
|
||||
# Create entity for user2
|
||||
create_user2 = f"""
|
||||
CREATE (e:{user2_label}:Technology {{name: 'Docker', type: 'technology'}})
|
||||
RETURN e
|
||||
"""
|
||||
await graph_service.neo4j.execute_query(create_user2)
|
||||
|
||||
# Get entities for user1 - should only see user1's entities
|
||||
entities_user1 = await get_entities_with_paths(graph_service, "user1")
|
||||
entity_names_user1 = [e["name"] for e in entities_user1]
|
||||
|
||||
# Verify isolation
|
||||
assert "Docker" in entity_names_user1
|
||||
# We can't verify the exact count without knowing what else is in the DB,
|
||||
# but we verified we can retrieve entities for user1
|
||||
|
||||
# Cleanup
|
||||
await graph_service.neo4j.execute_query(f"MATCH (e:{user1_label}) WHERE e.name = 'Docker' DETACH DELETE e")
|
||||
await graph_service.neo4j.execute_query(f"MATCH (e:{user2_label}) WHERE e.name = 'Docker' DETACH DELETE e")
|
||||
|
||||
|
||||
# ============================================================================
|
||||
# Cleanup
|
||||
# ============================================================================
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_cleanup_entity_linking_test_data(neo4j_client):
|
||||
"""Clean up all test data created by entity linking tests."""
|
||||
from src.core.multi_tenancy import get_neo4j_user_base_label
|
||||
|
||||
for user in [TEST_USER, "user1", "user2"]:
|
||||
user_label = get_neo4j_user_base_label(user)
|
||||
|
||||
cleanup_query = f"""
|
||||
MATCH (n:{user_label})
|
||||
WHERE n.page_id IN [9999, 9998]
|
||||
OR n.name IN ['Docker', 'Kubernetes', 'PostgreSQL']
|
||||
DETACH DELETE n
|
||||
"""
|
||||
await neo4j_client.execute_query(cleanup_query)
|
||||
|
||||
print(f"\n✓ Cleaned up entity linking test data")
|
||||
@@ -1,249 +0,0 @@
|
||||
"""
|
||||
Tests for GraphService - knowledge graph operations.
|
||||
|
||||
Tests cover:
|
||||
- Document node creation with tags
|
||||
- Entity-stub page skipping
|
||||
- Entity extraction
|
||||
|
||||
Run with: pytest tests/test_graph_service.py -v -s
|
||||
"""
|
||||
|
||||
import pytest
|
||||
import pytest_asyncio
|
||||
from unittest.mock import AsyncMock, MagicMock, patch
|
||||
from typing import AsyncGenerator
|
||||
|
||||
from src.services.graph_service import GraphService
|
||||
|
||||
|
||||
# Test constants
|
||||
TEST_USER = "graph-tester"
|
||||
TEST_PAGE_ID = 123
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def mock_neo4j():
|
||||
"""Mock Neo4j client."""
|
||||
mock = AsyncMock()
|
||||
mock.execute_query = AsyncMock(return_value=[{"d": {"page_id": TEST_PAGE_ID}}])
|
||||
return mock
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def mock_wiki():
|
||||
"""Mock Wiki.js client."""
|
||||
mock = AsyncMock()
|
||||
return mock
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def graph_service(mock_neo4j, mock_wiki):
|
||||
"""Get GraphService with mocked dependencies."""
|
||||
return GraphService(
|
||||
neo4j_client=mock_neo4j,
|
||||
wikijs_client=mock_wiki
|
||||
)
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def sample_page():
|
||||
"""Sample wiki page data."""
|
||||
return {
|
||||
"id": TEST_PAGE_ID,
|
||||
"title": "Test Page",
|
||||
"path": f"users/{TEST_USER}/technology/docker",
|
||||
"content": "Docker is a containerization platform. It uses containers to run applications.",
|
||||
"tags": ["technology", "docker", "containers"]
|
||||
}
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def sample_page_without_tags():
|
||||
"""Sample wiki page without tags."""
|
||||
return {
|
||||
"id": TEST_PAGE_ID,
|
||||
"title": "Test Page No Tags",
|
||||
"path": f"users/{TEST_USER}/misc/test",
|
||||
"content": "This is a test page with no tags.",
|
||||
"tags": []
|
||||
}
|
||||
|
||||
|
||||
class TestDocumentNodeCreation:
|
||||
"""Test Document node creation in Neo4j."""
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_document_node_includes_tags(
|
||||
self,
|
||||
graph_service,
|
||||
mock_neo4j,
|
||||
mock_wiki,
|
||||
sample_page
|
||||
):
|
||||
"""Test that Document node is created with tags property."""
|
||||
mock_wiki.get_page = AsyncMock(return_value=sample_page)
|
||||
|
||||
result = await graph_service.update_from_page(
|
||||
page_id=TEST_PAGE_ID,
|
||||
user=TEST_USER
|
||||
)
|
||||
|
||||
# Verify execute_query was called
|
||||
assert mock_neo4j.execute_query.called
|
||||
assert result.success is True
|
||||
|
||||
# Find the document creation query
|
||||
calls = mock_neo4j.execute_query.call_args_list
|
||||
doc_creation_call = None
|
||||
for call in calls:
|
||||
query = call[0][0] if call[0] else ""
|
||||
if "MERGE" in query and "Document" in query and "tags" in query:
|
||||
doc_creation_call = call
|
||||
break
|
||||
|
||||
assert doc_creation_call is not None, "Document creation query with tags not found"
|
||||
|
||||
# Verify tags are in the query parameters
|
||||
params = doc_creation_call[0][1] if len(doc_creation_call[0]) > 1 else {}
|
||||
assert "tags" in params
|
||||
assert params["tags"] == ["technology", "docker", "containers"]
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_document_node_with_empty_tags(
|
||||
self,
|
||||
graph_service,
|
||||
mock_neo4j,
|
||||
mock_wiki,
|
||||
sample_page_without_tags
|
||||
):
|
||||
"""Test Document node creation with empty tags list."""
|
||||
mock_wiki.get_page = AsyncMock(return_value=sample_page_without_tags)
|
||||
|
||||
result = await graph_service.update_from_page(
|
||||
page_id=TEST_PAGE_ID,
|
||||
user=TEST_USER
|
||||
)
|
||||
|
||||
assert result.success is True
|
||||
|
||||
# Find the document creation query
|
||||
calls = mock_neo4j.execute_query.call_args_list
|
||||
doc_creation_call = None
|
||||
for call in calls:
|
||||
query = call[0][0] if call[0] else ""
|
||||
if "MERGE" in query and "Document" in query:
|
||||
doc_creation_call = call
|
||||
break
|
||||
|
||||
assert doc_creation_call is not None
|
||||
params = doc_creation_call[0][1] if len(doc_creation_call[0]) > 1 else {}
|
||||
assert "tags" in params
|
||||
assert params["tags"] == []
|
||||
|
||||
|
||||
class TestEntityStubSkipping:
|
||||
"""Test that entity-stub pages are skipped."""
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_skip_entity_stub_pages(
|
||||
self,
|
||||
graph_service,
|
||||
mock_neo4j,
|
||||
mock_wiki
|
||||
):
|
||||
"""Test that entity-stub tagged pages skip entity extraction."""
|
||||
stub_page = {
|
||||
"id": TEST_PAGE_ID,
|
||||
"title": "Auto Entity",
|
||||
"path": f"users/{TEST_USER}/entities/test",
|
||||
"content": "Auto-generated content.",
|
||||
"tags": ["entity-stub", "auto-generated"]
|
||||
}
|
||||
mock_wiki.get_page = AsyncMock(return_value=stub_page)
|
||||
|
||||
result = await graph_service.update_from_page(
|
||||
page_id=TEST_PAGE_ID,
|
||||
user=TEST_USER
|
||||
)
|
||||
|
||||
# Should return success but skip processing
|
||||
assert result.success is True
|
||||
# Neo4j should NOT be called for entity-stub pages
|
||||
assert mock_neo4j.execute_query.call_count == 0
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_skip_auto_generated_pages(
|
||||
self,
|
||||
graph_service,
|
||||
mock_neo4j,
|
||||
mock_wiki
|
||||
):
|
||||
"""Test that auto-generated tagged pages skip entity extraction."""
|
||||
auto_page = {
|
||||
"id": TEST_PAGE_ID,
|
||||
"title": "Auto Page",
|
||||
"path": f"users/{TEST_USER}/auto/test",
|
||||
"content": "Auto-generated content.",
|
||||
"tags": ["auto-generated"]
|
||||
}
|
||||
mock_wiki.get_page = AsyncMock(return_value=auto_page)
|
||||
|
||||
result = await graph_service.update_from_page(
|
||||
page_id=TEST_PAGE_ID,
|
||||
user=TEST_USER
|
||||
)
|
||||
|
||||
assert result.success is True
|
||||
assert mock_neo4j.execute_query.call_count == 0
|
||||
|
||||
|
||||
class TestPageNotFound:
|
||||
"""Test handling of missing pages."""
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_page_not_found_returns_failure(
|
||||
self,
|
||||
graph_service,
|
||||
mock_wiki
|
||||
):
|
||||
"""Test that missing page returns failure result."""
|
||||
mock_wiki.get_page = AsyncMock(return_value=None)
|
||||
|
||||
result = await graph_service.update_from_page(
|
||||
page_id=999,
|
||||
user=TEST_USER
|
||||
)
|
||||
|
||||
# The service catches the exception and returns a failed result
|
||||
assert result.success is False
|
||||
assert result.error_message is not None
|
||||
assert "not found" in result.error_message.lower()
|
||||
|
||||
|
||||
class TestEntityExtraction:
|
||||
"""Test entity extraction from page content."""
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_creates_document_and_entities(
|
||||
self,
|
||||
graph_service,
|
||||
mock_neo4j,
|
||||
mock_wiki,
|
||||
sample_page
|
||||
):
|
||||
"""Test that document and entity nodes are created."""
|
||||
mock_wiki.get_page = AsyncMock(return_value=sample_page)
|
||||
|
||||
result = await graph_service.update_from_page(
|
||||
page_id=TEST_PAGE_ID,
|
||||
user=TEST_USER
|
||||
)
|
||||
|
||||
assert result.success is True
|
||||
# Should have called neo4j at least once (for document node)
|
||||
assert mock_neo4j.execute_query.called
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
pytest.main([__file__, "-v", "-s"])
|
||||
@@ -1,711 +0,0 @@
|
||||
"""
|
||||
Comprehensive tests for HybridRAG system.
|
||||
|
||||
Tests cover all 6 phases:
|
||||
- Phase 0: Query Enhancement (keyword/synonym extraction)
|
||||
- Phase 1: Parallel Retrieval (vector + graph + web)
|
||||
- Phase 2: RRF Fusion
|
||||
- Phase 3: Enrichment (related dossiers)
|
||||
- Phase 4: LLM Re-ranking
|
||||
- Phase 5: Context Formatting
|
||||
- Phase 6: Persistence (search storage)
|
||||
|
||||
Uses 'llm-tester' user to avoid contaminating production data.
|
||||
|
||||
Run with: pytest tests/test_hybrid_rag.py -v -s
|
||||
"""
|
||||
|
||||
import pytest
|
||||
import pytest_asyncio
|
||||
from typing import AsyncGenerator
|
||||
import json
|
||||
|
||||
from src.clients.neo4j_client import Neo4jClient
|
||||
from src.clients.qdrant_client import QdrantClientWrapper
|
||||
from src.clients.wikijs_client import WikiJSClient
|
||||
from src.clients.searxng_client import SearXNGClient
|
||||
from src.clients.ollama_client import OllamaClient
|
||||
from src.services.hybrid_rag_service import HybridRAGService
|
||||
from src.services.vector_service import VectorService
|
||||
from src.services.graph_service import GraphService
|
||||
from src.models.hybrid_rag import HybridRAGConfig, HybridRAGRequest
|
||||
from src.config import get_settings
|
||||
|
||||
# Test user to isolate test data
|
||||
TEST_USER = "llm-tester"
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def settings():
|
||||
"""Get application settings."""
|
||||
return get_settings()
|
||||
|
||||
|
||||
@pytest_asyncio.fixture
|
||||
async def neo4j_client(settings) -> AsyncGenerator[Neo4jClient, None]:
|
||||
"""Get connected Neo4j client."""
|
||||
client = Neo4jClient(
|
||||
uri=settings.neo4j_uri,
|
||||
user=settings.neo4j_user,
|
||||
password=settings.neo4j_password
|
||||
)
|
||||
await client.connect()
|
||||
yield client
|
||||
await client.close()
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def qdrant_client(settings) -> QdrantClientWrapper:
|
||||
"""Get Qdrant client."""
|
||||
return QdrantClientWrapper(url=settings.qdrant_url)
|
||||
|
||||
|
||||
@pytest_asyncio.fixture
|
||||
async def wiki_client(settings) -> AsyncGenerator[WikiJSClient, None]:
|
||||
"""Get Wiki.js client."""
|
||||
client = WikiJSClient(
|
||||
base_url=settings.wikijs_url,
|
||||
username=settings.wikijs_username,
|
||||
password=settings.wikijs_password
|
||||
)
|
||||
yield client
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def searxng_client(settings) -> SearXNGClient:
|
||||
"""Get SearXNG client."""
|
||||
return SearXNGClient(base_url=settings.searxng_url)
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def ollama_client(settings) -> OllamaClient:
|
||||
"""Get Ollama client."""
|
||||
return OllamaClient(base_url=settings.ollama_url)
|
||||
|
||||
|
||||
@pytest_asyncio.fixture
|
||||
async def vector_service(qdrant_client, wiki_client, ollama_client):
|
||||
"""Get VectorService instance."""
|
||||
return VectorService(qdrant_client, wiki_client, ollama_client)
|
||||
|
||||
|
||||
@pytest_asyncio.fixture
|
||||
async def graph_service(neo4j_client, wiki_client):
|
||||
"""Get GraphService instance."""
|
||||
return GraphService(neo4j_client, wiki_client)
|
||||
|
||||
|
||||
@pytest_asyncio.fixture
|
||||
async def hybrid_rag_service(
|
||||
vector_service,
|
||||
graph_service,
|
||||
searxng_client,
|
||||
ollama_client,
|
||||
settings
|
||||
):
|
||||
"""Get HybridRAGService instance."""
|
||||
return HybridRAGService(
|
||||
vector_service=vector_service,
|
||||
graph_service=graph_service,
|
||||
searxng_client=searxng_client,
|
||||
ollama_client=ollama_client,
|
||||
settings=settings
|
||||
)
|
||||
|
||||
|
||||
@pytest_asyncio.fixture
|
||||
async def test_wiki_page(wiki_client):
|
||||
"""
|
||||
Create test wiki page for llm-tester user.
|
||||
|
||||
Creates a page about Docker and Kubernetes for testing.
|
||||
"""
|
||||
from src.core.multi_tenancy import get_wikijs_namespace
|
||||
|
||||
namespace = get_wikijs_namespace(TEST_USER)
|
||||
path = f"{namespace}/testing/docker-kubernetes"
|
||||
|
||||
# Create test page
|
||||
page_data = {
|
||||
"title": "Docker and Kubernetes Testing",
|
||||
"path": path,
|
||||
"content": """# Docker and Kubernetes
|
||||
|
||||
Docker is a containerization platform that packages applications into containers.
|
||||
Kubernetes (k8s) is an orchestration platform for managing Docker containers at scale.
|
||||
|
||||
## Key Technologies
|
||||
- Docker: Container runtime
|
||||
- Kubernetes: Orchestration platform
|
||||
- Helm: Package manager for Kubernetes
|
||||
- kubectl: Command-line tool for k8s
|
||||
|
||||
## Use Cases
|
||||
Our infrastructure uses Docker containers orchestrated by Kubernetes clusters.
|
||||
We deploy microservices using Helm charts and manage them with kubectl.
|
||||
""",
|
||||
"description": "Test page for HybridRAG testing",
|
||||
"tags": ["testing", "infrastructure", "docker"]
|
||||
}
|
||||
|
||||
try:
|
||||
# Delete if exists
|
||||
existing = await wiki_client.search_pages(query="Docker and Kubernetes Testing")
|
||||
for page in existing:
|
||||
if page.get("path") == path:
|
||||
await wiki_client.delete_page(page["id"])
|
||||
|
||||
# Create new
|
||||
page = await wiki_client.create_page(**page_data)
|
||||
yield page
|
||||
|
||||
# Cleanup
|
||||
try:
|
||||
await wiki_client.delete_page(page["id"])
|
||||
except:
|
||||
pass
|
||||
except Exception as e:
|
||||
pytest.skip(f"Could not create test page: {e}")
|
||||
|
||||
|
||||
@pytest_asyncio.fixture
|
||||
async def test_graph_data(graph_service, test_wiki_page):
|
||||
"""
|
||||
Populate graph with test data for llm-tester.
|
||||
|
||||
Extracts entities from test page.
|
||||
"""
|
||||
try:
|
||||
summary = await graph_service.update_from_page(
|
||||
page_id=test_wiki_page["id"],
|
||||
user=TEST_USER
|
||||
)
|
||||
yield summary
|
||||
except Exception as e:
|
||||
pytest.skip(f"Could not populate graph: {e}")
|
||||
|
||||
|
||||
@pytest_asyncio.fixture
|
||||
async def test_vector_data(vector_service, test_wiki_page):
|
||||
"""
|
||||
Populate vector DB with test data for llm-tester.
|
||||
|
||||
Creates embeddings from test page.
|
||||
"""
|
||||
try:
|
||||
summary = await vector_service.update_from_page(
|
||||
page_id=test_wiki_page["id"],
|
||||
user=TEST_USER
|
||||
)
|
||||
yield summary
|
||||
except Exception as e:
|
||||
pytest.skip(f"Could not populate vectors: {e}")
|
||||
|
||||
|
||||
# ============================================================================
|
||||
# Unit Tests - Individual Components
|
||||
# ============================================================================
|
||||
|
||||
class TestRRFFusion:
|
||||
"""Test Reciprocal Rank Fusion algorithm."""
|
||||
|
||||
def test_rrf_single_source(self, hybrid_rag_service):
|
||||
"""Test RRF with single source."""
|
||||
results_by_source = {
|
||||
"vector": [
|
||||
{"page_id": 1, "title": "Doc 1", "content": "test"},
|
||||
{"page_id": 2, "title": "Doc 2", "content": "test"}
|
||||
]
|
||||
}
|
||||
|
||||
fused = hybrid_rag_service._reciprocal_rank_fusion(results_by_source, k=60)
|
||||
|
||||
assert len(fused) == 2
|
||||
assert fused[0]["rrf_score"] > fused[1]["rrf_score"] # Rank 1 > Rank 2
|
||||
assert fused[0]["sources"] == ["vector"]
|
||||
|
||||
def test_rrf_multiple_sources_same_doc(self, hybrid_rag_service):
|
||||
"""Test RRF with same document from multiple sources."""
|
||||
results_by_source = {
|
||||
"vector": [{"page_id": 1, "title": "Doc 1", "content": "test"}],
|
||||
"graph": [{"page_id": 1, "title": "Doc 1", "content": ""}],
|
||||
}
|
||||
|
||||
fused = hybrid_rag_service._reciprocal_rank_fusion(results_by_source, k=60)
|
||||
|
||||
assert len(fused) == 1 # Deduplicated
|
||||
assert len(fused[0]["sources"]) == 2 # Both sources
|
||||
assert "vector" in fused[0]["sources"]
|
||||
assert "graph" in fused[0]["sources"]
|
||||
# RRF score should be sum: 1/(60+1) + 1/(60+1)
|
||||
expected_score = 1/61 + 1/61
|
||||
assert abs(fused[0]["rrf_score"] - expected_score) < 0.001
|
||||
|
||||
def test_rrf_web_results(self, hybrid_rag_service):
|
||||
"""Test RRF with web results (URL-based)."""
|
||||
results_by_source = {
|
||||
"web": [
|
||||
{"url": "https://example.com/1", "title": "Web 1", "content": "test"},
|
||||
{"url": "https://example.com/2", "title": "Web 2", "content": "test"}
|
||||
]
|
||||
}
|
||||
|
||||
fused = hybrid_rag_service._reciprocal_rank_fusion(results_by_source, k=60)
|
||||
|
||||
assert len(fused) == 2
|
||||
assert fused[0]["result"]["url"] == "https://example.com/1"
|
||||
|
||||
|
||||
class TestContextFormatting:
|
||||
"""Test context formatting for LLM."""
|
||||
|
||||
def test_format_basic(self, hybrid_rag_service):
|
||||
"""Test basic context formatting."""
|
||||
from src.models.hybrid_rag import HybridRAGResult
|
||||
|
||||
results = [
|
||||
HybridRAGResult(
|
||||
source_type="vector",
|
||||
title="Test Document",
|
||||
content="This is test content for formatting",
|
||||
page_id=1,
|
||||
rrf_score=0.5,
|
||||
final_rank=1,
|
||||
sources=["vector"]
|
||||
)
|
||||
]
|
||||
|
||||
context = hybrid_rag_service._format_context_for_llm(results)
|
||||
|
||||
assert "Test Document" in context
|
||||
assert "[VECTOR]" in context
|
||||
assert "test content" in context
|
||||
|
||||
def test_format_with_related_dossiers(self, hybrid_rag_service):
|
||||
"""Test context formatting with related dossiers."""
|
||||
from src.models.hybrid_rag import HybridRAGResult, RelatedDossier
|
||||
|
||||
results = [
|
||||
HybridRAGResult(
|
||||
source_type="vector+graph",
|
||||
title="Test Document",
|
||||
content="Content",
|
||||
page_id=1,
|
||||
rrf_score=0.5,
|
||||
final_rank=1,
|
||||
sources=["vector", "graph"],
|
||||
related_dossiers=[
|
||||
RelatedDossier(
|
||||
page_id=2,
|
||||
title="Related Doc",
|
||||
path="/test/related",
|
||||
tag="infrastructure",
|
||||
shared_entities=5
|
||||
)
|
||||
]
|
||||
)
|
||||
]
|
||||
|
||||
context = hybrid_rag_service._format_context_for_llm(results)
|
||||
|
||||
assert "Related research: infrastructure" in context
|
||||
|
||||
|
||||
# ============================================================================
|
||||
# Integration Tests - Phase Testing
|
||||
# ============================================================================
|
||||
|
||||
class TestPhase0_QueryEnhancement:
|
||||
"""Test Phase 0: Query Enhancement (keyword/synonym extraction)."""
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_extract_keywords_basic(self, hybrid_rag_service):
|
||||
"""Test basic keyword extraction."""
|
||||
query = "Docker container orchestration with Kubernetes"
|
||||
|
||||
keywords_data = await hybrid_rag_service._extract_keywords_and_synonyms(query)
|
||||
|
||||
assert "core_keywords" in keywords_data
|
||||
assert "entities" in keywords_data
|
||||
assert "synonyms" in keywords_data
|
||||
assert "expansions" in keywords_data
|
||||
|
||||
# Should extract Docker and Kubernetes
|
||||
all_terms = (
|
||||
keywords_data["core_keywords"] +
|
||||
keywords_data["entities"]
|
||||
)
|
||||
assert any("docker" in term.lower() for term in all_terms)
|
||||
assert any("kubernetes" in term.lower() or "k8s" in term.lower() for term in all_terms)
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_extract_keywords_with_abbreviations(self, hybrid_rag_service):
|
||||
"""Test keyword extraction handles abbreviations."""
|
||||
query = "k8s cluster management"
|
||||
|
||||
keywords_data = await hybrid_rag_service._extract_keywords_and_synonyms(query)
|
||||
|
||||
# Should expand k8s to kubernetes
|
||||
all_data = json.dumps(keywords_data).lower()
|
||||
assert "k8s" in all_data or "kubernetes" in all_data
|
||||
|
||||
|
||||
class TestPhase1_ParallelRetrieval:
|
||||
"""Test Phase 1: Parallel Retrieval."""
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_parallel_retrieval_all_sources(
|
||||
self,
|
||||
hybrid_rag_service,
|
||||
test_wiki_page,
|
||||
test_graph_data,
|
||||
test_vector_data
|
||||
):
|
||||
"""Test parallel retrieval from all sources."""
|
||||
config = HybridRAGConfig(
|
||||
enable_vector=True,
|
||||
enable_graph=True,
|
||||
enable_web=True,
|
||||
vector_limit=5,
|
||||
graph_limit=5,
|
||||
web_limit=3
|
||||
)
|
||||
|
||||
keywords_data = {
|
||||
"core_keywords": ["docker", "kubernetes"],
|
||||
"entities": ["Docker", "Kubernetes"],
|
||||
"synonyms": {"docker": ["container"], "kubernetes": ["k8s"]},
|
||||
"expansions": {"k8s": ["kubernetes"]}
|
||||
}
|
||||
|
||||
results = await hybrid_rag_service._retrieve_parallel(
|
||||
query="docker kubernetes",
|
||||
user=TEST_USER,
|
||||
config=config,
|
||||
keywords_data=keywords_data
|
||||
)
|
||||
|
||||
assert "vector" in results
|
||||
assert "graph" in results
|
||||
assert "web" in results
|
||||
assert "timing" in results
|
||||
|
||||
# Should have timing for each source
|
||||
assert results["timing"]["vector_ms"] >= 0
|
||||
assert results["timing"]["graph_ms"] >= 0
|
||||
assert results["timing"]["web_ms"] >= 0
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_parallel_retrieval_graceful_degradation(self, hybrid_rag_service):
|
||||
"""Test graceful degradation when sources fail."""
|
||||
config = HybridRAGConfig(
|
||||
enable_vector=True,
|
||||
enable_graph=True,
|
||||
enable_web=True
|
||||
)
|
||||
|
||||
keywords_data = {"core_keywords": ["test"], "entities": [], "synonyms": {}, "expansions": {}}
|
||||
|
||||
# Even if some sources fail, should return results from working sources
|
||||
results = await hybrid_rag_service._retrieve_parallel(
|
||||
query="test query",
|
||||
user=TEST_USER,
|
||||
config=config,
|
||||
keywords_data=keywords_data
|
||||
)
|
||||
|
||||
# Should have all keys even if empty
|
||||
assert "vector" in results
|
||||
assert "graph" in results
|
||||
assert "web" in results
|
||||
|
||||
|
||||
class TestPhase3_Enrichment:
|
||||
"""Test Phase 3: Graph Enrichment."""
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_enrich_with_related_dossiers(
|
||||
self,
|
||||
hybrid_rag_service,
|
||||
graph_service,
|
||||
test_wiki_page,
|
||||
test_graph_data
|
||||
):
|
||||
"""Test enriching results with related dossiers."""
|
||||
# Create mock fused results
|
||||
fused_results = [
|
||||
{
|
||||
"result": {
|
||||
"page_id": test_wiki_page["id"],
|
||||
"title": test_wiki_page["title"],
|
||||
"content": "test"
|
||||
},
|
||||
"rrf_score": 0.5,
|
||||
"sources": ["vector"]
|
||||
}
|
||||
]
|
||||
|
||||
enriched = await hybrid_rag_service._enrich_with_related_dossiers(
|
||||
fused_results,
|
||||
user=TEST_USER
|
||||
)
|
||||
|
||||
assert len(enriched) == 1
|
||||
assert "related_dossiers" in enriched[0]
|
||||
# May or may not have related docs depending on graph state
|
||||
assert isinstance(enriched[0]["related_dossiers"], list)
|
||||
|
||||
|
||||
class TestPhase6_Persistence:
|
||||
"""Test Phase 6: Search Persistence."""
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_persist_search_creates_node(
|
||||
self,
|
||||
hybrid_rag_service,
|
||||
neo4j_client,
|
||||
test_wiki_page
|
||||
):
|
||||
"""Test that search persistence creates SearchQuery node."""
|
||||
keywords_data = {
|
||||
"core_keywords": ["docker", "kubernetes"],
|
||||
"entities": [],
|
||||
"synonyms": {},
|
||||
"expansions": {}
|
||||
}
|
||||
|
||||
raw_results = {
|
||||
"vector": [{"page_id": test_wiki_page["id"], "title": "Test", "content": "test"}],
|
||||
"graph": [],
|
||||
"web": []
|
||||
}
|
||||
|
||||
final_results = [
|
||||
{
|
||||
"result": {"page_id": test_wiki_page["id"], "title": "Test"},
|
||||
"rrf_score": 0.5,
|
||||
"final_rank": 1,
|
||||
"sources": ["vector"]
|
||||
}
|
||||
]
|
||||
|
||||
timing = {"total_ms": 1000}
|
||||
|
||||
search_id = await hybrid_rag_service._persist_search_for_librarian(
|
||||
query="test query",
|
||||
user=TEST_USER,
|
||||
keywords_data=keywords_data,
|
||||
raw_results=raw_results,
|
||||
final_results=final_results,
|
||||
timing=timing
|
||||
)
|
||||
|
||||
assert search_id is not None
|
||||
|
||||
# Verify SearchQuery node was created
|
||||
from src.core.multi_tenancy import get_neo4j_user_base_label
|
||||
user_label = get_neo4j_user_base_label(TEST_USER)
|
||||
|
||||
query = f"""
|
||||
MATCH (sq:{user_label}_SearchQuery:SearchQuery {{id: $search_id}})
|
||||
RETURN sq.query as query, sq.processed as processed
|
||||
"""
|
||||
|
||||
result = await neo4j_client.execute_query(query, {"search_id": search_id})
|
||||
assert len(result) == 1
|
||||
assert result[0]["query"] == "test query"
|
||||
assert result[0]["processed"] == False
|
||||
|
||||
# Cleanup
|
||||
cleanup_query = f"""
|
||||
MATCH (sq:{user_label}_SearchQuery:SearchQuery {{id: $search_id}})
|
||||
DETACH DELETE sq
|
||||
"""
|
||||
await neo4j_client.execute_query(cleanup_query, {"search_id": search_id})
|
||||
|
||||
|
||||
# ============================================================================
|
||||
# End-to-End Tests
|
||||
# ============================================================================
|
||||
|
||||
class TestHybridRAG_EndToEnd:
|
||||
"""End-to-end tests for complete HybridRAG flow."""
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_full_search_pipeline(
|
||||
self,
|
||||
hybrid_rag_service,
|
||||
test_wiki_page,
|
||||
test_graph_data,
|
||||
test_vector_data
|
||||
):
|
||||
"""
|
||||
Test complete HybridRAG search pipeline with all 6 phases.
|
||||
|
||||
This is the main end-to-end test that validates:
|
||||
- Phase 0: Query enhancement
|
||||
- Phase 1: Parallel retrieval
|
||||
- Phase 2: RRF fusion
|
||||
- Phase 3: Enrichment
|
||||
- Phase 4: Re-ranking
|
||||
- Phase 5: Context formatting
|
||||
- Phase 6: Persistence
|
||||
"""
|
||||
query = "How does Docker work with Kubernetes?"
|
||||
config = HybridRAGConfig(
|
||||
vector_limit=5,
|
||||
graph_limit=5,
|
||||
web_limit=3,
|
||||
enable_reranking=True,
|
||||
enable_enrichment=True,
|
||||
final_result_count=10
|
||||
)
|
||||
|
||||
# Execute full search
|
||||
response = await hybrid_rag_service.search(
|
||||
query=query,
|
||||
user=TEST_USER,
|
||||
config=config
|
||||
)
|
||||
|
||||
# Validate response structure
|
||||
assert response.query == query
|
||||
assert response.keywords is not None
|
||||
assert response.results is not None
|
||||
assert response.context is not None
|
||||
assert response.source_counts is not None
|
||||
assert response.total_results >= 0
|
||||
assert response.timing is not None
|
||||
assert response.config_used == config
|
||||
assert response.search_id is not None
|
||||
|
||||
# Validate timing breakdown
|
||||
assert response.timing.query_enhancement_ms >= 0
|
||||
assert response.timing.vector_ms >= 0
|
||||
assert response.timing.graph_ms >= 0
|
||||
assert response.timing.web_ms >= 0
|
||||
assert response.timing.fusion_ms >= 0
|
||||
assert response.timing.enrichment_ms >= 0
|
||||
assert response.timing.reranking_ms >= 0
|
||||
assert response.timing.persistence_ms >= 0
|
||||
assert response.timing.total_ms >= 0
|
||||
|
||||
# Validate keywords extraction
|
||||
assert len(response.keywords.core_keywords) > 0
|
||||
|
||||
# Validate context is formatted
|
||||
assert len(response.context) > 0
|
||||
|
||||
# Log results for inspection
|
||||
print(f"\n=== HybridRAG E2E Test Results ===")
|
||||
print(f"Query: {response.query}")
|
||||
print(f"Total Results: {response.total_results}")
|
||||
print(f"Source Counts: {response.source_counts}")
|
||||
print(f"Keywords: {response.keywords.core_keywords}")
|
||||
print(f"Total Time: {response.timing.total_ms:.0f}ms")
|
||||
print(f"Search ID: {response.search_id}")
|
||||
|
||||
if response.results:
|
||||
print(f"\nTop Result:")
|
||||
top = response.results[0]
|
||||
print(f" Title: {top.title}")
|
||||
print(f" Source: {top.source_type}")
|
||||
print(f" RRF Score: {top.rrf_score:.4f}")
|
||||
print(f" Rank: {top.final_rank}")
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_search_with_disabled_sources(
|
||||
self,
|
||||
hybrid_rag_service,
|
||||
test_wiki_page,
|
||||
test_vector_data
|
||||
):
|
||||
"""Test HybridRAG with some sources disabled."""
|
||||
config = HybridRAGConfig(
|
||||
enable_vector=True,
|
||||
enable_graph=False, # Disabled
|
||||
enable_web=False, # Disabled
|
||||
enable_reranking=False,
|
||||
final_result_count=5
|
||||
)
|
||||
|
||||
response = await hybrid_rag_service.search(
|
||||
query="docker containers",
|
||||
user=TEST_USER,
|
||||
config=config
|
||||
)
|
||||
|
||||
# Should only have vector results
|
||||
assert response.total_results >= 0
|
||||
if response.total_results > 0:
|
||||
assert all(
|
||||
"vector" in result.sources
|
||||
for result in response.results
|
||||
)
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_search_performance_target(
|
||||
self,
|
||||
hybrid_rag_service,
|
||||
test_wiki_page,
|
||||
test_graph_data,
|
||||
test_vector_data
|
||||
):
|
||||
"""Test that search completes within performance target (<3.5s)."""
|
||||
import time
|
||||
|
||||
config = HybridRAGConfig()
|
||||
|
||||
start = time.time()
|
||||
response = await hybrid_rag_service.search(
|
||||
query="kubernetes orchestration",
|
||||
user=TEST_USER,
|
||||
config=config
|
||||
)
|
||||
duration_ms = (time.time() - start) * 1000
|
||||
|
||||
print(f"\nPerformance: {duration_ms:.0f}ms (target: <3500ms)")
|
||||
|
||||
# Soft assertion - warn if exceeds target
|
||||
if duration_ms > 3500:
|
||||
print(f"WARNING: Search exceeded 3.5s target ({duration_ms:.0f}ms)")
|
||||
|
||||
|
||||
# ============================================================================
|
||||
# Cleanup Tests
|
||||
# ============================================================================
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_cleanup_test_data(neo4j_client, qdrant_client):
|
||||
"""
|
||||
Cleanup test data for llm-tester user.
|
||||
|
||||
Run this to clean up test data:
|
||||
pytest tests/test_hybrid_rag.py::test_cleanup_test_data -v -s
|
||||
"""
|
||||
from src.core.multi_tenancy import (
|
||||
get_neo4j_user_base_label,
|
||||
get_neo4j_user_label,
|
||||
get_qdrant_collection_name
|
||||
)
|
||||
|
||||
# Clean Neo4j
|
||||
user_base_label = get_neo4j_user_base_label(TEST_USER)
|
||||
user_doc_label = get_neo4j_user_label(TEST_USER)
|
||||
|
||||
# Delete all test user nodes
|
||||
delete_query = f"""
|
||||
MATCH (n)
|
||||
WHERE n:{user_base_label} OR n:{user_doc_label}
|
||||
DETACH DELETE n
|
||||
"""
|
||||
await neo4j_client.execute_query(delete_query, {})
|
||||
|
||||
# Clean Qdrant
|
||||
collection_name = get_qdrant_collection_name(TEST_USER)
|
||||
try:
|
||||
await qdrant_client.delete_collection(collection_name)
|
||||
except:
|
||||
pass
|
||||
|
||||
print(f"\n✓ Cleaned up test data for user: {TEST_USER}")
|
||||
@@ -1,314 +0,0 @@
|
||||
"""
|
||||
Tests for IngestionService - document ingestion operations.
|
||||
|
||||
Tests cover:
|
||||
- Single page ingestion
|
||||
- Batch ingestion
|
||||
- Full re-index (ingest_all_pages)
|
||||
- list_all_pages usage
|
||||
|
||||
Run with: pytest tests/test_ingestion.py -v -s
|
||||
"""
|
||||
|
||||
import pytest
|
||||
import pytest_asyncio
|
||||
from unittest.mock import AsyncMock, MagicMock, patch
|
||||
from typing import AsyncGenerator
|
||||
|
||||
from src.services.ingestion_service import IngestionService
|
||||
from src.models.ingestion import (
|
||||
IngestionRequest,
|
||||
IngestionResult,
|
||||
BatchIngestionRequest,
|
||||
BatchIngestionResult
|
||||
)
|
||||
|
||||
|
||||
# Test constants
|
||||
TEST_USER = "ingestion-tester"
|
||||
TEST_PAGE_ID = 456
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def mock_vector_service():
|
||||
"""Mock Vector service."""
|
||||
mock = AsyncMock()
|
||||
mock.update_from_page = AsyncMock(return_value=MagicMock(
|
||||
chunks_created=2,
|
||||
chunks_deleted=0,
|
||||
success=True
|
||||
))
|
||||
return mock
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def mock_graph_service():
|
||||
"""Mock Graph service."""
|
||||
mock = AsyncMock()
|
||||
mock.update_from_page = AsyncMock(return_value=MagicMock(
|
||||
entities_extracted=5,
|
||||
relationships_created=5,
|
||||
success=True
|
||||
))
|
||||
mock.create_entity_mention_links = AsyncMock(return_value=3)
|
||||
mock.get_all_entities = AsyncMock(return_value=[])
|
||||
return mock
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def mock_wiki_client():
|
||||
"""Mock Wiki.js client."""
|
||||
mock = AsyncMock()
|
||||
mock.get_page = AsyncMock(return_value={
|
||||
"id": TEST_PAGE_ID,
|
||||
"title": "Test Page",
|
||||
"path": f"users/{TEST_USER}/test",
|
||||
"content": "Test content here.",
|
||||
"tags": ["test"]
|
||||
})
|
||||
mock.list_all_pages = AsyncMock(return_value=[
|
||||
{"id": 1, "path": f"users/{TEST_USER}/page1", "title": "Page 1"},
|
||||
{"id": 2, "path": f"users/{TEST_USER}/page2", "title": "Page 2"},
|
||||
{"id": 3, "path": f"users/{TEST_USER}/page3", "title": "Page 3"},
|
||||
])
|
||||
return mock
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def ingestion_service(mock_vector_service, mock_graph_service, mock_wiki_client):
|
||||
"""Get IngestionService with mocked dependencies."""
|
||||
return IngestionService(
|
||||
vector_service=mock_vector_service,
|
||||
graph_service=mock_graph_service,
|
||||
wiki_client=mock_wiki_client
|
||||
)
|
||||
|
||||
|
||||
class TestSinglePageIngestion:
|
||||
"""Test single page ingestion."""
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_ingest_page_success(
|
||||
self,
|
||||
ingestion_service,
|
||||
mock_wiki_client
|
||||
):
|
||||
"""Test successful page ingestion."""
|
||||
result = await ingestion_service.ingest_page(
|
||||
page_id=TEST_PAGE_ID,
|
||||
user=TEST_USER
|
||||
)
|
||||
|
||||
assert result.success is True
|
||||
assert result.page_id == TEST_PAGE_ID
|
||||
mock_wiki_client.get_page.assert_called_once_with(TEST_PAGE_ID)
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_ingest_page_not_found(
|
||||
self,
|
||||
ingestion_service,
|
||||
mock_wiki_client
|
||||
):
|
||||
"""Test ingestion when page not found."""
|
||||
mock_wiki_client.get_page.return_value = None
|
||||
|
||||
result = await ingestion_service.ingest_page(
|
||||
page_id=999,
|
||||
user=TEST_USER
|
||||
)
|
||||
|
||||
assert result.success is False
|
||||
assert "not found" in result.error.lower()
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_ingest_page_skip_vectors(
|
||||
self,
|
||||
ingestion_service,
|
||||
mock_vector_service,
|
||||
mock_graph_service
|
||||
):
|
||||
"""Test ingestion with vectors skipped."""
|
||||
result = await ingestion_service.ingest_page(
|
||||
page_id=TEST_PAGE_ID,
|
||||
user=TEST_USER,
|
||||
skip_vectors=True
|
||||
)
|
||||
|
||||
assert result.success is True
|
||||
# Vector service should not be called
|
||||
mock_vector_service.update_from_page.assert_not_called()
|
||||
# Graph service should still be called
|
||||
mock_graph_service.update_from_page.assert_called_once()
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_ingest_page_skip_graph(
|
||||
self,
|
||||
ingestion_service,
|
||||
mock_vector_service,
|
||||
mock_graph_service
|
||||
):
|
||||
"""Test ingestion with graph skipped."""
|
||||
result = await ingestion_service.ingest_page(
|
||||
page_id=TEST_PAGE_ID,
|
||||
user=TEST_USER,
|
||||
skip_graph=True
|
||||
)
|
||||
|
||||
assert result.success is True
|
||||
# Vector service should be called
|
||||
mock_vector_service.update_from_page.assert_called_once()
|
||||
# Graph service should not be called
|
||||
mock_graph_service.update_from_page.assert_not_called()
|
||||
|
||||
|
||||
class TestBatchIngestion:
|
||||
"""Test batch page ingestion."""
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_ingest_batch_success(
|
||||
self,
|
||||
ingestion_service,
|
||||
mock_wiki_client
|
||||
):
|
||||
"""Test successful batch ingestion."""
|
||||
result = await ingestion_service.ingest_batch(
|
||||
page_ids=[1, 2, 3],
|
||||
user=TEST_USER,
|
||||
max_concurrent=2
|
||||
)
|
||||
|
||||
assert result.total_pages == 3
|
||||
assert result.successful == 3
|
||||
assert result.failed == 0
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_ingest_batch_with_failures(
|
||||
self,
|
||||
ingestion_service,
|
||||
mock_wiki_client
|
||||
):
|
||||
"""Test batch ingestion with some failures."""
|
||||
# Make page 2 not found
|
||||
def get_page_side_effect(page_id):
|
||||
if page_id == 2:
|
||||
return None
|
||||
return {
|
||||
"id": page_id,
|
||||
"title": f"Page {page_id}",
|
||||
"path": f"users/{TEST_USER}/page{page_id}",
|
||||
"content": "Content",
|
||||
"tags": []
|
||||
}
|
||||
|
||||
mock_wiki_client.get_page.side_effect = get_page_side_effect
|
||||
|
||||
result = await ingestion_service.ingest_batch(
|
||||
page_ids=[1, 2, 3],
|
||||
user=TEST_USER
|
||||
)
|
||||
|
||||
assert result.total_pages == 3
|
||||
assert result.successful == 2
|
||||
assert result.failed == 1
|
||||
|
||||
|
||||
class TestIngestAllPages:
|
||||
"""Test full re-index (ingest_all_pages)."""
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_ingest_all_uses_list_all_pages(
|
||||
self,
|
||||
ingestion_service,
|
||||
mock_wiki_client
|
||||
):
|
||||
"""Test that ingest_all_pages uses list_all_pages (not search)."""
|
||||
result = await ingestion_service.ingest_all_pages(
|
||||
user=TEST_USER
|
||||
)
|
||||
|
||||
# Should use list_all_pages, not search_pages
|
||||
mock_wiki_client.list_all_pages.assert_called_once()
|
||||
# Should have processed 3 pages from the mock
|
||||
assert result.total_pages == 3
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_ingest_all_with_path_prefix(
|
||||
self,
|
||||
ingestion_service,
|
||||
mock_wiki_client
|
||||
):
|
||||
"""Test ingest_all_pages with path prefix filter."""
|
||||
await ingestion_service.ingest_all_pages(
|
||||
user=TEST_USER,
|
||||
path_prefix=f"users/{TEST_USER}/technology"
|
||||
)
|
||||
|
||||
mock_wiki_client.list_all_pages.assert_called_once_with(
|
||||
path_prefix=f"users/{TEST_USER}/technology"
|
||||
)
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_ingest_all_empty_wiki(
|
||||
self,
|
||||
ingestion_service,
|
||||
mock_wiki_client
|
||||
):
|
||||
"""Test ingest_all_pages when no pages found."""
|
||||
mock_wiki_client.list_all_pages.return_value = []
|
||||
|
||||
result = await ingestion_service.ingest_all_pages(
|
||||
user=TEST_USER
|
||||
)
|
||||
|
||||
assert result.total_pages == 0
|
||||
assert result.successful == 0
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_ingest_all_respects_max_concurrent(
|
||||
self,
|
||||
ingestion_service,
|
||||
mock_wiki_client
|
||||
):
|
||||
"""Test that max_concurrent parameter is passed through."""
|
||||
# Create many pages
|
||||
mock_wiki_client.list_all_pages.return_value = [
|
||||
{"id": i, "path": f"users/{TEST_USER}/page{i}", "title": f"Page {i}"}
|
||||
for i in range(20)
|
||||
]
|
||||
|
||||
result = await ingestion_service.ingest_all_pages(
|
||||
user=TEST_USER,
|
||||
max_concurrent=5
|
||||
)
|
||||
|
||||
assert result.total_pages == 20
|
||||
|
||||
|
||||
class TestIngestionModels:
|
||||
"""Test ingestion request/response models."""
|
||||
|
||||
def test_ingestion_request_defaults(self):
|
||||
"""Test IngestionRequest default values."""
|
||||
request = IngestionRequest(
|
||||
page_id=123,
|
||||
user="testuser"
|
||||
)
|
||||
assert request.page_id == 123
|
||||
assert request.user == "testuser"
|
||||
assert request.force_refresh is False
|
||||
assert request.skip_vectors is False
|
||||
assert request.skip_graph is False
|
||||
|
||||
def test_batch_ingestion_request(self):
|
||||
"""Test BatchIngestionRequest."""
|
||||
request = BatchIngestionRequest(
|
||||
page_ids=[1, 2, 3],
|
||||
user="testuser",
|
||||
max_concurrent=5
|
||||
)
|
||||
assert len(request.page_ids) == 3
|
||||
assert request.max_concurrent == 5
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
pytest.main([__file__, "-v", "-s"])
|
||||
@@ -1,334 +0,0 @@
|
||||
"""Integration tests for Library Desk service clients.
|
||||
|
||||
These tests require actual service connectivity:
|
||||
- Neo4j running at bolt://neo4j:7687
|
||||
- Qdrant running at http://qdrant:6333
|
||||
- Wiki.js running at http://wiki:3000
|
||||
- SearXNG running at http://searxng:8080
|
||||
- Ollama running at http://ollama:11434
|
||||
|
||||
Run with: pytest tests/test_integration.py -v
|
||||
"""
|
||||
|
||||
import pytest
|
||||
import pytest_asyncio
|
||||
from typing import AsyncGenerator
|
||||
|
||||
from src.clients.neo4j_client import Neo4jClient
|
||||
from src.clients.qdrant_client import QdrantClientWrapper
|
||||
from src.clients.wikijs_client import WikiJSClient
|
||||
from src.clients.searxng_client import SearXNGClient
|
||||
from src.clients.ollama_client import OllamaClient
|
||||
from src.jobs.job_manager import JobManager, JobType, JobStatus
|
||||
from src.config import get_settings
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def settings():
|
||||
"""Get application settings."""
|
||||
return get_settings()
|
||||
|
||||
|
||||
@pytest_asyncio.fixture
|
||||
async def neo4j_client(settings) -> AsyncGenerator[Neo4jClient, None]:
|
||||
"""Get connected Neo4j client."""
|
||||
client = Neo4jClient(
|
||||
uri=settings.neo4j_uri,
|
||||
user=settings.neo4j_user,
|
||||
password=settings.neo4j_password
|
||||
)
|
||||
await client.connect()
|
||||
yield client
|
||||
await client.close()
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def qdrant_client(settings) -> QdrantClientWrapper:
|
||||
"""Get Qdrant client."""
|
||||
return QdrantClientWrapper(url=settings.qdrant_url)
|
||||
|
||||
|
||||
@pytest_asyncio.fixture
|
||||
async def wikijs_client(settings) -> AsyncGenerator[WikiJSClient, None]:
|
||||
"""Get Wiki.js client."""
|
||||
client = WikiJSClient(
|
||||
base_url=settings.wikijs_url,
|
||||
api_key=settings.wikijs_api_key
|
||||
)
|
||||
yield client
|
||||
await client.close()
|
||||
|
||||
|
||||
@pytest_asyncio.fixture
|
||||
async def searxng_client(settings) -> AsyncGenerator[SearXNGClient, None]:
|
||||
"""Get SearXNG client."""
|
||||
client = SearXNGClient(base_url=settings.searxng_url)
|
||||
yield client
|
||||
await client.close()
|
||||
|
||||
|
||||
@pytest_asyncio.fixture
|
||||
async def ollama_client(settings) -> AsyncGenerator[OllamaClient, None]:
|
||||
"""Get Ollama client."""
|
||||
client = OllamaClient(
|
||||
base_url=settings.ollama_url,
|
||||
model=settings.ollama_model
|
||||
)
|
||||
yield client
|
||||
await client.close()
|
||||
|
||||
|
||||
@pytest_asyncio.fixture
|
||||
async def job_manager(settings) -> AsyncGenerator[JobManager, None]:
|
||||
"""Get job manager."""
|
||||
manager = JobManager(redis_url=settings.redis_url)
|
||||
await manager.connect()
|
||||
yield manager
|
||||
await manager.close()
|
||||
|
||||
|
||||
class TestNeo4jIntegration:
|
||||
"""Test Neo4j connectivity and basic operations."""
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_connection(self, neo4j_client):
|
||||
"""Test Neo4j connection."""
|
||||
result = await neo4j_client.execute_query("RETURN 1 as test", {})
|
||||
assert result
|
||||
assert result[0]["test"] == 1
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_create_and_get_document(self, neo4j_client, test_user):
|
||||
"""Test creating and retrieving a document node."""
|
||||
doc_id = "test_doc_integration"
|
||||
|
||||
# Create document
|
||||
doc = await neo4j_client.create_document_node(
|
||||
user=test_user,
|
||||
doc_id=doc_id,
|
||||
properties={
|
||||
"title": "Integration Test Document",
|
||||
"source": "test"
|
||||
}
|
||||
)
|
||||
assert doc is not None
|
||||
assert doc["id"] == doc_id
|
||||
|
||||
# Retrieve document
|
||||
retrieved = await neo4j_client.get_document_node(test_user, doc_id)
|
||||
assert retrieved is not None
|
||||
assert retrieved["id"] == doc_id
|
||||
assert retrieved["title"] == "Integration Test Document"
|
||||
|
||||
# Cleanup
|
||||
deleted = await neo4j_client.delete_document_node(test_user, doc_id)
|
||||
assert deleted is True
|
||||
|
||||
|
||||
class TestQdrantIntegration:
|
||||
"""Test Qdrant connectivity and basic operations."""
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_collection_creation(self, qdrant_client, test_user):
|
||||
"""Test creating a collection."""
|
||||
await qdrant_client.ensure_collection(test_user)
|
||||
|
||||
collection_name = qdrant_client.get_collection_name(test_user)
|
||||
collections = qdrant_client.client.get_collections()
|
||||
collection_names = [c.name for c in collections.collections]
|
||||
|
||||
assert collection_name in collection_names
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_upsert_and_search(self, qdrant_client, test_user):
|
||||
"""Test upserting and searching chunks."""
|
||||
await qdrant_client.ensure_collection(test_user)
|
||||
|
||||
# Create test chunks with 768-dimensional embeddings
|
||||
chunks = [
|
||||
{"content": "Test chunk 1", "metadata": {}},
|
||||
{"content": "Test chunk 2", "metadata": {}}
|
||||
]
|
||||
embeddings = [
|
||||
[0.1] * 768,
|
||||
[0.2] * 768
|
||||
]
|
||||
|
||||
# Upsert
|
||||
count = await qdrant_client.upsert_document_chunks(
|
||||
user=test_user,
|
||||
doc_id="test_doc_qdrant",
|
||||
chunks=chunks,
|
||||
embeddings=embeddings
|
||||
)
|
||||
assert count == 2
|
||||
|
||||
# Search
|
||||
results = await qdrant_client.search(
|
||||
user=test_user,
|
||||
query_vector=[0.1] * 768,
|
||||
limit=5,
|
||||
score_threshold=0.0
|
||||
)
|
||||
assert len(results) > 0
|
||||
|
||||
# Cleanup
|
||||
deleted = await qdrant_client.delete_document(test_user, "test_doc_qdrant")
|
||||
assert deleted is True
|
||||
|
||||
|
||||
class TestWikiJSIntegration:
|
||||
"""Test Wiki.js connectivity and basic operations."""
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_list_pages(self, wikijs_client):
|
||||
"""Test listing pages."""
|
||||
pages = await wikijs_client.list_pages(limit=10)
|
||||
assert isinstance(pages, list)
|
||||
# May be empty if wiki is new
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_list_all_pages(self, wikijs_client):
|
||||
"""Test listing all pages with pagination support."""
|
||||
pages = await wikijs_client.list_all_pages(path_prefix="users/")
|
||||
assert isinstance(pages, list)
|
||||
# Verify each page has expected fields
|
||||
for page in pages[:5]: # Check first 5
|
||||
assert "id" in page
|
||||
assert "path" in page
|
||||
assert "title" in page
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_get_taxonomy_structure(self, wikijs_client):
|
||||
"""Test getting taxonomy structure for a user."""
|
||||
taxonomy = await wikijs_client.get_taxonomy_structure("users/jpmschweitzer")
|
||||
assert isinstance(taxonomy, dict)
|
||||
# Each key should be a category, value should be list of subcategories
|
||||
for category, subcategories in taxonomy.items():
|
||||
assert isinstance(category, str)
|
||||
assert isinstance(subcategories, list)
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_search_pages(self, wikijs_client):
|
||||
"""Test searching pages."""
|
||||
results = await wikijs_client.search_pages("test")
|
||||
assert isinstance(results, list)
|
||||
|
||||
|
||||
class TestSearXNGIntegration:
|
||||
"""Test SearXNG connectivity and search."""
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_general_search(self, searxng_client):
|
||||
"""Test general web search."""
|
||||
results = await searxng_client.search_general("python programming", limit=5)
|
||||
assert isinstance(results, list)
|
||||
if results:
|
||||
assert "title" in results[0]
|
||||
assert "url" in results[0]
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_search_with_suggestions(self, searxng_client):
|
||||
"""Test getting search suggestions."""
|
||||
suggestions = await searxng_client.get_suggestions("pytho")
|
||||
assert isinstance(suggestions, list)
|
||||
|
||||
|
||||
class TestOllamaIntegration:
|
||||
"""Test Ollama connectivity and embeddings."""
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_health_check(self, ollama_client):
|
||||
"""Test Ollama health check."""
|
||||
is_healthy = await ollama_client.health_check()
|
||||
# May be False if model not pulled
|
||||
assert isinstance(is_healthy, bool)
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_list_models(self, ollama_client):
|
||||
"""Test listing available models."""
|
||||
models = await ollama_client.list_models()
|
||||
assert isinstance(models, list)
|
||||
|
||||
@pytest.mark.asyncio
|
||||
@pytest.mark.skipif(
|
||||
True, # Skip by default as embeddings can be slow
|
||||
reason="Embedding generation is slow - enable manually if needed"
|
||||
)
|
||||
async def test_generate_embedding(self, ollama_client):
|
||||
"""Test generating a single embedding."""
|
||||
embedding = await ollama_client.embed("test text")
|
||||
if embedding:
|
||||
assert isinstance(embedding, list)
|
||||
assert len(embedding) > 0
|
||||
|
||||
|
||||
class TestJobManagerIntegration:
|
||||
"""Test Job Manager with Redis."""
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_create_and_get_job(self, job_manager, test_user):
|
||||
"""Test creating and retrieving a job."""
|
||||
# Create job
|
||||
job_id = await job_manager.create_job(
|
||||
job_type=JobType.DOCUMENT_INGESTION,
|
||||
user=test_user,
|
||||
parameters={"doc_url": "https://example.com/doc.pdf"}
|
||||
)
|
||||
assert job_id
|
||||
|
||||
# Get job
|
||||
job = await job_manager.get_job(job_id)
|
||||
assert job is not None
|
||||
assert job["job_id"] == job_id
|
||||
assert job["user"] == test_user
|
||||
assert job["status"] == JobStatus.QUEUED.value
|
||||
|
||||
# Update job
|
||||
await job_manager.update_job_status(
|
||||
job_id,
|
||||
JobStatus.PROCESSING,
|
||||
progress=50
|
||||
)
|
||||
|
||||
updated_job = await job_manager.get_job(job_id)
|
||||
assert updated_job["status"] == JobStatus.PROCESSING.value
|
||||
assert updated_job["progress"] == 50
|
||||
|
||||
# Cleanup
|
||||
deleted = await job_manager.delete_job(job_id, test_user)
|
||||
assert deleted is True
|
||||
|
||||
|
||||
class TestDependencyInjection:
|
||||
"""Test dependency injection and lifecycle management."""
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_startup_clients(self):
|
||||
"""Test client startup."""
|
||||
from src.core.dependencies import startup_clients
|
||||
|
||||
# Should not raise exceptions
|
||||
await startup_clients()
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_check_service_health(self):
|
||||
"""Test health check for all services."""
|
||||
from src.core.dependencies import check_service_health
|
||||
|
||||
health = await check_service_health()
|
||||
|
||||
assert isinstance(health, dict)
|
||||
assert "neo4j" in health
|
||||
assert "qdrant" in health
|
||||
assert "wikijs" in health
|
||||
assert "searxng" in health
|
||||
assert "ollama" in health
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_shutdown_clients(self):
|
||||
"""Test client shutdown."""
|
||||
from src.core.dependencies import shutdown_clients
|
||||
|
||||
# Should not raise exceptions
|
||||
await shutdown_clients()
|
||||
@@ -1,115 +0,0 @@
|
||||
"""Tests for multi-tenancy helpers."""
|
||||
|
||||
import pytest
|
||||
from src.core.multi_tenancy import (
|
||||
sanitize_user_id,
|
||||
get_qdrant_collection_name,
|
||||
get_wikijs_namespace,
|
||||
get_neo4j_user_label,
|
||||
validate_user_id,
|
||||
is_path_in_user_namespace,
|
||||
DEFAULT_USER
|
||||
)
|
||||
|
||||
|
||||
class TestSanitizeUserId:
|
||||
"""Test user ID sanitization."""
|
||||
|
||||
def test_lowercase_conversion(self):
|
||||
assert sanitize_user_id("JohnDoe") == "johndoe"
|
||||
|
||||
def test_email_conversion(self):
|
||||
assert sanitize_user_id("john@example.com") == "john_at_example_com"
|
||||
|
||||
def test_dot_conversion(self):
|
||||
assert sanitize_user_id("john.doe") == "john_doe"
|
||||
|
||||
def test_space_conversion(self):
|
||||
assert sanitize_user_id("John Doe") == "john_doe"
|
||||
|
||||
def test_special_chars_removal(self):
|
||||
assert sanitize_user_id("john-doe!") == "john_doe"
|
||||
|
||||
def test_consecutive_underscores(self):
|
||||
assert sanitize_user_id("john__doe") == "john_doe"
|
||||
|
||||
def test_leading_trailing_underscores(self):
|
||||
assert sanitize_user_id("_john_") == "john"
|
||||
|
||||
|
||||
class TestQdrantCollectionName:
|
||||
"""Test Qdrant collection name generation."""
|
||||
|
||||
def test_simple_user(self):
|
||||
assert get_qdrant_collection_name("jpmschweitzer") == "library_desk_jpmschweitzer"
|
||||
|
||||
def test_email_user(self):
|
||||
assert get_qdrant_collection_name("john@example.com") == "library_desk_john_at_example_com"
|
||||
|
||||
def test_default_user(self):
|
||||
assert get_qdrant_collection_name(DEFAULT_USER) == f"library_desk_{DEFAULT_USER}"
|
||||
|
||||
|
||||
class TestWikijsNamespace:
|
||||
"""Test Wiki.js namespace generation."""
|
||||
|
||||
def test_simple_user(self):
|
||||
assert get_wikijs_namespace("jpmschweitzer") == "/users/jpmschweitzer"
|
||||
|
||||
def test_email_user(self):
|
||||
assert get_wikijs_namespace("john@example.com") == "/users/john_at_example_com"
|
||||
|
||||
def test_starts_with_slash(self):
|
||||
namespace = get_wikijs_namespace("testuser")
|
||||
assert namespace.startswith("/")
|
||||
|
||||
|
||||
class TestNeo4jUserLabel:
|
||||
"""Test Neo4j user label generation."""
|
||||
|
||||
def test_simple_user(self):
|
||||
assert get_neo4j_user_label("jpmschweitzer") == "User_Jpmschweitzer_Document"
|
||||
|
||||
def test_email_user(self):
|
||||
result = get_neo4j_user_label("john@example.com")
|
||||
# Should be title case
|
||||
assert result == "User_John_At_Example_Com_Document"
|
||||
|
||||
def test_title_case(self):
|
||||
result = get_neo4j_user_label("john_doe")
|
||||
assert result == "User_John_Doe_Document"
|
||||
|
||||
|
||||
class TestValidateUserId:
|
||||
"""Test user ID validation."""
|
||||
|
||||
def test_valid_simple(self):
|
||||
assert validate_user_id("jpmschweitzer") is True
|
||||
|
||||
def test_valid_email(self):
|
||||
assert validate_user_id("john@example.com") is True
|
||||
|
||||
def test_empty_invalid(self):
|
||||
assert validate_user_id("") is False
|
||||
|
||||
def test_too_long_invalid(self):
|
||||
assert validate_user_id("a" * 101) is False
|
||||
|
||||
def test_no_alphanumeric_invalid(self):
|
||||
assert validate_user_id("___") is False
|
||||
|
||||
|
||||
class TestPathInNamespace:
|
||||
"""Test path namespace checking."""
|
||||
|
||||
def test_path_in_namespace(self):
|
||||
assert is_path_in_user_namespace("/users/jpmschweitzer/projects", "jpmschweitzer") is True
|
||||
|
||||
def test_path_not_in_namespace(self):
|
||||
assert is_path_in_user_namespace("/users/other/projects", "jpmschweitzer") is False
|
||||
|
||||
def test_public_path_not_in_namespace(self):
|
||||
assert is_path_in_user_namespace("/public/docs", "jpmschweitzer") is False
|
||||
|
||||
def test_root_path(self):
|
||||
assert is_path_in_user_namespace("/users/test", "test") is True
|
||||
@@ -1,435 +0,0 @@
|
||||
"""
|
||||
Tests for Wiki.js Change Listener
|
||||
|
||||
Tests the PostgreSQL NOTIFY/LISTEN change detection system including:
|
||||
- Database connection and listener startup
|
||||
- Notification handling (INSERT/UPDATE/DELETE)
|
||||
- Loop prevention (automated user filtering)
|
||||
- Debouncing (duplicate notification filtering)
|
||||
- Page processing
|
||||
"""
|
||||
|
||||
import pytest
|
||||
import asyncio
|
||||
from unittest.mock import AsyncMock, MagicMock, patch
|
||||
from datetime import datetime, timedelta
|
||||
|
||||
from src.services.wiki_change_listener import WikiChangeListener
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def mock_settings():
|
||||
"""Mock settings for testing."""
|
||||
settings = MagicMock()
|
||||
settings.wikijs_db_host = "postgres-shared"
|
||||
settings.wikijs_db_port = 5432
|
||||
settings.wikijs_db_user = "library_desk_listener"
|
||||
settings.wikijs_db_password = "test_password"
|
||||
settings.wikijs_db_name = "library"
|
||||
settings.wikijs_username = "librarian@schweitz.net"
|
||||
settings.wikijs_change_listener_debounce_seconds = 5
|
||||
return settings
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def listener(mock_settings):
|
||||
"""Create a WikiChangeListener instance with mocked settings."""
|
||||
with patch('src.services.wiki_change_listener.get_settings', return_value=mock_settings):
|
||||
return WikiChangeListener()
|
||||
|
||||
|
||||
class TestWikiChangeListener:
|
||||
"""Test suite for WikiChangeListener."""
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_listener_initialization(self, listener, mock_settings):
|
||||
"""Test that listener initializes with correct settings."""
|
||||
assert listener.settings == mock_settings
|
||||
assert listener.connection is None
|
||||
assert listener.running is False
|
||||
assert listener._debounce_seconds == 5
|
||||
assert len(listener._recent_notifications) == 0
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_automated_user_filtering(self, listener):
|
||||
"""Test that automated users are correctly identified."""
|
||||
# Automated users should be filtered
|
||||
assert listener._is_automated_user("librarian@schweitz.net") is True
|
||||
assert listener._is_automated_user("library-desk@system") is True
|
||||
assert listener._is_automated_user("automation@system") is True
|
||||
assert listener._is_automated_user("bot@system") is True
|
||||
|
||||
# Case insensitive
|
||||
assert listener._is_automated_user("LIBRARIAN@SCHWEITZ.NET") is True
|
||||
|
||||
# Regular users should not be filtered
|
||||
assert listener._is_automated_user("user@example.com") is False
|
||||
assert listener._is_automated_user("john@example.com") is False
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_debouncing_prevents_duplicates(self, listener):
|
||||
"""Test that debouncing prevents duplicate processing."""
|
||||
page_id = 123
|
||||
|
||||
# First notification - should not be filtered
|
||||
assert listener._is_recently_processed(page_id) is False
|
||||
|
||||
# Mark as processed
|
||||
listener._mark_as_processed(page_id)
|
||||
|
||||
# Immediate second notification - should be filtered (within debounce window)
|
||||
assert listener._is_recently_processed(page_id) is True
|
||||
|
||||
# Different page - should not be filtered
|
||||
assert listener._is_recently_processed(456) is False
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_debouncing_expires_after_window(self, listener):
|
||||
"""Test that debouncing expires after the configured time window."""
|
||||
page_id = 123
|
||||
|
||||
# Mark as processed with old timestamp (outside debounce window)
|
||||
listener._recent_notifications[page_id] = datetime.now() - timedelta(seconds=10)
|
||||
|
||||
# Should not be filtered anymore (10 seconds > 5 second debounce)
|
||||
assert listener._is_recently_processed(page_id) is False
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_mark_as_processed_cleanup(self, listener):
|
||||
"""Test that old entries are cleaned up to prevent memory growth."""
|
||||
# Add 101 entries to trigger cleanup (threshold is 100)
|
||||
for i in range(101):
|
||||
listener._mark_as_processed(i)
|
||||
|
||||
# Should only keep last 100 entries
|
||||
assert len(listener._recent_notifications) == 100
|
||||
|
||||
# Oldest entry (0) should be removed
|
||||
assert 0 not in listener._recent_notifications
|
||||
|
||||
# Newest entries should be kept
|
||||
assert 100 in listener._recent_notifications
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_notification_payload_parsing(self, listener):
|
||||
"""Test that notification payloads are correctly parsed."""
|
||||
mock_connection = AsyncMock()
|
||||
|
||||
with patch.object(listener, '_process_page_change', new_callable=AsyncMock) as mock_process:
|
||||
# Test UPDATE notification
|
||||
await listener._handle_notification(
|
||||
mock_connection, 1234, 'wiki_page_changes',
|
||||
'UPDATE:123:user@example.com'
|
||||
)
|
||||
|
||||
mock_process.assert_called_once()
|
||||
call_args = mock_process.call_args[1]
|
||||
assert call_args['page_id'] == 123
|
||||
assert call_args['event'] == 'page.update'
|
||||
assert call_args['user'] == 'user'
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_notification_operations_mapping(self, listener):
|
||||
"""Test that database operations map to correct webhook events."""
|
||||
mock_connection = AsyncMock()
|
||||
|
||||
with patch.object(listener, '_process_page_change', new_callable=AsyncMock) as mock_process:
|
||||
# INSERT -> page.create (use page 100)
|
||||
await listener._handle_notification(
|
||||
mock_connection, 1234, 'wiki_page_changes',
|
||||
'INSERT:100:user@example.com'
|
||||
)
|
||||
mock_process.assert_called_once()
|
||||
assert mock_process.call_args[1]['event'] == 'page.create'
|
||||
|
||||
mock_process.reset_mock()
|
||||
|
||||
# UPDATE -> page.update (use different page 200 to avoid debouncing)
|
||||
await listener._handle_notification(
|
||||
mock_connection, 1234, 'wiki_page_changes',
|
||||
'UPDATE:200:user@example.com'
|
||||
)
|
||||
mock_process.assert_called_once()
|
||||
assert mock_process.call_args[1]['event'] == 'page.update'
|
||||
|
||||
mock_process.reset_mock()
|
||||
|
||||
# DELETE -> page.delete (use different page 300 to avoid debouncing)
|
||||
await listener._handle_notification(
|
||||
mock_connection, 1234, 'wiki_page_changes',
|
||||
'DELETE:300:user@example.com'
|
||||
)
|
||||
mock_process.assert_called_once()
|
||||
assert mock_process.call_args[1]['event'] == 'page.delete'
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_automated_user_notification_filtered(self, listener):
|
||||
"""Test that notifications from automated users are filtered out."""
|
||||
mock_connection = AsyncMock()
|
||||
|
||||
with patch.object(listener, '_process_page_change', new_callable=AsyncMock) as mock_process:
|
||||
# Notification from automated user should be skipped
|
||||
await listener._handle_notification(
|
||||
mock_connection, 1234, 'wiki_page_changes',
|
||||
'UPDATE:123:librarian@schweitz.net'
|
||||
)
|
||||
|
||||
# Process should NOT be called
|
||||
mock_process.assert_not_called()
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_duplicate_notification_filtered(self, listener):
|
||||
"""Test that duplicate notifications within debounce window are filtered."""
|
||||
mock_connection = AsyncMock()
|
||||
page_id = 123
|
||||
|
||||
# Mark page as recently processed
|
||||
listener._mark_as_processed(page_id)
|
||||
|
||||
with patch.object(listener, '_process_page_change', new_callable=AsyncMock) as mock_process:
|
||||
# Duplicate notification should be skipped
|
||||
await listener._handle_notification(
|
||||
mock_connection, 1234, 'wiki_page_changes',
|
||||
f'UPDATE:{page_id}:user@example.com'
|
||||
)
|
||||
|
||||
# Process should NOT be called
|
||||
mock_process.assert_not_called()
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_invalid_notification_payload_handled(self, listener):
|
||||
"""Test that invalid notification payloads are handled gracefully."""
|
||||
mock_connection = AsyncMock()
|
||||
|
||||
with patch.object(listener, '_process_page_change', new_callable=AsyncMock) as mock_process:
|
||||
# Invalid payload (too few parts)
|
||||
await listener._handle_notification(
|
||||
mock_connection, 1234, 'wiki_page_changes',
|
||||
'INVALID:123' # Missing user email
|
||||
)
|
||||
|
||||
# Should not crash and should not process
|
||||
mock_process.assert_not_called()
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_email_to_username_extraction(self, listener):
|
||||
"""Test that user email is correctly extracted to username."""
|
||||
mock_connection = AsyncMock()
|
||||
|
||||
with patch.object(listener, '_process_page_change', new_callable=AsyncMock) as mock_process:
|
||||
await listener._handle_notification(
|
||||
mock_connection, 1234, 'wiki_page_changes',
|
||||
'UPDATE:123:john.doe@example.com'
|
||||
)
|
||||
|
||||
# Should extract 'john.doe' from email
|
||||
call_args = mock_process.call_args[1]
|
||||
assert call_args['user'] == 'john.doe'
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_email_without_at_sign_fallback(self, listener):
|
||||
"""Test fallback when email doesn't contain @ sign."""
|
||||
mock_connection = AsyncMock()
|
||||
|
||||
with patch.object(listener, '_process_page_change', new_callable=AsyncMock) as mock_process:
|
||||
await listener._handle_notification(
|
||||
mock_connection, 1234, 'wiki_page_changes',
|
||||
'UPDATE:123:invaliduser'
|
||||
)
|
||||
|
||||
# Should use default user
|
||||
call_args = mock_process.call_args[1]
|
||||
assert call_args['user'] == 'jpmschweitzer'
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_process_page_delete_calls_cleanup(self, listener):
|
||||
"""Test that DELETE events call cleanup_deleted_page."""
|
||||
with patch('src.routers.webhooks.cleanup_deleted_page', new_callable=AsyncMock) as mock_cleanup, \
|
||||
patch('src.services.wiki_change_listener.get_ingestion_service') as mock_service:
|
||||
|
||||
await listener._process_page_change(
|
||||
page_id=123,
|
||||
event='page.delete',
|
||||
user='testuser'
|
||||
)
|
||||
|
||||
mock_cleanup.assert_called_once()
|
||||
call_args = mock_cleanup.call_args[1]
|
||||
assert call_args['page_id'] == 123
|
||||
assert call_args['user'] == 'testuser'
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_process_page_create_calls_process_wiki_page_change(self, listener):
|
||||
"""Test that CREATE events call process_wiki_page_change."""
|
||||
mock_page = MagicMock()
|
||||
mock_page.title = "Test Page"
|
||||
|
||||
with patch('src.routers.webhooks.process_wiki_page_change', new_callable=AsyncMock) as mock_process, \
|
||||
patch('src.core.dependencies.get_wiki_service') as mock_wiki_service_factory, \
|
||||
patch('src.services.wiki_change_listener.get_ingestion_service'):
|
||||
|
||||
mock_wiki_service = AsyncMock()
|
||||
mock_wiki_service.get_page.return_value = mock_page
|
||||
mock_wiki_service_factory.return_value = mock_wiki_service
|
||||
|
||||
await listener._process_page_change(
|
||||
page_id=123,
|
||||
event='page.create',
|
||||
user='testuser'
|
||||
)
|
||||
|
||||
mock_process.assert_called_once()
|
||||
call_args = mock_process.call_args[1]
|
||||
assert call_args['page_id'] == 123
|
||||
assert call_args['page_title'] == "Test Page"
|
||||
assert call_args['event'] == 'page.create'
|
||||
assert call_args['user'] == 'testuser'
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_process_page_update_calls_process_wiki_page_change(self, listener):
|
||||
"""Test that UPDATE events call process_wiki_page_change."""
|
||||
mock_page = MagicMock()
|
||||
mock_page.title = "Updated Page"
|
||||
|
||||
with patch('src.routers.webhooks.process_wiki_page_change', new_callable=AsyncMock) as mock_process, \
|
||||
patch('src.core.dependencies.get_wiki_service') as mock_wiki_service_factory, \
|
||||
patch('src.services.wiki_change_listener.get_ingestion_service'):
|
||||
|
||||
mock_wiki_service = AsyncMock()
|
||||
mock_wiki_service.get_page.return_value = mock_page
|
||||
mock_wiki_service_factory.return_value = mock_wiki_service
|
||||
|
||||
await listener._process_page_change(
|
||||
page_id=456,
|
||||
event='page.update',
|
||||
user='testuser'
|
||||
)
|
||||
|
||||
mock_process.assert_called_once()
|
||||
call_args = mock_process.call_args[1]
|
||||
assert call_args['page_id'] == 456
|
||||
assert call_args['page_title'] == "Updated Page"
|
||||
assert call_args['event'] == 'page.update'
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_connection_lifecycle(self, listener, mock_settings):
|
||||
"""Test listener connection start and stop lifecycle."""
|
||||
mock_connection = AsyncMock()
|
||||
|
||||
with patch('src.services.wiki_change_listener.asyncpg.connect', return_value=mock_connection) as mock_connect:
|
||||
# Start listener
|
||||
await listener.start()
|
||||
|
||||
# Verify connection was established with correct parameters
|
||||
mock_connect.assert_called_once_with(
|
||||
host=mock_settings.wikijs_db_host,
|
||||
port=mock_settings.wikijs_db_port,
|
||||
user=mock_settings.wikijs_db_user,
|
||||
password=mock_settings.wikijs_db_password,
|
||||
database=mock_settings.wikijs_db_name
|
||||
)
|
||||
|
||||
# Verify listener was added
|
||||
mock_connection.add_listener.assert_called_once_with(
|
||||
'wiki_page_changes',
|
||||
listener._handle_notification
|
||||
)
|
||||
|
||||
assert listener.running is True
|
||||
assert listener.connection == mock_connection
|
||||
|
||||
# Stop listener
|
||||
await listener.stop()
|
||||
|
||||
# Verify listener was removed and connection closed
|
||||
mock_connection.remove_listener.assert_called_once()
|
||||
mock_connection.close.assert_called_once()
|
||||
assert listener.running is False
|
||||
|
||||
|
||||
class TestLoopPreventionScenarios:
|
||||
"""Integration tests for loop prevention scenarios."""
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_full_loop_prevention_flow(self, listener):
|
||||
"""
|
||||
Test complete loop prevention flow:
|
||||
1. User edits page -> Processes
|
||||
2. Entity linking updates page (as automated user) -> Filtered
|
||||
3. Rapid duplicate edits -> Debounced
|
||||
"""
|
||||
mock_connection = AsyncMock()
|
||||
|
||||
with patch.object(listener, '_process_page_change', new_callable=AsyncMock) as mock_process:
|
||||
# 1. User edit - should process
|
||||
await listener._handle_notification(
|
||||
mock_connection, 1234, 'wiki_page_changes',
|
||||
'UPDATE:100:user@example.com'
|
||||
)
|
||||
assert mock_process.call_count == 1
|
||||
|
||||
mock_process.reset_mock()
|
||||
|
||||
# 2. Automated edit (entity linking) - should be filtered
|
||||
await listener._handle_notification(
|
||||
mock_connection, 1234, 'wiki_page_changes',
|
||||
'UPDATE:100:librarian@schweitz.net'
|
||||
)
|
||||
assert mock_process.call_count == 0
|
||||
|
||||
# 3. Rapid duplicate from same user - should be debounced
|
||||
await listener._handle_notification(
|
||||
mock_connection, 1234, 'wiki_page_changes',
|
||||
'UPDATE:100:user@example.com'
|
||||
)
|
||||
assert mock_process.call_count == 0 # Debounced
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_different_pages_not_debounced(self, listener):
|
||||
"""Test that edits to different pages are not debounced."""
|
||||
mock_connection = AsyncMock()
|
||||
|
||||
with patch.object(listener, '_process_page_change', new_callable=AsyncMock) as mock_process:
|
||||
# Edit page 100
|
||||
await listener._handle_notification(
|
||||
mock_connection, 1234, 'wiki_page_changes',
|
||||
'UPDATE:100:user@example.com'
|
||||
)
|
||||
assert mock_process.call_count == 1
|
||||
|
||||
# Edit page 200 immediately - should NOT be debounced
|
||||
await listener._handle_notification(
|
||||
mock_connection, 1234, 'wiki_page_changes',
|
||||
'UPDATE:200:user@example.com'
|
||||
)
|
||||
assert mock_process.call_count == 2
|
||||
|
||||
|
||||
@pytest.mark.integration
|
||||
class TestWikiChangeListenerIntegration:
|
||||
"""
|
||||
Integration tests (require actual PostgreSQL connection).
|
||||
These tests are marked with @pytest.mark.integration and skipped by default.
|
||||
Run with: pytest -m integration
|
||||
"""
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_real_database_connection(self):
|
||||
"""Test connection to real PostgreSQL database (requires setup)."""
|
||||
pytest.skip("Requires actual PostgreSQL setup with triggers")
|
||||
|
||||
listener = WikiChangeListener()
|
||||
try:
|
||||
await listener.start()
|
||||
assert listener.running is True
|
||||
assert listener.connection is not None
|
||||
finally:
|
||||
await listener.stop()
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_real_notification_handling(self):
|
||||
"""Test handling real NOTIFY events from PostgreSQL."""
|
||||
pytest.skip("Requires actual PostgreSQL setup with triggers")
|
||||
|
||||
# This would test actual pg_notify() calls from triggers
|
||||
# and verify the listener receives and processes them
|
||||
+5
-88
@@ -4,41 +4,15 @@ version: '3.8'
|
||||
# Application Layer
|
||||
# Port: 8089 (HTTP)
|
||||
# GPU: No
|
||||
# Storage: SSD (venv, logs)
|
||||
# Source: https://git.schweitz.net/jpmschweitzer/library-desk
|
||||
|
||||
services:
|
||||
library-desk:
|
||||
image: python:3.12-slim
|
||||
image: git.schweitz.net/jpmschweitzer/library-desk:latest
|
||||
container_name: library-desk
|
||||
restart: unless-stopped
|
||||
ports:
|
||||
- "8089:8089" # FastAPI HTTP
|
||||
volumes:
|
||||
# Mount source code for live editing
|
||||
- /home/jpmschweitzer/Projects/portainer-core/services/library-desk:/app
|
||||
# Persist container's venv for fast restarts
|
||||
- /home/jpmschweitzer/docker-data/library-desk/venv:/venv
|
||||
|
||||
working_dir: /app
|
||||
|
||||
command: >
|
||||
sh -c "
|
||||
echo 'Installing system dependencies...' &&
|
||||
apt-get update -qq &&
|
||||
apt-get install -y --no-install-recommends curl >/dev/null 2>&1 &&
|
||||
rm -rf /var/lib/apt/lists/* &&
|
||||
echo 'Setting up Python environment...' &&
|
||||
if [ ! -f /venv/bin/python ]; then
|
||||
echo 'Initializing venv...' &&
|
||||
python3 -m venv --clear /venv;
|
||||
fi &&
|
||||
echo 'Upgrading pip...' &&
|
||||
/venv/bin/python -m pip install --upgrade pip --quiet &&
|
||||
echo 'Installing dependencies from requirements.txt...' &&
|
||||
/venv/bin/python -m pip install -r /app/requirements.txt --quiet &&
|
||||
echo 'Starting Library Desk API...' &&
|
||||
/venv/bin/python -m uvicorn src.main:app --host 0.0.0.0 --port 8089 --workers 2
|
||||
"
|
||||
environment:
|
||||
# API Configuration
|
||||
- LIBRARY_API_KEY=${LIBRARY_API_KEY}
|
||||
@@ -73,6 +47,8 @@ services:
|
||||
# Python Configuration
|
||||
- PYTHONUNBUFFERED=1
|
||||
- TZ=${TZ:-Europe/Amsterdam}
|
||||
labels:
|
||||
- "com.centurylinklabs.watchtower.enable=true"
|
||||
networks:
|
||||
- docker-dataplane
|
||||
deploy:
|
||||
@@ -82,7 +58,7 @@ services:
|
||||
limits:
|
||||
memory: 768M
|
||||
healthcheck:
|
||||
test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://localhost:8089/health')"]
|
||||
test: ["CMD", "curl", "-f", "http://localhost:8089/health"]
|
||||
interval: 30s
|
||||
timeout: 10s
|
||||
retries: 3
|
||||
@@ -92,62 +68,3 @@ networks:
|
||||
docker-dataplane:
|
||||
external: true
|
||||
name: docker-dataplane
|
||||
|
||||
# ⚠️ SECURITY WARNING:
|
||||
# Set these in environment variables before deploying:
|
||||
# - LIBRARY_API_KEY (generate with: openssl rand -hex 32)
|
||||
# - NEO4J_PASSWORD (from library-neo4j deployment)
|
||||
# - LIBRARY_DB_PASSWORD (from library-wiki deployment)
|
||||
# - WIKIJS_API_KEY (from Wiki.js admin panel → API Access)
|
||||
#
|
||||
# Prerequisites:
|
||||
# 1. Create /home/jpmschweitzer/Projects/portainer-core/services/library/services/front-desk/
|
||||
# 2. Create requirements.txt (see DEPLOYMENT.md Phase 4.3)
|
||||
# 3. Create main.py with FastAPI app (see DEPLOYMENT.md Phase 4.3)
|
||||
#
|
||||
# API Endpoints (once implemented):
|
||||
# Query:
|
||||
# POST /query/hybrid # HybridRAG (graph + vector + web)
|
||||
# POST /query/semantic # Vector search only
|
||||
# POST /query/graph # Graph traversal only
|
||||
# GET /query/related/{id} # Find related content
|
||||
#
|
||||
# Content Management:
|
||||
# POST /ingest/document # Index new document
|
||||
# POST /ingest/wiki-page # Sync Wiki.js page
|
||||
# POST /wiki/dossier # Create dossier (proxies to Wiki.js)
|
||||
# PUT /wiki/dossier/{id} # Update dossier
|
||||
# DELETE /wiki/dossier/{id} # Delete dossier
|
||||
#
|
||||
# Graph Operations:
|
||||
# GET /graph/entities # List entities
|
||||
# GET /graph/mindmap/{id} # Generate mind map for dossier
|
||||
# POST /graph/query # Execute Cypher query
|
||||
#
|
||||
# Deduplication:
|
||||
# POST /deduplicate/find # Find potential duplicates
|
||||
# POST /deduplicate/merge # Merge duplicate entities
|
||||
#
|
||||
# System:
|
||||
# GET /stats # System statistics
|
||||
# GET /health # Health check
|
||||
#
|
||||
# API Documentation:
|
||||
# - Interactive docs: http://192.168.86.149:8089/docs
|
||||
# - OpenAPI spec: http://192.168.86.149:8089/openapi.json
|
||||
#
|
||||
# Features:
|
||||
# - HybridRAG queries (Neo4j + Qdrant + SearXNG)
|
||||
# - Document ingestion and indexing
|
||||
# - Entity extraction and relationship mapping
|
||||
# - Mind map generation
|
||||
# - Deduplication detection
|
||||
# - Wiki.js API proxy
|
||||
#
|
||||
# Dependencies:
|
||||
# - Neo4j (knowledge graph)
|
||||
# - Qdrant (vector search)
|
||||
# - Wiki.js (wiki operations)
|
||||
# - SearXNG (web search)
|
||||
# - Ollama (embeddings)
|
||||
# - Redis (caching)
|
||||
|
||||
Reference in New Issue
Block a user