Compare commits

..
7 Commits
Author SHA1 Message Date
jpmschweitzerandClaude Opus 4.5 2552bfd1f9 fix: make Wiki.js API token optional for open GraphQL endpoints
Build and Push / build (release) Successful in 33s
The Wiki.js GraphQL API is accessible without authentication.
Make WIKI_GRAPHQL_API env var optional with empty default to fix
container startup failures.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2025-12-24 16:59:41 +01:00
jpmschweitzerandClaude Opus 4.5 2c35aec179 release: v1.4.0 - maintenance system and Wiki.js API token auth
Build and Push / build (release) Successful in 29s
Features:
- Maintenance router with index reconciliation
- Bidirectional orphan detection (vectors ↔ graph)
- Wiki.js API token authentication

See CHANGELOG.md for full details.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2025-12-24 16:36:54 +01:00
jpmschweitzerandClaude Opus 4.5 97bd52006f chore: add local development environment files
Development setup:
- .env.example: Template with all required environment variables
- CLAUDE.md: Claude Code agent instructions
- wakeup.sh: Local server startup script with auto-reload

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2025-12-24 16:35:05 +01:00
jpmschweitzerandClaude Opus 4.5 f139f518ff docs: add scheduler integration and memory system plan
Documentation updates:
- AGENTS.md: Added wakeup.sh usage and local testing instructions
- LIBRARIAN_INTEGRATION.md: Complete maintenance scheduler docs
  - Scheduled task configuration for reconcile-index
  - Endpoint specifications and response formats
- docs/MEMORY_SYSTEM_PLAN.md: Three-tier memory architecture
  - Volatile (Redis TTL) for ephemeral context
  - Documents (TBD) for git mirrors, PDFs, images
  - Knowledge (Wiki + Neo4j) for permanent research

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2025-12-24 16:33:06 +01:00
jpmschweitzerandClaude Opus 4.5 f756ca9490 feat: add maintenance system with index reconciliation
Complete maintenance subsystem for index health and cleanup:

Endpoints:
- GET /maintenance/health - lightweight (or detailed) health check
- POST /maintenance/cleanup/all - full orphan cleanup
- POST /maintenance/cleanup/vectors - purge orphan vector chunks
- POST /maintenance/cleanup/graph - purge orphan graph nodes
- POST /maintenance/reconcile-index - cleanup + reindex missing pages

Bidirectional orphan detection:
- find_documents_without_vectors() in GraphService
- find_chunks_without_graph_nodes() in VectorService

Redis integration:
- Tracks last_cleanup timestamp for scheduler visibility

Config additions:
- Document store, volatile cache, and maintenance settings
- VectorServiceDep and GraphServiceDep type aliases

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2025-12-24 16:22:39 +01:00
jpmschweitzerandClaude Opus 4.5 0e6c3619eb feat: add scroll and batch delete methods to Qdrant client
Add bulk operations needed for maintenance:
- scroll_all_points(): paginated iteration over all points
- delete_by_ids(): batch delete points by ID list

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2025-12-24 16:21:57 +01:00
jpmschweitzerandClaude Opus 4.5 9446d6bf9a refactor: switch Wiki.js client to API token authentication
Replace username/password login flow with simpler API token auth:
- Use WIKI_GRAPHQL_API environment variable for JWT token
- Remove login() method and session management
- Add list_pages() method for fetching all pages
- Keep legacy auth fields in config for backwards compatibility

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2025-12-24 16:21:39 +01:00
17 changed files with 2533 additions and 88 deletions
+23
View File
@@ -0,0 +1,23 @@
# Service URLs for local dev (pointing to your server)
TEST_HOST=192.168.86.149
WIKIJS_URL=http://192.168.86.149:8088
NEO4J_URI=bolt://192.168.86.149:7687
QDRANT_HOST=192.168.86.149
QDRANT_PORT=6333
OLLAMA_URL=http://192.168.86.149:11434
SEARXNG_URL=http://192.168.86.149:8080
REDIS_HOST=192.168.86.149
OLLAMA_MODEL=mistral-nemo-large:latest
OLLAMA_EMBEDDING_MODEL=nomic-embed-text
# Wiki.js auth
WIKIJS_USERNAME=librarian@schweitz.net
WIKIJS_PASSWORD=key_here
# Wiki.js GraphQL API token (generate from Admin → API Access)
WIKI_GRAPHQL_API=your_jwt_token_here
LIBRARY_API_KEY=key_here
NEO4J_PASSWORD=key_here
WIKIJS_DB_PASSWORD=key_here
SCHEDULER_API_KEY=key_here
+15
View File
@@ -51,6 +51,21 @@ When changes are ready for deployment:
---
### 🧪 Local Development Setup
* **Always test locally first** before committing and deploying. The build-deploy loop is slow.
* **Start the local server** with `./wakeup.sh` - logs are written to `logs/server.log` for easy tailing
* **Auto-reload**: The wakeup script runs uvicorn in reload mode - code changes are picked up automatically without restart (except for requirements.txt changes)
* **Test REST endpoints** against `http://localhost:8778` using curl or similar tools
* **Only deploy** when a phase or feature is complete and tested locally
* **Environment**: Copy `.env.example` to `.env` and configure for your local setup (Ollama, Redis, Neo4j, Qdrant, Wiki.js hosts)
* **Running tests**: Always use the venv explicitly to avoid environment mismatches:
```bash
.venv/bin/python -m pytest tests/ # All tests
.venv/bin/python -m pytest tests/ -v # Verbose output
```
---
## 2. FastAPI Architecture & Best Practices
*Reference: [FastAPI Best Practices](https://github.com/zhanymkanov/fastapi-best-practices)*
+41
View File
@@ -5,6 +5,47 @@ All notable changes to Library Desk will be documented in this file.
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
## [1.4.1] - 2025-12-24
### Fixed
- Wiki.js API token now optional - GraphQL API works without authentication
- Container startup failure when `WIKI_GRAPHQL_API` env var not set
## [1.4.0] - 2025-12-24
### Added
- **Maintenance Router** - New `/maintenance` endpoints for system health and cleanup
- `GET /maintenance/health` - Lightweight health check (detailed mode available)
- `POST /maintenance/cleanup/all` - Full orphan cleanup (vectors + graph)
- `POST /maintenance/cleanup/vectors` - Purge orphan vector chunks
- `POST /maintenance/cleanup/graph` - Purge orphan graph nodes
- `POST /maintenance/reconcile-index` - Combined cleanup + reindex missing pages
- **Bidirectional Orphan Detection** - Cross-validate vectors and graph nodes
- `find_documents_without_vectors()` - Graph nodes missing vector chunks
- `find_chunks_without_graph_nodes()` - Vector chunks missing graph nodes
- **Qdrant Client Methods** - Bulk operations for maintenance
- `scroll_all_points()` - Iterate all points with pagination
- `delete_by_ids()` - Batch delete by point IDs
- **Graph Service Cleanup** - Node deletion methods
- `delete_document_node()` - Remove document and relationships
- `delete_collection_node()` - Remove collection and contained documents
- `get_all_document_references()` - Get all document references for validation
- **Redis Timestamp Tracking** - `last_cleanup` timestamp for scheduler integration
- **Memory System Plan** - Documented three-tier architecture (volatile/documents/knowledge)
### Changed
- **Wiki.js Authentication** - Switched from username/password to API token
- New `WIKI_GRAPHQL_API` environment variable for JWT token
- Deprecated `WIKIJS_USERNAME` and `WIKIJS_PASSWORD` (kept for backwards compatibility)
- **Service Dependencies** - Added `VectorServiceDep` and `GraphServiceDep` type aliases
### Fixed
- Wiki.js client now properly handles API token auth without login flow
## [1.3.3] - 2025-12-23
### Added
+17
View File
@@ -0,0 +1,17 @@
# Claude Code Instructions
**MANDATORY: Read AGENTS.md instead of this file.**
This project uses a unified configuration file for all LLM coding agents.
## Instructions
1. **Read and follow AGENTS.md** - All project guidelines are located there
2. **Do not modify this file** - Only update AGENTS.md
3. **Do not create or modify other agent-specific files** - Use AGENTS.md as the single source of truth
This approach ensures consistent behavior across all LLM coding agents without managing separate configuration files.
---
If you need to update project guidelines, edit AGENTS.md, not this file.
+148
View File
@@ -455,6 +455,154 @@ LIBRARY_BATCH_SIZE=50
LIBRARY_SYNC_ENABLED=true
```
## Maintenance Tasks
### Index Reconciliation (Daily)
The `reconcile-index` endpoint performs full index maintenance:
1. **Cleanup Phase**: Remove orphaned data
- Vector chunks without wiki source
- Graph nodes without vectors (bidirectional)
- Vectors without graph nodes (bidirectional)
- Orphan entities (no MENTIONS relationships)
- Broken relationships
2. **Reindex Phase**: Index missing pages
- Wiki pages without vector embeddings
- Wiki pages without graph Document nodes
**Scheduler Task: `library_reconcile_index`**
```yaml
Task Name: library_reconcile_index
Description: Daily index reconciliation - cleanup orphans + reindex missing pages
Schedule: Daily at 04:00 (after library_sync at 03:30)
Priority: 10 (system maintenance)
Service: library
Executor: POST /maintenance/reconcile-index
Configuration:
- LIBRARY_DESK_URL: http://library-desk:8089
- LIBRARY_API_KEY: ${LIBRARY_API_KEY}
Parameters:
- user: jpmschweitzer
- dry_run: false
Outputs:
- Vector orphans purged
- Entity orphans purged
- Missing pages reindexed
```
### Maintenance Endpoints
| Endpoint | Method | Purpose |
|----------|--------|---------|
| `/maintenance/reconcile-index` | POST | **Recommended**: Full cleanup + reindex missing |
| `/maintenance/cleanup/all` | POST | Cleanup only (orphan removal) |
| `/maintenance/cleanup/vectors` | POST | Clean orphan vector chunks only |
| `/maintenance/cleanup/graph` | POST | Clean orphan entities & stale docs only |
| `/maintenance/health` | GET | Lightweight health check (for uptime monitoring) |
| `/maintenance/health?detailed=true` | GET | Full analysis with orphan counts |
| `/maintenance/reindex/{page_id}` | POST | Force re-index a specific page |
### Health Check Modes
**Lightweight (default)** - Use for frequent uptime checks (every 30s):
```bash
curl "http://library-desk:8089/maintenance/health?user=jpmschweitzer" \
-H "Authorization: Bearer ${LIBRARY_API_KEY}"
```
Returns only last cleanup timestamp and basic status (no database queries).
**Detailed** - Use for dashboards or before reconciliation:
```bash
curl "http://library-desk:8089/maintenance/health?user=jpmschweitzer&detailed=true" \
-H "Authorization: Bearer ${LIBRARY_API_KEY}"
```
Returns full orphan analysis (runs database queries).
### Example Reconcile Request
```bash
curl -X POST "http://library-desk:8089/maintenance/reconcile-index?user=jpmschweitzer" \
-H "Authorization: Bearer ${LIBRARY_API_KEY}"
```
### Example Response
```json
{
"success": true,
"cleanup": {
"success": true,
"vector_cleanup": {
"wiki_chunks": {"orphans_found": 5, "orphans_purged": 5},
"document_chunks": {"orphans_found": 0, "orphans_purged": 0},
"chunks_without_graph": {"orphans_found": 2, "orphans_purged": 2},
"total_chunks_scanned": 1250,
"total_orphans_purged": 7
},
"graph_cleanup": {
"orphan_entities": {"orphans_found": 3, "orphans_purged": 3},
"stale_wiki_documents": {"orphans_found": 1, "orphans_purged": 1},
"stale_store_documents": {"orphans_found": 0, "orphans_purged": 0},
"docs_without_vectors": {"orphans_found": 0, "orphans_purged": 0},
"broken_relationships_cleaned": 0
},
"total_duration_ms": 1523.5
},
"reindex_missing": {
"pages_without_vectors": 2,
"pages_without_graph": 1,
"pages_reindexed": 2,
"pages_failed": 0,
"failed_page_ids": [],
"duration_ms": 3421.2
},
"total_duration_ms": 4944.7
}
```
### Scheduler Integration Code
```python
# scheduler/src/tasks/library_maintenance.py
async def library_reconcile_index_task(user: str = "jpmschweitzer"):
"""Run daily Library Desk index reconciliation."""
async with httpx.AsyncClient() as client:
# Run reconcile-index (cleanup + reindex missing)
result = await client.post(
f"{LIBRARY_DESK_URL}/maintenance/reconcile-index",
params={"user": user, "dry_run": False},
headers={"Authorization": f"Bearer {LIBRARY_API_KEY}"},
timeout=600.0 # 10 minutes for large indexes
)
data = result.json()
# Log summary
cleanup = data["cleanup"]
reindex = data["reindex_missing"]
logger.info(
f"Reconcile complete: "
f"{cleanup['vector_cleanup']['total_orphans_purged']} vector orphans, "
f"{cleanup['graph_cleanup']['orphan_entities']['orphans_purged']} entity orphans, "
f"{reindex['pages_reindexed']} pages reindexed"
)
if reindex["pages_failed"] > 0:
logger.warning(f"Failed to reindex pages: {reindex['failed_page_ids']}")
return data
```
---
## Next Steps
1. Implement ingestion endpoints in Library Desk
+283
View File
@@ -0,0 +1,283 @@
# Memory Management System - Implementation Plan
## Overview
A three-tier memory architecture for Library Desk with intelligent orchestration:
| Tier | Storage | Purpose | TTL |
|------|---------|---------|-----|
| **Volatile** | Redis | Weather, news, financial, ephemeral context | 5min - 2hr |
| **Documents** | TBD (research) | Git mirrors, PDFs, video, images | Permanent |
| **Knowledge** | Wiki + Neo4j | Personal dossiers, research, summaries | Permanent |
**Implementation Priority**: Cleanup → Volatile → Documents
---
## Phase 1: Cleanup System Completion
### Current State
- **COMPLETE** - All Phase 1 tasks implemented
- Redis timestamp tracking for last cleanup
- Bidirectional orphan detection between vectors and graph
- Scheduler integration endpoints ready
### Tasks
#### 1.1 Add Scheduler Integration Points ✅
**Files**: `src/routers/maintenance.py`
- [x] Add `last_cleanup` timestamp tracking in Redis
- [x] Return cleanup stats in format scheduler can log
- [x] Added `RedisDep` to cleanup endpoints
#### 1.2 Bidirectional Orphan Detection ✅
**Files**: `src/services/graph_service.py`, `src/services/vector_service.py`
- [x] `find_documents_without_vectors()` - graph nodes with no vectors
- [x] `find_chunks_without_graph_nodes()` - vectors with no graph node
- [x] Updated maintenance endpoints to use bidirectional checks
- [x] Added `chunks_without_graph` and `docs_without_vectors` to response models
#### 1.3 Scheduler Configuration ✅
**Scheduler-side task definition:**
```json
{
"task_name": "library_reconcile_index",
"schedule": "0 4 * * *",
"endpoint": "POST /maintenance/reconcile-index?user=jpmschweitzer",
"description": "Daily index reconciliation - cleanup + reindex missing"
}
```
- [x] Documented in `LIBRARIAN_INTEGRATION.md`
- [x] Added `reconcile-index` endpoint (cleanup + reindex missing)
- [x] Lightweight health check mode for uptime monitoring
- [x] Detailed health check mode for dashboards
---
## Phase 2: Volatile Memory System
### Architecture
```
┌─────────────────┐ ┌──────────────┐ ┌─────────────────┐
│ Library-Desk │◄───│ Scheduler │───►│ External APIs │
│ │ │ │ │ (weather, news) │
│ VolatileCache │ │ Refresh │ └─────────────────┘
│ Service │ │ Jobs │
└────────┬────────┘ └──────────────┘
┌─────────────────┐
│ Redis │
│ (DB 4, TTL) │
└─────────────────┘
```
### Data Model
```python
class VolatileRecord(BaseModel):
key: str # e.g., "weather:rotterdam"
namespace: str # e.g., "weather", "news", "financial"
data: dict # Actual content
source: Optional[str] # Origin API/service
created_at: datetime
updated_at: datetime
ttl: int # Seconds until expiration
refresh_schedule: Optional[str] # Cron expression, if repeating
user: str # Multi-tenant isolation
```
**Key pattern**: `{user}:volatile:{namespace}:{key_hash}`
### Implementation Order: Integration-First
1. **Start with Consolidation Hook** - Understand data flow through existing system
2. **Build Service Layer** - VolatileCacheService with Redis operations
3. **Add API Endpoints** - REST interface for volatile data
4. **Biographer Integration** - Query user preferences for relevance
### Tasks
#### 2.1 Integrate with Consolidation (FIRST)
**New file**: `src/services/volatile_service.py`
```python
class VolatileCacheService:
async def get(user, namespace, key) -> Optional[VolatileRecord]
async def set(user, namespace, key, data, ttl, refresh_schedule=None)
async def delete(user, namespace, key)
async def list_namespace(user, namespace) -> List[str]
async def get_scheduled(user) -> List[VolatileRecord] # For scheduler
```
#### 2.2 Create Volatile API Router
**New file**: `src/routers/volatile.py`
| Endpoint | Method | Purpose |
|----------|--------|---------|
| `/volatile/{namespace}/{key}` | GET | Retrieve record |
| `/volatile/{namespace}/{key}` | POST | Store/update record |
| `/volatile/{namespace}/{key}` | DELETE | Remove record |
| `/volatile/{namespace}` | GET | List keys in namespace |
| `/volatile/scheduled` | GET | List records needing refresh |
| `/volatile/stats` | GET | Cache statistics |
#### 2.3 Integrate with Consolidation
**File**: `src/services/consolidation_service.py`
Add relevance trigger detection:
1. During consolidation, analyze search results for location/interest patterns
2. Query tatlock's Biographer collection for user preferences
3. If match found, create/update volatile refresh schedule
#### 2.4 Biographer Integration
**File**: `src/core/dependencies.py`
```python
def get_biographer_qdrant() -> QdrantClientWrapper:
"""Direct access to tatlock's Biographer collection."""
# Configure to connect to tatlock's Qdrant
```
#### 2.5 Scheduler-Side Configuration
Document required scheduler tasks:
```json
{
"task_name": "volatile_refresh",
"schedule": "*/15 * * * *",
"endpoint": "GET /volatile/scheduled",
"follow_up": "For each record, call refresh endpoint with record.refresh_schedule"
}
```
---
## Phase 3: Document Storage (Research + Implementation)
### Research Scope
Evaluate FOSS self-hosted options for:
- Git repository mirroring
- PDF/document storage with metadata
- Image/video blob storage
- Full-text search capability
**Constraints**:
- Must be self-hosted, Docker-deployable
- Performance is priority (can wrap complexity in API)
- No cloud dependencies
**Candidates to evaluate**:
1. MinIO (S3-compatible object storage) + metadata in Neo4j
2. Paperless-ngx (document management with OCR)
3. SeaweedFS (distributed file system)
4. Custom: filesystem + Neo4j metadata
### Category Descriptors
**Wiki page structure for document collections**:
```markdown
# FastAPI Documentation
## Overview
[LLM-generated summary from web search about FastAPI]
## Collection Statistics
- **Documents**: 342 files
- **Last Sync**: 2025-12-24 03:30 UTC
- **Source**: github.com/tiangolo/fastapi
- **Coverage**: API reference, tutorials, deployment guides
## What's Included
[LLM summary of collection contents based on document analysis]
## Related Topics
- [[Python Web Frameworks]]
- [[REST API Design]]
```
### Tasks
#### 3.1 Storage Research
**Deliverable**: Evaluation document comparing options
#### 3.2 Storage Service Implementation
**New file**: `src/services/document_store_service.py`
(Details pending research results)
#### 3.3 Category Descriptor Generation
**File**: `src/services/consolidation_service.py`
Add LLM-powered category descriptor generation:
1. Web search for topic overview
2. Analyze collection contents
3. Generate/update wiki page with template
---
## Files to Modify/Create
### Phase 1 (Cleanup)
- `src/routers/maintenance.py` - Add timestamp tracking
- `src/services/graph_service.py` - Bidirectional validation
- `src/services/vector_service.py` - Cross-reference checks
- `LIBRARIAN_INTEGRATION.md` - Scheduler config docs
### Phase 2 (Volatile)
- `src/services/volatile_service.py` - **NEW**
- `src/routers/volatile.py` - **NEW**
- `src/models/volatile.py` - **NEW**
- `src/core/dependencies.py` - Add Biographer client
- `src/services/consolidation_service.py` - Relevance triggers
- `tests/test_volatile.py` - **NEW**
### Phase 3 (Documents)
- `docs/DOCUMENT_STORAGE_RESEARCH.md` - **NEW**
- `src/services/document_store_service.py` - **NEW** (post-research)
- `src/routers/documents.py` - **NEW** (post-research)
---
## Resolved Design Decisions
1. **Biographer Qdrant**: Same Qdrant instance, different collection. Library-Desk queries directly.
2. **Scheduler API**: Has REST API for task registration. Library-Desk can programmatically create refresh schedules.
3. **External API calls**: Library-Desk routes through SearXNG for web search. Consider dedicated API integrations for high-value volatiles (weather, financial) for consistent quality.
---
## Future Consideration: Dedicated API Integrations
For volatile data where quality/consistency matters (weather, financial), consider:
- OpenWeatherMap API for weather (daily refresh cycle)
- Financial data API (Alpha Vantage, Yahoo Finance)
- News APIs (NewsAPI, GDELT)
- **NOS.nl** - Explicit source for Dutch news
This would live in a new `src/clients/` module with:
- `weather_client.py` - Daily refresh cycle
- `financial_client.py`
- `news_client.py` - Include NOS.nl scraper/API for Dutch coverage
These provide structured, reliable data vs. SearXNG web scraping. Implementation deferred to later phase.
---
## Refresh Schedules
**Note:** TTL should be longer than refresh interval to prevent data gaps.
| Volatile Type | TTL | Refresh Cycle | Refresh Interval | Sources |
|---------------|-----|---------------|------------------|---------|
| Weather | 86400s (24hr) | Daily | Every 24hr | OpenWeatherMap |
| Dutch News | 28800s (8hr) | 4x daily | Every 6hr | NOS.nl |
| Global News | 28800s (8hr) | 4x daily | Every 6hr | NewsAPI, GDELT |
| Financial | 600s (10min) | On-demand | N/A | Alpha Vantage |
**TTL Logic:**
- TTL = Refresh Interval × 1.5 (buffer for failed refreshes)
- On-demand data gets shorter TTL since it's fetched when needed
+1 -1
View File
@@ -1,6 +1,6 @@
[project]
name = "library-desk"
version = "1.3.3"
version = "1.4.1"
description = "Coordination service for The Library system - HybridRAG queries, document ingestion, entity extraction, and knowledge consolidation"
readme = "README.md"
requires-python = ">=3.12"
+91
View File
@@ -551,6 +551,97 @@ class QdrantClientWrapper:
logger.error(f"Search failed: {e}", exc_info=True)
return []
async def scroll_all_points(
self,
collection_name: str,
batch_size: int = 100,
with_payload: bool = True,
with_vectors: bool = False,
filter_conditions: Optional[Dict[str, Any]] = None
) -> List[Dict[str, Any]]:
"""
Scroll through all points in a collection.
Args:
collection_name: Collection name
batch_size: Number of points per batch
with_payload: Include payload in results
with_vectors: Include vectors in results
filter_conditions: Optional filter conditions
Returns:
List of all points with id and payload
"""
all_points = []
offset = None
# Build filter if provided
scroll_filter = None
if filter_conditions:
conditions = []
for key, value in filter_conditions.items():
conditions.append(
FieldCondition(key=key, match=MatchValue(value=value))
)
scroll_filter = Filter(must=conditions)
try:
while True:
points, next_offset = self.client.scroll(
collection_name=collection_name,
scroll_filter=scroll_filter,
limit=batch_size,
offset=offset,
with_payload=with_payload,
with_vectors=with_vectors
)
for point in points:
all_points.append({
"id": str(point.id),
"payload": dict(point.payload) if point.payload else {}
})
if next_offset is None:
break
offset = next_offset
return all_points
except Exception as e:
logger.error(f"Failed to scroll collection {collection_name}: {e}", exc_info=True)
return []
async def delete_by_ids(
self,
collection_name: str,
point_ids: List[str]
) -> int:
"""
Delete points by their IDs.
Args:
collection_name: Collection name
point_ids: List of point IDs to delete
Returns:
Number of points deleted
"""
if not point_ids:
return 0
try:
self.client.delete(
collection_name=collection_name,
points_selector=point_ids
)
logger.info(f"Deleted {len(point_ids)} points from {collection_name}")
return len(point_ids)
except Exception as e:
logger.error(f"Failed to delete points by IDs: {e}", exc_info=True)
return 0
async def list_collections(self) -> List[Dict[str, Any]]:
"""
List all collections with stats.
+13 -81
View File
@@ -20,96 +20,34 @@ class WikiJSClient:
Wiki.js GraphQL API client.
Documentation: https://docs.requarks.io/dev/api
Authentication: Username/password login to get user-specific JWT token
Authentication: API token (JWT) generated from Wiki.js admin panel
"""
def __init__(self, base_url: str, username: str, password: str):
def __init__(self, base_url: str, api_token: str):
"""
Initialize Wiki.js client.
Args:
base_url: Wiki.js base URL (e.g., "http://wiki:3000")
username: Wiki.js username (e.g., "librarian@schweitz.net")
password: Wiki.js password
api_token: Wiki.js API token (JWT from admin panel)
"""
self.base_url = base_url.rstrip("/")
self.graphql_url = f"{self.base_url}/graphql"
self.username = username
self.password = password
self.jwt_token: Optional[str] = None
self.api_token = api_token
self.client = httpx.AsyncClient(timeout=30.0)
logger.info(f"Initialized Wiki.js client: {base_url} (user: {username})")
auth_mode = "with API token" if api_token else "without auth (open API)"
logger.info(f"Initialized Wiki.js client: {base_url} ({auth_mode})")
async def close(self):
"""Close HTTP client"""
await self.client.aclose()
async def login(self) -> bool:
"""
Authenticate with Wiki.js using username/password.
Returns:
True if login successful, False otherwise
"""
login_mutation = """
mutation Login($username: String!, $password: String!, $strategy: String!) {
authentication {
login(username: $username, password: $password, strategy: $strategy) {
responseResult {
succeeded
errorCode
message
}
jwt
}
}
}
"""
variables = {
"username": self.username,
"password": self.password,
"strategy": "local"
}
try:
response = await self.client.post(
self.graphql_url,
headers={"Content-Type": "application/json"},
json={"query": login_mutation, "variables": variables}
)
response.raise_for_status()
result = response.json()
if "errors" in result:
logger.error(f"Login failed: {result['errors']}")
return False
login_result = result.get("data", {}).get("authentication", {}).get("login", {})
response_result = login_result.get("responseResult", {})
if not response_result.get("succeeded"):
logger.error(f"Login failed: {response_result.get('message')}")
return False
self.jwt_token = login_result.get("jwt")
if not self.jwt_token:
logger.error("Login succeeded but no JWT token received")
return False
logger.info(f"Successfully authenticated as {self.username}")
return True
except Exception as e:
logger.error(f"Login failed: {e}", exc_info=True)
return False
async def _ensure_authenticated(self):
"""Ensure we have a valid JWT token, login if needed."""
if not self.jwt_token:
success = await self.login()
if not success:
raise Exception("Failed to authenticate with Wiki.js")
def _get_headers(self) -> Dict[str, str]:
"""Get request headers, optionally including auth token."""
headers = {"Content-Type": "application/json"}
if self.api_token:
headers["Authorization"] = f"Bearer {self.api_token}"
return headers
async def _execute_query(
self,
@@ -129,18 +67,12 @@ class WikiJSClient:
Raises:
Exception: If query fails or returns errors
"""
# Ensure we're authenticated before making requests
await self._ensure_authenticated()
payload = {
"query": query,
"variables": variables or {}
}
headers = {
"Authorization": f"Bearer {self.jwt_token}",
"Content-Type": "application/json"
}
headers = self._get_headers()
try:
response = await self.client.post(
+19 -2
View File
@@ -43,8 +43,10 @@ class Settings(BaseSettings):
# Wiki.js Configuration
wikijs_url: str = Field(default="http://wiki:3000", description="Wiki.js URL")
wikijs_username: str = Field(..., description="Wiki.js username")
wikijs_password: str = Field(..., description="Wiki.js password")
wiki_graphql_api: str = Field(default="", description="Wiki.js GraphQL API token (optional - API may be open)")
# Legacy auth fields - kept for backwards compatibility but deprecated
wikijs_username: str = Field(default="", description="Wiki.js username (deprecated, use wiki_graphql_api)")
wikijs_password: str = Field(default="", description="Wiki.js password (deprecated, use wiki_graphql_api)")
# Wiki.js Database Configuration (for change listener)
wikijs_db_host: str = Field(default="postgres-shared", description="Wiki.js PostgreSQL host")
@@ -99,6 +101,21 @@ class Settings(BaseSettings):
content_extraction_timeout: int = Field(default=5, ge=1, le=30, description="Trafilatura per-URL timeout in seconds")
content_max_length: int = Field(default=2000, ge=500, le=10000, description="Max extracted content length per result")
# Document Store Configuration
document_store_enabled: bool = Field(default=True, description="Enable document store feature")
document_catalog_path_prefix: str = Field(default="docs", description="Wiki path prefix for catalog pages")
# Volatile Cache Configuration
volatile_cache_enabled: bool = Field(default=True, description="Enable volatile cache feature")
volatile_default_ttl: int = Field(default=3600, ge=60, le=86400, description="Default TTL in seconds")
volatile_weather_ttl: int = Field(default=1800, ge=60, le=7200, description="Weather data TTL in seconds")
volatile_news_ttl: int = Field(default=7200, ge=300, le=86400, description="News data TTL in seconds")
volatile_financial_ttl: int = Field(default=300, ge=60, le=3600, description="Financial data TTL in seconds")
# Maintenance Configuration
maintenance_orphan_cleanup_enabled: bool = Field(default=True, description="Enable automatic orphan cleanup")
maintenance_cleanup_batch_size: int = Field(default=100, ge=10, le=1000, description="Cleanup batch size")
@property
def qdrant_url(self) -> str:
"""Computed Qdrant URL."""
+11 -3
View File
@@ -76,13 +76,12 @@ def get_wikijs_client() -> WikiJSClient:
Get Wiki.js client singleton.
Returns:
Initialized Wiki.js GraphQL client with username/password auth
Initialized Wiki.js GraphQL client with API token auth
"""
settings = get_settings()
client = WikiJSClient(
base_url=settings.wikijs_url,
username=settings.wikijs_username,
password=settings.wikijs_password
api_token=settings.wiki_graphql_api
)
logger.debug("Created Wiki.js client instance")
return client
@@ -425,3 +424,12 @@ async def verify_api_key(
detail="Invalid API key"
)
return credentials.credentials
# Service type aliases for FastAPI endpoint dependencies
# These are defined after the factory functions
from src.services.vector_service import VectorService
from src.services.graph_service import GraphService
VectorServiceDep = Annotated[VectorService, Depends(get_vector_service)]
GraphServiceDep = Annotated[GraphService, Depends(get_graph_service)]
+3 -1
View File
@@ -50,7 +50,8 @@ app.add_middleware(
# Register routers
from src.routers import (
wiki, tools, graph, vector, hybrid_rag, consolidation,
ingestion, entity_linking, webhooks, rag_search, content
ingestion, entity_linking, webhooks, rag_search, content,
maintenance
)
app.include_router(wiki.router)
@@ -64,6 +65,7 @@ app.include_router(entity_linking.router)
app.include_router(webhooks.router)
app.include_router(rag_search.router)
app.include_router(content.router)
app.include_router(maintenance.router)
# Mount static files directory for Wiki.js integration scripts
static_dir = Path(__file__).parent.parent / "static"
+794
View File
@@ -0,0 +1,794 @@
"""
Maintenance router for Library Desk cleanup operations.
Provides endpoints to clean up orphaned data in vectors and graph:
- Orphan vector chunks (no matching page/document in graph)
- Orphan entities (no MENTIONS relationships)
- Stale documents (graph nodes with no matching wiki page)
- Broken relationships
"""
from fastapi import APIRouter, HTTPException, Depends, Query
from pydantic import BaseModel, Field
from typing import Optional, List, Dict, Any
import logging
import time
from src.services.vector_service import VectorService
from src.services.graph_service import GraphService
from src.core.dependencies import (
VectorServiceDep, GraphServiceDep, WikiJSDep, RedisDep,
verify_api_key
)
from datetime import datetime, timezone
logger = logging.getLogger(__name__)
router = APIRouter(prefix="/maintenance", tags=["Maintenance"])
# Redis key for tracking last cleanup timestamp
LAST_CLEANUP_KEY = "library:maintenance:last_cleanup:{user}"
async def _get_last_cleanup(redis, user: str) -> Optional[str]:
"""Get last cleanup timestamp from Redis."""
try:
key = LAST_CLEANUP_KEY.format(user=user)
return await redis.get(key)
except Exception as e:
logger.warning(f"Failed to get last cleanup timestamp: {e}")
return None
async def _set_last_cleanup(redis, user: str) -> None:
"""Store current timestamp as last cleanup time."""
try:
key = LAST_CLEANUP_KEY.format(user=user)
timestamp = datetime.now(timezone.utc).isoformat()
# Keep for 30 days
await redis.setex(key, 86400 * 30, timestamp)
logger.info(f"Recorded cleanup timestamp: {timestamp}")
except Exception as e:
logger.warning(f"Failed to store cleanup timestamp: {e}")
async def _find_unindexed_pages(
wiki_pages: List[Dict],
chunk_refs: List[Dict],
graph_docs: List[Dict]
) -> tuple[List[int], List[int]]:
"""
Find wiki pages that are missing from vectors or graph.
Returns:
Tuple of (pages_without_vectors, pages_without_graph)
"""
# Build sets of indexed page IDs
vectorized_page_ids = {
ref.get("page_id") for ref in chunk_refs
if ref.get("doc_type") == "wiki" and ref.get("page_id")
}
graphed_page_ids = {
doc.get("page_id") for doc in graph_docs
if doc.get("doc_type") == "wiki" and doc.get("page_id")
}
# Find wiki pages missing from each store
pages_without_vectors = []
pages_without_graph = []
for page in wiki_pages:
page_id = page.get("id")
if not page_id:
continue
if page_id not in vectorized_page_ids:
pages_without_vectors.append(page_id)
if page_id not in graphed_page_ids:
pages_without_graph.append(page_id)
return pages_without_vectors, pages_without_graph
async def _reindex_missing_pages(
page_ids: List[int],
user: str,
vector_service,
graph_service
) -> tuple[int, int, List[int]]:
"""
Reindex pages that are missing from vectors or graph.
Returns:
Tuple of (pages_reindexed, pages_failed, failed_page_ids)
"""
reindexed = 0
failed = 0
failed_ids = []
for page_id in page_ids:
try:
# Index to both stores
vector_result = await vector_service.update_from_page(page_id, user, force_refresh=True)
graph_result = await graph_service.update_from_page(page_id, user, force_refresh=True)
if vector_result.success and graph_result.success:
reindexed += 1
logger.info(f"Reindexed missing page {page_id}")
else:
failed += 1
failed_ids.append(page_id)
logger.warning(f"Failed to reindex page {page_id}: vector={vector_result.success}, graph={graph_result.success}")
except Exception as e:
failed += 1
failed_ids.append(page_id)
logger.error(f"Error reindexing page {page_id}: {e}")
return reindexed, failed, failed_ids
# ========== Response Models ==========
class CleanupResult(BaseModel):
"""Result of a cleanup operation."""
orphans_found: int = Field(default=0, description="Number of orphans detected")
orphans_purged: int = Field(default=0, description="Number of orphans deleted")
duration_ms: float = Field(description="Operation duration in milliseconds")
class VectorCleanupResponse(BaseModel):
"""Response from vector cleanup operation."""
success: bool
wiki_chunks: CleanupResult
document_chunks: CleanupResult
chunks_without_graph: CleanupResult # Vectors with no graph node
total_chunks_scanned: int
total_orphans_purged: int
duration_ms: float
class GraphCleanupResponse(BaseModel):
"""Response from graph cleanup operation."""
success: bool
orphan_entities: CleanupResult
stale_wiki_documents: CleanupResult
stale_store_documents: CleanupResult
docs_without_vectors: CleanupResult # Graph nodes with no vectors
broken_relationships_cleaned: int
duration_ms: float
class FullCleanupResponse(BaseModel):
"""Response from full cleanup operation."""
success: bool
vector_cleanup: VectorCleanupResponse
graph_cleanup: GraphCleanupResponse
total_duration_ms: float
class HealthCheckResponse(BaseModel):
"""Response from maintenance health check."""
status: str = Field(description="Health status: healthy, degraded, or unhealthy")
orphan_vector_count: int = Field(description="Number of orphan vector chunks (no source)")
orphan_entity_count: int = Field(description="Number of orphan entities")
stale_document_count: int = Field(description="Number of stale document nodes")
vectors_without_graph: int = Field(default=0, description="Vector chunks with no graph node")
docs_without_vectors: int = Field(default=0, description="Graph docs with no vectors")
unindexed_pages: int = Field(default=0, description="Wiki pages missing from indexes")
last_cleanup: Optional[str] = Field(default=None, description="Timestamp of last cleanup")
recommendations: List[str] = Field(default_factory=list)
class ReindexResponse(BaseModel):
"""Response from reindex operation."""
success: bool
page_id: int
vectors_deleted: int
vectors_created: int
graph_updated: bool
duration_ms: float
error: Optional[str] = None
class ReindexMissingResult(BaseModel):
"""Result of reindexing missing pages."""
pages_without_vectors: int = Field(description="Wiki pages with no vector embeddings")
pages_without_graph: int = Field(description="Wiki pages with no graph Document node")
pages_reindexed: int = Field(description="Pages successfully reindexed")
pages_failed: int = Field(description="Pages that failed to reindex")
failed_page_ids: List[int] = Field(default_factory=list)
duration_ms: float
class ReconcileIndexResponse(BaseModel):
"""Response from reconcile-index operation (cleanup + reindex-missing)."""
success: bool
cleanup: FullCleanupResponse
reindex_missing: ReindexMissingResult
total_duration_ms: float
# ========== Endpoints ==========
@router.post("/cleanup/vectors", response_model=VectorCleanupResponse)
async def cleanup_vectors(
user: str = Query(..., description="User identifier"),
dry_run: bool = Query(False, description="If true, only count orphans without deleting"),
vector_service: VectorServiceDep = None,
graph_service: GraphServiceDep = None,
wiki_client: WikiJSDep = None,
api_key: str = Depends(verify_api_key)
):
"""
Find and purge orphan vector chunks.
Orphan chunks are vector embeddings that reference:
- Wiki pages that no longer exist
- Document Store documents that no longer exist
- Chunks with no corresponding graph Document node (bidirectional check)
**Scheduler Task** - Recommended to run daily.
"""
start_time = time.time()
try:
# Get all vector chunk references
chunk_refs = await vector_service.get_all_chunk_references(user)
total_scanned = len(chunk_refs)
# Get all valid page IDs from wiki
wiki_pages = await wiki_client.list_all_pages()
valid_page_ids = {p.get("id") for p in wiki_pages if p.get("id")}
# Get all valid document references from graph
graph_docs = await graph_service.get_all_document_references(user)
valid_doc_ids = {d["document_id"] for d in graph_docs if d.get("document_id")}
# Find orphan wiki chunks (page_id not in wiki)
wiki_orphan_ids = []
doc_orphan_ids = []
for ref in chunk_refs:
doc_type = ref.get("doc_type", "wiki")
if doc_type == "wiki":
page_id = ref.get("page_id")
if page_id and page_id not in valid_page_ids:
wiki_orphan_ids.append(ref["chunk_id"])
else:
document_id = ref.get("document_id")
if document_id and document_id not in valid_doc_ids:
doc_orphan_ids.append(ref["chunk_id"])
# Bidirectional check: chunks with no graph node
chunks_without_graph = vector_service.find_chunks_without_graph_nodes(
chunk_refs, graph_docs
)
# Purge orphans if not dry run
wiki_purged = 0
doc_purged = 0
graph_orphans_purged = 0
if not dry_run:
if wiki_orphan_ids:
wiki_purged = await vector_service.purge_chunks_by_ids(user, wiki_orphan_ids)
if doc_orphan_ids:
doc_purged = await vector_service.purge_chunks_by_ids(user, doc_orphan_ids)
if chunks_without_graph:
graph_orphans_purged = await vector_service.purge_chunks_by_ids(
user, chunks_without_graph
)
duration_ms = (time.time() - start_time) * 1000
return VectorCleanupResponse(
success=True,
wiki_chunks=CleanupResult(
orphans_found=len(wiki_orphan_ids),
orphans_purged=wiki_purged,
duration_ms=duration_ms / 3
),
document_chunks=CleanupResult(
orphans_found=len(doc_orphan_ids),
orphans_purged=doc_purged,
duration_ms=duration_ms / 3
),
chunks_without_graph=CleanupResult(
orphans_found=len(chunks_without_graph),
orphans_purged=graph_orphans_purged,
duration_ms=duration_ms / 3
),
total_chunks_scanned=total_scanned,
total_orphans_purged=wiki_purged + doc_purged + graph_orphans_purged,
duration_ms=duration_ms
)
except Exception as e:
logger.error(f"Vector cleanup failed: {e}", exc_info=True)
raise HTTPException(status_code=500, detail=str(e))
@router.post("/cleanup/graph", response_model=GraphCleanupResponse)
async def cleanup_graph(
user: str = Query(..., description="User identifier"),
dry_run: bool = Query(False, description="If true, only count orphans without deleting"),
vector_service: VectorServiceDep = None,
graph_service: GraphServiceDep = None,
wiki_client: WikiJSDep = None,
api_key: str = Depends(verify_api_key)
):
"""
Find and purge orphan entities and stale documents from the graph.
Cleans up:
- Orphan entities (no MENTIONS relationships)
- Stale wiki Document nodes (page deleted from Wiki.js)
- Stale Document Store nodes (document deleted)
- Graph Document nodes with no corresponding vectors (bidirectional check)
- Broken FOUND relationships from SearchQuery nodes
"""
start_time = time.time()
try:
# 1. Find orphan entities
orphan_entities = await graph_service.find_orphan_entities(user)
entities_purged = 0
if not dry_run and orphan_entities:
entities_purged = await graph_service.purge_orphan_entities(user)
# 2. Find stale wiki documents
graph_docs = await graph_service.get_all_document_references(user)
wiki_docs = [d for d in graph_docs if d.get("doc_type") == "wiki" and d.get("page_id")]
# Get valid wiki page IDs
wiki_pages = await wiki_client.list_all_pages()
valid_page_ids = {p.get("id") for p in wiki_pages if p.get("id")}
stale_wiki_ids = [d["page_id"] for d in wiki_docs if d["page_id"] not in valid_page_ids]
wiki_docs_purged = 0
if not dry_run and stale_wiki_ids:
wiki_docs_purged = await graph_service.purge_stale_documents_by_ids(
user, page_ids=stale_wiki_ids
)
# 3. Find stale Document Store documents (these would be detected differently)
# For now, Document Store docs are only stale if the collection is deleted
# This will be more relevant once DocumentService exists
stale_store_docs = 0
store_docs_purged = 0
# 4. Bidirectional check: graph docs with no vectors
chunk_refs = await vector_service.get_all_chunk_references(user)
docs_without_vectors = await graph_service.find_documents_without_vectors(
user, chunk_refs
)
docs_without_vectors_purged = 0
if not dry_run and docs_without_vectors:
# Purge wiki docs without vectors
wiki_orphans = [d["page_id"] for d in docs_without_vectors
if d.get("doc_type") == "wiki" and d.get("page_id")]
doc_orphans = [d["document_id"] for d in docs_without_vectors
if d.get("doc_type") != "wiki" and d.get("document_id")]
if wiki_orphans:
docs_without_vectors_purged += await graph_service.purge_stale_documents_by_ids(
user, page_ids=wiki_orphans
)
if doc_orphans:
docs_without_vectors_purged += await graph_service.purge_stale_documents_by_ids(
user, document_ids=doc_orphans
)
# 5. Clean broken relationships
broken_rels_cleaned = 0
if not dry_run:
broken_rels_cleaned = await graph_service.cleanup_broken_relationships(user)
duration_ms = (time.time() - start_time) * 1000
return GraphCleanupResponse(
success=True,
orphan_entities=CleanupResult(
orphans_found=len(orphan_entities),
orphans_purged=entities_purged,
duration_ms=duration_ms / 5
),
stale_wiki_documents=CleanupResult(
orphans_found=len(stale_wiki_ids),
orphans_purged=wiki_docs_purged,
duration_ms=duration_ms / 5
),
stale_store_documents=CleanupResult(
orphans_found=stale_store_docs,
orphans_purged=store_docs_purged,
duration_ms=duration_ms / 5
),
docs_without_vectors=CleanupResult(
orphans_found=len(docs_without_vectors),
orphans_purged=docs_without_vectors_purged,
duration_ms=duration_ms / 5
),
broken_relationships_cleaned=broken_rels_cleaned,
duration_ms=duration_ms
)
except Exception as e:
logger.error(f"Graph cleanup failed: {e}", exc_info=True)
raise HTTPException(status_code=500, detail=str(e))
@router.post("/cleanup/all", response_model=FullCleanupResponse)
async def cleanup_all(
user: str = Query(..., description="User identifier"),
dry_run: bool = Query(False, description="If true, only count orphans without deleting"),
vector_service: VectorServiceDep = None,
graph_service: GraphServiceDep = None,
wiki_client: WikiJSDep = None,
redis: RedisDep = None,
api_key: str = Depends(verify_api_key)
):
"""
Full cleanup of vectors and graph.
Runs both vector and graph cleanup in sequence.
**Scheduler Task** - Recommended to run daily at low-traffic time.
**Scheduler Integration:**
```json
{
"task_name": "library_maintenance",
"schedule": "0 4 * * *",
"endpoint": "POST /maintenance/cleanup/all?user=jpmschweitzer",
"description": "Daily cleanup of orphan vectors and graph nodes"
}
```
"""
start_time = time.time()
try:
# Run vector cleanup
vector_result = await cleanup_vectors(
user=user,
dry_run=dry_run,
vector_service=vector_service,
graph_service=graph_service,
wiki_client=wiki_client,
api_key=api_key
)
# Run graph cleanup
graph_result = await cleanup_graph(
user=user,
dry_run=dry_run,
vector_service=vector_service,
graph_service=graph_service,
wiki_client=wiki_client,
api_key=api_key
)
total_duration_ms = (time.time() - start_time) * 1000
# Record cleanup timestamp (only if not dry run)
if not dry_run and redis:
await _set_last_cleanup(redis, user)
return FullCleanupResponse(
success=True,
vector_cleanup=vector_result,
graph_cleanup=graph_result,
total_duration_ms=total_duration_ms
)
except Exception as e:
logger.error(f"Full cleanup failed: {e}", exc_info=True)
raise HTTPException(status_code=500, detail=str(e))
@router.get("/health", response_model=HealthCheckResponse)
async def maintenance_health(
user: str = Query(..., description="User identifier"),
detailed: bool = Query(False, description="If true, run full orphan analysis (slower)"),
vector_service: VectorServiceDep = None,
graph_service: GraphServiceDep = None,
wiki_client: WikiJSDep = None,
redis: RedisDep = None,
api_key: str = Depends(verify_api_key)
):
"""
Health check for maintenance status.
**Lightweight mode (default)**: Returns last cleanup timestamp and basic status.
Use for frequent uptime checks (every 30s).
**Detailed mode (?detailed=true)**: Runs full orphan/unindexed analysis.
Use for dashboards or before running reconcile-index.
"""
try:
# Get last cleanup timestamp from Redis (lightweight)
last_cleanup = None
if redis:
last_cleanup = await _get_last_cleanup(redis, user)
# Lightweight mode - just return basic status
if not detailed:
return HealthCheckResponse(
status="healthy" if last_cleanup else "unknown",
orphan_vector_count=0,
orphan_entity_count=0,
stale_document_count=0,
vectors_without_graph=0,
docs_without_vectors=0,
unindexed_pages=0,
last_cleanup=last_cleanup,
recommendations=[] if last_cleanup else ["No cleanup recorded. Run POST /maintenance/reconcile-index"]
)
# Detailed mode - full analysis
recommendations = []
# Count orphan vector chunks
chunk_refs = await vector_service.get_all_chunk_references(user)
wiki_pages = await wiki_client.list_all_pages()
valid_page_ids = {p.get("id") for p in wiki_pages if p.get("id")}
orphan_vector_count = sum(
1 for ref in chunk_refs
if ref.get("doc_type") == "wiki"
and ref.get("page_id") not in valid_page_ids
)
if orphan_vector_count > 10:
recommendations.append(
f"Found {orphan_vector_count} orphan vector chunks. "
"Consider running POST /maintenance/cleanup/vectors"
)
# Count orphan entities
orphan_entities = await graph_service.find_orphan_entities(user)
orphan_entity_count = len(orphan_entities)
if orphan_entity_count > 5:
recommendations.append(
f"Found {orphan_entity_count} orphan entities. "
"Consider running POST /maintenance/cleanup/graph"
)
# Count stale documents
graph_docs = await graph_service.get_all_document_references(user)
wiki_docs = [d for d in graph_docs if d.get("doc_type") == "wiki" and d.get("page_id")]
stale_document_count = sum(1 for d in wiki_docs if d["page_id"] not in valid_page_ids)
if stale_document_count > 0:
recommendations.append(
f"Found {stale_document_count} stale Document nodes. "
"Consider running POST /maintenance/cleanup/graph"
)
# Bidirectional: vectors without graph nodes
vectors_without_graph = len(vector_service.find_chunks_without_graph_nodes(
chunk_refs, graph_docs
))
if vectors_without_graph > 5:
recommendations.append(
f"Found {vectors_without_graph} vectors without graph nodes. "
"Consider running POST /maintenance/cleanup/vectors"
)
# Bidirectional: graph docs without vectors
docs_without_vectors_list = await graph_service.find_documents_without_vectors(
user, chunk_refs
)
docs_without_vectors = len(docs_without_vectors_list)
if docs_without_vectors > 5:
recommendations.append(
f"Found {docs_without_vectors} graph docs without vectors. "
"Consider running POST /maintenance/cleanup/graph"
)
# Unindexed pages: wiki pages missing from vectors or graph
pages_without_vectors, pages_without_graph = await _find_unindexed_pages(
wiki_pages, chunk_refs, graph_docs
)
unindexed_pages = len(set(pages_without_vectors + pages_without_graph))
if unindexed_pages > 0:
recommendations.append(
f"Found {unindexed_pages} wiki pages not in indexes. "
"Consider running POST /maintenance/reconcile-index"
)
# Determine overall status
total_issues = (orphan_vector_count + orphan_entity_count + stale_document_count +
vectors_without_graph + docs_without_vectors + unindexed_pages)
if total_issues == 0:
status = "healthy"
elif total_issues < 20:
status = "degraded"
else:
status = "unhealthy"
# Get last cleanup timestamp from Redis
last_cleanup = None
if redis:
last_cleanup = await _get_last_cleanup(redis, user)
return HealthCheckResponse(
status=status,
orphan_vector_count=orphan_vector_count,
orphan_entity_count=orphan_entity_count,
stale_document_count=stale_document_count,
vectors_without_graph=vectors_without_graph,
docs_without_vectors=docs_without_vectors,
unindexed_pages=unindexed_pages,
last_cleanup=last_cleanup,
recommendations=recommendations
)
except Exception as e:
logger.error(f"Health check failed: {e}", exc_info=True)
return HealthCheckResponse(
status="unhealthy",
orphan_vector_count=-1,
orphan_entity_count=-1,
stale_document_count=-1,
vectors_without_graph=-1,
docs_without_vectors=-1,
unindexed_pages=-1,
recommendations=[f"Health check failed: {str(e)}"]
)
@router.post("/reindex/{page_id}", response_model=ReindexResponse)
async def reindex_page(
page_id: int,
user: str = Query(..., description="User identifier"),
vector_service: VectorServiceDep = None,
graph_service: GraphServiceDep = None,
api_key: str = Depends(verify_api_key)
):
"""
Force re-index a wiki page.
Deletes existing vectors and graph data, then re-ingests.
Useful for fixing corrupted or stale data for a specific page.
"""
start_time = time.time()
try:
# Delete existing vectors
vectors_deleted = await vector_service.delete_page_chunks(page_id, user)
# Delete and recreate graph node
await graph_service.delete_page(page_id, user)
# Re-ingest
vector_result = await vector_service.update_from_page(page_id, user, force_refresh=True)
graph_result = await graph_service.update_from_page(page_id, user, force_refresh=True)
duration_ms = (time.time() - start_time) * 1000
return ReindexResponse(
success=vector_result.success and graph_result.success,
page_id=page_id,
vectors_deleted=vectors_deleted,
vectors_created=vector_result.chunks_created,
graph_updated=graph_result.success,
duration_ms=duration_ms,
error=vector_result.error_message or graph_result.error_message
)
except Exception as e:
duration_ms = (time.time() - start_time) * 1000
logger.error(f"Reindex failed for page {page_id}: {e}", exc_info=True)
return ReindexResponse(
success=False,
page_id=page_id,
vectors_deleted=0,
vectors_created=0,
graph_updated=False,
duration_ms=duration_ms,
error=str(e)
)
@router.post("/reconcile-index", response_model=ReconcileIndexResponse)
async def reconcile_index(
user: str = Query(..., description="User identifier"),
dry_run: bool = Query(False, description="If true, only detect issues without fixing"),
vector_service: VectorServiceDep = None,
graph_service: GraphServiceDep = None,
wiki_client: WikiJSDep = None,
redis: RedisDep = None,
api_key: str = Depends(verify_api_key)
):
"""
Full index reconciliation: cleanup orphans + reindex missing pages.
This is the recommended daily maintenance endpoint. It:
1. Cleans up orphan vectors and graph nodes (data without sources)
2. Reindexes wiki pages that are missing from vectors or graph
**Scheduler Task** - Recommended to run daily at low-traffic time.
**Scheduler Integration:**
```json
{
"task_name": "library_reconcile_index",
"schedule": "0 4 * * *",
"endpoint": "POST /maintenance/reconcile-index?user=jpmschweitzer",
"description": "Daily index reconciliation - cleanup + reindex missing"
}
```
"""
start_time = time.time()
try:
# Phase 1: Run full cleanup
cleanup_result = await cleanup_all(
user=user,
dry_run=dry_run,
vector_service=vector_service,
graph_service=graph_service,
wiki_client=wiki_client,
redis=redis,
api_key=api_key
)
# Phase 2: Find and reindex missing pages
reindex_start = time.time()
# Get current state
wiki_pages = await wiki_client.list_all_pages()
chunk_refs = await vector_service.get_all_chunk_references(user)
graph_docs = await graph_service.get_all_document_references(user)
# Find pages missing from indexes
pages_without_vectors, pages_without_graph = await _find_unindexed_pages(
wiki_pages, chunk_refs, graph_docs
)
# Combine unique page IDs that need reindexing
missing_page_ids = list(set(pages_without_vectors + pages_without_graph))
# Reindex missing pages (unless dry run)
reindexed = 0
failed = 0
failed_ids = []
if not dry_run and missing_page_ids:
reindexed, failed, failed_ids = await _reindex_missing_pages(
missing_page_ids, user, vector_service, graph_service
)
reindex_duration = (time.time() - reindex_start) * 1000
total_duration = (time.time() - start_time) * 1000
# Record reconciliation timestamp
if not dry_run and redis:
await _set_last_cleanup(redis, user)
return ReconcileIndexResponse(
success=cleanup_result.success and failed == 0,
cleanup=cleanup_result,
reindex_missing=ReindexMissingResult(
pages_without_vectors=len(pages_without_vectors),
pages_without_graph=len(pages_without_graph),
pages_reindexed=reindexed,
pages_failed=failed,
failed_page_ids=failed_ids,
duration_ms=reindex_duration
),
total_duration_ms=total_duration
)
except Exception as e:
logger.error(f"Reconcile-index failed: {e}", exc_info=True)
raise HTTPException(status_code=500, detail=str(e))
+350
View File
@@ -1261,3 +1261,353 @@ Feel free to expand it with more details!
except Exception as e:
logger.error(f"Failed to create entity mentions: {e}", exc_info=True)
return 0
# ========== Cleanup Methods ==========
async def delete_document_node(
self,
document_id: str,
user: str
) -> int:
"""
Delete a Document Store document node and all its relationships.
Args:
document_id: Document UUID (Document Store)
user: User identifier
Returns:
Number of nodes deleted (1 if successful, 0 if not found)
"""
user_doc_label = get_neo4j_user_label(user)
delete_query = f"""
MATCH (d:{user_doc_label}:Document {{document_id: $document_id}})
DETACH DELETE d
RETURN count(d) as deleted_count
"""
try:
result = await self.neo4j.execute_query(
delete_query,
{"document_id": document_id}
)
deleted_count = result[0]["deleted_count"] if result else 0
if deleted_count > 0:
logger.info(f"Deleted Document node for document {document_id}")
else:
logger.warning(f"No Document node found for document {document_id}")
return deleted_count
except Exception as e:
logger.error(f"Failed to delete document {document_id} from graph: {e}", exc_info=True)
return 0
async def delete_collection_node(
self,
collection_id: str,
user: str
) -> int:
"""
Delete a DocumentCollection node and all contained documents.
Args:
collection_id: Collection UUID
user: User identifier
Returns:
Number of nodes deleted (collection + documents)
"""
user_doc_label = get_neo4j_user_label(user)
# Delete collection and all documents it contains
delete_query = f"""
MATCH (c:{user_doc_label}:DocumentCollection {{id: $collection_id}})
OPTIONAL MATCH (c)-[:CONTAINS]->(d:Document)
DETACH DELETE c, d
RETURN count(c) + count(d) as deleted_count
"""
try:
result = await self.neo4j.execute_query(
delete_query,
{"collection_id": collection_id}
)
deleted_count = result[0]["deleted_count"] if result else 0
logger.info(f"Deleted collection {collection_id} with {deleted_count} total nodes")
return deleted_count
except Exception as e:
logger.error(f"Failed to delete collection {collection_id}: {e}", exc_info=True)
return 0
async def find_orphan_entities(
self,
user: str
) -> List[Dict[str, Any]]:
"""
Find entities with no MENTIONS relationships (orphaned).
Args:
user: User identifier
Returns:
List of orphaned entities {id, name, type}
"""
from src.core.multi_tenancy import get_neo4j_user_base_label
user_base_label = get_neo4j_user_base_label(user)
query = f"""
MATCH (e:{user_base_label})
WHERE NOT e:Document
AND NOT e:DocumentCollection
AND NOT EXISTS {{ (d:Document)-[:MENTIONS]->(e) }}
RETURN elementId(e) as id, e.name as name, labels(e) as labels
"""
try:
results = await self.neo4j.execute_query(query, {})
orphans = []
for r in results:
labels = r.get("labels", [])
entity_type = next(
(l for l in labels if l != user_base_label),
"Unknown"
)
orphans.append({
"id": r["id"],
"name": r["name"],
"type": entity_type
})
logger.info(f"Found {len(orphans)} orphan entities for user {user}")
return orphans
except Exception as e:
logger.error(f"Failed to find orphan entities: {e}", exc_info=True)
return []
async def purge_orphan_entities(
self,
user: str
) -> int:
"""
Delete all orphaned entities (entities with no MENTIONS relationships).
Args:
user: User identifier
Returns:
Number of entities purged
"""
from src.core.multi_tenancy import get_neo4j_user_base_label
user_base_label = get_neo4j_user_base_label(user)
query = f"""
MATCH (e:{user_base_label})
WHERE NOT e:Document
AND NOT e:DocumentCollection
AND NOT EXISTS {{ (d:Document)-[:MENTIONS]->(e) }}
DETACH DELETE e
RETURN count(e) as purged_count
"""
try:
results = await self.neo4j.execute_query(query, {})
purged_count = results[0]["purged_count"] if results else 0
logger.info(f"Purged {purged_count} orphan entities for user {user}")
return purged_count
except Exception as e:
logger.error(f"Failed to purge orphan entities: {e}", exc_info=True)
return 0
async def get_all_document_references(
self,
user: str
) -> List[Dict[str, Any]]:
"""
Get all Document node references for orphan detection.
Returns page_id for wiki docs and document_id for Document Store docs.
Args:
user: User identifier
Returns:
List of document references {page_id, document_id, doc_type, title}
"""
user_doc_label = get_neo4j_user_label(user)
query = f"""
MATCH (d:{user_doc_label}:Document)
RETURN d.page_id as page_id,
d.document_id as document_id,
COALESCE(d.doc_type, 'wiki') as doc_type,
d.title as title
"""
try:
results = await self.neo4j.execute_query(query, {})
references = []
for r in results:
references.append({
"page_id": r.get("page_id"),
"document_id": r.get("document_id"),
"doc_type": r.get("doc_type", "wiki"),
"title": r.get("title")
})
logger.info(f"Found {len(references)} document references for user {user}")
return references
except Exception as e:
logger.error(f"Failed to get document references: {e}", exc_info=True)
return []
async def purge_stale_documents_by_ids(
self,
user: str,
page_ids: List[int] = None,
document_ids: List[str] = None
) -> int:
"""
Delete specific stale Document nodes by their IDs.
Args:
user: User identifier
page_ids: List of wiki page IDs to delete
document_ids: List of Document Store document IDs to delete
Returns:
Number of documents purged
"""
user_doc_label = get_neo4j_user_label(user)
total_purged = 0
try:
# Purge by page_id (wiki docs)
if page_ids:
query = f"""
MATCH (d:{user_doc_label}:Document)
WHERE d.page_id IN $page_ids
DETACH DELETE d
RETURN count(d) as purged_count
"""
results = await self.neo4j.execute_query(query, {"page_ids": page_ids})
count = results[0]["purged_count"] if results else 0
total_purged += count
logger.info(f"Purged {count} wiki Document nodes")
# Purge by document_id (Document Store docs)
if document_ids:
query = f"""
MATCH (d:{user_doc_label}:Document)
WHERE d.document_id IN $document_ids
DETACH DELETE d
RETURN count(d) as purged_count
"""
results = await self.neo4j.execute_query(query, {"document_ids": document_ids})
count = results[0]["purged_count"] if results else 0
total_purged += count
logger.info(f"Purged {count} Document Store Document nodes")
return total_purged
except Exception as e:
logger.error(f"Failed to purge stale documents: {e}", exc_info=True)
return 0
async def cleanup_broken_relationships(
self,
user: str
) -> int:
"""
Clean up broken FOUND relationships from SearchQuery nodes.
Removes relationships pointing to deleted documents.
Args:
user: User identifier
Returns:
Number of relationships cleaned
"""
query = """
MATCH (sq:SearchQuery)-[r:FOUND]->(d)
WHERE NOT EXISTS { (d) }
DELETE r
RETURN count(r) as cleaned_count
"""
try:
results = await self.neo4j.execute_query(query, {})
cleaned_count = results[0]["cleaned_count"] if results else 0
if cleaned_count > 0:
logger.info(f"Cleaned {cleaned_count} broken FOUND relationships")
return cleaned_count
except Exception as e:
logger.error(f"Failed to cleanup broken relationships: {e}", exc_info=True)
return 0
async def find_documents_without_vectors(
self,
user: str,
vector_references: List[Dict[str, Any]]
) -> List[Dict[str, Any]]:
"""
Find Document nodes that have no corresponding vectors.
Used for bidirectional orphan detection - graph nodes without vector data.
Args:
user: User identifier
vector_references: List of vector refs from VectorService.get_all_chunk_references()
Returns:
List of orphan documents {page_id, document_id, doc_type, title}
"""
# Get all graph document references
graph_docs = await self.get_all_document_references(user)
if not graph_docs:
return []
# Build sets of IDs that have vectors
vector_page_ids = {
ref.get("page_id") for ref in vector_references
if ref.get("doc_type") == "wiki" and ref.get("page_id")
}
vector_doc_ids = {
ref.get("document_id") for ref in vector_references
if ref.get("doc_type") != "wiki" and ref.get("document_id")
}
# Find graph docs with no vectors
orphans = []
for doc in graph_docs:
doc_type = doc.get("doc_type", "wiki")
if doc_type == "wiki":
page_id = doc.get("page_id")
if page_id and page_id not in vector_page_ids:
orphans.append(doc)
else:
document_id = doc.get("document_id")
if document_id and document_id not in vector_doc_ids:
orphans.append(doc)
logger.info(f"Found {len(orphans)} graph documents without vectors for user {user}")
return orphans
+186
View File
@@ -356,3 +356,189 @@ class VectorService:
collections=[],
total=0
)
# ========== Cleanup Methods ==========
async def delete_document_chunks(
self,
document_id: str,
user: str
) -> int:
"""
Delete all chunks for a document (Document Store).
Args:
document_id: Document UUID
user: User identifier
Returns:
Number of chunks deleted
"""
collection_name = get_qdrant_collection_name(user)
try:
deleted_count = await self.qdrant.delete_by_filter(
collection_name=collection_name,
filter_conditions={"document_id": document_id}
)
logger.info(f"Deleted chunks for document {document_id}")
return deleted_count
except Exception as e:
logger.error(f"Failed to delete chunks for document {document_id}: {e}", exc_info=True)
return 0
async def delete_collection_chunks(
self,
collection_id: str,
user: str
) -> int:
"""
Delete all chunks for a document collection.
Args:
collection_id: Collection UUID
user: User identifier
Returns:
Number of chunks deleted
"""
collection_name = get_qdrant_collection_name(user)
try:
deleted_count = await self.qdrant.delete_by_filter(
collection_name=collection_name,
filter_conditions={"collection_id": collection_id}
)
logger.info(f"Deleted chunks for collection {collection_id}")
return deleted_count
except Exception as e:
logger.error(f"Failed to delete chunks for collection {collection_id}: {e}", exc_info=True)
return 0
async def get_all_chunk_references(
self,
user: str
) -> List[Dict[str, Any]]:
"""
Get all chunk references for orphan detection.
Returns list of {id, page_id, document_id} for all chunks.
Args:
user: User identifier
Returns:
List of chunk references
"""
collection_name = get_qdrant_collection_name(user)
try:
# Check if collection exists
exists = await self.qdrant.collection_exists(collection_name)
if not exists:
return []
all_points = await self.qdrant.scroll_all_points(
collection_name=collection_name,
batch_size=100,
with_payload=True
)
references = []
for point in all_points:
payload = point.get("payload", {})
references.append({
"chunk_id": point["id"],
"page_id": payload.get("page_id"),
"document_id": payload.get("document_id"),
"collection_id": payload.get("collection_id"),
"doc_type": payload.get("doc_type", "wiki")
})
logger.info(f"Found {len(references)} chunks for user {user}")
return references
except Exception as e:
logger.error(f"Failed to get chunk references: {e}", exc_info=True)
return []
async def purge_chunks_by_ids(
self,
user: str,
chunk_ids: List[str]
) -> int:
"""
Delete specific chunks by their IDs.
Args:
user: User identifier
chunk_ids: List of chunk IDs to delete
Returns:
Number of chunks deleted
"""
if not chunk_ids:
return 0
collection_name = get_qdrant_collection_name(user)
try:
deleted_count = await self.qdrant.delete_by_ids(
collection_name=collection_name,
point_ids=chunk_ids
)
logger.info(f"Purged {deleted_count} orphan chunks for user {user}")
return deleted_count
except Exception as e:
logger.error(f"Failed to purge chunks: {e}", exc_info=True)
return 0
def find_chunks_without_graph_nodes(
self,
chunk_references: List[Dict[str, Any]],
graph_references: List[Dict[str, Any]]
) -> List[str]:
"""
Find vector chunks that have no corresponding graph Document node.
Used for bidirectional orphan detection - vectors without graph representation.
Args:
chunk_references: List from get_all_chunk_references()
graph_references: List from GraphService.get_all_document_references()
Returns:
List of orphan chunk IDs
"""
# Build sets of IDs that have graph nodes
graph_page_ids = {
ref.get("page_id") for ref in graph_references
if ref.get("doc_type") == "wiki" and ref.get("page_id")
}
graph_doc_ids = {
ref.get("document_id") for ref in graph_references
if ref.get("doc_type") != "wiki" and ref.get("document_id")
}
# Find chunks with no graph node
orphan_ids = []
for chunk in chunk_references:
doc_type = chunk.get("doc_type", "wiki")
if doc_type == "wiki":
page_id = chunk.get("page_id")
if page_id and page_id not in graph_page_ids:
orphan_ids.append(chunk["chunk_id"])
else:
document_id = chunk.get("document_id")
if document_id and document_id not in graph_doc_ids:
orphan_ids.append(chunk["chunk_id"])
logger.info(f"Found {len(orphan_ids)} vector chunks without graph nodes")
return orphan_ids
+488
View File
@@ -0,0 +1,488 @@
"""
Tests for maintenance router and cleanup functionality.
Tests cleanup of:
- Orphan vector chunks
- Orphan entities in graph
- Stale document nodes
"""
import pytest
from unittest.mock import AsyncMock, MagicMock, patch
from src.routers.maintenance import (
cleanup_vectors,
cleanup_graph,
cleanup_all,
maintenance_health,
reindex_page,
CleanupResult,
VectorCleanupResponse,
GraphCleanupResponse,
FullCleanupResponse,
HealthCheckResponse,
ReindexResponse
)
class TestCleanupResult:
"""Test CleanupResult model."""
def test_cleanup_result_defaults(self):
"""Test CleanupResult with default values."""
result = CleanupResult(duration_ms=100.0)
assert result.orphans_found == 0
assert result.orphans_purged == 0
assert result.duration_ms == 100.0
def test_cleanup_result_with_values(self):
"""Test CleanupResult with actual values."""
result = CleanupResult(
orphans_found=10,
orphans_purged=8,
duration_ms=250.5
)
assert result.orphans_found == 10
assert result.orphans_purged == 8
assert result.duration_ms == 250.5
class TestVectorCleanupResponse:
"""Test VectorCleanupResponse model."""
def test_vector_cleanup_response(self):
"""Test VectorCleanupResponse structure."""
response = VectorCleanupResponse(
success=True,
wiki_chunks=CleanupResult(orphans_found=5, orphans_purged=5, duration_ms=50),
document_chunks=CleanupResult(orphans_found=3, orphans_purged=3, duration_ms=50),
chunks_without_graph=CleanupResult(orphans_found=2, orphans_purged=2, duration_ms=50),
total_chunks_scanned=100,
total_orphans_purged=10,
duration_ms=100
)
assert response.success is True
assert response.wiki_chunks.orphans_found == 5
assert response.document_chunks.orphans_found == 3
assert response.chunks_without_graph.orphans_found == 2
assert response.total_orphans_purged == 10
class TestGraphCleanupResponse:
"""Test GraphCleanupResponse model."""
def test_graph_cleanup_response(self):
"""Test GraphCleanupResponse structure."""
response = GraphCleanupResponse(
success=True,
orphan_entities=CleanupResult(orphans_found=10, orphans_purged=10, duration_ms=25),
stale_wiki_documents=CleanupResult(orphans_found=2, orphans_purged=2, duration_ms=25),
stale_store_documents=CleanupResult(orphans_found=0, orphans_purged=0, duration_ms=25),
docs_without_vectors=CleanupResult(orphans_found=1, orphans_purged=1, duration_ms=25),
broken_relationships_cleaned=5,
duration_ms=100
)
assert response.success is True
assert response.orphan_entities.orphans_found == 10
assert response.docs_without_vectors.orphans_found == 1
assert response.broken_relationships_cleaned == 5
class TestHealthCheckResponse:
"""Test HealthCheckResponse model."""
def test_health_check_healthy(self):
"""Test healthy status."""
response = HealthCheckResponse(
status="healthy",
orphan_vector_count=0,
orphan_entity_count=0,
stale_document_count=0
)
assert response.status == "healthy"
assert response.recommendations == []
def test_health_check_degraded(self):
"""Test degraded status with recommendations."""
response = HealthCheckResponse(
status="degraded",
orphan_vector_count=15,
orphan_entity_count=3,
stale_document_count=0,
recommendations=[
"Found 15 orphan vector chunks. Consider running POST /maintenance/cleanup/vectors"
]
)
assert response.status == "degraded"
assert len(response.recommendations) == 1
@pytest.mark.asyncio
class TestVectorCleanup:
"""Test vector cleanup endpoint."""
async def test_cleanup_vectors_no_orphans(self):
"""Test cleanup when no orphans exist."""
# Mock services
vector_service = AsyncMock()
vector_service.get_all_chunk_references.return_value = [
{"chunk_id": "c1", "page_id": 1, "doc_type": "wiki"}
]
# find_chunks_without_graph_nodes is not async
vector_service.find_chunks_without_graph_nodes = MagicMock(return_value=[])
graph_service = AsyncMock()
graph_service.get_all_document_references.return_value = [
{"page_id": 1, "doc_type": "wiki", "title": "Test"}
]
wiki_client = AsyncMock()
wiki_client.list_all_pages.return_value = [{"id": 1, "path": "test"}]
# Call cleanup
result = await cleanup_vectors(
user="testuser",
dry_run=False,
vector_service=vector_service,
graph_service=graph_service,
wiki_client=wiki_client,
api_key="test"
)
assert result.success is True
assert result.wiki_chunks.orphans_found == 0
assert result.chunks_without_graph.orphans_found == 0
assert result.total_orphans_purged == 0
async def test_cleanup_vectors_with_orphans(self):
"""Test cleanup when orphans exist."""
# Mock services
vector_service = AsyncMock()
vector_service.get_all_chunk_references.return_value = [
{"chunk_id": "c1", "page_id": 1, "doc_type": "wiki"},
{"chunk_id": "c2", "page_id": 999, "doc_type": "wiki"}, # Orphan
{"chunk_id": "c3", "page_id": 999, "doc_type": "wiki"}, # Orphan
]
vector_service.purge_chunks_by_ids.return_value = 2
# find_chunks_without_graph_nodes is not async
vector_service.find_chunks_without_graph_nodes = MagicMock(return_value=[])
graph_service = AsyncMock()
graph_service.get_all_document_references.return_value = [
{"page_id": 1, "doc_type": "wiki", "title": "Test"}
]
wiki_client = AsyncMock()
wiki_client.list_all_pages.return_value = [{"id": 1, "path": "test"}]
# Call cleanup
result = await cleanup_vectors(
user="testuser",
dry_run=False,
vector_service=vector_service,
graph_service=graph_service,
wiki_client=wiki_client,
api_key="test"
)
assert result.success is True
assert result.wiki_chunks.orphans_found == 2
assert result.wiki_chunks.orphans_purged == 2
assert result.total_orphans_purged == 2
async def test_cleanup_vectors_dry_run(self):
"""Test cleanup dry run doesn't purge."""
# Mock services
vector_service = AsyncMock()
vector_service.get_all_chunk_references.return_value = [
{"chunk_id": "c1", "page_id": 999, "doc_type": "wiki"}, # Orphan
]
# find_chunks_without_graph_nodes is not async
vector_service.find_chunks_without_graph_nodes = MagicMock(return_value=[])
graph_service = AsyncMock()
graph_service.get_all_document_references.return_value = []
wiki_client = AsyncMock()
wiki_client.list_all_pages.return_value = []
# Call cleanup in dry run mode
result = await cleanup_vectors(
user="testuser",
dry_run=True,
vector_service=vector_service,
graph_service=graph_service,
wiki_client=wiki_client,
api_key="test"
)
assert result.success is True
assert result.wiki_chunks.orphans_found == 1
assert result.wiki_chunks.orphans_purged == 0 # Not purged due to dry run
vector_service.purge_chunks_by_ids.assert_not_called()
@pytest.mark.asyncio
class TestGraphCleanup:
"""Test graph cleanup endpoint."""
async def test_cleanup_graph_no_orphans(self):
"""Test cleanup when no orphans exist."""
vector_service = AsyncMock()
vector_service.get_all_chunk_references.return_value = [
{"chunk_id": "c1", "page_id": 1, "doc_type": "wiki"}
]
graph_service = AsyncMock()
graph_service.find_orphan_entities.return_value = []
graph_service.get_all_document_references.return_value = [
{"page_id": 1, "doc_type": "wiki", "title": "Test"}
]
graph_service.find_documents_without_vectors.return_value = []
graph_service.cleanup_broken_relationships.return_value = 0
wiki_client = AsyncMock()
wiki_client.list_all_pages.return_value = [{"id": 1, "path": "test"}]
result = await cleanup_graph(
user="testuser",
dry_run=False,
vector_service=vector_service,
graph_service=graph_service,
wiki_client=wiki_client,
api_key="test"
)
assert result.success is True
assert result.orphan_entities.orphans_found == 0
assert result.stale_wiki_documents.orphans_found == 0
assert result.docs_without_vectors.orphans_found == 0
async def test_cleanup_graph_with_orphan_entities(self):
"""Test cleanup of orphan entities."""
vector_service = AsyncMock()
vector_service.get_all_chunk_references.return_value = []
graph_service = AsyncMock()
graph_service.find_orphan_entities.return_value = [
{"id": "e1", "name": "Orphan1", "type": "Person"},
{"id": "e2", "name": "Orphan2", "type": "Technology"},
]
graph_service.purge_orphan_entities.return_value = 2
graph_service.get_all_document_references.return_value = []
graph_service.find_documents_without_vectors.return_value = []
graph_service.cleanup_broken_relationships.return_value = 0
wiki_client = AsyncMock()
wiki_client.list_all_pages.return_value = []
result = await cleanup_graph(
user="testuser",
dry_run=False,
vector_service=vector_service,
graph_service=graph_service,
wiki_client=wiki_client,
api_key="test"
)
assert result.success is True
assert result.orphan_entities.orphans_found == 2
assert result.orphan_entities.orphans_purged == 2
async def test_cleanup_graph_with_stale_documents(self):
"""Test cleanup of stale document nodes."""
vector_service = AsyncMock()
vector_service.get_all_chunk_references.return_value = [
{"chunk_id": "c1", "page_id": 1, "doc_type": "wiki"}
]
graph_service = AsyncMock()
graph_service.find_orphan_entities.return_value = []
graph_service.get_all_document_references.return_value = [
{"page_id": 1, "doc_type": "wiki", "title": "Exists"},
{"page_id": 999, "doc_type": "wiki", "title": "Deleted"}, # Stale
]
graph_service.find_documents_without_vectors.return_value = []
graph_service.purge_stale_documents_by_ids.return_value = 1
graph_service.cleanup_broken_relationships.return_value = 0
wiki_client = AsyncMock()
wiki_client.list_all_pages.return_value = [{"id": 1, "path": "test"}]
result = await cleanup_graph(
user="testuser",
dry_run=False,
vector_service=vector_service,
graph_service=graph_service,
wiki_client=wiki_client,
api_key="test"
)
assert result.success is True
assert result.stale_wiki_documents.orphans_found == 1
assert result.stale_wiki_documents.orphans_purged == 1
@pytest.mark.asyncio
class TestFullCleanup:
"""Test full cleanup endpoint."""
async def test_full_cleanup(self):
"""Test full cleanup runs both vector and graph cleanup."""
vector_service = AsyncMock()
vector_service.get_all_chunk_references.return_value = []
# find_chunks_without_graph_nodes is not async
vector_service.find_chunks_without_graph_nodes = MagicMock(return_value=[])
graph_service = AsyncMock()
graph_service.find_orphan_entities.return_value = []
graph_service.get_all_document_references.return_value = []
graph_service.find_documents_without_vectors.return_value = []
graph_service.cleanup_broken_relationships.return_value = 0
wiki_client = AsyncMock()
wiki_client.list_all_pages.return_value = []
result = await cleanup_all(
user="testuser",
dry_run=False,
vector_service=vector_service,
graph_service=graph_service,
wiki_client=wiki_client,
api_key="test"
)
assert result.success is True
assert result.vector_cleanup.success is True
assert result.graph_cleanup.success is True
@pytest.mark.asyncio
class TestMaintenanceHealth:
"""Test maintenance health endpoint."""
async def test_health_healthy(self):
"""Test healthy status when no orphans."""
vector_service = AsyncMock()
vector_service.get_all_chunk_references.return_value = []
# find_chunks_without_graph_nodes is not async
vector_service.find_chunks_without_graph_nodes = MagicMock(return_value=[])
graph_service = AsyncMock()
graph_service.find_orphan_entities.return_value = []
graph_service.get_all_document_references.return_value = []
graph_service.find_documents_without_vectors.return_value = []
wiki_client = AsyncMock()
wiki_client.list_all_pages.return_value = []
result = await maintenance_health(
user="testuser",
vector_service=vector_service,
graph_service=graph_service,
wiki_client=wiki_client,
api_key="test"
)
assert result.status == "healthy"
assert result.orphan_vector_count == 0
assert result.orphan_entity_count == 0
assert result.vectors_without_graph == 0
assert result.docs_without_vectors == 0
async def test_health_degraded(self):
"""Test degraded status with orphans."""
vector_service = AsyncMock()
vector_service.get_all_chunk_references.return_value = [
{"chunk_id": f"c{i}", "page_id": 999, "doc_type": "wiki"}
for i in range(15)
]
# find_chunks_without_graph_nodes is not async
vector_service.find_chunks_without_graph_nodes = MagicMock(return_value=[])
graph_service = AsyncMock()
graph_service.find_orphan_entities.return_value = [
{"id": f"e{i}", "name": f"Entity{i}", "type": "Entity"}
for i in range(3)
]
graph_service.get_all_document_references.return_value = []
graph_service.find_documents_without_vectors.return_value = []
wiki_client = AsyncMock()
wiki_client.list_all_pages.return_value = []
result = await maintenance_health(
user="testuser",
vector_service=vector_service,
graph_service=graph_service,
wiki_client=wiki_client,
api_key="test"
)
assert result.status == "degraded"
assert result.orphan_vector_count == 15
assert result.orphan_entity_count == 3
assert len(result.recommendations) >= 1
@pytest.mark.asyncio
class TestReindexPage:
"""Test reindex page endpoint."""
async def test_reindex_success(self):
"""Test successful page reindex."""
vector_service = AsyncMock()
vector_service.delete_page_chunks.return_value = 5
vector_service.update_from_page.return_value = MagicMock(
success=True,
chunks_created=6,
error_message=None
)
graph_service = AsyncMock()
graph_service.delete_page.return_value = 1
graph_service.update_from_page.return_value = MagicMock(
success=True,
error_message=None
)
result = await reindex_page(
page_id=123,
user="testuser",
vector_service=vector_service,
graph_service=graph_service,
api_key="test"
)
assert result.success is True
assert result.page_id == 123
assert result.vectors_deleted == 5
assert result.vectors_created == 6
assert result.graph_updated is True
async def test_reindex_failure(self):
"""Test reindex with failure."""
vector_service = AsyncMock()
vector_service.delete_page_chunks.return_value = 0
vector_service.update_from_page.return_value = MagicMock(
success=False,
chunks_created=0,
error_message="Page not found"
)
graph_service = AsyncMock()
graph_service.delete_page.return_value = 0
graph_service.update_from_page.return_value = MagicMock(
success=False,
error_message="Page not found"
)
result = await reindex_page(
page_id=999,
user="testuser",
vector_service=vector_service,
graph_service=graph_service,
api_key="test"
)
assert result.success is False
assert result.error == "Page not found"
Executable
+50
View File
@@ -0,0 +1,50 @@
#!/bin/bash
# Library-Desk Server Startup Script
set -e
# Colors for output
GREEN='\033[0;32m'
YELLOW='\033[1;33m'
RED='\033[0;31m'
NC='\033[0m' # No Color
echo -e "${GREEN}Starting Library-Desk server...${NC}"
# Check if port 8778 is already in use
if lsof -Pi :8778 -sTCP:LISTEN -t >/dev/null 2>&1 ; then
echo -e "${RED}Error: Port 8778 is already in use${NC}"
echo "Run: lsof -i :8778 to see what's using it"
echo "Or run: kill \$(lsof -t -i:8778) to stop it"
exit 1
fi
# Activate virtual environment if not already activated
if [ -z "$VIRTUAL_ENV" ]; then
if [ -d ".venv" ]; then
echo -e "${YELLOW}Activating virtual environment...${NC}"
source .venv/bin/activate
else
echo -e "${RED}Error: Virtual environment not found${NC}"
echo "Run: python -m venv .venv && source .venv/bin/activate && pip install -r requirements.txt"
exit 1
fi
fi
# Create logs directory if it doesn't exist
LOGS_DIR="logs"
mkdir -p "$LOGS_DIR"
# Clear/create log file
LOG_FILE="$LOGS_DIR/server.log"
> "$LOG_FILE"
echo -e "${YELLOW}Logs will be written to: ${LOG_FILE}${NC}"
# Start the server
echo -e "${GREEN}Starting uvicorn server on http://tower-of-joy:8778${NC}"
echo -e "${YELLOW}Press Ctrl+C to stop the server${NC}"
echo ""
uvicorn src.main:app --reload --host 0.0.0.0 --port 8778 2>&1 | tee "$LOG_FILE"