Initial commit: library-desk service extraction from portainer-core
Build and Push / build (release) Successful in 36s

This commit is contained in:
2025-12-11 17:28:23 +01:00
commit 95852190ba
59 changed files with 17146 additions and 0 deletions
+473
View File
@@ -0,0 +1,473 @@
# Librarian Integration Guide
**How The Scheduler (Librarian) interacts with Library Desk for documentation management**
## Overview
The Scheduler's documentation mirroring tasks feed content into The Library system via Library Desk API. This creates a knowledge graph and vector index of all documentation for semantic search and relationship discovery.
## Architecture Flow
```
┌─────────────────┐
│ The Scheduler │ (The Librarian)
│ (scheduler) │
└────────┬────────┘
│ 1. Mirror docs from sources
│ (GitHub, Gitea, etc.)
┌─────────────────┐
│ Gitea Repo │
│ docs-mirror/* │
└────────┬────────┘
│ 2. Ingest to Library
┌─────────────────┐
│ Library Desk │ (Coordination API)
│ (library-desk) │
└────────┬────────┘
│ 3. Process & Index
├──→ Neo4j (relationships)
├──→ Qdrant (embeddings)
└──→ Wiki.js (dossiers)
```
## Required Endpoints
### 1. Document Ingestion
**POST /ingest/document**
```json
{
"source": "github",
"repository": "anthropics/anthropic-cookbook",
"path": "skills/citation/guide.md",
"content": "# Citation Guide\n...",
"metadata": {
"commit_sha": "abc123",
"author": "Anthropic",
"updated_at": "2025-12-08T10:30:00Z",
"gitea_mirror_path": "docs-mirror/anthropic-cookbook/skills/citation/guide.md",
"language": "markdown",
"tags": ["skills", "citation", "prompting"]
}
}
```
**Response:**
```json
{
"document_id": "doc_abc123xyz",
"status": "processing",
"operations": {
"chunking": "pending",
"embedding": "pending",
"entity_extraction": "pending",
"graph_indexing": "pending"
},
"estimated_completion": "2025-12-08T10:30:15Z"
}
```
### 2. Batch Ingestion
**POST /ingest/batch**
```json
{
"source": "github",
"repository": "anthropics/anthropic-cookbook",
"documents": [
{
"path": "skills/citation/guide.md",
"content": "...",
"metadata": {...}
},
{
"path": "skills/summarization/techniques.md",
"content": "...",
"metadata": {...}
}
]
}
```
**Response:**
```json
{
"batch_id": "batch_xyz789",
"total_documents": 2,
"status": "processing",
"documents": [
{"document_id": "doc_1", "status": "queued"},
{"document_id": "doc_2", "status": "queued"}
]
}
```
### 3. Document Status Check
**GET /ingest/status/{document_id}**
**Response:**
```json
{
"document_id": "doc_abc123xyz",
"status": "completed",
"operations": {
"chunking": "completed",
"embedding": "completed",
"entity_extraction": "completed",
"graph_indexing": "completed"
},
"results": {
"chunks_created": 12,
"vectors_indexed": 12,
"entities_extracted": 8,
"relationships_created": 15
},
"completed_at": "2025-12-08T10:30:14Z"
}
```
### 4. Content Update Detection
**POST /ingest/check-updates**
```json
{
"documents": [
{
"path": "docs-mirror/anthropic-cookbook/skills/citation/guide.md",
"content_hash": "sha256:abc123...",
"updated_at": "2025-12-08T10:30:00Z"
}
]
}
```
**Response:**
```json
{
"updates_needed": [
{
"path": "docs-mirror/anthropic-cookbook/skills/citation/guide.md",
"reason": "content_changed",
"last_indexed": "2025-12-07T10:30:00Z",
"action": "re-index"
}
],
"up_to_date": [],
"new_documents": []
}
```
### 5. Deduplication Check
**POST /deduplicate/check**
```json
{
"document_id": "doc_abc123xyz",
"similarity_threshold": 0.85
}
```
**Response:**
```json
{
"duplicates": [
{
"document_id": "doc_def456",
"similarity": 0.92,
"path": "docs-mirror/claude-docs/citation-best-practices.md",
"overlap_summary": "Both documents cover citation formatting"
}
],
"suggestions": {
"action": "merge_or_cross_reference",
"confidence": 0.88
}
}
```
### 6. Repository Sync Status
**GET /ingest/repo-status/{repository_name}**
**Response:**
```json
{
"repository": "anthropic-cookbook",
"total_documents": 156,
"indexed_documents": 156,
"pending_updates": 0,
"failed_documents": 0,
"last_sync": "2025-12-08T03:00:00Z",
"next_scheduled_sync": "2025-12-09T03:00:00Z"
}
```
## Scheduler Integration Workflow
### Phase 1: Mirror Documentation (Current)
```python
# This already exists in the Scheduler
async def mirror_documentation():
"""Mirror docs from external sources to Gitea"""
repos = [
"anthropics/anthropic-cookbook",
"anthropics/prompt-eng-interactive-tutorial",
# ... etc
]
for repo in repos:
# Clone/pull to /docs-mirror/{repo-name}
await git_sync(repo, f"/docs-mirror/{repo}")
```
### Phase 2: Index to Library (New)
```python
async def index_to_library():
"""Send mirrored docs to Library Desk for indexing"""
# Get list of documents in docs-mirror
docs_path = Path("/docs-mirror")
for repo_dir in docs_path.iterdir():
if not repo_dir.is_dir():
continue
# Find all markdown files
md_files = list(repo_dir.rglob("*.md"))
# Check what needs updating
update_check = await check_library_updates(md_files)
if update_check["updates_needed"]:
# Batch ingest updated documents
await batch_ingest_documents(
repository=repo_dir.name,
documents=update_check["updates_needed"]
)
# Wait for processing to complete
await wait_for_batch_completion(batch_id)
# Check for duplicates
await check_and_resolve_duplicates(repo_dir.name)
```
### Phase 3: Monitor & Maintain
```python
async def maintain_library_index():
"""Periodic maintenance of Library index"""
# Check for orphaned entries (deleted from source)
await cleanup_orphaned_documents()
# Update embeddings if model changed
await refresh_embeddings_if_needed()
# Generate relationship maps for new content
await discover_document_relationships()
```
## Scheduler Task Definition
**New Task: `library_sync`**
```yaml
Task Name: library_sync
Description: Sync mirrored documentation to Library for indexing and search
Schedule: Daily at 03:30 (after doc mirroring at 03:00)
Priority: 15 (user maintenance)
Service: library
Executor: scheduler.tasks.library_tasks.sync_library_index
Dependencies:
- docs_mirror (must complete first)
Configuration:
- LIBRARY_DESK_URL: http://library-desk:8089
- LIBRARY_API_KEY: ${LIBRARY_API_KEY}
- BATCH_SIZE: 50
- CHECK_UPDATES_ONLY: true
Outputs:
- Documents indexed
- Duplicates found
- Relationships created
```
## API Client Example
```python
# scheduler/src/clients/library_desk.py
import httpx
from typing import List, Dict, Any
from pathlib import Path
class LibraryDeskClient:
def __init__(self, base_url: str, api_key: str):
self.base_url = base_url
self.headers = {
"Authorization": f"Bearer {api_key}",
"Content-Type": "application/json"
}
async def ingest_document(
self,
source: str,
repository: str,
path: str,
content: str,
metadata: Dict[str, Any]
) -> Dict[str, Any]:
"""Ingest a single document"""
async with httpx.AsyncClient() as client:
response = await client.post(
f"{self.base_url}/ingest/document",
headers=self.headers,
json={
"source": source,
"repository": repository,
"path": path,
"content": content,
"metadata": metadata
}
)
response.raise_for_status()
return response.json()
async def batch_ingest(
self,
source: str,
repository: str,
documents: List[Dict[str, Any]]
) -> Dict[str, Any]:
"""Ingest multiple documents"""
async with httpx.AsyncClient() as client:
response = await client.post(
f"{self.base_url}/ingest/batch",
headers=self.headers,
json={
"source": source,
"repository": repository,
"documents": documents
},
timeout=300.0 # 5 minutes for large batches
)
response.raise_for_status()
return response.json()
async def check_updates(
self,
documents: List[Dict[str, str]]
) -> Dict[str, Any]:
"""Check which documents need updating"""
async with httpx.AsyncClient() as client:
response = await client.post(
f"{self.base_url}/ingest/check-updates",
headers=self.headers,
json={"documents": documents}
)
response.raise_for_status()
return response.json()
async def get_repo_status(self, repository: str) -> Dict[str, Any]:
"""Get indexing status for a repository"""
async with httpx.AsyncClient() as client:
response = await client.get(
f"{self.base_url}/ingest/repo-status/{repository}",
headers=self.headers
)
response.raise_for_status()
return response.json()
```
## Typical Workflow Sequence
### Daily Documentation Sync (03:00-03:45)
1. **03:00** - Scheduler runs `docs_mirror` task
- Pulls latest from all configured repos
- Writes to `/docs-mirror/*`
2. **03:30** - Scheduler runs `library_sync` task
- Scans `/docs-mirror/` for changes
- Calls `POST /ingest/check-updates` with file hashes
- Gets list of updated/new documents
3. **03:31-03:40** - Batch ingestion
- Groups documents by repo
- Calls `POST /ingest/batch` for each repo
- Monitors `GET /ingest/status/{batch_id}`
4. **03:41-03:44** - Post-processing
- Calls `POST /deduplicate/check` for new docs
- Reviews duplicate suggestions
- Logs statistics to Scheduler database
5. **03:45** - Completion
- Scheduler marks task complete
- Sends summary to logs
- Updates next run time
## Error Handling
### Retry Strategy
```python
async def ingest_with_retry(document: Dict, max_retries: int = 3):
"""Ingest with exponential backoff"""
for attempt in range(max_retries):
try:
result = await library_client.ingest_document(**document)
return result
except httpx.TimeoutException:
if attempt < max_retries - 1:
await asyncio.sleep(2 ** attempt) # 1s, 2s, 4s
else:
# Log failure and continue
logger.error(f"Failed to ingest {document['path']} after {max_retries} attempts")
return None
```
### Graceful Degradation
- If Library Desk is down, queue documents for later ingestion
- Store failed ingestions in Scheduler database
- Retry failed ingestions on next run
## Metrics to Track
The Scheduler should track:
- Documents mirrored vs. documents indexed
- Average ingestion time per document
- Deduplication rate
- Failed ingestions
- Library Desk response times
These can be displayed in the Scheduler UI dashboard.
## Environment Variables
Add to Scheduler's environment:
```bash
# Library Integration
LIBRARY_DESK_URL=http://library-desk:8089
LIBRARY_API_KEY=${LIBRARY_API_KEY}
LIBRARY_BATCH_SIZE=50
LIBRARY_SYNC_ENABLED=true
```
## Next Steps
1. Implement ingestion endpoints in Library Desk
2. Add LibraryDeskClient to Scheduler
3. Create `library_sync` task in Scheduler
4. Test with small batch of documents
5. Monitor and tune performance
6. Expand to full documentation corpus
## Benefits
- **Automatic Knowledge Base**: All mirrored docs automatically indexed
- **Semantic Search**: Find docs by meaning, not just keywords
- **Relationship Discovery**: Understand connections between docs
- **Deduplication**: Identify overlapping content across repos
- **HybridRAG Ready**: Knowledge graph + vectors enable advanced AI queries