feat: implement check-updates, job-backed ingest status, and dedup scan

Replace the four stub endpoints with real implementations, all requiring
an explicit tenant user (Phase B rule):

- /ingest/check-updates: GraphService now records a SHA-256 content_hash
  on every Document node at ingestion time; the endpoint compares those
  stored hashes against current Wiki.js page content in one UNWIND Cypher
  query per tenant and returns changed/new/deleted page lists (entity-stub
  pages excluded, pre-hash-tracking documents flagged stored_hash_missing).
- /ingest/status/{job_id}: backed by the Redis JobManager; jobs are
  tenant-scoped (foreign jobs 404). /ingest/page, /ingest/batch and
  /ingest/all now create job records and return job_id.
- /ingest/repo-status/{repository}: wiki page count vs indexed Document
  nodes under users/{tenant}/{repository} plus tenant job stats.
- /deduplicate/check: tenant-scoped Qdrant similarity scan; chunk pairs
  above ~0.9 cosine from different pages grouped per page pair with best
  score and page references (read-only).

Supporting changes: get_job_manager dependency (+ shutdown close),
scroll_all_points can return vectors, VectorService.find_duplicate_pairs,
src/core/hashing.compute_content_hash. 13 new offline unit tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbFZyDvYksazX6nYQYZ67L
This commit is contained in:
2026-07-14 12:06:21 +02:00
co-authored by Claude Fable 5
parent eae39aff3e
commit 86051d8022
10 changed files with 857 additions and 59 deletions
+26
View File
@@ -175,6 +175,21 @@ def get_redis_client() -> aioredis.Redis:
return client
@lru_cache
def get_job_manager() -> "JobManager":
"""
Get Redis-backed JobManager singleton.
Returns:
JobManager for background job tracking (connects lazily)
"""
from src.jobs.job_manager import JobManager
settings = get_settings()
manager = JobManager(redis_url=settings.redis_url)
logger.debug(f"Created JobManager: {settings.redis_url}")
return manager
@lru_cache
def get_content_extractor() -> ContentExtractor:
"""
@@ -368,6 +383,9 @@ PaperlessDep = Annotated[PaperlessClient, Depends(get_paperless_client)]
SettingsClientDep = Annotated[SettingsClient, Depends(get_settings_client)]
SchedulerDep = Annotated[SchedulerClient, Depends(get_scheduler_client)]
from src.jobs.job_manager import JobManager # noqa: E402
JobManagerDep = Annotated[JobManager, Depends(get_job_manager)]
# External API provider dependencies
WeatherProviderDep = Annotated[OpenMeteoProvider, Depends(get_weather_provider)]
NewsProviderDep = Annotated[AggregatedNewsProvider, Depends(get_news_provider)]
@@ -527,6 +545,14 @@ async def shutdown_clients():
except Exception as e:
logger.error(f"Error closing scheduler client: {e}")
# Close job manager Redis connection
try:
job_manager = get_job_manager()
await job_manager.close()
logger.info("✓ JobManager closed")
except Exception as e:
logger.error(f"Error closing JobManager: {e}")
logger.info("Service clients shutdown complete")
+26
View File
@@ -0,0 +1,26 @@
"""
Content hashing helpers.
A single canonical hash implementation is used everywhere page content is
fingerprinted (Document nodes at ingestion time, /ingest/check-updates
comparisons) so hashes computed at different times are comparable.
"""
import hashlib
def compute_content_hash(content: str) -> str:
"""
Compute the canonical content hash for wiki page content.
Args:
content: Raw page content (markdown). None-safe: treated as "".
Returns:
Hex-encoded SHA-256 digest of the UTF-8 encoded content.
Examples:
>>> compute_content_hash("hello")
'2cf24dba5fb0a30e26e83b2ac5b9e29e1b161e5c1fa7425e73043362938b9824'
"""
return hashlib.sha256((content or "").encode("utf-8")).hexdigest()