perf: switch Qdrant to AsyncQdrantClient with explicit timeout
Every vector call ran on the sync QdrantClient inside async wrapper methods, blocking the FastAPI event loop per Qdrant round-trip. The wrapper now holds an AsyncQdrantClient (timeout via QDRANT_TIMEOUT, default 30s) and awaits all client calls; the wrapper API is unchanged. Call sites off the wrapper were fixed too: the HybridRAG document leg now uses the async search_vectors wrapper instead of the deprecated raw client.search (also fixing its call to the nonexistent ollama.embed_text which made the leg permanently report 'failed'), the health check awaits get_collections, and document_sync's raw delete/upsert calls are awaited (routed through wrappers in the next commit). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QbFZyDvYksazX6nYQYZ67L
This commit is contained in:
@@ -198,7 +198,7 @@ class DocumentSyncService:
|
||||
|
||||
# Delete existing chunks for this document
|
||||
try:
|
||||
self.qdrant.client.delete(
|
||||
await self.qdrant.client.delete(
|
||||
collection_name=collection,
|
||||
points_selector={
|
||||
"filter": {
|
||||
@@ -242,7 +242,7 @@ class DocumentSyncService:
|
||||
|
||||
# Upsert to Qdrant
|
||||
if points:
|
||||
self.qdrant.client.upsert(
|
||||
await self.qdrant.client.upsert(
|
||||
collection_name=collection,
|
||||
points=points
|
||||
)
|
||||
|
||||
@@ -485,34 +485,26 @@ JSON:"""
|
||||
return [], (time.time() - start) * 1000, None
|
||||
|
||||
# Get query embedding
|
||||
query_embedding = await self.vector.ollama.embed_text(query)
|
||||
query_embedding = await self.vector.ollama.embed(query)
|
||||
|
||||
# Search with filter for doc_type=document
|
||||
from qdrant_client.models import Filter, FieldCondition, MatchValue
|
||||
search_results = self.vector.qdrant.client.search(
|
||||
# Search via the async wrapper with doc_type=document filter
|
||||
search_results = await self.vector.qdrant.search_vectors(
|
||||
collection_name=collection_name,
|
||||
query_vector=query_embedding,
|
||||
limit=config.document_limit,
|
||||
score_threshold=config.document_threshold,
|
||||
query_filter=Filter(
|
||||
must=[
|
||||
FieldCondition(
|
||||
key="doc_type",
|
||||
match=MatchValue(value="document")
|
||||
)
|
||||
]
|
||||
)
|
||||
filter_conditions={"doc_type": "document"}
|
||||
)
|
||||
|
||||
# Format results
|
||||
formatted = []
|
||||
for r in search_results:
|
||||
payload = r.payload or {}
|
||||
payload = r.get("payload") or {}
|
||||
formatted.append({
|
||||
"paperless_id": payload.get("paperless_id"),
|
||||
"title": payload.get("title", "Untitled Document"),
|
||||
"content": payload.get("chunk_text", ""),
|
||||
"score": r.score,
|
||||
"score": r["score"],
|
||||
"correspondent": payload.get("correspondent"),
|
||||
"document_type": payload.get("document_type"),
|
||||
"tags": payload.get("tags", []),
|
||||
|
||||
Reference in New Issue
Block a user