ok... ok... I'll add it to git...
This commit is contained in:
@@ -0,0 +1,414 @@
|
||||
# Phase 2: Memory Systems Architecture
|
||||
|
||||
**Status:** In Progress
|
||||
**Started:** 2025-11-13
|
||||
**Phase Goal:** Persistent 3-tier conversation memory with automatic consolidation
|
||||
|
||||
## Overview
|
||||
|
||||
The memory system provides persistent, intelligent conversation context using a three-tier architecture:
|
||||
|
||||
1. **Tier 1 (Working Memory):** Fast in-memory buffer for recent turns
|
||||
2. **Tier 2 (Short-term):** SQLite database for summarized conversation history
|
||||
3. **Tier 3 (Long-term):** Qdrant vector store for semantic search across all conversations
|
||||
|
||||
## Architecture Diagram
|
||||
|
||||
```
|
||||
┌─────────────────────────────────────────────────────────────┐
|
||||
│ Chat Endpoint (/v1/chat/completions) │
|
||||
│ │
|
||||
│ 1. Accept user message │
|
||||
│ 2. Retrieve relevant memory from all tiers │
|
||||
│ 3. Build context: [Tier 1 + Tier 2 + Tier 3 semantic] │
|
||||
│ 4. Generate response with Ollama │
|
||||
│ 5. Store new turn in Tier 1 │
|
||||
│ 6. Trigger consolidation if needed │
|
||||
└─────────────────────────────────────────────────────────────┘
|
||||
│
|
||||
▼
|
||||
┌─────────────────────────────────────────────────────────────┐
|
||||
│ Memory Manager │
|
||||
│ │
|
||||
│ - Coordinates all 3 tiers │
|
||||
│ - Handles memory retrieval │
|
||||
│ - Triggers consolidation │
|
||||
│ - Manages conversation sessions │
|
||||
└─────────────────────────────────────────────────────────────┘
|
||||
│ │ │
|
||||
▼ ▼ ▼
|
||||
┌──────────────────┐ ┌─────────────────┐ ┌──────────────────┐
|
||||
│ Tier 1 │ │ Tier 2 │ │ Tier 3 │
|
||||
│ Buffer Memory │ │ SQLite Summary │ │ Qdrant Vectors │
|
||||
│ │ │ │ │ │
|
||||
│ • In-memory dict │ │ • memory.db │ │ • conversation_ │
|
||||
│ • Last 10 turns │ │ • Summaries │ │ memory │
|
||||
│ • < 1ms access │ │ • ~10ms access │ │ • Semantic │
|
||||
│ • Ephemeral │ │ • Persistent │ │ • ~50ms access │
|
||||
│ • ~5KB RAM │ │ • ~500KB/100 │ │ • ~1KB per turn │
|
||||
└──────────────────┘ └─────────────────┘ └──────────────────┘
|
||||
│ │ │
|
||||
└───────────────────┴────────────────────┘
|
||||
│
|
||||
▼
|
||||
┌─────────────────────────────┐
|
||||
│ Memory Consolidation │
|
||||
│ Service │
|
||||
│ │
|
||||
│ Triggers: │
|
||||
│ • Every 10 messages │
|
||||
│ • Token limit (2000) │
|
||||
│ • Conversation end │
|
||||
│ • Explicit save command │
|
||||
│ │
|
||||
│ Actions: │
|
||||
│ • Tier 1 → Tier 2 summary │
|
||||
│ • Tier 2 → Tier 3 embed │
|
||||
│ • Prune old Tier 1 data │
|
||||
└─────────────────────────────┘
|
||||
```
|
||||
|
||||
## Data Structures
|
||||
|
||||
### Tier 1: ConversationBufferMemory
|
||||
|
||||
```python
|
||||
{
|
||||
"conversation_id": "conv_123",
|
||||
"turns": [
|
||||
{
|
||||
"role": "user",
|
||||
"content": "What is FastAPI?",
|
||||
"timestamp": "2025-11-13T10:00:00Z",
|
||||
"turn_number": 1
|
||||
},
|
||||
{
|
||||
"role": "assistant",
|
||||
"content": "FastAPI is a modern Python web framework...",
|
||||
"timestamp": "2025-11-13T10:00:02Z",
|
||||
"turn_number": 2,
|
||||
"tokens": {"prompt": 15, "completion": 120, "total": 135}
|
||||
}
|
||||
],
|
||||
"metadata": {
|
||||
"created_at": "2025-11-13T10:00:00Z",
|
||||
"last_updated": "2025-11-13T10:00:02Z",
|
||||
"turn_count": 2,
|
||||
"total_tokens": 135
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### Tier 2: SQLite Schema
|
||||
|
||||
```sql
|
||||
-- conversations table
|
||||
CREATE TABLE conversations (
|
||||
id INTEGER PRIMARY KEY AUTOINCREMENT,
|
||||
conversation_id TEXT UNIQUE NOT NULL,
|
||||
user_id TEXT,
|
||||
created_at TIMESTAMP DEFAULT CURRENT_TIMESTAMP,
|
||||
last_message_at TIMESTAMP,
|
||||
turn_count INTEGER DEFAULT 0,
|
||||
total_tokens INTEGER DEFAULT 0,
|
||||
summary TEXT,
|
||||
status TEXT DEFAULT 'active' -- active, archived, deleted
|
||||
);
|
||||
|
||||
-- conversation_turns table
|
||||
CREATE TABLE conversation_turns (
|
||||
id INTEGER PRIMARY KEY AUTOINCREMENT,
|
||||
conversation_id TEXT NOT NULL,
|
||||
turn_number INTEGER NOT NULL,
|
||||
role TEXT NOT NULL, -- user, assistant, system
|
||||
content TEXT NOT NULL,
|
||||
timestamp TIMESTAMP DEFAULT CURRENT_TIMESTAMP,
|
||||
tokens_prompt INTEGER,
|
||||
tokens_completion INTEGER,
|
||||
tokens_total INTEGER,
|
||||
FOREIGN KEY (conversation_id) REFERENCES conversations(conversation_id),
|
||||
UNIQUE(conversation_id, turn_number)
|
||||
);
|
||||
|
||||
-- conversation_summaries table (for Tier 2 condensed storage)
|
||||
CREATE TABLE conversation_summaries (
|
||||
id INTEGER PRIMARY KEY AUTOINCREMENT,
|
||||
conversation_id TEXT NOT NULL,
|
||||
summary_text TEXT NOT NULL,
|
||||
turn_range_start INTEGER NOT NULL,
|
||||
turn_range_end INTEGER NOT NULL,
|
||||
created_at TIMESTAMP DEFAULT CURRENT_TIMESTAMP,
|
||||
token_count INTEGER,
|
||||
FOREIGN KEY (conversation_id) REFERENCES conversations(conversation_id)
|
||||
);
|
||||
|
||||
-- Indexes for performance
|
||||
CREATE INDEX idx_conversation_id ON conversation_turns(conversation_id);
|
||||
CREATE INDEX idx_timestamp ON conversation_turns(timestamp);
|
||||
CREATE INDEX idx_summary_conv ON conversation_summaries(conversation_id);
|
||||
```
|
||||
|
||||
### Tier 3: Qdrant Collection Schema
|
||||
|
||||
```python
|
||||
# Collection: conversation_memory
|
||||
{
|
||||
"collection_name": "conversation_memory",
|
||||
"vectors": {
|
||||
"size": 384, # all-MiniLM-L6-v2 embedding dimension
|
||||
"distance": "Cosine"
|
||||
},
|
||||
"payload_schema": {
|
||||
"conversation_id": "string",
|
||||
"turn_number": "integer",
|
||||
"role": "string",
|
||||
"content": "text",
|
||||
"timestamp": "datetime",
|
||||
"tokens": "integer",
|
||||
"summary": "text", # Optional condensed version
|
||||
"tags": ["string"] # e.g., ["question", "code", "technical"]
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## Memory Retrieval Flow
|
||||
|
||||
### Query: "What did we discuss about FastAPI?"
|
||||
|
||||
```python
|
||||
# 1. Tier 1: Check recent buffer (last 10 turns)
|
||||
tier1_results = buffer_memory.get_recent_turns(limit=10)
|
||||
# Returns last 10 turns if they exist
|
||||
|
||||
# 2. Tier 2: Check SQLite summaries
|
||||
tier2_results = sqlite_memory.search_summaries(
|
||||
conversation_id="conv_123",
|
||||
query="FastAPI discussion"
|
||||
)
|
||||
# Returns summaries containing "FastAPI"
|
||||
|
||||
# 3. Tier 3: Semantic search in Qdrant
|
||||
tier3_results = qdrant_memory.similarity_search(
|
||||
query="FastAPI discussion",
|
||||
limit=5,
|
||||
filter={"conversation_id": "conv_123"}
|
||||
)
|
||||
# Returns 5 most semantically similar turns
|
||||
|
||||
# 4. Merge and deduplicate
|
||||
context = merge_memory_results(tier1_results, tier2_results, tier3_results)
|
||||
|
||||
# 5. Build prompt with context
|
||||
prompt = build_prompt_with_memory(
|
||||
system_message="You are a helpful assistant",
|
||||
memory_context=context,
|
||||
user_message="What did we discuss about FastAPI?"
|
||||
)
|
||||
```
|
||||
|
||||
## Memory Consolidation Logic
|
||||
|
||||
### Trigger Conditions
|
||||
|
||||
```python
|
||||
class ConsolidationTrigger:
|
||||
MESSAGE_COUNT = 10 # Every 10 messages
|
||||
TOKEN_LIMIT = 2000 # When context > 2000 tokens
|
||||
CONVERSATION_END = True # End of conversation
|
||||
EXPLICIT_SAVE = True # User command: "remember this"
|
||||
TIME_ELAPSED = 3600 # 1 hour idle
|
||||
```
|
||||
|
||||
### Consolidation Process
|
||||
|
||||
```python
|
||||
async def consolidate_memory(conversation_id: str):
|
||||
"""
|
||||
Consolidate memory from Tier 1 → Tier 2 → Tier 3
|
||||
"""
|
||||
# 1. Get Tier 1 buffer
|
||||
buffer = tier1_memory.get_buffer(conversation_id)
|
||||
|
||||
if len(buffer.turns) >= 10:
|
||||
# 2. Summarize buffer using lightweight model
|
||||
summary = await summarize_conversation(
|
||||
turns=buffer.turns,
|
||||
model="gemma:7b"
|
||||
)
|
||||
|
||||
# 3. Store summary in Tier 2 (SQLite)
|
||||
tier2_memory.add_summary(
|
||||
conversation_id=conversation_id,
|
||||
summary=summary,
|
||||
turn_range=(buffer.turns[0].turn_number, buffer.turns[-1].turn_number)
|
||||
)
|
||||
|
||||
# 4. Embed individual turns to Tier 3 (Qdrant)
|
||||
for turn in buffer.turns:
|
||||
embedding = await embed_text(turn.content)
|
||||
tier3_memory.add_turn(
|
||||
conversation_id=conversation_id,
|
||||
turn=turn,
|
||||
embedding=embedding
|
||||
)
|
||||
|
||||
# 5. Prune Tier 1 buffer (keep only last 5 turns)
|
||||
tier1_memory.prune(conversation_id, keep_last=5)
|
||||
```
|
||||
|
||||
## File Structure
|
||||
|
||||
```
|
||||
services/core-api/src/
|
||||
├── memory/
|
||||
│ ├── __init__.py
|
||||
│ ├── base.py # Base memory classes
|
||||
│ ├── tier1_buffer.py # ConversationBufferMemory
|
||||
│ ├── tier2_sqlite.py # ConversationSummaryMemory
|
||||
│ ├── tier3_qdrant.py # VectorStoreRetrieverMemory
|
||||
│ ├── manager.py # MemoryManager (coordinates all tiers)
|
||||
│ ├── consolidation.py # Consolidation service
|
||||
│ └── schemas.py # Pydantic models
|
||||
├── api/
|
||||
│ └── v1/
|
||||
│ ├── chat.py # Updated with memory integration
|
||||
│ ├── memory.py # NEW: Memory API endpoints
|
||||
│ └── schemas.py # Updated with memory schemas
|
||||
├── models/
|
||||
│ ├── ollama_client.py # Existing
|
||||
│ └── embeddings.py # NEW: Embedding model client
|
||||
└── utils/
|
||||
└── database.py # NEW: SQLite utilities
|
||||
```
|
||||
|
||||
## API Endpoints (New)
|
||||
|
||||
### GET /v1/conversations
|
||||
List all conversations
|
||||
|
||||
### GET /v1/conversations/{conversation_id}
|
||||
Get conversation details and history
|
||||
|
||||
### GET /v1/conversations/{conversation_id}/turns
|
||||
Get all turns in a conversation
|
||||
|
||||
### POST /v1/conversations/{conversation_id}/search
|
||||
Semantic search within a conversation
|
||||
|
||||
### DELETE /v1/conversations/{conversation_id}
|
||||
Delete/archive a conversation
|
||||
|
||||
### POST /v1/conversations/{conversation_id}/consolidate
|
||||
Manually trigger memory consolidation
|
||||
|
||||
## Configuration Updates
|
||||
|
||||
```python
|
||||
# config.py additions
|
||||
class Settings(BaseSettings):
|
||||
# ... existing ...
|
||||
|
||||
# Memory Configuration
|
||||
memory_tier1_max_turns: int = 10
|
||||
memory_tier2_summary_threshold: int = 10
|
||||
memory_tier3_enabled: bool = True
|
||||
|
||||
# SQLite
|
||||
sqlite_database_path: str = "/app/data/memory.db"
|
||||
|
||||
# Qdrant
|
||||
qdrant_host: str = "qdrant"
|
||||
qdrant_port: int = 6333
|
||||
qdrant_collection_conversations: str = "conversation_memory"
|
||||
qdrant_collection_documents: str = "documents"
|
||||
qdrant_collection_user_facts: str = "user_facts"
|
||||
|
||||
# Embeddings
|
||||
embedding_model: str = "sentence-transformers/all-MiniLM-L6-v2"
|
||||
embedding_dimension: int = 384
|
||||
```
|
||||
|
||||
## Dependencies to Add
|
||||
|
||||
```txt
|
||||
# requirements.txt additions
|
||||
sqlalchemy==2.0.23 # SQLite ORM
|
||||
qdrant-client==1.7.0 # Qdrant Python client
|
||||
sentence-transformers==2.2.2 # Embedding models
|
||||
torch==2.1.0 # PyTorch (for embeddings)
|
||||
```
|
||||
|
||||
## Implementation Phases
|
||||
|
||||
### Phase 2.1: Tier 1 (Day 1)
|
||||
- ✅ Create base memory classes
|
||||
- ✅ Implement ConversationBufferMemory
|
||||
- ✅ Add basic memory schemas
|
||||
- ✅ Test in-memory storage and retrieval
|
||||
|
||||
### Phase 2.2: Tier 2 (Day 2)
|
||||
- ✅ Setup SQLite database
|
||||
- ✅ Create schema and migrations
|
||||
- ✅ Implement ConversationSummaryMemory
|
||||
- ✅ Add summarization using Ollama
|
||||
- ✅ Test persistence across restarts
|
||||
|
||||
### Phase 2.3: Tier 3 (Day 3)
|
||||
- ✅ Setup Qdrant collections
|
||||
- ✅ Implement embedding pipeline
|
||||
- ✅ Implement VectorStoreRetrieverMemory
|
||||
- ✅ Test semantic search
|
||||
- ✅ Test Qdrant connectivity
|
||||
|
||||
### Phase 2.4: Integration (Day 4)
|
||||
- ✅ Create MemoryManager
|
||||
- ✅ Implement consolidation service
|
||||
- ✅ Update /v1/chat/completions to use memory
|
||||
- ✅ Add memory API endpoints
|
||||
- ✅ Test end-to-end flow
|
||||
|
||||
### Phase 2.5: Testing & Polish (Day 5)
|
||||
- ✅ Comprehensive testing
|
||||
- ✅ Performance optimization
|
||||
- ✅ Memory leak checks
|
||||
- ✅ Documentation updates
|
||||
- ✅ Integration with Open WebUI
|
||||
|
||||
## Success Metrics
|
||||
|
||||
- **Tier 1 Performance:** < 1ms access time
|
||||
- **Tier 2 Performance:** < 10ms query time
|
||||
- **Tier 3 Performance:** < 50ms semantic search
|
||||
- **Memory Persistence:** 100% across container restarts
|
||||
- **Context Relevance:** Semantic search returns appropriate results
|
||||
- **Memory Growth:** Bounded growth with automatic pruning
|
||||
- **Container Restart:** Conversations resume with full context
|
||||
|
||||
## Testing Plan
|
||||
|
||||
1. **Unit Tests:**
|
||||
- Each tier independently
|
||||
- Consolidation logic
|
||||
- Memory retrieval
|
||||
|
||||
2. **Integration Tests:**
|
||||
- Full memory flow
|
||||
- Container restart persistence
|
||||
- Multi-conversation handling
|
||||
|
||||
3. **Performance Tests:**
|
||||
- 100 conversations
|
||||
- 1000 turns total
|
||||
- Memory usage monitoring
|
||||
- Query performance benchmarks
|
||||
|
||||
4. **User Acceptance:**
|
||||
- Start conversation
|
||||
- Restart container
|
||||
- Resume conversation with context
|
||||
- Ask about past discussions
|
||||
- Verify relevant recall
|
||||
|
||||
---
|
||||
|
||||
**Next Step:** Implement Tier 1 (ConversationBufferMemory)
|
||||
Reference in New Issue
Block a user