# Phase 2: Memory Systems Architecture **Status:** In Progress **Started:** 2025-11-13 **Phase Goal:** Persistent 3-tier conversation memory with automatic consolidation ## Overview The memory system provides persistent, intelligent conversation context using a three-tier architecture: 1. **Tier 1 (Working Memory):** Fast in-memory buffer for recent turns 2. **Tier 2 (Short-term):** SQLite database for summarized conversation history 3. **Tier 3 (Long-term):** Qdrant vector store for semantic search across all conversations ## Architecture Diagram ``` ┌─────────────────────────────────────────────────────────────┐ │ Chat Endpoint (/v1/chat/completions) │ │ │ │ 1. Accept user message │ │ 2. Retrieve relevant memory from all tiers │ │ 3. Build context: [Tier 1 + Tier 2 + Tier 3 semantic] │ │ 4. Generate response with Ollama │ │ 5. Store new turn in Tier 1 │ │ 6. Trigger consolidation if needed │ └─────────────────────────────────────────────────────────────┘ │ ▼ ┌─────────────────────────────────────────────────────────────┐ │ Memory Manager │ │ │ │ - Coordinates all 3 tiers │ │ - Handles memory retrieval │ │ - Triggers consolidation │ │ - Manages conversation sessions │ └─────────────────────────────────────────────────────────────┘ │ │ │ ▼ ▼ ▼ ┌──────────────────┐ ┌─────────────────┐ ┌──────────────────┐ │ Tier 1 │ │ Tier 2 │ │ Tier 3 │ │ Buffer Memory │ │ SQLite Summary │ │ Qdrant Vectors │ │ │ │ │ │ │ │ • In-memory dict │ │ • memory.db │ │ • conversation_ │ │ • Last 10 turns │ │ • Summaries │ │ memory │ │ • < 1ms access │ │ • ~10ms access │ │ • Semantic │ │ • Ephemeral │ │ • Persistent │ │ • ~50ms access │ │ • ~5KB RAM │ │ • ~500KB/100 │ │ • ~1KB per turn │ └──────────────────┘ └─────────────────┘ └──────────────────┘ │ │ │ └───────────────────┴────────────────────┘ │ ▼ ┌─────────────────────────────┐ │ Memory Consolidation │ │ Service │ │ │ │ Triggers: │ │ • Every 10 messages │ │ • Token limit (2000) │ │ • Conversation end │ │ • Explicit save command │ │ │ │ Actions: │ │ • Tier 1 → Tier 2 summary │ │ • Tier 2 → Tier 3 embed │ │ • Prune old Tier 1 data │ └─────────────────────────────┘ ``` ## Data Structures ### Tier 1: ConversationBufferMemory ```python { "conversation_id": "conv_123", "turns": [ { "role": "user", "content": "What is FastAPI?", "timestamp": "2025-11-13T10:00:00Z", "turn_number": 1 }, { "role": "assistant", "content": "FastAPI is a modern Python web framework...", "timestamp": "2025-11-13T10:00:02Z", "turn_number": 2, "tokens": {"prompt": 15, "completion": 120, "total": 135} } ], "metadata": { "created_at": "2025-11-13T10:00:00Z", "last_updated": "2025-11-13T10:00:02Z", "turn_count": 2, "total_tokens": 135 } } ``` ### Tier 2: SQLite Schema ```sql -- conversations table CREATE TABLE conversations ( id INTEGER PRIMARY KEY AUTOINCREMENT, conversation_id TEXT UNIQUE NOT NULL, user_id TEXT, created_at TIMESTAMP DEFAULT CURRENT_TIMESTAMP, last_message_at TIMESTAMP, turn_count INTEGER DEFAULT 0, total_tokens INTEGER DEFAULT 0, summary TEXT, status TEXT DEFAULT 'active' -- active, archived, deleted ); -- conversation_turns table CREATE TABLE conversation_turns ( id INTEGER PRIMARY KEY AUTOINCREMENT, conversation_id TEXT NOT NULL, turn_number INTEGER NOT NULL, role TEXT NOT NULL, -- user, assistant, system content TEXT NOT NULL, timestamp TIMESTAMP DEFAULT CURRENT_TIMESTAMP, tokens_prompt INTEGER, tokens_completion INTEGER, tokens_total INTEGER, FOREIGN KEY (conversation_id) REFERENCES conversations(conversation_id), UNIQUE(conversation_id, turn_number) ); -- conversation_summaries table (for Tier 2 condensed storage) CREATE TABLE conversation_summaries ( id INTEGER PRIMARY KEY AUTOINCREMENT, conversation_id TEXT NOT NULL, summary_text TEXT NOT NULL, turn_range_start INTEGER NOT NULL, turn_range_end INTEGER NOT NULL, created_at TIMESTAMP DEFAULT CURRENT_TIMESTAMP, token_count INTEGER, FOREIGN KEY (conversation_id) REFERENCES conversations(conversation_id) ); -- Indexes for performance CREATE INDEX idx_conversation_id ON conversation_turns(conversation_id); CREATE INDEX idx_timestamp ON conversation_turns(timestamp); CREATE INDEX idx_summary_conv ON conversation_summaries(conversation_id); ``` ### Tier 3: Qdrant Collection Schema ```python # Collection: conversation_memory { "collection_name": "conversation_memory", "vectors": { "size": 384, # all-MiniLM-L6-v2 embedding dimension "distance": "Cosine" }, "payload_schema": { "conversation_id": "string", "turn_number": "integer", "role": "string", "content": "text", "timestamp": "datetime", "tokens": "integer", "summary": "text", # Optional condensed version "tags": ["string"] # e.g., ["question", "code", "technical"] } } ``` ## Memory Retrieval Flow ### Query: "What did we discuss about FastAPI?" ```python # 1. Tier 1: Check recent buffer (last 10 turns) tier1_results = buffer_memory.get_recent_turns(limit=10) # Returns last 10 turns if they exist # 2. Tier 2: Check SQLite summaries tier2_results = sqlite_memory.search_summaries( conversation_id="conv_123", query="FastAPI discussion" ) # Returns summaries containing "FastAPI" # 3. Tier 3: Semantic search in Qdrant tier3_results = qdrant_memory.similarity_search( query="FastAPI discussion", limit=5, filter={"conversation_id": "conv_123"} ) # Returns 5 most semantically similar turns # 4. Merge and deduplicate context = merge_memory_results(tier1_results, tier2_results, tier3_results) # 5. Build prompt with context prompt = build_prompt_with_memory( system_message="You are a helpful assistant", memory_context=context, user_message="What did we discuss about FastAPI?" ) ``` ## Memory Consolidation Logic ### Trigger Conditions ```python class ConsolidationTrigger: MESSAGE_COUNT = 10 # Every 10 messages TOKEN_LIMIT = 2000 # When context > 2000 tokens CONVERSATION_END = True # End of conversation EXPLICIT_SAVE = True # User command: "remember this" TIME_ELAPSED = 3600 # 1 hour idle ``` ### Consolidation Process ```python async def consolidate_memory(conversation_id: str): """ Consolidate memory from Tier 1 → Tier 2 → Tier 3 """ # 1. Get Tier 1 buffer buffer = tier1_memory.get_buffer(conversation_id) if len(buffer.turns) >= 10: # 2. Summarize buffer using lightweight model summary = await summarize_conversation( turns=buffer.turns, model="gemma:7b" ) # 3. Store summary in Tier 2 (SQLite) tier2_memory.add_summary( conversation_id=conversation_id, summary=summary, turn_range=(buffer.turns[0].turn_number, buffer.turns[-1].turn_number) ) # 4. Embed individual turns to Tier 3 (Qdrant) for turn in buffer.turns: embedding = await embed_text(turn.content) tier3_memory.add_turn( conversation_id=conversation_id, turn=turn, embedding=embedding ) # 5. Prune Tier 1 buffer (keep only last 5 turns) tier1_memory.prune(conversation_id, keep_last=5) ``` ## File Structure ``` services/core-api/src/ ├── memory/ │ ├── __init__.py │ ├── base.py # Base memory classes │ ├── tier1_buffer.py # ConversationBufferMemory │ ├── tier2_sqlite.py # ConversationSummaryMemory │ ├── tier3_qdrant.py # VectorStoreRetrieverMemory │ ├── manager.py # MemoryManager (coordinates all tiers) │ ├── consolidation.py # Consolidation service │ └── schemas.py # Pydantic models ├── api/ │ └── v1/ │ ├── chat.py # Updated with memory integration │ ├── memory.py # NEW: Memory API endpoints │ └── schemas.py # Updated with memory schemas ├── models/ │ ├── ollama_client.py # Existing │ └── embeddings.py # NEW: Embedding model client └── utils/ └── database.py # NEW: SQLite utilities ``` ## API Endpoints (New) ### GET /v1/conversations List all conversations ### GET /v1/conversations/{conversation_id} Get conversation details and history ### GET /v1/conversations/{conversation_id}/turns Get all turns in a conversation ### POST /v1/conversations/{conversation_id}/search Semantic search within a conversation ### DELETE /v1/conversations/{conversation_id} Delete/archive a conversation ### POST /v1/conversations/{conversation_id}/consolidate Manually trigger memory consolidation ## Configuration Updates ```python # config.py additions class Settings(BaseSettings): # ... existing ... # Memory Configuration memory_tier1_max_turns: int = 10 memory_tier2_summary_threshold: int = 10 memory_tier3_enabled: bool = True # SQLite sqlite_database_path: str = "/app/data/memory.db" # Qdrant qdrant_host: str = "qdrant" qdrant_port: int = 6333 qdrant_collection_conversations: str = "conversation_memory" qdrant_collection_documents: str = "documents" qdrant_collection_user_facts: str = "user_facts" # Embeddings embedding_model: str = "sentence-transformers/all-MiniLM-L6-v2" embedding_dimension: int = 384 ``` ## Dependencies to Add ```txt # requirements.txt additions sqlalchemy==2.0.23 # SQLite ORM qdrant-client==1.7.0 # Qdrant Python client sentence-transformers==2.2.2 # Embedding models torch==2.1.0 # PyTorch (for embeddings) ``` ## Implementation Phases ### Phase 2.1: Tier 1 (Day 1) - ✅ Create base memory classes - ✅ Implement ConversationBufferMemory - ✅ Add basic memory schemas - ✅ Test in-memory storage and retrieval ### Phase 2.2: Tier 2 (Day 2) - ✅ Setup SQLite database - ✅ Create schema and migrations - ✅ Implement ConversationSummaryMemory - ✅ Add summarization using Ollama - ✅ Test persistence across restarts ### Phase 2.3: Tier 3 (Day 3) - ✅ Setup Qdrant collections - ✅ Implement embedding pipeline - ✅ Implement VectorStoreRetrieverMemory - ✅ Test semantic search - ✅ Test Qdrant connectivity ### Phase 2.4: Integration (Day 4) - ✅ Create MemoryManager - ✅ Implement consolidation service - ✅ Update /v1/chat/completions to use memory - ✅ Add memory API endpoints - ✅ Test end-to-end flow ### Phase 2.5: Testing & Polish (Day 5) - ✅ Comprehensive testing - ✅ Performance optimization - ✅ Memory leak checks - ✅ Documentation updates - ✅ Integration with Open WebUI ## Success Metrics - **Tier 1 Performance:** < 1ms access time - **Tier 2 Performance:** < 10ms query time - **Tier 3 Performance:** < 50ms semantic search - **Memory Persistence:** 100% across container restarts - **Context Relevance:** Semantic search returns appropriate results - **Memory Growth:** Bounded growth with automatic pruning - **Container Restart:** Conversations resume with full context ## Testing Plan 1. **Unit Tests:** - Each tier independently - Consolidation logic - Memory retrieval 2. **Integration Tests:** - Full memory flow - Container restart persistence - Multi-conversation handling 3. **Performance Tests:** - 100 conversations - 1000 turns total - Memory usage monitoring - Query performance benchmarks 4. **User Acceptance:** - Start conversation - Restart container - Resume conversation with context - Ask about past discussions - Verify relevant recall --- **Next Step:** Implement Tier 1 (ConversationBufferMemory)