feat(ai): complete Phase 2/3 documentation and memory system improvements
Phase completion and enhancement updates: ## Documentation Added - Phase 2 completion: Memory system implementation details - Phase 3 completion: Research capabilities and tool integration - Session documentation: Model testing, VRAM optimization analysis - Test results: Comprehensive prompt testing (v1_verbose: 87/100) - Tool logging implementation guide ## System Prompts - Added prompts.py with 7 tested variants for A/B testing - v1_verbose, v2_concise, v3_imperative, v4_minimal, etc. - Comprehensive testing results for each variant - Production-ready prompt selection guidance ## Memory System Enhancements - Multi-tenancy support: Added user_id parameter throughout - System message filtering: Don't store system messages in history - Improved conversation turn tracking with user isolation - Enhanced memory manager for better multi-user support ## AI Controller Improvements - Better memory integration with user_id support - Enhanced error handling for memory operations - Improved token tracking for usage monitoring - Skip system message storage (part of agent state) ## Portainer Client - Comprehensive API client (148 lines) - Stack management and service monitoring - Container operations with full error handling - Async support for all operations ## Architecture Documentation - Updated agent flow diagrams for ADK architecture - Enhanced core-api README with current setup - Updated Docker compose stack configuration - Complete testing and validation documentation
This commit is contained in:
@@ -0,0 +1,434 @@
|
||||
# Phase 3: Multi-Agent Workflows Implementation Plan
|
||||
|
||||
**Date**: 2025-11-24
|
||||
**Status**: 🎯 Ready to Start
|
||||
**Duration**: 1 week
|
||||
**Prerequisites**: ✅ Phase 1 Complete, ✅ Phase 2 Complete
|
||||
|
||||
## Overview
|
||||
|
||||
Implement LangGraph-based multi-agent system with intelligent routing. The current unified agent (mistral:7b) will become the orchestrator/router, delegating to specialist agents for complex tasks.
|
||||
|
||||
## Current State
|
||||
|
||||
**What We Have** ✅:
|
||||
- Unified agent with tool calling (mistral:7b)
|
||||
- Basic orchestration (LangGraph ReAct agent)
|
||||
- 7 working tools (list_services, get_service_details, list_domains, etc.)
|
||||
- Streaming responses with proper formatting
|
||||
- Memory system (Tier 1-3 with Qdrant)
|
||||
- OpenAI-compatible API
|
||||
|
||||
**Current Architecture**:
|
||||
```
|
||||
User → Core API → Unified Agent (mistral:7b) → Tools
|
||||
↓
|
||||
Memory (Buffer + Qdrant)
|
||||
```
|
||||
|
||||
## Target Architecture
|
||||
|
||||
```
|
||||
User → Core API → Router Agent (mistral:7b)
|
||||
↓
|
||||
┌─────────┴──────────┐
|
||||
▼ ▼
|
||||
Chat Agent Research Agent
|
||||
(mistral:7b) (mistral:7b + tools)
|
||||
│ │
|
||||
▼ ▼
|
||||
Memory System Web Search/Scraping
|
||||
```
|
||||
|
||||
**Future Expansion** (Phase 4+):
|
||||
```
|
||||
Router Agent
|
||||
├── Chat Agent (general conversation)
|
||||
├── Research Agent (web search + synthesis)
|
||||
├── Code Agent (codestral for programming)
|
||||
└── Tool Agent (infrastructure actions)
|
||||
```
|
||||
|
||||
## Implementation Strategy
|
||||
|
||||
### Option A: Extend Current Unified Agent (Recommended)
|
||||
|
||||
**Pros**:
|
||||
- ✅ Builds on working foundation
|
||||
- ✅ Minimal disruption
|
||||
- ✅ Can migrate gradually
|
||||
- ✅ Reuses existing streaming, memory, tools
|
||||
|
||||
**Cons**:
|
||||
- ⚠️ Slightly less separation than multi-agent
|
||||
- ⚠️ All in one orchestrator file
|
||||
|
||||
**Approach**: Add routing logic to existing unified agent to detect complex tasks and create sub-workflows.
|
||||
|
||||
### Option B: Full LangGraph Multi-Agent Rewrite
|
||||
|
||||
**Pros**:
|
||||
- ✅ Clean separation of agents
|
||||
- ✅ True multi-agent pattern
|
||||
- ✅ Easier to add new agents later
|
||||
|
||||
**Cons**:
|
||||
- ❌ Major rewrite
|
||||
- ❌ Risk breaking existing functionality
|
||||
- ❌ Complex state management
|
||||
- ❌ Harder to debug
|
||||
|
||||
**Approach**: Create separate agent modules, supervisor pattern, state graph.
|
||||
|
||||
**Decision**: **Use Option A** - Extend current unified agent with routing intelligence.
|
||||
|
||||
## Phase 3 Goals
|
||||
|
||||
### Core Goals
|
||||
1. **Intelligent Task Detection**: Automatically identify when a task needs research vs simple chat
|
||||
2. **Research Workflow**: Multi-step web search → scraping → synthesis for complex queries
|
||||
3. **Proper Context Passing**: Pass memory context to sub-workflows
|
||||
4. **Streaming Updates**: Show progress during multi-step research
|
||||
|
||||
### Non-Goals (Deferred to Phase 4)
|
||||
- ❌ Code specialist agent (codestral)
|
||||
- ❌ Tool executor agent (separate from router)
|
||||
- ❌ Model switching based on complexity
|
||||
- ❌ Supervisor pattern for agent coordination
|
||||
|
||||
## Implementation Tasks
|
||||
|
||||
### Task 1: Add Research Detection
|
||||
|
||||
**File**: `services/core-api/src/agent/orchestrator.py`
|
||||
|
||||
**Changes**:
|
||||
- Add system prompt instructions for research detection
|
||||
- Detect keywords: "research", "find information about", "look up", "what's the latest"
|
||||
- Detect follow-up tool usage patterns (web_search → web_scrape)
|
||||
|
||||
**Pseudo-code**:
|
||||
```python
|
||||
SYSTEM_PROMPT = """
|
||||
...existing prompt...
|
||||
|
||||
## Research Mode
|
||||
When user asks for current information, recent news, or complex topics requiring web search:
|
||||
1. Use web_search tool to find relevant sources
|
||||
2. Use web_scrape tool (via web_search) to extract content
|
||||
3. Synthesize information from multiple sources
|
||||
4. Cite sources in your response
|
||||
|
||||
Examples of research queries:
|
||||
- "What's the latest on [topic]?"
|
||||
- "Research [topic] for me"
|
||||
- "Find information about [topic]"
|
||||
- "What are people saying about [topic]?"
|
||||
"""
|
||||
```
|
||||
|
||||
**Success Criteria**:
|
||||
- Agent detects research queries correctly (>80% accuracy)
|
||||
- Automatically triggers web_search when needed
|
||||
- Follows up with synthesis
|
||||
|
||||
### Task 2: Improve Web Search Tool
|
||||
|
||||
**File**: `services/core-api/src/agent/tools.py`
|
||||
|
||||
**Current State**: We have `web_search` tool that fetches and extracts content from a URL.
|
||||
|
||||
**Enhancements Needed**:
|
||||
1. Add actual search capability (DuckDuckGo API)
|
||||
2. Return multiple results (not just one URL)
|
||||
3. Add trafilatura for better content extraction
|
||||
|
||||
**New Implementation**:
|
||||
```python
|
||||
@tool
|
||||
async def web_search(query: str, num_results: int = 3) -> str:
|
||||
"""
|
||||
Search the web using DuckDuckGo and extract content from top results.
|
||||
|
||||
Args:
|
||||
query: Search query
|
||||
num_results: Number of results to return (default 3)
|
||||
|
||||
Returns:
|
||||
Formatted results with titles, URLs, and content summaries
|
||||
"""
|
||||
from duckduckgo_search import DDGS
|
||||
|
||||
results = []
|
||||
with DDGS() as ddgs:
|
||||
search_results = list(ddgs.text(query, max_results=num_results))
|
||||
|
||||
for result in search_results:
|
||||
# Scrape each result
|
||||
content = await scrape_url(result['href'])
|
||||
results.append({
|
||||
'title': result['title'],
|
||||
'url': result['href'],
|
||||
'snippet': result['body'],
|
||||
'content': content[:500] # First 500 chars
|
||||
})
|
||||
|
||||
return format_search_results(results)
|
||||
```
|
||||
|
||||
**Dependencies**: Add to `requirements.txt`:
|
||||
```
|
||||
duckduckgo-search==4.1.1
|
||||
```
|
||||
|
||||
**Success Criteria**:
|
||||
- Returns 3+ search results
|
||||
- Each result has title, URL, snippet
|
||||
- Content extraction works for most sites
|
||||
|
||||
### Task 3: Add Research Workflow Pattern
|
||||
|
||||
**File**: `services/core-api/src/agent/orchestrator.py`
|
||||
|
||||
**Pattern**: Multi-step tool usage
|
||||
```
|
||||
1. User: "Research AI agent frameworks"
|
||||
2. Agent: [Thinking] This needs research...
|
||||
3. Agent: [Tool Call] web_search("AI agent frameworks 2025")
|
||||
4. Tool: Returns 3 results with content
|
||||
5. Agent: [Synthesizing] Based on search results...
|
||||
6. Agent: [Response] Here's what I found: ...
|
||||
```
|
||||
|
||||
**Implementation**: Already handled by LangGraph ReAct agent! Just need better tools.
|
||||
|
||||
**Success Criteria**:
|
||||
- Agent chains tool calls naturally
|
||||
- Synthesizes information from multiple sources
|
||||
- Cites sources in response
|
||||
|
||||
### Task 4: Add Progress Indicators for Research
|
||||
|
||||
**File**: `services/core-api/src/agent/streaming.py`
|
||||
|
||||
**Enhancement**: Add more granular status updates
|
||||
|
||||
**Current**:
|
||||
```python
|
||||
"[🔧 Using web_search...]"
|
||||
```
|
||||
|
||||
**Enhanced**:
|
||||
```python
|
||||
"[🔍 Searching web for: {query}...]"
|
||||
"[📄 Reading result 1/3...]"
|
||||
"[📄 Reading result 2/3...]"
|
||||
"[🧠 Synthesizing information...]"
|
||||
"[✓ Research complete]"
|
||||
```
|
||||
|
||||
**Implementation**: Enhance tool_call streaming messages
|
||||
|
||||
**Success Criteria**:
|
||||
- User sees progress during research
|
||||
- Clear indication of what's happening
|
||||
- Doesn't spam with too many updates
|
||||
|
||||
### Task 5: Test Research Workflows
|
||||
|
||||
**Test Queries**:
|
||||
1. "What's the latest news about AI?"
|
||||
2. "Research LangGraph vs CrewAI"
|
||||
3. "Find information about Mistral AI models"
|
||||
4. "What are people saying about Open WebUI?"
|
||||
5. "Look up Qdrant vector database features"
|
||||
|
||||
**Success Criteria**:
|
||||
- Agent uses web_search automatically
|
||||
- Returns multi-source synthesis
|
||||
- Cites URLs in response
|
||||
- Completes in <10 seconds
|
||||
|
||||
### Task 6: Add Research History to Memory
|
||||
|
||||
**File**: `services/core-api/src/memory/manager.py`
|
||||
|
||||
**Enhancement**: Tag research results in memory
|
||||
|
||||
**Schema Addition**:
|
||||
```python
|
||||
metadata = {
|
||||
"type": "research",
|
||||
"sources": ["url1", "url2", "url3"],
|
||||
"query": "original search query"
|
||||
}
|
||||
```
|
||||
|
||||
**Success Criteria**:
|
||||
- Research results stored in memory
|
||||
- Can recall previous research
|
||||
- Sources preserved for future reference
|
||||
|
||||
## Testing Plan
|
||||
|
||||
### Unit Tests
|
||||
```python
|
||||
# Test research detection
|
||||
def test_research_detection():
|
||||
queries = [
|
||||
("What's the weather?", False), # Not research
|
||||
("Research AI frameworks", True), # Is research
|
||||
("Find info about Kubernetes", True), # Is research
|
||||
]
|
||||
for query, expected in queries:
|
||||
assert is_research_query(query) == expected
|
||||
|
||||
# Test web search tool
|
||||
@pytest.mark.asyncio
|
||||
async def test_web_search():
|
||||
results = await web_search("LangGraph")
|
||||
assert len(results) >= 1
|
||||
assert "url" in results[0]
|
||||
assert "content" in results[0]
|
||||
```
|
||||
|
||||
### Integration Tests
|
||||
```python
|
||||
# Test research workflow
|
||||
@pytest.mark.asyncio
|
||||
async def test_research_workflow():
|
||||
agent = get_unified_agent()
|
||||
response = await agent.chat(
|
||||
"Research LangGraph for me",
|
||||
stream=False
|
||||
)
|
||||
|
||||
# Should have used web_search
|
||||
# Should have synthesized results
|
||||
# Should cite sources
|
||||
assert "http" in response # Has URLs
|
||||
assert len(response) > 200 # Detailed response
|
||||
```
|
||||
|
||||
### Manual Tests
|
||||
1. Ask research query in Open WebUI
|
||||
2. Verify agent searches web
|
||||
3. Verify progress indicators appear
|
||||
4. Verify synthesized response with sources
|
||||
5. Verify research saved to memory
|
||||
|
||||
## Dependencies
|
||||
|
||||
**New packages** needed:
|
||||
```
|
||||
# requirements.txt additions
|
||||
duckduckgo-search==4.1.1 # Web search
|
||||
```
|
||||
|
||||
**Existing packages** (already installed):
|
||||
```
|
||||
httpx==0.28.1 # HTTP client
|
||||
beautifulsoup4==4.12.3 # HTML parsing
|
||||
trafilatura==1.12.2 # Content extraction
|
||||
```
|
||||
|
||||
## Migration Plan
|
||||
|
||||
### Step 1: Add Dependencies
|
||||
```bash
|
||||
# Add to requirements.txt
|
||||
echo "duckduckgo-search==4.1.1" >> services/core-api/requirements.txt
|
||||
|
||||
# Rebuild container
|
||||
docker-compose -f stacks/core-api.yml build
|
||||
docker-compose -f stacks/core-api.yml up -d
|
||||
```
|
||||
|
||||
### Step 2: Implement Web Search Tool
|
||||
- Update `tools.py` with DuckDuckGo integration
|
||||
- Test independently
|
||||
- Add to agent's tool list (already automatic)
|
||||
|
||||
### Step 3: Update System Prompt
|
||||
- Add research detection instructions
|
||||
- Test with various queries
|
||||
- Tune detection accuracy
|
||||
|
||||
### Step 4: Enhance Streaming
|
||||
- Add research progress indicators
|
||||
- Test in Open WebUI
|
||||
- Ensure doesn't break existing functionality
|
||||
|
||||
### Step 5: Integration Testing
|
||||
- Test research workflows end-to-end
|
||||
- Verify memory storage
|
||||
- Verify source citations
|
||||
|
||||
### Step 6: User Acceptance
|
||||
- Ask user to test in Open WebUI
|
||||
- Gather feedback
|
||||
- Iterate on improvements
|
||||
|
||||
## Success Metrics
|
||||
|
||||
### Quantitative
|
||||
- **Research Detection Accuracy**: >80% (detects research queries correctly)
|
||||
- **Tool Chain Success**: >90% (completes multi-step research)
|
||||
- **Response Time**: <10s (average research query)
|
||||
- **Source Citations**: >90% (includes URLs in response)
|
||||
|
||||
### Qualitative
|
||||
- User feels agent is more capable
|
||||
- Research responses are comprehensive
|
||||
- Sources are relevant and recent
|
||||
- Progress indicators are helpful
|
||||
|
||||
## Risks & Mitigation
|
||||
|
||||
### Risk 1: Web Search Too Slow
|
||||
**Impact**: User experience degraded
|
||||
**Mitigation**:
|
||||
- Limit to 3 results max
|
||||
- Run scraping in parallel
|
||||
- Add timeout (10s)
|
||||
- Show progress to user
|
||||
|
||||
### Risk 2: Search Results Low Quality
|
||||
**Impact**: Agent gives poor answers
|
||||
**Mitigation**:
|
||||
- Use multiple search engines if needed
|
||||
- Implement result filtering
|
||||
- Let agent decide relevance
|
||||
- Allow user to refine query
|
||||
|
||||
### Risk 3: Breaking Existing Functionality
|
||||
**Impact**: Simple chat stops working
|
||||
**Mitigation**:
|
||||
- Test simple queries extensively
|
||||
- Keep research optional (agent decides)
|
||||
- Easy rollback (git revert)
|
||||
- Gradual deployment
|
||||
|
||||
## Phase 3 Completion Criteria
|
||||
|
||||
✅ **Phase 3 Complete** when:
|
||||
1. Web search tool integrated (DuckDuckGo)
|
||||
2. Agent detects research queries automatically
|
||||
3. Multi-step research workflows work
|
||||
4. Progress indicators show during research
|
||||
5. Research results cite sources
|
||||
6. Research stored in memory with metadata
|
||||
7. All tests pass
|
||||
8. User validates in Open WebUI
|
||||
|
||||
## Next Phase Preview
|
||||
|
||||
**Phase 4: Enhanced Tool Integration**
|
||||
- Infrastructure tools (restart services, check logs)
|
||||
- File operations (Nextcloud integration)
|
||||
- Calendar management (CalDAV)
|
||||
- Code agent (codestral specialist)
|
||||
|
||||
---
|
||||
|
||||
**Ready to start?** This phase should take ~1 week and builds directly on the working Phase 1+2 foundation.
|
||||
Reference in New Issue
Block a user