From 7b25985108bf7e77bdebd1d41b1c049ce8221b4f Mon Sep 17 00:00:00 2001 From: Jeroen Schweitzer Date: Sun, 7 Dec 2025 01:14:08 +0100 Subject: [PATCH] docs: expand Phase 2 roadmap with detailed Steward implementation plan MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Significantly expand the Steward system implementation plan with: Core Architecture: - Two-tier request flow diagram (Steward → Tatlock) - Detailed explanation of scope-narrowing principle 5 Major Deliverables: 1. Tool & Agent Registry System - Registry module with metadata schemas - Category-based organization - Dynamic discovery and loading 2. Steward PydanticAI Agent - Structured recommendation output - Request analysis and capability matching - Conservative tool/agent selection 3. Request Preprocessing Pipeline - Integration layer for Steward → Tatlock flow - Note formatting for recommendations - Tool scoping implementation 4. Real-Time Transparency - Stream Steward analysis to reasoning output - User visibility into resource planning 5. Model Efficiency Optimization - Shared base model to keep it hot in VRAM - Performance monitoring Implementation Strategy: - Week-by-week breakdown (7-8 weeks total) - Specific tasks and deliverables per week Enhanced Documentation: - Expanded success criteria (5 → 9 items) - Performance targets with quantified metrics - Risk mitigation strategies - Future enhancements roadmap Estimated effort increased from 3-4 weeks to 7-8 weeks to reflect comprehensive implementation scope with proper testing and optimization. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude Sonnet 4.5 --- IMPLEMENTATION_ROADMAP.md | 326 +++++++++++++++++++++++++++++++++++--- 1 file changed, 305 insertions(+), 21 deletions(-) diff --git a/IMPLEMENTATION_ROADMAP.md b/IMPLEMENTATION_ROADMAP.md index ad39ee4..909e5f5 100644 --- a/IMPLEMENTATION_ROADMAP.md +++ b/IMPLEMENTATION_ROADMAP.md @@ -92,35 +92,319 @@ Without real LLM integration, we can't meaningfully implement the Steward/Butler **Goal**: Implement the first-tier LLM call for tool/agent selection +**Purpose**: The Steward performs crucial preparatory work before Tatlock engages with a request. By analyzing incoming requests and determining which tools, services, and household staff members will be needed, the Steward creates a curated recommendation that streamlines Tatlock's work and prevents cognitive overload. + +### Core Architecture + +The Steward operates as the first tier in the two-tier request flow: + +``` +User Request → Orchestrator → Steward Analysis → Recommendations → Tatlock (with scoped tools/agents) +``` + +**Key Principle**: The Steward narrows the scope to only relevant capabilities, making Tatlock's decision-making cleaner and more focused. + ### Deliverables -1. **Orchestrator Framework** - - Python orchestrator service/module - - Request preprocessing pipeline - - Tool/agent registry system - - Recommendation format definition +#### 1. Tool & Agent Registry System -2. **Steward Agent Implementation** - - Steward prompt engineering - - Tool selection logic - - Agent recommendation generation - - Output format (note to Butler) +**Purpose**: Centralized catalog of all available capabilities for the Steward to recommend -3. **Tool Registry** - - Available tools catalog - - Tool capability descriptions - - Tool category organization - - Dynamic tool loading +**Implementation Details**: +- **Registry Module** (`src/core/registry.py`) + - Tool registration decorator pattern + - Agent registration with capability metadata + - Category-based organization (computation, information, automation, communication) + - Dynamic tool/agent discovery and loading + +- **Tool Metadata Schema** + ```python + { + "name": "calculator", + "category": "computation", + "description": "Safe mathematical expression evaluation", + "capabilities": ["arithmetic", "algebra", "trigonometry"], + "cost": "low", # computational cost indicator + "requires_network": false + } + ``` + +- **Agent Metadata Schema** + ```python + { + "name": "developer", + "role": "The Developer", + "category": "technical", + "description": "Software development assistance", + "domains": ["code_generation", "debugging", "architecture"], + "specialized_model": "codestral", # optional + "cost": "high" + } + ``` + +- **Registry API** + - `get_all_tools()` - List all available tools + - `get_all_agents()` - List all expert agents + - `get_by_category(category)` - Filter by category + - `search_by_capability(query)` - Semantic search (future: vector search) + +**Testing**: +- Unit tests for registration and retrieval +- Test dynamic loading of new tools/agents +- Validate metadata schemas + +#### 2. Steward PydanticAI Agent + +**Purpose**: First-tier LLM that analyzes requests and recommends relevant tools/agents + +**Implementation Details**: + +- **Agent Module** (`src/agents/steward.py`) + ```python + from pydantic_ai import Agent, RunContext + from pydantic import BaseModel + + class StewardRecommendation(BaseModel): + """Structured output from Steward analysis""" + recommended_tools: list[str] + recommended_agents: list[str] + reasoning: str + estimated_complexity: str # "simple", "moderate", "complex" + requires_multi_step: bool + + steward = Agent( + 'ollama:mistral-nemo', # Same base model as Tatlock + result_type=StewardRecommendation, + system_prompt="""...""" + ) + ``` + +- **System Prompt Engineering** + - Role: Estate steward responsible for efficient household coordination + - Task: Analyze requests to determine needed resources + - Output: Structured recommendations with reasoning + - Constraints: Be conservative (recommend only truly relevant capabilities) + - Context: Full registry of available tools and agents + +- **Steward Tools** + ```python + @steward.tool + def get_available_capabilities(ctx: RunContext) -> dict: + """Get catalog of all available tools and agents.""" + return { + "tools": registry.get_all_tools(), + "agents": registry.get_all_agents() + } + ``` + +- **Request Analysis Flow** + 1. Receive user request + 2. Query capability registry via tool + 3. Analyze request for required capabilities + 4. Generate structured recommendation + 5. Format as note to Tatlock + +**Testing**: +- Test various request types (simple, complex, multi-domain) +- Verify recommendations are relevant and not over-inclusive +- Test structured output parsing +- Validate reasoning quality + +#### 3. Request Preprocessing Pipeline + +**Purpose**: Integration layer that routes requests through Steward before Tatlock + +**Implementation Details**: + +- **Preprocessing Module** (`src/core/preprocessing.py`) + ```python + async def preprocess_request(user_request: str) -> EnrichedRequest: + """ + 1. Call Steward for analysis + 2. Get recommendations + 3. Enrich original request + 4. Return scoped context for Tatlock + """ + # Get Steward analysis + steward_result = await steward.run(user_request) + recommendations = steward_result.data + + # Create note to Tatlock + steward_note = format_steward_note(recommendations) + + # Build scoped tool/agent list + scoped_tools = get_scoped_tools(recommendations.recommended_tools) + scoped_agents = get_scoped_agents(recommendations.recommended_agents) + + return EnrichedRequest( + original_request=user_request, + steward_note=steward_note, + available_tools=scoped_tools, + available_agents=scoped_agents, + metadata=recommendations + ) + ``` + +- **Note Formatting** + ``` + === Internal Note from the Steward === + + Request Analysis: + {steward reasoning} + + Recommended Tools: + - calculator: For mathematical computations + - web_search: To find current information + + Recommended Household Staff: + - The Developer: For code generation assistance + + Estimated Complexity: moderate + =================================== + + [Original User Request] + ``` + +- **Orchestrator Integration** + - Modify `src/responses/service.py` to call preprocessing + - Prepend Steward note to request before sending to Tatlock + - Limit Tatlock's tool access to recommended tools only + - Stream Steward's reasoning to output + +**Testing**: +- Integration tests for full preprocessing flow +- Test request enrichment format +- Verify tool scoping works correctly +- Test streaming of Steward reasoning + +#### 4. Real-Time Transparency + +**Purpose**: Stream Steward's analysis to user's reasoning output + +**Implementation Details**: + +- **Streaming Integration** (`src/responses/streaming.py`) + - Add Steward analysis phase to stream + - Format as reasoning item + - Include recommendation summary + +- **Example Output to User**: + ``` + [Reasoning] + Consulting the Steward for resource planning... + + The Steward's Analysis: + - Request requires mathematical computation + - Need to verify current information via web search + - May benefit from Developer's code expertise + + Recommended: calculator, web_search, The Developer + + Proceeding with scoped resources... + ``` + +**Testing**: +- Test streaming of Steward analysis +- Verify formatting in Open WebUI +- Test error handling if Steward fails + +#### 5. Model Efficiency Optimization + +**Purpose**: Ensure the base model stays loaded in VRAM + +**Implementation Details**: + +- **Shared Model Configuration** + - Both Steward and Tatlock use `ollama:mistral-nemo` by default + - Sequential calls (Steward → Tatlock) keep model hot + - No reload delays between tiers + +- **Performance Monitoring** + - Log response times for Steward calls + - Track total request latency (Steward + Tatlock) + - Identify optimization opportunities + +**Testing**: +- Benchmark Steward → Tatlock call latency +- Verify model stays loaded between calls +- Test performance under load + +### Implementation Strategy + +#### Week 1-2: Foundation +- [ ] Design and implement registry system +- [ ] Create tool/agent metadata schemas +- [ ] Build registry API with tests +- [ ] Migrate existing tools to registry + +#### Week 3-4: Steward Agent +- [ ] Create Steward PydanticAI agent +- [ ] Engineer system prompt for analysis +- [ ] Implement structured recommendation output +- [ ] Add registry query tool +- [ ] Test with various request types + +#### Week 5-6: Integration +- [ ] Build request preprocessing pipeline +- [ ] Implement note formatting +- [ ] Integrate with Orchestrator +- [ ] Add streaming transparency +- [ ] Tool scoping for Tatlock + +#### Week 7: Testing & Refinement +- [ ] End-to-end integration tests +- [ ] Performance optimization +- [ ] Prompt refinement based on results +- [ ] Documentation and examples ### Success Criteria -- [ ] Steward analyzes incoming requests -- [ ] Produces tool/agent recommendations -- [ ] Recommendations formatted as prepended note -- [ ] Tool registry is queryable and extensible -- [ ] Steward output visible in reasoning stream + +- [ ] **Steward analyzes incoming requests** using PydanticAI agent +- [ ] **Produces structured recommendations** (tools, agents, reasoning) +- [ ] **Recommendations formatted as prepended note** to Tatlock +- [ ] **Tool registry is queryable and extensible** via clean API +- [ ] **Steward output visible in reasoning stream** for transparency +- [ ] **Only recommended tools available** to Tatlock (scoped context) +- [ ] **Base model stays loaded** between Steward and Tatlock calls +- [ ] **Recommendations are accurate** (not over/under-inclusive) +- [ ] **Integration tests pass** for full Steward → Tatlock flow + +### Performance Targets + +- **Steward Analysis Time**: < 2 seconds for typical requests +- **Total Added Latency**: < 3 seconds including streaming +- **Recommendation Accuracy**: > 90% relevance (manual evaluation) +- **Model Reload Delay**: 0 seconds (model stays hot) + +### Risk Mitigation + +**Risk**: Steward recommendations too broad (defeats purpose) +- Mitigation: Conservative prompt engineering, test with diverse requests, iterate + +**Risk**: Added latency unacceptable to users +- Mitigation: Stream Steward reasoning for transparency, optimize prompt, parallel processing where possible + +**Risk**: Tool registry becomes unwieldy +- Mitigation: Good categorization, semantic search (future), regular pruning + +**Risk**: Steward and Tatlock models compete for VRAM +- Mitigation: Use same base model, sequential calls, monitor memory + +### Future Enhancements (Post-Phase 2) + +- **Semantic Search**: Vector-based capability search instead of metadata lookup +- **Learning from Usage**: Track which recommendations work well, adjust over time +- **Confidence Scores**: Steward provides confidence for each recommendation +- **Request Classification**: Cache classifications for similar requests +- **Multi-Model Support**: Allow Steward to recommend specialized models for specific tasks ### Estimated Effort -**3-4 weeks** - Core intelligence routing + +**7-8 weeks** - Core intelligence routing with comprehensive implementation + +### Why Second? + +The Steward is the foundation of the household architecture. Without it, we'd need to expose all tools/agents to Tatlock, creating cognitive overload and poor decision-making. The Steward enables the focused expertise pattern that makes the whole system work. ---