Files
tatlock/IMPLEMENTATION_ROADMAP.md
T
jpmschweitzer fd459e9ffb docs: update documentation for v0.2.0 tools release
README.md:
- Add Tatlock agent capabilities and tool descriptions
- Add requirements section (Ollama, SearXNG setup)
- Add configuration examples for external services
- Add tool usage examples and philosophy
- Add troubleshooting for Ollama and SearXNG
- Update test statistics

AGENTS.md:
- Refactor for LLM development focus
- Add PydanticAI tool registration pattern
- Add tool implementation guidelines
- Remove project status, focus on development instructions

IMPLEMENTATION_ROADMAP.md:
- Mark Phase 1 as "MOSTLY COMPLETE"
- Update detailed completion status
- Update current state summary
2025-12-07 00:13:41 +01:00

16 KiB

Tatlock Implementation Roadmap

Reference: See PHILOSOPHY.md for the target architecture and vision

This document outlines the phased implementation plan to transform the current OpenAI-compatible API into the full Tatlock household butler system.

Current State (v0.1.1+ - Phase 1 Mostly Complete)

What we have:

  • The Orchestrator - FastAPI infrastructure layer
    • OpenAI-compatible API endpoints (Responses API + Chat Completions)
    • Streaming coordination and conversation management
    • Response format with reasoning support
    • Test infrastructure (131 tests, 81.78% coverage)
  • Tatlock Agent - Real PydanticAI integration
    • Connected to Ollama (mistral-nemo:latest)
    • British butler personality with research mindset
    • Streaming responses with reasoning
    • Tool calling framework functional
  • Permanent Tools
    • Calculator (safe mathematical expressions)
    • Date/Time toolkit (current time, relative dates, time differences)
    • Web search (SearXNG integration)
  • Mock agent (lorem-tester for testing)
  • Agent interface abstraction

What we need:

  • The Household - Full multi-agent coordination:
    • The Steward (first-tier request analysis)
    • Tatlock coordination layer (expert agent delegation)
    • Expert household staff agents (Librarian, Developer, Handyman, etc.)
  • Multi-tenant database architecture
  • Containerized service ecosystem
  • MCP (Model Context Protocol) integration
  • Dynamic model switching for specialized tasks

Phase 1: Real LLM Integration - PydanticAI + Tools

Goal: Connect to actual language models and establish the base plumbing

Note: Ollama is an external service dependency (already running separately)

Deliverables

  1. PydanticAI Integration

    • PydanticAI → Ollama connection
    • Agent creation patterns
    • Streaming response handling
    • Error handling and retries
  2. Convert Tatlock Agent

    • Convert Tatlock agent from mock to PydanticAI
    • British butler personality prompt
    • Research-oriented mindset
    • Streaming to reasoning output
    • Tool calling framework setup
  3. Permanent Tools

    • Calculator: Safe mathematical expression evaluation
    • Date/Time toolkit: Current time, relative dates, time differences
    • Web search: SearXNG integration (external service)
    • Tool registration with PydanticAI
  4. Testing Infrastructure

    • Integration tests with real LLM
    • Tool functionality tests
    • Response quality validation
    • 131 tests, 81.78% coverage

Success Criteria

  • PydanticAI agents can call Ollama (mistral-nemo:latest)
  • Streaming works end-to-end
  • Tool calling framework functional
  • Permanent tools working (calculator, date/time, search)
  • Tests pass with real LLM
  • Can switch models dynamically (e.g., Codestral for code)

Status

MOSTLY COMPLETE - Tatlock agent functional with permanent tools

Remaining Work

  • Dynamic model switching for specialized tasks (e.g., Codestral for coding)

Why First?

Without real LLM integration, we can't meaningfully implement the Steward/Butler pattern. Everything else depends on having actual AI agents working.


Phase 2: Orchestration Layer - The Steward

Goal: Implement the first-tier LLM call for tool/agent selection

Deliverables

  1. Orchestrator Framework

    • Python orchestrator service/module
    • Request preprocessing pipeline
    • Tool/agent registry system
    • Recommendation format definition
  2. Steward Agent Implementation

    • Steward prompt engineering
    • Tool selection logic
    • Agent recommendation generation
    • Output format (note to Butler)
  3. Tool Registry

    • Available tools catalog
    • Tool capability descriptions
    • Tool category organization
    • Dynamic tool loading

Success Criteria

  • Steward analyzes incoming requests
  • Produces tool/agent recommendations
  • Recommendations formatted as prepended note
  • Tool registry is queryable and extensible
  • Steward output visible in reasoning stream

Estimated Effort

3-4 weeks - Core intelligence routing


Phase 3: The Butler - Tatlock Agent

Goal: Implement the second-tier coordinator with personality within the existing Orchestrator infrastructure

Context: The Orchestrator (FastAPI infrastructure) already exists. This phase implements the real Tatlock PydanticAI agent to replace the current mock agent.

Deliverables

  1. Butler Agent (Tatlock)

    • PydanticAI agent implementation within Orchestrator
    • Personality prompt engineering (witty British butler)
    • Tool calling framework
    • Multi-agent coordination logic
  2. Scoped Tool Access

    • Filter tools based on Steward recommendations
    • Dynamic tool loading for Butler context
    • Tool execution framework
    • Result aggregation
  3. Real-Time Reasoning Output

    • Stream all Butler activities to reasoning output
    • Tool call progress indicators
    • Expert agent consultation messages
    • Wait time transparency

Success Criteria

  • Tatlock receives enriched requests (user + Steward notes)
  • Only recommended tools are available
  • Tatlock coordinates multiple tool calls
  • All actions streamed to reasoning output
  • Responses have consistent personality
  • Synthesizes multi-source results coherently

Estimated Effort

4-5 weeks - Complex coordination logic


Phase 4: Expert Household Staff - Core Agents

Goal: Implement the initial set of domain-specific expert agents

Priority Expert Agents

  1. The Librarian (Research & Knowledge Management) Priority

    • Research assistance and synthesis
    • Automatic research dossier generation
    • Knowledge base queries and organization
    • Reference management
    • Wiki integration (future: dedicated wiki container)
    • Mind map maintenance (future)
    • Rationale: Helps guide development priorities through better research
  2. The Developer (Software Development)

    • Code generation assistance
    • Debugging support
    • Documentation generation
    • Architecture guidance
    • Rationale: Directly supports building the system itself
  3. The Handyman (System Maintenance)

    • System status queries
    • Log analysis
    • Basic troubleshooting
    • Infrastructure monitoring
  4. The Secretary (Scheduling & Organization)

    • Calendar integration (placeholder)
    • Task management (placeholder)
    • Reminder system
    • Schedule conflict detection
  5. The Housekeeper (Home Automation)

    • Device control interface
    • Status queries
    • Automation triggers
    • Environmental monitoring

Each Agent Includes

  • Specialized prompt and personality
  • Domain-specific tools
  • MCP integration points (where applicable)
  • Integration with Butler orchestration

Success Criteria

  • Each agent implemented as separate module
  • Agents callable via tool framework
  • Agents use specialized prompts
  • Results integrate cleanly with Butler
  • Can invoke specialized models (e.g., Codestral for Developer)

Estimated Effort

6-8 weeks - Parallel development possible


Phase 5: Persistence Layer - Database & Multi-Tenancy

Goal: Add persistent storage and multi-user support when needed

Deliverables

  1. PostgreSQL Integration

    • Docker compose configuration for PostgreSQL
    • Database schema design with tenant isolation
    • Alembic migrations setup
    • SQLAlchemy models
  2. Multi-Tenant Architecture

    • Tenant identification middleware
    • Tenant-scoped database sessions
    • User authentication system (basic)
    • Per-tenant data isolation
  3. Core Data Models

    • Users and tenants
    • Conversations and messages (migrate from in-memory)
    • Agent interactions log
    • System configuration and preferences
  4. Migration Strategy

    • Gradual migration from in-memory to database
    • Backward compatibility during transition
    • Data export/import utilities

Success Criteria

  • PostgreSQL container running
  • Multiple users can authenticate separately
  • Each user sees only their own data
  • Conversations persist across restarts
  • Database migrations work correctly
  • Tests verify tenant isolation

Estimated Effort

3-4 weeks - Data layer foundation

Why Later?

The core orchestration (Steward → Butler → Experts) can work entirely with in-memory state. We only need database persistence when we want conversations to survive restarts and multiple users to have isolated experiences.


Phase 6: Extended Services Integration

Goal: Connect to additional supporting services

Services to Integrate

  1. Redis (Memory & Caching)

    • Docker compose setup
    • Conversation cache
    • Short-term memory
    • Session management
  2. Qdrant (Vector Storage)

    • Docker compose setup
    • Long-term memory embeddings
    • Semantic search
    • Conversation history vectors
  3. SearxNG (Web Search)

    • Docker compose setup
    • Search tool integration
    • Result processing
    • Privacy-preserving queries

Success Criteria

  • All services defined in docker-compose.yml
  • Services communicate correctly
  • Tatlock can invoke web search
  • Redis used for session data
  • Qdrant stores conversation embeddings
  • Ollama serves the base model

Estimated Effort

3-4 weeks - Infrastructure setup


Phase 7: MCP (Model Context Protocol) Integration

Goal: Enable rich tool integrations via MCP

Deliverables

  1. MCP Server Framework

    • MCP server implementation
    • Tool registration via MCP
    • Schema validation
    • Error handling
  2. MCP Client in Agents

    • PydanticAI MCP integration
    • Tool discovery from MCP servers
    • Dynamic tool loading
    • Result processing
  3. Initial MCP Tools

    • File system operations
    • Database queries
    • API integrations
    • System commands

Success Criteria

  • MCP server running
  • Tools exposed via MCP protocol
  • Agents can discover and use MCP tools
  • New tools addable without code changes
  • MCP tools visible in Steward recommendations

Estimated Effort

3-4 weeks - Standards-based integration


Phase 8: Advanced Memory & Context

Goal: Implement sophisticated memory and context management

Deliverables

  1. Long-Term Memory

    • Conversation embedding pipeline
    • Semantic search over history
    • Memory consolidation
    • Relevance ranking
  2. Context Management

    • Smart context window trimming
    • Conversation branching
    • Topic tracking
    • Memory retrieval integration
  3. Personalization

    • User preference learning
    • Interaction pattern analysis
    • Adaptive responses
    • Custom agent personalities per user

Success Criteria

  • Conversations automatically embedded to Qdrant
  • Relevant history retrieved for new requests
  • Context stays within model limits
  • User preferences affect responses
  • Memory improves over time

Estimated Effort

4-5 weeks - AI/ML heavy


Phase 9: Extended Household Staff

Goal: Add specialized agents for additional domains

Future Agents

  1. The Librarian (Knowledge Management)

    • Personal documentation indexing
    • Research assistance
    • Knowledge base queries
    • Reference management
  2. The Accountant (Financial Tracking)

    • Expense tracking
    • Budget monitoring
    • Financial reports
    • Transaction categorization
  3. The Chef (Meal Planning)

    • Recipe management
    • Meal planning
    • Nutrition tracking
    • Grocery lists
  4. Others as Needed

    • Domain-specific as requirements emerge

Success Criteria

  • Each new agent follows household pattern
  • Integrates with Steward/Butler flow
  • Has appropriate specialized tools
  • Documented in PHILOSOPHY.md updates

Estimated Effort

Ongoing - Add as needed


Phase 10: User Experience Refinement

Goal: Polish the interaction experience

Deliverables

  1. Personality Tuning

    • Refine Tatlock's wit and tone
    • Consistent household character
    • Cultural references appropriate
    • Humor that doesn't annoy
  2. Transparency Improvements

    • Better progress indicators
    • Clearer reasoning explanations
    • Informative wait messages
    • Error message clarity
  3. Performance Optimization

    • Response time improvements
    • Model loading optimization
    • Caching strategies
    • Streaming smoothness

Success Criteria

  • Users find Tatlock engaging
  • Wait times feel reasonable
  • Errors are understandable
  • System feels responsive

Estimated Effort

Ongoing - Continuous improvement


Phase 11: Production Hardening

Goal: Make the system production-ready for homelab deployment

Deliverables

  1. Deployment

    • Complete docker-compose stack
    • Environment configuration
    • Backup strategies
    • Update procedures
  2. Monitoring

    • Health checks
    • Performance metrics
    • Error tracking
    • Usage analytics
  3. Security

    • Authentication hardening
    • Rate limiting
    • Input validation
    • Audit logging
  4. Documentation

    • Installation guide
    • Configuration reference
    • Troubleshooting guide
    • Architecture documentation

Success Criteria

  • One-command deployment
  • System health is monitorable
  • Secure for homelab use
  • Well documented

Estimated Effort

3-4 weeks - Production polish


Dependencies Between Phases

Phase 1 (Ollama + PydanticAI) ← Foundation for all AI
    ↓
Phase 2 (Steward)
    ↓
Phase 3 (Butler/Tatlock)
    ↓
Phase 4 (Expert Agents) ← Phase 7 (MCP) can enhance
    ↓
Phase 5 (Database/Multi-Tenancy) ← Can be deferred
    ↓
Phase 6 (Extended Services) → Phase 8 (Advanced Memory)
    ↓
Phase 9 (Extended Staff) → Phase 10 (UX) → Phase 11 (Production)

Critical Path: Phases 1 → 2 → 3 → 4 must be sequential Can Be Deferred: Phase 5 (Database) until you need persistence Parallel Opportunities: Phase 6 and 7 can overlap; Phase 9 and 10 ongoing


Overall Timeline Estimate

Minimum Viable Household (Phases 1-4): 15-20 weeks

  • Working Steward → Butler → Expert Agents with real LLM
  • In-memory state (no persistence needed yet)
  • Core household functional

With Persistence (Phases 1-5): 18-24 weeks

  • Add database and multi-tenancy
  • Conversations survive restarts
  • Multiple users supported

Full-Featured System (Phases 1-9): 35-45 weeks

  • All services integrated
  • Advanced memory and context
  • Extended household staff

Production-Ready (All phases): 40-50 weeks

  • Polished UX
  • Hardened for homelab deployment
  • Fully documented

Note: Timeline assumes consistent part-time development effort


Success Metrics

Technical

  • System implements PHILOSOPHY.md patterns
  • All household roles functional
  • Multi-tenant isolation verified
  • Real-time reasoning transparency working
  • MCP integration complete

User Experience

  • Tatlock feels like interacting with a butler
  • Wait times are transparent and acceptable
  • Expert agents provide value in their domains
  • System is reliable and trustworthy

Architecture

  • Clean separation between household roles
  • Easy to add new agents/tools
  • Model efficiency (base model stays loaded)
  • Scales to household + friends usage

Risk Management

High Risk Items

  1. PydanticAI + Ollama integration complexity

    • Mitigation: Prototype early, iterate on connection layer
  2. Multi-agent coordination complexity

    • Mitigation: Start simple, add coordination gradually
  3. Model performance on homelab hardware

    • Mitigation: Model selection, quantization, optimization
  4. Prompt engineering for personality consistency

    • Mitigation: Extensive testing, user feedback, iteration

Medium Risk Items

  • MCP protocol adoption and tooling maturity
  • Vector embedding quality for memory
  • Home automation integration variability
  • User authentication security

Next Steps

  1. Immediate: Commit model name fix (Tatlock)
  2. Week 1-2: Begin Phase 1 (PostgreSQL + multi-tenancy design)
  3. Week 3: Parallel prototype of Steward agent
  4. Ongoing: Update this roadmap as we learn

Document Status: Active planning document Created: 2025-12-06 Last Updated: 2025-12-06