Compare commits

..
Author SHA1 Message Date
jpmschweitzerandClaude Opus 4.5 9d7ce399c8 fix(memory): biographer tool type hints for Ollama (v1.3.2)
Build and Push / build (release) Successful in 51s
- Change `str | None` to `str` with empty default for memory_type
- Remove `keywords` parameter from store_insight (auto-generated anyway)
- Ollama's OpenAI API doesn't handle union types with None properly

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2025-12-14 15:04:10 +01:00
jpmschweitzerandClaude Opus 4.5 8e38a568ef fix(memory): add biographer to delegation wrappers (v1.3.1)
Build and Push / build (release) Successful in 50s
- Add delegate_to_biographer to household registry delegation map
- Was returning raw tools which caused Ollama "invalid message content type: nil"
- Add Qdrant host/port to .env.example

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2025-12-14 14:56:32 +01:00
jpmschweitzerandClaude Opus 4.5 40663511b4 feat: memory system fixes and Redis config cleanup (v1.3.0)
Build and Push / build (release) Successful in 26s
- Fix Qdrant client to use query_points API (qdrant-client >= 1.10)
- Rename REDIS_DB to REDIS_BENCHMARK_DB for clarity
- Update Redis defaults to match stack allocation (benchmark=6, memory=1)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2025-12-14 14:36:56 +01:00
jpmschweitzer 8f008c7fd2 no longer needed 2025-12-14 14:05:07 +01:00
jpmschweitzerandClaude Opus 4.5 d207594e3c fix(deps): add missing pydantic-settings dependency
Build and Push / build (release) Successful in 50s
pydantic-ai-slim doesn't include pydantic-settings as a transitive
dependency like the full pydantic-ai package did.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2025-12-14 13:57:33 +01:00
jpmschweitzerandClaude Opus 4.5 523c5c43a0 feat(ci): trigger Watchtower update after image push
Build and Push / build (release) Successful in 1m56s
Automatically notify Watchtower to pull and deploy the new image
after a successful registry push.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2025-12-14 13:01:59 +01:00
jpmschweitzerandClaude Opus 4.5 822cdc9bf4 fix(ci): upgrade to build-push-action@v6, disable sbom
- Upgrade docker/build-push-action from v5 to v6
- Add sbom: false alongside provenance: false
- Update registry URL to internal domain

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2025-12-14 11:42:37 +01:00
jpmschweitzer 214dc4e725 fixed ci/cd network issue
Build and Push / build (release) Failing after 11s
2025-12-14 10:49:29 +01:00
jpmschweitzerandClaude Opus 4.5 acdde99a5c fix(ci): disable provenance for Gitea registry compatibility
Build and Push / build (release) Failing after 1m1s
Add provenance: false to docker/build-push-action to fix
"received unexpected HTTP status: 200 OK" error when pushing
to Gitea container registry.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2025-12-13 21:13:16 +01:00
jpmschweitzerandClaude Opus 4.5 4e6f1da4f3 chore: slim dependencies with pydantic-ai-slim[openai]
Build and Push / build (release) Failing after 1m3s
- Switch from pydantic-ai to pydantic-ai-slim[openai]
- Removes unused provider SDKs (anthropic, boto3, cohere, google, groq, huggingface)
- Production packages: 53 (down from ~158)
- Production footprint: 178MB
- Add DEPENDENCY_SLIM.md with rollback instructions

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2025-12-13 21:02:09 +01:00
jpmschweitzerandClaude Opus 4.5 7dd2c20e76 docs: update README and roadmap for v1.2.0
Build and Push / build (release) Failing after 1m47s
README.md:
- Add household staff table with current status
- Update requirements to list external services
- Add Redis, Qdrant to configuration section
- Update project structure with new modules
- Update version to 1.2.0

IMPLEMENTATION_ROADMAP.md:
- Update current state to v1.2.0
- Mark Phase 2 (Steward) as complete
- Mark Phase 3 (Butler coordination) as complete
- Update Phase 4 with Librarian and Biographer complete
- Mark Phase 6 (Services) as complete
- Update Phase 8 (Memory) with completed items
- Update next steps

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2025-12-13 19:29:30 +01:00
jpmschweitzerandClaude Opus 4.5 7426dd1ac3 feat: add Phase F.2 - The Biographer (memory agent)
Add The Biographer household member for user memory management:

Memory Service (direct access layer):
- src/core/memory_service.py for fast, LLM-free lookups
- Profile, preference, and fact management
- Session context with Redis caching
- Steward integration via prefetch_context()

The Biographer Agent:
- src/agents/biographer/ package with PydanticAI agent
- Discreet chronicler personality for privacy
- Tools: recall_semantic, list_memories, store_insight,
  update_profile, update_preference, forget_memory
- Registered with Household Registry on startup

Steward Integration:
- Memory context pre-fetch during analysis
- Profile/preferences included in Butler note
- Keyword-based context determination

Also includes:
- delegate_to_biographer() wrapper
- 34 new tests (capability + memory service)
- Version bump to 1.2.0

Documentation cleanup:
- Removed obsolete PHASE2_COMPLETE.md, PHASE2_PLAN.md
- Removed docs/library-desk-requirements.md
- Moved ORCHESTRATION_SCENARIOS.md to project root

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2025-12-13 19:20:18 +01:00
26 changed files with 2467 additions and 1977 deletions
+7 -2
View File
@@ -21,13 +21,18 @@ SEARXNG_TIMEOUT=30
# Redis Configuration # Redis Configuration
REDIS_HOST=redis-shared REDIS_HOST=redis-shared
REDIS_PORT=6379 REDIS_PORT=6379
REDIS_DB=1 REDIS_MEMORY_DB=1
REDIS_BENCHMARK_DB=6
REDIS_TIMEOUT=5 REDIS_TIMEOUT=5
# Qdrant Configuraton
QDRANT_HOST=qdrant
QDRANT_PORT=6333
# Logging # Logging
LOG_LEVEL=INFO LOG_LEVEL=INFO
ENABLE_BENCHMARKS=true ENABLE_BENCHMARKS=true
# Note: Log format is auto-selected based on ENVIRONMENT (console for dev, json for production) # Note: Log format is auto-selected based on ENVIRONMENT (console for dev, json for production)
# CORS (comma-separated list) # CORS (comma-separated list)
CORS_ORIGINS=* CORS_ORIGINS=["*"]
+12 -4
View File
@@ -13,15 +13,23 @@ jobs:
- name: Login to Gitea Registry - name: Login to Gitea Registry
uses: docker/login-action@v3 uses: docker/login-action@v3
with: with:
registry: git.schweitz.net registry: git.schweitz.internal
username: ${{ secrets.REGISTRY_USER }} username: ${{ secrets.REGISTRY_USER }}
password: ${{ secrets.REGISTRY_PASSWORD }} password: ${{ secrets.REGISTRY_PASSWORD }}
- name: Build and push - name: Build and push
uses: docker/build-push-action@v5 uses: docker/build-push-action@v6
with: with:
context: . context: .
push: true push: true
provenance: false
sbom: false
tags: | tags: |
git.schweitz.net/jpmschweitzer/tatlock:latest git.schweitz.internal/jpmschweitzer/tatlock:latest
git.schweitz.net/jpmschweitzer/tatlock:${{ github.ref_name }} git.schweitz.internal/jpmschweitzer/tatlock:${{ github.ref_name }}
- name: Trigger Watchtower update
if: success()
run: |
curl -sf -H "Authorization: Bearer ${{ secrets.WATCHTOWER_TOKEN }}" \
http://watchtower:8080/v1/update
+138 -1
View File
@@ -7,6 +7,142 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
## [Unreleased] ## [Unreleased]
## [1.3.2] - 2025-12-14
### Fixed
- **Memory**: Fix biographer tool type hints for Ollama compatibility (remove `| None` union types)
## [1.3.1] - 2025-12-14
### Fixed
- **Memory**: Add biographer to delegation wrappers (was returning raw tools causing Ollama error)
- **Config**: Add Qdrant host/port to .env.example
## [1.3.0] - 2025-12-14
### Fixed
- **Memory**: Update Qdrant client to use `query_points` API (qdrant-client >= 1.10)
### Changed
- **Config**: Rename `REDIS_DB` to `REDIS_BENCHMARK_DB` for clarity
- **Config**: Update Redis defaults to match stack allocation (benchmark=6, memory=1)
## [1.2.5] - 2025-12-14
### Fixed
- **Dependencies**: Add missing `pydantic-settings` (not included in pydantic-ai-slim)
## [1.2.4] - 2025-12-14
### Added
- **CI**: Trigger Watchtower update after successful image push
## [1.2.3] - 2025-12-14
### Fixed
- **CI**: Upgrade to build-push-action@v6, disable provenance and sbom for Gitea registry
## [1.2.2] - 2025-12-13
### Fixed
- **CI**: Add `provenance: false` to docker/build-push-action to fix Gitea registry push
## [1.2.1] - 2025-12-13
### Changed
- **Dependency slimming**: Switched from `pydantic-ai` to `pydantic-ai-slim[openai]`
- Removes unused LLM provider SDKs (anthropic, boto3, cohere, google-genai, groq, huggingface)
- Production packages: 53 (down from ~158)
- Production footprint: 178MB
- Tatlock uses Ollama via OpenAI-compatible API, so only `openai` extra is needed
- See `DEPENDENCY_SLIM.md` for rollback instructions
## [1.2.0] - 2025-12-13
### Added
#### Phase F: Memory System (The Biographer)
- **Memory Infrastructure** (Phase F.1):
- `src/core/context.py`: ContextVar-based request context for async-safe user/conversation tracking
- `get_user()`, `get_conversation_id()` helpers
- `RequestContext` manager for clean setup/teardown
- `src/core/multi_tenancy.py`: User ID sanitization and collection naming
- Per-user collection pattern: `memories_{user}`
- Redis key patterns: `session:{user}:{conv}`, `entities:{user}:{conv}`
- `src/core/embeddings.py`: Ollama embedding client
- nomic-embed-text model (768 dimensions)
- `embed()`, `embed_batch()`, `health_check()` methods
- `src/core/qdrant.py`: Qdrant vector database client
- `ensure_collection()`, `upsert_memory()`, `search_memories()`, `delete_memory()`
- Type-based filtering for memory queries
- `src/core/memory_cache.py`: Redis session memory cache
- Session context with 24h TTL (db=2, separate from benchmarks)
- Recent entities tracking per conversation
- **Memory Service** (Phase F.2a):
- `src/core/memory_service.py`: Direct access layer for fast, LLM-free memory lookups
- Profile methods: `get_profile()`, `set_profile()`
- Preference methods: `get_preference()`, `set_preference()`, `get_all_preferences()`
- Fact methods: `store_fact()`, `get_fact()`
- Session context: `get_session_context()`, `set_session_context()`, `update_session_context()`
- Steward integration: `prefetch_context()` for request preprocessing
- **The Biographer Agent** (Phase F.2b):
- `src/agents/biographer/`: Household memory keeper agent
- PydanticAI agent with discreet chronicler personality
- System prompt emphasizes privacy and accurate recall
- **Biographer Tools** (`src/agents/biographer/tools.py`):
- `recall_semantic`: Semantic search for memories by meaning
- `list_memories`: Browse stored memories by type
- `store_insight`: Record new facts from conversation
- `update_profile`: Update core profile fields (name, location, timezone)
- `update_preference`: Update user preferences (units, theme)
- `forget_memory`: Remove specific memories
- **Capability Registration**:
- `BIOGRAPHER_CAPABILITY` with context domain
- Automatic registration on startup
- Low cost (vector search, minimal LLM)
- **Delegation Wrapper**:
- `delegate_to_biographer()` in `src/agents/delegation.py`
- Async delegation with error handling
- **Steward Memory Integration**:
- Memory context pre-fetch during request analysis
- Profile and preferences included in Steward's note to Butler
- Keyword-based context determination (weather → location, time → timezone)
- **Configuration**:
- `QDRANT_HOST`, `QDRANT_PORT`, `QDRANT_EMBEDDING_DIM` (768)
- `OLLAMA_EMBEDDING_MODEL` (nomic-embed-text)
- `REDIS_MEMORY_DB` (2), `REDIS_MEMORY_TTL_HOURS` (24)
- **Test Suite**:
- 34 new tests for memory system
- Biographer capability tests (15 tests)
- Memory service tests (19 tests)
- **OpenAI Standard `user` Field**:
- Added `user` field to `ResponseRequest` schema
- Request context set at API entry point
- Propagates through async calls via ContextVar
### Changed
- Application startup now registers The Biographer with Household Registry
- Steward analysis includes memory context pre-fetch
- Librarian client methods now use `get_user()` from context (12 methods updated)
- Request router sets user/conversation context at entry
## [1.1.0] - 2025-12-11 ## [1.1.0] - 2025-12-11
### Added ### Added
@@ -390,7 +526,8 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
- CORS middleware - CORS middleware
- Exception handlers (OpenAI-compatible error format) - Exception handlers (OpenAI-compatible error format)
[Unreleased]: https://git.schweitz.net/jpmschweitzer/tatlock/compare/v1.1.0...main [Unreleased]: https://git.schweitz.net/jpmschweitzer/tatlock/compare/v1.2.0...main
[1.2.0]: https://git.schweitz.net/jpmschweitzer/tatlock/compare/v1.1.0...v1.2.0
[1.1.0]: https://git.schweitz.net/jpmschweitzer/tatlock/compare/v1.0.0a...v1.1.0 [1.1.0]: https://git.schweitz.net/jpmschweitzer/tatlock/compare/v1.0.0a...v1.1.0
[1.0.0a]: https://git.schweitz.net/jpmschweitzer/tatlock/compare/v0.2.5...v1.0.0a [1.0.0a]: https://git.schweitz.net/jpmschweitzer/tatlock/compare/v0.2.5...v1.0.0a
[0.2.5]: https://git.schweitz.net/jpmschweitzer/tatlock/compare/v0.2.0...v0.2.5 [0.2.5]: https://git.schweitz.net/jpmschweitzer/tatlock/compare/v0.2.0...v0.2.5
+124 -87
View File
@@ -4,35 +4,37 @@
This document outlines the phased implementation plan to transform the current OpenAI-compatible API into the full Tatlock household butler system. This document outlines the phased implementation plan to transform the current OpenAI-compatible API into the full Tatlock household butler system.
## Current State (v0.1.1+ - Phase 1 Mostly Complete) ## Current State (v1.2.0 - Phase F Complete)
**What we have**: **What we have**:
-**The Orchestrator** - FastAPI infrastructure layer -**The Orchestrator** - FastAPI infrastructure layer
- OpenAI-compatible API endpoints (Responses API + Chat Completions) - OpenAI-compatible API endpoints (Responses API + Chat Completions)
- Streaming coordination and conversation management - Streaming coordination and conversation management
- Response format with reasoning support - Response format with reasoning support
- Test infrastructure (131 tests, 81.78% coverage) - Test infrastructure (~400 tests)
-**Tatlock Agent** - Real PydanticAI integration -**Two-Tier Architecture**
- Connected to Ollama (mistral-nemo:latest) - The Steward analyzes requests and recommends capabilities
- British butler personality with research mindset - Tatlock coordinates execution with scoped tools
- Streaming responses with reasoning - Real-time streaming of analysis and reasoning
- Tool calling framework functional -**Household Staff**
-**Permanent Tools** - **Tatlock** (Butler): Primary interface with witty personality
- Calculator (safe mathematical expressions) - **The Steward**: Request analysis and capability recommendation
- Date/Time toolkit (current time, relative dates, time differences) - **The Librarian**: Research via library-desk HybridRAG + wiki
- Web search (SearXNG integration) - **The Biographer**: User memory, profiles, preferences, semantic recall
-**Core Tools**
- Calculator, Date/Time toolkit, Web search (SearXNG)
-**Memory System**
- Direct access layer (memory_service) for fast lookups
- Vector storage (Qdrant) for semantic recall
- Session cache (Redis) with 24h TTL
- Multi-tenancy via ContextVar
- ✅ Mock agent (lorem-tester for testing) - ✅ Mock agent (lorem-tester for testing)
- ✅ Agent interface abstraction
**What we need**: **What we need**:
- **The Household** - Full multi-agent coordination: - More household staff (Developer, Secretary, Handyman, Housekeeper)
- The Steward (first-tier request analysis)
- Tatlock coordination layer (expert agent delegation)
- Expert household staff agents (Librarian, Developer, Handyman, etc.)
- Multi-tenant database architecture
- Containerized service ecosystem
- MCP (Model Context Protocol) integration - MCP (Model Context Protocol) integration
- Dynamic model switching for specialized tasks - Dynamic model switching for specialized tasks
- Full multi-tenant database (PostgreSQL)
--- ---
@@ -359,15 +361,18 @@ User Request → Orchestrator → Steward Analysis → Recommendations → Tatlo
### Success Criteria ### Success Criteria
- [ ] **Steward analyzes incoming requests** using PydanticAI agent - [x] **Steward analyzes incoming requests** using PydanticAI agent
- [ ] **Produces structured recommendations** (tools, agents, reasoning) - [x] **Produces structured recommendations** (tools, agents, reasoning)
- [ ] **Recommendations formatted as prepended note** to Tatlock - [x] **Recommendations formatted as prepended note** to Tatlock
- [ ] **Tool registry is queryable and extensible** via clean API - [x] **Tool registry is queryable and extensible** via clean API
- [ ] **Steward output visible in reasoning stream** for transparency - [x] **Steward output visible in reasoning stream** for transparency
- [ ] **Only recommended tools available** to Tatlock (scoped context) - [x] **Only recommended tools available** to Tatlock (scoped context)
- [ ] **Base model stays loaded** between Steward and Tatlock calls - [x] **Base model stays loaded** between Steward and Tatlock calls
- [ ] **Recommendations are accurate** (not over/under-inclusive) - [x] **Recommendations are accurate** (not over/under-inclusive)
- [ ] **Integration tests pass** for full Steward → Tatlock flow - [x] **Integration tests pass** for full Steward → Tatlock flow
### Status
**✅ COMPLETE** (v0.2.5)
### Performance Targets ### Performance Targets
@@ -435,12 +440,15 @@ The Steward is the foundation of the household architecture. Without it, we'd ne
- Wait time transparency - Wait time transparency
### Success Criteria ### Success Criteria
- [ ] Tatlock receives enriched requests (user + Steward notes) - [x] Tatlock receives enriched requests (user + Steward notes)
- [ ] Only recommended tools are available - [x] Only recommended tools are available
- [ ] Tatlock coordinates multiple tool calls - [x] Tatlock coordinates multiple tool calls
- [ ] All actions streamed to reasoning output - [x] All actions streamed to reasoning output
- [ ] Responses have consistent personality - [x] Responses have consistent personality
- [ ] Synthesizes multi-source results coherently - [x] Synthesizes multi-source results coherently
### Status
**✅ COMPLETE** (v1.1.0)
### Estimated Effort ### Estimated Effort
**4-5 weeks** - Complex coordination logic **4-5 weeks** - Complex coordination logic
@@ -453,39 +461,44 @@ The Steward is the foundation of the household architecture. Without it, we'd ne
### Priority Expert Agents ### Priority Expert Agents
1. **The Librarian** (Research & Knowledge Management) **Priority** 1. **The Librarian** (Research & Knowledge Management) **COMPLETE** (v1.1.0)
- Research assistance and synthesis - Research assistance via library-desk HybridRAG
- Automatic research dossier generation - Wiki page management (search, create, update)
- Knowledge base queries and organization - Semantic vector search
- Reference management - Knowledge graph queries
- Wiki integration (future: dedicated wiki container) - Dossier browsing
- Mind map maintenance (future)
- *Rationale: Helps guide development priorities through better research*
2. **The Developer** (Software Development) 2. **The Biographer** (User Memory) ✅ **COMPLETE** (v1.2.0)
- User profile management (name, location, timezone)
- Preference storage (units, theme)
- Semantic memory recall ("What car do I drive?")
- Fact storage from conversations
- Session context caching
3. **The Developer** (Software Development) 🔜 **Planned**
- Code generation assistance - Code generation assistance
- Debugging support - Debugging support
- Documentation generation - Documentation generation
- Architecture guidance - Architecture guidance
- *Rationale: Directly supports building the system itself* - *Rationale: Directly supports building the system itself*
3. **The Handyman** (System Maintenance) 4. **The Handyman** (System Maintenance) 🔜 **Planned**
- System status queries - System status queries
- Log analysis - Log analysis
- Basic troubleshooting - Basic troubleshooting
- Infrastructure monitoring - Infrastructure monitoring
4. **The Secretary** (Scheduling & Organization) 5. **The Secretary** (Scheduling & Organization) 🔜 **Planned**
- Calendar integration (placeholder) - Calendar integration
- Task management (placeholder) - Task management
- Reminder system - Reminder system
- Schedule conflict detection - Schedule conflict detection
5. **The Housekeeper** (Home Automation) 6. **The Housekeeper** (Home Automation) 🔜 **Planned**
- Home Assistant integration
- Device control interface - Device control interface
- Status queries - Status queries
- Automation triggers - Automation triggers
- Environmental monitoring
### Each Agent Includes ### Each Agent Includes
- Specialized prompt and personality - Specialized prompt and personality
@@ -494,12 +507,15 @@ The Steward is the foundation of the household architecture. Without it, we'd ne
- Integration with Butler orchestration - Integration with Butler orchestration
### Success Criteria ### Success Criteria
- [ ] Each agent implemented as separate module - [x] Each agent implemented as separate module
- [ ] Agents callable via tool framework - [x] Agents callable via tool framework
- [ ] Agents use specialized prompts - [x] Agents use specialized prompts
- [ ] Results integrate cleanly with Butler - [x] Results integrate cleanly with Butler
- [ ] Can invoke specialized models (e.g., Codestral for Developer) - [ ] Can invoke specialized models (e.g., Codestral for Developer)
### Status
**🔶 PARTIAL** - Librarian and Biographer complete, others planned
### Estimated Effort ### Estimated Effort
**6-8 weeks** - Parallel development possible **6-8 weeks** - Parallel development possible
@@ -556,31 +572,37 @@ The core orchestration (Steward → Butler → Experts) can work entirely with i
### Services to Integrate ### Services to Integrate
1. **Redis (Memory & Caching)** 1. **Redis (Memory & Caching)** ✅ **COMPLETE** (v1.2.0)
- Docker compose setup - Benchmark storage (db=1)
- Conversation cache - Memory cache for sessions (db=2)
- Short-term memory - 24h TTL for session context
- Session management - Recent entities tracking
3. **Qdrant (Vector Storage)** 2. **Qdrant (Vector Storage)** ✅ **COMPLETE** (v1.2.0)
- Docker compose setup - Per-user memory collections
- Long-term memory embeddings - 768-dim nomic-embed-text vectors
- Semantic search - Semantic search for recall
- Conversation history vectors - Type-based filtering
4. **SearxNG (Web Search)** 3. **SearxNG (Web Search)** ✅ **COMPLETE** (v0.2.0)
- Docker compose setup
- Search tool integration - Search tool integration
- Result processing - Result processing
- Privacy-preserving queries - Privacy-preserving queries
4. **library-desk (Research API)** ✅ **COMPLETE** (v1.1.0)
- HybridRAG search
- Wiki management
- Knowledge graph queries
### Success Criteria ### Success Criteria
- [ ] All services defined in docker-compose.yml - [x] Services communicate correctly
- [ ] Services communicate correctly - [x] Tatlock can invoke web search
- [ ] Tatlock can invoke web search - [x] Redis used for session data
- [ ] Redis used for session data - [x] Qdrant stores user memories
- [ ] Qdrant stores conversation embeddings - [x] Ollama serves the base model
- [ ] Ollama serves the base model
### Status
**✅ COMPLETE** - All core services integrated
### Estimated Effort ### Estimated Effort
**3-4 weeks** - Infrastructure setup **3-4 weeks** - Infrastructure setup
@@ -629,33 +651,48 @@ The core orchestration (Steward → Butler → Experts) can work entirely with i
### Deliverables ### Deliverables
1. **Long-Term Memory** 1. **Long-Term Memory** ✅ **COMPLETE** (v1.2.0 - Phase F)
- Conversation embedding pipeline - Memory service for direct key-based access
- Semantic search over history - Qdrant vector storage for semantic recall
- Memory consolidation - Embedding via nomic-embed-text
- Relevance ranking - The Biographer agent for memory management
2. **Context Management** 2. **Session Memory** ✅ **COMPLETE** (v1.2.0)
- Redis session cache with 24h TTL
- Recent entities tracking
- Conversation context preservation
- Multi-tenancy via ContextVar
3. **Steward Integration** ✅ **COMPLETE** (v1.2.0)
- Memory pre-fetch during request analysis
- Profile/preferences included in context
- Keyword-based context determination
4. **Context Management** 🔜 **Future**
- Smart context window trimming - Smart context window trimming
- Conversation branching - Conversation branching
- Topic tracking - Topic tracking
- Memory retrieval integration - Memory retrieval integration
3. **Personalization** 5. **Personalization** 🔜 **Future**
- User preference learning - User preference learning
- Interaction pattern analysis - Interaction pattern analysis
- Adaptive responses - Adaptive responses
- Custom agent personalities per user - Custom agent personalities per user
### Success Criteria ### Success Criteria
- [x] User facts stored in Qdrant with semantic search
- [x] Profile and preferences accessible via memory_service
- [x] Session context cached in Redis
- [x] User preferences affect responses (via Steward pre-fetch)
- [ ] Conversations automatically embedded to Qdrant - [ ] Conversations automatically embedded to Qdrant
- [ ] Relevant history retrieved for new requests - [ ] Memory improves over time (learning from interactions)
- [ ] Context stays within model limits
- [ ] User preferences affect responses ### Status
- [ ] Memory improves over time **🔶 PARTIAL** - Core memory system complete, advanced features planned
### Estimated Effort ### Estimated Effort
**4-5 weeks** - AI/ML heavy **4-5 weeks** - AI/ML heavy (remaining work)
--- ---
@@ -871,13 +908,13 @@ Phase 9 (Extended Staff) → Phase 10 (UX) → Phase 11 (Production)
## Next Steps ## Next Steps
1. **Immediate**: Commit model name fix (Tatlock) 1. **Priority**: Implement The Developer agent for code assistance
2. **Week 1-2**: Begin Phase 1 (PostgreSQL + multi-tenancy design) 2. **Integration**: Add Home Assistant integration for The Housekeeper
3. **Week 3**: Parallel prototype of Steward agent 3. **Calendar**: Integrate scheduling service for The Secretary
4. **Ongoing**: Update this roadmap as we learn 4. **Ongoing**: Add more household staff as needed
--- ---
**Document Status**: Active planning document **Document Status**: Active planning document
**Created**: 2025-12-06 **Created**: 2025-12-06
**Last Updated**: 2025-12-06 **Last Updated**: 2025-12-13
-535
View File
@@ -1,535 +0,0 @@
# Phase 2 Completion Summary: The Steward
**Status**: ✅ COMPLETE
**Completed**: 2025-12-07
**Duration**: 1 day (accelerated from 7-week plan)
**Test Coverage**: 223 passing tests (99.5% pass rate)
---
## Executive Summary
Phase 2 successfully implements **The Steward** - a first-tier LLM agent that creates a two-tier architecture for intelligent request routing. The Steward analyzes incoming requests, identifies relevant household capabilities, and provides scoped tool recommendations to Tatlock (the Butler).
This architecture prevents cognitive overload by ensuring Tatlock only sees tools relevant to each specific request, while maintaining full conversation context awareness and providing complete observability through benchmarking and logging.
---
## Delivered Features
### 1. The Steward Agent ✅
**Location**: `src/agents/steward/`
- **Request Analysis**: Analyzes user requests with full conversation history
- **Capability Recommendation**: Recommends relevant household tools/capabilities
- **Context Awareness**: Identifies references to previous conversation turns
- **Complexity Assessment**: Estimates request complexity (simple/moderate/complex)
- **Missing Capability Detection**: Explicitly states when needed tools are unavailable
- **VRAM Efficiency**: Uses same Ollama model as Tatlock (mistral-nemo:latest)
**Key Files**:
- `agent.py`: Steward PydanticAI agent implementation
- `schemas.py`: `StewardRecommendation` and `ConversationContext` structures
- `service.py`: Service layer with logging and benchmarking
### 2. Household Registry ✅
**Location**: `src/core/household_registry.py`
- **Centralized Capability Management**: Single source of truth for household tools
- **Executive Summaries**: High-level capability descriptions for Steward/Butler coordination
- **PydanticAI Toolsets**: Native toolset composition and scoping
- **Domain Organization**: Tools organized by household member (e.g., `tatlock_core`)
- **Dynamic Tool Scoping**: Creates combined toolsets based on recommendations
**Architecture**:
```
HouseholdRegistry
├─ HouseholdMember (tatlock_core)
│ ├─ HouseholdCapability (summary)
│ └─ FunctionToolset (calculator, datetime, search)
├─ Future: HouseholdMember (librarian)
└─ Future: HouseholdMember (developer)
```
### 3. Request Preprocessing Pipeline ✅
**Location**: `src/core/preprocessing.py`
**4-Phase Flow**:
1. **Steward Analysis**: Analyzes request with full conversation history
2. **Tool Scoping**: Creates combined toolset from recommendations
3. **Note Formatting**: Prepares Steward note for Butler (invisible to user)
4. **Enrichment**: Returns `EnrichedRequest` with all context
**Integration**: Fully integrated with Responses API via `create_response_with_steward()`
### 4. Tool Usage Tracking ✅
**Location**: `src/core/tool_tracking.py`
**Capabilities**:
- Tracks recommended vs. actual tool usage
- Logs unexpected tool calls (not recommended but used)
- Logs unused recommendations (recommended but not used)
- Records timing data for each tool call
- Stores benchmarks to Redis for analysis
**Metrics Supported**:
- Precision: Recommended and used / All recommendations
- Recall: Recommended and used / All tool calls
- F1 Score: Harmonic mean of precision and recall
### 5. Streaming Transparency ✅
**Location**: `src/responses/streaming.py`
**Features**:
- Streams Steward's analysis first (reasoning summary deltas)
- Streams Tatlock's response second (output text deltas)
- Full SSE support with proper event types
- Conversation context visible in stream
- Missing capabilities warnings included
**Event Sequence**:
```
1. response.reasoning_summary_text.delta (Steward analysis)
2. response.reasoning_summary_text.done
3. response.output_text.delta (Tatlock response)
4. response.output_text.done
5. response.done (final response)
```
### 6. Structured Logging ✅
**Location**: `src/core/logging_config.py`
**Features**:
- JSON-formatted structured logging via `structlog`
- Operation timing via context managers (`log_operation`)
- Metadata enrichment for debugging
- Integrated with benchmark recording
- Machine-parseable output for analysis
### 7. Redis Benchmark Storage ✅
**Location**: `src/core/benchmarks.py`
**Features**:
- Cross-session performance metrics storage
- Time-series data with 30-day automatic expiry
- Operations tracked: `steward_analysis`, `tool_call`
- Queryable by operation type, time range, metadata
- Supports accuracy analysis (recommended vs. used)
**Benchmark Schema**:
- Timestamp, operation, duration, success/failure
- Steward-specific: recommendation_count, complexity
- Tool-specific: tool_name, was_recommended, was_actually_used
- Context: conversation_id, metadata dict
### 8. Benchmark Analysis Tools ✅
**Location**: `scripts/benchmark_analysis.py`
**CLI Features**:
```bash
# Steward performance over last 24 hours
python scripts/benchmark_analysis.py --operation steward_analysis --hours 24
# Tool recommendation accuracy over last 7 days
python scripts/benchmark_analysis.py --tool-accuracy --days 7
# Summary of all operations
python scripts/benchmark_analysis.py --summary --hours 1
```
**Metrics Provided**:
- Average Steward latency (target: < 2s)
- Success rate percentage
- Recommendation count distribution
- Complexity distribution
- Tool-specific accuracy (precision/recall/F1)
- Per-tool usage patterns
---
## Architecture
### Request Flow
```
User Request
Responses API (FastAPI)
┌─────────────────────────────────────────────┐
│ Preprocessing Pipeline │
│ ├─ Steward Agent │
│ │ ├─ Receives: Full conversation history │
│ │ ├─ Analyzes: Context + requirements │
│ │ ├─ Queries: Household registry │
│ │ └─ Returns: StewardRecommendation │
│ │ │
│ ├─ Create Scoped Toolset │
│ │ └─ CombinedToolset from capabilities │
│ │ │
│ └─ Format Steward Note │
│ └─ Context summary for Butler │
└─────────────────────────────────────────────┘
Tatlock Agent (Butler)
├─ Receives: Enriched request + note
├─ Tools: ONLY scoped recommendations
├─ Tracking: Tool usage monitored
└─ Context: Full conversation history
Response to User
├─ Steward's reasoning (streamed first)
└─ Tatlock's response (streamed second)
Background:
└─ Redis: Benchmarks + metrics
```
### Two-Tier Abstraction
**Tier 1: Executive Summaries (Steward/Butler coordination)**
```python
HouseholdCapability(
name="tatlock_core",
role="Butler's Core Tools",
category="core",
description="Mathematical calculation, date/time operations, web search",
domains=["computation", "information", "datetime"],
cost="low",
requires_network=True
)
```
**Tier 2: Implementation Details (Tool execution)**
```python
FunctionToolset containing:
- calculate(expression: str) -> str
- get_current_datetime(format_str: str) -> str
- calculate_time_offset(offset: str) -> str
- time_difference(date1: str, date2: str) -> str
- search_web(query: str, num_results: int) -> str
```
---
## Test Coverage
### Test Statistics
- **Total Tests**: 223 (219 passing, 1 pre-existing failure unrelated to Phase 2)
- **Pass Rate**: 99.5%
- **Coverage**: 77.6% overall
### Test Categories
#### Unit Tests ✅
- **Household Registry** (12 tests): Registration, retrieval, toolset composition
- **Steward Schemas** (11 tests): Data structures, formatting
- **Steward Service** (9 tests): Request analysis, context detection, capabilities
- **Preprocessing** (6 tests via integration): Request enrichment, tool scoping
#### Integration Tests ✅
- **Steward → Tatlock Flow** (6 tests):
- Simple math request
- Conversation history propagation
- No capabilities needed (conversational)
- Tool tracker integration
- Missing capabilities warning
- Conversation ID propagation
- **Streaming Integration** (4 tests):
- Basic streaming with Steward
- Conversation history in streaming
- Reasoning contains Steward analysis
- Missing capabilities in stream
### Key Test Files
- `tests/agents/steward/test_steward_schemas.py`
- `tests/agents/steward/test_steward_service.py`
- `tests/integration/test_steward_tatlock_integration.py`
- `tests/integration/test_steward_streaming.py`
---
## Technical Achievements
### 1. PydanticAI Native Patterns ✅
- `FunctionToolset` for tool grouping
- `CombinedToolset` for dynamic composition
- Decorator-based tool registration (`@agent.tool`)
- Structured outputs via Pydantic models (`StewardRecommendation`)
- Dependency injection for tracking (`RunContext[ToolCallTracker]`)
### 2. Tool Scoping Enforcement ✅
- Compile-time scoping via toolset creation
- Tools not even visible to LLM if not recommended
- Fresh agent instances with scoped tools only
- No runtime permission checks needed
### 3. Conversation Context Awareness ✅
- Steward sees FULL conversation history
- Identifies references to previous turns
- Provides contextual notes to Butler
- Example: "User mentioned Python debugging in turn 3"
### 4. Plain Text Approach ✅
- Steward returns natural language analysis
- Service layer parses for structured data
- Keyword extraction for capabilities
- Pattern matching for complexity and context
### 5. Observability ✅
- Structured logging for all operations
- Benchmark recording to Redis
- Tool usage tracking (recommended vs. actual)
- Cross-session performance analysis
---
## Performance Characteristics
### Latency (Estimated)
- **Steward Analysis**: ~1-2 seconds (single LLM call)
- **Tatlock Execution**: ~2-5 seconds (depends on tool usage)
- **Total Added Overhead**: ~1-2 seconds vs. direct Tatlock call
- **Streaming Transparency**: Steward reasoning visible immediately
### Resource Usage
- **VRAM**: Same model for both agents (mistral-nemo:latest)
- **Model Loading**: No additional model loads (efficient!)
- **Redis**: Minimal (benchmarks with 30-day expiry)
- **Network**: Only when web search tools used
### Accuracy Targets
- **Recommendation Precision**: > 90% (tools recommended and actually used)
- **Recommendation Recall**: > 90% (tools used were recommended)
- **False Positives**: < 10% (recommended but not used)
- **False Negatives**: < 10% (used but not recommended)
*Note: Actual metrics available via `scripts/benchmark_analysis.py` after production usage*
---
## Files Created
### Core Implementation
1. `src/core/household_registry.py` - Capability management
2. `src/core/preprocessing.py` - Request preprocessing pipeline
3. `src/core/tool_tracking.py` - Tool usage tracking
4. `src/core/logging_config.py` - Structured logging (M1)
5. `src/core/benchmarks.py` - Redis benchmark storage (M1)
### Steward Agent
6. `src/agents/steward/agent.py` - Steward PydanticAI agent
7. `src/agents/steward/schemas.py` - Data structures
8. `src/agents/steward/service.py` - Service layer
### Tatlock Core Organization
9. `src/agents/tatlock_core/tools.py` - Tool implementations (reorganized)
10. `src/agents/tatlock_core/toolset.py` - PydanticAI toolset
11. `src/agents/tatlock_core/capability.py` - Registry integration
### Tests
12. `tests/agents/steward/test_steward_schemas.py` - Schema tests
13. `tests/agents/steward/test_steward_service.py` - Service tests
14. `tests/integration/test_steward_tatlock_integration.py` - Full flow tests
15. `tests/integration/test_steward_streaming.py` - Streaming tests
### Tools & Documentation
16. `scripts/benchmark_analysis.py` - Performance analysis CLI
17. `PHASE2_PLAN.md` - Detailed implementation plan
18. `PHASE2_COMPLETE.md` - This completion summary
### Modified Files
- `src/agents/tatlock.py` - Added `run_with_scoped_tools()` method
- `src/responses/service.py` - Added `create_response_with_steward()`
- `src/responses/router.py` - Steward routing logic
- `src/responses/streaming.py` - Added `stream_response_with_steward()`
- `CHANGELOG.md` - Phase 2 documentation
---
## Success Metrics
### Technical ✅
- ✅ Household registry operational with executive summaries
- ✅ Steward produces structured recommendations
- ✅ Steward analyzes full conversation context
- ✅ Tool scoping enforced (Tatlock can't use non-recommended tools)
- ✅ Model efficiency preserved (no reload delays)
- ✅ Performance benchmarks recorded to Redis
- ✅ Tool usage tracking (recommended vs. actual)
- ✅ Streaming transparency implemented
### Observability ✅
- ✅ Structured logging (JSON format)
- ✅ Benchmark analysis tools available
- ✅ Tool recommendation accuracy measurable
- ✅ Cross-session performance trends visible
### Architectural ✅
- ✅ PydanticAI patterns followed (Toolsets, decorators, structured outputs)
- ✅ Clean separation: registry vs. agents vs. tools
- ✅ Two-tier abstraction working (summaries vs. details)
- ✅ Future-proof for expert agents (Phase 4)
### Testing ✅
- ✅ 223 tests passing (99.5% pass rate)
- ✅ Integration tests for full flow
- ✅ Streaming integration tests
- ✅ 77.6% test coverage maintained
---
## Usage Examples
### Non-Streaming Request
```python
from src.responses.service import create_response_with_steward
from src.responses.schemas import ResponseRequest
request = ResponseRequest(
model="tatlock",
input=[
{"role": "user", "content": "What's sqrt(144)?"}
],
metadata={"conversation_id": "conv_123"}
)
response = await create_response_with_steward(request)
# Response includes:
# 1. Steward's analysis (reasoning output)
# 2. Tatlock's answer (message output)
```
### Streaming Request
```python
from src.responses.streaming import StreamingCoordinator
coordinator = StreamingCoordinator()
async for event in coordinator.stream_response_with_steward(request):
if event.event == "response.reasoning_summary_text.delta":
print(f"Steward: {event.delta}", end="")
elif event.event == "response.output_text.delta":
print(f"Tatlock: {event.delta}", end="")
elif event.event == "response.done":
print(f"\nFinal response: {event.response.id}")
```
### Benchmark Analysis
```bash
# View Steward performance
python scripts/benchmark_analysis.py --operation steward_analysis --hours 24
# Analyze tool accuracy
python scripts/benchmark_analysis.py --tool-accuracy --days 7
# Get summary
python scripts/benchmark_analysis.py --summary --hours 1
```
---
## Future-Proofing for Phase 4
### Expert Agent Pattern (Ready to Use)
When adding The Librarian, The Developer, or other expert agents:
```
src/agents/librarian/
├── agent.py # Librarian PydanticAI agent
├── tools.py # Research, wiki, knowledge tools
├── toolset.py # PydanticAI toolset
└── capability.py # Registry integration
```
**Registration**:
```python
from src.core.household_registry import get_household_registry
registry = get_household_registry()
registry.register(
name="librarian",
capability=LIBRARIAN_CAPABILITY,
toolset=librarian_toolset,
agent=librarian_agent # For delegation
)
```
**Delegation from Tatlock** (Phase 4):
```python
@tatlock_agent.tool
async def consult_librarian(
ctx: RunContext[None],
research_query: str
) -> str:
"""Consult the Librarian for research assistance."""
return await librarian_agent.run(research_query, usage=ctx.usage)
```
---
## Lessons Learned
### What Went Well
1. **PydanticAI Integration**: Native toolset patterns work beautifully
2. **Two-Tier Architecture**: Clean separation between coordination and execution
3. **Plain Text Approach**: More flexible than structured output for Steward
4. **Test Coverage**: Comprehensive integration tests caught edge cases early
5. **Streaming**: SSE events provide excellent real-time transparency
### Challenges Overcome
1. **Schema vs. Agent OutputItems**: Fixed `_calculate_usage` to handle both types
2. **Registry Initialization**: Added fixtures to ensure registry available in tests
3. **Plain Text Parsing**: Keyword extraction works well but needs careful test mocking
4. **Complexity Substring Matching**: "Complexity:" contains "complex" - fixed test mocks
### Optimizations
1. **Single Model**: Using same Ollama model for both agents saves VRAM
2. **Sequential Execution**: No parallel LLM calls needed (Steward → Tatlock)
3. **Tool Scoping**: Fresh agent instances more reliable than runtime filtering
4. **Benchmark Expiry**: 30-day TTL prevents Redis bloat
---
## Next Steps
### Immediate
- Monitor Steward accuracy in production
- Collect real-world benchmarks
- Iterate on Steward prompt based on metrics
### Phase 3 (Optional)
- Web search delegation to The Librarian
- Enhanced research capabilities
- Multi-source information synthesis
### Phase 4
- Expert agent delegation (Librarian, Developer, etc.)
- Dynamic agent selection based on request
- Cross-agent collaboration patterns
---
## Conclusion
Phase 2 successfully delivers a production-ready two-tier architecture with The Steward managing intelligent request routing and tool scoping. The implementation is:
-**Complete**: All planned features delivered
-**Tested**: 223 tests with 99.5% pass rate
-**Observable**: Full logging and benchmarking
-**Efficient**: Single model, minimal overhead
-**Extensible**: Ready for expert agents in Phase 4
The Steward provides intelligent capability coordination while maintaining conversation context awareness, creating a foundation for scalable multi-agent collaboration in future phases.
**Phase 2 Status**: ✅ **COMPLETE**
---
**Document Version**: 1.0
**Created**: 2025-12-07
**Author**: Development Team
**Reference**: [PHASE2_PLAN.md](PHASE2_PLAN.md)
-865
View File
@@ -1,865 +0,0 @@
# Phase 2 Implementation Plan: The Steward
**Status**: Active Planning
**Created**: 2025-12-07
**Estimated Duration**: 4-5 weeks
**Goal**: Implement first-tier request analysis and household capability coordination
---
## Executive Summary
Phase 2 introduces **The Steward** - a first-tier LLM agent that analyzes incoming requests, identifies relevant household capabilities, and provides focused recommendations to Tatlock (the Butler). This creates a two-tier architecture that prevents cognitive overload and enables efficient tool/agent coordination.
### Key Deliverables
1. **Household Registry**: Centralized capability catalog with PydanticAI Toolsets
2. **Steward Agent**: Request analyzer with conversation context awareness
3. **Tool Scoping**: Dynamic toolset creation based on recommendations
4. **Observability**: Performance benchmarking and tool usage tracking via Redis
5. **Integration**: Full Steward → Tatlock request flow
---
## Core Architectural Principles
### 1. Household-Based Organization
- Each expert agent owns their tools in a domain directory
- Tools organized as functional clusters around capabilities
- Example: `src/agents/tatlock_core/` contains calculator, datetime, web search
### 2. Two-Tier Capability Abstraction
- **Executive Summary**: High-level capabilities for Steward/Butler coordination
- **Implementation Details**: Full tool specifications for household members
- Steward sees summaries, household members see full details
### 3. PydanticAI Native Patterns
- Use `FunctionToolset` and `CombinedToolset` for composition
- Decorator-based tool registration (`@agent.tool`)
- Structured outputs via Pydantic models
- Agent delegation pattern for expert agents (Phase 4)
### 4. Separate Registries
- **Household Registry**: Tools + capabilities (new in Phase 2)
- **Model Registry**: Agents/models (existing from Phase 1)
- Clean separation of concerns
### 5. Start Minimal
- Only 3 core Tatlock tools initially: calculator, datetime, web search
- No new tools until expert agents exist (Phase 4)
- Prove the pattern before expanding
---
## Implementation Milestones
### Milestone 1: Household Registry + Logging Infrastructure (Week 1-2)
#### Goal
Create a registry system that aggregates household capabilities using PydanticAI Toolsets and establish observability infrastructure.
#### Tasks
**1.1 Create Household Registry Module**
Location: `src/core/household_registry.py`
```python
from pydantic import BaseModel
from pydantic_ai import FunctionToolset, CombinedToolset
class HouseholdCapability(BaseModel):
"""Executive summary of a household member's capabilities."""
name: str # "tatlock_core", "librarian", "developer"
role: str # "Butler's Core Tools", "The Librarian"
category: str # "core", "research", "technical"
description: str # One-sentence description
domains: list[str] # ["computation", "information", "datetime"]
cost: str # "low", "medium", "high"
requires_network: bool
class HouseholdMember(BaseModel):
"""Full specification of a household member."""
capability: HouseholdCapability
toolset: FunctionToolset
agent: Agent | None = None # For expert agents in Phase 4
class HouseholdRegistry:
"""Registry of household capabilities and implementations."""
def __init__(self):
self._members: dict[str, HouseholdMember] = {}
def register(
self,
name: str,
capability: HouseholdCapability,
toolset: FunctionToolset,
agent: Agent | None = None
):
"""Register a household member."""
self._members[name] = HouseholdMember(
capability=capability,
toolset=toolset,
agent=agent
)
def get_all_capabilities(self) -> list[HouseholdCapability]:
"""Get executive summaries for Steward/Butler."""
return [m.capability for m in self._members.values()]
def get_scoped_toolset(self, names: list[str]) -> CombinedToolset:
"""Create combined toolset from recommended capabilities."""
toolsets = [self._members[name].toolset for name in names]
return CombinedToolset(toolsets)
# Global registry instance
household_registry = HouseholdRegistry()
```
**1.2 Reorganize Tatlock Core Tools**
Create domain-based organization:
```
src/agents/tatlock_core/
├── __init__.py
├── tools.py # Tool implementations (moved from src/agents/tools.py)
├── toolset.py # PydanticAI toolset registration
└── capability.py # Executive summary for registry
```
**1.3 Create Logging Infrastructure**
Location: `src/core/logging_config.py`
- Structured logging with `structlog`
- JSON format for machine parsing
- Operation timing and metadata tracking
- Context manager for automatic timing
**1.4 Create Redis Benchmark Storage**
Location: `src/core/benchmarks.py`
Features:
- Performance benchmark recording (Steward analysis, tool calls)
- Cross-session persistence via Redis
- Time-series storage with automatic expiry (30 days)
- Queryable metrics for analysis
Benchmark schema:
```python
class PerformanceBenchmark(BaseModel):
timestamp: datetime
operation: str # "steward_analysis", "tool_call"
duration_seconds: float
success: bool
# Steward-specific
recommendation_count: Optional[int]
confidence: Optional[float]
# Tool-specific
tool_name: Optional[str]
was_recommended: Optional[bool]
was_actually_used: Optional[bool]
# Context
conversation_id: Optional[str]
metadata: dict
```
**1.5 Testing**
- Test household registry registration and retrieval
- Test Toolset composition
- Test benchmark recording to Redis
- Test structured logging output
#### Success Criteria
- ✅ Household registry operational
- ✅ Tatlock core tools organized in domain directory
- ✅ Redis benchmarks working
- ✅ Structured logging functional
- ✅ Tests pass and maintain 80%+ coverage
---
### Milestone 2: Minimal Steward Agent with Context Analysis (Week 3-4)
#### Goal
Create a Steward agent that analyzes requests with full conversation context and recommends relevant household capabilities.
#### Tasks
**2.1 Create Steward Agent**
Location: `src/agents/steward/agent.py`
Structured output schema:
```python
class ConversationContext(BaseModel):
"""Contextual information from conversation history."""
has_previous_context: bool
relevant_turns: list[int] # 0-indexed turn numbers
context_summary: str # Summary for Butler
class StewardRecommendation(BaseModel):
"""Structured recommendation from Steward analysis."""
recommended_capabilities: list[str]
reasoning: str
estimated_complexity: Literal["simple", "moderate", "complex"]
conversation_context: ConversationContext
missing_capabilities: Optional[str] = None
```
Key features:
- Uses same model as Tatlock (`ollama:mistral-nemo`) for VRAM efficiency
- Receives FULL conversation history
- Queries household registry via tool
- Conservative recommendations (avoid over-inclusion)
- Explicit handling of missing capabilities
**2.2 Steward System Prompt**
Responsibilities:
1. **Capability Recommendation**: Query registry, recommend only necessary tools
2. **Conversation Analysis**: Identify references to previous topics
3. **Complexity Assessment**: Simple/moderate/complex classification
4. **Missing Capability Detection**: Suggest what's needed if no tools available
**2.3 Steward Service Layer with Logging**
Location: `src/agents/steward/service.py`
```python
async def analyze_request(
user_request: str,
conversation_history: list[dict] # FULL conversation
) -> StewardRecommendation:
"""Analyze request with full conversation context."""
async with log_operation("steward_analysis", {...}) as log_ctx:
result = await steward_agent.run(
user_request,
message_history=convert_to_pydantic_history(conversation_history),
usage_limits=UsageLimits(request_limit=3)
)
# Log and benchmark
log_ctx["recommendation_count"] = len(result.data.recommended_capabilities)
await benchmark_store.record(...)
return result.data
```
**2.4 Testing**
Test scenarios:
- Calculator request → recommends tatlock_core
- Simple greeting → recommends []
- Web search request → recommends tatlock_core
- Request referencing previous turn → identifies context
- Impossible request → returns missing_capabilities
#### Success Criteria
- ✅ Steward queries household registry successfully
- ✅ Produces structured recommendations
- ✅ Analyzes full conversation context
- ✅ Handles missing capabilities gracefully
- ✅ Conservative recommendations (> 90% accuracy)
- ✅ Benchmarks recorded to Redis
---
### Milestone 3: Request Preprocessing & Tool Tracking (Week 5-6)
#### Goal
Wire Steward into request flow, implement tool scoping, and track tool usage.
#### Tasks
**3.1 Create Preprocessing Pipeline**
Location: `src/core/preprocessing.py`
```python
@dataclass
class EnrichedRequest:
"""Request enriched with Steward's analysis."""
original_request: str
steward_note: str # Formatted note for Tatlock
scoped_toolset: CombinedToolset # Only recommended tools
recommendation: StewardRecommendation
steward_reasoning_output: str # For streaming to user
async def preprocess_request(
user_request: str,
conversation_history: list[dict] # FULL conversation
) -> EnrichedRequest:
"""Analyze via Steward and prepare scoped context."""
# Call Steward with full conversation
recommendation = await analyze_request(user_request, conversation_history)
# Format note to Tatlock (includes conversation context)
steward_note = format_steward_note(recommendation)
# Create scoped toolset
scoped_toolset = household_registry.get_scoped_toolset(
recommendation.recommended_capabilities
)
return EnrichedRequest(...)
```
Note formatting:
- Includes conversation context summary
- Highlights missing capabilities if applicable
- Provides complexity estimate
**3.2 Tool Usage Tracking**
Location: `src/core/tool_tracking.py`
```python
class ToolCallTracker:
"""Tracks tool calls for benchmarking."""
def __init__(self, recommended_tools: list[str]):
self.recommended_tools = set(recommended_tools)
self.actual_calls: dict[str, list[float]] = {}
async def track_call(self, tool_name: str, duration: float):
"""Record a tool call with timing."""
# Log if tool wasn't recommended
if tool_name not in self.recommended_tools:
logger.warning("tool_call_not_recommended", ...)
# Record benchmark to Redis
await benchmark_store.record(...)
async def finalize(self):
"""Log unused recommended tools."""
unused = self.recommended_tools - set(self.actual_calls.keys())
# Record benchmarks for unused tools
```
**3.3 Integrate with Responses API**
Modify `src/responses/service.py`:
```python
async def generate_response(request: ResponseRequest) -> ResponseOutput:
# Preprocess via Steward (with full conversation)
enriched = await preprocess_request(
user_message,
conversation_history=request.input[:-1]
)
# Run Tatlock with scoped tools and tracker
result = await run_tatlock_with_scoped_tools(
enriched.original_request,
enriched.steward_note,
enriched.scoped_toolset,
enriched.recommendation.recommended_capabilities, # For tracking
message_history,
usage_tracker
)
# Build response with Steward reasoning
return build_response_with_steward_reasoning(...)
```
**3.4 Update Tatlock Agent**
Location: `src/agents/tatlock.py`
```python
async def run_tatlock_with_scoped_tools(
user_request: str,
steward_note: str,
scoped_toolset: CombinedToolset,
recommended_tools: list[str],
message_history: list[dict],
usage: UsageeLimits
):
# Initialize tracker
tracker = ToolCallTracker(recommended_tools)
# Prepend Steward's note (invisible to user, visible to Tatlock)
enriched_prompt = f"{steward_note}\n\n{user_request}"
# Run with ONLY scoped tools
result = await tatlock_agent.run(
enriched_prompt,
message_history=convert_to_pydantic_history(message_history),
toolsets=[scoped_toolset], # Tool scoping enforced
deps=tracker, # For tracking
usage=usage
)
# Finalize tracking
await tracker.finalize()
return result
```
**3.5 Add Streaming Transparency**
Modify `src/responses/streaming.py`:
- Stream Steward's reasoning first
- Then stream Tatlock's response
- Include conversation context notes
- Format missing capabilities warnings
**3.6 Testing**
Integration tests:
- Full Steward → Tatlock flow
- Tool scoping enforcement (can't use non-recommended tools)
- Tool usage tracking (recommended vs. actual)
- Conversation context propagation
- Missing capabilities handling
#### Success Criteria
- ✅ Full request flow working (User → Steward → Tatlock)
- ✅ Steward reasoning visible in output stream
- ✅ Tool scoping enforced (only recommended tools available)
- ✅ Tool usage tracked and logged to Redis
- ✅ Conversation context passed through pipeline
- ✅ Integration tests pass end-to-end
---
### Milestone 4: Testing, Benchmarking & Refinement (Week 7)
#### Goal
Validate the system, optimize performance, refine prompts, and establish monitoring.
#### Tasks
**4.1 Comprehensive Testing**
Test categories:
- End-to-end integration tests (full request flow)
- Performance benchmarks (latency targets)
- Prompt refinement (recommendation accuracy)
- Edge cases (errors, timeouts, missing capabilities)
- Conversation context accuracy
**4.2 Performance Validation**
Targets:
- Steward analysis: < 2 seconds
- Total added latency: < 3 seconds
- Model stays hot in VRAM (no reload delays)
- Tool recommendation accuracy: > 90%
**4.3 Benchmark Analysis Tools**
Create `scripts/benchmark_analysis.py`:
```bash
# View Steward performance over last 24 hours
python scripts/benchmark_analysis.py --operation steward_analysis --hours 24
# Analyze tool recommendation accuracy
python scripts/benchmark_analysis.py --tool-accuracy --days 7
```
Metrics to track:
- Average Steward analysis time
- Recommendation count distribution
- Tool accuracy (recommended & used, recommended but unused, not recommended but used)
- Recommendation precision percentage
**4.4 Prompt Engineering**
Iterate on Steward system prompt:
- Test with diverse request types
- Tune conservativeness (balance false positives/negatives)
- Validate conversation context analysis
- Test missing capability detection
**4.5 Documentation**
Update documentation:
- README.md: Steward explanation and examples
- AGENTS.md: Household registration pattern
- IMPLEMENTATION_ROADMAP.md: Mark Phase 2 complete
- Add benchmark analysis guide
#### Success Criteria
- ✅ < 3 seconds added latency for Steward analysis
- ✅ > 90% recommendation accuracy (manual evaluation)
- ✅ All integration tests pass
- ✅ Benchmark tools functional
- ✅ Documentation complete and accurate
- ✅ Ready for Phase 3/4 (expert agents)
---
## Architecture Diagram
```
User Request
Orchestrator (FastAPI)
Preprocessing Pipeline
├─→ Steward Agent
│ ├─ Receives: FULL conversation history
│ ├─ Analyzes: Context, references, requirements
│ ├─ Queries: Household registry (capabilities)
│ ├─ Outputs: StewardRecommendation
│ │ ├─ recommended_capabilities: list[str]
│ │ ├─ conversation_context: ConversationContext
│ │ ├─ missing_capabilities: str | None
│ │ └─ reasoning: str
│ └─ Logs: Performance benchmarks → Redis
├─→ Create Scoped Toolset
│ └─ CombinedToolset from recommended capabilities
└─→ Format Steward Note
└─ Includes conversation context for Tatlock
Tatlock Agent (with scoped tools)
├─ Receives: Enriched request + Steward note
├─ Has access to: ONLY recommended tools
├─ Tool calls tracked: ToolCallTracker
└─ Logs: Tool usage benchmarks → Redis
Response to User
├─ Steward's reasoning (streamed first)
└─ Tatlock's response (streamed second)
Background:
└─ Redis: Performance benchmarks, tool usage analysis
```
---
## Design Decisions Summary
### 1. Logging & Performance Benchmarks
**Decision**: Full observability with Redis-backed benchmark storage
**Rationale**:
- Track Steward recommendations vs. Tatlock's actual tool usage
- Measure performance metrics (latency, token usage)
- Cross-session analysis for optimization
- Identify recommendation accuracy over time
### 2. Steward Fallback Behavior
**Decision**: Explicit missing capability communication
**Rationale**:
- No suitable tools → Steward states "missing capabilities" with description
- Can suggest what type of tool would be helpful
- Code errors → standard exception handlers (don't suppress real errors)
- Better UX than silent failures or defaulting to all tools
### 3. Conversation History for Steward
**Decision**: Steward sees FULL conversation, not just current turn
**Rationale**:
- Can identify references to previous topics
- Provides contextual notes to Butler
- "Two sets of eyes" on conversation
- Example: "User mentioned Python debugging in turn 3, relevant details: async code"
### 4. Registry Pattern
**Decision**: Separate Household Registry from Model Registry
**Rationale**:
- Tools belong to household members, not models
- Clean separation of concerns
- Executive summaries for coordination, details for execution
### 5. Tool Composition
**Decision**: PydanticAI FunctionToolset + CombinedToolset
**Rationale**:
- Native PydanticAI pattern
- Clean composition and filtering
- Dynamic scoping per request
### 6. Tool Scoping
**Decision**: Compile-time scoping via toolset creation
**Rationale**:
- Tools not even visible to LLM
- Cleaner than runtime permission checks
- Enforced at PydanticAI level
### 7. Organization
**Decision**: Domain-based household directories
**Rationale**:
- Each household member owns their tools
- Clear bounded contexts
- Example: `src/agents/tatlock_core/`, `src/agents/librarian/` (future)
---
## Infrastructure Requirements
### Redis Setup
Development (quick start):
```bash
# Docker (recommended)
docker run -d -p 6379:6379 --name tatlock-redis redis:7-alpine
# Or local installation
# macOS: brew install redis && brew services start redis
# Linux: sudo apt install redis-server && sudo systemctl start redis
```
Production (docker-compose.yml):
```yaml
services:
redis:
image: redis:7-alpine
ports:
- "6379:6379"
volumes:
- redis_data:/data
command: redis-server --appendonly yes
volumes:
redis_data:
```
### Dependencies Update
Add to `requirements.txt`:
```txt
redis[hiredis]>=5.0.0,<6.0.0
structlog>=24.1.0,<25.0.0
```
### Configuration
Add to `.env`:
```env
# Redis Configuration
REDIS_URL=redis://localhost:6379/1
# Logging
LOG_LEVEL=INFO
LOG_FORMAT=json
ENABLE_BENCHMARKS=true
```
---
## Timeline
**Week 1-2**: Household Registry + Logging Infrastructure
- Household registry with Toolsets
- Structured logging with structlog
- Redis benchmark storage
- Tatlock core reorganization
- Tests: Registry + benchmarking
**Week 3-4**: Steward Agent with Context Analysis
- Steward agent with conversation context
- ConversationContext in recommendations
- Missing capabilities handling
- Tests: Context analysis, missing capabilities
**Week 5-6**: Integration + Tool Tracking
- Request preprocessing with full conversation
- Tool usage tracking middleware
- Scoped toolset creation
- Streaming transparency
- Tests: Full flow + tool tracking
**Week 7**: Testing, Benchmarking & Refinement
- End-to-end integration tests
- Benchmark analysis tools
- Prompt refinement
- Performance validation
- Documentation updates
**Total: 4-5 weeks** (core implementation complete in 6 weeks, polish in week 7)
---
## Success Metrics
### Technical
- ✅ Household registry operational with executive summaries
- ✅ Steward produces accurate recommendations (> 90%)
- ✅ Steward analyzes full conversation context
- ✅ Tool scoping enforced (Tatlock can't use non-recommended tools)
- ✅ Model efficiency preserved (no reload delays)
- ✅ Added latency < 3 seconds
- ✅ Performance benchmarks recorded to Redis
- ✅ Tool usage tracking (recommended vs. actual)
### Observability
- ✅ Structured logging (JSON format)
- ✅ Benchmark analysis tools available
- ✅ Tool recommendation accuracy measurable
- ✅ Cross-session performance trends visible
### Error Handling
- ✅ Missing capabilities explicitly communicated
- ✅ Steward can guide user toward needed resources
- ✅ Code errors properly surfaced (not suppressed)
### Architectural
- ✅ PydanticAI patterns followed (Toolsets, decorators, structured outputs)
- ✅ Clean separation: registry vs. agents vs. tools
- ✅ Two-tier abstraction working (summaries vs. details)
- ✅ Future-proof for expert agents (Phase 4)
### Testing
- ✅ Maintain 80%+ test coverage
- ✅ Integration tests for full flow
- ✅ Performance benchmarks established
---
## Future-Proofing for Phase 4
### Expert Agent Pattern (Template)
When adding The Librarian, The Developer, etc., follow this structure:
```
src/agents/librarian/
├── __init__.py
├── agent.py # Librarian PydanticAI agent
├── tools.py # Librarian-specific tools (wiki, research, etc.)
├── toolset.py # PydanticAI toolset creation
└── capability.py # Executive summary for registry
```
Example capability registration:
```python
# capability.py
LIBRARIAN_CAPABILITY = HouseholdCapability(
name="librarian",
role="The Librarian",
category="research",
description="Research assistance, knowledge management, and information synthesis",
domains=["research", "knowledge_base", "documentation"],
cost="medium",
requires_network=True
)
def register_librarian():
household_registry.register(
name="librarian",
capability=LIBRARIAN_CAPABILITY,
toolset=librarian_toolset,
agent=librarian_agent # Expert agent for delegation
)
```
Tatlock delegation pattern (Phase 4):
```python
@tatlock_agent.tool
async def consult_librarian(
ctx: RunContext[None],
research_query: str
) -> str:
"""Consult the Librarian for research assistance."""
from src.agents.librarian.agent import librarian_agent
result = await librarian_agent.run(
research_query,
usage=ctx.usage # Aggregate usage
)
return result.data
```
---
## Risk Mitigation
### Identified Risks
1. **Steward recommendations too broad**
- Mitigation: Conservative prompt engineering, benchmark tracking, iterate based on false positives
2. **Added latency unacceptable**
- Mitigation: Stream Steward reasoning for transparency, optimize prompt, use same base model
3. **Tool registry becomes unwieldy**
- Mitigation: Good categorization, semantic search (future), regular pruning
4. **Model VRAM competition**
- Mitigation: Use same base model for Steward and Tatlock, sequential calls
5. **Redis dependency**
- Mitigation: Make benchmarking optional, graceful degradation if Redis unavailable
---
## Open Questions - RESOLVED
All major design questions have been resolved. See "Design Decisions Summary" section above.
---
## Next Steps
### Immediate (Today/This Week)
1. Set up Redis (Docker or local)
2. Create `src/core/logging_config.py` with structured logging
3. Create `src/core/benchmarks.py` with Redis storage
4. Add `redis` and `structlog` to requirements.txt
5. Create household registry skeleton
### Week 1-2
1. Complete household registry with Toolset integration
2. Reorganize Tatlock core tools into domain directory
3. Implement logging infrastructure
4. Write tests for registry + benchmarking
### Week 3-4
1. Create Steward agent with conversation context
2. Implement missing capabilities handling
3. Test context analysis accuracy
4. Iterate on system prompt
### Week 5-6
1. Build preprocessing pipeline
2. Integrate with Responses API
3. Implement tool tracking
4. Add streaming transparency
### Week 7
1. End-to-end testing
2. Benchmark analysis
3. Performance optimization
4. Documentation updates
---
## Document Status
**Status**: Active Planning Document
**Created**: 2025-12-07
**Last Updated**: 2025-12-07
**Version**: 1.0
**Next Review**: After Milestone 1 completion
---
**Reference Documents**:
- [PHILOSOPHY.md](PHILOSOPHY.md) - System vision and architecture
- [IMPLEMENTATION_ROADMAP.md](IMPLEMENTATION_ROADMAP.md) - Full project roadmap
- [AGENTS.md](AGENTS.md) - Agent development guidelines
- [README.md](README.md) - User documentation
+84 -40
View File
@@ -6,12 +6,25 @@ A privacy-first, offline-capable personal assistant system that coordinates spec
## Current Status ## Current Status
-**Production-ready testing API** with OpenAI Responses API format -**Production-ready API** with OpenAI Responses API format
-**Open WebUI integration** with reasoning bubbles (`<think>` tags) -**Open WebUI integration** with reasoning bubbles (`<think>` tags)
-**Conversation history** with auto-generated IDs and context management -**Two-tier architecture** - The Steward analyzes requests, Tatlock coordinates execution
-**Tatlock PydanticAI Agent** - Real LLM integration with Ollama + permanent tools -**Multi-agent coordination** - Expert household staff for specialized tasks
-**Permanent Tools** - Calculator, date/time toolkit, web search (SearXNG) -**Memory system** - User profile, preferences, and semantic recall
-**Comprehensive testing** - 131 tests, 81.78% coverage -**Comprehensive testing** - 399 tests with good coverage
### The Household Staff
| Agent | Role | Status |
|-------|------|--------|
| **Tatlock** | The Butler - Primary interface with witty personality | ✅ Active |
| **The Steward** | Request analysis and capability recommendation | ✅ Active |
| **The Librarian** | Research, wiki management, knowledge synthesis | ✅ Active |
| **The Biographer** | User memory - profiles, preferences, facts | ✅ Active |
| **The Developer** | Code assistance, debugging, architecture | 🔜 Planned |
| **The Secretary** | Scheduling, calendars, reminders | 🔜 Planned |
| **The Handyman** | System administration, monitoring | 🔜 Planned |
| **The Housekeeper** | Home automation (Home Assistant) | 🔜 Planned |
## Features ## Features
@@ -45,24 +58,27 @@ A privacy-first, offline-capable personal assistant system that coordinates spec
- Error triggers for testing (rate_limit, context_overflow) - Error triggers for testing (rate_limit, context_overflow)
- **Tatlock**: Real PydanticAI agent with butler personality - **Tatlock**: Real PydanticAI agent with butler personality
- **LLM Backend**: Ollama (mistral-nemo:latest) - **LLM Backend**: Ollama (mistral-nemo:latest by default)
- **Personality**: Witty British butler, research-oriented - **Personality**: Witty British butler, research-oriented
- **Permanent Tools**: - **Core Tools**:
- **Calculator**: Safe mathematical expression evaluation (arithmetic, algebra, trigonometry, logarithms) - **Calculator**: Safe mathematical expression evaluation
- **Date/Time Toolkit**: Current time, relative dates ("1 week ago"), time differences - **Date/Time Toolkit**: Current time, relative dates, time differences
- **Web Search**: Privacy-preserving search via SearXNG - **Web Search**: Privacy-preserving search via SearXNG
- **Capabilities**: Streaming, reasoning, tool calling - **Household Coordination**:
- **Phase**: Phase 1 - Basic Integration (full household coordination coming in future phases) - **The Steward**: Analyzes requests and recommends capabilities
- **The Librarian**: Research via library-desk HybridRAG + wiki
- **The Biographer**: User memory and preference management
- **Capabilities**: Streaming, reasoning, tool calling, multi-agent delegation
## Requirements ## Requirements
- Python 3.12+ (Python 3.12.11 recommended) - Python 3.12+ (Python 3.12.11 recommended)
- **Ollama** (for Tatlock agent): Running locally or network-accessible - **External Services** (must be running separately):
- Download: https://ollama.ai/ - **Ollama**: LLM inference (mistral-nemo:latest, nomic-embed-text)
- Model: `ollama pull mistral-nemo:latest` - **Redis**: Caching and session memory
- **SearXNG** (for web search tool): Optional but recommended - **Qdrant**: Vector storage for The Biographer's memory
- Docker: `docker run -d -p 8087:8080 searxng/searxng` - **SearXNG**: Web search (optional)
- Or use public instance (less private) - **library-desk**: Research API for The Librarian (optional)
## Quick Start ## Quick Start
@@ -251,15 +267,18 @@ Interactive documentation available at:
# Run all tests # Run all tests
pytest pytest
# Run unit tests only (no external services needed)
pytest --ignore=tests/e2e --ignore=tests/integration
# Run with coverage # Run with coverage
pytest --cov=src --cov-report=term-missing pytest --cov=src --cov-report=term-missing
# Current: 131 tests, 81.78% coverage # Current: ~400 tests
``` ```
**Test Categories:** **Test Categories:**
- Unit tests: Agent tools, streaming, schemas - Unit tests: Agent tools, capabilities, schemas, memory service
- Integration tests: Full API stack with real Ollama calls - Integration tests: Full API stack with real Ollama
- End-to-end tests: Chat completions, responses API - End-to-end tests: Chat completions, responses API
## Deployment ## Deployment
@@ -291,9 +310,25 @@ API_PORT=8000
# Ollama Configuration # Ollama Configuration
OLLAMA_HOST=http://localhost:11434 OLLAMA_HOST=http://localhost:11434
OLLAMA_DEFAULT_MODEL=mistral-nemo:latest OLLAMA_DEFAULT_MODEL=mistral-nemo:latest
OLLAMA_EMBEDDING_MODEL=nomic-embed-text
OLLAMA_TIMEOUT=120 OLLAMA_TIMEOUT=120
# SearXNG Configuration (for web search tool) # Redis Configuration
REDIS_HOST=localhost
REDIS_PORT=6379
REDIS_MEMORY_DB=2
REDIS_MEMORY_TTL_HOURS=24
# Qdrant Configuration (for memory)
QDRANT_HOST=localhost
QDRANT_PORT=6333
QDRANT_EMBEDDING_DIM=768
# Library-desk Configuration (for The Librarian)
LIBRARY_DESK_HOST=http://localhost:8089
LIBRARY_DESK_TIMEOUT=60
# SearXNG Configuration (for web search)
SEARXNG_HOST=http://localhost:8087 SEARXNG_HOST=http://localhost:8087
SEARXNG_TIMEOUT=30 SEARXNG_TIMEOUT=30
@@ -339,23 +374,32 @@ See `.env.example` for full configuration options.
``` ```
tatlock/ tatlock/
├── src/ ├── src/
│ ├── agents/ # Agent interface and implementations │ ├── agents/ # Agent implementations
│ │ ├── base.py # AgentInterface abstract class │ │ ├── biographer/ # The Biographer - memory management
│ │ ├── lorem_tester.py # Mock agent for testing │ │ ├── librarian/ # The Librarian - research & wiki
│ │ ├── tatlock.py # Real PydanticAI butler agent │ │ ├── steward/ # The Steward - request analysis
│ │ ├── tools.py # Permanent tools (calculator, date/time, search) │ │ ├── tatlock_core/ # Core butler tools
│ │ ── registry.py # Model registry │ │ ── tatlock.py # Tatlock PydanticAI agent
│ ├── responses/ # Responses API (primary endpoint) │ ├── coordination.py # Multi-agent coordination
│ ├── chat/ # Chat Completions wrapper │ ├── delegation.py # Expert delegation wrappers
├── models/ # Models listing │ └── protocol.py # Agent communication protocol
│ ├── core/ # Shared utilities and config │ ├── responses/ # Responses API (primary endpoint)
── main.py # Application entry point ── chat/ # Chat Completions wrapper
├── tests/ # Comprehensive test suite (131 tests) │ ├── models/ # Models listing
├── AGENTS.md # LLM agent development guidelines │ ├── core/ # Shared infrastructure
├── PHILOSOPHY.md # System vision and architecture │ │ ├── config.py # Configuration management
├── IMPLEMENTATION_ROADMAP.md # Development phases │ │ ├── context.py # Request context (ContextVar)
├── CHANGELOG.md # Version history │ │ ├── memory_service.py # Direct memory access
└── README.md # This file ├── memory_cache.py # Redis session cache
│ │ ├── embeddings.py # Ollama embedding client
│ │ ├── qdrant.py # Vector database client
│ │ └── multi_tenancy.py # User isolation utilities
│ └── main.py # Application entry point
├── tests/ # Comprehensive test suite
├── PHILOSOPHY.md # System vision and architecture
├── IMPLEMENTATION_ROADMAP.md # Development phases
├── CHANGELOG.md # Version history
└── README.md # This file
``` ```
## Development ## Development
@@ -388,8 +432,8 @@ For LLM agent development guidelines and architectural decisions, see [AGENTS.md
## Version ## Version
Current version: **0.2.5** - Phase 2: The Steward (Two-Tier Architecture) Current version: **1.3.2** - Biographer tool type hints fix
--- ---
**Note**: This is a production-ready testing API with mock responses. The architecture is designed for easy integration with real LLM backends (PydanticAI, Ollama, OpenAI, etc.). **Note**: Tatlock is a production-ready homelab butler. All household staff use PydanticAI with Ollama for local LLM inference.
-424
View File
@@ -1,424 +0,0 @@
# Library-Desk API Requirements for Tatlock Integration
## Overview
The Librarian agent in Tatlock needs additional endpoints in library-desk to support wiki page editing and content management. Currently, the API provides read operations but The Librarian needs write capabilities for:
- Creating new wiki pages
- Updating existing wiki pages (content, title, tags, description)
## Required Endpoints
### 1. Create Wiki Page (Already Exists)
**Endpoint:** `POST /wiki/pages`
This endpoint already exists and works correctly.
### 2. Update Wiki Page (Needs Enhancement)
**Endpoint:** `PUT /wiki/pages/{page_id}`
**Current Status:** May exist but needs verification that it supports partial updates.
**Required Behavior:**
- Accept partial updates (only provided fields should be updated)
- Support updating: `content`, `title`, `tags`, `description`
- Auto-update vector embeddings after content changes
- Auto-update knowledge graph after content changes
**Request Body:**
```json
{
"content": "# New Content\n\nOptional - only if changing content",
"title": "Optional - only if renaming",
"tags": ["optional", "list", "of", "new", "tags"],
"description": "Optional new description"
}
```
**Query Parameters:**
- `user`: User identifier for multi-tenancy (required)
**Response:**
```json
{
"id": 42,
"path": "/projects/example",
"title": "Updated Title",
"description": "Updated description",
"content": "# New Content...",
"tags": ["updated", "tags"],
"updated_at": "2024-01-15T10:30:00Z"
}
```
**Notes:**
- Should trigger background tasks to re-index vectors and refresh graph entities
- Should validate that user has access to the page (namespace check)
- Should preserve fields that are not provided in the request
## Use Cases for The Librarian
### Adding New Knowledge
When a user says "Add this to the wiki" or "Create a page about X":
- Librarian uses `POST /wiki/pages` to create the page
- Tags are assigned based on context (dossiers)
### Correcting Information
When a user says "Update the page about X" or "Fix this fact":
1. Librarian searches for the page with `GET /wiki/search`
2. Fetches full content with `GET /wiki/pages/{id}`
3. Updates with corrected content via `PUT /wiki/pages/{id}`
### Organizing Knowledge
When a user says "Add this page to the projects dossier":
- Librarian updates just the tags field via `PUT /wiki/pages/{id}`
## Integration Notes
- The Librarian will call these endpoints via HTTP from Tatlock
- Authentication uses Bearer token (LIBRARY_DESK_API_KEY)
- All operations are scoped to the user's namespace
- Background processing (vectors, graph) should not block the response
## Testing Checklist
- [ ] `PUT /wiki/pages/{page_id}` accepts partial updates
- [ ] Updating content triggers vector re-indexing
- [ ] Updating content triggers graph entity extraction
- [ ] Tags can be updated independently of content
- [ ] Description can be updated independently
- [ ] Title can be updated (with path remaining the same)
- [ ] User namespace validation works correctly
===== IMPLEMENTATION INSTRUCTIONS =========
# Librarian Wiki Integration Guide
This document provides implementation instructions for integrating the library-desk wiki endpoints into the Librarian agent (Tatlock).
## Available Endpoints
### 1. Create Wiki Page
**Endpoint:** `POST /wiki/pages`
Use this for simple page creation when the Librarian already has the content.
```python
async def create_wiki_page(
title: str,
path: str,
content: str,
tags: list[str],
description: str = "",
user: str = "default"
) -> dict:
"""Create a new wiki page."""
response = await http_client.post(
f"{LIBRARY_DESK_URL}/wiki/pages",
headers={"Authorization": f"Bearer {LIBRARY_DESK_API_KEY}"},
json={
"title": title,
"path": path,
"content": content,
"tags": tags,
"description": description,
"user": user
}
)
return response.json()
```
**When to use:**
- User provides specific content to add
- Librarian has already composed the content
- Simple note-taking or quick additions
---
### 2. Smart Create Wiki Page (Recommended for Research)
**Endpoint:** `POST /wiki/pages/smart-create`
Use this when the Librarian should research a topic before creating the page. This endpoint:
1. Searches existing wiki, knowledge graph, and web for context
2. Uses LLM to synthesize findings into structured content
3. Creates the page with proper attribution
4. Automatically links entities bidirectionally
```python
async def smart_create_wiki_page(
topic: str,
tags: list[str],
user: str = "default",
path: str | None = None,
include_web_research: bool = True,
include_wiki_search: bool = True
) -> dict:
"""Create a wiki page with HybridRAG research."""
response = await http_client.post(
f"{LIBRARY_DESK_URL}/wiki/pages/smart-create",
headers={"Authorization": f"Bearer {LIBRARY_DESK_API_KEY}"},
json={
"topic": topic,
"path": path, # Optional - auto-generated from topic if not provided
"tags": tags,
"user": user,
"include_web_research": include_web_research,
"include_wiki_search": include_wiki_search
}
)
return response.json()
```
**Response includes:**
```json
{
"page": {
"id": 123,
"path": "/users/jpmschweitzer/technology/docker-orchestration",
"title": "Docker orchestration",
"content": "# Docker Orchestration\n\n...",
"tags": ["technology", "devops"],
"created_at": "2024-01-15T10:30:00Z",
"updated_at": "2024-01-15T10:30:00Z"
},
"research_summary": {
"wiki_results": 3,
"web_results": 8,
"graph_entities": 5,
"keywords_extracted": 12,
"timing_ms": 4500
},
"sources_used": 11,
"search_id": "uuid-for-reference",
"entity_linking": {
"forward_links": 5,
"backward_links": 3,
"pages_updated": 2
}
}
```
**When to use:**
- User says "Create a page about X"
- User says "Add information about X to the wiki"
- Librarian needs to research before writing
- Topic benefits from context from existing knowledge
---
### 3. Update Wiki Page
**Endpoint:** `PUT /wiki/pages/{page_id}`
Use this for modifying existing pages. Supports partial updates.
```python
async def update_wiki_page(
page_id: int,
user: str = "default",
content: str | None = None,
title: str | None = None,
tags: list[str] | None = None,
description: str | None = None
) -> dict:
"""Update an existing wiki page (partial updates supported)."""
# Only include fields that are being updated
update_data = {}
if content is not None:
update_data["content"] = content
if title is not None:
update_data["title"] = title
if tags is not None:
update_data["tags"] = tags
if description is not None:
update_data["description"] = description
response = await http_client.put(
f"{LIBRARY_DESK_URL}/wiki/pages/{page_id}?user={user}",
headers={"Authorization": f"Bearer {LIBRARY_DESK_API_KEY}"},
json=update_data
)
return response.json()
```
**When to use:**
- User says "Update the page about X"
- User says "Fix this information"
- User says "Add this page to the projects dossier" (update tags only)
- Correcting or enhancing existing content
---
### 4. Search Wiki Pages
**Endpoint:** `GET /wiki/search`
Use this to find existing pages before updating.
```python
async def search_wiki(
query: str,
user: str = "default"
) -> dict:
"""Search wiki pages."""
response = await http_client.get(
f"{LIBRARY_DESK_URL}/wiki/search",
headers={"Authorization": f"Bearer {LIBRARY_DESK_API_KEY}"},
params={"q": query, "user": user}
)
return response.json()
```
---
### 5. Get Wiki Page
**Endpoint:** `GET /wiki/pages/{page_id}`
Use this to fetch full page content before editing.
```python
async def get_wiki_page(
page_id: int,
user: str = "default"
) -> dict:
"""Get a wiki page by ID."""
response = await http_client.get(
f"{LIBRARY_DESK_URL}/wiki/pages/{page_id}",
headers={"Authorization": f"Bearer {LIBRARY_DESK_API_KEY}"},
params={"user": user}
)
return response.json()
```
---
## Decision Flow for Librarian
```
User Request
┌─────────────────────────────────────────────┐
│ Does user want to CREATE or UPDATE a page? │
└─────────────────────────────────────────────┘
│ │
▼ ▼
CREATE UPDATE
│ │
▼ ▼
┌─────────────────┐ ┌──────────────────────┐
│ Does Librarian │ │ Search for the page │
│ need to research│ │ GET /wiki/search │
│ the topic? │ └──────────────────────┘
└─────────────────┘ │
│ │ ▼
▼ ▼ ┌──────────────────────┐
YES NO │ Get full page content│
│ │ │ GET /wiki/pages/{id} │
▼ ▼ └──────────────────────┘
┌─────────┐ ┌─────────┐ │
│ smart- │ │ POST │ ▼
│ create │ │ /wiki/ │ ┌──────────────────────┐
│ │ │ pages │ │ Update the page │
└─────────┘ └─────────┘ │ PUT /wiki/pages/{id} │
└──────────────────────┘
```
---
## Common Use Cases
### 1. "Create a page about Docker Compose"
```python
# Use smart-create for research-backed content
result = await smart_create_wiki_page(
topic="Docker Compose",
tags=["technology", "devops", "containers"],
user="jpmschweitzer"
)
# Returns page with synthesized content from wiki + web research
```
### 2. "Add this note to the wiki: Remember to renew SSL cert on Jan 15"
```python
# Use simple create for user-provided content
result = await create_wiki_page(
title="SSL Certificate Renewal Reminder",
path="/reminders/ssl-renewal",
content="# SSL Certificate Renewal\n\nRemember to renew SSL cert on Jan 15",
tags=["reminders", "infrastructure"],
user="jpmschweitzer"
)
```
### 3. "Update the page about my home server to add the new IP"
```python
# 1. Search for the page
search_results = await search_wiki("home server", user="jpmschweitzer")
page_id = search_results["results"][0]["id"]
# 2. Get current content
page = await get_wiki_page(page_id, user="jpmschweitzer")
# 3. Modify content (Librarian edits the markdown)
new_content = page["content"] + "\n\n## Updated IP\n\nNew IP: 192.168.1.100"
# 4. Update the page
result = await update_wiki_page(
page_id=page_id,
content=new_content,
user="jpmschweitzer"
)
```
### 4. "Add this page to the projects dossier"
```python
# Update only tags (partial update)
result = await update_wiki_page(
page_id=page_id,
tags=["projects", "existing-tag"], # Add "projects" tag
user="jpmschweitzer"
)
```
---
## Background Processing
All write operations trigger background tasks that:
1. **Vector Indexing:** Chunks content and generates embeddings in Qdrant
2. **Graph Extraction:** Extracts entities and creates Neo4j relationships
3. **Entity Linking:** (smart-create only) Links entities bidirectionally
These run asynchronously and don't block the API response.
---
## Authentication
All endpoints require Bearer token authentication:
```
Authorization: Bearer {LIBRARY_DESK_API_KEY}
```
---
## Multi-Tenancy
All operations are scoped to the user's namespace:
- Pages are stored under `/users/{user}/...`
- Vector collections are per-user: `library_desk_{user}`
- Graph nodes are labeled per-user: `User_{User}_Document`
Always pass the `user` parameter to ensure proper isolation.
+1 -1
View File
@@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta"
[project] [project]
name = "tatlock" name = "tatlock"
version = "1.1.0" version = "1.3.2"
description = "OpenAI-compatible API with Ollama backend" description = "OpenAI-compatible API with Ollama backend"
requires-python = ">=3.12" requires-python = ">=3.12"
dependencies = [] dependencies = []
+9 -3
View File
@@ -14,11 +14,17 @@ uvicorn[standard]>=0.38,<0.39
# Latest: 2.12.4 (Nov 5, 2025) - No known CVEs # Latest: 2.12.4 (Nov 5, 2025) - No known CVEs
pydantic>=2.11,<2.13 pydantic>=2.11,<2.13
# Pydantic settings for configuration management
# Required explicitly since pydantic-ai-slim doesn't include it
# Latest: 2.12.0 (Dec 2025) - No known CVEs
pydantic-settings>=2.12,<2.13
# AI/LLM integration # AI/LLM integration
# PydanticAI: Agent framework for using Pydantic with LLMs # PydanticAI: Agent framework for using Pydantic with LLMs
# Latest: 1.27.0 (Dec 5, 2025) - No known CVEs # Using slim version with only openai extra (Ollama uses OpenAI-compatible API)
# Supports Ollama backend out of the box # This avoids installing SDKs for anthropic, cohere, google, groq, huggingface, etc.
pydantic-ai>=1.27,<1.28 # See DEPENDENCY_SLIM.md for rollback instructions if this breaks
pydantic-ai-slim[openai]>=1.27,<1.28
# HTTP client for Ollama communication # HTTP client for Ollama communication
# Latest: 0.28.1 - No known CVEs # Latest: 0.28.1 - No known CVEs
+34
View File
@@ -0,0 +1,34 @@
"""
The Biographer - Expert for recording and recalling the user's story.
The Biographer serves as the household's memory keeper, responsible for:
- Recording and recalling facts about the user's life
- Storing personal information, preferences, and insights
- Answering questions like "What car do I drive?", "Where do I work?"
- Managing what the household knows and remembers
For direct key-based lookups (location, timezone, preferences),
use the memory_service instead - it's faster and doesn't require LLM.
The Biographer handles semantic, fuzzy queries.
"""
from src.agents.biographer.agent import (
get_biographer_agent,
run_biographer,
run_biographer_stream,
)
from src.agents.biographer.capability import (
BIOGRAPHER_CAPABILITY,
get_biographer_capability,
register_biographer,
unregister_biographer,
)
__all__ = [
"BIOGRAPHER_CAPABILITY",
"get_biographer_capability",
"get_biographer_agent",
"register_biographer",
"unregister_biographer",
"run_biographer",
"run_biographer_stream",
]
+273
View File
@@ -0,0 +1,273 @@
"""
The Biographer - Expert for recording and recalling the user's story.
A PydanticAI agent that serves as the household's memory keeper:
- Records facts about the user's life, work, and preferences
- Recalls information semantically ("What car do I drive?")
- Manages user profile and preferences
- Forgets information when requested
"""
from typing import Any, Optional
from pydantic_ai import Agent
from src.agents.biographer.tools import (
forget_memory,
list_memories,
recall_semantic,
store_insight,
update_preference,
update_profile,
)
from src.core.config import config
from src.core.logging_config import get_logger
logger = get_logger(__name__)
# The Biographer's system prompt
BIOGRAPHER_SYSTEM_PROMPT = """You are The Biographer, the household's memory keeper in the Tatlock estate.
Your role is to record, recall, and manage the story of the user's life:
- Personal facts (vehicle, pets, family members, hobbies, interests)
- Life details (employer, occupation, significant events)
- Profile information (name, location, timezone)
- Preferences (units, theme, communication style)
## Your Character
You are a discreet and attentive chronicler. Like a personal biographer who has been
with the household for years, you:
- Listen carefully and remember important details
- Recall information accurately when asked
- Never gossip or volunteer unnecessary information
- Respect privacy absolutely
- Acknowledge when you don't know something rather than guessing
## Your Tools
### Recalling the Story
- **recall_semantic**: Your primary tool for answering questions about the user
- "What car do I drive?" → searches for car-related memories
- "Where do I work?" → finds employment information
- Finds relevant memories even without exact keywords
- **list_memories**: Browse all recorded memories of a type
- Use when user asks "What do you know about me?"
- Shows everything you've recorded
### Recording New Details
- **store_insight**: Record new facts from conversation
- User says "My car is a Tesla" → store_insight("car", "Tesla Model 3")
- User says "I work at Acme" → store_insight("employer", "Acme Corp")
- Use for facts that don't fit standard profile fields
- **update_profile**: Update core biographical fields
- name, location, timezone only
- "I live in Amsterdam" → update_profile("location", "Amsterdam")
- **update_preference**: Record user preferences
- temperature_unit, distance_unit, theme, etc.
- "Use Celsius please" → update_preference("temperature_unit", "celsius")
### Managing Records
- **forget_memory**: Remove specific records
- User asks to forget something → honor immediately
- Information becomes outdated → remove it
## Guidelines
### What to Record
- Explicit statements: "I drive a Tesla", "My wife is Sarah"
- Corrections: "Actually, I moved to Berlin"
- Preferences: "I prefer metric units"
### What NOT to Record
- Sensitive data: passwords, financial details, health information
- Temporary information: "I'm tired today"
- Speculation or assumptions
### Responding to Tatlock
Your responses go to Tatlock (the butler) who synthesizes the final answer. Be:
- Direct and factual
- Clear about what you found or didn't find
- Structured for easy integration with other responses
When you don't have information:
"I have no record of the user's [topic]. Would you like me to record this information?"
When recalling:
"According to my records, [information]. This was recorded [source/when if available]."
"""
# Lazy initialization to avoid connection issues during imports
_biographer_agent: Optional[Agent[None, str]] = None
def _create_biographer_agent() -> Agent[None, str]:
"""Create The Biographer PydanticAI agent."""
# Import required classes for Ollama configuration
from pydantic_ai.models.openai import OpenAIChatModel
from pydantic_ai.providers.ollama import OllamaProvider
# PydanticAI expects Ollama base URL to end with /v1
clean_host = str(config.OLLAMA_HOST).rstrip('/')
base_url = f"{clean_host}/v1"
# Create Ollama model with provider
model = OpenAIChatModel(
model_name=config.OLLAMA_DEFAULT_MODEL,
provider=OllamaProvider(base_url=base_url)
)
agent: Agent[None, str] = Agent(
model=model,
system_prompt=BIOGRAPHER_SYSTEM_PROMPT,
retries=2,
)
# Register recall tools
agent.tool_plain(recall_semantic)
agent.tool_plain(list_memories)
# Register recording tools
agent.tool_plain(store_insight)
agent.tool_plain(update_profile)
agent.tool_plain(update_preference)
# Register management tools
agent.tool_plain(forget_memory)
logger.info(
"biographer_agent_created",
model=config.OLLAMA_DEFAULT_MODEL,
tool_count=6,
)
return agent
def get_biographer_agent() -> Agent[None, str]:
"""
Get The Biographer agent instance (lazy initialization).
Returns:
PydanticAI Agent configured for memory tasks
"""
global _biographer_agent
if _biographer_agent is None:
_biographer_agent = _create_biographer_agent()
return _biographer_agent
async def run_biographer(
task: str,
context: str = "",
message_history: Optional[list[Any]] = None,
) -> str:
"""
Execute a memory task with The Biographer.
This is the main entry point for delegating memory tasks
from Tatlock or other agents.
Args:
task: The memory task or question
context: Additional context from conversation
message_history: Optional conversation history
Returns:
Memory results or confirmation
Example:
result = await run_biographer(
task="What car do I drive?",
context="User is asking about their vehicle",
)
"""
agent = get_biographer_agent()
# Build prompt with context if provided
prompt = task
if context:
prompt = f"Context: {context}\n\nTask: {task}"
logger.info(
"biographer_task_started",
task=task[:100],
has_context=bool(context),
has_history=bool(message_history),
)
try:
result = await agent.run(
prompt,
message_history=message_history,
)
logger.info(
"biographer_task_completed",
task=task[:50],
output_length=len(result.output),
)
return result.output
except Exception as e:
logger.error(
"biographer_task_error",
task=task[:50],
error=str(e),
exc_info=True,
)
return f"The Biographer encountered an error: {str(e)}"
async def run_biographer_stream(
task: str,
context: str = "",
message_history: Optional[list[Any]] = None,
):
"""
Execute a memory task with streaming output.
Yields text deltas as The Biographer generates the response.
Args:
task: The memory task or question
context: Additional context from conversation
message_history: Optional conversation history
Yields:
str: Text deltas from the response
Example:
async for delta in run_biographer_stream("What do you know about me?"):
print(delta, end="", flush=True)
"""
agent = get_biographer_agent()
# Build prompt with context if provided
prompt = task
if context:
prompt = f"Context: {context}\n\nTask: {task}"
logger.info(
"biographer_stream_started",
task=task[:100],
)
try:
async with agent.run_stream(
prompt,
message_history=message_history,
) as response:
async for delta in response.stream_text(delta=True):
yield delta
logger.info("biographer_stream_completed", task=task[:50])
except Exception as e:
logger.error(
"biographer_stream_error",
task=task[:50],
error=str(e),
exc_info=True,
)
yield f"\n\nThe Biographer encountered an error: {str(e)}"
+88
View File
@@ -0,0 +1,88 @@
"""
Biographer capability registration for the Household Registry.
Defines The Biographer's capabilities and registers it as a
household member for coordination by the Steward and Tatlock.
"""
from src.agents.biographer.agent import get_biographer_agent
from src.agents.biographer.tools import BIOGRAPHER_TOOLS
from src.core.household_registry import (
HouseholdCapability,
get_household_registry,
)
from src.core.logging_config import get_logger
logger = get_logger(__name__)
# The Biographer's capability summary for Steward coordination
BIOGRAPHER_CAPABILITY = HouseholdCapability(
name="biographer",
role="The Biographer",
category="context",
description=(
"Memory keeper for the user's story: can RECALL personal facts "
"(car, job, family, pets), RECORD new information learned from "
"conversation, UPDATE profile (name, location, timezone) and "
"preferences (units, theme), and FORGET information when requested. "
"Use for: 'what car do I drive?', 'remember that I...', "
"'forget my...', 'what do you know about me?'"
),
domains=[
"remember",
"recall",
"forget",
"memory",
"preferences",
"profile",
"personal",
"know",
"about me",
"my",
],
cost="low", # Mostly vector search, minimal LLM
requires_network=False, # All local (Qdrant, Redis)
)
def get_biographer_capability() -> HouseholdCapability:
"""Get The Biographer's capability definition."""
return BIOGRAPHER_CAPABILITY
def register_biographer() -> None:
"""
Register The Biographer with the Household Registry.
This makes The Biographer available for:
- Steward recommendations (via capability summary)
- Tatlock delegation (via agent reference)
- Tool scoping (via tool list)
"""
registry = get_household_registry()
# Check if already registered
if "biographer" in registry:
logger.debug("biographer_already_registered")
return
registry.register(
name="biographer",
capability=BIOGRAPHER_CAPABILITY,
tools=BIOGRAPHER_TOOLS,
agent=get_biographer_agent(),
)
logger.info(
"biographer_registered",
role=BIOGRAPHER_CAPABILITY.role,
domains=BIOGRAPHER_CAPABILITY.domains,
tool_count=len(BIOGRAPHER_TOOLS),
)
def unregister_biographer() -> None:
"""Unregister The Biographer from the Household Registry."""
registry = get_household_registry()
registry.unregister("biographer")
logger.info("biographer_unregistered")
+457
View File
@@ -0,0 +1,457 @@
"""
Biographer tools for PydanticAI agent.
These tools enable The Biographer to record and recall the user's story:
- recall_semantic: Find memories by meaning/concept
- store_insight: Record new facts about the user
- list_memories: Browse recorded memories by type
- forget_memory: Remove specific memories
For direct key-based access (get/set profile, preferences),
use memory_service directly - these tools are for semantic queries.
"""
from src.core.context import get_user
from src.core.embeddings import get_embedding_client
from src.core.logging_config import get_logger
from src.core.memory_service import MemoryType, memory_service
from src.core.qdrant import get_qdrant_client
logger = get_logger(__name__)
# ============================================================================
# Semantic Recall
# ============================================================================
async def recall_semantic(
query: str,
memory_type: str = "",
limit: int = 5,
) -> str:
"""
Search memories by semantic similarity.
Use this to find memories that are conceptually related to
the query, even if exact words don't match. This is the main
tool for answering questions like "What car do I drive?" or
"What did I mention about my job?"
Args:
query: Natural language query to search for
memory_type: Optional filter: "user_profile", "preference", "learned_fact"
limit: Maximum memories to return (default: 5)
Returns:
Matching memories with their content and relevance scores
Examples:
recall_semantic("What is my car?")
recall_semantic("work preferences", memory_type="preference")
recall_semantic("family members")
"""
try:
user = get_user()
embedding_client = get_embedding_client()
qdrant = get_qdrant_client()
# Generate embedding for query
query_vector = await embedding_client.embed(query)
if not query_vector:
return "Unable to process query - embedding generation failed"
# Search memories
results = await qdrant.search_memories(
user=user,
query_vector=query_vector,
limit=limit,
memory_type=memory_type if memory_type else None,
)
if not results:
return f"No memories found related to '{query}'"
output_parts = [f"## Memories matching: {query}\n"]
for i, memory in enumerate(results, 1):
mem_type = memory.get("type", "unknown")
key = memory.get("key", "")
value = memory.get("value", "")
score = memory.get("score", 0.0)
source = memory.get("source", "unknown")
type_icon = {
"user_profile": "👤",
"preference": "⚙️",
"learned_fact": "💡",
}.get(mem_type, "📝")
output_parts.append(f"{i}. {type_icon} **{key}** (relevance: {score:.2f})")
output_parts.append(f" {value}")
output_parts.append(f" _Type: {mem_type}, Source: {source}_")
output_parts.append("")
logger.info(
"memory_recall_semantic",
query=query[:50],
result_count=len(results),
user=user,
)
return "\n".join(output_parts)
except Exception as e:
logger.error("memory_recall_semantic_error", error=str(e), query=query[:50])
return f"Error searching memories: {str(e)}"
# ============================================================================
# Store Memory
# ============================================================================
async def store_insight(
key: str,
value: str,
importance: float = 0.5,
) -> str:
"""
Store a new insight or learned fact about the user.
Use this when:
- User explicitly asks to remember something
- User shares personal information worth remembering
- You learn something from conversation that should persist
The memory will be stored with vector embedding for semantic search
and can be recalled later using recall_semantic.
Args:
key: Short identifier for the memory (e.g., "car", "employer", "pet")
value: The actual information to remember
importance: How important is this? 0.0 (trivial) to 1.0 (critical)
Returns:
Confirmation of stored memory
Examples:
store_insight("car", "User drives a Tesla Model 3")
store_insight("employer", "Works at Acme Corp as software engineer", importance=0.8)
"""
try:
# Auto-generate keywords from key and value
keywords = [key]
words = value.lower().split()
keywords.extend([w for w in words if len(w) > 4][:5])
success = await memory_service.store_fact(
key=key,
value=value,
keywords=keywords,
importance=importance,
source="conversation",
)
if success:
output_parts = [
"## Memory Stored",
f"**Key:** {key}",
f"**Value:** {value}",
f"**Keywords:** {', '.join(keywords)}",
f"**Importance:** {importance:.1f}",
"",
"_Memory is now searchable via semantic recall._"
]
logger.info(
"memory_store_insight",
key=key,
importance=importance,
user=get_user(),
)
return "\n".join(output_parts)
else:
return f"Failed to store memory for key '{key}'"
except Exception as e:
logger.error("memory_store_insight_error", error=str(e), key=key)
return f"Error storing memory: {str(e)}"
async def update_profile(
key: str,
value: str,
) -> str:
"""
Update user profile information.
Use this for core identity information:
- name, location, timezone
- language preferences
- occupation
Profile data has high importance and is used for context
by the Steward during request analysis.
Args:
key: Profile field (e.g., "name", "location", "timezone")
value: The value to set
Returns:
Confirmation of profile update
Examples:
update_profile("location", "Amsterdam, Netherlands")
update_profile("timezone", "Europe/Amsterdam")
update_profile("name", "John")
"""
try:
success = await memory_service.set_profile(
key=key,
value=value,
keywords=[key, "profile"],
)
if success:
output_parts = [
"## Profile Updated",
f"**{key}:** {value}",
"",
"_Profile data is automatically included in context._"
]
logger.info(
"memory_update_profile",
key=key,
user=get_user(),
)
return "\n".join(output_parts)
else:
return f"Failed to update profile field '{key}'"
except Exception as e:
logger.error("memory_update_profile_error", error=str(e), key=key)
return f"Error updating profile: {str(e)}"
async def update_preference(
key: str,
value: str,
) -> str:
"""
Update user preferences.
Use this for settings and preferences:
- temperature_unit (celsius/fahrenheit)
- distance_unit (metric/imperial)
- theme, language, etc.
Preferences are used by agents to customize responses.
Args:
key: Preference name (e.g., "temperature_unit", "theme")
value: Preference value
Returns:
Confirmation of preference update
Examples:
update_preference("temperature_unit", "celsius")
update_preference("distance_unit", "metric")
update_preference("theme", "dark")
"""
try:
success = await memory_service.set_preference(
key=key,
value=value,
)
if success:
output_parts = [
"## Preference Updated",
f"**{key}:** {value}",
"",
"_Preference will be applied to future responses._"
]
logger.info(
"memory_update_preference",
key=key,
user=get_user(),
)
return "\n".join(output_parts)
else:
return f"Failed to update preference '{key}'"
except Exception as e:
logger.error("memory_update_preference_error", error=str(e), key=key)
return f"Error updating preference: {str(e)}"
# ============================================================================
# List Memories
# ============================================================================
async def list_memories(
memory_type: str = "learned_fact",
limit: int = 20,
) -> str:
"""
List stored memories of a specific type.
Use this to browse what's stored in memory without
a specific search query.
Args:
memory_type: Type to list: "user_profile", "preference", "learned_fact"
limit: Maximum memories to return (default: 20)
Returns:
List of memories with their keys and values
Examples:
list_memories("user_profile")
list_memories("preference")
list_memories("learned_fact", limit=10)
"""
try:
user = get_user()
qdrant = get_qdrant_client()
# Convert string to MemoryType
try:
mem_type = MemoryType(memory_type)
except ValueError:
return f"Invalid memory type '{memory_type}'. Use: user_profile, preference, or learned_fact"
# Get all memories of type
results = qdrant._client.scroll(
collection_name=f"memories_{user}",
scroll_filter={
"must": [
{"key": "type", "match": {"value": memory_type}},
]
},
limit=limit,
with_payload=True,
with_vectors=False,
)
points, _ = results
if not points:
return f"No {memory_type} memories found"
type_icon = {
"user_profile": "👤",
"preference": "⚙️",
"learned_fact": "💡",
}.get(memory_type, "📝")
output_parts = [f"## {type_icon} {memory_type.replace('_', ' ').title()} Memories\n"]
for point in points:
payload = point.payload
key = payload.get("key", "unknown")
value = payload.get("value", "")
importance = payload.get("importance", 0.5)
output_parts.append(f"- **{key}**: {value}")
if importance > 0.7:
output_parts.append(f" _(importance: {importance:.1f})_")
logger.info(
"memory_list",
memory_type=memory_type,
count=len(points),
user=user,
)
return "\n".join(output_parts)
except Exception as e:
logger.error("memory_list_error", error=str(e), memory_type=memory_type)
return f"Error listing memories: {str(e)}"
# ============================================================================
# Forget Memory
# ============================================================================
async def forget_memory(
key: str,
memory_type: str = "learned_fact",
) -> str:
"""
Remove a specific memory.
Use this when:
- User asks to forget something
- Information is outdated or incorrect
- Privacy concerns
Args:
key: Key of the memory to forget
memory_type: Type of memory: "user_profile", "preference", "learned_fact"
Returns:
Confirmation of deletion
Examples:
forget_memory("old_car")
forget_memory("location", memory_type="user_profile")
forget_memory("theme", memory_type="preference")
"""
try:
# Convert string to MemoryType
try:
mem_type = MemoryType(memory_type)
except ValueError:
return f"Invalid memory type '{memory_type}'. Use: user_profile, preference, or learned_fact"
success = await memory_service.delete_memory(
key=key,
memory_type=mem_type,
)
if success:
output_parts = [
"## Memory Forgotten",
f"**Key:** {key}",
f"**Type:** {memory_type}",
"",
"_Memory has been removed._"
]
logger.info(
"memory_forget",
key=key,
memory_type=memory_type,
user=get_user(),
)
return "\n".join(output_parts)
else:
return f"Memory '{key}' not found or already deleted"
except Exception as e:
logger.error("memory_forget_error", error=str(e), key=key)
return f"Error forgetting memory: {str(e)}"
# ============================================================================
# Tool Collection for Registration
# ============================================================================
# All tools available to The Biographer
BIOGRAPHER_TOOLS = [
# Recall
recall_semantic,
list_memories,
# Record
store_insight,
update_profile,
update_preference,
# Manage
forget_memory,
]
+78 -1
View File
@@ -146,7 +146,84 @@ async def delegate_to_librarian(
) )
async def delegate_to_biographer(
task: str,
context: str = "",
) -> DelegationResult:
"""
Delegate a memory task to The Biographer.
The Biographer handles:
- Semantic recall ("What car do I drive?", "What's my job?")
- Recording new facts from conversation
- Profile updates (name, location, timezone)
- Preference updates (units, theme)
- Memory management (forget, list)
For direct key-based lookups (get location, get timezone), use
memory_service directly - it's faster and doesn't require LLM.
Args:
task: Clear description of what needs to be done.
Include the action verb (recall, remember, forget, etc.)
Example: "What car do I drive?"
Example: "Remember that I work at Acme Corp"
context: Additional context from the user's request or
conversation history
Returns:
DelegationResult with The Biographer's response
Example:
>>> result = await delegate_to_biographer(
... task="What do you know about my preferences?",
... context="User is asking about stored information",
... )
>>> if result.success:
... print(result.output)
"""
from src.agents.biographer.agent import run_biographer
logger.info(
"delegation_to_biographer_started",
task=task[:100],
has_context=bool(context),
)
try:
# Use run() not run_stream() - avoids Ollama bug
output = await run_biographer(task=task, context=context)
logger.info(
"delegation_to_biographer_completed",
task=task[:50],
output_length=len(output),
)
return DelegationResult(
expert_name="biographer",
task=task,
success=True,
output=output,
)
except Exception as e:
logger.error(
"delegation_to_biographer_error",
task=task[:50],
error=str(e),
exc_info=True,
)
return DelegationResult(
expert_name="biographer",
task=task,
success=False,
output="",
error=str(e),
)
# Future expert delegation wrappers will be added here: # Future expert delegation wrappers will be added here:
# - delegate_to_memory(task, context) -> DelegationResult
# - delegate_to_home_automation(task, context) -> DelegationResult # - delegate_to_home_automation(task, context) -> DelegationResult
# - delegate_to_developer(task, context) -> DelegationResult # - delegate_to_developer(task, context) -> DelegationResult
+22 -1
View File
@@ -4,7 +4,7 @@ Steward agent schemas.
Defines the structured output models for Steward's request analysis Defines the structured output models for Steward's request analysis
and capability recommendations. and capability recommendations.
""" """
from typing import Literal, Optional from typing import Any, Literal, Optional
from pydantic import BaseModel, Field from pydantic import BaseModel, Field
@@ -56,6 +56,10 @@ class StewardRecommendation(BaseModel):
default=None, default=None,
description="Description of capabilities that would be helpful but aren't available" description="Description of capabilities that would be helpful but aren't available"
) )
memory_context: dict[str, Any] = Field(
default_factory=dict,
description="Pre-fetched user context from memory (profile, preferences)"
)
def format_for_butler(self) -> str: def format_for_butler(self) -> str:
""" """
@@ -88,6 +92,23 @@ class StewardRecommendation(BaseModel):
if self.missing_capabilities: if self.missing_capabilities:
lines.append(f"⚠️ Missing: {self.missing_capabilities}") lines.append(f"⚠️ Missing: {self.missing_capabilities}")
# Memory context (user profile and preferences)
if self.memory_context:
profile = self.memory_context.get("profile", {})
preferences = self.memory_context.get("preferences", {})
if profile or preferences:
lines.append("-" * 40)
lines.append("User Context:")
if profile:
for key, value in profile.items():
lines.append(f"{key}: {value}")
if preferences:
prefs_str = ", ".join(f"{k}={v}" for k, v in preferences.items())
lines.append(f" • preferences: {prefs_str}")
lines.append("=" * 40) lines.append("=" * 40)
return "\n".join(lines) return "\n".join(lines)
+73 -2
View File
@@ -5,13 +5,15 @@ Provides high-level interface for request analysis with logging,
benchmarking, and error handling. benchmarking, and error handling.
Parses plain text recommendations into structured data. Parses plain text recommendations into structured data.
Includes memory pre-fetch for user context injection.
""" """
import re import re
from typing import Optional from typing import Any, Optional
from src.core.benchmarks import PerformanceBenchmark, get_benchmark_store from src.core.benchmarks import PerformanceBenchmark, get_benchmark_store
from src.core.household_registry import get_household_registry from src.core.household_registry import get_household_registry
from src.core.logging_config import get_logger, log_operation from src.core.logging_config import get_logger, log_operation
from src.core.memory_service import memory_service
from .agent import get_steward_agent from .agent import get_steward_agent
from .schemas import ConversationContext, StewardRecommendation from .schemas import ConversationContext, StewardRecommendation
@@ -147,6 +149,69 @@ def _extract_missing_capabilities(text: str) -> Optional[str]:
return None return None
async def _prefetch_memory_context(user_request: str) -> dict[str, Any]:
"""
Pre-fetch user context that might be needed for this request.
This is the "direct access" layer - fast lookups without LLM overhead.
Uses simple keyword matching to determine what context to fetch.
Args:
user_request: The user's request text
Returns:
Dict with profile and/or preferences data
Example:
>>> ctx = await _prefetch_memory_context("What's the weather?")
>>> ctx
{"profile": {"location": "Amsterdam"}}
"""
request_lower = user_request.lower()
# Determine what context might be needed based on keywords
profile_keys = []
# Location-related queries
if any(word in request_lower for word in [
"weather", "temperature", "forecast", "nearby", "local",
"directions", "distance", "map", "here"
]):
profile_keys.append("location")
# Time-related queries
if any(word in request_lower for word in [
"time", "schedule", "meeting", "appointment", "reminder",
"alarm", "when", "today", "tomorrow"
]):
profile_keys.append("timezone")
# Personal queries
if any(word in request_lower for word in [
"my name", "who am i", "about me"
]):
profile_keys.append("name")
# Always fetch preferences if they might affect response format
include_preferences = any(word in request_lower for word in [
"temperature", "weather", "convert", "unit", "format",
"celsius", "fahrenheit", "metric", "imperial"
])
try:
return await memory_service.prefetch_context(
include_profile=bool(profile_keys),
include_preferences=include_preferences,
profile_keys=profile_keys if profile_keys else None,
)
except Exception as e:
logger.warning(
"steward_prefetch_memory_failed",
error=str(e),
)
return {}
async def analyze_request( async def analyze_request(
user_request: str, user_request: str,
conversation_history: list[dict], conversation_history: list[dict],
@@ -186,6 +251,10 @@ async def analyze_request(
} }
) as log_ctx: ) as log_ctx:
try: try:
# Pre-fetch user context from memory (fast, no LLM)
memory_context = await _prefetch_memory_context(user_request)
log_ctx["memory_context_keys"] = list(memory_context.keys())
# Get Steward agent # Get Steward agent
steward = get_steward_agent() steward = get_steward_agent()
@@ -193,6 +262,7 @@ async def analyze_request(
"steward_analyzing_request", "steward_analyzing_request",
request=user_request, request=user_request,
history_turns=len(conversation_history), history_turns=len(conversation_history),
memory_context=bool(memory_context),
) )
# Get plain text analysis from Steward # Get plain text analysis from Steward
@@ -212,7 +282,8 @@ async def analyze_request(
reasoning=analysis_text, reasoning=analysis_text,
estimated_complexity=complexity, estimated_complexity=complexity,
conversation_context=context, conversation_context=context,
missing_capabilities=missing missing_capabilities=missing,
memory_context=memory_context,
) )
# Update log context with results # Update log context with results
+5 -5
View File
@@ -101,9 +101,9 @@ class Config(BaseSettings):
default=6379, default=6379,
description="Redis server port" description="Redis server port"
) )
REDIS_DB: int = Field( REDIS_BENCHMARK_DB: int = Field(
default=1, default=6,
description="Redis database number" description="Redis database number for benchmarks"
) )
REDIS_TIMEOUT: int = Field( REDIS_TIMEOUT: int = Field(
default=5, default=5,
@@ -146,7 +146,7 @@ class Config(BaseSettings):
# Redis Memory Database (separate from benchmarks) # Redis Memory Database (separate from benchmarks)
REDIS_MEMORY_DB: int = Field( REDIS_MEMORY_DB: int = Field(
default=2, default=1,
description="Redis database number for memory cache" description="Redis database number for memory cache"
) )
REDIS_MEMORY_TTL_HOURS: int = Field( REDIS_MEMORY_TTL_HOURS: int = Field(
@@ -170,7 +170,7 @@ class Config(BaseSettings):
@property @property
def redis_url(self) -> str: def redis_url(self) -> str:
"""Construct Redis connection URL for benchmarks.""" """Construct Redis connection URL for benchmarks."""
return f"redis://{self.REDIS_HOST}:{self.REDIS_PORT}/{self.REDIS_DB}" return f"redis://{self.REDIS_HOST}:{self.REDIS_PORT}/{self.REDIS_BENCHMARK_DB}"
@property @property
def redis_memory_url(self) -> str: def redis_memory_url(self) -> str:
+2 -2
View File
@@ -223,12 +223,12 @@ class HouseholdRegistry:
>>> # Returns: [delegate_to_librarian, calculate, datetime, ...] >>> # Returns: [delegate_to_librarian, calculate, datetime, ...]
>>> # Instead of: [hybrid_search, search_wiki, create_wiki_page, ... (16 tools)] >>> # Instead of: [hybrid_search, search_wiki, create_wiki_page, ... (16 tools)]
""" """
from src.agents.delegation import delegate_to_librarian from src.agents.delegation import delegate_to_librarian, delegate_to_biographer
# Map of expert names to their delegation wrappers # Map of expert names to their delegation wrappers
delegation_wrappers = { delegation_wrappers = {
"librarian": delegate_to_librarian, "librarian": delegate_to_librarian,
# Future: "memory": delegate_to_memory, "biographer": delegate_to_biographer,
# Future: "home_automation": delegate_to_home_automation, # Future: "home_automation": delegate_to_home_automation,
} }
+619
View File
@@ -0,0 +1,619 @@
"""
Memory service for direct key-based access.
Provides fast, LLM-free access to user memories for:
- Known-key lookups (location, timezone, preferences)
- Session context (current topic, recent entities)
- Structured storage (explicit user instructions)
This is the "direct access layer" - no LLM interpretation.
For semantic/fuzzy queries, use the Memory Agent instead.
Usage:
from src.core.memory_service import memory_service
# Get user's location (fast, no LLM)
location = await memory_service.get_profile("location")
# Set a preference
await memory_service.set_preference("temperature_unit", "celsius")
# Get session context
ctx = await memory_service.get_session_context(conversation_id)
"""
from datetime import datetime, timezone
from enum import Enum
from typing import Any
from pydantic import BaseModel, Field
from .config import config
from .context import get_user, get_conversation_id
from .embeddings import get_embedding_client
from .logging_config import get_logger
from .memory_cache import get_memory_cache
from .multi_tenancy import get_memory_collection_name
from .qdrant import get_qdrant_client
logger = get_logger(__name__)
class MemoryType(str, Enum):
"""Types of memories stored in Qdrant."""
USER_PROFILE = "user_profile" # Name, location, timezone
PREFERENCE = "preference" # Units, language, theme
LEARNED_FACT = "learned_fact" # "My car is a Tesla"
class MemoryRecord(BaseModel):
"""A memory record stored in Qdrant."""
id: str
type: MemoryType
key: str # e.g., "location", "timezone", "car"
value: str # The actual content
keywords: list[str] = Field(default_factory=list)
importance: float = 0.5 # 0.0 - 1.0
source: str = "explicit" # "explicit" | "inferred" | "conversation"
created_at: str = Field(default_factory=lambda: datetime.now(timezone.utc).isoformat())
updated_at: str = Field(default_factory=lambda: datetime.now(timezone.utc).isoformat())
class MemoryService:
"""
Direct access to user memories without LLM overhead.
Use this for:
- Known-key lookups: get_profile("location"), get_preference("units")
- Explicit storage: set_preference("theme", "dark")
- Session context: get_session_context(), update_session_context()
Do NOT use for:
- Fuzzy queries: "What car do I drive?" → Use Memory Agent
- Semantic recall: "What did I mention about X?" → Use Memory Agent
"""
def __init__(self):
"""Initialize memory service with lazy client loading."""
self._qdrant = None
self._embedding = None
self._cache = None
@property
def qdrant(self):
"""Lazy-load Qdrant client."""
if self._qdrant is None:
self._qdrant = get_qdrant_client()
return self._qdrant
@property
def embedding(self):
"""Lazy-load embedding client."""
if self._embedding is None:
self._embedding = get_embedding_client()
return self._embedding
@property
def cache(self):
"""Lazy-load Redis cache."""
if self._cache is None:
self._cache = get_memory_cache()
return self._cache
# =========================================================================
# Profile Methods (user_profile type)
# =========================================================================
async def get_profile(self, key: str, user: str | None = None) -> str | None:
"""
Get a user profile value by key.
Args:
key: Profile key (e.g., "location", "timezone", "name")
user: User ID (defaults to current request context)
Returns:
Profile value or None if not found
Example:
>>> location = await memory_service.get_profile("location")
>>> location
"Amsterdam, Netherlands"
"""
user = user or get_user()
return await self._get_memory(user, MemoryType.USER_PROFILE, key)
async def set_profile(
self,
key: str,
value: str,
user: str | None = None,
keywords: list[str] | None = None,
) -> bool:
"""
Set a user profile value.
Args:
key: Profile key (e.g., "location", "timezone")
value: Profile value
user: User ID (defaults to current request context)
keywords: Optional keywords for semantic search
Returns:
True if successful
Example:
>>> await memory_service.set_profile("location", "Amsterdam, Netherlands")
True
"""
user = user or get_user()
return await self._set_memory(
user=user,
memory_type=MemoryType.USER_PROFILE,
key=key,
value=value,
keywords=keywords or [key],
importance=0.9, # Profile data is important
)
# =========================================================================
# Preference Methods (preference type)
# =========================================================================
async def get_preference(self, key: str, user: str | None = None) -> str | None:
"""
Get a user preference by key.
Args:
key: Preference key (e.g., "temperature_unit", "language", "theme")
user: User ID (defaults to current request context)
Returns:
Preference value or None if not found
Example:
>>> units = await memory_service.get_preference("temperature_unit")
>>> units
"celsius"
"""
user = user or get_user()
return await self._get_memory(user, MemoryType.PREFERENCE, key)
async def set_preference(
self,
key: str,
value: str,
user: str | None = None,
) -> bool:
"""
Set a user preference.
Args:
key: Preference key
value: Preference value
user: User ID (defaults to current request context)
Returns:
True if successful
Example:
>>> await memory_service.set_preference("theme", "dark")
True
"""
user = user or get_user()
return await self._set_memory(
user=user,
memory_type=MemoryType.PREFERENCE,
key=key,
value=value,
keywords=[key, "preference"],
importance=0.7,
)
async def get_all_preferences(self, user: str | None = None) -> dict[str, str]:
"""
Get all preferences for a user.
Returns:
Dict of key -> value for all preferences
"""
user = user or get_user()
memories = await self._get_all_by_type(user, MemoryType.PREFERENCE)
return {m["key"]: m["value"] for m in memories}
# =========================================================================
# Learned Facts (learned_fact type) - for direct storage only
# =========================================================================
async def store_fact(
self,
key: str,
value: str,
user: str | None = None,
keywords: list[str] | None = None,
importance: float = 0.5,
source: str = "explicit",
) -> bool:
"""
Store a learned fact about the user.
Use this for explicit user statements like:
- "Remember that my car is a Tesla"
- "I work at Acme Corp"
For semantic extraction from conversation, use the Memory Agent.
Args:
key: Fact identifier (e.g., "car", "employer")
value: The fact content
user: User ID
keywords: Keywords for semantic search
importance: 0.0-1.0 importance score
source: "explicit" | "inferred" | "conversation"
Returns:
True if successful
"""
user = user or get_user()
return await self._set_memory(
user=user,
memory_type=MemoryType.LEARNED_FACT,
key=key,
value=value,
keywords=keywords or [key],
importance=importance,
source=source,
)
async def get_fact(self, key: str, user: str | None = None) -> str | None:
"""
Get a specific fact by key.
For semantic/fuzzy queries, use the Memory Agent.
"""
user = user or get_user()
return await self._get_memory(user, MemoryType.LEARNED_FACT, key)
# =========================================================================
# Session Context (Redis-backed, 24h TTL)
# =========================================================================
async def get_session_context(
self,
conversation_id: str | None = None,
user: str | None = None,
) -> dict[str, Any] | None:
"""
Get session context for current conversation.
Args:
conversation_id: Conversation ID (defaults to current context)
user: User ID (defaults to current context)
Returns:
Session context dict or None
"""
user = user or get_user()
conversation_id = conversation_id or get_conversation_id()
if not conversation_id:
return None
return await self.cache.get_session_context(user, conversation_id)
async def set_session_context(
self,
context: dict[str, Any],
conversation_id: str | None = None,
user: str | None = None,
) -> bool:
"""
Set session context for current conversation.
Args:
context: Context data to store
conversation_id: Conversation ID
user: User ID
Returns:
True if successful
"""
user = user or get_user()
conversation_id = conversation_id or get_conversation_id()
if not conversation_id:
logger.warning("memory_service_no_conversation_id")
return False
return await self.cache.set_session_context(user, conversation_id, context)
async def update_session_context(
self,
updates: dict[str, Any],
conversation_id: str | None = None,
user: str | None = None,
) -> bool:
"""
Update session context (merge with existing).
Args:
updates: Fields to update
conversation_id: Conversation ID
user: User ID
Returns:
True if successful
"""
user = user or get_user()
conversation_id = conversation_id or get_conversation_id()
if not conversation_id:
return False
return await self.cache.update_session_context(user, conversation_id, updates)
async def get_recent_entities(
self,
conversation_id: str | None = None,
user: str | None = None,
) -> list[str]:
"""
Get recently mentioned entities in conversation.
Returns:
List of entity names
"""
user = user or get_user()
conversation_id = conversation_id or get_conversation_id()
if not conversation_id:
return []
return await self.cache.get_recent_entities(user, conversation_id)
async def add_recent_entities(
self,
entities: list[str],
conversation_id: str | None = None,
user: str | None = None,
) -> bool:
"""
Add entities to recent entities set.
Args:
entities: Entity names to add
conversation_id: Conversation ID
user: User ID
Returns:
True if successful
"""
user = user or get_user()
conversation_id = conversation_id or get_conversation_id()
if not conversation_id:
return False
return await self.cache.add_recent_entities(user, conversation_id, entities)
# =========================================================================
# Bulk / Pre-fetch Methods (for Steward)
# =========================================================================
async def prefetch_context(
self,
user: str | None = None,
include_profile: bool = True,
include_preferences: bool = True,
profile_keys: list[str] | None = None,
) -> dict[str, Any]:
"""
Pre-fetch commonly needed context for Steward.
This is the main entry point for Steward to get user context
before analyzing a request.
Args:
user: User ID
include_profile: Include profile data
include_preferences: Include preferences
profile_keys: Specific profile keys to fetch (None = common ones)
Returns:
Dict with profile and preferences data
Example:
>>> ctx = await memory_service.prefetch_context()
>>> ctx
{
"profile": {"location": "Amsterdam", "timezone": "Europe/Amsterdam"},
"preferences": {"temperature_unit": "celsius"}
}
"""
user = user or get_user()
result: dict[str, Any] = {}
if include_profile:
profile_keys = profile_keys or ["location", "timezone", "name"]
profile = {}
for key in profile_keys:
value = await self.get_profile(key, user)
if value:
profile[key] = value
if profile:
result["profile"] = profile
if include_preferences:
preferences = await self.get_all_preferences(user)
if preferences:
result["preferences"] = preferences
logger.debug(
"memory_service_prefetch",
user=user,
profile_keys=list(result.get("profile", {}).keys()),
preference_keys=list(result.get("preferences", {}).keys()),
)
return result
# =========================================================================
# Internal Methods
# =========================================================================
async def _get_memory(
self,
user: str,
memory_type: MemoryType,
key: str,
) -> str | None:
"""Get a memory by type and key (exact match)."""
collection = get_memory_collection_name(user)
try:
# Search with filter for exact type + key match
# We use a dummy vector since we're filtering by payload
results = self.qdrant._client.scroll(
collection_name=collection,
scroll_filter={
"must": [
{"key": "type", "match": {"value": memory_type.value}},
{"key": "key", "match": {"value": key}},
]
},
limit=1,
with_payload=True,
with_vectors=False,
)
points, _ = results
if points:
return points[0].payload.get("value")
return None
except Exception as e:
logger.warning(
"memory_service_get_failed",
user=user,
type=memory_type.value,
key=key,
error=str(e),
)
return None
async def _set_memory(
self,
user: str,
memory_type: MemoryType,
key: str,
value: str,
keywords: list[str],
importance: float = 0.5,
source: str = "explicit",
) -> bool:
"""Set a memory (upsert by type + key)."""
try:
# Generate embedding for semantic search
embedding = await self.embedding.embed(f"{key}: {value}")
if not embedding:
logger.error("memory_service_embedding_failed", key=key)
return False
# Create memory ID from type + key for idempotent upserts
memory_id = f"{memory_type.value}:{key}"
payload = {
"type": memory_type.value,
"key": key,
"value": value,
"keywords": keywords,
"importance": importance,
"source": source,
"updated_at": datetime.now(timezone.utc).isoformat(),
}
result = await self.qdrant.upsert_memory(
user=user,
memory_id=memory_id,
vector=embedding,
payload=payload,
)
if result:
logger.debug(
"memory_service_set",
user=user,
type=memory_type.value,
key=key,
)
return True
return False
except Exception as e:
logger.error(
"memory_service_set_failed",
user=user,
type=memory_type.value,
key=key,
error=str(e),
)
return False
async def _get_all_by_type(
self,
user: str,
memory_type: MemoryType,
limit: int = 100,
) -> list[dict[str, Any]]:
"""Get all memories of a specific type."""
collection = get_memory_collection_name(user)
try:
results = self.qdrant._client.scroll(
collection_name=collection,
scroll_filter={
"must": [
{"key": "type", "match": {"value": memory_type.value}},
]
},
limit=limit,
with_payload=True,
with_vectors=False,
)
points, _ = results
return [p.payload for p in points]
except Exception as e:
logger.warning(
"memory_service_get_all_failed",
user=user,
type=memory_type.value,
error=str(e),
)
return []
async def delete_memory(
self,
key: str,
memory_type: MemoryType,
user: str | None = None,
) -> bool:
"""
Delete a specific memory.
Args:
key: Memory key
memory_type: Type of memory
user: User ID
Returns:
True if deleted
"""
user = user or get_user()
memory_id = f"{memory_type.value}:{key}"
return await self.qdrant.delete_memory(user, memory_id)
# Global service instance
memory_service = MemoryService()
+4 -4
View File
@@ -232,14 +232,14 @@ class MemoryQdrantClient:
] ]
) )
# Search # Search using new Query API (qdrant-client >= 1.10)
results = self._client.search( results = self._client.query_points(
collection_name=collection_name, collection_name=collection_name,
query_vector=query_vector, query=query_vector,
limit=limit, limit=limit,
query_filter=query_filter, query_filter=query_filter,
score_threshold=score_threshold, score_threshold=score_threshold,
) ).points
# Format results # Format results
memories = [] memories = []
+12
View File
@@ -5,6 +5,7 @@ Handles initialization of household registry and other startup tasks.
This module should be called during application startup to register This module should be called during application startup to register
all household members. all household members.
""" """
from src.agents.biographer import register_biographer
from src.agents.librarian import register_librarian from src.agents.librarian import register_librarian
from src.agents.tatlock_core import TATLOCK_CORE_CAPABILITY, tatlock_core_tools from src.agents.tatlock_core import TATLOCK_CORE_CAPABILITY, tatlock_core_tools
from src.core.household_registry import get_household_registry from src.core.household_registry import get_household_registry
@@ -23,6 +24,7 @@ def register_household_members():
Currently registers: Currently registers:
- tatlock_core: Butler's core tools (calculator, datetime, web search) - tatlock_core: Butler's core tools (calculator, datetime, web search)
- librarian: Research and knowledge management (Phase 3) - librarian: Research and knowledge management (Phase 3)
- biographer: User memory and context management (Phase F)
""" """
registry = get_household_registry() registry = get_household_registry()
@@ -52,6 +54,16 @@ def register_household_members():
error=str(e), error=str(e),
) )
# Register The Biographer (Phase F)
try:
register_biographer()
except Exception as e:
# Don't fail startup if Biographer registration fails
logger.warning(
"biographer_registration_failed",
error=str(e),
)
logger.info( logger.info(
"household_registration_complete", "household_registration_complete",
total_members=len(registry), total_members=len(registry),
+1
View File
@@ -0,0 +1 @@
"""Tests for The Biographer agent."""
+145
View File
@@ -0,0 +1,145 @@
"""
Tests for Biographer capability registration.
"""
import pytest
from unittest.mock import MagicMock, patch
from src.agents.biographer.capability import (
BIOGRAPHER_CAPABILITY,
get_biographer_capability,
register_biographer,
unregister_biographer,
)
from src.core.household_registry import HouseholdCapability
@pytest.mark.unit
class TestBiographerCapability:
"""Tests for the Biographer capability definition."""
def test_capability_is_household_capability(self):
"""Test capability is correct type."""
assert isinstance(BIOGRAPHER_CAPABILITY, HouseholdCapability)
def test_capability_name(self):
"""Test capability has correct name."""
assert BIOGRAPHER_CAPABILITY.name == "biographer"
def test_capability_role(self):
"""Test capability has correct role."""
assert BIOGRAPHER_CAPABILITY.role == "The Biographer"
def test_capability_category(self):
"""Test capability is in context category."""
assert BIOGRAPHER_CAPABILITY.category == "context"
def test_capability_domains(self):
"""Test capability covers expected domains."""
domains = BIOGRAPHER_CAPABILITY.domains
assert "remember" in domains
assert "recall" in domains
assert "forget" in domains
assert "memory" in domains
assert "preferences" in domains
assert "profile" in domains
def test_capability_does_not_require_network(self):
"""Test capability does not require network access."""
assert BIOGRAPHER_CAPABILITY.requires_network is False
def test_capability_low_cost(self):
"""Test capability has low cost (vector search, minimal LLM)."""
assert BIOGRAPHER_CAPABILITY.cost == "low"
def test_get_biographer_capability(self):
"""Test getter returns same capability."""
cap = get_biographer_capability()
assert cap is BIOGRAPHER_CAPABILITY
@pytest.mark.unit
class TestBiographerRegistration:
"""Tests for Biographer registration functions."""
def test_register_biographer(self):
"""Test registering biographer with registry."""
mock_registry = MagicMock()
mock_registry.__contains__ = MagicMock(return_value=False)
with patch(
"src.agents.biographer.capability.get_household_registry",
return_value=mock_registry,
):
with patch(
"src.agents.biographer.capability.get_biographer_agent"
) as mock_get_agent:
mock_agent = MagicMock()
mock_get_agent.return_value = mock_agent
register_biographer()
mock_registry.register.assert_called_once()
call_kwargs = mock_registry.register.call_args[1]
assert call_kwargs["name"] == "biographer"
assert call_kwargs["capability"] is BIOGRAPHER_CAPABILITY
assert call_kwargs["agent"] is mock_agent
def test_register_biographer_already_registered(self):
"""Test registering when already registered does nothing."""
mock_registry = MagicMock()
mock_registry.__contains__ = MagicMock(return_value=True)
with patch(
"src.agents.biographer.capability.get_household_registry",
return_value=mock_registry,
):
register_biographer()
# Should not call register since already registered
mock_registry.register.assert_not_called()
def test_unregister_biographer(self):
"""Test unregistering biographer from registry."""
mock_registry = MagicMock()
with patch(
"src.agents.biographer.capability.get_household_registry",
return_value=mock_registry,
):
unregister_biographer()
mock_registry.unregister.assert_called_once_with("biographer")
@pytest.mark.unit
class TestCapabilityDescription:
"""Tests for capability description."""
def test_description_mentions_recall(self):
"""Test description mentions recall capabilities."""
desc = BIOGRAPHER_CAPABILITY.description.lower()
assert "recall" in desc
def test_description_mentions_record(self):
"""Test description mentions recording capability."""
desc = BIOGRAPHER_CAPABILITY.description.lower()
assert "record" in desc
def test_description_mentions_forget(self):
"""Test description mentions forget capability."""
desc = BIOGRAPHER_CAPABILITY.description.lower()
assert "forget" in desc
def test_description_mentions_profile(self):
"""Test description mentions profile updates."""
desc = BIOGRAPHER_CAPABILITY.description.lower()
assert "profile" in desc
def test_description_mentions_preferences(self):
"""Test description mentions preferences."""
desc = BIOGRAPHER_CAPABILITY.description.lower()
assert "preferences" in desc
+279
View File
@@ -0,0 +1,279 @@
"""
Tests for the memory service (direct access layer).
"""
import pytest
from unittest.mock import MagicMock, patch, AsyncMock
from src.core.memory_service import (
MemoryService,
MemoryType,
MemoryRecord,
memory_service,
)
@pytest.mark.unit
class TestMemoryType:
"""Tests for MemoryType enum."""
def test_user_profile_type(self):
"""Test user_profile type exists."""
assert MemoryType.USER_PROFILE.value == "user_profile"
def test_preference_type(self):
"""Test preference type exists."""
assert MemoryType.PREFERENCE.value == "preference"
def test_learned_fact_type(self):
"""Test learned_fact type exists."""
assert MemoryType.LEARNED_FACT.value == "learned_fact"
@pytest.mark.unit
class TestMemoryRecord:
"""Tests for MemoryRecord model."""
def test_create_minimal_record(self):
"""Test creating record with minimal fields."""
record = MemoryRecord(
id="test_1",
type=MemoryType.USER_PROFILE,
key="location",
value="Amsterdam",
)
assert record.id == "test_1"
assert record.type == MemoryType.USER_PROFILE
assert record.key == "location"
assert record.value == "Amsterdam"
assert record.importance == 0.5 # Default
assert record.source == "explicit" # Default
def test_create_full_record(self):
"""Test creating record with all fields."""
record = MemoryRecord(
id="test_2",
type=MemoryType.LEARNED_FACT,
key="car",
value="Tesla Model 3",
keywords=["car", "vehicle", "tesla"],
importance=0.8,
source="conversation",
)
assert record.keywords == ["car", "vehicle", "tesla"]
assert record.importance == 0.8
assert record.source == "conversation"
@pytest.mark.unit
class TestMemoryServiceInit:
"""Tests for MemoryService initialization."""
def test_service_has_lazy_clients(self):
"""Test service initializes with lazy client loading."""
service = MemoryService()
assert service._qdrant is None
assert service._embedding is None
assert service._cache is None
def test_global_instance_exists(self):
"""Test global memory_service instance exists."""
assert memory_service is not None
assert isinstance(memory_service, MemoryService)
@pytest.mark.unit
class TestMemoryServiceProfileMethods:
"""Tests for profile-related methods."""
@pytest.mark.asyncio
async def test_get_profile_uses_context(self):
"""Test get_profile uses request context for user."""
service = MemoryService()
with patch.object(service, "_get_memory", new_callable=AsyncMock) as mock_get:
mock_get.return_value = "Amsterdam"
with patch("src.core.memory_service.get_user", return_value="testuser"):
result = await service.get_profile("location")
mock_get.assert_called_once_with("testuser", MemoryType.USER_PROFILE, "location")
assert result == "Amsterdam"
@pytest.mark.asyncio
async def test_get_profile_explicit_user(self):
"""Test get_profile with explicit user parameter."""
service = MemoryService()
with patch.object(service, "_get_memory", new_callable=AsyncMock) as mock_get:
mock_get.return_value = "Berlin"
result = await service.get_profile("location", user="otheruser")
mock_get.assert_called_once_with("otheruser", MemoryType.USER_PROFILE, "location")
assert result == "Berlin"
@pytest.mark.asyncio
async def test_set_profile_high_importance(self):
"""Test set_profile uses high importance (0.9)."""
service = MemoryService()
with patch.object(service, "_set_memory", new_callable=AsyncMock) as mock_set:
mock_set.return_value = True
with patch("src.core.memory_service.get_user", return_value="testuser"):
result = await service.set_profile("timezone", "Europe/Amsterdam")
call_kwargs = mock_set.call_args[1]
assert call_kwargs["importance"] == 0.9
assert result is True
@pytest.mark.unit
class TestMemoryServicePreferenceMethods:
"""Tests for preference-related methods."""
@pytest.mark.asyncio
async def test_get_preference(self):
"""Test get_preference retrieves correctly."""
service = MemoryService()
with patch.object(service, "_get_memory", new_callable=AsyncMock) as mock_get:
mock_get.return_value = "celsius"
with patch("src.core.memory_service.get_user", return_value="testuser"):
result = await service.get_preference("temperature_unit")
mock_get.assert_called_once_with("testuser", MemoryType.PREFERENCE, "temperature_unit")
assert result == "celsius"
@pytest.mark.asyncio
async def test_set_preference_medium_importance(self):
"""Test set_preference uses medium importance (0.7)."""
service = MemoryService()
with patch.object(service, "_set_memory", new_callable=AsyncMock) as mock_set:
mock_set.return_value = True
with patch("src.core.memory_service.get_user", return_value="testuser"):
result = await service.set_preference("theme", "dark")
call_kwargs = mock_set.call_args[1]
assert call_kwargs["importance"] == 0.7
@pytest.mark.unit
class TestMemoryServiceFactMethods:
"""Tests for fact-related methods."""
@pytest.mark.asyncio
async def test_store_fact_default_importance(self):
"""Test store_fact uses default importance (0.5)."""
service = MemoryService()
with patch.object(service, "_set_memory", new_callable=AsyncMock) as mock_set:
mock_set.return_value = True
with patch("src.core.memory_service.get_user", return_value="testuser"):
result = await service.store_fact("car", "Tesla Model 3")
call_kwargs = mock_set.call_args[1]
assert call_kwargs["importance"] == 0.5
@pytest.mark.asyncio
async def test_store_fact_custom_importance(self):
"""Test store_fact with custom importance."""
service = MemoryService()
with patch.object(service, "_set_memory", new_callable=AsyncMock) as mock_set:
mock_set.return_value = True
with patch("src.core.memory_service.get_user", return_value="testuser"):
result = await service.store_fact(
"employer",
"Acme Corp",
importance=0.8,
)
call_kwargs = mock_set.call_args[1]
assert call_kwargs["importance"] == 0.8
@pytest.mark.asyncio
async def test_get_fact(self):
"""Test get_fact retrieves correctly."""
service = MemoryService()
with patch.object(service, "_get_memory", new_callable=AsyncMock) as mock_get:
mock_get.return_value = "Tesla Model 3"
with patch("src.core.memory_service.get_user", return_value="testuser"):
result = await service.get_fact("car")
mock_get.assert_called_once_with("testuser", MemoryType.LEARNED_FACT, "car")
assert result == "Tesla Model 3"
@pytest.mark.unit
class TestMemoryServicePrefetch:
"""Tests for prefetch_context method."""
@pytest.mark.asyncio
async def test_prefetch_default_keys(self):
"""Test prefetch with default profile keys."""
service = MemoryService()
with patch.object(service, "get_profile", new_callable=AsyncMock) as mock_profile:
with patch.object(service, "get_all_preferences", new_callable=AsyncMock) as mock_prefs:
mock_profile.side_effect = [
"Amsterdam", # location
"Europe/Amsterdam", # timezone
"John", # name
]
mock_prefs.return_value = {"temperature_unit": "celsius"}
with patch("src.core.memory_service.get_user", return_value="testuser"):
result = await service.prefetch_context()
assert result["profile"]["location"] == "Amsterdam"
assert result["profile"]["timezone"] == "Europe/Amsterdam"
assert result["profile"]["name"] == "John"
assert result["preferences"]["temperature_unit"] == "celsius"
@pytest.mark.asyncio
async def test_prefetch_specific_keys(self):
"""Test prefetch with specific profile keys."""
service = MemoryService()
with patch.object(service, "get_profile", new_callable=AsyncMock) as mock_profile:
with patch.object(service, "get_all_preferences", new_callable=AsyncMock) as mock_prefs:
mock_profile.return_value = "Amsterdam"
mock_prefs.return_value = {}
with patch("src.core.memory_service.get_user", return_value="testuser"):
result = await service.prefetch_context(
profile_keys=["location"],
include_preferences=False,
)
# Should only fetch location
mock_profile.assert_called_once()
mock_prefs.assert_not_called()
@pytest.mark.asyncio
async def test_prefetch_no_profile(self):
"""Test prefetch without profile data."""
service = MemoryService()
with patch.object(service, "get_profile", new_callable=AsyncMock) as mock_profile:
with patch.object(service, "get_all_preferences", new_callable=AsyncMock) as mock_prefs:
mock_prefs.return_value = {"theme": "dark"}
with patch("src.core.memory_service.get_user", return_value="testuser"):
result = await service.prefetch_context(include_profile=False)
mock_profile.assert_not_called()
assert "profile" not in result
assert result["preferences"]["theme"] == "dark"