One agent doc per repo, and it is CLAUDE.md. Unlike elsewhere, the existing CLAUDE.md was not a stub -- it carried seven hard-won gotchas, all of which survive intact. AGENTS.md supplied the deployment and release material, minus its feature-branch mandate and its `git add -A` snippet, and minus its pointer to portainer-core, which is deprecated and must not be used as a source of infra facts. README.md and docs/philosophy.md linked to the retired file, so those pointers move with it. The new material is two traps that both make the runtime look like the opposite of what it is. A cold import inside the container loads src/anthropic but not src/ollama, and Ollama is the primary backend. The only import of src/ollama is a function-body one at src/anthropic/model_selector.py:230, while PREFER_CLOUD_BACKEND=false keeps the Claude path off. Read the module list naively and the disabled fallback looks live while the hot path looks dead. This matters because the Claude migration is abandoned and its remnants are supposed to read as vestigial, not as unfinished work; the doc carries the decision id so that reasoning is fetchable. Second, get_household_registry() in a fresh `docker exec python` returns zero members while the running app serves two models from it. It is populated at startup, so importing the singleton from outside the app and reading it as empty is a measurement error, not a finding. Co-Authored-By: Claude <noreply@anthropic.com>
442 lines
13 KiB
Markdown
442 lines
13 KiB
Markdown
# Tatlock - Your Homelab Butler
|
|
|
|
> **📖 For the complete system vision and architectural philosophy, see [docs/philosophy.md](docs/philosophy.md)**
|
|
|
|
A privacy-first, offline-capable personal assistant system that coordinates specialized AI agents to help with research, development, home automation, and daily organization.
|
|
|
|
## Current Status
|
|
|
|
- ✅ **Production-ready API** with OpenAI Responses API format
|
|
- ✅ **Open WebUI integration** with reasoning bubbles (`<think>` tags)
|
|
- ✅ **Two-tier architecture** - The Steward analyzes requests, Tatlock coordinates execution
|
|
- ✅ **Multi-agent coordination** - Expert household staff for specialized tasks
|
|
- ✅ **Memory system** - User profile, preferences, and semantic recall
|
|
- ✅ **Comprehensive testing** - 399 tests with good coverage
|
|
|
|
### The Household Staff
|
|
|
|
| Agent | Role | Status |
|
|
|-------|------|--------|
|
|
| **Tatlock** | The Butler - Primary interface with witty personality | ✅ Active |
|
|
| **The Steward** | Request analysis and capability recommendation | ✅ Active |
|
|
| **The Librarian** | Research, wiki management, knowledge synthesis | ✅ Active |
|
|
| **The Biographer** | User memory - profiles, preferences, facts | ✅ Active |
|
|
| **The Developer** | Code assistance, debugging, architecture | 🔜 Planned |
|
|
| **The Secretary** | Scheduling, calendars, reminders | 🔜 Planned |
|
|
| **The Handyman** | System administration, monitoring | 🔜 Planned |
|
|
| **The Housekeeper** | Home automation (Home Assistant) | 🔜 Planned |
|
|
|
|
## Features
|
|
|
|
### API Endpoints
|
|
|
|
- **Responses API** (`/v1/responses`) - OpenAI Responses API format with structured output
|
|
- Reasoning items for displaying thinking process
|
|
- Function call items for tool execution
|
|
- Message items for assistant responses
|
|
- Streaming and non-streaming support
|
|
|
|
- **Chat Completions** (`/v1/chat/completions`) - OpenAI Chat Completions compatibility
|
|
- Automatic reasoning conversion to `<think>` tags for Open WebUI
|
|
- Full OpenAI API compatibility
|
|
- Streaming support
|
|
|
|
- **Models** (`/v1/models`) - List available models
|
|
|
|
### Advanced Capabilities
|
|
|
|
- **Conversation History**: Auto-generated IDs, configurable max turns (default: 20)
|
|
- **Context Management**: Token counting, automatic trimming, usage statistics
|
|
- **Parameter Validation**: Temperature (0.0-2.0), reasoning effort levels, max tokens, stop sequences
|
|
- **Real-time Enforcement**: Stop sequence detection and max token limits during streaming
|
|
|
|
### Available Models
|
|
|
|
- **lorem-tester**: Full-featured mock agent with realistic behavior
|
|
- Configurable reasoning effort levels
|
|
- Random tool/function calls
|
|
- Error triggers for testing (rate_limit, context_overflow)
|
|
|
|
- **Tatlock**: Real PydanticAI agent with butler personality
|
|
- **LLM Backend**: Ollama (gemma4:e2b by default, local-first) with optional Claude fallback
|
|
- **Personality**: Witty British butler, research-oriented
|
|
- **Core Tools**:
|
|
- **Calculator**: Safe mathematical expression evaluation
|
|
- **Date/Time Toolkit**: Current time, relative dates, time differences
|
|
- **Web Search**: Privacy-preserving search via SearXNG
|
|
- **Household Coordination**:
|
|
- **The Steward**: Analyzes requests and recommends capabilities
|
|
- **The Librarian**: Research via library-desk HybridRAG + wiki
|
|
- **The Biographer**: User memory and preference management
|
|
- **Capabilities**: Streaming, reasoning, tool calling, multi-agent delegation
|
|
|
|
## Requirements
|
|
|
|
- Python 3.12+ (Python 3.12.11 recommended)
|
|
- **External Services** (must be running separately):
|
|
- **Ollama**: LLM inference (gemma4:e2b, nomic-embed-text)
|
|
- **Redis**: Caching and session memory
|
|
- **Qdrant**: Vector storage for The Biographer's memory
|
|
- **SearXNG**: Web search (optional)
|
|
- **library-desk**: Research API for The Librarian (optional)
|
|
|
|
## Quick Start
|
|
|
|
### Installation
|
|
|
|
```bash
|
|
# Clone the repository
|
|
git clone https://git.schweitz.net/jpmschweitzer/tatlock.git
|
|
cd tatlock
|
|
|
|
# Install dependencies
|
|
make setup
|
|
```
|
|
|
|
### Run the Server
|
|
|
|
```bash
|
|
uvicorn src.main:app --reload
|
|
```
|
|
|
|
API available at `http://localhost:8000`
|
|
|
|
## Usage Examples
|
|
|
|
### Responses API
|
|
|
|
Generate a response with reasoning:
|
|
|
|
```bash
|
|
curl http://localhost:8000/v1/responses \
|
|
-H "Content-Type: application/json" \
|
|
-d '{
|
|
"model": "lorem-tester",
|
|
"input": [
|
|
{"role": "user", "content": "Explain quantum computing"}
|
|
],
|
|
"reasoning": {
|
|
"effort": "medium",
|
|
"summary": "auto"
|
|
},
|
|
"max_output_tokens": 500,
|
|
"stream": false
|
|
}'
|
|
```
|
|
|
|
**Response Structure:**
|
|
```json
|
|
{
|
|
"id": "resp_abc123",
|
|
"object": "response",
|
|
"created_at": 1733529600,
|
|
"model": "lorem-tester",
|
|
"status": "completed",
|
|
"output": [
|
|
{
|
|
"type": "reasoning",
|
|
"summary": ["Analyzing the request...", "Considering quantum mechanics..."]
|
|
},
|
|
{
|
|
"type": "message",
|
|
"role": "assistant",
|
|
"content": [{"type": "output_text", "text": "Quantum computing uses..."}]
|
|
}
|
|
],
|
|
"usage": {
|
|
"input_tokens": 10,
|
|
"output_tokens": 50,
|
|
"reasoning_tokens": 20,
|
|
"total_tokens": 80
|
|
}
|
|
}
|
|
```
|
|
|
|
### Chat Completions (OpenAI-compatible)
|
|
|
|
```bash
|
|
curl http://localhost:8000/v1/chat/completions \
|
|
-H "Content-Type: application/json" \
|
|
-d '{
|
|
"model": "lorem-tester",
|
|
"messages": [
|
|
{"role": "user", "content": "Hello!"}
|
|
],
|
|
"temperature": 0.7,
|
|
"stream": true
|
|
}'
|
|
```
|
|
|
|
### List Models
|
|
|
|
```bash
|
|
curl http://localhost:8000/v1/models
|
|
```
|
|
|
|
### Conversation History
|
|
|
|
Optionally track conversations using metadata:
|
|
|
|
```bash
|
|
curl http://localhost:8000/v1/responses \
|
|
-H "Content-Type: application/json" \
|
|
-d '{
|
|
"model": "lorem-tester",
|
|
"input": [{"role": "user", "content": "Hello"}],
|
|
"metadata": {"conversation_id": "conv_abc123"}
|
|
}'
|
|
```
|
|
|
|
**Note**: Client must send full conversation history in `input` array (OpenAI compatible). Server optionally tracks via `metadata.conversation_id` for future features.
|
|
|
|
### Using Tatlock with Tools
|
|
|
|
Tatlock automatically uses his permanent tools when appropriate:
|
|
|
|
```bash
|
|
# Mathematical calculation
|
|
curl http://localhost:8000/v1/chat/completions \
|
|
-H "Content-Type: application/json" \
|
|
-d '{
|
|
"model": "Tatlock",
|
|
"messages": [{"role": "user", "content": "What is sqrt(144) + 25?"}]
|
|
}'
|
|
|
|
# Date/time queries
|
|
curl http://localhost:8000/v1/chat/completions \
|
|
-H "Content-Type: application/json" \
|
|
-d '{
|
|
"model": "Tatlock",
|
|
"messages": [{"role": "user", "content": "What was the date 2 weeks ago?"}]
|
|
}'
|
|
|
|
# Web search for current information
|
|
curl http://localhost:8000/v1/chat/completions \
|
|
-H "Content-Type: application/json" \
|
|
-d '{
|
|
"model": "Tatlock",
|
|
"messages": [{"role": "user", "content": "Search for recent Python 3.12 features"}]
|
|
}'
|
|
```
|
|
|
|
**Tatlock's Tool Usage Philosophy:**
|
|
- Uses calculator for ALL mathematics (even simple arithmetic)
|
|
- Uses date/time tools instead of guessing dates
|
|
- Searches for current/volatile information to verify facts
|
|
- Maintains a researcher's mindset with tool-assisted verification
|
|
|
|
## Open WebUI Integration
|
|
|
|
### Connection
|
|
|
|
If running Open WebUI in Docker and API on host:
|
|
|
|
```bash
|
|
# Use Docker bridge gateway IP
|
|
http://172.17.0.1:8000/v1/chat/completions
|
|
```
|
|
|
|
### Reasoning Display
|
|
|
|
The Chat Completions endpoint automatically:
|
|
1. Enables reasoning generation
|
|
2. Converts reasoning to `<think>` tags
|
|
3. Streams thinking before the response
|
|
|
|
Open WebUI displays this as thought bubbles separate from the main response.
|
|
|
|
### Testing Error Handling
|
|
|
|
Use special triggers in user messages:
|
|
- `"trigger_rate_limit"` - Simulates rate limit error
|
|
- `"trigger_context_overflow"` - Simulates context length error
|
|
|
|
## API Documentation
|
|
|
|
Interactive documentation available at:
|
|
- **Swagger UI**: `http://localhost:8000/docs`
|
|
- **ReDoc**: `http://localhost:8000/redoc`
|
|
|
|
## Testing
|
|
|
|
```bash
|
|
# Run all tests
|
|
pytest
|
|
|
|
# Run unit tests only (no external services needed)
|
|
pytest --ignore=tests/e2e --ignore=tests/integration --ignore=tests/contracts
|
|
|
|
# Wire-level contract tests against live service boundaries
|
|
make test-contracts
|
|
|
|
# Run with coverage
|
|
pytest --cov=src --cov-report=term-missing
|
|
|
|
# Current: ~400 tests
|
|
```
|
|
|
|
**Test Categories:**
|
|
- Unit tests: Agent tools, capabilities, schemas, memory service
|
|
- Integration tests: Full API stack with real Ollama
|
|
- End-to-end tests: Chat completions, responses API
|
|
|
|
## Deployment
|
|
|
|
### Production Server
|
|
|
|
```bash
|
|
# Multiple workers for production
|
|
uvicorn src.main:app --host 0.0.0.0 --port 8000 --workers 4
|
|
```
|
|
|
|
### Recommendations
|
|
|
|
- Use reverse proxy (nginx/caddy) for HTTPS
|
|
- Enable rate limiting
|
|
- Set up monitoring and logging
|
|
- Configure resource limits
|
|
- Use process manager (systemd/supervisor)
|
|
|
|
## Configuration
|
|
|
|
Create a `.env` file for custom configuration:
|
|
|
|
```env
|
|
# API Configuration
|
|
API_HOST=0.0.0.0
|
|
API_PORT=8000
|
|
|
|
# Ollama Configuration (primary backend)
|
|
OLLAMA_HOST=http://localhost:11434
|
|
OLLAMA_DEFAULT_MODEL=gemma4:e2b
|
|
OLLAMA_EMBEDDING_MODEL=nomic-embed-text
|
|
OLLAMA_TIMEOUT=120
|
|
|
|
# Claude fallback (optional; used when Ollama is down or PREFER_CLOUD_BACKEND=true)
|
|
# ANTHROPIC_API_KEY=sk-ant-api03-your-key-here
|
|
ANTHROPIC_MODEL=claude-sonnet-5
|
|
PREFER_CLOUD_BACKEND=false
|
|
|
|
# Redis Configuration
|
|
REDIS_HOST=localhost
|
|
REDIS_PORT=6379
|
|
REDIS_MEMORY_DB=2
|
|
REDIS_MEMORY_TTL_HOURS=24
|
|
|
|
# Qdrant Configuration (for memory)
|
|
QDRANT_HOST=localhost
|
|
QDRANT_PORT=6333
|
|
QDRANT_EMBEDDING_DIM=768
|
|
|
|
# Library-desk Configuration (for The Librarian)
|
|
LIBRARY_DESK_HOST=http://localhost:8089
|
|
LIBRARY_DESK_TIMEOUT=60
|
|
|
|
# SearXNG Configuration (for web search)
|
|
SEARXNG_HOST=http://localhost:8087
|
|
SEARXNG_TIMEOUT=30
|
|
|
|
# Logging
|
|
LOG_LEVEL=INFO
|
|
|
|
# CORS (default: allow all)
|
|
CORS_ORIGINS=["*"]
|
|
```
|
|
|
|
See `.env.example` for full configuration options.
|
|
|
|
## Troubleshooting
|
|
|
|
### Streaming not working
|
|
- Verify SSE-Starlette is installed
|
|
- Check client supports Server-Sent Events
|
|
- Test with: `pytest tests/responses/ -k streaming`
|
|
|
|
### Open WebUI can't connect
|
|
- Use Docker bridge gateway IP: `172.17.0.1:8000`
|
|
- Check firewall settings
|
|
- Verify server is running on `0.0.0.0`
|
|
|
|
### Reasoning not showing
|
|
- Ensure using Chat Completions endpoint (auto-enables reasoning)
|
|
- Or manually enable in Responses API: `"reasoning": {"effort": "medium", "summary": "auto"}`
|
|
- Check Open WebUI version supports `<think>` tags
|
|
|
|
### Tatlock agent errors
|
|
- Verify Ollama is running: `curl http://localhost:11434/api/tags`
|
|
- Check model is downloaded: `ollama list`
|
|
- Review environment variables: `OLLAMA_HOST`, `OLLAMA_DEFAULT_MODEL`
|
|
- Check logs: `tail -f logs/server.log`
|
|
|
|
### Web search not working
|
|
- Verify SearXNG is running: `curl http://localhost:8087/`
|
|
- Check `SEARXNG_HOST` environment variable
|
|
- SearXNG is optional - Tatlock will note if search is unavailable
|
|
|
|
## Project Structure
|
|
|
|
```
|
|
tatlock/
|
|
├── src/
|
|
│ ├── agents/ # Agent implementations
|
|
│ │ ├── biographer/ # The Biographer - memory management
|
|
│ │ ├── librarian/ # The Librarian - research & wiki
|
|
│ │ ├── steward/ # The Steward - request analysis
|
|
│ │ ├── tatlock_core/ # Core butler tools
|
|
│ │ ├── tatlock.py # Tatlock PydanticAI agent
|
|
│ │ ├── delegation.py # Expert delegation wrappers
|
|
│ │ └── protocol.py # Agent error protocol
|
|
│ ├── responses/ # Responses API (primary endpoint)
|
|
│ ├── chat/ # Chat Completions wrapper
|
|
│ ├── models/ # Models listing
|
|
│ ├── core/ # Shared infrastructure
|
|
│ │ ├── config.py # Configuration management
|
|
│ │ ├── context.py # Request context (ContextVar)
|
|
│ │ ├── memory_service.py # Direct memory access
|
|
│ │ ├── memory_cache.py # Redis session cache
|
|
│ │ ├── embeddings.py # Ollama embedding client
|
|
│ │ ├── qdrant.py # Vector database client
|
|
│ │ └── multi_tenancy.py # User isolation utilities
|
|
│ └── main.py # Application entry point
|
|
├── tests/ # Comprehensive test suite
|
|
├── docs/ # Project documentation
|
|
├── CHANGELOG.md # Version history
|
|
└── README.md # This file
|
|
```
|
|
|
|
## Development
|
|
|
|
For LLM agent development guidelines and architectural decisions, see [CLAUDE.md](CLAUDE.md).
|
|
|
|
## Contributing
|
|
|
|
1. Fork the repository
|
|
2. Create a feature branch
|
|
3. Make changes with tests
|
|
4. Ensure tests pass: `pytest`
|
|
5. Submit pull request
|
|
|
|
## Documentation
|
|
|
|
- **System Philosophy**: [docs/philosophy.md](docs/philosophy.md) - Vision, goals, and architectural patterns
|
|
- **Development Roadmap**: [docs/roadmap.md](docs/roadmap.md) - Open work and planned phases
|
|
- **Developer Guidelines**: [CLAUDE.md](CLAUDE.md) - LLM agent development patterns
|
|
- **Version History**: [CHANGELOG.md](CHANGELOG.md) - Changes and releases
|
|
|
|
### External References
|
|
- **OpenAI Responses API**: https://platform.openai.com/docs/api-reference/responses
|
|
- **FastAPI**: https://fastapi.tiangolo.com/
|
|
- **PydanticAI**: https://ai.pydantic.dev/
|
|
|
|
## License
|
|
|
|
[Add your license here]
|
|
|
|
## Version
|
|
|
|
Current version: see [CHANGELOG.md](CHANGELOG.md)
|
|
|
|
---
|
|
|
|
**Note**: Tatlock is a production-ready homelab butler. All household staff use PydanticAI with local Ollama inference (gemma4), with an optional Claude cloud fallback.
|