Files
jpmschweitzerandClaude f8059771ce docs: fold AGENTS.md into CLAUDE.md and record the backend traps
One agent doc per repo, and it is CLAUDE.md. Unlike elsewhere, the
existing CLAUDE.md was not a stub -- it carried seven hard-won gotchas,
all of which survive intact. AGENTS.md supplied the deployment and
release material, minus its feature-branch mandate and its `git add -A`
snippet, and minus its pointer to portainer-core, which is deprecated and
must not be used as a source of infra facts. README.md and
docs/philosophy.md linked to the retired file, so those pointers move
with it.

The new material is two traps that both make the runtime look like the
opposite of what it is.

A cold import inside the container loads src/anthropic but not
src/ollama, and Ollama is the primary backend. The only import of
src/ollama is a function-body one at src/anthropic/model_selector.py:230,
while PREFER_CLOUD_BACKEND=false keeps the Claude path off. Read the
module list naively and the disabled fallback looks live while the hot
path looks dead. This matters because the Claude migration is abandoned
and its remnants are supposed to read as vestigial, not as unfinished
work; the doc carries the decision id so that reasoning is fetchable.

Second, get_household_registry() in a fresh `docker exec python` returns
zero members while the running app serves two models from it. It is
populated at startup, so importing the singleton from outside the app and
reading it as empty is a measurement error, not a finding.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-09 03:16:29 +02:00

13 KiB

Tatlock - Your Homelab Butler

📖 For the complete system vision and architectural philosophy, see docs/philosophy.md

A privacy-first, offline-capable personal assistant system that coordinates specialized AI agents to help with research, development, home automation, and daily organization.

Current Status

  • Production-ready API with OpenAI Responses API format
  • Open WebUI integration with reasoning bubbles (<think> tags)
  • Two-tier architecture - The Steward analyzes requests, Tatlock coordinates execution
  • Multi-agent coordination - Expert household staff for specialized tasks
  • Memory system - User profile, preferences, and semantic recall
  • Comprehensive testing - 399 tests with good coverage

The Household Staff

Agent Role Status
Tatlock The Butler - Primary interface with witty personality Active
The Steward Request analysis and capability recommendation Active
The Librarian Research, wiki management, knowledge synthesis Active
The Biographer User memory - profiles, preferences, facts Active
The Developer Code assistance, debugging, architecture 🔜 Planned
The Secretary Scheduling, calendars, reminders 🔜 Planned
The Handyman System administration, monitoring 🔜 Planned
The Housekeeper Home automation (Home Assistant) 🔜 Planned

Features

API Endpoints

  • Responses API (/v1/responses) - OpenAI Responses API format with structured output

    • Reasoning items for displaying thinking process
    • Function call items for tool execution
    • Message items for assistant responses
    • Streaming and non-streaming support
  • Chat Completions (/v1/chat/completions) - OpenAI Chat Completions compatibility

    • Automatic reasoning conversion to <think> tags for Open WebUI
    • Full OpenAI API compatibility
    • Streaming support
  • Models (/v1/models) - List available models

Advanced Capabilities

  • Conversation History: Auto-generated IDs, configurable max turns (default: 20)
  • Context Management: Token counting, automatic trimming, usage statistics
  • Parameter Validation: Temperature (0.0-2.0), reasoning effort levels, max tokens, stop sequences
  • Real-time Enforcement: Stop sequence detection and max token limits during streaming

Available Models

  • lorem-tester: Full-featured mock agent with realistic behavior

    • Configurable reasoning effort levels
    • Random tool/function calls
    • Error triggers for testing (rate_limit, context_overflow)
  • Tatlock: Real PydanticAI agent with butler personality

    • LLM Backend: Ollama (gemma4:e2b by default, local-first) with optional Claude fallback
    • Personality: Witty British butler, research-oriented
    • Core Tools:
      • Calculator: Safe mathematical expression evaluation
      • Date/Time Toolkit: Current time, relative dates, time differences
      • Web Search: Privacy-preserving search via SearXNG
    • Household Coordination:
      • The Steward: Analyzes requests and recommends capabilities
      • The Librarian: Research via library-desk HybridRAG + wiki
      • The Biographer: User memory and preference management
    • Capabilities: Streaming, reasoning, tool calling, multi-agent delegation

Requirements

  • Python 3.12+ (Python 3.12.11 recommended)
  • External Services (must be running separately):
    • Ollama: LLM inference (gemma4:e2b, nomic-embed-text)
    • Redis: Caching and session memory
    • Qdrant: Vector storage for The Biographer's memory
    • SearXNG: Web search (optional)
    • library-desk: Research API for The Librarian (optional)

Quick Start

Installation

# Clone the repository
git clone https://git.schweitz.net/jpmschweitzer/tatlock.git
cd tatlock

# Install dependencies
make setup

Run the Server

uvicorn src.main:app --reload

API available at http://localhost:8000

Usage Examples

Responses API

Generate a response with reasoning:

curl http://localhost:8000/v1/responses \
  -H "Content-Type: application/json" \
  -d '{
    "model": "lorem-tester",
    "input": [
      {"role": "user", "content": "Explain quantum computing"}
    ],
    "reasoning": {
      "effort": "medium",
      "summary": "auto"
    },
    "max_output_tokens": 500,
    "stream": false
  }'

Response Structure:

{
  "id": "resp_abc123",
  "object": "response",
  "created_at": 1733529600,
  "model": "lorem-tester",
  "status": "completed",
  "output": [
    {
      "type": "reasoning",
      "summary": ["Analyzing the request...", "Considering quantum mechanics..."]
    },
    {
      "type": "message",
      "role": "assistant",
      "content": [{"type": "output_text", "text": "Quantum computing uses..."}]
    }
  ],
  "usage": {
    "input_tokens": 10,
    "output_tokens": 50,
    "reasoning_tokens": 20,
    "total_tokens": 80
  }
}

Chat Completions (OpenAI-compatible)

curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "lorem-tester",
    "messages": [
      {"role": "user", "content": "Hello!"}
    ],
    "temperature": 0.7,
    "stream": true
  }'

List Models

curl http://localhost:8000/v1/models

Conversation History

Optionally track conversations using metadata:

curl http://localhost:8000/v1/responses \
  -H "Content-Type: application/json" \
  -d '{
    "model": "lorem-tester",
    "input": [{"role": "user", "content": "Hello"}],
    "metadata": {"conversation_id": "conv_abc123"}
  }'

Note: Client must send full conversation history in input array (OpenAI compatible). Server optionally tracks via metadata.conversation_id for future features.

Using Tatlock with Tools

Tatlock automatically uses his permanent tools when appropriate:

# Mathematical calculation
curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Tatlock",
    "messages": [{"role": "user", "content": "What is sqrt(144) + 25?"}]
  }'

# Date/time queries
curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Tatlock",
    "messages": [{"role": "user", "content": "What was the date 2 weeks ago?"}]
  }'

# Web search for current information
curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Tatlock",
    "messages": [{"role": "user", "content": "Search for recent Python 3.12 features"}]
  }'

Tatlock's Tool Usage Philosophy:

  • Uses calculator for ALL mathematics (even simple arithmetic)
  • Uses date/time tools instead of guessing dates
  • Searches for current/volatile information to verify facts
  • Maintains a researcher's mindset with tool-assisted verification

Open WebUI Integration

Connection

If running Open WebUI in Docker and API on host:

# Use Docker bridge gateway IP
http://172.17.0.1:8000/v1/chat/completions

Reasoning Display

The Chat Completions endpoint automatically:

  1. Enables reasoning generation
  2. Converts reasoning to <think> tags
  3. Streams thinking before the response

Open WebUI displays this as thought bubbles separate from the main response.

Testing Error Handling

Use special triggers in user messages:

  • "trigger_rate_limit" - Simulates rate limit error
  • "trigger_context_overflow" - Simulates context length error

API Documentation

Interactive documentation available at:

  • Swagger UI: http://localhost:8000/docs
  • ReDoc: http://localhost:8000/redoc

Testing

# Run all tests
pytest

# Run unit tests only (no external services needed)
pytest --ignore=tests/e2e --ignore=tests/integration --ignore=tests/contracts

# Wire-level contract tests against live service boundaries
make test-contracts

# Run with coverage
pytest --cov=src --cov-report=term-missing

# Current: ~400 tests

Test Categories:

  • Unit tests: Agent tools, capabilities, schemas, memory service
  • Integration tests: Full API stack with real Ollama
  • End-to-end tests: Chat completions, responses API

Deployment

Production Server

# Multiple workers for production
uvicorn src.main:app --host 0.0.0.0 --port 8000 --workers 4

Recommendations

  • Use reverse proxy (nginx/caddy) for HTTPS
  • Enable rate limiting
  • Set up monitoring and logging
  • Configure resource limits
  • Use process manager (systemd/supervisor)

Configuration

Create a .env file for custom configuration:

# API Configuration
API_HOST=0.0.0.0
API_PORT=8000

# Ollama Configuration (primary backend)
OLLAMA_HOST=http://localhost:11434
OLLAMA_DEFAULT_MODEL=gemma4:e2b
OLLAMA_EMBEDDING_MODEL=nomic-embed-text
OLLAMA_TIMEOUT=120

# Claude fallback (optional; used when Ollama is down or PREFER_CLOUD_BACKEND=true)
# ANTHROPIC_API_KEY=sk-ant-api03-your-key-here
ANTHROPIC_MODEL=claude-sonnet-5
PREFER_CLOUD_BACKEND=false

# Redis Configuration
REDIS_HOST=localhost
REDIS_PORT=6379
REDIS_MEMORY_DB=2
REDIS_MEMORY_TTL_HOURS=24

# Qdrant Configuration (for memory)
QDRANT_HOST=localhost
QDRANT_PORT=6333
QDRANT_EMBEDDING_DIM=768

# Library-desk Configuration (for The Librarian)
LIBRARY_DESK_HOST=http://localhost:8089
LIBRARY_DESK_TIMEOUT=60

# SearXNG Configuration (for web search)
SEARXNG_HOST=http://localhost:8087
SEARXNG_TIMEOUT=30

# Logging
LOG_LEVEL=INFO

# CORS (default: allow all)
CORS_ORIGINS=["*"]

See .env.example for full configuration options.

Troubleshooting

Streaming not working

  • Verify SSE-Starlette is installed
  • Check client supports Server-Sent Events
  • Test with: pytest tests/responses/ -k streaming

Open WebUI can't connect

  • Use Docker bridge gateway IP: 172.17.0.1:8000
  • Check firewall settings
  • Verify server is running on 0.0.0.0

Reasoning not showing

  • Ensure using Chat Completions endpoint (auto-enables reasoning)
  • Or manually enable in Responses API: "reasoning": {"effort": "medium", "summary": "auto"}
  • Check Open WebUI version supports <think> tags

Tatlock agent errors

  • Verify Ollama is running: curl http://localhost:11434/api/tags
  • Check model is downloaded: ollama list
  • Review environment variables: OLLAMA_HOST, OLLAMA_DEFAULT_MODEL
  • Check logs: tail -f logs/server.log

Web search not working

  • Verify SearXNG is running: curl http://localhost:8087/
  • Check SEARXNG_HOST environment variable
  • SearXNG is optional - Tatlock will note if search is unavailable

Project Structure

tatlock/
├── src/
│   ├── agents/              # Agent implementations
│   │   ├── biographer/      # The Biographer - memory management
│   │   ├── librarian/       # The Librarian - research & wiki
│   │   ├── steward/         # The Steward - request analysis
│   │   ├── tatlock_core/    # Core butler tools
│   │   ├── tatlock.py       # Tatlock PydanticAI agent
│   │   ├── delegation.py    # Expert delegation wrappers
│   │   └── protocol.py      # Agent error protocol
│   ├── responses/           # Responses API (primary endpoint)
│   ├── chat/                # Chat Completions wrapper
│   ├── models/              # Models listing
│   ├── core/                # Shared infrastructure
│   │   ├── config.py        # Configuration management
│   │   ├── context.py       # Request context (ContextVar)
│   │   ├── memory_service.py # Direct memory access
│   │   ├── memory_cache.py  # Redis session cache
│   │   ├── embeddings.py    # Ollama embedding client
│   │   ├── qdrant.py        # Vector database client
│   │   └── multi_tenancy.py # User isolation utilities
│   └── main.py              # Application entry point
├── tests/                   # Comprehensive test suite
├── docs/                    # Project documentation
├── CHANGELOG.md             # Version history
└── README.md                # This file

Development

For LLM agent development guidelines and architectural decisions, see CLAUDE.md.

Contributing

  1. Fork the repository
  2. Create a feature branch
  3. Make changes with tests
  4. Ensure tests pass: pytest
  5. Submit pull request

Documentation

External References

License

[Add your license here]

Version

Current version: see CHANGELOG.md


Note: Tatlock is a production-ready homelab butler. All household staff use PydanticAI with local Ollama inference (gemma4), with an optional Claude cloud fallback.