Files
tatlock/CHANGELOG.md
T
jpmschweitzerandClaude Sonnet 4.5 505d284977 docs: update changelog for streaming fixes and E2E tests
Document streaming bug fixes and new E2E test suite in changelog.

**Added:**
- End-to-End test suite documentation (17 tests)
- OpenAI API spec compliance verification
- Tool usage indicators and flexible LLM assertions

**Fixed:**
- Streaming text repetition (delta mode implementation)
- Broken tool execution in streaming
- Invalid schema parameters
- Case sensitivity in model routing

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>
2025-12-07 15:41:34 +01:00

14 KiB

Changelog

All notable changes to this project will be documented in this file.

The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.

Unreleased

Added

Phase 2: The Steward (Two-Tier Architecture)

  • The Steward Agent: First-tier LLM agent for request analysis and capability recommendation

    • Analyzes requests with full conversation context awareness
    • Recommends relevant household capabilities for each request
    • Detects missing capabilities and provides guidance
    • Estimates request complexity (simple/moderate/complex)
    • Uses same Ollama model as Tatlock for VRAM efficiency
  • Household Registry: Centralized capability management system

    • HouseholdRegistry for registering capabilities and toolsets
    • HouseholdCapability executive summaries for coordination
    • HouseholdMember specifications with PydanticAI toolsets
    • Domain-based tool organization (e.g., src/agents/tatlock_core/)
    • Dynamic tool scoping per request
  • Request Preprocessing Pipeline: Steward → Tatlock flow integration

    • preprocess_request() orchestrates Steward analysis
    • Creates scoped toolsets based on recommendations
    • Formats Steward notes for Butler (conversation context included)
    • Integrated with Responses API via create_response_with_steward()
  • Tool Usage Tracking: Benchmarking and accuracy analysis

    • ToolCallTracker for monitoring recommended vs. actual tool usage
    • Tracks recommendation accuracy metrics
    • Records benchmarks to Redis for cross-session analysis
    • Supports precision/recall/F1 score calculation
  • Streaming Transparency: Real-time Steward analysis visibility

    • Streams Steward's reasoning as reasoning summary deltas
    • Streams Tatlock's response as output text deltas
    • Full SSE support for Steward + Tatlock flow
    • Conversation context and missing capabilities visible in stream
  • Structured Logging: Operation timing and metadata tracking

    • structlog-based JSON logging for machine parsing
    • Context managers for automatic operation timing
    • Metadata enrichment for debugging and analysis
    • Integrated with benchmark recording
  • Redis Benchmark Storage: Performance metrics persistence

    • Cross-session benchmark storage with 30-day expiry
    • Time-series metrics for Steward analysis and tool calls
    • Queryable by operation, time range, and metadata
    • Support for recommendation accuracy tracking
  • Benchmark Analysis Tools: Performance analysis CLI

    • scripts/benchmark_analysis.py for metric analysis
    • Steward performance statistics (latency, success rate, recommendations)
    • Tool recommendation accuracy analysis (precision, recall, F1)
    • Per-tool accuracy breakdown and duration statistics
  • End-to-End Test Suite: Comprehensive API integration tests

    • 17 E2E tests making real HTTP requests to running server
    • Tests for Chat Completions, Responses API, and streaming endpoints
    • OpenAI API spec compliance verification (format validation)
    • Steward preprocessing integration verification
    • Error handling tests (404, 422 status codes)
    • Flexible assertions for LLM output variance
    • Tool usage indicators: 🧮 (calculator), 🔍 (search), 🕐 (datetime)
    • Full documentation in tests/e2e/README.md

Phase 1 Enhancements

  • Conversation history support: Tatlock now remembers previous turns in multi-turn conversations
    • OpenAI-format messages converted to PydanticAI ModelRequest/ModelResponse objects
    • Full conversation context passed to agent via message_history parameter
    • Empty messages filtered to prevent Ollama errors
  • Tool call logging to reasoning output: Users can see what tools are doing in real-time
    • ToolCallTracker dependency system for per-request tool usage logging
    • Web search queries appear with 🔍 emoji (e.g., "🔍 Searching for: 'Python 3.13'")
    • Calculator expressions appear with 🧮 emoji (e.g., "🧮 Calculating: sqrt(144) + 25")
    • Date/time operations appear with 🕐 emoji (e.g., "🕐 Calculating date offset: 2 weeks ago")
    • Tool usage visible in <think> tags in Open WebUI

Changed

  • Architecture: Two-tier request flow (Steward analysis → Tatlock execution)
  • Tool Organization: Tatlock core tools reorganized into domain directory
  • Tool Scoping: Tatlock runs with dynamically scoped toolsets per request
  • Responses API: Integrated Steward preprocessing for all Tatlock requests
  • Streaming: Enhanced to include Steward reasoning transparency
  • Enhanced Tatlock agent with conversation memory capabilities
  • All tools now log their usage via RunContext dependencies
  • Improved debug logging for message history construction

Fixed

  • Streaming text repetition: Fixed text accumulation bug causing repetitive output in Open WebUI
    • Changed from accumulated text to delta mode (stream_text(delta=True))
    • Implemented proper run_with_scoped_tools_stream() using PydanticAI's run_stream()
    • Replaced artificial word-by-word chunking with real LLM deltas
  • Broken tool execution in streaming: Tools now execute properly in streaming mode
    • Previously showed raw JSON function calls instead of executed results
    • Now properly streams tool execution results
  • Invalid schema parameter: Removed invalid thinking parameter from ReasoningOutputItem
  • Case sensitivity in model routing: Model comparison now case-insensitive (.lower())
  • Conversation context now properly maintained across multiple turns
  • Tool usage transparency - users can see exactly what queries/calculations are being performed
  • Schema object handling in usage calculation (_calculate_usage reordered isinstance checks)

0.2.0 - 2025-12-06

Added

PydanticAI Integration (Phase 1)

  • Real Tatlock agent using PydanticAI with Ollama backend (mistral-nemo:latest)
  • British butler personality with research-oriented mindset
  • Lazy agent initialization to avoid connection issues in tests
  • Streaming response integration with reasoning output
  • Error handling for PydanticAI-specific exceptions

Permanent Tools (Phase 1)

  • Calculator tool (src/agents/tools.py):
    • Safe mathematical expression evaluation using restricted namespace
    • Support for arithmetic, algebra, trigonometry, logarithms
    • Math functions: sqrt, sin, cos, tan, log, exp, etc.
    • Constants: pi, e
    • Integer result formatting (removes unnecessary decimals)
  • Date/Time toolkit:
    • get_current_datetime: Current date/time in multiple formats
    • calculate_time_offset: Relative date calculations ("1 week ago", "2 months from now")
    • time_difference: Human-readable time differences between dates
  • Web Search tool:
    • SearXNG integration for privacy-preserving web search
    • Automatic fallback from production to localhost in development
    • Formatted search results with titles, URLs, and snippets
    • Configurable result limits (max 10)

Tool Framework

  • PydanticAI tool registration with @agent.tool decorator
  • Tool descriptions visible to LLM for intelligent usage
  • Async tool support for I/O operations
  • Error handling with string-based error messages
  • Tool usage guidelines in system prompt

Configuration

  • SearXNG configuration in src/core/config.py:
    • SEARXNG_HOST with development fallback
    • SEARXNG_TIMEOUT setting
  • Updated .env.example with SearXNG configuration
  • Ollama configuration documentation

Testing

  • 26 new tool tests (tests/agents/test_tools.py):
    • 7 calculator tests (arithmetic, functions, error handling)
    • 14 date/time tests (current time, offsets, differences)
    • 5 web search tests (mocked HTTP client)
  • Updated registry tests for tools capability
  • Total: 131 tests, 81.78% coverage (up from 95 tests, 78.95%)

Documentation

  • Comprehensive README.md updates:
    • Tatlock agent capabilities and tool descriptions
    • Requirements section with Ollama and SearXNG setup
    • Configuration examples for external services
    • Tool usage examples and philosophy
    • Troubleshooting for Ollama and SearXNG
    • Updated test statistics
  • AGENTS.md refactored for LLM development:
    • PydanticAI tool registration pattern
    • Tool implementation guidelines
    • Removed project status, focused on development instructions
  • IMPLEMENTATION_ROADMAP.md updates:
    • Phase 1 marked as "MOSTLY COMPLETE"
    • Detailed completion status for each deliverable
    • Updated current state summary

Changed

  • Tatlock agent converted from mock to real PydanticAI implementation
  • Tatlock capabilities updated: tools: True
  • Streaming coordination now handles chunk-based delivery (50 chars) to preserve markdown
  • Chat service streaming updated to preserve formatting
  • System prompt enhanced with tool usage guidelines and research mindset
  • Agent initialization changed to lazy pattern for better testability

Fixed

  • Text duplication bug in streaming responses (proper delta calculation)
  • Markdown formatting preservation in streamed responses
  • GeneratorExit errors from async context managers in generators
  • PydanticAI API usage (result.output instead of result.data)

0.1.1 - 2025-12-06

Added

Agent Interface (Phase 1)

  • Abstract AgentInterface base class for model abstraction
  • LoremTesterAgent: Full-featured mock agent with realistic behavior
    • Configurable reasoning effort levels (none, minimal, low, medium, high, xhigh)
    • Random tool/function call generation for testing
    • Error triggers: rate_limit, context_overflow, invalid_tool
    • Temperature-based response variation
  • TatlockAgent: Placeholder for future PydanticAI integration
  • ModelRegistry: Centralized model management and discovery
  • 18 agent tests with comprehensive coverage

Responses API (Phases 2, 3, 6)

  • OpenAI Responses API format with structured output (/v1/responses)
    • Reasoning items (thinking summaries with configurable effort)
    • Function call items (tool execution simulation)
    • Message items (assistant responses with output_text)
    • Streaming and non-streaming modes
  • Real-time streaming with SSE-Starlette
  • Conversation history management (Phase 3):
    • Hybrid client/server approach (client maintains state, server tracks)
    • Auto-generated deterministic conversation IDs from message hash
    • Configurable max turns with automatic trimming (default: 20)
    • Context window management with approximate token counting
    • Token usage statistics
    • Placeholder for future vector memory (Qdrant)
  • Advanced features (Phase 6):
    • Parameter validation with Pydantic field validators
    • Temperature: 0.0-2.0 range enforcement
    • Reasoning effort: 6 levels validation
    • Max output tokens: positive integer enforcement
    • Stop sequences: up to 4, non-empty strings
    • Real-time stop sequence detection during streaming
    • Real-time max tokens enforcement with token counting
  • 45 Responses API tests (router, error handling, history, advanced features)

Chat Completions Wrapper (Phase 5)

  • OpenAI Chat Completions compatibility layer (/v1/chat/completions)
  • Single source of truth architecture (wraps Responses API)
  • Automatic reasoning generation
  • Converts reasoning items to <think> tags for Open WebUI
  • Pipeline prefix preservation
  • System message support
  • Enhanced error types (RateLimitError, ContextLengthError)
  • 12 Chat Completions tests (router + streaming wrapper)

Application Infrastructure

  • FastAPI application factory pattern
  • CORS middleware with configurable origins
  • Global exception handlers:
    • AppException handler for custom errors
    • RequestValidationError handler for Pydantic validation
    • General exception handler for unexpected errors
  • Lifespan management for startup/shutdown
  • OpenAPI schema with interactive documentation
  • 16 main application tests

Testing Infrastructure

  • Comprehensive test suite: 95 tests, 78.95% coverage (up from 62%)
  • Async test support with pytest-asyncio
  • Test fixtures for sync and async clients
  • Integration tests for all API endpoints
  • Streaming functionality tests
  • Parameter validation tests
  • Error handling tests
  • Conversation history tests

Documentation

  • Complete README.md rewrite with hybrid architecture
  • Architecture diagrams and decision documentation
  • AGENTS.md with technical implementation details
  • API usage examples for all endpoints
  • Conversation history guide
  • Open WebUI integration instructions
  • Troubleshooting section
  • Implementation planning documents

Changed

  • Hybrid architecture with Responses API as primary endpoint
  • Chat Completions now wraps Responses API (no duplicate logic)
  • Enhanced error handling with OpenAI-compatible format
  • Improved streaming with word-by-word delivery
  • Better test organization with domain-based structure

Security

  • Minor version locking for all dependencies
  • All packages CVE-checked (as of 2025-12-06)
  • Environment variable protection via .gitignore
  • No known vulnerabilities in dependency tree
  • Input validation on all API endpoints

0.1.0 - 2025-12-06

Added

  • Project initialization
  • Python 3.12.11 environment
  • FastAPI 0.123.9 web framework
  • PydanticAI 1.27.0 dependency (ready for future integration)
  • Mock chat completions (lorem ipsum responses)
  • Mock model listing (mistral-nemo:latest)
  • Testing infrastructure (pytest, coverage, ruff, mypy)
  • Configuration management with pydantic-settings
  • CORS middleware
  • Exception handlers (OpenAI-compatible error format)