Files
tatlock/AGENTS.md
T

18 KiB

LLM Agent Instructions

This document contains instructions and documentation references for AI assistants working with this codebase.

📖 Important: Before working on this project, read PHILOSOPHY.md to understand the system vision, architectural patterns, and design goals. All development should work towards realizing those patterns.

Project Overview

This project implements an OpenAI-compatible API endpoint using FastAPI, with streaming support. Currently returns mock responses - infrastructure prepared for future Ollama/PydanticAI integration.

Current State: Production-ready testing API with Responses API and Open WebUI integration Future Integration: PydanticAI for real LLM agents (tatlock model placeholder ready)

Current Architecture (As of 2025-12-06)

This project implements the Orchestrator infrastructure layer with a hybrid API architecture:

Client (Open WebUI)
      ↓
Chat Completions (/v1/chat/completions) → Wrapper
      ↓
Responses API (/v1/responses) → Primary
      ↓
Agent Interface (lorem-tester, Tatlock)
      ↓
Mock Agents (lorem-tester) / Future: PydanticAI Agents (Tatlock, Steward, etc.)

Architectural Layers:

  1. The Orchestrator (Current Implementation)

    • FastAPI application providing the infrastructure
    • HTTP/SSE endpoints, streaming coordination
    • Conversation history and context management
    • OpenAI-compatible API surface
  2. Future: The Household (Phases 1-4)

    • Steward: First-tier LLM for request analysis (PydanticAI agent)
    • Tatlock: Second-tier LLM with butler personality (PydanticAI agent)
    • Expert Agents: Domain specialists (Librarian, Developer, Handyman, etc.)

Key Architectural Decisions:

  1. Single Source of Truth: All response generation happens in the Responses API

    • Structured output with reasoning, function_call, and message items
    • Real-time stop sequence and max tokens enforcement
    • Conversation history tracking
    • Context window management
  2. Chat Completions Wrapper: Provides compatibility without duplicating logic

    • Calls Responses API internally
    • Automatically enables reasoning generation
    • Converts reasoning items to <think> tags for Open WebUI
    • Maintains OpenAI-compatible format
  3. Agent Interface: Clean abstraction for multiple models

    • lorem-tester: Full-featured mock agent with realistic behavior
      • Reasoning summaries (adjustable effort levels)
      • Random tool/function calls
      • Error triggers for testing
      • Temperature variation
    • Tatlock: Advertised model name (currently mock, future: PydanticAI Butler agent)
  4. Hybrid Conversation History:

    • Client MUST send full context in input array (OpenAI compatible)
    • Server optionally tracks via metadata.conversation_id
    • Auto-generates deterministic IDs from first message
    • Supports future vector memory integration (Qdrant)

Why This Architecture?

  • Open WebUI Compatibility: Native Responses API support not yet in stable release
  • Future-Proof: Easy migration when Open WebUI adds native support
  • Testability: Full-featured mock agent (lorem-tester) for integration testing
  • Clean Separation: Responses API as stable core, wrappers can change

Components

  • FastAPI: Web framework for the API layer
  • SSE-Starlette: Server-Sent Events for streaming responses
  • Pydantic: Request/response validation with field validators
  • Agent Interface: Abstract base class for model implementations
  • Conversation History: Server-side tracking with configurable max turns
  • Context Window: Token counting and management
  • PydanticAI: Dependency installed, ready for tatlock agent implementation

Documentation References

Core Framework Documentation

FastAPI

Uvicorn

  • Official Documentation: https://www.uvicorn.org/
  • Version: 0.38.0 (Oct 2025)
  • Key Topics:
    • ASGI server configuration
    • Deployment settings
    • Logging and monitoring
    • SSL/TLS configuration

AI/LLM Integration

PydanticAI

Pydantic

  • Official Documentation: https://docs.pydantic.dev/latest/
  • Version: 2.11+ (Required for PydanticAI, currently using >=2.11,<2.13)
  • Key Topics:
    • Data validation and serialization
    • Field types and validators
    • Model configuration
    • JSON schema generation

HTTP and Streaming

HTTPX

  • Official Documentation: https://www.python-httpx.org/
  • Version: 0.28.1
  • Key Topics:
    • Async HTTP client for Ollama communication
    • Streaming responses
    • Timeout configuration
    • Connection pooling

SSE-Starlette

Ollama Integration

Ollama API

  • Official Documentation: https://github.com/ollama/ollama/blob/main/docs/api.md
  • Status: Async client implemented in src/ollama/client.py, ready for future integration
  • Key Topics (for future implementation):
    • REST API endpoints
    • Streaming responses
    • Model management
    • Generate and chat endpoints
    • Model configuration
  • Current Model Target: mistral-nemo:latest

OpenAI API Compatibility

OpenAI API Reference

  • Official Documentation: https://platform.openai.com/docs/api-reference
  • Implemented Endpoints:
    • /v1/responses - Responses API (PRIMARY) with structured output
      • Reasoning items (thinking summaries)
      • Function call items (tool execution)
      • Message items (assistant responses)
      • Full streaming support with SSE
      • Stop sequence detection
      • Max tokens enforcement
      • Conversation history tracking
    • /v1/chat/completions - Compatibility wrapper around Responses API
      • Converts reasoning to <think> tags for Open WebUI
      • Automatically enables reasoning generation
      • Maintains OpenAI-compatible format
      • Supports streaming and non-streaming
    • /v1/models - List available models (lorem-tester, tatlock)
  • Future Endpoints:
    • 🚧 /v1/completions - Text completion (legacy)
    • 🚧 /v1/embeddings - Text embeddings
  • Implemented Features:
    • Responses API Format:
      • Structured output items (reasoning, function_call, message)
      • Extended thinking support
      • Tool/function calling support
      • Streaming with multiple event types
    • Advanced Parameter Validation:
      • Temperature: 0.0-2.0 with Pydantic validators
      • Reasoning effort: none, minimal, low, medium, high, xhigh
      • Max output tokens: positive integer enforcement
      • Stop sequences: up to 4, non-empty strings
    • Conversation History:
      • Hybrid client/server approach
      • Auto-generated conversation IDs
      • Configurable max turns (default: 20)
      • Placeholder for vector memory
    • Context Management:
      • Approximate token counting (~4 chars/token)
      • Context window trimming
      • Usage statistics
    • Streaming Enforcement:
      • Real-time stop sequence detection
      • Real-time max tokens enforcement
      • Word-by-word streaming with delays
    • Error Handling:
      • Custom exception types (RateLimitError, ContextLengthError)
      • OpenAI-compatible error format
      • Error triggers in lorem-tester for testing
    • Testing Infrastructure:
      • 95 tests (78.95% coverage)
      • Unit tests for all components
      • Integration tests for API endpoints
      • Streaming tests for SSE functionality
      • Main application and wrapper tests

FastAPI Best Practices

This project follows best practices from github.com/zhanymkanov/fastapi-best-practices

Project Structure

Domain-Based Organization: Code is organized by domain/feature rather than by file type:

src/
├── agents/                # Agent interface and implementations
│   ├── base.py            # Abstract AgentInterface
│   ├── lorem_tester.py    # Full-featured mock agent
│   ├── tatlock.py         # Placeholder for real agent
│   └── registry.py        # ModelRegistry for agent management
├── responses/             # Responses API domain (PRIMARY)
│   ├── router.py          # POST /v1/responses endpoint
│   ├── schemas.py         # Request/response models with validators
│   ├── service.py         # Response generation logic
│   ├── streaming.py       # SSE streaming coordinator
│   ├── history.py         # Conversation history management
│   └── context.py         # Context window and token management
├── chat/                  # Chat Completions domain (WRAPPER)
│   ├── router.py          # POST /v1/chat/completions endpoint
│   ├── schemas.py         # Chat request/response models
│   ├── service.py         # Wraps Responses API, converts to <think> tags
│   ├── constants.py       # Chat constants (roles, finish reasons)
│   └── __init__.py
├── models/                # Models listing domain
│   ├── router.py          # GET /v1/models endpoint
│   ├── schemas.py         # Model schemas
│   ├── service.py         # Accesses ModelRegistry
│   └── __init__.py
├── core/                  # Shared utilities
│   ├── config.py          # Global configuration (BaseSettings)
│   ├── models.py          # Custom base Pydantic models
│   ├── exceptions.py      # Custom exceptions (RateLimitError, etc.)
│   ├── dependencies.py    # Shared dependencies
│   └── router.py          # Core routes (health, root)
├── ollama/                # Ollama client layer (not yet integrated)
│   ├── client.py          # Async Ollama HTTP client
│   └── schemas.py         # Ollama API models
└── main.py                # Application factory & configuration

Key Architectural Principles:

  • Single Source of Truth: Responses API handles all generation logic
  • Wrapper Pattern: Chat Completions wraps Responses API without duplicating code
  • Agent Abstraction: AgentInterface defines contract for all models
  • Domain Separation: Each domain has its own router, schemas, service
  • Service Layer: Business logic in services, not routers
  • Type Safety: Pydantic models for ALL request/response validation
  • Async First: All I/O operations use async/await

Async/Await Best Practices

Critical Understanding: FastAPI handles sync and async routes differently:

  • Async routes (async def): Called directly in event loop

    • Use ONLY for non-blocking operations
    • Perfect for await httpx.get(), database queries, file I/O
    • NEVER use blocking calls like time.sleep() - this blocks entire server
  • Sync routes (def): Run in thread pool

    • Use for CPU-intensive work or blocking SDKs
    • Blocking I/O won't freeze the event loop
    • Example: time.sleep(10) is safe here

Example:

@router.get("/terrible")
async def terrible():
    time.sleep(10)  # ❌ BLOCKS ENTIRE SERVER

@router.get("/good")
def good():
    time.sleep(10)  # ✅ Runs in thread pool

@router.get("/perfect")
async def perfect():
    await asyncio.sleep(10)  # ✅ Non-blocking async

For CPU-intensive tasks: Use separate worker processes (not threads) due to Python's GIL.

Pydantic Configuration

Custom Base Model: All schemas inherit from CustomBaseModel for consistent behavior:

# src/core/models.py
class CustomBaseModel(BaseModel):
    model_config = ConfigDict(
        json_encoders={datetime: datetime_to_iso_str},
        populate_by_name=True,
        use_enum_values=True,
        validate_assignment=True,
    )

    def serializable_dict(self, **kwargs):
        """Return dict with only JSON-serializable fields."""
        return jsonable_encoder(self.model_dump(**kwargs))

Benefits:

  • Consistent datetime serialization across all responses
  • Alias support for field name flexibility
  • Easy JSON encoding for logging/debugging

Decoupled Settings: Split configuration by domain instead of one monolithic file:

# src/core/config.py - Global settings
class Config(BaseSettings):
    DATABASE_URL: PostgresDsn
    ENVIRONMENT: Environment

# src/chat/config.py - Chat-specific settings
class ChatConfig(BaseSettings):
    MAX_TOKENS: int
    DEFAULT_TEMPERATURE: float

Dependency Injection Patterns

Validation with Dependencies: Use dependencies for complex validations:

async def valid_post_id(post_id: UUID4) -> dict:
    """Validate post exists in database."""
    post = await service.get_by_id(post_id)
    if not post:
        raise PostNotFound()
    return post

@router.get("/posts/{post_id}")
async def get_post(post: dict = Depends(valid_post_id)):
    return post  # Already validated!

Chaining Dependencies: Build reusable validation layers:

async def valid_owned_post(
    post: dict = Depends(valid_post_id),
    token_data: dict = Depends(parse_jwt_data),
) -> dict:
    if post["creator_id"] != token_data["user_id"]:
        raise UserNotOwner()
    return post

Dependency Caching: Dependencies are cached within request scope - FastAPI only executes each dependency once per request, even if used multiple times.

Application Factory Pattern

Main.py uses factory pattern for testability and configuration:

def create_application() -> FastAPI:
    """Create and configure FastAPI app."""
    app = FastAPI(title=config.APP_NAME)

    # Add middleware
    app.add_middleware(CORSMiddleware, ...)

    # Register exception handlers
    register_exception_handlers(app)

    # Include routers
    app.include_router(chat_router, prefix="/v1")

    return app

app = create_application()

Development Guidelines

Git Workflow

IMPORTANT: Do NOT handle git commits or pushes automatically. Wait for explicit user instruction before:

  • Running git add
  • Running git commit
  • Running git push
  • Creating or pushing tags

The user will manage git operations themselves unless they specifically request assistance.

Code Structure (Current Implementation)

  • Use async/await for ALL I/O operations (database, HTTP, file access)
  • Use sync (def) for blocking SDKs or CPU-intensive work
  • Implement proper error handling and logging
  • Follow dependency injection for validation and shared resources
  • Use Pydantic models for ALL request/response validation
  • Keep business logic in service modules, not routers
  • Domain-based project structure (not file-type based)

Security Considerations

  • Validate all inputs using Pydantic models
  • Use environment variables for sensitive configuration
  • Keep dependencies updated (all CVE-checked as of 2025-12-06)
  • Minor version locking for supply chain protection
  • 🚧 Implement rate limiting for API endpoints (future)
  • 🚧 Add authentication/API keys (future)

Testing (Current Coverage: 78.95%, 95 tests)

  • Integration tests for API endpoints
  • Streaming functionality with 20s timeout protection
  • Async test support with pytest-asyncio
  • Validate OpenAI API compatibility
  • Mock responses for all endpoints
  • Main application tests (CORS, exception handlers, lifespan)
  • Chat streaming wrapper tests
  • 🚧 Future: Mock Ollama responses when integrated

Configuration

  • Use .env files for local development
  • Document all environment variables in README
  • Provide sensible defaults where possible
  • BaseSettings from pydantic-settings
  • 🚧 Support container-based configuration (future)

Common Patterns

Streaming Response Pattern ( Implemented)

See src/chat/router.py for the current implementation:

from sse_starlette.sse import EventSourceResponse
from fastapi import FastAPI

async def event_generator():
    # Currently yields mock lorem ipsum chunks
    # Future: Stream from Ollama/PydanticAI
    yield {"data": chunk.model_dump_json()}
    yield {"data": "[DONE]"}

@app.post("/stream")
async def stream():
    return EventSourceResponse(event_generator())

PydanticAI Agent Pattern (🚧 Future Reference)

For future integration when connecting to Ollama:

from pydantic_ai import Agent

agent = Agent(
    'ollama:mistral-nemo',  # Target model
    # Configuration here
)

# Use the agent
result = await agent.run('Your prompt')

OpenAI-Compatible Response Format ( Implemented)

Current implementation in src/chat/schemas.py:

{
    "id": "chatcmpl-123",
    "object": "chat.completion.chunk",
    "created": 1234567890,
    "model": "mistral-nemo:latest",
    "choices": [{
        "index": 0,
        "delta": {"content": "response"},
        "finish_reason": None
    }]
}

Update Policy

This document should be updated when:

  • Package versions are upgraded
  • New major features are added
  • Breaking API changes occur
  • Security vulnerabilities are discovered

Last updated: 2025-12-06