Files
tatlock/AGENTS.md
T
jpmschweitzerandClaude 661db0a672 Update documentation with hybrid architecture
README.md:
- Complete rewrite with hybrid architecture documentation
- Architecture diagram showing wrapper pattern
- Detailed feature list for all implemented phases
- API usage examples for Responses and Chat Completions
- Conversation history usage guide
- Open WebUI integration instructions
- Comprehensive troubleshooting section
- Updated project structure
- Deployment considerations

AGENTS.md:
- Current architecture section (as of 2025-12-06)
- Hybrid architecture explanation
- Key architectural decisions documented
- Agent interface design patterns
- Conversation history approach
- Context window management
- Testing infrastructure details
- Implementation status updates
- Coverage statistics (78.95%, 95 tests)

Key Documentation Themes:
- Single source of truth: Responses API
- Wrapper pattern for Chat Completions
- Hybrid conversation history approach
- Clean agent abstraction
- Production-ready testing infrastructure

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-12-06 19:39:32 +01:00

16 KiB

LLM Agent Instructions

This document contains instructions and documentation references for AI assistants working with this codebase.

Project Overview

This project implements an OpenAI-compatible API endpoint using FastAPI, with streaming support. Currently returns mock responses - infrastructure prepared for future Ollama/PydanticAI integration.

Current State: Production-ready testing API with Responses API and Open WebUI integration Future Integration: PydanticAI for real LLM agents (tatlock model placeholder ready)

Current Architecture (As of 2025-12-06)

This project implements a hybrid architecture with the Responses API as the primary endpoint and Chat Completions as a compatibility wrapper:

Client (Open WebUI)
      ↓
Chat Completions (/v1/chat/completions) → Wrapper
      ↓
Responses API (/v1/responses) → Primary
      ↓
Agent Interface (lorem-tester, tatlock)

Key Architectural Decisions:

  1. Single Source of Truth: All response generation happens in the Responses API

    • Structured output with reasoning, function_call, and message items
    • Real-time stop sequence and max tokens enforcement
    • Conversation history tracking
    • Context window management
  2. Chat Completions Wrapper: Provides compatibility without duplicating logic

    • Calls Responses API internally
    • Automatically enables reasoning generation
    • Converts reasoning items to <think> tags for Open WebUI
    • Maintains OpenAI-compatible format
  3. Agent Interface: Clean abstraction for multiple models

    • lorem-tester: Full-featured mock agent with realistic behavior
      • Reasoning summaries (adjustable effort levels)
      • Random tool/function calls
      • Error triggers for testing
      • Temperature variation
    • tatlock: Placeholder for future PydanticAI agent
  4. Hybrid Conversation History:

    • Client MUST send full context in input array (OpenAI compatible)
    • Server optionally tracks via metadata.conversation_id
    • Auto-generates deterministic IDs from first message
    • Supports future vector memory integration (Qdrant)

Why This Architecture?

  • Open WebUI Compatibility: Native Responses API support not yet in stable release
  • Future-Proof: Easy migration when Open WebUI adds native support
  • Testability: Full-featured mock agent (lorem-tester) for integration testing
  • Clean Separation: Responses API as stable core, wrappers can change

Components

  • FastAPI: Web framework for the API layer
  • SSE-Starlette: Server-Sent Events for streaming responses
  • Pydantic: Request/response validation with field validators
  • Agent Interface: Abstract base class for model implementations
  • Conversation History: Server-side tracking with configurable max turns
  • Context Window: Token counting and management
  • PydanticAI: Dependency installed, ready for tatlock agent implementation

Documentation References

Core Framework Documentation

FastAPI

Uvicorn

  • Official Documentation: https://www.uvicorn.org/
  • Version: 0.38.0 (Oct 2025)
  • Key Topics:
    • ASGI server configuration
    • Deployment settings
    • Logging and monitoring
    • SSL/TLS configuration

AI/LLM Integration

PydanticAI

Pydantic

  • Official Documentation: https://docs.pydantic.dev/latest/
  • Version: 2.11+ (Required for PydanticAI, currently using >=2.11,<2.13)
  • Key Topics:
    • Data validation and serialization
    • Field types and validators
    • Model configuration
    • JSON schema generation

HTTP and Streaming

HTTPX

  • Official Documentation: https://www.python-httpx.org/
  • Version: 0.28.1
  • Key Topics:
    • Async HTTP client for Ollama communication
    • Streaming responses
    • Timeout configuration
    • Connection pooling

SSE-Starlette

Ollama Integration

Ollama API

  • Official Documentation: https://github.com/ollama/ollama/blob/main/docs/api.md
  • Status: Async client implemented in src/ollama/client.py, ready for future integration
  • Key Topics (for future implementation):
    • REST API endpoints
    • Streaming responses
    • Model management
    • Generate and chat endpoints
    • Model configuration
  • Current Model Target: mistral-nemo:latest

OpenAI API Compatibility

OpenAI API Reference

  • Official Documentation: https://platform.openai.com/docs/api-reference
  • Implemented Endpoints:
    • /v1/responses - Responses API (PRIMARY) with structured output
      • Reasoning items (thinking summaries)
      • Function call items (tool execution)
      • Message items (assistant responses)
      • Full streaming support with SSE
      • Stop sequence detection
      • Max tokens enforcement
      • Conversation history tracking
    • /v1/chat/completions - Compatibility wrapper around Responses API
      • Converts reasoning to <think> tags for Open WebUI
      • Automatically enables reasoning generation
      • Maintains OpenAI-compatible format
      • Supports streaming and non-streaming
    • /v1/models - List available models (lorem-tester, tatlock)
  • Future Endpoints:
    • 🚧 /v1/completions - Text completion (legacy)
    • 🚧 /v1/embeddings - Text embeddings
  • Implemented Features:
    • Responses API Format:
      • Structured output items (reasoning, function_call, message)
      • Extended thinking support
      • Tool/function calling support
      • Streaming with multiple event types
    • Advanced Parameter Validation:
      • Temperature: 0.0-2.0 with Pydantic validators
      • Reasoning effort: none, minimal, low, medium, high, xhigh
      • Max output tokens: positive integer enforcement
      • Stop sequences: up to 4, non-empty strings
    • Conversation History:
      • Hybrid client/server approach
      • Auto-generated conversation IDs
      • Configurable max turns (default: 20)
      • Placeholder for vector memory
    • Context Management:
      • Approximate token counting (~4 chars/token)
      • Context window trimming
      • Usage statistics
    • Streaming Enforcement:
      • Real-time stop sequence detection
      • Real-time max tokens enforcement
      • Word-by-word streaming with delays
    • Error Handling:
      • Custom exception types (RateLimitError, ContextLengthError)
      • OpenAI-compatible error format
      • Error triggers in lorem-tester for testing
    • Testing Infrastructure:
      • 75 tests (78.95% coverage)
      • Unit tests for all components
      • Integration tests for API endpoints
      • Streaming tests for SSE functionality

FastAPI Best Practices

This project follows best practices from github.com/zhanymkanov/fastapi-best-practices

Project Structure

Domain-Based Organization: Code is organized by domain/feature rather than by file type:

src/
├── agents/                # Agent interface and implementations
│   ├── base.py            # Abstract AgentInterface
│   ├── lorem_tester.py    # Full-featured mock agent
│   ├── tatlock.py         # Placeholder for real agent
│   └── registry.py        # ModelRegistry for agent management
├── responses/             # Responses API domain (PRIMARY)
│   ├── router.py          # POST /v1/responses endpoint
│   ├── schemas.py         # Request/response models with validators
│   ├── service.py         # Response generation logic
│   ├── streaming.py       # SSE streaming coordinator
│   ├── history.py         # Conversation history management
│   └── context.py         # Context window and token management
├── chat/                  # Chat Completions domain (WRAPPER)
│   ├── router.py          # POST /v1/chat/completions endpoint
│   ├── schemas.py         # Chat request/response models
│   ├── service.py         # Wraps Responses API, converts to <think> tags
│   ├── constants.py       # Chat constants (roles, finish reasons)
│   └── __init__.py
├── models/                # Models listing domain
│   ├── router.py          # GET /v1/models endpoint
│   ├── schemas.py         # Model schemas
│   ├── service.py         # Accesses ModelRegistry
│   └── __init__.py
├── core/                  # Shared utilities
│   ├── config.py          # Global configuration (BaseSettings)
│   ├── models.py          # Custom base Pydantic models
│   ├── exceptions.py      # Custom exceptions (RateLimitError, etc.)
│   ├── dependencies.py    # Shared dependencies
│   └── router.py          # Core routes (health, root)
├── ollama/                # Ollama client layer (not yet integrated)
│   ├── client.py          # Async Ollama HTTP client
│   └── schemas.py         # Ollama API models
└── main.py                # Application factory & configuration

Key Architectural Principles:

  • Single Source of Truth: Responses API handles all generation logic
  • Wrapper Pattern: Chat Completions wraps Responses API without duplicating code
  • Agent Abstraction: AgentInterface defines contract for all models
  • Domain Separation: Each domain has its own router, schemas, service
  • Service Layer: Business logic in services, not routers
  • Type Safety: Pydantic models for ALL request/response validation
  • Async First: All I/O operations use async/await

Async/Await Best Practices

Critical Understanding: FastAPI handles sync and async routes differently:

  • Async routes (async def): Called directly in event loop

    • Use ONLY for non-blocking operations
    • Perfect for await httpx.get(), database queries, file I/O
    • NEVER use blocking calls like time.sleep() - this blocks entire server
  • Sync routes (def): Run in thread pool

    • Use for CPU-intensive work or blocking SDKs
    • Blocking I/O won't freeze the event loop
    • Example: time.sleep(10) is safe here

Example:

@router.get("/terrible")
async def terrible():
    time.sleep(10)  # ❌ BLOCKS ENTIRE SERVER

@router.get("/good")
def good():
    time.sleep(10)  # ✅ Runs in thread pool

@router.get("/perfect")
async def perfect():
    await asyncio.sleep(10)  # ✅ Non-blocking async

For CPU-intensive tasks: Use separate worker processes (not threads) due to Python's GIL.

Pydantic Configuration

Custom Base Model: All schemas inherit from CustomBaseModel for consistent behavior:

# src/core/models.py
class CustomBaseModel(BaseModel):
    model_config = ConfigDict(
        json_encoders={datetime: datetime_to_iso_str},
        populate_by_name=True,
        use_enum_values=True,
        validate_assignment=True,
    )

    def serializable_dict(self, **kwargs):
        """Return dict with only JSON-serializable fields."""
        return jsonable_encoder(self.model_dump(**kwargs))

Benefits:

  • Consistent datetime serialization across all responses
  • Alias support for field name flexibility
  • Easy JSON encoding for logging/debugging

Decoupled Settings: Split configuration by domain instead of one monolithic file:

# src/core/config.py - Global settings
class Config(BaseSettings):
    DATABASE_URL: PostgresDsn
    ENVIRONMENT: Environment

# src/chat/config.py - Chat-specific settings
class ChatConfig(BaseSettings):
    MAX_TOKENS: int
    DEFAULT_TEMPERATURE: float

Dependency Injection Patterns

Validation with Dependencies: Use dependencies for complex validations:

async def valid_post_id(post_id: UUID4) -> dict:
    """Validate post exists in database."""
    post = await service.get_by_id(post_id)
    if not post:
        raise PostNotFound()
    return post

@router.get("/posts/{post_id}")
async def get_post(post: dict = Depends(valid_post_id)):
    return post  # Already validated!

Chaining Dependencies: Build reusable validation layers:

async def valid_owned_post(
    post: dict = Depends(valid_post_id),
    token_data: dict = Depends(parse_jwt_data),
) -> dict:
    if post["creator_id"] != token_data["user_id"]:
        raise UserNotOwner()
    return post

Dependency Caching: Dependencies are cached within request scope - FastAPI only executes each dependency once per request, even if used multiple times.

Application Factory Pattern

Main.py uses factory pattern for testability and configuration:

def create_application() -> FastAPI:
    """Create and configure FastAPI app."""
    app = FastAPI(title=config.APP_NAME)

    # Add middleware
    app.add_middleware(CORSMiddleware, ...)

    # Register exception handlers
    register_exception_handlers(app)

    # Include routers
    app.include_router(chat_router, prefix="/v1")

    return app

app = create_application()

Development Guidelines

Code Structure (Current Implementation)

  • Use async/await for ALL I/O operations (database, HTTP, file access)
  • Use sync (def) for blocking SDKs or CPU-intensive work
  • Implement proper error handling and logging
  • Follow dependency injection for validation and shared resources
  • Use Pydantic models for ALL request/response validation
  • Keep business logic in service modules, not routers
  • Domain-based project structure (not file-type based)

Security Considerations

  • Validate all inputs using Pydantic models
  • Use environment variables for sensitive configuration
  • Keep dependencies updated (all CVE-checked as of 2025-12-06)
  • Minor version locking for supply chain protection
  • 🚧 Implement rate limiting for API endpoints (future)
  • 🚧 Add authentication/API keys (future)

Testing (Current Coverage: 62%)

  • Integration tests for API endpoints
  • Streaming functionality with 20s timeout protection
  • Async test support with pytest-asyncio
  • Validate OpenAI API compatibility
  • Mock responses for all endpoints
  • 🚧 Future: Mock Ollama responses when integrated

Configuration

  • Use .env files for local development
  • Document all environment variables in README
  • Provide sensible defaults where possible
  • BaseSettings from pydantic-settings
  • 🚧 Support container-based configuration (future)

Common Patterns

Streaming Response Pattern ( Implemented)

See src/chat/router.py for the current implementation:

from sse_starlette.sse import EventSourceResponse
from fastapi import FastAPI

async def event_generator():
    # Currently yields mock lorem ipsum chunks
    # Future: Stream from Ollama/PydanticAI
    yield {"data": chunk.model_dump_json()}
    yield {"data": "[DONE]"}

@app.post("/stream")
async def stream():
    return EventSourceResponse(event_generator())

PydanticAI Agent Pattern (🚧 Future Reference)

For future integration when connecting to Ollama:

from pydantic_ai import Agent

agent = Agent(
    'ollama:mistral-nemo',  # Target model
    # Configuration here
)

# Use the agent
result = await agent.run('Your prompt')

OpenAI-Compatible Response Format ( Implemented)

Current implementation in src/chat/schemas.py:

{
    "id": "chatcmpl-123",
    "object": "chat.completion.chunk",
    "created": 1234567890,
    "model": "mistral-nemo:latest",
    "choices": [{
        "index": 0,
        "delta": {"content": "response"},
        "finish_reason": None
    }]
}

Update Policy

This document should be updated when:

  • Package versions are upgraded
  • New major features are added
  • Breaking API changes occur
  • Security vulnerabilities are discovered

Last updated: 2025-12-06