Created PHILOSOPHY.md to establish the foundational vision and architectural patterns for the Tatlock system. PHILOSOPHY.md: - Establishes Tatlock as a homelab butler coordinating expert agents - Defines the British household metaphor and two-tier architecture - Documents the Steward (request analysis) and Butler (orchestration) - Describes household staff roles (Handyman, Housekeeper, Secretary, Developer) - Explains real-time reasoning transparency for UX - Details model efficiency strategy (unified base model, specialized when needed) - Sets modification policy: only update for architectural deviations README.md: - Streamlined header with link to PHILOSOPHY.md - Simplified description to focus on practical usage - Updated documentation section to prioritize PHILOSOPHY.md - Maintained all usage examples and technical guides AGENTS.md: - Added prominent link to PHILOSOPHY.md at header - Emphasized that development should align with philosophy Documentation hierarchy: 1. PHILOSOPHY.md - Vision and architectural patterns (stable) 2. README.md - User guide and practical usage 3. AGENTS.md - LLM agent development guidelines 4. CHANGELOG.md - Version history 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>
17 KiB
LLM Agent Instructions
This document contains instructions and documentation references for AI assistants working with this codebase.
📖 Important: Before working on this project, read PHILOSOPHY.md to understand the system vision, architectural patterns, and design goals. All development should work towards realizing those patterns.
Project Overview
This project implements an OpenAI-compatible API endpoint using FastAPI, with streaming support. Currently returns mock responses - infrastructure prepared for future Ollama/PydanticAI integration.
Current State: Production-ready testing API with Responses API and Open WebUI integration Future Integration: PydanticAI for real LLM agents (tatlock model placeholder ready)
Current Architecture (As of 2025-12-06)
This project implements a hybrid architecture with the Responses API as the primary endpoint and Chat Completions as a compatibility wrapper:
Client (Open WebUI)
↓
Chat Completions (/v1/chat/completions) → Wrapper
↓
Responses API (/v1/responses) → Primary
↓
Agent Interface (lorem-tester, tatlock)
Key Architectural Decisions:
-
Single Source of Truth: All response generation happens in the Responses API
- Structured output with reasoning, function_call, and message items
- Real-time stop sequence and max tokens enforcement
- Conversation history tracking
- Context window management
-
Chat Completions Wrapper: Provides compatibility without duplicating logic
- Calls Responses API internally
- Automatically enables reasoning generation
- Converts reasoning items to
<think>tags for Open WebUI - Maintains OpenAI-compatible format
-
Agent Interface: Clean abstraction for multiple models
- lorem-tester: Full-featured mock agent with realistic behavior
- Reasoning summaries (adjustable effort levels)
- Random tool/function calls
- Error triggers for testing
- Temperature variation
- tatlock: Placeholder for future PydanticAI agent
- lorem-tester: Full-featured mock agent with realistic behavior
-
Hybrid Conversation History:
- Client MUST send full context in
inputarray (OpenAI compatible) - Server optionally tracks via
metadata.conversation_id - Auto-generates deterministic IDs from first message
- Supports future vector memory integration (Qdrant)
- Client MUST send full context in
Why This Architecture?
- Open WebUI Compatibility: Native Responses API support not yet in stable release
- Future-Proof: Easy migration when Open WebUI adds native support
- Testability: Full-featured mock agent (lorem-tester) for integration testing
- Clean Separation: Responses API as stable core, wrappers can change
Components
- FastAPI: Web framework for the API layer
- SSE-Starlette: Server-Sent Events for streaming responses
- Pydantic: Request/response validation with field validators
- Agent Interface: Abstract base class for model implementations
- Conversation History: Server-side tracking with configurable max turns
- Context Window: Token counting and management
- PydanticAI: Dependency installed, ready for tatlock agent implementation
Documentation References
Core Framework Documentation
FastAPI
- Official Documentation: https://fastapi.tiangolo.com/
- Version: 0.123.9 (Dec 2025)
- Key Topics:
- Path operations and routing
- Request/response models with Pydantic
- Dependency injection
- Background tasks
- WebSocket and streaming support
- PyPI: https://pypi.org/project/fastapi/
Uvicorn
- Official Documentation: https://www.uvicorn.org/
- Version: 0.38.0 (Oct 2025)
- Key Topics:
- ASGI server configuration
- Deployment settings
- Logging and monitoring
- SSL/TLS configuration
AI/LLM Integration
PydanticAI
- Official Documentation: https://ai.pydantic.dev/
- Version: 1.27.0 (Dec 2025)
- Status: Dependency installed, ready for future integration
- Key Topics (for future implementation):
- Agent creation and configuration
- LLM provider integration (Ollama support)
- Structured outputs with Pydantic
- Streaming responses
- Tool/function calling
- RunContext and dynamic configuration
- MCP server integration
- GitHub: https://github.com/pydantic/pydantic-ai
- PyPI: https://pypi.org/project/pydantic-ai/
Pydantic
- Official Documentation: https://docs.pydantic.dev/latest/
- Version: 2.11+ (Required for PydanticAI, currently using >=2.11,<2.13)
- Key Topics:
- Data validation and serialization
- Field types and validators
- Model configuration
- JSON schema generation
HTTP and Streaming
HTTPX
- Official Documentation: https://www.python-httpx.org/
- Version: 0.28.1
- Key Topics:
- Async HTTP client for Ollama communication
- Streaming responses
- Timeout configuration
- Connection pooling
SSE-Starlette
- GitHub: https://github.com/sysid/sse-starlette
- Version: 3.0.2 (Oct 2025)
- Key Topics:
- Server-Sent Events implementation
- Streaming event responses
- Integration with FastAPI/Starlette
Ollama Integration
Ollama API
- Official Documentation: https://github.com/ollama/ollama/blob/main/docs/api.md
- Status: Async client implemented in
src/ollama/client.py, ready for future integration - Key Topics (for future implementation):
- REST API endpoints
- Streaming responses
- Model management
- Generate and chat endpoints
- Model configuration
- Current Model Target: mistral-nemo:latest
OpenAI API Compatibility
OpenAI API Reference
- Official Documentation: https://platform.openai.com/docs/api-reference
- Implemented Endpoints:
- ✅
/v1/responses- Responses API (PRIMARY) with structured output- Reasoning items (thinking summaries)
- Function call items (tool execution)
- Message items (assistant responses)
- Full streaming support with SSE
- Stop sequence detection
- Max tokens enforcement
- Conversation history tracking
- ✅
/v1/chat/completions- Compatibility wrapper around Responses API- Converts reasoning to
<think>tags for Open WebUI - Automatically enables reasoning generation
- Maintains OpenAI-compatible format
- Supports streaming and non-streaming
- Converts reasoning to
- ✅
/v1/models- List available models (lorem-tester, tatlock)
- ✅
- Future Endpoints:
- 🚧
/v1/completions- Text completion (legacy) - 🚧
/v1/embeddings- Text embeddings
- 🚧
- Implemented Features:
- ✅ Responses API Format:
- Structured output items (reasoning, function_call, message)
- Extended thinking support
- Tool/function calling support
- Streaming with multiple event types
- ✅ Advanced Parameter Validation:
- Temperature: 0.0-2.0 with Pydantic validators
- Reasoning effort: none, minimal, low, medium, high, xhigh
- Max output tokens: positive integer enforcement
- Stop sequences: up to 4, non-empty strings
- ✅ Conversation History:
- Hybrid client/server approach
- Auto-generated conversation IDs
- Configurable max turns (default: 20)
- Placeholder for vector memory
- ✅ Context Management:
- Approximate token counting (~4 chars/token)
- Context window trimming
- Usage statistics
- ✅ Streaming Enforcement:
- Real-time stop sequence detection
- Real-time max tokens enforcement
- Word-by-word streaming with delays
- ✅ Error Handling:
- Custom exception types (RateLimitError, ContextLengthError)
- OpenAI-compatible error format
- Error triggers in lorem-tester for testing
- ✅ Testing Infrastructure:
- 95 tests (78.95% coverage)
- Unit tests for all components
- Integration tests for API endpoints
- Streaming tests for SSE functionality
- Main application and wrapper tests
- ✅ Responses API Format:
FastAPI Best Practices
This project follows best practices from github.com/zhanymkanov/fastapi-best-practices
Project Structure
Domain-Based Organization: Code is organized by domain/feature rather than by file type:
src/
├── agents/ # Agent interface and implementations
│ ├── base.py # Abstract AgentInterface
│ ├── lorem_tester.py # Full-featured mock agent
│ ├── tatlock.py # Placeholder for real agent
│ └── registry.py # ModelRegistry for agent management
├── responses/ # Responses API domain (PRIMARY)
│ ├── router.py # POST /v1/responses endpoint
│ ├── schemas.py # Request/response models with validators
│ ├── service.py # Response generation logic
│ ├── streaming.py # SSE streaming coordinator
│ ├── history.py # Conversation history management
│ └── context.py # Context window and token management
├── chat/ # Chat Completions domain (WRAPPER)
│ ├── router.py # POST /v1/chat/completions endpoint
│ ├── schemas.py # Chat request/response models
│ ├── service.py # Wraps Responses API, converts to <think> tags
│ ├── constants.py # Chat constants (roles, finish reasons)
│ └── __init__.py
├── models/ # Models listing domain
│ ├── router.py # GET /v1/models endpoint
│ ├── schemas.py # Model schemas
│ ├── service.py # Accesses ModelRegistry
│ └── __init__.py
├── core/ # Shared utilities
│ ├── config.py # Global configuration (BaseSettings)
│ ├── models.py # Custom base Pydantic models
│ ├── exceptions.py # Custom exceptions (RateLimitError, etc.)
│ ├── dependencies.py # Shared dependencies
│ └── router.py # Core routes (health, root)
├── ollama/ # Ollama client layer (not yet integrated)
│ ├── client.py # Async Ollama HTTP client
│ └── schemas.py # Ollama API models
└── main.py # Application factory & configuration
Key Architectural Principles:
- Single Source of Truth: Responses API handles all generation logic
- Wrapper Pattern: Chat Completions wraps Responses API without duplicating code
- Agent Abstraction: AgentInterface defines contract for all models
- Domain Separation: Each domain has its own router, schemas, service
- Service Layer: Business logic in services, not routers
- Type Safety: Pydantic models for ALL request/response validation
- Async First: All I/O operations use async/await
Async/Await Best Practices
Critical Understanding: FastAPI handles sync and async routes differently:
-
Async routes (
async def): Called directly in event loop- Use ONLY for non-blocking operations
- Perfect for
await httpx.get(), database queries, file I/O - NEVER use blocking calls like
time.sleep()- this blocks entire server
-
Sync routes (
def): Run in thread pool- Use for CPU-intensive work or blocking SDKs
- Blocking I/O won't freeze the event loop
- Example:
time.sleep(10)is safe here
Example:
@router.get("/terrible")
async def terrible():
time.sleep(10) # ❌ BLOCKS ENTIRE SERVER
@router.get("/good")
def good():
time.sleep(10) # ✅ Runs in thread pool
@router.get("/perfect")
async def perfect():
await asyncio.sleep(10) # ✅ Non-blocking async
For CPU-intensive tasks: Use separate worker processes (not threads) due to Python's GIL.
Pydantic Configuration
Custom Base Model: All schemas inherit from CustomBaseModel for consistent behavior:
# src/core/models.py
class CustomBaseModel(BaseModel):
model_config = ConfigDict(
json_encoders={datetime: datetime_to_iso_str},
populate_by_name=True,
use_enum_values=True,
validate_assignment=True,
)
def serializable_dict(self, **kwargs):
"""Return dict with only JSON-serializable fields."""
return jsonable_encoder(self.model_dump(**kwargs))
Benefits:
- Consistent datetime serialization across all responses
- Alias support for field name flexibility
- Easy JSON encoding for logging/debugging
Decoupled Settings: Split configuration by domain instead of one monolithic file:
# src/core/config.py - Global settings
class Config(BaseSettings):
DATABASE_URL: PostgresDsn
ENVIRONMENT: Environment
# src/chat/config.py - Chat-specific settings
class ChatConfig(BaseSettings):
MAX_TOKENS: int
DEFAULT_TEMPERATURE: float
Dependency Injection Patterns
Validation with Dependencies: Use dependencies for complex validations:
async def valid_post_id(post_id: UUID4) -> dict:
"""Validate post exists in database."""
post = await service.get_by_id(post_id)
if not post:
raise PostNotFound()
return post
@router.get("/posts/{post_id}")
async def get_post(post: dict = Depends(valid_post_id)):
return post # Already validated!
Chaining Dependencies: Build reusable validation layers:
async def valid_owned_post(
post: dict = Depends(valid_post_id),
token_data: dict = Depends(parse_jwt_data),
) -> dict:
if post["creator_id"] != token_data["user_id"]:
raise UserNotOwner()
return post
Dependency Caching: Dependencies are cached within request scope - FastAPI only executes each dependency once per request, even if used multiple times.
Application Factory Pattern
Main.py uses factory pattern for testability and configuration:
def create_application() -> FastAPI:
"""Create and configure FastAPI app."""
app = FastAPI(title=config.APP_NAME)
# Add middleware
app.add_middleware(CORSMiddleware, ...)
# Register exception handlers
register_exception_handlers(app)
# Include routers
app.include_router(chat_router, prefix="/v1")
return app
app = create_application()
Development Guidelines
Code Structure (Current Implementation)
- ✅ Use async/await for ALL I/O operations (database, HTTP, file access)
- ✅ Use sync (def) for blocking SDKs or CPU-intensive work
- ✅ Implement proper error handling and logging
- ✅ Follow dependency injection for validation and shared resources
- ✅ Use Pydantic models for ALL request/response validation
- ✅ Keep business logic in service modules, not routers
- ✅ Domain-based project structure (not file-type based)
Security Considerations
- ✅ Validate all inputs using Pydantic models
- ✅ Use environment variables for sensitive configuration
- ✅ Keep dependencies updated (all CVE-checked as of 2025-12-06)
- ✅ Minor version locking for supply chain protection
- 🚧 Implement rate limiting for API endpoints (future)
- 🚧 Add authentication/API keys (future)
Testing (Current Coverage: 78.95%, 95 tests)
- ✅ Integration tests for API endpoints
- ✅ Streaming functionality with 20s timeout protection
- ✅ Async test support with pytest-asyncio
- ✅ Validate OpenAI API compatibility
- ✅ Mock responses for all endpoints
- ✅ Main application tests (CORS, exception handlers, lifespan)
- ✅ Chat streaming wrapper tests
- 🚧 Future: Mock Ollama responses when integrated
Configuration
- ✅ Use
.envfiles for local development - ✅ Document all environment variables in README
- ✅ Provide sensible defaults where possible
- ✅ BaseSettings from pydantic-settings
- 🚧 Support container-based configuration (future)
Common Patterns
Streaming Response Pattern (✅ Implemented)
See src/chat/router.py for the current implementation:
from sse_starlette.sse import EventSourceResponse
from fastapi import FastAPI
async def event_generator():
# Currently yields mock lorem ipsum chunks
# Future: Stream from Ollama/PydanticAI
yield {"data": chunk.model_dump_json()}
yield {"data": "[DONE]"}
@app.post("/stream")
async def stream():
return EventSourceResponse(event_generator())
PydanticAI Agent Pattern (🚧 Future Reference)
For future integration when connecting to Ollama:
from pydantic_ai import Agent
agent = Agent(
'ollama:mistral-nemo', # Target model
# Configuration here
)
# Use the agent
result = await agent.run('Your prompt')
OpenAI-Compatible Response Format (✅ Implemented)
Current implementation in src/chat/schemas.py:
{
"id": "chatcmpl-123",
"object": "chat.completion.chunk",
"created": 1234567890,
"model": "mistral-nemo:latest",
"choices": [{
"index": 0,
"delta": {"content": "response"},
"finish_reason": None
}]
}
Update Policy
This document should be updated when:
- Package versions are upgraded
- New major features are added
- Breaking API changes occur
- Security vulnerabilities are discovered
Last updated: 2025-12-06