# LLM Agent Instructions This document contains instructions and documentation references for AI assistants working with this codebase. ## Project Overview This project implements an OpenAI-compatible API endpoint using FastAPI, with streaming support. Currently returns mock responses - infrastructure prepared for future Ollama/PydanticAI integration. **Current State**: Production-ready testing API with Responses API and Open WebUI integration **Future Integration**: PydanticAI for real LLM agents (tatlock model placeholder ready) ### Current Architecture (As of 2025-12-06) This project implements a **hybrid architecture** with the Responses API as the primary endpoint and Chat Completions as a compatibility wrapper: ``` Client (Open WebUI) ↓ Chat Completions (/v1/chat/completions) → Wrapper ↓ Responses API (/v1/responses) → Primary ↓ Agent Interface (lorem-tester, tatlock) ``` **Key Architectural Decisions:** 1. **Single Source of Truth**: All response generation happens in the Responses API - Structured output with reasoning, function_call, and message items - Real-time stop sequence and max tokens enforcement - Conversation history tracking - Context window management 2. **Chat Completions Wrapper**: Provides compatibility without duplicating logic - Calls Responses API internally - Automatically enables reasoning generation - Converts reasoning items to `` tags for Open WebUI - Maintains OpenAI-compatible format 3. **Agent Interface**: Clean abstraction for multiple models - **lorem-tester**: Full-featured mock agent with realistic behavior - Reasoning summaries (adjustable effort levels) - Random tool/function calls - Error triggers for testing - Temperature variation - **tatlock**: Placeholder for future PydanticAI agent 4. **Hybrid Conversation History**: - Client MUST send full context in `input` array (OpenAI compatible) - Server optionally tracks via `metadata.conversation_id` - Auto-generates deterministic IDs from first message - Supports future vector memory integration (Qdrant) **Why This Architecture?** - **Open WebUI Compatibility**: Native Responses API support not yet in stable release - **Future-Proof**: Easy migration when Open WebUI adds native support - **Testability**: Full-featured mock agent (lorem-tester) for integration testing - **Clean Separation**: Responses API as stable core, wrappers can change ### Components - **FastAPI**: Web framework for the API layer - **SSE-Starlette**: Server-Sent Events for streaming responses - **Pydantic**: Request/response validation with field validators - **Agent Interface**: Abstract base class for model implementations - **Conversation History**: Server-side tracking with configurable max turns - **Context Window**: Token counting and management - **PydanticAI**: Dependency installed, ready for tatlock agent implementation ## Documentation References ### Core Framework Documentation #### FastAPI - **Official Documentation**: https://fastapi.tiangolo.com/ - **Version**: 0.123.9 (Dec 2025) - **Key Topics**: - Path operations and routing - Request/response models with Pydantic - Dependency injection - Background tasks - WebSocket and streaming support - **PyPI**: https://pypi.org/project/fastapi/ #### Uvicorn - **Official Documentation**: https://www.uvicorn.org/ - **Version**: 0.38.0 (Oct 2025) - **Key Topics**: - ASGI server configuration - Deployment settings - Logging and monitoring - SSL/TLS configuration ### AI/LLM Integration #### PydanticAI - **Official Documentation**: https://ai.pydantic.dev/ - **Version**: 1.27.0 (Dec 2025) - **Status**: Dependency installed, ready for future integration - **Key Topics** (for future implementation): - Agent creation and configuration - LLM provider integration (Ollama support) - Structured outputs with Pydantic - Streaming responses - Tool/function calling - RunContext and dynamic configuration - MCP server integration - **GitHub**: https://github.com/pydantic/pydantic-ai - **PyPI**: https://pypi.org/project/pydantic-ai/ #### Pydantic - **Official Documentation**: https://docs.pydantic.dev/latest/ - **Version**: 2.11+ (Required for PydanticAI, currently using >=2.11,<2.13) - **Key Topics**: - Data validation and serialization - Field types and validators - Model configuration - JSON schema generation ### HTTP and Streaming #### HTTPX - **Official Documentation**: https://www.python-httpx.org/ - **Version**: 0.28.1 - **Key Topics**: - Async HTTP client for Ollama communication - Streaming responses - Timeout configuration - Connection pooling #### SSE-Starlette - **GitHub**: https://github.com/sysid/sse-starlette - **Version**: 3.0.2 (Oct 2025) - **Key Topics**: - Server-Sent Events implementation - Streaming event responses - Integration with FastAPI/Starlette ### Ollama Integration #### Ollama API - **Official Documentation**: https://github.com/ollama/ollama/blob/main/docs/api.md - **Status**: Async client implemented in `src/ollama/client.py`, ready for future integration - **Key Topics** (for future implementation): - REST API endpoints - Streaming responses - Model management - Generate and chat endpoints - Model configuration - **Current Model Target**: mistral-nemo:latest ### OpenAI API Compatibility #### OpenAI API Reference - **Official Documentation**: https://platform.openai.com/docs/api-reference - **Implemented Endpoints**: - ✅ `/v1/responses` - **Responses API (PRIMARY)** with structured output - Reasoning items (thinking summaries) - Function call items (tool execution) - Message items (assistant responses) - Full streaming support with SSE - Stop sequence detection - Max tokens enforcement - Conversation history tracking - ✅ `/v1/chat/completions` - **Compatibility wrapper** around Responses API - Converts reasoning to `` tags for Open WebUI - Automatically enables reasoning generation - Maintains OpenAI-compatible format - Supports streaming and non-streaming - ✅ `/v1/models` - List available models (lorem-tester, tatlock) - **Future Endpoints**: - 🚧 `/v1/completions` - Text completion (legacy) - 🚧 `/v1/embeddings` - Text embeddings - **Implemented Features**: - ✅ **Responses API Format**: - Structured output items (reasoning, function_call, message) - Extended thinking support - Tool/function calling support - Streaming with multiple event types - ✅ **Advanced Parameter Validation**: - Temperature: 0.0-2.0 with Pydantic validators - Reasoning effort: none, minimal, low, medium, high, xhigh - Max output tokens: positive integer enforcement - Stop sequences: up to 4, non-empty strings - ✅ **Conversation History**: - Hybrid client/server approach - Auto-generated conversation IDs - Configurable max turns (default: 20) - Placeholder for vector memory - ✅ **Context Management**: - Approximate token counting (~4 chars/token) - Context window trimming - Usage statistics - ✅ **Streaming Enforcement**: - Real-time stop sequence detection - Real-time max tokens enforcement - Word-by-word streaming with delays - ✅ **Error Handling**: - Custom exception types (RateLimitError, ContextLengthError) - OpenAI-compatible error format - Error triggers in lorem-tester for testing - ✅ **Testing Infrastructure**: - 75 tests (78.95% coverage) - Unit tests for all components - Integration tests for API endpoints - Streaming tests for SSE functionality ## FastAPI Best Practices This project follows best practices from [github.com/zhanymkanov/fastapi-best-practices](https://github.com/zhanymkanov/fastapi-best-practices) ### Project Structure **Domain-Based Organization**: Code is organized by domain/feature rather than by file type: ``` src/ ├── agents/ # Agent interface and implementations │ ├── base.py # Abstract AgentInterface │ ├── lorem_tester.py # Full-featured mock agent │ ├── tatlock.py # Placeholder for real agent │ └── registry.py # ModelRegistry for agent management ├── responses/ # Responses API domain (PRIMARY) │ ├── router.py # POST /v1/responses endpoint │ ├── schemas.py # Request/response models with validators │ ├── service.py # Response generation logic │ ├── streaming.py # SSE streaming coordinator │ ├── history.py # Conversation history management │ └── context.py # Context window and token management ├── chat/ # Chat Completions domain (WRAPPER) │ ├── router.py # POST /v1/chat/completions endpoint │ ├── schemas.py # Chat request/response models │ ├── service.py # Wraps Responses API, converts to tags │ ├── constants.py # Chat constants (roles, finish reasons) │ └── __init__.py ├── models/ # Models listing domain │ ├── router.py # GET /v1/models endpoint │ ├── schemas.py # Model schemas │ ├── service.py # Accesses ModelRegistry │ └── __init__.py ├── core/ # Shared utilities │ ├── config.py # Global configuration (BaseSettings) │ ├── models.py # Custom base Pydantic models │ ├── exceptions.py # Custom exceptions (RateLimitError, etc.) │ ├── dependencies.py # Shared dependencies │ └── router.py # Core routes (health, root) ├── ollama/ # Ollama client layer (not yet integrated) │ ├── client.py # Async Ollama HTTP client │ └── schemas.py # Ollama API models └── main.py # Application factory & configuration ``` **Key Architectural Principles**: - **Single Source of Truth**: Responses API handles all generation logic - **Wrapper Pattern**: Chat Completions wraps Responses API without duplicating code - **Agent Abstraction**: AgentInterface defines contract for all models - **Domain Separation**: Each domain has its own router, schemas, service - **Service Layer**: Business logic in services, not routers - **Type Safety**: Pydantic models for ALL request/response validation - **Async First**: All I/O operations use async/await ### Async/Await Best Practices **Critical Understanding**: FastAPI handles sync and async routes differently: - **Async routes** (`async def`): Called directly in event loop - Use ONLY for non-blocking operations - Perfect for `await httpx.get()`, database queries, file I/O - **NEVER** use blocking calls like `time.sleep()` - this blocks entire server - **Sync routes** (`def`): Run in thread pool - Use for CPU-intensive work or blocking SDKs - Blocking I/O won't freeze the event loop - Example: `time.sleep(10)` is safe here **Example**: ```python @router.get("/terrible") async def terrible(): time.sleep(10) # ❌ BLOCKS ENTIRE SERVER @router.get("/good") def good(): time.sleep(10) # ✅ Runs in thread pool @router.get("/perfect") async def perfect(): await asyncio.sleep(10) # ✅ Non-blocking async ``` **For CPU-intensive tasks**: Use separate worker processes (not threads) due to Python's GIL. ### Pydantic Configuration **Custom Base Model**: All schemas inherit from `CustomBaseModel` for consistent behavior: ```python # src/core/models.py class CustomBaseModel(BaseModel): model_config = ConfigDict( json_encoders={datetime: datetime_to_iso_str}, populate_by_name=True, use_enum_values=True, validate_assignment=True, ) def serializable_dict(self, **kwargs): """Return dict with only JSON-serializable fields.""" return jsonable_encoder(self.model_dump(**kwargs)) ``` **Benefits**: - Consistent datetime serialization across all responses - Alias support for field name flexibility - Easy JSON encoding for logging/debugging **Decoupled Settings**: Split configuration by domain instead of one monolithic file: ```python # src/core/config.py - Global settings class Config(BaseSettings): DATABASE_URL: PostgresDsn ENVIRONMENT: Environment # src/chat/config.py - Chat-specific settings class ChatConfig(BaseSettings): MAX_TOKENS: int DEFAULT_TEMPERATURE: float ``` ### Dependency Injection Patterns **Validation with Dependencies**: Use dependencies for complex validations: ```python async def valid_post_id(post_id: UUID4) -> dict: """Validate post exists in database.""" post = await service.get_by_id(post_id) if not post: raise PostNotFound() return post @router.get("/posts/{post_id}") async def get_post(post: dict = Depends(valid_post_id)): return post # Already validated! ``` **Chaining Dependencies**: Build reusable validation layers: ```python async def valid_owned_post( post: dict = Depends(valid_post_id), token_data: dict = Depends(parse_jwt_data), ) -> dict: if post["creator_id"] != token_data["user_id"]: raise UserNotOwner() return post ``` **Dependency Caching**: Dependencies are cached within request scope - FastAPI only executes each dependency once per request, even if used multiple times. ### Application Factory Pattern Main.py uses factory pattern for testability and configuration: ```python def create_application() -> FastAPI: """Create and configure FastAPI app.""" app = FastAPI(title=config.APP_NAME) # Add middleware app.add_middleware(CORSMiddleware, ...) # Register exception handlers register_exception_handlers(app) # Include routers app.include_router(chat_router, prefix="/v1") return app app = create_application() ``` ## Development Guidelines ### Code Structure (Current Implementation) - ✅ Use async/await for ALL I/O operations (database, HTTP, file access) - ✅ Use sync (def) for blocking SDKs or CPU-intensive work - ✅ Implement proper error handling and logging - ✅ Follow dependency injection for validation and shared resources - ✅ Use Pydantic models for ALL request/response validation - ✅ Keep business logic in service modules, not routers - ✅ Domain-based project structure (not file-type based) ### Security Considerations - ✅ Validate all inputs using Pydantic models - ✅ Use environment variables for sensitive configuration - ✅ Keep dependencies updated (all CVE-checked as of 2025-12-06) - ✅ Minor version locking for supply chain protection - 🚧 Implement rate limiting for API endpoints (future) - 🚧 Add authentication/API keys (future) ### Testing (Current Coverage: 62%) - ✅ Integration tests for API endpoints - ✅ Streaming functionality with 20s timeout protection - ✅ Async test support with pytest-asyncio - ✅ Validate OpenAI API compatibility - ✅ Mock responses for all endpoints - 🚧 Future: Mock Ollama responses when integrated ### Configuration - ✅ Use `.env` files for local development - ✅ Document all environment variables in README - ✅ Provide sensible defaults where possible - ✅ BaseSettings from pydantic-settings - 🚧 Support container-based configuration (future) ## Common Patterns ### Streaming Response Pattern (✅ Implemented) See `src/chat/router.py` for the current implementation: ```python from sse_starlette.sse import EventSourceResponse from fastapi import FastAPI async def event_generator(): # Currently yields mock lorem ipsum chunks # Future: Stream from Ollama/PydanticAI yield {"data": chunk.model_dump_json()} yield {"data": "[DONE]"} @app.post("/stream") async def stream(): return EventSourceResponse(event_generator()) ``` ### PydanticAI Agent Pattern (🚧 Future Reference) For future integration when connecting to Ollama: ```python from pydantic_ai import Agent agent = Agent( 'ollama:mistral-nemo', # Target model # Configuration here ) # Use the agent result = await agent.run('Your prompt') ``` ### OpenAI-Compatible Response Format (✅ Implemented) Current implementation in `src/chat/schemas.py`: ```python { "id": "chatcmpl-123", "object": "chat.completion.chunk", "created": 1234567890, "model": "mistral-nemo:latest", "choices": [{ "index": 0, "delta": {"content": "response"}, "finish_reason": None }] } ``` ## Update Policy This document should be updated when: - Package versions are upgraded - New major features are added - Breaking API changes occur - Security vulnerabilities are discovered Last updated: 2025-12-06