# LLM Agent Instructions This document contains instructions and documentation references for AI assistants working with this codebase. ## Project Overview This project implements an OpenAI-compatible API endpoint using FastAPI, with streaming support for LLM responses. The architecture consists of: - **FastAPI**: Web framework for the API layer - **PydanticAI**: Agent framework for LLM integration - **Ollama**: LLM backend running on a networked container - **Stream Coordinator**: Manages streaming responses in OpenAI-compatible format ## Documentation References ### Core Framework Documentation #### FastAPI - **Official Documentation**: https://fastapi.tiangolo.com/ - **Version**: 0.123.9 (Dec 2025) - **Key Topics**: - Path operations and routing - Request/response models with Pydantic - Dependency injection - Background tasks - WebSocket and streaming support - **PyPI**: https://pypi.org/project/fastapi/ #### Uvicorn - **Official Documentation**: https://www.uvicorn.org/ - **Version**: 0.38.0 (Oct 2025) - **Key Topics**: - ASGI server configuration - Deployment settings - Logging and monitoring - SSL/TLS configuration ### AI/LLM Integration #### PydanticAI - **Official Documentation**: https://ai.pydantic.dev/ - **Version**: 1.27.0 (Dec 2025) - **Key Topics**: - Agent creation and configuration - LLM provider integration (Ollama support) - Structured outputs with Pydantic - Streaming responses - Tool/function calling - RunContext and dynamic configuration - MCP server integration - **GitHub**: https://github.com/pydantic/pydantic-ai - **PyPI**: https://pypi.org/project/pydantic-ai/ #### Pydantic - **Official Documentation**: https://docs.pydantic.dev/latest/ - **Version**: 2.10+ (Required for PydanticAI) - **Key Topics**: - Data validation and serialization - Field types and validators - Model configuration - JSON schema generation ### HTTP and Streaming #### HTTPX - **Official Documentation**: https://www.python-httpx.org/ - **Version**: 0.28.1 - **Key Topics**: - Async HTTP client for Ollama communication - Streaming responses - Timeout configuration - Connection pooling #### SSE-Starlette - **GitHub**: https://github.com/sysid/sse-starlette - **Version**: 3.0.2 (Oct 2025) - **Key Topics**: - Server-Sent Events implementation - Streaming event responses - Integration with FastAPI/Starlette ### Ollama Integration #### Ollama API - **Official Documentation**: https://github.com/ollama/ollama/blob/main/docs/api.md - **Key Topics**: - REST API endpoints - Streaming responses - Model management - Generate and chat endpoints - Model configuration ### OpenAI API Compatibility #### OpenAI API Reference - **Official Documentation**: https://platform.openai.com/docs/api-reference - **Key Endpoints to Implement**: - `/v1/chat/completions` - Chat completion with streaming - `/v1/models` - List available models - `/v1/completions` - Text completion (legacy) - **Key Features**: - Streaming with Server-Sent Events - Message format compatibility - Response structure compatibility ## FastAPI Best Practices This project follows best practices from [github.com/zhanymkanov/fastapi-best-practices](https://github.com/zhanymkanov/fastapi-best-practices) ### Project Structure **Domain-Based Organization**: Code is organized by domain/feature rather than by file type: ``` src/ ├── chat/ # Chat completions domain │ ├── router.py # FastAPI routes │ ├── schemas.py # Pydantic request/response models │ ├── service.py # Business logic │ ├── dependencies.py # Domain-specific dependencies │ ├── constants.py # Domain constants │ └── __init__.py ├── models/ # Models listing domain │ ├── router.py │ ├── schemas.py │ ├── service.py │ └── __init__.py ├── core/ # Shared utilities │ ├── config.py # Global configuration │ ├── models.py # Custom base Pydantic models │ ├── exceptions.py # Global exceptions │ ├── dependencies.py # Shared dependencies │ └── router.py # Core routes (health, root) ├── ollama/ # Ollama client layer │ ├── client.py # Async Ollama HTTP client │ ├── schemas.py # Ollama API models │ └── __init__.py └── main.py # Application factory & configuration ``` **Key Principles**: - Each domain has its own router, schemas, models, service, etc. - Cross-domain imports use explicit naming: `from src.auth import constants as auth_constants` - Main.py focuses on configuration, middleware, and exception handlers - Business logic stays in service modules - Routes delegate to services for all business logic ### Async/Await Best Practices **Critical Understanding**: FastAPI handles sync and async routes differently: - **Async routes** (`async def`): Called directly in event loop - Use ONLY for non-blocking operations - Perfect for `await httpx.get()`, database queries, file I/O - **NEVER** use blocking calls like `time.sleep()` - this blocks entire server - **Sync routes** (`def`): Run in thread pool - Use for CPU-intensive work or blocking SDKs - Blocking I/O won't freeze the event loop - Example: `time.sleep(10)` is safe here **Example**: ```python @router.get("/terrible") async def terrible(): time.sleep(10) # ❌ BLOCKS ENTIRE SERVER @router.get("/good") def good(): time.sleep(10) # ✅ Runs in thread pool @router.get("/perfect") async def perfect(): await asyncio.sleep(10) # ✅ Non-blocking async ``` **For CPU-intensive tasks**: Use separate worker processes (not threads) due to Python's GIL. ### Pydantic Configuration **Custom Base Model**: All schemas inherit from `CustomBaseModel` for consistent behavior: ```python # src/core/models.py class CustomBaseModel(BaseModel): model_config = ConfigDict( json_encoders={datetime: datetime_to_iso_str}, populate_by_name=True, use_enum_values=True, validate_assignment=True, ) def serializable_dict(self, **kwargs): """Return dict with only JSON-serializable fields.""" return jsonable_encoder(self.model_dump(**kwargs)) ``` **Benefits**: - Consistent datetime serialization across all responses - Alias support for field name flexibility - Easy JSON encoding for logging/debugging **Decoupled Settings**: Split configuration by domain instead of one monolithic file: ```python # src/core/config.py - Global settings class Config(BaseSettings): DATABASE_URL: PostgresDsn ENVIRONMENT: Environment # src/chat/config.py - Chat-specific settings class ChatConfig(BaseSettings): MAX_TOKENS: int DEFAULT_TEMPERATURE: float ``` ### Dependency Injection Patterns **Validation with Dependencies**: Use dependencies for complex validations: ```python async def valid_post_id(post_id: UUID4) -> dict: """Validate post exists in database.""" post = await service.get_by_id(post_id) if not post: raise PostNotFound() return post @router.get("/posts/{post_id}") async def get_post(post: dict = Depends(valid_post_id)): return post # Already validated! ``` **Chaining Dependencies**: Build reusable validation layers: ```python async def valid_owned_post( post: dict = Depends(valid_post_id), token_data: dict = Depends(parse_jwt_data), ) -> dict: if post["creator_id"] != token_data["user_id"]: raise UserNotOwner() return post ``` **Dependency Caching**: Dependencies are cached within request scope - FastAPI only executes each dependency once per request, even if used multiple times. ### Application Factory Pattern Main.py uses factory pattern for testability and configuration: ```python def create_application() -> FastAPI: """Create and configure FastAPI app.""" app = FastAPI(title=config.APP_NAME) # Add middleware app.add_middleware(CORSMiddleware, ...) # Register exception handlers register_exception_handlers(app) # Include routers app.include_router(chat_router, prefix="/v1") return app app = create_application() ``` ## Development Guidelines ### Code Structure - Use async/await for ALL I/O operations (database, HTTP, file access) - Use sync (def) for blocking SDKs or CPU-intensive work - Implement proper error handling and logging - Follow dependency injection for validation and shared resources - Use Pydantic models for ALL request/response validation - Keep business logic in service modules, not routers ### Security Considerations - Validate all inputs using Pydantic models - Implement rate limiting for API endpoints - Use environment variables for sensitive configuration - Keep dependencies updated (check for CVEs regularly) ### Testing - Write integration tests for API endpoints - Test streaming functionality thoroughly - Mock Ollama responses for unit tests - Validate OpenAI API compatibility ### Configuration - Use `.env` files for local development - Document all environment variables in README - Provide sensible defaults where possible - Support container-based configuration ## Common Patterns ### Streaming Response Pattern ```python from sse_starlette.sse import EventSourceResponse from fastapi import FastAPI async def event_generator(): # Stream events from Ollama/PydanticAI yield {"data": "chunk1"} yield {"data": "chunk2"} @app.get("/stream") async def stream(): return EventSourceResponse(event_generator()) ``` ### PydanticAI Agent Pattern ```python from pydantic_ai import Agent agent = Agent( 'ollama:llama2', # Or other Ollama model # Configuration here ) # Use the agent result = await agent.run('Your prompt') ``` ### OpenAI-Compatible Response Format ```python { "id": "chatcmpl-123", "object": "chat.completion.chunk", "created": 1234567890, "model": "model-name", "choices": [{ "index": 0, "delta": {"content": "response"}, "finish_reason": None }] } ``` ## Update Policy This document should be updated when: - Package versions are upgraded - New major features are added - Breaking API changes occur - Security vulnerabilities are discovered Last updated: 2025-12-05