README.md: - Complete rewrite with hybrid architecture documentation - Architecture diagram showing wrapper pattern - Detailed feature list for all implemented phases - API usage examples for Responses and Chat Completions - Conversation history usage guide - Open WebUI integration instructions - Comprehensive troubleshooting section - Updated project structure - Deployment considerations AGENTS.md: - Current architecture section (as of 2025-12-06) - Hybrid architecture explanation - Key architectural decisions documented - Agent interface design patterns - Conversation history approach - Context window management - Testing infrastructure details - Implementation status updates - Coverage statistics (78.95%, 95 tests) Key Documentation Themes: - Single source of truth: Responses API - Wrapper pattern for Chat Completions - Hybrid conversation history approach - Clean agent abstraction - Production-ready testing infrastructure 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>
16 KiB
LLM Agent Instructions
This document contains instructions and documentation references for AI assistants working with this codebase.
Project Overview
This project implements an OpenAI-compatible API endpoint using FastAPI, with streaming support. Currently returns mock responses - infrastructure prepared for future Ollama/PydanticAI integration.
Current State: Production-ready testing API with Responses API and Open WebUI integration Future Integration: PydanticAI for real LLM agents (tatlock model placeholder ready)
Current Architecture (As of 2025-12-06)
This project implements a hybrid architecture with the Responses API as the primary endpoint and Chat Completions as a compatibility wrapper:
Client (Open WebUI)
↓
Chat Completions (/v1/chat/completions) → Wrapper
↓
Responses API (/v1/responses) → Primary
↓
Agent Interface (lorem-tester, tatlock)
Key Architectural Decisions:
-
Single Source of Truth: All response generation happens in the Responses API
- Structured output with reasoning, function_call, and message items
- Real-time stop sequence and max tokens enforcement
- Conversation history tracking
- Context window management
-
Chat Completions Wrapper: Provides compatibility without duplicating logic
- Calls Responses API internally
- Automatically enables reasoning generation
- Converts reasoning items to
<think>tags for Open WebUI - Maintains OpenAI-compatible format
-
Agent Interface: Clean abstraction for multiple models
- lorem-tester: Full-featured mock agent with realistic behavior
- Reasoning summaries (adjustable effort levels)
- Random tool/function calls
- Error triggers for testing
- Temperature variation
- tatlock: Placeholder for future PydanticAI agent
- lorem-tester: Full-featured mock agent with realistic behavior
-
Hybrid Conversation History:
- Client MUST send full context in
inputarray (OpenAI compatible) - Server optionally tracks via
metadata.conversation_id - Auto-generates deterministic IDs from first message
- Supports future vector memory integration (Qdrant)
- Client MUST send full context in
Why This Architecture?
- Open WebUI Compatibility: Native Responses API support not yet in stable release
- Future-Proof: Easy migration when Open WebUI adds native support
- Testability: Full-featured mock agent (lorem-tester) for integration testing
- Clean Separation: Responses API as stable core, wrappers can change
Components
- FastAPI: Web framework for the API layer
- SSE-Starlette: Server-Sent Events for streaming responses
- Pydantic: Request/response validation with field validators
- Agent Interface: Abstract base class for model implementations
- Conversation History: Server-side tracking with configurable max turns
- Context Window: Token counting and management
- PydanticAI: Dependency installed, ready for tatlock agent implementation
Documentation References
Core Framework Documentation
FastAPI
- Official Documentation: https://fastapi.tiangolo.com/
- Version: 0.123.9 (Dec 2025)
- Key Topics:
- Path operations and routing
- Request/response models with Pydantic
- Dependency injection
- Background tasks
- WebSocket and streaming support
- PyPI: https://pypi.org/project/fastapi/
Uvicorn
- Official Documentation: https://www.uvicorn.org/
- Version: 0.38.0 (Oct 2025)
- Key Topics:
- ASGI server configuration
- Deployment settings
- Logging and monitoring
- SSL/TLS configuration
AI/LLM Integration
PydanticAI
- Official Documentation: https://ai.pydantic.dev/
- Version: 1.27.0 (Dec 2025)
- Status: Dependency installed, ready for future integration
- Key Topics (for future implementation):
- Agent creation and configuration
- LLM provider integration (Ollama support)
- Structured outputs with Pydantic
- Streaming responses
- Tool/function calling
- RunContext and dynamic configuration
- MCP server integration
- GitHub: https://github.com/pydantic/pydantic-ai
- PyPI: https://pypi.org/project/pydantic-ai/
Pydantic
- Official Documentation: https://docs.pydantic.dev/latest/
- Version: 2.11+ (Required for PydanticAI, currently using >=2.11,<2.13)
- Key Topics:
- Data validation and serialization
- Field types and validators
- Model configuration
- JSON schema generation
HTTP and Streaming
HTTPX
- Official Documentation: https://www.python-httpx.org/
- Version: 0.28.1
- Key Topics:
- Async HTTP client for Ollama communication
- Streaming responses
- Timeout configuration
- Connection pooling
SSE-Starlette
- GitHub: https://github.com/sysid/sse-starlette
- Version: 3.0.2 (Oct 2025)
- Key Topics:
- Server-Sent Events implementation
- Streaming event responses
- Integration with FastAPI/Starlette
Ollama Integration
Ollama API
- Official Documentation: https://github.com/ollama/ollama/blob/main/docs/api.md
- Status: Async client implemented in
src/ollama/client.py, ready for future integration - Key Topics (for future implementation):
- REST API endpoints
- Streaming responses
- Model management
- Generate and chat endpoints
- Model configuration
- Current Model Target: mistral-nemo:latest
OpenAI API Compatibility
OpenAI API Reference
- Official Documentation: https://platform.openai.com/docs/api-reference
- Implemented Endpoints:
- ✅
/v1/responses- Responses API (PRIMARY) with structured output- Reasoning items (thinking summaries)
- Function call items (tool execution)
- Message items (assistant responses)
- Full streaming support with SSE
- Stop sequence detection
- Max tokens enforcement
- Conversation history tracking
- ✅
/v1/chat/completions- Compatibility wrapper around Responses API- Converts reasoning to
<think>tags for Open WebUI - Automatically enables reasoning generation
- Maintains OpenAI-compatible format
- Supports streaming and non-streaming
- Converts reasoning to
- ✅
/v1/models- List available models (lorem-tester, tatlock)
- ✅
- Future Endpoints:
- 🚧
/v1/completions- Text completion (legacy) - 🚧
/v1/embeddings- Text embeddings
- 🚧
- Implemented Features:
- ✅ Responses API Format:
- Structured output items (reasoning, function_call, message)
- Extended thinking support
- Tool/function calling support
- Streaming with multiple event types
- ✅ Advanced Parameter Validation:
- Temperature: 0.0-2.0 with Pydantic validators
- Reasoning effort: none, minimal, low, medium, high, xhigh
- Max output tokens: positive integer enforcement
- Stop sequences: up to 4, non-empty strings
- ✅ Conversation History:
- Hybrid client/server approach
- Auto-generated conversation IDs
- Configurable max turns (default: 20)
- Placeholder for vector memory
- ✅ Context Management:
- Approximate token counting (~4 chars/token)
- Context window trimming
- Usage statistics
- ✅ Streaming Enforcement:
- Real-time stop sequence detection
- Real-time max tokens enforcement
- Word-by-word streaming with delays
- ✅ Error Handling:
- Custom exception types (RateLimitError, ContextLengthError)
- OpenAI-compatible error format
- Error triggers in lorem-tester for testing
- ✅ Testing Infrastructure:
- 75 tests (78.95% coverage)
- Unit tests for all components
- Integration tests for API endpoints
- Streaming tests for SSE functionality
- ✅ Responses API Format:
FastAPI Best Practices
This project follows best practices from github.com/zhanymkanov/fastapi-best-practices
Project Structure
Domain-Based Organization: Code is organized by domain/feature rather than by file type:
src/
├── agents/ # Agent interface and implementations
│ ├── base.py # Abstract AgentInterface
│ ├── lorem_tester.py # Full-featured mock agent
│ ├── tatlock.py # Placeholder for real agent
│ └── registry.py # ModelRegistry for agent management
├── responses/ # Responses API domain (PRIMARY)
│ ├── router.py # POST /v1/responses endpoint
│ ├── schemas.py # Request/response models with validators
│ ├── service.py # Response generation logic
│ ├── streaming.py # SSE streaming coordinator
│ ├── history.py # Conversation history management
│ └── context.py # Context window and token management
├── chat/ # Chat Completions domain (WRAPPER)
│ ├── router.py # POST /v1/chat/completions endpoint
│ ├── schemas.py # Chat request/response models
│ ├── service.py # Wraps Responses API, converts to <think> tags
│ ├── constants.py # Chat constants (roles, finish reasons)
│ └── __init__.py
├── models/ # Models listing domain
│ ├── router.py # GET /v1/models endpoint
│ ├── schemas.py # Model schemas
│ ├── service.py # Accesses ModelRegistry
│ └── __init__.py
├── core/ # Shared utilities
│ ├── config.py # Global configuration (BaseSettings)
│ ├── models.py # Custom base Pydantic models
│ ├── exceptions.py # Custom exceptions (RateLimitError, etc.)
│ ├── dependencies.py # Shared dependencies
│ └── router.py # Core routes (health, root)
├── ollama/ # Ollama client layer (not yet integrated)
│ ├── client.py # Async Ollama HTTP client
│ └── schemas.py # Ollama API models
└── main.py # Application factory & configuration
Key Architectural Principles:
- Single Source of Truth: Responses API handles all generation logic
- Wrapper Pattern: Chat Completions wraps Responses API without duplicating code
- Agent Abstraction: AgentInterface defines contract for all models
- Domain Separation: Each domain has its own router, schemas, service
- Service Layer: Business logic in services, not routers
- Type Safety: Pydantic models for ALL request/response validation
- Async First: All I/O operations use async/await
Async/Await Best Practices
Critical Understanding: FastAPI handles sync and async routes differently:
-
Async routes (
async def): Called directly in event loop- Use ONLY for non-blocking operations
- Perfect for
await httpx.get(), database queries, file I/O - NEVER use blocking calls like
time.sleep()- this blocks entire server
-
Sync routes (
def): Run in thread pool- Use for CPU-intensive work or blocking SDKs
- Blocking I/O won't freeze the event loop
- Example:
time.sleep(10)is safe here
Example:
@router.get("/terrible")
async def terrible():
time.sleep(10) # ❌ BLOCKS ENTIRE SERVER
@router.get("/good")
def good():
time.sleep(10) # ✅ Runs in thread pool
@router.get("/perfect")
async def perfect():
await asyncio.sleep(10) # ✅ Non-blocking async
For CPU-intensive tasks: Use separate worker processes (not threads) due to Python's GIL.
Pydantic Configuration
Custom Base Model: All schemas inherit from CustomBaseModel for consistent behavior:
# src/core/models.py
class CustomBaseModel(BaseModel):
model_config = ConfigDict(
json_encoders={datetime: datetime_to_iso_str},
populate_by_name=True,
use_enum_values=True,
validate_assignment=True,
)
def serializable_dict(self, **kwargs):
"""Return dict with only JSON-serializable fields."""
return jsonable_encoder(self.model_dump(**kwargs))
Benefits:
- Consistent datetime serialization across all responses
- Alias support for field name flexibility
- Easy JSON encoding for logging/debugging
Decoupled Settings: Split configuration by domain instead of one monolithic file:
# src/core/config.py - Global settings
class Config(BaseSettings):
DATABASE_URL: PostgresDsn
ENVIRONMENT: Environment
# src/chat/config.py - Chat-specific settings
class ChatConfig(BaseSettings):
MAX_TOKENS: int
DEFAULT_TEMPERATURE: float
Dependency Injection Patterns
Validation with Dependencies: Use dependencies for complex validations:
async def valid_post_id(post_id: UUID4) -> dict:
"""Validate post exists in database."""
post = await service.get_by_id(post_id)
if not post:
raise PostNotFound()
return post
@router.get("/posts/{post_id}")
async def get_post(post: dict = Depends(valid_post_id)):
return post # Already validated!
Chaining Dependencies: Build reusable validation layers:
async def valid_owned_post(
post: dict = Depends(valid_post_id),
token_data: dict = Depends(parse_jwt_data),
) -> dict:
if post["creator_id"] != token_data["user_id"]:
raise UserNotOwner()
return post
Dependency Caching: Dependencies are cached within request scope - FastAPI only executes each dependency once per request, even if used multiple times.
Application Factory Pattern
Main.py uses factory pattern for testability and configuration:
def create_application() -> FastAPI:
"""Create and configure FastAPI app."""
app = FastAPI(title=config.APP_NAME)
# Add middleware
app.add_middleware(CORSMiddleware, ...)
# Register exception handlers
register_exception_handlers(app)
# Include routers
app.include_router(chat_router, prefix="/v1")
return app
app = create_application()
Development Guidelines
Code Structure (Current Implementation)
- ✅ Use async/await for ALL I/O operations (database, HTTP, file access)
- ✅ Use sync (def) for blocking SDKs or CPU-intensive work
- ✅ Implement proper error handling and logging
- ✅ Follow dependency injection for validation and shared resources
- ✅ Use Pydantic models for ALL request/response validation
- ✅ Keep business logic in service modules, not routers
- ✅ Domain-based project structure (not file-type based)
Security Considerations
- ✅ Validate all inputs using Pydantic models
- ✅ Use environment variables for sensitive configuration
- ✅ Keep dependencies updated (all CVE-checked as of 2025-12-06)
- ✅ Minor version locking for supply chain protection
- 🚧 Implement rate limiting for API endpoints (future)
- 🚧 Add authentication/API keys (future)
Testing (Current Coverage: 62%)
- ✅ Integration tests for API endpoints
- ✅ Streaming functionality with 20s timeout protection
- ✅ Async test support with pytest-asyncio
- ✅ Validate OpenAI API compatibility
- ✅ Mock responses for all endpoints
- 🚧 Future: Mock Ollama responses when integrated
Configuration
- ✅ Use
.envfiles for local development - ✅ Document all environment variables in README
- ✅ Provide sensible defaults where possible
- ✅ BaseSettings from pydantic-settings
- 🚧 Support container-based configuration (future)
Common Patterns
Streaming Response Pattern (✅ Implemented)
See src/chat/router.py for the current implementation:
from sse_starlette.sse import EventSourceResponse
from fastapi import FastAPI
async def event_generator():
# Currently yields mock lorem ipsum chunks
# Future: Stream from Ollama/PydanticAI
yield {"data": chunk.model_dump_json()}
yield {"data": "[DONE]"}
@app.post("/stream")
async def stream():
return EventSourceResponse(event_generator())
PydanticAI Agent Pattern (🚧 Future Reference)
For future integration when connecting to Ollama:
from pydantic_ai import Agent
agent = Agent(
'ollama:mistral-nemo', # Target model
# Configuration here
)
# Use the agent
result = await agent.run('Your prompt')
OpenAI-Compatible Response Format (✅ Implemented)
Current implementation in src/chat/schemas.py:
{
"id": "chatcmpl-123",
"object": "chat.completion.chunk",
"created": 1234567890,
"model": "mistral-nemo:latest",
"choices": [{
"index": 0,
"delta": {"content": "response"},
"finish_reason": None
}]
}
Update Policy
This document should be updated when:
- Package versions are upgraded
- New major features are added
- Breaking API changes occur
- Security vulnerabilities are discovered
Last updated: 2025-12-06