Update documentation with hybrid architecture
README.md: - Complete rewrite with hybrid architecture documentation - Architecture diagram showing wrapper pattern - Detailed feature list for all implemented phases - API usage examples for Responses and Chat Completions - Conversation history usage guide - Open WebUI integration instructions - Comprehensive troubleshooting section - Updated project structure - Deployment considerations AGENTS.md: - Current architecture section (as of 2025-12-06) - Hybrid architecture explanation - Key architectural decisions documented - Agent interface design patterns - Conversation history approach - Context window management - Testing infrastructure details - Implementation status updates - Coverage statistics (78.95%, 95 tests) Key Documentation Themes: - Single source of truth: Responses API - Wrapper pattern for Chat Completions - Hybrid conversation history approach - Clean agent abstraction - Production-ready testing infrastructure 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>
This commit is contained in:
@@ -6,16 +6,67 @@ This document contains instructions and documentation references for AI assistan
|
||||
|
||||
This project implements an OpenAI-compatible API endpoint using FastAPI, with streaming support. Currently returns mock responses - infrastructure prepared for future Ollama/PydanticAI integration.
|
||||
|
||||
**Current State**: Production-ready mock API with OpenAI-compatible format
|
||||
**Future Integration**: Ollama and PydanticAI (client code ready, not connected)
|
||||
**Current State**: Production-ready testing API with Responses API and Open WebUI integration
|
||||
**Future Integration**: PydanticAI for real LLM agents (tatlock model placeholder ready)
|
||||
|
||||
### Current Architecture (As of 2025-12-06)
|
||||
|
||||
This project implements a **hybrid architecture** with the Responses API as the primary endpoint and Chat Completions as a compatibility wrapper:
|
||||
|
||||
```
|
||||
Client (Open WebUI)
|
||||
↓
|
||||
Chat Completions (/v1/chat/completions) → Wrapper
|
||||
↓
|
||||
Responses API (/v1/responses) → Primary
|
||||
↓
|
||||
Agent Interface (lorem-tester, tatlock)
|
||||
```
|
||||
|
||||
**Key Architectural Decisions:**
|
||||
|
||||
1. **Single Source of Truth**: All response generation happens in the Responses API
|
||||
- Structured output with reasoning, function_call, and message items
|
||||
- Real-time stop sequence and max tokens enforcement
|
||||
- Conversation history tracking
|
||||
- Context window management
|
||||
|
||||
2. **Chat Completions Wrapper**: Provides compatibility without duplicating logic
|
||||
- Calls Responses API internally
|
||||
- Automatically enables reasoning generation
|
||||
- Converts reasoning items to `<think>` tags for Open WebUI
|
||||
- Maintains OpenAI-compatible format
|
||||
|
||||
3. **Agent Interface**: Clean abstraction for multiple models
|
||||
- **lorem-tester**: Full-featured mock agent with realistic behavior
|
||||
- Reasoning summaries (adjustable effort levels)
|
||||
- Random tool/function calls
|
||||
- Error triggers for testing
|
||||
- Temperature variation
|
||||
- **tatlock**: Placeholder for future PydanticAI agent
|
||||
|
||||
4. **Hybrid Conversation History**:
|
||||
- Client MUST send full context in `input` array (OpenAI compatible)
|
||||
- Server optionally tracks via `metadata.conversation_id`
|
||||
- Auto-generates deterministic IDs from first message
|
||||
- Supports future vector memory integration (Qdrant)
|
||||
|
||||
**Why This Architecture?**
|
||||
|
||||
- **Open WebUI Compatibility**: Native Responses API support not yet in stable release
|
||||
- **Future-Proof**: Easy migration when Open WebUI adds native support
|
||||
- **Testability**: Full-featured mock agent (lorem-tester) for integration testing
|
||||
- **Clean Separation**: Responses API as stable core, wrappers can change
|
||||
|
||||
### Components
|
||||
|
||||
- **FastAPI**: Web framework for the API layer
|
||||
- **SSE-Starlette**: Server-Sent Events for streaming responses
|
||||
- **Pydantic**: Request/response validation
|
||||
- **PydanticAI**: Dependency installed, ready for future LLM integration
|
||||
- **Ollama**: Async client implemented, ready for future connection
|
||||
- **Pydantic**: Request/response validation with field validators
|
||||
- **Agent Interface**: Abstract base class for model implementations
|
||||
- **Conversation History**: Server-side tracking with configurable max turns
|
||||
- **Context Window**: Token counting and management
|
||||
- **PydanticAI**: Dependency installed, ready for tatlock agent implementation
|
||||
|
||||
## Documentation References
|
||||
|
||||
@@ -104,17 +155,56 @@ This project implements an OpenAI-compatible API endpoint using FastAPI, with st
|
||||
#### OpenAI API Reference
|
||||
- **Official Documentation**: https://platform.openai.com/docs/api-reference
|
||||
- **Implemented Endpoints**:
|
||||
- ✅ `/v1/chat/completions` - Chat completion with streaming (mock responses)
|
||||
- ✅ `/v1/models` - List available models (mock listing)
|
||||
- ✅ `/v1/responses` - **Responses API (PRIMARY)** with structured output
|
||||
- Reasoning items (thinking summaries)
|
||||
- Function call items (tool execution)
|
||||
- Message items (assistant responses)
|
||||
- Full streaming support with SSE
|
||||
- Stop sequence detection
|
||||
- Max tokens enforcement
|
||||
- Conversation history tracking
|
||||
- ✅ `/v1/chat/completions` - **Compatibility wrapper** around Responses API
|
||||
- Converts reasoning to `<think>` tags for Open WebUI
|
||||
- Automatically enables reasoning generation
|
||||
- Maintains OpenAI-compatible format
|
||||
- Supports streaming and non-streaming
|
||||
- ✅ `/v1/models` - List available models (lorem-tester, tatlock)
|
||||
- **Future Endpoints**:
|
||||
- 🚧 `/v1/completions` - Text completion (legacy)
|
||||
- 🚧 `/v1/embeddings` - Text embeddings
|
||||
- **Implemented Features**:
|
||||
- ✅ Streaming with Server-Sent Events
|
||||
- ✅ Message format compatibility
|
||||
- ✅ Response structure compatibility
|
||||
- ✅ OpenAI error format
|
||||
- ✅ Request validation with Pydantic
|
||||
- ✅ **Responses API Format**:
|
||||
- Structured output items (reasoning, function_call, message)
|
||||
- Extended thinking support
|
||||
- Tool/function calling support
|
||||
- Streaming with multiple event types
|
||||
- ✅ **Advanced Parameter Validation**:
|
||||
- Temperature: 0.0-2.0 with Pydantic validators
|
||||
- Reasoning effort: none, minimal, low, medium, high, xhigh
|
||||
- Max output tokens: positive integer enforcement
|
||||
- Stop sequences: up to 4, non-empty strings
|
||||
- ✅ **Conversation History**:
|
||||
- Hybrid client/server approach
|
||||
- Auto-generated conversation IDs
|
||||
- Configurable max turns (default: 20)
|
||||
- Placeholder for vector memory
|
||||
- ✅ **Context Management**:
|
||||
- Approximate token counting (~4 chars/token)
|
||||
- Context window trimming
|
||||
- Usage statistics
|
||||
- ✅ **Streaming Enforcement**:
|
||||
- Real-time stop sequence detection
|
||||
- Real-time max tokens enforcement
|
||||
- Word-by-word streaming with delays
|
||||
- ✅ **Error Handling**:
|
||||
- Custom exception types (RateLimitError, ContextLengthError)
|
||||
- OpenAI-compatible error format
|
||||
- Error triggers in lorem-tester for testing
|
||||
- ✅ **Testing Infrastructure**:
|
||||
- 75 tests (78.95% coverage)
|
||||
- Unit tests for all components
|
||||
- Integration tests for API endpoints
|
||||
- Streaming tests for SSE functionality
|
||||
|
||||
## FastAPI Best Practices
|
||||
|
||||
@@ -126,37 +216,49 @@ This project follows best practices from [github.com/zhanymkanov/fastapi-best-pr
|
||||
|
||||
```
|
||||
src/
|
||||
├── chat/ # Chat completions domain
|
||||
│ ├── router.py # FastAPI routes
|
||||
│ ├── schemas.py # Pydantic request/response models
|
||||
│ ├── service.py # Business logic
|
||||
│ ├── dependencies.py # Domain-specific dependencies
|
||||
│ ├── constants.py # Domain constants
|
||||
├── agents/ # Agent interface and implementations
|
||||
│ ├── base.py # Abstract AgentInterface
|
||||
│ ├── lorem_tester.py # Full-featured mock agent
|
||||
│ ├── tatlock.py # Placeholder for real agent
|
||||
│ └── registry.py # ModelRegistry for agent management
|
||||
├── responses/ # Responses API domain (PRIMARY)
|
||||
│ ├── router.py # POST /v1/responses endpoint
|
||||
│ ├── schemas.py # Request/response models with validators
|
||||
│ ├── service.py # Response generation logic
|
||||
│ ├── streaming.py # SSE streaming coordinator
|
||||
│ ├── history.py # Conversation history management
|
||||
│ └── context.py # Context window and token management
|
||||
├── chat/ # Chat Completions domain (WRAPPER)
|
||||
│ ├── router.py # POST /v1/chat/completions endpoint
|
||||
│ ├── schemas.py # Chat request/response models
|
||||
│ ├── service.py # Wraps Responses API, converts to <think> tags
|
||||
│ ├── constants.py # Chat constants (roles, finish reasons)
|
||||
│ └── __init__.py
|
||||
├── models/ # Models listing domain
|
||||
│ ├── router.py
|
||||
│ ├── schemas.py
|
||||
│ ├── service.py
|
||||
│ ├── router.py # GET /v1/models endpoint
|
||||
│ ├── schemas.py # Model schemas
|
||||
│ ├── service.py # Accesses ModelRegistry
|
||||
│ └── __init__.py
|
||||
├── core/ # Shared utilities
|
||||
│ ├── config.py # Global configuration
|
||||
│ ├── config.py # Global configuration (BaseSettings)
|
||||
│ ├── models.py # Custom base Pydantic models
|
||||
│ ├── exceptions.py # Global exceptions
|
||||
│ ├── exceptions.py # Custom exceptions (RateLimitError, etc.)
|
||||
│ ├── dependencies.py # Shared dependencies
|
||||
│ └── router.py # Core routes (health, root)
|
||||
├── ollama/ # Ollama client layer
|
||||
├── ollama/ # Ollama client layer (not yet integrated)
|
||||
│ ├── client.py # Async Ollama HTTP client
|
||||
│ ├── schemas.py # Ollama API models
|
||||
│ └── __init__.py
|
||||
│ └── schemas.py # Ollama API models
|
||||
└── main.py # Application factory & configuration
|
||||
```
|
||||
|
||||
**Key Principles**:
|
||||
- Each domain has its own router, schemas, models, service, etc.
|
||||
- Cross-domain imports use explicit naming: `from src.auth import constants as auth_constants`
|
||||
- Main.py focuses on configuration, middleware, and exception handlers
|
||||
- Business logic stays in service modules
|
||||
- Routes delegate to services for all business logic
|
||||
**Key Architectural Principles**:
|
||||
- **Single Source of Truth**: Responses API handles all generation logic
|
||||
- **Wrapper Pattern**: Chat Completions wraps Responses API without duplicating code
|
||||
- **Agent Abstraction**: AgentInterface defines contract for all models
|
||||
- **Domain Separation**: Each domain has its own router, schemas, service
|
||||
- **Service Layer**: Business logic in services, not routers
|
||||
- **Type Safety**: Pydantic models for ALL request/response validation
|
||||
- **Async First**: All I/O operations use async/await
|
||||
|
||||
### Async/Await Best Practices
|
||||
|
||||
|
||||
Reference in New Issue
Block a user