Update documentation with hybrid architecture

README.md:
- Complete rewrite with hybrid architecture documentation
- Architecture diagram showing wrapper pattern
- Detailed feature list for all implemented phases
- API usage examples for Responses and Chat Completions
- Conversation history usage guide
- Open WebUI integration instructions
- Comprehensive troubleshooting section
- Updated project structure
- Deployment considerations

AGENTS.md:
- Current architecture section (as of 2025-12-06)
- Hybrid architecture explanation
- Key architectural decisions documented
- Agent interface design patterns
- Conversation history approach
- Context window management
- Testing infrastructure details
- Implementation status updates
- Coverage statistics (78.95%, 95 tests)

Key Documentation Themes:
- Single source of truth: Responses API
- Wrapper pattern for Chat Completions
- Hybrid conversation history approach
- Clean agent abstraction
- Production-ready testing infrastructure

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
This commit is contained in:
2025-12-06 19:39:32 +01:00
co-authored by Claude
parent 8b6cff2920
commit 661db0a672
2 changed files with 490 additions and 194 deletions
+356 -162
View File
@@ -1,36 +1,114 @@
# OpenAI-Compatible API
# Tatlock - OpenAI-Compatible API with Responses API
A FastAPI-based service that provides OpenAI-compatible API endpoints with streaming support. Currently returns mock responses - ready for future Ollama/PydanticAI integration.
A FastAPI-based service providing OpenAI-compatible API endpoints with full Responses API support, reasoning display, and streaming. Features a hybrid architecture with chat completions as a compatibility wrapper around the Responses API.
## Current Status
**✅ Production-ready mock API** with OpenAI-compatible format
**🚧 Ollama/PydanticAI integration** prepared but not connected
**✅ Production-ready testing API** with OpenAI Responses API format
**✅ Open WebUI integration** with reasoning bubbles (`<think>` tags)
**✅ Conversation history** with hybrid client/server approach
**🚧 PydanticAI integration** prepared for future real LLM connection
## Components
## Architecture Overview
- **FastAPI**: High-performance web framework providing the API layer
- **SSE-Starlette**: Server-Sent Events for streaming responses
- **Pydantic**: Type-safe request/response handling and validation
- **Ollama Client**: Async HTTP client prepared for future integration
- **PydanticAI**: Ready for LLM integration (not yet connected)
### Hybrid API Design
```
┌─────────────────────────────────────────┐
│ Client (Open WebUI, etc.) │
└────────┬────────────────────────────────┘
│
├──────────────────────────────────┐
│ │
v v
┌────────────────────┐ ┌──────────────────────┐
│ /v1/chat/ │ wrapper │ /v1/responses │
│ completions ├─────────>│ (Primary API) │
│ │ │ │
│ • OpenAI compat │ │ • Reasoning items │
│ • <think> tags │ │ • Function calls │
│ • Legacy support │ │ • Message items │
└────────────────────┘ └──────────┬───────────┘
│
v
┌──────────────────────┐
│ Agent Interface │
│ │
│ • lorem-tester │
│ • tatlock (future) │
└──────────────────────┘
```
**Key Architectural Decisions:**
- **Single Source of Truth**: Responses API handles all generation logic
- **Chat Completions Wrapper**: Converts Responses output to Chat format with `<think>` tags
- **Agent Interface**: Clean abstraction for multiple models (mock and real)
- **Hybrid History**: Client sends full context, server optionally tracks conversations
## Features
- ✅ OpenAI-compatible API endpoints (`/v1/chat/completions`, `/v1/models`)
- ✅ Streaming responses with Server-Sent Events (SSE)
- ✅ Type-safe request/response handling with Pydantic
- ✅ Async/await throughout for optimal performance
- ✅ Comprehensive test suite (62% coverage)
- ✅ Domain-based architecture following FastAPI best practices
- 🚧 Ollama integration (client ready, not connected)
- 🚧 PydanticAI integration (dependency installed, not connected)
### Core API
- ✅ **Responses API** (`/v1/responses`) - Primary endpoint with structured output
- Reasoning items (thinking/extended thinking)
- Function call items (tool execution)
- Message items (assistant responses)
- Streaming and non-streaming modes
- ✅ **Chat Completions API** (`/v1/chat/completions`) - Compatibility wrapper
- Converts reasoning to `<think>` tags for Open WebUI
- Maintains OpenAI-compatible format
- Wraps Responses API (single source of truth)
- ✅ **Models API** (`/v1/models`) - Lists available models
### Advanced Features
- ✅ **Conversation History Management**
- Hybrid approach: client maintains state, server tracks optionally
- Auto-generated conversation IDs from first message hash
- Configurable max turns (default: 20)
- Placeholder for future vector memory (Qdrant)
- ✅ **Context Window Management**
- Approximate token counting (~4 chars/token)
- Context trimming to fit model limits
- Token usage statistics
- ✅ **Parameter Validation**
- Temperature: 0.0-2.0
- Reasoning effort: none, minimal, low, medium, high, xhigh
- Max output tokens enforcement
- Stop sequences (up to 4)
- ✅ **Stop Sequence Detection**
- Real-time detection during streaming
- Stops generation immediately when encountered
- ✅ **Max Tokens Enforcement**
- Real-time token counting during streaming
- Stops when limit reached
### Testing Models
- ✅ **lorem-tester** - Full-featured mock agent
- Realistic reasoning summaries
- Random tool/function call generation
- Error triggers for testing (rate_limit, context_overflow)
- Temperature variation
- ✅ **tatlock** - Placeholder for real PydanticAI agent
### Open WebUI Integration
- ✅ **Reasoning Display** - Thinking bubbles shown separately from responses
- ✅ **Streaming Support** - Smooth word-by-word streaming
- ✅ **Error Handling** - Graceful error display
- ✅ **Model Selection** - Both models available in dropdown
## Components
- **FastAPI**: High-performance web framework
- **SSE-Starlette**: Server-Sent Events for streaming
- **Pydantic**: Type-safe request/response validation
- **Agent Interface**: Abstraction for multiple model backends
- **Conversation History**: Server-side tracking with hybrid approach
- **Context Window**: Token management and trimming
## Requirements
- Python 3.12+ (Python 3.12.11 recommended for security)
- No external dependencies required for mock API
- (Future: Network access to Ollama instance for LLM integration)
- Python 3.12+ (Python 3.12.11 recommended)
- No external dependencies for mock API
- (Future: Network access for PydanticAI integration)
## Installation
@@ -44,8 +122,8 @@ cd tatlock
### 2. Create a virtual environment
```bash
python -m venv venv
source venv/bin/activate # On Windows: venv\Scripts\activate
python -m venv .venv
source .venv/bin/activate # On Windows: .venv\Scripts\activate
```
### 3. Install dependencies
@@ -54,9 +132,9 @@ source venv/bin/activate # On Windows: venv\Scripts\activate
pip install -r requirements.txt
```
### 4. Configure environment variables (Optional)
### 4. Configure environment (Optional)
Create a `.env` file in the project root for custom configuration:
Create a `.env` file for custom configuration:
```env
# API Configuration
@@ -66,231 +144,348 @@ API_PORT=8000
# Logging
LOG_LEVEL=INFO
# Future Ollama Configuration (not yet integrated)
# OLLAMA_HOST=http://localhost:11434
# OLLAMA_DEFAULT_MODEL=mistral-nemo:latest
# OLLAMA_TIMEOUT=120
# Future: Add real LLM configuration here
```
**Note**: The API works with defaults. Environment variables are optional for customization. Ollama configuration is prepared but not currently used.
## Usage
### Start the development server
### Start the server
```bash
uvicorn src.main:app --reload
```
The API will be available at `http://localhost:8000`
API available at `http://localhost:8000`
### API Endpoints
#### Chat Completions (OpenAI-compatible)
#### Responses API (Primary)
Returns mock lorem ipsum responses:
OpenAI Responses API format with structured output:
```bash
curl http://localhost:8000/v1/responses \
-H "Content-Type: application/json" \
-d '{
"model": "lorem-tester",
"input": [
{"role": "user", "content": "Explain quantum computing"}
],
"reasoning": {
"effort": "medium",
"summary": "auto"
},
"max_output_tokens": 500,
"stop": ["END"],
"stream": false
}'
```
**Response Structure:**
```json
{
"id": "resp_abc123",
"object": "response",
"created_at": 1733529600,
"model": "lorem-tester",
"status": "completed",
"output": [
{
"type": "reasoning",
"id": "reasoning_xyz",
"summary": [
"Analyzing the user's request...",
"Considering quantum mechanics principles..."
]
},
{
"type": "message",
"id": "msg_def456",
"role": "assistant",
"content": [
{
"type": "output_text",
"text": "Quantum computing uses quantum mechanics..."
}
]
}
],
"usage": {
"input_tokens": 10,
"output_tokens": 50,
"reasoning_tokens": 20,
"total_tokens": 80
}
}
```
#### Chat Completions (Compatibility)
OpenAI-compatible format with `<think>` tags:
```bash
curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "mistral-nemo:latest",
"model": "lorem-tester",
"messages": [
{"role": "user", "content": "Hello, how are you?"}
{"role": "user", "content": "Hello!"}
],
"temperature": 0.7,
"stream": true
}'
```
#### List Models
**Note**: Chat Completions automatically enables reasoning and converts it to `<think>` tags for Open WebUI compatibility.
Returns mock model listing:
#### List Models
```bash
curl http://localhost:8000/v1/models
```
Response: `{"object": "list", "data": [{"id": "mistral-nemo:latest", ...}]}`
Returns:
```json
{
"object": "list",
"data": [
{
"id": "lorem-tester",
"object": "model",
"created": 1733529600,
"owned_by": "tatlock"
},
{
"id": "tatlock",
"object": "model",
"created": 1733529600,
"owned_by": "tatlock"
}
]
}
```
### Interactive API Documentation
### Conversation History
- Swagger UI: `http://localhost:8000/docs`
- ReDoc: `http://localhost:8000/redoc`
Optional conversation tracking via metadata:
## Security
```bash
curl http://localhost:8000/v1/responses \
-H "Content-Type: application/json" \
-d '{
"model": "lorem-tester",
"input": [
{"role": "user", "content": "Hello"}
],
"metadata": {
"conversation_id": "conv_abc123"
}
}'
```
### Version Locking Strategy
**Hybrid Approach:**
- Client MUST send full conversation history in `input` array (OpenAI compatible)
- Server optionally tracks via `metadata.conversation_id` (for analytics, future vector memory)
- Auto-generates conversation ID from first message hash if not provided
This project uses minor version locking (`>=X.Y,<X.(Y+1)`) to protect against supply chain attacks while allowing patch updates. All dependencies have been:
### Interactive Documentation
- Checked for known CVEs (as of 2025-12-06)
- Pinned to secure minor versions
- Documented with version rationale in `requirements.txt`
- **Swagger UI**: `http://localhost:8000/docs`
- **ReDoc**: `http://localhost:8000/redoc`
### CVE Status (2025-12-06)
## Open WebUI Integration
- **FastAPI 0.123.9**: No known vulnerabilities
- **Uvicorn 0.38.0**: No known vulnerabilities
- **PydanticAI 1.27.0**: No known vulnerabilities
- **HTTPX 0.28.1**: No known vulnerabilities
- **SSE-Starlette 3.0.2**: No known vulnerabilities
### Docker Networking
Regular security updates are recommended. Check for new versions monthly.
If running Open WebUI in Docker and API on host:
### Security Best Practices
```bash
# Use Docker bridge gateway IP
http://172.17.0.1:8000/v1/chat/completions
```
1. Never commit `.env` files
2. Use environment variables for sensitive configuration
3. Keep dependencies updated
4. Implement rate limiting in production
5. Use HTTPS in production environments
6. Validate all inputs with Pydantic models
### Reasoning Display
The Chat Completions wrapper automatically:
1. Enables reasoning generation
2. Converts reasoning items to `<think>` tags
3. Streams thinking before the actual response
Open WebUI displays this as:
- **Thought bubble** showing reasoning steps
- **Main response** showing the actual answer
### Testing Error Handling
Lorem-tester supports error triggers:
- **"trigger_rate_limit"** - Simulates rate limit error
- **"trigger_context_overflow"** - Simulates context length error
## Development
### Project Structure
Following [FastAPI best practices](https://github.com/zhanymkanov/fastapi-best-practices) with domain-based organization:
Following FastAPI best practices with domain-based organization:
```
tatlock/
├── src/
│ ├── chat/ # Chat completions domain
│ │ ├── router.py # OpenAI-compatible /v1/chat/completions
│ ├── agents/ # Agent interface and implementations
│ │ ├── base.py # Abstract AgentInterface
│ │ ├── lorem_tester.py # Full-featured mock agent
│ │ ├── tatlock.py # Placeholder for real agent
│ │ └── registry.py # Model registry
│ ├── responses/ # Responses API domain (PRIMARY)
│ │ ├── router.py # POST /v1/responses
│ │ ├── schemas.py # Request/response models
│ │ ├── service.py # Business logic (currently mock)
│ │ ├── dependencies.py # Route dependencies
│ │ └── constants.py # Domain constants
│ │ ├── service.py # Response generation logic
│ │ ├── streaming.py # SSE streaming coordinator
│ │ ├── history.py # Conversation history management
│ │ └── context.py # Context window management
│ ├── chat/ # Chat Completions domain (WRAPPER)
│ │ ├── router.py # POST /v1/chat/completions
│ │ ├── schemas.py # Chat request/response models
│ │ ├── service.py # Wraps Responses API
│ │ └── constants.py # Chat constants
│ ├── models/ # Models listing domain
│ │ ├── router.py # OpenAI-compatible /v1/models
│ │ ├── router.py # GET /v1/models
│ │ ├── schemas.py # Model schemas
│ │ └── service.py # Model list service (currently mock)
│ │ └── service.py # Model registry access
│ ├── core/ # Shared utilities
│ │ ├── config.py # Global configuration (BaseSettings)
│ │ ├── models.py # Custom Pydantic base models
│ │ ├── config.py # Configuration (BaseSettings)
│ │ ├── models.py # Custom Pydantic base
│ │ ├── exceptions.py # Custom exceptions
│ │ ├── router.py # Health check & root endpoints
│ │ └── dependencies.py # Shared dependencies
│ ├── ollama/ # Ollama client (ready, not integrated yet)
│ │ ├── client.py # Async HTTP client
│ │ └── schemas.py # Ollama API models
│ └── main.py # Application factory & configuration
├── requirements.txt # Python dependencies (pinned)
├── .env # Environment variables (git-ignored)
├── .gitignore # Git ignore rules
├── AGENTS.md # LLM agent documentation + best practices
└── README.md # This file
│ │ └── router.py # Health check endpoints
│ └── main.py # Application factory
├── tests/ # Comprehensive test suite
│ ├── agents/ # Agent tests
│ ├── responses/ # Responses API tests
│ ├── chat/ # Chat completions tests
│ ├── models/ # Models API tests
│ └── core/ # Core tests
├── requirements.txt # Dependencies (pinned)
├── .env # Environment variables
├── AGENTS.md # Agent documentation
├── CLEANUP_TODO.md # Architecture notes
└── README.md # This file
```
**Key Architectural Decisions**:
- **Domain-based** structure (not file-type based)
- **Separation of concerns**: Routers → Services → Clients
- **Factory pattern** in main.py for testability
- **Custom base models** for consistent serialization
- **Async-first** for all I/O operations
### Code Style
Following FastAPI best practices:
- **Async routes** for ALL I/O operations (HTTP, database, file access)
- **Sync routes** only for CPU-intensive work or blocking SDKs
- **Type hints** on all functions and class attributes
- **Pydantic models** for ALL request/response validation
- **Dependency injection** for validation and shared resources
- **Business logic** in service modules, NOT in routers
- Follow PEP 8 style guidelines
- Document complex logic with docstrings
### Current Implementation Status
The API is currently set up with **mock responses** for development:
**✅ Implemented**:
- OpenAI-compatible API structure
- `/v1/chat/completions` endpoint (returns lorem ipsum)
- `/v1/models` endpoint (returns mistral-nemo:latest)
- `/health` and `/` endpoints
- Streaming support with SSE
- Exception handling
- Configuration management
- Async Ollama client (ready, not connected)
**🚧 TODO** (future integration):
- Connect chat completions to Ollama/PydanticAI
- Implement actual model listing from Ollama
- Add authentication/API keys
- Rate limiting
- Usage tracking
- More OpenAI-compatible endpoints
### Testing
```bash
# Install test dependencies
pip install -r requirements-dev.txt
# Run tests
# Run all tests
pytest
# Run tests with coverage
# Run with coverage
pytest --cov=src --cov-report=term-missing
# Current coverage: 62%
# Current coverage: 78.95% (75 tests passing)
```
## Documentation
**Test Organization:**
- Unit tests for all components
- Integration tests for API endpoints
- Streaming tests for SSE functionality
- Error handling tests
- Advanced features tests (stop sequences, max tokens, validation)
- See `AGENTS.md` for LLM agent instructions and package documentation
- FastAPI docs: https://fastapi.tiangolo.com/
- PydanticAI docs: https://ai.pydantic.dev/
- Ollama API: https://github.com/ollama/ollama/blob/main/docs/api.md
### Code Style
- **Async-first**: All I/O operations use async/await
- **Type hints**: All functions fully typed
- **Pydantic validation**: All request/response validation
- **Domain separation**: Clear boundaries between components
- **Single responsibility**: Each module has one clear purpose
## Security
### Version Locking
Minor version locking (`>=X.Y,<X.(Y+1)`) for security:
- Allows patch updates
- Blocks potentially breaking minor updates
- All dependencies checked for CVEs (2025-12-06)
### Best Practices
1. Never commit `.env` files
2. Use environment variables for sensitive config
3. Keep dependencies updated monthly
4. Validate all inputs with Pydantic
5. Use HTTPS in production
6. Implement rate limiting
## Deployment
### Docker Deployment (Coming Soon)
### Production Server
```bash
docker-compose up -d
# Multiple workers for production
uvicorn src.main:app --host 0.0.0.0 --port 8000 --workers 4
```
### Production Considerations
### Considerations
- Use a production ASGI server (uvicorn with multiple workers)
- Enable HTTPS with reverse proxy (nginx/caddy)
- Implement rate limiting
- Use reverse proxy (nginx/caddy) for HTTPS
- Enable rate limiting (SlowAPI or similar)
- Set up monitoring and logging
- Use a process manager (systemd/supervisor)
- Configure proper resource limits
- (Future) Ensure reliable network connectivity to Ollama instance
- (Future) Consider Ollama failover/redundancy strategies
- Configure resource limits
- Use process manager (systemd/supervisor)
## Troubleshooting
### Common Issues
**Issue**: Import errors
- Ensure virtual environment is activated
- Reinstall dependencies: `pip install -r requirements.txt`
- Verify Python 3.12+ is being used
**Issue**: Streaming not working
**Streaming not working:**
- Verify SSE-Starlette is installed
- Check client supports Server-Sent Events
- Review browser/tool compatibility
- Check test suite: `pytest tests/chat/test_router.py -k streaming`
- Test with: `pytest tests/responses/ -k streaming`
**Issue**: Tests failing
**Open WebUI can't connect:**
- Use Docker bridge gateway IP: `172.17.0.1:8000`
- Check firewall settings
- Verify server is running on `0.0.0.0`
**Tests failing:**
- Install test dependencies: `pip install -r requirements-dev.txt`
- Check async test configuration in `pytest.ini`
- Run with verbose output: `pytest -v`
- Activate virtual environment
- Run with verbose: `pytest -v`
**Reasoning not showing:**
- Ensure using Chat Completions endpoint (auto-enables reasoning)
- Or manually enable in Responses API: `"reasoning": {"effort": "medium", "summary": "auto"}`
- Check Open WebUI version supports `<think>` tags
## Future Roadmap
### Short-term
- [ ] Connect tatlock model to real PydanticAI agent
- [ ] Implement vector memory (Qdrant integration)
- [ ] Add authentication/API keys
- [ ] Rate limiting middleware
### Long-term
- [ ] Multi-model support (OpenAI, Anthropic, etc.)
- [ ] Advanced conversation memory
- [ ] Tool/function calling integration
- [ ] Usage tracking and analytics
## Contributing
1. Fork the repository
2. Create a feature branch
3. Make your changes
4. Add tests
5. Submit a pull request
3. Make changes with tests
4. Ensure tests pass: `pytest`
5. Submit pull request
## Documentation
- **AGENTS.md**: Agent architecture and best practices
- **CLEANUP_TODO.md**: Architecture decisions and future considerations
- **CHANGELOG.md**: Version history
- OpenAI Responses API: https://platform.openai.com/docs/api-reference/responses
- FastAPI: https://fastapi.tiangolo.com/
- PydanticAI: https://ai.pydantic.dev/
## License
@@ -298,5 +493,4 @@ docker-compose up -d
---
For detailed changelog, see [CHANGELOG.md](CHANGELOG.md).
For AI assistant instructions and package documentation, see [AGENTS.md](AGENTS.md).
**Note**: This is a testing/development API with mock responses. The architecture is production-ready and designed for easy integration with real LLM backends (PydanticAI, Ollama, OpenAI, etc.).