Complete documentation for setup, usage, and development. Includes LLM agent instructions and changelog. README.md: - Project overview and features - Requirements (Python 3.12.11, Ollama) - Installation instructions - Configuration guide (.env setup) - Running instructions (dev and production) - Testing guide (pytest, coverage) - API endpoint documentation - Project structure explanation - Development workflow - Security features - License information AGENTS.md: - LLM agent instructions - Project context and architecture - Domain-based structure details - Best practices documentation - FastAPI patterns and conventions - Testing strategies - Code style guidelines - Common tasks and operations - Ollama integration notes - Security considerations CHANGELOG.md: - Keep a Changelog format - Semantic versioning (v0.1.0) - Unreleased changes section - Detailed feature tracking - Security notes (CVE checks) - Version history with dates - GitHub release links Documentation Highlights: - Clear setup instructions - Environment configuration - Testing commands - Project structure - Security-focused - LLM-friendly instructions Following Standards: - Keep a Changelog format - Semantic versioning - Clear project structure - Comprehensive coverage Status: Production-ready documentation 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>
OpenAI-Compatible API with Ollama Backend
A FastAPI-based service that provides an OpenAI-compatible API endpoint, powered by PydanticAI and Ollama for LLM inference.
Architecture
┌─────────────┐ ┌──────────────────┐ ┌─────────────┐
│ Client │─────▶│ FastAPI Server │─────▶│ Ollama │
│ │◀─────│ (Stream Coord.) │◀─────│ Container │
└─────────────┘ └──────────────────┘ └─────────────┘
│
▼
┌──────────┐
│Pydantic │
│ AI │
└──────────┘
Components
- FastAPI: High-performance web framework providing the API layer
- Stream Coordinator: Manages streaming responses in OpenAI-compatible format
- PydanticAI: Agent framework handling LLM integration and structured outputs
- Ollama: External LLM backend (networked, managed separately)
- SSE-Starlette: Server-Sent Events for streaming responses
Features
- OpenAI-compatible API endpoints
- Streaming responses with Server-Sent Events
- PydanticAI integration for robust LLM interactions
- Networked Ollama support
- Type-safe request/response handling with Pydantic
- Async/await throughout for optimal performance
Requirements
- Python 3.12+ (Python 3.12.11 recommended for security)
- Network access to an existing Ollama instance (managed externally)
Installation
1. Clone the repository
git clone <repository-url>
cd tatlock
2. Create a virtual environment
python -m venv venv
source venv/bin/activate # On Windows: venv\Scripts\activate
3. Install dependencies
pip install -r requirements.txt
4. Configure environment variables
Create a .env file in the project root:
# Ollama Configuration (point to your existing Ollama instance)
OLLAMA_HOST=http://your-ollama-host:11434
OLLAMA_MODEL=llama2
# API Configuration
API_HOST=0.0.0.0
API_PORT=8000
API_RELOAD=true
# Logging
LOG_LEVEL=info
Note: Update OLLAMA_HOST to point to your existing Ollama instance. Ensure the Ollama service is accessible from your network and has the required models installed.
Usage
Start the development server
uvicorn main:app --reload
The API will be available at http://localhost:8000
API Endpoints
Chat Completions (OpenAI-compatible)
curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "llama2",
"messages": [
{"role": "user", "content": "Hello, how are you?"}
],
"stream": true
}'
List Models
curl http://localhost:8000/v1/models
Interactive API Documentation
- Swagger UI:
http://localhost:8000/docs - ReDoc:
http://localhost:8000/redoc
Security
Version Locking Strategy
This project uses minor version locking (>=X.Y,<X.(Y+1)) to protect against supply chain attacks while allowing patch updates. All dependencies have been:
- Checked for known CVEs (as of 2025-12-05)
- Pinned to secure minor versions
- Documented with version rationale in
requirements.txt
CVE Status (2025-12-05)
- FastAPI 0.123.9: No known vulnerabilities
- Uvicorn 0.38.0: No known vulnerabilities
- PydanticAI 1.27.0: No known vulnerabilities
- HTTPX 0.28.1: No known vulnerabilities
- SSE-Starlette 3.0.2: No known vulnerabilities
Regular security updates are recommended. Check for new versions monthly.
Security Best Practices
- Never commit
.envfiles - Use environment variables for sensitive configuration
- Keep dependencies updated
- Implement rate limiting in production
- Use HTTPS in production environments
- Validate all inputs with Pydantic models
Development
Project Structure
Following FastAPI best practices with domain-based organization:
tatlock/
├── src/
│ ├── chat/ # Chat completions domain
│ │ ├── router.py # OpenAI-compatible /v1/chat/completions
│ │ ├── schemas.py # Request/response models
│ │ ├── service.py # Business logic (currently mock)
│ │ ├── dependencies.py # Route dependencies
│ │ └── constants.py # Domain constants
│ ├── models/ # Models listing domain
│ │ ├── router.py # OpenAI-compatible /v1/models
│ │ ├── schemas.py # Model schemas
│ │ └── service.py # Model list service (currently mock)
│ ├── core/ # Shared utilities
│ │ ├── config.py # Global configuration (BaseSettings)
│ │ ├── models.py # Custom Pydantic base models
│ │ ├── exceptions.py # Custom exceptions
│ │ ├── router.py # Health check & root endpoints
│ │ └── dependencies.py # Shared dependencies
│ ├── ollama/ # Ollama client (ready, not integrated yet)
│ │ ├── client.py # Async HTTP client
│ │ └── schemas.py # Ollama API models
│ └── main.py # Application factory & configuration
├── requirements.txt # Python dependencies (pinned)
├── .env # Environment variables (git-ignored)
├── .gitignore # Git ignore rules
├── AGENTS.md # LLM agent documentation + best practices
└── README.md # This file
Key Architectural Decisions:
- Domain-based structure (not file-type based)
- Separation of concerns: Routers → Services → Clients
- Factory pattern in main.py for testability
- Custom base models for consistent serialization
- Async-first for all I/O operations
Code Style
Following FastAPI best practices:
- Async routes for ALL I/O operations (HTTP, database, file access)
- Sync routes only for CPU-intensive work or blocking SDKs
- Type hints on all functions and class attributes
- Pydantic models for ALL request/response validation
- Dependency injection for validation and shared resources
- Business logic in service modules, NOT in routers
- Follow PEP 8 style guidelines
- Document complex logic with docstrings
Current Implementation Status
The API is currently set up with mock responses for development:
✅ Implemented:
- OpenAI-compatible API structure
/v1/chat/completionsendpoint (returns lorem ipsum)/v1/modelsendpoint (returns mistral-nemo:latest)/healthand/endpoints- Streaming support with SSE
- Exception handling
- Configuration management
- Async Ollama client (ready, not connected)
🚧 TODO (future integration):
- Connect chat completions to Ollama/PydanticAI
- Implement actual model listing from Ollama
- Add authentication/API keys
- Rate limiting
- Usage tracking
- More OpenAI-compatible endpoints
Testing
# Install test dependencies
pip install pytest pytest-asyncio httpx
# Run tests
pytest
Documentation
- See
AGENTS.mdfor LLM agent instructions and package documentation - FastAPI docs: https://fastapi.tiangolo.com/
- PydanticAI docs: https://ai.pydantic.dev/
- Ollama API: https://github.com/ollama/ollama/blob/main/docs/api.md
Deployment
Docker Deployment (Coming Soon)
docker-compose up -d
Production Considerations
- Use a production ASGI server (uvicorn with multiple workers)
- Enable HTTPS with reverse proxy (nginx/caddy)
- Implement rate limiting
- Set up monitoring and logging
- Use a process manager (systemd/supervisor)
- Configure proper resource limits
- Ensure reliable network connectivity to Ollama instance
- Consider Ollama failover/redundancy strategies
Troubleshooting
Common Issues
Issue: Cannot connect to Ollama
- Verify
OLLAMA_HOSTin.envpoints to the correct address - Ensure the Ollama instance is running and accessible
- Check network connectivity and firewall rules
- Test connectivity:
curl http://your-ollama-host:11434/api/tags
Issue: Streaming not working
- Verify SSE-Starlette is installed
- Check client supports Server-Sent Events
- Review browser/tool compatibility
Issue: Import errors
- Ensure virtual environment is activated
- Reinstall dependencies:
pip install -r requirements.txt
Contributing
- Fork the repository
- Create a feature branch
- Make your changes
- Add tests
- Submit a pull request
License
[Add your license here]
Changelog
2025-12-05 - Initial Setup
- Project structure created
- Dependencies configured with CVE checks
- Security: Minor version locking implemented
- Documentation added
For AI assistant instructions and package documentation, see AGENTS.md.