Clean up documentation to reflect current implementation
Remove confusing references to unimplemented features and clarify what's currently working vs prepared for future integration. README.md Changes: - Update title to reflect mock API (not "with Ollama Backend") - Remove architecture diagram showing Ollama/PydanticAI integration - Clarify current status section (mock API, integration prepared) - Fix uvicorn command: main:app → src.main:app - Update model examples: llama2 → mistral-nemo:latest - Mark Ollama requirements as future (not currently needed) - Update environment variables (Ollama config commented out) - Clarify API endpoints return mock responses - Update CVE check date: 2025-12-05 → 2025-12-06 - Fix testing section to use requirements-dev.txt - Remove Ollama troubleshooting (not connected yet) - Mark production Ollama considerations as future - Remove redundant changelog section (use CHANGELOG.md) AGENTS.md Changes: - Clarify project overview (mock API, not integrated) - Add status indicators to components section - Mark PydanticAI section as "for future implementation" - Mark Ollama section as "ready for future integration" - Add target model: mistral-nemo:latest - Update OpenAI compatibility section with implemented status - Fix Pydantic version: 2.10+ → 2.11+ (matches requirements) - Add implementation status to development guidelines - Mark common patterns as implemented vs future reference - Update CVE check date: 2025-12-05 → 2025-12-06 CHANGELOG.md Changes: - Clarify PydanticAI line: "for LLM integration" → "dependency (ready for future integration)" Key Improvements: - Clear distinction between implemented vs prepared features - No misleading references to Ollama/PydanticAI integration - Accurate model names (mistral-nemo:latest) - Correct command examples (src.main:app) - Proper date stamps (2025-12-06) - Removed confusing troubleshooting for unconnected services Status: Documentation now accurately reflects v0.1.0 mock API 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>
This commit is contained in:
@@ -4,12 +4,18 @@ This document contains instructions and documentation references for AI assistan
|
||||
|
||||
## Project Overview
|
||||
|
||||
This project implements an OpenAI-compatible API endpoint using FastAPI, with streaming support for LLM responses. The architecture consists of:
|
||||
This project implements an OpenAI-compatible API endpoint using FastAPI, with streaming support. Currently returns mock responses - infrastructure prepared for future Ollama/PydanticAI integration.
|
||||
|
||||
**Current State**: Production-ready mock API with OpenAI-compatible format
|
||||
**Future Integration**: Ollama and PydanticAI (client code ready, not connected)
|
||||
|
||||
### Components
|
||||
|
||||
- **FastAPI**: Web framework for the API layer
|
||||
- **PydanticAI**: Agent framework for LLM integration
|
||||
- **Ollama**: LLM backend running on a networked container
|
||||
- **Stream Coordinator**: Manages streaming responses in OpenAI-compatible format
|
||||
- **SSE-Starlette**: Server-Sent Events for streaming responses
|
||||
- **Pydantic**: Request/response validation
|
||||
- **PydanticAI**: Dependency installed, ready for future LLM integration
|
||||
- **Ollama**: Async client implemented, ready for future connection
|
||||
|
||||
## Documentation References
|
||||
|
||||
@@ -40,7 +46,8 @@ This project implements an OpenAI-compatible API endpoint using FastAPI, with st
|
||||
#### PydanticAI
|
||||
- **Official Documentation**: https://ai.pydantic.dev/
|
||||
- **Version**: 1.27.0 (Dec 2025)
|
||||
- **Key Topics**:
|
||||
- **Status**: Dependency installed, ready for future integration
|
||||
- **Key Topics** (for future implementation):
|
||||
- Agent creation and configuration
|
||||
- LLM provider integration (Ollama support)
|
||||
- Structured outputs with Pydantic
|
||||
@@ -53,7 +60,7 @@ This project implements an OpenAI-compatible API endpoint using FastAPI, with st
|
||||
|
||||
#### Pydantic
|
||||
- **Official Documentation**: https://docs.pydantic.dev/latest/
|
||||
- **Version**: 2.10+ (Required for PydanticAI)
|
||||
- **Version**: 2.11+ (Required for PydanticAI, currently using >=2.11,<2.13)
|
||||
- **Key Topics**:
|
||||
- Data validation and serialization
|
||||
- Field types and validators
|
||||
@@ -83,25 +90,31 @@ This project implements an OpenAI-compatible API endpoint using FastAPI, with st
|
||||
|
||||
#### Ollama API
|
||||
- **Official Documentation**: https://github.com/ollama/ollama/blob/main/docs/api.md
|
||||
- **Key Topics**:
|
||||
- **Status**: Async client implemented in `src/ollama/client.py`, ready for future integration
|
||||
- **Key Topics** (for future implementation):
|
||||
- REST API endpoints
|
||||
- Streaming responses
|
||||
- Model management
|
||||
- Generate and chat endpoints
|
||||
- Model configuration
|
||||
- **Current Model Target**: mistral-nemo:latest
|
||||
|
||||
### OpenAI API Compatibility
|
||||
|
||||
#### OpenAI API Reference
|
||||
- **Official Documentation**: https://platform.openai.com/docs/api-reference
|
||||
- **Key Endpoints to Implement**:
|
||||
- `/v1/chat/completions` - Chat completion with streaming
|
||||
- `/v1/models` - List available models
|
||||
- `/v1/completions` - Text completion (legacy)
|
||||
- **Key Features**:
|
||||
- Streaming with Server-Sent Events
|
||||
- Message format compatibility
|
||||
- Response structure compatibility
|
||||
- **Implemented Endpoints**:
|
||||
- ✅ `/v1/chat/completions` - Chat completion with streaming (mock responses)
|
||||
- ✅ `/v1/models` - List available models (mock listing)
|
||||
- **Future Endpoints**:
|
||||
- 🚧 `/v1/completions` - Text completion (legacy)
|
||||
- 🚧 `/v1/embeddings` - Text embeddings
|
||||
- **Implemented Features**:
|
||||
- ✅ Streaming with Server-Sent Events
|
||||
- ✅ Message format compatibility
|
||||
- ✅ Response structure compatibility
|
||||
- ✅ OpenAI error format
|
||||
- ✅ Request validation with Pydantic
|
||||
|
||||
## FastAPI Best Practices
|
||||
|
||||
@@ -270,55 +283,68 @@ app = create_application()
|
||||
|
||||
## Development Guidelines
|
||||
|
||||
### Code Structure
|
||||
- Use async/await for ALL I/O operations (database, HTTP, file access)
|
||||
- Use sync (def) for blocking SDKs or CPU-intensive work
|
||||
- Implement proper error handling and logging
|
||||
- Follow dependency injection for validation and shared resources
|
||||
- Use Pydantic models for ALL request/response validation
|
||||
- Keep business logic in service modules, not routers
|
||||
### Code Structure (Current Implementation)
|
||||
- ✅ Use async/await for ALL I/O operations (database, HTTP, file access)
|
||||
- ✅ Use sync (def) for blocking SDKs or CPU-intensive work
|
||||
- ✅ Implement proper error handling and logging
|
||||
- ✅ Follow dependency injection for validation and shared resources
|
||||
- ✅ Use Pydantic models for ALL request/response validation
|
||||
- ✅ Keep business logic in service modules, not routers
|
||||
- ✅ Domain-based project structure (not file-type based)
|
||||
|
||||
### Security Considerations
|
||||
- Validate all inputs using Pydantic models
|
||||
- Implement rate limiting for API endpoints
|
||||
- Use environment variables for sensitive configuration
|
||||
- Keep dependencies updated (check for CVEs regularly)
|
||||
- ✅ Validate all inputs using Pydantic models
|
||||
- ✅ Use environment variables for sensitive configuration
|
||||
- ✅ Keep dependencies updated (all CVE-checked as of 2025-12-06)
|
||||
- ✅ Minor version locking for supply chain protection
|
||||
- 🚧 Implement rate limiting for API endpoints (future)
|
||||
- 🚧 Add authentication/API keys (future)
|
||||
|
||||
### Testing
|
||||
- Write integration tests for API endpoints
|
||||
- Test streaming functionality thoroughly
|
||||
- Mock Ollama responses for unit tests
|
||||
- Validate OpenAI API compatibility
|
||||
### Testing (Current Coverage: 62%)
|
||||
- ✅ Integration tests for API endpoints
|
||||
- ✅ Streaming functionality with 20s timeout protection
|
||||
- ✅ Async test support with pytest-asyncio
|
||||
- ✅ Validate OpenAI API compatibility
|
||||
- ✅ Mock responses for all endpoints
|
||||
- 🚧 Future: Mock Ollama responses when integrated
|
||||
|
||||
### Configuration
|
||||
- Use `.env` files for local development
|
||||
- Document all environment variables in README
|
||||
- Provide sensible defaults where possible
|
||||
- Support container-based configuration
|
||||
- ✅ Use `.env` files for local development
|
||||
- ✅ Document all environment variables in README
|
||||
- ✅ Provide sensible defaults where possible
|
||||
- ✅ BaseSettings from pydantic-settings
|
||||
- 🚧 Support container-based configuration (future)
|
||||
|
||||
## Common Patterns
|
||||
|
||||
### Streaming Response Pattern
|
||||
### Streaming Response Pattern (✅ Implemented)
|
||||
|
||||
See `src/chat/router.py` for the current implementation:
|
||||
|
||||
```python
|
||||
from sse_starlette.sse import EventSourceResponse
|
||||
from fastapi import FastAPI
|
||||
|
||||
async def event_generator():
|
||||
# Stream events from Ollama/PydanticAI
|
||||
yield {"data": "chunk1"}
|
||||
yield {"data": "chunk2"}
|
||||
# Currently yields mock lorem ipsum chunks
|
||||
# Future: Stream from Ollama/PydanticAI
|
||||
yield {"data": chunk.model_dump_json()}
|
||||
yield {"data": "[DONE]"}
|
||||
|
||||
@app.get("/stream")
|
||||
@app.post("/stream")
|
||||
async def stream():
|
||||
return EventSourceResponse(event_generator())
|
||||
```
|
||||
|
||||
### PydanticAI Agent Pattern
|
||||
### PydanticAI Agent Pattern (🚧 Future Reference)
|
||||
|
||||
For future integration when connecting to Ollama:
|
||||
|
||||
```python
|
||||
from pydantic_ai import Agent
|
||||
|
||||
agent = Agent(
|
||||
'ollama:llama2', # Or other Ollama model
|
||||
'ollama:mistral-nemo', # Target model
|
||||
# Configuration here
|
||||
)
|
||||
|
||||
@@ -326,13 +352,16 @@ agent = Agent(
|
||||
result = await agent.run('Your prompt')
|
||||
```
|
||||
|
||||
### OpenAI-Compatible Response Format
|
||||
### OpenAI-Compatible Response Format (✅ Implemented)
|
||||
|
||||
Current implementation in `src/chat/schemas.py`:
|
||||
|
||||
```python
|
||||
{
|
||||
"id": "chatcmpl-123",
|
||||
"object": "chat.completion.chunk",
|
||||
"created": 1234567890,
|
||||
"model": "model-name",
|
||||
"model": "mistral-nemo:latest",
|
||||
"choices": [{
|
||||
"index": 0,
|
||||
"delta": {"content": "response"},
|
||||
@@ -349,4 +378,4 @@ This document should be updated when:
|
||||
- Breaking API changes occur
|
||||
- Security vulnerabilities are discovered
|
||||
|
||||
Last updated: 2025-12-05
|
||||
Last updated: 2025-12-06
|
||||
|
||||
+1
-1
@@ -33,7 +33,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
|
||||
- Project initialization
|
||||
- Python 3.12.11 environment
|
||||
- FastAPI 0.123.9 web framework
|
||||
- PydanticAI 1.27.0 for LLM integration
|
||||
- PydanticAI 1.27.0 dependency (ready for future integration)
|
||||
- Mock chat completions (lorem ipsum responses)
|
||||
- Mock model listing (mistral-nemo:latest)
|
||||
- Testing infrastructure (pytest, coverage, ruff, mypy)
|
||||
|
||||
@@ -1,43 +1,36 @@
|
||||
# OpenAI-Compatible API with Ollama Backend
|
||||
# OpenAI-Compatible API
|
||||
|
||||
A FastAPI-based service that provides an OpenAI-compatible API endpoint, powered by PydanticAI and Ollama for LLM inference.
|
||||
A FastAPI-based service that provides OpenAI-compatible API endpoints with streaming support. Currently returns mock responses - ready for future Ollama/PydanticAI integration.
|
||||
|
||||
## Architecture
|
||||
## Current Status
|
||||
|
||||
```
|
||||
┌─────────────┐ ┌──────────────────┐ ┌─────────────┐
|
||||
│ Client │─────▶│ FastAPI Server │─────▶│ Ollama │
|
||||
│ │◀─────│ (Stream Coord.) │◀─────│ Container │
|
||||
└─────────────┘ └──────────────────┘ └─────────────┘
|
||||
│
|
||||
▼
|
||||
┌──────────┐
|
||||
│Pydantic │
|
||||
│ AI │
|
||||
└──────────┘
|
||||
```
|
||||
**✅ Production-ready mock API** with OpenAI-compatible format
|
||||
**🚧 Ollama/PydanticAI integration** prepared but not connected
|
||||
|
||||
### Components
|
||||
## Components
|
||||
|
||||
- **FastAPI**: High-performance web framework providing the API layer
|
||||
- **Stream Coordinator**: Manages streaming responses in OpenAI-compatible format
|
||||
- **PydanticAI**: Agent framework handling LLM integration and structured outputs
|
||||
- **Ollama**: External LLM backend (networked, managed separately)
|
||||
- **SSE-Starlette**: Server-Sent Events for streaming responses
|
||||
- **Pydantic**: Type-safe request/response handling and validation
|
||||
- **Ollama Client**: Async HTTP client prepared for future integration
|
||||
- **PydanticAI**: Ready for LLM integration (not yet connected)
|
||||
|
||||
## Features
|
||||
|
||||
- OpenAI-compatible API endpoints
|
||||
- Streaming responses with Server-Sent Events
|
||||
- PydanticAI integration for robust LLM interactions
|
||||
- Networked Ollama support
|
||||
- Type-safe request/response handling with Pydantic
|
||||
- Async/await throughout for optimal performance
|
||||
- ✅ OpenAI-compatible API endpoints (`/v1/chat/completions`, `/v1/models`)
|
||||
- ✅ Streaming responses with Server-Sent Events (SSE)
|
||||
- ✅ Type-safe request/response handling with Pydantic
|
||||
- ✅ Async/await throughout for optimal performance
|
||||
- ✅ Comprehensive test suite (62% coverage)
|
||||
- ✅ Domain-based architecture following FastAPI best practices
|
||||
- 🚧 Ollama integration (client ready, not connected)
|
||||
- 🚧 PydanticAI integration (dependency installed, not connected)
|
||||
|
||||
## Requirements
|
||||
|
||||
- Python 3.12+ (Python 3.12.11 recommended for security)
|
||||
- Network access to an existing Ollama instance (managed externally)
|
||||
- No external dependencies required for mock API
|
||||
- (Future: Network access to Ollama instance for LLM integration)
|
||||
|
||||
## Installation
|
||||
|
||||
@@ -61,32 +54,32 @@ source venv/bin/activate # On Windows: venv\Scripts\activate
|
||||
pip install -r requirements.txt
|
||||
```
|
||||
|
||||
### 4. Configure environment variables
|
||||
### 4. Configure environment variables (Optional)
|
||||
|
||||
Create a `.env` file in the project root:
|
||||
Create a `.env` file in the project root for custom configuration:
|
||||
|
||||
```env
|
||||
# Ollama Configuration (point to your existing Ollama instance)
|
||||
OLLAMA_HOST=http://your-ollama-host:11434
|
||||
OLLAMA_MODEL=llama2
|
||||
|
||||
# API Configuration
|
||||
API_HOST=0.0.0.0
|
||||
API_PORT=8000
|
||||
API_RELOAD=true
|
||||
|
||||
# Logging
|
||||
LOG_LEVEL=info
|
||||
LOG_LEVEL=INFO
|
||||
|
||||
# Future Ollama Configuration (not yet integrated)
|
||||
# OLLAMA_HOST=http://localhost:11434
|
||||
# OLLAMA_DEFAULT_MODEL=mistral-nemo:latest
|
||||
# OLLAMA_TIMEOUT=120
|
||||
```
|
||||
|
||||
**Note**: Update `OLLAMA_HOST` to point to your existing Ollama instance. Ensure the Ollama service is accessible from your network and has the required models installed.
|
||||
**Note**: The API works with defaults. Environment variables are optional for customization. Ollama configuration is prepared but not currently used.
|
||||
|
||||
## Usage
|
||||
|
||||
### Start the development server
|
||||
|
||||
```bash
|
||||
uvicorn main:app --reload
|
||||
uvicorn src.main:app --reload
|
||||
```
|
||||
|
||||
The API will be available at `http://localhost:8000`
|
||||
@@ -95,11 +88,13 @@ The API will be available at `http://localhost:8000`
|
||||
|
||||
#### Chat Completions (OpenAI-compatible)
|
||||
|
||||
Returns mock lorem ipsum responses:
|
||||
|
||||
```bash
|
||||
curl http://localhost:8000/v1/chat/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "llama2",
|
||||
"model": "mistral-nemo:latest",
|
||||
"messages": [
|
||||
{"role": "user", "content": "Hello, how are you?"}
|
||||
],
|
||||
@@ -109,10 +104,14 @@ curl http://localhost:8000/v1/chat/completions \
|
||||
|
||||
#### List Models
|
||||
|
||||
Returns mock model listing:
|
||||
|
||||
```bash
|
||||
curl http://localhost:8000/v1/models
|
||||
```
|
||||
|
||||
Response: `{"object": "list", "data": [{"id": "mistral-nemo:latest", ...}]}`
|
||||
|
||||
### Interactive API Documentation
|
||||
|
||||
- Swagger UI: `http://localhost:8000/docs`
|
||||
@@ -124,11 +123,11 @@ curl http://localhost:8000/v1/models
|
||||
|
||||
This project uses minor version locking (`>=X.Y,<X.(Y+1)`) to protect against supply chain attacks while allowing patch updates. All dependencies have been:
|
||||
|
||||
- Checked for known CVEs (as of 2025-12-05)
|
||||
- Checked for known CVEs (as of 2025-12-06)
|
||||
- Pinned to secure minor versions
|
||||
- Documented with version rationale in `requirements.txt`
|
||||
|
||||
### CVE Status (2025-12-05)
|
||||
### CVE Status (2025-12-06)
|
||||
|
||||
- **FastAPI 0.123.9**: No known vulnerabilities
|
||||
- **Uvicorn 0.38.0**: No known vulnerabilities
|
||||
@@ -228,10 +227,15 @@ The API is currently set up with **mock responses** for development:
|
||||
|
||||
```bash
|
||||
# Install test dependencies
|
||||
pip install pytest pytest-asyncio httpx
|
||||
pip install -r requirements-dev.txt
|
||||
|
||||
# Run tests
|
||||
pytest
|
||||
|
||||
# Run tests with coverage
|
||||
pytest --cov=src --cov-report=term-missing
|
||||
|
||||
# Current coverage: 62%
|
||||
```
|
||||
|
||||
## Documentation
|
||||
@@ -257,27 +261,28 @@ docker-compose up -d
|
||||
- Set up monitoring and logging
|
||||
- Use a process manager (systemd/supervisor)
|
||||
- Configure proper resource limits
|
||||
- Ensure reliable network connectivity to Ollama instance
|
||||
- Consider Ollama failover/redundancy strategies
|
||||
- (Future) Ensure reliable network connectivity to Ollama instance
|
||||
- (Future) Consider Ollama failover/redundancy strategies
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
### Common Issues
|
||||
|
||||
**Issue**: Cannot connect to Ollama
|
||||
- Verify `OLLAMA_HOST` in `.env` points to the correct address
|
||||
- Ensure the Ollama instance is running and accessible
|
||||
- Check network connectivity and firewall rules
|
||||
- Test connectivity: `curl http://your-ollama-host:11434/api/tags`
|
||||
**Issue**: Import errors
|
||||
- Ensure virtual environment is activated
|
||||
- Reinstall dependencies: `pip install -r requirements.txt`
|
||||
- Verify Python 3.12+ is being used
|
||||
|
||||
**Issue**: Streaming not working
|
||||
- Verify SSE-Starlette is installed
|
||||
- Check client supports Server-Sent Events
|
||||
- Review browser/tool compatibility
|
||||
- Check test suite: `pytest tests/chat/test_router.py -k streaming`
|
||||
|
||||
**Issue**: Import errors
|
||||
- Ensure virtual environment is activated
|
||||
- Reinstall dependencies: `pip install -r requirements.txt`
|
||||
**Issue**: Tests failing
|
||||
- Install test dependencies: `pip install -r requirements-dev.txt`
|
||||
- Check async test configuration in `pytest.ini`
|
||||
- Run with verbose output: `pytest -v`
|
||||
|
||||
## Contributing
|
||||
|
||||
@@ -291,14 +296,7 @@ docker-compose up -d
|
||||
|
||||
[Add your license here]
|
||||
|
||||
## Changelog
|
||||
|
||||
### 2025-12-05 - Initial Setup
|
||||
- Project structure created
|
||||
- Dependencies configured with CVE checks
|
||||
- Security: Minor version locking implemented
|
||||
- Documentation added
|
||||
|
||||
---
|
||||
|
||||
For detailed changelog, see [CHANGELOG.md](CHANGELOG.md).
|
||||
For AI assistant instructions and package documentation, see [AGENTS.md](AGENTS.md).
|
||||
|
||||
Reference in New Issue
Block a user