Add Responses API with streaming, history, and advanced features

Implements Phases 2, 3, and 6: Complete Responses API implementation

Core API (Phase 2):
- OpenAI Responses API format with structured output items
- Streaming and non-streaming support via SSE-Starlette
- Reasoning items (thinking summaries)
- Function call items (tool execution)
- Message items (assistant responses)
- Router, schemas, service, and streaming coordinator

Conversation History (Phase 3):
- Hybrid client/server approach
- Auto-generated deterministic conversation IDs
- Configurable max turns with automatic trimming
- Context window management with token counting
- Token usage statistics
- Placeholder for future vector memory integration

Advanced Features (Phase 6):
- Parameter validation with Pydantic field validators:
  - Temperature: 0.0-2.0 range enforcement
  - Reasoning effort: 6 levels (none to xhigh)
  - Max output tokens: positive integer enforcement
  - Stop sequences: up to 4, non-empty strings
- Real-time stop sequence detection during streaming
- Real-time max tokens enforcement with token counting
- Graceful error handling and OpenAI-compatible error format

Testing:
- 9 unit tests for API endpoints and streaming
- 11 unit tests for error handling
- 13 unit tests for conversation history and context
- 12 unit tests for advanced features and validation
- Total: 45 tests with comprehensive coverage

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
This commit is contained in:
2025-12-06 19:38:00 +01:00
co-authored by Claude
parent 4e6ca4466b
commit ff6c3cf1b5
12 changed files with 2680 additions and 0 deletions
+116
View File
@@ -0,0 +1,116 @@
"""
Responses router.
OpenAI-compatible /v1/responses endpoint with streaming support.
"""
import logging
from fastapi import APIRouter, HTTPException
from sse_starlette.sse import EventSourceResponse
from src.responses import service
from src.responses.schemas import ResponseRequest, Response
from src.core.exceptions import ModelNotFoundError, AppException
logger = logging.getLogger(__name__)
router = APIRouter(prefix="/responses", tags=["responses"])
@router.post("", response_model=Response)
async def create_response(
request: ResponseRequest,
) -> Response | EventSourceResponse:
"""
Create a response using Responses API format.
Supports:
- Reasoning summaries (thinking/reasoning display)
- Function calling (tool usage)
- Streaming responses
- Multi-turn conversations
- Error handling
Args:
request: Response request with model, input, optional reasoning/tools
Returns:
Response object or SSE stream
Example non-streaming request:
POST /v1/responses
{
"model": "lorem-tester",
"input": [{"role": "user", "content": "Hello"}],
"reasoning": {"effort": "medium", "summary": "auto"},
"stream": false
}
Example streaming request:
POST /v1/responses
{
"model": "lorem-tester",
"input": [{"role": "user", "content": "Hello"}],
"stream": true
}
Response format (non-streaming):
{
"id": "resp_...",
"object": "response",
"created_at": 1733529600,
"model": "lorem-tester",
"status": "completed",
"output": [
{
"type": "reasoning",
"id": "rs_...",
"summary": ["Analyzing...", "Considering..."]
},
{
"type": "message",
"id": "msg_...",
"role": "assistant",
"content": [{"type": "output_text", "text": "Lorem ipsum..."}]
}
],
"usage": {
"input_tokens": 10,
"output_tokens": 50,
"reasoning_tokens": 20,
"total_tokens": 80
}
}
Streaming format (SSE):
event: response.reasoning_summary_text.delta
data: {"delta": "Analyzing..."}
event: response.output_text.delta
data: {"delta": "Lorem"}
event: response.done
data: {"response": {...}}
"""
logger.info(f"Response request for model: {request.model}")
try:
if request.stream:
logger.info("Streaming response requested")
return EventSourceResponse(
service.create_response_stream(request)
)
return await service.create_response(request)
except ModelNotFoundError as e:
logger.error(f"Model not found: {e}")
raise HTTPException(status_code=404, detail=str(e))
except AppException as e:
logger.error(f"Application error: {e}")
raise HTTPException(status_code=e.status_code, detail=e.message)
except Exception as e:
logger.error(f"Unexpected error: {e}", exc_info=True)
raise HTTPException(status_code=500, detail="Internal server error")