From 9f3eda869583411cd4931fe4f18785ca4133e8fc Mon Sep 17 00:00:00 2001 From: Jeroen Schweitzer Date: Sat, 6 Dec 2025 19:40:30 +0100 Subject: [PATCH] Add implementation planning and architecture documents MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit IMPLEMENTATION_PLAN.md: - Phase-by-phase implementation plan - Success criteria for each phase - Testing requirements - Dependencies and prerequisites CLEANUP_TODO.md: - Architecture decision log - Future considerations and trade-offs - Migration path notes - Technical debt tracking These documents provide context for implementation decisions and serve as a reference for future development. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude --- CLEANUP_TODO.md | 206 ++++++++ IMPLEMENTATION_PLAN.md | 1124 ++++++++++++++++++++++++++++++++++++++++ 2 files changed, 1330 insertions(+) create mode 100644 CLEANUP_TODO.md create mode 100644 IMPLEMENTATION_PLAN.md diff --git a/CLEANUP_TODO.md b/CLEANUP_TODO.md new file mode 100644 index 0000000..51a8dca --- /dev/null +++ b/CLEANUP_TODO.md @@ -0,0 +1,206 @@ +# Cleanup Todo: Chat Completions API Wrapper + +**Status**: CHANGED - Chat Completions is now a wrapper, not legacy code + +## New Architecture (As of 2025-12-06) + +Chat Completions API is now a **compatibility wrapper** around Responses API: +- Maintains single source of truth (Responses API) +- Converts Responses output to Chat Completions format +- Embeds reasoning in `` tags for Open WebUI +- Provides backward compatibility + +## When to Consider Cleanup + +Only remove Chat Completions wrapper if **ALL** of these conditions are met: + +1. ✅ Open WebUI natively supports Responses API +2. ✅ All other consumers migrated to `/v1/responses` +3. ✅ No tools/clients require `/v1/chat/completions` +4. ✅ Team consensus on removal + +**Current Status**: KEEP - Provides essential compatibility for Open WebUI + +--- + +## Optional Removal Checklist + +### Phase 1: Verify All Consumers Migrated + +- [ ] Open WebUI using `/v1/responses` natively +- [ ] No external tools using `/v1/chat/completions` +- [ ] Check server logs for chat completions usage +- [ ] Confirm with team/users + +### Phase 2: Deprecation Period + +- [ ] Add deprecation notice to `/v1/chat/completions` docs +- [ ] Update OpenAPI/Swagger docs with deprecation warning +- [ ] Log deprecation warnings on chat completions usage +- [ ] Maintain for at least one version cycle + +### Phase 3: Code Removal (If Approved) + +#### Remove Chat Domain + +- [ ] Delete `src/chat/` directory + - [ ] `src/chat/router.py` + - [ ] `src/chat/service.py` + - [ ] `src/chat/schemas.py` + - [ ] `src/chat/constants.py` + - [ ] Keep `src/chat/__init__.py` if needed for imports + +#### Update Main Application + +- [ ] Edit `src/main.py` + - [ ] Remove `from src.chat.router import router as chat_router` + - [ ] Remove `app.include_router(chat_router, prefix=config.API_PREFIX)` + +#### Remove Tests + +- [ ] Delete `tests/chat/` directory +- [ ] Remove chat fixtures from `tests/conftest.py` + +### Phase 4: Documentation Updates + +- [ ] Update README.md + - [ ] Remove `/v1/chat/completions` examples + - [ ] Mark as "Responses API only" + - [ ] Add migration guide link + +- [ ] Update CHANGELOG.md + ```markdown + ## [Unreleased] + + ### Removed + - Chat Completions API wrapper (deprecated in v0.X.0) + - All clients must now use POST /v1/responses + + ### Migration + - See MIGRATION.md for Responses API format + ``` + +- [ ] Create MIGRATION.md + - [ ] Show before/after examples + - [ ] Explain `` tags vs proper reasoning items + - [ ] Conversion guide + +--- + +## Current Recommended Action + +**DO NOT REMOVE** - The Chat Completions wrapper serves these purposes: + +1. **Open WebUI Compatibility** - Provides `` tag format +2. **Backward Compatibility** - Supports existing clients +3. **Migration Path** - Smooth transition to Responses API +4. **Minimal Cost** - Thin wrapper, low maintenance + +### Documentation Instead of Removal + +Update documentation to clarify architecture: + +- [x] Chat Completions is a wrapper around Responses API +- [ ] Document wrapper behavior in README.md +- [ ] Explain `` tag conversion in AGENTS.md +- [ ] Note single source of truth (Responses API) + +--- + +## Alternative: Keep as Compatibility Layer + +**Recommended approach**: Keep Chat Completions as permanent compatibility layer + +### Benefits + +- ✅ Supports multiple client types +- ✅ Single codebase maintains both formats +- ✅ No breaking changes for users +- ✅ Minimal maintenance burden + +### Maintenance + +- Update wrapper when Responses API changes +- Keep tests for both endpoints +- Document as compatibility layer +- Monitor usage metrics + +--- + +## Git Commit (If Documenting, Not Removing) + +```bash +git add -A +git commit -m "Document Chat Completions as Responses API wrapper + +Update documentation to clarify that Chat Completions endpoint is a +compatibility wrapper around Responses API, not legacy code. + +Architecture: +- Single source of truth: Responses API +- Chat Completions wraps Responses API +- Converts reasoning to tags for Open WebUI +- Maintains backward compatibility + +Documentation Updates: +- README.md: Explain wrapper architecture +- AGENTS.md: Document conversion behavior +- CLEANUP_TODO.md: Update removal guidance + +Status: Chat Completions is a supported compatibility layer + +🤖 Generated with [Claude Code](https://claude.com/claude-code) + +Co-Authored-By: Claude " +``` + +--- + +## Decision Criteria + +Ask these questions before removing: + +1. **Is anyone using `/v1/chat/completions`?** + - Check server logs + - Survey users/team + - Monitor for 30+ days + +2. **Does Open WebUI support Responses API natively?** + - Version with native support released? + - Tested and verified working? + - Migration path clear? + +3. **Is there value in multiple formats?** + - Some clients prefer Chat Completions + - Easier integration for simple use cases + - Lower barrier to entry + +4. **What's the maintenance cost?** + - Currently: LOW (thin wrapper) + - If removed: Users must update clients + - Cost/benefit analysis + +--- + +## Conclusion + +**Current Recommendation**: **KEEP** Chat Completions wrapper + +- Provides value (Open WebUI compatibility) +- Low maintenance cost +- Enables smooth migration +- No compelling reason to remove + +**Revisit** when: +- Open WebUI natively supports Responses API +- Usage metrics show zero chat completions traffic +- Team decides compatibility layer not needed + +--- + +## Notes + +- This replaces the original "remove chat completions" plan +- Architecture has evolved: wrapper is intentional, not legacy +- Removal is now **optional** and **conditional** +- Focus on documentation, not deletion diff --git a/IMPLEMENTATION_PLAN.md b/IMPLEMENTATION_PLAN.md new file mode 100644 index 0000000..f8700c1 --- /dev/null +++ b/IMPLEMENTATION_PLAN.md @@ -0,0 +1,1124 @@ +# Implementation Plan: Lorem Tester Agent with Responses API + +## Overview + +Build production-ready OpenAI Responses API (`/v1/responses`) with Lorem Tester agent. All infrastructure is real code - only mock at PydanticAI interface boundary. + +**API Format**: OpenAI Responses API (NOT Chat Completions) +**Timeline**: Phased implementation, each phase independently testable with Open WebUI + +--- + +## Architecture + +``` +┌─────────────────────────────────────────────────────────────┐ +│ Responses Router & Service (REAL implementation) │ +│ - Conversation history with optional conversation_id │ +│ - Streaming coordination (reasoning, tools, content) │ +│ - Output array builder (reasoning/function/message items) │ +│ - Thinking block formatting │ +│ - Error handling │ +└─────────────────────┬───────────────────────────────────────┘ + │ + ▼ + ┌────────────────────────┐ + │ Agent Interface │ ◄── Abstraction layer + └────────────────────────┘ + │ + ┌───────────┴───────────┐ + ▼ ▼ + ┌──────────┐ ┌─────────────┐ + │ Lorem │ │ Tatlock │ + │ Tester │ │ Agent │ + │ │ │ (future) │ + │ (Mock │ │ (Real │ + │ PydanticAI) │ PydanticAI) │ + └──────────┘ └─────────────┘ +``` + +--- + +## Phase 1: Agent Interface & Multi-Model Support + +### 1.1 Create Agent Interface Abstraction + +**File**: `src/agents/base.py` + +```python +from abc import ABC, abstractmethod +from typing import AsyncGenerator +from src.responses.schemas import OutputItem + +class AgentInterface(ABC): + """Abstract interface for all agents (Lorem Tester, Tatlock, etc.)""" + + @abstractmethod + async def generate_response( + self, + messages: list[dict], + **kwargs + ) -> AsyncGenerator[OutputItem, None]: + """ + Generate streaming response as output items. + Yields: OutputItem objects (reasoning, function_call, message) + """ + pass + + @abstractmethod + async def supports_tools(self) -> bool: + """Whether agent supports function calling""" + pass + + @abstractmethod + async def supports_reasoning(self) -> bool: + """Whether agent provides reasoning summaries""" + pass + + @abstractmethod + async def get_capabilities(self) -> dict: + """Return agent capabilities for model listing""" + pass +``` + +**File**: `src/agents/lorem_tester.py` + +```python +class LoremTesterAgent(AgentInterface): + """ + Mock agent that implements all Responses API features with Lorem Ipsum. + + Features: + - Reasoning summaries (mock step-by-step thinking) + - Function calling (mock tool execution) + - Multi-turn conversations + - Error scenarios (triggered by keywords) + + This is the ONLY mock part - everything else is real production code. + """ + + async def generate_response( + self, + messages: list[dict], + reasoning: dict | None = None, + tools: list[dict] | None = None, + **kwargs + ) -> AsyncGenerator[OutputItem, None]: + """ + Mock PydanticAI interface - generates fake reasoning and responses. + + Real implementation would call PydanticAI here. + """ + + # 1. Yield reasoning item if requested + if reasoning and reasoning.get("summary") == "auto": + yield OutputItem( + type="reasoning", + id=f"rs_{generate_id()}", + summary=self._generate_mock_reasoning(messages) + ) + + # 2. Randomly yield function calls (30% chance) + if tools and random.random() < 0.3: + yield OutputItem( + type="function_call", + id=f"fc_{generate_id()}", + name=random.choice([t["name"] for t in tools]), + arguments=self._generate_mock_args() + ) + + # 3. Yield final message + yield OutputItem( + type="message", + id=f"msg_{generate_id()}", + role="assistant", + content=[{ + "type": "output_text", + "text": self._generate_lorem_ipsum() + }] + ) + + def _generate_mock_reasoning(self, messages: list[dict]) -> list[str]: + """Generate fake reasoning steps""" + return [ + "Analyzing the user's request and context...", + "Considering available information and constraints...", + "Formulating a comprehensive response strategy...", + "Selecting appropriate lorem ipsum content..." + ] +``` + +**File**: `src/agents/tatlock.py` + +```python +class TatlockAgent(AgentInterface): + """Placeholder for future real agent with PydanticAI""" + + async def generate_response(self, messages, **kwargs): + """Minimal placeholder implementation""" + yield OutputItem( + type="message", + id=f"msg_{generate_id()}", + role="assistant", + content=[{ + "type": "output_text", + "text": "Tatlock agent not yet implemented" + }] + ) + + async def supports_tools(self) -> bool: + return False # Not yet + + async def supports_reasoning(self) -> bool: + return False # Not yet +``` + +### 1.2 Model Registry + +**File**: `src/agents/registry.py` + +```python +from src.agents.lorem_tester import LoremTesterAgent +from src.agents.tatlock import TatlockAgent + +class ModelRegistry: + """Central registry for all available models""" + + MODELS = { + "lorem-tester": { + "agent_class": LoremTesterAgent, + "capabilities": { + "streaming": True, + "reasoning": True, + "tools": True, + "vision": False, + "audio": False + }, + "description": "Testing agent with mock Responses API features", + "created": 1733529600, # 2025-12-06 + }, + "tatlock": { + "agent_class": TatlockAgent, + "capabilities": { + "streaming": True, + "reasoning": False, # Not yet + "tools": False, # Not yet + "vision": False, + "audio": False + }, + "description": "Tatlock reasoning agent (placeholder)", + "created": 1733529600, + } + } + + @classmethod + def get_agent(cls, model_id: str) -> AgentInterface: + """Instantiate agent for given model""" + if model_id not in cls.MODELS: + raise ValueError(f"Model {model_id} not found") + + agent_class = cls.MODELS[model_id]["agent_class"] + return agent_class() + + @classmethod + def list_models(cls) -> list[dict]: + """Return all models in OpenAI format""" + return [ + { + "id": model_id, + "object": "model", + "created": config["created"], + "owned_by": "tatlock", + "capabilities": config["capabilities"], + "description": config["description"] + } + for model_id, config in cls.MODELS.items() + ] +``` + +### 1.3 Update Models Endpoint + +**File**: `src/models/service.py` + +```python +from src.agents.registry import ModelRegistry + +async def list_models() -> dict: + """List all available models""" + return { + "object": "list", + "data": ModelRegistry.list_models() + } +``` + +**Tests**: +- Model registry returns both models +- Each model has correct capabilities +- GET /v1/models returns proper format +- Invalid model selection returns 404 + +--- + +## Phase 2: Responses API Core Structure + +### 2.1 Response Schemas + +**File**: `src/responses/schemas.py` + +```python +from src.core.models import CustomBaseModel +from typing import Literal + +class OutputTextContent(CustomBaseModel): + """Text content in message output""" + type: Literal["output_text"] = "output_text" + text: str + annotations: list[dict] = [] + +class MessageOutputItem(CustomBaseModel): + """Message item in output array""" + type: Literal["message"] = "message" + id: str + role: Literal["assistant"] = "assistant" + content: list[OutputTextContent] + status: Literal["completed"] | None = "completed" + +class ReasoningOutputItem(CustomBaseModel): + """Reasoning item in output array""" + type: Literal["reasoning"] = "reasoning" + id: str + summary: list[str] + status: Literal["completed"] | None = "completed" + +class FunctionCallOutputItem(CustomBaseModel): + """Function call item in output array""" + type: Literal["function_call"] = "function_call" + id: str + name: str + arguments: str # JSON string + status: Literal["completed"] | None = "completed" + +# Union type for output items +OutputItem = MessageOutputItem | ReasoningOutputItem | FunctionCallOutputItem + +class ResponseUsage(CustomBaseModel): + """Token usage statistics""" + input_tokens: int + output_tokens: int + reasoning_tokens: int = 0 + total_tokens: int + +class ResponseRequest(CustomBaseModel): + """Responses API request""" + model: str + input: list[dict] # Previous responses or messages + reasoning: dict | None = None # {"effort": "medium", "summary": "auto"} + tools: list[dict] | None = None + metadata: dict | None = None # Custom field for conversation_id + stream: bool = False + max_output_tokens: int | None = None + temperature: float = 1.0 + stop: list[str] | None = None + +class Response(CustomBaseModel): + """Complete response object""" + id: str + object: Literal["response"] = "response" + created_at: int + model: str + status: Literal["completed", "in_progress", "failed"] + output: list[OutputItem] + usage: ResponseUsage +``` + +### 2.2 Streaming Event Schemas + +**File**: `src/responses/streaming.py` + +```python +from enum import Enum + +class StreamEventType(str, Enum): + """Streaming event types""" + REASONING_SUMMARY_DELTA = "response.reasoning_summary_text.delta" + REASONING_SUMMARY_DONE = "response.reasoning_summary_text.done" + OUTPUT_TEXT_DELTA = "response.output_text.delta" + OUTPUT_TEXT_DONE = "response.output_text.done" + FUNCTION_CALL_DELTA = "response.function_call_arguments.delta" + FUNCTION_CALL_DONE = "response.function_call_arguments.done" + RESPONSE_DONE = "response.done" + +class StreamEvent(CustomBaseModel): + """Base streaming event""" + event: StreamEventType + data: dict + +# Specific event types +class ReasoningSummaryDelta(CustomBaseModel): + event: Literal[StreamEventType.REASONING_SUMMARY_DELTA] + delta: str + +class OutputTextDelta(CustomBaseModel): + event: Literal[StreamEventType.OUTPUT_TEXT_DELTA] + delta: str + +class FunctionCallDelta(CustomBaseModel): + event: Literal[StreamEventType.FUNCTION_CALL_DELTA] + delta: str + name: str | None = None # Only in first chunk + +class ResponseDone(CustomBaseModel): + event: Literal[StreamEventType.RESPONSE_DONE] + response: Response +``` + +### 2.3 Responses Router + +**File**: `src/responses/router.py` + +```python +from fastapi import APIRouter, HTTPException +from sse_starlette.sse import EventSourceResponse +from src.responses import service +from src.responses.schemas import ResponseRequest, Response + +router = APIRouter(prefix="/responses", tags=["responses"]) + +@router.post("", response_model=Response) +async def create_response( + request: ResponseRequest, +) -> Response | EventSourceResponse: + """ + Create a response using Responses API format. + + Supports: + - Reasoning summaries + - Function calling + - Streaming + - Multi-turn conversations + + Args: + request: Response request + + Returns: + Response object or SSE stream + """ + + if request.stream: + return EventSourceResponse( + service.create_response_stream(request) + ) + + return await service.create_response(request) +``` + +### 2.4 Responses Service + +**File**: `src/responses/service.py` + +```python +import time +from src.agents.registry import ModelRegistry +from src.responses.schemas import Response, OutputItem, ResponseUsage +from src.responses.streaming import StreamingCoordinator + +async def create_response(request: ResponseRequest) -> Response: + """Create non-streaming response""" + + # Get agent for model + agent = ModelRegistry.get_agent(request.model) + + # Collect all output items from agent + output_items = [] + async for item in agent.generate_response( + messages=request.input, + reasoning=request.reasoning, + tools=request.tools, + temperature=request.temperature, + max_tokens=request.max_output_tokens, + ): + output_items.append(item) + + # Calculate token usage (real counting for production) + usage = _calculate_usage(request.input, output_items) + + return Response( + id=f"resp_{generate_id()}", + created_at=int(time.time()), + model=request.model, + status="completed", + output=output_items, + usage=usage + ) + +async def create_response_stream(request: ResponseRequest): + """Create streaming response""" + + coordinator = StreamingCoordinator() + + async for event in coordinator.stream_response(request): + yield { + "event": event.event, + "data": event.model_dump_json() + } +``` + +**Tests**: +- Non-streaming response has correct structure +- Output array contains reasoning, function, message items +- Token usage calculated correctly +- Model selection works + +--- + +## Phase 3: Conversation History with Hybrid Approach + +### 3.1 Conversation Storage + +**File**: `src/responses/history.py` + +```python +from typing import Dict, List +import hashlib + +class ConversationHistory: + """ + Real conversation history management. + + Supports hybrid approach: + - Client sends full input array (OpenAI compatible) + - Optional conversation_id in metadata for server-side grouping + - Server can augment with vector memories (future) + """ + + def __init__(self, max_turns: int = 20): + self._conversations: Dict[str, List[Response]] = {} + self._max_turns = max_turns + + async def get_conversation_id( + self, + request: ResponseRequest + ) -> str: + """ + Get or generate conversation ID. + + Priority: + 1. metadata.conversation_id if provided + 2. Generate from first message hash + """ + if request.metadata and "conversation_id" in request.metadata: + return request.metadata["conversation_id"] + + # Generate deterministic ID from first message + first_msg = str(request.input[0]) if request.input else "" + return hashlib.sha256(first_msg.encode()).hexdigest()[:16] + + async def add_response( + self, + conversation_id: str, + response: Response + ): + """Add response to conversation history""" + if conversation_id not in self._conversations: + self._conversations[conversation_id] = [] + + self._conversations[conversation_id].append(response) + + # Trim old turns + await self._trim_history(conversation_id) + + async def get_history( + self, + conversation_id: str + ) -> List[Response]: + """Retrieve conversation history""" + return self._conversations.get(conversation_id, []) + + async def _trim_history(self, conversation_id: str): + """Keep only recent turns within limit""" + if len(self._conversations[conversation_id]) > self._max_turns: + self._conversations[conversation_id] = ( + self._conversations[conversation_id][-self._max_turns:] + ) + + # Future: Integration point for Qdrant vector memory + async def get_relevant_memories( + self, + conversation_id: str, + query: str + ) -> List[dict]: + """ + Retrieve relevant memories from vector store. + + TODO: Integrate Qdrant for semantic search + """ + return [] # Placeholder +``` + +### 3.2 Context Window Management + +**File**: `src/responses/context.py` + +```python +class ContextWindow: + """Manage token limits and context trimming""" + + def __init__(self, max_tokens: int = 4096): + self.max_tokens = max_tokens + + async def count_tokens(self, items: list) -> int: + """ + Real token counting. + + For now: approximate by character count + Future: Use tiktoken or similar + """ + total_chars = sum( + len(str(item)) + for item in items + ) + return total_chars // 4 # Rough approximation + + async def trim_to_fit( + self, + items: list, + reserve_tokens: int = 512 + ) -> list: + """Trim old items to fit context window""" + available = self.max_tokens - reserve_tokens + + # Start from most recent, work backwards + kept_items = [] + current_tokens = 0 + + for item in reversed(items): + item_tokens = await self.count_tokens([item]) + if current_tokens + item_tokens <= available: + kept_items.insert(0, item) + current_tokens += item_tokens + else: + break + + return kept_items +``` + +**Tests**: +- Conversation ID generation (with and without metadata) +- History storage and retrieval +- Context window trimming +- Token counting + +--- + +## Phase 4: Streaming Coordinator + +### 4.1 Streaming Implementation + +**File**: `src/responses/streaming.py` + +```python +class StreamingCoordinator: + """ + Real production streaming logic. + + Coordinates complex streaming: + 1. Reasoning summary chunks + 2. Function call arguments + 3. Output text chunks + 4. Error handling + 5. Final response event + """ + + async def stream_response( + self, + request: ResponseRequest + ) -> AsyncGenerator[StreamEvent, None]: + """Coordinate streaming from agent""" + + agent = ModelRegistry.get_agent(request.model) + + output_items = [] + current_reasoning = [] + current_function = None + current_message = "" + + try: + async for item in agent.generate_response( + messages=request.input, + reasoning=request.reasoning, + tools=request.tools, + **request.model_dump(exclude={"input", "reasoning", "tools"}) + ): + output_items.append(item) + + # Stream based on item type + if item.type == "reasoning": + # Stream reasoning summary + for step in item.summary: + yield ReasoningSummaryDelta( + event=StreamEventType.REASONING_SUMMARY_DELTA, + delta=step + ) + + yield StreamEvent( + event=StreamEventType.REASONING_SUMMARY_DONE, + data={} + ) + + elif item.type == "function_call": + # Stream function call arguments + # First chunk includes name + yield FunctionCallDelta( + event=StreamEventType.FUNCTION_CALL_DELTA, + name=item.name, + delta="" + ) + + # Stream arguments in chunks + args = item.arguments + chunk_size = 20 + for i in range(0, len(args), chunk_size): + yield FunctionCallDelta( + event=StreamEventType.FUNCTION_CALL_DELTA, + delta=args[i:i+chunk_size] + ) + + yield StreamEvent( + event=StreamEventType.FUNCTION_CALL_DONE, + data={} + ) + + elif item.type == "message": + # Stream output text + text = item.content[0].text + words = text.split() + + for word in words: + yield OutputTextDelta( + event=StreamEventType.OUTPUT_TEXT_DELTA, + delta=f"{word} " + ) + + # Simulate typing delay (for testing) + await asyncio.sleep(0.05) + + yield StreamEvent( + event=StreamEventType.OUTPUT_TEXT_DONE, + data={} + ) + + # Final response.done event with complete response + usage = _calculate_usage(request.input, output_items) + + final_response = Response( + id=f"resp_{generate_id()}", + created_at=int(time.time()), + model=request.model, + status="completed", + output=output_items, + usage=usage + ) + + yield ResponseDone( + event=StreamEventType.RESPONSE_DONE, + response=final_response + ) + + except Exception as e: + # Stream error + yield self._create_error_event(e) +``` + +**Tests**: +- Reasoning summary streaming +- Function call argument streaming +- Output text streaming word-by-word +- response.done event with complete response +- Error mid-stream handling + +--- + +## Phase 5: Error Handling + +### 5.1 Error Scenarios in Lorem Tester + +**File**: `src/agents/lorem_tester.py` (additions) + +```python +class LoremTesterAgent: + """ + Trigger errors based on keywords in input for testing: + - "trigger_rate_limit" → 429 rate limit error + - "trigger_timeout" → timeout after 121s + - "trigger_invalid_tool" → invalid tool call error + - "trigger_context_overflow" → context length exceeded + - "trigger_partial_failure" → error mid-stream + """ + + async def generate_response(self, messages, **kwargs): + # Check for error triggers + last_message = str(messages[-1]) if messages else "" + + if "trigger_rate_limit" in last_message: + raise RateLimitError("Rate limit exceeded (mock)") + + if "trigger_timeout" in last_message: + await asyncio.sleep(121) # Exceed default timeout + + if "trigger_context_overflow" in last_message: + raise ContextLengthError("Context length exceeded (mock)") + + if "trigger_partial_failure" in last_message: + # Yield some items then fail + yield OutputItem(...) + yield OutputItem(...) + raise APIError("Simulated mid-stream failure") + + # Normal flow... +``` + +### 5.2 Error Responses + +**File**: `src/responses/schemas.py` (additions) + +```python +class ErrorResponse(CustomBaseModel): + """Error response format""" + error: dict + +class ErrorDetail(CustomBaseModel): + type: str + message: str + code: int | None = None +``` + +**File**: `src/responses/service.py` (additions) + +```python +async def create_response_stream(request): + """Create streaming response with error handling""" + + coordinator = StreamingCoordinator() + + try: + async for event in coordinator.stream_response(request): + yield {"event": event.event, "data": event.model_dump_json()} + + except RateLimitError as e: + yield { + "event": "error", + "data": json.dumps({ + "type": "rate_limit_exceeded", + "message": str(e), + "code": 429 + }) + } + + except ContextLengthError as e: + yield { + "event": "error", + "data": json.dumps({ + "type": "context_length_exceeded", + "message": str(e), + "code": 400 + }) + } + + except Exception as e: + yield { + "event": "error", + "data": json.dumps({ + "type": "internal_error", + "message": str(e), + "code": 500 + }) + } +``` + +**Tests**: +- All error trigger keywords +- Mid-stream failure handling +- Error event format +- Graceful degradation + +--- + +## Phase 6: Advanced Features + +### 6.1 Stop Sequences + +```python +class StreamingCoordinator: + async def _check_stop_sequence( + self, + accumulated_text: str, + stop_sequences: list[str] + ) -> bool: + """Check if any stop sequence encountered""" + return any(seq in accumulated_text for seq in stop_sequences) +``` + +### 6.2 Token Limits + +```python +class StreamingCoordinator: + async def _enforce_max_tokens( + self, + token_count: int, + max_tokens: int | None + ) -> bool: + """Stop streaming if max tokens reached""" + if max_tokens and token_count >= max_tokens: + return True + return False +``` + +### 6.3 Parameter Validation + +**File**: `src/responses/schemas.py` (additions) + +```python +from pydantic import Field, field_validator + +class ResponseRequest(CustomBaseModel): + # ... existing fields + + temperature: float = Field( + default=1.0, + ge=0.0, + le=2.0, + description="Sampling temperature" + ) + + @field_validator('reasoning') + def validate_reasoning(cls, v): + if v is not None: + if 'effort' in v: + allowed = ['none', 'minimal', 'low', 'medium', 'high', 'xhigh'] + if v['effort'] not in allowed: + raise ValueError(f"reasoning.effort must be one of {allowed}") + return v +``` + +**Tests**: +- Stop sequence detection +- Max token enforcement +- Temperature validation +- Reasoning effort validation + +--- + +## Phase 7: Testing & Integration + +### 7.1 Test Suite Structure + +``` +tests/ +├── responses/ +│ ├── test_router.py # API endpoint tests +│ ├── test_service.py # Service layer tests +│ ├── test_streaming.py # Streaming tests +│ └── test_history.py # Conversation history tests +├── agents/ +│ ├── test_lorem_tester.py # Lorem Tester agent tests +│ ├── test_registry.py # Model registry tests +│ └── test_interface.py # Agent interface tests +└── integration/ + └── test_openwebui.py # Open WebUI integration tests +``` + +### 7.2 Open WebUI Integration Testing + +**Manual Test Checklist**: +- [ ] Both models appear in model dropdown +- [ ] Lorem Tester shows thinking/reasoning bubbles +- [ ] Reasoning displays separately from response +- [ ] Tool calls show in visual flow +- [ ] Tool execution progress visible +- [ ] Final answer clearly separated +- [ ] Conversation history maintained across turns +- [ ] Error messages display gracefully +- [ ] Streaming smooth and responsive +- [ ] All trigger keywords work for testing + +### 7.3 Documentation + +**File**: `OPENWEBUI_INTEGRATION.md` + +```markdown +# Open WebUI Integration Guide + +## Connection Setup + +1. In Open WebUI, go to Settings → Connections +2. Add new OpenAI-compatible API: + - Base URL: `http://localhost:8000/v1` + - API Key: (leave empty for now) + - API Type: OpenAI + +## Available Models + +### lorem-tester +Testing model with all Responses API features: +- Reasoning summaries (thinking bubbles) +- Function calling (mock tools) +- Error scenarios (trigger keywords) + +**Test Keywords**: +- Include "trigger_rate_limit" to test rate limit errors +- Include "trigger_timeout" to test timeout handling +- Include "trigger_partial_failure" to test mid-stream errors + +### tatlock +Placeholder for production agent (minimal implementation) + +## Visual Features + +Lorem Tester demonstrates: +- **Thinking Bubbles**: Separate reasoning display +- **Tool Progress**: Visual flow for function calls +- **Final Answer**: Clear separation from thinking +- **Error Handling**: Graceful error messages + +## Conversation History + +The API supports hybrid conversation tracking: +- Send full conversation in `input` array (OpenAI compatible) +- Optional: Include `metadata.conversation_id` for server-side grouping +- Server maintains history for context window management +- Future: Vector memory integration via conversation_id +``` + +--- + +## Phase 8: Cleanup Old API + +**File**: `CLEANUP_TODO.md` + +```markdown +# Cleanup Todo: Remove Chat Completions API + +After Responses API is fully tested and integrated, remove old API: + +## Code Removal + +- [ ] Delete `src/chat/` directory entirely + - router.py + - service.py + - schemas.py + - constants.py + - dependencies.py + +- [ ] Remove chat imports from `src/main.py` + - Remove chat router registration + - Remove `/v1/chat/completions` route + +- [ ] Update core router (`src/core/router.py`) + - Remove any chat-specific health checks + +## Documentation Updates + +- [ ] README.md + - Remove chat completions examples + - Update all examples to use Responses API + - Update architecture diagrams + - Update feature list + +- [ ] AGENTS.md + - Remove chat completions references + - Update all code examples to Responses API + - Update streaming patterns + +- [ ] CHANGELOG.md + - Add deprecation notice for v0.1.0 + - Document Responses API adoption + +## Test Cleanup + +- [ ] Delete `tests/chat/` directory +- [ ] Remove chat-related fixtures from `conftest.py` +- [ ] Update integration tests + +## Configuration Cleanup + +- [ ] Review `src/core/config.py` for chat-specific settings +- [ ] Remove unused constants + +## Verification + +- [ ] All tests pass without chat code +- [ ] Open WebUI integration still works +- [ ] Documentation is consistent +- [ ] No broken imports +``` + +--- + +## Project Structure (Final) + +``` +src/ +├── agents/ # Agent implementations +│ ├── base.py # AgentInterface abstraction +│ ├── lorem_tester.py # Mock agent with all features +│ ├── tatlock.py # Placeholder real agent +│ └── registry.py # Model registry +├── responses/ # Responses API domain +│ ├── router.py # POST /v1/responses +│ ├── schemas.py # Request/response/output items +│ ├── service.py # Response generation logic +│ ├── streaming.py # Streaming coordinator +│ ├── history.py # Conversation history +│ └── context.py # Context window management +├── models/ +│ ├── router.py # GET /v1/models +│ ├── schemas.py # Model schemas +│ └── service.py # Model listing +├── core/ +│ ├── config.py # Global configuration +│ ├── models.py # Custom Pydantic base +│ ├── exceptions.py # Custom exceptions +│ └── router.py # Health/root endpoints +└── main.py # Application factory + +tests/ +├── responses/ # Responses API tests +├── agents/ # Agent tests +├── models/ # Model tests +└── integration/ # Open WebUI integration + +docs/ +├── OPENWEBUI_INTEGRATION.md # Integration guide +├── CLEANUP_TODO.md # Old API removal checklist +└── IMPLEMENTATION_PLAN.md # This file +``` + +--- + +## Implementation Order + +1. **Phase 1** - Agent interface, registry, model selection +2. **Phase 2** - Responses API schemas and core service +3. **Phase 3** - Conversation history with hybrid approach +4. **Phase 4** - Streaming coordinator with all event types +5. **Phase 5** - Error handling and trigger scenarios +6. **Phase 6** - Advanced features (stop, tokens, validation) +7. **Phase 7** - Testing and Open WebUI integration +8. **Phase 8** - Clean out old Chat Completions API + +Each phase is independently testable with Open WebUI. + +--- + +## Questions Resolved + +1. ✅ **Conversation tracking**: Hybrid approach with optional conversation_id +2. ✅ **Persistence**: In-memory for now, Qdrant integration point ready +3. ✅ **Tool execution**: Mock only for Lorem Tester +4. ✅ **API format**: Responses API only (not Chat Completions) + +--- + +## Next Steps + +Ready to begin Phase 1: Agent Interface & Multi-Model Support