refactor(core-ai): comprehensive cleanup - PydanticAI only architecture

Remove all obsolete agent implementations and framework references.
Keep only PydanticAI (primary) and SimpleLiteLLM (fallback).

This cleanup eliminates confusion between multiple frameworks that were
tried during development (LangChain, LangGraph, ADK, OllamaNative) and
establishes PydanticAI as the single agent framework going forward.

BREAKING CHANGES:
- Removed OllamaNativeAgent - use PydanticAgent instead
- Removed /test/ollama-tools diagnostic endpoint
- Default /v1/chat/completions now uses PydanticAgent

Files Deleted (32 total):
- Obsolete agents: ollama_native_agent.py
- Diagnostic files: ARCHITECTURE.md, DIAGNOSTIC_RESULTS.md, PHASE*.md
- Legacy tools: src/tools.py
- Test files: test_ai_flow_quality.py, test_02/03 (diagnostic layers)
- Documentation: ADK_Ollama_Research.md, agent-flow-diagrams.md
- Session docs: 3 files with LangChain/LangGraph implementations
- Plans: 5 completed plans about obsolete frameworks
- Migration docs: MIGRATION_PLAN_LANGCHAIN_TO_ADK.md

Files Modified (8 total):
- main.py: Refactored to PydanticAI only (305 lines vs 457 before)
- agents/__init__.py: Removed OllamaNativeAgent exports
- README.md: Complete rewrite for PydanticAI architecture
- prompts.py: Updated for PydanticAI (infrastructure tool guidance)
- STATUS.md: Updated to v0.11.0-pydantic-ai
- CHANGELOG.md: Added v0.11.0 entry documenting cleanup
- plans/active/*.md: Updated to reference PydanticAI

Current Architecture:
- Framework: PydanticAI with native Ollama SDK
- Agents: PydanticAgent (primary) + SimpleLiteLLMAgent (fallback)
- Model: mistral-nemo:latest
- Tools: 6 core + 28+ OpenAPI-discovered
- Memory: 3-tier system with Qdrant
- VRAM: ~4-6GB

Lines Removed: ~3000+ lines of obsolete code

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
This commit is contained in:
2025-12-03 14:24:52 +01:00
co-authored by Claude
parent 96492cb1ed
commit 66f6e54fc3
32 changed files with 1078 additions and 10376 deletions
File diff suppressed because it is too large Load Diff
+126 -102
View File
@@ -37,7 +37,7 @@ Instead of configuring functions in Open WebUI (or any other UI), the Core API b
│
↓
┌─────────────────────────────────────────────────────────────┐
│ Agent Orchestrator (LangGraph) │
│ Agent Orchestrator (PydanticAI) │
│ ┌──────────────────────────────────────────────┐ │
│ │ Reasoning Loop: │ │
│ │ 1. Analyze user intent │ │
@@ -63,51 +63,59 @@ Instead of configuring functions in Open WebUI (or any other UI), the Core API b
## Implementation Options
### Option 1: LangGraph (Recommended)
### Option 1: PydanticAI (Current Implementation)
**Pros:**
- Built-in agent loops and tool calling
- State management for multi-step reasoning
- Streaming support for intermediate steps
- Well-documented patterns
- Active development
- Type-safe tool definitions with Pydantic models
- Built-in streaming support with structured output
- Native Ollama integration via HTTP API
- Lightweight and minimal dependencies
- Clear separation of concerns with dependency injection
- Excellent debugging with structured validation
**Cons:**
- Additional dependency (~50MB)
- Learning curve for LangGraph concepts
- Some overhead vs custom implementation
- Relatively new framework (less established patterns)
- Manual agent loop implementation required
- Less built-in state management compared to stateful frameworks
**Example flow:**
```python
from langgraph.prebuilt import create_react_agent
from langchain_core.tools import tool
from pydantic_ai import Agent, RunContext
from pydantic import BaseModel
@tool
def deploy_service(service_name: str, compose_yaml: str) -> str:
"""Deploy a containerized service via Portainer"""
# Use existing infrastructure controller
return portainer_client.deploy_stack(...)
class DeployServiceParams(BaseModel):
service_name: str
compose_yaml: str
@tool
def web_search(query: str) -> str:
"""Search the web and extract content"""
# Use existing web scraper
return scraper.scrape(...)
class WebSearchParams(BaseModel):
query: str
agent = create_react_agent(
model=ChatOllama(model="gemma:7b"),
tools=[deploy_service, web_search, ...],
state_modifier="You are a homelab infrastructure assistant..."
agent = Agent(
model="ollama:mistral-tools:7b",
system_prompt="You are a homelab infrastructure assistant...",
result_type=str
)
@agent.tool
async def deploy_service(ctx: RunContext[None], params: DeployServiceParams) -> str:
"""Deploy a containerized service via Portainer"""
return await portainer_client.deploy_stack(
params.service_name,
params.compose_yaml
)
@agent.tool
async def web_search(ctx: RunContext[None], params: WebSearchParams) -> str:
"""Search the web and extract content"""
return await scraper.scrape(params.query)
# Streaming with reasoning
for chunk in agent.stream({"messages": [user_message]}):
if "thinking" in chunk:
yield f"data: {json.dumps({'reasoning': chunk['thinking']})}\n\n"
if "tool_calls" in chunk:
yield f"data: {json.dumps({'tool': chunk['tool_calls'][0]['name']})}\n\n"
if "response" in chunk:
yield f"data: {json.dumps({'content': chunk['response']})}\n\n"
async with agent.run_stream(user_message) as stream:
async for chunk in stream.stream_text():
if chunk.type == "tool_call":
yield f"data: {json.dumps({'tool': chunk.tool_name})}\n\n"
elif chunk.type == "text":
yield f"data: {json.dumps({'content': chunk.content})}\n\n"
```
### Option 2: Custom Agent Loop
@@ -122,6 +130,7 @@ for chunk in agent.stream({"messages": [user_message]}):
- More code to maintain
- Need to implement tool calling protocol
- Reinventing some wheels
- Manual type validation
**Example flow:**
```python
@@ -147,26 +156,27 @@ class UnifiedAgent:
yield {"type": "content", "content": response}
```
### Option 3: Hybrid (LangChain Tools + Custom Orchestration)
### Option 3: Hybrid (PydanticAI + Custom Extensions)
Use LangChain's tool framework but custom agent loop:
- Leverage `@tool` decorator for easy tool definitions
- Custom routing logic for model selection
- Manual streaming control
Use PydanticAI's agent framework with custom enhancements:
- Leverage type-safe tool definitions
- Add custom routing logic for multi-model selection
- Enhanced streaming control for reasoning output
- Custom dependency injection for context management
## Recommended Approach: LangGraph with Custom Extensions
## Recommended Approach: PydanticAI with Custom Extensions
**Phase 1: Core Agent (Week 1)**
- Set up LangGraph agent with basic tools
- Set up PydanticAI agent with basic tools
- Implement streaming with reasoning output
- Wire up existing infrastructure tools
- Test with simple queries
**Phase 2: Advanced Routing (Week 2)**
- Multi-model routing (small for simple, large for complex)
- Parallel tool execution
- Error handling and retries
- Context management
- Parallel tool execution via async tools
- Error handling and retries with custom logic
- Context management using RunContext
**Phase 3: Multi-Modal (Week 3)**
- Image analysis (if needed)
@@ -179,23 +189,34 @@ Use LangChain's tool framework but custom agent loop:
### Tier 1: Infrastructure Tools (Existing)
```python
@tool
async def list_services() -> List[Dict]:
from pydantic_ai import Agent, RunContext
from pydantic import BaseModel
class ListServicesResult(BaseModel):
services: List[Dict]
class DeployServiceParams(BaseModel):
name: str
compose: str
@agent.tool
async def list_services(ctx: RunContext[None]) -> ListServicesResult:
"""List all running Docker services"""
return await portainer_client.list_containers()
services = await portainer_client.list_containers()
return ListServicesResult(services=services)
@tool
async def deploy_service(name: str, compose: str) -> str:
@agent.tool
async def deploy_service(ctx: RunContext[None], params: DeployServiceParams) -> str:
"""Deploy a new service from Docker Compose YAML"""
return await portainer_client.deploy_stack(name, compose)
return await portainer_client.deploy_stack(params.name, params.compose)
@tool
async def create_proxy(domain: str, target: str) -> str:
@agent.tool
async def create_proxy(ctx: RunContext[None], domain: str, target: str) -> str:
"""Create Nginx reverse proxy for a service"""
return await npm_client.create_proxy_host(domain, target)
@tool
async def check_service_health(service: str) -> Dict:
@agent.tool
async def check_service_health(ctx: RunContext[None], service: str) -> Dict:
"""Check if a service is healthy"""
return await kuma_client.get_monitor_status(service)
```
@@ -203,18 +224,18 @@ async def check_service_health(service: str) -> Dict:
### Tier 2: Knowledge Tools
```python
@tool
async def web_search(query: str) -> str:
@agent.tool
async def web_search(ctx: RunContext[None], query: str) -> str:
"""Search the web and extract main content"""
return await scraper.scrape_url(query)
@tool
async def query_memory(question: str) -> List[str]:
@agent.tool
async def query_memory(ctx: RunContext[None], question: str) -> List[str]:
"""Search conversation history for relevant context"""
return await memory.semantic_search(question)
@tool
async def read_documentation(topic: str) -> str:
@agent.tool
async def read_documentation(ctx: RunContext[None], topic: str) -> str:
"""Read project documentation"""
docs_path = f"/docs/{topic}.md"
return read_file(docs_path)
@@ -223,14 +244,14 @@ async def read_documentation(topic: str) -> str:
### Tier 3: Execution Tools (Future)
```python
@tool
async def execute_python(code: str) -> str:
@agent.tool
async def execute_python(ctx: RunContext[None], code: str) -> str:
"""Execute Python code in sandbox"""
# Future: Integrate code interpreter
pass
@tool
async def query_database(sql: str) -> List[Dict]:
@agent.tool
async def query_database(ctx: RunContext[None], sql: str) -> List[Dict]:
"""Query PostgreSQL database"""
# Future: Safe SQL execution
pass
@@ -246,7 +267,7 @@ async def query_database(sql: str) -> List[Dict]:
"type": "thinking", # or "tool_call", "content", "error"
"content": "Searching the web for nginx configuration...",
"tool": "web_search", # optional, if type is tool_call
"model": "gemma:7b" # optional, which model is being used
"model": "mistral-tools:7b" # optional, which model is being used
}
# Example stream
@@ -281,53 +302,56 @@ Open WebUI already supports streaming, we just need to format it correctly:
### Intent-Based Routing
```python
from pydantic_ai import Agent
class ModelRouter:
MODELS = {
"simple": "gemma:2b", # Fast, <100 tokens
"general": "gemma:7b", # Balanced
"expert": "mistral:7b", # Complex reasoning
"code": "codestral:latest" # Code tasks
"simple": "ollama:gemma:2b", # Fast, <100 tokens
"general": "ollama:gemma:7b", # Balanced
"expert": "ollama:mistral:7b", # Complex reasoning
"code": "ollama:codestral:latest" # Code tasks
}
async def select_model(self, message: str, context: str) -> str:
# Use lightweight model for routing decision
prompt = f"""Analyze this request and categorize:
routing_agent = Agent(
model="ollama:gemma:2b",
result_type=str,
system_prompt="""Analyze this request and categorize:
User: {message}
Context: {context}
Categories:
- simple: Greetings, basic facts, short answers
- general: Normal conversation, explanations
- expert: Complex reasoning, multi-step problems
- code: Programming tasks, debugging
Categories:
- simple: Greetings, basic facts, short answers
- general: Normal conversation, explanations
- expert: Complex reasoning, multi-step problems
- code: Programming tasks, debugging
Return ONLY the category.
"""
)
Return ONLY the category.
"""
category = await ollama.generate(model="gemma:2b", prompt=prompt)
return self.MODELS[category.strip()]
result = await routing_agent.run(f"User: {message}\nContext: {context}")
return self.MODELS[result.data.strip()]
```
## Next Steps
1. **Prototype LangGraph agent** (2-3 hours)
- Basic agent with 2-3 tools
- Streaming with thinking output
- Test with Open WebUI
1. **Enhance PydanticAI agent** (2-3 hours)
- Add more infrastructure tools
- Improve streaming with reasoning output
- Test with complex queries
2. **Integrate existing tools** (3-4 hours)
- Wrap infrastructure controller as tools
- Wrap web scraper as tool
- Test tool calling
2. **Integrate remaining tools** (3-4 hours)
- Migrate all infrastructure controller tools
- Add web scraper tool improvements
- Test multi-tool workflows
3. **Model routing** (2 hours)
- Implement intent analysis
- Add model selection logic
- Test performance
3. **Model routing enhancements** (2 hours)
- Refine intent analysis
- Add model selection metrics
- Test performance improvements
4. **Production deployment** (2 hours)
- Error handling
4. **Production hardening** (2 hours)
- Enhanced error handling
- Rate limiting
- Logging and monitoring
- Update API documentation
@@ -391,20 +415,20 @@ User: "How do I configure Headscale?"
## Technology Stack
- **Agent Framework:** LangGraph 0.2.x
- **LLM Integration:** LangChain-Ollama
- **Tool Framework:** LangChain Tools
- **Agent Framework:** PydanticAI
- **LLM Integration:** Native Ollama SDK (HTTP API)
- **Tool Framework:** PydanticAI Tools with Pydantic validation
- **Streaming:** SSE (Server-Sent Events)
- **State Management:** LangGraph StateGraph
- **State Management:** RunContext dependency injection
- **Memory:** Existing Qdrant integration
## Risk Mitigation
**Risk:** LangGraph adds complexity
- **Mitigation:** Start simple, add features incrementally
**Risk:** PydanticAI is relatively new
- **Mitigation:** Strong typing provides safety, active development community
**Risk:** Tool calling may be slow
- **Mitigation:** Parallel execution, caching, optimized tools
- **Mitigation:** Async tools enable parallel execution, caching, optimized tools
**Risk:** Reasoning output may be verbose
- **Mitigation:** Configurable verbosity, collapsible UI elements