refactor(core-ai): comprehensive cleanup - PydanticAI only architecture
Remove all obsolete agent implementations and framework references. Keep only PydanticAI (primary) and SimpleLiteLLM (fallback). This cleanup eliminates confusion between multiple frameworks that were tried during development (LangChain, LangGraph, ADK, OllamaNative) and establishes PydanticAI as the single agent framework going forward. BREAKING CHANGES: - Removed OllamaNativeAgent - use PydanticAgent instead - Removed /test/ollama-tools diagnostic endpoint - Default /v1/chat/completions now uses PydanticAgent Files Deleted (32 total): - Obsolete agents: ollama_native_agent.py - Diagnostic files: ARCHITECTURE.md, DIAGNOSTIC_RESULTS.md, PHASE*.md - Legacy tools: src/tools.py - Test files: test_ai_flow_quality.py, test_02/03 (diagnostic layers) - Documentation: ADK_Ollama_Research.md, agent-flow-diagrams.md - Session docs: 3 files with LangChain/LangGraph implementations - Plans: 5 completed plans about obsolete frameworks - Migration docs: MIGRATION_PLAN_LANGCHAIN_TO_ADK.md Files Modified (8 total): - main.py: Refactored to PydanticAI only (305 lines vs 457 before) - agents/__init__.py: Removed OllamaNativeAgent exports - README.md: Complete rewrite for PydanticAI architecture - prompts.py: Updated for PydanticAI (infrastructure tool guidance) - STATUS.md: Updated to v0.11.0-pydantic-ai - CHANGELOG.md: Added v0.11.0 entry documenting cleanup - plans/active/*.md: Updated to reference PydanticAI Current Architecture: - Framework: PydanticAI with native Ollama SDK - Agents: PydanticAgent (primary) + SimpleLiteLLMAgent (fallback) - Model: mistral-nemo:latest - Tools: 6 core + 28+ OpenAPI-discovered - Memory: 3-tier system with Qdrant - VRAM: ~4-6GB Lines Removed: ~3000+ lines of obsolete code 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>
This commit is contained in:
File diff suppressed because it is too large
Load Diff
@@ -37,7 +37,7 @@ Instead of configuring functions in Open WebUI (or any other UI), the Core API b
|
||||
│
|
||||
↓
|
||||
┌─────────────────────────────────────────────────────────────┐
|
||||
│ Agent Orchestrator (LangGraph) │
|
||||
│ Agent Orchestrator (PydanticAI) │
|
||||
│ ┌──────────────────────────────────────────────┐ │
|
||||
│ │ Reasoning Loop: │ │
|
||||
│ │ 1. Analyze user intent │ │
|
||||
@@ -63,51 +63,59 @@ Instead of configuring functions in Open WebUI (or any other UI), the Core API b
|
||||
|
||||
## Implementation Options
|
||||
|
||||
### Option 1: LangGraph (Recommended)
|
||||
### Option 1: PydanticAI (Current Implementation)
|
||||
|
||||
**Pros:**
|
||||
- Built-in agent loops and tool calling
|
||||
- State management for multi-step reasoning
|
||||
- Streaming support for intermediate steps
|
||||
- Well-documented patterns
|
||||
- Active development
|
||||
- Type-safe tool definitions with Pydantic models
|
||||
- Built-in streaming support with structured output
|
||||
- Native Ollama integration via HTTP API
|
||||
- Lightweight and minimal dependencies
|
||||
- Clear separation of concerns with dependency injection
|
||||
- Excellent debugging with structured validation
|
||||
|
||||
**Cons:**
|
||||
- Additional dependency (~50MB)
|
||||
- Learning curve for LangGraph concepts
|
||||
- Some overhead vs custom implementation
|
||||
- Relatively new framework (less established patterns)
|
||||
- Manual agent loop implementation required
|
||||
- Less built-in state management compared to stateful frameworks
|
||||
|
||||
**Example flow:**
|
||||
```python
|
||||
from langgraph.prebuilt import create_react_agent
|
||||
from langchain_core.tools import tool
|
||||
from pydantic_ai import Agent, RunContext
|
||||
from pydantic import BaseModel
|
||||
|
||||
@tool
|
||||
def deploy_service(service_name: str, compose_yaml: str) -> str:
|
||||
"""Deploy a containerized service via Portainer"""
|
||||
# Use existing infrastructure controller
|
||||
return portainer_client.deploy_stack(...)
|
||||
class DeployServiceParams(BaseModel):
|
||||
service_name: str
|
||||
compose_yaml: str
|
||||
|
||||
@tool
|
||||
def web_search(query: str) -> str:
|
||||
"""Search the web and extract content"""
|
||||
# Use existing web scraper
|
||||
return scraper.scrape(...)
|
||||
class WebSearchParams(BaseModel):
|
||||
query: str
|
||||
|
||||
agent = create_react_agent(
|
||||
model=ChatOllama(model="gemma:7b"),
|
||||
tools=[deploy_service, web_search, ...],
|
||||
state_modifier="You are a homelab infrastructure assistant..."
|
||||
agent = Agent(
|
||||
model="ollama:mistral-tools:7b",
|
||||
system_prompt="You are a homelab infrastructure assistant...",
|
||||
result_type=str
|
||||
)
|
||||
|
||||
@agent.tool
|
||||
async def deploy_service(ctx: RunContext[None], params: DeployServiceParams) -> str:
|
||||
"""Deploy a containerized service via Portainer"""
|
||||
return await portainer_client.deploy_stack(
|
||||
params.service_name,
|
||||
params.compose_yaml
|
||||
)
|
||||
|
||||
@agent.tool
|
||||
async def web_search(ctx: RunContext[None], params: WebSearchParams) -> str:
|
||||
"""Search the web and extract content"""
|
||||
return await scraper.scrape(params.query)
|
||||
|
||||
# Streaming with reasoning
|
||||
for chunk in agent.stream({"messages": [user_message]}):
|
||||
if "thinking" in chunk:
|
||||
yield f"data: {json.dumps({'reasoning': chunk['thinking']})}\n\n"
|
||||
if "tool_calls" in chunk:
|
||||
yield f"data: {json.dumps({'tool': chunk['tool_calls'][0]['name']})}\n\n"
|
||||
if "response" in chunk:
|
||||
yield f"data: {json.dumps({'content': chunk['response']})}\n\n"
|
||||
async with agent.run_stream(user_message) as stream:
|
||||
async for chunk in stream.stream_text():
|
||||
if chunk.type == "tool_call":
|
||||
yield f"data: {json.dumps({'tool': chunk.tool_name})}\n\n"
|
||||
elif chunk.type == "text":
|
||||
yield f"data: {json.dumps({'content': chunk.content})}\n\n"
|
||||
```
|
||||
|
||||
### Option 2: Custom Agent Loop
|
||||
@@ -122,6 +130,7 @@ for chunk in agent.stream({"messages": [user_message]}):
|
||||
- More code to maintain
|
||||
- Need to implement tool calling protocol
|
||||
- Reinventing some wheels
|
||||
- Manual type validation
|
||||
|
||||
**Example flow:**
|
||||
```python
|
||||
@@ -147,26 +156,27 @@ class UnifiedAgent:
|
||||
yield {"type": "content", "content": response}
|
||||
```
|
||||
|
||||
### Option 3: Hybrid (LangChain Tools + Custom Orchestration)
|
||||
### Option 3: Hybrid (PydanticAI + Custom Extensions)
|
||||
|
||||
Use LangChain's tool framework but custom agent loop:
|
||||
- Leverage `@tool` decorator for easy tool definitions
|
||||
- Custom routing logic for model selection
|
||||
- Manual streaming control
|
||||
Use PydanticAI's agent framework with custom enhancements:
|
||||
- Leverage type-safe tool definitions
|
||||
- Add custom routing logic for multi-model selection
|
||||
- Enhanced streaming control for reasoning output
|
||||
- Custom dependency injection for context management
|
||||
|
||||
## Recommended Approach: LangGraph with Custom Extensions
|
||||
## Recommended Approach: PydanticAI with Custom Extensions
|
||||
|
||||
**Phase 1: Core Agent (Week 1)**
|
||||
- Set up LangGraph agent with basic tools
|
||||
- Set up PydanticAI agent with basic tools
|
||||
- Implement streaming with reasoning output
|
||||
- Wire up existing infrastructure tools
|
||||
- Test with simple queries
|
||||
|
||||
**Phase 2: Advanced Routing (Week 2)**
|
||||
- Multi-model routing (small for simple, large for complex)
|
||||
- Parallel tool execution
|
||||
- Error handling and retries
|
||||
- Context management
|
||||
- Parallel tool execution via async tools
|
||||
- Error handling and retries with custom logic
|
||||
- Context management using RunContext
|
||||
|
||||
**Phase 3: Multi-Modal (Week 3)**
|
||||
- Image analysis (if needed)
|
||||
@@ -179,23 +189,34 @@ Use LangChain's tool framework but custom agent loop:
|
||||
### Tier 1: Infrastructure Tools (Existing)
|
||||
|
||||
```python
|
||||
@tool
|
||||
async def list_services() -> List[Dict]:
|
||||
from pydantic_ai import Agent, RunContext
|
||||
from pydantic import BaseModel
|
||||
|
||||
class ListServicesResult(BaseModel):
|
||||
services: List[Dict]
|
||||
|
||||
class DeployServiceParams(BaseModel):
|
||||
name: str
|
||||
compose: str
|
||||
|
||||
@agent.tool
|
||||
async def list_services(ctx: RunContext[None]) -> ListServicesResult:
|
||||
"""List all running Docker services"""
|
||||
return await portainer_client.list_containers()
|
||||
services = await portainer_client.list_containers()
|
||||
return ListServicesResult(services=services)
|
||||
|
||||
@tool
|
||||
async def deploy_service(name: str, compose: str) -> str:
|
||||
@agent.tool
|
||||
async def deploy_service(ctx: RunContext[None], params: DeployServiceParams) -> str:
|
||||
"""Deploy a new service from Docker Compose YAML"""
|
||||
return await portainer_client.deploy_stack(name, compose)
|
||||
return await portainer_client.deploy_stack(params.name, params.compose)
|
||||
|
||||
@tool
|
||||
async def create_proxy(domain: str, target: str) -> str:
|
||||
@agent.tool
|
||||
async def create_proxy(ctx: RunContext[None], domain: str, target: str) -> str:
|
||||
"""Create Nginx reverse proxy for a service"""
|
||||
return await npm_client.create_proxy_host(domain, target)
|
||||
|
||||
@tool
|
||||
async def check_service_health(service: str) -> Dict:
|
||||
@agent.tool
|
||||
async def check_service_health(ctx: RunContext[None], service: str) -> Dict:
|
||||
"""Check if a service is healthy"""
|
||||
return await kuma_client.get_monitor_status(service)
|
||||
```
|
||||
@@ -203,18 +224,18 @@ async def check_service_health(service: str) -> Dict:
|
||||
### Tier 2: Knowledge Tools
|
||||
|
||||
```python
|
||||
@tool
|
||||
async def web_search(query: str) -> str:
|
||||
@agent.tool
|
||||
async def web_search(ctx: RunContext[None], query: str) -> str:
|
||||
"""Search the web and extract main content"""
|
||||
return await scraper.scrape_url(query)
|
||||
|
||||
@tool
|
||||
async def query_memory(question: str) -> List[str]:
|
||||
@agent.tool
|
||||
async def query_memory(ctx: RunContext[None], question: str) -> List[str]:
|
||||
"""Search conversation history for relevant context"""
|
||||
return await memory.semantic_search(question)
|
||||
|
||||
@tool
|
||||
async def read_documentation(topic: str) -> str:
|
||||
@agent.tool
|
||||
async def read_documentation(ctx: RunContext[None], topic: str) -> str:
|
||||
"""Read project documentation"""
|
||||
docs_path = f"/docs/{topic}.md"
|
||||
return read_file(docs_path)
|
||||
@@ -223,14 +244,14 @@ async def read_documentation(topic: str) -> str:
|
||||
### Tier 3: Execution Tools (Future)
|
||||
|
||||
```python
|
||||
@tool
|
||||
async def execute_python(code: str) -> str:
|
||||
@agent.tool
|
||||
async def execute_python(ctx: RunContext[None], code: str) -> str:
|
||||
"""Execute Python code in sandbox"""
|
||||
# Future: Integrate code interpreter
|
||||
pass
|
||||
|
||||
@tool
|
||||
async def query_database(sql: str) -> List[Dict]:
|
||||
@agent.tool
|
||||
async def query_database(ctx: RunContext[None], sql: str) -> List[Dict]:
|
||||
"""Query PostgreSQL database"""
|
||||
# Future: Safe SQL execution
|
||||
pass
|
||||
@@ -246,7 +267,7 @@ async def query_database(sql: str) -> List[Dict]:
|
||||
"type": "thinking", # or "tool_call", "content", "error"
|
||||
"content": "Searching the web for nginx configuration...",
|
||||
"tool": "web_search", # optional, if type is tool_call
|
||||
"model": "gemma:7b" # optional, which model is being used
|
||||
"model": "mistral-tools:7b" # optional, which model is being used
|
||||
}
|
||||
|
||||
# Example stream
|
||||
@@ -281,53 +302,56 @@ Open WebUI already supports streaming, we just need to format it correctly:
|
||||
### Intent-Based Routing
|
||||
|
||||
```python
|
||||
from pydantic_ai import Agent
|
||||
|
||||
class ModelRouter:
|
||||
MODELS = {
|
||||
"simple": "gemma:2b", # Fast, <100 tokens
|
||||
"general": "gemma:7b", # Balanced
|
||||
"expert": "mistral:7b", # Complex reasoning
|
||||
"code": "codestral:latest" # Code tasks
|
||||
"simple": "ollama:gemma:2b", # Fast, <100 tokens
|
||||
"general": "ollama:gemma:7b", # Balanced
|
||||
"expert": "ollama:mistral:7b", # Complex reasoning
|
||||
"code": "ollama:codestral:latest" # Code tasks
|
||||
}
|
||||
|
||||
async def select_model(self, message: str, context: str) -> str:
|
||||
# Use lightweight model for routing decision
|
||||
prompt = f"""Analyze this request and categorize:
|
||||
routing_agent = Agent(
|
||||
model="ollama:gemma:2b",
|
||||
result_type=str,
|
||||
system_prompt="""Analyze this request and categorize:
|
||||
|
||||
User: {message}
|
||||
Context: {context}
|
||||
Categories:
|
||||
- simple: Greetings, basic facts, short answers
|
||||
- general: Normal conversation, explanations
|
||||
- expert: Complex reasoning, multi-step problems
|
||||
- code: Programming tasks, debugging
|
||||
|
||||
Categories:
|
||||
- simple: Greetings, basic facts, short answers
|
||||
- general: Normal conversation, explanations
|
||||
- expert: Complex reasoning, multi-step problems
|
||||
- code: Programming tasks, debugging
|
||||
Return ONLY the category.
|
||||
"""
|
||||
)
|
||||
|
||||
Return ONLY the category.
|
||||
"""
|
||||
|
||||
category = await ollama.generate(model="gemma:2b", prompt=prompt)
|
||||
return self.MODELS[category.strip()]
|
||||
result = await routing_agent.run(f"User: {message}\nContext: {context}")
|
||||
return self.MODELS[result.data.strip()]
|
||||
```
|
||||
|
||||
## Next Steps
|
||||
|
||||
1. **Prototype LangGraph agent** (2-3 hours)
|
||||
- Basic agent with 2-3 tools
|
||||
- Streaming with thinking output
|
||||
- Test with Open WebUI
|
||||
1. **Enhance PydanticAI agent** (2-3 hours)
|
||||
- Add more infrastructure tools
|
||||
- Improve streaming with reasoning output
|
||||
- Test with complex queries
|
||||
|
||||
2. **Integrate existing tools** (3-4 hours)
|
||||
- Wrap infrastructure controller as tools
|
||||
- Wrap web scraper as tool
|
||||
- Test tool calling
|
||||
2. **Integrate remaining tools** (3-4 hours)
|
||||
- Migrate all infrastructure controller tools
|
||||
- Add web scraper tool improvements
|
||||
- Test multi-tool workflows
|
||||
|
||||
3. **Model routing** (2 hours)
|
||||
- Implement intent analysis
|
||||
- Add model selection logic
|
||||
- Test performance
|
||||
3. **Model routing enhancements** (2 hours)
|
||||
- Refine intent analysis
|
||||
- Add model selection metrics
|
||||
- Test performance improvements
|
||||
|
||||
4. **Production deployment** (2 hours)
|
||||
- Error handling
|
||||
4. **Production hardening** (2 hours)
|
||||
- Enhanced error handling
|
||||
- Rate limiting
|
||||
- Logging and monitoring
|
||||
- Update API documentation
|
||||
@@ -391,20 +415,20 @@ User: "How do I configure Headscale?"
|
||||
|
||||
## Technology Stack
|
||||
|
||||
- **Agent Framework:** LangGraph 0.2.x
|
||||
- **LLM Integration:** LangChain-Ollama
|
||||
- **Tool Framework:** LangChain Tools
|
||||
- **Agent Framework:** PydanticAI
|
||||
- **LLM Integration:** Native Ollama SDK (HTTP API)
|
||||
- **Tool Framework:** PydanticAI Tools with Pydantic validation
|
||||
- **Streaming:** SSE (Server-Sent Events)
|
||||
- **State Management:** LangGraph StateGraph
|
||||
- **State Management:** RunContext dependency injection
|
||||
- **Memory:** Existing Qdrant integration
|
||||
|
||||
## Risk Mitigation
|
||||
|
||||
**Risk:** LangGraph adds complexity
|
||||
- **Mitigation:** Start simple, add features incrementally
|
||||
**Risk:** PydanticAI is relatively new
|
||||
- **Mitigation:** Strong typing provides safety, active development community
|
||||
|
||||
**Risk:** Tool calling may be slow
|
||||
- **Mitigation:** Parallel execution, caching, optimized tools
|
||||
- **Mitigation:** Async tools enable parallel execution, caching, optimized tools
|
||||
|
||||
**Risk:** Reasoning output may be verbose
|
||||
- **Mitigation:** Configurable verbosity, collapsible UI elements
|
||||
|
||||
Reference in New Issue
Block a user