Files
portainer-core/plans/active/unified-agent-architecture.md
T

15 KiB

Unified Agent Architecture Plan

Date: 2025-11-23 Objective: Build a single intelligent agent that handles all tool routing, multi-modal processing, and agentic reasoning internally, exposing one simple chat endpoint to any UI

Vision

Instead of configuring functions in Open WebUI (or any other UI), the Core API becomes an intelligent orchestrator that:

  1. Accepts simple chat messages - Just like talking to ChatGPT
  2. Internally routes to specialized tools/models - Infrastructure management, web search, code execution, etc.
  3. Streams reasoning/thinking - Shows what it's doing ("Searching the web...", "Querying database...", "Using expert model...")
  4. Returns unified responses - Combines results from multiple sources transparently

Benefits

✅ UI-agnostic - Works with Open WebUI, CLI, mobile apps, any client ✅ No configuration needed - Users just chat naturally ✅ Transparent reasoning - See what's happening under the hood ✅ Tool discovery - Agent decides when to use tools, not manual triggers ✅ Multi-modal support - Handle text, images, code, infrastructure queries ✅ Expert model routing - Use small models for simple tasks, large for complex

Architecture Overview

┌─────────────────────────────────────────────────────────────┐
│                     User Interface                          │
│         (Open WebUI, CLI, Mobile App, etc.)                 │
└──────────────────────┬──────────────────────────────────────┘
                       │ Simple chat: "Deploy nginx proxy"
                       ↓
┌─────────────────────────────────────────────────────────────┐
│              Core API - Unified Agent                       │
│  /v1/chat/completions (OpenAI-compatible endpoint)          │
└──────────────────────┬──────────────────────────────────────┘
                       │
                       ↓
┌─────────────────────────────────────────────────────────────┐
│           Agent Orchestrator (LangGraph)                    │
│  ┌──────────────────────────────────────────────┐           │
│  │  Reasoning Loop:                             │           │
│  │  1. Analyze user intent                      │           │
│  │  2. Select appropriate tool(s)               │           │
│  │  3. Execute tool calls                       │           │
│  │  4. Synthesize results                       │           │
│  │  5. Stream thinking/reasoning                │           │
│  └──────────────────────────────────────────────┘           │
└──────────────────────┬──────────────────────────────────────┘
                       │
        ┌──────────────┼──────────────┬──────────────┐
        │              │              │              │
        ↓              ↓              ↓              ↓
┌──────────────┐ ┌──────────┐ ┌──────────┐ ┌──────────────┐
│ Tool Catalog │ │ Models   │ │ Memory   │ │ Knowledge    │
│              │ │          │ │          │ │              │
│ • Infra Mgmt │ │ • Gemma  │ │ • Qdrant │ │ • Web Search │
│ • Web Scrape │ │ • Codestral│ │ • Buffer│ │ • Docs       │
│ • File Ops   │ │ • Mistral│ │          │ │              │
│ • Code Exec  │ │          │ │          │ │              │
└──────────────┘ └──────────┘ └──────────┘ └──────────────┘

Implementation Options

Pros:

  • Built-in agent loops and tool calling
  • State management for multi-step reasoning
  • Streaming support for intermediate steps
  • Well-documented patterns
  • Active development

Cons:

  • Additional dependency (~50MB)
  • Learning curve for LangGraph concepts
  • Some overhead vs custom implementation

Example flow:

from langgraph.prebuilt import create_react_agent
from langchain_core.tools import tool

@tool
def deploy_service(service_name: str, compose_yaml: str) -> str:
    """Deploy a containerized service via Portainer"""
    # Use existing infrastructure controller
    return portainer_client.deploy_stack(...)

@tool
def web_search(query: str) -> str:
    """Search the web and extract content"""
    # Use existing web scraper
    return scraper.scrape(...)

agent = create_react_agent(
    model=ChatOllama(model="gemma:7b"),
    tools=[deploy_service, web_search, ...],
    state_modifier="You are a homelab infrastructure assistant..."
)

# Streaming with reasoning
for chunk in agent.stream({"messages": [user_message]}):
    if "thinking" in chunk:
        yield f"data: {json.dumps({'reasoning': chunk['thinking']})}\n\n"
    if "tool_calls" in chunk:
        yield f"data: {json.dumps({'tool': chunk['tool_calls'][0]['name']})}\n\n"
    if "response" in chunk:
        yield f"data: {json.dumps({'content': chunk['response']})}\n\n"

Option 2: Custom Agent Loop

Pros:

  • Full control over behavior
  • Minimal dependencies
  • Optimized for specific use case
  • Easier to debug

Cons:

  • More code to maintain
  • Need to implement tool calling protocol
  • Reinventing some wheels

Example flow:

class UnifiedAgent:
    def __init__(self):
        self.tools = ToolCatalog()
        self.model = OllamaClient()

    async def process(self, user_message: str):
        # 1. Intent analysis
        yield {"type": "thinking", "content": "Analyzing your request..."}
        intent = await self.analyze_intent(user_message)

        # 2. Tool selection
        if intent.requires_tool:
            yield {"type": "thinking", "content": f"Using {intent.tool_name}..."}
            tool_result = await self.tools.execute(intent.tool_name, intent.params)

        # 3. Response generation
        yield {"type": "thinking", "content": "Generating response..."}
        response = await self.model.generate(context=tool_result)

        yield {"type": "content", "content": response}

Option 3: Hybrid (LangChain Tools + Custom Orchestration)

Use LangChain's tool framework but custom agent loop:

  • Leverage @tool decorator for easy tool definitions
  • Custom routing logic for model selection
  • Manual streaming control

Phase 1: Core Agent (Week 1)

  • Set up LangGraph agent with basic tools
  • Implement streaming with reasoning output
  • Wire up existing infrastructure tools
  • Test with simple queries

Phase 2: Advanced Routing (Week 2)

  • Multi-model routing (small for simple, large for complex)
  • Parallel tool execution
  • Error handling and retries
  • Context management

Phase 3: Multi-Modal (Week 3)

  • Image analysis (if needed)
  • Code execution sandbox
  • File operations
  • Database queries

Tool Catalog Design

Tier 1: Infrastructure Tools (Existing)

@tool
async def list_services() -> List[Dict]:
    """List all running Docker services"""
    return await portainer_client.list_containers()

@tool
async def deploy_service(name: str, compose: str) -> str:
    """Deploy a new service from Docker Compose YAML"""
    return await portainer_client.deploy_stack(name, compose)

@tool
async def create_proxy(domain: str, target: str) -> str:
    """Create Nginx reverse proxy for a service"""
    return await npm_client.create_proxy_host(domain, target)

@tool
async def check_service_health(service: str) -> Dict:
    """Check if a service is healthy"""
    return await kuma_client.get_monitor_status(service)

Tier 2: Knowledge Tools

@tool
async def web_search(query: str) -> str:
    """Search the web and extract main content"""
    return await scraper.scrape_url(query)

@tool
async def query_memory(question: str) -> List[str]:
    """Search conversation history for relevant context"""
    return await memory.semantic_search(question)

@tool
async def read_documentation(topic: str) -> str:
    """Read project documentation"""
    docs_path = f"/docs/{topic}.md"
    return read_file(docs_path)

Tier 3: Execution Tools (Future)

@tool
async def execute_python(code: str) -> str:
    """Execute Python code in sandbox"""
    # Future: Integrate code interpreter
    pass

@tool
async def query_database(sql: str) -> List[Dict]:
    """Query PostgreSQL database"""
    # Future: Safe SQL execution
    pass

Streaming Reasoning Output

SSE Format for Transparency

# Stream format
{
    "type": "thinking",  # or "tool_call", "content", "error"
    "content": "Searching the web for nginx configuration...",
    "tool": "web_search",  # optional, if type is tool_call
    "model": "gemma:7b"    # optional, which model is being used
}

# Example stream
data: {"type": "thinking", "content": "Analyzing your request..."}

data: {"type": "thinking", "content": "Detected infrastructure task"}

data: {"type": "tool_call", "tool": "list_services", "content": "Checking current services..."}

data: {"type": "thinking", "content": "Found 22 running services"}

data: {"type": "thinking", "content": "Using expert model for response..."}

data: {"type": "model_switch", "from": "gemma:2b", "to": "mistral:7b"}

data: {"type": "content", "content": "Here are your running services:\n\n..."}

data: [DONE]

Open WebUI Integration

Open WebUI already supports streaming, we just need to format it correctly:

// Open WebUI will render thinking/reasoning in a collapsible section
// Standard content renders as usual

Model Routing Strategy

Intent-Based Routing

class ModelRouter:
    MODELS = {
        "simple": "gemma:2b",      # Fast, <100 tokens
        "general": "gemma:7b",     # Balanced
        "expert": "mistral:7b",    # Complex reasoning
        "code": "codestral:latest" # Code tasks
    }

    async def select_model(self, message: str, context: str) -> str:
        # Use lightweight model for routing decision
        prompt = f"""Analyze this request and categorize:

        User: {message}
        Context: {context}

        Categories:
        - simple: Greetings, basic facts, short answers
        - general: Normal conversation, explanations
        - expert: Complex reasoning, multi-step problems
        - code: Programming tasks, debugging

        Return ONLY the category.
        """

        category = await ollama.generate(model="gemma:2b", prompt=prompt)
        return self.MODELS[category.strip()]

Next Steps

  1. Prototype LangGraph agent (2-3 hours)

    • Basic agent with 2-3 tools
    • Streaming with thinking output
    • Test with Open WebUI
  2. Integrate existing tools (3-4 hours)

    • Wrap infrastructure controller as tools
    • Wrap web scraper as tool
    • Test tool calling
  3. Model routing (2 hours)

    • Implement intent analysis
    • Add model selection logic
    • Test performance
  4. Production deployment (2 hours)

    • Error handling
    • Rate limiting
    • Logging and monitoring
    • Update API documentation

Total effort: ~12-15 hours (1-2 weeks of focused work)

Success Criteria

✅ User can chat naturally without configuring functions ✅ Agent automatically uses tools when appropriate ✅ Streaming shows what the agent is doing ✅ Works with Open WebUI without changes ✅ Can be used from CLI/API directly ✅ Performance is acceptable (<5s for tool-using responses) ✅ Errors are handled gracefully

Example User Flows

Flow 1: Infrastructure Query

User: "What services are currently running?"

[Thinking: Analyzing request...]
[Thinking: Detected infrastructure query]
[Tool Call: list_services - Fetching service list...]
[Thinking: Processing results...]
[Content: You have 22 services running:
- ollama (healthy)
- core-api (healthy)
- ...]

Flow 2: Complex Task

User: "Deploy an nginx proxy for my new blog at blog.schweitz.net"

[Thinking: Breaking down the task...]
[Thinking: Need to deploy nginx and configure NPM]
[Tool Call: deploy_service - Deploying nginx container...]
[Tool Call: create_proxy - Creating proxy host...]
[Thinking: Configuring SSL certificate...]
[Content: Done! Your blog is now accessible at https://blog.schweitz.net
- Nginx container: running
- SSL certificate: active
- Health check: passing]

Flow 3: Knowledge Query

User: "How do I configure Headscale?"

[Thinking: Checking documentation...]
[Tool Call: read_documentation(headscale)]
[Thinking: Extracting relevant steps...]
[Content: To configure Headscale on tower-of-joy:

1. Create a user: `headscale users create homelab`
2. Generate auth key: `headscale preauthkeys create...`
...]

Technology Stack

  • Agent Framework: LangGraph 0.2.x
  • LLM Integration: LangChain-Ollama
  • Tool Framework: LangChain Tools
  • Streaming: SSE (Server-Sent Events)
  • State Management: LangGraph StateGraph
  • Memory: Existing Qdrant integration

Risk Mitigation

Risk: LangGraph adds complexity

  • Mitigation: Start simple, add features incrementally

Risk: Tool calling may be slow

  • Mitigation: Parallel execution, caching, optimized tools

Risk: Reasoning output may be verbose

  • Mitigation: Configurable verbosity, collapsible UI elements

Risk: May not work with all UIs

  • Mitigation: Stick to OpenAI-compatible streaming format

Open Questions

  1. Should we support function calling format for backwards compatibility?
  2. How verbose should reasoning output be?
  3. Should we cache tool results?
  4. Do we need user confirmation for destructive operations?
  5. Should tools have permission levels based on user?

Ready to implement: Yes ✓ Estimated timeline: 1-2 weeks Priority: High (enables true agentic behavior)