Files
portainer-core/docs/sessions/2025-11-23-model-level-routing.md
T

9.7 KiB
Raw Blame History

Migration to Model-Level Tool Routing

Date: 2025-11-23 Status: Complete Impact: Simplified architecture, LLM decides tool usage

Summary

Removed application-level routing (use_agent parameter) in favor of model-level routing where mistral:7b autonomously decides whether to use tools or answer directly.

Architectural Change

Before (Application-Level Routing):

# AI Controller decides routing
if request.use_agent:
    → Route to agent (mistral:7b with tools)
else:
    → Direct Ollama call (any model)

Problem: Application layer must decide which queries need tools

After (Model-Level Routing):

# Always route through agent, LLM decides tool usage
→ Unified Agent (mistral:7b with tools)
   → LLM analyzes query autonomously
   → LLM decides: use tools OR answer directly

Solution: LLM understands context and decides intelligently

Why This is Better

✅ LLM Already Has This Capability

LangGraph's create_react_agent means:

  • mistral:7b sees available tools during generation
  • mistral:7b outputs tool calls when needed
  • mistral:7b answers directly when tools aren't needed
  • No application-level classification required

✅ Simpler Code

Removed:

  • use_agent: bool parameter from request schema
  • Conditional routing logic in ai_controller.py
  • Need to document when to use use_agent=true

Result: Single code path for all requests

✅ More Intelligent

The LLM understands nuance better than boolean flags:

Query LLM Decision Application Would Have
"What is Docker?" Answer directly (no tools) ❌ Might route wrong
"Is core-api running?" Use tool (needs real data) ✅ Correct
"List services and explain what Docker is" Use tool + knowledge ✅ Handles complexity

✅ Consistent UX

  • Always get thinking indicators [💭 Analyzing...]
  • Always see tool usage [🔧 Checking services...]
  • More transparent reasoning process

✅ Perfect for Homelab Context

  • Token usage doesn't matter - Running locally on Ollama (free)
  • Latency increase minimal - ~1-2s extra for simple queries
  • Flexibility matters more - Edge cases handled automatically

Implementation Changes

1. Removed use_agent Parameter

File: src/api/v1/schemas.py

# REMOVED
use_agent: bool = Field(
    default=True,
    description="Use intelligent agent with tool calling and reasoning (recommended)"
)

Now all requests go through agent by default.

2. Simplified AI Controller

File: src/controllers/ai_controller.py

# Before
if request.use_agent and AGENT_AVAILABLE:
    # Route to agent
else:
    # Direct Ollama

# After
if AGENT_AVAILABLE:
    try:
        # Always route through agent
        # mistral:7b decides tool usage
    except Exception as e:
        # Fallback to direct Ollama if agent fails

Added try-except for graceful fallback if agent initialization fails.

3. Maintained Fallback

If agent is unavailable or fails:

  • Falls back to direct Ollama call
  • Uses requested model (gemma:2b, gemma:7b, etc.)
  • No intelligent tool routing, just basic chat

How It Works

LangGraph ReAct Loop

User Query
    ↓
mistral:7b (with bound tools)
    ↓
[Thought] Analyze query + available tools
    ↓
[Decision] Does this need a tool?
    ├─→ NO  → Generate answer directly
    └─→ YES → Call tool(s) → Get results → Synthesize answer

The model sees tool descriptions and autonomously decides:

# Tools are bound to the LLM
llm_with_tools = ChatOllama(model="mistral:7b").bind_tools(tools)

# LLM output contains tool_calls if it wants to use tools
response = llm_with_tools.invoke(messages)

if response.tool_calls:
    # Execute tools
else:
    # Return answer directly

Key Point: The application doesn't decide tool usage - it just checks if the LLM outputted tool calls.

Test Results

All query types work correctly with mistral:7b deciding autonomously:

Test 1: Simple Math (No Tools)

Query: "What is 2+2?"
Response: "The sum of 2+2 is 4."
Tool Calls: None ✓
Time: ~2s

Test 2: Infrastructure Query (Needs Tools)

Query: "List all running services"
Response: [Detailed service list with ports]
Tool Calls: list_services ✓
Time: ~5s

Test 3: Knowledge Question (No Tools)

Query: "What is Docker?"
Response: [Detailed Docker explanation]
Tool Calls: None ✓
Time: ~2s

Test 4: Streaming with Tools

Query: "Check service health for core-api"
Stream: [💭 Analyzing...] → "To check the health status..."
Tool Calls: check_service_health ✓
Time: ~4s

Performance Impact

Latency Comparison

Query Type Before (use_agent=false) After (always agent) Delta
Simple math ~1s (gemma:2b direct) ~2s (mistral:7b) +1s
Knowledge ~1-2s (gemma:7b direct) ~2s (mistral:7b) ~0s
Tool needed ~5s (mistral:7b agent) ~5s (mistral:7b) 0s
Multi-tool ~10s (mistral:7b agent) ~10s (mistral:7b) 0s

Verdict: Minimal impact (<2s for simple queries), acceptable for homelab use

Token Usage

  • Agent adds reasoning tokens (~100-200 extra per request)
  • Impact: Zero (local Ollama, tokens are free)

Memory Usage

  • Consistent: Always uses mistral:7b (~4GB when loaded)
  • Before: Mixed (gemma:2b ~1GB, gemma:7b ~3GB, mistral:7b ~4GB)
  • Result: More predictable resource usage

Benefits Summary

Aspect Benefit
Code Complexity Reduced - single code path
Maintainability Improved - less conditional logic
Flexibility Increased - LLM handles edge cases
User Experience Consistent - always see reasoning
Performance Acceptable - ~1-2s increase for simple queries
Context Awareness Better - LLM understands nuance

OpenAI Compatibility

Still fully compatible with OpenAI clients:

# Works with any OpenAI-compatible client
curl -X POST http://api.schweitz.net/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-3.5-turbo",
    "messages": [{"role": "user", "content": "List services"}],
    "stream": true
  }'

No use_agent parameter needed - agent is transparent to client

Migration for Clients

Before

# Client had to know when to use agent
response = client.chat.completions.create(
    model="gpt-3.5-turbo",
    messages=[{"role": "user", "content": "List services"}],
    extra_body={"use_agent": True}  # Had to specify
)

After

# Client doesn't need to know about agent
response = client.chat.completions.create(
    model="gpt-3.5-turbo",
    messages=[{"role": "user", "content": "List services"}]
    # Agent automatically handles everything
)

Migration: Remove use_agent parameter from client code - it's ignored now

Fallback Behavior

If agent fails to initialize or encounters an error:

try:
    # Route through agent
    response = await agent.chat(...)
except Exception as e:
    logger.error(f"Agent failed, falling back to direct Ollama: {e}")
    # Fall through to direct Ollama call
    # Uses requested model without tool capabilities

Ensures service remains available even if agent has issues.

Research Findings

From LangChain/LangGraph best practices:

  1. Tool calling is model-level - LLMs natively support tool calling, application should just expose tools
  2. ReAct pattern - LangGraph's create_react_agent implements Reason+Act loop where LLM decides actions
  3. Simpler is better - Industry consensus is to let LLM decide tool usage rather than hardcode routing
  4. bind_tools() vs routing - Use bind_tools() for flexibility, use routing only when needed (cost, latency critical)

For homelab context where tokens are free and flexibility matters, model-level routing is the clear winner.

Future Enhancements

1. Model Routing (Optional)

Could add intelligent model selection:

# Agent detects task type
if task_type == "code":
    use codestral:latest
elif task_type == "analysis":
    use mixtral:8x7b
else:
    use mistral:7b (default)

2. Tool Result Caching

Cache infrastructure queries:

  • Service list (60s TTL)
  • Domain list (5min TTL)
  • Reduces repeated tool calls

3. Parallel Tool Execution

When agent needs multiple independent tools:

# Sequential: 3 tools × 2s = 6s
# Parallel: max(tool times) = ~2s

Documentation Updates Needed

  • Update API documentation to remove use_agent
  • Update Open WebUI integration guide
  • Add architecture diagrams showing model-level routing
  • Document test results and performance characteristics

Conclusion

Migration successful! The system now:

  • ✅ Uses model-level routing (LLM decides tool usage)
  • ✅ Simpler codebase (removed use_agent parameter)
  • ✅ More intelligent (LLM understands context)
  • ✅ Consistent UX (always see reasoning)
  • ✅ Maintains fallback (direct Ollama if agent fails)
  • ✅ OpenAI-compatible (clients don't need to change)

The agent is now transparent to users - they just chat naturally and mistral:7b intelligently decides when to use tools.