9.7 KiB
Migration to Model-Level Tool Routing
Date: 2025-11-23 Status: Complete Impact: Simplified architecture, LLM decides tool usage
Summary
Removed application-level routing (use_agent parameter) in favor of model-level routing where mistral:7b autonomously decides whether to use tools or answer directly.
Architectural Change
Before (Application-Level Routing):
# AI Controller decides routing
if request.use_agent:
→ Route to agent (mistral:7b with tools)
else:
→ Direct Ollama call (any model)
Problem: Application layer must decide which queries need tools
After (Model-Level Routing):
# Always route through agent, LLM decides tool usage
→ Unified Agent (mistral:7b with tools)
→ LLM analyzes query autonomously
→ LLM decides: use tools OR answer directly
Solution: LLM understands context and decides intelligently
Why This is Better
✅ LLM Already Has This Capability
LangGraph's create_react_agent means:
- mistral:7b sees available tools during generation
- mistral:7b outputs tool calls when needed
- mistral:7b answers directly when tools aren't needed
- No application-level classification required
✅ Simpler Code
Removed:
use_agent: boolparameter from request schema- Conditional routing logic in ai_controller.py
- Need to document when to use
use_agent=true
Result: Single code path for all requests
✅ More Intelligent
The LLM understands nuance better than boolean flags:
| Query | LLM Decision | Application Would Have |
|---|---|---|
| "What is Docker?" | Answer directly (no tools) | ❌ Might route wrong |
| "Is core-api running?" | Use tool (needs real data) | ✅ Correct |
| "List services and explain what Docker is" | Use tool + knowledge | ✅ Handles complexity |
✅ Consistent UX
- Always get thinking indicators
[💭 Analyzing...] - Always see tool usage
[🔧 Checking services...] - More transparent reasoning process
✅ Perfect for Homelab Context
- Token usage doesn't matter - Running locally on Ollama (free)
- Latency increase minimal - ~1-2s extra for simple queries
- Flexibility matters more - Edge cases handled automatically
Implementation Changes
1. Removed use_agent Parameter
File: src/api/v1/schemas.py
# REMOVED
use_agent: bool = Field(
default=True,
description="Use intelligent agent with tool calling and reasoning (recommended)"
)
Now all requests go through agent by default.
2. Simplified AI Controller
File: src/controllers/ai_controller.py
# Before
if request.use_agent and AGENT_AVAILABLE:
# Route to agent
else:
# Direct Ollama
# After
if AGENT_AVAILABLE:
try:
# Always route through agent
# mistral:7b decides tool usage
except Exception as e:
# Fallback to direct Ollama if agent fails
Added try-except for graceful fallback if agent initialization fails.
3. Maintained Fallback
If agent is unavailable or fails:
- Falls back to direct Ollama call
- Uses requested model (gemma:2b, gemma:7b, etc.)
- No intelligent tool routing, just basic chat
How It Works
LangGraph ReAct Loop
User Query
↓
mistral:7b (with bound tools)
↓
[Thought] Analyze query + available tools
↓
[Decision] Does this need a tool?
├─→ NO → Generate answer directly
└─→ YES → Call tool(s) → Get results → Synthesize answer
The model sees tool descriptions and autonomously decides:
# Tools are bound to the LLM
llm_with_tools = ChatOllama(model="mistral:7b").bind_tools(tools)
# LLM output contains tool_calls if it wants to use tools
response = llm_with_tools.invoke(messages)
if response.tool_calls:
# Execute tools
else:
# Return answer directly
Key Point: The application doesn't decide tool usage - it just checks if the LLM outputted tool calls.
Test Results
All query types work correctly with mistral:7b deciding autonomously:
Test 1: Simple Math (No Tools)
Query: "What is 2+2?"
Response: "The sum of 2+2 is 4."
Tool Calls: None ✓
Time: ~2s
Test 2: Infrastructure Query (Needs Tools)
Query: "List all running services"
Response: [Detailed service list with ports]
Tool Calls: list_services ✓
Time: ~5s
Test 3: Knowledge Question (No Tools)
Query: "What is Docker?"
Response: [Detailed Docker explanation]
Tool Calls: None ✓
Time: ~2s
Test 4: Streaming with Tools
Query: "Check service health for core-api"
Stream: [💭 Analyzing...] → "To check the health status..."
Tool Calls: check_service_health ✓
Time: ~4s
Performance Impact
Latency Comparison
| Query Type | Before (use_agent=false) | After (always agent) | Delta |
|---|---|---|---|
| Simple math | ~1s (gemma:2b direct) | ~2s (mistral:7b) | +1s |
| Knowledge | ~1-2s (gemma:7b direct) | ~2s (mistral:7b) | ~0s |
| Tool needed | ~5s (mistral:7b agent) | ~5s (mistral:7b) | 0s |
| Multi-tool | ~10s (mistral:7b agent) | ~10s (mistral:7b) | 0s |
Verdict: Minimal impact (<2s for simple queries), acceptable for homelab use
Token Usage
- Agent adds reasoning tokens (~100-200 extra per request)
- Impact: Zero (local Ollama, tokens are free)
Memory Usage
- Consistent: Always uses mistral:7b (~4GB when loaded)
- Before: Mixed (gemma:2b ~1GB, gemma:7b ~3GB, mistral:7b ~4GB)
- Result: More predictable resource usage
Benefits Summary
| Aspect | Benefit |
|---|---|
| Code Complexity | Reduced - single code path |
| Maintainability | Improved - less conditional logic |
| Flexibility | Increased - LLM handles edge cases |
| User Experience | Consistent - always see reasoning |
| Performance | Acceptable - ~1-2s increase for simple queries |
| Context Awareness | Better - LLM understands nuance |
OpenAI Compatibility
Still fully compatible with OpenAI clients:
# Works with any OpenAI-compatible client
curl -X POST http://api.schweitz.net/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-3.5-turbo",
"messages": [{"role": "user", "content": "List services"}],
"stream": true
}'
No use_agent parameter needed - agent is transparent to client
Migration for Clients
Before
# Client had to know when to use agent
response = client.chat.completions.create(
model="gpt-3.5-turbo",
messages=[{"role": "user", "content": "List services"}],
extra_body={"use_agent": True} # Had to specify
)
After
# Client doesn't need to know about agent
response = client.chat.completions.create(
model="gpt-3.5-turbo",
messages=[{"role": "user", "content": "List services"}]
# Agent automatically handles everything
)
Migration: Remove use_agent parameter from client code - it's ignored now
Fallback Behavior
If agent fails to initialize or encounters an error:
try:
# Route through agent
response = await agent.chat(...)
except Exception as e:
logger.error(f"Agent failed, falling back to direct Ollama: {e}")
# Fall through to direct Ollama call
# Uses requested model without tool capabilities
Ensures service remains available even if agent has issues.
Research Findings
From LangChain/LangGraph best practices:
- Tool calling is model-level - LLMs natively support tool calling, application should just expose tools
- ReAct pattern - LangGraph's
create_react_agentimplements Reason+Act loop where LLM decides actions - Simpler is better - Industry consensus is to let LLM decide tool usage rather than hardcode routing
bind_tools()vs routing - Usebind_tools()for flexibility, use routing only when needed (cost, latency critical)
For homelab context where tokens are free and flexibility matters, model-level routing is the clear winner.
Future Enhancements
1. Model Routing (Optional)
Could add intelligent model selection:
# Agent detects task type
if task_type == "code":
use codestral:latest
elif task_type == "analysis":
use mixtral:8x7b
else:
use mistral:7b (default)
2. Tool Result Caching
Cache infrastructure queries:
- Service list (60s TTL)
- Domain list (5min TTL)
- Reduces repeated tool calls
3. Parallel Tool Execution
When agent needs multiple independent tools:
# Sequential: 3 tools × 2s = 6s
# Parallel: max(tool times) = ~2s
Documentation Updates Needed
- Update API documentation to remove
use_agent - Update Open WebUI integration guide
- Add architecture diagrams showing model-level routing
- Document test results and performance characteristics
Conclusion
Migration successful! The system now:
- ✅ Uses model-level routing (LLM decides tool usage)
- ✅ Simpler codebase (removed
use_agentparameter) - ✅ More intelligent (LLM understands context)
- ✅ Consistent UX (always see reasoning)
- ✅ Maintains fallback (direct Ollama if agent fails)
- ✅ OpenAI-compatible (clients don't need to change)
The agent is now transparent to users - they just chat naturally and mistral:7b intelligently decides when to use tools.
Related Files
- AI Controller - Simplified routing
- Request Schema - Removed
use_agent - Agent Orchestrator - Unchanged (already did model-level)
- Agent Flow Diagrams - Visual architecture