Major architectural changes and improvements: ## ADK Framework Migration (v0.10.0) - Migrated from LangChain/LangGraph to Google ADK 1.3.0 with LiteLLM 1.80.5 - Improved tool calling reliability with local Ollama models - Converted all 10 tools to ADK async generator format - Updated streaming pipeline for ADK event system - Enhanced error handling and agent initialization ## Model Optimization - Switched from gemma3:12b (10GB VRAM) to gemma3:4b (4.8GB VRAM) - Reduced VRAM usage from 91% to 43% (5.4GB freed) - Optimized for production stability with memory headroom ## Health Check System Overhaul - Optimized /health/full: 6ms response (was 30s+) - Added model verification: confirms configured model is available - New /health/diagnostics endpoint with optional deep testing - Added currently loaded models tracking - Clear emoji status indicators (✅/❌/⚠️) - Fixed AGENT_AVAILABLE flag export for proper health reporting ## Ollama Client Enhancements - Added list_models() method for model inventory - Enhanced model verification in health checks - Better error handling and reporting ## Documentation Updates - Updated STATUS.md to v0.10.0-adk-migration - Comprehensive CHANGELOG.md entry with migration details - Updated PLANS.md showing Phase 4 complete - Updated ai-orchestrator-plan.md with ADK status - Added MIGRATION_PLAN_LANGCHAIN_TO_ADK.md - Added ADK_Ollama_Research.md with implementation analysis ## Technical Details - 10 tools: 7 infrastructure + 2 research + 1 response tool - Framework: Google ADK with UnifiedAgent pattern - System prompt: v7_adk_best_practice - Container health: Now passing Docker healthchecks - Response times: Simple queries ~0.3-1s, Research ~4-7s
13 KiB
Migration Plan: LangChain/LangGraph → Google ADK
Status: Draft - Awaiting Approval Created: 2025-11-25 Estimated Effort: Medium (4-6 hours) Risk Level: Medium
Executive Summary
Replace the current LangChain/LangGraph implementation with Google's Agent Development Kit (ADK) to achieve reliable tool calling with Ollama local models. The current setup fails because LangGraph's create_react_agent doesn't properly trigger tool calls with Mistral 7B, despite the model supporting tools at the API level.
Why This Migration
Current Issues:
- ❌ LangGraph agents not calling tools (empty
tool_calls: []) - ❌ Models hallucinating instead of using tools
- ❌ Gemma models not supported by LangChain (status 400)
- ❌ Only Mistral 7B passes tests, but fails in production
Expected Benefits:
- ✅ Proven Ollama + ADK integration (multiple 2025 examples)
- ✅ Works with Gemma 3, Mistral, Qwen models
- ✅ Model-agnostic architecture (future flexibility)
- ✅ Built-in streaming support
- ✅ Active development and Google backing
Current Architecture Analysis
Files to Modify/Replace
-
src/agent/orchestrator.py(278 lines)- Current: LangGraph
create_react_agentwith ChatOllama - Replace with: ADK Agent with LiteLLM
- Current: LangGraph
-
src/agent/tools.py(429 lines)- Current: LangChain
@tooldecorator - Migrate to: ADK tool format (async generators)
- Current: LangChain
-
src/agent/streaming.py(161 lines)- Current: Converts LangGraph output to SSE
- Update: Adapt for ADK streaming format
-
requirements.txt- Remove:
langgraph,langchain-*packages (5 packages) - Add:
google-adk,litellm(2 packages)
- Remove:
What Stays The Same
- ✅ API endpoints (
src/controllers/ai_controller.py) - minimal changes - ✅ Tool implementations - logic unchanged, only decorators
- ✅ Memory system - completely independent
- ✅ System prompts - reusable
- ✅ Frontend integration - SSE format preserved
Technical Implementation Plan
Phase 1: Dependencies & Setup
1.1 Update requirements.txt
Remove:
langgraph~=1.0.3
langchain~=1.0.8
langchain-community~=0.4.1
langchain-core~=1.1.0
langchain-ollama~=1.0.0
Add:
# Google ADK + Model Integration
google-adk~=1.3.0
litellm~=1.55.0
1.2 Environment Configuration
Add to .env or config:
OLLAMA_API_BASE="http://ollama:11434"
Estimated Time: 15 minutes
Phase 2: Tool Migration
2.1 Convert Tool Decorators
Before (LangChain):
from langchain_core.tools import tool
@tool
async def get_current_time() -> str:
"""Get the current date and time."""
# implementation
return result
After (ADK):
from google.adk.tools import Tool
from typing import AsyncGenerator
async def get_current_time() -> AsyncGenerator[str, None]:
"""Get the current date and time."""
# implementation
yield result
2.2 Tool List Format
Before:
ALL_TOOLS = [
list_services,
get_service_details,
# ...
]
After:
from google.adk.tools import Tool
ALL_TOOLS = [
Tool(
name="get_current_time",
description="Get the current date and time",
fn=get_current_time,
),
Tool(
name="list_services",
description="List all running Docker services",
fn=list_services,
),
# ... convert all 9 tools
]
Files to Modify:
src/agent/tools.py- Convert all 9 tools
Estimated Time: 1 hour
Phase 3: Agent Orchestrator Replacement
3.1 Create New ADK-Based Orchestrator
Key Changes:
- Model Initialization
# Replace ChatOllama with LiteLLM
from google.adk.models.lite_llm import LiteLlm
self.llm = LiteLlm(
model="ollama_chat/mistral:7b",
api_base="http://ollama:11434",
)
- Agent Creation
# Replace create_react_agent with ADK Agent
from google.adk.agents import Agent
self.agent = Agent(
model=self.llm,
name="tatlock",
description="British butler assistant",
instruction=get_prompt(variant), # Reuse system prompts!
tools=ALL_TOOLS,
)
- Streaming Interface
# Replace LangGraph astream with ADK streaming
async def chat(self, message: str, history: List[Dict]) -> AsyncIterator[Dict]:
# Convert history to ADK format
messages = self._build_messages(history, message)
# Stream from ADK agent
async for chunk in self.agent.run(messages, stream=True):
# Map ADK events to our format
yield self._map_chunk(chunk)
3.2 Event Mapping
ADK provides different event types than LangGraph:
tool_call_start→ map to{"type": "tool_call"}tool_call_end→ map to{"type": "tool_result"}content_delta→ map to{"type": "content"}
Files to Modify:
src/agent/orchestrator.py- Complete rewrite (keep interface)
Estimated Time: 2 hours
Phase 4: Streaming Adapter
4.1 Update SSE Converter
The stream_agent_to_sse function should continue to work with minimal changes since we maintain the same intermediate format:
{"type": "tool_call", "tool": "...", "content": "..."}
{"type": "content", "content": "..."}
ADK streaming will provide similar events, just need to map them correctly in the orchestrator.
Files to Modify:
src/agent/streaming.py- Minor adjustments only
Estimated Time: 30 minutes
Phase 5: Integration & Testing
5.1 Update Controller
Minimal changes needed in ai_controller.py:
- Import path changes (
from src.agent import get_unified_agent) - Everything else stays the same
5.2 Testing Checklist
Create comprehensive tests:
# test_adk_integration.py
async def test_basic_chat():
"""Test agent responds without tools"""
agent = get_unified_agent()
response = await agent.chat_completion("Hello!")
assert len(response) > 0
async def test_tool_calling():
"""Test agent calls get_current_time tool"""
agent = get_unified_agent()
response = await agent.chat_completion("What time is it?")
# Should contain actual time, not hallucination
assert "202" in response # Year should be in response
async def test_web_search():
"""Test web_search tool (will fail with DuckDuckGo rate limits)"""
agent = get_unified_agent()
response = await agent.chat_completion("What is the weather in Amsterdam?")
# Should attempt web search
assert response # At minimum, should respond
async def test_streaming():
"""Test streaming output"""
agent = get_unified_agent()
chunks = []
async for chunk in agent.chat("List the services", stream=True):
chunks.append(chunk)
assert len(chunks) > 0
assert any(c["type"] == "content" for c in chunks)
5.3 Model Testing Matrix
Test with multiple models to find best one:
| Model | Size | ADK Compatible? | Tool Calling? | Notes |
|---|---|---|---|---|
| mistral:7b | 4.4GB | ✅ (proven) | ✅ | Current choice |
| gemma3:4b | 3.3GB | ✅ (docs) | ✅ | Better for VRAM |
| qwen3:14b | ~8GB | ✅ (docs) | ✅ | If VRAM allows |
| gemma3:12b | 8.1GB | ✅ (docs) | ✅ | High capability |
Estimated Time: 1.5 hours
Phase 6: Deployment
6.1 Docker Rebuild
# Rebuild core-api service with new dependencies
docker-compose build core-api
docker-compose up -d core-api
6.2 Verification
- Check logs:
docker logs core-api --tail 50 - Test health endpoint
- Test via webui with "What time is it?"
- Monitor for tool calls in logs
6.3 Rollback Plan
Keep LangChain implementation in a git branch:
git checkout -b backup/langchain-implementation
git add -A && git commit -m "Backup before ADK migration"
git checkout main
# ... perform migration ...
# If issues: git checkout backup/langchain-implementation
Estimated Time: 30 minutes
Risk Assessment & Mitigation
High Risks
Risk 1: ADK Tool Calling Issues with Certain Models
- Evidence: GitHub issue #716 mentions "Ollama tool calling incompatible"
- Mitigation: Test multiple models, have fallback to Mistral 7B
- Impact: Could require model switching
Risk 2: Streaming Format Incompatibility
- Evidence: ADK streaming might emit different event structures
- Mitigation: Thorough mapping layer in orchestrator
- Impact: Could affect frontend display
Medium Risks
Risk 3: LiteLLM Configuration
- Evidence: Requires correct env vars and model naming
- Mitigation: Follow documented examples exactly
- Impact: Could cause initialization failures
Risk 4: Breaking Changes in Dependencies
- Evidence: New framework, different API paradigms
- Mitigation: Pin exact versions, test thoroughly
- Impact: Could require API adjustments
Low Risks
Risk 5: Memory System Integration
- Impact: Memory system is independent, should not be affected
- Mitigation: Keep interface the same
Alternative Approaches Considered
Option A: Fix LangGraph (Not Recommended)
- Try different system prompts
- Try different model configurations
- Why Not: Already tried, fundamental compatibility issue
Option B: Pydantic AI (Alternative)
- Pros: Simpler API, type safety
- Cons: Less documentation for Ollama, newer than ADK
- Verdict: ADK has better Ollama documentation
Option C: Custom Implementation
- Pros: Full control
- Cons: Reinventing wheel, more maintenance
- Verdict: ADK provides everything needed
Success Criteria
Must Have (Required for Success)
- ✅ Agent calls tools when appropriate (no hallucination)
- ✅
get_current_timeworks reliably - ✅ Streaming output preserved
- ✅ No regressions in memory system
- ✅ API endpoints unchanged
Should Have (Desired Outcomes)
- ✅ Works with Gemma 3 models (~4GB VRAM)
- ✅ Web search functional (when not rate-limited)
- ✅ Performance equivalent or better
- ✅ Clear logs showing tool calls
Nice to Have (Bonus)
- ✅ Support for multiple model providers
- ✅ Better error messages
- ✅ Reduced memory footprint
Timeline & Effort Estimate
| Phase | Time | Dependencies |
|---|---|---|
| 1. Dependencies | 15 min | None |
| 2. Tool Migration | 1 hour | Phase 1 |
| 3. Orchestrator | 2 hours | Phase 2 |
| 4. Streaming | 30 min | Phase 3 |
| 5. Testing | 1.5 hours | Phase 4 |
| 6. Deployment | 30 min | Phase 5 |
| Total | 5.5 hours | Sequential |
Add 1 hour buffer for unexpected issues = 6.5 hours total
Post-Migration Tasks
-
Documentation
- Update README with ADK setup instructions
- Document model compatibility matrix
- Add troubleshooting guide
-
Monitoring
- Watch for tool calling failures
- Monitor response quality
- Track VRAM usage
-
Optimization
- Test alternative models for better VRAM efficiency
- Tune system prompts for ADK
- Consider caching strategies
Key Resources
Google ADK Documentation:
Integration Examples:
Package Documentation:
Open Questions
-
Q: Which Ollama model works best with ADK tool calling?
- A: Test Mistral 7B, Gemma 3:4b, Qwen 3:14b in Phase 5
-
Q: Does ADK support custom SSE format for Open WebUI?
- A: Yes, we maintain the mapping layer in streaming.py
-
Q: Will system prompts need modification?
- A: Likely minor tweaks, but current prompts should mostly work
-
Q: What's the fallback if ADK also fails?
- A: Consider Pydantic AI or custom implementation
Approval Checklist
Before proceeding with implementation:
- User approves overall approach
- User agrees with Google ADK choice
- User confirms 6-hour timeline is acceptable
- User reviews risk assessment
- User approves model testing matrix
- User confirms rollback plan is sufficient
Next Steps After Approval:
- Create feature branch:
feature/migrate-to-adk - Begin Phase 1: Dependencies
- Document progress in this file
- Request code review after Phase 5
- Deploy to production after testing
Plan prepared by Claude Code Ready for user review and approval