Files
portainer-core/MIGRATION_PLAN_LANGCHAIN_TO_ADK.md
T
jpmschweitzer e3b451b7b0 feat(ai): complete ADK migration and optimize system health checks
Major architectural changes and improvements:

## ADK Framework Migration (v0.10.0)
- Migrated from LangChain/LangGraph to Google ADK 1.3.0 with LiteLLM 1.80.5
- Improved tool calling reliability with local Ollama models
- Converted all 10 tools to ADK async generator format
- Updated streaming pipeline for ADK event system
- Enhanced error handling and agent initialization

## Model Optimization
- Switched from gemma3:12b (10GB VRAM) to gemma3:4b (4.8GB VRAM)
- Reduced VRAM usage from 91% to 43% (5.4GB freed)
- Optimized for production stability with memory headroom

## Health Check System Overhaul
- Optimized /health/full: 6ms response (was 30s+)
- Added model verification: confirms configured model is available
- New /health/diagnostics endpoint with optional deep testing
- Added currently loaded models tracking
- Clear emoji status indicators (✅/❌/⚠️)
- Fixed AGENT_AVAILABLE flag export for proper health reporting

## Ollama Client Enhancements
- Added list_models() method for model inventory
- Enhanced model verification in health checks
- Better error handling and reporting

## Documentation Updates
- Updated STATUS.md to v0.10.0-adk-migration
- Comprehensive CHANGELOG.md entry with migration details
- Updated PLANS.md showing Phase 4 complete
- Updated ai-orchestrator-plan.md with ADK status
- Added MIGRATION_PLAN_LANGCHAIN_TO_ADK.md
- Added ADK_Ollama_Research.md with implementation analysis

## Technical Details
- 10 tools: 7 infrastructure + 2 research + 1 response tool
- Framework: Google ADK with UnifiedAgent pattern
- System prompt: v7_adk_best_practice
- Container health: Now passing Docker healthchecks
- Response times: Simple queries ~0.3-1s, Research ~4-7s
2025-11-26 08:36:50 +01:00

13 KiB

Migration Plan: LangChain/LangGraph → Google ADK

Status: Draft - Awaiting Approval Created: 2025-11-25 Estimated Effort: Medium (4-6 hours) Risk Level: Medium


Executive Summary

Replace the current LangChain/LangGraph implementation with Google's Agent Development Kit (ADK) to achieve reliable tool calling with Ollama local models. The current setup fails because LangGraph's create_react_agent doesn't properly trigger tool calls with Mistral 7B, despite the model supporting tools at the API level.

Why This Migration

Current Issues:

  • ❌ LangGraph agents not calling tools (empty tool_calls: [])
  • ❌ Models hallucinating instead of using tools
  • ❌ Gemma models not supported by LangChain (status 400)
  • ❌ Only Mistral 7B passes tests, but fails in production

Expected Benefits:

  • ✅ Proven Ollama + ADK integration (multiple 2025 examples)
  • ✅ Works with Gemma 3, Mistral, Qwen models
  • ✅ Model-agnostic architecture (future flexibility)
  • ✅ Built-in streaming support
  • ✅ Active development and Google backing

Current Architecture Analysis

Files to Modify/Replace

  1. src/agent/orchestrator.py (278 lines)

    • Current: LangGraph create_react_agent with ChatOllama
    • Replace with: ADK Agent with LiteLLM
  2. src/agent/tools.py (429 lines)

    • Current: LangChain @tool decorator
    • Migrate to: ADK tool format (async generators)
  3. src/agent/streaming.py (161 lines)

    • Current: Converts LangGraph output to SSE
    • Update: Adapt for ADK streaming format
  4. requirements.txt

    • Remove: langgraph, langchain-* packages (5 packages)
    • Add: google-adk, litellm (2 packages)

What Stays The Same

  • ✅ API endpoints (src/controllers/ai_controller.py) - minimal changes
  • ✅ Tool implementations - logic unchanged, only decorators
  • ✅ Memory system - completely independent
  • ✅ System prompts - reusable
  • ✅ Frontend integration - SSE format preserved

Technical Implementation Plan

Phase 1: Dependencies & Setup

1.1 Update requirements.txt

Remove:

langgraph~=1.0.3
langchain~=1.0.8
langchain-community~=0.4.1
langchain-core~=1.1.0
langchain-ollama~=1.0.0

Add:

# Google ADK + Model Integration
google-adk~=1.3.0
litellm~=1.55.0

1.2 Environment Configuration

Add to .env or config:

OLLAMA_API_BASE="http://ollama:11434"

Estimated Time: 15 minutes


Phase 2: Tool Migration

2.1 Convert Tool Decorators

Before (LangChain):

from langchain_core.tools import tool

@tool
async def get_current_time() -> str:
    """Get the current date and time."""
    # implementation
    return result

After (ADK):

from google.adk.tools import Tool
from typing import AsyncGenerator

async def get_current_time() -> AsyncGenerator[str, None]:
    """Get the current date and time."""
    # implementation
    yield result

2.2 Tool List Format

Before:

ALL_TOOLS = [
    list_services,
    get_service_details,
    # ...
]

After:

from google.adk.tools import Tool

ALL_TOOLS = [
    Tool(
        name="get_current_time",
        description="Get the current date and time",
        fn=get_current_time,
    ),
    Tool(
        name="list_services",
        description="List all running Docker services",
        fn=list_services,
    ),
    # ... convert all 9 tools
]

Files to Modify:

  • src/agent/tools.py - Convert all 9 tools

Estimated Time: 1 hour


Phase 3: Agent Orchestrator Replacement

3.1 Create New ADK-Based Orchestrator

Key Changes:

  1. Model Initialization
# Replace ChatOllama with LiteLLM
from google.adk.models.lite_llm import LiteLlm

self.llm = LiteLlm(
    model="ollama_chat/mistral:7b",
    api_base="http://ollama:11434",
)
  1. Agent Creation
# Replace create_react_agent with ADK Agent
from google.adk.agents import Agent

self.agent = Agent(
    model=self.llm,
    name="tatlock",
    description="British butler assistant",
    instruction=get_prompt(variant),  # Reuse system prompts!
    tools=ALL_TOOLS,
)
  1. Streaming Interface
# Replace LangGraph astream with ADK streaming
async def chat(self, message: str, history: List[Dict]) -> AsyncIterator[Dict]:
    # Convert history to ADK format
    messages = self._build_messages(history, message)

    # Stream from ADK agent
    async for chunk in self.agent.run(messages, stream=True):
        # Map ADK events to our format
        yield self._map_chunk(chunk)

3.2 Event Mapping

ADK provides different event types than LangGraph:

  • tool_call_start → map to {"type": "tool_call"}
  • tool_call_end → map to {"type": "tool_result"}
  • content_delta → map to {"type": "content"}

Files to Modify:

  • src/agent/orchestrator.py - Complete rewrite (keep interface)

Estimated Time: 2 hours


Phase 4: Streaming Adapter

4.1 Update SSE Converter

The stream_agent_to_sse function should continue to work with minimal changes since we maintain the same intermediate format:

{"type": "tool_call", "tool": "...", "content": "..."}
{"type": "content", "content": "..."}

ADK streaming will provide similar events, just need to map them correctly in the orchestrator.

Files to Modify:

  • src/agent/streaming.py - Minor adjustments only

Estimated Time: 30 minutes


Phase 5: Integration & Testing

5.1 Update Controller

Minimal changes needed in ai_controller.py:

  • Import path changes (from src.agent import get_unified_agent)
  • Everything else stays the same

5.2 Testing Checklist

Create comprehensive tests:

# test_adk_integration.py

async def test_basic_chat():
    """Test agent responds without tools"""
    agent = get_unified_agent()
    response = await agent.chat_completion("Hello!")
    assert len(response) > 0

async def test_tool_calling():
    """Test agent calls get_current_time tool"""
    agent = get_unified_agent()
    response = await agent.chat_completion("What time is it?")
    # Should contain actual time, not hallucination
    assert "202" in response  # Year should be in response

async def test_web_search():
    """Test web_search tool (will fail with DuckDuckGo rate limits)"""
    agent = get_unified_agent()
    response = await agent.chat_completion("What is the weather in Amsterdam?")
    # Should attempt web search
    assert response  # At minimum, should respond

async def test_streaming():
    """Test streaming output"""
    agent = get_unified_agent()
    chunks = []
    async for chunk in agent.chat("List the services", stream=True):
        chunks.append(chunk)
    assert len(chunks) > 0
    assert any(c["type"] == "content" for c in chunks)

5.3 Model Testing Matrix

Test with multiple models to find best one:

Model Size ADK Compatible? Tool Calling? Notes
mistral:7b 4.4GB ✅ (proven) ✅ Current choice
gemma3:4b 3.3GB ✅ (docs) ✅ Better for VRAM
qwen3:14b ~8GB ✅ (docs) ✅ If VRAM allows
gemma3:12b 8.1GB ✅ (docs) ✅ High capability

Estimated Time: 1.5 hours


Phase 6: Deployment

6.1 Docker Rebuild

# Rebuild core-api service with new dependencies
docker-compose build core-api
docker-compose up -d core-api

6.2 Verification

  1. Check logs: docker logs core-api --tail 50
  2. Test health endpoint
  3. Test via webui with "What time is it?"
  4. Monitor for tool calls in logs

6.3 Rollback Plan

Keep LangChain implementation in a git branch:

git checkout -b backup/langchain-implementation
git add -A && git commit -m "Backup before ADK migration"
git checkout main
# ... perform migration ...
# If issues: git checkout backup/langchain-implementation

Estimated Time: 30 minutes


Risk Assessment & Mitigation

High Risks

Risk 1: ADK Tool Calling Issues with Certain Models

  • Evidence: GitHub issue #716 mentions "Ollama tool calling incompatible"
  • Mitigation: Test multiple models, have fallback to Mistral 7B
  • Impact: Could require model switching

Risk 2: Streaming Format Incompatibility

  • Evidence: ADK streaming might emit different event structures
  • Mitigation: Thorough mapping layer in orchestrator
  • Impact: Could affect frontend display

Medium Risks

Risk 3: LiteLLM Configuration

  • Evidence: Requires correct env vars and model naming
  • Mitigation: Follow documented examples exactly
  • Impact: Could cause initialization failures

Risk 4: Breaking Changes in Dependencies

  • Evidence: New framework, different API paradigms
  • Mitigation: Pin exact versions, test thoroughly
  • Impact: Could require API adjustments

Low Risks

Risk 5: Memory System Integration

  • Impact: Memory system is independent, should not be affected
  • Mitigation: Keep interface the same

Alternative Approaches Considered

  • Try different system prompts
  • Try different model configurations
  • Why Not: Already tried, fundamental compatibility issue

Option B: Pydantic AI (Alternative)

  • Pros: Simpler API, type safety
  • Cons: Less documentation for Ollama, newer than ADK
  • Verdict: ADK has better Ollama documentation

Option C: Custom Implementation

  • Pros: Full control
  • Cons: Reinventing wheel, more maintenance
  • Verdict: ADK provides everything needed

Success Criteria

Must Have (Required for Success)

  1. ✅ Agent calls tools when appropriate (no hallucination)
  2. ✅ get_current_time works reliably
  3. ✅ Streaming output preserved
  4. ✅ No regressions in memory system
  5. ✅ API endpoints unchanged

Should Have (Desired Outcomes)

  1. ✅ Works with Gemma 3 models (~4GB VRAM)
  2. ✅ Web search functional (when not rate-limited)
  3. ✅ Performance equivalent or better
  4. ✅ Clear logs showing tool calls

Nice to Have (Bonus)

  1. ✅ Support for multiple model providers
  2. ✅ Better error messages
  3. ✅ Reduced memory footprint

Timeline & Effort Estimate

Phase Time Dependencies
1. Dependencies 15 min None
2. Tool Migration 1 hour Phase 1
3. Orchestrator 2 hours Phase 2
4. Streaming 30 min Phase 3
5. Testing 1.5 hours Phase 4
6. Deployment 30 min Phase 5
Total 5.5 hours Sequential

Add 1 hour buffer for unexpected issues = 6.5 hours total


Post-Migration Tasks

  1. Documentation

    • Update README with ADK setup instructions
    • Document model compatibility matrix
    • Add troubleshooting guide
  2. Monitoring

    • Watch for tool calling failures
    • Monitor response quality
    • Track VRAM usage
  3. Optimization

    • Test alternative models for better VRAM efficiency
    • Tune system prompts for ADK
    • Consider caching strategies

Key Resources

Google ADK Documentation:

Integration Examples:

Package Documentation:


Open Questions

  1. Q: Which Ollama model works best with ADK tool calling?

    • A: Test Mistral 7B, Gemma 3:4b, Qwen 3:14b in Phase 5
  2. Q: Does ADK support custom SSE format for Open WebUI?

    • A: Yes, we maintain the mapping layer in streaming.py
  3. Q: Will system prompts need modification?

    • A: Likely minor tweaks, but current prompts should mostly work
  4. Q: What's the fallback if ADK also fails?

    • A: Consider Pydantic AI or custom implementation

Approval Checklist

Before proceeding with implementation:

  • User approves overall approach
  • User agrees with Google ADK choice
  • User confirms 6-hour timeline is acceptable
  • User reviews risk assessment
  • User approves model testing matrix
  • User confirms rollback plan is sufficient

Next Steps After Approval:

  1. Create feature branch: feature/migrate-to-adk
  2. Begin Phase 1: Dependencies
  3. Document progress in this file
  4. Request code review after Phase 5
  5. Deploy to production after testing

Plan prepared by Claude Code Ready for user review and approval