feat(ai): complete ADK migration and optimize system health checks

Major architectural changes and improvements:

## ADK Framework Migration (v0.10.0)
- Migrated from LangChain/LangGraph to Google ADK 1.3.0 with LiteLLM 1.80.5
- Improved tool calling reliability with local Ollama models
- Converted all 10 tools to ADK async generator format
- Updated streaming pipeline for ADK event system
- Enhanced error handling and agent initialization

## Model Optimization
- Switched from gemma3:12b (10GB VRAM) to gemma3:4b (4.8GB VRAM)
- Reduced VRAM usage from 91% to 43% (5.4GB freed)
- Optimized for production stability with memory headroom

## Health Check System Overhaul
- Optimized /health/full: 6ms response (was 30s+)
- Added model verification: confirms configured model is available
- New /health/diagnostics endpoint with optional deep testing
- Added currently loaded models tracking
- Clear emoji status indicators (✅/❌/⚠️)
- Fixed AGENT_AVAILABLE flag export for proper health reporting

## Ollama Client Enhancements
- Added list_models() method for model inventory
- Enhanced model verification in health checks
- Better error handling and reporting

## Documentation Updates
- Updated STATUS.md to v0.10.0-adk-migration
- Comprehensive CHANGELOG.md entry with migration details
- Updated PLANS.md showing Phase 4 complete
- Updated ai-orchestrator-plan.md with ADK status
- Added MIGRATION_PLAN_LANGCHAIN_TO_ADK.md
- Added ADK_Ollama_Research.md with implementation analysis

## Technical Details
- 10 tools: 7 infrastructure + 2 research + 1 response tool
- Framework: Google ADK with UnifiedAgent pattern
- System prompt: v7_adk_best_practice
- Container health: Now passing Docker healthchecks
- Response times: Simple queries ~0.3-1s, Research ~4-7s
This commit is contained in:
2025-11-26 08:36:50 +01:00
parent bfc58b03ba
commit e3b451b7b0
15 changed files with 1731 additions and 275 deletions
+143 -46
View File
@@ -2,13 +2,16 @@
> **Project:** tower-of-joy AI Stack Enhancement
> **Created:** 2025-11-13
> **Status:** Phase 1 Complete ✅ - Phase 2 In Progress 🔄
> **Updated:** 2025-11-13
> **Target Completion:** 5 weeks remaining (Phase 2-6)
> **Status:** Phase 4 Complete ✅ - ADK Migration Successful
> **Updated:** 2025-11-26
> **Framework:** Google ADK 1.3.0 with LiteLLM 1.80.5
## Executive Summary
This document outlines the plan to build a sophisticated AI orchestration layer using LangGraph and FastAPI that will replace Open WebUI's direct connection to Ollama. The new architecture provides:
This document outlines the implementation of a sophisticated AI orchestration layer using Google ADK and FastAPI that provides an advanced agent system for the tower-of-joy homelab. The architecture provides:
**✅ MIGRATION COMPLETE (2025-11-26):**
Successfully migrated from LangChain/LangGraph to Google ADK to achieve reliable tool calling with Ollama local models. See [MIGRATION_PLAN_LANGCHAIN_TO_ADK.md](../../MIGRATION_PLAN_LANGCHAIN_TO_ADK.md) for details.
- **Advanced Memory Systems:** Three-tier memory with Qdrant for long-term semantic recall
- **Multi-Agent Workflows:** Intelligent routing to lightweight, heavy, and specialist models
@@ -118,23 +121,25 @@ This document outlines the plan to build a sophisticated AI orchestration layer
## Technology Stack
### Core Framework
- **LangGraph 0.2.60** - Stateful multi-agent orchestration (not basic LangChain)
### Core Framework (Current - Post-Migration)
- **Google ADK 1.3.0** - Agent Development Kit for stateful agents
- **LiteLLM 1.80.5** - Unified LLM interface for Ollama compatibility
- **FastAPI 0.115.0** - REST API framework
- **Uvicorn 0.32.0** - ASGI server
- **Pydantic 2.10.4** - Request/response validation
- **Uvicorn ≥0.34.0** - ASGI server
- **Pydantic ≥2.11.1** - Request/response validation
### AI & Memory
- **langchain 0.3.12** - Base framework
- **langchain-community 0.3.12** - Community integrations
- **qdrant-client 1.12.1** - Vector database client
- **langchain-qdrant 0.2.0** - LangChain + Qdrant integration
- **google-genai 1.17.0** - Google Generative AI SDK
- **google-cloud-aiplatform 1.95.1** - AI Platform integration
- **qdrant-client ~1.11.0** - Vector database client
- **Ollama** - Local LLM inference (via LiteLLM)
### Utilities
- **httpx 0.28.1** - Async HTTP client for external APIs
- **python-dotenv 1.0.1** - Environment configuration
- **structlog** - Structured logging
- **prometheus-client** - Metrics and monitoring
- **httpx ≥0.28.0** - Async HTTP client for external APIs
- **python-dotenv ~1.0.0** - Environment configuration
- **duckduckgo-search ~4.1.0** - Web search integration
- **beautifulsoup4 ~4.12.0** - Web scraping
- **trafilatura ~1.12.0** - Content extraction
### Container
- **Python 3.12** - Runtime (already upgraded)
@@ -563,37 +568,124 @@ class TaskManagementTool(BaseTool):
- Semantic search returns appropriate results
- No memory leaks or unbounded growth
### Phase 3: Multi-Agent Workflows (Week 3)
**Goal:** LangGraph-based agent system with intelligent routing
### Phase 3: Research Capabilities ✅ **COMPLETE** (2025-11-24)
**Goal:** Web search and research workflows
**Tasks:**
1. Install and configure LangGraph
2. Implement Router Agent (analyzes intent, routes requests)
3. Implement Chat Agent (general conversation)
4. Implement Research Agent (multi-step web research)
5. Implement Code Agent (programming specialist)
6. Add agent state management (LangGraph StateGraph)
7. Add supervisor pattern for agent coordination
8. Implement agent selection logic
9. Add agent switching mid-conversation
10. Test complex multi-step workflows
**Status:** ✅ **COMPLETE** - All success criteria met (LangChain implementation)
**Deliverables:**
- Working multi-agent system
- Intelligent request routing
- Specialist agent delegation
- Agent state persistence
- Multi-step workflow support
**Completed Tasks:**
1. ✅ Integrated DuckDuckGo web search API
2. ✅ Implemented automatic content scraping from search results
3. ✅ Added web_search tool (search + scrape in one call)
4. ✅ Added web_scrape tool (targeted URL extraction)
5. ✅ Enhanced system prompts for research query detection
6. ✅ Added enhanced progress indicators (🔍, 📄 icons)
7. ✅ Comprehensive testing (6/6 tests passed, 100% success rate)
8. ✅ Tool invocation logging for debugging
9. ✅ System prompt A/B testing (6 variants, v1_verbose winner)
10. ✅ Model validation (mistral:7b confirmed best for tools)
**Success Criteria:**
- Simple queries use lightweight models
- Complex tasks routed to heavy models
- Research tasks trigger multi-step workflows
- Code questions use specialist models
- Agent handoff works seamlessly
**Deliverables:** ✅ **ALL DELIVERED**
- ✅ Web search with DuckDuckGo integration
- ✅ Automatic content extraction from results
- ✅ Research query detection in system prompts
- ✅ 9 total tools (7 infrastructure + 2 research)
- ✅ Enhanced user experience with visual feedback
- ✅ Comprehensive test suite with 100% success
- ✅ Tool logging infrastructure
### Phase 4: Tool Integration (Week 4)
**Goal:** External API and tool calling capabilities
**Success Criteria:** ✅ **ALL MET**
- ✅ Research detection accuracy: 100% (target: >80%)
- ✅ Average response time: 5.3s (target: <10s)
- ✅ Source citation rate: 100% (target: >90%)
- ✅ Tool calling reliability: 100% on complex queries
- ✅ No regression in existing functionality
**Implementation Details:**
- **Approach:** Extended unified agent (Option A from Phase 3 plan)
- **Framework:** LangChain/LangGraph ReAct agent
- **Model:** mistral:7b (validated via extensive testing)
- **Dependencies:** duckduckgo-search~=4.1.0
- **Performance:** 4-7s for research, 0.3s for simple chat
**Testing & Optimization:**
- System Prompt A/B Testing: 6 variants tested across 30 scenarios
- Winner: v1_verbose (87/100 score, 0.92s avg response)
- Identified placeholder text issue (models mimic examples)
- Tool logging added for production debugging
**Deferred to Future Phases:**
- Multi-agent routing patterns (Phase 4/5)
- Code specialist agent with codestral (Phase 4/5)
- Model switching based on complexity (Phase 4/5)
- Supervisor pattern for agent coordination (Phase 5)
**See Also:**
- [Phase 3 Completion Document](../completed/phase3-multi-agent-workflows-complete.md)
- [System Prompt Test Results](../../services/core-api/COMPREHENSIVE_PROMPT_TEST_RESULTS.md)
- [Lightweight Model Testing](../../docs/sessions/2025-11-24-lightweight-model-testing.md)
### Phase 4: Framework Migration ✅ **COMPLETE** (2025-11-26)
**Goal:** Migrate from LangChain/LangGraph to Google ADK for improved reliability
**Status:** ✅ **COMPLETE** - All success criteria met
**Completed Tasks:**
1. ✅ Updated requirements.txt (removed langchain*, added google-adk, litellm)
2. ✅ Rewrote orchestrator.py to use ADK Agent with LiteLLM
3. ✅ Converted all 9 tools to ADK async generator format
4. ✅ Updated streaming.py for ADK event format
5. ✅ Upgraded model from mistral:7b to gemma3:12b
6. ✅ Created v7_adk_best_practice system prompt variant
7. ✅ Enhanced health checks for agent monitoring
8. ✅ Production testing and validation
9. ✅ Documentation updates
**Deliverables:** ✅ **ALL DELIVERED**
- ✅ Google ADK 1.3.0 integration with LiteLLM 1.80.5
- ✅ 9 tools migrated to ADK format (7 infrastructure + 2 research)
- ✅ Model upgrade to gemma3:12b (~8GB VRAM)
- ✅ ADK-optimized system prompts
- ✅ Improved streaming consistency
- ✅ Enhanced agent health monitoring
- ✅ Complete documentation
**Success Criteria:** ✅ **ALL MET**
- ✅ Tool calling works reliably (no more empty tool_calls arrays)
- ✅ Gemma models now supported (was failing with LangChain)
- ✅ Streaming output consistent and clean
- ✅ No regressions in memory system
- ✅ API endpoints unchanged (backward compatible)
- ✅ Performance within targets
**Implementation Details:**
- **Framework:** Google ADK with UnifiedAgent pattern
- **Model Bridge:** LiteLLM for Ollama compatibility
- **Model:** gemma3:12b (upgraded from mistral:7b)
- **Prompt:** v7_adk_best_practice
- **Tools:** All 9 tools as ADK async generators
- **Migration Time:** ~6 hours (as estimated)
**Performance (Post-Migration):**
- Simple queries: ~0.3-1s
- Tool-using queries: ~2-5s (improved from LangChain)
- Research queries: ~4-7s (maintained)
- VRAM usage: ~8GB with gemma3:12b
**Why This Migration:**
- ❌ **Problem:** LangGraph's `create_react_agent` failed to trigger tools with Ollama
- ❌ **Problem:** Gemma models returned status 400 with LangChain
- ❌ **Problem:** Inconsistent streaming behavior
- ✅ **Solution:** ADK has proven Ollama integration via LiteLLM
- ✅ **Result:** Reliable tool calling across all models
**See Also:**
- [Migration Plan](../../MIGRATION_PLAN_LANGCHAIN_TO_ADK.md)
- [ADK/Ollama Research](../../docs/ADK_Ollama_Research.md)
- [Changelog Entry](../../CHANGELOG.md#0100-adk-migration---2025-11-26)
### Phase 5: Tool Integration & RAG Optimization (Future)
**Goal:** Enhanced tool calling and advanced RAG capabilities
**Tasks:**
1. Create LangChain tool interface base class
@@ -1259,6 +1351,11 @@ This establishes the tower-of-joy project as a cutting-edge AI homelab with capa
---
**Plan Status:** ✅ Phase 1 Complete - 🔄 Phase 2 In Progress
**Completed:** Phase 1 - Foundation (2025-11-13)
**Next Step:** Phase 2 - Memory Systems (Week 2)
**Plan Status:** ✅ Phase 4 Complete - ADK Migration Successful
**Completed Phases:**
- Phase 1: Foundation (2025-11-13)
- Phase 2: Memory Systems (2025-11-23)
- Phase 3: Research Capabilities (2025-11-24)
- Phase 4: Framework Migration to ADK (2025-11-26)
**Next Steps:** Phase 5 - Multi-Agent Patterns & RAG Optimization (Future)