feat(ai): complete ADK migration and optimize system health checks

Major architectural changes and improvements:

## ADK Framework Migration (v0.10.0)
- Migrated from LangChain/LangGraph to Google ADK 1.3.0 with LiteLLM 1.80.5
- Improved tool calling reliability with local Ollama models
- Converted all 10 tools to ADK async generator format
- Updated streaming pipeline for ADK event system
- Enhanced error handling and agent initialization

## Model Optimization
- Switched from gemma3:12b (10GB VRAM) to gemma3:4b (4.8GB VRAM)
- Reduced VRAM usage from 91% to 43% (5.4GB freed)
- Optimized for production stability with memory headroom

## Health Check System Overhaul
- Optimized /health/full: 6ms response (was 30s+)
- Added model verification: confirms configured model is available
- New /health/diagnostics endpoint with optional deep testing
- Added currently loaded models tracking
- Clear emoji status indicators (✅/❌/⚠️)
- Fixed AGENT_AVAILABLE flag export for proper health reporting

## Ollama Client Enhancements
- Added list_models() method for model inventory
- Enhanced model verification in health checks
- Better error handling and reporting

## Documentation Updates
- Updated STATUS.md to v0.10.0-adk-migration
- Comprehensive CHANGELOG.md entry with migration details
- Updated PLANS.md showing Phase 4 complete
- Updated ai-orchestrator-plan.md with ADK status
- Added MIGRATION_PLAN_LANGCHAIN_TO_ADK.md
- Added ADK_Ollama_Research.md with implementation analysis

## Technical Details
- 10 tools: 7 infrastructure + 2 research + 1 response tool
- Framework: Google ADK with UnifiedAgent pattern
- System prompt: v7_adk_best_practice
- Container health: Now passing Docker healthchecks
- Response times: Simple queries ~0.3-1s, Research ~4-7s
This commit is contained in:
2025-11-26 08:36:50 +01:00
parent bfc58b03ba
commit e3b451b7b0
15 changed files with 1731 additions and 275 deletions
+66 -13
View File
@@ -1,12 +1,12 @@
# Project Status
> **Last Updated:** 2025-11-23
> **Version:** 0.8.2-authentik-api-protection
> **Last Updated:** 2025-11-26
> **Version:** 0.10.0-adk-migration
## Current Phase
**Active Work:** Security & SSO Implementation (Authentik Deployment)
**Status:** 🔄 **IN PROGRESS** - 2 Services Protected (Organizr + Core API)
**Active Work:** AI Infrastructure Optimization & System Hardening
**Status:** ✅ **STABLE** - ADK Migration Complete, All Systems Operational
See [PLANS.md](PLANS.md) for complete implementation roadmap and [CHANGELOG.md](CHANGELOG.md) for version history.
@@ -81,14 +81,67 @@ See [PLANS.md](PLANS.md) for complete implementation roadmap and [CHANGELOG.md](
- [ ] Replace ad-hoc shell scripts in `/stacks` with API endpoints
- [ ] Add CLI wrapper for common operations
### Priority 3: AI Orchestrator Phase 2 (Memory Systems) - DEFERRED
- [ ] Implement Tier 1: ConversationBufferMemory (in-memory, last 10 turns)
- [ ] Implement Tier 2: ConversationSummaryMemory (SQLite summaries)
- [ ] Integrate Tier 3: VectorStoreRetrieverMemory (Qdrant semantic search)
- [ ] Create Qdrant collections (conversation_memory, documents, user_facts)
- [ ] Implement memory consolidation service
- [ ] Add conversation history API endpoints
- [ ] Test memory persistence across container restarts
### Priority 3: AI Orchestrator Phase 2 (Memory Systems) ✅ COMPLETE
- [x] Implement Tier 1: ConversationBufferMemory (in-memory, last 10 turns)
- [x] Implement Tier 2/3: Unified Qdrant storage (persistent + semantic search)
- [x] Create Qdrant collection (core_api_conversations with 768d nomic-embed-text)
- [x] Implement auto-consolidation service (triggers at 10 turns)
- [x] Add memory persistence across container restarts
- [x] Implement dual-retrieval (buffer + Qdrant)
- [x] **Phase 2.5: Multi-Tenancy** (user_id isolation with default "llm-testuser")
**Implementation Details:**
- **Tier 1 (Buffer):** In-memory storage for last 10 turns (< 1ms access)
- **Tier 2/3 (Qdrant):** Unified persistent storage + semantic search (768d embeddings)
- **Auto-Consolidation:** Automatically moves buffer → Qdrant at 10 turns
- **Multi-Tenancy:** Single collection with user_id filtering (default: "llm-testuser")
- **Embedding Model:** nomic-embed-text (768 dimensions, via Ollama)
- **Memory Retrieval:** Dual-check buffer + Qdrant for cross-restart persistence
- **Status:** 32 points stored, tested with multiple users, recall working after restarts
### Priority 4: AI Orchestrator - Framework Migration ✅ COMPLETE (2025-11-26)
- [x] **Phase 3:** Research Capabilities (DuckDuckGo, web scraping) - COMPLETE (2025-11-24)
- [x] **Framework Migration:** LangChain/LangGraph → Google ADK - COMPLETE (2025-11-26)
- [x] Migrate agent orchestrator to Google ADK with LiteLLM
- [x] Convert all 9 tools to ADK async generator format
- [x] Update streaming pipeline for ADK event format
- [x] Switch model to gemma3:12b with ADK-optimized prompts
- [x] Implement comprehensive agent health checks
- [x] Update requirements.txt (remove langchain*, add google-adk)
- [x] Production testing and validation
**Current Implementation (as of 2025-11-26):**
- **Framework:** Google ADK 1.3.0 with LiteLLM 1.80.5 (migrated from LangChain)
- **Model:** gemma3:12b (upgraded from mistral:7b)
- **System Prompt:** v7_adk_best_practice (optimized for ADK)
- **Total Tool Count:** 9 tools (7 infrastructure + 2 research)
- **Tools Format:** ADK async generators with proper streaming support
- **Architecture:** UnifiedAgent with stateless sessions
**Tools Available:**
- **Infrastructure (7):** get_current_time, list_services, get_service_details, list_stacks, get_stack_details, list_npm_hosts, get_npm_host_details
- **Research (2):** web_search (DuckDuckGo + auto-scrape), web_scrape (targeted extraction)
**Migration Benefits Achieved:**
- ✅ Improved tool calling reliability with Ollama models
- ✅ Better streaming support with ADK event system
- ✅ Model flexibility (works with Gemma, Mistral, Qwen families)
- ✅ Cleaner architecture with unified agent pattern
- ✅ Production-ready health monitoring
**Performance Metrics (Post-Migration):**
- **Simple queries:** ~0.3-1s response time
- **Tool-using queries:** ~2-5s response time
- **Research queries:** ~4-7s response time
- **Tool calling success rate:** Monitoring in progress
- **VRAM usage:** ~8GB with gemma3:12b
**Optional Future Enhancements (deferred):**
- Multi-agent routing patterns (Phase 4+)
- Code specialist agent with codestral (Phase 4+)
- Time-based memory consolidation
- User filtering in Qdrant queries
- User management API endpoints
## Current Blockers
@@ -105,7 +158,7 @@ See [PLANS.md](PLANS.md) for complete implementation roadmap and [CHANGELOG.md](
| **Remote Access** | Working | Ready | 🟢 Headscale + NPM |
| **Firewall Active** | Yes | Yes | 🟢 UFW Configured |
| **Backups Configured** | Yes | Yes | 🟢 Daily @ 3 AM |
| **AI Orchestrator** | Phase 6 | Phase 1 ✅ | 🟡 Phase 2 Deferred |
| **AI Orchestrator** | Phase 6 | Phase 4 ✅ | 🟢 ADK Migration Complete |
| **SSO (Authentik)** | Phase 5 | Core Complete ✅ | 🟢 Organizr + Core API Protected |
## Quick Reference