ai-flow improvement / add langchain

This commit is contained in:
2025-11-23 14:51:19 +01:00
parent ade84f34d5
commit 5e734ad27f
25 changed files with 4867 additions and 51 deletions
+757
View File
@@ -0,0 +1,757 @@
# Agent Architecture Flow Diagrams
**Date**: 2025-11-23
**System**: Core API Unified Agent with LangGraph
This document shows the data flow through the agent system for various scenarios, including which models are used and how components interact.
---
## System Components Overview
```
┌─────────────────────────────────────────────────────────────────┐
│ Open WebUI │
│ (or any OpenAI client) │
└────────────────────────┬────────────────────────────────────────┘
│ POST /v1/chat/completions
│ {"use_agent": true/false}
▼
┌─────────────────────────────────────────────────────────────────┐
│ Core API (FastAPI) │
│ ┌──────────────────────────────────────────────────────────┐ │
│ │ AI Controller (ai_controller.py) │ │
│ │ • Routes to agent or direct LLM based on use_agent │ │
│ │ • Converts OpenAI format ↔ agent format │ │
│ └─────────┬────────────────────────────────────────┬───────┘ │
│ │ use_agent=false │ │
│ │ use_agent=true │ │
└────────────┼────────────────────────────────────────┼───────────┘
│ │
▼ ▼
┌────────────────┐ ┌──────────────────────┐
│ Direct to │ │ Unified Agent │
│ Ollama │ │ (orchestrator.py) │
│ (any model) │ │ • LangGraph ReAct │
└────────────────┘ │ • mistral:7b only │
│ • Tool calling │
└──────────┬───────────┘
│
┌──────────▼───────────┐
│ Agent Tools │
│ (tools.py) │
│ • Infrastructure │
│ • Web scraping │
│ • Documentation │
└──────────────────────┘
```
---
## Scenario 1: Simple Knowledge Prompt (No Tools Needed)
**User**: _"What is Docker?"_
```
┌──────────┐
│ User │ "What is Docker?"
└────┬─────┘
│ POST /v1/chat/completions
│ use_agent: true
▼
┌────────────────────────────────────────────┐
│ Core API - AI Controller │
│ │
│ 1. Parse request │
│ 2. Check use_agent flag → TRUE │
│ 3. Extract message & history │
└────┬───────────────────────────────────────┘
│
▼
┌────────────────────────────────────────────┐
│ Unified Agent (orchestrator.py) │
│ │
│ Model: mistral:7b (tool-calling capable) │
│ │
│ System Prompt: │
│ "You are a homelab assistant..." │
│ │
│ Available Tools: │
│ - list_services │
│ - web_search │
│ - read_documentation │
│ - ... [7 tools total] │
└────┬───────────────────────────────────────┘
│
│ Agent reasoning:
│ "This is general knowledge,
│ no tools needed"
│
▼
┌────────────────────────────────────────────┐
│ LangGraph ReAct Loop │
│ │
│ [Thought] Analyzing query... │
│ [Decision] Direct answer, no tools │
│ [Action] Generate response │
└────┬───────────────────────────────────────┘
│
▼
┌────────────────────────────────────────────┐
│ Ollama (mistral:7b) │
│ │
│ Generates: "Docker is a platform for │
│ containerizing applications..." │
└────┬───────────────────────────────────────┘
│
│ [💭 Analyzing...] (thinking)
│ "Docker is a platform..." (content)
│
▼
┌────────────────────────────────────────────┐
│ Stream to SSE Format │
│ (streaming.py) │
│ │
│ Converts to OpenAI SSE chunks: │
│ data: {"choices":[{"delta":{"content":""}}]}│
└────┬───────────────────────────────────────┘
│
▼
┌──────────┐
│ User │ Sees: [💭 Analyzing...] → response
└──────────┘
```
**Models Used**:
- `mistral:7b` (agent reasoning + response generation)
**Data Flow**:
1. Request → AI Controller
2. AI Controller → Unified Agent
3. Agent → mistral:7b (direct query, no tools)
4. mistral:7b → Response text
5. Agent → SSE formatter → User
---
## Scenario 2: Web Search Required
**User**: _"What's the weather in San Francisco?"_
```
┌──────────┐
│ User │ "What's the weather in SF?"
└────┬─────┘
│ use_agent: true
▼
┌────────────────────────────────────────────┐
│ AI Controller │
└────┬───────────────────────────────────────┘
│
▼
┌────────────────────────────────────────────┐
│ Unified Agent (mistral:7b) │
│ │
│ [Thought] Need real-time weather data │
│ [Decision] Use web_search tool │
│ [Action] Call web_search( │
│ url="https://wttr.in/san-francisco" │
│ ) │
└────┬───────────────────────────────────────┘
│
│ Tool call
│
▼
┌────────────────────────────────────────────┐
│ Tool: web_search (tools.py) │
│ │
│ 1. Fetch URL via httpx │
│ 2. Extract content (trafilatura) │
│ 3. Return text content │
└────┬───────────────────────────────────────┘
│
│ Tool result: "Current: 62°F, Cloudy..."
│
▼
┌────────────────────────────────────────────┐
│ Unified Agent (mistral:7b) │
│ │
│ [Observation] Got weather data │
│ [Thought] Format for user │
│ [Action] Generate final response │
└────┬───────────────────────────────────────┘
│
▼
┌────────────────────────────────────────────┐
│ Ollama (mistral:7b) │
│ │
│ Generates: "The weather in San Francisco │
│ is currently 62°F and cloudy..." │
└────┬───────────────────────────────────────┘
│
│ SSE stream:
│ [💭 Analyzing...] → [🔧 Searching web...] → [✓ Found data] → Response
│
▼
┌──────────┐
│ User │
└──────────┘
```
**Models Used**:
- `mistral:7b` (agent reasoning, tool selection, response synthesis)
**Data Flow**:
1. User → AI Controller → Agent
2. Agent analyzes → Decides to use `web_search`
3. Tool executes → Fetches web content
4. Tool result → Back to agent
5. Agent synthesizes → Final response
6. Stream to user with status indicators
**Components Involved**:
- AI Controller (routing)
- Unified Agent (orchestration)
- mistral:7b (reasoning at each step)
- web_search tool (httpx + trafilatura)
- SSE formatter (status indicators)
---
## Scenario 3: Code Generation from Swagger Docs
**User**: _"Write Python code to list all containers using the Core API"_
```
┌──────────┐
│ User │ "Write code to list containers"
└────┬─────┘
│
▼
┌────────────────────────────────────────────┐
│ AI Controller │
└────┬───────────────────────────────────────┘
│
▼
┌────────────────────────────────────────────────────────────┐
│ Unified Agent (mistral:7b) │
│ │
│ [Thought] Need API docs to write accurate code │
│ [Decision] Use read_documentation tool │
│ [Action] read_documentation("swagger") │
└────┬───────────────────────────────────────────────────────┘
│
│ Tool call
│
▼
┌────────────────────────────────────────────────────────────┐
│ Tool: read_documentation (tools.py) │
│ │
│ 1. Reads /app/docs/openapi.json │
│ 2. Searches for container-related endpoints │
│ 3. Returns relevant API specs │
└────┬───────────────────────────────────────────────────────┘
│
│ Returns: GET /infrastructure/containers endpoint spec
│
▼
┌────────────────────────────────────────────────────────────┐
│ Unified Agent (mistral:7b) │
│ │
│ [Observation] Found API endpoint details │
│ [Thought] Need to generate Python code │
│ [Decision] Could use code model for better quality │
│ │
│ ⚠️ Current: Uses mistral:7b for code generation │
│ 🔮 Future: Could route to codestral:latest │
└────┬───────────────────────────────────────────────────────┘
│
│ Generate code using API spec
│
▼
┌────────────────────────────────────────────────────────────┐
│ Ollama (mistral:7b) │
│ │
│ Synthesizes code based on: │
│ - API documentation │
│ - User request │
│ - Python best practices │
│ │
│ Output: │
│ ```python │
│ import httpx │
│ │
│ async def list_containers(): │
│ async with httpx.AsyncClient() as client: │
│ response = await client.get( │
│ "http://api.schweitz.net/infrastructure/..." │
│ ) │
│ return response.json() │
│ ``` │
└────┬───────────────────────────────────────────────────────┘
│
│ SSE stream:
│ [💭 Analyzing...] → [🔧 Reading docs...] → [✓ Found API] → Code output
│
▼
┌──────────┐
│ User │
└──────────┘
```
**Models Used**:
- `mistral:7b` (agent reasoning + code generation)
- **Future enhancement**: Could route to `codestral:latest` for code generation
**Data Flow**:
1. User → Agent
2. Agent → read_documentation tool
3. Tool → Reads OpenAPI spec from disk
4. Spec → Back to agent
5. Agent + spec → mistral:7b for code synthesis
6. Code → Stream to user
**Potential Optimization**:
```
┌────────────────────────────────────────────┐
│ Future: Model Routing │
│ │
│ Agent detects code generation request │
│ ↓ │
│ Routes to codestral:latest │
│ (instead of mistral:7b) │
│ ↓ │
│ Better code quality │
└────────────────────────────────────────────┘
```
---
## Scenario 4: Infrastructure Query
**User**: _"List all NPM proxy hosts and their domains"_
```
┌──────────┐
│ User │ "List NPM proxies and domains"
└────┬─────┘
│
▼
┌────────────────────────────────────────────┐
│ AI Controller │
└────┬───────────────────────────────────────┘
│
▼
┌──────────────────────────────────────────────────────────────┐
│ Unified Agent (mistral:7b) │
│ │
│ [Thought] User wants NPM proxy configuration │
│ [Decision] Use list_domains tool │
│ [Action] list_domains() │
└────┬─────────────────────────────────────────────────────────┘
│
│ Tool call
│
▼
┌──────────────────────────────────────────────────────────────┐
│ Tool: list_domains (tools.py) │
│ │
│ 1. Calls get_npm_client() │
│ 2. Makes request to NPM API: │
│ GET http://npm:81/api/nginx/proxy-hosts │
│ 3. Parses response │
│ 4. Extracts domain names & forwards │
└────┬─────────────────────────────────────────────────────────┘
│
│ Tool result:
│ [
│ {"domain": "home.schweitz.net", "forward": "organizr:80"},
│ {"domain": "api.schweitz.net", "forward": "core-api:8083"},
│ {"domain": "media.schweitz.net", "forward": "jellyfin:8096"},
│ ...
│ ]
│
▼
┌──────────────────────────────────────────────────────────────┐
│ Unified Agent (mistral:7b) │
│ │
│ [Observation] Got NPM proxy list │
│ [Thought] Format nicely for user │
│ [Action] Generate formatted response │
└────┬─────────────────────────────────────────────────────────┘
│
▼
┌──────────────────────────────────────────────────────────────┐
│ Ollama (mistral:7b) │
│ │
│ Synthesizes response: │
│ │
│ "Here are your NPM proxy hosts: │
│ │
│ 1. home.schweitz.net → organizr:80 │
│ 2. api.schweitz.net → core-api:8083 │
│ 3. media.schweitz.net → jellyfin:8096 │
│ ..." │
└────┬─────────────────────────────────────────────────────────┘
│
│ SSE stream:
│ [💭 Analyzing...] → [🔧 Querying NPM...] → [✓ Found 12 proxies] → Response
│
▼
┌──────────┐
│ User │
└──────────┘
Data Path Detail:
═══════════════════
User Request
↓
AI Controller
↓
Unified Agent (mistral:7b)
↓
list_domains tool
↓
NPM Client (npm_client.py)
↓
HTTP Request → NPM Container (nginx-proxy-manager:81)
↓
NPM API Response (JSON)
↓
Parsed data → Tool
↓
Tool result → Agent
↓
mistral:7b synthesizes
↓
Formatted response
↓
SSE Stream → User
```
**Models Used**:
- `mistral:7b` (all reasoning + synthesis)
**Components in Data Path**:
1. **AI Controller** - Request routing
2. **Unified Agent** - Orchestration & reasoning (mistral:7b)
3. **list_domains Tool** - Business logic wrapper
4. **NPM Client** - HTTP client to NPM API
5. **NPM Container** - Actual nginx proxy manager
6. **SSE Formatter** - Stream status indicators
**External Systems**:
- Nginx Proxy Manager API (port 81)
---
## Scenario 5: Multi-Tool Complex Query
**User**: _"Which services are unhealthy and need to be restarted?"_
```
┌──────────┐
│ User │ "Which services unhealthy?"
└────┬─────┘
│
▼
┌────────────────────────────────────────────────────────────────┐
│ Unified Agent (mistral:7b) - Multi-step reasoning │
│ │
│ STEP 1: [Thought] Need to check all services │
│ [Decision] Use list_services tool │
│ [Action] list_services() │
└────┬───────────────────────────────────────────────────────────┘
│
▼
┌────────────────────────────────────────────────────────────────┐
│ Tool: list_services → Portainer API │
│ │
│ Returns: [ │
│ {"name": "core-api", "status": "running"}, │
│ {"name": "jellyfin", "status": "running"}, │
│ {"name": "uptime-kuma", "status": "running"}, │
│ ... │
│ ] │
└────┬───────────────────────────────────────────────────────────┘
│
│ Result → Agent
│
▼
┌────────────────────────────────────────────────────────────────┐
│ Unified Agent (mistral:7b) │
│ │
│ STEP 2: [Observation] All services show "running" │
│ [Thought] Need health check details from monitoring │
│ [Decision] Use check_service_health for each │
│ [Action] Loop through services │
└────┬───────────────────────────────────────────────────────────┘
│
▼
┌────────────────────────────────────────────────────────────────┐
│ Tool: check_service_health (for each service) │
│ │
│ check_service_health("core-api") │
│ → Uptime Kuma API → {"status": "up", "ping": "23ms"} │
│ │
│ check_service_health("jellyfin") │
│ → Uptime Kuma API → {"status": "down", "ping": "timeout"} │
│ │
│ check_service_health("uptime-kuma") │
│ → Uptime Kuma API → {"status": "up", "ping": "5ms"} │
└────┬───────────────────────────────────────────────────────────┘
│
│ Results → Agent
│
▼
┌────────────────────────────────────────────────────────────────┐
│ Unified Agent (mistral:7b) │
│ │
│ STEP 3: [Observation] Jellyfin is down! │
│ [Thought] User asked which need restarting │
│ [Decision] Report findings │
│ [Action] Generate response with recommendation │
└────┬───────────────────────────────────────────────────────────┘
│
▼
┌────────────────────────────────────────────────────────────────┐
│ Ollama (mistral:7b) - Final synthesis │
│ │
│ "Based on health checks, Jellyfin (media.schweitz.net) is │
│ currently unhealthy and not responding to health probes. │
│ │
│ Recommendation: Restart the jellyfin service. │
│ │
│ Would you like me to restart it for you?" │
└────┬───────────────────────────────────────────────────────────┘
│
│ SSE stream with multiple status updates:
│ [💭 Analyzing...]
│ → [🔧 Listing services...]
│ → [✓ Found 15 services]
│ → [🔧 Checking health...]
│ → [✓ Checked 15 monitors]
│ → Response
│
▼
┌──────────┐
│ User │
└──────────┘
Multi-Tool Flow:
═══════════════
┌─────────────────┐
│ Agent Reasoning │
│ (mistral:7b) │
└────┬────────────┘
│
┌────▼─────────────────────────────────┐
│ ReAct Loop (LangGraph) │
│ │
│ Thought → Action → Observation │
│ ↓ ↓ ↑ │
│ Analyze Execute Process │
│ Tool Result │
└──────────────────────────────────────┘
│
┌────▼────┐ ┌────▼────┐ ┌────▼────┐
│ Tool 1 │ │ Tool 2 │ │ Tool 3 │
│ list_ │ │ check_ │ │ check_ │
│services │ │ health │ │ health │
│ │ │ (x15) │ │ ... │
└─────────┘ └─────────┘ └─────────┘
│ │ │
┌────▼────────────▼────────────▼────┐
│ External Systems │
│ • Portainer API │
│ • Uptime Kuma API │
└───────────────────────────────────┘
```
**Models Used**:
- `mistral:7b` (all reasoning, tool orchestration, synthesis)
**Tool Call Sequence**:
1. `list_services()` → Portainer → 15 services
2. Loop: `check_service_health(service)` × 15 → Uptime Kuma
3. Analyze results → Identify unhealthy
4. Synthesize recommendation
**Why Single Model Works**:
- mistral:7b maintains context across tool calls
- LangGraph manages the ReAct loop state
- Agent "thinks" between each tool call
- No model switching needed for multi-step reasoning
---
## Model Selection Summary
### Current Implementation:
| Scenario | Model Used | Reason |
|----------|-----------|--------|
| **Agent mode** (any query) | `mistral:7b` | Supports tool calling |
| **Direct chat** (use_agent=false) | User's choice | gemma:2b, gemma:7b, etc. |
| **Embeddings** | `nomic-embed-text` (via Ollama) | No local PyTorch needed |
### Why mistral:7b for Agent?
✅ **Supports tool calling** - Gemma/Gemma2 do not
✅ **Good reasoning** - Handles multi-step logic
✅ **Fast enough** - 7B parameters, ~2-5s responses
✅ **Available locally** - Already in Ollama
### Future Enhancements:
```
┌────────────────────────────────────────────┐
│ Potential Model Routing │
│ │
│ Task Type → Model │
│ ──────────────────────────────────── │
│ General reasoning → mistral:7b │
│ Code generation → codestral:latest │
│ Fast queries → gemma:2b │
│ Complex analysis → mixtral:8x7b │
│ Embeddings → nomic-embed-text │
└────────────────────────────────────────────┘
```
Could implement model routing in agent:
- Detect task type (code vs general vs analysis)
- Route to specialized model
- Return to mistral:7b for synthesis
---
## Component Communication Matrix
```
Core API Components
═══════════════════
┌─────────────┬──────────┬────────┬────────┬─────────┐
│ Component │ Mistral │ Ollama │ Tools │ External│
│ │ :7b │ API │ │ APIs │
├─────────────┼──────────┼────────┼────────┼─────────┤
│ AI │ │ ✓ │ │ │
│ Controller │ Routes │ Direct │ │ │
│ │ │ call │ │ │
├─────────────┼──────────┼────────┼────────┼─────────┤
│ Unified │ ✓ │ ✓ │ ✓ │ │
│ Agent │ Reasoning│ LLM │ Calls │ │
│ │ │ invoke │ │ │
├─────────────┼──────────┼────────┼────────┼─────────┤
│ Tools │ │ │ │ ✓ │
│ │ │ │ │ Portainer│
│ │ │ │ │ NPM, Kuma│
├─────────────┼──────────┼────────┼────────┼─────────┤
│ SSE │ │ │ ✓ │ │
│ Formatter │ │ │ Status │ │
│ │ │ │ events │ │
└─────────────┴──────────┴────────┴────────┴─────────┘
Legend:
═══════
✓ = Direct communication
Routes = Decision point, passes through
```
---
## Performance Characteristics
### Response Times (Typical):
| Scenario | Time to First Token | Total Time | Model Calls |
|----------|---------------------|------------|-------------|
| **Knowledge query** | ~500ms | 2-3s | 1 (mistral:7b) |
| **Single tool use** | ~500ms | 4-6s | 2 (reasoning + synthesis) |
| **Multi-tool query** | ~500ms | 8-15s | 3+ (reasoning per tool + synthesis) |
| **Code generation** | ~500ms | 5-10s | 2 (read docs + generate) |
### Streaming Benefits:
```
Without Streaming:
User waits → → → [silence] → → → Full response
With Streaming:
User sees → [💭 Thinking] → [🔧 Tool use] → [✓ Done] → Response chunks
↑ 500ms ↑ 2s ↑ 4s
```
User perceives faster response due to immediate feedback!
---
## Key Architectural Decisions
### ✅ Single Agent Model (mistral:7b)
**Pro**: Maintains context across tool calls, simpler architecture
**Con**: Can't leverage specialized models for specific tasks
### ✅ Ollama-Based Embeddings
**Pro**: No local PyTorch (~2GB saved), flexible model switching
**Con**: Network dependency on Ollama service
### ✅ OpenAI-Compatible API
**Pro**: Works with any OpenAI client, easy integration
**Con**: Must convert between formats
### ✅ Tool-Based Architecture
**Pro**: Extensible, clear separation of concerns
**Con**: Each tool call adds latency
### ✅ Streaming with Status Indicators
**Pro**: Transparent reasoning, better UX
**Con**: More complex implementation
---
## Future Optimizations
### 1. Model Routing
Add intelligence to route requests to specialized models:
- Code → `codestral:latest`
- Analysis → `mixtral:8x7b`
- Fast queries → `gemma:2b`
### 2. Tool Result Caching
Cache frequently-accessed infrastructure data:
- Service list (60s TTL)
- Domain list (5min TTL)
- Reduces tool call latency
### 3. Parallel Tool Execution
When independent tools needed:
```python
results = await asyncio.gather(
check_service_health("service1"),
check_service_health("service2"),
check_service_health("service3"),
)
```
Reduces 3×2s = 6s to ~2s
### 4. Smaller Agent Model
Try `gemma2:9b` or `qwen2.5:7b` if they support tools:
- Potentially faster inference
- Lower memory usage
---
## Conclusion
The unified agent architecture successfully:
- ✅ Routes all requests through single intelligent orchestrator
- ✅ Uses `mistral:7b` for tool-calling capability
- ✅ Maintains transparent reasoning via streaming
- ✅ Integrates with existing infrastructure (Portainer, NPM, Kuma)
- ✅ Works with any OpenAI-compatible client
- ✅ Saves ~2GB memory by using Ollama embeddings
Next steps: Test with Open WebUI and document usage for end users.