ai-flow improvement / add langchain
This commit is contained in:
@@ -0,0 +1,757 @@
|
||||
# Agent Architecture Flow Diagrams
|
||||
|
||||
**Date**: 2025-11-23
|
||||
**System**: Core API Unified Agent with LangGraph
|
||||
|
||||
This document shows the data flow through the agent system for various scenarios, including which models are used and how components interact.
|
||||
|
||||
---
|
||||
|
||||
## System Components Overview
|
||||
|
||||
```
|
||||
┌─────────────────────────────────────────────────────────────────┐
|
||||
│ Open WebUI │
|
||||
│ (or any OpenAI client) │
|
||||
└────────────────────────┬────────────────────────────────────────┘
|
||||
│ POST /v1/chat/completions
|
||||
│ {"use_agent": true/false}
|
||||
▼
|
||||
┌─────────────────────────────────────────────────────────────────┐
|
||||
│ Core API (FastAPI) │
|
||||
│ ┌──────────────────────────────────────────────────────────┐ │
|
||||
│ │ AI Controller (ai_controller.py) │ │
|
||||
│ │ • Routes to agent or direct LLM based on use_agent │ │
|
||||
│ │ • Converts OpenAI format ↔ agent format │ │
|
||||
│ └─────────┬────────────────────────────────────────┬───────┘ │
|
||||
│ │ use_agent=false │ │
|
||||
│ │ use_agent=true │ │
|
||||
└────────────┼────────────────────────────────────────┼───────────┘
|
||||
│ │
|
||||
▼ ▼
|
||||
┌────────────────┐ ┌──────────────────────┐
|
||||
│ Direct to │ │ Unified Agent │
|
||||
│ Ollama │ │ (orchestrator.py) │
|
||||
│ (any model) │ │ • LangGraph ReAct │
|
||||
└────────────────┘ │ • mistral:7b only │
|
||||
│ • Tool calling │
|
||||
└──────────┬───────────┘
|
||||
│
|
||||
┌──────────▼───────────┐
|
||||
│ Agent Tools │
|
||||
│ (tools.py) │
|
||||
│ • Infrastructure │
|
||||
│ • Web scraping │
|
||||
│ • Documentation │
|
||||
└──────────────────────┘
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Scenario 1: Simple Knowledge Prompt (No Tools Needed)
|
||||
|
||||
**User**: _"What is Docker?"_
|
||||
|
||||
```
|
||||
┌──────────┐
|
||||
│ User │ "What is Docker?"
|
||||
└────┬─────┘
|
||||
│ POST /v1/chat/completions
|
||||
│ use_agent: true
|
||||
▼
|
||||
┌────────────────────────────────────────────┐
|
||||
│ Core API - AI Controller │
|
||||
│ │
|
||||
│ 1. Parse request │
|
||||
│ 2. Check use_agent flag → TRUE │
|
||||
│ 3. Extract message & history │
|
||||
└────┬───────────────────────────────────────┘
|
||||
│
|
||||
▼
|
||||
┌────────────────────────────────────────────┐
|
||||
│ Unified Agent (orchestrator.py) │
|
||||
│ │
|
||||
│ Model: mistral:7b (tool-calling capable) │
|
||||
│ │
|
||||
│ System Prompt: │
|
||||
│ "You are a homelab assistant..." │
|
||||
│ │
|
||||
│ Available Tools: │
|
||||
│ - list_services │
|
||||
│ - web_search │
|
||||
│ - read_documentation │
|
||||
│ - ... [7 tools total] │
|
||||
└────┬───────────────────────────────────────┘
|
||||
│
|
||||
│ Agent reasoning:
|
||||
│ "This is general knowledge,
|
||||
│ no tools needed"
|
||||
│
|
||||
▼
|
||||
┌────────────────────────────────────────────┐
|
||||
│ LangGraph ReAct Loop │
|
||||
│ │
|
||||
│ [Thought] Analyzing query... │
|
||||
│ [Decision] Direct answer, no tools │
|
||||
│ [Action] Generate response │
|
||||
└────┬───────────────────────────────────────┘
|
||||
│
|
||||
▼
|
||||
┌────────────────────────────────────────────┐
|
||||
│ Ollama (mistral:7b) │
|
||||
│ │
|
||||
│ Generates: "Docker is a platform for │
|
||||
│ containerizing applications..." │
|
||||
└────┬───────────────────────────────────────┘
|
||||
│
|
||||
│ [💭 Analyzing...] (thinking)
|
||||
│ "Docker is a platform..." (content)
|
||||
│
|
||||
▼
|
||||
┌────────────────────────────────────────────┐
|
||||
│ Stream to SSE Format │
|
||||
│ (streaming.py) │
|
||||
│ │
|
||||
│ Converts to OpenAI SSE chunks: │
|
||||
│ data: {"choices":[{"delta":{"content":""}}]}│
|
||||
└────┬───────────────────────────────────────┘
|
||||
│
|
||||
▼
|
||||
┌──────────┐
|
||||
│ User │ Sees: [💭 Analyzing...] → response
|
||||
└──────────┘
|
||||
```
|
||||
|
||||
**Models Used**:
|
||||
- `mistral:7b` (agent reasoning + response generation)
|
||||
|
||||
**Data Flow**:
|
||||
1. Request → AI Controller
|
||||
2. AI Controller → Unified Agent
|
||||
3. Agent → mistral:7b (direct query, no tools)
|
||||
4. mistral:7b → Response text
|
||||
5. Agent → SSE formatter → User
|
||||
|
||||
---
|
||||
|
||||
## Scenario 2: Web Search Required
|
||||
|
||||
**User**: _"What's the weather in San Francisco?"_
|
||||
|
||||
```
|
||||
┌──────────┐
|
||||
│ User │ "What's the weather in SF?"
|
||||
└────┬─────┘
|
||||
│ use_agent: true
|
||||
▼
|
||||
┌────────────────────────────────────────────┐
|
||||
│ AI Controller │
|
||||
└────┬───────────────────────────────────────┘
|
||||
│
|
||||
▼
|
||||
┌────────────────────────────────────────────┐
|
||||
│ Unified Agent (mistral:7b) │
|
||||
│ │
|
||||
│ [Thought] Need real-time weather data │
|
||||
│ [Decision] Use web_search tool │
|
||||
│ [Action] Call web_search( │
|
||||
│ url="https://wttr.in/san-francisco" │
|
||||
│ ) │
|
||||
└────┬───────────────────────────────────────┘
|
||||
│
|
||||
│ Tool call
|
||||
│
|
||||
▼
|
||||
┌────────────────────────────────────────────┐
|
||||
│ Tool: web_search (tools.py) │
|
||||
│ │
|
||||
│ 1. Fetch URL via httpx │
|
||||
│ 2. Extract content (trafilatura) │
|
||||
│ 3. Return text content │
|
||||
└────┬───────────────────────────────────────┘
|
||||
│
|
||||
│ Tool result: "Current: 62°F, Cloudy..."
|
||||
│
|
||||
▼
|
||||
┌────────────────────────────────────────────┐
|
||||
│ Unified Agent (mistral:7b) │
|
||||
│ │
|
||||
│ [Observation] Got weather data │
|
||||
│ [Thought] Format for user │
|
||||
│ [Action] Generate final response │
|
||||
└────┬───────────────────────────────────────┘
|
||||
│
|
||||
▼
|
||||
┌────────────────────────────────────────────┐
|
||||
│ Ollama (mistral:7b) │
|
||||
│ │
|
||||
│ Generates: "The weather in San Francisco │
|
||||
│ is currently 62°F and cloudy..." │
|
||||
└────┬───────────────────────────────────────┘
|
||||
│
|
||||
│ SSE stream:
|
||||
│ [💭 Analyzing...] → [🔧 Searching web...] → [✓ Found data] → Response
|
||||
│
|
||||
▼
|
||||
┌──────────┐
|
||||
│ User │
|
||||
└──────────┘
|
||||
```
|
||||
|
||||
**Models Used**:
|
||||
- `mistral:7b` (agent reasoning, tool selection, response synthesis)
|
||||
|
||||
**Data Flow**:
|
||||
1. User → AI Controller → Agent
|
||||
2. Agent analyzes → Decides to use `web_search`
|
||||
3. Tool executes → Fetches web content
|
||||
4. Tool result → Back to agent
|
||||
5. Agent synthesizes → Final response
|
||||
6. Stream to user with status indicators
|
||||
|
||||
**Components Involved**:
|
||||
- AI Controller (routing)
|
||||
- Unified Agent (orchestration)
|
||||
- mistral:7b (reasoning at each step)
|
||||
- web_search tool (httpx + trafilatura)
|
||||
- SSE formatter (status indicators)
|
||||
|
||||
---
|
||||
|
||||
## Scenario 3: Code Generation from Swagger Docs
|
||||
|
||||
**User**: _"Write Python code to list all containers using the Core API"_
|
||||
|
||||
```
|
||||
┌──────────┐
|
||||
│ User │ "Write code to list containers"
|
||||
└────┬─────┘
|
||||
│
|
||||
▼
|
||||
┌────────────────────────────────────────────┐
|
||||
│ AI Controller │
|
||||
└────┬───────────────────────────────────────┘
|
||||
│
|
||||
▼
|
||||
┌────────────────────────────────────────────────────────────┐
|
||||
│ Unified Agent (mistral:7b) │
|
||||
│ │
|
||||
│ [Thought] Need API docs to write accurate code │
|
||||
│ [Decision] Use read_documentation tool │
|
||||
│ [Action] read_documentation("swagger") │
|
||||
└────┬───────────────────────────────────────────────────────┘
|
||||
│
|
||||
│ Tool call
|
||||
│
|
||||
▼
|
||||
┌────────────────────────────────────────────────────────────┐
|
||||
│ Tool: read_documentation (tools.py) │
|
||||
│ │
|
||||
│ 1. Reads /app/docs/openapi.json │
|
||||
│ 2. Searches for container-related endpoints │
|
||||
│ 3. Returns relevant API specs │
|
||||
└────┬───────────────────────────────────────────────────────┘
|
||||
│
|
||||
│ Returns: GET /infrastructure/containers endpoint spec
|
||||
│
|
||||
▼
|
||||
┌────────────────────────────────────────────────────────────┐
|
||||
│ Unified Agent (mistral:7b) │
|
||||
│ │
|
||||
│ [Observation] Found API endpoint details │
|
||||
│ [Thought] Need to generate Python code │
|
||||
│ [Decision] Could use code model for better quality │
|
||||
│ │
|
||||
│ ⚠️ Current: Uses mistral:7b for code generation │
|
||||
│ 🔮 Future: Could route to codestral:latest │
|
||||
└────┬───────────────────────────────────────────────────────┘
|
||||
│
|
||||
│ Generate code using API spec
|
||||
│
|
||||
▼
|
||||
┌────────────────────────────────────────────────────────────┐
|
||||
│ Ollama (mistral:7b) │
|
||||
│ │
|
||||
│ Synthesizes code based on: │
|
||||
│ - API documentation │
|
||||
│ - User request │
|
||||
│ - Python best practices │
|
||||
│ │
|
||||
│ Output: │
|
||||
│ ```python │
|
||||
│ import httpx │
|
||||
│ │
|
||||
│ async def list_containers(): │
|
||||
│ async with httpx.AsyncClient() as client: │
|
||||
│ response = await client.get( │
|
||||
│ "http://api.schweitz.net/infrastructure/..." │
|
||||
│ ) │
|
||||
│ return response.json() │
|
||||
│ ``` │
|
||||
└────┬───────────────────────────────────────────────────────┘
|
||||
│
|
||||
│ SSE stream:
|
||||
│ [💭 Analyzing...] → [🔧 Reading docs...] → [✓ Found API] → Code output
|
||||
│
|
||||
▼
|
||||
┌──────────┐
|
||||
│ User │
|
||||
└──────────┘
|
||||
```
|
||||
|
||||
**Models Used**:
|
||||
- `mistral:7b` (agent reasoning + code generation)
|
||||
- **Future enhancement**: Could route to `codestral:latest` for code generation
|
||||
|
||||
**Data Flow**:
|
||||
1. User → Agent
|
||||
2. Agent → read_documentation tool
|
||||
3. Tool → Reads OpenAPI spec from disk
|
||||
4. Spec → Back to agent
|
||||
5. Agent + spec → mistral:7b for code synthesis
|
||||
6. Code → Stream to user
|
||||
|
||||
**Potential Optimization**:
|
||||
```
|
||||
┌────────────────────────────────────────────┐
|
||||
│ Future: Model Routing │
|
||||
│ │
|
||||
│ Agent detects code generation request │
|
||||
│ ↓ │
|
||||
│ Routes to codestral:latest │
|
||||
│ (instead of mistral:7b) │
|
||||
│ ↓ │
|
||||
│ Better code quality │
|
||||
└────────────────────────────────────────────┘
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Scenario 4: Infrastructure Query
|
||||
|
||||
**User**: _"List all NPM proxy hosts and their domains"_
|
||||
|
||||
```
|
||||
┌──────────┐
|
||||
│ User │ "List NPM proxies and domains"
|
||||
└────┬─────┘
|
||||
│
|
||||
▼
|
||||
┌────────────────────────────────────────────┐
|
||||
│ AI Controller │
|
||||
└────┬───────────────────────────────────────┘
|
||||
│
|
||||
▼
|
||||
┌──────────────────────────────────────────────────────────────┐
|
||||
│ Unified Agent (mistral:7b) │
|
||||
│ │
|
||||
│ [Thought] User wants NPM proxy configuration │
|
||||
│ [Decision] Use list_domains tool │
|
||||
│ [Action] list_domains() │
|
||||
└────┬─────────────────────────────────────────────────────────┘
|
||||
│
|
||||
│ Tool call
|
||||
│
|
||||
▼
|
||||
┌──────────────────────────────────────────────────────────────┐
|
||||
│ Tool: list_domains (tools.py) │
|
||||
│ │
|
||||
│ 1. Calls get_npm_client() │
|
||||
│ 2. Makes request to NPM API: │
|
||||
│ GET http://npm:81/api/nginx/proxy-hosts │
|
||||
│ 3. Parses response │
|
||||
│ 4. Extracts domain names & forwards │
|
||||
└────┬─────────────────────────────────────────────────────────┘
|
||||
│
|
||||
│ Tool result:
|
||||
│ [
|
||||
│ {"domain": "home.schweitz.net", "forward": "organizr:80"},
|
||||
│ {"domain": "api.schweitz.net", "forward": "core-api:8083"},
|
||||
│ {"domain": "media.schweitz.net", "forward": "jellyfin:8096"},
|
||||
│ ...
|
||||
│ ]
|
||||
│
|
||||
▼
|
||||
┌──────────────────────────────────────────────────────────────┐
|
||||
│ Unified Agent (mistral:7b) │
|
||||
│ │
|
||||
│ [Observation] Got NPM proxy list │
|
||||
│ [Thought] Format nicely for user │
|
||||
│ [Action] Generate formatted response │
|
||||
└────┬─────────────────────────────────────────────────────────┘
|
||||
│
|
||||
▼
|
||||
┌──────────────────────────────────────────────────────────────┐
|
||||
│ Ollama (mistral:7b) │
|
||||
│ │
|
||||
│ Synthesizes response: │
|
||||
│ │
|
||||
│ "Here are your NPM proxy hosts: │
|
||||
│ │
|
||||
│ 1. home.schweitz.net → organizr:80 │
|
||||
│ 2. api.schweitz.net → core-api:8083 │
|
||||
│ 3. media.schweitz.net → jellyfin:8096 │
|
||||
│ ..." │
|
||||
└────┬─────────────────────────────────────────────────────────┘
|
||||
│
|
||||
│ SSE stream:
|
||||
│ [💭 Analyzing...] → [🔧 Querying NPM...] → [✓ Found 12 proxies] → Response
|
||||
│
|
||||
▼
|
||||
┌──────────┐
|
||||
│ User │
|
||||
└──────────┘
|
||||
|
||||
Data Path Detail:
|
||||
═══════════════════
|
||||
|
||||
User Request
|
||||
↓
|
||||
AI Controller
|
||||
↓
|
||||
Unified Agent (mistral:7b)
|
||||
↓
|
||||
list_domains tool
|
||||
↓
|
||||
NPM Client (npm_client.py)
|
||||
↓
|
||||
HTTP Request → NPM Container (nginx-proxy-manager:81)
|
||||
↓
|
||||
NPM API Response (JSON)
|
||||
↓
|
||||
Parsed data → Tool
|
||||
↓
|
||||
Tool result → Agent
|
||||
↓
|
||||
mistral:7b synthesizes
|
||||
↓
|
||||
Formatted response
|
||||
↓
|
||||
SSE Stream → User
|
||||
```
|
||||
|
||||
**Models Used**:
|
||||
- `mistral:7b` (all reasoning + synthesis)
|
||||
|
||||
**Components in Data Path**:
|
||||
1. **AI Controller** - Request routing
|
||||
2. **Unified Agent** - Orchestration & reasoning (mistral:7b)
|
||||
3. **list_domains Tool** - Business logic wrapper
|
||||
4. **NPM Client** - HTTP client to NPM API
|
||||
5. **NPM Container** - Actual nginx proxy manager
|
||||
6. **SSE Formatter** - Stream status indicators
|
||||
|
||||
**External Systems**:
|
||||
- Nginx Proxy Manager API (port 81)
|
||||
|
||||
---
|
||||
|
||||
## Scenario 5: Multi-Tool Complex Query
|
||||
|
||||
**User**: _"Which services are unhealthy and need to be restarted?"_
|
||||
|
||||
```
|
||||
┌──────────┐
|
||||
│ User │ "Which services unhealthy?"
|
||||
└────┬─────┘
|
||||
│
|
||||
▼
|
||||
┌────────────────────────────────────────────────────────────────┐
|
||||
│ Unified Agent (mistral:7b) - Multi-step reasoning │
|
||||
│ │
|
||||
│ STEP 1: [Thought] Need to check all services │
|
||||
│ [Decision] Use list_services tool │
|
||||
│ [Action] list_services() │
|
||||
└────┬───────────────────────────────────────────────────────────┘
|
||||
│
|
||||
▼
|
||||
┌────────────────────────────────────────────────────────────────┐
|
||||
│ Tool: list_services → Portainer API │
|
||||
│ │
|
||||
│ Returns: [ │
|
||||
│ {"name": "core-api", "status": "running"}, │
|
||||
│ {"name": "jellyfin", "status": "running"}, │
|
||||
│ {"name": "uptime-kuma", "status": "running"}, │
|
||||
│ ... │
|
||||
│ ] │
|
||||
└────┬───────────────────────────────────────────────────────────┘
|
||||
│
|
||||
│ Result → Agent
|
||||
│
|
||||
▼
|
||||
┌────────────────────────────────────────────────────────────────┐
|
||||
│ Unified Agent (mistral:7b) │
|
||||
│ │
|
||||
│ STEP 2: [Observation] All services show "running" │
|
||||
│ [Thought] Need health check details from monitoring │
|
||||
│ [Decision] Use check_service_health for each │
|
||||
│ [Action] Loop through services │
|
||||
└────┬───────────────────────────────────────────────────────────┘
|
||||
│
|
||||
▼
|
||||
┌────────────────────────────────────────────────────────────────┐
|
||||
│ Tool: check_service_health (for each service) │
|
||||
│ │
|
||||
│ check_service_health("core-api") │
|
||||
│ → Uptime Kuma API → {"status": "up", "ping": "23ms"} │
|
||||
│ │
|
||||
│ check_service_health("jellyfin") │
|
||||
│ → Uptime Kuma API → {"status": "down", "ping": "timeout"} │
|
||||
│ │
|
||||
│ check_service_health("uptime-kuma") │
|
||||
│ → Uptime Kuma API → {"status": "up", "ping": "5ms"} │
|
||||
└────┬───────────────────────────────────────────────────────────┘
|
||||
│
|
||||
│ Results → Agent
|
||||
│
|
||||
▼
|
||||
┌────────────────────────────────────────────────────────────────┐
|
||||
│ Unified Agent (mistral:7b) │
|
||||
│ │
|
||||
│ STEP 3: [Observation] Jellyfin is down! │
|
||||
│ [Thought] User asked which need restarting │
|
||||
│ [Decision] Report findings │
|
||||
│ [Action] Generate response with recommendation │
|
||||
└────┬───────────────────────────────────────────────────────────┘
|
||||
│
|
||||
▼
|
||||
┌────────────────────────────────────────────────────────────────┐
|
||||
│ Ollama (mistral:7b) - Final synthesis │
|
||||
│ │
|
||||
│ "Based on health checks, Jellyfin (media.schweitz.net) is │
|
||||
│ currently unhealthy and not responding to health probes. │
|
||||
│ │
|
||||
│ Recommendation: Restart the jellyfin service. │
|
||||
│ │
|
||||
│ Would you like me to restart it for you?" │
|
||||
└────┬───────────────────────────────────────────────────────────┘
|
||||
│
|
||||
│ SSE stream with multiple status updates:
|
||||
│ [💭 Analyzing...]
|
||||
│ → [🔧 Listing services...]
|
||||
│ → [✓ Found 15 services]
|
||||
│ → [🔧 Checking health...]
|
||||
│ → [✓ Checked 15 monitors]
|
||||
│ → Response
|
||||
│
|
||||
▼
|
||||
┌──────────┐
|
||||
│ User │
|
||||
└──────────┘
|
||||
|
||||
Multi-Tool Flow:
|
||||
═══════════════
|
||||
|
||||
┌─────────────────┐
|
||||
│ Agent Reasoning │
|
||||
│ (mistral:7b) │
|
||||
└────┬────────────┘
|
||||
│
|
||||
┌────▼─────────────────────────────────┐
|
||||
│ ReAct Loop (LangGraph) │
|
||||
│ │
|
||||
│ Thought → Action → Observation │
|
||||
│ ↓ ↓ ↑ │
|
||||
│ Analyze Execute Process │
|
||||
│ Tool Result │
|
||||
└──────────────────────────────────────┘
|
||||
│
|
||||
┌────▼────┐ ┌────▼────┐ ┌────▼────┐
|
||||
│ Tool 1 │ │ Tool 2 │ │ Tool 3 │
|
||||
│ list_ │ │ check_ │ │ check_ │
|
||||
│services │ │ health │ │ health │
|
||||
│ │ │ (x15) │ │ ... │
|
||||
└─────────┘ └─────────┘ └─────────┘
|
||||
│ │ │
|
||||
┌────▼────────────▼────────────▼────┐
|
||||
│ External Systems │
|
||||
│ • Portainer API │
|
||||
│ • Uptime Kuma API │
|
||||
└───────────────────────────────────┘
|
||||
```
|
||||
|
||||
**Models Used**:
|
||||
- `mistral:7b` (all reasoning, tool orchestration, synthesis)
|
||||
|
||||
**Tool Call Sequence**:
|
||||
1. `list_services()` → Portainer → 15 services
|
||||
2. Loop: `check_service_health(service)` × 15 → Uptime Kuma
|
||||
3. Analyze results → Identify unhealthy
|
||||
4. Synthesize recommendation
|
||||
|
||||
**Why Single Model Works**:
|
||||
- mistral:7b maintains context across tool calls
|
||||
- LangGraph manages the ReAct loop state
|
||||
- Agent "thinks" between each tool call
|
||||
- No model switching needed for multi-step reasoning
|
||||
|
||||
---
|
||||
|
||||
## Model Selection Summary
|
||||
|
||||
### Current Implementation:
|
||||
|
||||
| Scenario | Model Used | Reason |
|
||||
|----------|-----------|--------|
|
||||
| **Agent mode** (any query) | `mistral:7b` | Supports tool calling |
|
||||
| **Direct chat** (use_agent=false) | User's choice | gemma:2b, gemma:7b, etc. |
|
||||
| **Embeddings** | `nomic-embed-text` (via Ollama) | No local PyTorch needed |
|
||||
|
||||
### Why mistral:7b for Agent?
|
||||
|
||||
✅ **Supports tool calling** - Gemma/Gemma2 do not
|
||||
✅ **Good reasoning** - Handles multi-step logic
|
||||
✅ **Fast enough** - 7B parameters, ~2-5s responses
|
||||
✅ **Available locally** - Already in Ollama
|
||||
|
||||
### Future Enhancements:
|
||||
|
||||
```
|
||||
┌────────────────────────────────────────────┐
|
||||
│ Potential Model Routing │
|
||||
│ │
|
||||
│ Task Type → Model │
|
||||
│ ──────────────────────────────────── │
|
||||
│ General reasoning → mistral:7b │
|
||||
│ Code generation → codestral:latest │
|
||||
│ Fast queries → gemma:2b │
|
||||
│ Complex analysis → mixtral:8x7b │
|
||||
│ Embeddings → nomic-embed-text │
|
||||
└────────────────────────────────────────────┘
|
||||
```
|
||||
|
||||
Could implement model routing in agent:
|
||||
- Detect task type (code vs general vs analysis)
|
||||
- Route to specialized model
|
||||
- Return to mistral:7b for synthesis
|
||||
|
||||
---
|
||||
|
||||
## Component Communication Matrix
|
||||
|
||||
```
|
||||
Core API Components
|
||||
═══════════════════
|
||||
|
||||
┌─────────────┬──────────┬────────┬────────┬─────────┐
|
||||
│ Component │ Mistral │ Ollama │ Tools │ External│
|
||||
│ │ :7b │ API │ │ APIs │
|
||||
├─────────────┼──────────┼────────┼────────┼─────────┤
|
||||
│ AI │ │ ✓ │ │ │
|
||||
│ Controller │ Routes │ Direct │ │ │
|
||||
│ │ │ call │ │ │
|
||||
├─────────────┼──────────┼────────┼────────┼─────────┤
|
||||
│ Unified │ ✓ │ ✓ │ ✓ │ │
|
||||
│ Agent │ Reasoning│ LLM │ Calls │ │
|
||||
│ │ │ invoke │ │ │
|
||||
├─────────────┼──────────┼────────┼────────┼─────────┤
|
||||
│ Tools │ │ │ │ ✓ │
|
||||
│ │ │ │ │ Portainer│
|
||||
│ │ │ │ │ NPM, Kuma│
|
||||
├─────────────┼──────────┼────────┼────────┼─────────┤
|
||||
│ SSE │ │ │ ✓ │ │
|
||||
│ Formatter │ │ │ Status │ │
|
||||
│ │ │ │ events │ │
|
||||
└─────────────┴──────────┴────────┴────────┴─────────┘
|
||||
|
||||
Legend:
|
||||
═══════
|
||||
✓ = Direct communication
|
||||
Routes = Decision point, passes through
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Performance Characteristics
|
||||
|
||||
### Response Times (Typical):
|
||||
|
||||
| Scenario | Time to First Token | Total Time | Model Calls |
|
||||
|----------|---------------------|------------|-------------|
|
||||
| **Knowledge query** | ~500ms | 2-3s | 1 (mistral:7b) |
|
||||
| **Single tool use** | ~500ms | 4-6s | 2 (reasoning + synthesis) |
|
||||
| **Multi-tool query** | ~500ms | 8-15s | 3+ (reasoning per tool + synthesis) |
|
||||
| **Code generation** | ~500ms | 5-10s | 2 (read docs + generate) |
|
||||
|
||||
### Streaming Benefits:
|
||||
|
||||
```
|
||||
Without Streaming:
|
||||
User waits → → → [silence] → → → Full response
|
||||
|
||||
With Streaming:
|
||||
User sees → [💭 Thinking] → [🔧 Tool use] → [✓ Done] → Response chunks
|
||||
↑ 500ms ↑ 2s ↑ 4s
|
||||
```
|
||||
|
||||
User perceives faster response due to immediate feedback!
|
||||
|
||||
---
|
||||
|
||||
## Key Architectural Decisions
|
||||
|
||||
### ✅ Single Agent Model (mistral:7b)
|
||||
**Pro**: Maintains context across tool calls, simpler architecture
|
||||
**Con**: Can't leverage specialized models for specific tasks
|
||||
|
||||
### ✅ Ollama-Based Embeddings
|
||||
**Pro**: No local PyTorch (~2GB saved), flexible model switching
|
||||
**Con**: Network dependency on Ollama service
|
||||
|
||||
### ✅ OpenAI-Compatible API
|
||||
**Pro**: Works with any OpenAI client, easy integration
|
||||
**Con**: Must convert between formats
|
||||
|
||||
### ✅ Tool-Based Architecture
|
||||
**Pro**: Extensible, clear separation of concerns
|
||||
**Con**: Each tool call adds latency
|
||||
|
||||
### ✅ Streaming with Status Indicators
|
||||
**Pro**: Transparent reasoning, better UX
|
||||
**Con**: More complex implementation
|
||||
|
||||
---
|
||||
|
||||
## Future Optimizations
|
||||
|
||||
### 1. Model Routing
|
||||
Add intelligence to route requests to specialized models:
|
||||
- Code → `codestral:latest`
|
||||
- Analysis → `mixtral:8x7b`
|
||||
- Fast queries → `gemma:2b`
|
||||
|
||||
### 2. Tool Result Caching
|
||||
Cache frequently-accessed infrastructure data:
|
||||
- Service list (60s TTL)
|
||||
- Domain list (5min TTL)
|
||||
- Reduces tool call latency
|
||||
|
||||
### 3. Parallel Tool Execution
|
||||
When independent tools needed:
|
||||
```python
|
||||
results = await asyncio.gather(
|
||||
check_service_health("service1"),
|
||||
check_service_health("service2"),
|
||||
check_service_health("service3"),
|
||||
)
|
||||
```
|
||||
Reduces 3×2s = 6s to ~2s
|
||||
|
||||
### 4. Smaller Agent Model
|
||||
Try `gemma2:9b` or `qwen2.5:7b` if they support tools:
|
||||
- Potentially faster inference
|
||||
- Lower memory usage
|
||||
|
||||
---
|
||||
|
||||
## Conclusion
|
||||
|
||||
The unified agent architecture successfully:
|
||||
- ✅ Routes all requests through single intelligent orchestrator
|
||||
- ✅ Uses `mistral:7b` for tool-calling capability
|
||||
- ✅ Maintains transparent reasoning via streaming
|
||||
- ✅ Integrates with existing infrastructure (Portainer, NPM, Kuma)
|
||||
- ✅ Works with any OpenAI-compatible client
|
||||
- ✅ Saves ~2GB memory by using Ollama embeddings
|
||||
|
||||
Next steps: Test with Open WebUI and document usage for end users.
|
||||
Reference in New Issue
Block a user