Files
portainer-core/docs/sessions/2025-11-24-vram-optimization-strategy.md
T
jpmschweitzer 0c2c838766 feat(ai): complete Phase 2/3 documentation and memory system improvements
Phase completion and enhancement updates:

## Documentation Added
- Phase 2 completion: Memory system implementation details
- Phase 3 completion: Research capabilities and tool integration
- Session documentation: Model testing, VRAM optimization analysis
- Test results: Comprehensive prompt testing (v1_verbose: 87/100)
- Tool logging implementation guide

## System Prompts
- Added prompts.py with 7 tested variants for A/B testing
- v1_verbose, v2_concise, v3_imperative, v4_minimal, etc.
- Comprehensive testing results for each variant
- Production-ready prompt selection guidance

## Memory System Enhancements
- Multi-tenancy support: Added user_id parameter throughout
- System message filtering: Don't store system messages in history
- Improved conversation turn tracking with user isolation
- Enhanced memory manager for better multi-user support

## AI Controller Improvements
- Better memory integration with user_id support
- Enhanced error handling for memory operations
- Improved token tracking for usage monitoring
- Skip system message storage (part of agent state)

## Portainer Client
- Comprehensive API client (148 lines)
- Stack management and service monitoring
- Container operations with full error handling
- Async support for all operations

## Architecture Documentation
- Updated agent flow diagrams for ADK architecture
- Enhanced core-api README with current setup
- Updated Docker compose stack configuration
- Complete testing and validation documentation
2025-11-26 08:41:44 +01:00

392 lines
9.7 KiB
Markdown

# VRAM Optimization Strategy for Model Orchestration
**Date**: 2025-11-24
**Context**: Multi-model architecture with always-loaded orchestrator + expert models
**Hardware**: RTX 2080 Ti (11GB VRAM)
## Problem Statement
**Goal**: Keep orchestrator model always loaded in VRAM to prevent cold starts, while maximizing VRAM availability for expert models.
**Current State**:
- Orchestrator: `mistral:7b` (5.1GB VRAM)
- Free VRAM: 4.7GB
- Use case: Orchestrator decides → routes to expert models (codestral, etc.)
**Challenge**: mistral:7b consumes 45% of available VRAM, limiting expert model options.
## VRAM Budget Analysis
### Current Configuration
```
Total VRAM: 11.0 GB
├─ mistral:7b: 5.1 GB (46.4%) - Orchestrator
├─ Overhead: 1.2 GB (10.6%) - System/Ollama
└─ Available: 4.7 GB (42.7%) - For expert models
```
### Desired Configuration
```
Total VRAM: 11.0 GB
├─ Orchestrator: ??? GB (minimize)
├─ Expert Model: ??? GB (maximize)
└─ Overhead: 1.2 GB
```
## Solution Options
### Option 1: Accept gemma3-tools:1b Limitations ⚠️
**VRAM Savings**: 3.8GB (5.1GB → 1.3GB)
```
Orchestrator: gemma3-tools:1b (1.3GB)
Free for experts: 8.5GB
```
**Pros**:
- ✅ Massive VRAM savings (74% reduction)
- ✅ Leaves 8.5GB for expert models
- ✅ Can load codestral:22b (full size) + orchestrator simultaneously
**Cons**:
- ❌ 33% tool calling reliability
- ❌ Wrong tool selection
- ❌ Erratic responses (raw JSON output)
- ❌ Poor user experience
**Verdict**: ❌ **Not recommended** - Unreliability hurts more than VRAM savings help
---
### Option 2: Use Smaller Quantization of mistral:7b ✅ RECOMMENDED
Ollama supports multiple quantization levels. You're currently using Q4_K_M, but Q2 or Q3 exist.
**Available Quantizations**:
- Q2_K: ~2.5GB VRAM (70% quality retention, aggressive)
- Q3_K_M: ~3.2GB VRAM (80% quality, good balance)
- Q4_K_M: ~5.1GB VRAM (90% quality, current)
- Q5_K_M: ~6.2GB VRAM (95% quality)
- Q8: ~7.7GB VRAM (99% quality, near full precision)
**Recommended**: Pull `mistral:7b-instruct-q3_K_M`
```bash
# Pull lower quantization
ollama pull mistral:7b-instruct-q3_K_M
# Update Core API config
# services/core-api/.env
AGENT_MODEL=mistral:7b-instruct-q3_K_M
```
**New VRAM Budget**:
```
Orchestrator: mistral:7b Q3_K_M (3.2GB)
Free for experts: 6.6GB
Savings: 1.9GB (37% reduction)
```
**Pros**:
- ✅ 100% tool calling compatibility (same model architecture)
- ✅ 1.9GB VRAM savings
- ✅ Minimal quality loss (80% of full precision is fine for routing)
- ✅ Proven reliability maintained
**Cons**:
- ⚠️ Slightly lower response quality (acceptable for orchestration)
- ⚠️ May need testing to verify tool calling still works
**Verdict**: ✅ **Best option** - Balanced approach
---
### Option 3: Hybrid Orchestrator (Simple Router + mistral:7b) 🔮 ADVANCED
Use a **two-tier routing system**:
1. **Lightweight classifier** (gemma3-tools:1b) - Always loaded
2. **Full orchestrator** (mistral:7b) - Loaded on demand for complex queries
**Architecture**:
```python
# Tier 1: Fast classifier (always loaded)
if query_is_simple(message):
# Direct routing: "list services" → list_services tool
# Load time: 0ms (always in VRAM)
use_simple_router(gemma3-tools:1b)
else:
# Complex routing: multi-tool, reasoning needed
# Load time: ~2s (load mistral:7b)
use_full_orchestrator(mistral:7b)
```
**VRAM Budget**:
```
Tier 1 (always): gemma3-tools:1b (1.3GB)
Tier 2 (on-demand): mistral:7b (5.1GB, loaded when needed)
Free when Tier 1 only: 8.5GB
Free when both loaded: 3.4GB
```
**Pros**:
- ✅ 8.5GB free for expert models most of the time
- ✅ Only loads mistral:7b when truly needed
- ✅ Simple queries stay fast (no model swap)
**Cons**:
- ❌ Complex implementation (need query classifier)
- ❌ 2s latency spike when switching to Tier 2
- ❌ More failure modes (what if Tier 1 misclassifies?)
**Verdict**: 🔮 **Future enhancement** - Interesting but complex
---
### Option 4: Use Different Base Model 🔍 RESEARCH NEEDED
Look for other tool-capable models with better size/quality trade-offs.
**Candidates to research**:
- `qwen2.5:7b-instruct-q3` - Alibaba's model, claimed good tool support
- `llama3.2:3b-instruct` - Meta's latest, check if tool-capable
- `hermes3:3b` - Nous Research, specifically trained for function calling
**Action**: Test these if available in Ollama registry.
---
## Recommended Implementation: Option 2
### Step 1: Pull Q3 Quantization
```bash
# Check if Q3 variant exists
ollama list | grep mistral
# Pull Q3 quantization (if available)
ollama pull mistral:7b-instruct-q3_K_M
# OR manually create Q3 from modelfile
cat > /tmp/mistral-q3.Modelfile << 'EOF'
FROM mistral:7b
PARAMETER quantization Q3_K_M
EOF
ollama create mistral:7b-q3 -f /tmp/mistral-q3.Modelfile
```
### Step 2: Test Tool Calling with Q3
```bash
# Run our test script with Q3 variant
source .venv/bin/activate
python3 << 'PYEOF'
import asyncio
import httpx
async def test():
payload = {
"model": "mistral:7b-q3",
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "List all services"}
],
"tools": [{
"type": "function",
"function": {
"name": "list_services",
"description": "List all running services",
"parameters": {"type": "object", "properties": {}}
}
}]
}
async with httpx.AsyncClient(timeout=60) as client:
r = await client.post("http://localhost:11434/api/chat", json=payload)
data = r.json()
message = data.get("message", {})
if "tool_calls" in message:
print("✅ Q3 quantization: Tool calling WORKS")
print(f" Called: {message['tool_calls'][0]['function']['name']}")
else:
print("❌ Q3 quantization: Tool calling BROKEN")
print(f" Response: {message.get('content', '')[:100]}")
asyncio.run(test())
PYEOF
```
### Step 3: Update Core API Configuration
```bash
# services/core-api/.env
AGENT_MODEL=mistral:7b-q3
```
```bash
# Restart core-api to pick up new model
docker restart core-api
```
### Step 4: Verify VRAM Usage
```bash
# Check new VRAM allocation
curl -s http://localhost:11434/api/ps | jq '.models[] | {name, size_vram_gb: (.size_vram / 1024 / 1024 / 1024)}'
```
**Expected Result**:
```json
{
"name": "mistral:7b-q3",
"size_vram_gb": 3.2
}
```
---
## Expert Model Strategy
With ~6.6GB available after Q3 orchestrator, you can now fit:
### Option A: Single Large Expert
```
Orchestrator: mistral:7b-q3 (3.2GB)
Expert: codestral:22b-q2 (6GB)
Total: 9.2GB / 11GB
```
### Option B: Multiple Smaller Experts
```
Orchestrator: mistral:7b-q3 (3.2GB)
Expert 1: codegemma:7b (4GB) - Code generation
Expert 2: gemma3:4b (3GB) - General knowledge
Total: 10.2GB / 11GB (near full capacity)
```
### Option C: Dynamic Loading (Current Behavior)
```
Orchestrator: mistral:7b-q3 (3.2GB) - Always loaded
Expert: Load on demand (6.6GB available)
- codestral for code
- gemma3:12b for general
- Model swaps as needed
```
---
## Advanced: Ollama Keep Alive Configuration
Control how long models stay in VRAM:
```bash
# Keep orchestrator always loaded (never unload)
curl -X POST http://localhost:11434/api/generate \
-d '{
"model": "mistral:7b-q3",
"keep_alive": -1,
"prompt": "warm up"
}'
# Expert models: unload after 5 minutes idle
curl -X POST http://localhost:11434/api/generate \
-d '{
"model": "codestral:latest",
"keep_alive": "5m",
"prompt": "warm up"
}'
```
**Configuration in Core API**:
```python
# services/core-api/src/agent/orchestrator.py
self.llm = ChatOllama(
model=self.settings.agent_model, # mistral:7b-q3
base_url=self.settings.ollama_base_url,
temperature=0.7,
keep_alive=-1, # Never unload orchestrator
)
```
---
## Testing Checklist
Before switching to Q3 quantization:
- [ ] Pull or create Q3 variant
- [ ] Test tool calling functionality
- [ ] Test tool selection accuracy (list_services vs get_service_details)
- [ ] Test multi-tool workflows
- [ ] Compare response quality vs Q4
- [ ] Verify VRAM usage reduction
- [ ] Test with Open WebUI
- [ ] Monitor for any degradation
If Q3 shows issues:
- Try Q4_K_S (slightly smaller than Q4_K_M)
- Fall back to current Q4_K_M if necessary
---
## Alternative Models Research
If mistral Q3 proves insufficient, test these:
### qwen2.5:7b (Alibaba Cloud)
- Similar size to mistral
- Claimed excellent tool calling
- May have Q3/Q4 variants available
```bash
ollama pull qwen2.5:7b-instruct
# Test with our tool calling script
```
### hermes3:3b (Nous Research)
- Specifically trained for function calling
- 3B parameters (smaller than mistral)
- Check Ollama availability
```bash
ollama search hermes3
```
---
## Summary
**Immediate Action**: Pull `mistral:7b` with Q3_K_M quantization
```bash
# Check available quantizations
ollama show mistral:7b --modelfile
# Pull Q3 if available, or create from Q4
ollama pull mistral:7b-instruct-q3_K_M
```
**Expected Outcome**:
- VRAM savings: 1.9GB (5.1GB → 3.2GB)
- Tool calling: Should work (same architecture)
- Quality: 80% of Q4 (acceptable for routing logic)
- Expert model budget: 6.6GB (up from 4.7GB)
**Risk Mitigation**:
- Test thoroughly before production
- Keep Q4 variant as backup
- Monitor for quality degradation
**Long-term**:
- Research newer models (qwen2.5, hermes3)
- Consider hybrid routing if complexity justified
- Revisit when Ollama adds model multiplexing features
---
**Status**: Research complete, awaiting quantization testing
**Next Steps**: User decision on Q3 testing approach