Remove all obsolete agent implementations and framework references. Keep only PydanticAI (primary) and SimpleLiteLLM (fallback). This cleanup eliminates confusion between multiple frameworks that were tried during development (LangChain, LangGraph, ADK, OllamaNative) and establishes PydanticAI as the single agent framework going forward. BREAKING CHANGES: - Removed OllamaNativeAgent - use PydanticAgent instead - Removed /test/ollama-tools diagnostic endpoint - Default /v1/chat/completions now uses PydanticAgent Files Deleted (32 total): - Obsolete agents: ollama_native_agent.py - Diagnostic files: ARCHITECTURE.md, DIAGNOSTIC_RESULTS.md, PHASE*.md - Legacy tools: src/tools.py - Test files: test_ai_flow_quality.py, test_02/03 (diagnostic layers) - Documentation: ADK_Ollama_Research.md, agent-flow-diagrams.md - Session docs: 3 files with LangChain/LangGraph implementations - Plans: 5 completed plans about obsolete frameworks - Migration docs: MIGRATION_PLAN_LANGCHAIN_TO_ADK.md Files Modified (8 total): - main.py: Refactored to PydanticAI only (305 lines vs 457 before) - agents/__init__.py: Removed OllamaNativeAgent exports - README.md: Complete rewrite for PydanticAI architecture - prompts.py: Updated for PydanticAI (infrastructure tool guidance) - STATUS.md: Updated to v0.11.0-pydantic-ai - CHANGELOG.md: Added v0.11.0 entry documenting cleanup - plans/active/*.md: Updated to reference PydanticAI Current Architecture: - Framework: PydanticAI with native Ollama SDK - Agents: PydanticAgent (primary) + SimpleLiteLLMAgent (fallback) - Model: mistral-nemo:latest - Tools: 6 core + 28+ OpenAPI-discovered - Memory: 3-tier system with Qdrant - VRAM: ~4-6GB Lines Removed: ~3000+ lines of obsolete code 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>
Core-AI Service
AI agent service built on PydanticAI for infrastructure management and automation.
Architecture
Core-AI provides two agents with distinct capabilities:
┌─────────────────────────────────────────────────────────────┐
│ Core-AI Service │
├─────────────────────────────────────────────────────────────┤
│ │
│ ┌──────────────────┐ ┌──────────────────────┐ │
│ │ PydanticAgent │ │ SimpleLiteLLMAgent │ │
│ │ (Primary) │ │ (Fallback) │ │
│ │ │ │ │ │
│ │ • Tool calling │ │ • No tools │ │
│ │ • Memory (3-tier)│ │ • Direct LiteLLM │ │
│ │ • OpenAPI tools │ │ • Minimal overhead │ │
│ └────────┬─────────┘ └──────────┬───────────┘ │
│ │ │ │
│ └──────────┬───────────────────┘ │
│ │ │
│ ▼ │
│ ┌──────────────────────┐ │
│ │ PydanticAI Runtime │ │
│ │ (Ollama backend) │ │
│ └──────────────────────┘ │
└─────────────────────────────────────────────────────────────┘
PydanticAgent (Primary)
Endpoint: /v1/chat/completions (default)
Advanced agent using the PydanticAI framework with:
- Tool Calling: Automatic function calling with proper validation
- Memory System: 3-tier conversation memory (buffer + Qdrant)
- Local Tools: Time, calculations, web search (SearXNG)
- OpenAPI Tools: Auto-discovered from core-api infrastructure endpoints
- Streaming Support: Server-sent events for real-time responses
SimpleLiteLLMAgent (Fallback)
Endpoint: /v1/chat/simple
Lightweight agent for direct LLM interaction:
- No Tools: Pure conversational mode
- Direct LiteLLM: Minimal abstraction layer
- No Memory: Stateless request/response
- Low Latency: Fastest response times
Tool System
Local Tools
Built-in utilities available immediately (defined in src/tools/local.py):
get_current_time(timezone)- Timezone-aware time with IANA timezone supportget_current_date()- Current date in ISO formatcalculate(expression)- Safe mathematical calculationscalculate_date_difference(date1, date2)- Date arithmeticadd_days_to_date(date, days)- Date manipulationweb_search(query, category, max_results)- SearXNG metasearch integration
OpenAPI Discovery
Dynamically discovers infrastructure tools from core-api's OpenAPI spec:
- Auto-Discovery: Fetches
/openapi.jsonon startup - REST Mapping: Converts endpoints to callable functions
- Prefixed Names: Tools prefixed with service name (e.g.,
core-api__list_containers) - Type Safety: Preserves parameter types and validation
Configuration:
OPENAPI_ENABLED=true
OPENAPI_ENDPOINTS=http://core-api:8083/openapi.json
List Available Tools:
curl http://localhost:8086/v1/tools
Memory System
3-tier multi-tenant memory with per-user data isolation:
Tier 1: Conversation Buffer (RAM)
- Storage: In-memory per-user buffers
- Scope: Recent N turns (configurable, default: 10)
- Speed: Instant access
- Purpose: Fast context for ongoing conversations
Tier 2: Persistent Storage (Qdrant)
- Storage: Per-user Qdrant collections
- Scope: Complete conversation history
- Speed: Fast retrieval by conversation ID
- Purpose: Conversation continuity across sessions
Tier 3: Semantic Search (Qdrant)
- Storage: Same as Tier 2 with vector embeddings
- Scope: Cross-conversation semantic search
- Speed: Sub-second similarity search
- Purpose: Contextual recall across all user conversations
Multi-Tenancy
- Per-User Collections: Each user gets isolated Qdrant collection
- User ID Format: Sanitized email (
username_at_domain_com) - GDPR Compliance: Complete user data deletion support
- Automatic Isolation: No cross-user data leakage
Memory Configuration:
MEMORY_ENABLED=true
MEMORY_TIER1_SIZE=10
QDRANT_URL=http://qdrant:6333
EMBEDDING_MODEL=nomic-embed-text
DEFAULT_USER_ID=llmdefault_at_schweitz_net
API Endpoints
Chat Completions
POST /v1/chat/completions (Default: PydanticAI)
OpenAI-compatible chat endpoint using PydanticAgent.
Request:
{
"messages": [
{"role": "user", "content": "What containers are running?"}
],
"conversation_id": "optional-conversation-id",
"enable_tools": true,
"stream": false
}
Response:
{
"choices": [{
"index": 0,
"message": {
"role": "assistant",
"content": "I found 5 running containers..."
},
"finish_reason": "stop"
}],
"model": "pydantic",
"tools_enabled": true,
"tools_count": 12
}
Streaming: Set "stream": true for SSE response
Simple Chat
POST /v1/chat/simple
No-tools fallback endpoint using SimpleLiteLLMAgent.
Same request/response format as above, but tools_enabled will be false.
List Models
GET /v1/models
Returns available agent types:
pydantic- PydanticAgent (primary)simple- SimpleLiteLLMAgent (fallback)
List Tools
GET /v1/tools
Returns all available tools (local + discovered OpenAPI tools).
Health Check
GET /health
Service health status with agent availability.
Configuration
All configuration via environment variables (see src/config.py):
Core Settings
| Variable | Default | Description |
|---|---|---|
HOST |
0.0.0.0 |
Server host |
PORT |
8086 |
Server port |
LOG_LEVEL |
INFO |
Logging level |
Ollama Integration
| Variable | Default | Description |
|---|---|---|
OLLAMA_BASE_URL |
http://ollama:11434 |
Ollama API URL |
AGENT_MODEL |
mistral-nemo:latest |
Primary model (tool-calling optimized) |
OLLAMA_TIMEOUT |
300 |
Request timeout (seconds) |
System Prompts
| Variable | Default | Description |
|---|---|---|
SYSTEM_PROMPT_VARIANT |
minimal_agent |
Prompt for SimpleLiteLLMAgent |
PYDANTIC_SYSTEM_PROMPT_VARIANT |
pydantic_agent |
Prompt for PydanticAgent |
Tool Discovery
| Variable | Default | Description |
|---|---|---|
OPENAPI_ENABLED |
true |
Enable OpenAPI tool discovery |
OPENAPI_ENDPOINTS |
http://core-api:8083/openapi.json |
OpenAPI spec URLs (comma-separated) |
Memory System
| Variable | Default | Description |
|---|---|---|
MEMORY_ENABLED |
true |
Enable conversation memory |
MEMORY_TIER1_SIZE |
10 |
Max turns in RAM buffer |
QDRANT_URL |
http://qdrant:6333 |
Qdrant vector DB URL |
QDRANT_COLLECTION_PREFIX |
core_ai_user |
Prefix for user collections |
EMBEDDING_MODEL |
nomic-embed-text |
Ollama embedding model |
EMBEDDING_DIMENSION |
768 |
Embedding vector size |
DEFAULT_USER_ID |
llmdefault_at_schweitz_net |
Default user (until auth integration) |
Quick Start
1. Install Dependencies
pip install -r requirements.txt
2. Configure Environment
Create .env file:
OLLAMA_BASE_URL=http://ollama:11434
AGENT_MODEL=mistral-nemo:latest
QDRANT_URL=http://qdrant:6333
MEMORY_ENABLED=true
OPENAPI_ENABLED=true
3. Start Service
python main.py
Service available at http://localhost:8086
4. Test Chat
curl -X POST http://localhost:8086/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"messages": [
{"role": "user", "content": "What time is it in Amsterdam?"}
],
"enable_tools": true
}'
The agent will automatically use the get_current_time tool.
Testing
Unit Tests
# Run all tests
pytest tests/ -v
# Run specific test suite
pytest tests/test_ai_flow_quality.py -v
# Run with coverage
pytest tests/ --cov=src --cov-report=html
Integration Tests
Quality tests for end-to-end AI flows:
pytest tests/test_ai_flow_quality.py -v
See tests/QUALITY_TESTS.md for test documentation.
Manual Testing
# Test PydanticAgent (with tools)
curl -X POST http://localhost:8086/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"messages": [{"role": "user", "content": "Calculate 123 * 456"}]}'
# Test SimpleLiteLLMAgent (no tools)
curl -X POST http://localhost:8086/v1/chat/simple \
-H 'Content-Type: application/json' \
-d '{"messages": [{"role": "user", "content": "Hello!"}]}'
# List available tools
curl http://localhost:8086/v1/tools
# Health check
curl http://localhost:8086/health
Docker Deployment
Build
docker build -t core-ai:latest .
Run
docker run -d \
--name core-ai \
-p 8086:8086 \
-e OLLAMA_BASE_URL=http://ollama:11434 \
-e QDRANT_URL=http://qdrant:6333 \
-e AGENT_MODEL=mistral-nemo:latest \
--network docker-dataplane \
core-ai:latest
Using Docker Compose
docker-compose -f ../../stacks/core-ai.yml up
Project Structure
services/core-ai/
├── main.py # HTTP server (aiohttp)
├── src/
│ ├── agents/
│ │ ├── __init__.py # Agent exports
│ │ ├── pydantic_agent.py # PydanticAgent (primary)
│ │ └── simple.py # SimpleLiteLLMAgent (fallback)
│ ├── memory/
│ │ ├── manager.py # Multi-tenant memory manager
│ │ ├── tier1_buffer.py # RAM conversation buffer
│ │ ├── qdrant_memory.py # Qdrant persistent + semantic
│ │ ├── base.py # Base memory interfaces
│ │ └── schemas.py # Memory data schemas
│ ├── tools/
│ │ ├── local.py # Local utility tools
│ │ ├── openapi_discovery.py # OpenAPI tool discovery
│ │ └── registry.py # Tool registration system
│ ├── config.py # Configuration (Pydantic Settings)
│ ├── prompts.py # System prompts
│ └── utils.py # Utilities
├── tests/
│ ├── test_ai_flow_quality.py # End-to-end AI quality tests
│ └── QUALITY_TESTS.md # Test documentation
├── requirements.txt
├── Dockerfile
└── README.md (this file)
Development
Adding Local Tools
Edit src/tools/local.py:
from src.tools.registry import register_tool
@register_tool
async def my_new_tool(param: str) -> str:
"""
Tool description for LLM.
Args:
param: Parameter description
Returns:
Result description
"""
# Implementation
return f"Result: {param}"
Tool automatically available to PydanticAgent.
Adding OpenAPI Sources
Add endpoints to configuration:
OPENAPI_ENDPOINTS=http://core-api:8083/openapi.json,http://automation:8080/openapi.json
Tools auto-discovered on startup with service prefix:
core-api__list_containersautomation__deploy_stack
Modifying System Prompts
Edit src/prompts.py:
PROMPTS = {
"pydantic_agent": "Your custom PydanticAgent prompt...",
"minimal_agent": "Your custom SimpleLiteLLMAgent prompt..."
}
Update environment:
PYDANTIC_SYSTEM_PROMPT_VARIANT=pydantic_agent
Memory System Usage
Memory automatically managed per user:
from src.memory import get_memory_manager_for_user
# Get user's memory manager
memory = get_memory_manager_for_user(user_id="user_at_example_com")
# Memory automatically used by PydanticAgent when conversation_id provided
# See: src/agents/pydantic_agent.py
Troubleshooting
PydanticAI Not Available
Error: PydanticAI not available. Install with: pip install pydantic-ai
Solution:
pip install pydantic-ai
Tools Not Discovered
Issue: /v1/tools returns empty list or only local tools
Check:
- Verify
OPENAPI_ENABLED=true - Check core-api is running:
curl http://core-api:8083/openapi.json - Review logs for discovery errors:
docker logs core-ai
Memory Errors
Issue: Memory operations failing
Check:
- Verify Qdrant running:
curl http://qdrant:6333/collections - Check embedding model available:
docker exec ollama ollama list | grep nomic-embed-text - Review logs for initialization errors
Model Timeouts
Issue: Requests timing out
Solutions:
- Increase timeout:
OLLAMA_TIMEOUT=600 - Use smaller model:
AGENT_MODEL=mistral-tools:7b - Check GPU access:
docker exec ollama nvidia-smi
Tool Calling Failures
Issue: Agent not using tools correctly
Check:
- Verify model supports tool calling:
mistral-nemo,mistral-tools:7b - Test with
enable_tools=falseto isolate issue - Review tool logs: Look for
🔧 TOOL CALL:in logs
Model Recommendations
For Tool Calling (PydanticAgent)
- mistral-nemo:latest (default) - Best balance
- mistral-tools:7b - Faster, less accurate
- llama3.1:8b - Good alternative
For Simple Chat (SimpleLiteLLMAgent)
- gemma2:9b - Fast conversational
- llama3.2:3b - Minimal resources
- Any model works (no tool calling required)
Migration Notes
This service has migrated from:
- ADK (Agent Development Kit) → PydanticAI
- LangChain/LangGraph → PydanticAI native
- OllamaNativeAgent → Removed (superseded by PydanticAgent)
All references to these frameworks have been removed. The codebase now exclusively uses PydanticAI for agent orchestration.
Contributing
When making changes:
- Add tests in
tests/ - Update docstrings
- Test with both agents (
/v1/chat/completionsand/v1/chat/simple) - Verify tool discovery works
- Test memory persistence
License
Part of the portainer-core project.