jpmschweitzerandClaude 632b20febe feat(core-ai): implement Phase 1 of AI performance metrics system
Adds comprehensive in-memory metrics collection for monitoring AI agent
performance, tool execution, and system behavior.

New Components:
- src/metrics/collector.py: Thread-safe MetricsCollector class
  - Tracks agent requests (response times, errors, concurrency)
  - Tracks tool execution (calls, success/failure, durations)
  - Tracks memory system (tier1/tier2 hits, consolidations)
  - Calculates percentiles (p50, p95, p99) for performance analysis
  - Sliding window retention (1h detailed, 24h aggregated)

- src/metrics/decorators.py: Automatic instrumentation decorators
  - @track_tool_execution: Auto-tracks tool calls with metrics
  - @track_duration: Generic duration tracking decorator

- src/metrics/__init__.py: Module exports

API Endpoints:
- GET /metrics: Comprehensive performance metrics snapshot
- GET /metrics/errors: Recent request errors with timestamps
- GET /metrics/tool-failures: Recent tool execution failures
- POST /metrics/reset: Clear all metrics (admin endpoint)

Instrumentation:
- Enhanced main.py chat handlers with metrics tracking
- Modified tools/registry.py log_tool_call to track execution metrics
- All metrics recorded with proper error handling and context

Features:
- Thread-safe with threading.Lock for concurrent requests
- No database dependencies (in-memory only)
- Automatic cleanup of old data (sliding windows)
- Detailed statistics: avg, p50, p95, p99 response times
- Per-user tracking and request attribution
- Tool success rates and performance analysis

Tested and validated:
- All endpoints responding correctly
- Request metrics collected successfully
- Response time percentiles calculated correctly
- User tracking functional

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-12-03 22:08:03 +01:00
2025-11-14 15:31:25 +01:00
2025-11-20 12:08:57 +01:00
2025-11-14 15:31:25 +01:00
2025-11-14 15:31:25 +01:00
2025-11-20 12:08:57 +01:00
2025-11-14 15:31:25 +01:00
2025-11-20 10:17:09 +01:00

portainer-core

Self-hosted home server infrastructure with GPU-accelerated ML, AI orchestration, media streaming, and secure remote access

Main Dashboard: https://home.schweitz.net (Organizr)

Getting Started

Implementation Plans

Documentation Index

Architecture & Design

Operational Guides

Services

Reference

For AI Agents

Architecture Overview

┌─────────────────────────────────────────┐
│  Infrastructure Layer                   │
│  ├── Portainer (8080) - Container mgmt │
│  ├── PostgreSQL Shared (5432) - DB     │
│  ├── Redis Shared (6379) - Cache       │
│  ├── NPM (8000) - Reverse proxy        │
│  └── Ollama (11434) - ML models [GPU]  │
├─────────────────────────────────────────┤
│  Networking Layer                       │
│  ├── docker-dataplane - Service mesh   │
│  └── Headscale (8085) - VPN mesh       │
├─────────────────────────────────────────┤
│  Monitoring Layer                       │
│  ├── Uptime Kuma (3001) - Uptime       │
│  ├── Netdata (19999) - Metrics         │
│  └── Organizr (8084) - Dashboard       │
├─────────────────────────────────────────┤
│  Optimization Layer                     │
│  ├── Watchtower - Auto-updates         │
│  └── Maintenance - Automated backups   │
├─────────────────────────────────────────┤
│  Application Layer                      │
│  ├── Open WebUI (8081) - LLM chat UI   │
│  ├── Core API (8083) - Infra mgmt      │
│  ├── Jellyfin (8096) - Media [GPU]     │
│  ├── Nextcloud (8082) - Cloud storage  │
│  ├── Gitea (3002) - Git hosting        │
│  └── Samba (445) - File shares         │
└─────────────────────────────────────────┘

Project Structure

portainer-core/
├── plans/                  # Implementation plans
│   ├── active/            # Current development work
│   └── completed/         # Historical implementations
├── docs/                  # Documentation
│   ├── architecture/      # Design documents
│   ├── guides/            # Setup and operational guides
│   ├── services/          # Service-specific documentation
│   └── reference/         # Quick reference materials
├── stacks/                # Docker Compose files (version-controlled)
├── scripts/               # Maintenance automation
├── services/              # Service source code
│   ├── core-api/         # Infrastructure management API
│   └── ...
├── organizr-widgets/      # Dashboard widgets
├── AGENTS.md             # AI agent guidelines (single source of truth)
├── README.md             # This file (documentation index)
├── PLANS.md              # Implementation plan tracker
├── STATUS.md             # Current phase tracking
└── CHANGELOG.md          # Version history

Common Commands

Infrastructure Management

make status              # Show running containers and system status
make health              # Comprehensive health check
make gpu-check           # Verify GPU passthrough
make disk                # Disk usage report
make backup              # Backup Docker configs
make cleanup             # Clean unused Docker resources

Stack Management

make deploy-portainer    # Deploy Portainer
make deploy-ollama       # Deploy Ollama
make logs-ollama         # View Ollama logs
make update-jellyfin     # Update Jellyfin to latest
make stop-nextcloud      # Stop Nextcloud stack

See Stacks Reference for complete stack inventory and deployment procedures.

Development Setup

Python Environment

Some automation scripts require Python dependencies:

# Activate virtual environment
source .venv/bin/activate

# Install/update dependencies
pip install -r requirements.txt

# Deactivate when done
deactivate

Service Development

See individual service documentation:

Service Ports Reference

Service Port Description
Portainer 8080 Container management UI
Nginx Proxy Manager 8000 Reverse proxy admin
Open WebUI 8081 LLM chat interface
Nextcloud 8082 Cloud storage
Core API 8083 Infrastructure management API
Organizr 8084 Unified dashboard
Headscale 8085 VPN control server
Jellyfin 8096 Media streaming
Uptime Kuma 3001 Service monitoring
Gitea 3002 Git repository hosting
Gitea SSH 2222 Git SSH access
PostgreSQL Shared 5432 Shared database (internal)
Redis Shared 6379 Shared cache (internal)
Qdrant 6333, 6334 Vector database
Ollama 11434 ML model API
Netdata 19999 System monitoring

See Stacks Reference for complete port allocation.

GPU Services

Two services leverage the RTX 2080 Ti:

  1. Ollama - ML model inference (3B-13B parameter models)
  2. Jellyfin - Hardware video transcoding (NVENC)

See GPU Docker Configuration for setup.

Storage Strategy

SSD (Performance): /home/jpmschweitzer/docker-data/

  • Docker configs, databases, cache, container images

HDD (Capacity): /mnt/media/

  • Media files, user data, backups

See Stacks Reference for details.

Current Phase

Phase 2 of AI Orchestrator Enhancement (Memory Systems) 🔄 In Progress

See STATUS.md for detailed progress tracking.

Contributing

This is a personal infrastructure project. For AI agents:

  • Read AGENTS.md first - Mandatory guidelines
  • Follow conventional commit format
  • Test GPU access before deploying GPU services
  • Update STATUS.md when completing phases

Resources


Version: 0.7.1 Last Updated: 2025-11-20 System: tower-of-joy

S
Description
No description provided
Readme
3.4 MiB
Languages
Makefile 68.8%
Shell 31.2%