chore: prune obsolete documentation and consolidate
- Remove obsolete plans/ directory (AI orchestrator moved to separate services) - Remove docs/sessions/ (historical session notes) - Remove docs/reference/ (duplicated in CONTAINERS.md) - Remove docs/architecture/ (duplicated in CONTAINERS.md) - Remove maintenance stack and backup-procedures guide (decommissioned) - Remove scripts/ directory (unused) - Fold EXTERNAL_SERVICES.md into CONTAINERS.md - Update Nextcloud to reflect PostgreSQL shared (was MariaDB) - Clean up README.md links 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
This commit is contained in:
@@ -1,83 +0,0 @@
|
||||
# External Services
|
||||
|
||||
Services that have been extracted to their own repositories but are still part of the overall infrastructure.
|
||||
|
||||
---
|
||||
|
||||
## Scheduler
|
||||
|
||||
**Repository:** https://git.schweitz.net/jpmschweitzer/scheduler
|
||||
**Container Image:** `git.schweitz.net/jpmschweitzer/scheduler:latest`
|
||||
**Stack File:** `stacks/scheduler.yml` (remains in portainer-core)
|
||||
|
||||
### Overview
|
||||
The Scheduler service handles system-wide maintenance orchestration including:
|
||||
- Docker config backups
|
||||
- Documentation mirroring
|
||||
- Task automation via REST API
|
||||
- Integration with Library Desk for consolidation jobs
|
||||
|
||||
### Deployment
|
||||
- Container image built via Gitea Actions on release
|
||||
- Stack file in portainer-core defines volumes, environment, and network
|
||||
- Watchtower monitors for image updates
|
||||
|
||||
### Development
|
||||
To make changes:
|
||||
1. Clone: `git clone gitea:jpmschweitzer/scheduler.git`
|
||||
2. Make changes
|
||||
3. Create a release in Gitea to trigger build
|
||||
4. Watchtower will auto-update the running container
|
||||
|
||||
### API
|
||||
- Health: `http://scheduler:8090/health`
|
||||
- Tasks: `http://scheduler:8090/tasks` (requires API key)
|
||||
- Docs: `http://scheduler:8090/docs`
|
||||
|
||||
---
|
||||
|
||||
## Core-API
|
||||
|
||||
**Repository:** https://git.schweitz.net/jpmschweitzer/core-api
|
||||
**Container Image:** `git.schweitz.net/jpmschweitzer/core-api:latest`
|
||||
**Stack File:** `stacks/core-api.yml` (remains in portainer-core)
|
||||
|
||||
### Overview
|
||||
Core-API provides infrastructure orchestration and OpenAI-compatible API endpoints:
|
||||
- Infrastructure management (Portainer, NPM, Uptime Kuma integration)
|
||||
- Web scraping tools
|
||||
- AI metrics proxy
|
||||
- Service health monitoring
|
||||
|
||||
### Deployment
|
||||
- Container image built via Gitea Actions on release
|
||||
- Stack file in portainer-core defines volumes, environment, and network
|
||||
- Watchtower monitors for image updates
|
||||
|
||||
### Development
|
||||
To make changes:
|
||||
1. Clone: `git clone gitea:jpmschweitzer/core-api.git`
|
||||
2. Make changes
|
||||
3. Create a release in Gitea to trigger build
|
||||
4. Watchtower will auto-update the running container
|
||||
|
||||
### API
|
||||
- Health: `http://core-api:8083/health/full`
|
||||
- Docs: `http://core-api:8083/docs`
|
||||
|
||||
---
|
||||
|
||||
## Tatlock
|
||||
|
||||
**Repository:** https://git.schweitz.net/jpmschweitzer/tatlock
|
||||
**Location:** `/home/jpmschweitzer/Projects/tatlock`
|
||||
|
||||
### Overview
|
||||
Tatlock is a separate project maintained in its own repository.
|
||||
|
||||
---
|
||||
|
||||
## Future Migrations
|
||||
|
||||
The following services are planned for extraction:
|
||||
- **library-desk** - Knowledge management and consolidation service
|
||||
@@ -1,294 +0,0 @@
|
||||
# Shared Infrastructure Architecture
|
||||
|
||||
**Purpose:** Centralized PostgreSQL and Redis services for all homelab stacks
|
||||
**Benefits:** Resource efficiency, easier maintenance, unified backups, centralized monitoring
|
||||
|
||||
---
|
||||
|
||||
## Architecture Overview
|
||||
|
||||
```
|
||||
┌─────────────────────────────────────────────────────────────┐
|
||||
│ Application Stacks │
|
||||
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ ┌──────────┐ │
|
||||
│ │Authentik │ │ Gitea │ │ Organizr │ │ Future │ │
|
||||
│ │ │ │ │ │ │ │ Stack │ │
|
||||
│ └────┬─────┘ └────┬─────┘ └────┬─────┘ └────┬─────┘ │
|
||||
│ │ │ │ │ │
|
||||
└───────┼─────────────┼──────────────┼──────────────┼─────────┘
|
||||
│ │ │ │
|
||||
└─────────────┴──────────────┴──────────────┘
|
||||
│
|
||||
┌─────────────▼──────────────────────────────┐
|
||||
│ Unified Data Plane Network │
|
||||
│ (docker-dataplane) │
|
||||
└─────────────┬──────────────────────────────┘
|
||||
│
|
||||
┌─────────────┴──────────────┐
|
||||
│ │
|
||||
┌────▼──────┐ ┌────────▼────┐
|
||||
│PostgreSQL │ │ Redis │
|
||||
│ Shared │ │ Shared │
|
||||
│ │ │ │
|
||||
│ Databases:│ │ DB 0: Cache │
|
||||
│ - auth │ │ DB 1: Auth │
|
||||
│ - gitea │ │ DB 2: Gitea │
|
||||
│ - organizr│ │ DB 3-15: .. │
|
||||
│ - future │ │ │
|
||||
└───────────┘ └─────────────┘
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Design Principles
|
||||
|
||||
### 1. **Database Isolation**
|
||||
- Each application gets its own PostgreSQL database within the shared instance
|
||||
- Each application gets its own Redis database number (0-15)
|
||||
- Separate credentials per application for security
|
||||
|
||||
### 2. **Network Architecture**
|
||||
- **Unified network:** `docker-dataplane` (external, bridge)
|
||||
- All application containers connect to this single network
|
||||
- Simplified connectivity: services discover each other by container name
|
||||
- Replaces per-stack networks (ai-dataplane, nextcloud-network, etc.)
|
||||
|
||||
### 3. **Resource Allocation**
|
||||
- PostgreSQL: No hard limits (homelab resource availability)
|
||||
- Redis: No hard limits (lightweight Alpine image)
|
||||
- Shared instances more efficient than per-stack deployments
|
||||
|
||||
### 4. **Backup Strategy**
|
||||
- Single PostgreSQL backup covers all databases
|
||||
- Automated pg_dumpall for disaster recovery
|
||||
- Redis persistence: AOF + RDB snapshots
|
||||
|
||||
### 5. **Security Model**
|
||||
- Each app has dedicated PostgreSQL user with access only to its database
|
||||
- Redis AUTH with per-database passwords (optional)
|
||||
- Network-level isolation via Docker networks
|
||||
|
||||
---
|
||||
|
||||
## Database Allocation Plan
|
||||
|
||||
### PostgreSQL Databases
|
||||
|
||||
| Database Name | Application | User | Purpose |
|
||||
|---------------|-------------|------|---------|
|
||||
| `authentik` | Authentik | `authentik_user` | User/group/policy storage |
|
||||
| `gitea` | Gitea | `gitea_user` | Git repos, users, issues |
|
||||
| `organizr` | Organizr | `organizr_user` | Dashboard configuration and user data |
|
||||
| `future_app1` | TBD | `app1_user` | Reserved |
|
||||
| `future_app2` | TBD | `app2_user` | Reserved |
|
||||
|
||||
**Note:** Existing services stay as-is:
|
||||
- Nextcloud: MariaDB (existing, not migrated)
|
||||
- Others can migrate over time if beneficial
|
||||
|
||||
### Redis Database Numbers
|
||||
|
||||
| DB# | Application | Purpose |
|
||||
|-----|-------------|---------|
|
||||
| 0 | Authentik | Sessions, cache, message queue |
|
||||
| 1 | Available | Reserved for future applications |
|
||||
| 2 | Available | Reserved for future applications |
|
||||
| 3-15 | Available | Reserved for future applications |
|
||||
|
||||
**Note:** Each application uses a dedicated DB number to prevent key collisions while sharing the same Redis instance.
|
||||
|
||||
---
|
||||
|
||||
## Connection Configuration
|
||||
|
||||
### PostgreSQL Connection Strings
|
||||
|
||||
**From Docker containers:**
|
||||
```
|
||||
Host: postgres-shared
|
||||
Port: 5432
|
||||
Database: authentik
|
||||
User: authentik_user
|
||||
Password: <app-specific-password>
|
||||
```
|
||||
|
||||
**From host:**
|
||||
```
|
||||
Host: localhost
|
||||
Port: 5432
|
||||
Database: authentik
|
||||
User: authentik_user
|
||||
Password: <app-specific-password>
|
||||
```
|
||||
|
||||
### Redis Connection Strings
|
||||
|
||||
**From Docker containers:**
|
||||
```
|
||||
redis://redis-shared:6379/0 (for Authentik, DB 0)
|
||||
redis://redis-shared:6379/1 (for future apps, DB 1)
|
||||
redis://redis-shared:6379/2 (for future apps, DB 2)
|
||||
```
|
||||
|
||||
**From host:**
|
||||
```
|
||||
redis://localhost:6379/0
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Migration Strategy
|
||||
|
||||
### Phase 1: Deploy Shared Infrastructure ✅ **COMPLETE**
|
||||
1. ✅ Deployed `postgres-shared.yml` and `redis-shared.yml` via Portainer
|
||||
2. ✅ Verified PostgreSQL 17 and Redis 7 running on docker-dataplane
|
||||
3. ✅ Created initial databases and users (authentik, gitea)
|
||||
4. ✅ Both services monitored via Uptime Kuma
|
||||
|
||||
### Phase 2: New Services (Authentik) 🚧 **IN PROGRESS**
|
||||
1. ⏳ Deploy Authentik pointing to shared services
|
||||
2. ⏳ Test thoroughly
|
||||
3. ⏳ Validate no performance degradation
|
||||
|
||||
### Phase 3: Network Consolidation ✅ **COMPLETE**
|
||||
1. ✅ All services migrated to docker-dataplane network
|
||||
2. ✅ Removed 7 obsolete Docker networks
|
||||
3. ✅ 18 containers on unified network for service discovery
|
||||
|
||||
### Phase 4: Migrate Existing Services (Optional)
|
||||
1. **Gitea**: Already uses PostgreSQL
|
||||
- Export existing database
|
||||
- Create gitea database in shared PostgreSQL
|
||||
- Import data
|
||||
- Update Gitea stack to use shared PostgreSQL
|
||||
- Remove old gitea-db container
|
||||
|
||||
2. **Other services**: Evaluate case-by-case
|
||||
- Nextcloud: Keep MariaDB (complex migration, low benefit)
|
||||
- Future services: Use shared from day 1
|
||||
|
||||
---
|
||||
|
||||
## Advantages
|
||||
|
||||
✅ **Resource Efficiency**
|
||||
- One PostgreSQL instance: ~1GB RAM vs ~300MB per instance
|
||||
- Saves ~700MB RAM per additional service using PostgreSQL
|
||||
|
||||
✅ **Operational Simplicity**
|
||||
- Single backup process for all PostgreSQL databases
|
||||
- Centralized monitoring and health checks
|
||||
- Easier version upgrades (upgrade once, affects all)
|
||||
|
||||
✅ **Performance**
|
||||
- Shared connection pooling
|
||||
- Better resource utilization
|
||||
- Optimized caching with shared Redis
|
||||
|
||||
✅ **Scalability**
|
||||
- Add new applications without deploying new database instances
|
||||
- Up to 15 Redis databases (more than enough for homelab)
|
||||
|
||||
---
|
||||
|
||||
## Disadvantages & Mitigations
|
||||
|
||||
⚠️ **Single Point of Failure**
|
||||
- **Mitigation:** Health checks, automated restarts, regular backups
|
||||
- **Acceptable for homelab:** VPN access ensures admin can fix issues
|
||||
|
||||
⚠️ **Resource Contention**
|
||||
- **Mitigation:** PostgreSQL connection limits per database
|
||||
- **Mitigation:** Redis max memory policy (LRU eviction)
|
||||
- **Monitoring:** Track per-database usage
|
||||
|
||||
⚠️ **Version Lock-In**
|
||||
- **Mitigation:** Use latest stable PostgreSQL version (17)
|
||||
- **Mitigation:** Test upgrades in staging before production deployment
|
||||
|
||||
---
|
||||
|
||||
## Monitoring & Maintenance
|
||||
|
||||
### Health Checks
|
||||
- PostgreSQL: `pg_isready` every 30s
|
||||
- Redis: `redis-cli ping` every 30s
|
||||
- Application connectivity tests
|
||||
|
||||
### Uptime Kuma Integration ✅ **DEPLOYED**
|
||||
|
||||
Both shared services are monitored via Uptime Kuma with automatic monitor creation through the Core API:
|
||||
|
||||
**PostgreSQL Monitor** (ID 20):
|
||||
```bash
|
||||
curl -X POST http://192.168.86.149:8083/infrastructure/monitors \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"type": "postgres",
|
||||
"name": "PostgreSQL Shared",
|
||||
"interval": 60,
|
||||
"retryInterval": 60,
|
||||
"maxretries": 3,
|
||||
"notificationIDList": [],
|
||||
"accepted_statuscodes": ["200-299"],
|
||||
"databaseConnectionString": "postgres://postgres:<url-encoded-password>@postgres-shared:5432/postgres"
|
||||
}'
|
||||
```
|
||||
|
||||
**Redis Monitor** (ID 18):
|
||||
```bash
|
||||
curl -X POST http://192.168.86.149:8083/infrastructure/monitors \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"type": "port",
|
||||
"name": "Redis Shared - Port Check",
|
||||
"hostname": "redis-shared",
|
||||
"port": 6379,
|
||||
"interval": 60,
|
||||
"retryInterval": 60,
|
||||
"maxretries": 3,
|
||||
"notificationIDList": [],
|
||||
"accepted_statuscodes": ["200-299"]
|
||||
}'
|
||||
```
|
||||
|
||||
**Note:** When monitoring PostgreSQL with passwords containing special characters, URL-encode them (`/` → `%2F`, `=` → `%3D`).
|
||||
|
||||
### Backup Schedule
|
||||
- **PostgreSQL:** Manual pg_dump to `/backups/` volume (automated backups pending)
|
||||
- **Redis:** AOF persistence (real-time) enabled via `--appendonly yes`
|
||||
|
||||
### Performance Monitoring
|
||||
- Query: `SELECT datname, numbackends FROM pg_stat_database;` (active connections)
|
||||
- Redis: `INFO stats` (keyspace usage per database)
|
||||
- Uptime Kuma dashboard: Real-time availability tracking
|
||||
|
||||
### Upgrade Path
|
||||
1. Backup all databases
|
||||
2. Test upgrade with docker-compose override
|
||||
3. Deploy new version
|
||||
4. Verify all applications connect successfully
|
||||
5. Rollback if issues detected
|
||||
|
||||
---
|
||||
|
||||
## Implementation Status
|
||||
|
||||
1. ✅ Review architecture design
|
||||
2. ✅ Create `postgres-shared.yml` and `redis-shared.yml` stacks
|
||||
3. ✅ Deploy shared PostgreSQL 17 and Redis 7 via Portainer
|
||||
4. ✅ Create initial databases (authentik, gitea)
|
||||
5. ✅ Consolidate all services to docker-dataplane network
|
||||
6. ✅ Implement Uptime Kuma monitoring via Core API
|
||||
7. ✅ Document connection patterns and deployment procedures
|
||||
8. ⏳ Update `authentik.yml` to use shared services (pending)
|
||||
9. ⏳ Test Authentik with shared infrastructure (pending)
|
||||
|
||||
---
|
||||
|
||||
## Future Enhancements
|
||||
|
||||
- **PostgreSQL Read Replicas** (if needed for heavy read workloads)
|
||||
- **Redis Sentinel** (high availability, probably overkill for homelab)
|
||||
- **PgBouncer** (connection pooling if >100 connections needed)
|
||||
- **Prometheus + Grafana** (metrics visualization)
|
||||
@@ -1,219 +0,0 @@
|
||||
# Backup & Restore Procedures
|
||||
|
||||
## Overview
|
||||
|
||||
The `maintenance` container runs scheduled backup tasks using cron. It's a simple, reliable, "set and forget" solution.
|
||||
|
||||
**Current Backups:**
|
||||
- **Docker Configs:** Daily at 3 AM
|
||||
- **Retention:** 30 days
|
||||
- **Size:** ~94 MB per backup
|
||||
- **Location:** `/mnt/media/backups/docker-configs/`
|
||||
|
||||
**What's Backed Up:**
|
||||
- ✅ All Docker container configurations
|
||||
- ✅ Nginx Proxy Manager configs & SSL certificates
|
||||
- ✅ Headscale database & config
|
||||
- ✅ All dashboard settings (Heimdall, Organizr, Uptime Kuma)
|
||||
- ✅ All service configs
|
||||
- ❌ Ollama models (re-downloadable)
|
||||
- ❌ Cache files
|
||||
- ❌ Log files
|
||||
|
||||
## Automated Backups
|
||||
|
||||
**Schedule:** Daily at 3:00 AM (configured in crontab)
|
||||
|
||||
**View Backup Logs:**
|
||||
```bash
|
||||
# Real-time logs
|
||||
docker logs -f maintenance
|
||||
|
||||
# Backup script logs
|
||||
cat ~/docker-data/maintenance/logs/backup-configs.log
|
||||
```
|
||||
|
||||
**List Existing Backups:**
|
||||
```bash
|
||||
ls -lh /mnt/media/backups/docker-configs/
|
||||
```
|
||||
|
||||
## Manual Backup
|
||||
|
||||
Run a backup anytime:
|
||||
```bash
|
||||
docker exec maintenance /scripts/backup-configs.sh
|
||||
```
|
||||
|
||||
## Restore from Backup
|
||||
|
||||
### Full Restore
|
||||
|
||||
1. **Stop all containers:**
|
||||
```bash
|
||||
docker stop $(docker ps -aq)
|
||||
```
|
||||
|
||||
2. **Backup current state (just in case):**
|
||||
```bash
|
||||
mv ~/docker-data ~/docker-data.old
|
||||
```
|
||||
|
||||
3. **Extract backup:**
|
||||
```bash
|
||||
cd ~
|
||||
tar -xzf /mnt/media/backups/docker-configs/docker-configs-YYYYMMDD-HHMMSS.tar.gz
|
||||
```
|
||||
|
||||
4. **Restart containers:**
|
||||
```bash
|
||||
docker start $(docker ps -aq)
|
||||
```
|
||||
|
||||
5. **Verify services:**
|
||||
```bash
|
||||
docker ps
|
||||
```
|
||||
|
||||
### Selective Restore (Single Service)
|
||||
|
||||
Restore only one service's config (example: Headscale):
|
||||
|
||||
```bash
|
||||
# Extract only headscale directory
|
||||
tar -xzf /mnt/media/backups/docker-configs/docker-configs-20251111-221349.tar.gz \
|
||||
--strip-components=2 \
|
||||
-C ~/docker-data/ \
|
||||
docker-data/headscale
|
||||
|
||||
# Restart the service
|
||||
docker restart headscale
|
||||
```
|
||||
|
||||
## Adding New Maintenance Tasks
|
||||
|
||||
The maintenance container can run any scheduled task, not just backups.
|
||||
|
||||
### 1. Create New Script
|
||||
|
||||
```bash
|
||||
# Create script file
|
||||
nano ~/docker-data/maintenance/scripts/my-task.sh
|
||||
|
||||
# Make it executable
|
||||
chmod +x ~/docker-data/maintenance/scripts/my-task.sh
|
||||
```
|
||||
|
||||
### 2. Add to Crontab
|
||||
|
||||
```bash
|
||||
# Edit crontab
|
||||
nano ~/docker-data/maintenance/crontab
|
||||
|
||||
# Add your schedule (example: every Sunday at 4 AM)
|
||||
# 0 4 * * 0 /scripts/my-task.sh
|
||||
```
|
||||
|
||||
### 3. Restart Container
|
||||
|
||||
```bash
|
||||
docker restart maintenance
|
||||
```
|
||||
|
||||
### Examples of Future Tasks
|
||||
|
||||
- **Weekly cleanup:** Remove old Docker images
|
||||
- **Health checks:** Verify all services are responding
|
||||
- **Update checks:** Notify when container updates available
|
||||
- **Database optimization:** Compact/optimize databases
|
||||
- **SSL renewal checks:** Verify certificates are valid
|
||||
|
||||
## Testing Backup Integrity
|
||||
|
||||
Periodically test that backups can be restored:
|
||||
|
||||
```bash
|
||||
# Create test directory
|
||||
mkdir -p /tmp/backup-test
|
||||
|
||||
# Extract backup
|
||||
tar -xzf /mnt/media/backups/docker-configs/docker-configs-LATEST.tar.gz \
|
||||
-C /tmp/backup-test
|
||||
|
||||
# Verify contents
|
||||
ls -la /tmp/backup-test/docker-data/
|
||||
|
||||
# Clean up
|
||||
rm -rf /tmp/backup-test
|
||||
```
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
### Backup Not Running
|
||||
|
||||
**Check if container is running:**
|
||||
```bash
|
||||
docker ps | grep maintenance
|
||||
```
|
||||
|
||||
**Check cron logs:**
|
||||
```bash
|
||||
docker logs maintenance
|
||||
```
|
||||
|
||||
**Manually run backup to test:**
|
||||
```bash
|
||||
docker exec maintenance /scripts/backup-configs.sh
|
||||
```
|
||||
|
||||
### Backup Taking Too Long
|
||||
|
||||
- Check if exclusions are working (Ollama models should be excluded)
|
||||
- Monitor disk I/O: `iostat -x 1`
|
||||
- Check HDD health: `sudo smartctl -a /dev/sdb`
|
||||
|
||||
### Backup Disk Full
|
||||
|
||||
- Old backups auto-delete after 30 days
|
||||
- Manually remove old backups if needed:
|
||||
```bash
|
||||
# List backups by size
|
||||
du -h /mnt/media/backups/docker-configs/*
|
||||
|
||||
# Remove specific backup
|
||||
rm /mnt/media/backups/docker-configs/docker-configs-20251001-*.tar.gz
|
||||
```
|
||||
|
||||
### Restore Failed
|
||||
|
||||
1. Check backup file integrity:
|
||||
```bash
|
||||
tar -tzf /mnt/media/backups/docker-configs/backup-file.tar.gz > /dev/null
|
||||
```
|
||||
|
||||
2. If corrupted, try previous backup
|
||||
|
||||
3. Check disk space before restoring:
|
||||
```bash
|
||||
df -h ~/docker-data
|
||||
```
|
||||
|
||||
## Backup Storage
|
||||
|
||||
**Current Usage:**
|
||||
- ~94 MB per daily backup
|
||||
- 30 days retention = ~2.8 GB total
|
||||
- Stored on 3.6 TB HDD (plenty of space)
|
||||
|
||||
**Offsite Backups (Recommended):**
|
||||
|
||||
For extra protection, periodically copy backups to external drive:
|
||||
|
||||
```bash
|
||||
# Copy last 7 days to external drive
|
||||
rsync -av --progress /mnt/media/backups/docker-configs/ /mnt/external-drive/backups/
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
**Last Updated:** 2025-11-11
|
||||
@@ -1,343 +0,0 @@
|
||||
# Stack Automation Guide
|
||||
|
||||
## Overview
|
||||
|
||||
The `update-stack.sh` script enables programmatic stack updates via Portainer's REST API. This allows LLM agents (like Claude) and automation scripts to safely update Portainer stacks without requiring UI access.
|
||||
|
||||
## Quick Start
|
||||
|
||||
```bash
|
||||
# Navigate to stacks directory
|
||||
cd /home/jpmschweitzer/Projects/portainer-core/stacks
|
||||
|
||||
# Update a stack (interactive mode - first time)
|
||||
./update-stack.sh open-webui.yml
|
||||
|
||||
# Subsequent updates (uses stored token)
|
||||
./update-stack.sh open-webui.yml
|
||||
```
|
||||
|
||||
## How It Works
|
||||
|
||||
### Authentication Flow
|
||||
|
||||
1. **First Run:**
|
||||
- Prompts for Portainer username/password
|
||||
- Authenticates with Portainer API
|
||||
- Generates JWT access token
|
||||
- Saves token to `.portainer-token` (gitignored)
|
||||
|
||||
2. **Subsequent Runs:**
|
||||
- Reads token from `.portainer-token`
|
||||
- Uses token for API calls
|
||||
- No credential prompts needed
|
||||
|
||||
### Update Process
|
||||
|
||||
1. Reads YAML file from `stacks/` directory
|
||||
2. Authenticates with Portainer (or uses cached token)
|
||||
3. Looks up stack by name (filename without .yml)
|
||||
4. Sends updated stack configuration via API
|
||||
5. Portainer validates and applies changes
|
||||
|
||||
## Usage Modes
|
||||
|
||||
### Interactive Mode (Human Operators)
|
||||
|
||||
```bash
|
||||
./update-stack.sh open-webui.yml
|
||||
```
|
||||
|
||||
**First run prompts for:**
|
||||
- Portainer username
|
||||
- Portainer password
|
||||
|
||||
**Token persists for subsequent runs.**
|
||||
|
||||
### Non-Interactive Mode (Automation/LLM Agents)
|
||||
|
||||
```bash
|
||||
export PORTAINER_USERNAME="admin"
|
||||
export PORTAINER_PASSWORD="your-secure-password"
|
||||
./update-stack.sh open-webui.yml
|
||||
```
|
||||
|
||||
**Use this mode for:**
|
||||
- CI/CD pipelines
|
||||
- LLM agent workflows
|
||||
- Automated deployment scripts
|
||||
- Cron jobs
|
||||
|
||||
### Environment Variables
|
||||
|
||||
| Variable | Required | Default | Description |
|
||||
|----------|----------|---------|-------------|
|
||||
| `PORTAINER_URL` | No | `http://localhost:8080` | Portainer instance URL |
|
||||
| `PORTAINER_USERNAME` | Non-interactive only | - | Admin username |
|
||||
| `PORTAINER_PASSWORD` | Non-interactive only | - | Admin password |
|
||||
|
||||
## Examples
|
||||
|
||||
### Update Single Stack
|
||||
|
||||
```bash
|
||||
./update-stack.sh open-webui.yml
|
||||
```
|
||||
|
||||
### Update Multiple Stacks
|
||||
|
||||
```bash
|
||||
for stack in open-webui.yml ollama.yml core-api.yml; do
|
||||
./update-stack.sh "$stack"
|
||||
echo "---"
|
||||
done
|
||||
```
|
||||
|
||||
### LLM Agent Integration
|
||||
|
||||
```bash
|
||||
# Claude Code workflow example
|
||||
export PORTAINER_USERNAME="admin"
|
||||
export PORTAINER_PASSWORD="${PORTAINER_ADMIN_PASSWORD}" # from secure env
|
||||
|
||||
# Update stack after modifying YAML
|
||||
./update-stack.sh open-webui.yml
|
||||
|
||||
# Check result
|
||||
echo $? # 0 = success, 1 = failure
|
||||
```
|
||||
|
||||
### Remote Portainer Instance
|
||||
|
||||
```bash
|
||||
export PORTAINER_URL="https://portainer.example.com"
|
||||
./update-stack.sh my-stack.yml
|
||||
```
|
||||
|
||||
## Security Considerations
|
||||
|
||||
### Token Storage
|
||||
|
||||
- Token stored in `.portainer-token` (gitignored)
|
||||
- File permissions: `600` (owner read/write only)
|
||||
- Token expires based on Portainer settings (default: 8 hours)
|
||||
- Re-authentication automatic if token expires
|
||||
|
||||
### Credentials
|
||||
|
||||
**DO NOT:**
|
||||
- ❌ Commit `.portainer-token` to git
|
||||
- ❌ Hardcode passwords in scripts
|
||||
- ❌ Share tokens between users
|
||||
- ❌ Use root/admin account for automation (create dedicated API user)
|
||||
|
||||
**DO:**
|
||||
- ✅ Use environment variables for non-interactive mode
|
||||
- ✅ Store credentials in secure password manager
|
||||
- ✅ Create dedicated Portainer user for automation
|
||||
- ✅ Rotate passwords regularly
|
||||
- ✅ Use `.gitignore` to exclude token file
|
||||
|
||||
### Best Practices
|
||||
|
||||
1. **Create Automation User:**
|
||||
```
|
||||
Portainer → Users → Add User
|
||||
Username: portainer-automation
|
||||
Role: Environment Administrator (or custom)
|
||||
```
|
||||
|
||||
2. **Use Environment Variables:**
|
||||
```bash
|
||||
# In ~/.bashrc or secure environment
|
||||
export PORTAINER_USERNAME="portainer-automation"
|
||||
export PORTAINER_PASSWORD="$(pass show portainer/automation)" # from password manager
|
||||
```
|
||||
|
||||
3. **Restrict Permissions:**
|
||||
- Grant minimum required permissions
|
||||
- Limit to specific environments/stacks if possible
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
### Authentication Failed
|
||||
|
||||
```
|
||||
[ERROR] Failed to authenticate. Check credentials and try again.
|
||||
```
|
||||
|
||||
**Solutions:**
|
||||
- Verify username/password are correct
|
||||
- Check Portainer is accessible: `curl http://localhost:8080/api/status`
|
||||
- Ensure user has admin/environment admin role
|
||||
- Try removing `.portainer-token` and re-authenticating
|
||||
|
||||
### Stack Not Found
|
||||
|
||||
```
|
||||
[ERROR] Stack 'my-stack' not found in Portainer
|
||||
Available stacks:
|
||||
- open-webui
|
||||
- ollama
|
||||
- core-api
|
||||
```
|
||||
|
||||
**Solutions:**
|
||||
- Verify stack name matches filename (without .yml)
|
||||
- Check stack exists in Portainer UI
|
||||
- Stack name is case-sensitive
|
||||
- Create stack in Portainer first if it doesn't exist
|
||||
|
||||
### Connection Refused
|
||||
|
||||
```
|
||||
[ERROR] Failed to connect to Portainer at http://localhost:8080
|
||||
```
|
||||
|
||||
**Solutions:**
|
||||
- Check Portainer is running: `docker ps | grep portainer`
|
||||
- Verify port: Portainer default is 8080
|
||||
- Set `PORTAINER_URL` if using different port/host
|
||||
- Check firewall rules if accessing remote instance
|
||||
|
||||
### Token Expired
|
||||
|
||||
```
|
||||
[ERROR] Invalid authentication token
|
||||
```
|
||||
|
||||
**Solutions:**
|
||||
- Delete token file: `rm .portainer-token`
|
||||
- Re-run script to re-authenticate
|
||||
- Check Portainer token expiration settings
|
||||
|
||||
### YAML Validation Error
|
||||
|
||||
```
|
||||
[ERROR] Stack update failed: invalid compose file
|
||||
```
|
||||
|
||||
**Solutions:**
|
||||
- Validate YAML syntax: `yamllint open-webui.yml`
|
||||
- Check Docker Compose version compatibility
|
||||
- Review Portainer logs: `docker logs portainer`
|
||||
- Test with `docker compose config -f open-webui.yml`
|
||||
|
||||
## Integration with LLM Agents
|
||||
|
||||
### Claude Code Workflow
|
||||
|
||||
This script is designed to integrate seamlessly with Claude Code workflows:
|
||||
|
||||
1. **Agent modifies YAML file:**
|
||||
```python
|
||||
# Claude uses Edit tool to update open-webui.yml
|
||||
```
|
||||
|
||||
2. **Agent calls update script:**
|
||||
```bash
|
||||
cd /home/jpmschweitzer/Projects/portainer-core/stacks
|
||||
./update-stack.sh open-webui.yml
|
||||
```
|
||||
|
||||
3. **Agent verifies deployment:**
|
||||
```bash
|
||||
docker logs open-webui --tail 20
|
||||
curl http://localhost:82 # Verify service
|
||||
```
|
||||
|
||||
### Setting Up for Claude
|
||||
|
||||
Add to user profile or environment:
|
||||
|
||||
```bash
|
||||
# In ~/.bashrc or secure location
|
||||
export PORTAINER_USERNAME="admin"
|
||||
export PORTAINER_PASSWORD="your-secure-password"
|
||||
|
||||
# Or use password manager
|
||||
export PORTAINER_PASSWORD="$(pass show portainer/admin)"
|
||||
```
|
||||
|
||||
Then Claude can directly call:
|
||||
```bash
|
||||
./update-stack.sh <stack-file.yml>
|
||||
```
|
||||
|
||||
## Advanced Usage
|
||||
|
||||
### Custom Portainer URL
|
||||
|
||||
```bash
|
||||
# Connect to remote Portainer
|
||||
export PORTAINER_URL="https://portainer.mydomain.com"
|
||||
./update-stack.sh open-webui.yml
|
||||
```
|
||||
|
||||
### Token Management
|
||||
|
||||
```bash
|
||||
# View current token (for debugging)
|
||||
cat .portainer-token | base64 -d | jq
|
||||
|
||||
# Force re-authentication
|
||||
rm .portainer-token
|
||||
./update-stack.sh open-webui.yml
|
||||
|
||||
# Use specific token
|
||||
echo "your-jwt-token-here" > .portainer-token
|
||||
chmod 600 .portainer-token
|
||||
```
|
||||
|
||||
### Dry Run (Check Only)
|
||||
|
||||
```bash
|
||||
# Validate YAML before updating
|
||||
docker compose -f open-webui.yml config
|
||||
|
||||
# Check current stack status
|
||||
curl -s http://localhost:8080/api/stacks \
|
||||
-H "Authorization: Bearer $(cat .portainer-token)" | jq
|
||||
```
|
||||
|
||||
## API Reference
|
||||
|
||||
The script uses these Portainer API endpoints:
|
||||
|
||||
- `POST /api/auth` - Authenticate and get token
|
||||
- `GET /api/endpoints` - List Docker endpoints
|
||||
- `GET /api/stacks` - List all stacks
|
||||
- `PUT /api/stacks/{id}` - Update specific stack
|
||||
|
||||
For full API documentation: https://docs.portainer.io/api/docs
|
||||
|
||||
## Maintenance
|
||||
|
||||
### Regular Tasks
|
||||
|
||||
- **Monthly:** Rotate automation user password
|
||||
- **Quarterly:** Review and audit API access logs
|
||||
- **After incidents:** Revoke and regenerate tokens
|
||||
|
||||
### Token Rotation
|
||||
|
||||
```bash
|
||||
# Revoke old token (Portainer UI)
|
||||
Portainer → Users → [user] → Access Tokens → Revoke All
|
||||
|
||||
# Re-authenticate
|
||||
rm .portainer-token
|
||||
./update-stack.sh open-webui.yml
|
||||
```
|
||||
|
||||
## Support
|
||||
|
||||
For issues or questions:
|
||||
1. Check Portainer logs: `docker logs portainer`
|
||||
2. Review this guide's Troubleshooting section
|
||||
3. Check Portainer API docs: https://docs.portainer.io/api/docs
|
||||
4. Open issue in project repository
|
||||
|
||||
---
|
||||
|
||||
*Last updated: 2025-11-14*
|
||||
@@ -1,266 +0,0 @@
|
||||
# SYSTEM.md
|
||||
|
||||
**System documentation for LLM coding agents** - This file describes the computer system where this project resides, including hardware, OS, installed software, and environment details.
|
||||
|
||||
> Last updated: 2025-11-11
|
||||
|
||||
## System Overview
|
||||
|
||||
- **Hostname**: tower-of-joy
|
||||
- **User**: jpmschweitzer
|
||||
- **Home Directory**: /home/jpmschweitzer
|
||||
- **Project Location**: /home/jpmschweitzer/Projects/portainer-core
|
||||
|
||||
## Operating System
|
||||
|
||||
### Distribution
|
||||
- **OS**: Zorin OS 16.3
|
||||
- **Based on**: Ubuntu 20.04 (Focal Fossa)
|
||||
- **Kernel**: Linux 5.4.0-216-generic
|
||||
- **Architecture**: x86_64 (64-bit)
|
||||
|
||||
### Desktop Environment
|
||||
- **Display Server**: X11 (GDM)
|
||||
- **Desktop**: GNOME Shell (Zorin session mode)
|
||||
- **Session Manager**: gnome-session
|
||||
|
||||
### Locale & Timezone
|
||||
- **Language**: en_US.UTF-8
|
||||
- **Numeric/Time Format**: nl_NL.UTF-8
|
||||
- **Timezone**: Europe/Amsterdam (CET, +0100)
|
||||
|
||||
## Hardware Specifications
|
||||
|
||||
### CPU
|
||||
- **Model**: Intel Core i7-6700 @ 3.40GHz (6th Gen Skylake)
|
||||
- **Cores**: 4 physical cores, 8 threads (2 threads per core)
|
||||
- **Architecture**: x86_64
|
||||
- **Frequency**: 800 MHz - 4000 MHz (currently ~3666 MHz)
|
||||
- **Cache**:
|
||||
- L1d: 128 KiB
|
||||
- L1i: 128 KiB
|
||||
- L2: 1 MiB
|
||||
- L3: 8 MiB
|
||||
- **Virtualization**: VT-x supported
|
||||
- **Notable Flags**: AVX, AVX2, AES-NI, SSE4.1, SSE4.2, FMA
|
||||
|
||||
### Memory
|
||||
- **Total RAM**: 16 GiB
|
||||
- **Available**: ~9.6 GiB (typical)
|
||||
- **Swap**: 2.0 GiB
|
||||
|
||||
### Storage
|
||||
|
||||
**System has 2 disks with total capacity of 4.2TB:**
|
||||
|
||||
#### Disk 1: System SSD (/dev/sda)
|
||||
- **Model**: Crucial CT525MX300SSD1 (525GB SSD)
|
||||
- **Partition**: /dev/sda1
|
||||
- **Filesystem**: ext4
|
||||
- **Total Size**: 489 GB
|
||||
- **Used**: 92 GB (21%)
|
||||
- **Available**: 365 GB
|
||||
- **Mount Point**: `/` (root)
|
||||
- **Purpose**: Operating system, Docker containers, application data
|
||||
|
||||
#### Disk 2: Media HDD (/dev/sdb)
|
||||
- **Model**: Seagate IronWolf NE ST4000NE001 (4TB NAS-grade HDD)
|
||||
- **Total Size**: 3.7 TB
|
||||
- **Filesystem**: ext4
|
||||
- **Label**: "media"
|
||||
- **UUID**: f4300e91-3f51-45a0-b038-03335c5bd792
|
||||
- **Mount Status**: ⚠️ **Currently unmounted** (not in /etc/fstab)
|
||||
- **Purpose**: Media storage for Jellyfin, Nextcloud data, backups
|
||||
- **Drive Type**: NAS-optimized (24/7 operation, multi-user workloads)
|
||||
|
||||
**Total Storage Capacity**: 4.2 TB
|
||||
|
||||
### Graphics
|
||||
- **GPU**: NVIDIA GeForce RTX 2080 Ti (TU102, Rev. A)
|
||||
- **VRAM**: 11 GB (11018 MiB)
|
||||
- **Driver**: NVIDIA 470.256.02
|
||||
- **CUDA Version**: 11.4
|
||||
- **Bus**: PCIe 0a:00.0
|
||||
- **Current Usage**: ~390 MiB VRAM (mostly X11/GNOME)
|
||||
- **Power**: 260W TDP
|
||||
|
||||
**Note**: NVCC (CUDA compiler) is not currently in PATH, but CUDA drivers are installed.
|
||||
|
||||
## Development Tools & Languages
|
||||
|
||||
### Programming Languages
|
||||
|
||||
#### Python
|
||||
- **Version**: 3.8.10 (system default)
|
||||
- **pip**: 25.3 (Python 3.10 in user site-packages)
|
||||
- **Location**: /usr/bin/python3
|
||||
- **Python 2.x**: Not installed
|
||||
- **Virtual Environments**:
|
||||
- virtualenv: Not installed
|
||||
- Conda: Not installed
|
||||
- venv module: Available (built-in)
|
||||
|
||||
#### Node.js & JavaScript
|
||||
- **Node.js**: v24.11.0
|
||||
- **npm**: 11.6.1
|
||||
- **Version Manager**: NVM installed at /home/jpmschweitzer/.nvm
|
||||
|
||||
#### Java
|
||||
- **Version**: Java 21.0.4 LTS (Oracle JDK)
|
||||
- **Runtime**: Java(TM) SE Runtime Environment (build 21.0.4+8-LTS-274)
|
||||
- **VM**: Java HotSpot 64-Bit Server VM
|
||||
|
||||
#### C/C++
|
||||
- **GCC**: 9.4.0 (Ubuntu 9.4.0-1ubuntu1~20.04.2)
|
||||
- **Make**: GNU Make 4.2.1
|
||||
- **CMake**: Not installed
|
||||
|
||||
#### Other Languages
|
||||
- **Go**: Not installed
|
||||
- **Rust**: Not installed
|
||||
|
||||
### Version Control
|
||||
- **Git**: 2.25.1
|
||||
|
||||
### Containerization & Virtualization
|
||||
- **Docker**: 28.1.1, build 4eba377
|
||||
|
||||
### Editors & IDEs
|
||||
- **Vim**: 8.1 (2018 May 18)
|
||||
- **VS Code**: Not installed
|
||||
|
||||
### Command Line Tools
|
||||
- **Shell**: Bash 5.0.17
|
||||
- **curl**: 7.68.0
|
||||
- **wget**: 1.20.3
|
||||
- **SSH**: OpenSSH 8.2p1 Ubuntu-4ubuntu0.13
|
||||
|
||||
## GPU & CUDA Information
|
||||
|
||||
### NVIDIA GPU Details
|
||||
The system has an NVIDIA RTX 2080 Ti with CUDA support, suitable for:
|
||||
- Machine learning and deep learning workloads
|
||||
- CUDA-accelerated computing
|
||||
- GPU rendering and compute tasks
|
||||
- Parallel processing
|
||||
|
||||
### CUDA Configuration
|
||||
- **Driver Version**: 470.256.02
|
||||
- **CUDA Toolkit Version**: 11.4 (driver supports)
|
||||
- **Compute Capability**: 7.5 (Turing architecture)
|
||||
- **NVCC**: Not in PATH (may need manual setup)
|
||||
|
||||
### GPU Usage Considerations
|
||||
When working with GPU-accelerated code:
|
||||
- Ensure CUDA toolkit is properly installed if needed
|
||||
- Use appropriate CUDA version compatibility (11.4 or compatible)
|
||||
- PyTorch/TensorFlow should use CUDA 11.x compatible builds
|
||||
- Monitor VRAM usage (11 GB total, ~10.6 GB available for compute)
|
||||
|
||||
## System Capabilities & Recommendations
|
||||
|
||||
### Suitable For
|
||||
- **Web Development**: Node.js, npm available
|
||||
- **Python Development**: Python 3.8 with pip
|
||||
- **Java Development**: Java 21 LTS
|
||||
- **Machine Learning**: CUDA-capable GPU with 11GB VRAM
|
||||
- **Containerized Development**: Docker available
|
||||
- **Compiled Languages**: GCC toolchain available
|
||||
- **Media Server**: 3.7TB NAS-grade storage for Jellyfin/Plex
|
||||
- **NAS/File Server**: Seagate IronWolf drive optimized for 24/7 operation
|
||||
- **Cloud Storage**: Ample space for Nextcloud deployments
|
||||
- **Home Server**: Suitable for comprehensive home lab setup
|
||||
|
||||
### Limitations
|
||||
- No Rust toolchain (needs installation)
|
||||
- No Go compiler (needs installation)
|
||||
- CMake not installed (needed for some C/C++ projects)
|
||||
- VS Code not installed (Vim available as alternative)
|
||||
- CUDA compiler not in PATH
|
||||
|
||||
### Environment Notes
|
||||
- NVM is available for Node.js version management
|
||||
- Python 3.8 is the system default (older, consider pyenv for newer versions)
|
||||
- pip is installed in user site-packages (Python 3.10 version)
|
||||
- Docker is available for containerized workflows
|
||||
|
||||
## Package Management
|
||||
|
||||
### System Package Manager
|
||||
- **APT**: Available (Ubuntu/Debian package manager)
|
||||
- Use `sudo apt install <package>` for system packages
|
||||
|
||||
### Language-Specific Package Managers
|
||||
- **Python**: pip3 (25.3)
|
||||
- **Node.js**: npm (11.6.1), managed via NVM
|
||||
- **Java**: Maven/Gradle likely needed (not verified)
|
||||
|
||||
## Network Information
|
||||
- SSH client available (OpenSSH 8.2p1)
|
||||
- Standard network tools available (curl, wget)
|
||||
|
||||
## Usage Notes for LLM Agents
|
||||
|
||||
### Before Installing New Software
|
||||
1. Check if the tool is already installed using `which <command>`
|
||||
2. Check available disk space:
|
||||
- System SSD: 365 GB available (for OS and containers)
|
||||
- Media HDD: 3.7 TB available (currently unmounted - needs mounting)
|
||||
3. Use appropriate package manager (apt, pip, npm, etc.)
|
||||
4. Consider using Docker for isolated environments
|
||||
5. **Mount the 4TB media drive** before deploying data-intensive services:
|
||||
- Recommended mount point: `/mnt/media` or `/media/storage`
|
||||
- Add to `/etc/fstab` for automatic mounting on boot
|
||||
- UUID: `f4300e91-3f51-45a0-b038-03335c5bd792`
|
||||
|
||||
### GPU Development
|
||||
1. Verify CUDA toolkit path if developing GPU code
|
||||
2. Check GPU memory availability with `nvidia-smi`
|
||||
3. Use CUDA 11.x compatible libraries
|
||||
4. Monitor GPU utilization to avoid OOM errors
|
||||
|
||||
### Python Development
|
||||
1. System Python is 3.8.10 (older version)
|
||||
2. Consider using venv for project isolation
|
||||
3. pip is available but points to Python 3.10 libs in user space
|
||||
4. May need to install python3-venv: `sudo apt install python3-venv`
|
||||
|
||||
### Node.js Development
|
||||
1. NVM is installed for version management
|
||||
2. Current Node.js is v24.11.0 (latest as of 2024)
|
||||
3. npm 11.6.1 is available
|
||||
|
||||
### Docker Usage
|
||||
1. Docker 28.1.1 is installed
|
||||
2. Useful for consistent development environments
|
||||
3. Can isolate dependencies and avoid system conflicts
|
||||
|
||||
### Storage Management (Dual-Disk Setup)
|
||||
1. **System SSD (/dev/sda)**: Use for:
|
||||
- Operating system
|
||||
- Docker images and container configs
|
||||
- Application databases (small, performance-critical)
|
||||
- Cache directories
|
||||
|
||||
2. **Media HDD (/dev/sdb)**: Use for:
|
||||
- Jellyfin/Plex media libraries
|
||||
- Nextcloud user data
|
||||
- Backups and archives
|
||||
- Large file storage
|
||||
- Any data-intensive workloads
|
||||
|
||||
3. **Best Practices**:
|
||||
- Keep Docker container configs on SSD for performance
|
||||
- Store media files on HDD (they're sequential access, HDD is fine)
|
||||
- Use bind mounts to map HDD storage into containers
|
||||
- Example: `-v /mnt/media/jellyfin:/media:ro` in Docker
|
||||
|
||||
4. **Before First Use**:
|
||||
- Mount the media drive (see step 5 in "Before Installing New Software")
|
||||
- Verify mount with `df -h /mnt/media`
|
||||
- Set appropriate permissions: `sudo chown -R $USER:$USER /mnt/media`
|
||||
|
||||
---
|
||||
|
||||
*Generated automatically on 2025-11-11, updated with storage configuration*
|
||||
*For project-specific guidelines, see [AGENTS.md](./AGENTS.md)*
|
||||
@@ -1,200 +0,0 @@
|
||||
# Maintenance Scripts Reference
|
||||
|
||||
Shell scripts for common maintenance tasks located in `/scripts/`.
|
||||
|
||||
## Available Scripts
|
||||
|
||||
| Script | Description | Usage |
|
||||
|--------|-------------|-------|
|
||||
| `gpu-check.sh` | Verify GPU passthrough in containers | `./scripts/gpu-check.sh` |
|
||||
| `health-check.sh` | Check all services and report status | `./scripts/health-check.sh` |
|
||||
| `setup-kuma-monitors.sh` | Manual guide for configuring Uptime Kuma monitors | `./scripts/setup-kuma-monitors.sh` |
|
||||
| `setup-kuma-monitors.py` | **Automated** Uptime Kuma monitor setup via API | `source .venv/bin/activate && python3 scripts/setup-kuma-monitors.py` |
|
||||
| `backup-configs.sh` | Backup all Docker configs | `./scripts/backup-configs.sh` |
|
||||
| `disk-usage.sh` | Report disk usage for SSD and HDD | `./scripts/disk-usage.sh` |
|
||||
| `update-stacks.sh` | Pull latest images and update containers | `./scripts/update-stacks.sh <stack-name>` |
|
||||
| `cleanup.sh` | Clean up unused Docker resources | `./scripts/cleanup.sh` |
|
||||
|
||||
## Making Scripts Executable
|
||||
|
||||
```bash
|
||||
# Make all scripts executable
|
||||
chmod +x scripts/*.sh
|
||||
|
||||
# Or individually
|
||||
chmod +x scripts/health-check.sh
|
||||
```
|
||||
|
||||
## Scheduling with Cron
|
||||
|
||||
Add to crontab for automated maintenance:
|
||||
|
||||
```bash
|
||||
# Edit crontab
|
||||
crontab -e
|
||||
|
||||
# Examples:
|
||||
# Daily health check at 8 AM
|
||||
0 8 * * * /home/jpmschweitzer/Projects/portainer-core/scripts/health-check.sh >> /var/log/portainer-core-health.log 2>&1
|
||||
|
||||
# Weekly cleanup on Sunday at 3 AM
|
||||
0 3 * * 0 /home/jpmschweitzer/Projects/portainer-core/scripts/cleanup.sh
|
||||
|
||||
# Daily backup at 2 AM
|
||||
0 2 * * * /home/jpmschweitzer/Projects/portainer-core/scripts/backup-configs.sh
|
||||
```
|
||||
|
||||
## Script Details
|
||||
|
||||
### GPU Check (`gpu-check.sh`)
|
||||
|
||||
Verifies GPU passthrough is working in GPU-enabled containers (Ollama, Jellyfin).
|
||||
|
||||
**Usage:**
|
||||
```bash
|
||||
./scripts/gpu-check.sh
|
||||
```
|
||||
|
||||
**Output:**
|
||||
- Lists all running containers with GPU access
|
||||
- Runs `nvidia-smi` inside each container
|
||||
- Reports any containers that fail GPU detection
|
||||
|
||||
### Health Check (`health-check.sh`)
|
||||
|
||||
Checks status of all deployed services and generates a health report.
|
||||
|
||||
**Usage:**
|
||||
```bash
|
||||
./scripts/health-check.sh
|
||||
```
|
||||
|
||||
**Checks:**
|
||||
- Container running status
|
||||
- Container health status (if health check defined)
|
||||
- Port accessibility
|
||||
- Basic connectivity tests
|
||||
|
||||
### Uptime Kuma Monitor Setup
|
||||
|
||||
Two versions available:
|
||||
|
||||
**Manual Script (`setup-kuma-monitors.sh`):**
|
||||
- Interactive guide for adding monitors
|
||||
- Shows recommended settings for each service
|
||||
- Good for understanding monitor configuration
|
||||
|
||||
**Automated Script (`setup-kuma-monitors.py`):**
|
||||
- Python script using Uptime Kuma API
|
||||
- Automatically creates monitors for all services
|
||||
- Requires Uptime Kuma API key
|
||||
|
||||
**Usage:**
|
||||
```bash
|
||||
# Automated setup
|
||||
source .venv/bin/activate
|
||||
python3 scripts/setup-kuma-monitors.py
|
||||
```
|
||||
|
||||
### Backup Configs (`backup-configs.sh`)
|
||||
|
||||
Backs up Docker container configurations and important data.
|
||||
|
||||
**Usage:**
|
||||
```bash
|
||||
./scripts/backup-configs.sh
|
||||
```
|
||||
|
||||
**What it backs up:**
|
||||
- Docker Compose files from `/stacks/`
|
||||
- Container configs from `/home/jpmschweitzer/docker-data/`
|
||||
- Project documentation
|
||||
- Excludes large media files (those are backed up separately)
|
||||
|
||||
**Backup location:**
|
||||
- `/mnt/media/backups/portainer-core/`
|
||||
|
||||
See [Backup Procedures](../guides/backup-procedures.md) for comprehensive backup strategy.
|
||||
|
||||
### Disk Usage (`disk-usage.sh`)
|
||||
|
||||
Reports disk usage breakdown for SSD and HDD storage.
|
||||
|
||||
**Usage:**
|
||||
```bash
|
||||
./scripts/disk-usage.sh
|
||||
```
|
||||
|
||||
**Output:**
|
||||
- Total SSD usage (`/home/jpmschweitzer/docker-data/`)
|
||||
- Total HDD usage (`/mnt/media/`)
|
||||
- Per-service breakdown
|
||||
- Available space warnings
|
||||
|
||||
### Update Stacks (`update-stacks.sh`)
|
||||
|
||||
Pulls latest images and updates a specific stack.
|
||||
|
||||
**Usage:**
|
||||
```bash
|
||||
./scripts/update-stacks.sh <stack-name>
|
||||
|
||||
# Examples:
|
||||
./scripts/update-stacks.sh jellyfin
|
||||
./scripts/update-stacks.sh core-api
|
||||
```
|
||||
|
||||
**What it does:**
|
||||
1. Pulls latest images for the stack
|
||||
2. Stops containers gracefully
|
||||
3. Recreates containers with new images
|
||||
4. Removes old images
|
||||
5. Verifies containers started successfully
|
||||
|
||||
**Note:** Watchtower handles this automatically for most services. Use this script for manual updates or services excluded from Watchtower.
|
||||
|
||||
### Cleanup (`cleanup.sh`)
|
||||
|
||||
Cleans up unused Docker resources to free disk space.
|
||||
|
||||
**Usage:**
|
||||
```bash
|
||||
./scripts/cleanup.sh
|
||||
```
|
||||
|
||||
**What it removes:**
|
||||
- Stopped containers
|
||||
- Unused images
|
||||
- Dangling build cache
|
||||
- Unused volumes (with confirmation prompt)
|
||||
- Unused networks
|
||||
|
||||
**Warning:** Always review what will be removed before confirming volume deletion.
|
||||
|
||||
## Script Guidelines
|
||||
|
||||
All scripts follow these conventions:
|
||||
|
||||
- Include error handling and exit codes
|
||||
- Use absolute paths for reliability
|
||||
- Log output for debugging
|
||||
- Exit with status codes (0 = success, non-zero = failure)
|
||||
- Include help text with `-h` or `--help` flags
|
||||
- Non-destructive by default (ask before deleting)
|
||||
|
||||
## Creating New Scripts
|
||||
|
||||
When adding new maintenance scripts:
|
||||
|
||||
1. Place in `/scripts/` directory
|
||||
2. Use `.sh` extension for shell scripts
|
||||
3. Make executable: `chmod +x scripts/your-script.sh`
|
||||
4. Add to this documentation
|
||||
5. Include help text and error handling
|
||||
6. Test thoroughly before scheduling with cron
|
||||
|
||||
## Related Documentation
|
||||
|
||||
- [Backup Procedures](../guides/backup-procedures.md) - Comprehensive backup strategy
|
||||
- [Stacks Reference](stacks.md) - Stack deployment and management
|
||||
- [Automation Reference](AUTOMATION.md) - Portainer REST API automation
|
||||
@@ -1,185 +0,0 @@
|
||||
# Docker Compose Stacks Reference
|
||||
|
||||
Complete reference for all Docker Compose stacks in the portainer-core infrastructure.
|
||||
|
||||
## Deployment
|
||||
|
||||
See the [core-api OpenAPI documentation](http://localhost:8083/docs) for infrastructure management REST endpoints.
|
||||
|
||||
All stacks are located in the `/stacks/` directory and version-controlled.
|
||||
|
||||
## Stack Inventory
|
||||
|
||||
### Phase 1: Foundation
|
||||
|
||||
| Stack | File | Ports | GPU | Description |
|
||||
|-------|------|-------|-----|-------------|
|
||||
| **Portainer** | `portainer.yml` | 8080, 8443 | No | Container management UI |
|
||||
| **Nginx Proxy Manager** | `nginx-proxy-manager.yml` | 8000, 80, 443 | No | Reverse proxy and unified web interface |
|
||||
| **Ollama** | `ollama.yml` | 11434 | **Yes** | ML model serving with GPU acceleration |
|
||||
|
||||
### Phase 2: Networking
|
||||
|
||||
| Stack | File | Ports | GPU | Description |
|
||||
|-------|------|-------|-----|-------------|
|
||||
| **Headscale** | `headscale.yml` | 8085, 9090 | No | Self-hosted Tailscale control server |
|
||||
|
||||
### Phase 3: Monitoring
|
||||
|
||||
| Stack | File | Ports | GPU | Description |
|
||||
|-------|------|-------|-----|-------------|
|
||||
| **Uptime Kuma** | `uptime-kuma.yml` | 3001 | No | Service availability monitoring |
|
||||
| **Netdata** | `netdata.yml` | 19999 | No | Real-time system performance monitoring |
|
||||
| **Heimdall** | `heimdall.yml` | 8888, 8889 | No | Application dashboard |
|
||||
|
||||
### Phase 4: Optimization
|
||||
|
||||
| Stack | File | Ports | GPU | Description |
|
||||
|-------|------|-------|-----|-------------|
|
||||
| **Watchtower** | `watchtower.yml` | - | No | Automatic container updates |
|
||||
| **Duplicati** | `duplicati.yml` | 8200 | No | Backup solution |
|
||||
|
||||
### Applications
|
||||
|
||||
| Stack | File | Ports | GPU | Description |
|
||||
|-------|------|-------|-----|-------------|
|
||||
| **Jellyfin** | `jellyfin.yml` | 8096, 8920, 7359, 1900 | **Yes** | Media server with GPU transcoding |
|
||||
| **Nextcloud** | `nextcloud.yml` | 8082 | No | Cloud storage (uses shared PostgreSQL and Redis) |
|
||||
| **Gitea** | `gitea.yml` | 3002, 2222 | No | Git repository hosting (includes PostgreSQL) |
|
||||
| **Samba** | `samba.yml` | 139, 445 | No | Network file sharing |
|
||||
| **Open WebUI** | `open-webui.yml` | 8081 | No | AI chat interface with Ollama integration |
|
||||
| **Core API** | `core-api.yml` | 8083 | No | Infrastructure management and AI orchestration |
|
||||
| **Qdrant** | `qdrant.yml` | 6333, 6334 | No | Vector database for embeddings |
|
||||
| **Organizr** | `organizr.yml` | 8084 | No | Unified dashboard |
|
||||
|
||||
### Shared Infrastructure
|
||||
|
||||
| Stack | File | Ports | GPU | Description |
|
||||
|-------|------|-------|-----|-------------|
|
||||
| **PostgreSQL Shared** | `postgres-shared.yml` | 5432 | No | Shared database for Nextcloud |
|
||||
| **Redis Shared** | `redis-shared.yml` | 6379 | No | Shared cache for Nextcloud |
|
||||
|
||||
## Port Allocation
|
||||
|
||||
### Infrastructure Services (8000-8099)
|
||||
- 8000: Nginx Proxy Manager (unified web interface)
|
||||
- 8080: Portainer
|
||||
- 8081: Open WebUI
|
||||
- 8082: Nextcloud
|
||||
- 8083: Core API
|
||||
- 8084: Organizr
|
||||
- 8085: Headscale
|
||||
- 8096: Jellyfin
|
||||
|
||||
### Git & Development Services
|
||||
- 2222: Gitea SSH
|
||||
- 3002: Gitea HTTP
|
||||
|
||||
### Monitoring Services (3000-3999, 19000-19999)
|
||||
- 3001: Uptime Kuma
|
||||
- 8200: Duplicati
|
||||
- 8888: Heimdall
|
||||
- 19999: Netdata
|
||||
|
||||
### ML/API Services (11000+)
|
||||
- 11434: Ollama
|
||||
- 6333: Qdrant HTTP
|
||||
- 6334: Qdrant gRPC
|
||||
|
||||
### Database Services
|
||||
- 5432: PostgreSQL (shared)
|
||||
- 6379: Redis (shared)
|
||||
|
||||
### Network Services
|
||||
- 80: HTTP (NPM reverse proxy)
|
||||
- 443: HTTPS (NPM reverse proxy)
|
||||
- 139, 445: Samba/SMB
|
||||
- 9090: Headscale metrics
|
||||
|
||||
## Storage Convention
|
||||
|
||||
All stacks follow the dual-disk strategy:
|
||||
|
||||
**SSD (Performance):**
|
||||
- Configs: `/home/jpmschweitzer/docker-data/<service>/config`
|
||||
- Cache: `/home/jpmschweitzer/docker-data/<service>/cache`
|
||||
- Databases: `/home/jpmschweitzer/docker-data/<service>/db`
|
||||
|
||||
**HDD (Capacity):**
|
||||
- User content: `/mnt/media/<service>/data`
|
||||
- Media files: `/mnt/media/<service>/media`
|
||||
- Backups: `/mnt/media/backups/<service>`
|
||||
|
||||
See [Shared Infrastructure Architecture](../architecture/SHARED_INFRASTRUCTURE_ARCHITECTURE.md) for database and cache sharing details.
|
||||
|
||||
## GPU Services
|
||||
|
||||
Stacks requiring GPU access (marked with **Yes** above):
|
||||
- `ollama.yml` - ML model inference
|
||||
- `jellyfin.yml` - Hardware transcoding
|
||||
|
||||
**Prerequisites:**
|
||||
- NVIDIA Container Toolkit installed
|
||||
- GPU verified: `docker run --rm --gpus all nvidia/cuda:11.4.0-base-ubuntu20.04 nvidia-smi`
|
||||
|
||||
See [GPU Docker Configuration](../guides/gpu-docker-config.md) for setup details.
|
||||
|
||||
## Deployment Checklist
|
||||
|
||||
### Before Deploying
|
||||
|
||||
1. **Review environment variables** - Change default passwords!
|
||||
2. **Create directories** - Ensure volume paths exist
|
||||
3. **Check ports** - Verify no conflicts with existing services
|
||||
4. **GPU services** - Confirm NVIDIA toolkit installed
|
||||
5. **Update STATUS.md** - Plan the deployment
|
||||
|
||||
### After Deploying
|
||||
|
||||
1. **Test service** - Access web UI or API endpoint
|
||||
2. **Check logs** - `docker logs <container-name>`
|
||||
3. **Verify GPU** - `docker exec <container> nvidia-smi` (if applicable)
|
||||
4. **Update documentation** - Add to STATUS.md and CHANGELOG.md
|
||||
5. **Configure backup** - Add to Duplicati backup job
|
||||
6. **Add monitoring** - Configure Uptime Kuma checks
|
||||
|
||||
## Maintenance
|
||||
|
||||
### Update a Stack
|
||||
|
||||
```bash
|
||||
# Pull latest images
|
||||
docker compose -f stacks/<stack-name>.yml pull
|
||||
|
||||
# Recreate containers with new images
|
||||
docker compose -f stacks/<stack-name>.yml up -d
|
||||
|
||||
# Or let Watchtower handle it automatically
|
||||
```
|
||||
|
||||
### Backup Stack Configuration
|
||||
|
||||
Stacks are version-controlled in the `/stacks/` directory. Backup container data separately using the backup procedures.
|
||||
|
||||
See [Backup Procedures](../guides/backup-procedures.md) for details.
|
||||
|
||||
### Troubleshooting
|
||||
|
||||
- Container won't start: `docker logs <container-name>`
|
||||
- Port conflicts: `sudo netstat -tulpn | grep <port>`
|
||||
- Permission issues: Check volume path ownership
|
||||
- GPU not detected: Verify NVIDIA toolkit and restart Docker
|
||||
|
||||
## Automation
|
||||
|
||||
The project includes automation scripts for stack management:
|
||||
|
||||
- `update-stack.sh` - Pull and update specific stack
|
||||
- See [Automation Reference](AUTOMATION.md) for Portainer REST API usage
|
||||
|
||||
## Related Documentation
|
||||
|
||||
- [Container Reference](CONTAINERS.md) - Complete container profiles
|
||||
- [System Specifications](SYSTEM.md) - Hardware and software specs
|
||||
- [Shared Infrastructure Architecture](../architecture/SHARED_INFRASTRUCTURE_ARCHITECTURE.md) - Database/cache sharing
|
||||
- [Maintenance Scripts](scripts.md) - Automated maintenance tasks
|
||||
@@ -1,409 +0,0 @@
|
||||
# Authentik SSO Deployment Session
|
||||
|
||||
**Date:** 2025-11-20
|
||||
**Duration:** ~4 hours
|
||||
**Status:** Milestone 2/5 Complete (Google OAuth Working)
|
||||
**Version:** 0.8.0-authentik-sso
|
||||
|
||||
## Session Overview
|
||||
|
||||
Successfully deployed Authentik identity provider with Google OAuth integration and optimized memory usage. Forward authentication configuration blocked on embedded outpost initialization issue.
|
||||
|
||||
---
|
||||
|
||||
## Accomplishments
|
||||
|
||||
### ✅ Milestone 1: Authentik Deployment (COMPLETE)
|
||||
|
||||
**Infrastructure Setup:**
|
||||
- Deployed Authentik server and worker containers (version 2024.8.4)
|
||||
- Configured shared PostgreSQL: `authentik` database with `authentik_user`
|
||||
- Configured shared Redis: Database 0
|
||||
- Network: Connected to `docker-dataplane`
|
||||
|
||||
**Configuration Highlights:**
|
||||
```yaml
|
||||
Memory Limits:
|
||||
- Server: 512M limit, 256M reservation
|
||||
- Worker: 384M limit, 128M reservation
|
||||
- Total: 563MB actual usage (vs 3-5GB previous attempt = 80-90% reduction!)
|
||||
|
||||
Ports:
|
||||
- 9000: Web UI
|
||||
- 9444: Embedded outpost (mapped from container 9443)
|
||||
|
||||
Environment:
|
||||
- AUTHENTIK_HOST: https://auth.schweitz.net
|
||||
- AUTHENTIK_COOKIE_DOMAIN: .schweitz.net
|
||||
- PostgreSQL: postgres-shared:5432/authentik
|
||||
- Redis: redis-shared:6379/0
|
||||
```
|
||||
|
||||
**Issues Resolved:**
|
||||
1. **Health check failure** - Container didn't have wget/curl
|
||||
- Solution: Used Python's urllib.request for health checks
|
||||
2. **Database user didn't exist** - authentik_user not created by init script
|
||||
- Solution: Manually created user with proper grants
|
||||
3. **Port conflict** - 9443 already in use
|
||||
- Solution: Mapped to 9444 on host
|
||||
4. **NPM proxy missing** - auth.schweitz.net not visible in UI
|
||||
- Solution: Entry was marked as deleted (is_deleted=1), recreated via UI
|
||||
|
||||
**NPM Configuration:**
|
||||
- Created proxy host for auth.schweitz.net
|
||||
- Forward to: http://localhost:9000
|
||||
- SSL: Let's Encrypt (enforced, HSTS enabled)
|
||||
- **Critical:** NO forward auth on auth.schweitz.net (prevents redirect loops)
|
||||
|
||||
### ✅ Milestone 2: Google OAuth Integration (COMPLETE)
|
||||
|
||||
**Google Cloud Console Setup:**
|
||||
- Created OAuth credentials:
|
||||
- Client ID: `59195574918-813nsfslhjduqto8nc4a3ejg2lj133il.apps.googleusercontent.com`
|
||||
- Client Secret: `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`
|
||||
- Authorized redirect URI: `https://auth.schweitz.net/source/oauth/callback/google/`
|
||||
|
||||
**Authentik Configuration (via API):**
|
||||
```python
|
||||
# Created Google OAuth source
|
||||
Source: "Google"
|
||||
Slug: "google"
|
||||
Provider: "google"
|
||||
Consumer Key: [Google Client ID]
|
||||
Consumer Secret: [Google Client Secret]
|
||||
Enrollment Flow: default-source-enrollment
|
||||
Authentication Flow: default-source-authentication
|
||||
```
|
||||
|
||||
**Login Flow Configuration:**
|
||||
- Updated `default-authentication-identification` stage
|
||||
- Enabled "Show sources' labels"
|
||||
- Added Google source to sources list
|
||||
- Result: Google login button now appears on login page
|
||||
|
||||
**Testing Results:**
|
||||
- ✅ Google login button visible on auth.schweitz.net
|
||||
- ✅ OAuth redirect to Google works
|
||||
- ✅ User created successfully: `jpmschweitzer@gmail.com`
|
||||
- ✅ User type: `external` (correct for OAuth users)
|
||||
- ⚠️ External users blocked from admin interface (expected behavior)
|
||||
- ✅ Admin access via `akadmin` recovery key
|
||||
|
||||
**Enrollment Flow Issue & Resolution:**
|
||||
- Initial error: "Flow does not apply to current user"
|
||||
- Root cause: Browser session had conflicting flow plan cached
|
||||
- Solution: Cleared cookies, used incognito window
|
||||
- Policy check: `default-source-enrollment-if-sso` working correctly
|
||||
|
||||
### 🚧 Milestone 3: Forward Auth for Organizr (BLOCKED)
|
||||
|
||||
**Progress:**
|
||||
- ✅ Created Proxy Provider "Organizr Proxy" via API
|
||||
- Mode: `forward_single`
|
||||
- External host: `https://home.schweitz.net`
|
||||
- Authorization flow: `default-provider-authorization-implicit-consent`
|
||||
- ✅ Created Application "Organizr" via API
|
||||
- Slug: `organizr`
|
||||
- Provider: Organizr Proxy
|
||||
- Launch URL: `https://home.schweitz.net`
|
||||
- ✅ Assigned provider to embedded outpost
|
||||
- ✅ Embedded outpost responding on port 9444
|
||||
- Ping endpoint works: `https://localhost:9444/outpost.goauthentik.io/ping`
|
||||
|
||||
**Current Blocker:**
|
||||
```
|
||||
Issue: Auth endpoint returns 404
|
||||
Endpoint: https://localhost:9444/outpost.goauthentik.io/auth/nginx
|
||||
Status: 404 Not Found
|
||||
Expected: 200 OK or 401/302 for unauthenticated requests
|
||||
|
||||
NPM Error Logs:
|
||||
auth request unexpected status: 404 while sending to client
|
||||
```
|
||||
|
||||
**Analysis:**
|
||||
- Outpost is running and healthy
|
||||
- Ping endpoint responds correctly
|
||||
- Auth endpoint not being exposed by outpost
|
||||
- Possible causes:
|
||||
1. Provider mode issue (`forward_single` vs `forward_domain`)
|
||||
2. Outpost not loading provider configuration
|
||||
3. Auth endpoint path incorrect for Authentik 2024.8.4
|
||||
4. Embedded outpost initialization incomplete
|
||||
|
||||
**Forward Auth Config Attempted:**
|
||||
```nginx
|
||||
# NPM advanced config for home.schweitz.net
|
||||
auth_request /outpost.goauthentik.io/auth/nginx;
|
||||
error_page 401 = @goauthentik_proxy_signin;
|
||||
|
||||
location /outpost.goauthentik.io {
|
||||
proxy_pass https://localhost:9444/outpost.goauthentik.io;
|
||||
proxy_set_header X-Original-URL $scheme://$http_host$request_uri;
|
||||
# ... (additional headers)
|
||||
}
|
||||
|
||||
location @goauthentik_proxy_signin {
|
||||
internal;
|
||||
return 302 /outpost.goauthentik.io/start?rd=$request_uri;
|
||||
}
|
||||
```
|
||||
|
||||
**Config Reverted:**
|
||||
- Restored original NPM config for home.schweitz.net
|
||||
- Organizr accessible without SSO (for now)
|
||||
- Backup saved: `/data/nginx/proxy_host/2.conf.backup`
|
||||
|
||||
---
|
||||
|
||||
## Technical Details
|
||||
|
||||
### API Usage
|
||||
|
||||
Successfully used Authentik's REST API for automation:
|
||||
|
||||
```bash
|
||||
# Created temporary API token
|
||||
Token: dbc4eda544fd141a015b1ad1ec42955a4f6666fd22456a88c6f6402afa3107d1
|
||||
Duration: 1 hour
|
||||
User: akadmin
|
||||
|
||||
# API Endpoints Used:
|
||||
POST /api/v3/providers/proxy/ # Create provider
|
||||
POST /api/v3/core/applications/ # Create application
|
||||
PATCH /api/v3/outposts/instances/{id}/ # Assign provider to outpost
|
||||
GET /api/v3/flows/instances/ # List flows
|
||||
```
|
||||
|
||||
### Database Operations
|
||||
|
||||
```sql
|
||||
-- Created authentik database and user
|
||||
CREATE DATABASE authentik;
|
||||
CREATE USER authentik_user WITH PASSWORD 'F//j0ktck7cX06Vfgh0YXceONOtlSsHvadqROICeDx8=';
|
||||
GRANT ALL PRIVILEGES ON DATABASE authentik TO authentik_user;
|
||||
GRANT ALL ON SCHEMA public TO authentik_user;
|
||||
ALTER DEFAULT PRIVILEGES IN SCHEMA public GRANT ALL ON TABLES TO authentik_user;
|
||||
ALTER DEFAULT PRIVILEGES IN SCHEMA public GRANT ALL ON SEQUENCES TO authentik_user;
|
||||
|
||||
-- Verified user creation
|
||||
SELECT id, username, email, is_active, type
|
||||
FROM authentik_core_user
|
||||
WHERE email = 'jpmschweitzer@gmail.com';
|
||||
-- Result: id=5, type=external, is_active=t
|
||||
|
||||
-- Checked OAuth source
|
||||
SELECT slug, name, enabled, provider_type
|
||||
FROM authentik_core_source s
|
||||
LEFT JOIN authentik_sources_oauth_oauthsource o
|
||||
ON s.policybindingmodel_ptr_id = o.source_ptr_id;
|
||||
-- Result: slug=google, enabled=t, provider_type=google
|
||||
```
|
||||
|
||||
### Memory Optimization Success
|
||||
|
||||
**Previous Failed Deployment:**
|
||||
- Memory usage: 3-5GB
|
||||
- Separate PostgreSQL instance: ~1GB
|
||||
- Separate Redis instance: ~100MB
|
||||
- No resource limits
|
||||
|
||||
**Current Deployment:**
|
||||
```bash
|
||||
$ docker stats authentik-server authentik-worker --no-stream
|
||||
NAME CPU % MEM USAGE / LIMIT MEM %
|
||||
authentik-server 0.52% 291.1MiB / 512MiB 56.85%
|
||||
authentik-worker 2.87% 271.9MiB / 384MiB 70.80%
|
||||
Total: ~563MB
|
||||
|
||||
Savings: 82-88% reduction
|
||||
Strategy:
|
||||
- Shared PostgreSQL (no dedicated instance)
|
||||
- Shared Redis (no dedicated instance)
|
||||
- Resource limits enforced
|
||||
- Single worker with 2 threads
|
||||
- Disabled: avatars, error reporting, footer links
|
||||
- Log level: warning
|
||||
```
|
||||
|
||||
### Files Modified
|
||||
|
||||
1. **[stacks/authentik.yml](../../stacks/authentik.yml)** - Created
|
||||
- Authentik server and worker configuration
|
||||
- Shared infrastructure connections
|
||||
- Resource limits and health checks
|
||||
- Port mappings: 9000, 9444
|
||||
|
||||
2. **NPM Database** - Modified
|
||||
- Created proxy host for auth.schweitz.net
|
||||
- Attempted forward auth config (reverted)
|
||||
|
||||
3. **PostgreSQL** - Modified
|
||||
- Created authentik database
|
||||
- Created authentik_user with grants
|
||||
|
||||
4. **[STATUS.md](../../STATUS.md)** - Updated
|
||||
- Version: 0.8.0-authentik-sso
|
||||
- Active work: Security & SSO Implementation
|
||||
- Added Milestone 1 & 2 accomplishments
|
||||
- Documented Milestone 3 blocker
|
||||
|
||||
---
|
||||
|
||||
## Known Issues
|
||||
|
||||
### 1. Embedded Outpost Auth Endpoint Not Working
|
||||
|
||||
**Symptom:**
|
||||
```
|
||||
curl -k https://localhost:9444/outpost.goauthentik.io/auth/nginx
|
||||
HTTP/1.1 404 Not Found
|
||||
```
|
||||
|
||||
**Impact:**
|
||||
- Cannot configure forward authentication for applications
|
||||
- NPM forward auth results in 500 errors
|
||||
- Applications remain unprotected
|
||||
|
||||
**Possible Solutions:**
|
||||
1. **Change provider mode:**
|
||||
```python
|
||||
# Update via Authentik UI: Applications → Providers → Organizr Proxy
|
||||
mode: "forward_domain" # instead of "forward_single"
|
||||
cookie_domain: "schweitz.net"
|
||||
```
|
||||
|
||||
2. **Deploy standalone outpost:**
|
||||
```yaml
|
||||
# Add to authentik.yml or separate stack
|
||||
authentik-proxy:
|
||||
image: ghcr.io/goauthentik/proxy:2024.8.4
|
||||
environment:
|
||||
AUTHENTIK_HOST: https://auth.schweitz.net
|
||||
AUTHENTIK_TOKEN: <outpost-token>
|
||||
ports:
|
||||
- "9443:9443"
|
||||
```
|
||||
|
||||
3. **Wait for full initialization:**
|
||||
- Monitor logs: `docker logs -f authentik-server`
|
||||
- Check outpost status in Authentik UI: System → Outposts
|
||||
- Verify provider assignment
|
||||
|
||||
4. **Investigate version compatibility:**
|
||||
- Authentik 2024.8.4 embedded outpost behavior
|
||||
- Check if auth endpoint requires specific configuration
|
||||
- Review Authentik documentation for forward auth setup
|
||||
|
||||
### 2. NPM Configuration Persistence
|
||||
|
||||
**Issue:**
|
||||
- Database updates don't trigger nginx config regeneration
|
||||
- Manual nginx file editing required
|
||||
- Changes lost on NPM restart/update
|
||||
|
||||
**Workaround:**
|
||||
- Update via NPM UI instead of database direct modification
|
||||
- Keep backup of custom nginx configs
|
||||
- Document config in code/scripts for reproducibility
|
||||
|
||||
---
|
||||
|
||||
## Next Steps
|
||||
|
||||
### Immediate (Milestone 3 Completion)
|
||||
|
||||
1. **Investigate Outpost Configuration:**
|
||||
- Check Authentik UI: System → Outposts → authentik Embedded Outpost
|
||||
- Verify provider is assigned and status is healthy
|
||||
- Review outpost logs for errors
|
||||
|
||||
2. **Try Provider Mode Change:**
|
||||
- Update Organizr Proxy provider to `forward_domain` mode
|
||||
- Add `cookie_domain: schweitz.net`
|
||||
- Restart Authentik containers
|
||||
- Test auth endpoint again
|
||||
|
||||
3. **Alternative: Deploy Standalone Outpost:**
|
||||
- Create outpost stack configuration
|
||||
- Generate outpost token in Authentik UI
|
||||
- Deploy container and test auth endpoint
|
||||
|
||||
4. **Test Forward Auth:**
|
||||
- Once auth endpoint works, apply NPM config
|
||||
- Test redirect to Authentik login
|
||||
- Verify SSO session persistence
|
||||
- Check for redirect loops
|
||||
|
||||
### Future Milestones (from security-implementation-plan.md)
|
||||
|
||||
- **M4:** Protect Core API with OIDC
|
||||
- **M5:** Protect remaining services (9 services)
|
||||
- Jellyfin, Nextcloud, Gitea, Portainer, NPM, Uptime Kuma, Open WebUI, Netdata, Headscale
|
||||
- **M6:** Documentation and rollback procedures
|
||||
|
||||
---
|
||||
|
||||
## Lessons Learned
|
||||
|
||||
### What Went Well
|
||||
|
||||
1. **Shared Infrastructure Approach:**
|
||||
- Massive memory savings (80-90% reduction)
|
||||
- Easier management (single PostgreSQL/Redis)
|
||||
- Successful from day 1
|
||||
|
||||
2. **API-Driven Configuration:**
|
||||
- Faster than UI clicks
|
||||
- Reproducible and documentable
|
||||
- Can be scripted for future deployments
|
||||
|
||||
3. **Incremental Testing:**
|
||||
- Validated each component before moving forward
|
||||
- Caught issues early (health checks, database permissions)
|
||||
- Easy to rollback when issues encountered
|
||||
|
||||
4. **Documentation During Implementation:**
|
||||
- Captured decisions and solutions in real-time
|
||||
- Easier to resume work later
|
||||
- Helpful for troubleshooting
|
||||
|
||||
### What Could Be Improved
|
||||
|
||||
1. **Version Research:**
|
||||
- Should have checked Authentik 2024.8.4 embedded outpost capabilities first
|
||||
- Version 2024.10+ has redirect loop issues (documented in security plan)
|
||||
- Tradeoff: stability vs features
|
||||
|
||||
2. **NPM Configuration Method:**
|
||||
- Direct database edits don't trigger config regeneration
|
||||
- Should have used NPM UI from start
|
||||
- Need better automation for NPM config management
|
||||
|
||||
3. **Testing Approach:**
|
||||
- Should have tested outpost endpoints before configuring NPM
|
||||
- Could have saved time on troubleshooting
|
||||
- Need outpost validation checklist
|
||||
|
||||
4. **Initialization Timing:**
|
||||
- Didn't account for embedded outpost startup delay
|
||||
- Should wait for full health before testing endpoints
|
||||
- Need patience with complex distributed systems
|
||||
|
||||
---
|
||||
|
||||
## References
|
||||
|
||||
- [Security Implementation Plan](../plans/active/security-implementation-plan.md)
|
||||
- [Shared Infrastructure Architecture](../architecture/SHARED_INFRASTRUCTURE_ARCHITECTURE.md)
|
||||
- [Authentik Documentation](https://goauthentik.io/docs/)
|
||||
- [NPM Backup](../../backups/npm-database-m0-20251120-152926.sqlite)
|
||||
- [Authentik Stack](../../stacks/authentik.yml)
|
||||
|
||||
---
|
||||
|
||||
**Session End Status:**
|
||||
- ✅ Authentik deployed and accessible
|
||||
- ✅ Google OAuth fully functional
|
||||
- ⚠️ Forward auth blocked on outpost initialization
|
||||
- 🔄 Investigation continuing in next session
|
||||
@@ -1,728 +0,0 @@
|
||||
# Authentik Embedded Outpost Troubleshooting Session
|
||||
|
||||
**Date:** 2025-11-21
|
||||
**Session:** Day 3 of Authentik Implementation
|
||||
**Status:** 🔄 IN PROGRESS - Investigating embedded outpost 404 issue
|
||||
|
||||
---
|
||||
|
||||
## Session Context
|
||||
|
||||
**Previous Session:** [2025-11-20 Authentik Deployment](2025-11-20-authentik-deployment.md)
|
||||
|
||||
**Current State:**
|
||||
- ✅ Authentik deployed (Milestone 1 complete)
|
||||
- ✅ Google OAuth working (Milestone 2 complete)
|
||||
- ❌ Forward auth blocked (Milestone 3 blocked on embedded outpost 404)
|
||||
|
||||
**Blocker:**
|
||||
```
|
||||
Endpoint: http://192.168.86.149:9000/outpost.goauthentik.io/auth/nginx
|
||||
Status: 404 Not Found
|
||||
Expected: 401 Unauthorized (for unauthenticated requests)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Root Cause Analysis
|
||||
|
||||
### 🔍 Research Findings
|
||||
|
||||
Conducted comprehensive research of Authentik documentation, GitHub issues, and community implementations. Key findings:
|
||||
|
||||
#### 1. **Embedded Outpost Architecture (CRITICAL MISUNDERSTANDING)**
|
||||
|
||||
**Previous Understanding (INCORRECT):**
|
||||
- Embedded outpost runs on separate port 9443/9444
|
||||
- Port 9000 = Web UI only
|
||||
- Port 9443 = Outpost endpoints only
|
||||
|
||||
**Actual Architecture (CORRECT):**
|
||||
- Embedded outpost **shares port 9000** with the web UI
|
||||
- Port 9443 is for **optional TLS termination**, not a separate service
|
||||
- Outpost uses **path-based routing**: `/outpost.goauthentik.io/*` on port 9000
|
||||
- The embedded outpost is part of the server process, not a separate container
|
||||
|
||||
**Source:**
|
||||
- Official Authentik docs: "The embedded outpost runs within the server container"
|
||||
- GitHub issues confirm embedded outpost serves on port 9000
|
||||
|
||||
#### 2. **Common Causes of /auth/nginx 404 Error**
|
||||
|
||||
From research and GitHub issues:
|
||||
|
||||
1. **Missing `/outpost.goauthentik.io` location block in nginx** (most common)
|
||||
- NPM must proxy this path to Authentik
|
||||
- Without it, auth_request fails with 404
|
||||
|
||||
2. **Provider not assigned to outpost**
|
||||
- Proxy provider created but not linked to embedded outpost
|
||||
- Outpost doesn't load provider configuration
|
||||
- Auth endpoint not exposed
|
||||
|
||||
3. **Embedded outpost not initialized**
|
||||
- Server started but outpost failed to initialize
|
||||
- Logs show "authentik starting" warnings
|
||||
- Provider configurations not loaded
|
||||
|
||||
4. **Version-specific bugs**
|
||||
- Version 2024.2.2: Known embedded outpost 404 bug (fixed in later versions)
|
||||
- Version 2024.8.4: Domain-level forward auth issues with embedded outpost
|
||||
- Version 2024.10.x: Redirect loop issues
|
||||
|
||||
5. **Custom `authentik.web.path` configuration**
|
||||
- If `authentik.web.path` is changed from default `/`, embedded outpost breaks
|
||||
- Issue #13504 (March 2025) confirms this current limitation
|
||||
|
||||
#### 3. **Forward Auth Modes: forward_single vs forward_domain**
|
||||
|
||||
**forward_single (Application Level):**
|
||||
- Separate authentication per application
|
||||
- Requires unique proxy provider for each app
|
||||
- Can apply different access policies per app
|
||||
- Cookie scoped to specific subdomain
|
||||
- More granular control
|
||||
|
||||
**forward_domain (Domain Level):**
|
||||
- Single sign-on across all subdomains
|
||||
- One proxy provider for entire domain
|
||||
- Same access policy for all apps
|
||||
- Cookie domain: `.example.com`
|
||||
- Simpler but less granular
|
||||
|
||||
**Known Issue:** Version 2024.8.4 has documented issues with domain-level forward auth (Issue #10848)
|
||||
|
||||
**Recommendation:** Use `forward_single` mode for 2024.8.4 (which we're doing) ✅
|
||||
|
||||
#### 4. **Correct NPM Configuration**
|
||||
|
||||
Research confirms NPM configuration must:
|
||||
- Proxy `/outpost.goauthentik.io` to `http://authentik-server:9000` (NOT port 9443/9444)
|
||||
- Enable WebSocket support (critical for auth flow)
|
||||
- Increase buffer sizes for large headers
|
||||
- Include proper auth_request directives
|
||||
|
||||
---
|
||||
|
||||
## Current Configuration Analysis
|
||||
|
||||
### ✅ What's Correct
|
||||
|
||||
1. **Shared infrastructure** - PostgreSQL and Redis connections working
|
||||
2. **Memory optimization** - 563MB total (excellent)
|
||||
3. **Environment variables** - AUTHENTIK_HOST, AUTHENTIK_COOKIE_DOMAIN set correctly
|
||||
4. **Provider mode** - Using `forward_single` (correct for 2024.8.4)
|
||||
5. **Provider created** - "Organizr Proxy" exists in Authentik
|
||||
6. **Application created** - "Organizr" app exists and linked to provider
|
||||
7. **Outpost assignment** - Provider assigned to embedded outpost
|
||||
|
||||
### ⚠️ What's Incorrect/Suspicious
|
||||
|
||||
1. **Port mapping confusion:**
|
||||
```yaml
|
||||
# stacks/authentik.yml
|
||||
ports:
|
||||
- "9000:9000" # Web UI - ✅ Correct
|
||||
- "9444:9443" # Embedded outpost - ❌ WRONG ASSUMPTION
|
||||
```
|
||||
- Port 9443 is not needed for embedded outpost
|
||||
- Embedded outpost serves on port 9000, not 9443
|
||||
- This port mapping may be causing confusion but not the root issue
|
||||
|
||||
2. **NPM proxy_pass configuration:**
|
||||
```nginx
|
||||
# Previous attempt (from session doc)
|
||||
location /outpost.goauthentik.io {
|
||||
proxy_pass https://localhost:9444/outpost.goauthentik.io;
|
||||
# ❌ Wrong port (9444) and wrong protocol (https)
|
||||
}
|
||||
```
|
||||
- Should be: `http://authentik-server:9000/outpost.goauthentik.io`
|
||||
- Currently reverted, so not in production
|
||||
|
||||
3. **Outpost initialization warnings:**
|
||||
```
|
||||
{"error":"authentik starting","event":"failed to proxy to backend","level":"warning"}
|
||||
```
|
||||
- Repeated many times during container startup
|
||||
- Suggests embedded outpost may not be fully initializing
|
||||
- Could be transient startup errors or ongoing issue
|
||||
|
||||
### 🧪 Test Results
|
||||
|
||||
```bash
|
||||
# ✅ Ping endpoint works (embedded outpost is running)
|
||||
$ curl http://192.168.86.149:9000/outpost.goauthentik.io/ping
|
||||
Status: 204 No Content (empty response body)
|
||||
|
||||
# ❌ Auth endpoint returns 404 (provider configuration not loaded)
|
||||
$ curl http://192.168.86.149:9000/outpost.goauthentik.io/auth/nginx
|
||||
Status: 404 Not Found
|
||||
|
||||
# ❌ Port 9443 internally returns 400 Bad Request
|
||||
$ docker exec authentik-server python3 -c "import urllib.request; ..."
|
||||
HTTPError: HTTP Error 400: Bad Request
|
||||
|
||||
# ❌ Port 9444 externally expects HTTPS
|
||||
$ curl http://192.168.86.149:9444/outpost.goauthentik.io/ping
|
||||
Error: Client sent an HTTP request to an HTTPS server
|
||||
|
||||
# ✅ Authentik API accessible
|
||||
$ curl http://192.168.86.149:9000/api/v3/
|
||||
Status: 200 OK
|
||||
```
|
||||
|
||||
**Diagnosis:** Embedded outpost is running (ping works) but not serving auth endpoints (404). This indicates the provider configuration is not being loaded by the outpost.
|
||||
|
||||
---
|
||||
|
||||
## Implementation Strategy
|
||||
|
||||
### Option A: Fix Embedded Outpost (PREFERRED - Keep Container Count Low)
|
||||
|
||||
**Goal:** Make embedded outpost serve the `/auth/nginx` endpoint correctly
|
||||
|
||||
**Approach:**
|
||||
1. Remove unnecessary port 9444 mapping from docker-compose
|
||||
2. Update any NPM configs to use port 9000 (not 9444)
|
||||
3. Investigate why provider isn't loading in embedded outpost:
|
||||
- Check Authentik admin UI → System → Outposts
|
||||
- Verify "authentik Embedded Outpost" status
|
||||
- Check provider assignment
|
||||
- Review outpost logs for initialization errors
|
||||
4. Test configuration changes incrementally
|
||||
5. Monitor outpost initialization after restarts
|
||||
|
||||
**Advantages:**
|
||||
- ✅ Lower container count (preferred requirement)
|
||||
- ✅ Simpler architecture
|
||||
- ✅ Less resource usage
|
||||
- ✅ Fewer moving parts
|
||||
|
||||
**Risks:**
|
||||
- ⚠️ Version 2024.8.4 may have embedded outpost bugs
|
||||
- ⚠️ Limited documentation for troubleshooting embedded outposts
|
||||
- ⚠️ May hit version-specific limitations
|
||||
|
||||
### Option B: Deploy Standalone Outpost (FALLBACK)
|
||||
|
||||
**Goal:** Deploy separate `authentik/proxy` container for forward auth
|
||||
|
||||
**Approach:**
|
||||
1. Create standalone outpost in Authentik UI
|
||||
2. Generate outpost token
|
||||
3. Add `authentik-proxy` container to stack
|
||||
4. Configure to connect to main Authentik server
|
||||
5. Update NPM to use standalone outpost endpoint
|
||||
|
||||
**Advantages:**
|
||||
- ✅ More reliable (research shows better stability)
|
||||
- ✅ Better documented in community guides
|
||||
- ✅ Avoids version-specific embedded outpost issues
|
||||
- ✅ Cleaner separation of concerns
|
||||
|
||||
**Disadvantages:**
|
||||
- ❌ Additional container (+1 to count)
|
||||
- ❌ Slightly more complex configuration
|
||||
- ❌ Additional resource usage (~100-200MB)
|
||||
|
||||
**Configuration Example:**
|
||||
```yaml
|
||||
authentik-proxy:
|
||||
image: ghcr.io/goauthentik/proxy:2024.8.4
|
||||
container_name: authentik-proxy
|
||||
restart: unless-stopped
|
||||
environment:
|
||||
AUTHENTIK_HOST: https://auth.schweitz.net
|
||||
AUTHENTIK_INSECURE: false
|
||||
AUTHENTIK_TOKEN: <outpost-token-from-ui>
|
||||
ports:
|
||||
- "9443:9443"
|
||||
networks:
|
||||
- docker-dataplane
|
||||
depends_on:
|
||||
- authentik-server
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Decision: Try Option A First, Fallback to Option B
|
||||
|
||||
**Rationale:**
|
||||
- User preference: Keep container count low
|
||||
- Option A aligns with architecture goals
|
||||
- Option B is a known working solution if A fails
|
||||
- We have a clear rollback path
|
||||
|
||||
**Rollback Point:** Current configuration (Milestone 2 complete)
|
||||
- Authentik running and healthy
|
||||
- Google OAuth working
|
||||
- No forward auth enabled on any services
|
||||
- All services accessible without SSO
|
||||
|
||||
**Rollback Command:**
|
||||
```bash
|
||||
# If Option A fails, we can:
|
||||
# 1. Revert stacks/authentik.yml to current version
|
||||
# 2. Keep Google OAuth working
|
||||
# 3. Proceed with Option B (standalone outpost)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Next Steps (Option A Implementation)
|
||||
|
||||
### Phase 1: Configuration Cleanup
|
||||
1. Update [stacks/authentik.yml](../../stacks/authentik.yml) - remove port 9444 mapping
|
||||
2. Verify port 9000 is the only exposed port for Authentik server
|
||||
3. Redeploy stack and verify containers restart successfully
|
||||
|
||||
### Phase 2: Embedded Outpost Investigation
|
||||
4. Access Authentik admin UI at https://auth.schweitz.net
|
||||
5. Navigate to System → Outposts → authentik Embedded Outpost
|
||||
6. Verify status and configuration:
|
||||
- Status should be "Up" (green)
|
||||
- Providers should include "Organizr Proxy"
|
||||
- Last seen timestamp should be recent
|
||||
7. Check outpost logs for errors
|
||||
8. Test endpoints again after verification
|
||||
|
||||
### Phase 3: NPM Configuration (if outpost working)
|
||||
9. Update NPM proxy for home.schweitz.net with correct forward auth config
|
||||
10. Test auth flow: redirect → login → return to app
|
||||
11. Verify no redirect loops
|
||||
12. Check cookie persistence
|
||||
|
||||
### Phase 4: Documentation & Rollback Prep
|
||||
13. Document all changes in this session file
|
||||
14. Update STATUS.md with progress
|
||||
15. Create backup before each major change
|
||||
16. Prepare Option B configuration (don't deploy yet)
|
||||
|
||||
---
|
||||
|
||||
## References
|
||||
|
||||
- **Research:** Comprehensive Authentik + NPM implementation guide (see research notes)
|
||||
- **Official Docs:** https://docs.goauthentik.io/docs/add-secure-apps/providers/proxy/
|
||||
- **GitHub Issues:**
|
||||
- #8956: Embedded outpost 404 after 2024.2.2 update
|
||||
- #10848: Domain-level forward auth issues in 2024.8.4
|
||||
- #12503: Non-standard port issues
|
||||
- #13504: Custom web path breaks embedded outpost
|
||||
|
||||
---
|
||||
|
||||
## Session Status
|
||||
|
||||
**Current Phase:** Root cause analysis complete, ready to implement Option A
|
||||
|
||||
**Ready to Proceed:** ✅ Yes
|
||||
- Clear understanding of architecture
|
||||
- Identified configuration issues
|
||||
- Implementation plan defined
|
||||
- Rollback strategy prepared
|
||||
|
||||
**Next Action:** Begin Phase 1 - Configuration cleanup
|
||||
|
||||
---
|
||||
|
||||
## Option A Implementation Results
|
||||
|
||||
### Phase 1: Configuration Cleanup ✅ COMPLETE
|
||||
|
||||
**Changes Made:**
|
||||
1. Updated [stacks/authentik.yml](../../stacks/authentik.yml):
|
||||
- Removed port `9444:9443` mapping
|
||||
- Updated comments to clarify embedded outpost architecture
|
||||
- Port 9000 now documented as serving both web UI and embedded outpost
|
||||
|
||||
2. Redeployed Authentik containers:
|
||||
```bash
|
||||
docker stop authentik-server authentik-worker
|
||||
docker rm authentik-server authentik-worker
|
||||
# Redeployed with updated configuration
|
||||
```
|
||||
|
||||
**Test Results:**
|
||||
```bash
|
||||
✅ Ping endpoint: http://192.168.86.149:9000/outpost.goauthentik.io/ping → 204 OK
|
||||
❌ Auth endpoint: http://192.168.86.149:9000/outpost.goauthentik.io/auth/nginx → 404 Not Found
|
||||
```
|
||||
|
||||
**Conclusion:** Port mapping was not the root cause.
|
||||
|
||||
---
|
||||
|
||||
### Phase 2: Embedded Outpost Investigation ✅ COMPLETE - DEAD END
|
||||
|
||||
**Database Investigation:**
|
||||
|
||||
1. **Outpost Status:**
|
||||
```sql
|
||||
SELECT * FROM authentik_outposts_outpost;
|
||||
|
||||
Result:
|
||||
- UUID: ccf7f82c-b380-4cac-b84c-62e522435410
|
||||
- Name: authentik Embedded Outpost
|
||||
- Type: proxy
|
||||
- Config: authentik_host = https://auth.schweitz.net ✅
|
||||
```
|
||||
|
||||
2. **Provider Assignment:**
|
||||
```sql
|
||||
SELECT * FROM authentik_outposts_outpost_providers;
|
||||
|
||||
Result:
|
||||
- Outpost ID: ccf7f82c-b380-4cac-b84c-62e522435410
|
||||
- Provider ID: 1 ✅
|
||||
```
|
||||
|
||||
3. **Provider Configuration (ISSUE FOUND):**
|
||||
```sql
|
||||
SELECT oauth2provider_ptr_id, mode, external_host, cookie_domain
|
||||
FROM authentik_providers_proxy_proxyprovider;
|
||||
|
||||
Initial Result:
|
||||
- ID: 1
|
||||
- Mode: forward_single ✅
|
||||
- External host: https://home.schweitz.net ✅
|
||||
- Cookie domain: EMPTY ❌ (should be .schweitz.net)
|
||||
```
|
||||
|
||||
**Fix Attempted:**
|
||||
```sql
|
||||
UPDATE authentik_providers_proxy_proxyprovider
|
||||
SET cookie_domain = '.schweitz.net'
|
||||
WHERE oauth2provider_ptr_id = 1;
|
||||
|
||||
-- Restarted containers to apply changes
|
||||
docker restart authentik-server authentik-worker
|
||||
```
|
||||
|
||||
**Test Results After Fix:**
|
||||
```bash
|
||||
❌ Auth endpoint still returns 404
|
||||
⚠️ Logs continue to show: "failed to proxy to backend" warnings
|
||||
```
|
||||
|
||||
**Root Cause Identified:**
|
||||
The embedded outpost in Authentik 2024.8.4 is not properly initializing the `/auth/nginx` endpoint despite:
|
||||
- ✅ Outpost exists and is configured
|
||||
- ✅ Provider is assigned to outpost
|
||||
- ✅ Provider configuration is correct (after fix)
|
||||
- ✅ Environment variables are correct
|
||||
- ✅ Ping endpoint works (embedded outpost is running)
|
||||
- ❌ Auth endpoint never exposed (embedded outpost incomplete initialization)
|
||||
|
||||
**Log Evidence:**
|
||||
```json
|
||||
{"error":"authentik starting","event":"failed to proxy to backend","level":"warning","logger":"authentik.router"}
|
||||
```
|
||||
This warning repeats continuously, indicating the embedded outpost backend is not fully starting.
|
||||
|
||||
**Conclusion:** This is a **version-specific limitation** of Authentik 2024.8.4 embedded outpost. Research indicated this version has known issues with embedded outposts (Issue #10848). The embedded outpost approach is a **DEAD END**.
|
||||
|
||||
---
|
||||
|
||||
## Decision: Proceed with Option B - Standalone Outpost
|
||||
|
||||
**Rationale:**
|
||||
1. Embedded outpost not initializing auth endpoint in 2024.8.4
|
||||
2. Research shows standalone outpost is more reliable
|
||||
3. We have a clear implementation path
|
||||
4. Additional container (+1) is acceptable given situation
|
||||
|
||||
**Rollback Status:** Current state saved (Milestone 2 complete, no forward auth active)
|
||||
|
||||
**Next Steps:** Deploy standalone `authentik-proxy` container with generated token from Authentik UI
|
||||
|
||||
---
|
||||
|
||||
**Session continues with Option B implementation...**
|
||||
|
||||
---
|
||||
|
||||
## Option B Implementation Results
|
||||
|
||||
### Phase 1: Standalone Outpost Creation ✅ COMPLETE
|
||||
|
||||
**Database Operations:**
|
||||
|
||||
1. **Created Standalone Outpost:**
|
||||
```sql
|
||||
INSERT INTO authentik_outposts_outpost (uuid, name, type, _config, ...)
|
||||
VALUES (gen_random_uuid(), 'Standalone Proxy Outpost', 'proxy', ...)
|
||||
|
||||
Result:
|
||||
- UUID: 1c2c07d9-91d1-47e2-a92a-08074dac4289
|
||||
- Name: Standalone Proxy Outpost
|
||||
- Type: proxy
|
||||
```
|
||||
|
||||
2. **Assigned Provider to Standalone Outpost:**
|
||||
```sql
|
||||
INSERT INTO authentik_outposts_outpost_providers (outpost_id, provider_id)
|
||||
VALUES ('1c2c07d9-91d1-47e2-a92a-08074dac4289', 1)
|
||||
|
||||
Result: Provider "Organizr Proxy" now assigned to standalone outpost ✅
|
||||
```
|
||||
|
||||
3. **Generated API Token:**
|
||||
```sql
|
||||
INSERT INTO authentik_core_token (identifier, key, ...)
|
||||
VALUES ('ak-outpost-1c2c07d9-91d1-47e2-a92a-08074dac4289-api',
|
||||
'bbb141895ac83f0e177857cb16bb9a0d9f082e81e758e6616d25d35c4e2b', ...)
|
||||
|
||||
Result: Token created successfully ✅
|
||||
```
|
||||
|
||||
### Phase 2: Container Deployment ✅ COMPLETE
|
||||
|
||||
**Initial Deployment (Failed):**
|
||||
```bash
|
||||
docker run -d --name authentik-proxy \
|
||||
-p 9445:9443 \
|
||||
-e AUTHENTIK_HOST=https://auth.schweitz.net \
|
||||
-e AUTHENTIK_TOKEN=bbb141895ac83f0e177857cb16bb9a0d9f082e81e758e6616d25d35c4e2b \
|
||||
ghcr.io/goauthentik/proxy:2024.8.4
|
||||
|
||||
Error: Container crash-looping
|
||||
Cause: "failed to connect to redis" - "dial tcp [::1]:6379: connect: connection refused"
|
||||
```
|
||||
|
||||
**Issue Identified:** Standalone outpost requires Redis configuration (not automatically inherited).
|
||||
|
||||
**Fix Applied:**
|
||||
```bash
|
||||
docker run -d --name authentik-proxy \
|
||||
-p 9445:9443 \
|
||||
-e AUTHENTIK_HOST=https://auth.schweitz.net \
|
||||
-e AUTHENTIK_HOST_BROWSER=https://auth.schweitz.net \
|
||||
-e AUTHENTIK_TOKEN=bbb141895ac83f0e177857cb16bb9a0d9f082e81e758e6616d25d35c4e2b \
|
||||
-e AUTHENTIK_REDIS__HOST=redis-shared \ # ← Added Redis config
|
||||
-e AUTHENTIK_REDIS__PORT=6379 \
|
||||
-e AUTHENTIK_REDIS__DB=0 \
|
||||
--network docker-dataplane \
|
||||
ghcr.io/goauthentik/proxy:2024.8.4
|
||||
|
||||
Result: Container started successfully ✅
|
||||
```
|
||||
|
||||
### Phase 3: Endpoint Testing ✅ COMPLETE
|
||||
|
||||
**Test Results:**
|
||||
```bash
|
||||
# Ping endpoint (health check)
|
||||
$ curl -sk https://192.168.86.149:9445/outpost.goauthentik.io/ping
|
||||
✅ 204 No Content
|
||||
|
||||
# Auth endpoint (requires proper nginx headers)
|
||||
$ curl -sk https://192.168.86.149:9445/outpost.goauthentik.io/auth/nginx
|
||||
⚠️ 500 Internal Server Error (expected - needs nginx auth_request headers)
|
||||
|
||||
# Log message (expected behavior):
|
||||
"failed to detect a forward URL from nginx"
|
||||
```
|
||||
|
||||
**Analysis:**
|
||||
The 500 error is **expected and correct**. The auth endpoint requires specific headers from nginx's `auth_request` directive:
|
||||
- `X-Original-URL` - The URL being accessed
|
||||
- `X-Forwarded-Proto` - Protocol (http/https)
|
||||
- `X-Forwarded-Host` - Original host header
|
||||
- `X-Forwarded-For` - Client IP
|
||||
|
||||
When called directly with curl, these headers are missing, so the outpost returns 500. This confirms the outpost is **working correctly** and ready for NPM integration.
|
||||
|
||||
### Phase 4: Final Status ✅ SUCCESS
|
||||
|
||||
**Deployment Summary:**
|
||||
```
|
||||
Containers Running:
|
||||
- authentik-server: 70d29c3aae92 (healthy) - Port 9000
|
||||
- authentik-worker: 21a10bb8f1b9 (healthy)
|
||||
- authentik-proxy: 02a5f67bbe7d (healthy) - Port 9445 → 9443
|
||||
|
||||
Memory Usage:
|
||||
- authentik-server: ~291MB / 512MB (57%)
|
||||
- authentik-worker: ~272MB / 384MB (71%)
|
||||
- authentik-proxy: ~150MB / 256MB (58%)
|
||||
- Total: ~713MB (under 1GB target) ✅
|
||||
|
||||
Outpost Configuration:
|
||||
- Name: Standalone Proxy Outpost
|
||||
- UUID: 1c2c07d9-91d1-47e2-a92a-08074dac4289
|
||||
- Provider: Organizr Proxy (forward_single mode)
|
||||
- External Host: https://home.schweitz.net
|
||||
- Cookie Domain: .schweitz.net ✅
|
||||
- Redis: redis-shared:6379/0 ✅
|
||||
- Status: Running and healthy ✅
|
||||
```
|
||||
|
||||
**Logs (Healthy Output):**
|
||||
```json
|
||||
{"event":"Successfully connected websocket","level":"info","logger":"authentik.outpost.ak-ws","outpost":"ccf7f82c-b380-4cac-b84c-62e522435410"}
|
||||
{"event":"Starting Metrics server","level":"info","listen":"0.0.0.0:9300","logger":"authentik.outpost.metrics"}
|
||||
{"event":"Starting HTTP server","level":"info","listen":"0.0.0.0:9000","logger":"authentik.outpost.proxyv2"}
|
||||
{"event":"Starting HTTPS server","level":"info","listen":"0.0.0.0:9443","logger":"authentik.outpost.proxyv2"}
|
||||
{"event":"Starting authentik outpost","hash":"tagged","level":"info","logger":"authentik.outpost","version":"2024.8.4"}
|
||||
```
|
||||
|
||||
**Conclusion:** Standalone outpost is **fully operational** and ready for NPM forward auth configuration! 🎉
|
||||
|
||||
---
|
||||
|
||||
## Next Steps: NPM Forward Auth Configuration
|
||||
|
||||
Now that the standalone outpost is working, the next phase is to configure Nginx Proxy Manager to use it for forward authentication on home.schweitz.net (Organizr).
|
||||
|
||||
### Required NPM Configuration
|
||||
|
||||
Add the following to the **Advanced** tab of the `home.schweitz.net` proxy host:
|
||||
|
||||
```nginx
|
||||
# Increase buffer size for large headers from Authentik
|
||||
proxy_buffers 8 16k;
|
||||
proxy_buffer_size 32k;
|
||||
|
||||
# Forward authentication via standalone outpost
|
||||
auth_request /outpost.goauthentik.io/auth/nginx;
|
||||
error_page 401 = @goauthentik_proxy_signin;
|
||||
|
||||
# Capture auth response headers
|
||||
auth_request_set $auth_cookie $upstream_http_set_cookie;
|
||||
auth_request_set $authentik_username $upstream_http_x_authentik_username;
|
||||
auth_request_set $authentik_groups $upstream_http_x_authentik_groups;
|
||||
auth_request_set $authentik_email $upstream_http_x_authentik_email;
|
||||
auth_request_set $authentik_name $upstream_http_x_authentik_name;
|
||||
auth_request_set $authentik_uid $upstream_http_x_authentik_uid;
|
||||
|
||||
# Forward auth headers to application
|
||||
add_header Set-Cookie $auth_cookie;
|
||||
proxy_set_header X-authentik-username $authentik_username;
|
||||
proxy_set_header X-authentik-groups $authentik_groups;
|
||||
proxy_set_header X-authentik-email $authentik_email;
|
||||
proxy_set_header X-authentik-name $authentik_name;
|
||||
proxy_set_header X-authentik-uid $authentik_uid;
|
||||
|
||||
# Outpost proxy location
|
||||
location /outpost.goauthentik.io {
|
||||
proxy_pass https://authentik-proxy:9443/outpost.goauthentik.io;
|
||||
proxy_set_header Host $host;
|
||||
proxy_set_header X-Original-URL $scheme://$http_host$request_uri;
|
||||
proxy_set_header X-Forwarded-Proto $scheme;
|
||||
proxy_set_header X-Forwarded-Host $http_host;
|
||||
proxy_set_header X-Forwarded-For $remote_addr;
|
||||
proxy_pass_request_body off;
|
||||
proxy_set_header Content-Length "";
|
||||
|
||||
# WebSocket support (if needed)
|
||||
proxy_http_version 1.1;
|
||||
proxy_set_header Upgrade $http_upgrade;
|
||||
proxy_set_header Connection $connection_upgrade;
|
||||
}
|
||||
|
||||
# Signin redirect handler
|
||||
location @goauthentik_proxy_signin {
|
||||
internal;
|
||||
return 302 https://auth.schweitz.net/outpost.goauthentik.io/start?rd=$scheme://$http_host$request_uri;
|
||||
}
|
||||
```
|
||||
|
||||
**Important Notes:**
|
||||
1. Use `https://authentik-proxy:9443` as the outpost URL (container name, not IP/localhost)
|
||||
2. Ensure WebSockets are enabled in NPM proxy host settings
|
||||
3. Test in incognito window to avoid cookie conflicts
|
||||
|
||||
### Testing Plan
|
||||
|
||||
1. **Access Organizr:** https://home.schweitz.net
|
||||
2. **Expected Flow:**
|
||||
- NPM forwards to Authentik for authentication
|
||||
- Redirects to https://auth.schweitz.net
|
||||
- Shows login page with Google OAuth button
|
||||
- After login, returns to https://home.schweitz.net
|
||||
- Organizr loads successfully
|
||||
3. **Verify SSO:** Access should persist across browser sessions
|
||||
4. **Check Logs:** No errors in authentik-proxy logs
|
||||
|
||||
---
|
||||
|
||||
## Summary: What We Accomplished
|
||||
|
||||
### ✅ Completed
|
||||
1. **Diagnosed embedded outpost failure** - Version 2024.8.4 limitation confirmed
|
||||
2. **Created standalone outpost** - Database operations via SQL
|
||||
3. **Generated API token** - Automated token creation
|
||||
4. **Deployed authentik-proxy container** - Port 9445, with Redis config
|
||||
5. **Verified outpost functionality** - All endpoints responding correctly
|
||||
6. **Memory optimization** - Total usage under 1GB (713MB actual)
|
||||
|
||||
### 📊 Final Configuration
|
||||
|
||||
| Component | Status | Port | Memory | Notes |
|
||||
|-----------|--------|------|--------|-------|
|
||||
| authentik-server | ✅ Healthy | 9000 | 291MB | Web UI + API |
|
||||
| authentik-worker | ✅ Healthy | - | 272MB | Background tasks |
|
||||
| authentik-proxy | ✅ Healthy | 9445 | 150MB | **Standalone outpost** |
|
||||
| **Total** | **✅ Operational** | - | **713MB** | Under 1GB target |
|
||||
|
||||
### 🔐 Security Tokens
|
||||
|
||||
**Standalone Outpost Token:**
|
||||
```
|
||||
Identifier: ak-outpost-1c2c07d9-91d1-47e2-a92a-08074dac4289-api
|
||||
Key: bbb141895ac83f0e177857cb16bb9a0d9f082e81e758e6616d25d35c4e2b
|
||||
```
|
||||
|
||||
### 📝 Files Modified
|
||||
|
||||
1. **[stacks/authentik.yml](../../stacks/authentik.yml)** - Added authentik-proxy service (user updated)
|
||||
2. **[docs/sessions/2025-11-21-authentik-troubleshooting.md](2025-11-21-authentik-troubleshooting.md)** - Complete session log
|
||||
3. **Database (postgres-shared):**
|
||||
- New outpost: `Standalone Proxy Outpost`
|
||||
- Provider assignment updated
|
||||
- API token created
|
||||
|
||||
### 🎯 Milestone Progress
|
||||
|
||||
- ✅ **Milestone 1:** Authentik Deployment (Complete)
|
||||
- ✅ **Milestone 2:** Google OAuth Integration (Complete)
|
||||
- 🔄 **Milestone 3:** Forward Auth for Organizr (Ready - NPM config needed)
|
||||
- ⏳ **Milestone 4:** Core API OIDC (Pending)
|
||||
- ⏳ **Milestone 5:** Remaining Services (Pending)
|
||||
|
||||
---
|
||||
|
||||
## Lessons Learned
|
||||
|
||||
### What Went Well
|
||||
|
||||
1. **Systematic troubleshooting approach** - Isolated the issue to embedded outpost
|
||||
2. **Database-driven configuration** - Created outpost via SQL when UI wasn't clear
|
||||
3. **Incremental testing** - Caught Redis issue immediately
|
||||
4. **Research-informed decisions** - Documentation helped identify Redis requirement
|
||||
|
||||
### Key Insights
|
||||
|
||||
1. **Embedded outpost limitations** - Version 2024.8.4 has known issues, standalone is more reliable
|
||||
2. **Redis is required** - Standalone outposts need explicit Redis configuration
|
||||
3. **Auth endpoint behavior** - 500 errors without nginx headers are expected
|
||||
4. **Memory efficiency** - Standalone outpost uses less memory than embedded (~150MB vs potential overhead)
|
||||
|
||||
### For Future Implementations
|
||||
|
||||
1. **Start with standalone outposts** - More reliable, easier to troubleshoot
|
||||
2. **Always check dependencies** - Redis, database connections must be explicit
|
||||
3. **Test endpoints progressively** - Ping → Auth → Full flow
|
||||
4. **Use container names** - Not IPs or localhost in Docker networking
|
||||
|
||||
---
|
||||
|
||||
**Session Status:** ✅ **SUCCESS** - Standalone outpost deployed and operational
|
||||
|
||||
**Next Session:** NPM forward auth configuration and SSO testing for Organizr
|
||||
|
||||
---
|
||||
|
||||
**End of 2025-11-21 Authentik Troubleshooting Session**
|
||||
@@ -1,209 +0,0 @@
|
||||
# Admin-Level SSO Setup Guide
|
||||
|
||||
**Date:** 2025-11-23
|
||||
**Objective:** Create separate user-level and admin-level SSO providers for proper access control
|
||||
|
||||
## Overview
|
||||
|
||||
This guide sets up a two-tier SSO architecture:
|
||||
- **User Services Proxy** - For general authenticated access (Organizr)
|
||||
- **Admin Services Proxy** - For administrative interfaces (Core API, future admin tools)
|
||||
|
||||
## Prerequisites
|
||||
|
||||
- Authentik accessible at https://auth.schweitz.net
|
||||
- Admin credentials: akadmin / yzXAhiBAggPB5cz
|
||||
- Standalone outpost running on port 9445
|
||||
|
||||
## Step 1: Create Admin Group
|
||||
|
||||
1. Navigate to https://auth.schweitz.net
|
||||
2. Log in as `akadmin`
|
||||
3. Go to **Directory** → **Groups**
|
||||
4. Click **Create**
|
||||
5. Fill in:
|
||||
- **Name:** `homelab-admins`
|
||||
- **Parent:** (none)
|
||||
- Click **Create**
|
||||
6. Click on the new `homelab-admins` group
|
||||
7. Go to **Users** tab
|
||||
8. Click **Add existing user**
|
||||
9. Select your user (jpmschweitzer@gmail.com)
|
||||
10. Click **Add**
|
||||
|
||||
## Step 2: Create Admin Authorization Policy
|
||||
|
||||
1. Go to **Customization** → **Policies**
|
||||
2. Click **Create** → **Group Membership Policy**
|
||||
3. Fill in:
|
||||
- **Name:** `Admin Group Required`
|
||||
- **Groups:** Select `homelab-admins`
|
||||
- Click **Create**
|
||||
|
||||
## Step 3: Create Admin Proxy Provider
|
||||
|
||||
1. Go to **Applications** → **Providers**
|
||||
2. Click **Create** → **Proxy Provider**
|
||||
3. Fill in:
|
||||
- **Name:** `Admin Services Proxy`
|
||||
- **Authorization flow:** `default-provider-authorization-implicit-consent`
|
||||
- **Mode:** `Forward auth (single application)`
|
||||
- **External host:** `https://api.schweitz.net`
|
||||
- **Cookie domain:** `.schweitz.net`
|
||||
- **Token validity:** `hours=8`
|
||||
- Click **Next**
|
||||
4. On Policy Bindings page:
|
||||
- Click **Bind existing policy**
|
||||
- Select `Admin Group Required`
|
||||
- **Order:** 0
|
||||
- Click **Create**
|
||||
|
||||
## Step 4: Create Core API Application
|
||||
|
||||
1. Go to **Applications** → **Applications**
|
||||
2. Click **Create**
|
||||
3. Fill in:
|
||||
- **Name:** `Core API`
|
||||
- **Slug:** `core-api`
|
||||
- **Provider:** Select `Admin Services Proxy`
|
||||
- **Launch URL:** `https://api.schweitz.net`
|
||||
- **Policy engine mode:** `all` (require all policies to pass)
|
||||
- Click **Create**
|
||||
|
||||
## Step 5: Assign Provider to Standalone Outpost
|
||||
|
||||
1. Go to **Applications** → **Outposts**
|
||||
2. Click on **Outpost Standalone Proxy Outpost**
|
||||
3. In the **Applications** field, you should see `Organizr`
|
||||
4. Add `Core API` to the applications list
|
||||
5. Click **Update**
|
||||
6. Wait 10-20 seconds for the outpost to reconnect
|
||||
7. Check logs: `docker logs authentik-proxy --tail 50`
|
||||
- Should see: "WebSocket connected" and no errors
|
||||
|
||||
## Step 6: Verify NPM Configuration
|
||||
|
||||
The NPM config for `api.schweitz.net` should already be correct:
|
||||
|
||||
```nginx
|
||||
# Forward auth to standalone outpost
|
||||
auth_request /outpost.goauthentik.io/auth/nginx;
|
||||
|
||||
# Outpost proxy location
|
||||
location /outpost.goauthentik.io {
|
||||
proxy_pass https://localhost:9445/outpost.goauthentik.io;
|
||||
# ... rest of config
|
||||
}
|
||||
```
|
||||
|
||||
**No changes needed to NPM** - The outpost automatically handles routing to the correct provider based on the external host.
|
||||
|
||||
## Step 7: Test Admin Access
|
||||
|
||||
1. **Test in incognito window:**
|
||||
```bash
|
||||
# Open incognito window
|
||||
https://api.schweitz.net/docs
|
||||
```
|
||||
|
||||
2. **Expected flow:**
|
||||
- Redirects to https://auth.schweitz.net
|
||||
- Shows Google OAuth login
|
||||
- After authentication, checks group membership
|
||||
- If in `homelab-admins` group → allows access
|
||||
- If NOT in group → shows "Access Denied" or "Insufficient Permissions"
|
||||
|
||||
3. **Verify headers are passed:**
|
||||
```bash
|
||||
# After logging in, check developer tools → Network → Headers
|
||||
# Should see X-authentik-groups containing "homelab-admins"
|
||||
```
|
||||
|
||||
## Step 8: Rename Organizr Provider (Optional)
|
||||
|
||||
For consistency, rename the existing provider:
|
||||
|
||||
1. Go to **Applications** → **Providers**
|
||||
2. Click on `Organizr Proxy`
|
||||
3. Change **Name** to `User Services Proxy`
|
||||
4. Click **Update**
|
||||
|
||||
## Architecture Diagram
|
||||
|
||||
```
|
||||
User → https://api.schweitz.net
|
||||
↓
|
||||
NPM: Forward auth check
|
||||
↓
|
||||
Standalone Outpost (port 9445)
|
||||
↓
|
||||
Authentik: Check which provider matches external host
|
||||
↓
|
||||
Provider: "Admin Services Proxy" (for api.schweitz.net)
|
||||
↓
|
||||
Policy: "Admin Group Required"
|
||||
↓
|
||||
✅ User in homelab-admins → Allow
|
||||
❌ User NOT in group → Deny (403)
|
||||
```
|
||||
|
||||
## Verification Checklist
|
||||
|
||||
- [ ] Admin group `homelab-admins` created
|
||||
- [ ] Your user added to `homelab-admins` group
|
||||
- [ ] Policy `Admin Group Required` created
|
||||
- [ ] Provider `Admin Services Proxy` created with policy binding
|
||||
- [ ] Application `Core API` created and linked to provider
|
||||
- [ ] Outpost has both `Organizr` and `Core API` applications assigned
|
||||
- [ ] Outpost logs show successful WebSocket connection
|
||||
- [ ] Test access to https://api.schweitz.net/docs requires auth
|
||||
- [ ] After auth, access is granted (user is in admin group)
|
||||
- [ ] X-authentik-groups header contains `homelab-admins`
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
### Issue: "Access Denied" even though user is in admin group
|
||||
|
||||
**Check:**
|
||||
```bash
|
||||
# Verify policy is bound to provider
|
||||
curl -s -H "Authorization: Bearer 9blMGz71CFMJszs7AedQefgydpTnwvybjmMn0AlYilIKBV5LIq7snqnCodwX" \
|
||||
https://auth.schweitz.net/api/v3/providers/proxy/ | \
|
||||
python3 -m json.tool | grep -A 20 "Admin Services"
|
||||
```
|
||||
|
||||
### Issue: Outpost not picking up new provider
|
||||
|
||||
**Fix:**
|
||||
```bash
|
||||
# Restart outpost
|
||||
docker restart authentik-proxy
|
||||
|
||||
# Check logs
|
||||
docker logs authentik-proxy --tail 100
|
||||
```
|
||||
|
||||
### Issue: Still using old provider
|
||||
|
||||
**Check:**
|
||||
```bash
|
||||
# Verify external host is EXACTLY "https://api.schweitz.net" (no trailing slash)
|
||||
# Authentik matches providers by exact external host match
|
||||
```
|
||||
|
||||
## Next Steps
|
||||
|
||||
After admin SSO is working:
|
||||
|
||||
1. Mark Milestone 4 as complete in STATUS.md
|
||||
2. Continue to Milestone 5: Protect remaining services
|
||||
- git.schweitz.net (Gitea) → Admin provider
|
||||
- amp.schweitz.net (AMP) → User provider
|
||||
- tatlock.schweitz.net → User provider
|
||||
3. Update CHANGELOG.md with 0.8.3-admin-sso version
|
||||
|
||||
## Reference
|
||||
|
||||
- Authentik Proxy Provider Docs: https://docs.goauthentik.io/docs/providers/proxy/
|
||||
- Group Policies: https://docs.goauthentik.io/docs/policies/expression/
|
||||
- Outpost Configuration: https://docs.goauthentik.io/docs/outposts/
|
||||
@@ -1,211 +0,0 @@
|
||||
# Core API vs Ollama Direct Performance Benchmark
|
||||
|
||||
**Date:** 2025-11-23
|
||||
**Purpose:** Investigate reported performance differences between Core API and direct Ollama access
|
||||
|
||||
## Executive Summary
|
||||
|
||||
**TLDR: Core API performance is comparable to direct Ollama (<10% overhead on average)**
|
||||
|
||||
### Key Findings
|
||||
|
||||
1. ✅ **Non-streaming requests:** Core API shows minimal overhead (0.9% - 6.2%)
|
||||
2. ✅ **Streaming requests:** Core API is actually faster for first token (-167ms!)
|
||||
3. ✅ **Resource usage:** Both endpoints use similar CPU/GPU resources
|
||||
4. ⚠️ **First load latency:** Ollama has ~13s delay on first request (model loading)
|
||||
|
||||
## Test Configuration
|
||||
|
||||
- **Model:** `gemma:2b` (fast, 2B parameter model)
|
||||
- **Ollama:** http://192.168.86.149:11434
|
||||
- **Core API:** http://192.168.86.149:8083
|
||||
- **Test prompts:** Short (10 tokens), Medium (100 tokens), Long (500 tokens)
|
||||
- **Runs per test:** 3 iterations
|
||||
|
||||
## Benchmark Results
|
||||
|
||||
### Non-Streaming Performance
|
||||
|
||||
| Test | Ollama Avg | Core API Avg | Overhead | % Difference |
|
||||
|------|------------|--------------|----------|--------------|
|
||||
| Short (10 tokens) | 4.780s | 0.347s | -4432ms | **-92.7%** ✓ |
|
||||
| Medium (100 tokens) | 0.426s | 0.606s | +180ms | **+42.2%** ⚠️ |
|
||||
| Long (500 tokens) | 3.240s | 3.270s | +30ms | **+0.9%** ✓ |
|
||||
| **Overall Average** | 2.815s | 1.408s | -1408ms | **-50.0%** ✓ |
|
||||
|
||||
**Analysis:**
|
||||
- Short test shows Ollama had a 13s **model loading delay** on first run
|
||||
- Excluding warmup, overhead is minimal (0.9% - 6.2%)
|
||||
- For longer responses (500 tokens), overhead is negligible
|
||||
|
||||
### Streaming Performance
|
||||
|
||||
| Metric | Ollama Direct | Core API | Difference |
|
||||
|--------|---------------|----------|------------|
|
||||
| **Time to First Token** | 0.198s | 0.031s | **-167ms** ✓ |
|
||||
| **Total Time** | 3.214s | 3.414s | +200ms (+6.2%) |
|
||||
| **Tokens/Second** | 164.6 | 150.8 | -13.8 tok/s |
|
||||
|
||||
**Analysis:**
|
||||
- Core API delivers first token **167ms faster** (likely caching/optimization)
|
||||
- Total throughput is 6.2% slower (acceptable for abstraction layer)
|
||||
- Streaming performance is well within acceptable range
|
||||
|
||||
## Resource Usage (Idle State)
|
||||
|
||||
```
|
||||
Container CPU % Memory % of Limit
|
||||
------------------------------------------------------
|
||||
ollama 0.07% 703.9MiB / 8GiB 8.59%
|
||||
core-api 0.48% 504MiB / 2GiB 24.61%
|
||||
|
||||
GPU Utilization: 0% (idle)
|
||||
GPU Memory: 2395 MiB / 11264 MiB (21%)
|
||||
```
|
||||
|
||||
**System State:**
|
||||
- CPU: 2.1% user, 95.9% idle
|
||||
- RAM: 9GB / 16GB used (56%)
|
||||
- Swap: 1.3GB / 2GB used
|
||||
|
||||
## Performance Analysis
|
||||
|
||||
### Why is Core API Sometimes Faster?
|
||||
|
||||
The benchmark shows Core API is often comparable or even faster than direct Ollama. This seems counterintuitive, but here's why:
|
||||
|
||||
1. **Efficient FastAPI async handling** - Non-blocking I/O reduces overhead
|
||||
2. **Minimal middleware** - Only CORS and logging add <10ms
|
||||
3. **No heavy memory layer active** - Memory system exists but doesn't slow requests
|
||||
4. **HTTP connection pooling** - httpx AsyncClient reuses connections
|
||||
5. **Measurement variance** - Network/scheduling jitter affects sub-second measurements
|
||||
|
||||
### Where is the 42% Overhead in Medium Test?
|
||||
|
||||
The "medium" test showed +180ms overhead:
|
||||
- Ollama: 0.426s average
|
||||
- Core API: 0.606s average
|
||||
|
||||
**Root cause:** Likely serialization overhead for medium-length responses
|
||||
- Request parsing: JSON → Pydantic models
|
||||
- Response formatting: Ollama format → OpenAI format
|
||||
- SSE streaming setup (even for non-streaming requests)
|
||||
|
||||
**Impact:** Acceptable - only affects responses in 100-200 token range
|
||||
|
||||
### First Request Latency (13s)
|
||||
|
||||
The "short" test Run 1 showed Ollama taking 13.797s:
|
||||
- This is **model loading time** (cold start)
|
||||
- Ollama loads model into GPU memory on first request
|
||||
- Subsequent requests use cached model (0.2-0.3s)
|
||||
|
||||
**Not a Core API issue** - both endpoints experience this warmup delay
|
||||
|
||||
## Bottleneck Identification
|
||||
|
||||
Based on the benchmarks, here are the confirmed bottlenecks:
|
||||
|
||||
### ✓ NOT Bottlenecks (Performance is Good)
|
||||
|
||||
1. **Core API abstraction layer** - Adds <10% overhead
|
||||
2. **FastAPI framework** - Efficient async handling
|
||||
3. **JSON serialization** - Fast enough for this use case
|
||||
4. **Network hop** (client → Core API → Ollama) - Minimal latency
|
||||
|
||||
### ⚠️ Actual Bottlenecks (If You're Experiencing Slowness)
|
||||
|
||||
If you're experiencing poor performance, it's likely one of these:
|
||||
|
||||
1. **Client-side issues:**
|
||||
- Network latency to server
|
||||
- Client HTTP library blocking/synchronous calls
|
||||
- Browser tab throttling
|
||||
- Open WebUI buffering/rendering
|
||||
|
||||
2. **Model/GPU issues:**
|
||||
- Model not loaded (13s cold start)
|
||||
- GPU memory fragmentation
|
||||
- Other GPU processes competing (AMP, Jellyfin transcoding)
|
||||
|
||||
3. **System resources:**
|
||||
- 9GB RAM used (56%) - some swap pressure
|
||||
- CPU load from other services (AMP using 27% RAM)
|
||||
|
||||
## Recommendations
|
||||
|
||||
### For Current Setup (No Changes Needed)
|
||||
|
||||
✅ **Core API performance is GOOD** - Keep using it for:
|
||||
- OpenAI API compatibility
|
||||
- Open WebUI integration
|
||||
- Conversation memory features
|
||||
- Infrastructure automation
|
||||
|
||||
### If You Experience Slowness
|
||||
|
||||
1. **Check client-side:**
|
||||
```bash
|
||||
# Test direct from terminal
|
||||
time curl -X POST http://192.168.86.149:8083/v1/chat/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{"model": "gemma:2b", "messages": [{"role": "user", "content": "Hello"}]}'
|
||||
```
|
||||
|
||||
2. **Monitor GPU usage:**
|
||||
```bash
|
||||
watch -n 1 nvidia-smi
|
||||
# Check if GPU is loaded with other tasks
|
||||
```
|
||||
|
||||
3. **Check if model is loaded:**
|
||||
```bash
|
||||
curl http://192.168.86.149:11434/api/tags
|
||||
# First request after restart takes 13s to load model
|
||||
```
|
||||
|
||||
4. **Reduce concurrent GPU load:**
|
||||
- Don't use Jellyfin transcoding + AI chat simultaneously
|
||||
- AMP game servers may use GPU for some tasks
|
||||
|
||||
### Optional Optimizations (If Needed)
|
||||
|
||||
**For sub-second responses:**
|
||||
- Use `gemma:2b` instead of `gemma:7b` (3x faster, similar quality)
|
||||
- Pre-load model: `docker exec ollama ollama run gemma:2b "test"`
|
||||
|
||||
**For long conversations:**
|
||||
- Enable memory tier consolidation (already implemented)
|
||||
- Use streaming responses for better UX
|
||||
|
||||
**For API-heavy workloads:**
|
||||
- Increase Core API container CPU limit
|
||||
- Enable response caching for identical requests
|
||||
|
||||
## Conclusion
|
||||
|
||||
**The Core API is performing excellently.**
|
||||
|
||||
- Average overhead: <10%
|
||||
- Streaming first token: -167ms (faster!)
|
||||
- Resource usage: Minimal
|
||||
|
||||
If you're experiencing slow performance, it's likely:
|
||||
1. Client-side buffering/rendering (Open WebUI)
|
||||
2. Cold start model loading (first request)
|
||||
3. GPU contention with other services
|
||||
|
||||
The benchmark proves the abstraction layer is **not** the bottleneck.
|
||||
|
||||
## Test Scripts
|
||||
|
||||
Benchmark scripts are available at:
|
||||
- `/tmp/benchmark_ollama_vs_api.py` - Comprehensive non-streaming test
|
||||
- `/tmp/test_streaming_performance.py` - Streaming performance test
|
||||
- `/tmp/monitor_resources.sh` - System resource monitoring
|
||||
|
||||
To re-run:
|
||||
```bash
|
||||
python3 /tmp/benchmark_ollama_vs_api.py
|
||||
python3 /tmp/test_streaming_performance.py
|
||||
```
|
||||
@@ -1,347 +0,0 @@
|
||||
# VRAM Budget Analysis - Multi-Model Strategy
|
||||
|
||||
**Hardware**: RTX 2080 Ti (11GB VRAM)
|
||||
**Goal**: Keep orchestrator loaded + room for expert models
|
||||
|
||||
## Current Model Inventory
|
||||
|
||||
| Model | Size on Disk | VRAM When Loaded | Quantization |
|
||||
|-------|--------------|------------------|--------------|
|
||||
| **mistral:7b** | 4.4GB | ~5.1GB | Q4_K_M |
|
||||
| **mistral:7b Q4_K_S** | 4.1GB | ~4.7GB | Q4_K_S |
|
||||
| **mistral:7b Q3_K_M** | 3.5GB | ~4.0GB | Q3_K_M |
|
||||
| **codegemma:latest** | 5.0GB | ~5.8GB | Unknown |
|
||||
| **codestral:latest** | 12GB | ~13GB | Too large! |
|
||||
|
||||
## Key Finding: Q3 Removes Tool Support ❌
|
||||
|
||||
**Critical Issue**: The Q3_K_M quantization **removes tool calling capability**.
|
||||
|
||||
```
|
||||
mistral:7b Q4_K_M:
|
||||
Capabilities: completion, tools ✅
|
||||
|
||||
mistral:7b Q3_K_M:
|
||||
Capabilities: completion ❌ No tools!
|
||||
```
|
||||
|
||||
**This means**: You cannot use Q3 for the orchestrator. Tool calling requires Q4 or higher.
|
||||
|
||||
---
|
||||
|
||||
## Scenario Analysis
|
||||
|
||||
### Scenario 1: Current Setup (mistral:7b Q4_K_M)
|
||||
|
||||
```
|
||||
Total VRAM: 11.0 GB
|
||||
├─ mistral:7b Q4: 5.1 GB (46%) ← Orchestrator (always loaded)
|
||||
├─ Overhead: 1.2 GB (11%)
|
||||
└─ Available: 4.7 GB (43%) ← For expert models
|
||||
```
|
||||
|
||||
**What fits in 4.7GB free space:**
|
||||
- ✅ codegemma:latest (5.8GB) - **Does NOT fit** (need 5.8GB, have 4.7GB)
|
||||
- ❌ codestral:latest (13GB) - **Does NOT fit** (way too large)
|
||||
- ✅ gemma3:4b (4.5GB) - **Barely fits** (general purpose)
|
||||
- ✅ qwen2.5:3b (3.5GB) - **Fits comfortably** (if available)
|
||||
|
||||
**Reality Check**: You **cannot** load codegemma or codestral alongside mistral:7b Q4.
|
||||
|
||||
---
|
||||
|
||||
### Scenario 2: Slightly Smaller Q4 (mistral:7b-instruct-q4_K_S)
|
||||
|
||||
```
|
||||
Total VRAM: 11.0 GB
|
||||
├─ mistral:7b Q4_K_S: 4.7 GB (43%) ← Orchestrator (slightly smaller)
|
||||
├─ Overhead: 1.2 GB (11%)
|
||||
└─ Available: 5.1 GB (46%) ← For expert models
|
||||
```
|
||||
|
||||
**Savings**: 400MB (5.1GB → 4.7GB)
|
||||
|
||||
**What fits now:**
|
||||
- ⚠️ codegemma:latest (5.8GB) - **Still doesn't fit** (need 5.8GB, have 5.1GB)
|
||||
- ❌ codestral:latest (13GB) - **No chance**
|
||||
- ✅ gemma3:4b (4.5GB) - **Fits with room to spare**
|
||||
|
||||
**Benefit**: Not enough to matter. Still can't fit codegemma.
|
||||
|
||||
---
|
||||
|
||||
### Scenario 3: Dynamic Loading (Current Ollama Behavior)
|
||||
|
||||
**This is what Ollama already does by default!**
|
||||
|
||||
```
|
||||
Step 1: Only orchestrator loaded
|
||||
├─ mistral:7b Q4: 5.1 GB
|
||||
├─ Overhead: 1.2 GB
|
||||
└─ Available: 4.7 GB
|
||||
|
||||
Step 2: User requests code generation
|
||||
├─ Unload mistral:7b (-5.1GB)
|
||||
├─ Load codestral (+13GB) ← Swaps automatically
|
||||
└─ Available: 0 GB (codestral fills VRAM)
|
||||
|
||||
Step 3: Codestral finishes, times out
|
||||
├─ Unload codestral (-13GB)
|
||||
├─ Load mistral:7b (+5.1GB) ← Swaps back
|
||||
└─ Back to Step 1
|
||||
```
|
||||
|
||||
**How it works:**
|
||||
- Ollama has a `keep_alive` timer (default: 5 minutes)
|
||||
- When a model isn't used for 5min, it's unloaded from VRAM
|
||||
- When you request a different model, Ollama swaps them automatically
|
||||
|
||||
**Cold start times:**
|
||||
- Loading mistral:7b: ~2-3 seconds
|
||||
- Loading codestral:22b: ~8-10 seconds
|
||||
- Loading codegemma:9b: ~3-4 seconds
|
||||
|
||||
---
|
||||
|
||||
## The Math: Why Expert Models Don't Fit
|
||||
|
||||
Your 11GB VRAM budget breaks down like this:
|
||||
|
||||
```
|
||||
11GB total VRAM
|
||||
- 5.1GB orchestrator (mistral:7b Q4)
|
||||
- 1.2GB system overhead
|
||||
━━━━━━━━━━━━━━━━━━━━━━
|
||||
= 4.7GB available
|
||||
|
||||
But your expert models need:
|
||||
- codestral:22b = 13GB ❌ (needs 8GB more than you have)
|
||||
- codegemma:9b = 5.8GB ❌ (needs 1GB more than available)
|
||||
```
|
||||
|
||||
**Even if you use the smallest possible orchestrator:**
|
||||
```
|
||||
11GB total VRAM
|
||||
- 3.8GB orchestrator (gemma3-tools:1b, unreliable!)
|
||||
- 1.2GB system overhead
|
||||
━━━━━━━━━━━━━━━━━━━━━━
|
||||
= 6.0GB available
|
||||
|
||||
Still not enough for:
|
||||
- codestral:22b = 13GB ❌ (needs 7GB more)
|
||||
- codegemma:9b = 5.8GB ✅ (fits, but orchestrator is unreliable)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Reality: You Need Dynamic Loading
|
||||
|
||||
**Conclusion**: With 11GB VRAM, you **cannot** keep both:
|
||||
1. A reliable orchestrator (min 4.7GB for mistral Q4_K_S)
|
||||
2. Large expert models (5.8GB+ for code models)
|
||||
|
||||
**loaded simultaneously**.
|
||||
|
||||
### Option A: Accept Dynamic Loading (Recommended)
|
||||
|
||||
**Keep orchestrator loaded** with `keep_alive=-1`, but expert models swap in/out:
|
||||
|
||||
```python
|
||||
# In Core API orchestrator.py
|
||||
self.llm = ChatOllama(
|
||||
model="mistral:7b", # Use Q4_K_M or Q4_K_S
|
||||
keep_alive=-1, # Never unload orchestrator
|
||||
)
|
||||
|
||||
# When calling expert models:
|
||||
codestral_llm = ChatOllama(
|
||||
model="codestral:latest",
|
||||
keep_alive="5m", # Auto-unload after 5 min idle
|
||||
)
|
||||
```
|
||||
|
||||
**How it works in practice:**
|
||||
|
||||
1. **Orchestrator queries** (~80% of requests):
|
||||
- mistral:7b always in VRAM
|
||||
- Instant response (~0ms cold start)
|
||||
- Uses 5.1GB VRAM
|
||||
|
||||
2. **Code generation** (~20% of requests):
|
||||
- mistral:7b stays loaded initially
|
||||
- Ollama sees codestral request
|
||||
- **Unloads mistral** automatically
|
||||
- **Loads codestral** (8-10s cold start)
|
||||
- Codestral generates code
|
||||
- After 5min idle: **unloads codestral, reloads mistral**
|
||||
|
||||
**Trade-offs:**
|
||||
- ✅ Orchestrator instant most of the time
|
||||
- ⚠️ 8-10s cold start when switching to codestral (first code request)
|
||||
- ⚠️ 2-3s cold start when switching back to orchestrator (after codestral timeout)
|
||||
- ✅ Can use full-size expert models (codestral:22b, etc.)
|
||||
|
||||
---
|
||||
|
||||
### Option B: Use Smaller Expert Models
|
||||
|
||||
If cold starts are unacceptable, use smaller expert models that fit alongside orchestrator:
|
||||
|
||||
```
|
||||
Orchestrator: mistral:7b Q4_K_S (4.7GB)
|
||||
Expert: qwen2.5-coder:3b (3.5GB) ← Smaller code model
|
||||
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
|
||||
Total: 8.2GB + 1.2GB overhead = 9.4GB
|
||||
Available: 1.6GB buffer
|
||||
```
|
||||
|
||||
**Smaller code model options:**
|
||||
- `qwen2.5-coder:3b` (3.5GB) - Good for simple code tasks
|
||||
- `starcoder2:3b` (3.2GB) - Focused on code completion
|
||||
- `deepseek-coder:1.3b` (1.5GB) - Very small, lower quality
|
||||
|
||||
**Trade-offs:**
|
||||
- ✅ Both models always loaded (no cold starts)
|
||||
- ✅ Instant switching
|
||||
- ❌ Smaller models = lower code quality
|
||||
- ❌ Can't use top-tier models like codestral
|
||||
|
||||
---
|
||||
|
||||
### Option C: Upgrade GPU (Future)
|
||||
|
||||
If you want both instant orchestrator AND large expert models:
|
||||
|
||||
**RTX 4070 Ti (16GB VRAM):**
|
||||
```
|
||||
Total VRAM: 16GB
|
||||
├─ mistral:7b Q4: 5.1GB (32%)
|
||||
├─ codestral:22b: 8.0GB (50%) ← Quantized version
|
||||
├─ Overhead: 1.5GB (9%)
|
||||
└─ Available: 1.4GB (9%)
|
||||
```
|
||||
|
||||
**With 16GB, you can fit:**
|
||||
- Orchestrator + codestral Q4 (13GB total)
|
||||
- Orchestrator + codegemma (11GB total)
|
||||
- Orchestrator + multiple small experts
|
||||
|
||||
---
|
||||
|
||||
## Video/Image Models: The Situation
|
||||
|
||||
Video and image models are **MUCH larger** than text models:
|
||||
|
||||
### Image Generation Models:
|
||||
- **SDXL (Stable Diffusion XL)**: 6-7GB VRAM
|
||||
- **Flux.1**: 16-24GB VRAM (dev/schnell variants)
|
||||
- **SD 1.5**: 3-4GB VRAM (older, lower quality)
|
||||
|
||||
### Video Models:
|
||||
- **AnimateDiff**: 8-12GB VRAM
|
||||
- **Stable Video Diffusion**: 10-14GB VRAM
|
||||
- **CogVideoX**: 16-48GB VRAM
|
||||
|
||||
### Vision Models (Image Understanding):
|
||||
- **LLaVA 7B**: 6-7GB VRAM
|
||||
- **LLaVA 13B**: 10-12GB VRAM
|
||||
- **GPT-4V equivalent**: 12-16GB VRAM
|
||||
|
||||
**Reality Check for 11GB VRAM:**
|
||||
|
||||
```
|
||||
Scenario: Orchestrator + Vision Model
|
||||
├─ mistral:7b Q4: 5.1GB
|
||||
├─ LLaVA 7B: 6.5GB
|
||||
━━━━━━━━━━━━━━━━━━━━━━━
|
||||
Total needed: 11.6GB ❌ Doesn't fit!
|
||||
```
|
||||
|
||||
Even the smallest vision model (LLaVA 7B) won't fit alongside your orchestrator.
|
||||
|
||||
**For image/video generation**: You'd need to fully unload the orchestrator to make room.
|
||||
|
||||
---
|
||||
|
||||
## Recommendation: Hybrid Strategy
|
||||
|
||||
**For your 11GB VRAM constraint, I recommend:**
|
||||
|
||||
### 1. Keep Orchestrator Always Loaded
|
||||
```bash
|
||||
# mistral:7b Q4_K_M (5.1GB) with keep_alive=-1
|
||||
# Current setup, no changes needed
|
||||
```
|
||||
|
||||
### 2. Accept Dynamic Loading for Experts
|
||||
- **Code models**: Load on-demand (codestral, codegemma)
|
||||
- **Vision models**: Load on-demand (LLaVA)
|
||||
- **Image gen**: Load on-demand (SDXL)
|
||||
|
||||
### 3. Optimize with `keep_alive` Tuning
|
||||
|
||||
```python
|
||||
# Orchestrator: Never unload
|
||||
orchestrator = ChatOllama(model="mistral:7b", keep_alive=-1)
|
||||
|
||||
# Frequently used expert: Keep for 30min
|
||||
code_expert = ChatOllama(model="codegemma:9b", keep_alive="30m")
|
||||
|
||||
# Rarely used expert: Keep for 5min only
|
||||
vision_expert = ChatOllama(model="llava:7b", keep_alive="5m")
|
||||
```
|
||||
|
||||
**Result:**
|
||||
- Orchestrator: Always instant
|
||||
- Frequent code requests: 1st request has 3-4s cold start, then instant for 30min
|
||||
- Rare vision requests: 6-8s cold start each time
|
||||
|
||||
### 4. Monitor and Adjust
|
||||
|
||||
Track which expert models you use most:
|
||||
- If you do a LOT of coding → Keep codegemma loaded longer (`keep_alive="1h"`)
|
||||
- If coding is rare → Accept the cold start (`keep_alive="5m"`)
|
||||
|
||||
---
|
||||
|
||||
## Future-Proofing
|
||||
|
||||
**If you want to add image/video in the future:**
|
||||
|
||||
### Option 1: Offload to CPU (Slow)
|
||||
```bash
|
||||
# Run image generation on CPU (very slow, 5-10min per image)
|
||||
OLLAMA_NUM_GPU=0 ollama run stable-diffusion
|
||||
```
|
||||
|
||||
### Option 2: Dedicated GPU
|
||||
- Keep RTX 2080 Ti for text models (orchestrator + code)
|
||||
- Add second GPU for image/video (RTX 3060 12GB, ~$250 used)
|
||||
|
||||
### Option 3: Cloud Hybrid
|
||||
- Local: Text models (orchestrator, code, chat)
|
||||
- Cloud: Image/video generation (Replicate API, RunPod, etc.)
|
||||
- Cost: ~$0.002-0.01 per image
|
||||
|
||||
---
|
||||
|
||||
## Bottom Line
|
||||
|
||||
**Your VRAM situation:**
|
||||
|
||||
| Capability | Status | Notes |
|
||||
|------------|--------|-------|
|
||||
| **Keep orchestrator loaded** | ✅ Yes | 5.1GB with mistral:7b Q4 |
|
||||
| **+ codegemma simultaneously** | ❌ No | Need 5.8GB, have 4.7GB free |
|
||||
| **+ codestral simultaneously** | ❌ No | Need 13GB, have 4.7GB free |
|
||||
| **+ vision model simultaneously** | ❌ No | Need 6GB+, have 4.7GB free |
|
||||
| **Dynamic loading (swap models)** | ✅ Yes | 2-10s cold starts |
|
||||
| **Smaller experts simultaneously** | ✅ Maybe | With 3-4GB models only |
|
||||
|
||||
**Verdict**:
|
||||
- ✅ You CAN keep orchestrator always loaded
|
||||
- ⚠️ You CANNOT keep large experts loaded simultaneously
|
||||
- ✅ Dynamic loading works fine with acceptable cold start times
|
||||
- ❌ Image/video models won't fit even with dynamic loading (need GPU upgrade)
|
||||
|
||||
**Best approach**: Keep current setup (mistral:7b Q4 always loaded), accept dynamic swapping for expert models. It's what Ollama is designed to do, and 3-8s cold starts are acceptable for occasional expert model use.
|
||||
@@ -1,391 +0,0 @@
|
||||
# VRAM Optimization Strategy for Model Orchestration
|
||||
|
||||
**Date**: 2025-11-24
|
||||
**Context**: Multi-model architecture with always-loaded orchestrator + expert models
|
||||
**Hardware**: RTX 2080 Ti (11GB VRAM)
|
||||
|
||||
## Problem Statement
|
||||
|
||||
**Goal**: Keep orchestrator model always loaded in VRAM to prevent cold starts, while maximizing VRAM availability for expert models.
|
||||
|
||||
**Current State**:
|
||||
- Orchestrator: `mistral:7b` (5.1GB VRAM)
|
||||
- Free VRAM: 4.7GB
|
||||
- Use case: Orchestrator decides → routes to expert models (codestral, etc.)
|
||||
|
||||
**Challenge**: mistral:7b consumes 45% of available VRAM, limiting expert model options.
|
||||
|
||||
## VRAM Budget Analysis
|
||||
|
||||
### Current Configuration
|
||||
```
|
||||
Total VRAM: 11.0 GB
|
||||
├─ mistral:7b: 5.1 GB (46.4%) - Orchestrator
|
||||
├─ Overhead: 1.2 GB (10.6%) - System/Ollama
|
||||
└─ Available: 4.7 GB (42.7%) - For expert models
|
||||
```
|
||||
|
||||
### Desired Configuration
|
||||
```
|
||||
Total VRAM: 11.0 GB
|
||||
├─ Orchestrator: ??? GB (minimize)
|
||||
├─ Expert Model: ??? GB (maximize)
|
||||
└─ Overhead: 1.2 GB
|
||||
```
|
||||
|
||||
## Solution Options
|
||||
|
||||
### Option 1: Accept gemma3-tools:1b Limitations ⚠️
|
||||
|
||||
**VRAM Savings**: 3.8GB (5.1GB → 1.3GB)
|
||||
|
||||
```
|
||||
Orchestrator: gemma3-tools:1b (1.3GB)
|
||||
Free for experts: 8.5GB
|
||||
```
|
||||
|
||||
**Pros**:
|
||||
- ✅ Massive VRAM savings (74% reduction)
|
||||
- ✅ Leaves 8.5GB for expert models
|
||||
- ✅ Can load codestral:22b (full size) + orchestrator simultaneously
|
||||
|
||||
**Cons**:
|
||||
- ❌ 33% tool calling reliability
|
||||
- ❌ Wrong tool selection
|
||||
- ❌ Erratic responses (raw JSON output)
|
||||
- ❌ Poor user experience
|
||||
|
||||
**Verdict**: ❌ **Not recommended** - Unreliability hurts more than VRAM savings help
|
||||
|
||||
---
|
||||
|
||||
### Option 2: Use Smaller Quantization of mistral:7b ✅ RECOMMENDED
|
||||
|
||||
Ollama supports multiple quantization levels. You're currently using Q4_K_M, but Q2 or Q3 exist.
|
||||
|
||||
**Available Quantizations**:
|
||||
- Q2_K: ~2.5GB VRAM (70% quality retention, aggressive)
|
||||
- Q3_K_M: ~3.2GB VRAM (80% quality, good balance)
|
||||
- Q4_K_M: ~5.1GB VRAM (90% quality, current)
|
||||
- Q5_K_M: ~6.2GB VRAM (95% quality)
|
||||
- Q8: ~7.7GB VRAM (99% quality, near full precision)
|
||||
|
||||
**Recommended**: Pull `mistral:7b-instruct-q3_K_M`
|
||||
|
||||
```bash
|
||||
# Pull lower quantization
|
||||
ollama pull mistral:7b-instruct-q3_K_M
|
||||
|
||||
# Update Core API config
|
||||
# services/core-api/.env
|
||||
AGENT_MODEL=mistral:7b-instruct-q3_K_M
|
||||
```
|
||||
|
||||
**New VRAM Budget**:
|
||||
```
|
||||
Orchestrator: mistral:7b Q3_K_M (3.2GB)
|
||||
Free for experts: 6.6GB
|
||||
Savings: 1.9GB (37% reduction)
|
||||
```
|
||||
|
||||
**Pros**:
|
||||
- ✅ 100% tool calling compatibility (same model architecture)
|
||||
- ✅ 1.9GB VRAM savings
|
||||
- ✅ Minimal quality loss (80% of full precision is fine for routing)
|
||||
- ✅ Proven reliability maintained
|
||||
|
||||
**Cons**:
|
||||
- ⚠️ Slightly lower response quality (acceptable for orchestration)
|
||||
- ⚠️ May need testing to verify tool calling still works
|
||||
|
||||
**Verdict**: ✅ **Best option** - Balanced approach
|
||||
|
||||
---
|
||||
|
||||
### Option 3: Hybrid Orchestrator (Simple Router + mistral:7b) 🔮 ADVANCED
|
||||
|
||||
Use a **two-tier routing system**:
|
||||
1. **Lightweight classifier** (gemma3-tools:1b) - Always loaded
|
||||
2. **Full orchestrator** (mistral:7b) - Loaded on demand for complex queries
|
||||
|
||||
**Architecture**:
|
||||
```python
|
||||
# Tier 1: Fast classifier (always loaded)
|
||||
if query_is_simple(message):
|
||||
# Direct routing: "list services" → list_services tool
|
||||
# Load time: 0ms (always in VRAM)
|
||||
use_simple_router(gemma3-tools:1b)
|
||||
else:
|
||||
# Complex routing: multi-tool, reasoning needed
|
||||
# Load time: ~2s (load mistral:7b)
|
||||
use_full_orchestrator(mistral:7b)
|
||||
```
|
||||
|
||||
**VRAM Budget**:
|
||||
```
|
||||
Tier 1 (always): gemma3-tools:1b (1.3GB)
|
||||
Tier 2 (on-demand): mistral:7b (5.1GB, loaded when needed)
|
||||
Free when Tier 1 only: 8.5GB
|
||||
Free when both loaded: 3.4GB
|
||||
```
|
||||
|
||||
**Pros**:
|
||||
- ✅ 8.5GB free for expert models most of the time
|
||||
- ✅ Only loads mistral:7b when truly needed
|
||||
- ✅ Simple queries stay fast (no model swap)
|
||||
|
||||
**Cons**:
|
||||
- ❌ Complex implementation (need query classifier)
|
||||
- ❌ 2s latency spike when switching to Tier 2
|
||||
- ❌ More failure modes (what if Tier 1 misclassifies?)
|
||||
|
||||
**Verdict**: 🔮 **Future enhancement** - Interesting but complex
|
||||
|
||||
---
|
||||
|
||||
### Option 4: Use Different Base Model 🔍 RESEARCH NEEDED
|
||||
|
||||
Look for other tool-capable models with better size/quality trade-offs.
|
||||
|
||||
**Candidates to research**:
|
||||
- `qwen2.5:7b-instruct-q3` - Alibaba's model, claimed good tool support
|
||||
- `llama3.2:3b-instruct` - Meta's latest, check if tool-capable
|
||||
- `hermes3:3b` - Nous Research, specifically trained for function calling
|
||||
|
||||
**Action**: Test these if available in Ollama registry.
|
||||
|
||||
---
|
||||
|
||||
## Recommended Implementation: Option 2
|
||||
|
||||
### Step 1: Pull Q3 Quantization
|
||||
|
||||
```bash
|
||||
# Check if Q3 variant exists
|
||||
ollama list | grep mistral
|
||||
|
||||
# Pull Q3 quantization (if available)
|
||||
ollama pull mistral:7b-instruct-q3_K_M
|
||||
|
||||
# OR manually create Q3 from modelfile
|
||||
cat > /tmp/mistral-q3.Modelfile << 'EOF'
|
||||
FROM mistral:7b
|
||||
PARAMETER quantization Q3_K_M
|
||||
EOF
|
||||
|
||||
ollama create mistral:7b-q3 -f /tmp/mistral-q3.Modelfile
|
||||
```
|
||||
|
||||
### Step 2: Test Tool Calling with Q3
|
||||
|
||||
```bash
|
||||
# Run our test script with Q3 variant
|
||||
source .venv/bin/activate
|
||||
python3 << 'PYEOF'
|
||||
import asyncio
|
||||
import httpx
|
||||
|
||||
async def test():
|
||||
payload = {
|
||||
"model": "mistral:7b-q3",
|
||||
"messages": [
|
||||
{"role": "system", "content": "You are a helpful assistant."},
|
||||
{"role": "user", "content": "List all services"}
|
||||
],
|
||||
"tools": [{
|
||||
"type": "function",
|
||||
"function": {
|
||||
"name": "list_services",
|
||||
"description": "List all running services",
|
||||
"parameters": {"type": "object", "properties": {}}
|
||||
}
|
||||
}]
|
||||
}
|
||||
|
||||
async with httpx.AsyncClient(timeout=60) as client:
|
||||
r = await client.post("http://localhost:11434/api/chat", json=payload)
|
||||
data = r.json()
|
||||
message = data.get("message", {})
|
||||
|
||||
if "tool_calls" in message:
|
||||
print("✅ Q3 quantization: Tool calling WORKS")
|
||||
print(f" Called: {message['tool_calls'][0]['function']['name']}")
|
||||
else:
|
||||
print("❌ Q3 quantization: Tool calling BROKEN")
|
||||
print(f" Response: {message.get('content', '')[:100]}")
|
||||
|
||||
asyncio.run(test())
|
||||
PYEOF
|
||||
```
|
||||
|
||||
### Step 3: Update Core API Configuration
|
||||
|
||||
```bash
|
||||
# services/core-api/.env
|
||||
AGENT_MODEL=mistral:7b-q3
|
||||
```
|
||||
|
||||
```bash
|
||||
# Restart core-api to pick up new model
|
||||
docker restart core-api
|
||||
```
|
||||
|
||||
### Step 4: Verify VRAM Usage
|
||||
|
||||
```bash
|
||||
# Check new VRAM allocation
|
||||
curl -s http://localhost:11434/api/ps | jq '.models[] | {name, size_vram_gb: (.size_vram / 1024 / 1024 / 1024)}'
|
||||
```
|
||||
|
||||
**Expected Result**:
|
||||
```json
|
||||
{
|
||||
"name": "mistral:7b-q3",
|
||||
"size_vram_gb": 3.2
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Expert Model Strategy
|
||||
|
||||
With ~6.6GB available after Q3 orchestrator, you can now fit:
|
||||
|
||||
### Option A: Single Large Expert
|
||||
```
|
||||
Orchestrator: mistral:7b-q3 (3.2GB)
|
||||
Expert: codestral:22b-q2 (6GB)
|
||||
Total: 9.2GB / 11GB
|
||||
```
|
||||
|
||||
### Option B: Multiple Smaller Experts
|
||||
```
|
||||
Orchestrator: mistral:7b-q3 (3.2GB)
|
||||
Expert 1: codegemma:7b (4GB) - Code generation
|
||||
Expert 2: gemma3:4b (3GB) - General knowledge
|
||||
Total: 10.2GB / 11GB (near full capacity)
|
||||
```
|
||||
|
||||
### Option C: Dynamic Loading (Current Behavior)
|
||||
```
|
||||
Orchestrator: mistral:7b-q3 (3.2GB) - Always loaded
|
||||
Expert: Load on demand (6.6GB available)
|
||||
- codestral for code
|
||||
- gemma3:12b for general
|
||||
- Model swaps as needed
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Advanced: Ollama Keep Alive Configuration
|
||||
|
||||
Control how long models stay in VRAM:
|
||||
|
||||
```bash
|
||||
# Keep orchestrator always loaded (never unload)
|
||||
curl -X POST http://localhost:11434/api/generate \
|
||||
-d '{
|
||||
"model": "mistral:7b-q3",
|
||||
"keep_alive": -1,
|
||||
"prompt": "warm up"
|
||||
}'
|
||||
|
||||
# Expert models: unload after 5 minutes idle
|
||||
curl -X POST http://localhost:11434/api/generate \
|
||||
-d '{
|
||||
"model": "codestral:latest",
|
||||
"keep_alive": "5m",
|
||||
"prompt": "warm up"
|
||||
}'
|
||||
```
|
||||
|
||||
**Configuration in Core API**:
|
||||
```python
|
||||
# services/core-api/src/agent/orchestrator.py
|
||||
|
||||
self.llm = ChatOllama(
|
||||
model=self.settings.agent_model, # mistral:7b-q3
|
||||
base_url=self.settings.ollama_base_url,
|
||||
temperature=0.7,
|
||||
keep_alive=-1, # Never unload orchestrator
|
||||
)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Testing Checklist
|
||||
|
||||
Before switching to Q3 quantization:
|
||||
|
||||
- [ ] Pull or create Q3 variant
|
||||
- [ ] Test tool calling functionality
|
||||
- [ ] Test tool selection accuracy (list_services vs get_service_details)
|
||||
- [ ] Test multi-tool workflows
|
||||
- [ ] Compare response quality vs Q4
|
||||
- [ ] Verify VRAM usage reduction
|
||||
- [ ] Test with Open WebUI
|
||||
- [ ] Monitor for any degradation
|
||||
|
||||
If Q3 shows issues:
|
||||
- Try Q4_K_S (slightly smaller than Q4_K_M)
|
||||
- Fall back to current Q4_K_M if necessary
|
||||
|
||||
---
|
||||
|
||||
## Alternative Models Research
|
||||
|
||||
If mistral Q3 proves insufficient, test these:
|
||||
|
||||
### qwen2.5:7b (Alibaba Cloud)
|
||||
- Similar size to mistral
|
||||
- Claimed excellent tool calling
|
||||
- May have Q3/Q4 variants available
|
||||
|
||||
```bash
|
||||
ollama pull qwen2.5:7b-instruct
|
||||
# Test with our tool calling script
|
||||
```
|
||||
|
||||
### hermes3:3b (Nous Research)
|
||||
- Specifically trained for function calling
|
||||
- 3B parameters (smaller than mistral)
|
||||
- Check Ollama availability
|
||||
|
||||
```bash
|
||||
ollama search hermes3
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Summary
|
||||
|
||||
**Immediate Action**: Pull `mistral:7b` with Q3_K_M quantization
|
||||
|
||||
```bash
|
||||
# Check available quantizations
|
||||
ollama show mistral:7b --modelfile
|
||||
|
||||
# Pull Q3 if available, or create from Q4
|
||||
ollama pull mistral:7b-instruct-q3_K_M
|
||||
```
|
||||
|
||||
**Expected Outcome**:
|
||||
- VRAM savings: 1.9GB (5.1GB → 3.2GB)
|
||||
- Tool calling: Should work (same architecture)
|
||||
- Quality: 80% of Q4 (acceptable for routing logic)
|
||||
- Expert model budget: 6.6GB (up from 4.7GB)
|
||||
|
||||
**Risk Mitigation**:
|
||||
- Test thoroughly before production
|
||||
- Keep Q4 variant as backup
|
||||
- Monitor for quality degradation
|
||||
|
||||
**Long-term**:
|
||||
- Research newer models (qwen2.5, hermes3)
|
||||
- Consider hybrid routing if complexity justified
|
||||
- Revisit when Ollama adds model multiplexing features
|
||||
|
||||
---
|
||||
|
||||
**Status**: Research complete, awaiting quantization testing
|
||||
**Next Steps**: User decision on Q3 testing approach
|
||||
Reference in New Issue
Block a user