291 lines
15 KiB
Markdown
291 lines
15 KiB
Markdown
# tower-of-joy Project Status
|
|
|
|
> Last Updated: 2025-11-14
|
|
> Version: 0.7.1-gitea-deployment
|
|
|
|
## Current Phase
|
|
|
|
**Phase:** AI Orchestrator Development - Phase 1 (Foundation)
|
|
**Status:** ✅ **COMPLETED**
|
|
**Started:** 2025-11-13
|
|
**Completed:** 2025-11-13
|
|
|
|
**Next Phase:** AI Orchestrator - Phase 2 (Memory Systems)
|
|
**Status:** 🔄 **IN PROGRESS**
|
|
|
|
## Implementation Progress
|
|
|
|
### ✅ Completed
|
|
|
|
**Planning & Documentation:**
|
|
- [x] Research home server solutions (see containers/research.md)
|
|
- [x] Architecture decision: Portainer + Docker Compose
|
|
- [x] Implementation plan created (see containers/implementation-plan.md)
|
|
- [x] AGENTS.md populated with project-specific guidelines
|
|
- [x] Storage strategy defined (SSD for configs, HDD for content)
|
|
|
|
**Phase 1: Foundation Setup:**
|
|
- [x] Install NVIDIA Container Toolkit (v1.17.9-1 - downgraded for driver 470 compatibility)
|
|
- [x] Deploy Portainer (port 8001, host networking)
|
|
- [x] Configure GPU management (via docker-compose deploy configuration)
|
|
- [x] Mount 4TB media drive at /mnt/media
|
|
- [x] Configure AMP integration (kept on port 8080, no conflicts)
|
|
- [x] Add user to docker group
|
|
- [x] Deploy Nginx Proxy Manager (port 81, host networking)
|
|
- [x] Deploy Ollama ML infrastructure (port 11434, GPU-enabled)
|
|
- [x] Fix Docker networking issues (iptables FORWARD chain, host networking solution)
|
|
|
|
**Phase 2: Networking & External Access:**
|
|
- [x] Deploy Headscale (port 8085, mesh VPN control server)
|
|
- [x] Configure Headscale with custom 10.99.0.0/16 network range
|
|
- [x] Create homelab user and generate pre-auth keys
|
|
- [x] Document connection procedures for all device types
|
|
|
|
**Phase 3: Monitoring & Management:**
|
|
- [x] Deploy Uptime Kuma (port 3001, service uptime monitoring)
|
|
- [x] Deploy Netdata (port 19999, real-time system metrics)
|
|
- [x] Deploy Heimdall dashboard (port 8888, unified dashboard)
|
|
|
|
**Phase 4: Optimization & Security:**
|
|
- [x] Deploy Watchtower (automatic container updates, daily at 4 AM)
|
|
- [x] Configure Docker log rotation (10MB max, 3 files per container)
|
|
- [x] Configure UFW firewall (SSH, Tailscale, infrastructure services allowed)
|
|
- [x] Deploy maintenance container (scheduled backups & future maintenance tasks)
|
|
- [x] Configure automated Docker config backups (daily at 3 AM, 30-day retention, ~94MB/backup)
|
|
|
|
**Application Deployment:**
|
|
- [x] Deploy Open WebUI (LLM chat interface, port 82)
|
|
- [x] Integrate Open WebUI with Uptime Kuma (health monitoring)
|
|
- [x] Integrate Open WebUI with Organizr (dashboard tab + homepage)
|
|
- [x] Document Open WebUI in CONTAINERS.md
|
|
- [x] Deploy Core API (OpenAPI functions for Open WebUI, port 8083)
|
|
- [x] Integrate Core API with Uptime Kuma (health monitoring)
|
|
- [x] Integrate Core API with Organizr (dashboard tab)
|
|
- [x] Document Core API in CONTAINERS.md
|
|
- [x] Upgrade system to Python 3.12 (from EOL 3.8)
|
|
|
|
**AI Orchestrator Development (Phase 1):**
|
|
- [x] Implement OpenAI-compatible `/v1/chat/completions` endpoint
|
|
- [x] Implement OpenAI-compatible `/v1/models` endpoint
|
|
- [x] Add streaming support (Server-Sent Events format)
|
|
- [x] Add non-streaming response mode
|
|
- [x] Implement model aliasing system (gpt-3.5-turbo → gemma:7b, etc.)
|
|
- [x] Create Ollama client with connection pooling
|
|
- [x] Add proper request/response schemas (Pydantic models)
|
|
- [x] Deploy to ai-dataplane network with hot-reload
|
|
- [x] Test streaming and non-streaming responses
|
|
- [x] Update CONTAINERS.md documentation
|
|
- [x] Deploy Nextcloud (cloud storage and collaboration platform, port 8082)
|
|
- [x] Configure Nextcloud with MariaDB, Redis, and reverse proxy (https://cloud.schweitz.net)
|
|
- [x] Optimize Nextcloud (database indices, bigint conversion, cron background jobs)
|
|
- [x] Relocate Nextcloud cron to maintenance container
|
|
- [x] Integrate Nextcloud with Uptime Kuma (HTTP monitoring)
|
|
- [x] Integrate Nextcloud with Organizr (dashboard tab)
|
|
- [x] Document Nextcloud in CONTAINERS.md
|
|
- [x] Deploy Samba (network file sharing, ports 139/445)
|
|
- [x] Configure Samba shares (Media R/W, Downloads R/W, Backups R/O)
|
|
- [x] Disable host Samba service to prevent conflicts
|
|
- [x] Configure UFW firewall for Samba ports
|
|
- [x] Integrate Samba with Uptime Kuma (TCP port monitoring)
|
|
- [x] Document Samba in CONTAINERS.md
|
|
- [x] Fix Uptime Kuma network connectivity (multi-network bridge to all service networks)
|
|
- [x] Deploy Gitea (Git repository hosting, ports 3002/2222)
|
|
- [x] Configure Gitea with PostgreSQL database backend
|
|
- [x] Configure NPM reverse proxy for https://git.schweitz.net with Let's Encrypt SSL
|
|
- [x] Complete Gitea initial setup wizard and create admin account
|
|
- [x] Integrate Gitea with Uptime Kuma (HTTP monitoring)
|
|
- [x] Integrate Gitea with Organizr (dashboard tab)
|
|
- [x] Document Gitea in CONTAINERS.md
|
|
|
|
### 🔄 In Progress
|
|
|
|
**Priority 1: Core-API Refactoring & Infrastructure Management:**
|
|
- [ ] **Code Cleanup:** Restructure Core API into function-specific controller files
|
|
- [x] Create `/controllers` directory structure
|
|
- [x] Create `/clients` directory structure
|
|
- [x] Create `base.py` controller base class
|
|
- [x] Add infrastructure settings to `config.py` (Portainer, NPM, Kuma URLs/credentials)
|
|
- [ ] Separate AI Orchestrator logic into `ai_controller.py`
|
|
- [ ] Extract webscraper to `tools_controller.py`
|
|
- [ ] Create `health_controller.py` for monitoring endpoints
|
|
- [ ] Update imports and routing in `main.py`
|
|
- [x] **Infrastructure Management Controller:** Build automation API for service management
|
|
- [x] Create `infrastructure_controller.py` with read/list endpoints
|
|
- [x] **Portainer Integration:** HTTP client with access token authentication
|
|
- [x] `get_endpoints()` - List Docker environments
|
|
- [x] `get_stacks()` - List all stacks
|
|
- [x] `get_stack(id)` - Get stack details
|
|
- [x] `create_stack()` - Deploy from compose YAML
|
|
- [x] `update_stack()` - Update existing stack
|
|
- [x] `delete_stack()` - Remove stack
|
|
- [x] **NPM Integration:** HTTP client with JWT bearer token + auto-refresh
|
|
- [x] Token refresh mechanism (24h expiration handling)
|
|
- [x] `get_proxy_hosts()` - List all proxy hosts
|
|
- [x] `get_proxy_host(id)` - Get proxy details
|
|
- [x] `create_proxy_host()` - Create new proxy configuration
|
|
- [x] `get_certificates()` - List SSL certificates
|
|
- [x] `create_certificate()` - Request Let's Encrypt cert
|
|
- [x] **Read/List Endpoints Implemented:**
|
|
- [x] `GET /infrastructure/health` - Check Portainer/NPM connectivity
|
|
- [x] `GET /infrastructure/services` - List all deployed services
|
|
- [x] `GET /infrastructure/services/{name}` - Get service details
|
|
- [x] `GET /infrastructure/ports` - List allocated ports (skeleton)
|
|
- [x] `GET /infrastructure/domains` - List configured domains
|
|
- [ ] **Uptime Kuma Integration:** WebSocket client (deferred - complex Socket.IO)
|
|
- [ ] **Write Endpoints:** Deploy/update/delete operations
|
|
- [ ] Replace ad-hoc shell scripts in `/stacks` with API endpoints
|
|
- [ ] Add CLI wrapper for common operations
|
|
- [ ] Test endpoints with live infrastructure
|
|
|
|
**Priority 2: AI Orchestrator Enhancement (Phase 2 - Memory Systems):**
|
|
- [ ] Implement Tier 1: ConversationBufferMemory (in-memory, last 10 turns)
|
|
- [ ] Implement Tier 2: ConversationSummaryMemory (SQLite summaries)
|
|
- [ ] Integrate Tier 3: VectorStoreRetrieverMemory (Qdrant semantic search)
|
|
- [ ] Create Qdrant collections (conversation_memory, documents, user_facts)
|
|
- [ ] Implement memory consolidation service
|
|
- [ ] Add conversation history API endpoints
|
|
- [ ] Test memory persistence across container restarts
|
|
|
|
### 📋 Planned (After Current Work)
|
|
|
|
**AI Orchestrator Phases 3-6:**
|
|
- Phase 3: Multi-agent workflows with LangGraph (Router, Chat, Research, Code agents)
|
|
- Phase 4: Tool integration (web search, web scrape, document search)
|
|
- Phase 5: RAG & advanced memory (hybrid retrieval, document upload)
|
|
- Phase 6: Production hardening (metrics, monitoring, optimization)
|
|
|
|
### 🔮 Backlog (Future Enhancements)
|
|
|
|
**Infrastructure Consolidation & Technical Debt:**
|
|
1. **Centralized Database Container:** Consolidate SQLite databases from multiple services (Uptime Kuma, Organizr, etc.) into a single PostgreSQL/MySQL container for easier management and backups
|
|
2. **Maintenance Container Consolidation:** Migrate maintenance container cron jobs into Core API endpoints with scheduled triggers - consolidate custom code into single service
|
|
3. **Version Control Setup:** Initialize portainer-core repository in Gitea for proper version control, branching, and change tracking
|
|
4. **Disaster Recovery Strategy:** Design offsite backup solution with restore/bootstrap scripts for full tower-of-joy recreation on new hardware
|
|
|
|
**Post-Phase 6 Integrations:**
|
|
- Nextcloud integration (file search, calendar management)
|
|
- ComfyUI integration (image generation)
|
|
- Home Assistant integration (smart home control)
|
|
- Custom mobile apps (iOS/Android)
|
|
|
|
## Current Blockers
|
|
|
|
None - All core services deployed and operational.
|
|
|
|
## Next Steps
|
|
|
|
### Priority 1: Core-API Refactoring & Infrastructure Management
|
|
|
|
**Why Now:** Clean up technical debt before adding more features. Build proper infrastructure management API to eliminate ad-hoc scripts and enable programmatic service deployment.
|
|
|
|
**Immediate Actions:**
|
|
1. **Refactor Core-API structure** - Create controller-based architecture for maintainability
|
|
2. **Build Infrastructure Management API** - Automate Portainer/NPM/Kuma operations
|
|
3. **Replace shell scripts** - Migrate `/stacks/*.sh` to proper API endpoints with CLI wrappers
|
|
|
|
**Benefits:**
|
|
- Cleaner codebase for future AI Orchestrator development
|
|
- Automated service deployment and monitoring setup
|
|
- Programmatic infrastructure management (no more manual NPM/Kuma configuration)
|
|
- Foundation for self-managing homelab
|
|
|
|
### Priority 2: AI Orchestrator Phase 2 (Memory Systems)
|
|
|
|
- **Implement 3-tier memory architecture:** ConversationBufferMemory (Tier 1), ConversationSummaryMemory (Tier 2), Qdrant VectorStore (Tier 3)
|
|
- **Create Qdrant collections:** conversation_memory, documents, user_facts
|
|
- **Build memory consolidation pipeline:** Automatic summarization and vector embedding
|
|
- **Add conversation history endpoints:** Query and manage conversation memory
|
|
- **Test integration with Open WebUI:** Verify memory persistence and recall
|
|
|
|
### Optional: Enhanced Capabilities
|
|
|
|
- **Connect Devices to Headscale:** Set up additional devices on mesh VPN for remote access
|
|
- **Jellyfin Media Library:** Populate media libraries with content
|
|
- **Nextcloud Desktop Clients:** Install sync clients on workstations
|
|
- **External Monitoring:** Set up Uptime Kuma notifications (email, Discord, etc.)
|
|
- **Advanced Automation:** Expand maintenance container with additional scheduled tasks
|
|
|
|
## Key Metrics
|
|
|
|
| Metric | Target | Current | Status |
|
|
|--------|--------|---------|--------|
|
|
| **Containers Running** | 15+ | 19 | 🟢 All Services Operational |
|
|
| **GPU Accessible** | Yes | Yes | 🟢 Working |
|
|
| **Storage Mounted** | 4.2TB | 3.6TB (58% used) | 🟢 Mounted |
|
|
| **Services Accessible** | All | 19/19 | 🟢 Complete |
|
|
| **Remote Access** | Working | Ready | 🟢 Headscale + NPM |
|
|
| **Firewall Active** | Yes | Yes | 🟢 UFW Configured |
|
|
| **Backups Configured** | Yes | Yes | 🟢 Maintenance Container |
|
|
| **Cloud Storage** | Yes | Yes | 🟢 Nextcloud Deployed |
|
|
| **File Sharing** | Yes | Yes | 🟢 Samba Deployed |
|
|
| **AI Orchestrator** | Phase 6 | Phase 1 ✅ | 🟡 In Progress (Phase 2 next) |
|
|
|
|
## Version History
|
|
|
|
- **v0.7.1-gitea-deployment** (2025-11-14): Gitea Git service deployed with PostgreSQL, NPM reverse proxy (https://git.schweitz.net), SSH port 2222, full Uptime Kuma + Organizr integration
|
|
- **v0.7.0-ai-orchestrator-phase1** (2025-11-13): AI Orchestrator Phase 1 complete - OpenAI-compatible API (`/v1/chat/completions`, `/v1/models`) with model aliasing and streaming support
|
|
- **v0.6.0-applications** (2025-11-13): Nextcloud and Samba deployed - cloud storage, file sharing, multi-network Uptime Kuma integration
|
|
- **v0.5.2-core-api** (2025-11-13): Core API deployed for Open WebUI functions, Python 3.12 upgrade (from EOL 3.8)
|
|
- **v0.5.1-open-webui** (2025-11-12): Open WebUI deployed with built-in voice capabilities (local STT/TTS)
|
|
- **v0.5.0-optimization** (2025-11-11): Phase 4 complete - Optimization & security (Watchtower, UFW, log rotation, maintenance container)
|
|
- **v0.4.0-monitoring** (2025-11-11): Phase 3 complete - Monitoring stack deployed (Uptime Kuma, Netdata, Heimdall)
|
|
- **v0.3.0-networking** (2025-11-11): Phase 2 complete - Headscale deployed with 10.99.0.0/16 mesh network
|
|
- **v0.2.0-foundation** (2025-11-11): Phase 1 complete - Portainer, NPM, Ollama deployed with GPU support
|
|
- **v0.1.0-planning** (2025-11-11): Project initialized, research and planning complete
|
|
|
|
---
|
|
|
|
## Quick Reference
|
|
|
|
**Documentation:**
|
|
- Architecture research: `containers/research.md`
|
|
- Implementation plan: `containers/implementation-plan.md`
|
|
- Agent guidelines: `AGENTS.md`
|
|
- System details: `SYSTEM.md`
|
|
|
|
**Key Paths:**
|
|
- SSD configs: `/home/jpmschweitzer/docker-data/`
|
|
- HDD content: `/mnt/media/`
|
|
- Stacks: Managed in Portainer web UI
|
|
|
|
**Active Services & Ports:**
|
|
- **Portainer:** http://192.168.86.149:8001 (container management)
|
|
- **Nginx Proxy Manager:** http://192.168.86.149:81 (reverse proxy admin)
|
|
- **AMP:** http://192.168.86.149:8080 (game servers - native)
|
|
- **Ollama:** http://192.168.86.149:11434 (ML models API)
|
|
- **Headscale:** http://192.168.86.149:8085 (mesh VPN control server)
|
|
- **Uptime Kuma:** http://192.168.86.149:3001 (service uptime monitoring)
|
|
- **Netdata:** http://192.168.86.149:19999 (real-time system metrics)
|
|
- **Heimdall:** http://192.168.86.149:8888 (unified dashboard)
|
|
- **Watchtower:** (background service - automatic updates daily at 4 AM)
|
|
- **Maintenance:** (background service - automated backups & scheduled tasks)
|
|
|
|
**Application Services:**
|
|
- **Open WebUI:** http://192.168.86.149:82 (LLM chat interface)
|
|
- **Core API:** http://192.168.86.149:8083 (OpenAPI functions for Open WebUI)
|
|
- **Jellyfin:** http://192.168.86.149:8096 OR https://media.schweitz.net (GPU-accelerated media server)
|
|
- **Nextcloud:** http://192.168.86.149:8082 OR https://cloud.schweitz.net (cloud storage & collaboration)
|
|
- **Gitea:** http://192.168.86.149:3002 OR https://git.schweitz.net (Git repository hosting, SSH: port 2222)
|
|
- **Samba:** \\\\192.168.86.149 or \\\\tower-of-joy (network file shares: Media, Downloads, Backups)
|
|
|
|
---
|
|
|
|
*Update this file as you complete each phase and checkpoint*
|
|
|
|
---
|
|
|
|
## Recent Updates
|
|
|
|
### 2025-11-13 Evening
|
|
**Phase 1 Testing & Bug Fix:**
|
|
- ✅ Completed comprehensive testing of Phase 1 implementation
|
|
- ✅ Fixed model ID formatting issue (extra quotes in model names)
|
|
- ✅ All 10/10 tests passing
|
|
- ✅ Zero known issues remaining
|
|
- ✅ Performance: 245ms average response time
|
|
- ✅ 100% OpenAI API compatibility verified
|
|
- ✅ Created comprehensive test results document (docs/phase1-test-results.md)
|
|
|
|
**Status:** Phase 1 100% complete and production-ready
|
|
**Next:** Begin Phase 2 (Memory Systems) implementation
|