Introduces a comprehensive, multi-tiered memory system to provide conversation history and context for the AI agent. This lays the foundation for more stateful and intelligent interactions. Key components of this implementation: - **Multi-Tiered Memory Architecture:** - **Tier 1 (Working Memory):** A fast, in-memory buffer (`ConversationBufferMemory`) that holds the most recent turns of a conversation for immediate access. - **Tier 3 (Long-Term Memory):** A persistent, semantic search-based memory store using Qdrant (`QdrantConversationMemory`). It stores all conversation turns as vector embeddings, enabling long-term recall and similarity search. - **Qdrant Integration:** - The `qdrant-client` is added to manage collections and perform vector search operations. - Each user is assigned a dedicated Qdrant collection for multi-tenancy. - **Ollama Embedding Client:** - A new `OllamaEmbeddingClient` generates text embeddings via the Ollama API, replacing the need for local sentence-transformer models. This significantly reduces the service's dependency footprint. - **Configuration and Stack Updates:** - The `config.py` and `core-ai.yml` stack file are updated with new settings for enabling memory, configuring Qdrant, and specifying the embedding model. - **Utility and Schema Additions:** - New Pydantic schemas (`memory/schemas.py`) define the data structures for conversation turns and memory management. - Utility functions (`utils.py`) are added for user ID sanitization and collection naming. This feature enhances the agent's capabilities by allowing it to maintain context across multiple turns and sessions, leading to more coherent and relevant responses.
Docker Compose Stacks
This directory contains version-controlled Docker Compose files for all services in the tower-of-joy infrastructure.
Deployment
review the http://core-api/docs openapi documentation for infrastructure management REST endpoints.
Stack Inventory
Phase 1: Foundation
| Stack | File | Ports | GPU | Description |
|---|---|---|---|---|
| Portainer | portainer.yml |
8080, 8443 | No | Container management UI |
| Nginx Proxy Manager | nginx-proxy-manager.yml |
8000, 80, 443 | No | Reverse proxy and unified web interface |
| Ollama | ollama.yml |
11434 | Yes | ML model serving with GPU acceleration |
Phase 2: Networking
| Stack | File | Ports | GPU | Description |
|---|---|---|---|---|
| Headscale | headscale.yml |
8085, 9090 | No | Self-hosted Tailscale control server |
Phase 3: Monitoring
| Stack | File | Ports | GPU | Description |
|---|---|---|---|---|
| Uptime Kuma | uptime-kuma.yml |
3001 | No | Service availability monitoring |
| Netdata | netdata.yml |
19999 | No | Real-time system performance monitoring |
| Heimdall | heimdall.yml |
8888, 8889 | No | Application dashboard |
Phase 4: Optimization
| Stack | File | Ports | GPU | Description |
|---|---|---|---|---|
| Watchtower | watchtower.yml |
- | No | Automatic container updates |
| Duplicati | duplicati.yml |
8200 | No | Backup solution |
Backlog: Applications
| Stack | File | Ports | GPU | Description |
|---|---|---|---|---|
| Jellyfin | jellyfin.yml |
8096, 8920, 7359, 1900 | Yes | Media server with GPU transcoding |
| Nextcloud | nextcloud.yml |
8082 | No | Cloud storage (includes DB and Redis) |
| Gitea | gitea.yml |
3002, 2222 | No | Git repository hosting (includes PostgreSQL) |
| Samba | samba.yml |
139, 445 | No | Network file sharing |
Port Allocation
Infrastructure Services (8000-8099)
- 8000: Nginx Proxy Manager (unified web interface)
- 8080: Portainer
- 8081: AMP (game servers - existing)
- 8082: Nextcloud
- 8085: Headscale
- 8096: Jellyfin
Git & Development Services
- 2222: Gitea SSH
- 3002: Gitea HTTP
Monitoring Services (3000-3999, 19000-19999)
- 3001: Uptime Kuma
- 8200: Duplicati
- 8888: Heimdall
- 19999: Netdata
ML/API Services (11000+)
- 11434: Ollama
Network Services
- 80: HTTP (NPM reverse proxy)
- 443: HTTPS (NPM reverse proxy)
- 139, 445: Samba/SMB
- 9090: Headscale metrics
Storage Convention
All stacks follow the dual-disk strategy:
SSD (Performance):
- Configs:
/home/jpmschweitzer/docker-data/<service>/config - Cache:
/home/jpmschweitzer/docker-data/<service>/cache - Databases:
/home/jpmschweitzer/docker-data/<service>/db
HDD (Capacity):
- User content:
/mnt/media/<service>/data - Media files:
/mnt/media/<service>/media - Backups:
/mnt/media/backups/<service>
GPU Services
Stacks requiring GPU access (marked with Yes above):
ollama.yml- ML model inferencejellyfin.yml- Hardware transcoding
Prerequisites:
- NVIDIA Container Toolkit installed
- GPU verified:
docker run --rm --gpus all nvidia/cuda:11.4.0-base-ubuntu20.04 nvidia-smi
Before Deploying
- Review environment variables - Change default passwords!
- Create directories - Ensure volume paths exist
- Check ports - Verify no conflicts with existing services
- GPU services - Confirm NVIDIA toolkit installed
- Update STATUS.md - Mark stack as deployed when complete
After Deploying
- Test service - Access web UI or API endpoint
- Check logs -
docker logs <container-name> - Verify GPU -
docker exec <container> nvidia-smi(if applicable) - Update documentation - Add to STATUS.md and CHANGELOG.md
- Configure backup - Add to Duplicati backup job
Maintenance
Update a Stack
# Pull latest images
docker compose -f stacks/<stack-name>.yml pull
# Recreate containers with new images
docker compose -f stacks/<stack-name>.yml up -d
# Or let Watchtower handle it automatically
Backup Stack Configuration
# Stacks are version-controlled in this directory
# Backup container data separately (see scripts/backup.sh)
Troubleshooting
- Container won't start:
docker logs <container-name> - Port conflicts:
sudo netstat -tulpn | grep <port> - Permission issues: Check volume path ownership
- GPU not detected: Verify NVIDIA toolkit and restart Docker
For detailed implementation instructions, see containers/implementation-plan.md