Files
scheduler/TASK_REGISTRATION.md
jpmschweitzerandClaude 933196c2a2 feat(executors): back up Portainer's own state
Portainer keeps every stack definition, endpoint, user and access-control
rule in a BoltDB inside the portainer_data Docker volume. That volume
sits under /var/lib/docker/volumes/, and the daily config backup covers
~/docker-data and code-server-config only — so the thing that defines all
24 stacks was the one thing not backed up.

Calls Portainer's /api/backup rather than tarring the volume. BoltDB is a
single memory-mapped file, so copying it while Portainer writes can
capture a torn page; the API serialises a consistent snapshot.

A 200 whose body is not a readable archive is treated as failure. An
archive that will not open is worse than a missing one, because it looks
like a backup until the day it is needed. Writing that check found a real
gap in it: a truncated tar.gz raises EOFError, which is neither TarError
nor OSError, so the first version of the guard let it through.

Archives contain TLS certificates and private keys and are written 0600.
Retention only ever deletes files matching the exact name this executor
writes, so an unrelated archive left in the same directory survives.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-08 19:56:44 +02:00

362 lines
12 KiB
Markdown

# Scheduler Task Registration Guide
## Overview
The Scheduler is **DB-driven** (not YAML-based). Tasks are registered via REST API and stored in PostgreSQL.
## Quick Reference
### Register a New Task
```bash
curl -X POST http://192.168.86.149:8090/tasks \
-H "Authorization: Bearer $SCHEDULER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"task_name": "my_task",
"service": "my-service",
"executor": "rest_api",
"priority": 25,
"description": "Task description",
"minute": 0,
"hour": -1,
"config": {
"method": "POST",
"url": "http://service:port/endpoint",
"headers": {
"Authorization": "Bearer ${ENV_VAR}"
},
"body": {}
}
}'
```
### List All Tasks
```bash
curl http://192.168.86.149:8090/tasks \
-H "Authorization: Bearer $SCHEDULER_API_KEY"
```
### Trigger Task Manually
```bash
curl -X POST http://192.168.86.149:8090/tasks/{task_name}/trigger \
-H "Authorization: Bearer $SCHEDULER_API_KEY"
```
### View Task Execution History
```bash
curl http://192.168.86.149:8090/executions?task_name={task_name} \
-H "Authorization: Bearer $SCHEDULER_API_KEY"
```
## Schedule Patterns
Use cron-style fields where **`-1` means "every"**:
| Pattern | minute | hour | day_of_month | month | day_of_week | Description |
|---------|--------|------|--------------|-------|-------------|-------------|
| Every minute | -1 | -1 | -1 | -1 | -1 | Runs every minute |
| Hourly | 0 | -1 | -1 | -1 | -1 | Runs at minute 0 of every hour |
| Daily at 2 AM | 0 | 2 | -1 | -1 | -1 | Runs at 2:00 AM every day |
| Weekly (Monday 2 AM) | 0 | 2 | -1 | -1 | 0 | Runs at 2:00 AM every Monday |
| Monthly (1st at 2 AM) | 0 | 2 | 1 | -1 | -1 | Runs at 2:00 AM on 1st of month |
**Note:** `day_of_week` is 0-6 (Monday-Sunday)
## Priority Levels
| Priority Range | Level | Use Cases |
|----------------|-------|-----------|
| 1-5 | Emergency/System | Critical system tasks |
| 10-30 | User | User-initiated tasks |
| 40-70 | Maintenance | Background maintenance |
| 70+ | Low | Low priority cleanup |
## Executor Types
### `rest_api` Executor
Calls HTTP endpoints. Supports environment variable substitution in headers/body.
**Config Format:**
```json
{
"method": "POST|GET|PUT|DELETE|PATCH",
"url": "http://service:port/endpoint",
"headers": {
"Authorization": "Bearer ${ENV_VAR}",
"Content-Type": "application/json"
},
"body": {
"key": "value"
},
"timeout": 30,
"auth": {
"type": "bearer|basic|api_key",
"value": "${TOKEN_ENV_VAR}"
}
}
```
**Environment Variables:** Use `${VAR_NAME}` syntax to reference env vars from the scheduler container.
### Other Executors
The `executor` field is the module name under `src/executors/`. These are the
modules that actually exist:
- `config_backup_executor`: tar.gz backup of mounted directories, with retention
- `gcs_backup_executor`: offsite backup to Google Cloud Storage
- `doc_sync_executor`: mirror upstream docs into Gitea
- `gitea_release_cleanup_executor`: drop old Gitea releases, keeping the newest N
- `postgres_retention_executor`: delete rows past a retention window (see below)
- `docker_prune_executor`: reclaim Docker disk usage (see below)
- `portainer_backup_executor`: archive Portainer's own state via its backup API (see below)
- `example_executor`: demo/test
There is **no `shell` or `python` executor**. Earlier revisions of this document
listed them and they were never implemented; work needing a shell belongs either
in a purpose-built executor or on a host systemd timer.
#### `postgres_retention_executor`
Connects with the Scheduler's own Postgres credentials, overriding only the
database name, so the target database must grant `scheduler_user` SELECT and
DELETE on the table. Table and column names are validated against a strict
identifier pattern because they cannot be bound as query parameters.
```json
{
"database": "sysmon",
"table": "check_history",
"timestamp_column": "ts",
"retention_days": 30,
"dry_run": false
}
```
#### `portainer_backup_executor`
Portainer keeps every stack definition, endpoint, user and access-control rule in
a BoltDB inside the `portainer_data` Docker volume, which lives under
`/var/lib/docker/volumes/` and is **not** covered by the daily config backup.
This calls Portainer's `/api/backup` rather than tarring the volume: BoltDB is a
single memory-mapped file, so copying it live can capture a torn page.
The archive contains TLS certificates and private keys and is written `0600`. A
200 response whose body is not a readable archive is treated as a failure — an
archive that will not open is worse than a missing one, because it looks like a
backup until the day it is needed.
```json
{
"url": "${PORTAINER_URL}",
"api_key": "${PORTAINER_API_KEY}",
"output_dir": "/backups/portainer",
"retention_days": 30
}
```
Portainer runs host-networked, so a container name does not resolve; use the
host address. Requires `/mnt/media/backups/portainer` mounted into the container.
#### `docker_prune_executor`
Uses the docker socket already mounted into the container. Only the two stages
that discard regenerable data are on by default.
```json
{
"build_cache": true,
"dangling_images": true,
"unused_images": false,
"volumes": false,
"build_cache_until_hours": 168,
"dry_run": false
}
```
**`volumes` removes volumes belonging to merely-stopped containers, not just
orphaned ones.** Leave it off unless you have checked what is currently
unattached; on this host it is a plausible way to lose a database.
## Complete Task Schema
### Required Fields
- `task_name` (string): Unique task identifier
- `service` (string): Service that owns this task (e.g., "library-desk", "core-api")
- `executor` (string): Executor type ("rest_api", "shell", etc.)
- `priority` (int): Priority level (1-100)
### Optional Fields
- `description` (string): Human-readable description
- `minute` (int): -1 to 59, default: -1 (every minute)
- `hour` (int): -1 to 23, default: -1 (every hour)
- `day_of_month` (int): -1 to 31, default: -1 (every day)
- `month` (int): -1 to 12, default: -1 (every month)
- `day_of_week` (int): -1 to 6, default: -1 (every day)
- `enabled` (bool): Whether task is enabled, default: true
- `max_retries` (int): Max retry attempts, default: 3
- `timeout_seconds` (int): Execution timeout, default: 3600
- `config` (object): Executor-specific configuration
- `created_by` (string): Creator identifier, default: "api"
## Example: Librarian Knowledge Consolidation
This task processes unprocessed search queries and consolidates knowledge into wiki pages.
**Registration:**
```bash
curl -X POST http://192.168.86.149:8090/tasks \
-H "Authorization: Bearer $SCHEDULER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"task_name": "librarian_consolidation",
"service": "library-desk",
"executor": "rest_api",
"priority": 25,
"description": "Processes unprocessed search queries and consolidates knowledge",
"minute": 0,
"hour": -1,
"config": {
"method": "POST",
"url": "http://library-desk:8089/consolidation/run",
"headers": {
"Authorization": "Bearer ${LIBRARY_API_KEY}",
"Content-Type": "application/json"
},
"body": {
"dry_run": false,
"process_limit": 10,
"lookback_days": 7,
"min_web_results": 2
}
}
}'
```
**Schedule:** Runs hourly (at minute 0 of every hour)
**What it does:**
1. Queries Neo4j for unprocessed `SearchQuery` nodes
2. Analyzes web results with Ollama
3. Creates/updates wiki pages via WikiPageWriter
4. Extracts and adds new entities to knowledge graph
5. Marks searches as processed
## API Endpoints
| Method | Endpoint | Description |
|--------|----------|-------------|
| GET | `/tasks` | List all tasks |
| GET | `/tasks/{task_name}` | Get task details |
| POST | `/tasks` | Create new task |
| PUT | `/tasks/{task_name}` | Update task |
| DELETE | `/tasks/{task_name}` | Delete task |
| POST | `/tasks/{task_name}/trigger` | Manually trigger task |
| GET | `/executions` | View execution history |
| GET | `/stats` | Get scheduler statistics |
| GET | `/health` | Health check |
## Interactive Documentation
Visit **http://192.168.86.149:8090/docs** for interactive Swagger UI with:
- Full API documentation
- Request/response schemas
- Try-it-out functionality
- Example payloads
## Environment Variables
Get API keys from containers:
```bash
# Scheduler API key
docker exec scheduler printenv SCHEDULER_API_KEY
# Library Desk API key
docker exec library-desk printenv LIBRARY_API_KEY
```
## Task Management Tips
1. **Test with dry_run**: Set `enabled: false` initially or use dry_run in task config
2. **Monitor executions**: Check `/executions` endpoint regularly
3. **Start conservative**: Use longer intervals (hourly, daily) before going to frequent runs
4. **Use priorities wisely**: Reserve 1-10 for critical tasks, use 20-30 for user tasks
5. **Set timeouts appropriately**: Long-running tasks need higher timeout_seconds
## Troubleshooting
### Task not running
1. Check if task is enabled: `GET /tasks/{task_name}`
2. Check execution history: `GET /executions?task_name={task_name}`
3. Verify schedule matches current time
4. Check scheduler logs: `docker logs scheduler`
### Task failing
1. Check execution history for error messages
2. Verify endpoint is accessible from scheduler container
3. Test endpoint manually with curl
4. Check API keys and authentication
5. Verify executor configuration (especially for rest_api)
### Debug a task
```bash
# Get task details
curl http://192.168.86.149:8090/tasks/librarian_consolidation \
-H "Authorization: Bearer $SCHEDULER_API_KEY"
# View recent executions
curl http://192.168.86.149:8090/executions?task_name=librarian_consolidation&limit=5 \
-H "Authorization: Bearer $SCHEDULER_API_KEY"
# Trigger manually
curl -X POST http://192.168.86.149:8090/tasks/librarian_consolidation/trigger \
-H "Authorization: Bearer $SCHEDULER_API_KEY"
```
## Architecture
```
┌─────────────────────────────────────────────────────────┐
│ The Scheduler (AsyncIOScheduler) │
│ │
│ ┌────────────────────────────────────────────┐ │
│ │ Cron Job (every minute) │ │
│ │ ↓ │ │
│ │ Query PostgreSQL for tasks scheduled │ │
│ │ for current minute/hour/day │ │
│ │ ↓ │ │
│ │ Execute up to 5 tasks concurrently │ │
│ │ (based on priority) │ │
│ │ ↓ │ │
│ │ Call appropriate executor: │ │
│ │ - rest_api → HTTP request │ │
│ │ - shell → subprocess │ │
│ │ - python → exec │ │
│ └────────────────────────────────────────────┘ │
│ │
│ PostgreSQL: │
│ - scheduled_tasks (task definitions) │
│ - task_executions (execution history) │
└─────────────────────────────────────────────────────────┘
```
## Current Registered Tasks
As of 2025-12-09:
1. **backup_docker_configs_daily** - Daily at 3:05 AM
2. **librarian_consolidation** - Hourly (minute 0)
3. **test_example_task** - Every minute (testing)
4. **sync_fastapi_docs_monthly** - 11th of month at 4:00 AM
5. **sync_ollama_docs_monthly** - 12th of month at 4:00 AM