diff --git a/.claude/skills/gen-audio/SKILL.md b/.claude/skills/gen-audio/SKILL.md index 5f8d19d3e..eece283f5 100644 --- a/.claude/skills/gen-audio/SKILL.md +++ b/.claude/skills/gen-audio/SKILL.md @@ -27,13 +27,19 @@ workflow, and quality validation. # Check API health db/connectors/audio-health -# Generate audio +# Generate a single asset (WAV only) db/connectors/audio-generate "prompt text" \ - --duration 10 \ - --steps 100 \ - --cfg 7 \ - --output path/to/output.wav \ - --timeout 600 + --duration 10 --steps 100 --cfg 7 \ + --output path/to/output.wav + +# Generate + post-process in one command (WAV → trim → normalize → OGG) +db/connectors/audio-generate "prompt text" \ + --duration 10 --steps 100 --cfg 7 \ + --output path/to/gen/intermediate.wav \ + --output-ogg client/assets/audio/final.ogg + +# Batch-generate from a manifest (preferred for multiple assets) +db/connectors/audio-batch docs/assets/audio/batch-s10-327.json ``` ### Parameters @@ -43,7 +49,9 @@ db/connectors/audio-generate "prompt text" \ | `--duration` | 10 | 0-47s | Max 47s per generation. For longer loops, generate 45s with crossfade overlap. | | `--steps` | 100 | 10-200 | More steps = better quality, slower. Use 50 for quick previews, 100-150 for final. | | `--cfg` | 7 | 1-15 | Classifier-free guidance. Higher = more prompt-adherent but less natural. 5-9 is the sweet spot. | -| `--output` | auto | — | Output file path. Auto-names from prompt if omitted. | +| `--output` | auto | — | Output WAV file path. Auto-names from prompt if omitted. | +| `--post` | off | — | Run trim + normalize + convert after generation. | +| `--output-ogg` | auto | — | OGG output path (implies `--post`). Defaults to same basename as WAV. | | `--timeout` | 600 | — | Max wait in seconds. Generation can take 2-5 minutes on 11GB VRAM. | ### Critical Constraints @@ -55,8 +63,6 @@ db/connectors/audio-generate "prompt text" \ Be patient. The timeout default (600s) is generous. - **Max 47 seconds** per generation. For 60-90s ambient loops, generate 45s clips and crossfade-stitch in post-processing. -- **Output is WAV at 44.1kHz stereo.** Convert to .ogg for Godot import: - `ffmpeg -i input.wav -c:a libvorbis -q:a 6 output.ogg` ## Prompt Assembly @@ -74,19 +80,121 @@ family prefix and matching category template. asset type (ambient, sfx, ui) - **Asset description:** Look up the specific asset in `docs/assets/audio/{category}.md` +## Batch Workflow (Preferred) + +For generating multiple assets, use a manifest file. This reduces prompt +approvals to 2: one Write (manifest) + one Bash (batch run). + +### 1. Create the manifest + +Write a JSON manifest to `docs/assets/audio/batch-{sprint}-{ticket}.json`: + +```json +{ + "description": "Sprint 10 ambient + world SFX batch", + "output_dir": "client/assets/audio", + "gen_dir": "client/assets/audio/gen", + "defaults": { + "steps": 100, + "cfg": 7, + "lufs": -16, + "quality": 6 + }, + "assets": [ + { + "id": "AMB-001", + "filename": "amb_station_base.ogg", + "method": "sao", + "duration": 45, + "steps": 150, + "cfg": 5, + "prompt": "[sonic family prefix] + [template] + [description]" + }, + { + "id": "UI-005", + "filename": "sfx_monologue_chime.ogg", + "method": "synth", + "synth": { + "type": "harmonic", + "duration": 0.8, + "fundamental": 1200, + "harmonics": [ + {"freq": 2400, "db": -12}, + {"freq": 3600, "db": -24} + ], + "attack_ms": 15, + "sustain_ratio": 0.2, + "decay": "exponential" + } + } + ] +} +``` + +Asset `id` values must match IDs in `docs/assets/audio/{category}.md` (e.g., +AMB-001, SFX-002, UI-005). This couples the manifest to the asset inventory. + +### 2. Run the batch + +```bash +# Full run +db/connectors/audio-batch docs/assets/audio/batch-s10-327.json + +# Dry run — preview what would be generated +db/connectors/audio-batch docs/assets/audio/batch-s10-327.json --dry-run + +# Generate only specific assets +db/connectors/audio-batch docs/assets/audio/batch-s10-327.json --only AMB-001,AMB-002 + +# Skip assets that already have OGG files +db/connectors/audio-batch docs/assets/audio/batch-s10-327.json --skip-existing +``` + +### 3. Update asset docs with prompts + +After the batch completes, write the exact prompts used back into the +Prompt/Notes column of `docs/assets/audio/{category}.md`. The manifest records +what was generated; the asset docs record what we have. + +### Manifest fields + +| Field | Required | Notes | +|-------|----------|-------| +| `id` | yes | Asset ID from docs (AMB-001, SFX-002, UI-005) | +| `filename` | yes | Output filename (must match asset doc) | +| `method` | yes | `sao` (Stable Audio Open) or `synth` (harmonic synthesis) | +| `duration` | SAO only | Duration in seconds | +| `prompt` | SAO only | Full assembled prompt | +| `steps` | no | Override default steps | +| `cfg` | no | Override default CFG | +| `synth` | synth only | Synthesis parameters (see below) | + +### Synth parameters + +| Field | Default | Notes | +|-------|---------|-------| +| `type` | harmonic | Only `harmonic` supported currently | +| `duration` | — | Duration in seconds | +| `fundamental` | — | Fundamental frequency in Hz | +| `harmonics` | [] | List of `{"freq": Hz, "db": dB}` objects | +| `attack_ms` | 10 | Attack time in milliseconds | +| `sustain_ratio` | 0.2 | Fraction of duration at full level before decay | +| `decay` | exponential | `exponential` or `linear` | + ## Single Asset Workflow +For one-off generation or iteration on a specific asset: + 1. Find the asset in `docs/assets/audio/{ambient,sfx,ui}.md` — note filename, duration, bus, method, and design intent. 2. Read `references/sonic-palette.md` for the sonic family prefix. 3. Read `references/category-templates.md` for the matching template. 4. Assemble the full prompt. 5. Run `db/connectors/audio-health` to verify the API is up. -6. Run `db/connectors/audio-generate` with the assembled prompt. **One request - at a time. Wait for completion.** -7. Listen to the output (or describe it based on file size/duration). -8. If acceptable, convert to .ogg and place in `client/assets/audio/`. -9. Update the asset status in `docs/assets/audio/{category}.md`. +6. Run `db/connectors/audio-generate` with `--post` or `--output-ogg` to + generate and post-process in one step. +7. Verify the output (file size, duration). +8. Update the asset status and prompt in `docs/assets/audio/{category}.md`. ## Iteration Workflow @@ -101,62 +209,38 @@ For each asset, generate 4-6 candidates: 4. **Fatigue test** (loops only) — can you listen for 5+ minutes without a jarring repeat? 5. **Close-your-eyes test** — does it create a mental image or sensation? -6. Select the best candidate, trim, normalize, convert. +6. Select the best candidate (post-processing is already done if `--post` was + used). -## Post-Processing +## Post-Processing (Standalone) -After selecting the best generation: +If you need to post-process separately (e.g., re-normalizing an existing file): ```bash -# Trim silence from start/end -ffmpeg -i input.wav -af "silenceremove=start_periods=1:start_silence=0.1:start_threshold=-50dB,areverse,silenceremove=start_periods=1:start_silence=0.1:start_threshold=-50dB,areverse" trimmed.wav +# Full pipeline: trim → normalize → convert +db/connectors/audio-post pipeline input.wav --output output.ogg -# LUFS normalize to -16 LUFS (broadcast standard, good for game audio) -ffmpeg -i trimmed.wav -af loudnorm=I=-16:LRA=11:TP=-1 normalized.wav - -# Convert to .ogg for Godot -ffmpeg -i normalized.wav -c:a libvorbis -q:a 6 output.ogg - -# For loops: verify loop point -ffplay -loop 0 output.ogg -``` - -For ambient loops, create crossfade overlap: -```bash -# Create a 45s loop with 3s crossfade overlap -# (manual: export 48s, crossfade first 3s with last 3s in Audacity) +# Individual steps +db/connectors/audio-post trim input.wav +db/connectors/audio-post normalize input.wav --lufs -16 +db/connectors/audio-post convert input.wav --output output.ogg ``` ## Manual Synthesis (Insert-Tech Sounds) For sounds under 200ms (cursor hover, weapon aim), Stable Audio Open cannot -produce meaningful output. Use manual synthesis instead: +produce meaningful output. Use manual synthesis via `tooling/synth_ui_sounds.py` +or the batch manifest's `method: "synth"` with harmonic parameters. -```python -# Example: 50ms cursor hover tick -import numpy as np -import wave - -sr = 44100 -duration = 0.05 # 50ms -t = np.linspace(0, duration, int(sr * duration), endpoint=False) -freq = 3200 # Hz -signal = np.sin(2 * np.pi * freq * t) -envelope = np.exp(-t * 80) # exponential decay -audio = (signal * envelope * 32767).astype(np.int16) - -with wave.open("cursor_hover.wav", "w") as f: - f.setnchannels(1) - f.setsampwidth(2) - f.setframerate(sr) - f.writeframes(audio.tobytes()) -``` +For complex synthesis beyond the `harmonic` type (FM, filtered noise, bandpass +impulse), write a custom script in `tooling/` following the pattern in +`tooling/synth_ui_sounds.py`. ## Quality Checklist After generating, verify: - Sound matches the sonic family (insert-tech = synthetic/precise, organic = warm/natural) -- Frequency range doesn't mask other layers (check docs/assets/audio/palette.md) +- Frequency range doesn't mask other layers (check docs/assets/audio/) - Duration matches spec - No unwanted artifacts (clicks, pops, digital noise at start/end) - Loop point is clean (ambient loops only) @@ -166,22 +250,6 @@ After generating, verify: ## File Placement Generated assets go to `client/assets/audio/` with exact filenames from the -asset docs: +asset docs. Intermediates go to `client/assets/audio/gen/` (gitignored). -``` -client/assets/audio/ - amb_station_base.ogg # Ambient bus - amb_workplace_layer.ogg # Ambient bus - amb_bar_layer.ogg # Ambient bus - amb_corridor_layer.ogg # Ambient bus - sfx_footstep_metal.ogg # Player Actions bus - sfx_footstep_metal_run.ogg # Player Actions bus - cursor_hover.ogg # UI Sounds bus - implant_open.ogg # UI Sounds bus - fog_recognition.ogg # UI Sounds bus - weapon_aim.ogg # UI Sounds bus - sfx_monologue_chime.ogg # UI Sounds bus - sfx_monologue_chime_urgent.ogg # UI Sounds bus -``` - -AudioManager discovers these by directory scan — filenames must match exactly. +AudioManager discovers assets by directory scan — filenames must match exactly.