chore(skills): update gen-audio skill with batch workflow and --post docs
Documents audio-batch manifest format, synth parameters, and the --post flag as the preferred workflows. Batch reduces approval count from ~30 to 2 for multi-asset generation. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
This commit is contained in:
@@ -27,13 +27,19 @@ workflow, and quality validation.
|
||||
# Check API health
|
||||
db/connectors/audio-health
|
||||
|
||||
# Generate audio
|
||||
# Generate a single asset (WAV only)
|
||||
db/connectors/audio-generate "prompt text" \
|
||||
--duration 10 \
|
||||
--steps 100 \
|
||||
--cfg 7 \
|
||||
--output path/to/output.wav \
|
||||
--timeout 600
|
||||
--duration 10 --steps 100 --cfg 7 \
|
||||
--output path/to/output.wav
|
||||
|
||||
# Generate + post-process in one command (WAV → trim → normalize → OGG)
|
||||
db/connectors/audio-generate "prompt text" \
|
||||
--duration 10 --steps 100 --cfg 7 \
|
||||
--output path/to/gen/intermediate.wav \
|
||||
--output-ogg client/assets/audio/final.ogg
|
||||
|
||||
# Batch-generate from a manifest (preferred for multiple assets)
|
||||
db/connectors/audio-batch docs/assets/audio/batch-s10-327.json
|
||||
```
|
||||
|
||||
### Parameters
|
||||
@@ -43,7 +49,9 @@ db/connectors/audio-generate "prompt text" \
|
||||
| `--duration` | 10 | 0-47s | Max 47s per generation. For longer loops, generate 45s with crossfade overlap. |
|
||||
| `--steps` | 100 | 10-200 | More steps = better quality, slower. Use 50 for quick previews, 100-150 for final. |
|
||||
| `--cfg` | 7 | 1-15 | Classifier-free guidance. Higher = more prompt-adherent but less natural. 5-9 is the sweet spot. |
|
||||
| `--output` | auto | — | Output file path. Auto-names from prompt if omitted. |
|
||||
| `--output` | auto | — | Output WAV file path. Auto-names from prompt if omitted. |
|
||||
| `--post` | off | — | Run trim + normalize + convert after generation. |
|
||||
| `--output-ogg` | auto | — | OGG output path (implies `--post`). Defaults to same basename as WAV. |
|
||||
| `--timeout` | 600 | — | Max wait in seconds. Generation can take 2-5 minutes on 11GB VRAM. |
|
||||
|
||||
### Critical Constraints
|
||||
@@ -55,8 +63,6 @@ db/connectors/audio-generate "prompt text" \
|
||||
Be patient. The timeout default (600s) is generous.
|
||||
- **Max 47 seconds** per generation. For 60-90s ambient loops, generate 45s
|
||||
clips and crossfade-stitch in post-processing.
|
||||
- **Output is WAV at 44.1kHz stereo.** Convert to .ogg for Godot import:
|
||||
`ffmpeg -i input.wav -c:a libvorbis -q:a 6 output.ogg`
|
||||
|
||||
## Prompt Assembly
|
||||
|
||||
@@ -74,19 +80,121 @@ family prefix and matching category template.
|
||||
asset type (ambient, sfx, ui)
|
||||
- **Asset description:** Look up the specific asset in `docs/assets/audio/{category}.md`
|
||||
|
||||
## Batch Workflow (Preferred)
|
||||
|
||||
For generating multiple assets, use a manifest file. This reduces prompt
|
||||
approvals to 2: one Write (manifest) + one Bash (batch run).
|
||||
|
||||
### 1. Create the manifest
|
||||
|
||||
Write a JSON manifest to `docs/assets/audio/batch-{sprint}-{ticket}.json`:
|
||||
|
||||
```json
|
||||
{
|
||||
"description": "Sprint 10 ambient + world SFX batch",
|
||||
"output_dir": "client/assets/audio",
|
||||
"gen_dir": "client/assets/audio/gen",
|
||||
"defaults": {
|
||||
"steps": 100,
|
||||
"cfg": 7,
|
||||
"lufs": -16,
|
||||
"quality": 6
|
||||
},
|
||||
"assets": [
|
||||
{
|
||||
"id": "AMB-001",
|
||||
"filename": "amb_station_base.ogg",
|
||||
"method": "sao",
|
||||
"duration": 45,
|
||||
"steps": 150,
|
||||
"cfg": 5,
|
||||
"prompt": "[sonic family prefix] + [template] + [description]"
|
||||
},
|
||||
{
|
||||
"id": "UI-005",
|
||||
"filename": "sfx_monologue_chime.ogg",
|
||||
"method": "synth",
|
||||
"synth": {
|
||||
"type": "harmonic",
|
||||
"duration": 0.8,
|
||||
"fundamental": 1200,
|
||||
"harmonics": [
|
||||
{"freq": 2400, "db": -12},
|
||||
{"freq": 3600, "db": -24}
|
||||
],
|
||||
"attack_ms": 15,
|
||||
"sustain_ratio": 0.2,
|
||||
"decay": "exponential"
|
||||
}
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
Asset `id` values must match IDs in `docs/assets/audio/{category}.md` (e.g.,
|
||||
AMB-001, SFX-002, UI-005). This couples the manifest to the asset inventory.
|
||||
|
||||
### 2. Run the batch
|
||||
|
||||
```bash
|
||||
# Full run
|
||||
db/connectors/audio-batch docs/assets/audio/batch-s10-327.json
|
||||
|
||||
# Dry run — preview what would be generated
|
||||
db/connectors/audio-batch docs/assets/audio/batch-s10-327.json --dry-run
|
||||
|
||||
# Generate only specific assets
|
||||
db/connectors/audio-batch docs/assets/audio/batch-s10-327.json --only AMB-001,AMB-002
|
||||
|
||||
# Skip assets that already have OGG files
|
||||
db/connectors/audio-batch docs/assets/audio/batch-s10-327.json --skip-existing
|
||||
```
|
||||
|
||||
### 3. Update asset docs with prompts
|
||||
|
||||
After the batch completes, write the exact prompts used back into the
|
||||
Prompt/Notes column of `docs/assets/audio/{category}.md`. The manifest records
|
||||
what was generated; the asset docs record what we have.
|
||||
|
||||
### Manifest fields
|
||||
|
||||
| Field | Required | Notes |
|
||||
|-------|----------|-------|
|
||||
| `id` | yes | Asset ID from docs (AMB-001, SFX-002, UI-005) |
|
||||
| `filename` | yes | Output filename (must match asset doc) |
|
||||
| `method` | yes | `sao` (Stable Audio Open) or `synth` (harmonic synthesis) |
|
||||
| `duration` | SAO only | Duration in seconds |
|
||||
| `prompt` | SAO only | Full assembled prompt |
|
||||
| `steps` | no | Override default steps |
|
||||
| `cfg` | no | Override default CFG |
|
||||
| `synth` | synth only | Synthesis parameters (see below) |
|
||||
|
||||
### Synth parameters
|
||||
|
||||
| Field | Default | Notes |
|
||||
|-------|---------|-------|
|
||||
| `type` | harmonic | Only `harmonic` supported currently |
|
||||
| `duration` | — | Duration in seconds |
|
||||
| `fundamental` | — | Fundamental frequency in Hz |
|
||||
| `harmonics` | [] | List of `{"freq": Hz, "db": dB}` objects |
|
||||
| `attack_ms` | 10 | Attack time in milliseconds |
|
||||
| `sustain_ratio` | 0.2 | Fraction of duration at full level before decay |
|
||||
| `decay` | exponential | `exponential` or `linear` |
|
||||
|
||||
## Single Asset Workflow
|
||||
|
||||
For one-off generation or iteration on a specific asset:
|
||||
|
||||
1. Find the asset in `docs/assets/audio/{ambient,sfx,ui}.md` — note filename,
|
||||
duration, bus, method, and design intent.
|
||||
2. Read `references/sonic-palette.md` for the sonic family prefix.
|
||||
3. Read `references/category-templates.md` for the matching template.
|
||||
4. Assemble the full prompt.
|
||||
5. Run `db/connectors/audio-health` to verify the API is up.
|
||||
6. Run `db/connectors/audio-generate` with the assembled prompt. **One request
|
||||
at a time. Wait for completion.**
|
||||
7. Listen to the output (or describe it based on file size/duration).
|
||||
8. If acceptable, convert to .ogg and place in `client/assets/audio/`.
|
||||
9. Update the asset status in `docs/assets/audio/{category}.md`.
|
||||
6. Run `db/connectors/audio-generate` with `--post` or `--output-ogg` to
|
||||
generate and post-process in one step.
|
||||
7. Verify the output (file size, duration).
|
||||
8. Update the asset status and prompt in `docs/assets/audio/{category}.md`.
|
||||
|
||||
## Iteration Workflow
|
||||
|
||||
@@ -101,62 +209,38 @@ For each asset, generate 4-6 candidates:
|
||||
4. **Fatigue test** (loops only) — can you listen for 5+ minutes without a
|
||||
jarring repeat?
|
||||
5. **Close-your-eyes test** — does it create a mental image or sensation?
|
||||
6. Select the best candidate, trim, normalize, convert.
|
||||
6. Select the best candidate (post-processing is already done if `--post` was
|
||||
used).
|
||||
|
||||
## Post-Processing
|
||||
## Post-Processing (Standalone)
|
||||
|
||||
After selecting the best generation:
|
||||
If you need to post-process separately (e.g., re-normalizing an existing file):
|
||||
|
||||
```bash
|
||||
# Trim silence from start/end
|
||||
ffmpeg -i input.wav -af "silenceremove=start_periods=1:start_silence=0.1:start_threshold=-50dB,areverse,silenceremove=start_periods=1:start_silence=0.1:start_threshold=-50dB,areverse" trimmed.wav
|
||||
# Full pipeline: trim → normalize → convert
|
||||
db/connectors/audio-post pipeline input.wav --output output.ogg
|
||||
|
||||
# LUFS normalize to -16 LUFS (broadcast standard, good for game audio)
|
||||
ffmpeg -i trimmed.wav -af loudnorm=I=-16:LRA=11:TP=-1 normalized.wav
|
||||
|
||||
# Convert to .ogg for Godot
|
||||
ffmpeg -i normalized.wav -c:a libvorbis -q:a 6 output.ogg
|
||||
|
||||
# For loops: verify loop point
|
||||
ffplay -loop 0 output.ogg
|
||||
```
|
||||
|
||||
For ambient loops, create crossfade overlap:
|
||||
```bash
|
||||
# Create a 45s loop with 3s crossfade overlap
|
||||
# (manual: export 48s, crossfade first 3s with last 3s in Audacity)
|
||||
# Individual steps
|
||||
db/connectors/audio-post trim input.wav
|
||||
db/connectors/audio-post normalize input.wav --lufs -16
|
||||
db/connectors/audio-post convert input.wav --output output.ogg
|
||||
```
|
||||
|
||||
## Manual Synthesis (Insert-Tech Sounds)
|
||||
|
||||
For sounds under 200ms (cursor hover, weapon aim), Stable Audio Open cannot
|
||||
produce meaningful output. Use manual synthesis instead:
|
||||
produce meaningful output. Use manual synthesis via `tooling/synth_ui_sounds.py`
|
||||
or the batch manifest's `method: "synth"` with harmonic parameters.
|
||||
|
||||
```python
|
||||
# Example: 50ms cursor hover tick
|
||||
import numpy as np
|
||||
import wave
|
||||
|
||||
sr = 44100
|
||||
duration = 0.05 # 50ms
|
||||
t = np.linspace(0, duration, int(sr * duration), endpoint=False)
|
||||
freq = 3200 # Hz
|
||||
signal = np.sin(2 * np.pi * freq * t)
|
||||
envelope = np.exp(-t * 80) # exponential decay
|
||||
audio = (signal * envelope * 32767).astype(np.int16)
|
||||
|
||||
with wave.open("cursor_hover.wav", "w") as f:
|
||||
f.setnchannels(1)
|
||||
f.setsampwidth(2)
|
||||
f.setframerate(sr)
|
||||
f.writeframes(audio.tobytes())
|
||||
```
|
||||
For complex synthesis beyond the `harmonic` type (FM, filtered noise, bandpass
|
||||
impulse), write a custom script in `tooling/` following the pattern in
|
||||
`tooling/synth_ui_sounds.py`.
|
||||
|
||||
## Quality Checklist
|
||||
|
||||
After generating, verify:
|
||||
- Sound matches the sonic family (insert-tech = synthetic/precise, organic = warm/natural)
|
||||
- Frequency range doesn't mask other layers (check docs/assets/audio/palette.md)
|
||||
- Frequency range doesn't mask other layers (check docs/assets/audio/)
|
||||
- Duration matches spec
|
||||
- No unwanted artifacts (clicks, pops, digital noise at start/end)
|
||||
- Loop point is clean (ambient loops only)
|
||||
@@ -166,22 +250,6 @@ After generating, verify:
|
||||
## File Placement
|
||||
|
||||
Generated assets go to `client/assets/audio/` with exact filenames from the
|
||||
asset docs:
|
||||
asset docs. Intermediates go to `client/assets/audio/gen/` (gitignored).
|
||||
|
||||
```
|
||||
client/assets/audio/
|
||||
amb_station_base.ogg # Ambient bus
|
||||
amb_workplace_layer.ogg # Ambient bus
|
||||
amb_bar_layer.ogg # Ambient bus
|
||||
amb_corridor_layer.ogg # Ambient bus
|
||||
sfx_footstep_metal.ogg # Player Actions bus
|
||||
sfx_footstep_metal_run.ogg # Player Actions bus
|
||||
cursor_hover.ogg # UI Sounds bus
|
||||
implant_open.ogg # UI Sounds bus
|
||||
fog_recognition.ogg # UI Sounds bus
|
||||
weapon_aim.ogg # UI Sounds bus
|
||||
sfx_monologue_chime.ogg # UI Sounds bus
|
||||
sfx_monologue_chime_urgent.ogg # UI Sounds bus
|
||||
```
|
||||
|
||||
AudioManager discovers these by directory scan — filenames must match exactly.
|
||||
AudioManager discovers assets by directory scan — filenames must match exactly.
|
||||
|
||||
Reference in New Issue
Block a user