docs(decisions): record audio discussion outcomes D-067 through D-074

Recognition chime timing (D-067), 5-bus audio architecture (D-068),
audio dip profiles (D-069), confrontation as cognitive vulnerability
(D-070), monologue chime placeholder strategy (D-071), universal
conversation murmur (D-072), zone crossfade (D-073), hybrid audio
generation (D-074). Amends D-038 scope, resolves Q-014.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
This commit is contained in:
2026-02-16 01:03:02 +01:00
co-authored by Claude Opus 4.6
parent 8ec875385a
commit ac86c8a72b
10 changed files with 126 additions and 15 deletions
+65 -2
View File
@@ -1,6 +1,6 @@
# Perception Decisions
How the player observes and interacts with the world: camera, fog, line-of-sight, sound, monologue, perception modes.
How the player observes and interacts with the world: camera, fog, line-of-sight, sound, monologue, perception modes, audio dip profiles, cognitive delays.
---
@@ -273,6 +273,69 @@ How the player observes and interacts with the world: camera, fog, line-of-sight
- **Raised by:** Stig (UI spec + no portraits), Lead (20% height constraint + max-width directive)
- **Dissent:** Stig initially proposed 25% height and 50% width centered. Lead constrained to 20% height and max-width.
### D-067: Recognition chime fires at onset of cognitive delay
- **Date:** 2026-02-16
- **Decision:** The monologue recognition chime ([D-038](scope.md#d-038-audio-in-v01-scope--8-files-via-stable-audio-open) assets 7-8) fires at the ONSET of the cognitive delay ([D-060](#d-060-cognitive-delay-for-fog-recognition)), not at its completion. Chime duration 300-400ms, overlapping with the start of the delay. The chime is "unresolved" — it opens a question, doesn't answer one. Full sequence when recognizing an entity in fog: (1) hear/sense something → (2) chime plays → (3) 0.6s cognitive delay begins (concurrent with chime tail) → (4) monologue text appears during delay ("Those footsteps... that's Kael's walk") → (5) blob transitions to [D-033](#d-033-entity-color--relationship-to-player) color + silhouette feature → (6) recognition complete.
- **Rationale:** The chime marks the character's attention shifting, not the recognition completing. It creates an "unresolved" sensation — something is happening in the character's mind. Monologue text provides the resolution. This prevents the chime from feeling like a UI notification and instead makes it part of the character's cognitive process.
- **Resolves:** Q-014
- **Cross-reference:** Audio assets ([D-038](scope.md#d-038-audio-in-v01-scope--8-files-via-stable-audio-open)), cognitive delay ([D-060](#d-060-cognitive-delay-for-fog-recognition)), fog recognition ([D-059](#d-059-fog--shader-based-five-layers-knowledge-graph-driven)), monologue ([D-016](#d-016-internal-monologue-as-core-perceptionatmosphere-system))
- **Raised by:** Gestalt (proposal), confirmed by Team Leader (Jeroen)
- **Dissent:** None
### D-069: Audio dip profiles for dialogue and confrontation
- **Date:** 2026-02-16
- **Decision:** Three audio dip profiles modify bus volumes during player focus states, all relative to player slider settings (proportional, not absolute):
- **Dialogue dip:** Ambient -6 to -8dB, World SFX 0dB, Player Actions 0dB, UI 0dB. 300ms ease-in, 500ms ease-out. Reduces environmental noise to foreground conversation without muting world events.
- **Confrontation dip:** Ambient -10 to -12dB + low-pass filter (20kHz → 800Hz), World SFX -4 to -6dB (graduated — loud events like sprinting footsteps break through, quiet sneaking doesn't), Player Actions 0dB, UI 0dB. 500ms ease-in, 1000ms ease-out with filter sweep. Creates emotional/cognitive muffling per [D-070](#d-070-confrontation-as-cognitive-vulnerability).
- **ListeningFocus boost:** World SFX +2 to +3dB when player is stationary for 30+ ticks (eavesdrop bonus, per [D-071](#d-071-no-ambient-dip-for-eavesdropping--listeningfocus-boost)). Expresses heightened attention as audio gain.
- **Key behaviors:**
- All dB values are proportional to player slider setting, not absolute. If player sets World SFX to 50%, the -6dB dip applies to that 50% base.
- Dips are interruptible — walk-away or stance change kills current tween, starts new tween to neutral.
- Graduated World SFX dip during confrontation: loud, urgent events (sprinting, doors slamming, alarms) break through at reduced volume. Quiet, careful movement may not be noticed.
- **Rationale:** Dips model finite attention. Dialogue suppresses ambient noise (focus on conversation). Confrontation suppresses both ambient and peripheral sounds (emotional tunnel vision per [D-070](#d-070-confrontation-as-cognitive-vulnerability)). ListeningFocus boost rewards deliberate eavesdropping. All effects proportional to user volume settings respects player accessibility choices.
- **Cross-reference:** Audio architecture ([D-068](architecture.md#d-068-5-bus-audio-architecture)), confrontation vulnerability ([D-070](#d-070-confrontation-as-cognitive-vulnerability)), eavesdropping ([D-071](#d-071-no-ambient-dip-for-eavesdropping--listeningfocus-boost)), audio aesthetic ([D-074](content.md#d-074-audio-aesthetic-identity--insert-tech-vs-organic))
- **Raised by:** Inigo (initial spec), Gestalt (graduated approach + ListeningFocus boost), Tyre (implementation details), Ozzie (proportional dip to respect player settings)
- **Dissent:** None
### D-070: Confrontation as cognitive vulnerability
- **Date:** 2026-02-16
- **Decision:** Confrontation creates perceptual vulnerability through reduced ambient awareness. Design principle: "Emotional focus creates perceptual vulnerability." Player actions requiring cognitive focus (confrontation dialogue, future: deep terminal reading, complex insert queries) mechanically reduce ambient perception via audio dip ([D-069](#d-069-audio-dip-profiles-for-dialogue-and-confrontation)) and may suppress peripheral monologue triggers. This is NOT a punishment — it's a realistic consequence of finite attention. The player mitigates risk by choosing WHERE and WHEN to confront (safe location vs exposed corridor, early shift vs busy period).
- **Mechanical expression:**
- During confrontation: ambient sounds muffled, quiet NPC movement may go undetected (graduated World SFX dip means careful footsteps suppressed, sprinting footsteps still audible).
- Post-confrontation: delayed monologue may fire for missed events ("Wait — did someone just pass by?").
- World continues per [D-045](#d-045-art-direction--environmental-neutrality-strict-zero-shift) environmental indifference — no scripted ambushes, but NPCs on routine may coincidentally pass while player is focused.
- Vulnerability is probabilistic and emergent, not scripted.
- **Consistency principle:** Connects to [D-053](scope.md#d-053-movement-as-stance-toggle-system) sprint suppressing monologue (physical focus limits interpretation). Confrontation suppresses ambient perception (emotional focus limits peripheral awareness). High-focus activities suppress peripheral cognition — this is the shared rule.
- **Must feel "felt, not computed" (Paula):** Muffling is psychological tunnel vision, not a debuff tooltip. No UI indicator. Player realizes in retrospect: "I was so focused on Sera I didn't hear Kael walk past."
- **Rationale:** Makes WHERE and WHEN to confront meaningful decisions. Confronting in a private corner = safer. Confronting in a busy corridor during shift change = riskier. Player agency through spatial and temporal choice, not RNG.
- **Cross-reference:** Audio dip ([D-069](#d-069-audio-dip-profiles-for-dialogue-and-confrontation)), confrontation UI ([D-063](content.md#d-063-confrontation--same-box-different-weight)), environmental indifference ([D-045](#d-045-art-direction--environmental-neutrality-strict-zero-shift)), stance system ([D-053](scope.md#d-053-movement-as-stance-toggle-system)), monologue ([D-016](#d-016-internal-monologue-as-core-perceptionatmosphere-system))
- **Raised by:** Gestalt (vulnerability principle), Ozzie (graduated dip approach), Paula ("felt not computed" framing)
- **Dissent:** None
### D-071: No ambient dip for eavesdropping — ListeningFocus boost
- **Date:** 2026-02-16
- **Decision:** Overheard conversation (eavesdropping) does NOT apply dialogue dip ([D-069](#d-069-audio-dip-profiles-for-dialogue-and-confrontation)). Cognitive state is opposite of direct conversation: extracting signal from noise requires MORE ambient awareness, not less. Instead, eavesdrop quality improves via ListeningFocus boost (+2 to +3dB World SFX when stationary 30+ ticks). Eavesdrop information quality is: distance-dependent, stance-modified ([D-053](scope.md#d-053-movement-as-stance-toggle-system) Careful stance grants "tell notice bonus"), ListeningFocus-boosted, and ambient-noise-penalized.
- **Ambient zone affects eavesdrop difficulty:**
- Bar (high ambient) = conversations hidden in murmur, requires proximity + ListeningFocus.
- Workplace (moderate ambient) = conversations semi-conspicuous, distance-dependent.
- Corridor (low ambient) = conversations conspicuous, easy to overhear from distance.
- **Information confidence:** Overheard information enters knowledge graph at lower confidence (KnowsOf, not KnowsDetails per [D-041](architecture.md#d-041-knowledge-graph-data-model)) unless ListeningFocus + proximity allows high-quality capture.
- **Rationale:** Direct conversation = character focuses, ambient dips. Overheard conversation = character strains to hear, ambient stays or boosts. Opposite cognitive states produce opposite audio profiles. Creates meaningful difference between talking TO someone vs listening IN on someone.
- **Cross-reference:** Audio dip ([D-069](#d-069-audio-dip-profiles-for-dialogue-and-confrontation)), stance system ([D-053](scope.md#d-053-movement-as-stance-toggle-system)), knowledge graph ([D-041](architecture.md#d-041-knowledge-graph-data-model)), zone audio ([D-072](#d-072-universal-event-driven-conversation-murmur))
- **Raised by:** Gestalt (cognitive state framing), Paula (zone-conspicuousness model), unanimously endorsed
- **Dissent:** None
### D-072: Universal event-driven conversation murmur
- **Date:** 2026-02-16
- **Decision:** NPC-to-NPC conversations use a single universal murmur sound asset, not zone-specific variants. Zone ambient determines conspicuousness: bar ambient (high murmur density) hides conversations, corridor ambient (low background) makes conversations conspicuous. Maps to [D-047](#d-047-art-direction--two-tier-animation-system) two-tier behavior visibility: bar conversations are Tier 1 (clearly happening, content obscured), corridor conversations are Tier 2 (ambiguous — are they talking or just standing near each other?).
- **Asset split:**
- Bar ambient murmur: baked into `amb_bar_layer.ogg` (continuous background per [D-038](scope.md#d-038-audio-in-v01-scope--8-files-via-stable-audio-open)).
- NPC proximity murmur: separate event-driven asset, deferred to future sprint (not in Sprint 7 scope). When implemented, same asset plays everywhere; zone ambient determines audibility.
- **Rationale:** One murmur asset + zone-dependent conspicuousness creates the signal/noise dynamic naturally. Bar = high noise floor, conversations blend in. Corridor = low noise floor, conversations stand out. Simpler content production, richer emergent behavior.
- **Cross-reference:** Audio assets ([D-038](scope.md#d-038-audio-in-v01-scope--8-files-via-stable-audio-open)), two-tier animation ([D-047](#d-047-art-direction--two-tier-animation-system)), eavesdropping ([D-071](#d-071-no-ambient-dip-for-eavesdropping--listeningfocus-boost))
- **Raised by:** Paula (zone-conspicuousness model), Inigo (scoping to future sprint)
- **Dissent:** None
---
*23 decisions. Last updated: 2026-02-14*
*29 decisions. Last updated: 2026-02-16*