docs(workshops): LLM voice pipeline workshop — D-138, D-123 amended, D-124 superseded

3-round workshop (7 participants + Qatux + SI) deciding content generation
architecture for NPC observable behaviors and dialogue.

Key decisions:
- D-138: LLM re-voicing pipeline (Gemma 2B Q4, llama-cpp-rs, bundled)
- Behaviors + dialogue both re-voiced; tells always passthrough
- Tells as read-only context inputs shaping surrounding content tone
- Cache-as-determinism, separate thread pools, layered hardware detection
- Two-spike validation: plumbing first, then integration
- D-123 amended (authoring tool + runtime enhancement)
- D-124 superseded (door walked through)
- Q-012 and Q-057 resolved

Artifacts: Krenn injectors v2, NI-1-5, culture template, 6 dialogue
constraints, tell-tone injectors, spike payloads, 12-risk register.
9 tickets created (#638-#647).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
This commit is contained in:
2026-03-07 13:31:27 +01:00
co-authored by Claude Opus 4.6
parent 6e6a3c1304
commit 9f91fd077f
30 changed files with 7283 additions and 12 deletions
+1 -1
View File
@@ -12,7 +12,7 @@ Cross-domain decisions live in one file with cross-reference notes in related fi
|------|--------|-----------|
| [architecture.md](architecture.md) | Technical foundation | D-008, D-009, D-010, D-012, D-020, D-026, D-030, D-031, D-041, D-042, D-054, D-055, D-066, D-068, D-073, D-085, D-088, D-094, D-096, D-097, D-099, D-100, D-101, D-102, D-103, D-106, D-108, D-109, D-113, D-133, D-134, D-135, D-136, D-137 |
| [perception.md](perception.md) | Player observation | D-011, D-015, D-016, D-017, D-018, D-019, D-033, D-035, D-043, D-044, D-045, D-046, D-047, D-048, D-049, D-052, D-056, D-057, D-058, D-059, D-060, D-061, D-067, D-069, D-070, D-071, D-072, D-076, D-077, D-078, D-086 |
| [content.md](content.md) | NPC, dialogue, templates | D-023, D-024, D-025, D-028, D-029, D-032, D-034, D-035, D-036, D-037, D-050, D-062, D-063, D-064, D-074, D-075, D-084, D-090, D-092, D-093, D-095, D-098, D-104, D-105, D-107, D-121, D-122, D-123, D-124, D-125, D-126, D-127, D-128, D-129, D-130, D-131, D-132 |
| [content.md](content.md) | NPC, dialogue, templates | D-023, D-024, D-025, D-028, D-029, D-032, D-034, D-035, D-036, D-037, D-050, D-062, D-063, D-064, D-074, D-075, D-084, D-090, D-092, D-093, D-095, D-098, D-104, D-105, D-107, D-121, D-122, D-123, D-124, D-125, D-126, D-127, D-128, D-129, D-130, D-131, D-132, D-138 |
| [scope.md](scope.md) | Game concept, prototype | D-001, D-003, D-005, D-006, D-007, D-013, D-014, D-027, D-038, D-039, D-051, D-053, D-065, D-087, D-089, D-091, D-114, D-115, D-116, D-117, D-118, D-119, D-120 |
| [process.md](process.md) | Team, workflow | D-004, D-021, D-022, D-040 |
| [questions.md](questions.md) | Open questions (index) | Q-001 through Q-054 |
+40 -10
View File
@@ -314,22 +314,29 @@ How narrative, NPCs, and world content are created: content tiers, NPC generatio
- **Supersedes:** Named NPC assignments in [D-034](#d-034-the-friend--production-level-npc-pattern) (Kael/Sera as hand-authored characters — see amendment on D-034)
- **Cross-reference:** [D-123](#d-123-generative-ai-for-npc-content-templating-via-culture-vectors) (AI templating), [D-129](#d-129-npc-personality-traits--behavior-first-relationships-codified-for-systems) (NPC personality model)
### D-123: Generative AI for NPC content templating via culture vectors
### D-123: Generative AI for NPC content — build-time authoring tool and runtime voice pipeline
- **Date:** 2026-03-05
- **Decision:** NPC content (dialogue pools, voice, vocabulary) is generated using generative AI with culture vectors, tone, and accent prompts as constraints. Culture vectors are the primary prompt constraint — they prevent the AI pipeline from defaulting to genre conventions. The AI pipeline is an authoring tool for content assembly, not a runtime system. Limited vocabulary acceptable at first; AI templating scales content as the generator matures.
- **Rationale:** The copy pool for all-generated NPCs at scale is enormous. Generative AI with culture-vector constraints is the only viable path to populating it without hand-authoring every line. Culture profiles (Miri prerequisite) become the primary authoring deliverable feeding the pipeline.
- **Source:** Where's the Fun? Workshop, Round 4 Interview, Decision 8
- **Date (amended):** 2026-03-07
- **Decision:** The AI pipeline operates in two distinct modes with different safety profiles:
- **Build-time mode (authoring tool):** Content generated at build time for baked hub zones. Subject to mandatory human review before shipping. AI as an accelerated authoring tool producing content humans review and approve.
- **Runtime mode (background enhancement):** Content generated during gameplay for non-baked zones, via a background inference queue, when "AI-Enhanced Dialogue" is enabled. Not human-reviewed per line. Safety provided by three layers: (1) base-text-as-fallback — always present and complete; (2) build-time-validated injectors — only pre-validated prompts used, never ad-hoc; (3) runtime contamination filter — lightweight check before content is served.
- **Non-negotiable constraints (both modes):** Culture vectors are the primary prompt constraint. The AI does not default to genre conventions. Authorial control governs what the LLM may and may not produce through injector clauses, negative constraints, and pipeline routing rules. The AI pipeline applies voice to authored semantic content; it does not generate narrative decisions, base text, tell behaviors, secret-tier dialogue (D-028 Layer 3), or anchor lines (D-092). These categories are always authored and always served as-authored.
- **Rationale:** Full pipeline (behaviors + dialogue) is the correct scope. A system that voices observed behavior but not spoken dialogue creates register whiplash at the highest-investment moment of player engagement. Build-time mode preserves the human-review safety model. Runtime mode enables scaling to the generated world with base-text fallback as the permanent safety net.
- **Source:** Where's the Fun? Workshop (original); LLM Voice Pipeline Workshop (amendment)
- **Raised by:** Team Leader (Jeroen)
- **Dissent:** None
- **Dissent:** None on amendment
- **Amended by:** [D-138](#d-138-llm-re-voicing-pipeline-for-npc-voice) (LLM Voice Pipeline Workshop, 2026-03-07)
- **Cross-reference:** [D-121](#d-121-voice-is-culture-driven--job-as-modifier) (culture-primary voice), [D-128](#d-128-culture-implicit-in-starting-location--krenn-system-equals-krenn-culture) (culture profile as generator input)
### D-124: In-game ollama for live NPC dialogue — deferred, door open
### D-124: In-game ollama for live NPC dialogue — ~~deferred~~ SUPERSEDED by D-138
- **Date:** 2026-03-05
- **Decision:** Running a dressed-down version of ollama in-game for live NPC dialogue is possible and interesting, but deferred. The door is explicitly left open — this is not a rejected alternative, it is a future investigation item. For v0.2, NPC dialogue uses template-assembled content (D-123). Live in-game AI dialogue is post-proof-of-life.
- **Rationale:** Live AI dialogue requires solving NPC quality floor, performance, and determinism questions that are out of scope for the generator proof-of-life. Deferred until the base generator is proven solid.
- **Source:** Where's the Fun? Workshop, Round 4 Interview, Decision 9
- **Date (superseded):** 2026-03-07
- **Decision:** ~~Running a dressed-down version of ollama in-game for live NPC dialogue is possible and interesting, but deferred.~~ **Superseded by [D-138](#d-138-llm-re-voicing-pipeline-for-npc-voice).** The in-game AI system uses `llama-cpp-rs` (not ollama) with GGUF Q4_K_M quantization, bundled with the game, running background inference via an isolated thread pool. The key constraint from D-124 remains binding through D-123 (amended): this system does not drive live narrative decisions. It applies voice to authored semantic content.
- **Rationale:** The LLM Voice Pipeline Workshop (2026-03-07) walked through the door D-124 left open. The quality, performance, and determinism questions D-124 cited as blockers are addressed by cache-as-determinism, base-text fallback, and layered hardware detection.
- **Source:** Where's the Fun? Workshop (original); LLM Voice Pipeline Workshop (supersession)
- **Raised by:** Team Leader (Jeroen)
- **Dissent:** None
- **Superseded by:** [D-138](#d-138-llm-re-voicing-pipeline-for-npc-voice)
### D-125: World is quietly responsive — gradient of caring by social proximity
- **Date:** 2026-03-05
@@ -400,6 +407,29 @@ How narrative, NPCs, and world content are created: content tiers, NPC generatio
- **Dissent:** None
- **Cross-reference:** [D-129](#d-129-npc-personality--traits--behavior-first-relationships-codified-for-systems) (relationships as consequence substrate)
### D-138: LLM Re-voicing Pipeline for NPC Voice
- **Date:** 2026-03-07
- **Decision:** NPC observable behaviors and dialogue are processed through an LLM re-voicing pipeline that translates culture-neutral semantic base text into character-voiced output. The pipeline is a background runtime enhancement, not a live generation system. Tell behaviors are base-text passthrough — always. Active tell state influences the re-voicing prompt for surrounding content (tells are read-only inputs to the LLM, never LLM outputs). The game is complete and functional without the pipeline; it is an enhancement that elevates voice quality for players with sufficient hardware.
- **Architecture:**
- **Model:** Gemma 2 2B (Q4_K_M, ~1.5GB), bundled with game. Phi-3 (MIT) as fallback. No Chinese-origin models.
- **Runtime:** `llama-cpp-rs` with GGUF format. Separate inference thread pool at below-normal priority.
- **Content tiers:** Baked (hub zones, build-time, human-reviewed) → Pre-voiced (background queue, priority-ordered) → Base text fallback (always present).
- **Tell treatment:** Passthrough always. Tell state flows into re-voicing prompts as universal tone injectors. Cultural flavor is conditional and additive — humans are humans first; micro-expressions and body language must remain universally recognizable. Per-culture tell-tone tables are optional enrichment, not a launch requirement.
- **Determinism:** Cache-as-determinism. LLM generates once per seed; result cached. Cache lookup is deterministic.
- **Caching:** 6 variants per line (neutral + 5 TellCategory states). Key: `(npc_stable_id, line_id, tell_state, culture_id)`.
- **Hardware:** "AI-Enhanced Dialogue" toggle. Layered detection: RAM check → TPT benchmark → recommendation. No hard minimum spec floor. Player can always override.
- **Distribution:** Model bundled in game install (~1.5GB).
- **Protected categories (never re-voiced):** Tell behaviors, secret-tier dialogue (D-028 Layer 3), anchor lines (D-092), relationship-specific lines naming third parties.
- **Validation:** Two-spike strategy. Spike 1: Rust `sr-voice` CLI + manual prompt testing (Jeroen/Mellanie/Paula). Spike 2: full pipeline integration.
- **Rationale:** D-122 (all NPCs generated) and D-128 (culture implicit in starting location) require NPC voice to scale across zones and cultures without O(R×Z×C) hand-authoring. The re-voicing model is the only architecture that scales while preserving content quality. Base-text fallback ensures the game is complete without the pipeline.
- **Source:** LLM Voice Pipeline Workshop (2026-03-07)
- **Raised by:** Team Leader (Jeroen), with Gestalt, Tyre, Paula, Mellanie, Ozzie, Miri, Troblum
- **Dissent:** Miri flagged concern about cultural philosophy at 2B model size — addressed via hybrid injector format (instruction + example pairs) and spike validation.
- **Amends:** [D-123](#d-123-generative-ai-for-npc-content--build-time-authoring-tool-and-runtime-voice-pipeline) (scope extended from authoring tool to authoring + runtime)
- **Supersedes:** [D-124](#d-124-in-game-ollama-for-live-npc-dialogue--deferred-superseded-by-d-138) (in-game AI no longer deferred)
- **Resolves:** Q-057 (composable behavior generation), Q-012 (generation expansion method)
- **Cross-reference:** [D-010](architecture.md#d-010) (information boundaries), [D-121](#d-121-voice-is-culture-driven--job-as-modifier) (culture-primary voice), [D-122](#d-122-all-npcs-generated--named-npcs-deferred) (all NPCs generated), [D-128](#d-128-culture-implicit-in-starting-location--krenn-system-equals-krenn-culture) (culture as generator input), [D-029](#d-029-population-entanglement-ratio--305020) (NPC tier model), [D-092](perception.md#d-092) (anchor lines)
---
*37 decisions. Last updated: 2026-03-05 (D-121–D-132 added; D-023, D-024, D-028, D-029, D-032, D-034, D-036 amended; D-032 superseded — Where's the Fun? Workshop)*
*38 decisions. Last updated: 2026-03-07 (D-138 added; D-123 amended; D-124 superseded — LLM Voice Pipeline Workshop)*
+2 -1
View File
@@ -10,10 +10,11 @@ Narrative, NPCs, dialogue, templates, setting, worldbuilding, and storyteller me
- **Assigned to:** Gestalt, Nigel
### Q-012: Generation expansion method for dialogue
- **Status:** Open
- **Status:** Resolved
- **Question:** How does the 4x generation expansion pass work? LLM-based, template-based, or rule-based? Affects how base lines are authored — LLM needs style-strong anchors; rules need substitution patterns.
- **Assigned to:** Gestalt, Mellanie
- **Source:** Content Gap Analysis Workshop (Mellanie R2)
- **Resolution:** LLM-based re-voicing via bundled Gemma 2B Q4. Culture-neutral semantic base text is the LLM seed; culture injectors + trait modifiers + tell-context tone shape the output. Resolved by [D-138](content.md#d-138-llm-re-voicing-pipeline-for-npc-voice) (LLM Voice Pipeline Workshop, 2026-03-07).
### Q-013: Line previewer temporal progression
- **Status:** Open