Files
settled-reach/docs/workshops/llm-voice-pipeline/round-2-proposals.md
T
jpmschweitzer 23d9ff0a58 Merge remote-tracking branch 'origin/main' into planning
# Conflicts:
#	CHANGELOG.md
#	content/_meta/README.md
#	content/_meta/npc-authoring-style-guide.md
#	wiki/_templates/cultural-group.md
#	wiki/_templates/institution.md
#	wiki/_templates/star-system.md
#	wiki/characters/devra.md
#	wiki/characters/drin.md
#	wiki/characters/harek.md
#	wiki/characters/lera-sessik.md
#	wiki/characters/maret-korr.md
#	wiki/characters/naia-tamm.md
#	wiki/characters/nils-davan.md
#	wiki/characters/pell.md
#	wiki/characters/renn.md
#	wiki/characters/resha.md
#	wiki/characters/sabel.md
#	wiki/characters/sera-venn.md
#	wiki/characters/torek-lintar.md
#	wiki/characters/voss.md
#	wiki/star-systems/krenn/index.md
2026-03-14 00:24:53 +01:00

124 lines
7.7 KiB
Markdown

---
title: "Round 2 Proposals"
description: "Synthesised proposals from round 1 outputs for convergent evaluation in round 2"
type: workshop
status: archived
workshop: llm-voice-pipeline
agent: ""
round: 2
created: 2026-03-07
---
# LLM Voice Pipeline Workshop — Round 2 Proposals
**Compiled by:** Team Lead (synthesis of Round 1 outputs)
**Round:** 2 — Convergent Evaluation
**Date:** 2026-03-07
---
## Preamble: What Round 1 Settled
All 7 participants favor Option 3 (LLM re-voicing). Round 2 does not revisit that choice. Instead, it presents three concrete architecture variants that resolve the open tensions from Round 1.
**D-123/D-124 framing note:** D-123 (2026-03-05) frames generative AI as "an authoring tool, not a runtime system." D-124 defers in-game live AI but explicitly "leaves the door open." All three proposals below walk through that door to varying degrees. Each proposal must state how it amends D-123 and supersedes D-124.
**Resolved inputs for all proposals:**
- Runtime: `llama-cpp-rs` with GGUF Q4_K_M quantization (C-6)
- Determinism: cache-as-determinism — generate once per seed, cache result (C-7)
- Thread isolation: separate thread pool for inference vs. world generation (C-8)
- Base text is the fallback and the LLM seed (C-5)
- Existing zone RON content (~50 lines/role) survives as base text seeds (C-9)
---
## Proposal A: Conservative — Behaviors Only, Tells Locked
**Scope:** Re-voice observable behaviors only. Dialogue deferred to a future spike.
**Tell protection:** Base-text passthrough. Tells are never sent to the LLM. A new `tell_behaviors: Vec<String>` field is added to the NPC data model alongside `observable_behaviors`. The pipeline checks the field tag, not the content. Tell lines ship as-authored in all cases.
**Model:** Gemma 2 2B (Q4_K_M, ~1.5GB). Single candidate — no fallback model needed because behaviors are short-form (5-15 words) and the prompt is simple.
**Injector model:** Culture injector clauses (10-20 per culture) + trait modifier (1 per trait) + mood tag. Total injector budget: ~150 tokens. Authored by copy team, sourced from culture RON `speech` fields. Negative injectors required (explicit NOT-lists for lore contamination).
**Content tiers:**
- Baked: Sova Transit District + other hub zones, pre-voiced at build time
- Pre-voiced: background queue, prioritized by proximity and plot-criticality
- Fallback: base text (always available, always complete)
**D-123 amendment:** D-123 is amended to "authoring tool AND background runtime enhancement." D-124 is superseded — this IS the in-game AI system, scoped to behaviors.
**Pros:** Smallest risk surface. Behaviors are short-form, easy to validate. Tell safety is absolute (passthrough). Spike is simple: one model, one prompt template, one content type.
**Cons:** Leaves dialogue scaling unsolved. Dialogue is arguably the higher-value target for re-voicing. May feel like a half-measure if ambient behaviors get culture voice but dialogue stays template-assembled.
---
## Proposal B: Two-Track — Behaviors + Tells with Semantic Core
**Scope:** Re-voice observable behaviors AND tell behaviors, with different pipelines. Dialogue deferred.
**Tell protection:** Constrained re-voicing via `semantic_core` tag (Gestalt's proposal). The `Tell` struct gains a `semantic_core: String` field that names the phenomenon the tell must preserve (e.g., `"avoidance_behavior"`, `"nervous_fidget"`, `"concealment_tell"`). The re-voicing prompt includes an explicit constraint: `PRESERVE: {semantic_core}. Culture-voice the expression, not the phenomenon.`
**Model:** Gemma 2 2B (Q4_K_M). Two prompt templates: free re-voicing (ambient behaviors) and constrained re-voicing (tells).
**Injector model:** Same as Proposal A (150 token budget), plus semantic core constraints for tells (~20 additional tokens per tell).
**Content tiers:** Same as Proposal A.
**D-123 amendment:** Same as Proposal A.
**Pros:** Tells get culture voice (a Van Maanen's Star tell reads differently from a Sovari tell — richer world). The semantic core constraint is testable: spike can measure whether the phenomenon survives re-voicing. Two-track architecture is future-proof for dialogue.
**Cons:** Constrained re-voicing is harder to validate than passthrough. If the model fails to preserve the semantic core, the tell is corrupted — and the failure is subtle (not missing, just wrong). Requires the copy team to author semantic core labels for every tell type.
---
## Proposal C: Full Pipeline — Behaviors + Dialogue, Tells Locked
**Scope:** Re-voice observable behaviors AND dialogue. Tells pass through untouched (same as Proposal A).
**Tell protection:** Base-text passthrough (same as Proposal A). Tells are never re-voiced.
**Model:** Gemma 2 2B for behaviors; potentially Qwen2.5-3B (Q4, ~2.0GB) for dialogue if 2B quality is insufficient for longer-form output. Spike tests both on behaviors and dialogue separately.
**Dialogue re-voicing:** Dialogue lines already have access tier and trust tier tags (Paula's observation). The re-voicing prompt includes these as constraints alongside culture injectors. Dialogue is longer-form (15-40 words) and requires more prompt context (relationship state, conversation topic).
**Injector model:** Culture injectors (150 tokens) + dialogue context (relationship, access tier, trust tier — ~80 additional tokens). Total prompt budget ~400-500 tokens for dialogue.
**Content tiers:** Same as A, but baked content includes pre-voiced dialogue for hub NPCs.
**D-123 amendment:** D-123 is fully superseded. The AI pipeline is both an authoring tool and a runtime system. D-124 is superseded.
**Pros:** Solves the full content scaling problem in one architecture. Dialogue is where culture voice matters most to players (what NPCs SAY). Avoids the half-measure feeling of behaviors-only.
**Cons:** Larger spike scope. Dialogue quality at 2B may not meet the bar — may force a model size increase (3B) which tightens RAM. Two content types means two prompt templates, two validation passes, two quality bars. More can go wrong.
---
## Resolution Matrix
Each participant should evaluate all three proposals against their domain and answer:
| Question | Your answer |
|----------|-------------|
| Which proposal do you recommend? | A / B / C |
| Are there blockers in your recommended proposal? | Yes/No + details |
| Can you live with each of the other two proposals? | Yes/No per proposal |
| What is the minimum change to your non-preferred proposals that would make them acceptable? | |
### Open questions to resolve in Round 2:
**Q-R1-01 (tell literacy model):** Is the player's tell literacy cross-NPC grammar or fresh-each-time? — Gestalt, answer this. It determines whether Proposal B's constrained re-voicing is safe.
**Q-R1-02 (scope):** Proposals A and B defer dialogue; Proposal C includes it. — Tyre, Paula, assess feasibility and quality risk for each.
**Q-R1-03 (tell data model):** All three proposals require `tell_behaviors` as a separate field. This is now a prerequisite, not a question. — Gestalt, Tyre, confirm this is implementable.
**Q-R1-04 (token budget):** 150 tokens for culture injectors in A/B, 400-500 for dialogue in C. — Miri, is 150 sufficient for cultural philosophy? Troblum, what's the throughput impact of 500-token prompts vs 150?
**Q-R1-05 (minimum CPU spec):** This workshop cannot define it — it's a product decision. For Round 2, assume: 4-core CPU from 2019 or later (e.g., Intel i5-9400, Ryzen 5 3600). — Troblum, is this sufficient for Gemma 2B Q4 background inference?
**D-123 tension (Mellanie):** All three proposals amend or supersede D-123. — Mellanie, is the proposed amendment language acceptable? Paula, does this conflict with any narrative architecture constraints?