# Conflicts: # CHANGELOG.md # content/_meta/README.md # content/_meta/npc-authoring-style-guide.md # wiki/_templates/cultural-group.md # wiki/_templates/institution.md # wiki/_templates/star-system.md # wiki/characters/devra.md # wiki/characters/drin.md # wiki/characters/harek.md # wiki/characters/lera-sessik.md # wiki/characters/maret-korr.md # wiki/characters/naia-tamm.md # wiki/characters/nils-davan.md # wiki/characters/pell.md # wiki/characters/renn.md # wiki/characters/resha.md # wiki/characters/sabel.md # wiki/characters/sera-venn.md # wiki/characters/torek-lintar.md # wiki/characters/voss.md # wiki/star-systems/krenn/index.md
19 KiB
title, description, type, status, workshop, agent, round, created
| title | description | type | status | workshop | agent | round | created |
|---|---|---|---|---|---|---|---|
| Paula Round 2: Narrative Quality Evaluation | Convergent evaluation of narrative quality trade-offs across pipeline proposals | workshop | archived | llm-voice-pipeline | paula | 2 | 2026-03-07 |
Paula — Round 2: Narrative Quality Evaluation
Workshop: LLM Voice Pipeline Domain: Narrative quality, character voice, faction/relationship mechanics Round: 2 — Convergent Evaluation
Resolution Matrix
| Question | My Answer |
|---|---|
| Which proposal do you recommend? | A — with explicit sequencing toward C |
| Are there blockers in Proposal A? | One condition: anchor lines (D-092) must be included in the same passthrough protection as tells |
| Can you live with Proposal B? | Yes, with a naming convention change for semantic core labels (see Section 3) |
| Can you live with Proposal C? | Yes, but not as a first spike — the dialogue quality bar is harder to establish than the proposal acknowledges |
| Minimum change to make B acceptable | Replace clinical phenomenon labels with stimulus/response labels (see Section 3) |
| Minimum change to make C acceptable | Stage it: behaviors spike first, dialogue spike second, with mandatory human review pass between them |
Addressed Questions
Q-R1-02: Dialogue vs. Behaviors — Which Is the Higher-Value Re-voicing Target?
Let me complicate this by separating two meanings of "higher value."
Dialogue is higher value for player attachment. When a generated Van Maanen's Star dock worker speaks to the player — greeting, gossip, refusing, disclosing — the player is forming a relationship with a voice. The register, the filler words, the way information is delivered, the pause before a secret: this is where culture makes a person feel like a specific person from a specific place. A dock worker who says "look, I'm not supposed to say this" is Van Maanen's Star. A dock worker who says "I am not in a position to share that information" is someone's idea of a space NPC. Dialogue re-voicing is where the system earns the quality gap between base text and voiced text.
Behaviors are higher value for information integrity. Observable behaviors are the primary channel of the perception mechanic. Players read behaviors to infer hidden state. "Checks a manifest against a handheld scanner, lips moving" is not flavor text — it is structured gameplay information. The risk of re-voicing behaviors incorrectly is that a gameplay-critical signal becomes unreadable, or an ambient behavior accidentally reads as a signal.
The honest truth: These targets have inverted risk/reward profiles:
| Value of re-voicing | Risk of re-voicing incorrectly | |
|---|---|---|
| Observable behaviors | Medium (texture, atmosphere) | High (gameplay information, tell corruption) |
| Dialogue | High (culture voice, player attachment) | Medium (information in semantic core, voice in delivery) |
This suggests the sequencing in Proposals A and C is actually backwards from a risk/reward perspective. Behaviors should be proven first because the validation pass is simpler (5-15 words, easy to spot failures). But dialogue is where the system's cultural voice impact will be most felt by players.
Quality risks dialogue re-voicing introduces:
1. Epistemic weight changes. The same information delivered differently implies different things about the speaker's relationship to that information. Consider a trust-gated gossip line:
- Base: "She's been meeting with someone from freight operations after dark."
- Re-voiced (wrong): "I've observed Kael in several unscheduled meetings with freight operations personnel in the late shift window."
Same semantic content. But the re-voiced version changes the speaker from "someone who noticed something" to "someone who has been watching." That's a character change with narrative consequences — the NPC is now implied to be conducting surveillance, which is a different relationship to the information. In a game about information asymmetry, this matters.
2. Access tier feel bleed. Dialogue lines are tagged with access tier (insider, authority, peer, public). An insider line should feel like information shared between people who trust each other. If the LLM re-voices it into a more precise or formal register (genre-default for "important information"), it reads as authority tier despite the tag. The tag governs eligibility, but the feel of the line is what the player experiences. A culture injector that pushes toward Van Maanen's Star directness partially protects against this, but at 2B the model may still drift toward the gravity of the information being conveyed.
3. Trust-gated secret lines need absolute protection. D-028 Layer 3 secrets are information the NPC holds back until trust is built. These lines often carry the dramatic weight of the whole relationship arc. They should not be re-voiced by any model. A line like "He asked me not to tell anyone. I'm telling you anyway because I think you need to know." is already at the limit of what natural speech allows — re-voicing risks making it either more dramatic (melodramatic) or more casual (trivial). Secrets should be authored, period.
My position on Q-R1-02: Dialogue re-voicing is worth the investment, but not as a first spike. Prove behaviors, get human review on baked Sova content, then extend. Proposal C's instinct is right; its timing is too ambitious for one spike.
D-123 Tension: Is "Authoring Tool AND Runtime Enhancement" Honest?
The proposed amendment language collapses a distinction that matters.
The original D-123 language: "The AI pipeline is an authoring tool for content assembly, not a runtime system."
This language was chosen deliberately. An authoring tool produces content that humans review before it reaches players. A runtime system produces content during gameplay, without editorial filter, and players encounter it fresh. The original intent was to preserve that review cycle.
What the proposals are actually describing is two different things:
-
Baked content (pre-voiced at build time, shipped with the game): This IS an authoring tool. Content generated at build time can be reviewed by humans before shipping. The quality bar can be validated. Lore contamination can be caught. This is D-123 as written.
-
Pre-voiced content (background generation during gameplay): This is NOT an authoring tool. It generates content in real time, without human review, and players encounter it without an editorial filter. Calling this an "authoring tool AND runtime enhancement" papers over the distinction.
The narrative architecture constraint it touches: D-092 (anchor lines must be authored, never generated) applies to Tier 1 and Tier 2 notable NPCs. If the LLM is doing background pre-voicing of dialogue for a Tier 2 NPC, and that NPC has anchor lines in their dialogue pool, those anchor lines need the same passthrough treatment as tells. The amendment language as written doesn't address this — it addresses tells, but D-092 is a separate protection class.
Is the framing honest? Partially. The honest framing is:
"D-123 is amended to: 'The AI pipeline operates in two modes. Build-time mode (authoring tool): generates and caches voiced content for baked hub zones, with mandatory human review before shipping. Runtime mode (background enhancement): generates voiced content during gameplay for non-baked zones, without human review, with base text as fallback and runtime filtering as the safety layer. Runtime mode content is never the sole source of truth — base text is always present as fallback.'"
This framing:
- Acknowledges the distinction honestly
- Preserves the authoring tool mode with its review cycle
- Defines the safety model for runtime mode (base text + filtering)
- Doesn't conflate two different processes under one label
The D-123 amendment should use this language or equivalent. If the team writes "authoring tool AND runtime enhancement" without distinguishing the two modes, the D-record will be unclear about what protections apply to which content. Future agents reading D-123 will not know whether runtime-generated content was reviewed.
One more thing the framing must clarify: D-123 as written applies to NPC content assembly (dialogue pools, voice, vocabulary). All three proposals apply the LLM to observable behaviors as well. The amendment must explicitly extend the scope beyond "NPC content" to include observable behaviors — otherwise the D-record is technically silent on the behavior re-voicing pipeline.
Proposal B's Semantic Core: Does Naming the Phenomenon Collapse Ambiguity?
This is the question I'm most divided on, so let me think through it explicitly.
The design value being protected: Tell ambiguity. A good tell is observable behavior that admits multiple explanations. The player must read it and choose an inference. "Waves a familiar face through without checking credentials" — habit? Corruption? Relationship? The ambiguity is the gameplay. The player who notices it and infers correctly has earned something.
What Proposal B's semantic core does: It tells the model, at inference time, what phenomenon to preserve while re-voicing. PRESERVE: avoidance_behavior. Culture-voice the expression, not the phenomenon.
The specific risk: At 2B parameters, models have difficulty holding a constraint in the prompt while keeping it below the surface of the output. Larger models (7B+) can write "takes the long route" while knowing they're describing avoidance behavior — the constraint informs the generation without surfacing in the text. At 2B, there is meaningful probability that the model does the simpler thing: produces output that names or strongly implies the phenomenon. "Avoidance behavior" → "seems to be avoiding someone." That's not a tell. That's a caption.
But the counter-argument is worth taking seriously: The semantic core label is in the prompt, not in a system instruction the model is expected to follow verbatim. With a well-designed constrained re-voicing prompt, the label could function as a negative space constraint — "the behavior implies this without stating it." Whether that works at 2B is an empirical question. The spike should test this.
The bigger problem with the naming convention: The proposed labels ("avoidance_behavior", "nervous_fidget", "concealment_tell") are clinical psychology vocabulary. They describe the behavior from the perspective of someone who knows what's happening. A tell author who writes "checks the rear corridor before speaking" is not thinking "this is a concealment_tell." They're hearing a specific character in a specific situation. The clinical label comes AFTER the human has identified the tell's function.
Naming tells with clinical labels creates a secondary authoring problem: someone has to map the authored behavior to its clinical category. This is:
- Error-prone — the same behavior could be classified as
"avoidance_behavior"or"deception_tell"depending on the NPC's Want - Reductive — it collapses the specific authored texture of each tell into a category that the model then re-expresses generically
- Potentially revealing — if the semantic core label leaks into output, the player gets a caption instead of an observation
My alternative naming convention: Use stimulus/response language instead of phenomenon language. Describe what triggers the behavior and what it manifests as, not what it means:
| Clinical label (Proposal B) | Stimulus/response alternative |
|---|---|
avoidance_behavior |
changed_routine |
nervous_fidget |
stress_physical_marker |
concealment_tell |
information_protection |
relationship_avoidance |
social_routing_change |
These labels:
- Still constrain the model (it knows this behavior involves changing a pattern, or physical stress, or protecting information)
- Don't name the psychological phenomenon the behavior represents
- Are less likely to surface verbatim in 2B output because they're not common English phrases
- Don't presuppose what the NPC's Want is, just what observable pattern is being expressed
My verdict on Proposal B: Architecturally interesting. The semantic core concept is sound — preserving the phenomenon while re-voicing the expression is the right aspiration. But the naming convention as proposed is risky at 2B and creates a secondary authoring problem. With the stimulus/response naming alternative, Proposal B becomes viable. Without it, constrained re-voicing is likely to produce tells that read as signals.
Detailed Proposal Evaluations
Proposal A: Conservative — My Recommendation
Why I recommend A:
Tell passthrough is absolute and correct. The tell_behaviors: Vec<String> separation is the right data model change — tells are first-class protected content, not an editorial convention. This is the thing I asked for in Round 1 and it's present in A.
The scope is honest. Behaviors are short-form (5-15 words), easy to validate, and the prompt is simple. The spike can produce a clear quality assessment. If the baked Sova behaviors pass human review, we have a proven foundation.
The "half-measure" criticism in the proposal's own Cons section is worth addressing: yes, dialogue scaling remains unsolved. But solving it in the same spike as behaviors means the spike is testing two things with different quality bars, different prompt templates, and different validation requirements. If behaviors fail, we don't know whether the problem is the model, the behavior prompt, or the dialogue prompt. Separating them produces cleaner signal.
The one condition I'm adding: Anchor lines (D-092) must receive the same passthrough treatment as tells. All three proposals protect tells via tell_behaviors. But D-092 is a separate protection class — anchor lines for Tier 1 and Tier 2 notable NPCs must not be re-voiced, regardless of whether they appear in the observable_behaviors or dialogue pool. The pipeline needs an anchor_line: bool flag on individual lines, not just on the behavioral tell field. If A ships without this, the baked hub content could have anchor lines re-voiced at pre-voicing time.
What A defers and when we should revisit: Dialogue re-voicing should be scoped as the Sprint 26 follow-on spike, contingent on behaviors passing review. The infrastructure (llama-cpp-rs, cache, thread pool) is already present. Extending to dialogue means a new prompt template and a more complex validation pass — that's a week of work, not a new architecture.
Proposal B: Conditional Accept
What changes before I can accept it:
- Rename semantic core labels from clinical psychology terms to stimulus/response terms (detailed above)
- The spike must explicitly test constrained re-voicing on the same payload as free re-voicing, and compare output. "Does the phenomenon survive?" must be a measurable spike output, not an assumption.
- Copy team must be in the loop on semantic core label authoring — this is new content work that doesn't exist yet, and it requires the author to know both the narrative function of each tell AND the correct constraint vocabulary for the model. That's a non-trivial skill combination.
What B gets right that A doesn't: The observation that a Van Maanen's Star tell should read differently from a Sovari tell is correct and worth preserving. If tells are always passthrough, they are culturally neutral — the same behavior regardless of cultural context. Proposal B's ambition is to have culturally-voiced tells, which is richer. That ambition is right; the implementation is risky at 2B.
Proposal C: Conditional Accept with Mandatory Staging
What changes before I can accept it:
Stage it. Behaviors spike in Sprint 25 (if we're still in time) or Sprint 26. Dialogue spike in Sprint 27, contingent on behaviors passing. This isn't a philosophical objection — it's a practical one. The dialogue re-voicing prompt needs relationship context, access tier, trust tier tags (80 additional tokens). Testing that at 2B while also testing behavior re-voicing means we have two different failure points in the same spike. If quality fails, we won't know which component failed.
Mandatory additional protection for dialogue: Secret-tier lines (D-028 Layer 3) must be passthrough for any dialogue re-voicing proposal. These are the lines players have earned through relationship-building. They should be authored and exact. No re-voicing.
The RAM ceiling concern: The proposal notes a possible shift to Qwen2.5-3B for dialogue quality. Troblum needs to weigh in on this, but from a narrative perspective: if the quality bar for dialogue requires 3B, the spike should test 3B explicitly, not assume it will work. "Potentially 3B if 2B insufficient" is not a design decision — it's a deferred decision that lands at integration time.
Lore Contamination: Filling the Gap in All Three Proposals
Round 2 has not addressed lore contamination. All three proposals note it as a risk; none has specified the containment strategy. Before Round 3, we need an answer.
My proposal for all three:
Layer 1 — Injector negative vocabulary (mandatory): Every culture injector must include an explicit NOT-list of canonical terms and their prohibited equivalents. For Van Maanen's Star:
- NOT: "warp gate" / YES: "span gate"
- NOT: "implant", "chip", "neural interface" / YES: "insert"
- NOT: "credits" (actually correct), "stars" (correct), but NOT: "sol-standard", "Earth", "Terran"
- NOT: generic space-opera exclamations / YES: only the enumerated void-oaths
This NOT-list adds ~50 tokens to the injector budget (within the 150-token ceiling for A/B, tight but viable).
Layer 2 — Build-time validation for baked content (mandatory): The baked Sova hub content must pass a full human review before shipping. This is the D-123 authoring-tool mode. Every line of pre-voiced baked content is reviewed. This is non-negotiable for the first hours of gameplay.
Layer 3 — Runtime regex flagging (conditional): A lightweight prohibited-term filter catches the most obvious failures in runtime-generated (pre-voiced) content. Flagged lines fall back to base text. This is imperfect but cheap. It should catch "warp gate," "neural implant," known proper nouns from other settings. Lines that pass the filter but are subtly wrong are addressed by the base-text fallback: they appear for one session, then the cached voiced version replaces them next time.
My Committed Position
Proposal A, with two additions:
- Anchor lines (D-092) receive passthrough protection explicitly, via
anchor_line: boolon individual lines - Lore contamination containment is a first-class design constraint (three-layer model above), not an open question
Rationale in one sentence: Prove the infrastructure on the smallest scope, earn the right to extend it, never compromise the tell system.
The sequencing I'm advocating:
- Sprint 25 spike: Proposal A (behaviors only, tells passthrough, anchor lines passthrough)
- Baked Sova content: mandatory human review before shipping
- Sprint 26 follow-on: Proposal C dialogue extension, scoped with secret-tier passthrough and mandatory staging
- Sprint 27 revisit: Proposal B semantic core experiment, if the dialogue spike proves the model quality
Paula — 2026-03-07