Files
settled-reach/docs/workshops/llm-voice-pipeline/paula-round1.md
T
jpmschweitzerandClaude Opus 4.6 a301fdab47 fix(content): purge all remaining Krenn references
Replace every occurrence of "Krenn" with "Van Maanen's Star" (or
contextual variants like VMS for locale codes, Van Maanen for proper
noun contexts). Covers CHANGELOG, briefings, workshop docs, sprint
briefings, environmental text examples, templates, ticker content,
and architecture docs.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-16 01:24:04 +01:00

21 KiB
Raw Blame History

title, description, type, status, workshop, agent, round, created
title description type status workshop agent round created
Paula Round 1: Narrative Quality Inventory Narrative quality inventory covering character voice, faction, and relationship mechanics workshop archived llm-voice-pipeline paula 1 2026-03-07

Paula — Round 1: Narrative Quality Inventory

Workshop: LLM Voice Pipeline Domain: Narrative quality, character voice, faction/relationship mechanics Round: 1 — Inventory (divergent)


Preamble: What I Read

I read the full workshop brief, proposed-llm-voice.md, both zone RON files (rural-zone-spec.ron, industrial-zone-spec.ron), culture-van-maanens-star.ron, the generator spike (generator_spike.rs), the NpcBlueprint struct (blueprint.rs), and decisions D-010, D-023, D-024, D-028, D-029, D-034, D-090, D-092, D-121, D-122, D-123, D-124, D-128. Also Q-012 and Q-033.


1. The Three Options — Narrative Quality Assessment

Option 1: Hand-authored pools (current)

What it does well: The zone RON files demonstrate what quality looks like at the top of this approach. The behaviors are complete gestures with cultural specificity embedded:

"wipes grease on the thigh of her coveralls between jobs" "sits in the shade of the gatehouse with a local newsline" "laughs at something a technician says, then catches herself and goes quiet"

These work because they are compositional wholes. The specificity is not decorative — it is the content. "Wipes grease on coveralls" tells you she's manual labor. "Thigh of her coveralls" tells you this is habitual and unself-conscious. "Between jobs" tells you there is no downtime — work is the state she returns to.

What it cannot do: At O(R×Z×C) scale, this approach requires reimagining every behavior from scratch per culture. A second culture's rural mechanic doesn't just use different words — she has a different physical relationship to her tools, a different social relationship to the person she's working for, a different set of gestures that register competence. You can't template that. D-122 (all NPCs generated) combined with any non-trivial number of cultures and zones makes this approach logistically impossible.

Verdict: Not viable at scale. But it establishes the quality floor that everything else is measured against.


Option 2: Composable primitives (Q-057)

What it does well: Nothing that I can see, beyond implementability. And I want to be careful here — I'm not dismissing systems complexity, I'm making a specific claim about narrative texture.

The decomposition problem: The behaviors in the zone RON files work precisely because they resist decomposition. Try it:

"laughs at something a technician says, then catches herself and goes quiet"

What is the action? Laughing. What is the cultural modifier? Catching herself. What is the context tag? Foreman-technician interaction. Now reassemble from components: [laugh_action] + [self_correction_modifier] + [authority_suppression_tag] → "laughs and then stops."

The reassembled version is grammatically correct and semantically equivalent. It is also emotionally empty. The original line works because of "catches herself" — the comma pause, the specificity of the suppression, the choice of "quiet" over "serious" or "professional." These are not modifiers on a verb. They are the verb.

The grammar-to-sentence problem: Composable primitives are a grammar. Grammars produce grammatically valid sentences; they do not produce specifically good ones. The hand-authored behaviors are good because a human looked at a Van Maanen's Star foreman and heard a specific voice. That act of hearing cannot be parameterized.

Verdict: Produces mechanical output. Creates an engine more complex than LLM re-voicing without the quality upside. I'm skeptical this approach can sustain narrative depth. That said — if the spike proves me wrong (some decompositions produce surprisingly specific output), I'd want to revisit.


Option 3: LLM re-voicing

This is the option with the most promise and the most risk, and the two are inseparable.

What the proposal gets right: The i18n analogy is apt. Culture-neutral base text as en-base, culture-voiced text as en-VMS-DIRECT. The injector clause model (10-20 per culture) scales in the right direction. The progressive enhancement framing — base text is functional, voiced text is premium — is elegant and de-risks hardware concerns.

What the proposal is missing: It was written before the generator spike added Want/State, relationship behaviors, and the perception mechanic. These systems change the calculus substantially. See Section 3.

Verdict for narrative quality: Conditionally viable. Viable for Tier 3 ambient and non-tell Tier 2 behaviors. Not viable without explicit protection for semantic load-bearing content. The hybrid is mandatory — not optional — and the protected zones must be a first-class design constraint, not an afterthought.


2. Culture-Specific Vocabulary at 2B Model Size

Let me complicate this with what I actually see in culture-van-maanens-star.ron.

The Van Maanen's Star speech register is not what the proposal assumes

The proposed-llm-voice.md uses this as the Van Maanen's Star injector clause example:

"Your speech is formal and avoids contractions."

This is wrong for Van Maanen's Star. The actual Van Maanen's Star register from culture-van-maanens-star.ron:

  • Register: "direct, minimal pleasantries, gets to the point"
  • Filler words: "look", "right", "yeah", "so", "listen"
  • Greetings: "hey", "shift treating you alright?", "all good?"
  • Farewells: "shift's calling", "gotta move"

Van Maanen's Star is informal, clipped, and working-class. "Formal and avoids contractions" describes Commonwealth institutional culture or perhaps a Sheldon family retainer. It is the opposite of Van Maanen's Star. This error in the example injector is not a minor slip — it reveals that the injector authoring requires actual knowledge of the culture RON, not a generic characterization.

The void-oaths are the hard test

The exclamations ("void take it", "blood and void", "cold vacuum", "void's sake") are the cultural vocabulary most at risk from a 2B model. These phrases exist nowhere in any training corpus. A 2B model instructed to "include Van Maanen's Star cultural exclamations" has two failure modes:

  1. Invents generic space-opera profanity ("stars and void," "by the black," etc.) — readable but not canonical
  2. Produces nothing — defaults to vanilla emotional beats with no exclamations

The solution is enumeration, not instruction. The injector clause cannot say "use void-oaths appropriate to the Van Maanen's Star culture." It must say: "When expressing strong emotion, use ONLY these phrases: void take it, blood and void, cold vacuum, void's sake, damn all, stars." The specific phrases must be injected as a closed vocabulary list, not as a stylistic instruction.

Formality levels within Van Maanen's Star

The RON file captures one formality level (social register). But real cultures have register variation — the same Van Maanen's Star farmer talks differently to their supervisor than to their shift partner than to an outsider. Can injector clauses capture this gradient reliably at 2B?

My assessment: at 2B, probably not reliably. The model can handle one register per culture injector. If we need register variation within a culture (which we will need for relationship-specific dialogue), that variation should be authored at the line level (access tier tags: insider vs authority) rather than asked of the LLM.


3. Re-voicing and the 30/50/20 Tier Model

This is where the Tier 2 boundary becomes load-bearing.

Tier 3 (30% flat wallpaper): Full LLM re-voicing is appropriate

These NPCs carry no semantic load. They are texture. "A dock worker moves freight containers." The base text is functional, and LLM re-voicing can produce cultural flavor without risk. If the LLM slightly mishandles the register, the damage is aesthetically suboptimal, not gameplay-breaking. This is where the pipeline earns its cost.

Tier 2 (50% mundane triangles): Conditional

Tier 2 NPCs carry relationship information that leaks through behavior. A mechanic who borrows tools from a neighbor and returns them without being asked is showing something about her relationship to that neighbor. If the LLM re-voices "borrows a tool from a neighbor and returns it without being asked" into "retrieves equipment from a colleague" — the relationship signal is gone.

The rule I'd propose for Tier 2: Behaviors that contain a named or implied social target (another NPC, a specific relationship) must not be re-voiced. They should be authored. Behaviors that describe an isolated role action (running diagnostics, patching pipe) can be re-voiced.

The practical test: if removing the behavior from context and reading it alone still produces a complete social meaning, it should be protected. "Returns it without being asked" means something about character without any context. "Runs diagnostics on a console" only means something in context.

Tier 1 (20% entangled with intrigue): No LLM re-voicing

Tier 1 NPCs include triangle members and anyone whose behavior is a tell for hidden internal state. These behaviors must be:

  • Authored with precise semantic intent
  • Marked as protected from re-voicing
  • Treated as anchor-line-equivalent per D-092

The FRIEND pattern (D-034) is the extreme case. The FRIEND's observable contradiction — "meeting with unknown contact in restricted corridor" — cannot be re-voiced. Any variation in phrasing changes the information the player receives. Is it "unknown" or "unfamiliar"? Is it "restricted" or "secure"? These words are not stylistic — they encode the player's knowledge state.

The tier boundary: a concrete proposal

Tier Population Re-voicing
Tier 3 ambient 30% Full LLM re-voicing
Tier 2, generic role behaviors ~35% LLM re-voicing with culture injectors
Tier 2, relationship-specific behaviors ~15% Authored or human-reviewed post-generation
Tier 1, non-tell content ~15% Human-reviewed post-generation, not LLM
Tells (all tiers) All NPCs with a Want Protected — never re-voiced
Anchor lines (D-092) Tier 1 and 2 notable NPCs Protected — authored

This is not a clean tier-boundary — it's a behavior-class boundary that applies across tiers. The question "is this a tell?" is more important than the question "what tier is this NPC?"


4. Observable Behaviors vs. Dialogue — What Should the LLM Touch?

The workshop brief asks whether LLM re-voicing should apply to observable behaviors (what you SEE) or dialogue (what NPCs SAY) or both.

The honest truth is these are fundamentally different problems.

Observable behaviors (what you SEE)

Observable behaviors are gameplay information in the perception system. Players read behaviors to infer state. The read→notice→follow→discover sequence (D-027) runs on behaviors. When a player observes a foreman "sits alone in the break room rubbing the back of her neck, datapad face-down on the table" — they are receiving structured information: isolation, stress, concealment.

LLM re-voicing of observable behaviors requires knowing what the behavior means mechanically before deciding whether it can be re-voiced. This is a semantic load problem. The base text "checks credentials at the gate" is safe to re-voice. The base text "waves a familiar face through without checking credentials" is a tell (routine/secret axis) and cannot be re-voiced without potentially losing "without checking" — the specific departure from procedure that makes it an investigative signal.

My position: Observable behaviors should be re-voiced only when:

  1. The behavior is not a tell (not connected to the NPC's Want/State)
  2. The behavior does not name or imply a specific social relationship
  3. The behavior has been reviewed and marked as re-voicing-eligible in the data model

Dialogue (what NPCs SAY)

Dialogue is more appropriate for LLM re-voicing because:

  1. The semantic core (D-028 base line) preserves gameplay-critical information
  2. Cultural voice is the natural value-add (how someone says "you need a keycard" is pure register)
  3. The access tier and trust tier tags already filter what information can be conveyed
  4. Relationship-specific information is handled by the pool selection system, not the individual line

But even here: trust-gated secret lines should be authored. "She changed the subject. Fast." (the monologue beat for a withheld secret) cannot be re-voiced without losing the pause that carries the weight.

My position: Dialogue is the primary candidate for LLM re-voicing. Observable behaviors require a protected-class marker before any re-voicing pass.


5. Preventing Lore-Breaking Content

This is the risk I'd rank highest after tell preservation, because lore contamination is invisible until someone notices it.

What failure looks like

A 2B model instructed to "voice a Van Maanen's Star dock worker" has seen Star Wars, Firefly, Dune, and ten thousand pieces of space opera. It will default toward genre conventions when the injector clauses don't constrain it. Failure modes:

  1. Wrong exclamations: "By the stars," "What in the void" — plausible-sounding but not canonical Van Maanen's Star vocabulary
  2. Wrong social references: References to "the Empire," "the Alliance," "credits" (actually correct) or "sol-standard time" — things that don't exist in the Commonwealth
  3. Wrong technology register: Describing a span gate as a "warp gate" or "jump point," describing inserts as "chips" or "implants" — adjacent to canon but not canon
  4. Wrong socioeconomic register: Treating a dock worker as aspirationally middle-class (genre convention) rather than working-class pragmatic (Van Maanen's Star reality)

The containment strategy

Three layers of prevention:

Layer 1 — Closed vocabulary in injectors: Canonical terms must be injected as closed lists. The injector does not say "use appropriate space-travel terminology." It says: "The following terms are ALWAYS used: insert (neural interface), span gate (interstellar gate), void (space), The Ring (horizon station). NEVER use: warp gate, implant, jump drive, hyperspace, stargate."

Layer 2 — Build-time validation for baked content: The baked hub system content (Sova Transit District) is generated at build time and can be validated by human review before ship. This is the highest-risk content (players' first hours) and should have a full human review pass regardless of pipeline.

Layer 3 — Runtime sampling and flagging: A lightweight rule-based filter (regex against a prohibited terms list) can catch obvious failures before content is served. Flagged lines fall back to base text. This is imperfect but cheap.

What's missing from the proposal: The proposed-llm-voice.md does not mention lore contamination at all. This is a significant gap. The proposal treats the LLM as a stylistic localization engine, but localization engines operate on canonical source text. The LLM has a training distribution that pulls toward genre conventions. Without explicit containment, contamination is not a risk — it is a certainty at volume.


6. What Breaks If We Choose the Wrong Option

If we choose hand-authored only:

  • Immediate: D-122 (all NPCs generated) becomes incompatible with content availability. A generated world full of NPCs with empty behavior pools produces a dead place, not a living one. The generator spike output looks compelling because behaviors exist. Without them, it's a list of names and traits.
  • At scale: The copy team cannot author behaviors for more than 2-3 culture/zone combinations before the sprint budget runs out. The game stalls at Van Maanen's Star/rural + Van Maanen's Star/industrial.
  • What survives: The quality model is still the reference. Even if we move to LLM re-voicing, the hand-authored zone RON behaviors are the gold standard the pipeline is calibrated against.

If we choose composable primitives:

  • Narrative texture collapses: Players stop noticing NPCs. The behaviors become grammatically correct descriptions of actions — "a dock worker loads freight, greets passersby, and monitors the gate." Readable, but not a person.
  • The perception mechanic degrades: If behaviors are assembled from generic components, the signals players use to READ NPCs become harder to distinguish from noise. Tells need to read as specific; composed behaviors are generic by construction.
  • Cross-culture quality drops: The whole point of composable primitives is culture modifier + role action = culture-specific output. But the culture modifier in a composable system is a vocabulary adjustment. "Adjusting vocabulary" is not the same as "sounding like you live in this place." The cultural texture in the Van Maanen's Star RON is not in the vocabulary — it's in the social texture of the behaviors (who you wave through at the gate, how you borrow tools).

If we choose naive LLM re-voicing (no protected zones):

  • Tells become unstable: A tell that was authored as "wipes her hands without making eye contact" might be re-voiced as "quickly cleans up and avoids looking at anyone." Same semantic content, different investigative readability. The player who sees "without making eye contact" has a clue. The player who sees "avoids looking at anyone" has an obvious tell. The ambiguity that makes the perception mechanic rewarding disappears.
  • Relationship behaviors lose specificity: Behaviors that reference a specific social dynamic ("waves a familiar face through without checking credentials") become generic ("allows known workers to pass without verification"). The social information encoded in "familiar face" — the recognition, the implied history — is stripped by the normalization.
  • Replayability is paradoxically hurt: If the LLM re-voices with non-deterministic variance (and without seed-fixed caching it will), the same NPC has different behaviors on Tuesday than Monday. Within a run this is tolerable. Across saves or reloads it's a continuity problem for investigative inference.

7. My One Question Before I Can Commit

Are tells (want-leaking behaviors) a first-class concept in the NpcBlueprint data model, with a field that marks them as protected from re-voicing?

The generator spike has gen_want and gen_want_tell referenced in the workshop brief's "what's new" section, but the current NpcBlueprint struct has only observable_behaviors: Vec<String> — a flat list. There is no semantic distinction between a generic role behavior ("tends crops in the field") and a tell ("checks credentials at the gate, then waves a familiar face through without checking the one behind them").

If tells are not distinguishable from ambient behaviors in the data model, then:

  • Any re-voicing architecture that doesn't know which behaviors are tells will apply the same treatment to both
  • Authors cannot mark tells as protected — there's nowhere to put the flag
  • Build-time validation of baked content has no basis for flagging tell-variants

This is not a blocking question for the hybrid architecture direction — I can commit to LLM re-voicing + protected zones as the right approach. But it is a blocking question for implementation: before the spike, I need to know whether "protected behavior" is a data model feature or an editorial convention enforced by human review. The answer changes what the injector system and cache format need to support.


Summary Position

Question My Answer
Which option best serves narrative quality? Hybrid: LLM re-voicing for Tier 3 and non-sensitive Tier 2, with explicit protected zones for tells, relationship behaviors, and anchor lines
Can injectors capture Van Maanen's Star void-oaths at 2B? Only if injectors enumerate the specific phrases as a closed list, not as stylistic instruction
Tier 2 boundary? Behavior-class boundary, not tier boundary: protected = tells + social-target-naming behaviors; re-voiceable = isolated role actions
SEE or SAY or both? SAY first (more appropriate for cultural re-voicing). SEE only after behavior protection is a data model feature
Lore contamination prevention? Layer 1 closed vocabulary, Layer 2 build-time human review on baked content, Layer 3 runtime regex flagging
What breaks if we choose wrong? Hand-authored: content stalls at Van Maanen's Star. Composable: texture collapses. Naive LLM: tells degrade, perception mechanic loses precision
My committed question Are tells a first-class protected field in NpcBlueprint, or editorial convention only?

The honest summary: The proposal in proposed-llm-voice.md is the right direction, written before the systems that make the direction dangerous were built. The Want/State layer and the perception mechanic changed the calculus. The architecture needs a protected behavior class before narrative quality is safe. But the scale argument is correct and the hybrid approach is viable. I'm not a blocker on this — I'm asking for one data model guarantee.


Paula — 2026-03-07