Files
settled-reach/docs/workshops/llm-voice-pipeline/miri-round1.md
T
jpmschweitzer 23d9ff0a58 Merge remote-tracking branch 'origin/main' into planning
# Conflicts:
#	CHANGELOG.md
#	content/_meta/README.md
#	content/_meta/npc-authoring-style-guide.md
#	wiki/_templates/cultural-group.md
#	wiki/_templates/institution.md
#	wiki/_templates/star-system.md
#	wiki/characters/devra.md
#	wiki/characters/drin.md
#	wiki/characters/harek.md
#	wiki/characters/lera-sessik.md
#	wiki/characters/maret-korr.md
#	wiki/characters/naia-tamm.md
#	wiki/characters/nils-davan.md
#	wiki/characters/pell.md
#	wiki/characters/renn.md
#	wiki/characters/resha.md
#	wiki/characters/sabel.md
#	wiki/characters/sera-venn.md
#	wiki/characters/torek-lintar.md
#	wiki/characters/voss.md
#	wiki/star-systems/krenn/index.md
2026-03-14 00:24:53 +01:00

18 KiB

title, description, type, status, workshop, agent, round, created
title description type status workshop agent round created
Miri Round 1: World Consistency Inventory World consistency inventory assessing lore contamination and IP originality risks workshop archived llm-voice-pipeline miri 1 2026-03-07

Round 1: World Consistency Inventory — Miri

Workshop: LLM Voice Pipeline Domain: Worldbuilding / Setting Consistency / IP Originality Round: 1 (Inventory)


Opening position

Let me check this against what we've established before I endorse anything.

I've read the full proposal, the generator spike, both zone specs, the Van Maanen's Star culture profile, and the relevant D-records (D-121, D-122, D-123, D-128). My position: Option 3 (LLM re-voicing) is correct in direction, but the proposal as written underestimates the injection complexity required to preserve cultural distinctiveness at 2B model size. The base-text-as-fallback architecture is worldbuilding-sound. The injector clause system, as currently described, is not deep enough to produce Van Maanen's Star voices rather than generic SF working-class voices.

This is fixable. It is not a reason to reject Option 3. But it needs to be flagged clearly before the spike is designed.


1. Which option best preserves world consistency and cultural distinctiveness?

Option 1 — Hand-authored pools (current)

Setting note: This is the highest-fidelity option for Van Maanen's Star specifically, but it encodes a trap. We've invested enough authoring to make Van Maanen's Star feel like a place. Adding a second culture — say, a station culture with different history and different relationship to the void — requires rebuilding the entire content layer from scratch. Option 1 preserves what we have; it cannot scale to what the Reach requires.

The cultural distinctiveness of Van Maanen's Star in the current zone files is not accidental. Lines like "wipes grease on the thigh of her coveralls between jobs" and "explains a repair in clipped shorthand without looking up" are specific and earned. That specificity comes from Mellanie understanding Van Maanen's Star culture well enough to author from the inside. You cannot template that away. What you can do is provide the LLM with enough cultural context that it produces output that doesn't contradict it.

Verdict: Best quality, worst scalability. Viable only for Van Maanen's Star, only for v0.2.

Option 2 — Composable primitives

Setting note: The cultural markers system already exists in the NpcBlueprint — speech_register, filler_words, greeting. These are well-defined discrete items. The problem is that composable primitives can only assemble vocabulary; they cannot assemble worldview.

"Void take it" is in the culture RON as an exclamation. A composable system can insert it correctly when an NPC exclaims. But it cannot decide that a Van Maanen's Star character, when stressed, says "cold vacuum" rather than "void take it" — that requires understanding the emotional register each phrase carries. Van Maanen's Star speech is working-class and compressed, not working-class and verbose. Composition engines tend toward additive assembly; Van Maanen's Star culture requires compression and omission.

More seriously: composable primitives cannot prevent the cultural void-oath from being inserted in a context where it reads wrong. The behavior "hauls produce to the market stall before the morning exchange opens" doesn't naturaly carry an exclamation — but a template that tries to add cultural flavor might generate something like "hauls produce, grumbling 'void take it' at the weight." That's not wrong vocabulary. It's wrong register.

Verdict: Sufficient for vocabulary, insufficient for worldview. Produces Van Maanen's Star-vocabulary characters that don't feel Van Maanen's Star.

Setting note: The architecture of this option is sound worldbuilding. The semantic base text as the gameplay layer and the voiced text as the enhancement layer maps cleanly to how the setting works diegetically — the world is always legible; the insert just adds resolution. A player who plays without AI enhancement experiences a functional Van Maanen's Star world; one with it enabled hears the grain.

The injector clause structure — culture as baseline, personality as flavor, mood as override — matches the cultural hierarchy we've established (D-121: voice is culture-driven, job as modifier). This is not a coincidence; it's a correct abstraction of what the existing culture RON encodes.

My concern is specifically about 2B-class model capability. See Section 2.

Verdict: Correct direction. Quality ceiling depends on injector depth and model instruction-following capability.


2. Can injector clauses preserve culture-specific vocabulary at 2B model size?

This is where I need to be cautious.

The proposal describes the Van Maanen's Star cultural injector as: "Your speech is formal and avoids contractions."

That is the wrong injector for Van Maanen's Star. Van Maanen's Star speech is not formal. It is direct-informal. Formal-without-contractions describes a completely different culture. This example injector reads like a placeholder written for a generic "culture adds formality" slot. If this is the actual injector that ships, we will produce NPCs who sound like junior civil servants, not people who live in a pressurized box 180 years from Earth.

The actual Van Maanen's Star cultural injector needs to encode:

  1. Register: Direct but not hostile. Short because time is genuinely scarce, not because they're unfriendly.
  2. Oath vocabulary: Void-adjacent exclamations only. "Void take it," "cold vacuum," "blood and void." Not divine oaths. Not secular Earth oaths ("damn it," "hell," "crap"). Space is the threat that kills you, not a metaphysical abstraction.
  3. Community anchors: Crew, shift, and street as emotional reference points. Not family in the traditional sense. Not institution. The people you'd bleed for are the people on your shift.
  4. Competence signaling: Respect is earned through doing the work. Characters signal this through precision of observation and action, not through status talk.
  5. Negative space: What NOT to say. No quips. No banter-for-banter's-sake. No Firefly register. No "sir"/"ma'am" deference culture. No references to political institutions by name.

This is approximately 200-300 words of instruction. At 2B model size, the effective instruction-following window for stylistic constraints is uncertain. Small models are known to:

  • Anchor to the most-represented working-class register in training data (which is contemporary American/British English)
  • Treat unfamiliar vocabulary ("void take it") as errors and smooth them to standard alternatives
  • Flatten cultural subtlety under pressure from the base text's neutral English

My assessment: A 2B model can probably preserve oath vocabulary if the injector explicitly lists the terms and instructs their use. It cannot reliably preserve the philosophy behind the vocabulary. The difference between a character who says "void take it" because they were told to and a character who says it because space genuinely terrifies them — that lives in tone and context, not in word selection.

Practical floor: Injector clauses can guarantee correct oath vocabulary and correct register description. They cannot guarantee that the model uses them with correct Van Maanen's Star weight. The spike must explicitly test oath preservation and register accuracy, not just fluency.


3. How do we prevent the LLM from introducing lore-breaking content?

Setting note — I need to enumerate what can actually go wrong here, because "lore contamination" is too vague to design against.

Type A: Franchise bleed

At 2B, the model's working-class SF character register draws heavily from training data: Firefly, The Expanse, Babylon 5, Mass Effect ambient NPCs. These feel like the Settled Reach superficially (space, working class, pragmatic) but are not it. Indicators:

  • Firefly register: "Shiny," quippy banter, frontier-town affect
  • The Expanse register: Belt creole vocabulary, anti-inner-planets resentment framing
  • Mass Effect register: Military protocol, "Commander/Spectre" deference vocabulary

Guard: Negative injectors. The cultural injector should include explicit NOT-lists: "Do not use military rank terms. Do not use contractions as markers of informality. Do not produce quips or banter." This is unusual prompting but necessary at small model sizes.

Type B: Anachronistic technology

The model knows what generic SF NPCs talk about. Wormholes are in the Settled Reach vocabulary — good. Holoscreens, jump drives, FTL ships, blasters — not in the Settled Reach. "Insert" is the correct term for neural implants; the model may substitute "neural link," "implant," "chip," "interface." The span gate is the correct term; the model may produce "wormhole portal," "jump gate," "stargate."

Guard: Terminology whitelist in injectors. This is a short list: insert, span gate, horizon gate, void, the Reach. Instruction: "Only use the following terms for technology and infrastructure: [list]." This must be verified explicitly in the spike.

Type C: Setting-neutral social structures

The model may produce NPCs who reference senators, admirals, corporations, megacities — structures that exist in generic SF but not in the Settled Reach's specific institutional topology. For Van Maanen's Star, the relevant institutions are the Commission (distant authority, suspect), the shift structure (immediate authority, respected if competent), and the local community (primary loyalty).

Guard: Injector should specify institutional vocabulary. "When referencing authority, use: Commission, shift lead, port authority. Do not use: government, military, senate, council, corporation."

Type D: Social register bleed

Working-class characters in English-language training data sound like contemporary Earth working class. Van Maanen's Star working class has 180 years of post-Earth cultural evolution in an enclosed artificial environment. The biggest surface tell is: contemporary Earth profanity and social reference. A Van Maanen's Star character should not reference sports, religion, nationalism, or other Earth-rooted social fabric. The model will produce these because they are statistically dominant in training data for working-class dialogue.

Guard: Explicit exclusion in injectors. "Do not reference religion, sports, nationality, or Earth-origin social structures."

Type E: Want/Tell contamination

This is the most dangerous type. The Want/tell system generates deliberately ambiguous behavioral signals — the player is supposed to read them, not have them explained. If the LLM re-voices a Want tell, it might either: (a) neutralize the ambiguity into a flat description, or (b) over-explain it into an obvious broadcast.

Base text: "checks the vault door twice before walking away" Bad re-voice A (neutralized): "walks past the vault door" — tell removed entirely Bad re-voice B (over-explained): "lingers nervously near the vault door, clearly worried about something inside" — tell made too explicit

Guard: Want tells must be in the protected-content category. They are not candidates for re-voicing. The base text for a Want tell IS the player-facing text. This needs to be a hard architectural boundary.


4. How does re-voicing interact with the cultural markers system?

The current CulturalMarkers struct carries:

  • speech_register (a string)
  • filler_words (a vec of strings)
  • greeting (a string)

These are already discrete, enumerable, culture-authored items. They are the output of the generator, not the LLM injector input. This creates a possible alignment problem.

If the LLM injector says "use filler word 'look'" but the NPC's generated CulturalMarkers.filler_words contains ["right", "yeah"] — which governs? The struct was built from the culture RON with randomness applied. The injector is built from the culture RON directly.

More importantly: the cultural markers system is already doing what Option 3 proposes, for vocabulary. It is assigning culture-specific vocabulary to individual NPCs. The LLM injector would be a second layer doing the same thing at the prose level.

My recommendation: The cultural markers struct should be the source of truth for the LLM injector's per-NPC vocabulary. When constructing the injector prompt, pull filler_words, greeting, and speech_register from the NPC's generated blueprint, not from the culture RON directly. This ensures the voiced output is consistent with what the blueprint already specifies, and avoids the dual-source problem.

This also means the LLM injector for vocabulary is zero-additional-authoring — it reads from the already-generated NpcBlueprint.


5. What breaks if we choose the wrong option?

If we choose Option 1 (hand-authored only)

Setting cost: Van Maanen's Star is permanently the only culture with full coverage. Every other culture the team needs — and the Reach requires multiple cultures for the investigation mechanics to work — starts from nothing. The IP originality problem is managed by authoring, but only for Van Maanen's Star. Everywhere else defaults to generic SF.

More importantly: D-128 says culture is implicit in location. As we add locations, we add culture requirements. Option 1 makes every new location a content crisis.

If we choose Option 2 (composable primitives, no LLM)

Setting cost: Characters produce the right vocabulary in the wrong contexts. A composable system that assembles "void take it" as a cultural marker will insert it wherever the culture modifier fires, regardless of whether the character is mildly inconvenienced or confronting existential danger. Van Maanen's Star exclamations are calibrated by severity — "stars" is mild, "blood and void" is serious. Template assembly has no severity model.

More practically: composable primitives produce dialogue that reads as assembled. The player will notice the seams. The immersive sim depends on NPCs feeling like inhabitants. Assembled dialogue breaks that.

If we choose Option 3 poorly (LLM with shallow injectors)

Setting cost: Franchise bleed at scale. Every NPC sounds vaguely like a Space Western/Military SF character. The Settled Reach stops feeling like its own place and starts feeling like a mashup of recognizable genre influences. This is the IP originality failure mode — not copyright infringement, but creative dissolution. If a reader could point at any random NPC and say "that's The Expanse," we've failed.

The secondary failure: void-oaths become decorative. If the model uses "void take it" and "damn it" interchangeably based on training data frequency, the oath stops carrying worldbuilding weight. The player stops reading it as a setting signal.

If we choose Option 3 correctly (LLM with deep injectors, protected tells)

Residual risk: The seam between base-text and voiced-text may be perceptible when the player first encounters a slow-generated NPC. From a worldbuilding perspective, this is survivable — the base text is designed to be legible, not broken. But the transition needs to be invisible. If a player sees the base text and the voiced text in close succession (e.g., first visit vs. return visit after pre-voicing completes), the delta in quality might draw attention to the system rather than the world.


6. One question I need answered before I can commit

What is the effective token budget for cultural injector clauses in the final prompt construction?

This is the binding constraint for everything I've described. If the total prompt is structured as:

[Task instruction] + [Base text] + [Mood injector] + [Personality injectors] + [Cultural injector] + [Format instruction]

...then the cultural injector is competing for space with everything else. At 2B model size, very long prompts produce worse instruction-following, not better. The cultural injector I described above — register, oath vocabulary, community anchors, competence signaling, exclusions — requires approximately 200-300 words to encode Van Maanen's Star accurately. If the budget is 50-80 tokens, we can specify register and list the oaths but nothing else. If it's 200+ tokens, we can encode the cultural philosophy.

The quality ceiling of culture preservation in this system is directly determined by prompt budget. I cannot assess whether Option 3 can preserve Van Maanen's Star cultural distinctiveness until Tyre tells me how many tokens the cultural injector can consume without degrading output quality at 2B model size.

If the answer is "under 100 tokens," we need to revisit the injector architecture and consider culture-specific few-shot examples rather than instruction-only injectors. Few-shot examples may produce better Van Maanen's Star register than instructions about Van Maanen's Star register — but they consume more tokens and require authoring examples for each culture.


Summary position

Criterion Option 1 Option 2 Option 3 (hybrid)
Van Maanen's Star cultural distinctiveness Excellent Adequate Good, if injectors are deep
Multi-culture scalability Poor Moderate Excellent
Void-oath preservation Guaranteed Vocabulary only Depends on model + budget
Lore contamination risk None Low Moderate (franchise bleed)
Want/tell protection Guaranteed Guaranteed Requires explicit protection
IP originality Guaranteed Guaranteed Requires negative injectors

Recommended path: Option 3 (LLM re-voicing) with:

  1. Want tells and relationship-specific behaviors as explicitly protected, non-re-voiced content
  2. Cultural injectors sourced from NpcBlueprint.cultural_markers (not re-derived from culture RON)
  3. Negative injectors (NOT-lists) as a first-class component of cultural injection
  4. Spike must explicitly test oath preservation and franchise-bleed resistance, not just fluency
  5. The current placeholder Van Maanen's Star injector ("formal, avoids contractions") must be replaced before any quality assessment is valid

The base-text-as-fallback architecture is correct and worldbuilding-sound. The progressive enhancement model maps cleanly to how the setting works. My only blocker is knowing the prompt token budget before I can assess whether deep injectors are feasible at 2B.