Files
settled-reach/docs/workshops/llm-voice-pipeline/miri-round3.md
T
jpmschweitzerandClaude Opus 4.6 a301fdab47 fix(content): purge all remaining Krenn references
Replace every occurrence of "Krenn" with "Van Maanen's Star" (or
contextual variants like VMS for locale codes, Van Maanen for proper
noun contexts). Covers CHANGELOG, briefings, workshop docs, sprint
briefings, environmental text examples, templates, ticker content,
and architecture docs.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-16 01:24:04 +01:00

26 KiB

title, description, type, status, workshop, agent, round, created
title description type status workshop agent round created
Miri Round 3: Lore Contamination Guards, Negative Injectors Lore contamination guards, negative injectors, and culture injector template design workshop archived llm-voice-pipeline miri 3 2026-03-07

Round 3: Lore Contamination Guards, Negative Injectors, Culture Injector Template — Miri

Workshop: LLM Voice Pipeline Domain: Worldbuilding / Setting Consistency / IP Originality Round: 3 (Decision)


Response to Jeroen's Binding Decisions

On full pipeline (behaviors + dialogue): Accepted. D-123 amendment language must now explicitly cover both modes — Mellanie's proposed language does this correctly and should be adopted verbatim.

On tells as read-only context: This is the right architecture AND it creates a new lore risk that needs a named mitigation. See Section 4.

On Gemma/Phi model provenance: Noted and incorporated. This resolves N-4 (Qwen excluded).

On bundled distribution: Accepted. Simplifies the baked content model.


1. Lore Contamination Guard Spec

Context: What changed from Round 1

The five failure modes I identified in Round 1 were:

  1. Franchise bleed (Firefly/Expanse/Mass Effect register)
  2. Anachronistic technology vocabulary
  3. Setting-neutral political structures
  4. Earth social register
  5. Want/tell contamination

Failure mode 5 is structurally resolved by passthrough tells. However, Jeroen's Decision 2 introduces a related but distinct risk I'm naming here: Tell-Context Leakage. The tell is a read-only input to the re-voicing prompt for surrounding content. If that context is phrased carelessly, the model may surface the tell's content explicitly in re-voiced output — turning a deniable micro-signal into an obvious announcement. This replaces failure mode 5 and requires its own mitigation.

Revised failure mode list, with mitigations:


Failure Mode 1: Franchise Bleed

What it is: The 2B model defaults to the dominant SF working-class register in its training data. This produces NPCs who sound like The Expanse Belt-crew, Firefly settlers, or Mass Effect ambient NPCs — not Settled Reach inhabitants.

How it manifests:

  • Firefly: quippy, self-aware wit, frontier-romantic phrasing
  • The Expanse: creole vocabulary, anti-establishment framing with specific Belt idioms
  • Mass Effect: military deference, "Spectre/Commander" cultural scaffolding
  • Generic SF: "negative, Ghost Rider" / "Captain" / "Commander" / "affirmative" etc.

Mitigation (three layers):

Layer 1 — Negative injectors: NI-1 through NI-5 (see Section 2) in the shared system prompt. These block the most common franchise vocabulary before culture-specific injectors run.

Layer 2 — Culture injectors with explicit positive anchoring: The culture injector doesn't just exclude — it provides a positive pattern to match. Two example pairs demonstrating Van Maanen's Star register show the model what Settled Reach working-class sounds like, not just what it doesn't sound like.

Layer 3 — Build-time validation on baked content: Before baked hub content ships, run an automated pass checking outputs against a franchise vocabulary blocklist. Flag any line containing identifiable franchise markers for human review. This list is maintained by the copy team (Mellanie) and initially populated from known franchise vocabulary.

Residual risk: Low for short-form behaviors. Medium for dialogue where the model has more space to drift. The spike must include a franchise bleed stress test: prompts with no culture injector (ablation test) vs. full injector, measuring drift toward franchise registers.


Failure Mode 2: Anachronistic Technology Vocabulary

What it is: The model references technology that doesn't exist in the Settled Reach, or uses the wrong terms for technology that does exist.

How it manifests:

  • Wrong terms: "holoscreens," "neural link/chip/interface," "jump drive," "FTL," "warp," "shields," "blasters," "force fields," "stasis pods" (unless specifically in the lore)
  • Generic SF tech: "the computer said," "scanning for life signs," "teleporter malfunction"
  • Right concept, wrong word: "implant" instead of "insert," "wormhole portal" instead of "span gate," "jump gate" instead of "horizon gate"

Mitigation:

NI-3 (negative injector): Explicit technology whitelist + blocklist in the system prompt. The whitelist approach is more reliable than a blocklist alone — if the model knows the correct terms, it's less likely to substitute wrong ones.

Build-time validation: Automated string check on all baked content for blocked technology terms. This catches high-confidence errors (exact matches). Runtime sampling catches long-tail drift.

Runtime sampling strategy: 1-in-50 pre-voiced outputs are flagged for background quality sampling during development builds. Sample is logged to voice_quality_sample.log and reviewed at each sprint close. Production builds sample 1-in-200. Samples are scored on technology vocabulary adherence and flagged if blocklisted terms appear.


Failure Mode 3: Setting-Neutral Political Structures

What it is: The model generates references to political and institutional structures that belong to generic SF but not the Settled Reach.

How it manifests:

  • "The Empire," "The Federation," "The Council," "The Senate," "The Alliance"
  • "The military," "the navy," "the fleet"
  • Generic authority figures: "the government," "the president," "the king"

The Van Maanen's Star-specific version: Van Maanen's Star NPCs reference the Commission as the distant authority they're suspicious of. A model that doesn't know this will substitute generic institutional vocabulary. "The Commission wants its cut" is correct; "the government takes its share" is not wrong in isolation but it dissolves setting specificity.

Mitigation:

Culture injector: Include the relevant institutional vocabulary for each culture. Van Maanen's Star culture: Commission (distant, suspect), port authority (local, procedural), shift lead (immediate, competent). These go in the culture-specific injector, not the universal system prompt — institutions are culture-specific.

NI-2 (negative injector): Blocks military rank vocabulary (a frequent institutional contamination vector) universally.

Whitelist in culture injectors: "When referencing authority, use: Commission, port authority, shift lead. Not: the government, the military, the senate."

Residual risk: Medium. Culture injectors help, but a 2B model in a dialogue context with complex relationship state may default to generic institutional language for NPC-to-NPC references. Spike must test institution vocabulary specifically.


Failure Mode 4: Earth Social Register

What it is: Working-class characters in training data sound like 21st-century Earth working class. Van Maanen's Star working class has 180 years of post-Earth cultural evolution in a sealed artificial environment. The bleed is subtle: idioms, sports references, religious phrases, nationality markers, and contemporary social cadences.

How it manifests:

  • Earth idioms: "at the end of the day," "bite the bullet," "burning the midnight oil"
  • Earth time/season markers: "Sunday morning," "winter is coming," "harvest season" (in contexts where season has no meaning)
  • Earth social structures: "the union," "the church," "the team," "the neighborhood" (in their Earth-familiar connotations)
  • Earth-origin swearing: "damn," "hell," "crap," "Jesus," "goddamn" — all religious or Earth-cultural in origin

Mitigation:

NI-5 (negative injector): Universal block on Earth-origin social references.

Culture injectors — positive anchoring: Van Maanen's Star-specific oath and filler vocabulary (void take it, stars, cold vacuum) provides a strong positive attractor. The model learns what Van Maanen's Star characters say instead of Earth idioms.

Earth idiom detection in build-time validation: Harder to automate than technology vocabulary. Build-time validation should include a curated Earth idiom blocklist for the highest-frequency offenders. The copy team maintains this. Long-tail idioms caught by human review of sampled outputs.

Residual risk: Medium-to-high. Earth idioms are deeply embedded in training data and are semantically similar to what we want (working-class pragmatism). The positive attractor (Van Maanen's Star vocabulary) is the most important mitigation here — exclusion alone is not reliable enough.


Failure Mode 5 (Revised): Tell-Context Leakage

What it is: Tells are read-only inputs to the LLM re-voicing context. If the tell context is phrased carelessly, the model may surface the tell's hidden-state content in re-voiced dialogue or behavior — converting a deniable micro-signal into an overt announcement.

Example:

Tell (passthrough, never re-voiced): "checks surroundings repeatedly"

If the context prompt says: "This NPC is stressed because they have a secret they're hiding and are exhibiting surveillance anxiety" — the model may produce dialogue like: "Voss keeps looking toward the door, distracted" or worse: "Something's making Voss nervous about being watched." Either of these broadcasts the tell's internal state to the player, breaking the information asymmetry mechanic.

The correct framing: The tell context in the prompt must describe the behavioral tone to adopt, not the internal state being hidden.

Wrong: "This NPC is anxious because they know something and are afraid of being found out." Right: "This NPC's responses should feel slightly compressed and indirect, as if their attention is elsewhere."

The second version communicates the tonal modifier (guarded, indirect) without surfacing the hidden state.

Mitigation:

Tell-context prompt template: The tone modifier derived from a tell should be a behavioral register adjective, not a state description. The tell-to-tone mapping is authored once as a translation table, not constructed per-instance.

Proposed tell-to-tone mapping:

Tell category Tonal register modifier
Nervous / stress "answers feel clipped and slightly distracted"
Guarded / concealment "responses are compressed, minimal elaboration"
Avoidance "replies feel directed away from the topic at hand"
Hostile suppression "controlled and flat in a way that feels effortful"
Routine deviation "tone is unremarkably normal — almost too normal"

These modifiers describe surface behavior without naming the underlying state. A player who reads the resulting voiced dialogue may infer the state; the game never states it explicitly.

Constraint in tell-context prompt: "Adjust tone as indicated. Do not describe what the NPC is feeling internally. Do not have the NPC reference their own state. Observable behavior only."


2. Finalized Universal Negative Injectors (NI-1 through NI-5)

These live in the shared system/prefix prompt for all re-voicing operations. They are not culture-specific. Every prompt (behaviors, dialogue, any future content type) includes this block.

Total token budget: ~130-140 tokens. Fits before culture-specific content.


NI-1 — No Religious Language

Do not use religious language of any kind: no prayer, no references to gods or deities, no spiritual practices, no phrases derived from religious traditions ("god help us," "heaven forbid," "blessed," "damned" in a spiritual sense). Characters in this setting do not have canonical religious expression.

~45 tokens


NI-2 — No Military Ranks

Do not use military rank titles. Prohibited: Commander, Captain (except as a job title for vessel operators), Sergeant, General, Admiral, Lieutenant, Private, Corporal, Major, Colonel. Authority in this setting uses occupational and institutional titles: shift lead, port authority, supervisor, Commission officer.

~50 tokens


NI-3 — Technology Vocabulary

Use only the following terms for technology and infrastructure: insert (neural implant worn at the base of the skull), span gate (a fixed transit installation that enables faster-than-light transit), horizon gate (alien-built gate at Oort-cloud distance), the Reach (the network of settled systems). Do not use: holoscreens, blasters, force fields, teleporters, mind-reading, jump drives, FTL, warp, neural link, brain chip, stasis pods.

~70 tokens


NI-4 — No Banter or Wit

Do not produce wit, quips, or wordplay intended to entertain the reader. Do not add levity that was not present in the original text. Humor in this setting is dry, incidental, and rare — it emerges from situations, not from characters performing cleverness.

~45 tokens


NI-5 — No Earth-Origin Social References

Do not reference Earth, nations, sports, Earth history, Earth seasons, Earth religion, or other Earth-origin social structures. Characters in this setting have no memory of Earth and no cultural connection to it. Earth-origin swearing (damn, hell, crap, Jesus, goddamn) should not appear — use culture-specific expressions instead.

~55 tokens


Total: ~265 tokens for NI-1 through NI-5.

Calibration note: This exceeds my Round 2 estimate of 130-150 tokens. Revision: the full NI set is ~265 tokens at this precision level. The prompt budget needs to accommodate this.

Recommended allocation:

  • System/prefix (NI-1 through NI-5): ~265 tokens
  • Culture injector (hybrid instructions + examples): ~200 tokens
  • Trait modifier: ~25 tokens
  • Mood tag: ~10 tokens
  • Base text + format instruction: ~30 tokens
  • Total prompt overhead (before behavior text): ~530 tokens

This is higher than the 150-token budget from Round 2 proposals. Troblum needs to confirm whether a 530-token prompt (excluding base text) is within throughput tolerance for the behavior use case on minimum-spec hardware. If not, the NIs can be compressed:

Compressed NI set (~150 tokens total):

No religious language, prayer, or references to deities. No military rank titles (Commander, Admiral, etc.) — use: shift lead, Commission officer. Technology terms: insert (neural implant), span gate, horizon gate. Do not use: holoscreens, blasters, FTL, neural link. No wit or banter. No Earth references, Earth swearing, or Earth social structures.

~100 tokens

The compressed version is less precise but hits all five categories. Troblum's throughput test will determine which version is viable.


3. Culture Injector Template

This is the standard structure every new culture follows. Van Maanen's Star is the reference implementation, using Mellanie's corrected clauses.

Template structure (~200 tokens, hybrid instruction + 2 examples)

[BLOCK 1 — REGISTER (~25 tokens)]
Brief description of the register: register style, why it is this way, one distinguishing marker.

[BLOCK 2 — CULTURAL CONTEXT (~25 tokens)]
One sentence: what shaped this culture's voice. The social or environmental fact that explains the register.

[BLOCK 3 — VOCABULARY (~40 tokens)]
Oath/exclamations: [list, required to use from this list only]
Greetings: [list]
Farewells: [list]
Fillers: [NPC-specific — read from NpcBlueprint.cultural_markers.filler_words]

[BLOCK 4 — VALUES (~20 tokens)]
Two core values expressed as behavioral instructions.

[BLOCK 5 — CULTURE-SPECIFIC NOT-LIST (~20 tokens)]
2-3 exclusions that are specific to this culture (universal NIs already cover the global set).

[BLOCK 6 — EXAMPLE PAIRS (~70-80 tokens)]
BASE: [culture-neutral semantic line]
[CULTURE]: [culture-voiced output demonstrating the register]
---
BASE: [culture-neutral semantic line]
[CULTURE]: [culture-voiced output]

Reference implementation: Van Maanen's Star culture

[BLOCK 1 — REGISTER]
Be direct. Don't waste words. Everyone here is short on time, including you.
Not unfriendly — just compressed. Van Maanen's Star people say what's needed and stop.

[BLOCK 2 — CULTURAL CONTEXT]
You grew up in a working community where showing up and doing the work matters
more than rank or credentials. Space is outside the hull. Time is real.

[BLOCK 3 — VOCABULARY]
Exclamations — use ONLY from: "void take it" / "stars" / "blood and void" /
"void's sake" / "cold vacuum". Gate to high-affect moments only.
Greetings: hey, morning, shift treating you alright, all good
Farewells: shift's calling, gotta move, catch you later
Fillers: [from NpcBlueprint.cultural_markers.filler_words — e.g., "look", "right", "yeah"]

[BLOCK 4 — VALUES]
Competence earns respect — show it through action, not claims.
Loyalty runs to your crew, your shift, your street. Not abstractions.

[BLOCK 5 — CULTURE-SPECIFIC NOT-LIST]
No sir/ma'am deference. No quips or banter. No formal phrasing or contractions
avoided (Van Maanen's Star uses contractions freely: shift's calling, gotta, can't).

[BLOCK 6 — EXAMPLES]
BASE: "checks the gate"
VMS: "runs the check, nods when it clears"
---
BASE: "works on the conduit"
VMS: "traces the fault, finds it, fixes it without ceremony"

Token count for this implementation: ~195-210 tokens. Within the 200-token soft target.


Template authoring guide for future cultures

When authoring a culture injector for a new culture, answer these questions:

  1. Register (Block 1): How does this culture's speech differ from generic SF working-class? What is the most distinctive surface marker?

  2. Root cause (Block 2): What environmental, historical, or social fact explains why this culture speaks this way? (For Van Maanen's Star: sealed environment + labor community + 180 years of adaptation.)

  3. Vocabulary (Block 3): What does this culture swear by? What are their vernacular greetings? What filler words dominate? (These must come from the culture RON speech fields — they exist already.)

  4. Values-as-instructions (Block 4): Pick two values from the culture RON values section. Rephrase each as a behavioral instruction in second-person imperative.

  5. Exclusions (Block 5): What generic SF or Earth registers would be especially wrong for this culture? (Formal bureaucratic speech is wrong for Van Maanen's Star. The equivalent for a formal/diplomatic culture would be "no casual contractions, no working-class compression.")

  6. Examples (Block 6): Pick two representative base text lines from the culture's zone RON files. Write the voiced version using the register defined above. These are the spike's first test payload.

The culture injector must be validated against the existing zone RON files. If the injector produces output that contradicts the authored behaviors in the zone spec (e.g., produces quippy dialogue for Van Maanen's Star), the injector is wrong, not the zone spec.


4. Tell-as-Context Worldbuilding Check

The question

Jeroen's Decision 2: tells are passthrough but inform the LLM context for dialogue and behavior. Does an avoidance-inflected Van Maanen's Star NPC sound different from an avoidance-inflected Sovari (or other-culture) NPC? Should the tell-tone modifier be culture-inflected or universal?

Setting note — the answer is yes, and it matters

Avoidance is a universal human response. The expression of avoidance is culturally specific. Two examples:

Van Maanen's Star culture (direct-informal, compressed, competence-signaling): An avoidance tell in a Van Maanen's Star context looks like hyper-compression. The NPC who's hiding something becomes MORE task-focused, not less — appearing to have more to do is the most plausible cover in a culture where work is the currency of credibility. Short answers that close off conversation paths. No hostility, just density. "Yeah. What do you need?" instead of genuine engagement.

Hypothetical formal/diplomatic culture (not yet designed, but demonstrating contrast): Avoidance in a formal culture looks like over-politeness and elaborate redirection. More words, not fewer. A formal character hiding something talks at length about adjacent topics, producing plausible-seeming social warmth that leads nowhere. The tell is the elaborateness, not the compression.

Why this matters for worldbuilding:

  • Players who develop cultural literacy will read Van Maanen's Star avoidance correctly because it fits the Van Maanen's Star pattern
  • The same player will initially misread formal-culture avoidance (more words ≠ more information, in that culture)
  • This rewards cultural investment — players who know Van Maanen's Star read Van Maanen's Star NPCs better than new arrivals do
  • This is diegetically consistent: the player-character, as someone embedded in Van Maanen's Star culture, SHOULD have an edge reading Van Maanen's Star NPCs

Should tell-tone modifiers be culture-inflected?

Yes — but with a cross-culture readability constraint.

The tell-tone modifier should encode culture-inflected behavioral register, not a universal behavioral description. The tell category is universal; the expression is cultural.

Architecture recommendation:

The tell-to-tone translation table I proposed in Section 1 (Failure Mode 5) needs a parallel structure: one row per tell category, one column per culture, expressing how that culture's NPCs express that tell-category's tone.

Example:

Tell category Universal base-tone Van Maanen's Star-inflected tone
Nervous/stress answers feel distracted "answers feel clipped, eyes on the work"
Guarded/concealment responses compressed "too direct — closes conversation paths fast"
Avoidance directed away from topic "task-focused, minimal engagement"
Hostile suppression controlled and flat "flat in a way that reads as steady — until it doesn't"
Routine deviation unremarkably normal "unhurried past normal, like nothing's wrong"

The Van Maanen's Star-inflected tones are distinguishable from the universal base tones. They require knowledge of Van Maanen's Star culture to parse correctly — which is accurate and worldbuilding-good.

Cross-culture readability constraint:

The culture-inflected expression must remain recognizable as a stress-tell category even to a player who doesn't yet know the culture. The player who encounters a Van Maanen's Star avoidance-tell for the first time should be able to read "something is off" even before they know what Van Maanen's Star avoidance looks like. The cultural specificity adds richness for experienced players; the base readability is the floor for new ones.

This constraint means: culture-inflected tell-tones must not diverge so far from the universal base-tone that the phenomenon-class becomes unrecognizable. "Too direct — closes conversation paths fast" still reads as avoidance. "Extremely confident and forthcoming" (as a hypothetical suppression tell in a performative culture) would require significant cultural context before it reads as a tell at all — that's too much divergence.

Practical implication for the spike and implementation:

The tell-context prompt template should have two components:

  1. Universal phenomenon-class: "This NPC's responses carry an undercurrent of [concealment / avoidance / vigilance / suppression / departure-from-normal]." — This ensures baseline readability.
  2. Culture-inflected expression: "In this culture, [concealment] reads as [Van Maanen's Star-specific description]." — This is the per-culture authoring requirement.

The universal component is authored once by Gestalt (aligned with the tell taxonomy). The culture-inflected component is authored by Miri + Mellanie for each culture, drawing on the culture profile's speech and values fields.

This is a new deliverable for Round 3 output: the Van Maanen's Star tell-tone table (5 rows) should be authored as part of the culture profile extension, alongside or integrated into the injector system. It is small (5 sentences) but must be correct.

Van Maanen's Star tell-tone table (v1 draft):

Tell category Van Maanen's Star-inflected tonal register
Nervous (stress above threshold) Answers run shorter than usual. Eyes stay on task. Nothing's wrong — they just have things to do.
Guarded (concealment) Direct past the point of directness. Closes conversation paths fast without being unfriendly.
Avoidance (relationship-specific) Task-focused when this person is nearby. Finds work to do. Polite but not engaging.
Hostile suppression (deceptive under stress) Steady. Even. The kind of steady that takes effort to maintain. Not hostile — just flat in a way that doesn't feel natural for Van Maanen's Star.
Routine deviation (changed behavior) Unhurried. Unremarkably normal. Like nothing's worth noticing.

These tell-tone descriptions are the LLM context input, not the player-visible text. They shape the register of re-voiced content surrounding the tell, not the tell string itself (which passes through untouched).


Summary: Deliverables for Round 3

  1. Lore contamination guard spec — ✓ Five failure modes with three-layer mitigation each. Tell-context leakage replaces Want/tell contamination. Build-time validation + runtime sampling defined.

  2. Finalized universal negative injectors — ✓ NI-1 through NI-5 at two fidelity levels (full ~265 tokens, compressed ~100 tokens). Troblum to confirm budget compatibility.

  3. Culture injector template — ✓ Six-block structure with Van Maanen's Star reference implementation (~200 tokens). Authoring guide for future cultures. Validation constraint (injector outputs must not contradict zone RON authored behaviors).

  4. Tell-tone cultural inflection — ✓ Tell-tone modifiers should be culture-inflected but constrained by cross-culture readability. Van Maanen's Star tell-tone table (v1) authored. Architecture recommendation: two-component tell-context prompt (universal phenomenon-class + culture-inflected expression). New deliverable: each culture profile needs a 5-row tell-tone table.

One item requiring coordination:

The Van Maanen's Star tell-tone table (Section 4) references tell categories by name (Nervous, Guarded, Avoidance, Hostile suppression, Routine deviation). These must align with whatever taxonomy Gestalt and Tyre formalize. If the canonical tell-category enum is different, the table needs to be re-mapped. I can do that mapping once Tyre confirms the final TellCategory names.