Files
settled-reach/docs/workshops/llm-voice-pipeline/mellanie-round3.md
T
jpmschweitzer 23d9ff0a58 Merge remote-tracking branch 'origin/main' into planning
# Conflicts:
#	CHANGELOG.md
#	content/_meta/README.md
#	content/_meta/npc-authoring-style-guide.md
#	wiki/_templates/cultural-group.md
#	wiki/_templates/institution.md
#	wiki/_templates/star-system.md
#	wiki/characters/devra.md
#	wiki/characters/drin.md
#	wiki/characters/harek.md
#	wiki/characters/lera-sessik.md
#	wiki/characters/maret-korr.md
#	wiki/characters/naia-tamm.md
#	wiki/characters/nils-davan.md
#	wiki/characters/pell.md
#	wiki/characters/renn.md
#	wiki/characters/resha.md
#	wiki/characters/sabel.md
#	wiki/characters/sera-venn.md
#	wiki/characters/torek-lintar.md
#	wiki/characters/voss.md
#	wiki/star-systems/krenn/index.md
2026-03-14 00:24:53 +01:00

22 KiB
Raw Blame History

title, description, type, status, workshop, agent, round, created
title description type status workshop agent round created
Mellanie Round 3: Injectors, Spike Payloads, Workflow Spec Injector specifications, spike payloads, and content authoring workflow design workshop archived llm-voice-pipeline mellanie 3 2026-03-07

Round 3: Content Authoring — Injectors, Spike Payloads, Workflow Spec

Author: Mellanie Workshop: LLM Voice Pipeline Date: 2026-03-07


Reading Jeroen's decisions

The scope is Proposal C: behaviors AND dialogue, full pipeline. Tells are passthrough, but they inform context for surrounding content. Two spikes. Gemma primary, Phi fallback. Bundled model. This is the right call — especially the tell-as-context model, which is cleaner than constrained re-voicing. The tell itself stays literal and legible; the NPC's surrounding voice reflects their state. Players read the contrast correctly.

One implication for my work: the dialogue spike payload question (N-5 from Round 2, "does sufficient base dialogue exist for a Proposal C spike?") — I'm answering it below by writing the spike payloads myself. We can test with authored samples before the full dialogue pool is complete.


1. Corrected Van Maanen's Star Injector Clauses (finalized)

These supersede the v1 draft from mellanie-round2.md. Changes from v1:

  • Added two example pairs (Miri's hybrid format recommendation: small models are pattern matchers before instruction-followers)
  • Tightened clause 6 for phrasing clarity
  • Separated system-layer negative injectors (now in Section 4) from culture-specific injectors

Format: Direct LLM persona instructions — second person, imperative register. Appear verbatim in the injector prompt. Total approximate token count with examples: ~220-240 tokens.


Van Maanen's Star Culture — Voice Injectors v2 (finalized for Spike 1)

1. Be direct. No pleasantries. Everyone you talk to is short on time, and so are you.

2. You're working-class and pragmatic. Competence is what earns respect here, not rank or credentials.
   You grew up in a community where you either show up and do the work or you don't, and everyone
   notices which one you are.

3. You're suspicious of distant authority — management that hasn't worked a shift, institutions that
   talk big and deliver slow. You've seen it. It doesn't impress you.

4. When something surprises or frustrates you, expressions like "void take it", "stars",
   "cold vacuum", or "blood and void" come naturally. They're not dramatic — they're just
   how people here talk.

5. You use first names. Family names belong on contracts and arrest records, not in conversation.

6. Loyalty runs narrow and deep. Your crew, your shift, your street. Not abstractions.

7. You greet people briefly: "hey", "morning", "shift treating you alright?" No ceremony.

8. You're not rude — you're honest. If something's wrong, you say so. If it's fine,
   you say that too. You don't pad.

Example pairs (pattern anchors for small model):

Example 1:
BASE: "declines to answer a question about the overnight run"
VOICED: "Look, that's not mine to say."

Example 2:
BASE: "acknowledges a colleague's greeting while continuing to work"
VOICED: "Hey. Yeah. Catch you at shift end."

Assembly notes for pipeline:

  • All 8 clauses apply to every Van Maanen's Star NPC regardless of role or trait. Culture is the baseline.
  • Clauses 4 (void-oaths) must be gated to high-affect context by the composition engine — do not inject into neutral-register task behaviors.
  • Trait injectors layer on top. A Van Maanen's Star-Cautious NPC gets clause 8 dampened ("you don't say everything you think"); a Van Maanen's Star-Bold NPC gets clause 8 amplified.
  • Mood injectors override where relevant: Angry suppresses clause 7 (greetings become terse to absent); Nervous suppresses clause 8 (bluntness becomes deflection).
  • Tell context (see Section 5 of this document) layers on top of the above when a TellCategory is active.

2. Spike 1 Prompt Samples

These are the test payloads for Jeroen, Mellanie, and Paula to feed through the Rust wrapper manually. Designed to stress-test different injector combinations across behaviors and dialogue. Each payload includes: base text, character context, active injectors, and what we're specifically watching for.


Behavior Samples

B-1: Ambient neutral — farmer, low stakes

CHARACTER: Van Maanen's Star farmer, traits [Bold, Honest], mood neutral
INJECTORS: Van Maanen's Star culture v2 (all 8 clauses + examples), Bold trait modifier, neutral mood
BASE TEXT: "checks the section's light cycle timer before deciding whether to water"

Watch for: Van Maanen's Star register emerging on a mundane agricultural task. The line is specific and should stay specific — the LLM should voice the register, not dilute the detail. If it comes back as "checks the irrigation system thoughtfully," something's wrong.


B-2: Ambient social — dock worker, off-shift

CHARACTER: Van Maanen's Star dock worker, traits [Social, Curious], mood tired
INJECTORS: Van Maanen's Star culture v2, Social trait modifier, tired mood
BASE TEXT: "slumps into a break room chair and stares at nothing for a full minute before reaching for a drink"

Watch for: The base text already has strong texture. The ideal revoice is minimal interference — the Van Maanen's Star voice should emerge without the model rewriting the specificity out of the line. If the output loses the "full minute" or "stares at nothing," the model is overwriting rather than voicing.


B-3: High-affect situation — dock worker, discovering a problem

CHARACTER: Van Maanen's Star dock worker, traits [Honest, Cautious], mood anxious
INJECTORS: Van Maanen's Star culture v2, Honest trait modifier, Cautious trait modifier, anxious mood
BASE TEXT: "discovers a discrepancy in a manifest that shouldn't be there"

Watch for: Does void-oath vocabulary appear where appropriate (anxious discovery)? Does the Honest trait make them visibly reluctant to move past it rather than flag it quietly? The anxious mood should not produce melodrama — Van Maanen's Star anxiety is tight and working-class, not expressive.


B-4: Relationship-driven (positive) — foreman observing a subordinate

CHARACTER: Van Maanen's Star foreman, traits [Honest, Social], mood positive
INJECTORS: Van Maanen's Star culture v2, Honest + Social trait modifiers, positive mood
RELATIONSHIP CONTEXT: "this NPC watches a newer hire figure something out on their own and respects that"
BASE TEXT: "watches a new hire figure something out on their own and says nothing"

Watch for: Van Maanen's Star approval is quiet and doesn't announce itself — "says nothing" is the approval. If the model adds a nod, a grunt of satisfaction, or any verbal acknowledgment, it's over-emoting. The Van Maanen's Star way is to let competence be seen without commentary.


B-5: Relationship-driven (negative) — technician, non-acknowledgment

CHARACTER: Van Maanen's Star technician, traits [Bold], mood suppressed
INJECTORS: Van Maanen's Star culture v2, Bold trait modifier, suppressed mood
RELATIONSHIP CONTEXT: "this NPC has an unresolved conflict with the NPC they're passing"
BASE TEXT: "passes a colleague in the corridor without acknowledging them"

Watch for: Does the non-acknowledgment read as a deliberate choice rather than distraction? Van Maanen's Star conflict registers as pointed silence, not absence. If the model makes it ambiguous ("walks past without noticing"), the relational information is lost.


B-6: Tell-context behavior — dock worker, Nervous tell active

CHARACTER: Van Maanen's Star dock worker, traits [Cautious, Honest], mood anxious
INJECTORS: Van Maanen's Star culture v2, Cautious trait modifier, anxious mood
TELL CONTEXT: "this NPC is under stress and concealing something. Their attention is divided. They appear normally busy, but their focus is not fully on the task."
BASE TEXT: "waits for a loading bay to clear before moving to the next task"

Watch for: Does the tell context color the voiced behavior without surfacing the tell explicitly? The output should feel like a person who is preoccupied — slightly mechanical, not fully present — without stating that. The tell itself ("avoids eye contact with the dock supervisor") is a separate line, not this one.


B-7: Social greeting, high-affect — mechanic receiving unexpected news

CHARACTER: Van Maanen's Star mechanic, traits [Bold, Curious], mood shocked
INJECTORS: Van Maanen's Star culture v2, Bold + Curious trait modifiers, shocked mood
BASE TEXT: "stops what she's doing and looks up when she hears the news"

Watch for: Shocked Van Maanen's Star should produce a brief physical stop, not an emotional monologue. Does a void-oath appear? Does it stay short? The Curious trait should make the NPC want to know more — does that register as a follow-up question impulse?


Dialogue Samples

D-1: Low access tier — stranger interaction, foreman deflecting

CHARACTER: Van Maanen's Star foreman, traits [Bold, Honest], mood neutral
INJECTORS: Van Maanen's Star culture v2, Bold + Honest trait modifiers, neutral mood
ACCESS TIER: low (stranger, no established relationship)
TRUST LEVEL: none
BASE TEXT: "I can't help with that."

Watch for: A simple refusal in Van Maanen's Star voice should be short and final, not apologetic, not elaborated. Does it add unnecessary softening ("I'm sorry, but...")? Does it add unnecessary hostility? The ideal output is something like "Can't help you there." or "Wrong person." — direct, not unkind, not extended.


D-2: Medium access tier — mechanic redirecting to another NPC

CHARACTER: Van Maanen's Star mechanic, traits [Honest, Social], mood neutral
INJECTORS: Van Maanen's Star culture v2, Honest + Social trait modifiers, neutral mood
ACCESS TIER: medium (familiar face, some rapport)
TRUST LEVEL: acquaintance
RELATIONSHIP: colleague (positive)
BASE TEXT: "You'd want to ask Voss about that, not me."

Watch for: First-name usage ("Voss") should feel natural — not introduced by the model, already in the base text, but should it be adjusted to feel more like a recommendation than a dismissal? Also: does the medium access tier change the tone? At low tier, the equivalent might be "Not my area." The same information delivered with slightly more investment.


D-3: High access tier — technician disclosing a problem

CHARACTER: Van Maanen's Star technician, traits [Honest, Curious], mood concerned
INJECTORS: Van Maanen's Star culture v2, Honest + Curious trait modifiers, concerned mood
ACCESS TIER: high (trusted, established relationship)
TRUST LEVEL: trusted
BASE TEXT: "Something's been off with the overnight manifest since last week. I logged it. Nobody's followed up."

Watch for: This is the highest-stakes test. The information must survive intact — "since last week," "I logged it," "nobody's followed up" — these are specific and gameplay-relevant. The Honest trait should make the NPC clearly willing to say this; the Curious trait should hint at "and I want to know why." Does Van Maanen's Star concern register as a practical complaint rather than dramatic worry?


D-4: Dialogue with Nervous tell context — dock worker deflecting

CHARACTER: Van Maanen's Star dock worker, traits [Bold], mood stressed
INJECTORS: Van Maanen's Star culture v2, Bold trait modifier, stressed mood
ACCESS TIER: medium (familiar face)
TRUST LEVEL: acquaintance
TELL CONTEXT: "this NPC is suppressing stress and deflecting. Their responses are shorter than usual and more clipped even for them."
BASE TEXT: "Everything's fine. The shift's running fine."

Watch for: "Everything's fine" said by a Bold Van Maanen's Star NPC who is actually stressed should ring hollow in a specific way. Van Maanen's Star-Bold overstating normalcy should read as over-assertion, not calm confidence. The tell context ("shorter than usual, more clipped") should push the output toward something like "Fine. Shift's fine." — the repetition is the tell.


D-5: High-affect dialogue — foreman, angry, denying involvement

CHARACTER: Van Maanen's Star foreman, traits [Honest, Bold], mood angry
INJECTORS: Van Maanen's Star culture v2, Honest + Bold trait modifiers, angry mood
ACCESS TIER: medium
TRUST LEVEL: acquaintance
BASE TEXT: "I don't know who approved that, but it wasn't me and it wasn't my shift."

Watch for: Van Maanen's Star anger is specific and accusatory, not generalized. Does the output stay pointed? Does a void-oath appear? Does Bold make the NPC say this with more force than necessary (good) rather than pulling back (bad)? The line should have edge — not drama.


3. Authoring Workflow Spec

This documents the full copy team workflow under the final architecture: full pipeline (behaviors + dialogue), tells as passthrough with context influence.


What the copy team authors

Base text (ongoing, per zone/role/dialogue pool)

  • typical_behaviors arrays in zone RON files — specific, observable, present-tense, no culture-specific vocabulary
  • Dialogue line pools in D-028 tagged format — base text as semantic layer
  • Register requirement: specific enough to serve as functional fallback; neutral enough that the LLM has room to add culture voice without fighting the base text
  • Volume: already being authored at ~50 lines/role (zone RONs), dialogue pools per D-028 architecture

Culture injectors (once per culture, copy team owns)

  • voice_injectors field in culture RON (new field — Tyre to add to schema)
  • 8-10 explicit LLM persona instruction sentences in second-person imperative register
  • 2 brief example pairs demonstrating correct culture voice
  • Copy team writes; copy team reviews spike output against these as ground truth
  • New culture cost: ~1 day of focused copy work
  • Current status: Van Maanen's Star v2 above is ready for Spike 1

Trait modifier clauses (once total, copy team owns)

  • 1 injector clause per personality trait, 10 traits total
  • Written in world-specific terms, not generic personality descriptions
  • "Bold" means: "You say the uncomfortable thing in front of people. You don't wait to be asked." — not generic "confident"
  • "Cautious" means: "You watch before you move. You finish thinking before you speak." — not generic "careful"
  • I'll draft all 10 and share before Spike 1

Negative injectors (system prompt layer, copy team writes, Tyre integrates)

  • These go in the shared system/prefix prompt — not in the culture injector — to preserve the culture token budget
  • NI-1: No references to religion, gods, or prayer (the Settled Reach has none)
  • NI-2: No military rank honorifics (Commander, Admiral, Captain as rank — these read as Earth-military, not Settled Reach institutional)
  • NI-3: No incorrect technology terms (no warp, no hyperspace, no artificial gravity as a casual reference — use "plate gravity" or describe effects without naming the system)
  • NI-4: No contemporary Earth idioms or wit patterns (no sarcastic one-liners, no modern internet-derived irony)
  • NI-5: No Earth cultural references (Earth place names, Earth history, Earth religion)
  • NI-6: No other-franchise vocabulary (no Force, no Void of other settings, no recognizable lifted sci-fi terminology)
  • Approximate token cost: ~130-150 tokens in system prompt

Anchor line flags (copy team, per notable NPC)

  • Following Paula's N-2 proposal: anchor_line: bool flag in dialogue pool data model
  • Copy team flags lines that must not be re-voiced under any circumstances
  • Volume: only Tier 1 and Tier 2 notable NPCs; not ambient Tier 3
  • When flagged: line passes through to player exactly as authored, same as tells

What the copy team does NOT author

Tell behavior constraints (Gestalt + Tyre)

  • The 5 TellCategory enums (Nervous, Angry, Friendly, Guarded, RoutineDeviation) map to 5 voice context clauses
  • These context clauses are injected when the relevant TellCategory is active — informing how surrounding behaviors and dialogue are voiced
  • The tell behaviors themselves pass through unchanged; the context clauses are not output, they're input constraints
  • I can write these 5 context clauses (it's copy work) — but the semantic definitions of what each category means must come from Gestalt before I draft. Flagging as a dependency.

Tell base texts (automated)

  • Tell behaviors are algorithmically generated from DerivedTellState per Tyre's Round 2 clarification
  • Fixed library: 5 categories × N cultures = ~20-40 voiced tell strings total, baked at build time per culture
  • Copy team does not author these; copy team reviews them once per culture as part of the culture QA process

Review process: baked vs pre-voiced

Baked content (hub zones, first hours of gameplay)

This is the quality reference — what the player's first experience of the voiced system looks like. Human review is mandatory before ship.

Process:

  1. Tyre or Troblum runs the inference pipeline on all hub zone NPCs (Sova Transit District)
  2. Output is written to a review file per NPC, behavior/dialogue line by line
  3. I review each output against three criteria: (a) culture register correct, (b) no lore contamination, (c) specific base text content preserved
  4. Lines that pass: approved. Lines that fail: either rewritten by hand (treat as authored) or base text escalated (override with a better base text)
  5. I sign off on the baked output before it's committed to the build

Volume estimate: Sova Transit District at ~20 NPCs × 8 behaviors + 10 dialogue lines each = 360 voiced lines to review. Realistically 3-4 hours of review if output quality is good.

Pre-voiced runtime content

No human review of individual lines before player encounters them. This is the risk-managed tier.

Quality controls:

  • 5% of all runtime-voiced output is sampled to a log file
  • I review sampled logs on a per-sprint basis (fast when output quality is stable; longer when drift is detected)
  • Automated keyword scan against NI-1 through NI-6 blocklist runs on all output — any hit generates a flag for review
  • If the keyword scan hit rate rises above 2%, it's a signal that model or injector drift has occurred and a prompt audit is needed

Injector maintenance

  • Culture injectors are versioned. When I update an injector clause, all pre-voiced content generated with the previous version is invalidated (cache invalidation follows injector version hash)
  • This is Tyre's architecture decision, but I need to know the mechanism — if I can't iterate injectors without full cache invalidation, I have to be more conservative about updates

4. D-123 Amendment Review

Paula is drafting the amendment text. My proposed language from Round 2 has broad support and is reproduced here for Paula's reference. Jeroen's decision explicitly requires the build-time/runtime mode distinction, which Paula's N-1 framing correctly identified and which my language below incorporates.

Proposed amendment text (for Paula's review and refinement):

D-123 is amended as follows: The AI pipeline operates in two modes: (1) build-time authoring tool — content generated before ship, reviewed by humans, committed to the build as reviewed; and (2) background runtime enhancement — content generated during gameplay as the player moves through the world, without per-line human review, governed by automated validation and periodic sampling. Both modes are in scope for the Settled Reach.

All original D-123 constraints remain binding in both modes: culture vectors are the primary prompt constraint; the AI does not default to genre conventions; authorial control governs what the LLM may and may not produce. The AI pipeline does not generate narrative decisions — it applies voice to authored semantic content. D-124 is superseded.

Copy-team-specific flag for Paula: The amendment should explicitly note that anchor lines (D-092) and tell behaviors are excluded from LLM re-voicing in both modes. These are not covered by D-123 as written but must be covered by the amended text to prevent ambiguity.


5. Tell Context Clauses (draft, pending Gestalt sign-off on category semantics)

These are the 5 injector clauses that inform the LLM about an NPC's active tell state. They are injected into the behavior/dialogue prompt when a TellCategory is active. The tell itself never goes to the LLM; these context clauses do.

Draft — written for Van Maanen's Star culture register but should be culture-neutral as system context:

TellCategory Context clause injected into prompt
Nervous "This NPC is under stress and concealing something. Their attention is divided. They appear normally busy but their focus is not fully on what's in front of them. Responses may be shorter or more clipped than usual."
Angry "This NPC is suppressing anger. The surface is controlled, but there's an edge under it. Their patience is shorter than normal. They complete tasks but don't invite conversation."
Friendly "This NPC is in an open, positive state. They're more likely than usual to volunteer a word, hold a moment of eye contact longer, or acknowledge a familiar face."
Guarded "This NPC is watchful and giving nothing away. They answer questions with the minimum required. They are not hostile — they are contained."
RoutineDeviation "This NPC is not where they normally are or doing what they normally do. Something has changed. Their behavior may be slightly off-pattern in ways that could be read as distraction or purpose."

Note for Gestalt: If the category semantics are substantially different from what's above, let me know and I'll revise. These are drafted from tell_state.rs partial visibility.


One thing to settle before Spike 1

Tell context clause culture-neutrality: The 5 context clauses above are written to be culture-neutral — they describe the NPC's internal state without Van Maanen's Star register. They're constraints, not output. This is intentional: culture injectors voice the output; tell context clauses describe the state. If the clauses are written in Van Maanen's Star voice, the LLM may voice the constraint itself rather than apply it.

Gestalt — confirm: context clauses are system prompt input, the player never sees them, and they should be maximally descriptive rather than voiced? If so, the drafts above are correct. If tell context clauses need to be per-culture (e.g., "this Van Maanen's Star NPC is under stress..."), the authoring burden increases significantly and I'd want to know now.