Files
settled-reach/docs/workshops/llm-voice-pipeline/mellanie-round1.md
T
jpmschweitzer 23d9ff0a58 Merge remote-tracking branch 'origin/main' into planning
# Conflicts:
#	CHANGELOG.md
#	content/_meta/README.md
#	content/_meta/npc-authoring-style-guide.md
#	wiki/_templates/cultural-group.md
#	wiki/_templates/institution.md
#	wiki/_templates/star-system.md
#	wiki/characters/devra.md
#	wiki/characters/drin.md
#	wiki/characters/harek.md
#	wiki/characters/lera-sessik.md
#	wiki/characters/maret-korr.md
#	wiki/characters/naia-tamm.md
#	wiki/characters/nils-davan.md
#	wiki/characters/pell.md
#	wiki/characters/renn.md
#	wiki/characters/resha.md
#	wiki/characters/sabel.md
#	wiki/characters/sera-venn.md
#	wiki/characters/torek-lintar.md
#	wiki/characters/voss.md
#	wiki/star-systems/krenn/index.md
2026-03-14 00:24:53 +01:00

13 KiB
Raw Blame History

title, description, type, status, workshop, agent, round, created
title description type status workshop agent round created
Mellanie Round 1: Content Authoring Inventory Content authoring inventory assessing LLM pipeline impact on copy workflow workshop archived llm-voice-pipeline mellanie 1 2026-03-07

Round 1: Content Authoring Inventory

Author: Mellanie Workshop: LLM Voice Pipeline Date: 2026-03-07


Reading notes before I start

I went back to the source: rural-zone-spec.ron, industrial-zone-spec.ron, culture-van-maanens-star.ron, the generator spike, and the relevant D-records. The question this workshop is actually asking is not "do we use an LLM?" — D-123 already decided generative AI is in the pipeline. The question is: what model, what scope, and what does the copy team own versus the system?


1. Which option produces the best authoring workflow?

Option 3 (LLM re-voicing) — with specific constraints.

Here's why Option 2 (composable primitives) is the wrong tool for the copy team: it shifts authoring from writing character voice to writing a grammar engine. "Role actions + culture modifiers + context tags" is a data schema problem, not a copywriting problem. The copy team writes sentences that breathe. Composition engines produce sentences that compile. Players notice the difference.

Here's why Option 1 (hand-authored) is already failing: the RON work from #630 — fifty lines per role — is good. It's exactly the quality we want. But we renamed those files from rural-zone-spec.ron to van-maanens-star-rural-zone.ron to make explicit what we already knew: every zone file is really culture × zone content, authored from scratch. Adding a second culture means authoring from scratch again. The math doesn't work.

Option 3 works because the copy team's current output is already the right input. The behavior lines in rural-zone-spec.ron — "tends rows of low-growing crops with a long-handled hoe," "patches a cracked irrigation pipe with strips of bonding tape" — these are semantic lines. Specific, observable, functional. They don't need to be culture-neutral to be LLM seeds; they need to be specific enough that the LLM has something real to revoice. They already are.

What Option 3 adds for the copy team: a thin injector authoring layer, once per culture. Five to ten culture injector clauses for Van Maanen's Star. We write them once; they voice every Van Maanen's Star NPC. That is the leverage point.


2. What does the authoring workflow look like?

Three layers, each with a clear owner:

Layer 1: Base text (copy team, per-role, per-zone)

This is already being written. The typical_behaviors arrays in zone RON files. No format change needed at this layer. The copy team continues writing specific, observable, present-tense action lines. Rules:

  • No culture-specific vocabulary (exclamations, Van Maanen's Star idioms) — those belong to the voiced tier
  • Specific enough to be evocative as fallback; not so culture-loaded that the LLM is fighting the base text
  • The existing lines in industrial-zone-spec.ron and rural-zone-spec.ron are already at the right register

Layer 2: Culture injectors (copy team, per-culture, authored once)

This is the new work. Currently culture-van-maanens-star.ron has a speech section: register, filler_words, greetings, farewells, exclamations. These were designed as generator inputs, not LLM injector instructions. They're useful source material, but "direct, minimal pleasantries, gets to the point" is a description of a voice, not an instruction to an LLM.

The copy team should author 5-10 explicit injector clauses per culture — sentences written directly as LLM persona instructions. Not derived automatically from the existing culture RON fields; the existing fields weren't designed for this. Written from scratch by the copy team, once per culture, living in a new voice_injectors field in the culture RON.

Example injectors for Van Maanen's Star (draft):

  • "Your speech is direct. No pleasantries. Get to the point because everyone's short on time."
  • "You are community-oriented and pragmatic. You trust people who show up and do the work."
  • "You are suspicious of distant authority and institutional rank. Competence is what earns respect."
  • "You might use words like 'void take it', 'stars', or 'cold vacuum' when surprised or frustrated."
  • "You use first names. Family names belong to forms and arrest records."

These are what the LLM receives as persona context. The copy team writes them; the pipeline uses them verbatim.

Layer 3: Personality injectors (copy team, per-trait, authored once)

The proposal mentions 10-20 personality trait injector clauses. These should be authored by the copy team, not auto-derived from trait names. "Bold" in the Settled Reach is not generic confidence — it's Van Maanen's Star-bold, which reads as directness and willingness to say an uncomfortable thing in front of people. The trait injectors need to be written with the world in mind.

Ten traits = ten injector clauses. One-time cost, high leverage.


3. Does the base text need to change for LLM re-voicing?

Do not strip culture vocabulary from base text. The fallback experience depends on it.

Players who run with AI-Enhanced Dialogue OFF see base text. If we strip it to minimal semantics — "tends crops," "checks manifest" — the fallback reads as placeholder text. The current RON lines ("tends rows of low-growing crops with a long-handled hoe") are functional prose. They do real work as fallback.

What needs to change is authorial awareness, not the format:

  • Base text should avoid culture-specific vocabulary, which should live only in the voiced tier
  • Base text should avoid first-person register (it's observable behavior, third-person present)
  • Tells embedded in behaviors need to be structurally separable — see section 5 below

The one RON format addition I'd propose: an optional voice_injectors field on the culture RON (not the zone RON) for the explicit LLM persona clauses. Everything else stays.


4. How do we quality-control LLM output?

Three failure modes, three responses:

Failure mode A — Lore contamination. The LLM introduces references, technologies, or cultural facts that don't exist in the Settled Reach. ("The Imperial Fleet," "FTL drives," real-world idioms.)

Response: I'll write a blocklist of excluded vocabulary and genre conventions — a short document the validation pass uses. Baked content (hub zones) gets human spot-check of all LLM output before ship. This is manageable because baked zones are finite. Build-time validation catches hard violations; human review catches drift.

Failure mode B — Voice drift. The LLM drifts from Van Maanen's Star register toward generic sci-fi. All NPCs start sounding the same.

Response: Per-culture ground-truth examples. I'll write 20-30 "this is what good Van Maanen's Star-voiced output looks like" examples per culture, used as LLM few-shot examples and as QA reference. Runtime content gets sampled at 5% and logged for periodic review. Not every line, but enough to detect systemic drift.

Failure mode C — Injected exclamation in wrong context. Personality injectors applied mechanically produce jarring results: "Void take it, the manifest checks out." The cultural exclamation was injected without situational awareness.

Response: The composition engine (proposal section 5) needs a context gate on culture exclamation injectors — they should only fire in high-affect situations, not neutral task behaviors. This is a systems concern but the copy team can flag which base-text lines are neutral-register and which are emotionally charged, helping the injector assembly logic.

On the tell system specifically: tell-adjacent lines require stricter QA than general behaviors. See section 5.


5. How much of the #630 work survives?

By option:

Option Survival rate What changes
Option 1 (hand-authored) 100% Nothing. The work is the product.
Option 2 (composable) 20-30% Lines become raw material for extracting primitives. Significant rewrite in a different authoring grammar.
Option 3 (LLM re-voicing) ~90% Lines become base text. Minor cleanup for register consistency. The culture-specific vocabulary moves to injectors.

The existing rural-zone-spec.ron and industrial-zone-spec.ron lines are already good LLM seeds. "Slumps into a break room chair and stares at nothing for a full minute before reaching for a drink" — that's specific, evocative, and has real situational texture. The LLM can revoice the register; it can't manufacture that specificity. The copy team's investment in specificity survives.

The only category that needs authoring review: behaviors that contain Van Maanen's Star cultural vocabulary should be flagged and either cleaned to neutral base text or moved to an explicit voiced_behaviors array for baked zones where copy team authors the voiced variant directly.


6. What breaks if we choose the wrong option?

Choose Option 1: The copy team becomes the hard scaling wall. By the time we have three cultures and five zone types, we need 750+ authored behavior lines before any dialogue. Every new zone type, every cultural variant, every sprint with new NPCs requires fresh hand-authored sentences. The content team becomes a bottleneck that grows with the world. Q-012 stays open forever.

Choose Option 2: The composition engine produces grammatically correct but voice-flat output. "Bold dock worker at Van Maanen's Star industrial zone performs checking manifest with direct confidence." Players notice that NPCs sound assembled. The tell system suffers most — tells that need to read as natural behavior start reading as labeled states. The copy team's skill set (voice, rhythm, specificity) doesn't map to grammar-authoring. We'd be asking them to work in a medium they don't think in.

Choose Option 3 with bad injectors: All NPCs converge to a middle-ground voice. The injectors become wallpaper — the LLM acknowledges them and ignores them at small model sizes. This is the most specific risk at 2B-class models: if the model can't hold culture register AND personality AND situational context simultaneously, it defaults to something legible but generic. The spike should test injector faithfulness specifically, not just fluency.

Choose Option 3 without protecting tells: Tell-bearing behaviors get revoiced like any other line. A tell that was authored as "checks exits habitually" might become "seems to always know where the exits are" for one generation and "glances toward the doors every few minutes" for another. Both are informative, but they're not the same signal. Players on different seeds encounter different phrasings of the same tell. Is that a feature (phrasing variance = feel of a living world) or a bug (the tell mechanic is information delivery, not poetry)? This needs a decision before we commit.


7. One question before I can commit

Can tells be structurally separated from regular behaviors in the RON schema?

Specifically: is there a tell_behaviors field (or equivalent) that the LLM pipeline treats differently from typical_behaviors? Or are tell-carrying lines mixed into the same array?

If tells are mixed in, the pipeline has no way to distinguish "this line is flavor" from "this line is information." The LLM will revoice both, and information fidelity becomes a probabilistic bet, not an authoring guarantee.

If they're separable, the copy team can author tell lines to be LLM-resistant — short, specific, verb-first ("checks exits habitually"), with clear behavioral focus — and flag them for pass-through or constrained revoicing. That's an authoring problem I can solve.

If they're not separable, the architecture needs to answer: is observable-behavior revoicing even safe for tells? Or does the LLM only revoice dialogue, and tells always use base text?

This is a Gestalt + Tyre question. Their answer determines whether I can sign off on Option 3 for the full behavior pipeline, or only for dialogue.


Summary position

Recommend Option 3 (LLM re-voicing). The copy team's existing work is already structured correctly for this pipeline. The per-culture injector authoring is low-volume, high-leverage, and within the copy team's skill set. Base text format needs no structural change — only authorial discipline about register.

The one unresolved blocker: tell-line structural separation. If that's solvable at the schema level, I'm in.

The thing that would break Option 3 fatally: if 2B-class models can't reliably preserve culture register when personality and mood injectors are also active. That's the spike's core test. I'll need to write test payloads — give me the model candidates and I'll produce the semantic lines and injector combos to run.