# Conflicts: # CHANGELOG.md # content/_meta/README.md # content/_meta/npc-authoring-style-guide.md # wiki/_templates/cultural-group.md # wiki/_templates/institution.md # wiki/_templates/star-system.md # wiki/characters/devra.md # wiki/characters/drin.md # wiki/characters/harek.md # wiki/characters/lera-sessik.md # wiki/characters/maret-korr.md # wiki/characters/naia-tamm.md # wiki/characters/nils-davan.md # wiki/characters/pell.md # wiki/characters/renn.md # wiki/characters/resha.md # wiki/characters/sabel.md # wiki/characters/sera-venn.md # wiki/characters/torek-lintar.md # wiki/characters/voss.md # wiki/star-systems/krenn/index.md
11 KiB
title, description, type, status, workshop, agent, round, created
| title | description | type | status | workshop | agent | round | created |
|---|---|---|---|---|---|---|---|
| Ozzie Round 1: Player Experience Inventory | Player experience inventory evaluating wow factor of LLM voice pipeline options | workshop | archived | llm-voice-pipeline | ozzie | 1 | 2026-03-07 |
Round 1 — Ozzie: Player Experience Inventory
Workshop: LLM Voice Pipeline Role: Player experience / wow factor advocate Round: 1 (Divergent Inventory)
My gut reaction to the three options
I read the behavior pools in rural-zone-spec.ron. Then I looked at the hardcoded base texts in the spike binary. That comparison IS the whole conversation.
Hand-authored pool (current):
"holds eye contact through a long pause, waiting for the price to land" "wipes grease on the thigh of her coveralls between jobs" "calls across a field to a neighbor without looking up from work"
Hardcoded base texts in the spike:
"tends crops in the field" "checks credentials at the gate" "watches foot traffic from market stall"
That gap is enormous. The base texts look like placeholder copy. They look like the developer left notes for the writer. "Tends crops in the field" is what you write when you're sketching the system. "Tends rows of low-growing crops with a long-handled hoe" is what the player actually sees and believes.
This tells me ONE thing before I evaluate anything else: the base text spec needs a major rethink before this architecture can work. Right now, "base text" means rough draft. For the re-voicing model to work, base text has to mean something different — it has to mean complete, evocative, and deliberately minimal. Not broken. Not placeholder. Intentionally spare in the way that Van Maanen's Star culture is spare.
More on this below. But it's the issue I'm going to fight for hardest.
Which option best serves player experience?
Option 1: Hand-authored pools — I want this. I can't have it.
The quality ceiling is exactly where it needs to be. The rural zone file reads like real people. The trader who "holds eye contact through a long pause, waiting for the price to land" — I believe that person. That's the game I want to play.
But the math kills it. O(R x Z x C) means that every new culture or zone type the team adds is authoring from scratch. We can't have Van Maanen's Star-rural AND Van Maanen's Star-industrial AND Sova-industrial AND a third culture's rural variant without a team of writers and years of budget. The game has to grow. This option doesn't grow.
Option 2: Composable primitives — I'm scared.
Composed text FEELS composed. Players feel it in their bones even when they can't name it. "Greets you warmly because [Social] + [Rural context]" produces something like "nods a friendly greeting to people passing by" — which is grammatically correct and soul-dead. The hand-authored version would be "calls across a field to a neighbor without looking up from work." Same beat, completely different texture.
The risk is real: composable systems produce text that reads like it was assembled, because it was. The seam is visible. Players stop believing the NPCs are people. When players stop believing the NPCs are people, THE FRIEND doesn't work. The contradiction doesn't land. The whole detective loop falls apart.
I'd fight hard against pure composable as our primary model.
Option 3: LLM re-voicing — YES, with conditions.
This is the only path that scales to the world we want to build AND has a shot at preserving the quality ceiling. The i18n analogy is right. The architecture is right. The implementation plan (baked + pre-voiced + fallback) is right.
But it comes with three serious player experience risks that I need the team to address before I'll commit. See below.
The three things that will make or break player experience
1. The base text problem — this is critical
The base text is THE FALLBACK EXPERIENCE. Every player on minimal hardware sees it. Every player who gets ahead of the pre-voicing queue sees it. Every player who turns AI-Enhanced Dialogue off sees it.
Right now, base texts look like design notes. That has to change.
Base text must be: complete, self-contained, evocative, and deliberately minimal. Not a stub. Not a placeholder. A different register — sparse and functional, like a stage direction — but never rough.
Think about it this way: if a theater does a stripped-down version of a play, the stripped-down version still has to WORK. It's not lesser. It's the same story told differently. That's what base text needs to be.
"Tends crops in the field" needs to become something like "works a row of low crops with steady, unhurried hands." Still culture-neutral. Still LLM-seedable. But not draft copy.
This is authoring work. It's not free. But it's the foundation the whole architecture rests on. If the fallback experience feels broken, we've built a system where players feel punished for having modest hardware. That's a terrible message.
2. The Want tell problem — this one scares me most
The brief flags this and it's RIGHT to flag it. The Want/State layer is the core of the detection game. The player reads behaviors to infer hidden internal state. The tell is the mechanic.
If the LLM re-voices a tell and changes its semantic content, we've broken the game.
Here's the exact failure mode: an NPC whose Want is [MONEY] has a tell behavior — let's say they're a guard who "glances at the freight container being logged without checking in." The LLM re-voices this as "keeps an eye on the dock traffic" (Cautious cultural voice) or "watches the loading operation with professional attention" (Honest cultural voice). Both could be innocent. Both could be the tell. Now the player can't read it.
Tells must be locked. They should not go through the LLM re-voicing pass. They either: a) Pass through to the player as base text (intentionally culture-neutral, which actually works — tells feel MORE legible when they're stripped of cultural noise) b) Have their own separate re-voicing pass with TIGHTER constraints that preserve semantic content
I lean toward option (a). A tell that's culture-neutral IS more suspicious — it stands out. The Van Maanen's Star guard speaks direct and minimal. If they suddenly have a moment of strange stillness with the freight, that's MORE readable as a tell, not less. The culture voice actually makes tells blend in. The base voice makes them pop.
This could be a feature, not a bug. But it needs to be a decision, not an accident.
3. The AI-Enhanced Dialogue toggle — the perception problem is real
Two quality tiers means players on better hardware get a richer game. That's a real fairness issue and a real messaging problem.
But it's solvable. The solution is: don't frame it as tiers. Frame it as modes.
- AI-Enhanced Dialogue OFF: "Classic voice mode — clean, direct, full gameplay."
- AI-Enhanced Dialogue ON: "Enhanced voice mode — character-voiced, culturally textured."
Neither is "better." They're different aesthetic experiences. The functionality is identical. If we nail the base text quality (see point 1), this framing is honest.
The real danger: if we ship base text that feels like placeholder, players on low-end hardware feel cheated. If we ship base text that feels intentional and complete, they have a different experience, not a worse one.
Framing and base text quality are the two levers. Both are doable.
How large is the baked cache? Does it matter?
For hub systems (Sova Transit District): rough estimate. ~50 behaviors per role x 4 roles x 2-3 zone types = 400-600 base behaviors to voice. Each voiced output is maybe 30-80 words. At plain text, that's ~30-50KB of voiced content per hub. Even if we're verbose with metadata, we're talking low megabytes for the full first-hours baked cache.
That's nothing. Modern games ship 50GB of asset data. A few MB of voiced NPC text is below perception threshold for install size.
The per-seed cache is the wildcard. If players run 10 seeds and every seed caches voiced content for every zone they visit, that could balloon. We need a cache size cap and eviction policy. But for the baked hub content? Not a problem.
What breaks if we choose the wrong option?
If we choose hand-authored only: The game can't grow. We ship Sova Transit District beautifully and then we can't add a second culture. Every expansion is a writer-years investment. The generator spike becomes a curiosity, not a product.
If we choose composable primitives: Players feel it immediately. The NPCs stop being people. The FRIEND arc breaks because Kael needs to feel like a real person for his contradiction to hurt. A composed NPC doesn't generate that attachment. This option quietly poisons every emotional beat in the game.
If we choose LLM re-voicing without solving the tell problem: The detection mechanic degrades. Players can't reliably read tells. They learn to distrust the behavior text. Instead of reading NPCs like a detective, they start ignoring NPC behaviors as noise. THAT'S THE GAME WE BUILT. If we make the behavior layer untrustworthy, we have no game.
If we choose LLM re-voicing without fixing base text: We ship with a fallback experience that feels broken. Players on low-end hardware (which is most players) feel like they're playing the rough draft. They review the game as unfinished. We lose them before they get to the good parts.
My recommendation
LLM re-voicing, but with two pre-conditions that are non-negotiable from a player experience standpoint:
-
Base text elevation pass — the copy team needs to rewrite all base texts to "complete-and-spare" quality before this architecture goes into production. Not longer. Not more detailed. Better. The goal is: base text reads like intentional minimalism, not like a draft.
-
Tells are locked or separately controlled — Want tells do not go through the general re-voicing pass. They are either served as base text (my preference — it makes them MORE detectable, which is a design upside) or given a constrained re-voicing pass that preserves semantic content. This is a systems decision, but it has to be decided before implementation.
If those two conditions are met, this architecture can give us the world we want to build.
My one question before I commit
Are Want tells embedded in the same behavior text strings that go through LLM re-voicing, or are they a separate data channel?
If tells are mixed into the general behavior pool and indistinguishable from flavor text at the data level, we have a serious problem. The LLM won't know which lines to preserve and which to style. Every tell is at risk of paraphrase.
If tells are tagged, separated, or handled through a different pipeline, I'm comfortable proceeding.
That answer determines whether the re-voicing architecture is safe for the core mechanic. Everything else is solvable. This one I need Gestalt and Tyre to answer.
Ozzie out. Someone tell me when something explodes.