Files
settled-reach/docs/workshops/llm-voice-pipeline/ozzie-round3.md
T
jpmschweitzer 23d9ff0a58 Merge remote-tracking branch 'origin/main' into planning
# Conflicts:
#	CHANGELOG.md
#	content/_meta/README.md
#	content/_meta/npc-authoring-style-guide.md
#	wiki/_templates/cultural-group.md
#	wiki/_templates/institution.md
#	wiki/_templates/star-system.md
#	wiki/characters/devra.md
#	wiki/characters/drin.md
#	wiki/characters/harek.md
#	wiki/characters/lera-sessik.md
#	wiki/characters/maret-korr.md
#	wiki/characters/naia-tamm.md
#	wiki/characters/nils-davan.md
#	wiki/characters/pell.md
#	wiki/characters/renn.md
#	wiki/characters/resha.md
#	wiki/characters/sabel.md
#	wiki/characters/sera-venn.md
#	wiki/characters/torek-lintar.md
#	wiki/characters/voss.md
#	wiki/star-systems/krenn/index.md
2026-03-14 00:24:53 +01:00

15 KiB

title, description, type, status, workshop, agent, round, created
title description type status workshop agent round created
Ozzie Round 3: Player Experience Spec Player experience specification and validation criteria for voice pipeline output workshop archived llm-voice-pipeline ozzie 3 2026-03-07

Round 3 — Ozzie: Player Experience Spec

Workshop: LLM Voice Pipeline Role: Player experience / wow factor advocate Round: 3 (Commitment + Spec)


The binding decisions are good decisions

Full pipeline. Both behaviors and dialogue. Tells locked as base text with context influence. This is what I voted for. Let me write the spec for what the player actually experiences.


1. Quality seam mitigation

How tell contrast reads to the player

Jeroen's tell treatment is elegant: tells are read-only inputs to the LLM, never outputs. The tell itself ships as authored base text. The NPC's surrounding dialogue and behavior are shaped by the tell's presence.

What this means in practice:

An NPC with an avoidance tell behaves like an avoiding person — their voiced dialogue is hesitant, their ambient behaviors are distanced — but the tell itself stands apart. Clinical. Observational. "Keeps their back to the loading bay entrance when the foreman speaks." Everything else is Van Maanen's Star-voiced and inhabited. That line is a field note.

This is a feature. Call it intentional. Here's why it works:

The player's experience of the game world is dual-layered. They are both a character in the world (receiving culture-voiced content, feeling the texture of place) AND a detective reading the world (parsing evidence, noting anomalies). The tell is the moment when the detective layer activates. A tell in base text says: pay attention. this is evidence. The shift in register is the shift in mode.

The tell doesn't sound like the NPC's culture. It sounds like the player's investigation log. That's the right sound.

The risk: If base texts are elevated (which they must be — this is still non-negotiable), the distinction holds. If base texts read as rough drafts, the tell sounds like an unfinished line, not an evidence marker. The player's response is "this text is worse" instead of "this is a clue." See section 4 for the quality bar.

The mid-session transition: base text → voiced text

This is the subtler UX challenge. Pre-voicing catches up in the background. An NPC the player saw in base text at first encounter is voiced the next time they look. What happens at that seam?

The bad version: The player noticed Torek's line was "waits at a loading bay with arms crossed" (base text). They come back and it reads "leans at the bay entrance, arms folded, watching the dock traffic with the patience of someone who has done this for twenty years." They feel the difference. They wonder if something changed. They might think the game updated the NPC's state.

The problem: State-change confusion. "Did I do something that made Torek different?" No. The voicing caught up. But the player can't know that.

The solution: don't let the player see the same line change. The transition from base text to voiced text should never happen on a line the player has already read in the current session. Options:

  1. Session lock: Once a player has seen a base text line, that line stays as base text for the rest of the session. Voiced version appears next session. Clean. No jarring transitions.

  2. Zone re-entry rule: Base text is shown on first entry to a zone this session. If the player leaves and re-enters, voiced content is shown if available. This is natural — the player moved away, things changed, they returned. Re-entry provides a diegetic cover for the transition.

  3. Soft labels (not recommended): Show a subtle indicator when voiced content is available. This is the worst option — it tells the player the system exists, which breaks immersion and prompts them to think about the technology instead of the world.

My recommendation: Zone re-entry rule. It's the most natural. A player who's in a zone, reads base text, and leaves has already contextualized those NPCs. When they return, slightly different phrasing reads as: they've changed, or I'm perceiving them differently now. That's good. That's the game.

For the tells: Tells never change. Ever. Passthrough in all cases, all sessions. The tell is the anchor. Ambient lines can shift on re-entry. Tells don't.

A note on tells and context-influenced dialogue

Jeroen's "tells inform the LLM context" is the right call. If the player engages an NPC who has an active avoidance tell, the NPC's dialogue should feel avoiding — not because the tell text changes, but because the whole person is avoiding. This is how real people work. The tell is a symptom; the character is the disease.

For the player this creates a moment I'm very excited about: they see the tell (base text, stands out), they engage the NPC in dialogue (Van Maanen's Star-voiced, hesitant, deflecting), they feel the avoidance everywhere. The tell is confirmed by the conversation. THAT'S the detective loop. Evidence → engagement → confirmation.


2. Toggle UX — "AI-Enhanced Dialogue"

The framing problem

"AI-Enhanced Dialogue" sounds like: the real game is on, and you can turn it off if your hardware is bad. That's the wrong message. We need language that says: this is a choice, not a hardware penalty.

Settings screen copy

Toggle label: Character Voice Mode

State — Mode A (standard text, LLM off):

Standard — NPCs speak and act in clear, direct text. Full gameplay, any hardware.

State — Mode B (LLM active):

Enhanced — NPCs speak and act in their own voice — culturally textured, personality-inflected. Requires background processing.

Supporting note (shown below the toggle):

Both modes are complete experiences. Standard mode is intentional design, not a fallback. Some players prefer it.

Why these words

"Character Voice Mode" frames the toggle as a stylistic choice, not a quality gate. "Standard" and "Enhanced" are value-neutral — one isn't lesser. "Clear, direct text" is a positive description of base text, not an apology for it. "Full gameplay" assures players that no content is gated. "Intentional design, not a fallback" — that last line is defensive but necessary. We will have players who read reviews saying "the AI dialogue is the real experience" and feel cheated if they can't run it. This line gives them permission to enjoy the standard mode.

DO NOT use:

  • "AI-Enhanced Dialogue" as the label (sounds like a tier upgrade)
  • "Fallback" anywhere in player-facing copy
  • "Limited" or "Basic" to describe standard mode
  • "Performance Mode" (implies compromise)

First-run experience

If the player has never launched the game before, and the hardware detection recommends standard mode (see section 3), the first-run UX should present the toggle with the recommendation already applied but not yet confirmed. The player makes an active choice — they don't get defaulted into standard mode silently.

Character Voice Mode

[Enhanced] is available on your hardware, but we recommend [Standard]
for smooth performance. You can change this any time in Settings.

[Use Standard]  [Use Enhanced Anyway]

No shame on either button. "Enhanced Anyway" is not positioned as a warning — just as an informed choice.


3. Hardware detection UX — what the player sees at each stage

Layer 1: RAM check (silent)

The player never sees this. It's a pre-launch check. If the system has insufficient RAM to load the model (sub-4GB available after game load), Character Voice Mode defaults to Standard and is greyed out in Settings with a tooltip:

"Character Voice Mode requires additional memory to run. Close background applications and restart to enable."

No shame. No "your hardware is too old." Just: not enough memory right now, here's what to do.

Layer 2: Time-per-token benchmark (visible, one-time)

First time the player enables Enhanced mode, a brief benchmark runs. This is unavoidable — we have to know if inference is usable. Make it feel like the game doing something useful, not the game testing the player's machine.

Loading screen framing:

Preparing character voices...

That's it. No "benchmarking your hardware." No "testing inference speed." From the player's perspective, the game is getting characters ready. Which is true.

After the benchmark, one of two states:

If inference is fast enough: No message. Character Voice Mode activates. The player never learns a benchmark happened.

If inference is below threshold: A short, non-alarming pop-up:

Character voices are running slowly on your hardware.

[Standard mode] will give you a smoother experience with the same full
gameplay. You can switch to [Enhanced] at any time from Settings.

[Switch to Standard]  [Keep Enhanced]

Key decisions in this copy:

  • "Running slowly" — honest, not condescending. Doesn't say "your computer is slow."
  • "Same full gameplay" — the reassurance again. Keeps hitting this.
  • "At any time" — gives them an exit. They're not locked into the slower experience.
  • "Keep Enhanced" — respects player autonomy. If they want to run it slow, that's their call.

Layer 3: Ongoing recommendation (very light touch)

If the player keeps Enhanced mode running and the queue is consistently behind (player moves faster than pre-voicing, sees base text frequently), we could surface a suggestion — but only once, only if they've seen base text fallback more than N times in a session.

One-time soft nudge (appears in a settings-adjacent notification, not a modal):

You've been seeing Standard voice text more often — Enhanced mode is
running behind on your hardware. Switch to Standard in Settings for
a consistent experience.

After this nudge, never show it again for the session. Never show it again at all if the player dismisses it. This is a suggestion, not a nag.

What we absolutely do not do:

  • Pop-up modals mid-gameplay
  • Repeat warnings
  • Change the setting without player action
  • Say anything that implies the player made a bad choice by keeping Enhanced

4. Base text elevation criteria

This is the most important deliverable in my Round 3 output because it determines whether the whole architecture works. Everything — tell contrast, toggle UX, fallback experience — rests on base text being good.

The quality bar, stated plainly

Base text should read as deliberately sparse observation — complete, evocative, and culturally neutral. Not a rough draft. Not a placeholder. An intentionally minimal form, like a stage direction that fully serves the scene.

The test: read the base text line in isolation and ask — does this feel like a person doing something real? If yes, it's at the bar. If it feels like a note-to-self, a design stub, or a sentence that's waiting to be finished, it's below the bar.

Examples: placeholder copy vs. deliberately spare

Role: Farmer

Placeholder Deliberately spare
"tends crops in the field" "works a crop row with slow, unhurried passes"
"does farm work" "checks seedling trays in a low prefab greenhouse"
"harvests produce" "lifts a crate of produce onto a flatbed, tests the weight, adjusts"

The placeholder tells you what job the person has. The deliberately spare version shows you a moment that implies the job, the pace, and something about the person.

Role: Militia / Security

Placeholder Deliberately spare
"checks credentials at the gate" "holds out a hand for credentials without looking up from the gate log"
"patrols the area" "walks the fence line at an even pace, eyes ahead"
"watches for trouble" "sits in the gatehouse shade with a newsline, one eye on the road"

Role: Trader

Placeholder Deliberately spare
"sells goods at stall" "squares goods on a fold-out display with small, deliberate adjustments"
"watches customers" "leans back on a stool and watches foot traffic, says nothing"
"haggles with buyers" "holds the pause after a counteroffer, not moving"

For dialogue (these standards apply equally):

Placeholder Deliberately spare
"I don't know anything about that." "That's not something I know anything about."
"Things are difficult lately." "It's been a rough few shifts."
"You should be careful here." "Watch yourself around here."

The differences:

  • Placeholder reads generic, applicable to any character anywhere
  • Deliberately spare reads specific, even if culturally neutral — it has rhythm, it has implied manner
  • Deliberately spare is still short — it's not adding words, it's finding better words

The test the copy team should apply

For every base text line, ask three questions:

  1. Does this show a moment, not a category? ("holds the pause" > "waits")
  2. Could you imagine a specific person doing this? Not "a guard" — a guard with weight, with habit
  3. Would you be okay reading this as the only text the player sees? Not "this is a fine draft," but "this IS the experience for some players"

If any answer is no, the line needs work.

Scale note

The copy team doesn't need to elevate every line at once. Priority order:

  1. Hub zones (Sova Transit District) — these are baked, always visible, represent the quality floor
  2. Plot-critical NPC roles — foremen, guards, traders in story-adjacent locations
  3. Tells — always passthrough, always highest priority for elevation (they're the mechanic)
  4. Ambient roles in non-hub zones — lowest urgency; pre-voicing will catch these

My summary statement for the D-record

The player experience architecture for the LLM voice pipeline rests on three interdependent pillars:

1. Base text is a designed aesthetic, not a fallback. It reads as deliberately spare observation. Standard mode is a complete experience. The copy team must author base texts to this bar, not to a rough-draft bar.

2. Tell contrast is intentional. Tells in base text read as detective observations against culture-voiced ambient content. This is not a seam — it is a designed register shift that signals "pay attention here." Preserve this distinction in all documentation, all onboarding, all QA.

3. Player autonomy is respected at every hardware decision. The game never makes choices for the player. It recommends. It explains. It never shames. The toggle exists in Settings at all times. The player can always override.

These three pillars are the spec. If the D-record captures them, the implementation team has what they need.


Done. Let's build the thing.