# Conflicts: # CHANGELOG.md # content/_meta/README.md # content/_meta/npc-authoring-style-guide.md # wiki/_templates/cultural-group.md # wiki/_templates/institution.md # wiki/_templates/star-system.md # wiki/characters/devra.md # wiki/characters/drin.md # wiki/characters/harek.md # wiki/characters/lera-sessik.md # wiki/characters/maret-korr.md # wiki/characters/naia-tamm.md # wiki/characters/nils-davan.md # wiki/characters/pell.md # wiki/characters/renn.md # wiki/characters/resha.md # wiki/characters/sabel.md # wiki/characters/sera-venn.md # wiki/characters/torek-lintar.md # wiki/characters/voss.md # wiki/star-systems/krenn/index.md
12 KiB
title, description, type, status, workshop, agent, round, created
| title | description | type | status | workshop | agent | round | created |
|---|---|---|---|---|---|---|---|
| Ozzie Round 2: Player Experience Evaluation | Convergent evaluation of player experience impact across pipeline proposals | workshop | archived | llm-voice-pipeline | ozzie | 2 | 2026-03-07 |
Round 2 — Ozzie: Player Experience Evaluation
Workshop: LLM Voice Pipeline Role: Player experience / wow factor advocate Round: 2 (Convergent Evaluation)
First: the three specific questions I was asked
Tell contrast — does it work?
Proposal A scenario: An NPC has three ambient behaviors, all culture-voiced. Then a tell — passing through as base text, culture-neutral.
In my gut: yes. Here's why.
When everything around the tell is richly textured — "wipes grease on the thigh of her coveralls between jobs," "borrows a tool from a neighbor and returns it without being asked" — the tell reads in a different register. Clinical. Observational. Like the player's own voice noting something. "She glances at the freight container being logged without checking in." That sentence doesn't sound like the NPC's world. It sounds like an investigator's field note.
That's exactly right. That's the detective game. The player isn't watching the NPC perform; the player is READING the world for evidence. A tell that sounds like an observation rather than a performance is a tell that invites investigation. The register difference is the tell's signal.
BUT. This only works if base texts are elevated. If "glances at the freight container" lives alongside "tends crops in the field" — i.e., if base texts look like rough drafts — then the contrast doesn't read as designed intentionality. It reads as: this line has worse writing than the others. And a player smart enough to pick up on that starts pattern-matching on text quality instead of semantic content. They find tells by spotting the worse prose. That breaks the mechanic entirely.
My verdict on tell contrast: It works. Contingent on base text elevation. I flagged this in Round 1 and I'm flagging it again. This is the load-bearing condition for the whole architecture. Not just for tells — for the entire fallback experience.
Proposal B risk — how bad is a constrained re-voicing failure?
Let me be specific about what failure looks like.
Source tell: "looks away when Kael's name comes up" (semantic core: avoidance_behavior)
Good constrained re-voice (Van Maanen's Star culture, Bold trait):
"goes quiet when Kael comes up — just for a beat, then moves on"
Still avoidance. Van Maanen's Star directness preserved. The tell pops.
Borderline constrained re-voice:
"doesn't have much to say about Kael"
Ambiguous. Could be innocent. Player might dismiss it. The tell is WEAKENED, not destroyed — but weakened tells mean players miss clues, and missing clues means the detective game gets harder in the wrong ways (not "I missed evidence" but "the evidence wasn't readable").
Failed constrained re-voice:
"seems to have a thing about Kael"
TOO explicit. The mystery collapses. The player gets handed the answer instead of discovering it. This is WORSE than missing the tell.
Worst case:
"seems distracted around the cargo manifests"
The target got lost. The tell preserved avoidance behavior but lost the relationship component (Kael → cargo manifests). The player gets a partial, misleading clue. They go looking for cargo manifest anomalies instead of watching Kael.
The worst case is the misleading partial. A missing tell is recoverable — the player replays, looks harder, finds the other evidence. A misleading tell sends players on a wrong track. That's not a missed clue, that's the game being unfair.
How bad is the damage vs the benefit?
The benefit is real and significant. A Van Maanen's Star NPC whose avoidance reads as "goes real quiet, then moves on" hits differently than a Sovari NPC whose avoidance reads as something more ceremonially formal. Cultural voice on tells makes the world feel consistent. The detective puzzle is richer if you have to READ through the culture voice to find the signal.
But the failure mode is subtle and hard to catch at scale. Baked content gets validation. Pre-voiced content at runtime — every NPC the player encounters, every seed, every zone they reach before the queue finishes — that's too much to validate exhaustively.
My verdict: Proposal B is the higher-ceiling option and I find it genuinely exciting. But it REQUIRES the spike to demonstrate constrained re-voicing reliability before I'll recommend it for tells. If the spike shows >95% semantic core preservation across diverse test payloads, I'm in. If it shows 85%, we're shipping corrupted tells into production and I'll fight against it.
Define the success bar before the spike, not after.
Dialogue gap — is it noticeable? Does it matter?
YES. And it matters more than behaviors.
Here's the thing: observable behaviors are what the player reads about the NPC from across the room. Dialogue is what the NPC says to the player's FACE. When the relationship is most direct, when the player is most invested, when the character is supposed to feel most real — that's when dialogue fires.
If behaviors are richly culture-voiced and dialogue falls back to template patterns, the gap is at its most jarring exactly when it most needs to hold. The NPC who "wipes grease on the thigh of her coveralls between jobs" then says "Hello. Do you have a question? I can assist you." That's a whiplash moment. The player's belief collapses.
Does it matter? IT'S THE ONLY THING THAT MATTERS when the player is in conversation.
That said — Proposal C has a scope problem that's real. Dialogue re-voicing is harder, longer-form, requires more context, and might need a larger model (3B). Troblum and Tyre need to answer whether that's feasible.
But from a player experience standpoint: if we ship Proposals A or B, we should be honest that we're shipping half the experience. Behaviors without dialogue is an incomplete culture voice. The NPC speaks in one voice when observed and another when approached. Players will notice. It won't break the game but it will break immersion at the moments that should be strongest.
Full proposal evaluation
Proposal A: Conservative — Behaviors Only, Tells Locked
Player experience verdict: Strong foundation. Clean risk surface. The tell passthrough works (with base text elevation). The behaviors-only scope is a real limitation but it's honest and shippable.
My concern: This is a great v1 that could feel incomplete. "The NPCs talk like themselves but speak like form letters" is a real player complaint waiting to happen.
Blocker: None, given base text elevation. Without base text elevation, the fallback experience is broken.
Can I live with it? Yes. If we ship Proposal A with a clear path to dialogue re-voicing in v0.3, this is responsible scope management.
Proposal B: Two-Track — Behaviors + Tells with Semantic Core
Player experience verdict: The highest ceiling, the most interesting result. Culture-voiced tells is the thing I didn't know I wanted until I thought about it. A BOLD Van Maanen's Star NPC's avoidance tell reads completely differently from a Cautious one. That's detective-game richness.
My concern: Constrained re-voicing failure is the scariest failure mode in this whole architecture. Not because it breaks the game loudly — because it breaks it quietly. Players can't tell the tell got corrupted. They just get a worse, less-fair experience.
Blocker: Spike success bar must be defined before implementation. If the spike doesn't hit the bar, this proposal should fall back to Proposal A tell handling (passthrough). The two-track architecture should be designed so the tell track can be switched to passthrough without rebuilding everything.
Can I live with it? Yes, with that caveat.
Proposal C: Full Pipeline — Behaviors + Dialogue, Tells Locked
Player experience verdict: This is the RIGHT architecture. Dialogue is where culture voice has the highest impact. The tell safety (passthrough, same as A) means no tell corruption risk. The scope is larger but the payoff justifies it.
My concern: Quality at 2B for dialogue. Behaviors are 5-15 words. Dialogue is 15-40 words with relationship context. A model that handles behaviors gracefully might hallucinate on dialogue. If dialogue quality fails, players experience the worst possible seam — culture-voiced observation but broken dialogue. That's worse than Proposal A.
Blocker: The spike MUST test dialogue quality separately from behavior quality. Don't average them. If behaviors pass at 2B and dialogue doesn't, we don't ship dialogue re-voicing — we fall back to Proposal A scope and wait for a dialogue-safe model.
Can I live with it? Yes — this is my preferred outcome if the spike validates dialogue quality.
Resolution Matrix
| Question | Answer |
|---|---|
| Which proposal do you recommend? | C, with A as fallback |
| Are there blockers in your recommended proposal? | Yes: dialogue quality at 2B is unvalidated. The spike must test dialogue separately. |
| Can you live with Proposal A? | Yes. Clean, safe, shippable. Missing dialogue is a real gap but honest about scope. |
| Can you live with Proposal B? | Yes, if spike defines and hits a success bar for constrained re-voicing. Requires the tell track to be switchable to passthrough without an architecture rebuild. |
| Minimum change to make A acceptable | Base text elevation pass by copy team. Without this, fallback experience reads as unfinished. |
| Minimum change to make B acceptable | Pre-defined spike success bar (I'd say >95% semantic core preservation). Fallback: if bar not met, tells revert to passthrough. |
My vote
Proposal C is the architecture we should build.
Here's the player experience argument in plain terms: the world has to feel like one thing. Not "rich when observed, functional when approached." Not "textured from a distance, generic up close." One thing. The culture voice has to be everywhere, or the player stops believing in it the moment it matters most — when they're talking to THE FRIEND, when Kael is deflecting, when the contradiction lands.
Proposal A ships half of that. It's responsible. It's safe. But "the NPC observes in Van Maanen's Star voice and speaks in form-letter voice" is a seam the player will feel.
Proposal C is the full promise. It's more work. It's a larger spike. But it's the promise we made when we said the world would feel inhabited.
My recommended implementation order if C is chosen:
- Spike tests behavior quality first (short-form, lower risk)
- Spike tests dialogue quality separately (longer-form, higher risk)
- If behavior quality passes and dialogue fails — ship Proposal A scope, iterate
- If both pass — ship Proposal C
- Tell track: passthrough in any case (don't add constrained re-voicing risk in the same sprint)
That gives us a decision tree out of the spike, not a binary pass/fail.
One thing I need the team to resolve before Round 3
What's the spike's dialogue test payload?
Behavior testing is easy — we have ~50 lines per role in the zone specs. Dialogue testing requires actual dialogue samples with relationship state and access tier context. Do we have those? If the copy team is still writing base dialogue, the spike can't test dialogue quality yet.
If we don't have dialogue samples for the spike, Proposal C can't be validated this sprint. That means A is the implementable choice now, with C as the target for next sprint.
Paula and Mellanie should answer this. Not a blocker for the architecture decision, but it determines which proposal we can actually ship.
This is the best problem we've had. We're arguing about which part of "fully alive world" to build first. I'll take that fight any day.