# Conflicts: # CHANGELOG.md # content/_meta/README.md # content/_meta/npc-authoring-style-guide.md # wiki/_templates/cultural-group.md # wiki/_templates/institution.md # wiki/_templates/star-system.md # wiki/characters/devra.md # wiki/characters/drin.md # wiki/characters/harek.md # wiki/characters/lera-sessik.md # wiki/characters/maret-korr.md # wiki/characters/naia-tamm.md # wiki/characters/nils-davan.md # wiki/characters/pell.md # wiki/characters/renn.md # wiki/characters/resha.md # wiki/characters/sabel.md # wiki/characters/sera-venn.md # wiki/characters/torek-lintar.md # wiki/characters/voss.md # wiki/star-systems/krenn/index.md
12 KiB
title, description, type, status, workshop, agent, round, created
| title | description | type | status | workshop | agent | round | created |
|---|---|---|---|---|---|---|---|
| Mellanie Round 2: Content Authoring Evaluation | Content authoring evaluation of D-123 amendment language and injector design | workshop | archived | llm-voice-pipeline | mellanie | 2 | 2026-03-07 |
Round 2: Content Authoring Evaluation
Author: Mellanie Workshop: LLM Voice Pipeline Date: 2026-03-07
D-123 tension: is the amendment language acceptable?
Short answer: Yes for A and B. Needs revision for C.
The amendment "authoring tool AND background runtime enhancement" is accurate and acceptable for Proposals A and B. What D-123 was protecting against was live, autonomous, player-prompt-driven generation — an LLM improvising narrative outside authorial control. Background pre-voicing doesn't do that. The LLM receives authored base text, authored injectors, and authored constraints. It applies register. That's closer to a pipeline tool that happens to run on the player's machine than to a runtime AI system in the dangerous sense.
D-123's core principles survive the amendment:
- Culture vectors as primary prompt constraint — preserved
- AI doesn't default to genre conventions — preserved (that's what injectors + negative lists are for)
- Authorial control over what the LLM can and can't do — preserved
The only thing that changes: "not a runtime system" → "background runtime enhancement when AI-Enhanced Dialogue is ON."
Proposal C's language is a problem. "D-123 is fully superseded" implies throwing out the whole decision. The runtime restriction is the only bit that needs to change. The rest of D-123 — culture vectors primary, no genre-convention defaults, authorial constraints binding — needs to stay in force for Proposal C just as much as for A and B. If "fully superseded" means we're free to ignore culture vectors and let the LLM rephrase however it wants for dialogue, that's a regression, not an improvement.
Proposed amendment language for all three proposals:
D-123 is amended as follows: "The AI pipeline is an authoring tool for content assembly AND a background runtime enhancement when AI-Enhanced Dialogue is enabled. All other constraints remain binding: culture vectors are the primary prompt constraint, the AI does not default to genre conventions, and authorial control governs what the LLM may and may not produce. The AI pipeline does not drive live narrative decisions — it applies voice to authored semantic content."
This covers all three proposals. No proposal fully supersedes D-123; they all amend it.
One specific correction to flag: proposed-llm-voice.md Section 4 gives as an example Van Maanen's Star Culture injector: "Your speech is formal and avoids contractions." This is wrong. Van Maanen's Star is direct and working-class, not formal. It uses contractions constantly ("shift's calling", "gotta move", "can't get there from here"). If this example injector shipped as-is, every Van Maanen's Star NPC would sound like a mid-level bureaucrat. The injector drafts in this document (below) correct this.
Resolution matrix
| Question | Answer |
|---|---|
| Which proposal do I recommend? | B |
| Blockers in Proposal B? | One: semantic core labels require careful per-tell authoring — doable but copy team needs a definition of the full tell taxonomy first (from Gestalt/Tyre) |
| Can I live with Proposal A? | Yes — clean, safe, and the architecture supports adding B later |
| Can I live with Proposal C? | Yes, with amended D-123 language and hard requirement for human review of all baked dialogue output |
| Minimum change to A to make it acceptable | Nothing — A is already acceptable |
| Minimum change to C to make it acceptable | (1) Amend D-123 language as above, (2) require human review sign-off on baked dialogue before ship, (3) treat 2B dialogue quality as a spike gate — if it fails, scope back to A/B |
Authoring load by proposal
Proposal A: culture injectors + base text only
New copy work:
- Culture injector clauses: 5-10 per culture (~8 for Van Maanen's Star — see drafts below)
- Trait modifier clauses: 1 per trait, 10 traits (10 sentences total)
- Negative injectors / lore contamination blocklist: ~20-30 excluded terms and genre phrases (one-time, I own this)
- Tell behavior flagging: just identifying which existing behaviors are tells, no new writing required — the passthrough system handles the rest
Ongoing work:
- Per-culture injectors when new cultures are added (same one-time cost per culture)
- Blocklist maintenance as new lore contamination patterns are identified
Volume estimate: ~2-3 days of focused copy work to stand up Van Maanen's Star completely. Each additional culture: ~1 day.
Assessment: This is the right authoring load for the copy team. Low volume, permanent leverage.
Proposal B: + semantic core labels for tells
Additional new copy work beyond A:
semantic_corelabels for each tell type: e.g.,"avoidance_behavior","nervous_fidget","concealment_tell","hostile_suppression","knowledge_gap_tell"- These aren't just labels — they're constraints that must precisely name the phenomenon the tell must preserve
- I can draft these, but I need the full tell taxonomy first: how many tell categories, what are the behavioral expressions per category? The
tell_state.rsshows 5 categories (Nervous, Angry, and others). I need the full enumeration from Gestalt/Tyre. - Estimated: 15-25 semantic core labels, plus documentation of what each means for the LLM constraint
Assessment: Moderate additional work, high value. The semantic core label is a copy team artifact — it requires understanding both the narrative intent (what the tell is communicating to the player) and the LLM instruction (what must survive revoicing). This is exactly the kind of precision work the copy team should own, not generate automatically. The tell taxonomy spec from Gestalt blocks me here.
Proposal C: + dialogue injector context
Additional new copy work beyond A:
- Dialogue context fields in the prompt (relationship, access tier, trust tier) already exist as tags in the D-028/D-035 taxonomy — copy team doesn't author new tags, just validates the existing tags are being passed correctly
- But: baked dialogue for hub NPCs requires human review before ship — this is the real load
- How many dialogue lines per hub NPC? If Sova Transit District has ~20 ambient NPCs × 10 dialogue lines each, that's 200 voiced lines to review at bake time
- At realistic review speed (read, judge, flag or approve), 200 lines takes a day
- This is recurring cost for each new baked zone, not one-time
- Two validation passes (behavior + dialogue) instead of one
Assessment: The additional authoring load isn't in writing — the tags exist. It's in reviewing LLM dialogue output at bake time, which is labor-intensive if dialogue quality at 2B is inconsistent. If the model is reliable, review is fast. If it drifts, review becomes a bottleneck that grows with every new baked zone.
My recommendation: Proposal B
Why B over A: Culture-voiced tells are worth having. A Van Maanen's Star NPC who's nervous about a secret should express that nervousness in a Van Maanen's Star-flavored way — not a generic sci-fi way. Proposal B enables this. The semantic_core constraint is the right mechanism: it tells the LLM what phenomenon to preserve, not how to express it. That's good architecture.
Why B over C: Dialogue at 2B is the high-risk bet. Behaviors are short-form (5-15 words), the prompt is simple, failure is obvious and recoverable. Dialogue is longer, the prompt is more complex, and a subtle failure — dialogue that's fluent but slightly off-register — is harder to catch. Behaviors first; if the model proves itself, add dialogue.
The spike should include a B-gate: After validating behavior re-voicing (Proposal A tests), run a constrained re-voicing test with semantic_core on 5 tell behaviors. If the phenomenon survives in all 5 cases, we've validated B. If not, we ship A and add B when we have a stronger model.
If the spike fails for B's constrained tells: Fall back to A. The architecture supports it — tell_behaviors is a passthrough field regardless, and the semantic_core is an optional constraint layer on top.
Van Maanen's Star injector clauses — corrected drafts
The example in proposed-llm-voice.md ("Your speech is formal and avoids contractions") describes the opposite of Van Maanen's Star culture. These are the corrected injectors.
Note on format: These are written as direct LLM persona instructions — second person, imperative register. They should appear verbatim in the injector prompt, not as description-of-description.
Van Maanen's Star Culture — Voice Injectors (v1, for spike validation)
-
"Be direct. Don't waste words. Everyone you talk to is short on time, including you."
-
"You're working-class and pragmatic. You grew up in a community where you either show up and do the work, or you don't — and everyone notices which one you are."
-
"You don't trust distant authority. Management that hasn't worked a shift, institutions that talk big and deliver slow, credentials without competence — you've seen all of it, and it doesn't impress you."
-
"When something surprises or frustrates you, expressions like 'void take it', 'stars', 'cold vacuum', or 'blood and void' come naturally. They're not dramatic — they're just how people here talk."
-
"You use first names. Family names belong on contracts, registrations, and arrest records. Not in conversation."
-
"Loyalty runs narrow and deep. Your crew, your shift, your street. Not abstractions."
-
"You greet people briefly: 'hey', 'morning', 'shift treating you alright?' No ceremony."
-
"You're not rude — you're honest. If something's wrong, you say so. If something's fine, you say that too. You don't pad."
Usage notes for the injector assembly system:
- All 8 clauses should be included for every Van Maanen's Star NPC regardless of role or trait. Culture is the baseline register.
- Trait injectors layer on top: a Van Maanen's Star-Bold NPC gets clause 8 amplified; a Van Maanen's Star-Cautious NPC gets clause 8 dampened slightly.
- Mood injectors override where relevant: Van Maanen's Star-Angry should suppress the directness of clause 8 toward bluntness; Van Maanen's Star-Nervous should suppress clause 3's confidence.
- Do NOT use clause 4 (void-oaths) in neutral-register behaviors. Gate it to high-affect contexts. This is the composition engine's responsibility, but flag it explicitly so the system doesn't inject "void take it" into "checks a manifest."
Tell taxonomy blocker
I need the following from Gestalt + Tyre before I can write semantic core labels for Proposal B:
- Full enumeration of tell categories (I see 5 in
tell_state.rsbut only partially — what are all five?) - Whether tell categories map 1:1 to semantic core labels or whether one category can have multiple labels (e.g., "Nervous" might express as
nervous_fidget,avoidance_behavior, orconcealment_telldepending on context — are these separate semantic cores or one?) - Confirmation that
tell_behaviors: Vec<String>is the accepted schema field name — I'll use this in the semantic core label documentation
Once I have the tell taxonomy, I can draft all semantic core labels within a day. They're not long — they're precise.
One thing that should not be left open
The proposed-llm-voice.md lists Gemma 2B and Phi-3-mini as spike candidates. The updated proposals specify Gemma 2 2B (Q4_K_M) as the resolved candidate (C-6). I want to confirm: is the spike still testing both models, or just Gemma 2 2B?
From a copy team perspective, the spike test payloads I'll write will work with either model — I'll produce semantic lines + injector combos, not model-specific prompts. But if we're testing both, I want to write payloads that stress-test injector faithfulness specifically, because that's where 2B models tend to drift. Tell me what you need and I'll have test payloads ready.