# Conflicts: # CHANGELOG.md # content/_meta/README.md # content/_meta/npc-authoring-style-guide.md # wiki/_templates/cultural-group.md # wiki/_templates/institution.md # wiki/_templates/star-system.md # wiki/characters/devra.md # wiki/characters/drin.md # wiki/characters/harek.md # wiki/characters/lera-sessik.md # wiki/characters/maret-korr.md # wiki/characters/naia-tamm.md # wiki/characters/nils-davan.md # wiki/characters/pell.md # wiki/characters/renn.md # wiki/characters/resha.md # wiki/characters/sabel.md # wiki/characters/sera-venn.md # wiki/characters/torek-lintar.md # wiki/characters/voss.md # wiki/star-systems/krenn/index.md
25 KiB
title, description, type, status, workshop, agent, round, created
| title | description | type | status | workshop | agent | round | created |
|---|---|---|---|---|---|---|---|
| Troblum Round 3: Hardware Detection, Distribution, Risk Register | Hardware detection spec, distribution strategy, and infrastructure risk register | workshop | archived | llm-voice-pipeline | troblum | 3 | 2026-03-07 |
Troblum — Round 3: Hardware Detection, Distribution, Risk Register
Workshop: LLM Voice Pipeline Domain: Infrastructure / Performance Round: 3 — Decision & Implementation Spec Date: 2026-03-07
1. Hardware Detection Spec: Three-Layer System
Jeroen's design: layer 1 (can the model load?), layer 2 (is inference fast enough?), layer 3 (recommendation threshold with player override). Here is the concrete specification for each layer.
Layer 1: RAM Check
When: Triggered on first "AI-Enhanced Dialogue" enable per session. Also re-checked on game resume if the feature was previously enabled but the game was suspended.
What it checks: Available free physical RAM at the moment of enabling.
Threshold:
| Condition | Action |
|---|---|
| Free RAM ≥ 2.0 GB | Pass — proceed to Layer 2 |
| Free RAM 1.6–2.0 GB | Marginal — warn, offer to proceed (player may have closed other applications) |
| Free RAM < 1.6 GB | Fail — feature disabled, message shown |
Rationale for 2.0 GB threshold: Gemma 2B Q4_K_M requires ~1.5 GB for weights + ~100-150 MB for KV cache at typical context lengths = ~1.65 GB peak. The 2.0 GB threshold provides ~350 MB margin for OS overhead and inference worker stack space. If the system is marginal (1.6-2.0 GB), we warn but don't refuse — the player may be able to free RAM by closing browser tabs.
Message on fail: "AI-Enhanced Dialogue requires 2 GB of free memory to run. Your system currently has [X] GB available. Close other applications and try again, or leave the setting off — the game is complete either way."
No hard minimum — if they have enough RAM, they can try.
Layer 2: Time-Per-Token (TPT) Benchmark
When: Immediately after Layer 1 passes, model is loaded (this is the same operation — the model must be loaded for the benchmark, and the load itself is the heaviest part). Benchmark runs once per installation. Result is cached in user config. Player can force a re-benchmark via settings.
What it measures: Synthetic inference run. Prompt: 150-token system prompt (universal negative injectors + minimal Van Maanen's Star culture injector + one base text seed). Output: measure wall-clock time for 20 tokens of generation. Tokens/sec = 20 ÷ elapsed_seconds.
Why 20 tokens: Fast enough to not feel like a loading screen (2-4 seconds on good hardware, 6-20 seconds on minimum spec). Long enough to average out single-token timing noise.
Implementation:
// Pseudocode — actual benchmark function in inference_worker.rs
fn run_tpt_benchmark(model: &LlamaModel) -> f32 {
let prompt = benchmark_prompt(); // hardcoded 150-token synthetic prompt
let start = Instant::now();
let result = model.generate(prompt, max_tokens: 20, temperature: 0.0);
let elapsed = start.elapsed().as_secs_f32();
20.0 / elapsed // tokens per second
}
Temperature 0.0 for the benchmark (greedy decoding) — deterministic, consistent across runs.
Result stored in: {user_data}/ai-dialogue-config.json:
{
"benchmark_tps": 7.4,
"benchmark_date": "2026-03-07",
"model_version": "gemma-2b-q4_k_m-v1.0",
"recommendation": "green"
}
Layer 3: Recommendation Thresholds
| TPT result | Status | Player message |
|---|---|---|
| ≥ 6 t/s | Green — full experience | No message. Feature enables silently. |
| 3–6 t/s | Yellow — partial experience | "Your system is running at [X] tokens/sec. Pre-voicing will work for main characters and key scenes. Background NPCs may appear in base text until the queue catches up. Continue?" |
| < 3 t/s | Red — recommend off | "Your system is running at [X] tokens/sec. Pre-voicing may not keep up with gameplay — you'll often see the unvoiced text. We recommend leaving this off, but the choice is yours." |
Threshold rationale:
At 6 t/s: rural zone (12 behavior tasks × ~4s each) pre-voices in ~48 seconds. Industrial zone (30 tasks) in ~2 minutes. Both complete comfortably before most immersive-sim player interactions.
At 3 t/s: rural zone pre-voices in ~1.6 minutes, industrial in ~4 minutes. Workable for players who move slowly and spend 10+ minutes per zone. Not workable for transit zones or fast-moving players. This is the boundary where the experience degrades from "seamless" to "sometimes base text."
At < 3 t/s: industrial zone takes 8+ minutes for behaviors alone. Even plot-critical P0 NPCs may not finish pre-voicing before the player reaches them. The feature produces no improvement over base text in practice. Recommend off.
Player override: Player can always proceed against the recommendation. The warning is a single dialog — "continue anyway / turn off." If they continue, the feature enables. No further nagging. They chose.
Ongoing TPT monitoring: The inference worker tracks a moving average of time-per-token during active inference (window: last 10 generation tasks). If sustained degradation exceeds 40% from the benchmark baseline (thermal throttling, background OS load), the feature status indicator in settings changes to yellow with a note: "Performance has dropped. Consider suspending AI dialogue." Not a forced disable — information only.
2. Bundled Distribution Plan
Jeroen decided: model ships with the game. ~1.5 GB is acceptable. No optional download.
Install Directory Structure
SettledReach/
├── game.exe (or settled-reach.x86_64 on Linux)
├── SettledReach.pck (Godot asset bundle)
├── models/
│ └── voice-pipeline/
│ ├── gemma-2b-q4_k_m.gguf (~1.5 GB — model weights)
│ └── model-manifest.json (version, checksum, performance profile)
├── data/
│ └── baked-voice/
│ ├── sova-transit-district.voicecache (pre-voiced hub content)
│ └── [other hub zones].voicecache
└── [other game files]
Why models/ is separate from data/: The model file is a large binary blob that doesn't follow standard asset versioning. Keeping it separate makes it clear to players (and antivirus software) what the file is, makes patch targeting unambiguous, and prevents the asset pipeline from trying to process it.
Why data/baked-voice/ is separate from models/: Baked voiced content is game data, not the model. It ships as compressed text records (not model weights) and is human-reviewed. It's versioned with the game, not with the model.
Model File Format
GGUF (the native llama.cpp format). Single file. Self-describing metadata header contains model architecture, quantization scheme, and vocabulary.
Checksum verification on load: the inference backend reads the model file's SHA256 hash and compares against model-manifest.json. Mismatch = log error, disable feature, surface message: "The AI dialogue model file may be corrupted. Reinstall the game to restore it."
// model-manifest.json
{
"model_id": "gemma-2b-q4_k_m",
"version": "1.0",
"sha256": "a3f9b2...",
"min_game_version": "0.2.0",
"params_billions": 2.506,
"quantization": "Q4_K_M",
"size_bytes": 1611661312,
"performance_reference": {
"i5_9400_tps": 8.0,
"ryzen_5_3600_tps": 10.5,
"m1_metal_tps": 22.0
},
"notes": "Gemma 2B by Google. Apache 2.0 + Google Gemma Terms of Use."
}
Loading Mechanism
Cold start: At game launch, no model is loaded. The inference backend is not initialized. Cold start performance is completely unaffected by the model's presence on disk.
On "AI-Enhanced Dialogue" enable: Layer 1 RAM check → Layer 2 benchmark (loads model, runs 20-token test) → result cached → feature active. Model stays resident in the inference worker's memory for the session.
On disable during session: Model is unloaded immediately. Memory freed. If re-enabled same session, model is reloaded (skips benchmark, uses cached result).
Model handle ownership: The inference worker thread owns the model context (via llama-cpp-rs bindings). No other system holds a pointer to the model. Queue interaction is through a Rust MPSC channel: other systems send VoicingTask structs, worker returns VoicingResult structs via callback channel. The model is never touched from outside the worker thread.
Baked Content Pipeline
Baked hub content is generated at build time, human-reviewed, and committed to source control as compressed text. The process:
make voice-bake— a build-time Make target- Target checks: model file exists, SHA256 matches manifest, game version is stamped
- Runs inference locally on the build machine against all hub NPC blueprint data
- Writes
data/baked-voice/*.voicecachefiles (compressed JSON: NPC stable ID → voiced lines) - Human review required before commit. The reviewer checks: oath vocab, register accuracy, franchise bleed. This is Paula and Mellanie's job, not automated.
- Once reviewed and committed, the baked cache ships in the game package
The baked content is generated once per model version × game version. It does not regenerate automatically. When the model is updated (e.g., a better quantization version), make voice-bake runs again, humans re-review, and the new baked cache commits.
Model Updates
A model update requires a game patch. The patch replaces the GGUF file in models/voice-pipeline/. Patch size = model size (~1.5 GB for a full replacement). Delta patches on binary GGUF files are not feasible — the file format is not delta-friendly.
Recommendation for Steam/distribution: mark the model file in the depot manifest as its own depot chunk, so Steam's delta update system can detect "model file unchanged" and skip re-downloading it when other game files change. This keeps routine game patches small even when the model directory is present.
For the initial v0.2 release: one model file, no update history to manage. This only becomes relevant in later versions.
Platform Notes
| Platform | Notes |
|---|---|
| Windows | Model ships in install directory. llama.cpp uses AVX2 (auto-detected). |
| macOS (Apple Silicon) | Metal acceleration via llama.cpp Metal backend — 15-30 t/s expected on M1/M2. Hardware detection benchmark will score green on all Apple Silicon. |
| Linux | Same structure as Windows. Vulkan backend available if libvulkan present. |
| Steam Deck | AMD RDNA2, Vulkan available. Expect 6-10 t/s. Should score green. |
3. Full Risk Register
Compiled from all three rounds. Severity: HIGH / MEDIUM / LOW. Status: OPEN / MITIGATED / ACCEPTED.
R-001: Quality floor — LLM output below reference bar (HIGH, OPEN)
What: 2B model produces voiced content that sounds worse than the base text it enhances. Players notice the degradation. The "enhancement" is a downgrade.
Scenario: Gemma 2B Q4 produces generic, culture-neutral phrasing that strips the specificity from well-authored base text. A farmer who "hauls produce to the market stall before the morning exchange opens" becomes "carries goods to the market" in voiced output.
Mitigation:
- Spike 1 quality gate: oath vocab >95%, register accuracy by blind review consensus, franchise bleed <2%
- If Spike 1 fails the quality gate, ship base text only — no voiced content. The system is built; it's just not enabled.
- Miri's hybrid format recommendation (80-90 tokens instruction + 120-140 tokens example pairs) directly addresses the "2B follows vocabulary lists, not cultural philosophy" concern. The spike tests both formats.
Residual risk: MEDIUM — there is no guarantee 2B quality meets the bar until Spike 1 runs. This is the primary unknown.
R-002: RAM pressure — OOM after passing Layer 1 (MEDIUM, MITIGATED)
What: Layer 1 RAM check passes at enable time. Later in the session, OS allocates memory for other operations (zone loading, asset streaming), leaving insufficient memory for inference. Model KV cache evicted, inference crashes.
Mitigation:
- Layer 1 threshold includes 350 MB safety margin above minimum model need
- Inference worker catches allocation failures and disables gracefully (returns base text for remaining session, logs error)
- Ongoing monitoring: inference worker tracks peak memory usage per session; if it approaches system limit, suspends inference preemptively
Residual risk: LOW — the margin and graceful handling cover this. Not a crash risk in normal operation.
R-003: Thermal throttling — TPT degrades from benchmark (MEDIUM-HIGH, MITIGATED)
What: Layer 2 benchmark runs on a cool CPU and scores 8 t/s (green). After 20 minutes of gameplay, CPU reaches thermal limit. Real inference speed drops to 4 t/s without the feature status changing.
Scenario: Laptop under sustained gaming load. Processor throttles from 3.5 GHz to 2.0 GHz. Player sees more base text than expected; thinks the feature is broken.
Mitigation:
- Inference worker maintains moving average TPT over last 10 tasks
- If sustained degradation >40% from benchmark baseline: settings status changes to yellow, tooltip explains thermal degradation, offers to suspend
- Not a forced disable — the player observes and decides
- Queue scheduler already pauses inference during zone transitions (C-8). This reduces sustained load by creating thermal recovery windows.
Residual risk: LOW — the mitigation handles it gracefully. Base text covers the degraded output transparently.
R-004: Lore contamination — franchise bleed (MEDIUM-HIGH, MITIGATED)
What: 2B model produces references to things that don't exist in the Settled Reach: Earth place names, modern idioms, wrong-era technology, other-franchise vocabulary.
Scenario: Van Maanen's Star dock worker says "worth a king's ransom" or references "checking his phone" or names a tool that doesn't exist in the setting.
Mitigation:
- Universal negative injectors (NI-1 through NI-5) in shared system/prefix prompt layer covering: religious terms, military ranks, anachronistic tech, banter/wit register, Earth geography
- Miri's NIs are the primary defense. The spike measures bleed rate against this baseline.
- Baked content: human review before commit (Paula, Mellanie). The hub zones have a complete editorial pass.
- Runtime content: automated blocklist scan on generated lines before caching. Terms in the blocklist trigger regeneration with a harder negative constraint. After 3 failures, fall back to base text for that line and log for review.
- Sampling: on each game build, 5% of runtime-generated lines from that build's test run are reviewed manually. Systemic bleed is caught before reaching players.
Residual risk: MEDIUM — individual franchise-bleed lines can reach players in runtime-generated content even with mitigations. The blocklist cannot anticipate every possible contamination. This is an ongoing operational concern, not a launch blocker.
R-005: Lore contamination — wrong culture register (MEDIUM, MITIGATED)
What: LLM applies a culture-neutral or wrong-culture register despite injectors. Every NPC sounds the same regardless of culture. Void-oaths absent. Direct register ignored.
Scenario: Injectors are correctly authored but the 2B model's context window pressure causes it to drop cultural constraints by token 200 of a 500-token dialogue prompt.
Mitigation:
- Oath vocabulary is an objective, measurable metric (void-oaths appear or they don't). Tracked per line.
- Hybrid format (instruction + examples) increases register reliability for small models (Miri's Round 2 finding)
- Spike 1 measures this directly: same prompts through both format variants, blind review of results
- If culture injectors fail to hold across diverse prompts in Spike 1: ship behaviors only (Proposal A mechanism) where prompts are short enough to stay within reliable context
Residual risk: MEDIUM for dialogue (long prompts); LOW for behaviors (short prompts, simpler context). The sequencing (behaviors first, dialogue later) is already the adopted architecture, which directly manages this risk.
R-006: Cache invalidation failure (LOW, MITIGATED)
What: Voiced content cached under stale key survives into a session where it's wrong (injector update, culture mod change, NPC relationship change).
Mitigation:
- Cache key:
hash(seed + zone_id + npc_stable_id + base_text + injector_set_version + model_version) - Any change to any input field changes the key — old entry is effectively dead (never looked up)
- Cache is append-only with TTL sweep (stale entries cleaned on game launch, not mid-session)
- Injector versioning: injector set is hashed on load; if injector files change, injector_set_version changes, all dependent cache entries become unreachable
Residual risk: LOW — hash-based invalidation is robust if the key is correctly defined. The main correctness requirement is that base_text is included in the key — so if the copy team improves the base text, old voiced versions don't survive.
R-007: Install size distribution friction (HIGH → ACCEPTED)
What: 1.5 GB model bundled in base install creates itch.io file limit problems (2 GB per file), slow downloads for players who don't use the feature, and perception issues.
Resolution: Jeroen decided bundled. This risk is accepted. Mitigation for itch.io: split installer into two files (base game + model pack), both downloadable from the game's itch.io page. Player downloads both; installer merges them. Not elegant but functional.
Residual risk: LOW — accepted by the project lead. Distribution packaging must account for the split-installer approach on platforms with file size limits.
R-008: Model provenance and licensing (MEDIUM, PARTIALLY MITIGATED)
What: Model license changes or becomes incompatible with commercial game distribution. Or: platform policies evolve to prohibit AI-generated content.
Gemma 2B status: Apache 2.0 + Google Gemma Terms of Use. Allows commercial distribution when bundled. No royalties. Restrictions: no misrepresenting model origin, no use to train competing models. Compatible with game distribution as of 2026-03-07.
Phi-3-mini status (fallback): MIT license. No restrictions beyond standard MIT.
Qwen status: RESOLVED. No Chinese models. Qwen is off the table by Jeroen's decision.
Mitigation:
- License terms are reviewed at each game version update (tracked in
model-manifest.jsonnotes field) - Phi-3-mini is a tested fallback — if Gemma's terms change unfavorably, we have a tested alternative
- "AI-Enhanced Dialogue" is optional. If the feature must be removed, base text remains and gameplay is unaffected.
Residual risk: MEDIUM — AI model licensing in commercial games is a new and evolving space. Monitor, don't ignore.
R-009: Save compatibility / voiced text drift (LOW-MEDIUM, MITIGATED)
What: Player reloads a save from a previous session. The voiced text they heard is different this time (model update cleared cache, or player installed on a new machine).
Mitigation:
- Cache is persistent across sessions in user data directory (
{user_data}/voice-cache/) - Cache is never automatically cleared on game update — only on model version change (injector_set_version or model_version in key changes)
- On model update: old cache entries become unreachable (key changes). Regeneration happens lazily in the background. Player may briefly see base text for previously-voiced content. This is the same as a first-install experience — acceptable.
Residual risk: LOW — player accepts that game updates may change content. Identical to the experience of a translation update in a localized game.
R-010: Inference worker crash or hang (MEDIUM, MITIGATED)
What: The llama-cpp-rs C FFI layer crashes (OOM, bad model file, unexpected input). Or: model enters a degenerate generation loop and never finishes.
Mitigation:
- Per-task timeout: 60 seconds maximum per voicing task. If exceeded, cancel task, return base text, log timeout with task parameters.
- Worker restart on crash: inference worker is a supervised Rust task. On panic, supervisor restarts it (model reload required, ~3 seconds). If 3 crashes in one session: feature auto-disables with message.
- Degenerate generation: llama.cpp's sampler handles repetition penalty; set repetition_penalty ≥ 1.1 to prevent repetition loops. Max tokens hard limit per task (50 for behaviors, 75 for dialogue) prevents infinite generation.
- SIGABRT/segfault in C layer: the worker process (if the inference is in a subprocess) isolates the crash from the game. If inline via FFI, the crash propagates to the game process — this is the main risk. Mitigation: careful OOM handling in Tyre's wrapper; never let the model load fail silently.
Residual risk: MEDIUM — C FFI is inherently riskier than pure Rust. The timeout and restart mitigations reduce impact but don't eliminate the underlying risk.
R-011: Model misclassification in spike planning (LOW → RESOLVED)
See section 4 (Phi-3 classification correction) below. Resolved in writing.
R-012: Baked content pipeline divergence (LOW-MEDIUM, MITIGATED)
What: Build-time voicing runs with a different model version, different injectors, or different prompts than the runtime voicing. Baked hub content sounds different from runtime-generated content. Quality cliff at the hub/world boundary.
Mitigation:
make voice-baketarget reads model version frommodel-manifest.jsonand fails if it doesn't match the expected version for this game build- Baked content is generated with the same prompt templates as runtime (not a special build-time path)
- The only difference is human review (baked goes through editorial; runtime does not)
- Voice cache format is identical: baked and runtime caches use the same schema, same key format
Residual risk: LOW — the build target enforces model version consistency. Process risk (someone forgets to run the bake after a model update) is addressed by making the bake a required CI check before the game package is built.
4. Phi-3 Classification Correction
This is in writing: Phi-3-mini is a 3.8 billion parameter model. It is not "2B class."
The original proposed-llm-voice.md document lists "Phi-3-mini" alongside "Gemma 2B" as two "2B class" candidates. This is incorrect. At 3.8B parameters, Phi-3-mini is approximately 52% larger than Gemma 2B (2.5B params).
Consequences for Spike 1:
| Metric | Gemma 2B Q4_K_M | Phi-3-mini Q4_K_M |
|---|---|---|
| Model weights in RAM | ~1.5 GB | ~2.2 GB |
| KV cache (at 512t context) | ~100 MB | ~130 MB |
| Total RAM footprint | ~1.6 GB | ~2.35 GB |
| Layer 1 RAM threshold needed | 2.0 GB free | 2.7 GB free |
| Decode speed (i5-9400 class) | 7–9 t/s | 4–6 t/s |
| Decode speed (Ryzen 5 3600) | 9–12 t/s | 6–8 t/s |
Phi-3-mini is slower because it has more parameters — each decode step reads more model weight data from RAM, hitting memory bandwidth harder even though both models share the same Q4 compression.
What this means for Spike 1:
Spike 1 is plumbing + quality with no game integration. RAM and throughput don't matter for Spike 1 — the model is loaded on a development machine, prompts are fed manually, outputs are evaluated. Spike 1 can and should test both models on the same prompts.
Phi-3-mini's advantage is real: at 3.8B parameters with Microsoft's instruction-tuning focus, it follows multi-constraint prompts more reliably than Gemma 2B. For dialogue re-voicing (long context, multiple simultaneous constraints: culture register + relationship state + access tier + negative injectors), this quality advantage may be decisive.
What this means for Spike 2:
If Phi-3-mini wins Spike 1 on quality, Spike 2 integration must account for:
- Layer 1 threshold: raise to 2.7 GB free RAM minimum (not 2.0 GB)
- Layer 2 benchmark: expect 4–6 t/s on i5-9400; yellow threshold triggers more often → more players see the "partial experience" warning → more players may disable the feature
- Zone pre-voicing times: ~30–45% longer across the board vs. Gemma 2B
If Gemma 2B meets the quality bar in Spike 1 (which is the primary hypothesis): Phi-3-mini remains a tested fallback for model provenance scenarios, not the primary model.
Bottom line: Test both. Pick the one that passes the quality bar. Know that Phi-3-mini's throughput penalty is real and will affect the yellow/green threshold distribution in production.
Summary
| Deliverable | Status |
|---|---|
| Hardware detection spec (3-layer) | Complete — thresholds defined, RAM check, TPT benchmark, recommendation tiers |
| Bundled distribution plan | Complete — directory structure, load mechanism, baked pipeline, update strategy |
| Risk register | Complete — 12 risks across all three rounds, severity and mitigation for each |
| Phi-3 classification | Confirmed in writing: Phi-3-mini = 3.8B, not 2B. RAM and throughput implications documented. |
Troblum, 2026-03-07