--- title: "Tyre Round 1: Technical Feasibility Inventory" description: "Technical feasibility inventory assessing architecture options for voice pipeline" type: workshop status: archived workshop: llm-voice-pipeline agent: "tyre" round: 1 created: 2026-03-07 --- # Tyre Round 1: Technical Feasibility Inventory **Domain:** Technical architecture **Input documents reviewed:** workshop-brief.md, proposed-llm-voice.md, generator_spike.rs, blueprint.rs, culture-van-maanens-star.ron, rural-zone-spec.ron, industrial-zone-spec.ron, D-010, D-024, D-121, D-122, D-123, D-128, Q-057, Q-012 --- ## 1. The Three Options — Technical Assessment ### Option 1: Hand-authored pools (current) **Difficulty tier: Easy to build, impossible to scale.** The current system works. `RoleSpec.typical_behaviors` is a `Vec`, the generator draws from it with Fisher-Yates, done. Zero runtime complexity. But the brief nails the problem: O(R x Z x C) content. Right now we have ~7 behaviors per role across 2 zones and 1 culture. Adding a second culture doubles the authoring. Adding a station zone type triples it. By the time we have 4 cultures and 5 zone types we're looking at ~700 hand-authored behavior strings just for ambient behaviors, before dialogue. The copy team already flagged this (Q-057). Technically trivial. Content-impossible at scale. Not viable as the sole strategy. ### Option 2: Composable primitives (Q-057) **Difficulty tier: Medium to build, moderate to scale, high risk of mechanical output.** The idea: decompose "tends rows of low-growing crops with a long-handled hoe" into `[action:tends] [object:crops] [tool:hoe] [manner:practiced]` and recombine with culture modifiers. This is a string assembly engine — essentially a sophisticated template system. Technical assessment: - **Build cost:** 2-3 sprints for the composition engine, tag taxonomy, and modifier system. - **Maintenance cost:** High. Every new combination needs QA. The tag taxonomy becomes a coordination bottleneck (see Q-049 ObjectTag co-maintenance problem — same class of issue). - **Output quality ceiling:** Mechanical. "Tends crops with a long-handled hoe in a direct, unhurried manner" reads like a sentence assembled from parts, because it was. The Sprint 25 spike proved that *specific, authored phrasing* is what makes behaviors legible — "wipes grease on the thigh of her coveralls between jobs" cannot be composed from primitives without losing the detail that makes it human. - **Integration:** Fits cleanly into the existing `typical_behaviors: Vec` — the composition engine produces strings, same as hand-authoring. No architectural change needed downstream. Feasible but produces the wrong output. The quality floor is too low for what the spike proved works. ### Option 3: LLM re-voicing **Difficulty tier: Challenging but doable. Let me be honest about what this means technically.** The i18n analogy is elegant and architecturally sound. Base text as both seed and fallback is a clean design that eliminates the dual-authoring problem. But "ship an LLM with the game" is not a small sentence. Let me break down what this actually requires: **What's actually easier than it sounds:** - The prompt engineering. The injector clause system maps directly to data we already have: `CultureProfile.speech`, `NpcBlueprint.traits`, `NpcWant`. The prompt is a structured assembly of existing data fields + a base text string. This is well-defined work, not open-ended AI research. - The cache/fallback model. Base text IS the fallback — no separate system needed. Cache is a string-keyed lookup: `(seed, zone, culture, npc_id, behavior_index) -> voiced_string`. Fits naturally into our existing RON/MessagePack pipeline. - Integration with the generator. `NpcBlueprint.observable_behaviors` is already `Vec`. Re-voicing replaces strings in-place. The rest of the pipeline (perception, observer, wire format) doesn't know or care whether the string was hand-authored, composed, or LLM-generated. **What's harder than it sounds:** - Model selection and bundling (see section 2). - Determinism guarantees (see section 3). - Memory budget on minimum spec (see section 2). --- ## 2. Model Selection and Inference Wrapper ### Hardware constraint: the real bottleneck Minimum spec from the brief: integrated GPU, 8GB RAM shared with game. Let me be precise about what this means. The game already claims memory: - Godot client: ~300-500MB (renderer, assets, scene tree) - Rust server process: ~100-200MB (ECS, generation, world state) - OS overhead: ~1-1.5GB - **Available for LLM: ~5-6GB absolute max, realistically 3-4GB to avoid pressure** A 2B parameter model in Q4 quantization: ~1.2-1.5GB. That fits. A 3B model in Q4: ~1.8-2.2GB. Tight but possible. Anything larger is out. ### Model candidates (2026 landscape) The proposal mentions Gemma 2B and Phi-3-mini. Let me update for what's actually available now and what matters for our specific task: | Model | Parameters | Q4 Size | Task fit | Notes | |-------|-----------|---------|----------|-------| | Gemma 2 2B | 2.6B | ~1.5GB | Good | Strong instruction following, multilingual base helps with "dialect" tasks | | Phi-3-mini | 3.8B | ~2.2GB | Better quality, tight on RAM | Microsoft's dense model, excellent reasoning per parameter | | Qwen2.5-1.5B | 1.5B | ~0.9GB | Adequate | Smallest viable option, leaves most RAM headroom | | SmolLM2-1.7B | 1.7B | ~1.0GB | Worth testing | Hugging Face, specifically designed for on-device | | Gemma 2 2B (Q3) | 2.6B | ~1.1GB | Testing needed | Aggressive quantization may hurt style consistency | **My recommendation:** Spike with Gemma 2 2B (Q4) as primary candidate, Qwen2.5-1.5B as fallback. The task is stylistic rephrasing, not reasoning — a 2B model should handle it. But the spike must validate this empirically. If a 2B model can't reliably preserve void-oaths and speech register, we have a problem. ### Rust inference wrapper Three serious options for shipping an LLM in a Rust binary: **Option A: llama.cpp via llama-cpp-rs bindings** - Maturity: High. Battle-tested across hundreds of apps. GGUF format is the standard for quantized models. - Binary size impact: ~5-8MB for the llama.cpp static library. - Startup cost: Model load from disk takes 1-3 seconds (acceptable — happens once at game start or first inference request). - GPU acceleration: Optional CUDA/Metal/Vulkan backends. CPU-only fallback works. Important: on integrated GPU systems, the CPU path may actually be faster than competing for shared GPU memory with Godot's renderer. - **My recommendation.** It's the boring choice, and boring is correct here. **Option B: candle (Hugging Face Rust ML framework)** - Pure Rust, no C++ dependency. Smaller binary footprint (~2-3MB). - Less mature for production inference. Quantization support is narrower. - Advantage: no cross-compilation headaches with C++ toolchains. - Risk: fewer model format options, less community optimization. **Option C: burn (Rust ML framework)** - Pure Rust, very early. Not production-ready for inference of transformer models at the scale we need. - Would require manual model conversion work. - **Not recommended for v0.2.** **Verdict:** llama-cpp-rs with GGUF models. It's proven, it handles quantization correctly, and the binary size impact is acceptable. We wrap it in a thin Rust crate (`sr-voice` or similar) that exposes exactly one function: `revoice(base_text: &str, context: &VoiceContext) -> String`. ### Binary size and distribution impact | Component | Size | |-----------|------| | llama.cpp static lib | ~5-8MB | | GGUF model (Q4, 2B) | ~1.2-1.5GB | | Baked voice cache (hub zones) | ~5-20MB (text only, compresses well) | | **Total distribution impact** | **~1.3-1.6GB** | This is significant but not unusual for a modern game. The model ships as a separate asset, not baked into the binary. Players who disable "AI-Enhanced Dialogue" could theoretically skip the download (future optimization, not v0.2). --- ## 3. Determinism — the D-010 Problem *cracks knuckles* — This is where it gets interesting. D-010 principle 4 mandates BTreeMap everywhere for determinism. Same seed = same world. LLM inference is inherently non-deterministic across: - Different hardware (floating point rounding) - Different quantization levels - Different batch sizes - Different llama.cpp versions **The proposal's answer — generate once per seed, cache the result — is correct but needs formalization.** ### Cache-as-determinism model The LLM does NOT run during gameplay simulation ticks. It runs during world generation (baked or lazy pre-voicing). The output is cached. From that point forward, the cached string is deterministic — it's just a lookup. ``` Generation time: base_text + context -> LLM -> voiced_text -> cache Game time: cache_key -> voiced_text (deterministic lookup) ``` **Cache key structure:** ``` (world_seed: u64, culture_id: &str, zone_type: &str, npc_stable_id: StableId, behavior_index: u8) ``` This means: - Same seed on the same machine = same voiced text (LLM output cached on first generation) - Same seed on different machines = potentially different voiced text (acceptable — the base text is identical, only the stylistic variation differs) - **Want tells and relationship behaviors: generated by the Rust simulation, then re-voiced.** The tell content is deterministic (SimRng-seeded). The voiced phrasing is cached. The gameplay-critical information (the tell exists, it references a specific person) is in the base text, not added by the LLM. ### What must NOT be re-voiced This is critical. Some strings carry precise gameplay information: | Content type | Re-voice? | Why | |-------------|-----------|-----| | Role behaviors ("tends crops") | Yes | Flavor text, no gameplay info loss | | Want tells ("watches the room in the glass of a nearby surface") | **Carefully** | The tell IS the gameplay. Re-voicing must preserve the observable action. Restrict LLM to style/voice changes, not semantic changes. | | Relationship behaviors ("talks past Rask without making eye contact") | **Carefully** | The named target and the social signal must survive re-voicing. | | AvoidingSomeone tells with named targets | **No** | Format string with `{name}` substitution. Re-voicing risks losing the name reference. | | Dialogue (future) | Yes | Culture voice is the primary enhancement target | The safe rule: **if the string contains a proper noun reference to another NPC, pass it through untouched.** The LLM can re-voice generic role actions freely. --- ## 4. Pre-voicing Queue and Lazy Generation Integration ### Same thread pool or separate? **Separate.** Here's why: The world generator (zone skeletons, NPC blueprints, tile placement) is CPU-bound Rust running on the server process. It uses `SimRng` and must be deterministic. It runs during zone loading and produces `SpikeOutput`/`NpcBlueprint` data. The voice pipeline is I/O-bound (model loading) then CPU-bound (inference), non-deterministic, and operates on generator *output*. It should run in its own thread pool with: - A bounded work queue (e.g., `crossbeam-channel` with capacity 256) - Priority ordering: P0 (plot-critical) > P1 (semi-unique) > P2 (ambient) - Backpressure: if the queue is full, new items wait — the game continues with base text ### Integration with lazy world generation ``` Player enters zone trigger area -> World generator produces NpcBlueprints (deterministic, fast) -> NPC entities spawn with base_text behaviors (immediate, playable) -> Voice queue receives (blueprint, culture, zone_context) work items -> Voice worker processes queue in background -> Completed items update the behavior cache -> Next perception tick: observer reads voiced text from cache instead of base text ``` The key insight: **re-voicing is an asynchronous enhancement, not a blocking dependency.** The game is always playable with base text. Voiced text replaces it when ready. The observer system (`ObserverSnapshot`) already reads behavior strings from a cache — we just add a "voiced version available?" check. ### Latency budget For background generation on minimum-spec hardware (CPU-only inference on a 2B model): - Per-behavior re-voicing: ~200-500ms per inference call (short input, short output) - Per-NPC (2 behaviors): ~400ms-1s - Per-zone (10 NPCs): ~4-10 seconds - **Adjacent zone pre-voicing while player is in current zone: easily achievable.** Player spends minutes in a zone; pre-voicing the next zone takes seconds. On higher-spec hardware with GPU acceleration: 5-10x faster. Negligible. --- ## 5. Cache Format and Invalidation ### Format MessagePack (D-020) for consistency with the rest of the pipeline. The voice cache is a flat map: ```rust struct VoiceCache { /// (seed, zone, culture, npc_id, behavior_idx) -> voiced string entries: BTreeMap, /// Model version used to generate these entries model_version: String, /// Cache format version for migration format_version: u8, } ``` Stored per-zone as `.msgpack` files alongside save data. Baked caches for hub zones ship as game assets. ### Invalidation rules | Event | Invalidation scope | Rationale | |-------|-------------------|-----------| | Seed change (new game) | Full regen | Different world = different NPCs = different voiced text | | Culture mod added | Culture-scoped regen | Culture injectors changed, all culture-specific voicing is stale | | Model update (game patch) | Full regen | Different model = potentially different output | | Zone spec change (content patch) | Zone-scoped regen | Base text pool changed | | NPC relationship change (runtime) | Single NPC regen | Relationship behavior text changed | **Save file interaction:** Voice cache is NOT part of the save file. It's a derived artifact that can be regenerated. Save files remain small (D-020). The cache lives in a separate directory (`user://voice_cache/`). --- ## 6. What Breaks If We Choose Wrong ### If we choose Option 1 (hand-authored only): - Content team hits a wall at 3 cultures x 5 zone types. D-122 (all NPCs generated) becomes impossible to staff. Q-057 remains open forever because the answer is "write more strings manually." - **Risk level: Project-blocking.** The generator-first vision (D-114, D-117) requires content at scale. Hand-authoring doesn't scale. ### If we choose Option 2 (composable primitives): - We build a composition engine that produces output below the quality bar the Sprint 25 spike established. The spike proved that *specific authored detail* makes behaviors readable — "wipes grease on the thigh of her coveralls" is not composable. We'd ship a technically sophisticated system that produces bland output. - **Risk level: Quality-killing.** The game reads as procedurally generated in the pejorative sense. ### If we choose Option 3 (LLM re-voicing) and the model can't do it: - If 2B models can't reliably preserve void-oaths, speech register, and relationship references, we've built an inference pipeline for nothing. The fallback is base text — playable but bland. - **Risk level: Recoverable.** The base text fallback means the game ships either way. The infrastructure cost (~3 sprints for the wrapper + queue + cache) is the sunk cost. - **Additional risk:** Install size. ~1.5GB for a model some players won't use. Needs to be a separate, optional download. ### If we choose a hybrid (my recommendation) and it's overengineered: - We build both the base text pipeline (already exists) and the re-voicing layer. If the LLM layer underperforms, we disable it and ship with base text + hand-authored hub content. - **Risk level: Minimal.** The hybrid architecture is inherently risk-managed by the fallback design. --- ## 7. My Position: Option 3 (LLM re-voicing) with hybrid fallback *cracks knuckles* — Let me be direct. **The architecture in proposed-llm-voice.md is sound.** The i18n model is elegant — a single design that solves content scaling, quality floor, hardware flexibility, and player choice simultaneously. That's rare. When a single architectural decision solves four problems, you take it. The technical risk is real but bounded: 1. The spike will tell us if 2B models can handle the task (1-2 sprints to answer) 2. The fallback (base text) means we ship either way 3. The infrastructure (llama-cpp-rs wrapper, queue, cache) is well-understood systems programming, not research **Effort estimate:** - Spike (model evaluation + prompt testing): 1-2 sprints - Rust inference wrapper (`sr-voice` crate): 1 sprint - Pre-voicing queue + cache system: 1 sprint - Integration with observer/generator pipeline: 1 sprint - Baked cache generation for hub zones: 0.5 sprint - **Total: 4.5-5.5 sprints** (parallelizable — spike can run while other v0.2 work continues) **Scope-wise, this means** the LLM voice pipeline is a ~6 sprint investment that replaces what would otherwise be an unbounded content authoring commitment. That's a good trade. --- ## 8. One Question I Need Answered **Does the re-voicing target observable behaviors only, or dialogue too?** The workshop brief asks this (question 1) and the answer has major architectural implications: - **Behaviors only:** The re-voicing input is the `observable_behaviors: Vec` from `NpcBlueprint`. Short strings (5-15 words), no conversation context needed. A 2B model handles this easily. Prompt is simple: "Rephrase this action description in [culture] voice with [personality] manner." - **Behaviors + dialogue:** The dialogue system (D-028) has its own tagged line pools, conversation context, and layer model. Re-voicing dialogue means the LLM needs conversation history, speaker/listener context, and longer output. This pushes toward 3B+ models, longer inference times, and significantly more complex prompting. My technical recommendation: **start with behaviors only.** Validate the pipeline end-to-end on the simpler case. Extend to dialogue in a follow-up sprint if the model and quality pass the spike criteria. The architecture supports both — the `revoice()` function doesn't care what kind of string it processes — but the prompt engineering and quality bar are different. This question determines whether the spike tests short-form re-voicing (behaviors) or long-form (dialogue), which affects model selection, latency budgets, and the spike's success criteria.