Files
settled-reach/docs/workshops/llm-voice-pipeline/troblum-round1.md
T
jpmschweitzerandClaude Opus 4.6 2a6a023196 docs(docs): add frontmatter to llm-voice-pipeline workshop
Standardized YAML frontmatter on all 28 files.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-13 23:40:59 +01:00

19 KiB
Raw Blame History

title, description, type, status, workshop, agent, round, created
title description type status workshop agent round created
Troblum Round 1: Infrastructure and Performance Inventory Infrastructure and performance inventory for LLM voice pipeline hardware requirements workshop archived llm-voice-pipeline troblum 1 2026-03-07

Troblum — Round 1: Infrastructure & Performance Inventory

Workshop: LLM Voice Pipeline Domain: Infrastructure / Performance Date: 2026-03-07


Preamble

I read the full brief, the proposed architecture document, the generator spike output, the blueprint structs, both zone RON files, the culture file, and the relevant decisions. I have numbers. The numbers are not all good. Here is what I found.


1. Which Option Is Viable From an Infrastructure Perspective?

Option 1 (hand-authored): Infrastructure cost is zero. The scaling problem is authoring bandwidth, not compute. No runtime risk. This is the "safe" answer that produces the content wall the workshop exists to solve.

Option 2 (composable primitives): Infrastructure cost is also near-zero. String assembly at runtime is trivially cheap — microseconds per call, no memory overhead worth measuring. The complexity lives in the composition engine logic, not the hardware. Infrastructure has no objection here.

Option 3 (LLM re-voicing): Infrastructure has strong opinions, and they are conditional. This option is viable ONLY if the model size and inference runtime are chosen correctly. The margin for error is real. Details follow.


2. Memory Budget: 2B Model on 8GB RAM Shared With the Game

This is the most load-bearing constraint and the brief is too vague about it. Here is what I can say precisely:

Model memory footprint at different quantization levels

Model Params FP16 RAM INT8 RAM Q4_K_M RAM
Gemma 2B 2.5B 5.0 GB 2.5 GB ~1.5 GB
Phi-3-mini 3.8B 7.6 GB 3.8 GB ~2.2 GB
SmolLM2-1.7B 1.7B 3.4 GB 1.7 GB ~1.0 GB
Qwen2.5-1.5B 1.5B 3.0 GB 1.5 GB ~0.9 GB

FP16 and INT8 are not viable for minimum-spec hardware. At 8GB total RAM with a game running, you need Q4 quantization (GGUF format) as a hard requirement. No exceptions.

Game RAM baseline at minimum spec

  • OS overhead: ~1.0–1.5 GB (Windows 10 minimum)
  • Godot 4 client: ~400–700 MB (scene tree, textures, audio)
  • Rust server subprocess: ~150–300 MB (ECS world, simulation state)
  • Lazy world generation working set: ~100–200 MB (chunk buffers)
  • Available headroom: ~5.3–6.3 GB

At Q4_K_M, Gemma 2B (~1.5 GB) fits comfortably. Phi-3-mini at Q4 (~2.2 GB) also fits but leaves less margin. Inference also requires a KV cache during generation — at typical context lengths of 512 tokens, this adds ~50–100 MB. So the full runtime cost of Gemma 2B Q4 is approximately 1.6 GB, and Phi-3-mini Q4 is approximately 2.3 GB.

Both fit. But Phi-3-mini leaves ~100–200 MB less margin for memory spikes during zone transitions when both world generation and inference might be active simultaneously.

VRAM on integrated GPUs

This question contains a false premise. Integrated GPUs (Intel UHD, AMD Radeon integrated) do not have dedicated VRAM. They operate under Unified Memory Architecture (UMA) — iGPU and CPU share the same physical RAM pool. The "8 GB RAM + integrated GPU" configuration means the same 8 GB pool is divided between everything.

Practical implication: model offloading to iGPU does not free up RAM; it consumes more of the same RAM through graphics/compute allocation. On Windows, iGPU typically carves out 512 MB–2 GB for its driver state. Factor this into your budget.

For discrete GPU at minimum spec (e.g., 4 GB VRAM budget laptop): full model offload to VRAM at Q4 is feasible for Gemma 2B and leaves CPU memory largely untouched. This is the best case scenario and not what minimum-spec means for our purposes.

Conclusion: We are designing for CPU-only inference at Q4 quantization. All performance estimates below assume this.


3. Inference Latency: What Throughput Can We Expect?

CPU-only, Q4_K_M quantization

Per-token generation speed depends primarily on CPU memory bandwidth and available cores. Real benchmarks from llama.cpp community testing (the only realistic runtime option — see section 6):

Hardware class CPU example Tokens/sec (Q4_K_M, 2B model)
2022+ mid-range Core i5-1235U (10 core, AVX2) 8–14 t/s
2019–2021 mid-range Core i5-8265U (4 core, AVX2) 5–8 t/s
2017–2019 budget Core i3-7100U (2 core, AVX2) 3–5 t/s
Minimum conceivable Core i3-6006U (2 core, no AVX2) 1–3 t/s

At 3 tokens/sec (minimum viable), generating a 30-token voiced behavior line requires 10 seconds. At 10 tokens/sec, that's 3 seconds.

Task volume estimate for background pre-voicing

From the zone specs, a typical zone has 4 roles × several NPCs. At population_density 2–6, the generator produces 2–12 NPCs per zone. If each NPC gets 2 behavior lines of ~30 tokens output each, with a 150-token input prompt (system prompt + injectors + semantic line):

Zone size NPCs Lines Q4 @ 5 t/s Q4 @ 3 t/s
Rural (density 2) 2–4 8–16 ~1–3 min ~2–5 min
Industrial (density 6) 6–12 24–48 ~3–10 min ~5–16 min

This is workable IF the player spends >5 minutes per zone, which is consistent with the immersive-sim design. It is NOT workable if zone transitions happen quickly (sprint-through navigation, fast travel, etc.).

On minimum-conceivable hardware at 1–2 t/s, even a small zone may never finish pre-voicing before the player leaves. That player will always see base text. This must be acceptable — and the proposal says it is (graceful fallback). Fine. But "AI-Enhanced Dialogue" as a toggle will be effectively non-functional on that hardware class regardless of the setting.

iGPU acceleration via Vulkan

llama.cpp supports Vulkan for GPU-accelerated inference. On Intel UHD Graphics (integrated), partial layer offloading (8–16 layers of a 28-layer 2B model) can yield 1.5–2× speedup. This brings 3 t/s → 5–6 t/s on otherwise-marginal hardware. Not guaranteed, requires driver support, and adds build complexity (Vulkan SDK dependency, shader compilation).

Verdict: Vulkan GPU acceleration on iGPU is worth investigating for a later optimization pass, but do not build the queue scheduler assuming it will be available. Design for CPU-only, treat iGPU as a bonus.


4. Binary Size: Bundled Model + Inference Runtime

This is the number I'm most concerned about, and the proposal document does not address it directly.

Inference runtime overhead

  • llama.cpp compiled as shared library: ~8–20 MB depending on feature flags (BLAS, CUDA, Vulkan backends)
  • Rust bindings (llama-cpp-rs or equivalent): ~2–5 MB additional
  • No runtime JVM, Python interpreter, or other large runtimes

Runtime overhead is acceptable: ~15–25 MB added to install.

Model file size

Model GGUF Q4_K_M
Gemma 2B ~1.5 GB
Phi-3-mini (3.8B) ~2.2 GB
SmolLM2-1.7B ~1.0 GB

For reference: typical indie game install sizes are 2–8 GB. Adding 1–2 GB for the AI model is a 25–100% install size increase on the low end. For Steam and GOG distribution this is uncomfortable but not impossible. For itch.io with a free tier, 2 GB per file upload is a hard limit.

Mitigation options:

  1. Separate DLC/download: Ship game without model, offer it as a free optional download for "AI-Enhanced Dialogue." Preserves base install size. Adds post-install friction.
  2. Streaming download on first enable: Player enables the toggle; game downloads the model on demand. Breaks the "no cloud" guarantee from the proposal.
  3. Accept the size: Bundle the model, ship it as part of the installer. No user friction, single package. Increases minimum download by ~1.5 GB.

Option 3 is the cleanest from a player experience standpoint. Whether the project is willing to accept that distribution overhead is a business decision, not a technical one. I note it here because the proposal doesn't mention it at all.


5. The Pre-Voicing Queue: Priority Scheduling With Lazy World Generation

The concurrency problem

Both systems are CPU-bound and memory-bandwidth-heavy:

  • World generation: procedural computation, RON deserialization, ECS entity creation. Not trivially parallelizable with inference.
  • LLM inference: sequential token generation, constant streaming of model weights through CPU cache. Cache-evicting everything else in L3 during a full inference pass.

These two workloads in the same thread pool will contend for:

  • L3 cache (inference evicts world-gen data; world-gen thrash reloads model weights mid-inference)
  • Memory bandwidth (DDR bandwidth is a shared resource; both saturate it)

Recommendation: separate thread pools with explicit priority control.

Thread pool A (world generation): 2–4 threads, standard priority
Thread pool B (LLM inference): 1 thread, below-normal OS priority

One inference thread is correct. llama.cpp parallelizes across CPU cores internally via thread count parameter. Set llama.cpp threads to physical_cores - world_gen_threads.

Resource contention risks

  1. Zone transition spike: Player moves between zones. World generation fires (new district skeleton, NPC spawn, asset loading). Simultaneously, the inference queue for the new zone fires. Both peak at the same moment. Mitigation: pause the inference queue during active zone transitions. Resume after the world generation burst subsides (detectable via a "zone settled" signal from the world gen system).

  2. Thermal throttling on laptops: Sustained LLM inference at 100% CPU generates heat. On thin laptops with aggressive thermal throttle (common on minimum-spec hardware), inference speed degrades over time. A pre-voicing session that benchmarks at 6 t/s at minute 1 may be running at 3 t/s by minute 5 due to throttle. Budget accordingly; add 50% latency margin to all estimates.

  3. Low battery / power saver mode: Windows and macOS aggressively throttle CPU on battery at power saver settings. Inference tokens/sec can drop by 60–70% in these modes. The inference queue must detect this and pause. A simple mechanism: monitor time-per-token; if it exceeds a threshold (e.g., 1 second/token), suspend inference and set a flag for the player.

Queue integration with the existing lazy world gen pipeline

The proposal describes anticipation-based pre-voicing: when the player signals intent to move to a new area, the queue populates. This is the same signal that triggers lazy world generation. Both systems want to act on the same event.

I'd recommend that world generation feeds the voicing queue as a subscriber: world gen completes a zone's NPC generation → publishes ZonePopulated event → voicing queue picks up the NpcBlueprint list and schedules voicing tasks. This avoids the voicing queue having to separately track which zones are generated.


6. Rust Inference Wrapper: Crate Candidates

The proposal says "lightweight Rust wrapper, not ollama." Correct instinct. Here are the options with honest assessments:

Crate Backend Status Notes
llama-cpp-rs llama.cpp (C FFI) Active, maintained Best CPU performance, GGUF support, quantization. Build complexity: requires C compiler, large build.
llama-cpp-2 llama.cpp (C FFI) Active Alternative binding, similar profile.
candle Pure Rust Active (HuggingFace) No C dependency. Performance: 2–3× slower than llama.cpp for CPU inference. No GGUF native support — requires safetensors format. Bigger models in RAM (no quantization parity with GGUF).
burn Pure Rust Research-grade Not production-viable for this use case.
ort ONNX Runtime Active Requires ONNX model conversion. Good performance via optimized runtime. Adds 50–100 MB ONNX runtime dependency.
llm (Rustformers) Custom Abandoned Do not use. Last commit 2023.

My recommendation: llama-cpp-rs (or llama-cpp-2) against a pinned llama.cpp version.

The performance gap with candle is too wide to accept given the minimum-spec constraints. At 3 t/s on minimum spec with llama.cpp, candle would put us at 1–1.5 t/s — non-functional for any background generation purpose. Pure Rust is a nice property; usable inference speed is a required property.

Startup cost: llama.cpp model load from GGUF on cold start is 2–5 seconds for a 2B model from SSD. Factor this into the first-run experience. The model should be loaded lazily (on first inference request) not at game startup.


7. Cache Size Estimates Per Seed

Addressed directly: cache size is not a concern.

Voiced text is just text. Average voiced line: ~50 bytes. Generous estimate for a full playthrough:

  • 100 zones × 20 NPCs × 5 behavior lines per NPC = 10,000 lines
  • At 80 bytes average (line + metadata key): ~800 KB per seed

Even at 10 seeds cached: ~8 MB. SQLite with one row per (seed, zone_id, npc_stable_id, behavior_id, injector_hash) → voiced_text would handle this trivially.

Cache invalidation:

  • Seed changes → full cache for that seed is stale. Drop the seed's partition.
  • Culture mod changes → hash the injector set per NPC; if injector hash changes, invalidate that NPC's lines.
  • Game version changes → bundle a format version tag; on version mismatch, full regeneration.

The cache can live in the user's local data directory alongside save files. No special handling needed.

Hub pre-baked content shipped in the game binary: Same math, smaller scope. 5 hub zones × 30 NPCs × 10 lines = 1,500 lines × 80 bytes = 120 KB. Negligible. Ship it as a compressed asset bundle.


8. What Breaks If We Choose the Wrong Option?

If we choose Option 3 and make wrong infrastructure decisions:

Wrong model size (too large): Phi-3-mini at FP16 on a minimum-spec 8GB machine = OOM. Game crash or forced kill of inference process. Player experience: game appears to freeze, then silently disables AI dialogue without explanation. This is the worst failure mode. Guard against it by enforcing Q4 quantization as the only supported format and documenting a VRAM floor check at feature enable.

Wrong inference runtime (candle instead of llama.cpp): Halved throughput means the feature is effectively non-functional on hardware below the median spec. Players will enable "AI-Enhanced Dialogue" and see no improvement because the queue never catches up. They will correctly conclude the feature is broken. Ship llama.cpp bindings or don't ship the feature.

No thread pool isolation: Inference runs in the same pool as world generation. During zone transitions, both peak simultaneously. Frame time spikes. Stutters. On minimum-spec hardware, this is audible as audio dropout (Godot's audio thread starved by CPU contention). Players will perceive this as a game bug, not an AI feature.

No thermal/power awareness in the queue: Sustained inference on a throttled laptop CPU generates noise (fan) and reduces battery life visibly. Players disable the feature not because of quality but because their laptop gets hot. Addressable with simple token-rate monitoring and auto-suspend.

No battery/power-saver detection: At minimum, suspend inference when Windows reports power saver mode or when time-per-token exceeds 2 seconds.

If we choose Option 2 (composable primitives):

Infrastructure risk is near-zero. Quality risk is Mellanie and Paula's problem. From my seat: no objection.

If we choose Option 1 (hand-authored):

Infrastructure risk is zero. Content authoring is the only wall. Not my domain.


9. The Phi-3-Mini Problem

The proposal document lists "Phi-3-mini" as a candidate alongside Gemma 2B and calls them both "2B class." This is incorrect.

Phi-3-mini is 3.8 billion parameters — nearly 2× the size of Gemma 2B. At Q4_K_M, it requires ~2.2 GB RAM versus Gemma 2B's ~1.5 GB. At FP16, Phi-3-mini requires 7.6 GB — functionally all of available RAM on a minimum-spec machine before the OS, game, or any other process has touched it.

More relevantly: Phi-3-mini also runs at 60–70% of Gemma 2B's inference speed on equivalent hardware due to higher parameter count. On minimum-spec hardware, this is the difference between "background generation keeps up" and "background generation never finishes."

Phi-3-mini's advantage is quality — it genuinely outperforms Gemma 2B on instruction following, which matters for multi-constraint injector prompts. The tradeoff is real. But calling them both "2B class" and treating them as equivalent candidates for a minimum-spec hardware budget is a mistake that will produce surprising results in the spike.

The spike should test both on actual minimum-spec hardware (or hardware equivalent), not just benchmarked on developer machines.


10. One Question I Need Answered Before Committing

What is the exact CPU specification for minimum-spec hardware?

The brief says "8 GB RAM, integrated GPU." It does not define the CPU. This is not a minor detail.

The difference between an Intel Core i3-7100U (2017, 2 cores, AVX2) and a Core i5-1235U (2022, 10 cores, AVX2) is 3× in inference throughput at Q4. On the i3-7100U, Gemma 2B at 3 t/s means a 12-NPC rural zone takes ~10 minutes to pre-voice. On the i5-1235U, it takes ~3 minutes. One of those is "fine for background pre-voicing while the player is active in a zone." The other one isn't.

The answer to this question determines:

  • Whether Gemma 2B is our ceiling or our floor
  • Whether we need a 1B-class model (SmolLM2-1.7B or similar) to hit functional performance on minimum hardware
  • Whether iGPU acceleration via Vulkan is worth the engineering investment
  • How aggressively the queue scheduler needs to pace itself

Until this is defined, all infrastructure feasibility assessments for Option 3 are conditional. The framework is sound. The numbers are good given specific hardware. I need the hardware defined.


Summary Table

Concern Option 1 Option 2 Option 3 (LLM)
RAM impact None None +1.5–2.2 GB at Q4
Install size impact None None +1.0–2.2 GB
Inference latency N/A Microseconds 3–14 t/s (CPU, Q4)
Zone pre-voice time N/A Instant 1–16 min depending on hardware
Thread contention risk None None High without isolation
Thermal/battery risk None None Real on minimum-spec laptops
Infrastructure risk None None Conditional (manageable with right choices)

Option 3 is not a "just add an LLM" decision. It is a systems engineering problem that requires: specific model selection validated on actual minimum-spec hardware, a correctly isolated inference thread pool, power-state awareness in the queue scheduler, and a distribution strategy for the model file. None of these are unsolvable. All of them require explicit design decisions before I can sign off.


Troblum, 2026-03-07