Files
settled-reach/docs/workshops/test-architecture/round-1-notes.md
T
jpmschweitzerandClaude Opus 4.6 a87c95a6eb docs(workshops): complete QA test architecture workshop
3-round workshop with 7 agents (Tyre, Dudley, Stig, Hoshe,
Justine, Gestalt, Ozzie) plus Qatux documenting. Produced:

- 59-item prioritized test backlog (60 tickets under epic #455)
- Gauntlet test world spec: 7 rooms + hub, 48 entities
- Test client binary spec (tooling/test-client/)
- Determinism fixes (3 patches, ~22 lines)
- Server --test-mode + --port 0 design
- Content cross-reference validation (9 checks)
- make pre-pr pipeline (6-step)
- 38 client tests prioritized
- Anti-tedium features (reset plate, hub teleport, WRONG button)
- Human tester walkthrough
- CI pipeline design (deferred but documented)

Sprint 8 scope: ~17.75 team-days across 26 tickets.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-17 14:00:15 +01:00

564 lines
30 KiB
Markdown

# QA Strategy & Test Architecture Workshop — Round 1 Notes
**Workshop:** QA Strategy & Test Architecture
**Round:** 1 (Analysis)
**Date:** 2026-02-17
**Documenter:** Qatux
**Participants:** Gestalt, Dudley, Tyre, Stig, Hoshe, Justine, Ozzie
---
## Executive Summary
Round 1 produced strong consensus on the major architectural questions and a wealth of concrete specifications. All 7 participants delivered detailed, code-level proposals. The key finding: **the simulation is remarkably close to deterministic already** — Tyre and Dudley independently audited the codebase and found only 2-4 targeted fixes needed, totaling ~15 lines of code. The workshop also surfaced 5 new Gauntlet rooms (from Gestalt), 36 map-agnostic invariants, a 41-value boundary test matrix (from Hoshe), and a concrete CI pipeline design (from Justine).
**Five decisions require lead confirmation before Round 2 can proceed.** See [Decisions Requiring Confirmation](#decisions-requiring-confirmation) below.
---
## Track 1: Test World Design (Gauntlet)
### Participants: Gestalt (T1), Dudley (T1), Tyre (T1), Ozzie (T1)
### 1.1 Gauntlet Implementation Strategy
**Consensus: Hybrid YAML content pack + Rust ECS state injection.**
Both Tyre and Dudley independently proposed the same hybrid approach — strong convergence.
| Aspect | YAML content pack | Rust builder function |
|--------|-------------------|----------------------|
| Room geometry / walls | WalkabilityMap (Dudley: Rust builder; Tyre: "YAML for geometry") | — |
| NPC/item placement | `content/gauntlet/` directory, loaded by existing ContentPlugin | — |
| Knowledge graph states | — | Post-spawn injection in `inject_gauntlet_state()` |
| Trust tiers, contradictions | — | Programmatic: `player_kg.set_relationship(...)` |
| Named coordinates | `gauntlet_coords` constants module | Referenced by tests |
**Minor tension on geometry:** Dudley proposes WalkabilityMap construction in Rust (`server/src/test_world/rooms.rs`). Tyre proposes "YAML for geometry + entity placement." The difference is whether room walls are YAML or code. **Both agree** the content loader handles NPCs; the question is map geometry specifically.
**Proposed by:** Dudley and Tyre (independently converged)
**Estimated effort:** 2-3 days for first 3 rooms, ~0.5 day per additional room (Tyre)
### 1.2 Gauntlet Layout: Hub-and-Spoke with Cross-Cuts
**Proposed by:** Ozzie
Ozzie argued strongly against both pure linear and pure hub-and-spoke layouts. Recommendation: **central hub** with connections to every room, plus **intentional cross-room corridors** for testing system combinations.
Key cross-cut corridors proposed:
- Crowd Plaza <-> Occlusion Corridor (sprint through crowd into LOS testing)
- Fog Theater <-> Dialogue Room (cognitive delay + dialogue interaction)
- Inventory Warehouse <-> Interaction Gallery (full inventory + interaction verbs)
- Pause Chamber connects to all rooms
No dissent on this from other participants.
### 1.3 Missing Gauntlet Rooms
**Proposed by:** Gestalt
Gestalt identified 5 rooms missing from the workshop brief's proposed list, each targeting system *combinations* not tested by isolated rooms:
| Room | Priority | Systems Combined | Key Mechanic Tested |
|------|----------|-----------------|---------------------|
| **Eavesdrop Alcove** | P1 | D-071 + D-072 + D-053 | ListeningFocus, zone ambient, Careful stance perception |
| **Confrontation Stage** | P0 | D-070 + D-069 + D-018 | Confrontation as cognitive vulnerability, audio dip, post-confrontation delayed monologue |
| **Sprint Gauntlet** | P1 | D-055 + D-053 + D-060 | Sprint suppression, anomaly survival, interaction buffer clearing |
| **Sound Lab** | P0 | D-018 + D-059 + D-067 | Three-range sound model, sound pings, recognition chime sequence |
| **Decay Observatory** | P2 | D-041 + D-059 + D-011 | Knowledge decay timing, fog representation degradation |
Additionally, Gestalt proposed a **Shift Change Room** as the combined stress test — 15 NPCs simultaneously changing state (arriving, leaving, conversing, entering/exiting fog) within a 10-tick window. This replaces simple density testing with *simultaneous system activation* testing.
Gestalt also defined 6 **cross-room transition scenarios** testing system combinations at room boundaries (sprint exit, peripheral interact, cognitive delay + dialogue, confrontation + sound, pause during delay, walk-away sprint).
### 1.4 Gauntlet Signage and Tester UX
**Proposed by:** Ozzie
Three-layer signage system:
1. **In-world station signage** — diegetic labels ("PERCEPTION CALIBRATION BAY"), always visible
2. **Debug overlay (F3 toggle)** — floating entity labels with full state data
3. **Printed checklist** — auto-generated from room metadata, markdown format
Stig's complementary proposal: checklists co-located with Gauntlet room YAML definitions. `make checklist` generates `gauntlet-checklist.md`. New rooms automatically extend the checklist.
### 1.5 Anti-Tedium Features
**Proposed by:** Ozzie (emphatic — "this is where the Gauntlet lives or dies")
| Feature | Description | Priority |
|---------|-------------|----------|
| Room Reset Triggers | Floor plate at entrance resets room to tick-0 state | Mandatory |
| Instant Hub Teleport | Hotkey returns to Central Hub | Mandatory |
| Room Timer + PB | Gamified testing with speedrun incentive | Nice-to-have |
| State Inspector (F3) | Extended debug overlay with full ECS data | High |
| "WRONG" Button (F12) | One-press bug report: captures full game state + 60 ticks of history | High |
| Auto-Checklist Progress | Tracks which checklist items verified this session | Nice-to-have |
### 1.6 Reserved Zone Gate
**Proposed by:** Tyre
Minimum provision: empty room in map YAML + `ZoneTransition` trait stub + future-contract test file (`zone_transition_contracts.rs` with `#[test] #[ignore]`). Zero runtime code. ~0.5 day effort.
### 1.7 Performance Budget
**Proposed by:** Tyre
| Threshold | Tick Time | Action |
|-----------|-----------|--------|
| Target | ≤5ms | Normal operation |
| Warning | >8ms | Investigate |
| Error | >20ms | Test failure (>40% of frame budget) |
| Critical | >50ms | Tick overflow, simulation can't keep up |
CI performance test: 100 ticks, measure p50/p95/max, fail if p95 > 20ms.
---
## Track 2: Deterministic Gameplay
### Participants: Tyre (T2), Dudley (T2)
### 2.1 Determinism Audit — Independent Convergence
Tyre and Dudley **independently** audited the critical path and reached **highly aligned conclusions**. The simulation is close to deterministic already. Both identified the same primary issue: `HashSet` iteration in `observer/mod.rs`.
**Combined non-determinism findings:**
| # | Location | Issue | Tyre | Dudley | Severity |
|---|----------|-------|------|--------|----------|
| 1 | `observer/mod.rs:134` | Sprint anomaly iterates `HashSet<u64>` — "first Contradicted match" is non-deterministic | CRITICAL — Fix #1 | Identified same pattern | CRITICAL |
| 2 | `observer/mod.rs:197` | `visible_tiles` Vec built from HashSet iteration, non-deterministic ordering | MEDIUM — Fix #2 | Identified as "cosmetically non-deterministic" | MEDIUM |
| 3 | `movement.rs:288` | `validate_movement` tie-breaking is implicit (archetype order) | LOW — Fix #3 (document) | "no test for which one wins" — proposes sort by Entity bits | LOW-MEDIUM |
| 4 | `interaction.rs:191` | Equidistant NPC ordering | Not flagged | Identified | LOW |
**Tyre's Fix #4 (client audit):** Need to verify Godot fog/tile rendering doesn't depend on `visible_tiles` Vec ordering. Delegated to Stig for Round 2.
**SimRng status:** Both agree the SimRng (ChaCha20) is correctly seeded and consumed in a deterministic order — IF system execution order is fixed. Dudley identified that monologue systems lack explicit `.after()` constraints and could run in either order. Tyre confirmed system ordering is explicit in `SimulationPlugin::build()` but monologue ordering needs pinning.
### 2.2 Minimum Change Set for Determinism
**Consensus: 4 targeted fixes, ~15-30 lines total.**
| Fix | Change | Lines | Blocks |
|-----|--------|-------|--------|
| `visible_ids: HashSet → BTreeSet` | `observer/mod.rs` | ~3 | Deterministic sprint anomaly detection |
| Sort `visible_tiles` by `(x, y, z)` | `observer/mod.rs` | ~1 | Canonical wire format for golden files |
| Pin monologue system ordering | `simulation/mod.rs` | ~3 | Deterministic RNG consumption |
| Sort movers by Entity bits (or document) | `movement.rs` | ~5 | Deterministic tie-breaking |
**Explicitly NOT required** (both agree):
- `WalkabilityMap.chunks: HashMap` — point lookups only, never iterated
- `validate_movement` occupied HashMap — containment checks only
- `visible_positions: HashSet` — membership checks only
- Fixed-point math — simulation is integer-only; f32 in vision_cone.rs is IEEE 754 deterministic same-platform
### 2.3 Determinism Test Specification
**Proposed by:** Tyre (with concrete code)
```
Test: gauntlet_deterministic_replay
Setup: Build headless App with SimulationPlugin + BridgePlugin (no TCP)
Input: Seed 42, fixed input sequence, run to tick 50
Assert: Two identical runs produce identical ObserverSnapshot
```
Both agree the test is feasible once the 4 fixes are applied. Estimated 1-2 days.
**Key requirement:** `ObserverSnapshot` must derive `PartialEq` or compare via canonical serialized form (since it contains f32 fields).
---
## Track 3: Test Automation & Text Renderer
### Participants: Tyre (T3), Dudley (T3), Stig (T3)
### 3.1 Test Client Architecture
**Consensus: Integration test harness in `server/tests/`, not a separate binary.**
Tyre proposed extending the `game_loop.rs` pattern. The `GauntletRunner` helper wraps a headless bevy App, pre-loads inputs, ticks, and extracts snapshots. No TCP needed for most tests; TCP used only for Layer 3 IPC tests.
```
server/tests/gauntlet/
mod.rs, helpers.rs, replay.rs
gauntlet_inventory.rs, gauntlet_occlusion.rs, etc.
```
### 3.2 Text Renderer
**Consensus: Server-side formatter on ObserverSnapshot, output to stdout.**
All three participants agree the text renderer should:
- Live server-side (no Godot dependency)
- Format ObserverSnapshot (not raw ECS state)
- Be activated via `--text-render` flag (Stig) or `--text-mode` flag (Tyre)
- Output to stdout for piping/diffing
**Missing data for text rendering** (identified by Dudley):
1. Wall positions blocking LOS paths — not in snapshot (only visible tiles)
2. Entity display names — snapshot has `entity_id` but no name
3. Explored vs. unexplored tile counts — client-side accumulation, not in snapshot
Dudley proposes an optional `blocked_entities: Vec<BlockedEntity>` field, populated only when `debug_los: true`.
**Ozzie's additions** — the text renderer MUST also show:
- Sound events (D-018 three-range model) — "40% of the perception system"
- Fog layer TYPE per entity (D-059 five layers, not just "remembered")
- Cognitive delay state with timing (D-060)
- Monologue content (not just "active")
- Entity colors with semantic labels ("#4a9ebb (Neutral)")
- What's blocked and WHY ("wall at (17,10) blocks LOS to npc:hidden-1")
**Stig's "no" on client `--text-mode`:** Text rendering is purely server-side. Client gets `--verbose` flag for one-line-per-frame debug output only.
### 3.3 Assertion Language
**Consensus: Rust test macros with named helpers. No custom DSL.**
Tyre argued against a custom DSL: "requires a parser, error handling, debugging tools, documentation. It's a language design project. Premature for v0.1." The same expressiveness comes from Rust helpers:
```rust
fn assert_entity_visible(snapshot: &ObserverSnapshot, name: &str, pos: TilePosition);
fn assert_entity_not_visible(snapshot: &ObserverSnapshot, name: &str);
fn assert_inventory_count(snapshot: &ObserverSnapshot, expected: usize);
fn assert_interaction_available(snapshot: &ObserverSnapshot, entity: &str, verb: &str);
// ... ~10 assertion helpers total
```
DSL can be revisited as a sprint 10+ luxury if needed.
### 3.4 Client-Side Assertion Targets
**Proposed by:** Stig
~32 new gdUnit4 test functions across 6 categories, all structural scene-tree assertions (no pixel comparison):
| Category | Test Count | Key Assertions |
|----------|-----------|----------------|
| Camera system | 7 | Position after ready, anchored flag, smoothing, follows movement, static during pause |
| Entity rendering | 7 | Peripheral alpha=0.5, forward alpha=1.0, D-033 colors, lifecycle |
| Z-layer ordering | 4 | Fog rect z=900, fog entities z=950, insert canvas=10, UI canvas=20 |
| Fog shader state | 3 | Visibility texture updates, exploration persistence (255→128), peripheral dimming (180) |
| UI elements | 8 | Monologue consumed once, not lost on overwrite (bug #5), interaction list, dialogue, inventory grid |
| Entity lerp | 3 | Target set on update, snap on first appearance, convergence |
### 3.5 Scripted Replay
**Proposed by:** Dudley
InputQueue pre-loading is straightforward — `push()` with tick-ordering enforcement. Streaming `ReplayFeeder` handles sequences >1000 inputs (InputQueue capacity). Wait conditions use tick-budget polling:
```rust
enum WaitCondition {
RecognitionComplete { entity_name: String },
MonologueFired { monologue_id: String },
EntityVisible { entity_name: String },
Ticks(u64),
}
```
Every wait condition gets a tick budget (e.g., 1000 ticks max) to prevent infinite loops.
---
## Track 4: Serialization & Integration Testing
### Participants: Hoshe (T4), Dudley (T4)
### 4.1 MessagePack Boundary Value Matrix
**Proposed by:** Hoshe (comprehensive specification)
41 boundary test values covering every format transition in the MessagePack integer encoding. Traces every branch in `messagepack.gd:69-95` against the MessagePack spec.
Key ranges: positive fixint (0-127), uint 8 (128-255), int 16 (256-32767), uint 16 (32768-65535), int 32 (65536-2^31-1), uint 32 (2^31-2^32-1), int 64 (2^32+), and corresponding negative boundaries.
**Critical finding — asymmetric encoding:** GDScript encoder uses int_16 for positive values 256-32767; Rust encoder (rmp_serde) uses uint_16 for the same values. Both are spec-valid but produce different bytes. Implications:
- Byte-for-byte golden file comparison between GDScript and Rust will FAIL for values 256-32767
- Golden files must be direction-specific (Rust-canonical vs GDScript-canonical)
- Decoders on both sides MUST accept both signed and unsigned encodings
### 4.2 Boundary Test Placement
**Proposed by:** Hoshe
| Layer | Location | Speed | Frequency |
|-------|----------|-------|-----------|
| Encode-only (GDScript) | `test_msgpack_boundaries.gd` | ~100ms | Every commit |
| Encode-only (Rust) | `serialization.rs` extension | ~50ms | Every commit |
| Decode cross-language | Fixtures: Rust→GDScript and GDScript→Rust | ~2s | Every PR |
| Full roundtrip | `bridge_tcp.rs` extension over TCP | ~5s | Nightly/pre-merge |
### 4.3 gen_fixtures.rs Extension
**Proposed by:** Hoshe
Add `generate_boundary_fixtures()` producing:
1. Raw integer boundary fixtures (one `.msgpack` per boundary value)
2. Snapshot boundary fixtures (snapshots with tick values at critical boundaries: 127, 128, 32768, 65536)
Client verifies via `test_boundary_fixture_snapshot_tick_128()`.
### 4.4 Golden File Pipeline
**Consensus between Hoshe and Tyre:** ObserverSnapshot at fixed player positions, not full ECS world state.
Hoshe's pipeline:
1. Server generates canonical snapshots (`cargo test --test gen_gauntlet_golden -- --ignored`)
2. `.msgpack` + `.json` companion (human-readable diff)
3. CI compares fresh generation against committed golden files
4. Field-by-field comparison with actionable error messages (not just "binary files differ")
Tyre adds: **multi-position golden files** — one per room, from the room's designated observer position. ~4-6 golden files for the initial Gauntlet.
**Justine proposes:** Golden files checked into repo (version-controlled alongside code). Changes visible in PR diffs. Sorted JSON for human readability.
### 4.5 Layer 3 Test (Real Subprocess Integration)
**Proposed by:** Hoshe (full code specification)
```
Test: server_subprocess_sends_snapshot_on_connect
Setup: cargo build server, launch as child process with --test-mode --port 0
Steps: Send PlayerInput, read ObserverSnapshot with 5s timeout
Assert: version, tick, entity count, player entity kind
Duration: <10 seconds
```
Requires two server features: `--test-mode` flag (loads Gauntlet, fixed seed, exits after disconnect) and `--port 0` support (random available port printed to stdout).
### 4.6 Pause Guard Tests
**Proposed by:** Dudley
8 test cases for `process_player_input` pause filtering — currently only 1 test exists (`process_input_pause_sets_paused`), which doesn't verify movement is actually discarded.
| Test | Priority |
|------|----------|
| `movement_discarded_while_paused` | P0 — direct bug #3 regression |
| `unpause_accepted_while_paused` | P0 |
| `stance_toggle_allowed_while_paused` | P1 |
| `interact_allowed_while_paused` | P1 |
| `multiple_movements_in_paused_batch_all_discarded` | P1 |
| `pause_unpause_roundtrip_with_movement` | P1 |
### 4.7 EntityRegistry Lifecycle Tests
**Proposed by:** Dudley
5 new test cases for entity lifecycle. Most critical: `old_stable_id_not_resolvable_after_unregister` — prevents a despawned entity's StableId from resolving to a recycled entity, which would corrupt the knowledge graph.
### 4.8 Bridge Deserialization Robustness
**Proposed by:** Dudley
Current `tcp.rs` rejects the ENTIRE `Vec<PlayerInput>` batch if any single input fails to deserialize. Proposal: per-input deserialization with skip-and-log. However, Dudley notes the alternative: since both sides are co-versioned (D-020), batch failure is a hard programming error, and the current behavior is acceptable IF boundary value tests prevent malformed encoding from shipping.
---
## Track 5: Content Scaling & CI Pipeline
### Participants: Hoshe (T5), Justine (T5), Tyre (T5)
### 5.1 CI Pipeline Design
**Consensus between Hoshe and Justine on a 3-tier pipeline:**
| Tier | Trigger | Duration Budget | Contents |
|------|---------|----------------|----------|
| **Commit** | Every push | <2 min | lint-server, lint-client, validate-content, check-fact-ids |
| **PR** (merge gate) | PR opened/updated | <10 min | Commit tier + build both, test-server, test-client, fixture staleness check |
| **Nightly** | Scheduled daily | <30 min | PR tier + Layer 3 subprocess, golden file regen + diff, content load-test, performance benchmark |
Justine provided complete Gitea Actions YAML. Hoshe provided the same tier structure independently.
**CI runner:** Self-hosted strongly recommended (Gitea already self-hosted at `git.schweitz.internal`). Enables stable performance baselines and pre-installed Godot.
**Merge policy** (Justine):
- **BLOCKER:** Server tests, client tests, lint, content validation
- **WARNING:** Golden file changed (requires reviewer ack), fixture regeneration needed
- **INFO:** Performance delta, build size increase
### 5.2 Godot in CI
**Proposed by:** Justine
Download official headless binary (not container). ~80MB cached between runs. `--headless` + `--ignoreHeadlessMode` for gdUnit4. Pre-install on self-hosted runner to eliminate download step.
**Open concern:** Need to verify `test_fog_shader.gd` and `test_rendering.gd` pass headless before gating merges on them.
### 5.3 Performance Regression Detection
**Proposed by:** Justine
Median of 5 runs, relative threshold against committed baseline (`tests/perf/baseline.json`):
- <15% delta: PASS
- 15-30% delta: WARNING
- >30% delta: FAIL
Baseline updated manually in PRs that legitimately change performance. Self-hosted runner essential for reducing noise.
### 5.4 Golden File Diff Format
**Proposed by:** Justine
Structured field-by-field diff with semantic grouping (not raw binary diff, not full dump). Categories: changed fields, added entities, removed entities, unchanged count. Posted as PR comment via `tea comment`.
### 5.5 Content Validation Layers
**Consensus across all three:**
| Layer | What | Status | Owner |
|-------|------|--------|-------|
| 1. Schema validation | YAML structure | Exists (`make validate-content`) | — |
| 2. Cross-reference validation | Entity refs resolve, fact_ids exist | Partial | Tyre proposes `validate_cross_references()` |
| 3. Runtime "boot and tick" | Load content, tick 10, no panic | Missing | Hoshe proposes `content_produces_valid_snapshot()` |
| 4. Regression snapshots | Golden file comparison | Missing | Tied to Gauntlet golden file pipeline |
| 5. Stress test | 100 ticks with max-NPC pack | Missing | Nightly tier |
Hoshe adds **district capacity check** (warn if >15 NPCs, the Crowd Plaza boundary).
### 5.6 Content Scaling Test
**Proposed by:** Hoshe
Comparative testing: load baseline content → tick 10 → snapshot, then load baseline + 1 extra NPC → tick 10 → snapshot. Assert modified is superset of baseline, all original NPCs present, no tick >50ms.
Scaling matrix covers: +1 NPC, +5 NPCs, +15 NPCs, +1 location, +1 item.
---
## Map-Agnostic Invariants
**Proposed by:** Gestalt (36 invariants)
Organized into 4 categories. These must hold for ANY valid map — Gauntlet, procedural, or hand-crafted.
### Structural (INV-S01 through INV-S08)
Player spawn reachable, NPC spawns reachable, NPC routine paths valid, no entity inside geometry, 2x2 geometry minimum (D-066), zone boundary coherence, door bidirectionality.
### Perception (INV-P01 through INV-P05)
Vision cone at spawn (>0 visible tiles), LOS symmetry, fog layer ordering (clear ⊂ peripheral ⊂ deep ⊂ unexplored), insert independence (D-048), sound propagation coherence (sound range ≥ vision range).
### Population (INV-C01 through INV-C08)
Minimum nearby NPC (≥1 within 10 tiles of spawn), social site minimum, StableId uniqueness, KG reference validity, D-033 color validity, monologue trigger reachability, dialogue pool non-empty, perception mode consistency.
### Simulation (INV-T01 through INV-T08)
Deterministic replay, tick budget (≤100ms), snapshot delivery (one per tick, ordered), pause coherence, input ordering, SimRng consumption order, cognitive delay monotonicity, knowledge decay timing.
Tyre proposed 5 of these independently for dynamic map fuzzy testing (reachability, NPC minimum, no overlap, walkability coherence, content reference integrity). Full alignment with Gestalt's larger set.
---
## Decisions Requiring Confirmation
These emerged as consensus positions but require explicit lead sign-off before Round 2 implementation planning.
| # | Decision | Proposed By | Consensus | Dissent |
|---|----------|-------------|-----------|---------|
| **WS-D1** | Commit to determinism now (4 targeted fixes, ~15 lines) | Tyre, Dudley | Unanimous | None — both auditors agree cost is trivial, D-010 mandates it |
| **WS-D2** | Gauntlet as hybrid YAML + Rust inject (not pure YAML or pure Rust) | Tyre, Dudley | Strong (independently converged) | Minor: geometry in YAML (Tyre) vs. geometry in Rust (Dudley) — needs alignment |
| **WS-D3** | Assertion language as Rust macros with named helpers (no custom DSL) | Tyre | Unanimous | None — all agree DSL is premature |
| **WS-D4** | Golden file format as ObserverSnapshot at fixed positions (not full ECS world state) | Tyre, Hoshe, Justine | Unanimous | None |
| **WS-D5** | Text renderer lives server-side, formats ObserverSnapshot to stdout | Tyre, Stig, Dudley | Unanimous | None — Stig explicitly vetoed client-side `--text-mode` |
---
## Open Questions for Round 2
### Cross-Agent Questions (raised during Round 1)
| ID | From | To | Question |
|----|------|----|----------|
| OQ-01 | Tyre | Stig | Does the Godot client depend on `visible_tiles` Vec ordering? (Fix #2 severity) |
| OQ-02 | Tyre | Dudley | Which bevy queries beyond the critical path use iteration order that affects output? |
| OQ-03 | Hoshe | Tyre/Justine | Gitea at `git.schweitz.internal` — are Gitea Actions enabled? Is a self-hosted runner available? |
| OQ-04 | Hoshe | Tyre | Golden file JSON companion format — full ObserverSnapshot or reduced "diff-friendly" format? |
| OQ-05 | Hoshe | Justine | Fixture staleness check (`make fixtures && git diff --exit-code`) — robust enough, or need content-addressed hashing? |
| OQ-06 | Hoshe | Dudley | Does `rmp_serde::from_slice::<u64>()` accept int_16-encoded positive values (256-32767)? (GDScript/Rust encoding asymmetry) |
| OQ-07 | Hoshe | Dudley | Server binary — does it support `--test-mode` and `--port 0` flags? Layer 3 depends on this. |
| OQ-08 | Stig | Dudley | Text renderer needs room name from player coordinates. Does Gauntlet content include room bounds? |
| OQ-09 | Stig | Tyre | Text renderer should format from ObserverSnapshot (not raw ECS). Agree? |
| OQ-10 | Stig | Hoshe | Fog texture byte values (255/180/128/0) — should these be named constants in FogState? |
| OQ-11 | Justine | — | Client test stability in headless mode — do `test_fog_shader.gd` and `test_rendering.gd` pass headless? |
| OQ-12 | Dudley | Tyre | `WalkabilityMap.chunks: HashMap` — convert to BTreeMap preemptively, or wait until chunk iteration is needed? |
### Unresolved Architectural Questions
| ID | Question | Raised By | Impact |
|----|----------|-----------|--------|
| UQ-01 | Bridge deserialization: per-input skip-and-log vs. batch failure? | Dudley | Robustness vs. simplicity tradeoff |
| UQ-02 | Gauntlet geometry: YAML vs. Rust WalkabilityMap builder? | Tyre/Dudley | Minor — needs alignment in Round 2 |
| UQ-03 | Text renderer missing data: how to expose wall-blocking-LOS info? | Dudley | New optional field on ObserverSnapshot? Performance cost? |
| UQ-04 | Fixture path `../client/tests/fixtures/msgpack/` — fragile relative path? | Justine | CI robustness |
---
## Gaps Identified
Issues not fully addressed by Round 1 that need exploration in Round 2 or later:
1. **Save/load interaction with determinism.** Deterministic replay enables "serialize inputs + seed" save format. Nobody addressed how this interacts with the save/load architecture (workshop brief listed but not yet held).
2. **Multiplayer implications of determinism.** Tyre mentions "future multiplayer sync" as a benefit but the networking implications aren't explored. Oscar (networking) is not a workshop participant.
3. **Content authoring workflow.** The Gauntlet needs content (NPCs, dialogue, monologue). Who writes it? How is it reviewed? Mellanie (copywriter) is not a participant.
4. **Visual verification beyond text renderer.** Ozzie's debug overlay (F3) and "WRONG" button (F12) are excellent UX proposals but have no implementation specs. These are client-side features (Stig's domain) not yet designed.
5. **Nightly test infrastructure.** The nightly tier (Layer 3, golden file, stress test) requires server binary artifacts in CI. The build/cache/reuse workflow is sketched (Justine) but not specified.
6. **The 5 new rooms (Gestalt) lack entity placement specs.** Gestalt provided layout concepts and named entities but not precise tile coordinates. These need to be pinned before Gauntlet YAML can be written.
---
## Priority-Ordered Action Items (from all participants)
Combined priority ranking across all tracks. Items at the same priority are independent and can be parallelized.
| Priority | Item | Track | Owner(s) | Effort |
|----------|------|-------|----------|--------|
| **P0** | Fix `visible_ids` HashSet → BTreeSet | T2 | Dudley | 3 lines |
| **P0** | Sort `visible_tiles` in ObserverSnapshot | T2 | Dudley | 1 line |
| **P0** | Pin monologue system ordering (`.after()`) | T2 | Dudley | 3 lines |
| **P0** | Bug #3 regression test: `movement_discarded_while_paused` | T4 | Dudley | ~30 lines |
| **P1** | Build Gauntlet hybrid loader (YAML + Rust inject) | T1 | Dudley, Tyre | 2-3 days |
| **P1** | Build `GauntletRunner` test harness | T3 | Tyre | 1-2 days |
| **P1** | Write determinism regression test | T2 | Tyre | 1 day |
| **P1** | Boundary value test matrix (41 values, both languages) | T4 | Hoshe | 1-2 days |
| **P1** | gen_fixtures.rs boundary extension | T4 | Hoshe | 1 day |
| **P1** | Pause guard test suite (8 tests) | T4 | Dudley | 1 day |
| **P1** | EntityRegistry lifecycle tests (5 tests) | T4 | Dudley | 0.5 day |
| **P2** | CI pipeline (Gitea Actions YAML) | T5 | Justine | 1-2 days |
| **P2** | Golden file generation + comparison tool | T5 | Hoshe, Tyre | 1 day |
| **P2** | Text renderer for ObserverSnapshot | T3 | Tyre | 0.5-1 day |
| **P2** | Cross-reference content validator | T5 | Tyre | 2 days |
| **P2** | Client-side assertion suite (32 tests) | T3 | Stig | 2-3 days |
| **P2** | Layer 3 subprocess test | T4 | Hoshe | 1-2 days |
| **P3** | Zone Gate architectural provision | T1 | Tyre | 0.5 day |
| **P3** | Fuzzy invariant tests for dynamic maps | T5 | Tyre | 1 day |
| **P3** | Content scaling test suite | T5 | Hoshe | 1-2 days |
| **P3** | Map-agnostic invariant tests (36 invariants) | T1 | Gestalt spec, impl TBD | 2-3 days |
---
## Points of Agreement (strong consensus)
1. **Determinism is cheap and necessary.** Both auditors confirm ~15 lines of fixes. No dissent.
2. **ObserverSnapshot is the right golden file format.** Unanimously preferred over full ECS world state.
3. **Text renderer lives server-side.** No client-side rendering path. Stig explicit.
4. **Rust assertions, not custom DSL.** Universal agreement that DSL is premature.
5. **Hub-and-spoke Gauntlet layout with cross-cuts.** Ozzie's proposal went unchallenged.
6. **3-tier CI pipeline (commit/PR/nightly).** Hoshe and Justine independently designed the same tiers.
7. **Content validation needs a runtime "boot and tick" layer.** Schema-only is insufficient.
## Points of Tension (minor, resolvable in Round 2)
1. **Gauntlet geometry: YAML vs. Rust.** Tyre says YAML; Dudley says Rust builder. Both agree on the hybrid approach for everything else. This is a narrow question about WalkabilityMap construction.
2. **validate_movement tie-breaking:** Tyre says "document as invariant" (accept either winner). Dudley says "sort by Entity bits" (force deterministic winner). Low stakes — the existing test accepts either.
3. **Bridge batch deserialization:** Dudley raises per-input skip-and-log as desirable but acknowledges the current batch-failure behavior is acceptable with test coverage. No strong disagreement.
4. **Gauntlet room count:** The workshop brief proposes 7 rooms + 1 reserved. Gestalt adds 5 more + 1 stress room. Total would be 13-14 rooms. Scope needs to be managed — which rooms are Sprint 7-8 vs. later?