Files
settled-reach/docs/workshops/test-architecture/tyre-round3.md
T
jpmschweitzerandClaude Opus 4.6 563a295a90 docs(docs): add frontmatter to test-architecture workshop
Standardized YAML frontmatter on all 20 files.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-14 00:01:23 +01:00

44 KiB

title, description, type, status, workshop, agent, round, created
title description type status workshop agent round created
Tyre — Round 3: Sprint 8 Implementation Plan, Test Client Final Spec, Roadmap Tyre's sprint 8 implementation plan with final test client spec and multi-sprint roadmap workshop archived test-architecture tyre 3 2026-02-17

Tyre — Round 3: Sprint 8 Implementation Plan, Test Client Final Spec, Roadmap

Workshop: QA Strategy & Test Architecture Round: 3 (Prioritization & Final Specs) Date: 2026-02-17 Inputs: All 7 agent Round 2 outputs + Qatux round-2-notes.md synthesis OVERRIDE APPLIED: Lead overruled server/src/bin/ placement. Test client binary moves to tooling/test-client/ as a separate Rust crate.


0. Lead Override: Test Client Location — Architectural Analysis

The lead overruled my Round 2 recommendation of server/src/bin/test_client.rs. The test client binary goes in tooling/test-client/ as a standalone Rust crate. Let me be honest about what this means technically.

What Changes

Aspect Round 2 (server/src/bin/) Round 3 (tooling/test-client/)
Type sharing Free — same crate Path dependency: settled-reach-server = { path = "../../server" }
Build command cargo build --bin settled-reach-test-client cd tooling/test-client && cargo build
Binary location server/target/debug/settled-reach-test-client tooling/test-client/target/debug/settled-reach-test-client
Compile time Zero incremental cost (shared compilation) First build pulls in full server dep tree (~30-60s extra for bevy). Incremental ~2-5s.
Cross-binary tests cargo test in server builds both binaries Requires Makefile orchestration to build both
Crate management None New Cargo.toml, new Cargo.lock

Why This Works

The server crate already has a lib.rs that publicly exports everything the test client needs:

// server/src/lib.rs — already exists
pub mod bridge;      // → bridge::types::*, bridge::framing::*
pub mod knowledge;   // → knowledge::types::RelationshipState, EntityVisibility, etc.
pub mod simulation;  // → simulation::time::DayPhase, TickRate

The bridge types used by the test client (ObserverSnapshot, PlayerInput, read_framed, write_framed) are already pub. No server-side changes needed for the test client to import them.

Dependency Cost

The test client crate will transitively depend on the entire server dependency tree:

settled-reach-test-client
  └─ settled-reach-server (path)
       ├─ bevy_ecs 0.18
       ├─ bevy_app 0.18
       ├─ rmp-serde 1
       ├─ serde 1
       ├─ pathfinding 4
       ├─ rand 0.9
       └─ ... (~15 transitive deps)

The test client only needs serde, rmp-serde, and the wire types. The bevy dependency is dead weight — pulled in because ObjectType derives Component and SnapshotBuffer derives Resource in bridge/types.rs.

Sprint 8 pragmatic choice: Accept the heavy dependency. Bevy compiles once and caches. Incremental test client builds are fast (~2-5s). Binary size is larger than necessary but DCE removes unused code.

Sprint 9+ option: Extract a settled-reach-protocol crate containing only bridge types + framing (no bevy). Both server and test client depend on it. This is the right long-term answer but premature for Sprint 8. Added to roadmap as R-25.

Why the Override Is Architecturally Sound

Despite my Round 2 preference, the lead's reasoning holds:

  1. tooling/ establishes a clear convention — project tools that aren't the game server or game client live here. The test client is a development tool, not a game component.
  2. Separation prevents accidental coupling. A server/src/bin/ test client could accidentally use server internals (private modules, internal state). A separate crate can only use what lib.rs exports.
  3. Independent release cycle. The test client can version independently and add its own dependencies (crossterm, clap) without affecting the server's Cargo.toml.
  4. Existing pattern. tooling/content-converter/ and tooling/line-previewer/ are already standalone Rust crates in this directory.

1. Sprint 8 Implementation Plan

Sprint 8 ships infrastructure — the plumbing. The Gauntlet rooms are content that flows through the plumbing. Sprint 8 ships the pipes, Sprint 9 fills them.

Dependency Graph

S8-1: Determinism fixes ──────────────────────────────┐
                                                        ├─→ S8-5: make pre-pr
S8-2: Content cross-ref validation ───────────────────┤
                                                        │
S8-3: --test-mode + --port 0 ──→ S8-4: Test client ───┤
                                  binary MVP            │
                                       │                │
                                       └─→ S8-6: Layer 3 test wiring
                                                        │
S8-7: Pause guard tests ─────────────────────────────┘ (parallel, no deps)
S8-8: Determinism regression tests ── (after S8-1)
S8-9: EntityRegistry lifecycle tests ── (parallel, no deps)
S8-10: Fixture staleness check ── (after S8-5)

Ordered Implementation

# Item Owner Effort Depends On Deliverable
S8-1 Determinism fixes (A, B, D) Dudley 0.5 day Nothing 3 fixes, ~40 lines across 4 files. Fix C already done.
S8-2 Content cross-reference validation Justine 1 day Nothing 9 checks added to tooling/validate-content. Python only.
S8-3 --test-mode + --port 0 server flags Dudley 0.5 day Nothing Modified main.rs: flag parsing, LISTENING:{port} stdout, fixed seed 42, accept_on(listener).
S8-4 Test client binary MVP Dudley 2.5-3.5 days S8-3 tooling/test-client/ — new crate, connect, receive, text render, golden compare, replay. See Section 2 for full spec. Increased 0.5d from R2 estimate for crate setup + dependency wiring.
S8-5 make pre-pr chain Justine 0.5 day S8-1, S8-2 Makefile target: lint → build → test → validate → fixtures. Plus branch-specific variants.
S8-6 Layer 3 test wiring Dudley 0.5 day S8-3, S8-4 make test-layer3 — builds both binaries, runs subprocess integration test.
S8-7 Pause guard tests (6 tests) Dudley 0.5 day Nothing All 6 Hoshe-validated gaps. P0: movement_discarded_while_paused, unpause_accepted_while_paused.
S8-8 Determinism regression tests Dudley 0.5 day S8-1 Per-fix regression tests: sorted tiles, sorted entities, deterministic mover winner.
S8-9 EntityRegistry lifecycle tests (3 tests) Dudley 0.25 day Nothing P0: register_respawn_no_stale_mapping. P1: register_after_unregister_gets_new_id, unregister_unknown_entity_is_noop.
S8-10 Fixture staleness check Justine 0.25 day S8-5 make fixtures-check target. BLOCKER policy (not WARNING).

Critical Path

S8-3 (0.5d) → S8-4 (2.5-3.5d) → S8-6 (0.5d) = 3.5-4.5 days

The test client binary remains the long pole, slightly longer due to new crate setup. Everything else proceeds in parallel.

Sprint 8 Total Effort

Track Items Effort
Server (Dudley) S8-1, S8-3, S8-4, S8-6, S8-7, S8-8, S8-9 ~5.5-7 days
Tooling (Justine) S8-2, S8-5, S8-10 ~1.75 days
Total 10 items ~7.5-9 team-days

No client (Stig) work in Sprint 8 for this track. Client test writing starts Sprint 9 once fixtures are flowing through make pre-pr.

What Explicitly Does NOT Ship in Sprint 8

  • Gauntlet room content (rooms 1-14)
  • Anti-tedium features (room reset, hub teleport, WRONG button)
  • Crossterm live terminal display
  • CI automation (Gitea Actions)
  • Performance baselines
  • Encoding asymmetry cross-language tests
  • Client tests (Stig's 38+)
  • Cross-room transition scenarios
  • Protocol crate extraction

All Sprint 9+ scope. See Section 3.


2. Test Client Binary — Final Build-Ready Specification

This is the definitive spec. It combines Tyre R2 architecture + Ozzie R2 UX research + Dudley R2 server-side requirements + the lead's override on binary location. Implementers should treat this as the contract.

2.1 Crate Location and Structure

Location: tooling/test-client/

tooling/test-client/
├── Cargo.toml
├── Cargo.lock
└── src/
    ├── main.rs           # CLI parsing, connection loop, output dispatch
    ├── golden.rs         # Golden file comparison (JSON field diff)
    └── replay.rs         # JSONL replay file loading

tooling/test-client/Cargo.toml:

[package]
name = "settled-reach-test-client"
version = "0.1.0"
edition = "2021"
description = "Headless test client for Gauntlet QA verification. Connects to the game server via TCP, receives ObserverSnapshots, renders as text, compares against golden files."

[[bin]]
name = "settled-reach-test-client"
path = "src/main.rs"

[dependencies]
# Server crate provides bridge types + framing
settled-reach-server = { path = "../../server" }

# CLI
clap = { version = "4", features = ["derive"] }

# Golden file comparison (JSON)
serde_json = "1"
serde = { version = "1", features = ["derive"] }

# Wire protocol (shared with server — version must match)
rmp-serde = "1"

Build: cd tooling/test-client && cargo build

Or from project root via Makefile:

build-test-client:
	cd tooling/test-client && cargo build

build-test-client-release:
	cd tooling/test-client && cargo build --release

2.2 Import Paths

The test client imports shared types from the server crate's library:

// tooling/test-client/src/main.rs

// Wire types
use settled_reach_server::bridge::types::{
    ObserverSnapshot, PlayerInput, PlayerAction,
    VisibleEntity, EntityKind, VisibilitySector,
    GameTime, MovementStance, NearbyInteraction,
    MonologueEvent, PendingRecognitionWire,
};
use settled_reach_server::knowledge::types::{
    RelationshipState, EntityVisibility,
};
use settled_reach_server::simulation::time::{DayPhase, TickRate};

// Framing protocol
use settled_reach_server::bridge::framing::{read_framed, write_framed};

// Text renderer (library function in server crate)
use settled_reach_server::bridge::text_renderer::format_snapshot_text;

Dudley action required: Ensure server/src/bridge/mod.rs adds pub mod text_renderer; when the text renderer is implemented. All other exports already exist.

2.3 CLI Interface (Sprint 8 MVP)

settled-reach-test-client [OPTIONS]

CONNECTION:
  --connect <host:port>     Server address (default: 127.0.0.1:9876)

INPUT:
  --replay <file.jsonl>     Send inputs from file (one JSON PlayerInput per line)
                            Empty lines = idle tick (no input sent)

OUTPUT (mutually exclusive):
  --text                    Render each snapshot as structured text to stdout (default)
  --json                    Dump each snapshot as JSON to stdout (for golden file generation)
  --quiet                   No output, assertions only (for CI)

ASSERTIONS:
  --golden <file.json>      Compare FINAL snapshot against golden file, exit 1 on diff
  --ticks <N>               Disconnect after N ticks (default: unlimited)

EXIT CODES:
  0                         Success (all assertions passed, or no assertions)
  1                         Golden file mismatch (diff printed to stderr)
  2                         Connection error or protocol error

Sprint 9+ CLI additions (NOT Sprint 8):

  --interactive             Read inputs from stdin (human-in-the-loop)
  --live                    Crossterm live-updating terminal display (Ozzie spec)
  --log <file>              Append text output to session log
  --history-buffer <N>      Ring buffer depth for WRONG captures (default: 60)
  --checklist <file.yaml>   Load checklist for auto-tracking

2.4 Connection Protocol

┌─────────────────────────┐              ┌──────────────────────────────┐
│ settled-reach-server    │     TCP      │ settled-reach-test-client    │
│                         │◄────────────►│                              │
│ --test-mode --port 0    │   framed     │ --connect 127.0.0.1:{port}  │
│ --seed 42               │   msgpack    │ --replay inputs.jsonl        │
│                         │              │ --ticks 50 --golden f.json   │
└─────────────────────────┘              └──────────────────────────────┘
 server/target/debug/                     tooling/test-client/target/debug/
 settled-reach-server                     settled-reach-test-client

Startup sequence:

  1. Server binds TCP socket, prints LISTENING:{port}\n to stdout, flushes, blocks on accept
  2. Test client reads LISTENING:{port} from server's stdout (when spawned by Layer 3 test) or connects to known port via --connect
  3. Test client calls TcpStream::connect(addr)
  4. Server's accept_on(listener) returns — game loop starts

Per-tick protocol:

  1. Server runs one tick of simulation
  2. Server sends ObserverSnapshot via write_framed (length-prefixed MessagePack)
  3. Test client receives snapshot via read_framed + rmp_serde::from_slice::<ObserverSnapshot>
  4. Test client formats output (text/json/quiet)
  5. If --replay: test client sends next line's Vec<PlayerInput> via write_framed + rmp_serde::to_vec
  6. If no more replay lines: send empty Vec<PlayerInput> (idle tick)

Important: Use rmp_serde::to_vec (array format), NOT to_vec_named (map format). This matches GDScript's encoding. Dudley confirmed in R2 cross-review.

Shutdown:

After --ticks N snapshots received:

  1. If --golden: compare final snapshot, print diff to stderr if mismatch, exit 1
  2. Drop TCP connection (triggers server shutdown in --test-mode)
  3. Exit 0

2.5 Text Output Format (Sprint 8 MVP)

=== Tick 42 | Player (15,10) facing East | Stance: Walk | TickRate: Full ===
Game time: Day 0, 04:12 (Morning)
Room: Occlusion Corridor
Entities (5):
  npc:100     (18,10) Forward  rel:Neutral     vis:Visible      d=3
  npc:101     (20,10) Forward  rel:Unknown     vis:Remembered   d=5
  obj:200     (16,9)  Forward  rel:n/a         vis:Visible      d=1
  npc:102     (12,8)  Periph   rel:Hostile      vis:Visible      d=5
  npc:103     (22,14) Periph   rel:Friendly     vis:Visible      d=9
Pending recognitions: 1 [npc:104 at (19,12) 3/8 ticks]
Tiles: 31 visible
Interactions (2):
  npc:100 [Talk(1), ExamineNpc(2)] distance=3
  obj:200 [Read(1), Observe(2)] distance=1
Inventory: 2/9 [item:300(slot-0), item:301(slot-3)]
Monologue: "Something about this manifest doesn't add up."
===

Format rules:

Rule Detail
Entity labels kind:entity_id (e.g., npc:100, obj:200). No display names on wire.
Entity sort By distance from player (nearest first). Ties broken by entity_id.
Tick separators === lines for clean diff between ticks.
Positions Integer tile coords. f32 render offset is irrelevant for testing.
Room name From room_at(player_pos) using Gauntlet coordinate constants. Shows Room: (unknown) if outside all room bounds.
Sector labels Forward, Periph — mapped from VisibilitySector enum. (Behind entities are absent from snapshot.)
Relationship labels Unknown, Known, Friendly, PersonOfInterest, Hostile. Mapped from RelationshipState.
Visibility labels Visible, Remembered, Fogged. Mapped from EntityVisibility.
Distance Manhattan distance in tiles from player to entity.
Missing data Sections with no data are omitted (no Interactions (0): noise).
Pending recognitions From ObserverSnapshot.pending_recognitions — shows entity, position, remaining/total ticks.

2.6 Text Renderer — Library Function in Server Crate

The text renderer lives in the server crate's library (not in the test client). This is deliberate:

  • Server integration tests can call it for debug output
  • The test client imports it via the path dependency
  • The server binary never references it — zero bloat

File: server/src/bridge/text_renderer.rs

// server/src/bridge/text_renderer.rs
// Library code — callable by test client crate AND server integration tests.
// The server binary never references this module.

use std::fmt::Write;
use crate::bridge::types::*;
use crate::knowledge::types::{EntityVisibility, RelationshipState};

/// Format an ObserverSnapshot as structured text for human verification.
pub fn format_snapshot_text(snapshot: &ObserverSnapshot) -> String {
    let mut out = String::with_capacity(2048);

    // Find player entity for position reference
    let player = snapshot.entities.iter()
        .find(|e| matches!(e.kind, EntityKind::Player));
    let (px, py) = player
        .map(|p| (p.x as i32, p.y as i32))
        .unwrap_or((-1, -1));

    // Header
    writeln!(out, "=== Tick {} | Player ({},{}) facing {:?} | Stance: {:?} | TickRate: {:?} ===",
        snapshot.tick, px, py,
        snapshot.player_facing,
        snapshot.player_stance,
        snapshot.game_time.tick_rate,
    ).ok();

    // Game time
    let minutes = snapshot.game_time.time_of_day % 60;
    let hours = (snapshot.game_time.time_of_day / 60) % 24;
    writeln!(out, "Game time: Day {}, {:02}:{:02} ({:?})",
        snapshot.game_time.day, hours, minutes,
        snapshot.game_time.day_phase,
    ).ok();

    // Room name (from Gauntlet constants — returns "(unknown)" if not in any room)
    #[cfg(feature = "test-world")]
    {
        let room_name = crate::test_world::constants::room_at_position(px, py)
            .map(|r| r.name)
            .unwrap_or("(unknown)");
        writeln!(out, "Room: {}", room_name).ok();
    }

    // Non-player entities sorted by distance then entity_id
    let mut entities: Vec<&VisibleEntity> = snapshot.entities.iter()
        .filter(|e| !matches!(e.kind, EntityKind::Player))
        .collect();
    entities.sort_by_key(|e| {
        let dist = (e.x as i32 - px).unsigned_abs() + (e.y as i32 - py).unsigned_abs();
        (dist, e.entity_id)
    });

    if !entities.is_empty() {
        writeln!(out, "Entities ({}):", entities.len()).ok();
        for e in &entities {
            let dist = (e.x as i32 - px).unsigned_abs() + (e.y as i32 - py).unsigned_abs();
            writeln!(out, "  {}:{:<8} ({},{}) {:<8} rel:{:<16} vis:{:<12} d={}",
                kind_label(e.kind), e.entity_id,
                e.x as i32, e.y as i32,
                sector_label(e.visibility),
                relationship_label(e.relationship),
                observation_label(e.observation),
                dist,
            ).ok();
        }
    }

    // Pending recognitions
    if !snapshot.pending_recognitions.is_empty() {
        write!(out, "Pending recognitions: {}", snapshot.pending_recognitions.len()).ok();
        for pr in &snapshot.pending_recognitions {
            write!(out, " [npc:{} at ({},{}) {}/{} ticks]",
                pr.entity_id, pr.x as i32, pr.y as i32,
                pr.total_delay_ticks - pr.remaining_ticks, pr.total_delay_ticks,
            ).ok();
        }
        writeln!(out).ok();
    }

    // Tiles
    writeln!(out, "Tiles: {} visible", snapshot.visible_tiles.len()).ok();

    // Interactions
    if !snapshot.nearby_interactions.is_empty() {
        writeln!(out, "Interactions ({}):", snapshot.nearby_interactions.len()).ok();
        for ni in &snapshot.nearby_interactions {
            let verbs: Vec<String> = ni.verbs.iter()
                .map(|v| format!("{}({})", v.label, v.priority))
                .collect();
            writeln!(out, "  {}:{} [{}] distance={}",
                kind_label(ni.entity_type), ni.entity_id,
                verbs.join(", "), ni.distance,
            ).ok();
        }
    }

    // Inventory
    if !snapshot.player_inventory.is_empty() {
        let slots: Vec<String> = snapshot.player_inventory.iter()
            .map(|item| format!("{}(slot-{})", item.name, item.slot))
            .collect();
        writeln!(out, "Inventory: {}/9 [{}]",
            snapshot.player_inventory.len(), slots.join(", "),
        ).ok();
    }

    // Monologue
    if let Some(ref mono) = snapshot.current_monologue {
        writeln!(out, "Monologue: \"{}\"", mono.text).ok();
    }

    writeln!(out, "===").ok();
    out
}

fn kind_label(kind: EntityKind) -> &'static str {
    match kind {
        EntityKind::Player => "player",
        EntityKind::Npc => "npc",
        EntityKind::Object => "obj",
        EntityKind::Terrain => "terrain",
    }
}

fn sector_label(sector: VisibilitySector) -> &'static str {
    match sector {
        VisibilitySector::Forward => "Forward",
        VisibilitySector::Peripheral => "Periph",
    }
}

fn relationship_label(rel: RelationshipState) -> &'static str {
    match rel {
        RelationshipState::Unknown => "Unknown",
        RelationshipState::Known => "Known",
        RelationshipState::Friendly => "Friendly",
        RelationshipState::PersonOfInterest => "POI",
        RelationshipState::Hostile => "Hostile",
    }
}

fn observation_label(obs: EntityVisibility) -> &'static str {
    match obs {
        EntityVisibility::Visible => "Visible",
        EntityVisibility::Remembered => "Remembered",
        EntityVisibility::Fogged => "Fogged",
    }
}

Dudley action items for server crate:

  1. Add pub mod text_renderer; to server/src/bridge/mod.rs
  2. Add pub mod test_world; to server/src/lib.rs (for Gauntlet constants — can be feature-gated behind #[cfg(feature = "test-world")] if desired)

2.7 Golden File Comparison

Format: JSON with sorted keys, pretty-printed.

Generation:

# Generate golden file for tick 5 of the proof room (no input)
cd tooling/test-client
cargo run -- --connect 127.0.0.1:9876 --ticks 5 --json | jq -S . > ../../tests/fixtures/golden/proof_room_tick5.json

Comparison algorithm:

// tooling/test-client/src/golden.rs

use settled_reach_server::bridge::types::ObserverSnapshot;
use serde_json::Value;
use std::path::Path;

pub struct FieldDiff {
    pub path: String,
    pub expected: String,
    pub actual: String,
}

pub fn compare_golden(actual: &ObserverSnapshot, golden_path: &Path) -> Result<Vec<FieldDiff>, String> {
    let golden_str = std::fs::read_to_string(golden_path)
        .map_err(|e| format!("failed to read golden file: {}", e))?;
    let golden: ObserverSnapshot = serde_json::from_str(&golden_str)
        .map_err(|e| format!("failed to parse golden file: {}", e))?;

    let actual_json = serde_json::to_value(actual)
        .map_err(|e| format!("failed to serialize actual: {}", e))?;
    let golden_json = serde_json::to_value(&golden)
        .map_err(|e| format!("failed to serialize golden: {}", e))?;

    Ok(diff_values(&actual_json, &golden_json, String::new()))
}

fn diff_values(actual: &Value, expected: &Value, path: String) -> Vec<FieldDiff> {
    if actual == expected {
        return vec![];
    }
    match (actual, expected) {
        (Value::Object(a), Value::Object(e)) => {
            let mut diffs = Vec::new();
            for (key, eval) in e {
                let child_path = format!("{}.{}", path, key);
                match a.get(key) {
                    Some(aval) => diffs.extend(diff_values(aval, eval, child_path)),
                    None => diffs.push(FieldDiff {
                        path: child_path,
                        expected: format!("{}", eval),
                        actual: "MISSING".into(),
                    }),
                }
            }
            for key in a.keys() {
                if !e.contains_key(key) {
                    diffs.push(FieldDiff {
                        path: format!("{}.{}", path, key),
                        expected: "ABSENT".into(),
                        actual: format!("{}", a[key]),
                    });
                }
            }
            diffs
        }
        (Value::Array(a), Value::Array(e)) => {
            let mut diffs = Vec::new();
            let len = a.len().max(e.len());
            for i in 0..len {
                let child_path = format!("{}[{}]", path, i);
                match (a.get(i), e.get(i)) {
                    (Some(av), Some(ev)) => diffs.extend(diff_values(av, ev, child_path)),
                    (Some(av), None) => diffs.push(FieldDiff {
                        path: child_path, expected: "ABSENT".into(), actual: format!("{}", av),
                    }),
                    (None, Some(ev)) => diffs.push(FieldDiff {
                        path: child_path, expected: format!("{}", ev), actual: "MISSING".into(),
                    }),
                    (None, None) => {}
                }
            }
            diffs
        }
        _ => vec![FieldDiff {
            path,
            expected: format!("{}", expected),
            actual: format!("{}", actual),
        }],
    }
}

Diff output format (to stderr on mismatch):

GOLDEN FILE MISMATCH: proof_room_tick5.json
  .entities[2].x: expected 18, got 19 (CHANGED)
  .entities[2].relationship: expected "Neutral", got "Hostile" (CHANGED)
  .visible_tiles: 31 expected, 29 actual
    [29]: expected {"x":22,"y":10,...}, got MISSING
    [30]: expected {"x":22,"y":11,...}, got MISSING
  .current_monologue: expected {"text":"Quiet shift.",...}, got MISSING

Why JSON not MessagePack: Human-readable in git diff. Sorted keys = deterministic output. JSON round-trips through serde_json without precision loss for ObserverSnapshot types (u64, f32, strings, enums).

2.8 Replay File Format

Format: JSONL (one JSON line per tick)

Each line is a JSON array of PlayerInput objects:

[{"tick":0,"action":"MoveNorth"}]
[{"tick":1,"action":"MoveNorth"}]
[]
[{"tick":3,"action":{"Interact":{"target_entity_id":100,"verb":"Talk"}}}]
[{"tick":4,"action":"ToggleStanceUp"}]

Empty array [] = idle tick (no player input).

Generation: Manual creation for targeted tests. Future: record from Godot client or interactive test client.

2.9 Layer 3 Integration Test

With the test client in a separate crate, Layer 3 testing requires building two independent binaries. A Makefile target handles orchestration:

# Makefile (project root)

# Build both binaries needed for Layer 3 testing
build-layer3: build-server build-test-client

# Run the Layer 3 subprocess integration test
test-layer3: build-layer3
	cd tooling/test-client && cargo test --test layer3 -- --ignored

The Layer 3 test lives in the test client crate (not the server crate) because the test client crate already depends on the server crate, giving access to shared types:

// tooling/test-client/tests/layer3.rs

use std::process::{Command, Stdio};
use std::io::{BufRead, BufReader};
use std::path::PathBuf;

/// Locate the server binary. Built via `make build-server`.
fn server_binary() -> PathBuf {
    // Navigate from tooling/test-client/ up to project root, then to server binary
    let project_root = PathBuf::from(env!("CARGO_MANIFEST_DIR"))
        .join("..").join("..");
    let binary = project_root.join("server/target/debug/settled-reach-server");
    assert!(binary.exists(),
        "Server binary not found at {:?}. Run `make build-server` first.", binary);
    binary
}

/// Locate the test client binary. Built via `make build-test-client`.
fn test_client_binary() -> PathBuf {
    let project_root = PathBuf::from(env!("CARGO_MANIFEST_DIR"))
        .join("..").join("..");
    let binary = project_root
        .join("tooling/test-client/target/debug/settled-reach-test-client");
    assert!(binary.exists(),
        "Test client binary not found. Run `make build-test-client` first.");
    binary
}

#[test]
#[ignore] // Slow — run via `make test-layer3`
fn server_and_test_client_subprocess_roundtrip() {
    // 1. Launch server with --test-mode --port 0
    let mut server = Command::new(server_binary())
        .args(["--test-mode", "--port", "0"])
        .stdout(Stdio::piped())
        .stderr(Stdio::null())
        .spawn()
        .expect("failed to launch server");

    // 2. Read port from server stdout
    let stdout = server.stdout.take().unwrap();
    let mut reader = BufReader::new(stdout);
    let mut line = String::new();
    reader.read_line(&mut line).expect("failed to read LISTENING line");
    let port: u16 = line.trim()
        .strip_prefix("LISTENING:")
        .expect("expected LISTENING:{port}")
        .parse()
        .expect("invalid port");

    // 3. Launch test client: connect, receive 5 ticks, compare golden
    let project_root = PathBuf::from(env!("CARGO_MANIFEST_DIR"))
        .join("..").join("..");
    let golden = project_root.join("tests/fixtures/golden/proof_room_tick5.json");

    let client_status = Command::new(test_client_binary())
        .args([
            "--connect", &format!("127.0.0.1:{}", port),
            "--ticks", "5",
            "--golden", golden.to_str().unwrap(),
            "--quiet",
        ])
        .status()
        .expect("failed to launch test client");

    // 4. Assert test client passed
    assert!(client_status.success(), "test client golden file mismatch (exit code: {})",
        client_status.code().unwrap_or(-1));

    // 5. Server should exit after client disconnect (--test-mode behavior)
    let server_exit = server.wait().expect("failed to wait for server");
    assert!(server_exit.success(), "server exited with error");
}

2.10 File Inventory

File Sprint Description
tooling/test-client/Cargo.toml 8 New crate: depends on server lib, clap, serde_json
tooling/test-client/src/main.rs 8 CLI parsing, TCP connection, tick loop, output dispatch
tooling/test-client/src/golden.rs 8 Golden file JSON comparison with field-by-field diff
tooling/test-client/src/replay.rs 8 JSONL replay file loading
tooling/test-client/tests/layer3.rs 8 Subprocess integration test (S8-6)
server/src/bridge/text_renderer.rs 8 format_snapshot_text() library function (callable by test client + server tests)
server/src/bridge/mod.rs 8 Add pub mod text_renderer;
server/src/test_world/constants.rs 8 GauntletRoom struct, ROOMS array, room_at_position() function
tests/fixtures/golden/proof_room_tick5.json 8 First golden file (proof room, no input, 5 ticks)

2A. tooling/ Directory Structure

The lead asked for a proposed structure. Here's the current state and the proposed additions.

Current State

tooling/
├── .gitkeep
├── check-fact-ids              # Shell script — validates fact IDs in content
├── content-converter/          # Standalone Rust crate — YAML→RON conversion
│   ├── Cargo.toml
│   ├── Cargo.lock
│   └── src/
│       ├── main.rs
│       └── types.rs
├── db-backup                   # Shell script — database backup
├── db-install                  # Shell script — database setup
├── install-godot               # Shell script — Godot installation
├── install-rust                # Shell script — Rust toolchain setup
├── line-previewer/             # Standalone Rust crate — dialogue/monologue preview
│   ├── Cargo.toml
│   ├── Cargo.lock
│   └── src/
│       ├── main.rs
│       ├── pipeline.rs
│       └── types.rs
├── synth_ui_sounds.py          # Python script — audio synthesis
└── validate-content            # Python script — content schema validation

Convention (Already Established)

  • Directories = Rust crates (each with own Cargo.toml and Cargo.lock)
  • Files = Standalone scripts (shell, Python — no build step)

This convention is clean and consistent. No restructure needed.

Proposed Addition

tooling/
├── ... (all existing files unchanged)
├── test-client/                # NEW — Rust crate, depends on server lib
│   ├── Cargo.toml
│   ├── Cargo.lock
│   └── src/
│       ├── main.rs
│       ├── golden.rs
│       └── replay.rs
└── ...

The test client follows the same standalone-crate pattern as content-converter/ and line-previewer/. It differs in one way: it has a path dependency on the server crate (../../server), while the other two crates are fully independent. This is the architectural trade-off of sharing bridge types instead of duplicating them.

Future Considerations

Milestone Potential Change
Sprint 9+ (if compile times hurt) Extract settled-reach-protocol crate to tooling/protocol/. Both server and test-client depend on it. Eliminates bevy transitive dep for test client.
When Cargo workspace makes sense Root Cargo.toml with workspace.members = ["server", "tooling/test-client", "tooling/protocol"]. Shared target/ directory, shared Cargo.lock, faster builds.
When more Rust tools arrive Consider if the standalone pattern still scales. Workspace becomes more attractive at 4+ crates.

No action needed now. The standalone pattern works for 3 Rust crates.


3. Sprint 9+ Roadmap

Everything that's NOT Sprint 8, ranked by implementation value. Tier 1 = next sprint. Tier 2 = sprint after. Tier 3 = when needed.

Tier 1: Sprint 9 (High Value — Unblocked by Sprint 8)

# Item Owner Effort Depends On Value
R-01 Gauntlet rooms 1-4 (Inventory, Occlusion, Interaction, Crowd) Content + Dudley 3-4 days S8-3 (--test-mode) Unlocks room-specific testing. These 4 cover the broadest system range.
R-02 Room reset trigger Dudley 1.5 days R-01 Dudley's R2 design is implementation-ready. RoomResetTrigger component, RoomSnapshots resource, execute_room_reset system.
R-03 Hub teleport Dudley 0.5 day R-01 PlayerAction::TeleportToHub, instant camera snap. Simplest anti-tedium feature.
R-04 Client tests P0-P1 (10 tests) Stig 2-3 days S8-10 (fixtures) Bug #5 monologue regression, Bug #2 camera, fog shader (3), entity lifecycle, pending recognition blob. Highest-value client coverage.
R-05 Determinism golden files (per-room) Dudley 1 day R-01, S8-4 One golden file per room: 10-tick idle snapshot. The "did anything break?" safety net.
R-06 Fog byte constants Stig 0.25 day Nothing VIS_HIDDEN=0, VIS_PERIPHERAL=180, VIS_FORWARD=255, EXP_UNEXPLORED=0, EXP_EXPLORED=128, EXP_VISIBLE=255. Replace magic numbers in fog shader.
R-07 Encoding asymmetry tests (4-direction) Hoshe + Stig 1.5 days S8-10 Hoshe's R2 spec: Rust→GDScript fixtures, GDScript→Rust fixtures, raw byte round-trips both directions.

Sprint 9 total: ~10-12 team-days

Tier 2: Sprint 10 (Medium Value)

# Item Owner Effort Value
R-08 WRONG button MVP Stig + Dudley 2 days Bug capture during manual testing. F12 hotkey, snapshot + text dump, human description prompt.
R-09 Gauntlet rooms 5-8 (Fog Theater, Dialogue, Pause, Zone Gate) Content + Dudley 3-4 days Perception + dialogue system testing. Zone Gate is reserved/stub.
R-10 Room timer + personal bests Ozzie spec, Stig impl 1 day Anti-tedium. Session stats to tests/gauntlet-stats.json.
R-11 Checklist auto-tracking Stig + Dudley 2 days Merged Ozzie+Stig YAML format. make checklist generates markdown. Test client loads conditions.
R-12 Performance baselines Justine 1.5 days tests/perf/baseline.json, tooling/perf-measure, 15%/30% thresholds. Machine-tagged.
R-13 Client tests P2 (16 tests) Stig 3 days Remaining camera (5), entity alpha+color (4), UI elements (7).
R-14 Enhanced test client terminal Dudley 2 days Ozzie's crossterm live-updating layout. Sound, cognition, checklist sections.

Sprint 10 total: ~15-17 team-days (may split across 2 sprints)

Tier 3: Sprint 11+ (Build When Needed)

# Item Effort Trigger
R-15 Gauntlet rooms 9-14 (Eavesdrop, Confrontation, Sprint, Sound Lab, Decay, Shift Change) 4-6 days When audio + cognitive delay systems are implemented
R-16 Cross-room transition scenarios (T1-T8) 2 days When cross-cuts are built (see Q3 answer below)
R-17 CI automation (Gitea Actions) 1 day When lead greenlights. Workflow runs make ci. PR tier first.
R-18 Content scaling stress tests 1 day When content volume exceeds proof-of-concept size
R-19 blocked_entities debug field on ObserverSnapshot 1 day When LOS debugging becomes a bottleneck for testers
R-20 Client test headless stability (OQ-11) 0.5 day Before making client tests a merge gate in CI
R-21 Client tests P3 (12 tests) 2 days Z-layer ordering (4), entity lerp (3), remaining additions (5)
R-22 Hoshe's 3 additional pause guard tests 0.5 day P2: set_tick_rate_while_paused, perception_mode_while_paused, interact_take_while_paused
R-23 WRONG button full capture (ring buffer, replay seed, world digest) 1.5 days When human testers file enough bugs to justify the investment
R-24 Content cross-reference bidirectional relationship warnings 0.5 day When relationship asymmetries cause real content bugs
R-25 Extract settled-reach-protocol crate 1-2 days When test client compile times become a pain point, or when a 4th Rust crate needs bridge types

4. Remaining Question Answers

Q5: Test Client Binary Location — OVERRULED, NEW ANSWER

Answer: tooling/test-client/ as a standalone Rust crate.

The lead overruled my Round 2 recommendation of server/src/bin/test_client.rs. See Section 0 for the full architectural analysis of why the override works and the trade-offs involved.

The test client crate:

  • Lives at tooling/test-client/ alongside content-converter/ and line-previewer/
  • Has a path dependency on the server crate: settled-reach-server = { path = "../../server" }
  • Imports bridge types, framing functions, and text renderer from the server's library
  • Has its own Cargo.toml, Cargo.lock, target/ directory

Q3: Are 4 Cross-Cuts Too Many for Sprint 8 Gauntlet MVP?

Answer: Sprint 8 has zero cross-cuts. Zero Gauntlet rooms. The question is moot for Sprint 8 — but the map DESIGN should include them.

Let me be honest about what this means technically.

Sprint 8 scope: Sprint 8 ships infrastructure only (see Section 1). The --test-mode flag falls back to the existing proof room. There are no Gauntlet rooms in Sprint 8. Therefore, there are no cross-cuts in Sprint 8.

Sprint 9 scope (first rooms): Rooms 1-4 ship as hub-and-spoke. No cross-cuts needed. Each room tests its own systems independently. The hub is the only navigation between rooms.

Sprint 10+ scope (cross-cuts): Cross-cuts become valuable when we test transitions (Gestalt's T1-T8 scenarios). At that point:

Gestalt's Proposed Cross-Cut Value Sprint
Plaza → Occlusion Corridor (T1: Sprint Exit) HIGH — tests sprint buffer clear + LOS recalc 10
Fog Theater → Dialogue Room (T2: Fog into Dialogue) HIGH — tests state preservation across modes 10
Confrontation → Eavesdrop (T5) MEDIUM — audio system transition 11+
Sound Lab → Fog Theater (T7) LOW — can be tested via hub + 2 teleports 11+

Recommendation: Design all 4 cross-cuts in the Gauntlet map layout NOW (Gestalt's job). Build the first 2 (Plaza→Occlusion, Fog→Dialogue) in Sprint 10 when rooms 5-8 ship. Build the remaining 2 when rooms 9-14 ship. Cross-cuts are additive — adding a doorway between two rooms doesn't change either room's layout.

Gestalt's sub-question about consolidating Sound Lab into Occlusion Corridor: Don't. Sound Lab tests D-018's three-range sound model, which is orthogonal to LOS/shadowcasting. Consolidating them makes the room too complex to isolate failures. Keep them separate.

R2-OQ-03 (from Hoshe): Should make pre-pr include make content-ron?

Answer: Yes, add it to make pre-pr but NOT to make pre-pr-server.

pre-pr: lint build test validate-content content-ron fixtures-check
pre-pr-server: lint-server build-server test-server fixtures-check
pre-pr-client: lint-client build-client test-client
pre-pr-content: validate-content content-ron

Budget impact: content-ron takes <5s for current content volume. Negligible.

R2-OQ-09 (from Gestalt): Room ordering affects entity StableId assignment?

Answer: Yes, and it's handled by spawn order in the setup function.

Gauntlet entity spawn order must be deterministic. Since --test-mode uses seed 42, the SimRng is fixed. Entity spawn order is determined by the setup_gauntlet_world() function's code order, not YAML file ordering. As long as the setup function spawns entities in a fixed order (room 1 entities, then room 2, etc.), StableIds are deterministic.

The rule: Gauntlet setup spawns entities in room-number order, entities within a room in a canonical order defined in the constants module. YAML room files are loaded in alphabetical order. If YAML files define entity lists, those lists must be ordered (not sets/maps with non-deterministic iteration).

R2-OQ-10 (from Gestalt): Per-room reset sufficient, or need full server restart?

Answer: Per-room reset is sufficient for Sprint 9-10. Full restart is kill + relaunch, which already works.

Room reset restores entity state within a room. Full reset = restart the server (the tester kills the test client, server exits on disconnect, tester relaunches both). With --test-mode --seed 42, a fresh server launch is identical to the previous one.

Don't build a "full server reset" command. The operating system already provides one: process restart.


5. Risk Register

Risk Likelihood Impact Mitigation
Test client crate setup + bevy compile takes longer than expected Medium +0.5-1 day on S8-4 Estimate already padded (2.5-3.5d vs original 2-3d). First bevy compile is one-time cost.
Cross-binary Layer 3 test is flaky (race conditions, path issues) Medium Delays S8-6 LISTENING:{port} protocol + explicit binary path discovery via CARGO_MANIFEST_DIR.
Content cross-ref validation finds many existing errors Medium Sprint 8 effort bloat fixing content bugs Run validation as WARNING first sprint, promote to ERROR in Sprint 9.
Client test headless stability (OQ-11) blocks CI Low Can't automate client tests in CI Deferred to Tier 3. Manual make pre-pr-client is the interim gate.
Fixture staleness BLOCKER policy frustrates developers Low Devs skip make pre-pr Education. The alternative (false-positive client tests) is worse.
Standalone crate pattern doesn't scale past 4 Rust crates Low Build fragmentation, dep version drift Evaluate Cargo workspace at R-25 (protocol crate extraction).

Summary

Deliverable Status
Lead override analysis (Section 0) — why tooling/test-client/ works, trade-offs, dependency cost Complete
Sprint 8 implementation plan (10 items, ~7.5-9 team-days, dependency graph) Complete
Test client binary final spec (new crate structure, imports, CLI, protocol, text format, golden diff, Layer 3, file inventory) Complete
tooling/ directory structure proposal (Section 2A) Complete
Sprint 9+ roadmap (25 items across 3 tiers, including R-25 protocol extraction) Complete
Q5 (test client location): tooling/test-client/ — override applied Updated
Q3 (cross-cuts): 0 in Sprint 8, design all 4 now, build first 2 in Sprint 10 Answered
R2-OQ-03 (content-ron in pre-pr): Yes, in full pre-pr only Answered
R2-OQ-09 (room ordering): Fixed spawn order in setup function Answered
R2-OQ-10 (full restart): Kill + relaunch, don't build a command Answered