3-round workshop with 7 agents (Tyre, Dudley, Stig, Hoshe, Justine, Gestalt, Ozzie) plus Qatux documenting. Produced: - 59-item prioritized test backlog (60 tickets under epic #455) - Gauntlet test world spec: 7 rooms + hub, 48 entities - Test client binary spec (tooling/test-client/) - Determinism fixes (3 patches, ~22 lines) - Server --test-mode + --port 0 design - Content cross-reference validation (9 checks) - make pre-pr pipeline (6-step) - 38 client tests prioritized - Anti-tedium features (reset plate, hub teleport, WRONG button) - Human tester walkthrough - CI pipeline design (deferred but documented) Sprint 8 scope: ~17.75 team-days across 26 tickets. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
26 KiB
Dudley — Round 2: Determinism Fixes, Server Flags, Cross-Review
Workshop: QA Strategy & Test Architecture Tracks: 2 (Determinism), 3 (Automation), 4 (Serialization) Date: 2026-02-17 Round 2 scope: Concrete code changes, server flag design, open question answers, cross-review
1. Determinism Fixes — Concrete Code Changes
Four fixes to achieve deterministic single-player, single-thread simulation. Estimated total: ~40 lines changed across 4 files.
Fix A: BTreeSet for visible_ids in observer snapshot
Problem: filter_visible_entities in observer/mod.rs:217-219 uses HashSet<u64> for visible_ids. The HashSet itself is only used for membership tests (contains), but the entities: Vec<VisibleEntity> output is built by iterating all_entities.iter() — bevy query order is non-deterministic. Additionally, collect_remembered_entities iterates visible_ids for its exclusion check — safe (membership test), but visible_positions in VisibilityGeometry is also a HashSet and feeds into the same exclusion logic.
However, the real non-determinism in visible_tiles comes from upstream: the shadowcasting algorithm outputs HashSet<(i32, i32)> (shadowcast.rs:99-100), which apply_vision_cone iterates via fov.visible_tiles() (shadowcast.rs:52-53). This HashSet::iter() order feeds directly into VisibilityGeometry.visible_tiles, making the entire tile list non-deterministically ordered.
Fix: Sort visible_tiles in NaturalVision::compute_geometry() after collection. Change visible_ids to BTreeSet for consistency.
File: server/src/perception/query.rs
// Line 8 — BEFORE:
use std::collections::{HashMap, HashSet};
// AFTER:
use std::collections::{BTreeSet, HashMap, HashSet};
// Line 24 — BEFORE:
pub visible_positions: HashSet<(i32, i32)>,
// AFTER:
pub visible_positions: BTreeSet<(i32, i32)>,
File: server/src/perception/query.rs, inside NaturalVision::compute_geometry()
// Lines 67-97 — BEFORE:
let visible_tiles = cone_tiles
.iter()
.map(|&(x, y, sector)| {
// ...
})
.collect();
let visible_positions = cone_tiles.iter().map(|&(x, y, _)| (x, y)).collect();
let sector_lookup = cone_tiles
.iter()
.map(|&(x, y, sector)| ((x, y), sector))
.collect();
// AFTER:
let mut visible_tiles: Vec<_> = cone_tiles
.iter()
.map(|&(x, y, sector)| {
let tile_kind = if walkability.can_move_to(&TilePosition::new(x, y, z)) {
TileKind::Floor
} else {
TileKind::Wall
};
VisibleTile {
x,
y,
z,
visibility: sector,
tile_kind,
}
})
.collect();
// Sort for deterministic snapshot ordering
visible_tiles.sort_by_key(|t| (t.x, t.y));
let visible_positions: BTreeSet<(i32, i32)> =
cone_tiles.iter().map(|&(x, y, _)| (x, y)).collect();
let sector_lookup = cone_tiles
.iter()
.map(|&(x, y, sector)| ((x, y), sector))
.collect();
File: server/src/perception/observer/mod.rs
// Line 10 — BEFORE:
use std::collections::HashSet;
// AFTER:
use std::collections::{BTreeSet, HashSet};
// Line 217-219 — BEFORE:
fn filter_visible_entities(
// ...
) -> (Vec<VisibleEntity>, HashSet<u64>) {
let mut entities = Vec::new();
let mut visible_ids: HashSet<u64> = HashSet::new();
// AFTER:
fn filter_visible_entities(
// ...
) -> (Vec<VisibleEntity>, BTreeSet<u64>) {
let mut entities = Vec::new();
let mut visible_ids: BTreeSet<u64> = BTreeSet::new();
// Line 284-285 — BEFORE:
fn collect_remembered_entities(
// ...
visible_ids: &HashSet<u64>,
visible_positions: &HashSet<(i32, i32)>,
// AFTER:
fn collect_remembered_entities(
// ...
visible_ids: &BTreeSet<u64>,
visible_positions: &BTreeSet<(i32, i32)>,
Impact: visible_tiles in the snapshot is now sorted by (x, y). visible_ids is a BTreeSet (deterministic iteration for any future use). visible_positions is a BTreeSet (deterministic membership tests — same result as HashSet, but consistent type choice). The sector_lookup HashMap stays — it's point-lookup only via .get(), never iterated.
Fix B: Sort visible entities in snapshot
Problem: filter_visible_entities builds the entities: Vec<VisibleEntity> list by iterating all_entities.iter() — bevy query iteration order is archetype-based and not stable across runs. The snapshot entity list order is therefore non-deterministic.
Fix: Sort entities by entity_id after collection.
File: server/src/perception/observer/mod.rs, inside compute_observer_snapshot()
// After line 126 (after collect_remembered_entities call) — ADD:
entities.sort_by_key(|e| {
// Primary sort: entity_id for deterministic ordering
// Secondary sort not needed — entity_id is unique
e.entity_id
});
Insert this between the collect_remembered_entities call and the sprint anomaly detection block.
Fix C: Pin monologue system order (CORRECTION — already done)
Round 1 correction: In my Round 1 analysis, I stated that monologue systems lack explicit ordering. This was wrong. After reading bridge/mod.rs:168-175, I can see the ordering is already explicit:
crate::simulation::monologue::trigger_monologue
.after(crate::simulation::movement::validate_movement),
crate::simulation::monologue::process_sprint_anomaly_monologue
.after(crate::simulation::monologue::trigger_monologue),
No change needed. The BridgePlugin already pins both monologue systems with .after() constraints relative to validate_movement and each other. The SimRng consumption order is deterministic. I apologize for the Round 1 error — I only checked SimulationPlugin in simulation/mod.rs and missed the ordering in BridgePlugin.
Fix D: Sort movers in validate_movement
Problem: validate_movement in movement.rs:288 iterates movers.iter_mut() — bevy query order is non-deterministic. When two entities move to the same tile in the same tick, which one "wins" depends on iteration order. The existing test (movement.rs:556-590) correctly tests that exactly one wins, but not which one.
Fix: Collect movers into a Vec, sort by Entity bits, then process in order.
File: server/src/simulation/movement.rs
// Lines 288-316 — BEFORE:
for (entity, intent, mut position, presence) in movers.iter_mut() {
let target = &intent.target;
let layer = presence.copied().unwrap_or_default();
let slot = (*target, layer);
if !map.can_move_to(target) {
tracing::trace!("Entity {:?} blocked by terrain at {:?}", entity, target);
} else if occupied.contains_key(&slot) {
// ...
} else {
// ... move logic
}
commands.entity(entity).remove::<MoveIntent>();
}
// AFTER:
// Collect and sort for deterministic processing order (D-010 principle 4).
// Entity::to_bits() provides a stable ordering within a single run.
let mut mover_list: Vec<_> = movers.iter_mut().collect();
mover_list.sort_by_key(|(entity, _, _, _)| entity.to_bits());
for (entity, intent, mut position, presence) in mover_list {
let target = &intent.target;
let layer = presence.copied().unwrap_or_default();
let slot = (*target, layer);
if !map.can_move_to(target) {
tracing::trace!("Entity {:?} blocked by terrain at {:?}", entity, target);
} else if occupied.contains_key(&slot) {
tracing::trace!(
"Entity {:?} blocked by entity at {:?} (layer {:?})",
entity,
target,
layer
);
} else {
tracing::trace!(
"Entity {:?} moving from {:?} to {:?} (layer {:?})",
entity,
*position,
target,
layer
);
occupied.remove(&(*position, layer));
*position = *target;
occupied.insert(slot, entity);
}
commands.entity(entity).remove::<MoveIntent>();
}
Note on entity.to_bits(): This provides stable ordering within a single run (same spawn order → same Entity bits → same processing order). Across save/load cycles, Entity bits may differ, but StableId lookup would be needed. For the Gauntlet (fixed spawn order, no save/load), to_bits() is sufficient. If cross-session determinism is needed later, sort by registry.to_stable(entity).map(|s| s.0) instead.
Test update: The existing test validate_movement_two_movers_same_target_first_wins can now be strengthened to assert WHICH entity wins (the one with the lower Entity bits), making it a determinism regression test.
2. Server --test-mode Flag Design
What it does:
| Aspect | Behavior |
|---|---|
| Content | Loads the Gauntlet content pack (content/gauntlet/) instead of campaign content. Falls back to the current hardcoded proof room if Gauntlet content doesn't exist yet. |
| Seed | Fixed seed (42) for deterministic replay. |
| Stdout | Prints LISTENING:{port} to stdout after binding the TCP socket, before accepting a connection. This is the signal the test client uses to discover the port. |
| Shutdown | Exits after first client disconnects (existing behavior — ServerRunning flag already handles this). |
| Logging | Defaults to warn level instead of debug to reduce noise in test output. Test can override with RUST_LOG. |
Minimum change to main.rs:
// server/src/main.rs — proposed changes
fn main() {
// Parse CLI args
let args: Vec<String> = std::env::args().collect();
let test_mode = args.iter().any(|a| a == "--test-mode");
let port = args.iter()
.position(|a| a == "--port")
.and_then(|i| args.get(i + 1))
.map(|s| s.as_str());
// Initialize tracing — quieter in test mode
let default_filter = if test_mode {
"settled_reach_server=warn"
} else {
"settled_reach_server=debug"
};
tracing_subscriber::registry()
.with(
tracing_subscriber::EnvFilter::try_from_default_env()
.unwrap_or_else(|_| default_filter.into()),
)
.with(tracing_subscriber::fmt::layer())
.init();
// Resolve address
let addr = if let Some(p) = port {
format!("127.0.0.1:{}", p)
} else {
std::env::args()
.nth(1)
.filter(|a| !a.starts_with("--")) // Don't treat flags as address
.or_else(|| std::env::var("SR_ADDR").ok())
.unwrap_or_else(|| "127.0.0.1:9876".to_string())
};
// Bind FIRST, print port, THEN accept.
// Splitting bind/accept is critical for --port 0: we need
// the OS-assigned port number before the client connects.
let listener = std::net::TcpListener::bind(&addr).unwrap_or_else(|e| {
eprintln!("Failed to bind {}: {}", addr, e);
std::process::exit(1);
});
let actual_port = listener.local_addr().unwrap().port();
// Print the port signal for test client discovery.
// Flush to ensure the test client reads it before we block on accept().
println!("LISTENING:{}", actual_port);
use std::io::Write;
std::io::stdout().flush().ok();
tracing::info!("Waiting for client connection on port {}", actual_port);
let bridge = TcpBridge::accept_on(listener).unwrap_or_else(|e| {
tracing::error!("Failed to accept: {}", e);
std::process::exit(1);
});
// ... rest of main.rs (app setup, game loop) ...
// Seed: use fixed seed in test mode
let seed = if test_mode { 42 } else { 0 };
// Replace: .insert_resource(rng::SimRng::new(0))
// With: passed into SimulationPlugin or inserted directly
}
Key design decisions:
TcpBridge::accept_on(listener)already exists (tcp.rs:67-97). This separates bind from accept, which is exactly what we need for port discovery.- The
LISTENING:{port}signal is printed to stdout (not stderr, not tracing) so the test client can parse it reliably without tracing noise. --port 0triggers OS port assignment. The actual port is read fromlistener.local_addr().- The first positional argument is currently treated as the address. With
--port, we switch to flag-based parsing. Both modes coexist —--porttakes precedence.
3. --port 0 Support
How it works with the current TCP bridge setup:
The current main.rs:47 calls TcpBridge::accept(&addr) which both binds AND accepts in one call. For --port 0, we need to:
- Bind the listener (to get the OS-assigned port)
- Print the port
- Accept on the bound listener
TcpBridge::accept_on(listener) already exists for exactly this purpose (it was added for test race condition prevention — tcp.rs:67, "Avoids race conditions in tests by separating bind from accept").
No changes needed to TcpBridge. The only change is in main.rs to call TcpBridge::accept_on() instead of TcpBridge::accept(). See the code above — the proposed main.rs already uses accept_on.
Interaction with the test client:
Test client Server
| |
| bind("127.0.0.1:0")
| stdout: "LISTENING:54321"
| accept() — blocks
| |
| <-- read stdout, parse port |
| connect("127.0.0.1:54321") ------>|
| accept() returns
| <--- snapshot ---------- |
| --- input -------------> |
| ... |
4. OQ-2 Answer: Does rmp_serde accept int_16-encoded positive values for u64?
Yes. Verified by tracing through the rmp-serde 1.3.1 source code.
When rmp_serde::Deserializer::deserialize_u64() encounters a Marker::I16 byte (0xd1), it calls visitor.visit_i16(value). Serde's default visit_i16 implementation for u64 deserialization performs u64::try_from(value), which succeeds for any non-negative i16 value.
Chain of evidence:
deserialize_u64()delegates toany_num()(rmp-serde decode.rs)any_num()matches on the marker:Marker::I16 => visitor.visit_i16(rd.read_data_i16()?)- Serde's
Visitor::visit_i16for u64 callsself.visit_i64(v as i64) - Serde's
Visitor::visit_i64for u64 callsu64::try_from(v)— succeeds for positive values
This means: GDScript encoding 256 as int_16 (0xd1, 0x01, 0x00) will deserialize correctly into Rust's u64 field. The asymmetric encoding between GDScript and Rust (Hoshe's observation in T4-H1) does NOT cause a deserialization failure.
However, negative int_16 values (e.g., -129) will fail u64::try_from() and produce a serde error. This is the expected behavior — a negative tick is always a bug.
Recommendation: No code change needed. The existing deserialization handles this correctly. Hoshe's cross-language boundary tests should explicitly cover values 256-32767 deserialized from int_16 into u64 fields to document this guarantee.
5. OQ-7 Answer: Room Name Lookup
Question: Where does the test client get room names — from Gauntlet coordinate bounds or content metadata?
Answer: Gauntlet coordinate bounds, defined as constants in a shared module.
The Gauntlet is a fixed-layout map where every room has known coordinate ranges. The test client doesn't need to query the server for room names — it knows the map layout at compile time.
// server/src/test_world/constants.rs (also importable by test client binary)
pub struct GauntletRoom {
pub name: &'static str,
pub bounds: (TilePosition, TilePosition), // top-left, bottom-right
pub observer_position: TilePosition, // fixed position for golden file snapshots
}
pub const ROOMS: &[GauntletRoom] = &[
GauntletRoom {
name: "inventory_warehouse",
bounds: (TilePosition::new(0, 0, 0), TilePosition::new(15, 15, 0)),
observer_position: TilePosition::new(8, 8, 0),
},
GauntletRoom {
name: "occlusion_corridor",
bounds: (TilePosition::new(16, 0, 0), TilePosition::new(31, 15, 0)),
observer_position: TilePosition::new(20, 8, 0),
},
// ... etc
];
/// Look up the room a position falls in.
pub fn room_at(pos: &TilePosition) -> Option<&'static GauntletRoom> {
ROOMS.iter().find(|r| {
pos.x >= r.bounds.0.x && pos.x <= r.bounds.1.x
&& pos.y >= r.bounds.0.y && pos.y <= r.bounds.1.y
&& pos.z >= r.bounds.0.z && pos.z <= r.bounds.1.z
})
}
The test client uses room_at(player_position) to label rooms in text output. The server doesn't need a room-naming API — the coordinate bounds are the source of truth.
Why not content metadata? The Gauntlet rooms aren't "locations" in the content system sense — they're mechanical test environments. Content locations (transit district, docking bay) use a different system (content/locations/). The Gauntlet constants module is purpose-built for the test world and shouldn't be conflated with game content.
6. Room Reset Triggers (Anti-Tedium)
Use case: A human tester walks through the Gauntlet, completes the Inventory Warehouse tests, then wants to re-test. Instead of restarting the server, they step on a trigger tile to reset the room to tick-0 state.
Server-Side Mechanism
Proposal: Reset trigger tiles that restore per-room entity state from a stored snapshot.
// server/src/test_world/reset.rs
/// Component marking a tile as a room reset trigger.
#[derive(Component)]
pub struct RoomResetTrigger {
pub room_name: &'static str,
}
/// Stored tick-0 state for a Gauntlet room's entities.
/// Captured immediately after setup, before any simulation ticks.
#[derive(Resource)]
pub struct RoomSnapshots {
/// room_name -> list of (StableId, component bundle snapshot)
snapshots: BTreeMap<String, Vec<EntitySnapshot>>,
}
#[derive(Clone)]
struct EntitySnapshot {
stable_id: StableId,
position: TilePosition,
// All resettable component values captured at tick 0:
knowledge_state: Option<KnowledgeState>,
relationship: Option<RelationshipState>,
confidence: Option<KnowledgeConfidence>,
inventory_items: Vec<(StableId, String, u8)>, // item_sid, name, slot
// Extend as needed per room type
}
Reset Flow
- At Gauntlet setup (tick 0): After
setup_gauntlet_world()completes, capture each room's entity state intoRoomSnapshots. - Trigger detection system: Each tick, check if the player is on a
RoomResetTriggertile. (Simple position check, runs aftervalidate_movement.) - Reset execution: When triggered:
- Move all entities in the room back to their tick-0 positions
- Restore knowledge graph entries to tick-0 state
- Remove all inventory items picked up from the room, re-place them
- Reset the player's KnowledgeGraph entries for entities in this room
- Reset MonologueState cooldowns for the room
/// System: detect player stepping on reset trigger tiles.
pub fn detect_room_reset(
player_query: Query<&TilePosition, With<PlayerCharacter>>,
triggers: Query<(&TilePosition, &RoomResetTrigger)>,
mut reset_events: EventWriter<RoomResetEvent>,
) {
let Ok(player_pos) = player_query.single() else { return };
for (trigger_pos, trigger) in triggers.iter() {
if player_pos == trigger_pos {
reset_events.send(RoomResetEvent {
room_name: trigger.room_name.to_string(),
});
}
}
}
/// System: execute room reset from stored snapshot.
pub fn execute_room_reset(
mut events: EventReader<RoomResetEvent>,
snapshots: Res<RoomSnapshots>,
registry: Res<EntityRegistry>,
mut positions: Query<&mut TilePosition>,
// ... other component queries for full reset
) {
for event in events.read() {
let Some(room_snapshot) = snapshots.snapshots.get(&event.room_name) else {
tracing::warn!("No snapshot for room {}", event.room_name);
continue;
};
for entity_snap in room_snapshot {
let Some(entity) = registry.to_entity(&entity_snap.stable_id) else {
continue;
};
if let Ok(mut pos) = positions.get_mut(entity) {
*pos = entity_snap.position;
}
// Restore other components...
}
tracing::info!("Room {} reset to tick-0 state", event.room_name);
}
}
ECS Implications
-
Despawned entities: If the tester picked up an item (which removes TilePosition and adds CarriedBy), the reset must reverse this — remove CarriedBy/InventorySlot, re-add TilePosition. This is straightforward with
Commands. -
Knowledge graph: The player's KG entries for room entities need selective reset. This requires knowing which StableIds belong to which room — the
GauntletRoomconstants provide this mapping. -
No new entity spawning: Reset restores existing entities, it doesn't despawn/respawn. This avoids Entity ID recycling issues and keeps the EntityRegistry stable.
-
Trigger debounce: The reset trigger should have a cooldown (e.g., 10 ticks) to prevent re-triggering while the player walks across the trigger tile.
-
Test mode only:
RoomResetTriggerentities and the reset systems are only added in--test-mode. They don't exist in production.
7. Cross-Review: Hoshe's Layer 3 Test
Hoshe's proposal (from hoshe-round1.md, T4-H5): Launch the server binary as a child process with --test-mode --port 0, connect via TCP, send one input, receive one snapshot, assert fields.
Feasibility Assessment
Verdict: Feasible with the server flag changes from section 2. No architectural blockers.
Detailed review:
| Aspect | Hoshe's Proposal | Feasibility | Notes |
|---|---|---|---|
| Build server binary | cargo build --bin settled-reach-server |
Works | Binary name is correct per Cargo.toml |
| Launch as subprocess | Command::new("target/debug/settled-reach-server") |
Works | Standard Rust subprocess API |
--test-mode flag |
Args: ["--test-mode", "--port", "0"] |
Needs implementation | See section 2 above |
| Port discovery | Parse LISTENING:{port} from stdout |
Works with section 2 changes | Current main.rs doesn't print this — section 2 adds it |
| TCP connect | TcpStream::connect(format!("127.0.0.1:{}", port)) |
Works | Standard TCP |
| Send PlayerInput | write_framed with rmp_serde::to_vec_named |
Works | But note: Hoshe serializes Vec<PlayerInput>, which is correct — the bridge expects a Vec, not a single input |
| Receive snapshot | read_framed → rmp_serde::from_slice::<ObserverSnapshot> |
Works | The bridge sends snapshots as single ObserverSnapshot, not Vec |
| Assertions | version, tick, entity count, player entity | All valid | Snapshot at tick 0 will have the proof room entities |
| Cleanup | Drop connection, kill process | Works | server.kill().ok() is correct — graceful shutdown via ServerRunning flag also works when TCP disconnects |
Issues Found
-
Timing sensitivity. Hoshe's test sends one input and expects one snapshot. But the server's game loop runs at 20 tps with
thread::sleep(50ms). After accepting the connection, the server enters the game loop. The firstapp.update()callsreceive_bridge_inputswhich reads the TCP socket (non-blocking). If the test client hasn't sent its input yet, the server produces a snapshot with tick=0 and no input processed. The client may receive this "empty" snapshot before its input is processed.Fix: The test should loop, reading snapshots until it gets one where its input has been processed (e.g.,
snapshot.tick >= 1), or accept that tick=0 snapshots are valid (they contain the player entity and visible NPCs even without input).Hoshe's assertion
snapshot.tick == 0is actually fine for the first snapshot. The test doesn't need to verify input processing — it just needs to verify the snapshot arrives. Input processing is tested at Layer 2. -
Serialization format mismatch. Hoshe writes
rmp_serde::to_vec_named(&inputs)but the bridge receives withrmp_serde::from_slice::<Vec<PlayerInput>>. Theto_vec_namedfunction produces MessagePack maps (with field names), whileto_vecproduces arrays. SincePlayerInputderivesDeserialize,from_slicehandles both formats — but the GDScript client sendsto_vec(array) format. For consistency, the test should usermp_serde::to_vec(&inputs)to match the real client's encoding. -
read_framed import. Hoshe references
read_framedandwrite_framedfrom the bridge module. These are inserver/src/bridge/framing.rsand take&mut impl Read/&mut impl Write. The test will need to import them viasettled_reach_server::bridge::framing::{read_framed, write_framed}— this requires the bridge framing module to be public. Currently,framingispub mod framinginbridge/mod.rs:9, so it's accessible.
Summary
Hoshe's Layer 3 test design is sound and feasible once the --test-mode and --port 0 server flags are implemented. The test correctly identifies what Layer 3 catches that Layer 2 doesn't (subprocess I/O blocking, binary startup issues). Two minor adjustments recommended: (1) use to_vec instead of to_vec_named for input serialization, (2) tolerate receiving an "initial" snapshot before the input is processed.
Summary of Round 2 Deliverables
| # | Deliverable | Status |
|---|---|---|
| 1 | 4 determinism fixes with exact code changes | Complete (Fix A-D; Fix C was already done) |
| 2 | --test-mode flag design + main.rs changes |
Complete |
| 3 | --port 0 support design |
Complete (uses existing accept_on) |
| 4 | OQ-2: rmp_serde int_16→u64 | Answered: Yes, works correctly |
| 5 | OQ-7: Room name lookup | Answered: Gauntlet coordinate constants |
| 6 | Room reset trigger mechanism | Complete |
| 7 | Cross-review of Hoshe's Layer 3 test | Complete — feasible with 2 minor adjustments |