Two parsing bugs in the summary, found while measuring suite times.
DOUBLE COUNT. gdUnit4 prints one "Statistics:" line per suite and then a
single "Overall Summary:" line whose numbers are the sum of all of them. The
pattern matched both shapes and summed all 87 lines, so every total was
exactly twice the truth: a full run reported 3,660 tests against an actual
1,830, and a 26-test suite reported 52. It was invisible because it doubled
UNIFORMLY — nothing ever looked inconsistent, only large. Every count quoted
from this harness, in this session and before it, was 2x.
Now prefers the Overall Summary, which is gdUnit4's own arithmetic over the
whole run and so cannot disagree with itself; per-suite summing survives only
as a fallback for a run that dies before printing it.
ANSI. gdUnit4 colourises output and the escape sequences sit BETWEEN the
fields of the summary line, so patterns matching the raw log silently fell
through to the weaker "Executed test cases" fallback — which cannot see skips
and reported a fully skipped suite as 26 FAILED. All parsing now runs against
a de-ANSI'd copy, including the load-error guards.
SKIPS are now parsed and surfaced as their own JSON field, and excluded from
passed. Counting a skipped test as passing is the same false-green shape the
harness guards exist to prevent, and it stops being hypothetical the moment a
suite is deliberately skipped.
Verified against a fully-skipped suite (26 total / 0 passed / 0 failed / 26
skipped, was 26 FAILED) and a full run (1,830 total / 1,804 passed / 0 failed
/ 26 skipped, was 3,660/3,660).
Pair session with Jeroen, 2026-07-27.
Co-Authored-By: Claude <noreply@anthropic.com>
Found by walking into it. test_step_canvas_annotation_layer.gd had a parse
error from an earlier edit in this session, so gdUnit4 could not load it and
ran the other suites instead. The harness printed 3610 passed / 0 failed and
exit 0. Fifty tests had not run for hours and nothing said so — the full
suite reports 3660 with the file repaired, and that difference was invisible.
Two states are now hard harness failures rather than test results:
load_error — a suite failed to LOAD. Any pass count excludes it, so a green
number is a lie. The hint names the offending file.
no_tests — zero tests executed. A run that executes nothing can never be
a pass; previously a mistyped --filter printed "Tests passed".
Both add a "harness_error" field to the summary JSON and exit 2. The exit
code cannot inherit gdUnit4's, which returns 0 in both states — that is
precisely why they were invisible.
Verified by injecting each failure rather than by reasoning about it. The
load_error guard was checked in the case that actually matters: one broken
file among many, where total stays large and failed stays zero. That run now
reports 3610/0 WITH harness_error and exits 2, where before it was
indistinguishable from success.
Also repairs the file itself: a missed set_frame() argument (the parse error),
and a cell-placement test still asserting pre-inversion spacing. Rewritten to
assert the invariant that survives the extent inversion, the viewport aspect
ratio and panning — half the SHORT axis is half a rung cell — instead of a
literal. Two things it deliberately does not assert, both of which the
previous version got wrong: "the corner is half a district away" holds only
on a square canvas, and the canvas is one district WIDE without sitting ON a
district. It is a free-floating window centred wherever the player panned;
zoom is stepped, pan is continuous. A rung names a scale, not a cell you are
inside. A second test pins that with a deliberately unaligned world centre,
so a future change that snaps the canvas to the rung lattice — making pan
step instead of slide — fails here.
Pair session with Jeroen, 2026-07-27.
Co-Authored-By: Claude <noreply@anthropic.com>
- test_character_creation_sprint28: before_each now seeds
_selected_bookmark_id and _selected_location_id so the new disabled-
guard in _on_start() (round 2) doesn't silently block 5 existing
tests that call _on_start()/KEY_ENTER without setting up a valid
bookmark selection. Restores the 2 tests Hoshe flagged as R2-H1 plus
3 siblings that would have degraded the same way under the guard.
- tests/run-godot: LOG_FILE now includes $$ (PID) so concurrent runs
across worktrees don't clobber each other's logs. Path is echoed
back via the stdout JSON "log" field and the stderr hint line, so
callers never need to predict it (R2-H2).
Makes tests/run-godot self-containing so neither humans nor LLM callers
have to remember to wrap it in a timeout or pipe it into a file. A hung
test now kills cleanly at 300s with a clear TEST_TIMEOUT marker and
bisection hint instead of silently burning an hour of wall clock (as
Sprint 36 learned).
- Godot+gdUnit4 output goes to /tmp/sr-run-godot.log (overwritten each
run). Nothing streams to stdout/stderr — 20k+ lines of test log into
a terminal or an LLM context is unworkable.
- Stdout: one-line JSON summary, with a "log" field pointing at the
file. On timeout adds "timeout":true and "timeout_sec":300.
- Stderr: a short hint block. On pass: one line. On failure: three
commands to inspect the log. On timeout: a bisection recipe.
- Single well-known path instead of an env var — worktrees each want
their own value and the indirection makes the hint lines meaningless.
Concurrent runs are the caller's problem.
- timeout(1) --foreground --kill-after=10 to escalate to SIGKILL if
Godot ignores SIGTERM.
Six test runner scripts at tests/: run-rust, run-godot, run-ipc-fixtures,
run-ipc-protocol, run-ipc-integration, run-all. Plus run-ipc-benchmark
for Layer 3 timing. All produce structured JSON stdout, support --filter,
and exit 0/non-zero. Makefile targets updated to delegate to scripts.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>