fix(client): cold-body Pending responses reach the retry path — the launch-shape starvation root cause

The true cause of the black first launch, server-confirmed after four
disproven theories: a stone-cold body answers the FIRST window request
with a whole-response status 'Pending' (only the whole-body cache-hit
branch sets Ready), and on_response()'s very first check — status !=
Ready -> return — swallowed it before the retry machinery could run.
retries stayed 0 forever; the DERIVING state never resolved. Second
connections worked by luck (the first request warms the whole-body
cache, so they read Ready and take the healthy path). Both tile fan-out
AND single-window first-descents were affected — one shared function,
one fix: branch on the outer status first (the atlas_generation_proxy
reference shape): Pending -> retry, Ready -> existing null-window retry,
NotFound/Error -> give up immediately (the principled give-up policy,
replacing the elapsed-retries ceiling).

Hardening in the same round: deterministic exponential backoff (0.5s
doubling, 4s cap) + per-tile stagger (0.1s * index — six tiles retry at
0.5/0.6/0.7/0.8/0.9/1.0s, strictly-increasing asserted, not jittered);
MAX_RETRIES 20->30 (~110s horizon under backoff).

Regressions are wire-accurate by construction: whole-response Pending
(body_id only, no center/n/granularity — verified against the server's
own response construction) delivered through the REAL fan-out
(tile_set._on_atlas_layers_received, never tile.on_response directly),
mirrored at single-window level. Two first-draft tests that passed with
the bug reverted were caught and strengthened before reporting; every
fix and hardening piece revert-verified independently (7+1 failing
tests without them).
This commit is contained in:
2026-07-22 20:56:46 +02:00
parent ddc4483462
commit 3891c83206
4 changed files with 427 additions and 17 deletions
@@ -38,8 +38,34 @@ const AtlasWindowCache := preload("res://ui/implant/apps/atlas/atlas_window_cach
const DISTRICT_WINDOW_DEFAULT_N: int = 32
const DEBOUNCE_DELAY: float = 0.15 # 150ms, §4/§5
const RETRY_DELAY: float = 0.5 # matches atlas_generation_proxy.gd's GEN_RETRY_DELAY
const MAX_RETRIES: int = 20 # ~10s ceiling, matches atlas_generation_proxy.gd's GEN_MAX_RETRIES
## Cold-start dossier (PR #192 round 3): the retry-on-PENDING loop used to be
## a flat RETRY_DELAY=0.5s / MAX_RETRIES=20 (~10s ceiling), copied verbatim
## from atlas_generation_proxy.gd's Layer1 poll — a DIFFERENT, typically
## faster derive. The real starvation bug turned out to be the status-gate
## fix in on_response() (see that function's own doc) — this backoff/stagger
## work is HARDENING landed alongside it, not the fix itself: once the
## status-gate fix makes 6 independent tiles all correctly retry on a
## whole-response Pending, they do so in perfect lockstep (all six went
## pending at entry within the same frame, so all six retry timers fire
## within the same frame too) — six re-requests every RETRY_DELAY, in sync,
## is exactly the "storm" shape worth damping even though it isn't what
## caused the starvation. Exponential backoff (INITIAL_RETRY_DELAY doubling
## to MAX_RETRY_DELAY) plus a DETERMINISTIC per-tile stagger
## (STAGGER_STEP * stagger_index, set once by the owning AtlasWindowTileSet
## at construction — see `_stagger_index`) spread that pulse into a trickle:
## tile 0 retries at 0.5s, tile 1 at 0.6s, tile 2 at 0.7s, etc. — deterministic
## and directly assertable in a test, not a randomized jitter a test would
## have to tolerance-check. Backoff ALSO buys a much longer wall-clock window
## from a modest MAX_RETRIES increase (~110s at 30 retries, see
## _retry_delay_for()'s own doc) without ever polling aggressively for that
## whole span. A fast (already-warm) response still resolves on retry #1,
## unaffected — backoff/stagger only matter once a request is genuinely
## still pending past the first cycle.
const INITIAL_RETRY_DELAY: float = 0.5 # first retry, matches the old flat RETRY_DELAY
const MAX_RETRY_DELAY: float = 4.0 # backoff ceiling — never polls slower than this
const STAGGER_STEP: float = 0.1 # per-tile-index offset — tile i retries STAGGER_STEP*i later
const MAX_RETRIES: int = 30 # ~110s wall-clock at the backoff schedule above
## T-1150 struct/key plumbing: legacy int granularity — district is the
## default for every caller that doesn't request quarter/Region explicitly.
@@ -89,6 +115,12 @@ var _min_wl_m: int = DEFAULT_MIN_WL_M
var _pending: bool = false
var _retries: int = 0
var _debounce_timer: Timer = null
## Cold-start dossier round 3: deterministic per-request stagger index for
## the retry backoff (see STAGGER_STEP's own doc) — 0 for the single-window
## viewer's own request (no fan-out, nothing to desync from), the tile's own
## index (0..5) for a tile-set-owned request (AtlasWindowTileSet.enter()
## sets this once at construction, right after AtlasWindowRequest.new()).
var _stagger_index: int = 0
func _init(owner_ref = null) -> void:
@@ -288,25 +320,44 @@ func _on_debounce_timeout() -> void:
## real server, always) -> v2 is the ONLY granularity comparison; absent (a
## hypothetically old, pre-T-1152 server) -> fall back to the legacy
## comparison alone, matching this object's own pre-T-1152 behavior exactly.
## PR #192 cold-start round 3: the coordinator's live cold-server capture
## (retries=0, pending=true, forever) exposed that the OLD version of this
## function returned unconditionally whenever the WHOLE response's status
## wasn't "Ready" — treating a cold body's `status: "Pending"` (the FIRST
## request against a whole-body cache miss, before ANY layer including the
## window has even been queued — `serve_district_window`/`get_or_generate()`
## in server/src/atlas/layer_proxy.rs) identically to `NotFound`/`Error`: a
## silent no-op, never reaching the retry-scheduling code at all. Confirmed
## server-side: `status: Ready` is set ONLY on the whole-body cache-HIT
## branch, entirely independent of whether the WINDOW itself has resolved —
## so a cold body's first-ever window request gets `Pending` at the OUTER
## layer, while a body someone has already warmed (a later connection, or
## this SAME connection's own re-request once its own AnalyzeBody has
## landed) gets `Ready` with `district_window: null` inside it, correctly
## reaching the retry branch below. Same "still generating" signal, two
## different wire shapes depending on which cache warmed first — the fix is
## to treat BOTH as the identical retry-worthy state, matching
## atlas_generation_proxy.gd's own on_response() `match` shape exactly
## (Ready -> handle, Pending -> retry, NotFound/Error -> give up now, not
## after MAX_RETRIES: a real error is never going to resolve by waiting).
func on_response(response: Dictionary) -> void:
if str(response.get("body_id", "")) != _body_id:
return
if str(response.get("status", "")) != "Ready":
return # Pending/NotFound/Error on the WHOLE response — not a window signal either way
var status := str(response.get("status", ""))
if status == "Pending":
_retry_if_pending()
return
if status != "Ready":
_pending = false # NotFound / Error — a real failure, not a queue wait; give up now
return
var window: Variant = response.get("district_window")
if window == null:
# §1: an as-yet-underived window rides as `district_window: None` inside
# a Ready response — this is the "still generating" signal, not an
# error. Re-poll until the background derive lands or the retry
# ceiling is hit (queue-based serving, PR #185 — the response lands
# on a LATER tick, never this same round-trip).
if not _pending:
return
if _retries < MAX_RETRIES:
_retries += 1
_schedule_retry()
else:
_pending = false # gave up — caller's border-fade / empty state persists
# §1: an as-yet-underived window rides as `district_window: None`
# inside an OUTER-Ready response — the whole-body cache already
# warmed, but this specific window hasn't derived yet. Same
# "still generating" signal the outer-Pending branch above handles,
# just the OTHER wire shape it can arrive in.
_retry_if_pending()
return
var w: Dictionary = window
@@ -328,6 +379,21 @@ func on_response(response: Dictionary) -> void:
window_ready.emit(w)
## Shared "still generating, re-poll" logic for BOTH wire shapes on_response()
## can see it in (outer status=="Pending", or inner district_window==null
## inside an outer Ready) — re-request until the derive lands or the retry
## ceiling is hit (queue-based serving, PR #185 — the response lands on a
## LATER tick, never this same round-trip).
func _retry_if_pending() -> void:
if not _pending:
return
if _retries < MAX_RETRIES:
_retries += 1
_schedule_retry()
else:
_pending = false # gave up — caller's border-fade / empty state persists
## The granularity half of on_response()'s staleness check, split out for the
## v2-authoritative-when-present precedence rule (see on_response()'s own
## doc for the full live-round rationale). Presence, not value, is the
@@ -343,8 +409,23 @@ func _echoed_granularity_matches(w: Dictionary) -> bool:
return echoed_granularity == _granularity
## Pure: the exponential-backoff delay for retry attempt number `retry_count`
## (1-indexed — the FIRST retry, right after the initial request's own
## PENDING answer, uses `retry_count=1`), staggered by `stagger_index`
## (STAGGER_STEP*stagger_index added on top — deterministic, not randomized,
## so a test can assert the exact delay sequence for tile N directly). Split
## out from _schedule_retry() as a pure function for the same reason every
## other formula in this file is: directly unit-testable without a live
## Timer/SceneTree.
static func _retry_delay_for(retry_count: int, stagger_index: int) -> float:
var base: float = INITIAL_RETRY_DELAY * pow(2.0, float(maxi(retry_count - 1, 0)))
var capped: float = minf(base, MAX_RETRY_DELAY)
return capped + STAGGER_STEP * float(stagger_index)
func _schedule_retry() -> void:
var timer := get_tree().create_timer(RETRY_DELAY)
var delay: float = _retry_delay_for(_retries, _stagger_index)
var timer := get_tree().create_timer(delay)
timer.timeout.connect(
func() -> void:
if _pending:
@@ -97,6 +97,17 @@ func enter(body_id: String, body_radius_km: float) -> void:
var center: Vector2i = centers[i]
var request = AtlasWindowRequest.new(self)
request.name = "Tile%d" % i
# Cold-start dossier round 3 hardening: deterministic per-tile retry
# stagger (STAGGER_STEP*i) — without it, all 6 tiles go pending in the
# same frame and retry in perfect lockstep, a request pulse every
# RETRY_DELAY instead of a spread trickle. Set BEFORE request_now()
# so it's already in place for the very first retry, if one fires.
# Cold-start dossier round 3 hardening: deterministic per-tile retry
# stagger (STAGGER_STEP*i) — without it, all 6 tiles go pending in the
# same frame and retry in perfect lockstep, a request pulse every
# RETRY_DELAY instead of a spread trickle. Set BEFORE request_now()
# so it's already in place for the very first retry, if one fires.
request._stagger_index = i
add_child(request)
var tile_index := i # capture by value for the lambda below
request.window_ready.connect(