fix(client): cold-body Pending responses reach the retry path — the launch-shape starvation root cause
The true cause of the black first launch, server-confirmed after four disproven theories: a stone-cold body answers the FIRST window request with a whole-response status 'Pending' (only the whole-body cache-hit branch sets Ready), and on_response()'s very first check — status != Ready -> return — swallowed it before the retry machinery could run. retries stayed 0 forever; the DERIVING state never resolved. Second connections worked by luck (the first request warms the whole-body cache, so they read Ready and take the healthy path). Both tile fan-out AND single-window first-descents were affected — one shared function, one fix: branch on the outer status first (the atlas_generation_proxy reference shape): Pending -> retry, Ready -> existing null-window retry, NotFound/Error -> give up immediately (the principled give-up policy, replacing the elapsed-retries ceiling). Hardening in the same round: deterministic exponential backoff (0.5s doubling, 4s cap) + per-tile stagger (0.1s * index — six tiles retry at 0.5/0.6/0.7/0.8/0.9/1.0s, strictly-increasing asserted, not jittered); MAX_RETRIES 20->30 (~110s horizon under backoff). Regressions are wire-accurate by construction: whole-response Pending (body_id only, no center/n/granularity — verified against the server's own response construction) delivered through the REAL fan-out (tile_set._on_atlas_layers_received, never tile.on_response directly), mirrored at single-window level. Two first-draft tests that passed with the bug reverted were caught and strengthened before reporting; every fix and hardening piece revert-verified independently (7+1 failing tests without them).
This commit is contained in:
@@ -38,8 +38,34 @@ const AtlasWindowCache := preload("res://ui/implant/apps/atlas/atlas_window_cach
|
||||
|
||||
const DISTRICT_WINDOW_DEFAULT_N: int = 32
|
||||
const DEBOUNCE_DELAY: float = 0.15 # 150ms, §4/§5
|
||||
const RETRY_DELAY: float = 0.5 # matches atlas_generation_proxy.gd's GEN_RETRY_DELAY
|
||||
const MAX_RETRIES: int = 20 # ~10s ceiling, matches atlas_generation_proxy.gd's GEN_MAX_RETRIES
|
||||
|
||||
## Cold-start dossier (PR #192 round 3): the retry-on-PENDING loop used to be
|
||||
## a flat RETRY_DELAY=0.5s / MAX_RETRIES=20 (~10s ceiling), copied verbatim
|
||||
## from atlas_generation_proxy.gd's Layer1 poll — a DIFFERENT, typically
|
||||
## faster derive. The real starvation bug turned out to be the status-gate
|
||||
## fix in on_response() (see that function's own doc) — this backoff/stagger
|
||||
## work is HARDENING landed alongside it, not the fix itself: once the
|
||||
## status-gate fix makes 6 independent tiles all correctly retry on a
|
||||
## whole-response Pending, they do so in perfect lockstep (all six went
|
||||
## pending at entry within the same frame, so all six retry timers fire
|
||||
## within the same frame too) — six re-requests every RETRY_DELAY, in sync,
|
||||
## is exactly the "storm" shape worth damping even though it isn't what
|
||||
## caused the starvation. Exponential backoff (INITIAL_RETRY_DELAY doubling
|
||||
## to MAX_RETRY_DELAY) plus a DETERMINISTIC per-tile stagger
|
||||
## (STAGGER_STEP * stagger_index, set once by the owning AtlasWindowTileSet
|
||||
## at construction — see `_stagger_index`) spread that pulse into a trickle:
|
||||
## tile 0 retries at 0.5s, tile 1 at 0.6s, tile 2 at 0.7s, etc. — deterministic
|
||||
## and directly assertable in a test, not a randomized jitter a test would
|
||||
## have to tolerance-check. Backoff ALSO buys a much longer wall-clock window
|
||||
## from a modest MAX_RETRIES increase (~110s at 30 retries, see
|
||||
## _retry_delay_for()'s own doc) without ever polling aggressively for that
|
||||
## whole span. A fast (already-warm) response still resolves on retry #1,
|
||||
## unaffected — backoff/stagger only matter once a request is genuinely
|
||||
## still pending past the first cycle.
|
||||
const INITIAL_RETRY_DELAY: float = 0.5 # first retry, matches the old flat RETRY_DELAY
|
||||
const MAX_RETRY_DELAY: float = 4.0 # backoff ceiling — never polls slower than this
|
||||
const STAGGER_STEP: float = 0.1 # per-tile-index offset — tile i retries STAGGER_STEP*i later
|
||||
const MAX_RETRIES: int = 30 # ~110s wall-clock at the backoff schedule above
|
||||
|
||||
## T-1150 struct/key plumbing: legacy int granularity — district is the
|
||||
## default for every caller that doesn't request quarter/Region explicitly.
|
||||
@@ -89,6 +115,12 @@ var _min_wl_m: int = DEFAULT_MIN_WL_M
|
||||
var _pending: bool = false
|
||||
var _retries: int = 0
|
||||
var _debounce_timer: Timer = null
|
||||
## Cold-start dossier round 3: deterministic per-request stagger index for
|
||||
## the retry backoff (see STAGGER_STEP's own doc) — 0 for the single-window
|
||||
## viewer's own request (no fan-out, nothing to desync from), the tile's own
|
||||
## index (0..5) for a tile-set-owned request (AtlasWindowTileSet.enter()
|
||||
## sets this once at construction, right after AtlasWindowRequest.new()).
|
||||
var _stagger_index: int = 0
|
||||
|
||||
|
||||
func _init(owner_ref = null) -> void:
|
||||
@@ -288,25 +320,44 @@ func _on_debounce_timeout() -> void:
|
||||
## real server, always) -> v2 is the ONLY granularity comparison; absent (a
|
||||
## hypothetically old, pre-T-1152 server) -> fall back to the legacy
|
||||
## comparison alone, matching this object's own pre-T-1152 behavior exactly.
|
||||
## PR #192 cold-start round 3: the coordinator's live cold-server capture
|
||||
## (retries=0, pending=true, forever) exposed that the OLD version of this
|
||||
## function returned unconditionally whenever the WHOLE response's status
|
||||
## wasn't "Ready" — treating a cold body's `status: "Pending"` (the FIRST
|
||||
## request against a whole-body cache miss, before ANY layer including the
|
||||
## window has even been queued — `serve_district_window`/`get_or_generate()`
|
||||
## in server/src/atlas/layer_proxy.rs) identically to `NotFound`/`Error`: a
|
||||
## silent no-op, never reaching the retry-scheduling code at all. Confirmed
|
||||
## server-side: `status: Ready` is set ONLY on the whole-body cache-HIT
|
||||
## branch, entirely independent of whether the WINDOW itself has resolved —
|
||||
## so a cold body's first-ever window request gets `Pending` at the OUTER
|
||||
## layer, while a body someone has already warmed (a later connection, or
|
||||
## this SAME connection's own re-request once its own AnalyzeBody has
|
||||
## landed) gets `Ready` with `district_window: null` inside it, correctly
|
||||
## reaching the retry branch below. Same "still generating" signal, two
|
||||
## different wire shapes depending on which cache warmed first — the fix is
|
||||
## to treat BOTH as the identical retry-worthy state, matching
|
||||
## atlas_generation_proxy.gd's own on_response() `match` shape exactly
|
||||
## (Ready -> handle, Pending -> retry, NotFound/Error -> give up now, not
|
||||
## after MAX_RETRIES: a real error is never going to resolve by waiting).
|
||||
func on_response(response: Dictionary) -> void:
|
||||
if str(response.get("body_id", "")) != _body_id:
|
||||
return
|
||||
if str(response.get("status", "")) != "Ready":
|
||||
return # Pending/NotFound/Error on the WHOLE response — not a window signal either way
|
||||
var status := str(response.get("status", ""))
|
||||
if status == "Pending":
|
||||
_retry_if_pending()
|
||||
return
|
||||
if status != "Ready":
|
||||
_pending = false # NotFound / Error — a real failure, not a queue wait; give up now
|
||||
return
|
||||
var window: Variant = response.get("district_window")
|
||||
if window == null:
|
||||
# §1: an as-yet-underived window rides as `district_window: None` inside
|
||||
# a Ready response — this is the "still generating" signal, not an
|
||||
# error. Re-poll until the background derive lands or the retry
|
||||
# ceiling is hit (queue-based serving, PR #185 — the response lands
|
||||
# on a LATER tick, never this same round-trip).
|
||||
if not _pending:
|
||||
return
|
||||
if _retries < MAX_RETRIES:
|
||||
_retries += 1
|
||||
_schedule_retry()
|
||||
else:
|
||||
_pending = false # gave up — caller's border-fade / empty state persists
|
||||
# §1: an as-yet-underived window rides as `district_window: None`
|
||||
# inside an OUTER-Ready response — the whole-body cache already
|
||||
# warmed, but this specific window hasn't derived yet. Same
|
||||
# "still generating" signal the outer-Pending branch above handles,
|
||||
# just the OTHER wire shape it can arrive in.
|
||||
_retry_if_pending()
|
||||
return
|
||||
|
||||
var w: Dictionary = window
|
||||
@@ -328,6 +379,21 @@ func on_response(response: Dictionary) -> void:
|
||||
window_ready.emit(w)
|
||||
|
||||
|
||||
## Shared "still generating, re-poll" logic for BOTH wire shapes on_response()
|
||||
## can see it in (outer status=="Pending", or inner district_window==null
|
||||
## inside an outer Ready) — re-request until the derive lands or the retry
|
||||
## ceiling is hit (queue-based serving, PR #185 — the response lands on a
|
||||
## LATER tick, never this same round-trip).
|
||||
func _retry_if_pending() -> void:
|
||||
if not _pending:
|
||||
return
|
||||
if _retries < MAX_RETRIES:
|
||||
_retries += 1
|
||||
_schedule_retry()
|
||||
else:
|
||||
_pending = false # gave up — caller's border-fade / empty state persists
|
||||
|
||||
|
||||
## The granularity half of on_response()'s staleness check, split out for the
|
||||
## v2-authoritative-when-present precedence rule (see on_response()'s own
|
||||
## doc for the full live-round rationale). Presence, not value, is the
|
||||
@@ -343,8 +409,23 @@ func _echoed_granularity_matches(w: Dictionary) -> bool:
|
||||
return echoed_granularity == _granularity
|
||||
|
||||
|
||||
## Pure: the exponential-backoff delay for retry attempt number `retry_count`
|
||||
## (1-indexed — the FIRST retry, right after the initial request's own
|
||||
## PENDING answer, uses `retry_count=1`), staggered by `stagger_index`
|
||||
## (STAGGER_STEP*stagger_index added on top — deterministic, not randomized,
|
||||
## so a test can assert the exact delay sequence for tile N directly). Split
|
||||
## out from _schedule_retry() as a pure function for the same reason every
|
||||
## other formula in this file is: directly unit-testable without a live
|
||||
## Timer/SceneTree.
|
||||
static func _retry_delay_for(retry_count: int, stagger_index: int) -> float:
|
||||
var base: float = INITIAL_RETRY_DELAY * pow(2.0, float(maxi(retry_count - 1, 0)))
|
||||
var capped: float = minf(base, MAX_RETRY_DELAY)
|
||||
return capped + STAGGER_STEP * float(stagger_index)
|
||||
|
||||
|
||||
func _schedule_retry() -> void:
|
||||
var timer := get_tree().create_timer(RETRY_DELAY)
|
||||
var delay: float = _retry_delay_for(_retries, _stagger_index)
|
||||
var timer := get_tree().create_timer(delay)
|
||||
timer.timeout.connect(
|
||||
func() -> void:
|
||||
if _pending:
|
||||
|
||||
@@ -97,6 +97,17 @@ func enter(body_id: String, body_radius_km: float) -> void:
|
||||
var center: Vector2i = centers[i]
|
||||
var request = AtlasWindowRequest.new(self)
|
||||
request.name = "Tile%d" % i
|
||||
# Cold-start dossier round 3 hardening: deterministic per-tile retry
|
||||
# stagger (STAGGER_STEP*i) — without it, all 6 tiles go pending in the
|
||||
# same frame and retry in perfect lockstep, a request pulse every
|
||||
# RETRY_DELAY instead of a spread trickle. Set BEFORE request_now()
|
||||
# so it's already in place for the very first retry, if one fires.
|
||||
# Cold-start dossier round 3 hardening: deterministic per-tile retry
|
||||
# stagger (STAGGER_STEP*i) — without it, all 6 tiles go pending in the
|
||||
# same frame and retry in perfect lockstep, a request pulse every
|
||||
# RETRY_DELAY instead of a spread trickle. Set BEFORE request_now()
|
||||
# so it's already in place for the very first retry, if one fires.
|
||||
request._stagger_index = i
|
||||
add_child(request)
|
||||
var tile_index := i # capture by value for the lambda below
|
||||
request.window_ready.connect(
|
||||
|
||||
Reference in New Issue
Block a user