Files
settled-reach/docs/workshops/body-map-viewer/stig-round1.md
T
jpmschweitzerandClaude Fable 5 29c22cb728 docs(meta): body-map-viewer workshop — rounds, measurements, outcomes, as-built briefing
The complete workshop record: four round-1 positions, five round-2 syntheses
(incl. Troblum's adversarial pass with addendum + final scorecard — all seven
findings resolved), both lead interviews, Qatux's round notes and the 8-section
workshop-outcomes.md (the lakes message-crossing documented as process
history), measurement ⑥ (set_pixel/c1) + the population-survey and chunk/S2
addenda in the measurement docs, the brief's appendix updated through ⑥, and
architecture-briefing-final.md — Jeroen's outline written back as-built
(six-level ladder, lakes, ~9MB resident global tier). README row: Complete.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-24 10:57:45 +02:00

20 KiB
Raw Blame History

title, description, type, status, workshop, created, owner
title description type status workshop created owner
Body Map Viewer — Stig's Round 1 Position Client map component structure, step-cross UX, overlay compositing, and the client cache store decision workshop active body-map-viewer 2026-07-25 Stig

Stig — Round 1 Position

Answering the four questions in my slice. Grounded in the measured appendix (①–⑤) and the current client/ui/implant/apps/atlas/ cluster as it exists on disk today, not as remembered from the T-1143 doc.

1. The client map-drawing component

Kill _canvas.scale. There is no more zoom-scaled canvas. Today's model — _canvas.position/_canvas.scale on a Node2D, _apply_transform() setting _canvas.scale = Vector2(_view_zoom, _view_zoom), every nature-overlay draw_* call wrapped in _zs()/_zs_stroke()/_zs_ring_radius() compensation — is the entire error class named in the brief (§2 of prep-grounding). The premise "server determines content, client draws map-art" doesn't just permit retiring that model, it removes the only reason it existed: the scale node was there to let one held composite serve a continuous zoom range. Stepped zoom with a server-resolved canvas per step means the client never needs to render the same data at two different implied resolutions — each step gets its own texture, drawn near enough to 1:1 that "zoom" stops being a client-side transform of derived geometry at all.

Structure: two sibling layers under one step-canvas owner, replacing the single _canvas Node2D.

  • Terrain/classification layer — RTT, texel-exact. One ImageTexture per held step canvas, built server-side-derived / client-colorized, drawn via draw_texture_rect at the tunable px ratio (1:1 ideal, up to 5×5 px/gridunit fallback per the resolution tunable). This is not a new pattern — it's _rebuild_texture_if_needed/_build_tile_texture in atlas_window_overlay.gd generalized from "per-tile mosaic composite" to "the one and only terrain path." The tile-mosaic code (_draw_tile_mosaic/_draw_one_tile/_tile_texture_cache) already proves the RTT-per-cell-block shape works and composites cleanly at orbital rest state — it's not being invented, it's being promoted to universal.
  • Screen-space annotation layer — unscaled sibling, literal px. Everything currently living in atlas_window_geometry_nature.gd's _zs-wrapped draw_polyline/draw_circle/draw_arc calls (rivers, settlement glyphs, POI rings, mouth markers) moves to a sibling Node2D that is never scaled. Positions are world→screen transformed per-frame (cheap — it's a linear map, not a re-derivation); sizes are constants or class-driven constants, never divided by a zoom factor. This deletes _zs(), _zs_stroke(), _zs_ring_radius(), and the entire "did I remember to wrap this call site" failure mode — by construction, not by discipline.

What retires:

  • _canvas.scale/_canvas.position transform model and _apply_transform() as currently shaped (a replacement "step-canvas anchor" concept survives, see step-cross below, but it's not a continuous scale).
  • The _zs/_zs_stroke/_zs_ring_radius compensation family in atlas_window_geometry_nature.gd — wholesale, once the sibling layer lands.
  • select_rung()'s coverage-ceiling walk (Tyre's call to make formally, but it's dead the moment the step index replaces it — my component doesn't need a rung selector, it needs a step index).
  • The AtlasViewer/orbital-mosaic-vs-window split as two code paths — under the stepped model, global zoom is just step 0, not a structurally different viewer. (T-1157's dead-goldens problem is a direct consequence of this split existing today; the redesign should collapse it, not re-target goldens at the same fork.)

What survives as-is:

  • atlas_window_tile_set.gd's LRU-by-key cache shape (see §4 — it's the right skeleton for the new client cache, wrong eviction policy alone).
  • _filter_for_granularity_v2() — NEAREST vs LINEAR sampling filter choice per rung is still a real question at RTT scale-ratio draw time, independent of the transform model change.
  • atlas_window_water_clip.gd, atlas_marker_overlay.gd, atlas_legend_panel.gd, atlas_overlay_bar.gd — these are compositing/UI chrome, not the transform mechanism; they consume whatever the new layer pair exposes and shouldn't need structural rewrites, only call-site updates.

T-1158 (viewer decomposition into input/orchestration/canonical-frame clusters) is subsumed, not scheduled separately — the canonical-frame state machine T-1158 wanted to extract is the exact thing that changes shape under stepped zoom (no more continuous zoom floor, no more _canonical_fit_zoom() fitting a float zoom — the "canonical frame" becomes "step 0"). Extracting the old cluster now would be extracting code about to be deleted. Recommend cancelling T-1158 with the supersession note and letting the new component's structure be decided fresh as part of this implementation (naturally three pieces again — input/pan, step-cross orchestration, and the two draw layers — but that's a consequence of good decomposition, not a retained ticket).

2. Step-cross experience

Hold-fetch-swap, not blend-fetch-swap, as the baseline — a morph is a separate, optional cosmetic layer on top.

Sequence on a scroll-step:

  1. Player scrolls one notch. Compute the new step's data-canvas bounds (cursor-anchored — the center the new canvas should be requested around is the world point under the cursor, matching the surviving entry-seam behavior).
  2. If that canvas is already in the client cache (§4), swap immediately — this is the common case for backtracking (zoom out then back in) and for revisiting a spot, and it's why the cache matters for feel, not just bandwidth.
  3. If not cached, hold the current step's texture displayed, unscaled, while the fetch is in flight — this is the "between-step magnification" red flag's actual mechanism: showing the coarser canvas magnified to fill the new step's viewport for the fetch duration. Measurement ① confirms this window is short: District-class server derive is sub-millisecond served, and even an uncached 330K-gridunit step canvas is ~75 ms derive + ~5 ms PNG-encode server-side (measurement ④'s "derivation cost for context" row) — the hold interval is double-digit milliseconds, not a visible stall.
  4. On arrival: decode, build/update the ImageTexture (measurement ⑤: worst case 8.3M px update is ~4.34.6 ms median, comfortably under one frame), swap the terrain layer, re-run the annotation layer's world→screen transform against the new step's bounds.

ImageTexture upload cost is a non-gating input, confirmed by ⑤. Every canvas size the workshop cites uploads in single-digit ms with no observed frame-budget break, including the frame-delta-spike case. This means the step-count/canvas-size decision (Dudley's/Tyre's call) can be made on derivation cost and wire size alone — upload is not a constraint that trades off against them. One concrete implementation note from ⑤: prefer texture.update() reuse over fresh create_from_image() per step, not for raw speed (reuse is marginally slower in the raw numbers — 4.34.6 ms vs 3.2 ms median at 8.3M px) but because it avoids per-step Texture object churn on the RenderingServer side that the microbenchmark doesn't capture. And: use L8 wherever a plane is genuinely single-channel — 49× cheaper than RGBA8 at every size in ⑤, a free win if any wire field (elevation, a grayscale classification pass) can ship single-channel before the client colorizes it.

The morph/tween is real but strictly cosmetic, and I'd sequence it after the hold-fetch-swap baseline ships, not with it. A cross-fade or scale-tween between the held step-N texture and the arriving step-(N+1) texture, purely in screen space, never touching derived data — this is squarely inside "GPU is presentation only." It softens the perceptual jump D-166's amendment has to own honestly (red flag 1) without pretending to be continuous zoom. I'd implement it as an optional CanvasItem alpha/scale tween gated behind a toggle, not a hard requirement — if it turns out ugly or distracting at real step factors, dropping it costs nothing structurally because it never participates in the data path.

Client-side cache is what makes repeated step-crossing (the actual common case — players hunt around a region, not monotonically zoom in once) feel instant. Per premise 9: a canvas for a fixed seed never changes, so once fetched it is valid forever for that exact (body, step, center) key — the atlas_window_tile_set.gd LRU shape already assumes exactly this ("no freshness check, no TTL, no invalidation path — the only reason an entry leaves is capacity pressure"). The new cache needs the same discipline plus the sim-state carve-out premise 9 calls out (frozen/flooded needs a shorter TTL layered on top of the otherwise-permanent geometry entries) — see §4 for the store shape.

3. Overlay compositing under RTT — shader vs CPU (the c1 decision)

I'd ship CPU set_pixel coloring for the terrain layer first, with a concrete, named follow-up to shader-side compositing once the toggle set grows — not because shaders are wrong, but because the migration risk right now is elsewhere and CPU coloring is what's already proven in-tree.

Reasoning:

  • Today's code already does this and it works. _build_tile_texture/ _rebuild_texture_if_needed build the Image cell-by-cell via set_pixel from typed field values (_cell_color/_base_cell_color/_temp_cell_color/ _moisture_cell_color/_veg_cell_color/_apply_glaciation — five toggle overlays already implemented this way). This is the temperature/moisture/ vegetation/glaciation toggle set the question asks about, already shipped on the CPU path. Moving the transform model (RTT-as-universal, no more scaled canvas) is already the workshop's biggest client-side change; I don't want to also change the coloring mechanism in the same pass without a measured reason to.
  • The measured numbers don't force the shader answer. Nothing in ①–⑤ prices CPU set_pixel coloring cost at step-canvas scale (330K8.3M cells) — that's a real gap, and I'd flag it as a follow-up measurement before committing either way at the deep end of the ladder. What ⑤ does show is that the upload step (which happens regardless of who colors the pixels) is cheap; the open question is purely "how long does building the Image take in GDScript at 330K+ cells," which is untested. Given that gap, defaulting to the known-working CPU path and measuring before the largest canvas sizes ship is the honest sequencing — not picking shaders on the strength of "should be faster" without a number.
  • Where shaders clearly win, and where I'd schedule the follow-up: the moment T-1175's per-vertex river tapering/width-grammar work starts (source tapering, tributary-join width ramps), that work wants shader-side or at minimum Polygon2D-strip geometry regardless of what the terrain layer does — draw_polyline is single-width by construction and can't taper. That's a screen-space-annotation-layer concern (§1), not a terrain-raster concern, and it's already gated on "after the nature layer stands" per T-1175's own scheduling note. If shader-side terrain compositing is adopted later (multiple data textures — morphology/elev/temp/moisture/vegetation/ glaciation — sampled and colorized in a fragment shader), the toggle overlays and T-1175's tapering become the same mechanism for free, which is a genuine architectural win — I'm not against it, I'm against committing to it on zero cost data for the actual bottleneck (CPU set_pixel at 8.3M cells) when a known-good fallback exists.
  • Recommendation for the ticket plan: ship CPU coloring behind the new RTT structure for the initial step-ladder cutover (lowest risk, matches proven code), file a measurement ticket for GDScript Image.set_pixel cost at 330K/2.07M/8.3M cells (the missing sixth measurement), and treat the shader migration as the natural landing spot for T-1175 rather than a precondition for the ladder itself.

4. The client cache store — in-memory vs disk-backed

Answering the three candidate shapes against the fact base (SQLite ships server-side only; client has gdUnit4 + messagepack, no SQLite addon today).

My answer: (iii) plain FileAccess cache dir + index, not (i) server-side SQLite and not (ii) godot-sqlite. In that order of preference, for these reasons:

Against (i) server-side SQLite cache DB. The "client-primary reads as local-machine-primary" argument is real — a warm local subprocess is fast — but it quietly relocates a client-side design decision onto the server's process boundary and turns every cache read into an IPC round-trip instead of an in-process Godot call. It also risks exactly the thing the asset-pipeline golden rule and D-227 both warn about: a second SQLite file living next to systems.db invites confusion about which one is canonical, even with a different filename, because the mental model "the server owns the DB" now has two meanings (canonical snapshot vs cache). D-227's discipline needs to be re-proven at a new file rather than reusing an already-well-understood boundary. Not fatal, but it's the shape with the most governance surface area for the least architectural gain — premise 9 explicitly frames this as a client-side cache question, and routing it through the server subprocess undercuts the "client-primary" framing it's supposed to answer.

Against (ii) godot-sqlite addon. A real client-side store, and if the cache needed relational queries (joins, filtered scans across many keyed records) I'd pick this without hesitation. It doesn't — the access pattern is point lookups by a composite key (body_id:step:center:granularity_v2, extending atlas_window_tile_set.gd's existing make_key() shape almost unchanged) plus a periodic TTL sweep over sim-state-tagged entries. That's a key-value store with expiry, not a relational workload. Pulling in a new compiled addon (build/platform surface, version-pin maintenance, another thing that can fail to load headless in CI) to do a job FileAccess + Time.get_unix_time_from_system() already does natively is the kind of dependency I'd only take if the simpler shape measurably couldn't do the job. It can.

For (iii) plain FileAccess cache dir + index — the shape:

  • Directory layout: one file per cached step-canvas entry under user://atlas_cache/<body_id>/, named by a hash or the same key string atlas_window_tile_set.gd already builds (make_key() extended with the step index) — content is the already-decided wire encoding (PNG-per-field per measurement ④, the clear winner on size and speed), so the disk file IS the wire payload, no re-encoding for storage.
  • Index: one small JSON or binary manifest (user://atlas_cache/index. <body_id>.dat or similar) mapping key → {written_at, last_read_at, kind: geometry|sim_state, size_bytes}. Loaded once per body-open, held in memory as a Dictionary — this is the same LRU-touch shape atlas_window_tile_set. gd already implements (erase+reinsert = move-to-MRU), just backed by files on disk instead of values in the dictionary, and the dictionary now stores metadata + a FileAccess path instead of the raw window Dictionary.
  • Two-tier eviction, matching premise 9's split exactly:
    • Geometry entries (morphology/elev/moisture/vegetation/glaciation/ height, everything D-227's determinism guarantee covers): LRU-evict-only, no TTL, same as today's atlas_window_tile_set.gd — a canvas for a fixed seed never changes, so "stale" isn't a concept that applies. Global step (step 0) entries get a retention floor (never evicted by the LRU sweep, only by an explicit clear/uninstall path) — this is the concrete mechanism for "always keep the global level" from Jeroen's outline, sized by red flag 2's number (~1.6 MB/body at 5×5 sampling; even at all ~273 bodies resident that's ~440 MB on disk, which is a completely different budget conversation than 440 MB in process memory — disk is cheap, RAM/VRAM residency is the thing that needed the ceiling).
    • Sim-state-tagged fields (frozen/flooded, whatever the map-time-axis ruling lands on): explicit TTL sweep, a periodic (not per-frame — on body-open and on a coarse timer) pass over the index removing entries past their TTL regardless of LRU recency. This is the literal self-cleaning-records behavior Jeroen asked for ("on disk probably, but with self cleaning of cache records"), and it's the one place this store needs logic atlas_window_tile_set.gd doesn't have today.
  • Never the source of truth, enforced structurally, not just by comment: if the cache directory is deleted, the client's only behavior change is re-fetching from the server — same guarantee D-227 gives server-side, now proven at the client tier too. Boot-time behavior: missing/corrupt index → treat as empty cache, don't crash, don't block first paint.
  • Sizing from the wire table (④): at the PNG-per-field encoding (the clear winner — smallest and fastest at every size), a 330K-gridunit step canvas is ~638 KB; a full global-tier set across ~273 bodies at the 5×5 fallback resolution is the ~440 MB figure red flag 2 already computed. That number is disk-budget-safe on any target platform (modern discretionary disk cache budgets for a locally-installed game are routinely in the multi-GB range) in a way it is not RAM-safe — which is the strongest single argument for "disk-backed, not memory-only": premise 9's "always keep global" instruction is only affordable at all if it's a disk tier, not permanent process residency. An in-memory-only cache would have to either violate "always keep global" (evict it under memory pressure, defeating the snappy-navigation goal) or accept the ~440 MB RAM floor red flag 2 already calls infeasible. Disk removes that tension entirely.
  • In-memory tier still exists, layered on top, not replaced. The atlas_window_tile_set.gd-style in-process Dictionary cache remains the hot path for "the step I'm looking at right now and its immediate neighbors" — a session-scoped, small (current DEFAULT_MAX_ENTRIES = 24 is a reasonable starting point, tunable) LRU exactly as it works today. The disk store is the tier below it: a miss in memory checks disk before going to the server. Three tiers total, cheapest-first: in-memory Dictionary → disk FileAccess store → server (which has its own tiers, Dudley's to answer). This layering is what makes "client-primary reads as local-machine-primary" actually true without needing the server-SQLite detour — the disk tier gets you nearly all of that latency win already, in-process, no IPC.

Net recommendation: (iii), layered under the existing in-memory LRU as a second tier, two eviction policies (retention-floor+LRU for geometry, TTL-swept for sim-state), sized directly from measurement ④'s table. No new Godot addon, no second SQLite file, no server round-trip for a client-local question.

Open items I'm carrying into round 2

  • The missing sixth measurement: GDScript Image.set_pixel coloring cost at 330K/2.07M/8.3M cells — needed before the c1 shader-vs-CPU call can be made with the same rigor as everything else in this brief. I'd rather name this gap now than let CPU-coloring become an unexamined default at the largest canvas sizes.
  • Step-count/canvas-size is Dudley's/Tyre's number to land, but it directly sizes my annotation-layer redraw frequency and cache key space — I'll adjust §1/§4 sizing once that lands in the lead interview.
  • Phase-5 reuse question (Jeroen's #4): I'd scope the map-drawing component as Atlas-only for this pass. The RTT-terrain + screen-space-annotation split is a reasonable shape for a future in-world viewport too, but Phase-5 in-world rendering has different constraints ( per-frame redraw, character-relative camera, no step-cross discontinuity to design around) that I don't want to speculatively design against now — that would be exactly the kind of later-phase drag the cascade discipline warns against. Build it right for Atlas; revisit reuse when Phase 5 actually starts.