--- title: "Body Map Viewer — Stig's Round 1 Position" description: "Client map component structure, step-cross UX, overlay compositing, and the client cache store decision" type: workshop status: active workshop: body-map-viewer created: 2026-07-25 owner: Stig --- # Stig — Round 1 Position Answering the four questions in my slice. Grounded in the measured appendix (①–⑤) and the current `client/ui/implant/apps/atlas/` cluster as it exists on disk today, not as remembered from the T-1143 doc. ## 1. The client map-drawing component **Kill `_canvas.scale`. There is no more zoom-scaled canvas.** Today's model — `_canvas.position`/`_canvas.scale` on a Node2D, `_apply_transform()` setting `_canvas.scale = Vector2(_view_zoom, _view_zoom)`, every nature-overlay `draw_*` call wrapped in `_zs()`/`_zs_stroke()`/`_zs_ring_radius()` compensation — is the entire error class named in the brief (§2 of prep-grounding). The premise "server determines content, client draws map-art" doesn't just permit retiring that model, it removes the only reason it existed: the scale node was there to let one held composite serve a continuous zoom range. Stepped zoom with a server-resolved canvas per step means the client never needs to render the *same* data at two different implied resolutions — each step gets its own texture, drawn near enough to 1:1 that "zoom" stops being a client-side transform of derived geometry at all. **Structure: two sibling layers under one step-canvas owner, replacing the single `_canvas` Node2D.** - **Terrain/classification layer — RTT, texel-exact.** One `ImageTexture` per held step canvas, built server-side-derived / client-colorized, drawn via `draw_texture_rect` at the tunable px ratio (1:1 ideal, up to 5×5 px/gridunit fallback per the resolution tunable). This is **not a new pattern** — it's `_rebuild_texture_if_needed`/`_build_tile_texture` in `atlas_window_overlay.gd` generalized from "per-tile mosaic composite" to "the one and only terrain path." The tile-mosaic code (`_draw_tile_mosaic`/`_draw_one_tile`/`_tile_texture_cache`) already proves the RTT-per-cell-block shape works and composites cleanly at orbital rest state — it's not being invented, it's being promoted to universal. - **Screen-space annotation layer — unscaled sibling, literal px.** Everything currently living in `atlas_window_geometry_nature.gd`'s `_zs`-wrapped `draw_polyline`/`draw_circle`/`draw_arc` calls (rivers, settlement glyphs, POI rings, mouth markers) moves to a sibling `Node2D` that is **never scaled**. Positions are world→screen transformed per-frame (cheap — it's a linear map, not a re-derivation); sizes are constants or class-driven constants, never divided by a zoom factor. This deletes `_zs()`, `_zs_stroke()`, `_zs_ring_radius()`, and the entire "did I remember to wrap this call site" failure mode — by construction, not by discipline. **What retires:** - `_canvas.scale`/`_canvas.position` transform model and `_apply_transform()` as currently shaped (a replacement "step-canvas anchor" concept survives, see step-cross below, but it's not a continuous scale). - The `_zs`/`_zs_stroke`/`_zs_ring_radius` compensation family in `atlas_window_geometry_nature.gd` — wholesale, once the sibling layer lands. - `select_rung()`'s coverage-ceiling walk (Tyre's call to make formally, but it's dead the moment the step index replaces it — my component doesn't need a rung selector, it needs a step index). - The `AtlasViewer`/orbital-mosaic-vs-window split as two *code paths* — under the stepped model, global zoom is just step 0, not a structurally different viewer. (T-1157's dead-goldens problem is a direct consequence of this split existing today; the redesign should collapse it, not re-target goldens at the same fork.) **What survives as-is:** - `atlas_window_tile_set.gd`'s LRU-by-key cache *shape* (see §4 — it's the right skeleton for the new client cache, wrong eviction policy alone). - `_filter_for_granularity_v2()` — NEAREST vs LINEAR sampling filter choice per rung is still a real question at RTT scale-ratio draw time, independent of the transform model change. - `atlas_window_water_clip.gd`, `atlas_marker_overlay.gd`, `atlas_legend_panel.gd`, `atlas_overlay_bar.gd` — these are compositing/UI chrome, not the transform mechanism; they consume whatever the new layer pair exposes and shouldn't need structural rewrites, only call-site updates. T-1158 (viewer decomposition into input/orchestration/canonical-frame clusters) is **subsumed, not scheduled separately** — the canonical-frame state machine T-1158 wanted to extract is the exact thing that changes shape under stepped zoom (no more continuous zoom floor, no more `_canonical_fit_zoom()` fitting a float zoom — the "canonical frame" becomes "step 0"). Extracting the old cluster now would be extracting code about to be deleted. Recommend cancelling T-1158 with the supersession note and letting the new component's structure be decided fresh as part of this implementation (naturally three pieces again — input/pan, step-cross orchestration, and the two draw layers — but that's a consequence of good decomposition, not a retained ticket). ## 2. Step-cross experience **Hold-fetch-swap, not blend-fetch-swap, as the baseline — a morph is a separate, optional cosmetic layer on top.** Sequence on a scroll-step: 1. Player scrolls one notch. Compute the new step's data-canvas bounds (cursor-anchored — the center the new canvas should be requested around is the world point under the cursor, matching the surviving entry-seam behavior). 2. If that canvas is already in the client cache (§4), swap immediately — this is the common case for backtracking (zoom out then back in) and for revisiting a spot, and it's why the cache matters for *feel*, not just bandwidth. 3. If not cached, **hold the current step's texture displayed, unscaled, while the fetch is in flight** — this is the "between-step magnification" red flag's actual mechanism: showing the coarser canvas magnified to fill the new step's viewport for the fetch duration. Measurement ① confirms this window is short: District-class server derive is sub-millisecond served, and even an uncached 330K-gridunit step canvas is ~75 ms derive + ~5 ms PNG-encode server-side (measurement ④'s "derivation cost for context" row) — the hold interval is double-digit milliseconds, not a visible stall. 4. On arrival: decode, build/update the `ImageTexture` (measurement ⑤: worst case 8.3M px update is ~4.3–4.6 ms median, comfortably under one frame), swap the terrain layer, re-run the annotation layer's world→screen transform against the new step's bounds. **ImageTexture upload cost is a non-gating input, confirmed by ⑤.** Every canvas size the workshop cites uploads in single-digit ms with no observed frame-budget break, including the frame-delta-spike case. This means the step-count/canvas-size decision (Dudley's/Tyre's call) can be made on derivation cost and wire size alone — upload is not a constraint that trades off against them. One concrete implementation note from ⑤: **prefer `texture.update()` reuse over fresh `create_from_image()` per step**, not for raw speed (reuse is marginally *slower* in the raw numbers — 4.3–4.6 ms vs 3.2 ms median at 8.3M px) but because it avoids per-step Texture object churn on the RenderingServer side that the microbenchmark doesn't capture. And: **use L8 wherever a plane is genuinely single-channel** — 4–9× cheaper than RGBA8 at every size in ⑤, a free win if any wire field (elevation, a grayscale classification pass) can ship single-channel before the client colorizes it. **The morph/tween is real but strictly cosmetic, and I'd sequence it after the hold-fetch-swap baseline ships, not with it.** A cross-fade or scale-tween between the held step-N texture and the arriving step-(N+1) texture, purely in screen space, never touching derived data — this is squarely inside "GPU is presentation only." It softens the perceptual jump D-166's amendment has to own honestly (red flag 1) without pretending to be continuous zoom. I'd implement it as an optional `CanvasItem` alpha/scale tween gated behind a toggle, not a hard requirement — if it turns out ugly or distracting at real step factors, dropping it costs nothing structurally because it never participates in the data path. **Client-side cache is what makes repeated step-crossing (the actual common case — players hunt around a region, not monotonically zoom in once) feel instant.** Per premise 9: a canvas for a fixed seed never changes, so once fetched it is valid forever for that exact (body, step, center) key — the `atlas_window_tile_set.gd` LRU shape already assumes exactly this ("no freshness check, no TTL, no invalidation path — the only reason an entry leaves is capacity pressure"). The new cache needs the same discipline **plus** the sim-state carve-out premise 9 calls out (frozen/flooded needs a shorter TTL layered on top of the otherwise-permanent geometry entries) — see §4 for the store shape. ## 3. Overlay compositing under RTT — shader vs CPU (the c1 decision) **I'd ship CPU `set_pixel` coloring for the terrain layer first, with a concrete, named follow-up to shader-side compositing once the toggle set grows — not because shaders are wrong, but because the migration risk right now is elsewhere and CPU coloring is what's already proven in-tree.** Reasoning: - **Today's code already does this and it works.** `_build_tile_texture`/ `_rebuild_texture_if_needed` build the `Image` cell-by-cell via `set_pixel` from typed field values (`_cell_color`/`_base_cell_color`/`_temp_cell_color`/ `_moisture_cell_color`/`_veg_cell_color`/`_apply_glaciation` — five toggle overlays already implemented this way). This is the temperature/moisture/ vegetation/glaciation toggle set the question asks about, already shipped on the CPU path. Moving the *transform model* (RTT-as-universal, no more scaled canvas) is already the workshop's biggest client-side change; I don't want to also change the *coloring* mechanism in the same pass without a measured reason to. - **The measured numbers don't force the shader answer.** Nothing in ①–⑤ prices CPU `set_pixel` coloring cost at step-canvas scale (330K–8.3M cells) — that's a real gap, and I'd flag it as a follow-up measurement before committing either way at the deep end of the ladder. What ⑤ *does* show is that the upload step (which happens regardless of who colors the pixels) is cheap; the open question is purely "how long does building the `Image` take in GDScript at 330K+ cells," which is untested. Given that gap, defaulting to the known-working CPU path and measuring before the largest canvas sizes ship is the honest sequencing — not picking shaders on the strength of "should be faster" without a number. - **Where shaders clearly win, and where I'd schedule the follow-up:** the moment T-1175's per-vertex river tapering/width-grammar work starts (source tapering, tributary-join width ramps), that work wants shader-side or at minimum `Polygon2D`-strip geometry regardless of what the terrain layer does — `draw_polyline` is single-width by construction and can't taper. That's a screen-space-annotation-layer concern (§1), not a terrain-raster concern, and it's already gated on "after the nature layer stands" per T-1175's own scheduling note. If shader-side terrain compositing is adopted later (multiple data textures — morphology/elev/temp/moisture/vegetation/ glaciation — sampled and colorized in a fragment shader), the toggle overlays and T-1175's tapering become the same mechanism for free, which is a genuine architectural win — I'm not against it, I'm against committing to it on zero cost data for the actual bottleneck (CPU set_pixel at 8.3M cells) when a known-good fallback exists. - **Recommendation for the ticket plan:** ship CPU coloring behind the new RTT structure for the initial step-ladder cutover (lowest risk, matches proven code), file a measurement ticket for GDScript `Image.set_pixel` cost at 330K/2.07M/8.3M cells (the missing sixth measurement), and treat the shader migration as the natural landing spot for T-1175 rather than a precondition for the ladder itself. ## 4. The client cache store — in-memory vs disk-backed Answering the three candidate shapes against the fact base (SQLite ships server-side only; client has gdUnit4 + messagepack, no SQLite addon today). **My answer: (iii) plain `FileAccess` cache dir + index, not (i) server-side SQLite and not (ii) godot-sqlite.** In that order of preference, for these reasons: **Against (i) server-side SQLite cache DB.** The "client-primary reads as local-machine-primary" argument is real — a warm local subprocess is fast — but it quietly relocates a client-side design decision onto the server's process boundary and turns every cache read into an IPC round-trip instead of an in-process Godot call. It also risks exactly the thing the asset-pipeline golden rule and D-227 both warn about: a second SQLite file living next to `systems.db` invites confusion about which one is canonical, even with a different filename, because the *mental model* "the server owns the DB" now has two meanings (canonical snapshot vs cache). D-227's discipline needs to be re-proven at a new file rather than reusing an already-well-understood boundary. Not fatal, but it's the shape with the most governance surface area for the least architectural gain — premise 9 explicitly frames this as a *client*-side cache question, and routing it through the server subprocess undercuts the "client-primary" framing it's supposed to answer. **Against (ii) godot-sqlite addon.** A real client-side store, and if the cache needed relational queries (joins, filtered scans across many keyed records) I'd pick this without hesitation. It doesn't — the access pattern is point lookups by a composite key (`body_id:step:center:granularity_v2`, extending `atlas_window_tile_set.gd`'s existing `make_key()` shape almost unchanged) plus a periodic TTL sweep over sim-state-tagged entries. That's a key-value store with expiry, not a relational workload. Pulling in a new compiled addon (build/platform surface, version-pin maintenance, another thing that can fail to load headless in CI) to do a job `FileAccess` + `Time.get_unix_time_from_system()` already does natively is the kind of dependency I'd only take if the simpler shape measurably couldn't do the job. It can. **For (iii) plain `FileAccess` cache dir + index — the shape:** - **Directory layout:** one file per cached step-canvas entry under `user://atlas_cache//`, named by a hash or the same key string `atlas_window_tile_set.gd` already builds (`make_key()` extended with the step index) — content is the already-decided wire encoding (PNG-per-field per measurement ④, the clear winner on size *and* speed), so the disk file IS the wire payload, no re-encoding for storage. - **Index:** one small JSON or binary manifest (`user://atlas_cache/index. .dat` or similar) mapping key → `{written_at, last_read_at, kind: geometry|sim_state, size_bytes}`. Loaded once per body-open, held in memory as a `Dictionary` — this is the same LRU-touch shape `atlas_window_tile_set. gd` already implements (erase+reinsert = move-to-MRU), just backed by files on disk instead of values in the dictionary, and the dictionary now stores metadata + a `FileAccess` path instead of the raw window `Dictionary`. - **Two-tier eviction, matching premise 9's split exactly:** - **Geometry entries (morphology/elev/moisture/vegetation/glaciation/ height, everything D-227's determinism guarantee covers): LRU-evict-only, no TTL**, same as today's `atlas_window_tile_set.gd` — a canvas for a fixed seed never changes, so "stale" isn't a concept that applies. Global step (step 0) entries get a **retention floor** (never evicted by the LRU sweep, only by an explicit clear/uninstall path) — this is the concrete mechanism for "always keep the global level" from Jeroen's outline, sized by red flag 2's number (~1.6 MB/body at 5×5 sampling; even at all ~273 bodies resident that's ~440 MB on disk, which is a completely different budget conversation than 440 MB in *process memory* — disk is cheap, RAM/VRAM residency is the thing that needed the ceiling). - **Sim-state-tagged fields (frozen/flooded, whatever the map-time-axis ruling lands on): explicit TTL sweep**, a periodic (not per-frame — on body-open and on a coarse timer) pass over the index removing entries past their TTL regardless of LRU recency. This is the literal self-cleaning-records behavior Jeroen asked for ("on disk probably, but with self cleaning of cache records"), and it's the one place this store needs logic `atlas_window_tile_set.gd` doesn't have today. - **Never the source of truth, enforced structurally, not just by comment:** if the cache directory is deleted, the client's only behavior change is re-fetching from the server — same guarantee D-227 gives server-side, now proven at the client tier too. Boot-time behavior: missing/corrupt index → treat as empty cache, don't crash, don't block first paint. - **Sizing from the wire table (④):** at the PNG-per-field encoding (the clear winner — smallest and fastest at every size), a 330K-gridunit step canvas is ~638 KB; a full global-tier set across ~273 bodies at the 5×5 fallback resolution is the ~440 MB figure red flag 2 already computed. That number is **disk-budget-safe on any target platform** (modern discretionary disk cache budgets for a locally-installed game are routinely in the multi-GB range) in a way it is not RAM-safe — which is the strongest single argument for "disk-backed, not memory-only": premise 9's "always keep global" instruction is only affordable at all if it's a disk tier, not permanent process residency. An in-memory-only cache would have to either violate "always keep global" (evict it under memory pressure, defeating the snappy-navigation goal) or accept the ~440 MB RAM floor red flag 2 already calls infeasible. Disk removes that tension entirely. - **In-memory tier still exists, layered on top, not replaced.** The `atlas_window_tile_set.gd`-style in-process Dictionary cache remains the hot path for "the step I'm looking at right now and its immediate neighbors" — a session-scoped, small (current `DEFAULT_MAX_ENTRIES = 24` is a reasonable starting point, tunable) LRU exactly as it works today. The disk store is the tier below it: a miss in memory checks disk before going to the server. Three tiers total, cheapest-first: in-memory Dictionary → disk `FileAccess` store → server (which has its own tiers, Dudley's to answer). This layering is what makes "client-primary reads as local-machine-primary" actually true without needing the server-SQLite detour — the disk tier gets you nearly all of that latency win already, in-process, no IPC. **Net recommendation:** (iii), layered under the existing in-memory LRU as a second tier, two eviction policies (retention-floor+LRU for geometry, TTL-swept for sim-state), sized directly from measurement ④'s table. No new Godot addon, no second SQLite file, no server round-trip for a client-local question. ## Open items I'm carrying into round 2 - **The missing sixth measurement:** GDScript `Image.set_pixel` coloring cost at 330K/2.07M/8.3M cells — needed before the c1 shader-vs-CPU call can be made with the same rigor as everything else in this brief. I'd rather name this gap now than let CPU-coloring become an unexamined default at the largest canvas sizes. - **Step-count/canvas-size** is Dudley's/Tyre's number to land, but it directly sizes my annotation-layer redraw frequency and cache key space — I'll adjust §1/§4 sizing once that lands in the lead interview. - **Phase-5 reuse question (Jeroen's #4):** I'd scope the map-drawing component as **Atlas-only for this pass**. The RTT-terrain + screen-space-annotation split is a reasonable shape for a future in-world viewport too, but Phase-5 in-world rendering has different constraints ( per-frame redraw, character-relative camera, no step-cross discontinuity to design around) that I don't want to speculatively design against now — that would be exactly the kind of later-phase drag the cascade discipline warns against. Build it right for Atlas; revisit reuse when Phase 5 actually starts.