Files
settled-reach/docs/workshops/body-map-viewer/stig-round1.md
T
jpmschweitzerandClaude Fable 5 29c22cb728 docs(meta): body-map-viewer workshop — rounds, measurements, outcomes, as-built briefing
The complete workshop record: four round-1 positions, five round-2 syntheses
(incl. Troblum's adversarial pass with addendum + final scorecard — all seven
findings resolved), both lead interviews, Qatux's round notes and the 8-section
workshop-outcomes.md (the lakes message-crossing documented as process
history), measurement ⑥ (set_pixel/c1) + the population-survey and chunk/S2
addenda in the measurement docs, the brief's appendix updated through ⑥, and
architecture-briefing-final.md — Jeroen's outline written back as-built
(six-level ladder, lakes, ~9MB resident global tier). README row: Complete.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-24 10:57:45 +02:00

334 lines
20 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
---
title: "Body Map Viewer — Stig's Round 1 Position"
description: "Client map component structure, step-cross UX, overlay compositing, and the client cache store decision"
type: workshop
status: active
workshop: body-map-viewer
created: 2026-07-25
owner: Stig
---
# Stig — Round 1 Position
Answering the four questions in my slice. Grounded in the measured appendix
(①–⑤) and the current `client/ui/implant/apps/atlas/` cluster as it exists on
disk today, not as remembered from the T-1143 doc.
## 1. The client map-drawing component
**Kill `_canvas.scale`. There is no more zoom-scaled canvas.** Today's model —
`_canvas.position`/`_canvas.scale` on a Node2D, `_apply_transform()` setting
`_canvas.scale = Vector2(_view_zoom, _view_zoom)`, every nature-overlay
`draw_*` call wrapped in `_zs()`/`_zs_stroke()`/`_zs_ring_radius()` compensation
— is the entire error class named in the brief (§2 of prep-grounding). The
premise "server determines content, client draws map-art" doesn't just permit
retiring that model, it removes the only reason it existed: the scale node was
there to let one held composite serve a continuous zoom range. Stepped zoom
with a server-resolved canvas per step means the client never needs to render
the *same* data at two different implied resolutions — each step gets its own
texture, drawn near enough to 1:1 that "zoom" stops being a client-side
transform of derived geometry at all.
**Structure: two sibling layers under one step-canvas owner, replacing the
single `_canvas` Node2D.**
- **Terrain/classification layer — RTT, texel-exact.** One `ImageTexture` per
held step canvas, built server-side-derived / client-colorized, drawn via
`draw_texture_rect` at the tunable px ratio (1:1 ideal, up to 5×5 px/gridunit
fallback per the resolution tunable). This is **not a new pattern** — it's
`_rebuild_texture_if_needed`/`_build_tile_texture` in
`atlas_window_overlay.gd` generalized from "per-tile mosaic composite" to
"the one and only terrain path." The tile-mosaic code
(`_draw_tile_mosaic`/`_draw_one_tile`/`_tile_texture_cache`) already proves
the RTT-per-cell-block shape works and composites cleanly at orbital rest
state — it's not being invented, it's being promoted to universal.
- **Screen-space annotation layer — unscaled sibling, literal px.** Everything
currently living in `atlas_window_geometry_nature.gd`'s `_zs`-wrapped
`draw_polyline`/`draw_circle`/`draw_arc` calls (rivers, settlement glyphs,
POI rings, mouth markers) moves to a sibling `Node2D` that is **never
scaled**. Positions are world→screen transformed per-frame (cheap — it's a
linear map, not a re-derivation); sizes are constants or class-driven
constants, never divided by a zoom factor. This deletes `_zs()`,
`_zs_stroke()`, `_zs_ring_radius()`, and the entire "did I remember to wrap
this call site" failure mode — by construction, not by discipline.
**What retires:**
- `_canvas.scale`/`_canvas.position` transform model and `_apply_transform()`
as currently shaped (a replacement "step-canvas anchor" concept survives,
see step-cross below, but it's not a continuous scale).
- The `_zs`/`_zs_stroke`/`_zs_ring_radius` compensation family in
`atlas_window_geometry_nature.gd` — wholesale, once the sibling layer lands.
- `select_rung()`'s coverage-ceiling walk (Tyre's call to make formally, but
it's dead the moment the step index replaces it — my component doesn't need
a rung selector, it needs a step index).
- The `AtlasViewer`/orbital-mosaic-vs-window split as two *code paths* — under
the stepped model, global zoom is just step 0, not a structurally different
viewer. (T-1157's dead-goldens problem is a direct consequence of this split
existing today; the redesign should collapse it, not re-target goldens at
the same fork.)
**What survives as-is:**
- `atlas_window_tile_set.gd`'s LRU-by-key cache *shape* (see §4 — it's the
right skeleton for the new client cache, wrong eviction policy alone).
- `_filter_for_granularity_v2()` — NEAREST vs LINEAR sampling filter choice
per rung is still a real question at RTT scale-ratio draw time, independent
of the transform model change.
- `atlas_window_water_clip.gd`, `atlas_marker_overlay.gd`,
`atlas_legend_panel.gd`, `atlas_overlay_bar.gd` — these are compositing/UI
chrome, not the transform mechanism; they consume whatever the new layer
pair exposes and shouldn't need structural rewrites, only call-site updates.
T-1158 (viewer decomposition into input/orchestration/canonical-frame
clusters) is **subsumed, not scheduled separately** — the canonical-frame
state machine T-1158 wanted to extract is the exact thing that changes shape
under stepped zoom (no more continuous zoom floor, no more
`_canonical_fit_zoom()` fitting a float zoom — the "canonical frame" becomes
"step 0"). Extracting the old cluster now would be extracting code about to be
deleted. Recommend cancelling T-1158 with the supersession note and letting
the new component's structure be decided fresh as part of this implementation
(naturally three pieces again — input/pan, step-cross orchestration, and the
two draw layers — but that's a consequence of good decomposition, not a
retained ticket).
## 2. Step-cross experience
**Hold-fetch-swap, not blend-fetch-swap, as the baseline — a morph is a
separate, optional cosmetic layer on top.**
Sequence on a scroll-step:
1. Player scrolls one notch. Compute the new step's data-canvas bounds
(cursor-anchored — the center the new canvas should be requested around is
the world point under the cursor, matching the surviving entry-seam
behavior).
2. If that canvas is already in the client cache (§4), swap immediately —
this is the common case for backtracking (zoom out then back in) and for
revisiting a spot, and it's why the cache matters for *feel*, not just
bandwidth.
3. If not cached, **hold the current step's texture displayed, unscaled,
while the fetch is in flight** — this is the "between-step magnification"
red flag's actual mechanism: showing the coarser canvas magnified to fill
the new step's viewport for the fetch duration. Measurement ① confirms
this window is short: District-class server derive is sub-millisecond
served, and even an uncached 330K-gridunit step canvas is ~75 ms
derive + ~5 ms PNG-encode server-side (measurement ④'s "derivation cost
for context" row) — the hold interval is double-digit milliseconds, not a
visible stall.
4. On arrival: decode, build/update the `ImageTexture` (measurement ⑤:
worst case 8.3M px update is ~4.3–4.6 ms median, comfortably under one
frame), swap the terrain layer, re-run the annotation layer's
world→screen transform against the new step's bounds.
**ImageTexture upload cost is a non-gating input, confirmed by ⑤.** Every
canvas size the workshop cites uploads in single-digit ms with no observed
frame-budget break, including the frame-delta-spike case. This means the
step-count/canvas-size decision (Dudley's/Tyre's call) can be made on
derivation cost and wire size alone — upload is not a constraint that trades
off against them. One concrete implementation note from ⑤: **prefer
`texture.update()` reuse over fresh `create_from_image()` per step**, not for
raw speed (reuse is marginally *slower* in the raw numbers — 4.3–4.6 ms vs
3.2 ms median at 8.3M px) but because it avoids per-step Texture object churn
on the RenderingServer side that the microbenchmark doesn't capture. And:
**use L8 wherever a plane is genuinely single-channel** — 4–9× cheaper than
RGBA8 at every size in ⑤, a free win if any wire field (elevation, a
grayscale classification pass) can ship single-channel before the client
colorizes it.
**The morph/tween is real but strictly cosmetic, and I'd sequence it after
the hold-fetch-swap baseline ships, not with it.** A cross-fade or scale-tween
between the held step-N texture and the arriving step-(N+1) texture, purely
in screen space, never touching derived data — this is squarely inside "GPU is
presentation only." It softens the perceptual jump D-166's amendment has to
own honestly (red flag 1) without pretending to be continuous zoom. I'd
implement it as an optional `CanvasItem` alpha/scale tween gated behind a
toggle, not a hard requirement — if it turns out ugly or distracting at real
step factors, dropping it costs nothing structurally because it never
participates in the data path.
**Client-side cache is what makes repeated step-crossing (the actual common
case — players hunt around a region, not monotonically zoom in once) feel
instant.** Per premise 9: a canvas for a fixed seed never changes, so once
fetched it is valid forever for that exact (body, step, center) key — the
`atlas_window_tile_set.gd` LRU shape already assumes exactly this ("no
freshness check, no TTL, no invalidation path — the only reason an entry
leaves is capacity pressure"). The new cache needs the same discipline **plus**
the sim-state carve-out premise 9 calls out (frozen/flooded needs a shorter
TTL layered on top of the otherwise-permanent geometry entries) — see §4 for
the store shape.
## 3. Overlay compositing under RTT — shader vs CPU (the c1 decision)
**I'd ship CPU `set_pixel` coloring for the terrain layer first, with a
concrete, named follow-up to shader-side compositing once the toggle set
grows — not because shaders are wrong, but because the migration risk right
now is elsewhere and CPU coloring is what's already proven in-tree.**
Reasoning:
- **Today's code already does this and it works.** `_build_tile_texture`/
`_rebuild_texture_if_needed` build the `Image` cell-by-cell via `set_pixel`
from typed field values (`_cell_color`/`_base_cell_color`/`_temp_cell_color`/
`_moisture_cell_color`/`_veg_cell_color`/`_apply_glaciation` — five toggle
overlays already implemented this way). This is the temperature/moisture/
vegetation/glaciation toggle set the question asks about, already shipped
on the CPU path. Moving the *transform model* (RTT-as-universal, no more
scaled canvas) is already the workshop's biggest client-side change; I don't
want to also change the *coloring* mechanism in the same pass without a
measured reason to.
- **The measured numbers don't force the shader answer.** Nothing in ①–⑤
prices CPU `set_pixel` coloring cost at step-canvas scale (330K–8.3M
cells) — that's a real gap, and I'd flag it as a follow-up measurement
before committing either way at the deep end of the ladder. What ⑤ *does*
show is that the upload step (which happens regardless of who colors the
pixels) is cheap; the open question is purely "how long does building the
`Image` take in GDScript at 330K+ cells," which is untested. Given that gap,
defaulting to the known-working CPU path and measuring before the largest
canvas sizes ship is the honest sequencing — not picking shaders on the
strength of "should be faster" without a number.
- **Where shaders clearly win, and where I'd schedule the follow-up:** the
moment T-1175's per-vertex river tapering/width-grammar work starts (source
tapering, tributary-join width ramps), that work wants shader-side or at
minimum `Polygon2D`-strip geometry regardless of what the terrain layer
does — `draw_polyline` is single-width by construction and can't taper.
That's a screen-space-annotation-layer concern (§1), not a terrain-raster
concern, and it's already gated on "after the nature layer stands" per
T-1175's own scheduling note. If shader-side terrain compositing is adopted
later (multiple data textures — morphology/elev/temp/moisture/vegetation/
glaciation — sampled and colorized in a fragment shader), the toggle
overlays and T-1175's tapering become the same mechanism for free, which is
a genuine architectural win — I'm not against it, I'm against committing to
it on zero cost data for the actual bottleneck (CPU set_pixel at 8.3M
cells) when a known-good fallback exists.
- **Recommendation for the ticket plan:** ship CPU coloring behind the new RTT
structure for the initial step-ladder cutover (lowest risk, matches proven
code), file a measurement ticket for GDScript `Image.set_pixel` cost at
330K/2.07M/8.3M cells (the missing sixth measurement), and treat the shader
migration as the natural landing spot for T-1175 rather than a
precondition for the ladder itself.
## 4. The client cache store — in-memory vs disk-backed
Answering the three candidate shapes against the fact base (SQLite ships
server-side only; client has gdUnit4 + messagepack, no SQLite addon today).
**My answer: (iii) plain `FileAccess` cache dir + index, not (i) server-side
SQLite and not (ii) godot-sqlite.** In that order of preference, for these
reasons:
**Against (i) server-side SQLite cache DB.** The "client-primary reads as
local-machine-primary" argument is real — a warm local subprocess is fast —
but it quietly relocates a client-side design decision onto the server's
process boundary and turns every cache read into an IPC round-trip instead of
an in-process Godot call. It also risks exactly the thing the asset-pipeline
golden rule and D-227 both warn about: a second SQLite file living next to
`systems.db` invites confusion about which one is canonical, even with a
different filename, because the *mental model* "the server owns the DB" now
has two meanings (canonical snapshot vs cache). D-227's discipline needs to
be re-proven at a new file rather than reusing an already-well-understood
boundary. Not fatal, but it's the shape with the most governance surface area
for the least architectural gain — premise 9 explicitly frames this as a
*client*-side cache question, and routing it through the server subprocess
undercuts the "client-primary" framing it's supposed to answer.
**Against (ii) godot-sqlite addon.** A real client-side store, and if the
cache needed relational queries (joins, filtered scans across many keyed
records) I'd pick this without hesitation. It doesn't — the access pattern is
point lookups by a composite key (`body_id:step:center:granularity_v2`,
extending `atlas_window_tile_set.gd`'s existing `make_key()` shape almost
unchanged) plus a periodic TTL sweep over sim-state-tagged entries. That's a
key-value store with expiry, not a relational workload. Pulling in a new
compiled addon (build/platform surface, version-pin maintenance, another
thing that can fail to load headless in CI) to do a job `FileAccess` +
`Time.get_unix_time_from_system()` already does natively is the kind of
dependency I'd only take if the simpler shape measurably couldn't do the job.
It can.
**For (iii) plain `FileAccess` cache dir + index — the shape:**
- **Directory layout:** one file per cached step-canvas entry under
`user://atlas_cache/<body_id>/`, named by a hash or the same key string
`atlas_window_tile_set.gd` already builds (`make_key()` extended with the
step index) — content is the already-decided wire encoding (PNG-per-field
per measurement ④, the clear winner on size *and* speed), so the disk file
IS the wire payload, no re-encoding for storage.
- **Index:** one small JSON or binary manifest (`user://atlas_cache/index.
<body_id>.dat` or similar) mapping key → `{written_at, last_read_at, kind:
geometry|sim_state, size_bytes}`. Loaded once per body-open, held in memory
as a `Dictionary` — this is the same LRU-touch shape `atlas_window_tile_set.
gd` already implements (erase+reinsert = move-to-MRU), just backed by files
on disk instead of values in the dictionary, and the dictionary now stores
metadata + a `FileAccess` path instead of the raw window `Dictionary`.
- **Two-tier eviction, matching premise 9's split exactly:**
- **Geometry entries (morphology/elev/moisture/vegetation/glaciation/
height, everything D-227's determinism guarantee covers): LRU-evict-only,
no TTL**, same as today's `atlas_window_tile_set.gd` — a canvas for a
fixed seed never changes, so "stale" isn't a concept that applies. Global
step (step 0) entries get a **retention floor** (never evicted by the LRU
sweep, only by an explicit clear/uninstall path) — this is the concrete
mechanism for "always keep the global level" from Jeroen's outline, sized
by red flag 2's number (~1.6 MB/body at 5×5 sampling; even at all ~273
bodies resident that's ~440 MB on disk, which is a completely different
budget conversation than 440 MB in *process memory* — disk is cheap,
RAM/VRAM residency is the thing that needed the ceiling).
- **Sim-state-tagged fields (frozen/flooded, whatever the map-time-axis
ruling lands on): explicit TTL sweep**, a periodic (not per-frame — on
body-open and on a coarse timer) pass over the index removing entries
past their TTL regardless of LRU recency. This is the literal
self-cleaning-records behavior Jeroen asked for ("on disk probably, but
with self cleaning of cache records"), and it's the one place this store
needs logic `atlas_window_tile_set.gd` doesn't have today.
- **Never the source of truth, enforced structurally, not just by
comment:** if the cache directory is deleted, the client's only behavior
change is re-fetching from the server — same guarantee D-227 gives
server-side, now proven at the client tier too. Boot-time behavior:
missing/corrupt index → treat as empty cache, don't crash, don't block
first paint.
- **Sizing from the wire table (④):** at the PNG-per-field encoding
(the clear winner — smallest and fastest at every size), a 330K-gridunit
step canvas is ~638 KB; a full global-tier set across ~273 bodies at the
5×5 fallback resolution is the ~440 MB figure red flag 2 already computed.
That number is **disk-budget-safe on any target platform** (modern
discretionary disk cache budgets for a locally-installed game are routinely
in the multi-GB range) in a way it is not RAM-safe — which is the
strongest single argument for "disk-backed, not memory-only": premise 9's
"always keep global" instruction is only affordable at all if it's a disk
tier, not permanent process residency. An in-memory-only cache would have
to either violate "always keep global" (evict it under memory pressure,
defeating the snappy-navigation goal) or accept the ~440 MB RAM floor red
flag 2 already calls infeasible. Disk removes that tension entirely.
- **In-memory tier still exists, layered on top, not replaced.** The
`atlas_window_tile_set.gd`-style in-process Dictionary cache remains the
hot path for "the step I'm looking at right now and its immediate
neighbors" — a session-scoped, small (current `DEFAULT_MAX_ENTRIES = 24`
is a reasonable starting point, tunable) LRU exactly as it works today.
The disk store is the tier below it: a miss in memory checks disk before
going to the server. Three tiers total, cheapest-first: in-memory Dictionary
→ disk `FileAccess` store → server (which has its own tiers, Dudley's to
answer). This layering is what makes "client-primary reads as
local-machine-primary" actually true without needing the server-SQLite
detour — the disk tier gets you nearly all of that latency win already,
in-process, no IPC.
**Net recommendation:** (iii), layered under the existing in-memory LRU as a
second tier, two eviction policies (retention-floor+LRU for geometry,
TTL-swept for sim-state), sized directly from measurement ④'s table. No new
Godot addon, no second SQLite file, no server round-trip for a client-local
question.
## Open items I'm carrying into round 2
- **The missing sixth measurement:** GDScript `Image.set_pixel` coloring cost
at 330K/2.07M/8.3M cells — needed before the c1 shader-vs-CPU call can be
made with the same rigor as everything else in this brief. I'd rather name
this gap now than let CPU-coloring become an unexamined default at the
largest canvas sizes.
- **Step-count/canvas-size** is Dudley's/Tyre's number to land, but it
directly sizes my annotation-layer redraw frequency and cache key space —
I'll adjust §1/§4 sizing once that lands in the lead interview.
- **Phase-5 reuse question (Jeroen's #4):** I'd scope the map-drawing
component as **Atlas-only for this pass**. The RTT-terrain +
screen-space-annotation split is a reasonable shape for a future in-world
viewport too, but Phase-5 in-world rendering has different constraints (
per-frame redraw, character-relative camera, no step-cross discontinuity to
design around) that I don't want to speculatively design against now — that
would be exactly the kind of later-phase drag the cascade discipline warns
against. Build it right for Atlas; revisit reuse when Phase 5 actually
starts.