The complete workshop record: four round-1 positions, five round-2 syntheses (incl. Troblum's adversarial pass with addendum + final scorecard — all seven findings resolved), both lead interviews, Qatux's round notes and the 8-section workshop-outcomes.md (the lakes message-crossing documented as process history), measurement ⑥ (set_pixel/c1) + the population-survey and chunk/S2 addenda in the measurement docs, the brief's appendix updated through ⑥, and architecture-briefing-final.md — Jeroen's outline written back as-built (six-level ladder, lakes, ~9MB resident global tier). README row: Complete. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
352 lines
22 KiB
Markdown
352 lines
22 KiB
Markdown
---
|
||
title: "Body Map Viewer — Dudley Round 1"
|
||
description: "Server derivation position: hydrology algorithm + cliff representation, step-canvas generation + canonical-vs-viewport, compute-chunk partitioning, cache tiers"
|
||
workshop: body-map-viewer
|
||
round: 1
|
||
owner: Dudley
|
||
status: complete
|
||
decision_refs: [D-166, D-226, D-227, D-243, D-203]
|
||
---
|
||
|
||
# Dudley — Round 1 Position
|
||
|
||
All four numbers I was gated on are in the appendix and I helped produce three
|
||
of them (①②③, plus half of ④'s framing). I'm not going to re-derive them here —
|
||
I'll cite and build on them. Nothing below is extrapolated past what
|
||
`t1177-hydrology.md`, `t1178-t1154-derive-bench.md`, and `t1179-wire-table.md`
|
||
measured. Where I lean on Stig's ⑤ for the render-side implication of a
|
||
server decision, I'll say so explicitly.
|
||
|
||
## Question 1 — Hydrology: algorithm, cost, cliff representation
|
||
|
||
**Algorithm: priority-flood (Barnes/Planchon-Darboux class) + Dijkstra overflow
|
||
routing, exactly as prototyped in T-1177.** Min-heap flood fill from
|
||
below-sea-level + grid-edge seeds finds each basin's true spill point; ascending-
|
||
spill-level basin processing lets a resolved lower basin be a valid target for a
|
||
higher basin's overflow search; carving applies only where peak flow accumulation
|
||
on the carved path clears `RIVER_THRESHOLD = 200`. This is a **pure function**
|
||
(`solve(elevation, width, height, sea_level, climate) -> HydrologyResult`) — no
|
||
tick loop, no RNG, no wall-clock — which is the only shape D-010 and D-227 permit
|
||
for anything that later feeds a save-relevant derivation. I'm recommending we
|
||
adopt it as-is; the prototype module (`server/src/atlas/hydrology_equilibrium.rs`)
|
||
is production-shaped, not throwaway.
|
||
|
||
**Cost, both regimes measured:**
|
||
|
||
- **Single-body solve** (the real production shape — hydrology settles once per
|
||
body, like `drainage::analyze` does today): 512×256 (real GJ1c working grid) ≈
|
||
24 ms single-threaded. This is cheap enough to not need parallelizing internally
|
||
at all.
|
||
- **All 273 bodies, Rayon `par_iter` across bodies** (not within one solve — see
|
||
below for why): **0.7–0.8 s total**, both initial and stability-re-run passes
|
||
agree within 1-3%. This is the number that answers "can we settle hydrology for
|
||
the whole Reach at world-open without it being felt": yes, comfortably, and it
|
||
parallelizes on the *correct* axis.
|
||
- **8.3M-cell single grid** (the stress ceiling, not a real request shape): ~5.7 s
|
||
single-threaded. I want this number on record because it's the honest "computer
|
||
catches fire" case Tyre's implications doc asked for — but no production path
|
||
needs to pay it. Hydrology is a **per-body Layer-1 property**, computed once at
|
||
512×256 (or whatever the working-grid resolution ends up being — see the
|
||
step-canvas discussion below for why 512×256 stays the right size even in the
|
||
stepped model), not re-solved per step-canvas or per zoom step. It is exactly
|
||
the kind of thing D-203's `BodyWorldState` LRU cache already exists to hold.
|
||
|
||
**Why `solve()` isn't internally parallelized, and why that's fine.** The
|
||
priority-flood fill and the Dijkstra overflow search are both globally sequential
|
||
by construction — one shared min-heap with strict pop order, architecturally
|
||
identical to `road_graph.rs`'s own unparallelized `astar`. A basin's outcome can
|
||
depend on a lower basin having already resolved (the chaining case in the
|
||
`DownstreamTarget` model), so cells are not independent pure functions the way
|
||
`layer_proxy.rs`'s row-chunked derive is (D-227's per-cell independence, which is
|
||
*exactly* what lets that path parallelize so cleanly — see Q3). Hydrology doesn't
|
||
have that property internally. It doesn't need to: the real workload is "273
|
||
independent bodies," and that parallelizes trivially and is the number that
|
||
matters.
|
||
|
||
**Cliff/multi-height representation (red flag 4).** Proposal, grounded directly
|
||
in what the solver already computes, not chosen independently of the
|
||
measurement: **dominant height (`elevation`, unchanged) + `channel_depth: u16`
|
||
(quantized) + `cliff_edge: bool`.**
|
||
|
||
- `elevation` keeps carrying the gridunit's rim/dominant/walkable height — zero
|
||
change to the existing single-height contract for the non-gorge case (which is
|
||
nearly every case — see below).
|
||
- `channel_depth` is a direct, non-lossy carry of `channel_depth_scaled` — the
|
||
solver already produces this number, no re-derivation downstream.
|
||
- `cliff_edge` marks a rim/discontinuity cell so the client map-art function
|
||
renders a cliff-face transition instead of a smooth gradient at that boundary.
|
||
|
||
I considered and rejected a min/max height pair: it loses the *shape* of the
|
||
transition (point-drop-at-one-edge vs. spans-the-whole-gridunit) that
|
||
depth+edge-flag preserves implicitly, and it's a worse fit to what the solver
|
||
actually outputs — it would require synthesizing two heights from one measured
|
||
depth value for no informational gain.
|
||
|
||
**The honest finding that should reframe red flag 4's urgency: gorge carving is
|
||
structurally rare.** All three production-scale benches (512×256, 768×432,
|
||
3.84M... 8.3M) carved **zero cells**. This isn't a bug — I traced it to a real
|
||
structural property: priority-flood finds the *true* global minimum rim of a
|
||
basin, so the first cell any overflow search reaches outside a sealed lake is, by
|
||
construction, never higher than that basin's own spill level. A single sealed
|
||
basin, however jagged its rim, always carves zero. Genuine carving needs a
|
||
narrower geometry — two independently-sealed basins connected by a corridor
|
||
exactly one cell wide in both dimensions and higher than both basins' own rims —
|
||
which is real terrain but rare at continental working-grid resolution. I
|
||
constructed four fixture attempts trying to force it and each one correctly
|
||
diagnosed as "basins merged" or "spill level absorbed the connector," not a code
|
||
defect (`server/src/atlas/hydrology_equilibrium.rs`, "Gorge carving" test
|
||
section, four named tests covering the mechanism directly since an end-to-end
|
||
production-scale trigger proved impractical to construct honestly).
|
||
|
||
**What this means for the wire-cost side of red flag 4:** the field is real and
|
||
needed for correctness, but it will be populated `false`/`0` for the overwhelming
|
||
majority of gridunits at every canvas size measured. The wire-cost argument
|
||
against including it is weaker than the red flag assumed — a rarely-nonzero
|
||
`bool` + quantized `u16` is cheap under any of the encodings in ④'s table (PNG's
|
||
DEFLATE in particular loves a field that's constant almost everywhere). I'd
|
||
rather we carry it at Phase-4 Atlas scope than defer cliffs to Phase-5 purely on
|
||
a cost argument that the measurement doesn't actually support — Tyre and
|
||
Araminta should weigh in on whether "rare but real, cheap to carry" changes the
|
||
Phase-4-vs-Phase-5 call red flag 4 posed; I'm only settling the representation
|
||
and the frequency, not the scope ruling.
|
||
|
||
## Question 2 — Step-canvas generation and canonical-vs-viewport (red flag 3)
|
||
|
||
**Recommendation: viewport-sized canvases at every step, no canonical fixed
|
||
canvas anywhere on the ladder. This isn't just the cheaper cache shape — at the
|
||
deep steps it's the only one that stays legal under D-226(d).**
|
||
|
||
The reasoning, stated plainly because red flag 3 called this "the sharpest thing
|
||
in the whole design" and I don't think that's overstated:
|
||
|
||
- A **canonical** fixed-size canvas at a given step's spacing would have to cover
|
||
enough of the body to serve every possible viewport at that step — at shallow
|
||
steps (region/district spacing) that's affordable (it's what the orbital tile
|
||
mosaic already does today, per-tile ImageTextures at capped density). At the
|
||
**deepest** step (10 px/tile, 1 m spacing), a canonical canvas covering enough
|
||
ground to be reusable across viewports is a near-whole-body metre-resolution
|
||
derivation — which is precisely what D-226(d)'s surviving whole-body-planetary-
|
||
layer prohibition forbids, independent of whether the cost is affordable.
|
||
- A **viewport-sized** canvas at 1 m spacing is exactly the 216×384 m (82,944
|
||
cell) shape I measured in T-1154: **17 ms parallel, comfortably interactive**,
|
||
and by construction never a whole-body layer — it's bounded by what's on
|
||
screen, the same "windowed viewport" carve-out D-226(d) already opened for the
|
||
district/quarter rungs (T-1124 amendment). This is the *only* canvas policy
|
||
that keeps the deep end of the ladder inside the boundary that survives this
|
||
workshop's governance delta.
|
||
|
||
**Practical shape at each step: viewport-sized, sized to a fixed pixel budget
|
||
(the same "land on a 4K window" instinct from Jeroen's outline), not to the
|
||
literal current window size.** Concretely: server renders each step's data
|
||
canvas at a fixed canvas-pixel budget (e.g. 3840×2160, or whatever the
|
||
round-2 wire-contract decision lands on) regardless of the client's actual
|
||
monitor resolution, the same way the current windowed-family already treats
|
||
`WIRE_CAP_CELLS` as a budget, not a literal viewport echo. This gives headroom
|
||
for larger monitors without re-deriving per-monitor, and it's the shape all
|
||
four of my measured cost numbers (330K/2.07M/8.3M) already assume — I didn't
|
||
measure a "canonical whole-body" shape at any step because I don't think one
|
||
should exist past the region/orbital rungs.
|
||
|
||
**Cost at each rung, all MEASURED not extrapolated (own numbers, T-1178/T-1154):**
|
||
|
||
| Step spacing | 330K canvas (parallel) | 2.07M canvas (parallel) | 8.3M canvas (parallel) | Realistic viewport canvas |
|
||
|---|---:|---:|---:|---|
|
||
| District (2,048 m) | 63.7 ms | 394.6 ms | 1,827 ms | — |
|
||
| Block (128 m) | same band, ~190-220 ns/cell | same band | same band | — |
|
||
| Tile (1 m) | same band | same band | same band | **82,944 cells (216×384 m) → 17 ms** |
|
||
|
||
The per-cell parallel rate (~190–220 ns/cell) is flat from 330K to 8.3M cells at
|
||
every spacing I tested (District through Tile) — this is the T-1143-killing gap
|
||
closed: there is no cliff where chunking overhead or cache pressure eats the
|
||
win at scale. **Cost does not gate any step in the ladder.** What gates the deep
|
||
steps is the wire carrier (④: even the best encoding is 21×–563× the existing
|
||
30 KB windowed-payload cap) and the canonical-vs-viewport governance boundary
|
||
above — neither is a cost question my measurements can resolve, both of which
|
||
I've now priced and named explicitly for Araminta/Tyre's round-2 synthesis.
|
||
|
||
**One caveat I want on record precisely because it's easy to gloss over:** the
|
||
8.3M-cell numbers above are for a *canvas that size*, not a *viewport that
|
||
size*. A 3840×2160 canvas at 1 m spacing covers only 3.84 km × 2.16 km of
|
||
ground — tiny. If "8.3M cells" is ever read as "a viewport-sized Tile-rung
|
||
canvas," that's wrong; the viewport-sized Tile canvas is the 83K-cell number
|
||
(216×384 m), not the 8.3M one. The 8.3M row exists to answer "does the derive
|
||
path degrade at that many cells" (no), not "is an 8.3M-cell canvas a realistic
|
||
Tile-step request" (it categorically is not, by the viewport-sizing policy
|
||
above).
|
||
|
||
## Question 3 — Compute-chunk partitioning across the Rayon queue
|
||
|
||
**The existing row-chunked `into_par_iter()` loop in `build_district_window_layer`
|
||
(`layer_proxy.rs:1564-1586`) is the mechanism, and it holds at every size and
|
||
every spacing on the ladder — this is the single most load-bearing "measured,
|
||
not extrapolated" result across all four gates.** Same per-cell rate (190–220
|
||
ns/cell parallel) confirmed at 330K, 2.07M, and 8.3M cells, at District, Block,
|
||
and Tile spacing, across three independently-run fixtures (my own synthetic
|
||
gradient body, the real GJ1c body, and Araminta's independent GJ338Bd run for
|
||
④) — three cross-validating measurements landing in the same band is a robust
|
||
finding, not a single-run artifact.
|
||
|
||
**Step boundaries as compute-chunk boundaries (premise 8) is compatible with
|
||
this mechanism as-is, with one caveat.** The row-chunking happens *within* one
|
||
step's derive call — it's how a single step canvas's cells get distributed
|
||
across the 16 Rayon workers, not how work is split *between* steps. Premise 8's
|
||
"step boundaries = compute-chunk boundaries" reads correctly as: each step is
|
||
one `build_district_window_layer`-shaped call (one derive request, internally
|
||
row-chunked), and steps are never partially computed or streamed mid-derive —
|
||
you get a whole step canvas or you're still waiting. That matches the
|
||
hold-fetch-swap model Stig's ⑤ measured render costs for.
|
||
|
||
**The caveat: courses are excluded from every chunked-cost number I measured
|
||
at 8.3M and at the deep-step 83K shape.** The row-chunked *replica* loop used
|
||
for every rectangular (16:9) canvas measurement is courses-empty by
|
||
construction (H2 in `t1178-t1154-derive-bench.md`) — it has no `RiverNetwork`
|
||
wiring at all. The real production path (square windows, real river network)
|
||
measured courses-inclusive at real density: **+0.09–0.21 ms against a ~5 ms
|
||
District-cap baseline (under 5%)**, cross-checked at 18 real courses in a
|
||
331,776-cell GJ1c window (195.0 ns/cell, within 2% of the courses-sparse
|
||
synthetic number). So the chunking mechanism's cost story holds either way —
|
||
courses add a small, bounded tax, not a different order of magnitude — but
|
||
anyone implementing step-canvas serving should wire courses into the row-chunk
|
||
loop from day one rather than treating them as a bolt-on, since the deep-step
|
||
83K number I reported (17 ms) technically excludes them and I don't want that
|
||
caveat to get lost between this document and an implementation ticket.
|
||
|
||
**One thing I did NOT measure and want flagged for round 2/implementation:**
|
||
whether chunking should change shape when a step canvas also needs to serve as
|
||
seed input for the next tier down (Jeroen's outline: "this at the same time
|
||
serves as seed information for the deeper cascade"). Today's row-chunked loop
|
||
produces one flat output array — using that array as *input* to a finer step's
|
||
derive is a data-flow question (does the finer step re-derive from the same
|
||
`(seed, position)` pure function, or does it consume the coarser step's output
|
||
values directly?), not a partitioning-cost question. My measurements assume
|
||
each step re-derives from `(seed, position)` independently per D-227's
|
||
derive-don't-store discipline (the coarser canvas is not an input the finer
|
||
canvas depends on for correctness, only for display continuity/pre-fetch
|
||
ordering) — if round 2 wants the finer tier to actually consume the coarser
|
||
tier's values as a literal input, that's a different architecture than what I
|
||
benchmarked and needs its own cost pass.
|
||
|
||
## Question 4 — Cache tiers server-side
|
||
|
||
**The global-tier cost, computed from measured numbers, not a vibe (red flag
|
||
2):** Jeroen's "always keep the global level" instinct is affordable **if
|
||
"global level" means the orbital/region-spacing rung, not a per-step cache of
|
||
every rung at every zoom the player has visited.** Using T-1179's measured
|
||
6.00 bytes/cell raw rate (or ~0.32× that with PNG-per-field, ~1.9 bytes/cell
|
||
effective):
|
||
|
||
- At district-spacing global canvas (~330K cells is the deep end of what
|
||
"orbital" would ever need — the real orbital rung is coarser, region-spaced,
|
||
which is fewer cells, not more): raw ≈ 1.99 MB/body, PNG-encoded ≈ 638 KB/body.
|
||
- **Across all ~273 bodies, raw: ~543 MB. PNG-encoded: ~174 MB.** Consistent
|
||
with red flag 2's own back-of-envelope (~440 MB at 5×5, ~13 GB at 1×1) — my
|
||
numbers land lower because I'm using the *measured* PNG-encoded rate rather
|
||
than the brief's pre-measurement estimate, and because "global" for the
|
||
orbital rung should be region-spaced (far fewer cells than the district-
|
||
spacing canvas I priced above as a conservative upper bound).
|
||
|
||
**Recommendation: "always keep the global level" = a resident, disk-backed
|
||
per-body cache at the orbital/region rung only, never a permanent allocation
|
||
for finer rungs.** ~174 MB (PNG-encoded, all 273 bodies) is a reasonable
|
||
resident/disk budget for "atlas navigation snappy after first calc" — this is
|
||
squarely a caching decision, not a live-memory one, and D-227's "transient
|
||
evictable cache" framing applies even to a tier that's practically
|
||
never-evicted in normal play: it must still be *derivable*, not
|
||
authoritative, so a cold-start or corrupted cache is a recompute, never data
|
||
loss. This is the same shape D-203's existing `BodyWorldState` LRU already
|
||
established server-side (50-body cap, ~5 MB budget, oldest-`last_accessed`
|
||
eviction, pinned current-location/neighbors) — I'd extend that same
|
||
resource/pattern to cover the orbital-rung step canvas rather than invent a
|
||
parallel cache mechanism. Concretely: `BodyWorldState` (or a sibling resource
|
||
following its exact shape) gains an `orbital_canvas: Option<EncodedCanvas>`
|
||
field, populated lazily on first Atlas-open for that body by the existing
|
||
Rayon background-queue population path (D-206), evicted only on the existing
|
||
LRU discipline (which for 273 bodies at ~638 KB PNG-encoded each barely
|
||
pressures a 50-entry cap sized for a much smaller ~100 KB/entry Layer-1
|
||
budget today — this may argue for either a higher cap or a split resource;
|
||
I'd defer that split to whoever implements it, since it's a tuning question
|
||
once the byte budget above is accepted, not an architecture one).
|
||
|
||
**Finer rungs (district/block/tile step canvases): TTL(detail, time,
|
||
distance), never "always keep."** This is D-227's concrete eviction policy
|
||
made specific:
|
||
|
||
- **detail** (rung depth) is the dominant term — deeper rungs get shorter TTL,
|
||
monotonically, because they're both cheaper to re-derive (my own measured
|
||
~190–220 ns/cell parallel holds flat across every rung, so re-derivation
|
||
cost does NOT argue for longer retention at deeper rungs the way it might in
|
||
a system where finer detail was expensive) and cover less ground (so a
|
||
given player session visits far more distinct deep-rung windows than
|
||
shallow ones, meaning the *hit rate* on a long TTL at a deep rung would be
|
||
low anyway — cache entries at Tile spacing are viewport-sized and
|
||
effectively single-use across a normal pan).
|
||
- **time** and **distance** (player/viewport focus, both character location
|
||
AND atlas-viewport-as-player-focus per premise 1) modulate the base TTL the
|
||
same way the existing FIFO-256 window cache + client 24-entry LRU already
|
||
behave, generalized: entries far from current focus (either kind) decay
|
||
faster; entries at the current focus are effectively pinned the way D-203
|
||
pins current-location/neighbor bodies.
|
||
- **Concretely, I'd propose the TTL formula as a multiplicative discount**:
|
||
`ttl = BASE_TTL[rung] × time_decay(age) × distance_decay(distance_from_focus)`,
|
||
with `BASE_TTL` the only per-rung tunable and the two decay functions shared
|
||
across rungs — this keeps the policy legible (one knob per rung, two shared
|
||
curves) rather than a per-rung bespoke formula, which matters for the same
|
||
reason the classification pipeline avoids per-rung decision-boundary tables
|
||
(T-1150 §6 precedent Araminta's ② question cites). I'm not committing to
|
||
exact `BASE_TTL` numbers here — that's a tuning pass once the wire contract
|
||
and step count are settled in round 2, not a round-1 architecture call.
|
||
|
||
**What I'm explicitly NOT proposing:** a disk-backed persistent tier for
|
||
anything finer than the orbital/global rung. Every finer-rung canvas is cheap
|
||
enough to re-derive (my own flat ~200 ns/cell parallel number holds at every
|
||
size and spacing I tested) that a disk tier for them would be optimizing a
|
||
cost that isn't the bottleneck — the bottleneck at finer rungs is wire size
|
||
(④) and step-cross latency (Stig's ⑤: sub-5ms upload, not a texture-upload
|
||
problem either), not derive cost. A disk tier for deep rungs would add
|
||
complexity (self-cleaning sweep, disk I/O latency on a cache miss) to solve a
|
||
problem the measurements say doesn't exist. This is my answer to Stig's
|
||
question 4's first candidate shape ((i) server-side SQLite cache DB) as it
|
||
applies to the *global* tier specifically: I think a simple in-process
|
||
resource (the D-203-shaped extension above) is sufficient for the orbital
|
||
rung's ~174 MB budget without reaching for SQLite at all — SQLite would only
|
||
earn its keep if the budget were large enough to need database-grade
|
||
eviction bookkeeping, and 273 rows at one field each doesn't need that. If
|
||
Stig's cross-boundary analysis (server-primary vs. client-primary reads) finds
|
||
a reason SQLite specifically is preferable for other reasons (client-facing
|
||
query shape, cross-session persistence semantics), I'd defer to that — my
|
||
claim is narrower: the *cost* doesn't force SQLite, a plain resource would
|
||
work fine at this size.
|
||
|
||
## Summary for round 2 / Jeroen's interview
|
||
|
||
1. **Hydrology is settled-viable, cheap, and mostly non-carving.** Ship the
|
||
T-1177 prototype's algorithm as-is; solve once per body (like
|
||
`drainage::analyze` today), hold in the D-203-shaped cache, never re-solve
|
||
per step or per request.
|
||
2. **Cliff representation is `channel_depth` + `cliff_edge` alongside
|
||
unchanged `elevation`** — cheap to carry given how rare carving measures
|
||
as being; Phase-4-vs-Phase-5 scope call is Tyre/Araminta's, not mine to
|
||
settle alone.
|
||
3. **Viewport-sized canvases at every step, always — no canonical fixed
|
||
canvas past the orbital/region rung.** This is the only policy that keeps
|
||
the deep ladder inside D-226(d); cost was never the constraint here,
|
||
governance is.
|
||
4. **The row-chunked par_iter mechanism already in `layer_proxy.rs` is the
|
||
answer to compute-chunk partitioning** — it holds flat from 330K to 8.3M
|
||
cells at every spacing, at three independent cross-validations. Wire
|
||
courses into it from the start; don't let the courses-empty replica-loop
|
||
measurement become the shipped assumption.
|
||
5. **Cache tiers: one resident/disk global tier at the orbital rung
|
||
(~174 MB PNG-encoded across all bodies, D-203-shaped resource extension),
|
||
TTL(detail, time, distance) for everything finer, no disk tier below
|
||
orbital.** The wire carrier (④) and the canonical-vs-viewport boundary,
|
||
not derive cost, are what gate the deep steps — cost is a solved problem
|
||
at every rung I measured.
|
||
|
||
None of the four things I was gated on turned into a "computer catches fire"
|
||
result. The genuinely open items for round 2 are Araminta's wire contract
|
||
(the tagged-envelope call is now unavoidable per ④, I agree with that
|
||
reading), Tyre's gridunit↔D-243 snap-or-float ruling, and the Phase-4-vs-
|
||
Phase-5 cliff scope call — none of which are cost questions my measurements
|
||
can resolve alone.
|