Files
settled-reach/docs/workshops/body-map-viewer/dudley-round2.md
T
jpmschweitzerandClaude Fable 5 29c22cb728 docs(meta): body-map-viewer workshop — rounds, measurements, outcomes, as-built briefing
The complete workshop record: four round-1 positions, five round-2 syntheses
(incl. Troblum's adversarial pass with addendum + final scorecard — all seven
findings resolved), both lead interviews, Qatux's round notes and the 8-section
workshop-outcomes.md (the lakes message-crossing documented as process
history), measurement ⑥ (set_pixel/c1) + the population-survey and chunk/S2
addenda in the measurement docs, the brief's appendix updated through ⑥, and
architecture-briefing-final.md — Jeroen's outline written back as-built
(six-level ladder, lakes, ~9MB resident global tier). README row: Complete.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-24 10:57:45 +02:00

769 lines
47 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
---
title: "Body Map Viewer — Dudley Round 2"
description: "Envelope mechanics (wire/serving design), seed-chaining ruling, step-ladder tables, and final cache-tier spec incorporating Jeroen's storage-eviction amendment"
workshop: body-map-viewer
round: 2
owner: Dudley
status: complete
decision_refs: [D-166, D-225, D-226, D-227, D-243, D-203]
---
# Dudley — Round 2 Position
Read `lead-interview-1.md` and all three other round-1 documents first. Four
items in the coordinator's split: envelope mechanics (wire/serving half —
Tyre owns the D-record text), the seed-chaining ruling (closing round-1 open
item 8), step-ladder tables (with Tyre — producing the tables Jeroen rules
from at interview 2), and the final cache-tier spec incorporating Jeroen's
storage-eviction amendment.
**Coordination note:** I first attempted to reach Tyre directly (SendMessage)
to split envelope-mechanics drafting live and he wasn't addressable yet, so I
designed the wire/serving half unilaterally, grounded directly in the actual
`decode_inbound`/`Inbound` code (`server/src/bridge/mod.rs`). Tyre reached me
mid-draft with his own three concrete questions (message shape, legacy
coexistence, step-0 carrier) plus his snap-rule constraints and three
candidate step-ladder skeletons (Options A/B/C, all D-243-disciplined,
correcting an error in my own first pass — see §(c)). This document is
revised to answer his questions directly and replaces my original two tables
with cost fills against his skeletons. Araminta also relayed her dense/sparse
envelope-framing question (via the coordinator) — answered at the end of
§(a). All three of us are converging live; this is the reconciled version.
---
## (a) Envelope mechanics — wire/serving design
### The constraint this has to satisfy, stated precisely
I read the actual code before designing this, not just the doc summaries.
`server/src/bridge/mod.rs`'s `decode_inbound` demuxes five inbound map/array
shapes today (`Inputs`, `AtlasRequest`, `StarMapRequest`, `CityNamesRequest`,
`BrowseRequest`), and the module's own doc comment (lines 67-75) already
names this exact moment:
> "The next genuinely NEW inbound shape (a sixth) must migrate the channel to
> the tagged-envelope framing D-225 deferred — do not add a sixth probe."
A step-canvas request is unambiguously a sixth new shape. So this isn't a
judgment call about *whether* to migrate — the ceiling was pre-declared and
this is the shape that trips it. What I'm designing is the concrete *how*.
### What "tagged envelope" already means in this codebase — don't invent a new mechanism
Reading `Inbound`'s existing three newer variants (`StarMapRequest`,
`CityNamesRequest`, `BrowseRequest`), the "tag" D-225's 2026-06-12 amendment
asked for is **already partially implemented**, just not generalized: each of
those three structs carries a mandatory boolean discriminator field the
others don't have at all (`star_map: bool`, `city_names: bool`, `browse:
bool`), and `decode_inbound`'s `ShapeProbe` defensively checks that at most
one discriminator is present in any given frame, rejecting ambiguous unions
outright (PR #176 review H1). This is a **tagged envelope in substance** — a
required marker field distinguishing shapes — it just hasn't needed a shared
name or a unified probe struct because five shapes fit in one hand-rolled
`ShapeProbe`.
**My recommendation: extend this exact pattern for the sixth shape, don't
invent a parallel envelope format.**
```rust
/// A step-canvas data-canvas request (body-map-viewer workshop, T-1176
/// follow-on). Discriminator field `step_canvas: true` — same tagged-marker
/// pattern as `star_map`/`city_names`/`browse`.
#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct StepCanvasRequest {
pub step_canvas: bool, // mandatory discriminator, always true when present
pub body_id: String,
pub step_index: u32, // the discrete step (§(c) below), never a raw spacing float
pub center: (i64, i64), // world-metres, snapped to the step's D-243 rung grid (Tyre's snap ruling)
pub extent: (u32, u32), // canvas px budget, e.g. 3840x2160 — fixed budget, not literal viewport echo
pub min_wl_m: u32, // octave cutoff, quantized band (existing window_min_wl_m precedent)
}
```
Add `StepCanvasRequest(StepCanvasRequest)` as a sixth `Inbound` variant,
extend `ShapeProbe` with a `step_canvas: Option<IgnoredAny>` field, and add it
to the mutual-exclusivity sum and the try-order chain in `decode_inbound`.
This is a small, mechanical, additive change — the demux *mechanism* doesn't
need a redesign, because the five-shape ceiling was never about the
mechanism breaking, it was a **discipline marker** ("stop adding shapes this
way past five, go do the real thing"). The "real thing" it points at, once I
traced the actual code, turns out to be exactly the pattern already in place
— which is good news, not a corner cut: it means the migration is
low-risk, not a rewrite.
**Response side is the actual new shape, and it does need its own message —
not a field.** `AtlasLayerResponse`'s own doc (`layer_proxy.rs:605-631`) is
explicit that a windowed-family second field is "a dedicated response
message by rule, not a second `Option`" (D-226 T-1124 §2). A step-canvas
response is emphatically that case — Tyre's amendment text already rules
this (§1d of his round-1 doc). Concretely:
```rust
/// Response to a StepCanvasRequest. Deliberately NOT a field on
/// AtlasLayerResponse (D-226 T-1124 §2's windowed-family ceiling names this
/// exact case) — a wholly separate response type outside that family.
#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct StepCanvasResponse {
pub body_id: String,
pub step_index: u32,
pub center: (i64, i64), // echoed, same staleness-guard pattern as district_window
pub status: StepCanvasStatus, // Ready | Pending | Error(String) — mirrors the existing three-way status enums
pub canvas: Option<EncodedStepCanvas>,
}
/// One dense field, PNG-per-field encoded (T-1179's measured winner).
pub struct EncodedField {
pub png_bytes: Vec<u8>,
}
pub struct EncodedStepCanvas {
pub width: u32,
pub height: u32,
pub morphology: EncodedField,
pub elev_q: EncodedField,
pub temp_dc: EncodedField,
pub moisture_q: EncodedField,
pub vegetation: EncodedField,
pub glaciation: EncodedField,
pub settlement_id: EncodedField, // Araminta's new dense field
pub courses: Vec<RiverCourse>, // sparse, MessagePack-native (unchanged shape)
pub cliffs: Vec<CliffSegment>, // sparse, MessagePack-native — see §(a) note on the Dudley/Araminta tension below
}
```
`SimBridge` gains `fn send_step_canvas_response(&self, resp:
&StepCanvasResponse) -> Result<(), BridgeError>`, parallel to the existing
five `send_*` methods — same pattern `TcpBridge` already implements for
`send_atlas_response`/`send_star_map_response`/etc. This is the entire
"envelope migration" in mechanical terms: **one more tagged inbound variant,
one more dedicated outbound method, both following patterns already proven
five times over in this exact file.** I want this stated plainly because
Tyre's round-1 text (correctly) frames the migration as "expected scope,
challenging but doable" at the governance level — at the *code* level, once
you actually read `bridge/mod.rs`, it's closer to "doable" than
"challenging." The challenge is elsewhere (the client-side rebuild Stig owns,
the cache-tier and step-ladder work below), not in extending this demux.
**Note on the cliff-field dense-vs-sparse tension (round-1 conflict #1,
`round-1-notes.md` §3):** I'm adopting Araminta's sparse `Vec<CliffSegment>`
shape in the struct above, not my round-1 text's implicit dense-array
framing. Reading her argument again against my own hydrology numbers, she's
right and I was underspecified, not actually disagreeing: I only ever
proposed the *field set* (`channel_depth` + `cliff_edge`), never committed to
dense-array-shape explicitly, and given the measured rarity (zero carved
cells at every production-scale bench I ran), a mostly-empty `Vec` costs
strictly less than a mostly-zero dense array under every encoding in T-1179's
table — sparse wins outright here, no tradeoff to weigh. Tyre's "new arrays"
language in his synthesis was, I think, imprecise rather than a third
position — resolving as: **`cliffs: Vec<CliffSegment>` parallel to
`courses`, both sparse, both MessagePack-native, neither PNG-encoded.**
### Legacy `district_window` / `AtlasLayerRequest` coexistence during migration
**No retirement message, no deprecation flag, no dual-write.** The legacy
windowed carrier keeps working completely unchanged — same struct, same
`serve_district_window`, same cache key, same wire bytes — for as long as the
*current* Atlas viewer (`AtlasWindowViewer`, the thing Stig's round-1 doc
describes retiring client-side) keeps requesting it. The server has no
concept of "the old viewer" — it just answers whichever request shape
arrives. Once Stig's stepped viewer ships and the old `AtlasWindowViewer`
code is deleted client-side, the server simply stops receiving
`AtlasLayerRequest.window_center`-populated requests, and `district_window`
quietly goes cold — no server-side cleanup ticket needed, per D-005/D-192's
co-ship guarantee (client and server ship together, so there's never a
window where an old client talks to a new server or vice versa). This is the
same "eviction → recompute, always valid" property D-227 already gives every
cache entry; an unrequested carrier isn't a liability, it's just unused code
someone can delete in a follow-up cleanup ticket once the cutover is
confirmed complete, not a coexistence *problem* requiring design work now.
**One thing worth flagging for the ticket plan (not a design gap, a
sequencing note):** the migration should ship the new `StepCanvasRequest`
path fully working, verified via the live Gauntlet pipeline (project
convention), *before* the client cuts over — i.e. the server-side sixth-shape
addition and the client-side viewer rewrite don't have to land in the same
PR, since the old carrier keeps serving the old client throughout. This
lowers implementation risk (the migration doesn't need a flag-day cutover)
and is worth Tyre/the lead naming explicitly in the ticket plan.
### Direct answers to Tyre's three questions
**1. Message shape: discriminator field on the existing stream, not a new
wrapper/framing layer.** Concretely, `step_canvas: bool` on a new
`StepCanvasRequest` struct, extending the exact `ShapeProbe` pattern that
already governs `star_map`/`city_names`/`browse` (`server/src/bridge/mod.rs`,
read directly — see above). This *is* "a required marker field," which is
literally the alternative-satisfying language D-225's 2026-06-12 amendment
used ("a tagged envelope (or a required marker field)") — I'd steer away from
inventing a wrapper/framing layer on top of the existing stream, because the
stream itself doesn't need restructuring, only one more disambiguated shape.
Same answer for the response side: a new `StepCanvasResponse` type +
`SimBridge::send_step_canvas_response`, not a wrapped/multiplexed response
envelope. My reasoning for this over a more general `envelope_version`/
`msg_type` integer tag: zero new demux machinery, proven five times already,
and I don't have a seventh shape waiting that would justify generalizing the
mechanism now — if you want to formally name a reusable envelope *type* in
the record text for whatever comes after this one, that's a framing choice
the record can state independent of what actually ships in code (the code
only needs one more discriminated variant either way).
**2. Legacy coexistence: `district_window` survives unchanged, and I'd let
the envelope subsume only the NEW step-canvas traffic, not migrate existing
windowed traffic onto it.** Reasoning: `district_window`'s current consumer
(the existing `AtlasWindowViewer`) isn't being redesigned by this workshop —
it's being *replaced* by Stig's new stepped viewer. There's no reason to
migrate the old carrier's wire shape when the old carrier's only client is
also being deleted; migrating it would be extra engineering (touch the
legacy struct, the legacy cache key, the legacy client decode path) for a
code path with a defined end-of-life, not an extended future. Your lean
matches mine exactly — keep `district_window` alive, byte-for-byte, for
whatever old-viewer traffic exists until the client cutover, then it goes
cold and gets deleted in a follow-up cleanup ticket. One carrier "cleaner in
the abstract" isn't worth the extra migration surface for a carrier that's
already scheduled for deletion.
**3. Step-0/Region rides the NEW envelope, not the legacy whole-body
path.** Now that you and Araminta have confirmed Region = step-0, it's
generated by the *same* mechanism every other step uses (a `StepCanvasRequest`
at `step_index: 0`, D-243 rung = Region, viewport-sized per the capped-tile
mosaic shape T-1143 §4 already established) — not a special case riding
`district_grid`/`region_grid`'s existing whole-body `Option` fields on
`AtlasLayerResponse`. Putting step-0 on the legacy whole-body family would
mean the new stepped client has to speak two different response protocols
depending on which step it's on (the new envelope for steps 1-5, the old
family for step 0) for no architectural gain — one request shape, one
response shape, across the whole ladder is simpler for both sides and is
what "step boundaries = compute-chunk boundaries" (premise 8) already implies
uniformly. The existing `region_grid`/whole-body layers keep serving whatever
non-stepped consumers still want them (if any survive the cutover) — they're
just not what step-0 of the new ladder uses.
### Araminta's dense/sparse framing question — answered
Relayed via the coordinator: does the tagged envelope frame the step-canvas
response as (a) one flat tagged message with N typed fields (her default,
today's `DistrictWindowLayer` shape scaled up), or (b) dense/sparse split
into separately-tagged sub-messages enabling partial/progressive delivery?
**Answer: (a), one flat message — I don't think (b) is forced by anything in
my design, and I'd argue against it for this pass specifically.** My
`StepCanvasResponse` struct above is exactly her default assumption: one
`Ready | Pending | Error` response carrying all eight fields (six dense +
`courses` + `cliffs`) together, matching the whole-payload-together
precedent D-225 already established for `Layer1Output` ("the layers are
produced together in one drainage pass... per-layer requests save no compute
and only add round-trips") and the countervailing case (b) would trade
against isn't free:
- **Cost argument against (b):** every field in the response comes off the
*same* row-chunked derive pass (`build_district_window_layer`'s loop
computes all six dense fields per cell in one traversal, per my own
measured numbers) — splitting delivery into sub-messages doesn't save any
server-side compute, it only adds round-trips and partial-state handling
on both ends for a payload that's already fully computed by the time the
*first* sub-message could go out. This is the identical argument D-225 used
against per-layer whole-body requests, and it applies with the same force
here.
- **The progressive-paint UX case is real but belongs one layer up, not
in the wire protocol.** If "classification planes first, courses/cliffs
after" is worth doing for perceived latency, it's cheap to get without
splitting the *server response* — the client can paint the terrain RTT
layer as soon as the (whole) response decodes and defer building the
annotation layer's draw calls by a frame or two, since courses/cliffs are
a tiny fraction of total decode time (sparse lists, a few KB, against a
multi-hundred-KB-to-tens-of-MB dense payload per Table B below). That's a
client-side rendering-order choice Stig can make freely without any
server-side protocol complexity — no reason to pay demux/partial-delivery
cost on the wire for a UX win available for free on the receiving end.
- **Where (b) would earn its keep and doesn't yet:** if a future
measurement showed the dense-field derive pass and the sparse-list
derive pass (courses/cliffs) were on meaningfully different cost/latency
timelines — e.g. if course invention turned out to be slow enough to want
to ship the terrain first and stream courses in later — that would be a
real argument for (b). Nothing measured this workshop shows that: courses
add "+0.09–0.21 ms against a ~5 ms baseline" per my round-1 citation of
`bench_course_cost_on_vs_off` — negligible, not staggered. I'd revisit
this if a future profiling pass finds a real timeline split, but I'm not
designing speculative complexity against a cost gap that isn't there.
**Net: her default holds. One flat tagged `StepCanvasResponse`, all fields
together, matching D-225's whole-payload precedent — she can write the
client-facing field spec against shape (a) without qualification from my
side.**
---
## (b) Seed-chaining ruling — independent re-derivation
**Ruling: each step re-derives independently from `(seed, position)`. A
step's output is never consumed as literal input by a finer step's
derivation.** This closes round-1 open item 8, and it's not a new
architectural call — it's D-227 applied to a question I should have answered
more explicitly in round 1 instead of flagging as unmeasured.
### The argument
**D-227 purity is the whole case.** `subtile(x,y,z) = derive(seed, atlas,
position)` — every derived value is a pure function of the fixed inputs, not
of some other derived value's cached output. If a finer step consumed a
coarser step's *computed* array as an input (rather than re-deriving from the
same `(seed, position)` the coarser step also started from), the finer step's
correctness would depend on the coarser step having been computed first, with
a specific set of parameters, and cached — which means:
1. **A cache-dependency chain**, exactly what D-227's "evictable, cache is a
bonus, never truth" discipline forbids. If the coarser step's cache entry
is evicted (my own TTL policy in §(d) below deliberately evicts sub-global
geometry), a finer step that depends on it as literal input either fails
to derive or must first regenerate the coarser step — silently
reintroducing exactly the "step boundaries as compute-chunk boundaries"
complexity premise 8 was trying to keep simple, plus a correctness risk
if the regenerated coarser step doesn't byte-match what was evicted (it
would, since both are pure functions of the same seed — but now *proving*
that becomes load-bearing instead of automatic).
2. **A determinism ordering hazard.** Two clients requesting the same finer
step from a cold cache, in different orders relative to any coarser-step
requests, would need to agree on exactly which coarser-step values fed
the derivation — this is solvable, but it's solving a problem that
doesn't exist if every step independently calls `derive_at_metres(seed,
position, this_step's_own_min_wl_m)`.
3. **It contradicts my own measured cost model.** All of T-1178/T-1154's
numbers (the flat ~190–220 ns/cell parallel rate at every rung) are for
independent per-cell derivation calls — `derive_window_cell` never reads a
neighboring cell's *derived* value, let alone a different rung's. If a
finer step needed the coarser step's actual array as input, that's a
different, unmeasured computational shape (a dependency graph between
rasters, not a row-chunked independent-cell map) — my GO verdicts at
every rung do not cover that shape.
**What Jeroen's outline phrase actually means, read against the code that
exists today.** "This at the same time serves as seed information for the
deeper cascade" — I read this as describing the existing, already-shipped
pattern in `district_profile.rs`/`detail_scatter.rs`: a district's
`invent_primitives` call reads the *region baseline* (a coarser, independent
derivation) as one of several deterministic inputs alongside the fixed seed
and position, the same way `voxel_relief` already contributes to
`elev_q`/`slope_q` at Block spacing today (confirmed in my own T-1154
measurement doc). That's not "consume the coarser step's computed array
element-for-element as this step's starting point" — it's "the coarser
rung's own *independently-derivable* baseline value at this position is one
of several `(seed, position, ...)`-keyed inputs to the finer rung's own
independent derivation call." Both are pure functions; the finer one just
happens to call the coarser one's derivation function internally as a
sub-computation, not read its *cached* output. This distinction is exactly
what keeps D-227 intact: `derive_district(seed, pos)` calling
`region_baseline_at_district(seed, pos)` internally is still one pure
function of `(seed, pos)` — no cache dependency, no ordering hazard, because
the "coarser step" being referenced is re-executed as code, not looked up as
data.
**Concretely, the ruling is: "consuming the coarser output" already happens,
and it happens the *only* way D-227 permits — as a nested pure-function
call, evaluated fresh every time, never as a read from the coarser step's
cache entry.** This is not new work; the region-baseline-feeds-district
pattern is the precedent, and step-canvas generation should follow it
exactly: a Tile-step's derive call may internally call the same
`region_baseline_at_metres`/`derive_at_metres`-family functions a Region-step
call would, at the same `(seed, position)`, but it never reads
`StepCanvasResponse` bytes from a prior request as an input.
### What this settles and what it doesn't
- **Settles:** my own cost numbers (measurement ②/③) are validated as the
correct model for step-canvas generation — no re-measurement needed,
because "independent re-derivation, coarser rung called as a nested
function" is exactly the shape those benchmarks already exercise
(`derive_at_metres` calling into the same octave-primitive functions at
every rung).
- **Doesn't settle:** whether every finer step's derivation *should* call
every coarser rung's function as a sub-computation (a "does Tile-rung
derivation call Block-rung, District-rung, AND Region-rung functions in
sequence, or just the immediately-coarser one?" question) — that's an
algorithm-composition detail for whoever implements `derive_at_metres`'s
Tile/Block extension, not an architecture question this ruling needs to
answer. The architecture-level answer is just: however many coarser
functions get called, they're called fresh, never read from a response
cache.
---
## (c) Step-ladder tables — with Tyre
**Correction to my own first pass, before the tables:** my original draft of
this section used a "Tile-adjacent (4 m)" spacing that is **not a D-243
rung** — I conflated a T-1154 benchmark data point (which measured 4 m only
as a probe value, not a named ladder level) with a real step. Tyre's message
caught this by construction — his three skeletons are built strictly from
D-243's actual six rungs (voxel 1 m, chunk 64 m, block 128 m, quarter 512 m,
district 2,048 m, region 204.8 km) and nothing else, which is the right
discipline and the one I should have applied the first time. The tables
below use only real D-243 rungs, per his skeletons, with my cost fills.
**Measurement ⑥ status:** not received — checked
`docs/workshops/body-map-viewer/measurements/`, still only ①–⑤ present as of
this document. Cost fills below are ①–④ only; ⑥ affects Stig's shader-vs-CPU
call, not these tables.
**Canvas-extent convention for every row below:** the fixed 3840×2160 px
budget you proposed, viewport-sized (never canonical past step 0) — using
your suggested convention directly, not a different one.
### Cost fills against your three skeletons
**Option A — one-rung-per-step (6 steps): Region → District → Quarter →
Block → chunk (64 m) → voxel (1 m).**
| Step | D-243 rung | Spacing | Step factor (from prev) | Viewport extent @ 3840×2160 canvas | Cells | Derive cost (parallel) | Wire cost (PNG-per-field) |
|---|---|---:|---:|---|---:|---:|---:|
| 0 | Region | 204.8 km | — | whole body (capped-tile mosaic, T-1143 §4) | 8,294,400 | 1,827 ms (measured) | 16.88 MB (measured) |
| 1 | District | 2,048 m | 100× | 7,864 × 4,424 km | 8,294,400 | 1,827 ms (measured) | 16.88 MB (measured) |
| 2 | Quarter | 512 m | 4× | 1,966 × 1,106 km | 8,294,400 | ~1,827 ms (same band, T-1154 confirms Quarter costs the same per-cell rate as District) | ~16.88 MB |
| 3 | Block | 128 m | 4× | 491 × 276 km | 8,294,400 | ~1,827 ms (measured directly at Block spacing, T-1154: same ~190–220 ns/cell parallel band) | ~16.88 MB |
| 4 | chunk | 64 m | 2× | 246 × 138 km | 8,294,400 | **UNMEASURED — flagging honestly, not filling with a guess.** T-1154 tested Block (128 m), Tile-adjacent (4 m — an ad hoc probe, not this rung), and Tile (1 m). 64 m spacing was never benched directly. The cutoff-mechanism finding (T-1154: `VOXEL_OCTAVE_WAVELENGTHS_M` bottoms at 128 m, so nothing truncates below Block) strongly implies chunk costs the *same* flat ~1.8 µs/cell single-thread / ~200 ns/cell parallel rate every other rung in this band does — but "strongly implies" is not "measured," and I'm not reporting a number I didn't run. | Same caveat — implied ~16.88 MB by the flat-rate pattern, not measured. |
| 5 (deepest) | voxel | 1 m | 64× | **8.3M cells at 1 m spacing = 3.84 × 2.16 km — this is the wrong deep-step shape**, see the general note below | 8,294,400 | same caveat as row 4 pattern, ~1,827 ms if it held | ~16.88 MB |
**Option A's real problem isn't the missing chunk measurement — it's that a
fixed 3840×2160 canvas at voxel (1 m) spacing is oversized for what a real
viewport needs at 10 px/tile** (3840 px ÷ 10 px/tile = 384 m — my own T-1154
"realistic deep-step canvas" used exactly this reasoning to arrive at
216×384 m, 82,944 cells, not an 8.3M-cell canvas). Applying your fixed-canvas
convention literally at every step breaks down at the deepest step
specifically because the fixed-px-budget convention and the display-ratio
target (10 px/tile) can't both hold at 1 m spacing simultaneously — this
isn't a flaw in your skeleton, it's a flaw in "same canvas budget at every
step" as a blanket rule, which is a genuinely useful thing for these tables
to expose. I'd apply the fixed-3840×2160 convention at every step **except**
the deepest, where the canvas must instead be sized to the display-ratio
contract (my round-1 T-1154 number: 216×384 m, 17 ms, ~486 KB) — noted
consistently in all three tables below, not just this one, since it's a
convention-level correction, not a skeleton-specific one.
**Option B — skip-chunk (5 steps): Region → District → Quarter → Block →
voxel (1 m).**
| Step | D-243 rung | Spacing | Step factor | Viewport extent @ 3840×2160 (except deepest) | Cells | Derive cost (parallel) | Wire cost |
|---|---|---:|---:|---|---:|---:|---:|
| 0 | Region | 204.8 km | — | whole body | 8,294,400 | 1,827 ms (measured) | 16.88 MB (measured) |
| 1 | District | 2,048 m | 100× | 7,864 × 4,424 km | 8,294,400 | 1,827 ms (measured) | 16.88 MB (measured) |
| 2 | Quarter | 512 m | 4× | 1,966 × 1,106 km | 8,294,400 | ~1,827 ms (T-1154 same-band confirmation) | ~16.88 MB |
| 3 | Block | 128 m | 4× | 491 × 276 km | 8,294,400 | ~1,827 ms (measured directly, T-1154) | ~16.88 MB |
| 4 (deepest) | voxel | 1 m | **128×** | 216 × 384 m (display-ratio-sized, per the correction above, NOT the fixed-canvas convention) | 82,944 | **17 ms (measured directly, T-1154)** | **~486 KB (measured rate, T-1179)** |
**This is the skeleton I'd recommend, and the missing-chunk-measurement
problem is exactly why: every single row in Option B is either directly
measured or in the same measured cost band as a directly-measured rung —
there is no row I have to caveat as unmeasured.** That's not a coincidence
of which skeleton I like; it's a direct consequence of chunk (64 m) never
having been benched, and Option B is the one skeleton among your three that
doesn't need it. If chunk's cost does turn out to match the flat-rate pattern
(likely, per the cutoff-mechanism reasoning above), Option A costs the same
as Option B row-for-row at every *other* rung anyway — the only real
difference is Option A has six fetches per full zoom-out-to-in traversal
where Option B has five, and pays one unmeasured/probably-redundant rung to
get there, since chunk sits between Block and voxel with a huge factor
either way (128× Block→voxel in Option A's own 64m→1m final leg, vs 128×
Block→voxel directly in Option B — the chunk step doesn't actually *reduce*
the biggest jump in the ladder, D-243 names chunk as "stream/derive unit,"
not a natural *display* rung, which is exactly Tyre's own hesitation about
it in his message).
**Option C — coarse-doubled (4 steps): Region → Quarter → Block → voxel,
skipping District.**
| Step | D-243 rung | Spacing | Step factor | Viewport extent | Cells | Derive cost | Wire cost |
|---|---|---:|---:|---|---:|---:|---:|
| 0 | Region | 204.8 km | — | whole body | 8,294,400 | 1,827 ms (measured) | 16.88 MB (measured) |
| 1 | Quarter | 512 m | **400×** | 1,966 × 1,106 km | 8,294,400 | ~1,827 ms (same band) | ~16.88 MB |
| 2 | Block | 128 m | 4× | 491 × 276 km | 8,294,400 | ~1,827 ms (measured) | ~16.88 MB |
| 3 (deepest) | voxel | 1 m | 128× | 216 × 384 m (display-ratio-sized) | 82,944 | 17 ms (measured) | ~486 KB (measured) |
**Cost-wise, Option C is not cheaper than B — the per-step cost numbers are
identical to B's (same rungs at the same measured rate), it just skips
District.** So the choice between B and C is **not a cost question at
all** — every row I can cost is the same regardless of which skeleton wins.
It's purely the UX-pacing question your message already correctly separated
out (my answer to your question 2 below).
### Answering your three direct questions
**1. Which skeleton do the costs favor?** None of them are uncomfortable —
every derive cost across all three options is comfortably interactive (17 ms
to 1.8 s, the same band my round-1 doc already established as safe against
any step-cross tolerance). **The costs don't pick a winner; they only
disqualify chunk (64 m) from being *load-bearing* for the decision**, because
it's the one rung nobody has actually measured. That makes Option B (which
never needs chunk) the *cleanest* recommendation on evidentiary grounds, not
because A or C are unaffordable — I'd frame this to Jeroen as "B is the
option with zero asterisks," not "A and C are too expensive."
**2. Does the D-243 factor unevenness (100×, 4×, 4×, 128×) matter for
serving/precache, or is it purely UX pacing?** Both, but asymmetrically —
**it matters for serving in exactly one place (the Region→District 100× jump
at step 0→1), and is purely UX pacing everywhere else.** Reasoning: my cache
tiers (§(d) below) already treat Region as a categorically different tier
(keep-always, global) from every step below it (storage-evictable) — the
100× factor at the top of the ladder lines up with a real architectural
seam that already exists in the cache design, not just a display jump. The
4×/4×/128× factors *within* the sub-global tier don't correspond to any
serving-side seam — District, Quarter, and Block are all "sub-global
geometry," evicted by the same storage-TTL mechanism, served by the same
row-chunked derive path, at the same flat per-cell cost. So: the top jump is
architecturally real (it's the global/sub-global cache boundary already
established for other reasons); every jump below it is exactly what you
said — a UX-pacing question about how gradual the zoom feels, with no
serving-side consequence I can find. This is useful to tell Jeroen directly:
he can pick B vs C on feel alone for the lower rungs without worrying he's
picking a cost or caching regression.
**3. Viewport metre-extent assumption per step?** Confirmed: your fixed
3840×2160 canvas budget, with the one correction above (the deepest step
must use the display-ratio-sized canvas — 216×384 m at 10 px/tile — not the
fixed-px-budget convention, because the two conventions are mutually
incompatible exactly at voxel spacing, per Option A's row 5 above). I'd
apply "fixed 3840×2160 px budget" as the rule for every step except the
last, and "sized to the display-ratio contract" as the rule for the last
step specifically — both tables B and C already reflect this split.
### Confirmation from Stig's measurement ⑥ (relayed, landed after the tables above were drafted)
Stig's `Image.set_pixel` colorize cost measurement (77.5 ns/cell flat across
330K–8.3M cells, 25.7 ms at 330K, beating `PackedByteArray`-direct ~2×) closes
his round-1 open item and independently confirms c1 as CPU-first — not
something my tables need to change for, but his own px-band recommendation
(1×1 at Block/Tile "full fidelity," 1×1 preferred at Quarter/District as long
as the realistic canvas stays under ~2M cells, ~5×5 fallback reserved for
shallow/orbital steps where canvas *extent* is what grows, not density) is
worth stating explicitly against my tables because **it's the same policy my
cost fills already assume, from a different, client-side cost driver.**
Every row in Options A/B/C above already uses 1 gridunit-per-pixel (cell
count = canvas px count) at every step from Region through Block — I never
invoked the ~5×5 fallback anywhere in my server-side cost fills, because
nothing in my derive-cost numbers forced it (the flat ~190–220 ns/cell
parallel rate holds at every canvas size I measured, so there was no
server-side reason to sample coarser-than-1:1 at any step). Stig's finding
adds the client-side half of that same argument: colorize cost is
**cell-count-driven, not display-density-driven**, so downsampling to 5×5 at
a step that doesn't need it wouldn't even save client-side coloring time
proportionally — the cost is paid per gridunit regardless of how many screen
pixels each one covers. **Two independent cost models (server derive, client
colorize) land on the same policy: 1×1 wherever the resulting cell count is
affordable (every step through Block, per both his ~2M-cell comfort line and
my own flat-rate measurements), reserving the coarser ratio only for the one
place canvas *extent* — not density — is actually the pressure (the
whole-body Region/orbital step, T-1143 §4's capped-tile mosaic).** This
confirms, not revises, every table above; I'm noting it because two
independently-measured cost models agreeing on a policy neither one was
designed to argue for is exactly the kind of convergence worth flagging to
Jeroen rather than leaving as a coincidence buried in two separate documents.
---
## (d) Cache tiers — final spec incorporating Jeroen's storage-eviction amendment
Jeroen's amendment (verbatim, `lead-interview-1.md` ruling 2): *"we still may
also want to evict non global level geometry based on time to save storage
for planets the player visits but never goes back to."* Captured there as:
**staleness-eviction and storage-eviction are distinct axes** — geometry
never goes stale (re-derivable, byte-identical forever), but sub-global
geometry still gets evicted on time-since-last-visit as a storage-budget
policy, independent of whether it's "correct." The global tier alone is
keep-always.
This sharpens, rather than replaces, my round-1 TTL proposal — I had
conflated "TTL" as a single staleness-driven mechanism; Jeroen's ruling
correctly separates it into two policies with different triggers and
different consequences on a miss.
### Tier 1 — Global (Region/orbital rung): keep-always, my round-1 number stands
**~174 MB PNG-encoded across all ~273 bodies** (recomputed check against
Table B's step-0 numbers above: step 0's capped-tile mosaic shape, per
T-1143 §4, is smaller per-body than a full district-spacing canvas — my
round-1 174 MB used district-spacing as a conservative upper bound, so 174 MB
remains a safe ceiling, not an underestimate, for whatever the actual
step-0/orbital canvas shape lands on).
- **D-203-shaped resource extension** — `BodyWorldState` (or a sibling
resource matching its exact pattern) gains an `orbital_canvas:
Option<EncodedStepCanvas>` field.
- **Never evicted by time-since-last-visit.** This is the one tier Jeroen's
amendment explicitly excludes ("the global tier alone is keep-always") —
"always keep the global level" from the original outline stands unmodified
by the amendment; the amendment only sharpens what happens *below* this
tier.
- **Still an evictable cache in D-227's sense, not stored truth** — a
corrupted or manually-cleared global-tier entry is a recompute (~24 ms
hydrology + the derive cost from Table B row 0), never data loss. "Keep-
always" is a *policy choice about when eviction runs*, not an exemption
from D-227's "cache, never source of truth" discipline.
### Tier 2 — Sub-global geometry (District/Quarter/Block/Tile step canvases): dual-axis eviction
This is where Jeroen's amendment lands. Two independent axes, evaluated
separately, either one sufficient to evict an entry:
**Axis 1 — staleness: NEVER for geometry.** Morphology, elevation, moisture,
vegetation, glaciation, settlement presence, river/cliff geometry — every
field D-227 covers as `derive(seed, position)` — is byte-identical on every
recompute. There is no staleness concept for these fields at all; an entry
that exists is always correct, forever, regardless of age. This part of my
round-1 proposal (the `time_decay`/`distance_decay` multiplicative TTL
formula) was **wrong to frame as staleness** — re-reading Jeroen's ruling, I
was solving "when does this become wrong" for data that is never wrong. What
I actually needed was axis 2.
**Axis 2 — storage-budget eviction: time-since-last-visit, a distinct
mechanism.** A per-entry `last_accessed` timestamp (the exact field D-203's
`BodyWorldState` already carries — I'm reusing the pattern, not inventing a
new one), swept periodically (not per-tick — a coarse background sweep,
matching the "not blocking for user output" premise). An entry whose
`last_accessed` exceeds a per-rung threshold is evicted **for storage
reasons, not correctness reasons** — the distinction matters operationally:
a storage-eviction miss is silently identical in cost to a cold-start miss
(re-derive at Table B's measured cost, 17 ms–1.8 s depending on rung), it is
never treated as an error or a "the data might be stale, re-verify"
condition, because there is nothing to verify — it's a pure function, it
recomputes to the same bytes every time.
**Concrete formula, revised from round 1:**
```
evict_if: time_since_last_visit(entry) > STORAGE_TTL[rung]
```
One term, not the three-term multiplicative formula I proposed in round 1 —
Jeroen's amendment removes the need for `distance_decay` and `detail`-scaling
as *separate* factors once staleness is off the table: what's left is purely
"how long has it been since anyone looked at this," which is
time-since-last-visit alone. `STORAGE_TTL[rung]` is still the one per-rung
tunable (deeper rungs plausibly get a shorter floor, since they're both
cheaper to regenerate and cover less ground per entry — same reasoning as my
round-1 doc, just now attached to a storage-thrift rationale rather than a
staleness one), but the mechanism collapses to a single, legible sweep rather
than a compound scoring formula. This is a simpler design than what I
proposed in round 1, which I'll flag plainly: **Jeroen's amendment made my
original proposal simpler, not more complex** — separating the two axes
removed a term instead of adding one, because "distance from current focus"
turns out to be redundant with "time since last visited" for a
storage-thrift purpose (a location far from current focus that the player
*just* visited doesn't need evicting for storage reasons yet; one they
haven't touched in a long session, regardless of current distance, does) —
distance matters for *prioritizing what to precache next*, which is a
different mechanism (smart precache, premise 1) than *what to evict for
storage*, and conflating them in one formula was my round-1 mistake.
**What "player visits a planet, never goes back" concretely evicts under
this policy:** every sub-global step-canvas entry for that body — District
through Tile, every step index, every center the player's viewport ever
requested — ages out independently on its own `last_accessed`. The global
(Region/orbital) entry for that same body does NOT evict (Tier 1's
keep-always rule), so "atlas navigation snappy after first calc" survives
even for an abandoned body — the player can still open the Atlas and see the
planet at global zoom instantly; only the finer-grained "I was standing at
this exact district window" entries age out, which is exactly the storage
being reclaimed (per-step canvases are the bulk of the byte cost — District
alone is 16.88 MB per cached window per Table B, vastly larger than the
174 MB *total* global tier across all 273 bodies).
**Sim-state fields (frozen/flooded) — the map-time-axis TTL, separate from
both axes above, unaffected by this amendment.** Jeroen's lead-interview
ruling on the map time axis (a different ruling from the storage-eviction
amendment, but adjacent) already established: static geometry cached
indefinitely-fresh, frozen/flooded carried as separately-cached short-TTL
planes re-requested as sim time advances. This is a genuine staleness TTL
(the sim state actually can become wrong as game time passes), and it's
orthogonal to the storage-eviction axis above — a frozen/flooded plane has
both a staleness TTL (it goes wrong after some sim-time interval) AND is
subject to the same storage-eviction sweep as its parent geometry entry (if
nobody's visited the body in a long time, evict the whole cached window,
sim-state plane included, for storage reasons — regenerating both on the
next visit is cheap either way).
### Tier 3 — client-side (Stig's, composing not competing, per lead-interview ratification)
Restating only to confirm compatibility, not re-designing Stig's tier:
Stig's `FileAccess` disk cache dir + in-memory LRU (round-1 §4) already
implements the identical two-axis split independently — his "geometry:
LRU-evict-only, no TTL" is my Axis 1 (never stale) plus his own LRU acting as
*his* storage-thrift mechanism (client disk budget, not server RAM budget);
his "sim-state: explicit TTL sweep" is the same map-time-axis TTL described
above. **The two tiers were designed independently and landed on the same
two-axis shape** — worth naming as convergent validation of Jeroen's
amendment being the right cut, not just a server-side patch.
### Summary table
| Tier | Scope | Staleness eviction | Storage eviction | Size |
|---|---|---|---|---|
| 1 — Global | Region/orbital canvas, per body | Never (D-227 pure) | **Never** (Jeroen: keep-always) | ~174 MB, all bodies |
| 2 — Sub-global geometry | District→Tile step canvases | Never (D-227 pure) | **`time_since_last_visit > STORAGE_TTL[rung]`** | unbounded resident, bounded by sweep |
| 2b — Sim-state planes | Frozen/flooded overlays | **Yes** — real staleness TTL (map-time-axis ruling) | Same sweep as parent geometry entry | small, per-window |
| 3 — Client disk/memory | Stig's `FileAccess` + in-memory LRU | Mirrors tier 2/2b split exactly | Client-local LRU + disk budget | ~440 MB ceiling at 5×5, disk-safe |
**Note on the "5×5" figure in the row above:** that's Stig's round-1 citation
of red flag 2's original conservative estimate for the *global* tier
specifically (all 273 bodies' Region/orbital canvases, disk-resident) — it
predates his own measurement ⑥ px-band refinement above and isn't affected by
it, since ⑥'s finding is about *sub-global* steps (Block/Quarter/District)
staying 1×1, not about the global tier's own sizing. My own Tier 1 number
(~174 MB, PNG-encoded) is the tighter, measured figure for the same tier —
both numbers describe the same "keep-always" budget from two different
starting estimates (his conservative pre-④ ceiling vs. my post-④ measured
rate); not a live disagreement, just two vintages of the same calculation
worth reconciling to one number (mine) before this goes to filing.
---
## Summary for interview 2
1. **Envelope mechanics are lower-risk than the governance framing might
suggest** — the tagged-marker pattern the sixth shape needs is already
proven five times in `bridge/mod.rs`; this is one more `Inbound` variant
and one more `SimBridge` method, not a demux redesign. Legacy
`district_window` needs zero coexistence engineering — it just keeps
working until nothing requests it. Step-0/Region rides the same new
envelope as every other step, not the legacy whole-body family. One flat
tagged response (Araminta's default shape (a)), not a dense/sparse split
— the progressive-paint UX case is real but cheaper to get client-side
than by adding wire-protocol complexity for a cost split the measurements
don't show.
2. **Seed-chaining: independent re-derivation, ruled and argued from D-227
purity** — a finer step may call a coarser rung's derivation *function*
internally (the existing region-baseline-feeds-district precedent), but
never reads a coarser step's cached *response* as literal input. My
round-1 cost numbers are validated as the correct model for this shape.
3. **Three step-ladder skeletons costed against Tyre's D-243-disciplined
options; Option B (skip-chunk, 5 steps: Region→District→Quarter→Block→
voxel) recommended** — the only skeleton where every row is directly
measured or in a directly-measured cost band, because chunk (64 m) was
never benched and Option B is the one option that doesn't need it. Costs
don't disqualify Option A or C (nothing is unaffordable), they just can't
fully justify chunk as its own step. The D-243 factor unevenness matters
for serving at exactly one seam (Region→District, which lines up with the
global/sub-global cache boundary) and is pure UX pacing everywhere below
it — Jeroen can pick B vs C on feel alone for the lower rungs. My own
first draft of this section used an invented 4 m "Tile-adjacent" spacing
that isn't a real D-243 rung — corrected once Tyre's message caught it;
the deepest step also needs a display-ratio-sized canvas (216×384 m, not
the fixed-3840×2160 convention that works at every shallower step).
4. **Cache tiers finalized**: Jeroen's amendment simplified my round-1
three-term TTL formula into a cleaner two-axis model — geometry never
goes stale (removed from the eviction question entirely), sub-global
geometry evicts purely on time-since-last-visit for storage thrift, the
global tier alone is keep-always, sim-state planes keep their own
separate real staleness TTL. Stig's independently-designed client tier
landed on the identical two-axis shape, which I read as confirmation the
amendment cut the problem correctly.
Nothing in this round changed a round-1 cost number. What changed is
precision: the cliff wire-shape tension resolves to Araminta's sparse
framing, the seed-chaining question resolves to independent re-derivation
with a named precedent, the cache-tier formula gets simpler (not more
complex) once staleness and storage-thrift stop being conflated, and the
step-ladder section itself got one honest correction (the invented 4 m
spacing) caught by live coordination with Tyre rather than by me catching it
alone — worth naming, since the correction is exactly what round 2's
synthesis format is supposed to produce.