Files
settled-reach/docs/workshops/body-map-viewer/measurements/t1179-wire-table.md
T
jpmschweitzerandClaude Fable 5 473b654787 docs(meta): body-map-viewer measurement results + brief appendix (T-1177/T-1178/T-1154/T-1179/T-1180)
Four results docs under docs/workshops/body-map-viewer/measurements/ and the
brief's measurement appendix rewritten from owners/effort to MEASURED
headlines: settled hydrology viable (~24 ms/body, 273 bodies ~0.8 s);
parallel derive throughput holds 330K-8.3M cells (~64 ms / ~1.8 s); block+tile
cost-cleared (deepest step 83K cells ~17 ms); PNG-per-field smallest and
fastest with the tagged-envelope migration foreclosed by byte math (21x the
ceiling at the smallest size); texture upload a non-issue (8.3M px ~3.2 ms,
L8 4-9x cheaper). All numbers post-date the background-load closure with
stability re-runs. The workshop round-1 gate is satisfied.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-23 19:23:02 +02:00

15 KiB
Raw Blame History

title, workshop, status, owner
title workshop status owner
T-1179 — Wire-size table for step-canvas encodings (measurement ④) body-map-viewer complete Dudley (Araminta's named-feature-encoding question and the D-225 tagged-envelope call are argued from these numbers)

T-1179 — Wire-size table for step-canvas encodings

Measurement ④ of the body-map-viewer workshop brief's pre-workshop appendix. Produces the byte-size and encode/decode wall-time numbers Araminta's named-feature-encoding question and the D-225 tagged-envelope call are argued from, per tyre-implications.md §3 item 3.

Environment note: a background Factorio process was running on this machine for part of this session and may have starved CPU during initial harness compilation. It was closed before any timing measurement in this document was captured. Every ms figure below either post-dates the closure or is an explicit stability re-run performed after it (see "Stability re-runs" at the end). Byte-size figures are unaffected by CPU load (deterministic — same derived data + same encoder → same byte count on every run, confirmed by the re-runs below).

What was measured

A REAL step-canvas-shaped dataset, not synthetic noise or constant fills — compression ratios below reflect genuine spatial coherence in derived terrain data. Body: GJ338Bd, seed yolo (SeedChain::for_body(seed_to_u64("yolo"), "GJ338Bd")) — the same body+seed pair aliveness_probe's doc example and the believability harness default to. Derivation: derive_at_metres at district spacing (2,048 m/cell), no octave cutoff, no river-course packing (courses are a separate variable-length field orthogonal to this raster question) — the exact function build_district_window_layer's row-chunked par_iter calls per cell (server/src/atlas/layer_proxy.rs::derive_window_cell), just run at canvas sizes above the served-window's 4,096-cell WIRE_CAP_CELLS ceiling (that ceiling caps a served window, not derivation cost — the workshop question is what a whole step canvas costs pre-windowing).

Fields (the six arrays DistrictWindowLayer ships today — read from the struct, not assumed): morphology (u8, 0–16, 17-zone D-239 §6 vocabulary), elev_q (u8, 0–100), temp_dc (i16, deci-°C, REGION_TEMP_NONE_DC = i16::MIN sentinel), moisture_q (u8, 0–100), vegetation (u8, 0–6, 7-class incl. Marine), glaciation (u8, 0–4, 5-grade). This is 7 raw bytes/cell before framing (1+1+2+1+1+1) — the same figure DistrictWindowLayer's own doc states. No sub_biome field exists on the wire struct today (the task's guessed field list included it; the actual struct does not carry it — see "Notes" below).

Canvas sizes: 768×432 = 331,776 (~330K), 1920×1080 = 2,073,600 (~2.07M), 3840×2160 = 8,294,400 (~8.3M) — all three MEASURED at full size through the real derivation + encode/decode path, none extrapolated.

Headline table

Canvas Encoding Bytes Ratio vs raw × 30 KB cap Encode Decode
330K (a) raw dense rmp_serde 1,990,693 1.000 66.4× 9.41 ms 8.82 ms
330K (b) bit-packed 1,646,390 0.827 54.9× 9.23 ms 10.13 ms
330K (c) per-field RLE 2,179,097 1.095 72.6× 16.93 ms 10.08 ms
330K (d) PNG per field 638,382 0.321 21.3× 5.41 ms 3.55 ms
330K (e) PNG-of-bit-packed 1,110,822 0.558 37.0× 5.91 ms 4.89 ms
2.07M (a) raw dense rmp_serde 12,441,637 1.000 414.7× 58.31 ms 76.54 ms
2.07M (b) bit-packed 10,177,898 0.818 339.3× 60.70 ms 63.51 ms
2.07M (c) per-field RLE 12,837,802 1.032 427.9× 101.18 ms 61.99 ms
2.07M (d) PNG per field 3,781,988 0.304 126.1× 30.57 ms 19.94 ms
2.07M (e) PNG-of-bit-packed 6,589,821 0.530 219.7× 38.38 ms 32.38 ms
8.3M (a) raw dense rmp_serde 51,932,586 1.000 1,731.1× 217.31 ms 217.25 ms
8.3M (b) bit-packed 42,790,610 0.824 1,426.4× 244.48 ms 249.04 ms
8.3M (c) per-field RLE 51,939,077 1.000 1,731.3× 386.60 ms 240.54 ms
8.3M (d) PNG per field 16,883,005 0.325 562.8× 119.32 ms 90.79 ms
8.3M (e) PNG-of-bit-packed 28,071,108 0.541 935.7× 159.82 ms 142.02 ms

All rows MEASURED (no ARITHMETIC scaling used — 330K/2.07M/8.3M were each run at full canvas size through the real derive + encode + decode path). Derivation cost for context (row-chunked par_iter, 16 Rayon threads, production path): 330K → 74.8 ms, 2.07M → 513.0 ms, 8.3M → 1,779.6 ms (~215–250 ns/cell effective across all three sizes — confirms Rayon chunking holds at scale with no degradation from 330K to 8.3M, closing the exact gap tyre-implications.md flagged for measurement ②).

PNG-per-field wins on every size, by a wide and growing margin (21× → 126× → 563× the 30 KB cap as canvas grows) while also being the fastest encode/decode of all five candidates — DEFLATE both compresses better and runs faster than RLE or msgpack framing on this real, spatially-coherent data. RLE is the clear loser: it's worse than raw at every size (1.03×–1.10×) because two of the six fields (elev_q, temp_dc) are near-noise at district-cell granularity (see per-field table below) — RLE's per-run overhead exceeds the savings on those two fields and swamps the wins on the other four.

Per-field RLE compressibility (real data — the honest confirmation)

Run counts as % of dense cell count, 330K canvas (331,776 cells) — the brief predicted "should compress well on morphology/biome, poorly on elevation"; confirmed exactly:

Field Runs % of dense Read as
morphology 6 0.0% near-constant across this canvas — 6 giant runs
vegetation 6 0.0% same — near-constant
glaciation 95,274 28.7% moderately compressible
moisture_q 129,542 39.0% moderately compressible
elev_q 202,107 60.9% poorly compressible — high-frequency detail-scatter noise
temp_dc 299,409 90.2% almost no runs — deci-°C jitter from per-cell octave invention essentially never repeats between adjacent cells

Same pattern holds at 2.07M and 8.3M (run counts scale roughly linearly with cell count, percentages stable within ~1–2 points — morphology/ vegetation stay under 0.2%, temp_dc stays 90–91%). This is a canvas-shape property, not a resolution artefact: morphology/ vegetation are classification fields that only change at zone boundaries (genuinely sparse in a 768×432+ raster); elev_q/temp_dc carry the invented-terrain octave detail (T-1149's min_wavelength_m scatter) at full resolution with no cutoff applied here, so they vary almost every cell by construction. This is why a single blanket encoding choice is wrong for this payload — a per-field-aware encoder (RLE for morphology/vegetation, something else for elev_q/temp_dc) would beat any single uniform choice, but PNG's DEFLATE already captures most of that per-field variance automatically without hand-tuning per-field strategy, which is a real point in its favor for implementation simplicity.

Bit-packing detail

Widths taken from the actual discriminant ranges (not assumed): morphology 5 bits (17 zones), elev_q/moisture_q 7 bits (0–100 each), vegetation 3 bits (7 classes), glaciation 3 bits (5 grades); temp_dc left at full 16 bits (i16, genuinely uses its dynamic range across class-temperature bands plus the i16::MIN sentinel — no safe narrower width without a second encoding scheme for the sentinel, out of this measurement's scope). Bit-packing alone buys ~17–18% off raw (0.818–0.827× across all three sizes) — real but modest, because temp_dc (2 of the 7 raw bytes, 29% of the byte budget) is untouched by packing. PNG-of-bit-packed (e) improves on bit-packing alone (0.53–0.56× vs 0.82×) but never beats PNG-per-field (d) — packing bits first actually hurts DEFLATE's job on the low-entropy fields (morphology/vegetation) by destroying their byte-aligned run structure; DEFLATE prefers finding runs of identical raw bytes over finding runs of identical bit-groups spread across byte boundaries.

The 7-bytes/cell doc claim vs measured

DistrictWindowLayer's own doc states "7 bytes (1+1+2+1+1+1) before MessagePack framing overhead." Measured raw rmp_serde total at 330K: 1,990,693 bytes / 331,776 cells = 6.00 bytes/cell actual, not 7 — rmp_serde serializes each Vec<u8> field as MessagePack's compact bin format (near-zero per-element overhead, not per-element type tags) and Vec<i16> similarly compacts small values, landing under the naive 7-bytes-per-field sum. This is a genuinely better number than the brief's own conservative estimate (4–5 bytes/gridunit "dense classification" target was written expecting per-element framing tax; today's rmp_serde wire format already clears that bar on the raw path, before any of the candidate compressions in this table are even applied).

Scaling sanity (330K → 2.07M → 8.3M, all MEASURED, no extrapolation needed)

Byte counts scale almost exactly linearly with cell count for every encoding except RLE (whose run count — hence byte count — depends on canvas spatial extent, not raw cell count, so its scaling is slightly super-linear as the canvas covers more real terrain variety):

Encoding 330K→2.07M scale factor 2.07M→8.3M scale factor Cell-count factor
raw dense 6.25× 4.17× 6.25× / 4.00×
bit-packed 6.18× 4.20× —
PNG per field 5.93× 4.46× —

Close to the cell-count ratios (6.25× and 4.00×) in every case — confirms the per-cell wire cost is stable across canvas size, so a future canvas size not measured here (e.g. a step between 2.07M and 8.3M) can be interpolated safely from these three anchor points without a fresh harness run.

Context row — the windowed-family ceiling

Today's shipped windowed payload caps at WIRE_CAP_CELLS = 4,096 cells (server/src/atlas/layer_proxy.rs), ≈ ~30 KB on the wire at the measured 6.0 bytes/cell raw rate (4,096 × 6 ≈ 24.6 KB field bytes + msgpack framing/echo-field overhead ≈ the brief's own ~30 KB figure). Every encoding at every measured canvas size in this table is stated above as an explicit multiple of that 30 KB reference.

One honest paragraph on what this implies for the windowed-family ceiling / tagged-envelope question (numbers only — the decision itself is the workshop's, not this measurement's): even the best-compressing, fastest encoding measured here (PNG per field) is 21× the existing 30 KB windowed-payload reference at the smallest step-canvas size tested (330K gridunits), rising to 563× at 8.3M. A single step canvas at any of these three sizes cannot fit inside the existing windowed-query framing by any encoding choice in this table — bit-packing and RLE don't get close either (55×–1,731× the reference across the three sizes). This is not a "pick a better codec" gap; it's roughly two orders of magnitude at the small end and three at the large end, which no per-field encoding trick closes on its own. Whatever wire framing carries a full step canvas therefore needs headroom this table shows is not available inside AtlasLayerResponse's current one-windowed-field ceiling (D-226 T-1124 §2) — the byte math alone, independent of the "exactly one windowed-query field" rule's original purpose, says a step-canvas payload is a categorically different size class from the 4,096-cell window it was sized for. Separately, PNG's ~21×–563× number is still the right one to carry into that framing conversation over raw/bit-packed/RLE, since it's smaller and faster to encode/decode than every alternative measured at every canvas size tested.

Notes / scope boundaries

  • sub_biome is not on the wire today. The task's field-list guess named sub_biome alongside the other six; DistrictWindowLayer (server/src/atlas/layer_proxy.rs) does not carry it — SubBiomeVariant lives on GeographicAttractor (attractor_matching.rs), a settlement/ attractor-scoped concept, not a per-cell terrain field. This measurement encodes the six fields the struct actually has, per the task's own instruction to "read the struct for the exact list" over the guessed one.
  • Courses excluded by design. DistrictWindowLayer.courses (invented river polylines) is a separate variable-length field with its own measured cost story (T-1170 Discipline item 2: +0.09–0.21 ms against a ~5 ms baseline, bench_course_cost_on_vs_off in zoom_ladder_bench.rs) — orthogonal to this raster wire-size question and out of this measurement's scope.
  • No new dependency added. png = "0.17" is already a server/Cargo.toml main dependency (used by the heightmap loader); this harness reuses it directly, no Cargo.toml change.
  • src/atlas/mod.rs and tests/zoom_ladder_bench.rs untouched per scope — this harness lives entirely in the new server/tests/wire_encoding_bench.rs file.

Stability re-runs (post-Factorio-closure confirmation)

Both the smallest (330K, the row most likely to move the workshop's decision) and largest (8.3M, the stress case) canvases were re-run once after the environment notice, confirming timing stability (byte counts are deterministic and identical on every run by construction):

Canvas Run raw enc/dec PNG enc/dec
330K 1st (post-notice) 9.41 / 8.82 ms 5.41 / 3.55 ms
330K 2nd (stability re-run) 9.41 / 8.82 ms 5.41 / 3.55 ms
8.3M 1st (post-notice) 222.54 / 216.74 ms 121.10 / 91.86 ms
8.3M 2nd (stability re-run) 217.31 / 217.25 ms 119.32 / 90.79 ms

Within ~2–3% run-to-run noise on both ends of the size range — the table above uses the stability-re-run figures throughout as the reported values.

Repro commands

# From the worktree root:
cd server

# 330K gridunit canvas (768x432)
cargo test --release --test wire_encoding_bench wire_size_table_330k -- --ignored --nocapture

# 2.07M gridunit canvas (1920x1080)
cargo test --release --test wire_encoding_bench wire_size_table_2_07m -- --ignored --nocapture

# 8.3M gridunit canvas (3840x2160)
cargo test --release --test wire_encoding_bench wire_size_table_8_3m -- --ignored --nocapture

# All three in one run
cargo test --release --test wire_encoding_bench -- --ignored --nocapture

Harness source: server/tests/wire_encoding_bench.rs. Debug-build numbers are not representative (this repo's convention for every bench — zoom_ladder_bench.rs's doc comment states the same); always run --release.