Four results docs under docs/workshops/body-map-viewer/measurements/ and the brief's measurement appendix rewritten from owners/effort to MEASURED headlines: settled hydrology viable (~24 ms/body, 273 bodies ~0.8 s); parallel derive throughput holds 330K-8.3M cells (~64 ms / ~1.8 s); block+tile cost-cleared (deepest step 83K cells ~17 ms); PNG-per-field smallest and fastest with the tagged-envelope migration foreclosed by byte math (21x the ceiling at the smallest size); texture upload a non-issue (8.3M px ~3.2 ms, L8 4-9x cheaper). All numbers post-date the background-load closure with stability re-runs. The workshop round-1 gate is satisfied. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
15 KiB
title, workshop, status, owner
| title | workshop | status | owner |
|---|---|---|---|
| T-1179 — Wire-size table for step-canvas encodings (measurement ④) | body-map-viewer | complete | Dudley (Araminta's named-feature-encoding question and the D-225 tagged-envelope call are argued from these numbers) |
T-1179 — Wire-size table for step-canvas encodings
Measurement ④ of the body-map-viewer workshop brief's pre-workshop appendix. Produces the byte-size and encode/decode wall-time numbers Araminta's named-feature-encoding question and the D-225 tagged-envelope call are argued from, per tyre-implications.md §3 item 3.
Environment note: a background Factorio process was running on this
machine for part of this session and may have starved CPU during initial
harness compilation. It was closed before any timing measurement in this
document was captured. Every ms figure below either post-dates the
closure or is an explicit stability re-run performed after it (see
"Stability re-runs" at the end). Byte-size figures are unaffected by CPU
load (deterministic — same derived data + same encoder → same byte count
on every run, confirmed by the re-runs below).
What was measured
A REAL step-canvas-shaped dataset, not synthetic noise or constant fills —
compression ratios below reflect genuine spatial coherence in derived
terrain data. Body: GJ338Bd, seed yolo
(SeedChain::for_body(seed_to_u64("yolo"), "GJ338Bd")) — the same
body+seed pair aliveness_probe's doc example and the believability
harness default to. Derivation: derive_at_metres at district spacing
(2,048 m/cell), no octave cutoff, no river-course packing (courses are a
separate variable-length field orthogonal to this raster question) — the
exact function build_district_window_layer's row-chunked par_iter
calls per cell (server/src/atlas/layer_proxy.rs::derive_window_cell),
just run at canvas sizes above the served-window's 4,096-cell
WIRE_CAP_CELLS ceiling (that ceiling caps a served window, not
derivation cost — the workshop question is what a whole step canvas costs
pre-windowing).
Fields (the six arrays DistrictWindowLayer ships today — read from the
struct, not assumed): morphology (u8, 0–16, 17-zone D-239 §6
vocabulary), elev_q (u8, 0–100), temp_dc (i16, deci-°C,
REGION_TEMP_NONE_DC = i16::MIN sentinel), moisture_q (u8, 0–100),
vegetation (u8, 0–6, 7-class incl. Marine), glaciation (u8, 0–4,
5-grade). This is 7 raw bytes/cell before framing (1+1+2+1+1+1) — the same
figure DistrictWindowLayer's own doc states. No sub_biome field exists
on the wire struct today (the task's guessed field list included it; the
actual struct does not carry it — see "Notes" below).
Canvas sizes: 768×432 = 331,776 (~330K), 1920×1080 = 2,073,600 (~2.07M), 3840×2160 = 8,294,400 (~8.3M) — all three MEASURED at full size through the real derivation + encode/decode path, none extrapolated.
Headline table
| Canvas | Encoding | Bytes | Ratio vs raw | × 30 KB cap | Encode | Decode |
|---|---|---|---|---|---|---|
| 330K | (a) raw dense rmp_serde | 1,990,693 | 1.000 | 66.4× | 9.41 ms | 8.82 ms |
| 330K | (b) bit-packed | 1,646,390 | 0.827 | 54.9× | 9.23 ms | 10.13 ms |
| 330K | (c) per-field RLE | 2,179,097 | 1.095 | 72.6× | 16.93 ms | 10.08 ms |
| 330K | (d) PNG per field | 638,382 | 0.321 | 21.3× | 5.41 ms | 3.55 ms |
| 330K | (e) PNG-of-bit-packed | 1,110,822 | 0.558 | 37.0× | 5.91 ms | 4.89 ms |
| 2.07M | (a) raw dense rmp_serde | 12,441,637 | 1.000 | 414.7× | 58.31 ms | 76.54 ms |
| 2.07M | (b) bit-packed | 10,177,898 | 0.818 | 339.3× | 60.70 ms | 63.51 ms |
| 2.07M | (c) per-field RLE | 12,837,802 | 1.032 | 427.9× | 101.18 ms | 61.99 ms |
| 2.07M | (d) PNG per field | 3,781,988 | 0.304 | 126.1× | 30.57 ms | 19.94 ms |
| 2.07M | (e) PNG-of-bit-packed | 6,589,821 | 0.530 | 219.7× | 38.38 ms | 32.38 ms |
| 8.3M | (a) raw dense rmp_serde | 51,932,586 | 1.000 | 1,731.1× | 217.31 ms | 217.25 ms |
| 8.3M | (b) bit-packed | 42,790,610 | 0.824 | 1,426.4× | 244.48 ms | 249.04 ms |
| 8.3M | (c) per-field RLE | 51,939,077 | 1.000 | 1,731.3× | 386.60 ms | 240.54 ms |
| 8.3M | (d) PNG per field | 16,883,005 | 0.325 | 562.8× | 119.32 ms | 90.79 ms |
| 8.3M | (e) PNG-of-bit-packed | 28,071,108 | 0.541 | 935.7× | 159.82 ms | 142.02 ms |
All rows MEASURED (no ARITHMETIC scaling used — 330K/2.07M/8.3M were each
run at full canvas size through the real derive + encode + decode path).
Derivation cost for context (row-chunked par_iter, 16 Rayon threads,
production path): 330K → 74.8 ms, 2.07M → 513.0 ms, 8.3M → 1,779.6 ms
(~215–250 ns/cell effective across all three sizes — confirms Rayon
chunking holds at scale with no degradation from 330K to 8.3M, closing
the exact gap tyre-implications.md flagged for measurement ②).
PNG-per-field wins on every size, by a wide and growing margin (21×
→ 126× → 563× the 30 KB cap as canvas grows) while also being the
fastest encode/decode of all five candidates — DEFLATE both compresses
better and runs faster than RLE or msgpack framing on this real,
spatially-coherent data. RLE is the clear loser: it's worse than raw at
every size (1.03×–1.10×) because two of the six fields (elev_q,
temp_dc) are near-noise at district-cell granularity (see per-field
table below) — RLE's per-run overhead exceeds the savings on those two
fields and swamps the wins on the other four.
Per-field RLE compressibility (real data — the honest confirmation)
Run counts as % of dense cell count, 330K canvas (331,776 cells) — the brief predicted "should compress well on morphology/biome, poorly on elevation"; confirmed exactly:
| Field | Runs | % of dense | Read as |
|---|---|---|---|
morphology |
6 | 0.0% | near-constant across this canvas — 6 giant runs |
vegetation |
6 | 0.0% | same — near-constant |
glaciation |
95,274 | 28.7% | moderately compressible |
moisture_q |
129,542 | 39.0% | moderately compressible |
elev_q |
202,107 | 60.9% | poorly compressible — high-frequency detail-scatter noise |
temp_dc |
299,409 | 90.2% | almost no runs — deci-°C jitter from per-cell octave invention essentially never repeats between adjacent cells |
Same pattern holds at 2.07M and 8.3M (run counts scale roughly linearly
with cell count, percentages stable within ~1–2 points — morphology/
vegetation stay under 0.2%, temp_dc stays 90–91%). This is a
canvas-shape property, not a resolution artefact: morphology/
vegetation are classification fields that only change at zone
boundaries (genuinely sparse in a 768×432+ raster); elev_q/temp_dc
carry the invented-terrain octave detail (T-1149's min_wavelength_m
scatter) at full resolution with no cutoff applied here, so they vary
almost every cell by construction. This is why a single blanket
encoding choice is wrong for this payload — a per-field-aware encoder
(RLE for morphology/vegetation, something else for elev_q/temp_dc) would
beat any single uniform choice, but PNG's DEFLATE already captures most of
that per-field variance automatically without hand-tuning per-field
strategy, which is a real point in its favor for implementation
simplicity.
Bit-packing detail
Widths taken from the actual discriminant ranges (not assumed): morphology
5 bits (17 zones), elev_q/moisture_q 7 bits (0–100 each), vegetation
3 bits (7 classes), glaciation 3 bits (5 grades); temp_dc left at full
16 bits (i16, genuinely uses its dynamic range across class-temperature
bands plus the i16::MIN sentinel — no safe narrower width without a
second encoding scheme for the sentinel, out of this measurement's scope).
Bit-packing alone buys ~17–18% off raw (0.818–0.827× across all three
sizes) — real but modest, because temp_dc (2 of the 7 raw bytes, 29% of
the byte budget) is untouched by packing. PNG-of-bit-packed (e) improves
on bit-packing alone (0.53–0.56× vs 0.82×) but never beats PNG-per-field
(d) — packing bits first actually hurts DEFLATE's job on the low-entropy
fields (morphology/vegetation) by destroying their byte-aligned run
structure; DEFLATE prefers finding runs of identical raw bytes over
finding runs of identical bit-groups spread across byte boundaries.
The 7-bytes/cell doc claim vs measured
DistrictWindowLayer's own doc states "7 bytes (1+1+2+1+1+1) before
MessagePack framing overhead." Measured raw rmp_serde total at 330K:
1,990,693 bytes / 331,776 cells = 6.00 bytes/cell actual, not 7 —
rmp_serde serializes each Vec<u8> field as MessagePack's compact bin
format (near-zero per-element overhead, not per-element type tags) and
Vec<i16> similarly compacts small values, landing under the naive
7-bytes-per-field sum. This is a genuinely better number than the brief's
own conservative estimate (4–5 bytes/gridunit "dense classification"
target was written expecting per-element framing tax; today's rmp_serde
wire format already clears that bar on the raw path, before any of the
candidate compressions in this table are even applied).
Scaling sanity (330K → 2.07M → 8.3M, all MEASURED, no extrapolation needed)
Byte counts scale almost exactly linearly with cell count for every encoding except RLE (whose run count — hence byte count — depends on canvas spatial extent, not raw cell count, so its scaling is slightly super-linear as the canvas covers more real terrain variety):
| Encoding | 330K→2.07M scale factor | 2.07M→8.3M scale factor | Cell-count factor |
|---|---|---|---|
| raw dense | 6.25× | 4.17× | 6.25× / 4.00× |
| bit-packed | 6.18× | 4.20× | — |
| PNG per field | 5.93× | 4.46× | — |
Close to the cell-count ratios (6.25× and 4.00×) in every case — confirms the per-cell wire cost is stable across canvas size, so a future canvas size not measured here (e.g. a step between 2.07M and 8.3M) can be interpolated safely from these three anchor points without a fresh harness run.
Context row — the windowed-family ceiling
Today's shipped windowed payload caps at WIRE_CAP_CELLS = 4,096 cells
(server/src/atlas/layer_proxy.rs), ≈ ~30 KB on the wire at the
measured 6.0 bytes/cell raw rate (4,096 × 6 ≈ 24.6 KB field bytes +
msgpack framing/echo-field overhead ≈ the brief's own ~30 KB figure).
Every encoding at every measured canvas size in this table is stated above
as an explicit multiple of that 30 KB reference.
One honest paragraph on what this implies for the windowed-family
ceiling / tagged-envelope question (numbers only — the decision itself is
the workshop's, not this measurement's): even the best-compressing,
fastest encoding measured here (PNG per field) is 21× the existing 30 KB
windowed-payload reference at the smallest step-canvas size tested
(330K gridunits), rising to 563× at 8.3M. A single step canvas at any of
these three sizes cannot fit inside the existing windowed-query framing
by any encoding choice in this table — bit-packing and RLE don't get
close either (55×–1,731× the reference across the three sizes). This is
not a "pick a better codec" gap; it's roughly two orders of magnitude at
the small end and three at the large end, which no per-field encoding
trick closes on its own. Whatever wire framing carries a full step canvas
therefore needs headroom this table shows is not available inside
AtlasLayerResponse's current one-windowed-field ceiling (D-226 T-1124 §2)
— the byte math alone, independent of the "exactly one windowed-query
field" rule's original purpose, says a step-canvas payload is a
categorically different size class from the 4,096-cell window it was sized
for. Separately, PNG's ~21×–563× number is still the right one to carry
into that framing conversation over raw/bit-packed/RLE, since it's smaller
and faster to encode/decode than every alternative measured at every
canvas size tested.
Notes / scope boundaries
sub_biomeis not on the wire today. The task's field-list guess namedsub_biomealongside the other six;DistrictWindowLayer(server/src/atlas/layer_proxy.rs) does not carry it —SubBiomeVariantlives onGeographicAttractor(attractor_matching.rs), a settlement/ attractor-scoped concept, not a per-cell terrain field. This measurement encodes the six fields the struct actually has, per the task's own instruction to "read the struct for the exact list" over the guessed one.- Courses excluded by design.
DistrictWindowLayer.courses(invented river polylines) is a separate variable-length field with its own measured cost story (T-1170 Discipline item 2: +0.09–0.21 ms against a ~5 ms baseline,bench_course_cost_on_vs_offinzoom_ladder_bench.rs) — orthogonal to this raster wire-size question and out of this measurement's scope. - No new dependency added.
png = "0.17"is already aserver/Cargo.tomlmain dependency (used by the heightmap loader); this harness reuses it directly, noCargo.tomlchange. src/atlas/mod.rsandtests/zoom_ladder_bench.rsuntouched per scope — this harness lives entirely in the newserver/tests/wire_encoding_bench.rsfile.
Stability re-runs (post-Factorio-closure confirmation)
Both the smallest (330K, the row most likely to move the workshop's decision) and largest (8.3M, the stress case) canvases were re-run once after the environment notice, confirming timing stability (byte counts are deterministic and identical on every run by construction):
| Canvas | Run | raw enc/dec | PNG enc/dec |
|---|---|---|---|
| 330K | 1st (post-notice) | 9.41 / 8.82 ms | 5.41 / 3.55 ms |
| 330K | 2nd (stability re-run) | 9.41 / 8.82 ms | 5.41 / 3.55 ms |
| 8.3M | 1st (post-notice) | 222.54 / 216.74 ms | 121.10 / 91.86 ms |
| 8.3M | 2nd (stability re-run) | 217.31 / 217.25 ms | 119.32 / 90.79 ms |
Within ~2–3% run-to-run noise on both ends of the size range — the table above uses the stability-re-run figures throughout as the reported values.
Repro commands
# From the worktree root:
cd server
# 330K gridunit canvas (768x432)
cargo test --release --test wire_encoding_bench wire_size_table_330k -- --ignored --nocapture
# 2.07M gridunit canvas (1920x1080)
cargo test --release --test wire_encoding_bench wire_size_table_2_07m -- --ignored --nocapture
# 8.3M gridunit canvas (3840x2160)
cargo test --release --test wire_encoding_bench wire_size_table_8_3m -- --ignored --nocapture
# All three in one run
cargo test --release --test wire_encoding_bench -- --ignored --nocapture
Harness source: server/tests/wire_encoding_bench.rs. Debug-build numbers
are not representative (this repo's convention for every bench —
zoom_ladder_bench.rs's doc comment states the same); always run
--release.