Files
settled-reach/docs/workshops/body-map-viewer/measurements/t-setpixel-c1.md
T
jpmschweitzerandClaude Fable 5 29c22cb728 docs(meta): body-map-viewer workshop — rounds, measurements, outcomes, as-built briefing
The complete workshop record: four round-1 positions, five round-2 syntheses
(incl. Troblum's adversarial pass with addendum + final scorecard — all seven
findings resolved), both lead interviews, Qatux's round notes and the 8-section
workshop-outcomes.md (the lakes message-crossing documented as process
history), measurement ⑥ (set_pixel/c1) + the population-survey and chunk/S2
addenda in the measurement docs, the brief's appendix updated through ⑥, and
architecture-briefing-final.md — Jeroen's outline written back as-built
(six-level ladder, lakes, ~9MB resident global tier). README row: Complete.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-24 10:57:45 +02:00

9.4 KiB
Raw Blame History

title, ticket, owner, workshop, status
title ticket owner workshop status
Measurement ⑥: GDScript CPU-colorize cost — Image.set_pixel vs PackedByteArray-direct (330K–8.3M cells) T-1176 (body-map-viewer workshop, round 2 synthesis) Stig body-map-viewer complete

Measurement ⑥ — CPU-colorize cost (the c1 input)

What this answers

Round 1 (stig-round1.md §3) recommended CPU-side coloring for the terrain RTT layer as the default, explicitly flagged as pending this exact number, and named it the missing sixth measurement. Jeroen ruled at lead interview 1: proceed, run ⑥ now, so the c1 shader-vs-CPU call lands with a number instead of a "known-good fallback" hedge.

This measures the classify-byte → palette-lookup → write-into-Image loop — the CPU-side half of the terrain-layer build, at the same three canvas sizes (330K / 2.07M / 8.3M gridunits) everything else in the appendix was measured at — so it can be compared directly against ④'s encode/decode numbers and ①'s derivation numbers on the same size axis.

Method

Headless is correct here, unlike ⑤. ⑤ (ImageTexture upload) required a real GPU-backed display because Godot's headless dummy renderer fakes GPU texture uploads — timing that path headless would have measured nothing. This measurement has no GPU/rendering-driver dependency at all: Image.set_pixel() and PackedByteArray element writes are ordinary CPU-side data-structure operations, and Image.create()/ Image.create_from_data() allocate a plain in-memory buffer — none of it touches the RenderingServer or a swapchain. Confirmed by inspection of the actual production code path this mirrors (atlas_window_overlay.gd::_rebuild_texture_if_needed, which itself never does anything GPU-side until the separate ImageTexture.create_from_image call — the exact boundary ⑤ already priced). Run via godot --headless --path client --script res://tmp_drive_t_setpixel_c1.gd from the main checkout, no DISPLAY needed.

  • Driver: a temporary SceneTree script (client/tmp_drive_t_setpixel_c1.gd, deleted after this measurement — not committed, same convention as ⑤'s tmp_drive_t1180.gd).
  • Two candidate implementations, since the delta between them is the c1 decision input (per the task instruction):
    • (A) Image.set_pixel(x, y, Color) per cell — the exact shape atlas_window_overlay.gd's _rebuild_texture_if_needed() / _build_tile_texture() use today: read a classification byte from a per-cell array, look up a Color in a palette table, call img.set_pixel().
    • (B) Direct PackedByteArray buffer write + Image.create_from_data — skip the per-pixel method call and Color object construction; look up a 4-byte RGBA quad and write it directly into a flat buffer at (row*width+col)*4, then hand the whole buffer to Image.create_from_data() once. Same output format (Image.FORMAT_RGBA8) as path A and as today's production code.
  • Realistic colorize loop, not a synthetic fill. The per-cell classification array is a 17-zone D-239 §6 morphology-vocabulary spread (RandomNumberGenerator, fixed seed 1234, randi_range(0, 16) per cell) — every class actually gets looked up across the canvas, exercising real branchy palette-array access rather than a constant/degenerate case that would flatter either path.
  • 12 iterations per case, 2 discarded as warmup, 10 kept. Median, p95, min, max computed over the 10 kept samples. Timed with Time.get_ticks_usec() immediately around each colorize call (palette-array construction and the synthetic classification array are built once, outside the timed loop).
  • Run twice back-to-back for stability confirmation (below).

Environment

Same machine as ①–⑤ (16-core, Rayon 14, no GPU involvement in this particular measurement since it's headless). No background load running during this measurement.

Results — run A

size path median_ms p95_ms min_ms max_ms
330K A_set_pixel 25.695 27.852 24.644 27.852
330K B_byte_buffer 51.033 51.510 49.246 51.510
2.07M A_set_pixel 165.170 174.554 158.683 174.554
2.07M B_byte_buffer 319.174 333.268 308.201 333.268
8.3M A_set_pixel 643.514 655.528 634.316 655.528
8.3M B_byte_buffer 1261.385 1284.790 1245.232 1284.790

Results — run B (stability re-run)

size path median_ms p95_ms min_ms max_ms
330K A_set_pixel 26.186 26.892 25.238 26.892
330K B_byte_buffer 52.566 55.523 49.407 55.523
2.07M A_set_pixel 161.351 164.131 155.692 164.131
2.07M B_byte_buffer 316.721 329.357 310.894 329.357
8.3M A_set_pixel 641.453 742.036 637.308 742.036
8.3M B_byte_buffer 1266.614 1291.422 1252.194 1291.422

Run A and run B agree within ~2–3% on every case (the one outlier, 8.3M A_set_pixel p95 at 742 ms vs 656 ms in run A, is a single high sample in a 10-kept-sample set — the median, the number this doc leads with, moves by under 0.3ms between runs). Same stability profile as ⑤'s own two-run confirmation.

Headline numbers (using run A, run B confirms stability)

size path median ns/cell
330K (331,776 cells) set_pixel 25.7 ms 77.5 ns/cell
330K byte-buffer 51.0 ms 153.8 ns/cell
2.07M (2,073,600 cells) set_pixel 165.2 ms 79.7 ns/cell
2.07M byte-buffer 319.2 ms 154.0 ns/cell
8.3M (8,294,400 cells) set_pixel 643.5 ms 77.6 ns/cell
8.3M byte-buffer 1,261.4 ms 152.1 ns/cell

Per-cell rate is flat across all three sizes for both paths (~78 ns/cell for set_pixel, ~153 ns/cell for byte-buffer, both within ~3% across the 25× cell-count range from 330K to 8.3M) — no cliff, matching the "flat rate holds at scale" pattern every other measurement in this appendix (①②③④⑤) also found. This is a clean, linearly-scaling GDScript-interpreter cost, not an algorithmic blowup.

Image.set_pixel is ~2× FASTER than the direct PackedByteArray write — the opposite of the naive assumption. I expected the raw-buffer path to win by skipping Color object construction and the set_pixel method-call overhead; measured, it loses by roughly 2×, consistently at every size. Read on this: GDScript's own per-element PackedByteArray indexed write (buf[off] = ..., four separate indexed writes per cell in path B) carries enough per-access interpreter overhead of its own that it outweighs whatever set_pixel's internal Color-to-RGBA8-conversion cost is; set_pixel is presumably a single, more-optimized engine-side call per pixel rather than four separate GDScript-level array-index operations. This is a genuinely useful finding for implementation: prefer Image.set_pixel over hand- rolled buffer writes in GDScript — the "avoid the method-call" instinct that's often correct in compiled languages does not hold here.

What this means for the frame/interaction budget

643 ms at the largest canvas size (8.3M cells) is not free — this is the first CPU number in the whole appendix that DOES land near a budget that matters, and it changes the shape of the c1 recommendation.

Context against the rest of the appendix, same 8.3M-cell canvas:

  • Server derivation (②): ~1.7–1.8 s parallel — the client colorize cost (0.64 s) is roughly a third of the server's own derive time, not a rounding error against it.
  • Wire decode, PNG-per-field (④): ~91 ms — colorize is ~7× the decode cost at this size.
  • ImageTexture upload (⑤): ~3–5 ms — colorize is ~130–200× the upload cost at this size.

At the 330K size (the workshop's own "acceptable fallback" resolution, 5×5 px/gridunit), set_pixel colorize costs 25.7 ms — comfortably interactive on its own (single-digit frame budgets, well under any reasonable "player is waiting for this step" tolerance), and this is the size that matters most for the deep/mid steps of the ladder per Dudley's round-1 recommendation (viewport-sized canvases, not the 8.3M "stress ceiling" shape — his own round-1 doc is explicit that 8.3M is a canvas-size stress test, not a realistic per-step request shape). At the sizes the ladder will actually request in steady-state play, CPU colorize is cheap. It only becomes a real cost at the largest canvas sizes this appendix measured as an upper-bound stress case, which per Dudley's own viewport- sizing recommendation should rarely if ever be requested as a literal step-canvas payload.

Implication for the hold-fetch-swap sequence (§2 of my round-1 doc)

Colorize happens once per arrived step canvas, in the hold-fetch-swap sequence's "on arrival: decode, colorize, upload, swap" chain. At 330K cells (the realistic per-step size), colorize (25.7 ms) is now the second most expensive step in that chain after server derivation itself, ahead of both decode (④: ~3.6 ms at 330K) and upload (⑤: ~0.2 ms at 330K, create_from_image RGBA8) — worth naming explicitly since round 1 didn't have this number and could have under-priced the arrival-side cost.

Repro

# Driver script was client/tmp_drive_t_setpixel_c1.gd (temporary, deleted
# after this measurement — not in the committed tree). Launch command used:
cd /var/mnt/data/projects/settled-reach
godot --headless --path client --script res://tmp_drive_t_setpixel_c1.gd