The complete workshop record: four round-1 positions, five round-2 syntheses (incl. Troblum's adversarial pass with addendum + final scorecard — all seven findings resolved), both lead interviews, Qatux's round notes and the 8-section workshop-outcomes.md (the lakes message-crossing documented as process history), measurement ⑥ (set_pixel/c1) + the population-survey and chunk/S2 addenda in the measurement docs, the brief's appendix updated through ⑥, and architecture-briefing-final.md — Jeroen's outline written back as-built (six-level ladder, lakes, ~9MB resident global tier). README row: Complete. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
9.4 KiB
title, ticket, owner, workshop, status
| title | ticket | owner | workshop | status |
|---|---|---|---|---|
| Measurement ⑥: GDScript CPU-colorize cost — Image.set_pixel vs PackedByteArray-direct (330K–8.3M cells) | T-1176 (body-map-viewer workshop, round 2 synthesis) | Stig | body-map-viewer | complete |
Measurement ⑥ — CPU-colorize cost (the c1 input)
What this answers
Round 1 (stig-round1.md §3) recommended CPU-side coloring for the terrain
RTT layer as the default, explicitly flagged as pending this exact
number, and named it the missing sixth measurement. Jeroen ruled at lead
interview 1: proceed, run ⑥ now, so the c1 shader-vs-CPU call lands with a
number instead of a "known-good fallback" hedge.
This measures the classify-byte → palette-lookup → write-into-Image loop —
the CPU-side half of the terrain-layer build, at the same three canvas sizes
(330K / 2.07M / 8.3M gridunits) everything else in the appendix was measured
at — so it can be compared directly against ④'s encode/decode numbers and
①'s derivation numbers on the same size axis.
Method
Headless is correct here, unlike ⑤. ⑤ (ImageTexture upload) required a
real GPU-backed display because Godot's headless dummy renderer fakes GPU
texture uploads — timing that path headless would have measured nothing.
This measurement has no GPU/rendering-driver dependency at all:
Image.set_pixel() and PackedByteArray element writes are ordinary
CPU-side data-structure operations, and Image.create()/
Image.create_from_data() allocate a plain in-memory buffer — none of it
touches the RenderingServer or a swapchain. Confirmed by inspection of the
actual production code path this mirrors
(atlas_window_overlay.gd::_rebuild_texture_if_needed, which itself never
does anything GPU-side until the separate ImageTexture.create_from_image
call — the exact boundary ⑤ already priced). Run via
godot --headless --path client --script res://tmp_drive_t_setpixel_c1.gd
from the main checkout, no DISPLAY needed.
- Driver: a temporary
SceneTreescript (client/tmp_drive_t_setpixel_c1.gd, deleted after this measurement — not committed, same convention as ⑤'stmp_drive_t1180.gd). - Two candidate implementations, since the delta between them is the c1
decision input (per the task instruction):
- (A)
Image.set_pixel(x, y, Color)per cell — the exact shapeatlas_window_overlay.gd's_rebuild_texture_if_needed()/_build_tile_texture()use today: read a classification byte from a per-cell array, look up aColorin a palette table, callimg.set_pixel(). - (B) Direct
PackedByteArraybuffer write +Image.create_from_data— skip the per-pixel method call andColorobject construction; look up a 4-byte RGBA quad and write it directly into a flat buffer at(row*width+col)*4, then hand the whole buffer toImage.create_from_data()once. Same output format (Image.FORMAT_RGBA8) as path A and as today's production code.
- (A)
- Realistic colorize loop, not a synthetic fill. The per-cell
classification array is a 17-zone D-239 §6 morphology-vocabulary spread
(
RandomNumberGenerator, fixed seed 1234,randi_range(0, 16)per cell) — every class actually gets looked up across the canvas, exercising real branchy palette-array access rather than a constant/degenerate case that would flatter either path. - 12 iterations per case, 2 discarded as warmup, 10 kept. Median, p95, min,
max computed over the 10 kept samples. Timed with
Time.get_ticks_usec()immediately around each colorize call (palette-array construction and the synthetic classification array are built once, outside the timed loop). - Run twice back-to-back for stability confirmation (below).
Environment
Same machine as ①–⑤ (16-core, Rayon 14, no GPU involvement in this particular measurement since it's headless). No background load running during this measurement.
Results — run A
| size | path | median_ms | p95_ms | min_ms | max_ms |
|---|---|---|---|---|---|
| 330K | A_set_pixel | 25.695 | 27.852 | 24.644 | 27.852 |
| 330K | B_byte_buffer | 51.033 | 51.510 | 49.246 | 51.510 |
| 2.07M | A_set_pixel | 165.170 | 174.554 | 158.683 | 174.554 |
| 2.07M | B_byte_buffer | 319.174 | 333.268 | 308.201 | 333.268 |
| 8.3M | A_set_pixel | 643.514 | 655.528 | 634.316 | 655.528 |
| 8.3M | B_byte_buffer | 1261.385 | 1284.790 | 1245.232 | 1284.790 |
Results — run B (stability re-run)
| size | path | median_ms | p95_ms | min_ms | max_ms |
|---|---|---|---|---|---|
| 330K | A_set_pixel | 26.186 | 26.892 | 25.238 | 26.892 |
| 330K | B_byte_buffer | 52.566 | 55.523 | 49.407 | 55.523 |
| 2.07M | A_set_pixel | 161.351 | 164.131 | 155.692 | 164.131 |
| 2.07M | B_byte_buffer | 316.721 | 329.357 | 310.894 | 329.357 |
| 8.3M | A_set_pixel | 641.453 | 742.036 | 637.308 | 742.036 |
| 8.3M | B_byte_buffer | 1266.614 | 1291.422 | 1252.194 | 1291.422 |
Run A and run B agree within ~2–3% on every case (the one outlier, 8.3M
A_set_pixel p95 at 742 ms vs 656 ms in run A, is a single high sample in a
10-kept-sample set — the median, the number this doc leads with, moves by
under 0.3ms between runs). Same stability profile as ⑤'s own two-run
confirmation.
Headline numbers (using run A, run B confirms stability)
| size | path | median | ns/cell |
|---|---|---|---|
| 330K (331,776 cells) | set_pixel |
25.7 ms | 77.5 ns/cell |
| 330K | byte-buffer | 51.0 ms | 153.8 ns/cell |
| 2.07M (2,073,600 cells) | set_pixel |
165.2 ms | 79.7 ns/cell |
| 2.07M | byte-buffer | 319.2 ms | 154.0 ns/cell |
| 8.3M (8,294,400 cells) | set_pixel |
643.5 ms | 77.6 ns/cell |
| 8.3M | byte-buffer | 1,261.4 ms | 152.1 ns/cell |
Per-cell rate is flat across all three sizes for both paths (~78 ns/cell
for set_pixel, ~153 ns/cell for byte-buffer, both within ~3% across the
25× cell-count range from 330K to 8.3M) — no cliff, matching the "flat rate
holds at scale" pattern every other measurement in this appendix (①②③④⑤) also
found. This is a clean, linearly-scaling GDScript-interpreter cost, not an
algorithmic blowup.
Image.set_pixel is ~2× FASTER than the direct PackedByteArray write —
the opposite of the naive assumption. I expected the raw-buffer path to
win by skipping Color object construction and the set_pixel method-call
overhead; measured, it loses by roughly 2×, consistently at every size. Read
on this: GDScript's own per-element PackedByteArray indexed write
(buf[off] = ..., four separate indexed writes per cell in path B) carries
enough per-access interpreter overhead of its own that it outweighs whatever
set_pixel's internal Color-to-RGBA8-conversion cost is; set_pixel is
presumably a single, more-optimized engine-side call per pixel rather than
four separate GDScript-level array-index operations. This is a genuinely
useful finding for implementation: prefer Image.set_pixel over hand-
rolled buffer writes in GDScript — the "avoid the method-call" instinct
that's often correct in compiled languages does not hold here.
What this means for the frame/interaction budget
643 ms at the largest canvas size (8.3M cells) is not free — this is the first CPU number in the whole appendix that DOES land near a budget that matters, and it changes the shape of the c1 recommendation.
Context against the rest of the appendix, same 8.3M-cell canvas:
- Server derivation (②): ~1.7–1.8 s parallel — the client colorize cost (0.64 s) is roughly a third of the server's own derive time, not a rounding error against it.
- Wire decode, PNG-per-field (④): ~91 ms — colorize is ~7× the decode cost at this size.
- ImageTexture upload (⑤): ~3–5 ms — colorize is ~130–200× the upload cost at this size.
At the 330K size (the workshop's own "acceptable fallback" resolution,
5×5 px/gridunit), set_pixel colorize costs 25.7 ms — comfortably
interactive on its own (single-digit frame budgets, well under any
reasonable "player is waiting for this step" tolerance), and this is the
size that matters most for the deep/mid steps of the ladder per Dudley's
round-1 recommendation (viewport-sized canvases, not the 8.3M "stress
ceiling" shape — his own round-1 doc is explicit that 8.3M is a canvas-size
stress test, not a realistic per-step request shape). At the sizes the
ladder will actually request in steady-state play, CPU colorize is cheap.
It only becomes a real cost at the largest canvas sizes this appendix
measured as an upper-bound stress case, which per Dudley's own viewport-
sizing recommendation should rarely if ever be requested as a literal
step-canvas payload.
Implication for the hold-fetch-swap sequence (§2 of my round-1 doc)
Colorize happens once per arrived step canvas, in the hold-fetch-swap sequence's "on arrival: decode, colorize, upload, swap" chain. At 330K cells (the realistic per-step size), colorize (25.7 ms) is now the second most expensive step in that chain after server derivation itself, ahead of both decode (④: ~3.6 ms at 330K) and upload (⑤: ~0.2 ms at 330K, create_from_image RGBA8) — worth naming explicitly since round 1 didn't have this number and could have under-priced the arrival-side cost.
Repro
# Driver script was client/tmp_drive_t_setpixel_c1.gd (temporary, deleted
# after this measurement — not in the committed tree). Launch command used:
cd /var/mnt/data/projects/settled-reach
godot --headless --path client --script res://tmp_drive_t_setpixel_c1.gd