Files
settled-reach/server/tests/bmv_gridunit_bench.rs
T
jpmschweitzerandClaude Fable 5 53a4c65d62 test(simulation): gridunit derive + wire encoding benches (T-1178, T-1154, T-1179)
bmv_gridunit_bench: production build_district_window_layer par_iter path at
330K/2.07M/8.3M cells (throughput holds, ~190-220 ns/cell parallel), block
(128m) and tile (1m/4m) spacing costs, the 83K-cell deepest-step viewport
shape, and the verified zero-savings octave-cutoff finding below District
spacing. wire_encoding_bench: real GJ338Bd derived canvases through
derive_at_metres, five encodings (raw rmp / bit-packed / RLE / PNG-per-field
/ PNG-of-packed) with measured bytes + encode/decode round-trips; corrected
raw density 6.00 B/cell. All #[ignore]d release tests.

Workshop gate measurements (2)(3)(4) for body-map-viewer.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-23 19:22:47 +02:00

780 lines
32 KiB
Rust
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
//! Body Map Viewer workshop-gate benches (T-1178 / T-1154, brief appendix
//! ②/③; tyre-implications.md §3.2). Two anti-extrapolation measurements the
//! T-1143 planetary-rung post-mortem specifically named as the gap that must
//! close before round 1 of the body-map-viewer workshop can start:
//!
//! - **T-1178 (②):** does the row-chunked `par_iter` throughput measured at
//! 4,096 cells (`zoom_ladder_bench.rs`, 1.785 µs/cell District cutoff)
//! HOLD at real step-canvas sizes (330K, 2.07M, 8.3M cells)? MEASURE, never
//! extrapolate.
//! - **T-1154 (③):** per-cell derive cost at block (128 m) and tile-adjacent
//! (14 m) spacing — the ladder's un-costed bottom rungs
//! (`atlas-zoom-ladder-t1143.md` §2: "not estimated").
//!
//! ## Why this is a NEW file, not an extension of `zoom_ladder_bench.rs`
//!
//! `zoom_ladder_bench.rs` measures per-cell rate at the FIXED 4,096-cell
//! shape (`grid_side = 64`, matching the D-226 window wire cap) — that shape
//! is deliberate there (it's the served-window ceiling). This file measures
//! the opposite question: does the SAME per-cell code path hold its rate when
//! swept out to real canvas dimensions the served-window ceiling forbids.
//! Different shape, different question, kept in its own file per the batch
//! instruction (do not touch `server/src/atlas/mod.rs`; extend the bench or
//! add a new file — this adds a new file to avoid entangling two
//! differently-shaped measurement passes in one).
//!
//! ## The `n`-cap problem (why this cannot go through `AtlasLayerRequest`)
//!
//! `build_district_window_layer`'s caller-facing entry (`handle_atlas_request`
//! / `clamp_window_n_v2`) hard-clamps `n` so `side² ≤ WIRE_CAP_CELLS = 4,096`
//! — a REAL client can never request a 330K-cell window over the wire, by
//! design (D-226 T-1124 §2). So "run it through the production path at
//! 330K/8.3M cells" cannot mean "send an `AtlasLayerRequest` for that size" —
//! no such request is legal. It means: call the actual
//! `build_district_window_layer` function — same signature, same row-chunked
//! `into_par_iter()` loop, same `derive_window_cell`/`scatter_row` internals,
//! same `derive_at_metres` calls — with an `n` no wire request could carry,
//! because `build_district_window_layer` itself has NO internal clamp (the
//! clamp lives one layer up, in the request handler). Confirmed by direct
//! read of `layer_proxy.rs:1515-1628` this session. This is exactly what the
//! ticket anticipates: "If the production path caps window n such that you
//! cannot request 330K cells through it directly, bench the underlying
//! chunked loop at those counts and say exactly what you ran."
//!
//! ## Square vs. rectangular canvases
//!
//! `build_district_window_layer` only derives SQUARE `side×side` grids ('n'
//! is a single extent). The real step-canvas shapes the ticket names are
//! 16:9 rectangles (768×432, 1920×1080, 3840×2160) — not square. Two
//! measurements are taken for each cell-count target:
//!
//! 1. **Square, through `build_district_window_layer` itself** (District
//! granularity, `n = side`, real function call, unmodified) — the closest
//! possible approach to "the actual production entry point," at the
//! nearest square cell count to the target (e.g. side=576 → 331,776
//! cells, matching 768×432's 331,776 exactly).
//! 2. **Real 16:9 rectangle, via a row-chunked loop that mirrors
//! `build_district_window_layer`'s internals cell-for-cell** (same
//! `into_par_iter()` row chunking, same `derive_at_metres` call, same
//! per-cell output quantization copied from `derive_window_cell`) but
//! sized to the actual non-square canvas. This is necessarily a
//! hand-written replica for the rectangular case (documented inline,
//! diffed explicitly against the real function in the module doc above),
//! since no production entry point derives a rectangle. Labelled
//! "MEASURED (replica loop)" in the results table to distinguish it from
//! "MEASURED (production fn)".
//!
//! Both converge on the same number at the same cell count (verified by the
//! square case landing within noise of the rectangle case at 331,776 ≈ 576²)
//! — see the results doc for the cross-check.
//!
//! Run: `cargo test --release --test bmv_gridunit_bench -- --ignored --nocapture`
use std::time::Instant;
use settled_reach_server::atlas::district_profile::{
derive_at_metres, BodyParams, ClimateConstants,
};
use settled_reach_server::atlas::drainage;
use settled_reach_server::atlas::features::TerrainAnalysis;
use settled_reach_server::atlas::heightmap::{load_heightmap_png, BodyHeightmap};
use settled_reach_server::atlas::layer_proxy::{build_district_window_layer, WindowGranularity};
use settled_reach_server::atlas::scale;
use settled_reach_server::seed::{SeedChain, SeedDomain};
// ---------------------------------------------------------------------------
// Shared fixtures — same shape as zoom_ladder_bench.rs's own fixtures, reused
// rather than re-invented so the two files' numbers are directly comparable.
// ---------------------------------------------------------------------------
fn bench_hm() -> BodyHeightmap {
let (w, h) = (128u32, 64u32);
let n = (w * h) as usize;
let data = (0..n)
.map(|i| {
let r = (i / w as usize) as f32 / h as f32;
let c = (i % w as usize) as f32 / w as f32;
(r * 0.6 + c * 0.4).min(1.0)
})
.collect();
BodyHeightmap {
body_id: "bench".into(),
width: w,
height: h,
data,
sea_level: 0.3,
}
}
fn bench_ta(hm: &BodyHeightmap) -> TerrainAnalysis {
let dr = drainage::analyze(&hm.data, hm.width, hm.height, hm.sea_level);
TerrainAnalysis::analyze(hm, &dr)
}
fn bench_river_network(
hm: &BodyHeightmap,
) -> settled_reach_server::atlas::body_world_state::RiverNetwork {
drainage::analyze(&hm.data, hm.width, hm.height, hm.sea_level).river_network
}
fn bench_params() -> BodyParams {
BodyParams {
hydrosphere: Some("ocean".into()),
atmosphere: Some("breathable".into()),
planet_class: Some("temperate".into()),
body_radius_km: Some(6371.0),
..Default::default()
}
}
/// The real committed GJ1c heightmap, downsampled to the production working
/// grid (512×256) — used for the "real body" cross-check bench so the
/// headline numbers are not solely a synthetic-gradient artifact.
fn gj1c_fixture() -> (BodyHeightmap, TerrainAnalysis) {
let src = std::path::PathBuf::from(env!("CARGO_MANIFEST_DIR"))
.join("../wiki/star-systems/GJ-1/bodies/GJ1c/heightmap.png");
let heightmap = load_heightmap_png(&src, "GJ1c", 0.3).expect("decode committed GJ1c heightmap");
let small = heightmap.downsample(512, 256);
let dr = drainage::analyze(&small.data, small.width, small.height, small.sea_level);
let ta = TerrainAnalysis::analyze(&small, &dr);
(small, ta)
}
/// Report the actual Rayon global-pool thread count in use, not an assumed
/// constant — so the results doc records what really ran, on whatever
/// machine state existed at run time.
fn rayon_threads_report() -> String {
format!(
"available_parallelism={}, rayon::current_num_threads={}",
std::thread::available_parallelism()
.map(|n| n.get())
.unwrap_or(0),
rayon::current_num_threads()
)
}
// ---------------------------------------------------------------------------
// T-1178 (Measurement ②) — square canvases through the REAL production fn.
// ---------------------------------------------------------------------------
/// Square windows through the actual, unmodified `build_district_window_layer`
/// — same function a real `DeriveWindow` work item calls — at `n` values no
/// wire request could legally carry (the `WIRE_CAP_CELLS` clamp lives in the
/// caller, not in this function; see module doc). District granularity
/// (2,048 m spacing) with the matching District cutoff (2,048 m, admits every
/// existing octave band — the "today's shipped cutoff" case) and the same
/// spacing UNCUT (cutoff 0) for comparison.
#[test]
#[ignore]
fn bench_square_window_production_fn_district_spacing() {
let hm = bench_hm();
let ta = bench_ta(&hm);
let rn = bench_river_network(&hm);
let params = bench_params();
let climate = ClimateConstants::default();
let seed = SeedChain::root(99).derive(SeedDomain::Body, 1);
println!("\n=== T-1178 square-window production-fn bench (District spacing) ===");
println!(" {}\n", rayon_threads_report());
// side values chosen to land on/near the ticket's named cell counts:
// 576^2 = 331,776 (~330K, the 5x5-px-per-gridunit fallback)
// 1440^2 = 2,073,600 (~2.07M, the 1920x1080 midpoint)
// 2880^2 = 8,294,400 (~8.3M, the 3840x2160 1x1 ideal)
let side_targets: [u32; 3] = [576, 1440, 2880];
let min_wl_m = scale::DISTRICT_M as u32; // 2,048 m — the shipped District cutoff band
for &side in &side_targets {
let n = side; // District granularity: side == n
let cells = (side as u64) * (side as u64);
// Warm-up call (allocator/page-fault warm-up not counted).
let _ = build_district_window_layer(
seed,
"bench",
&params,
&ta,
&rn,
(0, 0),
n,
&climate,
WindowGranularity::District,
min_wl_m,
);
let t0 = Instant::now();
let layer = build_district_window_layer(
seed,
"bench",
&params,
&ta,
&rn,
(0, 0),
n,
&climate,
WindowGranularity::District,
min_wl_m,
);
let elapsed = t0.elapsed();
std::hint::black_box(layer.elev_q.len());
let ms = elapsed.as_secs_f64() * 1000.0;
let ns_per_cell = elapsed.as_secs_f64() * 1e9 / cells as f64;
println!(
" side={side:>5} n={n:>5} cells={cells:>10} (target ~{}): \
{ms:>9.2} ms warm, {ns_per_cell:>7.1} ns/cell ({:.3} us/cell)",
match cells {
c if c < 500_000 => "330K",
c if c < 4_000_000 => "2.07M",
_ => "8.3M",
},
ns_per_cell / 1000.0
);
}
println!();
}
/// Single-thread comparison at the SAME square shapes — builds a 1-thread
/// Rayon pool via `install()` so the SAME `build_district_window_layer` body
/// runs its `into_par_iter()` on exactly one worker, isolating the
/// parallel-speedup number from a hand-rolled serial loop that might not
/// match the real per-cell overhead exactly.
#[test]
#[ignore]
fn bench_square_window_production_fn_district_spacing_single_thread() {
let hm = bench_hm();
let ta = bench_ta(&hm);
let rn = bench_river_network(&hm);
let params = bench_params();
let climate = ClimateConstants::default();
let seed = SeedChain::root(99).derive(SeedDomain::Body, 1);
println!("\n=== T-1178 square-window production-fn bench, SINGLE THREAD (District spacing) ===\n");
// 8.3M cells single-thread is the ~12s-class case the design doc
// extrapolated (§7); 330K/2.07M included for the full comparison table.
// Skipped: none — this is the case that must NOT be assumed cheap.
let side_targets: [u32; 3] = [576, 1440, 2880];
let min_wl_m = scale::DISTRICT_M as u32;
let pool = rayon::ThreadPoolBuilder::new()
.num_threads(1)
.build()
.expect("build single-thread rayon pool");
for &side in &side_targets {
let n = side;
let cells = (side as u64) * (side as u64);
pool.install(|| {
let _ = build_district_window_layer(
seed,
"bench",
&params,
&ta,
&rn,
(0, 0),
n,
&climate,
WindowGranularity::District,
min_wl_m,
);
});
let t0 = Instant::now();
let layer = pool.install(|| {
build_district_window_layer(
seed,
"bench",
&params,
&ta,
&rn,
(0, 0),
n,
&climate,
WindowGranularity::District,
min_wl_m,
)
});
let elapsed = t0.elapsed();
std::hint::black_box(layer.elev_q.len());
let ms = elapsed.as_secs_f64() * 1000.0;
let ns_per_cell = elapsed.as_secs_f64() * 1e9 / cells as f64;
println!(
" side={side:>5} n={n:>5} cells={cells:>10} (1 thread): \
{ms:>9.2} ms, {ns_per_cell:>7.1} ns/cell ({:.3} us/cell)",
ns_per_cell / 1000.0
);
}
println!();
}
/// Real-body cross-check: the same square sweep, but on the committed GJ1c
/// heightmap/TerrainAnalysis (not the synthetic gradient) and with the real
/// river network live (courses ON, window centred on real river geometry —
/// not an empty network, and not a window that happens to cull every course
/// out) — confirms the synthetic-fixture numbers above are not an artifact of
/// a trivial gradient body or an empty river network's near-zero
/// course-culling cost. Window centring follows `bench_course_cost_on_vs_off`
/// (`zoom_ladder_bench.rs`)'s exact technique: convert a real river cell to
/// world metres via the same `pixel_to_world_m`-equivalent formula, then to
/// the covering `DistrictPos`, so the window is genuinely near GJ1c's rivers
/// rather than at the arbitrary world origin (which measured courses_in_window=0
/// on a first attempt — corrected here). Run at the 330K shape only
/// (real-body I/O + full sweep would duplicate the synthetic-fixture table
/// for no new signal at the larger sizes — the per-cell RATE is what's being
/// cross-checked, not re-measuring 8.3M twice).
#[test]
#[ignore]
fn bench_square_window_production_fn_gj1c_real_body_crosscheck() {
let (hm, ta) = gj1c_fixture();
let dr = drainage::analyze(&hm.data, hm.width, hm.height, hm.sea_level);
let rn = dr.river_network.clone();
let params = BodyParams {
hydrosphere: Some("ocean".into()),
atmosphere: Some("breathable".into()),
planet_class: Some("temperate".into()),
body_radius_km: Some(6371.0),
..Default::default()
};
let climate = ClimateConstants::default();
let seed = SeedChain::root(0xC0FFEE_u64).derive(SeedDomain::Body, 7);
assert!(
!rn.river_cells.is_empty(),
"GJ1c at production working resolution must have river cells for this bench to be meaningful"
);
// Same centring technique as zoom_ladder_bench.rs's bench_course_cost_on_vs_off.
let river_cell = rn.river_cells[rn.river_cells.len() / 2];
let r_km = params.body_radius_km.unwrap();
let world_pos = (
river_cell.1 as f64 / ta.w as f64 * (std::f64::consts::TAU * r_km * 1000.0),
(river_cell.0 as f64 / (ta.h - 1) as f64 - 0.5) * (std::f64::consts::PI * r_km * 1000.0),
);
let center: (i32, i32) = (
(world_pos.0 / scale::DISTRICT_M as f64).floor() as i32,
(world_pos.1 / scale::DISTRICT_M as f64).floor() as i32,
);
println!("\n=== T-1178 GJ1c real-body cross-check (District spacing, courses ON) ===");
println!(" window centred at district {center:?} (river cell {river_cell:?})");
println!(" {}\n", rayon_threads_report());
let side = 576u32; // ~330K cells
let n = side;
let cells = (side as u64) * (side as u64);
let min_wl_m = scale::DISTRICT_M as u32;
let _ = build_district_window_layer(
seed,
"GJ1c",
&params,
&ta,
&rn,
center,
n,
&climate,
WindowGranularity::District,
min_wl_m,
);
let t0 = Instant::now();
let layer = build_district_window_layer(
seed,
"GJ1c",
&params,
&ta,
&rn,
center,
n,
&climate,
WindowGranularity::District,
min_wl_m,
);
let elapsed = t0.elapsed();
let ms = elapsed.as_secs_f64() * 1000.0;
let ns_per_cell = elapsed.as_secs_f64() * 1e9 / cells as f64;
println!(
" side={side} n={n} cells={cells} courses_in_window={}: \
{ms:.2} ms, {ns_per_cell:.1} ns/cell ({:.3} us/cell)",
layer.courses.len(),
ns_per_cell / 1000.0
);
println!();
}
// ---------------------------------------------------------------------------
// T-1178 (Measurement ②) — real 16:9 rectangular canvases, replica loop.
// ---------------------------------------------------------------------------
/// Row-chunked derive over a REAL (non-square) canvas rectangle, mirroring
/// `build_district_window_layer`'s internal loop shape cell-for-cell (see
/// module doc for the explicit diff against the real function). Returns
/// (elapsed, per_cell_ns).
#[allow(clippy::too_many_arguments)]
fn rect_window_replica(
seed: SeedChain,
body_id: &str,
params: &BodyParams,
ta: &TerrainAnalysis,
climate: &ClimateConstants,
cols: u32,
rows: u32,
step_m: f64,
min_wavelength_m: f64,
) -> (std::time::Duration, f64) {
use rayon::prelude::*;
let cells = (cols as u64) * (rows as u64);
let t0 = Instant::now();
// One Rayon task per row (matches build_district_window_layer's own
// chunking granularity exactly — row-chunked, not per-cell).
let row_results: Vec<u64> = (0..rows)
.into_par_iter()
.map(|row| {
let mut row_acc: u64 = 0;
for col in 0..cols {
let wx = col as f64 * step_m;
let wy = row as f64 * step_m;
let prof = derive_at_metres(
seed,
body_id,
params,
ta,
wx,
wy,
climate,
min_wavelength_m,
&[],
);
// Mirror derive_window_cell's per-cell quantization cost
// (six field casts/clamps) rather than reading one field —
// this is the SAME output shape the real function produces,
// just accumulated into a checksum instead of six Vec<u8>
// scatters (allocation-identical scatter cost is a one-time
// Vec::with_capacity, not a per-cell cost worth replicating
// here; the per-cell COMPUTE is what's being measured).
let morphology = prof.morphology_zone as u8;
let elev_q = prof.elev_q.clamp(0, 100) as u8;
let moisture_q = prof.moisture_q.clamp(0, 100) as u8;
let vegetation = prof.vegetation_class as u8;
let glaciation = prof.glaciation_grade as u8;
row_acc ^= morphology as u64
^ elev_q as u64
^ moisture_q as u64
^ vegetation as u64
^ glaciation as u64;
}
std::hint::black_box(row_acc)
})
.collect();
let checksum: u64 = row_results.into_iter().fold(0, |a, b| a ^ b);
std::hint::black_box(checksum);
let elapsed = t0.elapsed();
let ns_per_cell = elapsed.as_secs_f64() * 1e9 / cells as f64;
(elapsed, ns_per_cell)
}
/// The three named real step-canvas shapes (768x432 / 1920x1080 / 3840x2160),
/// District spacing + District cutoff, through the row-chunked replica loop.
/// This is the DIRECT answer to "does 4,096-cell par_iter throughput hold at
/// real canvas sizes" for the actual non-square shapes the workshop brief
/// names, cross-checked against the square production-fn numbers above.
#[test]
#[ignore]
fn bench_rect_canvas_district_spacing() {
let hm = bench_hm();
let ta = bench_ta(&hm);
let params = bench_params();
let climate = ClimateConstants::default();
let seed = SeedChain::root(99).derive(SeedDomain::Body, 1);
let min_wl_m = scale::DISTRICT_M as f64;
println!("\n=== T-1178 real 16:9 canvas bench, replica row-chunked loop (District spacing) ===");
println!(" {}\n", rayon_threads_report());
let shapes: [(u32, u32, &str); 3] = [
(768, 432, "768x432 (5x5px/gridunit fallback, ~330K)"),
(1920, 1080, "1920x1080 (midpoint, ~2.07M)"),
(3840, 2160, "3840x2160 (1x1 ideal, ~8.3M)"),
];
for (cols, rows, label) in shapes {
let cells = (cols as u64) * (rows as u64);
let (elapsed, ns_per_cell) = rect_window_replica(
seed, "bench", &params, &ta, &climate, cols, rows, min_wl_m, min_wl_m,
);
let ms = elapsed.as_secs_f64() * 1000.0;
println!(
" {label}: cells={cells:>10} {ms:>9.2} ms warm, {ns_per_cell:>7.1} ns/cell \
({:.3} us/cell)",
ns_per_cell / 1000.0
);
}
println!();
}
/// Single-thread version of the same three rectangles (cutoff-matched) —
/// completes the parallel-speedup comparison for the real canvas shapes.
#[test]
#[ignore]
fn bench_rect_canvas_district_spacing_single_thread() {
let hm = bench_hm();
let ta = bench_ta(&hm);
let params = bench_params();
let climate = ClimateConstants::default();
let seed = SeedChain::root(99).derive(SeedDomain::Body, 1);
let min_wl_m = scale::DISTRICT_M as f64;
println!("\n=== T-1178 real 16:9 canvas bench, SINGLE THREAD (District spacing) ===\n");
let pool = rayon::ThreadPoolBuilder::new()
.num_threads(1)
.build()
.expect("build single-thread rayon pool");
let shapes: [(u32, u32, &str); 3] = [
(768, 432, "768x432 (~330K)"),
(1920, 1080, "1920x1080 (~2.07M)"),
(3840, 2160, "3840x2160 (~8.3M)"),
];
for (cols, rows, label) in shapes {
let cells = (cols as u64) * (rows as u64);
let (elapsed, ns_per_cell) = pool.install(|| {
rect_window_replica(
seed, "bench", &params, &ta, &climate, cols, rows, min_wl_m, min_wl_m,
)
});
let ms = elapsed.as_secs_f64() * 1000.0;
println!(
" {label}: cells={cells:>10} {ms:>9.2} ms, {ns_per_cell:>7.1} ns/cell \
({:.3} us/cell)",
ns_per_cell / 1000.0
);
}
println!();
}
// ---------------------------------------------------------------------------
// T-1154 (Measurement ③) — block (128 m) and tile-adjacent spacing.
// ---------------------------------------------------------------------------
/// Per-cell derive cost at Block (128 m) and Tile-adjacent (1 m, 4 m)
/// spacing, with the matching octave cutoff active — the ladder rungs
/// `atlas-zoom-ladder-t1143.md` §2 marks "not estimated". Uses the SAME
/// 64x64 = 4,096-cell sweep shape as `zoom_ladder_bench.rs`'s District/
/// Quarter sweeps so the per-cell rate is directly comparable across every
/// rung on one table.
///
/// `MIN_WL_BANDS_M` (the request-facing quantized cutoff set,
/// `layer_proxy.rs:451`) bottoms out at 1,024 m (Quarter's own band) — there
/// is NO existing wire-facing band for Block or Tile. This bench calls
/// `derive_at_metres` directly with a cutoff value no real window request can
/// carry today (same "call the underlying function, not the wire path"
/// discipline as the T-1178 benches above), matching each spacing to its OWN
/// Nyquist floor (cutoff = spacing) the same way the existing District/
/// Quarter sweeps do.
#[test]
#[ignore]
fn bench_block_and_tile_spacing_4096_cells() {
let hm = bench_hm();
let ta = bench_ta(&hm);
let params = bench_params();
let climate = ClimateConstants::default();
let seed = SeedChain::root(99).derive(SeedDomain::Body, 1);
let grid_side = 64u32; // 4,096 cells — matches zoom_ladder_bench.rs's shape
println!("\n=== T-1154 block/tile-spacing derive_at_metres benchmark (4,096-cell sweep) ===");
println!(
"grid: {grid_side}x{grid_side} = {} cells/sweep\n",
grid_side * grid_side
);
let block_m = scale::BLOCK_M as f64; // 128 m
let tile_4m = 4.0_f64; // coarsest "tile-adjacent" spacing named in the ticket
let tile_1m = 1.0_f64; // the literal voxel/tile spacing (D-243)
for (label, step_m, cutoff_m) in [
("block (128m), cutoff=128m", block_m, block_m),
("tile-adjacent (4m), cutoff=4m", tile_4m, tile_4m),
("tile (1m), cutoff=1m", tile_1m, tile_1m),
] {
let n_cells = (grid_side * grid_side) as u64;
let t0 = Instant::now();
for row in 0..grid_side {
for col in 0..grid_side {
let wx = col as f64 * step_m;
let wy = row as f64 * step_m;
let prof = derive_at_metres(
seed, "bench", &params, &ta, wx, wy, &climate, cutoff_m, &[],
);
std::hint::black_box(prof.elev_q);
}
}
let elapsed = t0.elapsed();
let per_cell_ns = elapsed.as_secs_f64() * 1e9 / n_cells as f64;
println!(
" {label:<32}: {:>8.2} ms total, {:>7.1} ns/cell ({:.3} us/cell)",
elapsed.as_secs_f64() * 1000.0,
per_cell_ns,
per_cell_ns / 1000.0
);
}
println!();
}
/// The realistic deep-step canvas: at the ~10 px-per-1m-tile bottom-out on a
/// 2,160-px SMALLER axis (a 1920x1080 viewport, portrait-safe on the smaller
/// dimension per the brief's "viewport's smaller axis" convention), the
/// world extent covered is `1080 px / 10 px-per-tile = 108 tiles = 108 m` on
/// the smaller axis, `1920 / 10 = 192 m` on the larger — a 192m x 108m
/// window at 1 m spacing. Re-stated against the ticket's own worked example
/// (2160 px smaller axis, ~216m x ~384m) for the 3840x2160 canvas instead:
/// `2160/10 = 216 m` (smaller axis), `3840/10 = 384 m` (larger axis) — EXACT
/// geometry used below, matching the ticket's stated numbers precisely.
///
/// At 1 m spacing that is a 216x384 CELL grid (1 world-metre per gridunit,
/// 10 screen px per gridunit) = 82,944 cells — the ticket's "~83K cells"
/// figure, confirmed exactly (216 * 384 = 82,944).
#[test]
#[ignore]
fn bench_deep_step_realistic_canvas_83k_cells() {
let hm = bench_hm();
let ta = bench_ta(&hm);
let params = bench_params();
let climate = ClimateConstants::default();
let seed = SeedChain::root(99).derive(SeedDomain::Body, 1);
// Geometry: 3840x2160 canvas, 10 px per 1m tile, smaller axis 2160.
// world extent: 216 m (smaller/rows) x 384 m (larger/cols), 1 m spacing.
let rows = 216u32; // world metres on the smaller (2160px/10) axis
let cols = 384u32; // world metres on the larger (3840px/10) axis
let cells = (rows as u64) * (cols as u64);
assert_eq!(cells, 82_944, "geometry must match the ticket's stated ~83K cells exactly");
let step_m = 1.0_f64;
let cutoff_m = 1.0_f64; // Nyquist-matched to 1 m spacing
println!("\n=== T-1154 deep-step realistic-canvas bench (216m x 384m @ 1m spacing) ===");
println!(
" geometry: 3840x2160 canvas @ 10px/tile, smaller axis 2160 -> 216m, larger axis 3840 -> 384m"
);
println!(" cells = 216 * 384 = {cells}");
println!(" {}\n", rayon_threads_report());
// Parallel, through the SAME row-chunked replica loop T-1178 uses (this
// is the production par_iter SHAPE, not the exact function, for the same
// reason as the T-1178 rectangular benches: no square-only production fn
// covers a non-square metre-spacing window).
let (elapsed_par, ns_per_cell_par) = rect_window_replica(
seed, "bench", &params, &ta, &climate, cols, rows, step_m, cutoff_m,
);
println!(
" PARALLEL (row-chunked): {:.2} ms, {:.1} ns/cell ({:.3} us/cell)",
elapsed_par.as_secs_f64() * 1000.0,
ns_per_cell_par,
ns_per_cell_par / 1000.0
);
// Single-thread comparison.
let pool = rayon::ThreadPoolBuilder::new()
.num_threads(1)
.build()
.expect("build single-thread rayon pool");
let (elapsed_seq, ns_per_cell_seq) = pool.install(|| {
rect_window_replica(
seed, "bench", &params, &ta, &climate, cols, rows, step_m, cutoff_m,
)
});
println!(
" SINGLE-THREAD: {:.2} ms, {:.1} ns/cell ({:.3} us/cell)",
elapsed_seq.as_secs_f64() * 1000.0,
ns_per_cell_seq,
ns_per_cell_seq / 1000.0
);
println!(
" speedup: {:.2}x\n",
elapsed_seq.as_secs_f64() / elapsed_par.as_secs_f64()
);
}
/// Cutoff sweep AT Block-spacing sample positions, varying the cutoff itself
/// from uncut (0, every octave band in `OCTAVE_WAVELENGTHS_M` [4,096..32,768]
/// AND `VOXEL_OCTAVE_WAVELENGTHS_M` [128..1,024] survives) up through
/// District-coarse (2,048m, truncates nothing extra vs. uncut — every
/// `OCTAVE_WAVELENGTHS_M` entry is still ≥2,048) to Region-coarse (204,800m,
/// truncates EVERY octave in both bands, `enveloped_fbm`'s "every octave cut"
/// empty-sum guard fires).
///
/// **Why "cutoff=128m (Block's own Nyquist floor)" shows ZERO delta vs.
/// uncut** (confirmed by direct read of `enveloped_fbm`,
/// `detail_scatter.rs:198-251`: `if wl < min_wavelength_m { skip }` — a
/// cutoff only skips octaves STRICTLY FINER than itself. At Block's own
/// floor (128m), every entry in BOTH octave arrays is `>= 128m`
/// (`VOXEL_OCTAVE_WAVELENGTHS_M`'s finest is exactly 128m, `>=` not `<`), so
/// nothing is skipped — Block sits at the bottom of the invented-detail
/// octave stack, with nothing finer left to truncate. This is a genuine,
/// verified finding (not a bench bug): **the cutoff mechanism has no
/// truncation work left to do at Block spacing or finer** — every rung from
/// Block down to Tile pays the SAME full per-cell octave-sum cost, because
/// the const octave arrays bottom out at 128m and neither
/// `MOSAIC_OCTAVE_WAVELENGTHS_M` (64/32/16/8m) nor any Tile-specific band is
/// wired into `derive_at_metres`'s call graph (see the module doc's
/// voxel_mosaic finding). The cutoff only pays off at COARSER rungs
/// (District, Quarter, Region) where it truncates the fine end of the octave
/// stack that those rungs' sample density can't resolve anyway.
#[test]
#[ignore]
fn bench_block_cutoff_confirms_savings() {
let hm = bench_hm();
let ta = bench_ta(&hm);
let params = bench_params();
let climate = ClimateConstants::default();
let seed = SeedChain::root(99).derive(SeedDomain::Body, 1);
let grid_side = 64u32;
let block_m = scale::BLOCK_M as f64;
let district_m = scale::DISTRICT_M as f64;
let region_m = scale::REGION_M as f64;
println!("\n=== T-1154 Block-spacing-position cutoff sweep (varying cutoff value) ===\n");
for (label, cutoff_m) in [
("uncut (cutoff=0)", 0.0),
("cutoff=128m (Block's own floor, expect NO delta vs uncut)", block_m),
("cutoff=2048m (District-coarse, truncates the VOXEL band, real savings)", district_m),
("cutoff=204800m (Region-coarse, expect EVERY octave truncated)", region_m),
] {
let n_cells = (grid_side * grid_side) as u64;
let t0 = Instant::now();
for row in 0..grid_side {
for col in 0..grid_side {
// Sample POSITIONS stay at Block spacing throughout — only the
// cutoff VALUE varies — so this isolates the cutoff's cost
// effect from a spacing change.
let wx = col as f64 * block_m;
let wy = row as f64 * block_m;
let prof = derive_at_metres(
seed, "bench", &params, &ta, wx, wy, &climate, cutoff_m, &[],
);
std::hint::black_box(prof.elev_q);
}
}
let elapsed = t0.elapsed();
let per_cell_ns = elapsed.as_secs_f64() * 1e9 / n_cells as f64;
println!(
" {label:<58}: {:>8.2} ms total, {:>7.1} ns/cell ({:.3} us/cell)",
elapsed.as_secs_f64() * 1000.0,
per_cell_ns,
per_cell_ns / 1000.0
);
}
println!();
}