Substantial quality pass on gemma_naming.py driven by user review of
the first real-mode smoke test output. The earlier run produced names
that read too sci-fi / epic-fantasy / same-y: Aureus, Aetheria,
Stellaris, Nexus, Elysium. Root cause analysis + fixes:
1. Runtime timestamps. The log prefix is now
`[HH:MM:SS +00h03m]` — clock time plus elapsed-since-start. Gives
the user an at-a-glance sense of how long the run has been going
without scrolling back to the banner.
2. System / body headers. When the loop enters a new system it prints
`── SYSTEM K/N GJ 71 — Tau Ceti (hop 0)`. Each body line now
shows `GJ71c (Threshold)` if the body has a proper_name in
systems.db, so the log reads like a tour of the reach rather than
a wall of body_id slugs. Preserved (already-named) bodies now log
a compact "(skip — N names already set)" line so progress is
visible even when no inference happened.
3. Prompt grounding overhaul. The old few-shot examples were all
classical/epic (Wolcott Beck, Nakamura Stream, Ribeiro do Sal,
Drayton Spine) which biased Gemma 2 2B toward Latin/Greek
coinages. New preambles use the shape:
"Settlers named X after themselves, after what they saw, or
after places back home. Most names are mundane, short, and
direct — a surname, a compass direction, a feature, a
practical description. Classical or epic names are rare."
Combined with grounded example pools, Gemma now produces names
like "Cooper's Creek", "Western Ridge", "The Highroad",
"Blackwood Creek", "Dustbowl".
4. Core corridor relabel. The "core" palette inflection was
"institutional Latin / pan-Anglo / Gateway-era", which pattern-
matched in Gemma's training data to "make up Latin-sounding
words" (→ Ardenia, Aurelia, Stellaris). Now it's
"administrative English / Gateway-era" and the outputs are
prosaic — Port Dundas, East Ridge, Meridian, Landing.
5. Rotating few-shot example pools. Each feature type now has 5-7
pools of 5-6 examples each. `_build_prompt()` picks a pool
deterministically per (body_id, local_id, attempt) so:
- Same feature always gets the same prompt (determinism preserved).
- Neighbouring features on the same body get different prompts
(output variance — the sampler doesn't collapse to a single
mode when you ask for 16 mountain names in a row).
- Retries rotate to a new pool, not just a bumped seed, giving
dedup failures a clean second attempt.
6. Cosmopolitan cultural variety in the examples. Earlier pools only
showed British/Australian, Korean/Japanese, Portuguese/Swahili,
German/Dutch/Nordic axes — the four reach corridors. Gemma learned
"names come in four flavours". New pools span Dutch, Nordic,
Italian, French, Polish, Hungarian, Czech, Spanish, Russian,
Finnish, Greek, Irish, Japanese, and British — teaching the model
that names can be any real Earth cultural register, not just the
corridor label. The result: actual Dutch names (Egelantier,
Hochland, Van Damhoeve), actual Nordic (Lundstad, Brygga),
actual Italian (Borgo Marconi, Piazza Nuova), etc.
7. First-name possessive pools. Per user feedback, settler naming
includes both surnames ("Cooper's Creek") and first names
("Clifford's Bay", "Maura's Run", "Yuki's Pool"). Each feature
type now has a dedicated first-name-possessive pool in addition
to the existing surname pool — the two rotate alongside so both
patterns show up without either dominating.
8. One "classical/Latinate" pool per feature type (≈17% of calls
given 5-7 pools per type). Keeps occasional Latin flavour without
making it dominant — the user explicitly noted that replacing
one pattern with another "is never a clean fix for a randomizer."
9. Earth-name blocklist expansion. The Gemma 2 model reached for
real European names ("Weser", "Rhine", "Reykjavik") in the first
real run. Added 21 European rivers (Rhine, Weser, Elbe, Oder,
Vistula, Loire, Rhône, Douro, Tagus, Ebro, Po, Arno, Tiber, …)
and 25 Nordic/Eastern European cities (Reykjavik, Oslo, Gdansk,
Krakow, Prague, Warsaw, Budapest, Belgrade, …). Case-insensitive
"The <name>" stripping still applies so "The Great Divide" also
matches "Great Divide".
Combined smoke test after these changes (10 real-mode prompts across
core + west_reach):
- core: Port Dundas, The Backbone, Dustbowl, Blackwood Creek
- west_reach: Egelantier, Hochland, Der Rücken, Lundstad, Klipfjord
- no placeholder residue, no markdown, no 5+ word outputs.
--shard is gone (dead code since GPU contention killed parallelism).
Resume semantics are still free: re-run the same command and
already-named bodies skip via the preserved path.
Parallelism via two concurrent sr-voice subprocesses does not work on
this ROCm + llama-cpp-rs setup — launching a second instance poisons
the first one's GPU context (both fall back to 0% GPU / 50% CPU
busy-loop and stop making progress). Verified empirically: single
shard runs cleanly at ~1.2s/feature, two shards deadlock.
Without a working parallel path, --shard is dead weight. Resume
semantics were already free: the pipeline skips bodies whose
markers.json has non-empty name fields (preserved path), so a
killed run re-starts just by re-running the same command.
Simplifications:
- Remove --shard argument and all slicing logic.
- Remove banner_shard / shard_offset / shard_n / shard_m plumbing.
- Rename internal total_shard_systems → total_systems.
- Default --log path is now .tmp/gemma_naming.log (was conditional
on --shard). Pass `--log -` to disable file logging.
- Startup banner now prints a one-line resume reminder so the user
can see at a glance that a killed run is recoverable.
New end-to-end pipeline that walks every markers.json in the reach and
fills empty `name` fields using the Gemma 2 voice pipeline via
`sr-voice serve --stdio`. Per D-191 §4: the same Gemma 2 pipeline the
client uses for NPC voicing also produces the atlas content, which is
dual-purposed as a quality test of the LLM plumbing.
Pipeline per body (hop-ordered, core-first):
1. Load markers.json; identify feature records whose `name` is
blank (null or ""). Hand-authored names are never overwritten;
the 6 template bodies and any partial authoring stay put.
2. Look up body context (planet_class, settlement_pattern,
cultural_corridor, population, economic_role) from systems.db.
3. Build a short corridor-aware few-shot prompt per feature type.
Prompts carry 3 concrete `Style: X. Answer: Y` examples so
Gemma 2 2B completes a pattern instead of generating to an
open-ended instruction — this is the single biggest lever
against placeholder echoes on a small model.
4. Stream the prompt into a long-lived sr-voice subprocess, read
the JSONL response, post-process (strip markdown, label
prefixes, brackets, reject 5+ word outputs and placeholder
tokens), check the earth-name blocklist, check per-(corridor,
feature_type) + per-body dedup, check the per-stem cap, retry
up to 3 times with a bumped seed.
5. On persistent failure, fall back to a deterministic palette
generator so every feature ends up with a name.
6. Write markers.json atomically and refresh atlas_* DB rows via
sync_markers_to_db. Commit the DB per body so a crash loses
at most one body of state.
7. Restart the sr-voice subprocess every `--refresh` requests
(default: 200) to prevent KV-cache context bleed.
Core design decisions:
- Determinism: per-(world_seed, body_id, feature_local_id, attempt)
seed so the full run is reproducible.
- Ordering: bodies are processed in ascending `hop_distance_from_gateway`
so core bodies get first pick at every unique Gemma output and
outer sectors fall into the palette fallback when they lose the
dedup race.
- Dedup scope: (cultural_corridor, feature_type) across the run,
PLUS a per-body cross-type set so the same name can't be a river
AND an ocean AND a mountain on the same world. Hand-authored names
are seeded into both sets on load so templates win priority.
- Stem cap: each non-generic root token (e.g. 'Arcturus', 'Meridian')
may appear at most `--stem-cap` times across the full run (default
20), preventing single-word runaway. Fallback names bypass the cap.
- Earth blocklist: 181 curated entries covering major Earth cities,
mountains, rivers, oceans, historical/colonial spellings, and
Greek/Roman mythology that reads too literally. Prefixed variants
('Nouveau Paris', 'New Tokyo') explicitly allowed per the product
intent that Earth-echo names are fine but must not dominate.
Leading 'The ' is stripped before comparison so 'The Great Divide'
also matches.
Operational features:
- `--shard N/M` slices the body list into M partitions for parallel
runs. Two terminals × `--shard 0/2` + `--shard 1/2` fits the
~2.5 GB/instance VRAM footprint twice under the 50% cap on a
16 GB AMD GPU and roughly halves wall time.
- `--log PATH` writes a timestamped tee of every status line to a
file. Default: `.tmp/gemma_naming.shard{N}of{M}.log` when a
non-trivial shard is in use.
- SQLite `PRAGMA journal_mode=WAL` + `busy_timeout=15000` so two
concurrent shards serialize writes without lock errors.
- Per-body progress lines report `body K/N`, `sys K/N`, and
`hop=H` so the user can watch core sectors finish first.
- Each body logs the new names it produced per feature type so the
user can eyeball quality as the run progresses.
- Checkpoint summary every 25 bodies: cumulative names, rate,
ETA — gives the log regular scroll points.
- `--mock` uses `server/sr-voice/mock-stdio.sh` for dry-fire
pipeline validation without a model load (tested end-to-end).
Supporting files:
- `tooling/planet-gen/earth_blocklist.txt` — 181 curated entries.
- `tooling/db/backfill_cultural_corridor.py` — one-off migration
that fills the `cultural_corridor` column on both `star_systems`
and `bodies` from the `geographic_sector` values. Before this
pass, 99.4% of rows (3221/3240) had a NULL cultural_corridor
despite `wiki_sync.py` being aware of the column — the wiki
index.md files only carry the sector header, which was never
propagated to the DB column. Idempotent, safe to re-run after
any wiki_sync rebuild, explicit transaction wrapper with
rollback on failure.
Full batch runtime estimate: ~20 hours single-shard / ~10 hours
double-shard on this hardware. Smoke tests across five hardened
iterations (v1–v5) on GJ71b/c/d/d-1/e confirm the pipeline produces
clean, varied, culturally-coherent names with zero post-processing
residue.
- ON DELETE CASCADE added to every atlas_* foreign key (atlas_body_grids,
atlas_cities, atlas_roads, atlas_railroads, atlas_pois, atlas_rivers,
atlas_oceans, atlas_mountain_ranges). Previously, deleting a body from
the bodies table or NULL-ing its terrain_reference would leave orphan
atlas rows forever — sync_markers_to_db only cleans up for bodies it
re-processes. The existing atlas tables in systems.db were dropped and
recreated with the new constraint; FK list now reports CASCADE.
- Atlas DDL deduplicated. systems-schema.sql is now the single source of
truth, bracketed by `-- BEGIN ATLAS INDEX` / `-- END ATLAS INDEX`
markers. generate_atlas.py reads that block via `_load_atlas_schema()`
and applies it at runtime, so there is no second copy of the DDL to
keep in sync. Adding a column requires one edit, not two.
- Uniqueness guard on city coordinates. `_enforce_unique_city_coords`
runs at the end of `place_cities` and deterministically perturbs any
duplicate (row, col) via a fixed spiral walk to the first free
walkable land cell. Rare in practice but the MST collapses to a
zero-distance edge otherwise, producing an empty A* path and silently
dropping the road.
- Grid header validation. `load_markers` now raises `AtlasGridMismatch`
if the loaded `grid: {w, h}` header does not match `GRID_W`/`GRID_H`.
Both the incremental-skip path and the regenerate path route through
this loader, so a hand-authored template shipping a different grid
size fails loud with a per-body error rather than silently producing
half-scale coordinates.
- Unused `seed_rng` parameter removed from `_analyse_terrain`. The
function is RNG-free (continent flood-fill, habitability scoring,
river-mouth dedup, cost grid — all pure functions of terrain). The
false API contract made it look like terrain analysis consumed RNG
state and had to be sequenced with downstream RNG use.
- `_score_capital_sites` river-mouth bonus now builds one sparse
accumulator with all mouth points set at once and runs a single
`gaussian_filter` call, instead of O(n_mouths) filter calls over
single-point images.
- `binary_dilation(analysis["land_mask"] == False)` replaced with the
idiomatic `~analysis["land_mask"]`, matching the convention used
elsewhere in the file.
- `atlas-generate` Makefile target now guards on
`SELECT COUNT(*) FROM bodies WHERE terrain_reference IS NOT NULL`.
On a fresh DB that count is 0 and the generator previously exited
"success" after processing zero bodies. The target now fails loud
with a pointer to `populate_terrain_reference.py`.
- `main.rs` SimRng defensive re-insertion gains a long comment
explaining the exact plugin-ordering hazard it guards against, so
future readers don't treat the line as dead code. Tied to #826.
Implements the Phase 3 atlas content generator per D-191 §3, §8, and §9.
Pipeline per body (terrain-aware, deterministic per seed + body):
1. Simulate terrain via planet_simulation.simulate().
2. Analyse continents (flood-fill), habitability (temp/moisture/slope +
coastal bonus), river mouths, and a terrain A* cost grid.
3. Place cities sequentially — capital first (habitability + river-mouth
bias), then corridor growth via multi-source Dijkstra, quadrant-spread
penalty after 2 cities in a quadrant, port-on-new-continent bonus at
cities 3–4. ±25% noise for seed variation.
4. Generate roads and railroads as an MST over city positions, with
A* paths on the terrain cost grid (rail follows roads where possible).
5. Place a transit POI at the capital (15% chance to scatter to a
secondary city).
Output (canonical markers.json schema, pixel space per D-191 §8):
- cities: {id, name, kind, center:[r,c], population}
- roads: {id, name, kind, path:[[r,c],...]}
- railroads: {id, name, kind, path:[[r,c],...]}
- pois: {id, name, kind, center:[r,c]}
- existing rivers/oceans/mountain_ranges preserved untouched.
City names are left empty for gemma_naming.py (#833). Body population is
split across cities with geometric decay (capital ~50%, each subsequent
city half the previous). The 6 hand-authored bodies (Lendel, Edict,
Vuurkloof, Røros, Cairnside, Estrade) are detected by existing
`cities` and skipped for regeneration; their markers are still synced
to the DB index below.
Atlas index in systems.db (new):
- atlas_body_grids, atlas_cities, atlas_roads, atlas_railroads,
atlas_pois, atlas_rivers, atlas_oceans, atlas_mountain_ranges
- Scalar metadata mirror of every markers.json — the implant atlas app
and development queries can lookup cities/POIs/features without
scanning 267 JSON files. Polyline geometry stays in the markers.json
files next to the heightmaps (used by the renderer); the DB only
stores filterable scalar fields plus `point_count` as a length proxy.
- Schema lives in server/data/systems-schema.sql; generate_atlas.py
mirrors the CREATE TABLE IF NOT EXISTS block so it runs against any
DB state (matches the economy-db importer pattern).
- Populated and refreshed on every run. Each body's rows are deleted
and reinserted deterministically — no stale state.
Also fixes a pre-existing WIP bug in the quadrant-saturation penalty
loop (a stray outer `for r in range(GRID_H)` with unreachable breaks
meant only the NW quadrant was ever checked).
Runtime: 280s for all 267 inhabited bodies on a single core. 265 bodies
updated this run, 6 hand-authored bodies synced to DB without
regeneration.
Atlas index after run:
atlas_cities 329 (15 hand-authored + 314 awaiting #833)
atlas_roads 46
atlas_railroads 44
atlas_pois 287
atlas_rivers 2034
atlas_oceans 696
atlas_mountain_ranges 1953
atlas_body_grids 267
Adds tooling/planet-gen/populate_terrain_reference.py and runs it against
systems.db. Resolves each body's expected wiki heightmap path (repo-root
relative) and writes it into bodies.terrain_reference. Missing heightmaps
are logged for remediation.
Result: 2380/3240 bodies populated, 860 still missing heightmaps. This
unblocks generate_atlas.py (#832) for every body that has a heightmap.
Custom pipeline for GJ-0 (Sol) that imports real NASA/USGS
planetary data instead of procedural generation. Produces the
same output format (heightmap.png, globe.png, markers.json).
Real data bodies:
- Earth: ETOPO2022 elevation + WorldClim climate + 14 rivers
- Mars: MOLA DEM + ferric biome classes + terraformed water
- Luna: LOLA DEM + lunar biome palette
Procedural fallback for Mercury, Venus, Phobos, Deimos.
Synthetic elevation from albedo for Io, Europa, Ganymede,
Callisto, Titan, Enceladus. Gas giants use existing renderer.
New biome classes 34-36 (ferric_dust/highland/lowland) for
Mars iron oxide surface. Earth features: 50 cities (smart
scatter by continent), 15 named rivers, oceans, mountains.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
arctan2(hz, hx) wrapped longitude counter-clockwise, mirroring
east and west on the globe. Changed to arctan2(hx, hz) which
increases eastward (right on screen). Added +0.5 offset to center
the view on 0° longitude, keeping the dateline seam on the back.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
1. Add 5 extended atmosphere colors to biomes.toml (globe renderer reads
from toml, not body_def) — cold_arid, hot_arid, tropical, boreal,
temperate_terminator. Remove dead martian entry. Slightly differentiate
colors from base classes.
2. Fix profile table raw identifiers — apply .replace('_', ' ') to class
field in wiki body pages.
3. Fix legend text overflow — truncate labels with ellipsis when step
size is too narrow for full text.
4. Fix title panel underscores — use .replace('_', ' ') instead of
.replace('_ringed', '').
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
1. Fix pclass.title() underscore bug in scaffold headings (216+ files)
2. Add 5 extended classes to batch.py valid_classes set
3. Add atmosphere rim colors for cold_arid/hot_arid/tropical/boreal/temperate_terminator
4. Add cloud/tilt ranges for extended classes (boreal 55-75%, tropical 60-80%)
5. Fix legend overflow at 1024px + deduplicate rainforest labels
6. Move _check_habitability() inside loop (was only checking last body)
7. Preserve gas_giant_ringed distinction in profile table
8. Drop meaningless terrain fields from gas giant frontmatter
9. Update stale docstrings/comments for 1024x512 default
10. Add infernal ring color
All lookup tables (CLASS_TILT, CLASS_CLOUD, CLASS_POLAR_ICE, CLASS_GEOTHERMAL,
CLASS_OBLATENESS, atmo_colors, tectonic_map, substrate_map) now include the
5 extended planet classes. Body content requires full regeneration.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The atomic .tmp→rename pattern caused silent FileNotFoundError on some
bodies. Removed in favour of direct writes — resume logic already handles
interrupted runs. Fixed --heightmap-size CLI flag which was silently
ignored due to Python default parameter binding. Changed default heightmap
resolution from 4096x2048 to 1024x512 (native simulation grid — no
information gain from upscaling).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
"Biome" describes per-zone vegetation classification (Whittaker table).
"Planet class" describes overall planetary character. The conflation
caused the planet generator to misclassify ~270 bodies as barren.
Scope: systems.db column, schema SQL, Rust atlas code, wiki table
headers (Biome → Class), atlas proposal JSONs, all docs/decisions,
tooling scripts. Also normalizes atmosphere vocabulary (breathable →
standard) and expands planet class mapping to all 26 wiki values.
Unknown classes default to temperate for modder safety.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Remove tracked kallast_terrain.npy (HIGH — binary in git)
- Amend D-086: stance + interaction icons delivered, not deferred
- Create ticket #818 for icon_tint.gdshader (client team)
- Fix gas giant profile table: suppress gravity/land%/hydrosphere
- Fix atmosphere_color: null when atmosphere is none
- Fix gas_giant_ringed display as gas_giant in profile table
- Re-scaffold + regenerate Ran system with fixes
.import sidecars: not applicable — gitignored by design (client/**/*.import).
Godot auto-generates on first run.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>