Finalizes #849 core-world atlas cohesion: GJ0d (Earth/Sol) markers.json
cleaned of erroneous data, refine_log updated with Sol body gap notes,
atlas_quality_analysis.py added for ongoing metric tracking.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Per R-012: delete conversation.rs, both overheard content files, and
remove all 6 wire-up points (social_plugin, bridge/types, monologue,
voice/integration). Protocol version 22 → 23. Scope confirmed by
#842 audit — npc/ and content/global/ untouched. Surviving NPC
components (NpcName, NpcColorIndex, NpcConversation) migrated to
simulation/npc_components.rs for use by D-080 knowledge propagation.
Also applies pre-existing cargo fmt debt (names.rs and 4 others).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Addresses blocker comments from PR review:
- Delete star_map.gd and star_map.tscn — dead implant/map/starchart
HudGroups registration that should have landed with the atlas
unification (D-191 criterion 1)
- Remove test_star_map_is_accessible_from_insert_ui and
test_star_map_scene_exists from test_sprint30.gd — D-191 supersedes
the insert-UI accessibility pattern
- ImplantRegistry._scan() now detects default_key collisions
(first-wins with push_warning) and validates default_mode against
HudGroups.Mode enum (skip + warn on invalid)
star_map_data.json remains — still used by atlas_app and
economics_app for system index lookups.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Names all 33 null-name auto-detected Sol features: Earth ocean + 3 rivers,
24 Luna mountain ranges (real IAU lunar mountain names), 4 Mars mountains,
1 Europa mountain. All using real-world geographic names. Cross-reference
arcs added on Mars (Hellas-, Chryse-) and Europa (Conamara-, Pwyll-).
Adds sol_name_fixes.py for reproducible Sol feature naming. Updates refine
log to mark Sol complete with full audit metrics for all 6 touched systems.
DB synced: GJ0d, GJ0d-1, GJ0e, GJ0f-2 (all Sol inhabited bodies).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Systematic sweep eliminated all city name cross-body collisions across the
273 inhabited bodies. Started from Forum Veritas/Jade Harbor/Fort Iron
clusters identified during the #849 analysis pass.
Strategy: use world proper_name as capital city name wherever unique.
For worlds sharing a proper_name, author corridor-appropriate alternates.
All edits synced to atlas_cities via generate_atlas.py --body.
Before: 119+ cross-body city collisions, worst-case ×20 (Jade Harbor)
After: 0 cross-body city collisions
Clusters eliminated: Forum Veritas ×10, Jade Harbor ×19, Fort Iron ×10,
Eisenstadt ×7, Fjordheim/Fjordholm ×6 each, Eisenberg/Eisenfels/Hanseong ×5
each, Ridge Marker ×5, plus 20+ smaller clusters down to ×2.
River/ocean collisions (Rio Grande ×23 rivers, Steinbruch ×19, etc.)
remain — these affect uninhabited secondary bodies at scale and require a
dedicated batch-script pass.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Remove overheard.ron (1629 lines) and overheard.yaml.deprecated. D-078 overheard
system is retired — the content and production pipeline for ambient NPC dialogue
is deferred until the world is walkable (Phase 6). Deep module interdependencies
(perception, simulation, bridge) mean the server-side plumbing stays in place;
only the content files with no live consumers are removed.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
CultureResolver with Arc<Mutex<Connection>> over systems.db (SQLITE_OPEN_READ_ONLY).
3-pass lookup: system_id → body_id (COALESCE parent fallback) → station_id.
CultureResolverResource registered in main.rs with graceful warn-on-missing.
BookmarkRegistry.build_catalog() uses resolver for allowed_locations_cultures.
8 unit tests including concurrent safety. SQLite fixture at
server/src/knowledge/fixtures/culture_test.db.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
White-glove name pass on the five highest-traffic inhabited bodies in the
Ran (GJ 144) and Tau Ceti (GJ 71) systems. All markers.json edits synced
to atlas_* tables via generate_atlas.py.
GJ144d Kallast (2B pop): 4 fixes — "Aldren Pass" river renamed to
Randalfoss (avoids cross-system stem collision with Lendel's "Aldren");
two generic oceans renamed (Keldmere, Seterfjord — the latter cross-refs
mountain Seterfjellet); POI renamed to "Kallast Gate Terminal".
Established cross-ref arcs: Rán-, Seter-, Keld-.
GJ144e Vethis (1.2B pop): 9 fixes — 4 river renames (1 cardinal, 1
earth-echo, 2 generics), 1 ocean (Ash- overuse → Veth Mere), 3 mountain
renames (2 generics, 1 Ash- overuse). Established arcs: Grey- (4 names),
Thorn- (2), Kel- (3), Veth- (3), Ash- (2, down from 3).
GJ71c Threshold (600M pop): 1 fix — river "Aethelred" (Anglo-Saxon)
replaced with "Gaius" to complete the all-Latin survey-team arc (Octavius,
Septimus, Quintus, Valeria, Marcus, Gaius).
GJ71d Arden (500M pop): 2 fixes — "Concordia Hall" city renamed "The
Praxis" (Concordia = GJ71c ocean, cross-body stem collision); "Basilica
Nova" river renamed "Via Principia" (exact name match with GJ71c POI).
GJ71d-1 Verantis (20M pop): no name changes — mountains already updated
in prior pass (The Lateranum, The Curia Magna, etc.); DB sync only.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Adds two reusable scripts for the core-world hand-refine pass:
- atlas_cohesion_audit.py: SQL analysis against atlas_* tables. Reports
empty names, lazy/generic outputs, cardinal direction density, earth-echo
concentration, same-body cross-feature stem duplicates, and cross-body
stem collisions within a system. Supports --system, --body, --db flags.
Baseline run ranked Ran and Tau Ceti as highest-priority targets.
- apply_name_fixes.py: Applies curated name replacement tables to
markers.json files (name fields only; geometry preserved). Supports
--dry-run. After running, caller syncs DB via generate_atlas.py --body.
- refine_log_849.md: Hand-refine log documenting each body touched, the
rationale per change, cross-reference arcs established, and systems
flagged as blocked or needing follow-up (Sol, Barnard's Star, Proxima).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Introduces the "faux mobile OS" framing: ImplantApp base class,
ImplantNavStack, ImplantAppManifest (app.tres), and ImplantRegistry
autoload. Moddability is a first-class design driver — apps are
droppable directories discovered at startup, main.gd key routing is
manifest-driven, and D-169 primitives stay data-shape agnostic.
Phasing: full pattern lands in the Sprint 36 atlas refactor PR
(#844); Intents dispatcher and DataChannels seam are sketched but
deferred; shipped-build mod discovery stays Phase 6+.
Includes review checklist for #844 and nav-stack edge cases.
Unreleased entries for Sprint 36 client: #844 atlas unification, #724
version on loading screen, #722 --help on DB wrappers, plus the enum
renumber and symmetric back-nav tweak.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Loading screen reads the client version from project.yaml (root version
field) and displays it alongside Protocol.PROTOCOL_VERSION at the bottom
of the overlay: "v0.1.35 · protocol 21".
Falls back to "?.?.?" if project.yaml is missing or unreadable (e.g.
when run from an exported pck where the relative path is unavailable).
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Both wrappers previously silently treated --help as a SQL comment and
returned empty result JSON. They now intercept --help/-h before
delegating to the Python connector and print proper usage text with
the correct JSON key names (affected_rows, not rows_affected).
ticket, sprint, and decision already supported --help — no change.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Per D-191, the atlas is the star map extended downward — not a separate
app. AtlasPanel now owns the full Reach → system → planet → heightmap
zoom hierarchy as a single HudGroups app (implant/map).
- Add REACH_MAP as Level 0 of AtlasPanel's zoom hierarchy; renumber the
enum so higher index = deeper zoom
- Port hop-ring rendering (pan/zoom, system markers, hover/info, sector
layout) from star_map.gd into AtlasPanel methods
- Change AtlasPanel.APP_PATH from "implant/map/atlas" to "implant/map"
- main.gd: KEY_M toggles unified atlas; KEY_A binding removed
- Symmetric nav: ORBITAL_DIAGRAM back goes to REACH_MAP (not
SYSTEM_PICKER), matching the forward skip
- hud_groups.gd docstring documents the unified path and flags the
legacy starchart path as kept-for-compat (retirement tracked in #852)
star_map.gd's HudGroups registration stays live but inert — no key
binding reaches it. Full retirement follows in #852 after a sprint of
soak on the unified panel.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Sprint goal: close Phase 3 Atlas (unified nav chain, brand corps,
content refinement) and establish Phase 4 foundations (bookmark system,
location-culture resolution, character creation skeleton).
15 tickets assigned across server (7), client (5), copy (3).
Briefings written for all four teams. DB backup updated.
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
- Fix "Gemma 2 GGUF" in user-facing error message (line 2073)
- Fix gemma2.gguf in docstring usage example (line 32)
- Fix O(N) _is_duplicate: pre-build lowercase shadow sets for O(1) lookup
- Expand vestigial note to enumerate full ~750-line dead island boundaries
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Hoshe:
- Mark --dump-prompts / name_feature() as vestigial with TODO note
- Fix --refresh help string: 200 → 1000 (matches actual default)
- Fix _RIVER_POOLS comment numbering: Pool 6 before Pool 5 → correct order
- Remove dead first-pass code in fix_fewshot_bleed.py
- _CAPTURE_FILE leak noted in vestigial TODO
Tyre:
- Fix stale "Gemma 2" strings in banner, argparse description, model help
- Note dead code for cleanup pass (name_feature ~700 lines)
Hoshe (prune):
- prune_atlas_features.py: named features sort before unnamed, preventing
silent discard of hand-authored names during pruning
naming_core:
- v0.2: few-shot blocklist, stricter is_valid_name (min 3 chars, no digits,
no brackets), prompt fragment rejection expanded
Miri clarification: the 261 "empty-string" files contain only roads (37)
and railroads (37) — infrastructure features never in naming scope. All
cities/rivers/oceans/mountains/POIs are clean.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Full architecture doc covering the Gemma 4 batch naming pipeline:
pipeline stages, cultural registers, body ordering, known limitations,
QA process, and extension guide.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Replace the one-at-a-time Gemma 2 naming pipeline with a batch-oriented
Gemma 4 E2B pipeline. Key changes:
- naming_core.py: shared library with Levenshtein distinctiveness ranking,
batch prompt building, mood injection pool, name validation, and
adjacent-register refill logic
- Wiki-grounded register selection: per-system LLM call picks the cultural
register based on wiki/GTTR content instead of hash randomizer
- Batch naming: requests N*2 names per call, ranks by word-average
Levenshtein distance, fills quota from most-distinct candidates
- Mood pool: 13 emotional seeds randomized per-body for vocabulary
divergence (ambition, fear, isolation, defiance, etc.)
- Adjacent-register refill: when primary register exhausts, automatically
switches to next corridor substyle
- Inhabited-first body ordering: habitable worlds get first pick of
register vocabulary, barren moons get leftovers
- Process group cleanup: SIGTERM/SIGKILL the full distrobox chain on
subprocess refresh to prevent GPU zombie processes
- qa_naming.py: QA report, fix_fewshot_bleed.py: post-hoc fix script
- test_batch_naming.py, test_register_selection.py: test harnesses
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Individual dedup/blocklist/placeholder/empty rejections that recover
on the next attempt are now silent. Only the skipped: summary line
prints when all 5 attempts fail. Subprocess errors still print
immediately (those indicate a real problem).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Bumped max_attempts from 3 to 5 — with per-system dedup and no stem
cap, the remaining dedup hits are mostly per-body collisions which
a couple extra attempts with rotated pools can escape.
Bumped --refresh default from 200 to 1000. Fewer subprocess restarts
= fewer model reloads via distrobox. KV-cache bleed risk is lower
now that the validation gauntlet is lighter.
Reverted the batch-prompt experiment — Gemma 2 2B drifts on
multi-line output; individual calls are more reliable.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The stem cap (--stem-cap 20) was rejecting valid names because common
feature-type vocabulary tokens like "ridge", "hill", "range" hit the
cap after ~200 bodies and blocked all subsequent names containing
them. With sub-style rotation already providing variety, the cap was
doing more harm than good. Removed entirely.
Cross-body dedup narrowed from (hop, corridor, feature_type) to
(system_id, feature_type). Two rivers in the same system can't share
a name; two rivers in different systems can. This matches how
settlers actually name things — they don't coordinate with other
star systems.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Replace single-inflection corridor palettes with lists of sub-styles.
Each system picks one deterministically via hash(system_id), so all
bodies in the same system share a cultural register but neighbouring
systems get different registers.
Core corridor splits into 6 sub-styles (English rural, British
colonial, US rural, US cosmopolitan, classical/institutional,
Australian/NZ). North/south/east/west reach each get 5 sub-styles
covering their cultural spectrum. Deep frontier gets 3 (founder-name,
surveyor-descriptive, outpost-functional).
This multiplies Gemma's effective vocabulary per corridor by the
sub-style count, dramatically reducing dedup pressure. A 6-style
core corridor means each sub-style serves ~4 systems instead of 24,
so "The Ridge" exhausts after ~4 systems, not ~24.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
When all 3 LLM attempts are rejected (dedup, blocklist, etc.),
name_feature now returns None instead of a deterministic palette
fallback. process_body leaves the name as null in markers.json.
The preserved path (_is_blank) treats null as unnamed, so a fill
round (re-running the script) picks up only the skipped features
with a fresh corpus — zero dedup pressure from the first pass. The
fill round can use a different seed, slower prompt, or a different
backend entirely (e.g. Haiku).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Changed dedup key from (corridor, feature_type) to
(hop, corridor, feature_type). Systems at the same gate-hop distance
in the same corridor are near neighbors and shouldn't share feature
names; systems at different hops can. This prevents corpus exhaustion
where Gemma's narrow range-name distribution ("The Ridge", "Blackwood
Range") collides after ~20 bodies and drives fallback rates toward
100%.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The upstream terrain pipeline assigns sparse IDs (range_1, range_50,
range_29...) and the prune pass drops entries but keeps original IDs.
This leaves 2394 bodies with non-sequential IDs across mountain_ranges,
rivers, and oceans.
Renumbered all feature IDs to sequential {prefix}_0, {prefix}_1, ...
preserving sort order. 24021 IDs fixed across 2394 bodies. No name
or geometry data changed — only the id field.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Retry rejection lines (dedup, blocklist, placeholder, stem_cap, empty,
error) now print unconditionally, not only under --verbose. The
fallback line also includes a tally of the rejection reasons that
exhausted all attempts, e.g.:
fallback: GJ144e-1/range_43 → 'Kirkwood Spine' [blocklist=2 dedup=1]
Diagnostic run on 20 bodies confirms dedup is the primary fallback
driver. Gemma converges on a narrow set of range names ("The Ridge",
"Blackwood Range", "The Spine") that collide across bodies in the
same corridor. Blocklist catches "Thames" and "The Great Divide"
correctly. Zero stem-cap or subprocess-error fallbacks observed.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The gfx1201 preflight check piped `strings` into grep, which fails
silently on a Bazzite host where binutils is not installed and
`strings` is not on PATH. `grep -a` reads the binary directly as
text, works everywhere grep exists, and produces the same result.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Wraps gemma_naming.py with the validated overnight recipe: gfx1201
ROCm binary path, distrobox reach-build for libhipblas at runtime,
timestamped log under .tmp/.
Preflight checks: binary exists and is executable, model present,
reach-build container exists, binary strings contains gfx1201 kernels.
Fails fast on any missing prerequisite so a broken build can't waste
an overnight window. Script takes no arguments; anything passed is
rejected so a stray --help can't accidentally launch the pipeline.
Estimate ~4-6 h for ~26k features across 2394 bodies at 74 t/s on an
RX 9070. Safe to interrupt and resume — preserved path skips
already-named bodies.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Three additions unlocked by the gfx1201 ROCm debug session.
1. _find_sr_voice() resolves the default binary path to
~/Projects/settled-reach/binaries/sr-voice-rocm (persistent across
worktree lifetimes) with a legacy fallback to the main workdir's
cargo target dir. Matches #850's plan to ship platform binaries
outside the repo.
2. --distrobox <name> wraps the sr-voice subprocess in
`distrobox enter <name> --` when the built binary depends on libs
that only exist inside a dev container (libhipblas.so.2 on a
Bazzite host). Stdio JSONL protocol flows through unchanged.
3. --dump-prompts PATH captures the attempt-0 prompt for every
feature as JSONL without calling an LLM. Force --mock and
short-circuit name_feature to return a unique deterministic
placeholder. Used to feed the same prompt set to alternate
backends (Haiku agent, other models) for offline A/B comparison
of naming quality independent of the sampling backend.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
LlamaModelParams::default() sets n_gpu_layers=0, so even with --features
rocm the model ran entirely on CPU at ~19 t/s. Setting n_gpu_layers to a
large sentinel value asks llama.cpp to offload every layer the model
has; llama.cpp clamps to the real count (27 for Gemma 2 2B). Observed
throughput jumps from 19 t/s to 74 t/s on an RX 9070 once the ROCm
binary is also compiled for gfx1201 (see tooling commit).
Also adds server/sr-voice/.gitignore so locally-built binaries don't
sneak into the worktree. Release binaries ship out-of-tree per #850.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Previous runs (both the GPU-contention kill and the anglophone-only
interrupt) left 19 hop 0-1 core bodies with stale generator output in
their markers.json files. Those bodies were being skipped via the
preserved path on relaunch, which meant the gttr-context fix
(commit a5fbce4c) would never touch them — exactly the set of
high-visibility systems that benefits most from per-system cohesion.
Reset to origin/main (clean null-name state) + re-prune to the 8/6
caps. Hand-authored templates (Edict, Estrade, Vuurkloof, Lendel,
Cairnside, Røros) explicitly excluded from the reset list and
verified intact (2-4 named cities each, untouched).
After this commit only the 6 hand-authored templates have
populated names in wiki/star-systems/. The entire rest of the reach
is clean and will be freshly named by the next gemma_naming.py run
with the full gttr + cosmopolitan + grounded few-shot + rotating
pool stack.
Bodies reset:
- Ran (GJ 144): all 9 bodies
- Sirius (GJ 244A): GJ244Ab, c, e-1, e-2 (not Ad, that's Edict)
- ACB (GJ 559B): GJ559Bb
- Tau Ceti (GJ 71): GJ71b, c, d, d-1, e
Two related quality fixes observed mid-run on Sirius + ACB + Ran:
1) Cosmopolitan corridor palettes. The six corridor inflection labels
were single-culture dominant ("administrative English / Gateway-era",
"British / Australian / Irish", "Korean/Japanese/Taiwanese", etc).
Gemma 2 2B interpreted these as "produce ONLY in this register" and
every core body came out anglophone, every east_reach body came out
East Asian. The real Earth diaspora in the setting is cosmopolitan —
a British surveyor on an east_reach moon still names a river after
their aunt in Dorset. The labels now spell out the dominant register
AND explicitly invite cross-cultural variety so Gemma samples from
the full few-shot pool instead of collapsing to one culture.
2) Per-system gttr context (the big one). The gttr.md files under
wiki/star-systems/<slug>/gttr.md already carry a vivid one-sentence
characterisation of every system — "where the rules live", "forty
years old and still in the draft", "the most connected system in
the Reach", "grandparents owned the land". This is a far stronger
cultural signal than the corridor inflection alone.
New column `star_systems.gttr_hook` stores a pre-extracted 45-word
hook per system. `tooling/db/populate_gttr_hook.py` parses each
gttr.md, regex-matches the first `**NAME**` paragraph, normalises
whitespace, truncates softly at a word cap, and stores it. Covers
all 301 systems (full coverage). Idempotent, safe to re-run after
any wiki update. Explicit transaction wrapper.
gemma_naming.py loads the hook cache at startup via
`load_system_gttr_hooks` and threads `system_hook` plus the system
and body proper names through process_body → name_feature →
_build_prompt. The prompt now carries:
System: <proper_name>. Planet: <body_name>.
About the system: <gttr_hook>
Style: British. Answer: Cooper's Creek
Style: Dutch. Answer: Meijer Beek
...
Real-mode smoke on 10 cases across 4 contrasting systems shows the
hook is doing exactly what it should. Sample output on the same
body_id / local_id pairs:
Tau Ceti (cosmopolitan hub) → Oakham River, Riverwood, Bridle Way
Ran (old-family agricultural) → Hart's Well, Blackwood Ridge
ACB (Lattice Commission seat) → Greenhaven, Rudge Brook
Posto Avançado (PT frontier dead-end) → Rio Preto, Serra de Caxias, Cunha's Cove
Posto Avançado went from "likely-English under the old corridor-only
prompt" to actual Portuguese names with a real Brazilian place stem
(Caxias), because the hook explicitly mentions wave_5 Portuguese
founders and frontier dead-end context. The gttr cultural one-liner
is the single strongest lever available for per-system cohesion —
this was the mono-culture issue observed in the first run, now fixed.
Token cost: ~60-90 extra tokens per prompt (hook + ident line).
Inference slowdown: ~5-10% per call. Acceptable for the quality gain.
Also restores 10 markers.json files that were stale from the aborted
run just killed — they were all core bodies at hop 0-1 which benefit
most from the gttr-context upgrade, so re-running them with the new
prompt is worth the ~3 minutes of re-inference.
Two fixes from observing the first Gemma 2 batch run on Sirius:
1) Prune oversized feature counts. The upstream terrain pipeline emits
every distinct mountain cluster as a separate `mountain_range` and
every flowing path as a separate `river`. At the atlas generator's
512×256 grid this produced bodies with 40-80 named ranges and
10-15 rivers — noise, not information. A single planet with 48
ridges isn't richer, it's unparseable.
`tooling/planet-gen/prune_atlas_features.py` walks every
`markers.json` under `wiki/star-systems/`, ranks each feature type
by a size proxy, and keeps only the top N:
- mountain_ranges: sorted by `area_cells`, top 8 per body
- rivers: sorted by path length, top 6 per body
- oceans / cities / pois: untouched (already small, or
hand-authored by generate_atlas.py)
Sol (GJ-0) is hardcoded-excluded from pruning so the hand-authored
Earth / Mars / moon content stays untouched.
Each pruned body gets its atlas_* rows re-synced via
`sync_markers_to_db` so the DB mirror stays consistent. Bodies
whose wiki folder has no matching row in `bodies` (14 pre-existing
orphans like GJ1156h-1, GJ34Ah-2, …) are pruned in-file but skip
the DB sync to avoid FK violations on atlas_body_grids.
First run results:
bodies scanned: 2394
bodies pruned: 1513
mountain ranges dropped: 11640
rivers dropped: 1382
Safe to re-run — idempotent when a body is already within the caps.
2) Grounded cosmopolitan fallback palette. When Gemma's 3 retries
all fail (dedup, blocklist, stem-cap, placeholder), the code falls
to `_FALLBACK_STEMS[corridor]`. The old table had 10 stems per
corridor, all Latin-institutional (Meridian, Concord, Prefecture,
Cardinal, Lumen, Foro, Tabula, Vox, Axis, Senatus), which produced
the same-y `Axis Spine / Axis Ridge / Axis Heights / Axis Scarp`
clusters the user flagged on Sirius — exactly the old epic-Latin
register the few-shot pools were rewritten to avoid.
Fallbacks now draw from a 30-45 stem grounded cosmopolitan list
per corridor matching the few-shot pool intent:
- core: 45 stems (Ashfield, Bellview, Cedarbrook,
Fairmont, Ironwood, Kirkwood, Linden, Meridian,
Northfield, Riverside, Westbrook, …)
- north_reach: 40 stems (Ashford, Bellfield, Clifford, Drayton,
Elmhurst, Garner, Holmwood, Kelsworth, …)
- west_reach: 35 stems (Altdorf, Bergfjord, Eikhof, Hoogland,
Järvenpää, Kloosterdam, Nieuwpoort, Sørholm,
Svarteberg, Torsfell, Voorhout, Weserhof, Östby, …)
- east_reach: 35 stems (Aomori, Baektu, Chōshi, Fukagawa,
Hanyang, Izumi, Takamine, Yurigawa, …)
- south_reach: 36 stems (Alves, Brandão, Évora, Gomes, Ribeiro,
Serra, Várzea, Hlanganani, Kilimi, …)
- deep_frontier: 30 stems (Okafor, Stenner, Weller, Kellogg,
Stonebrook, Dustgate, Blackwater, …)
Per-feature suffix lists also expanded (e.g. river suffixes now
include Brook, Stream, Flow, Creek on top of the original Run /
Water / Beck / Rill / Course). Net effect: 300-450 unique fallback
combinations per (corridor, feature_type), up from 50, in the same
grounded register the few-shot pools teach.
Also preserves aliases `inner_corridor`, `inner_orbit`, and
`sol-gateway-axis` as legacy-compatible keys pointing at the
administrative-English palette.
Combined effect on the next run:
- ~45% fewer features to name (pruned 13k/52k)
- ~9× more fallback variety per corridor when fallback does trigger
- Same grounding overhaul from the previous commit, now reaching
into the safety-net path
Substantial quality pass on gemma_naming.py driven by user review of
the first real-mode smoke test output. The earlier run produced names
that read too sci-fi / epic-fantasy / same-y: Aureus, Aetheria,
Stellaris, Nexus, Elysium. Root cause analysis + fixes:
1. Runtime timestamps. The log prefix is now
`[HH:MM:SS +00h03m]` — clock time plus elapsed-since-start. Gives
the user an at-a-glance sense of how long the run has been going
without scrolling back to the banner.
2. System / body headers. When the loop enters a new system it prints
`── SYSTEM K/N GJ 71 — Tau Ceti (hop 0)`. Each body line now
shows `GJ71c (Threshold)` if the body has a proper_name in
systems.db, so the log reads like a tour of the reach rather than
a wall of body_id slugs. Preserved (already-named) bodies now log
a compact "(skip — N names already set)" line so progress is
visible even when no inference happened.
3. Prompt grounding overhaul. The old few-shot examples were all
classical/epic (Wolcott Beck, Nakamura Stream, Ribeiro do Sal,
Drayton Spine) which biased Gemma 2 2B toward Latin/Greek
coinages. New preambles use the shape:
"Settlers named X after themselves, after what they saw, or
after places back home. Most names are mundane, short, and
direct — a surname, a compass direction, a feature, a
practical description. Classical or epic names are rare."
Combined with grounded example pools, Gemma now produces names
like "Cooper's Creek", "Western Ridge", "The Highroad",
"Blackwood Creek", "Dustbowl".
4. Core corridor relabel. The "core" palette inflection was
"institutional Latin / pan-Anglo / Gateway-era", which pattern-
matched in Gemma's training data to "make up Latin-sounding
words" (→ Ardenia, Aurelia, Stellaris). Now it's
"administrative English / Gateway-era" and the outputs are
prosaic — Port Dundas, East Ridge, Meridian, Landing.
5. Rotating few-shot example pools. Each feature type now has 5-7
pools of 5-6 examples each. `_build_prompt()` picks a pool
deterministically per (body_id, local_id, attempt) so:
- Same feature always gets the same prompt (determinism preserved).
- Neighbouring features on the same body get different prompts
(output variance — the sampler doesn't collapse to a single
mode when you ask for 16 mountain names in a row).
- Retries rotate to a new pool, not just a bumped seed, giving
dedup failures a clean second attempt.
6. Cosmopolitan cultural variety in the examples. Earlier pools only
showed British/Australian, Korean/Japanese, Portuguese/Swahili,
German/Dutch/Nordic axes — the four reach corridors. Gemma learned
"names come in four flavours". New pools span Dutch, Nordic,
Italian, French, Polish, Hungarian, Czech, Spanish, Russian,
Finnish, Greek, Irish, Japanese, and British — teaching the model
that names can be any real Earth cultural register, not just the
corridor label. The result: actual Dutch names (Egelantier,
Hochland, Van Damhoeve), actual Nordic (Lundstad, Brygga),
actual Italian (Borgo Marconi, Piazza Nuova), etc.
7. First-name possessive pools. Per user feedback, settler naming
includes both surnames ("Cooper's Creek") and first names
("Clifford's Bay", "Maura's Run", "Yuki's Pool"). Each feature
type now has a dedicated first-name-possessive pool in addition
to the existing surname pool — the two rotate alongside so both
patterns show up without either dominating.
8. One "classical/Latinate" pool per feature type (≈17% of calls
given 5-7 pools per type). Keeps occasional Latin flavour without
making it dominant — the user explicitly noted that replacing
one pattern with another "is never a clean fix for a randomizer."
9. Earth-name blocklist expansion. The Gemma 2 model reached for
real European names ("Weser", "Rhine", "Reykjavik") in the first
real run. Added 21 European rivers (Rhine, Weser, Elbe, Oder,
Vistula, Loire, Rhône, Douro, Tagus, Ebro, Po, Arno, Tiber, …)
and 25 Nordic/Eastern European cities (Reykjavik, Oslo, Gdansk,
Krakow, Prague, Warsaw, Budapest, Belgrade, …). Case-insensitive
"The <name>" stripping still applies so "The Great Divide" also
matches "Great Divide".
Combined smoke test after these changes (10 real-mode prompts across
core + west_reach):
- core: Port Dundas, The Backbone, Dustbowl, Blackwood Creek
- west_reach: Egelantier, Hochland, Der Rücken, Lundstad, Klipfjord
- no placeholder residue, no markdown, no 5+ word outputs.
--shard is gone (dead code since GPU contention killed parallelism).
Resume semantics are still free: re-run the same command and
already-named bodies skip via the preserved path.