Files
settled-reach/tooling/db/backfill_cultural_corridor.py
T
jpmschweitzer 1efbbcaea9 feat(tooling): gemma_naming.py batch naming pipeline for atlas (#833)
New end-to-end pipeline that walks every markers.json in the reach and
fills empty `name` fields using the Gemma 2 voice pipeline via
`sr-voice serve --stdio`. Per D-191 §4: the same Gemma 2 pipeline the
client uses for NPC voicing also produces the atlas content, which is
dual-purposed as a quality test of the LLM plumbing.

Pipeline per body (hop-ordered, core-first):
  1. Load markers.json; identify feature records whose `name` is
     blank (null or ""). Hand-authored names are never overwritten;
     the 6 template bodies and any partial authoring stay put.
  2. Look up body context (planet_class, settlement_pattern,
     cultural_corridor, population, economic_role) from systems.db.
  3. Build a short corridor-aware few-shot prompt per feature type.
     Prompts carry 3 concrete `Style: X.   Answer: Y` examples so
     Gemma 2 2B completes a pattern instead of generating to an
     open-ended instruction — this is the single biggest lever
     against placeholder echoes on a small model.
  4. Stream the prompt into a long-lived sr-voice subprocess, read
     the JSONL response, post-process (strip markdown, label
     prefixes, brackets, reject 5+ word outputs and placeholder
     tokens), check the earth-name blocklist, check per-(corridor,
     feature_type) + per-body dedup, check the per-stem cap, retry
     up to 3 times with a bumped seed.
  5. On persistent failure, fall back to a deterministic palette
     generator so every feature ends up with a name.
  6. Write markers.json atomically and refresh atlas_* DB rows via
     sync_markers_to_db. Commit the DB per body so a crash loses
     at most one body of state.
  7. Restart the sr-voice subprocess every `--refresh` requests
     (default: 200) to prevent KV-cache context bleed.

Core design decisions:
- Determinism: per-(world_seed, body_id, feature_local_id, attempt)
  seed so the full run is reproducible.
- Ordering: bodies are processed in ascending `hop_distance_from_gateway`
  so core bodies get first pick at every unique Gemma output and
  outer sectors fall into the palette fallback when they lose the
  dedup race.
- Dedup scope: (cultural_corridor, feature_type) across the run,
  PLUS a per-body cross-type set so the same name can't be a river
  AND an ocean AND a mountain on the same world. Hand-authored names
  are seeded into both sets on load so templates win priority.
- Stem cap: each non-generic root token (e.g. 'Arcturus', 'Meridian')
  may appear at most `--stem-cap` times across the full run (default
  20), preventing single-word runaway. Fallback names bypass the cap.
- Earth blocklist: 181 curated entries covering major Earth cities,
  mountains, rivers, oceans, historical/colonial spellings, and
  Greek/Roman mythology that reads too literally. Prefixed variants
  ('Nouveau Paris', 'New Tokyo') explicitly allowed per the product
  intent that Earth-echo names are fine but must not dominate.
  Leading 'The ' is stripped before comparison so 'The Great Divide'
  also matches.

Operational features:
- `--shard N/M` slices the body list into M partitions for parallel
  runs. Two terminals × `--shard 0/2` + `--shard 1/2` fits the
  ~2.5 GB/instance VRAM footprint twice under the 50% cap on a
  16 GB AMD GPU and roughly halves wall time.
- `--log PATH` writes a timestamped tee of every status line to a
  file. Default: `.tmp/gemma_naming.shard{N}of{M}.log` when a
  non-trivial shard is in use.
- SQLite `PRAGMA journal_mode=WAL` + `busy_timeout=15000` so two
  concurrent shards serialize writes without lock errors.
- Per-body progress lines report `body K/N`, `sys K/N`, and
  `hop=H` so the user can watch core sectors finish first.
- Each body logs the new names it produced per feature type so the
  user can eyeball quality as the run progresses.
- Checkpoint summary every 25 bodies: cumulative names, rate,
  ETA — gives the log regular scroll points.
- `--mock` uses `server/sr-voice/mock-stdio.sh` for dry-fire
  pipeline validation without a model load (tested end-to-end).

Supporting files:
- `tooling/planet-gen/earth_blocklist.txt` — 181 curated entries.
- `tooling/db/backfill_cultural_corridor.py` — one-off migration
  that fills the `cultural_corridor` column on both `star_systems`
  and `bodies` from the `geographic_sector` values. Before this
  pass, 99.4% of rows (3221/3240) had a NULL cultural_corridor
  despite `wiki_sync.py` being aware of the column — the wiki
  index.md files only carry the sector header, which was never
  propagated to the DB column. Idempotent, safe to re-run after
  any wiki_sync rebuild, explicit transaction wrapper with
  rollback on failure.

Full batch runtime estimate: ~20 hours single-shard / ~10 hours
double-shard on this hardware. Smoke tests across five hardened
iterations (v1–v5) on GJ71b/c/d/d-1/e confirm the pipeline produces
clean, varied, culturally-coherent names with zero post-processing
residue.
2026-04-15 10:54:28 +02:00

163 lines
5.6 KiB
Python
Executable File

#!/usr/bin/env python3
"""
Backfill star_systems.cultural_corridor and bodies.cultural_corridor.
The schema has a `cultural_corridor` column on both tables, but
`wiki_sync.py` never populated it from the wiki index.md files — the
cultural/geographic identity of each system lives in
`star_systems.geographic_sector` instead (values: core, north_reach,
south_reach, east_reach, west_reach, deep_frontier). These two fields
refer to the same concept: which arc of the reach the system belongs
to. Leaving `cultural_corridor` NULL on 99%+ of rows defeats every
downstream consumer that actually wants to filter by corridor
(gemma_naming.py, future atlas UI queries, narrative tools).
This script treats `geographic_sector` as the source of truth and
copies it into `cultural_corridor`:
star_systems.cultural_corridor := star_systems.geographic_sector
WHERE cultural_corridor IS NULL
bodies.cultural_corridor := parent star_systems.cultural_corridor
WHERE bodies.cultural_corridor IS NULL
It is safe to re-run — idempotent, NULL-only updates, explicit
transaction wrapper so a crash never leaves a half-populated state.
Run it after any `wiki_sync.py` pass that creates fresh systems.db
rows.
Usage:
tooling/db/backfill_cultural_corridor.py
tooling/db/backfill_cultural_corridor.py --db path/to/systems.db
tooling/db/backfill_cultural_corridor.py --dry-run
Decisions: D-191 (atlas pipeline — downstream consumer)
"""
import argparse
import sqlite3
import sys
from pathlib import Path
SCRIPT_DIR = Path(__file__).resolve().parent
REPO_ROOT = (SCRIPT_DIR / ".." / "..").resolve()
DB_PATH = REPO_ROOT / "server" / "data" / "systems.db"
def main():
parser = argparse.ArgumentParser(
description="Backfill cultural_corridor on star_systems and bodies"
)
parser.add_argument("--db", default=str(DB_PATH), help="Path to systems.db")
parser.add_argument(
"--dry-run",
action="store_true",
help="Report counts without writing",
)
args = parser.parse_args()
db_path = Path(args.db)
if not db_path.exists():
print(f"error: {db_path} not found", file=sys.stderr)
sys.exit(1)
conn = sqlite3.connect(str(db_path))
conn.execute("PRAGMA foreign_keys=ON")
# Counts before.
before_systems_null = conn.execute(
"SELECT COUNT(*) FROM star_systems WHERE cultural_corridor IS NULL"
).fetchone()[0]
before_bodies_null = conn.execute(
"SELECT COUNT(*) FROM bodies WHERE cultural_corridor IS NULL"
).fetchone()[0]
print(f"\n cultural_corridor backfill")
print(f" DB: {db_path}")
if args.dry_run:
print(f" Mode: DRY RUN")
print()
print(f" Before:")
print(f" star_systems.cultural_corridor NULL: {before_systems_null}")
print(f" bodies.cultural_corridor NULL: {before_bodies_null}")
conn.execute("BEGIN")
try:
# 1. Star systems — copy geographic_sector into cultural_corridor
# where the latter is still NULL. If geographic_sector is also
# NULL, leave cultural_corridor NULL — there is nothing to
# copy and a bogus placeholder is worse than honest NULL.
sys_rows_updated = conn.execute(
"""
UPDATE star_systems
SET cultural_corridor = geographic_sector
WHERE cultural_corridor IS NULL
AND geographic_sector IS NOT NULL
"""
).rowcount
# 2. Bodies — inherit from the parent star_systems row.
body_rows_updated = conn.execute(
"""
UPDATE bodies
SET cultural_corridor = (
SELECT s.cultural_corridor
FROM star_systems s
WHERE s.system_id = bodies.system_id
)
WHERE cultural_corridor IS NULL
AND EXISTS (
SELECT 1
FROM star_systems s
WHERE s.system_id = bodies.system_id
AND s.cultural_corridor IS NOT NULL
)
"""
).rowcount
if args.dry_run:
conn.rollback()
print()
print(f" Would update:")
print(f" star_systems: {sys_rows_updated}")
print(f" bodies: {body_rows_updated}")
print(f"\n Dry run — no changes written.")
else:
conn.commit()
print()
print(f" Updated:")
print(f" star_systems: {sys_rows_updated}")
print(f" bodies: {body_rows_updated}")
# Counts after.
after_systems_null = conn.execute(
"SELECT COUNT(*) FROM star_systems WHERE cultural_corridor IS NULL"
).fetchone()[0]
after_bodies_null = conn.execute(
"SELECT COUNT(*) FROM bodies WHERE cultural_corridor IS NULL"
).fetchone()[0]
print()
print(f" After:")
print(f" star_systems.cultural_corridor NULL: {after_systems_null}")
print(f" bodies.cultural_corridor NULL: {after_bodies_null}")
# Show the distribution so the outcome is visible.
print()
print(f" star_systems.cultural_corridor distribution:")
for corridor, count in conn.execute(
"SELECT cultural_corridor, COUNT(*) FROM star_systems "
"GROUP BY cultural_corridor ORDER BY COUNT(*) DESC"
).fetchall():
print(f" {corridor!r}: {count}")
except BaseException:
conn.rollback()
conn.close()
raise
conn.close()
print()
if __name__ == "__main__":
main()