New end-to-end pipeline that walks every markers.json in the reach and
fills empty `name` fields using the Gemma 2 voice pipeline via
`sr-voice serve --stdio`. Per D-191 §4: the same Gemma 2 pipeline the
client uses for NPC voicing also produces the atlas content, which is
dual-purposed as a quality test of the LLM plumbing.
Pipeline per body (hop-ordered, core-first):
1. Load markers.json; identify feature records whose `name` is
blank (null or ""). Hand-authored names are never overwritten;
the 6 template bodies and any partial authoring stay put.
2. Look up body context (planet_class, settlement_pattern,
cultural_corridor, population, economic_role) from systems.db.
3. Build a short corridor-aware few-shot prompt per feature type.
Prompts carry 3 concrete `Style: X. Answer: Y` examples so
Gemma 2 2B completes a pattern instead of generating to an
open-ended instruction — this is the single biggest lever
against placeholder echoes on a small model.
4. Stream the prompt into a long-lived sr-voice subprocess, read
the JSONL response, post-process (strip markdown, label
prefixes, brackets, reject 5+ word outputs and placeholder
tokens), check the earth-name blocklist, check per-(corridor,
feature_type) + per-body dedup, check the per-stem cap, retry
up to 3 times with a bumped seed.
5. On persistent failure, fall back to a deterministic palette
generator so every feature ends up with a name.
6. Write markers.json atomically and refresh atlas_* DB rows via
sync_markers_to_db. Commit the DB per body so a crash loses
at most one body of state.
7. Restart the sr-voice subprocess every `--refresh` requests
(default: 200) to prevent KV-cache context bleed.
Core design decisions:
- Determinism: per-(world_seed, body_id, feature_local_id, attempt)
seed so the full run is reproducible.
- Ordering: bodies are processed in ascending `hop_distance_from_gateway`
so core bodies get first pick at every unique Gemma output and
outer sectors fall into the palette fallback when they lose the
dedup race.
- Dedup scope: (cultural_corridor, feature_type) across the run,
PLUS a per-body cross-type set so the same name can't be a river
AND an ocean AND a mountain on the same world. Hand-authored names
are seeded into both sets on load so templates win priority.
- Stem cap: each non-generic root token (e.g. 'Arcturus', 'Meridian')
may appear at most `--stem-cap` times across the full run (default
20), preventing single-word runaway. Fallback names bypass the cap.
- Earth blocklist: 181 curated entries covering major Earth cities,
mountains, rivers, oceans, historical/colonial spellings, and
Greek/Roman mythology that reads too literally. Prefixed variants
('Nouveau Paris', 'New Tokyo') explicitly allowed per the product
intent that Earth-echo names are fine but must not dominate.
Leading 'The ' is stripped before comparison so 'The Great Divide'
also matches.
Operational features:
- `--shard N/M` slices the body list into M partitions for parallel
runs. Two terminals × `--shard 0/2` + `--shard 1/2` fits the
~2.5 GB/instance VRAM footprint twice under the 50% cap on a
16 GB AMD GPU and roughly halves wall time.
- `--log PATH` writes a timestamped tee of every status line to a
file. Default: `.tmp/gemma_naming.shard{N}of{M}.log` when a
non-trivial shard is in use.
- SQLite `PRAGMA journal_mode=WAL` + `busy_timeout=15000` so two
concurrent shards serialize writes without lock errors.
- Per-body progress lines report `body K/N`, `sys K/N`, and
`hop=H` so the user can watch core sectors finish first.
- Each body logs the new names it produced per feature type so the
user can eyeball quality as the run progresses.
- Checkpoint summary every 25 bodies: cumulative names, rate,
ETA — gives the log regular scroll points.
- `--mock` uses `server/sr-voice/mock-stdio.sh` for dry-fire
pipeline validation without a model load (tested end-to-end).
Supporting files:
- `tooling/planet-gen/earth_blocklist.txt` — 181 curated entries.
- `tooling/db/backfill_cultural_corridor.py` — one-off migration
that fills the `cultural_corridor` column on both `star_systems`
and `bodies` from the `geographic_sector` values. Before this
pass, 99.4% of rows (3221/3240) had a NULL cultural_corridor
despite `wiki_sync.py` being aware of the column — the wiki
index.md files only carry the sector header, which was never
propagated to the DB column. Idempotent, safe to re-run after
any wiki_sync rebuild, explicit transaction wrapper with
rollback on failure.
Full batch runtime estimate: ~20 hours single-shard / ~10 hours
double-shard on this hardware. Smoke tests across five hardened
iterations (v1–v5) on GJ71b/c/d/d-1/e confirm the pipeline produces
clean, varied, culturally-coherent names with zero post-processing
residue.
163 lines
5.6 KiB
Python
Executable File
163 lines
5.6 KiB
Python
Executable File
#!/usr/bin/env python3
|
|
"""
|
|
Backfill star_systems.cultural_corridor and bodies.cultural_corridor.
|
|
|
|
The schema has a `cultural_corridor` column on both tables, but
|
|
`wiki_sync.py` never populated it from the wiki index.md files — the
|
|
cultural/geographic identity of each system lives in
|
|
`star_systems.geographic_sector` instead (values: core, north_reach,
|
|
south_reach, east_reach, west_reach, deep_frontier). These two fields
|
|
refer to the same concept: which arc of the reach the system belongs
|
|
to. Leaving `cultural_corridor` NULL on 99%+ of rows defeats every
|
|
downstream consumer that actually wants to filter by corridor
|
|
(gemma_naming.py, future atlas UI queries, narrative tools).
|
|
|
|
This script treats `geographic_sector` as the source of truth and
|
|
copies it into `cultural_corridor`:
|
|
|
|
star_systems.cultural_corridor := star_systems.geographic_sector
|
|
WHERE cultural_corridor IS NULL
|
|
|
|
bodies.cultural_corridor := parent star_systems.cultural_corridor
|
|
WHERE bodies.cultural_corridor IS NULL
|
|
|
|
It is safe to re-run — idempotent, NULL-only updates, explicit
|
|
transaction wrapper so a crash never leaves a half-populated state.
|
|
Run it after any `wiki_sync.py` pass that creates fresh systems.db
|
|
rows.
|
|
|
|
Usage:
|
|
tooling/db/backfill_cultural_corridor.py
|
|
tooling/db/backfill_cultural_corridor.py --db path/to/systems.db
|
|
tooling/db/backfill_cultural_corridor.py --dry-run
|
|
|
|
Decisions: D-191 (atlas pipeline — downstream consumer)
|
|
"""
|
|
|
|
import argparse
|
|
import sqlite3
|
|
import sys
|
|
from pathlib import Path
|
|
|
|
SCRIPT_DIR = Path(__file__).resolve().parent
|
|
REPO_ROOT = (SCRIPT_DIR / ".." / "..").resolve()
|
|
DB_PATH = REPO_ROOT / "server" / "data" / "systems.db"
|
|
|
|
|
|
def main():
|
|
parser = argparse.ArgumentParser(
|
|
description="Backfill cultural_corridor on star_systems and bodies"
|
|
)
|
|
parser.add_argument("--db", default=str(DB_PATH), help="Path to systems.db")
|
|
parser.add_argument(
|
|
"--dry-run",
|
|
action="store_true",
|
|
help="Report counts without writing",
|
|
)
|
|
args = parser.parse_args()
|
|
|
|
db_path = Path(args.db)
|
|
if not db_path.exists():
|
|
print(f"error: {db_path} not found", file=sys.stderr)
|
|
sys.exit(1)
|
|
|
|
conn = sqlite3.connect(str(db_path))
|
|
conn.execute("PRAGMA foreign_keys=ON")
|
|
|
|
# Counts before.
|
|
before_systems_null = conn.execute(
|
|
"SELECT COUNT(*) FROM star_systems WHERE cultural_corridor IS NULL"
|
|
).fetchone()[0]
|
|
before_bodies_null = conn.execute(
|
|
"SELECT COUNT(*) FROM bodies WHERE cultural_corridor IS NULL"
|
|
).fetchone()[0]
|
|
|
|
print(f"\n cultural_corridor backfill")
|
|
print(f" DB: {db_path}")
|
|
if args.dry_run:
|
|
print(f" Mode: DRY RUN")
|
|
print()
|
|
print(f" Before:")
|
|
print(f" star_systems.cultural_corridor NULL: {before_systems_null}")
|
|
print(f" bodies.cultural_corridor NULL: {before_bodies_null}")
|
|
|
|
conn.execute("BEGIN")
|
|
try:
|
|
# 1. Star systems — copy geographic_sector into cultural_corridor
|
|
# where the latter is still NULL. If geographic_sector is also
|
|
# NULL, leave cultural_corridor NULL — there is nothing to
|
|
# copy and a bogus placeholder is worse than honest NULL.
|
|
sys_rows_updated = conn.execute(
|
|
"""
|
|
UPDATE star_systems
|
|
SET cultural_corridor = geographic_sector
|
|
WHERE cultural_corridor IS NULL
|
|
AND geographic_sector IS NOT NULL
|
|
"""
|
|
).rowcount
|
|
|
|
# 2. Bodies — inherit from the parent star_systems row.
|
|
body_rows_updated = conn.execute(
|
|
"""
|
|
UPDATE bodies
|
|
SET cultural_corridor = (
|
|
SELECT s.cultural_corridor
|
|
FROM star_systems s
|
|
WHERE s.system_id = bodies.system_id
|
|
)
|
|
WHERE cultural_corridor IS NULL
|
|
AND EXISTS (
|
|
SELECT 1
|
|
FROM star_systems s
|
|
WHERE s.system_id = bodies.system_id
|
|
AND s.cultural_corridor IS NOT NULL
|
|
)
|
|
"""
|
|
).rowcount
|
|
|
|
if args.dry_run:
|
|
conn.rollback()
|
|
print()
|
|
print(f" Would update:")
|
|
print(f" star_systems: {sys_rows_updated}")
|
|
print(f" bodies: {body_rows_updated}")
|
|
print(f"\n Dry run — no changes written.")
|
|
else:
|
|
conn.commit()
|
|
print()
|
|
print(f" Updated:")
|
|
print(f" star_systems: {sys_rows_updated}")
|
|
print(f" bodies: {body_rows_updated}")
|
|
|
|
# Counts after.
|
|
after_systems_null = conn.execute(
|
|
"SELECT COUNT(*) FROM star_systems WHERE cultural_corridor IS NULL"
|
|
).fetchone()[0]
|
|
after_bodies_null = conn.execute(
|
|
"SELECT COUNT(*) FROM bodies WHERE cultural_corridor IS NULL"
|
|
).fetchone()[0]
|
|
print()
|
|
print(f" After:")
|
|
print(f" star_systems.cultural_corridor NULL: {after_systems_null}")
|
|
print(f" bodies.cultural_corridor NULL: {after_bodies_null}")
|
|
|
|
# Show the distribution so the outcome is visible.
|
|
print()
|
|
print(f" star_systems.cultural_corridor distribution:")
|
|
for corridor, count in conn.execute(
|
|
"SELECT cultural_corridor, COUNT(*) FROM star_systems "
|
|
"GROUP BY cultural_corridor ORDER BY COUNT(*) DESC"
|
|
).fetchall():
|
|
print(f" {corridor!r}: {count}")
|
|
except BaseException:
|
|
conn.rollback()
|
|
conn.close()
|
|
raise
|
|
|
|
conn.close()
|
|
print()
|
|
|
|
|
|
if __name__ == "__main__":
|
|
main()
|