Files
settled-reach/tooling/economy-db/import_economics.py
T
jpmschweitzerandClaude Opus 4.8 21793c1a94 data(economics): #1016 cultural-write pass + canonical 47-value heritage migration
Completes the deferred cultural half of #1016 and migrates the heritage
vocabulary to the canonical 47-value taxonomy (D-237).

Migration (mechanical):
- _CULTURAL_HERITAGE replaced with the canonical 47-value set
  (docs/.../heritage-taxonomy-draft.md). Renamed the 4 shipped hero
  values to canonical: afrikaans_cape->afrikaans, french_provencal->
  french, italian_northern->italian, norse_compact->nordic.

Content pass (authored from GTTR/name heritage triage, 6-sector fan-out):
- +139 new cultural_specialization pins and 1 change (Altmark
  cosmopolitan->financial_technocratic), taking coverage 32 -> 171
  across 54 distinct values. Each pin is grounded in an explicit GTTR
  founding-community statement or system/feature name etymology, applied
  only where founding heritage DIVERGES from the corridor baseline
  (D-167/D-232); corridor-typical systems left NULL.

Distribution is now balanced and diverse — portuguese 15, german 12,
vietnamese 10, afrikaans 10 (no longer dominant post-#1019), then a long
tail filling the previously-missing slots (welsh, herero, czech, akan,
hungarian, konkan, afro_brazilian, cape_verdean, shona, arab/persian/
turkic, etc.).

Judgment calls (flagged for review):
- Kruger 60 + Pedra Seca -> xhosa (mixed SA founding; adds diversity).
  Pedra Seca has a Portuguese name but Afrikaans/Xhosa GTTR founding —
  name/heritage mismatch noted for a future content fix.
- Kampala Gate / Moyale / Mwangaza -> swahili as the nearest token for
  East-African-interior heritage; taxonomy may want a dedicated value.
- Kept hero pins Kensho=scholarly, Keid=scholarly, Nyrheim=nordic.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 10:06:23 +02:00

2276 lines
93 KiB
Python
Executable File
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
#!/usr/bin/env python3
"""
Import economics data into systems.db.
Reads TOML/JSON source files and populates the economics tables:
- gate_links from docs/design/star-map.json (335 edges, bidirectional)
- commodities from wiki/economics/commodities.toml (36 types)
- production_chains + chain_inputs from wiki/economics/production_chains.toml
- currency_zone on star_systems (default TRACTUS_PRIMARY)
- gate_energy_connected on star_systems (D-186: false for MARK_PRIMARY zones)
- corporations from wiki/corporations/*.md (sync + insert new records)
- corp_presence from wiki/corporations/*.md (headquarters location data)
Validation (hard errors, non-zero exit on any failure):
- Wiki corporation names must match DB proper_name records (D-182 sync constraint)
- Chain completeness: every intermediate commodity has at least one production chain
- Commodity coverage: 3+ corporations per major commodity type (D-175)
- System coverage: 1+ corporation per inhabited system with population > 100K (D-175)
Usage:
python3 tooling/economy-db/import_economics.py
python3 tooling/economy-db/import_economics.py --dry-run
python3 tooling/economy-db/import_economics.py --db path/to/systems.db
"""
import argparse
import glob
import hashlib
import json
import re
import sqlite3
import sys
import tomllib
from pathlib import Path
REPO_ROOT = Path(__file__).resolve().parent.parent.parent
# Import shared schema version constant (#888) — single source of truth in tooling/schema_version.py
sys.path.insert(0, str(REPO_ROOT / "tooling"))
from schema_version import SCHEMA_VERSION # noqa: E402
DB_PATH = REPO_ROOT / "server" / "data" / "systems.db"
STAR_MAP = REPO_ROOT / "docs" / "design" / "star-map.json"
COMMODITIES_TOML = REPO_ROOT / "wiki" / "economics" / "commodities.toml"
CHAINS_TOML = REPO_ROOT / "wiki" / "economics" / "production_chains.toml"
SPECIALIZATION_VOCAB_TOML = REPO_ROOT / "wiki" / "economics" / "specialization_vocabulary.toml"
SYSTEM_SPECIALIZATION_TOML = REPO_ROOT / "wiki" / "economics" / "system_specialization.toml"
SCHEMA_SQL = REPO_ROOT / "server" / "data" / "systems-schema.sql"
CORPORATIONS_DIR = REPO_ROOT / "wiki" / "corporations"
WIKI_STAR_SYSTEMS = REPO_ROOT / "wiki" / "star-systems"
BRANDS_TOML = REPO_ROOT / "wiki" / "economics" / "corporations" / "brands.toml"
GENERATED_BRANDS_TOML = REPO_ROOT / "wiki" / "economics" / "corporations" / "generated_brands.toml"
# Rust sources for the generate_brands subroutine. import_economics shells out to
# tooling/generate-brands as part of its normal flow (see regenerate_brands()), so
# both Rust files contribute to this script's effective source SHA: any change to
# either must invalidate the meta stamp even though Python hasn't changed.
GENERATE_BRANDS_RS = REPO_ROOT / "server" / "src" / "bin" / "generate_brands" / "main.rs"
GENERATE_BRANDS_NAMES_RS = REPO_ROOT / "server" / "src" / "bin" / "generate_brands" / "names.rs"
GENERATE_BRANDS_WRAPPER = REPO_ROOT / "tooling" / "generate-brands"
def _file_sha1(*paths: Path) -> str:
"""Return SHA-1 hex of the concatenated content of one or more files.
Files are sorted by path for determinism. Missing files raise FileNotFoundError
rather than silently contributing an empty-string hash — a ghost SHA masks real
breakage (review comment H2: da39a3ee… convergence could produce vacuous passes).
"""
h = hashlib.sha1()
for p in sorted(paths):
if not p.exists():
raise FileNotFoundError(f"generator source not found: {p}")
h.update(p.read_bytes())
return h.hexdigest()
# Canonical source set for import_economics' meta stamp. Covers its own .py file
# plus the Rust binary it invokes (generate_brands main.rs + names.rs + wrapper
# script) so any change to the brand generation pipeline flips the stamp. Keep
# this list in sync with GENERATOR_SOURCES["import_economics"] in
# tooling/check-systems-db-stamp.
IMPORT_ECONOMICS_SOURCES: tuple[Path, ...] = (
Path(__file__),
GENERATE_BRANDS_RS,
GENERATE_BRANDS_NAMES_RS,
GENERATE_BRANDS_WRAPPER,
REPO_ROOT / "tooling" / "schema_version.py",
# D-237 authored specialization layer: these data TOMLs feed the DB, so a
# change to either must flip the stamp and force a regen (#1013). Mirror in
# GENERATOR_SOURCES["import_economics"] in tooling/check-systems-db-stamp.
SPECIALIZATION_VOCAB_TOML,
SYSTEM_SPECIALIZATION_TOML,
)
def _write_stamp(conn: sqlite3.Connection, generator_name: str, *source_files: Path) -> None:
"""Upsert a row in the meta table recording this generator's current source SHA.
Called after every successful non-dry-run commit. Idempotent: running
twice on the same sources writes the same sha with an updated timestamp.
Only one stamp is written by this module: ``import_economics``, whose source
set includes the Rust binary it invokes (see IMPORT_ECONOMICS_SOURCES).
generate_atlas writes its own stamp. generate_brands does NOT write a stamp
of its own — it's a subroutine of import_economics, not an independent DB
writer (PR #136 review T2/H3).
The meta table is created by the MIGRATION_SQL block above; this
function assumes it exists (caller must run migrations first).
"""
schema_sha = _file_sha1(SCHEMA_SQL)
generator_sha = _file_sha1(*source_files)
conn.execute(
"""INSERT OR REPLACE INTO meta
(generator_name, schema_version, schema_sha, generator_sha, generated_at)
VALUES (?, ?, ?, ?, datetime('now'))""",
(generator_name, SCHEMA_VERSION, schema_sha, generator_sha),
)
def regenerate_brands() -> None:
"""Run the Rust generate_brands binary to refresh generated_brands.toml.
Invoked as the first step of import_economics' main flow so the TOML on disk
always matches the current Rust source before the Python import reads it.
This replaces the former split (tooling/generate-brands run separately by
make regen-db) with a single, coherent brand pipeline owned by one stamp.
The wrapper script builds the binary on demand and runs it with the default
canonical seed=1; callers that need non-canonical seeds must still invoke
the wrapper directly (experimentation only — committed output must be seed=1).
"""
import subprocess
if not GENERATE_BRANDS_WRAPPER.exists():
raise FileNotFoundError(
f"generate_brands wrapper not found at {GENERATE_BRANDS_WRAPPER}"
)
print(" [pre/10] Running generate_brands (Rust) to refresh generated_brands.toml...")
result = subprocess.run(
[str(GENERATE_BRANDS_WRAPPER)],
cwd=str(REPO_ROOT),
capture_output=True,
text=True,
)
if result.returncode != 0:
print(result.stdout, file=sys.stderr)
print(result.stderr, file=sys.stderr)
raise _ImportAborted()
# Print the Rust binary's own summary lines (brands generated, coverage).
# Indent so they fold under the pre-step heading.
for line in result.stdout.splitlines():
if line.strip():
print(f" {line}")
class _ImportAborted(Exception):
"""Raised internally by main() when a validation step wants a clean
rollback + exit 1. Caught only by main(); error messages are printed
before raising so the user sees them."""
# ---------------------------------------------------------------------------
# Schema migration — add new tables and columns to existing DB
# ---------------------------------------------------------------------------
MIGRATION_SQL = """
-- Economics tables (idempotent — safe to re-run)
-- Brand layer tables (D-189, #827)
CREATE TABLE IF NOT EXISTS brand_products (
brand_product_id TEXT PRIMARY KEY,
corp_id TEXT NOT NULL REFERENCES corporations(corp_id),
product_name TEXT NOT NULL,
brand_category TEXT NOT NULL,
value_trajectory TEXT NOT NULL,
scarcity_class TEXT NOT NULL,
product_subcategory TEXT,
base_premium_multiplier REAL NOT NULL DEFAULT 1.0,
premium_floor REAL NOT NULL DEFAULT 0.0,
origin_system TEXT REFERENCES star_systems(system_id),
terroir_locked INTEGER NOT NULL DEFAULT 0,
currency_denomination TEXT NOT NULL DEFAULT 'tractus',
shadow_viable INTEGER NOT NULL DEFAULT 0,
brand_tier TEXT NOT NULL,
halo_brand_id TEXT REFERENCES brand_products(brand_product_id),
price_tier TEXT,
updated_at TEXT DEFAULT (datetime('now'))
);
CREATE TABLE IF NOT EXISTS brand_inputs (
brand_product_id TEXT NOT NULL REFERENCES brand_products(brand_product_id),
commodity_id TEXT NOT NULL REFERENCES commodities(commodity_id),
quantity REAL NOT NULL,
PRIMARY KEY (brand_product_id, commodity_id)
);
CREATE TABLE IF NOT EXISTS system_fiscal (
system_id TEXT PRIMARY KEY REFERENCES star_systems(system_id),
corp_tax_rate REAL NOT NULL DEFAULT 0.22,
collection_efficiency REAL NOT NULL DEFAULT 1.0,
updated_at TEXT DEFAULT (datetime('now'))
);
CREATE TABLE IF NOT EXISTS corp_financial_state (
corp_id TEXT PRIMARY KEY REFERENCES corporations(corp_id),
health_metric REAL NOT NULL DEFAULT 1.0,
updated_at TEXT DEFAULT (datetime('now'))
);
CREATE TABLE IF NOT EXISTS corp_lifecycle_events (
event_id INTEGER PRIMARY KEY AUTOINCREMENT,
corp_id TEXT NOT NULL REFERENCES corporations(corp_id),
event_type TEXT NOT NULL,
event_tick INTEGER NOT NULL DEFAULT 0,
event_data TEXT,
created_at TEXT DEFAULT (datetime('now'))
);
CREATE INDEX IF NOT EXISTS idx_brand_products_corp_category ON brand_products(corp_id, brand_category);
CREATE INDEX IF NOT EXISTS idx_brand_products_origin ON brand_products(origin_system);
CREATE INDEX IF NOT EXISTS idx_brand_products_tier ON brand_products(brand_tier);
CREATE INDEX IF NOT EXISTS idx_brand_inputs_commodity ON brand_inputs(commodity_id);
CREATE INDEX IF NOT EXISTS idx_corp_lifecycle_events_corp ON corp_lifecycle_events(corp_id);
CREATE TABLE IF NOT EXISTS gate_links (
from_system_id TEXT NOT NULL REFERENCES star_systems(system_id),
to_system_id TEXT NOT NULL REFERENCES star_systems(system_id),
PRIMARY KEY (from_system_id, to_system_id)
);
CREATE TABLE IF NOT EXISTS commodities (
commodity_id TEXT PRIMARY KEY,
name TEXT NOT NULL,
tier TEXT NOT NULL,
elasticity TEXT NOT NULL,
base_price REAL NOT NULL,
bulk_class TEXT,
unit TEXT,
production_ubiquity TEXT,
demand_model TEXT,
commission_certifiable INTEGER DEFAULT 0,
compact_contested INTEGER DEFAULT 0,
shadow_viable INTEGER DEFAULT 0,
panic_threshold_weeks INTEGER DEFAULT 0,
description TEXT,
updated_at TEXT DEFAULT (datetime('now'))
);
-- D-237 authored specialization layer vocabulary (must follow commodities for FK).
-- Mirrors the canonical DDL in systems-schema.sql; here so the migration path
-- (existing DBs) gets the table, not just fresh systems-schema.sql builds.
CREATE TABLE IF NOT EXISTS specialization_vocabulary (
specialization_id TEXT PRIMARY KEY,
commodity_id TEXT NOT NULL REFERENCES commodities(commodity_id),
production_ubiquity_override TEXT,
bulk_class_projected TEXT NOT NULL,
production_ubiquity_projected TEXT NOT NULL,
description TEXT NOT NULL
);
CREATE INDEX IF NOT EXISTS idx_spec_vocab_commodity ON specialization_vocabulary(commodity_id);
CREATE TABLE IF NOT EXISTS production_chains (
chain_id TEXT PRIMARY KEY,
output_commodity_id TEXT NOT NULL REFERENCES commodities(commodity_id),
output_quantity REAL NOT NULL DEFAULT 1.0,
location_bound INTEGER DEFAULT 0,
description TEXT,
updated_at TEXT DEFAULT (datetime('now'))
);
CREATE TABLE IF NOT EXISTS chain_inputs (
chain_id TEXT NOT NULL REFERENCES production_chains(chain_id),
input_commodity_id TEXT NOT NULL REFERENCES commodities(commodity_id),
quantity REAL NOT NULL,
PRIMARY KEY (chain_id, input_commodity_id)
);
CREATE TABLE IF NOT EXISTS corp_presence (
corp_id TEXT NOT NULL REFERENCES corporations(corp_id),
location_id TEXT NOT NULL,
location_type TEXT NOT NULL,
primary_operation TEXT,
updated_at TEXT DEFAULT (datetime('now')),
PRIMARY KEY (corp_id, location_id)
);
-- New indexes
CREATE INDEX IF NOT EXISTS idx_star_systems_currency ON star_systems(currency_zone);
CREATE INDEX IF NOT EXISTS idx_gate_links_from ON gate_links(from_system_id);
CREATE INDEX IF NOT EXISTS idx_gate_links_to ON gate_links(to_system_id);
CREATE INDEX IF NOT EXISTS idx_commodities_tier ON commodities(tier);
CREATE INDEX IF NOT EXISTS idx_production_chains_output ON production_chains(output_commodity_id);
CREATE INDEX IF NOT EXISTS idx_chain_inputs_commodity ON chain_inputs(input_commodity_id);
CREATE INDEX IF NOT EXISTS idx_corp_presence_corp ON corp_presence(corp_id);
CREATE INDEX IF NOT EXISTS idx_corp_presence_location ON corp_presence(location_id);
-- Generator metadata stamp (#855, #856)
CREATE TABLE IF NOT EXISTS meta (
generator_name TEXT PRIMARY KEY,
schema_version TEXT NOT NULL,
generator_sha TEXT NOT NULL,
generated_at TEXT NOT NULL DEFAULT (datetime('now'))
);
-- Drop the pre-merge 'generate_brands' stamp row if it exists (PR #136 review T2/H3).
-- The Rust brand binary is now a subroutine of import_economics — its source
-- SHA contributes to the 'import_economics' stamp — so it no longer merits its
-- own meta row. This DELETE makes the check-systems-db-stamp "unknown generator"
-- path (fail-closed per T6) compatible with older DBs that still have the row.
DELETE FROM meta WHERE generator_name = 'generate_brands';
-- Drop the retired 'generate_atlas' stamp row if it exists (#951, D-223).
-- The atlas geometry generator was retired; import_economics now owns the
-- atlas index, so generate_atlas no longer merits its own meta row. Without
-- this DELETE, check-systems-db-stamp's fail-closed "unknown generator" path
-- (T6) would reject any committed DB that still carries the old row.
DELETE FROM meta WHERE generator_name = 'generate_atlas';
-- Drop the retired heightmap BLOB table (D-202 amended, #963): canonical
-- elevation is now a per-body 16-bit grayscale heightmap.png file, not a DB
-- BLOB. The Rust loader reads the PNG; nothing reads this table anymore.
DROP TABLE IF EXISTS atlas_body_heightmaps;
-- City name reservations (D-207, #902)
CREATE TABLE IF NOT EXISTS atlas_city_names (
id INTEGER PRIMARY KEY AUTOINCREMENT,
body_id TEXT NOT NULL REFERENCES bodies(body_id) ON DELETE CASCADE,
name TEXT NOT NULL,
kind TEXT NOT NULL DEFAULT 'city',
economic_role TEXT NOT NULL,
population INTEGER NOT NULL,
settlement_class TEXT,
corp_id TEXT REFERENCES corporations(corp_id),
reserved INTEGER NOT NULL DEFAULT 0,
updated_at TEXT NOT NULL DEFAULT (datetime('now'))
);
CREATE INDEX IF NOT EXISTS idx_atlas_city_names_body ON atlas_city_names(body_id);
CREATE INDEX IF NOT EXISTS idx_atlas_city_names_kind ON atlas_city_names(kind);
CREATE INDEX IF NOT EXISTS idx_atlas_city_names_corp ON atlas_city_names(corp_id);
-- Geographic feature name reservations (#903)
CREATE TABLE IF NOT EXISTS atlas_feature_names (
id INTEGER PRIMARY KEY AUTOINCREMENT,
body_id TEXT NOT NULL REFERENCES bodies(body_id) ON DELETE CASCADE,
name TEXT NOT NULL,
feature_type TEXT NOT NULL,
priority INTEGER NOT NULL DEFAULT 0,
updated_at TEXT NOT NULL DEFAULT (datetime('now'))
);
CREATE INDEX IF NOT EXISTS idx_atlas_feature_names_body ON atlas_feature_names(body_id);
-- Province boundaries (D-205, #904)
CREATE TABLE IF NOT EXISTS atlas_province_boundaries (
body_id TEXT NOT NULL REFERENCES bodies(body_id) ON DELETE CASCADE,
basin_id INTEGER NOT NULL,
path TEXT NOT NULL,
area_pct REAL NOT NULL,
PRIMARY KEY (body_id, basin_id)
);
CREATE INDEX IF NOT EXISTS idx_atlas_province_boundaries_body ON atlas_province_boundaries(body_id);
-- City positions — attractor-matched placement output (D-211, #34)
CREATE TABLE IF NOT EXISTS atlas_city_positions (
city_names_id INTEGER PRIMARY KEY REFERENCES atlas_city_names(id) ON DELETE CASCADE,
body_id TEXT NOT NULL REFERENCES bodies(body_id) ON DELETE CASCADE,
row INTEGER NOT NULL,
col INTEGER NOT NULL,
attractor_type TEXT NOT NULL,
score REAL NOT NULL
);
CREATE INDEX IF NOT EXISTS idx_atlas_city_positions_body ON atlas_city_positions(body_id);
-- Normalize bodies.economic_role to the D-194 canonical 10-value set (#911).
-- Idempotent: each UPDATE is a no-op if the old value is already gone.
UPDATE bodies SET economic_role = 'agricultural' WHERE economic_role IN ('agriculture', 'mixed-agriculture');
UPDATE bodies SET economic_role = 'extraction' WHERE economic_role IN ('mining', 'resource_extraction', 'energy');
UPDATE bodies SET economic_role = 'transit_hub' WHERE economic_role = 'transit';
UPDATE bodies SET economic_role = 'service_mixed' WHERE economic_role IN ('commercial', 'coordination');
UPDATE bodies SET economic_role = 'residential' WHERE economic_role = 'frontier';
-- Backfill bodies.founding_age_years for all inhabited bodies in player scope (D-216 amendment,
-- ticket #1000). Idempotent: WHERE clause limits to NULL rows, so a re-run is a no-op.
--
-- Strategy: COALESCE(events-first, wave-fallback)
-- events-first: the system's colonial_charter event age_years (authored; 9 systems in live DB).
-- system_history is keyed PRIMARY KEY on system_id, so the wave subquery returns
-- at most one row; LIMIT 1 is used on historical_events for safety (max one
-- colonial_charter per system in the data).
-- wave-fallback: canonical founding-edge of each settlement_wave range per D-216:
-- wave_1=600, wave_2=500, wave_3=300, wave_4=100, wave_5=40.
--
-- Exclusions (founding_age_years stays NULL):
-- origin — Sol; out of player scope per D-236.
-- unsettled — 24 systems with inhabited=0; the EXISTS guard also excludes them.
UPDATE bodies
SET founding_age_years = COALESCE(
(
SELECT he.age_years
FROM historical_events he
WHERE he.system_id = bodies.system_id
AND he.event_type = 'colonial_charter'
ORDER BY he.sort_order ASC
LIMIT 1
),
(
SELECT CASE sh.settlement_wave
WHEN 'wave_1' THEN 600
WHEN 'wave_2' THEN 500
WHEN 'wave_3' THEN 300
WHEN 'wave_4' THEN 100
WHEN 'wave_5' THEN 40
ELSE NULL -- unexpected settlement_wave: add a WHEN above
END
FROM system_history sh
WHERE sh.system_id = bodies.system_id
)
)
WHERE founding_age_years IS NULL
AND inhabited = 1
AND EXISTS (
SELECT 1
FROM system_history sh2
WHERE sh2.system_id = bodies.system_id
AND sh2.settlement_wave NOT IN ('origin', 'unsettled')
);
"""
# Columns to add to existing tables (ALTER TABLE is idempotent via try/except)
COLUMN_MIGRATIONS = [
("star_systems", "currency_zone", "TEXT DEFAULT 'TRACTUS_PRIMARY'"),
("star_systems", "gate_energy_connected", "INTEGER DEFAULT 1"),
("corporations", "behavioral_archetype", "TEXT"),
("corporations", "supply_chain_role", "TEXT"),
("corporations", "shadow_economy_access", "INTEGER DEFAULT 0"),
("brand_products", "price_tier", "TEXT"),
("bodies", "body_radius_km", "REAL"), # D-204 — physical radius in km, nullable
("meta", "schema_sha", "TEXT"),
("atlas_city_names", "settlement_class", "TEXT"), # D-196 — NULL until placement (#37)
("system_economy", "economic_specialization", "TEXT"), # D-237 — authored specialization layer
("system_economy", "cultural_specialization", "TEXT"), # D-237 — authored specialization layer
]
def _add_column(conn: sqlite3.Connection, table: str, col: str, col_type: str):
"""Add a column if it doesn't exist. SQLite has no IF NOT EXISTS for ALTER."""
try:
conn.execute(f"ALTER TABLE {table} ADD COLUMN {col} {col_type}")
except sqlite3.OperationalError as e:
if "duplicate column" in str(e).lower():
pass # already exists
else:
raise
# ---------------------------------------------------------------------------
# Gate links
# ---------------------------------------------------------------------------
def import_gate_links(conn: sqlite3.Connection, dry_run: bool) -> int:
with open(STAR_MAP) as f:
data = json.load(f)
edges = data["edges"]
system_ids = {r[0] for r in conn.execute("SELECT system_id FROM star_systems").fetchall()}
rows = []
skipped = []
for a, b in edges:
if a not in system_ids:
skipped.append(a)
continue
if b not in system_ids:
skipped.append(b)
continue
rows.append((a, b))
rows.append((b, a))
if skipped:
unique_skipped = sorted(set(skipped))
print(f" warning: {len(unique_skipped)} system(s) in star-map.json not in DB: {unique_skipped[:5]}...")
if not dry_run:
conn.executemany(
"INSERT OR IGNORE INTO gate_links (from_system_id, to_system_id) VALUES (?, ?)",
rows,
)
return len(rows)
# ---------------------------------------------------------------------------
# Commodities
# ---------------------------------------------------------------------------
def import_commodities(conn: sqlite3.Connection, dry_run: bool) -> int:
with open(COMMODITIES_TOML, "rb") as f:
data = tomllib.load(f)
rows = []
for cid, c in data.items():
rows.append((
cid,
c["name"],
c["tier"],
c["elasticity"],
c["base_price"],
c.get("bulk_class"),
c.get("unit"),
c.get("production_ubiquity"),
c.get("demand_model"),
int(c.get("commission_certifiable", False)),
int(c.get("compact_contested", False)),
int(c.get("shadow_viable", False)),
c.get("panic_threshold_weeks", 0),
c.get("description"),
))
if not dry_run:
conn.executemany(
"""INSERT INTO commodities (
commodity_id, name, tier, elasticity, base_price,
bulk_class, unit, production_ubiquity, demand_model,
commission_certifiable, compact_contested, shadow_viable,
panic_threshold_weeks, description
) VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)""",
rows,
)
return len(rows)
# ---------------------------------------------------------------------------
# Production chains + inputs
# ---------------------------------------------------------------------------
def import_chains(conn: sqlite3.Connection, dry_run: bool) -> tuple[int, int]:
with open(CHAINS_TOML, "rb") as f:
data = tomllib.load(f)
chain_rows = []
input_rows = []
for chain_id, c in data.items():
chain_rows.append((
chain_id,
c["output"],
c.get("output_quantity", 1.0),
int(c.get("location_bound", False)),
c.get("description"),
))
for inp in c.get("inputs", []):
input_rows.append((
chain_id,
inp["commodity"],
inp["quantity"],
))
if not dry_run:
conn.executemany(
"""INSERT INTO production_chains (
chain_id, output_commodity_id, output_quantity,
location_bound, description
) VALUES (?, ?, ?, ?, ?)""",
chain_rows,
)
conn.executemany(
"""INSERT INTO chain_inputs (
chain_id, input_commodity_id, quantity
) VALUES (?, ?, ?)""",
input_rows,
)
return len(chain_rows), len(input_rows)
# ---------------------------------------------------------------------------
# Currency zones
# ---------------------------------------------------------------------------
def set_currency_zones(conn: sqlite3.Connection, dry_run: bool) -> dict:
"""Set currency_zone on star_systems from wiki/economics/currency_zones.toml.
Default: TRACTUS_PRIMARY. Sol (GJ 0): MIXED (set before file is read).
MARK_PRIMARY and MIXED assignments come from the TOML file (D-172).
"""
if dry_run:
return {"TRACTUS_PRIMARY": "all", "MIXED": "GJ 0 + toml"}
# Default everything to TRACTUS_PRIMARY
conn.execute("UPDATE star_systems SET currency_zone = 'TRACTUS_PRIMARY'")
# Sol system is MIXED (Earth legacy currency presence — set before TOML load)
conn.execute("UPDATE star_systems SET currency_zone = 'MIXED' WHERE system_id = 'GJ 0'")
# Load MARK_PRIMARY and MIXED assignments from authored TOML (D-172)
zones_path = REPO_ROOT / "wiki" / "economics" / "currency_zones.toml"
if zones_path.exists():
import tomllib # Python 3.11+
with open(zones_path, "rb") as f:
zones = tomllib.load(f)
mark_ids = [entry["system_id"] for entry in zones.get("mark_primary", [])]
mixed_ids = [entry["system_id"] for entry in zones.get("mixed", [])]
for sid in mark_ids:
conn.execute(
"UPDATE star_systems SET currency_zone = 'MARK_PRIMARY' WHERE system_id = ?",
(sid,),
)
for sid in mixed_ids:
conn.execute(
"UPDATE star_systems SET currency_zone = 'MIXED' WHERE system_id = ?",
(sid,),
)
else:
print(" warning: wiki/economics/currency_zones.toml not found — "
"all systems default to TRACTUS_PRIMARY / Sol to MIXED")
counts = {}
for row in conn.execute("SELECT currency_zone, COUNT(*) FROM star_systems GROUP BY currency_zone"):
counts[row[0]] = row[1]
return counts
# ---------------------------------------------------------------------------
# Gate energy connectivity (D-186)
# ---------------------------------------------------------------------------
def set_gate_energy(conn: sqlite3.Connection, dry_run: bool) -> dict:
"""Set gate_energy_connected on star_systems based on currency_zone.
MARK_PRIMARY zones default to false (Compact refused Gate Corp dependency).
All other zones default to true.
"""
if dry_run:
return {"on_grid": "non-MARK_PRIMARY", "off_grid": "MARK_PRIMARY"}
# Default: all systems on-grid
conn.execute("UPDATE star_systems SET gate_energy_connected = 1 WHERE gate_energy_connected IS NULL")
# MARK_PRIMARY zones are off-grid (Compact energy sovereignty)
conn.execute("UPDATE star_systems SET gate_energy_connected = 0 WHERE currency_zone = 'MARK_PRIMARY'")
counts = {}
for row in conn.execute(
"SELECT gate_energy_connected, COUNT(*) FROM star_systems GROUP BY gate_energy_connected"
):
label = "on_grid" if row[0] == 1 else "off_grid"
counts[label] = row[1]
return counts
# ---------------------------------------------------------------------------
# Corporation wiki parsing
# ---------------------------------------------------------------------------
def _parse_corp_frontmatter(path: Path) -> dict | None:
"""Parse YAML frontmatter from a wiki corporation markdown file."""
text = path.read_text()
lines = text.split("\n")
if not lines or lines[0].strip() != "---":
return None
end_idx = None
for i, line in enumerate(lines[1:], 1):
if line.strip() == "---":
end_idx = i
break
if end_idx is None:
return None
fm: dict = {}
for line in lines[1:end_idx]:
if ":" not in line:
continue
key, _, val = line.partition(":")
key = key.strip()
val = val.strip()
if val.startswith("[") and val.endswith("]"):
items = [x.strip().strip('"').strip("'") for x in val[1:-1].split(",")]
fm[key] = [item for item in items if item]
else:
fm[key] = val.strip('"').strip("'")
return fm
def load_wiki_corps() -> list[dict]:
"""Load all wiki corporation files. Returns list of parsed corp records."""
corps = []
for md_file in sorted(CORPORATIONS_DIR.glob("*.md")):
if md_file.name == "index.md":
continue
fm = _parse_corp_frontmatter(md_file)
if not fm or not fm.get("slug") or not fm.get("title"):
continue
hq = fm.get("headquarters", "")
m = re.search(r"\(([^)]+)\)", hq)
system_id = m.group(1) if m else None
corps.append({
"corp_id": fm["slug"],
"proper_name": fm["title"],
"system_id": system_id,
"tags": fm.get("tags", []),
"scope": fm.get("scope", ""),
})
return corps
# ---------------------------------------------------------------------------
# Corporation sync (D-182: wiki is source of truth)
# ---------------------------------------------------------------------------
def sync_corporations(
conn: sqlite3.Connection, wiki_corps: list[dict], dry_run: bool
) -> list[str]:
"""Sync wiki corps to DB. Hard error on proper_name divergence (D-182).
Returns list of error strings. Inserts corps that exist in wiki but not DB.
Corps that exist only in DB (legacy records) are left untouched.
headquarters_system is only written if the system_id exists in star_systems
(to avoid FK violations when atlas hasn't yet registered the system).
"""
errors: list[str] = []
existing = {
r[0]: r[1]
for r in conn.execute("SELECT corp_id, proper_name FROM corporations").fetchall()
}
valid_systems = {
r[0] for r in conn.execute("SELECT system_id FROM star_systems").fetchall()
}
to_insert = []
for corp in wiki_corps:
corp_id = corp["corp_id"]
proper_name = corp["proper_name"]
if corp_id in existing:
if existing[corp_id] != proper_name:
errors.append(
f"name divergence: corp_id='{corp_id}' "
f"wiki='{proper_name}' db='{existing[corp_id]}'"
)
else:
system_id = corp.get("system_id")
hq_system = system_id if system_id and system_id in valid_systems else None
if system_id and system_id not in valid_systems:
print(f" warning: {corp_id} HQ system '{system_id}' not in DB, "
f"headquarters_system set to NULL")
to_insert.append((
corp_id,
proper_name,
"corporation",
corp.get("scope") or None,
hq_system,
))
if not dry_run and not errors:
conn.executemany(
"""INSERT OR IGNORE INTO corporations
(corp_id, proper_name, corp_type, scope, headquarters_system)
VALUES (?, ?, ?, ?, ?)""",
to_insert,
)
return errors
# ---------------------------------------------------------------------------
# Corp presence population
# ---------------------------------------------------------------------------
def _resolve_hq_location(
conn: sqlite3.Connection,
system_id: str,
headquarters_body: str | None,
) -> tuple[str, str] | None:
"""Resolve a corp's HQ to a (location_id, location_type) pair.
Resolution order:
1. Use headquarters_body from corporations table if set (body or station).
2. Most-populated body in the system.
3. Any body in the system.
4. Any station in the system.
Returns None if no body or station found.
"""
if headquarters_body:
# Determine whether it's a body or station
body = conn.execute(
"SELECT body_id FROM bodies WHERE body_id = ?", (headquarters_body,)
).fetchone()
if body:
return (headquarters_body, "body")
station = conn.execute(
"SELECT station_id FROM stations WHERE station_id = ?",
(headquarters_body,),
).fetchone()
if station:
return (headquarters_body, "station")
# Most-populated body
body = conn.execute(
"""SELECT body_id FROM bodies WHERE system_id = ?
ORDER BY population DESC LIMIT 1""",
(system_id,),
).fetchone()
if body:
return (body[0], "body")
# Any station
station = conn.execute(
"SELECT station_id FROM stations WHERE system_id = ? LIMIT 1",
(system_id,),
).fetchone()
if station:
return (station[0], "station")
return None
def import_corp_presence(
conn: sqlite3.Connection,
wiki_corps: list[dict],
commodity_ids: set[str],
dry_run: bool,
) -> int:
"""Populate corp_presence from wiki headquarters data.
Each corporation gets one presence row at its headquarters body or station.
location_type is 'body' or 'station' per schema (D-182).
primary_operation is set to the first commodity tag matching a known commodity ID.
"""
valid_systems = {
r[0] for r in conn.execute("SELECT system_id FROM star_systems").fetchall()
}
# Load headquarters_body from corporations table (set during import)
hq_body_map: dict[str, str | None] = {
r[0]: r[1]
for r in conn.execute(
"SELECT corp_id, headquarters_body FROM corporations"
).fetchall()
}
rows = []
skipped = []
for corp in wiki_corps:
system_id = corp.get("system_id")
if not system_id:
skipped.append(f"{corp['corp_id']} (no headquarters system parsed)")
continue
if system_id not in valid_systems:
skipped.append(f"{corp['corp_id']} (system '{system_id}' not in DB)")
continue
hq_body = hq_body_map.get(corp["corp_id"])
location = _resolve_hq_location(conn, system_id, hq_body)
if not location:
skipped.append(
f"{corp['corp_id']} (no body/station found in system '{system_id}')"
)
continue
location_id, location_type = location
primary_op = next(
(tag for tag in corp.get("tags", []) if tag in commodity_ids), None
)
rows.append((corp["corp_id"], location_id, location_type, primary_op))
if skipped:
for s in skipped:
print(f" warning: skipped corp_presence for {s}")
if not dry_run:
conn.execute("DELETE FROM corp_presence")
conn.executemany(
"""INSERT OR IGNORE INTO corp_presence
(corp_id, location_id, location_type, primary_operation)
VALUES (?, ?, ?, ?)""",
rows,
)
return len(rows)
# ---------------------------------------------------------------------------
# Validation
# ---------------------------------------------------------------------------
def validate(conn: sqlite3.Connection) -> list[str]:
"""Validate structural integrity of imported data.
Checks FK integrity, chain commodity references, and chain completeness.
These are hard blockers — broken data must not be committed.
Coverage validation (commodity/system thresholds) is separate and runs
after commit via _validate_commodity_coverage() and _validate_system_coverage().
"""
errors = []
# FK integrity
fk_issues = conn.execute("PRAGMA foreign_key_check").fetchall()
if fk_issues:
for issue in fk_issues[:10]:
errors.append(f"FK violation: table={issue[0]} rowid={issue[1]} "
f"parent={issue[2]} fkid={issue[3]}")
# Chain inputs reference valid commodities
orphan_inputs = conn.execute("""
SELECT ci.chain_id, ci.input_commodity_id
FROM chain_inputs ci
LEFT JOIN commodities c ON ci.input_commodity_id = c.commodity_id
WHERE c.commodity_id IS NULL
""").fetchall()
for chain_id, cid in orphan_inputs:
errors.append(f"chain_inputs: chain '{chain_id}' references unknown commodity '{cid}'")
# Chain outputs reference valid commodities
orphan_outputs = conn.execute("""
SELECT pc.chain_id, pc.output_commodity_id
FROM production_chains pc
LEFT JOIN commodities c ON pc.output_commodity_id = c.commodity_id
WHERE c.commodity_id IS NULL
""").fetchall()
for chain_id, cid in orphan_outputs:
errors.append(f"production_chains: chain '{chain_id}' outputs unknown commodity '{cid}'")
# economic_role must be one of the D-194 canonical 10 values
valid_roles = {
'manufacturing', 'financial', 'agricultural', 'extraction',
'service_mixed', 'institutional', 'transit_hub', 'research',
'military', 'residential',
}
bad_roles = conn.execute("""
SELECT DISTINCT economic_role, COUNT(*) as cnt
FROM bodies
WHERE economic_role IS NOT NULL
AND economic_role NOT IN (
'manufacturing', 'financial', 'agricultural', 'extraction',
'service_mixed', 'institutional', 'transit_hub', 'research',
'military', 'residential'
)
GROUP BY economic_role
""").fetchall()
for role, cnt in bad_roles:
errors.append(
f"bodies.economic_role: non-canonical value '{role}' on {cnt} row(s) — "
f"valid values: {sorted(valid_roles)}"
)
# Chain completeness: every intermediate commodity must have at least one producer
missing_chains = conn.execute("""
SELECT c.commodity_id, c.name
FROM commodities c
WHERE c.tier = 'intermediate'
AND c.commodity_id NOT IN (SELECT output_commodity_id FROM production_chains)
ORDER BY c.commodity_id
""").fetchall()
for cid, name in missing_chains:
errors.append(f"chain completeness: no production chain produces intermediate '{cid}' ({name})")
return errors
def _validate_commodity_coverage(
conn: sqlite3.Connection, wiki_corps: list[dict], commodity_ids: set[str]
) -> list[str]:
"""3+ corporations per major commodity type (raw + intermediate). D-175."""
errors: list[str] = []
major = [
r[0]
for r in conn.execute(
"SELECT commodity_id FROM commodities "
"WHERE tier IN ('raw', 'intermediate') ORDER BY commodity_id"
).fetchall()
]
# Build commodity → corp set from wiki tags filtered to known commodity IDs
coverage: dict[str, set[str]] = {cid: set() for cid in major}
for corp in wiki_corps:
for tag in corp.get("tags", []):
if tag in coverage:
coverage[tag].add(corp["corp_id"])
for cid in major:
n = len(coverage[cid])
if n < 3:
corp_list = sorted(coverage[cid]) if coverage[cid] else ["none"]
errors.append(
f"commodity coverage: '{cid}' has {n}/3 corp(s) — {corp_list}"
)
return errors
def _validate_system_coverage(
conn: sqlite3.Connection, wiki_corps: list[dict]
) -> list[str]:
"""1+ corporation per inhabited system with population > 100K. D-175.
Uses wiki_corps headquarters data (not DB corp_presence) so this check
is accurate in both dry-run and real-run modes.
"""
covered = {c["system_id"] for c in wiki_corps if c.get("system_id")}
populated = conn.execute("""
SELECT se.system_id, ss.proper_name, se.population
FROM system_economy se
JOIN star_systems ss ON se.system_id = ss.system_id
WHERE se.population > 100000
ORDER BY se.system_id
""").fetchall()
return [
f"system coverage: no corp presence in '{sid}' ({name}, pop={pop:,})"
for sid, name, pop in populated
if sid not in covered
]
# ---------------------------------------------------------------------------
# Brand layer import (D-189, #827)
# ---------------------------------------------------------------------------
VALID_BRAND_CATEGORIES = {
"terroir", "heritage_craft", "tech_premium", "cultural",
"service_premium", "commodity_branded", "design_heritage", "platform_catalogue",
}
VALID_VALUE_TRAJECTORIES = {"appreciating", "depreciating", "timeless"}
VALID_SCARCITY_CLASSES = {"capped", "constrained", "scalable", "unlimited"}
VALID_BRAND_TIERS = {"halo", "volume"}
VALID_CURRENCY_DENOMINATIONS = {"tractus", "mark", "mixed", "sol_adjacent"}
VALID_PRICE_TIERS = {"mass", "premium", "luxury", "flagship", "institutional"}
def _load_brand_file(path) -> tuple[list, list]:
"""Load brand_products and brand_inputs from a TOML file. Returns empty lists if missing."""
if not path.exists():
return [], []
with open(path, "rb") as f:
data = tomllib.load(f)
return data.get("brand_products", []), data.get("brand_inputs", [])
def import_brands(
conn: sqlite3.Connection, dry_run: bool
) -> tuple[int, int]:
"""Import brand_products and brand_inputs from brands.toml and generated_brands.toml.
Hand-authored brands (brands.toml) are imported first; generated brands
(generated_brands.toml, produced by `tooling/generate-brands`) are merged in.
Returns (n_products, n_inputs).
"""
if not BRANDS_TOML.exists():
print(" warning: brands.toml not found — brand layer skipped")
return 0, 0
products_authored, inputs_authored = _load_brand_file(BRANDS_TOML)
products_generated, inputs_generated = _load_brand_file(GENERATED_BRANDS_TOML)
if products_generated:
print(f" merging {len(products_generated)} generated brand_products from generated_brands.toml")
products = products_authored + products_generated
inputs = inputs_authored + inputs_generated
product_rows = []
for p in products:
product_rows.append((
p["brand_product_id"],
p["corp_id"],
p["product_name"],
p["brand_category"],
p["value_trajectory"],
p["scarcity_class"],
p.get("product_subcategory"),
p.get("base_premium_multiplier", 1.0),
p.get("premium_floor", 0.0),
p.get("origin_system"),
int(p.get("terroir_locked", False)),
p.get("currency_denomination", "tractus"),
int(p.get("shadow_viable", False)),
p["brand_tier"],
p.get("halo_brand_id"),
p.get("price_tier"),
))
input_rows = []
for inp in inputs:
input_rows.append((
inp["brand_product_id"],
inp["commodity_id"],
inp["quantity"],
))
if not dry_run:
conn.executemany(
"""INSERT OR REPLACE INTO brand_products (
brand_product_id, corp_id, product_name, brand_category,
value_trajectory, scarcity_class, product_subcategory,
base_premium_multiplier, premium_floor, origin_system,
terroir_locked, currency_denomination, shadow_viable,
brand_tier, halo_brand_id, price_tier
) VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)""",
product_rows,
)
conn.executemany(
"""INSERT OR REPLACE INTO brand_inputs
(brand_product_id, commodity_id, quantity) VALUES (?, ?, ?)""",
input_rows,
)
return len(product_rows), len(input_rows)
def import_system_fiscal(conn: sqlite3.Connection, dry_run: bool) -> int:
"""Populate system_fiscal with hardcoded Phase 2 values.
Phase 2 values (NOT derived from D-189 §6 yet):
- corp_tax_rate = 0.22 (flat default)
- collection_efficiency = 0.85 (mid-reach average placeholder)
The D-189 §6 formula `collection_efficiency = 1.0 - shadow_economy_intensity × 0.6`
is deliberately NOT implemented here — `shadow_economy_intensity` is not
yet per-system in the DB (pending the shadow_economy.toml pipeline). When
that pipeline lands, replace the hardcoded 0.85 with the derivation and
wire `shadow_economy_intensity` through the SELECT. Tracked as a Phase 3
follow-up.
"""
inhabited = conn.execute("""
SELECT ss.system_id, COALESCE(se.population, 0)
FROM star_systems ss
LEFT JOIN system_economy se ON ss.system_id = se.system_id
WHERE ss.inhabited_planet_count > 0 OR se.population > 0
ORDER BY ss.system_id
""").fetchall()
PHASE2_CORP_TAX_RATE = 0.22
PHASE2_COLLECTION_EFFICIENCY = 0.85
rows = [
(system_id, PHASE2_CORP_TAX_RATE, PHASE2_COLLECTION_EFFICIENCY)
for system_id, _pop in inhabited
]
if not dry_run:
conn.executemany(
"""INSERT OR IGNORE INTO system_fiscal
(system_id, corp_tax_rate, collection_efficiency) VALUES (?, ?, ?)""",
rows,
)
return len(rows)
def populate_body_radius_km(conn: sqlite3.Connection, dry_run: bool) -> int:
"""Populate body_radius_km column from planet_class fallback (D-204, #910).
Applies the fallback lookup table to rows where body_radius_km IS NULL.
Does not overwrite rows where body_radius_km is already set (authoritative data).
Fallback values (km):
super_earth -> 8000
earth_like -> 6371
earth -> 6371 (alternate spelling)
sub_earth -> 4500
ocean_world -> 6500
arid -> 5800
frozen -> 4500
ice_world -> 3000
barren -> 4500
volcanic -> 5500
gas_giant -> 0 (no settlements, skip)
moon -> 1737
other/unknown -> 6371 (Earth default)
"""
PLANET_CLASS_RADIUS = {
"super_earth": 8000.0,
"earth_like": 6371.0,
"earth": 6371.0,
"sub_earth": 3500.0,
"ocean_world": 6500.0,
"arid": 5800.0,
"frozen": 3500.0,
"ice_world": 3000.0,
"barren": 3500.0,
"volcanic": 5500.0,
"temperate": 6371.0,
"moon": 1737.0,
}
DEFAULT_RADIUS = 6371.0
SKIP_RADIUS_TYPES = {"oort_cloud", "asteroid_belt"}
GAS_GIANT_RADIUS = {
"gas_giant": 50000.0,
"ice_giant": 25000.0,
}
GAS_GIANT_DEFAULT = 45000.0
GAS_GIANT_SCATTER = 0.20 # ±20%
rows = conn.execute(
"SELECT body_id, planet_class, body_type, mass_class FROM bodies WHERE body_radius_km IS NULL"
).fetchall()
import hashlib
SCATTER_FRACTION = 0.15 # ±15% for rocky bodies
updates = []
for body_id, planet_class, body_type, mass_class in rows:
if body_type in SKIP_RADIUS_TYPES:
continue
h = int(hashlib.sha256(body_id.encode()).hexdigest()[:8], 16)
scatter_val = (h / 0xFFFFFFFF) * 2.0 - 1.0 # [-1.0, 1.0]
if body_type == "gas_giant":
mc = (mass_class or "").lower()
base_radius = GAS_GIANT_RADIUS.get(mc, GAS_GIANT_DEFAULT)
radius = round(base_radius * (1.0 + scatter_val * GAS_GIANT_SCATTER), 1)
elif body_type == "moon":
base_radius = 1400.0
scatter_frac = 0.86 # ±86% → ~1962604 km
radius = round(base_radius * (1.0 + scatter_val * scatter_frac), 1)
else:
base_radius = PLANET_CLASS_RADIUS.get(
(planet_class or "").lower(), DEFAULT_RADIUS
)
radius = round(base_radius * (1.0 + scatter_val * SCATTER_FRACTION), 1)
updates.append((radius, body_id))
if not dry_run and updates:
conn.executemany(
"UPDATE bodies SET body_radius_km = ? WHERE body_id = ?", updates
)
return len(updates)
# Atlas geometry index tables (D-191). These hold computed positions — city
# centres, road/river/rail polylines, ocean/mountain extents. Under D-223 the
# Python atlas geometry generator was retired (#951); the deterministic
# server-side cascade (Phase 4) is the sole producer of this geometry. We keep
# the tables (the Atlas viewer #960 and the cascade read them) but empty them on
# every regen so the committed DB carries no stale prototype geometry — the
# empty tables are the gap the server cascade fills.
_ATLAS_GEOMETRY_TABLES = (
"atlas_cities",
"atlas_roads",
"atlas_railroads",
"atlas_pois",
"atlas_rivers",
"atlas_oceans",
"atlas_mountain_ranges",
"atlas_body_grids",
)
_ATLAS_INDEX_BEGIN_MARKER = "-- BEGIN ATLAS INDEX"
_ATLAS_INDEX_END_MARKER = "-- END ATLAS INDEX"
def ensure_atlas_index_schema(conn: sqlite3.Connection, dry_run: bool) -> None:
"""Apply the canonical atlas_* DDL and empty the geometry tables (D-223, #951).
systems-schema.sql is the single source of truth for the atlas index tables
(the BEGIN/END ATLAS INDEX block). The retired generate_atlas.py used to
apply this block; import_economics now owns it, since it is the only
regen-db generator that touches systems.db's atlas tables. The block is all
CREATE ... IF NOT EXISTS, so applying it on the committed DB is a no-op and
on a fresh DB it creates the geometry tables.
After ensuring the schema, the geometry tables are cleared: their geometry
now comes from the server cascade, not from authored markers (D-223).
"""
text = SCHEMA_SQL.read_text()
try:
start = text.index(_ATLAS_INDEX_BEGIN_MARKER)
end = text.index(_ATLAS_INDEX_END_MARKER, start)
except ValueError as e:
raise RuntimeError(
f"systems-schema.sql is missing the {_ATLAS_INDEX_BEGIN_MARKER}/"
f"{_ATLAS_INDEX_END_MARKER} block — has the schema been restructured?"
) from e
conn.executescript(text[start:end])
if not dry_run:
for table in _ATLAS_GEOMETRY_TABLES:
conn.execute(f"DELETE FROM {table}")
def populate_atlas_city_names(conn: sqlite3.Connection, dry_run: bool) -> int:
"""Populate atlas_city_names from the names-only markers.json pool (D-223, #951).
Scans wiki/star-systems/*/bodies/*/markers.json for the flavoured city
name pool at `names.cities` and inserts one atlas_city_names row per name:
- body_id : directory name (e.g. GJ0e)
- name : pooled city name
- kind : 'city' — capital is chosen at placement (#955)
- economic_role : inherited from bodies.economic_role; fallback 'mixed'
- population : 0 — assigned by the server cascade at placement (#955)
- corp_id : NULL — populated by populate_atlas_city_names_corps (#909)
- reserved : 0
markers.json is a names-only flavoured pool (D-223): it carries no geometry
or population. The deterministic server cascade attaches these names to
computed settlements and assigns population/kind/position at placement time;
this importer just loads the pool.
Deterministic rebuild: clears atlas_city_names first (the FK cascade clears
atlas_city_positions), so re-runs are idempotent — there is no UNIQUE on
(body_id, name), so without the clear a re-run would accumulate duplicates.
Skips body directories not found in the bodies table (missing FK).
"""
# Build body_id -> economic_role map
body_roles: dict[str, str] = {}
for body_id, role in conn.execute(
"SELECT body_id, economic_role FROM bodies"
).fetchall():
body_roles[body_id] = role or "mixed"
valid_body_ids: set[str] = set(body_roles.keys())
# Sol (system 'GJ 0') is permanently exempt from the normal generators
# (D-223, #951): its bodies use real Earth/Mars/Luna geography via
# sol_import.py and keep geometry-bearing markers.json as preserved config.
# Sol names come from its own (future) scripted integration, not the names
# pool — skip Sol bodies here regardless of their markers format.
sol_body_ids: set[str] = {
r[0] for r in conn.execute(
"SELECT body_id FROM bodies WHERE system_id = 'GJ 0'"
).fetchall()
}
rows: list[tuple] = []
skipped_bodies: list[str] = []
pattern = str(WIKI_STAR_SYSTEMS / "*" / "bodies" / "*" / "markers.json")
for markers_path in sorted(glob.glob(pattern)):
body_id = markers_path.split("/bodies/")[1].split("/")[0]
if body_id not in valid_body_ids or body_id in sol_body_ids:
if body_id not in valid_body_ids:
skipped_bodies.append(body_id)
continue
with open(markers_path) as fh:
data = json.load(fh)
names_pool = (data.get("names") or {}).get("cities") or []
economic_role = body_roles[body_id]
for raw_name in names_pool:
name = (raw_name or "").strip()
if not name:
continue
# kind defaults to 'city'; population 0 until placement (#955).
rows.append((body_id, name, "city", economic_role, 0))
if skipped_bodies:
unique = sorted(set(skipped_bodies))
print(f" warning: {len(unique)} body dirs not in DB — skipped: {unique[:5]}")
if not dry_run:
conn.execute("DELETE FROM atlas_city_names")
if rows:
conn.executemany(
"""INSERT INTO atlas_city_names
(body_id, name, kind, economic_role, population)
VALUES (?, ?, ?, ?, ?)""",
rows,
)
return len(rows)
def populate_atlas_city_names_corps(conn: sqlite3.Connection, dry_run: bool) -> tuple[int, int]:
"""Cross-reference corp HQ city names into atlas_city_names (D-207, #909).
For each corporation with a parseable headquarters field ("City (SYSTEM_ID)"):
- If atlas_city_names already has a row with matching name on a body in that
system: UPDATE the row to set corp_id.
- Otherwise: INSERT a reserved row (reserved=1) so the name is protected.
Attaches to the most-populated body in the system (fallback: any body).
Returns (n_updated, n_inserted).
"""
# Build system_id -> sorted bodies (by population desc, then body_id)
sys_bodies: dict[str, list[tuple[int, str, str]]] = {}
for body_id, sys_id, pop, role in conn.execute(
"SELECT body_id, system_id, COALESCE(population, 0), COALESCE(economic_role, 'mixed') FROM bodies"
).fetchall():
sys_bodies.setdefault(sys_id, []).append((pop, body_id, role))
for v in sys_bodies.values():
v.sort(key=lambda x: (-x[0], x[1]))
# Build (body_id, name_lower) -> id index for existing atlas_city_names rows
existing: dict[tuple[str, str], int] = {}
body_to_sys: dict[str, str] = {
r[0]: r[1]
for r in conn.execute("SELECT body_id, system_id FROM bodies").fetchall()
}
for row_id, body_id, name in conn.execute(
"SELECT id, body_id, name FROM atlas_city_names"
).fetchall():
existing[(body_id, name.lower())] = row_id
# Build system_id -> set of body_ids for quick lookup
sys_body_ids: dict[str, set[str]] = {}
for body_id, sys_id in body_to_sys.items():
sys_body_ids.setdefault(sys_id, set()).add(body_id)
updated: list[tuple[str, int]] = [] # (corp_id, atlas_row_id)
inserted: list[tuple] = [] # insert rows
for corp_id, headquarters_system in conn.execute(
"SELECT corp_id, headquarters_system FROM corporations WHERE headquarters_system IS NOT NULL"
).fetchall():
# Retrieve original headquarters string from wiki to get city name
md_file = CORPORATIONS_DIR / f"{corp_id}.md"
if not md_file.exists():
continue
hq_raw = ""
with open(md_file) as f:
in_fm = False
for line in f:
if line.strip() == "---":
if not in_fm:
in_fm = True
continue
else:
break
if in_fm and line.startswith("headquarters:"):
hq_raw = line.split(":", 1)[1].strip().strip('"')
break
if not hq_raw:
continue
m = re.search(r"\(([^)]+)\)", hq_raw)
city_name = hq_raw[: m.start()].strip() if m else hq_raw.strip()
if not city_name:
continue
# Try to find a matching atlas_city_names row in the same system
body_ids_in_sys = sys_body_ids.get(headquarters_system, set())
match_id: int | None = None
# sorted() for determinism: on a name collision across bodies in the
# same system, set iteration order is not stable (D-010 #4).
for body_id in sorted(body_ids_in_sys):
key = (body_id, city_name.lower())
if key in existing:
match_id = existing[key]
break
if match_id is not None:
updated.append((corp_id, match_id))
else:
# Sol (system 'GJ 0') is exempt from the normal generators (D-223,
# #951) — do not synthesize a reserved corp-HQ row on a Sol body;
# Sol's atlas data comes from its own scripted integration.
if headquarters_system == "GJ 0":
continue
# Insert a reserved row on the most-populated body in the system
candidates = sys_bodies.get(headquarters_system, [])
if not candidates:
continue
_, target_body_id, body_role = candidates[0]
inserted.append((target_body_id, city_name, "city", body_role, 0, corp_id, 1))
if not dry_run:
for corp_id, row_id in updated:
conn.execute(
"UPDATE atlas_city_names SET corp_id = ? WHERE id = ?",
(corp_id, row_id),
)
if inserted:
conn.executemany(
"""INSERT OR IGNORE INTO atlas_city_names
(body_id, name, kind, economic_role, population, corp_id, reserved)
VALUES (?, ?, ?, ?, ?, ?, ?)""",
inserted,
)
return len(updated), len(inserted)
def validate_brands(conn: sqlite3.Connection) -> list[str]:
"""Brand layer structural validation rules V-B01 through V-B06.
V-B01: Every brand_products row has a valid corp_id (FK to corporations).
V-B02: Every brand_inputs row has valid brand_product_id and commodity_id FKs.
V-B03: Every halo brand has at least one brand_inputs entry (demand stub must consume).
V-B04: Every volume tier must reference an existing halo brand_product_id.
V-B05: No brand_product_id is used as halo_brand_id by a non-volume-tier product.
V-B06: Every enum column (brand_category, value_trajectory, scarcity_class,
brand_tier, currency_denomination) is a member of its VALID_* set.
"""
errors: list[str] = []
# V-B01: brand_products → corporations FK
orphan_corps = conn.execute("""
SELECT bp.brand_product_id, bp.corp_id
FROM brand_products bp
LEFT JOIN corporations c ON bp.corp_id = c.corp_id
WHERE c.corp_id IS NULL
""").fetchall()
for pid, corp_id in orphan_corps:
errors.append(
f"V-B01: brand_product '{pid}' references unknown corp_id '{corp_id}'"
)
# V-B02: brand_inputs → brand_products and brand_inputs → commodities FKs
orphan_inputs_bp = conn.execute("""
SELECT bi.brand_product_id, bi.commodity_id
FROM brand_inputs bi
LEFT JOIN brand_products bp ON bi.brand_product_id = bp.brand_product_id
WHERE bp.brand_product_id IS NULL
""").fetchall()
for pid, cid in orphan_inputs_bp:
errors.append(
f"V-B02: brand_inputs row ({pid}, {cid}) references unknown brand_product_id"
)
orphan_inputs_comm = conn.execute("""
SELECT bi.brand_product_id, bi.commodity_id
FROM brand_inputs bi
LEFT JOIN commodities c ON bi.commodity_id = c.commodity_id
WHERE c.commodity_id IS NULL
""").fetchall()
for pid, cid in orphan_inputs_comm:
errors.append(
f"V-B02: brand_inputs row ({pid}, {cid}) references unknown commodity_id '{cid}'"
)
# V-B03: every halo brand has at least one brand_inputs entry
halo_no_inputs = conn.execute("""
SELECT bp.brand_product_id
FROM brand_products bp
WHERE bp.brand_tier = 'halo'
AND bp.brand_product_id NOT IN (SELECT brand_product_id FROM brand_inputs)
""").fetchall()
for (pid,) in halo_no_inputs:
errors.append(
f"V-B03: halo brand '{pid}' has no brand_inputs entries "
f"(must consume at least one commodity as a demand node)"
)
# V-B04: volume tiers reference valid halo_brand_id
volume_bad_halo = conn.execute("""
SELECT bp.brand_product_id, bp.halo_brand_id
FROM brand_products bp
WHERE bp.brand_tier = 'volume'
AND (bp.halo_brand_id IS NULL
OR bp.halo_brand_id NOT IN (SELECT brand_product_id FROM brand_products))
""").fetchall()
for pid, halo_id in volume_bad_halo:
errors.append(
f"V-B04: volume brand '{pid}' has invalid halo_brand_id '{halo_id}'"
)
# V-B05: halo_brand_id must only point to halo-tier products
halo_points_to_non_halo = conn.execute("""
SELECT child.brand_product_id, child.halo_brand_id, parent.brand_tier
FROM brand_products child
JOIN brand_products parent ON child.halo_brand_id = parent.brand_product_id
WHERE child.brand_tier = 'volume'
AND parent.brand_tier != 'halo'
""").fetchall()
for child_id, halo_id, parent_tier in halo_points_to_non_halo:
errors.append(
f"V-B05: volume brand '{child_id}' points to '{halo_id}' "
f"which has brand_tier='{parent_tier}', not 'halo'"
)
# V-B06: every enum column is in its VALID_* set. The SQL columns are
# plain TEXT without CHECK constraints, so a typo like `terrior` would
# otherwise silently import.
enum_checks = [
("brand_category", VALID_BRAND_CATEGORIES),
("value_trajectory", VALID_VALUE_TRAJECTORIES),
("scarcity_class", VALID_SCARCITY_CLASSES),
("brand_tier", VALID_BRAND_TIERS),
("currency_denomination", VALID_CURRENCY_DENOMINATIONS),
("price_tier", VALID_PRICE_TIERS),
]
for column, valid_set in enum_checks:
bad = conn.execute(
f"SELECT brand_product_id, {column} FROM brand_products"
).fetchall()
for pid, value in bad:
if value is None:
continue # nullable columns (e.g. price_tier) may be unset
if value not in valid_set:
errors.append(
f"V-B06: brand_product '{pid}' has {column}='{value}' — "
f"must be one of {sorted(valid_set)}"
)
return errors
# ---------------------------------------------------------------------------
# System specialization — D-237 authored layer
# ---------------------------------------------------------------------------
# Authoritative D-233 projected enums. The full CI guardrail suite (V-SES-*)
# lands in #1015; this function performs the FK + enum sanity the import itself
# needs to stay sound. NOTE for #1015: the D-237 "equal-or-higher" override rule
# must NOT hard-fail HUB specializations (shipbuilding, transit_hub) whose local
# production_ubiquity_projected is intentionally below their commodity's global
# default — see specialization_vocabulary.toml header.
_BULK_CLASSES = {"BulkSolid", "BulkLiquid", "PrecisionDense", "Perishable", "NonPhysical"}
_PRODUCTION_UBIQUITY = {"Ubiquitous", "Common", "Specialist", "MonopolySource"}
_FACTION_VOCAB = {
"concord_assembly", "compact", "compact_sympathetic", "syndic_dominant",
"veil_institute", "independent", "disputed", "mixed",
}
# Combined cultural_specialization vocabulary (D-237; miri-round3 §2). Two value
# kinds in one column: activity/character and founding-heritage. EXTENSIBLE — the
# #1016 content pass adds heritage values here as more GTTR systems are reviewed;
# add the new value to this set and CI accepts it. V-SES-04 validates against it.
_CULTURAL_ACTIVITY = {
"scholarly", "artistic", "institutional", "commercial", "agrarian",
"industrial_heritage", "medical_elite", "ecological", "military",
"financial_technocratic", "cosmopolitan", "compact_cooperative",
}
# Canonical 47-value heritage taxonomy (#1016, D-237). Source of truth:
# docs/workshops/system-economic-specialization/heritage-taxonomy-draft.md.
# Real-world people/nationality granularity, lowercase_snake. Heritage wins over
# activity values when both apply. Pin only when a system's founding heritage
# DIVERGES from its corridor baseline (D-167/D-232); corridor-typical systems
# stay NULL and take the corridor default.
_CULTURAL_HERITAGE = {
# British Isles & Anglo-diaspora
"anglo", "scottish", "irish", "welsh",
# Iberian, Latin & Lusophone
"portuguese", "brazilian", "afro_brazilian", "cape_verdean", "angolan",
"sao_tomean", "spanish", "canarian", "italian", "french",
# Northern / Central / Eastern European
"german", "dutch", "nordic", "finnish", "polish", "czech", "russian",
"luxembourgish", "hungarian",
# Sub-Saharan African
"afrikaans", "cape_malay", "zulu", "xhosa", "herero", "shona", "swahili",
"igbo", "yoruba", "hausa", "akan",
# South Asian
"indian", "bengali", "punjabi", "konkan",
# East & Southeast Asian
"chinese", "korean", "japanese", "vietnamese", "tagalog",
# Pacific
"maori",
# Middle East / North Africa / Central Asia
"arab", "persian", "turkic",
}
_CULTURAL_VOCAB = _CULTURAL_ACTIVITY | _CULTURAL_HERITAGE
# Catalog production_ubiquity concentration ranking for V-SES-03 (override may
# only be >= the commodity's global default). regional ≈ common tier.
_UBIQUITY_RANK = {
"ubiquitous": 0, "common": 1, "regional": 1, "concentrated": 2, "monopolistic": 3,
}
def import_system_specialization(conn: sqlite3.Connection, dry_run: bool,
strict: bool = False) -> dict:
"""Import + validate the D-237 authored specialization layer (#1013, #1015).
Reads two TOMLs:
- specialization_vocabulary.toml -> specialization_vocabulary table
(FK-validated against commodities; MUST run after import_commodities).
- system_specialization.toml -> UPSERTs economic_specialization +
cultural_specialization onto system_economy, and dominant_faction onto
system_factions, for authored (hero) systems only.
Authored lore wins; unauthored systems are left NULL for the generator's
heuristic fallback. economic_specialization / cultural_specialization are
owned exclusively by this importer, so they are cleared to NULL first for
idempotency (a removed stanza must not leave a stale value). dominant_faction
is shared with other derivation paths, so it is ONLY overwritten for systems
present in the TOML (per #1013) — never globally cleared.
CI guardrails (#1015):
Always-on hard errors (abort): V-SES-01 (econ value in vocab), V-SES-03
(override >= catalog concentration), V-SES-04 (cultural value in vocab),
V-SES-05 (faction in vocab), V-SES-06 (vocab commodity FK), plus unknown
system_id.
Completeness gates V-SES-02 (every inhabited system resolves a non-null
economic value) and V-FAC-01 (every inhabited named system has authored
dominant_faction) are HARD only under `strict` — their preconditions are
the #1014 fallback (blocked by #982) and the #1016 content pass. Until
those land, they emit warnings; flip `--strict-specialization` on once
both are complete so regen-db enforces them.
Soft warnings (W-SES-*, W-FAC-*) and the coverage report always print.
Returns a coverage dict for the caller's report. Raises _ImportAborted on a
hard validation failure.
"""
with open(SPECIALIZATION_VOCAB_TOML, "rb") as f:
vocab = tomllib.load(f)
with open(SYSTEM_SPECIALIZATION_TOML, "rb") as f:
systems = tomllib.load(f)
commodity_pu = {
r[0]: r[1] for r in conn.execute(
"SELECT commodity_id, production_ubiquity FROM commodities"
).fetchall()
}
commodity_ids = set(commodity_pu)
system_ids = {
r[0] for r in conn.execute("SELECT system_id FROM star_systems").fetchall()
}
economy_system_ids = {
r[0] for r in conn.execute("SELECT system_id FROM system_economy").fetchall()
}
faction_system_ids = {
r[0] for r in conn.execute("SELECT system_id FROM system_factions").fetchall()
}
errors: list[str] = []
# --- Vocabulary: FK + enum validation (V-SES-06, V-SES-03) ----------
vocab_rows = []
for spec_id, v in vocab.items():
cid = v.get("commodity_id")
if cid not in commodity_ids: # V-SES-06
errors.append(
f"V-SES-06: specialization_vocabulary '{spec_id}': commodity_id "
f"'{cid}' not in commodities catalog"
)
bc = v.get("bulk_class_projected")
if bc not in _BULK_CLASSES:
errors.append(
f"specialization_vocabulary '{spec_id}': bulk_class_projected "
f"'{bc}' invalid (expected one of {sorted(_BULK_CLASSES)})"
)
pu = v.get("production_ubiquity_projected")
if pu not in _PRODUCTION_UBIQUITY:
errors.append(
f"specialization_vocabulary '{spec_id}': "
f"production_ubiquity_projected '{pu}' invalid "
f"(expected one of {sorted(_PRODUCTION_UBIQUITY)})"
)
override = v.get("production_ubiquity_override") or None # "" -> NULL
# V-SES-03: a non-empty override may only raise (or equal) the
# commodity's global concentration — never claim a globally scarce good
# is locally more common. Empty override = HUB value (intentionally
# projects below catalog; exempt — see vocab TOML header).
if override is not None and cid in commodity_pu:
cat_rank = _UBIQUITY_RANK.get(commodity_pu[cid], -1)
ovr_rank = _UBIQUITY_RANK.get(override, -1)
if ovr_rank < 0:
errors.append(
f"V-SES-03: specialization_vocabulary '{spec_id}': "
f"production_ubiquity_override '{override}' not a catalog term"
)
elif ovr_rank < cat_rank:
errors.append(
f"V-SES-03: specialization_vocabulary '{spec_id}': override "
f"'{override}' is less concentrated than commodity "
f"'{cid}' catalog default '{commodity_pu[cid]}' — incoherent"
)
vocab_rows.append((spec_id, cid, override, bc, pu, v.get("description", "")))
valid_spec_ids = set(vocab.keys())
# --- System stanzas: id + value validation (V-SES-01/04/05) ---------
for sid, s in systems.items():
if sid not in system_ids:
errors.append(
f"system_specialization '{sid}': not a known star_systems.system_id"
)
es = s.get("economic_specialization")
if es is not None and es not in valid_spec_ids: # V-SES-01
errors.append(
f"V-SES-01: system_specialization '{sid}': economic_specialization "
f"'{es}' not in specialization_vocabulary"
)
cs = s.get("cultural_specialization")
if cs is not None and cs not in _CULTURAL_VOCAB: # V-SES-04
errors.append(
f"V-SES-04: system_specialization '{sid}': cultural_specialization "
f"'{cs}' not in the activity+heritage vocabulary "
f"(add new heritage values to _CULTURAL_HERITAGE)"
)
df = s.get("dominant_faction")
if df is not None and df not in _FACTION_VOCAB: # V-SES-05
errors.append(
f"V-SES-05: system_specialization '{sid}': dominant_faction "
f"'{df}' invalid (expected one of {sorted(_FACTION_VOCAB)})"
)
if errors:
print(f" SPECIALIZATION ERRORS ({len(errors)}):")
for e in errors:
print(f" - {e}")
raise _ImportAborted()
coverage = {
"vocab": len(vocab_rows),
"economic": 0,
"cultural": 0,
"faction": 0,
"missing_economy_row": [],
"missing_faction_row": [],
}
if dry_run:
for sid, s in systems.items():
coverage["economic"] += 1 if s.get("economic_specialization") else 0
coverage["cultural"] += 1 if s.get("cultural_specialization") else 0
coverage["faction"] += 1 if s.get("dominant_faction") else 0
_specialization_checks(conn, vocab, systems, strict)
return coverage
# --- Repopulate vocabulary table ------------------------------------
conn.execute("DELETE FROM specialization_vocabulary")
conn.executemany(
"""INSERT INTO specialization_vocabulary (
specialization_id, commodity_id, production_ubiquity_override,
bulk_class_projected, production_ubiquity_projected, description
) VALUES (?, ?, ?, ?, ?, ?)""",
vocab_rows,
)
# --- Clear importer-owned columns (idempotency) ---------------------
conn.execute(
"UPDATE system_economy SET economic_specialization = NULL, "
"cultural_specialization = NULL"
)
# --- UPSERT per-system authored fields ------------------------------
for sid, s in systems.items():
es = s.get("economic_specialization")
cs = s.get("cultural_specialization")
if sid in economy_system_ids:
conn.execute(
"UPDATE system_economy SET economic_specialization = ?, "
"cultural_specialization = ? WHERE system_id = ?",
(es, cs, sid),
)
coverage["economic"] += 1 if es else 0
coverage["cultural"] += 1 if cs else 0
else:
coverage["missing_economy_row"].append(sid)
df = s.get("dominant_faction")
if df is not None:
if sid in faction_system_ids:
conn.execute(
"UPDATE system_factions SET dominant_faction = ? "
"WHERE system_id = ?",
(df, sid),
)
coverage["faction"] += 1
else:
coverage["missing_faction_row"].append(sid)
# --- Completeness gates + soft warnings + coverage report -----------
# Run AFTER the UPSERTs so they see the freshly-written DB state.
_specialization_checks(conn, vocab, systems, strict)
return coverage
def _specialization_checks(conn, vocab, systems, strict):
"""V-SES-02 / V-FAC-01 completeness gates, soft warnings, coverage report.
Reads the post-UPSERT DB state. Gates are warnings unless `strict` (their
preconditions — the #1014 fallback and the #1016 content pass — are not yet
in place). Raises _ImportAborted only when strict and a gate fails.
"""
gate_failures: list[str] = []
warnings: list[str] = []
# Population: integer where present. Inhabited = population > 0.
inhabited = [
(r[0], r[1]) for r in conn.execute(
"SELECT system_id, population FROM system_economy "
"WHERE population IS NOT NULL AND population > 0"
).fetchall()
]
econ = {
r[0]: r[1] for r in conn.execute(
"SELECT system_id, economic_specialization FROM system_economy"
).fetchall()
}
cult = {
r[0]: r[1] for r in conn.execute(
"SELECT system_id, cultural_specialization FROM system_economy"
).fetchall()
}
faction = {
r[0]: r[1] for r in conn.execute(
"SELECT system_id, dominant_faction FROM system_factions"
).fetchall()
}
currency = {
r[0]: r[1] for r in conn.execute(
"SELECT system_id, currency_zone FROM star_systems"
).fetchall()
}
# "Named" = has authored GTTR identity (proper_name or gttr_hook).
named = {
r[0] for r in conn.execute(
"SELECT system_id FROM star_systems "
"WHERE (proper_name IS NOT NULL AND proper_name != '') "
" OR (gttr_hook IS NOT NULL AND gttr_hook != '')"
).fetchall()
}
# Tier-1 monopolist corp HQ presence per system (best-effort; tables may be
# sparse pre-#1016). primary_operation/headquarters live on corp_presence.
hq_systems = set()
try:
hq_systems = {
r[0] for r in conn.execute(
"SELECT DISTINCT location_id FROM corp_presence "
"WHERE primary_operation IS NOT NULL"
).fetchall()
}
except sqlite3.OperationalError:
pass
vocab_pu = {k: (v.get("production_ubiquity_projected")) for k, v in vocab.items()}
vocab_commodity = {k: v.get("commodity_id") for k, v in vocab.items()}
# V-SES-02: every inhabited system must resolve a non-null economic value
# (authored here, or via the #1014 fallback once it exists).
for sid, _pop in inhabited:
if not econ.get(sid):
gate_failures.append(
f"V-SES-02: inhabited system '{sid}' has no economic_specialization "
f"(authored or fallback)"
)
# V-FAC-01: every inhabited NAMED system must have an authored faction.
for sid, _pop in inhabited:
if sid in named and not faction.get(sid):
gate_failures.append(
f"V-FAC-01: inhabited named system '{sid}' has no dominant_faction"
)
# --- Soft warnings --------------------------------------------------
# W-SES-01: MonopolySource systems for D-177 human review.
monopoly = [
sid for sid, e in econ.items()
if e and vocab_pu.get(e) == "MonopolySource"
]
if monopoly:
warnings.append(f"W-SES-01: MonopolySource systems [D-177 review]: {sorted(monopoly)}")
# W-SES-02: >25% of authored-economic systems share one value.
from collections import Counter
econ_counts = Counter(e for e in econ.values() if e)
n_authored_econ = sum(econ_counts.values())
if n_authored_econ:
for val, cnt in econ_counts.items():
if cnt > 0.25 * n_authored_econ and cnt > 2:
warnings.append(
f"W-SES-02: '{val}' covers {cnt}/{n_authored_econ} "
f"({100*cnt//n_authored_econ}%) of authored-economic systems"
)
# W-SES-08: estate_farming + large population (probably breadbasket).
for sid, pop in inhabited:
if econ.get(sid) == "estate_farming" and pop and pop > 5_000_000:
warnings.append(
f"W-SES-08: '{sid}' is estate_farming with population {pop} "
f"(probably breadbasket)"
)
# W-FAC-01: compact + tractus_primary currency (D-172 violation).
# W-FAC-04: compact_sympathetic + mark_primary (may be full member).
for sid, f in faction.items():
if not f:
continue
cz = (currency.get(sid) or "").upper()
if f == "compact" and cz == "TRACTUS_PRIMARY":
warnings.append(f"W-FAC-01: '{sid}' compact + TRACTUS_PRIMARY currency (D-172)")
if f == "compact_sympathetic" and cz == "MARK_PRIMARY":
warnings.append(f"W-FAC-04: '{sid}' compact_sympathetic + MARK_PRIMARY (may be full member)")
# W-FAC-02: compact + MonopolySource extraction (Compact self-sufficiency),
# excluding terroir_* / marble_monopoly (lore-sanctioned monopolies).
_exempt = {"marble_monopoly", "terroir_agriculture", "terroir_spirits", "terroir_organics"}
for sid, f in faction.items():
e = econ.get(sid)
if f == "compact" and e and vocab_pu.get(e) == "MonopolySource" and e not in _exempt:
warnings.append(f"W-FAC-02: '{sid}' compact + MonopolySource '{e}' (self-sufficiency doctrine)")
# W-FAC-03: syndic_dominant + no corp HQ presence (ungrounded pin).
for sid, f in faction.items():
if f == "syndic_dominant" and sid not in hq_systems:
warnings.append(f"W-FAC-03: '{sid}' syndic_dominant but no corp HQ in corp_presence (ungrounded)")
# --- Coverage report (always) ---------------------------------------
n_inhabited = len(inhabited)
n_named = len(named)
n_econ = sum(1 for e in econ.values() if e)
n_cult = sum(1 for c in cult.values() if c)
n_fac = sum(1 for f in faction.values() if f)
bulk_dist = Counter()
pu_dist = Counter()
for e in econ.values():
if e and e in vocab:
bulk_dist[vocab[e].get("bulk_class_projected")] += 1
pu_dist[vocab[e].get("production_ubiquity_projected")] += 1
print(" Specialization coverage (D-237):")
print(f" economic: {n_econ} authored ({n_inhabited} inhabited; rest on fallback once #1014 lands)")
print(f" cultural: {n_cult} authored ({n_named} named; rest on corridor default)")
print(f" faction: {n_fac} authored ({n_named} named; rest on derivation)")
print(f" BulkClass: " + " | ".join(f"{k}={v}" for k, v in sorted(bulk_dist.items())))
print(f" ProductionUbiquity: " + " | ".join(f"{k}={v}" for k, v in sorted(pu_dist.items())))
if warnings:
print(f" Specialization warnings ({len(warnings)}):")
for w in warnings:
print(f" - {w}")
if gate_failures:
if strict:
print(f" SPECIALIZATION COMPLETENESS FAILURES ({len(gate_failures)}) [strict]:")
for g in gate_failures:
print(f" - {g}")
raise _ImportAborted()
else:
print(
f" Specialization completeness: {len(gate_failures)} gate item(s) "
f"pending (#1014 fallback / #1016 content pass) — warnings only until "
f"--strict-specialization"
)
# ---------------------------------------------------------------------------
# Main
# ---------------------------------------------------------------------------
def main():
parser = argparse.ArgumentParser(description="Import economics data into systems.db")
parser.add_argument("--db", default=str(DB_PATH), help="Path to systems.db")
parser.add_argument("--dry-run", action="store_true", help="Validate without writing")
parser.add_argument(
"--strict-specialization", action="store_true",
help="Treat D-237 completeness gates (V-SES-02, V-FAC-01) as hard errors. "
"Off by default until the #1014 fallback and #1016 content pass land.",
)
args = parser.parse_args()
db_path = Path(args.db)
if not db_path.exists():
print(f"error: {db_path} not found", file=sys.stderr)
sys.exit(1)
print(f"\n Economics Import Pipeline")
print(f" DB: {db_path}")
if args.dry_run:
print(f" Mode: DRY RUN")
print()
# Load wiki corps before opening DB — allows early exit on parse failures
print(" Loading wiki corporations...")
wiki_corps = load_wiki_corps()
print(f" {len(wiki_corps)} corporation files parsed")
# Regenerate generated_brands.toml via the Rust binary before the Python
# import reads it. Single pipeline, single stamp — resolves review T2/H3
# ("on-behalf stamping" coupling) by folding brand generation into
# import_economics' flow rather than having the caller (Makefile / user)
# remember to run it first. Skipped on --dry-run to avoid a disk
# side-effect during validation.
if not args.dry_run:
try:
regenerate_brands()
except _ImportAborted:
sys.exit(1)
conn = sqlite3.connect(str(db_path))
conn.execute("PRAGMA foreign_keys=ON")
# The clear-then-reimport cycle below runs as a single explicit
# transaction. Any crash, validation error, or KeyboardInterrupt
# between the first DELETE and the final commit rolls everything
# back — the DB never ends up half-cleared with stale rows in some
# tables and empty rows in others. On success we commit exactly
# once, immediately after structural validation passes.
conn.execute("BEGIN")
try:
# 1. Migrate schema (idempotent, inside the tx so a crash here
# leaves no half-applied ALTER TABLE.)
print(" [1/10] Schema migration...")
for table, col, col_type in COLUMN_MIGRATIONS:
_add_column(conn, table, col, col_type)
conn.executescript(MIGRATION_SQL)
# Atlas index tables: apply canonical DDL + empty geometry (D-223, #951)
ensure_atlas_index_schema(conn, args.dry_run)
print(" tables and columns ready")
# Clear economics tables in FK-safe order (children before parents)
# corp_presence cleared here; corporations table is append-only
# (never cleared).
if not args.dry_run:
conn.execute("DELETE FROM corp_presence")
conn.execute("DELETE FROM brand_inputs")
conn.execute("DELETE FROM brand_products")
conn.execute("DELETE FROM system_fiscal")
conn.execute("DELETE FROM corp_financial_state")
conn.execute("DELETE FROM corp_lifecycle_events")
conn.execute("DELETE FROM chain_inputs")
conn.execute("DELETE FROM production_chains")
# specialization_vocabulary FK-references commodities — clear it
# before commodities so the FK-on delete does not fail (D-237).
conn.execute("DELETE FROM specialization_vocabulary")
conn.execute("DELETE FROM commodities")
conn.execute("DELETE FROM gate_links")
# 2. Gate links
print(" [2/10] Importing gate links...")
n_links = import_gate_links(conn, args.dry_run)
print(f" {n_links} rows (bidirectional)")
# 3. Commodities
print(" [3/10] Importing commodities...")
n_commodities = import_commodities(conn, args.dry_run)
print(f" {n_commodities} commodities")
# 4. Production chains
print(" [4/10] Importing production chains...")
n_chains, n_inputs = import_chains(conn, args.dry_run)
print(f" {n_chains} chains, {n_inputs} inputs")
# 4b. System specialization (D-237 authored layer) — after commodities
# (FK) and before currency zones. UPSERTs onto pre-existing
# system_economy / system_factions rows; unauthored systems stay NULL.
print(" [4b/10] Importing system specialization (D-237)...")
spec = import_system_specialization(
conn, args.dry_run, strict=args.strict_specialization
)
print(
f" vocab {spec['vocab']} | economic {spec['economic']} | "
f"cultural {spec['cultural']} | faction {spec['faction']}"
)
if spec["missing_economy_row"]:
print(
f" WARNING: {len(spec['missing_economy_row'])} authored "
f"system(s) lack a system_economy row (values dropped): "
f"{spec['missing_economy_row']}"
)
if spec["missing_faction_row"]:
print(
f" WARNING: {len(spec['missing_faction_row'])} authored "
f"system(s) lack a system_factions row (faction dropped): "
f"{spec['missing_faction_row']}"
)
# 5. Currency zones
print(" [5/10] Setting currency zones...")
zones = set_currency_zones(conn, args.dry_run)
for zone, count in sorted(zones.items()):
print(f" {zone}: {count}")
# 6. Gate energy connectivity (D-186) — must run after currency zones
print(" [6/10] Setting gate energy connectivity...")
energy = set_gate_energy(conn, args.dry_run)
for label, count in sorted(energy.items()):
print(f" {label}: {count}")
# 7. Sync corporations from wiki (D-182: hard error on name divergence)
print(" [7/10] Syncing corporations...")
corp_errors = sync_corporations(conn, wiki_corps, args.dry_run)
if corp_errors:
print(f" ERRORS ({len(corp_errors)}) — name divergence detected (D-182):")
for e in corp_errors:
print(f" - {e}")
print(" Fix: update wiki title or DB proper_name to match, then re-run.")
raise _ImportAborted()
n_db_corps = conn.execute("SELECT COUNT(*) FROM corporations").fetchone()[0]
print(f" {n_db_corps} corporations in DB ({len(wiki_corps)} from wiki)")
# 8. Corp presence from wiki headquarters data
print(" [8/10] Importing corp presence...")
commodity_ids = {
r[0] for r in conn.execute("SELECT commodity_id FROM commodities").fetchall()
}
n_presence = import_corp_presence(conn, wiki_corps, commodity_ids, args.dry_run)
print(f" {n_presence} corp_presence rows")
# 9. Brand products and inputs (D-189, #827)
print(" [9/10] Importing brand products and inputs...")
n_brands, n_brand_inputs = import_brands(conn, args.dry_run)
print(f" {n_brands} brand_products, {n_brand_inputs} brand_inputs")
# 10. System fiscal parameters (D-189 section 6)
print(" [10/13] Populating system_fiscal...")
n_fiscal = import_system_fiscal(conn, args.dry_run)
print(f" {n_fiscal} system_fiscal rows")
# 11. body_radius_km fallback from planet_class (D-204, #910)
print(" [11/13] Populating body_radius_km fallback...")
n_radius = populate_body_radius_km(conn, args.dry_run)
print(f" {n_radius} bodies updated")
# 12. atlas_city_names from wiki markers.json (D-207, #908)
print(" [12/13] Populating atlas_city_names from wiki content...")
n_cities = populate_atlas_city_names(conn, args.dry_run)
print(f" {n_cities} city name rows")
# 13. atlas_city_names corp HQ cross-reference (D-207, #909)
print(" [13/13] Cross-referencing corp HQ cities into atlas_city_names...")
n_updated, n_inserted = populate_atlas_city_names_corps(conn, args.dry_run)
print(f" {n_updated} rows updated, {n_inserted} reserved rows inserted")
# Validate structural integrity (FK, chain refs, chain completeness).
# These errors indicate broken imported data — do NOT commit.
print("\n Validating structural integrity...")
struct_errors = validate(conn)
struct_errors.extend(validate_brands(conn))
if struct_errors:
print(f" STRUCTURAL ERRORS ({len(struct_errors)}) — rolling back:")
for e in struct_errors:
print(f" - {e}")
raise _ImportAborted()
print(" FK integrity, chain completeness, and brand layer (V-B01V-B06) OK")
# Commit all imported data (corps, presence, etc.) before coverage check.
# Coverage validation is a Phase 2 gate (D-175) — data should be persisted
# so tools can query it and report gaps clearly.
if not args.dry_run:
conn.commit()
print(" Data committed.")
else:
# Dry-run: leave the transaction open so the coverage check below
# can still SELECT against the in-memory imported data. The
# transaction is discarded when conn.close() runs on exit.
print(" Dry run — no changes written.")
except _ImportAborted:
conn.rollback()
conn.close()
sys.exit(1)
except BaseException:
# Any other exception (KeyboardInterrupt, MemoryError, DB error,
# programmer error) triggers a rollback so the DB is never left in
# a half-imported state. Re-raise so the user sees the traceback.
conn.rollback()
conn.close()
raise
# Stamp generator metadata (#855, #856): record source SHAs so the
# pre-push hook can detect stale DB snapshots. Written BEFORE the
# coverage gate — the stamp records generator execution (code version),
# not data completeness. Coverage gaps (#860) are pre-existing data
# issues and must not prevent the stamp from landing.
if not args.dry_run:
try:
_write_stamp(conn, "import_economics", *IMPORT_ECONOMICS_SOURCES)
conn.commit()
print(" Stamped: import_economics (covers brand pipeline Rust sources)")
except Exception as exc: # noqa: BLE001
print(f" WARNING: failed to write generator stamp: {exc}", file=sys.stderr)
# Validate coverage (hard errors per D-175, but after commit so data is usable).
print("\n Validating coverage (D-175 Phase 2 gate)...")
coverage_errors: list[str] = []
commodity_ids_for_coverage = {
r[0] for r in conn.execute("SELECT commodity_id FROM commodities").fetchall()
}
coverage_errors.extend(
_validate_commodity_coverage(conn, wiki_corps, commodity_ids_for_coverage)
)
coverage_errors.extend(_validate_system_coverage(conn, wiki_corps))
if coverage_errors:
print(f" COVERAGE ERRORS ({len(coverage_errors)}) — Phase 2 gate not met:")
for e in coverage_errors:
print(f" - {e}")
print("\n Data committed but Phase 2 gate is NOT met. "
"Add corporations to meet coverage thresholds and re-run.")
conn.close()
sys.exit(2) # exit 2 = coverage warning (data+stamp committed); exit 1 = real error
else:
print(" All coverage thresholds met — Phase 2 gate PASSED.")
conn.close()
print(f"\n Done: {n_links} gate_links, {n_commodities} commodities, "
f"{n_chains} chains, {n_inputs} inputs, {n_presence} corp_presence, "
f"{n_brands} brand_products, {n_brand_inputs} brand_inputs, "
f"{n_fiscal} system_fiscal\n")
if __name__ == "__main__":
main()