Files
settled-reach/docs/wiki-structure-findings.md
T
jpmschweitzerandClaude Opus 5 c25af8d753 docs(meta): make the wiki seed reachable — blind-prediction experiment and its fix
An experiment, at Jeroen's request: predict how the wiki seed data is structured
WITHOUT reading it, seal the prediction, then score it. The prediction is
a1addf7e2, committed before wiki/ was opened so it could not be retrofitted.

The score, against a rule fixed in advance:

RIGHT — markdown + YAML frontmatter, TOML for economics tables, the body path
shape, more trees than the three I had seen.

WRONG — "a source, never an output". That holds for 253 pages and is backwards
for 3,262: star-systems/ is GENERATED from systems.db by tooling/db/wiki_sync.py,
its <!-- READ-ONLY --> blocks are renders, and body frontmatter IS the body
definition rather than a description of one. Also wrong: "probably no schema
docs" — there are 17 templates, ten authoring guides, economics/schema.md and a
GOVERNANCE.md that states the ownership models plainly.

ABSENT (the expensive bucket) — the two ownership models running in OPPOSITE
directions; the GTTR prose channel; terrain.npz/globe.png; that stations and
districts have NO wiki directories; that `description` frontmatter exists so
agents can filter before loading; and the scale, 11,864 files.

ROOT CAUSE, and it is not missing documentation. The wiki documents itself well.
It was unreachable: wiki/ appears in NEITHER CLAUDE.md's Project Structure block
NOR .claude/rules/project-structure.md, the annotated tree whose entire job is
orienting an agent. The largest tree in the repo — the seed for the whole Reach —
was invisible from both files a session reads first. Every item in the absent
bucket follows from that one omission. The proof is this session: it spent three
days fixing Ferrath's terrain rendering and never once saw
wiki/star-systems/GJ-820B/bodies/GJ820Bc/index.md, the file that defines Ferrath.

Fixed here: wiki/ enters both structure documents with the ownership split stated
where it will be read, and Skill(wiki) carries the traps — never hand-edit a
READ-ONLY block or body frontmatter, stations have no directories, the id is
spelled two ways, editing corp PROSE stales systems.db, and absent variance is
often deliberate rather than a gap.

That last point cost two false findings in one measurement and is worth the
warning: chemosynthetic:false on every body is a namespace reservation for
dextro-DNA-style biochemistry once geology and nature spawn to the 1x1m pixel,
and enabled:false on ~65% is staged rollout — clean planet types first, generator
scripts for the rest after. Both read as defects without the roadmap.

Also measured, since the seed's job is to supply variance: continuous axes are
rich (unique seed per body, 460-716 distinct values across orbit/tilt/ice/land)
while the categoricals that gate morphology are concentrated (68% tectonics low,
51% planet_class frozen). Filed as T-1244 with the design question stated first —
whether the distribution is intended — rather than as a defect.

Method caveat recorded in the findings: the aggregator reads scalar frontmatter
only, and atmosphere_color's "100% null" was a parser artefact, not a finding.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-20 00:31:06 +02:00

11 KiB
Raw Blame History

Findings: how the wiki seed data is actually structured

2026-08-20. Scored against wiki-structure-blind-prediction.md, which was committed before wiki/ was opened. Scoring rule was fixed in advance: right, wrong, absent — with absent (a load-bearing thing I did not know existed) treated as the expensive bucket, because that is what a future session walks into blind.

The one-line version

The wiki is 93% machine-generated and 7% authored, and for the generated 93% the direction of truth runs DB → wiki, which is the opposite of what the prediction assumed. The single most useful sentence for a new session is: most of the wiki is an OUTPUT you must not hand-edit; a small flat set is an INPUT that generators read.

Scored

Right

  • wiki/ is the authored world-content root; systems.db is a build artefact. (True for the flat categories only — see Wrong.)
  • Markdown + YAML frontmatter, with the frontmatter carrying machine-read fields.
  • TOML for tabular economics data; JSON (markers.json) for name pools.
  • star-systems/GJ-{id}/bodies/{body_id}/ with heightmap.png, reliefmap.png, markers.json. (All [SEEN], so no credit.)
  • System dirs hyphenate (GJ-820B), body dirs do not (GJ820Bc). [SEEN]
  • More content trees than the three I had seen. There are 16.
  • Frontmatter is partly load-bearing and partly decorative, with no strong schema gate on the authored side.

Wrong

  1. "A source, never an output." Wrong for 93% of the wiki. star-systems/ pages are GENERATED and carry <!-- READ-ONLY --> sections rendered from systems.db and star-map.json by tooling/db/wiki_sync.py. Authored prose lives in designated sections that survive regeneration. Hand-editing a read-only section is silently reverted on the next sync.
  2. "Bodies may not have their own pages." They do — ~2,600 of them — and their frontmatter IS the body definition, i.e. generator INPUT rather than description. wiki/star-systems/GJ-820B/bodies/GJ820Bc/index.md carries seed: 3669559456, land_fraction, axial_tilt_deg, star type. This is the file that defines Ferrath, the body whose Atlas rendering this session spent its entire length fixing, and the session never knew it existed.
  3. "Probably no schema documentation." There is wiki/GOVERNANCE.md, wiki/_templates/ (17 templates, one per entity kind), ten authoring guides under wiki/authoring/, and wiki/economics/schema.md. The documentation is not missing. It is unreachable from where an agent starts.
  4. "One generator (import_economics)." At least six touch the wiki: wiki_sync.py, scaffold_bodies.py, body_definition_parser.py, populate_gttr_hook.py, populate_terrain_reference.py, plus the atlas sync, with atlas_cohesion_audit.py auditing.

Absent — the expensive bucket

Things that are load-bearing and that I had no model of at all:

  1. The two ownership models, and that they run in opposite directions. wiki/GOVERNANCE.md states it plainly; nothing an agent reads at session start points there. This is the fact most likely to cause damage: an agent "fixing" a system page edits a read-only block and the work vanishes.
  2. The GTTR prose layergttr.md per system, gttr-{location}.md per station, gttr-{corridor}.md per corridor, feeding gttr_hook via populate_gttr_hook.py. An entire authored content channel, invisible to me.
  3. terrain.npz and globe.png per body, alongside the two PNGs I knew.
  4. Stations and districts have NO wiki directories — they are systems.db rows rendered into each system page's table, with prose in flat gttr-{slug}.md. I would have gone looking for wiki/stations/.
  5. The description frontmatter field exists for AGENTS: GOVERNANCE.md says "agents filter by description before loading full files". There is a designed affordance for context-aware loading, and no agent-facing doc mentions it.
  6. Scale: 301 systems, ~2,600 body pages, 11,864 files. The prediction had no sense of magnitude, which matters because it decides whether you grep or query.

A second-order finding: I documented a tree I had not opened

Jeroen, mid-pass: "did you scan the organizations economics and corporations dirs and such? do you know what is in it or do you have statistical scans of it?"

Statistical scans. At that point I had read index.md, GOVERNANCE.md, directory listings, one body page and a scalar aggregate over 2,625 body files — and had opened zero corporation, faction, economics, concept or triangle files. Everything written above about the flat categories came from GOVERNANCE.md DESCRIBING them. I was one commit from shipping a skill instructing others how to work in trees I had never looked inside.

Reading them changed material facts:

  • Four categories are empty scaffolding. institutions/, species/, cultural-groups/, lore/ hold 0 files while having templates. The distribution is lopsided: corporations/ 156, economics/ 17, factions/ 9, technology/ 7, contraband/ and concepts/ 4, triangles/ 2.
  • Corp frontmatter carries 14 fields, not the 3 the prediction guessed — including scope, faction_type, tags, cross_refs, decision_refs.
  • economics/ is an economic MODEL, not a set of lookup vocabularies: 37 authored commodity pages, production_chains.toml with real recipes (2t metallic_ore + 0.3 fusion_fuel -> 1t refined_metals), currency zones, shadow economy, behavioural/lore archetypes, and brand generation. The earlier draft of the skill called it "TOML vocabularies + schema.md", which is wrong in a way that would mislead.
  • specialization_vocabulary.toml is a shared id-space read by both system_economy.economic_specialization and corporations.corp_specialization — the latter authored in corp page frontmatter, not in a TOML.

The lesson generalises past this experiment: a statistical scan tells you the SHAPE of a tree and nothing about its CONTENT, and the two are easy to conflate when the statistics are detailed enough to feel like knowledge. The check that catches it is the one Jeroen applied — "have you opened one?"

Root cause of the gap

Not missing documentation — missing reachability. The wiki documents itself well (GOVERNANCE.md, templates, authoring guides, economics/schema.md). But:

  • wiki/ does not appear in CLAUDE.md's Project Structure block.
  • wiki/ does not appear in .claude/rules/project-structure.md, the detailed annotated tree that exists precisely to orient an agent. Its only match for "wiki" is the parenthetical in db/ # Asset/connector scripts (audio, image, trellis, wiki).

So the single largest content tree in the repo — 11,864 files, the seed data for the entire Reach — is absent from both documents an agent reads first. Everything in the absent bucket follows from that one omission.

The seed is a parameter space, not a reference work

Jeroen's framing (2026-08-20), which corrects the emphasis above: most of this is written once at dev time, and its job is to be the seed for variance in the generator. Not documentation to maintain — the authored input the cascade varies from. Read that way, "is the wiki well documented" is the wrong question; "does the seed supply variance, and does an agent know it IS a seed" is the right one.

Measured variance across 2,625 body definitions

Continuous axes are genuinely rich:

field distinct values
seed 2,625 — unique per body
distance_au 716
polar_ice_lat 691
axial_tilt_deg 578
land_fraction 460

Categorical axes — the ones that gate morphology and biome branches — are heavily concentrated:

field distinct dominant value
tectonics 4 67.8% low
atmosphere 4 66.3% none
planet_class 13 51.1% frozen
substrate 3 51.1% ice
geothermal_flux 3 91.2% low

name: null on ~89% of bodies is real (531 of 600 sampled), consistent with "220 named" of 301 systems.

Method caveat, stated because the number was wrong once: the aggregator reads scalar frontmatter only. atmosphere_color first read as "100% null", which was a PARSER ARTEFACT — it is a YAML list, and the scalar regex matched only its key line. Ferrath carries a real colour triple. Any list-valued field in that table would be similarly wrong; the scalar rows are sound.

Two readings that looked like defects and are not

Both were flagged on first pass and both are deliberate. They are recorded because the mistake is the point: an agent lacking the roadmap reads a zero-variance field as a bug, and may "helpfully" populate or delete it.

  • chemosynthetic: false on 100% of bodies is a NAMESPACE RESERVATION for dextro-DNA-style biochemistry, for when geology and nature spawn down to the 1×1 m pixel. It is not meant to vary yet.
  • enabled: false on ~65% is staged rollout, not dead content: the clean planet types are learned first, then generator scripts are added for the other types and the playable count rises.

What the economics tree is actually for

The economics TOMLs are not merely DB rows for a trade sim. When geology and nature spawn to the 1×1 m pixel, the economics information generates world content — production chains and specializations decide what is physically there. That makes wiki/economics/ a seed for the same generator, one tier up from the body definitions, and raises the stakes on its vocabularies.

The open question this leaves

A 68%-low-tectonics world is mostly flat BY THE SEED. This session spent its full length making the Atlas render terrain legibly, and composition.rs already says "where the terrain is flat, flat is the honest answer". So there is a ceiling on what renderer work can deliver, and it is set here rather than in the client. Whether that distribution is intended (most worlds ARE dull; the interesting ones stand out) or an artefact of the authoring pass is a question for the seed, not the renderer. Ticketed separately.

What this predicts

A session asked to change world content will, with no prompting otherwise:

  • assume the wiki is hand-authored throughout, and hand-edit a generated page;
  • not know wiki_sync.py exists, so not know the edit is reverted rather than wrong;
  • go looking for stations under wiki/stations/;
  • miss the GTTR channel entirely when asked for location prose;
  • and read whole files instead of filtering on description, burning context the wiki was explicitly designed to save.

That is the same shape as this month's expensive bugs: the information existed, nothing pointed at it, and the failure was silent.