Files
settled-reach/docs/gemma-architecture-flavor-methodology.md
jpmschweitzerandClaude Fable 5 a6c2ff4740 docs(assets): Gemma architecture-flavor methodology — bounded corpus distillation (T-992)
Stage-1 register selection over the current wiki tree (Q-107 assumption
stated), stage-2 TEXT-FIELDS-ONLY structured proposals (tag, label,
cultural_description, candidate allow/block tags) — humans own all numeric
eligibility/weights (catalog-wide CI guardrails can't be satisfied
per-template); Araminta owns visual_bundle. No era-bands (Q-106 dissolved);
trauma_decay_rate stale-schema note. Mirrors docs/gemma-naming-methodology.md
incl. the sr-voice-tooling serving pattern. Worked example: GJ-4053 granite
dry-stack accept; GJ-1248/GJ-285/GJ-588 as correct stage-1 rejections.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-08 09:15:56 +02:00

18 KiB
Raw Permalink Blame History

Gemma Architecture-Flavor Methodology (D-232, T-992)

Status: design methodology, not yet implemented. Analogous in structure and intent to docs/gemma-naming-methodology.md (D-223) — the retired pipeline that named atlas features from wiki prose — but the two are solving different problems. Naming filled thousands of empty name fields at generation time. This pass proposes a small number of new rows in an authored, frozen catalog, once, with a human in the loop before anything ships. There is no runtime Gemma call anywhere in the architecture- flavor draw (D-232 phase 1-3): trait_selection is a deterministic SeedChain-seeded pick from the catalog. Gemma's only role is helping grow the catalog file between authoring passes.


1. Scope — bounded corpus-distillation, not per-body generation

D-232 round 3 retired the round-2 per-body Gemma-array model in favor of a shared trait-template catalog. The catalog's content is split into two tracks (T-977 scope note, 2026-05-26):

  • Core hand-authored set (T-1005, shipped) — Miri (cultural meaning + eligibility) and Araminta (visual_bundle) authored 28 templates directly: 10 cross_corridor, 8 baseline, 10 heritage (see the corridor-pool breakdown in wiki/economics/architecture_trait_catalog.toml). This sits within D-232's stated floor (~25-35) but below its target (40-45).
  • This pass — a bounded Gemma corpus-distillation run: one read of all wiki prose → propose candidate templates → human-gated review, growing the catalog from 28 toward the 40-45 target. It runs after the core set exists and reads it as context (see §4.1).

What this explicitly is not:

  • Not per-body open-vocabulary generation. There is no Gemma call in the skeleton-generation or chunk-fill path. The round-2 model (a Gemma call per body, producing a bespoke culture) is the thing D-232 retired.
  • Not a recurring or CI-triggered job. It runs manually, whenever the catalog is judged to need growth (new corridor, new archetype, a corpus content addition worth distilling) — not on every wiki edit.
  • Not a source of any numeric field. See §5.

Superseded stage-2 schema note: the original T-977/T-992 ticket text (round 2) described a stage-2 output of {tag, allow, block, weight, trauma_decay_rate}. Neither weight nor trauma_decay_rate survive to D-232 round 3 — there is no per-template decay rate anywhere in the shipped schema or traits.py, and Gemma never emits a weight (§5-§6 explain why). Treat any reference to trauma_decay_rate outside this note as stale round-2 prose.


2. Model & runtime — reuse the naming pipeline's binary split

docs/architecture/atlas-naming-pipeline.md §2 documents the split the retired naming pipeline used, and the same reasoning applies here, more strongly:

  • Two binaries exist. The in-repo sr-voice serve --model <path> --stdio (server/sr-voice/) is the game runtime — Gemma 2, small context (--ctx-size, historically 1024 tokens for naming), tuned for low-latency in-game NPC dialogue, no network port (Gemma 2 T&C compliance). The naming pipeline instead used sr-voice-tooling — a separate binary built against Gemma 4 (GGUF), compiled for ROCm/HIP (gfx1201), living outside the repo (~/Projects/settled-reach/binaries/sr-voice-tooling) and run inside the reach-build distrobox. The split exists because tooling work trades latency for richer output and a bigger context window — exactly the trade this pass wants.
  • Recommendation: reuse sr-voice-tooling / Gemma 4, not the in-game runtime. This pass needs to read a full index.md page (several hundred to ~1000+ words — see §3), not the naming pipeline's hand-compressed 30-45 word GTTR hook. Gemma 2's ~1024-token context was already tight for the compressed naming case; reading full system pages needs the larger context the tooling binary provides.
  • Protocol: identical stdio JSONL, unchanged. Both binaries speak the same newline-delimited JSON contract (server/sr-voice/src/stdio.rs, docs/architecture/atlas-naming-pipeline.md §2): one request per line, {"prompt": "...", "seed": N}, one response per line, {"text": "..."} or {"error": "..."}. No new protocol work is needed to point a new script at either binary.
  • Call volume is much smaller, so KV-cache bleed matters less — but the fix is the same if it shows up. The naming pipeline processed ~2,400 bodies × 5 feature types and needed --refresh-driven subprocess restarts to stop a long-lived model echoing earlier completions. This pass makes on the order of ~300 stage-1 calls (one per system) plus one stage-2 call per surfaced candidate (expected: a small fraction of 300, since most systems should say "no distinct tradition" — see §4.1). If output diversity degrades across a long run, apply the same lesson: restart the subprocess periodically (e.g. every system boundary) to flush context.
  • Few-shot over instructions carries over unchanged. The naming pipeline's model-fit lesson — Gemma is far better at pattern completion than instruction-following — applies to both stages here. Prompts should be worked examples ("system prose → verdict + name", "flagged prose → structured JSON"), not bare instructions.

3. Source material — the wiki tree, and the Q-107 assumption

Stage 1 reads wiki/star-systems/*/index.md — the system-level wiki page (System Profile, cultural/economic prose, Faction Notes, Silence Topic, Narrative Hook). All 302 systems have one.

This is an explicit assumption, not a permanent contract. Q-107 (open — wiki → Atlas content-set consolidation) proposes moving all wiki content (pages, heightmaps, markers, architecture-flavor source) into one consolidated, first-class content tree. If Q-107 resolves and the content moves, only stage 1's file glob needs updating — the extraction prompt, the structured-output schema, and the human-review gate are all location-invariant, per D-232's own note that "the generator-facing tables are invariant to [source location]."

Companion gttr.md / per-body gttr-<slug>.md files (present for ~301 of 302 systems) carry additional narrative color and were the naming pipeline's secondary source for register selection. This pass scopes stage 1 to index.md only — per the T-992 refinement. Folding gttr*.md in as a secondary source is a plausible future expansion, not required for the first pass.

Q-107's audit note (2026-06-12) already flags the kind of prose this pass exists to consume — distinctive material/craft traditions that currently exist only as prose and have no catalog counterpart:

System index.md location Tradition
GJ-4053 (Dzimbahwe) lines 22-26 granite dry-stack, mortarless masonry
GJ-1248 line 30 peat-wood carving
GJ-285 lines 28-38 basalt crater farms
GJ-588 lines 30-34 structural timber

None of these four appear as a distinct template in the shipped 28 (see the worked example, §8).


4. Two-stage flow

4.1 Stage 1 — candidate surfacing (per-system read → shortlist)

This differs from the naming pipeline's stage 1 (pick one of N predefined cultural registers). Here Gemma performs signal detection, not classification: for each system, it reads the index.md prose plus a compact context block listing the existing catalog's tag + label + cultural_description (28 entries — enough to ground "already covered" but short enough to fit context), and is asked one question — does this system's prose describe a coherent, distinctive architectural/material tradition not already represented in the catalog, and if so, name it and quote the grounding passage.

The expected answer for most systems is no — most systems should draw from an existing corridor baseline or heritage template, not spawn their own. A system earning a "yes" needs a real, holistic, catalog-worthy signal (the Dzimbahwe stone tradition — not "this system likes blue paint").

Output: a shortlist of (system_id, candidate_name, grounding_quote) tuples, much smaller than the 302-system corpus.

4.2 Stage 2 — structured proposal (shortlist → schema)

For each stage-1 hit, a second call turns the flagged prose into the structured proposal (§5). Candidate allow_tags/block_tags are drawn only from the catalog's existing shared material palette (the ObjectTag vocabulary enumerated in the catalog file's header comment — walls, roofs, facade, street lists) so proposals stay pre-constrained to compatible vocabulary. Minting a genuinely new ObjectTag is a separate, human-only vocabulary-authoring step (T-977(d), Miri+Araminta co-maintained) — Gemma proposes combinations of existing tags, never new tags.


5. Structured output schema — text fields only

{
  "tag": "granite_drystack_vernacular",
  "label": "Granite Dry-Stack Vernacular",
  "cultural_description": "...",
  "candidate_allow_tags": ["stone_wall", "..."],
  "candidate_block_tags": ["glass_curtain_wall", "..."],
  "source_system": "GJ-4053",
  "grounding_quote": "..."
}

Gemma never sets a numeric or eligibility field. This is not a scoping choice made for simplicity — it's forced by the shape of the guardrails: the catalog-wide CI checks (V-TT-01: ≥5 eligible templates per BulkClass post-gate; V-TT-02: no template >60% of its eligible pool's weight, tooling/economy-db/economy_import/traits.py) are properties of the whole catalog, not of any one template in isolation. A per-template extraction pass, looking at one system's prose, has no way to know whether adding this candidate at base_weight = 9000 would push some BulkClass's weighted pool over the 60% line — only a human with the whole catalog in view can safely assign eligibility gates and weights. This is the same channel-separation principle D-232/D-233 already establish for economics vs. culture (economics hard-gates what a building is; the catalog only weights how it's characterized), extended one level further:

Field(s) Set by Why
tag, label, cultural_description Gemma, human-reviewed text extraction is the point of this pass
candidate allow_tags / block_tags Gemma, human-corrected pre-constrained to the existing ObjectTag palette
corridor_pool, geographic_sector Miri needs whole-pool view (baseline vs. heritage; D-237 heritage-value alignment)
bulk_class_gate, production_ubiquity_gate, min_prosperity_bps Miri hard-gate eligibility — guardrail-sensitive
base_weight, weight_mods Miri (Nigel spot-checks V-TT-02) pool-weight math the guardrails enforce
zone_affinity Miri district fit, informed by cultural_description
era_scope Miri maintenance/wear scope (D-217), not a tech ladder
visual_bundle Araminta D-235 fallback-hierarchy authoring

Field-naming note: D-232's decision text names these fields allow / block; the shipped catalog TOML and traits.py use allow_tags / block_tags (T-1005 implementation). This doc uses the shipped names throughout, since that's the schema any accepted proposal has to match — the decision record's prose is the one out of step here and could use a tidy-up note pointing at the implemented names.


6. The human-review gate

Reviewers: Miri (cultural meaning, eligibility, weights — same role as T-1005), Araminta (visual_bundle), Nigel (spot-checks CI guardrail math on request; not required per-proposal since the importer re-validates automatically at bake time — see §7).

Acceptance criteria, applied to every candidate before it becomes a [templates.*] stanza:

  1. Genuine distinctivenesscultural_description reads as a coherent, holistic relationship between wall/roof/facade/street (never an axis-scramble), and isn't a restatement of an existing template under a new name.
  2. Dedup against the existing catalog — check both the tag (mechanical, unique key — traits.py will already reject a collision) and the label/cultural_description (semantic — a near-duplicate cultural voice under a different label doesn't earn a second template; fold the distinguishing detail into an existing template's allow_tags instead, or reject outright).
  3. Vocabulary compatibility — candidate allow_tags/block_tags reference only tags already in the shared ObjectTag palette; a genuinely new material/facade/roof tag routes to the separate vocabulary-authoring step, not this pass.
  4. Corridor fit — Miri assigns corridor_pool + geographic_sector; "corridors are tendencies, not borders" (D-232) — a candidate doesn't need to be exclusive to one corridor to be accepted as a heritage sub-pool entry.
  5. CI guardrails as a post-authoring gate, not a pass/fail on Gemma's output. Once Miri has assigned base_weight/weight_mods/gates, make regen-db re-runs populate_trait_templates() (tooling/economy-db/economy_import/traits.py), which re-validates V-TT-01 and V-TT-02 across the whole catalog, not just the new entry. A proposal can clear 1-4 and still get rebalanced or rejected at bake time if it happens to push a BulkClass pool over the 60% line — that's expected, not a process failure, and it's exactly why Gemma never sets these numbers itself (§5).

Nothing is committed to architecture_trait_catalog.toml until it clears 1-4 and a make regen-db / make check-systems-db dry run comes back clean.


7. Flow into the catalog

  1. Stage 1 + 2 produce N candidate proposals (one JSON object per stage-1 hit that survived stage 2).
  2. Miri reviews against §6's criteria 1-4; rejects, merges, or accepts each.
  3. Accepted proposals become new [templates.<tag>] TOML stanzas, hand-added to wiki/economics/architecture_trait_catalog.toml in the matching section (the file's existing sections: A cross-corridor pool, B per-corridor baseline, C heritage sub-pools) with Miri's numeric fields filled in.
  4. Araminta adds the matching visual_bundle, including the D-235 fallback parent for every new specific tag (a new template ships with a working generic-parent fallback from day one, same as the shipped 28).
  5. make regen-db bakes the change via import_economics.pytraits.py:populate_trait_templates(); V-TT-01/V-TT-02 re-run automatically (§6.5). make check-systems-db confirms the meta stamp.
  6. Commit the TOML change and the regenerated systems.db together — the ordinary asset-pipeline rule (.claude/rules/asset-pipeline.md): edit sources, regen, commit both.

No path lets this pass write to systems.db directly. The golden rule is unchanged: edit sources, never the DB.


8. Worked example — GJ-4053 (Dzimbahwe)

  • Stage 1 flags GJ-4053: the prose describes a 400-year granite dry-stack tradition ("great house of stone... dry-stacked without mortar") and a distinct sculptural tradition, neither represented by an existing template. The closest existing templates — civic_monumental (stone_wall, vaulted_roof, colonnade, institutional/ceremonial) and the south_reach heritage pool (iberian_hacienda, atlantic_creole_maritime — Iberian and Afro-Atlantic, not southern-African stone) — don't fit.
  • Stage 2 produces a proposal along these lines: tag: south_african_stone_vernacular, label: "Southern African Stone Vernacular", cultural_description drawn from the "great house of stone" language, candidate_allow_tags: [stone_wall, ...], candidate_block_tags: [glass_curtain_wall, ...], source_system: GJ-4053, plus the grounding quote.
  • Miri reviews: distinct from the existing south_reach heritage pool (criterion 2 passes) — accepts as a new south_reach heritage sub-pool entry, assigns min_prosperity_bps, weight_mods (e.g. founding_age.ancient), zone_affinity (Residential-leaning, matching the domestic/sculptural framing).
  • Araminta authors the visual_bundle (wall = stone_wall, roof/facade choices, fallback chain to generic_wall/generic_roof).

The other three systems Q-107's audit note flags are useful for calibrating stage 1's expected "no" rate, not as three more automatic templates: GJ-285 (basalt crater farms) and GJ-1248 (peat-wood carving / Moorlesen) are agricultural and craft-object traditions, not building-architecture ones — a correctly-run stage 1 would surface them as signal for the materials/crafts side of Q-107's consolidation (feeding atlas_body_trait_bias hero-pin authoring against an existing template), not propose a new [templates.*] stanza. GJ-588 (structural timber) is architecture-adjacent, but its outer north_reach setting already has a timber-forward baseline (north_anglo_frontier: timber_wall, pitched_roof, regular_facade) — the expected stage-1 outcome there is "no distinct tradition, corridor baseline already covers it," which is the correct, common result, not a missed opportunity. Only GJ-4053's stone tradition clears the bar because no existing template — baseline or heritage, any corridor — matches it.


9. Open follow-ups

  • Q-107 (open) — source-location consolidation. If it resolves, update §3's file glob; nothing else in this doc changes.
  • Filing the actual pass — this doc defines the methodology only (T-992's scope). No ticket yet schedules an execution run; file one under T-977 when the catalog is judged to need growth toward the 40-45 target.