Commit Graph
3011 Commits
Author SHA1 Message Date
jpmschweitzerandClaude Opus 5 6daaa2235f docs(governance): D-263 — domains mirror the implant apps, and atlas is one ladder
reach's domain names should not be a fresh taxonomy. Where the game already
presents something to the player, the CLI takes that name and that shape: what
you browse in-game is what you generate and inspect from the terminal.

That splits domains in two. atlas, ledger and wiki mirror implant apps and
follow their structure. check, validate, godot, visual, jobs and dev mirror
nothing — no app exists for a lint gate, and inventing a player-facing framing
for one would be worse than having none.

The first consequence corrects a contradiction rather than a preference. D-191
already says "Atlas is the star map extended downward, not a separate app —
implant/map at different zoom levels", four rungs from Reach map to regional.
The domain map had atlas, starmap and planet as peers, which would have
presented as three unrelated things what the game presents as one descent.
Generation now nests by rung; authoring and inspection verbs stay flat on
atlas, because they act on the whole thing rather than a rung.

The second is a rename with the same reasoning: db becomes ledger, after the UI
component that will aggregate economics — markets, wealth, transactions, the
economic counterpart to what the Atlas offers for topography. db named a
storage layer nobody looks at.

One caution recorded because the words collide. D-191's MVP criterion 7 says
"Atlas is read-only (no verbs execute from map)". That governs the app. The
atlas tooling writes — it commits proposals, mutates fields, syncs the wiki —
and a later reader must not take the app's constraint as licence to delete the
authoring verbs.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-02 14:15:03 +02:00
jpmschweitzerandClaude Opus 5 a384ec0c7c feat(config): T-1283 — the godot and visual domains
reach godot parse-sweep / cold-parse, reach visual diff / blank-check /
thumbnail. Five scripts retired, and the callers rewired — tests/run-visual
invoked three of them by path at four sites, which is a wider blast radius than
the make targets were.

The godot pair were grep pipelines encoding five hard-won lessons as comments
nobody could test. They are Python filters now, with the reasons attached, and
the engine invocation is a guarded exec. Verified on the real client: 229
scripts, clean.

Their three not-ok states stay distinct, because only one is a verdict about
the code. An engine that crashed or is missing is not a parse failure —
reporting it as one blames the tree for a broken toolchain. A sweep that
emitted no completion marker checked nothing, and zero errors from a check that
never ran reads as clean, which is the false-green the sweep exists to close.
The deliberate asymmetry between the two checks is preserved and documented:
cold-parse filters "Cannot infer the type", the sweep does not, because that
suppression is why cold-parse stayed silent about a helper that genuinely does
not parse.

All three visual scripts carried the same root bug as validate-checklist:
Path(__file__).parent.parent, correct at tooling/ and two levels too deep at
tooling/domains/visual. Fixed during the move rather than after, having learned
that it fails silently — paths resolve to nothing, the work appears to have
nothing to do, and the tool reports success. Three domains now where that would
have shipped a false pass.

Two bugs my own transformation introduced, both found by running rather than
reading. Multi-line print(..., file=sys.stderr) became console.event(...,
file=sys.stderr), and console puts unknown kwargs into the payload — a file
object would have reached json.dumps at the exact moment something was already
being reported as an error. And the replacement script wrote escaped quotes
into three files. Mechanical transformations need mechanical verification.

sys.exit removed from four sites: a service must not end the process.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-02 13:59:29 +02:00
jpmschweitzerandClaude Opus 5 7f20bd303b feat(config): T-1282 — the validate domain, and a move that broke a root
reach validate content / checklist / ron / name-collisions. The three old
scripts are retired, their make targets with them.

Print statements go through the logging sink rather than a collector. The
validators emit their findings as console events as they run, so a long content
validation streams instead of going quiet and dumping at the end — the message
strings and their order are unchanged, only the destination. That also
satisfies the conformance rule forbidding print() in the package, which is what
forced the question.

validate-ron was three languages deep: bash dispatching on a flag, a Python
heredoc doing collision detection, cargo run for schema validation. Logic
embedded in a shell string cannot be imported, tested, or found by anything
that indexes Python, so it became Python; the cargo call became a guarded exec.
It also split into two verbs, because --check-name-collisions answered a
different question from the default path: whether the SET of cultures is
coherent, versus whether ONE file is well-formed.

The move broke something, quietly, which is the point of doing these one at a
time. validate-checklist computed ROOT as Path(__file__).parent.parent — the
repo root while it lived at tooling/validate-checklist, and tooling/domains
once moved. Both its schema and gauntlet paths silently repointed at nothing,
the gauntlet directory "did not exist", and it reported success having checked
zero files. Caught by running it beside the original: old exit 1, new exit 0.
Now config.repo_root(), and load_schema raises ReachError instead of calling
sys.exit, which a service must not do.

Parity on the live tree: content reproduces the original byte for byte
including its counts, name-collisions likewise. Tests pin what those runs
cannot reach — the detection path, since the repo currently has no collisions,
and the argument errors.

Two things found and left alone: validate-content FAILS on the live tree with
13 missing schemas, pre-existing and unrelated to this port; and the ticket's
claim that validate-content sits in the pre-commit hook is wrong — that hook
runs only check-fact-ids and pql decisions validate, so there was no shared
edit to coordinate.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-02 12:48:56 +02:00
jpmschweitzerandClaude Opus 5 cb5d3f1335 refactor(config): T-1281 — the check domain retires its five scripts
All five gates are ported, tested against their failure paths, and the
originals are gone. reach check is the only way to run them.

Parity first, then deletion. Every case in test_check.py began as a parity case
running the new implementation beside the script it replaced; that evidence is
in the ticket. With the scripts retired there is nothing left to compare
against, so the assertions become the spec and the file drops its "_parity"
name. A parity test is scaffolding with a defined lifetime — keeping one after
its subject is deleted would mean keeping the subject alive to be compared
with, which is the opposite of a migration.

Two gates could not be parity-tested in a fixture at all, and both reasons are
findings rather than obstacles. canvas-version: canvas_sources globs from a
__file__ root while the service resolves git through config.repo_root(), so a
fixture would diff one tree and glob another — real history is used instead,
including two genuine instances of the regression the gate exists to catch.
systems-db-stamp: generator_sources raises at IMPORT time when the economy-db
tree is absent, so the old script died before reaching any logic in every
fixture. The ported service imports it lazily and after the absent/unstamped
checks, which is exactly why those states are testable now and were not before.

Hooks rewired: pre-commit runs reach check fact-ids, pre-push runs the other
four. Both pass --no-input, because a hook has no TTY and a prompt there does
not wait, it crashes. Both guard on `command -v reach` and skip with a message
rather than blocking every commit on a missing tool.

Make targets are RETIRED, not wrapped, per the D-263 split — with the mapping
left as a comment where they used to be. Wrapping would leave two ways to
invoke each gate, and reach --help would stop being the answer to "what tooling
exists" while the Makefile remained a competing index. pre-pr-validate and
pre-pr-content keep their orchestration role and lose the individual target.

Sprint archives and workshop notes still name the old paths and are left alone:
they record what was true when written.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-31 20:41:33 +02:00
jpmschweitzerandClaude Opus 5 793a5239cd test(config): T-1281 — canvas-version parity, using its own history as fixture
Five cases, including the failure path that matters — this is the gate with
five shipped regressions behind it.

A fixture repo does not work here, and finding out why exposed a real
inconsistency in the port: canvas_sources globs from a __file__-derived root,
so under SR_REPO_ROOT the service would diff the fixture while globbing the
real tree. Git goes through config.repo_root(); the registry does not. Harmless
in production since they are the same repo, but it is the same
no-root-override asymmetry the domain map noted about the old scripts, now
inside the new code. Not fixed here — making it dynamic means restructuring six
module-level constants in a module the old script still imports.

Real history is the better fixture anyway: both implementations see identical
input, nothing is mutated, and nothing can drift from the thing it models. Two
of the cases are genuine historical instances of the regression this gate
exists to catch — T-1237 and T-1194 both changed canvas generation and were
bumped only after the fact. The history that produced the check, used as its
own test.

The test also asserts no changed file is dropped from the failure message. That
list is the actionable half; "something changed" without saying what leaves the
reader to re-derive the intersection by hand.

Proven to fail by truncating the touched-file set, which reported both the
exit-code divergence and all three omitted filenames by name. Each case also
asserts the OLD script still behaves as the case claims, so a rewritten history
would say so rather than silently checking nothing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-31 20:07:26 +02:00
jpmschweitzerandClaude Opus 5 a3cbc478a0 feat(config): T-1281 — canvas-version, and typer's other rich path
All five gates now live in the check domain. canvas-version produces
byte-identical output to the original on the live tree.

It is the first real consumer of core/process.run. The git calls pass
check=False deliberately: a git failure here is not an error to report but a
signal that there is nothing to compare, since a fresh clone with no remote is
a legitimate state rather than a broken one. The argv-list and missing-binary
guards still apply.

Its two skips are kept distinct from its pass. NO_BASE and DIFF_FAILED exit 0,
as does CLEAN — but only CLEAN means the gate actually looked at something.
Collapsing them would hide a gate that had silently stopped running, which for
this check in particular is the exact failure it exists to prevent.

Found a second rich path while a NameError was rendering as a full-width
box-drawn traceback: typer's pretty-exception handler is a different mechanism
from rich_markup_mode, and setting one does nothing about the other. Same log
pollution T-1259 thought it had closed, arriving through another door and
landing in the worst place — a hook log at the moment something has already
gone wrong. pretty_exceptions_enable=False now on the root and on every domain
built by cli.domain().

test_canvas_version_check.py moves with the code it guards. It had been loading
the extensionless script through a SourceFileLoader and reaching canvas_sources
by sys.path insert, both only because tooling/ was not importable. Second
instance of that debt evaporating on contact. What it asserts is unchanged,
which is the point: diff_has_version_bump was kept pure in the port so its six
properties still hold without constructing git history.

Also restores an import the check router dropped in T-1267 when it moved to
cli.domain() — caught by running the command rather than by reading it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-31 18:22:25 +02:00
jpmschweitzerandClaude Opus 5 05bf1732d4 feat(config): T-1281 — dataflow-graph and systems-db-stamp join the check domain
Both were already Python, so these are moves rather than rewrites, and both
produce byte-identical output to their originals on the live tree with the same
exit codes.

The E402 debt evaporated on contact, which is the first concrete evidence for
T-1274's premise. check-systems-db-stamp reached generator_sources through a
sys.path.insert and a noqa suppression, because tooling/ was not a package. It
now imports as `from tooling import generator_sources` — no hack, no
suppression.

The stamp gate's six failure modes are preserved as a StampState enum rather
than collapsed into pass/fail, because they carry different remedies and one
carries a different exit code: UNSTAMPED exits 2 while every other failure
exits 1, and the pre-push hook has relied on that distinction since T-857.

One deliberate behavioural difference, flagged rather than hidden: the old
stamp script was silent on success unless given --verbose, and the new one
always prints its verdict. No fact is lost, so parity holds, and it makes the
gate consistent with client-version and dataflow-graph which both always print
— the old script was the odd one out. Its per-command --verbose gives way to
the global one, which is the consolidation this initiative is for.

Also corrects a claim in the ticket itself: check-dataflow-graph.py does not
parse git output, it globs the filesystem. Only check-canvas-version parses
git, so only that fixture needs a real repo.

Still open and recorded as such: check-canvas-version, and parity tests for
these two — both were verified side by side on the live tree, which proves the
happy path and nothing else.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-31 17:55:28 +02:00
jpmschweitzerandClaude Opus 5 afe2328182 feat(config): core/process.run — one guarded exec, and the logic in Python
Clarifies the rewrite decision to what it actually meant: rewriting the bash in
Python does not mean reimplementing the operating system. A guarded exec is the
right answer for rustup, curl, unzip, git, godot, blender. What must become
Python is the LOGIC — which version is wanted, whether it is already present,
what the output means, what to do when it fails. The test of a correct port is
not whether it calls anything external, but whether the decisions can be
exercised without performing them.

Delivered ahead of the remaining ports because every one of them needs it.
core/process.run is the single sanctioned exec, and each of its guards exists
because a per-domain subprocess call is precisely where that guard goes
missing:

- An argv list, never a shell string. A string is rejected outright rather than
  helpfully split, since the helpful split is the vulnerability.
- shell=False always.
- A non-zero exit becomes a ReachError naming the command, carrying its output,
  and preserving its exit code — not a CalledProcessError traceback at someone
  who wanted to know the next step.
- A missing binary reports what to install. FileNotFoundError names the path
  that was not found, which is the less useful half of the answer.

All four verified against real commands, including a genuine git failure
relaying exit 128.

A conformance invariant keeps the door single: nothing outside core/process.py
may import subprocess or call os.system/popen/exec*. Proven to fail by
importing subprocess into a domain service.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-31 17:48:22 +02:00
jpmschweitzerandClaude Opus 5 b88791705c feat(config): T-1281 — check fact-ids, the first bash rewrite
89 lines of grep/sed pipeline become a service returning a FactIdCheck and a
router that renders it. Parity on the live tree is exact: both implementations
print "check-fact-ids: OK — 6 references validated against 61 canonical facts"
and exit 0. The matching counts are the real evidence — a line-matching regex
that differed from the grep chain even slightly would move 6 or 61.

Kept line-matched rather than YAML-parsed on purpose. Parsing properly would
change which lines count: anchors, merge keys and multi-document files would
start contributing ids the old check never saw. That is a different check
wearing the same name, and a port is not the place to make it.

Three parity cases: ok, unknown fact_id, and the advisory mode where the
catalogs hold no definitions and the gate deliberately exits 0 — failing every
commit until they are populated would teach people to bypass the hook, and a
gate people route around protects nothing.

Proven to fail by removing the entity-attributes.yaml exclusion, and caught in
a way worth noting: not by the assertion aimed at it, but by the advisory case,
where including that file made the catalog non-empty so the new implementation
enforced while the old stayed advisory. A real behavioural divergence, surfaced
by exit code.

Retirement waits for the whole domain, per the per-domain rule — three gates
remain. It also resolves a tension: the parity test copies the old script into
its fixture, so deleting the script early would delete the test's own subject.
A parity test is scaffolding with a defined lifetime.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-31 17:38:54 +02:00
jpmschweitzerandClaude Opus 5 e359cfa841 docs(governance): D-263 — half the tooling is bash, and it gets rewritten
Found on starting the first port: 17 of the 33 tooling executables are bash,
about 900 lines. Both this record and the domain map had assumed a Python tree,
so those are rewrites rather than moves — a materially larger epic than T-1250
was written for.

Decided: rewrite them, do not wrap them. Wrapping would achieve one door while
leaving half the CLI surface outside the contract — no @command, no remedy on
failure, no streaming, no testable service. reach --help would then list verbs
that behave differently from the ones beside them, which is worse than two
doors, because the inconsistency is invisible until something fails.

The cost lands unevenly and the record says where. The grep-pipeline scripts
compute verdicts and gain most from becoming services. The environment scripts
— install-godot, install-rust, worktree-setup — gain least and carry the most
regression risk, because downloading a specific Godot build or driving rustup
is awkward to exercise in a gate. For those, port the decision logic into a
testable service and keep the irreducible external calls behind core/process: a
rewrite that cannot be tested has to be trusted instead, and trusting an
installer is how a working environment becomes an unreproducible one.

The domain map gains the inventory by shape, and a rule that every per-domain
ticket states which of its sources are bash — since that is what turns a port
from mechanical into a rewrite needing its own parity evidence.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-31 17:33:42 +02:00
jpmschweitzerandClaude Opus 5 de69bd70b4 feat(config): T-1280 — job log retention, and the corpse that would never die
Pruning happens at spawn time rather than on a schedule: a retention pass that
depends on someone remembering to run it is one that silently never happens.
reach jobs prune is the explicit escape hatch for reclaiming space now.

The cap was measured rather than guessed, which is why this ticket ran last. A
chatty short job writes ~1.8 KB across its three files, so 100 jobs is
single-digit megabytes even if a generator emits per-body progress — inside
.cache/, where being wrong costs disk and never data. SR_JOB_KEEP overrides it.

The interesting part is what "a running job is never pruned" has to mean. Not
"the file says running" — a process killed outright never updates its own
status, so that reading would make every crashed job immortal. Those are
exactly the ones that accumulate, so the naive rule produces the opposite of
retention: the only logs that never go away are the ones nobody wants. The
check consults the process table instead.

Verified both directions. Live, a running 30-second job survived a prune to
--keep 1. Pinned with a fixture holding a finished job, a corpse (record says
running, pid gone), and a genuinely live one — asserting the live one survives
and the corpse does not. Proven to fail by dropping the liveness check.

One false alarm worth recording: my first live test looked exactly like the bug,
showing a running job pruned. It was not — my commands ran two minutes apart, so
the "20-second" job had finished long before. The test was invalid, not the
guard. A timing-sensitive check across separate shell turns proves nothing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-31 17:27:25 +02:00
jpmschweitzerandClaude Opus 5 3b211a5450 feat(config): T-1279 — a detached failure reaches its caller
The non-negotiable from D-263, pointed at its worst hiding place: a foreground
command that swallows a failure at least does it in front of someone, while a
background runner that reports "started" and loses the failure does it where
nothing is watching.

Testing the two timing cases the ticket names — fails before the parent exits,
fails long after — needs a command slow enough to tell them apart, and every
verb in reach finishes in milliseconds. So `reach dev selftest` exists: emits
progress for N seconds, then optionally fails with a chosen code. A genuine
diagnostic rather than a test hook, in the dev domain the map already planned,
and the only way to answer "does streaming work here, can I tail it, does a
failure survive detach" by observation instead of argument.

The slow case is the one that proves the design. --detach returned in 75ms
while the child ran six seconds, so the parent was demonstrably gone long
before the child failed — and wait still relayed exit 7. That is the half of
the recording path only this case reaches, and why T-1277 moved completion
recording into the child.

Also pinned: --detach exits 0 for starting and SAYS "not succeeded" in words,
which the test asserts on rather than trusting the code to be read correctly;
a failed job nobody waited on shows as failed in jobs list; and every event a
detached job emits carries its job id.

Closed T-1278's open gap in passing — jobs log --follow had never run against a
genuinely long job because none existed. It now has: attached mid-flight,
streamed the remaining steps live, and caught the final verdict after the job
ended.

Proven to fail by making effective_exit_code always return 0 — the trap itself.
Both timing cases failed by name.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-31 17:14:23 +02:00
jpmschweitzerandClaude Opus 5 6f08cc9156 feat(config): T-1278 — the jobs domain, and typer.Exit is not a SystemExit
reach jobs list / status / log --follow / wait. A domain rather than core/,
because these verbs carry logic and state: they reconcile recorded status
against process liveness, tail a file from an offset, and relay an exit code.

Found a latent bug in already-committed code before building on it. typer.Exit
is a RuntimeError, not a SystemExit, so @handle_errors caught it like any other
unexpected exception: `raise typer.Exit(3)` inside a decorated command printed
"unexpected Exit: 3" and exited 1, silently discarding the requested code.
Nothing hit it because the check router had been converted to ReachError — but
jobs wait needs exactly this and it is what anyone would naturally write. Added
core/errors.ReachExit as the sanctioned control-flow exit, passed straight
through with no verdict. ReachError would have been wrong twice: a failure
verdict for a command that worked, and a demand for a fix= where there is no
remedy.

Reconciliation proved out on a real corpse rather than a simulated one — the
job stranded by the T-1277 bug, status "running" with its process long gone,
now reports as died. DIED is derived, never recorded, because a process killed
outright cannot write its own ending. It relays 137, never 0: a died job has no
exit code of its own and borrowing success points the exit-0 trap straight at
whatever gated on the run.

Second UTC bug of the same family as T-1276's: jobs list reported a job started
minutes earlier as running for 133m, because _parse used mktime on a UTC stamp
and silently added the offset to every duration.

console.render() is public now, so jobs log replays stored events through the
same path a live run prints them — a second renderer would drift, and the
divergence would surface exactly when someone is reading a log to find out what
went wrong.

test_jobs.py closes the gap T-1257 named: D-263 claims services are callable
without a CLI round trip, and nothing had ever demonstrated it, which left the
layering as unverified decoration. Every test here calls the service directly.

Not yet exercised, and said plainly: log --follow against a genuinely
long-running job. Nothing in reach runs long enough to tail yet. The offset
mechanics underneath are tested; the live loop waits for a slow domain.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-31 17:02:36 +02:00
jpmschweitzerandClaude Opus 5 c924b0934e feat(config): T-1277 — detach, and a failed job that looked busy
core/process.py spawns a child that outlives its parent: its own session, so a
signal to the parent's group or a timeout kill does not take the work with it;
re-execing reach by BARE NAME, because an absolute path would freeze the child
to whichever checkout was current at spawn time and silently run the wrong
source after a repoint; and streams kept separate exactly as in the foreground,
events to <id>.jsonl and real output to <id>.out.

Testing a case the ticket did not name found a real hole. Recording completion
inside @command looked right and was wrong: a child that fails BEFORE any
command runs — bad arguments, an unknown verb, an import error — never reaches
that decorator. `reach --detach check bogus` left its metadata reading
"running" forever with the process long gone. That is the exit-0 trap wearing a
new disguise and worse than the original, because a failed job that looks busy
sits somewhere nobody is watching, and a caller polling for completion would
wait indefinitely on something that failed in milliseconds.

So completion is recorded at the PROCESS's exit instead. main.py gains main(),
wrapping cli() in a single try/finally, and the entry point moves to
main:main. Every exit path now passes through one place. Removed from @command
rather than left in both — two writers of one field is how they drift.

Verified on three paths: success records done/0, a real drift failure records
failed/1, and the parse failure that exposed the hole now records failed/2.

One narrow conformance exemption, with its reason inline so it does not read as
an oversight: the no-domain-imports-core.jobs invariant fired on main.py,
correctly by its letter and wrongly by its purpose. main.py is not a command;
it is the entry point, and it already owns --detach.

Still open, and carried to T-1278: a child killed outright cannot record
anything, so jobs list must reconcile against process liveness rather than
trusting the file.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-31 16:46:19 +02:00
jpmschweitzerandClaude Opus 5 5d83e1d2eb feat(config): T-1276 — every invocation is a job, carried ambiently
Streaming as a decorator, first half. Each invocation of reach gets an id and
every event it emits is tagged with it, which is what will let a detached run's
log be read back and what correlates the lines of a run that streamed for nine
minutes. No command signature changed and no command imports core.jobs — that
is the point, per the D-263 amendment: a command must not know jobs exist,
because the alternative is call-site discipline wearing a different hat.

A ContextVar rather than a module global. A global is correct only until
something runs two invocations in one process — which a test harness or a
future batch verb does immediately, and which would then interleave two jobs'
events under one id with nothing reporting an error.

The job context is the OUTERMOST wrapper, and it has to be. @logged emits from
its finally and @handle_errors emits its verdict while unwinding, so a context
established inside either would already be reset by the time the two most
important events are written — leaving them the only untagged lines in the log,
and they are precisely the ones a detached run gets read back for.

Fixed in passing: the job id used local time while every event's ts is UTC, so
an id read 155327 beside its own first log line reading 13:53:27. Two hours
apart reads as a logging bug every time someone correlates them by eye.

New conformance invariant — nothing outside core/ may import core.jobs. My
first version of it inspected only the module path, so it missed
`from tooling.core import jobs`, where the name is in the import LIST and which
is the form anyone would actually write. It passed while checking nothing.
Rewritten to catch all three reachable forms and then verified by committing a
real violation, which it named by file and line.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-31 15:59:39 +02:00
jpmschweitzerandClaude Opus 5 b5beda0df7 feat(config): T-1275 — bare reach is discovery, so it exits 0
Bare `reach` and bare `reach <domain>` printed help and exited 2, Click's
usage-error convention. Running reach with no arguments is the DISCOVERY
action — it is how the tool gets learned from nothing — and a caller that
branches on exit status would read its own onboarding as a failure. Now they
exit 0.

D-263's exit-code contract is untouched: it governs failures, and printing a
command list is not one. Verified across the whole matrix, because this change
flirts with the exit-0 trap that record opens with — bare 0, bare domain 0,
--help 0, unknown domain 2, unknown verb 2, real failure 1. All five are now
pinned as a sixth conformance invariant, since an exit code regresses silently
and nothing else would notice. Proven to fail by putting the 2 back.

The implementation also collapses a duplicated class. core/cli.py holds
ReachGroup with both shared behaviours — no-args-prints-help-and-exits-0, and
unknown-name-enumerates — and LazyDomainGroup now extends it instead of
subclassing TyperGroup directly, keeping only the laziness and the
domain-specific wording. The enumeration logic previously existed twice in
slightly different forms, which is how the root and the domains would have
drifted into disagreeing about their own conventions.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-31 15:49:08 +02:00
jpmschweitzerandClaude Opus 5 91e25a3e7a docs(governance): D-263 — the primary user is an agent, and that changes things
Stated plainly because the record was quietly assuming otherwise: Jeroen runs
make and plays the game; the caller typing reach all day is Claude.

It resolves several arguments in the opposite direction from human-CLI
instinct. --help is a discovery mechanism rather than documentation, since it
is how the tool gets relearned from nothing every session — which makes the
domain list and closed-set enumeration load-bearing rather than polish. Output
volume is a context cost, so quiet-by-default is right for a better reason than
not spamming a hook. Latency matters less than legibility: nobody drums their
fingers at 300 ms, but a multi-minute silence is expensive because a wedge is
indistinguishable from work. And errors that name the next command are the
highest-value requirement here, because the reader is usually deciding what to
run next — "no" costs a whole exploratory turn.

One correction follows directly. D-263 had scoped streaming to "callers with no
escape — a human terminal, a Makefile, a git hook", reasoning that Claude
Code's background mode already solved the timeout for agents. That got the
audience backwards. Background mode solves the timeout and nothing else: it
returns when the process exits, so a nine-minute wedge still looks exactly like
nine minutes of work. Streaming is what makes a long run legible while it runs,
and reattach is worth most to the caller whose attention is not continuous.
Both are primary-user features, not fallbacks.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-31 15:41:29 +02:00
jpmschweitzerandClaude Opus 5 6b31111cd2 docs(governance): D-263 — make and reach split by kind, streaming as a decorator
Two decisions taken before the 160-file move, because both change what the
move produces.

The Makefile has 84 targets and is today's front door, so "one CLI for all
repo tooling" was not yet true. The split is by what a target DOES: make keeps
genuine build and test orchestration, and targets that are really tooling
wrappers are retired in favour of reach verbs — retired, not wrapped. A
wrapper leaves two ways to invoke every tool, and then reach --help stops
being the answer to "what tooling exists" because the Makefile is still a
competing index. Two doors is the condition this record exists to end, so
keeping both would defeat it while looking like caution.

Streaming becomes a decorator rather than an API commands call. @command
already wraps every invocation, and that is exactly the seam where job
identity, progress correlation and detach belong: the decorator assigns the
job id, tags the events, and forks on --detach. A command must not know that
jobs exist. The alternative — each command opening a job and remembering to
close it — is call-site discipline wearing a different hat, and it fails the
same way the fortieth command into a porting session, with the failure
vanishing from the log and nothing to indicate anything is missing. Logging
and error handling are decorators for this reason; streaming is the third
cross-cutting concern, not a special case.

Consequent resequencing: T-1264 lands before the T-1250 move, so every ported
command arrives already streaming. Old scripts now retire per domain as each
port passes its parity test, rather than in one sweep at the end — a
continuous shrink, instead of months where every tool exists twice and an edit
can land in the dead copy.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-31 15:38:26 +02:00
jpmschweitzerandClaude Opus 5 91a57b8304 docs(config): T-1271 — the domain map, before anything moves
Every Python file and executable in tooling/ assigned to one of 15 domains,
with the ambiguous cases carrying their reasoning. The per-domain port tickets
are written from this rather than guessed, so their boundaries do not have to
be renegotiated halfway through a 160-file move.

Three things counting turned up that reading would not have.

The Blender carve-out is 35 files, not the 13 visible at top level — 22 more
are inside garment-fit/, which turns out to be a payload directory wearing a
domain's name. The epic said 35 and an earlier survey of mine said 14; the
epic was right. That is not cosmetic: `character` is a far smaller domain than
directory sizes imply, and a port ticket written from the listing would have
been wrong about both it and the carve-out.

The "28 singleton prefixes" were an artefact of splitting filenames on the
first token, which scattered coherent families — sculpt-star-map,
tune-star-map-topology and generate-star-map* are one group counted as three
orphans. Counting families instead, the genuinely ambiguous set is small
enough to enumerate with reasons.

And tooling/db/ is misnamed: it holds the audio/image/Trellis connectors and
wiki_sync, while the actual database work is in economy-db/. Naming a domain
after that directory would have carried the misnomer forward.

Judgment calls settled with reasons, since each sets a precedent. Registries
stay data rather than becoming verbs nobody would type. Gate tests do not
become a `test` domain implying a runner that does not exist. pql-migrate is
provenance — archived, not deleted and not importable. `pr` is a domain the
epic omitted, kept out of `dev` so dev does not become the drawer everything
ambiguous goes into. And `atlas` is overloaded across three unrelated places —
map data, terrain quality analysis, and systems.db index tables — which stay
with their owners rather than being collected into a domain whose only common
thread is a noun.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-31 15:26:55 +02:00
jpmschweitzerandClaude Opus 5 49fa6ada95 feat(config): T-1249 — the contract is a decorator, and now a test
Every non-zero exit names the command that would fix it, and still exits
non-zero. Both halves matter; the second is the one that gets lost, because a
tool that explains itself beautifully and exits 0 looks MORE correct while
having silently disabled its own gate.

core/errors.py holds ReachError(message, fix=) and @handle_errors.
core/logging.py holds @logged, emitting through console rather than a second
sink — one output path, so there is nothing to drift. core/command.py composes
them, and the order is load-bearing: handle_errors wraps logged, so the logger
sees the original exception. Inverted, every failure would be recorded as
"SystemExit" and the log would say nothing about what went wrong while looking
like it worked.

core/ raises SystemExit, not typer.Exit. A service must be callable from a
test, another service, or a future second front end, and an exception type that
only makes sense inside a CLI leaks the transport into every layer.

The check router is retrofitted off its hand-rolled verdict-and-exit pattern —
exactly the boilerplate this removes — and test_check_parity.py passes
unchanged across the retrofit. That test predates the decorators and pins exit
codes against the old script, so it is independent evidence, not a test tuned
to match new behaviour.

Unknown domains and unknown verbs now enumerate what exists instead of only
saying no. That needed a shared group class, which collided with "no typer
outside main.py and router.py" — resolved by sharpening the invariant rather
than breaking it, since its purpose is that a SERVICE never knows it was called
from a CLI. Transport now lives in main.py, router.py and core/cli.py; never in
service.py, schemas.py or helpers.py. The upside is that cli.domain() carries
the settings that were previously per-router decisions, including the
load-bearing rich_markup_mode=None that one forgetful domain could have undone.

test_conformance.py makes five invariants executable, AST-based rather than
grep. Scoped to the package, not the 123 legacy scripts — and deliberately so:
as T-1250 moves each script into domains/, it lands inside the scope and the
rules start applying automatically, so the test's reach grows with the
migration.

Proven to fail before being trusted: removing @command and removing a fix= each
produced a failure naming the file, the line and the reason.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-31 15:03:34 +02:00
jpmschweitzerandClaude Opus 5 1eb30a1460 chore(config): T-1263 — one permission rule for the whole tool surface
Bash(reach) and Bash(reach *) join .claude/settings.json beside the pql pair.
Two entries, not the one the ticket asked for: a rule ending in " *" does not
match the bare word, and bare `reach` is a real invocation now that it prints
the domain list. pql, make, cargo test and ruff check each carry a bare-form
entry alongside the wildcard for exactly this reason, and adding only the
wildcard would have left `reach` prompting while `reach check ...` did not.

This is the line Q-124 was actually filed about. Ten hand-written
Bash(tooling/...) entries each cover a single script and every unlisted tool
prompts; one command with subcommands is one rule covering everything. The ten
stay for now — the old scripts are still the working tools until T-1253.

On verification, since the ticket warned specifically against declaring this
done on the wrong evidence: real calls run clean, but that is NOT proof the
rule matched. The same calls succeeded before the rule existed — there was no
Bash(reach ...) entry in either settings file and no blanket grant — so the
session was already permitting them and the observation cannot distinguish "the
rule matched" from "the rule was never consulted". settings.json is read at
session start, so this cannot be self-verified from the session that wrote it.
Proof is a later session, in a prompting mode, where reach runs without asking.

One accepted limitation, documented rather than worked around: rules
prefix-match the whole command string, so an env-prefixed call like
SR_REPO_ROOT=... reach ... will still prompt. An environment override is a real
departure from normal invocation; the ordinary form is what needs to be
frictionless.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-31 13:52:45 +02:00
jpmschweitzerandClaude Opus 5 5cdb3e9327 feat(config): T-1262 — parity is facts and exit codes, not bytes
schemas.py becomes pydantic, so the reference domain is the normal pattern
rather than an exception carrying a footnote. Frozen: a result is a statement
about what was found, and nothing downstream should edit the finding on its way
to being reported. pydantic stays off the --help path — test_lazy_domains still
passes, which is precisely the assertion that it loads with the domain and not
with the CLI.

The acceptance criterion could not be met as written, and that is the finding
worth keeping. It asked for byte-for-byte parity with the old script; D-263 was
amended after this ticket to give reach a streaming model that puts the verdict
on stderr, while the old script writes its success line to stdout. Measured:
the text is byte-identical in text mode, only the stream differs. Matching both
would mean abandoning streaming or special-casing every ported gate.

So parity is redefined, and it is stronger than bytes where it counts: exit
codes match exactly, no fact the old message carried is lost, and failures name
a remedy as a structured field. That governs every port in T-1251, not just
this one, so it is in D-263 rather than only here.

test_check_parity.py runs three paths — ok, drift, missing file — through both
implementations and compares. It builds a throwaway fixture repo and copies the
OLD script into it, because that script resolves its root from __file__ and has
no override; the new command just takes SR_REPO_ROOT. That asymmetry is part of
why the port earns its keep. It also asserts the failing paths actually exit
non-zero, without which "the exit codes matched" would be vacuous for two
checks that both silently pass.

Proven to fail twice before being trusted. Once by accident: the first version
asserted the yaml version appears on every failing path, which the old script
does not report when the client file is missing — the test was wrong, not the
code, and it now derives expected facts from what the old output actually
contains. Once on purpose: mutating the router to drop a version made it fail
and name the missing fact.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-31 13:45:39 +02:00
jpmschweitzerandClaude Opus 5 f4cca69cab feat(config): T-1261 — reach is a bare name on PATH, in every context
`uv tool install --editable` puts the executable in ~/.local/bin rather than
.venv/bin, which is the difference between a command that works everywhere and
one that works only under an activated venv. Agents and git hooks never
activate one.

Verified in the three contexts that matter, with a negative control so the
passes discriminate: a stripped non-interactive shell, a REAL git hook process
(via git -c core.hooksPath ... hook run pre-push, not a simulation), and an
agent Bash call — all with VIRTUAL_ENV unset. With ~/.local/bin removed from
PATH the same check reports NOT-FOUND, so this is not passing because a venv
happens to be active.

Found a silent interpreter fork while doing it, which is this initiative's own
failure mode wearing a different hat. uv tool install without --python picked
CPython 3.11 for the tool environment while .venv and system python are 3.14 —
uv selects the lowest interpreter satisfying requires-python. reach would have
run on one interpreter and the test scripts on another, with different wheels
for numpy/scipy/PIL, and future 3.12+ syntax would break the tool while the
venv stayed green. PYTHON_VERSION now pins both.

make setup-venv is rebuilt on uv, per the T-1258 finding that it called
.venv/bin/pip against a venv that has no pip. The first fix was wrong too:
plain `uv venv` fails on an existing venv, so the target was not idempotent
where the version it replaced had been. Caught by running it twice instead of
dry-running it — which is how the original rotted unnoticed.

make install-reach self-checks that reach is actually on PATH afterwards
rather than assuming it. make reach-repoint gives a name to the situation
where uv keeps resolving a deleted worktree: reach still runs, edits in the
main checkout do nothing, and there is no error message.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 14:36:31 +02:00
jpmschweitzerandClaude Opus 5 b9d81ac694 feat(config): T-1260 — reach lists its domains without importing them
`reach --help` renders from a declaration table and imports nothing. The cost
of help is now flat as the registry grows, which is the property that has to
hold going from one domain to a dozen.

The trap is real and was confirmed in typer's vendored source rather than
assumed from upstream Click: TyperGroup.format_commands loops over
list_commands calling get_command on each, purely to read a short help string
off the loaded command. With lazy loading underneath, that imports every
domain in the registry to render --help — while the output looks entirely
correct. Nothing observable changes; only the import graph does.

So the test asserts on sys.modules, and it was proven to fail before being
trusted. Disabling the format_commands override made it fail and name the
cause, listing all five leaked check modules. It also carries a positive
control — invoking a domain must import its service — because without one,
"nothing was imported" would pass equally for a loader that is simply broken,
and it fails on an empty registry, which would otherwise satisfy everything
vacuously.

The check domain is created here because the test needs a subject: a stub
raising NotImplementedError would have been committed dead code. That takes
the port out of T-1262, which is rescoped to what it still owns — pydantic
schemas, byte-for-byte output parity on the drift path, and the failure
tests. The old tooling/check-client-version script stays in place and stays
wired to the pre-push hook; the deprecation window is deliberate.

One Typer behaviour worth knowing before every future domain: a single-command
app collapses into a bare command, so `reach check client-version` failed with
"unexpected extra argument" until the router got a callback. Same mechanism as
the root callback, different symptom.

Help now works at every level, closing item 5 of T-1248.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-25 07:31:15 +02:00
jpmschweitzerandClaude Opus 5 8d64800fe9 feat(config): T-1259 — reach is a real command, and Typer vendors Click
`reach --help` runs from the console entrypoint in 80 ms. typer 0.27.1 and
pydantic 2.13.4 join the dependencies, both CVE-checked against NVD, OSV and
the GitHub Advisory Database.

The design in the ticket did not survive contact. It specified a click.Group
root, on the reasoning that it would keep typer off the --help path — but
typer vendors Click as of 0.26.0, so there is no top-level click package to
import and no supported way to extract typer's internal one. A click.Group
root hosting Typer sub-apps would put two Click implementations in one
process. The root is therefore a typer.Typer, and lazy registration will go
through the supported typer.Typer(cls=...) surface with a TyperGroup
subclass. T-1260 is corrected to match.

The callback is not decoration: a Typer root with no commands AND no callback
raises at build time, and lazy registration means no command is ever eager.
The ticket claimed a zero-command root always raises — half right, and the
half that matters is that a callback makes it legal.

rich_markup_mode=None is load-bearing rather than cosmetic. It takes an empty
--help from 168 ms to 74 ms, and keeps rich and pygments off the import path
entirely rather than merely skipping the render. It also stops typer drawing
box-art help, which it does even when stdout is a pipe — that would have put
box-drawing characters into every hook log and agent capture. typer-slim was
considered and rejected: deprecated since 0.22.0, now a shallow wrapper that
installs all of typer.

D-263 amended: the feels-instant ceiling goes from 250 ms to 500 ms. A ceiling
is not a typical and most invocations sit far below it; the tighter number was
buying discipline that the import-graph assertion enforces better. Stay smart
about what loads, stop worrying about tightness.

Security, checked 2026-08-23. typer has no advisories on record. pydantic
2.13.4 clears PYSEC-2026-1812 (email-regex ReDoS, fixed in 2.4.0) — and the
2026 SSRF advisories CVE-2026-25580 and CVE-2026-54249 are against
pydantic-ai, a different package that is not a dependency here, recorded in
pyproject so the next sweep does not re-panic. Transitively, pygments 2.21.0
clears CVE-2026-4539.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-23 14:01:32 +02:00
jpmschweitzerandClaude Opus 5 559f3d82dc chore(config): T-1258 — tooling/ becomes an importable package
The skeleton the reach CLI hangs off. Nothing moves yet: this adds the
package, the bounded core/, and explicit setuptools discovery.

core/console.py is the single output path, and the split it enforces is the
whole design — stdout carries the command's actual output so `reach ... | jq`
keeps working, stderr carries the event stream as JSONL. Rendering happens at
the sink: a terminal gets human text, anything else gets raw JSONL, so a live
view and a job log are one artefact in two presentations. Emitting is
optional — the gates emit nothing — and verdict() prints once, last, carrying
its remedy as a structured field.

core/config.py resolves the repo root from __file__ against a project.yaml
sentinel, with an SR_REPO_ROOT override. No subprocess and no git call: this
is on the gate path, and cwd is not a reliable signal anyway since a hook runs
from the root and an agent call may not. Both paths are validated, because a
silent fallback is how you end up editing one checkout and checking another.

Discovery is configured explicitly rather than left to flat-layout
auto-discovery, which would have had to choose between erroring on the
ambiguity and quietly shipping client/ or docs/. Verified: top_level.txt
contains exactly "tooling".

Verified beyond the happy path — the sentinel rejects SR_REPO_ROOT=/tmp and
names both remedies; debug events are suppressed at the default threshold
while the verdict is not; stdout stays clean with stderr redirected away; and
the three unconditional push-gate checks still pass now that tooling/ is a
package, which was the real regression risk.

Two findings recorded on the tickets. make setup-venv is stale — it calls
.venv/bin/pip, but the venv was created by uv and has no pip, so the recorded
procedure and the actual state have already diverged (T-1261 owns the fix).
And settled-reach-tooling had never actually been installed: site-packages
held the dependencies but no dist-info, which follows from there being no
__init__.py to expose. This is the first commit where `import tooling` means
anything.

.venv/ was only ignored via .git/info/exclude, which is machine-local, so a
fresh clone or a new worktree did not ignore it at all.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-20 18:53:48 +02:00
jpmschweitzerandClaude Opus 5 5df8afedb9 docs(governance): D-263 — output parity over timing, and commands that stream
Three amendments, all from pressure-testing the record against how the CLI
will actually be used.

The ~104 ms push-gate ceiling is withdrawn. It was the summed cost of three
single-sample timings, imported as a requirement without asking who pays —
and who pays is the pre-push hook, which already runs cargo test or the
gdUnit4 suite on any code push. A few hundred milliseconds is invisible
there, and on a governance-only push the whole hook is about a second. The
criterion is OUTPUT parity: a ported check must produce the same output and
the same exit code as the script it replaces, and is not required to be as
fast. What replaces the ratchet is a ceiling with headroom — under ~250 ms to
feel instant. Lazy registration stays mandatory, justified by the real
threat rather than by parity: scipy.ndimage alone is 275 ms, and an eager
entrypoint would pay ~460 ms before executing a line of its own.

That budget change removed the only argument for keeping pydantic out of the
gate domain, so the carve-out goes with it. One fewer exception, and the
reference implementation is now the normal pattern rather than a footnote.

Commands also stream. The gates are milliseconds but the generators are
minutes, and an agent Bash call gives up at two and sends nothing. Detaching
alone would fix the timeout and keep the silence; streaming fixes the part
that costs real time — you learn a generator is wedged at minute one instead
of minute nine. JSONL events on stderr, stdout reserved for actual output,
rendering at the sink so a job log and a live terminal are one artefact in
two presentations. Reattach is a byte offset into an append-only file, which
is why there is deliberately no daemon.

The trap, recorded because it would quietly undo the thing this record cares
most about: streaming is ADDITIVE to the failure contract. A remedy emitted
at line 400 of 900 is printed and invisible, so the verdict still prints
once, last.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-20 18:53:29 +02:00
jpmschweitzerandClaude Opus 5 bbd64307ab docs(governance): D-263 — one CLI named reach, and Q-124 answered
Q-124 asked whether the 123-file Python tooling should be retooled into a
Rust CLI. The answer is no, and it is a costing rather than a preference.
All three frictions it names — per-script permission prompts, the venv/PATH
split between interactive and non-interactive shells, and interpreter
startup paid four times per push — are packaging problems, and one bare
command on PATH with lazy subcommand loading fixes all three. Rust would
additionally owe a numerical-equivalence proof on the planet-gen path,
whose heightmaps are committed build artefacts with goldens standing on
them: a large one-time cost to avoid a small recurring one, paid in the
currency the project can least afford to spend.

D-263 fixes the shape. tooling/ becomes an installable package behind the
`reach` command: a routing-only main.py, every domain under domains/<name>/
split router/service/schemas/helpers, a core/ bounded on day one to what
has no domain, logging and error handling attached as decorators rather
than call-site discipline, and pydantic confined to domain schemas —
measured at 87 ms against a whole gate check of 20-46 ms, which is why it
must never reach the push path. Failures carry the command that fixes them
and keep their exit code; a tool that explains itself and exits 0 silently
disables its own gate.

R-014 records the Rust option as costed down, not argued down, with the
condition under which it is worth reopening. T-1247 files the work as
eight dependency-ordered epics; only the skeleton is unblocked.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-20 02:32:21 +02:00
jpmschweitzerandClaude Opus 5 284ce847c3 docs(governance): Q-124 — one door, domain split, and failures that teach
Jeroen's shape for the tooling CLI: move the Python into a package with a
proper domain split, one door that answers everything with help, and errors
that hand back instructions rather than a status.

The domain split turns out to be discoverable rather than invented. tooling/ is
85 top-level entries — 37 loose .py, ~36 extensionless executables, 11 dirs of
which only 6 hold anything — across four coexisting naming conventions. But the
domains are already encoded as filename prefixes: blender x14, atlas x8,
generate x7, check x7, then visual/validate/test x3 and
godot/garment/pql/install x2. Those prefixes are the subcommand groups, which
is what makes the consolidation mechanical enough to be safe.

Two constraints recorded against "a new prompt not an error code", because
taken literally each would break something:

- Exit codes stay. Four of these run in the pre-push hook, which fails a push
  ONLY by non-zero exit; a tool that explains itself and exits 0 silently
  disables its own gate. That exact failure was observed in clide today, where
  unsupported-format, no-such-file and unknown-subsystem all returned 0.
  So: code AND message, never either/or.
- It must not become literally interactive. Agents and git hooks have no TTY,
  and the tea scar is already written down — its prompts "crash in Claude Code
  (no TTY)", which is why every tea call passes all flags explicitly. Any
  prompt must be TTY-gated and suppressible.

pql was cited as the precedent and measured rather than assumed. The principle
holds there for unknown subcommands (full usage dump) and not for invalid
values: `ticket status <id> nonsense` says invalid without naming the six legal
values it knows, `ticket new` says "accepts 2 arg(s)" without naming which two.
The gap is the closed sets, and it is the more common failure. Logged upstream
as pql T-112 rather than worked around here — the bar for our CLI is the
stronger one: whenever the accepted set is known, print it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 02:06:25 +02:00
jpmschweitzerandClaude Opus 5 3a640f91f7 docs(governance): Q-124 — Typer costs the cheap option down, and moves the target
Jeroen raised Typer as the Python-CLI option. Costing it changed what the
question is actually about.

The repo is already most of the way there: pyproject.toml exists, `make
setup-venv` already does `pip install -e ".[dev]"`, and 22 tooling files
already use argparse. What is missing is a single line — there is no
[project.scripts] entry at all, so no console entrypoint exists. This is
consolidation, not authorship, and it resolves the largest friction (per-script
permission prompts) for one allowlist entry.

But the framework is the second decision, not the first. A [project.scripts]
entrypoint lands in .venv/bin/, which is on PATH only when the venv is
activated — and agents and git hooks never activate it. That is the same split
VENV_PY already papers over in the Makefile, and precisely the failure recorded
for tea: an absolute path breaks the Bash(tea *) rule and prompts every time,
fixed only by a bare name on PATH. So the deliverable is "one bare command
reliably on PATH" (uv tool / pipx into ~/.local/bin, or a symlink), and a Typer
app behind an absolute venv path would solve nothing.

Two honest costs recorded against it: Typer and Click are further venv
dependencies, so it does not help the venv friction at all; and a single
entrypoint importing every subcommand eagerly would pay all 123 modules'
import cost on every invocation, four times per push. Lazy subcommand
registration is therefore mandatory rather than an optimisation, and must be
measured before and after.

Net: this looks like the answer for the check/gate family and the day-to-day
scripts, and it leaves the numpy/scipy/PIL planet-gen path alone — the part a
Rust port would have had to prove numerical equivalence for. T-1246 updated to
start here.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 02:00:39 +02:00
jpmschweitzerandClaude Opus 5 6949f800dc docs(governance): D-262 — the wiki generator flow has one canonical map
The relationship between wiki/, the generators, systems.db and the runtime is
a directed graph with two edges running opposite to the obvious direction and
one running backwards into its own producer. Prose renders that badly: every
document that has described it states a single ownership direction and is
therefore wrong about part of the tree. D-262 makes the diagram the source of
truth and points CLAUDE.md, Skill(wiki), project-structure.md and
wiki/GOVERNANCE.md at it.

The correction that matters most: body pages were described everywhere as
machine-owned and reverted on sync. They are not. scaffold_bodies.py writes
one once and never overwrites it, and import_economics then reads that
frontmatter directly as input — so a hand-edit is not reverted, it is obeyed,
and silently changes world generation. Worse than being overwritten, and the
actual reason GOVERNANCE.md forbids the edit.

New: tooling/check-dataflow-graph.py, wired into the Makefile and the pre-push
hook. It asserts every repo path named in a hand-authored diagram still
resolves — and its docstring states plainly what it cannot do: verify that an
edge still MEANS what it says. If wiki_sync.py stopped writing body pages
tomorrow, every path would still exist and the check would still pass. Edge
semantics stay a human check against the tool's source, so nobody reads a green
gate as a verified map.

Verified by breaking it: pointing one label at a moved path fails with exit 1
naming that path; restoring it passes. Building the checker also caught two
real vaguenesses in the diagram — "GJ-*/index.md" and "bodies/{id}/index.md"
were written without their wiki/star-systems/ prefix, which is precisely the
ambiguity this map exists to remove. Generated star-map .d2 files are excluded
by name; their correctness belongs to their generator under D-223.

Also files Q-124 + T-1246 (tooling): whether the 123 Python files under
tooling/ should become one Rust CLI of pql's calibre. The friction is real and
mostly not about the language — the permission gate prefix-matches whole
command strings and a blanket Bash(python3 *) grant is forbidden, so each tool
prompts near-individually, while a single binary is one allowlist entry. The
record requires pricing the cheap alternative (a Python dispatcher entrypoint)
before recommending Rust, and flags the hard constraint: import_economics is
stamped by source SHA, so any port must keep that contract intact through the
transition rather than disabled during it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 01:56:12 +02:00
jpmschweitzerandClaude Opus 5 0817befcba docs(diagrams): SVG replaces PNG, and a map of the wiki generator flow
d2 emits SVG natively; its PNG path wants a ~150 MB headless-Chromium
download and prompts interactively, so every PNG here was produced by an
out-of-band magick step. There is no Chromium on this system. Dropping PNG
removes the dependency rather than trading one format for another, and cuts
docs/diagrams/ from 17 MB to 3.3 MB. SVG renders in Gitea and in clide
(`clide draw --file <path>`, which takes .d2 source directly), and diffs as
text.

One PNG is kept on purpose: design/star-map-concentric.png has no .d2 source.

Also renders the 7 star-map .d2 files for the first time. star-map-plan.md
listed their renders as a deliverable in March and the step never ran; the
new `make check-diagrams` is what surfaced it.

New: docs/diagrams/data-flow/wiki-generator-flow.d2 — which way the arrows
point for any file under wiki/. Every edge was read in the tool's own source
rather than inferred. It records the trap that keeps costing us: scaffold_bodies.py
writes a body page once and never overwrites it, and the generator then reads
that frontmatter directly — so a hand-edit there is not reverted, it is obeyed,
and silently changes world generation.

Two rendering traps found the expensive way and now written down:

- A d2 `|md` block becomes an SVG <foreignObject>. ImageMagick and flutter_svg
  both silently drop it, so the legend was in the file and invisible in every
  viewer except a browser. Plain labels render as real <text> everywhere.
- Container boxes fight the layout engine. Grouping nodes whose flow-depths
  differ forces long edge routes; this diagram went from an unreadable 2.4:1
  sprawl to a legible 0.75:1 by deleting five containers and changing nothing
  else. Colour classes carry the grouping instead.

make diagrams / make check-diagrams render and gate. Repo-specific rules in
.claude/rules/diagrams.md; d2 syntax and the traps live in the user-scope
d2-diagram skill, whose PNG default was flipped to SVG to match.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 01:49:19 +02:00
jpmschweitzerandClaude Opus 5 34e5d7b636 docs(meta): body pages are obeyed, not reverted — correcting the wiki skill again
Jeroen asked whether I had read the python that writes the frontmatter. I had not — only grepped it. Reading body_definition_parser.py and scaffold_bodies.py properly overturned what I had written twice today.

scaffold_bodies.py NEVER OVERWRITES ('Only creates files that don't exist yet... existing body index.md files are skipped'), and the generator reads that frontmatter directly. So a hand-edited body page is not reverted, it is OBEYED, and it silently changes world generation — worse than being overwritten, and the actual reason GOVERNANCE.md forbids it. System pages behave the opposite way: wiki_sync.py re-renders their READ-ONLY blocks, so edits there ARE reverted. Three cases, not two.

It also explains T-1244's whole measurement: body_definition_parser resolves each field override > direct read > derived > inferred > SEEDED RANDOM. Continuous axes vary because they fall to the random tier; categorical axes are concentrated because they are read from the bodies table. The variance question belongs to the atlas CLI catalog, not the wiki.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-20 00:53:00 +02:00
jpmschweitzerandClaude Opus 5 0dc68dc1b8 docs(meta): fork-for-sidequests rule + wiki skill corrections from the cold test
The cold test worked as an experiment: a fresh agent with Skill(wiki) cited it first, refused to hand-edit body frontmatter, knew the corp regen-db stamp trap, and knew corp_specialization is missing from its own template. It also found four things the skill had wrong or missing, all verified before folding in: body frontmatter is a MIDDLE layer (atlas CLI -> systems.db -> scaffold writes the page -> import_economics reads it back), not the origin GOVERNANCE.md implies; the four empty categories are Q-118, an open scope question rather than an invitation; some bodies are visual-regression goldens and nothing in wiki/ says so; and status is editorial, not an import gate. Also: check current state before editing, since the test's own task described a change that was already true.

T-1244 corrected in the same pass — tectonics is derived from planet_class via a lookup (body_definition_parser.py:563), so the measured 68% 'low' is a projection of the class distribution, not an authoring choice. The ticket's question changed accordingly.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-20 00:46:04 +02:00
jpmschweitzerandClaude Opus 5 9fb74bdf22 chore(meta): record the mechanism behind the roadmap gap on T-1245
Jeroen: 'I tend to restrict future side quests to not confuse your context.' A reasonable practice with a bad side effect — intent stays conversational and reaches the repo by accident. Resolution recorded as two channels rather than more sharing: working context stays narrow, forward intent gets FILED. Carries a design constraint into the initiative's form investigation — weight options by cost-to-APPEND, since these facts surface mid-bug and a ritual will not get used.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-20 00:40:21 +02:00
jpmschweitzerandClaude Opus 5 4de90529ae chore(meta): file T-1245 — a roadmap that carries intent, not just work order
The cascade owns order, the ticket tree owns decomposition, the DQR tree owns individual rulings; none answer what the game is going to be. Evidence it is a real gap: three roadmap-level facts surfaced in one conversation on 2026-08-20 that exist in no artefact, and two of them were written up as suspected defects by an agent reading carefully, because nothing recorded them as intent. Pickup instructions make epics an OUTPUT of a harvest/interview/investigate-form/propose pass, explicitly not an input.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-20 00:38:40 +02:00
jpmschweitzerandClaude Opus 5 c25af8d753 docs(meta): make the wiki seed reachable — blind-prediction experiment and its fix
An experiment, at Jeroen's request: predict how the wiki seed data is structured
WITHOUT reading it, seal the prediction, then score it. The prediction is
a1addf7e2, committed before wiki/ was opened so it could not be retrofitted.

The score, against a rule fixed in advance:

RIGHT — markdown + YAML frontmatter, TOML for economics tables, the body path
shape, more trees than the three I had seen.

WRONG — "a source, never an output". That holds for 253 pages and is backwards
for 3,262: star-systems/ is GENERATED from systems.db by tooling/db/wiki_sync.py,
its <!-- READ-ONLY --> blocks are renders, and body frontmatter IS the body
definition rather than a description of one. Also wrong: "probably no schema
docs" — there are 17 templates, ten authoring guides, economics/schema.md and a
GOVERNANCE.md that states the ownership models plainly.

ABSENT (the expensive bucket) — the two ownership models running in OPPOSITE
directions; the GTTR prose channel; terrain.npz/globe.png; that stations and
districts have NO wiki directories; that `description` frontmatter exists so
agents can filter before loading; and the scale, 11,864 files.

ROOT CAUSE, and it is not missing documentation. The wiki documents itself well.
It was unreachable: wiki/ appears in NEITHER CLAUDE.md's Project Structure block
NOR .claude/rules/project-structure.md, the annotated tree whose entire job is
orienting an agent. The largest tree in the repo — the seed for the whole Reach —
was invisible from both files a session reads first. Every item in the absent
bucket follows from that one omission. The proof is this session: it spent three
days fixing Ferrath's terrain rendering and never once saw
wiki/star-systems/GJ-820B/bodies/GJ820Bc/index.md, the file that defines Ferrath.

Fixed here: wiki/ enters both structure documents with the ownership split stated
where it will be read, and Skill(wiki) carries the traps — never hand-edit a
READ-ONLY block or body frontmatter, stations have no directories, the id is
spelled two ways, editing corp PROSE stales systems.db, and absent variance is
often deliberate rather than a gap.

That last point cost two false findings in one measurement and is worth the
warning: chemosynthetic:false on every body is a namespace reservation for
dextro-DNA-style biochemistry once geology and nature spawn to the 1x1m pixel,
and enabled:false on ~65% is staged rollout — clean planet types first, generator
scripts for the rest after. Both read as defects without the roadmap.

Also measured, since the seed's job is to supply variance: continuous axes are
rich (unique seed per body, 460-716 distinct values across orbit/tilt/ice/land)
while the categoricals that gate morphology are concentrated (68% tectonics low,
51% planet_class frozen). Filed as T-1244 with the design question stated first —
whether the distribution is intended — rather than as a defect.

Method caveat recorded in the findings: the aggregator reads scalar frontmatter
only, and atmosphere_color's "100% null" was a parser artefact, not a finding.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-20 00:31:06 +02:00
jpmschweitzerandClaude Opus 5 a1addf7e21 docs(meta): seal a blind prediction of the wiki seed structure before reading it
Written before opening wiki/, committed first so the prediction is timestamped and cannot be retrofitted once the answer is known. Contamination declared inline: paths already seen this session via generator_sources.py and visual_scenarios.gd are marked [SEEN] and discounted. Scoring rule fixed in advance, with 'load-bearing thing I did not know existed' as the bucket that decides what gets documented.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-20 00:18:07 +02:00
jpmschweitzerandClaude Opus 5 fab70edd5a feat(ui): the stipple says WHERE the ground is broken, not just that it is (T-1194)
RimWorld's technique 4, the reference this ticket names, works by CONTRAST: the
Rockies and Appalachians carry dense hatching and the Great Plains carry none.
Ferrath's Global carried an even wash of dots over every landmass instead —
texture present, information absent.

The cause was a constant measured on the wrong rung. RUGGEDNESS_FULL_SCALE_Q = 4
comes from Region ("d4 2.38"), but the elev_q path it gates only ever RUNS at the
orbital rungs, where relief_q is flat and the stipple falls back to it. Ferrath
Global measures a mean 1-cell elev_q gradient of 0.99, so a coherent 4-cell
baseline reaches ~4 and saturates the constant exactly: every land cell read as
fully rugged.

It now normalizes against the canvas's own measured gradient — the same
self-calibrating shape T-1240 gave the hillshade, and for the same reason: one
constant cannot serve rungs whose sample spacing differs by four orders of
magnitude. A contrast curve rides on top, because even unsaturated the linear
reading puts ordinary ground mid-range and paints grain everywhere.

Measured on Global, before -> after: 1,679 -> 1,900 distinct colours, 145.90 ->
147.39 lum spread, and the pale uplands now stipple visibly denser than the
lowlands beside them.

Recorded because it cost two wrong turns: I first guessed saturation, then
talked myself out of it after measuring a 1-CELL gradient (0.99) against a full
scale meant for the 4-CELL baseline, and shipped a contrast curve alone — which
measurably did nothing (1,679 -> 1,698), because a curve cannot separate values
already clamped to 1.0. The gamma is kept; it does its job now that there is a
range to curve. The elev_q gradient joins relief_q's in the capture readout, so
the next person tuning this can read the number instead of guessing at it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-20 00:13:51 +02:00
jpmschweitzerandClaude Opus 5 646db131a9 chore(meta): close T-1240 — Region reads as terrain, acceptance ladder shot
relief_grad across the cold ladder: Global 0.00, Region 1.08, District 0.30, Quarter 0.07. Closure note records the two premises the ticket got wrong (Nyquist is the aliasing limit, not a legibility one) and the coast-warp trade taken at Region, with its reversal path.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-18 20:49:24 +02:00
jpmschweitzerandClaude Opus 5 9b146f9e1f fix(simulation): derive at the octaves a rung can actually reconstruct (T-1240)
Region rendered as fine uniform stucco while District and Quarter, on identical
code, read as terrain. The cause was sampling: `min_wl_m` arrives as an LOD
request and defaults to 0, so every invented octave contributed at every rung.
MIN_WL_BANDS_M was meant to be the floor but is built from the rung's CELL SIZE
(2 x DISTRICT_M), which stopped being the sample spacing at the D-255 extent
inversion — a rung fixes EXTENT now and spacing falls out of the canvas size.
The bands were off by roughly the cell count, and the served path never consulted
them anyway.

The cutoff is now derived from the resolved spacing, which is what this ticket
asked for. Two things had to be measured rather than reasoned to get it right,
and both corrected me.

FIRST: the field was the culprit, not the renderer. I attributed the stucco to
the client stipple painting noise onto a smooth field. Surfacing the terrain
layer's own mean |relief_q gradient| in the capture readout settled it in one
shot: Region 18.24 steps per cell — 144 m of relief between NEIGHBOURING cells —
against District's 0.30 and Quarter's 0.07. The server was sending noise. That
diagnostic ships here for the same reason `plane_variety` did in T-1213: a noisy
field and a renderer inventing noise look identical, and one number separates
them.

SECOND: Nyquist is the wrong threshold. The first version floored at 2 x spacing,
the aliasing limit, and Region barely moved (56.16 -> 59.73 lum spread, gradient
still 18.24) because 2 samples per cycle is unaliased but renders jagged. The
rungs that already worked say what the real bar is: District reconstructs its
finest surviving octave at 34 samples per cycle, Quarter at 135. At 8x, Region
goes to 1.08 gradient and 70.01 spread, and shows ridges and valleys.

THE TRADE, taken deliberately and recorded in the tests: an 8x floor also
truncates the coast warp's 2,048 and 1,024 m octaves at Region, the band T-1160
added for "one coastline at every rung". An earlier test here asserted that band
must survive; it now asserts the opposite. Same reasoning as the relief: a
1,024 m coastline wiggle at 379.3 m per cell is 2.7 samples per cycle, so drawing
it draws noise rather than coastline character — a rung cannot show shape finer
than its own cell. The warp is amplitude-capped sub-pixel on the working grid, so
what is lost is small. If a future pass wants the warp exempt, the fix is a
relief-only floor threaded through derive_at_metres, NOT a lower multiple, which
takes the stucco back.

Global is exempt: its floor would be ~70 km and would truncate the whole warp
band, and it needs none — the orbital derive leaves relief_q flat at 50. District
(3.79 m spacing) and Quarter (0.948 m) floor below every octave in play and
derive byte-identically, which their own test pins.

Cache-safe by construction: the floor is a pure function of (rung, extent,
body_radius), all three already in the step-canvas cache key. 0.4.12 is required
anyway — this changes derived BYTES at Region, so a 0.4.11 entry holds a field
this build would never produce.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 12:43:46 +02:00
jpmschweitzerandClaude Opus 5 6d4c9e92c2 feat(ui): the copse and the rocky outcrop, drawn per cover class (T-1213)
D-258 invariant 2 asks the map to show the minority the orbital summary
suppressed — clearings, marsh, rock, scrub inside a cell that reads "forest" from
space. composition.rs goes to real trouble to invent it, and the conservation
harness measures it: over a District patch on Ferrath the tally is {Barren: 175,
Forest: 16209}, so 1.07% of that ground is exposed rock.

The renderer was averaging it back out. One test, `veg >= Scrub`, at one density
and one strength: Forest, Scrub and both Riparian classes drew the IDENTICAL
mark, and Barren drew none at all. A wood looked like scrub, and bare rock was
invisible by construction — on the rungs whose whole purpose is to show what the
summary hid.

Each class now has its own grammar, differing on the three axes a mark has:
density (how much of the class's ground carries it), strength (how far it moves
the base colour), and lattice (the block size marks are decided on — bigger reads
as a clump, smaller as grain). Rock LIGHTENS where everything else darkens, which
is the point rather than a flourish: bare stone catching the light is the one
cover type brighter than the ground around it, so it separates from vegetation by
sign alone and can never read as "denser plants". It gets its own ScatterField
salt so an outcrop does not preferentially land where a copse already did.

Tuned against captures, not guessed. A first pass gave Forest a 4x4 lattice at
52% density, which produced visibly axis-aligned dark SQUARES — a 4-cell block is
8 screen px at District — and read as an artefact laid over the hillshade. It
also mistook the background for a feature: Forest is 98.9% of this frame, so
marking half of it dark is not "the occasional copse", it is a second colour
layer. Pulled back to grain at a 2x2 lattice, and the occasional thing is now the
thing that catches the eye: the outcrops read as scattered pale clusters of
exposed ground, exactly the "occasional copse/tree/rocky outcropping" that was
missing.

District holds its form through the change (lum p1-p99 74.15, against 74.43
before the cover marks and 13.72 when the rung was flat).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 09:04:32 +02:00
jpmschweitzerandClaude Opus 5 8787ee1844 feat(ui): hillshade the deep rungs — form from light across slope (T-1213)
relief_q now reaches the renderer, and the first pass spent it on brightness:
lighten where the ground is high, darken where it is low. That moved the numbers
(District 13.72 -> 77.01 lum spread) and still looked like moss, because the eye
does not read landform from absolute brightness. It reads it from light falling
ACROSS a gradient — height-shading gives a rise and a fall the same tone, so no
ridge ever reads as a ridge.

So relief drives a proper hillshade: the local gradient of the field dotted with
a light from the upper-left. The light direction is not a free choice; lit from
the lower-right the brain inverts the read and valleys pop out as ridges.

THE SCALE IS MEASURED PER CANVAS, not fixed, and the first attempt at this failed
exactly the way this file already warned a fixed gradient constant would (see
RUGGEDNESS_BASELINE_CELLS: "the same 4-cell delta reads 21.86 at Region and 0.08
at District"). With a constant full-scale of 8:

    Region 81.72 but District 77.01 -> 20.01, Quarter 42.56 -> 16.44

because at District's 3.8 m per cell neighbouring cells barely differ. The
terrain layer now measures each canvas's own mean |gradient| once per rebuild and
the hillshade normalizes against it, so one constant works at every rung.

Ladder (tooling/atlas-flatness, lum p1-p99), flat -> shipped:

    Global    145.69 -> 145.69   unchanged; relief_q is flat 50 at orbital
    Region     33.59 ->  54.30
    District   13.72 ->  74.43
    Quarter    11.01 ->  73.72

District and Quarter now read as terrain — ridgelines, valleys, and the stipple
organised into contour-like bands. Judged by eye on the captures, not by the
metric alone.

Stipple full-scale 25 -> 60. The old value was calibrated against a relief_q that
never arrived, so it was tuned to the elev_q fallback; with the real plane nearly
every land cell earned a mark and Region read as static (17,599 distinct colours,
more than twice Global's, for a quarter of the legibility). Form comes from the
hillshade now; the stipple is grain on top of it.

REGION IS NOT FIXED, and the cause is T-1240 rather than this change. It renders
as fine uniform stucco: a Region cell is 379 m of ground while the relief field's
content sits in the 128-1024 m band, so the field is at or below Nyquist and the
gradient the hillshade reads is aliasing, not slope. min_wl_m defaults to 0 on
the served path, so nothing truncates the octaves Region cannot resolve — which
is precisely what T-1240 proposes to fix. That ticket said the stale cutoff was
"currently inert"; it is now the thing capping Region, and T-1240 is updated
with the measurement.

Three tests, on direction rather than magnitude so tuning does not rewrite them:
a hill's west flank lit and east flank shadowed, a uniform field shading nothing,
and the canvas edge not drawing a rim. That last one is a bug this nearly
shipped: `_l8_value` returns 0 out of bounds and 0 on relief_q means MAXIMUM
HOLLOW, so sampling off-canvas posts a full-scale false gradient all the way
round the frame. The sample position is clamped instead.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-16 18:30:21 +02:00
jpmschweitzerandClaude Opus 5 b9cd26429e chore(meta): file T-1243 — fog perf test flakes under gate load
The pre-push gate rejected the T-1213 push on a wall-clock fog budget (0.606 vs 0.5 ms) that passes 23/23 in isolation on the same build. Second hardening cycle for the same failure mode: min-of-7 defends against one slow sample, not the sustained core saturation the gate itself creates by running cargo and tooling suites immediately before it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-16 16:37:58 +02:00
jpmschweitzerandClaude Opus 5 3ec35b87c8 fix(client): the deep rungs were flat because relief_q fell off the wire (T-1213)
`relief_q` is the one field with signal below District — elev_q's 80 m steps
quantise sub-district detail away, which is precisely why relief_q was invented.
The server has encoded it since 5eb394b36 and the terrain layer has asked for it
by name ever since. step_canvas_protocol.gd's decode dictionary never listed the
key, so `canvas.get("relief_q")` was always null and the plane arrived nowhere.
The server half of that change landed; the protocol half did not.

That is the whole reason Region and below rendered as a flat wash. Measured plane
variety at District before the fix:

    {morphology: 1, elev_q: 11, relief_q: 0, moisture_q: 25, vegetation: 3}

A 0 there means ABSENT, not constant — a distinction the capture could not make
until this commit adds it, and the reason two earlier sessions read the flatness
as a missing generator rather than a missing key.

Also spends the field properly. It drove a stipple PROBABILITY only, so a ridge
and a plain differed in dot density, which at one pixel per cell reads as noise;
and `_ruggedness()` took absf(relief_q - 50), discarding the sign the server
deliberately preserved ("a hollow and a rise are different ground... the reverse
is not recoverable"). Relief now shades continuously and signed — rises lighten,
hollows darken — UNDER the stipple rather than instead of it. Ruggedness
(unsigned) and elevation (signed) are different questions and both are worth
asking.

Ladder, before -> after (tooling/atlas-flatness, lum p1-p99):

    Global    145.69 -> 145.69   unchanged, correct: relief_q is flat 50 at
                                 orbital rungs by construction
    Region     33.59 ->  71.01   2.1x
    District   13.72 ->  77.01   5.6x
    Quarter    11.01 ->  42.56   3.9x

Structure retention Global->Quarter: 7.6% -> 29%.

NOT finished, and the ticket says so: Region now reads as heavy speckle, because
ruggedness is real data instead of an elev_q-gradient fallback and far more cells
earn a mark than the T-1194 tuning assumed; District reads as soft blobby relief,
form without directionality. Both are grammar/tuning follow-ups on a channel that
finally carries signal.

0.4.9 is a REQUIRED bump. The disk cache stores the DECODED canvas, so every
earlier entry physically lacks the field and would keep rendering flat against a
build that reads it — the first bump in this series where a warm cache is wrong
about CONTENT, not merely stale. tooling/canvas_sources.py gains
step_canvas_protocol.gd for the same reason: it decides which planes exist, the
cache stores its output, and the T-1242 gate would not have flagged this fix
while the registry stopped at ui/.../step_canvas/.

Regression cover: every protocol test passed throughout the weeks the plane was
missing, because each asserted a field it already knew about and none asserted
the SET. There is now a test walking all eight dense planes of EncodedStepCanvas,
verified by disabling the fix and watching it fail by name.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-16 15:22:50 +02:00
jpmschweitzerandClaude Opus 5 e5224b1a44 chore(meta): 0.4.8 — the T-1242 gate's first false positive, paid not dodged
A test-only edit to composition.rs tripped the canvas-generation gate, which is path-based and cannot tell an assertion fix from a generator change. Bumped rather than excepted: the ruling is that a false positive costs one round of cache misses and a false negative costs a week. Second no-op bump in two days, noted in project.yaml so the rate is visible if it becomes noise.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-16 13:03:09 +02:00
jpmschweitzerandClaude Opus 5 869837f728 test(simulation): the conservation gate's monoculture check was a tautology (T-1213)
D-258 invariant 2 says descending the ladder must reveal COMPOSITION — a cell
reading forest must be able to contain the clearings and rock the vote
suppressed. One assertion stood behind that, and it read:

    assert!(tally.len() > 1 || share == 1.0, ...)

A single-class tally has a 100% share by definition, so both branches are always
satisfiable: the check could never fail, including in the exact case its own
message names, "or nothing was composed". The invariant had a test and no gate.

Split into the two bounds the invariant actually has, because it is two-sided:
conservation caps how much may be invented (majority > 50%, already asserted) and
composition sets a floor on how little (minority >= 0.1%). Verified by raising
the floor to 2% and watching it fail on the measured 1.07%, then restoring it —
the floor is a tripwire for "did anything happen", deliberately far below the
measurement rather than tuned to it.

Measured at the descent ladder's own anchor on Ferrath:
  conservation: majority class 3 at 98.9% across 2 classes {1: 175, 3: 16209}

So composition IS working in the data and conservation holds. The map is flat
anyway, and tooling/atlas-flatness (added here) says why the eye was not enough:

    rung      distinct   lum p1-p99
    Global        1581       145.69
    Region        2923        33.59
    District        53        13.72
    Quarter         46        11.01

Region carries almost TWICE Global's distinct-colour count while holding a
quarter of its structure — the dither pass adds colour noise, not information, so
a colour-count metric would have called the flattest rung the richest. Structure
falls ~92% from Global to Quarter.

The cause is a channel mismatch rather than a missing generator: composition
perturbs moisture_q/slope_q, and the base map draws morphology hue x elev_q
lightness. The ladder scenarios pass no overlays deliberately, so the composed
fields are never rendered in the very shots that judge this work. Recorded on
T-1213 with the three ways forward; the choice touches D-258 and is Jeroen's.

The gate is still #[ignore]d — noted on the ticket as worth moving into a harness
that runs, since believability and window-derivation already load real bodies in
the normal cargo test path.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-16 12:35:19 +02:00
jpmschweitzerandClaude Opus 5 6e6218d654 chore(meta): close T-1242
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-15 00:07:01 +02:00
jpmschweitzerandClaude Opus 5 48fee8a0b6 feat(config): make the canvas-generation/version pairing a gate, not a habit (T-1242)
project.yaml's version is the Atlas disk cache's only invalidation signal, and
nothing enforced that changing canvas GENERATION also moved it. It broke five
times -- 0.4.2 lake_margin_q, 0.4.3 coast_warp_px, 0.4.4 the extent inversion,
0.4.5 the Global sentinel, 0.4.6 one-course-per-river -- each bumped only after
someone noticed a wrong map. The failure is invisible to its author: it needs a
warm cache to reproduce, so a cold checkout looks fine. T-1239 is the last one,
and it took eight days.

tooling/canvas_sources.py is the path registry; tooling/check-canvas-version
rejects a push that touches those paths without moving project.yaml's version
line. Wired into the pre-push hook, `make check-canvas-version`, and, for the
parsing units, `make test-tooling`.

Verified against real history rather than a synthetic branch: run over
4e503c356 -- the commit that actually caused T-1239 -- the gate rejects and names
the three files. Run over the commits that DID bump (bdea71953, 39f0fd8c5, and
T-1239's own fix), it passes.

The registry is globbed, not hand-listed. step_canvas.rs imports ten sibling
modules and those import more, so a traced closure would be stale within a month,
and stale here is silent. It over-includes on purpose: a false positive costs one
bump and one round of cache misses, a false negative costs another week of a
wrong map -- the ticket's own ruling.

Two deliberate calls worth naming. The registry includes ITSELF, which closes the
narrowing hole: remove a path and change that same path in one push, and the gate
still fires because the registry file is in the set. And there is no override
flag -- it would be reached for exactly when someone is certain their change is
harmless, which is the reasoning behind all five regressions.

Version bumped 0.4.6 -> 0.4.7 with NO canvas-generation change: self-inclusion
means adding the registry trips its own rule. Spent rather than special-cased,
because the first exception is how a rule like this dies.

The units cover the property no branch run can show -- that editing project.yaml's
comment block, which quotes old version NUMBERS directly above the field, is not
a bump -- plus a registry-coverage test naming the files each of the five known
regressions touched, so a future narrowing past them fails loudly.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-15 00:06:32 +02:00
jpmschweitzerandClaude Opus 5 a1568d27c1 chore(meta): close T-1241
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-14 23:35:41 +02:00