5d83e1d2eb210e0952a17f43d9db1a77c8f30619
2997
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
5d83e1d2eb |
feat(config): T-1276 — every invocation is a job, carried ambiently
Streaming as a decorator, first half. Each invocation of reach gets an id and every event it emits is tagged with it, which is what will let a detached run's log be read back and what correlates the lines of a run that streamed for nine minutes. No command signature changed and no command imports core.jobs — that is the point, per the D-263 amendment: a command must not know jobs exist, because the alternative is call-site discipline wearing a different hat. A ContextVar rather than a module global. A global is correct only until something runs two invocations in one process — which a test harness or a future batch verb does immediately, and which would then interleave two jobs' events under one id with nothing reporting an error. The job context is the OUTERMOST wrapper, and it has to be. @logged emits from its finally and @handle_errors emits its verdict while unwinding, so a context established inside either would already be reset by the time the two most important events are written — leaving them the only untagged lines in the log, and they are precisely the ones a detached run gets read back for. Fixed in passing: the job id used local time while every event's ts is UTC, so an id read 155327 beside its own first log line reading 13:53:27. Two hours apart reads as a logging bug every time someone correlates them by eye. New conformance invariant — nothing outside core/ may import core.jobs. My first version of it inspected only the module path, so it missed `from tooling.core import jobs`, where the name is in the import LIST and which is the form anyone would actually write. It passed while checking nothing. Rewritten to catch all three reachable forms and then verified by committing a real violation, which it named by file and line. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
b5beda0df7 |
feat(config): T-1275 — bare reach is discovery, so it exits 0
Bare `reach` and bare `reach <domain>` printed help and exited 2, Click's usage-error convention. Running reach with no arguments is the DISCOVERY action — it is how the tool gets learned from nothing — and a caller that branches on exit status would read its own onboarding as a failure. Now they exit 0. D-263's exit-code contract is untouched: it governs failures, and printing a command list is not one. Verified across the whole matrix, because this change flirts with the exit-0 trap that record opens with — bare 0, bare domain 0, --help 0, unknown domain 2, unknown verb 2, real failure 1. All five are now pinned as a sixth conformance invariant, since an exit code regresses silently and nothing else would notice. Proven to fail by putting the 2 back. The implementation also collapses a duplicated class. core/cli.py holds ReachGroup with both shared behaviours — no-args-prints-help-and-exits-0, and unknown-name-enumerates — and LazyDomainGroup now extends it instead of subclassing TyperGroup directly, keeping only the laziness and the domain-specific wording. The enumeration logic previously existed twice in slightly different forms, which is how the root and the domains would have drifted into disagreeing about their own conventions. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
91e25a3e7a |
docs(governance): D-263 — the primary user is an agent, and that changes things
Stated plainly because the record was quietly assuming otherwise: Jeroen runs make and plays the game; the caller typing reach all day is Claude. It resolves several arguments in the opposite direction from human-CLI instinct. --help is a discovery mechanism rather than documentation, since it is how the tool gets relearned from nothing every session — which makes the domain list and closed-set enumeration load-bearing rather than polish. Output volume is a context cost, so quiet-by-default is right for a better reason than not spamming a hook. Latency matters less than legibility: nobody drums their fingers at 300 ms, but a multi-minute silence is expensive because a wedge is indistinguishable from work. And errors that name the next command are the highest-value requirement here, because the reader is usually deciding what to run next — "no" costs a whole exploratory turn. One correction follows directly. D-263 had scoped streaming to "callers with no escape — a human terminal, a Makefile, a git hook", reasoning that Claude Code's background mode already solved the timeout for agents. That got the audience backwards. Background mode solves the timeout and nothing else: it returns when the process exits, so a nine-minute wedge still looks exactly like nine minutes of work. Streaming is what makes a long run legible while it runs, and reattach is worth most to the caller whose attention is not continuous. Both are primary-user features, not fallbacks. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
6b31111cd2 |
docs(governance): D-263 — make and reach split by kind, streaming as a decorator
Two decisions taken before the 160-file move, because both change what the move produces. The Makefile has 84 targets and is today's front door, so "one CLI for all repo tooling" was not yet true. The split is by what a target DOES: make keeps genuine build and test orchestration, and targets that are really tooling wrappers are retired in favour of reach verbs — retired, not wrapped. A wrapper leaves two ways to invoke every tool, and then reach --help stops being the answer to "what tooling exists" because the Makefile is still a competing index. Two doors is the condition this record exists to end, so keeping both would defeat it while looking like caution. Streaming becomes a decorator rather than an API commands call. @command already wraps every invocation, and that is exactly the seam where job identity, progress correlation and detach belong: the decorator assigns the job id, tags the events, and forks on --detach. A command must not know that jobs exist. The alternative — each command opening a job and remembering to close it — is call-site discipline wearing a different hat, and it fails the same way the fortieth command into a porting session, with the failure vanishing from the log and nothing to indicate anything is missing. Logging and error handling are decorators for this reason; streaming is the third cross-cutting concern, not a special case. Consequent resequencing: T-1264 lands before the T-1250 move, so every ported command arrives already streaming. Old scripts now retire per domain as each port passes its parity test, rather than in one sweep at the end — a continuous shrink, instead of months where every tool exists twice and an edit can land in the dead copy. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
91a57b8304 |
docs(config): T-1271 — the domain map, before anything moves
Every Python file and executable in tooling/ assigned to one of 15 domains, with the ambiguous cases carrying their reasoning. The per-domain port tickets are written from this rather than guessed, so their boundaries do not have to be renegotiated halfway through a 160-file move. Three things counting turned up that reading would not have. The Blender carve-out is 35 files, not the 13 visible at top level — 22 more are inside garment-fit/, which turns out to be a payload directory wearing a domain's name. The epic said 35 and an earlier survey of mine said 14; the epic was right. That is not cosmetic: `character` is a far smaller domain than directory sizes imply, and a port ticket written from the listing would have been wrong about both it and the carve-out. The "28 singleton prefixes" were an artefact of splitting filenames on the first token, which scattered coherent families — sculpt-star-map, tune-star-map-topology and generate-star-map* are one group counted as three orphans. Counting families instead, the genuinely ambiguous set is small enough to enumerate with reasons. And tooling/db/ is misnamed: it holds the audio/image/Trellis connectors and wiki_sync, while the actual database work is in economy-db/. Naming a domain after that directory would have carried the misnomer forward. Judgment calls settled with reasons, since each sets a precedent. Registries stay data rather than becoming verbs nobody would type. Gate tests do not become a `test` domain implying a runner that does not exist. pql-migrate is provenance — archived, not deleted and not importable. `pr` is a domain the epic omitted, kept out of `dev` so dev does not become the drawer everything ambiguous goes into. And `atlas` is overloaded across three unrelated places — map data, terrain quality analysis, and systems.db index tables — which stay with their owners rather than being collected into a domain whose only common thread is a noun. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
49fa6ada95 |
feat(config): T-1249 — the contract is a decorator, and now a test
Every non-zero exit names the command that would fix it, and still exits non-zero. Both halves matter; the second is the one that gets lost, because a tool that explains itself beautifully and exits 0 looks MORE correct while having silently disabled its own gate. core/errors.py holds ReachError(message, fix=) and @handle_errors. core/logging.py holds @logged, emitting through console rather than a second sink — one output path, so there is nothing to drift. core/command.py composes them, and the order is load-bearing: handle_errors wraps logged, so the logger sees the original exception. Inverted, every failure would be recorded as "SystemExit" and the log would say nothing about what went wrong while looking like it worked. core/ raises SystemExit, not typer.Exit. A service must be callable from a test, another service, or a future second front end, and an exception type that only makes sense inside a CLI leaks the transport into every layer. The check router is retrofitted off its hand-rolled verdict-and-exit pattern — exactly the boilerplate this removes — and test_check_parity.py passes unchanged across the retrofit. That test predates the decorators and pins exit codes against the old script, so it is independent evidence, not a test tuned to match new behaviour. Unknown domains and unknown verbs now enumerate what exists instead of only saying no. That needed a shared group class, which collided with "no typer outside main.py and router.py" — resolved by sharpening the invariant rather than breaking it, since its purpose is that a SERVICE never knows it was called from a CLI. Transport now lives in main.py, router.py and core/cli.py; never in service.py, schemas.py or helpers.py. The upside is that cli.domain() carries the settings that were previously per-router decisions, including the load-bearing rich_markup_mode=None that one forgetful domain could have undone. test_conformance.py makes five invariants executable, AST-based rather than grep. Scoped to the package, not the 123 legacy scripts — and deliberately so: as T-1250 moves each script into domains/, it lands inside the scope and the rules start applying automatically, so the test's reach grows with the migration. Proven to fail before being trusted: removing @command and removing a fix= each produced a failure naming the file, the line and the reason. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
1eb30a1460 |
chore(config): T-1263 — one permission rule for the whole tool surface
Bash(reach) and Bash(reach *) join .claude/settings.json beside the pql pair. Two entries, not the one the ticket asked for: a rule ending in " *" does not match the bare word, and bare `reach` is a real invocation now that it prints the domain list. pql, make, cargo test and ruff check each carry a bare-form entry alongside the wildcard for exactly this reason, and adding only the wildcard would have left `reach` prompting while `reach check ...` did not. This is the line Q-124 was actually filed about. Ten hand-written Bash(tooling/...) entries each cover a single script and every unlisted tool prompts; one command with subcommands is one rule covering everything. The ten stay for now — the old scripts are still the working tools until T-1253. On verification, since the ticket warned specifically against declaring this done on the wrong evidence: real calls run clean, but that is NOT proof the rule matched. The same calls succeeded before the rule existed — there was no Bash(reach ...) entry in either settings file and no blanket grant — so the session was already permitting them and the observation cannot distinguish "the rule matched" from "the rule was never consulted". settings.json is read at session start, so this cannot be self-verified from the session that wrote it. Proof is a later session, in a prompting mode, where reach runs without asking. One accepted limitation, documented rather than worked around: rules prefix-match the whole command string, so an env-prefixed call like SR_REPO_ROOT=... reach ... will still prompt. An environment override is a real departure from normal invocation; the ordinary form is what needs to be frictionless. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
5cdb3e9327 |
feat(config): T-1262 — parity is facts and exit codes, not bytes
schemas.py becomes pydantic, so the reference domain is the normal pattern rather than an exception carrying a footnote. Frozen: a result is a statement about what was found, and nothing downstream should edit the finding on its way to being reported. pydantic stays off the --help path — test_lazy_domains still passes, which is precisely the assertion that it loads with the domain and not with the CLI. The acceptance criterion could not be met as written, and that is the finding worth keeping. It asked for byte-for-byte parity with the old script; D-263 was amended after this ticket to give reach a streaming model that puts the verdict on stderr, while the old script writes its success line to stdout. Measured: the text is byte-identical in text mode, only the stream differs. Matching both would mean abandoning streaming or special-casing every ported gate. So parity is redefined, and it is stronger than bytes where it counts: exit codes match exactly, no fact the old message carried is lost, and failures name a remedy as a structured field. That governs every port in T-1251, not just this one, so it is in D-263 rather than only here. test_check_parity.py runs three paths — ok, drift, missing file — through both implementations and compares. It builds a throwaway fixture repo and copies the OLD script into it, because that script resolves its root from __file__ and has no override; the new command just takes SR_REPO_ROOT. That asymmetry is part of why the port earns its keep. It also asserts the failing paths actually exit non-zero, without which "the exit codes matched" would be vacuous for two checks that both silently pass. Proven to fail twice before being trusted. Once by accident: the first version asserted the yaml version appears on every failing path, which the old script does not report when the client file is missing — the test was wrong, not the code, and it now derives expected facts from what the old output actually contains. Once on purpose: mutating the router to drop a version made it fail and name the missing fact. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
f4cca69cab |
feat(config): T-1261 — reach is a bare name on PATH, in every context
`uv tool install --editable` puts the executable in ~/.local/bin rather than .venv/bin, which is the difference between a command that works everywhere and one that works only under an activated venv. Agents and git hooks never activate one. Verified in the three contexts that matter, with a negative control so the passes discriminate: a stripped non-interactive shell, a REAL git hook process (via git -c core.hooksPath ... hook run pre-push, not a simulation), and an agent Bash call — all with VIRTUAL_ENV unset. With ~/.local/bin removed from PATH the same check reports NOT-FOUND, so this is not passing because a venv happens to be active. Found a silent interpreter fork while doing it, which is this initiative's own failure mode wearing a different hat. uv tool install without --python picked CPython 3.11 for the tool environment while .venv and system python are 3.14 — uv selects the lowest interpreter satisfying requires-python. reach would have run on one interpreter and the test scripts on another, with different wheels for numpy/scipy/PIL, and future 3.12+ syntax would break the tool while the venv stayed green. PYTHON_VERSION now pins both. make setup-venv is rebuilt on uv, per the T-1258 finding that it called .venv/bin/pip against a venv that has no pip. The first fix was wrong too: plain `uv venv` fails on an existing venv, so the target was not idempotent where the version it replaced had been. Caught by running it twice instead of dry-running it — which is how the original rotted unnoticed. make install-reach self-checks that reach is actually on PATH afterwards rather than assuming it. make reach-repoint gives a name to the situation where uv keeps resolving a deleted worktree: reach still runs, edits in the main checkout do nothing, and there is no error message. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
b9d81ac694 |
feat(config): T-1260 — reach lists its domains without importing them
`reach --help` renders from a declaration table and imports nothing. The cost of help is now flat as the registry grows, which is the property that has to hold going from one domain to a dozen. The trap is real and was confirmed in typer's vendored source rather than assumed from upstream Click: TyperGroup.format_commands loops over list_commands calling get_command on each, purely to read a short help string off the loaded command. With lazy loading underneath, that imports every domain in the registry to render --help — while the output looks entirely correct. Nothing observable changes; only the import graph does. So the test asserts on sys.modules, and it was proven to fail before being trusted. Disabling the format_commands override made it fail and name the cause, listing all five leaked check modules. It also carries a positive control — invoking a domain must import its service — because without one, "nothing was imported" would pass equally for a loader that is simply broken, and it fails on an empty registry, which would otherwise satisfy everything vacuously. The check domain is created here because the test needs a subject: a stub raising NotImplementedError would have been committed dead code. That takes the port out of T-1262, which is rescoped to what it still owns — pydantic schemas, byte-for-byte output parity on the drift path, and the failure tests. The old tooling/check-client-version script stays in place and stays wired to the pre-push hook; the deprecation window is deliberate. One Typer behaviour worth knowing before every future domain: a single-command app collapses into a bare command, so `reach check client-version` failed with "unexpected extra argument" until the router got a callback. Same mechanism as the root callback, different symptom. Help now works at every level, closing item 5 of T-1248. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
8d64800fe9 |
feat(config): T-1259 — reach is a real command, and Typer vendors Click
`reach --help` runs from the console entrypoint in 80 ms. typer 0.27.1 and pydantic 2.13.4 join the dependencies, both CVE-checked against NVD, OSV and the GitHub Advisory Database. The design in the ticket did not survive contact. It specified a click.Group root, on the reasoning that it would keep typer off the --help path — but typer vendors Click as of 0.26.0, so there is no top-level click package to import and no supported way to extract typer's internal one. A click.Group root hosting Typer sub-apps would put two Click implementations in one process. The root is therefore a typer.Typer, and lazy registration will go through the supported typer.Typer(cls=...) surface with a TyperGroup subclass. T-1260 is corrected to match. The callback is not decoration: a Typer root with no commands AND no callback raises at build time, and lazy registration means no command is ever eager. The ticket claimed a zero-command root always raises — half right, and the half that matters is that a callback makes it legal. rich_markup_mode=None is load-bearing rather than cosmetic. It takes an empty --help from 168 ms to 74 ms, and keeps rich and pygments off the import path entirely rather than merely skipping the render. It also stops typer drawing box-art help, which it does even when stdout is a pipe — that would have put box-drawing characters into every hook log and agent capture. typer-slim was considered and rejected: deprecated since 0.22.0, now a shallow wrapper that installs all of typer. D-263 amended: the feels-instant ceiling goes from 250 ms to 500 ms. A ceiling is not a typical and most invocations sit far below it; the tighter number was buying discipline that the import-graph assertion enforces better. Stay smart about what loads, stop worrying about tightness. Security, checked 2026-08-23. typer has no advisories on record. pydantic 2.13.4 clears PYSEC-2026-1812 (email-regex ReDoS, fixed in 2.4.0) — and the 2026 SSRF advisories CVE-2026-25580 and CVE-2026-54249 are against pydantic-ai, a different package that is not a dependency here, recorded in pyproject so the next sweep does not re-panic. Transitively, pygments 2.21.0 clears CVE-2026-4539. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
559f3d82dc |
chore(config): T-1258 — tooling/ becomes an importable package
The skeleton the reach CLI hangs off. Nothing moves yet: this adds the package, the bounded core/, and explicit setuptools discovery. core/console.py is the single output path, and the split it enforces is the whole design — stdout carries the command's actual output so `reach ... | jq` keeps working, stderr carries the event stream as JSONL. Rendering happens at the sink: a terminal gets human text, anything else gets raw JSONL, so a live view and a job log are one artefact in two presentations. Emitting is optional — the gates emit nothing — and verdict() prints once, last, carrying its remedy as a structured field. core/config.py resolves the repo root from __file__ against a project.yaml sentinel, with an SR_REPO_ROOT override. No subprocess and no git call: this is on the gate path, and cwd is not a reliable signal anyway since a hook runs from the root and an agent call may not. Both paths are validated, because a silent fallback is how you end up editing one checkout and checking another. Discovery is configured explicitly rather than left to flat-layout auto-discovery, which would have had to choose between erroring on the ambiguity and quietly shipping client/ or docs/. Verified: top_level.txt contains exactly "tooling". Verified beyond the happy path — the sentinel rejects SR_REPO_ROOT=/tmp and names both remedies; debug events are suppressed at the default threshold while the verdict is not; stdout stays clean with stderr redirected away; and the three unconditional push-gate checks still pass now that tooling/ is a package, which was the real regression risk. Two findings recorded on the tickets. make setup-venv is stale — it calls .venv/bin/pip, but the venv was created by uv and has no pip, so the recorded procedure and the actual state have already diverged (T-1261 owns the fix). And settled-reach-tooling had never actually been installed: site-packages held the dependencies but no dist-info, which follows from there being no __init__.py to expose. This is the first commit where `import tooling` means anything. .venv/ was only ignored via .git/info/exclude, which is machine-local, so a fresh clone or a new worktree did not ignore it at all. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
5df8afedb9 |
docs(governance): D-263 — output parity over timing, and commands that stream
Three amendments, all from pressure-testing the record against how the CLI will actually be used. The ~104 ms push-gate ceiling is withdrawn. It was the summed cost of three single-sample timings, imported as a requirement without asking who pays — and who pays is the pre-push hook, which already runs cargo test or the gdUnit4 suite on any code push. A few hundred milliseconds is invisible there, and on a governance-only push the whole hook is about a second. The criterion is OUTPUT parity: a ported check must produce the same output and the same exit code as the script it replaces, and is not required to be as fast. What replaces the ratchet is a ceiling with headroom — under ~250 ms to feel instant. Lazy registration stays mandatory, justified by the real threat rather than by parity: scipy.ndimage alone is 275 ms, and an eager entrypoint would pay ~460 ms before executing a line of its own. That budget change removed the only argument for keeping pydantic out of the gate domain, so the carve-out goes with it. One fewer exception, and the reference implementation is now the normal pattern rather than a footnote. Commands also stream. The gates are milliseconds but the generators are minutes, and an agent Bash call gives up at two and sends nothing. Detaching alone would fix the timeout and keep the silence; streaming fixes the part that costs real time — you learn a generator is wedged at minute one instead of minute nine. JSONL events on stderr, stdout reserved for actual output, rendering at the sink so a job log and a live terminal are one artefact in two presentations. Reattach is a byte offset into an append-only file, which is why there is deliberately no daemon. The trap, recorded because it would quietly undo the thing this record cares most about: streaming is ADDITIVE to the failure contract. A remedy emitted at line 400 of 900 is printed and invisible, so the verdict still prints once, last. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
bbd64307ab |
docs(governance): D-263 — one CLI named reach, and Q-124 answered
Q-124 asked whether the 123-file Python tooling should be retooled into a Rust CLI. The answer is no, and it is a costing rather than a preference. All three frictions it names — per-script permission prompts, the venv/PATH split between interactive and non-interactive shells, and interpreter startup paid four times per push — are packaging problems, and one bare command on PATH with lazy subcommand loading fixes all three. Rust would additionally owe a numerical-equivalence proof on the planet-gen path, whose heightmaps are committed build artefacts with goldens standing on them: a large one-time cost to avoid a small recurring one, paid in the currency the project can least afford to spend. D-263 fixes the shape. tooling/ becomes an installable package behind the `reach` command: a routing-only main.py, every domain under domains/<name>/ split router/service/schemas/helpers, a core/ bounded on day one to what has no domain, logging and error handling attached as decorators rather than call-site discipline, and pydantic confined to domain schemas — measured at 87 ms against a whole gate check of 20-46 ms, which is why it must never reach the push path. Failures carry the command that fixes them and keep their exit code; a tool that explains itself and exits 0 silently disables its own gate. R-014 records the Rust option as costed down, not argued down, with the condition under which it is worth reopening. T-1247 files the work as eight dependency-ordered epics; only the skeleton is unblocked. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
284ce847c3 |
docs(governance): Q-124 — one door, domain split, and failures that teach
Jeroen's shape for the tooling CLI: move the Python into a package with a proper domain split, one door that answers everything with help, and errors that hand back instructions rather than a status. The domain split turns out to be discoverable rather than invented. tooling/ is 85 top-level entries — 37 loose .py, ~36 extensionless executables, 11 dirs of which only 6 hold anything — across four coexisting naming conventions. But the domains are already encoded as filename prefixes: blender x14, atlas x8, generate x7, check x7, then visual/validate/test x3 and godot/garment/pql/install x2. Those prefixes are the subcommand groups, which is what makes the consolidation mechanical enough to be safe. Two constraints recorded against "a new prompt not an error code", because taken literally each would break something: - Exit codes stay. Four of these run in the pre-push hook, which fails a push ONLY by non-zero exit; a tool that explains itself and exits 0 silently disables its own gate. That exact failure was observed in clide today, where unsupported-format, no-such-file and unknown-subsystem all returned 0. So: code AND message, never either/or. - It must not become literally interactive. Agents and git hooks have no TTY, and the tea scar is already written down — its prompts "crash in Claude Code (no TTY)", which is why every tea call passes all flags explicitly. Any prompt must be TTY-gated and suppressible. pql was cited as the precedent and measured rather than assumed. The principle holds there for unknown subcommands (full usage dump) and not for invalid values: `ticket status <id> nonsense` says invalid without naming the six legal values it knows, `ticket new` says "accepts 2 arg(s)" without naming which two. The gap is the closed sets, and it is the more common failure. Logged upstream as pql T-112 rather than worked around here — the bar for our CLI is the stronger one: whenever the accepted set is known, print it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
3a640f91f7 |
docs(governance): Q-124 — Typer costs the cheap option down, and moves the target
Jeroen raised Typer as the Python-CLI option. Costing it changed what the question is actually about. The repo is already most of the way there: pyproject.toml exists, `make setup-venv` already does `pip install -e ".[dev]"`, and 22 tooling files already use argparse. What is missing is a single line — there is no [project.scripts] entry at all, so no console entrypoint exists. This is consolidation, not authorship, and it resolves the largest friction (per-script permission prompts) for one allowlist entry. But the framework is the second decision, not the first. A [project.scripts] entrypoint lands in .venv/bin/, which is on PATH only when the venv is activated — and agents and git hooks never activate it. That is the same split VENV_PY already papers over in the Makefile, and precisely the failure recorded for tea: an absolute path breaks the Bash(tea *) rule and prompts every time, fixed only by a bare name on PATH. So the deliverable is "one bare command reliably on PATH" (uv tool / pipx into ~/.local/bin, or a symlink), and a Typer app behind an absolute venv path would solve nothing. Two honest costs recorded against it: Typer and Click are further venv dependencies, so it does not help the venv friction at all; and a single entrypoint importing every subcommand eagerly would pay all 123 modules' import cost on every invocation, four times per push. Lazy subcommand registration is therefore mandatory rather than an optimisation, and must be measured before and after. Net: this looks like the answer for the check/gate family and the day-to-day scripts, and it leaves the numpy/scipy/PIL planet-gen path alone — the part a Rust port would have had to prove numerical equivalence for. T-1246 updated to start here. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
6949f800dc |
docs(governance): D-262 — the wiki generator flow has one canonical map
The relationship between wiki/, the generators, systems.db and the runtime is
a directed graph with two edges running opposite to the obvious direction and
one running backwards into its own producer. Prose renders that badly: every
document that has described it states a single ownership direction and is
therefore wrong about part of the tree. D-262 makes the diagram the source of
truth and points CLAUDE.md, Skill(wiki), project-structure.md and
wiki/GOVERNANCE.md at it.
The correction that matters most: body pages were described everywhere as
machine-owned and reverted on sync. They are not. scaffold_bodies.py writes
one once and never overwrites it, and import_economics then reads that
frontmatter directly as input — so a hand-edit is not reverted, it is obeyed,
and silently changes world generation. Worse than being overwritten, and the
actual reason GOVERNANCE.md forbids the edit.
New: tooling/check-dataflow-graph.py, wired into the Makefile and the pre-push
hook. It asserts every repo path named in a hand-authored diagram still
resolves — and its docstring states plainly what it cannot do: verify that an
edge still MEANS what it says. If wiki_sync.py stopped writing body pages
tomorrow, every path would still exist and the check would still pass. Edge
semantics stay a human check against the tool's source, so nobody reads a green
gate as a verified map.
Verified by breaking it: pointing one label at a moved path fails with exit 1
naming that path; restoring it passes. Building the checker also caught two
real vaguenesses in the diagram — "GJ-*/index.md" and "bodies/{id}/index.md"
were written without their wiki/star-systems/ prefix, which is precisely the
ambiguity this map exists to remove. Generated star-map .d2 files are excluded
by name; their correctness belongs to their generator under D-223.
Also files Q-124 + T-1246 (tooling): whether the 123 Python files under
tooling/ should become one Rust CLI of pql's calibre. The friction is real and
mostly not about the language — the permission gate prefix-matches whole
command strings and a blanket Bash(python3 *) grant is forbidden, so each tool
prompts near-individually, while a single binary is one allowlist entry. The
record requires pricing the cheap alternative (a Python dispatcher entrypoint)
before recommending Rust, and flags the hard constraint: import_economics is
stamped by source SHA, so any port must keep that contract intact through the
transition rather than disabled during it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
0817befcba |
docs(diagrams): SVG replaces PNG, and a map of the wiki generator flow
d2 emits SVG natively; its PNG path wants a ~150 MB headless-Chromium download and prompts interactively, so every PNG here was produced by an out-of-band magick step. There is no Chromium on this system. Dropping PNG removes the dependency rather than trading one format for another, and cuts docs/diagrams/ from 17 MB to 3.3 MB. SVG renders in Gitea and in clide (`clide draw --file <path>`, which takes .d2 source directly), and diffs as text. One PNG is kept on purpose: design/star-map-concentric.png has no .d2 source. Also renders the 7 star-map .d2 files for the first time. star-map-plan.md listed their renders as a deliverable in March and the step never ran; the new `make check-diagrams` is what surfaced it. New: docs/diagrams/data-flow/wiki-generator-flow.d2 — which way the arrows point for any file under wiki/. Every edge was read in the tool's own source rather than inferred. It records the trap that keeps costing us: scaffold_bodies.py writes a body page once and never overwrites it, and the generator then reads that frontmatter directly — so a hand-edit there is not reverted, it is obeyed, and silently changes world generation. Two rendering traps found the expensive way and now written down: - A d2 `|md` block becomes an SVG <foreignObject>. ImageMagick and flutter_svg both silently drop it, so the legend was in the file and invisible in every viewer except a browser. Plain labels render as real <text> everywhere. - Container boxes fight the layout engine. Grouping nodes whose flow-depths differ forces long edge routes; this diagram went from an unreadable 2.4:1 sprawl to a legible 0.75:1 by deleting five containers and changing nothing else. Colour classes carry the grouping instead. make diagrams / make check-diagrams render and gate. Repo-specific rules in .claude/rules/diagrams.md; d2 syntax and the traps live in the user-scope d2-diagram skill, whose PNG default was flipped to SVG to match. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
34e5d7b636 |
docs(meta): body pages are obeyed, not reverted — correcting the wiki skill again
Jeroen asked whether I had read the python that writes the frontmatter. I had not — only grepped it. Reading body_definition_parser.py and scaffold_bodies.py properly overturned what I had written twice today.
scaffold_bodies.py NEVER OVERWRITES ('Only creates files that don't exist yet... existing body index.md files are skipped'), and the generator reads that frontmatter directly. So a hand-edited body page is not reverted, it is OBEYED, and it silently changes world generation — worse than being overwritten, and the actual reason GOVERNANCE.md forbids it. System pages behave the opposite way: wiki_sync.py re-renders their READ-ONLY blocks, so edits there ARE reverted. Three cases, not two.
It also explains T-1244's whole measurement: body_definition_parser resolves each field override > direct read > derived > inferred > SEEDED RANDOM. Continuous axes vary because they fall to the random tier; categorical axes are concentrated because they are read from the bodies table. The variance question belongs to the atlas CLI catalog, not the wiki.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
0dc68dc1b8 |
docs(meta): fork-for-sidequests rule + wiki skill corrections from the cold test
The cold test worked as an experiment: a fresh agent with Skill(wiki) cited it first, refused to hand-edit body frontmatter, knew the corp regen-db stamp trap, and knew corp_specialization is missing from its own template. It also found four things the skill had wrong or missing, all verified before folding in: body frontmatter is a MIDDLE layer (atlas CLI -> systems.db -> scaffold writes the page -> import_economics reads it back), not the origin GOVERNANCE.md implies; the four empty categories are Q-118, an open scope question rather than an invitation; some bodies are visual-regression goldens and nothing in wiki/ says so; and status is editorial, not an import gate. Also: check current state before editing, since the test's own task described a change that was already true. T-1244 corrected in the same pass — tectonics is derived from planet_class via a lookup (body_definition_parser.py:563), so the measured 68% 'low' is a projection of the class distribution, not an authoring choice. The ticket's question changed accordingly. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
9fb74bdf22 |
chore(meta): record the mechanism behind the roadmap gap on T-1245
Jeroen: 'I tend to restrict future side quests to not confuse your context.' A reasonable practice with a bad side effect — intent stays conversational and reaches the repo by accident. Resolution recorded as two channels rather than more sharing: working context stays narrow, forward intent gets FILED. Carries a design constraint into the initiative's form investigation — weight options by cost-to-APPEND, since these facts surface mid-bug and a ritual will not get used. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
4de90529ae |
chore(meta): file T-1245 — a roadmap that carries intent, not just work order
The cascade owns order, the ticket tree owns decomposition, the DQR tree owns individual rulings; none answer what the game is going to be. Evidence it is a real gap: three roadmap-level facts surfaced in one conversation on 2026-08-20 that exist in no artefact, and two of them were written up as suspected defects by an agent reading carefully, because nothing recorded them as intent. Pickup instructions make epics an OUTPUT of a harvest/interview/investigate-form/propose pass, explicitly not an input. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c25af8d753 |
docs(meta): make the wiki seed reachable — blind-prediction experiment and its fix
An experiment, at Jeroen's request: predict how the wiki seed data is structured
WITHOUT reading it, seal the prediction, then score it. The prediction is
|
||
|
|
a1addf7e21 |
docs(meta): seal a blind prediction of the wiki seed structure before reading it
Written before opening wiki/, committed first so the prediction is timestamped and cannot be retrofitted once the answer is known. Contamination declared inline: paths already seen this session via generator_sources.py and visual_scenarios.gd are marked [SEEN] and discounted. Scoring rule fixed in advance, with 'load-bearing thing I did not know existed' as the bucket that decides what gets documented. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
fab70edd5a |
feat(ui): the stipple says WHERE the ground is broken, not just that it is (T-1194)
RimWorld's technique 4, the reference this ticket names, works by CONTRAST: the
Rockies and Appalachians carry dense hatching and the Great Plains carry none.
Ferrath's Global carried an even wash of dots over every landmass instead —
texture present, information absent.
The cause was a constant measured on the wrong rung. RUGGEDNESS_FULL_SCALE_Q = 4
comes from Region ("d4 2.38"), but the elev_q path it gates only ever RUNS at the
orbital rungs, where relief_q is flat and the stipple falls back to it. Ferrath
Global measures a mean 1-cell elev_q gradient of 0.99, so a coherent 4-cell
baseline reaches ~4 and saturates the constant exactly: every land cell read as
fully rugged.
It now normalizes against the canvas's own measured gradient — the same
self-calibrating shape T-1240 gave the hillshade, and for the same reason: one
constant cannot serve rungs whose sample spacing differs by four orders of
magnitude. A contrast curve rides on top, because even unsaturated the linear
reading puts ordinary ground mid-range and paints grain everywhere.
Measured on Global, before -> after: 1,679 -> 1,900 distinct colours, 145.90 ->
147.39 lum spread, and the pale uplands now stipple visibly denser than the
lowlands beside them.
Recorded because it cost two wrong turns: I first guessed saturation, then
talked myself out of it after measuring a 1-CELL gradient (0.99) against a full
scale meant for the 4-CELL baseline, and shipped a contrast curve alone — which
measurably did nothing (1,679 -> 1,698), because a curve cannot separate values
already clamped to 1.0. The gamma is kept; it does its job now that there is a
range to curve. The elev_q gradient joins relief_q's in the capture readout, so
the next person tuning this can read the number instead of guessing at it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
646db131a9 |
chore(meta): close T-1240 — Region reads as terrain, acceptance ladder shot
relief_grad across the cold ladder: Global 0.00, Region 1.08, District 0.30, Quarter 0.07. Closure note records the two premises the ticket got wrong (Nyquist is the aliasing limit, not a legibility one) and the coast-warp trade taken at Region, with its reversal path. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
9b146f9e1f |
fix(simulation): derive at the octaves a rung can actually reconstruct (T-1240)
Region rendered as fine uniform stucco while District and Quarter, on identical code, read as terrain. The cause was sampling: `min_wl_m` arrives as an LOD request and defaults to 0, so every invented octave contributed at every rung. MIN_WL_BANDS_M was meant to be the floor but is built from the rung's CELL SIZE (2 x DISTRICT_M), which stopped being the sample spacing at the D-255 extent inversion — a rung fixes EXTENT now and spacing falls out of the canvas size. The bands were off by roughly the cell count, and the served path never consulted them anyway. The cutoff is now derived from the resolved spacing, which is what this ticket asked for. Two things had to be measured rather than reasoned to get it right, and both corrected me. FIRST: the field was the culprit, not the renderer. I attributed the stucco to the client stipple painting noise onto a smooth field. Surfacing the terrain layer's own mean |relief_q gradient| in the capture readout settled it in one shot: Region 18.24 steps per cell — 144 m of relief between NEIGHBOURING cells — against District's 0.30 and Quarter's 0.07. The server was sending noise. That diagnostic ships here for the same reason `plane_variety` did in T-1213: a noisy field and a renderer inventing noise look identical, and one number separates them. SECOND: Nyquist is the wrong threshold. The first version floored at 2 x spacing, the aliasing limit, and Region barely moved (56.16 -> 59.73 lum spread, gradient still 18.24) because 2 samples per cycle is unaliased but renders jagged. The rungs that already worked say what the real bar is: District reconstructs its finest surviving octave at 34 samples per cycle, Quarter at 135. At 8x, Region goes to 1.08 gradient and 70.01 spread, and shows ridges and valleys. THE TRADE, taken deliberately and recorded in the tests: an 8x floor also truncates the coast warp's 2,048 and 1,024 m octaves at Region, the band T-1160 added for "one coastline at every rung". An earlier test here asserted that band must survive; it now asserts the opposite. Same reasoning as the relief: a 1,024 m coastline wiggle at 379.3 m per cell is 2.7 samples per cycle, so drawing it draws noise rather than coastline character — a rung cannot show shape finer than its own cell. The warp is amplitude-capped sub-pixel on the working grid, so what is lost is small. If a future pass wants the warp exempt, the fix is a relief-only floor threaded through derive_at_metres, NOT a lower multiple, which takes the stucco back. Global is exempt: its floor would be ~70 km and would truncate the whole warp band, and it needs none — the orbital derive leaves relief_q flat at 50. District (3.79 m spacing) and Quarter (0.948 m) floor below every octave in play and derive byte-identically, which their own test pins. Cache-safe by construction: the floor is a pure function of (rung, extent, body_radius), all three already in the step-canvas cache key. 0.4.12 is required anyway — this changes derived BYTES at Region, so a 0.4.11 entry holds a field this build would never produce. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
6d4c9e92c2 |
feat(ui): the copse and the rocky outcrop, drawn per cover class (T-1213)
D-258 invariant 2 asks the map to show the minority the orbital summary
suppressed — clearings, marsh, rock, scrub inside a cell that reads "forest" from
space. composition.rs goes to real trouble to invent it, and the conservation
harness measures it: over a District patch on Ferrath the tally is {Barren: 175,
Forest: 16209}, so 1.07% of that ground is exposed rock.
The renderer was averaging it back out. One test, `veg >= Scrub`, at one density
and one strength: Forest, Scrub and both Riparian classes drew the IDENTICAL
mark, and Barren drew none at all. A wood looked like scrub, and bare rock was
invisible by construction — on the rungs whose whole purpose is to show what the
summary hid.
Each class now has its own grammar, differing on the three axes a mark has:
density (how much of the class's ground carries it), strength (how far it moves
the base colour), and lattice (the block size marks are decided on — bigger reads
as a clump, smaller as grain). Rock LIGHTENS where everything else darkens, which
is the point rather than a flourish: bare stone catching the light is the one
cover type brighter than the ground around it, so it separates from vegetation by
sign alone and can never read as "denser plants". It gets its own ScatterField
salt so an outcrop does not preferentially land where a copse already did.
Tuned against captures, not guessed. A first pass gave Forest a 4x4 lattice at
52% density, which produced visibly axis-aligned dark SQUARES — a 4-cell block is
8 screen px at District — and read as an artefact laid over the hillshade. It
also mistook the background for a feature: Forest is 98.9% of this frame, so
marking half of it dark is not "the occasional copse", it is a second colour
layer. Pulled back to grain at a 2x2 lattice, and the occasional thing is now the
thing that catches the eye: the outcrops read as scattered pale clusters of
exposed ground, exactly the "occasional copse/tree/rocky outcropping" that was
missing.
District holds its form through the change (lum p1-p99 74.15, against 74.43
before the cover marks and 13.72 when the rung was flat).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
8787ee1844 |
feat(ui): hillshade the deep rungs — form from light across slope (T-1213)
relief_q now reaches the renderer, and the first pass spent it on brightness:
lighten where the ground is high, darken where it is low. That moved the numbers
(District 13.72 -> 77.01 lum spread) and still looked like moss, because the eye
does not read landform from absolute brightness. It reads it from light falling
ACROSS a gradient — height-shading gives a rise and a fall the same tone, so no
ridge ever reads as a ridge.
So relief drives a proper hillshade: the local gradient of the field dotted with
a light from the upper-left. The light direction is not a free choice; lit from
the lower-right the brain inverts the read and valleys pop out as ridges.
THE SCALE IS MEASURED PER CANVAS, not fixed, and the first attempt at this failed
exactly the way this file already warned a fixed gradient constant would (see
RUGGEDNESS_BASELINE_CELLS: "the same 4-cell delta reads 21.86 at Region and 0.08
at District"). With a constant full-scale of 8:
Region 81.72 but District 77.01 -> 20.01, Quarter 42.56 -> 16.44
because at District's 3.8 m per cell neighbouring cells barely differ. The
terrain layer now measures each canvas's own mean |gradient| once per rebuild and
the hillshade normalizes against it, so one constant works at every rung.
Ladder (tooling/atlas-flatness, lum p1-p99), flat -> shipped:
Global 145.69 -> 145.69 unchanged; relief_q is flat 50 at orbital
Region 33.59 -> 54.30
District 13.72 -> 74.43
Quarter 11.01 -> 73.72
District and Quarter now read as terrain — ridgelines, valleys, and the stipple
organised into contour-like bands. Judged by eye on the captures, not by the
metric alone.
Stipple full-scale 25 -> 60. The old value was calibrated against a relief_q that
never arrived, so it was tuned to the elev_q fallback; with the real plane nearly
every land cell earned a mark and Region read as static (17,599 distinct colours,
more than twice Global's, for a quarter of the legibility). Form comes from the
hillshade now; the stipple is grain on top of it.
REGION IS NOT FIXED, and the cause is T-1240 rather than this change. It renders
as fine uniform stucco: a Region cell is 379 m of ground while the relief field's
content sits in the 128-1024 m band, so the field is at or below Nyquist and the
gradient the hillshade reads is aliasing, not slope. min_wl_m defaults to 0 on
the served path, so nothing truncates the octaves Region cannot resolve — which
is precisely what T-1240 proposes to fix. That ticket said the stale cutoff was
"currently inert"; it is now the thing capping Region, and T-1240 is updated
with the measurement.
Three tests, on direction rather than magnitude so tuning does not rewrite them:
a hill's west flank lit and east flank shadowed, a uniform field shading nothing,
and the canvas edge not drawing a rim. That last one is a bug this nearly
shipped: `_l8_value` returns 0 out of bounds and 0 on relief_q means MAXIMUM
HOLLOW, so sampling off-canvas posts a full-scale false gradient all the way
round the frame. The sample position is clamped instead.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
b9cd26429e |
chore(meta): file T-1243 — fog perf test flakes under gate load
The pre-push gate rejected the T-1213 push on a wall-clock fog budget (0.606 vs 0.5 ms) that passes 23/23 in isolation on the same build. Second hardening cycle for the same failure mode: min-of-7 defends against one slow sample, not the sustained core saturation the gate itself creates by running cargo and tooling suites immediately before it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
3ec35b87c8 |
fix(client): the deep rungs were flat because relief_q fell off the wire (T-1213)
`relief_q` is the one field with signal below District — elev_q's 80 m steps
quantise sub-district detail away, which is precisely why relief_q was invented.
The server has encoded it since
|
||
|
|
e5224b1a44 |
chore(meta): 0.4.8 — the T-1242 gate's first false positive, paid not dodged
A test-only edit to composition.rs tripped the canvas-generation gate, which is path-based and cannot tell an assertion fix from a generator change. Bumped rather than excepted: the ruling is that a false positive costs one round of cache misses and a false negative costs a week. Second no-op bump in two days, noted in project.yaml so the rate is visible if it becomes noise. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
869837f728 |
test(simulation): the conservation gate's monoculture check was a tautology (T-1213)
D-258 invariant 2 says descending the ladder must reveal COMPOSITION — a cell
reading forest must be able to contain the clearings and rock the vote
suppressed. One assertion stood behind that, and it read:
assert!(tally.len() > 1 || share == 1.0, ...)
A single-class tally has a 100% share by definition, so both branches are always
satisfiable: the check could never fail, including in the exact case its own
message names, "or nothing was composed". The invariant had a test and no gate.
Split into the two bounds the invariant actually has, because it is two-sided:
conservation caps how much may be invented (majority > 50%, already asserted) and
composition sets a floor on how little (minority >= 0.1%). Verified by raising
the floor to 2% and watching it fail on the measured 1.07%, then restoring it —
the floor is a tripwire for "did anything happen", deliberately far below the
measurement rather than tuned to it.
Measured at the descent ladder's own anchor on Ferrath:
conservation: majority class 3 at 98.9% across 2 classes {1: 175, 3: 16209}
So composition IS working in the data and conservation holds. The map is flat
anyway, and tooling/atlas-flatness (added here) says why the eye was not enough:
rung distinct lum p1-p99
Global 1581 145.69
Region 2923 33.59
District 53 13.72
Quarter 46 11.01
Region carries almost TWICE Global's distinct-colour count while holding a
quarter of its structure — the dither pass adds colour noise, not information, so
a colour-count metric would have called the flattest rung the richest. Structure
falls ~92% from Global to Quarter.
The cause is a channel mismatch rather than a missing generator: composition
perturbs moisture_q/slope_q, and the base map draws morphology hue x elev_q
lightness. The ladder scenarios pass no overlays deliberately, so the composed
fields are never rendered in the very shots that judge this work. Recorded on
T-1213 with the three ways forward; the choice touches D-258 and is Jeroen's.
The gate is still #[ignore]d — noted on the ticket as worth moving into a harness
that runs, since believability and window-derivation already load real bodies in
the normal cargo test path.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
6e6218d654 |
chore(meta): close T-1242
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
48fee8a0b6 |
feat(config): make the canvas-generation/version pairing a gate, not a habit (T-1242)
project.yaml's version is the Atlas disk cache's only invalidation signal, and nothing enforced that changing canvas GENERATION also moved it. It broke five times -- 0.4.2 lake_margin_q, 0.4.3 coast_warp_px, 0.4.4 the extent inversion, 0.4.5 the Global sentinel, 0.4.6 one-course-per-river -- each bumped only after someone noticed a wrong map. The failure is invisible to its author: it needs a warm cache to reproduce, so a cold checkout looks fine. T-1239 is the last one, and it took eight days. tooling/canvas_sources.py is the path registry; tooling/check-canvas-version rejects a push that touches those paths without moving project.yaml's version line. Wired into the pre-push hook, `make check-canvas-version`, and, for the parsing units, `make test-tooling`. Verified against real history rather than a synthetic branch: run over |
||
|
|
a1568d27c1 |
chore(meta): close T-1241
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
086d9ed56e |
fix(client): bake the version into the build, so an export can invalidate its cache (T-1241)
current_schema_version() line-scanned res://../project.yaml at runtime. That resolves to the repo root in a dev run and to nothing in an exported build, so a shipped game got the "?.?.?" fallback every time. Since that tag is the Atlas disk cache's ONLY invalidation signal, every exported build stamped and compared the same sentinel: a canvas cached by one build would be served by every later build, forever. T-1239 is what that failure looks like once it happens. loading_screen.gd carried a byte-for-byte copy of the same function, so the version shown to the player was "?.?.?" in exactly the builds where a version string is worth showing. Both call sites now share client/scripts/build_version.gd, which reads application/config/version out of ProjectSettings — a value Godot bakes into the PCK, identical in the editor and in an export by construction rather than by luck. No file IO, no fallback branch. project.yaml stays the source of truth (CLAUDE.md); client/project.godot mirrors it. A mirror nobody checks would be worse than the bug it replaces -- the old code failed loudly everywhere, a stale mirror fails silently -- so tooling/check-client-version compares the two and the pre-push hook runs it unconditionally. Not gated on "were those files in this push": drift persists on main once introduced, and gating would let an existing drift ride along. The test this replaces asserted that current_schema_version() did not return its fallback, and passed -- in the one environment where the code under test worked. Three tests now pin the property that actually matters: a real version, sourced from the baked setting, matching project.yaml. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
e8d522b482 |
chore(meta): close T-1239, file T-1241/T-1242 follow-ups
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
07ed2a47ab |
fix(client): the Atlas was replaying a cache from a build that no longer existed (T-1239)
Ferrath's Global map drew no rivers at native resolution: 375 courses arrived and 0 were drawn. The report suspected the D-261 length cull or the water truncation. Both were innocent, and so was the renderer. The client served the canvas from its own disk cache (T-1183). Every payload for GJ820Bc predated T-1237 ( |
||
|
|
16348e2e89 |
chore(meta): drop the no-op Write() twins from the permission lists
A Write(<path>) permission rule matches nothing. File permission checks consult only Edit(<path>) rules, which already cover every file-editing tool — Write, Edit and NotebookEdit alike. Claude Code now warns about the dead shape at session start. All eight removed here sat directly beside their Edit() twin, so the allow grant over the repo tree and the ask gates guarding settings and hook files kept working throughout. Behaviour is unchanged. That ask block remains the pattern worth copying to the other repos in this tree — it is the only one that stops an agent quietly widening its own permissions, and it has to be ask rather than deny to stay fixable. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
dd0f9804b5 |
chore(meta): install pql replication hooks and cover the format marker
This repo had no .pql/hooks/ at all — the four replication hooks were never installed, because pql's installer used to ignore a redirected core.hooksPath. It works with .config/hooks now, so init prepends a two-line shim to each hook that sources the pql half. Existing hook bodies are untouched; the shim goes above them. What this buys: post-merge now runs `pql plan upgrade`, so a pull that brings in a newer changelog format migrates it forward automatically instead of replaying under superseded rules. .gitattributes gains a rule for changelog files at the root of .pql/changelog/. The existing `**/*.sql` pattern requires a directory component and so did not match the new 0000-format.sql marker, which would have made it a merge conflict rather than a union merge. Note for a follow-up: the hand-folded pql block in .config/hooks/post-merge (lines ~10-12) is now redundant with the shim, so plan import and decisions sync each run twice per pull. Both are idempotent, so this is waste rather than breakage — but that block and its stale "installer is dead" comment can be dropped. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
962cbe83a7 |
chore(meta): migrate changelog to format 2.0.0, recovering 69 descriptions
pql 2.0.0 versions the changelog file format and carries older ones forward. The rewrite touches only the inline conflict guard on each line, which moved from a content-hash tiebreak to append position (3992 lines in, 3992 out — no row data altered). This repo carried real damage from the old rule. A ticket created and appended to within one wall-clock second produced two changelog rows tied on updated_at, and the hash decided the winner — arbitrarily, and on every replay, so the loss reappeared on each fresh clone and branch switch. Replaying the pre-upgrade changelog and diffing all 1232 tickets against the repaired state: 69 tickets gained description text, none lost any, 13270 characters recovered in total. Six had no description at all. T-1057, where this was first noticed, keeps the description a session hand-recovered from ticket_history in July; its later updated_at means the tie no longer decides it. The workaround scaffolding in that field can be tidied whenever convenient. plan rebuild --verify reports zero rows lost. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
3a38322ce9 | Merge remote-tracking branch 'origin/ocean-guard-synthetic' | ||
|
|
d5e617eff6 |
docs(simulation): PR #218 round 3 — the round-2 fix outran its own documentation
Three findings, all doc-accuracy, and all the same root cause: folding the spacing predicate into the ring walk changed what three comments describe, and two of those comments were written by this same PR one round earlier. TYRE 1 — road_graph.rs's T-1206 gap-closure comment cited `nearest_land_cell`, which round 2 made `#[cfg(test)]`. A reader chasing that name lands on a test-only function and reasonably wonders whether they are looking at dead code. Repointed to `nearest_cell_matching`, and the paragraph's closing claim that "T-1206 guarantees the placement pixel is land" is corrected: it has been land-AND-spacing-or-skip since round 2. TYRE 2 — `max_land_search_ring`'s doc named the same test-only wrapper as the thing that walks the bound. It now names the production consumer and both callers. TYRE 3 — the D-211 amendment was written in round 1, before round 2 existed, and still described a land-only correction. It now carries a dated refinement recording what the code actually does: the walk satisfies BOTH of step 4's promises in one search, and SKIP therefore also fires where land exists but none of it clears spacing within the bound. The no-re-decision conclusion is unaffected — position remains a deterministic, non-fabricated function of seed and terrain — and the refinement notes the spacing promise is step 4's alone, since Tier A/B/C placements sit on their matched attractor and were never subject to it. HOSHE's three findings were the same three hunks, observed uncommitted while the review ran: accurate content, but not in the branch tip, so the PR would have merged a governance record that misdescribes its own commit. That is this commit. Both reviewers independently confirmed what the round-2 fix claims. Tyre traced the ring geometry and tie-break order by hand against the spacing predicate; Hoshe re-ran the full 267-body corpus scan live (850s) and reproduced the figures exactly — 267 bodies, 267 reaching Layer 3, 344 placements, 109 synthetic, 0 in water, 0 spacing violations. The shared-ring-search-helper retraction is confirmed and settled, with NEW grounds rather than a restatement: round 2 strengthened the case for keeping them separate, since this walk is now parameterized by an arbitrary predicate over native u16 terrain coordinates while road_graph's is a RouteGrid method over downsampled routing cells with a fixed cost test and an unrelated bound. 22 module tests green; clippy and fmt clean. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
6cfd829445 |
feat(simulation): sub-cell composition — a summarised cell can contain its minority (T-1213)
D-258 invariant 2: descending the ladder must reveal COMPOSITION, so a cell reading "forest" globally contains the clearings, marsh, rock and scrub the vote suppressed. Below Global it contained nothing: morphology carried NINE zones at Global and exactly ONE (AlluvialPlain) at every rung under it, and vegetation collapsed to one class from District down. WHY THE EXISTING TIERS COULD NOT DO IT. vegetation_invention already perturbs moisture with a massif band and a texture band — but that pair was built for cross-rung coherence and is weighted 70/30 specifically so "the texture term alone can never outweigh the massif term". It is designed NOT to change a verdict, which is the exact opposite of what composition needs. And it could not simply be turned up. Measured: 50.2% of the texture field's amplitude sits in its 32,768 m octave alone, and everything at or below 2,048 m holds 5.9% of the total. Across a District window only that 5.9% varies, which after the 30% weight and a ~29-point ceiling swings moisture by +/-0.51 points against vegetation gates 5-15 points apart. Nothing could ever cross one. That is the geometric series, not a tuning shortfall — raising the ceiling enough to matter at District would make the field violent at Region. So a third tier carries the fine band ALONE, normalized to its own full swing: quiet where the coarse tiers are loud, loud where they have nothing left to say. It feeds BOTH classification inputs, because moisture alone would have left morphology just as flat. CONSERVATION IS THE BOUND, not weighting (D-258 invariant 3). The field is zero-mean, so a downsample returns the summary it was added to. Pinned two ways: a field-level zero-mean test, and a real-terrain test that derives a District-sized patch and asserts the majority vegetation class survives. RARE INCLUSIONS, and this was a correction. The smooth term is a gentle sway around the base, so it can only flip a verdict where the ground already sits near a gate — which made deep-in-class ground immune, and the conservation test duly measured a patch that was 100% Forest. A monoculture is the flat map this ticket exists to fix, one scale down. Jeroen: "maybe a dense forest should still sometimes produce a clearing or a rocky outcropping." D-258 says CONTAIN, not border on. A sparse high-contrast term now rides on top — thresholded value noise so inclusions are connected blobs rather than stray speckled cells. The same patch now reads 98.9% Forest with 1.1% Barren outcrops. The moisture half of an inclusion obeys the envelope rule (a world with no patchiness ceiling grows no glades — caught by the zero-ceiling test, which the first version failed by putting damp pockets on airless rock); the slope half does not, because an outcrop is geology and a dead world is exactly where bare rock should break the surface. The slope ceiling is 6, not the 15 first written. Measured against the window-derivation fixtures, base slope_q on ordinary ground is 2-6, so +/-15 did not vary the signal but REPLACED it — one fixture moved 6 -> 19 and two coastal samples flipped to Wetland on invented slope alone. The morphology gates are far apart because a cliff coast is a real landform; composition must let marginal ground fall both ways, never manufacture a fjord on a flood plain. NOT applied at the orbital rung, which is envelope-only by design and documents that it never invents slope — pinned by derive_orbital_at_metres_never_invents_slope, which caught the first version. Both goldens move in the IMPROVING direction, checked before updating rather than blind-refreshed: GJ338Bd moisture 75->85 distinct, slope 28->34 (range 0-50), materials 3->4 GJ244Ad moisture 25->42, slope 15->22, morphology zones 6->7, materials 2->3 Ladder effect (was -> now): Region moisture 25->39; District moisture 3->17 and vegetation 1->2; Quarter moisture 3->8. 45 server suites green, clippy clean. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
43a267439b |
feat(client): ScatterField — reusable seeded scatter for client-side paint
Extracted from the T-1194 stipple, which was the first of a family: graffiti placement, cracks in textures, drifting cloud cover — presentation decisions that must look the same when the player returns to a place, and which the simulation has no opinion about and should not be burdened with. THE LINE IT DRAWS. It answers "how is this drawn", never "what is here". A cell's biome, a settlement's position, whether a wall exists — those are world data, derived once by the server and sampled everywhere (D-255(f) mechanism B), and the player eventually stands on them; inventing those here would put the map and the ground in disagreement. Stated on the class so the next consumer does not have to re-derive it: if the answer changes what is THERE it is not a ScatterField question; if it only changes how it is DRAWN, it is. Bit-identity with the server's Rust noise is explicitly NOT a requirement (Jeroen: "a seed is a seed and the functional intended outcome is repetition here"). Nothing here is compared against a server value or round-tripped through a save, so the contract is stability across sessions, not agreement across languages — which is precisely why paint belongs on this side: it buys visual density with no cross-language determinism burden. Seeded from GameState.world_seed, so two playthroughs scatter differently and one playthrough is stable forever. API: domain() resolves a name to a salt ONCE (the first consumer runs ~700,000 times per canvas rebuild, so the hot calls take an int, never a string); value/chance/pick/jitter for discrete marks; smooth() for continuous fields like cloud cover; an optional time axis for animation. Domains keep consumers uncorrelated — without them graffiti and cracks at the same wall coordinate would mark identical spots and read as one artefact. The tests pin the CONTRACT, not the numbers — freezing outputs would make any future improvement to the mixer a breaking change for no gain. They caught a real defect immediately: (-x, -y) collided with (x, y), because negated coordinates produce negated products and the sign-bit mask folded the pair together, mirroring every mark west and south of the origin onto its north-east counterpart. Not an edge case — the descent ladder's own anchor sits at y = -5,675,959. Fixed by zigzag-encoding coordinates before mixing. 1853 client tests, 0 failed (15 new). Global capture re-verified unchanged after migrating the stipple onto the service. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
5eb394b36f |
feat(simulation): relief_q — a local relief signal the deep rungs can resolve (T-1213)
The District and Quarter rungs rendered as flat colour, and the cause was not the biome work everyone assumed. Measured on Ferrath through the production canvas builder: at District the mean |elev_q delta| between neighbouring gridunits is 0.02, and NOT ONE PAIR in a 1290x540 frame differs by 2. elev_q spans 0-100 across the body's whole 8 km elevation range, so ONE STEP IS 80 METRES. A District canvas covers 2,048 m of ground, where the rolling relief a walker navigates by is metres to tens of metres -- a fraction of a single step. The sub-district detail IS generated (invent_primitives' scatter and relief bands compute it) and then rounded away. Confirmed by running the diagnostic with the octave cutoff disabled: still 0.02. relief_q carries that same invented fine component against a scale chosen to resolve it: 0-100 about a flat 50, RELIEF_FULL_SCALE_M = 400 m either side, so 8 m per step -- ten times finer than elev_q. elev_q keeps its body-absolute meaning and the Atlas legend stays true. Measured effect, elev_q vs relief_q (distinct values / mean 4-cell delta): Region 49 / 2.38 -> 101 / 21.86 District 10 / 0.08 -> 35 / 0.35 Quarter 8 / 0.02 -> 19 / 0.06 FIXED metre scale, never per-canvas normalization: the value for a piece of ground must not depend on what else is in frame, or the same hillside changes tone as the viewer pans. And it excludes elev_pct deliberately -- this is the departure from the surrounding land, not height above sea level; including the base would re-introduce the body-scale dominance that makes elev_q unusable down here. 50 at the orbital rungs, which skip invent_primitives by design. Nothing is lost: Global and Region still have varied elev_q (101 and 49 distinct values), and the client takes whichever field carries signal via a max, with no rung-name branching. The client's ruggedness driver changes with it. It was an elev_q GRADIENT, which cannot work across rungs -- the same 4-cell delta reads 21.86 at Region and 0.08 at District, so any single full-scale constant either saturates one or vanishes on the other. relief_q states relief outright, so |relief_q - 50| is the answer directly and a fixed metre scale is immune to that by construction. An absent plane reads FLAT, not zero -- 0 on this field means maximum relief BELOW flat, so a payload without it would have stippled the entire map. That is reachable: the field is #[serde(default)] so old-shape payloads decode. Two colorize tests whose fixtures predate the plane caught it. 2004 server tests, 1838 client tests, 0 failed. clippy clean. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
566b566519 |
feat(ui): relief and vegetation texture over the terrain hue (T-1194)
RimWorld technique 4 from the reference map: texture as data, not decoration.
WHY IT WAS FLAT. The base layer reads hue from morphology and lightness from
elev_q. Below Global, morphology resolves to exactly ONE zone per canvas, so
the frame became a single colour whose only variation was a lightness ramp too
subtle to see. Measured on Ferrath at Region: 1 morphology zone, but 49
distinct elev_q values. The information was already on the wire and arriving —
the renderer was discarding it by expressing it in lightness alone.
TWO MARKS, NOT ONE SLIDER — the ticket's design question (b), settled by
looking at the reference rather than reasoning about it:
- RELIEF stipple: fine, dense, darker, keyed to RUGGEDNESS not height. The
reference's high flat plains carry none while its ranges are dense with it,
so the driver is the local elev_q gradient; a high plateau stays clean.
- VEGETATION blotch: coarser, softer, marked on a half-frequency lattice so
it reads as patches rather than a second speckle at the same pitch.
Inline in the existing per-cell loop (question (a)) and always-on, base layer
only (question (c)). The TMP/MST/VEG toggles are ANALYTIC reads — stippling a
temperature ramp would corrupt the quantity being read.
THE BASELINE IS MEASURED, NOT GUESSED, and the first attempt got it wrong: a
1-cell ruggedness delta samples mostly quantization noise, reads
near-identically everywhere, and rendered as uniform static over flat green —
grain, not structure. The gradient saturates by about 4 cells (Region: d1 1.40,
d4 2.38, d8 2.41, d16 2.51), so the baseline is 4 and the full scale 4.
Verified by capture at native resolution: on Global the stipple now
concentrates on rugged ground and leaves plains clean.
WATER TAKES NEITHER MARK, and gets a flat tone. An earlier version excluded
Lake alone and stippled the entire ocean — the one surface with no relief to
express. Both open-water zones are excluded now.
The ocean also stops shading by elev_q, which is the same argument T-1188
already made for lakes and never applied here: elev_q on a water cell is the
bedrock UNDER the water, not the surface, so shading the sea by it paints
seabed relief nobody can see. Near a coast that bedrock rises steeply and
quantizes hard, which is exactly where it showed — a pale, pixellated, broken
fringe hugging every shore (Jeroen, on the capture). One tone for the sea reads
as water and lets the coastline be the edge.
D-255(e)-legal throughout: texture-space dithering of already-derived per-cell
values, decided per server cell by a hash of its own coordinates and values —
no sample invented between cells, identical on cache hit and miss.
Two colorize tests updated: both asserted the old ocean shading incidentally
while testing zero-fill/no-crash. Property under test unchanged.
1838 client tests, 0 failed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
8eb5cb9ff7 |
feat(ui): the Atlas ladder bottoms out at Quarter (D-255, T-1213)
Quarter becomes the deepest navigable rung. Block and Chunk leave the ladder. The rule, from the amendment: the deepest Atlas rung is the one at which a screen pixel shows one subtile. At the uniform 2x2 px display ratio a 3440x1440 window gives 540 gridunits on the short axis, so Quarter's 512 m extent draws 0.948 m per gridunit -- about one voxel per gridunit and one 0.5 m subtile per pixel. Block (0.237) and Chunk (0.119) magnify beneath the finest datum that can exist, and measured as exactly that on Ferrath: one morphology zone, one vegetation class, an unbroken colour field. They were not missing a feature; there was nothing left to show them. CHUNK ITSELF IS UNTOUCHED. It remains D-243's 64 m stream/derive unit and is where Phase 5 derives first-person walkable content -- D-012's load-around-the-player is expressed in chunks. Block remains the 128 m generator planning unit. Both keep their enum variants, their extent_m answers and their wire vocabulary. What was retired is the claim that a MAP of one is worth looking at. DEEP_RUNGS moved with the floor, and this is the part worth reading twice. It is a mandatory D-255(d) hardening: a per-body retention cap bounding how much ground an exhaustive pan can hold resident at fine spacing. Left as [Block, Chunk] it would have guarded rungs no client can request -- dead code -- while the accumulation gap silently re-opened under Quarter, now the finest navigable rung at ~0.95 m per gridunit. It is now [District, Quarter]. A control that names its targets by rung has to follow the ladder when the ladder moves. Six tests pinned the old floor and were updated rather than deleted, since each was protecting a real property: the clamp tests now clamp at Quarter, and the disk-cache tests use Region for "shallow" (District is capped now) and Quarter for "deep". One new test pins the distinction the change turns on -- the retired rungs are absent from RUNG_LADDER but still present in RUNG_EXTENT_M, because viewability was retired, not vocabulary. Also removes the four capture scenarios for the retired rungs, including the two blank goldens that had been passing against blank captures. 1838 client tests, 0 failed. Server clippy clean, step_canvas suite green. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
b429f633e6 |
docs(meta): D-255 — the ladder floor is one subtile per pixel (T-1213)
Measured during T-1213, through the production canvas builder on Ferrath: at Quarter and below, morphology collapses to ONE zone and vegetation to ONE class. The uniform frames in the 2026-08-06 descent ladder were those rungs drawing exactly what they contain. The cause is arithmetic, not a missing feature. At the uniform 2x2 px display ratio a 3440x1440 window gives 540 gridunits on the short axis, so: Quarter 512 m -> 0.948 m/gridunit -> 0.474 m/px ~1 subtile per pixel Block 128 m -> 0.237 m/gridunit 4 gridunits per voxel Chunk 64 m -> 0.119 m/gridunit 8 gridunits per voxel Block and Chunk magnify beneath the finest datum that can exist, so they can only ever draw one voxel larger. Quarter lands within 5% of one subtile per pixel and becomes the floor. Stated as a rule so it survives the constants moving: the deepest Atlas rung is the one at which a screen pixel shows one subtile. It is derived from the data model rather than chosen, and it moves automatically if the subtile does. WHAT THIS IS NOT. Chunk remains the 64 m stream/derive unit of D-243 and stays vital — it is what Phase 5 derives first-person walkable content on, and D-012's load-around-the-player is expressed in chunks. Block remains the 128 m generator planning unit. Only Atlas VIEWABILITY is retired; the containment ladder is untouched. This record governs what the map draws, not what the generator builds. The justification is the Atlas's purpose (Jeroen): it exists to give the player information, and a rung earns its place by answering a question the rung above cannot. Once a pixel is a subtile there is no finer datum to answer with. The resulting Global -> Region -> District -> Quarter steps at ~93x -> 100x -> 4x. That unevenness is NOT from this change — the rungs removed were 4x and 2x steps carrying no information — it is D-243's one non-power-of-2 rung, and T-1218 already exists to re-balance it. A compensating rung above Region was considered and declined here; it belongs with that ticket. CLAUDE.md's cascade line updated in the same commit. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |