merge: reconcile Wave 1.1 with post-PR40 lab

Merge canonical lab 9557b8d5909eb4a885c3bf49e19a65dd904f8c1d exactly once.
Retain invocation journal ownership and lineage, provider terminal ordering,
teacher handoff, framed DONE handling, and canonical authority/Ajax routing.

Combine dynamic dispatch receipts with lab policy forwarding. Adapt native
shell/patch evidence, explicit TUI verifiers, and artifact recovery presentation.
Refresh generated configuration source links and strengthen adapter regressions.

Validation: focused 2118 passed; Wave 1.1 script 2291 passed; broad runtime
5649 passed; full pytest 11581 passed, 53 skipped, 2 xfailed, 6 subtests passed.
Compileall 1689 Python files; syntax 279 JS and 82 MJS files; diff and
conflict-marker checks passed.
This commit is contained in:
Alexandre Teixeira
2026-10-01 09:09:55 +01:00
304 changed files with 41744 additions and 24384 deletions
+89
View File
@@ -0,0 +1,89 @@
# Ref parity audit
`scripts/ref_parity_audit.py` reports which commits on one git ref left no trace
in another, and which files exist on one and not the other. It is read-only: it
runs `git log`, `git show`, `git diff`, `git grep`, `git ls-tree` and
`git merge-base`, writes nothing to the repository, touches no remote, and does
not import the application package.
## Why it exists
`lab` and the public `dev` line share only the repository's first commit as a
merge base, so `git log lab..dev` lists thousands of commits — nearly all of
which are in fact present on both sides, having arrived under different SHAs. A
plain log tells you nothing about what is actually missing.
The question that matters before `lab` becomes a release is narrower: is there a
fix on the public line that never reached `lab`? This script answers that by
sampling distinctive added lines from each commit and searching the other tree
for them.
## Running it
```bash
git remote add public https://github.com/odysseus-dev/odysseus.git # once
git fetch public dev --no-tags
scripts/ref_parity_audit.py --source public/dev --target lab --since 2026-08-10
```
Roughly 30 seconds for a 100-commit window; it grows linearly, so bound a wide
audit with `--since`. Add `--format json` for a machine-readable report and
`--output PATH` to write it to a file.
| Flag | Effect |
|---|---|
| `--source REF` | The ref whose commits are audited. Required. |
| `--target REF` | The ref searched for traces of them. Required. |
| `--since` / `--until` | Bound the commit range. Both filter **committer** date, which is also the date the report prints. |
| `--traversal linear` | Default. Individual authored commits, merges dropped. Finds a fix that arrived on a side branch. |
| `--traversal first-parent` | One row per merge into the source branch, which reads as one row per merged pull request. |
| `--probes N` | Probe lines sampled per commit, default 4. |
| `--exclude GLOB` | Extra path glob whose lines are not used as probes. Repeatable. |
| `--no-default-excludes` | Drop the built-in vendored / lockfile / binary exclusions. |
| `--top N` | Rows shown per file list, default 50. |
| `--repo PATH` | Repository to run in. Defaults to this checkout. |
## How a verdict is reached
For each commit in `target..source`, the script takes the patch with no context
lines, collects the added lines, drops the ones from vendored code, committed
build output, lockfiles and binaries, and keeps those that are at least 24
characters long and name at least two distinct identifiers. It ranks what is
left by how many distinct identifiers each line carries (length breaks ties),
takes the top `--probes`, and searches the whole target tree for each one with
`git grep --fixed-strings`.
Probes are stripped of leading and trailing whitespace, so a change that was
re-indented on the target still counts as present. The whole target tree is
searched, not the same file, because a ported fix routinely moves.
| Verdict | Meaning |
|---|---|
| **absent** | No probe found anywhere in the target. Treat as a real gap and read the diff. |
| **partial** | Some probes found. **Inconclusive.** A line can be rewritten by a refactor on the target and still be the same change. |
| **present** | Every probe found. The change is almost certainly there in some form. |
| **no-probe** | Nothing to sample: a deletion-only commit, or one touching only excluded paths. No verdict. |
## What is exact and what is a heuristic
**Exact:** the two file-presence lists. They come from `git ls-tree` on both
refs, so a file in "on the source and not the target" is definitely not there.
**Heuristic:** every commit verdict. It samples at most four lines out of a
diff that may be hundreds, and a probe can be absent because the area was
refactored rather than because the change was never made.
The two complement each other in a specific and useful way. A commit that reads
**present** while one of the files it added shows up in the source-only list is
almost always a fix whose production change was reproduced on the target without
its test. The line sampling cannot see that; the presence diff can.
Read the diff before porting anything. The verdicts say where to look, not what
to do.
## Tests
`tests/test_ref_parity_audit.py`. The end-to-end cases build a throwaway
repository with two branches off one root, so the verdicts come from git's own
`grep` and `diff` rather than from a fake.
@@ -0,0 +1,153 @@
# Wave 1.1 final post-PR40 reconciliation
This is the one-time local reconciliation of completed Wave 1.1 with the
authoritative post-PR40 lab commit. It does not start another runtime wave.
## Verified starting state
- Wave branch: `feature/agent-runtime-wave-1-1`.
- Original Wave HEAD: `63457367aeed431b2c48967988259e5861f19916`, clean.
- Canonical branch: `lab`.
- Canonical HEAD: `9557b8d5909eb4a885c3bf49e19a65dd904f8c1d`, clean.
- Merge base: `f0761641a12b63e401960f596d3d1be8fc90fbea`.
- Divergence: 10 Wave-only commits and 47 lab-only commits.
- Changed-file overlap: `src/agent_loop.py`, `src/tool_execution.py`,
`tests/test_tool_policy.py`, and `tests/README.md`.
The Wave-only commits were `d57d5c58`, `dfeab64a`, `ae2445d6`, `7d84f3fe`,
`1470dbb2`, `32830918`, `ba29afb9`, `bdfcbc0a`, `70cbaf81`, and `63457367`.
Their completed behavior is retained. The canonical worktree is read-only;
the exact canonical SHA was merged once with `--no-ff --no-commit`.
## Semantic integration
The only textual conflict was in `src/tool_execution.py`, where Wave 1.1
wrapped dynamic dispatch with `dispatched(...)` and lab added `disabled_tools`
and `tool_policy` forwarding. The resolution retains both inside the wrapper.
Lab's new owner-aware image-generation dispatch also receives that wrapper.
The image regression checks that explicit denial never invokes the backend,
actual dispatch has an execution identity, and a backend without an explicit
exit code does not manufacture an authoritative success receipt.
Broad validation exposed narrow adapter incompatibilities beyond the textual
conflict. Native host-shell JSON now uses the same decoded command classification
as journal evidence. The exact existing TUI interpreter-selection string is
shared with the evidence parser: a following foreground verifier keeps its
exit status, while generic conditional discovery, help/collection modes,
variable arguments, and status-masking tails remain insufficient test proof.
The generated interpreter-selection command itself is unchanged.
The generated environment reference is refreshed with the canonical generator
so its source-location links match the reconciled code.
Structured native patch arguments retain artifact targets. A pre-edit
inspection cannot invalidate a later passing executable verifier, but still
cannot verify the edited artifact by itself; failed post-edit inspections
remain failures. Artifact recovery's terminal round-text revisions retract buffered rejected drafts
before presentation; their replacement prose is still gated by journal
evidence. Explicit final-response events retain precedence, safe reasoning
survives, and provider-error partials and diagnostics retain their ordering.
Existing subprocess doubles now carry PIDs. Execution simulations use the
existing receipt-aware test helper. Contract tests assert the additional
completion-decision event and retain their no-inference/no-execution spies.
TUI tests retain tool order, retry behavior, and positive explicit-verifier
coverage while additionally rejecting completion from an opaque fallback.
Round-control fixtures explicitly fail unconfigured direct-provider synthesis
instead of contacting their fake endpoint, and supply the synthetic context
window while retaining real compaction logic. Conversational round provenance is
preserved outside artifact recovery.
The agent-loop changes merged automatically: lab's weather relevance and
policy-gated browser fallback coexist with Wave's action receipts, completion
gate, and deferred teacher handoff. The fallback dispatcher runs inside the
current invocation's journal. No generic tool floor was restored.
`src/agent_runs.py`, `routes/chat_routes.py`, `static/js/chat.js`,
`static/js/chatRenderer.js`, `src/tool_policy.py`, `src/tool_capabilities.py`,
`src/turn_contract.py`, `src/model_profiles.py`, and
`src/clean_agent_preview.py` retain the exact canonical lab content.
## Identity audit
These classifications describe every relevant identity use across the
detached-run manager, chat routes/browser consumers, completion gate, journal,
teacher handoff, and existing server-owned security provenance.
| Class | Uses and boundary |
| --- | --- |
| 1. Live/detached stream-run identity | `agent_runs._Run.run_id`, `get_run_id`, and the chat response's `X-Odysseus-Run-Id` identify the detached stream. The browser's `_streamRunIds` is populated from the response header. |
| 2. Stop/resume/replay identity | `expected_run_id` in `stop` and `request_finish`, route request headers, `_postExactStop`, the finish-editor request, `streamRunId`, and `resumeRunId` refer to that same detached stream. `subscribe` binds the exact `_Run` object returned by start/resume. |
| 3. Stream metrics/cost identity | `_metricsCostRecordId` uses the header-derived stream ID plus `primary`/`teacher`; `metrics._costRecordId` and the cost renderer's local `runId` refer to this accounting key. Neither uses terminal metadata's journal `run_id`. |
| 4. Logical nested invocation identity | `ActionJournal.run_id` is generated per completion-gated invocation. `action_id` is derived from it. The completion gate's terminal metadata `run_id` identifies this logical invocation. Existing `ToolRunSecurityContext.run_id` and `origin_run_id` values identify separate server-owned invocation/skill provenance operations; they are neither stream IDs nor journal lineage. |
| 5. ActionJournal parent/child identity | `ActionJournal.parent_run_id`, the gate's parent lookup, `_parent_run_id`, `request_teacher_takeover`'s captured parent ID, and `run_teacher_inline(parent_run_id=...)` link journal invocations. The completion metadata's `parent_run_id` preserves that lineage. |
No invocation ID is passed to stream stop/finish/replay APIs. No stream ID is
inserted into journal lineage. A new detached-stream regression creates nested
gates, rejects both journal IDs at stop/finish, accepts the stream ID for finish,
and verifies identical replay and unchanged journal metadata.
## Runtime invariants and final lab behavior
Every gated invocation creates a distinct journal, including children using
the same workspace. Journal and action bindings restore on normal unwind,
exception, cancellation, and generator close. Child awaiting/exhausted/error
state cannot rewrite the parent's completion decision or receipts.
The teacher adapter runs after the student gate closes. It forwards the parent
turn contract, tool policy, disabled tools, plan, client runtime context, and
external-untrusted-context restriction. Teacher execution receives a new
journal whose parent is the student invocation. Inner terminal frames are
consumed; only the outer adapter emits final termination. Exact framed
`data: [DONE]` events are distinguished from ordinary content containing the
literal marker.
Provider failures retain live events, then safe partial content when present,
then a non-completing decision, terminal metadata, and the original error last,
without DONE. A bare error remains a bare error. Completion gating does not
add provider calls or turn missing evidence into extra provider rounds.
Lab's server-owned authority remains narrower than inventory or availability.
Transcription, OCR, tasks, browser fallback, request-specific capability
selection, compact contracts, and provider-compatible tool choice retain the
canonical implementation. Model ID `Ajax` selects the Odysseus compact profile;
its selected schema boundary survives compatible `auto` tool choice, explicit
no-tools remains explicit, and transport remains OpenAI-compatible. No
benchmark-runner code was independently edited or executed.
## Validation records
The current requirements were installed in an isolated environment under this
worktree's ignored `.cache/wave1-1-reconciliation` directory. The shell's
unrelated `python` environment was not used for the accepted validation.
Canonical full pytest uses the repository's default data directory and allows
dotenv loading so research-path and setup tests can exercise their own fixtures;
the focused Wave script retains its explicit runtime isolation settings.
Optional live Ajax tests retain their opt-in skips; no live model or benchmark
run is part of this reconciliation.
- [Focused tests](validation/wave-1-1-reconciliation-focused.txt)
- [Wave 1.1 validation script](validation/wave-1-1-reconciliation-wave-validation.txt)
- [Broad affected runtime suite](validation/wave-1-1-reconciliation-broad.txt)
- [Canonical full pytest](validation/wave-1-1-reconciliation-pytest.txt)
- [Compileall, JS/MJS syntax, diff checks, and conflict-marker scan](validation/wave-1-1-reconciliation-gates.txt)
The focused records include the final relevant rerun after the reconciliation
audit was written. Full pytest and canonical static gates run afterward. The
local merge is committed only after the required checks pass. No push, PR,
deployment, or later-wave work is authorized by this reconciliation.
## Maestrum limitations encountered
The normal read-only pre-merge comparison stalled without a completion or
failure payload; its execution cell was terminated and the investigation was
not retried. Exact-path inspection proceeded using `local_only` with
`scope_mode="worktree"`.
The Context Firewall rejected an unbounded `git diff --cached --check` command
and withheld raw log output after the inspection allowance was exhausted.
Requests for ignored `.log` files were rejected with
`scope_rejected: ignored_by_git`. Unignored `.txt` validation records were
subsequently admitted by exact path. Canonical checks themselves run as
validation operations and record their exit status in the admitted gate log.
No epoch waiting or alternative worker mechanism was used.
@@ -0,0 +1,7 @@
Python compileall: 1689 tracked files; 0 failures
JS syntax: 279 tracked files; 0 failures
MJS syntax: 82 tracked files; 0 failures
git diff --check: exit 0
git diff --cached --check: exit 0
git diff HEAD --check: exit 0
Conflict-marker scan: 2377 tracked files; 0 matches