website/configuration-reference.md pinned every ODYSSEUS_* read to a
`path:line`, and test_env_reference fails when the committed page differs from
the generator output. So any change that adds or removes a line above an env
read in any scanned file has to regenerate the page, even when no variable
changed. Since the page landed on 2026-09-30, 24 of the 34 commits touching it
changed nothing but line numbers, and all three open PRs that touch it today do
the same - which also makes them conflict with each other on that file.
The Read in column now names the file, and "+N more" counts other files rather
than other read sites. The page still goes stale, and the test still fails,
when a variable is added, removed, moves to another file, or changes default.
check_notes still reports path:line on stderr, where it is useful and never
committed.
Request authority admits whole tool families from the user's request, so a
request to read email also admits send_email, delete_email and bulk_email,
and agent processes inherit the host network. With the gate defaulting to
off, an instruction injected through an email or a fetched page reaches those
tools with no other check; dev refuses them today.
Default the gate on, keep ODYSSEUS_TOOL_APPROVAL_GATE=0 as the opt-out, and
pin the production default with a test that imports the module in a fresh
interpreter. Four routing tests written for the opt-out posture now set it
explicitly.
Use Map-backed cost ledgers so externally derived session and run identifiers never cross ordinary object prototype semantics. Preserve the existing JSON storage format and extend browser and isolated ledger regressions for replay, overflow, legacy data, and reserved keys.
Require caller-supplied model endpoints to resolve through enabled owner-visible registrations, harden session path encoding, and remove the SVG title HTML parsing sink.
Document opt-in local workers and the per-process runtime ownership model.
Record the two parallel-only failures found and fixed in test
infrastructure. Record the measured results on the final code: serial
oracle 496.3s, -n 2 267.6s (1.85x), and two green -n 4 runs at 155.6s mean
(3.19x), all with identical skip and xfail sets and no leaked processes,
listeners, state, or runtime roots. Recommend -n 4 locally and explain
why -n auto was not run. The full serial run remains the release oracle.
The shared static server handled one connection at a time. Chromium can
open a speculative connection and never send a request, so every queued
request waited behind it. Under parallel load a computed-style capture's
navigation stalled for 30s and failed. Under CPU saturation, 4 of 12
captures stalled for about 29s each.
Serve each connection on a daemon thread. The existing serve-this-worktree
test now holds a silent connection open while it fetches, and times out
against the serial server. The configuration reference's recorded source
lines are unchanged.
Moving TMPDIR into the private runtime root left pytest's default
<TMPDIR>/pytest-of-<user>/pytest-<n> beneath it. With xdist's popen-gw<n>
the real-tmux witness bound a 110-byte socket path, over Linux's 107-byte
sun_path limit, so it failed under every worker count while passing
serially.
The controller now roots basetemp at the private root's pytest directory;
xdist hands workers popen-gw<n> beneath it. An explicit --basetemp wins.
A tmux-independent witness binds a socket at the same path budget.
Integrates the merged and frozen Wave 3 lab commit
b1666951faf8285054e1ca90f11533b0fb53fb57 with a normal merge, preserving
every Wave 4 commit unchanged.
Conflict: src/agent_runtime/resources.py. Wave 3's _control_plane_snapshot()
/ _control_plane_path(path, *, snapshot=None) split is kept. The snapshot adds
the effect-store directories to its prefix set after the recursive job-dir
inventory and no longer references path (the auto-merged prefix check would
have raised NameError there). _control_plane_path calls _aliases_effect_store
after its os.stat, only for multiply linked files, so single-link files never
list the store.
Semantic reconciliation (no textual conflict): bg_monitor keeps launch
settlement right after the first successful validate_job and before the
authority check, with Wave 3's post-drain revalidation intact. The
deleted-session branch, terminal before linkage validation, now settles a
validated launch too: that job is later pruned and its publication retired,
which would otherwise leave its effect RUNNING. Regression tests cover the
snapshot form of the effect-store check and both deleted-session linkage
outcomes.
Real-seam coverage for each corrective fix, each checked by mutation:
requested edit/patch states (CRLF-exact, unrelated change contradicts,
partial read and underivable targets stay unverified, superseded effects are
history); directory and launch-index fsync order observed via real fsync
targets; dispatch refused when the directory fsync fails; independent
objects, threads and processes never reuse positions; settle-once and
recovery against another writer; torn-tail repair; unbound tools cannot
manufacture RUNNING/cleanup/facts or settle launches; external effects never
complete as satisfied, are always disclosed, and passing tests stay test
facts; verifier staleness and RUNNING launches without obligations;
known-scope child effects leave unrelated parent evidence fresh; browser page
refusal survives a matching approval and child authority with no claim, no
execution id and no producer call; effect-store hardlinks are caught without
scanning the store.
Replaces the uncommitted tests that asserted a CONTENT_CHANGED predicate and
blocking on any RUNNING effect.
Effect obligations were consulted only for declared artifacts, and reported
external success could be presented as done. Now, regardless of declared
artifacts:
- the latest effect on any changed file contradicted by a fresh readback
fails the run (a superseded earlier effect is history, not a contradiction);
- a passing verifier followed by an effect that may have changed state
without settled evidence is stale (BLOCKED);
- executed external effects that are not VERIFIED cap the decision at
UNVERIFIED, and the answer always carries server-authored facts for them
("reported success; any external change it made was not independently
verified", "reported failure", "unknown outcome").
The disclosure is structural and does not depend on recognizing the model's
wording. When it is the only change, the model's answer events are released
unchanged and the disclosure follows as one delta (and in round_texts).
Prose filtering is also tightened (remote verbs are mutation claims, an
unnamed "I updated it" cannot borrow the single required artifact, bare
"Done." is a terminal claim beside unverified external effects). A passing
verifier still supports test claims; it never speaks for the external effect.
Replaces the uncommitted attempt that blocked every run with any RUNNING
effect: a background launch with no declared obligations completes
UNVERIFIED.
The dispatcher imports src.agent_tools at call time. Binding TOOL_HANDLERS at
test-module import left patches on a stale dict after another test reloaded
the module, so two tests failed only in full-suite order.
Only the native detached launch (bg_job_id) or a bridge's explicit detachment
marks an operation's own work as RUNNING. A listing that reports some other
download/model/job as running settled normally; treating it as running left
the claim pending forever and could block required artifacts.
EffectHistory now indexes outcomes per effect once, removing a cubic scan in
assessment over long run lineages.
Adds adversarial coverage: browser page operations stay fail-closed through
the real dispatcher with effects enabled (no claim, never dispatched),
scheduler task triggers stay unverified admission, and assessment scales.
The background monitor already validates the exact Wave 3 job linkage
(job_from_record + validate_job) before continuing a session. At that point it
now records the job's settlement against the durable launch claim through the
launch-generation index, using typed lifecycle facts from the server-owned
record. Settlement is idempotent across deferred retries, is execution
evidence only, and never reads the delivered output: the injected report stays
attributed content. Failure to record leaves the claim running/unknown and
never blocks the follow-up.
Dispatcher seam: mark_dispatch, which runs inside the live Wave 3 binding
scope immediately before backend invocation, now durably claims a possible
effect before execution_id is assigned. If the claim cannot be persisted the
action stays undispatched and the dispatcher returns BLOCKED; dispatched()
closes the never-awaited coroutine. record_action appends the outcome
(including cancellation/interruption) and admitted-read observations before
the receipt reduction drops producer facts.
Adapters consume only the bound operations the dispatcher admitted:
filesystem bindings give exact scope and predicates (write_file content digest
after fence unwrapping, apply_patch add/delete, edit existence); bash/python
launches have unknown scope with the launch generation as lineage; job kills
scope the exact job and its processes; owned operations scope their exact
revisioned records; external backends are claimed as external and never
verified by acknowledgement; browser session_info yields session lifecycle
observations only, and a page binding is never effect scope. Complete
read_file re-reads the exact bound source to digest it; offset/limit,
truncation, extraction and listings are partial. Background launches stay
RUNNING until an admitted read of the exact job generation (via a durable
launch index, across continuation runs) reports settlement.
Producer seams: typed job lifecycle facts on manage_bg_jobs reads/kills, a
structured timed_out flag on containment timeouts, and mutation_attempted on
in-place write_file/edit_file failures after truncation.
Completion: the existing EvidenceLedger consumes effect assessments through a
single helper used for the decision, ask_user and prose filtering. A required
artifact is unsettled by a later unresolved effect that may have touched it,
a fresh contradicting readback fails the decision, and partial reads no longer
count as artifact validation. Ordinary conversation and read-only turns are
unchanged; no second completion policy is introduced.