Require caller-supplied model endpoints to resolve through enabled owner-visible registrations, harden session path encoding, and remove the SVG title HTML parsing sink.
Regenerate website/configuration-reference.md with
scripts/generate_env_reference.py. The negative-web correction in
96e82562 inserted five lines in src/agent_loop.py ahead of the
ODYSSEUS_MAX_VISUAL_EVIDENCE_IMAGES and _FRAMES reads, so their recorded
locations move from 15329/15361 to 15334/15366. No variable, default or
description changed.
Document opt-in local workers and the per-process runtime ownership model.
Record the two parallel-only failures found and fixed in test
infrastructure. Record the measured results on the final code: serial
oracle 496.3s, -n 2 267.6s (1.85x), and two green -n 4 runs at 155.6s mean
(3.19x), all with identical skip and xfail sets and no leaked processes,
listeners, state, or runtime roots. Recommend -n 4 locally and explain
why -n auto was not run. The full serial run remains the release oracle.
The shared static server handled one connection at a time. Chromium can
open a speculative connection and never send a request, so every queued
request waited behind it. Under parallel load a computed-style capture's
navigation stalled for 30s and failed. Under CPU saturation, 4 of 12
captures stalled for about 29s each.
Serve each connection on a daemon thread. The existing serve-this-worktree
test now holds a silent connection open while it fetches, and times out
against the serial server. The configuration reference's recorded source
lines are unchanged.
Moving TMPDIR into the private runtime root left pytest's default
<TMPDIR>/pytest-of-<user>/pytest-<n> beneath it. With xdist's popen-gw<n>
the real-tmux witness bound a 110-byte socket path, over Linux's 107-byte
sun_path limit, so it failed under every worker count while passing
serially.
The controller now roots basetemp at the private root's pytest directory;
xdist hands workers popen-gw<n> beneath it. An explicit --basetemp wins.
A tmux-independent witness binds a socket at the same path budget.
Integrates the merged and frozen Wave 3 lab commit
b1666951faf8285054e1ca90f11533b0fb53fb57 with a normal merge, preserving
every Wave 4 commit unchanged.
Conflict: src/agent_runtime/resources.py. Wave 3's _control_plane_snapshot()
/ _control_plane_path(path, *, snapshot=None) split is kept. The snapshot adds
the effect-store directories to its prefix set after the recursive job-dir
inventory and no longer references path (the auto-merged prefix check would
have raised NameError there). _control_plane_path calls _aliases_effect_store
after its os.stat, only for multiply linked files, so single-link files never
list the store.
Semantic reconciliation (no textual conflict): bg_monitor keeps launch
settlement right after the first successful validate_job and before the
authority check, with Wave 3's post-drain revalidation intact. The
deleted-session branch, terminal before linkage validation, now settles a
validated launch too: that job is later pruned and its publication retired,
which would otherwise leave its effect RUNNING. Regression tests cover the
snapshot form of the effect-store check and both deleted-session linkage
outcomes.
Real-seam coverage for each corrective fix, each checked by mutation:
requested edit/patch states (CRLF-exact, unrelated change contradicts,
partial read and underivable targets stay unverified, superseded effects are
history); directory and launch-index fsync order observed via real fsync
targets; dispatch refused when the directory fsync fails; independent
objects, threads and processes never reuse positions; settle-once and
recovery against another writer; torn-tail repair; unbound tools cannot
manufacture RUNNING/cleanup/facts or settle launches; external effects never
complete as satisfied, are always disclosed, and passing tests stay test
facts; verifier staleness and RUNNING launches without obligations;
known-scope child effects leave unrelated parent evidence fresh; browser page
refusal survives a matching approval and child authority with no claim, no
execution id and no producer call; effect-store hardlinks are caught without
scanning the store.
Replaces the uncommitted tests that asserted a CONTENT_CHANGED predicate and
blocking on any RUNNING effect.
A hardlink into the effect store was protected only by the log's own nlink
refusal, which an agent could undo by removing the alias after writing
through it. Adding the store to the recursive control-plane inventory would
make every path check cost grow with accumulated runs. _aliases_effect_store
instead uses the store's invariants: logs and launch indexes refuse
st_nlink != 1 and the store is flat, so only a multiply linked regular file
on the store's device is compared by inode against one non-recursive listing.
Single-link files cost nothing and the store is never rglob-inventoried. The
helper takes a stat result so it plugs into Wave 3's scan-local snapshot after
the rebase.
Effect obligations were consulted only for declared artifacts, and reported
external success could be presented as done. Now, regardless of declared
artifacts:
- the latest effect on any changed file contradicted by a fresh readback
fails the run (a superseded earlier effect is history, not a contradiction);
- a passing verifier followed by an effect that may have changed state
without settled evidence is stale (BLOCKED);
- executed external effects that are not VERIFIED cap the decision at
UNVERIFIED, and the answer always carries server-authored facts for them
("reported success; any external change it made was not independently
verified", "reported failure", "unknown outcome").
The disclosure is structural and does not depend on recognizing the model's
wording. When it is the only change, the model's answer events are released
unchanged and the disclosure follows as one delta (and in round_texts).
Prose filtering is also tightened (remote verbs are mutation claims, an
unnamed "I updated it" cannot borrow the single required artifact, bare
"Done." is a terminal claim beside unverified external effects). A passing
verifier still supports test claims; it never speaks for the external effect.
Replaces the uncommitted attempt that blocked every run with any RUNNING
effect: a background launch with no declared obligations completes
UNVERIFIED.
Result-dictionary keys could set lifecycle state for any producer: a dynamic
or registry tool returning bg_job_id/detached became RUNNING, teardown became
verified cleanup, and timed_out/failure_kind/mutation_attempted/containment
were copied from untrusted results. Facts are now scoped to the producer the
dispatcher actually bound. An unbound tool contributes its exit status alone.
RUNNING requires a bound process producer (and an exact launch reservation
for bg_job_id), cleanup is attested only by a bound process producer, and job
observations and launch settlement only by a bound manage_bg_jobs operation
on exactly one Wave 3-validated job. External/remote-acknowledged facts come
from the captured ExternalResource, not from the result.
Sequence positions were allocated from each EffectLog object's in-memory
counter, so two objects, threads or processes could reuse a position or
settle one effect twice; replay then failed closed for the whole log. Every
append now takes an exclusive flock, merges the durable records other writers
appended (truncating a torn tail a crashed writer left), allocates from that
merged tail, rejects records the merged history makes invalid (a second
settlement, recovery of a claim another writer settled or marked running),
then appends, fsyncs and releases. history() merges others' records under a
shared lock. An incremental consistency index keeps appends O(1).
The first append of each log object fsyncs the log's directory, and every
directory created for it is fsynced in its parent, all under the lock before
the claim returns. A failed write or directory fsync truncates the record
back, so dispatch is refused and nothing unacknowledged is later merged. The
launch index writes and fsyncs a temp file, replaces it, then fsyncs the
directory. flock and directory fsync are POSIX-only and not claimed
elsewhere.
edit_file and apply_patch update claims asserted only existence (or, in the
uncommitted corrective attempt, any content change), so an unrelated write
could verify them. Each filesystem postcondition is now the exact content the
producer's own transformation writes from the identity-checked pre-state:
edit_file through the extracted pure _edit_file_text (no newline
translation), apply_patch updates through _apply_patch_hunks on the
universal-newline pre-state. An oversized, replaced or undecodable pre-state,
a non-matching hunk, or an underivable write_file body leaves the whole claim
without postconditions (UNVERIFIED) instead of letting derivable targets
verify the operation or falling back to existence.
The dispatcher imports src.agent_tools at call time. Binding TOOL_HANDLERS at
test-module import left patches on a stale dict after another test reloaded
the module, so two tests failed only in full-suite order.
Records the decision to recreate rather than cherry-pick 9012e208, the claim,
outcome, observation, invalidation and verification model, the durable log and
replay design, per-family adapters, completion-gate integration, browser and
background preservation, and residual P2 limitations.
Only the native detached launch (bg_job_id) or a bridge's explicit detachment
marks an operation's own work as RUNNING. A listing that reports some other
download/model/job as running settled normally; treating it as running left
the claim pending forever and could block required artifacts.
EffectHistory now indexes outcomes per effect once, removing a cubic scan in
assessment over long run lineages.
Adds adversarial coverage: browser page operations stay fail-closed through
the real dispatcher with effects enabled (no claim, never dispatched),
scheduler task triggers stay unverified admission, and assessment scales.
The background monitor already validates the exact Wave 3 job linkage
(job_from_record + validate_job) before continuing a session. At that point it
now records the job's settlement against the durable launch claim through the
launch-generation index, using typed lifecycle facts from the server-owned
record. Settlement is idempotent across deferred retries, is execution
evidence only, and never reads the delivered output: the injected report stays
attributed content. Failure to record leaves the claim running/unknown and
never blocks the follow-up.
Dispatcher seam: mark_dispatch, which runs inside the live Wave 3 binding
scope immediately before backend invocation, now durably claims a possible
effect before execution_id is assigned. If the claim cannot be persisted the
action stays undispatched and the dispatcher returns BLOCKED; dispatched()
closes the never-awaited coroutine. record_action appends the outcome
(including cancellation/interruption) and admitted-read observations before
the receipt reduction drops producer facts.
Adapters consume only the bound operations the dispatcher admitted:
filesystem bindings give exact scope and predicates (write_file content digest
after fence unwrapping, apply_patch add/delete, edit existence); bash/python
launches have unknown scope with the launch generation as lineage; job kills
scope the exact job and its processes; owned operations scope their exact
revisioned records; external backends are claimed as external and never
verified by acknowledgement; browser session_info yields session lifecycle
observations only, and a page binding is never effect scope. Complete
read_file re-reads the exact bound source to digest it; offset/limit,
truncation, extraction and listings are partial. Background launches stay
RUNNING until an admitted read of the exact job generation (via a durable
launch index, across continuation runs) reports settlement.
Producer seams: typed job lifecycle facts on manage_bg_jobs reads/kills, a
structured timed_out flag on containment timeouts, and mutation_attempted on
in-place write_file/edit_file failures after truncation.
Completion: the existing EvidenceLedger consumes effect assessments through a
single helper used for the decision, ask_user and prose filtering. A required
artifact is unsettled by a later unresolved effect that may have touched it,
a fresh contradicting readback fails the decision, and partial reads no longer
count as artifact validation. Ordinary conversation and read-only turns are
unchanged; no second completion policy is introduced.
Claims are fsynced to a per-lineage JSONL log before a caller may invoke a
backend; a persistence failure raises EffectPersistenceError (a
ResourceIdentityError) so dispatch fails closed. Outcomes and observations are
appended; nothing is rewritten. Reload validates every record strictly, ignores
only a torn final write, and fails closed on corruption, forgery or hardlink
aliasing. recover_interrupted appends INTERRUPTED/possible-impact outcomes for
claims that never settled and leaves RUNNING background effects alone.
The effect store is added to Wave 3 control-plane paths (prefix check only;
the log itself refuses aliased files), so filesystem tools cannot forge it.
Tests redirect the store to a session tmp directory.
Recreate (rather than cherry-pick 9012e208) the effect/provenance foundation.
The historical types used opaque string resource keys, a may_have_changed flag
defaulting to no impact, and a single status mixing execution and verification.
Claims, outcomes and observations now reference only typed Wave 3 identities
(filesystem root/inode/ancestor chain, process PID+start token, job generation,
owned revision, external incarnation, browser session incarnation). Browser
page resources are refused. Known no-op is limited to refusal before
invocation; unknown scope stays conservative; verification is derived from
fresh, complete, post-settlement readback checked against an explicit
predicate, and unknown execution with matching state is reported as observed
state without causation. Cleanup is recorded separately from effect outcome.