The #6503 merge added _control_plane_path() to _is_denied_tool_path(),
which grep, glob and ls call per enumerated entry. Without a snapshot
argument every call rebuilt _control_plane_snapshot() and stat'ed every
protected state file again. A 5,000-file grep went from 0.75 s to ~7 s,
and around 20k files it hits the grep timeout.
Thread an optional snapshot through _is_denied_tool_path() and
_can_traverse_tool_path(), and build it once per scan, the way
process_resources already does for its launch-boundary walk. Root
resolution and bound-resource resolution still observe fresh state.
The compact preview reduces every pre-execution tool error to a fixed
allowlist so exception text cannot reach the client. Messages our own
validators build per call are not on that list, so the model got "The
tool call could not be validated..." instead of the pdf_extract path
hint, the email uid provenance error, the artifact-Python target
message, or which schema argument was wrong. That is the default path
for the tool-profile models, Qwen3.5-9B included.
Each validator site now records its message in a per-call
validator_error before raising, and only that text is returned
verbatim. A schema failure gets a hint rebuilt from the offered schema
and the call's own arguments (missing property, expected type, allowed
values), never from the caught ValidationError. Everything else,
including all execution failures, still goes through
_public_preview_tool_error.
test_allows_user_content_the_app_hands_to_the_model is parametrized on
UPLOAD_DIR and friends, which sit under the per-worker data dir. pytest
used those paths as test ids, so every xdist worker collected different
names and `pytest -n 4` (documented in tests/README.md) aborted at
collection. CI shards with --shard and never saw it.
website/configuration-reference.md pinned every ODYSSEUS_* read to a
`path:line`, and test_env_reference fails when the committed page differs from
the generator output. So any change that adds or removes a line above an env
read in any scanned file has to regenerate the page, even when no variable
changed. Since the page landed on 2026-09-30, 24 of the 34 commits touching it
changed nothing but line numbers, and all three open PRs that touch it today do
the same - which also makes them conflict with each other on that file.
The Read in column now names the file, and "+N more" counts other files rather
than other read sites. The page still goes stale, and the test still fails,
when a variable is added, removed, moves to another file, or changes default.
check_notes still reports path:line on stderr, where it is useful and never
committed.
Request authority admits whole tool families from the user's request, so a
request to read email also admits send_email, delete_email and bulk_email,
and agent processes inherit the host network. With the gate defaulting to
off, an instruction injected through an email or a fetched page reaches those
tools with no other check; dev refuses them today.
Default the gate on, keep ODYSSEUS_TOOL_APPROVAL_GATE=0 as the opt-out, and
pin the production default with a test that imports the module in a fresh
interpreter. Four routing tests written for the opt-out posture now set it
explicitly.
Use Map-backed cost ledgers so externally derived session and run identifiers never cross ordinary object prototype semantics. Preserve the existing JSON storage format and extend browser and isolated ledger regressions for replay, overflow, legacy data, and reserved keys.
Require caller-supplied model endpoints to resolve through enabled owner-visible registrations, harden session path encoding, and remove the SVG title HTML parsing sink.
Document opt-in local workers and the per-process runtime ownership model.
Record the two parallel-only failures found and fixed in test
infrastructure. Record the measured results on the final code: serial
oracle 496.3s, -n 2 267.6s (1.85x), and two green -n 4 runs at 155.6s mean
(3.19x), all with identical skip and xfail sets and no leaked processes,
listeners, state, or runtime roots. Recommend -n 4 locally and explain
why -n auto was not run. The full serial run remains the release oracle.
The shared static server handled one connection at a time. Chromium can
open a speculative connection and never send a request, so every queued
request waited behind it. Under parallel load a computed-style capture's
navigation stalled for 30s and failed. Under CPU saturation, 4 of 12
captures stalled for about 29s each.
Serve each connection on a daemon thread. The existing serve-this-worktree
test now holds a silent connection open while it fetches, and times out
against the serial server. The configuration reference's recorded source
lines are unchanged.
Moving TMPDIR into the private runtime root left pytest's default
<TMPDIR>/pytest-of-<user>/pytest-<n> beneath it. With xdist's popen-gw<n>
the real-tmux witness bound a 110-byte socket path, over Linux's 107-byte
sun_path limit, so it failed under every worker count while passing
serially.
The controller now roots basetemp at the private root's pytest directory;
xdist hands workers popen-gw<n> beneath it. An explicit --basetemp wins.
A tmux-independent witness binds a socket at the same path budget.
Integrates the merged and frozen Wave 3 lab commit
b1666951faf8285054e1ca90f11533b0fb53fb57 with a normal merge, preserving
every Wave 4 commit unchanged.
Conflict: src/agent_runtime/resources.py. Wave 3's _control_plane_snapshot()
/ _control_plane_path(path, *, snapshot=None) split is kept. The snapshot adds
the effect-store directories to its prefix set after the recursive job-dir
inventory and no longer references path (the auto-merged prefix check would
have raised NameError there). _control_plane_path calls _aliases_effect_store
after its os.stat, only for multiply linked files, so single-link files never
list the store.
Semantic reconciliation (no textual conflict): bg_monitor keeps launch
settlement right after the first successful validate_job and before the
authority check, with Wave 3's post-drain revalidation intact. The
deleted-session branch, terminal before linkage validation, now settles a
validated launch too: that job is later pruned and its publication retired,
which would otherwise leave its effect RUNNING. Regression tests cover the
snapshot form of the effect-store check and both deleted-session linkage
outcomes.
Real-seam coverage for each corrective fix, each checked by mutation:
requested edit/patch states (CRLF-exact, unrelated change contradicts,
partial read and underivable targets stay unverified, superseded effects are
history); directory and launch-index fsync order observed via real fsync
targets; dispatch refused when the directory fsync fails; independent
objects, threads and processes never reuse positions; settle-once and
recovery against another writer; torn-tail repair; unbound tools cannot
manufacture RUNNING/cleanup/facts or settle launches; external effects never
complete as satisfied, are always disclosed, and passing tests stay test
facts; verifier staleness and RUNNING launches without obligations;
known-scope child effects leave unrelated parent evidence fresh; browser page
refusal survives a matching approval and child authority with no claim, no
execution id and no producer call; effect-store hardlinks are caught without
scanning the store.
Replaces the uncommitted tests that asserted a CONTENT_CHANGED predicate and
blocking on any RUNNING effect.
Effect obligations were consulted only for declared artifacts, and reported
external success could be presented as done. Now, regardless of declared
artifacts:
- the latest effect on any changed file contradicted by a fresh readback
fails the run (a superseded earlier effect is history, not a contradiction);
- a passing verifier followed by an effect that may have changed state
without settled evidence is stale (BLOCKED);
- executed external effects that are not VERIFIED cap the decision at
UNVERIFIED, and the answer always carries server-authored facts for them
("reported success; any external change it made was not independently
verified", "reported failure", "unknown outcome").
The disclosure is structural and does not depend on recognizing the model's
wording. When it is the only change, the model's answer events are released
unchanged and the disclosure follows as one delta (and in round_texts).
Prose filtering is also tightened (remote verbs are mutation claims, an
unnamed "I updated it" cannot borrow the single required artifact, bare
"Done." is a terminal claim beside unverified external effects). A passing
verifier still supports test claims; it never speaks for the external effect.
Replaces the uncommitted attempt that blocked every run with any RUNNING
effect: a background launch with no declared obligations completes
UNVERIFIED.