"Is this path inside that root" is asked in twenty places in this tree and
answered twenty times by a locally written realpath/commonpath pair. Nine test
files exist because nine call sites each needed their own proof. Each one is
defensible alone; together they are the defect, because the boundary has no
single definition and a site that gets a detail wrong is wrong by itself.
src/path_confinement.py is that definition, and it settles the details the
copies disagreed on. Both sides get canonicalized: comparing a realpath-ed
candidate against a root that was only abspath-ed is the macOS /tmp ->
/private/tmp mismatch that has already produced a false failure here, and
canonicalizing one side is worse than canonicalizing neither. commonpath rather
than startswith, because /a/bc begins with /a/b and is not inside it. A relative
candidate joins the root rather than os.getcwd(), which is whatever directory
the server happens to be running in. NUL and newline are refused with a reason
instead of caught by a bare `except Exception` and reported as an ordinary
escape. Eighteen call sites go through it now. It deliberately does not decide
whether a path is sensitive -- that deny list answers "allowed" rather than
"inside", and it stays with src/tool_execution, which owns it. The one
commonpath left in the tree, in src/workspace_paths.py, stays: that function
translates a host path into a container path, so canonicalizing either side
would change the relative path it computes and break the mapping. It is not a
confinement check.
Two of those sites were weaker than the rest and are fixed rather than moved.
The email attachment check used abspath, which folds `..` but does not resolve
symlinks, so a symlink written into the extraction directory passed it and was
then read through. The skill-reference guard compared a realpath-ed target
against a raw dirname, so on a host where the skills tree is reached through a
symlink the two sides never matched and the guard could not fire.
The execution boundary had two separate holes.
The workspace namespace bound /home and /mnt read-write. On the one platform
where that namespace engages at all, a command inside it reaches outside the
workspace and writes to the user's home directory -- measured by running this
argv on a Linux host with working bubblewrap, not inferred from the source.
Binding the user's whole home directory into a workspace-confinement namespace
gives back most of what the namespace was for. Both are read-only now. The
workspace is also bound writable at its real host path, not only at /workspace:
BashTool's own /tmp redirect rewrites `/tmp/` to `<agent_cwd()>/.tmp/` before
the namespace is built, so the command bwrap receives already names the real
path, and those writes previously landed only because the workspace happened to
sit under the writable /home.
`namespaced or _replace_workspace_alias(...)` chose between a mount namespace
and a regex with nothing in the result saying which one ran. The fallback
rewrites the literal token /workspace in the command string, so a command that
never mentions /workspace is untouched by it and runs on the host unrestricted
-- which is every agent shell command on macOS. Both tools now ask
containment.probe() instead of each deciding for itself, and every bash and
python result carries a containment block naming the mechanism and stating
whether the filesystem dimension actually held. Under enforcing mode the
command is not run and the result says so.
That block reports the filesystem dimension only, and says so in a
reported_dimensions field. The probe knows this host could also give a process
group and a real wall clock, but these two tools still assemble their own
create_subprocess_* call and pass neither, so listing those dimensions would be
exactly the false claim src/containment.py calls worse than an honest absence.
probe() is new on src/containment.py: the same mechanism table and the same
arithmetic as acquire(), stopping before the side effects. acquire() is the
wrong shape for a decision -- it writes a durable grant record, and a record
whose pid is never filled in and whose release() never runs is an entry a
restart reaper keeps finding.
CONTAINMENT_MODE stays report_only. Flipping it refuses every agent shell
command on macOS and on any Linux host without bubblewrap, which is a product
decision rather than a code one.
Smaller things in the same area: the /tmp redirect's makedirs was unguarded, so
a read-only workspace turned a command that merely mentioned `/tmp/` into an
OSError traceback instead of a tool error; it degrades now. WORKSPACE_MOUNT
moved to src/constants.py so the namespace and the path resolvers read one
definition of the contract rather than two. The ".tmp" dirname got a constant,
since it appeared in both tool paths.
One generated artifact moved with it: website/configuration-reference.md pins
the source line where each ODYSSEUS_* variable is read, and three of those
shifted. Regenerated with scripts/generate_env_reference.py; the diff is line
numbers only.
Three existing tests changed. test_workspace_artifact_tool_floor asserted that
an unsafe interpreter prefix produces no `--ro-bind <prefix> <prefix>`, which
now fires on /home because /home is legitimately a read-only base mount.
Asserting the absence of a literal flag string cannot distinguish "the prefix
was rejected" from "the argv mounted that root itself", so it compares the argv
against the no-prefix baseline instead: an unsafe prefix must add nothing.
The Windows bash test asserted dict equality on the
whole result, which makes adding a field to every bash result impossible without
touching a test about tmux; it asserts the shape now. The personal-dir symlink
test grepped the resolver's source for the literal "os.path.realpath", which is
gone because the resolution moved into the shared boundary -- it keeps the
negative assertion that the closure must not grow its own abspath check again,
and the behavioural half now runs against the boundary, where it covers every
call site instead of one closure.
Not verified: the bubblewrap argv is asserted, not executed. There is no bwrap
on macOS, and in Docker it needs --privileged to work at all -- default and
seccomp=unconfined both fail with "Creating new namespace failed", and
--cap-add=SYS_ADMIN fails at pivot_root. The Python tool's
needs_virtual_namespace gate means ordinary Python code gets no namespace even
on a Linux host that could provide one; that is reported now but deliberately
not changed, because it alters the Linux Python path on every call and cannot be
checked from here.
A recorded pid is a claim, not a handle. The containment grant store, the
background-job store and the Cookbook task list all outlive the process that
wrote them — deliberately, so a restart keeps a job and its result — and the
kernel reuses pids. Any teardown driven off one of those records can therefore
land on a process we never started. ODY-86 was exactly this, and the Cookbook
survivor sweep still terminated any process whose full command line matched a
tracked one, which is the same mistake spelled differently.
Identity is (pid, start token). The token comes from /proc/<pid>/stat on Linux,
ps -o lstart= on macOS and the BSDs, and GetProcessTimes on Windows; the kernel
will not hand a pid to a process that started earlier, so comparing the token
recorded at launch against the token read now answers "is this still ours"
without a handle or a supervisor. verify() returns owned, gone, foreign or
unverifiable, and only owned permits a signal.
Keeping "unverifiable" out of the other two is the point. Process inspection
has broken off Linux four times here — ODY-70, -86, -94, -99 — every time
because an absent mechanism read as a successful answer. Folding it into "ours"
signals strangers; folding it into "gone" abandons live processes. It is a
containment failure and every caller treats it as one.
Wired into the three places that signal:
- containment.release() gates a grant recovered from the durable store, and
leaves an in-process teardown alone, where the caller holds the child and no
identity question arises. The verdict lands on the record, so "why is this
grant still here" is answerable afterwards.
- A startup reaper. Nothing read either store before, so a crashed run left
every grant permanently active and every job permanently running, and the
first thing to touch such a record was a teardown aimed at a reassigned pid.
The two stores get opposite treatment: an orphaned grant has no caller left
and is torn down, while a detached job is documented to survive a restart and
is only corrected, never killed.
- The Cookbook sweep takes its ownership from the tmux pane's process tree,
captured before the kill destroys the only link between a surviving model
server and the session that started it. A process that merely matches the
tracked command line is now reported rather than killed: the Cookbook
composed that command line, so an identical one is just as likely to be a
server the user started by hand. The sweep also runs on hosts with no procfs
instead of silently skipping, and says so when it could not look at all.
_session_alive collapsed every OSError from killpg(pgid, 0) into "the
group is gone". EPERM means the opposite — the group answered the probe
but holds a process we may not signal — so a session we could not touch
was reported as contained, and a timed-out command that left children
running said it had terminated cleanly.
Resolving PTY_KILL_ESCALATION also named signal.SIGKILL unconditionally,
which does not exist on native Windows. app.py imports this module at
start-up, so that turned a POSIX-only teardown detail into the whole app
failing to import there.
/api/shell/stream starts its PTY child under os.setsid, so the child
leads its own session and process group. The timeout, client-disconnect
and error paths all called proc.kill(), which signals only the group
leader. Creating a group and then signalling only its leader is strictly
worse than never creating one: the descendants are detached from the
server's group as well, so nothing else will ever reach them, while the
route reports "Command timed out after Ns" and exit_code -1 as if the
command were gone.
The kernel's controlling-terminal SIGHUP hid this for well-behaved
children, which is why it reads as working. Anything that ignores
SIGHUP — a nohup'ed job, a daemon, a process that means to outlive its
terminal — survives the kill indefinitely.
Signal the whole group instead, escalate to SIGKILL if it outlives the
grace period, and confirm it is actually gone. The timeout response now
says so when containment could not be established rather than claiming
a clean kill it did not get.
Agent-reachable execution has 25 independent spawn sites and no single
place deciding where a process runs or under what limits. All three
consequences are visible on this SHA. When bwrap is absent the workspace
namespace degrades to a regex that rewrites /workspace to the real path,
and nothing in the tool result says which one you got. No spawn site
passes start_new_session, so a wall-clock kill reaches the wrapper shell
and leaves its backgrounded grandchildren running while reporting the
process killed. Teardown stops at SIGTERM without ever checking death.
src/containment.py gives those paths one boundary. acquire() establishes
containment or refuses -- a string rewrite is not a mechanism it can
select -- and the grant states which dimensions actually hold, which were
best-effort and are missing, and which were required and are missing.
run() enforces the wall clock and the output cap. release() signals the
process group, escalates to SIGKILL, and reports dead only for a group it
observed go empty.
Containment never sees the command: acquire() takes a workspace and
limits, and the command text only reaches run(). Nothing in a request can
widen a boundary it is never shown.
CONTAINMENT_MODE chooses between refusing an unestablishable required
dimension and recording it. It ships report-only, so landing this changes
no behaviour on a host without bwrap -- which is every host today.
No call sites move here; they follow on this branch. The configuration
reference is regenerated because the page records src/constants.py line
numbers and the new path constant shifts two of them.
The Windows branch of `_create_bash_subprocess` spawned Git Bash with
neither pipes nor the env it was handed. `proc.stdout` and `proc.stderr`
came back `None`, so `_run_subprocess_streaming`'s reader returned
immediately and the Bash tool reported `"(no output)"` alongside the real
exit code — while the child inherited the server's own stdout/stderr and
wrote agent command output into the console and the launchd/Docker logs.
The `env` parameter was accepted and never used, so `PATH`, `VIRTUAL_ENV`,
`HOME`, `TMPDIR` and the configured import paths carried in
`ctx["subproc_env"]` never reached the child on Windows, even though every
POSIX path applies them.
Spawn it the way the POSIX path at `:688` already does: `stdin=DEVNULL`,
`stdout=PIPE`, `stderr=PIPE`, `env=env`.
`website/configuration-reference.md` is generated from source line numbers,
so the four added lines shift one entry; regenerated with
`scripts/generate_env_reference.py`.
- Headless consumers (task scheduler, background follow-up) now treat a
completion-gate final_response as the authoritative answer instead of
collecting deltas only. A gated replacement no longer leaves scheduled
output empty, which used to trigger an extra, ungated grace-summary
model call.
- The scheduler closes the agent stream with contextlib.aclosing, so the
approval-pause break unwinds the gate's journal and teacher-takeover
context in its own task. Chained runs no longer inherit a stale
parent_run_id, and later finalization no longer raises ContextVar
reset errors.
- On provider error, the completion gate applies the live answer's
statement filter to persisted round_texts. Diagnostics and the failure
note survive; claims rejected by the gate cannot reappear on reload.
Snapshot current maintainer-preview application changes and regression fixtures for integration into lab. Excludes local runtime data, evaluation outputs and source backups. Focused Python regression selection: 140 passed; full suite not certified.
Nothing in the UI said which build was loaded. /api/version has reported
version, build and source_commit since the harness started versioning itself
apart from the public semver, but the only way to read it was to curl the
endpoint — so "is the preview actually running the commit I just merged?" took
a terminal to answer.
Pins a footer under the settings sidebar nav showing the registered version
(plus the harness build when it differs) and the short source commit, with the
full hash on hover. It sits outside the nav's scroll container so it stays at
the bottom-left, and it is .admin-only, so syncAdminVisibility() hides it from
non-admins the same way it hides the Admin nav group.
The commit resolves at import via `git rev-parse HEAD` and is the string
"unknown" when that fails — a read-only Docker tree with no .git. The footer
treats "unknown" as absent and stays hidden when nothing is left to show,
rather than printing it. The collapsed rail and the two narrow tab-rail
layouts hide it too: neither has a bottom-left to write in.
Completes the lane alteixeira20 asked for. Test-only: tests/ and test helpers,
no production runtime code, no benchmark fixtures or allowlists.
Supplied workspace context must not produce a clarification. Pins
_looks_like_unattended_clarification on four shapes that hand the decision
back ("could you please share", "shall I", "which approach do you prefer")
and three ordinary answers that must not trip it.
Repeated update_plan is not the turn's work. ask_user and update_plan are
permitted on nearly every turn, so if they counted as execution a model could
loop on them and look busy. Pins that _tool_rejection_reason does not
advertise either as an available tool, and that update_plan is permitted
without ever being in required.
Request-scoped tool authority. _request_scoped_allowed_tool_names must not
make an undeclared tool executable; the native-terminal widening is pinned
separately so it stays opt-in rather than drifting into the default.
Foreign-process safety. The Chrome sweep matches this runtime's own profile
prefix, so a fake procfs with our pid, a user's ordinary Chrome and another
worktree's agent browser must leave exactly two of the three alone.
14 passed, 2 xfailed. The xfails are the negative-wording cases from the first
commit that do not hold yet.
Four more tests in the same family as the /tmp ones this change already
fixes, and they hide for the same reason: the failure depends on how
long $TMPDIR happens to be.
tests/test_shell_routes.py::TestHostDockerAccess (three) and
tests/test_cookbook_docker_access.py::test_container_opt_in_with_unix_
socket_is_allowed each bind an AF_UNIX socket at tmp_path/"docker.sock".
macOS gives sun_path 104 bytes including the terminator. pytest roots
tmp_path at $TMPDIR, which on a stock Mac is a 49-character
/var/folders/<2>/<30>/T/; add pytest-of-<user>/pytest-<n>/ and the
test's own name and the bind path is 115 bytes before the filename.
OSError: AF_UNIX path too long
Linux allows 108 and roots $TMPDIR at /tmp, so CI never sees it. Under a
shortened $TMPDIR the path lands at exactly 103 and passes — until
pytest's run counter reaches two digits and it becomes 104. That is why
the ledger's counts did not include these: they were measured somewhere
the path fit.
Adds tests/helpers/unix_sockets.bound_unix_socket, which binds under a
short directory and asserts the length before it tries, so the next
socket fixture fails with a sentence rather than an errno. Records the
trap in KNOWN_FAILURES.md along with the instruction to re-measure with
the default $TMPDIR.
Verified end to end against a local Qwen3.5-9B Q4_K_M that when the user
explicitly enables web for the turn, none of the three phrasings withholds
web_search, web_fetch or private_browser, including the one these tests record
as held. The held case holds on the inferred path only, where no toggle is set
and the runtime decides from intent.
That distinction was missing and the file read as a stronger claim than the
measurement supports. An explicit toggle beating an inferred negative may be
the intended semantics, so it is recorded rather than asserted.
The route now answers 400 with a message naming the expected shape.
admin.js was taught to print `data.detail`; the Unified Integrations
form in settings.js still printed `Failed (400)` and dropped it.
That gap is exactly where the new validation bites. The client-side
JSON.parse guard added alongside it catches unparseable input, so the
only values that reach the route's 400 are ones that parse but are not
a list — `"npx"`, `{}`, `null` — and for those the status code alone
tells the user nothing about what is wrong with what they typed.
Adds source-level coverage for both forms; the PR changed two JS files
with no test on either.
Splitting style.css deleted it, and two files outside static/ still
named it:
- scripts/verify_background_research_cards.mjs injected
`<link rel="stylesheet" href="/static/style.css">` into the page it
builds. A stylesheet that 404s does not fail — the script kept
checking card layout against an unstyled page and kept reporting
pass. It now reads the shell's <link> tags out of index.html, the way
tests/css_snapshot/capture.mjs already does, so the set cannot drift
out from under it again.
- tests/css_snapshot/bench.html's hand-open fallback linked the same
deleted file. Replaced with the ordered set index.html ships.
test_every_stylesheet_referenced_by_shipped_html_exists only walked
static/*.html, which is why neither showed up. Extend it over the bench
page: the bench's whole output is computed styles, so a dead link there
is worth more than an unstyled page nobody looks at.
_process_is_alive used os.kill(pid, 0). That probe is POSIX-only:
CPython's Windows os.kill calls TerminateProcess(handle, sig) for any
signal other than CTRL_C/CTRL_BREAK, so it terminates the process it is
asked about. This function is only reached when there is no procfs to
read a command line from, which is exactly the macOS and Windows case
the rest of this change exists to handle.
core/platform_compat.py already owns that probe and documents the
hazard; its module docstring asks callers to import from there rather
than spell a POSIX-only call out locally. Delegate to it.
pid_alive answers False where os.kill raises PermissionError — a live
process owned by another user. Both call sites want that reading: the
sweep only unlinks a pid file it wrote itself, and a pid it cannot
confirm is not the daemon it is looking for.