Agent-reachable execution has 25 independent spawn sites and no single
place deciding where a process runs or under what limits. All three
consequences are visible on this SHA. When bwrap is absent the workspace
namespace degrades to a regex that rewrites /workspace to the real path,
and nothing in the tool result says which one you got. No spawn site
passes start_new_session, so a wall-clock kill reaches the wrapper shell
and leaves its backgrounded grandchildren running while reporting the
process killed. Teardown stops at SIGTERM without ever checking death.
src/containment.py gives those paths one boundary. acquire() establishes
containment or refuses -- a string rewrite is not a mechanism it can
select -- and the grant states which dimensions actually hold, which were
best-effort and are missing, and which were required and are missing.
run() enforces the wall clock and the output cap. release() signals the
process group, escalates to SIGKILL, and reports dead only for a group it
observed go empty.
Containment never sees the command: acquire() takes a workspace and
limits, and the command text only reaches run(). Nothing in a request can
widen a boundary it is never shown.
CONTAINMENT_MODE chooses between refusing an unestablishable required
dimension and recording it. It ships report-only, so landing this changes
no behaviour on a host without bwrap -- which is every host today.
No call sites move here; they follow on this branch. The configuration
reference is regenerated because the page records src/constants.py line
numbers and the new path constant shifts two of them.
- Headless consumers (task scheduler, background follow-up) now treat a
completion-gate final_response as the authoritative answer instead of
collecting deltas only. A gated replacement no longer leaves scheduled
output empty, which used to trigger an extra, ungated grace-summary
model call.
- The scheduler closes the agent stream with contextlib.aclosing, so the
approval-pause break unwinds the gate's journal and teacher-takeover
context in its own task. Chained runs no longer inherit a stale
parent_run_id, and later finalization no longer raises ContextVar
reset errors.
- On provider error, the completion gate applies the live answer's
statement filter to persisted round_texts. Diagnostics and the failure
note survive; claims rejected by the gate cannot reappear on reload.
Snapshot current maintainer-preview application changes and regression fixtures for integration into lab. Excludes local runtime data, evaluation outputs and source backups. Focused Python regression selection: 140 passed; full suite not certified.
_process_is_alive used os.kill(pid, 0). That probe is POSIX-only:
CPython's Windows os.kill calls TerminateProcess(handle, sig) for any
signal other than CTRL_C/CTRL_BREAK, so it terminates the process it is
asked about. This function is only reached when there is no procfs to
read a command line from, which is exactly the macOS and Windows case
the rest of this change exists to handle.
core/platform_compat.py already owns that probe and documents the
hazard; its module docstring asks callers to import from there rather
than spell a POSIX-only call out locally. Delegate to it.
pid_alive answers False where os.kill raises PermissionError — a live
process owned by another user. Both call sites want that reading: the
sweep only unlinks a pid file it wrote itself, and a pid it cannot
confirm is not the daemon it is looking for.
`_cached` deduplicates the scheduler's outbound fetches — Miniflux unread
counts and MCP tool snapshots — by parking every concurrent caller on one
shared Future. Two cancellation paths left that Future stranded.
A waiter awaited the shared Future directly, so cancelling the waiter
cancelled the Future the owner and every other waiter were using. It now
awaits through `asyncio.shield`.
The owner removed its pending entry inside the success and `except Exception`
branches. `CancelledError` is a `BaseException`, so it took neither: the key
stayed in `_shared_cache_pending` pointing at a Future nobody would ever
resolve, and every later caller for that key waited forever. Cleanup moves to
a `finally` that is synchronous on purpose, and the owner cancels its own
Future so current waiters wake while a later caller can still retry.
Ported from public `dev` (`ce04dc1d`, #6174 upstream), with its test.
The host cache was gated on the list being non-empty, so "queried fine, no
eligible peers" looked exactly like a cold cache and every caller paid for
another `tailscale status --json` — a subprocess with a 5s timeout.
Gate on the timestamp instead. Failures still leave the timestamp unset, so a
missing binary, a non-zero exit or unparseable output stays retryable rather
than being cached for the full TTL.
Ported from public `dev` (`affaee1e`, #6228 upstream), with its test.
`_cookbook_kill_session` kills the tmux session, then sweeps /proc for
model servers that survived the SIGHUP. The sweep had no guard, so on
macOS and Windows `os.listdir("/proc")` raised FileNotFoundError after
the kill had already succeeded. The function's outer except turned that
into `{"error": "...No such file or directory: '/proc'", "exit_code": 1}`
and skipped the state write that marks the session stopped — the agent
is told a stop failed that actually worked.
Guard the sweep with a procfs check, the way `_scan_running_model_processes`
already does a few hundred lines up. The root and the check now live in
`core/platform_compat`, which is where OS differences belong and which
makes both branches patchable from a test on either kind of host.
Adds a class-level guard test: the third instance of this defect, and two
of the three were found by reading source rather than by a test.
_terminate_owned_daemon() read /proc/<pid>/cmdline and, on FileNotFoundError,
unlinked the pid file on the stated assumption that "the daemon may have
exited". Off Linux that file is always missing, so the branch always fired:
the pid file of a live daemon was deleted and the daemon itself never killed.
_owned_daemon_exists() swallowed the same error and therefore always returned
False, which is precisely the state its own docstring warns about, since a
close against an unrecognised session can bootstrap a fresh daemon and wait on
its browser indefinitely.
Demonstrated on macOS before the change: a pid file holding a live pid is
removed by _terminate_owned_daemon() and _owned_daemon_exists() reports False.
After it, the file survives and the daemon is reported present.
_process_command_line() now returns None for "this host cannot tell" and
_process_is_alive() answers the separate question of whether the pid exists.
Without procfs we decline to kill a process we cannot confirm is ours, and we
only forget a pid file once the pid is genuinely gone. Linux behaviour is
unchanged: the command-line identity check still gates both paths.
_terminate_owned_chrome() walked Path("/proc") unconditionally, so on macOS
and Windows iterdir() raised FileNotFoundError out of private-browser session
shutdown. src/tools/cookbook.py already guards the same kind of scan with
os.path.isdir("/proc").
The sweep only reclaims Chrome trees that agent-browser reparented, so it is
an optimisation rather than a correctness requirement: degrade to a no-op
rather than failing the whole shutdown path.
_PROC_ROOT is a module attribute so both branches are testable on either kind
of host. The procfs-present path had no coverage at all before this.
The R09 scholarly hardening routes every hop through check_outbound_url with
block_private=True. On a DNS64/NAT64 network, an IPv4-only host can resolve
through the RFC 6052 Well-Known Prefix. For example, export.arxiv.org resolved
to 64:ff9b::924b:5b2a in the reproduced environment.
CPython classifies that outer IPv6 prefix as reserved, so the URL guard rejected
the request before examining the effective IPv4 destination.
Decode addresses in exactly 64:ff9b::/96 to their embedded IPv4 destination and
evaluate that destination under the strict outbound policy.
The translated target is always checked with private-address blocking enabled.
This prevents the NAT64 prefix from becoming a path to loopback, private,
shared/CGNAT, link-local, multicast, unspecified, or other non-global IPv4
space.
Network-specific translation prefixes are not decoded. In particular,
64:ff9b:1::/48 remains subject to the existing IPv6 policy.
Coverage includes:
- public IPv4 destinations embedded through the RFC 6052 prefix
- loopback, link-local, private, CGNAT, multicast, unspecified, benchmark, and
TEST-NET rejection
- network-specific NAT64 prefixes remaining undecoded
- deterministic scholarly provider tests without live DNS dependence
- redirect query parameter isolation
- relative redirect resolution
- HTTP error fallback behavior
- editor draft GET and DELETE behavior remaining unaffected
R09, R11, and R01 were independently audited and otherwise left unchanged.
- Enforce outbound URL policy on all hops via bounded manual redirects in scholarly lookups
- Guard declared Content-Length in EditorDraftRoute before request body parsing
- Keep scholarly lookup timeouts and budgets as internal constants rather than surface env vars
- Narrow TUI threat-model description around demonstrable host shell bridge behavior
- Add behavioral regression tests for redirect security and pre-parsing body size guards
The arXiv and OpenAlex title resolvers called two hardcoded endpoints with a
hand-written "Odysseus/0.20" User-Agent, no outbound-URL policy, and a full
12s timeout each. APP_VERSION was already 1.0.3, so the agent string was wrong
the moment it was written, and a scholarly query could hold a user-facing
search open for the sum of all three hops.
Move the endpoints to named constants, build the User-Agent from APP_VERSION,
run both calls through check_outbound_url (they follow redirects, so the final
host is not the one in the constant), and give the whole SearXNG -> OpenAlex ->
arXiv chain one shared wall-clock budget.
The budget is a ContextVar rather than a parameter so the existing test doubles
for _openalex_title_results and _arxiv_title_results keep working unchanged.