Completes the lane alteixeira20 asked for. Test-only: tests/ and test helpers,
no production runtime code, no benchmark fixtures or allowlists.
Supplied workspace context must not produce a clarification. Pins
_looks_like_unattended_clarification on four shapes that hand the decision
back ("could you please share", "shall I", "which approach do you prefer")
and three ordinary answers that must not trip it.
Repeated update_plan is not the turn's work. ask_user and update_plan are
permitted on nearly every turn, so if they counted as execution a model could
loop on them and look busy. Pins that _tool_rejection_reason does not
advertise either as an available tool, and that update_plan is permitted
without ever being in required.
Request-scoped tool authority. _request_scoped_allowed_tool_names must not
make an undeclared tool executable; the native-terminal widening is pinned
separately so it stays opt-in rather than drifting into the default.
Foreign-process safety. The Chrome sweep matches this runtime's own profile
prefix, so a fake procfs with our pid, a user's ordinary Chrome and another
worktree's agent browser must leave exactly two of the three alone.
14 passed, 2 xfailed. The xfails are the negative-wording cases from the first
commit that do not hold yet.
Four more tests in the same family as the /tmp ones this change already
fixes, and they hide for the same reason: the failure depends on how
long $TMPDIR happens to be.
tests/test_shell_routes.py::TestHostDockerAccess (three) and
tests/test_cookbook_docker_access.py::test_container_opt_in_with_unix_
socket_is_allowed each bind an AF_UNIX socket at tmp_path/"docker.sock".
macOS gives sun_path 104 bytes including the terminator. pytest roots
tmp_path at $TMPDIR, which on a stock Mac is a 49-character
/var/folders/<2>/<30>/T/; add pytest-of-<user>/pytest-<n>/ and the
test's own name and the bind path is 115 bytes before the filename.
OSError: AF_UNIX path too long
Linux allows 108 and roots $TMPDIR at /tmp, so CI never sees it. Under a
shortened $TMPDIR the path lands at exactly 103 and passes — until
pytest's run counter reaches two digits and it becomes 104. That is why
the ledger's counts did not include these: they were measured somewhere
the path fit.
Adds tests/helpers/unix_sockets.bound_unix_socket, which binds under a
short directory and asserts the length before it tries, so the next
socket fixture fails with a sentence rather than an errno. Records the
trap in KNOWN_FAILURES.md along with the instruction to re-measure with
the default $TMPDIR.
Verified end to end against a local Qwen3.5-9B Q4_K_M that when the user
explicitly enables web for the turn, none of the three phrasings withholds
web_search, web_fetch or private_browser, including the one these tests record
as held. The held case holds on the inferred path only, where no toggle is set
and the runtime decides from intent.
That distinction was missing and the file read as a stronger claim than the
measurement supports. An explicit toggle beating an inferred negative may be
the intended semantics, so it is recorded rather than asserted.
The move left three documents naming `routes/email_routes.py`,
`routes/email_helpers.py` and `routes/email_pollers.py` as where the
code is. The shims keep those import paths working, so nothing breaks —
but each of those files is now seventeen lines that redirect, and a
reader sent there finds no email code at all.
Follows the phrasing specs/persistence.md already uses for the
contacts and vault subpackages: name the canonical path and note the
shims.
The route now answers 400 with a message naming the expected shape.
admin.js was taught to print `data.detail`; the Unified Integrations
form in settings.js still printed `Failed (400)` and dropped it.
That gap is exactly where the new validation bites. The client-side
JSON.parse guard added alongside it catches unparseable input, so the
only values that reach the route's 400 are ones that parse but are not
a list — `"npx"`, `{}`, `null` — and for those the status code alone
tells the user nothing about what is wrong with what they typed.
Adds source-level coverage for both forms; the PR changed two JS files
with no test on either.
Splitting style.css deleted it, and two files outside static/ still
named it:
- scripts/verify_background_research_cards.mjs injected
`<link rel="stylesheet" href="/static/style.css">` into the page it
builds. A stylesheet that 404s does not fail — the script kept
checking card layout against an unstyled page and kept reporting
pass. It now reads the shell's <link> tags out of index.html, the way
tests/css_snapshot/capture.mjs already does, so the set cannot drift
out from under it again.
- tests/css_snapshot/bench.html's hand-open fallback linked the same
deleted file. Replaced with the ordered set index.html ships.
test_every_stylesheet_referenced_by_shipped_html_exists only walked
static/*.html, which is why neither showed up. Extend it over the bench
page: the bench's whole output is computed styles, so a dead link there
is worth more than an unstyled page nobody looks at.
First slice of the runtime regression lane. Drives stream_agent_loop with a
fake model and asserts on the tools the runtime offers, which is its decision
about what the turn may do. No production runtime code is touched and no
benchmark fixture or allowlist is imported.
The fake-model pattern is the one tests/test_tool_policy.py already uses:
patch stream_llm_with_fallback and inspect the tools kwarg.
Measured on lab@c499c01b, negative web wording is only partially detected:
"Answer from memory only, don't search online." web tools withheld
"Summarise what you already know. Do not search the web." web tools OFFERED
"No web search please, just tell me what you know..." web tools OFFERED
The case that holds is a plain regression guard. The two that do not are
xfail(strict=True): they run on every suite, document the target, and fail the
moment the behaviour lands so the marker gets removed rather than lingering.
A positive control keeps the guard from being satisfied by removing the web
tools altogether.
`_refresh_token_cache` rebuilt the bearer-token map in two steps, `clear()`
then `update()`. Between them the dict a concurrent reader was already holding
was empty, so a valid token landing in that window found zero candidates and
got a 401. The refresh runs on a worker thread via `to_thread`, so the window
is real rather than theoretical.
Build the new map and rebind the name. The reader dereferences the global once
and then holds a map that is complete — the previous one if it read early, the
new one if it read late, never a half-built one. `app.state._token_cache` is
rebound with it, because it was bound once at startup and would otherwise point
at the abandoned dict.
Ported from public `dev` (`984337b3`), with its test.
The upstream test does not pin the fix: both of its concurrency cases pass
against the pre-fix code, because landing a GIL switch inside a window a few
bytecodes wide does not happen across 100 refreshes. They are kept as written
and `TestRefreshLeavesTheReadersMapAlone` is added next to them, stating the
same invariant at object level — a map a reader already holds is not mutated
by a later refresh — which fails on the pre-fix code without depending on
thread scheduling.
`_cached` deduplicates the scheduler's outbound fetches — Miniflux unread
counts and MCP tool snapshots — by parking every concurrent caller on one
shared Future. Two cancellation paths left that Future stranded.
A waiter awaited the shared Future directly, so cancelling the waiter
cancelled the Future the owner and every other waiter were using. It now
awaits through `asyncio.shield`.
The owner removed its pending entry inside the success and `except Exception`
branches. `CancelledError` is a `BaseException`, so it took neither: the key
stayed in `_shared_cache_pending` pointing at a Future nobody would ever
resolve, and every later caller for that key waited forever. Cleanup moves to
a `finally` that is synchronous on purpose, and the owner cancels its own
Future so current waiters wake while a later caller can still retry.
Ported from public `dev` (`ce04dc1d`, #6174 upstream), with its test.
The configuration surface was undiscoverable. .env.example has three active
lines, and of the ODYSSEUS_* variables the code actually reads, most appear
nowhere in .env.example, docs/, website/ or README.md - including several that
change security-relevant behaviour (ODYSSEUS_BROWSER_NO_SANDBOX,
ODYSSEUS_ALLOW_PRIVATE_CALDAV, ODYSSEUS_ENABLE_HOST_DOCKER,
ODYSSEUS_MCP_ALLOWED_COMMANDS). Every question about one of them lands in the
issue tracker.
A hand-written page would drift within a month, so the page is generated:
- scripts/generate_env_reference.py walks the Python sources, collects each read
with the default it falls back to and the file it is read in, groups by area,
and marks the internal variables rather than omitting them.
- website/configuration-reference.md is the generated output, wired into the
Pages layout and linked from setup.md and .env.example. .env.example stays a
short deployment-level example and links onward rather than growing.
- tests/test_env_reference.py regenerates and compares, so adding a variable
without documenting it fails the suite. That is the point: the current state
happened because nothing objected.
Finding the reads needs more than one pattern. Two families - the upload caps in
src/upload_limits.py and the media-ingress overrides in src/media_ingress.py -
are read through helper functions, so the generator detects env-reader helpers
rather than hardcoding a list. Others span two lines, hold the variable name in
a module constant, or read through a mapping passed in as an argument. One lives
inside a string literal, in the Ollama probe script routes/cookbook_helpers.py
builds line by line. A line-based grep for os.environ.get("ODYSSEUS_ finds 70 of
the 98 the generator finds; the page reports that gap and recomputes it on every
run so the claim cannot go stale.
No application behaviour changes.
#25 fixed two sys.modules writes in test_auth_regressions.py that left empty
stub modules behind for the rest of the session, breaking 23 tests under one
collection order while the full suite stayed green. The class is wider than
that file, and an audit is the wrong answer to it: nothing stops the next one,
and the failure it causes lands on an unrelated test in a different file.
So this is a guard instead. An autouse fixture in the root conftest snapshots
which src.* / core.* names are bound to a bare ModuleType, and fails any test
that adds one. "Bare" is the same test the clear_fake_* helpers already use -
a plain types.ModuleType with no on-disk __file__. MagicMock stand-ins are out
of scope: they answer every attribute, so they fail at the point of use rather
than silently, and several files install them deliberately.
Three details that matter:
- It lives in the root conftest, so it is set up before any test-module
fixture and torn down after all of them. A stub a test's own teardown
removes is not reported.
- It drops the leaked entries as well as reporting them, so the failure stays
on the test that introduced it instead of cascading through the rest of the
run.
- It only reports stubs added during the test. Import state the session starts
with, including this conftest's own src.database stub, is left alone.
It found one beyond #25 on the first full run: _stub_heavy in
test_scheduler_restart_doublefire.py leaks the same five src.* modules as the
test #25 fixed, via sys.modules.setdefault. It already receives monkeypatch,
so the fix is to register through it. Fixed here because the guard has to land
green.
Full suite, macOS, default collection order:
this branch 10655 passed, 6 failed, 6 skipped 406s
lab 10655 passed, 6 failed, 6 skipped 371s
Same six either way, which is the point - none of this is visible in the
default order. Four are pre-existing macOS environment failures:
test_glob_confined_e2e and the two test_code_nav_tools document cases resolve
/tmp to /private/tmp, and
test_real_socket_falls_back_from_dead_first_to_live_second is connect-refused
timing on real sockets. The other two are the rich-colour and ffmpeg items
from the same ledger, fixed on their own branches.
Not verified: Linux, and any collection order other than the default. The
guard is order-independent by construction - it compares before and after
within a single test - but I have only run the default order.
test_inspect_media_exports_final_decodable_frame_at_exact_duration exported to
/workspace/final.webp and asserted exit_code == 0. WebP encoding is an ffmpeg
build option, not something this project requires - Homebrew's macOS ffmpeg is
built without it:
ffmpeg -encoders | grep -ic webp -> 0
so the tool returns "ffmpeg still extraction failed: ... Encoder not found"
and the test fails on the build rather than on the code under test.
The test's subject is the final frame being decodable at the exact duration,
which has nothing to do with the container. It now writes a PNG, and the two
things the WebP path was implicitly covering are split out and each guarded on
what is actually present:
- test_inspect_media_exports_a_webp_still - skipped unless ffmpeg reports a
webp encoder, and now asserts the file really is WebP rather than merely
non-empty.
- test_inspect_media_reports_a_missing_encoder_instead_of_crashing - runs only
where the encoder is absent, pinning the behaviour that surfaced this: the
tool reports ffmpeg's failure as a tool error and writes no partial file.
The product code has the same assumption and I left it alone. inspect_media
accepts any suffix in _IMAGE_SUFFIXES and hands the path to ffmpeg, so a .webp
request on a build without libwebp fails with ffmpeg's own message. That is a
poor message, not a crash or a corrupt file, and pre-validating the encoder
list is a separate change.
pytest tests/test_inspect_media_tool.py
68 passed, 2 skipped in 22.92s (1 failed, 66 passed, 1 skipped before)
Not verified: the WebP success path. This machine has no webp encoder, so
test_inspect_media_exports_a_webp_still skips here and has only been checked
for collection, not for a passing run.
test_rich_colors_follow_theme_and_undo_as_one_edit fails identically on every
macOS run, timing out after 30s waiting for a span that never appears. It was
written off as timing noise twice. It is not flaky - it is two Linux-only
input conventions, and the product code is fine.
Control+click: macOS delivers a Control-modified primary click as contextmenu,
not click. Instrumenting the Lemon swatch shows the button receiving
pointerdown, mousedown, contextmenu, pointerup, mouseup - and no click, so the
menu item's handler never runs and no highlight is applied. Control was never
meaningful here anyway; the palette item has no modifier behaviour. The two
calls now use a plain click.
Control+Z: the editor's undo accelerator is Cmd+Z on macOS. With the clicks
fixed, both undo assertions still failed until the presses became
ControlOrMeta+Z, which Playwright maps per platform.
Both fixes are portable - a plain click and ControlOrMeta are unchanged on
Linux, where this test already passes.
pytest tests/test_document_rich_color_reset_and_contrast.py
2 passed in 2.11s (1 failed, 1 passed before)
Not verified: Linux. I only have macOS here, so the claim that this stays
green on CI rests on the modifier being a no-op there, not on a run.
`add_server` wrapped `json.loads(args)` in a bare `except` that fell back to
`[]`, so an Args value that is not JSON — a bare path, which is what the form's
placeholder invites people to type — registered the server and spawned the
stdio subprocess with an empty argv. Nothing surfaced the loss: the POST
returned 200 and the row persisted with `"args": []`.
The route now returns 400 for an unparseable value, and also for valid JSON of
the wrong shape: `args=5` reached `StdioServerParameters(args=5)` and raised an
unhandled TypeError in the error formatter's `" ".join(...)`.
Both form clients mirror the guard instead of leaving the user to read a 400
they cannot see. `settings.js` stops silently defaulting a bad Args value, and
`admin.js` gains the same client-side parse check plus a `res.ok` branch so a
server-side rejection is not reported as a connection failure.
Ported from public `dev` (`9d5c0319`, #6215 upstream, fixing #6211), with its
test. The `admin.js` hunks are inert on `lab` — `initMcpForm` early-returns
because that form's markup is not in this build — and are carried anyway to
keep the two lines from diverging further.
The host cache was gated on the list being non-empty, so "queried fine, no
eligible peers" looked exactly like a cold cache and every caller paid for
another `tailscale status --json` — a subprocess with a 5s timeout.
Gate on the timestamp instead. Failures still leave the timestamp unset, so a
missing binary, a non-zero exit or unparseable output stays retryable rather
than being cached for the full TTL.
Ported from public `dev` (`affaee1e`, #6228 upstream), with its test.
`_cookbook_kill_session` kills the tmux session, then sweeps /proc for
model servers that survived the SIGHUP. The sweep had no guard, so on
macOS and Windows `os.listdir("/proc")` raised FileNotFoundError after
the kill had already succeeded. The function's outer except turned that
into `{"error": "...No such file or directory: '/proc'", "exit_code": 1}`
and skipped the state write that marks the session stopped — the agent
is told a stop failed that actually worked.
Guard the sweep with a procfs check, the way `_scan_running_model_processes`
already does a few hundred lines up. The root and the check now live in
`core/platform_compat`, which is where OS differences belong and which
makes both branches patchable from a test on either kind of host.
Adds a class-level guard test: the third instance of this defect, and two
of the three were found by reading source rather than by a test.
The decomposition lanes have two safety nets and neither covers the
product. The checkpoint benchmark measures the agent runtime; the
computed-style snapshot pins the CSS. Nothing checked that Notes,
Calendar, Documents, Email, Memory, Cookbook or Settings still worked
after a route package moved or a 17,000-line module was split - and the
unit suite does not, since a byte-identical file move can break tests
that pass on the base branch with CI green throughout. The 28 Playwright
specs we do have are all under tests/e2e/photo-editor/ and no workflow
runs them.
scripts/odysseus-smoke boots this worktree through `odysseus dev` and
runs tests/smoke/: one scenario per area, each asserting a user-visible
outcome rather than a status code. Models come from a deterministic
OpenAI-compatible stub on an ephemeral loopback port; email reuses the
existing ODYSSEUS_EMAIL_FIXTURE path rather than inventing a second
mechanism. No scenario touches a live endpoint or the network.
The report is a per-area table that prints the areas the suite does not
cover next to the ones it does, and builds its rows from the registry
rather than from what happened to run, so an area cannot go missing by
having its module deleted or renamed. Under a plain pytest with nothing
booted every scenario skips with the reason, so the full suite stays
green.
`lab` and the public `dev` line share only the repository's first commit as a
merge base, so `git log lab..dev` lists two thousand commits that are almost
all already present on both sides under different SHAs. Nothing in that output
says which public fixes never reached `lab`, which is the only question that
matters before `lab` becomes a release.
`scripts/ref_parity_audit.py` samples the most distinctive added lines from
each commit in the range and searches the other tree for them with
`git grep -F`, then reports the file-level presence diff. The two complement
each other: a commit whose probes are all found while one of the files it added
is missing from the target is a fix whose production change was reproduced
without its test, which the line sampling alone cannot see.
Probes are stripped of indentation and searched for anywhere in the tree, so a
port that moved or was re-indented still reads as present. Verdicts are
absent / partial / present / no-probe, and the report says outright that the
commit verdicts are a heuristic while the two file lists are exact.
Read-only by construction: `git log`, `show`, `diff`, `grep`, `ls-tree` and
`merge-base` only, no remote access, and it does not import the app package.
Three of the six recorded failures were the same test bug: an unresolved
/tmp path compared against a resolved /private/tmp one. macOS makes /tmp a
symlink, so a fixture built with tempfile.mkdtemp(dir="/tmp") and a code path
that resolves what it reports disagree about a file both found correctly.
test_code_nav_tools builds its fixture unresolved and compares it against the
reported path. One realpath fixes both of its failures.
test_glob_confined_e2e is the same cause through a longer route: it mixed
os.path.realpath(ws) with an unresolved secret directory, so relpath emitted
"../../../../tmp/<absolute path>" and the assertion that the absolute path was
absent from the output matched it as a substring. Resolving the secret
directory puts both sides in one tree and the relative path stays short.
macOS full suite goes from 6 failures to 3. The remaining three are an ffmpeg
build without a WebP encoder, a socket test that needs a fast connection
refusal, and the rich-text colour test that is still unexplained.
The ledger is updated in the same change so it does not describe failures that
no longer happen.