Commit Graph
1371 Commits
Author SHA1 Message Date
Léo 2a540f2acc fix(runtime): enforce workspace confinement in one place
"Is this path inside that root" is asked in twenty places in this tree and
answered twenty times by a locally written realpath/commonpath pair. Nine test
files exist because nine call sites each needed their own proof. Each one is
defensible alone; together they are the defect, because the boundary has no
single definition and a site that gets a detail wrong is wrong by itself.

src/path_confinement.py is that definition, and it settles the details the
copies disagreed on. Both sides get canonicalized: comparing a realpath-ed
candidate against a root that was only abspath-ed is the macOS /tmp ->
/private/tmp mismatch that has already produced a false failure here, and
canonicalizing one side is worse than canonicalizing neither. commonpath rather
than startswith, because /a/bc begins with /a/b and is not inside it. A relative
candidate joins the root rather than os.getcwd(), which is whatever directory
the server happens to be running in. NUL and newline are refused with a reason
instead of caught by a bare `except Exception` and reported as an ordinary
escape. Eighteen call sites go through it now. It deliberately does not decide
whether a path is sensitive -- that deny list answers "allowed" rather than
"inside", and it stays with src/tool_execution, which owns it. The one
commonpath left in the tree, in src/workspace_paths.py, stays: that function
translates a host path into a container path, so canonicalizing either side
would change the relative path it computes and break the mapping. It is not a
confinement check.

Two of those sites were weaker than the rest and are fixed rather than moved.
The email attachment check used abspath, which folds `..` but does not resolve
symlinks, so a symlink written into the extraction directory passed it and was
then read through. The skill-reference guard compared a realpath-ed target
against a raw dirname, so on a host where the skills tree is reached through a
symlink the two sides never matched and the guard could not fire.

The execution boundary had two separate holes.

The workspace namespace bound /home and /mnt read-write. On the one platform
where that namespace engages at all, a command inside it reaches outside the
workspace and writes to the user's home directory -- measured by running this
argv on a Linux host with working bubblewrap, not inferred from the source.
Binding the user's whole home directory into a workspace-confinement namespace
gives back most of what the namespace was for. Both are read-only now. The
workspace is also bound writable at its real host path, not only at /workspace:
BashTool's own /tmp redirect rewrites `/tmp/` to `<agent_cwd()>/.tmp/` before
the namespace is built, so the command bwrap receives already names the real
path, and those writes previously landed only because the workspace happened to
sit under the writable /home.

`namespaced or _replace_workspace_alias(...)` chose between a mount namespace
and a regex with nothing in the result saying which one ran. The fallback
rewrites the literal token /workspace in the command string, so a command that
never mentions /workspace is untouched by it and runs on the host unrestricted
-- which is every agent shell command on macOS. Both tools now ask
containment.probe() instead of each deciding for itself, and every bash and
python result carries a containment block naming the mechanism and stating
whether the filesystem dimension actually held. Under enforcing mode the
command is not run and the result says so.

That block reports the filesystem dimension only, and says so in a
reported_dimensions field. The probe knows this host could also give a process
group and a real wall clock, but these two tools still assemble their own
create_subprocess_* call and pass neither, so listing those dimensions would be
exactly the false claim src/containment.py calls worse than an honest absence.

probe() is new on src/containment.py: the same mechanism table and the same
arithmetic as acquire(), stopping before the side effects. acquire() is the
wrong shape for a decision -- it writes a durable grant record, and a record
whose pid is never filled in and whose release() never runs is an entry a
restart reaper keeps finding.

CONTAINMENT_MODE stays report_only. Flipping it refuses every agent shell
command on macOS and on any Linux host without bubblewrap, which is a product
decision rather than a code one.

Smaller things in the same area: the /tmp redirect's makedirs was unguarded, so
a read-only workspace turned a command that merely mentioned `/tmp/` into an
OSError traceback instead of a tool error; it degrades now. WORKSPACE_MOUNT
moved to src/constants.py so the namespace and the path resolvers read one
definition of the contract rather than two. The ".tmp" dirname got a constant,
since it appeared in both tool paths.

One generated artifact moved with it: website/configuration-reference.md pins
the source line where each ODYSSEUS_* variable is read, and three of those
shifted. Regenerated with scripts/generate_env_reference.py; the diff is line
numbers only.

Three existing tests changed. test_workspace_artifact_tool_floor asserted that
an unsafe interpreter prefix produces no `--ro-bind <prefix> <prefix>`, which
now fires on /home because /home is legitimately a read-only base mount.
Asserting the absence of a literal flag string cannot distinguish "the prefix
was rejected" from "the argv mounted that root itself", so it compares the argv
against the no-prefix baseline instead: an unsafe prefix must add nothing.

The Windows bash test asserted dict equality on the
whole result, which makes adding a field to every bash result impossible without
touching a test about tmux; it asserts the shape now. The personal-dir symlink
test grepped the resolver's source for the literal "os.path.realpath", which is
gone because the resolution moved into the shared boundary -- it keeps the
negative assertion that the closure must not grow its own abspath check again,
and the behavioural half now runs against the boundary, where it covers every
call site instead of one closure.

Not verified: the bubblewrap argv is asserted, not executed. There is no bwrap
on macOS, and in Docker it needs --privileged to work at all -- default and
seccomp=unconfined both fail with "Creating new namespace failed", and
--cap-add=SYS_ADMIN fails at pivot_root. The Python tool's
needs_virtual_namespace gate means ordinary Python code gets no namespace even
on a Linux host that could provide one; that is reported now but deliberately
not changed, because it alters the Linux Python path on every call and cannot be
checked from here.
2026-10-01 19:45:59 +02:00
Léo c004a26d46 fix(runtime): verify process identity before any teardown signal
A recorded pid is a claim, not a handle. The containment grant store, the
background-job store and the Cookbook task list all outlive the process that
wrote them — deliberately, so a restart keeps a job and its result — and the
kernel reuses pids. Any teardown driven off one of those records can therefore
land on a process we never started. ODY-86 was exactly this, and the Cookbook
survivor sweep still terminated any process whose full command line matched a
tracked one, which is the same mistake spelled differently.

Identity is (pid, start token). The token comes from /proc/<pid>/stat on Linux,
ps -o lstart= on macOS and the BSDs, and GetProcessTimes on Windows; the kernel
will not hand a pid to a process that started earlier, so comparing the token
recorded at launch against the token read now answers "is this still ours"
without a handle or a supervisor. verify() returns owned, gone, foreign or
unverifiable, and only owned permits a signal.

Keeping "unverifiable" out of the other two is the point. Process inspection
has broken off Linux four times here — ODY-70, -86, -94, -99 — every time
because an absent mechanism read as a successful answer. Folding it into "ours"
signals strangers; folding it into "gone" abandons live processes. It is a
containment failure and every caller treats it as one.

Wired into the three places that signal:

- containment.release() gates a grant recovered from the durable store, and
  leaves an in-process teardown alone, where the caller holds the child and no
  identity question arises. The verdict lands on the record, so "why is this
  grant still here" is answerable afterwards.

- A startup reaper. Nothing read either store before, so a crashed run left
  every grant permanently active and every job permanently running, and the
  first thing to touch such a record was a teardown aimed at a reassigned pid.
  The two stores get opposite treatment: an orphaned grant has no caller left
  and is torn down, while a detached job is documented to survive a restart and
  is only corrected, never killed.

- The Cookbook sweep takes its ownership from the tmux pane's process tree,
  captured before the kill destroys the only link between a surviving model
  server and the session that started it. A process that merely matches the
  tracked command line is now reported rather than killed: the Cookbook
  composed that command line, so an identical one is just as likely to be a
  server the user started by hand. The sweep also runs on hosts with no procfs
  instead of silently skipping, and says so when it could not look at all.
2026-10-01 18:55:50 +02:00
Léo 33fa27b4c8 feat(runtime): add the containment boundary and its failure contract
Agent-reachable execution has 25 independent spawn sites and no single
place deciding where a process runs or under what limits. All three
consequences are visible on this SHA. When bwrap is absent the workspace
namespace degrades to a regex that rewrites /workspace to the real path,
and nothing in the tool result says which one you got. No spawn site
passes start_new_session, so a wall-clock kill reaches the wrapper shell
and leaves its backgrounded grandchildren running while reporting the
process killed. Teardown stops at SIGTERM without ever checking death.

src/containment.py gives those paths one boundary. acquire() establishes
containment or refuses -- a string rewrite is not a mechanism it can
select -- and the grant states which dimensions actually hold, which were
best-effort and are missing, and which were required and are missing.
run() enforces the wall clock and the output cap. release() signals the
process group, escalates to SIGKILL, and reports dead only for a group it
observed go empty.

Containment never sees the command: acquire() takes a workspace and
limits, and the command text only reaches run(). Nothing in a request can
widen a boundary it is never shown.

CONTAINMENT_MODE chooses between refusing an unestablishable required
dimension and recording it. It ships report-only, so landing this changes
no behaviour on a host without bwrap -- which is every host today.

No call sites move here; they follow on this branch. The configuration
reference is regenerated because the page records src/constants.py line
numbers and the new path constant shifts two of them.
2026-10-01 17:55:33 +02:00
Alexandre Teixeira d49071bbec fix: close Wave 1.1 completion-gate audit findings
- Headless consumers (task scheduler, background follow-up) now treat a
  completion-gate final_response as the authoritative answer instead of
  collecting deltas only. A gated replacement no longer leaves scheduled
  output empty, which used to trigger an extra, ungated grace-summary
  model call.
- The scheduler closes the agent stream with contextlib.aclosing, so the
  approval-pause break unwinds the gate's journal and teacher-takeover
  context in its own task. Chained runs no longer inherit a stale
  parent_run_id, and later finalization no longer raises ContextVar
  reset errors.
- On provider error, the completion gate applies the live answer's
  statement filter to persisted round_texts. Diagnostics and the failure
  note survive; claims rejected by the gate cannot reappear on reload.
2026-10-01 14:49:32 +01:00
Alexandre Teixeira f4793696f4 merge: reconcile Wave 1.1 with post-PR40 lab
Merge canonical lab 9557b8d5909eb4a885c3bf49e19a65dd904f8c1d exactly once.
Retain invocation journal ownership and lineage, provider terminal ordering,
teacher handoff, framed DONE handling, and canonical authority/Ajax routing.

Combine dynamic dispatch receipts with lab policy forwarding. Adapt native
shell/patch evidence, explicit TUI verifiers, and artifact recovery presentation.
Refresh generated configuration source links and strengthen adapter regressions.

Validation: focused 2118 passed; Wave 1.1 script 2291 passed; broad runtime
5649 passed; full pytest 11581 passed, 53 skipped, 2 xfailed, 6 subtests passed.
Compileall 1689 Python files; syntax 279 JS and 82 MJS files; diff and
conflict-marker checks passed.
2026-10-01 09:09:55 +01:00
Alexandre Teixeira 5dfe1c353c test: stabilize final PR 40 validation gates 2026-10-01 06:39:22 +01:00
Alexandre Teixeira f74a262f73 merge: reconcile PR 40 with current lab
Integrate lab fff55a78 into PR #40 (cc25d5ba). Lab's modular email
backend/frontend, modular settings, split stylesheets (static/style.css
stays deleted), procfs compatibility, and request-scoped TurnContract
authority win; PR #40's routing classifiers, editor/email/task features,
and style.css changes are ported into lab's module and stylesheet homes.

Integration fixes:
- settings/api.js imports ui.js under its canonical versioned URL
- browser observations keep legacy CAPTCHA/access-block evidence
- artifact turns do not re-trigger broad-web research recovery
- env reference documents PR test-tool variables; page regenerated

PR #40 defects surfaced by lab gates and fixed here:
- web_fetch generic schema drops top-level anyOf (OpenAI contract);
  the compact preview contract still requires url or urls
- get_weather registered as a brokered network read
- new lazy editor modules precached for offline use
- SearXNG pin mirrored into GPU standalone compose files
- image model picker again skips offline endpoints

Tests updated where PR #40 changed behaviour on purpose, and PR tests
moved onto lab's document_source helpers.
2026-10-01 05:03:58 +01:00
Alexandre Teixeira d63932f1ad Merge lab into fix/procfs-pid-file-liveness 2026-10-01 03:10:10 +01:00
Alexandre Teixeira 6901c56192 Merge lab into feat/admin-build-provenance 2026-10-01 02:55:27 +01:00
Alexandre Teixeira c6f690a27b Merge lab into docs/env-configuration-reference 2026-10-01 02:52:37 +01:00
Alexandre Teixeira bdccb1ddb1 Merge lab into test/95-release-smoke-suite 2026-10-01 02:45:49 +01:00
Alexandre Teixeira ef0d96a3ae Merge lab into audit/ref-parity 2026-10-01 02:45:10 +01:00
Alexandre Teixeira 8c7e3a9411 Merge lab into test/runtime-behavior-regressions 2026-10-01 02:42:22 +01:00
Alexandre Teixeira 45e1e2e7d9 Merge lab into fix/token-cache-atomic-swap 2026-10-01 02:41:41 +01:00
pewdiepie-archdaemon 2e8413a54a Preserve preview harness, editor, email and task improvements
Snapshot current maintainer-preview application changes and regression fixtures for integration into lab. Excludes local runtime data, evaluation outputs and source backups. Focused Python regression selection: 140 passed; full suite not certified.
2026-10-01 01:34:26 +00:00
Alexandre Teixeira b0bc0b8c40 Merge lab into fix/6174-singleflight-cancel 2026-10-01 02:27:49 +01:00
Alexandre Teixeira 799adbbde6 Merge lab into refactor/email-library-package 2026-10-01 02:22:23 +01:00
Alexandre Teixeira 9477616e05 Merge lab into refactor/routes-email-subpackage 2026-10-01 02:15:36 +01:00
Alexandre Teixeira aedec7d005 fix(runtime): isolate nested invocation ownership 2026-10-01 02:11:53 +01:00
Alexandre Teixeira 33387d4a4c Merge lab into refactor/settings-shell-modules 2026-10-01 01:41:27 +01:00
Alexandre Teixeira 4c122de880 fix(runtime): scope completion claims to execution obligations 2026-10-01 01:36:18 +01:00
Alexandre Teixeira 9ee9205bfc merge: sync lab after stylesheet split 2026-10-01 01:08:15 +01:00
Alexandre Teixeira 86212bfe99 Merge lab into refactor/settings-panel-modules 2026-10-01 01:03:49 +01:00
Alexandre Teixeira 23fdc26079 Merge lab into refactor/split-style-css 2026-10-01 00:48:58 +01:00
Alexandre Teixeira 466a6b323a fix(runtime): preserve provider error terminal ordering 2026-10-01 00:33:46 +01:00
Alexandre Teixeira fe80eb795e merge: reconcile reviewed lab baseline for runtime wave 1.1 2026-10-01 00:25:42 +01:00
Alexandre Teixeira 40dbe17a95 Merge lab into test/document-module-set-helper
Resolve the overlap with #19 by preserving the whole-cascade stylesheet
helpers alongside #31's document-source/module-set helpers.

Maintainer validation:
- document/module contract and composition guards: 15 passed
- all 54 touched Python test modules: 366 passed
- py_compile: clean
- git diff --check: clean
2026-09-30 22:56:56 +01:00
Léo 40678bc466 feat(settings): show the running build's version and commit in the admin panel
Nothing in the UI said which build was loaded. /api/version has reported
version, build and source_commit since the harness started versioning itself
apart from the public semver, but the only way to read it was to curl the
endpoint — so "is the preview actually running the commit I just merged?" took
a terminal to answer.

Pins a footer under the settings sidebar nav showing the registered version
(plus the harness build when it differs) and the short source commit, with the
full hash on hover. It sits outside the nav's scroll container so it stays at
the bottom-left, and it is .admin-only, so syncAdminVisibility() hides it from
non-admins the same way it hides the Admin nav group.

The commit resolves at import via `git rev-parse HEAD` and is the string
"unknown" when that fails — a read-only Docker tree with no .git. The footer
treats "unknown" as absent and stays hidden when nothing is left to show,
rather than printing it. The collapsed rail and the two narrow tab-rail
layouts hide it too: neither has a bottom-left to write in.
2026-09-30 22:59:58 +02:00
Alexandre Teixeira 741f5dff2b Merge pull request #19 from o3LL/refactor/tests-read-the-whole-cascade
refactor(tests): read the whole cascade instead of style.css alone
2026-09-30 19:27:58 +01:00
Alexandre Teixeira 6ee51fee72 Merge pull request #18 from o3LL/fix/duplicate-keyframes
fix(css): collapse duplicate @keyframes names to the definition that wins
2026-09-30 19:18:13 +01:00
Alexandre Teixeira a969a68016 Merge pull request #34 from o3LL/test/macos-failures-and-stub-leak-guard
test: fix environment-dependent failures and guard module-stub leaks
2026-09-30 19:16:43 +01:00
Alexandre Teixeira 56484df737 test(media): detect ffmpeg encoders by codec alias 2026-09-30 19:16:29 +01:00
Alexandre Teixeira eaddc03729 Merge pull request #26 from o3LL/fix/tmp-realpath-test-bugs
fix(tests): resolve temp paths consistently on macOS
2026-09-30 18:21:04 +01:00
Alexandre Teixeira bfae249120 Merge pull request #28 from o3LL/fix/cookbook-stop-procfs-guard
fix(cookbook): skip the pid sweep when the host has no procfs
2026-09-30 18:07:07 +01:00
Alexandre Teixeira 054df080af Merge pull request #32 from o3LL/fix/6215-mcp-args-validation
fix(mcp): reject malformed Args on Add MCP Server instead of defaulting to []
2026-09-30 17:18:51 +01:00
Léo bf5d8e4001 test(runtime): cover the remaining four requested behaviours
Completes the lane alteixeira20 asked for. Test-only: tests/ and test helpers,
no production runtime code, no benchmark fixtures or allowlists.

Supplied workspace context must not produce a clarification. Pins
_looks_like_unattended_clarification on four shapes that hand the decision
back ("could you please share", "shall I", "which approach do you prefer")
and three ordinary answers that must not trip it.

Repeated update_plan is not the turn's work. ask_user and update_plan are
permitted on nearly every turn, so if they counted as execution a model could
loop on them and look busy. Pins that _tool_rejection_reason does not
advertise either as an available tool, and that update_plan is permitted
without ever being in required.

Request-scoped tool authority. _request_scoped_allowed_tool_names must not
make an undeclared tool executable; the native-terminal widening is pinned
separately so it stays opt-in rather than drifting into the default.

Foreign-process safety. The Chrome sweep matches this runtime's own profile
prefix, so a fake procfs with our pid, a user's ordinary Chrome and another
worktree's agent browser must leave exactly two of the three alone.

14 passed, 2 xfailed. The xfails are the negative-wording cases from the first
commit that do not hold yet.
2026-09-30 17:38:57 +02:00
Léo 9738405310 fix(tests): bind the docker-socket fixtures somewhere sun_path fits
Four more tests in the same family as the /tmp ones this change already
fixes, and they hide for the same reason: the failure depends on how
long $TMPDIR happens to be.

tests/test_shell_routes.py::TestHostDockerAccess (three) and
tests/test_cookbook_docker_access.py::test_container_opt_in_with_unix_
socket_is_allowed each bind an AF_UNIX socket at tmp_path/"docker.sock".
macOS gives sun_path 104 bytes including the terminator. pytest roots
tmp_path at $TMPDIR, which on a stock Mac is a 49-character
/var/folders/<2>/<30>/T/; add pytest-of-<user>/pytest-<n>/ and the
test's own name and the bind path is 115 bytes before the filename.

    OSError: AF_UNIX path too long

Linux allows 108 and roots $TMPDIR at /tmp, so CI never sees it. Under a
shortened $TMPDIR the path lands at exactly 103 and passes — until
pytest's run counter reaches two digits and it becomes 104. That is why
the ledger's counts did not include these: they were measured somewhere
the path fit.

Adds tests/helpers/unix_sockets.bound_unix_socket, which binds under a
short directory and asserts the length before it tries, so the next
socket fixture fails with a sentence rather than an errno. Records the
trap in KNOWN_FAILURES.md along with the instruction to re-measure with
the default $TMPDIR.
2026-09-30 17:35:06 +02:00
Léo d423632559 test(runtime): scope the negative-wording claim to the inferred path
Verified end to end against a local Qwen3.5-9B Q4_K_M that when the user
explicitly enables web for the turn, none of the three phrasings withholds
web_search, web_fetch or private_browser, including the one these tests record
as held. The held case holds on the inferred path only, where no toggle is set
and the runtime decides from intent.

That distinction was missing and the file read as a stronger claim than the
measurement supports. An explicit toggle beating an inferred negative may be
the intended semantics, so it is recorded rather than asserted.
2026-09-30 17:19:07 +02:00
Léo 5f18767528 fix(mcp): show the route's rejection reason on the Integrations form too
The route now answers 400 with a message naming the expected shape.
admin.js was taught to print `data.detail`; the Unified Integrations
form in settings.js still printed `Failed (400)` and dropped it.

That gap is exactly where the new validation bites. The client-side
JSON.parse guard added alongside it catches unparseable input, so the
only values that reach the route's 400 are ones that parse but are not
a list — `"npx"`, `{}`, `null` — and for those the status code alone
tells the user nothing about what is wrong with what they typed.

Adds source-level coverage for both forms; the PR changed two JS files
with no test on either.
2026-09-30 17:15:52 +02:00
Léo 51e09a1e32 fix(css): repoint the two stylesheet links the split left behind
Splitting style.css deleted it, and two files outside static/ still
named it:

- scripts/verify_background_research_cards.mjs injected
  `<link rel="stylesheet" href="/static/style.css">` into the page it
  builds. A stylesheet that 404s does not fail — the script kept
  checking card layout against an unstyled page and kept reporting
  pass. It now reads the shell's <link> tags out of index.html, the way
  tests/css_snapshot/capture.mjs already does, so the set cannot drift
  out from under it again.
- tests/css_snapshot/bench.html's hand-open fallback linked the same
  deleted file. Replaced with the ordered set index.html ships.

test_every_stylesheet_referenced_by_shipped_html_exists only walked
static/*.html, which is why neither showed up. Extend it over the bench
page: the bench's whole output is computed styles, so a dead link there
is worth more than an unstyled page nobody looks at.
2026-09-30 17:14:52 +02:00
Léo 5ce2394adb fix(browser): probe pid liveness through the platform-safe helper
_process_is_alive used os.kill(pid, 0). That probe is POSIX-only:
CPython's Windows os.kill calls TerminateProcess(handle, sig) for any
signal other than CTRL_C/CTRL_BREAK, so it terminates the process it is
asked about. This function is only reached when there is no procfs to
read a command line from, which is exactly the macOS and Windows case
the rest of this change exists to handle.

core/platform_compat.py already owns that probe and documents the
hazard; its module docstring asks callers to import from there rather
than spell a POSIX-only call out locally. Delegate to it.

pid_alive answers False where os.kill raises PermissionError — a live
process owned by another user. Both call sites want that reading: the
sweep only unlinks a pid file it wrote itself, and a pid it cannot
confirm is not the daemon it is looking for.
2026-09-30 17:13:29 +02:00
Léo 5fd9114882 test(runtime): pin negative capability wording, test-only
First slice of the runtime regression lane. Drives stream_agent_loop with a
fake model and asserts on the tools the runtime offers, which is its decision
about what the turn may do. No production runtime code is touched and no
benchmark fixture or allowlist is imported.

The fake-model pattern is the one tests/test_tool_policy.py already uses:
patch stream_llm_with_fallback and inspect the tools kwarg.

Measured on lab@c499c01b, negative web wording is only partially detected:

  "Answer from memory only, don't search online."            web tools withheld
  "Summarise what you already know. Do not search the web."  web tools OFFERED
  "No web search please, just tell me what you know..."      web tools OFFERED

The case that holds is a plain regression guard. The two that do not are
xfail(strict=True): they run on every suite, document the target, and fail the
moment the behaviour lands so the marker gets removed rather than lingering.
A positive control keeps the guard from being satisfied by removing the web
tools altogether.
2026-09-30 17:11:10 +02:00
Léo 5e3e153fba fix(auth): swap the API token cache atomically instead of clearing it
`_refresh_token_cache` rebuilt the bearer-token map in two steps, `clear()`
then `update()`. Between them the dict a concurrent reader was already holding
was empty, so a valid token landing in that window found zero candidates and
got a 401. The refresh runs on a worker thread via `to_thread`, so the window
is real rather than theoretical.

Build the new map and rebind the name. The reader dereferences the global once
and then holds a map that is complete — the previous one if it read early, the
new one if it read late, never a half-built one. `app.state._token_cache` is
rebound with it, because it was bound once at startup and would otherwise point
at the abandoned dict.

Ported from public `dev` (`984337b3`), with its test.

The upstream test does not pin the fix: both of its concurrency cases pass
against the pre-fix code, because landing a GIL switch inside a window a few
bytecodes wide does not happen across 100 refreshes. They are kept as written
and `TestRefreshLeavesTheReadersMapAlone` is added next to them, stating the
same invariant at object level — a map a reader already holds is not mutated
by a later refresh — which fails on the pre-fix code without depending on
thread scheduling.
2026-09-30 16:07:12 +02:00
Léo 6e4b3aa5bd fix(tasks): clean up the singleflight cache on cancellation
`_cached` deduplicates the scheduler's outbound fetches — Miniflux unread
counts and MCP tool snapshots — by parking every concurrent caller on one
shared Future. Two cancellation paths left that Future stranded.

A waiter awaited the shared Future directly, so cancelling the waiter
cancelled the Future the owner and every other waiter were using. It now
awaits through `asyncio.shield`.

The owner removed its pending entry inside the success and `except Exception`
branches. `CancelledError` is a `BaseException`, so it took neither: the key
stayed in `_shared_cache_pending` pointing at a Future nobody would ever
resolve, and every later caller for that key waited forever. Cleanup moves to
a `finally` that is synchronous on purpose, and the owner cancels its own
Future so current waiters wake while a later caller can still retry.

Ported from public `dev` (`ce04dc1d`, #6174 upstream), with its test.
2026-09-30 16:05:09 +02:00
Léo 5c8611fba6 docs: generate the ODYSSEUS_* configuration reference from the source
The configuration surface was undiscoverable. .env.example has three active
lines, and of the ODYSSEUS_* variables the code actually reads, most appear
nowhere in .env.example, docs/, website/ or README.md - including several that
change security-relevant behaviour (ODYSSEUS_BROWSER_NO_SANDBOX,
ODYSSEUS_ALLOW_PRIVATE_CALDAV, ODYSSEUS_ENABLE_HOST_DOCKER,
ODYSSEUS_MCP_ALLOWED_COMMANDS). Every question about one of them lands in the
issue tracker.

A hand-written page would drift within a month, so the page is generated:

- scripts/generate_env_reference.py walks the Python sources, collects each read
  with the default it falls back to and the file it is read in, groups by area,
  and marks the internal variables rather than omitting them.
- website/configuration-reference.md is the generated output, wired into the
  Pages layout and linked from setup.md and .env.example. .env.example stays a
  short deployment-level example and links onward rather than growing.
- tests/test_env_reference.py regenerates and compares, so adding a variable
  without documenting it fails the suite. That is the point: the current state
  happened because nothing objected.

Finding the reads needs more than one pattern. Two families - the upload caps in
src/upload_limits.py and the media-ingress overrides in src/media_ingress.py -
are read through helper functions, so the generator detects env-reader helpers
rather than hardcoding a list. Others span two lines, hold the variable name in
a module constant, or read through a mapping passed in as an argument. One lives
inside a string literal, in the Ollama probe script routes/cookbook_helpers.py
builds line by line. A line-based grep for os.environ.get("ODYSSEUS_ finds 70 of
the 98 the generator finds; the page reports that gap and recomputes it on every
run so the claim cannot go stale.

No application behaviour changes.
2026-09-30 15:11:19 +02:00
Léo fba6f73260 test: fail the test that leaks a bare src/core module stub
#25 fixed two sys.modules writes in test_auth_regressions.py that left empty
stub modules behind for the rest of the session, breaking 23 tests under one
collection order while the full suite stayed green. The class is wider than
that file, and an audit is the wrong answer to it: nothing stops the next one,
and the failure it causes lands on an unrelated test in a different file.

So this is a guard instead. An autouse fixture in the root conftest snapshots
which src.* / core.* names are bound to a bare ModuleType, and fails any test
that adds one. "Bare" is the same test the clear_fake_* helpers already use -
a plain types.ModuleType with no on-disk __file__. MagicMock stand-ins are out
of scope: they answer every attribute, so they fail at the point of use rather
than silently, and several files install them deliberately.

Three details that matter:

- It lives in the root conftest, so it is set up before any test-module
  fixture and torn down after all of them. A stub a test's own teardown
  removes is not reported.
- It drops the leaked entries as well as reporting them, so the failure stays
  on the test that introduced it instead of cascading through the rest of the
  run.
- It only reports stubs added during the test. Import state the session starts
  with, including this conftest's own src.database stub, is left alone.

It found one beyond #25 on the first full run: _stub_heavy in
test_scheduler_restart_doublefire.py leaks the same five src.* modules as the
test #25 fixed, via sys.modules.setdefault. It already receives monkeypatch,
so the fix is to register through it. Fixed here because the guard has to land
green.

Full suite, macOS, default collection order:

  this branch   10655 passed, 6 failed, 6 skipped   406s
  lab           10655 passed, 6 failed, 6 skipped   371s

Same six either way, which is the point - none of this is visible in the
default order. Four are pre-existing macOS environment failures:
test_glob_confined_e2e and the two test_code_nav_tools document cases resolve
/tmp to /private/tmp, and
test_real_socket_falls_back_from_dead_first_to_live_second is connect-refused
timing on real sockets. The other two are the rich-colour and ffmpeg items
from the same ledger, fixed on their own branches.

Not verified: Linux, and any collection order other than the default. The
guard is order-independent by construction - it compares before and after
within a single test - but I have only run the default order.
2026-09-30 13:02:34 +02:00
Léo eb98aa6dc2 test(media): stop asserting an optional ffmpeg webp encoder
test_inspect_media_exports_final_decodable_frame_at_exact_duration exported to
/workspace/final.webp and asserted exit_code == 0. WebP encoding is an ffmpeg
build option, not something this project requires - Homebrew's macOS ffmpeg is
built without it:

  ffmpeg -encoders | grep -ic webp   ->  0

so the tool returns "ffmpeg still extraction failed: ... Encoder not found"
and the test fails on the build rather than on the code under test.

The test's subject is the final frame being decodable at the exact duration,
which has nothing to do with the container. It now writes a PNG, and the two
things the WebP path was implicitly covering are split out and each guarded on
what is actually present:

- test_inspect_media_exports_a_webp_still - skipped unless ffmpeg reports a
  webp encoder, and now asserts the file really is WebP rather than merely
  non-empty.
- test_inspect_media_reports_a_missing_encoder_instead_of_crashing - runs only
  where the encoder is absent, pinning the behaviour that surfaced this: the
  tool reports ffmpeg's failure as a tool error and writes no partial file.

The product code has the same assumption and I left it alone. inspect_media
accepts any suffix in _IMAGE_SUFFIXES and hands the path to ffmpeg, so a .webp
request on a build without libwebp fails with ffmpeg's own message. That is a
poor message, not a crash or a corrupt file, and pre-validating the encoder
list is a separate change.

  pytest tests/test_inspect_media_tool.py
  68 passed, 2 skipped in 22.92s      (1 failed, 66 passed, 1 skipped before)

Not verified: the WebP success path. This machine has no webp encoder, so
test_inspect_media_exports_a_webp_still skips here and has only been checked
for collection, not for a passing run.
2026-09-30 13:02:34 +02:00
Léo 24e428d9cb test(document): use platform-correct input in the rich color test
test_rich_colors_follow_theme_and_undo_as_one_edit fails identically on every
macOS run, timing out after 30s waiting for a span that never appears. It was
written off as timing noise twice. It is not flaky - it is two Linux-only
input conventions, and the product code is fine.

Control+click: macOS delivers a Control-modified primary click as contextmenu,
not click. Instrumenting the Lemon swatch shows the button receiving
pointerdown, mousedown, contextmenu, pointerup, mouseup - and no click, so the
menu item's handler never runs and no highlight is applied. Control was never
meaningful here anyway; the palette item has no modifier behaviour. The two
calls now use a plain click.

Control+Z: the editor's undo accelerator is Cmd+Z on macOS. With the clicks
fixed, both undo assertions still failed until the presses became
ControlOrMeta+Z, which Playwright maps per platform.

Both fixes are portable - a plain click and ControlOrMeta are unchanged on
Linux, where this test already passes.

  pytest tests/test_document_rich_color_reset_and_contrast.py
  2 passed in 2.11s       (1 failed, 1 passed before)

Not verified: Linux. I only have macOS here, so the claim that this stays
green on CI rests on the modifier being a no-op there, not on a run.
2026-09-30 13:02:34 +02:00
Léo 01b8ac5fea fix(mcp): reject malformed Args on Add MCP Server instead of defaulting to []
`add_server` wrapped `json.loads(args)` in a bare `except` that fell back to
`[]`, so an Args value that is not JSON — a bare path, which is what the form's
placeholder invites people to type — registered the server and spawned the
stdio subprocess with an empty argv. Nothing surfaced the loss: the POST
returned 200 and the row persisted with `"args": []`.

The route now returns 400 for an unparseable value, and also for valid JSON of
the wrong shape: `args=5` reached `StdioServerParameters(args=5)` and raised an
unhandled TypeError in the error formatter's `" ".join(...)`.

Both form clients mirror the guard instead of leaving the user to read a 400
they cannot see. `settings.js` stops silently defaulting a bad Args value, and
`admin.js` gains the same client-side parse check plus a `res.ok` branch so a
server-side rejection is not reported as a connection failure.

Ported from public `dev` (`9d5c0319`, #6215 upstream, fixing #6211), with its
test. The `admin.js` hunks are inert on `lab` — `initMcpForm` early-returns
because that form's markup is not in this build — and are carried anyway to
keep the two lines from diverging further.
2026-09-30 12:54:09 +02:00
Léo f8269a829f fix(discovery): cache a successful but empty Tailscale lookup
The host cache was gated on the list being non-empty, so "queried fine, no
eligible peers" looked exactly like a cold cache and every caller paid for
another `tailscale status --json` — a subprocess with a 5s timeout.

Gate on the timestamp instead. Failures still leave the timestamp unset, so a
missing binary, a non-zero exit or unparseable output stays retryable rather
than being cached for the full TTL.

Ported from public `dev` (`affaee1e`, #6228 upstream), with its test.
2026-09-30 12:14:03 +02:00