Commit Graph
2440 Commits
Author SHA1 Message Date
Alexandre Teixeira 11ecf46abb Merge pull request #30 from o3LL/fix/6228-tailscale-empty-lookup-cache
fix(discovery): cache a successful but empty Tailscale lookup
2026-09-30 17:05:16 +01:00
Léo bf5d8e4001 test(runtime): cover the remaining four requested behaviours
Completes the lane alteixeira20 asked for. Test-only: tests/ and test helpers,
no production runtime code, no benchmark fixtures or allowlists.

Supplied workspace context must not produce a clarification. Pins
_looks_like_unattended_clarification on four shapes that hand the decision
back ("could you please share", "shall I", "which approach do you prefer")
and three ordinary answers that must not trip it.

Repeated update_plan is not the turn's work. ask_user and update_plan are
permitted on nearly every turn, so if they counted as execution a model could
loop on them and look busy. Pins that _tool_rejection_reason does not
advertise either as an available tool, and that update_plan is permitted
without ever being in required.

Request-scoped tool authority. _request_scoped_allowed_tool_names must not
make an undeclared tool executable; the native-terminal widening is pinned
separately so it stays opt-in rather than drifting into the default.

Foreign-process safety. The Chrome sweep matches this runtime's own profile
prefix, so a fake procfs with our pid, a user's ordinary Chrome and another
worktree's agent browser must leave exactly two of the three alone.

14 passed, 2 xfailed. The xfails are the negative-wording cases from the first
commit that do not hold yet.
2026-09-30 17:38:57 +02:00
Léo 9738405310 fix(tests): bind the docker-socket fixtures somewhere sun_path fits
Four more tests in the same family as the /tmp ones this change already
fixes, and they hide for the same reason: the failure depends on how
long $TMPDIR happens to be.

tests/test_shell_routes.py::TestHostDockerAccess (three) and
tests/test_cookbook_docker_access.py::test_container_opt_in_with_unix_
socket_is_allowed each bind an AF_UNIX socket at tmp_path/"docker.sock".
macOS gives sun_path 104 bytes including the terminator. pytest roots
tmp_path at $TMPDIR, which on a stock Mac is a 49-character
/var/folders/<2>/<30>/T/; add pytest-of-<user>/pytest-<n>/ and the
test's own name and the bind path is 115 bytes before the filename.

    OSError: AF_UNIX path too long

Linux allows 108 and roots $TMPDIR at /tmp, so CI never sees it. Under a
shortened $TMPDIR the path lands at exactly 103 and passes — until
pytest's run counter reaches two digits and it becomes 104. That is why
the ledger's counts did not include these: they were measured somewhere
the path fit.

Adds tests/helpers/unix_sockets.bound_unix_socket, which binds under a
short directory and asserts the length before it tries, so the next
socket fixture fails with a sentence rather than an errno. Records the
trap in KNOWN_FAILURES.md along with the instruction to re-measure with
the default $TMPDIR.
2026-09-30 17:35:06 +02:00
Léo d423632559 test(runtime): scope the negative-wording claim to the inferred path
Verified end to end against a local Qwen3.5-9B Q4_K_M that when the user
explicitly enables web for the turn, none of the three phrasings withholds
web_search, web_fetch or private_browser, including the one these tests record
as held. The held case holds on the inferred path only, where no toggle is set
and the runtime decides from intent.

That distinction was missing and the file read as a stronger claim than the
measurement supports. An explicit toggle beating an inferred negative may be
the intended semantics, so it is recorded rather than asserted.
2026-09-30 17:19:07 +02:00
Léo 1408bce0ab docs(email): point the specs at the canonical module paths
The move left three documents naming `routes/email_routes.py`,
`routes/email_helpers.py` and `routes/email_pollers.py` as where the
code is. The shims keep those import paths working, so nothing breaks —
but each of those files is now seventeen lines that redirect, and a
reader sent there finds no email code at all.

Follows the phrasing specs/persistence.md already uses for the
contacts and vault subpackages: name the canonical path and note the
shims.
2026-09-30 17:17:39 +02:00
Léo 5f18767528 fix(mcp): show the route's rejection reason on the Integrations form too
The route now answers 400 with a message naming the expected shape.
admin.js was taught to print `data.detail`; the Unified Integrations
form in settings.js still printed `Failed (400)` and dropped it.

That gap is exactly where the new validation bites. The client-side
JSON.parse guard added alongside it catches unparseable input, so the
only values that reach the route's 400 are ones that parse but are not
a list — `"npx"`, `{}`, `null` — and for those the status code alone
tells the user nothing about what is wrong with what they typed.

Adds source-level coverage for both forms; the PR changed two JS files
with no test on either.
2026-09-30 17:15:52 +02:00
Léo 51e09a1e32 fix(css): repoint the two stylesheet links the split left behind
Splitting style.css deleted it, and two files outside static/ still
named it:

- scripts/verify_background_research_cards.mjs injected
  `<link rel="stylesheet" href="/static/style.css">` into the page it
  builds. A stylesheet that 404s does not fail — the script kept
  checking card layout against an unstyled page and kept reporting
  pass. It now reads the shell's <link> tags out of index.html, the way
  tests/css_snapshot/capture.mjs already does, so the set cannot drift
  out from under it again.
- tests/css_snapshot/bench.html's hand-open fallback linked the same
  deleted file. Replaced with the ordered set index.html ships.

test_every_stylesheet_referenced_by_shipped_html_exists only walked
static/*.html, which is why neither showed up. Extend it over the bench
page: the bench's whole output is computed styles, so a dead link there
is worth more than an unstyled page nobody looks at.
2026-09-30 17:14:52 +02:00
Léo 5ce2394adb fix(browser): probe pid liveness through the platform-safe helper
_process_is_alive used os.kill(pid, 0). That probe is POSIX-only:
CPython's Windows os.kill calls TerminateProcess(handle, sig) for any
signal other than CTRL_C/CTRL_BREAK, so it terminates the process it is
asked about. This function is only reached when there is no procfs to
read a command line from, which is exactly the macOS and Windows case
the rest of this change exists to handle.

core/platform_compat.py already owns that probe and documents the
hazard; its module docstring asks callers to import from there rather
than spell a POSIX-only call out locally. Delegate to it.

pid_alive answers False where os.kill raises PermissionError — a live
process owned by another user. Both call sites want that reading: the
sweep only unlinks a pid file it wrote itself, and a pid it cannot
confirm is not the daemon it is looking for.
2026-09-30 17:13:29 +02:00
Léo 5fd9114882 test(runtime): pin negative capability wording, test-only
First slice of the runtime regression lane. Drives stream_agent_loop with a
fake model and asserts on the tools the runtime offers, which is its decision
about what the turn may do. No production runtime code is touched and no
benchmark fixture or allowlist is imported.

The fake-model pattern is the one tests/test_tool_policy.py already uses:
patch stream_llm_with_fallback and inspect the tools kwarg.

Measured on lab@c499c01b, negative web wording is only partially detected:

  "Answer from memory only, don't search online."            web tools withheld
  "Summarise what you already know. Do not search the web."  web tools OFFERED
  "No web search please, just tell me what you know..."      web tools OFFERED

The case that holds is a plain regression guard. The two that do not are
xfail(strict=True): they run on every suite, document the target, and fail the
moment the behaviour lands so the marker gets removed rather than lingering.
A positive control keeps the guard from being satisfied by removing the web
tools altogether.
2026-09-30 17:11:10 +02:00
Léo 5e3e153fba fix(auth): swap the API token cache atomically instead of clearing it
`_refresh_token_cache` rebuilt the bearer-token map in two steps, `clear()`
then `update()`. Between them the dict a concurrent reader was already holding
was empty, so a valid token landing in that window found zero candidates and
got a 401. The refresh runs on a worker thread via `to_thread`, so the window
is real rather than theoretical.

Build the new map and rebind the name. The reader dereferences the global once
and then holds a map that is complete — the previous one if it read early, the
new one if it read late, never a half-built one. `app.state._token_cache` is
rebound with it, because it was bound once at startup and would otherwise point
at the abandoned dict.

Ported from public `dev` (`984337b3`), with its test.

The upstream test does not pin the fix: both of its concurrency cases pass
against the pre-fix code, because landing a GIL switch inside a window a few
bytecodes wide does not happen across 100 refreshes. They are kept as written
and `TestRefreshLeavesTheReadersMapAlone` is added next to them, stating the
same invariant at object level — a map a reader already holds is not mutated
by a later refresh — which fails on the pre-fix code without depending on
thread scheduling.
2026-09-30 16:07:12 +02:00
Léo 6e4b3aa5bd fix(tasks): clean up the singleflight cache on cancellation
`_cached` deduplicates the scheduler's outbound fetches — Miniflux unread
counts and MCP tool snapshots — by parking every concurrent caller on one
shared Future. Two cancellation paths left that Future stranded.

A waiter awaited the shared Future directly, so cancelling the waiter
cancelled the Future the owner and every other waiter were using. It now
awaits through `asyncio.shield`.

The owner removed its pending entry inside the success and `except Exception`
branches. `CancelledError` is a `BaseException`, so it took neither: the key
stayed in `_shared_cache_pending` pointing at a Future nobody would ever
resolve, and every later caller for that key waited forever. Cleanup moves to
a `finally` that is synchronous on purpose, and the owner cancels its own
Future so current waiters wake while a later caller can still retry.

Ported from public `dev` (`ce04dc1d`, #6174 upstream), with its test.
2026-09-30 16:05:09 +02:00
Léo 5c8611fba6 docs: generate the ODYSSEUS_* configuration reference from the source
The configuration surface was undiscoverable. .env.example has three active
lines, and of the ODYSSEUS_* variables the code actually reads, most appear
nowhere in .env.example, docs/, website/ or README.md - including several that
change security-relevant behaviour (ODYSSEUS_BROWSER_NO_SANDBOX,
ODYSSEUS_ALLOW_PRIVATE_CALDAV, ODYSSEUS_ENABLE_HOST_DOCKER,
ODYSSEUS_MCP_ALLOWED_COMMANDS). Every question about one of them lands in the
issue tracker.

A hand-written page would drift within a month, so the page is generated:

- scripts/generate_env_reference.py walks the Python sources, collects each read
  with the default it falls back to and the file it is read in, groups by area,
  and marks the internal variables rather than omitting them.
- website/configuration-reference.md is the generated output, wired into the
  Pages layout and linked from setup.md and .env.example. .env.example stays a
  short deployment-level example and links onward rather than growing.
- tests/test_env_reference.py regenerates and compares, so adding a variable
  without documenting it fails the suite. That is the point: the current state
  happened because nothing objected.

Finding the reads needs more than one pattern. Two families - the upload caps in
src/upload_limits.py and the media-ingress overrides in src/media_ingress.py -
are read through helper functions, so the generator detects env-reader helpers
rather than hardcoding a list. Others span two lines, hold the variable name in
a module constant, or read through a mapping passed in as an argument. One lives
inside a string literal, in the Ollama probe script routes/cookbook_helpers.py
builds line by line. A line-based grep for os.environ.get("ODYSSEUS_ finds 70 of
the 98 the generator finds; the page reports that gap and recomputes it on every
run so the claim cannot go stale.

No application behaviour changes.
2026-09-30 15:11:19 +02:00
Léo fba6f73260 test: fail the test that leaks a bare src/core module stub
#25 fixed two sys.modules writes in test_auth_regressions.py that left empty
stub modules behind for the rest of the session, breaking 23 tests under one
collection order while the full suite stayed green. The class is wider than
that file, and an audit is the wrong answer to it: nothing stops the next one,
and the failure it causes lands on an unrelated test in a different file.

So this is a guard instead. An autouse fixture in the root conftest snapshots
which src.* / core.* names are bound to a bare ModuleType, and fails any test
that adds one. "Bare" is the same test the clear_fake_* helpers already use -
a plain types.ModuleType with no on-disk __file__. MagicMock stand-ins are out
of scope: they answer every attribute, so they fail at the point of use rather
than silently, and several files install them deliberately.

Three details that matter:

- It lives in the root conftest, so it is set up before any test-module
  fixture and torn down after all of them. A stub a test's own teardown
  removes is not reported.
- It drops the leaked entries as well as reporting them, so the failure stays
  on the test that introduced it instead of cascading through the rest of the
  run.
- It only reports stubs added during the test. Import state the session starts
  with, including this conftest's own src.database stub, is left alone.

It found one beyond #25 on the first full run: _stub_heavy in
test_scheduler_restart_doublefire.py leaks the same five src.* modules as the
test #25 fixed, via sys.modules.setdefault. It already receives monkeypatch,
so the fix is to register through it. Fixed here because the guard has to land
green.

Full suite, macOS, default collection order:

  this branch   10655 passed, 6 failed, 6 skipped   406s
  lab           10655 passed, 6 failed, 6 skipped   371s

Same six either way, which is the point - none of this is visible in the
default order. Four are pre-existing macOS environment failures:
test_glob_confined_e2e and the two test_code_nav_tools document cases resolve
/tmp to /private/tmp, and
test_real_socket_falls_back_from_dead_first_to_live_second is connect-refused
timing on real sockets. The other two are the rich-colour and ffmpeg items
from the same ledger, fixed on their own branches.

Not verified: Linux, and any collection order other than the default. The
guard is order-independent by construction - it compares before and after
within a single test - but I have only run the default order.
2026-09-30 13:02:34 +02:00
Léo eb98aa6dc2 test(media): stop asserting an optional ffmpeg webp encoder
test_inspect_media_exports_final_decodable_frame_at_exact_duration exported to
/workspace/final.webp and asserted exit_code == 0. WebP encoding is an ffmpeg
build option, not something this project requires - Homebrew's macOS ffmpeg is
built without it:

  ffmpeg -encoders | grep -ic webp   ->  0

so the tool returns "ffmpeg still extraction failed: ... Encoder not found"
and the test fails on the build rather than on the code under test.

The test's subject is the final frame being decodable at the exact duration,
which has nothing to do with the container. It now writes a PNG, and the two
things the WebP path was implicitly covering are split out and each guarded on
what is actually present:

- test_inspect_media_exports_a_webp_still - skipped unless ffmpeg reports a
  webp encoder, and now asserts the file really is WebP rather than merely
  non-empty.
- test_inspect_media_reports_a_missing_encoder_instead_of_crashing - runs only
  where the encoder is absent, pinning the behaviour that surfaced this: the
  tool reports ffmpeg's failure as a tool error and writes no partial file.

The product code has the same assumption and I left it alone. inspect_media
accepts any suffix in _IMAGE_SUFFIXES and hands the path to ffmpeg, so a .webp
request on a build without libwebp fails with ffmpeg's own message. That is a
poor message, not a crash or a corrupt file, and pre-validating the encoder
list is a separate change.

  pytest tests/test_inspect_media_tool.py
  68 passed, 2 skipped in 22.92s      (1 failed, 66 passed, 1 skipped before)

Not verified: the WebP success path. This machine has no webp encoder, so
test_inspect_media_exports_a_webp_still skips here and has only been checked
for collection, not for a passing run.
2026-09-30 13:02:34 +02:00
Léo 24e428d9cb test(document): use platform-correct input in the rich color test
test_rich_colors_follow_theme_and_undo_as_one_edit fails identically on every
macOS run, timing out after 30s waiting for a span that never appears. It was
written off as timing noise twice. It is not flaky - it is two Linux-only
input conventions, and the product code is fine.

Control+click: macOS delivers a Control-modified primary click as contextmenu,
not click. Instrumenting the Lemon swatch shows the button receiving
pointerdown, mousedown, contextmenu, pointerup, mouseup - and no click, so the
menu item's handler never runs and no highlight is applied. Control was never
meaningful here anyway; the palette item has no modifier behaviour. The two
calls now use a plain click.

Control+Z: the editor's undo accelerator is Cmd+Z on macOS. With the clicks
fixed, both undo assertions still failed until the presses became
ControlOrMeta+Z, which Playwright maps per platform.

Both fixes are portable - a plain click and ControlOrMeta are unchanged on
Linux, where this test already passes.

  pytest tests/test_document_rich_color_reset_and_contrast.py
  2 passed in 2.11s       (1 failed, 1 passed before)

Not verified: Linux. I only have macOS here, so the claim that this stays
green on CI rests on the modifier being a no-op there, not on a run.
2026-09-30 13:02:34 +02:00
Léo 01b8ac5fea fix(mcp): reject malformed Args on Add MCP Server instead of defaulting to []
`add_server` wrapped `json.loads(args)` in a bare `except` that fell back to
`[]`, so an Args value that is not JSON — a bare path, which is what the form's
placeholder invites people to type — registered the server and spawned the
stdio subprocess with an empty argv. Nothing surfaced the loss: the POST
returned 200 and the row persisted with `"args": []`.

The route now returns 400 for an unparseable value, and also for valid JSON of
the wrong shape: `args=5` reached `StdioServerParameters(args=5)` and raised an
unhandled TypeError in the error formatter's `" ".join(...)`.

Both form clients mirror the guard instead of leaving the user to read a 400
they cannot see. `settings.js` stops silently defaulting a bad Args value, and
`admin.js` gains the same client-side parse check plus a `res.ok` branch so a
server-side rejection is not reported as a connection failure.

Ported from public `dev` (`9d5c0319`, #6215 upstream, fixing #6211), with its
test. The `admin.js` hunks are inert on `lab` — `initMcpForm` early-returns
because that form's markup is not in this build — and are carried anyway to
keep the two lines from diverging further.
2026-09-30 12:54:09 +02:00
Léo f8269a829f fix(discovery): cache a successful but empty Tailscale lookup
The host cache was gated on the list being non-empty, so "queried fine, no
eligible peers" looked exactly like a cold cache and every caller paid for
another `tailscale status --json` — a subprocess with a 5s timeout.

Gate on the timestamp instead. Failures still leave the timestamp unset, so a
missing binary, a non-zero exit or unparseable output stays retryable rather
than being cached for the full TTL.

Ported from public `dev` (`affaee1e`, #6228 upstream), with its test.
2026-09-30 12:14:03 +02:00
Léo 30ef7652c0 fix(cookbook): skip the pid sweep when the host has no procfs
`_cookbook_kill_session` kills the tmux session, then sweeps /proc for
model servers that survived the SIGHUP. The sweep had no guard, so on
macOS and Windows `os.listdir("/proc")` raised FileNotFoundError after
the kill had already succeeded. The function's outer except turned that
into `{"error": "...No such file or directory: '/proc'", "exit_code": 1}`
and skipped the state write that marks the session stopped — the agent
is told a stop failed that actually worked.

Guard the sweep with a procfs check, the way `_scan_running_model_processes`
already does a few hundred lines up. The root and the check now live in
`core/platform_compat`, which is where OS differences belong and which
makes both branches patchable from a test on either kind of host.

Adds a class-level guard test: the third instance of this defect, and two
of the three were found by reading source rather than by a test.
2026-09-30 12:13:04 +02:00
Léo 12f74ec9ea test(smoke): add a release smoke suite over every advertised feature area
The decomposition lanes have two safety nets and neither covers the
product. The checkpoint benchmark measures the agent runtime; the
computed-style snapshot pins the CSS. Nothing checked that Notes,
Calendar, Documents, Email, Memory, Cookbook or Settings still worked
after a route package moved or a 17,000-line module was split - and the
unit suite does not, since a byte-identical file move can break tests
that pass on the base branch with CI green throughout. The 28 Playwright
specs we do have are all under tests/e2e/photo-editor/ and no workflow
runs them.

scripts/odysseus-smoke boots this worktree through `odysseus dev` and
runs tests/smoke/: one scenario per area, each asserting a user-visible
outcome rather than a status code. Models come from a deterministic
OpenAI-compatible stub on an ephemeral loopback port; email reuses the
existing ODYSSEUS_EMAIL_FIXTURE path rather than inventing a second
mechanism. No scenario touches a live endpoint or the network.

The report is a per-area table that prints the areas the suite does not
cover next to the ones it does, and builds its rows from the registry
rather than from what happened to run, so an area cannot go missing by
having its module deleted or renamed. Under a plain pytest with nothing
booted every scenario skips with the reason, so the full suite stays
green.
2026-09-30 12:12:58 +02:00
Léo 10cb8fc669 chore(tools): add a read-only ref-to-ref parity audit
`lab` and the public `dev` line share only the repository's first commit as a
merge base, so `git log lab..dev` lists two thousand commits that are almost
all already present on both sides under different SHAs. Nothing in that output
says which public fixes never reached `lab`, which is the only question that
matters before `lab` becomes a release.

`scripts/ref_parity_audit.py` samples the most distinctive added lines from
each commit in the range and searches the other tree for them with
`git grep -F`, then reports the file-level presence diff. The two complement
each other: a commit whose probes are all found while one of the files it added
is missing from the target is a fix whose production change was reproduced
without its test, which the line sampling alone cannot see.

Probes are stripped of indentation and searched for anywhere in the tree, so a
port that moved or was re-indented still reads as present. Verdicts are
absent / partial / present / no-probe, and the report says outright that the
commit verdicts are a heuristic while the two file lists are exact.

Read-only by construction: `git log`, `show`, `diff`, `grep`, `ls-tree` and
`merge-base` only, no remote access, and it does not import the app package.
2026-09-30 11:53:31 +02:00
Léo 32d9dbc267 fix(tests): resolve temp paths consistently on macOS
Three of the six recorded failures were the same test bug: an unresolved
/tmp path compared against a resolved /private/tmp one. macOS makes /tmp a
symlink, so a fixture built with tempfile.mkdtemp(dir="/tmp") and a code path
that resolves what it reports disagree about a file both found correctly.

test_code_nav_tools builds its fixture unresolved and compares it against the
reported path. One realpath fixes both of its failures.

test_glob_confined_e2e is the same cause through a longer route: it mixed
os.path.realpath(ws) with an unresolved secret directory, so relpath emitted
"../../../../tmp/<absolute path>" and the assertion that the absolute path was
absent from the output matched it as a substring. Resolving the secret
directory puts both sides in one tree and the relative path stays short.

macOS full suite goes from 6 failures to 3. The remaining three are an ffmpeg
build without a WebP encoder, a socket test that needs a fast connection
refusal, and the rich-text colour test that is still unexplained.

The ledger is updated in the same change so it does not describe failures that
no longer happen.
2026-09-30 10:56:58 +02:00
Léo eabdf84669 fix(tests): stop test_auth_regressions leaking stub modules
Running test_auth_regressions.py before the email modules failed 23 tests
that pass in isolation:

  pytest -p no:randomly tests/test_auth_regressions.py \
                        tests/test_email_urgency_checkpoint.py
  23 failed, 15 passed       (38 passed in the reverse order)

Every failure was ImportError "cannot import name X (unknown location)"
against a module already in sys.modules, which is what an empty stub module
looks like to a later import.

Two writes leaked, both in this file:

- test_pop_notifications_owner_filtered inserted five empty stub modules with
  a bare sys.modules[name] = mod and never removed them.
- _ensure_stub wrote its stub into sys.modules itself, so the autouse
  fixture's monkeypatch.setitem three lines later captured that stub as the
  value to restore. The fixture looked like it cleaned up and could not.

Both now go through monkeypatch, including the parent-package stub and the
attribute wiring, so everything is undone at teardown. The redundant setitem
calls in the fixture are gone: re-setting a key whose stub is already
installed is what made the leak invisible.

Pre-existing on lab, not introduced by any open PR. Verified against the
email subpackage branch too: identical numbers there before this fix.
2026-09-30 10:36:00 +02:00
Léo b7002c0fe3 fix(email): import spinnerModule in the modules that use it
Six of the seven modules call `spinnerModule.createWhirlpool(...)` and none
of them imported it. Everything still parsed, every module still loaded, and
the Email Settings page threw `ReferenceError: spinnerModule is not defined`
the moment `settings.js` mounted it — the panel rendered its loading spinner
and stopped there.

The cause is worth writing down because it is the same shape as the bug: the
import statements were collected with a pattern that was not anchored to the
start of a line, and the header comment says "the old import path keeps
resolving". That prose matched first and swallowed the real
`import spinnerModule from '../spinner.js';` that followed it.

`test_no_module_uses_a_package_name_it_never_bound` closes the class. Nothing
else here can: `node --check` parses without resolving scope, and loading a
module does not run the function body where the throw lives. It checks the
package's own vocabulary — every name any module in it binds — rather than
trying to model the browser's globals, so it has no false positives and still
catches the one mistake a split actually makes: the declaration stays behind
and the use moves.
2026-09-30 09:55:39 +02:00
Léo 1c5f60539f refactor(settings): move the shell out of settings.js into settings/
settings.js is 5,721 lines and the registry/navigation/search/sidebar/
lifecycle primitives already live in static/js/settings/. What was left
behind in the coordinator was the layer above them: what happens when a
panel becomes active, where an admin-managed tab is handed to admin.js,
which elements are admin-only, the Appearance window fade, and the
public open/close. That layer reached module-global `modalEl` and
`initialized` directly, so none of it could be exercised without
booting every panel in the file — and every panel I eventually move out
would have to route back through it.

Three modules, no behavior change:

  shell.js       panel-activation side effects, the admin handoff,
                 .admin-only visibility, open()/close(). Takes what it
                 needs from settings.js as injected callbacks, the same
                 shape bindSettingsNavigation() already uses, so it
                 holds no panel state.
  peek.js        the Appearance window fade and its toggle. It is window
                 chrome rather than Appearance panel data, and it has to
                 be cleared when the user leaves that panel.
  oauthReturn.js the once-per-load return path from the Google OAuth
                 redirect. It was an IIFE running at module evaluation
                 in the middle of a 5,700-line file.

settings.js keeps open/close/syncAdminVisibility as exports, so every
caller (app.js, calendar.js, chatStream.js, gallery.js, admin.js,
modelPicker.js, slashCommands.js, chatRenderer.js, emailLibrary.js) is
untouched. 5,721 -> 5,583 lines; the settings/ modules go 887 -> 1,107.

The real-ESM coordinator smoke now links the three new files and asserts
what moved: admin-only elements hidden for a non-admin and shown for an
admin, an admin-managed tab click handed to admin.js without a second
local activation, and the Peek fade applying on Appearance and clearing
when the user navigates away. The OAuth test follows its handler to the
new file and additionally pins the coordinator wiring, since "uses the
module-local open()" is now a property of the seam rather than of one
source slice.

No new module needs a cache-busting query or an sw.js precache entry:
the existing settings/ submodules have neither, they load transitively
from settings.js's versioned URL, and sw.js serves JS network-first.

admin.js stays where it is. It has no shell to extract — open() and
close() already delegate to settingsModule, and its 4,122 lines are all
panel code. That is a panel split, not this one.
2026-09-30 09:45:23 +02:00
Léo 345ce0a9ec test(document): read the editor through a module-set helper before it is split
static/js/document.js is 17,579 lines and is about to be decomposed behind a
re-export wrapper. 51 test files read it off disk and grep it as text, and 28
of those slice it with `src.split("function a", 1)[1].split("function b", 1)[0]`
-- "the region between a and b", which only means what the test intends while a
and b are neighbours in one file. Several also hard-code the file's two-space
indentation, which no extracted module reproduces. Left alone, the first
extraction makes those assertions cover the wrong region, and an `x in region`
check passes while covering more than it was written for.

tests/helpers/document_source is the one place that names the file now:

  - document_source() is the entry plus everything under static/js/document/,
    so a membership assertion keeps finding its subject wherever it lands;
  - function_body()/declaration() locate a construct by name in whichever
    module defines it and end at its real closing brace, so neither moving it
    nor moving its neighbour changes the region.

The rewrite only collapses a slice when the old terminator sat at the
construct's end. 28 slices deliberately span a whole family of functions --
everything from _docxHexColor to exportAsDocx -- and collapsing one to its
first member drops what the assertions look for, so those stay as they are and
are listed in KNOWN_ADJACENCY_SLICES, to be converted as each family becomes a
module. That list may only shrink.

Two guards come with it:

  - test_document_source_test_hygiene fails on a direct read of the entry file
    and on any new adjacency slice;
  - test_frontend_module_graph resolves every relative import under static/
    (718 of them, none broken today) and requires the document module set to
    stay in the sw.js precache, since the worker fetches the URLs it lists and
    not what they import.

test_document_module_api pins the 38 default-export keys and 29 named exports
by loading the module in a browser and reading what it actually exports, rather
than grepping for the literal object -- after extraction that object may be
assembled from imports, and a source-shape check would pass while the export
was broken.

No JavaScript moves here. static/ is untouched.
2026-09-30 09:39:05 +02:00
Léo c5db6ae7fd refactor(email): split the email library into seven modules
`emailLibrary/index.js` goes from 11,375 lines to 6,187. What came out:

- `settingsPage.js` (879) — the Email Settings view, its form and controls,
  away/auto-reply including the calendar-event sync, and the display
  preferences. The inline-image preference lives here rather than with the
  renderer because the settings form owns writing it.
- `unsubscribe.js` (1,224) — the bulk-unsubscribe review flow. One exported
  entry point, its own localStorage keys, its own agent-tool-output listeners.
- `reader.js` (676) — opening an email as a docked tab or a floating window,
  plus the AI summary panel both of them share.
- `menus.js` (855) — the reader More menu, the card kebab menu, the bulk
  Actions menu and `_bulkAction`.
- `bodyRender.js` (871) — plain and threaded body rendering, inline MIME
  images, quote folding.
- `attachments.js` (491) — attachment chips and the deferred load for messages
  whose attachment list was not in the list response.
- `aiReply.js` (449) — AI-reply entry points, the per-message context draft,
  the translate and remind submenus.

The extracted modules import back from `index.js`, so the graph has cycles.
That is safe for hoisted function declarations and unsafe for a value read
during evaluation, so nothing crosses a module boundary except functions:
`API_BASE` is re-declared per module, the way `emailInbox.js` and
`emailShared.js` already do it, and `_autoReplyRefreshSeq` moves into
`state.js` because the settings page and the unread-badge refresh both write
it and an imported binding is read-only. The module-graph test enters the
package at each module in turn, which is the order that would expose a
dead-zone read.

Two test helpers grew while doing this. `js_function_source` replaces three
marker-pair slices ("from this signature down to that comment") whose end
marker had moved into another module — the slice ran past the function and
kept passing against the wrong text. Its first implementation balanced braces
by walking characters and ended `_toggleCardPreview` 18 lines early, because
the apostrophe in `// that's a scroll, not a nav` opened a string that ate the
braces after it. It now keys on the invariant the file actually holds: a
top-level declaration starts at column 0, so its closing brace is the next
lone `}` at column 0.
2026-09-30 09:26:34 +02:00
Léo ab5a08a6f9 refactor(email): move emailLibrary.js into static/js/emailLibrary/
The email library was 11,375 lines in one file, the second-largest JS
module in the repo. `static/js/emailLibrary/` already held four extracted
helpers, so the package existed; the bulk of the code just was not in it.

The implementation moves to `emailLibrary/index.js` and the old path
becomes a re-export wrapper. Five call sites import that path, four of
them dynamically with a `?v=` string, and `sw.js` caches URLs verbatim,
so a wrapper is what makes the move need no coordinated edit to any of
them.

The test side is the part worth reviewing. 58 tests read
`static/js/emailLibrary.js` as text. Pointing them at
`emailLibrary/index.js` would buy one move and break again on the next
one, which is exactly what happened to the stylesheet tests (they now go
through `tests/helpers/stylesheets.py`). So the same shape:
`tests/helpers/js_modules.py` reads the whole package, and assertions
stop caring which module a function sits in.

`tests/test_email_library_module_graph_js.py` is new. This frontend has
no module-graph validation, and a package fails in ways a single file
cannot: a wrapper that drops an export is `undefined` at call time rather
than an error at load time, and a module that reads a `const` across an
import cycle throws only when that module is entered first. It pins the
wrapper's surface against the entry module's, evaluates every module on
its own in a browser, and requires both import paths to hand out one
instance.

It also caught a live gap while being written:
`emailLibrary/replyRecipients.js` is imported by `emailInbox.js`, an app
shell module, and was never in the `sw.js` precache.
2026-09-30 09:08:21 +02:00
Léo 8a95ec9299 refactor(settings): extract four panels into settings/ modules
static/js/settings.js was 5,721 lines behind a four-name public surface.
static/js/settings/ already existed with dom, registry, search, sidebar,
navigation and lifecycle, so this continues that package rather than
inventing a layout.

Moved verbatim, 407 lines:

  settings/speech.js        initTtsSettings, initSttSettings
  settings/writingStyle.js  initDocumentWritingStyle
  settings/imageModels.js   initImageSettings
  settings/agent.js         initAgentSettings
  settings/api.js           postSettings, lifted from _postSettings

These five were chosen because their only dependencies outside themselves
were el/byId, _postSettings and sortModelIds. Panels with wider reach stay
put: initEmailAccountsSettings, for instance, pulls a 64-declaration closure
covering most of the file, and splitting that is a design change rather than
a move.

settings.js keeps its public surface exactly: open, close,
refreshAiModelEndpoints and the default export, verified in a browser.

Three test updates the move required:

- tests/helpers/test_settings_shell_coordinator.mjs allowlists the real
  modules the coordinator may import, and rejected the new ones. That guard
  working is the reason to trust the rest of this diff.
- two source-introspection tests read settings.js for behaviour that now
  lives in a panel module. They read the whole settings surface now, so the
  next extraction does not break them again.
2026-09-30 09:04:20 +02:00
Alexandre Teixeira 3748a621d1 Merge current lab into agent runtime decomposition 2026-09-29 19:15:54 +01:00
Léo 4b4ae592da refactor(routes): move the email modules into routes/email/
routes/email_routes.py (7,194 lines), routes/email_helpers.py (2,025) and
routes/email_pollers.py (1,764) were the largest flat email area left in
routes/. They now live in routes/email/, following the pattern fifteen route
areas already use: the canonical module under the package, a shim at the old
path that replaces itself in sys.modules with the canonical object.

The shim matters more here than the move. Dozens of email tests monkeypatch
module attributes, and several pop the module from sys.modules and re-import
it. Without identity between routes.email_routes and routes.email.email_routes
those patches would apply to a different object than the app uses.

Intra-package imports keep the flat path, matching routes/document/ and
routes/gallery/: the shim resolves them and the diff stays a move.

Eight path references in five test modules read these files as source rather
than importing them, so they now read the canonical location. That is the same
adjustment every previous conversion made, and it is why the note/document
shims mention source-introspection tests explicitly.

No behaviour change. Verified with the full suite and under several explicit
collection orders, since a byte-identical routes move has broken tests under
one ordering before.
2026-09-29 18:21:18 +02:00
Léo b5d1505582 docs(tests): record the known full-suite failures
specs/testing-devops.md lists "no canonical full-suite known-failing/flaky
ledger" as a gap. Without one a first local run is uninterpretable: you cannot
tell a regression from a platform artifact, so you either chase a non-bug or
ignore a real one.

Six failures on macOS against lab@c499c01b, each with its cause and a verdict
rather than a blanket "environmental":

- three compare an unresolved /tmp path against a resolved /private/tmp one.
  Those are test bugs and the file says so.
- one asserts ffmpeg exit 0 for a .webp still, which is a build option Homebrew
  does not always carry. Needs a skip or a PNG fallback.
- one opens real sockets and needs a fast connection refusal. Environmental.
- one Playwright colour-contrast test had been written off as a flake. It is
  not: three consecutive runs failed identically at ~31s. Recorded as
  unexplained and possibly a real defect, because calling it noise is what
  stopped anyone looking.

Also documents the prerequisites, since most surprise failures are a missing
npm ci rather than anything here, and the CHROMADB_PORT precaution: the client
reaches Chroma over HTTP regardless of the data directory, so a test run can
attach to a store holding real data.
2026-09-29 18:00:12 +02:00
Léo a4715b77a5 refactor(css): name the fragments by area instead of by number
style-part-01..08 were arbitrary 6,000-line cuts, so the names said nothing
and a reader had no way to guess which file held a rule.

The original stylesheet turns out to be roughly area-grouped already, so
cutting where the content changes rather than every 6,000 lines produces
boundaries worth naming. Classifying every top-level rule by selector prefix
and smoothing over a 40-rule window gives 17 stable runs: agent-chat, compare,
memory, documents, admin-settings, skills, gallery, cookbook, tasks,
image-editor, email, notes, calendar, research.

The numeric prefix stays because load order is load-bearing, and it also
disambiguates the areas that appear twice: the original interleaves, and
merging non-adjacent blocks of the same area would reorder the cascade.

The names describe where a file sits, not a claim that it holds every rule for
that area or only rules for it. The header in each file says so, because that
is the misreading this naming invites.

Reconstruction was checked against lab's style.css before rewriting: identical
after whitespace normalisation, all 47,528 lines accounted for. The computed
style baseline still reproduces exactly.
2026-09-29 17:50:04 +02:00
Léo 03d55dfe3c refactor(css): split style.css into ordered fragments
static/style.css was 47,530 lines. It is now nine files under static/css/:
tokens.css holds the :root custom properties, light theme and density classes,
and style-part-01..08 carry the rest in their original order. Nothing was
rewritten; every line moved verbatim and the <link> order in index.html
reproduces the original file byte for byte.

Cut points are brace-depth zero and outside block comments, so no rule or
comment is split. All 47,531 source lines are accounted for across the
fragments.

The computed-style baseline recorded before the split is reproduced exactly,
which is the evidence that the cascade is unchanged rather than an argument
that it should be.

Three things the split broke and this fixes:

- The harness self-test swapped two conflicting .attach-strip blocks inside
  style.css to prove the digest is order-sensitive. It rewrote one hardcoded
  URL, so with the file gone it silently measured an unmodified page. It now
  searches every stylesheet, rewrites whichever holds the pair, and fails
  loudly if none does.
- Four tests asserted the cache-bust token by matching /static/style.css?v=.
  They check the real invariant now, that every app stylesheet shares one
  token with app.js, through stylesheet_cache_version().
- The three panel stylesheets carried their own tokens, so a browser could
  hold half an old cascade and half a new one. All app stylesheets now bust
  together.
2026-09-29 17:37:25 +02:00
Léo 1323fab0bd refactor(tests): read the whole cascade instead of style.css alone
static/style.css no longer holds every rule. 79 rules for the document,
gallery and editor panels now live in static/css/, yet 27 browser tests still
built their synthetic page with a single <link> to style.css, and 51 more read
that one file as though it were the whole cascade. Those tests kept passing
while covering less: a rule that moved became invisible to the assertion that
was meant to pin it.

Python readers now call tests.helpers.stylesheets.app_css(), and synthetic
pages are built from stylesheet_link_tags() so they load exactly what
index.html loads, in the same order. The helper already existed; this moves
the remaining callers onto it.

test_portal_dropdown_z_js parametrised over a file list including style.css to
assert an absence. Checking a negative against one file of a split stylesheet
is how a moved rule escapes, so the CSS case now checks the concatenation.

Two guards keep it from coming back: one fails on any test reading
static/style.css directly, the other on any synthetic page linking it alone.
Both name the helper to use.

No production code changes. Suite is unchanged at 6 pre-existing failures.
2026-09-29 17:05:29 +02:00
Léo 7a732989fd fix(css): collapse duplicate @keyframes names to the definition that wins
Duplicate @keyframes names resolve last-wins across the whole cascade, so
every definition but the last was dead code that still read as live at its
call site. Two of the five duplicated names differed from the winner:

  fadeIn          style.css had an opacity-only variant before the one that
                  adds translateY, so every consumer was already sliding
  research-pulse  a background-colour pulse sat before the opacity/scale one
                  that actually runs

The other three (spin, loading-bounce, pulse) were byte-identical repeats.

Removing the losing definitions changes nothing rendered, which the computed
style snapshot confirms: the committed baseline is reproduced exactly across
all three pages and 24 variants. Deciding that a consumer wanted the plain
fade, or the background pulse, would be a visual change and belongs in its own
PR with screenshots.

This also unblocks the mechanical stylesheet split. While two definitions of a
name differed, neither could be moved: relocating either changes which one is
last, and therefore changes behaviour. A regression test now fails on any
duplicate name so the trap cannot come back.
2026-09-29 16:36:01 +02:00
Léo 6cda92080c fix(browser): stop treating a missing cmdline as proof the daemon exited
_terminate_owned_daemon() read /proc/<pid>/cmdline and, on FileNotFoundError,
unlinked the pid file on the stated assumption that "the daemon may have
exited". Off Linux that file is always missing, so the branch always fired:
the pid file of a live daemon was deleted and the daemon itself never killed.
_owned_daemon_exists() swallowed the same error and therefore always returned
False, which is precisely the state its own docstring warns about, since a
close against an unrecognised session can bootstrap a fresh daemon and wait on
its browser indefinitely.

Demonstrated on macOS before the change: a pid file holding a live pid is
removed by _terminate_owned_daemon() and _owned_daemon_exists() reports False.
After it, the file survives and the daemon is reported present.

_process_command_line() now returns None for "this host cannot tell" and
_process_is_alive() answers the separate question of whether the pid exists.
Without procfs we decline to kill a process we cannot confirm is ours, and we
only forget a pid file once the pid is genuinely gone. Linux behaviour is
unchanged: the command-line identity check still gates both paths.
2026-09-29 16:27:22 +02:00
Alexandre Teixeira 6105702901 Merge pull request #16 from o3LL/lane/decomposition-and-test-isolation
refactor(lane): decompose style.css, pin it with a computed-style harness, and isolate worktree runs
2026-09-29 14:54:36 +01:00
Alexandre Teixeira 65c35c8211 Merge pull request #15 from o3LL/fix/test-static-port-isolation
test(harness): bind the static test server to an ephemeral port
2026-09-29 14:54:18 +01:00
Alexandre Teixeira 956671a731 Merge pull request #14 from o3LL/fix/proc-guard-private-browser
fix(browser): skip the Chrome sweep when the host has no procfs
2026-09-29 14:54:00 +01:00
Alexandre Teixeira aafd8a98eb Merge pull request #13 from o3LL/v1/css-split-3
refactor(css): move cookbook, research, memory and settings styles to their own file
2026-09-29 14:53:44 +01:00
Léo 4318974903 test(css): read the static origin at call time, not at import
The snapshot harness read ODYSSEUS_TEST_STATIC_ORIGIN into a module constant.
That env var is published by the session static-server fixture, which runs
after collection has already imported the module, so the constant always held
the 7011 fallback and the capture connected to a port nothing was listening on:

    route.fetch: connect ECONNREFUSED 127.0.0.1:7011

Neither branch is wrong on its own. The harness was written while the fixture
still bound a fixed 7011, and the ephemeral-port change removed that port. The
two only disagree once they are in the same tree, which is what this integration
branch is for.

With this, the baseline recorded before the stylesheet split is reproduced
exactly after it, so the split is confirmed to preserve computed styles rather
than only argued to.
2026-09-29 10:38:06 +02:00
Léo 822f7d80e4 Merge branch 'feat/odysseus-dev-worktree-boot' into lane/decomposition-and-test-isolation 2026-09-29 10:36:18 +02:00
Léo ac5e1a600d Merge branch 'fix/test-static-port-isolation' into lane/decomposition-and-test-isolation 2026-09-29 10:36:18 +02:00
Léo f8e9d00256 Merge branch 'fix/proc-guard-private-browser' into lane/decomposition-and-test-isolation 2026-09-29 10:36:18 +02:00
Léo 16c0fe0975 Merge remote-tracking branch 'fork/v1/css-split-3' into lane/decomposition-and-test-isolation 2026-09-29 10:36:18 +02:00
Alexandre Teixeira cea8ed297e fix(runtime): reject unobserved verification and preserve safe reasoning 2026-09-26 14:24:29 +01:00
Alexandre Teixeira b241bb3a7b fix(runtime): close completion stream bypass and preserve explanations 2026-09-26 14:13:40 +01:00
Alexandre Teixeira eaa5668baa feat(runtime): record execution evidence and gate completion 2026-09-26 13:48:12 +01:00
Alexandre Teixeira ea847e6c0b test(runtime): establish isolated validation and comparison gates 2026-09-26 13:13:07 +01:00
Léo 255f7c46c1 feat(scripts): add odysseus-dev, a per-worktree isolated boot
start-macos.sh is the single-instance launcher and adopts whatever is
already listening: an open ChromaDB port is a resource it reuses. With
one checkout that is right. With several, it means a scratch worktree
silently attaching to another checkout's vector store, and the script
reports it as a success.

odysseus-dev is the sibling that owns isolation instead. Ports are
derived from the worktree path, so two checkouts never collide and one
checkout always gets the same URL. A ChromaDB this worktree did not
start is refused, never adopted — we start our own or fall closed to
keyword mode and say which. The data dir, database and browser-MCP
cache live under .odysseus-dev/, leaving data/ to a normal launch. A
checkout wired into launchd or systemd will not boot at all, and the
ports the project already means something by (7000, 7011, 7860, 8100)
are refused even when asked for explicitly.

Readiness is /api/ready rather than a TCP accept: the port accepting
connections says nothing about the database or a writable data dir.
That endpoint is not auth-exempt, so the tool owns a dev admin account,
hands it to setup.py and prints it.

--from-pr N fetches pull/N/head into its own worktree and boots it,
borrowing a venv so a PR is a few seconds rather than a pip install.

start-macos.sh is untouched: it is what the LaunchAgent runs.
2026-09-25 17:35:47 +02:00