Own each agent-browser session as a browser tree: the daemon's POSIX
session, its runtime files and its Chrome profile. Timeouts, launch
failures, bootstrap recovery, cancellation and shutdown clean that tree
and verify nothing survives, instead of killing only the daemon and
orphaning Chrome. Per-call cleanup no longer sweeps every Chrome under
the runtime TMPDIR.
Sessionless calls get an ephemeral browser closed before returning.
Actions on one session are serialized. Recovery is bounded by one
deadline with at most one retry for local HTML open, and the retry flag
is no longer model-visible. Observations after a failed navigation are
marked stale. read URL navigates and extracts in one batch because
agent-browser has no read command. Results carry a browser_lifecycle
receipt with stages, timings, ownership and cleanup evidence.
Playwright MCP tool calls are bounded by
ODYSSEUS_BROWSER_MCP_CALL_TIMEOUT_S and are not retried. research_navigator
now passes timeout_ms.
Snapshot current maintainer-preview application changes and regression fixtures for integration into lab. Excludes local runtime data, evaluation outputs and source backups. Focused Python regression selection: 140 passed; full suite not certified.
The move left three documents naming `routes/email_routes.py`,
`routes/email_helpers.py` and `routes/email_pollers.py` as where the
code is. The shims keep those import paths working, so nothing breaks —
but each of those files is now seventeen lines that redirect, and a
reader sent there finds no email code at all.
Follows the phrasing specs/persistence.md already uses for the
contacts and vault subpackages: name the canonical path and note the
shims.
Splitting style.css deleted it, and two files outside static/ still
named it:
- scripts/verify_background_research_cards.mjs injected
`<link rel="stylesheet" href="/static/style.css">` into the page it
builds. A stylesheet that 404s does not fail — the script kept
checking card layout against an unstyled page and kept reporting
pass. It now reads the shell's <link> tags out of index.html, the way
tests/css_snapshot/capture.mjs already does, so the set cannot drift
out from under it again.
- tests/css_snapshot/bench.html's hand-open fallback linked the same
deleted file. Replaced with the ordered set index.html ships.
test_every_stylesheet_referenced_by_shipped_html_exists only walked
static/*.html, which is why neither showed up. Extend it over the bench
page: the bench's whole output is computed styles, so a dead link there
is worth more than an unstyled page nobody looks at.
The configuration surface was undiscoverable. .env.example has three active
lines, and of the ODYSSEUS_* variables the code actually reads, most appear
nowhere in .env.example, docs/, website/ or README.md - including several that
change security-relevant behaviour (ODYSSEUS_BROWSER_NO_SANDBOX,
ODYSSEUS_ALLOW_PRIVATE_CALDAV, ODYSSEUS_ENABLE_HOST_DOCKER,
ODYSSEUS_MCP_ALLOWED_COMMANDS). Every question about one of them lands in the
issue tracker.
A hand-written page would drift within a month, so the page is generated:
- scripts/generate_env_reference.py walks the Python sources, collects each read
with the default it falls back to and the file it is read in, groups by area,
and marks the internal variables rather than omitting them.
- website/configuration-reference.md is the generated output, wired into the
Pages layout and linked from setup.md and .env.example. .env.example stays a
short deployment-level example and links onward rather than growing.
- tests/test_env_reference.py regenerates and compares, so adding a variable
without documenting it fails the suite. That is the point: the current state
happened because nothing objected.
Finding the reads needs more than one pattern. Two families - the upload caps in
src/upload_limits.py and the media-ingress overrides in src/media_ingress.py -
are read through helper functions, so the generator detects env-reader helpers
rather than hardcoding a list. Others span two lines, hold the variable name in
a module constant, or read through a mapping passed in as an argument. One lives
inside a string literal, in the Ollama probe script routes/cookbook_helpers.py
builds line by line. A line-based grep for os.environ.get("ODYSSEUS_ finds 70 of
the 98 the generator finds; the page reports that gap and recomputes it on every
run so the claim cannot go stale.
No application behaviour changes.
The decomposition lanes have two safety nets and neither covers the
product. The checkpoint benchmark measures the agent runtime; the
computed-style snapshot pins the CSS. Nothing checked that Notes,
Calendar, Documents, Email, Memory, Cookbook or Settings still worked
after a route package moved or a 17,000-line module was split - and the
unit suite does not, since a byte-identical file move can break tests
that pass on the base branch with CI green throughout. The 28 Playwright
specs we do have are all under tests/e2e/photo-editor/ and no workflow
runs them.
scripts/odysseus-smoke boots this worktree through `odysseus dev` and
runs tests/smoke/: one scenario per area, each asserting a user-visible
outcome rather than a status code. Models come from a deterministic
OpenAI-compatible stub on an ephemeral loopback port; email reuses the
existing ODYSSEUS_EMAIL_FIXTURE path rather than inventing a second
mechanism. No scenario touches a live endpoint or the network.
The report is a per-area table that prints the areas the suite does not
cover next to the ones it does, and builds its rows from the registry
rather than from what happened to run, so an area cannot go missing by
having its module deleted or renamed. Under a plain pytest with nothing
booted every scenario skips with the reason, so the full suite stays
green.
`lab` and the public `dev` line share only the repository's first commit as a
merge base, so `git log lab..dev` lists two thousand commits that are almost
all already present on both sides under different SHAs. Nothing in that output
says which public fixes never reached `lab`, which is the only question that
matters before `lab` becomes a release.
`scripts/ref_parity_audit.py` samples the most distinctive added lines from
each commit in the range and searches the other tree for them with
`git grep -F`, then reports the file-level presence diff. The two complement
each other: a commit whose probes are all found while one of the files it added
is missing from the target is a fix whose production change was reproduced
without its test, which the line sampling alone cannot see.
Probes are stripped of indentation and searched for anywhere in the tree, so a
port that moved or was re-indented still reads as present. Verdicts are
absent / partial / present / no-probe, and the report says outright that the
commit verdicts are a heuristic while the two file lists are exact.
Read-only by construction: `git log`, `show`, `diff`, `grep`, `ls-tree` and
`merge-base` only, no remote access, and it does not import the app package.
start-macos.sh is the single-instance launcher and adopts whatever is
already listening: an open ChromaDB port is a resource it reuses. With
one checkout that is right. With several, it means a scratch worktree
silently attaching to another checkout's vector store, and the script
reports it as a success.
odysseus-dev is the sibling that owns isolation instead. Ports are
derived from the worktree path, so two checkouts never collide and one
checkout always gets the same URL. A ChromaDB this worktree did not
start is refused, never adopted — we start our own or fall closed to
keyword mode and say which. The data dir, database and browser-MCP
cache live under .odysseus-dev/, leaving data/ to a normal launch. A
checkout wired into launchd or systemd will not boot at all, and the
ports the project already means something by (7000, 7011, 7860, 8100)
are refused even when asked for explicitly.
Readiness is /api/ready rather than a TCP accept: the port accepting
connections says nothing about the database or a writable data dir.
That endpoint is not auth-exempt, so the tool owns a dev admin account,
hands it to setup.py and prints it.
--from-pr N fetches pull/N/head into its own worktree and boots it,
borrowing a venv so a PR is a few seconds rather than a pip install.
start-macos.sh is untouched: it is what the LaunchAgent runs.
static/style.css is 51,425 lines in one file. Hundreds of selectors are
declared more than once and !important is used throughout, so the rendered
result is a function of source order. Extracting a block into its own file
changes that order, and nothing in the suite would notice - which makes a
51k-line split unfalsifiable and "looks fine to me" the only available
evidence.
This moves no CSS. It captures getComputedStyle over a fixed inventory of
676 elements across three pages, four viewports, both themes and the three
density modes - 16,224 element snapshots - hashes them, and compares against
tests/css_snapshot/baseline.json. A capture takes about 21 seconds.
The bench page synthesises one element per selector from an evidence-driven
list: every selector declared more than once in style.css that can be
expressed as a static compound chain, plus a curated set per feature area.
Redeclared selectors are the ones a reorder can flip. The bench loads
whatever stylesheets index.html ships, so it keeps measuring the real set
once the file is split. tests/test_css_computed_style_snapshot.py also carries
a self-test that swaps two conflicting .attach-strip declarations and asserts
the digest moves, so the harness cannot silently stop watching.
The second half is the asset-manifest check specs/frontend.md asks for,
scoped to stylesheets: every stylesheet referenced by shipped HTML and by the
sw.js precache exists, and index.html and sw.js agree on the ?v= string. They
hardcode it independently today, so a split that updates one and not the
other ships an offline cache nobody notices until a plane.
scripts/mlx_image_server.py resolved the model per request
(`req.model or _args.model`) on both /v1/images/generations and
/v1/images/edits, so the caller chose which model was served.
`_is_hidream()` is a substring test and `_snapshot_path()` accepts either a
local directory or a Hugging Face repo id, so a caller-supplied string
selected the HiDream branch and then supplied the directory it runs
`scripts/hidream_o1/generate_hidream_o1_mlx.py` from, under sys.executable.
The server has no auth, and the Cookbook binds it to 0.0.0.0 whenever it is
serving to a remote host, so one POST executed attacker code on the serving
host.
Both paths now use `_args.model`. The request field is still accepted for
OpenAI wire compatibility and ignored, matching scripts/diffusion_server.py,
and Odysseus already sends the served model's own id, so this is a no-op for
legitimate callers. /v1/images/harmonize already pinned.
Regression tests cover both endpoints, the local-directory and
Hugging-Face-repo halves, and that a server actually launched with a HiDream
model still serves it. Three of the four fail on the unfixed code.