#25 fixed two sys.modules writes in test_auth_regressions.py that left empty
stub modules behind for the rest of the session, breaking 23 tests under one
collection order while the full suite stayed green. The class is wider than
that file, and an audit is the wrong answer to it: nothing stops the next one,
and the failure it causes lands on an unrelated test in a different file.
So this is a guard instead. An autouse fixture in the root conftest snapshots
which src.* / core.* names are bound to a bare ModuleType, and fails any test
that adds one. "Bare" is the same test the clear_fake_* helpers already use -
a plain types.ModuleType with no on-disk __file__. MagicMock stand-ins are out
of scope: they answer every attribute, so they fail at the point of use rather
than silently, and several files install them deliberately.
Three details that matter:
- It lives in the root conftest, so it is set up before any test-module
fixture and torn down after all of them. A stub a test's own teardown
removes is not reported.
- It drops the leaked entries as well as reporting them, so the failure stays
on the test that introduced it instead of cascading through the rest of the
run.
- It only reports stubs added during the test. Import state the session starts
with, including this conftest's own src.database stub, is left alone.
It found one beyond #25 on the first full run: _stub_heavy in
test_scheduler_restart_doublefire.py leaks the same five src.* modules as the
test #25 fixed, via sys.modules.setdefault. It already receives monkeypatch,
so the fix is to register through it. Fixed here because the guard has to land
green.
Full suite, macOS, default collection order:
this branch 10655 passed, 6 failed, 6 skipped 406s
lab 10655 passed, 6 failed, 6 skipped 371s
Same six either way, which is the point - none of this is visible in the
default order. Four are pre-existing macOS environment failures:
test_glob_confined_e2e and the two test_code_nav_tools document cases resolve
/tmp to /private/tmp, and
test_real_socket_falls_back_from_dead_first_to_live_second is connect-refused
timing on real sockets. The other two are the rich-colour and ffmpeg items
from the same ledger, fixed on their own branches.
Not verified: Linux, and any collection order other than the default. The
guard is order-independent by construction - it compares before and after
within a single test - but I have only run the default order.
test_inspect_media_exports_final_decodable_frame_at_exact_duration exported to
/workspace/final.webp and asserted exit_code == 0. WebP encoding is an ffmpeg
build option, not something this project requires - Homebrew's macOS ffmpeg is
built without it:
ffmpeg -encoders | grep -ic webp -> 0
so the tool returns "ffmpeg still extraction failed: ... Encoder not found"
and the test fails on the build rather than on the code under test.
The test's subject is the final frame being decodable at the exact duration,
which has nothing to do with the container. It now writes a PNG, and the two
things the WebP path was implicitly covering are split out and each guarded on
what is actually present:
- test_inspect_media_exports_a_webp_still - skipped unless ffmpeg reports a
webp encoder, and now asserts the file really is WebP rather than merely
non-empty.
- test_inspect_media_reports_a_missing_encoder_instead_of_crashing - runs only
where the encoder is absent, pinning the behaviour that surfaced this: the
tool reports ffmpeg's failure as a tool error and writes no partial file.
The product code has the same assumption and I left it alone. inspect_media
accepts any suffix in _IMAGE_SUFFIXES and hands the path to ffmpeg, so a .webp
request on a build without libwebp fails with ffmpeg's own message. That is a
poor message, not a crash or a corrupt file, and pre-validating the encoder
list is a separate change.
pytest tests/test_inspect_media_tool.py
68 passed, 2 skipped in 22.92s (1 failed, 66 passed, 1 skipped before)
Not verified: the WebP success path. This machine has no webp encoder, so
test_inspect_media_exports_a_webp_still skips here and has only been checked
for collection, not for a passing run.
test_rich_colors_follow_theme_and_undo_as_one_edit fails identically on every
macOS run, timing out after 30s waiting for a span that never appears. It was
written off as timing noise twice. It is not flaky - it is two Linux-only
input conventions, and the product code is fine.
Control+click: macOS delivers a Control-modified primary click as contextmenu,
not click. Instrumenting the Lemon swatch shows the button receiving
pointerdown, mousedown, contextmenu, pointerup, mouseup - and no click, so the
menu item's handler never runs and no highlight is applied. Control was never
meaningful here anyway; the palette item has no modifier behaviour. The two
calls now use a plain click.
Control+Z: the editor's undo accelerator is Cmd+Z on macOS. With the clicks
fixed, both undo assertions still failed until the presses became
ControlOrMeta+Z, which Playwright maps per platform.
Both fixes are portable - a plain click and ControlOrMeta are unchanged on
Linux, where this test already passes.
pytest tests/test_document_rich_color_reset_and_contrast.py
2 passed in 2.11s (1 failed, 1 passed before)
Not verified: Linux. I only have macOS here, so the claim that this stays
green on CI rests on the modifier being a no-op there, not on a run.
Running test_auth_regressions.py before the email modules failed 23 tests
that pass in isolation:
pytest -p no:randomly tests/test_auth_regressions.py \
tests/test_email_urgency_checkpoint.py
23 failed, 15 passed (38 passed in the reverse order)
Every failure was ImportError "cannot import name X (unknown location)"
against a module already in sys.modules, which is what an empty stub module
looks like to a later import.
Two writes leaked, both in this file:
- test_pop_notifications_owner_filtered inserted five empty stub modules with
a bare sys.modules[name] = mod and never removed them.
- _ensure_stub wrote its stub into sys.modules itself, so the autouse
fixture's monkeypatch.setitem three lines later captured that stub as the
value to restore. The fixture looked like it cleaned up and could not.
Both now go through monkeypatch, including the parent-package stub and the
attribute wiring, so everything is undone at teardown. The redundant setitem
calls in the fixture are gone: re-setting a key whose stub is already
installed is what made the leak invisible.
Pre-existing on lab, not introduced by any open PR. Verified against the
email subpackage branch too: identical numbers there before this fix.
The snapshot harness read ODYSSEUS_TEST_STATIC_ORIGIN into a module constant.
That env var is published by the session static-server fixture, which runs
after collection has already imported the module, so the constant always held
the 7011 fallback and the capture connected to a port nothing was listening on:
route.fetch: connect ECONNREFUSED 127.0.0.1:7011
Neither branch is wrong on its own. The harness was written while the fixture
still bound a fixed 7011, and the ephemeral-port change removed that port. The
two only disagree once they are in the same tree, which is what this integration
branch is for.
With this, the baseline recorded before the stylesheet split is reproduced
exactly after it, so the split is confirmed to preserve computed styles rather
than only argued to.
start-macos.sh is the single-instance launcher and adopts whatever is
already listening: an open ChromaDB port is a resource it reuses. With
one checkout that is right. With several, it means a scratch worktree
silently attaching to another checkout's vector store, and the script
reports it as a success.
odysseus-dev is the sibling that owns isolation instead. Ports are
derived from the worktree path, so two checkouts never collide and one
checkout always gets the same URL. A ChromaDB this worktree did not
start is refused, never adopted — we start our own or fall closed to
keyword mode and say which. The data dir, database and browser-MCP
cache live under .odysseus-dev/, leaving data/ to a normal launch. A
checkout wired into launchd or systemd will not boot at all, and the
ports the project already means something by (7000, 7011, 7860, 8100)
are refused even when asked for explicitly.
Readiness is /api/ready rather than a TCP accept: the port accepting
connections says nothing about the database or a writable data dir.
That endpoint is not auth-exempt, so the tool owns a dev admin account,
hands it to setup.py and prints it.
--from-pr N fetches pull/N/head into its own worktree and boots it,
borrowing a venv so a PR is a few seconds rather than a pip install.
start-macos.sh is untouched: it is what the LaunchAgent runs.
static/style.css is 51,425 lines in one file. Hundreds of selectors are
declared more than once and !important is used throughout, so the rendered
result is a function of source order. Extracting a block into its own file
changes that order, and nothing in the suite would notice - which makes a
51k-line split unfalsifiable and "looks fine to me" the only available
evidence.
This moves no CSS. It captures getComputedStyle over a fixed inventory of
676 elements across three pages, four viewports, both themes and the three
density modes - 16,224 element snapshots - hashes them, and compares against
tests/css_snapshot/baseline.json. A capture takes about 21 seconds.
The bench page synthesises one element per selector from an evidence-driven
list: every selector declared more than once in style.css that can be
expressed as a static compound chain, plus a curated set per feature area.
Redeclared selectors are the ones a reorder can flip. The bench loads
whatever stylesheets index.html ships, so it keeps measuring the real set
once the file is split. tests/test_css_computed_style_snapshot.py also carries
a self-test that swaps two conflicting .attach-strip declarations and asserts
the digest moves, so the harness cannot silently stop watching.
The second half is the asset-manifest check specs/frontend.md asks for,
scoped to stylesheets: every stylesheet referenced by shipped HTML and by the
sw.js precache exists, and index.html and sw.js agree on the ?v= string. They
hardcode it independently today, so a split that updates one and not the
other ships an offline cache nobody notices until a plane.
The session-scoped autouse fixture bound 127.0.0.1:7011 and raised when the
port was taken. Because it is autouse, that raise errored every collected
test rather than the browser ones: a second worktree running its own suite
produced 10,612 errors, none of them about the code under test. 7011 is also
the application's own default port, so the suite could not run while a local
instance was up.
Bind port 0 instead and publish the resulting origin as
ODYSSEUS_TEST_STATIC_ORIGIN. The browser tests shell out to node, which
inherits the environment, so the snippets read process.env rather than
hardcoding a port. ODYSSEUS_TEST_STATIC_PORT still pins one when something
outside pytest has to reach the server; that is the only path that can now
fail to bind, and it fails with a message that says so.
Two concurrent full runs from one checkout now both pass. Only the docx export
snippet is an rf-string, so it is the only one whose JS braces needed doubling.
_terminate_owned_chrome() walked Path("/proc") unconditionally, so on macOS
and Windows iterdir() raised FileNotFoundError out of private-browser session
shutdown. src/tools/cookbook.py already guards the same kind of scan with
os.path.isdir("/proc").
The sweep only reclaims Chrome trees that agent-browser reparented, so it is
an optimisation rather than a correctness requirement: degrade to a no-op
rather than failing the whole shutdown path.
_PROC_ROOT is a module attribute so both branches are testable on either kind
of host. The procfs-present path had no coverage at all before this.
The R09 scholarly hardening routes every hop through check_outbound_url with
block_private=True. On a DNS64/NAT64 network, an IPv4-only host can resolve
through the RFC 6052 Well-Known Prefix. For example, export.arxiv.org resolved
to 64:ff9b::924b:5b2a in the reproduced environment.
CPython classifies that outer IPv6 prefix as reserved, so the URL guard rejected
the request before examining the effective IPv4 destination.
Decode addresses in exactly 64:ff9b::/96 to their embedded IPv4 destination and
evaluate that destination under the strict outbound policy.
The translated target is always checked with private-address blocking enabled.
This prevents the NAT64 prefix from becoming a path to loopback, private,
shared/CGNAT, link-local, multicast, unspecified, or other non-global IPv4
space.
Network-specific translation prefixes are not decoded. In particular,
64:ff9b:1::/48 remains subject to the existing IPv6 policy.
Coverage includes:
- public IPv4 destinations embedded through the RFC 6052 prefix
- loopback, link-local, private, CGNAT, multicast, unspecified, benchmark, and
TEST-NET rejection
- network-specific NAT64 prefixes remaining undecoded
- deterministic scholarly provider tests without live DNS dependence
- redirect query parameter isolation
- relative redirect resolution
- HTTP error fallback behavior
- editor draft GET and DELETE behavior remaining unaffected
R09, R11, and R01 were independently audited and otherwise left unchanged.
static/style.css is 51,425 lines and holds every panel's styles, so two
people working on unrelated panels still edit the same file. This is the
first of three steps that give each panel group its own stylesheet. It moves
the document library, image gallery and image editor rules that can move
into static/css/documents-gallery-editor.css.
The new file is loaded eagerly from index.html immediately after style.css,
so the cascade is the concatenation of the two in that order, which is the
order those rules already had. Nothing is deferred and nothing renders
differently.
Rules moved verbatim. A rule stays in style.css when its selector group
covers more than one panel, when the class is shared design language used
from modules outside the panel, or when moving it would flip which of two
equally specific rules wins on an element the markup puts both classes on.
That last rule is what keeps this change small: most panel-prefixed rules on
this branch compete with a shared component rule somewhere later in the file.
Nineteen tests read static/style.css directly and go red the moment a rule
they assert on moves, which says nothing about the page. They now go through
tests/helpers/stylesheets.app_css(), which concatenates the stylesheets in
the order index.html loads them. The helper also exposes the <link> markup
for tests that build a synthetic page through Playwright.
- Enforce outbound URL policy on all hops via bounded manual redirects in scholarly lookups
- Guard declared Content-Length in EditorDraftRoute before request body parsing
- Keep scholarly lookup timeouts and budgets as internal constants rather than surface env vars
- Narrow TUI threat-model description around demonstrable host shell bridge behavior
- Add behavioral regression tests for redirect security and pre-parsing body size guards
Covers both behaviours end to end: a refused outbound URL and an exhausted
budget must skip the request entirely rather than fall through to httpx, the
remaining budget caps each hop's timeout, and an over-ceiling Content-Length
returns 413 without the body being parsed.
Verified against the pre-fix tree: the User-Agent was hardcoded, the endpoints
were literals, check_outbound_url was absent, both hops carried independent
12.0s timeouts, and an oversized declared body returned 422 after parsing
rather than 413 before it.