The decomposition lanes have two safety nets and neither covers the
product. The checkpoint benchmark measures the agent runtime; the
computed-style snapshot pins the CSS. Nothing checked that Notes,
Calendar, Documents, Email, Memory, Cookbook or Settings still worked
after a route package moved or a 17,000-line module was split - and the
unit suite does not, since a byte-identical file move can break tests
that pass on the base branch with CI green throughout. The 28 Playwright
specs we do have are all under tests/e2e/photo-editor/ and no workflow
runs them.
scripts/odysseus-smoke boots this worktree through `odysseus dev` and
runs tests/smoke/: one scenario per area, each asserting a user-visible
outcome rather than a status code. Models come from a deterministic
OpenAI-compatible stub on an ephemeral loopback port; email reuses the
existing ODYSSEUS_EMAIL_FIXTURE path rather than inventing a second
mechanism. No scenario touches a live endpoint or the network.
The report is a per-area table that prints the areas the suite does not
cover next to the ones it does, and builds its rows from the registry
rather than from what happened to run, so an area cannot go missing by
having its module deleted or renamed. Under a plain pytest with nothing
booted every scenario skips with the reason, so the full suite stays
green.
specs/testing-devops.md lists "no canonical full-suite known-failing/flaky
ledger" as a gap. Without one a first local run is uninterpretable: you cannot
tell a regression from a platform artifact, so you either chase a non-bug or
ignore a real one.
Six failures on macOS against lab@c499c01b, each with its cause and a verdict
rather than a blanket "environmental":
- three compare an unresolved /tmp path against a resolved /private/tmp one.
Those are test bugs and the file says so.
- one asserts ffmpeg exit 0 for a .webp still, which is a build option Homebrew
does not always carry. Needs a skip or a PNG fallback.
- one opens real sockets and needs a fast connection refusal. Environmental.
- one Playwright colour-contrast test had been written off as a flake. It is
not: three consecutive runs failed identically at ~31s. Recorded as
unexplained and possibly a real defect, because calling it noise is what
stopped anyone looking.
Also documents the prerequisites, since most surprise failures are a missing
npm ci rather than anything here, and the CHROMADB_PORT precaution: the client
reaches Chroma over HTTP regardless of the data directory, so a test run can
attach to a store holding real data.
static/style.css is 51,425 lines in one file. Hundreds of selectors are
declared more than once and !important is used throughout, so the rendered
result is a function of source order. Extracting a block into its own file
changes that order, and nothing in the suite would notice - which makes a
51k-line split unfalsifiable and "looks fine to me" the only available
evidence.
This moves no CSS. It captures getComputedStyle over a fixed inventory of
676 elements across three pages, four viewports, both themes and the three
density modes - 16,224 element snapshots - hashes them, and compares against
tests/css_snapshot/baseline.json. A capture takes about 21 seconds.
The bench page synthesises one element per selector from an evidence-driven
list: every selector declared more than once in style.css that can be
expressed as a static compound chain, plus a curated set per feature area.
Redeclared selectors are the ones a reorder can flip. The bench loads
whatever stylesheets index.html ships, so it keeps measuring the real set
once the file is split. tests/test_css_computed_style_snapshot.py also carries
a self-test that swaps two conflicting .attach-strip declarations and asserts
the digest moves, so the harness cannot silently stop watching.
The second half is the asset-manifest check specs/frontend.md asks for,
scoped to stylesheets: every stylesheet referenced by shipped HTML and by the
sw.js precache exists, and index.html and sw.js agree on the ?v= string. They
hardcode it independently today, so a split that updates one and not the
other ships an offline cache nobody notices until a plane.
static/style.css is 51,425 lines and holds every panel's styles, so two
people working on unrelated panels still edit the same file. This is the
first of three steps that give each panel group its own stylesheet. It moves
the document library, image gallery and image editor rules that can move
into static/css/documents-gallery-editor.css.
The new file is loaded eagerly from index.html immediately after style.css,
so the cascade is the concatenation of the two in that order, which is the
order those rules already had. Nothing is deferred and nothing renders
differently.
Rules moved verbatim. A rule stays in style.css when its selector group
covers more than one panel, when the class is shared design language used
from modules outside the panel, or when moving it would flip which of two
equally specific rules wins on an element the markup puts both classes on.
That last rule is what keeps this change small: most panel-prefixed rules on
this branch compete with a shared component rule somewhere later in the file.
Nineteen tests read static/style.css directly and go red the moment a rule
they assert on moves, which says nothing about the page. They now go through
tests/helpers/stylesheets.app_css(), which concatenates the stylesheets in
the order index.html loads them. The helper also exposes the <link> markup
for tests that build a synthetic page through Playwright.
Documentation-only PR continuing #2523. Adds tests/README.md to document helper conventions, validation expectations, and the next test-suite refactor phase.