Commit Graph
23 Commits
Author SHA1 Message Date
Léo 2bda788311 ci: run the pytest suite as four parallel shards
The suite is 12.1k tests in a single CI job — nearly six minutes of pytest
that every push and every PR waits on in one block, on top of a setup step
that already installs npm, Playwright, FFmpeg and bubblewrap. Split it into
four sections the matrix runs in parallel: 352s becomes a 98s longest pole
locally.

Shards partition by test *file*, not by the area_* taxonomy markers. Those
markers do not partition the suite - a file can carry a hand-applied area_*
mark on top of the one conftest derives from its filename, so a marker-based
split would run those tests in more than one section. Assignment is a total
function of the file path instead, so every file lands in exactly one shard
and the four together run every test exactly once.

Sharding deselects rather than narrowing collection, so every test module is
still imported, in the same order, in every shard. The import-time stubbing
in conftest and the session-scoped static server behave identically whether
the suite runs whole or in sections - this suite has known collection-order
coupling and splitting by path would have walked into it.

Balance uses the existing `slow` marker as its weight signal rather than a
committed duration table that would go stale unnoticed. Files pack
heaviest-first into the lightest shard, which is deterministic for a given
file set, so every parallel job computes the same plan from the same commit.

Verified: the four shards together reproduce the full run exactly - 12141
tests selected across the four, and the same 59 failures, 78 skips and 2
xfails, by node ID and not merely by count.
2026-10-02 18:42:10 +02:00
Alexandre Teixeira f4793696f4 merge: reconcile Wave 1.1 with post-PR40 lab
Merge canonical lab 9557b8d5909eb4a885c3bf49e19a65dd904f8c1d exactly once.
Retain invocation journal ownership and lineage, provider terminal ordering,
teacher handoff, framed DONE handling, and canonical authority/Ajax routing.

Combine dynamic dispatch receipts with lab policy forwarding. Adapt native
shell/patch evidence, explicit TUI verifiers, and artifact recovery presentation.
Refresh generated configuration source links and strengthen adapter regressions.

Validation: focused 2118 passed; Wave 1.1 script 2291 passed; broad runtime
5649 passed; full pytest 11581 passed, 53 skipped, 2 xfailed, 6 subtests passed.
Compileall 1689 Python files; syntax 279 JS and 82 MJS files; diff and
conflict-marker checks passed.
2026-10-01 09:09:55 +01:00
Alexandre Teixeira bdccb1ddb1 Merge lab into test/95-release-smoke-suite 2026-10-01 02:45:49 +01:00
Alexandre Teixeira 9ee9205bfc merge: sync lab after stylesheet split 2026-10-01 01:08:15 +01:00
Alexandre Teixeira 23fdc26079 Merge lab into refactor/split-style-css 2026-10-01 00:48:58 +01:00
Alexandre Teixeira fe80eb795e merge: reconcile reviewed lab baseline for runtime wave 1.1 2026-10-01 00:25:42 +01:00
Léo 12f74ec9ea test(smoke): add a release smoke suite over every advertised feature area
The decomposition lanes have two safety nets and neither covers the
product. The checkpoint benchmark measures the agent runtime; the
computed-style snapshot pins the CSS. Nothing checked that Notes,
Calendar, Documents, Email, Memory, Cookbook or Settings still worked
after a route package moved or a 17,000-line module was split - and the
unit suite does not, since a byte-identical file move can break tests
that pass on the base branch with CI green throughout. The 28 Playwright
specs we do have are all under tests/e2e/photo-editor/ and no workflow
runs them.

scripts/odysseus-smoke boots this worktree through `odysseus dev` and
runs tests/smoke/: one scenario per area, each asserting a user-visible
outcome rather than a status code. Models come from a deterministic
OpenAI-compatible stub on an ephemeral loopback port; email reuses the
existing ODYSSEUS_EMAIL_FIXTURE path rather than inventing a second
mechanism. No scenario touches a live endpoint or the network.

The report is a per-area table that prints the areas the suite does not
cover next to the ones it does, and builds its rows from the registry
rather than from what happened to run, so an area cannot go missing by
having its module deleted or renamed. Under a plain pytest with nothing
booted every scenario skips with the reason, so the full suite stays
green.
2026-09-30 12:12:58 +02:00
Alexandre Teixeira 3748a621d1 Merge current lab into agent runtime decomposition 2026-09-29 19:15:54 +01:00
Léo b5d1505582 docs(tests): record the known full-suite failures
specs/testing-devops.md lists "no canonical full-suite known-failing/flaky
ledger" as a gap. Without one a first local run is uninterpretable: you cannot
tell a regression from a platform artifact, so you either chase a non-bug or
ignore a real one.

Six failures on macOS against lab@c499c01b, each with its cause and a verdict
rather than a blanket "environmental":

- three compare an unresolved /tmp path against a resolved /private/tmp one.
  Those are test bugs and the file says so.
- one asserts ffmpeg exit 0 for a .webp still, which is a build option Homebrew
  does not always carry. Needs a skip or a PNG fallback.
- one opens real sockets and needs a fast connection refusal. Environmental.
- one Playwright colour-contrast test had been written off as a flake. It is
  not: three consecutive runs failed identically at ~31s. Recorded as
  unexplained and possibly a real defect, because calling it noise is what
  stopped anyone looking.

Also documents the prerequisites, since most surprise failures are a missing
npm ci rather than anything here, and the CHROMADB_PORT precaution: the client
reaches Chroma over HTTP regardless of the data directory, so a test run can
attach to a store holding real data.
2026-09-29 18:00:12 +02:00
Léo 16c0fe0975 Merge remote-tracking branch 'fork/v1/css-split-3' into lane/decomposition-and-test-isolation 2026-09-29 10:36:18 +02:00
Alexandre Teixeira ea847e6c0b test(runtime): establish isolated validation and comparison gates 2026-09-26 13:13:07 +01:00
Léo 1863de33c2 test(css): pin computed styles against a committed baseline
static/style.css is 51,425 lines in one file. Hundreds of selectors are
declared more than once and !important is used throughout, so the rendered
result is a function of source order. Extracting a block into its own file
changes that order, and nothing in the suite would notice - which makes a
51k-line split unfalsifiable and "looks fine to me" the only available
evidence.

This moves no CSS. It captures getComputedStyle over a fixed inventory of
676 elements across three pages, four viewports, both themes and the three
density modes - 16,224 element snapshots - hashes them, and compares against
tests/css_snapshot/baseline.json. A capture takes about 21 seconds.

The bench page synthesises one element per selector from an evidence-driven
list: every selector declared more than once in style.css that can be
expressed as a static compound chain, plus a curated set per feature area.
Redeclared selectors are the ones a reorder can flip. The bench loads
whatever stylesheets index.html ships, so it keeps measuring the real set
once the file is split. tests/test_css_computed_style_snapshot.py also carries
a self-test that swaps two conflicting .attach-strip declarations and asserts
the digest moves, so the harness cannot silently stop watching.

The second half is the asset-manifest check specs/frontend.md asks for,
scoped to stylesheets: every stylesheet referenced by shipped HTML and by the
sw.js precache exists, and index.html and sw.js agree on the ?v= string. They
hardcode it independently today, so a split that updates one and not the
other ships an offline cache nobody notices until a plane.
2026-09-25 11:32:46 +02:00
Léo e3eb7137b7 refactor(css): move document, gallery and image-editor styles to their own file
static/style.css is 51,425 lines and holds every panel's styles, so two
people working on unrelated panels still edit the same file. This is the
first of three steps that give each panel group its own stylesheet. It moves
the document library, image gallery and image editor rules that can move
into static/css/documents-gallery-editor.css.

The new file is loaded eagerly from index.html immediately after style.css,
so the cascade is the concatenation of the two in that order, which is the
order those rules already had. Nothing is deferred and nothing renders
differently.

Rules moved verbatim. A rule stays in style.css when its selector group
covers more than one panel, when the class is shared design language used
from modules outside the panel, or when moving it would flip which of two
equally specific rules wins on an element the markup puts both classes on.
That last rule is what keeps this change small: most panel-prefixed rules on
this branch compete with a shared component rule somewhere later in the file.

Nineteen tests read static/style.css directly and go red the moment a rule
they assert on moves, which says nothing about the page. They now go through
tests/helpers/stylesheets.app_css(), which concatenates the stylesheets in
the order index.html loads them. The helper also exposes the <link> markup
for tests that build a synthetic page through Playwright.
2026-09-23 16:12:51 +02:00
RaresKeY 2007e92a25 fix(devops): harden docker config defaults (#4349) 2026-06-16 04:03:43 +01:00
Alexandre Teixeira 0afea5db5c test: add report-only order-sensitivity runner (#3982)
* test: add report-only order-sensitivity runner

* test: report cwd in order-sensitivity runner
2026-06-15 15:49:47 +09:00
Alexandre Teixeira 4a0c778317 test: mark first slow tests from duration evidence (#3711) 2026-06-10 01:07:38 +02:00
Alexandre Teixeira b3d7477a17 test: pilot core database stub helper (#3685) 2026-06-09 22:23:33 +02:00
Alexandre Teixeira cf585f4dd3 test: add fast lane and duration visibility (#3659) 2026-06-09 20:11:47 +02:00
Alexandre Teixeira 02f930a3c5 test: add focused test selection runner (#3556) 2026-06-09 17:03:47 +02:00
Alexandre Teixeira ce80fd0a92 test(taxonomy): auto-mark tests by area and sub-area (#3491) 2026-06-09 01:13:28 +02:00
Alexandre Teixeira 2b6ff994da docs(tests): define testing standard and taxonomy (#3372) 2026-06-08 01:15:47 +02:00
Alexandre Teixeira 9d90be15df refactor(tests): add temp sqlite helper (#2930) 2026-06-07 23:44:16 +02:00
Alexandre Teixeira 29020b92e3 docs(tests): document helper conventions
Documentation-only PR continuing #2523. Adds tests/README.md to document helper conventions, validation expectations, and the next test-suite refactor phase.
2026-06-05 14:04:10 +01:00