Four more tests in the same family as the /tmp ones this change already
fixes, and they hide for the same reason: the failure depends on how
long $TMPDIR happens to be.
tests/test_shell_routes.py::TestHostDockerAccess (three) and
tests/test_cookbook_docker_access.py::test_container_opt_in_with_unix_
socket_is_allowed each bind an AF_UNIX socket at tmp_path/"docker.sock".
macOS gives sun_path 104 bytes including the terminator. pytest roots
tmp_path at $TMPDIR, which on a stock Mac is a 49-character
/var/folders/<2>/<30>/T/; add pytest-of-<user>/pytest-<n>/ and the
test's own name and the bind path is 115 bytes before the filename.
OSError: AF_UNIX path too long
Linux allows 108 and roots $TMPDIR at /tmp, so CI never sees it. Under a
shortened $TMPDIR the path lands at exactly 103 and passes — until
pytest's run counter reaches two digits and it becomes 104. That is why
the ledger's counts did not include these: they were measured somewhere
the path fit.
Adds tests/helpers/unix_sockets.bound_unix_socket, which binds under a
short directory and asserts the length before it tries, so the next
socket fixture fails with a sentence rather than an errno. Records the
trap in KNOWN_FAILURES.md along with the instruction to re-measure with
the default $TMPDIR.
The route now answers 400 with a message naming the expected shape.
admin.js was taught to print `data.detail`; the Unified Integrations
form in settings.js still printed `Failed (400)` and dropped it.
That gap is exactly where the new validation bites. The client-side
JSON.parse guard added alongside it catches unparseable input, so the
only values that reach the route's 400 are ones that parse but are not
a list — `"npx"`, `{}`, `null` — and for those the status code alone
tells the user nothing about what is wrong with what they typed.
Adds source-level coverage for both forms; the PR changed two JS files
with no test on either.
Splitting style.css deleted it, and two files outside static/ still
named it:
- scripts/verify_background_research_cards.mjs injected
`<link rel="stylesheet" href="/static/style.css">` into the page it
builds. A stylesheet that 404s does not fail — the script kept
checking card layout against an unstyled page and kept reporting
pass. It now reads the shell's <link> tags out of index.html, the way
tests/css_snapshot/capture.mjs already does, so the set cannot drift
out from under it again.
- tests/css_snapshot/bench.html's hand-open fallback linked the same
deleted file. Replaced with the ordered set index.html ships.
test_every_stylesheet_referenced_by_shipped_html_exists only walked
static/*.html, which is why neither showed up. Extend it over the bench
page: the bench's whole output is computed styles, so a dead link there
is worth more than an unstyled page nobody looks at.
#25 fixed two sys.modules writes in test_auth_regressions.py that left empty
stub modules behind for the rest of the session, breaking 23 tests under one
collection order while the full suite stayed green. The class is wider than
that file, and an audit is the wrong answer to it: nothing stops the next one,
and the failure it causes lands on an unrelated test in a different file.
So this is a guard instead. An autouse fixture in the root conftest snapshots
which src.* / core.* names are bound to a bare ModuleType, and fails any test
that adds one. "Bare" is the same test the clear_fake_* helpers already use -
a plain types.ModuleType with no on-disk __file__. MagicMock stand-ins are out
of scope: they answer every attribute, so they fail at the point of use rather
than silently, and several files install them deliberately.
Three details that matter:
- It lives in the root conftest, so it is set up before any test-module
fixture and torn down after all of them. A stub a test's own teardown
removes is not reported.
- It drops the leaked entries as well as reporting them, so the failure stays
on the test that introduced it instead of cascading through the rest of the
run.
- It only reports stubs added during the test. Import state the session starts
with, including this conftest's own src.database stub, is left alone.
It found one beyond #25 on the first full run: _stub_heavy in
test_scheduler_restart_doublefire.py leaks the same five src.* modules as the
test #25 fixed, via sys.modules.setdefault. It already receives monkeypatch,
so the fix is to register through it. Fixed here because the guard has to land
green.
Full suite, macOS, default collection order:
this branch 10655 passed, 6 failed, 6 skipped 406s
lab 10655 passed, 6 failed, 6 skipped 371s
Same six either way, which is the point - none of this is visible in the
default order. Four are pre-existing macOS environment failures:
test_glob_confined_e2e and the two test_code_nav_tools document cases resolve
/tmp to /private/tmp, and
test_real_socket_falls_back_from_dead_first_to_live_second is connect-refused
timing on real sockets. The other two are the rich-colour and ffmpeg items
from the same ledger, fixed on their own branches.
Not verified: Linux, and any collection order other than the default. The
guard is order-independent by construction - it compares before and after
within a single test - but I have only run the default order.
test_inspect_media_exports_final_decodable_frame_at_exact_duration exported to
/workspace/final.webp and asserted exit_code == 0. WebP encoding is an ffmpeg
build option, not something this project requires - Homebrew's macOS ffmpeg is
built without it:
ffmpeg -encoders | grep -ic webp -> 0
so the tool returns "ffmpeg still extraction failed: ... Encoder not found"
and the test fails on the build rather than on the code under test.
The test's subject is the final frame being decodable at the exact duration,
which has nothing to do with the container. It now writes a PNG, and the two
things the WebP path was implicitly covering are split out and each guarded on
what is actually present:
- test_inspect_media_exports_a_webp_still - skipped unless ffmpeg reports a
webp encoder, and now asserts the file really is WebP rather than merely
non-empty.
- test_inspect_media_reports_a_missing_encoder_instead_of_crashing - runs only
where the encoder is absent, pinning the behaviour that surfaced this: the
tool reports ffmpeg's failure as a tool error and writes no partial file.
The product code has the same assumption and I left it alone. inspect_media
accepts any suffix in _IMAGE_SUFFIXES and hands the path to ffmpeg, so a .webp
request on a build without libwebp fails with ffmpeg's own message. That is a
poor message, not a crash or a corrupt file, and pre-validating the encoder
list is a separate change.
pytest tests/test_inspect_media_tool.py
68 passed, 2 skipped in 22.92s (1 failed, 66 passed, 1 skipped before)
Not verified: the WebP success path. This machine has no webp encoder, so
test_inspect_media_exports_a_webp_still skips here and has only been checked
for collection, not for a passing run.
test_rich_colors_follow_theme_and_undo_as_one_edit fails identically on every
macOS run, timing out after 30s waiting for a span that never appears. It was
written off as timing noise twice. It is not flaky - it is two Linux-only
input conventions, and the product code is fine.
Control+click: macOS delivers a Control-modified primary click as contextmenu,
not click. Instrumenting the Lemon swatch shows the button receiving
pointerdown, mousedown, contextmenu, pointerup, mouseup - and no click, so the
menu item's handler never runs and no highlight is applied. Control was never
meaningful here anyway; the palette item has no modifier behaviour. The two
calls now use a plain click.
Control+Z: the editor's undo accelerator is Cmd+Z on macOS. With the clicks
fixed, both undo assertions still failed until the presses became
ControlOrMeta+Z, which Playwright maps per platform.
Both fixes are portable - a plain click and ControlOrMeta are unchanged on
Linux, where this test already passes.
pytest tests/test_document_rich_color_reset_and_contrast.py
2 passed in 2.11s (1 failed, 1 passed before)
Not verified: Linux. I only have macOS here, so the claim that this stays
green on CI rests on the modifier being a no-op there, not on a run.
`add_server` wrapped `json.loads(args)` in a bare `except` that fell back to
`[]`, so an Args value that is not JSON — a bare path, which is what the form's
placeholder invites people to type — registered the server and spawned the
stdio subprocess with an empty argv. Nothing surfaced the loss: the POST
returned 200 and the row persisted with `"args": []`.
The route now returns 400 for an unparseable value, and also for valid JSON of
the wrong shape: `args=5` reached `StdioServerParameters(args=5)` and raised an
unhandled TypeError in the error formatter's `" ".join(...)`.
Both form clients mirror the guard instead of leaving the user to read a 400
they cannot see. `settings.js` stops silently defaulting a bad Args value, and
`admin.js` gains the same client-side parse check plus a `res.ok` branch so a
server-side rejection is not reported as a connection failure.
Ported from public `dev` (`9d5c0319`, #6215 upstream, fixing #6211), with its
test. The `admin.js` hunks are inert on `lab` — `initMcpForm` early-returns
because that form's markup is not in this build — and are carried anyway to
keep the two lines from diverging further.
The host cache was gated on the list being non-empty, so "queried fine, no
eligible peers" looked exactly like a cold cache and every caller paid for
another `tailscale status --json` — a subprocess with a 5s timeout.
Gate on the timestamp instead. Failures still leave the timestamp unset, so a
missing binary, a non-zero exit or unparseable output stays retryable rather
than being cached for the full TTL.
Ported from public `dev` (`affaee1e`, #6228 upstream), with its test.
`_cookbook_kill_session` kills the tmux session, then sweeps /proc for
model servers that survived the SIGHUP. The sweep had no guard, so on
macOS and Windows `os.listdir("/proc")` raised FileNotFoundError after
the kill had already succeeded. The function's outer except turned that
into `{"error": "...No such file or directory: '/proc'", "exit_code": 1}`
and skipped the state write that marks the session stopped — the agent
is told a stop failed that actually worked.
Guard the sweep with a procfs check, the way `_scan_running_model_processes`
already does a few hundred lines up. The root and the check now live in
`core/platform_compat`, which is where OS differences belong and which
makes both branches patchable from a test on either kind of host.
Adds a class-level guard test: the third instance of this defect, and two
of the three were found by reading source rather than by a test.
Three of the six recorded failures were the same test bug: an unresolved
/tmp path compared against a resolved /private/tmp one. macOS makes /tmp a
symlink, so a fixture built with tempfile.mkdtemp(dir="/tmp") and a code path
that resolves what it reports disagree about a file both found correctly.
test_code_nav_tools builds its fixture unresolved and compares it against the
reported path. One realpath fixes both of its failures.
test_glob_confined_e2e is the same cause through a longer route: it mixed
os.path.realpath(ws) with an unresolved secret directory, so relpath emitted
"../../../../tmp/<absolute path>" and the assertion that the absolute path was
absent from the output matched it as a substring. Resolving the secret
directory puts both sides in one tree and the relative path stays short.
macOS full suite goes from 6 failures to 3. The remaining three are an ffmpeg
build without a WebP encoder, a socket test that needs a fast connection
refusal, and the rich-text colour test that is still unexplained.
The ledger is updated in the same change so it does not describe failures that
no longer happen.
Running test_auth_regressions.py before the email modules failed 23 tests
that pass in isolation:
pytest -p no:randomly tests/test_auth_regressions.py \
tests/test_email_urgency_checkpoint.py
23 failed, 15 passed (38 passed in the reverse order)
Every failure was ImportError "cannot import name X (unknown location)"
against a module already in sys.modules, which is what an empty stub module
looks like to a later import.
Two writes leaked, both in this file:
- test_pop_notifications_owner_filtered inserted five empty stub modules with
a bare sys.modules[name] = mod and never removed them.
- _ensure_stub wrote its stub into sys.modules itself, so the autouse
fixture's monkeypatch.setitem three lines later captured that stub as the
value to restore. The fixture looked like it cleaned up and could not.
Both now go through monkeypatch, including the parent-package stub and the
attribute wiring, so everything is undone at teardown. The redundant setitem
calls in the fixture are gone: re-setting a key whose stub is already
installed is what made the leak invisible.
Pre-existing on lab, not introduced by any open PR. Verified against the
email subpackage branch too: identical numbers there before this fix.
static/js/document.js is 17,579 lines and is about to be decomposed behind a
re-export wrapper. 51 test files read it off disk and grep it as text, and 28
of those slice it with `src.split("function a", 1)[1].split("function b", 1)[0]`
-- "the region between a and b", which only means what the test intends while a
and b are neighbours in one file. Several also hard-code the file's two-space
indentation, which no extracted module reproduces. Left alone, the first
extraction makes those assertions cover the wrong region, and an `x in region`
check passes while covering more than it was written for.
tests/helpers/document_source is the one place that names the file now:
- document_source() is the entry plus everything under static/js/document/,
so a membership assertion keeps finding its subject wherever it lands;
- function_body()/declaration() locate a construct by name in whichever
module defines it and end at its real closing brace, so neither moving it
nor moving its neighbour changes the region.
The rewrite only collapses a slice when the old terminator sat at the
construct's end. 28 slices deliberately span a whole family of functions --
everything from _docxHexColor to exportAsDocx -- and collapsing one to its
first member drops what the assertions look for, so those stay as they are and
are listed in KNOWN_ADJACENCY_SLICES, to be converted as each family becomes a
module. That list may only shrink.
Two guards come with it:
- test_document_source_test_hygiene fails on a direct read of the entry file
and on any new adjacency slice;
- test_frontend_module_graph resolves every relative import under static/
(718 of them, none broken today) and requires the document module set to
stay in the sw.js precache, since the worker fetches the URLs it lists and
not what they import.
test_document_module_api pins the 38 default-export keys and 29 named exports
by loading the module in a browser and reading what it actually exports, rather
than grepping for the literal object -- after extraction that object may be
assembled from imports, and a source-shape check would pass while the export
was broken.
No JavaScript moves here. static/ is untouched.
specs/testing-devops.md lists "no canonical full-suite known-failing/flaky
ledger" as a gap. Without one a first local run is uninterpretable: you cannot
tell a regression from a platform artifact, so you either chase a non-bug or
ignore a real one.
Six failures on macOS against lab@c499c01b, each with its cause and a verdict
rather than a blanket "environmental":
- three compare an unresolved /tmp path against a resolved /private/tmp one.
Those are test bugs and the file says so.
- one asserts ffmpeg exit 0 for a .webp still, which is a build option Homebrew
does not always carry. Needs a skip or a PNG fallback.
- one opens real sockets and needs a fast connection refusal. Environmental.
- one Playwright colour-contrast test had been written off as a flake. It is
not: three consecutive runs failed identically at ~31s. Recorded as
unexplained and possibly a real defect, because calling it noise is what
stopped anyone looking.
Also documents the prerequisites, since most surprise failures are a missing
npm ci rather than anything here, and the CHROMADB_PORT precaution: the client
reaches Chroma over HTTP regardless of the data directory, so a test run can
attach to a store holding real data.
style-part-01..08 were arbitrary 6,000-line cuts, so the names said nothing
and a reader had no way to guess which file held a rule.
The original stylesheet turns out to be roughly area-grouped already, so
cutting where the content changes rather than every 6,000 lines produces
boundaries worth naming. Classifying every top-level rule by selector prefix
and smoothing over a 40-rule window gives 17 stable runs: agent-chat, compare,
memory, documents, admin-settings, skills, gallery, cookbook, tasks,
image-editor, email, notes, calendar, research.
The numeric prefix stays because load order is load-bearing, and it also
disambiguates the areas that appear twice: the original interleaves, and
merging non-adjacent blocks of the same area would reorder the cascade.
The names describe where a file sits, not a claim that it holds every rule for
that area or only rules for it. The header in each file says so, because that
is the misreading this naming invites.
Reconstruction was checked against lab's style.css before rewriting: identical
after whitespace normalisation, all 47,528 lines accounted for. The computed
style baseline still reproduces exactly.
static/style.css was 47,530 lines. It is now nine files under static/css/:
tokens.css holds the :root custom properties, light theme and density classes,
and style-part-01..08 carry the rest in their original order. Nothing was
rewritten; every line moved verbatim and the <link> order in index.html
reproduces the original file byte for byte.
Cut points are brace-depth zero and outside block comments, so no rule or
comment is split. All 47,531 source lines are accounted for across the
fragments.
The computed-style baseline recorded before the split is reproduced exactly,
which is the evidence that the cascade is unchanged rather than an argument
that it should be.
Three things the split broke and this fixes:
- The harness self-test swapped two conflicting .attach-strip blocks inside
style.css to prove the digest is order-sensitive. It rewrote one hardcoded
URL, so with the file gone it silently measured an unmodified page. It now
searches every stylesheet, rewrites whichever holds the pair, and fails
loudly if none does.
- Four tests asserted the cache-bust token by matching /static/style.css?v=.
They check the real invariant now, that every app stylesheet shares one
token with app.js, through stylesheet_cache_version().
- The three panel stylesheets carried their own tokens, so a browser could
hold half an old cascade and half a new one. All app stylesheets now bust
together.
static/style.css no longer holds every rule. 79 rules for the document,
gallery and editor panels now live in static/css/, yet 27 browser tests still
built their synthetic page with a single <link> to style.css, and 51 more read
that one file as though it were the whole cascade. Those tests kept passing
while covering less: a rule that moved became invisible to the assertion that
was meant to pin it.
Python readers now call tests.helpers.stylesheets.app_css(), and synthetic
pages are built from stylesheet_link_tags() so they load exactly what
index.html loads, in the same order. The helper already existed; this moves
the remaining callers onto it.
test_portal_dropdown_z_js parametrised over a file list including style.css to
assert an absence. Checking a negative against one file of a split stylesheet
is how a moved rule escapes, so the CSS case now checks the concatenation.
Two guards keep it from coming back: one fails on any test reading
static/style.css directly, the other on any synthetic page linking it alone.
Both name the helper to use.
No production code changes. Suite is unchanged at 6 pre-existing failures.
Duplicate @keyframes names resolve last-wins across the whole cascade, so
every definition but the last was dead code that still read as live at its
call site. Two of the five duplicated names differed from the winner:
fadeIn style.css had an opacity-only variant before the one that
adds translateY, so every consumer was already sliding
research-pulse a background-colour pulse sat before the opacity/scale one
that actually runs
The other three (spin, loading-bounce, pulse) were byte-identical repeats.
Removing the losing definitions changes nothing rendered, which the computed
style snapshot confirms: the committed baseline is reproduced exactly across
all three pages and 24 variants. Deciding that a consumer wanted the plain
fade, or the background pulse, would be a visual change and belongs in its own
PR with screenshots.
This also unblocks the mechanical stylesheet split. While two definitions of a
name differed, neither could be moved: relocating either changes which one is
last, and therefore changes behaviour. A regression test now fails on any
duplicate name so the trap cannot come back.
The snapshot harness read ODYSSEUS_TEST_STATIC_ORIGIN into a module constant.
That env var is published by the session static-server fixture, which runs
after collection has already imported the module, so the constant always held
the 7011 fallback and the capture connected to a port nothing was listening on:
route.fetch: connect ECONNREFUSED 127.0.0.1:7011
Neither branch is wrong on its own. The harness was written while the fixture
still bound a fixed 7011, and the ephemeral-port change removed that port. The
two only disagree once they are in the same tree, which is what this integration
branch is for.
With this, the baseline recorded before the stylesheet split is reproduced
exactly after it, so the split is confirmed to preserve computed styles rather
than only argued to.
start-macos.sh is the single-instance launcher and adopts whatever is
already listening: an open ChromaDB port is a resource it reuses. With
one checkout that is right. With several, it means a scratch worktree
silently attaching to another checkout's vector store, and the script
reports it as a success.
odysseus-dev is the sibling that owns isolation instead. Ports are
derived from the worktree path, so two checkouts never collide and one
checkout always gets the same URL. A ChromaDB this worktree did not
start is refused, never adopted — we start our own or fall closed to
keyword mode and say which. The data dir, database and browser-MCP
cache live under .odysseus-dev/, leaving data/ to a normal launch. A
checkout wired into launchd or systemd will not boot at all, and the
ports the project already means something by (7000, 7011, 7860, 8100)
are refused even when asked for explicitly.
Readiness is /api/ready rather than a TCP accept: the port accepting
connections says nothing about the database or a writable data dir.
That endpoint is not auth-exempt, so the tool owns a dev admin account,
hands it to setup.py and prints it.
--from-pr N fetches pull/N/head into its own worktree and boots it,
borrowing a venv so a PR is a few seconds rather than a pip install.
start-macos.sh is untouched: it is what the LaunchAgent runs.
static/style.css is 51,425 lines in one file. Hundreds of selectors are
declared more than once and !important is used throughout, so the rendered
result is a function of source order. Extracting a block into its own file
changes that order, and nothing in the suite would notice - which makes a
51k-line split unfalsifiable and "looks fine to me" the only available
evidence.
This moves no CSS. It captures getComputedStyle over a fixed inventory of
676 elements across three pages, four viewports, both themes and the three
density modes - 16,224 element snapshots - hashes them, and compares against
tests/css_snapshot/baseline.json. A capture takes about 21 seconds.
The bench page synthesises one element per selector from an evidence-driven
list: every selector declared more than once in style.css that can be
expressed as a static compound chain, plus a curated set per feature area.
Redeclared selectors are the ones a reorder can flip. The bench loads
whatever stylesheets index.html ships, so it keeps measuring the real set
once the file is split. tests/test_css_computed_style_snapshot.py also carries
a self-test that swaps two conflicting .attach-strip declarations and asserts
the digest moves, so the harness cannot silently stop watching.
The second half is the asset-manifest check specs/frontend.md asks for,
scoped to stylesheets: every stylesheet referenced by shipped HTML and by the
sw.js precache exists, and index.html and sw.js agree on the ?v= string. They
hardcode it independently today, so a split that updates one and not the
other ships an offline cache nobody notices until a plane.
The session-scoped autouse fixture bound 127.0.0.1:7011 and raised when the
port was taken. Because it is autouse, that raise errored every collected
test rather than the browser ones: a second worktree running its own suite
produced 10,612 errors, none of them about the code under test. 7011 is also
the application's own default port, so the suite could not run while a local
instance was up.
Bind port 0 instead and publish the resulting origin as
ODYSSEUS_TEST_STATIC_ORIGIN. The browser tests shell out to node, which
inherits the environment, so the snippets read process.env rather than
hardcoding a port. ODYSSEUS_TEST_STATIC_PORT still pins one when something
outside pytest has to reach the server; that is the only path that can now
fail to bind, and it fails with a message that says so.
Two concurrent full runs from one checkout now both pass. Only the docx export
snippet is an rf-string, so it is the only one whose JS braces needed doubling.
_terminate_owned_chrome() walked Path("/proc") unconditionally, so on macOS
and Windows iterdir() raised FileNotFoundError out of private-browser session
shutdown. src/tools/cookbook.py already guards the same kind of scan with
os.path.isdir("/proc").
The sweep only reclaims Chrome trees that agent-browser reparented, so it is
an optimisation rather than a correctness requirement: degrade to a no-op
rather than failing the whole shutdown path.
_PROC_ROOT is a module attribute so both branches are testable on either kind
of host. The procfs-present path had no coverage at all before this.
The R09 scholarly hardening routes every hop through check_outbound_url with
block_private=True. On a DNS64/NAT64 network, an IPv4-only host can resolve
through the RFC 6052 Well-Known Prefix. For example, export.arxiv.org resolved
to 64:ff9b::924b:5b2a in the reproduced environment.
CPython classifies that outer IPv6 prefix as reserved, so the URL guard rejected
the request before examining the effective IPv4 destination.
Decode addresses in exactly 64:ff9b::/96 to their embedded IPv4 destination and
evaluate that destination under the strict outbound policy.
The translated target is always checked with private-address blocking enabled.
This prevents the NAT64 prefix from becoming a path to loopback, private,
shared/CGNAT, link-local, multicast, unspecified, or other non-global IPv4
space.
Network-specific translation prefixes are not decoded. In particular,
64:ff9b:1::/48 remains subject to the existing IPv6 policy.
Coverage includes:
- public IPv4 destinations embedded through the RFC 6052 prefix
- loopback, link-local, private, CGNAT, multicast, unspecified, benchmark, and
TEST-NET rejection
- network-specific NAT64 prefixes remaining undecoded
- deterministic scholarly provider tests without live DNS dependence
- redirect query parameter isolation
- relative redirect resolution
- HTTP error fallback behavior
- editor draft GET and DELETE behavior remaining unaffected
R09, R11, and R01 were independently audited and otherwise left unchanged.
Last step of splitting static/style.css by panel. This moves the cookbook,
deep research, memory and settings rules that can move into
static/css/cookbook-research-memory-settings.css.
After the three steps style.css still holds the shell and every rule whose
move would change the cascade, which on this branch is most of the file.
Eager ordered links, verbatim moves, no rendering change.
Second step of splitting static/style.css by panel. This moves the email,
calendar, notes and tasks rules that can move into
static/css/email-calendar-notes-tasks.css.
Same shape as the first step: eager ordered links, verbatim moves, no
rendering change. The header comment in each split file lists the load order
and the earlier file's header is updated so all of them agree.
static/style.css is 51,425 lines and holds every panel's styles, so two
people working on unrelated panels still edit the same file. This is the
first of three steps that give each panel group its own stylesheet. It moves
the document library, image gallery and image editor rules that can move
into static/css/documents-gallery-editor.css.
The new file is loaded eagerly from index.html immediately after style.css,
so the cascade is the concatenation of the two in that order, which is the
order those rules already had. Nothing is deferred and nothing renders
differently.
Rules moved verbatim. A rule stays in style.css when its selector group
covers more than one panel, when the class is shared design language used
from modules outside the panel, or when moving it would flip which of two
equally specific rules wins on an element the markup puts both classes on.
That last rule is what keeps this change small: most panel-prefixed rules on
this branch compete with a shared component rule somewhere later in the file.
Nineteen tests read static/style.css directly and go red the moment a rule
they assert on moves, which says nothing about the page. They now go through
tests/helpers/stylesheets.app_css(), which concatenates the stylesheets in
the order index.html loads them. The helper also exposes the <link> markup
for tests that build a synthetic page through Playwright.
- Enforce outbound URL policy on all hops via bounded manual redirects in scholarly lookups
- Guard declared Content-Length in EditorDraftRoute before request body parsing
- Keep scholarly lookup timeouts and budgets as internal constants rather than surface env vars
- Narrow TUI threat-model description around demonstrable host shell bridge behavior
- Add behavioral regression tests for redirect security and pre-parsing body size guards