Commit Graph
2302 Commits
Author SHA1 Message Date
Léo bf5d8e4001 test(runtime): cover the remaining four requested behaviours
Completes the lane alteixeira20 asked for. Test-only: tests/ and test helpers,
no production runtime code, no benchmark fixtures or allowlists.

Supplied workspace context must not produce a clarification. Pins
_looks_like_unattended_clarification on four shapes that hand the decision
back ("could you please share", "shall I", "which approach do you prefer")
and three ordinary answers that must not trip it.

Repeated update_plan is not the turn's work. ask_user and update_plan are
permitted on nearly every turn, so if they counted as execution a model could
loop on them and look busy. Pins that _tool_rejection_reason does not
advertise either as an available tool, and that update_plan is permitted
without ever being in required.

Request-scoped tool authority. _request_scoped_allowed_tool_names must not
make an undeclared tool executable; the native-terminal widening is pinned
separately so it stays opt-in rather than drifting into the default.

Foreign-process safety. The Chrome sweep matches this runtime's own profile
prefix, so a fake procfs with our pid, a user's ordinary Chrome and another
worktree's agent browser must leave exactly two of the three alone.

14 passed, 2 xfailed. The xfails are the negative-wording cases from the first
commit that do not hold yet.
2026-09-30 17:38:57 +02:00
Léo d423632559 test(runtime): scope the negative-wording claim to the inferred path
Verified end to end against a local Qwen3.5-9B Q4_K_M that when the user
explicitly enables web for the turn, none of the three phrasings withholds
web_search, web_fetch or private_browser, including the one these tests record
as held. The held case holds on the inferred path only, where no toggle is set
and the runtime decides from intent.

That distinction was missing and the file read as a stronger claim than the
measurement supports. An explicit toggle beating an inferred negative may be
the intended semantics, so it is recorded rather than asserted.
2026-09-30 17:19:07 +02:00
Léo 5fd9114882 test(runtime): pin negative capability wording, test-only
First slice of the runtime regression lane. Drives stream_agent_loop with a
fake model and asserts on the tools the runtime offers, which is its decision
about what the turn may do. No production runtime code is touched and no
benchmark fixture or allowlist is imported.

The fake-model pattern is the one tests/test_tool_policy.py already uses:
patch stream_llm_with_fallback and inspect the tools kwarg.

Measured on lab@c499c01b, negative web wording is only partially detected:

  "Answer from memory only, don't search online."            web tools withheld
  "Summarise what you already know. Do not search the web."  web tools OFFERED
  "No web search please, just tell me what you know..."      web tools OFFERED

The case that holds is a plain regression guard. The two that do not are
xfail(strict=True): they run on every suite, document the target, and fail the
moment the behaviour lands so the marker gets removed rather than lingering.
A positive control keeps the guard from being satisfied by removing the web
tools altogether.
2026-09-30 17:11:10 +02:00
Alexandre Teixeira 6105702901 Merge pull request #16 from o3LL/lane/decomposition-and-test-isolation
refactor(lane): decompose style.css, pin it with a computed-style harness, and isolate worktree runs
2026-09-29 14:54:36 +01:00
Alexandre Teixeira 65c35c8211 Merge pull request #15 from o3LL/fix/test-static-port-isolation
test(harness): bind the static test server to an ephemeral port
2026-09-29 14:54:18 +01:00
Alexandre Teixeira 956671a731 Merge pull request #14 from o3LL/fix/proc-guard-private-browser
fix(browser): skip the Chrome sweep when the host has no procfs
2026-09-29 14:54:00 +01:00
Alexandre Teixeira aafd8a98eb Merge pull request #13 from o3LL/v1/css-split-3
refactor(css): move cookbook, research, memory and settings styles to their own file
2026-09-29 14:53:44 +01:00
Léo 4318974903 test(css): read the static origin at call time, not at import
The snapshot harness read ODYSSEUS_TEST_STATIC_ORIGIN into a module constant.
That env var is published by the session static-server fixture, which runs
after collection has already imported the module, so the constant always held
the 7011 fallback and the capture connected to a port nothing was listening on:

    route.fetch: connect ECONNREFUSED 127.0.0.1:7011

Neither branch is wrong on its own. The harness was written while the fixture
still bound a fixed 7011, and the ephemeral-port change removed that port. The
two only disagree once they are in the same tree, which is what this integration
branch is for.

With this, the baseline recorded before the stylesheet split is reproduced
exactly after it, so the split is confirmed to preserve computed styles rather
than only argued to.
2026-09-29 10:38:06 +02:00
Léo 822f7d80e4 Merge branch 'feat/odysseus-dev-worktree-boot' into lane/decomposition-and-test-isolation 2026-09-29 10:36:18 +02:00
Léo ac5e1a600d Merge branch 'fix/test-static-port-isolation' into lane/decomposition-and-test-isolation 2026-09-29 10:36:18 +02:00
Léo f8e9d00256 Merge branch 'fix/proc-guard-private-browser' into lane/decomposition-and-test-isolation 2026-09-29 10:36:18 +02:00
Léo 16c0fe0975 Merge remote-tracking branch 'fork/v1/css-split-3' into lane/decomposition-and-test-isolation 2026-09-29 10:36:18 +02:00
Léo 255f7c46c1 feat(scripts): add odysseus-dev, a per-worktree isolated boot
start-macos.sh is the single-instance launcher and adopts whatever is
already listening: an open ChromaDB port is a resource it reuses. With
one checkout that is right. With several, it means a scratch worktree
silently attaching to another checkout's vector store, and the script
reports it as a success.

odysseus-dev is the sibling that owns isolation instead. Ports are
derived from the worktree path, so two checkouts never collide and one
checkout always gets the same URL. A ChromaDB this worktree did not
start is refused, never adopted — we start our own or fall closed to
keyword mode and say which. The data dir, database and browser-MCP
cache live under .odysseus-dev/, leaving data/ to a normal launch. A
checkout wired into launchd or systemd will not boot at all, and the
ports the project already means something by (7000, 7011, 7860, 8100)
are refused even when asked for explicitly.

Readiness is /api/ready rather than a TCP accept: the port accepting
connections says nothing about the database or a writable data dir.
That endpoint is not auth-exempt, so the tool owns a dev admin account,
hands it to setup.py and prints it.

--from-pr N fetches pull/N/head into its own worktree and boots it,
borrowing a venv so a PR is a few seconds rather than a pip install.

start-macos.sh is untouched: it is what the LaunchAgent runs.
2026-09-25 17:35:47 +02:00
Léo 1863de33c2 test(css): pin computed styles against a committed baseline
static/style.css is 51,425 lines in one file. Hundreds of selectors are
declared more than once and !important is used throughout, so the rendered
result is a function of source order. Extracting a block into its own file
changes that order, and nothing in the suite would notice - which makes a
51k-line split unfalsifiable and "looks fine to me" the only available
evidence.

This moves no CSS. It captures getComputedStyle over a fixed inventory of
676 elements across three pages, four viewports, both themes and the three
density modes - 16,224 element snapshots - hashes them, and compares against
tests/css_snapshot/baseline.json. A capture takes about 21 seconds.

The bench page synthesises one element per selector from an evidence-driven
list: every selector declared more than once in style.css that can be
expressed as a static compound chain, plus a curated set per feature area.
Redeclared selectors are the ones a reorder can flip. The bench loads
whatever stylesheets index.html ships, so it keeps measuring the real set
once the file is split. tests/test_css_computed_style_snapshot.py also carries
a self-test that swaps two conflicting .attach-strip declarations and asserts
the digest moves, so the harness cannot silently stop watching.

The second half is the asset-manifest check specs/frontend.md asks for,
scoped to stylesheets: every stylesheet referenced by shipped HTML and by the
sw.js precache exists, and index.html and sw.js agree on the ?v= string. They
hardcode it independently today, so a split that updates one and not the
other ships an offline cache nobody notices until a plane.
2026-09-25 11:32:46 +02:00
Léo 9b2185f1cc test(harness): bind the static test server to an ephemeral port
The session-scoped autouse fixture bound 127.0.0.1:7011 and raised when the
port was taken. Because it is autouse, that raise errored every collected
test rather than the browser ones: a second worktree running its own suite
produced 10,612 errors, none of them about the code under test. 7011 is also
the application's own default port, so the suite could not run while a local
instance was up.

Bind port 0 instead and publish the resulting origin as
ODYSSEUS_TEST_STATIC_ORIGIN. The browser tests shell out to node, which
inherits the environment, so the snippets read process.env rather than
hardcoding a port. ODYSSEUS_TEST_STATIC_PORT still pins one when something
outside pytest has to reach the server; that is the only path that can now
fail to bind, and it fails with a message that says so.

Two concurrent full runs from one checkout now both pass. Only the docx export
snippet is an rf-string, so it is the only one whose JS braces needed doubling.
2026-09-25 10:45:58 +02:00
Léo 8b85e11fb4 fix(browser): skip the Chrome sweep when the host has no procfs
_terminate_owned_chrome() walked Path("/proc") unconditionally, so on macOS
and Windows iterdir() raised FileNotFoundError out of private-browser session
shutdown. src/tools/cookbook.py already guards the same kind of scan with
os.path.isdir("/proc").

The sweep only reclaims Chrome trees that agent-browser reparented, so it is
an optimisation rather than a correctness requirement: degrade to a no-op
rather than failing the whole shutdown path.

_PROC_ROOT is a module attribute so both branches are testable on either kind
of host. The procfs-present path had no coverage at all before this.
2026-09-25 10:30:49 +02:00
Alexandre Teixeira 0e07d9a675 Merge pull request #10 from o3LL/fix/review-20260923-bound-outbound-and-draft-limits
fix(security): harden scholarly lookups and editor draft limits
2026-09-24 12:22:04 +01:00
Alexandre Teixeira 11ad4e2739 fix(url-safety): resolve NAT64 well-known prefix to IPv4 target
The R09 scholarly hardening routes every hop through check_outbound_url with
block_private=True. On a DNS64/NAT64 network, an IPv4-only host can resolve
through the RFC 6052 Well-Known Prefix. For example, export.arxiv.org resolved
to 64:ff9b::924b:5b2a in the reproduced environment.

CPython classifies that outer IPv6 prefix as reserved, so the URL guard rejected
the request before examining the effective IPv4 destination.

Decode addresses in exactly 64:ff9b::/96 to their embedded IPv4 destination and
evaluate that destination under the strict outbound policy.

The translated target is always checked with private-address blocking enabled.
This prevents the NAT64 prefix from becoming a path to loopback, private,
shared/CGNAT, link-local, multicast, unspecified, or other non-global IPv4
space.

Network-specific translation prefixes are not decoded. In particular,
64:ff9b:1::/48 remains subject to the existing IPv6 policy.

Coverage includes:

- public IPv4 destinations embedded through the RFC 6052 prefix
- loopback, link-local, private, CGNAT, multicast, unspecified, benchmark, and
  TEST-NET rejection
- network-specific NAT64 prefixes remaining undecoded
- deterministic scholarly provider tests without live DNS dependence
- redirect query parameter isolation
- relative redirect resolution
- HTTP error fallback behavior
- editor draft GET and DELETE behavior remaining unaffected

R09, R11, and R01 were independently audited and otherwise left unchanged.
2026-09-23 15:32:04 +01:00
Léo 3807cb22e7 refactor(css): move cookbook, research, memory and settings styles to their own file
Last step of splitting static/style.css by panel. This moves the cookbook,
deep research, memory and settings rules that can move into
static/css/cookbook-research-memory-settings.css.

After the three steps style.css still holds the shell and every rule whose
move would change the cascade, which on this branch is most of the file.
Eager ordered links, verbatim moves, no rendering change.
2026-09-23 16:12:53 +02:00
Léo acefcbec47 refactor(css): move email, calendar, notes and task styles to their own file
Second step of splitting static/style.css by panel. This moves the email,
calendar, notes and tasks rules that can move into
static/css/email-calendar-notes-tasks.css.

Same shape as the first step: eager ordered links, verbatim moves, no
rendering change. The header comment in each split file lists the load order
and the earlier file's header is updated so all of them agree.
2026-09-23 16:12:52 +02:00
Léo e3eb7137b7 refactor(css): move document, gallery and image-editor styles to their own file
static/style.css is 51,425 lines and holds every panel's styles, so two
people working on unrelated panels still edit the same file. This is the
first of three steps that give each panel group its own stylesheet. It moves
the document library, image gallery and image editor rules that can move
into static/css/documents-gallery-editor.css.

The new file is loaded eagerly from index.html immediately after style.css,
so the cascade is the concatenation of the two in that order, which is the
order those rules already had. Nothing is deferred and nothing renders
differently.

Rules moved verbatim. A rule stays in style.css when its selector group
covers more than one panel, when the class is shared design language used
from modules outside the panel, or when moving it would flip which of two
equally specific rules wins on an element the markup puts both classes on.
That last rule is what keeps this change small: most panel-prefixed rules on
this branch compete with a shared component rule somewhere later in the file.

Nineteen tests read static/style.css directly and go red the moment a rule
they assert on moves, which says nothing about the page. They now go through
tests/helpers/stylesheets.app_css(), which concatenates the stylesheets in
the order index.html loads them. The helper also exposes the <link> markup
for tests that build a synthetic page through Playwright.
2026-09-23 16:12:51 +02:00
Alexandre Teixeira 0784cd2fed fix(review): harden scholarly redirects, editor draft limits, and threat model
- Enforce outbound URL policy on all hops via bounded manual redirects in scholarly lookups
- Guard declared Content-Length in EditorDraftRoute before request body parsing
- Keep scholarly lookup timeouts and budgets as internal constants rather than surface env vars
- Narrow TUI threat-model description around demonstrable host shell bridge behavior
- Add behavioral regression tests for redirect security and pre-parsing body size guards
2026-09-23 12:53:06 +01:00
Alexandre Teixeira c90a81dbdc Merge commit '14afd3afb6274ff986733ddff992001e19c80e29' into fix/pr10-merge-ready 2026-09-23 12:29:27 +01:00
Alexandre Teixeira f577934777 fix(ci): restore green lab baseline before runtime W1 (#9)
Validated locally and on GitHub Actions before integration into lab.
2026-09-23 10:16:54 +01:00
Léo 1b75fdb438 test: pin the 2026-09-23 review fixes
Covers both behaviours end to end: a refused outbound URL and an exhausted
budget must skip the request entirely rather than fall through to httpx, the
remaining budget caps each hop's timeout, and an over-ceiling Content-Length
returns 413 without the body being parsed.

Verified against the pre-fix tree: the User-Agent was hardcoded, the endpoints
were literals, check_outbound_url was absent, both hops carried independent
12.0s timeouts, and an oversized declared body returned 422 after parsing
rather than 413 before it.
2026-09-23 11:09:48 +02:00
Léo 130b49f75d docs(security): document the tool approval gate and its default
The post-external-context gate is off unless ODYSSEUS_TOOL_APPROVAL_GATE is
set, but THREAT_MODEL.md described only the untrusted-context wrapper. A reader
auditing the prompt-injection posture would reasonably assume the gate was
active.

State the default, what it blocks when enabled, the two deliberate exemptions,
and note in Known Gaps that the compensating control for the missing shell
sandbox is off by default.
2026-09-23 11:09:48 +02:00
Léo 3412c4212e fix(editor-drafts): refuse oversized drafts before parsing the body
The 256 MiB ceiling was only enforced inside _dump_payload, which runs after
FastAPI has parsed the request and after json.dumps has re-serialised it. By
that point the payload has been materialised several times over, so the limit
rejected an allocation it had already paid for.

Check Content-Length first, on both write routes. A request that omits or lies
about the header still reaches the exact byte count, which remains
authoritative.
2026-09-23 11:09:48 +02:00
Léo edc94a244e fix(search): bound and police the scholarly metadata lookups
The arXiv and OpenAlex title resolvers called two hardcoded endpoints with a
hand-written "Odysseus/0.20" User-Agent, no outbound-URL policy, and a full
12s timeout each. APP_VERSION was already 1.0.3, so the agent string was wrong
the moment it was written, and a scholarly query could hold a user-facing
search open for the sum of all three hops.

Move the endpoints to named constants, build the User-Agent from APP_VERSION,
run both calls through check_outbound_url (they follow redirects, so the final
host is not the one in the constant), and give the whole SearXNG -> OpenAlex ->
arXiv chain one shared wall-clock budget.

The budget is a ContextVar rather than a parameter so the existing test doubles
for _openalex_title_results and _arxiv_title_results keep working unchanged.
2026-09-23 11:09:48 +02:00
Alexandre Teixeira 771d48f147 ci: install FFmpeg for media integration tests 2026-09-23 01:08:20 +01:00
Alexandre Teixeira 3cabbb9bca test(agent): restore exact turn-policy regressions 2026-09-23 00:14:30 +01:00
Alexandre Teixeira c0b71a5cef fix(tools): constrain Python environment namespace mount 2026-09-23 00:10:50 +01:00
Alexandre Teixeira 0012baa17c ci(trivy): free build cache before image scan 2026-09-22 23:27:29 +01:00
Alexandre Teixeira c0a1cecc00 test(baseline): align UI and model profile expectations 2026-09-22 23:27:17 +01:00
Alexandre Teixeira 9bc1cdcab0 test(harness): isolate imports and stabilize browser and DNS fixtures 2026-09-22 23:27:08 +01:00
Alexandre Teixeira 4520b4f6f0 fix(ui): restore rich text controls and accessible research map 2026-09-22 23:26:43 +01:00
Alexandre Teixeira b9cccb93ac fix(tools): expose active Python environment in workspace namespace 2026-09-22 23:26:28 +01:00
Alexandre Teixeira 79a55fac38 fix(agent): preserve focused turn contracts and verified completion 2026-09-22 23:26:16 +01:00
Alexandre Teixeira 128102ef26 test(runtime): isolate tool execution module state 2026-09-22 18:17:42 +01:00
Alexandre Teixeira 5568d7d631 fix(sw): complete editor panel precache 2026-09-22 18:09:04 +01:00
Alexandre Teixeira 83088f8811 ci: provision browser dependencies for pytest 2026-09-22 18:08:16 +01:00
Alexandre Teixeira ebcc6c524f test(database): restore model endpoint isolation 2026-09-22 18:07:09 +01:00
Alexandre Teixeira dbba4978ee test(harness): support cache-busted frontend module imports 2026-09-22 18:06:18 +01:00
Alexandre Teixeira 7e798fe925 Merge pull request #6 from pewdiepie-archdaemon/pr/alteixeira20/subscription-provider-ux
feat(provider): improve subscription UX and lazy discovery
2026-09-22 14:02:55 +01:00
Alexandre Teixeira 8bb48780a1 feat(provider): add lazy Featherless model discovery 2026-09-22 13:12:19 +01:00
Alexandre Teixeira 2b2bf0fb90 fix(ui): refine provider endpoint controls 2026-09-22 13:12:19 +01:00
Alexandre Teixeira 46c8451a29 fix(ui): bust cache for collapsible subscription usage 2026-09-22 13:12:19 +01:00
Alexandre Teixeira 411559913c feat(ui): refine subscription model and usage controls 2026-09-22 13:12:19 +01:00
Alexandre Teixeira ed7ccfd584 feat(provider): support multiple ChatGPT subscriptions with usage 2026-09-22 13:12:19 +01:00
Alexandre Teixeira 45330097b8 Merge pull request #7 from pewdiepie-archdaemon/fix/maintainer-harness-ci-reproducibility-v1
fix(ci): make maintainer harness and lab workflow portable
2026-09-22 12:44:20 +01:00
Alexandre Teixeira dd13f53507 ci: support maintainer lab pull requests 2026-09-22 12:39:14 +01:00