Commit Graph
1378 Commits
Author SHA1 Message Date
Alexandre Teixeira cdbcb44cc9 fix(runtime): handle synthetic requests without app scope and update env reference
- Narrowly guard _request_privileges() in routes/chat_routes.py against
  synthetic requests lacking scope['app'] or auth manager state, safely
  returning empty privileges without granting agent privileges.
- Add focused regression test in tests/test_context_resolution_route.py
  verifying that requests without app scope do not crash and cannot gain
  agent privileges or qualify for compact preview runtime.
- Regenerate website/configuration-reference.md mechanically to align with
  current source line numbers.
2026-10-01 23:38:45 +01:00
Alexandre Teixeira 57fe9946c2 refactor(runtime): one compact-runtime selection rule for route and dispatch
The chat route repeated the compact (clean v3) eligibility decision inline
to prepare the turn's context resolution, while the agent loop dispatched
on the contract stamp set by a separate, later condition. The two could
drift, and already disagreed for a user whose privileges demote the turn
to plain chat: the route prepared a compact resolution that no compact
runtime used.

src/agent_runtime/runtime_selection.py (no imports) now owns the rule:

- uses_compact_preview_runtime(): clean route requested, contract policy
  enabled, agent mode, agent permitted, not an image generation session.
- is_compact_preview_contract() and COMPACT_PREVIEW_MODE for the stamp.

The route evaluates the rule once, before context preparation, where all
of its facts are final (the agent privilege is read through the same
_request_privileges helper the later enforcement uses). That one value
gates the typed context resolution and is the _clean_v3_preview flag that
stamps the contract; inside the agent-contract branch it equals the
previous condition, so stamping behavior is unchanged. The agent loop
dispatches through is_compact_preview_contract(), and the compact runtime's
MODE is the shared constant.

A route-level matrix drives the real agent loop and asserts that route
preparation and compact dispatch agree for compact, escalated, configured
compact/full, regular, TUI, privilege-denied and image-generation turns.
2026-10-01 22:46:08 +01:00
Alexandre Teixeira 0054557027 fix(runtime): resolve compact-turn context once at the chat route
The first checkpoint removed terminal-metrics discovery, but a normal
compact chat turn still ran two context systems: build_chat_context's
legacy untyped lookup (directly or inside maybe_compact) and the typed
resolver inside stream_preview.

Resolve the typed ContextResolution once, at the chat route, before
build_chat_context, using the session's provider credentials. The
predicate mirrors _clean_v3_preview; every input it needs is known at
that point and the native-workspace term cannot veto a requested clean
route. The same object then:

- sizes legacy history shaping in build_chat_context through a new
  maybe_compact(context_length=...) override, so no legacy probe runs;
  an unknown window still shapes with DEFAULT_CONTEXT but gains no
  provenance;
- crosses stream_agent_loop (one new parameter, forwarded only at the
  compact dispatch) into stream_preview, which reuses it and probes only
  for callers that arrive without one or with one bound to another
  route.

ContextResolution now records the endpoint and model it describes
(endpoint URL excluded from repr and metrics). The bare legacy
context_length is never converted into typed evidence.

Credential scoping: origins compare with default ports normalized, an
empty host is never trusted, and the probe client never follows
redirects. Tests cover the configured origin, the server-resolved
Tailscale form, scheme/port/lookalike/userinfo/path origins, redirects,
and secret-free errors, logs and metrics.

The conftest guard now replaces only the resolver's I/O edges (HTTP
client and DNS-capable URL building) instead of the whole probe, and
exposes a context_probe_ledger fixture, so route integration tests run
the real resolver offline and can count metadata requests.
2026-10-01 22:31:21 +01:00
Alexandre Teixeira 837fbfd0ea feat(runtime): resolve compact-runtime context window at turn preparation
The compact (clean v3) runtime had no effective context window: it learned a
limit only reactively from a provider 400/413 and its terminal metrics carried
no context_length. PR #41 addressed the reporting gap by probing provider
metadata between the last model byte and [DONE], unauthenticated, and folded
known-table and endpoint evidence into one "known" flag.

Resolve the window once, before the first model request, instead:

- src/agent_runtime/context_resolution.py adds a typed ContextResolution
  (effective value, evidence class, source, all observations, conflicts,
  provider_io, cached, secret-free probe errors). Evidence classes stay
  distinct: runtime_confirmed (llama.cpp /slots, /props, or a limit the
  provider stated this turn), provider_advertised (models catalog),
  operator_declared (client_runtime_context.model_context_window),
  known_table, unknown (0, never a default).
- Selection is deterministic: runtime beats provider beats table; an
  operator declaration caps measured evidence and replaces weaker evidence.
  Disagreements are recorded as conflicts; a declaration below a measured
  value is a cap, above it a contradiction.
- The provider probe forwards the turn's credentials only to the provider's
  own origin, runs URL resolution off the event loop, is bounded by one
  deadline, never raises, and caches remote results per credential
  fingerprint (shorter TTL for failures; local servers are re-probed).
- stream_preview resolves at preparation (or accepts a supplied resolution),
  seeds the proactive trim budget from it when evidence is not unknown, and
  terminal metrics report only the stored resolution plus any limit the
  provider stated during the turn. Metrics perform no discovery.

src/agent_loop.py and the regular runtime's legacy model_context probe are
unchanged. A conftest guard keeps tests that drive the compact runtime with
placeholder endpoints from performing real DNS/HTTP lookups.
2026-10-01 21:59:53 +01:00
Alexandre Teixeira a183ec025b Merge pull request #44 from o3LL/fix/pty-session-group-teardown
fix(shell): kill the PTY command's whole session on timeout
2026-10-01 20:25:11 +01:00
Alexandre Teixeira 6f007ca55d Merge pull request #45 from o3LL/fix/windows-bash-pipes-env
fix(agent): capture Windows Bash output and pass the subprocess env
2026-10-01 20:20:45 +01:00
Alexandre Teixeira 9d0257134f feat(runtime): enforce server request authority 2026-10-01 18:35:41 +01:00
Léo 2429805a45 fix(shell): only treat ESRCH as proof a PTY session is gone
_session_alive collapsed every OSError from killpg(pgid, 0) into "the
group is gone". EPERM means the opposite — the group answered the probe
but holds a process we may not signal — so a session we could not touch
was reported as contained, and a timed-out command that left children
running said it had terminated cleanly.

Resolving PTY_KILL_ESCALATION also named signal.SIGKILL unconditionally,
which does not exist on native Windows. app.py imports this module at
start-up, so that turned a POSIX-only teardown detail into the whole app
failing to import there.
2026-10-01 18:46:19 +02:00
Léo f49e09e59a fix(shell): kill the PTY command's whole session on timeout
/api/shell/stream starts its PTY child under os.setsid, so the child
leads its own session and process group. The timeout, client-disconnect
and error paths all called proc.kill(), which signals only the group
leader. Creating a group and then signalling only its leader is strictly
worse than never creating one: the descendants are detached from the
server's group as well, so nothing else will ever reach them, while the
route reports "Command timed out after Ns" and exit_code -1 as if the
command were gone.

The kernel's controlling-terminal SIGHUP hid this for well-behaved
children, which is why it reads as working. Anything that ignores
SIGHUP — a nohup'ed job, a daemon, a process that means to outlive its
terminal — survives the kill indefinitely.

Signal the whole group instead, escalate to SIGKILL if it outlives the
grace period, and confirm it is actually gone. The timeout response now
says so when containment could not be established rather than claiming
a clean kill it did not get.
2026-10-01 18:43:28 +02:00
Léo 34f01c0b58 fix(agent): capture Windows Bash output and pass the subprocess env
The Windows branch of `_create_bash_subprocess` spawned Git Bash with
neither pipes nor the env it was handed. `proc.stdout` and `proc.stderr`
came back `None`, so `_run_subprocess_streaming`'s reader returned
immediately and the Bash tool reported `"(no output)"` alongside the real
exit code — while the child inherited the server's own stdout/stderr and
wrote agent command output into the console and the launchd/Docker logs.

The `env` parameter was accepted and never used, so `PATH`, `VIRTUAL_ENV`,
`HOME`, `TMPDIR` and the configured import paths carried in
`ctx["subproc_env"]` never reached the child on Windows, even though every
POSIX path applies them.

Spawn it the way the POSIX path at `:688` already does: `stdin=DEVNULL`,
`stdout=PIPE`, `stderr=PIPE`, `env=env`.

`website/configuration-reference.md` is generated from source line numbers,
so the four added lines shift one entry; regenerated with
`scripts/generate_env_reference.py`.
2026-10-01 16:48:05 +02:00
Alexandre Teixeira d49071bbec fix: close Wave 1.1 completion-gate audit findings
- Headless consumers (task scheduler, background follow-up) now treat a
  completion-gate final_response as the authoritative answer instead of
  collecting deltas only. A gated replacement no longer leaves scheduled
  output empty, which used to trigger an extra, ungated grace-summary
  model call.
- The scheduler closes the agent stream with contextlib.aclosing, so the
  approval-pause break unwinds the gate's journal and teacher-takeover
  context in its own task. Chained runs no longer inherit a stale
  parent_run_id, and later finalization no longer raises ContextVar
  reset errors.
- On provider error, the completion gate applies the live answer's
  statement filter to persisted round_texts. Diagnostics and the failure
  note survive; claims rejected by the gate cannot reappear on reload.
2026-10-01 14:49:32 +01:00
Alexandre Teixeira f4793696f4 merge: reconcile Wave 1.1 with post-PR40 lab
Merge canonical lab 9557b8d5909eb4a885c3bf49e19a65dd904f8c1d exactly once.
Retain invocation journal ownership and lineage, provider terminal ordering,
teacher handoff, framed DONE handling, and canonical authority/Ajax routing.

Combine dynamic dispatch receipts with lab policy forwarding. Adapt native
shell/patch evidence, explicit TUI verifiers, and artifact recovery presentation.
Refresh generated configuration source links and strengthen adapter regressions.

Validation: focused 2118 passed; Wave 1.1 script 2291 passed; broad runtime
5649 passed; full pytest 11581 passed, 53 skipped, 2 xfailed, 6 subtests passed.
Compileall 1689 Python files; syntax 279 JS and 82 MJS files; diff and
conflict-marker checks passed.
2026-10-01 09:09:55 +01:00
Alexandre Teixeira 5dfe1c353c test: stabilize final PR 40 validation gates 2026-10-01 06:39:22 +01:00
Alexandre Teixeira f74a262f73 merge: reconcile PR 40 with current lab
Integrate lab fff55a78 into PR #40 (cc25d5ba). Lab's modular email
backend/frontend, modular settings, split stylesheets (static/style.css
stays deleted), procfs compatibility, and request-scoped TurnContract
authority win; PR #40's routing classifiers, editor/email/task features,
and style.css changes are ported into lab's module and stylesheet homes.

Integration fixes:
- settings/api.js imports ui.js under its canonical versioned URL
- browser observations keep legacy CAPTCHA/access-block evidence
- artifact turns do not re-trigger broad-web research recovery
- env reference documents PR test-tool variables; page regenerated

PR #40 defects surfaced by lab gates and fixed here:
- web_fetch generic schema drops top-level anyOf (OpenAI contract);
  the compact preview contract still requires url or urls
- get_weather registered as a brokered network read
- new lazy editor modules precached for offline use
- SearXNG pin mirrored into GPU standalone compose files
- image model picker again skips offline endpoints

Tests updated where PR #40 changed behaviour on purpose, and PR tests
moved onto lab's document_source helpers.
2026-10-01 05:03:58 +01:00
Alexandre Teixeira d63932f1ad Merge lab into fix/procfs-pid-file-liveness 2026-10-01 03:10:10 +01:00
Alexandre Teixeira 6901c56192 Merge lab into feat/admin-build-provenance 2026-10-01 02:55:27 +01:00
Alexandre Teixeira c6f690a27b Merge lab into docs/env-configuration-reference 2026-10-01 02:52:37 +01:00
Alexandre Teixeira bdccb1ddb1 Merge lab into test/95-release-smoke-suite 2026-10-01 02:45:49 +01:00
Alexandre Teixeira ef0d96a3ae Merge lab into audit/ref-parity 2026-10-01 02:45:10 +01:00
Alexandre Teixeira 8c7e3a9411 Merge lab into test/runtime-behavior-regressions 2026-10-01 02:42:22 +01:00
Alexandre Teixeira 45e1e2e7d9 Merge lab into fix/token-cache-atomic-swap 2026-10-01 02:41:41 +01:00
pewdiepie-archdaemon 2e8413a54a Preserve preview harness, editor, email and task improvements
Snapshot current maintainer-preview application changes and regression fixtures for integration into lab. Excludes local runtime data, evaluation outputs and source backups. Focused Python regression selection: 140 passed; full suite not certified.
2026-10-01 01:34:26 +00:00
Alexandre Teixeira b0bc0b8c40 Merge lab into fix/6174-singleflight-cancel 2026-10-01 02:27:49 +01:00
Alexandre Teixeira 799adbbde6 Merge lab into refactor/email-library-package 2026-10-01 02:22:23 +01:00
Alexandre Teixeira 9477616e05 Merge lab into refactor/routes-email-subpackage 2026-10-01 02:15:36 +01:00
Alexandre Teixeira aedec7d005 fix(runtime): isolate nested invocation ownership 2026-10-01 02:11:53 +01:00
Alexandre Teixeira 33387d4a4c Merge lab into refactor/settings-shell-modules 2026-10-01 01:41:27 +01:00
Alexandre Teixeira 4c122de880 fix(runtime): scope completion claims to execution obligations 2026-10-01 01:36:18 +01:00
Alexandre Teixeira 9ee9205bfc merge: sync lab after stylesheet split 2026-10-01 01:08:15 +01:00
Alexandre Teixeira 86212bfe99 Merge lab into refactor/settings-panel-modules 2026-10-01 01:03:49 +01:00
Alexandre Teixeira 23fdc26079 Merge lab into refactor/split-style-css 2026-10-01 00:48:58 +01:00
Alexandre Teixeira 466a6b323a fix(runtime): preserve provider error terminal ordering 2026-10-01 00:33:46 +01:00
Alexandre Teixeira fe80eb795e merge: reconcile reviewed lab baseline for runtime wave 1.1 2026-10-01 00:25:42 +01:00
Alexandre Teixeira 40dbe17a95 Merge lab into test/document-module-set-helper
Resolve the overlap with #19 by preserving the whole-cascade stylesheet
helpers alongside #31's document-source/module-set helpers.

Maintainer validation:
- document/module contract and composition guards: 15 passed
- all 54 touched Python test modules: 366 passed
- py_compile: clean
- git diff --check: clean
2026-09-30 22:56:56 +01:00
Léo 40678bc466 feat(settings): show the running build's version and commit in the admin panel
Nothing in the UI said which build was loaded. /api/version has reported
version, build and source_commit since the harness started versioning itself
apart from the public semver, but the only way to read it was to curl the
endpoint — so "is the preview actually running the commit I just merged?" took
a terminal to answer.

Pins a footer under the settings sidebar nav showing the registered version
(plus the harness build when it differs) and the short source commit, with the
full hash on hover. It sits outside the nav's scroll container so it stays at
the bottom-left, and it is .admin-only, so syncAdminVisibility() hides it from
non-admins the same way it hides the Admin nav group.

The commit resolves at import via `git rev-parse HEAD` and is the string
"unknown" when that fails — a read-only Docker tree with no .git. The footer
treats "unknown" as absent and stays hidden when nothing is left to show,
rather than printing it. The collapsed rail and the two narrow tab-rail
layouts hide it too: neither has a bottom-left to write in.
2026-09-30 22:59:58 +02:00
Alexandre Teixeira 741f5dff2b Merge pull request #19 from o3LL/refactor/tests-read-the-whole-cascade
refactor(tests): read the whole cascade instead of style.css alone
2026-09-30 19:27:58 +01:00
Alexandre Teixeira 6ee51fee72 Merge pull request #18 from o3LL/fix/duplicate-keyframes
fix(css): collapse duplicate @keyframes names to the definition that wins
2026-09-30 19:18:13 +01:00
Alexandre Teixeira a969a68016 Merge pull request #34 from o3LL/test/macos-failures-and-stub-leak-guard
test: fix environment-dependent failures and guard module-stub leaks
2026-09-30 19:16:43 +01:00
Alexandre Teixeira 56484df737 test(media): detect ffmpeg encoders by codec alias 2026-09-30 19:16:29 +01:00
Alexandre Teixeira eaddc03729 Merge pull request #26 from o3LL/fix/tmp-realpath-test-bugs
fix(tests): resolve temp paths consistently on macOS
2026-09-30 18:21:04 +01:00
Alexandre Teixeira bfae249120 Merge pull request #28 from o3LL/fix/cookbook-stop-procfs-guard
fix(cookbook): skip the pid sweep when the host has no procfs
2026-09-30 18:07:07 +01:00
Alexandre Teixeira 054df080af Merge pull request #32 from o3LL/fix/6215-mcp-args-validation
fix(mcp): reject malformed Args on Add MCP Server instead of defaulting to []
2026-09-30 17:18:51 +01:00
Léo bf5d8e4001 test(runtime): cover the remaining four requested behaviours
Completes the lane alteixeira20 asked for. Test-only: tests/ and test helpers,
no production runtime code, no benchmark fixtures or allowlists.

Supplied workspace context must not produce a clarification. Pins
_looks_like_unattended_clarification on four shapes that hand the decision
back ("could you please share", "shall I", "which approach do you prefer")
and three ordinary answers that must not trip it.

Repeated update_plan is not the turn's work. ask_user and update_plan are
permitted on nearly every turn, so if they counted as execution a model could
loop on them and look busy. Pins that _tool_rejection_reason does not
advertise either as an available tool, and that update_plan is permitted
without ever being in required.

Request-scoped tool authority. _request_scoped_allowed_tool_names must not
make an undeclared tool executable; the native-terminal widening is pinned
separately so it stays opt-in rather than drifting into the default.

Foreign-process safety. The Chrome sweep matches this runtime's own profile
prefix, so a fake procfs with our pid, a user's ordinary Chrome and another
worktree's agent browser must leave exactly two of the three alone.

14 passed, 2 xfailed. The xfails are the negative-wording cases from the first
commit that do not hold yet.
2026-09-30 17:38:57 +02:00
Léo 9738405310 fix(tests): bind the docker-socket fixtures somewhere sun_path fits
Four more tests in the same family as the /tmp ones this change already
fixes, and they hide for the same reason: the failure depends on how
long $TMPDIR happens to be.

tests/test_shell_routes.py::TestHostDockerAccess (three) and
tests/test_cookbook_docker_access.py::test_container_opt_in_with_unix_
socket_is_allowed each bind an AF_UNIX socket at tmp_path/"docker.sock".
macOS gives sun_path 104 bytes including the terminator. pytest roots
tmp_path at $TMPDIR, which on a stock Mac is a 49-character
/var/folders/<2>/<30>/T/; add pytest-of-<user>/pytest-<n>/ and the
test's own name and the bind path is 115 bytes before the filename.

    OSError: AF_UNIX path too long

Linux allows 108 and roots $TMPDIR at /tmp, so CI never sees it. Under a
shortened $TMPDIR the path lands at exactly 103 and passes — until
pytest's run counter reaches two digits and it becomes 104. That is why
the ledger's counts did not include these: they were measured somewhere
the path fit.

Adds tests/helpers/unix_sockets.bound_unix_socket, which binds under a
short directory and asserts the length before it tries, so the next
socket fixture fails with a sentence rather than an errno. Records the
trap in KNOWN_FAILURES.md along with the instruction to re-measure with
the default $TMPDIR.
2026-09-30 17:35:06 +02:00
Léo d423632559 test(runtime): scope the negative-wording claim to the inferred path
Verified end to end against a local Qwen3.5-9B Q4_K_M that when the user
explicitly enables web for the turn, none of the three phrasings withholds
web_search, web_fetch or private_browser, including the one these tests record
as held. The held case holds on the inferred path only, where no toggle is set
and the runtime decides from intent.

That distinction was missing and the file read as a stronger claim than the
measurement supports. An explicit toggle beating an inferred negative may be
the intended semantics, so it is recorded rather than asserted.
2026-09-30 17:19:07 +02:00
Léo 5f18767528 fix(mcp): show the route's rejection reason on the Integrations form too
The route now answers 400 with a message naming the expected shape.
admin.js was taught to print `data.detail`; the Unified Integrations
form in settings.js still printed `Failed (400)` and dropped it.

That gap is exactly where the new validation bites. The client-side
JSON.parse guard added alongside it catches unparseable input, so the
only values that reach the route's 400 are ones that parse but are not
a list — `"npx"`, `{}`, `null` — and for those the status code alone
tells the user nothing about what is wrong with what they typed.

Adds source-level coverage for both forms; the PR changed two JS files
with no test on either.
2026-09-30 17:15:52 +02:00
Léo 51e09a1e32 fix(css): repoint the two stylesheet links the split left behind
Splitting style.css deleted it, and two files outside static/ still
named it:

- scripts/verify_background_research_cards.mjs injected
  `<link rel="stylesheet" href="/static/style.css">` into the page it
  builds. A stylesheet that 404s does not fail — the script kept
  checking card layout against an unstyled page and kept reporting
  pass. It now reads the shell's <link> tags out of index.html, the way
  tests/css_snapshot/capture.mjs already does, so the set cannot drift
  out from under it again.
- tests/css_snapshot/bench.html's hand-open fallback linked the same
  deleted file. Replaced with the ordered set index.html ships.

test_every_stylesheet_referenced_by_shipped_html_exists only walked
static/*.html, which is why neither showed up. Extend it over the bench
page: the bench's whole output is computed styles, so a dead link there
is worth more than an unstyled page nobody looks at.
2026-09-30 17:14:52 +02:00
Léo 5ce2394adb fix(browser): probe pid liveness through the platform-safe helper
_process_is_alive used os.kill(pid, 0). That probe is POSIX-only:
CPython's Windows os.kill calls TerminateProcess(handle, sig) for any
signal other than CTRL_C/CTRL_BREAK, so it terminates the process it is
asked about. This function is only reached when there is no procfs to
read a command line from, which is exactly the macOS and Windows case
the rest of this change exists to handle.

core/platform_compat.py already owns that probe and documents the
hazard; its module docstring asks callers to import from there rather
than spell a POSIX-only call out locally. Delegate to it.

pid_alive answers False where os.kill raises PermissionError — a live
process owned by another user. Both call sites want that reading: the
sweep only unlinks a pid file it wrote itself, and a pid it cannot
confirm is not the daemon it is looking for.
2026-09-30 17:13:29 +02:00
Léo 5fd9114882 test(runtime): pin negative capability wording, test-only
First slice of the runtime regression lane. Drives stream_agent_loop with a
fake model and asserts on the tools the runtime offers, which is its decision
about what the turn may do. No production runtime code is touched and no
benchmark fixture or allowlist is imported.

The fake-model pattern is the one tests/test_tool_policy.py already uses:
patch stream_llm_with_fallback and inspect the tools kwarg.

Measured on lab@c499c01b, negative web wording is only partially detected:

  "Answer from memory only, don't search online."            web tools withheld
  "Summarise what you already know. Do not search the web."  web tools OFFERED
  "No web search please, just tell me what you know..."      web tools OFFERED

The case that holds is a plain regression guard. The two that do not are
xfail(strict=True): they run on every suite, document the target, and fail the
moment the behaviour lands so the marker gets removed rather than lingering.
A positive control keeps the guard from being satisfied by removing the web
tools altogether.
2026-09-30 17:11:10 +02:00
Léo 5e3e153fba fix(auth): swap the API token cache atomically instead of clearing it
`_refresh_token_cache` rebuilt the bearer-token map in two steps, `clear()`
then `update()`. Between them the dict a concurrent reader was already holding
was empty, so a valid token landing in that window found zero candidates and
got a 401. The refresh runs on a worker thread via `to_thread`, so the window
is real rather than theoretical.

Build the new map and rebind the name. The reader dereferences the global once
and then holds a map that is complete — the previous one if it read early, the
new one if it read late, never a half-built one. `app.state._token_cache` is
rebound with it, because it was bound once at startup and would otherwise point
at the abandoned dict.

Ported from public `dev` (`984337b3`), with its test.

The upstream test does not pin the fix: both of its concurrency cases pass
against the pre-fix code, because landing a GIL switch inside a window a few
bytecodes wide does not happen across 100 refreshes. They are kept as written
and `TestRefreshLeavesTheReadersMapAlone` is added next to them, stating the
same invariant at object level — a map a reader already holds is not mutated
by a later refresh — which fails on the pre-fix code without depending on
thread scheduling.
2026-09-30 16:07:12 +02:00