Commit Graph
2450 Commits
Author SHA1 Message Date
Alexandre Teixeira 75243fe0b0 fix(runtime): restrict running effects to launch results and index history
Only the native detached launch (bg_job_id) or a bridge's explicit detachment
marks an operation's own work as RUNNING. A listing that reports some other
download/model/job as running settled normally; treating it as running left
the claim pending forever and could block required artifacts.

EffectHistory now indexes outcomes per effect once, removing a cubic scan in
assessment over long run lineages.

Adds adversarial coverage: browser page operations stay fail-closed through
the real dispatcher with effects enabled (no claim, never dispatched),
scheduler task triggers stay unverified admission, and assessment scales.
2026-10-02 20:13:24 +01:00
Alexandre Teixeira 3953ea2444 feat(runtime): settle background launch effects from validated job lifecycle
The background monitor already validates the exact Wave 3 job linkage
(job_from_record + validate_job) before continuing a session. At that point it
now records the job's settlement against the durable launch claim through the
launch-generation index, using typed lifecycle facts from the server-owned
record. Settlement is idempotent across deferred retries, is execution
evidence only, and never reads the delivered output: the injected report stays
attributed content. Failure to record leaves the claim running/unknown and
never blocks the follow-up.
2026-10-02 20:10:46 +01:00
Alexandre Teixeira 8402c388b4 feat(runtime): claim effects before dispatch and gate completion on them
Dispatcher seam: mark_dispatch, which runs inside the live Wave 3 binding
scope immediately before backend invocation, now durably claims a possible
effect before execution_id is assigned. If the claim cannot be persisted the
action stays undispatched and the dispatcher returns BLOCKED; dispatched()
closes the never-awaited coroutine. record_action appends the outcome
(including cancellation/interruption) and admitted-read observations before
the receipt reduction drops producer facts.

Adapters consume only the bound operations the dispatcher admitted:
filesystem bindings give exact scope and predicates (write_file content digest
after fence unwrapping, apply_patch add/delete, edit existence); bash/python
launches have unknown scope with the launch generation as lineage; job kills
scope the exact job and its processes; owned operations scope their exact
revisioned records; external backends are claimed as external and never
verified by acknowledgement; browser session_info yields session lifecycle
observations only, and a page binding is never effect scope. Complete
read_file re-reads the exact bound source to digest it; offset/limit,
truncation, extraction and listings are partial. Background launches stay
RUNNING until an admitted read of the exact job generation (via a durable
launch index, across continuation runs) reports settlement.

Producer seams: typed job lifecycle facts on manage_bg_jobs reads/kills, a
structured timed_out flag on containment timeouts, and mutation_attempted on
in-place write_file/edit_file failures after truncation.

Completion: the existing EvidenceLedger consumes effect assessments through a
single helper used for the decision, ask_user and prose filtering. A required
artifact is unsettled by a later unresolved effect that may have touched it,
a fresh contradicting readback fails the decision, and partial reads no longer
count as artifact validation. Ordinary conversation and read-only turns are
unchanged; no second completion policy is introduced.
2026-10-02 20:09:35 +01:00
Alexandre Teixeira 9c4ed24296 feat(runtime): add durable append-only effect log with interrupted replay
Claims are fsynced to a per-lineage JSONL log before a caller may invoke a
backend; a persistence failure raises EffectPersistenceError (a
ResourceIdentityError) so dispatch fails closed. Outcomes and observations are
appended; nothing is rewritten. Reload validates every record strictly, ignores
only a torn final write, and fails closed on corruption, forgery or hardlink
aliasing. recover_interrupted appends INTERRUPTED/possible-impact outcomes for
claims that never settled and leaves RUNNING background effects alone.

The effect store is added to Wave 3 control-plane paths (prefix check only;
the log itself refuses aliased files), so filesystem tools cannot forge it.
Tests redirect the store to a session tmp directory.
2026-10-02 20:02:31 +01:00
Alexandre Teixeira c57aa7ad5c feat(runtime): recreate Wave 4 effect contracts on exact Wave 3 resources
Recreate (rather than cherry-pick 9012e208) the effect/provenance foundation.
The historical types used opaque string resource keys, a may_have_changed flag
defaulting to no impact, and a single status mixing execution and verification.

Claims, outcomes and observations now reference only typed Wave 3 identities
(filesystem root/inode/ancestor chain, process PID+start token, job generation,
owned revision, external incarnation, browser session incarnation). Browser
page resources are refused. Known no-op is limited to refusal before
invocation; unknown scope stays conservative; verification is derived from
fresh, complete, post-settlement readback checked against an explicit
predicate, and unknown execution with matching state is reported as observed
state without causation. Cleanup is recorded separately from effect outcome.
2026-10-02 19:45:39 +01:00
Alexandre Teixeira 5dce6ae238 docs(config): refresh generated environment reference 2026-10-02 19:07:06 +01:00
Alexandre Teixeira b7182fff54 test(runtime): clean Wave 3 validation warnings 2026-10-02 18:57:53 +01:00
Alexandre Teixeira 7afa6e524d docs(validation): document Wave 3 corrective pass results and test classifications
Record metrics, triage, invariants, and classifications for the Wave 3
corrective implementation pass. Confirms 0 Wave 3 regressions remaining,
with 12358 tests passing across the repository.
2026-10-02 18:54:09 +01:00
Alexandre Teixeira 872888aa4a test(runtime): migrate Wave 3 legacy test suites to resource authority contracts
Migrate 28 legacy test failures to exercise behavior under valid sealed
RequestAuthority, native process reservations, sealed filesystem roots,
and external bridge contexts, or assert fail-closed unscoped behavior.
Preserves all design invariants without weakening production authority.
2026-10-02 18:54:09 +01:00
Alexandre Teixeira 29c31a4b24 test(runtime): verify deterministic refusal of unscoped remote scheduled SSH
Assert that raw scheduled remote SSH workloads without an exact external
backend binding fail closed deterministically with:
'Remote scheduled workload requires an exact external backend binding.'
2026-10-02 18:54:09 +01:00
Alexandre Teixeira 2a3d0c67d6 feat(security): sanitize subprocess environment inheritance and scrub credentials
Replace unscrubbed os.environ inheritance with an explicit allowlist
(_SAFE_SUBPROCESS_VARS) containing only variables necessary for bash/python
execution (PATH, locale, terminal, Python virtualenv/site-packages, Windows
essentials) and regex-based credential scrubbing (_SENSITIVE_PATTERN) to
prevent provider tokens, database URLs, and API keys from leaking into child
processes.
2026-10-02 18:54:09 +01:00
Alexandre Teixeira cc14151d10 fix(browser): clean up owned browser daemons on shutdown after cancellation
Shutdown cleanup must not depend on active record.session capability,
which is cleared on cancellation. Guard cleanup by socket directory
presence so all owned daemons are terminated.
2026-10-02 18:54:09 +01:00
Alexandre Teixeira 896e1f8233 fix(test-isolation): prevent scheduler test from poisoning database globals
Use monkeypatch.setattr for engine, SessionLocal, ScheduledTask, and TaskRun
in _setup_isolated_db to ensure pytest restores real database engine state on
teardown, preventing downstream test failures like no such table: documents.
2026-10-02 18:54:09 +01:00
Alexandre Teixeira ab89e3274a fix(runtime): exclude stale process resources during authority intersection
Catch ResourceIdentityError during intersection so that normal process exit
or background job termination does not crash child authority creation.
Stale or unverifiable observations are conservatively excluded from the
resulting authority while maintaining identity verification and preventing
PID reuse or renewal.
2026-10-02 18:54:09 +01:00
Alexandre Teixeira 525ae76df3 docs(runtime): correct Wave 3 validation xfails 2026-10-02 18:54:09 +01:00
Alexandre Teixeira e175bea752 feat(runtime): bind browser resources to authority 2026-10-02 18:54:09 +01:00
Alexandre Teixeira b648f9ddbe fix(runtime): isolate historical job lookup from PID reuse 2026-10-02 18:54:09 +01:00
Alexandre Teixeira db41d7e822 feat(runtime): bind process and job resources to authority 2026-10-02 18:54:09 +01:00
Alexandre Teixeira 7b8ac6f631 Merge lab after verified process lifecycle 2026-10-02 18:54:09 +01:00
Alexandre Teixeira 1c43f01c50 Merge pull request #59 from o3LL/ci/shard-pytest-suite
ci: run the pytest suite as four parallel shards
2026-10-02 18:48:57 +01:00
Léo 2bda788311 ci: run the pytest suite as four parallel shards
The suite is 12.1k tests in a single CI job — nearly six minutes of pytest
that every push and every PR waits on in one block, on top of a setup step
that already installs npm, Playwright, FFmpeg and bubblewrap. Split it into
four sections the matrix runs in parallel: 352s becomes a 98s longest pole
locally.

Shards partition by test *file*, not by the area_* taxonomy markers. Those
markers do not partition the suite - a file can carry a hand-applied area_*
mark on top of the one conftest derives from its filename, so a marker-based
split would run those tests in more than one section. Assignment is a total
function of the file path instead, so every file lands in exactly one shard
and the four together run every test exactly once.

Sharding deselects rather than narrowing collection, so every test module is
still imported, in the same order, in every shard. The import-time stubbing
in conftest and the session-scoped static server behave identically whether
the suite runs whole or in sections - this suite has known collection-order
coupling and splitting by path would have walked into it.

Balance uses the existing `slow` marker as its weight signal rather than a
committed duration table that would go stale unnoticed. Files pack
heaviest-first into the lightest shard, which is deterministic for a given
file set, so every parallel job computes the same plan from the same commit.

Verified: the four shards together reproduce the full run exactly - 12141
tests selected across the four, and the same 59 failures, 78 skips and 2
xfails, by node ID and not merely by count.
2026-10-02 18:42:10 +02:00
Alexandre Teixeira 08c2dc881e Merge pull request #54 from pewdiepie-archdaemon/feature/runtime-process-lifecycle
refactor(runtime): centralize verified process lifecycle
2026-10-02 10:21:22 +01:00
Alexandre Teixeira dabd32adb7 Merge lab after CI containment baseline 2026-10-02 03:11:21 +01:00
Alexandre Teixeira 87cd4524fa Merge pull request #55 from pewdiepie-archdaemon/fix/ci-functional-bwrap
ci: enable functional containment in pytest
2026-10-02 03:11:14 +01:00
Alexandre Teixeira 571f685ad5 feat(runtime): bind remote and owned resources to authority 2026-10-02 03:00:46 +01:00
Alexandre Teixeira 42ab4acc1b ci: enable functional containment in pytest 2026-10-02 02:44:19 +01:00
Alexandre Teixeira 1bf45b9fed fix(runtime): reconcile lifecycle CI contracts
- Regenerate website/configuration-reference.md: Wave 5B moved the
  ODYSSEUS_BROWSER_SCREENSHOT_DIR read in web_tools.py (3458 -> 3479).
- Give the Chrome sweep regression fixture a real process identity (stat
  start time, boot id, process_ownership.PROC_ROOT). The sweep now signals
  only verified identities; the old cmdline-only fixture borrowed the
  identity of whatever real process held pid 101 on the host, so it passed
  or failed depending on the machine.
- Import pytest in test_workspace_artifact_tool_floor.py: its existing
  bubblewrap capability skip raised NameError on hosts without functional
  namespaces.

No production code changes. Required containment still fails closed.
2026-10-02 02:23:07 +01:00
Alexandre Teixeira b4b5412cda fix(runtime): bind PTY teardown to spawn identity
Capture the PTY leader's ProcessIdentity and its own session group
immediately after spawn, while the child is held unreaped, and drive
teardown from that frozen record instead of re-deriving the group from
proc.pid. Every signal re-verifies the leader: OWNED and still leading
the group signals the group; GONE signals only the recorded group, never
the pid; FOREIGN proves the group's lifetime ended and nothing is
signalled; UNVERIFIABLE or a missing spawn identity signals nothing.
The server's own process group is never recorded or signalled.
2026-10-02 01:33:46 +01:00
Alexandre Teixeira f9aa2818c4 refactor(runtime): centralize verified process lifecycle
Extract the generic process lifecycle layer (src/process_lifecycle.py)
shared by runtime-owned subprocesses: process identity (pid + boot-bound
start token), identity-bound observation, group and pidfd probes, the
TERM -> verify -> KILL -> verify escalation with re-gating before
escalation, identity-scoped sweeps, and the termination receipt.

Containment, the PTY shell, the Cookbook survivor sweep, the browser
lifecycle, web_tools browser cleanup, kill_process_tree and the startup
reaper consume it while keeping their own ownership semantics.

Safety corrections:
- browser membership and identity are bound in one snapshot; no identity
  is recaptured after membership is decided
- web_tools legacy pid-file and profile-match kills signal only verified
  identities; browser CLI groups only while their spawn identity verifies
- Cookbook and legacy-tmux descendant capture bind membership to identity
- PTY teardown never signals the server's own process group
- unverifiable processes are reported, never signalled
2026-10-02 01:27:08 +01:00
Alexandre Teixeira ba3e631d58 feat(runtime): bind native filesystem operations to resource identities 2026-10-02 01:19:32 +01:00
Alexandre Teixeira 7aa891e5e6 Merge pull request #53 from pewdiepie-archdaemon/feature/runtime-containment
feat(runtime): enforce shared native execution containment
2026-10-02 00:44:43 +01:00
Alexandre Teixeira 58d8414867 Merge remote-tracking branch 'preview-upstream/lab' into feature/runtime-containment 2026-10-02 00:27:20 +01:00
Alexandre Teixeira fb49493e77 Merge pull request #52 from pewdiepie-archdaemon/feature/runtime-context-resolution
feat(runtime): resolve compact context once at turn preparation
2026-10-02 00:21:52 +01:00
Alexandre Teixeira 293bd07fa9 Merge remote-tracking branch 'preview-upstream/lab' into feature/runtime-containment 2026-10-02 00:09:38 +01:00
Alexandre Teixeira 037c5f51bf fix(runtime): resolve Wave 3-S audit findings (P2-1, P2-2, P2-3)
- P2-1: reject non-process PID values (None, 0, negative integers) in pid_alive
  without invoking underlying process probe.
- P2-2: truthfully represent production external bridge executions as uncontained,
  external, non-authoritative grants carrying sanitized endpoint metadata.
- P2-3: reject writable_extra overlay bindings over protected system roots and
  their descendants while preserving legitimate scratch destinations.
2026-10-02 00:07:44 +01:00
Alexandre Teixeira cdbcb44cc9 fix(runtime): handle synthetic requests without app scope and update env reference
- Narrowly guard _request_privileges() in routes/chat_routes.py against
  synthetic requests lacking scope['app'] or auth manager state, safely
  returning empty privileges without granting agent privileges.
- Add focused regression test in tests/test_context_resolution_route.py
  verifying that requests without app scope do not crash and cannot gain
  agent privileges or qualify for compact preview runtime.
- Regenerate website/configuration-reference.md mechanically to align with
  current source line numbers.
2026-10-01 23:38:45 +01:00
Alexandre Teixeira 57143a14ba fix(runtime): verify namespace init death and record Wave 3-S validation 2026-10-01 23:17:01 +01:00
Alexandre Teixeira 57fe9946c2 refactor(runtime): one compact-runtime selection rule for route and dispatch
The chat route repeated the compact (clean v3) eligibility decision inline
to prepare the turn's context resolution, while the agent loop dispatched
on the contract stamp set by a separate, later condition. The two could
drift, and already disagreed for a user whose privileges demote the turn
to plain chat: the route prepared a compact resolution that no compact
runtime used.

src/agent_runtime/runtime_selection.py (no imports) now owns the rule:

- uses_compact_preview_runtime(): clean route requested, contract policy
  enabled, agent mode, agent permitted, not an image generation session.
- is_compact_preview_contract() and COMPACT_PREVIEW_MODE for the stamp.

The route evaluates the rule once, before context preparation, where all
of its facts are final (the agent privilege is read through the same
_request_privileges helper the later enforcement uses). That one value
gates the typed context resolution and is the _clean_v3_preview flag that
stamps the contract; inside the agent-contract branch it equals the
previous condition, so stamping behavior is unchanged. The agent loop
dispatches through is_compact_preview_contract(), and the compact runtime's
MODE is the shared constant.

A route-level matrix drives the real agent loop and asserts that route
preparation and compact dispatch agree for compact, escalated, configured
compact/full, regular, TUI, privilege-denied and image-generation turns.
2026-10-01 22:46:08 +01:00
Alexandre Teixeira 0054557027 fix(runtime): resolve compact-turn context once at the chat route
The first checkpoint removed terminal-metrics discovery, but a normal
compact chat turn still ran two context systems: build_chat_context's
legacy untyped lookup (directly or inside maybe_compact) and the typed
resolver inside stream_preview.

Resolve the typed ContextResolution once, at the chat route, before
build_chat_context, using the session's provider credentials. The
predicate mirrors _clean_v3_preview; every input it needs is known at
that point and the native-workspace term cannot veto a requested clean
route. The same object then:

- sizes legacy history shaping in build_chat_context through a new
  maybe_compact(context_length=...) override, so no legacy probe runs;
  an unknown window still shapes with DEFAULT_CONTEXT but gains no
  provenance;
- crosses stream_agent_loop (one new parameter, forwarded only at the
  compact dispatch) into stream_preview, which reuses it and probes only
  for callers that arrive without one or with one bound to another
  route.

ContextResolution now records the endpoint and model it describes
(endpoint URL excluded from repr and metrics). The bare legacy
context_length is never converted into typed evidence.

Credential scoping: origins compare with default ports normalized, an
empty host is never trusted, and the probe client never follows
redirects. Tests cover the configured origin, the server-resolved
Tailscale form, scheme/port/lookalike/userinfo/path origins, redirects,
and secret-free errors, logs and metrics.

The conftest guard now replaces only the resolver's I/O edges (HTTP
client and DNS-capable URL building) instead of the whole probe, and
exposes a context_probe_ledger fixture, so route integration tests run
the real resolver offline and can count metadata requests.
2026-10-01 22:31:21 +01:00
Alexandre Teixeira 0dec375631 fix(runtime): release and report failed background supervisor setup 2026-10-01 22:26:36 +01:00
Alexandre Teixeira 63fe66e85e fix(runtime): enforce functional containment after closing execution bypasses (ODY-141) 2026-10-01 22:11:10 +01:00
Alexandre Teixeira 837fbfd0ea feat(runtime): resolve compact-runtime context window at turn preparation
The compact (clean v3) runtime had no effective context window: it learned a
limit only reactively from a provider 400/413 and its terminal metrics carried
no context_length. PR #41 addressed the reporting gap by probing provider
metadata between the last model byte and [DONE], unauthenticated, and folded
known-table and endpoint evidence into one "known" flag.

Resolve the window once, before the first model request, instead:

- src/agent_runtime/context_resolution.py adds a typed ContextResolution
  (effective value, evidence class, source, all observations, conflicts,
  provider_io, cached, secret-free probe errors). Evidence classes stay
  distinct: runtime_confirmed (llama.cpp /slots, /props, or a limit the
  provider stated this turn), provider_advertised (models catalog),
  operator_declared (client_runtime_context.model_context_window),
  known_table, unknown (0, never a default).
- Selection is deterministic: runtime beats provider beats table; an
  operator declaration caps measured evidence and replaces weaker evidence.
  Disagreements are recorded as conflicts; a declaration below a measured
  value is a cap, above it a contradiction.
- The provider probe forwards the turn's credentials only to the provider's
  own origin, runs URL resolution off the event loop, is bounded by one
  deadline, never raises, and caches remote results per credential
  fingerprint (shorter TTL for failures; local servers are re-probed).
- stream_preview resolves at preparation (or accepts a supplied resolution),
  seeds the proactive trim budget from it when evidence is not unknown, and
  terminal metrics report only the stored resolution plus any limit the
  provider stated during the turn. Metrics perform no discovery.

src/agent_loop.py and the regular runtime's legacy model_context probe are
unchanged. A conftest guard keeps tests that drive the compact runtime with
placeholder endpoints from performing real DNS/HTTP lookups.
2026-10-01 21:59:53 +01:00
Alexandre Teixeira 5c3fbb4132 Merge pull request #51 from pewdiepie-archdaemon/feature/browser-lifecycle
feat(browser): add deterministic browser lifecycle
2026-10-01 21:51:06 +01:00
Alexandre Teixeira 4efb85ee33 fix(runtime): replace tmux pane capture with explicit bounded output (ODY-150) 2026-10-01 21:43:17 +01:00
Alexandre Teixeira 85cfe59b15 fix(runtime): retire and reap owned agent tmux sessions (ODY-147) 2026-10-01 21:34:23 +01:00
Alexandre Teixeira 7892f7f650 fix(runtime): contain and supervise detached Bash jobs (ODY-145) 2026-10-01 21:23:21 +01:00
Alexandre Teixeira ba7c733741 test(runtime): verify a true sibling is hidden from Python namespace 2026-10-01 21:05:42 +01:00
Alexandre Teixeira bf0623a77b fix(runtime): contain every native Python execution (ODY-143) 2026-10-01 21:04:12 +01:00
Alexandre Teixeira 255bff1f76 fix(browser): derive batch navigation outcome from command rows
A batch whose open succeeded but whose later command failed was recorded
as a failed navigation, so a following observation was wrongly labelled
stale. Use the per-command rows; when the outcome cannot be determined,
treat the page as unknown instead of claiming either result.
2026-10-01 21:03:00 +01:00
Alexandre Teixeira 9bb2424e65 fix(runtime): fold native execution into shared containment (ODY-152) 2026-10-01 20:59:49 +01:00