Commit Graph
26 Commits
Author SHA1 Message Date
Alexandre Teixeira b23c6d40b3 fix(effects): require evidence for external completion claims
Effect obligations were consulted only for declared artifacts, and reported
external success could be presented as done. Now, regardless of declared
artifacts:

- the latest effect on any changed file contradicted by a fresh readback
  fails the run (a superseded earlier effect is history, not a contradiction);
- a passing verifier followed by an effect that may have changed state
  without settled evidence is stale (BLOCKED);
- executed external effects that are not VERIFIED cap the decision at
  UNVERIFIED, and the answer always carries server-authored facts for them
  ("reported success; any external change it made was not independently
  verified", "reported failure", "unknown outcome").

The disclosure is structural and does not depend on recognizing the model's
wording. When it is the only change, the model's answer events are released
unchanged and the disclosure follows as one delta (and in round_texts).
Prose filtering is also tightened (remote verbs are mutation claims, an
unnamed "I updated it" cannot borrow the single required artifact, bare
"Done." is a terminal claim beside unverified external effects). A passing
verifier still supports test claims; it never speaks for the external effect.

Replaces the uncommitted attempt that blocked every run with any RUNNING
effect: a background launch with no declared obligations completes
UNVERIFIED.
2026-10-03 00:58:32 +01:00
Alexandre Teixeira 682b44a3ec fix(effects): preserve producer trust boundaries
Result-dictionary keys could set lifecycle state for any producer: a dynamic
or registry tool returning bg_job_id/detached became RUNNING, teardown became
verified cleanup, and timed_out/failure_kind/mutation_attempted/containment
were copied from untrusted results. Facts are now scoped to the producer the
dispatcher actually bound. An unbound tool contributes its exit status alone.
RUNNING requires a bound process producer (and an exact launch reservation
for bg_job_id), cleanup is attested only by a bound process producer, and job
observations and launch settlement only by a bound manage_bg_jobs operation
on exactly one Wave 3-validated job. External/remote-acknowledged facts come
from the captured ExternalResource, not from the result.
2026-10-03 00:58:32 +01:00
Alexandre Teixeira 04da2f04e8 fix(effects): make journal sequencing crash and concurrency safe
Sequence positions were allocated from each EffectLog object's in-memory
counter, so two objects, threads or processes could reuse a position or
settle one effect twice; replay then failed closed for the whole log. Every
append now takes an exclusive flock, merges the durable records other writers
appended (truncating a torn tail a crashed writer left), allocates from that
merged tail, rejects records the merged history makes invalid (a second
settlement, recovery of a claim another writer settled or marked running),
then appends, fsyncs and releases. history() merges others' records under a
shared lock. An incremental consistency index keeps appends O(1).

The first append of each log object fsyncs the log's directory, and every
directory created for it is fsynced in its parent, all under the lock before
the claim returns. A failed write or directory fsync truncates the record
back, so dispatch is refused and nothing unacknowledged is later merged. The
launch index writes and fsyncs a temp file, replaces it, then fsyncs the
directory. flock and directory fsync are POSIX-only and not claimed
elsewhere.
2026-10-03 00:58:32 +01:00
Alexandre Teixeira f79c2aba7e fix(effects): make postconditions prove intended mutations
edit_file and apply_patch update claims asserted only existence (or, in the
uncommitted corrective attempt, any content change), so an unrelated write
could verify them. Each filesystem postcondition is now the exact content the
producer's own transformation writes from the identity-checked pre-state:
edit_file through the extracted pure _edit_file_text (no newline
translation), apply_patch updates through _apply_patch_hunks on the
universal-newline pre-state. An oversized, replaced or undecodable pre-state,
a non-matching hunk, or an underivable write_file body leaves the whole claim
without postconditions (UNVERIFIED) instead of letting derivable targets
verify the operation or falling back to existence.
2026-10-03 00:58:11 +01:00
Alexandre Teixeira 75243fe0b0 fix(runtime): restrict running effects to launch results and index history
Only the native detached launch (bg_job_id) or a bridge's explicit detachment
marks an operation's own work as RUNNING. A listing that reports some other
download/model/job as running settled normally; treating it as running left
the claim pending forever and could block required artifacts.

EffectHistory now indexes outcomes per effect once, removing a cubic scan in
assessment over long run lineages.

Adds adversarial coverage: browser page operations stay fail-closed through
the real dispatcher with effects enabled (no claim, never dispatched),
scheduler task triggers stay unverified admission, and assessment scales.
2026-10-02 20:13:24 +01:00
Alexandre Teixeira 3953ea2444 feat(runtime): settle background launch effects from validated job lifecycle
The background monitor already validates the exact Wave 3 job linkage
(job_from_record + validate_job) before continuing a session. At that point it
now records the job's settlement against the durable launch claim through the
launch-generation index, using typed lifecycle facts from the server-owned
record. Settlement is idempotent across deferred retries, is execution
evidence only, and never reads the delivered output: the injected report stays
attributed content. Failure to record leaves the claim running/unknown and
never blocks the follow-up.
2026-10-02 20:10:46 +01:00
Alexandre Teixeira 8402c388b4 feat(runtime): claim effects before dispatch and gate completion on them
Dispatcher seam: mark_dispatch, which runs inside the live Wave 3 binding
scope immediately before backend invocation, now durably claims a possible
effect before execution_id is assigned. If the claim cannot be persisted the
action stays undispatched and the dispatcher returns BLOCKED; dispatched()
closes the never-awaited coroutine. record_action appends the outcome
(including cancellation/interruption) and admitted-read observations before
the receipt reduction drops producer facts.

Adapters consume only the bound operations the dispatcher admitted:
filesystem bindings give exact scope and predicates (write_file content digest
after fence unwrapping, apply_patch add/delete, edit existence); bash/python
launches have unknown scope with the launch generation as lineage; job kills
scope the exact job and its processes; owned operations scope their exact
revisioned records; external backends are claimed as external and never
verified by acknowledgement; browser session_info yields session lifecycle
observations only, and a page binding is never effect scope. Complete
read_file re-reads the exact bound source to digest it; offset/limit,
truncation, extraction and listings are partial. Background launches stay
RUNNING until an admitted read of the exact job generation (via a durable
launch index, across continuation runs) reports settlement.

Producer seams: typed job lifecycle facts on manage_bg_jobs reads/kills, a
structured timed_out flag on containment timeouts, and mutation_attempted on
in-place write_file/edit_file failures after truncation.

Completion: the existing EvidenceLedger consumes effect assessments through a
single helper used for the decision, ask_user and prose filtering. A required
artifact is unsettled by a later unresolved effect that may have touched it,
a fresh contradicting readback fails the decision, and partial reads no longer
count as artifact validation. Ordinary conversation and read-only turns are
unchanged; no second completion policy is introduced.
2026-10-02 20:09:35 +01:00
Alexandre Teixeira 9c4ed24296 feat(runtime): add durable append-only effect log with interrupted replay
Claims are fsynced to a per-lineage JSONL log before a caller may invoke a
backend; a persistence failure raises EffectPersistenceError (a
ResourceIdentityError) so dispatch fails closed. Outcomes and observations are
appended; nothing is rewritten. Reload validates every record strictly, ignores
only a torn final write, and fails closed on corruption, forgery or hardlink
aliasing. recover_interrupted appends INTERRUPTED/possible-impact outcomes for
claims that never settled and leaves RUNNING background effects alone.

The effect store is added to Wave 3 control-plane paths (prefix check only;
the log itself refuses aliased files), so filesystem tools cannot forge it.
Tests redirect the store to a session tmp directory.
2026-10-02 20:02:31 +01:00
Alexandre Teixeira c57aa7ad5c feat(runtime): recreate Wave 4 effect contracts on exact Wave 3 resources
Recreate (rather than cherry-pick 9012e208) the effect/provenance foundation.
The historical types used opaque string resource keys, a may_have_changed flag
defaulting to no impact, and a single status mixing execution and verification.

Claims, outcomes and observations now reference only typed Wave 3 identities
(filesystem root/inode/ancestor chain, process PID+start token, job generation,
owned revision, external incarnation, browser session incarnation). Browser
page resources are refused. Known no-op is limited to refusal before
invocation; unknown scope stays conservative; verification is derived from
fresh, complete, post-settlement readback checked against an explicit
predicate, and unknown execution with matching state is reported as observed
state without causation. Cleanup is recorded separately from effect outcome.
2026-10-02 19:45:39 +01:00
Alexandre Teixeira ab89e3274a fix(runtime): exclude stale process resources during authority intersection
Catch ResourceIdentityError during intersection so that normal process exit
or background job termination does not crash child authority creation.
Stale or unverifiable observations are conservatively excluded from the
resulting authority while maintaining identity verification and preventing
PID reuse or renewal.
2026-10-02 18:54:09 +01:00
Alexandre Teixeira e175bea752 feat(runtime): bind browser resources to authority 2026-10-02 18:54:09 +01:00
Alexandre Teixeira db41d7e822 feat(runtime): bind process and job resources to authority 2026-10-02 18:54:09 +01:00
Alexandre Teixeira 571f685ad5 feat(runtime): bind remote and owned resources to authority 2026-10-02 03:00:46 +01:00
Alexandre Teixeira ba3e631d58 feat(runtime): bind native filesystem operations to resource identities 2026-10-02 01:19:32 +01:00
Alexandre Teixeira 57fe9946c2 refactor(runtime): one compact-runtime selection rule for route and dispatch
The chat route repeated the compact (clean v3) eligibility decision inline
to prepare the turn's context resolution, while the agent loop dispatched
on the contract stamp set by a separate, later condition. The two could
drift, and already disagreed for a user whose privileges demote the turn
to plain chat: the route prepared a compact resolution that no compact
runtime used.

src/agent_runtime/runtime_selection.py (no imports) now owns the rule:

- uses_compact_preview_runtime(): clean route requested, contract policy
  enabled, agent mode, agent permitted, not an image generation session.
- is_compact_preview_contract() and COMPACT_PREVIEW_MODE for the stamp.

The route evaluates the rule once, before context preparation, where all
of its facts are final (the agent privilege is read through the same
_request_privileges helper the later enforcement uses). That one value
gates the typed context resolution and is the _clean_v3_preview flag that
stamps the contract; inside the agent-contract branch it equals the
previous condition, so stamping behavior is unchanged. The agent loop
dispatches through is_compact_preview_contract(), and the compact runtime's
MODE is the shared constant.

A route-level matrix drives the real agent loop and asserts that route
preparation and compact dispatch agree for compact, escalated, configured
compact/full, regular, TUI, privilege-denied and image-generation turns.
2026-10-01 22:46:08 +01:00
Alexandre Teixeira 0054557027 fix(runtime): resolve compact-turn context once at the chat route
The first checkpoint removed terminal-metrics discovery, but a normal
compact chat turn still ran two context systems: build_chat_context's
legacy untyped lookup (directly or inside maybe_compact) and the typed
resolver inside stream_preview.

Resolve the typed ContextResolution once, at the chat route, before
build_chat_context, using the session's provider credentials. The
predicate mirrors _clean_v3_preview; every input it needs is known at
that point and the native-workspace term cannot veto a requested clean
route. The same object then:

- sizes legacy history shaping in build_chat_context through a new
  maybe_compact(context_length=...) override, so no legacy probe runs;
  an unknown window still shapes with DEFAULT_CONTEXT but gains no
  provenance;
- crosses stream_agent_loop (one new parameter, forwarded only at the
  compact dispatch) into stream_preview, which reuses it and probes only
  for callers that arrive without one or with one bound to another
  route.

ContextResolution now records the endpoint and model it describes
(endpoint URL excluded from repr and metrics). The bare legacy
context_length is never converted into typed evidence.

Credential scoping: origins compare with default ports normalized, an
empty host is never trusted, and the probe client never follows
redirects. Tests cover the configured origin, the server-resolved
Tailscale form, scheme/port/lookalike/userinfo/path origins, redirects,
and secret-free errors, logs and metrics.

The conftest guard now replaces only the resolver's I/O edges (HTTP
client and DNS-capable URL building) instead of the whole probe, and
exposes a context_probe_ledger fixture, so route integration tests run
the real resolver offline and can count metadata requests.
2026-10-01 22:31:21 +01:00
Alexandre Teixeira 837fbfd0ea feat(runtime): resolve compact-runtime context window at turn preparation
The compact (clean v3) runtime had no effective context window: it learned a
limit only reactively from a provider 400/413 and its terminal metrics carried
no context_length. PR #41 addressed the reporting gap by probing provider
metadata between the last model byte and [DONE], unauthenticated, and folded
known-table and endpoint evidence into one "known" flag.

Resolve the window once, before the first model request, instead:

- src/agent_runtime/context_resolution.py adds a typed ContextResolution
  (effective value, evidence class, source, all observations, conflicts,
  provider_io, cached, secret-free probe errors). Evidence classes stay
  distinct: runtime_confirmed (llama.cpp /slots, /props, or a limit the
  provider stated this turn), provider_advertised (models catalog),
  operator_declared (client_runtime_context.model_context_window),
  known_table, unknown (0, never a default).
- Selection is deterministic: runtime beats provider beats table; an
  operator declaration caps measured evidence and replaces weaker evidence.
  Disagreements are recorded as conflicts; a declaration below a measured
  value is a cap, above it a contradiction.
- The provider probe forwards the turn's credentials only to the provider's
  own origin, runs URL resolution off the event loop, is bounded by one
  deadline, never raises, and caches remote results per credential
  fingerprint (shorter TTL for failures; local servers are re-probed).
- stream_preview resolves at preparation (or accepts a supplied resolution),
  seeds the proactive trim budget from it when evidence is not unknown, and
  terminal metrics report only the stored resolution plus any limit the
  provider stated during the turn. Metrics perform no discovery.

src/agent_loop.py and the regular runtime's legacy model_context probe are
unchanged. A conftest guard keeps tests that drive the compact runtime with
placeholder endpoints from performing real DNS/HTTP lookups.
2026-10-01 21:59:53 +01:00
Alexandre Teixeira 9d0257134f feat(runtime): enforce server request authority 2026-10-01 18:35:41 +01:00
Alexandre Teixeira d49071bbec fix: close Wave 1.1 completion-gate audit findings
- Headless consumers (task scheduler, background follow-up) now treat a
  completion-gate final_response as the authoritative answer instead of
  collecting deltas only. A gated replacement no longer leaves scheduled
  output empty, which used to trigger an extra, ungated grace-summary
  model call.
- The scheduler closes the agent stream with contextlib.aclosing, so the
  approval-pause break unwinds the gate's journal and teacher-takeover
  context in its own task. Chained runs no longer inherit a stale
  parent_run_id, and later finalization no longer raises ContextVar
  reset errors.
- On provider error, the completion gate applies the live answer's
  statement filter to persisted round_texts. Diagnostics and the failure
  note survive; claims rejected by the gate cannot reappear on reload.
2026-10-01 14:49:32 +01:00
Alexandre Teixeira f4793696f4 merge: reconcile Wave 1.1 with post-PR40 lab
Merge canonical lab 9557b8d5909eb4a885c3bf49e19a65dd904f8c1d exactly once.
Retain invocation journal ownership and lineage, provider terminal ordering,
teacher handoff, framed DONE handling, and canonical authority/Ajax routing.

Combine dynamic dispatch receipts with lab policy forwarding. Adapt native
shell/patch evidence, explicit TUI verifiers, and artifact recovery presentation.
Refresh generated configuration source links and strengthen adapter regressions.

Validation: focused 2118 passed; Wave 1.1 script 2291 passed; broad runtime
5649 passed; full pytest 11581 passed, 53 skipped, 2 xfailed, 6 subtests passed.
Compileall 1689 Python files; syntax 279 JS and 82 MJS files; diff and
conflict-marker checks passed.
2026-10-01 09:09:55 +01:00
Alexandre Teixeira aedec7d005 fix(runtime): isolate nested invocation ownership 2026-10-01 02:11:53 +01:00
Alexandre Teixeira 4c122de880 fix(runtime): scope completion claims to execution obligations 2026-10-01 01:36:18 +01:00
Alexandre Teixeira 466a6b323a fix(runtime): preserve provider error terminal ordering 2026-10-01 00:33:46 +01:00
Alexandre Teixeira cea8ed297e fix(runtime): reject unobserved verification and preserve safe reasoning 2026-09-26 14:24:29 +01:00
Alexandre Teixeira b241bb3a7b fix(runtime): close completion stream bypass and preserve explanations 2026-09-26 14:13:40 +01:00
Alexandre Teixeira eaa5668baa feat(runtime): record execution evidence and gate completion 2026-09-26 13:48:12 +01:00