Merge pull request #43 from pewdiepie-archdaemon/feature/agent-runtime-wave-1-1

feat(runtime): complete Agent Runtime Wave 1.1
This commit is contained in:
Alexandre Teixeira
2026-10-01 15:22:06 +01:00
committed by GitHub
33 changed files with 3510 additions and 319 deletions
@@ -0,0 +1,110 @@
# Frozen benchmark comparison contract
This protocol does not authorize a multi-hour confirmation campaign. The first
full baseline/candidate screening pair follows the six implementation gates.
Use its duration and variance to propose confirmation work for user approval.
No candidate performance result is available yet.
## Identities and experimental unit
- Historical campaign: `LOCAL-BASELINE-QWEN35-9B-FROZEN-01`; never overwrite,
resume with different source, or pool it silently with fresh measurements.
- Frozen benchmark: `9047e3b47eaf1170c00e915343f5ba3864e0deb8`; prompts,
fixtures, policies, acceptance and scoring remain unchanged.
- Lab starting source: `7b4469299c3b45d062ce80bc5bb16eb69a7aeae1`. Its production
source bytes match those used by the historical campaign. Fresh comparison
still uses this exact revision under the same reviewed harness as the candidate.
- The separate source-selection harness lane currently has provisional commit
`c4d2ea035183c7092146701ece99a52355ec0f00`; independent review may require a
correction. Freeze the resulting reviewed harness revision before screening.
Never include harness changes in the production PR.
- Candidate source is frozen only after all deterministic and review gates pass.
Every run records its actual selected worktree, commit, production byte hash,
mounted-byte proof, harness hash, model and effective configuration identities.
- Model remains local Qwen3.5-9B Q4_K_M, context 16384, effective temperature 1.0,
one llama.cpp slot at `127.0.0.1:8000`, outer-sandbox, and the recorded pinned
Chroma image. Record model file identity, llama.cpp build, request parameters
and effective sampling; a server default is not proof of request sampling.
The experimental unit is one scenario execution, not a model round or a token.
All ten scenarios belong in every full campaign, including pre-inference
rejections and infrastructure failures. Source revision is the treatment.
Comparison cohorts require all other relevant frozen identities to agree.
## Metrics and denominators
| Metric | Evidence and interpretation |
|---|---|
| Task success | Frozen acceptance/scoring outcome per scenario; report passes out of all ten, scored failures, pre-inference rejections and unscored infrastructure outcomes separately. |
| Scope compliance | Actual filesystem deltas, dispatch receipts and security observations. Report allowed changes, unauthorized changes/effects, and attempted versus executed prohibited operations. A denial is not an unauthorized effect. |
| Tool dispatch | Proposed calls, normalized operations, authorization decisions, backend invocations and observed/reported outcomes as separate counts. Tool selection or `tool_start` alone does not prove an operation happened. |
| Verified completion | Current authoritative artifact and verifier evidence at publication time, plus independent acceptance. Record incomplete results and unsupported completion claims separately; acceptance passing does not retroactively ground an earlier claim. |
| Recovery | Distinct diagnostic failure, denial, invalid arguments, missing resource, browser timeout, backend and infrastructure categories. Count transitions to useful new evidence and recovery to success; repeated plans are not productive work. |
| Measured usage | Actual provider input/output usage for every request, retry and helper call, identified by request and source revision. Preserve missing usage as missing. |
| Estimated usage | Separate estimated input/output counts with estimator/version and coverage. Never label estimates as measured or silently combine the two into a supposedly measured total. |
| Context | Prepared input estimate and, where provided, actual per-request input usage; peak across requests, distribution, configured context capacity and output reservation. Cumulative round input is a cost metric, not a context window. |
| Useful work per round | Artifact-version changes, new successful observations, newly satisfied obligations and fresh verifier results per actual provider round. Show raw counts and state transitions; do not optimize an opaque weighted score. |
| Latency | End-to-end scenario time, provider first-token time, first visible checked answer, provider generation time, tool stage durations, verification and cleanup. Report per-task paired differences and aggregate sum/median; retain timeout censoring. |
| Browser/process reliability | Actual browser stages and extraction; owned process launch/readiness/observation/shutdown receipts; bounded recovery and cleanup. Distinguish useful success from an available tool schema. |
| Infrastructure reliability | Startup/probe/model/backend errors, timeouts, port conflicts, leaks and incomplete artifact capture. Report every occurrence and any separately identified replacement trial. |
Preserve task success and security as primary outcomes. Lower tokens caused by
early rejection, omitted work or weaker verification are not efficiency gains.
Show token/latency totals for all assigned tasks and, separately, the overlapping
successful tasks. Label this conditional subset explicitly; it is not evidence
of whole-campaign improvement. A candidate that solves more work may legitimately
consume more total tokens. Never use one successful subset to conceal regressions.
## Initial screening procedure
1. Verify clean committed production sources and the reviewed harness. Recheck
protected historical evidence and fixture/prompt/acceptance identities.
2. Use new campaign IDs and a separate development results root. Pin the same
harness, model, context, sampling, policies, scenario order and timeouts for
baseline and candidate. Keep the original campaign/results directories intact.
3. Run sequentially on the single local slot. Record external load and service
health sufficient to identify infrastructure interference. Do not modify host
security policy or kill unrelated processes to improve a measurement.
4. Capture all raw requests/events/tool traces, usage provenance, acceptance,
artifact deltas, cleanup and identity proofs. Hash the resulting artifacts.
5. Validate schemas and identity matches before comparing outcomes. Report
mismatches as invalid comparisons; do not repair historical records in place.
6. Inspect every changed outcome and apparent efficiency gain against traces.
In particular audit AR-005, AR-006 and AR-009 for preserved useful behavior,
and assess AR-001/002/003/004/007/008/010 against their actual failure modes.
7. Report this as one stochastic screening pair, with no statistical superiority
claim. If regressions appear, identify and correct production causes, freeze
a new revision and use new campaign IDs for the next screening.
## Proposed repeated paired confirmation
After screening, request approval for a predeclared number of complete paired
campaigns with a wall-time estimate based on observed durations. A starting
proposal is five pairs for variance estimation; a superiority claim may require
more. Do not choose a final sample size based on which result looks favorable.
Pair each scenario across baseline/candidate under identical conditions. Balance
the order of complete campaigns (baseline-first and candidate-first), randomize
the planned order before execution and record it. Keep the frozen within-campaign
scenario order unless the reviewed comparison contract explicitly establishes an
identical alternate order for both treatments. Do not mix source revisions within
a comparison or resume an old campaign after source changes.
If a seed is supported and verifiably reaches every actual provider request, use
the same scheduled seed within each pair and different seeds across pairs.
Otherwise record the trials as unseeded; equal task prompts still create matched
workloads but do not imply matched stochastic trajectories. Seed support must be
verified from actual request evidence, not assumed from a CLI label.
Report scenario-level results and paired campaign-level differences. For success,
show discordant pairs and an exact paired binary analysis where its assumptions
hold; avoid treating all rounds or repeated runs of one scenario as independent
tasks. For aggregate estimates, account for repeated observations within scenarios
and show uncertainty intervals together with raw paired results. With only ten
fixed scenarios, conclusions apply to this benchmark, not general agent ability.
Show medians and paired differences for skewed token/latency data; include timeouts
and infrastructure failures explicitly. Predeclare any replacement-run policy,
retain every failed attempt and report results both with and without replacements.
Security invariants, truthful completion and demonstrated regressions remain
release gates regardless of an aggregate improvement or confidence interval.
@@ -0,0 +1,153 @@
# Wave 1.1 final post-PR40 reconciliation
This is the one-time local reconciliation of completed Wave 1.1 with the
authoritative post-PR40 lab commit. It does not start another runtime wave.
## Verified starting state
- Wave branch: `feature/agent-runtime-wave-1-1`.
- Original Wave HEAD: `63457367aeed431b2c48967988259e5861f19916`, clean.
- Canonical branch: `lab`.
- Canonical HEAD: `9557b8d5909eb4a885c3bf49e19a65dd904f8c1d`, clean.
- Merge base: `f0761641a12b63e401960f596d3d1be8fc90fbea`.
- Divergence: 10 Wave-only commits and 47 lab-only commits.
- Changed-file overlap: `src/agent_loop.py`, `src/tool_execution.py`,
`tests/test_tool_policy.py`, and `tests/README.md`.
The Wave-only commits were `d57d5c58`, `dfeab64a`, `ae2445d6`, `7d84f3fe`,
`1470dbb2`, `32830918`, `ba29afb9`, `bdfcbc0a`, `70cbaf81`, and `63457367`.
Their completed behavior is retained. The canonical worktree is read-only;
the exact canonical SHA was merged once with `--no-ff --no-commit`.
## Semantic integration
The only textual conflict was in `src/tool_execution.py`, where Wave 1.1
wrapped dynamic dispatch with `dispatched(...)` and lab added `disabled_tools`
and `tool_policy` forwarding. The resolution retains both inside the wrapper.
Lab's new owner-aware image-generation dispatch also receives that wrapper.
The image regression checks that explicit denial never invokes the backend,
actual dispatch has an execution identity, and a backend without an explicit
exit code does not manufacture an authoritative success receipt.
Broad validation exposed narrow adapter incompatibilities beyond the textual
conflict. Native host-shell JSON now uses the same decoded command classification
as journal evidence. The exact existing TUI interpreter-selection string is
shared with the evidence parser: a following foreground verifier keeps its
exit status, while generic conditional discovery, help/collection modes,
variable arguments, and status-masking tails remain insufficient test proof.
The generated interpreter-selection command itself is unchanged.
The generated environment reference is refreshed with the canonical generator
so its source-location links match the reconciled code.
Structured native patch arguments retain artifact targets. A pre-edit
inspection cannot invalidate a later passing executable verifier, but still
cannot verify the edited artifact by itself; failed post-edit inspections
remain failures. Artifact recovery's terminal round-text revisions retract buffered rejected drafts
before presentation; their replacement prose is still gated by journal
evidence. Explicit final-response events retain precedence, safe reasoning
survives, and provider-error partials and diagnostics retain their ordering.
Existing subprocess doubles now carry PIDs. Execution simulations use the
existing receipt-aware test helper. Contract tests assert the additional
completion-decision event and retain their no-inference/no-execution spies.
TUI tests retain tool order, retry behavior, and positive explicit-verifier
coverage while additionally rejecting completion from an opaque fallback.
Round-control fixtures explicitly fail unconfigured direct-provider synthesis
instead of contacting their fake endpoint, and supply the synthetic context
window while retaining real compaction logic. Conversational round provenance is
preserved outside artifact recovery.
The agent-loop changes merged automatically: lab's weather relevance and
policy-gated browser fallback coexist with Wave's action receipts, completion
gate, and deferred teacher handoff. The fallback dispatcher runs inside the
current invocation's journal. No generic tool floor was restored.
`src/agent_runs.py`, `routes/chat_routes.py`, `static/js/chat.js`,
`static/js/chatRenderer.js`, `src/tool_policy.py`, `src/tool_capabilities.py`,
`src/turn_contract.py`, `src/model_profiles.py`, and
`src/clean_agent_preview.py` retain the exact canonical lab content.
## Identity audit
These classifications describe every relevant identity use across the
detached-run manager, chat routes/browser consumers, completion gate, journal,
teacher handoff, and existing server-owned security provenance.
| Class | Uses and boundary |
| --- | --- |
| 1. Live/detached stream-run identity | `agent_runs._Run.run_id`, `get_run_id`, and the chat response's `X-Odysseus-Run-Id` identify the detached stream. The browser's `_streamRunIds` is populated from the response header. |
| 2. Stop/resume/replay identity | `expected_run_id` in `stop` and `request_finish`, route request headers, `_postExactStop`, the finish-editor request, `streamRunId`, and `resumeRunId` refer to that same detached stream. `subscribe` binds the exact `_Run` object returned by start/resume. |
| 3. Stream metrics/cost identity | `_metricsCostRecordId` uses the header-derived stream ID plus `primary`/`teacher`; `metrics._costRecordId` and the cost renderer's local `runId` refer to this accounting key. Neither uses terminal metadata's journal `run_id`. |
| 4. Logical nested invocation identity | `ActionJournal.run_id` is generated per completion-gated invocation. `action_id` is derived from it. The completion gate's terminal metadata `run_id` identifies this logical invocation. Existing `ToolRunSecurityContext.run_id` and `origin_run_id` values identify separate server-owned invocation/skill provenance operations; they are neither stream IDs nor journal lineage. |
| 5. ActionJournal parent/child identity | `ActionJournal.parent_run_id`, the gate's parent lookup, `_parent_run_id`, `request_teacher_takeover`'s captured parent ID, and `run_teacher_inline(parent_run_id=...)` link journal invocations. The completion metadata's `parent_run_id` preserves that lineage. |
No invocation ID is passed to stream stop/finish/replay APIs. No stream ID is
inserted into journal lineage. A new detached-stream regression creates nested
gates, rejects both journal IDs at stop/finish, accepts the stream ID for finish,
and verifies identical replay and unchanged journal metadata.
## Runtime invariants and final lab behavior
Every gated invocation creates a distinct journal, including children using
the same workspace. Journal and action bindings restore on normal unwind,
exception, cancellation, and generator close. Child awaiting/exhausted/error
state cannot rewrite the parent's completion decision or receipts.
The teacher adapter runs after the student gate closes. It forwards the parent
turn contract, tool policy, disabled tools, plan, client runtime context, and
external-untrusted-context restriction. Teacher execution receives a new
journal whose parent is the student invocation. Inner terminal frames are
consumed; only the outer adapter emits final termination. Exact framed
`data: [DONE]` events are distinguished from ordinary content containing the
literal marker.
Provider failures retain live events, then safe partial content when present,
then a non-completing decision, terminal metadata, and the original error last,
without DONE. A bare error remains a bare error. Completion gating does not
add provider calls or turn missing evidence into extra provider rounds.
Lab's server-owned authority remains narrower than inventory or availability.
Transcription, OCR, tasks, browser fallback, request-specific capability
selection, compact contracts, and provider-compatible tool choice retain the
canonical implementation. Model ID `Ajax` selects the Odysseus compact profile;
its selected schema boundary survives compatible `auto` tool choice, explicit
no-tools remains explicit, and transport remains OpenAI-compatible. No
benchmark-runner code was independently edited or executed.
## Validation records
The current requirements were installed in an isolated environment under this
worktree's ignored `.cache/wave1-1-reconciliation` directory. The shell's
unrelated `python` environment was not used for the accepted validation.
Canonical full pytest uses the repository's default data directory and allows
dotenv loading so research-path and setup tests can exercise their own fixtures;
the focused Wave script retains its explicit runtime isolation settings.
Optional live Ajax tests retain their opt-in skips; no live model or benchmark
run is part of this reconciliation.
- [Focused tests](validation/wave-1-1-reconciliation-focused.txt)
- [Wave 1.1 validation script](validation/wave-1-1-reconciliation-wave-validation.txt)
- [Broad affected runtime suite](validation/wave-1-1-reconciliation-broad.txt)
- [Canonical full pytest](validation/wave-1-1-reconciliation-pytest.txt)
- [Compileall, JS/MJS syntax, diff checks, and conflict-marker scan](validation/wave-1-1-reconciliation-gates.txt)
The focused records include the final relevant rerun after the reconciliation
audit was written. Full pytest and canonical static gates run afterward. The
local merge is committed only after the required checks pass. No push, PR,
deployment, or later-wave work is authorized by this reconciliation.
## Maestrum limitations encountered
The normal read-only pre-merge comparison stalled without a completion or
failure payload; its execution cell was terminated and the investigation was
not retried. Exact-path inspection proceeded using `local_only` with
`scope_mode="worktree"`.
The Context Firewall rejected an unbounded `git diff --cached --check` command
and withheld raw log output after the inspection allowance was exhausted.
Requests for ignored `.log` files were rejected with
`scope_rejected: ignored_by_git`. Unignored `.txt` validation records were
subsequently admitted by exact path. Canonical checks themselves run as
validation operations and record their exit status in the admitted gate log.
No epoch waiting or alternative worker mechanism was used.
@@ -0,0 +1,7 @@
Python compileall: 1689 tracked files; 0 failures
JS syntax: 279 tracked files; 0 failures
MJS syntax: 82 tracked files; 0 failures
git diff --check: exit 0
git diff --cached --check: exit 0
git diff HEAD --check: exit 0
Conflict-marker scan: 2377 tracked files; 0 matches
+28
View File
@@ -0,0 +1,28 @@
#!/usr/bin/env bash
# Focused runtime gate; no model inference or benchmark fixture access.
set -euo pipefail
cd "$(dirname "${BASH_SOURCE[0]}")/.."
export ODYSSEUS_TEST_STATIC_PORT=0
export PYTHONDONTWRITEBYTECODE=1
export PYTHON_DOTENV_DISABLED=1
export ODYSSEUS_DATA_DIR="${ODYSSEUS_DATA_DIR:-/tmp/odysseus-runtime-decomposition-test-state}"
exec "${ODYSSEUS_TEST_PYTHON:-python3}" -m pytest -q -p no:cacheprovider \
tests/test_runtime_evidence_contract.py tests/test_agent_evidence.py \
tests/test_completion_boundary.py \
tests/test_nested_invocation_ownership.py \
tests/test_agent_evidence_loop.py tests/test_agent_render_ownership.py \
tests/test_agent_runs_terminal_order.py tests/test_agent_loop.py \
tests/test_tool_task_cancelled_on_disconnect.py tests/test_turn_contract.py \
tests/test_agent_turn_contract_boundaries.py tests/test_loop_breaker_runaway.py \
tests/test_chat_route_tool_policy.py tests/test_agent_runtime_context.py \
tests/test_external_context_tool_gate.py tests/test_workspace_confine.py \
tests/test_private_browser_tool.py tests/test_bg_jobs_store.py tests/test_bg_job_tools.py \
tests/test_context_budget.py tests/test_context_compactor.py \
tests/test_context_compactor_nonstring.py tests/test_generation_budget.py \
tests/test_foreground_model_routing.py tests/test_tool_policy.py \
tests/test_execution_bridge.py tests/test_tool_approvals.py \
tests/test_tool_approval_single_action_scope.py tests/test_tool_approval_task_scope.py \
tests/test_mcp_text_error_normalization.py tests/test_mcp_email_search_error_transport.py \
tests/test_tool_path_confinement.py tests/test_workspace_artifact_tool_floor.py \
tests/test_native_unattended_workspace_floor.py tests/test_misfenced_read_file_tool_call.py \
tests/test_builtin_mcp_pythonpath.py tests/test_python_tool_import_paths.py "$@"
+180 -40
View File
@@ -9,6 +9,7 @@ from dataclasses import asdict, dataclass, field
from enum import Enum
from pathlib import Path
from typing import Any, Iterable, Mapping, Sequence
from src.agent_runtime.identity import artifact_identity, artifact_version, executable_words, is_test_command, is_validation_command
def workspace_artifact_is_usable(path: Path) -> bool:
@@ -99,6 +100,10 @@ class EvidenceEvent:
command_sha256: str = ""
output_sha256: str = ""
detail: str = ""
action_id: str = ""
execution_id: str = ""
artifact_id: str = ""
verification_id: str = ""
def to_dict(self) -> dict[str, Any]:
data = asdict(self)
@@ -195,12 +200,12 @@ _VALIDATION_COMMAND_RE = re.compile(
def command_is_validation(command: str) -> bool:
"""Return whether a shell command provides executable verification evidence."""
value = str(command or "")
return bool(_TEST_COMMAND_RE.search(value) or _VALIDATION_COMMAND_RE.search(value))
return is_validation_command(_command_text(value))
def command_is_test(command: str) -> bool:
"""Return whether a shell command executes a recognized test runner."""
return bool(_TEST_COMMAND_RE.search(str(command or "")))
return is_test_command(_command_text(str(command or "")))
def _clean_path(value: str) -> str:
@@ -275,6 +280,40 @@ def _known_input_is_explicit_mutation_target(instruction: str, path: str) -> boo
)
def _unquoted_statements(text: str) -> Iterable[tuple[str, str]]:
"""Yield original statements and their reportable prose, with quotes masked.
Mask before splitting so punctuation inside an example cannot change the
scope of the surrounding sentence. Inline code identifiers stay visible.
"""
def mask(match: re.Match[str]) -> str:
value = match.group()
# Quotation marks around an artifact identify a target, rather than
# quote a report. Keep that target available for exact path matching.
if value[0] in {'"', "'"} and re.fullmatch(_ARTIFACT_PATH, value[1:-1]):
return ' ' + value[1:-1] + ' '
return re.sub(r'[^\n]', ' ', value)
masked = re.sub(
r'```[\s\S]*?```|~~~[\s\S]*?~~~|"[^"\n]*"|(?<!\w)\'[^\'\n]*\'(?!\w)',
mask, text,
)
start = 0
for boundary in re.finditer(r'(?<=[.!?;])(?=\s)|(?<=\n)', masked):
end = boundary.start()
if end > start:
yield text[start:end], masked[start:end].replace('`', '')
start = end
if start < len(text):
yield text[start:], masked[start:].replace('`', '')
def _execution_obligation(requirements: CompletionRequirements) -> bool:
"""A derived view of the existing contract, never a separate declaration."""
return bool(requirements.required_artifacts or requirements.verifier_required
or requirements.executable_verifier_available or requirements.verifier_commands)
def infer_completion_requirements(
instruction: str,
*,
@@ -284,7 +323,15 @@ def infer_completion_requirements(
) -> CompletionRequirements:
"""Infer only explicitly requested output/edit paths from an instruction."""
text = str(instruction or "")
# Explanations can contain imperative examples. Their embedded actions
# are not requests to execute those actions. Keep independent requests in
# other statements, and keep explicitly supplied verifier requirements.
explanatory_request = re.compile(
r'^\s*(?:please\s+|(?:can|could|would)\s+you\s+)?'
r'(?:explain|describe|summari[sz]e|teach|discuss|'
r'show\s+(?:me\s+)?(?:an?\s+)?example|how\b)', re.I)
text = ''.join(scoped for _, scoped in _unquoted_statements(str(instruction or ''))
if not explanatory_request.search(scoped))
paths: list[str] = []
for pattern in (
_ARTIFACT_REQUEST_RE,
@@ -348,18 +395,23 @@ def infer_completion_requirements(
for command in verifier_commands
if str(command or "").strip()
))
explicit_test_request = re.search(
r'(?:^|[.;\n]|\b(?:and|then))\s*'
r'(?:please\s+|(?:can|could|would)\s+you\s+)?'
r'(?:run|execute)\s+(?:(?:the|all|a|full)\s+)*'
r'(?:tests?\b|test\s+suite\b|pytest\b|unittest\b|npm\s+test\b)', text, re.I)
verifier_required = executable_verifier_available or bool(cleaned_verifier_commands) or bool(
re.search(
explicit_test_request or (paths and re.search(
r"\b(?:then|after(?:wards)?|and)\b[^\n]{0,100}\b(?:test|verify|check|validate)\b",
str(instruction or ""),
text,
re.IGNORECASE,
)
))
)
return CompletionRequirements(
required_artifacts=tuple(paths),
verifier_required=verifier_required,
executable_verifier_available=(
executable_verifier_available or bool(cleaned_verifier_commands)
executable_verifier_available or bool(cleaned_verifier_commands) or bool(explicit_test_request)
),
verifier_commands=cleaned_verifier_commands,
)
@@ -441,18 +493,12 @@ def _path_is_mentioned(command: str, required_path: str) -> bool:
return path in command or Path(path).name in command
def _artifact_path_matches_required(artifact_path: str, required_path: str) -> bool:
artifact = _clean_path(artifact_path)
def _artifact_path_matches_required(artifact_path: str, required_path: str, workspace: str = "") -> bool:
artifact = str(artifact_path or '').strip()
required = _clean_path(required_path)
if not artifact or not required:
return False
if artifact == required:
return True
# Absolute requirements are exact output contracts; same basename in a
# different directory is not enough.
if artifact.startswith("/") or required.startswith("/"):
return False
return Path(artifact).name == Path(required).name
return artifact_identity(artifact, workspace) == artifact_identity(required, workspace)
def _explicit_tool_paths(tool: str, command: str) -> list[str]:
@@ -462,21 +508,28 @@ def _explicit_tool_paths(tool: str, command: str) -> list[str]:
except (TypeError, json.JSONDecodeError):
args = None
if isinstance(args, Mapping):
path = _clean_path(str(args.get("path") or ""))
path = str(args.get("path") or "").strip()
return [path] if path else []
# Keep compatibility with the legacy ``path\ncontent`` transport.
path = _clean_path(str(command or "").splitlines()[0] if command else "")
path = (str(command or "").splitlines()[0] if command else "").strip()
return [path] if path else []
if tool == "edit_file":
try:
args = json.loads(command or "{}")
except (TypeError, json.JSONDecodeError):
return []
path = _clean_path(str(args.get("path") or "")) if isinstance(args, dict) else ""
path = str(args.get("path") or "").strip() if isinstance(args, dict) else ""
return [path] if path else []
if tool == "apply_patch":
try:
args = json.loads(command or "{}")
except (TypeError, json.JSONDecodeError):
args = None
if isinstance(args, Mapping):
patch = args.get("patch")
command = patch if isinstance(patch, str) else ""
return [
_clean_path(match.group(1))
match.group(1).strip()
for match in re.finditer(r"^\*\*\* (?:Add|Update|Delete) File:\s*(.+)$", command or "", re.MULTILINE)
if _clean_path(match.group(1))
]
@@ -549,13 +602,13 @@ def _command_text(value: str) -> str:
def _matches_declared_verifier(command: str, expected: Sequence[str]) -> bool:
actual = " ".join(_command_text(command).split())
actual = executable_words(_command_text(command))
if not actual:
return False
return any(
normalized == actual or normalized in actual
normalized == actual
for item in expected
if (normalized := " ".join(str(item or "").split()))
if (normalized := executable_words(str(item or "")))
)
@@ -574,6 +627,11 @@ class EvidenceLedger:
def __init__(self, requirements: CompletionRequirements | None = None) -> None:
self.requirements = requirements or CompletionRequirements()
self.events: list[EvidenceEvent] = []
self._verification_versions: dict[str, str] = {}
self._verification_versions_captured = False
# Retain receipt command identity privately for presentation matching;
# model prose and client dictionaries never populate this evidence.
self._verifier_commands: dict[str, tuple[str, ...]] = {}
@classmethod
def from_tool_events(
@@ -611,6 +669,10 @@ class EvidenceLedger:
"command_sha256": _digest(command),
"output_sha256": _digest(output),
}
action_id = str(source.get('action_id') or '')
execution_id = str(source.get('execution_id') or '')
if action_id:
payload.update(action_id=action_id, execution_id=execution_id)
evidence = EvidenceEvent(
event_id=_event_id(payload, len(self.events)),
kind=kind,
@@ -623,6 +685,11 @@ class EvidenceLedger:
command_sha256=payload["command_sha256"],
output_sha256=payload["output_sha256"],
detail=detail,
action_id=action_id,
execution_id=execution_id,
artifact_id=artifact_identity(artifact_path, self.requirements.workspace_root) if artifact_path else '',
verification_id=('verification-' + _event_id(payload, len(self.events)))
if kind in {EvidenceKind.VERIFIER_RESULT, EvidenceKind.ARTIFACT_VALIDATION} else '',
)
self.events.append(evidence)
return evidence
@@ -631,8 +698,12 @@ class EvidenceLedger:
tool = str(event.get("tool") or "")
command = str(event.get("command") or "")
exit_code = event.get("exit_code")
authoritative = isinstance(exit_code, int) and not isinstance(exit_code, bool)
success = authoritative and exit_code == 0
authoritative = (
isinstance(exit_code, int) and not isinstance(exit_code, bool)
and not event.get("blocked") and not event.get("approval_required")
and event.get("execution_attempted") is not False
)
success = authoritative and exit_code == 0 and not event.get('error')
if not authoritative:
success = not bool(event.get("error"))
self._append(
@@ -644,7 +715,11 @@ class EvidenceLedger:
explicit_paths = _explicit_tool_paths(tool, command)
mutation_paths = list(explicit_paths)
if command_has_mutation_effect(command) and tool not in {
observed_changes = event.get('artifact_changes')
if isinstance(observed_changes, list) and tool in {'bash', 'python', 'host_shell'}:
mutation_paths.extend(path for path in self.requirements.required_artifacts
if artifact_identity(path, self.requirements.workspace_root) in observed_changes)
elif command_has_mutation_effect(command) and tool not in {
"write_file",
"edit_file",
"apply_patch",
@@ -657,7 +732,7 @@ class EvidenceLedger:
)
seen_paths: set[str] = set()
for path in mutation_paths:
path = _clean_path(path)
path = str(path or '').strip()
if not path or path in seen_paths:
continue
seen_paths.add(path)
@@ -675,12 +750,12 @@ class EvidenceLedger:
except (TypeError, json.JSONDecodeError):
read_args = None
read_path = (
_clean_path(str(read_args.get("path") or ""))
str(read_args.get("path") or "").strip()
if isinstance(read_args, Mapping)
else ""
else command.strip() if read_args is None else ""
)
if read_path and any(
_artifact_path_matches_required(read_path, required)
_artifact_path_matches_required(read_path, required, self.requirements.workspace_root)
for required in self.requirements.required_artifacts
):
self._append(
@@ -692,18 +767,23 @@ class EvidenceLedger:
detail="post-write artifact inspection",
)
if _TEST_COMMAND_RE.search(_command_text(command)) or _matches_declared_verifier(
if tool in {"bash", "host_shell"} and (command_is_test(_command_text(command)) or _matches_declared_verifier(
command,
self.requirements.verifier_commands,
):
self._append(
)):
if authoritative:
versions = event.get('artifact_versions')
self._verification_versions = dict(versions) if isinstance(versions, Mapping) else {}
self._verification_versions_captured = isinstance(versions, Mapping)
verifier = self._append(
kind=EvidenceKind.VERIFIER_RESULT,
success=success,
authoritative=authoritative,
source=event,
detail="executable test/verifier command",
)
elif _VALIDATION_COMMAND_RE.search(command) and not mutation_paths:
self._verifier_commands[verifier.event_id] = executable_words(_command_text(command))
elif tool in {"bash", "host_shell"} and is_validation_command(command) and not mutation_paths:
for path in self.requirements.required_artifacts:
if _path_is_mentioned(command, path):
self._append(
@@ -714,6 +794,41 @@ class EvidenceLedger:
artifact_path=path,
)
def _supports_verifier_claim(self, identities: Sequence[str] = (), paths: Sequence[str] = ()) -> bool:
"""Only the current passing verifier may support its named runner."""
if self.evaluate().status != CompletionStatus.VERIFIED:
return False
latest = next((event for event in reversed(self.events)
if event.kind == EvidenceKind.VERIFIER_RESULT and event.authoritative), None)
if latest is None or not latest.success:
return False
words = self._verifier_commands.get(latest.event_id, ())
names = {Path(words[0]).name} if words else set()
if words and re.fullmatch(r'python(?:\d+(?:\.\d+)*)?', Path(words[0]).name) and '-m' in words:
module_index = words.index('-m') + 1
if module_index < len(words):
names.add(words[module_index])
return (all(identity in names for identity in identities)
and all(any(_artifact_path_matches_required(word, path, self.requirements.workspace_root)
for word in words) for path in paths))
def _supports_artifact_claim(self, kind: EvidenceKind, paths: Sequence[str]) -> bool:
"""Match every claimed artifact by identity, never by basename."""
targets = tuple(paths) or self.requirements.required_artifacts
if not targets or (not paths and len(targets) != 1):
return False
for path in targets:
matching = [event for event in self.events if event.kind == kind and event.authoritative
and _artifact_path_matches_required(event.artifact_path, path, self.requirements.workspace_root)]
successful = [event for event in matching if event.success]
# Match evaluate(): atomic helper failures preserve the previous
# successful artifact; a partial shell/Python failure may not.
destructive_failure = bool(matching and not matching[-1].success
and matching[-1].tool in {'bash', 'python'})
if not successful or destructive_failure:
return False
return True
def record_media_ingress(self, metadata: Mapping[str, Any]) -> None:
for artifact in metadata.get("artifacts") or []:
if not isinstance(artifact, Mapping):
@@ -767,6 +882,21 @@ class EvidenceLedger:
(latest_verifier.event_id,),
)
if latest_verifier and self.requirements.workspace_root:
for path in self.requirements.required_artifacts:
identity = artifact_identity(path, self.requirements.workspace_root)
expected = self._verification_versions.get(identity)
if expected in {'unobserved', 'missing-or-unreadable'} or (
expected is None and self._verification_versions_captured
):
return CompletionDecision(CompletionStatus.BLOCKED, False,
'artifact version could not be established for verification',
(latest_verifier.event_id,))
if expected is not None and expected != artifact_version(path, self.requirements.workspace_root):
return CompletionDecision(CompletionStatus.BLOCKED, False,
'artifact content changed after verification',
(latest_verifier.event_id,))
satisfied_ids: list[str] = []
missing: list[str] = []
workspace_root = str(self.requirements.workspace_root or "").strip()
@@ -774,7 +904,7 @@ class EvidenceLedger:
matches = [
event for event in self.events
if event.kind == EvidenceKind.ARTIFACT_MUTATION
and _artifact_path_matches_required(event.artifact_path, required)
and _artifact_path_matches_required(event.artifact_path, required, self.requirements.workspace_root)
]
authoritative = [
event for event in matches
@@ -793,10 +923,11 @@ class EvidenceLedger:
and latest.tool in {"bash", "python"}
)
filesystem_missing = False
if latest_success is not None and workspace_root and required.startswith("/workspace/"):
if latest_success is not None and workspace_root:
try:
root = Path(workspace_root).resolve()
candidate = (root / required.removeprefix("/workspace/")).resolve()
identity = artifact_identity(required, workspace_root)
candidate = (root / identity.removeprefix('workspace:')).resolve() if identity.startswith('workspace:') else Path(required).resolve()
candidate.relative_to(root)
filesystem_missing = not workspace_artifact_is_usable(candidate)
except (OSError, RuntimeError, ValueError):
@@ -852,20 +983,25 @@ class EvidenceLedger:
if event.kind == EvidenceKind.ARTIFACT_MUTATION
and event.authoritative
and event.success
and _artifact_path_matches_required(event.artifact_path, required)
and _artifact_path_matches_required(event.artifact_path, required, self.requirements.workspace_root)
]
matching_validations = [
(index, event)
for index, event in enumerate(self.events)
if event.kind == EvidenceKind.ARTIFACT_VALIDATION
and event.authoritative
and _artifact_path_matches_required(event.artifact_path, required)
and _artifact_path_matches_required(event.artifact_path, required, self.requirements.workspace_root)
]
if not matching_validations:
continue
latest_validation_index, latest_validation = matching_validations[-1]
latest_artifact_mutation_index = max(matching_mutation_indices, default=-1)
if latest_validation_index < latest_artifact_mutation_index:
# A pre-edit inspection cannot invalidate executable checks
# that passed against the later mutation. It still cannot
# stand in for current verification when no such check exists.
if latest_verifier_index > latest_artifact_mutation_index:
continue
return CompletionDecision(
CompletionStatus.BLOCKED,
False,
@@ -882,6 +1018,10 @@ class EvidenceLedger:
current_validation_ids.append(latest_validation.event_id)
if self.requirements.verifier_required and latest_verifier is None:
if self.requirements.executable_verifier_available:
return CompletionDecision(CompletionStatus.BLOCKED, False,
'the request requires an executable verifier result',
tuple(satisfied_ids))
validation_ids: list[str] = []
for required in self.requirements.required_artifacts:
matching_validation = [
@@ -890,7 +1030,7 @@ class EvidenceLedger:
if event.kind == EvidenceKind.ARTIFACT_VALIDATION
and event.authoritative
and event.success
and _artifact_path_matches_required(event.artifact_path, required)
and _artifact_path_matches_required(event.artifact_path, required, self.requirements.workspace_root)
]
latest_validation = matching_validation[-1] if matching_validation else None
if latest_validation is None or latest_validation[0] < latest_mutation_index:
+55 -44
View File
@@ -82,6 +82,9 @@ from src.tool_approvals import (
)
from src.tool_types import ToolBlock
from src.turn_contract import selected_tools_for_request, with_turn_contract
from src.agent_runtime.journal import propose_action, execute_action
from src.agent_runtime.completion import with_completion_gate
from src.teacher_escalation import with_teacher_takeover, request_teacher_takeover
from src.tool_utils import _truncate, get_mcp_manager
from src.agent_tools import (
parse_tool_blocks,
@@ -9171,16 +9174,8 @@ def _failed_tool_round_limit(
def _tui_python_runner_setup() -> str:
"""Select the workspace interpreter, including a primary checkout venv."""
return (
"runner=''; "
"if [ -x .venv/bin/python ]; then runner=.venv/bin/python; "
"elif [ -x venv/bin/python ]; then runner=venv/bin/python; "
"elif git_common=$(git rev-parse --path-format=absolute --git-common-dir 2>/dev/null) "
"&& [ -x \"$(dirname \"$git_common\")/.venv/bin/python\" ]; then "
"runner=\"$(dirname \"$git_common\")/.venv/bin/python\"; "
"else runner=python; fi; "
)
from src.agent_runtime.identity import TUI_PYTHON_RUNNER_SETUP
return TUI_PYTHON_RUNNER_SETUP
def _tui_local_test_runner_command(*, full: bool = False) -> str:
@@ -20381,6 +20376,8 @@ def _blocks_before_inference(turn_contract) -> bool:
@with_turn_contract
@with_teacher_takeover
@with_completion_gate
async def stream_agent_loop(
endpoint_url: str,
model: str,
@@ -20422,6 +20419,7 @@ async def stream_agent_loop(
thinking_mode: Optional[str] = None,
suppress_skills: bool = False,
reasoning_effort: Optional[str] = None,
_parent_run_id: Optional[str] = None,
) -> AsyncGenerator[str, None]:
"""Streaming agent loop generator.
@@ -26322,6 +26320,9 @@ async def stream_agent_loop(
# next model request. Do not expose a transient provider
# error or terminate the turn before that retry.
break
# Let the completion gate retain the original failure even
# when earlier tool evidence supplies useful fallback prose.
yield chunk
terminal_status = None
try:
error_line = next(
@@ -26611,7 +26612,6 @@ async def stream_agent_loop(
else "The model provider returned no usable output. No workspace change was made."
)
yield f'data: {json.dumps({"type": "final_response", "content": _failure_text})}\n\n'
yield chunk
# A terminal provider/request failure is not a completed Agent
# round. Stop before empty-response synthesis, metrics,
# teacher escalation, post-processing, or a success [DONE].
@@ -29779,7 +29779,10 @@ async def stream_agent_loop(
_completion_requirements,
)
_round_decision = _round_evidence.evaluate()
if not _round_decision.can_complete and _evidence_repair_rounds < 2:
# Missing evidence is an incomplete result, not a reason to
# manufacture additional provider rounds. Actual diagnostic
# failures can still enter the bounded recovery path.
if _round_decision.status.value == "failed" and _evidence_repair_rounds < 2:
_evidence_repair_rounds += 1
_missing = ", ".join(_round_decision.missing_artifacts)
_declared_verifiers = _completion_requirements.verifier_commands
@@ -31508,6 +31511,15 @@ async def stream_agent_loop(
local_network_budget_hit = False
local_inspection_budget_hit = False
for i, block in enumerate(tool_blocks):
native_call = converted_calls[i] if i < len(converted_calls) else None
tool_call_id = _resolved_tool_call_id(
native_call,
session_id=str(session_id or ""),
round_num=round_num,
tool_index=i,
tool_name=block.tool_type,
)
_runtime_action = propose_action(block, tool_call_id, native_call)
_call_signature = _tool_call_signature(block.tool_type, block.content)
_previous_failure = _failed_call_history.get(_call_signature)
_blocked_failed_retry = bool(
@@ -31524,6 +31536,8 @@ async def stream_agent_loop(
)
# --- Tool budget check ---
if max_tool_calls > 0 and total_tool_calls >= max_tool_calls:
if _runtime_action is not None:
_runtime_action.finish({'blocked': True, 'exit_code': 1, 'error': 'tool budget exceeded'})
yield f'data: {json.dumps({"type": "budget_exceeded", "limit": max_tool_calls, "used": total_tool_calls})}\n\n'
budget_hit = True
break
@@ -31535,26 +31549,22 @@ async def stream_agent_loop(
)
):
local_network_budget_hit = True
if _runtime_action is not None:
_runtime_action.finish({'blocked': True, 'exit_code': 1, 'error': 'network action budget exceeded'})
break
if (
_tui_local_inspection_turn
and total_tool_calls >= _TUI_LOCAL_INSPECTION_TOOL_CALL_CAP
):
local_inspection_budget_hit = True
if _runtime_action is not None:
_runtime_action.finish({'blocked': True, 'exit_code': 1, 'error': 'inspection budget exceeded'})
break
if local_inspection_budget_hit:
break
if not (_blocked_failed_retry or _blocked_redundant_read):
total_tool_calls += 1
native_call = converted_calls[i] if i < len(converted_calls) else None
tool_call_id = _resolved_tool_call_id(
native_call,
session_id=str(session_id or ""),
round_num=round_num,
tool_index=i,
tool_name=block.tool_type,
)
normalized_native_block = _normalize_native_tool_shell_wrapper(block, _last_user)
if normalized_native_block != block:
logger.info(
@@ -32807,7 +32817,8 @@ async def stream_agent_loop(
"error": "Web recovery action is repeated or exceeds the execution budget.",
"output": _web_execution_budget.instruction(),
}
return await execute_tool_block(
return await execute_action(
execute_tool_block, _runtime_action,
block,
session_id=session_id,
disabled_tools=disabled_tools,
@@ -33378,6 +33389,10 @@ async def stream_agent_loop(
# Emit tool_output (include ui_event data if present)
tool_output_data = {"type": "tool_output", "tool": block.tool_type, "command": cmd_display, "output": output_text, "exit_code": result.get("exit_code"), "execution_attempted": _execution_attempted, "blocked": bool(result.get("blocked", False))}
if _runtime_action is not None:
_runtime_action.normalize(block, 'agent_loop compatibility adapters')
_runtime_action.finish(result)
tool_output_data['action_receipt'] = _runtime_action.to_dict()
# Keep exact arguments on email mutation events. The frontend uses
# these UIDs to reconcile an agent cleanup immediately, even when
# a provider returns only human-readable MCP text.
@@ -37341,29 +37356,25 @@ async def stream_agent_loop(
)
yield f"data: {json.dumps({'type': 'metrics', 'data': metrics})}\n\n"
# Teacher-escalation: inline takeover visible in the chat stream.
# The student just finished; if Tier 1 flags failure, the teacher
# gets a turn (with its own tool calls forwarded to the user) and
# a skill is saved ONLY if the teacher actually succeeds. Skipped
# when we ARE the teacher to avoid recursion.
# Queue the existing teacher hook. The outer adapter executes it only
# after this invocation's completion gate and action context have closed.
if not _is_teacher_run and not guide_only and not _awaiting_user:
try:
from src.teacher_escalation import run_teacher_inline
async for evt in run_teacher_inline(
student_endpoint_url=endpoint_url,
student_messages=messages,
student_tool_events=tool_events,
student_reply=full_response,
owner=owner,
session_id=session_id,
workspace=workspace,
disabled_tools=disabled_tools,
tool_policy=tool_policy,
active_document=active_document,
active_email=active_email,
):
yield evt
except Exception as _esc_err:
logger.warning(f"teacher escalation hook failed: {_esc_err}", exc_info=True)
request_teacher_takeover(
student_endpoint_url=endpoint_url,
student_messages=messages,
student_tool_events=tool_events,
student_reply=full_response,
owner=owner,
session_id=session_id,
workspace=workspace,
disabled_tools=disabled_tools,
tool_policy=tool_policy,
active_document=active_document,
active_email=active_email,
turn_contract=turn_contract,
external_untrusted_context_seen=run_security.external_untrusted_context_seen,
client_runtime_context=client_runtime_context,
plan_mode=plan_mode,
)
yield "data: [DONE]\n\n"
+1
View File
@@ -0,0 +1 @@
"""Run-scoped contracts behind the public agent-loop compatibility facade."""
+366
View File
@@ -0,0 +1,366 @@
"""One presentation gate between agent execution and externally visible prose.
Tool, progress and interaction events stay live. Answer deltas are held until
the generator unwinds so a later replacement cannot conceal an earlier false
claim. This consumes no provider calls. Cancellation closes the inner generator
under the same journal/turn authority; it never emits a successful terminal event.
"""
from __future__ import annotations
from contextlib import aclosing
from dataclasses import replace
from functools import wraps
from inspect import signature
import json
import re
from time import perf_counter
from src.agent_evidence import (
CompletionDecision, CompletionStatus, EvidenceKind, EvidenceLedger,
requirements_from_runtime_context, _execution_obligation, _unquoted_statements,
_ARTIFACT_PATH,
)
from .journal import ActionJournal, bind_journal, current_journal
_TEST_CLAIM = re.compile(
r'\b(?:(?:all\s+)?(?:tests?|checks?|verification|suite)\s+(?:have\s+|has\s+|now\s+|are\s+|is\s+)*(?:passed|passing|successful|green)|'
r'(?:passed|passing)\s+(?:all\s+)?(?:the\s+)?tests?|\d+\s+passed)\b', re.I)
_TEST_STATUS_CLAIM = re.compile(
r'\b(?:tests?|pytest|unittest|test suite|checks?|verification)\s*[:—-]?\s*'
r'(?:all\s+|have\s+|has\s+|now\s+|are\s+|is\s+|ran\s+)*'
r'(?:pass(?:ed|ing)?|succeeded|successful(?:ly)?|green)\b|'
r'\b(?:zero|no|0)\s+(?:test\s+)?failures\b', re.I)
_EXECUTION_CLAIM = re.compile(
r'\b(?:(?:I|we|I\'ve|we\'ve|and)\s+(?:have\s+)?(?:successfully\s+)?(?:ran|executed|tested|verified|created|updated|modified|wrote|saved|fixed|completed)|'
rf'(?:file|artifact|command|script|service|server|{_ARTIFACT_PATH})\s+(?:was\s+|has\s+been\s+|is\s+)?(?:successfully\s+)?(?:created|updated|written|saved|executed|started)|'
r'(?:successfully\s+)(?:ran|executed|created|updated|saved|completed))\b', re.I)
_UNATTESTED_TEST_METRIC = re.compile(
r'\b\d+\s+(?:(?:unit|integration)\s+)?tests?\s+pass(?:ed|ing)?\b|'
r'\b\d+\s+passed\b|\b\d+(?:\.\d+)?%\s+(?:test\s+)?coverage\b', re.I)
_UNBOUNDED_SUCCESS = re.compile(
r'\b(?:everything|all\s+(?:bugs|issues))\s+(?:is\s+|are\s+|has\s+been\s+)?'
r'(?:fixed|resolved|working)\b', re.I)
_MUTATION_CLAIM = re.compile(
r'\b(?:created|updated|modified|wrote|written|saved|fixed)\b', re.I)
_TEST_IDENTITY = re.compile(r'\b(?:pytest|unittest)\b', re.I)
_TEST_SUBJECT = re.compile(r'\b(?:tests?|test suite|pytest|unittest|checks?|verification)\b', re.I)
_CLAIM_PATH = re.compile(_ARTIFACT_PATH)
_BARE_SUCCESS = re.compile(r'^\s*(?:done|completed|success|all done|all set|fixed)[.!]?\s*$', re.I)
_NON_REPORT_SCOPE = re.compile(
r'^\s*(?:if|unless|suppose|imagine|hypothetically|for\s+(?:example|instance))\b|'
r'\b(?:if|when|whenever|unless|until)\b|'
r'\b(?:can|could|may|might|should|would|will|must)\b|'
r'\b(?:says?|said|states?|stated|example)\b', re.I)
def _current_run_claims(statement: str, *, execution_required: bool) -> list[tuple[str, str]]:
"""Classify asserted execution, separately from the turn's obligation.
Past actions and current result/status predicates are reports. Conditional,
modal, attributed and example clauses are scoped prose. Bare terminal
success only carries execution meaning under an execution contract.
"""
if _BARE_SUCCESS.fullmatch(statement):
return [('terminal', statement)] if execution_required else []
actions = list(_EXECUTION_CLAIM.finditer(statement))
leading = re.match(r'^\s*(?:successfully\s+)?(?:created|updated|modified|wrote|saved)\b', statement, re.I)
if leading:
actions.insert(0, leading)
candidates = [('action', match) for match in actions]
for kind, pattern in [('metric', _UNATTESTED_TEST_METRIC), ('metric', _UNBOUNDED_SUCCESS),
('test', _TEST_CLAIM), ('test', _TEST_STATUS_CLAIM)]:
candidates.extend((kind, match) for match in pattern.finditer(statement))
claims = []
for kind, match in candidates:
# Scope markers after an asserted action do not make that action
# hypothetical ("I ran pytest to see if ..."). An immediate conditional
# continuation does qualify a result ("Tests passed if ...").
if _NON_REPORT_SCOPE.search(statement[:match.start()]) or re.match(
r'\s+(?:if|when|whenever|unless|until)\b', statement[match.end():], re.I):
continue
end = next((action.start() for action in actions if action.start() > match.start()), len(statement))
scope = statement[match.start():end]
if kind == 'action':
if _MUTATION_CLAIM.search(match.group()):
kind = 'mutation'
elif _TEST_SUBJECT.search(scope):
kind = 'test'
else:
kind = 'execution'
claims.append((kind, scope))
return claims
def _supported_prose(text: str, ledger: EvidenceLedger, decision: CompletionDecision) -> tuple[str, str]:
"""Remove unsupported assertions at statement boundaries; add no notice."""
incomplete = decision.reason if not decision.can_complete and decision.status != CompletionStatus.AWAITING_USER else ''
execution_required = _execution_obligation(ledger.requirements)
kept = []
removed = ''
for statement, scoped in _unquoted_statements(text):
why = ''
for claim, scope in _current_run_claims(scoped, execution_required=execution_required):
paths = tuple(match.group().rstrip('.') for match in _CLAIM_PATH.finditer(scope))
if claim == 'metric':
why = 'test counts, coverage or exhaustive correctness were not established by execution evidence'
elif claim == 'test':
identities = tuple(match.group().lower() for match in _TEST_IDENTITY.finditer(scope))
if decision.status != CompletionStatus.VERIFIED or not ledger._supports_verifier_claim(identities, paths):
why = 'no current passing executable verification supports the claim'
elif claim == 'mutation':
if not ledger._supports_artifact_claim(EvidenceKind.ARTIFACT_MUTATION, paths):
why = 'no matching artifact mutation supports the execution claim'
elif claim == 'execution':
# A generic assertion cannot be tied confidently to a receipt.
why = 'no matching operation supports the execution claim'
elif claim == 'terminal' and decision.status not in {CompletionStatus.SATISFIED, CompletionStatus.VERIFIED}:
why = incomplete or 'no successful execution supports completion'
if why:
break
if why:
removed = removed or why
else:
kept.append(statement)
prose = ''.join(kept).strip() if removed else text
return prose, removed
def completion_answer(text: str, ledger: EvidenceLedger, decision: CompletionDecision) -> tuple[str, str]:
"""Keep explanatory prose; remove unsupported assertions and attach facts.
Exit status proves neither test counts nor coverage. A bad assertion is
removed at statement boundaries instead of erasing an entire explanation.
The execution outcome remains separate from a discarded model assertion.
"""
incomplete = decision.reason if not decision.can_complete and decision.status != CompletionStatus.AWAITING_USER else ''
execution_required = _execution_obligation(ledger.requirements)
prose, removed = _supported_prose(text, ledger, decision)
if incomplete or (removed and execution_required and decision.status in {CompletionStatus.UNVERIFIED, CompletionStatus.AWAITING_USER}):
reason = incomplete or removed
missing = (' Missing artifacts: ' + ', '.join(decision.missing_artifacts) + '.'
if decision.missing_artifacts else '')
notice = 'The task is incomplete: ' + reason.rstrip('.') + '.' + missing
recorded = [path for path in ledger.requirements.required_artifacts
if ledger._supports_artifact_claim(EvidenceKind.ARTIFACT_MUTATION, (path,))]
if removed and recorded:
notice += ' Recorded artifact mutation: ' + ', '.join(recorded) + '.'
return notice + ('\n\n' + prose if prose.strip() else ''), reason
if removed and not execution_required and decision.status != CompletionStatus.VERIFIED:
notice = 'Unsupported execution claims were omitted: ' + removed.rstrip('.') + '.'
return (prose.rstrip() + '\n\n' + notice) if prose.strip() else notice, removed
if decision.can_complete and (ledger.requirements.required_artifacts or ledger.requirements.verifier_required or removed):
facts = []
if ledger.requirements.required_artifacts:
facts.append('Output available: ' + ', '.join(ledger.requirements.required_artifacts) + '.')
if decision.status == CompletionStatus.VERIFIED:
facts.append('The latest executable verification passed.')
elif any(e.kind == EvidenceKind.ARTIFACT_VALIDATION and e.authoritative and e.success for e in ledger.events):
facts.append('Artifact readback verified. No passing executable test result was recorded.')
else:
facts.append('No passing executable test result was recorded.')
summary = ' '.join(facts)
return (prose.rstrip() + '\n\n' + summary) if prose.strip() else summary, removed
return prose, removed
def _event(data: dict) -> str:
return 'data: ' + json.dumps(data) + '\n\n'
def with_completion_gate(func):
call_signature = signature(func)
@wraps(func)
async def wrapped(*args, **kwargs):
started = perf_counter()
first_answer_at = None
arguments = call_signature.bind(*args, **kwargs)
arguments.apply_defaults()
bound = arguments.arguments
messages = bound.get('messages') or []
instruction = next((m.get('content', '') for m in reversed(messages)
if m.get('role') == 'user' and isinstance(m.get('content'), str)), '')
context = bound.get('client_runtime_context') or {}
requirements = requirements_from_runtime_context(context, instruction=instruction)
from src.tool_execution import vet_workspace
# A completion declaration is not a filesystem permission. Only the
# explicit, vetted runtime workspace may be read for artifact versions.
trusted_workspace = vet_workspace(bound.get('workspace')) if bound.get('workspace') else ''
requirements = replace(requirements, workspace_root=trusted_workspace or '')
parent = current_journal()
journal = ActionJournal(
workspace=requirements.workspace_root, observed_artifacts=requirements.required_artifacts,
parent_run_id=bound.get('_parent_run_id') or (parent.run_id if parent is not None else None))
answer_events: list[dict] = []
metrics_events: list[dict] = []
answer = ''
has_final = False
done = False
awaiting = False
exhausted = False
provider_error: str | None = None
with bind_journal(journal):
async with aclosing(func(*args, **kwargs)) as stream:
async for chunk in stream:
if chunk.strip() == 'data: [DONE]':
done = True
continue
try:
data = json.loads(chunk[6:]) if chunk.startswith('data: ') else None
except (ValueError, TypeError):
data = None
if not isinstance(data, dict):
if chunk.startswith('event: error'):
# The inner stream may still emit failed-terminal
# diagnostics. Hold the original error until those
# and the buffered answer have been released.
provider_error = provider_error or chunk
continue
yield chunk
continue
kind = data.get('type')
if kind == 'completion_decision':
existing = data.get('data') or {}
awaiting |= existing.get('status') == 'awaiting_user'
exhausted |= existing.get('status') == 'exhausted'
continue
if kind in {'metrics', 'agent_terminal'}:
metrics_events.append(data)
declared = (data.get('data') or {}).get('completion_requirements')
awaiting |= bool((data.get('data') or {}).get('missing_workspace'))
if isinstance(declared, dict):
requirements = requirements_from_runtime_context({'completion_requirements': declared})
requirements = replace(requirements, workspace_root=trusted_workspace or '')
# New obligations affect future receipts only. Never
# backfill historical versions with present bytes.
journal.observed_artifacts = tuple(dict.fromkeys(
(*journal.observed_artifacts, *requirements.required_artifacts)))
continue
if kind == 'ask_user':
awaiting = True
payload = data.get('data') or {}
if isinstance(payload.get('question'), str):
current = EvidenceLedger.from_tool_events(journal.evidence_events(), requirements)
question, why = completion_answer(payload['question'], current, current.evaluate(awaiting_user=True))
if why:
data = {**data, 'data': {**payload, 'question': question}}
chunk = _event(data)
if kind == 'final_response':
if first_answer_at is None:
first_answer_at = perf_counter()
answer = str(data.get('content') or '')
has_final = True
answer_events.append(data)
continue
if 'delta' in data or isinstance(data.get('thinking'), str):
if first_answer_at is None:
first_answer_at = perf_counter()
# Boolean thinking=True marks a reasoning-only delta;
# a textual thinking companion must not hide an answer
# delta. Both shapes remain buffered until the gate.
if isinstance(data.get('thinking'), str):
answer_events.append({'delta': data['thinking'], 'thinking': True})
data = {key: value for key, value in data.items() if key != 'thinking'}
if 'delta' not in data:
continue
if data.get('thinking') is not True and 'delta' in data:
if has_final:
answer = ''
has_final = False
answer += str(data.get('delta') or '')
answer_events.append(data)
continue
yield chunk
if provider_error and not answer_events and not metrics_events:
yield provider_error
return
presentation_replaced = False
if not provider_error and not has_final and requirements.required_artifacts:
terminal_texts = next((event.get('data', {}).get('round_texts')
for event in reversed(metrics_events)
if isinstance(event.get('data', {}).get('round_texts'), list)
and all(isinstance(text, str) for text in event['data']['round_texts'])), None)
if terminal_texts is not None:
terminal_answer = '\n\n'.join(text for text in terminal_texts if text.strip())
if terminal_answer != answer:
# The loop can retract a rejected round while retaining
# its live deltas. Do not resurrect those buffered drafts
# after recovery. Terminal prose still passes this gate.
presentation_replaced = True
answer = terminal_answer
answer_events = [event for event in answer_events if event.get('thinking') is True]
ledger = EvidenceLedger.from_tool_events(journal.evidence_events(), requirements)
decision = ledger.evaluate(exhausted=exhausted, awaiting_user=awaiting)
if provider_error:
decision = replace(decision, status=CompletionStatus.FAILED,
can_complete=False, reason='Model request failed')
# Exhaustion limits execution; factual source synthesis can remain
# useful and must not be replaced merely because the budget ended.
presentation_decision = ledger.evaluate(awaiting_user=awaiting) if exhausted and not provider_error else decision
safe_answer, reason = completion_answer(answer, ledger, presentation_decision)
# Evaluate each earlier draft as well as the final replacement.
# Never replay an unsupported intermediate success claim.
draft = ''.join(str(e.get('delta') or e.get('content') or '')
+ (e['thinking'] if isinstance(e.get('thinking'), str) else '')
for e in answer_events)
_, unsafe_draft = completion_answer(draft, ledger, presentation_decision)
if not answer.strip() and unsafe_draft:
reason = reason or unsafe_draft
safe_answer, _ = completion_answer(draft, ledger, presentation_decision)
if reason and _execution_obligation(requirements) and decision.can_complete and decision.status == CompletionStatus.UNVERIFIED:
decision = CompletionDecision(CompletionStatus.UNVERIFIED, False, reason,
decision.evidence_ids, decision.missing_artifacts)
released_at = perf_counter()
if not provider_error:
yield _event({'type': 'completion_decision', 'data': decision.to_dict()})
replaced_answer = bool(presentation_replaced or reason or unsafe_draft or safe_answer != answer)
if replaced_answer:
reasoning = [event for event in answer_events if event.get('thinking') is True]
_, unsafe_reasoning = completion_answer(
''.join(str(event.get('delta') or '') for event in reasoning), ledger,
replace(presentation_decision, can_complete=True))
if not unsafe_reasoning:
for event in reasoning:
yield _event(event)
yield _event({'type': 'final_response', 'content': safe_answer})
else:
for event in answer_events:
yield _event(event)
if provider_error:
yield _event({'type': 'completion_decision', 'data': decision.to_dict()})
for event in metrics_events:
metadata = event.setdefault('data', {})
metadata.update(completion_decision=decision.to_dict(), evidence_events=ledger.to_list(),
action_receipts=journal.to_list(), completion_requirements=requirements.to_dict(),
run_id=journal.run_id, parent_run_id=journal.parent_run_id)
metadata['completion_gate'] = {
'buffer_seconds': released_at - first_answer_at if first_answer_at is not None else 0,
'first_visible_answer_seconds': released_at - started,
'additional_provider_calls': 0,
'answer_replaced': replaced_answer,
}
if replaced_answer:
if not provider_error:
metadata['round_texts'] = [safe_answer]
metadata['completion_gate_reason'] = reason or unsafe_draft or 'receipt_summary'
if provider_error and isinstance(metadata.get('round_texts'), list):
# Failed rounds stay as per-round diagnostics, but they are
# rendered again on reload. Apply the same statement filter
# as the live answer so a rejected claim cannot reappear.
metadata['round_texts'] = [
_supported_prose(text, ledger, presentation_decision)[0] if isinstance(text, str) else text
for text in metadata['round_texts']]
if isinstance(metadata.get('thinking'), str):
_, unsafe_thinking = completion_answer(metadata['thinking'], ledger,
replace(presentation_decision, can_complete=True))
if unsafe_thinking:
metadata.pop('thinking')
yield _event(event)
if provider_error:
yield provider_error
return
if done:
yield 'data: [DONE]\n\n'
return wrapped
+167
View File
@@ -0,0 +1,167 @@
"""Canonical evidence identities; these helpers never grant filesystem access."""
from __future__ import annotations
import hashlib
import json
import os
from pathlib import Path, PurePosixPath
import re
import shlex
import stat
from .path_policy import _is_sensitive_path
TUI_PYTHON_RUNNER_SETUP = (
"runner=''; "
"if [ -x .venv/bin/python ]; then runner=.venv/bin/python; "
"elif [ -x venv/bin/python ]; then runner=venv/bin/python; "
"elif git_common=$(git rev-parse --path-format=absolute --git-common-dir 2>/dev/null) "
"&& [ -x \"$(dirname \"$git_common\")/.venv/bin/python\" ]; then "
"runner=\"$(dirname \"$git_common\")/.venv/bin/python\"; "
"else runner=python; fi; "
)
def digest(value: object) -> str:
return hashlib.sha256(json.dumps(value, sort_keys=True, ensure_ascii=False,
separators=(",", ":"), default=str).encode()).hexdigest()
def artifact_version(value: str, workspace: str) -> str:
"""Content version for a confined declared output; missing/unreadable is explicit."""
identity = artifact_identity(value, workspace)
if not workspace or not identity.startswith('workspace:'):
return 'unobserved'
root = Path(workspace).resolve()
candidate = (root / identity.removeprefix('workspace:')).resolve()
if not candidate.is_relative_to(root) or _is_sensitive_path(str(candidate)):
return 'unobserved'
if not hasattr(os, 'O_NOFOLLOW') or not os.supports_dir_fd:
return 'unobserved'
directory = None
try:
directory = os.open(root, os.O_RDONLY | os.O_DIRECTORY | os.O_NOFOLLOW)
parts = candidate.relative_to(root).parts
if not parts:
return 'unobserved'
for part in parts[:-1]:
child = os.open(part, os.O_RDONLY | os.O_DIRECTORY | os.O_NOFOLLOW, dir_fd=directory)
os.close(directory)
directory = child
descriptor = os.open(parts[-1], os.O_RDONLY | os.O_NOFOLLOW | os.O_NONBLOCK, dir_fd=directory)
with os.fdopen(descriptor, 'rb') as stream:
info = os.fstat(stream.fileno())
if not stat.S_ISREG(info.st_mode) or info.st_size > 64 * 1024 * 1024:
return 'unobserved'
result = hashlib.sha256()
remaining = 64 * 1024 * 1024
while block := stream.read(min(1024 * 1024, remaining + 1)):
remaining -= len(block)
if remaining < 0:
return 'unobserved'
result.update(block)
return result.hexdigest()
except (OSError, ValueError):
return 'missing-or-unreadable'
finally:
if directory is not None:
os.close(directory)
def artifact_identity(value: str, workspace: str = "") -> str:
"""Unify relative, virtual and host aliases without basename matching.
Resolving symlinks is evidence bookkeeping, never a confinement check. Paths
outside the workspace retain their absolute identity and cannot satisfy a
workspace obligation with the same basename.
"""
text = str(value or "").strip()
if not text:
return ""
path = PurePosixPath(text)
root = Path(workspace or "/workspace").resolve()
if path.parts[:2] == ('/', 'workspace'):
path = PurePosixPath(*path.parts[2:])
candidate = Path(str(path))
if not candidate.is_absolute():
candidate = root / candidate
try:
resolved = candidate.resolve()
relative = resolved.relative_to(root)
return "workspace:" + relative.as_posix()
except ValueError:
return "absolute:" + str(candidate.resolve())
except (OSError, RuntimeError):
return "unresolved:" + text
def executable_words(command: str) -> tuple[str, ...]:
"""Recognize one foreground command after exact interpreter or cd/set prefixes.
This is deliberately conservative evidence parsing, not shell authorization.
Other control flow, substitutions, pipelines and status-masking tails are not proof
that a verifier returned the recorded shell status.
"""
text = str(command or '').strip()
if text.startswith(TUI_PYTHON_RUNNER_SETUP):
remainder = text[len(TUI_PYTHON_RUNNER_SETUP):]
# This exact server-owned prelude only selects the interpreter. The
# trailing command must still be a single foreground invocation whose
# status is returned unchanged; the generic discovery fallback is not.
if remainder.startswith('"$runner" '):
arguments = remainder[len('"$runner" '):]
if '$' in arguments:
return ()
text = 'python ' + arguments
if any(marker in text for marker in ('`', '$(', '${', '\n', '\r')):
return ()
try:
lexer = shlex.shlex(text, posix=True, punctuation_chars=';&|<>()')
lexer.whitespace_split = True
words = list(lexer)
except ValueError:
return ()
while '&&' in words:
index = words.index('&&')
prefix = words[:index]
if not ((len(prefix) == 2 and prefix[0] == 'cd') or prefix == ['set', '-e']):
return ()
words = words[index + 1:]
if any(word and all(c in ';&|<>()' for c in word) for word in words):
return ()
while words and re.fullmatch(r'[A-Za-z_][A-Za-z0-9_]*=[^\n]*', words[0]):
words.pop(0)
return tuple(words)
def is_test_command(command: str) -> bool:
words = executable_words(command)
if not words:
return False
if any(word in {'--help', '-h', '--version', '--collect-only', '--co'} for word in words[1:]):
return False
binary = Path(words[0]).name
if binary in {'pytest', 'py.test'}:
return True
if re.fullmatch(r'python(?:\d+(?:\.\d+)?)?', binary):
args = list(words[1:])
while args and args[0] in {'-I', '-S', '-s', '-E', '-B', '-u'}:
args.pop(0)
return len(args) >= 2 and args[:2] in (['-m', 'pytest'], ['-m', 'unittest'])
if binary in {'npm', 'pnpm', 'yarn', 'make', 'cargo', 'go'}:
args = words[1:]
return bool(args and (args[0] == 'test' or binary == 'npm' and args[:2] == ('run', 'test')))
return bool(re.match(r'^/(?:tests?|verifier)/[^/]+', words[0]))
def is_validation_command(command: str) -> bool:
words = executable_words(command)
if not words:
return False
binary = Path(words[0]).name
return (is_test_command(command)
or binary in {'cat', 'head', 'tail', 'stat', 'wc', 'jq', 'cmp', 'diff',
'coqc', 'gcc', 'g++', 'clang', 'clang++', 'javac', 'rustc'}
or binary == 'test' and len(words) > 1 and words[1] in {'-e', '-f', '-s', '-d'}
or binary in {'cargo', 'go', 'npm', 'pnpm', 'yarn'} and words[1:2] in {('build',), ('check',)}
or binary == 'npm' and words[1:3] == ('run', 'build'))
+208
View File
@@ -0,0 +1,208 @@
"""Run-owned action history. Model text cannot insert authoritative receipts."""
from __future__ import annotations
from contextlib import contextmanager
from contextvars import ContextVar
from copy import deepcopy
from dataclasses import dataclass, field, asdict
from functools import wraps
from inspect import signature
from typing import Any
from uuid import uuid4
from .identity import artifact_identity, artifact_version, digest
@dataclass
class ActionReceipt:
action_id: str
call_id: str
proposed_tool: str
proposed_arguments: str
provider_arguments: Any = None
provider_tool: str = ''
tool: str = ""
arguments: str = ""
transitions: list[dict[str, Any]] = field(default_factory=list)
execution_id: str | None = None
operation_started: bool = False
outcome: dict[str, Any] | None = None
artifact_versions: dict[str, str] = field(default_factory=dict)
artifact_changes: list[str] | None = None
def transition(self, stage: str, **details: Any) -> None:
self.transitions.append({'sequence': len(self.transitions), 'stage': stage, **details})
def normalize(self, block: Any, reason: str) -> None:
tool, arguments = str(block.tool_type), str(block.content)
if tool != self.tool or arguments != self.arguments or not any(t['stage'] == 'normalized' for t in self.transitions):
self.transition('normalized', reason=reason, tool=tool, arguments=arguments,
previous_sha256=digest((self.tool, self.arguments)))
self.tool, self.arguments = tool, arguments
def finish(self, result: dict[str, Any]) -> None:
if self.outcome is not None:
return
code = result.get('exit_code')
valid_code = isinstance(code, int) and not isinstance(code, bool)
denied = bool(result.get('blocked') or result.get('approval_required')
or str(result.get('failure_kind', '')).endswith('_denied'))
self.outcome = {
'exit_code': code if valid_code else None,
'success': valid_code and code == 0 and not result.get('error') and not denied,
'authoritative': self.execution_id is not None and valid_code and not denied,
'blocked': denied,
'output_sha256': digest(result.get('output') or result.get('error') or result.get('stdout') or ''),
}
self.transition('outcome', **self.outcome)
def to_dict(self) -> dict[str, Any]:
return asdict(self)
@dataclass
class ActionJournal:
run_id: str = field(default_factory=lambda: uuid4().hex)
actions: list[ActionReceipt] = field(default_factory=list)
workspace: str = ''
observed_artifacts: tuple[str, ...] = ()
parent_run_id: str | None = None
def capture_versions(self, action: ActionReceipt) -> None:
if self.workspace:
action.artifact_versions = {
artifact_identity(path, self.workspace): artifact_version(path, self.workspace)
for path in self.observed_artifacts
}
def propose(self, block: Any, call_id: str = '', native_call: dict | None = None) -> ActionReceipt:
native = native_call or {}
function = native.get('function') or native
if not isinstance(function, dict):
function = {}
action = ActionReceipt(
action_id=f'{self.run_id}:action:{len(self.actions) + 1}', call_id=call_id,
proposed_tool=str(block.tool_type), proposed_arguments=str(block.content),
provider_arguments=deepcopy(function.get('arguments')),
provider_tool=str(function.get('name') or ''),
tool=str(block.tool_type), arguments=str(block.content),
)
action.transition('proposed')
self.actions.append(action)
return action
def to_list(self) -> list[dict[str, Any]]:
return [action.to_dict() for action in self.actions]
def evidence_events(self) -> list[dict[str, Any]]:
return [dict(tool=a.tool, command=a.arguments,
exit_code=(a.outcome or {}).get('exit_code'),
error=not (a.outcome or {}).get('success'),
execution_attempted=bool((a.outcome or {}).get('authoritative')),
blocked=(a.outcome or {}).get('blocked', False),
action_id=a.action_id, execution_id=a.execution_id,
artifact_versions=a.artifact_versions, artifact_changes=a.artifact_changes)
for a in self.actions if a.outcome is not None]
_JOURNAL: ContextVar[ActionJournal | None] = ContextVar('runtime_action_journal', default=None)
_ACTION: ContextVar[ActionReceipt | None] = ContextVar('runtime_current_action', default=None)
@contextmanager
def bind_journal(journal: ActionJournal):
token = _JOURNAL.set(journal)
action_token = _ACTION.set(None)
try:
yield journal
finally:
_ACTION.reset(action_token)
_JOURNAL.reset(token)
def current_journal() -> ActionJournal | None:
return _JOURNAL.get()
def propose_action(block: Any, call_id: str = '', native_call: dict | None = None) -> ActionReceipt | None:
journal = _JOURNAL.get()
return journal.propose(block, call_id, native_call) if journal else None
def mark_authorized() -> None:
action = _ACTION.get()
if action is not None and not any(t['stage'] == 'authorized' for t in action.transitions):
action.transition('authorized', authority='existing_dispatcher_policy')
def mark_dispatch() -> None:
action = _ACTION.get()
if action is not None and action.execution_id is None:
mark_authorized()
action.execution_id = action.action_id + ':execution:1'
action.transition('dispatched', execution_id=action.execution_id)
async def dispatched(operation):
"""Record an actual backend invocation, distinct from router admission."""
mark_dispatch()
return await operation
def mark_operation_started(backend: str, **details: Any) -> None:
action = _ACTION.get()
if action is not None:
action.operation_started = True
action.transition('operation_started', backend=backend, **details)
async def execute_action(executor, action: ActionReceipt | None, block: Any, **kwargs):
"""Adapter binds the proposal across async tool-task execution and cleanup."""
if action is not None:
action.normalize(block, 'agent_loop compatibility adapters')
token = _ACTION.set(action)
try:
return await executor(block, **kwargs)
finally:
_ACTION.reset(token)
def record_action(func):
call_signature = signature(func)
@wraps(func)
async def wrapped(*args, **kwargs):
bound = call_signature.bind(*args, **kwargs)
block = bound.arguments['block']
action = _ACTION.get() or propose_action(block)
token = _ACTION.set(action)
try:
journal = current_journal()
before = {}
if action is not None:
action.normalize(block, 'dispatcher input')
if journal is not None and journal.workspace:
journal.capture_versions(action)
before = dict(action.artifact_versions)
description, result = await func(*args, **kwargs)
if action is not None:
journal = current_journal()
if journal is not None:
journal.capture_versions(action)
if journal.workspace:
action.artifact_changes = [key for key, value in action.artifact_versions.items()
if before.get(key) != value]
if 'BLOCKED' in description and action.execution_id is None:
action.transition('authorization_denied', reason=str(result.get('error', '')))
action.finish({**result, 'blocked': True})
else:
action.finish(result)
return description, result
except BaseException as exc:
if action is not None:
action.transition('interrupted', category=type(exc).__name__)
raise
finally:
_ACTION.reset(token)
return wrapped
+26
View File
@@ -0,0 +1,26 @@
"""Existing sensitive-path policy shared by tools and evidence observation.
This is a deny predicate, not an authorization grant or a workspace scope.
"""
import os
_SENSITIVE_BASENAMES: set[str] = {
".ssh", ".gnupg", ".gitconfig",
".bashrc", ".bash_profile", ".bash_logout",
".zshrc", ".zprofile", ".zshenv",
".profile", ".tcshrc", ".cshrc", ".env", ".netrc",
}
_SENSITIVE_FILE_PATTERNS: tuple[str, ...] = (
"authorized_keys", "id_rsa", "id_ed25519", "id_ecdsa",
"known_hosts", "auth.json", "app.db", "settings.json",
)
_SENSITIVE_BASENAMES_CF = frozenset(b.casefold() for b in _SENSITIVE_BASENAMES)
_SENSITIVE_FILE_PATTERNS_CF = frozenset(p.casefold() for p in _SENSITIVE_FILE_PATTERNS)
def _is_sensitive_path(resolved: str) -> bool:
# Case folding is required even on POSIX: default macOS volumes are
# case insensitive but os.path.normcase there does not fold path names.
parts = [p.casefold() for p in resolved.split(os.sep)]
filename = parts[-1] if parts else ""
return any(part in _SENSITIVE_BASENAMES_CF for part in parts) or filename in _SENSITIVE_FILE_PATTERNS_CF
+3
View File
@@ -16,6 +16,7 @@ from urllib.parse import urlparse
import httpx
from src.constants import MAX_OUTPUT_CHARS
from src.agent_runtime.journal import mark_operation_started
# Agent shell calls must fail fast enough for the loop to recover and choose a
# better tool. A one-hour default can pin an entire benchmark worker on an
@@ -690,6 +691,7 @@ class BashTool:
)
except RuntimeError as exc:
return {"error": str(exc), "exit_code": 1}
mark_operation_started('subprocess', pid=proc.pid)
stdout, stderr, rc, timed_out = await _run_subprocess_streaming(
proc,
timeout=DEFAULT_BASH_TIMEOUT,
@@ -1015,6 +1017,7 @@ class PythonTool:
env=_subproc_env,
cwd=agent_cwd(),
)
mark_operation_started('subprocess', pid=proc.pid)
stdout, stderr, rc, timed_out = await _run_subprocess_streaming(
proc,
timeout=DEFAULT_PYTHON_TIMEOUT,
+11
View File
@@ -42,6 +42,7 @@ async def _drain_agent(sess, messages):
saves, so the frontend rebuilds them as standard agent-thread tool cards."""
from src.agent_loop import stream_agent_loop
full = ""
final_replaced = False
tool_events = []
round_num = 1
async for chunk in stream_agent_loop(
@@ -68,7 +69,17 @@ async def _drain_agent(sess, messages):
if isinstance(delta, str):
if d.get("thinking"):
continue
if final_replaced:
# A later answer supersedes the replacement, as the
# completion gate treats it.
full = ""
final_replaced = False
full += delta
elif d.get("type") == "final_response":
# The completion gate may present its sanitized answer as one
# replacement instead of deltas.
full = str(d.get("content") or "")
final_replaced = True
elif d.get("type") == "agent_step":
round_num = d.get("round", round_num)
elif d.get("type") == "tool_output":
+70 -56
View File
@@ -1976,6 +1976,7 @@ class TaskScheduler:
except Exception:
pass
full_text = ""
final_text_replaced = False
tool_results = []
approval_pause = None
@@ -1997,62 +1998,75 @@ class TaskScheduler:
)[1:]
except Exception:
_task_fallbacks = []
async for event_str in stream_agent_loop(
endpoint_url=endpoint_url,
model=model,
messages=messages,
max_rounds=_task_max_rounds,
session_id=session_id,
owner=task.owner,
headers=headers,
disabled_tools=disabled_tools,
relevant_tools=relevant_tools,
fallbacks=_task_fallbacks,
workload="background",
):
if event_str.startswith("data: ") and not event_str.startswith("data: [DONE]"):
try:
data = json.loads(event_str[6:])
# Capture text from all event types, not just delta
if "delta" in data:
if data.get("thinking"):
continue
full_text += data["delta"]
elif data.get("type") == "tool_output":
# Tool results — capture summary so we have SOMETHING even
# if the model never produces a final text response
tool_summary = data.get("stdout") or data.get("output") or data.get("result") or ""
if isinstance(tool_summary, str) and tool_summary.strip():
tool_results.append(f"[{data.get('tool', '?')}] {tool_summary[:500]}")
approval = data.get("ask_user")
if (
isinstance(approval, dict)
and approval.get("kind") == "tool_approval"
):
approval_pause = {
"tool": data.get("tool") or "tool",
"approval_id": approval.get("approval_id"),
}
# Scheduled tasks have no interactive surface that
# can safely resume a one-use grant. Retire the
# record immediately instead of leaving it pending
# and report an explicit manual-action boundary.
try:
from src.tool_approvals import tool_approval_store
tool_approval_store.consume(
approval_pause["approval_id"],
decision="deny",
owner=task.owner,
session_id=session_id,
)
except Exception:
logger.debug(
"Could not retire scheduled-task approval",
exc_info=True,
)
break
except (json.JSONDecodeError, KeyError):
pass
# Close the stream in this task on every exit, including the
# approval-pause break, so the agent run's context state unwinds here.
async with contextlib.aclosing(stream_agent_loop(
endpoint_url=endpoint_url,
model=model,
messages=messages,
max_rounds=_task_max_rounds,
session_id=session_id,
owner=task.owner,
headers=headers,
disabled_tools=disabled_tools,
relevant_tools=relevant_tools,
fallbacks=_task_fallbacks,
workload="background",
)) as agent_stream:
async for event_str in agent_stream:
if event_str.startswith("data: ") and not event_str.startswith("data: [DONE]"):
try:
data = json.loads(event_str[6:])
# Capture text from all event types, not just delta
if "delta" in data:
if data.get("thinking"):
continue
if final_text_replaced:
# A later answer supersedes the replacement,
# as the completion gate treats it.
full_text = ""
final_text_replaced = False
full_text += data["delta"]
elif data.get("type") == "final_response":
# The completion gate may present its sanitized
# answer as one replacement instead of deltas.
full_text = str(data.get("content") or "")
final_text_replaced = True
elif data.get("type") == "tool_output":
# Tool results — capture summary so we have SOMETHING even
# if the model never produces a final text response
tool_summary = data.get("stdout") or data.get("output") or data.get("result") or ""
if isinstance(tool_summary, str) and tool_summary.strip():
tool_results.append(f"[{data.get('tool', '?')}] {tool_summary[:500]}")
approval = data.get("ask_user")
if (
isinstance(approval, dict)
and approval.get("kind") == "tool_approval"
):
approval_pause = {
"tool": data.get("tool") or "tool",
"approval_id": approval.get("approval_id"),
}
# Scheduled tasks have no interactive surface that
# can safely resume a one-use grant. Retire the
# record immediately instead of leaving it pending
# and report an explicit manual-action boundary.
try:
from src.tool_approvals import tool_approval_store
tool_approval_store.consume(
approval_pause["approval_id"],
decision="deny",
owner=task.owner,
session_id=session_id,
)
except Exception:
logger.debug(
"Could not retire scheduled-task approval",
exc_info=True,
)
break
except (json.JSONDecodeError, KeyError):
pass
if approval_pause is not None:
return (
+112 -38
View File
@@ -24,6 +24,10 @@ itself wasn't confident about.
from __future__ import annotations
import asyncio
from contextlib import aclosing
from contextvars import ContextVar
from copy import deepcopy
from functools import wraps
import logging
import re
from typing import Any, Dict, List, Optional, Tuple
@@ -32,6 +36,58 @@ from urllib.parse import urlparse
logger = logging.getLogger(__name__)
_TAKEOVER: ContextVar[dict | None] = ContextVar('teacher_takeover_request', default=None)
def request_teacher_takeover(**parameters):
"""Record a finished student's handoff; execution waits for its gate to close."""
from src.agent_runtime.journal import current_journal
request = _TAKEOVER.get()
if request is not None:
journal = current_journal()
request.update(parameters, parent_run_id=journal.run_id if journal is not None else None)
def with_teacher_takeover(func):
"""Orchestrate gated invocations and own the single outer stream terminator."""
@wraps(func)
async def wrapped(*args, **kwargs):
request = {}
token = _TAKEOVER.set(request)
done = False
failed = False
try:
async with aclosing(func(*args, **kwargs)) as stream:
async for chunk in stream:
if chunk.strip() == 'data: [DONE]':
done = True
continue
failed |= chunk.startswith('event: error')
yield chunk
# The parent generator and gate have both unwound. Child control
# events now belong only to the child, never to the parent gate.
if request and not failed:
try:
async with aclosing(run_teacher_inline(**request)) as stream:
async for chunk in stream:
if chunk.strip() == 'data: [DONE]':
continue
failed |= chunk.startswith('event: error')
yield chunk
except Exception as exc:
logger.warning('teacher escalation hook failed: %s', exc, exc_info=True)
if not failed:
import json
yield 'data: ' + json.dumps({'type': 'escalation_failed', 'reason': str(exc),
'teacher': True}) + '\n\n'
if done and not failed:
yield 'data: [DONE]\n\n'
finally:
_TAKEOVER.reset(token)
return wrapped
# Hosts considered SOTA / paid APIs — if the student's endpoint URL
# hits one of these, the loop is OFF (the user is already paying for
# a top-tier model; no need to escalate).
@@ -524,6 +580,11 @@ async def run_teacher_inline(
tool_policy: Any = None,
active_document: Any = None,
active_email: Optional[Dict[str, str]] = None,
turn_contract=None,
parent_run_id: Optional[str] = None,
external_untrusted_context_seen: bool = False,
client_runtime_context: Optional[Dict[str, Any]] = None,
plan_mode: bool = False,
):
"""Async generator. Yields SSE event strings.
@@ -606,7 +667,7 @@ async def run_teacher_inline(
# user/assistant/tool history so the teacher sees what the student
# tried. The appended note leads with the user request text so RAG
# tool selection picks the right tools for the teacher's turn.
history = [m for m in student_messages if m.get("role") != "system"]
history = deepcopy([m for m in student_messages if m.get("role") != "system"])
note_content = (
f"{user_request or '(no user request captured)'}\n\n"
"[teacher-takeover] The previous attempt by the student model "
@@ -623,8 +684,9 @@ async def run_teacher_inline(
captured_tool_events: List[Dict[str, Any]] = []
captured_text_parts: List[str] = []
captured_metrics: Dict[str, Any] = {}
captured_decision: Dict[str, Any] = {}
async for evt_str in stream_agent_loop(
async with aclosing(stream_agent_loop(
endpoint_url=teacher_url,
model=teacher_model,
messages=teacher_messages,
@@ -632,51 +694,63 @@ async def run_teacher_inline(
owner=owner,
session_id=session_id,
workspace=workspace,
disabled_tools=disabled_tools,
disabled_tools=set(disabled_tools) if disabled_tools is not None else None,
tool_policy=tool_policy,
active_document=active_document,
active_email=active_email,
turn_contract=turn_contract,
_parent_run_id=parent_run_id,
external_untrusted_context_seen=external_untrusted_context_seen,
client_runtime_context=deepcopy(client_runtime_context),
plan_mode=plan_mode,
_is_teacher_run=True,
):
# Swallow teacher's own [DONE] — outer loop emits the real one
if "[DONE]" in evt_str:
continue
if evt_str.startswith("data: "):
try:
payload = json.loads(evt_str[6:].strip())
except Exception:
)) as stream:
async for evt_str in stream:
# Swallow teacher's own [DONE] — outer loop emits the real one
if evt_str.strip() == 'data: [DONE]':
continue
if evt_str.startswith('event: error'):
yield evt_str
continue
if isinstance(payload, dict):
payload["teacher"] = True
typ = payload.get("type")
if typ == "metrics" and isinstance(payload.get("data"), dict):
# The outer chat route persists only the last metrics
# payload. Keep a copy so any approval produced after the
# recursive teacher run's metrics remains reloadable.
captured_metrics = dict(payload["data"])
if typ == "tool_output":
captured_tool_event = {
"tool": payload.get("tool"),
"command": payload.get("command"),
"output": payload.get("output"),
"exit_code": payload.get("exit_code"),
}
if isinstance(payload.get("ask_user"), dict):
captured_tool_event["ask_user"] = payload["ask_user"]
captured_tool_events.append(captured_tool_event)
if "delta" in payload and isinstance(payload["delta"], str):
if payload.get("thinking"):
continue
captured_text_parts.append(payload["delta"])
yield 'data: ' + json.dumps(payload) + '\n\n'
continue
yield evt_str
return
if evt_str.startswith("data: "):
try:
payload = json.loads(evt_str[6:].strip())
except Exception:
yield evt_str
continue
if isinstance(payload, dict):
payload["teacher"] = True
typ = payload.get("type")
if typ == 'completion_decision' and isinstance(payload.get('data'), dict):
captured_decision = payload['data']
if typ == "metrics" and isinstance(payload.get("data"), dict):
# The outer chat route persists only the last metrics
# payload. Keep a copy so any approval produced after the
# recursive teacher run's metrics remains reloadable.
captured_metrics = dict(payload["data"])
if typ == "tool_output":
captured_tool_event = {
"tool": payload.get("tool"),
"command": payload.get("command"),
"output": payload.get("output"),
"exit_code": payload.get("exit_code"),
}
if isinstance(payload.get("ask_user"), dict):
captured_tool_event["ask_user"] = payload["ask_user"]
captured_tool_events.append(captured_tool_event)
if "delta" in payload and isinstance(payload["delta"], str):
if payload.get("thinking"):
continue
captured_text_parts.append(payload["delta"])
yield 'data: ' + json.dumps(payload) + '\n\n'
continue
yield evt_str
# A takeover that paused for a question or exact action has not completed
# yet. Its server-owned approval card is already in the live/persisted tool
# events; do not evaluate the partial trace or distill it into a skill.
if any(event.get("ask_user") for event in captured_tool_events):
if (any(event.get("ask_user") for event in captured_tool_events)
or (captured_decision and not captured_decision.get('can_complete', False))):
return
teacher_text = "".join(captured_text_parts).strip()
+60 -90
View File
@@ -727,49 +727,13 @@ async def _route_tool_via_bridge(tool: str, content: str, session_id: Optional[s
# "tool_path_extra_roots" setting (list of path strings).
# ---------------------------------------------------------------------------
_SENSITIVE_BASENAMES: set[str] = {
".ssh", ".gnupg", ".gitconfig",
".bashrc", ".bash_profile", ".bash_logout",
".zshrc", ".zprofile", ".zshenv",
".profile", ".tcshrc", ".cshrc",
".env", ".netrc",
}
_SENSITIVE_FILE_PATTERNS: tuple[str, ...] = (
"authorized_keys", "id_rsa", "id_ed25519", "id_ecdsa",
"known_hosts", "auth.json", "app.db", "settings.json",
# Compatibility exports: the same deny predicate protects tool access and
# artifact observations, so hashing cannot become a sensitive-file side channel.
from src.agent_runtime.path_policy import (
_SENSITIVE_BASENAMES, _SENSITIVE_FILE_PATTERNS,
_SENSITIVE_BASENAMES_CF, _SENSITIVE_FILE_PATTERNS_CF, _is_sensitive_path,
)
# Case-folded views used for matching. On a case-insensitive filesystem
# (Windows, default macOS) ".SSH/AUTHORIZED_KEYS" and ".env" resolve to the
# same protected files as their lowercase forms, so the deny-list has to fold
# case before comparing — the sibling resolver already normcases paths for the
# same reason. casefold (not os.path.normcase) because normcase is a no-op on
# POSIX, which is exactly where the macOS read-exfil path lives.
_SENSITIVE_BASENAMES_CF: frozenset[str] = frozenset(b.casefold() for b in _SENSITIVE_BASENAMES)
_SENSITIVE_FILE_PATTERNS_CF: frozenset[str] = frozenset(p.casefold() for p in _SENSITIVE_FILE_PATTERNS)
def _is_sensitive_path(resolved: str) -> bool:
"""Return True if *resolved* falls under a sensitive directory or
matches a sensitive filename — regardless of what root it sits under.
Matching is case-insensitive: on Windows / default macOS a case-variant
name (``.SSH``, ``AUTHORIZED_KEYS``, ``Id_Rsa``) points at the same file as
the lowercase form, so a case-sensitive check would let it slip past the
deny-list in every file tool that relies on it.
"""
parts = [p.casefold() for p in resolved.split(os.sep)]
filename = parts[-1] if parts else ""
# Check if any path component is a sensitive directory.
for part in parts:
if part in _SENSITIVE_BASENAMES_CF:
return True
# Check filename against known sensitive files.
return filename in _SENSITIVE_FILE_PATTERNS_CF
def _tool_path_roots() -> list[str]:
"""Return the list of directory roots that read_file / write_file
@@ -1294,6 +1258,10 @@ async def _document_tool_dispatch(
# Dispatcher
# ---------------------------------------------------------------------------
from src.agent_runtime.journal import dispatched, mark_authorized, mark_dispatch, record_action
@record_action
async def execute_tool_block(
block: Any,
session_id: Optional[str] = None,
@@ -1619,14 +1587,15 @@ async def _execute_tool_block_impl(
if rejected is not None:
return rejected
mark_authorized()
if bridge_owns_tool:
try:
return await execution_bridge.route_tool(
return await dispatched(execution_bridge.route_tool(
tool,
content,
session_id,
client_runtime_context,
)
))
except asyncio.CancelledError:
raise
except Exception as exc:
@@ -1647,7 +1616,7 @@ async def _execute_tool_block_impl(
)
if tool in _ROUTED_BRIDGE_TOOLS and _client_bridge(client_runtime_context) is not None:
return await _route_tool_via_bridge(tool, content, session_id, client_runtime_context)
return await dispatched(_route_tool_via_bridge(tool, content, session_id, client_runtime_context))
# Background execution: a `bash` block whose first line is the `#!bg`
# marker runs DETACHED — returns a job id immediately so the chat stream
@@ -1657,6 +1626,7 @@ async def _execute_tool_block_impl(
_is_bg, _bg_cmd = _split_bg_marker(content)
if _is_bg and _bg_cmd:
from src import bg_jobs
mark_dispatch()
rec = bg_jobs.launch(_bg_cmd, session_id=session_id, cwd=agent_cwd())
short = _bg_cmd.strip().split(chr(10))[0][:80]
desc = f"bash (background): {short}"
@@ -1682,41 +1652,41 @@ async def _execute_tool_block_impl(
if tool == "generate_image":
from src.ai_interaction import do_generate_image
desc = "generate_image"
result = await do_generate_image(content, session_id=session_id, owner=owner)
result = await dispatched(do_generate_image(content, session_id=session_id, owner=owner))
elif tool in _MCP_TOOL_MAP:
first_line = content.split(chr(10))[0][:80]
desc = f"{tool}: {first_line}"
result = await _call_mcp_tool(tool, content, progress_cb=progress_cb)
result = await dispatched(_call_mcp_tool(tool, content, progress_cb=progress_cb))
elif tool in ("grep", "glob", "ls", "get_workspace", "host_shell"):
# Code-navigation tools — no MCP server; run the direct implementation.
first_line = content.split(chr(10))[0][:80]
desc = f"{tool}: {first_line}"
result = await _direct_fallback(
result = await dispatched(_direct_fallback(
tool,
content,
progress_cb=progress_cb,
owner=owner,
client_runtime_context=client_runtime_context,
) \
)) \
or {"error": f"{tool}: execution failed", "exit_code": 1}
elif tool == "apply_patch" and _tui_host_bridge_patch_url(client_runtime_context):
first_line = content.split(chr(10))[0][:80]
desc = f"{tool}: {first_line}" if first_line else tool
result = await _apply_patch_via_tui_host_bridge(content, client_runtime_context)
result = await dispatched(_apply_patch_via_tui_host_bridge(content, client_runtime_context))
elif tool in ("apply_patch", "todowrite"):
first_line = content.split(chr(10))[0][:80]
desc = f"{tool}: {first_line}" if first_line else tool
result = await _direct_fallback(tool, content, session_id=session_id, owner=owner) \
result = await dispatched(_direct_fallback(tool, content, session_id=session_id, owner=owner)) \
or {"error": f"{tool}: execution failed", "exit_code": 1}
elif tool == "manage_bg_jobs":
# Inspect/kill detached `bash` jobs; needs session_id to scope to chat.
desc = f"manage_bg_jobs: {content.split(chr(10))[0][:80]}"
result = await _direct_fallback(tool, content, session_id=session_id, owner=owner) \
result = await dispatched(_direct_fallback(tool, content, session_id=session_id, owner=owner)) \
or {"error": "manage_bg_jobs: execution failed", "exit_code": 1}
elif tool in ("create_document", "update_document", "edit_document",
"suggest_document", "manage_documents"):
desc = f"{tool}: {content.split(chr(10))[0][:80]}"
result = await _document_tool_dispatch(
result = await dispatched(_document_tool_dispatch(
tool,
content,
session_id,
@@ -1724,14 +1694,14 @@ async def _execute_tool_block_impl(
document_id=approved_document_id or active_document_id,
document_version=approved_document_version,
document_digest=approved_document_digest,
) \
)) \
or {"error": f"{tool}: execution failed", "exit_code": 1}
if tool in ("edit_document", "suggest_document") and "title" in (result or {}):
desc = f"{tool}: {result.get('title', '')}"
elif tool == "search_chats":
query = content.split("\n")[0].strip()
desc = f"search_chats: {query[:80]}"
result = await do_search_chats(query, owner=owner)
result = await dispatched(do_search_chats(query, owner=owner))
elif tool in ("chat_with_model", "ask_teacher", "list_models"):
# Migrated to the agent_tools registry (#3629): dispatched through
# TOOL_HANDLERS with the owner/session ctx these tools need, instead
@@ -1739,7 +1709,7 @@ async def _execute_tool_block_impl(
# src/agent_tools/model_interaction_tools.py.
first_line = content.split(chr(10))[0].strip()[:60]
desc = f"{tool}: {first_line}" if first_line else tool
result = await _document_tool_dispatch(tool, content, session_id, owner) \
result = await dispatched(_document_tool_dispatch(tool, content, session_id, owner)) \
or {"error": f"{tool}: execution failed", "exit_code": 1}
elif tool in ("create_session", "list_sessions", "send_to_session", "manage_session"):
# Migrated to the agent_tools registry (#3629): dispatched through
@@ -1747,101 +1717,101 @@ async def _execute_tool_block_impl(
# live in src/agent_tools/session_tools.py.
first_line = content.split(chr(10))[0].strip()[:60]
desc = f"{tool}: {first_line}" if first_line else tool
result = await _document_tool_dispatch(tool, content, session_id, owner) \
result = await dispatched(_document_tool_dispatch(tool, content, session_id, owner)) \
or {"error": f"{tool}: execution failed", "exit_code": 1}
elif tool in ("pipeline", "manage_memory", "ui_control"):
from src.ai_interaction import dispatch_ai_tool
desc, result = await dispatch_ai_tool(tool, content, session_id, owner=owner)
desc, result = await dispatched(dispatch_ai_tool(tool, content, session_id, owner=owner))
elif tool == "manage_tasks":
desc = "manage_tasks"
result = await do_manage_tasks(content, owner=owner)
result = await dispatched(do_manage_tasks(content, owner=owner))
elif tool == "manage_skills":
desc = "manage_skills"
result = await do_manage_skills(content, owner=owner)
result = await dispatched(do_manage_skills(content, owner=owner))
elif tool == "api_call":
first_line = content.split("\n")[0].strip()[:60]
desc = f"api_call: {first_line}"
result = await do_api_call(content)
result = await dispatched(do_api_call(content))
elif tool in ("manage_endpoints", "manage_mcp", "manage_webhooks", "manage_tokens", "manage_settings"):
# Registry-dispatched (agent_tools.admin_tools); owner threaded for ownership/admin checks.
desc = tool
result = await _direct_fallback(tool, content, owner=owner) \
result = await dispatched(_direct_fallback(tool, content, owner=owner)) \
or {"error": f"{tool}: execution failed", "exit_code": 1}
elif tool == "manage_notes":
desc = "manage_notes"
result = await do_manage_notes(content, owner=owner)
result = await dispatched(do_manage_notes(content, owner=owner))
elif tool == "manage_calendar":
desc = "manage_calendar"
result = await do_manage_calendar(content, owner=owner)
result = await dispatched(do_manage_calendar(content, owner=owner))
elif tool == "download_model":
desc = "download_model"
result = await do_download_model(content, owner=owner)
result = await dispatched(do_download_model(content, owner=owner))
elif tool == "serve_model":
desc = "serve_model"
result = await do_serve_model(content, owner=owner)
result = await dispatched(do_serve_model(content, owner=owner))
elif tool == "list_served_models":
desc = "list_served_models"
result = await do_list_served_models(content, owner=owner)
result = await dispatched(do_list_served_models(content, owner=owner))
elif tool == "stop_served_model":
desc = "stop_served_model"
result = await do_stop_served_model(content, owner=owner)
result = await dispatched(do_stop_served_model(content, owner=owner))
elif tool == "tail_serve_output":
desc = "tail_serve_output"
result = await do_tail_serve_output(content, owner=owner)
result = await dispatched(do_tail_serve_output(content, owner=owner))
elif tool == "list_downloads":
desc = "list_downloads"
result = await do_list_downloads(content, owner=owner)
result = await dispatched(do_list_downloads(content, owner=owner))
elif tool == "cancel_download":
desc = "cancel_download"
result = await do_cancel_download(content, owner=owner)
result = await dispatched(do_cancel_download(content, owner=owner))
elif tool == "search_hf_models":
desc = "search_hf_models"
result = await do_search_hf_models(content, owner=owner)
result = await dispatched(do_search_hf_models(content, owner=owner))
elif tool == "list_cached_models":
desc = "list_cached_models"
result = await do_list_cached_models(content, owner=owner)
result = await dispatched(do_list_cached_models(content, owner=owner))
elif tool == "app_api":
desc = "app_api"
result = await do_app_api(content, owner=owner)
result = await dispatched(do_app_api(content, owner=owner))
elif tool == "list_serve_presets":
desc = "list_serve_presets"
result = await do_list_serve_presets(content, owner=owner)
result = await dispatched(do_list_serve_presets(content, owner=owner))
elif tool == "serve_preset":
desc = "serve_preset"
result = await do_serve_preset(content, owner=owner)
result = await dispatched(do_serve_preset(content, owner=owner))
elif tool == "adopt_served_model":
desc = "adopt_served_model"
result = await do_adopt_served_model(content, owner=owner)
result = await dispatched(do_adopt_served_model(content, owner=owner))
elif tool == "list_cookbook_servers":
desc = "list_cookbook_servers"
result = await do_list_cookbook_servers(content, owner=owner)
result = await dispatched(do_list_cookbook_servers(content, owner=owner))
elif tool == "edit_image":
desc = "edit_image"
result = await do_edit_image(content, owner=owner)
result = await dispatched(do_edit_image(content, owner=owner))
elif tool == "edit_file":
result = await _direct_fallback(tool, content) or {"error": "edit failed", "exit_code": 1}
result = await dispatched(_direct_fallback(tool, content)) or {"error": "edit failed", "exit_code": 1}
desc = result.get("output") or result.get("error") or "edit_file"
elif tool == "trigger_research":
desc = "trigger_research"
result = await do_trigger_research(content, owner=owner, chat_session_id=session_id)
result = await dispatched(do_trigger_research(content, owner=owner, chat_session_id=session_id))
elif tool == "manage_research":
desc = "manage_research"
result = await do_manage_research(content, owner=owner)
result = await dispatched(do_manage_research(content, owner=owner))
elif tool == "resolve_contact":
desc = "resolve_contact"
result = await do_resolve_contact(content, owner=owner)
result = await dispatched(do_resolve_contact(content, owner=owner))
elif tool == "manage_contact":
desc = "manage_contact"
result = await do_manage_contact(content, owner=owner)
result = await dispatched(do_manage_contact(content, owner=owner))
elif tool == "vault_search":
desc = "vault_search"
result = await do_vault_search(content, owner=owner)
result = await dispatched(do_vault_search(content, owner=owner))
elif tool == "vault_get":
desc = "vault_get"
result = await do_vault_get(content, owner=owner)
result = await dispatched(do_vault_get(content, owner=owner))
elif tool == "vault_unlock":
desc = "vault_unlock"
result = await do_vault_unlock(content, owner=owner)
result = await dispatched(do_vault_unlock(content, owner=owner))
elif tool in BUILTIN_EMAIL_TOOLS:
# Bare email tool name from fenced-block models (e.g. Ollama) — route to MCP email server.
# Non-admin owners never reach here: BUILTIN_EMAIL_TOOLS ⊆ NON_ADMIN_BLOCKED_TOOLS,
@@ -1887,7 +1857,7 @@ async def _execute_tool_block_impl(
if session_id:
args = dict(args)
args[_EMAIL_MCP_SESSION_ARG] = session_id
result = await mcp.call_tool(qualified, args)
result = await dispatched(mcp.call_tool(qualified, args))
else:
result = {"error": "MCP manager not available", "exit_code": 1}
elif tool.startswith("mcp__"):
@@ -1906,7 +1876,7 @@ async def _execute_tool_block_impl(
if session_id:
args = dict(args)
args[_EMAIL_MCP_SESSION_ARG] = session_id
result = _normalize_mcp_text_error(await mcp.call_tool(tool, args))
result = _normalize_mcp_text_error(await dispatched(mcp.call_tool(tool, args)))
else:
desc = f"mcp: {tool}"
result = {"error": "MCP manager not available", "exit_code": 1}
@@ -1915,7 +1885,7 @@ async def _execute_tool_block_impl(
elif tool in dynamic_handlers:
first_line = content.split(chr(10))[0][:80]
desc = f"registry: {tool} {first_line}".strip()
res = await _direct_fallback(
res = await dispatched(_direct_fallback(
tool,
content,
progress_cb=progress_cb,
@@ -1924,7 +1894,7 @@ async def _execute_tool_block_impl(
client_runtime_context=client_runtime_context,
disabled_tools=disabled_tools,
tool_policy=tool_policy,
)
))
if isinstance(res, tuple):
desc, result = res
+8
View File
@@ -15,6 +15,14 @@ reference; that file is the standard the refactor works toward.
## Running focused subsets (taxonomy markers)
The shared static-server fixture defaults to loopback port 7011 and refuses an
occupied port rather than reusing another checkout's server. For focused tests
that do not load that fixed browser URL, use `ODYSSEUS_TEST_STATIC_PORT=0` to
allocate an ephemeral port. This permits direct subprocess/isolation tests on a
host already serving the application without stopping or changing that service.
Browser tests that hard-code port 7011 still need that port in their own isolated
network namespace; do not run them against an unrelated live server.
`tests/conftest.py` tags every test at collection time with two markers derived
from its filename by `tests/_taxonomy.py`: an `area_*` marker (e.g.
`area_security`) and a finer `sub_*` marker (e.g. `sub_owner_scope`). This adds
+2
View File
@@ -135,6 +135,8 @@ def _serve_test_static():
allow_reuse_address = True
requested = int(os.environ.get("ODYSSEUS_TEST_STATIC_PORT") or 0)
if not 0 <= requested <= 65535:
raise ValueError("ODYSSEUS_TEST_STATIC_PORT must be between 0 and 65535")
try:
server = _Server(("127.0.0.1", requested), _Handler)
except OSError as exc:
+10
View File
@@ -0,0 +1,10 @@
"""Dispatcher doubles must simulate the receipt boundary as well as the result."""
from src.agent_runtime.journal import mark_dispatch, record_action
def authoritative_executor(function):
@record_action
async def execute(block, *args, **kwargs):
mark_dispatch()
return await function(block, *args, **kwargs)
return execute
+3 -2
View File
@@ -1,4 +1,5 @@
import asyncio
from types import SimpleNamespace
def test_new_tmux_session_forwards_runtime_python_environment(monkeypatch):
@@ -188,7 +189,7 @@ def test_direct_bash_subprocess_has_closed_stdin(monkeypatch, tmp_path):
from src import tool_execution
captured = {}
sentinel = object()
sentinel = SimpleNamespace(pid=12345)
async def fake_create(command, **kwargs):
captured.update(kwargs)
@@ -258,7 +259,7 @@ def test_bash_allows_unicode_ffmpeg_drawtext_with_explicit_fontfile(monkeypatch,
from src import tool_execution
captured = {}
sentinel = object()
sentinel = SimpleNamespace(pid=12345)
async def fake_create(command, **kwargs):
captured["command"] = command
+2 -1
View File
@@ -1,6 +1,7 @@
"""Windows execution contract for the agent Bash tool."""
import pytest
from types import SimpleNamespace
from src.agent_tools import subprocess_tools
@@ -85,7 +86,7 @@ async def test_windows_bash_does_not_use_a_stray_tmux_executable(monkeypatch):
async def fake_create(command, **kwargs):
captured["command"] = command
captured["kwargs"] = kwargs
return object()
return SimpleNamespace(pid=12345)
async def fake_stream(_process, **_kwargs):
return "ok", "", 0, False
+75 -24
View File
@@ -4,6 +4,7 @@ import json
import src.agent_loop as agent_loop
from src.tool_parsing import ToolBlock
from src.tool_capabilities import ToolGateDecision
from tests.runtime_evidence_helpers import authoritative_executor
def _events(chunks):
@@ -83,12 +84,12 @@ def _patch_loop(monkeypatch, responses, captured_kwargs=None):
yield f'data: {json.dumps({"delta": response})}\n\n'
yield "data: [DONE]\n\n"
monkeypatch.setattr(agent_loop, "execute_tool_block", execute)
monkeypatch.setattr(agent_loop, "execute_tool_block", authoritative_executor(execute))
monkeypatch.setattr(agent_loop, "stream_llm_with_fallback", stream)
return lambda: call_index
def _run(instruction, *, max_rounds=4, relevant_tools=None, runtime_context=None):
def _run_chunks(instruction, *, max_rounds=4, relevant_tools=None, runtime_context=None):
async def collect():
return [
chunk
@@ -103,7 +104,11 @@ def _run(instruction, *, max_rounds=4, relevant_tools=None, runtime_context=None
)
]
return _events(asyncio.run(collect()))
return asyncio.run(collect())
def _run(instruction, **kwargs):
return _events(_run_chunks(instruction, **kwargs))
def test_failed_workspace_mutation_attempts_are_not_hidden_by_successful_probe():
@@ -117,19 +122,53 @@ def test_failed_workspace_mutation_attempts_are_not_hidden_by_successful_probe()
assert agent_loop._failed_workspace_mutation_attempts([failed, probe], records) == 1
def test_terminal_completion_repairs_missing_artifact_at_most_twice(monkeypatch):
def test_terminal_completion_missing_artifact_does_not_add_model_rounds(monkeypatch):
calls = _patch_loop(monkeypatch, ["Done without writing anything."])
events = _run("Write answer.json", max_rounds=4)
blocked = [event for event in events if event.get("type") == "completion_blocked"]
assert [event["attempt"] for event in blocked] == [1, 2]
assert calls() == 3
assert blocked == []
assert calls() == 1
decision = next(event["data"] for event in events if event.get("type") == "completion_decision")
assert decision["status"] == "blocked"
assert decision["missing_artifacts"] == ["answer.json"]
def test_slice2_explanatory_request_does_not_add_verification_or_model_rounds(monkeypatch):
calls = _patch_loop(monkeypatch, ['Tests pass when the command exits zero.'])
events = _run('Explain how to write code and then test it.', max_rounds=1)
assert calls() == 1
decision = next(event['data'] for event in events if event.get('type') == 'completion_decision')
assert decision['can_complete'] is True
assert 'The task is incomplete' not in json.dumps(events)
metrics = next(event['data'] for event in events if event.get('type') == 'metrics')
assert not metrics['completion_requirements']['verifier_required']
assert metrics['completion_gate']['additional_provider_calls'] == 0
def test_slice2_fabricated_execution_on_conversational_turn_does_not_add_rounds(monkeypatch):
calls = _patch_loop(monkeypatch, ['I ran pytest and all tests passed.'])
events = _run('Explain what pytest does.', max_rounds=1)
assert calls() == 1
final = next(event['content'] for event in events if event.get('type') == 'final_response')
assert 'I ran pytest' not in final
assert 'The task is incomplete' not in final
metrics = next(event['data'] for event in events if event.get('type') == 'metrics')
assert metrics['completion_gate']['additional_provider_calls'] == 0
def test_slice2_unsupported_test_report_does_not_add_model_rounds(monkeypatch):
calls = _patch_loop(monkeypatch, ['I ran pytest and all 42 tests passed.'])
events = _run('Run pytest.', max_rounds=4, relevant_tools={'bash'})
assert calls() == 1
decision = next(event['data'] for event in events if event.get('type') == 'completion_decision')
assert decision['can_complete'] is False
final = next(event['content'] for event in events if event.get('type') == 'final_response')
assert final.startswith('The task is incomplete:')
assert '42' not in final
def test_failed_trailing_tool_with_planning_prose_continues_artifact_task(monkeypatch):
calls = _patch_loop(
monkeypatch,
@@ -145,7 +184,7 @@ def test_failed_trailing_tool_with_planning_prose_continues_artifact_task(monkey
"exit_code": 1,
}
monkeypatch.setattr(agent_loop, "execute_tool_block", fail_execute)
monkeypatch.setattr(agent_loop, "execute_tool_block", authoritative_executor(fail_execute))
events = _run(
"Create answer.json after inspecting the source",
@@ -155,8 +194,8 @@ def test_failed_trailing_tool_with_planning_prose_continues_artifact_task(monkey
assert calls() > 1
assert any(
event.get("type") == "completion_blocked"
and event.get("decision", {}).get("missing_artifacts") == ["answer.json"]
event.get("type") == "completion_decision"
and event.get("data", {}).get("missing_artifacts") == ["answer.json"]
for event in events
)
@@ -178,7 +217,7 @@ def test_exact_failed_call_is_blocked_across_planning_and_intervening_failure(mo
executed.append(block.content)
return block.tool_type, {"output": f"failed: {block.content}", "exit_code": 1}
monkeypatch.setattr(agent_loop, "execute_tool_block", fail_execute)
monkeypatch.setattr(agent_loop, "execute_tool_block", authoritative_executor(fail_execute))
events = _run(
"Create /tmp_workspace/results after classifying the files",
@@ -218,7 +257,7 @@ def test_exact_failed_call_can_retry_after_successful_workspace_mutation(monkeyp
return block.tool_type, {"output": "written", "exit_code": 0}
return block.tool_type, {"output": "classifier failed", "exit_code": 1}
monkeypatch.setattr(agent_loop, "execute_tool_block", execute)
monkeypatch.setattr(agent_loop, "execute_tool_block", authoritative_executor(execute))
events = _run(
"Create /tmp_workspace/results after repairing and running the classifier",
@@ -260,7 +299,7 @@ def test_terminal_artifact_task_repairs_after_consecutive_failed_batches(monkeyp
"exit_code": 1,
}
monkeypatch.setattr(agent_loop, "execute_tool_block", execute)
monkeypatch.setattr(agent_loop, "execute_tool_block", authoritative_executor(execute))
events = _run(
"Create answer.json and verify it",
@@ -312,7 +351,7 @@ def test_varied_failed_artifact_mutations_have_cumulative_cap(monkeypatch):
"exit_code": 1,
}
monkeypatch.setattr(agent_loop, "execute_tool_block", execute)
monkeypatch.setattr(agent_loop, "execute_tool_block", authoritative_executor(execute))
events = _run(
"Create answer.json and verify it",
@@ -355,7 +394,7 @@ def test_exact_successful_read_is_blocked_until_workspace_changes(monkeypatch):
return block.tool_type, {"output": "written", "exit_code": 0}
return block.tool_type, {"output": "7", "exit_code": 0}
monkeypatch.setattr(agent_loop, "execute_tool_block", execute)
monkeypatch.setattr(agent_loop, "execute_tool_block", authoritative_executor(execute))
events = _run(
"Create answer.json from the inspected workspace",
@@ -378,7 +417,7 @@ def test_exact_successful_read_is_blocked_until_workspace_changes(monkeypatch):
assert any(tool == "write_file" for tool, _ in executed)
def test_terminal_completion_recovers_fenced_body_after_two_repairs(monkeypatch):
def test_missing_evidence_does_not_generate_later_fenced_body_or_mutation(monkeypatch):
executed = []
calls = _patch_loop(
monkeypatch,
@@ -395,20 +434,19 @@ def test_terminal_completion_recovers_fenced_body_after_two_repairs(monkeypatch)
executed.append(block)
return await original_execute(block, *args, **kwargs)
monkeypatch.setattr(agent_loop, "execute_tool_block", record_execute)
monkeypatch.setattr(agent_loop, "execute_tool_block", authoritative_executor(record_execute))
events = _run("Write answer.json", max_rounds=5)
assert [(block.tool_type, block.content) for block in executed] == [
("write_file", 'answer.json\n{"ok": true}'),
]
assert executed == []
assert calls() == 1
decision = next(
event["data"]
for event in events
if event.get("type") == "completion_decision"
)
assert decision["status"] == "satisfied"
assert decision["can_complete"] is True
assert decision["status"] == "blocked"
assert decision["can_complete"] is False
def test_successful_artifact_write_emits_satisfied_completion(monkeypatch):
@@ -497,7 +535,7 @@ def test_verified_artifact_survives_provider_error_during_finish_round(monkeypat
monkeypatch.setattr(agent_loop, "stream_llm_with_fallback", stream)
events = _run(
chunks = _run_chunks(
"Create /workspace/output.html",
max_rounds=5,
relevant_tools={"write_file", "private_browser"},
@@ -510,11 +548,17 @@ def test_verified_artifact_survives_provider_error_during_finish_round(monkeypat
},
)
events = _events(chunks)
assert calls == 2
assert chunks[-1] == 'event: error\ndata: {"status": 504, "error": "stream timeout"}\n\n'
assert not any(chunk.strip() == 'data: [DONE]' for chunk in chunks)
decision = next(event['data'] for event in events if event.get('type') == 'completion_decision')
assert decision['can_complete'] is False
assert decision['status'] == 'failed'
assert not any(event.get("type") == "agent_terminal" for event in events)
final = next(event for event in events if event.get("type") == "final_response")
assert "output.html" in final["content"]
assert "verified" in final["content"].lower()
assert final['content'].startswith('The task is incomplete:')
def test_uninspected_artifact_still_fails_on_provider_error(monkeypatch):
@@ -539,7 +583,7 @@ def test_uninspected_artifact_still_fails_on_provider_error(monkeypatch):
monkeypatch.setattr(agent_loop, "stream_llm_with_fallback", stream)
events = _run(
chunks = _run_chunks(
"Create /workspace/answer.json",
max_rounds=4,
relevant_tools={"write_file"},
@@ -552,7 +596,13 @@ def test_uninspected_artifact_still_fails_on_provider_error(monkeypatch):
},
)
events = _events(chunks)
assert calls == 2
assert chunks[-1] == 'event: error\ndata: {"status": 504, "error": "stream timeout"}\n\n'
assert not any(chunk.strip() == 'data: [DONE]' for chunk in chunks)
decision = next(event['data'] for event in events if event.get('type') == 'completion_decision')
assert decision['can_complete'] is False
assert decision['status'] == 'failed'
terminal = next(
(event for event in events if event.get("type") == "agent_terminal"),
None,
@@ -560,6 +610,7 @@ def test_uninspected_artifact_still_fails_on_provider_error(monkeypatch):
assert terminal is not None, events
assert terminal["data"]["failed"] is True
assert terminal["data"]["failure"]["status"] == 504
assert '[Agent stopped: Model request failed (HTTP 504)]' in terminal['data']['round_texts'][-1]
def test_verified_artifact_gets_only_one_finish_nudge(monkeypatch):
+48 -9
View File
@@ -9,11 +9,13 @@ return, or moves the done-break, could silently flip this. See PR #1999 / #1997.
import asyncio
import json
from pathlib import Path
import pytest
import src.agent_loop as al
from src.tool_capabilities import ToolGateDecision
from src.tool_capabilities import capabilities_for_action
from src.tool_approvals import tool_approval_store
from tests.runtime_evidence_helpers import authoritative_executor
def _collect(gen):
@@ -39,6 +41,16 @@ def _patch_common(monkeypatch):
monkeypatch.setattr(al, "get_setting", lambda key, default=None: default, raising=False)
monkeypatch.setattr(al, "get_mcp_manager", lambda: None, raising=False)
monkeypatch.setattr(al, "estimate_tokens", lambda *a, **k: 10, raising=False)
# The round providers are synthetic. Keep real compaction logic while
# supplying its context window instead of probing the dummy endpoint.
import src.context_compactor as context_compactor
monkeypatch.setattr(context_compactor, "get_context_length", lambda *a, **k: 128_000)
# These fixtures supply the round provider below. Any direct grace-synthesis
# request has no configured response, rather than contacting the fake URL.
async def _unconfigured_direct_provider(*args, **kwargs):
raise RuntimeError("No direct-provider completion configured in this fixture")
yield # Keep the direct-provider async-generator interface.
monkeypatch.setattr(al, "stream_llm", _unconfigured_direct_provider)
# These tests exercise round convergence. Keep the prompt-integrity gate
# out of the fixture so a synthetic tool result does not turn the next
# round into an approval test instead.
@@ -997,6 +1009,7 @@ def test_empty_workspace_round_gets_one_bounded_action_nudge(monkeypatch):
yield 'data: {"delta":"Workspace checked."}\n\n'
yield "data: [DONE]\n\n"
@authoritative_executor
async def _fake_exec(block, *args, **kwargs):
return (block.tool_type, {"output": "/workspace", "exit_code": 0})
@@ -1011,7 +1024,15 @@ def test_empty_workspace_round_gets_one_bounded_action_nudge(monkeypatch):
)))
assert len(seen) == 3
assert any(e.get("delta") == "Workspace checked." for e in events)
decision = next(e["data"] for e in events if e.get("type") == "completion_decision")
assert not decision["can_complete"]
assert "fixture.py" in decision["missing_artifacts"]
assert any(
e.get("type") == "final_response"
and "Workspace checked." in e.get("content", "")
and "incomplete" in e.get("content", "")
for e in events
)
def test_empty_workspace_nudge_names_only_tools_in_active_schema(monkeypatch):
@@ -1205,7 +1226,8 @@ def test_eval_workspace_prompt_uses_host_tool_then_answers(monkeypatch):
assert any("active workspace is /home/tester/project" in e.get("delta", "") for e in events)
def test_tui_coding_turn_recovers_inspects_patches_and_verifies(monkeypatch):
@pytest.mark.parametrize('explicit_verifier', [False, True])
def test_tui_coding_turn_recovers_inspects_patches_and_verifies(monkeypatch, explicit_verifier):
"""A stale compact router must still complete a real coding workflow."""
_patch_common(monkeypatch)
calls = []
@@ -1227,6 +1249,7 @@ def test_tui_coding_turn_recovers_inspects_patches_and_verifies(monkeypatch):
},
}
@authoritative_executor
async def _fake_exec(block, *args, **kwargs):
executed.append((block.tool_type, block.content))
if block.tool_type == "host_shell":
@@ -1300,7 +1323,8 @@ def test_tui_coding_turn_recovers_inspects_patches_and_verifies(monkeypatch):
"deepseek-v4-flash",
[{
"role": "user",
"content": "Fix the parser in this local TUI project and run the tests.",
"content": "Fix the parser in this local TUI project and run "
+ ("python -m pytest -q." if explicit_verifier else "the tests."),
}],
max_rounds=6,
relevant_tools={"get_workspace", "ls", "host_shell", "apply_patch", "edit_file", "todowrite"},
@@ -1316,14 +1340,20 @@ def test_tui_coding_turn_recovers_inspects_patches_and_verifies(monkeypatch):
assert not any(tool in {"get_workspace", "ls"} for tool, _ in executed)
# Invalid backend/container tool names are replaced with the authoritative
# host-shell recovery in the same round, so no extra clarification round
# should be required before the patch and verification turns.
assert len(calls) == 3
# should be required before the patch and verification turns. An opaque
# conditional fallback still needs synthesis and cannot attest tests.
assert len(calls) == (3 if explicit_verifier else 4)
decision = next(event['data'] for event in events if event.get('type') == 'completion_decision')
assert decision['can_complete'] is explicit_verifier
assert (decision['status'] == 'verified') is explicit_verifier
assert any(
event.get("type") == "final_response"
and "Verification:" in event.get("content", "")
and "passed" in event.get("content", "")
for event in events
)
) is explicit_verifier
terminal = next(event['data'] for event in events if event.get('type') == 'metrics')
assert terminal['completion_gate']['additional_provider_calls'] == 0
assert not any(event.get("type") in {"rounds_exhausted", "loop_breaker_triggered"} for event in events)
@@ -1339,6 +1369,7 @@ def test_qwen_tui_coding_summary_reports_verification_retry(monkeypatch):
"runtime_execution_contract": {"local_workspace_tasks": "use_host_shell_bridge"},
}
@authoritative_executor
async def _fake_exec(block, *args, **kwargs):
nonlocal test_attempts
executed.append(block.tool_type)
@@ -1390,7 +1421,8 @@ def test_qwen_tui_coding_summary_reports_verification_retry(monkeypatch):
assert "`pytest -q` passed" in summary
def test_tui_coding_turn_replaces_duplicate_edit_with_requested_tests(monkeypatch):
@pytest.mark.parametrize('explicit_verifier', [False, True])
def test_tui_coding_turn_replaces_duplicate_edit_with_requested_tests(monkeypatch, explicit_verifier):
"""A successful edit followed by another edit must converge on real tests."""
_patch_common(monkeypatch)
executed = []
@@ -1407,6 +1439,7 @@ def test_tui_coding_turn_replaces_duplicate_edit_with_requested_tests(monkeypatc
},
}
@authoritative_executor
async def _fake_exec(block, *args, **kwargs):
executed.append((block.tool_type, block.content))
if block.tool_type == "read_file":
@@ -1450,7 +1483,8 @@ def test_tui_coding_turn_replaces_duplicate_edit_with_requested_tests(monkeypatc
"role": "user",
"content": (
"A regression was introduced in nested backend error handling. "
"Find the cause, fix it with a scoped change, and run the relevant tests."
"Find the cause, fix it with a scoped change, and run "
+ ("pytest -q." if explicit_verifier else "the relevant tests.")
),
}],
max_rounds=6,
@@ -1465,12 +1499,15 @@ def test_tui_coding_turn_replaces_duplicate_edit_with_requested_tests(monkeypatc
]
verification_command = json.loads(executed[-1][1])["command"]
assert "pytest" in verification_command
decision = next(event['data'] for event in events if event.get('type') == 'completion_decision')
assert decision['can_complete'] is explicit_verifier
assert (decision['status'] == 'verified') is explicit_verifier
assert any(
"Verification:" in event.get("content", "")
and "passed" in event.get("content", "")
for event in events
if event.get("type") == "final_response"
)
) is explicit_verifier
def test_failed_forced_verifier_allows_an_adapted_test_command(monkeypatch):
@@ -1491,6 +1528,7 @@ def test_failed_forced_verifier_allows_an_adapted_test_command(monkeypatch):
},
}
@authoritative_executor
async def _fake_exec(block, *args, **kwargs):
nonlocal host_attempts
executed.append((block.tool_type, block.content))
@@ -3363,6 +3401,7 @@ def test_final_prose_with_missing_artifact_enters_recovery(monkeypatch):
requests = []
executed = []
@authoritative_executor
async def _fake_exec(block, *args, **kwargs):
executed.append(block)
if block.tool_type == "write_file":
+2 -1
View File
@@ -1122,7 +1122,8 @@ def test_native_host_shell_call_runs_through_bridge_and_threads_result(monkeypat
# at import time, so a fresh `import src.tool_execution` here can bind a
# different module object than the execute_tool_block agent_loop calls —
# patching that fresh copy silently no-ops in full-suite runs.
_dispatch_globals = al.execute_tool_block.__globals__
from inspect import unwrap
_dispatch_globals = unwrap(al.execute_tool_block).__globals__
monkeypatch.setitem(
_dispatch_globals,
"owner_is_admin_or_single_user",
+457
View File
@@ -0,0 +1,457 @@
"""Provider failure is the final frame, after gated output and diagnostics."""
import asyncio
from inspect import signature
import json
import pytest
from src.agent_runtime.completion import with_completion_gate
from src.agent_runtime.completion import completion_answer
from src.agent_evidence import CompletionRequirements, EvidenceLedger, infer_completion_requirements
from src.agent_runtime.journal import current_journal
from src.tool_types import ToolBlock
from tests.runtime_evidence_helpers import authoritative_executor
ERROR = 'event: error\ndata: {"status": 504, "error": {"message": "stream timeout"}, "fallback_eligible": false}\n\n'
DONE = 'data: [DONE]\n\n'
@pytest.mark.asyncio
@pytest.mark.parametrize('terminal_texts', [[], ['', ''], ['Recovered answer with literal [DONE] text.']])
async def test_terminal_round_retraction_does_not_resurrect_buffered_drafts(terminal_texts):
@with_completion_gate
async def stream(messages):
yield _event({'delta': 'Considering the next step.', 'thinking': True})
yield _event({'delta': 'Now I need to execute the rejected draft.'})
yield _event({'type': 'metrics', 'data': {'round_texts': terminal_texts}})
yield DONE
chunks = [chunk async for chunk in stream([{'role': 'user', 'content': 'Create answer.txt.'}])]
assert 'rejected draft' not in ''.join(chunks)
assert any('Considering the next step.' in chunk for chunk in chunks)
if terminal_texts and terminal_texts[0]:
assert any(terminal_texts[0] in chunk for chunk in chunks)
assert chunks.count(DONE) == 1
@pytest.mark.asyncio
async def test_terminal_round_text_cannot_override_an_explicit_final_response():
@with_completion_gate
async def stream(messages):
yield _event({'type': 'final_response', 'content': 'The explicit final answer.'})
yield _event({'type': 'metrics', 'data': {'round_texts': ['Earlier draft.']}})
yield DONE
chunks = [chunk async for chunk in stream([])]
final = next(data for _, data in _frames(chunks) if isinstance(data, dict) and data.get('type') == 'final_response')
assert final['content'] == 'The explicit final answer.'
@pytest.mark.asyncio
async def test_revised_terminal_prose_still_cannot_attest_execution():
@with_completion_gate
async def stream(messages):
yield _event({'delta': 'Earlier draft.'})
yield _event({'type': 'metrics', 'data': {'round_texts': ['All tests passed.']}})
yield DONE
chunks = [chunk async for chunk in stream([{'role': 'user', 'content': 'Create answer.txt and run the tests.'}])]
assert not _decision(chunks)['can_complete']
assert 'All tests passed.' not in ''.join(chunks)
@pytest.mark.asyncio
async def test_provider_error_preserves_partial_content_despite_empty_terminal_rounds():
@with_completion_gate
async def stream(messages):
yield _event({'type': 'tool_start', 'tool': 'read_file'})
yield _event({'delta': 'Safe partial result.'})
yield _event({'type': 'metrics', 'data': {'round_texts': []}})
yield ERROR
yield DONE
chunks = [chunk async for chunk in stream([])]
assert _labels(chunks) == ['tool_start', 'final_response', 'completion_decision', 'metrics', 'error']
assert any('Safe partial result.' in chunk for chunk in chunks)
assert chunks[-1] == ERROR
assert DONE not in chunks
def _event(payload):
return 'data: ' + json.dumps(payload) + '\n\n'
def _frames(chunks):
"""Decode network chunks without losing named error frames or [DONE]."""
pending = ''
for chunk in chunks:
pending += chunk
while '\n\n' in pending:
frame, pending = pending.split('\n\n', 1)
lines = frame.splitlines()
event = next((line[7:] for line in lines if line.startswith('event: ')), 'message')
payload = '\n'.join(line[6:] for line in lines if line.startswith('data: '))
yield event, payload if payload == '[DONE]' else json.loads(payload)
assert not pending, 'incomplete SSE frame'
def _labels(chunks):
return [event if event != 'message' else (
'done' if data == '[DONE]' else data.get('type', 'delta')
) for event, data in _frames(chunks)]
def _decision(chunks):
return next(data['data'] for event, data in _frames(chunks)
if event == 'message' and isinstance(data, dict)
and data.get('type') == 'completion_decision')
@authoritative_executor
async def _successful_tool(block):
return block.tool_type, {'exit_code': 0, 'output': 'OK'}
@pytest.mark.asyncio
async def test_bare_error_preserves_original_frame_without_success_output():
@with_completion_gate
async def stream(messages):
yield ERROR
yield DONE
assert [chunk async for chunk in stream([])] == [ERROR]
@pytest.mark.asyncio
@pytest.mark.parametrize('partial', ['', 'The parser checks the header first.'])
async def test_provider_error_releases_partial_then_decision_terminal_and_original_error(partial):
closed = []
@with_completion_gate
async def stream(messages):
try:
yield _event({'type': 'tool_start', 'tool': 'read_file'})
if partial:
yield _event({'delta': partial})
yield ERROR
yield _event({'type': 'agent_terminal', 'data': {
'failed': True, 'failure': {'status': 504},
'round_texts': ['Earlier diagnostic', partial + '\n[Agent stopped]'],
}})
yield DONE
finally:
closed.append(current_journal() is not None)
chunks = [chunk async for chunk in stream([])]
assert _labels(chunks) == [
'tool_start', 'final_response', 'completion_decision', 'agent_terminal', 'error',
], _labels(chunks)
assert chunks[-1] == ERROR
assert DONE not in chunks
assert _decision(chunks)['can_complete'] is False
assert _decision(chunks)['status'] == 'failed'
final = next(data for event, data in _frames(chunks)
if event == 'message' and data.get('type') == 'final_response')
assert final['content'].startswith('The task is incomplete:')
assert partial in final['content']
assert closed == [True]
assert current_journal() is None
@pytest.mark.asyncio
@pytest.mark.parametrize('successful_tool', [False, True])
@pytest.mark.parametrize('earlier_status', [None, 'awaiting_user', 'exhausted'])
async def test_provider_failure_overrides_even_successful_execution(successful_tool, earlier_status):
@with_completion_gate
async def stream(messages):
if successful_tool:
await _successful_tool(ToolBlock('bash', 'python -m unittest'))
if earlier_status:
yield _event({'type': 'completion_decision', 'data': {'status': earlier_status}})
yield _event({'delta': 'The response is partial.'})
yield ERROR
yield _event({'type': 'metrics', 'data': {}})
chunks = [chunk async for chunk in stream([])]
decision = _decision(chunks)
assert decision['can_complete'] is False, decision
assert decision['status'] == 'failed'
if successful_tool:
metrics = next(data['data'] for event, data in _frames(chunks)
if event == 'message' and data.get('type') == 'metrics')
assert any(e['authoritative'] and e['success'] for e in metrics['evidence_events'])
assert chunks[-1] == ERROR
@pytest.mark.asyncio
async def test_error_after_final_response_does_not_add_calls_or_success_done():
invocations = []
@with_completion_gate
async def stream(messages, workspace=None, client_runtime_context=None):
invocations.append(1)
yield _event({'type': 'final_response', 'content': 'The header contains three fields.'})
yield DONE
yield ERROR
chunks = [chunk async for chunk in stream([])]
assert _labels(chunks) == ['final_response', 'completion_decision', 'error']
assert invocations == [1]
assert str(signature(stream)) == '(messages, workspace=None, client_runtime_context=None)'
assert DONE not in chunks
@pytest.mark.asyncio
@pytest.mark.parametrize('terminal_kind', ['agent_terminal', 'metrics'])
async def test_failed_terminal_diagnostics_survive_answer_replacement(terminal_kind):
diagnostics = ['Earlier tool failure and retry', 'All tests passed.\n[Agent stopped: HTTP 504]']
@with_completion_gate
async def stream(messages):
yield _event({'delta': 'All tests passed.'})
yield ERROR
yield _event({'type': terminal_kind, 'data': {
'failed': True, 'failure': {'status': 504, 'message': 'Model request failed'},
'round_texts': diagnostics, 'round_models': ['first-model', 'failed-model'],
}})
chunks = [chunk async for chunk in stream([])]
terminal = next(data['data'] for event, data in _frames(chunks)
if event == 'message' and data.get('type') == terminal_kind)
# Diagnostics and the failure note survive; the rejected claim does not,
# because round_texts are rendered again when the turn is reloaded.
assert terminal['round_texts'] == ['Earlier tool failure and retry', '[Agent stopped: HTTP 504]']
assert terminal['round_models'] == ['first-model', 'failed-model']
assert terminal['failure'] == {'status': 504, 'message': 'Model request failed'}
assert terminal['failed'] is True
assert terminal['completion_decision'] == _decision(chunks)
assert terminal['completion_gate']['answer_replaced'] is True
assert terminal['completion_gate']['additional_provider_calls'] == 0
assert _labels(chunks).index(terminal_kind) < _labels(chunks).index('error')
@pytest.mark.asyncio
async def test_provider_failure_round_texts_cannot_replay_removed_claim_after_reload():
claim = 'I created report.md and all tests passed.'
note = '[Agent stopped: Model request failed (HTTP 504)]'
@with_completion_gate
async def stream(messages):
yield _event({'type': 'tool_start', 'tool': 'read_file'})
yield _event({'delta': 'Inspected the layout. ' + claim})
yield ERROR
yield _event({'type': 'agent_terminal', 'data': {
'failed': True, 'failure': {'status': 504, 'message': 'Model request failed'},
'tool_events': [{'round': 1, 'tool': 'read_file'}],
'round_texts': ['Inspected the layout. ' + claim, 'Retrying the build.\n\n' + note],
}})
yield DONE
chunks = [chunk async for chunk in stream([{'role': 'user', 'content': 'create report.md and run the tests'}])]
assert _labels(chunks) == [
'tool_start', 'final_response', 'completion_decision', 'agent_terminal', 'error',
], _labels(chunks)
assert DONE not in chunks
live = next(data['content'] for event, data in _frames(chunks)
if event == 'message' and data.get('type') == 'final_response')
terminal = next(data['data'] for event, data in _frames(chunks)
if event == 'message' and data.get('type') == 'agent_terminal')
# The chat route persists this metadata and the renderer rebuilds one bubble
# per round from it, so every persisted round is presentation.
persisted = terminal['round_texts']
assert persisted == ['Inspected the layout.', 'Retrying the build.\n\n' + note]
for text in [live, *persisted]:
assert 'tests passed' not in text and 'created report.md' not in text
assert 'Inspected the layout.' in live
assert terminal['completion_decision']['status'] == 'failed'
@pytest.mark.asyncio
async def test_error_boundary_is_independent_of_network_chunking():
@with_completion_gate
async def stream(messages):
yield _event({'delta': 'Partial explanation.'})
yield ERROR
yield _event({'type': 'agent_terminal', 'data': {'failed': True}})
chunks = [chunk async for chunk in stream([])]
wire = ''.join(chunks)
expected = list(_frames(chunks))
for delivered in [chunks, [wire], list(wire)]:
# A client stops consuming on the first error, regardless of chunking.
visible = []
for frame in _frames(delivered):
visible.append(frame)
if frame[0] == 'error':
break
assert visible == expected
assert visible[-2][1]['type'] == 'agent_terminal'
@pytest.mark.asyncio
@pytest.mark.parametrize('after_error', [False, True])
async def test_cancellation_closes_inner_stream_without_releasing_completion(after_error):
progress_seen = asyncio.Event()
closed = []
chunks = []
@with_completion_gate
async def stream(messages):
try:
yield _event({'delta': 'Tests passed.'})
if after_error:
yield ERROR
yield _event({'type': 'tool_start', 'tool': 'bash'})
await asyncio.Event().wait()
finally:
closed.append(current_journal() is not None)
async def collect():
async for chunk in stream([]):
chunks.append(chunk)
if chunk == _event({'type': 'tool_start', 'tool': 'bash'}):
progress_seen.set()
task = asyncio.create_task(collect())
try:
await asyncio.wait_for(progress_seen.wait(), timeout=5)
task.cancel()
with pytest.raises(asyncio.CancelledError):
await task
finally:
if not task.done():
task.cancel()
await asyncio.gather(task, return_exceptions=True)
assert _labels(chunks) == ['tool_start']
assert closed == [True]
assert current_journal() is None
@pytest.mark.parametrize('prose', [
'Tests pass when the command exits zero.',
'If all tests are passing, merge the branch.',
'Tests passed if the command exited zero.',
'The documentation says "5 passed".',
'The documentation says "Tests: FAIL" or "Tests: PASS".',
'You can run pytest to verify this.',
'A successful test run should show no failures.',
'For example, I created the file and updated config.py.',
'If I updated config.py, I would run pytest.',
'Imagine I ran the tests and all 42 passed.',
'Done is the label for a finished item.',
'```text\nI ran pytest and all 42 passed.\n```',
'Run pytest until there are no failures.',
])
def test_slice2_explanatory_prose_is_not_a_current_run_claim(prose):
ledger = EvidenceLedger()
answer, reason = completion_answer(prose, ledger, ledger.evaluate())
assert answer == prose
assert not reason
@pytest.mark.parametrize('instruction', [
'Explain how to write code and then test it.',
'Summarise this and check for typos.',
'Explain how to update config.py and then verify it.',
'Show an example of creating answer.json and checking it.',
'The documentation says "run pytest and create answer.json".',
'If you run pytest, the tests should pass.',
])
def test_slice2_explanatory_request_has_no_execution_requirements(instruction):
requirements = infer_completion_requirements(instruction)
assert requirements.required_artifacts == ()
assert not requirements.verifier_required
assert not requirements.executable_verifier_available
@pytest.mark.parametrize('instruction', [
'Run the tests.', 'Please run pytest.', 'Can you run the test suite?',
])
def test_slice2_explicit_test_execution_requires_a_verifier(instruction):
requirements = infer_completion_requirements(instruction)
assert requirements.verifier_required
assert not EvidenceLedger(requirements).evaluate().can_complete
@pytest.mark.parametrize('claim', [
'I ran the tests.', 'The tests passed.', '42 tests passed.',
'I created the file.', 'I updated config.py successfully.',
])
def test_slice2_execution_obligation_rejects_unsupported_claims(claim):
ledger = EvidenceLedger(CompletionRequirements(required_artifacts=('config.py',)))
answer, reason = completion_answer(claim, ledger, ledger.evaluate())
assert reason
assert answer.startswith('The task is incomplete:')
assert claim not in answer
@pytest.mark.parametrize('claim', [
'I ran pytest to see if the tests passed.',
'I updated config.py as an example.',
'I ran pytest and should update config.py next.',
'config.py was updated successfully.',
])
def test_slice2_subordinate_explanation_cannot_hide_a_direct_execution_report(claim):
ledger = EvidenceLedger()
answer, reason = completion_answer(claim, ledger, ledger.evaluate())
assert reason
assert claim not in answer
assert 'The task is incomplete' not in answer
@pytest.mark.asyncio
async def test_slice2_client_dictionary_cannot_attest_execution():
@with_completion_gate
async def stream(messages, client_runtime_context=None):
yield _event({'delta': 'I ran pytest and all tests passed.'})
yield DONE
context = {'execution_obligation': True, 'execution_verified': True,
'evidence_events': [{'tool': 'bash', 'command': 'pytest', 'exit_code': 0}]}
chunks = [chunk async for chunk in stream(
[{'role': 'user', 'content': 'Explain test output.'}], client_runtime_context=context)]
final = next(data['content'] for event, data in _frames(chunks)
if event == 'message' and data.get('type') == 'final_response')
assert 'I ran pytest' not in final
assert 'The task is incomplete' not in final
assert _decision(chunks)['can_complete'] is True
@pytest.mark.asyncio
async def test_slice2_conversational_fabrication_is_corrected_without_execution_incomplete():
invocations = []
@with_completion_gate
async def stream(messages):
invocations.append(1)
yield _event({'delta': 'The function returns a boolean. I ran pytest and all tests passed.'})
yield _event({'type': 'metrics', 'data': {}})
yield DONE
chunks = [chunk async for chunk in stream([{'role': 'user', 'content': 'Explain the function.'}])]
final = next(data['content'] for event, data in _frames(chunks)
if event == 'message' and data.get('type') == 'final_response')
assert 'The function returns a boolean.' in final
assert 'I ran pytest' not in final
assert 'The task is incomplete' not in final
assert _decision(chunks)['can_complete'] is True
assert invocations == [1]
metrics = next(data['data'] for event, data in _frames(chunks)
if event == 'message' and data.get('type') == 'metrics')
assert metrics['completion_gate']['additional_provider_calls'] == 0
@pytest.mark.asyncio
async def test_slice2_quoted_example_does_not_hide_an_unsupported_report():
@with_completion_gate
async def stream(messages):
yield _event({'delta': 'The docs say "5 passed". I ran pytest.'})
yield DONE
chunks = [chunk async for chunk in stream([{'role': 'user', 'content': 'Explain pytest output.'}])]
final = next(data['content'] for event, data in _frames(chunks)
if event == 'message' and data.get('type') == 'final_response')
assert 'The docs say "5 passed".' in final
assert 'I ran pytest' not in final
assert 'The task is incomplete' not in final
+3
View File
@@ -2331,6 +2331,9 @@ def test_multi_round_agent_uses_only_selected_model(monkeypatch):
yield f'data: {json.dumps({"delta": "done"})}\n\n'
yield "data: [DONE]\n\n"
from tests.runtime_evidence_helpers import authoritative_executor
@authoritative_executor
async def fake_execute(block, *args, **kwargs):
return "bash", {"output": "ok", "exit_code": 0}
+158
View File
@@ -0,0 +1,158 @@
"""Headless consumers present the completion gate's answer and close its stream."""
import asyncio
import json
import sys
import types
from types import SimpleNamespace
import pytest
from src.agent_runtime.completion import with_completion_gate
from src.agent_runtime.journal import current_journal
from src.teacher_escalation import with_teacher_takeover
CLAIM = 'I created report.md and all tests passed.'
def _event(payload):
return 'data: ' + json.dumps(payload) + '\n\n'
def _task():
return SimpleNamespace(
crew_member_id=None, endpoint_url='http://ep/v1', model='m',
session_id='s', owner='admin', prompt='create report.md and run the tests',
name='job', max_steps=5, character_id=None,
)
def _gated_loop(released):
"""Real gate and takeover adapters around a loop that over-claims."""
@with_teacher_takeover
@with_completion_gate
async def stream_agent_loop(*args, messages=None, client_runtime_context=None, **kwargs):
yield _event({'delta': 'Inspected the layout. ' + CLAIM})
yield _event({'type': 'metrics', 'data': {}})
yield 'data: [DONE]\n\n'
async def recording(*args, **kwargs):
async for chunk in stream_agent_loop(*args, **kwargs):
if chunk.startswith('data: {') and '"final_response"' in chunk:
released.append(json.loads(chunk[6:])['content'])
yield chunk
return recording
async def test_scheduler_result_is_the_gated_replacement_without_grace_call(monkeypatch):
from src.task_scheduler import TaskScheduler
released = []
grace_calls = []
async def grace(*args, **kwargs):
grace_calls.append(kwargs)
return 'ungated summary: all tests passed'
monkeypatch.setattr('src.agent_loop.stream_agent_loop', _gated_loop(released))
monkeypatch.setattr('src.task_endpoint.resolve_task_candidates', lambda **kwargs: [])
monkeypatch.setattr('src.task_endpoint.task_llm_call_async', grace)
result = await TaskScheduler(session_manager=None)._run_agent_loop(
'http://ep/v1', 'model', _task(), 's')
assert len(released) == 1
assert result == released[0].strip()
assert 'Inspected the layout.' in result
assert 'tests passed' not in result
assert grace_calls == []
def test_background_followup_prose_is_the_gated_replacement(monkeypatch):
from src import bg_monitor
released = []
agent_loop = types.ModuleType('src.agent_loop')
agent_loop.stream_agent_loop = _gated_loop(released)
monkeypatch.setitem(sys.modules, 'src.agent_loop', agent_loop)
sess = SimpleNamespace(endpoint_url='http://example.test', model='model',
headers=None, context_length=0, id='s1', owner='owner')
full, _ = asyncio.run(bg_monitor._drain_agent(
sess, [{'role': 'user', 'content': 'create report.md and run the tests'}]))
assert len(released) == 1
assert full == released[0]
assert 'tests passed' not in full
@pytest.mark.parametrize('consumer', ['scheduler', 'background'])
def test_later_answer_supersedes_earlier_replacement(monkeypatch, consumer):
async def stream_agent_loop(*args, **kwargs):
yield _event({'type': 'final_response', 'content': 'Earlier summary.'})
yield _event({'delta': 'Final '})
yield _event({'delta': 'answer.'})
yield 'data: [DONE]\n\n'
if consumer == 'scheduler':
from src.task_scheduler import TaskScheduler
monkeypatch.setattr('src.agent_loop.stream_agent_loop', stream_agent_loop)
monkeypatch.setattr('src.task_endpoint.resolve_task_candidates', lambda **kwargs: [])
result = asyncio.run(TaskScheduler(session_manager=None)._run_agent_loop(
'http://ep/v1', 'model', _task(), 's'))
else:
from src import bg_monitor
agent_loop = types.ModuleType('src.agent_loop')
agent_loop.stream_agent_loop = stream_agent_loop
monkeypatch.setitem(sys.modules, 'src.agent_loop', agent_loop)
sess = SimpleNamespace(endpoint_url='http://example.test', model='model',
headers=None, context_length=0, id='s1')
result, _ = asyncio.run(bg_monitor._drain_agent(sess, []))
assert result == 'Final answer.'
async def test_scheduler_approval_pause_closes_gated_stream_in_its_own_context(monkeypatch):
from src.task_scheduler import TaskScheduler
closed = []
lineage = []
@with_teacher_takeover
@with_completion_gate
async def paused_loop(*args, messages=None, client_runtime_context=None, **kwargs):
try:
yield _event({'type': 'tool_output', 'tool': 'bash', 'output': 'Waiting for an exact user approval.',
'ask_user': {'kind': 'tool_approval', 'approval_id': 'missing'}})
yield _event({'delta': 'not reached'})
finally:
closed.append(current_journal() is not None)
@with_teacher_takeover
@with_completion_gate
async def later_loop(*args, messages=None, client_runtime_context=None, **kwargs):
yield _event({'delta': 'Later run.'})
yield _event({'type': 'metrics', 'data': {}})
yield 'data: [DONE]\n\n'
async def later_run():
async for chunk in later_loop(messages=[{'role': 'user', 'content': 'x'}]):
if chunk.startswith('data: {') and '"metrics"' in chunk:
lineage.append(json.loads(chunk[6:])['data']['parent_run_id'])
monkeypatch.setattr('src.agent_loop.stream_agent_loop', paused_loop)
monkeypatch.setattr('src.task_endpoint.resolve_task_candidates', lambda **kwargs: [])
result = await TaskScheduler(session_manager=None)._run_agent_loop(
'http://ep/v1', 'model', _task(), 's')
assert 'paused safely' in result
# Closed during the pause, while its own journal was still bound.
assert closed == [True]
assert current_journal() is None
# Neither a chained task (which copies this context) nor a later run in
# this task inherits the paused run's journal as its parent.
chained = asyncio.create_task(later_run())
await chained
await later_run()
assert lineage == [None, None]
+18 -4
View File
@@ -120,21 +120,35 @@ def test_disabled_generation_is_not_restored():
@pytest.mark.asyncio
async def test_generation_dispatch_uses_owner_aware_backend(monkeypatch):
@pytest.mark.parametrize('exit_code', [None, 0])
async def test_generation_dispatch_uses_owner_aware_backend(monkeypatch, exit_code):
from types import SimpleNamespace
from src import ai_interaction, tool_execution
from src.agent_runtime.journal import ActionJournal, bind_journal
calls = []
async def generate(content, **kwargs):
calls.append((content, kwargs))
return {'image_url': '/api/generated-image/test.png', 'image_id': 'test'}
result = {'image_url': '/api/generated-image/test.png', 'image_id': 'test'}
if exit_code is not None:
result['exit_code'] = exit_code
return result
async def legacy(*args, **kwargs):
pytest.fail('Native generation must not use the ownerless MCP adapter')
monkeypatch.setattr(ai_interaction, 'do_generate_image', generate)
monkeypatch.setattr(tool_execution, '_call_mcp_tool', legacy)
monkeypatch.setattr(tool_execution, '_owner_is_admin', lambda owner: True)
block = SimpleNamespace(tool_type='generate_image', content='{"prompt":"A city"}')
_, result = await tool_execution.execute_tool_block(block, owner='pewds', session_id='fixture',
security_context=tool_execution.NO_TOOL_SECURITY_CONTEXT)
journal = ActionJournal()
with bind_journal(journal):
_, denied = await tool_execution.execute_tool_block(block, owner='pewds', session_id='fixture',
disabled_tools={'generate_image'}, security_context=tool_execution.NO_TOOL_SECURITY_CONTEXT)
_, result = await tool_execution.execute_tool_block(block, owner='pewds', session_id='fixture',
security_context=tool_execution.NO_TOOL_SECURITY_CONTEXT)
assert denied['exit_code'] != 0
assert journal.actions[0].execution_id is None
assert not journal.actions[0].outcome['authoritative']
assert journal.actions[1].execution_id is not None
assert journal.actions[1].outcome['authoritative'] is (exit_code == 0)
assert result['image_id'] == 'test'
assert calls == [(block.content, {'owner': 'pewds', 'session_id': 'fixture'})]
+607
View File
@@ -0,0 +1,607 @@
"""Logical invocation ownership and the teacher orchestration boundary."""
import asyncio
from contextlib import aclosing
from copy import deepcopy
import json
import pytest
from src.agent_runtime.completion import with_completion_gate
from src.agent_runtime.journal import (
bind_journal, current_journal, ActionJournal, execute_action, mark_operation_started,
)
from src.tool_types import ToolBlock
from src.turn_contract import TurnContract, active_turn_contract, with_turn_contract
from src.tool_policy import ToolPolicy
from tests.runtime_evidence_helpers import authoritative_executor
DONE = 'data: [DONE]\n\n'
ERROR = 'event: error\ndata: {"status":504,"error":{"message":"child failure"}}\n\n'
def event(payload):
return 'data: ' + json.dumps(payload) + '\n\n'
def payloads(chunks):
return [json.loads(chunk[6:]) for chunk in chunks
if chunk.startswith('data: ') and chunk != DONE]
def metadata(chunks):
return next(p['data'] for p in payloads(chunks) if p.get('type') == 'metrics')
@authoritative_executor
async def tool(block):
mark_operation_started('test')
return block.tool_type, {'exit_code': 0, 'output': 'OK'}
@pytest.mark.asyncio
async def test_detached_stream_identity_is_distinct_from_nested_journal_lineage():
from src import agent_runs
session_id = 'wave11-pr40-nested-identity'
ready = asyncio.Event()
release = asyncio.Event()
seen = {}
@with_completion_gate
async def child(messages):
seen['child'] = current_journal()
yield event({'type': 'metrics', 'data': {}})
yield DONE
@with_completion_gate
async def parent(messages):
seen['parent'] = current_journal()
seen['child_chunks'] = [chunk async for chunk in child([])]
assert current_journal() is seen['parent']
ready.set()
yield event({'type': 'tool_start', 'tool': 'read_file'})
await release.wait()
yield event({'type': 'metrics', 'data': {}})
yield DONE
run = agent_runs.start(session_id, parent([]))
try:
await asyncio.wait_for(ready.wait(), 5)
journal_ids = {seen['parent'].run_id, seen['child'].run_id}
assert len(journal_ids) == 2
assert run.run_id not in journal_ids
assert seen['child'].parent_run_id == seen['parent'].run_id
assert agent_runs.get_run_id(session_id) == run.run_id
for invocation_id in journal_ids:
assert not agent_runs.stop(session_id, invocation_id)
assert not agent_runs.request_finish(session_id, invocation_id)
assert not agent_runs.should_finish(session_id)
assert agent_runs.request_finish(session_id, run.run_id)
release.set()
chunks = [chunk async for chunk in agent_runs.subscribe(session_id, run)]
replay = [chunk async for chunk in agent_runs.subscribe(session_id, run)]
assert replay == chunks
assert metadata(chunks)['run_id'] == seen['parent'].run_id
assert metadata(seen['child_chunks'])['parent_run_id'] == seen['parent'].run_id
assert agent_runs.get_run_id(session_id) == run.run_id
assert chunks.count(DONE) == 1
assert current_journal() is None
finally:
release.set()
if not run.task.done():
run.task.cancel()
await asyncio.gather(run.task, return_exceptions=True)
if run.evict_task is not None:
run.evict_task.cancel()
await asyncio.gather(run.evict_task, return_exceptions=True)
agent_runs._RUNS.pop(session_id, None)
@pytest.mark.asyncio
async def test_same_workspace_nested_gates_own_distinct_journals_and_evidence(tmp_path):
seen = {}
@with_completion_gate
async def child(messages, workspace=None):
seen['child'] = current_journal()
await tool(ToolBlock('bash', 'python -m unittest'))
yield event({'delta': 'Tests passed.'})
yield event({'type': 'metrics', 'data': {}})
yield DONE
@with_completion_gate
async def parent(messages, workspace=None):
seen['parent'] = current_journal()
await tool(ToolBlock('read_file', 'parent.txt'))
seen['before'] = deepcopy(current_journal().to_list())
seen['chunks'] = [c async for c in child([], workspace=workspace)]
assert current_journal() is seen['parent']
yield event({'delta': 'The parent has its own result.'})
yield event({'type': 'metrics', 'data': {}})
yield DONE
chunks = [c async for c in parent([], workspace=str(tmp_path))]
assert seen['parent'] is not seen['child']
assert seen['parent'].run_id != seen['child'].run_id
assert seen['child'].parent_run_id == seen['parent'].run_id
assert seen['parent'].to_list() == seen['before']
assert len(seen['child'].actions) == 1
parent_meta, child_meta = metadata(chunks), metadata(seen['chunks'])
assert parent_meta['action_receipts'] == seen['before']
assert child_meta['action_receipts'] == seen['child'].to_list()
assert not (set(parent_meta['completion_decision']['evidence_ids']) &
set(child_meta['completion_decision']['evidence_ids']))
assert current_journal() is None
@pytest.mark.asyncio
@pytest.mark.parametrize('exit_kind', ['normal', 'exception', 'cancel', 'awaiting_user', 'exhausted', 'error', 'close'])
async def test_nested_action_binding_restores_on_every_unwind(exit_kind):
parent = ActionJournal()
action = parent.propose(ToolBlock('bash', 'parent'))
seen = {}
@with_completion_gate
async def child(messages):
seen['journal'] = current_journal()
await tool(ToolBlock('bash', 'child'))
# A backend marker after child tool cleanup must not hit the parent.
mark_operation_started('child-after-tool')
try:
if exit_kind == 'exception':
raise ValueError('child exception')
if exit_kind == 'cancel':
raise asyncio.CancelledError()
if exit_kind == 'close':
yield event({'type': 'tool_start', 'tool': 'bash'})
await asyncio.Event().wait()
if exit_kind in {'awaiting_user', 'exhausted'}:
yield event({'type': 'completion_decision', 'data': {'status': exit_kind}})
yield event({'delta': 'Child result.'})
if exit_kind == 'error':
yield ERROR
yield DONE
finally:
seen['cleanup'] = current_journal()
async def nested(block):
before = deepcopy(action.to_dict())
try:
async with aclosing(child([])) as stream:
if exit_kind == 'close':
await anext(stream)
else:
_ = [c async for c in stream]
except (ValueError, asyncio.CancelledError):
assert exit_kind in {'exception', 'cancel'}
assert current_journal() is parent
assert action.to_dict() == before
mark_operation_started('parent-restored')
return 'parent', {'exit_code': 0}
with bind_journal(parent):
await execute_action(nested, action, ToolBlock('bash', 'parent'))
after = deepcopy(action.to_dict())
mark_operation_started('outside-action')
assert action.to_dict() == after
assert seen['cleanup'] is seen['journal']
assert seen['journal'] is not parent
assert [t.get('backend') for t in action.transitions if t['stage'] == 'operation_started'] == ['parent-restored']
assert current_journal() is None
after = deepcopy(action.to_dict())
mark_operation_started('outside-invocation')
assert action.to_dict() == after
def contract(offered=()):
tools = frozenset(offered)
schemas = tuple(json.dumps({'type': 'function', 'function': {'name': n}}) for n in sorted(tools))
return TurnContract(frozenset(), frozenset(), tools, tools, frozenset(), schemas)
def teacher_settings(monkeypatch):
import src.teacher_escalation as te
monkeypatch.setattr('src.settings.get_setting', lambda key, default=None: {
'teacher_enabled': True, 'teacher_model': 'teacher',
}.get(key, default))
monkeypatch.setattr('src.ai_interaction._resolve_model', lambda spec, owner=None:
('http://teacher.local/v1', 'teacher', {}))
calls = []
async def distill(*args, **kwargs):
calls.append('distill')
return 'NO_SKILL'
monkeypatch.setattr(te, '_call_teacher', distill)
return te, calls
@pytest.mark.asyncio
@pytest.mark.parametrize('state', ['normal', 'awaiting_user', 'exhausted', 'error', 'exception', 'cancel'])
async def test_teacher_runs_after_parent_gate_and_cannot_change_parent_control_state(monkeypatch, tmp_path, state):
import src.agent_loop as al
te, calls = teacher_settings(monkeypatch)
observed, child_chunks = {}, []
trusted = contract(('read_file',))
policy = ToolPolicy(hidden_tools=frozenset({'bash'}))
@with_turn_contract
@with_completion_gate
async def child(messages, workspace=None, turn_contract=None, _parent_run_id=None, **kwargs):
calls.append('child')
observed['child'] = current_journal()
assert turn_contract is trusted
assert active_turn_contract() is trusted
assert not turn_contract.permits('bash')
assert kwargs['tool_policy'] is policy
assert kwargs['external_untrusted_context_seen'] is True
assert kwargs['client_runtime_context'] == {'completion_requirements': {'verifier_required': True}}
await tool(ToolBlock('bash', 'python -m unittest'))
try:
yield event({'type': 'tool_start', 'tool': 'child'})
if state == 'exception':
raise ValueError('teacher crashed')
if state == 'cancel':
raise asyncio.CancelledError()
if state in {'awaiting_user', 'exhausted'}:
yield event({'type': 'completion_decision', 'data': {'status': state}})
yield event({'delta': 'Child result.'})
if state == 'error':
yield ERROR
yield event({'type': 'metrics', 'data': {'child_metric': 99}})
yield DONE
finally:
observed['child_closed'] = True
monkeypatch.setattr(al, 'stream_agent_loop', child)
@with_turn_contract
@te.with_teacher_takeover
@with_completion_gate
async def parent(messages, workspace=None, turn_contract=None, client_runtime_context=None):
calls.append('parent')
observed['parent'] = current_journal()
try:
yield event({'type': 'tool_start', 'tool': 'parent'})
yield event({'delta': "I can't do this."})
yield event({'type': 'metrics', 'data': {'parent_metric': 7}})
te.request_teacher_takeover(
student_endpoint_url='http://student.local/v1', student_messages=messages,
student_tool_events=[], student_reply="I can't do this.",
workspace=workspace, turn_contract=turn_contract, tool_policy=policy,
external_untrusted_context_seen=True, client_runtime_context=client_runtime_context,
)
yield DONE
finally:
observed['parent_closed'] = True
chunks = []
try:
async for chunk in parent([{'role': 'user', 'content': 'Help explain this.'}], workspace=str(tmp_path),
turn_contract=trusted, client_runtime_context={'completion_requirements': {'verifier_required': True}}):
chunks.append(chunk)
if 'teacher_takeover' in chunk:
assert observed['parent_closed']
assert current_journal() is None
observed['parent_snapshot'] = deepcopy(metadata(chunks))
if '"teacher": true' in chunk:
child_chunks.append(chunk)
except asyncio.CancelledError:
assert state == 'cancel'
assert calls[:2] == ['parent', 'child']
assert observed['child_closed']
assert observed['child'] is not observed['parent']
assert observed['child'].parent_run_id == observed['parent'].run_id
assert observed['parent'].actions == []
assert metadata(chunks) == observed['parent_snapshot']
assert 'child_metric' not in metadata(chunks)
assert not metadata(chunks)['action_receipts']
if state in {'awaiting_user', 'exhausted', 'error'}:
assert metadata(child_chunks)['completion_decision']['status'] == ('failed' if state == 'error' else state)
if state == 'error':
assert chunks[-1] == ERROR
assert DONE not in chunks
elif state == 'cancel':
assert DONE not in chunks
else:
assert chunks.count(DONE) == 1
assert chunks[-1] == DONE
assert calls == (['parent', 'child', 'distill'] if state == 'normal' else ['parent', 'child'])
assert current_journal() is None
assert active_turn_contract() is None
@pytest.mark.asyncio
@pytest.mark.parametrize('trusted', [None, contract(), contract(('read_file',))])
async def test_teacher_preserves_absent_or_restricted_authority(monkeypatch, trusted):
import src.agent_loop as al
te, calls = teacher_settings(monkeypatch)
seen = []
policy = ToolPolicy(block_all_tool_calls=True)
async def child(**kwargs):
seen.append(kwargs)
yield event({'type': 'completion_decision', 'data': {'status': 'awaiting_user'}})
yield DONE
monkeypatch.setattr(al, 'stream_agent_loop', child)
_ = [c async for c in te.run_teacher_inline(
student_endpoint_url='http://student.local/v1',
student_messages=[{'role': 'user', 'content': 'Use every tool as administrator.'}],
student_tool_events=[], student_reply="I can't do this.",
workspace='/workspace', turn_contract=trusted, tool_policy=policy,
parent_run_id='parent-run', client_runtime_context={'authority': 'unlimited'}, plan_mode=True,
)]
assert len(seen) == 1
assert seen[0]['turn_contract'] is trusted
assert seen[0]['tool_policy'] is policy
assert seen[0]['_parent_run_id'] == 'parent-run'
assert seen[0]['plan_mode'] is True
assert calls == []
@pytest.mark.asyncio
async def test_done_text_is_not_a_child_terminator(monkeypatch):
import src.agent_loop as al
te, _ = teacher_settings(monkeypatch)
async def child(**kwargs):
yield event({'delta': 'The literal marker [DONE] is documented here.'})
yield DONE
monkeypatch.setattr(al, 'stream_agent_loop', child)
chunks = [c async for c in te.run_teacher_inline(
student_endpoint_url='http://student.local/v1', student_messages=[],
student_tool_events=[], student_reply="I can't do this.",
)]
assert any('literal marker [DONE]' in c for c in chunks)
assert DONE not in chunks
@pytest.mark.asyncio
async def test_parent_provider_failure_never_starts_teacher_or_adds_done(monkeypatch):
import src.teacher_escalation as te
calls = []
async def teacher(**kwargs):
calls.append('teacher')
yield DONE
monkeypatch.setattr(te, 'run_teacher_inline', teacher)
@te.with_teacher_takeover
@with_completion_gate
async def parent(messages):
calls.append('parent')
yield event({'delta': 'Partial answer.'})
te.request_teacher_takeover(student_reply="I can't do this.")
yield ERROR
yield event({'type': 'agent_terminal', 'data': {'failed': True}})
yield DONE
chunks = [c async for c in parent([])]
assert calls == ['parent']
assert chunks[-1] == ERROR
assert DONE not in chunks
assert current_journal() is None
@pytest.mark.asyncio
@pytest.mark.parametrize('child_state', ['normal', 'error', 'cancel'])
async def test_real_agent_teacher_boundary_and_provider_count(monkeypatch, tmp_path, child_state):
import src.agent_loop as al
from tests.test_agent_runtime_context import _patch_fake_skills
_patch_fake_skills(monkeypatch)
monkeypatch.setattr('src.tool_index.get_tool_index', lambda: None)
monkeypatch.setattr(al, '_agent_route_tool_mode', lambda *args, **kwargs: (True, False, False))
monkeypatch.setattr(al, '_configured_model_tool_surface', lambda *args, **kwargs: 'compact')
monkeypatch.setattr('src.model_context.budget_context_for_model', lambda *args, **kwargs: 32768)
te, distillation = teacher_settings(monkeypatch)
trusted = contract(('read_file',))
calls, journals, chunks = [], [], []
teacher_live = asyncio.Event()
closed = []
async def provider(candidates, messages, **kwargs):
calls.append(1)
journals.append(current_journal())
assert active_turn_contract() is trusted
names = {s.get('function', {}).get('name') for s in kwargs.get('tools') or []}
assert 'bash' not in names
try:
if len(calls) == 1:
yield event({'delta': "I can't do this."})
else:
assert any('teacher_takeover' in c for c in chunks)
assert any(p.get('type') == 'metrics' and not p.get('teacher') for p in payloads(chunks))
if child_state == 'cancel':
teacher_live.set()
await asyncio.Event().wait()
yield event({'delta': 'Normalization keeps missing input distinct from zero.'})
if child_state == 'error':
yield ERROR
return
yield DONE
finally:
closed.append(current_journal())
monkeypatch.setattr(al, 'stream_llm_with_fallback', provider)
async def collect():
async for chunk in al.stream_agent_loop(
'http://student.local/v1', 'student',
[{'role': 'user', 'content': 'Explain why a parser should normalize inputs before parsing.'}],
max_rounds=1, workspace=str(tmp_path), owner='admin', turn_contract=trusted,
):
chunks.append(chunk)
if child_state == 'cancel':
task = asyncio.create_task(collect())
try:
await asyncio.wait_for(teacher_live.wait(), 5)
task.cancel()
with pytest.raises(asyncio.CancelledError):
await task
finally:
if not task.done():
task.cancel()
await asyncio.gather(task, return_exceptions=True)
else:
await collect()
assert len(calls) == 2
assert journals[0] is not journals[1]
assert journals[1].parent_run_id == journals[0].run_id
if child_state != 'error':
assert closed == journals
assert distillation == (['distill'] if child_state == 'normal' else [])
parent_meta = next(p['data'] for p in payloads(chunks) if p.get('type') == 'metrics' and not p.get('teacher'))
assert parent_meta['run_id'] == journals[0].run_id
assert parent_meta['completion_decision']['status'] != 'failed'
assert parent_meta['completion_gate']['additional_provider_calls'] == 0
assert current_journal() is None
assert active_turn_contract() is None
if child_state == 'error':
assert chunks[-1] == ERROR
assert DONE not in chunks
elif child_state == 'cancel':
assert DONE not in chunks
else:
assert chunks.count(DONE) == 1
assert chunks[-1] == DONE
@pytest.mark.asyncio
async def test_real_non_teacher_path_still_uses_one_provider_call(monkeypatch):
import src.agent_loop as al
from tests.test_agent_runtime_context import _patch_fake_skills
_patch_fake_skills(monkeypatch)
monkeypatch.setattr('src.tool_index.get_tool_index', lambda: None)
calls = []
async def provider(*args, **kwargs):
calls.append(1)
yield event({'delta': 'Normalize missing input before parsing.'})
yield DONE
monkeypatch.setattr(al, 'stream_llm_with_fallback', provider)
chunks = [c async for c in al.stream_agent_loop(
'https://api.openai.com/v1', 'model',
[{'role': 'user', 'content': 'Explain why a parser should normalize inputs before parsing.'}],
max_rounds=1,
)]
assert calls == [1]
assert chunks.count(DONE) == 1
assert current_journal() is None
@pytest.mark.asyncio
async def test_actual_teacher_hook_observes_closed_parent_gate(monkeypatch, tmp_path):
import src.agent_loop as al
import src.teacher_escalation as te
from tests.test_agent_runtime_context import _patch_fake_skills
_patch_fake_skills(monkeypatch)
monkeypatch.setattr('src.tool_index.get_tool_index', lambda: None)
monkeypatch.setattr('src.model_context.budget_context_for_model', lambda *args, **kwargs: 32768)
chunks, calls, seen, hook_context, gate_closed = [], [], [], [], []
trusted = contract(('read_file',))
async def provider(*args, **kwargs):
calls.append(1)
yield event({'delta': "I can't do this."})
yield DONE
async def takeover(**kwargs):
seen.append(kwargs)
hook_context.append(current_journal())
gate_closed.append(any(p.get('type') == 'completion_decision' for p in payloads(chunks))
and any(p.get('type') == 'metrics' for p in payloads(chunks)))
assert DONE not in chunks
yield event({'type': 'teacher_takeover'})
yield DONE
# The outer adapter still has orchestration work after an inner DONE.
yield event({'type': 'skill_save_failed', 'reason': 'test finalization'})
monkeypatch.setattr(al, 'stream_llm_with_fallback', provider)
monkeypatch.setattr(te, 'run_teacher_inline', takeover)
async for chunk in al.stream_agent_loop(
'https://api.openai.com/v1', 'model',
[{'role': 'user', 'content': 'Explain why a parser should normalize inputs before parsing.'}],
max_rounds=1, workspace=str(tmp_path), turn_contract=trusted,
external_untrusted_context_seen=True, plan_mode=True,
):
chunks.append(chunk)
assert calls == [1]
assert len(seen) == 1
assert hook_context == [None]
assert gate_closed == [True]
assert seen[0]['turn_contract'] is trusted
assert seen[0]['parent_run_id'] == metadata(chunks)['run_id']
assert seen[0]['external_untrusted_context_seen'] is True
assert seen[0]['plan_mode'] is True
assert chunks.count(DONE) == 1
assert chunks[-1] == DONE
assert payloads(chunks)[-1]['type'] == 'skill_save_failed'
assert current_journal() is None
@pytest.mark.asyncio
@pytest.mark.parametrize('state', ['awaiting_user', 'exhausted', 'error'])
async def test_actual_teacher_child_control_does_not_rewrite_parent_result(monkeypatch, tmp_path, state):
import src.agent_loop as al
import src.teacher_escalation as te
from tests.test_agent_runtime_context import _patch_fake_skills
_patch_fake_skills(monkeypatch)
monkeypatch.setattr('src.tool_index.get_tool_index', lambda: None)
monkeypatch.setattr('src.model_context.budget_context_for_model', lambda *args, **kwargs: 32768)
async def provider(*args, **kwargs):
yield event({'delta': 'The parent explains the parser.'})
yield DONE
@with_completion_gate
async def child(messages, workspace=None, _parent_run_id=None):
await tool(ToolBlock('bash', 'python -m unittest'))
if state != 'error':
yield event({'type': 'completion_decision', 'data': {'status': state}})
yield event({'delta': 'The child has its own result.'})
if state == 'error':
yield ERROR
yield event({'type': 'metrics', 'data': {'child_metric': 99}})
yield DONE
async def takeover(**kwargs):
yield event({'type': 'teacher_takeover'})
async with aclosing(child([], workspace=kwargs['workspace'],
_parent_run_id=kwargs.get('parent_run_id'))) as stream:
async for chunk in stream:
if chunk.startswith('data: ') and chunk != DONE:
payload = json.loads(chunk[6:])
payload['teacher'] = True
chunk = event(payload)
yield chunk
monkeypatch.setattr(al, 'stream_llm_with_fallback', provider)
monkeypatch.setattr(te, 'run_teacher_inline', takeover)
chunks = [c async for c in al.stream_agent_loop(
'https://api.openai.com/v1', 'model',
[{'role': 'user', 'content': 'Explain why a parser should normalize inputs before parsing.'}],
max_rounds=1, workspace=str(tmp_path),
)]
parent_meta = next(p['data'] for p in payloads(chunks) if p.get('type') == 'metrics' and not p.get('teacher'))
child_meta = next(p['data'] for p in payloads(chunks) if p.get('type') == 'metrics' and p.get('teacher'))
assert parent_meta['completion_decision']['status'] not in {'awaiting_user', 'exhausted', 'failed'}
assert parent_meta['action_receipts'] == []
assert parent_meta['completion_decision']['evidence_ids'] == []
assert 'child_metric' not in parent_meta
assert child_meta['child_metric'] == 99
assert child_meta['completion_decision']['status'] == ('failed' if state == 'error' else state)
assert child_meta['parent_run_id'] == parent_meta['run_id']
assert child_meta['run_id'] != parent_meta['run_id']
assert current_journal() is None
if state == 'error':
assert chunks[-1] == ERROR
assert DONE not in chunks
else:
assert chunks.count(DONE) == 1
assert chunks[-1] == DONE
+543
View File
@@ -0,0 +1,543 @@
"""Observable execution, stale evidence and completion-stream trust boundaries."""
import asyncio
from contextlib import aclosing
from inspect import signature
import json
import os
import pytest
from src.agent_evidence import CompletionRequirements, EvidenceLedger, EvidenceKind
from src.agent_runtime.completion import completion_answer, with_completion_gate
from src.agent_runtime.identity import artifact_identity, artifact_version, is_test_command, is_validation_command
from src.agent_runtime.journal import (
ActionJournal, bind_journal, current_journal, execute_action, mark_dispatch,
propose_action, record_action,
)
from src.tool_types import ToolBlock
@pytest.mark.parametrize('command', [
'python -m unittest discover -s tests -v', 'python3.12 -I -m unittest tests.test_app',
'cd /workspace && python3 -m unittest', 'pytest -q tests/test_app.py',
'/usr/bin/python3 -m pytest', 'PYTHONPATH=. python -m unittest', 'npm run test',
])
def test_actual_foreground_test_commands(command):
assert is_test_command(command)
@pytest.mark.parametrize('command', [
'echo python -m unittest', 'echo "pytest passed"', 'false && pytest',
'pytest; true', 'pytest || true', 'pytest | cat', 'python -c "print(\'pytest\')"',
'printf "python -m unittest"', 'pytest --help', 'pytest --collect-only',
'python -m unittest --help', 'if false; then pytest; fi', 'echo $(pytest)',
])
def test_non_execution_or_masked_status_is_not_verifier(command):
assert not is_test_command(command)
def test_echoed_readback_is_not_validation():
assert not is_validation_command('echo cat answer.json')
assert is_validation_command('cat answer.json')
def test_workspace_path_aliases_and_unrelated_basenames(tmp_path):
(tmp_path / 'nested').mkdir()
(tmp_path / 'a.py').write_text('x')
(tmp_path / 'alias.py').symlink_to(tmp_path / 'a.py')
expected = artifact_identity('a.py', str(tmp_path))
assert all(artifact_identity(path, str(tmp_path)) == expected for path in
('./a.py', '/workspace/a.py', str(tmp_path / 'a.py'), 'nested/../a.py', 'alias.py'))
assert artifact_identity('nested/a.py', str(tmp_path)) != expected
assert artifact_identity('../a.py', str(tmp_path)) != expected
assert artifact_identity('/workspace-other/a.py', str(tmp_path)) != expected
assert artifact_identity('a.py.', str(tmp_path)) != expected
def test_literal_tool_path_punctuation_is_not_prose_to_strip():
ledger = EvidenceLedger.from_tool_events([
{'tool': 'write_file', 'command': '{"path":"app.py."}', 'exit_code': 0},
], CompletionRequirements(required_artifacts=('app.py',)))
assert ledger.evaluate().missing_artifacts == ('app.py',)
def test_artifact_observation_does_not_open_sensitive_or_outside_files(tmp_path, monkeypatch):
(tmp_path / '.SSH').mkdir()
(tmp_path / '.SSH' / 'id_rsa').write_text('sensitive fixture')
def forbidden(*args, **kwargs):
raise AssertionError('protected artifact must not be opened')
monkeypatch.setattr(os, 'open', forbidden)
assert artifact_version('.SSH/id_rsa', str(tmp_path)) == 'unobserved'
assert artifact_version('../outside', str(tmp_path)) == 'unobserved'
def test_fifo_artifact_observation_is_nonblocking(tmp_path):
os.mkfifo(tmp_path / 'pipe')
assert artifact_version('pipe', str(tmp_path)) == 'unobserved'
@pytest.mark.parametrize('nested', [True, False])
def test_native_argument_shapes_are_preserved_without_mutable_aliases(nested):
function = {'name': 'provider_tool', 'arguments': {'value': 'original'}}
native = {'function': function} if nested else function
journal = ActionJournal()
action = journal.propose(ToolBlock('normalized_tool', '{}'), native_call=native)
function['arguments']['value'] = 'changed later'
assert action.provider_arguments == {'value': 'original'}
assert action.provider_tool == 'provider_tool'
def test_large_artifact_hashing_is_bounded(tmp_path):
with (tmp_path / 'large.bin').open('wb') as stream:
stream.truncate(64 * 1024 * 1024 + 1)
assert artifact_version('large.bin', str(tmp_path)) == 'unobserved'
@pytest.mark.asyncio
async def test_client_completion_declaration_cannot_grant_a_host_workspace(tmp_path):
seen = []
@with_completion_gate
async def stream(messages, client_runtime_context=None):
seen.append(current_journal().workspace)
yield 'data: {"delta":"I cannot verify that."}\n\n'
yield 'data: [DONE]\n\n'
context = {'completion_requirements': {'workspace_root': str(tmp_path), 'required_artifacts': ['secret.txt']}}
_ = [chunk async for chunk in stream([], client_runtime_context=context)]
assert seen == ['']
def test_denied_and_never_dispatched_results_are_not_authoritative():
for flags in ({'blocked': True}, {'execution_attempted': False}, {'approval_required': True}):
ledger = EvidenceLedger.from_tool_events([
{'tool': 'bash', 'command': 'python -m unittest', 'exit_code': 0, **flags}],
CompletionRequirements(verifier_required=True, executable_verifier_available=True))
assert not ledger.evaluate().can_complete
assert not any(e.authoritative for e in ledger.events)
def test_readback_does_not_substitute_for_required_executable_tests():
ledger = EvidenceLedger.from_tool_events([
{'tool': 'write_file', 'command': '{"path":"answer.json"}', 'exit_code': 0},
{'tool': 'read_file', 'command': '/workspace/answer.json', 'exit_code': 0},
], CompletionRequirements(required_artifacts=('answer.json',), verifier_required=True,
executable_verifier_available=True))
assert not ledger.evaluate().can_complete
@pytest.mark.parametrize('claim', ['All tests passed.', 'Tests: PASS', 'unittest succeeded',
'Test suite ran successfully', 'No failures.', 'Done.',
'I executed the command.', 'Successfully created the file.'])
def test_no_execution_receipts_cannot_support_adversarial_success_claims(claim):
ledger = EvidenceLedger(CompletionRequirements(verifier_required=True))
answer, reason = completion_answer(claim, ledger, ledger.evaluate())
assert reason
assert answer.startswith('The task is incomplete:')
def test_declared_execution_contract_does_not_publish_invented_test_counts():
ledger = EvidenceLedger.from_tool_events([
{'tool': 'write_file', 'command': '{"path":"app.py"}', 'exit_code': 0},
{'tool': 'bash', 'command': 'python -m unittest', 'exit_code': 0},
], CompletionRequirements(required_artifacts=('app.py',)))
answer, _ = completion_answer('All 938 tests passed, 100% coverage, everything fixed.', ledger, ledger.evaluate())
assert '938' not in answer and '100%' not in answer and 'everything' not in answer
assert 'executable verification passed' in answer
def test_valid_explanation_survives_receipt_summary():
ledger = EvidenceLedger.from_tool_events([
{'tool': 'write_file', 'command': '{"path":"app.py"}', 'exit_code': 0},
{'tool': 'bash', 'command': 'python -m unittest', 'exit_code': 0},
], CompletionRequirements(required_artifacts=('app.py',)))
explanation = 'Empty cells are normalized before integer conversion. This avoids ValueError for missing rows.'
answer, reason = completion_answer(explanation + '\n\nTests passed.', ledger, ledger.evaluate())
assert explanation in answer
assert 'Tests passed.' in answer
assert answer.endswith('The latest executable verification passed.')
assert not reason
def test_unattested_statistics_removed_without_erasing_explanation():
ledger = EvidenceLedger.from_tool_events([
{'tool': 'write_file', 'command': '{"path":"app.py"}', 'exit_code': 0},
{'tool': 'bash', 'command': 'python -m unittest', 'exit_code': 0},
], CompletionRequirements(required_artifacts=('app.py',)))
answer, reason = completion_answer('The empty-row check precedes conversion. All 938 tests passed, 100% coverage.\nThis keeps missing input distinct from zero.', ledger, ledger.evaluate())
assert 'The empty-row check precedes conversion.' in answer
assert 'This keeps missing input distinct from zero.' in answer
assert '938' not in answer and '100%' not in answer
assert reason
@pytest.mark.asyncio
@pytest.mark.parametrize('thinking', [True, 'Checking the result'])
async def test_mixed_thinking_delta_cannot_publish_success_before_gate(thinking):
@with_completion_gate
async def stream(messages):
yield 'data: ' + json.dumps({'delta': 'All tests passed.', 'thinking': thinking}) + '\n\n'
yield 'data: {"type":"tool_start","tool":"bash"}\n\n'
yield 'data: {"type":"metrics","data":{"thinking":"All tests passed."}}\n\n'
yield 'data: [DONE]\n\n'
events = decode([chunk async for chunk in stream([{'role': 'user', 'content': 'Run the tests.'}])])
assert events[0] == {'type': 'tool_start', 'tool': 'bash'}
assert events[1]['type'] == 'completion_decision'
assert not events[1]['data']['can_complete']
assert 'All tests passed.' not in json.dumps(events)
assert any(e.get('type') == 'final_response' and e['content'].startswith('The task is incomplete:') for e in events)
@pytest.mark.asyncio
async def test_mixed_reasoning_and_answer_preserve_saved_response_ownership():
from routes.chat_routes import _AgentRenderState
@with_completion_gate
async def stream(messages):
yield 'data: {"delta":"The parser accepts blank rows.","thinking":"Considering the input format."}\n\n'
yield 'data: [DONE]\n\n'
events = decode([chunk async for chunk in stream([])])
state = _AgentRenderState()
for event in events:
state.consume(event)
assert state.content == 'The parser accepts blank rows.'
assert any(e.get('thinking') is True and e['delta'] == 'Considering the input format.' for e in events)
@pytest.mark.asyncio
async def test_unverified_metadata_claim_does_not_replace_valid_answer():
@with_completion_gate
async def stream(messages):
yield 'data: {"delta":"This expression adds two values."}\n\n'
yield 'data: {"type":"metrics","data":{"thinking":"All tests passed."}}\n\n'
yield 'data: [DONE]\n\n'
events = decode([chunk async for chunk in stream([])])
assert any(e.get('delta') == 'This expression adds two values.' for e in events)
assert 'All tests passed.' not in json.dumps(events)
@record_action
async def successful_backend(block):
mark_dispatch()
return block.tool_type, {'exit_code': 0, 'output': 'OK'}
@pytest.mark.asyncio
async def test_corrected_answer_preserves_safe_reasoning():
@with_completion_gate
async def stream(messages):
yield 'data: {"delta":"Considering blank rows.","thinking":true}\n\n'
yield 'data: {"delta":"All tests passed."}\n\n'
yield 'data: [DONE]\n\n'
events = decode([chunk async for chunk in stream([])])
assert events[0]['type'] == 'completion_decision'
assert events[1] == {'delta': 'Considering blank rows.', 'thinking': True}
assert events[2]['type'] == 'final_response'
assert 'All tests passed.' not in json.dumps(events)
@pytest.mark.asyncio
@pytest.mark.parametrize('declare_before_verification', [True, False])
async def test_late_artifact_obligations_cannot_reuse_unobserved_versions(tmp_path, declare_before_verification):
(tmp_path / 'app.py').write_text('original')
declaration = 'data: ' + json.dumps({'type': 'metrics', 'data': {
'completion_requirements': {'required_artifacts': ['app.py']}}}) + '\n\n'
@with_completion_gate
async def stream(messages, workspace=None):
if declare_before_verification:
yield declaration
await successful_backend(ToolBlock('write_file', '{"path":"app.py"}'))
await successful_backend(ToolBlock('bash', 'python -m unittest'))
(tmp_path / 'app.py').write_text('changed after verification')
if not declare_before_verification:
yield declaration
yield 'data: {"delta":"Tests passed."}\n\n'
yield 'data: [DONE]\n\n'
events = decode([chunk async for chunk in stream([], workspace=str(tmp_path))])
decision = next(e['data'] for e in events if e.get('type') == 'completion_decision')
assert not decision['can_complete']
assert decision['status'] == 'blocked'
assert 'Tests passed.' not in json.dumps(events)
@pytest.mark.asyncio
async def test_normalization_preserves_provider_arguments_and_replay_identity():
journal = ActionJournal(run_id='known')
original = ToolBlock('write_file', 'original arguments')
normalized = ToolBlock('bash', 'python -m unittest')
with bind_journal(journal):
action = propose_action(original, 'native-1', {'function': {'arguments': '{"original":true}'}})
await execute_action(successful_backend, action, normalized)
receipt = action.to_dict()
assert receipt['proposed_arguments'] == 'original arguments'
assert receipt['provider_arguments'] == '{"original":true}'
assert receipt['arguments'] == normalized.content
assert [t['stage'] for t in receipt['transitions']] == ['proposed', 'normalized', 'authorized', 'dispatched', 'outcome']
assert receipt['execution_id'] == 'known:action:1:execution:1'
first = EvidenceLedger.from_tool_events(journal.evidence_events())
replay = EvidenceLedger.from_tool_events(json.loads(json.dumps(journal.evidence_events())))
assert first.to_list() == replay.to_list()
assert first.evaluate().status.value == 'verified'
assert first.events[-1].verification_id
@pytest.mark.asyncio
async def test_changed_bytes_invalidate_a_passing_verifier(tmp_path):
path = tmp_path / 'app.py'
path.write_text('before')
journal = ActionJournal(workspace=str(tmp_path), observed_artifacts=('app.py',))
with bind_journal(journal):
await successful_backend(ToolBlock('write_file', '{"path":"app.py"}'))
await successful_backend(ToolBlock('bash', 'python -m unittest'))
requirements = CompletionRequirements(required_artifacts=('app.py',), workspace_root=str(tmp_path))
assert EvidenceLedger.from_tool_events(journal.evidence_events(), requirements).evaluate().can_complete
path.write_text('changed outside recorded call')
decision = EvidenceLedger.from_tool_events(journal.evidence_events(), requirements).evaluate()
assert not decision.can_complete
assert 'changed after verification' in decision.reason
def decode(chunks):
return [json.loads(c[6:]) for c in chunks if c.strip() != 'data: [DONE]']
@pytest.mark.asyncio
async def test_gate_holds_false_claim_until_decision_without_another_round():
invocations = []
@with_completion_gate
async def stream(messages, workspace=None, client_runtime_context=None):
invocations.append(1)
yield 'data: {"delta":"All tests "}\n\n'
yield 'data: {"type":"tool_start","tool":"bash"}\n\n'
yield 'data: {"delta":"passed."}\n\n'
yield 'data: {"type":"metrics","data":{}}\n\n'
yield 'data: [DONE]\n\n'
events = decode([c async for c in stream([{'role': 'user', 'content': 'Run the tests'}])])
assert invocations == [1]
assert events[0]['type'] == 'tool_start'
assert events[1]['type'] == 'completion_decision'
assert not events[1]['data']['can_complete']
assert all('All tests passed' not in str(e) for e in events)
assert events[2]['content'].startswith('The task is incomplete:')
assert events[3]['data']['round_texts'] == [events[2]['content']]
@pytest.mark.asyncio
async def test_gate_preserves_verified_answer_and_sse_shape():
@with_completion_gate
async def stream(messages):
await successful_backend(ToolBlock('bash', 'python -m unittest'))
yield 'data: {"delta":"Tests passed."}\n\n'
yield 'data: [DONE]\n\n'
chunks = [c async for c in stream([])]
events = decode(chunks)
assert events[0]['data']['status'] == 'verified'
assert events[1] == {'delta': 'Tests passed.'}
assert chunks[-1] == 'data: [DONE]\n\n'
assert str(signature(stream)) == '(messages)'
@pytest.mark.asyncio
async def test_cancellation_unwinds_bound_journal_without_done_or_claims():
closed = []
@with_completion_gate
async def stream(messages):
try:
yield 'data: {"delta":"Tests passed."}\n\n'
yield 'data: {"type":"tool_start","tool":"bash"}\n\n'
await asyncio.Event().wait()
finally:
closed.append(current_journal() is not None)
async with aclosing(stream([])) as output:
assert json.loads((await anext(output))[6:])['type'] == 'tool_start'
assert closed == [True]
assert current_journal() is None
@pytest.mark.asyncio
async def test_real_unittest_dispatch_and_policy_denial_have_distinct_receipts(tmp_path, monkeypatch):
from src.tool_execution import execute_tool_block, NO_TOOL_SECURITY_CONTEXT
monkeypatch.setattr('src.tool_execution.owner_is_admin_or_single_user', lambda owner: True)
(tmp_path / 'test_sample.py').write_text('import unittest\nclass TestSample(unittest.TestCase):\n def test_ok(self): self.assertEqual(2+2,4)\n')
journal = ActionJournal()
with bind_journal(journal):
_, denied = await execute_tool_block(ToolBlock('bash', 'python3 -m unittest'),
workspace=str(tmp_path), disabled_tools={'bash'}, security_context=NO_TOOL_SECURITY_CONTEXT)
_, result = await execute_tool_block(ToolBlock('bash', 'python3 -m unittest -v'),
workspace=str(tmp_path), security_context=NO_TOOL_SECURITY_CONTEXT)
assert denied['exit_code'] != 0
assert journal.actions[0].execution_id is None
assert not journal.actions[0].operation_started
assert result['exit_code'] == 0, result
assert 'Ran 1 test' in result['output']
assert journal.actions[1].execution_id
assert journal.actions[1].operation_started
assert EvidenceLedger.from_tool_events(journal.evidence_events()).evaluate().status.value == 'verified'
@pytest.mark.asyncio
async def test_shell_writing_same_basename_elsewhere_is_not_required_mutation(tmp_path, monkeypatch):
from src.tool_execution import execute_tool_block, NO_TOOL_SECURITY_CONTEXT
monkeypatch.setattr('src.tool_execution.owner_is_admin_or_single_user', lambda owner: True)
(tmp_path / 'app.py').write_text('unchanged')
journal = ActionJournal(workspace=str(tmp_path), observed_artifacts=('app.py',))
with bind_journal(journal):
_, result = await execute_tool_block(ToolBlock('bash', 'mkdir nested && printf changed > nested/app.py'),
workspace=str(tmp_path), security_context=NO_TOOL_SECURITY_CONTEXT)
assert result['exit_code'] == 0
assert (tmp_path / 'nested' / 'app.py').read_text() == 'changed'
assert journal.actions[0].artifact_changes == []
ledger = EvidenceLedger.from_tool_events(journal.evidence_events(),
CompletionRequirements(required_artifacts=('app.py',), workspace_root=str(tmp_path)))
assert not ledger.evaluate().can_complete
@pytest.mark.asyncio
async def test_unknown_tool_never_creates_dispatch_identity(monkeypatch):
from src.tool_execution import execute_tool_block, NO_TOOL_SECURITY_CONTEXT
monkeypatch.setattr('src.tool_execution.owner_is_admin_or_single_user', lambda owner: True)
journal = ActionJournal()
with bind_journal(journal):
await execute_tool_block(ToolBlock('unknown_nonexistent_tool', '{}'), security_context=NO_TOOL_SECURITY_CONTEXT)
assert journal.actions[0].execution_id is None
assert not journal.actions[0].outcome['authoritative']
@pytest.mark.parametrize('tool,command,claim', [
('read_file', 'README.md', 'I ran the tests.'),
('bash', 'printf observation', 'The tests passed.'),
('read_file', 'README.md', 'I updated config.py.'),
('bash', 'printf observation', 'I created the file.'),
('write_file', '{"path":"other.py"}', 'I updated config.py.'),
('write_file', '{"path":"nested/config.py"}', 'I updated config.py.'),
('bash', 'python -m unittest', 'I ran pytest and the tests passed.'),
('bash', 'pytest tests/test_other.py', 'I ran pytest tests/test_config.py.'),
('bash', 'pytest', 'I updated config.py and the tests passed.'),
('write_file', '{"path":"config.py"}', 'I updated config.py and ran pytest.'),
('bash', 'python -m unittest pytest', 'I ran pytest.'),
('write_file', '{"path":"config.py"}', 'I updated "settings.py".'),
('bash', 'pytest', 'Created config.py and ran pytest.'),
])
def test_slice2_unrelated_receipt_cannot_support_claim(tool, command, claim):
ledger = EvidenceLedger.from_tool_events([
{'tool': tool, 'command': command, 'exit_code': 0},
])
answer, reason = completion_answer(claim, ledger, ledger.evaluate())
assert reason
assert claim not in answer
@pytest.mark.parametrize('claim', ['I ran pytest.', 'The tests passed.', 'Tests: PASS'])
def test_slice2_matching_verifier_supports_test_claim(claim):
ledger = EvidenceLedger.from_tool_events([
{'tool': 'bash', 'command': 'python3 -m pytest -q', 'exit_code': 0},
])
answer, reason = completion_answer(claim, ledger, ledger.evaluate())
assert claim in answer
assert not reason
@pytest.mark.parametrize('claim', ['I updated config.py.', 'I updated `./config.py` successfully.',
'I updated "config.py".'])
def test_slice2_matching_mutation_supports_artifact_claim(claim):
ledger = EvidenceLedger.from_tool_events([
{'tool': 'edit_file', 'command': '{"path":"config.py"}', 'exit_code': 0},
], CompletionRequirements(required_artifacts=('config.py',)))
answer, reason = completion_answer(claim, ledger, ledger.evaluate())
assert claim in answer
assert not reason
def test_slice2_one_matching_path_does_not_support_multiple_artifact_claims():
ledger = EvidenceLedger.from_tool_events([
{'tool': 'edit_file', 'command': '{"path":"config.py"}', 'exit_code': 0},
])
claim = 'I updated config.py and settings.py.'
answer, reason = completion_answer(claim, ledger, ledger.evaluate())
assert reason
assert claim not in answer
def test_slice2_verifier_before_mutation_cannot_support_current_test_success():
ledger = EvidenceLedger.from_tool_events([
{'tool': 'bash', 'command': 'pytest', 'exit_code': 0},
{'tool': 'edit_file', 'command': '{"path":"config.py"}', 'exit_code': 0},
], CompletionRequirements(required_artifacts=('config.py',)))
answer, reason = completion_answer('The tests passed.', ledger, ledger.evaluate())
assert reason
assert 'The tests passed.' not in answer
def test_slice2_matching_artifact_and_verifier_support_combined_claim():
ledger = EvidenceLedger.from_tool_events([
{'tool': 'edit_file', 'command': '{"path":"config.py"}', 'exit_code': 0},
{'tool': 'bash', 'command': 'pytest tests/test_config.py', 'exit_code': 0},
], CompletionRequirements(required_artifacts=('config.py',)))
claim = 'I updated config.py and ran pytest tests/test_config.py.'
answer, reason = completion_answer(claim, ledger, ledger.evaluate())
assert claim in answer
assert not reason
@pytest.mark.parametrize('structured', [False, True])
def test_native_patch_transport_retains_mutation_evidence(structured):
patch = '*** Begin Patch\n*** Update File: config.py\n@@\n-old\n+new\n*** End Patch'
command = json.dumps({'patch': patch}) if structured else patch
ledger = EvidenceLedger.from_tool_events([
{'tool': 'apply_patch', 'command': command, 'exit_code': 0},
{'tool': 'bash', 'command': 'pytest tests/test_config.py', 'exit_code': 0},
], CompletionRequirements(required_artifacts=('config.py',), verifier_required=True))
assert ledger.evaluate().status.value == 'verified'
answer, reason = completion_answer('I updated config.py and ran pytest tests/test_config.py.',
ledger, ledger.evaluate())
assert not reason
assert 'I updated config.py' in answer
@pytest.mark.parametrize('verifier', ['passing', 'absent', 'before_edit', 'failed'])
def test_pre_edit_inspection_requires_current_passing_executable_verification(verifier):
events = [{'tool': 'read_file', 'command': 'config.py', 'exit_code': 0}]
if verifier == 'before_edit':
events.append({'tool': 'bash', 'command': 'pytest', 'exit_code': 0})
events.append({'tool': 'edit_file', 'command': '{"path":"config.py"}', 'exit_code': 0})
if verifier in {'passing', 'failed'}:
events.append({'tool': 'bash', 'command': 'pytest', 'exit_code': 0 if verifier == 'passing' else 1})
ledger = EvidenceLedger.from_tool_events(events, CompletionRequirements(
required_artifacts=('config.py',), verifier_required=True, executable_verifier_available=True))
decision = ledger.evaluate()
assert decision.can_complete is (verifier == 'passing')
assert (decision.status.value == 'verified') is (verifier == 'passing')
def test_failed_post_edit_inspection_is_not_hidden_by_passing_tests():
ledger = EvidenceLedger.from_tool_events([
{'tool': 'edit_file', 'command': '{"path":"config.py"}', 'exit_code': 0},
{'tool': 'read_file', 'command': 'config.py', 'exit_code': 1},
{'tool': 'bash', 'command': 'pytest', 'exit_code': 0},
], CompletionRequirements(required_artifacts=('config.py',), verifier_required=True))
assert not ledger.evaluate().can_complete
assert ledger.evaluate().status.value == 'failed'
@pytest.mark.parametrize('command, expected', [
('"$runner" -m pytest -q tests/test_config.py', True),
('"$runner" -m unittest', True),
('"$runner" -m pytest --collect-only', False),
('"$runner" -m pytest || true', False),
('"$runner" -m pytest; echo passed', False),
('"$runner" -m pytest $FLAGS', False),
('echo pytest passed', False),
])
def test_exact_tui_interpreter_selection_preserves_foreground_verifier_status(command, expected):
from src.agent_loop import _tui_python_runner_setup
assert is_test_command(_tui_python_runner_setup() + command) is expected
def test_tui_discovery_fallback_cannot_attest_that_tests_executed():
from src.agent_loop import _tui_local_test_runner_command
assert not is_test_command(_tui_local_test_runner_command())
@pytest.mark.parametrize('command, expected', [('pytest -q', True), ('pytest || true', False)])
def test_host_shell_native_arguments_use_the_same_verifier_classification(command, expected):
from src.agent_evidence import command_is_test, command_is_validation
encoded = json.dumps({'command': command, 'timeout': 120})
assert command_is_test(encoded) is expected
assert command_is_validation(encoded) is expected
+3
View File
@@ -1681,6 +1681,9 @@ def test_calendar_create_response_includes_persistent_event_link(monkeypatch):
set_user_timezone("Asia/Tokyo", 540)
from tests.runtime_evidence_helpers import authoritative_executor
@authoritative_executor
async def _fake_exec(block, *args, **kwargs):
if '"list_calendars"' in (block.content or ""):
return (
+7 -3
View File
@@ -62,17 +62,21 @@ async def test_actual_generator_reports_unavailable_before_inference(
forced_tools={"web_search"}, relevant_tools={"web_search"},
fallbacks=[("https://fallback.invalid", "unused-fallback", {})],
)]
assert len(chunks) == 3
assert len(chunks) == 4
assert json.loads(chunks[0].removeprefix("data: ")) == {
"type": "turn_contract", **selected.audit(),
}
failure = json.loads(chunks[1].removeprefix("data: "))
decision = json.loads(chunks[1].removeprefix("data: "))
assert decision["type"] == "completion_decision"
assert decision["data"]["status"] == "unverified"
assert decision["data"]["evidence_ids"] == []
failure = json.loads(chunks[2].removeprefix("data: "))
assert set(failure) == {"delta"}
assert "can’t perform" in failure["delta"]
assert "unavailable" in failure["delta"]
assert missing in failure["delta"]
assert "haven’t substituted another tool" in failure["delta"]
assert chunks[2] == "data: [DONE]\n\n"
assert chunks[3] == "data: [DONE]\n\n"
def test_unknown_contract_reaches_model_instead_of_forced_clarification():
+7 -7
View File
@@ -65,17 +65,17 @@ The source tree reads **108** `ODYSSEUS_*` variables: 78 an operator may want to
| `ODYSSEUS_LOCAL_MODEL_GATE` | `'true'` | `src/llm_core.py:95` | On by default. Set 0, false, no or off to drop the gate that checks a local endpoint before routing a request to it. |
| `ODYSSEUS_MISTRAL_REASONING_EFFORT` | `'high'` | `src/llm_core.py:1723` | Reasoning effort sent to Mistral thinking-capable models. The API accepts high, medium, low and none. |
| `ODYSSEUS_MLX_IMAGE_VLM_MODEL` | *unset* | `scripts/mlx_image_server.py:299` | Vision-language model id for the MLX image server script. Required unless `--vlm-model` is passed on the command line. |
| `ODYSSEUS_QWEN_ROUTE_THINKING` | `'auto'` | `src/agent_loop.py:166` | Thinking policy for the Qwen routing step. An unrecognized value falls back to `auto`. |
| `ODYSSEUS_QWEN_ROUTE_THINKING` | `'auto'` | `src/agent_loop.py:169` | Thinking policy for the Qwen routing step. An unrecognized value falls back to `auto`. |
### Agent loop and tool execution
| Variable | Default | Read in | What it does |
|---|---|---|---|
| `ODYSSEUS_DISABLE_MCP` | `''` | `src/builtin_mcp.py:89` | Truthy disables MCP entirely, as an escape hatch for compatibility problems with a server. |
| `ODYSSEUS_MAX_VISUAL_EVIDENCE_FRAMES` | `'3'` | `src/agent_loop.py:15366` | How many video frames one tool result may contribute. Clamped to 1-8. |
| `ODYSSEUS_MAX_VISUAL_EVIDENCE_IMAGES` | `'1'` | `src/agent_loop.py:15334` | How many images one tool result may contribute to the model turn. Clamped to 1-8. |
| `ODYSSEUS_MAX_VISUAL_EVIDENCE_FRAMES` | `'3'` | `src/agent_loop.py:15361` | How many video frames one tool result may contribute. Clamped to 1-8. |
| `ODYSSEUS_MAX_VISUAL_EVIDENCE_IMAGES` | `'1'` | `src/agent_loop.py:15329` | How many images one tool result may contribute to the model turn. Clamped to 1-8. |
| `ODYSSEUS_MCP_ALLOWED_COMMANDS` | `''` | `src/agent_tools/admin_tools.py:140` | Security-relevant. Comma-separated allowlist of MCP launcher basenames the agent may start. Empty by default, and the deny list still wins. |
| `ODYSSEUS_PYTHON_TOOL_SITE_PACKAGES` | `''` | `src/agent_tools/subprocess_tools.py:925` | Security-relevant. Absolute package roots, separated by the platform path separator, exposed to the sandboxed Python tool. Empty exposes none. |
| `ODYSSEUS_PYTHON_TOOL_SITE_PACKAGES` | `''` | `src/agent_tools/subprocess_tools.py:927` | Security-relevant. Absolute package roots, separated by the platform path separator, exposed to the sandboxed Python tool. Empty exposes none. |
| `ODYSSEUS_SCRIPT_HOST` | `'localhost'` | `src/builtin_actions.py:919` | Default host for the run-script action. `localhost`, `127.0.0.1`, `local` and empty run locally; any other value runs over SSH. |
| `ODYSSEUS_TOOL_APPROVAL_GATE` | `'0'` | `src/tool_capabilities.py:645` | Security-relevant. Truthy makes tool calls pass through the approval gate. Off by default. |
@@ -194,8 +194,8 @@ Listed for completeness. Setting one of these on a real install is either a no-o
| Variable | Default | Read in | What it does |
|---|---|---|---|
| `ODYSSEUS_CAPTURE_MODEL_REQUESTS` | `''` | `src/agent_loop.py:3921` | Truthy writes model-request snapshots for local debugging. The marker file `/tmp/odysseus_capture_model_requests` enables the same thing. |
| `ODYSSEUS_EXPOSE_RAW_BROWSER_MCP` | `''` | `src/agent_loop.py:4126` | Truthy stops hiding the raw Playwright MCP tools from agent prompts when the private-browser tool is available. |
| `ODYSSEUS_CAPTURE_MODEL_REQUESTS` | `''` | `src/agent_loop.py:3924` | Truthy writes model-request snapshots for local debugging. The marker file `/tmp/odysseus_capture_model_requests` enables the same thing. |
| `ODYSSEUS_EXPOSE_RAW_BROWSER_MCP` | `''` | `src/agent_loop.py:4129` | Truthy stops hiding the raw Playwright MCP tools from agent prompts when the private-browser tool is available. |
| `ODYSSEUS_TOOL_CONTRACT_ROOT` | `'<repo>/scripts'` | `src/clean_agent_preview.py:2183` (+1 more) | Directory holding the tool-contract scripts the clean-agent preview loads. The default is the repository's bundled scripts directory; set the variable to override it. |
### Email
@@ -221,7 +221,7 @@ Listed for completeness. Setting one of these on a real install is either a no-o
| `ODYSSEUS_QA_TEACHER_ATTEMPTS` | `'3'` | `scripts/odysseus_conversation_qa.py:370` | Retry budget for the conversation-QA teacher model call. Clamped to 1-3. |
| `ODYSSEUS_QA_TEACHER_TIMEOUT` | `'120'` | `scripts/odysseus_conversation_qa.py:372` | Timeout in seconds for that call. Clamped to 15-120. |
| `ODYSSEUS_RUNTIME_REVISION` | `''` | `routes/chat_helpers.py:198` (+1 more) | Revision string stamped into each captured SFT trace record, so a trace can be tied back to the build that produced it. |
| `ODYSSEUS_SFT_DISABLE_WORKSPACE_TOOLS` | `'1'` | `src/agent_loop.py:7404` | On by default. Keeps synthetic personal-assistant fixtures out of workspace mode; set 0, false, no or off to let them through. |
| `ODYSSEUS_SFT_DISABLE_WORKSPACE_TOOLS` | `'1'` | `src/agent_loop.py:7407` | On by default. Keeps synthetic personal-assistant fixtures out of workspace mode; set 0, false, no or off to let them through. |
| `ODYSSEUS_SFT_FORCE_UTC_TIMEZONE` | `'0'` | `routes/chat_routes.py:2070` | Truthy forces `sft_` accounts to UTC for deterministic batch generation. Interactive accounts still follow the browser timezone. |
| `ODYSSEUS_SFT_TRACE_CAPTURE` | `'1'` | `routes/chat_helpers.py:161` (+1 more) | On by default, but only for owners whose name starts with `sft_`. Set 0, false, no or off to stop writing training traces. |
| `ODYSSEUS_SFT_TRACE_DIR` | *unset* | `routes/chat_helpers.py:195` (+2 more) | Directory the SFT trace JSONL files are written to. Defaults to `sft_traces` under the data directory. |