Squash Odysseus development history

This commit is contained in:
pewdiepie-archdaemon
2026-09-11 06:04:19 +00:00
parent c9dd68d890
commit 84aa9a91de
871 changed files with 265870 additions and 27854 deletions
+75
View File
@@ -0,0 +1,75 @@
# Agent turn contract
Scope: product Agent turns on 7011. Environment-owned native/TUI bridges retain
their existing execution contract. No model weights or training settings change.
## Boundaries
1. `src/turn_contract.py` classifies capabilities, including explicit compound
requests and referential follow-ups. Classification is selection, not permission.
2. `routes/chat_routes.py` resolves toggles, privileges, global/plan/incognito
restrictions, fixture restrictions and available schema inventory before
freezing the offered set. Web enabled alone does not select web tools.
3. `TurnContract` checks `required <= offered <= executable`, stores immutable
serialized schema copies, and records unavailable requirements. An unavailable
request stops without inference or substitution; unknown actions ask for clarity.
Exact account-discovery requests narrow selection to account metadata only;
compounds retain their declared family scope. Media operations declare their
existing tool dependencies rather than falling back to shell generation.
4. The agent's prompt/schema route and fallback use that same logical scope.
Native versus textual serialization remains model-specific. Answer-only phases
can suppress tool calls without granting a different scope.
Contract turns preserve the already-compacted conversation and tool-call/result
IDs. The standalone specialist prompt's latest-message-only behavior is not used
for these product turns. Prompt domains also come from the contract.
Accepted in-scope calls retain their model-provided arguments and native IDs;
the explicit-intent fallback must not overwrite them with the whole user turn.
5. The context-bound dispatcher checks membership **and** existing runtime policy,
owner restrictions and exact-action approvals. A contract is not authorization
to bypass those gates. Contract work bypasses terminating legacy shortcuts.
6. `_AgentRenderState` explicitly identifies streamed versus canonical output.
Later synthesis transfers ownership with turn-scoped replacement. The frontend
reconciles visible DOM, not just accumulated strings; tool evidence is retained.
Ownership is included in saved metrics and `message_saved` events.
History and resume honor replacement scope. Single-capability turns retain
canonical output: an always-synthesize trial caused a live notes loop and was
reverted. Compound turns cannot terminate after only one capability's result.
## Verification
Use the project's configured Python environment, not an unrelated system Python:
```sh
python -m pytest -q \
tests/test_turn_contract.py tests/test_turn_contract_integration.py \
tests/test_agent_turn_contract_boundaries.py tests/test_turn_rendering_js.py \
tests/test_contract_prompt_conversation.py tests/test_product_turn_contract_route.py \
tests/test_contract_explicit_fallback.py \
tests/test_history_resume_rendering_js.py \
tests/test_chat_route_tool_policy.py tests/test_tool_policy.py \
tests/test_frontend_module_version_parity.py
node scripts/verify_agent_turn_contract.mjs --max-turns 80 --total-ms 900000
```
The browser verifier uses `sft_alex_creator` and actual 7011 Agent controls. It
captures request toggles, SSE contract/tool events, visible output and persisted
history. Ten families have four initial/follow-up Web-toggle combinations.
Blocked or unrun cases are not passes. Email requires verified fixture isolation;
do not enable global fixture mode on the user's live service to make a test pass.
## Remaining limits
- Classification is deterministic and vocabulary-based, not a proof of semantic
understanding. Add independent behavior examples for confirmed misses.
- Schema registration and policy permission do not guarantee a remote provider
stays healthy throughout a turn. Runtime failure must remain visible.
- Separate tool/argument errors, tool-service failures, rendering failures and
verifier defects in reports. Do not infer model accuracy from routing alone.
- Canonical summaries can still ignore presentation constraints such as a
requested item count. Do not count those as full functional passes. Forcing an
extra model round is not a validated general repair for this deployed model.
- Keep all imports of a local JS module on the same URL identity. Distinct query
versions instantiate separate module state even when source files are identical.
Live baseline and current matrix results are in `reports/agent-turn-contract-*`.
The implementation is not a claim that every family has passed live verification.
+55
View File
@@ -0,0 +1,55 @@
# Background research → originating chat
Chat `trigger_research` calls carry a **dispatcher-supplied** `origin_chat_id`.
The research start route verifies chat ownership before registering a durable
`background_tool_jobs` row and starting the existing research service. Panel
jobs have no origin and never inject a chat reply.
- Chat default: **2 rounds**, 120-second *soft* research budget. Explicit
deeper/Auto rounds regain the normal research time budget. Panel defaults
remain unchanged. This is not a guaranteed two-minute wall-clock deadline.
- A completion callback stores the report and sources. A startup worker also
reconciles missed callbacks and research errors/restarts.
- When the origin has no active foreground/detached run, its model summarizes
the report with thinking off and no tools. An outer 75-second deadline also
bounds model-slot waits. If synthesis is unavailable, deliver an honest
notice plus the report link; preserve the evidence for follow-ups.
- Message and delivery marker commit in one transaction with a deterministic
message ID. Report context is stored in server message metadata and injected
as untrusted evidence in regular and compact model history. Long excerpts
are explicitly marked; the saved full research report remains accessible.
- The browser polls owner-scoped `/api/research/chat-jobs/{chat_id}`, appending
unseen message IDs only when that chat is current and not streaming. No
transcript replacement or forced navigation. Reloaded history deduplicates.
- Chat uses the existing agent-thread rail and expandable rows. The compact
header shows status and a right-aligned BG task label with the shared whirlpool
while running; expanding reveals topic, phase/round, source count and report
link. Rows update in place, preserving expansion/focus while chat streams.
Completed rows remain visible; zero-source runs show a warning, not success.
Progress polling excludes reports and internal fields.
Other tools are **not automatically backgrounded**. The durable handoff can be
reused, but each future producer needs explicit launch/result/permission wiring.
## Verification
```sh
<configured-path> -q tests/test_background_tool_jobs.py tests/test_research_chat_runtime.py
node --test tests/backgroundToolJobs.test.mjs
node scripts/verify_background_delivery_isolation.mjs
node scripts/verify_background_research_cards.mjs
node scripts/verify_background_research_chat.mjs
```
The last script uses disposable `sft_alex_creator` chats and real research/model
calls, then removes only its own reports/chats. Do not use real-user mutations.
It checks two-round launch, continued chat, automatic arrival, no transcript
rebuild/duplicates, reload, and a follow-up. Inspect retained report excerpts
and generated summary when it fails; do not equate job launch with good research.
Initial live runs verified delivery/navigation/follow-ups but exposed a summary
attempt-count bug (fixed: helper requires **1 attempt**, not `max_retries=0`).
A later full run was interrupted by an inference endpoint outage. The corrected
summary path separately passed a real-model evidence/limitations/citation probe.
All targeted Python tests passed (441); real DOM isolation checks passed. A clean
full live run with useful retrieved evidence remains to be recorded.
+73
View File
@@ -0,0 +1,73 @@
# Typo-tolerant tool routing audit
The 9B SFT model was not retrained. This audit targets the earlier harness
stage that decides which complete tool families the model is allowed to see.
## Method
- Source prompts: real `sft_alex_creator` sessions from `a37dcb3b-...` onward.
- Labels: recorded single-family tool calls, excluding mixed/ambiguous traces.
- Variants: deletion, adjacent transposition, duplicated character,
keyboard-neighbor substitution, and accidental word split.
- Split: deterministic SHA-256 assignment before scoring (75% dev, 25% blind).
- Safety: static routing only; no historical mutation or send action is replayed.
- Acceptance: at least 95% blind exact-family accuracy and below 1% blind
wrong-family authorization. Abstention is measured separately.
## Results
| Router | Dev family supplied | Blind family supplied | Blind exact | Blind wrong-family |
|---|---:|---:|---:|---:|
| Previous exact rules | 63.64% | 65.69% | — | — |
| Conservative fuzzy fallback r4 | 96.31% | 98.31% | 96.62% | 0.00% |
| Final router + safe-read repair | 98.31% | 98.73% | 97.05% | 0.00% |
The fallback runs only for action/lookup-shaped requests, resolves exactly one
nearby family term, and abstains on ambiguity. Conceptual questions remain
tool-free. Complete family schemas are still selected by the immutable turn
contract; fuzzy matching never chooses an individual tool or its arguments.
Authoritative machine reports:
- `reports/typo-tool-routing-baseline-20260909.json`
- `reports/typo-tool-routing-fuzzy-r4-20260909.json`
- `reports/typo-tool-routing-final-20260909.json`
- `reports/post-followup-agent-80-20260909.json`
- `reports/post-typo-routing-agent-80-20260909.json`
- `reports/live-typo-agent-20-20260909.json`
- `reports/live-typo-unresolved-r3-20260909.json`
- `reports/live-typo-agent-final-20-20260909.json`
- `reports/post-typo-safe-read-agent-final-80-20260909.json`
## Live 7011 findings
The post-deployment standard matrix passed 80/80 through the real Agent UI.
The first read-only typo matrix then attempted 17 of 20 planned turns before
its total-time limit. Initial Notes, Calendar, Email, Tasks, Documents, and
Cookbook calls passed. Completed failing turns still had the correct family
and required tool in `turn_contract.offered`; the 9B model sometimes answered
without calling that offered tool. Memory and Search also exposed timeouts.
This separates three failure classes:
1. **Tool injection:** addressed by conservative fuzzy family routing; blind
exact routing is 96.62% with zero blind wrong-family authorizations.
2. **Required read execution:** a correctly offered safe list/refresh tool can
still be skipped by the model, especially after a typo or on “list those
again” follow-ups. This should be handled by the generic deterministic
safe-read path, not additional prompt-specific hints.
3. **Runtime timeout:** Search and one Memory follow-up require loop/backend
diagnosis. A timeout is not counted as a model-accuracy or routing result.
The generic safe-read parser and search-family precedence were then repaired.
The previously unresolved Calendar, Email, Search, and Shell/Files cases passed
8/8. The complete typo matrix passed 20/20, including initial requests and
follow-ups for all ten families. The final standard Agent UI compatibility
matrix passed 80/80 across family, Web-toggle, and follow-up combinations.
The broad routing regression suite passed 458 tests. The model was not
retrained and no DeepSeek API was used: the measured defect was in harness
family selection and deterministic safe-read execution, upstream of the
model. All 1,535 unique labeled historical turns were statically audited to
mine failure categories. Historical write/send/delete actions were not replayed
against live data; live verification used the deduplicated read-only matrices.
+26
View File
@@ -0,0 +1,26 @@
# Skills lifecycle
The UI exposes All, Built-in, Approved, and Draft. Draft includes archived
records so they remain inspectable and recoverable. Built-ins are not audited.
Approved means published, passing, at the configured confidence threshold,
and not marked unnecessary. Baseline speed measurements remain evidence, not
an additional hidden UI approval gate.
Automatic audits process at most eight eligible records at a time, oldest first.
New records are eligible immediately; inconclusive checks retry after a day;
failed repairs retry after a week. Passed, duplicate-skipped, and archived records
are excluded. Existing daily Skills Audit tasks drive this queue. Their quiet
window deferrals propagate to the scheduler rather than becoming task failures.
Automatic runs use background model scheduling. Existing self-repair and teacher
repair stages remain in place; failed candidates remain drafts.
The skill index advertises short descriptions; the agent loads a relevant full
procedure on demand and applies already-injected procedures directly. Extraction
prefers verified discoveries and specific workarounds over routine tool usage.
Reference reviewed: NousResearch/hermes-agent, MIT license, commit
cfdbbb6e35010ace89fbe8243ee82fa4de143e10, cloned to
<configured-path> In particular tools/skills_tool.py and
agent/prompt_builder.py use progressive disclosure and task-triggered procedure
loading. These changes adapt that approach to Odysseus's existing registry;
no Hermes implementation code was copied.