mirror of
https://github.com/pewdiepie-archdaemon/odysseus.git
synced 2026-10-06 06:52:20 +02:00
Squash Odysseus development history
This commit is contained in:
@@ -0,0 +1,75 @@
|
||||
# Agent turn contract
|
||||
|
||||
Scope: product Agent turns on 7011. Environment-owned native/TUI bridges retain
|
||||
their existing execution contract. No model weights or training settings change.
|
||||
|
||||
## Boundaries
|
||||
|
||||
1. `src/turn_contract.py` classifies capabilities, including explicit compound
|
||||
requests and referential follow-ups. Classification is selection, not permission.
|
||||
2. `routes/chat_routes.py` resolves toggles, privileges, global/plan/incognito
|
||||
restrictions, fixture restrictions and available schema inventory before
|
||||
freezing the offered set. Web enabled alone does not select web tools.
|
||||
3. `TurnContract` checks `required <= offered <= executable`, stores immutable
|
||||
serialized schema copies, and records unavailable requirements. An unavailable
|
||||
request stops without inference or substitution; unknown actions ask for clarity.
|
||||
Exact account-discovery requests narrow selection to account metadata only;
|
||||
compounds retain their declared family scope. Media operations declare their
|
||||
existing tool dependencies rather than falling back to shell generation.
|
||||
4. The agent's prompt/schema route and fallback use that same logical scope.
|
||||
Native versus textual serialization remains model-specific. Answer-only phases
|
||||
can suppress tool calls without granting a different scope.
|
||||
Contract turns preserve the already-compacted conversation and tool-call/result
|
||||
IDs. The standalone specialist prompt's latest-message-only behavior is not used
|
||||
for these product turns. Prompt domains also come from the contract.
|
||||
Accepted in-scope calls retain their model-provided arguments and native IDs;
|
||||
the explicit-intent fallback must not overwrite them with the whole user turn.
|
||||
5. The context-bound dispatcher checks membership **and** existing runtime policy,
|
||||
owner restrictions and exact-action approvals. A contract is not authorization
|
||||
to bypass those gates. Contract work bypasses terminating legacy shortcuts.
|
||||
6. `_AgentRenderState` explicitly identifies streamed versus canonical output.
|
||||
Later synthesis transfers ownership with turn-scoped replacement. The frontend
|
||||
reconciles visible DOM, not just accumulated strings; tool evidence is retained.
|
||||
Ownership is included in saved metrics and `message_saved` events.
|
||||
History and resume honor replacement scope. Single-capability turns retain
|
||||
canonical output: an always-synthesize trial caused a live notes loop and was
|
||||
reverted. Compound turns cannot terminate after only one capability's result.
|
||||
|
||||
## Verification
|
||||
|
||||
Use the project's configured Python environment, not an unrelated system Python:
|
||||
|
||||
```sh
|
||||
python -m pytest -q \
|
||||
tests/test_turn_contract.py tests/test_turn_contract_integration.py \
|
||||
tests/test_agent_turn_contract_boundaries.py tests/test_turn_rendering_js.py \
|
||||
tests/test_contract_prompt_conversation.py tests/test_product_turn_contract_route.py \
|
||||
tests/test_contract_explicit_fallback.py \
|
||||
tests/test_history_resume_rendering_js.py \
|
||||
tests/test_chat_route_tool_policy.py tests/test_tool_policy.py \
|
||||
tests/test_frontend_module_version_parity.py
|
||||
node scripts/verify_agent_turn_contract.mjs --max-turns 80 --total-ms 900000
|
||||
```
|
||||
|
||||
The browser verifier uses `sft_alex_creator` and actual 7011 Agent controls. It
|
||||
captures request toggles, SSE contract/tool events, visible output and persisted
|
||||
history. Ten families have four initial/follow-up Web-toggle combinations.
|
||||
Blocked or unrun cases are not passes. Email requires verified fixture isolation;
|
||||
do not enable global fixture mode on the user's live service to make a test pass.
|
||||
|
||||
## Remaining limits
|
||||
|
||||
- Classification is deterministic and vocabulary-based, not a proof of semantic
|
||||
understanding. Add independent behavior examples for confirmed misses.
|
||||
- Schema registration and policy permission do not guarantee a remote provider
|
||||
stays healthy throughout a turn. Runtime failure must remain visible.
|
||||
- Separate tool/argument errors, tool-service failures, rendering failures and
|
||||
verifier defects in reports. Do not infer model accuracy from routing alone.
|
||||
- Canonical summaries can still ignore presentation constraints such as a
|
||||
requested item count. Do not count those as full functional passes. Forcing an
|
||||
extra model round is not a validated general repair for this deployed model.
|
||||
- Keep all imports of a local JS module on the same URL identity. Distinct query
|
||||
versions instantiate separate module state even when source files are identical.
|
||||
|
||||
Live baseline and current matrix results are in `reports/agent-turn-contract-*`.
|
||||
The implementation is not a claim that every family has passed live verification.
|
||||
@@ -0,0 +1,55 @@
|
||||
# Background research → originating chat
|
||||
|
||||
Chat `trigger_research` calls carry a **dispatcher-supplied** `origin_chat_id`.
|
||||
The research start route verifies chat ownership before registering a durable
|
||||
`background_tool_jobs` row and starting the existing research service. Panel
|
||||
jobs have no origin and never inject a chat reply.
|
||||
|
||||
- Chat default: **2 rounds**, 120-second *soft* research budget. Explicit
|
||||
deeper/Auto rounds regain the normal research time budget. Panel defaults
|
||||
remain unchanged. This is not a guaranteed two-minute wall-clock deadline.
|
||||
- A completion callback stores the report and sources. A startup worker also
|
||||
reconciles missed callbacks and research errors/restarts.
|
||||
- When the origin has no active foreground/detached run, its model summarizes
|
||||
the report with thinking off and no tools. An outer 75-second deadline also
|
||||
bounds model-slot waits. If synthesis is unavailable, deliver an honest
|
||||
notice plus the report link; preserve the evidence for follow-ups.
|
||||
- Message and delivery marker commit in one transaction with a deterministic
|
||||
message ID. Report context is stored in server message metadata and injected
|
||||
as untrusted evidence in regular and compact model history. Long excerpts
|
||||
are explicitly marked; the saved full research report remains accessible.
|
||||
- The browser polls owner-scoped `/api/research/chat-jobs/{chat_id}`, appending
|
||||
unseen message IDs only when that chat is current and not streaming. No
|
||||
transcript replacement or forced navigation. Reloaded history deduplicates.
|
||||
- Chat uses the existing agent-thread rail and expandable rows. The compact
|
||||
header shows status and a right-aligned BG task label with the shared whirlpool
|
||||
while running; expanding reveals topic, phase/round, source count and report
|
||||
link. Rows update in place, preserving expansion/focus while chat streams.
|
||||
Completed rows remain visible; zero-source runs show a warning, not success.
|
||||
Progress polling excludes reports and internal fields.
|
||||
|
||||
Other tools are **not automatically backgrounded**. The durable handoff can be
|
||||
reused, but each future producer needs explicit launch/result/permission wiring.
|
||||
|
||||
## Verification
|
||||
|
||||
```sh
|
||||
<configured-path> -q tests/test_background_tool_jobs.py tests/test_research_chat_runtime.py
|
||||
node --test tests/backgroundToolJobs.test.mjs
|
||||
node scripts/verify_background_delivery_isolation.mjs
|
||||
node scripts/verify_background_research_cards.mjs
|
||||
node scripts/verify_background_research_chat.mjs
|
||||
```
|
||||
|
||||
The last script uses disposable `sft_alex_creator` chats and real research/model
|
||||
calls, then removes only its own reports/chats. Do not use real-user mutations.
|
||||
It checks two-round launch, continued chat, automatic arrival, no transcript
|
||||
rebuild/duplicates, reload, and a follow-up. Inspect retained report excerpts
|
||||
and generated summary when it fails; do not equate job launch with good research.
|
||||
|
||||
Initial live runs verified delivery/navigation/follow-ups but exposed a summary
|
||||
attempt-count bug (fixed: helper requires **1 attempt**, not `max_retries=0`).
|
||||
A later full run was interrupted by an inference endpoint outage. The corrected
|
||||
summary path separately passed a real-model evidence/limitations/citation probe.
|
||||
All targeted Python tests passed (441); real DOM isolation checks passed. A clean
|
||||
full live run with useful retrieved evidence remains to be recorded.
|
||||
@@ -0,0 +1,73 @@
|
||||
# Typo-tolerant tool routing audit
|
||||
|
||||
The 9B SFT model was not retrained. This audit targets the earlier harness
|
||||
stage that decides which complete tool families the model is allowed to see.
|
||||
|
||||
## Method
|
||||
|
||||
- Source prompts: real `sft_alex_creator` sessions from `a37dcb3b-...` onward.
|
||||
- Labels: recorded single-family tool calls, excluding mixed/ambiguous traces.
|
||||
- Variants: deletion, adjacent transposition, duplicated character,
|
||||
keyboard-neighbor substitution, and accidental word split.
|
||||
- Split: deterministic SHA-256 assignment before scoring (75% dev, 25% blind).
|
||||
- Safety: static routing only; no historical mutation or send action is replayed.
|
||||
- Acceptance: at least 95% blind exact-family accuracy and below 1% blind
|
||||
wrong-family authorization. Abstention is measured separately.
|
||||
|
||||
## Results
|
||||
|
||||
| Router | Dev family supplied | Blind family supplied | Blind exact | Blind wrong-family |
|
||||
|---|---:|---:|---:|---:|
|
||||
| Previous exact rules | 63.64% | 65.69% | — | — |
|
||||
| Conservative fuzzy fallback r4 | 96.31% | 98.31% | 96.62% | 0.00% |
|
||||
| Final router + safe-read repair | 98.31% | 98.73% | 97.05% | 0.00% |
|
||||
|
||||
The fallback runs only for action/lookup-shaped requests, resolves exactly one
|
||||
nearby family term, and abstains on ambiguity. Conceptual questions remain
|
||||
tool-free. Complete family schemas are still selected by the immutable turn
|
||||
contract; fuzzy matching never chooses an individual tool or its arguments.
|
||||
|
||||
Authoritative machine reports:
|
||||
|
||||
- `reports/typo-tool-routing-baseline-20260909.json`
|
||||
- `reports/typo-tool-routing-fuzzy-r4-20260909.json`
|
||||
- `reports/typo-tool-routing-final-20260909.json`
|
||||
- `reports/post-followup-agent-80-20260909.json`
|
||||
- `reports/post-typo-routing-agent-80-20260909.json`
|
||||
- `reports/live-typo-agent-20-20260909.json`
|
||||
- `reports/live-typo-unresolved-r3-20260909.json`
|
||||
- `reports/live-typo-agent-final-20-20260909.json`
|
||||
- `reports/post-typo-safe-read-agent-final-80-20260909.json`
|
||||
|
||||
## Live 7011 findings
|
||||
|
||||
The post-deployment standard matrix passed 80/80 through the real Agent UI.
|
||||
The first read-only typo matrix then attempted 17 of 20 planned turns before
|
||||
its total-time limit. Initial Notes, Calendar, Email, Tasks, Documents, and
|
||||
Cookbook calls passed. Completed failing turns still had the correct family
|
||||
and required tool in `turn_contract.offered`; the 9B model sometimes answered
|
||||
without calling that offered tool. Memory and Search also exposed timeouts.
|
||||
|
||||
This separates three failure classes:
|
||||
|
||||
1. **Tool injection:** addressed by conservative fuzzy family routing; blind
|
||||
exact routing is 96.62% with zero blind wrong-family authorizations.
|
||||
2. **Required read execution:** a correctly offered safe list/refresh tool can
|
||||
still be skipped by the model, especially after a typo or on “list those
|
||||
again” follow-ups. This should be handled by the generic deterministic
|
||||
safe-read path, not additional prompt-specific hints.
|
||||
3. **Runtime timeout:** Search and one Memory follow-up require loop/backend
|
||||
diagnosis. A timeout is not counted as a model-accuracy or routing result.
|
||||
|
||||
The generic safe-read parser and search-family precedence were then repaired.
|
||||
The previously unresolved Calendar, Email, Search, and Shell/Files cases passed
|
||||
8/8. The complete typo matrix passed 20/20, including initial requests and
|
||||
follow-ups for all ten families. The final standard Agent UI compatibility
|
||||
matrix passed 80/80 across family, Web-toggle, and follow-up combinations.
|
||||
|
||||
The broad routing regression suite passed 458 tests. The model was not
|
||||
retrained and no DeepSeek API was used: the measured defect was in harness
|
||||
family selection and deterministic safe-read execution, upstream of the
|
||||
model. All 1,535 unique labeled historical turns were statically audited to
|
||||
mine failure categories. Historical write/send/delete actions were not replayed
|
||||
against live data; live verification used the deduplicated read-only matrices.
|
||||
@@ -0,0 +1,26 @@
|
||||
# Skills lifecycle
|
||||
|
||||
The UI exposes All, Built-in, Approved, and Draft. Draft includes archived
|
||||
records so they remain inspectable and recoverable. Built-ins are not audited.
|
||||
Approved means published, passing, at the configured confidence threshold,
|
||||
and not marked unnecessary. Baseline speed measurements remain evidence, not
|
||||
an additional hidden UI approval gate.
|
||||
|
||||
Automatic audits process at most eight eligible records at a time, oldest first.
|
||||
New records are eligible immediately; inconclusive checks retry after a day;
|
||||
failed repairs retry after a week. Passed, duplicate-skipped, and archived records
|
||||
are excluded. Existing daily Skills Audit tasks drive this queue. Their quiet
|
||||
window deferrals propagate to the scheduler rather than becoming task failures.
|
||||
Automatic runs use background model scheduling. Existing self-repair and teacher
|
||||
repair stages remain in place; failed candidates remain drafts.
|
||||
|
||||
The skill index advertises short descriptions; the agent loads a relevant full
|
||||
procedure on demand and applies already-injected procedures directly. Extraction
|
||||
prefers verified discoveries and specific workarounds over routine tool usage.
|
||||
|
||||
Reference reviewed: NousResearch/hermes-agent, MIT license, commit
|
||||
cfdbbb6e35010ace89fbe8243ee82fa4de143e10, cloned to
|
||||
<configured-path> In particular tools/skills_tool.py and
|
||||
agent/prompt_builder.py use progressive disclosure and task-triggered procedure
|
||||
loading. These changes adapt that approach to Odysseus's existing registry;
|
||||
no Hermes implementation code was copied.
|
||||
Reference in New Issue
Block a user