mirror of
https://github.com/pewdiepie-archdaemon/odysseus.git
synced 2026-10-06 15:02:20 +02:00
Record wording robustness probes and source-loss replay evidence
This commit is contained in:
@@ -143,6 +143,14 @@ The Firefox trace showed HTTP-200 access-challenge pages treated as article evid
|
||||
|
||||
Earlier `fb669cde` added query-focused extractive passages to preserve relevant evidence beyond page prefixes. `f7532bd3` stopped appending an invented current year to evergreen reference queries. Latest suite covers 23 conversations, not 23 validated successes.
|
||||
|
||||
Additional matched wording probes (2026-09-17): verifier now includes polished, casual, and misspelled versions of the same official-release request, plus a correction-only control. `reports/clean-v3-search-quality-2026-09-17T21-57-22-320Z.json` completed all four with no mechanical failures; this is not a quality pass. Polished and casual answers gave conflicting latest-release versions, and the casual answer omitted the requested source link. Correction-only returned corrected text without research. Verify claims against captured sources before accepting any release answer.
|
||||
|
||||
Fixed fictional evidence diagnostic `reports/fixed-search-evidence-20260917T215415214156.json` also demonstrates an answer-level defect independent of live retrieval: the model correctly quoted measured and advertised battery durations but incorrectly said their rankings matched. The typo comparison omitted the requested price difference, while the polished comparison supplied it correctly. Tools in this probe are intercepted; these are not live-web benchmark results.
|
||||
|
||||
`2cf1c319` fixes an upstream evidence-loss boundary found by those wording probes: WebSearchTool prefix-truncated the full report at 10,000 characters before the runtime balanced excerpts at 8,000. Long early pages erased later CONTENT blocks permanently. The shared compactor now runs before the tool transport cap and again at the runtime budget. New tool-through-runtime regression failed before the patch (only early pages survived) and passes with all five page bodies and original source metadata preserved. Related runtime/routing suites: 1,277 passed; search provider/source-index/query suites: 90 passed. Deployed on 7011, readiness 302. Live replay report `reports/clean-v3-search-quality-2026-09-17T22-01-30-981Z.json` requires completion and manual review; this is not yet a factual answer-quality win.
|
||||
|
||||
The 22:01 live replay is terminal (four mechanically valid conversations, not four quality passes). Casual release lookup now receives CONTENT 1–5 instead of only 1–2; the evidence-preservation fix is exercised in production. Polished lookup quotes a date present in its retrieved release index and provides a link, but casual lookup still invents a different date not supported by its retrieved older-release pages. Both runs take roughly 15–16 seconds. Thus source preservation is validated; consistency, follow-up verification and claim grounding remain unresolved. No overall quality promotion.
|
||||
|
||||
1. Finish and manually audit all 16 conversations; inspect claim/source alignment, request completion, follow-up referents, and latency.
|
||||
2. Distinguish provider emptiness from model query drift and unsupported synthesis. Do not label every weak answer a routing defect.
|
||||
3. Preserve explicit user source constraints even when model queries omit them; do not infer official provenance from URL appearance.
|
||||
|
||||
@@ -0,0 +1,81 @@
|
||||
"""Real model + canonical runtime; fixed fictional search evidence, no web writes.
|
||||
|
||||
Diagnostic only: these results are not a real-web benchmark score. Every tool
|
||||
execution is intercepted in this process, so no public/example URL is fetched.
|
||||
"""
|
||||
import asyncio
|
||||
import json
|
||||
import os
|
||||
import sys
|
||||
import time
|
||||
from datetime import datetime, timezone
|
||||
from pathlib import Path
|
||||
|
||||
sys.path.insert(0, str(Path(__file__).resolve().parents[1]))
|
||||
import src.clean_agent_preview as runtime
|
||||
from src.tool_policy import ToolPolicy
|
||||
from src.tool_schemas import FUNCTION_TOOL_SCHEMAS
|
||||
from src.turn_contract import resolve_full_inventory_contract
|
||||
|
||||
SOURCES = {
|
||||
'https://example.org/alder': ('Alder fictional logger specification',
|
||||
'Alder costs $37 per device. It connects through USB only; it does not support Wi-Fi. '
|
||||
'The manufacturer advertises 40 hours of battery life. An independent test measured 26 hours. '
|
||||
'Both values refer to the same device; advertising is not a measured result.'),
|
||||
'https://example.org/birch': ('Birch fictional logger specification',
|
||||
'Birch costs $52 per device and supports Wi-Fi and USB. The manufacturer advertises '
|
||||
'32 hours of battery life; the same independent test measured 29 hours. '
|
||||
'No shipping cost or warranty duration was supplied for either product.'),
|
||||
}
|
||||
REPORT = '```sources\n' + '\n'.join(f'[{i}] {title}\n {url}' for i, (url, (title, _)) in enumerate(SOURCES.items(), 1)) + '\n```\nQuery: fictional logger specifications\n'
|
||||
REPORT += '\n'.join(f'\n[CONTENT {i}] From: {url}\nTitle: {title}\n------------------------------\n{body}' for i, (url, (title, body)) in enumerate(SOURCES.items(), 1))
|
||||
CASES = [
|
||||
('comparison', 'Search for the fictional Alder and Birch logger specifications. Compare price, Wi-Fi support and measured battery life. How much more does Birch cost? Cite sources.'),
|
||||
('comparison-typo', 'serch alder vs birch loggers, price diffrence wifi and tested battry life? sources pls'),
|
||||
('evidence-boundary', 'Look up the fictional Alder and Birch loggers. Which lasts longer in the independent test, and is that the same ranking as the advertised battery life? What are their warranty durations? Cite sources.'),
|
||||
]
|
||||
|
||||
async def execute(block, **kwargs):
|
||||
if block.tool_type == 'web_search':
|
||||
return 'fixed search', {'output': REPORT, 'exit_code': 0, 'evidence_status': 'available'}
|
||||
if block.tool_type == 'web_fetch':
|
||||
args = json.loads(block.content)
|
||||
url = args.get('url')
|
||||
if url in SOURCES:
|
||||
title, body = SOURCES[url]
|
||||
return 'fixed fetch', {'output': f'# {title}\nSource: {url}\n{body}', 'exit_code': 0}
|
||||
return 'fixed fixture', {'error': 'No fixture for this tool or URL. Do not invent evidence.', 'exit_code': 1}
|
||||
|
||||
async def main():
|
||||
schemas = [s for s in FUNCTION_TOOL_SCHEMAS if s['function']['name'] in {'web_search', 'web_fetch'}]
|
||||
contract = resolve_full_inventory_contract(schemas=schemas, policy=ToolPolicy())
|
||||
original = runtime.execute_tool_block
|
||||
runtime.execute_tool_block = execute
|
||||
results = []
|
||||
path = Path('reports') / ('fixed-search-evidence-' + datetime.now(timezone.utc).strftime('%Y%m%dT%H%M%S%f') + '.json')
|
||||
path.parent.mkdir(exist_ok=True)
|
||||
try:
|
||||
for name, prompt in CASES:
|
||||
started = time.monotonic()
|
||||
events = []
|
||||
async for chunk in runtime.stream_preview(
|
||||
endpoint_url=os.environ["ENDPOINT_URL"],
|
||||
model=os.environ.get('MODEL', 'model-f'), headers={},
|
||||
messages=[{'role': 'user', 'content': prompt}], turn_contract=contract,
|
||||
session_id='fixed-search-evidence', owner='sft_alex_creator',
|
||||
disabled_tools=set(), tool_policy=ToolPolicy(), max_rounds=8,
|
||||
):
|
||||
if chunk.startswith('data: ') and '[DONE]' not in chunk:
|
||||
events.append(json.loads(chunk[6:]))
|
||||
finals = [e['content'] for e in events if e.get('type') == 'final_response']
|
||||
answer = finals[-1] if finals else ''.join(e.get('delta', '') for e in events)
|
||||
result = {'name': name, 'prompt': prompt, 'seconds': time.monotonic()-started, 'answer': answer, 'events': events}
|
||||
results.append(result)
|
||||
path.write_text(json.dumps({'fixture_sources': SOURCES, 'diagnostic_only': True, 'results': results}, indent=2) + '\n')
|
||||
print(json.dumps({k:v for k,v in result.items() if k != 'events'}), flush=True)
|
||||
finally:
|
||||
runtime.execute_tool_block = original
|
||||
print(path, flush=True)
|
||||
|
||||
if __name__ == '__main__':
|
||||
asyncio.run(main())
|
||||
@@ -137,6 +137,10 @@ try {
|
||||
['no-web-missing-object', ['please find that'], false],
|
||||
['no-web-missing-price', ['what about its price?'], false],
|
||||
['grounded-lookup-followup', ['My next question is about the Python release schedule. For now, just acknowledge; do not search or save anything.', 'can u look it up'], [false, true]],
|
||||
['official-release-polished', ['Find the latest stable Python release on the official website. Give its version, release date, and source link.'], true],
|
||||
['official-release-casual', ['whats the newest stable python? version + date + official link pls'], true],
|
||||
['official-release-misspelled', ['whats teh newst stable pythno? verison date n offical link pls'], true],
|
||||
['no-web-quoted-release', ['Correct spelling only: whats teh newst stable pythno? verison date n offical link pls'], false],
|
||||
];
|
||||
async function runCase([name, prompts, needsWeb]) {
|
||||
const scenario = { name, status: 'running', turns: [] };
|
||||
|
||||
Reference in New Issue
Block a user