mirror of
https://github.com/pewdiepie-archdaemon/odysseus.git
synced 2026-10-06 06:52:20 +02:00
Record controlled missing-referent and schema-omission diagnostics
This commit is contained in:
@@ -75,6 +75,12 @@ Further provider inspection found that the news-to-general fallback dropped the
|
|||||||
|
|
||||||
### Additional informal/multi-part live checks
|
### Additional informal/multi-part live checks
|
||||||
|
|
||||||
|
Ambiguity controls with exact compact search schema/canonical base prompt: `reports/search-tool-choice-probe-1789681406232.json` reproduces invented game-release query even with auto choice. Adding an explicit missing-subject clarification instruction (`1789681407057`) merely changes invented query to France's capital. No network tool was executed by these probes.
|
||||||
|
|
||||||
|
`reports/search-tool-choice-probe-1789681431330.json`: tool_choice none while retaining schemas emits raw tool-call markup on two ambiguous prompts. Omitting schemas entirely with the same clarification instruction (`1789681457760`) yields correct clarification questions on all three (“can u look it up”, “please find that”, “what about its price?”), 0.9–1.2 seconds. This supports testing a genuinely context-aware missing-referent boundary, not blanket tool removal or phrase-only blocking of valid follow-ups. No production ambiguity guard deployed yet.
|
||||||
|
|
||||||
|
Casual citation replay `reports/clean-v3-search-quality-2026-09-17T21-42-14-297Z.json` failed quality: 29.3s/three rounds and final claims it cannot compare without data despite retrieved evidence. Citation recognition alone did not solve synthesis and must not be described as a successful answer-quality fix.
|
||||||
|
|
||||||
Extraction replay `reports/clean-v3-search-quality-2026-09-17T21-40-44-367Z.json` completed. Misspelled battery request remained too shallow/uncited (17.1s/two rounds), although it now framed cost reduction conditionally rather than saying batteries are categorically cheaper/safer. `8cd14e3f` recognizes standalone trailing citation requests such as “sources pls” and tests negative topic/negation cases. 926 relevant regressions passed. Deployed after the replay ended; live validation still pending. Concurrent agent committed the unrelated artifact-path regression as `b0a1f7fd`; that edit was not included in our commits.
|
Extraction replay `reports/clean-v3-search-quality-2026-09-17T21-40-44-367Z.json` completed. Misspelled battery request remained too shallow/uncited (17.1s/two rounds), although it now framed cost reduction conditionally rather than saying batteries are categorically cheaper/safer. `8cd14e3f` recognizes standalone trailing citation requests such as “sources pls” and tests negative topic/negation cases. 926 relevant regressions passed. Deployed after the replay ended; live validation still pending. Concurrent agent committed the unrelated artifact-path regression as `b0a1f7fd`; that edit was not included in our commits.
|
||||||
|
|
||||||
`reports/clean-v3-search-quality-2026-09-17T21-37-47-091Z.json` remains weak: weekly news took 44.5s and ended by asking the user to open/scroll the page; misspelled battery comparison took 17.6s and gave shallow uncited claims. Its evidence had substantial tag/related-post/reference noise. A fresh inspection of the actual battery page found one article nested within main. Extraction now prefers a single substantive article over its surrounding main wrapper, while multiple article listings preserve main context. Real fetch: 2,413 characters, comparison retained, related posts/comment form absent. 54 extraction/observation tests pass. This does not validate the article's claims: its cost discussion is internally inconsistent, so the model must still qualify/corroborate it. Another agent's unrelated workspace-path test in `tests/test_clean_agent_preview.py` was left untouched and uncommitted by this work.
|
`reports/clean-v3-search-quality-2026-09-17T21-37-47-091Z.json` remains weak: weekly news took 44.5s and ended by asking the user to open/scroll the page; misspelled battery comparison took 17.6s and gave shallow uncited claims. Its evidence had substantial tag/related-post/reference noise. A fresh inspection of the actual battery page found one article nested within main. Extraction now prefers a single substantive article over its surrounding main wrapper, while multiple article listings preserve main context. Real fetch: 2,413 characters, comparison retained, related posts/comment form absent. 54 extraction/observation tests pass. This does not validate the article's claims: its cost discussion is internally inconsistent, so the model must still qualify/corroborate it. Another agent's unrelated workspace-path test in `tests/test_clean_agent_preview.py` was left untouched and uncommitted by this work.
|
||||||
|
|||||||
@@ -6,7 +6,7 @@ const tools = JSON.parse(execFileSync((process.env.PYTHON || "python3"), ['-c',
|
|||||||
'import json; from src.clean_agent_preview import compact_schemas; from src.tool_schemas import FUNCTION_TOOL_SCHEMAS; print(json.dumps(compact_schemas([s for s in FUNCTION_TOOL_SCHEMAS if s["function"]["name"] == "web_search"])))'], {encoding:'utf8'}));
|
'import json; from src.clean_agent_preview import compact_schemas; from src.tool_schemas import FUNCTION_TOOL_SCHEMAS; print(json.dumps(compact_schemas([s for s in FUNCTION_TOOL_SCHEMAS if s["function"]["name"] == "web_search"])))'], {encoding:'utf8'}));
|
||||||
const model = process.env.MODEL || 'model-f';
|
const model = process.env.MODEL || 'model-f';
|
||||||
const endpoint = process.env.ENDPOINT_URL || (() => { throw new Error("ENDPOINT_URL is required"); })();
|
const endpoint = process.env.ENDPOINT_URL || (() => { throw new Error("ENDPOINT_URL is required"); })();
|
||||||
const prompts = ['Catch me up on the biggest AI developments this week. Explain why they matter and link your sources.', 'serch latest ai news pls'];
|
const prompts = process.env.PROBE_PROMPTS ? JSON.parse(process.env.PROBE_PROMPTS) : ['Catch me up on the biggest AI developments this week. Explain why they matter and link your sources.', 'serch latest ai news pls'];
|
||||||
const results = [];
|
const results = [];
|
||||||
let system = 'You are Odysseus. Use web_search to find current information relevant to the user request.';
|
let system = 'You are Odysseus. Use web_search to find current information relevant to the user request.';
|
||||||
if (process.env.CANONICAL_SYSTEM === '1') {
|
if (process.env.CANONICAL_SYSTEM === '1') {
|
||||||
@@ -24,16 +24,17 @@ shell_clause = 'Shell commands are disabled. '
|
|||||||
print(eval(compile(ast.Expression(assignment.value), '<canonical-system-expression>', 'eval')))
|
print(eval(compile(ast.Expression(assignment.value), '<canonical-system-expression>', 'eval')))
|
||||||
`], {input:fs.readFileSync('src/clean_agent_preview.py','utf8'),encoding:'utf8'}).trim();
|
`], {input:fs.readFileSync('src/clean_agent_preview.py','utf8'),encoding:'utf8'}).trim();
|
||||||
}
|
}
|
||||||
|
system += process.env.SYSTEM_APPEND || '';
|
||||||
for (const prompt of prompts) {
|
for (const prompt of prompts) {
|
||||||
for (const choice of ['auto', 'required', {type:'function', function:{name:'web_search'}}]) {
|
for (const choice of (process.env.PROBE_CHOICES ? JSON.parse(process.env.PROBE_CHOICES) : ['auto', 'required', {type:'function', function:{name:'web_search'}}])) {
|
||||||
const started = performance.now();
|
const started = performance.now();
|
||||||
const response = await fetch(endpoint, {method:'POST', headers:{'Content-Type':'application/json'}, signal:AbortSignal.timeout(90000),
|
const response = await fetch(endpoint, {method:'POST', headers:{'Content-Type':'application/json'}, signal:AbortSignal.timeout(90000),
|
||||||
body:JSON.stringify({model, messages:[{role:'system',content:system},{role:'user',content:prompt}], tools, tool_choice:choice, temperature:0, max_tokens:256, stream:false, chat_template_kwargs:{enable_thinking:false}})});
|
body:JSON.stringify({model, messages:[{role:'system',content:system},{role:'user',content:prompt}], ...(process.env.OMIT_TOOLS === '1' ? {} : {tools, tool_choice:choice}), temperature:0, max_tokens:256, stream:false, chat_template_kwargs:{enable_thinking:false}})});
|
||||||
const data = await response.json();
|
const data = await response.json();
|
||||||
const result = {prompt,choice,status:response.status,seconds:(performance.now()-started)/1000,message:data.choices?.[0]?.message,error:data.error};
|
const result = {prompt,choice,status:response.status,seconds:(performance.now()-started)/1000,message:data.choices?.[0]?.message,error:data.error};
|
||||||
results.push(result); console.log(JSON.stringify(result));
|
results.push(result); console.log(JSON.stringify(result));
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
const target = `reports/search-tool-choice-probe-${Date.now()}.json`;
|
const target = `reports/search-tool-choice-probe-${Date.now()}.json`;
|
||||||
fs.writeFileSync(target, JSON.stringify({model,system,tools,results},null,2)+'\n');
|
fs.writeFileSync(target, JSON.stringify({model,system,tools_omitted:process.env.OMIT_TOOLS === '1',tools,results},null,2)+'\n');
|
||||||
console.log(target);
|
console.log(target);
|
||||||
|
|||||||
Reference in New Issue
Block a user