Record controlled missing-referent and schema-omission diagnostics

This commit is contained in:
pewdiepie-archdaemon
2026-09-17 21:44:46 +00:00
parent e066b5d4dc
commit f8fc4ec4f7
2 changed files with 11 additions and 4 deletions
+6
View File
@@ -75,6 +75,12 @@ Further provider inspection found that the news-to-general fallback dropped the
### Additional informal/multi-part live checks
Ambiguity controls with exact compact search schema/canonical base prompt: `reports/search-tool-choice-probe-1789681406232.json` reproduces invented game-release query even with auto choice. Adding an explicit missing-subject clarification instruction (`1789681407057`) merely changes invented query to France's capital. No network tool was executed by these probes.
`reports/search-tool-choice-probe-1789681431330.json`: tool_choice none while retaining schemas emits raw tool-call markup on two ambiguous prompts. Omitting schemas entirely with the same clarification instruction (`1789681457760`) yields correct clarification questions on all three (“can u look it up”, “please find that”, “what about its price?”), 0.9–1.2 seconds. This supports testing a genuinely context-aware missing-referent boundary, not blanket tool removal or phrase-only blocking of valid follow-ups. No production ambiguity guard deployed yet.
Casual citation replay `reports/clean-v3-search-quality-2026-09-17T21-42-14-297Z.json` failed quality: 29.3s/three rounds and final claims it cannot compare without data despite retrieved evidence. Citation recognition alone did not solve synthesis and must not be described as a successful answer-quality fix.
Extraction replay `reports/clean-v3-search-quality-2026-09-17T21-40-44-367Z.json` completed. Misspelled battery request remained too shallow/uncited (17.1s/two rounds), although it now framed cost reduction conditionally rather than saying batteries are categorically cheaper/safer. `8cd14e3f` recognizes standalone trailing citation requests such as “sources pls” and tests negative topic/negation cases. 926 relevant regressions passed. Deployed after the replay ended; live validation still pending. Concurrent agent committed the unrelated artifact-path regression as `b0a1f7fd`; that edit was not included in our commits.
`reports/clean-v3-search-quality-2026-09-17T21-37-47-091Z.json` remains weak: weekly news took 44.5s and ended by asking the user to open/scroll the page; misspelled battery comparison took 17.6s and gave shallow uncited claims. Its evidence had substantial tag/related-post/reference noise. A fresh inspection of the actual battery page found one article nested within main. Extraction now prefers a single substantive article over its surrounding main wrapper, while multiple article listings preserve main context. Real fetch: 2,413 characters, comparison retained, related posts/comment form absent. 54 extraction/observation tests pass. This does not validate the article's claims: its cost discussion is internally inconsistent, so the model must still qualify/corroborate it. Another agent's unrelated workspace-path test in `tests/test_clean_agent_preview.py` was left untouched and uncommitted by this work.
+5 -4
View File
@@ -6,7 +6,7 @@ const tools = JSON.parse(execFileSync((process.env.PYTHON || "python3"), ['-c',
'import json; from src.clean_agent_preview import compact_schemas; from src.tool_schemas import FUNCTION_TOOL_SCHEMAS; print(json.dumps(compact_schemas([s for s in FUNCTION_TOOL_SCHEMAS if s["function"]["name"] == "web_search"])))'], {encoding:'utf8'}));
const model = process.env.MODEL || 'model-f';
const endpoint = process.env.ENDPOINT_URL || (() => { throw new Error("ENDPOINT_URL is required"); })();
const prompts = ['Catch me up on the biggest AI developments this week. Explain why they matter and link your sources.', 'serch latest ai news pls'];
const prompts = process.env.PROBE_PROMPTS ? JSON.parse(process.env.PROBE_PROMPTS) : ['Catch me up on the biggest AI developments this week. Explain why they matter and link your sources.', 'serch latest ai news pls'];
const results = [];
let system = 'You are Odysseus. Use web_search to find current information relevant to the user request.';
if (process.env.CANONICAL_SYSTEM === '1') {
@@ -24,16 +24,17 @@ shell_clause = 'Shell commands are disabled. '
print(eval(compile(ast.Expression(assignment.value), '<canonical-system-expression>', 'eval')))
`], {input:fs.readFileSync('src/clean_agent_preview.py','utf8'),encoding:'utf8'}).trim();
}
system += process.env.SYSTEM_APPEND || '';
for (const prompt of prompts) {
for (const choice of ['auto', 'required', {type:'function', function:{name:'web_search'}}]) {
for (const choice of (process.env.PROBE_CHOICES ? JSON.parse(process.env.PROBE_CHOICES) : ['auto', 'required', {type:'function', function:{name:'web_search'}}])) {
const started = performance.now();
const response = await fetch(endpoint, {method:'POST', headers:{'Content-Type':'application/json'}, signal:AbortSignal.timeout(90000),
body:JSON.stringify({model, messages:[{role:'system',content:system},{role:'user',content:prompt}], tools, tool_choice:choice, temperature:0, max_tokens:256, stream:false, chat_template_kwargs:{enable_thinking:false}})});
body:JSON.stringify({model, messages:[{role:'system',content:system},{role:'user',content:prompt}], ...(process.env.OMIT_TOOLS === '1' ? {} : {tools, tool_choice:choice}), temperature:0, max_tokens:256, stream:false, chat_template_kwargs:{enable_thinking:false}})});
const data = await response.json();
const result = {prompt,choice,status:response.status,seconds:(performance.now()-started)/1000,message:data.choices?.[0]?.message,error:data.error};
results.push(result); console.log(JSON.stringify(result));
}
}
const target = `reports/search-tool-choice-probe-${Date.now()}.json`;
fs.writeFileSync(target, JSON.stringify({model,system,tools,results},null,2)+'\n');
fs.writeFileSync(target, JSON.stringify({model,system,tools_omitted:process.env.OMIT_TOOLS === '1',tools,results},null,2)+'\n');
console.log(target);