diff --git a/docs/search-quality-audit-20260917.md b/docs/search-quality-audit-20260917.md index d9886549d..306c45623 100644 --- a/docs/search-quality-audit-20260917.md +++ b/docs/search-quality-audit-20260917.md @@ -75,6 +75,10 @@ Further provider inspection found that the news-to-general fallback dropped the ### Additional informal/multi-part live checks +The pre-dispatch sweep `reports/clean-v3-search-quality-2026-09-17T21-27-27-686Z.json` finished all 23 conversations. Automated summary: one mechanics failure, 22 awaiting quality review—not 22 quality passes. Manual review: arithmetic/greeting and basic no-search browser explanation succeeded; supplied-text edits avoided tools, but proofreading retained a tense error and correction-only omitted part of the sentence. Research still failed through shallow answers, missing links, unsupported latest-version claims, premature source-only rendering, and round-limit exhaustion. Context-free “can u look it up” invented a game-release query. These are open failures, not a promotion result. + +After terminal completion, restarted 7011 at `bcd52b8c` to deploy required-single-search dispatch and explicit link-only synthesis bypass. HTTP readiness returned 302. Targeted three-conversation replay launched (weekly news, typo news with follow-up, short typo search); source/claim accuracy and real latency still require review. No training or model checkpoint changes were made. + The broad sweep exposed an independent synthesis bypass: “more about the second story, with sources” was rendered as a single source link. Source-only detection was the absence of several explanation keywords rather than a positive link-only command. `87d1edaf` requires a complete explicit link-return request before deterministic source-only rendering; ordinary follow-up explanation remains model synthesis. Related suites: 1,237 passed, followed by 35 focused tests including runtime preservation of explanatory answers. Pending deployment together with forced-search dispatch while the original sweep finishes. Canonical-system confirmation `reports/search-tool-choice-probe-1789680586200.json`: auto and required supplied queries for both prompts; named search choice omitted query in both (and typo prompt emitted `command`). The same compact schema and model were used. Implemented forced-search dispatch as one offered web_search schema with required choice, preserving the forced-tool intent and original schema. Other tool choices remain unchanged. 1,230 routing/runtime regressions pass. **Not deployed yet:** the pre-change 23-conversation sweep remains active (nine conversations complete at this checkpoint); wait for its terminal state before restart and paired replay. This is a demonstrated argument-generation difference, not yet an end-to-end quality/speed win.