From d924ae3e26d080b8a317d2d03248b9fafa00e8ad Mon Sep 17 00:00:00 2001 From: pewdiepie-archdaemon Date: Thu, 17 Sep 2026 21:37:34 +0000 Subject: [PATCH] Record live dispatch improvement and source-retrieval exhaustion defect --- docs/search-quality-audit-20260917.md | 4 ++++ 1 file changed, 4 insertions(+) diff --git a/docs/search-quality-audit-20260917.md b/docs/search-quality-audit-20260917.md index 306c45623..47a61637d 100644 --- a/docs/search-quality-audit-20260917.md +++ b/docs/search-quality-audit-20260917.md @@ -75,6 +75,10 @@ Further provider inspection found that the news-to-general fallback dropped the ### Additional informal/multi-part live checks +Post-dispatch `reports/clean-v3-search-quality-2026-09-17T21-34-22-848Z.json` completed: all recorded searches had nonempty queries, though extra `command` arguments remained. Short typo news took 34.7s/five rounds versus prior 69.8s/eight rounds; non-frozen retrieval/concurrency prevent treating this as a statistical speed gain. Weekly news synthesized in 53.1s but relied on shallow snippets. The second-story follow-up now fetched the relevant article and explained it (29.3s/two rounds), instead of deterministic link-only output. Recorded article supports its main open-weight/WAICO/Kimi/MAZU points; broader factual corroboration not established. + +The weekly trace exposed two empty follow-up searches disabling all tools despite earlier discovered source URLs. Search exhaustion now suppresses further search while preserving fetch/browser if sources exist; zero-source exhaustion still ends tool use. A stream regression executes discovery → two empty follow-ups → successful fetch. Related suites: 1,240 passed. Targeted weekly-news replay pending after deployment. + The pre-dispatch sweep `reports/clean-v3-search-quality-2026-09-17T21-27-27-686Z.json` finished all 23 conversations. Automated summary: one mechanics failure, 22 awaiting quality review—not 22 quality passes. Manual review: arithmetic/greeting and basic no-search browser explanation succeeded; supplied-text edits avoided tools, but proofreading retained a tense error and correction-only omitted part of the sentence. Research still failed through shallow answers, missing links, unsupported latest-version claims, premature source-only rendering, and round-limit exhaustion. Context-free “can u look it up” invented a game-release query. These are open failures, not a promotion result. After terminal completion, restarted 7011 at `bcd52b8c` to deploy required-single-search dispatch and explicit link-only synthesis bypass. HTTP readiness returned 302. Targeted three-conversation replay launched (weekly news, typo news with follow-up, short typo search); source/claim accuracy and real latency still require review. No training or model checkpoint changes were made.