Record wording robustness probes and source-loss replay evidence

This commit is contained in:
pewdiepie-archdaemon
2026-09-17 22:02:34 +00:00
parent 144c8a3dd6
commit 4de1a4b9bb
3 changed files with 93 additions and 0 deletions
+8
View File
@@ -143,6 +143,14 @@ The Firefox trace showed HTTP-200 access-challenge pages treated as article evid
Earlier `fb669cde` added query-focused extractive passages to preserve relevant evidence beyond page prefixes. `f7532bd3` stopped appending an invented current year to evergreen reference queries. Latest suite covers 23 conversations, not 23 validated successes.
Additional matched wording probes (2026-09-17): verifier now includes polished, casual, and misspelled versions of the same official-release request, plus a correction-only control. `reports/clean-v3-search-quality-2026-09-17T21-57-22-320Z.json` completed all four with no mechanical failures; this is not a quality pass. Polished and casual answers gave conflicting latest-release versions, and the casual answer omitted the requested source link. Correction-only returned corrected text without research. Verify claims against captured sources before accepting any release answer.
Fixed fictional evidence diagnostic `reports/fixed-search-evidence-20260917T215415214156.json` also demonstrates an answer-level defect independent of live retrieval: the model correctly quoted measured and advertised battery durations but incorrectly said their rankings matched. The typo comparison omitted the requested price difference, while the polished comparison supplied it correctly. Tools in this probe are intercepted; these are not live-web benchmark results.
`2cf1c319` fixes an upstream evidence-loss boundary found by those wording probes: WebSearchTool prefix-truncated the full report at 10,000 characters before the runtime balanced excerpts at 8,000. Long early pages erased later CONTENT blocks permanently. The shared compactor now runs before the tool transport cap and again at the runtime budget. New tool-through-runtime regression failed before the patch (only early pages survived) and passes with all five page bodies and original source metadata preserved. Related runtime/routing suites: 1,277 passed; search provider/source-index/query suites: 90 passed. Deployed on 7011, readiness 302. Live replay report `reports/clean-v3-search-quality-2026-09-17T22-01-30-981Z.json` requires completion and manual review; this is not yet a factual answer-quality win.
The 22:01 live replay is terminal (four mechanically valid conversations, not four quality passes). Casual release lookup now receives CONTENT 1–5 instead of only 1–2; the evidence-preservation fix is exercised in production. Polished lookup quotes a date present in its retrieved release index and provides a link, but casual lookup still invents a different date not supported by its retrieved older-release pages. Both runs take roughly 15–16 seconds. Thus source preservation is validated; consistency, follow-up verification and claim grounding remain unresolved. No overall quality promotion.
1. Finish and manually audit all 16 conversations; inspect claim/source alignment, request completion, follow-up referents, and latency.
2. Distinguish provider emptiness from model query drift and unsupported synthesis. Do not label every weak answer a routing defect.
3. Preserve explicit user source constraints even when model queries omit them; do not infer official provenance from URL appearance.