mirror of
https://github.com/pewdiepie-archdaemon/odysseus.git
synced 2026-10-07 23:42:21 +02:00
169 lines
35 KiB
Markdown
169 lines
35 KiB
Markdown
# Search quality audit — September 17, 2026
|
||
|
||
Status: **not solved; no quality promotion claimed.** Model F, Odysseus 7011.
|
||
|
||
## Confirmed harness defects corrected
|
||
|
||
- `1771a6f2`: provider results could violate an explicit `site:` scope. Enforce host/subdomain boundaries, reject deceptive URLs, and avoid query relaxation that drops constraints.
|
||
- `85249454`: prefix-only observation truncation could remove later fetched pages. Share the existing 8,000-character budget across source excerpts, retaining attribution and removing duplicate summaries.
|
||
- `2811b6d5`: successful retrieval forced final synthesis regardless of evidence sufficiency. Keep source inspection available; retain discovery/call bounds.
|
||
- `4571e8d2`: model rewrites could lose explicit news intent. Preserve it in queries. HTML extraction now prefers semantic containers, removes navigation, and avoids emitting nested subtrees repeatedly. Extraction cache namespace changed to prevent old extracted bodies masking this fix.
|
||
|
||
## Live evidence, not just test counts
|
||
|
||
Local ignored reports contain public prompts, bounded tool evidence, final answers and per-turn latency:
|
||
|
||
- `reports/clean-v3-search-quality-2026-09-17T20-14-39-062Z.json`: domain filtering stopped unrelated domains for explicitly scoped queries, but Python answer still mismatched its citation. A natural-language “only python.org” constraint was omitted by the model's query. Evidence-reuse follow-up did not search again. A conceptual browser question returned an announcement rather than an explanation.
|
||
- `reports/clean-v3-search-quality-2026-09-17T20-20-56-923Z.json`: Python answer still cited a Python 2.7 page for a 3.14 claim; short news request took 43.3 seconds and ended with generic text and links, not a briefing.
|
||
- `reports/clean-v3-search-quality-2026-09-17T20-24-08-580Z.json`: full 16-conversation suite launched after `4571e8d2`; review is in progress. Early failures include vague AI news despite substantive fetched reports, unsupported browser comparison after two empty searches, and a manual request answered with directions but no link. Simple arithmetic and greeting succeeded in approximately 4.4 seconds without tools.
|
||
|
||
A separate direct endpoint control supplied two short **fictional** reports to Model F (temperature 0, thinking disabled, max_tokens 700). In 4.54 seconds it correctly summarized the parental-consent rule and the speech model's 4-to-12-language change, with the two supplied URLs. This proves only that the model can use short, clean supplied evidence; it does not validate real search or isolate every harness/model interaction.
|
||
|
||
Latest extraction/query regression run: 1,215 passing tests. Passing mechanics or length checks are **not** evidence of factual correctness.
|
||
|
||
### Completed variety run and matched synthesis probe
|
||
|
||
The 16 conversations completed (19 user turns). The run does **not** establish good search quality: examples include irrelevant battery citations, generic or unsupported news, missing manual links, poor source-seeking follow-ups, and a spelling correction incorrectly refused as an operation. Arithmetic, greeting, and the simple browser explanation were clear successes. Evidence reuse avoided another call, but answer quality remained limited.
|
||
|
||
`reports/search-synthesis-probe-1789676901820.json` reuses the exact first news turn's two public evidence outputs, temperature 0, max_tokens 768, thinking disabled. A short research-specific system prompt produced concrete stories in both user-evidence (18.91s) and tool-evidence (10.22s) placement; tool-evidence still supplied only one citation for multiple stories. This is not a fully isolated live-harness A/B: system prompt, prior assistant messages, tool availability, and recovery history also differ. Do not infer a unique cause from this control.
|
||
|
||
Further code inspection identified **automatic citation fabrication by the harness**: web search results were inserted into `entity_result_links`, then appended after model synthesis without claim support verification. Broad answers also received automatic source lists. Removing these paths preserves calendar/research-object navigation links and explicit source-only lookup results. A runtime regression test checks that an old-release search result is not attached as the citation for a latest-release answer. Earlier wrong citations therefore cannot be attributed solely to the model.
|
||
|
||
A temporary loopback relay captured zero requests because registered endpoint IDs override submitted URLs. It was shut down and removed. Endpoint record `1518b6ee` was checked read-only and does map to the same `19211` Model F used by the direct probe. Future evidence capture must respect that registered routing rather than claiming an unused proxy observed traffic.
|
||
|
||
### Sampling and system-prompt controls
|
||
|
||
`reports/search-synthesis-probe-1789677241866.json` used the actual canonical base system-prompt expression with the same tool-evidence messages and no tools offered. It still produced concrete news stories (7.96s), although citations were missing. Therefore the base system prompt alone does **not** explain the live failures; do not replace it on the earlier short-prompt comparison alone.
|
||
|
||
Code inspection found a sampling mismatch: UI default temperature is 1.0; the model-name-based deterministic override recognizes Odysseus/Ajax names, not `model-f`, even though that endpoint explicitly uses compact tool mode. Direct controls used temperature 0. Added an explicit per-test-session temperature option to the verifier and confirmed its persistence in the database. No global or existing user-session defaults changed.
|
||
|
||
Temperature-0 live run: `reports/clean-v3-search-quality-2026-09-17T20-35-39-951Z.json`. News became more concrete, but some claims/citations still need verification; browser comparison still had empty search evidence, and spelling correction was still incorrectly refused. Latency was 41.5s for news, 30.0s for its follow-up, 16.3s for comparison, and 6.6s for spelling. This does not demonstrate an overall quality/speed fix. Search results were not frozen, so this is diagnostic rather than a clean statistical A/B.
|
||
|
||
Post-citation-fix live replay `reports/clean-v3-search-quality-2026-09-17T20-34-07-667Z.json` returned a Python version in 15.9s without appending the unrelated Python 2.7 citation. It still omitted a useful supporting link, so the requested answer is not fully satisfactory.
|
||
|
||
### Supplied-text boundary and date-filter investigation
|
||
|
||
`reports/clean-v3-search-quality-2026-09-17T20-38-19-942Z.json` captured the actual denial for the spelling task: `manage_calendar`, `write_family_not_authorized`. The safety guard was correct; supplied text was being mistaken for operation intent. `e9993b65` introduces a shared explicit text-transformation boundary used by selection, write authority, and the compact offered-tool surface. `4f2cffb1` applies it to the independent document-review completion shortcut too. Ordinary external-editor requests remain outside this narrow classification.
|
||
|
||
Live reports `20-40-14-629Z` and `20-42-04-396Z`: spelling became “I received the calendar invite”; translation no longer called search/email; proofreading no longer demanded an open document. All made zero tool calls. **Proofreading still left a tense error** (“I have deleted ... yesterday”), so this demonstrates a routing/control fix, not full model correctness. Regression suite: 1,207 passed.
|
||
|
||
A direct paired SearXNG query `Firefox Chrome privacy features` returned five results without a publication window (4.02s), and zero with `time_filter=month` (7.04s). Returned pages were mostly generic Firefox pages, so this does not prove adequate comparison evidence. It does show an overly restrictive window can cause avoidable emptiness. Next retrieval work must distinguish current-valid documentation from recently published articles, without silently widening explicit user date restrictions.
|
||
|
||
### Publication-date repair
|
||
|
||
`54abfb9f` shares publication-intent inference between argument repair and the search tool. It removes model-invented windows from reference lookups without requested publication dates, preserves named user windows, stops provider day-to-week widening, and carries explicit filters through metadata/timeout paths. A date-filtered scholarly lookup no longer bypasses the provider through the unfiltered direct-title shortcut. Broader regression run: 1,293 passed.
|
||
|
||
Temperature-0 replay: `reports/clean-v3-search-quality-2026-09-17T20-47-17-632Z.json` (three conversations, four turns). The Firefox/Chrome comparison now retrieved sources and produced a substantive answer (44.3s) instead of the preceding empty-search refusal (16.3s). This is not a validated accuracy win: several current-feature claims still need support checks. Sony's actual official manuals page appeared in evidence; the 17.5s final omitted its link. Mozilla documentation lookup still failed to identify the requested page (14.6s), and its Chrome follow-up supplied an unverified URL (17.6s). No overall promotion claimed.
|
||
|
||
### Explicit source-link completion
|
||
|
||
`ba67ad26` adds one bounded evidence-grounded completion check when the user explicitly requested links but a searched answer omitted them. It does not append a search result as a citation; the model must select an evidenced URL or state the source was not found. This shares the existing answer-recovery budget. Source-request drafts are buffered to avoid displaying the incomplete draft as the final answer.
|
||
|
||
Live `reports/clean-v3-search-quality-2026-09-17T20-51-12-931Z.json`: the Sony lookup now returns the exact official manuals-page URL seen in evidence (18.9s, three rounds, one search), versus omitting it in the preceding 17.5s run. This is a successful link-completion replay, not a statistical latency result. The Python task failed on a model-added month filter; `5cf17293` extends reference-date semantics to version/release lookups and allows a corrected query to identify reference intent while the user's own wording remains authoritative for date constraints. Regression run: 1,226 passed; live version replay pending.
|
||
|
||
### Empty-result latency and relevance audit
|
||
|
||
`b7ed9e58` removes duplicate same-provider requests after a completed empty/irrelevant result set in both search orchestrators. Transport exceptions retain one retry; failure followed by empty response is reported as empty, not a stale transport error. Tests verify exact provider call sequences.
|
||
|
||
`8c090102` prevents a temporal qualifier such as “latest 2026” from being treated as a product model number when filtering documentation. Actual model numbers remain required. It also records effective temperature/output limits in runtime metrics; public test reports now retain the native trace so recovery behavior can be inspected rather than guessed. Regression suite: 1,232 passed.
|
||
|
||
`reports/clean-v3-search-quality-2026-09-17T20-58-28-001Z.json` confirms temperature 0 and max output 768. Mozilla lookup took 10.9s but still failed to find the requested page; Chrome follow-up took 22.6s and linked the generic Chrome homepage, not a proper comparison. These are **not quality passes**. Earlier short-news run `20-55-39-393Z` did perform a follow-up search based on a first-result story and synthesized a concrete answer in 37.1s; factual completeness still needs review. Neither run proves a statistical latency improvement.
|
||
|
||
Further provider inspection found that the news-to-general fallback dropped the date window even after the initial news request retained it. The fallback now inherits constraints and only activates for an actual news-category request (not an explicitly selected general engine). Narrow provider/filter tests: 72 passed.
|
||
|
||
## Outstanding work
|
||
|
||
### Additional informal/multi-part live checks
|
||
|
||
Completion-order replay `reports/clean-v3-search-quality-2026-09-17T21-50-39-701Z.json`: weekly news now performs search → follow-ups → fetch → browser, but ends with inaccessible-source limitation (42.3s/eight rounds), not a completed briefing. Short daily-news answer is substantive/cited but takes 59.1s and has a suspect input/output-pricing sentence requiring evidence audit. Do not promote either based only on workflow/length.
|
||
|
||
Fixed a separate fallback invariant: JSON-provider exceptions previously invoked HTML search without date/category/language/engine constraints. HTML transport now inherits these constraints and omits only format; mock failure regression confirms the exact request parameters across transports. 81 provider/publication/query tests pass. This is a deterministic contract fix, not a demonstrated live answer improvement.
|
||
|
||
Clean context replay `reports/clean-v3-search-quality-2026-09-17T21-48-59-742Z.json` passed the specific context invariant: setup acknowledged without tools/saving (5.17s), “can u look it up” searched Python release schedule (15.39s). Final answer remained generic, so this verifies referent/routing preservation rather than a complete source-rich research answer.
|
||
|
||
Completion ordering now decides whether research expansion is still due before citation/contentless-answer repairs. Previously weekly news performed a tool-free citation rewrite then demanded more search, wasting a round and placing contradictory instructions in history. The regression matrix covers source-requested/non-source-requested, embedded/no embedded article, and empty/successful follow-up search. 1,262 tests passed. Live weekly-news replay pending after deployment.
|
||
|
||
`reports/clean-v3-search-quality-2026-09-17T21-46-32-630Z.json`: all three context-free referential prompts asked sensible clarification questions, zero tools, 7.2–8.2s UI latency. Grounded follow-up searched the correct Python topic, but setup wording “Remember…” also created test-owner memory `5f3eab27-99f1-45cb-8c81-7fb66420b296`. Removed only that exact ID after API owner/text verification; subsequent GET returned 404. Its text remains recoverable in the report. Revised setup explicitly forbids saving, and launched a clean follow-up replay. Never count that setup mutation as a no-tool pass.
|
||
|
||
Answer-style controls `reports/search-synthesis-probe-1789681656044.json` (Firefox) and `1789681683794.json` (battery) replace only the canonical concise-answer sentence with completeness/uncertainty guidance. Results were mixed: Firefox became shorter; battery answer remained broad and introduced unsupported sustainability/cost assertions. No production prompt change made. More prose or links alone is not a factual-quality improvement.
|
||
|
||
`a8646d86` adds missing-subject clarification for complete referential lookup requests only when history has no prior user turn/assistant/tool evidence and there is no active editor, attachment/image, or native workspace. It omits tool schemas and asks the model to clarify; explicit subjects and context-bearing follow-ups retain normal routing. 1,257 related tests pass. Added live no-context variants and a same-wording follow-up with an established Python topic; four-case replay launched after deployment. This is conservative coverage of unresolved references, not a claim to resolve all linguistic ambiguity.
|
||
|
||
Ambiguity controls with exact compact search schema/canonical base prompt: `reports/search-tool-choice-probe-1789681406232.json` reproduces invented game-release query even with auto choice. Adding an explicit missing-subject clarification instruction (`1789681407057`) merely changes invented query to France's capital. No network tool was executed by these probes.
|
||
|
||
`reports/search-tool-choice-probe-1789681431330.json`: tool_choice none while retaining schemas emits raw tool-call markup on two ambiguous prompts. Omitting schemas entirely with the same clarification instruction (`1789681457760`) yields correct clarification questions on all three (“can u look it up”, “please find that”, “what about its price?”), 0.9–1.2 seconds. This supports testing a genuinely context-aware missing-referent boundary, not blanket tool removal or phrase-only blocking of valid follow-ups. No production ambiguity guard deployed yet.
|
||
|
||
Casual citation replay `reports/clean-v3-search-quality-2026-09-17T21-42-14-297Z.json` failed quality: 29.3s/three rounds and final claims it cannot compare without data despite retrieved evidence. Citation recognition alone did not solve synthesis and must not be described as a successful answer-quality fix.
|
||
|
||
Extraction replay `reports/clean-v3-search-quality-2026-09-17T21-40-44-367Z.json` completed. Misspelled battery request remained too shallow/uncited (17.1s/two rounds), although it now framed cost reduction conditionally rather than saying batteries are categorically cheaper/safer. `8cd14e3f` recognizes standalone trailing citation requests such as “sources pls” and tests negative topic/negation cases. 926 relevant regressions passed. Deployed after the replay ended; live validation still pending. Concurrent agent committed the unrelated artifact-path regression as `b0a1f7fd`; that edit was not included in our commits.
|
||
|
||
`reports/clean-v3-search-quality-2026-09-17T21-37-47-091Z.json` remains weak: weekly news took 44.5s and ended by asking the user to open/scroll the page; misspelled battery comparison took 17.6s and gave shallow uncited claims. Its evidence had substantial tag/related-post/reference noise. A fresh inspection of the actual battery page found one article nested within main. Extraction now prefers a single substantive article over its surrounding main wrapper, while multiple article listings preserve main context. Real fetch: 2,413 characters, comparison retained, related posts/comment form absent. 54 extraction/observation tests pass. This does not validate the article's claims: its cost discussion is internally inconsistent, so the model must still qualify/corroborate it. Another agent's unrelated workspace-path test in `tests/test_clean_agent_preview.py` was left untouched and uncommitted by this work.
|
||
|
||
Post-dispatch `reports/clean-v3-search-quality-2026-09-17T21-34-22-848Z.json` completed: all recorded searches had nonempty queries, though extra `command` arguments remained. Short typo news took 34.7s/five rounds versus prior 69.8s/eight rounds; non-frozen retrieval/concurrency prevent treating this as a statistical speed gain. Weekly news synthesized in 53.1s but relied on shallow snippets. The second-story follow-up now fetched the relevant article and explained it (29.3s/two rounds), instead of deterministic link-only output. Recorded article supports its main open-weight/WAICO/Kimi/MAZU points; broader factual corroboration not established.
|
||
|
||
The weekly trace exposed two empty follow-up searches disabling all tools despite earlier discovered source URLs. Search exhaustion now suppresses further search while preserving fetch/browser if sources exist; zero-source exhaustion still ends tool use. A stream regression executes discovery → two empty follow-ups → successful fetch. Related suites: 1,240 passed. Targeted weekly-news replay pending after deployment.
|
||
|
||
The pre-dispatch sweep `reports/clean-v3-search-quality-2026-09-17T21-27-27-686Z.json` finished all 23 conversations. Automated summary: one mechanics failure, 22 awaiting quality review—not 22 quality passes. Manual review: arithmetic/greeting and basic no-search browser explanation succeeded; supplied-text edits avoided tools, but proofreading retained a tense error and correction-only omitted part of the sentence. Research still failed through shallow answers, missing links, unsupported latest-version claims, premature source-only rendering, and round-limit exhaustion. Context-free “can u look it up” invented a game-release query. These are open failures, not a promotion result.
|
||
|
||
After terminal completion, restarted 7011 at `bcd52b8c` to deploy required-single-search dispatch and explicit link-only synthesis bypass. HTTP readiness returned 302. Targeted three-conversation replay launched (weekly news, typo news with follow-up, short typo search); source/claim accuracy and real latency still require review. No training or model checkpoint changes were made.
|
||
|
||
The broad sweep exposed an independent synthesis bypass: “more about the second story, with sources” was rendered as a single source link. Source-only detection was the absence of several explanation keywords rather than a positive link-only command. `87d1edaf` requires a complete explicit link-return request before deterministic source-only rendering; ordinary follow-up explanation remains model synthesis. Related suites: 1,237 passed, followed by 35 focused tests including runtime preservation of explanatory answers. Pending deployment together with forced-search dispatch while the original sweep finishes.
|
||
|
||
Canonical-system confirmation `reports/search-tool-choice-probe-1789680586200.json`: auto and required supplied queries for both prompts; named search choice omitted query in both (and typo prompt emitted `command`). The same compact schema and model were used. Implemented forced-search dispatch as one offered web_search schema with required choice, preserving the forced-tool intent and original schema. Other tool choices remain unchanged. 1,230 routing/runtime regressions pass. **Not deployed yet:** the pre-change 23-conversation sweep remains active (nine conversations complete at this checkpoint); wait for its terminal state before restart and paired replay. This is a demonstrated argument-generation difference, not yet an end-to-end quality/speed win.
|
||
|
||
Full 23-conversation regression launched on `6000b718`/current deployed harness: `reports/clean-v3-search-quality-2026-09-17T21-27-27-686Z.json`. Active handle recorded in session; do not restart based on elapsed observation time.
|
||
|
||
Read-only tool-choice control `reports/search-tool-choice-probe-1789680509169.json` uses the exact compact web_search schema, a short system prompt, identical user prompts/temperature/model, and never executes emitted calls. For both weekly-news and typo-news prompts, auto/required emitted nonempty queries. Forced named mode emitted an extraneous `command` field in both; typo-news omitted query entirely. Six calls are preliminary evidence of tool-choice/schema behavior, not proof of a universal backend defect or a production fix. Next test should use the canonical harness system/history before changing dispatch. Probe script saves full schemas and public emitted calls for reproducibility.
|
||
|
||
`reports/search-synthesis-probe-1789680321559.json` compares identical saved native tool history with/without `_harness_control` messages, same canonical base prompt, no offered tools. Full trace: short answer without links, 2.59s. Controls removed: longer answer with links, 6.39s, but introduced a Do Not Track URL not established by the recorded evidence. This is not grounds to remove recovery controls wholesale or claim a factual quality win.
|
||
|
||
Weekly-news replay `reports/clean-v3-search-quality-2026-09-17T21-24-11-361Z.json` corrected the missing query but still returned no evidence (16.35s). Direct simultaneous provider control with exact query `AI developments this week`, `time_filter=week`: general returned zero, news five. `3ea5a348` recognizes time-qualified developments as news intent while retaining general routing for tutorials, historical discussion, software versions and documentation. Provider/publication/query-relaxation tests: 80 passed. Live weekly-news replay launched after deployment; returned results still require relevance/source review.
|
||
|
||
Timed `reports/clean-v3-search-quality-2026-09-17T21-22-21-304Z.json`: Firefox 31.0s/five rounds, tool execution 5.779s; news 69.8s/eight rounds, tool execution 1.672s. Remaining time includes inference, streaming and orchestration—not proven pure GPU time. The source-link retry still failed on Firefox. Main observed delay is outside tool execution, not search-provider time in these cached runs.
|
||
|
||
`4b6a9721` applies the explicit-query requirement to initial calls too: the weekly-news trace had copied a whole compound request into a missing query. Regression suites: 1,249 passed. Weekly-news replay launched. `reports/search-synthesis-probe-1789680251422.json` feeds the saved Firefox evidence to the same model without live recovery history/offered tools: canonical-system answer took 4.98s and concise research-system answer 4.66s; both supplied a link. This proves the model can emit the link in simplified context, not factual correctness—the linked support page was access-blocked, and some feature assertions still need grounding. Do not infer the system prompt alone or lack of tool schemas uniquely explains the difference.
|
||
|
||
`reports/clean-v3-search-quality-2026-09-17T21-20-16-154Z.json`: rejecting fabricated follow-up queries did not yield a latency win; typo-news used eight rounds/57.6 seconds and still lacked source URLs. Natural weekly-news request returned a failure in 27.0 seconds. Do not claim speed improvement. `7a3e1567` exposes measured execution time per tool (separate from total runtime) in live reports, and recognizes explicit imperative source requests such as “link the instructions” that the prior link-noun patterns missed. Related suites: 1,226 passed; timed news/privacy replay running. Additional completion retries are not a substitute for auditing the underlying answer generation.
|
||
|
||
`reports/clean-v3-search-quality-2026-09-17T21-18-22-350Z.json` remains unsatisfactory: Mozilla lookup 13.9s failed to locate documentation, multi-part privacy request 22.0s omitted requested links and details, Chrome follow-up 28.8s supplied generic homepages instead of comparison. Do not promote based on mechanics.
|
||
|
||
News trace inspection found another synthetic harness distortion: a missing follow-up query was filled with the original user text plus “corroborating analysis authoritative sources.” This reintroduced misspellings and returned no evidence. `63488776` instead raises an explicit argument error asking for an evidence-based follow-up. This avoids an invented network query but does not yet prove reduced total latency or successful model repair. Related suites: 371 passed; live typo-news and natural-news replay launched.
|
||
|
||
News replay `reports/clean-v3-search-quality-2026-09-17T21-15-25-609Z.json` completed: the short misspelled request now synthesizes rather than exhausting the contradictory breadth loop, but takes 47.4 seconds; “ai news today” takes 62.6 seconds and omits actual source URLs. Neither is an accuracy/latency pass. Broad source/claim alignment still requires review.
|
||
|
||
Browser evidence handling now recognizes a structured challenge-page title followed by an empty snapshot, without treating ordinary empty pages or articles with that title as challenges. Failure of both transports for one source no longer forces tool-free completion of the entire research request. Regression exercises failed static fetch → blocked browser → successful alternate fetch. Related suites: 1,218 passed; live replay pending. The older keyword-based gate detector remains broader than the new structured check and needs false-positive audit.
|
||
|
||
Targeted replay `reports/clean-v3-search-quality-2026-09-17T21-13-13-145Z.json`: correction-only request made zero tool calls (7.3s), but only corrected some words rather than returning the whole corrected sentence. Firefox (26.5s) now follows failed `web_fetch` with `private_browser`, proving recovery was exercised. The browser still returned a challenge title and empty snapshot; the final omitted links and was incomplete. Browser navigation success must not be conflated with successful evidence acquisition.
|
||
|
||
The preceding short-news trace exposed contradictory harness controls: “no more tools” was followed twice by a demand to search again because the breadth check counted successful searches, not attempted follow-ups. Breadth recovery is now one-shot, only before a second attempt and before terminal search completion. A stream regression covers a successful first search and empty second search, preserving the final answer rather than demanding endless breadth. Related suites: 1,226 passed. Live replay remains required.
|
||
|
||
`reports/clean-v3-search-quality-2026-09-17T21-09-49-686Z.json` completed five additional cases. No overall quality pass: short misspelled news took 48.9 seconds and exhausted research without synthesis; Firefox instructions took 32.2 seconds and omitted requested links; a context-free “can u look it up” invented a game-release topic; correction-only text incorrectly triggered news research. The Python false-premise answer rejected Python 9.0, but its extra latest-version claim still needs source verification.
|
||
|
||
The Firefox trace showed HTTP-200 access-challenge pages treated as article evidence. `d1db1353` classifies short interstitials using corroborating title/body signals, emits an explicit fetch failure with recovery guidance, leaves ordinary articles intact, and avoids caching transient challenges. `44b56a46` preserves the supplied-text boundary for correction-only phrasing. Combined regression run: 1,249 passed. Both deployed; live targeted replay pending. Neither unit tests nor deployment establishes improved research quality.
|
||
|
||
Earlier `fb669cde` added query-focused extractive passages to preserve relevant evidence beyond page prefixes. `f7532bd3` stopped appending an invented current year to evergreen reference queries. Latest suite covers 23 conversations, not 23 validated successes.
|
||
|
||
Additional matched wording probes (2026-09-17): verifier now includes polished, casual, and misspelled versions of the same official-release request, plus a correction-only control. `reports/clean-v3-search-quality-2026-09-17T21-57-22-320Z.json` completed all four with no mechanical failures; this is not a quality pass. Polished and casual answers gave conflicting latest-release versions, and the casual answer omitted the requested source link. Correction-only returned corrected text without research. Verify claims against captured sources before accepting any release answer.
|
||
|
||
Fixed fictional evidence diagnostic `reports/fixed-search-evidence-20260917T215415214156.json` also demonstrates an answer-level defect independent of live retrieval: the model correctly quoted measured and advertised battery durations but incorrectly said their rankings matched. The typo comparison omitted the requested price difference, while the polished comparison supplied it correctly. Tools in this probe are intercepted; these are not live-web benchmark results.
|
||
|
||
`2cf1c319` fixes an upstream evidence-loss boundary found by those wording probes: WebSearchTool prefix-truncated the full report at 10,000 characters before the runtime balanced excerpts at 8,000. Long early pages erased later CONTENT blocks permanently. The shared compactor now runs before the tool transport cap and again at the runtime budget. New tool-through-runtime regression failed before the patch (only early pages survived) and passes with all five page bodies and original source metadata preserved. Related runtime/routing suites: 1,277 passed; search provider/source-index/query suites: 90 passed. Deployed on 7011, readiness 302. Live replay report `reports/clean-v3-search-quality-2026-09-17T22-01-30-981Z.json` requires completion and manual review; this is not yet a factual answer-quality win.
|
||
|
||
The 22:01 live replay is terminal (four mechanically valid conversations, not four quality passes). Casual release lookup now receives CONTENT 1–5 instead of only 1–2; the evidence-preservation fix is exercised in production. Polished lookup quotes a date present in its retrieved release index and provides a link, but casual lookup still invents a different date not supported by its retrieved older-release pages. Both runs take roughly 15–16 seconds. Thus source preservation is validated; consistency, follow-up verification and claim grounding remain unresolved. No overall quality promotion.
|
||
|
||
Streaming repair `7cdd7886`: user session 6067439a-0e23-4c8c-8c1f-410f3b3acf94 exposed that search/citation requests deliberately buffered every text chunk until final completion. Removed this buffering while retaining canonical final replacement after completion checks. A failing-before/passing-after regression asserts first delta delivery before upstream completion; recovery tests now require visible drafts but clean canonical replacement. Runtime/routing and browser-rendering suites: 1,265 passed. Deployed on 7011. Live news replay `reports/clean-v3-search-quality-2026-09-17T22-11-07-276Z.json` emitted 100 text deltas, no runtime error, 38.7 seconds; it hit the round limit and appended the limit notice, so this validates streamed transport, not research quality or clean completed-answer reconciliation.
|
||
|
||
Completed-answer streaming replay `reports/clean-v3-search-quality-2026-09-17T22-12-10-396Z.json`: 36 text deltas followed by one canonical final response, no runtime error, 12.1 seconds. This exercises both progressive delivery and final reconciliation in the live UI request path.
|
||
|
||
`b12181a1` addresses user session cac81b51-f5b9-40b7-a9e7-245d12af4d91: a streamed draft and its recovered answer remained in separate bubbles because unscoped streamed finals only deduplicated identical text. Corrected-draft first deltas and canonical research finals now explicitly replace prose across the current turn, preserving tool activity. A browser regression verifies one remaining answer with heading/bold/link structure and the same tool node. The original stored answer had plain paragraphs, not lost Markdown; system guidance now asks for headings or bold topic labels and descriptive links for multi-topic research while leaving simple answers brief. Related tests: 1,265 passed. Deployed on 7011; live Japan replay pending manual DOM review in reports/clean-v3-search-quality-2026-09-17T22-25-47-509Z.json.
|
||
|
||
The Japan replay completed: 688 streamed chunks, one canonical final, one visible answer body, nine rendered bold elements and three links. This validates single-bubble final reconciliation and actual Markdown rendering. It took 48.1 seconds, and source/claim quality remains separately unverified; formatting is not evidence of factual correctness or a speed improvement.
|
||
|
||
User follow-up cf01e358-34a0-481f-9e65-b8f4a266a196 showed streamed answer → extra search and plain prose at temperature 1.0. `df4435a1` moves the known broad-research follow-up prerequisite before model generation, suppresses prose during forced tool selection, and explicitly retains readable layout instructions in recovery synthesis. Runtime/routing/rendering tests: 1,265 passed. Replay `reports/clean-v3-search-quality-2026-09-17T22-46-41-660Z.json` uses actual temperature 1.0: Sweden and casual Japan each execute two searches before any streamed answer, end with one visible answer, and render emphasis and links. Sweden: 33.1s, 322 chunks, bold title plus italic topic labels; Japan: 46.0s, 590 chunks, ten bold spans. Not an overall research-quality pass: Sweden misses major national-news breadth and Japan includes an internally inconsistent country comparison. Continue factual-grounding/relevance audit independently from presentation validation.
|
||
|
||
1. Finish and manually audit all 16 conversations; inspect claim/source alignment, request completion, follow-up referents, and latency.
|
||
2. Distinguish provider emptiness from model query drift and unsupported synthesis. Do not label every weak answer a routing defect.
|
||
3. Preserve explicit user source constraints even when model queries omit them; do not infer official provenance from URL appearance.
|
||
4. Investigate why clean short evidence is used correctly in the direct control but substantive live sources produce vague or unsupported answers. Use matched inputs before changing training or adding more completion heuristics.
|
||
5. Keep failures visible. Do not count long answers, citation lists, or successful tool execution as completed research.
|