From 1eb1afc88b355091abfcde1bb664b6c138b92252 Mon Sep 17 00:00:00 2001 From: pewdiepie-archdaemon Date: Thu, 17 Sep 2026 22:48:03 +0000 Subject: [PATCH] Verify research-before-streaming at interactive temperature --- docs/search-quality-audit-20260917.md | 2 ++ tests/test_clean_agent_preview.py | 2 ++ 2 files changed, 4 insertions(+) diff --git a/docs/search-quality-audit-20260917.md b/docs/search-quality-audit-20260917.md index 2c8fe7183..7a0201e8a 100644 --- a/docs/search-quality-audit-20260917.md +++ b/docs/search-quality-audit-20260917.md @@ -159,6 +159,8 @@ Completed-answer streaming replay `reports/clean-v3-search-quality-2026-09-17T22 The Japan replay completed: 688 streamed chunks, one canonical final, one visible answer body, nine rendered bold elements and three links. This validates single-bubble final reconciliation and actual Markdown rendering. It took 48.1 seconds, and source/claim quality remains separately unverified; formatting is not evidence of factual correctness or a speed improvement. +User follow-up cf01e358-34a0-481f-9e65-b8f4a266a196 showed streamed answer → extra search and plain prose at temperature 1.0. `df4435a1` moves the known broad-research follow-up prerequisite before model generation, suppresses prose during forced tool selection, and explicitly retains readable layout instructions in recovery synthesis. Runtime/routing/rendering tests: 1,265 passed. Replay `reports/clean-v3-search-quality-2026-09-17T22-46-41-660Z.json` uses actual temperature 1.0: Sweden and casual Japan each execute two searches before any streamed answer, end with one visible answer, and render emphasis and links. Sweden: 33.1s, 322 chunks, bold title plus italic topic labels; Japan: 46.0s, 590 chunks, ten bold spans. Not an overall research-quality pass: Sweden misses major national-news breadth and Japan includes an internally inconsistent country comparison. Continue factual-grounding/relevance audit independently from presentation validation. + 1. Finish and manually audit all 16 conversations; inspect claim/source alignment, request completion, follow-up referents, and latency. 2. Distinguish provider emptiness from model query drift and unsupported synthesis. Do not label every weak answer a routing defect. 3. Preserve explicit user source constraints even when model queries omit them; do not infer official provenance from URL appearance. diff --git a/tests/test_clean_agent_preview.py b/tests/test_clean_agent_preview.py index 451ca0f78..5d8878e66 100644 --- a/tests/test_clean_agent_preview.py +++ b/tests/test_clean_agent_preview.py @@ -1307,6 +1307,8 @@ async def test_stream_research_prerequisite_precedes_broad_web_answer(monkeypatc for event in events ) assert sum(event.get('reason') == 'research_before_synthesis' for event in events) == 1 + phase_index = next(i for i, event in enumerate(events) if event.get('reason') == 'research_before_synthesis') + assert not any(event.get('delta') for event in events[:phase_index]) assert any( event.get('type') == 'final_response' and 'fuller evidence-based briefing' in event.get('content', '')