feat: make Ollama/gemma4 the primary backend with Claude as fallback
Rolls back the claudification backend preference: PREFER_CLOUD_BACKEND now defaults to false, resolve_backend() picks Ollama first and uses Claude when explicitly preferred or when the new Ollama startup health check fails. The Steward retries mid-request failures on the other backend in both directions. Also hardens the fallback itself: Anthropic SDK imports are lazy so a broken anthropic package degrades to Ollama-only instead of crashing at import time (root cause of the production outage since April), anthropic is pinned to a pydantic-ai-1.27-compatible range, ANTHROPIC_MODEL defaults to claude-sonnet-5 (sonnet-4-20250514 retired 2026-06-15), sampling parameters are stripped from Claude calls (Sonnet 5 rejects them), and the Steward timeout is configurable (STEWARD_TIMEOUT, default 60s) since gemma4 needs ~35s warm for analysis. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
@@ -204,13 +204,13 @@ async def run_housekeeper(
|
||||
)
|
||||
|
||||
try:
|
||||
# Use temperature 0.1 for slight exploration
|
||||
from pydantic_ai.settings import ModelSettings
|
||||
# Temperature 0.1 for slight exploration (skipped on Claude backend)
|
||||
from src.anthropic.model_selector import get_sampling_settings
|
||||
|
||||
result = await agent.run(
|
||||
prompt,
|
||||
message_history=message_history,
|
||||
model_settings=ModelSettings(temperature=0.1),
|
||||
model_settings=get_sampling_settings(0.1),
|
||||
)
|
||||
|
||||
logger.info(
|
||||
@@ -266,13 +266,13 @@ async def run_housekeeper_stream(
|
||||
)
|
||||
|
||||
try:
|
||||
# Use temperature 0.1 for slight exploration
|
||||
from pydantic_ai.settings import ModelSettings
|
||||
# Temperature 0.1 for slight exploration (skipped on Claude backend)
|
||||
from src.anthropic.model_selector import get_sampling_settings
|
||||
|
||||
async with agent.run_stream(
|
||||
prompt,
|
||||
message_history=message_history,
|
||||
model_settings=ModelSettings(temperature=0.1),
|
||||
model_settings=get_sampling_settings(0.1),
|
||||
) as response:
|
||||
async for delta in response.stream_text(delta=True):
|
||||
yield delta
|
||||
|
||||
Reference in New Issue
Block a user