The compact (clean v3) runtime had no effective context window: it learned a
limit only reactively from a provider 400/413 and its terminal metrics carried
no context_length. PR #41 addressed the reporting gap by probing provider
metadata between the last model byte and [DONE], unauthenticated, and folded
known-table and endpoint evidence into one "known" flag.
Resolve the window once, before the first model request, instead:
- src/agent_runtime/context_resolution.py adds a typed ContextResolution
(effective value, evidence class, source, all observations, conflicts,
provider_io, cached, secret-free probe errors). Evidence classes stay
distinct: runtime_confirmed (llama.cpp /slots, /props, or a limit the
provider stated this turn), provider_advertised (models catalog),
operator_declared (client_runtime_context.model_context_window),
known_table, unknown (0, never a default).
- Selection is deterministic: runtime beats provider beats table; an
operator declaration caps measured evidence and replaces weaker evidence.
Disagreements are recorded as conflicts; a declaration below a measured
value is a cap, above it a contradiction.
- The provider probe forwards the turn's credentials only to the provider's
own origin, runs URL resolution off the event loop, is bounded by one
deadline, never raises, and caches remote results per credential
fingerprint (shorter TTL for failures; local servers are re-probed).
- stream_preview resolves at preparation (or accepts a supplied resolution),
seeds the proactive trim budget from it when evidence is not unknown, and
terminal metrics report only the stored resolution plus any limit the
provider stated during the turn. Metrics perform no discovery.
src/agent_loop.py and the regular runtime's legacy model_context probe are
unchanged. A conftest guard keeps tests that drive the compact runtime with
placeholder endpoints from performing real DNS/HTTP lookups.