feat(backend): BACKEND_SLOT_PINNING — per-phase engine slot ownership
Each pipeline phase owns one llama-server slot (steward 0, orchestrator 1, synthesizer 2), carried as id_slot in extra_body through the same mechanism tool_choice already uses, so the phase's stable prompt prefix stays in that slot's KV cache and a turn re-prefills only its new tokens. Off by default; a no-op on the Claude backend and ignored by Ollama, so the flag is safe on any backend and the cutover itself stays a pure env swap. The merge helper preserves existing extra_body keys — mutation-checked (dropping the merge fails exactly the test written for it). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
@@ -259,6 +259,28 @@ def get_tool_choice_settings() -> ModelSettings:
|
||||
return ModelSettings(extra_body={"tool_choice": "required"})
|
||||
|
||||
|
||||
def with_slot_pinning(settings: ModelSettings | None, slot: int) -> ModelSettings | None:
|
||||
"""
|
||||
Merge llama-server slot pinning into model settings when enabled.
|
||||
|
||||
Each pipeline phase owns one engine slot (steward=0, orchestrator=1,
|
||||
synthesizer=2), so the phase's stable prompt prefix stays in that
|
||||
slot's KV cache and a turn re-prefills only its new tokens. Off by
|
||||
default (BACKEND_SLOT_PINNING); a no-op on the Claude backend, and
|
||||
Ollama ignores the field, so enabling it is safe on any backend.
|
||||
"""
|
||||
if not config.BACKEND_SLOT_PINNING or resolve_backend() == "claude":
|
||||
return settings
|
||||
|
||||
from pydantic_ai.settings import ModelSettings
|
||||
|
||||
merged = dict(settings or {})
|
||||
extra_body = dict(merged.get("extra_body") or {})
|
||||
extra_body["id_slot"] = slot
|
||||
merged["extra_body"] = extra_body
|
||||
return ModelSettings(**merged)
|
||||
|
||||
|
||||
def get_sampling_settings(temperature: float) -> ModelSettings:
|
||||
"""
|
||||
Get model_settings with a sampling temperature where the backend allows it.
|
||||
|
||||
Reference in New Issue
Block a user