Harden maintainer-preview harness and review fixes

Unify conversational domain routing, preserve artifact completion evidence, replace provisional tool-round prose with terminal synthesis, and resolve verified maintainer review findings across search, frontend module identity, path policy, configuration, and built-in skill startup.
This commit is contained in:
pewdiepie-archdaemon
2026-09-21 06:54:03 +00:00
parent d7cad0621f
commit 297ad19248
54 changed files with 1223 additions and 110 deletions
+223 -13
View File
@@ -7672,7 +7672,7 @@ If a calendar create/update request lacks a required date, time, or target event
"send_to_session": "- ```send_to_session``` — Send a message to another session. Line 1 = session_id, rest = message. Use for orchestrating work across sessions.",
"search_chats": "- ```search_chats``` — Search past session transcripts for direct conversation evidence. Use when user asks 'did we discuss X?', 'find the conversation about Y', or when prior chat context is more appropriate than persistent memory.",
"pipeline": "- ```pipeline``` — Run a multi-step AI pipeline. Args (JSON) with ordered steps, each specifying a model and prompt. Use for complex workflows.",
"ui_control": "- ```ui_control``` — Control the UI: toggle tools on/off, OPEN PANELS, open email reply drafts, switch models, change themes. Commands: `toggle <name> on/off` (names: bash/shell, web/search, research, incognito, document_editor/documents), `open_panel <name>` (panels: documents, gallery, calendar/schedule, email, sessions, notes, memories/brain, skills, settings, theme, cookbook), `open_panel calendar month|week|year|agenda [YYYY-MM or YYYY-MM-DD]` (open calendar directly to a view/range), `open_email_reply <uid> <folder> <reply|reply-all|ai-reply> <body text>` (opens an email compose document pre-filled with body, DOES NOT send; use this for normal “write/draft a reply saying X” requests), `set_mode agent/chat`, `switch_model <name>`, `set_theme <preset>`, `create_theme <name> <bg> <fg> <panel> <border> <accent>` (optional key=val for advanced colors AND background effects: bgPattern=<none|dots|synapse|rain|constellations|perlin-flow|petals|sparkles|embers>, bgEffectColor=#RRGGBB, bgEffectIntensity=<num>, bgEffectSize=<num>, frosted=true|false). \"open calendar\" / \"open schedule\" / \"open documents\" / \"open library\" / \"show gallery\" / \"open inbox\" / \"open notes\" / \"open theme\" / \"open cookbook\" all map to `open_panel <name>`. Built-in theme presets: dark, light, midnight, paper, cyberpunk, retrowave, forest, ocean, ume, copper, terminal, organs, lavender, gpt, claude, cute. For any other vibe/name, use create_theme.",
"ui_control": "- ```ui_control``` — Control the UI: toggle tools on/off, OPEN PANELS, open email reply drafts, switch models, change themes. Commands: `toggle <name> on/off` (names: bash/shell, web/search, research, incognito, document_editor/documents), `open_panel <name>` (panels: documents, gallery, calendar/schedule, email, sessions, notes, memories/brain, skills, settings, theme, cookbook), `open_panel calendar month|week|year|agenda [YYYY-MM or YYYY-MM-DD]` (open calendar directly to a view/range), `open_email_reply <uid> <folder> <reply|reply-all|ai-reply> <body text>` (opens an email compose document pre-filled with body, DOES NOT send; use this for normal “write/draft a reply saying X” requests), `set_mode agent/chat`, `switch_model <name>`, `set_theme <preset>`, `create_theme <name> <bg> <fg> <panel> <border> <accent>` (optional key=val for advanced colors AND background effects: bgPattern=<none|dots|synapse|rain|constellations|perlin-flow|petals|sparkles|embers>, bgEffectColor=#RRGGBB, bgEffectIntensity=<num>, bgEffectSize=<num>, frosted=true|false). \"open calendar\" / \"open schedule\" / \"open documents\" / \"open library\" / \"show gallery\" / \"open inbox\" / \"open notes\" / \"open theme\" / \"open cookbook\" all map to `open_panel <name>`. Built-in theme presets: dark, light, midnight, cyberpunk, retrowave, forest, ocean, ume, terminal, organs, gpt, claude, cute, eclipse, porcelain, arcade, blueprint, monolith, yoyo. For any other vibe/name, use create_theme.",
"ask_user": "- ```ask_user``` — Ask the user a question when the task is genuinely ambiguous and the answer changes what you do next (pick an approach, confirm an assumption, choose a target). Args (JSON): {\"question\": \"...\", \"options\": [{\"label\": \"...\", \"description\": \"...\"?}, ...], \"multi\": false?}. 2-6 options. The user gets clickable buttons; calling this ENDS your turn and their choice comes back as your next message. For open-ended missing data such as an exact calendar date, include an \"Exact date\" option and ask the user to type the date; do not invent arbitrary choices. Prefer sensible defaults — only ask when you truly can't proceed well without their input.",
"update_plan": "- ```update_plan``` — While executing an approved plan, write the plan back: tick steps done or revise them. Args (JSON): {\"plan\": \"- [x] done step\\n- [ ] next step\"}. Always pass the COMPLETE checklist, not a diff. Call it after finishing each step (mark it `- [x]`) and whenever the user asks to change the plan. The user's docked plan window updates live. Does nothing if there's no active plan.",
"list_served_models": "- ```list_served_models``` — Show what the Cookbook (LLM-serving subsystem) is currently running. NO args. Use this for ANY 'what's running' / 'what's serving' / 'show my cookbook' / 'is anything up' query. DO NOT shell out (`ps aux`, `docker ps`, etc.) — this tool is the source of truth. Failed serve tasks include recent logs plus diagnosis/retry suggestions; use those suggestions to call `serve_model` again with an adjusted command when appropriate.",
@@ -13695,6 +13695,76 @@ def _is_odysseus_qwen_native(model: str) -> bool:
return bool(re.search(r"\bqwen3(?:\.?(?:6|8))-27b-(?:mlx|fp8)(?:\b|[-_/])", value))
def _is_deepseek_flash_vision_model(model: str) -> bool:
"""Recognize the provider's exact vision-capable Flash variant."""
value = str(model or "").strip().lower().rstrip("/")
return value.rsplit("/", 1)[-1] == "deepseek-flash"
def _deepseek_flash_visual_continuation(
request_messages: Sequence[Mapping[str, Any]],
direct_user_text: str,
) -> Optional[list[dict]]:
"""Flatten one post-tool visual turn for DeepSeek Flash.
The hosted Flash vision path can reason over pixels and emit native tool
calls from a fresh multimodal request. It currently returns an empty,
length-terminated response when the image follows assistant/tool-call
history. Collapse only requests carrying explicit tool visual evidence;
ordinary text and later tool rounds retain their full history.
"""
newest_visual: Optional[Mapping[str, Any]] = None
for message in request_messages or ():
metadata = message.get("metadata") or {}
content = message.get("content")
if (
message.get("role") == "user"
and isinstance(metadata, Mapping)
and metadata.get("source") == "tool visual evidence"
and isinstance(content, list)
and any(
isinstance(block, Mapping) and block.get("type") == "image_url"
for block in content
)
):
newest_visual = message
if newest_visual is None:
return None
systems = [
dict(message)
for message in request_messages or ()
if message.get("role") == "system"
]
visual_content = newest_visual.get("content") or []
images = [
dict(block)
for block in visual_content
if isinstance(block, Mapping) and block.get("type") == "image_url"
]
if not images:
return None
task = str(direct_user_text or "").strip()
evidence_text = "\n".join(
str(block.get("text") or "").strip()
for block in visual_content
if isinstance(block, Mapping)
and block.get("type") == "text"
and str(block.get("text") or "").strip()
)
instruction = (
(f"{task}\n\n" if task else "")
+ (f"{evidence_text}\n\n" if evidence_text else "")
+ "The requested visual evidence is attached below. Analyze these pixels "
"directly and continue the task using downstream tools. Do not request "
"another inspection of this same view."
)
return systems + [{
"role": "user",
"content": [{"type": "text", "text": instruction}, *images],
}]
def _ody_qwen_temperature_cap(temperature):
"""Force-cap odysseus-qwen3 sampling; the finetune destabilizes above 0.2.
@@ -15663,9 +15733,30 @@ def _requested_post_edit_verification(text: str) -> bool:
return False
if _requested_verification_command(value):
return True
if re.search(
r"\b(?:inspect|review|check|verify|read(?:\s+it)?\s+back)\b"
r".{0,100}\b(?:saved|written|created|output|file|artifact)\b",
value,
re.IGNORECASE | re.DOTALL,
):
return True
return bool(re.search(
r"\b(?:then|after(?:wards)?|and)\b.{0,100}\b(?:run|execute|test|verify|check|build|compile|lint)\b"
r"|\b(?:run|execute|test|verify|check|build|compile|lint)\b.{0,100}\b(?:after|once|when)\b",
r"\b(?:then|after(?:wards)?|and)\b.{0,100}\b(?:run|execute|test|verify|check|inspect|review|read(?:\s+it)?\s+back|build|compile|lint)\b"
r"|\b(?:run|execute|test|verify|check|inspect|review|read(?:\s+it)?\s+back|build|compile|lint)\b.{0,100}\b(?:after|once|when)\b",
value,
re.IGNORECASE | re.DOTALL,
))
def _requested_artifact_readback(text: str) -> bool:
"""Whether verification specifically asks to inspect the saved artifact."""
value = str(text or "")
return bool(re.search(
r"\b(?:inspect|review|check|verify|read(?:\s+it)?\s+back)\b"
r".{0,100}\b(?:saved|written|created|output|file|artifact)\b"
r"|\b(?:saved|written|created|output|file|artifact)\b"
r".{0,100}\b(?:inspect|review|check|verify|read(?:\s+it)?\s+back)\b",
value,
re.IGNORECASE | re.DOTALL,
))
@@ -15703,6 +15794,32 @@ def _first_explicit_workspace_file(text: str) -> str:
return _clean_file_edit_value(str(match.group("path") or "").strip().rstrip("."))
def _read_file_block_path(content) -> str:
"""Path argument of a read_file tool block, JSON args or bare text."""
text = str(content or "").strip()
try:
args = json.loads(text)
if isinstance(args, dict):
return str(args.get("path") or "").strip()
except (TypeError, ValueError, json.JSONDecodeError):
pass
return text.splitlines()[0].strip() if text else ""
def _read_file_targets_artifact(content, target) -> bool:
"""True when a read_file block reads the artifact awaiting verification.
Reading the *input* named earlier in the same prompt must not satisfy a
request to verify the written output.
"""
if not target:
return False
path = _read_file_block_path(content)
if not path:
return False
return path == str(target) or Path(path).name == Path(str(target)).name
def _explicit_workspace_files(text: str) -> list[str]:
"""Return concrete source/test paths named in a workspace request."""
paths: list[str] = []
@@ -20297,6 +20414,7 @@ async def stream_agent_loop(
external_tool_schemas=external_tool_schemas,
max_tokens=max_tokens,
max_rounds=max_rounds,
max_tool_calls=max_tool_calls,
temperature=temperature,
):
yield chunk
@@ -22847,6 +22965,16 @@ async def stream_agent_loop(
]
+ _declared_native_artifacts
))
# A forced read-back must target a declared *output*.
# _workspace_artifacts is in prompt order, so index 0 is the input for
# the ordinary "read /workspace/in/x, write /workspace/out/y, then
# check the saved file" shape. Declared required_artifacts are
# authoritative outputs; otherwise prefer the last named path, which is
# the deliverable in that phrasing, over the first.
_artifact_readback_target = (
_declared_native_artifacts[0] if _declared_native_artifacts
else (_workspace_artifacts[-1] if _workspace_artifacts else None)
)
_artifact_creation_requested = bool(
(workspace or _native_artifact_runtime)
and _workspace_artifacts
@@ -23156,6 +23284,16 @@ async def stream_agent_loop(
"[agent-context] final trimmed request lost direct user turn; restoring it before provider call: %r",
_last_user[:160],
)
_trimmed_visual_evidence = [
message
for message in trimmed_messages
if (
isinstance(message, dict)
and message.get("role") == "user"
and (message.get("metadata") or {}).get("source")
== "tool visual evidence"
)
]
trimmed_messages = [
message for message in trimmed_messages
if not (
@@ -23164,7 +23302,14 @@ async def stream_agent_loop(
and (message.get("metadata") or {}).get("trusted") is False
and (message.get("metadata") or {}).get("source")
)
] + [{"role": "user", "content": _last_user}]
] + [
{"role": "user", "content": _last_user},
# Keep the newest tool pixels after the restored task
# text. Provider sanitization merges these consecutive
# user turns into one final multimodal turn; hosted
# vision APIs may ignore images stranded in an older turn.
*_trimmed_visual_evidence,
]
after_trim_tokens = estimate_tokens(trimmed_messages)
if after_trim_tokens < before_trim_tokens:
logger.info(
@@ -24019,6 +24164,7 @@ async def stream_agent_loop(
_single_execution_bound = _request_forbids_execution_retry(_last_user)
_execution_tool_attempts: dict[str, int] = {}
_post_edit_verification_required = _requested_post_edit_verification(_last_user)
_artifact_readback_requested = _requested_artifact_readback(_last_user)
_post_edit_verification_command = _requested_verification_command(_last_user)
if _post_edit_verification_required and not _post_edit_verification_command and _tui_test_request:
_post_edit_verification_command = _tui_local_fallback_shell_command(
@@ -25265,7 +25411,19 @@ async def stream_agent_loop(
candidate_model,
state["messages"],
)
state["request_messages"] = request_messages
deepseek_visual_messages = None
if _is_deepseek_flash_vision_model(candidate_model):
deepseek_visual_messages = _deepseek_flash_visual_continuation(
request_messages,
_last_user,
)
if deepseek_visual_messages is not None:
request_messages = deepseek_visual_messages
logger.info(
"[agent] flattened DeepSeek Flash post-tool visual "
"continuation and suppressed redundant inspect_media"
)
state["request_messages"] = request_messages
_last_route_request_messages = request_messages
state["context_length"] = _route_context_lengths.get(
(candidate_url, candidate_model),
@@ -25274,6 +25432,12 @@ async def stream_agent_loop(
_last_route_context_length = state["context_length"]
run_security.observe_messages(request_messages)
candidate_tools = _tool_schemas_for_route(state)
if deepseek_visual_messages is not None:
candidate_tools = [
schema
for schema in candidate_tools or ()
if schema.get("function", {}).get("name") != "inspect_media"
]
state["tools"] = candidate_tools
from src.generation_budget import fit_output_token_budget
@@ -27170,7 +27334,31 @@ async def stream_agent_loop(
native_tool_calls = []
used_native = False
logger.info("[agent] normalized inspection follow-up to one edit_file call")
if (
elif (
_artifact_readback_requested
and _post_effectful_mutation_done
and not _post_edit_verification_completed
and not _post_edit_verification_force_attempted
and _artifact_readback_target
):
# The user explicitly asked to inspect the saved artifact. Once a
# write succeeds, normalize one bounded read-back rather than
# letting a weak router reopen source-media inspection forever.
# Chained onto the preceding branches: an already-normalized
# authorized edit must not be overwritten by this read.
tool_blocks = [ToolBlock(
"read_file",
json.dumps({"path": _artifact_readback_target}),
)]
converted_calls = []
native_tool_calls = []
used_native = False
_post_edit_verification_force_attempted = True
logger.info(
"[agent] normalized post-edit artifact verification to read_file: %s",
_artifact_readback_target,
)
elif (
_post_edit_verification_nudge_sent
and (_post_effectful_mutation_done or _inspection_edit_completed or _file_creation_completed)
and not _post_edit_verification_completed
@@ -29453,9 +29641,13 @@ async def stream_agent_loop(
"role": "system",
"content": (
"The requested file edit succeeded, but the user also asked "
"for verification. Do that now with one concrete tool call "
"using the requested command (host_shell), then summarize. "
"Do not stop after the edit."
"for verification. "
+ (
"Read the saved output artifact now with read_file, then summarize. "
if _artifact_readback_requested
else "Do that now with one concrete tool call using the requested command (host_shell), then summarize. "
)
+ "Do not stop after the edit."
),
})
yield f'data: {json.dumps({"type": "agent_step", "round": round_num + 1})}\n\n'
@@ -34531,7 +34723,7 @@ async def stream_agent_loop(
]
_workspace_read_requires_mutation = True
if (
block.tool_type in {"host_shell", "bash", "python"}
block.tool_type in {"host_shell", "bash", "python", "read_file"}
and tool_result_is_successful(result)
and (
_post_edit_verification_nudge_sent
@@ -34539,6 +34731,20 @@ async def stream_agent_loop(
)
and (
not _post_edit_verification_required
or (
block.tool_type == "read_file"
and _artifact_readback_requested
and _post_effectful_mutation_done
# Only the artifact under verification counts. When no
# target could be resolved, fall back to the previous
# any-read behaviour so the turn cannot deadlock.
and (
not _artifact_readback_target
or _read_file_targets_artifact(
block.content, _artifact_readback_target
)
)
)
or (
_tui_test_request
and command_is_test(block.content)
@@ -35033,9 +35239,13 @@ async def stream_agent_loop(
"role": "system",
"content": (
"The requested file edit succeeded, but the user also asked "
"for verification. Do that now with one concrete tool call "
"using the requested command (host_shell), then summarize. "
"Do not stop after the edit."
"for verification. "
+ (
"Read the saved output artifact now with read_file, then summarize. "
if _artifact_readback_requested
else "Do that now with one concrete tool call using the requested command (host_shell), then summarize. "
)
+ "Do not stop after the edit."
),
})
yield f'data: {json.dumps({"type": "agent_step", "round": round_num + 1})}\n\n'
+4 -4
View File
@@ -695,7 +695,7 @@ async def do_ui_control(content: str, session_id: Optional[str] = None, owner: O
toggle <name> <on|off> — Toggle a setting (web, bash, rag, research, incognito, document_editor)
set_mode <agent|chat> — Switch between agent and chat mode
switch_model <model> — Change the model for the current session
set_theme <preset> — Apply a built-in theme preset (dark, light, midnight, paper, cyberpunk, retrowave, forest, ocean, ume, copper, terminal, organs, lavender, gpt, claude, cute)
set_theme <preset> — Apply a built-in theme preset (dark, light, midnight, cyberpunk, retrowave, forest, ocean, ume, terminal, organs, gpt, claude, cute, eclipse, porcelain, arcade, blueprint, monolith, yoyo)
create_theme <name> <bg> <fg> <panel> <border> <accent> [key=val ...] — Create custom theme. Optional key=val: advanced color overrides AND background effects: bgPattern=<none|dots|synapse|rain|constellations|perlin-flow|petals|sparkles|embers>, bgEffectColor=#RRGGBB, bgEffectIntensity=<num>, bgEffectSize=<num>, frosted=true|false
get_theme — Return the last server-synchronized theme for this user
open_panel <name> [view] — Open a panel; Cookbook views are download/models, launch/serve, active/running, dependencies, settings
@@ -798,9 +798,9 @@ async def do_ui_control(content: str, session_id: Optional[str] = None, owner: O
# Also check user's custom themes stored in prefs.
# Must match the THEMES keys in static/js/theme.js.
known_presets = [
"dark", "light", "midnight", "paper", "cyberpunk", "retrowave",
"forest", "ocean", "ume", "copper", "terminal", "organs",
"lavender", "gpt", "claude", "cute",
"dark", "light", "midnight", "cyberpunk", "retrowave", "forest",
"ocean", "ume", "terminal", "organs", "gpt", "claude", "cute",
"eclipse", "porcelain", "arcade", "blueprint", "monolith", "yoyo",
]
custom_themes = {}
try:
+3
View File
@@ -50,6 +50,9 @@ _VISION_MODEL_KEYWORDS = (
# Qwen3.5 is a natively multimodal family even when a served-model alias
# omits the traditional "VL" suffix (for example qwen35-9b-base-native).
"qwen3.5", "qwen3_5", "qwen35",
# The hosted Flash alias accepts images despite lacking a vision/VL suffix.
# Keep this exact: deepseek-v4-pro on the same provider is text-only.
"deepseek-flash",
# multimodal families whose names don't contain "vision"/"vl" but DO accept
# images — without these the image is silently dropped for common Ollama tags
# like gemma3:4b or gemma4:12b (issue #1274). Gemma 3/4 (4b+), Llama 4 (all),
+43 -6
View File
@@ -55,6 +55,9 @@ NATIVE_ROUND_LIMIT = 64
# and interaction are separate observable actions.
INTERACTIVE_TOOL_CALL_LIMIT = 18
INTERACTIVE_BROWSER_TOOL_CALL_LIMIT = 30
# "Unlimited" as a comparable int: every call site tests `calls < limit`, so a
# sentinel avoids threading an Optional through the whole preview loop.
UNLIMITED_TOOL_CALL_LIMIT = 1_000_000
INTERACTIVE_ROUND_LIMIT = 8
# Multi-record research tasks routinely need several search/fetch/inspection
# pairs before an artifact can be grounded. Preserve twelve calls for writing,
@@ -1468,15 +1471,41 @@ def standalone_social_turn(text):
def interactive_execution_limit(max_rounds):
"""Bound interactive turns independently of long-running native jobs."""
"""Honor the WebUI agent-step setting for the compact preview loop.
A configured finite budget is the user's explicit instruction and is
honored up to the same 200 ceiling the settings endpoint enforces.
INTERACTIVE_ROUND_LIMIT remains the fallback when no budget is resolvable
(adaptive ``None`` mode or a malformed value), so a turn still terminates.
"""
if max_rounds is None:
return INTERACTIVE_ROUND_LIMIT
try:
return max(1, min(int(max_rounds), INTERACTIVE_ROUND_LIMIT))
return max(1, min(int(max_rounds), 200))
except (TypeError, ValueError):
return INTERACTIVE_ROUND_LIMIT
def interactive_tool_call_limit(max_tool_calls, *, browser_offered=False):
"""Honor the configured agent tool-call budget; 0 means unlimited.
Matches the main agent loop, which treats ``max_tool_calls <= 0`` as
unbounded. The INTERACTIVE_* constants remain the fallback for a
malformed value.
"""
default = (
INTERACTIVE_BROWSER_TOOL_CALL_LIMIT if browser_offered
else INTERACTIVE_TOOL_CALL_LIMIT
)
try:
budget = int(max_tool_calls)
except (TypeError, ValueError):
return default
if budget <= 0:
return UNLIMITED_TOOL_CALL_LIMIT
return budget
def runtime_required_artifacts(user_text, client_runtime_context):
"""Use runner-declared outputs, falling back to prompt inference."""
context = client_runtime_context if isinstance(client_runtime_context, dict) else {}
@@ -4314,6 +4343,7 @@ async def stream_preview(*, endpoint_url, model, messages, headers, turn_contrac
history_session=None, external_untrusted_context_seen=False,
active_document=None, active_email=None, workspace=None,
client_runtime_context=None, max_tokens=768, max_rounds=8,
max_tool_calls=0,
external_tool_schemas=None, temperature=0.0,
**ignored):
from src.generation_sampling import validate_temperature
@@ -4577,10 +4607,12 @@ async def stream_preview(*, endpoint_url, model, messages, headers, turn_contrac
'Open the document you want reviewed, then ask for inline suggestions again.'
)
round_limit = interactive_execution_limit(max_rounds)
tool_call_limit = (
INTERACTIVE_BROWSER_TOOL_CALL_LIMIT
if any(canonical(schema['function']['name']) == 'private_browser' for schema in offered)
else INTERACTIVE_TOOL_CALL_LIMIT
tool_call_limit = interactive_tool_call_limit(
max_tool_calls,
browser_offered=any(
canonical(schema['function']['name']) == 'private_browser'
for schema in offered
),
)
if native_workspace_enabled:
try:
@@ -4929,6 +4961,11 @@ async def stream_preview(*, endpoint_url, model, messages, headers, turn_contrac
if content and not prior_summary_answer:
yield event({'delta': content})
proposed = [pending[i] for i in sorted(pending)]
# A lead-in emitted before a tool call is live progress, not
# part of the terminal answer. Replace that draft when the
# eventual synthesis begins instead of concatenating both.
if proposed and streamed_round_text:
replace_streamed_draft_on_finish = True
unexecutable_dsml_completion = False
if not proposed and 'DSML' in content:
offered_by_canonical = {
+1
View File
@@ -6,6 +6,7 @@ import subprocess
from src.runtime_paths import get_app_root, get_default_data_dir
APP_VERSION = "1.0.3"
BUILTIN_SKILLS_DIR = os.path.join(get_app_root(), "resources", "skills")
# Identifies the private maintainer-preview build without changing the public
# application semver used by release and readiness checks. Keep the API/UI
# value tied to HARNESS_VERSION so a version bump cannot leave the running
+15
View File
@@ -178,6 +178,21 @@ class FastEmbedClient:
except Exception as _e:
logger.debug("embedding cache symlink-heal skipped: %s", _e)
kwargs = {"model_name": self.model, "cache_dir": cache_dir}
# Isolated evaluation and worker fleets can run many Odysseus
# processes on one host. FastEmbed otherwise lets ONNX Runtime size
# a thread pool from the whole machine for every process, which can
# create hundreds of threads per worker and starve inference. Keep
# the existing default for normal installs, but allow operators to
# bound that pool explicitly.
raw_threads = os.getenv("FASTEMBED_THREADS", "").strip()
if raw_threads:
try:
threads = int(raw_threads)
except ValueError as exc:
raise ValueError("FASTEMBED_THREADS must be an integer") from exc
if not 1 <= threads <= 256:
raise ValueError("FASTEMBED_THREADS must be between 1 and 256")
kwargs["threads"] = threads
self._embedding = TextEmbedding(**kwargs)
self._dim: Optional[int] = None
self.url = "local://fastembed"
+1 -1
View File
@@ -737,7 +737,7 @@ _SENSITIVE_BASENAMES: set[str] = {
_SENSITIVE_FILE_PATTERNS: tuple[str, ...] = (
"authorized_keys", "id_rsa", "id_ed25519", "id_ecdsa",
"known_hosts",
"known_hosts", "auth.json", "app.db", "settings.json",
)
# Case-folded views used for matching. On a case-insensitive filesystem
+7 -1
View File
@@ -76,7 +76,13 @@ def select_experiment_inventory(inventory, routed, history, mode, *, user_text='
and routed.required_read_operation is None
and any(canonical_tool(name) == 'edit_image' for name in routed.required)
)
if not explicit_image_edit:
if not explicit_image_edit and not (
mode == MODEL_CHOICE_MODE and families
and routed.required_read_operation is None
):
# A concrete current-turn route owns the inventory. Reintroducing
# previously used families here lets an explicit domain switch retain
# stale authority and lets a resolved follow-up drift into shell.
families.update(recently_executed_families(
history, user_turns=6, maximum=3,
include_failed_attempts=mode == MODEL_CHOICE_MODE,
+1 -1
View File
@@ -895,7 +895,7 @@ FUNCTION_TOOL_SCHEMAS = [
"type": "function",
"function": {
"name": "ui_control",
"description": "Control the user interface. Actions: toggle (turn tools on/off), open_panel (open a modal: documents/library, gallery, calendar/schedule, email, sessions, notes, memories/brain, skills, settings, theme, cookbook; calendar supports month/week/year/agenda plus a date; Cookbook supports models/download, launch/serve, active/running, dependencies, and settings views), open_email_reply (legacy UI-only reply opener; prefer email MCP draft_email_reply for assistant-written reply drafts so a normal document-backed email draft is created), set_mode, switch_model, set_theme (built-in presets: dark, light, midnight, paper, cyberpunk, retrowave, forest, ocean, ume, copper, terminal, organs, lavender, gpt, claude, cute), create_theme (CREATE any custom theme with a name + colors object — pick distinctive, evocative hex colors that match the requested aesthetic, NOT generic defaults. The theme auto-applies after creation), get_theme, and get_toggles. When a user asks for ANY theme not in the built-in preset list, ALWAYS use create_theme.",
"description": "Control the user interface. Actions: toggle (turn tools on/off), open_panel (open a modal: documents/library, gallery, calendar/schedule, email, sessions, notes, memories/brain, skills, settings, theme, cookbook; calendar supports month/week/year/agenda plus a date; Cookbook supports models/download, launch/serve, active/running, dependencies, and settings views), open_email_reply (legacy UI-only reply opener; prefer email MCP draft_email_reply for assistant-written reply drafts so a normal document-backed email draft is created), set_mode, switch_model, set_theme (built-in presets: dark, light, midnight, cyberpunk, retrowave, forest, ocean, ume, terminal, organs, gpt, claude, cute, eclipse, porcelain, arcade, blueprint, monolith, yoyo), create_theme (CREATE any custom theme with a name + colors object — pick distinctive, evocative hex colors that match the requested aesthetic, NOT generic defaults. The theme auto-applies after creation), get_theme, and get_toggles. When a user asks for ANY theme not in the built-in preset list, ALWAYS use create_theme.",
"parameters": {
"type": "object",
"properties": {
+147 -1
View File
@@ -4078,6 +4078,82 @@ def recently_executed_families(history: Iterable, *, user_turns: int = 6,
return tuple(found)
_CONTINUITY_STOP_WORDS = frozenset({
"a", "about", "an", "and", "are", "at", "be", "but", "can", "could",
"did", "do", "does", "for", "from", "get", "have", "how", "i", "in",
"is", "it", "look", "me", "my", "not", "of", "on", "or", "please",
"search", "searched", "searching", "see", "show", "that", "the", "them",
"there", "these", "this", "those", "to", "u", "was", "what", "when",
"where", "which", "why", "with", "you", "your", "whats", "what's",
"cant", "can't", "cannot", "dont", "don't", "doesnt", "doesn't",
})
def _subject_tokens(value: object) -> frozenset[str]:
"""Return content-bearing tokens for conversation-subject continuity."""
return frozenset(
token for token in re.findall(r"[\w'-]+", str(value or "").casefold())
if len(token) > 2 and token not in _CONTINUITY_STOP_WORDS
)
def _immediate_prior_user_subject_tokens(history: Iterable) -> frozenset[str]:
rows = tuple(history or ())
seen_assistant = False
for row in reversed(rows):
role = row.get("role") if isinstance(row, dict) else getattr(row, "role", "")
if role == "assistant" and not seen_assistant:
seen_assistant = True
continue
if seen_assistant and role == "user":
content = row.get("content", "") if isinstance(row, dict) else getattr(row, "content", "")
return _subject_tokens(content)
return frozenset()
def immediately_established_family(message: str, history: Iterable) -> str | None:
"""Resolve an elliptical follow-up against the immediately proven domain.
Tool events provide the typed domain; subject-token overlap only determines
whether the new sentence continues that turn. This deliberately does not
infer authority from older turns or from model prose.
"""
rows = tuple(history or ())
assistant_index = None
families: set[str] = set()
for index in range(len(rows) - 1, -1, -1):
row = rows[index]
role = row.get("role") if isinstance(row, dict) else getattr(row, "role", "")
if role != "assistant":
continue
assistant_index = index
metadata = row.get("metadata") if isinstance(row, dict) else getattr(row, "metadata", None)
if isinstance(metadata, str):
try:
metadata = json.loads(metadata)
except (TypeError, json.JSONDecodeError):
metadata = {}
for event in (metadata or {}).get("tool_events") or ():
if event.get("error") is True or event.get("exit_code") not in (None, 0):
continue
families.update(_families_for_tool(canonical_tool(event.get("tool", ""))))
break
if assistant_index is None or len(families) != 1:
return None
prior_user_text = ""
for row in reversed(rows[:assistant_index]):
role = row.get("role") if isinstance(row, dict) else getattr(row, "role", "")
if role == "user":
prior_user_text = row.get("content", "") if isinstance(row, dict) else getattr(row, "content", "")
break
if not prior_user_text:
return None
if _subject_tokens(message) & _subject_tokens(prior_user_text):
return next(iter(families))
return None
def recently_read_gallery(history: Iterable, *, user_turns: int = 6) -> bool:
"""Whether a recent successful app_api call established gallery context."""
turns = 0
@@ -4140,6 +4216,58 @@ def recently_read_gallery(history: Iterable, *, user_turns: int = 4) -> bool:
return False
# Personal-data product nouns. A broad-briefing phrase ("what's new",
# "give me an update", "news") must not out-rank these: the user is asking
# about their own store, not the open Web. Scoped to a first-person
# possessive so open-web subjects that merely borrow a product noun
# ("the latest events in Kyiv") keep their Web route.
_PERSONAL_STORE_NOUNS = (
r"(?:e?mails?|inbox|mailbox|calendar|calender|events?|appointments?|"
r"meetings?|agenda|notes?|checklists?|tasks?|todos?|documents?|docs?|"
r"memor(?:y|ies)|contacts?|skills?|sessions?|chats?|conversations?)"
)
_PERSONAL_STORE_SUBJECT = re.compile(
rf"\b(?:my|our)\b(?:\s+\w+){{0,2}}\s+{_PERSONAL_STORE_NOUNS}\b|"
rf"\b(?:inbox|mailbox)\b",
re.I,
)
_PERSONAL_STORE_FAMILY = (
(("email", "emails", "mail", "mails", "inbox", "mailbox"), "email"),
(("calendar", "calender", "event", "events", "appointment", "appointments",
"meeting", "meetings", "agenda"), "calendar"),
(("note", "notes", "checklist", "checklists"), "notes"),
(("task", "tasks", "todo", "todos"), "tasks"),
(("document", "documents", "doc", "docs"), "documents"),
(("memory", "memories"), "memory"),
(("contact", "contacts"), "contacts"),
(("skill", "skills"), "skills"),
(("session", "sessions", "chat", "chats", "conversation", "conversations"),
"sessions"),
)
def names_personal_store(message: str) -> bool:
"""True when the request names the user's own data store."""
return bool(_PERSONAL_STORE_SUBJECT.search(str(message or "")))
def personal_store_families(message: str) -> frozenset[str]:
"""Families for the user's own stores named in a broad-briefing request.
A briefing phrase must resolve to the named store rather than falling
through to an empty inventory, which would offer no tools at all.
"""
families: set[str] = set()
for match in _PERSONAL_STORE_SUBJECT.finditer(str(message or "")):
matched = match.group(0).lower()
for nouns, family in _PERSONAL_STORE_FAMILY:
if any(re.search(rf"\b{noun}\b", matched) for noun in nouns):
families.add(family)
return frozenset(families)
def broad_web_briefing_request(message: str) -> bool:
"""Recognize requests that need broad, current, multi-source Web evidence."""
text = _normalize_request_lead(message)
@@ -4180,6 +4308,18 @@ def requested_capabilities(message: str, history: Iterable = (), *, active_docum
if lead := _CONVERSATIONAL_ACTION_LEAD.fullmatch(text):
text = lead["request"].strip()
history = tuple(history)
repeated_subject = _subject_tokens(text) & _immediate_prior_user_subject_tokens(history)
scope_text = " ".join(
token for token in re.findall(r"[\w'-]+", text)
if token.casefold() not in repeated_subject
)
newly_named_families = {
family for family, pattern in _FAMILY_WORDS.items()
if re.search(pattern, scope_text, re.I)
}
established_family = immediately_established_family(text, history)
if established_family and not newly_named_families:
return frozenset({established_family})
concrete_urls = re.findall(r"\bhttps?://[^\s<>\"']+", raw_text, re.I)
workspace_media = re.search(
r"(?:file://)?/workspace/[^\s`\"']+\."
@@ -4250,6 +4390,9 @@ def requested_capabilities(message: str, history: Iterable = (), *, active_docum
broad_web_briefing_request(text)
and not re.search(r"\b(?:research|investigate|deep[ -]?dive)\b", text, re.I)
):
_personal = personal_store_families(text)
if _personal:
return _personal
return frozenset({"search_browser"})
if re.search(r"\b(?:web_search|web_fetch)\b", raw_text, re.I):
# Explicit native-tool requests are stronger than incidental domain
@@ -4267,7 +4410,10 @@ def requested_capabilities(message: str, history: Iterable = (), *, active_docum
if (
re.search(r"\b(?:latest|recent|current|today(?:'s)?)\b", text, re.I)
and re.search(r"\b(?:info(?:rmation)?|news|nees|updates?)\b", text, re.I)
and not names_personal_store(text)
):
# A named personal store out-ranks the broad-briefing route; the
# guard above lets those fall through to the family grammar.
# Broad current-information requests still require live Web evidence.
# Keep the common ``nees`` typo because a missed route leaves the model
# with no way to answer and encourages it to ask unnecessary questions.
@@ -5792,7 +5938,7 @@ def resolve_turn_contract(*, capabilities: Iterable[str], schemas: Iterable[dict
if (
message is not None
and selected_tools is not None
and set(selected_tools) & {"web_search", "web_fetch"}
and selected & {"web_search", "web_fetch"}
):
# Browser is not core. It is a bounded recovery capability for a web
# turn when static search/fetch cannot read the named site.