Merge commit 'refs/phase3/pre-ajax/publication-tip' into integration/pre-ajax-release

# Conflicts:
#	routes/chat_routes.py
#	routes/session_routes.py
#	src/agent_loop.py
#	src/agent_tools/filesystem_tools.py
#	src/teacher_escalation.py
#	src/tool_capabilities.py
#	src/tool_execution.py
#	tests/test_mcp_add_server_args_validation.py
#	tests/test_token_cache_atomic_swap.py
This commit is contained in:
Alexandre Teixeira
2026-10-05 15:59:59 +01:00
1395 changed files with 360455 additions and 105938 deletions
BIN
View File
Binary file not shown.
BIN
View File
Binary file not shown.
Binary file not shown.
+281
View File
@@ -0,0 +1,281 @@
---
layout: default
---
# Configuration reference: ODYSSEUS_* environment variables
<!-- This page is generated from the source tree by `scripts/generate_env_reference.py`. Do not edit it by hand: run the script instead. `tests/test_env_reference.py` fails when the committed page and the source disagree, or when a variable is read without an entry in the script's notes table. -->
Odysseus reads its runtime configuration from the Settings UI. The
environment variables below are the deployment-level escape hatches underneath
that: they are read directly from the process environment, mostly at import or
startup, and most installs never need any of them.
`.env.example` stays a short, deployment-level example on purpose - `APP_BIND`,
`APP_PORT`, `AUTH_ENABLED`, `DATABASE_URL` and a pre-seeded admin password. This
page is the complete list, which is a different job.
Truthiness is not uniform across the codebase. Where a variable is described as
"truthy" the read accepts `1`, `true`, `yes` and usually `on`; where it is
described as a switch that turns something off, the read rejects `0`, `false`,
`no` and `off` and treats everything else as on. The `Default` column is the
value the code falls back to when the variable is unset, quoted from the source.
The source tree reads **112** `ODYSSEUS_*` variables: 81 an operator may want to set, and 31 that are internal - sentinels, fixture switches, capture hooks and development tooling. The internal ones are listed too, in their own section, so this page can be checked against the source mechanically.
> This page is generated. Edit `scripts/generate_env_reference.py` and
> re-run it; `tests/test_env_reference.py` enforces that the committed page
> matches the source.
## Variables you can set
### Deployment and first run
| Variable | Default | Read in | What it does |
|---|---|---|---|
| `ODYSSEUS_ADMIN_PASSWORD` | `''` | `setup.py:102` (+4 more) | Password for the admin account created on first run. Setup refuses a value shorter than its minimum length rather than silently falling back. |
| `ODYSSEUS_ADMIN_USER` | `''` | `setup.py:101` (+2 more) | Username for the admin account created on first run. Setup uses env vars first, then an interactive prompt, then a random password. |
| `ODYSSEUS_ALLOW_OLLAMA_CLI_SCAN` | *unset* | `routes/cookbook_helpers.py:544` (+1 more) | On Windows only, set truthy to let the Cookbook dependency probe shell out to `ollama list`. Ignored on other platforms, where the scan always runs. |
| `ODYSSEUS_CONTAINER_NETWORK_MODE` | `''` | `app.py:1022` (+1 more) | Declares the container's Docker network mode. Set to `host` to skip host-gateway probing when discovering local model endpoints. |
| `ODYSSEUS_ENABLE_HOST_DOCKER` | `''` | `src/host_docker_access.py:41` | Security-relevant. Must be exactly `true` before tools may use a mounted host Docker socket, and the socket itself must exist. |
| `ODYSSEUS_INPROCESS_POLLERS` | `'1'` | `routes/email/email_pollers.py:1722` | The same off switch for the in-process email pollers, when `odysseus-mail poll-scheduled` is the sole external driver. |
| `ODYSSEUS_INPROCESS_TASKS` | `'1'` | `app.py:1311` | Set to 0, false, no or off to stop the in-process scheduled-task runner, for deployments where an external worker drives task firing. |
| `ODYSSEUS_MODEL_KEEPALIVE` | `''` | `app.py:1218` | Opt-in periodic model keep-alive pings. Off by default: the ping path runs model discovery, so stale LAN endpoints add background pressure. |
| `ODYSSEUS_REQUIRE_TOOL_INDEX_READY` | `''` | `src/readiness.py:61` | Set truthy to make semantic tool-index readiness gate startup. Off by default so an install stays available on deterministic tool selection. |
| `ODYSSEUS_SKIP_ADMIN_PROMPT` | *unset* | `setup.py:112` | Any non-empty value suppresses the interactive admin-credential prompt even on a TTY, for unattended installs. |
| `ODYSSEUS_SLOW_REQUEST_LOG_SECONDS` | `'0.75'` | `app.py:239` | Request duration in seconds above which the middleware logs a slow-request warning. |
| `ODYSSEUS_STARTUP_WARMUPS` | `''` | `app.py:1192` | Opt-in startup pings of the configured model endpoints. Off by default because they compete with the first seconds of UI use. |
| `ODYSSEUS_TOOL_INDEX_PREWARM` | `'1'` | `src/tool_index.py:772` | Set to 0, false, no or off to skip background initialization of semantic tool retrieval at startup. |
### Data directories and paths
| Variable | Default | Read in | What it does |
|---|---|---|---|
| `ODYSSEUS_DATA_DIR` | `get_default_data_dir()` | `src/constants.py:56` (+1 more) | Root directory for every persisted file. Prefer this over the per-path overrides; the rest of `src/constants.py` derives from it. |
| `ODYSSEUS_MAIL_ATTACHMENTS_DIR` | `os.path.join(DATA_DIR, 'mail-attachments')` | `src/constants.py:105` | Dedicated override for the mail attachment store, which otherwise lives under the data directory. |
### Model routing and providers
| Variable | Default | Read in | What it does |
|---|---|---|---|
| `ODYSSEUS_COPILOT_API_VERSION` | `'2026-06-01'` | `src/copilot.py:39` | Dated API-version header the Copilot models and chat endpoints require. |
| `ODYSSEUS_COPILOT_CLIENT_ID` | `'01ab8ac9400c4e429b23'` | `src/copilot.py:34` | GitHub OAuth client id for the Copilot device flow. The default is the public VS Code client id; override it only with your own allow-listed app. |
| `ODYSSEUS_DEEPSEEK_REASONING_EFFORT` | `'high'` | `src/llm_core.py:1727` | Reasoning effort for DeepSeek. Only `high` and `max` are accepted; any other value falls back to the default. |
| `ODYSSEUS_FIRST_TOKEN_TIMEOUT` | `''` | `src/llm_core.py:223` | Seconds to wait for the first streamed token from a local endpoint before failing. Unset keeps the generous read timeout, which makes a stalled backend look like a hung agent. |
| `ODYSSEUS_LOCAL_MODEL_GATE` | `'true'` | `src/llm_core.py:95` | On by default. Set 0, false, no or off to drop the gate that checks a local endpoint before routing a request to it. |
| `ODYSSEUS_MISTRAL_REASONING_EFFORT` | `'high'` | `src/llm_core.py:1723` | Reasoning effort sent to Mistral thinking-capable models. The API accepts high, medium, low and none. |
| `ODYSSEUS_MLX_IMAGE_VLM_MODEL` | *unset* | `scripts/mlx_image_server.py:299` | Vision-language model id for the MLX image server script. Required unless `--vlm-model` is passed on the command line. |
| `ODYSSEUS_QWEN_ROUTE_THINKING` | `'auto'` | `src/agent_loop.py:170` | Thinking policy for the Qwen routing step. An unrecognized value falls back to `auto`. |
### Agent loop and tool execution
| Variable | Default | Read in | What it does |
|---|---|---|---|
| `ODYSSEUS_DISABLE_MCP` | `''` | `src/builtin_mcp.py:89` | Truthy disables MCP entirely, as an escape hatch for compatibility problems with a server. |
| `ODYSSEUS_MAX_VISUAL_EVIDENCE_FRAMES` | `'3'` | `src/agent_loop.py:15366` | How many video frames one tool result may contribute. Clamped to 1-8. |
| `ODYSSEUS_MAX_VISUAL_EVIDENCE_IMAGES` | `'1'` | `src/agent_loop.py:15334` | How many images one tool result may contribute to the model turn. Clamped to 1-8. |
| `ODYSSEUS_MCP_ALLOWED_COMMANDS` | `''` | `src/agent_tools/admin_tools.py:140` | Security-relevant. Comma-separated allowlist of MCP launcher basenames the agent may start. Empty by default, and the deny list still wins. |
| `ODYSSEUS_PYTHON_TOOL_SITE_PACKAGES` | `''` | `src/agent_runtime/process_resources.py:59` (+2 more) | Security-relevant. Absolute package roots, separated by the platform path separator, exposed to the sandboxed Python tool. Empty exposes none. |
| `ODYSSEUS_SCRIPT_HOST` | `'localhost'` | `src/builtin_actions.py:925` | Default host for the run-script action. `localhost`, `127.0.0.1`, `local` and empty run locally; any other value runs over SSH. |
| `ODYSSEUS_TOOL_APPROVAL_GATE` | `'0'` | `src/tool_capabilities.py:645` | Security-relevant. Truthy makes tool calls pass through the approval gate. Off by default. |
### Browser automation
| Variable | Default | Read in | What it does |
|---|---|---|---|
| `ODYSSEUS_BROWSER_EXECUTABLE` | `''` | `src/builtin_mcp.py:114` | Absolute path to the Chrome or Chromium binary. Empty searches the usual names, then lets Playwright MCP pick its own browser. |
| `ODYSSEUS_BROWSER_ISOLATED` | `'1'` | `src/builtin_mcp.py:139` | Security-relevant. On by default, adding `--isolated` so each browser session starts clean. Set 0, false or no to keep a persistent profile. |
| `ODYSSEUS_BROWSER_MCP_CACHE` | `os.path.join(base_dir, 'data', 'local', 'playwright-mcp-cache')` | `src/builtin_mcp.py:229` | Cache directory handed to the browser MCP server, so its npm download survives a container rebuild. |
| `ODYSSEUS_BROWSER_MCP_CALL_TIMEOUT_S` | `'90'` | `src/mcp_manager.py:27` | Upper bound in seconds for one browser MCP tool call. A call that exceeds it fails without being retried. |
| `ODYSSEUS_BROWSER_MCP_REQUIRE_CACHE` | `''` | `src/builtin_mcp.py:90` | Truthy refuses to start the browser MCP server unless its npm package is already in the npx cache, instead of installing it at startup. |
| `ODYSSEUS_BROWSER_NAMESPACE` | `'odysseus-ui'` | `src/agent_tools/web_tools.py:100` (+1 more) | Namespace for the detached agent-browser daemon's pid files, so two runtimes on one machine do not terminate each other's browsers. |
| `ODYSSEUS_BROWSER_NO_SANDBOX` | `'1'` | `src/builtin_mcp.py:142` | Security-relevant. On by default, adding `--no-sandbox` because the Docker image cannot use the Chromium sandbox. Set 0, false or no to keep it. |
| `ODYSSEUS_BROWSER_SCREENSHOT_DIR` | *unset* | `src/agent_tools/web_tools.py:2666` | Where private-browser screenshots are written. Falls back to the container path, then the system temp directory. |
### Container and workspace mounts
| Variable | Default | Read in | What it does |
|---|---|---|---|
| `ODYSSEUS_WORKSPACE_CONTAINER_ROOT` | *unset* | `src/workspace_paths.py:31` | Container-side root that the host root maps onto. An `or` fallback, not a read default, supplies `/workspace` when it is unset. |
| `ODYSSEUS_WORKSPACE_DEFAULT` | `''` | `routes/workspace_routes.py:97` | Default workspace path the admin-only workspace route reports. Empty means no default is configured. |
| `ODYSSEUS_WORKSPACE_HOST_ROOT` | *unset* | `src/workspace_paths.py:29` | Single host-side root, paired with the container root below. Simpler than the explicit mount list when there is only one mount. |
| `ODYSSEUS_WORKSPACE_MOUNTS` | `''` | `src/workspace_paths.py:19` | `host=container` path pairs separated by commas or semicolons, so the agent can translate a container path back to the host path a user typed. |
### Email
| Variable | Default | Read in | What it does |
|---|---|---|---|
| `ODYSSEUS_DOCUMENT_OWNER` | `''` | `mcp_servers/email_server.py:208` | Owner stamped on documents the email MCP server creates. Stdio MCP tools get no authenticated user, so without this a draft is invisible. |
| `ODYSSEUS_IMAP_TIMEOUT_SECONDS` | *unset* | `routes/email/email_helpers.py:1163` | IMAP socket timeout in seconds, clamped to 5-300. A non-numeric value falls back to 30 rather than failing. |
### Calendar, notes and single-user mode
| Variable | Default | Read in | What it does |
|---|---|---|---|
| `ODYSSEUS_ALLOW_PRIVATE_CALDAV` | `'0'` | `src/caldav_sync.py:52` | Security-relevant. Truthy lets CalDAV sync reach private and link-local addresses. Off by default; this is an SSRF guard. |
| `ODYSSEUS_FALLBACK_OWNER` | `'owner@localhost'` | `routes/calendar_routes.py:65` | Owner address that single-user mode attributes an unauthenticated request to. Only reachable while single-user mode is on. |
| `ODYSSEUS_SINGLE_USER` | `'1'` | `routes/calendar_routes.py:66` | Security-relevant. On by default. Set to 0 on a real multi-user install so unauthenticated calendar writes are rejected rather than absorbed. |
### Upload and media limits
| Variable | Default | Read in | What it does |
|---|---|---|---|
| `ODYSSEUS_CHAT_UPLOAD_MAX_BYTES` | `10 * 1024 * 1024` | `src/upload_limits.py:33` | Maximum bytes accepted for a chat attachment. |
| `ODYSSEUS_EDITOR_DRAFT_MAX_BYTES` | `256 * 1024 * 1024` | `src/upload_limits.py:47` | Maximum bytes accepted for a saved editor draft. |
| `ODYSSEUS_EMAIL_COMPOSE_UPLOAD_MAX_BYTES` | `25 * 1024 * 1024` | `src/upload_limits.py:56` | Maximum bytes accepted for an attachment added while composing mail. |
| `ODYSSEUS_GALLERY_TRANSFORM_UPLOAD_MAX_BYTES` | `25 * 1024 * 1024` | `src/upload_limits.py:44` (+1 more) | Maximum bytes accepted for an image handed to a gallery transform. |
| `ODYSSEUS_GALLERY_UPLOAD_MAX_BYTES` | `100 * 1024 * 1024` | `src/upload_limits.py:41` (+1 more) | Maximum bytes accepted for a gallery upload. |
| `ODYSSEUS_ICS_MAX_BYTES` | `10 * 1024 * 1024` | `src/upload_limits.py:62` | Maximum bytes accepted for an imported ICS file. |
| `ODYSSEUS_MEDIA_FRAME_TIMEOUT` | `30` | `src/media_ingress.py:146` | Seconds allowed for extracting frames from a video before giving up. |
| `ODYSSEUS_MEDIA_MAX_AUDIO_BYTES` | `32 * 1024 * 1024` | `src/media_ingress.py:129` | Largest source audio file the media pipeline will read. |
| `ODYSSEUS_MEDIA_MAX_DIMENSION` | `1600` | `src/media_ingress.py:138` | Longest edge in pixels an image is resized down to before encoding. |
| `ODYSSEUS_MEDIA_MAX_DOCUMENT_BYTES` | `32 * 1024 * 1024` | `src/media_ingress.py:126` | Largest source document the media pipeline will read. |
| `ODYSSEUS_MEDIA_MAX_DOCUMENT_CHARS` | `24000` | `src/media_ingress.py:135` | How many characters of an ingested document are inlined into the turn. |
| `ODYSSEUS_MEDIA_MAX_ENCODED_BYTES` | `24 * 1024 * 1024` | `src/media_ingress.py:132` | Budget for the encoded payload handed to the model, counted cumulatively across one turn's attachments rather than per file. |
| `ODYSSEUS_MEDIA_MAX_FILES` | `4` | `src/media_ingress.py:119` | How many local media files one agent turn may ingest. |
| `ODYSSEUS_MEDIA_MAX_IMAGE_BYTES` | `12 * 1024 * 1024` | `src/media_ingress.py:120` | Largest source image the media pipeline will read. |
| `ODYSSEUS_MEDIA_MAX_PIXELS` | `40000000` | `src/media_ingress.py:139` | Total pixel budget for a source image, as a decompression-bomb guard. |
| `ODYSSEUS_MEDIA_MAX_VIDEO_BYTES` | `128 * 1024 * 1024` | `src/media_ingress.py:123` | Largest source video the media pipeline will read. |
| `ODYSSEUS_MEDIA_MAX_VIDEO_FRAMES` | `8` | `src/media_ingress.py:140` | How many frames are sampled from a video. |
| `ODYSSEUS_MEDIA_PROBE_TIMEOUT` | `15` | `src/media_ingress.py:143` | Seconds allowed for probing a video's metadata before giving up. |
| `ODYSSEUS_MEMORY_IMPORT_MAX_BYTES` | `10 * 1024 * 1024` | `src/upload_limits.py:50` (+1 more) | Maximum bytes accepted for a memory import file. |
| `ODYSSEUS_PERSONAL_UPLOAD_MAX_BYTES` | `25 * 1024 * 1024` | `src/upload_limits.py:53` (+1 more) | Maximum bytes accepted for a personal-documents upload. |
| `ODYSSEUS_STT_MAX_AUDIO_BYTES` | `25 * 1024 * 1024` | `src/upload_limits.py:59` | Maximum bytes accepted for an audio file submitted for transcription. |
### Search
| Variable | Default | Read in | What it does |
|---|---|---|---|
| `ODYSSEUS_SEARCH_PROVIDER` | `''` | `services/search/providers.py:43` | Forces the search provider, overriding the Settings value. Empty keeps the UI authoritative, which is what a normal install wants. |
### Memory and skills
| Variable | Default | Read in | What it does |
|---|---|---|---|
| `ODYSSEUS_MCP_MEMORY_OWNER` | *unset* | `src/mcp_manager.py:190` | Application owner binding for the configured memory MCP backend. Takes precedence over ODYSSEUS_MEMORY_OWNER; missing ownership fails closed. |
| `ODYSSEUS_MEMORY_OWNER` | *unset* | `src/mcp_manager.py:190` | Fallback application owner binding for the memory MCP backend. This configuration identifies ownership; it does not grant read or egress authority. |
| `ODYSSEUS_SKILL_SEMANTIC_RETRIEVAL` | `'1'` | `services/memory/skills.py:796` | On by default. Set 0, false, no or off to fall back to keyword-only skill retrieval when no vector store is reachable. |
| `ODYSSEUS_SKILL_SEMANTIC_THRESHOLD` | `'0.4'` | `services/memory/skills.py:807` | Minimum semantic score a skill needs to be retrieved. A non-numeric value falls back to the default. |
### Speech and vision models
| Variable | Default | Read in | What it does |
|---|---|---|---|
| `ODYSSEUS_GROUNDING_MODEL` | `'google/owlvit-base-patch32'` | `routes/gallery/gallery_routes.py:96` | Object-grounding model id the gallery loads for text-driven selection. |
| `ODYSSEUS_SAM_MODEL` | `'facebook/sam-vit-base'` | `routes/gallery/gallery_routes.py:60` | Segmentation model id the gallery loads for subject selection. |
| `ODYSSEUS_STT_MODEL` | *unset* | `src/agent_tools/media_tools.py:2189` | Default speech-to-text model for media transcription when the tool call does not name one. |
| `ODYSSEUS_TTS_CACHE_MAX_BYTES` | `500 * 1024 * 1024` | `services/tts/tts_service.py:47` | Cap on the synthesized-speech cache. A non-numeric value falls back to the default. |
### Auth and internal API
| Variable | Default | Read in | What it does |
|---|---|---|---|
| `ODYSSEUS_INTERNAL_BASE` | *unset* | `src/constants.py:192` | Base URL the in-app tool layer uses for loopback HTTP calls. Set it when the app is not reachable at the port it thinks it is bound to. |
| `ODYSSEUS_INTERNAL_TOKEN` | *unset* | `core/middleware.py:20` | Security-relevant. Token that lets the in-app tool layer reach admin-gated routes over loopback. Unset generates a fresh per-process token, which is what you want unless something outside the process needs the same value. |
### Integrations (Claude, Codex)
| Variable | Default | Read in | What it does |
|---|---|---|---|
| `ODYSSEUS_API_TOKEN` | `''` | `integrations/claude/skills/odysseus/scripts/odysseus_api.py:40` (+1 more) | API token those scripts authenticate with. Both this and the URL are required; the scripts name whichever is missing. |
| `ODYSSEUS_URL` | `''` | `integrations/claude/skills/odysseus/scripts/odysseus_api.py:39` (+1 more) | Base URL of the Odysseus instance the bundled integration scripts call. |
## Internal and development-only variables
Listed for completeness. Setting one of these on a real install is either a no-op or a way to break something quietly.
### Model routing and providers
| Variable | Default | Read in | What it does |
|---|---|---|---|
| `ODYSSEUS_COPILOT_EDITOR_VERSION` | `'Odysseus/1.0'` | `src/copilot.py:54` | Editor-version header presented to the Copilot API. Kept stable on purpose. |
| `ODYSSEUS_COPILOT_INTEGRATION_ID` | `'vscode-chat'` | `src/copilot.py:51` | Integration id presented to the Copilot API. Kept stable on purpose. |
| `ODYSSEUS_COPILOT_USER_AGENT` | `'Odysseus/1.0'` | `src/copilot.py:48` | Editor-like User-Agent presented to the Copilot API. Kept stable on purpose. |
| `ODYSSEUS_DEBUG_LLM_SHAPE` | `''` | `src/llm_core.py:3552` (+1 more) | Truthy logs the shape of streamed provider chunks. Debugging aid for provider response parsing. |
### Agent loop and tool execution
| Variable | Default | Read in | What it does |
|---|---|---|---|
| `ODYSSEUS_CAPTURE_MODEL_REQUESTS` | `''` | `src/agent_loop.py:3925` | Truthy writes model-request snapshots for local debugging. The marker file `/tmp/odysseus_capture_model_requests` enables the same thing. |
| `ODYSSEUS_EXPOSE_RAW_BROWSER_MCP` | `''` | `src/agent_loop.py:4130` | Truthy stops hiding the raw Playwright MCP tools from agent prompts when the private-browser tool is available. |
| `ODYSSEUS_TOOL_CONTRACT_ROOT` | `'<repo>/scripts'` | `src/clean_agent_preview.py:2184` (+1 more) | Directory holding the tool-contract scripts the clean-agent preview loads. The default is the repository's bundled scripts directory; set the variable to override it. |
### Email
| Variable | Default | Read in | What it does |
|---|---|---|---|
| `ODYSSEUS_EMAIL_FIXTURE` | *unset* | `mcp_servers/email_server.py:1122` (+7 more) | Exactly `1`, plus a fixture file on disk, makes the email MCP server serve fixtures instead of a real mailbox. |
### Testing, capture and development tooling
| Variable | Default | Read in | What it does |
|---|---|---|---|
| `ODYSSEUS_AJAX_TEST_URL` | *unset* | `tests/test_ajax_email_live.py:17` (+4 more) | Chat-completions URL of a live Ajax endpoint. Unset skips the opt-in live Ajax email tests. |
| `ODYSSEUS_BROWSER_LIVE_CONTRACT` | *unset* | `tests/test_browser_producer_live_contract.py:19` | Set 1 only in the allowlisted release Docker environment to run the browser producer contract tests. Does not enable browser page operations. |
| `ODYSSEUS_EDITOR_ACTIONS` | `','.join([*actions, 'edit', 'update'])` | `tests/tools/editor_writing_smoke.py:71` | Comma-separated writing actions the editor-writing smoke tool runs. Unset runs every action plus edit and update. |
| `ODYSSEUS_EDITOR_MAX_TOKENS` | `'4096'` | `tests/tools/editor_writing_smoke.py:110` | Completion token limit for each editor-writing smoke request. |
| `ODYSSEUS_EDITOR_RICH_FIXTURE` | *unset* | `tests/tools/editor_writing_smoke.py:80` | Set to 1 to run the editor-writing smoke tool against a rich-text document fixture instead of Markdown. |
| `ODYSSEUS_EDITOR_TEST_ENDPOINT` | *unset* | `tests/tools/editor_writing_smoke.py:110` (+1 more) | Chat-completions URL the opt-in editor-writing and organizer smoke tools drive. Both tools require it. |
| `ODYSSEUS_EDITOR_TRACE` | *unset* | `tests/tools/editor_writing_smoke.py:140` | Any non-empty value prints every stream event after each editor-writing smoke action. |
| `ODYSSEUS_ORGANIZER_AUTO_CHOICE` | *unset* | `tests/tools/organizer_smoke.py:199` | With organizer tracing on, any non-empty value replaces forced tool choice with auto on traced requests. |
| `ODYSSEUS_ORGANIZER_CASES` | *unset* | `tests/tools/organizer_smoke.py:223` (+1 more) | Comma-separated organizer smoke case names to run. Unset runs every case. |
| `ODYSSEUS_ORGANIZER_TRACE` | *unset* | `tests/tools/organizer_smoke.py:193` | Any non-empty value prints each provider request the organizer smoke tool sends. |
| `ODYSSEUS_ORGANIZER_TRACE_MESSAGES` | *unset* | `tests/tools/organizer_smoke.py:203` | With organizer tracing on, any non-empty value also prints the request messages. |
| `ODYSSEUS_QA_TEACHER_ATTEMPTS` | `'3'` | `scripts/odysseus_conversation_qa.py:370` | Retry budget for the conversation-QA teacher model call. Clamped to 1-3. |
| `ODYSSEUS_QA_TEACHER_TIMEOUT` | `'120'` | `scripts/odysseus_conversation_qa.py:372` | Timeout in seconds for that call. Clamped to 15-120. |
| `ODYSSEUS_RUNTIME_REVISION` | `''` | `routes/chat_helpers.py:198` (+1 more) | Revision string stamped into each captured SFT trace record, so a trace can be tied back to the build that produced it. |
| `ODYSSEUS_SFT_DISABLE_WORKSPACE_TOOLS` | `'1'` | `src/agent_loop.py:7408` | On by default. Keeps synthetic personal-assistant fixtures out of workspace mode; set 0, false, no or off to let them through. |
| `ODYSSEUS_SFT_FORCE_UTC_TIMEZONE` | `'0'` | `routes/chat_routes.py:2097` | Truthy forces `sft_` accounts to UTC for deterministic batch generation. Interactive accounts still follow the browser timezone. |
| `ODYSSEUS_SFT_TRACE_CAPTURE` | `'1'` | `routes/chat_helpers.py:161` (+1 more) | On by default, but only for owners whose name starts with `sft_`. Set 0, false, no or off to stop writing training traces. |
| `ODYSSEUS_SFT_TRACE_DIR` | *unset* | `routes/chat_helpers.py:195` (+2 more) | Directory the SFT trace JSONL files are written to. Defaults to `sft_traces` under the data directory. |
| `ODYSSEUS_SKIP_RUN_HINT` | *unset* | `setup.py:284` | Any non-empty value suppresses the `start the server with` hint at the end of setup. `start-macos.sh` sets it because it starts the server itself. |
| `ODYSSEUS_TEST_STATIC_ORIGIN` | *unset* | `scripts/css_snapshot.py:254` (+6 more) | Origin an already-running static server is serving the repository from, so snapshot tooling reuses it instead of starting its own. |
| `ODYSSEUS_TEST_STATIC_PORT` | *unset* | `tests/conftest.py:201` | Fixed port for the test suite's static server. Unset takes an ephemeral port, which is what keeps parallel runs from colliding. |
### Build and release metadata
| Variable | Default | Read in | What it does |
|---|---|---|---|
| `ODYSSEUS_BUILD_VERSION` | `''` | `src/constants.py:15` | Overrides the build-version string the API and UI report, without touching the public application version. |
| `ODYSSEUS_SOURCE_COMMIT` | `''` | `src/constants.py:33` | Overrides the source commit reported for runtime provenance, for builds that ship without a git directory. |
## How this page is generated
The generator walks the Python sources under `app.py`, `launcher.py`, `setup.py`, `companion`, `config`, `core`, `integrations`, `mcp_servers`, `routes`, `scripts`, `services`, `src`, `tests` and finds
reads three ways, because no single pattern covers the codebase:
- Direct reads: `os.getenv(...)`, and any `.get` / `.setdefault` / `.pop` call or
subscript keyed by an `ODYSSEUS_*` literal, including names held in a
module-level constant. The receiver is not required to be `os.environ`, because
several call sites read through a mapping passed in as an argument
(`src/tool_index.py`, `src/host_docker_access.py`).
- Calls to env-reader helpers - any function that forwards one of its own
parameters to an environment read. This is detected rather than hardcoded, so a
new helper needs no change here. It is what finds the upload caps in
`src/upload_limits.py` and the media-ingress overrides in
`src/media_ingress.py`.
- A regex sweep of the raw file text, for reads the AST cannot see.
`routes/cookbook_helpers.py` builds an Ollama probe script as a list of source
lines, so one read lives inside a string literal.
The three passes are not redundancy. A line-based grep for a direct
`os.environ.get("ODYSSEUS_...` call finds 82 of the 112 variables on this
page. What it misses is reads through an env-reader helper, reads whose call
spans more than one line, reads whose variable name is held in a module
constant, and reads through a mapping passed in as an argument - which is the
whole reason this page is generated rather than maintained.
The `Default` column shows the expression as written, with one level of
indirection resolved: a module-level constant and a dataclass field default are
replaced by the literal they hold, so `defaults.max_media_files` shows as `4`.
Anything computed at import time - `get_default_data_dir()` - is shown as
written, because that is the honest answer. A few call sites supply their
fallback with `or` rather than a default argument; those show as *unset* and say
so in the last column.
Regenerate it with:
```bash
python3 scripts/generate_env_reference.py
```
Binary file not shown.
Binary file not shown.
+20 -195
View File
@@ -273,87 +273,12 @@
.shot .frame-dots { position: absolute; top: 10px; left: 12px; display: flex; gap: 5px; }
.shot .frame-dots i { width: 8px; height: 8px; border-radius: 50%; background: #39414d; display: inline-block; }
/* Previews — expanding hover carousel that plays a video on hover/tap */
.previews { display: flex; align-items: center; gap: 12px; height: 480px; max-width: 1000px; margin: 36px auto 0; }
.preview-panel {
position: relative; flex: 1 1 0; min-width: 0; height: 360px; overflow: hidden;
border: 1px solid var(--border); border-radius: var(--radius); cursor: pointer;
background: linear-gradient(180deg, var(--panel), var(--panel2));
transition: flex-grow .5s cubic-bezier(.2,.7,.2,1), height .5s cubic-bezier(.2,.7,.2,1), border-color .25s ease;
}
.previews:hover .preview-panel { flex-grow: 0.55; height: 300px; }
.preview-panel:hover, .preview-panel:focus-visible, .preview-panel.is-active { flex-grow: 3.4 !important; height: 480px !important; border-color: var(--accent); }
.preview-panel .ph {
position: absolute; inset: 0; display: flex; flex-direction: column;
align-items: center; justify-content: center; gap: 10px;
color: var(--muted); font-size: 12.5px; opacity: 0.7; text-align: center; padding: 8px;
}
.preview-panel video {
position: absolute; inset: 0; width: 100%; height: 100%; object-fit: cover;
z-index: 1; opacity: 0; transition: opacity .3s ease; background: transparent;
}
.preview-panel.has-video video { opacity: 1; }
/* These clips have their action on the left, so show the left edge instead of
the centered crop. */
.preview-panel:has(source[src="document.webm"]) video,
.preview-panel:has(source[src="notes.webm"]) video { object-position: right center; }
.preview-panel .label {
position: absolute; z-index: 2; left: 0; right: 0; bottom: 0; padding: 14px 16px;
background: linear-gradient(0deg, rgba(0,0,0,0.82), transparent);
color: var(--heading);
display: flex; flex-direction: column; align-items: flex-start; gap: 4px;
}
.preview-panel .label .t { display: flex; align-items: center; gap: 8px; white-space: nowrap; font-weight: 700; font-size: 14px; }
.preview-panel .label .ico { color: var(--accent); flex-shrink: 0; }
.preview-panel .label .desc {
font-weight: 400; font-size: 12.5px; line-height: 1.35; color: rgba(255,255,255,0.82);
white-space: normal; max-height: 0; opacity: 0; overflow: hidden;
transition: max-height .4s ease, opacity .4s ease;
}
.preview-panel:hover .label .desc, .preview-panel:focus-visible .label .desc, .preview-panel.is-active .label .desc { max-height: 64px; opacity: 1; }
@media (max-width: 760px) {
.previews { flex-direction: column; height: auto; touch-action: pan-y; }
.preview-panel { height: 190px; flex: none; width: 100%; }
.preview-panel.is-active { height: 280px !important; }
.previews:hover .preview-panel, .preview-panel:hover { flex: none !important; }
.preview-panel .label .desc { max-height: 64px; opacity: 1; }
}
/* Fullscreen video background for a section — treated as an ambient, cinematic
backdrop (soft blur + slow drift) so it sets a mood without fighting the copy. */
.has-bg-video { position: relative; overflow: hidden; }
.has-bg-video .sec-bg {
position: absolute; inset: 0; width: 100%; height: 100%;
object-fit: cover; z-index: 0; pointer-events: none;
/* blur softens the busy frame; the extra scale hides the blurred edges */
filter: blur(4px) saturate(1.08) brightness(0.92);
transform: scale(1.12);
transform-origin: 55% 45%;
animation: bg-drift 36s ease-in-out infinite alternate;
will-change: transform;
}
@keyframes bg-drift {
from { transform: scale(1.12) translate(0, 0); }
to { transform: scale(1.2) translate(-2.5%, -1.5%); }
}
.has-bg-video .sec-bg-tint {
position: absolute; inset: 0; z-index: 1; pointer-events: none;
background:
radial-gradient(900px 520px at 78% 18%, rgba(224,108,117,0.16), transparent 60%),
radial-gradient(760px 520px at 8% 88%, rgba(53,90,102,0.30), transparent 58%),
linear-gradient(180deg, rgba(17,17,17,0.86), rgba(17,17,17,0.62) 42%, rgba(17,17,17,0.92)),
radial-gradient(1200px 680px at 50% 46%, rgba(17,17,17,0.18), rgba(17,17,17,0.74));
}
.has-bg-video .wrap { position: relative; z-index: 2; }
/* Lift the copy off the moving backdrop. */
.has-bg-video .eyebrow,
.has-bg-video .h { text-shadow: 0 2px 22px rgba(0,0,0,0.7); }
.has-bg-video .sub { color: #b9e6f4; text-shadow: 0 1px 14px rgba(0,0,0,0.75); }
.hero.has-bg-video h1, .hero.has-bg-video .wordmark,
.hero.has-bg-video .lede, .hero.has-bg-video .slogan { text-shadow: 0 2px 22px rgba(0,0,0,0.72); }
@media (prefers-reduced-motion: reduce) {
.has-bg-video .sec-bg { animation: none; transform: scale(1.12); }
}
/* Feature descriptions */
.feature-descriptions { display: grid; grid-template-columns: repeat(auto-fit, minmax(220px, 1fr)); gap: 16px; margin-top: 36px; }
.feature-description { padding: 20px; border: 1px solid var(--border); border-radius: 12px; background: var(--panel); }
.feature-description .t { display: flex; align-items: center; gap: 8px; font-weight: 700; }
.feature-description .ico { color: var(--accent); }
.feature-description .desc { display: block; margin-top: 8px; color: var(--muted); line-height: 1.6; }
/* Get started */
.start {
@@ -598,62 +523,28 @@
</div>
</section>
<!-- PREVIEWS — hover/tap to expand + play -->
<section id="previews">
<!-- FEATURE DESCRIPTIONS -->
<section id="features-overview">
<div class="wrap">
<div class="center">
<div class="eyebrow"><svg viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M2 12s3.6-7 10-7 10 7 10 7-3.6 7-10 7-10-7-10-7z"/><circle cx="12" cy="12" r="3"/></svg>See it in action</div>
<h2 class="h">Hover or tap to take a closer look</h2>
<p class="sub center">Each panel expands and plays its preview when you hover or tap it. Swipe on mobile to move through them.</p>
<h2 class="h">Explore the workspace</h2>
<p class="sub center">Tools for local AI and everyday work.</p>
</div>
<div class="previews">
<div class="preview-panel" tabindex="0">
<div class="ph"><svg width="30" height="30" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="1.6" stroke-linecap="round" stroke-linejoin="round"><path d="M21 15a2 2 0 0 1-2 2H7l-4 4V5a2 2 0 0 1 2-2h14a2 2 0 0 1 2 2z"/></svg><span>[ Chat &amp; Agents ]</span></div>
<video muted loop playsinline preload="none"><source src="chat.webm" type="video/webm"><source src="chat.mp4" type="video/mp4"></video>
<div class="label"><span class="t"><svg class="ico" width="16" height="16" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M21 15a2 2 0 0 1-2 2H7l-4 4V5a2 2 0 0 1 2-2h14a2 2 0 0 1 2 2z"/></svg>Chat &amp; Agents</span><span class="desc">Talk to any local model, or give it tools and let the agent run.</span></div>
</div>
<div class="preview-panel" tabindex="0">
<div class="ph"><svg width="30" height="30" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="1.6" stroke-linecap="round" stroke-linejoin="round"><path d="M12 2 2 7l10 5 10-5-10-5zM2 17l10 5 10-5M2 12l10 5 10-5"/></svg><span>[ Cookbook ]</span></div>
<div class="label"><span class="t"><svg class="ico" width="16" height="16" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M12 2 2 7l10 5 10-5-10-5zM2 17l10 5 10-5M2 12l10 5 10-5"/></svg>Cookbook</span><span class="desc">Download, serve, and manage local models across your machines.</span></div>
</div>
<div class="preview-panel" tabindex="0">
<div class="ph"><svg width="30" height="30" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="1.6" stroke-linecap="round" stroke-linejoin="round"><circle cx="11" cy="11" r="7"/><path d="m21 21-4.35-4.35"/></svg><span>[ Deep Research ]</span></div>
<video muted loop playsinline preload="none"><source src="research.webm" type="video/webm"><source src="research.mp4" type="video/mp4"></video>
<div class="label"><span class="t"><svg class="ico" width="16" height="16" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><circle cx="11" cy="11" r="7"/><path d="m21 21-4.35-4.35"/></svg>Deep Research</span><span class="desc">Ask once: it searches, reads sources, and writes back a cited report.</span></div>
</div>
<div class="preview-panel" tabindex="0">
<div class="ph"><svg width="30" height="30" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="1.6" stroke-linecap="round" stroke-linejoin="round"><rect x="3" y="4" width="7" height="16" rx="1"/><rect x="14" y="4" width="7" height="16" rx="1"/></svg><span>[ Compare ]</span></div>
<video muted loop playsinline preload="none"><source src="compare.webm" type="video/webm"><source src="compare.mp4" type="video/mp4"></video>
<div class="label"><span class="t"><svg class="ico" width="16" height="16" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><rect x="3" y="4" width="7" height="16" rx="1"/><rect x="14" y="4" width="7" height="16" rx="1"/></svg>Compare</span><span class="desc">Send one prompt to many models at once and watch them answer side by side.</span></div>
</div>
<div class="preview-panel" tabindex="0">
<div class="ph"><svg width="30" height="30" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="1.6" stroke-linecap="round" stroke-linejoin="round"><path d="M14 2H6a2 2 0 0 0-2 2v16a2 2 0 0 0 2 2h12a2 2 0 0 0 2-2V8z"/><path d="M14 2v6h6"/><path d="M16 13H8M16 17H8M10 9H8"/></svg><span>[ Documents ]</span></div>
<video muted loop playsinline preload="none"><source src="document.webm" type="video/webm"><source src="document.mp4" type="video/mp4"></video>
<div class="label"><span class="t"><svg class="ico" width="16" height="16" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M14 2H6a2 2 0 0 0-2 2v16a2 2 0 0 0 2 2h12a2 2 0 0 0 2-2V8z"/><path d="M14 2v6h6"/><path d="M16 13H8M16 17H8M10 9H8"/></svg>Documents</span><span class="desc">A document editor that puts you first — work on what you want, with AI help when you want it.</span></div>
</div>
<div class="preview-panel" tabindex="0">
<div class="ph"><svg width="30" height="30" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="1.6" stroke-linecap="round" stroke-linejoin="round"><path d="m3 7 2 2 4-4"/><path d="m3 17 2 2 4-4"/><path d="M13 6h8M13 18h8"/></svg><span>[ Notes &amp; Tasks ]</span></div>
<video muted loop playsinline preload="none"><source src="notes.webm" type="video/webm"><source src="notes.mp4" type="video/mp4"></video>
<div class="label"><span class="t"><svg class="ico" width="16" height="16" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="m3 7 2 2 4-4"/><path d="m3 17 2 2 4-4"/><path d="M13 6h8M13 18h8"/></svg>Notes &amp; Tasks</span><span class="desc">Capture notes and to-dos, or let scheduled agents work and brief you after.</span></div>
</div>
<div class="preview-panel" tabindex="0">
<div class="ph"><svg width="30" height="30" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="1.6" stroke-linecap="round" stroke-linejoin="round"><rect x="3" y="3" width="18" height="18" rx="2"/><circle cx="9" cy="9" r="2"/><path d="m21 15-3.6-3.6a2 2 0 0 0-2.8 0L6 21"/></svg><span>[ Image Gallery ]</span></div>
<video muted loop playsinline preload="none"><source src="gallery.webm" type="video/webm"><source src="gallery.mp4" type="video/mp4"></video>
<div class="label"><span class="t"><svg class="ico" width="16" height="16" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><rect x="3" y="3" width="18" height="18" rx="2"/><circle cx="9" cy="9" r="2"/><path d="m21 15-3.6-3.6a2 2 0 0 0-2.8 0L6 21"/></svg>Image Gallery</span><span class="desc">Generate, edit, remove backgrounds, and inpaint in your own gallery.</span></div>
</div>
<div class="preview-panel" tabindex="0">
<div class="ph"><svg width="30" height="30" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="1.6" stroke-linecap="round" stroke-linejoin="round"><path d="M12 2.7 6.3 8.4a8 8 0 1 0 11.4 0z"/></svg><span>[ Themes ]</span></div>
<video muted loop playsinline preload="none"><source src="theme.webm" type="video/webm"><source src="theme.mp4" type="video/mp4"></video>
<div class="label"><span class="t"><svg class="ico" width="16" height="16" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M12 2.7 6.3 8.4a8 8 0 1 0 11.4 0z"/></svg>Themes</span><span class="desc">Restyle and make it yours — edit your own, or ask the agent to make one.</span></div>
</div>
<div class="feature-descriptions">
<div class="feature-description"><span class="t"><svg class="ico" width="16" height="16" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M21 15a2 2 0 0 1-2 2H7l-4 4V5a2 2 0 0 1 2-2h14a2 2 0 0 1 2 2z"/></svg>Chat &amp; Agents</span><span class="desc">Talk to any local model, or give it tools and let the agent run.</span></div>
<div class="feature-description"><span class="t"><svg class="ico" width="16" height="16" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M12 2 2 7l10 5 10-5-10-5zM2 17l10 5 10-5M2 12l10 5 10-5"/></svg>Cookbook</span><span class="desc">Download, serve, and manage local models across your machines.</span></div>
<div class="feature-description"><span class="t"><svg class="ico" width="16" height="16" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><circle cx="11" cy="11" r="7"/><path d="m21 21-4.35-4.35"/></svg>Deep Research</span><span class="desc">Ask once: it searches, reads sources, and writes back a cited report.</span></div>
<div class="feature-description"><span class="t"><svg class="ico" width="16" height="16" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><rect x="3" y="4" width="7" height="16" rx="1"/><rect x="14" y="4" width="7" height="16" rx="1"/></svg>Compare</span><span class="desc">Send one prompt to many models at once and watch them answer side by side.</span></div>
<div class="feature-description"><span class="t"><svg class="ico" width="16" height="16" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M14 2H6a2 2 0 0 0-2 2v16a2 2 0 0 0 2 2h12a2 2 0 0 0 2-2V8z"/><path d="M14 2v6h6"/><path d="M16 13H8M16 17H8M10 9H8"/></svg>Documents</span><span class="desc">A document editor that puts you first — work on what you want, with AI help when you want it.</span></div>
<div class="feature-description"><span class="t"><svg class="ico" width="16" height="16" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="m3 7 2 2 4-4"/><path d="m3 17 2 2 4-4"/><path d="M13 6h8M13 18h8"/></svg>Notes &amp; Tasks</span><span class="desc">Capture notes and to-dos, or let scheduled agents work and brief you after.</span></div>
<div class="feature-description"><span class="t"><svg class="ico" width="16" height="16" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><rect x="3" y="3" width="18" height="18" rx="2"/><circle cx="9" cy="9" r="2"/><path d="m21 15-3.6-3.6a2 2 0 0 0-2.8 0L6 21"/></svg>Image Gallery</span><span class="desc">Generate, edit, remove backgrounds, and inpaint in your own gallery.</span></div>
<div class="feature-description"><span class="t"><svg class="ico" width="16" height="16" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M12 2.7 6.3 8.4a8 8 0 1 0 11.4 0z"/></svg>Themes</span><span class="desc">Restyle and make it yours — edit your own, or ask the agent to make one.</span></div>
</div>
</div>
</section>
<!-- HOW IT STARTED -->
<section id="how" class="has-bg-video">
<video class="sec-bg" autoplay muted loop playsinline preload="auto"><source src="bg.webm" type="video/webm"><source src="bg.mp4" type="video/mp4"></video>
<div class="sec-bg-tint"></div>
<section id="how">
<div class="wrap">
<div class="eyebrow"><svg viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><circle cx="12" cy="12" r="9"/><path d="m15.6 8.4-2.1 5.1-5.1 2.1 2.1-5.1z"/></svg>How it actually started</div>
<h2 class="h">Uncompromised local LLM experience.</h2>
@@ -790,72 +681,6 @@
}
})();
// Previews: hovering/tapping a panel expands it (CSS) and plays its video; the
// video only becomes visible once it actually starts playing, so missing
// files just leave the labeled placeholder.
(function () {
var panels = [].slice.call(document.querySelectorAll('.preview-panel'));
if (!panels.length) return;
var active = -1;
function playPanel(p) {
var v = p.querySelector('video');
if (!v) return;
var pr = v.play();
if (pr && pr.catch) pr.catch(function () {});
}
function pausePanel(p) {
var v = p.querySelector('video');
if (v) v.pause();
}
function setActive(i, shouldPlay) {
active = (i + panels.length) % panels.length;
panels.forEach(function (panel, k) {
var on = k === active;
panel.classList.toggle('is-active', on);
panel.setAttribute('aria-expanded', on ? 'true' : 'false');
if (!on) pausePanel(panel);
});
if (shouldPlay !== false) playPanel(panels[active]);
}
panels.forEach(function (p, i) {
var v = p.querySelector('video');
if (v) {
v.addEventListener('playing', function () { p.classList.add('has-video'); });
v.addEventListener('pause', function () { /* keep last frame */ });
}
p.setAttribute('aria-expanded', 'false');
p.addEventListener('mouseenter', function () { setActive(i); });
p.addEventListener('focus', function () { setActive(i); });
p.addEventListener('mouseleave', function () {
if (!window.matchMedia || !window.matchMedia('(hover: none)').matches) {
p.classList.remove('is-active');
p.setAttribute('aria-expanded', 'false');
pausePanel(p);
}
});
p.addEventListener('blur', function () { pausePanel(p); });
p.addEventListener('click', function () { setActive(i); });
});
var strip = document.querySelector('.previews');
var sx = null, sy = null;
if (strip) {
strip.addEventListener('touchstart', function (e) {
if (!e.touches.length) return;
sx = e.touches[0].clientX;
sy = e.touches[0].clientY;
}, { passive: true });
strip.addEventListener('touchend', function (e) {
if (sx === null || sy === null || !e.changedTouches.length) return;
var dx = e.changedTouches[0].clientX - sx;
var dy = e.changedTouches[0].clientY - sy;
if (Math.abs(dx) > 42 && Math.abs(dx) > Math.abs(dy) * 1.25) {
setActive((active < 0 ? 0 : active) + (dx < 0 ? 1 : -1));
}
sx = sy = null;
}, { passive: true });
}
})();
// Domino reveal: fade/slide each section in as it scrolls into view.
(function () {
var els = document.querySelectorAll('.hero, section');
Binary file not shown.
@@ -0,0 +1,66 @@
# Original Harness Capability Policy
Odysseus Original is the canonical product and the only agent orchestrator.
Specialized runners are evidence sources, not merge targets.
## Architecture
The supported shape is:
1. One Original agent loop owns prompting, tool selection, policy, evidence,
recovery, compaction, and completion.
2. Capability contracts describe what the current environment can do.
3. Execution bridges or MCP servers perform work in the owning environment.
4. Personal tools remain available in interactive sessions but are excluded
from external terminal contracts unless explicitly provided.
## Capability Decisions
| Capability | Decision | Canonical form |
| --- | --- | --- |
| Workspace execution | Keep | Request-scoped `AgentExecutionBridge`, moving toward a Workspace MCP boundary |
| File mutation | Keep | `write_file`, `edit_file`, and `apply_patch` with evidence recording |
| Process recovery | Keep | Bounded polling, timeout, exit status, and stale-process recovery behind the execution contract |
| Repeated-action recovery | Keep | Detect identical calls, unchanged successful results, and failed batches separately |
| Completion | Keep | Required artifacts and executable verifier evidence; no evaluator-specific shortcuts |
| Context control | Keep | Deterministic compaction with retained user evidence and bounded tool output |
| Search | Keep | Existing private SearXNG path |
| Browser interaction | Keep | Existing private browser for rendered pages, sessions, clicks, and screenshots |
| Personal tools | Keep, isolated | Separate personal capability surface; never substitute editor documents for workspace files |
| Tool discovery | Keep | Existing tool RAG and capability-aware selection |
| Media ingress | Keep | Bounded image, audio, video-frame, and document ingestion with hashes and trace metadata |
| Visual verification | Candidate | Generic source grounding and rendered-artifact checks, gated by available media capabilities |
| Document access | Candidate | Structured PDF and Office extraction through `read_file` or a Workspace MCP implementation |
| Skills | Keep, generic only | Procedures for artifact completion, recovery, verification, development, media evidence, and research |
## Rejected Merges
Do not merge:
- another agent loop, submit controller, or conversation state machine;
- category names, task identifiers, fixed workspace paths, expected answers,
grader behavior, scoring rules, or evaluator prompts;
- per-category turn thresholds or phase transitions;
- terminal multiplexer control when the environment already exposes direct
process execution;
- repeated prompt guards that do not add new executable evidence;
- browser or search replacements for capabilities Original already owns.
## Merge Gate
A mechanism may enter Original only when all of these are true:
1. It is useful outside the evaluation that revealed it.
2. It is selected by an environment capability, not a task name or path.
3. It fits inside the Original loop or behind an execution boundary.
4. Its success or failure produces traceable evidence.
5. It has focused regressions and a representative held-out run.
6. It removes or contains complexity instead of adding an overlapping mode.
## Current Priority
The next capability work is reliable structured document access. Failed
binary-file inspections must remain failures; they must not be interpreted as
unchanged evidence or trigger early artifact synthesis. Once grounding is
reliable, add generic rendered-artifact verification without importing any
specialized task phases.
+346
View File
@@ -0,0 +1,346 @@
# Photo Editor Product Audit
Date: 2026-08-30
## Executive Verdict
Odysseus already has the structure of a real layered raster editor. It is not a
mockup: layers, nested groups, masks, selections, retained text, blending,
history, document geometry, project recovery, controlled export, and several AI
workflows operate on an editable document model.
It is **roughly 78% of a dependable everyday photo editor**, but only **about
35% of a professional Photoshop/Photopea alternative**. Those are deliberately
separate scores. The first target needs complete, trustworthy common workflows;
the second also needs non-destructive sources and filters, color management,
professional file interchange, vector/path tooling, automation, and scale.
The main product gap is no longer basic layer infrastructure. It is the absence
of a strong photo-correction and retouching workflow, combined with insufficient
proof that every existing operation preserves pixels, masks, text, and project
state across desktop and touch.
## Current Scorecard
| Area | Everyday readiness | Current assessment |
| --- | ---: | --- |
| Canvas navigation and precision | 90% | Pan/zoom, fit and 1:1, rulers, guides, grid, snapping, and numeric geometry are present |
| Layers and compositing | 96% | Raster/text layers, nested groups, clipping, masks, blend modes, multi-select, locks, subtree reorder, and shared transforms are strong |
| Selections and masks | 97% | Marquee, lasso, wand, SAM, Quick Mask, named selections, affine selection transform, and linked/unlinked layer-mask positioning work |
| Text | 65% | Retained text exists; paragraph layout, tracking, font status, stronger hit testing, and mobile proof do not |
| Painting and retouching | 70% | Brush, eraser, clone, source-free and sampled healing, smudge, retained linear/radial gradients, fill, AI inpaint, background removal, sharpen, eyedropper, dodge/burn, and brush presets exist; advanced raster retouching remains |
| Photo correction | 65% | Retained levels, curves, exposure, white balance, hue/saturation, vibrance, shadows/highlights, color balance, selective color, gradient map, brightness/contrast, and histogram controls exist; camera/lens correction remains |
| Non-destructive editing | 60% | Text and adjustment metadata are retained, image imports and pasted selections create source-backed placed layers, raster layers can be converted to sources, and transforms/replacement preserve source pixels; linked instances and richer source editing remain |
| Save, recovery, and export | 88% | Versioned project recovery and PNG/JPEG/WebP export are strong; pixel-equivalence, metadata, and color-profile policies remain |
| Mobile editing | 45% | Responsive UI, touch navigation, and non-overlapping tool/layer sheets are tested; core editing gestures still need broader coverage |
| Performance and color fidelity | 30% | Safety limits exist; workers/tiles, stress evidence, ICC handling, and high-bit-depth editing do not |
## Verified Baseline
The current implementation was checked on 2026-08-30:
- **72 editor-focused unit tests pass.** Coverage includes the v12 document
model, migrations, validation, geometry, selections, masks, mask offsets,
groups, clipping, multi-select, transforms, retained text, history limits,
guides, and export settings.
- **81 Chromium Playwright workflows pass.** They cover the layered core path,
group masks and reorder, clipping, independent locks, multi-selection,
nested groups, shared transforms, linked/unlinked mask movement and reopen,
exact draft reopen, export dimensions, project recovery, Quick Mask, precise
marquee geometry, transformed selections, named selections, and mobile touch
crop, selection, brush, and transform gestures, plus mobile mask persistence,
movement, and PNG export.
- Mobile editor refresh now has a dedicated recovery gate that restores the
active draft, editor tab, saved status, canvas, and layer stack.
- The adjustment popup has a 320px phone-width regression gate: slider rows stay
inside the sheet and the Apply/Cancel actions remain reachable.
- The Gradient Map adjustment has a dedicated 320px gate: its color controls
collapse to one responsive column and remain inside the adjustment sheet.
- Every retained adjustment popup now has a 320px matrix check for viewport
bounds, horizontal overflow, and reachable Apply/Cancel actions.
- The adjustment export matrix compares decoded PNG pixels against the visible
composite for every retained adjustment family, including Curves and
Gradient Map.
- The adjustment compositing workflow also compares export pixels when a
retained adjustment is clipped, masked, opacity-modified, undone/redone, and
reopened from a saved draft.
- The retained-adjustment export workflow passes in both Chromium and Firefox;
the Playwright harness now supports selecting `chromium`, `firefox`, or
`webkit` through `PHOTO_EDITOR_E2E_BROWSER`.
- Retained-effect previews now fall back cleanly when a worker cannot be
created, and stale worker errors no longer trigger an unnecessary full-size
synchronous render.
- The mobile layer-sheet workflow also verifies touch mask editing, undo/redo,
and mask persistence after browser reload.
- The merge-fidelity workflow compares the rendered composite before and after
Merge All with retained effects and adjustment layers, ensuring those edits
are baked into the resulting raster instead of being dropped.
- The grouped-effects workflow compares a retained group effect against the
exported PNG pixel-for-pixel and verifies its parameters survive draft reopen.
- The export preview workflow verifies matte pixels are cleared when switching
back to transparency.
- The export dialog workflow restores focus to the Save control after closing,
including the menu-launched export path.
- The core workflow now compares SHA-256 digests before and after draft reopen,
requires exact decoded pixels for native PNG export, and bounds premultiplied
pixel error for resized PNG output.
- The linked-mask regression gate verifies that an unlinked mask remains at a
fixed document position when its parent layer moves or transforms. Brush,
Wand, and Fill now resolve independently moved masks from the same origin.
This is solid Chromium/mouse evidence with focused touch coverage for crop,
selection, brush, and transform gestures, plus a focused Firefox export check.
It does not establish full Safari/Firefox compatibility, complete touch
reliability, large-document responsiveness, metadata or ICC fidelity, or
pixel-equivalent exports for every format.
## What Already Works
| Capability | Implementation status |
| --- | --- |
| Layered document | Raster and retained-text layers, nested groups, opacity, visibility, 16 Canvas2D blend modes, clipping, duplicate, merge, and hierarchy-aware reorder |
| Layer control | Multi-select, shared-bounds transforms, full/pixel/transparency/position locks, group masks, paintable layer masks, and linked/unlinked mask position |
| Selection model | Rectangle/ellipse marquee, lasso, Magic Wand, SAM, new/add/subtract/intersect, animated boundary, exact geometry, move/scale/rotate/flip, nudge, invert, Quick Mask, reselect, and named selections |
| Paint and AI | Brush, eraser, clone stamp, selection fill, inpaint, background removal, SAM, harmonize, upscale, denoise, face enhancement, and style operations |
| Geometry | Crop, image resize, canvas resize, rotate, flip, move, transforms, rulers, guides, grid, and snapping |
| Recovery | Validated v12+ project format, explicit migrations, bounded autosave/history, partial corrupt-project recovery, and exact server-draft reopen |
| Delivery | Previewed PNG/JPEG/WebP export with quality, dimensions, aspect lock, transparency/matte, filename, and gallery-copy flow |
## Release Blockers
These are the gaps that prevent calling the editor dependable today.
### P0: Trust Existing Operations
1. **Pixel-equivalent export proof is still too narrow.**
The core layered workflow now proves exact native PNG pixels and bounded
resized-PNG error, and format metadata is covered for JPEG/WebP. Grouped
blending, clipping, matte pixels, and color-profile behavior still need the
same evidence.
2. **Touch editing is only partially release-tested.**
Pan and pinch primitives exist, and paint, selection, crop, and transform
now have focused browser coverage, but broader mobile overlap and
inaccessible-control assertions are still needed around edge cases.
3. **Operation contracts are incomplete.**
Every destructive operation needs an explicit test matrix for raster layers,
retained text, linked/unlinked masks, group masks, selections, locks, clipped
layers, and multi-selection. The independently positioned mask bug found in
Brush/Wand/Fill demonstrates why shared coordinate helpers are required.
4. **Large-document behavior is bounded, not proven.**
Full-canvas Canvas2D compositing and RGBA history snapshots have hard limits,
but there is no stress suite, cancellation contract, worker/offscreen path,
memory telemetry, or degraded-preview strategy.
5. **Color and metadata behavior is undefined.**
Import relies on browser decoding and export writes a new bitmap. Users are
not told whether ICC, EXIF/IPTC, orientation, DPI, or location metadata is
honored, normalized, preserved, or stripped.
## Missing Everyday Features
These have higher value than adding more isolated AI tools.
### P1: Photo Correction
- Composite and per-channel histogram with clipping warnings
- Adjustment presets, reset, and non-destructive before/after compare are implemented
- Actual adjustment layers that affect content below, can be clipped/grouped,
and have their own masks; the current per-raster-layer stack is not equivalent
### P1: Retouching and Paint Ergonomics
- Richer multi-stop and radial gradient controls for raster content
- Healing Brush distinct from source-free Spot Healing and Clone Stamp
- Blur and Smudge brushes
- Content-aware fill workspace built on the existing inpaint capability
### P1: Geometry, Text, and Layer Workflow
- Image Size supports staged pixel/percentage resizing, aspect locking, and interpolation choice
- Canvas Size supports staged pixel/percentage bounds with anchor control
- Align and distribute selected layers (implemented in the multi-selection bar)
- Optional canvas auto-select for visible layers (group-level hit testing remains)
- Paragraph text boxes, tracking, vertical alignment, font loading/fallback
status, and more reliable text hit testing
- Mask density, non-destructive feather, invert, disable, link/unlink position
behavior, mask-only inspection, and applying true layer masks are implemented;
applying a group mask still requires an explicit flattening workflow
## Missing Professional Features
These define the gap to Photopea/Photoshop rather than blocking a credible v1.
### P2: Non-Destructive Core
- Embedded source layers / smart-object equivalent (toolbar/gallery/drop imports, pasted selections, and manual raster-to-source conversion now exist; richer source editing remains)
- Editable transform matrices that preserve original pixels through repeated
scale and rotate operations are now covered for placed and converted raster layers
- Linked instances and replace-source workflow
- Editable filter stacks with visibility, opacity, reorder, masks, and cached
previews
- Blur, sharpen, denoise, high pass, lens correction, and perspective correction
as retained filters
- Non-destructive transform masks, including perspective/warp later
This is the most important architectural gap. Photopea's Smart Objects retain a
separate source so repeated transforms can be recalculated without cumulative
loss, and its Smart Filters remain editable. Krita similarly models transform
and filter masks as non-destructive layer children.
### P2: File and Color Fidelity
- Explicit sRGB conversion and ICC profile awareness
- 16-bit processing before considering 32-bit/HDR
- EXIF/IPTC preservation or intentional stripping controls
- Reliable HEIC/TIFF handling and a RAW handoff/development path
- PSD import/export feasibility and a published compatibility matrix
- DPI/PPI and print-size metadata
### P2: Vector and Layout Work
- Shape layers for rectangle, ellipse, line, and custom paths
- Pen tool, editable Bezier paths, vector masks, and path-based selections
- Layer styles such as stroke, shadow, glow, and overlays
- Channels panel and channel operations
- Artboards only if multi-output design work is a product goal
### P3: Production Workflow
- Actions/macros, batch processing, and batch export
- Templates, reusable presets, and layer comps
- Soft proofing, gamut warning, and print output
- Plugin/filter extension surface
- Version history beyond the local bounded undo stack
## UX Audit
1. **The toolbar prioritizes AI before correction fundamentals.** Healing,
Eyedropper, Gradient, and Curves should be as discoverable as SAM and Inpaint.
2. **The distinction between masks is still cognitively expensive.** Selection,
Quick Mask, AI masks, layer masks, and group masks need consistent names,
thumbnails, active states, and properties rather than relying on sub-row
position alone.
3. **Properties are fragmented.** Tool controls, layer adjustments, transform
values, and mask settings should use one contextual Properties area. This
reduces modal popups and makes the selected target obvious.
4. **The editor needs clearer destructive-action signaling.** Blur, rasterize,
merge, and applied transforms should say when source pixels will be replaced,
with a one-step duplicate/convert-to-source option where appropriate.
5. **Mobile needs a deliberate mode.** Shrinking desktop controls is not enough;
canvas-first editing needs bottom-sheet properties, stable touch targets,
stylus behavior, and predictable two-finger navigation while a tool is active.
The tool and Layers sheets now claim the viewport exclusively, and mobile
layer/mask rows have stable non-scrolling layouts. Paint, crop, transform,
mask save/reopen, mask movement, and layer reorder gestures are covered;
mobile export now has viewport and download coverage, while broader
multi-tool touch workflows remain.
## Credible V1 Definition
A dependable everyday editor is reached when all of these workflows pass as a
single checked-in desktop and touch gate:
1. Import a common web image, correct exposure/color, retouch a blemish, crop,
resize, add text, and export at a chosen size and quality.
2. Build a layered composition with nested groups, clipping, a linked mask and
an independently positioned mask, then close and reopen without state loss.
3. Apply, cancel, undo, and redo each geometry/filter operation without changing
unrelated pixels or retained metadata.
4. Compare the flattened visible composite, reopened composite, and exported
bitmap within a documented pixel tolerance.
5. Complete the same core workflow with mouse and touch, with no inaccessible
controls, accidental page gestures, or silent partial edits.
6. Reject oversized, corrupt, or unsupported documents with a useful warning
while retaining every recoverable layer.
## Recommended Build Order
### Milestone 1: Reliability Gate
- Expand composite-versus-export pixel tests across groups, clipping,
adjustments, transparency/mattes, JPEG, and WebP.
- Add the operation/target matrix for masks, text, groups, clipping, locks, and
multi-select.
- Add Chromium touch/mobile workflows and basic Firefox/WebKit smoke coverage.
- Centralize document/layer/mask coordinate conversion.
- Define import/export color and metadata policy.
**Exit:** a mixed 20-edit document survives undo/redo, close/reopen, and export
with equivalent visible pixels and editable state.
### Milestone 2: Everyday Photo Workflow
- Extend raster retained gradients with richer stop editing
alongside the existing Eyedropper, Gradient, Healing, and Smudge tools.
- Add Curves, Exposure, White Balance, Vibrance, and Shadows/Highlights.
- Finish mask properties, Image/Canvas Size dialogs, text layout, alignment, and
brush presets.
**Exit:** crop, correction, blemish removal, annotation, transparent assets, and
social-image composition can be completed locally without another editor.
### Milestone 3: Non-Destructive Editing
- Introduce source layers and editable transform matrices.
- Promote adjustments into real adjustment layers.
- Add retained filter stacks and filter masks.
**Exit:** normal experimentation no longer requires manually duplicating layers
to protect the original pixels.
### Milestone 4: Interchange and Scale
- Add worker/offscreen rendering, cancellation, telemetry, and stress tests.
- Implement color-profile and metadata policy.
- Run a PSD/HEIC/TIFF/RAW feasibility spike and publish compatibility limits.
## Architecture Direction
The current `layer.canvas + optional metadata` model is reaching its limit.
Before adding adjustment layers, vectors, or source layers, move to a typed
document contract:
```text
Layer = RasterLayer | TextLayer | GroupLayer | AdjustmentLayer | SourceLayer
common: id, name, visible, opacity, blendMode, locks, masks, transform
raster: mutable pixel surface
text: content and typography
group: ordered child ids
adjustment: operation, parameters, clipping, mask
source: immutable embedded pixels, editable transform, filter stack
```
Rendering, serialization, history, geometry, duplicate, merge, and thumbnails
should dispatch through that contract. Coordinate conversion should likewise be
centralized around document, layer, mask, and viewport spaces instead of being
reimplemented inside tools.
History should evolve toward commands plus periodic checkpoints. Bounded full
RGBA snapshots prevent runaway memory today, but remain expensive for large
documents and awkward for retained non-raster layer types.
## Benchmark Basis
This audit uses mature editors as behavior references, not as a requirement to
clone every feature:
- [Photopea feature map](https://www.photopea.com/learn/)
- [Photopea masks and mask properties](https://www.photopea.com/learn/masks)
- [Photopea adjustment layers and Smart Filters](https://www.photopea.com/learn/adjustments-filters)
- [Photopea Smart Objects](https://www.photopea.com/learn/smart-objects)
- [Krita transform masks](https://docs.krita.org/en/reference_manual/layers_and_masks/transformation_masks.html)
- [Krita non-destructive filters](https://docs.krita.org/en/reference_manual/filters.html)
- [Photoshop color-adjustment workflow](https://helpx.adobe.com/photoshop/using/color-adjustments.html)
- [Photoshop Content-Aware Fill](https://helpx.adobe.com/photoshop/desktop/apply-painting-techniques/fill-objects-selections-layers/content-aware-fills.html)
## Immediate Next Slice
Extend the new fidelity/mobile gate across **grouped blending, adjustments,
and real touch paint/transform gestures** next. Selection copy, cut, paste, and
single-step undo now have checked coverage.
After that, implement **Gradient + Healing Brush** as one vertical
everyday-photo slice rather than adding disconnected controls.
Binary file not shown.
+7 -1
View File
@@ -736,6 +736,12 @@ Key settings:
All upload-limit vars are validated (must be a positive integer) and optional; an invalid value fails fast at startup.
The table above is the short list. The source tree reads a lot more `ODYSSEUS_*`
variables than this - browser automation, model routing, workspace mounts, media
limits, a few security switches - and the complete list, generated from the code
with the default each one falls back to, is in the
[configuration reference](configuration-reference.md).
### Built-in MCP servers (optional setup)
Odysseus auto-registers a few built-in MCP servers at startup. The npx-based ones (currently the browser server, `@playwright/mcp`) only start when their npm package is already in the local npx cache. If a package isn't cached, that server is skipped with a startup log message explaining what to do, so a fresh install does not block on a multi-minute npm download or hang if Playwright system deps are missing.
@@ -756,7 +762,7 @@ src/ llm_core, agent_loop, agent_tools, chat_processor, search/
routes/ chat, session, document, memory, model … endpoints
services/ docs, memory, search, hwfit (Cookbook) …
static/ index.html + app.js + style.css + js/ (modular front-end)
website/ landing page (index.html) + preview clips
website/ text-only landing page (index.html)
```
## Data
@@ -0,0 +1,220 @@
# SFT Expansion Launch Ledger
## Immutable Inputs
- Approved seed manifest: `tmp/sft_expansion_20260830/frozen_v1/seed_manifest.json`
- Approved seed turns: `tmp/sft_expansion_20260830/frozen_v1/approved_trace.jsonl`
- Environment inventory: `tmp/sft_expansion_20260830/environment_inventories.json`
- Seed baseline: 561 sessions, 1,372 approved turns
- Expansion selection: `tmp/sft_expansion_20260830/selected_200_manifest.json`
- Selection size: 50 owner-bound families, projected to 200 environment cases
## Acceptance Policy
1. Generate only from inventory facts and marker-scoped reversible mutations.
2. Reject unsupported tools, temporal contradictions, vague calendar mutation times, and compound mutation-plus-verification prompts.
3. Run each turn through the real 7011 agent surface with Kimi K3.
4. Require the expected tool and manager action, successful tool output, a nonempty answer, and no internal narration or false-unavailability claim.
5. Restore owner state after every case and delete mechanically failed sessions.
6. Review every mechanically passing session with DeepSeek.
7. Retain only `keep` verdicts scoring at least 80. Delete all other sessions and trace rows; repair by regenerating and replaying, never by inventing missing tool evidence.
8. Preserve the source seed family's train/validation/test split.
## Pilot Record
- Harness regression fixed: contextual calendar entries no longer route to `manage_tasks`.
- Harness regression fixed: contextual calendar reads require fresh `manage_calendar` evidence.
- Runner regression fixed: a manager tool name alone no longer passes; create/list/update/delete actions are checked.
- Generator regression fixed: calendar mutations require an exact time or `ask_user`.
- Unsafe global skill-directory rollback replaced with marker-scoped cleanup.
- Omar calendar lifecycle: deterministic pass, DeepSeek `keep`, 96/100, retained.
- Ambiguous Omar predecessor: DeepSeek `repair`, 60/100, deleted from app and trace.
- Maya ambiguous-date lifecycle: deterministic reject, deleted before review.
## Current Corpus
- Build: `tmp/sft_expansion_20260830/corpus_v2`
- Train: 994 turns
- Validation: 157 turns
- Test: 135 turns
- Sessions: 562
- Retained expansion sessions: 1
- Seed-family leakage: 0
## Clean Source Baseline
- Immutable source remains unchanged: `tmp/sft_expansion_20260830/frozen_v1/approved_trace.jsonl`
- Deterministic hygiene input: 561 sessions, 1,372 turns
- Deterministic hygiene result: 421 sessions, 859 turns retained; 140 sessions rejected
- Full DeepSeek semantic review: 336 keep, 77 repair, 8 delete
- Kimi repair result: 73 validated repaired sessions; 4 additional sessions excluded
- Clean derivative: `tmp/sft_expansion_20260830/clean_v2/approved_trace.jsonl`
- Clean derivative size: 409 sessions, 834 turns
- Post-repair deterministic recheck: 409/409 sessions pass
- Style contract: `docs/sft-style-contract.md`
- Assembly report: `tmp/sft_expansion_20260830/clean_v2/assembly_report.json`
- Semantic verdicts: `tmp/sft_expansion_20260830/clean_v1/deepseek_verdicts.jsonl`
The clean source does not yet support the intended balanced expansion by itself.
Notes has one clean source session, tasks has two, and session-management tools
are absent. Add and review natural live workflows for those domains before the
200-case cross-environment run.
## Commands
```bash
.venv/bin/python scripts/generate_sft_environment_expansion.py \
--seed-manifest tmp/sft_expansion_20260830/selected_200_manifest.json \
--inventories tmp/sft_expansion_20260830/environment_inventories.json \
--out tmp/sft_expansion_20260830/generated_200_cases.json \
--workers 4 --timeout 180 --retries 1
.venv/bin/python scripts/run_sft_environment_expansion.py \
--cases tmp/sft_expansion_20260830/generated_200_cases.json \
--out tmp/sft_expansion_20260830/generated_200_run.json
.venv/bin/python scripts/review_sft_environment_expansion.py \
--run tmp/sft_expansion_20260830/generated_200_run.json \
--out tmp/sft_expansion_20260830/generated_200_review.json
.venv/bin/python scripts/build_sft_expansion_splits.py \
--manifest tmp/sft_expansion_20260830/frozen_v1/seed_manifest.json \
--approved-trace tmp/sft_expansion_20260830/frozen_v1/approved_trace.jsonl \
--review tmp/sft_expansion_20260830/pilot_atomic_exact_v1.review.json \
--review tmp/sft_expansion_20260830/generated_200_review.json \
--out-dir tmp/sft_expansion_20260830/corpus_expanded
```
# All-Tools No-Thinking Validation Update
## Matched-control finding
The previously reported email LoRA speed advantage does not reproduce when the
raw base model and LoRA are served in the same vLLM process with the same GPU,
tool payload, prompt, cache settings, and `enable_thinking=false`.
- Original 24-case email suite, 14 compact tools:
- raw Qwen3.5-9B base: 23/24, 0.914s average
- all-tools iteration-2 LoRA: 23/24, 0.874s average
- Earlier base result of 9.655s was therefore confounded by serving conditions.
- The validated latency win is **thinking off + compact schemas**. LoRA has not
yet shown a material independent speed gain.
## All-tools iterations
- Iteration 1: 365 train sessions, 40 validation sessions, 92 steps, LR 5e-7,
one epoch. Eval loss 0.8364.
- Iteration 2: same corpus, 184 steps, LR 2e-6, two epochs. Eval loss 0.7433.
- Matched 40-case held-out routing result for base, v67, and iteration 2 was
effectively identical: 85% exact tool selection, 85% required-field
presence, 70% constrained required-argument accuracy, and no reasoning
leakage.
- Iteration 2 adapter weights differ from v67 (relative L2 delta 1.02%), so the
tie is not caused by an unchanged adapter file.
## Decision
Do not promote iteration 1 or 2 as a speed improvement. Before another train:
1. Build a larger, balanced, genuinely unseen eval across sparse tool families.
2. Expand sparse training families with multi-turn, context-dependent traces.
3. Evaluate whether LoRA preserves accuracy under more aggressive schema
pruning than base; that is the plausible route to an indirect latency win.
4. Keep no-thinking routing and compact schema selection as harness features,
independent of model promotion.
## Critical Qwen3.5 LoRA Serving Failure
### Symptom
vLLM 0.22 accepted `--enable-lora`, listed each adapter under `/v1/models`, and
logged `Loaded new LoRA adapter`, but Qwen3.5 text-tool adapters were not applied
to inference. Base, v67, iteration 3, and iteration 4 produced byte-identical tool
calls across 177 cases. A direct log-probability probe also returned numerically
identical values for base and every dynamic adapter.
Do **not** treat model registration, successful HTTP responses, different adapter
files, or latency differences as evidence that a LoRA is active.
### Minimal activation check
Before any benchmark, send the same deterministic request to base and adapter
models with `temperature=0`, `logprobs=true`, and thinking disabled. Compare the
token log-probabilities, not only generated text. Identical probabilities across
several prompts mean the adapter path is inactive.
Observed inactive result:
```text
base: cob -0.0484729, alt -0.0000684
v67: cob -0.0484729, alt -0.0000684
iteration4: cob -0.0484729, alt -0.0000684
```
Disabling prefix caching and enabling eager execution did not fix this.
### Merge trap
The tool adapters were trained with `AutoModelForCausalLM`, which gives Qwen3.5
text keys under:
```text
base_model.model.model.layers.*
```
The old merge script auto-selected `AutoModelForImageTextToText`, whose language
tower expects:
```text
base_model.model.model.language_model.layers.*
```
PEFT warned about missing adapter keys, then still wrote a model. That output was
effectively unmodified. A successful `save_pretrained` is therefore not proof of
a successful merge. Treat any missing-adapter-key warning as a hard failure.
Merging with the causal loader applied the adapter, but produced a
`Qwen3_5TextConfig` model that vLLM's production Qwen3.5 multimodal loader refused.
### Working merge path
`scripts/merge_hf_lora.py` now supports the production-safe bridge:
1. Load the full Qwen3.5 model with the image/multimodal loader.
2. Remap causal adapter keys from `model.layers` to
`model.language_model.layers`.
3. Exclude the visual tower from PEFT target matching.
4. Merge and save the full model plus tokenizer and processor metadata.
```bash
python scripts/merge_hf_lora.py \
--model-loader image \
--remap-qwen35-causal-adapter \
--base /mnt/HADES/models/Qwen3.5-9B \
--adapter /path/to/final_adapter \
--output /path/to/merged-full-bf16
```
Verify there are no missing adapter keys, serve the merged model as a standalone
model, and repeat the log-probability activation check.
Observed active merged result:
```text
v67 merged: cob -0.471937001
iteration4 merged: cob -0.319356978
```
### Corrected valid benchmark
The first valid merged-model all-tools benchmark showed:
- v67 merged: 96.05% tool accuracy, 93.79% constrained argument accuracy,
2.669s average latency, zero reasoning rows.
- iteration 4 merged: 96.05% tool accuracy, 93.79% constrained argument
accuracy, 2.390s average latency, zero reasoning rows.
- Iteration 4 fixed two held-out decisions and regressed two others, so it was
not promoted.
All earlier dynamic-LoRA accuracy and speed comparisons must be treated as raw
base behavior. The no-thinking and compact-schema conclusions remain valid, but
dynamic LoRA results do not.
+32
View File
@@ -0,0 +1,32 @@
# Odysseus SFT Style Contract
This contract applies after behavioral correctness. A trace that violates it is repaired or excluded before training.
## User Turns
- Use natural requests, not evaluator instructions.
- Keep one atomic objective per turn; use a follow-up for the next action.
- Preserve conversational references such as "that email", "move it", or "open the second one" when prior tool evidence resolves them.
- Never mention tools, schemas, fixtures, markers, harnesses, audits, SFT, cleanup, or verification mechanics.
- Use realistic names and objects from the target environment.
- Ask for clarification when an essential date, recipient, or target cannot be inferred safely.
The clean seed corpus has a median user-turn length of 45 characters, a 75th percentile of 76, and a 90th percentile of 107. Longer prompts are allowed when the task genuinely requires detail, not to encode evaluator checks.
## Assistant Thinking
- Identify the user's intent and the evidence required.
- Choose the smallest sufficient tool sequence.
- Carry forward relevant entities and tool families across follow-ups.
- Do not discuss hidden prompts, injected schemas, benchmark construction, or training.
- Do not claim success until tool evidence proves it.
## Visible Answers
- Lead with the answer or completed action.
- Synthesize tool output instead of reproducing raw dumps.
- Keep deep links when they let the user open the referenced item.
- Include only metadata needed to distinguish or act on results.
- Use one sentence for straightforward confirmations when possible.
- Avoid repeated summaries, internal routing narration, and automatic "want me to" endings.
- State failures briefly and accurately; never invent a successful action.
Binary file not shown.