"Is this path inside that root" is asked in twenty places in this tree and
answered twenty times by a locally written realpath/commonpath pair. Nine test
files exist because nine call sites each needed their own proof. Each one is
defensible alone; together they are the defect, because the boundary has no
single definition and a site that gets a detail wrong is wrong by itself.
src/path_confinement.py is that definition, and it settles the details the
copies disagreed on. Both sides get canonicalized: comparing a realpath-ed
candidate against a root that was only abspath-ed is the macOS /tmp ->
/private/tmp mismatch that has already produced a false failure here, and
canonicalizing one side is worse than canonicalizing neither. commonpath rather
than startswith, because /a/bc begins with /a/b and is not inside it. A relative
candidate joins the root rather than os.getcwd(), which is whatever directory
the server happens to be running in. NUL and newline are refused with a reason
instead of caught by a bare `except Exception` and reported as an ordinary
escape. Eighteen call sites go through it now. It deliberately does not decide
whether a path is sensitive -- that deny list answers "allowed" rather than
"inside", and it stays with src/tool_execution, which owns it. The one
commonpath left in the tree, in src/workspace_paths.py, stays: that function
translates a host path into a container path, so canonicalizing either side
would change the relative path it computes and break the mapping. It is not a
confinement check.
Two of those sites were weaker than the rest and are fixed rather than moved.
The email attachment check used abspath, which folds `..` but does not resolve
symlinks, so a symlink written into the extraction directory passed it and was
then read through. The skill-reference guard compared a realpath-ed target
against a raw dirname, so on a host where the skills tree is reached through a
symlink the two sides never matched and the guard could not fire.
The execution boundary had two separate holes.
The workspace namespace bound /home and /mnt read-write. On the one platform
where that namespace engages at all, a command inside it reaches outside the
workspace and writes to the user's home directory -- measured by running this
argv on a Linux host with working bubblewrap, not inferred from the source.
Binding the user's whole home directory into a workspace-confinement namespace
gives back most of what the namespace was for. Both are read-only now. The
workspace is also bound writable at its real host path, not only at /workspace:
BashTool's own /tmp redirect rewrites `/tmp/` to `<agent_cwd()>/.tmp/` before
the namespace is built, so the command bwrap receives already names the real
path, and those writes previously landed only because the workspace happened to
sit under the writable /home.
`namespaced or _replace_workspace_alias(...)` chose between a mount namespace
and a regex with nothing in the result saying which one ran. The fallback
rewrites the literal token /workspace in the command string, so a command that
never mentions /workspace is untouched by it and runs on the host unrestricted
-- which is every agent shell command on macOS. Both tools now ask
containment.probe() instead of each deciding for itself, and every bash and
python result carries a containment block naming the mechanism and stating
whether the filesystem dimension actually held. Under enforcing mode the
command is not run and the result says so.
That block reports the filesystem dimension only, and says so in a
reported_dimensions field. The probe knows this host could also give a process
group and a real wall clock, but these two tools still assemble their own
create_subprocess_* call and pass neither, so listing those dimensions would be
exactly the false claim src/containment.py calls worse than an honest absence.
probe() is new on src/containment.py: the same mechanism table and the same
arithmetic as acquire(), stopping before the side effects. acquire() is the
wrong shape for a decision -- it writes a durable grant record, and a record
whose pid is never filled in and whose release() never runs is an entry a
restart reaper keeps finding.
CONTAINMENT_MODE stays report_only. Flipping it refuses every agent shell
command on macOS and on any Linux host without bubblewrap, which is a product
decision rather than a code one.
Smaller things in the same area: the /tmp redirect's makedirs was unguarded, so
a read-only workspace turned a command that merely mentioned `/tmp/` into an
OSError traceback instead of a tool error; it degrades now. WORKSPACE_MOUNT
moved to src/constants.py so the namespace and the path resolvers read one
definition of the contract rather than two. The ".tmp" dirname got a constant,
since it appeared in both tool paths.
One generated artifact moved with it: website/configuration-reference.md pins
the source line where each ODYSSEUS_* variable is read, and three of those
shifted. Regenerated with scripts/generate_env_reference.py; the diff is line
numbers only.
Three existing tests changed. test_workspace_artifact_tool_floor asserted that
an unsafe interpreter prefix produces no `--ro-bind <prefix> <prefix>`, which
now fires on /home because /home is legitimately a read-only base mount.
Asserting the absence of a literal flag string cannot distinguish "the prefix
was rejected" from "the argv mounted that root itself", so it compares the argv
against the no-prefix baseline instead: an unsafe prefix must add nothing.
The Windows bash test asserted dict equality on the
whole result, which makes adding a field to every bash result impossible without
touching a test about tmux; it asserts the shape now. The personal-dir symlink
test grepped the resolver's source for the literal "os.path.realpath", which is
gone because the resolution moved into the shared boundary -- it keeps the
negative assertion that the closure must not grow its own abspath check again,
and the behavioural half now runs against the boundary, where it covers every
call site instead of one closure.
Not verified: the bubblewrap argv is asserted, not executed. There is no bwrap
on macOS, and in Docker it needs --privileged to work at all -- default and
seccomp=unconfined both fail with "Creating new namespace failed", and
--cap-add=SYS_ADMIN fails at pivot_root. The Python tool's
needs_virtual_namespace gate means ordinary Python code gets no namespace even
on a Linux host that could provide one; that is reported now but deliberately
not changed, because it alters the Linux Python path on every call and cannot be
checked from here.
Snapshot current maintainer-preview application changes and regression fixtures for integration into lab. Excludes local runtime data, evaluation outputs and source backups. Focused Python regression selection: 140 passed; full suite not certified.
- Enforce outbound URL policy on all hops via bounded manual redirects in scholarly lookups
- Guard declared Content-Length in EditorDraftRoute before request body parsing
- Keep scholarly lookup timeouts and budgets as internal constants rather than surface env vars
- Narrow TUI threat-model description around demonstrable host shell bridge behavior
- Add behavioral regression tests for redirect security and pre-parsing body size guards
The arXiv and OpenAlex title resolvers called two hardcoded endpoints with a
hand-written "Odysseus/0.20" User-Agent, no outbound-URL policy, and a full
12s timeout each. APP_VERSION was already 1.0.3, so the agent string was wrong
the moment it was written, and a scholarly query could hold a user-facing
search open for the sum of all three hops.
Move the endpoints to named constants, build the User-Agent from APP_VERSION,
run both calls through check_outbound_url (they follow redirects, so the final
host is not the one in the constant), and give the whole SearXNG -> OpenAlex ->
arXiv chain one shared wall-clock budget.
The budget is a ContextVar rather than a parameter so the existing test doubles
for _openalex_title_results and _arxiv_title_results keep working unchanged.
* fix(skill-importer): validate URL scheme and improve skills.sh handling
* fix(skill-importer): enhance DNS resolution and SSRF protection in fetch URL handling
* fix(url-safety): add allowed_dist parameter to check_outbound_url for flexible private blocking
* test(skill-importer): add comprehensive tests for URL parsing and outbound checks
* ensure newline at end of file in test_check_outbound_url_allows_public_ip
* fix(skill-importer): improve TLS certificate handling in _get_checked function
* fix(skill-importer): enhance _check_fetch_url to handle both hostnames and full URLs
* fix(skill-importer): enhance parse_skill_source to support skills.sh URLs in path and netloc
* fix(skill-importer): simplify skills.sh hostname check in parse_skill_source
* fix(skill-importer): enhance parse_skill_source to identify skills.sh URLs in path and handle localhost/IP addresses
* fix(skill-importer): enhance _resolve_and_check_url to validate all resolved IP addresses and prevent TOCTOU vulnerabilities
* fix(skill-importer): enhance parse_skill_source to support schemeless GitHub and skills.sh URLs
* fix(memory): resolve CodeQL URL sanitization warning and restore _check_fetch_url test alias
* fix(memory): pin skill fetch sockets without rewriting URLs
* fix(memory): reject unsupported skill wrapper hosts
* refactor(url-safety): remove unused importer exception
* test(memory): keep redirect regression hermetic
* test(dns-rebinding): add test for _PinnedTransport to ensure connection to pinned IP
* fix(skill-importer): enhance skills.sh support to extract GitHub links from page content
* fix(skill-importer): improve URL scheme validation for GitHub and skills.sh links
* fix(skills): reject unusable skill URLs instead of guessing
Resolving a skills.sh link by scraping the first github.com URL out of
the page body cannot work. Skill pages only ever link the repository
root, never the skill's subdirectory, so every skill in a repo resolved
to the same bundle: importing skills.sh/anthropics/skills/pdf walked the
whole monorepo, saturated the 64-file cap, and installed algorithmic-art
behind an ok:true response. Restore the redirect-target unwrap and fail
with a message that says what to do instead.
Also report the real reason a URL is rejected. The scheme check keyed off
"://" appearing anywhere in the string, so a supplied-but-unusable URL
came back as "URL is required", and a schemeless URL carrying "://" in
its query was reported as an unsupported scheme. Key off the parsed
scheme and let opaque schemes (mailto:, javascript:) and a schemeless
host:port fall through to the host check.
* test(skills): tighten the real-socket pinning regression
The handler swallowed its own exceptions, so a failure inside it
surfaced as a confusing assertion on the captured client address.
Record the exception and assert on it, run the thread as a daemon, and
close the listening socket from the test so a hang cannot outlive the
run. Also drop the duplicate ipaddress import and the missing newline.
* fix(skills): require exact GitHub skill URLs
* test(skills): read complete pinned request headers
---------
Co-authored-by: RaresKeY <158580472+RaresKeY@users.noreply.github.com>
Co-authored-by: Léo <leograndcontact@gmail.com>
_emit_scalar quotes a frontmatter scalar with json.dumps when it holds
punctuation that would change how the line reads back. _parse_scalar undid that
with a bare raw[1:-1]: it stripped the quotes but never decoded the escapes. So
a description containing ü was written as the escape sequence \u00fc, read
back with that escape still sitting literally in the value, and re-escaped on
the next save. The backslash run doubles every save, so a non-English skill
description degrades into backslash noise after a few edits, and the escapes are
shown verbatim in the skills list and the /skills catalog.
This is not limited to non-ASCII. Any description containing a quote takes the
same path, since the quote is itself what forces the quoted form.
Make the two halves symmetric: emit with ensure_ascii=False, since SKILL.md is
UTF-8 at both ends (skills.py reads it, atomic_write_text writes it) and the
ASCII-escaped form bought nothing; and parse double-quoted scalars with
json.loads, falling back to the previous literal reading when the value is not
valid JSON. Files already corrupted heal one level per load.
ensure_ascii=False on its own would open a smaller hole. json.dumps escapes
every C0 control character but passes NEL, LINE SEPARATOR and PARAGRAPH
SEPARATOR through literally, and parse_frontmatter reads one scalar per line via
str.splitlines(), which breaks on all three. Re-escape those three, and add them
plus the remaining splitlines characters to the set that forces a quoted scalar,
so none of them can reach the file bare.
Co-authored-by: Alexandre Teixeira <111787685+alteixeira20@users.noreply.github.com>
* fix(memory): don't let an unreadable store get overwritten with an empty one
load_all() answered a failed read the same way it answered an empty store:
with []. Every mutation path is a read-modify-write (load the whole file,
change it, save it back), so a failed read became
load_all() -> [] -> [].append(new) -> save([new])
and save() is atomic, so the replacement stuck.
The case that actually destroys data is a store that is READABLE but not
parseable - a truncated file, or one holding {} instead of []. Nothing
obstructs the write, so adding a memory returns HTTP 200 and every memory
already stored is gone. Verified end-to-end against a running instance: on the
current code a truncated memory.json plus one add leaves the file holding only
the new entry. Truncation is reachable - core/database.py rewrites memory.json
during migration with a plain open(.., "w") + json.dump, which is not atomic.
A live exclusive lock is not the dangerous case: it blocks the read and the
os.replace alike, so the save fails too and the store survives. That path
currently 500s and loses nothing.
_read_entries() now returns [] only when the file genuinely does not exist and
raises MemoryStoreUnreadable for every other failure, including a store that
parses but is not a JSON array. load_all() keeps the old lenient behaviour so
display, search and context injection still degrade quietly instead of
breaking chat. The read-modify-write callers switch to load_all_for_update(),
which propagates the error: the memory routes turn it into a 503 and change
nothing, backup import refuses rather than saving only the incoming rows, and
auto-extraction and the audit merge skip the write. The audit merge mattered
most - it rebuilds the whole file from one owner's slice plus everyone else's
rows, so an empty read there dropped every other tenant's memories.
The corrupt-JSON path still gets its one shot at the legacy memory.txt
migration before raising, so that recovery is unchanged.
The two updated fakes gained load_all_for_update because the real class has it;
MagicMock would otherwise hand the import path a Mock instead of the seeded list.
Fixes#5673
* fix(memory): fail closed on the remaining read-modify-write add paths
The strict loader landed with the routes, the backup import and the extractor
converted, but three read-modify-write sinks still called load_all(), which
degrades an unreadable store to []. Two of them are the paths users actually
reach, so the data loss in #5673 stayed reproducible:
- src/ai_interaction.py do_manage_memory, action "add" — reached from ordinary
chat via src/tool_execution.py:793 -> dispatch_ai_tool. "Remember that I
prefer X" against an unreadable store wrote a one-entry file over it and
reported success.
- mcp_servers/memory_server.py, action "add" — the same shape through
_scope_entries(), registered as a built-in in src/builtin_mcp.py.
- src/memory_provider.py NativeMemoryProvider.remember and .delete — wired
into app state in src/app_initializer.py but not consumed outside tests yet,
converted here so the pattern is uniform before it goes live.
The MCP server takes _scope_entries(for_update=True) so list keeps the lenient
read. The edit and delete branches on both tool paths were already fail-closed
by accident — an empty view matches nothing and returns before the save — so
they are left alone.
The three new tests drive the real entry points rather than replaying the
shape, and use a truncated store, which is the case that reads back fine so
nothing stops the save. Each asserts memory.json is byte-identical afterwards;
all three fail on the previous commit with the store overwritten.
* fix(skills): replace deprecated utcnow in skill timestamp helper
_now_iso() builds the 'created' value in skill frontmatter. datetime.utcnow()
returns a naive datetime and has been deprecated since Python 3.12, scheduled
for removal. Switch to the timezone-aware datetime.now(timezone.utc), keeping
the serialized YYYY-MM-DDTHH:MM:SSZ shape unchanged so existing skill files
keep parsing.
timezone.utc is used rather than the datetime.UTC alias, which is 3.11+ only.
Adds regression tests covering the deprecation, the serialized shape, and
UTC correctness under a non-UTC local timezone -- the last guards against a
bare datetime.now(), which yields the same shape but local wall time.
Fixes#5697
* test(skills): skip timezone mutation where unsupported
---------
Co-authored-by: Alexandre Teixeira <alexandremagteixeira@gmail.com>
* Harden skill importer against SSRF: block private targets + revalidate redirects per hop
The skill importer validated only the initial URL with the lenient SSRF guard
(block_private=False) and then fetched with follow_redirects=True, so a 3xx to
an internal/metadata address (169.254.169.254, 127.0.0.1, RFC-1918) was still
connected to — inconsistent with the hardened services/search/content.py
:_get_public_url path.
Add a _get_checked() helper that follows redirects manually and re-runs the
SSRF guard with block_private=True on every hop, and route all three fetch
sites (skills.sh unwrap, _fetch_bytes, _list_github_dir) through it. GitHub's
own redirects and the final-host _assert_github_url checks are preserved.
Adds hermetic regression tests (IP-literal hosts, faked HTTP layer) and updates
the existing mock signature for the new block_private kwarg.
Defense-in-depth: the endpoint is admin-gated (require_admin) and admins are
trusted per THREAT_MODEL.md, so this is not a cross-boundary vulnerability.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test: enforce follow_redirects=False invariant in mock client
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
_extract_entities used the ASCII-only class [A-Z][a-zA-Z]+ to pull name
entities from a query, so non-ASCII names were dropped ("İstanbul",
"Zürich" yielded nothing) or shredded ("São Paulo" -> only "Paulo"),
degrading query enhancement for non-English/accented searches. Match
Unicode words and keep the alphabetic, uppercase-initial ones; ASCII
behaviour (the word boundary already excludes camelCase mid-word
capitals) is unchanged.
* fix(search): pin DNS-validated fetch connections
Rebase the DNS-rebinding SSRF fix onto current dev after search content moved behind the services.search.content canonical module.
Integrate the pinned httpcore NetworkBackend/BaseTransport approach with the current size-capped Client.stream fetch path, preserving Host/SNI semantics while forcing TCP connect to the already validated resolved IP.
Keep src.search.content as the compatibility wrapper and preserve existing OG-image http(s) behavior; this avoids reintroducing the unrelated scope changes that previously blocked review.
Add the explicit httpcore>=1.0,<2.0 requirement used by the public httpcore NetworkBackend and ConnectionPool APIs.
* test(search): restore and rebase DNS rebinding regressions
Keep the current security regression coverage that the stale PR branch had deleted, including auth-disabled localhost bypass and Ollama cookbook hardening tests.
Carry forward the DNS-rebinding coverage for private resolve blocking, pinned TCP connect behavior, Host header preservation, redirect revalidation, and the BaseTransport/public-httpcore static guard.
Update redirect tests to mock the current Client.stream-based capped fetch path rather than the older httpx.stream/get path.
* test(search): adapt size-cap fetch tests to pinned client stream
The DNS-rebinding repair moved _get_public_url from the module-level httpx.stream shortcut to httpx.Client(...).stream(...) so the fetch can use the pinned transport.
Keep the existing size-cap test fakes by routing Client.stream through the monkeypatched httpx.stream only when a test has installed that fake; otherwise fall back to a real Client.
This fixes the CI failures in tests/test_web_fetch_size_caps.py without touching unrelated upload-handler atomicity behavior, which is already flaky on clean origin/dev.
---------
Co-authored-by: Alexandre Teixeira <alexandremagteixeira@gmail.com>
Add official Gemma 4 12B-it plus QAT-INT4/INT8 catalog entries (with their
GGUF sources), QAT quantization support across the quant tables and the
prequantized-prefix list, and the missing RTX 3050 / 3050 Ti memory
bandwidth so speed estimates stop falling back to the generic cuda value.
Remote Cookbook hwfit probes failed on Windows hosts because the PowerShell script was sent as nested -Command quoting through OpenSSH. Use -EncodedCommand for remote probes, auto-detect platform when omitted (including Darwin for Mac SSH hosts), and return a clearer error when SSH works but the probe fails.
Co-authored-by: Cursor <cursoragent@cursor.com>
Three classes of incorrect detection fixed:
(1) AMD GPU + no ROCm installed (e.g. Strix Halo) was reported as
backend=rocm everywhere, so launch commands emitted
HIP_VISIBLE_DEVICES (silent no-op on Vulkan) and the from-source
build path failed. Both _probe_amd_sysfs (routes/cookbook_routes)
and _detect_amd (services/hwfit/hardware) now probe rocminfo /
hipconfig / vulkaninfo at detection time and report vulkan when
only Vulkan is present.
(2) Build helper was picking the CUDA branch on AMD hosts whenever a
stray pip-installed nvcc was on PATH (vLLM wheels carry one
without libcudart). Added _odysseus_has_nvidia_hw() that checks
nvidia-smi / /dev/nvidia* / lspci, and gates both the nvcc PATH
augmentation and the CUDA elif branch on real hardware.
(3) Build chain reordered to ROCm/HIP > CUDA > Vulkan > CPU. Vulkan
tier added between CUDA and CPU as a portable fallback for hosts
with a GPU but no native toolchain (the common Strix Halo case).
Same _append_llama_cpp_linux_accel_build_lines also auto-attempts
sudo -n apt/pacman/dnf install of cmake/build-essential/git when
they are missing, surfacing a clear no-passwordless-sudo warning
otherwise.
* fix(hwfit): use CPU fallback for cpu_only speed estimates
* fix(hwfit): preserve ARM fallback for cpu_only estimates
---------
Co-authored-by: Cata <cata@bigjohn.local>