Commit Graph
1472 Commits
Author SHA1 Message Date
pewdiepie-archdaemon a48ab46f7d Enforce search source restrictions and expand live prompt coverage 2026-09-17 20:14:32 +00:00
pewdiepie-archdaemon 7fa6c93fb3 separate search freshness from news category selection 2026-09-17 20:08:58 +00:00
pewdiepie-archdaemon 566202dcad require domain evidence for official source shortcut 2026-09-17 20:07:51 +00:00
pewdiepie-archdaemon 6fb1ce7238 withhold tools for standalone social turns 2026-09-17 20:05:30 +00:00
pewdiepie-archdaemon dc4a38f3bc preserve manual formats and require discovery for evidence requests 2026-09-17 20:04:51 +00:00
pewdiepie-archdaemon 2d938b1650 recognize casual news requests and prevent premature source-only completion 2026-09-17 20:03:32 +00:00
pewdiepie-archdaemon a7576f721c verify embedded articles skip redundant retrieval rounds 2026-09-17 19:58:06 +00:00
pewdiepie-archdaemon 02200b2874 reuse embedded search articles and reject browser error pages 2026-09-17 19:57:20 +00:00
pewdiepie-archdaemon f27070c469 buffer rejected search drafts and cap interactive rounds 2026-09-17 19:53:40 +00:00
pewdiepie-archdaemon ff4d01a55e ground fresh searches and verify official documents 2026-09-17 19:49:28 +00:00
pewdiepie-archdaemon 8c7ea10520 ground current lookups and recover failed fetches 2026-09-17 19:42:38 +00:00
pewdiepie-archdaemon dfd4d5a80e preserve evidence after bounded web search 2026-09-17 19:35:54 +00:00
pewdiepie-archdaemon 295578514e route current events through bounded news search 2026-09-17 19:26:04 +00:00
pewdiepie-archdaemon e9a117df0d bound optional search fallback latency 2026-09-17 19:22:51 +00:00
pewdiepie-archdaemon 6b135eab73 finish resource lookups from verified search metadata 2026-09-17 19:18:05 +00:00
pewdiepie-archdaemon e49264ae47 preserve product identity in manual search 2026-09-17 19:15:39 +00:00
pewdiepie-archdaemon e48ea98d2f broaden empty searches through resilient providers 2026-09-17 19:12:28 +00:00
pewdiepie-archdaemon e22cf4b500 bound search attempts by research depth 2026-09-17 19:07:22 +00:00
pewdiepie-archdaemon 483fdc7054 require linked broad research synthesis 2026-09-17 19:04:08 +00:00
pewdiepie-archdaemon e3107a645d enforce distinct research refinement 2026-09-17 19:02:11 +00:00
pewdiepie-archdaemon 6121442658 unify broad web briefing semantics 2026-09-17 18:58:31 +00:00
pewdiepie-archdaemon be48da1146 recognize natural broad current queries 2026-09-17 18:55:19 +00:00
pewdiepie-archdaemon cdc0734347 require breadth for broad current research 2026-09-17 18:51:56 +00:00
pewdiepie-archdaemon 69a1b97c96 recover rendered pages from fetch boilerplate 2026-09-17 18:49:59 +00:00
pewdiepie-archdaemon 4d7c9d44d1 keep compact agent core tools available 2026-09-17 18:47:04 +00:00
pewdiepie-archdaemon 25ff725c1d repair broad web research recovery 2026-09-17 18:36:12 +00:00
pewdiepie-archdaemon 0aa470b095 recover synthesis after unavailable web search 2026-09-17 13:34:36 +00:00
pewdiepie-archdaemon 337a47d27d skip redundant synthesis after email actions 2026-09-17 13:21:11 +00:00
pewdiepie-archdaemon 10d637f8b9 ground research synthesis in retrieved source urls 2026-09-17 10:50:20 +00:00
pewdiepie-archdaemon 8ae0c31666 bound web research to retrieval and synthesis 2026-09-17 10:41:14 +00:00
pewdiepie-archdaemon 218d762427 Consolidate Odysseus agent harness and tool contracts 2026-09-17 10:07:40 +00:00
isharak7m 6001f82019 test: rewrite to exercise actual production _refresh_token_cache
The previous test used a _SharedCache simulation that proved the atomic
swap pattern works but didn't exercise the real app.py code. This rewrite
imports app.py with AUTH_ENABLED=true, mocks SessionLocal and logger,
creates a real AuthManager user, and calls the actual _refresh_token_cache()
while concurrent readers access the actual _token_cache global.

7 tests: single row, multiple prefixes, empty DB, app.state sync, dirty
flag cleared, 4 concurrent readers x 100 refreshes (zero empty reads),
and 50 create/revoke churn cycles with concurrent readers.
2026-09-13 19:19:12 +05:30
isharak7m 0b6d44890f test: add regression for token cache atomic swap race condition
Exercises the concurrent reader/writer scenario that the atomic swap
fix in app.py addresses. Uses a _SharedCache helper that mirrors the
module-level _token_cache global — both reader and writer access the
same .current reference, so the GIL-atomic swap is properly tested.

6 tests: swap correctness, concurrent readers (4 threads x 100 refreshes,
zero empty reads), app.state sync, multiple prefixes, empty DB, and
concurrent refresh from 4 threads.
2026-09-13 18:46:27 +05:30
Amir Fathi 9d5c031914 fix(mcp): reject malformed Args on Add MCP Server instead of silently defaulting to [] (#6215)
* fix(mcp): reject malformed Args on Add MCP Server instead of silently defaulting to []

* test(mcp): pass every Form param add_server reads past args validation

CI's pytest run showed test_add_server_still_accepts_valid_json_args and
test_add_server_still_defaults_empty_args_to_empty_list failing with
TypeError: the JSON object must be str, bytes or bytearray, not Form.

Calling the endpoint function directly bypasses FastAPI's dependency
resolution, so an unpassed Form(...) parameter (url, oauth_file,
oauth_config) arrives as the Form marker object itself rather than its
declared default, and add_server's later `if oauth_file:` check reads
that marker as truthy. The malformed-args test never hit this because it
raises before reaching that code. Not a production bug: a real HTTP
request resolves these through FastAPI before add_server ever runs.

* fix(mcp): reject non-list args and surface the new 400 in the Admin panel

o3LL's review on #6215 found two gaps in the args validation this PR adds:
the Admin panel posts to the same /api/mcp/servers endpoint but never
validates Args client-side, so the new 400 falls into the generic failure
branch and shows "Added but connection failed: unknown". Mirror the same
JSON.parse guard settings.js already has.

Also add an isinstance(list) check next to the existing JSON parse, since
valid-but-wrong-shaped JSON (args=5) reaches StdioServerParameters(args=5)
and 500s in the error formatter. Pre-existing on dev, same validation site
this PR already touches.

* fix(admin): surface the server's 400 detail instead of a generic connection-failed message

The Admin add-server handler read needs_oauth/connected/error but never
res.ok, so a request rejected by the isinstance(list) check added for
#6211 (args=5, a valid-JSON-but-non-list value the client-side JSON.parse
guard cannot catch) fell into the same-shape else branch as a successful
add whose connection attempt failed, and the form fields were cleared as
if the server had accepted it.
2026-09-11 15:36:41 +02:00
pewdiepie-archdaemon 84aa9a91de Squash Odysseus development history 2026-09-11 06:04:19 +00:00
nopozandRaresKeY 934d23c0be Merge commit from fork
* fix(security): keep agent file tools out of the app state directory

The agent's read tools (read_file, grep, glob, ls) resolved model-supplied
paths against a root list whose first entry was the whole data directory.
That directory holds the session store, the auth database, the app
encryption key and the settings file, so prompt-injected content could ask
for any of them. No approval prompt stood in the way: reads are classified
read_workspace and pass the untrusted-context gate untouched, which is
correct for reading a workspace and wrong for reading the app's own state.

The agent gets data/agent_workspace/ instead, and the subprocess cwd and
HOME move with it so bash and read_file agree on where scratch files live.

The deny itself is a property of the path, not of the root it arrived
through, because three routes reach the same bytes and closing only the
first leaves the other two working:

  - the default root list
  - a workspace bound at or above the data directory, which vet_workspace
    accepted and chat_routes auto-binds from a path named in the message
  - a tool_path_extra_roots setting covering the data directory

_resolve_search_root also returned the workspace root unchecked when the
path was empty, so a bare ls enumerated the directory whatever the deny
list said. It now resolves that case through the same guards.

A containment rule rather than a filename deny list, so state files added
later are covered without anyone remembering to list them, and so a user's
own settings.json or app.db inside a real workspace is not caught.

Four directories of user content stay readable, because the application
hands their paths to the model and tells it to open them: the chat upload
manifest, downloaded mail attachments, personal docs (which covers the
runbook) and personal uploads.

* fix: enforce state deny during recursive file search

* fix: bound protected filesystem searches

* fix(security): reject inode aliases and workspace redirects

* fix(security): harden partitioned agent searches

* fix(security): report fallback worker exits promptly

* fix(security): clean up search readers and retain relative data roots

---------

Co-authored-by: RaresKeY <158580472+RaresKeY@users.noreply.github.com>
2026-09-05 19:21:12 +02:00
nopozandRaresKeY f88e2d1f7f Merge commit from fork
* fix(security): stop API tokens reaching privileged agent tools

A bearer API token resolves to the human who minted it, and minting is admin-only, so every owner-keyed privilege check in the agent path answers "admin". A token issued for a narrow integration therefore reached bash and python with the authority of the account that created it.

Three independent routes to that sink, each closed here.

The token could answer its own tool-approval prompt. An approval records that a person authorized one dangerous action, and a token cannot make that statement, so /api/chat_stream now refuses an approval resume from a bearer caller.

The chat-session grant was reconstructable from caller-supplied message metadata. Two routes persist a metadata blob on the caller's behalf, so the shape of a resolved approval card could be written straight into a transcript and was then read back as authority. The server now signs the grant when it resolves an approval and verifies that signature when reading it back, binding it to the chat and the approval it was issued for. Both routes also drop server-owned keys from an inbound blob.

A run driven by a token inherited its owner's tool set. Such a run is now capped at the non-admin policy regardless of who minted the credential, which holds even where no approval is raised at all.

The human path is unchanged: a browser session still receives the prompt, still approves, and a granted chat-session scope still carries to later turns in that chat.

Scope enforcement across the wider route surface is a separate gap and is not addressed here.

* fix scoped chat delegation boundaries

* fix(auth): reject malformed chat approval signatures

---------

Co-authored-by: RaresKeY <158580472+RaresKeY@users.noreply.github.com>
2026-09-05 19:20:49 +02:00
rauljuaandRaul c7a8637475 fix(docker): repair app cache parent ownership (#6158)
* fix(docker): repair app cache parent ownership

* fix(docker): avoid walking mounted model cache

* test(docker): exercise nested cache ownership

---------

Co-authored-by: Raul <9117159+raultcj@users.noreply.github.com>
2026-09-05 18:05:38 +02:00
VykosandClaude affaee1e66 fix(discovery): cache a successful but empty Tailscale lookup (#6228)
The host cache was gated on the list being non-empty, so "queried fine, no
eligible peers" looked exactly like a cold cache and every caller paid for
another `tailscale status --json` — a subprocess with a 5s timeout.

Gate on the timestamp instead. Failures still leave the timestamp unset, so a
missing binary, a non-zero exit or unparseable output stays retryable rather
than being cached for the full TTL.

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-02 12:05:01 +02:00
daixiheguu ce04dc1db4 fix(tasks): clean up singleflight cache on cancellation (#6174)
Signed-off-by: daixiheguu <daixihegu@outlook.com>
2026-09-01 18:34:50 +02:00
RaresKeY c9dd68d890 refactor(docs): separate Pages site source (#6176)
* refactor(docs): separate Pages site source

* fix(docs): preserve published guide pages

* fix(ci): run asset ownership tests for site changes

* fix(docs): track future website media

* fix(ci): let Pages deployments finish

* build(deps): update Pages checkout action

* fix(ci): follow moved setup guide

* fix(docs): repair published setup guide

* fix(docs): retarget preview encoder
2026-08-27 10:20:36 +02:00
nopoz d0d8edf5d8 Merge commit from fork
scripts/mlx_image_server.py resolved the model per request
(`req.model or _args.model`) on both /v1/images/generations and
/v1/images/edits, so the caller chose which model was served.

`_is_hidream()` is a substring test and `_snapshot_path()` accepts either a
local directory or a Hugging Face repo id, so a caller-supplied string
selected the HiDream branch and then supplied the directory it runs
`scripts/hidream_o1/generate_hidream_o1_mlx.py` from, under sys.executable.
The server has no auth, and the Cookbook binds it to 0.0.0.0 whenever it is
serving to a remote host, so one POST executed attacker code on the serving
host.

Both paths now use `_args.model`. The request field is still accepted for
OpenAI wire compatibility and ignored, matching scripts/diffusion_server.py,
and Odysseus already sends the served model's own id, so this is a no-op for
legitimate callers. /v1/images/harmonize already pinned.

Regression tests cover both endpoints, the local-directory and
Hugging-Face-repo halves, and that a server actually launched with a HiDream
model still serves it. Three of the four fail on the unfixed code.
2026-08-24 17:38:40 +02:00
Joeseph Grey b4d12932a9 fix(agent): drop the empty assistant turn from an approved-action replay (#6124)
The approved-action replay appends the sealed tool result with no assistant
prose for that round, which produced an assistant message with content "".
Anthropic's Messages API rejects a non-final assistant message with empty
content, so a resumed turn after a tool approval failed before the model saw
the result. A turn carrying neither prose nor reasoning has nothing to say to
any provider, so it is no longer appended. A round with prose, and a
reasoning-only round that DeepSeek thinking mode needs, both still append.
2026-08-20 13:06:22 +02:00
Nikhil Chaudhary 85297cee44 fix(core): clean up orphaned temp files on atomic write failure (#6068)
* fix(core): clean up orphaned temp files on atomic write failure

* fixed reviewer suggestion

* removed whitespace
2026-08-19 17:38:24 +02:00
RaresKeYandLéo 981652358e fix(agent): allow remaining actions for an approved task (#6113)
* fix(agent): allow remaining actions for an approved task

* fix(agent): make approval continuation control-only

* fix(ci): preserve approval taint and cache-buster contract

* fix(ui): keep tool approvals in current chat

* fix(ui): route tool approvals through chat submit

* test(ui): pin approval submit routing

* fix(agent): complete approval denial flow

* fix(ui): avoid duplicate ask-user close icon

* fix(agent): retain approved tool in continuation set

* revert(ui): keep PR 6113 scoped to approval continuation

* fix(agent): add task and chat approval scopes

* fix(ui): prevent duplicate ask-user close icon

* feat(ui): add ask-user option shortcuts

* fix(compare): route ask-user choices per pane

* fix(agent): keep skill-test approvals to a single action

The chat card now reuses the wire value `approve` to mean chat-session
scope, and `consume()` returned `allow_remaining_actions=True` for it
unconditionally. The skill-test approval route was never updated: it still
sends `approve` meaning "once", and its button still reads "Allow once",
but the grant it got back set `approval_gate_bypassed` for the rest of the
resumed run. That surface wraps the skill body and every transcript byte
as untrusted context, so it is the last place where one click should
ungate everything that follows.

Give `consume()` an explicit `allow_continuation` flag. Callers that own a
resumable chat keep the scope the user picked; callers that do not — the
skill tester, unattended audits — get SINGLE_ACTION and the gate re-arms
behind the sealed action, which is what their label promises.

* fix(ui): cache-bust every module the approval click depends on

chatStream.js, compare/index.js and compare/stream.js all changed
behaviour but kept their old `?v=`, while chat.js and chatRenderer.js were
bumped. A returning browser therefore serves the new chat.js — which now
deliberately leaves the composer empty and clicks the send button — next to
the cached chatStream.js that has no interceptor. With an empty composer
that button sits at `data-mode="newchat"`, so the click opens a new chat
and the approval is dropped.

Bump the three, and version compare/stream.js's chatRenderer import to
match everyone else's so the ask_user keydown listener binds to one module
instance instead of two.

* fix(ui): keep the digit shortcuts off tool approval cards

With an approval card on screen and focus anywhere outside an input, a bare
`1` fired `approve_task` — the widest of the three grants — with no
modifier and no confirmation. That card is the one control whose entire
purpose is deliberate consent after untrusted context influenced the run,
and Deny sits at 3.

Label the card with its kind and skip the shortcut for approvals. Ordinary
ask_user questions keep 1-3.

* fix(compare): restore a pane's ask_user card instead of dropping the choice

renderAskUserCard removes the card as soon as onSubmit accepts, but the
resume loop gave up silently after 10s if the originating stream still owned
the pane. The user saw the click land, the card vanish, and nothing happen,
with no way to get it back.

Re-render the card on that deadline and say why. The reroll case still
returns without sending — that choice belongs to a stream that no longer
exists.

* refactor(chat): drop the unreachable deny branch

`if decision != "deny"` is always true — the deny path returns a
StreamingResponse a few lines above. It reads as if deny still falls
through to the toggle restore.

---------

Co-authored-by: Léo <leograndcontact@gmail.com>
2026-08-19 08:01:34 -06:00
Utkarsh AdhranandRaresKeY 5c835014ac fix(time): prefer IANA timezone name over offset (#6122)
* fix(time): prefer IANA timezone name over offset

When both headers are present, resolve x-tz-name with ZoneInfo and ignore
a conflicting numeric offset. The prompt label uses the resolved zone so
name and UTC offset cannot disagree.

Related: #6111

* test(calendar): cover IANA timezone precedence

---------

Co-authored-by: RaresKeY <158580472+RaresKeY@users.noreply.github.com>
2026-08-19 12:56:07 +02:00
Dividesbyzer0 43682d4e2e fix(cookbook): activate local Windows venv in bash runner (#5734) 2026-08-18 16:19:33 +02:00
RaresKeY 032967af4b fix(models): show API models by default (#6089) 2026-08-17 13:41:04 +02:00
RaresKeY 0e03aea134 fix(models): align API model checkbox state (#6087) 2026-08-17 11:09:02 +01:00
Joeseph Grey 2a6b09b968 Merge pull request #6081 from ydonghao/refactor/routes-task-to-subdir
refactor(routes): move task domain into routes/task/ subpackage
2026-08-16 22:29:43 -06:00