mirror of
https://github.com/pewdiepie-archdaemon/odysseus.git
synced 2026-10-06 15:02:20 +02:00
"Is this path inside that root" is asked in twenty places in this tree and answered twenty times by a locally written realpath/commonpath pair. Nine test files exist because nine call sites each needed their own proof. Each one is defensible alone; together they are the defect, because the boundary has no single definition and a site that gets a detail wrong is wrong by itself. src/path_confinement.py is that definition, and it settles the details the copies disagreed on. Both sides get canonicalized: comparing a realpath-ed candidate against a root that was only abspath-ed is the macOS /tmp -> /private/tmp mismatch that has already produced a false failure here, and canonicalizing one side is worse than canonicalizing neither. commonpath rather than startswith, because /a/bc begins with /a/b and is not inside it. A relative candidate joins the root rather than os.getcwd(), which is whatever directory the server happens to be running in. NUL and newline are refused with a reason instead of caught by a bare `except Exception` and reported as an ordinary escape. Eighteen call sites go through it now. It deliberately does not decide whether a path is sensitive -- that deny list answers "allowed" rather than "inside", and it stays with src/tool_execution, which owns it. The one commonpath left in the tree, in src/workspace_paths.py, stays: that function translates a host path into a container path, so canonicalizing either side would change the relative path it computes and break the mapping. It is not a confinement check. Two of those sites were weaker than the rest and are fixed rather than moved. The email attachment check used abspath, which folds `..` but does not resolve symlinks, so a symlink written into the extraction directory passed it and was then read through. The skill-reference guard compared a realpath-ed target against a raw dirname, so on a host where the skills tree is reached through a symlink the two sides never matched and the guard could not fire. The execution boundary had two separate holes. The workspace namespace bound /home and /mnt read-write. On the one platform where that namespace engages at all, a command inside it reaches outside the workspace and writes to the user's home directory -- measured by running this argv on a Linux host with working bubblewrap, not inferred from the source. Binding the user's whole home directory into a workspace-confinement namespace gives back most of what the namespace was for. Both are read-only now. The workspace is also bound writable at its real host path, not only at /workspace: BashTool's own /tmp redirect rewrites `/tmp/` to `<agent_cwd()>/.tmp/` before the namespace is built, so the command bwrap receives already names the real path, and those writes previously landed only because the workspace happened to sit under the writable /home. `namespaced or _replace_workspace_alias(...)` chose between a mount namespace and a regex with nothing in the result saying which one ran. The fallback rewrites the literal token /workspace in the command string, so a command that never mentions /workspace is untouched by it and runs on the host unrestricted -- which is every agent shell command on macOS. Both tools now ask containment.probe() instead of each deciding for itself, and every bash and python result carries a containment block naming the mechanism and stating whether the filesystem dimension actually held. Under enforcing mode the command is not run and the result says so. That block reports the filesystem dimension only, and says so in a reported_dimensions field. The probe knows this host could also give a process group and a real wall clock, but these two tools still assemble their own create_subprocess_* call and pass neither, so listing those dimensions would be exactly the false claim src/containment.py calls worse than an honest absence. probe() is new on src/containment.py: the same mechanism table and the same arithmetic as acquire(), stopping before the side effects. acquire() is the wrong shape for a decision -- it writes a durable grant record, and a record whose pid is never filled in and whose release() never runs is an entry a restart reaper keeps finding. CONTAINMENT_MODE stays report_only. Flipping it refuses every agent shell command on macOS and on any Linux host without bubblewrap, which is a product decision rather than a code one. Smaller things in the same area: the /tmp redirect's makedirs was unguarded, so a read-only workspace turned a command that merely mentioned `/tmp/` into an OSError traceback instead of a tool error; it degrades now. WORKSPACE_MOUNT moved to src/constants.py so the namespace and the path resolvers read one definition of the contract rather than two. The ".tmp" dirname got a constant, since it appeared in both tool paths. One generated artifact moved with it: website/configuration-reference.md pins the source line where each ODYSSEUS_* variable is read, and three of those shifted. Regenerated with scripts/generate_env_reference.py; the diff is line numbers only. Three existing tests changed. test_workspace_artifact_tool_floor asserted that an unsafe interpreter prefix produces no `--ro-bind <prefix> <prefix>`, which now fires on /home because /home is legitimately a read-only base mount. Asserting the absence of a literal flag string cannot distinguish "the prefix was rejected" from "the argv mounted that root itself", so it compares the argv against the no-prefix baseline instead: an unsafe prefix must add nothing. The Windows bash test asserted dict equality on the whole result, which makes adding a field to every bash result impossible without touching a test about tmux; it asserts the shape now. The personal-dir symlink test grepped the resolver's source for the literal "os.path.realpath", which is gone because the resolution moved into the shared boundary -- it keeps the negative assertion that the closure must not grow its own abspath check again, and the behavioural half now runs against the boundary, where it covers every call site instead of one closure. Not verified: the bubblewrap argv is asserted, not executed. There is no bwrap on macOS, and in Docker it needs --privileged to work at all -- default and seccomp=unconfined both fail with "Creating new namespace failed", and --cap-add=SYS_ADMIN fails at pivot_root. The Python tool's needs_virtual_namespace gate means ordinary Python code gets no namespace even on a Linux host that could provide one; that is reported now but deliberately not changed, because it alters the Linux Python path on every call and cannot be checked from here.
240 lines
8.0 KiB
Python
240 lines
8.0 KiB
Python
"""document_helpers.py — Pydantic models, doc serializers, owner gating, file-locator helpers shared with document_routes.py."""
|
|
|
|
"""Document routes — CRUD for living documents with version history."""
|
|
|
|
import logging
|
|
import os
|
|
import re
|
|
from typing import Any, Dict, Optional
|
|
|
|
from fastapi import HTTPException, Request
|
|
from pydantic import BaseModel
|
|
|
|
from core.database import Document, DocumentVersion
|
|
from core.database import Session as DbSession
|
|
from src.auth_helpers import _auth_disabled
|
|
from src.path_confinement import is_inside
|
|
from src.upload_handler import UploadHandler
|
|
|
|
logger = logging.getLogger(__name__)
|
|
|
|
|
|
# ---- Request schemas ----
|
|
|
|
class DocumentCreate(BaseModel):
|
|
session_id: Optional[str] = None
|
|
title: str = "Untitled"
|
|
language: Optional[str] = None
|
|
content: str = ""
|
|
|
|
class DocumentUpdate(BaseModel):
|
|
content: str
|
|
summary: Optional[str] = None
|
|
force_version: bool = False
|
|
|
|
class DocumentPatch(BaseModel):
|
|
title: Optional[str] = None
|
|
language: Optional[str] = None
|
|
session_id: Optional[str] = None # link/unlink document to a session
|
|
|
|
|
|
# ---- Helpers ----
|
|
|
|
def _doc_to_dict(doc: Document) -> Dict[str, Any]:
|
|
return {
|
|
"id": doc.id,
|
|
"session_id": doc.session_id,
|
|
"title": doc.title,
|
|
"language": doc.language,
|
|
"current_content": doc.current_content,
|
|
"version_count": doc.version_count,
|
|
"is_active": doc.is_active,
|
|
"archived": bool(getattr(doc, "archived", False)),
|
|
"created_at": (doc.created_at.isoformat() + "Z") if doc.created_at else None,
|
|
"updated_at": (doc.updated_at.isoformat() + "Z") if doc.updated_at else None,
|
|
# Source-email provenance (set when doc was created from an email
|
|
# attachment) — drives the "Send signed reply" menu item.
|
|
"source_email_uid": getattr(doc, "source_email_uid", None),
|
|
"source_email_folder": getattr(doc, "source_email_folder", None),
|
|
"source_email_account_id": getattr(doc, "source_email_account_id", None),
|
|
"source_email_message_id": getattr(doc, "source_email_message_id", None),
|
|
}
|
|
|
|
def _version_to_dict(v: DocumentVersion) -> Dict[str, Any]:
|
|
return {
|
|
"id": v.id,
|
|
"document_id": v.document_id,
|
|
"version_number": v.version_number,
|
|
"content": v.content,
|
|
"summary": v.summary,
|
|
"source": v.source,
|
|
"created_at": v.created_at.isoformat() if v.created_at else None,
|
|
}
|
|
|
|
|
|
def _verify_doc_owner(db, doc: Document, user: str):
|
|
"""Verify `user` owns this document. Raise 404 if not.
|
|
|
|
Documents now carry their own `owner` column, so a doc whose session
|
|
was deleted (session_id → NULL) can still prove ownership and stay
|
|
openable / cloneable. We trust that column first and only fall back to
|
|
the session join for any not-yet-backfilled legacy row.
|
|
"""
|
|
if user is None:
|
|
if _auth_disabled():
|
|
return # Single-user / no-auth mode: allow access
|
|
raise HTTPException(403, "Authentication required")
|
|
if doc.owner is not None:
|
|
if doc.owner != user:
|
|
raise HTTPException(404, "Document not found")
|
|
return
|
|
# Legacy fallback: derive ownership from the linked session.
|
|
if not doc.session_id:
|
|
raise HTTPException(404, "Document not found")
|
|
session = db.query(DbSession).filter(DbSession.id == doc.session_id).first()
|
|
if not session or session.owner != user:
|
|
raise HTTPException(404, "Document not found")
|
|
|
|
|
|
def _owner_session_filter(q, user):
|
|
"""Restrict a documents query to those owned by `user`.
|
|
|
|
Documents now carry their own `owner` column (backfilled at boot from
|
|
the linked session, or assigned to the admin user for legacy/orphaned
|
|
docs). We filter on that directly rather than on a session join, so a
|
|
document whose session was deleted (session_id → NULL) still shows up
|
|
for its owner instead of silently vanishing from the Library + search.
|
|
|
|
The owner backfill runs in init_db before the app serves requests, so
|
|
by the time this filter is live there are no NULL-owner rows to leak;
|
|
we therefore match the owner strictly for authenticated callers."""
|
|
if not user:
|
|
if user == "" or _auth_disabled():
|
|
return q
|
|
return q.filter(False)
|
|
return q.filter(Document.owner == user)
|
|
|
|
|
|
|
|
def _slug(name: str) -> str:
|
|
"""Filesystem-friendly version of a document title.
|
|
|
|
Whitespace becomes underscores; other unsafe punctuation is dropped.
|
|
Preserves letters, digits, dot, hyphen, underscore. Idempotent.
|
|
"""
|
|
import re as _re
|
|
s = (name or "").strip()
|
|
# Drop the trailing extension if the title happens to include one
|
|
s = _re.sub(r'\.pdf$', '', s, flags=_re.IGNORECASE)
|
|
s = _re.sub(r'\s+', '_', s)
|
|
s = _re.sub(r'[^A-Za-z0-9._-]', '', s)
|
|
s = _re.sub(r'_+', '_', s).strip('_')
|
|
return s or "form"
|
|
|
|
|
|
# DPI scale for the interactive PDF view. ~150 DPI (2x of 72 PDF user-units).
|
|
_PDF_RENDER_SCALE = 2.0
|
|
|
|
|
|
def _upload_path_inside(upload_dir: str, path: str) -> bool:
|
|
return is_inside(upload_dir, path)
|
|
|
|
|
|
def _resolve_user_upload_path(
|
|
upload_handler: Any,
|
|
upload_id: str,
|
|
owner: Optional[str],
|
|
auth_manager=None,
|
|
) -> Optional[str]:
|
|
"""Resolve an upload id to a filesystem path the caller may read."""
|
|
if upload_handler is None:
|
|
return None
|
|
resolved = upload_handler.resolve_upload(
|
|
upload_id,
|
|
owner=owner,
|
|
auth_manager=auth_manager,
|
|
)
|
|
if not isinstance(resolved, dict) or not resolved:
|
|
return None
|
|
path = resolved.get("path")
|
|
upload_dir = getattr(upload_handler, "upload_dir", None)
|
|
if path and upload_dir and not _upload_path_inside(upload_dir, path):
|
|
logger.warning("Upload path outside upload directory: %s", path)
|
|
return None
|
|
return path
|
|
|
|
|
|
def _locate_upload(
|
|
upload_dir: str,
|
|
file_id: str,
|
|
owner: Optional[str] = None,
|
|
auth_manager=None,
|
|
upload_handler: Any = None,
|
|
):
|
|
"""Find an upload by its filename ID via UploadHandler.resolve_upload."""
|
|
if upload_handler is None:
|
|
from src.upload_handler import UploadHandler
|
|
|
|
base_dir = os.path.dirname(os.path.abspath(upload_dir))
|
|
upload_handler = UploadHandler(base_dir, upload_dir)
|
|
return _resolve_user_upload_path(upload_handler, file_id, owner, auth_manager)
|
|
|
|
|
|
def _assert_pdf_marker_upload_owned(
|
|
request: Request,
|
|
content: str,
|
|
user: Optional[str],
|
|
upload_handler: Any,
|
|
) -> None:
|
|
"""Reject document content whose pdf_source marker points at another user's upload."""
|
|
if upload_handler is None:
|
|
return
|
|
from src.pdf_form_doc import find_source_upload_id
|
|
|
|
upload_id = find_source_upload_id(content or "")
|
|
if not upload_id:
|
|
return
|
|
auth_manager = getattr(getattr(request.app, "state", None), "auth_manager", None)
|
|
if not _resolve_user_upload_path(upload_handler, upload_id, user, auth_manager):
|
|
raise HTTPException(
|
|
400,
|
|
"Document PDF marker references an upload you do not own",
|
|
)
|
|
|
|
|
|
def _derive_title(content: str) -> str:
|
|
"""Derive a title from document content."""
|
|
import re
|
|
if not isinstance(content, str):
|
|
return "Untitled"
|
|
text = content.strip()
|
|
if not text:
|
|
return "Untitled"
|
|
|
|
# Markdown header
|
|
md = re.match(r'^#{1,3}\s+(.+)', text, re.MULTILINE)
|
|
if md:
|
|
title = md.group(1).strip()
|
|
if len(title) > 50:
|
|
title = title[:48] + "…"
|
|
return title
|
|
|
|
# HTML heading
|
|
html = re.search(r'<h[1-3][^>]*>([^<]+)</h[1-3]>', text, re.IGNORECASE)
|
|
if html:
|
|
title = html.group(1).strip()
|
|
if len(title) > 50:
|
|
title = title[:48] + "…"
|
|
return title
|
|
|
|
# First non-empty line (if short enough)
|
|
for line in text.split('\n'):
|
|
line = line.strip()
|
|
if line and 2 <= len(line) <= 60:
|
|
title = re.sub(r'[:#*`]+$', '', line).strip()
|
|
if title and len(title) > 50:
|
|
title = title[:48] + "…"
|
|
return title or "Untitled"
|
|
|
|
return "Untitled"
|