mirror of
https://github.com/pewdiepie-archdaemon/odysseus.git
synced 2026-10-06 23:12:22 +02:00
"Is this path inside that root" is asked in twenty places in this tree and answered twenty times by a locally written realpath/commonpath pair. Nine test files exist because nine call sites each needed their own proof. Each one is defensible alone; together they are the defect, because the boundary has no single definition and a site that gets a detail wrong is wrong by itself. src/path_confinement.py is that definition, and it settles the details the copies disagreed on. Both sides get canonicalized: comparing a realpath-ed candidate against a root that was only abspath-ed is the macOS /tmp -> /private/tmp mismatch that has already produced a false failure here, and canonicalizing one side is worse than canonicalizing neither. commonpath rather than startswith, because /a/bc begins with /a/b and is not inside it. A relative candidate joins the root rather than os.getcwd(), which is whatever directory the server happens to be running in. NUL and newline are refused with a reason instead of caught by a bare `except Exception` and reported as an ordinary escape. Eighteen call sites go through it now. It deliberately does not decide whether a path is sensitive -- that deny list answers "allowed" rather than "inside", and it stays with src/tool_execution, which owns it. The one commonpath left in the tree, in src/workspace_paths.py, stays: that function translates a host path into a container path, so canonicalizing either side would change the relative path it computes and break the mapping. It is not a confinement check. Two of those sites were weaker than the rest and are fixed rather than moved. The email attachment check used abspath, which folds `..` but does not resolve symlinks, so a symlink written into the extraction directory passed it and was then read through. The skill-reference guard compared a realpath-ed target against a raw dirname, so on a host where the skills tree is reached through a symlink the two sides never matched and the guard could not fire. The execution boundary had two separate holes. The workspace namespace bound /home and /mnt read-write. On the one platform where that namespace engages at all, a command inside it reaches outside the workspace and writes to the user's home directory -- measured by running this argv on a Linux host with working bubblewrap, not inferred from the source. Binding the user's whole home directory into a workspace-confinement namespace gives back most of what the namespace was for. Both are read-only now. The workspace is also bound writable at its real host path, not only at /workspace: BashTool's own /tmp redirect rewrites `/tmp/` to `<agent_cwd()>/.tmp/` before the namespace is built, so the command bwrap receives already names the real path, and those writes previously landed only because the workspace happened to sit under the writable /home. `namespaced or _replace_workspace_alias(...)` chose between a mount namespace and a regex with nothing in the result saying which one ran. The fallback rewrites the literal token /workspace in the command string, so a command that never mentions /workspace is untouched by it and runs on the host unrestricted -- which is every agent shell command on macOS. Both tools now ask containment.probe() instead of each deciding for itself, and every bash and python result carries a containment block naming the mechanism and stating whether the filesystem dimension actually held. Under enforcing mode the command is not run and the result says so. That block reports the filesystem dimension only, and says so in a reported_dimensions field. The probe knows this host could also give a process group and a real wall clock, but these two tools still assemble their own create_subprocess_* call and pass neither, so listing those dimensions would be exactly the false claim src/containment.py calls worse than an honest absence. probe() is new on src/containment.py: the same mechanism table and the same arithmetic as acquire(), stopping before the side effects. acquire() is the wrong shape for a decision -- it writes a durable grant record, and a record whose pid is never filled in and whose release() never runs is an entry a restart reaper keeps finding. CONTAINMENT_MODE stays report_only. Flipping it refuses every agent shell command on macOS and on any Linux host without bubblewrap, which is a product decision rather than a code one. Smaller things in the same area: the /tmp redirect's makedirs was unguarded, so a read-only workspace turned a command that merely mentioned `/tmp/` into an OSError traceback instead of a tool error; it degrades now. WORKSPACE_MOUNT moved to src/constants.py so the namespace and the path resolvers read one definition of the contract rather than two. The ".tmp" dirname got a constant, since it appeared in both tool paths. One generated artifact moved with it: website/configuration-reference.md pins the source line where each ODYSSEUS_* variable is read, and three of those shifted. Regenerated with scripts/generate_env_reference.py; the diff is line numbers only. Three existing tests changed. test_workspace_artifact_tool_floor asserted that an unsafe interpreter prefix produces no `--ro-bind <prefix> <prefix>`, which now fires on /home because /home is legitimately a read-only base mount. Asserting the absence of a literal flag string cannot distinguish "the prefix was rejected" from "the argv mounted that root itself", so it compares the argv against the no-prefix baseline instead: an unsafe prefix must add nothing. The Windows bash test asserted dict equality on the whole result, which makes adding a field to every bash result impossible without touching a test about tmux; it asserts the shape now. The personal-dir symlink test grepped the resolver's source for the literal "os.path.realpath", which is gone because the resolution moved into the shared boundary -- it keeps the negative assertion that the closure must not grow its own abspath check again, and the behavioural half now runs against the boundary, where it covers every call site instead of one closure. Not verified: the bubblewrap argv is asserted, not executed. There is no bwrap on macOS, and in Docker it needs --privileged to work at all -- default and seccomp=unconfined both fail with "Creating new namespace failed", and --cap-add=SYS_ADMIN fails at pivot_root. The Python tool's needs_virtual_namespace gate means ordinary Python code gets no namespace even on a Linux host that could provide one; that is reported now but deliberately not changed, because it alters the Linux Python path on every call and cannot be checked from here.
841 lines
36 KiB
Python
841 lines
36 KiB
Python
# services/memory/skills.py
|
|
"""Skills storage layer.
|
|
|
|
Skills live on disk as `data/skills/<category>/<name>/SKILL.md` files with
|
|
YAML frontmatter and a structured markdown body (When to Use / Procedure /
|
|
Pitfalls / Verification). See `skill_format.py` for the format.
|
|
|
|
Usage counters (`uses`, `last_used`) live in a sidecar
|
|
`data/skills/_usage.json` keyed by owner plus skill name so the SKILL.md
|
|
content doesn't churn on every retrieval.
|
|
|
|
Ownership: skills declare `owner: <username>` in frontmatter. Single-user
|
|
deployments can leave that blank.
|
|
|
|
This module also retains a JSON fallback for any legacy `data/skills.json`
|
|
entries — they're surfaced as read-only `Skill` objects so old data still
|
|
loads while a user migrates them to disk.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import json
|
|
import logging
|
|
import os
|
|
import time
|
|
from typing import Dict, Iterable, List, Optional
|
|
|
|
from src.path_confinement import confine
|
|
|
|
from .skill_format import Skill, slugify
|
|
|
|
logger = logging.getLogger(__name__)
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# Token / similarity helpers (kept for the relevance fallback)
|
|
# ---------------------------------------------------------------------------
|
|
|
|
def _tokenize(text: str) -> set:
|
|
return {w.strip('.,!?";:()[]') for w in (text or "").lower().split() if len(w) > 1}
|
|
|
|
|
|
def _jaccard(a: set, b: set) -> float:
|
|
if not a or not b:
|
|
return 0.0
|
|
return len(a & b) / len(a | b)
|
|
|
|
|
|
def _to_float(x, default: float = 0.0) -> float:
|
|
"""Coerce a possibly hand-edited frontmatter value to float without
|
|
raising — a blank or non-numeric `confidence:` in a SKILL.md must not
|
|
blow up retrieval or eviction."""
|
|
try:
|
|
return float(x)
|
|
except (TypeError, ValueError):
|
|
return default
|
|
|
|
|
|
def _approval_policy(owner: Optional[str]) -> tuple[bool, float]:
|
|
"""Read the user's automatic skill-approval gate without breaking retrieval."""
|
|
try:
|
|
from routes.prefs_routes import _load_for_user
|
|
prefs = _load_for_user(owner) or {}
|
|
except Exception:
|
|
prefs = {}
|
|
try:
|
|
from src.settings import get_setting
|
|
default_minimum = float(get_setting("skill_autosave_min_confidence", 0.85))
|
|
except Exception:
|
|
default_minimum = 0.85
|
|
try:
|
|
minimum = float(prefs.get("skill_min_confidence", default_minimum))
|
|
except (TypeError, ValueError):
|
|
minimum = default_minimum
|
|
return bool(prefs.get("auto_approve_skills", True)), max(0.0, min(1.0, minimum))
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# SkillsManager
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
class SkillsManager:
|
|
"""Read/write SKILL.md files under <data_dir>/skills/."""
|
|
|
|
def __init__(self, data_dir: str):
|
|
self.data_dir = data_dir
|
|
self.skills_root = os.path.join(data_dir, "skills")
|
|
self.usage_file = os.path.join(self.skills_root, "_usage.json")
|
|
self.legacy_file = os.path.join(data_dir, "skills.json") # back-compat
|
|
os.makedirs(self.skills_root, exist_ok=True)
|
|
|
|
# ----------------------------------------------------------------------
|
|
# Path helpers
|
|
# ----------------------------------------------------------------------
|
|
|
|
def _skill_dir(self, category: str, name: str) -> str:
|
|
cat = slugify(category or "general", fallback="general")
|
|
nm = slugify(name, fallback="skill")
|
|
return os.path.join(self.skills_root, cat, nm)
|
|
|
|
def _skill_file(self, category: str, name: str) -> str:
|
|
return os.path.join(self._skill_dir(category, name), "SKILL.md")
|
|
|
|
# ----------------------------------------------------------------------
|
|
# Usage sidecar
|
|
# ----------------------------------------------------------------------
|
|
|
|
def _load_usage(self) -> Dict[str, Dict]:
|
|
if not os.path.exists(self.usage_file):
|
|
return {}
|
|
try:
|
|
with open(self.usage_file, encoding="utf-8") as f:
|
|
d = json.load(f)
|
|
return d if isinstance(d, dict) else {}
|
|
except Exception:
|
|
return {}
|
|
|
|
def _save_usage(self, usage: Dict[str, Dict]) -> None:
|
|
try:
|
|
from core.atomic_io import atomic_write_json
|
|
atomic_write_json(self.usage_file, usage, indent=2)
|
|
except Exception:
|
|
tmp = self.usage_file + ".tmp"
|
|
with open(tmp, "w", encoding="utf-8") as f:
|
|
json.dump(usage, f, indent=2)
|
|
os.replace(tmp, self.usage_file)
|
|
|
|
@staticmethod
|
|
def _usage_key(name: str, owner: Optional[str] = None) -> str:
|
|
# Skill names are not globally unique once multiple owners are present.
|
|
# Keep the usage sidecar keyed the same way the skill file is scoped.
|
|
return f"{owner}::{name}" if owner else name
|
|
|
|
def _usage_entry(self, usage: Dict[str, Dict], name: str, owner: Optional[str] = None) -> Dict:
|
|
key = self._usage_key(name, owner)
|
|
entry = usage.get(key)
|
|
if isinstance(entry, dict):
|
|
return entry
|
|
return {}
|
|
|
|
def set_audit(self, name: str, verdict: str, by_teacher: bool = False,
|
|
worker_model: str = "", teacher_model: str = "",
|
|
owner: Optional[str] = None, saved_turns: Optional[int] = None,
|
|
saved_tool_calls: Optional[int] = None,
|
|
baseline_verdict: Optional[str] = None,
|
|
usefulness: Optional[float] = None,
|
|
audit_summary: Optional[str] = None) -> None:
|
|
"""Record the last test/audit result for a skill in the usage sidecar
|
|
(so it surfaces in load() without touching SKILL.md). Drives the
|
|
'verified' check + teacher mark on the card."""
|
|
import time as _t
|
|
usage = self._load_usage()
|
|
key = self._usage_key(name, owner)
|
|
e = usage.setdefault(key, {"uses": 0, "last_used": None})
|
|
e["audit_verdict"] = verdict
|
|
# Replace, rather than retain, the explanation from a previous run.
|
|
e["audit_summary"] = str(audit_summary or "")[:2000]
|
|
# Version 2 fixes audit-arm isolation and separates functional success
|
|
# from baseline utility. Legacy inconclusive results are not evidence
|
|
# under that protocol and should be eligible for a clean re-audit.
|
|
e["audit_version"] = 2
|
|
e["audit_by_teacher"] = bool(by_teacher)
|
|
if worker_model:
|
|
e["audit_worker_model"] = worker_model
|
|
if teacher_model:
|
|
e["audit_teacher_model"] = teacher_model
|
|
if saved_turns is not None:
|
|
try:
|
|
e["saved_turns"] = int(saved_turns)
|
|
except (TypeError, ValueError):
|
|
e.pop("saved_turns", None)
|
|
if saved_tool_calls is not None:
|
|
try:
|
|
e["saved_tool_calls"] = int(saved_tool_calls)
|
|
except (TypeError, ValueError):
|
|
e.pop("saved_tool_calls", None)
|
|
if baseline_verdict is not None:
|
|
e["baseline_verdict"] = str(baseline_verdict or "unknown")
|
|
if usefulness is not None:
|
|
try:
|
|
e["usefulness"] = float(usefulness)
|
|
except (TypeError, ValueError):
|
|
e.pop("usefulness", None)
|
|
e["audited_at"] = _t.time()
|
|
self._save_usage(usage)
|
|
|
|
def set_necessity(self, name: str, necessary: bool,
|
|
redundant_with=None, reason: str = "",
|
|
owner: Optional[str] = None) -> None:
|
|
"""Record the advisory 'is this skill necessary?' judgment in the usage
|
|
sidecar. Surfaced on the card as a flag; never acts on the skill."""
|
|
usage = self._load_usage()
|
|
key = self._usage_key(name, owner)
|
|
e = usage.setdefault(key, {"uses": 0, "last_used": None})
|
|
e["necessity"] = {
|
|
"necessary": bool(necessary),
|
|
"redundant_with": list(redundant_with or []),
|
|
"reason": str(reason or ""),
|
|
}
|
|
self._save_usage(usage)
|
|
|
|
# ----------------------------------------------------------------------
|
|
# Disk scan
|
|
# ----------------------------------------------------------------------
|
|
|
|
def _iter_skill_files(self) -> Iterable[str]:
|
|
if not os.path.isdir(self.skills_root):
|
|
return
|
|
for root, _dirs, files in os.walk(self.skills_root, followlinks=False):
|
|
if "SKILL.md" in files:
|
|
yield os.path.join(root, "SKILL.md")
|
|
|
|
def _read_skill(self, path: str) -> Optional[Skill]:
|
|
try:
|
|
with open(path, encoding="utf-8") as f:
|
|
text = f.read()
|
|
return Skill.from_markdown(text, path=path)
|
|
except Exception as e:
|
|
logger.warning(f"Failed to parse {path}: {e}")
|
|
return None
|
|
|
|
def _write_skill(self, sk: Skill) -> str:
|
|
path = self._skill_file(sk.category or "general", sk.name)
|
|
os.makedirs(os.path.dirname(path), exist_ok=True)
|
|
from core.atomic_io import atomic_write_text
|
|
atomic_write_text(path, sk.to_markdown())
|
|
sk.path = path
|
|
return path
|
|
|
|
def sync_builtin_skill(self, skill: Skill) -> str:
|
|
"""Persist a trusted built-in skill during startup synchronization."""
|
|
return self._write_skill(skill)
|
|
|
|
def backfill_owner(self, primary_owner: str, valid_owners: Optional[set[str]] = None) -> int:
|
|
"""Assign legacy/unclaimed skill files to the primary owner.
|
|
|
|
Skills are disk-backed, so the DB legacy-owner migration cannot fix
|
|
them. If strict owner filtering is enabled and SKILL.md files have no
|
|
owner or an owner from a deleted/test account, the UI appears empty even
|
|
though files still exist. This mirrors the DB legacy-owner sweep.
|
|
"""
|
|
primary_owner = (primary_owner or "").strip()
|
|
if not primary_owner:
|
|
return 0
|
|
valid_owners = set(valid_owners or [])
|
|
changed = 0
|
|
for path in self._iter_skill_files():
|
|
sk = self._read_skill(path)
|
|
if not sk:
|
|
continue
|
|
if sk.source == "builtin":
|
|
continue
|
|
owner = (sk.owner or "").strip()
|
|
if owner == primary_owner:
|
|
continue
|
|
if owner and owner in valid_owners:
|
|
continue
|
|
sk.owner = primary_owner
|
|
try:
|
|
self._write_skill(sk)
|
|
changed += 1
|
|
except Exception as e:
|
|
logger.warning("Failed to backfill owner for skill %s: %s", sk.name, e)
|
|
return changed
|
|
|
|
# ----------------------------------------------------------------------
|
|
# Public API — keeps the old method names so callers don't break
|
|
# ----------------------------------------------------------------------
|
|
|
|
def load_all(self) -> List[Dict]:
|
|
"""Return every skill as a plain dict, plus any legacy JSON entries."""
|
|
usage = self._load_usage()
|
|
out: List[Dict] = []
|
|
seen_names: set[str] = set()
|
|
for path in self._iter_skill_files():
|
|
sk = self._read_skill(path)
|
|
if not sk:
|
|
continue
|
|
d = sk.to_dict()
|
|
u = self._usage_entry(usage, sk.name, sk.owner)
|
|
d["uses"] = int(u.get("uses", 0))
|
|
d["last_used"] = u.get("last_used")
|
|
audit_verdict = u.get("audit_verdict")
|
|
try:
|
|
audit_version = int(u.get("audit_version") or 0)
|
|
except (TypeError, ValueError):
|
|
audit_version = 0
|
|
if audit_verdict == "inconclusive" and audit_version < 2:
|
|
audit_verdict = None
|
|
d["audit_verdict"] = audit_verdict
|
|
d["audit_summary"] = u.get("audit_summary", "") if audit_verdict else ""
|
|
d["audit_version"] = audit_version
|
|
d["audit_by_teacher"] = bool(u.get("audit_by_teacher"))
|
|
d["audit_worker_model"] = u.get("audit_worker_model")
|
|
d["audit_teacher_model"] = u.get("audit_teacher_model")
|
|
d["audited_at"] = u.get("audited_at") if audit_verdict else None
|
|
d["saved_turns"] = u.get("saved_turns")
|
|
d["saved_tool_calls"] = u.get("saved_tool_calls")
|
|
d["baseline_verdict"] = u.get("baseline_verdict")
|
|
d["usefulness"] = u.get("usefulness")
|
|
d["necessity"] = u.get("necessity")
|
|
out.append(d)
|
|
seen_names.add(sk.name)
|
|
# Legacy JSON entries — surfaced as draft, not editable from new flow
|
|
if os.path.exists(self.legacy_file):
|
|
try:
|
|
with open(self.legacy_file, encoding="utf-8") as f:
|
|
legacy = json.load(f)
|
|
if isinstance(legacy, list):
|
|
for row in legacy:
|
|
if not isinstance(row, dict):
|
|
continue
|
|
name = slugify(row.get("title") or row.get("id") or "skill")
|
|
if name in seen_names:
|
|
continue
|
|
out.append({
|
|
"id": row.get("id") or name,
|
|
"name": name,
|
|
"description": row.get("title", ""),
|
|
"version": "0.0.1",
|
|
"category": "legacy",
|
|
"tags": row.get("tags") or [],
|
|
"status": row.get("status") or "draft",
|
|
"confidence": row.get("confidence", 0.5),
|
|
"source": row.get("source", "imported"),
|
|
"owner": row.get("owner"),
|
|
"when_to_use": row.get("problem", ""),
|
|
"procedure": row.get("steps") or [],
|
|
"pitfalls": [],
|
|
"verification": [],
|
|
"body_extra": row.get("solution", ""),
|
|
"title": row.get("title", ""),
|
|
"problem": row.get("problem", ""),
|
|
"solution": row.get("solution", ""),
|
|
"steps": row.get("steps") or [],
|
|
"uses": row.get("uses", 0),
|
|
"last_used": row.get("last_used"),
|
|
"_legacy": True,
|
|
})
|
|
except Exception:
|
|
pass
|
|
return out
|
|
|
|
def load(self, owner: Optional[str] = None) -> List[Dict]:
|
|
entries = self.load_all()
|
|
if owner is None:
|
|
return entries
|
|
# SECURITY: strict ownership filter. The previous predicate also
|
|
# included skills with NO owner field (`not s.get("owner")`), which
|
|
# leaked legacy / un-stamped skills to every authenticated user.
|
|
# Hide them now; the owner needs to be backfilled on disk if those
|
|
# skills should be visible to a specific user.
|
|
return [
|
|
s for s in entries
|
|
if s.get("owner") == owner
|
|
or (s.get("source") == "builtin" and not s.get("owner"))
|
|
]
|
|
|
|
# ----------------------------------------------------------------------
|
|
# CRUD — disk-backed
|
|
# ----------------------------------------------------------------------
|
|
|
|
def add_skill(
|
|
self,
|
|
title: str = "",
|
|
problem: str = "",
|
|
solution: str = "",
|
|
steps: Optional[List[str]] = None,
|
|
tags: Optional[List[str]] = None,
|
|
source: str = "learned",
|
|
teacher_model: Optional[str] = None,
|
|
confidence: float = 0.8,
|
|
session_id: Optional[str] = None,
|
|
owner: Optional[str] = None,
|
|
# New-schema fields (optional; fall back to old shape if absent)
|
|
name: Optional[str] = None,
|
|
description: Optional[str] = None,
|
|
category: str = "general",
|
|
when_to_use: Optional[str] = None,
|
|
procedure: Optional[List[str]] = None,
|
|
pitfalls: Optional[List[str]] = None,
|
|
verification: Optional[List[str]] = None,
|
|
platforms: Optional[List[str]] = None,
|
|
requires_toolsets: Optional[List[str]] = None,
|
|
fallback_for_toolsets: Optional[List[str]] = None,
|
|
status: str = "draft",
|
|
version: str = "1.0.0",
|
|
) -> Dict:
|
|
# Normalize name
|
|
nm = slugify(name or title or description or "skill")
|
|
|
|
# Free dedup-at-creation (always, no API): for LLM-authored skills,
|
|
# skip if a near-identical skill already exists (Jaccard over
|
|
# name+description+when_to_use+procedure). User-authored skills are
|
|
# never auto-skipped — a human asked for it. The every-X AI audit
|
|
# handles the fuzzier near-duplicates this cheap check won't catch.
|
|
_all = self.load_all()
|
|
_dedup_pool = _all if owner is None else [s for s in _all if s.get("owner") == owner]
|
|
if source != "user":
|
|
cand = _tokenize(" ".join([
|
|
nm, (description or title or ""),
|
|
(when_to_use if when_to_use is not None else (problem or "")),
|
|
" ".join(procedure if procedure is not None else (steps or [])),
|
|
]))
|
|
if cand:
|
|
for s in _dedup_pool:
|
|
ex = _tokenize(" ".join([
|
|
s.get("name", ""), s.get("description", ""),
|
|
s.get("when_to_use", ""),
|
|
" ".join(s.get("procedure", []) or []),
|
|
]))
|
|
if _jaccard(cand, ex) >= 0.82:
|
|
# Near-identical — don't grow the library; bump the
|
|
# existing skill's usage and return it so the caller
|
|
# knows it already exists.
|
|
try:
|
|
self.record_use(s["name"], owner=s.get("owner"))
|
|
except Exception:
|
|
pass
|
|
return {**s, "_deduped": True, "_duplicate_of": s.get("name")}
|
|
|
|
# Avoid clobbering an existing skill with the same name
|
|
existing = {s["name"] for s in _all}
|
|
base = nm
|
|
i = 2
|
|
while nm in existing:
|
|
nm = f"{base}-{i}"
|
|
i += 1
|
|
|
|
sk = Skill(
|
|
name=nm,
|
|
description=(description or title or "").strip(),
|
|
version=version,
|
|
category=category or "general",
|
|
tags=list(tags or []),
|
|
platforms=list(platforms or []),
|
|
requires_toolsets=list(requires_toolsets or []),
|
|
fallback_for_toolsets=list(fallback_for_toolsets or []),
|
|
status=status or "draft",
|
|
confidence=float(confidence),
|
|
source=source,
|
|
teacher_model=teacher_model,
|
|
owner=owner,
|
|
when_to_use=(when_to_use if when_to_use is not None else (problem or "")),
|
|
procedure=list(procedure if procedure is not None else (steps or [])),
|
|
pitfalls=list(pitfalls or []),
|
|
verification=list(verification or []),
|
|
body_extra=(solution if solution and not procedure else ""),
|
|
)
|
|
self._write_skill(sk)
|
|
|
|
return sk.to_dict()
|
|
|
|
def import_bundle_from_files(
|
|
self,
|
|
files: Dict[str, str],
|
|
*,
|
|
owner: Optional[str] = None,
|
|
source_url: str = "",
|
|
category: str = "imported",
|
|
) -> Dict:
|
|
"""Install a fetched skill bundle (relative path → text) under skills/."""
|
|
from .skill_importer import SkillImportError, pick_skill_md, _safe_relpath
|
|
from core.atomic_io import atomic_write_text
|
|
|
|
if not files:
|
|
raise SkillImportError("empty bundle")
|
|
_rel, skill_md = pick_skill_md(files)
|
|
sk = Skill.from_markdown(skill_md)
|
|
nm = slugify(sk.name or _rel.split("/")[-2] or "skill")
|
|
cat = slugify(category or sk.category or "imported", fallback="imported")
|
|
|
|
existing = {s["name"] for s in self.load_all()}
|
|
base = nm
|
|
i = 2
|
|
while nm in existing:
|
|
nm = f"{base}-{i}"
|
|
i += 1
|
|
|
|
skill_dir = self._skill_dir(cat, nm)
|
|
os.makedirs(skill_dir, exist_ok=True)
|
|
|
|
# Preserve bundle layout (templates/, references/, etc.) under the skill dir.
|
|
for rel, content in files.items():
|
|
safe = _safe_relpath(rel)
|
|
dest = os.path.join(skill_dir, safe)
|
|
os.makedirs(os.path.dirname(dest), exist_ok=True)
|
|
atomic_write_text(dest, content)
|
|
|
|
sk.name = nm
|
|
sk.category = cat
|
|
sk.owner = owner
|
|
sk.source = "imported"
|
|
if source_url:
|
|
extra = (sk.body_extra or "").strip()
|
|
note = f"Imported from {source_url}"
|
|
sk.body_extra = f"{extra}\n\n{note}".strip() if extra else note
|
|
atomic_write_text(self._skill_file(cat, nm), sk.to_markdown())
|
|
sk.path = self._skill_file(cat, nm)
|
|
return sk.to_dict()
|
|
|
|
def update_skill(self, skill_id: str, updates: Dict, owner: Optional[str] = None) -> bool:
|
|
"""`skill_id` is the slug name. Allows updating any field plus
|
|
renames if `name` changes (file is moved on disk).
|
|
|
|
The call is owner-scoped: it matches a skill on disk only if
|
|
`skill.owner == owner` (string compare; both empty-string and
|
|
None mean "ownerless"). When `owner is None` (the default), the
|
|
call only matches skills whose own `owner` field is empty —
|
|
callers that want to edit an owned skill must pass the matching
|
|
owner explicitly. This prevents a caller with one owner from
|
|
mutating a file owned by another user that happens to share
|
|
the same slug across category directories. The `owner` key in
|
|
`updates` is also ignored — ownership is not an editable field
|
|
via this path; rename or admin tooling is required for that.
|
|
"""
|
|
for path in self._iter_skill_files():
|
|
sk = self._read_skill(path)
|
|
if not sk or sk.name != skill_id:
|
|
continue
|
|
if (sk.owner or "") != (owner or ""):
|
|
continue
|
|
|
|
old_dir = os.path.dirname(path)
|
|
|
|
scalar_keys = (
|
|
"description", "version", "category", "status", "confidence",
|
|
"source", "teacher_model", "when_to_use",
|
|
"body_extra",
|
|
)
|
|
for k in scalar_keys:
|
|
if k in updates:
|
|
setattr(sk, k, updates[k])
|
|
list_keys = ("tags", "procedure", "pitfalls", "verification",
|
|
"platforms", "requires_toolsets", "fallback_for_toolsets")
|
|
for k in list_keys:
|
|
if k in updates:
|
|
setattr(sk, k, list(updates[k] or []))
|
|
|
|
# Old-schema field aliases
|
|
if "title" in updates and "description" not in updates:
|
|
sk.description = updates["title"]
|
|
if "problem" in updates and "when_to_use" not in updates:
|
|
sk.when_to_use = updates["problem"]
|
|
if "solution" in updates and "body_extra" not in updates and not sk.procedure:
|
|
sk.body_extra = updates["solution"]
|
|
if "steps" in updates and "procedure" not in updates:
|
|
sk.procedure = list(updates["steps"] or [])
|
|
|
|
# Rename
|
|
new_name = slugify(updates.get("name") or sk.name)
|
|
if new_name != sk.name:
|
|
sk.name = new_name
|
|
|
|
# Write to potentially new path
|
|
new_path = self._skill_file(sk.category, sk.name)
|
|
if new_path != path:
|
|
# Move the whole skill directory if rename or recategorize
|
|
new_dir = os.path.dirname(new_path)
|
|
if os.path.isdir(new_dir):
|
|
logger.warning(f"Skill rename target exists: {new_dir}")
|
|
return False
|
|
os.makedirs(os.path.dirname(new_dir), exist_ok=True)
|
|
os.rename(old_dir, new_dir)
|
|
# Also rename usage key
|
|
usage = self._load_usage()
|
|
old_usage_key = self._usage_key(skill_id, sk.owner)
|
|
if old_usage_key in usage:
|
|
usage[self._usage_key(sk.name, sk.owner)] = usage.pop(old_usage_key)
|
|
self._save_usage(usage)
|
|
self._write_skill(sk)
|
|
return True
|
|
return False
|
|
|
|
def delete_skill(self, skill_id: str, owner: Optional[str] = None) -> bool:
|
|
for path in self._iter_skill_files():
|
|
sk = self._read_skill(path)
|
|
if not sk or sk.name != skill_id:
|
|
continue
|
|
if (sk.owner or "") != (owner or ""):
|
|
continue
|
|
skill_dir = os.path.dirname(path)
|
|
try:
|
|
# Remove the whole skill dir
|
|
for root, dirs, files in os.walk(skill_dir, topdown=False):
|
|
for f in files:
|
|
os.remove(os.path.join(root, f))
|
|
for d in dirs:
|
|
os.rmdir(os.path.join(root, d))
|
|
os.rmdir(skill_dir)
|
|
except Exception as e:
|
|
logger.warning(f"Failed to remove skill dir {skill_dir}: {e}")
|
|
return False
|
|
usage = self._load_usage()
|
|
usage_key = self._usage_key(skill_id, sk.owner)
|
|
if usage_key in usage:
|
|
del usage[usage_key]
|
|
self._save_usage(usage)
|
|
return True
|
|
return False
|
|
|
|
def record_use(self, skill_id: str, owner: Optional[str] = None) -> None:
|
|
usage = self._load_usage()
|
|
key = self._usage_key(skill_id, owner)
|
|
entry = usage.setdefault(key, {"uses": 0, "last_used": None})
|
|
entry["uses"] = int(entry.get("uses", 0)) + 1
|
|
entry["last_used"] = int(time.time())
|
|
self._save_usage(usage)
|
|
|
|
# ----------------------------------------------------------------------
|
|
# Reading a single skill (used by the skill_view tool)
|
|
# ----------------------------------------------------------------------
|
|
|
|
def read_skill_md(self, name: str, owner: Optional[str] = None) -> Optional[str]:
|
|
for path in self._iter_skill_files():
|
|
sk = self._read_skill(path)
|
|
if not sk or sk.name != name:
|
|
continue
|
|
# Built-in skills are shared, ownerless procedures. ``load``
|
|
# exposes them to every owner, so direct progressive-disclosure
|
|
# reads must apply the same visibility rule as the index/list
|
|
# path. Previously a built-in appeared in `list` but `view`
|
|
# returned not-found for authenticated users.
|
|
if not (
|
|
(sk.owner or "") == (owner or "")
|
|
or (sk.source == "builtin" and not (sk.owner or ""))
|
|
):
|
|
continue
|
|
try:
|
|
with open(path, encoding="utf-8") as f:
|
|
return f.read()
|
|
except Exception:
|
|
return None
|
|
return None
|
|
|
|
def read_skill_reference(self, name: str, ref_path: str, owner: Optional[str] = None) -> Optional[str]:
|
|
"""Read a sub-file under the skill's directory (references/, etc).
|
|
Refuses path traversal."""
|
|
for path in self._iter_skill_files():
|
|
sk = self._read_skill(path)
|
|
if not sk or sk.name != name:
|
|
continue
|
|
if not (
|
|
(sk.owner or "") == (owner or "")
|
|
or (sk.source == "builtin" and not (sk.owner or ""))
|
|
):
|
|
continue
|
|
# allow_root=False refuses the skill directory itself. The old
|
|
# guard compared a realpath-ed target against a raw dirname, so on
|
|
# a host where the skills tree is reached through a symlink (macOS
|
|
# /tmp -> /private/tmp) the two sides never matched and the guard
|
|
# could not fire.
|
|
try:
|
|
target = confine(os.path.dirname(path), ref_path, allow_root=False)
|
|
except (ValueError, OSError):
|
|
return None
|
|
if not os.path.isfile(target):
|
|
return None
|
|
try:
|
|
with open(target, encoding="utf-8") as f:
|
|
return f.read()
|
|
except Exception:
|
|
return None
|
|
return None
|
|
|
|
# ----------------------------------------------------------------------
|
|
# Index — the lightweight summary injected into the system prompt
|
|
# ----------------------------------------------------------------------
|
|
|
|
def index_for(
|
|
self,
|
|
owner: Optional[str] = None,
|
|
*,
|
|
active_toolsets: Optional[List[str]] = None,
|
|
platform: Optional[str] = None,
|
|
) -> List[Dict]:
|
|
"""Return the `[{name, description, category, status}]` list the
|
|
agent sees in its system prompt.
|
|
|
|
Includes built-ins plus user skills that have passed their audit and
|
|
meet the owner's current automatic-approval threshold. A persistent
|
|
``published`` flag is not sufficient: a changed threshold or a legacy
|
|
record must not make an unaudited skill eligible for prompt injection.
|
|
"""
|
|
auto_approve, min_confidence = _approval_policy(owner)
|
|
out = []
|
|
for s in self.load(owner=owner):
|
|
status = s.get("status")
|
|
# Published + None (pre-status legacy) always included.
|
|
# Drafts only if the teacher wrote them.
|
|
if status not in ("published", None):
|
|
if status == "draft" and s.get("source") == "teacher-escalation":
|
|
pass # let it through
|
|
else:
|
|
continue
|
|
# A stale published record must not remain injectable after an
|
|
# audit has recorded a failure. Inconclusive is not a failure.
|
|
audit_verdict = str(s.get("audit_verdict") or "").lower()
|
|
if audit_verdict in {"needs_work", "fail"}:
|
|
continue
|
|
if s.get("source") != "builtin" and auto_approve:
|
|
if status != "published" or audit_verdict != "pass":
|
|
continue
|
|
if _to_float(s.get("confidence"), 0.0) < min_confidence:
|
|
continue
|
|
necessity = s.get("necessity") or {}
|
|
if isinstance(necessity, dict) and necessity.get("necessary") is False:
|
|
continue
|
|
# Platform gating
|
|
if platform and s.get("platforms") and platform not in s["platforms"]:
|
|
continue
|
|
# requires_toolsets: hide unless every required toolset is active.
|
|
# active_toolsets=None means the caller doesn't know the active
|
|
# set (API listings, chat preface) — don't gate in that case;
|
|
# only an explicit list filters.
|
|
req = s.get("requires_toolsets") or []
|
|
if req and active_toolsets is not None and not all(t in active_toolsets for t in req):
|
|
continue
|
|
# fallback_for_toolsets: hide when any of those toolsets is active
|
|
fb = s.get("fallback_for_toolsets") or []
|
|
if fb and active_toolsets and any(t in active_toolsets for t in fb):
|
|
continue
|
|
out.append({
|
|
"name": s["name"],
|
|
"description": s.get("description") or s.get("title", ""),
|
|
"category": s.get("category", "general"),
|
|
"status": status or "published",
|
|
})
|
|
out.sort(key=lambda x: (x["category"], x["name"]))
|
|
return out
|
|
|
|
# ----------------------------------------------------------------------
|
|
# Relevance search (kept for the existing /api/skills/search endpoint
|
|
# and the `manage_skills` action="search"). Now operates on the new
|
|
# field set.
|
|
# ----------------------------------------------------------------------
|
|
|
|
def get_relevant_skills(
|
|
self,
|
|
query: str,
|
|
skills: Optional[List[Dict]] = None,
|
|
threshold: float = 0.3,
|
|
max_items: int = 5,
|
|
min_confidence: float = 0.0,
|
|
available_toolsets: Optional[Iterable[str]] = None,
|
|
platform: Optional[str] = None,
|
|
) -> List[Dict]:
|
|
if skills is None:
|
|
skills = self.load_all()
|
|
if not skills or not query.strip():
|
|
return []
|
|
# Consider published AND draft skills for relevance retrieval.
|
|
# The teacher-escalation loop writes new skills as drafts; the
|
|
# whole point is for the student to find them on the next try
|
|
# without a manual publish click. The UI flags teacher-written
|
|
# entries with a 🎓 badge so users can demote / delete bad
|
|
# ones when they spot them.
|
|
skills = [
|
|
s for s in skills
|
|
if s.get("status") in ("published", "draft")
|
|
and str(s.get("audit_verdict") or "").lower()
|
|
not in {"needs_work", "fail", "skipped"}
|
|
]
|
|
available = set(available_toolsets) if available_toolsets is not None else None
|
|
if available is not None:
|
|
skills = [
|
|
skill for skill in skills
|
|
if all(tool in available for tool in (skill.get("requires_toolsets") or []))
|
|
and not any(tool in available for tool in (skill.get("fallback_for_toolsets") or []))
|
|
]
|
|
if platform:
|
|
skills = [
|
|
skill for skill in skills
|
|
if not skill.get("platforms") or platform in skill.get("platforms", [])
|
|
]
|
|
# Prompt injection is fail-closed for user skills. Built-ins are
|
|
# shipped procedures; every other skill needs a passing audit and a
|
|
# confidence score at the user's current threshold.
|
|
if min_confidence > 0:
|
|
def _passes(s):
|
|
if s.get("source") == "builtin":
|
|
return True
|
|
return (
|
|
s.get("status") == "published"
|
|
and str(s.get("audit_verdict") or "").lower() == "pass"
|
|
and _to_float(s.get("confidence"), 0.0) >= min_confidence
|
|
)
|
|
skills = [s for s in skills if _passes(s)]
|
|
if not skills:
|
|
return []
|
|
|
|
query_tokens = _tokenize(query)
|
|
semantic_scores: Dict[int, float] = {}
|
|
semantic_enabled = str(
|
|
os.environ.get("ODYSSEUS_SKILL_SEMANTIC_RETRIEVAL", "1")
|
|
).strip().lower() not in {"0", "false", "no", "off"}
|
|
if semantic_enabled:
|
|
try:
|
|
from src.skill_index import semantic_skill_scores
|
|
|
|
semantic_scores = semantic_skill_scores(query, skills)
|
|
except Exception as exc:
|
|
logger.debug("Semantic skill retrieval unavailable: %s", exc)
|
|
try:
|
|
semantic_threshold = float(
|
|
os.environ.get("ODYSSEUS_SKILL_SEMANTIC_THRESHOLD", "0.4")
|
|
)
|
|
except (TypeError, ValueError):
|
|
semantic_threshold = 0.4
|
|
semantic_threshold = max(-1.0, min(1.0, semantic_threshold))
|
|
|
|
scored = []
|
|
for position, sk in enumerate(skills):
|
|
text = " ".join([
|
|
sk.get("name", ""),
|
|
sk.get("description", ""),
|
|
sk.get("when_to_use", ""),
|
|
" ".join(sk.get("tags", []) or []),
|
|
" ".join(sk.get("procedure", []) or []),
|
|
])
|
|
lexical_score = _jaccard(query_tokens, _tokenize(text))
|
|
for tag in sk.get("tags", []) or []:
|
|
# Match tags as whole tokens, not substrings: `tag in query`
|
|
# boosted e.g. a "ai" tag for any query containing "email".
|
|
tag_tokens = _tokenize(tag)
|
|
if tag_tokens and tag_tokens <= query_tokens:
|
|
lexical_score = max(lexical_score, 0.3) * 1.3
|
|
if query.lower() in (sk.get("description") or "").lower():
|
|
lexical_score = max(lexical_score, 0.6)
|
|
semantic_score = semantic_scores.get(position, -1.0)
|
|
if lexical_score < threshold and semantic_score < semantic_threshold:
|
|
continue
|
|
score = max(lexical_score, semantic_score)
|
|
score *= 1.0 + _to_float(sk.get("confidence"), 0.5) * 0.1
|
|
if sk.get("uses", 0) > 0:
|
|
score *= 1.05
|
|
scored.append((score, sk))
|
|
scored.sort(key=lambda x: x[0], reverse=True)
|
|
return [sk for _, sk in scored[:max_items]]
|