fix(runtime): verify process identity before any teardown signal

A recorded pid is a claim, not a handle. The containment grant store, the
background-job store and the Cookbook task list all outlive the process that
wrote them — deliberately, so a restart keeps a job and its result — and the
kernel reuses pids. Any teardown driven off one of those records can therefore
land on a process we never started. ODY-86 was exactly this, and the Cookbook
survivor sweep still terminated any process whose full command line matched a
tracked one, which is the same mistake spelled differently.

Identity is (pid, start token). The token comes from /proc/<pid>/stat on Linux,
ps -o lstart= on macOS and the BSDs, and GetProcessTimes on Windows; the kernel
will not hand a pid to a process that started earlier, so comparing the token
recorded at launch against the token read now answers "is this still ours"
without a handle or a supervisor. verify() returns owned, gone, foreign or
unverifiable, and only owned permits a signal.

Keeping "unverifiable" out of the other two is the point. Process inspection
has broken off Linux four times here — ODY-70, -86, -94, -99 — every time
because an absent mechanism read as a successful answer. Folding it into "ours"
signals strangers; folding it into "gone" abandons live processes. It is a
containment failure and every caller treats it as one.

Wired into the three places that signal:

- containment.release() gates a grant recovered from the durable store, and
  leaves an in-process teardown alone, where the caller holds the child and no
  identity question arises. The verdict lands on the record, so "why is this
  grant still here" is answerable afterwards.

- A startup reaper. Nothing read either store before, so a crashed run left
  every grant permanently active and every job permanently running, and the
  first thing to touch such a record was a teardown aimed at a reassigned pid.
  The two stores get opposite treatment: an orphaned grant has no caller left
  and is torn down, while a detached job is documented to survive a restart and
  is only corrected, never killed.

- The Cookbook sweep takes its ownership from the tmux pane's process tree,
  captured before the kill destroys the only link between a surviving model
  server and the session that started it. A process that merely matches the
  tracked command line is now reported rather than killed: the Cookbook
  composed that command line, so an identical one is just as likely to be a
  server the user started by hand. The sweep also runs on hosts with no procfs
  instead of silently skipping, and says so when it could not look at all.
This commit is contained in:
Léo
2026-10-01 18:55:50 +02:00
parent 33fa27b4c8
commit c004a26d46
9 changed files with 1908 additions and 67 deletions
+158
View File
@@ -0,0 +1,158 @@
"""Startup reconciliation for processes a previous run left behind.
Two stores in this tree outlive the process that wrote them, on purpose:
``data/containment_grants.json`` so a restart can reap rather than orphan, and
``data/bg_jobs.json`` so a restart never loses a detached job or its result.
Until now nothing read either of them at startup. A crashed or restarted server
therefore left every grant permanently "active" and every background job
permanently "running", and the first thing to touch one of those records was a
teardown aimed at a pid that had been reassigned in the meantime.
This module runs once, during startup, before anything of this run exists. That
timing is what makes its rules safe: every record it sees was written by an
earlier run, so "I cannot identify this process" is information about a previous
run's child and not about one of ours.
The two stores get **opposite** treatment, which is the whole reason this is a
module and not a loop:
* A **containment grant** is tied to a tool call that no longer has a caller.
A live process under an abandoned grant is by definition an orphan, so it is
torn down.
* A **background job** is detached deliberately and is documented to survive a
uvicorn restart. Killing one here would break the feature, so its record is
only corrected, never reaped. What gets fixed is identity: a job whose pid now
belongs to someone else is retired so that nothing later signals the stranger.
Fail closed in both: a signal requires a positive identity from
:mod:`src.process_ownership`, and every other verdict is recorded rather than
acted on. Containment that cannot identify its target is not containment, and
the honest failure is a visible orphan rather than a dead bystander.
"""
from __future__ import annotations
import logging
from typing import Any, Dict
from src import process_ownership
logger = logging.getLogger(__name__)
def reap_containment_grants() -> Dict[str, Any]:
"""Tear down or retire every grant a previous run left active.
Per grant: a verified live process is torn down through
:func:`src.containment.reap_record`; a grant whose process is gone is
dropped; a grant naming a pid that is now someone else's is dropped
*without a signal*, because the only thing left to do with it is stop
believing it. A grant that cannot be verified at all is **kept**, so the
orphan stays visible in ``active_grants()`` instead of being quietly
written off as handled.
"""
from src import containment
report: Dict[str, Any] = {
"seen": 0, "torn_down": 0, "already_gone": 0,
"foreign": 0, "unverifiable": 0, "failed": 0,
}
try:
records = containment.active_grants()
except Exception:
logger.warning("process_reaper: containment grant store unreadable", exc_info=True)
return report
for record in records:
report["seen"] += 1
grant_id = str(record.get("id") or "")
if record.get("external"):
# Nothing local ever ran, so there is nothing local to reap.
containment.forget(grant_id)
report["already_gone"] += 1
continue
verdict = process_ownership.verify_record(record)
if verdict == process_ownership.GONE:
containment.forget(grant_id)
report["already_gone"] += 1
continue
if verdict == process_ownership.FOREIGN:
logger.warning(
"process_reaper: grant %s named pid %s, which now belongs to a "
"different process; dropping the record unsignalled",
grant_id, record.get("pid"),
)
containment.forget(grant_id)
report["foreign"] += 1
continue
if verdict == process_ownership.UNVERIFIABLE:
logger.error(
"process_reaper: grant %s (pid %s, owner %s) cannot be verified "
"via %s; leaving it active and unsignalled — this is a "
"containment failure, not a clean start",
grant_id, record.get("pid"), record.get("owner"),
process_ownership.inspection_mechanism(),
)
report["unverifiable"] += 1
continue
try:
outcome = containment.reap_record(record)
except Exception:
logger.warning("process_reaper: tearing down grant %s failed", grant_id, exc_info=True)
report["failed"] += 1
continue
if outcome.dead:
containment.forget(grant_id)
report["torn_down"] += 1
else:
logger.error(
"process_reaper: grant %s survived teardown; survivors=%s",
grant_id, list(outcome.survivors),
)
report["failed"] += 1
return report
def reap_bg_jobs() -> Dict[str, Any]:
"""Correct the identity of background jobs a previous run launched.
Deliberately kills nothing: a ``#!bg`` job is detached so that it outlives
the request *and* the server, and the store exists so its result is still
collected afterwards. The defect being closed is narrower — a record whose
pid has been reassigned will be signalled by the max-runtime reaper an hour
later, and that signal lands on whatever now holds the pid.
"""
from src import bg_jobs
try:
return bg_jobs.disown_unverified()
except Exception:
logger.warning("process_reaper: background job store unreadable", exc_info=True)
return {"seen": 0, "retired": 0, "kept": 0}
def reap_orphans() -> Dict[str, Any]:
"""Run both reconciliations. Returns a report; raises nothing.
Blocking: a teardown escalates SIGTERM → grace → SIGKILL and waits for the
process to actually go. Call it off the event loop.
"""
report = {
"mechanism": process_ownership.inspection_mechanism(),
"grants": reap_containment_grants(),
"bg_jobs": reap_bg_jobs(),
}
if report["mechanism"] == process_ownership.MECHANISM_NONE:
logger.error(
"process_reaper: this host offers no process inspection; no orphan "
"from a previous run can be identified or reaped"
)
logger.info("process_reaper: startup reconciliation %s", report)
return report
async def reap_orphans_at_startup() -> Dict[str, Any]:
""":func:`reap_orphans` off the event loop, for an app startup task."""
import asyncio
return await asyncio.to_thread(reap_orphans)