Extract the generic process lifecycle layer (src/process_lifecycle.py)
shared by runtime-owned subprocesses: process identity (pid + boot-bound
start token), identity-bound observation, group and pidfd probes, the
TERM -> verify -> KILL -> verify escalation with re-gating before
escalation, identity-scoped sweeps, and the termination receipt.
Containment, the PTY shell, the Cookbook survivor sweep, the browser
lifecycle, web_tools browser cleanup, kill_process_tree and the startup
reaper consume it while keeping their own ownership semantics.
Safety corrections:
- browser membership and identity are bound in one snapshot; no identity
is recaptured after membership is decided
- web_tools legacy pid-file and profile-match kills signal only verified
identities; browser CLI groups only while their spawn identity verifies
- Cookbook and legacy-tmux descendant capture bind membership to identity
- PTY teardown never signals the server's own process group
- unverifiable processes are reported, never signalled
A recorded pid is a claim, not a handle. The containment grant store, the
background-job store and the Cookbook task list all outlive the process that
wrote them — deliberately, so a restart keeps a job and its result — and the
kernel reuses pids. Any teardown driven off one of those records can therefore
land on a process we never started. ODY-86 was exactly this, and the Cookbook
survivor sweep still terminated any process whose full command line matched a
tracked one, which is the same mistake spelled differently.
Identity is (pid, start token). The token comes from /proc/<pid>/stat on Linux,
ps -o lstart= on macOS and the BSDs, and GetProcessTimes on Windows; the kernel
will not hand a pid to a process that started earlier, so comparing the token
recorded at launch against the token read now answers "is this still ours"
without a handle or a supervisor. verify() returns owned, gone, foreign or
unverifiable, and only owned permits a signal.
Keeping "unverifiable" out of the other two is the point. Process inspection
has broken off Linux four times here — ODY-70, -86, -94, -99 — every time
because an absent mechanism read as a successful answer. Folding it into "ours"
signals strangers; folding it into "gone" abandons live processes. It is a
containment failure and every caller treats it as one.
Wired into the three places that signal:
- containment.release() gates a grant recovered from the durable store, and
leaves an in-process teardown alone, where the caller holds the child and no
identity question arises. The verdict lands on the record, so "why is this
grant still here" is answerable afterwards.
- A startup reaper. Nothing read either store before, so a crashed run left
every grant permanently active and every job permanently running, and the
first thing to touch such a record was a teardown aimed at a reassigned pid.
The two stores get opposite treatment: an orphaned grant has no caller left
and is torn down, while a detached job is documented to survive a restart and
is only corrected, never killed.
- The Cookbook sweep takes its ownership from the tmux pane's process tree,
captured before the kill destroys the only link between a surviving model
server and the session that started it. A process that merely matches the
tracked command line is now reported rather than killed: the Cookbook
composed that command line, so an identical one is just as likely to be a
server the user started by hand. The sweep also runs on hosts with no procfs
instead of silently skipping, and says so when it could not look at all.