The watchdog orphans a working step when its subagent finishes first: it checks the subagent's session, not the step's #141
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
What happened
In run 44 (
run_01M44PHE3XQ59GBHGRW9STWFDR, the second baseline for #139),openspec.apply[3]attempt 1 was orphaned at 01:36:11Z with "osfd died while this attempt was running". osfd had not died: it had been up on one boot since 00:12:23Z. Braid's own liveness watchdog killed a healthy attempt.From osfd's log and the run's
oc.db(UTC):…DMAvfotZ) starts atestsubagent,ses_ef64cb413ffekh1OvfuRzWAJ7d, "Run front-end test and build".openspec.apply[3] is running but nothing is running it — session ses_ef64cb413ffekh1OvfuRzWAJ7d has been missing from the busy map of a server that answers for 20s while the step is still running; orphaning it for §12Why
recovery._session_id, which the watchdog uses, takes the session of the attempt's newest logged event, so that a budget respawn's new session is found. But a subagent's events are logged on its parent's step under the subagent's own session id. Once a subagent finishes, the newest event can name the subagent. The watchdog then checks the subagent:The working primary was never asked about.
Any step can hit this when a subagent finishes and the primary then stays silent for more than about 20 seconds, such as during a long write or while it waits for an inference slot.
Fix
oc.child.startedevent announced for the attempt. Keep newest-first, so a respawn's session still wins.Related: #70 (recovery continues in place), #69 (a step waiting on its own tool call).
Fixed in
watchdog-primary-session. Deployed as25e898dat 02:00:58Z and archived in3431a23.oc.child.startedevent announced for that attempt, nested subagents included. It still takes the newest of the step's own, so a budget respawn's new session still wins. One lookup serves four places:the watchdog found nothing running it: …with what it saw, the same sentence it logs. "osfd died while this attempt was running" stays for attempts a restart interrupted.Verified
testsubagent finishes, then the primary is busy and silent past the grace. They fail on the old lookup and pass on the fix. The existing respawn test still passes.Run 44 was abandoned to deploy this. The baseline for #139 is now run 45, which started on this build.
An update since the fix: