A restart of osfd throws away running steps' work: recovery should continue in place, as resume does #70

Closed
opened 2026-09-27 00:02:14 -04:00 by cmoriarty · 1 comment
Owner

Problem

A restart of osfd throws away the work of every agent step that was running. On 2026-09-26 at 23:11 the deploy of #68 restarted osfd while openspec.apply[1] of run_01M3GB3KN38YT9TZJ7KA7RQ5M0 was 15 minutes in. Recovery closed attempt 1 as Orphaned, and attempt 2's entry checkpoint is the same tree as attempt 1's (105ce4f8d9): the worktree was rolled back, and the step started over from nothing.

That's the path #60 left alone. #60 made a resume continue in place: keep the worktree, and start a new session seeded with the last handoff plus git status/git diff. Its design listed "restart recovery and the stall watchdog keep rolling back" as a non-goal. Every deploy that doesn't wait for idle, and every crash, pays for it.

Why continuing in place is right here

A restart kills the run's opencode serve too (it runs inside the osfd container), so the old session can't be reattached. But the worktree survives the restart intact. That's exactly the state resume continues from.

Proposal

  • When recovery puts an agent step back after a restart (the Orphaned path, and the recovery gate's "resume" answer), mark its ready event continue: true, as resume_run does. _run_attempt then skips the rollback, and the executor seeds the session with the step's latest recorded handoff and the instruction to read git status and git diff first. That machinery is all #60's, reused as it is.
  • The stall watchdog's redo keeps rolling back: a stall means the attempt went wrong, not that it was interrupted.

Done when

  • An agent step running when osfd restarts continues in its next attempt with the worktree as it was, seeded with its last handoff if it has one.
  • A stall redo still rolls back.
## Problem A restart of osfd throws away the work of every agent step that was running. On 2026-09-26 at 23:11 the deploy of #68 restarted osfd while `openspec.apply[1]` of `run_01M3GB3KN38YT9TZJ7KA7RQ5M0` was 15 minutes in. Recovery closed attempt 1 as `Orphaned`, and attempt 2's entry checkpoint is the same tree as attempt 1's (`105ce4f8d9`): the worktree was rolled back, and the step started over from nothing. That's the path #60 left alone. #60 made a **resume** continue in place: keep the worktree, and start a new session seeded with the last handoff plus `git status`/`git diff`. Its design listed "restart recovery and the stall watchdog keep rolling back" as a non-goal. Every deploy that doesn't wait for idle, and every crash, pays for it. ## Why continuing in place is right here A restart kills the run's `opencode serve` too (it runs inside the osfd container), so the old session can't be reattached. But the worktree survives the restart intact. That's exactly the state resume continues from. ## Proposal - When recovery puts an agent step back after a restart (the `Orphaned` path, and the recovery gate's "resume" answer), mark its ready event `continue: true`, as `resume_run` does. `_run_attempt` then skips the rollback, and the executor seeds the session with the step's latest recorded handoff and the instruction to read `git status` and `git diff` first. That machinery is all #60's, reused as it is. - The stall watchdog's redo keeps rolling back: a stall means the attempt went wrong, not that it was interrupted. ## Done when - An agent step running when osfd restarts continues in its next attempt with the worktree as it was, seeded with its last handoff if it has one. - A stall redo still rolls back.
Author
Owner

Shipped in PR #72 (850dbab), live on production since 2026-09-27 01:57. Archived as openspec/changes/archive/2026-09-27-restart-continues-in-place (5cd985a), with restart-recovery updated.

What shipped:

  • When the session dies with osfd ("died mid-turn"), recovery now rules resume for every agent step. resume and reprompt put continue: true on the ready event. The next attempt keeps the worktree, and its session is seeded with the last handoff and told to read git status/git diff first. That's #60's machinery, reused as it was.
  • A recovery gate answered Resume continues the same way.
  • These still roll back: a wedged session, a stall redo, Redo from entry, and scripted steps.
  • resume_safe is retired. It no longer decided anything, since resume and rollback_retry both rolled back. The built-in pipelines and the generated apply chain drop it, and the loader ignores it where an existing .osf/pipeline.yaml still writes it.
  • Also: tools/test.sh full with OSF_URL used to die on macOS bash 3.2 before its browser step. That's fixed.

Checked on production: I restarted osfd at 02:38:47 while openspec.apply[4] of run_01M3GB3KN38YT9TZJ7KA7RQ5M0 had been editing, in attempt 4 (git diff HEAD fingerprint 97eb98f, 20 files; entry tree dc07cd6).

  • Recovery ruled resume, and the ready event carried continue.
  • Attempt 5 opened continued: true, with entry tree 4ec9c58, not dc07cd6.
  • The worktree was still 97eb98f.

An earlier restart at 02:01 took the same path, but that attempt had only been reading, so it couldn't show that no rollback happened.

Found on the way: #73 (the digest showed a dead attempt's tool call as running).

Shipped in PR #72 (`850dbab`), live on production since 2026-09-27 01:57. Archived as `openspec/changes/archive/2026-09-27-restart-continues-in-place` (`5cd985a`), with `restart-recovery` updated. **What shipped:** - When the session dies with osfd ("died mid-turn"), recovery now rules `resume` for every agent step. `resume` and `reprompt` put `continue: true` on the ready event. The next attempt keeps the worktree, and its session is seeded with the last handoff and told to read `git status`/`git diff` first. That's #60's machinery, reused as it was. - A recovery gate answered **Resume** continues the same way. - These still roll back: a wedged session, a stall redo, **Redo from entry**, and scripted steps. - `resume_safe` is retired. It no longer decided anything, since `resume` and `rollback_retry` both rolled back. The built-in pipelines and the generated apply chain drop it, and the loader ignores it where an existing `.osf/pipeline.yaml` still writes it. - Also: `tools/test.sh full` with `OSF_URL` used to die on macOS bash 3.2 before its browser step. That's fixed. **Checked on production:** I restarted osfd at 02:38:47 while `openspec.apply[4]` of `run_01M3GB3KN38YT9TZJ7KA7RQ5M0` had been editing, in attempt 4 (`git diff HEAD` fingerprint `97eb98f`, 20 files; entry tree `dc07cd6`). - Recovery ruled `resume`, and the ready event carried `continue`. - Attempt 5 opened `continued: true`, with entry tree `4ec9c58`, not `dc07cd6`. - The worktree was still `97eb98f`. An earlier restart at 02:01 took the same path, but that attempt had only been reading, so it couldn't show that no rollback happened. **Found on the way:** #73 (the digest showed a dead attempt's tool call as running).
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
cmoriarty/braid#70
No description provided.