A failed step waits for a person: try to fix it and resume it by itself, a bounded number of times #119

Closed
opened 2026-10-01 23:39:41 -04:00 by cmoriarty · 1 comment
Owner

Run run_01M3WX1VW43K77HWNQ328RBGJK (see #116) failed at agent.test.browser at 02:51 UTC and stayed failed. Nothing tried again, although the step had passed checks C1 to C8 and its work was all in the worktree. Resume is a button, so every failure waits for someone to press it.

Asked for: when an agent step fails for a reason a second try can fix, Braid tries again by itself, a bounded number of times, before it lets the run fail.

  • Which failures. The step ran out of time, ran out of context or was cut off at the output limit, or it ended without writing the file its check names. Not a report that is there and fails its check (a browser check that says fail), a rejected gate, an error in the pipeline itself, or a step with an irreversible effect: those want a person.
  • How. The step goes back as a new attempt that continues from the worktree, as Resume does, before anything after it is skipped. Its prompt carries a briefing: what failed and why, what the last attempt said and what its subagents reported, so finished work is not redone. A step that ran out of time gets twice the time.
  • Bounds. Two tries per step and eight per run. A setting on the Settings page sets the tries; 0 turns it off.
  • Seen. Each try is a row in the transcript, and a failure that will be tried again is notified quietly, so nobody is woken for it.
Run run_01M3WX1VW43K77HWNQ328RBGJK (see #116) failed at agent.test.browser at 02:51 UTC and stayed failed. Nothing tried again, although the step had passed checks C1 to C8 and its work was all in the worktree. Resume is a button, so every failure waits for someone to press it. Asked for: when an agent step fails for a reason a second try can fix, Braid tries again by itself, a bounded number of times, before it lets the run fail. - **Which failures.** The step ran out of time, ran out of context or was cut off at the output limit, or it ended without writing the file its check names. Not a report that is there and fails its check (a browser check that says fail), a rejected gate, an error in the pipeline itself, or a step with an irreversible effect: those want a person. - **How.** The step goes back as a new attempt that continues from the worktree, as Resume does, before anything after it is skipped. Its prompt carries a briefing: what failed and why, what the last attempt said and what its subagents reported, so finished work is not redone. A step that ran out of time gets twice the time. - **Bounds.** Two tries per step and eight per run. A setting on the Settings page sets the tries; 0 turns it off. - **Seen.** Each try is a row in the transcript, and a failure that will be tried again is notified quietly, so nobody is woken for it.
Author
Owner

Shipped and deployed, and archived as failure-mediation. Commits ea9ae67, 9e1f3d5 and 91ad99f, live since ac2caaf (2026-10-02 05:12 UTC), the last in 3bc80ce (05:48 UTC).

When an agent step's attempt fails, the engine now puts it back as a new attempt that continues in the worktree, as Resume does, before the steps after it are skipped and before the run is judged, so the run stays running. The failure carries the plan (mediation on osf.step.failed) and the tick carries it out (osf.step.mediated), so a restart in between loses nothing.

  • Retried: a time limit, a context budget, an output limit, and a completion check that failed. Not retried: an attempt the operator stopped, a rejected gate, an error in the pipeline or an executor error whose cause is unknown, a scripted step, an attempt with an irreversible effect. agent.test.browser and agent.code-review carry extra: {mediate: missing}, so a report that is there and says fail stays failed for a person; only a missing report is retried (or a time-out).
  • Bounds: the new setting "Retries of a failed step" (0 to 5, 2 by default, 0 turns it off) per step, and four times it per run. It is read when a step fails, so it can be changed under a running run.
  • A try that follows a time-out gets twice the time, never more than four times the node's own (3,600, 7,200, 14,400 s). The limit is on the log, so a run on a pinned pipeline gets it too.
  • The new attempt's prompt carries a briefing: which try it is, the failure in its own words, the files that were missing, the agent's last message, and the last two reports of its subagents (which is where "C1 to C8 all passed" was), and to write the files first. A person's Resume of a failed step is briefed the same way, without a try number. A failure that will be retried is not notified; the last one and the failed run are.
  • The transcript shows ↻ trying agent.test.browser again (1 of 2): stopped at its 1h time limit; …; 2h this time in amber after the failure it answers.

Seen on production: run #36 was resumed after the deploy. Its agent.test.browser started as attempt 2 continuing in place, with its deadline recorded and the briefing in its first prompt, skipped the checks that had passed, and wrote its report in 21 minutes where it had run out of its hour before. The run went on to merge cmoriarty/scratch#13 into develop and close cmoriarty/scratch#5. The engine's own retry has not had a failure to retry on production yet; it ran end to end against a real opencode and a fake model (retry, briefing on the wire, the row, the setting changed under a running run).

One thing found on the way, fixed in 91ad99f: the briefing quotes subagent reports, and one of them said every browser tool was pinned to one tab, which was true until #118 shipped that night. The primary copied that into the next subagent's task as a fact. The briefing now says after its reports that they describe what their writers had to work with then. It does not diagnose: a model that reads the failure first would be the next tier, behind this one.

Shipped and deployed, and archived as `failure-mediation`. Commits ea9ae67, 9e1f3d5 and 91ad99f, live since ac2caaf (2026-10-02 05:12 UTC), the last in 3bc80ce (05:48 UTC). When an agent step's attempt fails, the engine now puts it back as a new attempt that continues in the worktree, as Resume does, before the steps after it are skipped and before the run is judged, so the run stays `running`. The failure carries the plan (`mediation` on `osf.step.failed`) and the tick carries it out (`osf.step.mediated`), so a restart in between loses nothing. - Retried: a time limit, a context budget, an output limit, and a completion check that failed. Not retried: an attempt the operator stopped, a rejected gate, an error in the pipeline or an executor error whose cause is unknown, a scripted step, an attempt with an irreversible effect. `agent.test.browser` and `agent.code-review` carry `extra: {mediate: missing}`, so a report that is there and says `fail` stays failed for a person; only a missing report is retried (or a time-out). - Bounds: the new setting "Retries of a failed step" (0 to 5, 2 by default, 0 turns it off) per step, and four times it per run. It is read when a step fails, so it can be changed under a running run. - A try that follows a time-out gets twice the time, never more than four times the node's own (3,600, 7,200, 14,400 s). The limit is on the log, so a run on a pinned pipeline gets it too. - The new attempt's prompt carries a briefing: which try it is, the failure in its own words, the files that were missing, the agent's last message, and the last two reports of its subagents (which is where "C1 to C8 all passed" was), and to write the files first. A person's Resume of a failed step is briefed the same way, without a try number. A failure that will be retried is not notified; the last one and the failed run are. - The transcript shows `↻ trying agent.test.browser again (1 of 2): stopped at its 1h time limit; …; 2h this time` in amber after the failure it answers. Seen on production: run #36 was resumed after the deploy. Its `agent.test.browser` started as attempt 2 continuing in place, with its deadline recorded and the briefing in its first prompt, skipped the checks that had passed, and wrote its report in 21 minutes where it had run out of its hour before. The run went on to merge cmoriarty/scratch#13 into develop and close cmoriarty/scratch#5. The engine's own retry has not had a failure to retry on production yet; it ran end to end against a real opencode and a fake model (retry, briefing on the wire, the row, the setting changed under a running run). One thing found on the way, fixed in 91ad99f: the briefing quotes subagent reports, and one of them said every browser tool was pinned to one tab, which was true until #118 shipped that night. The primary copied that into the next subagent's task as a fact. The briefing now says after its reports that they describe what their writers had to work with then. It does not diagnose: a model that reads the failure first would be the next tier, behind this one.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
cmoriarty/braid#119
No description provided.