A failed step waits for a person: try to fix it and resume it by itself, a bounded number of times #119
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Run run_01M3WX1VW43K77HWNQ328RBGJK (see #116) failed at agent.test.browser at 02:51 UTC and stayed failed. Nothing tried again, although the step had passed checks C1 to C8 and its work was all in the worktree. Resume is a button, so every failure waits for someone to press it.
Asked for: when an agent step fails for a reason a second try can fix, Braid tries again by itself, a bounded number of times, before it lets the run fail.
Shipped and deployed, and archived as
failure-mediation. Commitsea9ae67,9e1f3d5and91ad99f, live sinceac2caaf(2026-10-02 05:12 UTC), the last in3bc80ce(05:48 UTC).When an agent step's attempt fails, the engine now puts it back as a new attempt that continues in the worktree, as Resume does, before the steps after it are skipped and before the run is judged, so the run stays
running. The failure carries the plan (mediationonosf.step.failed) and the tick carries it out (osf.step.mediated), so a restart in between loses nothing.agent.test.browserandagent.code-reviewcarryextra: {mediate: missing}, so a report that is there and saysfailstays failed for a person; only a missing report is retried (or a time-out).↻ trying agent.test.browser again (1 of 2): stopped at its 1h time limit; …; 2h this timein amber after the failure it answers.Seen on production: run #36 was resumed after the deploy. Its
agent.test.browserstarted as attempt 2 continuing in place, with its deadline recorded and the briefing in its first prompt, skipped the checks that had passed, and wrote its report in 21 minutes where it had run out of its hour before. The run went on to merge cmoriarty/scratch#13 into develop and close cmoriarty/scratch#5. The engine's own retry has not had a failure to retry on production yet; it ran end to end against a real opencode and a fake model (retry, briefing on the wire, the row, the setting changed under a running run).One thing found on the way, fixed in
91ad99f: the briefing quotes subagent reports, and one of them said every browser tool was pinned to one tab, which was true until #118 shipped that night. The primary copied that into the next subagent's task as a fact. The briefing now says after its reports that they describe what their writers had to work with then. It does not diagnose: a model that reads the failure first would be the next tier, behind this one.