Rename "Abandon run" to "Stop run", and let a stopped run be resumed #51

Closed
opened 2026-09-26 16:16:34 -04:00 by cmoriarty · 2 comments
Owner

Problem

Redo on a step of an abandoned run looks like it works and does nothing. On soundcheck#32's run (run_01M3F8PTF3EWBCB721BP869E9N, minimalist) I abandoned it during openspec.apply[4], then wanted to redo that step with a larger context per agent.

  • The Redo button is shown and POST /api/steps/{id}/redo accepts it: step_redo never checks the run's status.
  • The scheduler supersedes the attempt, rolls the worktree back to the step's entry checkpoint, and marks the step ready.
  • Nothing then dispatches it: the tick only reads runs in running/blocked (scheduler.py ~L841). Abandoning also cancelled every downstream step and stopped the run's opencode serve.

The step sits at ready forever, and the run cannot be continued. Starting a new run from the issue throws away everything that already succeeded.

Proposal

  1. "Abandon run" becomes "Stop run". Same effect as today (running steps stopped, pending ones cancelled, the run's opencode server shut down), but the name no longer says it is final. The status word changes to match (stopped), with existing abandoned runs read as stopped.
  2. A stopped run can be resumed. A Resume run button in the stage header, beside where Stop was, and on the ledger row:
    • the run goes back to running and its opencode serve is restarted;
    • steps cancelled by the stop go back to pending, and the step that was stopped mid-attempt gets a fresh attempt (the stopped one is kept as superseded, as a redo does);
    • steps that had already succeeded stay succeeded.
  3. Redo on a stopped run resumes it. Redoing a step of a stopped run resumes the run and redoes that step, instead of silently parking it. Redo stays refused (409, and hidden in the UI) on runs that succeeded.
  4. Archiving still applies to stopped runs; an archived run must be restored before it can be resumed.

Open questions

  • Should Resume re-read the current settings (e.g. a new context per agent), or keep the settings the run started with? My intent in the case above was to pick up the new setting, but the settings spec says a run keeps the settings it started with.
  • Is a failed run also resumable, or only stopped ones?

Done when

  • Stopping then resuming a run continues it from where it stopped, in a real browser.
  • Redo on a stopped run's step actually runs the step.
  • No button in the console accepts an action the scheduler will not carry out.
## Problem Redo on a step of an abandoned run looks like it works and does nothing. On soundcheck#32's run (`run_01M3F8PTF3EWBCB721BP869E9N`, minimalist) I abandoned it during `openspec.apply[4]`, then wanted to redo that step with a larger context per agent. - The Redo button is shown and `POST /api/steps/{id}/redo` accepts it: `step_redo` never checks the run's status. - The scheduler supersedes the attempt, **rolls the worktree back** to the step's entry checkpoint, and marks the step `ready`. - Nothing then dispatches it: the tick only reads runs in `running`/`blocked` (`scheduler.py` ~L841). Abandoning also cancelled every downstream step and stopped the run's `opencode serve`. The step sits at `ready` forever, and the run cannot be continued. Starting a new run from the issue throws away everything that already succeeded. ## Proposal 1. **"Abandon run" becomes "Stop run".** Same effect as today (running steps stopped, pending ones cancelled, the run's opencode server shut down), but the name no longer says it is final. The status word changes to match (`stopped`), with existing `abandoned` runs read as stopped. 2. **A stopped run can be resumed.** A **Resume run** button in the stage header, beside where Stop was, and on the ledger row: - the run goes back to `running` and its `opencode serve` is restarted; - steps cancelled by the stop go back to `pending`, and the step that was stopped mid-attempt gets a fresh attempt (the stopped one is kept as superseded, as a redo does); - steps that had already succeeded stay succeeded. 3. **Redo on a stopped run resumes it.** Redoing a step of a stopped run resumes the run and redoes that step, instead of silently parking it. Redo stays refused (409, and hidden in the UI) on runs that `succeeded`. 4. Archiving still applies to stopped runs; an archived run must be restored before it can be resumed. ## Open questions - Should Resume re-read the current settings (e.g. a new context per agent), or keep the settings the run started with? My intent in the case above was to pick up the new setting, but the settings spec says a run keeps the settings it started with. - Is a `failed` run also resumable, or only stopped ones? ## Done when - Stopping then resuming a run continues it from where it stopped, in a real browser. - Redo on a stopped run's step actually runs the step. - No button in the console accepts an action the scheduler will not carry out.
Author
Owner

Also in scope: redesign the Redo confirmation

The redo stage (RedoStage in ui/src/components/StepInspector.tsx) needs a redesign, and it belongs here because it shares words with Stop and Resume. What's wrong with it today:

  • Two type systems. Prose paragraphs use the human font at one size, and the section bodies are monospace at another. The eyebrow headings are a third style. Nothing lines up.
  • Text-dense. Every section is a full sentence, even when the answer is "none": "Nothing has left this machine on this attempt — no push, no comment, no pull request." and "Nothing below this has succeeded yet." An empty section should be quiet, or absent.
  • Confusing numbers. "0 files were present at entry, and anything newer goes" reads like a count of what will be deleted, which it isn't.
  • Stale words. It says attempt 1 "is abandoned", which will collide with Stop. It should say superseded.
  • Design notes shown as UI. The footer ("mode: reset · the session is never reused — a failed attempt left in context re-anchors a model whose named weakness is forgetfulness") is design rationale, not something to act on. It belongs behind an ⓘ.

What it should be

One screen answering three questions at a glance, in the console's own type:

  1. What happens: attempt 1 → attempt 2, from a clean session.
  2. What it undoes: the rollback target (checkpoint, time) and what gets removed, as a short list, or "nothing" when that's the case.
  3. What it can't undo: irreversible effects and downstream steps that will go stale, shown only when there are any, and loudly when there are.

Then the optional reason and the confirm button. On a stopped run, the button reads "Resume and redo". Explanations go behind ⓘ tips, as elsewhere after #49.

## Also in scope: redesign the Redo confirmation The redo stage (`RedoStage` in `ui/src/components/StepInspector.tsx`) needs a redesign, and it belongs here because it shares words with Stop and Resume. What's wrong with it today: - **Two type systems.** Prose paragraphs use the human font at one size, and the section bodies are monospace at another. The eyebrow headings are a third style. Nothing lines up. - **Text-dense.** Every section is a full sentence, even when the answer is "none": "Nothing has left this machine on this attempt — no push, no comment, no pull request." and "Nothing below this has succeeded yet." An empty section should be quiet, or absent. - **Confusing numbers.** "0 files were present at entry, and anything newer goes" reads like a count of what will be deleted, which it isn't. - **Stale words.** It says attempt 1 "is abandoned", which will collide with Stop. It should say superseded. - **Design notes shown as UI.** The footer ("mode: reset · the session is never reused — a failed attempt left in context re-anchors a model whose named weakness is forgetfulness") is design rationale, not something to act on. It belongs behind an ⓘ. ### What it should be One screen answering three questions at a glance, in the console's own type: 1. **What happens:** attempt 1 → attempt 2, from a clean session. 2. **What it undoes:** the rollback target (checkpoint, time) and what gets removed, as a short list, or "nothing" when that's the case. 3. **What it can't undo:** irreversible effects and downstream steps that will go stale, shown only when there are any, and loudly when there are. Then the optional reason and the confirm button. On a stopped run, the button reads "Resume and redo". Explanations go behind ⓘ tips, as elsewhere after #49.
Author
Owner

Shipped in #53 (7c9da17) and archived as openspec/changes/archive/2026-09-26-stop-and-resume-runs.

  • Stop run replaces Abandon run; a stopped run shows as stopped (POST /api/runs/{id}/stop; /abandon kept). Stored status and event names are unchanged, so old runs replay as before.
  • Resume run (stage header, ledger row, POST /api/runs/{id}/resume) re-opens a stopped or failed run and puts back to pending what the stop ended, plus failed and unreachable steps of a failed run. Each runs as a new attempt from its entry checkpoint; succeeded steps stay. Refused for succeeded, live, archived and reclaimed runs, with a reason.
  • Redo on a stopped run now continues the whole run; the button reads "Resume and redo".
  • Redo confirmation redesigned: what happens / what it undoes / what it cannot undo, one type style, empty sections left out, the mechanism behind ⓘ.
  • Open questions, as decided: a resumed step uses the current context per agent (the run keeps its backend and model), and failed runs are resumable as well as stopped ones.

Root cause, for the record: redo already re-opened ended runs, but a stop cancelled every later step and nothing ever un-cancelled them, so the run dead-ended one step after the redo.

Shipped in #53 (7c9da17) and archived as `openspec/changes/archive/2026-09-26-stop-and-resume-runs`. - **Stop run** replaces Abandon run; a stopped run shows as **stopped** (`POST /api/runs/{id}/stop`; `/abandon` kept). Stored status and event names are unchanged, so old runs replay as before. - **Resume run** (stage header, ledger row, `POST /api/runs/{id}/resume`) re-opens a stopped or failed run and puts back to pending what the stop ended, plus failed and unreachable steps of a failed run. Each runs as a new attempt from its entry checkpoint; succeeded steps stay. Refused for succeeded, live, archived and reclaimed runs, with a reason. - **Redo on a stopped run** now continues the whole run; the button reads "Resume and redo". - **Redo confirmation** redesigned: what happens / what it undoes / what it cannot undo, one type style, empty sections left out, the mechanism behind ⓘ. - Open questions, as decided: a resumed step uses the *current* context per agent (the run keeps its backend and model), and failed runs are resumable as well as stopped ones. Root cause, for the record: redo already re-opened ended runs, but a stop cancelled every later step and nothing ever un-cancelled them, so the run dead-ended one step after the redo.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
cmoriarty/braid#51
No description provided.