A node added to a built-in pipeline stalls runs already in flight #102

Closed
opened 2026-09-28 08:59:16 -04:00 by cmoriarty · 2 comments
Owner

Built-in pipelines are built from Braid's code, not read from the run's base commit. So after a deploy, a run admitted earlier loads the new graph, and its step rows were made from the old one.

If the new graph has a node the run has no row for (for example openspec.repair from validation-repair, or script.plan-tick), _promote never promotes anything that needs it (if any(need not in by_node for need in node.needs): continue). The run waits for ever with no reason shown.

No run has hit this yet; it was noticed while shipping validation-repair, and the deploys were timed so no run was mid-graph.

Options:

  1. Pin the built-in graph a run was admitted with, for example by recording the YAML on osf.run.started and loading that on recovery and redo.
  2. When a loaded graph has a node with no row, add its row (pending) on the next tick.

Option 1 keeps a run's process fixed for its life, which is what named-pipelines already promises for repository pipelines.

Built-in pipelines are built from Braid's code, not read from the run's base commit. So after a deploy, a run admitted earlier loads the new graph, and its step rows were made from the old one. If the new graph has a node the run has no row for (for example `openspec.repair` from `validation-repair`, or `script.plan-tick`), `_promote` never promotes anything that needs it (`if any(need not in by_node for need in node.needs): continue`). The run waits for ever with no reason shown. No run has hit this yet; it was noticed while shipping `validation-repair`, and the deploys were timed so no run was mid-graph. Options: 1. Pin the built-in graph a run was admitted with, for example by recording the YAML on `osf.run.started` and loading that on recovery and redo. 2. When a loaded graph has a node with no row, add its row (pending) on the next tick. Option 1 keeps a run's process fixed for its life, which is what named-pipelines already promises for repository pipelines.
Author
Owner

There's a second symptom of the same cause, seen on run #29 (git-flow on cmoriarty/scratch). After the 1d596b7 (#107) and 3346eae (#109) deploys, five of its finished steps were flagged "inputs changed", and the console offered "Re-run 5 stale steps":

Step Its node needed Now needs Changed by
git.commit.proposal script.budget script.gitignore.proposal #107
openspec.apply[1–3] script.review-outcome script.deps #109
agent.test.unit agent.reconcile-tasks script.deps.tests #109

recompute_inputs_digests recomputes each succeeded step's inputs digest from the dependencies its node has in the graph loaded now. The run has no rows for the new nodes, so the digest comes out different, and the step is marked stale although nothing it read changed. agent.reconcile-tasks, whose needs didn't change, wasn't flagged.

Re-running would have repeated about an hour of agent work for nothing. It also couldn't have unstuck the run: git.commit.feature had come to need agent.code-review (#108), for which the run had no row either, so it stayed pending until the run was abandoned.

Option 1, pinning the graph a run was admitted with, fixes both symptoms.

There's a second symptom of the same cause, seen on run #29 (git-flow on `cmoriarty/scratch`). After the 1d596b7 (#107) and 3346eae (#109) deploys, five of its finished steps were flagged **"inputs changed"**, and the console offered "Re-run 5 stale steps": | Step | Its node needed | Now needs | Changed by | |---|---|---|---| | `git.commit.proposal` | `script.budget` | `script.gitignore.proposal` | #107 | | `openspec.apply[1–3]` | `script.review-outcome` | `script.deps` | #109 | | `agent.test.unit` | `agent.reconcile-tasks` | `script.deps.tests` | #109 | `recompute_inputs_digests` recomputes each succeeded step's inputs digest from the dependencies its node has *in the graph loaded now*. The run has no rows for the new nodes, so the digest comes out different, and the step is marked stale although nothing it read changed. `agent.reconcile-tasks`, whose needs didn't change, wasn't flagged. Re-running would have repeated about an hour of agent work for nothing. It also couldn't have unstuck the run: `git.commit.feature` had come to need `agent.code-review` (#108), for which the run had no row either, so it stayed pending until the run was abandoned. Option 1, pinning the graph a run was admitted with, fixes both symptoms.
Author
Owner

Fixed in dab6969 with option 1, as the OpenSpec change pin-run-pipeline. It's archived in 75bb426, updating the named-pipelines requirement "The run records its pipeline and recovery replays it".

  • What's recorded: admission now records a built-in pipeline's text in osf.run.started, projected onto the new run.pipeline_yaml (migration 8). After a restart, and on redo, recovery and resume, Scheduler._pipeline loads that text first. A deploy that adds, removes or rewires built-in steps then changes new runs only. A run in flight keeps its graph, its finished steps keep their inputs, and nothing stalls.
  • Unchanged: a repository's own pipeline was already read at the recorded sha.
  • Fallbacks: runs admitted before this have no record and keep the old behaviour. So does a record this build can no longer load, for example one naming a prompt it has removed, and that case is logged.

Verified:

  • Unit tests: a scheduler test reproduces both symptoms on an unrecorded run after a simulated deploy (a step flagged stale, the feature commit never promoted), and neither on a recorded run. With the fix disabled, the tests fail.
  • Local osfd: a quick-fix run was restarted mid-fix onto a build whose quick-fix had a new step before git.commit.feature. It finished its feature commit on its recorded graph, with no step flagged stale, while a run started on the new build got the new step.
  • Production, after deploying dab6969: a run admitted after the deploy has its pipeline text recorded (12,552 characters). #34, admitted before it, has none, as expected.
Fixed in dab6969 with option 1, as the OpenSpec change `pin-run-pipeline`. It's archived in 75bb426, updating the named-pipelines requirement "The run records its pipeline and recovery replays it". - **What's recorded:** admission now records a built-in pipeline's text in `osf.run.started`, projected onto the new `run.pipeline_yaml` (migration 8). After a restart, and on redo, recovery and resume, `Scheduler._pipeline` loads that text first. A deploy that adds, removes or rewires built-in steps then changes new runs only. A run in flight keeps its graph, its finished steps keep their inputs, and nothing stalls. - **Unchanged:** a repository's own pipeline was already read at the recorded sha. - **Fallbacks:** runs admitted before this have no record and keep the old behaviour. So does a record this build can no longer load, for example one naming a prompt it has removed, and that case is logged. **Verified:** - **Unit tests:** a scheduler test reproduces both symptoms on an unrecorded run after a simulated deploy (a step flagged stale, the feature commit never promoted), and neither on a recorded run. With the fix disabled, the tests fail. - **Local osfd:** a quick-fix run was restarted mid-fix onto a build whose quick-fix had a new step before `git.commit.feature`. It finished its feature commit on its recorded graph, with no step flagged stale, while a run started on the new build got the new step. - **Production, after deploying dab6969:** a run admitted after the deploy has its pipeline text recorded (12,552 characters). #34, admitted before it, has none, as expected.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
cmoriarty/braid#102
No description provided.