Verification steps redo each other: reconcile, the unit tests and the code review re-read and re-run the same change #125

Closed
opened 2026-10-02 15:50:17 -04:00 by cmoriarty · 1 comment
Owner

Split off from #121 (finding 3 of the scratch-run analysis, docs/design/measurements/2026-10-02-scratch-run-efficiency.md).

  • agent.reconcile-tasks (7–12 min) changed nothing in each of the last five scratch runs (23, 25, 26, 36 and 37).
  • agent.test.unit (7–14 min) mostly found that the apply steps had already written the tests, and touched 0–3 files.
  • agent.code-review then re-reads the change and re-runs the suites again.
  • Run 37 ran the full backend suite 13 times and the frontend suite or build 9 times. Its reconcile also ran a browser check, the job agent.test.browser does next.

Together these steps take 25–35 min of every feature run, most of it reasoning, and the suites themselves take seconds.

Ideas:

  • A deterministic reconcile: when every task is ticked and the suite is green at the current tree, skip the agent step, or run it with no thinking.
  • One record of the last suite result for a tree, which later steps read instead of running the suite again.
  • Run the code review alongside the browser check instead of after it.

Needs a design and a measurement.

Split off from #121 (finding 3 of the scratch-run analysis, `docs/design/measurements/2026-10-02-scratch-run-efficiency.md`). - `agent.reconcile-tasks` (7–12 min) changed nothing in each of the last five scratch runs (23, 25, 26, 36 and 37). - `agent.test.unit` (7–14 min) mostly found that the apply steps had already written the tests, and touched 0–3 files. - `agent.code-review` then re-reads the change and re-runs the suites again. - Run 37 ran the full backend suite 13 times and the frontend suite or build 9 times. Its reconcile also ran a browser check, the job `agent.test.browser` does next. Together these steps take 25–35 min of every feature run, most of it reasoning, and the suites themselves take seconds. Ideas: - A deterministic reconcile: when every task is ticked and the suite is green at the current tree, skip the agent step, or run it with no thinking. - One record of the last suite result for a tree, which later steps read instead of running the suite again. - Run the code review alongside the browser check instead of after it. Needs a design and a measurement.
Author
Owner

Shipped and live on production (9ae119c, deployed 2026-10-03). Against the three ideas:

1. A deterministic reconcile — shipped. agent.reconcile-tasks has when: tasks_unfinished, which reads tasks.md. With every task ticked the step is skipped as not needed, with the reason every task is ticked (12 of 12), and script.deps.tests follows at once. A missing or empty tasks.md, or a task outside any numbered group left unticked, runs the step. The "and the suite is green at the current tree" half could not be used: reconcile runs before the unit lane, so there is no green suite for the tree yet. On production's run 41 the condition reads every task is ticked (12 of 12): its reconcile (4.9 minutes, changed nothing) would have been skipped.

2. One record of the last suite result for a tree — shipped for the code review. The unit lane's report is that record, and it now names the code it ran against: code_sha is a digest of the commit and of the content of every change to a file outside .osf/. Unlike the tree digest the freshness check uses, it changes when an already-modified file is edited again, and it does not change when Braid writes its own records; suites-green and its freshness check are untouched. The code review's prompt says, from it, when the suite already passed on exactly the code it is reading (the command, how long it took, and that the lane's lint is a separate matter) and not to run it before changing something; when the code changed since, it asks for one run. Whenever the review runs the suite it is told to run it once, in one place and not in a test subagent as well, under a timeout of four times the lane's duration (at least a minute): run 41's review spent about 20 of its 30.8 minutes on a suite the lane ran in 18 seconds, killed twice by the tool's limits. The unit-test step is told where the scenarios stand (how many have a tagged test and which do not, the lane's own arithmetic; 31 of 31 on run 41's worktree, the figure its report recorded). Not recorded: the apply steps' own suite runs.

3. The code review alongside the browser check — not done. Both edit one worktree, so the review would read a diff the browser check is still changing; overlapping independent steps is #128, whose list has this pair. Keeping the browser check out of reconcile shipped with #127.

Checked: unit tests for the condition, the digest (an edit to an already-modified file, a new file, a commit, a write under .osf/), the scenario status against the lane's report, and both prompts; the fast and full lanes; the reconcile condition end to end in a local osfd (ticked: skipped with the reason; unticked: runs); the prompts read for a real change; and on production through the deployed code. To measure in the next scratch run: reconcile (none), the unit step and the code review against 25–35 minutes between them, and the number of suite runs in the code review's transcript against run 37's 13 backend runs.

Change record: openspec/changes/archive/2026-10-03-verification-without-repeats.

Shipped and live on production (`9ae119c`, deployed 2026-10-03). Against the three ideas: **1. A deterministic reconcile — shipped.** `agent.reconcile-tasks` has `when: tasks_unfinished`, which reads `tasks.md`. With every task ticked the step is skipped as not needed, with the reason `every task is ticked (12 of 12)`, and `script.deps.tests` follows at once. A missing or empty `tasks.md`, or a task outside any numbered group left unticked, runs the step. The "and the suite is green at the current tree" half could not be used: reconcile runs before the unit lane, so there is no green suite for the tree yet. On production's run 41 the condition reads `every task is ticked (12 of 12)`: its reconcile (4.9 minutes, changed nothing) would have been skipped. **2. One record of the last suite result for a tree — shipped for the code review.** The unit lane's report is that record, and it now names the code it ran against: `code_sha` is a digest of the commit and of the content of every change to a file outside `.osf/`. Unlike the tree digest the freshness check uses, it changes when an already-modified file is edited again, and it does not change when Braid writes its own records; `suites-green` and its freshness check are untouched. The code review's prompt says, from it, when the suite already passed on exactly the code it is reading (the command, how long it took, and that the lane's lint is a separate matter) and not to run it before changing something; when the code changed since, it asks for one run. Whenever the review runs the suite it is told to run it once, in one place and not in a `test` subagent as well, under a `timeout` of four times the lane's duration (at least a minute): run 41's review spent about 20 of its 30.8 minutes on a suite the lane ran in 18 seconds, killed twice by the tool's limits. The unit-test step is told where the scenarios stand (how many have a tagged test and which do not, the lane's own arithmetic; 31 of 31 on run 41's worktree, the figure its report recorded). Not recorded: the apply steps' own suite runs. **3. The code review alongside the browser check — not done.** Both edit one worktree, so the review would read a diff the browser check is still changing; overlapping independent steps is #128, whose list has this pair. Keeping the browser check out of reconcile shipped with #127. Checked: unit tests for the condition, the digest (an edit to an already-modified file, a new file, a commit, a write under `.osf/`), the scenario status against the lane's report, and both prompts; the fast and full lanes; the reconcile condition end to end in a local osfd (ticked: skipped with the reason; unticked: runs); the prompts read for a real change; and on production through the deployed code. **To measure in the next scratch run**: reconcile (none), the unit step and the code review against 25–35 minutes between them, and the number of suite runs in the code review's transcript against run 37's 13 backend runs. Change record: `openspec/changes/archive/2026-10-03-verification-without-repeats`.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
cmoriarty/braid#125
No description provided.