A run's tests and fixes are spread over five agent sessions, and none of them judges the result: a hello world took 74 minutes and did not work #139

Open
opened 2026-10-03 17:25:00 -04:00 by cmoriarty · 4 comments
Owner

Baseline

Run 42 (run_01M418YAE03J9HCPRD5V5DXY56), pipeline quick-fix, repo cmoriarty/helloworld, 2026-10-03. Brief: "The smallest web app ever! Build a tiny web app that when you click a button, hello world appears. Just animate it using clever ascii art to make it look visually interesting."

It ended succeeded, merged, in 74.3 min. What was delivered does not work:

  1. Opened as a file it does nothing. index.html loads <script type="module" src="./helloworld.js">, which Chrome and Brave block over file:// (CORS error). The app works only behind node server.js.
  2. Served, the art is unreadable. After the type-in, a shimmer wave shifts every column up or down with wrap-around, tearing the letters. "HELLO WORLD" is not legible in the settled frame.

Both passed every check in the run.

Where the 74 minutes went

Step Time Notes
fix.implement 26 min app written in ~6 min; ~15 min in three browser-subagent checks of its own; found and fixed a first-click crash
agent.test.unit 6 min reread, one regression test, mutation check
agent.test.e2e 13 + 22 min attempt 1 wrote Node HTTP tests; the lane needs Playwright artifacts, so it failed on lane-passed:e2e. Attempt 2 (a different session) reread everything, probed for browsers, installed Playwright, rewrote the suite, edited server.js and helloworld.js again. ~22 of the 35 min were reasoning
agent.test.browser 5 min the same check a fourth time
fix.summarize 1.4 min

About 6 of the 74 minutes were spent writing the product.

What went wrong

  1. Every check measured a proxy that cannot fail. The e2e test, the browser step and all three of fix.implement's subagents asserted "exactly 107 # on screen". The shimmer permutes cells, so the count is invariant even when the letters are torn. It passed seven times. Four screenshots were taken in the browser step and none was judged.
  2. The evidence contradicted the verdict. The second subagent quoted a frame with nearly blank rows and still reported "clearly spelling HELLO WORLD".
  3. The implementer graded itself. The checks that passed were launched by the agent that wrote the code, and later steps reran its oracle.
  4. Development and testing are tangled. The unit, e2e and browser steps edit application code; fix.implement runs its own browser loop. Five sessions each build, test and fix.
  5. The lane contract surfaced only as a failure. Attempt 1 of the e2e step was not told the lane needs playwright-report/ or test-results/; 13 min lost.
  6. Nothing opened the deliverable the way a person would. The judge starts the app with the app tool (a served URL) by construction. No step asked how the user would run "the smallest web app".
  7. .gitignore came too late. script.gitignore.feature runs after the tests, so 35 Playwright report and trace files were committed (the summary admits it).
  8. The brief's intent was never checked. "Visually interesting" was not turned into anything a step could assert, and "the text is legible" was never a criterion.

Proposal (to be explored)

Replace the repeated build-test-fix sessions with a cycle the structure enforces:

  1. Implement: one agent writes code and its tests, running only targeted tests. It does not run browser loops of its own.
  2. Verify, run by the engine: a scripted step runs the lanes and writes the report (osf.pipeline.tests run ... already exists). No model is involved. Agents never grade themselves.
  3. A bounded fix loop: on failure, a fix step gets only the failing report; then verify runs again. The cap and the strict-progress condition are what LoopSpec (types.py) was designed for; the scheduler does not read it today, and the revise loop is hand-unrolled instead.
  4. An independent judge, in a fresh session that sees only the brief, the deliverable and screenshots. It opens the deliverable the way the brief implies and the README documents (for a static page: file:// as well as a server), and it judges the picture, not a count. It reports findings and does not fix; findings route back into the fix loop.
  5. Acceptance in the plan: the fix note or proposal states what the judge must see, in a form a check can assert ("the text HELLO WORLD is readable in the settled frame"). A count of ink is not an accepted oracle.
  6. Session continuity: keep the implementer's session across implement → verify → fix (attempt 2 of the e2e step spent most of its first ~13 min re-learning what attempt 1 knew, and fix.implement started three cold subagent sessions). Use fresh sessions only for the judge. Check how a mediation retry seeds history first: its session id differs from the failed attempt's.
  7. Tell the implementer the lane contract up front, and create .gitignore before the first test run.

Questions for the exploration

  • What did each verification step change across the last runs, and what would an engine-run verify plus a fix loop have missed? (Run 25's unit step picked up the test lens's gaps.)
  • Can the scenario-coverage check replace what agent.test.unit adds today?
  • Does the loop need engine support first (LoopSpec), or can it ship unrolled as the revise loop did?
  • Related: #138 (the lint and coverage decide nothing), #121 (efficiency findings 2 and 3: verification steps redo each other), #125 (reconcile skip), #128 (parallel steps).

Baseline for the change

Re-run the same brief on quick-fix and compare: wall clock, agent sessions, the number of times the suite is run, whether the delivered app opens from file://, and whether the settled art is legible.

## Baseline Run 42 (`run_01M418YAE03J9HCPRD5V5DXY56`), pipeline `quick-fix`, repo `cmoriarty/helloworld`, 2026-10-03. Brief: "The smallest web app ever! Build a tiny web app that when you click a button, hello world appears. Just animate it using clever ascii art to make it look visually interesting." It ended `succeeded`, merged, in 74.3 min. What was delivered does not work: 1. **Opened as a file it does nothing.** `index.html` loads `<script type="module" src="./helloworld.js">`, which Chrome and Brave block over `file://` (CORS error). The app works only behind `node server.js`. 2. **Served, the art is unreadable.** After the type-in, a shimmer wave shifts every column up or down with wrap-around, tearing the letters. "HELLO WORLD" is not legible in the settled frame. Both passed every check in the run. ## Where the 74 minutes went | Step | Time | Notes | |---|---|---| | `fix.implement` | 26 min | app written in ~6 min; ~15 min in three browser-subagent checks of its own; found and fixed a first-click crash | | `agent.test.unit` | 6 min | reread, one regression test, mutation check | | `agent.test.e2e` | 13 + 22 min | attempt 1 wrote Node HTTP tests; the lane needs Playwright artifacts, so it failed on `lane-passed:e2e`. Attempt 2 (a different session) reread everything, probed for browsers, installed Playwright, rewrote the suite, edited `server.js` and `helloworld.js` again. ~22 of the 35 min were reasoning | | `agent.test.browser` | 5 min | the same check a fourth time | | `fix.summarize` | 1.4 min | | About 6 of the 74 minutes were spent writing the product. ## What went wrong 1. **Every check measured a proxy that cannot fail.** The e2e test, the browser step and all three of `fix.implement`'s subagents asserted "exactly 107 `#` on screen". The shimmer permutes cells, so the count is invariant even when the letters are torn. It passed seven times. Four screenshots were taken in the browser step and none was judged. 2. **The evidence contradicted the verdict.** The second subagent quoted a frame with nearly blank rows and still reported "clearly spelling HELLO WORLD". 3. **The implementer graded itself.** The checks that passed were launched by the agent that wrote the code, and later steps reran its oracle. 4. **Development and testing are tangled.** The unit, e2e and browser steps edit application code; `fix.implement` runs its own browser loop. Five sessions each build, test and fix. 5. **The lane contract surfaced only as a failure.** Attempt 1 of the e2e step was not told the lane needs `playwright-report/` or `test-results/`; 13 min lost. 6. **Nothing opened the deliverable the way a person would.** The judge starts the app with the app tool (a served URL) by construction. No step asked how the user would run "the smallest web app". 7. **`.gitignore` came too late.** `script.gitignore.feature` runs after the tests, so 35 Playwright report and trace files were committed (the summary admits it). 8. **The brief's intent was never checked.** "Visually interesting" was not turned into anything a step could assert, and "the text is legible" was never a criterion. ## Proposal (to be explored) Replace the repeated build-test-fix sessions with a cycle the structure enforces: 1. **Implement**: one agent writes code and its tests, running only targeted tests. It does not run browser loops of its own. 2. **Verify, run by the engine**: a scripted step runs the lanes and writes the report (`osf.pipeline.tests run ...` already exists). No model is involved. Agents never grade themselves. 3. **A bounded fix loop**: on failure, a fix step gets only the failing report; then verify runs again. The cap and the strict-progress condition are what `LoopSpec` (`types.py`) was designed for; the scheduler does not read it today, and the revise loop is hand-unrolled instead. 4. **An independent judge**, in a fresh session that sees only the brief, the deliverable and screenshots. It opens the deliverable the way the brief implies and the README documents (for a static page: `file://` as well as a server), and it judges the picture, not a count. It reports findings and does not fix; findings route back into the fix loop. 5. **Acceptance in the plan**: the fix note or proposal states what the judge must see, in a form a check can assert ("the text HELLO WORLD is readable in the settled frame"). A count of ink is not an accepted oracle. 6. **Session continuity**: keep the implementer's session across implement → verify → fix (attempt 2 of the e2e step spent most of its first ~13 min re-learning what attempt 1 knew, and `fix.implement` started three cold subagent sessions). Use fresh sessions only for the judge. Check how a mediation retry seeds history first: its session id differs from the failed attempt's. 7. **Tell the implementer the lane contract up front**, and **create `.gitignore` before the first test run**. ## Questions for the exploration - What did each verification step change across the last runs, and what would an engine-run verify plus a fix loop have missed? (Run 25's unit step picked up the test lens's gaps.) - Can the scenario-coverage check replace what `agent.test.unit` adds today? - Does the loop need engine support first (`LoopSpec`), or can it ship unrolled as the revise loop did? - Related: #138 (the lint and coverage decide nothing), #121 (efficiency findings 2 and 3: verification steps redo each other), #125 (reconcile skip), #128 (parallel steps). ## Baseline for the change Re-run the same brief on `quick-fix` and compare: wall clock, agent sessions, the number of times the suite is run, whether the delivered app opens from `file://`, and whether the settled art is legible.
Author
Owner

First change shipped: fix-loop-with-judge, on quick-fix and git-flow-quick-fix. Deployed as 16f9163 (deploy 101) and archived in 813eebb. The issue stays open for the second change, which carries the loop to the other pipelines.

What a quick fix does now

  • One implementer. fix.implement writes the code and its tests, and runs only the tests it touched.
    • Its brief states the lane contract up front: the unit command in .osf/test/unit.command, and Playwright's output for a browser-test lane. The step cannot finish without the unit command.
    • Its note ends with ## Acceptance: what a person must see, A1, A2 and so on. A count of characters is not a criterion.
    • It does not open a browser.
  • The engine checks. script.verify runs the lanes with no model, streams them, and records the round. It stops a suite that is silent for ten minutes. A red lane is data, not a failed step.
  • A judge looks. When the lanes pass and the fix changed something a browser shows, agent.judge runs in a fresh session.
    • It sees the brief and the acceptance criteria, not the fix's account of itself or its tests.
    • It opens a static page from disk by its file:// URL, and the app served. Playwright MCP refused every file: URL before; it now runs with --allow-unrestricted-file-access.
    • It reads the screenshots itself and reports per criterion with evidence.
    • A failed criterion fails the round, whatever the overall verdict says.
    • It fixes nothing: a judge that edits the code fails its step.
  • The implementer fixes. fix.repair gets only what failed, in a fork of the implementer's opencode session. Each attempt forks afresh, and the copied history stays out of the step's record.
  • Bounded.
    • At most two fix rounds, unrolled in the graph like the review loop.
    • The loop stops early when a round fails the same way twice.
    • script.fix-outcome passes only when the last round passed on exactly the code being committed; otherwise the run stops for a person, with the reason.
  • Also.
    • The .gitignore step runs before the first lane, and a run's clone ignores test-results/ and playwright-report/; the forge's Node template ignores neither.
    • The summary says how the fix was checked and runs no tests.

The exploration's questions

  • What the test steps changed. In the quick fixes, they edited application code in 4 of their 7 runs; the browser checks edited nothing in 4.9 to 50.8 min. In the OpenSpec scratch runs 22 to 40, the unit steps added tests in 6 of 8 and never touched application code, and the four browser checks edited nothing in 165 min.
  • Does the loop need LoopSpec? No: it ships unrolled, as the review loop did.
  • Can scenario coverage replace agent.test.unit? Quick fixes have no scenarios, so this is left to the second change.

Baseline: run 42's brief again, as run 43 (new project cmoriarty/helloworld-139)

Run 42 Run 43
Wall clock 74.3 min 24.2 min
Agent sessions 13 5 (implementer, judge with 2 browser subagents, summary)
Test runs by agents 23 8
Lane runs by the engine 3 1
Opens from disk no yes, no console errors
Settled art rows torn, unreadable letter shapes intact and stable
  • The implementer took 14 min, against 26, with an inline classic script and a test asserting the exact settled frame.
  • The judge took 8 min and passed every criterion. No fix round was needed, so the fork ran end to end only locally, against the pinned opencode 1.18.21: round 1 red, the fix round's first request carrying the implementer's 7 messages, round 2 green.

Two things the judge got only partly right, for the second change's brief:

  • The art. Its shimmer draws some letter pixels as ' and :, so HELLO WORLD reads, but with effort. The judge's "reads unmistakably" came from binarising the frame, which a person does not do.
  • Test runs. It ran npm test, which its brief does not ask for.

Verified

  • Fast lane: 2,559 backend, 552 vitest, 163 browser tests.
  • Full lane: 2,567, 552 and 180.
  • On production, in the container: quick-fix renders with the loop, the MCP has the flag, its self-check passes, and the served console has the new steps.

Next: the second change, for thorough, git-flow, minimalist, git-flow-minimalist and git-flow-hotfix. It covers whose session a fix round forks when apply is fanned out, the delta specs' scenarios as the judge's criteria, what the scenario-test step keeps, and the code review's place in the loop.

First change shipped: `fix-loop-with-judge`, on `quick-fix` and `git-flow-quick-fix`. Deployed as 16f9163 (deploy 101) and archived in 813eebb. The issue stays open for the second change, which carries the loop to the other pipelines. **What a quick fix does now** - **One implementer.** `fix.implement` writes the code and its tests, and runs only the tests it touched. - Its brief states the lane contract up front: the unit command in `.osf/test/unit.command`, and Playwright's output for a browser-test lane. The step cannot finish without the unit command. - Its note ends with `## Acceptance`: what a person must see, A1, A2 and so on. A count of characters is not a criterion. - It does not open a browser. - **The engine checks.** `script.verify` runs the lanes with no model, streams them, and records the round. It stops a suite that is silent for ten minutes. A red lane is data, not a failed step. - **A judge looks.** When the lanes pass and the fix changed something a browser shows, `agent.judge` runs in a fresh session. - It sees the brief and the acceptance criteria, not the fix's account of itself or its tests. - It opens a static page from disk by its `file://` URL, and the app served. Playwright MCP refused every `file:` URL before; it now runs with `--allow-unrestricted-file-access`. - It reads the screenshots itself and reports per criterion with evidence. - A failed criterion fails the round, whatever the overall verdict says. - It fixes nothing: a judge that edits the code fails its step. - **The implementer fixes.** `fix.repair` gets only what failed, in a fork of the implementer's opencode session. Each attempt forks afresh, and the copied history stays out of the step's record. - **Bounded.** - At most two fix rounds, unrolled in the graph like the review loop. - The loop stops early when a round fails the same way twice. - `script.fix-outcome` passes only when the last round passed on exactly the code being committed; otherwise the run stops for a person, with the reason. - **Also.** - The `.gitignore` step runs before the first lane, and a run's clone ignores `test-results/` and `playwright-report/`; the forge's Node template ignores neither. - The summary says how the fix was checked and runs no tests. **The exploration's questions** - **What the test steps changed.** In the quick fixes, they edited application code in 4 of their 7 runs; the browser checks edited nothing in 4.9 to 50.8 min. In the OpenSpec scratch runs 22 to 40, the unit steps added tests in 6 of 8 and never touched application code, and the four browser checks edited nothing in 165 min. - **Does the loop need LoopSpec?** No: it ships unrolled, as the review loop did. - **Can scenario coverage replace agent.test.unit?** Quick fixes have no scenarios, so this is left to the second change. **Baseline: run 42's brief again, as run 43** (new project `cmoriarty/helloworld-139`) | | Run 42 | Run 43 | |---|---|---| | Wall clock | 74.3 min | 24.2 min | | Agent sessions | 13 | 5 (implementer, judge with 2 `browser` subagents, summary) | | Test runs by agents | 23 | 8 | | Lane runs by the engine | 3 | 1 | | Opens from disk | no | yes, no console errors | | Settled art | rows torn, unreadable | letter shapes intact and stable | - The implementer took 14 min, against 26, with an inline classic script and a test asserting the exact settled frame. - The judge took 8 min and passed every criterion. No fix round was needed, so the fork ran end to end only locally, against the pinned opencode 1.18.21: round 1 red, the fix round's first request carrying the implementer's 7 messages, round 2 green. **Two things the judge got only partly right, for the second change's brief:** - **The art.** Its shimmer draws some letter pixels as `'` and `:`, so HELLO WORLD reads, but with effort. The judge's "reads unmistakably" came from binarising the frame, which a person does not do. - **Test runs.** It ran `npm test`, which its brief does not ask for. **Verified** - Fast lane: 2,559 backend, 552 vitest, 163 browser tests. - Full lane: 2,567, 552 and 180. - On production, in the container: quick-fix renders with the loop, the MCP has the flag, its self-check passes, and the served console has the new steps. **Next:** the second change, for `thorough`, `git-flow`, `minimalist`, `git-flow-minimalist` and `git-flow-hotfix`. It covers whose session a fix round forks when apply is fanned out, the delta specs' scenarios as the judge's criteria, what the scenario-test step keeps, and the code review's place in the loop.
Author
Owner

Second change shipped: fix-loop-in-openspec-pipelines. The fix loop now checks every pipeline that writes code: thorough, git-flow, minimalist, git-flow-minimalist and git-flow-hotfix, as well as the two quick fixes. Deployed as 0af9d0f, with the console's VERIFY & FIX band after it (68e51b1, 9743066), and archived in openspec/changes/archive/2026-10-04-fix-loop-in-openspec-pipelines.

The steps are named here as #142 has since renamed them: agent.write-spec-tests (was agent.test.scenarios), script.run-tests (script.verify), agent.fix (fix.repair), script.confirm-checks-passed (script.fix-outcome) and agent.implement (fix.implement).

What an OpenSpec run does now, after the apply steps

  • The apply steps are the implementer. Each one's brief states the lane contract, as agent.implement's does. The first group to find .osf/test/unit.command missing writes it, and an apply step cannot finish without a unit command. They run only the tests they touch, and leave the browser to the judge.
  • agent.write-spec-tests replaces agent.test.unit. It runs only when a scenario of the change has no tagged test, as the lane counts them. It writes those tests and may change nothing but tests. A test that finds a bug stays red, and round 1 sends it to the fix round. An edit outside tests stops the step for a person.
  • The code review comes before the loop, in thorough and git-flow, so round 1 checks what it changed. It runs only the tests its own changes touch.
  • The loop is the quick fixes' loop, defined once in the default graph. The engine runs the lanes, a judge looks when a browser shows the change, and at most two fix rounds follow. A fix round continues the session of the apply step that finished last, may fix any group's code, and leaves the change's documents alone.
  • A change's judge judges its delta specs' scenarios a person can see, under <capability>/<scenario>, the tag the tests carry. It adds only what a person would plainly see is broken.
  • Every judge, quick fixes included, now reads text as drawn (run 43's judge binarised the art) and runs no tests (run 43's ran npm test).
  • The old test steps (agent.test.unit, agent.test.e2e, agent.test.browser and the scripted twins) left the built-in pipelines. They still run in a repository's own pipeline.
  • The console draws the loop under a VERIFY & FIX band between APPLY and the pull request's band, and only the band holding the step the run is at shows where the run is.

Found on the way

  • Locally, end to end, a scenario-test step that also fixed the code was retried by the engine, carrying on from its own worktree. Its condition, asked again, then found every scenario tagged and skipped it, and the edit reached the commit. The step now waits for a person on such an edit.
  • #141: the watchdog judged a step by the session of a subagent that had finished first, and orphaned a working apply step in run 44. Fixed and deployed (25e898d).

Baseline: not measured yet. The comparison with run 41 (315 min, 138 of them after the apply steps) needs a run that reaches the commit, on cmoriarty/scratch-139 (a copy of scratch at run 41's start) with run 41's brief. Three runs stopped before the loop:

  • run 44 at its third apply step, on #141;
  • run 45 before the apply steps, stopped to deploy the VERIFY & FIX band;
  • run 46 at its third apply step, when trogdor was taken for vLLM testing.

The apply steps that finished (groups 1 and 2 of runs 44 and 46) passed the new unit-command check on their first attempt. That is all the real-model evidence so far.

Verified

  • Fast lane: 2,585 backend, 554 vitest, 166 browser tests. Full lane: 2,593, 554 and 183.
  • Locally, end to end, with the pinned opencode 1.18.21 and the fake model:
    • an apply step without the unit command was failed and retried;
    • round 1 was red on the farewell bug the scenario test exposed;
    • the fix round's first request carried group 2's apply session, not group 1's;
    • round 2 was green, and the commit followed.
  • On production, in the container: thorough, minimalist and git-flow-hotfix have the scenario tests and the loop, and no old test step.

This issue stays open until the baseline run is in.

Second change shipped: `fix-loop-in-openspec-pipelines`. The fix loop now checks every pipeline that writes code: `thorough`, `git-flow`, `minimalist`, `git-flow-minimalist` and `git-flow-hotfix`, as well as the two quick fixes. Deployed as 0af9d0f, with the console's VERIFY & FIX band after it (68e51b1, 9743066), and archived in `openspec/changes/archive/2026-10-04-fix-loop-in-openspec-pipelines`. The steps are named here as #142 has since renamed them: `agent.write-spec-tests` (was `agent.test.scenarios`), `script.run-tests` (`script.verify`), `agent.fix` (`fix.repair`), `script.confirm-checks-passed` (`script.fix-outcome`) and `agent.implement` (`fix.implement`). **What an OpenSpec run does now, after the apply steps** - **The apply steps are the implementer.** Each one's brief states the lane contract, as `agent.implement`'s does. The first group to find `.osf/test/unit.command` missing writes it, and an apply step cannot finish without a unit command. They run only the tests they touch, and leave the browser to the judge. - **`agent.write-spec-tests` replaces `agent.test.unit`.** It runs only when a scenario of the change has no tagged test, as the lane counts them. It writes those tests and may change nothing but tests. A test that finds a bug stays red, and round 1 sends it to the fix round. An edit outside tests stops the step for a person. - **The code review comes before the loop**, in `thorough` and `git-flow`, so round 1 checks what it changed. It runs only the tests its own changes touch. - **The loop is the quick fixes' loop**, defined once in the default graph. The engine runs the lanes, a judge looks when a browser shows the change, and at most two fix rounds follow. A fix round continues the session of the apply step that finished last, may fix any group's code, and leaves the change's documents alone. - **A change's judge judges its delta specs' scenarios** a person can see, under `<capability>/<scenario>`, the tag the tests carry. It adds only what a person would plainly see is broken. - **Every judge**, quick fixes included, now reads text as drawn (run 43's judge binarised the art) and runs no tests (run 43's ran `npm test`). - **The old test steps** (`agent.test.unit`, `agent.test.e2e`, `agent.test.browser` and the scripted twins) left the built-in pipelines. They still run in a repository's own pipeline. - **The console** draws the loop under a VERIFY & FIX band between APPLY and the pull request's band, and only the band holding the step the run is at shows where the run is. **Found on the way** - Locally, end to end, a scenario-test step that also fixed the code was retried by the engine, carrying on from its own worktree. Its condition, asked again, then found every scenario tagged and skipped it, and the edit reached the commit. The step now waits for a person on such an edit. - #141: the watchdog judged a step by the session of a subagent that had finished first, and orphaned a working apply step in run 44. Fixed and deployed (25e898d). **Baseline: not measured yet.** The comparison with run 41 (315 min, 138 of them after the apply steps) needs a run that reaches the commit, on `cmoriarty/scratch-139` (a copy of scratch at run 41's start) with run 41's brief. Three runs stopped before the loop: - run 44 at its third apply step, on #141; - run 45 before the apply steps, stopped to deploy the VERIFY & FIX band; - run 46 at its third apply step, when trogdor was taken for vLLM testing. The apply steps that finished (groups 1 and 2 of runs 44 and 46) passed the new unit-command check on their first attempt. That is all the real-model evidence so far. **Verified** - Fast lane: 2,585 backend, 554 vitest, 166 browser tests. Full lane: 2,593, 554 and 183. - Locally, end to end, with the pinned opencode 1.18.21 and the fake model: - an apply step without the unit command was failed and retried; - round 1 was red on the farewell bug the scenario test exposed; - the fix round's first request carried group 2's apply session, not group 1's; - round 2 was green, and the commit followed. - On production, in the container: `thorough`, `minimalist` and `git-flow-hotfix` have the scenario tests and the loop, and no old test step. This issue stays open until the baseline run is in.
Author
Owner

The baseline is under way as run 47 (run_01M46705KJ27XX2VXKFJVNCM16), started 2026-10-05 14:20Z on eb1161c. It is the same setup as runs 44 to 46: git-flow, cmoriarty/scratch-139 at run 41's start (88be018), and run 41's brief.

trogdor changed since run 41. It moved from vLLM 0.28.0 to 0.31.0 + PCIe, with the same model (qwen3.8-27b):

0.28.0 (run 41) 0.31.0 (run 47)
Decode, short context 58.1 t/s 66.8 t/s
Decode at 20k 55.6 t/s 63.4 t/s
4-stream aggregate 156.6 t/s 170.6 t/s
Cold 18k prefill 10.5 s 10.6 s

The wall clock would therefore drop by up to about 15% with no change of ours. The comparison leads with measures that don't depend on decode speed: model turns, generated tokens (output plus reasoning, from each turn's token counts), sessions, attempts and test runs. Run 41, measured the same way:

Run 41 Turns Generated tokens
Whole run 633 743,218
From the end of the apply steps to the feature commit 326 331,910
of which agent.test.browser 204 224,399
of which agent.code-review 53 56,814
of which agent.test.unit 50 38,956

At 14:37Z run 47's change had validated on the first try, and its four review lenses were running. The results will follow when it finishes.

**The baseline is under way as run 47** (`run_01M46705KJ27XX2VXKFJVNCM16`), started 2026-10-05 14:20Z on eb1161c. It is the same setup as runs 44 to 46: `git-flow`, `cmoriarty/scratch-139` at run 41's start (88be018), and run 41's brief. **trogdor changed since run 41.** It moved from vLLM 0.28.0 to 0.31.0 + PCIe, with the same model (qwen3.8-27b): | | 0.28.0 (run 41) | 0.31.0 (run 47) | |---|---|---| | Decode, short context | 58.1 t/s | 66.8 t/s | | Decode at 20k | 55.6 t/s | 63.4 t/s | | 4-stream aggregate | 156.6 t/s | 170.6 t/s | | Cold 18k prefill | 10.5 s | 10.6 s | The wall clock would therefore drop by up to about 15% with no change of ours. The comparison leads with measures that don't depend on decode speed: model turns, generated tokens (output plus reasoning, from each turn's token counts), sessions, attempts and test runs. Run 41, measured the same way: | Run 41 | Turns | Generated tokens | |---|---|---| | Whole run | 633 | 743,218 | | From the end of the apply steps to the feature commit | 326 | 331,910 | | of which `agent.test.browser` | 204 | 224,399 | | of which `agent.code-review` | 53 | 56,814 | | of which `agent.test.unit` | 50 | 38,956 | At 14:37Z run 47's change had validated on the first try, and its four review lenses were running. The results will follow when it finishes.
Author
Owner

Run 48: the baseline again, after #144 and #145. git-flow, with run 41's brief and starting commit, on vLLM 0.31 (decoding about 15% faster than run 41's 0.28). It did not finish: it failed in round 2's judge (#150). Up to there:

run 41 run 48
wall clock 315 min, merged 322.5 min, failed in the judge (27.5 min of it held on an empty question, #147)
model turns 633 544
generated tokens 743,218 660,994
after the apply steps: turns / tokens 326 / 331,910, to the commit 362 / 444,875, to the failure
agent steps / attempts 19 / 20 18 / 19
subagents 17 6
agent test runs 42 16

What the loop changed. Fewer sessions and test runs:

  • 6 subagents against 17, and 16 agent test runs against 42;
  • the lanes ran as scripts: round 2's took 20 s;
  • the fix round was told exactly what failed.

What it cost. The judge dominates. Round 2's judge took 190 minutes over two attempts, with 303 turns and 371,260 generated tokens: 83% of everything generated after apply.

  • Its first attempt played the game for its whole two hours and wrote no report.
  • Its second wrote a valid report (#145's fix held), but a log left by the first failed its "code unchanged" check.

What the run found:

  • #147: an empty question held the run for 27½ minutes;
  • #148: the unit command's python was Braid's own interpreter, so round 1 checked nothing, and the fix that corrected it failed its own check;
  • #150: the judge's cost, its lost first attempt and its stray log.

The simulated runs (#149, pushed tonight) cover #145, #147 and #148, so those fixes are checked in seconds before another baseline. #150 will get one.

#139 stays open: the loop's shape works, but a baseline has to finish before the comparison means anything, and #150 comes first.

**Run 48: the baseline again, after #144 and #145.** git-flow, with run 41's brief and starting commit, on vLLM 0.31 (decoding about 15% faster than run 41's 0.28). It did not finish: it failed in round 2's judge (#150). Up to there: | | run 41 | run 48 | | --- | --- | --- | | wall clock | 315 min, merged | 322.5 min, failed in the judge (27.5 min of it held on an empty question, #147) | | model turns | 633 | 544 | | generated tokens | 743,218 | 660,994 | | after the apply steps: turns / tokens | 326 / 331,910, to the commit | 362 / 444,875, to the failure | | agent steps / attempts | 19 / 20 | 18 / 19 | | subagents | 17 | 6 | | agent test runs | 42 | 16 | **What the loop changed.** Fewer sessions and test runs: - 6 subagents against 17, and 16 agent test runs against 42; - the lanes ran as scripts: round 2's took 20 s; - the fix round was told exactly what failed. **What it cost.** The judge dominates. Round 2's judge took 190 minutes over two attempts, with 303 turns and 371,260 generated tokens: 83% of everything generated after apply. - Its first attempt played the game for its whole two hours and wrote no report. - Its second wrote a valid report (#145's fix held), but a log left by the first failed its "code unchanged" check. **What the run found:** - #147: an empty question held the run for 27½ minutes; - #148: the unit command's `python` was Braid's own interpreter, so round 1 checked nothing, and the fix that corrected it failed its own check; - #150: the judge's cost, its lost first attempt and its stray log. The simulated runs (#149, pushed tonight) cover #145, #147 and #148, so those fixes are checked in seconds before another baseline. #150 will get one. #139 stays open: the loop's shape works, but a baseline has to finish before the comparison means anything, and #150 comes first.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
cmoriarty/braid#139
No description provided.