A run's tests and fixes are spread over five agent sessions, and none of them judges the result: a hello world took 74 minutes and did not work #139
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Baseline
Run 42 (
run_01M418YAE03J9HCPRD5V5DXY56), pipelinequick-fix, repocmoriarty/helloworld, 2026-10-03. Brief: "The smallest web app ever! Build a tiny web app that when you click a button, hello world appears. Just animate it using clever ascii art to make it look visually interesting."It ended
succeeded, merged, in 74.3 min. What was delivered does not work:index.htmlloads<script type="module" src="./helloworld.js">, which Chrome and Brave block overfile://(CORS error). The app works only behindnode server.js.Both passed every check in the run.
Where the 74 minutes went
fix.implementagent.test.unitagent.test.e2elane-passed:e2e. Attempt 2 (a different session) reread everything, probed for browsers, installed Playwright, rewrote the suite, editedserver.jsandhelloworld.jsagain. ~22 of the 35 min were reasoningagent.test.browserfix.summarizeAbout 6 of the 74 minutes were spent writing the product.
What went wrong
fix.implement's subagents asserted "exactly 107#on screen". The shimmer permutes cells, so the count is invariant even when the letters are torn. It passed seven times. Four screenshots were taken in the browser step and none was judged.fix.implementruns its own browser loop. Five sessions each build, test and fix.playwright-report/ortest-results/; 13 min lost..gitignorecame too late.script.gitignore.featureruns after the tests, so 35 Playwright report and trace files were committed (the summary admits it).Proposal (to be explored)
Replace the repeated build-test-fix sessions with a cycle the structure enforces:
osf.pipeline.tests run ...already exists). No model is involved. Agents never grade themselves.LoopSpec(types.py) was designed for; the scheduler does not read it today, and the revise loop is hand-unrolled instead.file://as well as a server), and it judges the picture, not a count. It reports findings and does not fix; findings route back into the fix loop.fix.implementstarted three cold subagent sessions). Use fresh sessions only for the judge. Check how a mediation retry seeds history first: its session id differs from the failed attempt's..gitignorebefore the first test run.Questions for the exploration
agent.test.unitadds today?LoopSpec), or can it ship unrolled as the revise loop did?Baseline for the change
Re-run the same brief on
quick-fixand compare: wall clock, agent sessions, the number of times the suite is run, whether the delivered app opens fromfile://, and whether the settled art is legible.First change shipped:
fix-loop-with-judge, onquick-fixandgit-flow-quick-fix. Deployed as16f9163(deploy 101) and archived in813eebb. The issue stays open for the second change, which carries the loop to the other pipelines.What a quick fix does now
fix.implementwrites the code and its tests, and runs only the tests it touched..osf/test/unit.command, and Playwright's output for a browser-test lane. The step cannot finish without the unit command.## Acceptance: what a person must see, A1, A2 and so on. A count of characters is not a criterion.script.verifyruns the lanes with no model, streams them, and records the round. It stops a suite that is silent for ten minutes. A red lane is data, not a failed step.agent.judgeruns in a fresh session.file://URL, and the app served. Playwright MCP refused everyfile:URL before; it now runs with--allow-unrestricted-file-access.fix.repairgets only what failed, in a fork of the implementer's opencode session. Each attempt forks afresh, and the copied history stays out of the step's record.script.fix-outcomepasses only when the last round passed on exactly the code being committed; otherwise the run stops for a person, with the reason..gitignorestep runs before the first lane, and a run's clone ignorestest-results/andplaywright-report/; the forge's Node template ignores neither.The exploration's questions
Baseline: run 42's brief again, as run 43 (new project
cmoriarty/helloworld-139)browsersubagents, summary)Two things the judge got only partly right, for the second change's brief:
'and:, so HELLO WORLD reads, but with effort. The judge's "reads unmistakably" came from binarising the frame, which a person does not do.npm test, which its brief does not ask for.Verified
Next: the second change, for
thorough,git-flow,minimalist,git-flow-minimalistandgit-flow-hotfix. It covers whose session a fix round forks when apply is fanned out, the delta specs' scenarios as the judge's criteria, what the scenario-test step keeps, and the code review's place in the loop.Second change shipped:
fix-loop-in-openspec-pipelines. The fix loop now checks every pipeline that writes code:thorough,git-flow,minimalist,git-flow-minimalistandgit-flow-hotfix, as well as the two quick fixes. Deployed as0af9d0f, with the console's VERIFY & FIX band after it (68e51b1,9743066), and archived inopenspec/changes/archive/2026-10-04-fix-loop-in-openspec-pipelines.The steps are named here as #142 has since renamed them:
agent.write-spec-tests(wasagent.test.scenarios),script.run-tests(script.verify),agent.fix(fix.repair),script.confirm-checks-passed(script.fix-outcome) andagent.implement(fix.implement).What an OpenSpec run does now, after the apply steps
agent.implement's does. The first group to find.osf/test/unit.commandmissing writes it, and an apply step cannot finish without a unit command. They run only the tests they touch, and leave the browser to the judge.agent.write-spec-testsreplacesagent.test.unit. It runs only when a scenario of the change has no tagged test, as the lane counts them. It writes those tests and may change nothing but tests. A test that finds a bug stays red, and round 1 sends it to the fix round. An edit outside tests stops the step for a person.thoroughandgit-flow, so round 1 checks what it changed. It runs only the tests its own changes touch.<capability>/<scenario>, the tag the tests carry. It adds only what a person would plainly see is broken.npm test).agent.test.unit,agent.test.e2e,agent.test.browserand the scripted twins) left the built-in pipelines. They still run in a repository's own pipeline.Found on the way
25e898d).Baseline: not measured yet. The comparison with run 41 (315 min, 138 of them after the apply steps) needs a run that reaches the commit, on
cmoriarty/scratch-139(a copy of scratch at run 41's start) with run 41's brief. Three runs stopped before the loop:The apply steps that finished (groups 1 and 2 of runs 44 and 46) passed the new unit-command check on their first attempt. That is all the real-model evidence so far.
Verified
thorough,minimalistandgit-flow-hotfixhave the scenario tests and the loop, and no old test step.This issue stays open until the baseline run is in.
The baseline is under way as run 47 (
run_01M46705KJ27XX2VXKFJVNCM16), started 2026-10-05 14:20Z oneb1161c. It is the same setup as runs 44 to 46:git-flow,cmoriarty/scratch-139at run 41's start (88be018), and run 41's brief.trogdor changed since run 41. It moved from vLLM 0.28.0 to 0.31.0 + PCIe, with the same model (qwen3.8-27b):
The wall clock would therefore drop by up to about 15% with no change of ours. The comparison leads with measures that don't depend on decode speed: model turns, generated tokens (output plus reasoning, from each turn's token counts), sessions, attempts and test runs. Run 41, measured the same way:
agent.test.browseragent.code-reviewagent.test.unitAt 14:37Z run 47's change had validated on the first try, and its four review lenses were running. The results will follow when it finishes.
Run 48: the baseline again, after #144 and #145. git-flow, with run 41's brief and starting commit, on vLLM 0.31 (decoding about 15% faster than run 41's 0.28). It did not finish: it failed in round 2's judge (#150). Up to there:
What the loop changed. Fewer sessions and test runs:
What it cost. The judge dominates. Round 2's judge took 190 minutes over two attempts, with 303 turns and 371,260 generated tokens: 83% of everything generated after apply.
What the run found:
pythonwas Braid's own interpreter, so round 1 checked nothing, and the fix that corrected it failed its own check;The simulated runs (#149, pushed tonight) cover #145, #147 and #148, so those fixes are checked in seconds before another baseline. #150 will get one.
#139 stays open: the loop's shape works, but a baseline has to finish before the comparison means anything, and #150 comes first.