A pipeline can only be tested by a real run: simulate whole runs in minutes, with a scripted model, locally and in the deploy's image #149

Closed
opened 2026-10-05 19:31:30 -04:00 by cmoriarty · 1 comment
Owner

The problem

The only way to see a change to a pipeline work is a real run on the GPU host: run 48 took over two hours to reach its fix loop. Tonight's run found these, each one a whole run late:

  • #145: the judge's report failed its check in every OpenSpec run, and the loop then trusted it.
  • #144: the first apply step wrote no unit command.
  • #147: an empty question from an agent held the run for 27½ minutes.
  • #148: a unit command saying python ran on Braid's own interpreter. The fix round that corrected it failed its own check, and the console called both round 1 steps "not needed".

None of these needs a real model to show. Each is Braid's own logic, prompts, environment or console meeting a particular agent behaviour, and a scripted model can play that behaviour.

What exists

  • osf-fakeinfer is an OpenAI-compatible fake model with scripted turns, and tools/demo.sh starts a local osfd, opencode and fakeinfer.
  • On 10-04, for #139, a throwaway script drove an OpenSpec-shaped run through apply, the scenario tests and the fix loop against them: 20 steps in 22 seconds. It was never kept.
  • fakeinfer serves its turns in one global sequence. That works for one session at a time, not for a pipeline whose reviews and apply groups run side by side.

What is wanted

  • Simulated runs of the shipped pipelines, not copies of them, on a real osfd and a real opencode, with fakeinfer as the model.
  • A scenario says what the model does at each step and attempt (files written, commands run, questions asked, replies), what the target repository is, and what the run must end with: step states and the reasons recorded, events, files, and what the model was sent.
  • One scenario per bug above, so each fix lands with the run that shows it fixed and keeps it fixed.
  • A test lane that runs them locally in minutes.
  • The deploy's self-check runs them in the new image before it is tagged latest, which is where #148's interpreter problem lives.
**The problem** The only way to see a change to a pipeline work is a real run on the GPU host: run 48 took over two hours to reach its fix loop. Tonight's run found these, each one a whole run late: - #145: the judge's report failed its check in every OpenSpec run, and the loop then trusted it. - #144: the first apply step wrote no unit command. - #147: an empty question from an agent held the run for 27½ minutes. - #148: a unit command saying `python` ran on Braid's own interpreter. The fix round that corrected it failed its own check, and the console called both round 1 steps "not needed". None of these needs a real model to show. Each is Braid's own logic, prompts, environment or console meeting a particular agent behaviour, and a scripted model can play that behaviour. **What exists** - `osf-fakeinfer` is an OpenAI-compatible fake model with scripted turns, and `tools/demo.sh` starts a local osfd, opencode and fakeinfer. - On 10-04, for #139, a throwaway script drove an OpenSpec-shaped run through apply, the scenario tests and the fix loop against them: 20 steps in 22 seconds. It was never kept. - fakeinfer serves its turns in one global sequence. That works for one session at a time, not for a pipeline whose reviews and apply groups run side by side. **What is wanted** - Simulated runs of the shipped pipelines, not copies of them, on a real osfd and a real opencode, with fakeinfer as the model. - A scenario says what the model does at each step and attempt (files written, commands run, questions asked, replies), what the target repository is, and what the run must end with: step states and the reasons recorded, events, files, and what the model was sent. - One scenario per bug above, so each fix lands with the run that shows it fixed and keeps it fixed. - A test lane that runs them locally in minutes. - The deploy's self-check runs them in the new image before it is tagged `latest`, which is where #148's interpreter problem lives.
Author
Owner

Shipped, deployed and archived.

What changed (39f32a3; archived in 6883e8c)

  • python -m osf.sim runs a shipped pipeline, unchanged, through a real osfd and the pinned opencode. A scripted model and a fake forge stand in for the GPU host and Forgejo, and each scenario gets its own stack on its own ports.

  • The fake model knows whose each request is. opencode names its session on every request (x-session-id), and osfd records which attempt of which step opened that session, so steps side by side each follow their own script. A request it cannot place fails the scenario; it never answers with a default.

  • The fake forge serves git over HTTP and the REST calls Braid makes. Any other call fails the scenario, naming the endpoint.

  • Nine scenarios in tests/sim/:

    • git-flow and quick-fix, from brief to merge;
    • a red lane, then a fix;
    • #145: the judge's long criterion ids, and a report that fails its check, corrected on retry;
    • #144: an apply step that forgets the unit command.

    These pass. Three more are known failures, written for the behaviour their fix will bring: #147's empty question, and #148's python lane and its lane-command-only fix. They're reported but don't fail the lane; when a fix makes one pass, the lane fails until its marker is removed.

  • ./tools/test.sh sim runs them; the full lane runs them last. On a laptop the pinned opencode is installed into .cache/ once.

  • braid-selfcheck runs them in every newly built image, before latest is tagged.

Verified

  • Laptop: nine scenarios in about 70 s, four at a time. git-flow takes about 40 s from brief to merged pull request (54 steps), and quick-fix about 15 s.
  • With #145's 20-character id limit put back, judge-long-ids fails in 16 s: "agent.judge: failed … $.criteria[0].id: longer than 20". With the brief wording from before #144, apply-forgets-unit-command fails on the brief.
  • In the image (a local build, then this deploy): the same nine results, and /tmp left as found. A deliberately broken scenario fails the self-check, naming the scenario and the step.
  • Fast lane 2,662 backend tests; full lane 2,670 backend, 567 console unit and 183 browser, plus the simulated runs.
  • The deploy (Actions run 113): its "Self-check the new image" step ran the scenarios inside the new image in 4m05s and ended braid-selfcheck: ok. Production served 39f32a3 about 4m45s after the push.

One design change from the proposal: routing by the session instead of by the brief's text, which removed the 19 per-step matching rules. The proposal, spec, design and tasks were updated before the code was written.

Shipped, deployed and archived. **What changed** (39f32a3; archived in 6883e8c) - `python -m osf.sim` runs a shipped pipeline, unchanged, through a real osfd and the pinned opencode. A scripted model and a fake forge stand in for the GPU host and Forgejo, and each scenario gets its own stack on its own ports. - The fake model knows whose each request is. opencode names its session on every request (`x-session-id`), and osfd records which attempt of which step opened that session, so steps side by side each follow their own script. A request it cannot place fails the scenario; it never answers with a default. - The fake forge serves git over HTTP and the REST calls Braid makes. Any other call fails the scenario, naming the endpoint. - Nine scenarios in `tests/sim/`: - git-flow and quick-fix, from brief to merge; - a red lane, then a fix; - #145: the judge's long criterion ids, and a report that fails its check, corrected on retry; - #144: an apply step that forgets the unit command. These pass. Three more are **known failures**, written for the behaviour their fix will bring: #147's empty question, and #148's `python` lane and its lane-command-only fix. They're reported but don't fail the lane; when a fix makes one pass, the lane fails until its marker is removed. - `./tools/test.sh sim` runs them; the full lane runs them last. On a laptop the pinned opencode is installed into `.cache/` once. - `braid-selfcheck` runs them in every newly built image, before `latest` is tagged. **Verified** - Laptop: nine scenarios in about 70 s, four at a time. git-flow takes about 40 s from brief to merged pull request (54 steps), and quick-fix about 15 s. - With #145's 20-character id limit put back, `judge-long-ids` fails in 16 s: "agent.judge: failed … $.criteria[0].id: longer than 20". With the brief wording from before #144, `apply-forgets-unit-command` fails on the brief. - In the image (a local build, then this deploy): the same nine results, and `/tmp` left as found. A deliberately broken scenario fails the self-check, naming the scenario and the step. - Fast lane 2,662 backend tests; full lane 2,670 backend, 567 console unit and 183 browser, plus the simulated runs. - The deploy (Actions run 113): its "Self-check the new image" step ran the scenarios inside the new image in 4m05s and ended `braid-selfcheck: ok`. Production served 39f32a3 about 4m45s after the push. **One design change from the proposal:** routing by the session instead of by the brief's text, which removed the 19 per-step matching rules. The proposal, spec, design and tasks were updated before the code was written.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
cmoriarty/braid#149
No description provided.