A pipeline can only be tested by a real run: simulate whole runs in minutes, with a scripted model, locally and in the deploy's image #149
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
The problem
The only way to see a change to a pipeline work is a real run on the GPU host: run 48 took over two hours to reach its fix loop. Tonight's run found these, each one a whole run late:
pythonran on Braid's own interpreter. The fix round that corrected it failed its own check, and the console called both round 1 steps "not needed".None of these needs a real model to show. Each is Braid's own logic, prompts, environment or console meeting a particular agent behaviour, and a scripted model can play that behaviour.
What exists
osf-fakeinferis an OpenAI-compatible fake model with scripted turns, andtools/demo.shstarts a local osfd, opencode and fakeinfer.What is wanted
latest, which is where #148's interpreter problem lives.Shipped, deployed and archived.
What changed (
39f32a3; archived in6883e8c)python -m osf.simruns a shipped pipeline, unchanged, through a real osfd and the pinned opencode. A scripted model and a fake forge stand in for the GPU host and Forgejo, and each scenario gets its own stack on its own ports.The fake model knows whose each request is. opencode names its session on every request (
x-session-id), and osfd records which attempt of which step opened that session, so steps side by side each follow their own script. A request it cannot place fails the scenario; it never answers with a default.The fake forge serves git over HTTP and the REST calls Braid makes. Any other call fails the scenario, naming the endpoint.
Nine scenarios in
tests/sim/:These pass. Three more are known failures, written for the behaviour their fix will bring: #147's empty question, and #148's
pythonlane and its lane-command-only fix. They're reported but don't fail the lane; when a fix makes one pass, the lane fails until its marker is removed../tools/test.sh simruns them; the full lane runs them last. On a laptop the pinned opencode is installed into.cache/once.braid-selfcheckruns them in every newly built image, beforelatestis tagged.Verified
judge-long-idsfails in 16 s: "agent.judge: failed … $.criteria[0].id: longer than 20". With the brief wording from before #144,apply-forgets-unit-commandfails on the brief./tmpleft as found. A deliberately broken scenario fails the self-check, naming the scenario and the step.braid-selfcheck: ok. Production served39f32a3about 4m45s after the push.One design change from the proposal: routing by the session instead of by the brief's text, which removed the 19 per-step matching rules. The proposal, spec, design and tasks were updated before the code was written.
pythonruns on Braid's own interpreter, and the fix round that corrects it fails its check #148