A unit command that says python runs on Braid's own interpreter, and the fix round that corrects it fails its check #148

Closed
opened 2026-10-05 18:59:14 -04:00 by cmoriarty · 3 comments
Owner

What happened in run 48

  • openspec.apply[1] wrote .osf/test/unit.command the way the project's CI runs its tests: python -m pytest backend/tests && npm --prefix frontend test. Every agent in the run ran its own tests with backend/.venv/bin/python.
  • Round 1's unit lane resolved python to Braid's own interpreter, /opt/openspecflow/app/.venv/bin/python, which has no pytest. It failed in 0.1 s (No module named pytest). No test ran, and the judge was skipped as not due.
  • agent.fix read that, rewrote the command to backend/.venv/bin/python -m pytest backend/tests && npm --prefix frontend test, checked that it exits 0, and wrote its note. Its completion check then failed: "the code has not changed since round 1's check, so nothing was fixed".
  • The retry re-read the step's condition. Round 1's note was on disk, so the fix was no longer due, and the step ended "not needed" with a predicate failure.
  • Round 2's lanes ran the corrected command and passed: 119 backend and 82 frontend tests.

The run recovered, but round 1 checked nothing, and the step that fixed it shows as not needed.

Why

  1. The image puts Braid's virtualenv first on PATH (Dockerfile: ENV PATH=/opt/openspecflow/app/.venv/bin:${PATH}), and PATH passes through to every child: scripted steps, the lanes and the agents' shells. So a project's bare python, pip or pytest is Braid's.
  2. The brief asks for the project's own command for its unit tests. The project's CI says python -m pytest, which is right on CI after setup-python and pip install, and wrong in the lane.
  3. since-verify:changed compares the code without .osf/. A fix whose whole correction is the lane command cannot pass it, though the lane command is part of what the loop checks.
  4. A fix round whose check failed is skipped on retry once its note exists. Through a retry, since-verify:changed decides nothing, as the judge's check did not before #145.

How often: production's log has 15 unit commands written by agents since 09-15. Three are a bare python -m pytest …: one on 09-29, and two on 10-05, including this run.

Proposal

  • The lanes run a project's command with Braid's own virtualenv off PATH, and with the project's environments first: the virtualenvs script.dependencies made (backend/.venv/bin here) and node_modules/.bin. Then python -m pytest means the project's Python, as it does on CI. Braid's own scripted steps keep their interpreter.
  • A fix round's "changed" check counts .osf/test/*.command along with the code.
  • The loop's state takes a fix round's note as done only when its check would hold (something changed since the round's lanes ran), so a failed fix round is retried rather than skipped.
**What happened in run 48** - `openspec.apply[1]` wrote `.osf/test/unit.command` the way the project's CI runs its tests: `python -m pytest backend/tests && npm --prefix frontend test`. Every agent in the run ran its own tests with `backend/.venv/bin/python`. - Round 1's unit lane resolved `python` to Braid's own interpreter, `/opt/openspecflow/app/.venv/bin/python`, which has no pytest. It failed in 0.1 s (`No module named pytest`). No test ran, and the judge was skipped as not due. - `agent.fix` read that, rewrote the command to `backend/.venv/bin/python -m pytest backend/tests && npm --prefix frontend test`, checked that it exits 0, and wrote its note. Its completion check then failed: "the code has not changed since round 1's check, so nothing was fixed". - The retry re-read the step's condition. Round 1's note was on disk, so the fix was no longer due, and the step ended "not needed" with a predicate failure. - Round 2's lanes ran the corrected command and passed: 119 backend and 82 frontend tests. The run recovered, but round 1 checked nothing, and the step that fixed it shows as not needed. **Why** 1. The image puts Braid's virtualenv first on `PATH` (`Dockerfile`: `ENV PATH=/opt/openspecflow/app/.venv/bin:${PATH}`), and `PATH` passes through to every child: scripted steps, the lanes and the agents' shells. So a project's bare `python`, `pip` or `pytest` is Braid's. 2. The brief asks for the project's own command for its unit tests. The project's CI says `python -m pytest`, which is right on CI after `setup-python` and `pip install`, and wrong in the lane. 3. `since-verify:changed` compares the code without `.osf/`. A fix whose whole correction is the lane command cannot pass it, though the lane command is part of what the loop checks. 4. A fix round whose check failed is skipped on retry once its note exists. Through a retry, `since-verify:changed` decides nothing, as the judge's check did not before #145. **How often:** production's log has 15 unit commands written by agents since 09-15. Three are a bare `python -m pytest …`: one on 09-29, and two on 10-05, including this run. **Proposal** - The lanes run a project's command with Braid's own virtualenv off `PATH`, and with the project's environments first: the virtualenvs `script.dependencies` made (`backend/.venv/bin` here) and `node_modules/.bin`. Then `python -m pytest` means the project's Python, as it does on CI. Braid's own scripted steps keep their interpreter. - A fix round's "changed" check counts `.osf/test/*.command` along with the code. - The loop's state takes a fix round's note as done only when its check would hold (something changed since the round's lanes ran), so a failed fix round is retried rather than skipped.
Author
Owner

The console, too. Two rows of run 48's VERIFY & FIX band read wrong:

  • Round 1's agent.judge was skipped as designed: with the lanes red, the fix goes first, and the judge looks once they pass. Braid recorded why: "the fix loop's next step is fix in round 1: the unit lane failed: the suite exited 1". The spine draws it struck through as "not needed".
  • Round 1's agent.fix ran for 2m 52s and made the fix that let round 2 pass, then was skipped on retry. The spine draws it struck through as "not needed", and its step view says "This step was not needed on this run, so it never ran."

Both come from one rule. The spine (Spine.tsx) and the step view (Transcript.tsx) word every skip_kind: not_needed the same way, whatever reason the skip recorded.

Proposal:

  • A skipped step of the fix loop shows the reason it recorded, in a few words (the judge: "lanes failed"), with the whole reason in the hover and the step view.
  • A step that ran is never struck through or described as never having run. It shows its time and what happened: here, that its check failed and the loop moved on to round 2.
**The console, too.** Two rows of run 48's VERIFY & FIX band read wrong: - Round 1's `agent.judge` was skipped as designed: with the lanes red, the fix goes first, and the judge looks once they pass. Braid recorded why: "the fix loop's next step is fix in round 1: the unit lane failed: the suite exited 1". The spine draws it struck through as "not needed". - Round 1's `agent.fix` ran for 2m 52s and made the fix that let round 2 pass, then was skipped on retry. The spine draws it struck through as "not needed", and its step view says "This step was not needed on this run, so it never ran." Both come from one rule. The spine (`Spine.tsx`) and the step view (`Transcript.tsx`) word every `skip_kind: not_needed` the same way, whatever reason the skip recorded. Proposal: - A skipped step of the fix loop shows the reason it recorded, in a few words (the judge: "lanes failed"), with the whole reason in the hover and the step view. - A step that ran is never struck through or described as never having run. It shows its time and what happened: here, that its check failed and the loop moved on to round 2.
Author
Owner

Two simulated runs now show this bug (#149), both known failures of #148:

  • tests/sim/python-on-path.yaml: the unit command says python, and the tests need a module only the project's environment has. Round 1's unit lane should pass.
  • tests/sim/fix-only-lane-command.yaml: the fix round's whole correction is the unit command. It should succeed, not be skipped.

The fix removes both markers; the sim lane fails until it does. Run them with python -m osf.sim python-on-path fix-only-lane-command.

For the console part, python -m osf.sim --keep red-lane-then-fix leaves a run up with round 1's judge skipped for a red lane, to check its row against.

Two simulated runs now show this bug (#149), both known failures of #148: - `tests/sim/python-on-path.yaml`: the unit command says `python`, and the tests need a module only the project's environment has. Round 1's unit lane should pass. - `tests/sim/fix-only-lane-command.yaml`: the fix round's whole correction is the unit command. It should succeed, not be skipped. The fix removes both markers; the sim lane fails until it does. Run them with `python -m osf.sim python-on-path fix-only-lane-command`. For the console part, `python -m osf.sim --keep red-lane-then-fix` leaves a run up with round 1's judge skipped for a red lane, to check its row against.
Author
Owner

Shipped in two changes, both live and archived.

1. lanes-use-the-projects-environment (61ff22a)

  • A lane runs with Braid's own virtualenv off PATH and the project's environments first: each .venv/bin and node_modules/.bin that script.dependencies makes. So python -m pytest in a lane is the project's Python, as on CI. This covers the fix loop's lanes, an OpenSpec change's unit-test step, and apply test. Braid's own steps call $OSF_PYTHON and are unchanged.
    • The cause: the image puts Braid's virtualenv on PATH without setting VIRTUAL_ENV, so the existing removal of it never ran.
  • The fix loop's "changed since the check" and "unchanged since the check" count the lane commands, so a fix whose whole correction is .osf/test/unit.command passes its check.
  • A fix round whose note is there but which changed nothing is still due, so its retry runs instead of being skipped.
  • python-on-path and fix-only-lane-command pass; both known_failure markers are gone, and each scenario fails again with its fix reverted.

2. fix-loop-skip-says-why (e67c248, live with adb75b7)

  • The loop gives a few words for each of its states. A skipped fix-loop step shows them in place of "not needed" (e.g. lanes failed, fix made, checks passed), with the whole reason in the hover.
  • A step that ran and was skipped on its retry is no longer struck through, keeps its time, and its step view says "This step ran, then was skipped: …".
  • Migration 10 adds step.skip_reason and step.skip_note, and the API adds ran. Run 48's agent.fix now reads ran: true. Its recorded reason stays empty until an osfd rebuild, because it predates the migration.
  • Checked with Playwright on the fixture, and in a real browser on a kept red-lane-then-fix simulated run.

Deploy note: two deploys failed the self-check on a /tmp/tmp.* left behind. The new listing named it: Chromium's profile, from the browser check, recreated after its scratch directory was removed. adb75b7 stops every process naming that directory before removing it.

Shipped in two changes, both live and archived. **1. `lanes-use-the-projects-environment`** (61ff22a) - A lane runs with Braid's own virtualenv off `PATH` and the project's environments first: each `.venv/bin` and `node_modules/.bin` that `script.dependencies` makes. So `python -m pytest` in a lane is the project's Python, as on CI. This covers the fix loop's lanes, an OpenSpec change's unit-test step, and `apply test`. Braid's own steps call `$OSF_PYTHON` and are unchanged. - The cause: the image puts Braid's virtualenv on `PATH` without setting `VIRTUAL_ENV`, so the existing removal of it never ran. - The fix loop's "changed since the check" and "unchanged since the check" count the lane commands, so a fix whose whole correction is `.osf/test/unit.command` passes its check. - A fix round whose note is there but which changed nothing is still due, so its retry runs instead of being skipped. - `python-on-path` and `fix-only-lane-command` pass; both `known_failure` markers are gone, and each scenario fails again with its fix reverted. **2. `fix-loop-skip-says-why`** (e67c248, live with adb75b7) - The loop gives a few words for each of its states. A skipped fix-loop step shows them in place of "not needed" (e.g. `lanes failed`, `fix made`, `checks passed`), with the whole reason in the hover. - A step that ran and was skipped on its retry is no longer struck through, keeps its time, and its step view says "This step ran, then was skipped: …". - Migration 10 adds `step.skip_reason` and `step.skip_note`, and the API adds `ran`. Run 48's `agent.fix` now reads `ran: true`. Its recorded reason stays empty until an `osfd rebuild`, because it predates the migration. - Checked with Playwright on the fixture, and in a real browser on a kept `red-lane-then-fix` simulated run. **Deploy note:** two deploys failed the self-check on a `/tmp/tmp.*` left behind. The new listing named it: Chromium's profile, from the browser check, recreated after its scratch directory was removed. adb75b7 stops every process naming that directory before removing it.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
cmoriarty/braid#148
No description provided.