The judge's report is rejected in every OpenSpec run, and the loop then skips the judge and fixes what the report found #145

Closed
opened 2026-10-05 14:47:24 -04:00 by cmoriarty · 1 comment
Owner

Run 47's agent.judge spent 75 minutes playing a whole ten-hand game in a browser, with five live clients. All 28 of its criteria passed or were unclear. Then the step ended "not needed", and the loop sent a fix round after a defect that isn't one. Three bugs in #139's loop did this, in the order they bit:

  1. The report's schema rejects a change's criterion ids. #139 asks a change's judge to name each criterion <capability>/<scenario>, the tag the tests carry. The judge did, with ids of 29 to 76 characters. JUDGE_SCHEMA still caps an id at 20 characters, a limit set for the quick fixes' A1, A2. So every OpenSpec run's judge report fails its check: $.criteria[0].id: longer than 20.
  2. The loop trusts a report its check rejected. The retry re-reads the step's condition, and osf.pipeline.fixloop.state read the rejected judge.json as a verdict. So the judge was "not needed", the retry never ran, and the rejected report drove the loop.
  3. A served app is asked to open from disk. The loop lists any index.html one folder down as a page to open from disk. frontend/index.html is a Vite entry that loads /src/main.tsx, so it works only served, and served it did work. The judge recorded the file:// opening as not working, while calling the failure "inherent to the stack, not a defect", and that one entry made the round fail. Round 1's fix then spent 17 minutes changing five app files to make it open from disk, and round 2's judge took another 71 minutes.

Run 47 was stopped in round 2.

Asked for

  • a criterion id long enough for a scenario's name;
  • the loop reads only a report that passes its check, so a rejected report is retried, not acted on;
  • opening from disk is asked only of a page meant to be opened that way, never of a dev server's entry, and a served app is judged served.
Run 47's `agent.judge` spent 75 minutes playing a whole ten-hand game in a browser, with five live clients. All 28 of its criteria passed or were unclear. Then the step ended "not needed", and the loop sent a fix round after a defect that isn't one. Three bugs in #139's loop did this, in the order they bit: 1. **The report's schema rejects a change's criterion ids.** #139 asks a change's judge to name each criterion `<capability>/<scenario>`, the tag the tests carry. The judge did, with ids of 29 to 76 characters. `JUDGE_SCHEMA` still caps an id at 20 characters, a limit set for the quick fixes' `A1`, `A2`. So every OpenSpec run's judge report fails its check: `$.criteria[0].id: longer than 20`. 2. **The loop trusts a report its check rejected.** The retry re-reads the step's condition, and `osf.pipeline.fixloop.state` read the rejected `judge.json` as a verdict. So the judge was "not needed", the retry never ran, and the rejected report drove the loop. 3. **A served app is asked to open from disk.** The loop lists any `index.html` one folder down as a page to open from disk. `frontend/index.html` is a Vite entry that loads `/src/main.tsx`, so it works only served, and served it did work. The judge recorded the `file://` opening as not working, while calling the failure "inherent to the stack, not a defect", and that one entry made the round fail. Round 1's fix then spent 17 minutes changing five app files to make it open from disk, and round 2's judge took another 71 minutes. Run 47 was stopped in round 2. **Asked for** - a criterion id long enough for a scenario's name; - the loop reads only a report that passes its check, so a rejected report is retried, not acted on; - opening from disk is asked only of a page meant to be opened that way, never of a dev server's entry, and a served app is judged served.
Author
Owner

Shipped, deployed and archived.

What changed (578e003; archived in 0447d51)

  • A criterion id may be up to 200 characters, so a change's <capability>/<scenario> ids pass the judge's schema.
  • The loop reads a judge's report only once it passes its check. A report that fails counts as not written: the judge is still due, its retry runs, and no fix round starts on it.
  • A retried judge whose report failed its check is told so, with the errors, and asked to correct the report's form, keeping its verdicts and evidence.
  • A page that loads its code from a dev server (/src/…, or a .ts, .tsx or .jsx script) is no longer listed to open from disk, so a served app is judged served.
  • Specs: deliverable-judge and fix-loop.

Verified

  • Unit tests for each. In production, run 47's own report passes the schema, and its Vite entry is no longer listed to open from disk.
  • Run 48's round-2 judge wrote a report that passed its check on its first write, and judged the app served (Vite and the backend), not from disk.
  • The new simulated runs (#149) guard it:
    • with the id limit set back to 20, judge-long-ids fails in 16 s, naming the error;
    • judge-report-retried shows the retry told its report is there, and correcting it.

Run 48 still failed in that judge, for another reason. Its first attempt stopped at two hours with no report, and left dryrun1.log in the worktree. The second attempt's report was valid, but the stray log failed its "code unchanged" check. That is #150.

Shipped, deployed and archived. **What changed** (578e003; archived in 0447d51) - A criterion id may be up to 200 characters, so a change's `<capability>/<scenario>` ids pass the judge's schema. - The loop reads a judge's report only once it passes its check. A report that fails counts as not written: the judge is still due, its retry runs, and no fix round starts on it. - A retried judge whose report failed its check is told so, with the errors, and asked to correct the report's form, keeping its verdicts and evidence. - A page that loads its code from a dev server (`/src/…`, or a `.ts`, `.tsx` or `.jsx` script) is no longer listed to open from disk, so a served app is judged served. - Specs: `deliverable-judge` and `fix-loop`. **Verified** - Unit tests for each. In production, run 47's own report passes the schema, and its Vite entry is no longer listed to open from disk. - Run 48's round-2 judge wrote a report that passed its check on its first write, and judged the app served (Vite and the backend), not from disk. - The new simulated runs (#149) guard it: - with the id limit set back to 20, `judge-long-ids` fails in 16 s, naming the error; - `judge-report-retried` shows the retry told its report is there, and correcting it. **Run 48 still failed in that judge, for another reason.** Its first attempt stopped at two hours with no report, and left `dryrun1.log` in the worktree. The second attempt's report was valid, but the stray log failed its "code unchanged" check. That is #150.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
cmoriarty/braid#145
No description provided.