Run 48's judge spent 190 minutes, its first attempt left no report, and a stray log in the worktree failed the run #150

Closed
opened 2026-10-05 22:10:52 -04:00 by cmoriarty · 1 comment
Owner

What happened in run 48 (git-flow, round 2's agent.judge.r2)

  • Attempt 1 (two-hour limit). The judge wrote its own game driver, judge_drive.py, and played the game with bots and a browser.
    • Its first dry run was … judge_drive.py … | tee dryrun1.log, run from the worktree, so dryrun1.log landed in the repository. Its later dry runs wrote under $TMPDIR.
    • It stopped at its two-hour limit with no report.
  • Attempt 2 (a four-hour limit, the worktree kept).
    • It wrote .osf/fix/r2/judge.json, which passed its check (#145), and judged the app served.
    • Its completion check then failed on since-verify:unchanged: "the code changed after round 2's check: dryrun1.log", the file attempt 1 left.
    • The retries were used up, and the run failed at 02:07Z, before its pull request.
  • Cost. The judge took 190 minutes, 303 model turns and 371,260 generated tokens. That is 83% of what the run generated after its apply steps (444,875).

Why

  1. Nothing keeps a judge's scratch files out of the worktree, and its check counts any untracked file as a code change.
  2. A retry keeps the worktree, so a later attempt inherits an earlier one's stray files, and its brief does not mention them.
  3. The first attempt spent its whole two hours exploring and driving the game, and wrote nothing down. What it had found was lost with it.

Proposal

  • The judge's brief says where scratch files go ($TMPDIR).
  • A judge whose check failed only on changed files is retried with a brief that names the files, asks it to remove them, and keeps its report. This is #145's retry, for this failure.
  • The judge writes its report as it goes, criterion by criterion. Then a stop at the time limit keeps what it judged, and the last call before its deadline asks for the report it has.
  • A simulated run for it (#149): a judge that leaves a scratch file in the worktree.
**What happened in run 48** (git-flow, round 2's `agent.judge.r2`) - **Attempt 1** (two-hour limit). The judge wrote its own game driver, `judge_drive.py`, and played the game with bots and a browser. - Its first dry run was `… judge_drive.py … | tee dryrun1.log`, run from the worktree, so `dryrun1.log` landed in the repository. Its later dry runs wrote under `$TMPDIR`. - It stopped at its two-hour limit with no report. - **Attempt 2** (a four-hour limit, the worktree kept). - It wrote `.osf/fix/r2/judge.json`, which passed its check (#145), and judged the app served. - Its completion check then failed on `since-verify:unchanged`: "the code changed after round 2's check: dryrun1.log", the file attempt 1 left. - The retries were used up, and the run failed at 02:07Z, before its pull request. - **Cost.** The judge took 190 minutes, 303 model turns and 371,260 generated tokens. That is 83% of what the run generated after its apply steps (444,875). **Why** 1. Nothing keeps a judge's scratch files out of the worktree, and its check counts any untracked file as a code change. 2. A retry keeps the worktree, so a later attempt inherits an earlier one's stray files, and its brief does not mention them. 3. The first attempt spent its whole two hours exploring and driving the game, and wrote nothing down. What it had found was lost with it. **Proposal** - The judge's brief says where scratch files go (`$TMPDIR`). - A judge whose check failed only on changed files is retried with a brief that names the files, asks it to remove them, and keeps its report. This is #145's retry, for this failure. - The judge writes its report as it goes, criterion by criterion. Then a stop at the time limit keeps what it judged, and the last call before its deadline asks for the report it has. - A simulated run for it (#149): a judge that leaves a scratch file in the worktree.
Author
Owner

Shipped in 8bce10d, live, archived as judge-keeps-its-work.

What changed

  • Scratch files. The judge's brief puts scratch files (drivers, logs, captures) under $TMPDIR by absolute path. Its report is the only file it writes in the repository.
  • The report as it goes. The judge writes judge.json once it has opened the deliverable, with "finished": false, and adds each criterion as it judges it. An unfinished report fails its check, so the loop takes no verdict from it. Its retry is told which criteria are judged, keeps them, and judges only the rest. A report without finished reads as finished.
  • A judge that left files is retried. In run 48 the stray log ended the step at once: the judge's node is retried only for a missing or malformed report, so the retry count was not what stopped it. A check that failed only on files added after the lanes ran is now retried. The retry's brief names the files, asks the judge to delete them, and keeps its report. A judge that edited a file the lanes ran against still stops for a person.
  • Two engine fixes this needed.
    • Mediation now reads only the checks that failed. The passing report check beside the failed one had blocked the retry.
    • The loop keeps a passing judge due while files it added are in the worktree. Otherwise the retry is skipped as not needed and the closing check fails on the log.

Verified: two new simulated runs, judge-stray-file (run 48's dryrun1.log) and judge-unfinished. Both were written first as known failures, now pass, and fail again with their fix reverted. The deploy's self-check ran all 11 scenarios.

Not changed, for you to decide: how long a judge may take, or how much of a multi-player app it drives. Run 48's judge spent 190 minutes and 371,260 generated tokens. Today the limit is 2 hours, doubled on a timed-out retry. The background is in #127.

Known gap: a judge that fails the deliverable and also leaves a stray file goes on to a fix round, and the file can end up in that round's code. The $TMPDIR rule is what prevents it.

Shipped in 8bce10d, live, archived as `judge-keeps-its-work`. **What changed** - **Scratch files.** The judge's brief puts scratch files (drivers, logs, captures) under `$TMPDIR` by absolute path. Its report is the only file it writes in the repository. - **The report as it goes.** The judge writes `judge.json` once it has opened the deliverable, with `"finished": false`, and adds each criterion as it judges it. An unfinished report fails its check, so the loop takes no verdict from it. Its retry is told which criteria are judged, keeps them, and judges only the rest. A report without `finished` reads as finished. - **A judge that left files is retried.** In run 48 the stray log ended the step at once: the judge's node is retried only for a missing or malformed report, so the retry count was not what stopped it. A check that failed only on files *added* after the lanes ran is now retried. The retry's brief names the files, asks the judge to delete them, and keeps its report. A judge that edited a file the lanes ran against still stops for a person. - **Two engine fixes this needed.** - Mediation now reads only the checks that failed. The passing report check beside the failed one had blocked the retry. - The loop keeps a passing judge due while files it added are in the worktree. Otherwise the retry is skipped as not needed and the closing check fails on the log. **Verified:** two new simulated runs, `judge-stray-file` (run 48's `dryrun1.log`) and `judge-unfinished`. Both were written first as known failures, now pass, and fail again with their fix reverted. The deploy's self-check ran all 11 scenarios. **Not changed, for you to decide:** how long a judge may take, or how much of a multi-player app it drives. Run 48's judge spent 190 minutes and 371,260 generated tokens. Today the limit is 2 hours, doubled on a timed-out retry. The background is in #127. **Known gap:** a judge that *fails* the deliverable and also leaves a stray file goes on to a fix round, and the file can end up in that round's code. The `$TMPDIR` rule is what prevents it.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
cmoriarty/braid#150
No description provided.