The browser check of a multi-player app takes 40–80 min: drive the state with a script and check what people see #127

Closed
opened 2026-10-02 15:51:01 -04:00 by cmoriarty · 1 comment
Owner

Split off from #121 (finding 7 of the scratch-run analysis, docs/design/measurements/2026-10-02-scratch-run-efficiency.md).

agent.test.browser took 38 min in run 37 and 80 min over two attempts in run 36 on scratch, a five-player card game.

  • In run 37 the browser subagent ran out of steps on a full 10-hand game (about 105 moves). The primary then thought for 6.4 and 4.5 min and wrote a websocket script to play the other four seats while the browser played one.
  • In run 36 the subagent's first run (21 min) failed on a disconnected "ghost" seat its own tab handling had left. It blamed "an external connection"; the primary found later that it was the subagent's own discarded tab.
  • Each browser brief re-reads the app's source to learn how seats work, which the primary already knows.
  • Run 37's agent.reconcile-tasks ran a browser check too, the job agent.test.browser does next.

Ideas:

  • Tell the check to verify what a person sees at states that are cheap to reach, with the state driven by a script or the API rather than by clicking through a whole game.
  • Hand the subagent the app facts the primary gathered: how a page joins, which controls matter.
  • Keep browser checks out of reconcile.

Needs a design, and maybe a hook in the project (a seeded short game) that the brief can point at.

Split off from #121 (finding 7 of the scratch-run analysis, `docs/design/measurements/2026-10-02-scratch-run-efficiency.md`). `agent.test.browser` took 38 min in run 37 and 80 min over two attempts in run 36 on scratch, a five-player card game. - In run 37 the browser subagent ran out of steps on a full 10-hand game (about 105 moves). The primary then thought for 6.4 and 4.5 min and wrote a websocket script to play the other four seats while the browser played one. - In run 36 the subagent's first run (21 min) failed on a disconnected "ghost" seat its own tab handling had left. It blamed "an external connection"; the primary found later that it was the subagent's own discarded tab. - Each browser brief re-reads the app's source to learn how seats work, which the primary already knows. - Run 37's `agent.reconcile-tasks` ran a browser check too, the job `agent.test.browser` does next. Ideas: - Tell the check to verify what a person sees at states that are cheap to reach, with the state driven by a script or the API rather than by clicking through a whole game. - Hand the subagent the app facts the primary gathered: how a page joins, which controls matter. - Keep browser checks out of reconcile. Needs a design, and maybe a hook in the project (a seeded short game) that the brief can point at.
Author
Owner

Shipped and live on production (9ae119c, deployed 2026-10-03). The effect on time is not known until a scratch run with a UI change measures it. Against the three ideas:

1. Check what a person sees at states that are cheap to reach — shipped. The browser check's prompt now asks the agent to decide, for each behaviour, whether its state is already there when the page loads, a short click path away, or many actions away, and to reach the last one itself by the first of these the project allows: a documented seed or fixture command, the project's own test client or helpers (it names AGENTS.md, the README, the package scripts, the scripts directory and the project's tests as where to look), or a script of its own over the API or websocket the page uses that plays every person but the one browser holds (kept under $TMPDIR, so not committed). For that one person it may give browser one short driver for one tab, run once; a driver in each of several tabs, or one that waits on another tab, is ruled out, and so is changing the app to make a state easier to reach. browser is told it has about 40 turns and that a check may need about 25 actions (both read from the subagent's definition); it looks at the state it is given, reports UNCLEAR naming the state to script when the steps are longer, and does not write loops across tabs. A browser that reports it reached its maximum steps is not resumed to carry on the same flow (run 39 did that seven times); the agent starts a shorter check from a scripted state, or reports the check unclear. Each check's note says whether its state was loaded, clicked or scripted.

2. Hand the subagent the app facts — shipped. Every brief carries what the agent already knows: the page's URL with its path or query (a room code, a seat), the controls and text that matter, what passing shows, what not to touch; the subagent is told to work from the brief and read the source only for what it leaves out. Tabs: one per person, closed when done with, and each brief starts by closing the tabs an earlier task left, because the browser outlives a task (run 41's second call found four stale ones); a seat that looks taken is probably an earlier tab of the check (run 36's ghost seat).

3. Keep browser checks out of reconcile — shipped. Reconcile's prompt and both unit lanes' are told not to check the UI in a browser and not to start browser when the run has it; with browser switched off in settings they do not mention it. The shared delegation guidance, where an apply step that changed the UI may take a look, is not changed on this point.

The project hook. This repository documents its own in AGENTS.md: window.braid, and the fixture backend (ui/dev/fakeosfd.py) with a named run for each state and POST /_fake/... endpoints that move a state (a test keeps the names true). For scratch, the hook would be its own test client (test_full_game.py, play_pinned_hand) or a seeded short game, which the prompt tells the agent to look for first and to prefer to a script of its own.

Checked: tests for the prompt (the game case, the short path, hooks, drivers, too-long checks, not resuming, what a brief carries, tabs, the note) and for the subagent's text; the fast and full lanes; the prompt read for a five-player game whose delta spec describes the end-of-game screen, which leads to a script that plays the other seats and at most one short driver, not five tabs; and on production in the container, the browser subagent's rendered text is the new one while reader, explore and test hash identically to before. To measure in the next scratch run with a UI change: the step's duration against 38 and 80 minutes, the number of browser calls, and whether a script played the other seats.

Change record: openspec/changes/archive/2026-10-03-browser-check-drives-state.

Shipped and live on production (`9ae119c`, deployed 2026-10-03). The effect on time is not known until a scratch run with a UI change measures it. Against the three ideas: **1. Check what a person sees at states that are cheap to reach — shipped.** The browser check's prompt now asks the agent to decide, for each behaviour, whether its state is already there when the page loads, a short click path away, or many actions away, and to reach the last one itself by the first of these the project allows: a documented seed or fixture command, the project's own test client or helpers (it names `AGENTS.md`, the README, the package scripts, the scripts directory and the project's tests as where to look), or a script of its own over the API or websocket the page uses that plays every person but the one `browser` holds (kept under `$TMPDIR`, so not committed). For that one person it may give `browser` one short driver for one tab, run once; a driver in each of several tabs, or one that waits on another tab, is ruled out, and so is changing the app to make a state easier to reach. `browser` is told it has about 40 turns and that a check may need about 25 actions (both read from the subagent's definition); it looks at the state it is given, reports `UNCLEAR` naming the state to script when the steps are longer, and does not write loops across tabs. A `browser` that reports it reached its maximum steps is not resumed to carry on the same flow (run 39 did that seven times); the agent starts a shorter check from a scripted state, or reports the check `unclear`. Each check's `note` says whether its state was loaded, clicked or scripted. **2. Hand the subagent the app facts — shipped.** Every brief carries what the agent already knows: the page's URL with its path or query (a room code, a seat), the controls and text that matter, what passing shows, what not to touch; the subagent is told to work from the brief and read the source only for what it leaves out. Tabs: one per person, closed when done with, and each brief starts by closing the tabs an earlier task left, because the browser outlives a `task` (run 41's second call found four stale ones); a seat that looks taken is probably an earlier tab of the check (run 36's ghost seat). **3. Keep browser checks out of reconcile — shipped.** Reconcile's prompt and both unit lanes' are told not to check the UI in a browser and not to start `browser` when the run has it; with `browser` switched off in settings they do not mention it. The shared delegation guidance, where an apply step that changed the UI may take a look, is not changed on this point. **The project hook.** This repository documents its own in `AGENTS.md`: `window.braid`, and the fixture backend (`ui/dev/fakeosfd.py`) with a named run for each state and `POST /_fake/...` endpoints that move a state (a test keeps the names true). For scratch, the hook would be its own test client (`test_full_game.py`, `play_pinned_hand`) or a seeded short game, which the prompt tells the agent to look for first and to prefer to a script of its own. Checked: tests for the prompt (the game case, the short path, hooks, drivers, too-long checks, not resuming, what a brief carries, tabs, the note) and for the subagent's text; the fast and full lanes; the prompt read for a five-player game whose delta spec describes the end-of-game screen, which leads to a script that plays the other seats and at most one short driver, not five tabs; and on production in the container, the `browser` subagent's rendered text is the new one while `reader`, `explore` and `test` hash identically to before. **To measure in the next scratch run with a UI change**: the step's duration against 38 and 80 minutes, the number of `browser` calls, and whether a script played the other seats. Change record: `openspec/changes/archive/2026-10-03-browser-check-drives-state`.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
cmoriarty/braid#127
No description provided.