Give Braid's test subagent a browser it can drive (Playwright MCP) #17

Closed
opened 2026-09-13 23:47:09 -04:00 by cmoriarty · 1 comment
Owner

Braid's agents can run Playwright test files (npx playwright test through bash), but they cannot drive a browser: open a page, click through it, read the console and look at a screenshot of what is actually on screen. That is what verifying a UI change needs, and what the operator's own opencode test subagent does.

Why

  • On openspec-flow#11's runs, the verification task said to preview the change "in a real browser … (a screenshot is fine)". The agent wrote throwaway Playwright scripts through bash, and in the first run reported "this model can't read PNGs": the run's model entry does not declare image input.
  • A scripted spec checks what someone thought to assert. A driven browser plus a screenshot the agent reads catches layout, clipping, overlap and console errors nobody wrote an assertion for.

What exists

  • src/osf/engine/shape.py can already attach @playwright/mcp to a node (--headless --output-dir <screenshots>), and playwright is in KNOWN_MCP_SERVERS. No node uses it, the package is unpinned (@latest), and there is no tool allowlist because its cost was never measured.
  • The operator's test subagent (~/.config/opencode/opencode.jsonc) gets playwright* with browser_snapshot (the whole-page accessibility tree, ~20k tokens a call), run_code_unsafe, file_upload, drop and tabs disabled, and works from find + screenshots it reads itself.

Proposal

Once the apply-delegates-to-subagents change lands (a test subagent with no editing, background runs, image input declared on the model):

  • Attach Playwright MCP to the test subagent only, pinned to a version, with a measured allowlist (no browser_snapshot), screenshots written under the worktree's gitignored ui/test-results/.
  • The primary starts the dev server or preview on the slot's PREVIEW_PORT; test gets the URL, a scripted scenario and pass criteria, and returns PASS / FAIL / UNCLEAR with screenshot paths and quoted console errors.
  • One browser per run at a time; the transcript shows the screenshots test took under its delegation header.

Open questions

  • Is npx fetching the MCP package at run time acceptable, or should it be pre-installed on the host?
  • The token floor of the allowlisted tools, measured the way MCP_ALLOWLIST measured forgejo's.
  • Headless Chromium on the host running osfd: already present for the UI's own Playwright suite, but not under a run's isolated HOME.
Braid's agents can run Playwright **test files** (`npx playwright test` through bash), but they cannot **drive a browser**: open a page, click through it, read the console and look at a screenshot of what is actually on screen. That is what verifying a UI change needs, and what the operator's own opencode `test` subagent does. ## Why - On openspec-flow#11's runs, the verification task said to preview the change "in a real browser … (a screenshot is fine)". The agent wrote throwaway Playwright scripts through bash, and in the first run reported "this model can't read PNGs": the run's model entry does not declare image input. - A scripted spec checks what someone thought to assert. A driven browser plus a screenshot the agent reads catches layout, clipping, overlap and console errors nobody wrote an assertion for. ## What exists - `src/osf/engine/shape.py` can already attach `@playwright/mcp` to a node (`--headless --output-dir <screenshots>`), and `playwright` is in `KNOWN_MCP_SERVERS`. No node uses it, the package is unpinned (`@latest`), and there is no tool allowlist because its cost was never measured. - The operator's `test` subagent (`~/.config/opencode/opencode.jsonc`) gets `playwright*` with `browser_snapshot` (the whole-page accessibility tree, ~20k tokens a call), `run_code_unsafe`, `file_upload`, `drop` and `tabs` disabled, and works from `find` + screenshots it reads itself. ## Proposal Once the `apply-delegates-to-subagents` change lands (a `test` subagent with no editing, background runs, image input declared on the model): - Attach Playwright MCP to the `test` subagent only, pinned to a version, with a measured allowlist (no `browser_snapshot`), screenshots written under the worktree's gitignored `ui/test-results/`. - The primary starts the dev server or preview on the slot's `PREVIEW_PORT`; `test` gets the URL, a scripted scenario and pass criteria, and returns PASS / FAIL / UNCLEAR with screenshot paths and quoted console errors. - One browser per run at a time; the transcript shows the screenshots `test` took under its delegation header. ## Open questions - Is `npx` fetching the MCP package at run time acceptable, or should it be pre-installed on the host? - The token floor of the allowlisted tools, measured the way `MCP_ALLOWLIST` measured forgejo's. - Headless Chromium on the host running osfd: already present for the UI's own Playwright suite, but not under a run's isolated `HOME`.
Author
Owner

Shipped in #91 (OpenSpec change 2026-09-27-browser-subagent, archived on the branch); this issue closes when it merges.

How it differs from the proposal here:

  • A separate browser subagent, not test. test runs suites often, and the browser tools would be re-sent on every one of its turns; a narrow prompt also suits a 27B better. browser has no shell and no edits.
  • Allowlist, measured on opencode 1.18.21: 13 of Playwright MCP's 24 tools, 9,251 chars of schema instead of 17,678. No browser_snapshot; find returns the refs click takes. Also out: run_code_unsafe, file upload, drag/drop, tabs, network.
  • Tool routing: a task child inherits its parent prompt's tool map, so a prompt-level playwright*: false took the browser off the subagent too. Playwright is now switched off per agent in the run config.
  • Open questions answered: the server is pinned (@playwright/mcp@0.0.80, same Chromium 1243 as the image) and pre-installed in the image, not fetched by npx. Headless Chromium runs fine under the run's isolated HOME. Screenshots go to <run>/screenshots, outside the worktree, not ui/test-results/ (a relative filename wrote into the worktree, so the run plugin drops it).
  • One browser at a time is enforced by the run plugin, which refuses a second browser task while one is busy.
  • The primary serves the app on PREVIEW_PORT and hands browser the URL, steps and pass criteria.
  • Screenshots in the transcript: shown under each browser · screenshot call in the subagents panel.
  • Beyond the issue: a pipeline step, agent.test.browser, in default, git-flow, minimalist, git-flow-hotfix and quick-fix. It runs when UI files changed, the app can be served and the run has browser, and a report still saying fail fails the step.

Not yet verified: qwen3.8-27b actually driving the browser (every test used a scripted fake model), and the image build with the new install.

Shipped in #91 (OpenSpec change `2026-09-27-browser-subagent`, archived on the branch); this issue closes when it merges. How it differs from the proposal here: - **A separate `browser` subagent, not `test`.** `test` runs suites often, and the browser tools would be re-sent on every one of its turns; a narrow prompt also suits a 27B better. `browser` has no shell and no edits. - **Allowlist, measured on opencode 1.18.21:** 13 of Playwright MCP's 24 tools, 9,251 chars of schema instead of 17,678. No `browser_snapshot`; `find` returns the refs `click` takes. Also out: `run_code_unsafe`, file upload, drag/drop, tabs, network. - **Tool routing:** a `task` child inherits its parent prompt's tool map, so a prompt-level `playwright*: false` took the browser off the subagent too. Playwright is now switched off per agent in the run config. - **Open questions answered:** the server is pinned (`@playwright/mcp@0.0.80`, same Chromium 1243 as the image) and **pre-installed in the image**, not fetched by npx. Headless Chromium runs fine under the run's isolated HOME. Screenshots go to `<run>/screenshots`, outside the worktree, not `ui/test-results/` (a relative filename wrote into the worktree, so the run plugin drops it). - **One browser at a time** is enforced by the run plugin, which refuses a second `browser` task while one is busy. - **The primary serves the app** on `PREVIEW_PORT` and hands `browser` the URL, steps and pass criteria. - **Screenshots in the transcript:** shown under each `browser · screenshot` call in the subagents panel. - **Beyond the issue:** a pipeline step, `agent.test.browser`, in default, git-flow, minimalist, git-flow-hotfix and quick-fix. It runs when UI files changed, the app can be served and the run has `browser`, and a report still saying fail fails the step. Not yet verified: qwen3.8-27b actually driving the browser (every test used a scripted fake model), and the image build with the new install.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
cmoriarty/braid#17
No description provided.