feat(agents): a browser subagent checks UI changes in a real browser, and a pipeline step uses it (#17) #91
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "claude/playwright-test-subagent"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Closes #17.
Braid's agents could run Playwright spec files but not drive a browser and look at the result. This adds a fourth read-only subagent,
browser, that drives headless Chromium through Playwright MCP, and a pipeline step,agent.test.browser, that uses it to check a UI change before the commit.Changes
browsersubagent (on by default): 13 of Playwright MCP's 24 tools (9,251 chars of schema instead of 17,678), no whole-page snapshot, no shell, no edits. It reports PASS / FAIL / UNCLEAR with screenshot names and quoted console errors. No other agent gets the browser.taskchild inherits its parent prompt's tool map, so Playwright is switched off in each agent's config rather than in any prompt;tools_forno longer sendsplaywright*.@playwright/mcp@0.0.80(Chromium 1243, the image's revision), pre-installed in the image (OSF_PLAYWRIGHT_MCP), headless and isolated, screenshots to<run>/screenshots. The self-check fails on a version drift.filenamefrom Playwright calls (a relative one wrote into the worktree) and refuses a secondbrowsertask while one is busy.$PREVIEW_PORTand givebrowserthe URL.agent.test.browser: indefault,git-flow,minimalist,git-flow-hotfixandquick-fix, after the last test lane; runs only when UI files changed, the app can be served and the run hasbrowser(newbrowser_checkcondition). It writes.osf/test/browser.json; a report still saying fail fails the step.browser · <action>in the subagents panel with each screenshot under its call, served byGET /api/runs/{id}/screenshots/{name}(screenshot-shaped names only).mainafteragent-step-names; the step followsagent.test.unit/agent.test.e2e.Verification
OSF_URLat a--statusfixture) pass on the merged tree: backend 1832, UI unit 431, browser 99/99.tests/integration/test_browser_live.py(-m needs_opencode) against opencode 1.18.21: the primary has no browser tools,browserhas exactly the allowlist, a second browser task is refused while the first reports, and the screenshot lands outside the worktree. It fails when the prompt-level glob or the filename drop is put back.agent-view.spec.ts"flows a few characters a frame" fail once on a loaded machine; it passed 3/3 alone and in every later run.Not verified
npm install -g.browser, so its check is skipped untilbrowseris ticked in settings.🤖 Generated with Claude Code
Braid's agents could run Playwright spec files through bash but could not drive a browser and look at what was on screen. Now a fourth read-only subagent, `browser`, drives headless Chromium through Playwright MCP, and `agent.test.browser` has it check a UI change before the commit. - Measured on opencode 1.18.21: a task child inherits its parent prompt's tool map, so Playwright is switched off per agent in the run config, not in any prompt. `browser` gets 13 of the server's 24 tools (9,251 chars of schema, not 17,678); no whole-page snapshot. - The server is pinned to @playwright/mcp@0.0.80 (Chromium 1243, as the image installs), pre-installed in the image, headless and isolated, and writes screenshots to <run>/screenshots. - The run's plugin drops `filename` from Playwright calls (one wrote into the worktree) and refuses a second `browser` task while one is busy. - Implementation steps are told to serve the app on $PREVIEW_PORT and hand `browser` the URL. - agent.test.browser runs in default, git-flow, minimalist, git-flow-hotfix and quick-fix when UI files changed, the app can be served and the run has `browser`; a report still saying fail fails the step. - The console titles browser calls `browser · <action>` and shows each screenshot under its call, served by GET /api/runs/{id}/screenshots/{name}. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>