feat(ui): every console action has a key and a name, and the tests split into a fast lane and a full one (#30, #39) #58

Merged
cmoriarty merged 1 commit from feat/keyboard-commands-and-fast-tests into main 2026-09-26 18:40:10 -04:00
Owner

Closes #30, closes #39. OpenSpec change: openspec/changes/keyboard-commands-and-fast-tests (proposal, specs, design, tasks).

Keyboard commands (#30)

Shopify's point in the article is an agent-addressable app: an agent drives the app by named commands instead of reading the layout or the accessibility tree. Braid now works that way:

  • Every console action is a named command with keys. g r / g i / g m / g s go to runs, inbox, inference and settings; n new run; ] / [ next and previous run; j/k (and ↓/↑) walk the steps; g a jumps to the step the run is on; t / v / d / r open the transcript, review, details and redo lenses; Y accepts a failed step; h toggles thinking; a opens the subagents panel; i focuses the composer; S / R stop and resume the run. Esc closes the lens, then leaves the step, then the run.
  • A command is available exactly when its button is on screen. Redo and Stop only open their confirmations. Keys never fire while you type or while a dialog is open.
  • ? opens a sheet built from the same table the keys are dispatched from, so it can't drift. Commands the current view doesn't offer are dimmed.
  • window.braid: commands(), run(id) → whether it ran, and state() → view, run, step, lens, the run's steps with statuses, and open gates. It grants nothing a click can't do.

Measured (e2e/keyboard.spec.ts, the same six-move tour, median of 3): keys 94 ms against clicks 255 ms alone, and 241–360 ms against 595–617 ms with other specs running alongside. That's roughly 2–3× in Playwright, where locating an element is cheap. For an agent driving a browser through tool calls, the saving is larger: one key press or braid.run replaces a page read plus a find plus a click, and braid.state() replaces reading the page to check the result.

Test lanes (#39)

before after
backend verification 344 s 30 s (fast lane 28 s)
browser specs 61 s 17 s fast lane, 63 s full
everything, fast lane ~7 min 57 s
everything, full lane — 110 s
  • ./tools/test.sh is the fast lane: backend on every core (pytest-xdist), UI types, vitest, build, and the Playwright specs without @slow. ./tools/test.sh full runs everything, plus the @live specs when OSF_URL is set. CLAUDE.md, openspec/config.yaml, docs/guide.md and .osf/test/unit.command now point at the lanes.
  • Waste removed: five liveness tests waited out the real 30 s REDO_CANCEL_WAIT_S against a fake executor that never stops (190 s of the 344 s). Five executor tests ended through the real 5 s quiesce or the 3 s abort confirmation.
  • Kept, tagged @slow: 14 browser specs whose time is the measurement (streaming and layout observation windows). The two real-osfd specs are tagged @live.
  • The 5 s rule is enforced by a report: tests/fast_lane_budget.py names, at the end of every backend run, any test over 5 s that isn't marked slow.
  • Fixed on the way: settings.spec.ts keyed the fixture's saved settings by test id alone, so a second run against a fixture server left running failed. Each run now gets its own settings.
  • More Playwright workers made the browser lane slower and flakier (8 workers: 21 s, flaky, against 13 s with 4), because the fixture is a single Python process. The default stays.

Deploy (#39)

Deploy #25's log: all base images were CACHED, but every runtime RUN re-executed (apt 20 s, npm -g 41 s, Chromium 74 s, uv sync 10 s). The cause was ARG BRAID_COMMIT declared at the top of the stage: Docker passes a build arg to every later RUN, so each commit was a cache miss for all of them. The rebuilt ~1 GB of layers was also why the push took 2 m 14 s.

  • ARG BRAID_COMMIT now sits after the last RUN. Python dependencies install from the lock file before COPY src.
  • Verified locally: a build with only a new commit had every layer CACHED (0.5 s); a source change re-ran only the project install and the final step (1.5 s). braid-selfcheck passes and reports the commit in the env and the label.
  • wait-for-idle.sh confirms an idle production 5 s after the first zero instead of a full 30 s interval (still two readings). wait-for-commit.sh polls every 2 s instead of 10. Both have tests.
  • Expected: a source-only deploy drops from ~6 min to roughly a minute. The first deploy after this merge still rebuilds everything once, because the Dockerfile itself changed.

Verification

  • ./tools/test.sh: passed in 57 s. ./tools/test.sh full: passed in 110 s (1,741 backend, 394 vitest, 76 Playwright).
  • @live against the demo osfd on :8710: step-pane passes. Both follow-the-tail specs fail exactly as they did on main before this change: the running run opens on a view with no transcript. Not addressed here.
  • The shortcut sheet, j/k, lenses and Esc layering were checked by hand in a real browser against the fixture.

After merge: /opsx:archive keyboard-commands-and-fast-tests, then comment on and close #30 and #39.

🤖 Generated with Claude Code

Closes #30, closes #39. OpenSpec change: `openspec/changes/keyboard-commands-and-fast-tests` (proposal, specs, design, tasks). ## Keyboard commands (#30) Shopify's point in the article is an *agent-addressable* app: an agent drives the app by named commands instead of reading the layout or the accessibility tree. Braid now works that way: - **Every console action is a named command with keys.** `g r` / `g i` / `g m` / `g s` go to runs, inbox, inference and settings; `n` new run; `]` / `[` next and previous run; `j`/`k` (and ↓/↑) walk the steps; `g a` jumps to the step the run is on; `t` / `v` / `d` / `r` open the transcript, review, details and redo lenses; `Y` accepts a failed step; `h` toggles thinking; `a` opens the subagents panel; `i` focuses the composer; `S` / `R` stop and resume the run. `Esc` closes the lens, then leaves the step, then the run. - A command is available exactly when its button is on screen. Redo and Stop only *open* their confirmations. Keys never fire while you type or while a dialog is open. - **`?`** opens a sheet built from the same table the keys are dispatched from, so it can't drift. Commands the current view doesn't offer are dimmed. - **`window.braid`**: `commands()`, `run(id)` → whether it ran, and `state()` → view, run, step, lens, the run's steps with statuses, and open gates. It grants nothing a click can't do. **Measured** (`e2e/keyboard.spec.ts`, the same six-move tour, median of 3): **keys 94 ms against clicks 255 ms** alone, and 241–360 ms against 595–617 ms with other specs running alongside. That's roughly 2–3× in Playwright, where locating an element is cheap. For an agent driving a browser through tool calls, the saving is larger: one key press or `braid.run` replaces a page read plus a find plus a click, and `braid.state()` replaces reading the page to check the result. ## Test lanes (#39) | | before | after | |---|---|---| | backend verification | 344 s | 30 s (fast lane 28 s) | | browser specs | 61 s | 17 s fast lane, 63 s full | | everything, fast lane | ~7 min | **57 s** | | everything, full lane | — | 110 s | - `./tools/test.sh` is the fast lane: backend on every core (pytest-xdist), UI types, vitest, build, and the Playwright specs without `@slow`. `./tools/test.sh full` runs everything, plus the `@live` specs when `OSF_URL` is set. CLAUDE.md, `openspec/config.yaml`, `docs/guide.md` and `.osf/test/unit.command` now point at the lanes. - **Waste removed:** five liveness tests waited out the real 30 s `REDO_CANCEL_WAIT_S` against a fake executor that never stops (190 s of the 344 s). Five executor tests ended through the real 5 s quiesce or the 3 s abort confirmation. - **Kept, tagged `@slow`:** 14 browser specs whose time *is* the measurement (streaming and layout observation windows). The two real-osfd specs are tagged `@live`. - **The 5 s rule is enforced by a report:** `tests/fast_lane_budget.py` names, at the end of every backend run, any test over 5 s that isn't marked slow. - **Fixed on the way:** `settings.spec.ts` keyed the fixture's saved settings by test id alone, so a second run against a fixture server left running failed. Each run now gets its own settings. - More Playwright workers made the browser lane slower and flakier (8 workers: 21 s, flaky, against 13 s with 4), because the fixture is a single Python process. The default stays. ## Deploy (#39) Deploy #25's log: all base images were CACHED, but every runtime `RUN` re-executed (apt 20 s, `npm -g` 41 s, Chromium 74 s, `uv sync` 10 s). The cause was `ARG BRAID_COMMIT` declared at the top of the stage: Docker passes a build arg to every later `RUN`, so each commit was a cache miss for all of them. The rebuilt ~1 GB of layers was also why the push took 2 m 14 s. - `ARG BRAID_COMMIT` now sits after the last `RUN`. Python dependencies install from the lock file before `COPY src`. - **Verified locally:** a build with only a new commit had every layer CACHED (0.5 s); a source change re-ran only the project install and the final step (1.5 s). `braid-selfcheck` passes and reports the commit in the env and the label. - `wait-for-idle.sh` confirms an idle production 5 s after the first zero instead of a full 30 s interval (still two readings). `wait-for-commit.sh` polls every 2 s instead of 10. Both have tests. - Expected: a source-only deploy drops from ~6 min to roughly a minute. The first deploy after this merge still rebuilds everything once, because the Dockerfile itself changed. ## Verification - `./tools/test.sh`: passed in 57 s. `./tools/test.sh full`: passed in 110 s (1,741 backend, 394 vitest, 76 Playwright). - `@live` against the demo osfd on :8710: `step-pane` passes. Both `follow-the-tail` specs fail exactly as they did on `main` before this change: the running run opens on a view with no transcript. Not addressed here. - The shortcut sheet, `j`/`k`, lenses and Esc layering were checked by hand in a real browser against the fixture. After merge: `/opsx:archive keyboard-commands-and-fast-tests`, then comment on and close #30 and #39. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
Keyboard commands (#30): every operator action is a named command with keys (g i, j/k,
d, r, S, ...), `?` shows a sheet built from the same table, and window.braid lists,
runs and reports them so an agent can drive the console without reading the page. A tour
of six moves takes 94-360 ms by keys against 255-617 ms by clicks.

Test lanes (#39): tools/test.sh runs the fast lane (slow tests left out, backend on every
core) or the full lane. The backend verification went from 344 s to 30 s: the liveness
tests waited out a real 30 s redo timeout against a fake that never stops, and the suite
now runs under pytest-xdist. The browser lane went from 61 s to 17 s by tagging the
observation-window specs @slow. The backend names any unmarked test over 5 s.

Deploy (#39): ARG BRAID_COMMIT moves after the last RUN, so a new commit no longer
rebuilds and re-pushes apt, the Node tools, Chromium and the Python dependencies; an
idle production is confirmed in 5 s instead of 30 s, and the commit is polled every 2 s.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
cmoriarty deleted branch feat/keyboard-commands-and-fast-tests 2026-09-26 18:40:10 -04:00
Sign in to join this conversation.
No reviewers
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
cmoriarty/braid!58
No description provided.