Implementation steps delegate reading and testing to background subagents #18

Closed
opened 2026-09-13 23:48:21 -04:00 by cmoriarty · 2 comments
Owner

Braid's implementation steps (apply.group, tasks.reconcile, test.unit) do everything in one agent context: read whole files, run test suites and absorb their output, then hit the 112k-token alarm and are handed off to a fresh session that starts over.

On openspec-flow#11's run, apply.group[2] (verification) read 152k characters of files — one 56 KB fixture, whole — plus 32k characters of suite output, and was handed off at 113k tokens. Braid exposes opencode's task tool to these steps but defines no subagents, so neither group made a single delegation call.

The operator's own opencode setup works the other way: the primary gets as much context as it can use, and read-only subagents run in the background to read, look at images and run tests, returning short verdicts.

Proposal

Planned as the OpenSpec change apply-delegates-to-subagents (proposal, specs, design and tasks written, not yet applied):

  • A fixed, Braid-owned roster of read-only subagents for implementation steps: reader (digest files, logs, images), explore (where things are, who uses them), test (run the given commands, return a verdict). None can edit or delegate; opencode's general-purpose subagent is not offered.
  • Background subagents enabled in the run, with a run-home plugin that makes task's background argument a boolean (the model sends the string "True", which fails validation).
  • The implementation prompts open with when to delegate, and the primary is handed its group's code (or the line ranges its tasks cite) up front.
  • Two correctness fixes found while exploring: the context alarm and meter become per session (today a subagent's reading would trip the primary's handoff), and a handoff waits, bounded, for running subagents so their reports are not lost.
  • The model's image input is declared, so reader can look at screenshots.
  • UI: each subagent's collapsible header in the agent view shows its agent, task, background marker, state, elapsed time and token reading; its live text streams under the header; its report stays visible when collapsed; the stage header counts running subagents.

Out of scope: a browser for test (openspec-flow#17), subagents for the proposal and review steps, parallel worktrees.

Risk

Background subagents are experimental in opencode 1.18.30. The change starts with a spike on the real server; if they are unreliable, delegation runs in the foreground and the rest of the change still applies.

Braid's implementation steps (`apply.group`, `tasks.reconcile`, `test.unit`) do everything in one agent context: read whole files, run test suites and absorb their output, then hit the 112k-token alarm and are handed off to a fresh session that starts over. On openspec-flow#11's run, `apply.group[2]` (verification) read 152k characters of files — one 56 KB fixture, whole — plus 32k characters of suite output, and was handed off at 113k tokens. Braid exposes opencode's `task` tool to these steps but defines no subagents, so neither group made a single delegation call. The operator's own opencode setup works the other way: the primary gets as much context as it can use, and read-only subagents run in the background to read, look at images and run tests, returning short verdicts. ## Proposal Planned as the OpenSpec change `apply-delegates-to-subagents` (proposal, specs, design and tasks written, not yet applied): - A fixed, Braid-owned roster of read-only subagents for implementation steps: `reader` (digest files, logs, images), `explore` (where things are, who uses them), `test` (run the given commands, return a verdict). None can edit or delegate; opencode's general-purpose subagent is not offered. - Background subagents enabled in the run, with a run-home plugin that makes `task`'s `background` argument a boolean (the model sends the string "True", which fails validation). - The implementation prompts open with when to delegate, and the primary is handed its group's code (or the line ranges its tasks cite) up front. - Two correctness fixes found while exploring: the context alarm and meter become per session (today a subagent's reading would trip the primary's handoff), and a handoff waits, bounded, for running subagents so their reports are not lost. - The model's image input is declared, so `reader` can look at screenshots. - UI: each subagent's collapsible header in the agent view shows its agent, task, background marker, state, elapsed time and token reading; its live text streams under the header; its report stays visible when collapsed; the stage header counts running subagents. Out of scope: a browser for `test` (openspec-flow#17), subagents for the proposal and review steps, parallel worktrees. ## Risk Background subagents are experimental in opencode 1.18.30. The change starts with a spike on the real server; if they are unreliable, delegation runs in the foreground and the rest of the change still applies.
Author
Owner

Shipped as the OpenSpec change apply-delegates-to-subagents, archived as openspec/changes/archive/2026-09-14-apply-delegates-to-subagents/; agent-delegation is now a main spec and agent-transcript gained four requirements. Commits 412cea4, dc7238f, 2232728, fbd0c1d, 183b4f5.

Runtime

  • Every run's opencode config offers three read-only subagents to implementation steps: reader (files, logs, images), explore (where/who), test (runs the given command, returns PASS/FAIL). No edit, write, patch or delegation for any of them; bash only for test; the general-purpose subagent is hidden. The model declares image input.
  • Background subagents are on (OPENCODE_EXPERIMENTAL_BACKGROUND_SUBAGENTS), with a one-argument plugin in the run home that retypes task.background from "True".
  • The context alarm and meter are per session, so a subagent's reading cannot hand off its primary, and its idle is not the primary's.
  • A budget handoff waits up to 5 minutes for running subagents (their reports go to the session that started them), then stops and names any still running.

Prompts — apply.group, tasks.reconcile and test.unit open with one delegation policy, and are handed their group's code (cited ranges for large files).

Agent view — a delegation is one header: agent, task, background marker, running/finished/stopped, elapsed, its own context, and its report (visible while collapsed). Each subagent's lines and live text sit under its own header; the stage header counts running subagents. A late delta no longer resurrects live text after its part settled.

Verified

  • Spike on opencode 1.18.30: a background explore returned at once, the primary kept working, the result arrived ~39 s later as a synthetic user message.
  • Live server: the task menu lists exactly explore, reader, test; reader told to create a file reported it has only glob/grep/read/skill, and the file did not exist.
  • pytest 1498 passed; vitest 328; Playwright 16 passed (new e2e/delegation.spec.ts covers every header state, live text per subagent, collapse, report, count), three consecutive runs.

Not yet measured: the effect on a real run (delegation calls, primary peak context, wall time vs #11's apply.group[2] at 113k and a handoff). Braid runs are paused until #16, #15, #14, #12 and #9 are done; I will post the numbers here from the run after them.

Shipped as the OpenSpec change `apply-delegates-to-subagents`, archived as `openspec/changes/archive/2026-09-14-apply-delegates-to-subagents/`; `agent-delegation` is now a main spec and `agent-transcript` gained four requirements. Commits 412cea4, dc7238f, 2232728, fbd0c1d, 183b4f5. **Runtime** - Every run's opencode config offers three read-only subagents to implementation steps: `reader` (files, logs, images), `explore` (where/who), `test` (runs the given command, returns PASS/FAIL). No edit, write, patch or delegation for any of them; bash only for `test`; the general-purpose subagent is hidden. The model declares image input. - Background subagents are on (`OPENCODE_EXPERIMENTAL_BACKGROUND_SUBAGENTS`), with a one-argument plugin in the run home that retypes `task.background` from `"True"`. - The context alarm and meter are per session, so a subagent's reading cannot hand off its primary, and its idle is not the primary's. - A budget handoff waits up to 5 minutes for running subagents (their reports go to the session that started them), then stops and names any still running. **Prompts** — `apply.group`, `tasks.reconcile` and `test.unit` open with one delegation policy, and are handed their group's code (cited ranges for large files). **Agent view** — a delegation is one header: agent, task, background marker, running/finished/stopped, elapsed, its own context, and its report (visible while collapsed). Each subagent's lines and live text sit under its own header; the stage header counts running subagents. A late delta no longer resurrects live text after its part settled. **Verified** - Spike on opencode 1.18.30: a background `explore` returned at once, the primary kept working, the result arrived ~39 s later as a synthetic user message. - Live server: the task menu lists exactly explore, reader, test; `reader` told to create a file reported it has only glob/grep/read/skill, and the file did not exist. - pytest 1498 passed; vitest 328; Playwright 16 passed (new `e2e/delegation.spec.ts` covers every header state, live text per subagent, collapse, report, count), three consecutive runs. **Not yet measured:** the effect on a real run (delegation calls, primary peak context, wall time vs #11's `apply.group[2]` at 113k and a handoff). Braid runs are paused until #16, #15, #14, #12 and #9 are done; I will post the numbers here from the run after them.
Author
Owner

Delegation, measured on a full run (openspec-flow#11, 2026-09-14, PR #19)

First end-to-end run with apply-delegates-to-subagents (roster reader/explore/test, background, per-session gauge, 112k alarm / 150k hard). The run went issue → PR in 4h10m, gates answered with defaults.

step attempt wall delegations primary peak per session
apply.group[1] 1 22 min 0 112.7k → handoff → 32k
apply.group[2] 1 30 min 1 (test, bg) 125.9k (handoff turn kept working, see fix 1) — lost to an osfd restart
apply.group[2] 2 27 min 1 (test, bg) 117.8k → handoff → 33k
apply.group[3] 1 8 min 3 72k, no handoff
tasks.reconcile 1 13.5 min 4 86k, no handoff
test.unit 1 15 min 3 85k — orphaned by the watchdog (fix 2)
test.unit 3 63 min 6 145.6k → 113.5k → 60k (2 handoffs)

For comparison, #11's previous run hit the alarm in apply.group[2] at 113k with no subagents at all.

What worked

  • Test runs are delegated: every suite/typecheck/build run in groups 2–3, reconcile and test.unit went to a background test subagent, and the reports landed (PASS · 17 files, 342 tests, the delegation card renders them).
  • Verification-shaped steps stayed small: group 3 at 72k, reconcile at 86k, no handoffs.
  • Handoffs are clean now: a 5–14k character handoff, a fresh session at ~20–30k, the step finishes.

What did not

  • Reading is not delegated. apply.group[1] made 0 delegations and read 13 files itself (fakeosfd.py, format.py, digest.py…) to 112k before writing any code; group 2 did the same with the 1,900-line fakeosfd.py. The policy text is not changing this model's habit of reading whole files. Worth trying next: hand large files over by reference only (outline + line ranges) and say plainly "any file over ~400 lines: ask reader"; or a plugin-level nudge when read targets a large file.
  • test.unit is the heaviest step: the target repo is Braid itself, so it spent most of its context reading Braid's own lane/lint internals to decide how its verdict is computed.

Bugs found and fixed during the run (all on main, with tests)

  1. 525abaa — the handoff prompt was sent to a busy session, so opencode folded it into the turn in flight and it kept the node's tools: the model ran npx playwright test after "Stop working" and grew 108k → 126k. Now the in-flight turn is stopped first; verified live (the next handoff made no tool calls and stayed at 112k).
  2. a890b37 — when test.unit's turn ended and the orchestrator ran the declared suite (~6 min), the watchdog saw an idle session, orphaned the step after 20 s and rolled its tests back. The lane run is now an operation the watchdog honours; verified live (the 6-minute lane completed).
  3. f0a6142 — during a handoff the primary cannot be paused while its background subagents finish (measured: aborting a session aborts its background subagents too, opencode 1.18.30), and test.unit grew 115k → 145.6k inside that wait, 4.4k from the hard limit. The wait now also ends once the primary is halfway from alarm to hard limit.
  4. 5b17e49 — the transcript cut every answer over 2,000 characters mid-markdown with no "show more" (a per-line cap applied to the whole block).
## Delegation, measured on a full run (openspec-flow#11, 2026-09-14, PR #19) First end-to-end run with `apply-delegates-to-subagents` (roster reader/explore/test, background, per-session gauge, 112k alarm / 150k hard). The run went issue → PR in 4h10m, gates answered with defaults. | step | attempt | wall | delegations | primary peak per session | |---|---|---|---|---| | apply.group[1] | 1 | 22 min | **0** | 112.7k → handoff → 32k | | apply.group[2] | 1 | 30 min | 1 (`test`, bg) | 125.9k (handoff turn kept working, see fix 1) — lost to an osfd restart | | apply.group[2] | 2 | 27 min | 1 (`test`, bg) | 117.8k → handoff → 33k | | apply.group[3] | 1 | 8 min | 3 | 72k, no handoff | | tasks.reconcile | 1 | 13.5 min | 4 | 86k, no handoff | | test.unit | 1 | 15 min | 3 | 85k — orphaned by the watchdog (fix 2) | | test.unit | 3 | 63 min | 6 | 145.6k → 113.5k → 60k (2 handoffs) | For comparison, #11's previous run hit the alarm in `apply.group[2]` at 113k with no subagents at all. **What worked** - **Test runs are delegated**: every suite/typecheck/build run in groups 2–3, reconcile and test.unit went to a background `test` subagent, and the reports landed (`PASS · 17 files, 342 tests`, the delegation card renders them). - Verification-shaped steps stayed small: group 3 at 72k, reconcile at 86k, no handoffs. - Handoffs are clean now: a 5–14k character handoff, a fresh session at ~20–30k, the step finishes. **What did not** - **Reading is not delegated.** `apply.group[1]` made 0 delegations and read 13 files itself (`fakeosfd.py`, `format.py`, `digest.py`…) to 112k before writing any code; group 2 did the same with the 1,900-line `fakeosfd.py`. The policy text is not changing this model's habit of reading whole files. Worth trying next: hand large files over by reference only (outline + line ranges) and say plainly "any file over ~400 lines: ask `reader`"; or a plugin-level nudge when `read` targets a large file. - `test.unit` is the heaviest step: the target repo is Braid itself, so it spent most of its context reading Braid's own lane/lint internals to decide how its verdict is computed. **Bugs found and fixed during the run** (all on `main`, with tests) 1. `525abaa` — the handoff prompt was sent to a busy session, so opencode folded it into the turn in flight and it kept the node's tools: the model ran `npx playwright test` after "Stop working" and grew 108k → 126k. Now the in-flight turn is stopped first; verified live (the next handoff made no tool calls and stayed at 112k). 2. `a890b37` — when test.unit's turn ended and the orchestrator ran the declared suite (~6 min), the watchdog saw an idle session, orphaned the step after 20 s and rolled its tests back. The lane run is now an operation the watchdog honours; verified live (the 6-minute lane completed). 3. `f0a6142` — during a handoff the primary cannot be paused while its background subagents finish (**measured: aborting a session aborts its background subagents too**, opencode 1.18.30), and `test.unit` grew 115k → 145.6k inside that wait, 4.4k from the hard limit. The wait now also ends once the primary is halfway from alarm to hard limit. 4. `5b17e49` — the transcript cut every answer over 2,000 characters mid-markdown with no "show more" (a per-line cap applied to the whole block).
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
cmoriarty/braid#18
No description provided.