Implementation steps delegate reading and testing to background subagents #18
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Braid's implementation steps (
apply.group,tasks.reconcile,test.unit) do everything in one agent context: read whole files, run test suites and absorb their output, then hit the 112k-token alarm and are handed off to a fresh session that starts over.On openspec-flow#11's run,
apply.group[2](verification) read 152k characters of files — one 56 KB fixture, whole — plus 32k characters of suite output, and was handed off at 113k tokens. Braid exposes opencode'stasktool to these steps but defines no subagents, so neither group made a single delegation call.The operator's own opencode setup works the other way: the primary gets as much context as it can use, and read-only subagents run in the background to read, look at images and run tests, returning short verdicts.
Proposal
Planned as the OpenSpec change
apply-delegates-to-subagents(proposal, specs, design and tasks written, not yet applied):reader(digest files, logs, images),explore(where things are, who uses them),test(run the given commands, return a verdict). None can edit or delegate; opencode's general-purpose subagent is not offered.task'sbackgroundargument a boolean (the model sends the string "True", which fails validation).readercan look at screenshots.Out of scope: a browser for
test(openspec-flow#17), subagents for the proposal and review steps, parallel worktrees.Risk
Background subagents are experimental in opencode 1.18.30. The change starts with a spike on the real server; if they are unreliable, delegation runs in the foreground and the rest of the change still applies.
Shipped as the OpenSpec change
apply-delegates-to-subagents, archived asopenspec/changes/archive/2026-09-14-apply-delegates-to-subagents/;agent-delegationis now a main spec andagent-transcriptgained four requirements. Commits412cea4,dc7238f,2232728,fbd0c1d,183b4f5.Runtime
reader(files, logs, images),explore(where/who),test(runs the given command, returns PASS/FAIL). No edit, write, patch or delegation for any of them; bash only fortest; the general-purpose subagent is hidden. The model declares image input.OPENCODE_EXPERIMENTAL_BACKGROUND_SUBAGENTS), with a one-argument plugin in the run home that retypestask.backgroundfrom"True".Prompts —
apply.group,tasks.reconcileandtest.unitopen with one delegation policy, and are handed their group's code (cited ranges for large files).Agent view — a delegation is one header: agent, task, background marker, running/finished/stopped, elapsed, its own context, and its report (visible while collapsed). Each subagent's lines and live text sit under its own header; the stage header counts running subagents. A late delta no longer resurrects live text after its part settled.
Verified
explorereturned at once, the primary kept working, the result arrived ~39 s later as a synthetic user message.readertold to create a file reported it has only glob/grep/read/skill, and the file did not exist.e2e/delegation.spec.tscovers every header state, live text per subagent, collapse, report, count), three consecutive runs.Not yet measured: the effect on a real run (delegation calls, primary peak context, wall time vs #11's
apply.group[2]at 113k and a handoff). Braid runs are paused until #16, #15, #14, #12 and #9 are done; I will post the numbers here from the run after them.Delegation, measured on a full run (openspec-flow#11, 2026-09-14, PR #19)
First end-to-end run with
apply-delegates-to-subagents(roster reader/explore/test, background, per-session gauge, 112k alarm / 150k hard). The run went issue → PR in 4h10m, gates answered with defaults.test, bg)test, bg)For comparison, #11's previous run hit the alarm in
apply.group[2]at 113k with no subagents at all.What worked
testsubagent, and the reports landed (PASS · 17 files, 342 tests, the delegation card renders them).What did not
apply.group[1]made 0 delegations and read 13 files itself (fakeosfd.py,format.py,digest.py…) to 112k before writing any code; group 2 did the same with the 1,900-linefakeosfd.py. The policy text is not changing this model's habit of reading whole files. Worth trying next: hand large files over by reference only (outline + line ranges) and say plainly "any file over ~400 lines: askreader"; or a plugin-level nudge whenreadtargets a large file.test.unitis the heaviest step: the target repo is Braid itself, so it spent most of its context reading Braid's own lane/lint internals to decide how its verdict is computed.Bugs found and fixed during the run (all on
main, with tests)525abaa— the handoff prompt was sent to a busy session, so opencode folded it into the turn in flight and it kept the node's tools: the model rannpx playwright testafter "Stop working" and grew 108k → 126k. Now the in-flight turn is stopped first; verified live (the next handoff made no tool calls and stayed at 112k).a890b37— when test.unit's turn ended and the orchestrator ran the declared suite (~6 min), the watchdog saw an idle session, orphaned the step after 20 s and rolled its tests back. The lane run is now an operation the watchdog honours; verified live (the 6-minute lane completed).f0a6142— during a handoff the primary cannot be paused while its background subagents finish (measured: aborting a session aborts its background subagents too, opencode 1.18.30), andtest.unitgrew 115k → 145.6k inside that wait, 4.4k from the hard limit. The wait now also ends once the primary is halfway from alarm to hard limit.5b17e49— the transcript cut every answer over 2,000 characters mid-markdown with no "show more" (a per-line cap applied to the whole block).