A single-agent OpenSpec pipeline: explore, propose, apply and archive in one forked session, for a model that serves one session at a time #151

Open
opened 2026-10-06 01:09:08 -04:00 by cmoriarty · 3 comments
Owner

Why

I want to try a bigger model, hosted with llama.cpp, that can serve only one session at a time. Braid's OpenSpec pipelines spread a change over many sessions:

  • the proposal is split across one session per artifact;
  • apply is grouped by a token budget and fanned out;
  • four review lenses read the proposal in parallel.

Most of that suits a fast model with many parallel slots. Used by hand, OpenSpec is four steps, often in one session: explore, propose, apply, archive.

Proposal: a new experimental pipeline (working name solo, with git-flow-solo)

  • Steps: explore, then propose in one step (proposal, specs, design and tasks), then strict validation with a repair if needed, then the proposal gate, then one apply step, then the fix loop, then the summary, the pull request, the merge gate, the merge and the archive.
  • Not included: no per-artifact propose steps, no budget estimate and no grouped or fanned-out apply, no review lenses or review rounds, no code review, no agent test-writing step, and no archive gate.
  • One session, carried forward. Every agent step forks the session of the agent step before it (fork_session_of, #139), so the change's context carries from explore to the PR description. When the session reaches the context alarm, the existing handoff hands it to a fresh session, and the next step forks that one. Compaction stays the backstop.
  • The fix loop keeps the engine's lanes, the judge and at most two fix rounds, all forked.
  • Subagents are a choice made when starting the run. With an engine that serves one session, turn them off, so only the main session ever asks the model. With a server that runs several sessions, keep them on to save the main session's context.
    • With them off, there is no browser subagent, so the judge is not due and the loop is the lanes and the fix rounds.

To decide in the proposal

  • Whether the judge forks the implementer's session like every other step, or keeps a session of its own. It was made independent because the agent that wrote the code graded it in run 42.
  • What the proposal gate's "Request changes" does when there is no review round to send it to.
**Why** I want to try a bigger model, hosted with llama.cpp, that can serve only one session at a time. Braid's OpenSpec pipelines spread a change over many sessions: - the proposal is split across one session per artifact; - apply is grouped by a token budget and fanned out; - four review lenses read the proposal in parallel. Most of that suits a fast model with many parallel slots. Used by hand, OpenSpec is four steps, often in one session: explore, propose, apply, archive. **Proposal: a new experimental pipeline (working name `solo`, with `git-flow-solo`)** - **Steps:** explore, then propose in one step (proposal, specs, design and tasks), then strict validation with a repair if needed, then the proposal gate, then one apply step, then the fix loop, then the summary, the pull request, the merge gate, the merge and the archive. - **Not included:** no per-artifact propose steps, no budget estimate and no grouped or fanned-out apply, no review lenses or review rounds, no code review, no agent test-writing step, and no archive gate. - **One session, carried forward.** Every agent step forks the session of the agent step before it (`fork_session_of`, #139), so the change's context carries from explore to the PR description. When the session reaches the context alarm, the existing handoff hands it to a fresh session, and the next step forks that one. Compaction stays the backstop. - **The fix loop** keeps the engine's lanes, the judge and at most two fix rounds, all forked. - **Subagents are a choice made when starting the run.** With an engine that serves one session, turn them off, so only the main session ever asks the model. With a server that runs several sessions, keep them on to save the main session's context. - With them off, there is no `browser` subagent, so the judge is not due and the loop is the lanes and the fix rounds. **To decide in the proposal** - Whether the judge forks the implementer's session like every other step, or keeps a session of its own. It was made independent because the agent that wrote the code graded it in run 42. - What the proposal gate's "Request changes" does when there is no review round to send it to.
Author
Owner

Shipped in fe0b60a, live, archived as solo-pipeline.

What there is

  • solo and git-flow-solo, experimental built-ins:
    • explore;
    • one propose step that writes the proposal, specs, design and tasks;
    • strict validation, and a repair only if it fails;
    • the proposal gate;
    • one apply step for the whole task list, then a reconcile step only if a task is left unticked;
    • the fix loop;
    • the summary, the pull request, the merge gate, the merge and the archive.
  • Not included: no per-artifact propose steps, no budget step or fan-out, no review lenses, no code review, no agent test-writing step, and no archive gate.
  • One session carried through. Every agent step but explore and the judges forks the session of the step before it, newest first. The alarm's handoff still applies, and the next step forks the fresh session it made.
  • The judges keep their own session. The agent that wrote the code doesn't grade it (run 42). They still run one at a time.
  • "Request changes" at the proposal gate runs a revise step in a fork, with the person's words, and asks the gate again, at most twice.
  • Subagents are chosen in New Run, for solo pipelines only. The choice starts from Off: the run gets no subagents, so only the main session asks the model. With no browser subagent no judge is due, so the fix loop is the lanes and the fix rounds. On uses the roster in Settings.

Verified

  • Two new simulated runs: solo (subagents off, checking each step forks the one before) and solo-request-changes. The deploy's self-check ran both.
  • A kept solo run's opencode config has only the main agent.
  • Production's pipeline list offers both.

Not yet: no run on a real model. For llama.cpp with one slot, also set the concurrent runs in Settings to 1, and the context per agent to the model's window.

Shipped in fe0b60a, live, archived as `solo-pipeline`. **What there is** - **`solo` and `git-flow-solo`, experimental built-ins:** - explore; - one propose step that writes the proposal, specs, design and tasks; - strict validation, and a repair only if it fails; - the proposal gate; - one apply step for the whole task list, then a reconcile step only if a task is left unticked; - the fix loop; - the summary, the pull request, the merge gate, the merge and the archive. - **Not included:** no per-artifact propose steps, no budget step or fan-out, no review lenses, no code review, no agent test-writing step, and no archive gate. - **One session carried through.** Every agent step but explore and the judges forks the session of the step before it, newest first. The alarm's handoff still applies, and the next step forks the fresh session it made. - **The judges keep their own session.** The agent that wrote the code doesn't grade it (run 42). They still run one at a time. - **"Request changes" at the proposal gate** runs a revise step in a fork, with the person's words, and asks the gate again, at most twice. - **Subagents are chosen in New Run, for solo pipelines only.** The choice starts from Off: the run gets no subagents, so only the main session asks the model. With no `browser` subagent no judge is due, so the fix loop is the lanes and the fix rounds. On uses the roster in Settings. **Verified** - Two new simulated runs: `solo` (subagents off, checking each step forks the one before) and `solo-request-changes`. The deploy's self-check ran both. - A kept `solo` run's opencode config has only the main agent. - Production's pipeline list offers both. **Not yet:** no run on a real model. For llama.cpp with one slot, also set the concurrent runs in Settings to 1, and the context per agent to the model's window.
Author
Owner

First production run on solo (run 50, run_01M47X2Y9F5VYVGMFTHYT1EKHT)

The setup:

  • the same task as runs 41 and 48: run 48's brief, from scratch-139 at 88be018, on a solo-test branch so develop keeps the #139 baseline's start;
  • subagents off, as for a one-slot engine;
  • the model on trogdor.

Result: succeeded in 73 minutes, merged and archived. Every step took one attempt: no failure, no retry, no question.

  • explore 2.5 min, then propose 11.4 min (the whole change in one step; it validated first time, so no repair).
  • The proposal gate was auto-approved, and no revise was needed.
  • apply 49.6 min, one step for the whole task list.
  • The tests lane passed on the first round: 110 backend tests and 42 frontend tests, run with the project's own backend/.venv.
  • No judge was due, because with subagents off there is no browser subagent. So no fix rounds, and the result was not opened in a browser.
  • summary 7.6 min.

The carried session

  • propose forked explore (14 messages inherited), apply forked propose (39), and the summary forked apply (132).
  • No step reached the context alarm, so the whole run stayed in one chain with no handoff.

Against the baseline (#139)

run 41 (git-flow) run 48 (git-flow) run 50 (solo)
wall 315 min, merged 322.5 min, failed 73 min, merged
model turns 633 544 131
generated tokens 743,218 660,994 121,810
agent steps / attempts 19 / 20 18 / 19 4 / 4
subagents 17 6 0

Wall clock is across the vLLM 0.31 upgrade, so turns and tokens are the fairer comparison. Run 50 also skipped the judge, the review lenses and the code review by design. Whether five people can play a whole game over the network was checked only by the tests, not in a browser.

Worth a look: the summary step's 7.6 minutes for 2 turns. It is most likely the 132 inherited messages being read again.

**First production run on `solo`** (run 50, `run_01M47X2Y9F5VYVGMFTHYT1EKHT`) The setup: - the same task as runs 41 and 48: run 48's brief, from scratch-139 at 88be018, on a `solo-test` branch so `develop` keeps the #139 baseline's start; - subagents off, as for a one-slot engine; - the model on trogdor. **Result:** succeeded in 73 minutes, merged and archived. Every step took one attempt: no failure, no retry, no question. - explore 2.5 min, then propose 11.4 min (the whole change in one step; it validated first time, so no repair). - The proposal gate was auto-approved, and no revise was needed. - apply 49.6 min, one step for the whole task list. - The tests lane passed on the first round: 110 backend tests and 42 frontend tests, run with the project's own `backend/.venv`. - No judge was due, because with subagents off there is no `browser` subagent. So no fix rounds, and the result was not opened in a browser. - summary 7.6 min. **The carried session** - propose forked explore (14 messages inherited), apply forked propose (39), and the summary forked apply (132). - No step reached the context alarm, so the whole run stayed in one chain with no handoff. **Against the baseline (#139)** | | run 41 (`git-flow`) | run 48 (`git-flow`) | run 50 (`solo`) | | --- | --- | --- | --- | | wall | 315 min, merged | 322.5 min, failed | 73 min, merged | | model turns | 633 | 544 | 131 | | generated tokens | 743,218 | 660,994 | 121,810 | | agent steps / attempts | 19 / 20 | 18 / 19 | 4 / 4 | | subagents | 17 | 6 | 0 | Wall clock is across the vLLM 0.31 upgrade, so turns and tokens are the fairer comparison. Run 50 also skipped the judge, the review lenses and the code review by design. Whether five people can play a whole game over the network was checked only by the tests, not in a browser. **Worth a look:** the summary step's 7.6 minutes for 2 turns. It is most likely the 132 inherited messages being read again.
Author
Owner

Follow-up: the browser check should be mandatory in solo, and the loose ends to pick up later. Reopened so these stay visible. Trogdor is going off for now.

1. Make the browser check mandatory

Run 50 merged with nothing opening the game in a browser. With subagents off, no browser subagent exists, so fixloop.judge_due finds no judge due. "Five play a whole game over the network" was checked only by the 110 backend and 42 frontend tests.

  • The judge, or a browser check, should run in every solo run that changes something a browser shows, whatever the subagents choice.
  • One way: give the judge's own session the Playwright MCP directly when subagents are off, rather than through the browser subagent. The judge already has a session of its own and runs after apply, so it still asks the model one session at a time.
  • Whatever the judge finds then goes to the forked fix rounds as it does today.
  • The New Run text for Off ("no judge opens the app in a browser") changes with it.

2. Not yet exercised on production

  • A real single-slot model. The llama.cpp setup needs:
    • concurrent runs set to 1 in Settings;
    • context per agent set to the model's window.
  • The handoff inside a forked chain. Run 50 never reached the context alarm. A model with a smaller window is the first that will hand a forked step to a fresh session, and the next step should fork that one.
  • "Request changes" at the proposal gate. This is verified only in the simulated run solo-request-changes.

3. Worth a look

  • The summary step's time. agent.summarize took 7.6 minutes for 2 turns. It forked apply's session and inherited 132 messages, most likely read again in full. On a slower one-slot model that cost grows. Two ways to go:
    • give the summary a fresh session, as it had before solo;
    • check whether the prefix cache is being missed.

4. Open decisions from tonight's issues

  • #147: should waiting on an agent's question pause the step's time limit?
  • #150: should a judge's time be bounded, or should it drive less of a multi-player app? Run 48's judge spent 190 minutes.
  • #150, known gap: a judge that fails the deliverable and also leaves a stray file goes on to a fix round, where the file can land in the code. Only the brief's $TMPDIR rule prevents it.

5. State to pick up from

  • Production: fab99fb, with #147, #148, #150, #151 and #153 live.
  • #139 baseline: still needed. Run 49 (git-flow) was stopped during apply.
    • scratch-139's develop is still at 88be018, the start of runs 41 and 48.
    • Run 50's merged work is on the separate branch solo-test.
  • Run 50 (run_01M47X2Y9F5VYVGMFTHYT1EKHT): solo, subagents off. It succeeded in 73 minutes, every step on its first attempt, with 131 turns and 121,810 generated tokens. Runs 48 and 41 took 544 and 633 turns.
  • Older runs: steps skipped before migration 10 show no reason until an osfd rebuild.
  • Flaky browser specs: two seen under machine load, both passing alone. ui-polish-103 "another step opens at its own tail" failed at load 10, and resting-pointer failed at load 12.
**Follow-up: the browser check should be mandatory in `solo`, and the loose ends to pick up later.** Reopened so these stay visible. Trogdor is going off for now. ### 1. Make the browser check mandatory Run 50 merged with nothing opening the game in a browser. With subagents off, no `browser` subagent exists, so `fixloop.judge_due` finds no judge due. "Five play a whole game over the network" was checked only by the 110 backend and 42 frontend tests. - The judge, or a browser check, should run in every `solo` run that changes something a browser shows, whatever the subagents choice. - **One way:** give the judge's own session the Playwright MCP directly when subagents are off, rather than through the `browser` subagent. The judge already has a session of its own and runs after apply, so it still asks the model one session at a time. - Whatever the judge finds then goes to the forked fix rounds as it does today. - The New Run text for Off ("no judge opens the app in a browser") changes with it. ### 2. Not yet exercised on production - **A real single-slot model.** The llama.cpp setup needs: - concurrent runs set to 1 in Settings; - context per agent set to the model's window. - **The handoff inside a forked chain.** Run 50 never reached the context alarm. A model with a smaller window is the first that will hand a forked step to a fresh session, and the next step should fork that one. - **"Request changes" at the proposal gate.** This is verified only in the simulated run `solo-request-changes`. ### 3. Worth a look - **The summary step's time.** `agent.summarize` took 7.6 minutes for 2 turns. It forked apply's session and inherited 132 messages, most likely read again in full. On a slower one-slot model that cost grows. Two ways to go: - give the summary a fresh session, as it had before `solo`; - check whether the prefix cache is being missed. ### 4. Open decisions from tonight's issues - **#147:** should waiting on an agent's question pause the step's time limit? - **#150:** should a judge's time be bounded, or should it drive less of a multi-player app? Run 48's judge spent 190 minutes. - **#150, known gap:** a judge that fails the deliverable and also leaves a stray file goes on to a fix round, where the file can land in the code. Only the brief's `$TMPDIR` rule prevents it. ### 5. State to pick up from - **Production:** fab99fb, with #147, #148, #150, #151 and #153 live. - **#139 baseline:** still needed. Run 49 (`git-flow`) was stopped during apply. - scratch-139's `develop` is still at 88be018, the start of runs 41 and 48. - Run 50's merged work is on the separate branch `solo-test`. - **Run 50 (`run_01M47X2Y9F5VYVGMFTHYT1EKHT`):** `solo`, subagents off. It succeeded in 73 minutes, every step on its first attempt, with 131 turns and 121,810 generated tokens. Runs 48 and 41 took 544 and 633 turns. - **Older runs:** steps skipped before migration 10 show no reason until an `osfd rebuild`. - **Flaky browser specs:** two seen under machine load, both passing alone. `ui-polish-103` "another step opens at its own tail" failed at load 10, and `resting-pointer` failed at load 12.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
cmoriarty/braid#151
No description provided.