Agents spend most of their time thinking: cap the reasoning per turn, or switch it off for mechanical steps #124

Closed
opened 2026-10-02 15:49:55 -04:00 by cmoriarty · 2 comments
Owner

Split off from #121 (finding 1 of the scratch-run analysis, docs/design/measurements/2026-10-02-scratch-run-efficiency.md).

Inside agent steps, 79% of the time is the model generating, and 79% of what it generates is reasoning: 2.17M reasoning tokens against 0.59M visible ones over the eight finished scratch runs. Reasoning took 918 min. 129 single thoughts of 2 min or more make up 56% of it, and the longest ran 15.8 min (114k characters, run 36 agent.review[test]). They are not loops. They are exhaustive deliberation, such as a reviewer walking candidate issues A to Z. Reviewers spend 91% of their output on reasoning.

Capping each thought at about 2 min (about 5k tokens) would have cut about 250 min of reasoning from the six feature runs, 91 min of it on steps that run one after another.

Options:

  • No thinking for mechanical work: agent.summarize, agent.reconcile-tasks, the test subagent. Braid already has a no-thinking variant for the handoff turn (#60).
  • A thinking budget per turn for the rest, set per node.

Needs a measurement before it ships: the same brief run with and without the cap, comparing wall clock and the quality of the result (review findings, failing tests, redos).

Split off from #121 (finding 1 of the scratch-run analysis, `docs/design/measurements/2026-10-02-scratch-run-efficiency.md`). Inside agent steps, 79% of the time is the model generating, and 79% of what it generates is reasoning: 2.17M reasoning tokens against 0.59M visible ones over the eight finished scratch runs. Reasoning took 918 min. 129 single thoughts of 2 min or more make up 56% of it, and the longest ran 15.8 min (114k characters, run 36 `agent.review[test]`). They are not loops. They are exhaustive deliberation, such as a reviewer walking candidate issues A to Z. Reviewers spend 91% of their output on reasoning. Capping each thought at about 2 min (about 5k tokens) would have cut about 250 min of reasoning from the six feature runs, 91 min of it on steps that run one after another. Options: - No thinking for mechanical work: `agent.summarize`, `agent.reconcile-tasks`, the `test` subagent. Braid already has a no-thinking variant for the handoff turn (#60). - A thinking budget per turn for the rest, set per node. Needs a measurement before it ships: the same brief run with and without the cap, comparing wall clock and the quality of the result (review findings, failing tests, redos).
Author
Owner

Measured: run 40, at reasoning effort medium, against run 39, at the template's xhigh. Both ran the same brief (cmoriarty/scratch#6) with #121's fixes. Run 40 was stopped during its second apply step, so this covers planning, review and the first apply step.

Step Run 39 (xhigh): wall / reasoning Run 40 (medium)
From start to the first apply step 72 min 27 min
openspec.propose.spec ×3 28 / 22 min 7 / 3 min
openspec.propose.tasks 20 / 15 min 2 / 1 min
openspec.propose.design 11 / 8 min not run
agent.review ×4, round 1 50 / 42 min 40 / 34 min
openspec.revise and agent.review.r2 ×4 39 min not run: the loop ended after round 1
openspec.apply[1] 15 min 82 min (see below)

apply[1] doesn't measure medium. It measures #136. The first turn thought for 11 min (31,950 tokens) and hit the 32k output limit. The cut-off recovery then switched thinking off for the other 290 turns, and the agent brute-forced its test fixtures: 82 inline scripts and 54 test runs. apply[2] reasoned briefly on every turn, as medium intends.

Recommendation:

  • Keep medium as the default. Planning reached apply 45 min sooner, and the review loop needed one round instead of two.
  • Fix #136 before judging apply. A first turn can still think up to the output limit at medium, and the recovery must not switch thinking off for the rest of the step.
  • Then measure apply again on the same brief, and consider levels per step. Reviews still spend 85% of their time reasoning, so low there is the next candidate.

This is one pair of runs, so treat the size of the effect as approximate.

Measured: run 40, at reasoning effort `medium`, against run 39, at the template's xhigh. Both ran the same brief (cmoriarty/scratch#6) with #121's fixes. Run 40 was stopped during its second apply step, so this covers planning, review and the first apply step. | Step | Run 39 (xhigh): wall / reasoning | Run 40 (medium) | |---|---|---| | From start to the first apply step | 72 min | 27 min | | `openspec.propose.spec` ×3 | 28 / 22 min | 7 / 3 min | | `openspec.propose.tasks` | 20 / 15 min | 2 / 1 min | | `openspec.propose.design` | 11 / 8 min | not run | | `agent.review` ×4, round 1 | 50 / 42 min | 40 / 34 min | | `openspec.revise` and `agent.review.r2` ×4 | 39 min | not run: the loop ended after round 1 | | `openspec.apply[1]` | 15 min | 82 min (see below) | **`apply[1]` doesn't measure `medium`. It measures #136.** The first turn thought for 11 min (31,950 tokens) and hit the 32k output limit. The cut-off recovery then switched thinking off for the other 290 turns, and the agent brute-forced its test fixtures: 82 inline scripts and 54 test runs. `apply[2]` reasoned briefly on every turn, as `medium` intends. **Recommendation:** - **Keep `medium` as the default.** Planning reached apply 45 min sooner, and the review loop needed one round instead of two. - **Fix #136 before judging apply.** A first turn can still think up to the output limit at `medium`, and the recovery must not switch thinking off for the rest of the step. - **Then measure apply again on the same brief**, and consider levels per step. Reviews still spend 85% of their time reasoning, so `low` there is the next candidate. This is one pair of runs, so treat the size of the effect as approximate.
Author
Owner

Shipped in 5ce3bc4 (proposal 8c06369) and deployed on 2026-10-02 at 22:34 (2ec079a). Archived as openspec/changes/archive/2026-10-03-reasoning-effort-setting.

What changed

  • Settings → Reasoning effort has three levels:
    • high sends nothing, which keeps the template's own xhigh line;
    • medium, the default, leaves that line out;
    • low asks for brief thinking.
  • Every request gets it. It is sent as the run model's chat_template_kwargs, so it reaches every request of a run, subagents and the planning backend included, on vLLM, SGLang and llama.cpp. The handoff's no-think variant still merges in beside it.
  • A run keeps the effort it started with, and runs from before the setting keep high.

Measured (comment above): at medium, planning reached apply 45 min sooner (27 against 72 min), and the review loop needed one round instead of two.

Follow-ups

  • #136: the output-limit recovery switched thinking off for the rest of apply[1], and a thought should be capped below the output limit (thinking_token_budget).
  • Then levels per step (low for reviews, which still spend 85% of their time reasoning), once apply has been measured again after #136.
Shipped in 5ce3bc4 (proposal 8c06369) and deployed on 2026-10-02 at 22:34 (2ec079a). Archived as `openspec/changes/archive/2026-10-03-reasoning-effort-setting`. **What changed** - **Settings → Reasoning effort has three levels:** - `high` sends nothing, which keeps the template's own xhigh line; - `medium`, the default, leaves that line out; - `low` asks for brief thinking. - **Every request gets it.** It is sent as the run model's `chat_template_kwargs`, so it reaches every request of a run, subagents and the planning backend included, on vLLM, SGLang and llama.cpp. The handoff's no-think variant still merges in beside it. - **A run keeps the effort it started with,** and runs from before the setting keep `high`. **Measured** (comment above): at `medium`, planning reached apply 45 min sooner (27 against 72 min), and the review loop needed one round instead of two. **Follow-ups** - #136: the output-limit recovery switched thinking off for the rest of `apply[1]`, and a thought should be capped below the output limit (`thinking_token_budget`). - Then levels per step (`low` for reviews, which still spend 85% of their time reasoning), once apply has been measured again after #136.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
cmoriarty/braid#124
No description provided.