Agents spend most of their time thinking: cap the reasoning per turn, or switch it off for mechanical steps #124
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Split off from #121 (finding 1 of the scratch-run analysis,
docs/design/measurements/2026-10-02-scratch-run-efficiency.md).Inside agent steps, 79% of the time is the model generating, and 79% of what it generates is reasoning: 2.17M reasoning tokens against 0.59M visible ones over the eight finished scratch runs. Reasoning took 918 min. 129 single thoughts of 2 min or more make up 56% of it, and the longest ran 15.8 min (114k characters, run 36
agent.review[test]). They are not loops. They are exhaustive deliberation, such as a reviewer walking candidate issues A to Z. Reviewers spend 91% of their output on reasoning.Capping each thought at about 2 min (about 5k tokens) would have cut about 250 min of reasoning from the six feature runs, 91 min of it on steps that run one after another.
Options:
agent.summarize,agent.reconcile-tasks, thetestsubagent. Braid already has a no-thinking variant for the handoff turn (#60).Needs a measurement before it ships: the same brief run with and without the cap, comparing wall clock and the quality of the result (review findings, failing tests, redos).
.arrives, and the agent-view spec fails under load #130Measured: run 40, at reasoning effort
medium, against run 39, at the template's xhigh. Both ran the same brief (cmoriarty/scratch#6) with #121's fixes. Run 40 was stopped during its second apply step, so this covers planning, review and the first apply step.openspec.propose.spec×3openspec.propose.tasksopenspec.propose.designagent.review×4, round 1openspec.reviseandagent.review.r2×4openspec.apply[1]apply[1]doesn't measuremedium. It measures #136. The first turn thought for 11 min (31,950 tokens) and hit the 32k output limit. The cut-off recovery then switched thinking off for the other 290 turns, and the agent brute-forced its test fixtures: 82 inline scripts and 54 test runs.apply[2]reasoned briefly on every turn, asmediumintends.Recommendation:
mediumas the default. Planning reached apply 45 min sooner, and the review loop needed one round instead of two.medium, and the recovery must not switch thinking off for the rest of the step.lowthere is the next candidate.This is one pair of runs, so treat the size of the effect as approximate.
Shipped in
5ce3bc4(proposal8c06369) and deployed on 2026-10-02 at 22:34 (2ec079a). Archived asopenspec/changes/archive/2026-10-03-reasoning-effort-setting.What changed
highsends nothing, which keeps the template's own xhigh line;medium, the default, leaves that line out;lowasks for brief thinking.chat_template_kwargs, so it reaches every request of a run, subagents and the planning backend included, on vLLM, SGLang and llama.cpp. The handoff's no-think variant still merges in beside it.high.Measured (comment above): at
medium, planning reached apply 45 min sooner (27 against 72 min), and the review loop needed one round instead of two.Follow-ups
apply[1], and a thought should be capped below the output limit (thinking_token_budget).lowfor reviews, which still spend 85% of their time reasoning), once apply has been measured again after #136..arrives, and the agent-view spec fails under load #130