After an output-limit cut-off, the rest of the step runs without thinking #136
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Run 40 was the #124 measurement run on cmoriarty/scratch#6. Its
openspec.apply[1]took 82 min, where run 39's took 15:medium) and hit the 32,000-token output limit with nothing written (finish: length)._cut_off(#101,src/osf/engine/executors.py) re-prompted once with "keep any reasoning brief", using thenothinkvariant and the node's own tools. The code means one turn ("One more turn — told it was cut off, without thinking"). But opencode keeps a prompt's variant for every model call until the agent stops, so thinking stayed off for the other 290 turns (71 min), with 0 reasoning tokens on every one.In the runs collected, there were 5 cut-offs. The 4 in run 22, under the old 20k limit, left tails without thinking of 6–37 turns (2–9 min).
What to change
chat_template_kwargs: {reasoning_effort: "low"}instead ofenable_thinking: false. #124 found that Qwen3.8's template turnslowinto "Keep your thinking brief and focused, moving directly to the conclusion". The rest of the step then keeps thinking, briefly.thinking_token_budget, which is in vLLM 0.28.0's request schema though its enforcement is unverified. For example, 24k against the 32k limit. This is the backstop #124 deferred.last_callprompt near the time limit. It sends the same no-think variant with tools. It is meant for the last minutes, so its tail is short, but check it.To verify, run an e2e against the fake model (fakeinfer) in which the first turn ends with
finish: length. The requests after the recovery prompt should carryreasoning_effort: lowand noenable_thinking: false.Related: #101, #124, #60.
Shipped in
b1baa44andf9fb87d. Deployed in deploy 97 (76eaaae, 01:47 on 2026-10-03). Archived in3b488ccasbrief-thinking-after-cut-off.briefvariant (chat_template_kwargs.reasoning_effort: low) instead ofnothink. opencode keeps a prompt's variant until the agent stops, so the rest of the step goes on thinking, but briefly. The handoff turn and the last call before the deadline still usenothink.thinking_token_budget, set to the output limit less 4,096 with a floor of 1,024. That is 27,904 at 32,000, so a single thought ends with room left to answer. vLLM 0.28.0 enforces it: a probe with a budget of 64 reasoned for 63 tokens and still answered. Other backends get no budget, because it was only probed on vLLM.briefandnothinkand carriesreasoning_effort: mediumandthinking_token_budget: 27904.brief.Not yet seen on production: a real cut-off going through the new recovery. The next scratch run will show one if it happens. Run 40's apply[1] is the case to compare against.