After an output-limit cut-off, the rest of the step runs without thinking #136

Closed
opened 2026-10-03 00:47:04 -04:00 by cmoriarty · 1 comment
Owner

Run 40 was the #124 measurement run on cmoriarty/scratch#6. Its openspec.apply[1] took 82 min, where run 39's took 15:

  1. The first turn was cut off. It reasoned for 11 min (31,950 reasoning tokens at reasoning effort medium) and hit the 32,000-token output limit with nothing written (finish: length).
  2. The recovery turned thinking off for good. _cut_off (#101, src/osf/engine/executors.py) re-prompted once with "keep any reasoning brief", using the nothink variant and the node's own tools. The code means one turn ("One more turn — told it was cut off, without thinking"). But opencode keeps a prompt's variant for every model call until the agent stops, so thinking stayed off for the other 290 turns (71 min), with 0 reasoning tokens on every one.
  3. Without thinking, the agent brute-forced. It wrote tests that need exact outcomes (winners, ties), then searched for card deals that produce them by trial and error:
    • 82 throwaway inline Python scripts;
    • 54 test runs: all 6 failing for the first 30 min, all passing only at minute 79;
    • 94 edits;
    • 124k output tokens, against 10k in run 39.
  4. Every turn got slower. Context averaged 171k and peaked at 297k (run 39: 81k and 101k), and waiting for first tokens took 15 min.

In the runs collected, there were 5 cut-offs. The 4 in run 22, under the old 20k limit, left tails without thinking of 6–37 turns (2–9 min).

What to change

  • Keep thinking on after a cut-off, but brief. Send the cut-off recovery with a variant of chat_template_kwargs: {reasoning_effort: "low"} instead of enable_thinking: false. #124 found that Qwen3.8's template turns low into "Keep your thinking brief and focused, moving directly to the conclusion". The rest of the step then keeps thinking, briefly.
  • Prevent the cut-off. Cap a single thought below the output limit with vLLM's thinking_token_budget, which is in vLLM 0.28.0's request schema though its enforcement is unverified. For example, 24k against the 32k limit. This is the backstop #124 deferred.
  • Check the last_call prompt near the time limit. It sends the same no-think variant with tools. It is meant for the last minutes, so its tail is short, but check it.

To verify, run an e2e against the fake model (fakeinfer) in which the first turn ends with finish: length. The requests after the recovery prompt should carry reasoning_effort: low and no enable_thinking: false.

Related: #101, #124, #60.

Run 40 was the #124 measurement run on cmoriarty/scratch#6. Its `openspec.apply[1]` took 82 min, where run 39's took 15: 1. **The first turn was cut off.** It reasoned for 11 min (31,950 reasoning tokens at reasoning effort `medium`) and hit the 32,000-token output limit with nothing written (`finish: length`). 2. **The recovery turned thinking off for good.** `_cut_off` (#101, `src/osf/engine/executors.py`) re-prompted once with "keep any reasoning brief", using the `nothink` variant and the node's own tools. The code means one turn ("One more turn — told it was cut off, without thinking"). But opencode keeps a prompt's variant for every model call until the agent stops, so thinking stayed off for the other 290 turns (71 min), with 0 reasoning tokens on every one. 3. **Without thinking, the agent brute-forced.** It wrote tests that need exact outcomes (winners, ties), then searched for card deals that produce them by trial and error: - 82 throwaway inline Python scripts; - 54 test runs: all 6 failing for the first 30 min, all passing only at minute 79; - 94 edits; - 124k output tokens, against 10k in run 39. 4. **Every turn got slower.** Context averaged 171k and peaked at 297k (run 39: 81k and 101k), and waiting for first tokens took 15 min. In the runs collected, there were 5 cut-offs. The 4 in run 22, under the old 20k limit, left tails without thinking of 6–37 turns (2–9 min). ## What to change - **Keep thinking on after a cut-off, but brief.** Send the cut-off recovery with a variant of `chat_template_kwargs: {reasoning_effort: "low"}` instead of `enable_thinking: false`. #124 found that Qwen3.8's template turns `low` into "Keep your thinking brief and focused, moving directly to the conclusion". The rest of the step then keeps thinking, briefly. - **Prevent the cut-off.** Cap a single thought below the output limit with vLLM's `thinking_token_budget`, which is in vLLM 0.28.0's request schema though its enforcement is unverified. For example, 24k against the 32k limit. This is the backstop #124 deferred. - **Check the `last_call` prompt near the time limit.** It sends the same no-think variant with tools. It is meant for the last minutes, so its tail is short, but check it. To verify, run an e2e against the fake model (fakeinfer) in which the first turn ends with `finish: length`. The requests after the recovery prompt should carry `reasoning_effort: low` and no `enable_thinking: false`. Related: #101, #124, #60.
Author
Owner

Shipped in b1baa44 and f9fb87d. Deployed in deploy 97 (76eaaae, 01:47 on 2026-10-03). Archived in 3b488cc as brief-thinking-after-cut-off.

  • The re-prompt after a cut-off now names a new brief variant (chat_template_kwargs.reasoning_effort: low) instead of nothink. opencode keeps a prompt's variant until the agent stops, so the rest of the step goes on thinking, but briefly. The handoff turn and the last call before the deadline still use nothink.
  • On vLLM, each model's options also carry thinking_token_budget, set to the output limit less 4,096 with a floor of 1,024. That is 27,904 at 32,000, so a single thought ends with room left to answer. vLLM 0.28.0 enforces it: a probe with a budget of 64 reasoned for 63 tokens and still answered. Other backends get no budget, because it was only probed on vLLM.
  • Checked on production after the deploy:
    • A new run's config, rendered from production's settings, declares brief and nothink and carries reasoning_effort: medium and thinking_token_budget: 27904.
    • Two prompts went through opencode 1.18.21 to trogdor and both answered: 31 reasoning tokens with no variant, and 17 with brief.

Not yet seen on production: a real cut-off going through the new recovery. The next scratch run will show one if it happens. Run 40's apply[1] is the case to compare against.

Shipped in b1baa44 and f9fb87d. Deployed in deploy 97 (76eaaae, 01:47 on 2026-10-03). Archived in 3b488cc as `brief-thinking-after-cut-off`. - **The re-prompt after a cut-off** now names a new `brief` variant (`chat_template_kwargs.reasoning_effort: low`) instead of `nothink`. opencode keeps a prompt's variant until the agent stops, so the rest of the step goes on thinking, but briefly. The handoff turn and the last call before the deadline still use `nothink`. - **On vLLM**, each model's options also carry `thinking_token_budget`, set to the output limit less 4,096 with a floor of 1,024. That is 27,904 at 32,000, so a single thought ends with room left to answer. vLLM 0.28.0 enforces it: a probe with a budget of 64 reasoned for 63 tokens and still answered. Other backends get no budget, because it was only probed on vLLM. - **Checked on production** after the deploy: - A new run's config, rendered from production's settings, declares `brief` and `nothink` and carries `reasoning_effort: medium` and `thinking_token_budget: 27904`. - Two prompts went through opencode 1.18.21 to trogdor and both answered: 31 reasoning tokens with no variant, and 17 with `brief`. Not yet seen on production: a real cut-off going through the new recovery. The next scratch run will show one if it happens. Run 40's apply[1] is the case to compare against.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
cmoriarty/braid#136
No description provided.