Handoffs come too early: derive the context ladder from the real ceiling, carry handoffs across attempts, try YaRN #60

Closed
opened 2026-09-26 18:39:10 -04:00 by cmoriarty · 1 comment
Owner

Problem

Handoffs are the biggest slowdown in a run. The context per agent was raised to 262,144, but steps still hand off long before that. On production, run run_01M3FS9FFXEHTY5WTHX36V4Z5Y, step fix.implement (stp_01M3FS9JXBQJN4A21P8HBWMY0S), commit 2064a2f:

  • Attempt 1's session grew steadily from 10k to 196,115 over about 35 minutes. At that point osf.budget.alarm fired (alarm: 195734), the session wrote a 10,051-character handoff, and a fresh session started from it.
  • The respawned session reached 104k, then an interrupt aborted it. The step failed its predicate (no non-empty file matched: .osf/fix.md), and attempt 2 started cold, without the handoff. It re-read the repository and was at 129k within 10 minutes.
  • About 113k of the 196k came from tool output: 52 read calls came to about 91k tokens (two reads under src/osf/pipeline/ were 19k and 10k), and 28 bash calls added about 15k. The rest is many small turns of 100–2,000 tokens each.

The engine behaved as specified. The ladder is the problem.

Findings

  1. The alarm is a fixed 75% of the context per agent. Settings.alarm = H × 112/150 and allowance = H × 144/150 (src/osf/settings.py). Those ratios were set when H was 150k, to leave 38k for a measured worst-case overshoot of 32,435. At H = 262,144 the same ratio leaves 66k unused.
  2. The overshoot the margin was sized for is gone. That 32k was measured before the busy turn was aborted ahead of the handoff prompt. In this run the handoff turn read 190,444, which is below the 196,115 alarm reading. What remains is one turn's growth (recorded p99 11k, max 29k) plus the handoff's own output (3,426 here).
  3. The hard limit H = 262,144 cannot be reached. The run config declares limit.output: 32768 (MODEL_OUTPUT_LIMIT), and opencode asks for up to that much output on each request. vLLM rejects any request where prompt plus max_tokens exceeds --max-model-len 262144, so real input is capped at about 230k. The research agrees: every recorded context drop clusters at 229,751–230,137. So a step that gets past the alarm without a working handoff fails with a provider error at about 230k, and the compaction fallback, which triggers only at peak >= H, never runs. validate() allows agent_context up to model_context; the limit should be model_context − output reservation. To confirm on the wire: the max_tokens opencode actually sends.
  4. The output reservation is larger than it needs to be. Recorded generation is p99 6–9k per request and max 27.8k across 8,686 calls. opencode 1.18.21 also reads OPENCODE_EXPERIMENTAL_OUTPUT_TOKEN_MAX, which is a second way to lower it.
  5. A handoff does not carry across attempts. An interrupt, redo or failed predicate starts the next attempt from nothing, even when a handoff was written minutes earlier. Attempt 2 above threw away 196k of context.
  6. Minor, related to #57: attempt.peak_input_tokens for attempt 1 reads 104,129 (only the respawned session), while budget_ledger reads 196,115.

Explore all five

  1. Base the ladder on the ceiling, not a ratio. ceiling = model_context − output_reservation; H = min(agent_context, ceiling); A = H − margin, with the margin a fixed measured number (about 30k: one worst turn plus the handoff output) rather than 25% of H. Clamp agent_context to the ceiling in settings validation, and show the ceiling on the settings page.

  2. Lower the output reservation from 32,768 to 16,384 (or make it a setting). At 262,144 this moves the ceiling from about 230k to about 246k and the alarm to about 214k. Measure how often a turn reaches the cap.

  3. Carry the handoff across attempts. A new attempt of the same step, after an interrupt, redo, resume or failed predicate, starts from the last handoff written for that step, together with git diff, and does not start cold.

  4. Slow the growth. Find out whether opencode's tool-output pruning (compaction.prune) runs with compaction.auto: false and whether it helps. Try steering agents toward ranged reads, grep and the reader subagent over whole-file reads of large files.

  5. Experiment: raise vLLM's window with YaRN on trogdor. For example, --max-model-len 524288 with a YaRN rope_scaling factor of 2 over the native 262,144. The operator is willing to run this. What to measure:

    • wall time per run and handoffs per run, before and after, on the same pipeline and brief;
    • model mistake rates past 232k, the largest context observed without degradation (EFFECTIVE_WINDOW);
    • KV preemptions and slot throughput: the 971,740-token pool holds six sequences only up to about 162k each, so fewer run concurrently above that;
    • whether YaRN (static scaling in vLLM) hurts quality on short prompts.

    Braid needs model_context and the ladder in (1) to follow the new window with no other code changes.

Acceptance

  • The handoff point for a given setting is derived from the model window, the output reservation and a fixed margin, and the settings page shows all three.
  • agent_context cannot be set above what vLLM will accept.
  • A new attempt of a step that has a handoff starts from it.
  • A before/after measurement of the YaRN experiment is written up under docs/design/measurements/.
## Problem Handoffs are the biggest slowdown in a run. The context per agent was raised to 262,144, but steps still hand off long before that. On production, run `run_01M3FS9FFXEHTY5WTHX36V4Z5Y`, step `fix.implement` (`stp_01M3FS9JXBQJN4A21P8HBWMY0S`), commit `2064a2f`: - Attempt 1's session grew steadily from 10k to **196,115** over about 35 minutes. At that point `osf.budget.alarm` fired (`alarm: 195734`), the session wrote a 10,051-character handoff, and a fresh session started from it. - The respawned session reached 104k, then an interrupt aborted it. The step failed its predicate (`no non-empty file matched: .osf/fix.md`), and **attempt 2 started cold, without the handoff**. It re-read the repository and was at 129k within 10 minutes. - About 113k of the 196k came from tool output: 52 `read` calls came to about 91k tokens (two reads under `src/osf/pipeline/` were 19k and 10k), and 28 `bash` calls added about 15k. The rest is many small turns of 100–2,000 tokens each. The engine behaved as specified. The ladder is the problem. ## Findings 1. **The alarm is a fixed 75% of the context per agent.** `Settings.alarm = H × 112/150` and `allowance = H × 144/150` (`src/osf/settings.py`). Those ratios were set when H was 150k, to leave 38k for a measured worst-case overshoot of 32,435. At H = 262,144 the same ratio leaves 66k unused. 2. **The overshoot the margin was sized for is gone.** That 32k was measured before the busy turn was aborted ahead of the handoff prompt. In this run the handoff turn read 190,444, which is *below* the 196,115 alarm reading. What remains is one turn's growth (recorded p99 11k, max 29k) plus the handoff's own output (3,426 here). 3. **The hard limit H = 262,144 cannot be reached.** The run config declares `limit.output: 32768` (`MODEL_OUTPUT_LIMIT`), and opencode asks for up to that much output on each request. vLLM rejects any request where prompt plus `max_tokens` exceeds `--max-model-len 262144`, so real input is capped at about 230k. The research agrees: every recorded context drop clusters at 229,751–230,137. So a step that gets past the alarm without a working handoff fails with a provider error at about 230k, and the compaction fallback, which triggers only at `peak >= H`, never runs. `validate()` allows `agent_context` up to `model_context`; the limit should be `model_context − output reservation`. *To confirm on the wire: the `max_tokens` opencode actually sends.* 4. **The output reservation is larger than it needs to be.** Recorded generation is p99 6–9k per request and max 27.8k across 8,686 calls. opencode 1.18.21 also reads `OPENCODE_EXPERIMENTAL_OUTPUT_TOKEN_MAX`, which is a second way to lower it. 5. **A handoff does not carry across attempts.** An interrupt, redo or failed predicate starts the next attempt from nothing, even when a handoff was written minutes earlier. Attempt 2 above threw away 196k of context. 6. Minor, related to #57: `attempt.peak_input_tokens` for attempt 1 reads 104,129 (only the respawned session), while `budget_ledger` reads 196,115. ## Explore all five 1. **Base the ladder on the ceiling, not a ratio.** `ceiling = model_context − output_reservation`; `H = min(agent_context, ceiling)`; `A = H − margin`, with the margin a fixed measured number (about 30k: one worst turn plus the handoff output) rather than 25% of H. Clamp `agent_context` to the ceiling in settings validation, and show the ceiling on the settings page. 2. **Lower the output reservation** from 32,768 to 16,384 (or make it a setting). At 262,144 this moves the ceiling from about 230k to about 246k and the alarm to about 214k. Measure how often a turn reaches the cap. 3. **Carry the handoff across attempts.** A new attempt of the same step, after an interrupt, redo, resume or failed predicate, starts from the last handoff written for that step, together with `git diff`, and does not start cold. 4. **Slow the growth.** Find out whether opencode's tool-output pruning (`compaction.prune`) runs with `compaction.auto: false` and whether it helps. Try steering agents toward ranged reads, `grep` and the `reader` subagent over whole-file `read`s of large files. 5. **Experiment: raise vLLM's window with YaRN on trogdor.** For example, `--max-model-len 524288` with a YaRN `rope_scaling` factor of 2 over the native 262,144. The operator is willing to run this. What to measure: - wall time per run and handoffs per run, before and after, on the same pipeline and brief; - model mistake rates past 232k, the largest context observed without degradation (`EFFECTIVE_WINDOW`); - KV preemptions and slot throughput: the 971,740-token pool holds six sequences only up to about 162k each, so fewer run concurrently above that; - whether YaRN (static scaling in vLLM) hurts quality on *short* prompts. Braid needs `model_context` and the ladder in (1) to follow the new window with no other code changes. ## Acceptance - The handoff point for a given setting is derived from the model window, the output reservation and a fixed margin, and the settings page shows all three. - `agent_context` cannot be set above what vLLM will accept. - A new attempt of a step that has a handoff starts from it. - A before/after measurement of the YaRN experiment is written up under `docs/design/measurements/`.
Author
Owner

Shipped in #62 (1ea5d99, archived as 2026-09-26-context-ceiling-and-continue), deployed as 798f838.

1. The ladder comes from the real ceiling. ceiling = window − min(output reservation, 32,000), hard = min(context per agent, ceiling), alarm = hard − 30,000 (or a quarter of hard, if smaller), allowance = hard − 6,000 at the default margin. Defaults are 150,000 / 120,000 / 144,000. The ceiling is checked when settings are saved and clamped where it's used, so a value saved before this change is never silently dropped. The settings page shows the ceiling, the hand-off point and the limit as you type.

2. The output reservation is a setting, default 20,480 (it was 32,768), written into each run's limit.output.

3. Resume keeps the work. A resumed agent step continues without a rollback, seeded with its last handoff (now recorded on osf.budget.respawned) plus git status/git diff. Redo still starts over, and scripted steps still roll back.

4. Slower growth. opencode's compaction.prune turned out to be a dead end: it runs only when a prompt loop ends and never touches the last two user turns. Instead, Braid's plugin stops a read with no line range at 500 lines, and the implement prompts say to grep first. The handoff turn runs without thinking through an opencode model variant (chat_template_kwargs: {enable_thinking: false}), which I checked on the wire against opencode 1.18.21 and trogdor. It applies only on vLLM, SGLang and llama.cpp.

5. YaRN 2x has been live on trogdor since 2026-09-26 19:16 EDT (trog c2a7567, --max-model-len 524288). A 280k-token request was accepted and recalled a fact planted at the start. Production settings are a 524,288 window with 400,000 per agent, which gives an alarm at 370,000 on runs provisioned from now on. docs/design/measurements/yarn-2x-PLAN.md holds the baseline (29 handoffs, 1.81 per run, a mean handoff turn of 317 s, 153 min of handoff turns in total) and yarn-2x-baseline.py, which re-measures it with --since 1790464592.

Verification: the full test lane passed (1,778 backend, 406 UI unit and 79 browser tests; the two @live specs were not run), and the settings page was checked in a browser.

Shipped in #62 (`1ea5d99`, archived as `2026-09-26-context-ceiling-and-continue`), deployed as `798f838`. **1. The ladder comes from the real ceiling.** `ceiling = window − min(output reservation, 32,000)`, hard = `min(context per agent, ceiling)`, alarm = hard − 30,000 (or a quarter of hard, if smaller), allowance = hard − 6,000 at the default margin. Defaults are 150,000 / 120,000 / 144,000. The ceiling is checked when settings are saved and clamped where it's used, so a value saved before this change is never silently dropped. The settings page shows the ceiling, the hand-off point and the limit as you type. **2. The output reservation is a setting**, default 20,480 (it was 32,768), written into each run's `limit.output`. **3. Resume keeps the work.** A resumed agent step continues without a rollback, seeded with its last handoff (now recorded on `osf.budget.respawned`) plus `git status`/`git diff`. Redo still starts over, and scripted steps still roll back. **4. Slower growth.** opencode's `compaction.prune` turned out to be a dead end: it runs only when a prompt loop ends and never touches the last two user turns. Instead, Braid's plugin stops a `read` with no line range at 500 lines, and the implement prompts say to `grep` first. The handoff turn runs **without thinking** through an opencode model variant (`chat_template_kwargs: {enable_thinking: false}`), which I checked on the wire against opencode 1.18.21 and trogdor. It applies only on vLLM, SGLang and llama.cpp. **5. YaRN 2x** has been live on trogdor since 2026-09-26 19:16 EDT (trog `c2a7567`, `--max-model-len 524288`). A 280k-token request was accepted and recalled a fact planted at the start. Production settings are a 524,288 window with 400,000 per agent, which gives an alarm at 370,000 on runs provisioned from now on. `docs/design/measurements/yarn-2x-PLAN.md` holds the baseline (29 handoffs, 1.81 per run, a mean handoff turn of 317 s, 153 min of handoff turns in total) and `yarn-2x-baseline.py`, which re-measures it with `--since 1790464592`. Verification: the full test lane passed (1,778 backend, 406 UI unit and 79 browser tests; the two `@live` specs were not run), and the settings page was checked in a browser.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
cmoriarty/braid#60
No description provided.