Handoffs come too early: derive the context ladder from the real ceiling, carry handoffs across attempts, try YaRN #60
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Problem
Handoffs are the biggest slowdown in a run. The context per agent was raised to 262,144, but steps still hand off long before that. On production, run
run_01M3FS9FFXEHTY5WTHX36V4Z5Y, stepfix.implement(stp_01M3FS9JXBQJN4A21P8HBWMY0S), commit2064a2f:osf.budget.alarmfired (alarm: 195734), the session wrote a 10,051-character handoff, and a fresh session started from it.no non-empty file matched: .osf/fix.md), and attempt 2 started cold, without the handoff. It re-read the repository and was at 129k within 10 minutes.readcalls came to about 91k tokens (two reads undersrc/osf/pipeline/were 19k and 10k), and 28bashcalls added about 15k. The rest is many small turns of 100–2,000 tokens each.The engine behaved as specified. The ladder is the problem.
Findings
Settings.alarm = H × 112/150andallowance = H × 144/150(src/osf/settings.py). Those ratios were set when H was 150k, to leave 38k for a measured worst-case overshoot of 32,435. At H = 262,144 the same ratio leaves 66k unused.limit.output: 32768(MODEL_OUTPUT_LIMIT), and opencode asks for up to that much output on each request. vLLM rejects any request where prompt plusmax_tokensexceeds--max-model-len 262144, so real input is capped at about 230k. The research agrees: every recorded context drop clusters at 229,751–230,137. So a step that gets past the alarm without a working handoff fails with a provider error at about 230k, and the compaction fallback, which triggers only atpeak >= H, never runs.validate()allowsagent_contextup tomodel_context; the limit should bemodel_context − output reservation. To confirm on the wire: themax_tokensopencode actually sends.OPENCODE_EXPERIMENTAL_OUTPUT_TOKEN_MAX, which is a second way to lower it.attempt.peak_input_tokensfor attempt 1 reads 104,129 (only the respawned session), whilebudget_ledgerreads 196,115.Explore all five
Base the ladder on the ceiling, not a ratio.
ceiling = model_context − output_reservation;H = min(agent_context, ceiling);A = H − margin, with the margin a fixed measured number (about 30k: one worst turn plus the handoff output) rather than 25% of H. Clampagent_contextto the ceiling in settings validation, and show the ceiling on the settings page.Lower the output reservation from 32,768 to 16,384 (or make it a setting). At 262,144 this moves the ceiling from about 230k to about 246k and the alarm to about 214k. Measure how often a turn reaches the cap.
Carry the handoff across attempts. A new attempt of the same step, after an interrupt, redo, resume or failed predicate, starts from the last handoff written for that step, together with
git diff, and does not start cold.Slow the growth. Find out whether opencode's tool-output pruning (
compaction.prune) runs withcompaction.auto: falseand whether it helps. Try steering agents toward ranged reads,grepand thereadersubagent over whole-filereads of large files.Experiment: raise vLLM's window with YaRN on trogdor. For example,
--max-model-len 524288with a YaRNrope_scalingfactor of 2 over the native 262,144. The operator is willing to run this. What to measure:EFFECTIVE_WINDOW);Braid needs
model_contextand the ladder in (1) to follow the new window with no other code changes.Acceptance
agent_contextcannot be set above what vLLM will accept.docs/design/measurements/.Shipped in #62 (
1ea5d99, archived as2026-09-26-context-ceiling-and-continue), deployed as798f838.1. The ladder comes from the real ceiling.
ceiling = window − min(output reservation, 32,000), hard =min(context per agent, ceiling), alarm = hard − 30,000 (or a quarter of hard, if smaller), allowance = hard − 6,000 at the default margin. Defaults are 150,000 / 120,000 / 144,000. The ceiling is checked when settings are saved and clamped where it's used, so a value saved before this change is never silently dropped. The settings page shows the ceiling, the hand-off point and the limit as you type.2. The output reservation is a setting, default 20,480 (it was 32,768), written into each run's
limit.output.3. Resume keeps the work. A resumed agent step continues without a rollback, seeded with its last handoff (now recorded on
osf.budget.respawned) plusgit status/git diff. Redo still starts over, and scripted steps still roll back.4. Slower growth. opencode's
compaction.pruneturned out to be a dead end: it runs only when a prompt loop ends and never touches the last two user turns. Instead, Braid's plugin stops areadwith no line range at 500 lines, and the implement prompts say togrepfirst. The handoff turn runs without thinking through an opencode model variant (chat_template_kwargs: {enable_thinking: false}), which I checked on the wire against opencode 1.18.21 and trogdor. It applies only on vLLM, SGLang and llama.cpp.5. YaRN 2x has been live on trogdor since 2026-09-26 19:16 EDT (trog
c2a7567,--max-model-len 524288). A 280k-token request was accepted and recalled a fact planted at the start. Production settings are a 524,288 window with 400,000 per agent, which gives an alarm at 370,000 on runs provisioned from now on.docs/design/measurements/yarn-2x-PLAN.mdholds the baseline (29 handoffs, 1.81 per run, a mean handoff turn of 317 s, 153 min of handoff turns in total) andyarn-2x-baseline.py, which re-measures it with--since 1790464592.Verification: the full test lane passed (1,778 backend, 406 UI unit and 79 browser tests; the two
@livespecs were not run), and the settings page was checked in a browser.