feat(budget): the handoff point comes from what the model server accepts, and a resumed step keeps its work (#60) #62

Merged
cmoriarty merged 2 commits from feat/context-ceiling-60 into main 2026-09-26 19:57:56 -04:00
Owner

Fixes #60.

Handoffs were the slowest part of a run, and raising the context per agent to 262,144 didn't make them rarer. The alarm was a fixed 75% of the hard limit (195,734 at 262,144). The hard limit itself could never be reached: opencode asks vLLM for max_tokens of about 32,000 on every request (2,219 of 2,222 on trogdor), and vLLM refuses prompt plus max_tokens above its window. So the compaction backstop at the hard limit never fired.

What changes

  • The ladder comes from the ceiling (osf.engine.budget.ladder, the one implementation, mirrored by ladderOf in the UI with the same test cases on both sides):

    • ceiling = window − min(output reservation, 32,000)
    • hard limit = min(context per agent, ceiling)
    • alarm = hard limit − 30,000 (or a quarter of the hard limit, if smaller): room for the largest turn on record, 28,983
    • allowance = hard limit − a fifth of that margin

    Defaults: 150,000 hard / 120,000 alarm / 144,000 allowance (the alarm was 112,000). An attempt is capped by its run's own window and reservation.

  • The ceiling is checked on save only. read_settings drops any saved value that fails validate, so production's saved 262,144 is clamped where it is used instead of silently turning back into 150,000. There's a test for exactly that.

  • The output reservation is a setting, default 20,480 (it was a fixed 32,768). It's written into each run's limit.output, and a run keeps its value. The settings page shows the ceiling, the hand-off point and the limit as they are typed, plus a "Limited to N by the ceiling" note when the context per agent is clamped.

  • Resume keeps the work. A resumed stopped or failed agent step no longer rolls its worktree back. On production, a stop then a resume put fix.implement back on its entry tree and erased about an hour of edits. The new attempt starts from the step's last handoff (the respawn record now carries the text) plus git status/git diff. Redo is still the clean restart, and scripted steps still roll back.

  • Slower growth, cheaper handoffs.

    • A read with no line range stops at 500 lines, via Braid's opencode plugin, tested in Node. Whole-file reads were 91k of a 196k attempt.
    • The handoff turn runs without thinking on vLLM, SGLang and llama.cpp, through an opencode model variant carrying chat_template_kwargs. I checked it on the wire against opencode 1.18.21 and trogdor: only that prompt carries it. The handoff turn averaged 5.3 minutes over 29 production handoffs.
  • YaRN 2x plan: docs/design/measurements/yarn-2x-PLAN.md pre-registers the experiment (trogdor has run at 524,288 since 2026-09-26 19:16 EDT), with its baseline and the query that re-measures it.

Verification

  • ./tools/test.sh full: 1,778 backend tests, 406 UI unit tests and 79 browser tests passed. The two @live specs (step pane, transcript tail, not touched here) were not run: no local osfd with runs was up.
  • Checked the settings page in a real browser: the ceiling moves as the reservation is typed, and a context per agent above the ceiling is marked and refused.
  • openspec validate --specs --strict: all 27 specs pass. The change is archived in this PR, with its deltas synced into settings, context-budget and run-resume.

On deploy

Production's saved context per agent (400,000) with the saved 524,288 window gives hard 400,000 / alarm 370,000 for runs provisioned afterwards. Runs provisioned earlier keep their own window and a 32,000 reservation.

🤖 Generated with Claude Code

Fixes #60. Handoffs were the slowest part of a run, and raising the context per agent to 262,144 didn't make them rarer. The alarm was a fixed 75% of the hard limit (195,734 at 262,144). The hard limit itself could never be reached: opencode asks vLLM for `max_tokens` of about 32,000 on every request (2,219 of 2,222 on trogdor), and vLLM refuses prompt plus `max_tokens` above its window. So the compaction backstop at the hard limit never fired. ## What changes - **The ladder comes from the ceiling** (`osf.engine.budget.ladder`, the one implementation, mirrored by `ladderOf` in the UI with the same test cases on both sides): - `ceiling = window − min(output reservation, 32,000)` - hard limit = `min(context per agent, ceiling)` - alarm = hard limit − 30,000 (or a quarter of the hard limit, if smaller): room for the largest turn on record, 28,983 - allowance = hard limit − a fifth of that margin Defaults: 150,000 hard / 120,000 alarm / 144,000 allowance (the alarm was 112,000). An attempt is capped by its run's own window and reservation. - **The ceiling is checked on save only.** `read_settings` drops any saved value that fails `validate`, so production's saved 262,144 is clamped where it is used instead of silently turning back into 150,000. There's a test for exactly that. - **The output reservation is a setting**, default 20,480 (it was a fixed 32,768). It's written into each run's `limit.output`, and a run keeps its value. The settings page shows the ceiling, the hand-off point and the limit as they are typed, plus a "Limited to N by the ceiling" note when the context per agent is clamped. - **Resume keeps the work.** A resumed stopped or failed agent step no longer rolls its worktree back. On production, a stop then a resume put `fix.implement` back on its entry tree and erased about an hour of edits. The new attempt starts from the step's last handoff (the respawn record now carries the text) plus `git status`/`git diff`. Redo is still the clean restart, and scripted steps still roll back. - **Slower growth, cheaper handoffs.** - A `read` with no line range stops at 500 lines, via Braid's opencode plugin, tested in Node. Whole-file reads were 91k of a 196k attempt. - The handoff turn runs without thinking on vLLM, SGLang and llama.cpp, through an opencode model variant carrying `chat_template_kwargs`. I checked it on the wire against opencode 1.18.21 and trogdor: only that prompt carries it. The handoff turn averaged 5.3 minutes over 29 production handoffs. - **YaRN 2x plan:** `docs/design/measurements/yarn-2x-PLAN.md` pre-registers the experiment (trogdor has run at 524,288 since 2026-09-26 19:16 EDT), with its baseline and the query that re-measures it. ## Verification - `./tools/test.sh full`: 1,778 backend tests, 406 UI unit tests and 79 browser tests passed. The two `@live` specs (step pane, transcript tail, not touched here) were not run: no local osfd with runs was up. - Checked the settings page in a real browser: the ceiling moves as the reservation is typed, and a context per agent above the ceiling is marked and refused. - `openspec validate --specs --strict`: all 27 specs pass. The change is archived in this PR, with its deltas synced into `settings`, `context-budget` and `run-resume`. ## On deploy Production's saved context per agent (400,000) with the saved 524,288 window gives hard 400,000 / alarm 370,000 for runs provisioned afterwards. Runs provisioned earlier keep their own window and a 32,000 reservation. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
Handoffs were the slowest part of a run, and raising the context per agent to 262,144 did not
make them rarer the way the setting promised. The alarm was a fixed 75% of the hard limit, so
262,144 handed off at 195,734. The hard limit itself could never be reached: opencode asks
vLLM for max_tokens of 32,000 on every request (2,219 of 2,222 requests on trogdor), and vLLM
refuses prompt plus max_tokens above its window, so the compaction backstop at H never fired.

The ladder is now derived by osf.engine.budget.ladder: ceiling = window - min(output
reservation, 32,000); H = min(context per agent, ceiling); the alarm sits 30,000 below H (a
quarter of H when that is smaller), room for the largest turn on record (28,983); the
allowance sits a fifth of that below H. The default ladder is 150,000 / 120,000 / 144,000. An
attempt is capped by its run's own window and reservation. The ceiling is checked on save
only: read_settings drops a saved value that fails validate, and production's saved 262,144
must be clamped where it is used, not turned back into 150,000.

The output reservation is a setting (default 20,480, was a fixed 32,768), written into each
run's opencode config and kept by the run. The settings page shows the ceiling, the hand-off
point and the limit as they are typed.

Resuming a run no longer rolls a stopped or failed agent step back: on production a stop then
a resume put fix.implement back on its entry tree and erased an hour of edits. The new attempt
keeps the worktree and is seeded with the step's last handoff, which the respawn record now
carries, and with git status and git diff. Redo is still the clean restart.

Two changes slow the growth itself. A read with no line range stops at 500 lines, because
whole-file reads were 91k of a 196k attempt. The handoff turn runs without thinking, through
an opencode model variant carrying chat_template_kwargs, measured on the wire against
opencode 1.18.21. It averaged 5.3 minutes over 29 production handoffs.

docs/design/measurements/yarn-2x-PLAN.md pre-registers the YaRN 2x experiment on trogdor
(524,288 since 2026-09-26 19:16 EDT) with its baseline and the query that measures it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Archive context-ceiling-and-continue
All checks were successful
deploy / deploy (push) Successful in 59s
798f838b09
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign in to join this conversation.
No reviewers
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
cmoriarty/braid!62
No description provided.