feat(budget): the handoff point comes from what the model server accepts, and a resumed step keeps its work (#60) #62
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "feat/context-ceiling-60"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Fixes #60.
Handoffs were the slowest part of a run, and raising the context per agent to 262,144 didn't make them rarer. The alarm was a fixed 75% of the hard limit (195,734 at 262,144). The hard limit itself could never be reached: opencode asks vLLM for
max_tokensof about 32,000 on every request (2,219 of 2,222 on trogdor), and vLLM refuses prompt plusmax_tokensabove its window. So the compaction backstop at the hard limit never fired.What changes
The ladder comes from the ceiling (
osf.engine.budget.ladder, the one implementation, mirrored byladderOfin the UI with the same test cases on both sides):ceiling = window − min(output reservation, 32,000)min(context per agent, ceiling)Defaults: 150,000 hard / 120,000 alarm / 144,000 allowance (the alarm was 112,000). An attempt is capped by its run's own window and reservation.
The ceiling is checked on save only.
read_settingsdrops any saved value that failsvalidate, so production's saved 262,144 is clamped where it is used instead of silently turning back into 150,000. There's a test for exactly that.The output reservation is a setting, default 20,480 (it was a fixed 32,768). It's written into each run's
limit.output, and a run keeps its value. The settings page shows the ceiling, the hand-off point and the limit as they are typed, plus a "Limited to N by the ceiling" note when the context per agent is clamped.Resume keeps the work. A resumed stopped or failed agent step no longer rolls its worktree back. On production, a stop then a resume put
fix.implementback on its entry tree and erased about an hour of edits. The new attempt starts from the step's last handoff (the respawn record now carries the text) plusgit status/git diff. Redo is still the clean restart, and scripted steps still roll back.Slower growth, cheaper handoffs.
readwith no line range stops at 500 lines, via Braid's opencode plugin, tested in Node. Whole-file reads were 91k of a 196k attempt.chat_template_kwargs. I checked it on the wire against opencode 1.18.21 and trogdor: only that prompt carries it. The handoff turn averaged 5.3 minutes over 29 production handoffs.YaRN 2x plan:
docs/design/measurements/yarn-2x-PLAN.mdpre-registers the experiment (trogdor has run at 524,288 since 2026-09-26 19:16 EDT), with its baseline and the query that re-measures it.Verification
./tools/test.sh full: 1,778 backend tests, 406 UI unit tests and 79 browser tests passed. The two@livespecs (step pane, transcript tail, not touched here) were not run: no local osfd with runs was up.openspec validate --specs --strict: all 27 specs pass. The change is archived in this PR, with its deltas synced intosettings,context-budgetandrun-resume.On deploy
Production's saved context per agent (400,000) with the saved 524,288 window gives hard 400,000 / alarm 370,000 for runs provisioned afterwards. Runs provisioned earlier keep their own window and a 32,000 reservation.
🤖 Generated with Claude Code