Apply runs one step per tasks.md heading: pack small groups into one session, up to about 150k tokens #132
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
openspec.applyfans out one step per numbered heading oftasks.md(fanout_over: .osf/budget.json#groups[].id). The estimator only splits a group that is too big for the allowance, and with a 400k context per agent that allowance is now 394k tokens. So the number of apply steps is whatever the tasks writer chose, usually 3 or 4, and they run one after another in the run's worktree. Each step starts a session and orients itself again.Measured
Runs 22, 23, 25, 26, 36, 37 and run 39 (scratch, the last six feature runs before #121 and the first after it):
23 apply steps in 7 runs, 3 or 4 per run; 1 was redone.
184 of the 417 apply minutes (44%) came before each step's first code edit. In run 39, after #121's brief fixes, it was 51%.
Context per step: the first prompt is 17–48k and the peak is 91–158k, against a limit of 400k. One session doing every group would peak at about 153–372k.
Decode speed by context, on turns with no other stream running:
Time to first token stays near 1.8 s throughout, thanks to prefix caching. No apply step went past 200k.
So starting a session costs almost nothing, but orienting does. One session per change would orient only once, but it would spend most of its turns at 150–300k, decoding 15–25% slower, with the quality risk of a 27B model at 300k and more work lost if it fails late.
Proposal
Related: #124 (reasoning effort). #128 is the bigger lever: groups that do not depend on each other could run side by side in separate worktrees.
I measured this before building it, on runs 22, 23, 25, 26, 36, 37 and 39, and packing isn't worth it now:
These figures assume perfect estimates, using each step's real opening context and peak. They also count a joined group's whole orientation as saved, which overstates it: a merged session still has to read and think about each group's code.
The estimates are far from perfect. The estimator has never been calibrated:
osf.budget.estimated, so all 64 apply rows ofbudget_ledgerlack an estimate.script.budgetnever callscalibrate().budget.jsonin runs 36, 37, 39 and 40, the actual peaks are 1.3–2.75× the p50 (about 2.1× at the median), and above the p80 in 7 of 11 groups.So this issue is narrowed to closing that loop (change
record-apply-estimates):script.budgetestimate with the ledger's calibration.Packing is dropped, and the numbers above are here for when it comes up again. The bigger gain for apply is running independent groups side by side, which belongs to #128.
Shipped in
b58c34d(proposal536bfa7) and deployed on 2026-10-03 at 01:16 as part of7b6a101. Archived asopenspec/changes/archive/2026-10-03-record-apply-estimates.What changed
osf.budget.estimatedwith its group's estimate frombudget.json. Itsbudget_ledgerrow then holdsest_p50,est_p80andallowancebeside the attempt's peak. A missing file or group records nothing.script.budgetestimates with the record's calibration. osfd fits it from those rows (shapeimpl, shrunk toward the bootstrap by n/(n+10), applying from five samples) and passes it to scripted steps asOSF_CALIBRATION_JSON.budget.jsonnames the calibration's source and n.Verified on production
apply[2]attempt recordedest_p5037,739,est_p8061,461 andallowance394,000..arrives, and the agent-view spec fails under load #130