Apply runs one step per tasks.md heading: pack small groups into one session, up to about 150k tokens #132

Closed
opened 2026-10-02 23:17:03 -04:00 by cmoriarty · 2 comments
Owner

openspec.apply fans out one step per numbered heading of tasks.md (fanout_over: .osf/budget.json#groups[].id). The estimator only splits a group that is too big for the allowance, and with a 400k context per agent that allowance is now 394k tokens. So the number of apply steps is whatever the tasks writer chose, usually 3 or 4, and they run one after another in the run's worktree. Each step starts a session and orients itself again.

Measured

Runs 22, 23, 25, 26, 36, 37 and run 39 (scratch, the last six feature runs before #121 and the first after it):

  • 23 apply steps in 7 runs, 3 or 4 per run; 1 was redone.

  • 184 of the 417 apply minutes (44%) came before each step's first code edit. In run 39, after #121's brief fixes, it was 51%.

  • Context per step: the first prompt is 17–48k and the peak is 91–158k, against a limit of 400k. One session doing every group would peak at about 153–372k.

  • Decode speed by context, on turns with no other stream running:

    • 51 tok/s under 50k;
    • 47 tok/s at 50–100k;
    • 42 tok/s at 100–150k;
    • 40 tok/s at 150–200k.

    Time to first token stays near 1.8 s throughout, thanks to prefix caching. No apply step went past 200k.

So starting a session costs almost nothing, but orienting does. One session per change would orient only once, but it would spend most of its turns at 150–300k, decoding 15–25% slower, with the quality risk of a 27B model at 300k and more work lost if it fails late.

Proposal

  • Let the estimate pack consecutive task groups into one apply step while the step's predicted peak stays under a target of about 150k, which is still in the fast range. Most changes would then run as 2 apply steps instead of 3 or 4.
  • Keep each group's own completion check and the groups' order.
  • Measure apply's wall time and orientation per run against run 39 and the #124 measurement run, both on cmoriarty/scratch#6. Wait for #124's result first, because most orientation time is thinking, which #124 changes.

Related: #124 (reasoning effort). #128 is the bigger lever: groups that do not depend on each other could run side by side in separate worktrees.

`openspec.apply` fans out one step per numbered heading of `tasks.md` (`fanout_over: .osf/budget.json#groups[].id`). The estimator only splits a group that is too big for the allowance, and with a 400k context per agent that allowance is now 394k tokens. So the number of apply steps is whatever the tasks writer chose, usually 3 or 4, and they run one after another in the run's worktree. Each step starts a session and orients itself again. ## Measured Runs 22, 23, 25, 26, 36, 37 and run 39 (scratch, the last six feature runs before #121 and the first after it): - 23 apply steps in 7 runs, 3 or 4 per run; 1 was redone. - 184 of the 417 apply minutes (44%) came before each step's first code edit. In run 39, after #121's brief fixes, it was 51%. - Context per step: the first prompt is 17–48k and the peak is 91–158k, against a limit of 400k. One session doing every group would peak at about 153–372k. - Decode speed by context, on turns with no other stream running: - 51 tok/s under 50k; - 47 tok/s at 50–100k; - 42 tok/s at 100–150k; - 40 tok/s at 150–200k. Time to first token stays near 1.8 s throughout, thanks to prefix caching. No apply step went past 200k. So starting a session costs almost nothing, but orienting does. One session per change would orient only once, but it would spend most of its turns at 150–300k, decoding 15–25% slower, with the quality risk of a 27B model at 300k and more work lost if it fails late. ## Proposal - Let the estimate pack consecutive task groups into one apply step while the step's predicted peak stays under a target of about 150k, which is still in the fast range. Most changes would then run as 2 apply steps instead of 3 or 4. - Keep each group's own completion check and the groups' order. - Measure apply's wall time and orientation per run against run 39 and the #124 measurement run, both on cmoriarty/scratch#6. Wait for #124's result first, because most orientation time is thinking, which #124 changes. Related: #124 (reasoning effort). #128 is the bigger lever: groups that do not depend on each other could run side by side in separate worktrees.
Author
Owner

I measured this before building it, on runs 22, 23, 25, 26, 36, 37 and 39, and packing isn't worth it now:

Packing target Apply steps (7 runs) At most saved
150k 23 → 19 10 min in total, about 1.4 min per run
200k 23 → 15 40 min, about 6 min per run
250k 23 → 10 77 min, about 11 min per run, before slower decoding above 150k takes its share

These figures assume perfect estimates, using each step's real opening context and peak. They also count a joined group's whole orientation as saved, which overstates it: a merged session still has to read and think about each group's code.

The estimates are far from perfect. The estimator has never been calibrated:

  • Nothing emits osf.budget.estimated, so all 64 apply rows of budget_ledger lack an estimate.
  • script.budget never calls calibrate().
  • Against budget.json in runs 36, 37, 39 and 40, the actual peaks are 1.3–2.75× the p50 (about 2.1× at the median), and above the p80 in 7 of 11 groups.
  • Run 39's smallest estimate was its largest session.

So this issue is narrowed to closing that loop (change record-apply-estimates):

  • record each apply attempt's estimate beside its peak;
  • have script.budget estimate with the ledger's calibration.

Packing is dropped, and the numbers above are here for when it comes up again. The bigger gain for apply is running independent groups side by side, which belongs to #128.

I measured this before building it, on runs 22, 23, 25, 26, 36, 37 and 39, and packing isn't worth it now: | Packing target | Apply steps (7 runs) | At most saved | |---|---|---| | 150k | 23 → 19 | 10 min in total, about 1.4 min per run | | 200k | 23 → 15 | 40 min, about 6 min per run | | 250k | 23 → 10 | 77 min, about 11 min per run, before slower decoding above 150k takes its share | These figures assume perfect estimates, using each step's real opening context and peak. They also count a joined group's whole orientation as saved, which overstates it: a merged session still has to read and think about each group's code. **The estimates are far from perfect.** The estimator has never been calibrated: - Nothing emits `osf.budget.estimated`, so all 64 apply rows of `budget_ledger` lack an estimate. - `script.budget` never calls `calibrate()`. - Against `budget.json` in runs 36, 37, 39 and 40, the actual peaks are 1.3–2.75× the p50 (about 2.1× at the median), and above the p80 in 7 of 11 groups. - Run 39's smallest estimate was its largest session. **So this issue is narrowed to closing that loop** (change `record-apply-estimates`): - record each apply attempt's estimate beside its peak; - have `script.budget` estimate with the ledger's calibration. Packing is dropped, and the numbers above are here for when it comes up again. The bigger gain for apply is running independent groups side by side, which belongs to #128.
Author
Owner

Shipped in b58c34d (proposal 536bfa7) and deployed on 2026-10-03 at 01:16 as part of 7b6a101. Archived as openspec/changes/archive/2026-10-03-record-apply-estimates.

What changed

  • Each apply attempt records its estimate. When it opens, the agent executor emits osf.budget.estimated with its group's estimate from budget.json. Its budget_ledger row then holds est_p50, est_p80 and allowance beside the attempt's peak. A missing file or group records nothing.
  • script.budget estimates with the record's calibration. osfd fits it from those rows (shape impl, shrunk toward the bootstrap by n/(n+10), applying from five samples) and passes it to scripted steps as OSF_CALIBRATION_JSON. budget.json names the calibration's source and n.
  • Packing, this issue's first idea, was measured and dropped. The numbers are in the comment above.

Verified on production

  • Run 40's resumed apply[2] attempt recorded est_p50 37,739, est_p80 61,461 and allowance 394,000.
  • The console's Step details draws the "estimated 61k" marker and "VS ESTIMATE −29% under estimate".
  • The calibration starts applying once five attempts are recorded, about two runs. The unit tests cover that path.
Shipped in b58c34d (proposal 536bfa7) and deployed on 2026-10-03 at 01:16 as part of 7b6a101. Archived as `openspec/changes/archive/2026-10-03-record-apply-estimates`. **What changed** - **Each apply attempt records its estimate.** When it opens, the agent executor emits `osf.budget.estimated` with its group's estimate from `budget.json`. Its `budget_ledger` row then holds `est_p50`, `est_p80` and `allowance` beside the attempt's peak. A missing file or group records nothing. - **`script.budget` estimates with the record's calibration.** osfd fits it from those rows (shape `impl`, shrunk toward the bootstrap by n/(n+10), applying from five samples) and passes it to scripted steps as `OSF_CALIBRATION_JSON`. `budget.json` names the calibration's source and n. - **Packing, this issue's first idea, was measured and dropped.** The numbers are in the comment above. **Verified on production** - Run 40's resumed `apply[2]` attempt recorded `est_p50` 37,739, `est_p80` 61,461 and `allowance` 394,000. - The console's Step details draws the "estimated 61k" marker and "VS ESTIMATE −29% under estimate". - The calibration starts applying once five attempts are recorded, about two runs. The unit tests cover that path.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
cmoriarty/braid#132
No description provided.