Archive won't delete an old run's files: a stale opencode pid reads as its server still running, and a slow footprint gives up after 15 s #154

Open
opened 2026-10-06 01:50:21 -04:00 by cmoriarty · 1 comment
Owner

Archiving run 22 (a run that succeeded days ago) offers "Also delete its files on disk (749 MB)" greyed out, with the warning "The agent server for this run is still running (pid 31)." Nothing of run 22 is running.

Cause

run.oc_pid is written when a run's opencode server starts (RUN_SERVER_STARTED, engine/runtime.py), and nothing ever clears it. _reap_finished stops the server and logs "opencode stopped", but it doesn't record that, and the pid is lost to a restart anyway. The archive footprint (footprint in api/archive.py) then asks process_alive(oc_pid), which is os.kill(pid, 0). That reports any process with that number as alive, including a thread: on Linux, kill accepts a thread id.

On production on 2026-10-06, 49 runs had an oc_pid. Exactly one answered: run 22's 31, which is a thread of osfd itself (/proc/31/cmdline is osfd serve; ps lists osfd as pid 16). Which old run gets blocked changes with every restart, and inside the container small pids get reused constantly.

Want

  • A run's server pid stops counting once the server is gone. Record the stop (when the server is reaped, and when boot recovery finds that a run's server has died) and clear oc_pid, or keep the pid together with the boot it belongs to, so that a pid from an earlier boot is never taken as alive.
  • The archive check asks whether that run's opencode server is running: the pid is a process (not a thread), and its command line is opencode serve with that run's directory or port. A pid that is osfd, or anything else, isn't the run's server.
  • A run that is not live (succeeded, failed, abandoned) and whose server is not running can delete its files.
  • Look at the other pid checks, /api/status's _pid_alive (api/app.py) and the watchdog's, for the same false positive. recovery.process_alive is deliberately crude for a step that died mid-run, and that stays as it is. This issue is only about runs that are over.
Archiving run 22 (a run that succeeded days ago) offers "Also delete its files on disk (749 MB)" greyed out, with the warning "The agent server for this run is still running (pid 31)." Nothing of run 22 is running. ## Cause `run.oc_pid` is written when a run's opencode server starts (`RUN_SERVER_STARTED`, `engine/runtime.py`), and nothing ever clears it. `_reap_finished` stops the server and logs "opencode stopped", but it doesn't record that, and the pid is lost to a restart anyway. The archive footprint (`footprint` in `api/archive.py`) then asks `process_alive(oc_pid)`, which is `os.kill(pid, 0)`. That reports any process with that number as alive, including a thread: on Linux, `kill` accepts a thread id. On production on 2026-10-06, 49 runs had an `oc_pid`. Exactly one answered: run 22's 31, which is a thread of osfd itself (`/proc/31/cmdline` is `osfd serve`; `ps` lists osfd as pid 16). Which old run gets blocked changes with every restart, and inside the container small pids get reused constantly. ## Want - A run's server pid stops counting once the server is gone. Record the stop (when the server is reaped, and when boot recovery finds that a run's server has died) and clear `oc_pid`, or keep the pid together with the boot it belongs to, so that a pid from an earlier boot is never taken as alive. - The archive check asks whether *that run's* opencode server is running: the pid is a process (not a thread), and its command line is `opencode serve` with that run's directory or port. A pid that is osfd, or anything else, isn't the run's server. - A run that is not live (succeeded, failed, abandoned) and whose server is not running can delete its files. - Look at the other pid checks, `/api/status`'s `_pid_alive` (`api/app.py`) and the watchdog's, for the same false positive. `recovery.process_alive` is deliberately crude for a step that died mid-run, and that stays as it is. This issue is only about runs that are over.
Author
Owner

A second way the archive dialog loses the delete: the footprint times out.

Archiving run 26 (run_01M3NCK3GZYRMQJ3CTVA0A906S, succeeded) showed: "Could not read what it holds on disk (osfd did not answer within 15 s. If other Braid tabs are open, close them: the browser may be out of connections to osfd.), so archiving leaves its files where they are." The checkbox was gone, and the dialog offers no retry.

A few minutes later, GET /api/runs/run_01M3NCK3GZYRMQJ3CTVA0A906S/footprint answered in 1.7 s, twice: 1.3 GB, reclaimable: true, no reason. So the request was slow at that moment, not refused.

What was going on then: six runs had just been archived back to back between 05:48 and 05:50 UTC, each deleting 170 MB to 1.2 GB (shutil.rmtree in a thread). The footprint's _size lstats every file in the run's folder, which for this run is 1.3 GB of mostly small files (caches, installs). Both run off the event loop, so the likely cause is a cold walk competing for the disk with the previous archive's delete. It may also have been the browser's connection limit, as the message suggests; that has not been checked.

Want (in addition to the pid part)

  • The dialog doesn't silently fall back to "leave the files" when the footprint is slow. It keeps waiting and shows that it's measuring, or it offers to try again, so archiving many runs in a row doesn't quietly keep their files.
  • Deleting doesn't depend on knowing the exact size first. The size is shown when it's known (or the last size measured is used), and the safety checks (path inside the runs root, branch kept, server not running) are what decide whether the files can be deleted, not the size walk.
  • Only blame the browser's connections when that is really the cause.
A second way the archive dialog loses the delete: **the footprint times out.** Archiving run 26 (`run_01M3NCK3GZYRMQJ3CTVA0A906S`, succeeded) showed: "Could not read what it holds on disk (osfd did not answer within 15 s. If other Braid tabs are open, close them: the browser may be out of connections to osfd.), so archiving leaves its files where they are." The checkbox was gone, and the dialog offers no retry. A few minutes later, `GET /api/runs/run_01M3NCK3GZYRMQJ3CTVA0A906S/footprint` answered in 1.7 s, twice: 1.3 GB, `reclaimable: true`, no reason. So the request was slow at that moment, not refused. What was going on then: six runs had just been archived back to back between 05:48 and 05:50 UTC, each deleting 170 MB to 1.2 GB (`shutil.rmtree` in a thread). The footprint's `_size` `lstat`s every file in the run's folder, which for this run is 1.3 GB of mostly small files (caches, installs). Both run off the event loop, so the likely cause is a cold walk competing for the disk with the previous archive's delete. It may also have been the browser's connection limit, as the message suggests; that has not been checked. ## Want (in addition to the pid part) - The dialog doesn't silently fall back to "leave the files" when the footprint is slow. It keeps waiting and shows that it's measuring, or it offers to try again, so archiving many runs in a row doesn't quietly keep their files. - Deleting doesn't depend on knowing the exact size first. The size is shown when it's known (or the last size measured is used), and the safety checks (path inside the runs root, branch kept, server not running) are what decide whether the files can be deleted, not the size walk. - Only blame the browser's connections when that is really the cause.
cmoriarty changed title from Archive won't delete an old run's files: a stale opencode pid, now one of osfd's own threads, reads as its server still running to Archive won't delete an old run's files: a stale opencode pid reads as its server still running, and a slow footprint gives up after 15 s 2026-10-06 01:51:39 -04:00
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
cmoriarty/braid#154
No description provided.