Archive won't delete an old run's files: a stale opencode pid reads as its server still running, and a slow footprint gives up after 15 s #154
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Archiving run 22 (a run that succeeded days ago) offers "Also delete its files on disk (749 MB)" greyed out, with the warning "The agent server for this run is still running (pid 31)." Nothing of run 22 is running.
Cause
run.oc_pidis written when a run's opencode server starts (RUN_SERVER_STARTED,engine/runtime.py), and nothing ever clears it._reap_finishedstops the server and logs "opencode stopped", but it doesn't record that, and the pid is lost to a restart anyway. The archive footprint (footprintinapi/archive.py) then asksprocess_alive(oc_pid), which isos.kill(pid, 0). That reports any process with that number as alive, including a thread: on Linux,killaccepts a thread id.On production on 2026-10-06, 49 runs had an
oc_pid. Exactly one answered: run 22's 31, which is a thread of osfd itself (/proc/31/cmdlineisosfd serve;pslists osfd as pid 16). Which old run gets blocked changes with every restart, and inside the container small pids get reused constantly.Want
oc_pid, or keep the pid together with the boot it belongs to, so that a pid from an earlier boot is never taken as alive.opencode servewith that run's directory or port. A pid that is osfd, or anything else, isn't the run's server./api/status's_pid_alive(api/app.py) and the watchdog's, for the same false positive.recovery.process_aliveis deliberately crude for a step that died mid-run, and that stays as it is. This issue is only about runs that are over.A second way the archive dialog loses the delete: the footprint times out.
Archiving run 26 (
run_01M3NCK3GZYRMQJ3CTVA0A906S, succeeded) showed: "Could not read what it holds on disk (osfd did not answer within 15 s. If other Braid tabs are open, close them: the browser may be out of connections to osfd.), so archiving leaves its files where they are." The checkbox was gone, and the dialog offers no retry.A few minutes later,
GET /api/runs/run_01M3NCK3GZYRMQJ3CTVA0A906S/footprintanswered in 1.7 s, twice: 1.3 GB,reclaimable: true, no reason. So the request was slow at that moment, not refused.What was going on then: six runs had just been archived back to back between 05:48 and 05:50 UTC, each deleting 170 MB to 1.2 GB (
shutil.rmtreein a thread). The footprint's_sizelstats every file in the run's folder, which for this run is 1.3 GB of mostly small files (caches, installs). Both run off the event loop, so the likely cause is a cold walk competing for the disk with the previous archive's delete. It may also have been the browser's connection limit, as the message suggests; that has not been checked.Want (in addition to the pid part)
Archive won't delete an old run's files: a stale opencode pid, now one of osfd's own threads, reads as its server still runningto Archive won't delete an old run's files: a stale opencode pid reads as its server still running, and a slow footprint gives up after 15 s