A running command reads as silence: show what is running, and don't redo a step waiting on its own tool call #69
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Problem
On
run_01M3GB3KN38YT9TZJ7KA7RQ5M0(soundcheck#32), anopenspec.applyagent rannpx vitest run src/lib/server/ops/maintenance.test.ts src/lib/server/artists/resolve.test.ts 2>&1 | tail -12for several minutes. While it ran:Cause
Silence is measured from the step's last event (
StepActivityinsrc/osf/format.py; the watchdog's_examinealso counts streamed tokens). A shell command produces no events while it runs, so a step waiting on its ownbashreads exactly like a step that has stopped. Yet opencode already records the tool part asstate.status: running, with itsinput(the command, whichtool_targetextracts) andtime.start. The digest and the watchdog just don't use it.The same cause can kill a long command
The stall watchdog (
src/osf/engine/watchdog.py, the loop overreport.checked) redoes an agent step whose session is busy and which has been silent for more thanSTALL_REDO_S= 480 s. A redo rolls the worktree back to the entry checkpoint. It already skips a step waiting on a subagent (child_busy) and a step waiting on the operator (status notrunning), but not a step waiting on its own tool call. opencode'sbashtool allows up to 10 minutes per call (the agent can raise the 2-minute default), and a test suite or build that prints only at the end produces nothing until it exits. So any command running quietly past 8 minutes gets its step redone mid-command, and the work is thrown away. The dogfood evidence behind 480 s ("longest legitimate silence 215 s") predates agents running whole suites themselves.Proposal
pendingorrunning, the step is not silent, and its phrase names what is running, e.g.Running npx vitest run … · 3m 12s, timed from the tool's owntime.start.running_tool: {tool, target, started_at}. The spine row showsRunning a command · 3m 12sand puts the full command in its tooltip. The stage header says the same.runningis itself a wedge, and the stall redo applies as today.Done when
bashcall has been running for 5 minutes readsRunning … · 5m, notsilent for 5m, in the spine and the stage header.bashcall runs 9 minutes is not redone by the watchdog. One whose tool call is stillrunningafter the ceiling plus the margin is.Shipped in #71 (
558560e, archived as2026-09-27-running-tool-is-not-silence), deployed asbd3902fat 01:07 EDT.running_tool,running_targetandrunning_for_s, timed from the call's own start.Running npx vitest run src/lib/… · 3m 12s. The tooltip has the whole command line, how long it has run, and the step's time.TOOL_CALL_CEILING_S(20 minutes). A call still running past that is a wedge, and the redo applies.Restart recovery still rolls running steps back: #70.
Shipped in PR #71 (
558560e), deployed to production atbd3902fon 2026-09-27 at 01:07.What shipped:
running_tool,running_targetandrunning_for_s.Checked live: I captured
/api/stream?run=run_01M3GB3KN38YT9TZJ7KA7RQ5M0for 90 s after the deploy. Digest frames foropenspec.apply[4]carriedrunning_toolforreadandgrepcalls with their targets (.../scripts/extract.ts,.../server/ops), andrunning_for_scounted up.One defect found in that check: a call that was in flight when the restart killed attempt 1 kept showing as
bashrunning, with no target, while attempt 2 worked. The digest's fold spans attempts. Only the display is affected: the watchdog's read is attempt-scoped. Filed as #73.