A step that hits its time limit fails with no reason: name the cause, show it in the console, and show the clock #116
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
On production, run run_01M3WX1VW43K77HWNQ328RBGJK (git-flow, cmoriarty/scratch#5, run #36 in the console) failed at agent.test.browser, and the console gave no reason.
The step ran its whole one-hour limit (
timeout_s3600) and the engine stopped it at 3,603 s, before the agent wrote.osf/test/browser.json. The record says only that the file is missing or empty. The attempt row'serror_classanderror_textare null; the cause,reason: "timed_out", is only on theosf.attempt.closedevent, and nothing in the UI reads it. The console shows✗ interrupted, which reads like a person stopped it, then✗ agent.test.browser failed. Step details says "#1 failed". The UI never renderserror_text, a running step shows no clock against its limit, and the failure email was not sent (no SMTP credentials on the box).It is the first timeout in production to fail a step (1 of 650 attempts; the other, on 2026-09-28, still passed its file check). The output limit had the same shape in #101 and was fixed for that cause only.
Asked for:
error_classanderror_texton the attempt ("stopped at its 1h limit after 1h 00m"), and the step's failure leads with it, as #101 did for the output limit.error_textunder the failed row in the transcript, on the attempt in Step details, and as a one-line reason beside the Failed chip. When the engine stopped the step, the row before the stop says so instead of✗ interrupted. The notification headline names the step, not its id.Shipped and deployed, and archived as
step-failure-cause. Commitsf8405daand543e54d, live sinceac2caaf(2026-10-02 05:12 UTC); production now runs3bc80ce.timed_out,budget_exceeded) and its failure leads with it:stopped at its 1h time limit; .osf/test/browser.json is missing or empty. The engine appendsosf.step.timed_outbefore it aborts a step, so the console no longer draws that abort asinterrupted, which read as the operator's own Stop.osf.step.deadline, and again whenever the clock restarts; migration 9 addsattempt.timeout_sandattempt.deadline_at). Step details has a time gauge, and the spine's line adds· 13m leftin the last quarter of the limit, never earlier than the last quarter of an hour.Checked on production after the deploy: the database is at version 9 (backed up first), the console of run #36 loads, and its resumed
agent.test.browsershowed59m left of 1hin Step details. Attempts from before this carry no reason and no gauge.