Status light says offline when 5+ agent steps stream at once: the browser's 6-connection limit is used up #48

Closed
opened 2026-09-26 12:46:18 -04:00 by cmoriarty · 1 comment
Owner

The status light says offline while osfd is healthy and the run is progressing.

What happened

On 2026-09-26 at 12:39, run run_01M3F8PTF3EWBCB721BP869E9N (soundcheck#32, minimalist) fanned out to five agent steps at once: four openspec.propose.spec[...] steps and openspec.propose.design. The web app then reported the backend offline. At the same moment:

  • curl /api/status answered 200 in 13 ms, with state: running and the backend reachable
  • /healthz was ok: true and showed sse: {a: 1, b: 5}

Cause

osfd is served by uvicorn over HTTP/1.1 (no HTTP/2, no proxy), and browsers allow at most 6 connections per host over HTTP/1.1.

The UI holds two kinds of long-lived connection:

  • stream A, /api/stream: one
  • stream B, /api/steps/{id}/deltas: one per running agent step on the open run, capped by MAX_LIVE_STREAMS = 8 (ui/src/state/store.ts, syncLiveStreams)

With five agent steps running, that is 1 + 5 = 6 open streams, which uses every connection the browser has. Every ordinary request then waits for a free connection, including the 5 s /api/status poll (useStatus in ui/src/components/chrome.tsx). After OSFD_UNREACHABLE_MS = 12_000 without an answer, braidState returns offline. Snapshots, gate answers and composer sends would stall the same way.

It clears once fewer than five agent steps run, or when the tab is closed or switched to another run. The fan-out in minimalist and in the default pipeline's review lenses can reach this, and the cap of 8 assumed it never would.

Fix options

  1. Quick: keep long-lived connections at or below 4, leaving 2 for requests. Cap stream B at 3, always including the focused step; the rest fall back to the 4 Hz digest.
  2. Proper: one multiplexed live-text stream, e.g. /api/deltas?steps=a,b,c, reopened when the set changes. That is two connections total regardless of fan-out, and every running step stays live. It needs a server-side change.
  3. Also possible: serve over HTTP/2 (e.g. behind a proxy), which removes the 6-connection limit.

Done when

  • With 6 agent steps running on the open run, /api/status keeps answering in the browser and the light never says offline.
  • A Playwright test drives a run with 6+ running agent steps (the fake osfd) and asserts the light stays connected and a gate answer still goes through.
The status light says **offline** while osfd is healthy and the run is progressing. ## What happened On 2026-09-26 at 12:39, run `run_01M3F8PTF3EWBCB721BP869E9N` (soundcheck#32, `minimalist`) fanned out to five agent steps at once: four `openspec.propose.spec[...]` steps and `openspec.propose.design`. The web app then reported the backend offline. At the same moment: - `curl /api/status` answered 200 in 13 ms, with `state: running` and the backend reachable - `/healthz` was `ok: true` and showed `sse: {a: 1, b: 5}` ## Cause osfd is served by uvicorn over **HTTP/1.1** (no HTTP/2, no proxy), and browsers allow at most **6 connections per host** over HTTP/1.1. The UI holds two kinds of long-lived connection: - stream A, `/api/stream`: one - stream B, `/api/steps/{id}/deltas`: one per running agent step on the open run, capped by `MAX_LIVE_STREAMS = 8` (`ui/src/state/store.ts`, `syncLiveStreams`) With five agent steps running, that is 1 + 5 = 6 open streams, which uses every connection the browser has. Every ordinary request then waits for a free connection, including the 5 s `/api/status` poll (`useStatus` in `ui/src/components/chrome.tsx`). After `OSFD_UNREACHABLE_MS = 12_000` without an answer, `braidState` returns `offline`. Snapshots, gate answers and composer sends would stall the same way. It clears once fewer than five agent steps run, or when the tab is closed or switched to another run. The fan-out in `minimalist` and in the default pipeline's review lenses can reach this, and the cap of 8 assumed it never would. ## Fix options 1. **Quick:** keep long-lived connections at or below 4, leaving 2 for requests. Cap stream B at 3, always including the focused step; the rest fall back to the 4 Hz digest. 2. **Proper:** one multiplexed live-text stream, e.g. `/api/deltas?steps=a,b,c`, reopened when the set changes. That is two connections total regardless of fan-out, and every running step stays live. It needs a server-side change. 3. Also possible: serve over HTTP/2 (e.g. behind a proxy), which removes the 6-connection limit. ## Done when - With 6 agent steps running on the open run, `/api/status` keeps answering in the browser and the light never says offline. - A Playwright test drives a run with 6+ running agent steps (the fake osfd) and asserts the light stays connected and a gate answer still goes through.
Author
Owner

Shipped in #68 (7c715ec, archived as 2026-09-26-one-delta-stream), deployed as fea8788 at 23:11 EDT.

Hit on production on 2026-09-26 (run_01M3GB3KN38YT9TZJ7KA7RQ5M0): a fan-out of five agent steps used all six connections. The light said offline, steps showed as silent, step details failed with "signal is aborted without reason", and the page wouldn't reload. A second fault made it last: a step's delta stream never ended with the step, and /healthz showed five delta streams for finished steps.

  • GET /api/deltas?steps=a,b,c carries every running agent step's live text on one connection (option 2 above). A page holds two long-lived connections however many agents run.
  • Delta streams end with an end frame once none of their steps can still write, checked at connect and every heartbeat, and the console does not reconnect. The one-step route stays for pages loaded before the deploy, and ends the same way.
  • A request osfd doesn't answer in time says so, and suggests closing other Braid tabs.
  • e2e/connection-budget.spec.ts runs six agents at once on the fixture and asserts one delta stream plus answering requests. Chromium enforces the six-connection limit, so it reproduces the bug for real.

Checked on production after the deploy: a stream for a finished step answered end and closed in 9 ms, and /healthz showed sse: {a: 0, b: 0}.

Shipped in #68 (`7c715ec`, archived as `2026-09-26-one-delta-stream`), deployed as `fea8788` at 23:11 EDT. Hit on production on 2026-09-26 (`run_01M3GB3KN38YT9TZJ7KA7RQ5M0`): a fan-out of five agent steps used all six connections. The light said offline, steps showed as silent, step details failed with "signal is aborted without reason", and the page wouldn't reload. A second fault made it last: a step's delta stream never ended with the step, and `/healthz` showed five delta streams for finished steps. - `GET /api/deltas?steps=a,b,c` carries every running agent step's live text on **one** connection (option 2 above). A page holds two long-lived connections however many agents run. - Delta streams end with an `end` frame once none of their steps can still write, checked at connect and every heartbeat, and the console does not reconnect. The one-step route stays for pages loaded before the deploy, and ends the same way. - A request osfd doesn't answer in time says so, and suggests closing other Braid tabs. - `e2e/connection-budget.spec.ts` runs six agents at once on the fixture and asserts one delta stream plus answering requests. Chromium enforces the six-connection limit, so it reproduces the bug for real. Checked on production after the deploy: a stream for a finished step answered `end` and closed in 9 ms, and `/healthz` showed `sse: {a: 0, b: 0}`.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
cmoriarty/braid#48
No description provided.