Two needs_opencode live tests are stale: the agent set predates the subagents, and the shape floors have drifted #97

Closed
opened 2026-09-27 18:09:40 -04:00 by cmoriarty · 1 comment
Owner

With the pinned opencode 1.18.21 on PATH, pytest -m needs_opencode tests/integration/ has two failures on main. Both fail the same way before and after #89, so they are stale tests, not regressions.

  1. test_only_the_builtin_agents_exist_in_an_isolated_home expects agent names to be a subset of opencode's own built-in agents. Since 412cea4 (2026-09-13), the run config also declares Braid's reader and test subagents, so the set is larger. The allowed set should include the roster the config declares (SUBAGENT_NAMES), and the test should still catch an operator agent leaking in.

  2. test_each_shape_puts_its_measured_tools_on_the_wire measures the implement floor at 33,195 characters against MEASURED_FLOOR_CHARS[IMPL] = 32,341 (tolerance 200). The lean floor sits at 22,736 against 22,545, just inside the tolerance. Re-measured per tool:

    • task: 3,899 → 4,562. Its description now lists the explore, reader and test roster (412cea4).
    • bash: 5,312 → 5,425. Its description now names the run's own TMPDIR (#36), so the floor moves with the length of the run's path.
    • edit still measures 3,049, which is edit plus write (1,996 + 1,053) because turning edit off removes write too. It is unchanged.

The fix updates MEASURED_FLOOR_CHARS, TOOL_COST_CHARS and the osf.engine.shape docstring table, and the unit tests derived from them. The lean saving is now 10,459 characters, or 3,486 tokens a turn. shape.floor_tokens is used only by tests, so budgeting doesn't change. config.FLOOR_LEAN/FLOOR_IMPL (7,515 / 11,009), the budget estimator's inputs, come from a separate osfd measurement and are left as they are.

With the pinned opencode 1.18.21 on PATH, `pytest -m needs_opencode tests/integration/` has two failures on main. Both fail the same way before and after #89, so they are stale tests, not regressions. 1. **`test_only_the_builtin_agents_exist_in_an_isolated_home`** expects agent names to be a subset of opencode's own built-in agents. Since 412cea4 (2026-09-13), the run config also declares Braid's `reader` and `test` subagents, so the set is larger. The allowed set should include the roster the config declares (`SUBAGENT_NAMES`), and the test should still catch an operator agent leaking in. 2. **`test_each_shape_puts_its_measured_tools_on_the_wire`** measures the implement floor at 33,195 characters against `MEASURED_FLOOR_CHARS[IMPL]` = 32,341 (tolerance 200). The lean floor sits at 22,736 against 22,545, just inside the tolerance. Re-measured per tool: - `task`: 3,899 → 4,562. Its description now lists the `explore`, `reader` and `test` roster (412cea4). - `bash`: 5,312 → 5,425. Its description now names the run's own `TMPDIR` (#36), so the floor moves with the length of the run's path. - `edit` still measures 3,049, which is `edit` plus `write` (1,996 + 1,053) because turning `edit` off removes `write` too. It is unchanged. The fix updates `MEASURED_FLOOR_CHARS`, `TOOL_COST_CHARS` and the `osf.engine.shape` docstring table, and the unit tests derived from them. The lean saving is now 10,459 characters, or 3,486 tokens a turn. `shape.floor_tokens` is used only by tests, so budgeting doesn't change. `config.FLOOR_LEAN`/`FLOOR_IMPL` (7,515 / 11,009), the budget estimator's inputs, come from a separate osfd measurement and are left as they are.
Author
Owner

Fixed and pushed to main in 2ee5069.

Rebasing onto main brought in #17, which made a third test stale. The numbers here supersede the ones in the description:

  • test_only_the_builtin_agents_exist_in_an_isolated_home now allows opencode's own agents plus SUBAGENT_NAMES (reader, explore, test, browser).
  • test_the_run_is_isolated_from_the_operators_opencode and test_operator_config_leaks_in_without_home_isolation now expect the run's own Playwright MCP server (RUN_MCP = {"playwright"}, from #17) and nothing else.
  • test_each_shape_puts_its_measured_tools_on_the_wire: re-measured on 1.18.21 under the test's own path.
    • The implement floor is 33,606 (was 32,341) and the lean floor is 22,736 (was 22,545).
    • task is 4,973, because its description lists the four-subagent roster. bash is 5,445, because its description names the run's TMPDIR.
    • The lean saving is 10,870 characters, or 3,623 tokens a turn, which equals impl − lean exactly.
    • The failure message now prints the measured floor.

config.FLOOR_LEAN/FLOOR_IMPL, the budget estimator's inputs, are unchanged.

Results: -m needs_opencode tests/integration/ passes 23 of 23 with 1.18.21, and ./tools/test.sh full passes (backend 1,858, UI 444, browser 108; no osfd, so no @live specs).

Fixed and pushed to main in 2ee5069. Rebasing onto main brought in #17, which made a third test stale. The numbers here supersede the ones in the description: - **`test_only_the_builtin_agents_exist_in_an_isolated_home`** now allows opencode's own agents plus `SUBAGENT_NAMES` (`reader`, `explore`, `test`, `browser`). - **`test_the_run_is_isolated_from_the_operators_opencode`** and **`test_operator_config_leaks_in_without_home_isolation`** now expect the run's own Playwright MCP server (`RUN_MCP = {"playwright"}`, from #17) and nothing else. - **`test_each_shape_puts_its_measured_tools_on_the_wire`**: re-measured on 1.18.21 under the test's own path. - The implement floor is **33,606** (was 32,341) and the lean floor is **22,736** (was 22,545). - `task` is 4,973, because its description lists the four-subagent roster. `bash` is 5,445, because its description names the run's `TMPDIR`. - The lean saving is 10,870 characters, or 3,623 tokens a turn, which equals impl − lean exactly. - The failure message now prints the measured floor. `config.FLOOR_LEAN`/`FLOOR_IMPL`, the budget estimator's inputs, are unchanged. Results: `-m needs_opencode tests/integration/` passes 23 of 23 with 1.18.21, and `./tools/test.sh full` passes (backend 1,858, UI 444, browser 108; no osfd, so no `@live` specs).
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
cmoriarty/braid#97
No description provided.