The test lint and scenario coverage decide nothing, so a run can merge flagged tests and untested scenarios #138

Open
opened 2026-10-03 11:50:56 -04:00 by cmoriarty · 1 comment
Owner

Since #137, agent.test.unit and agent.test.e2e complete on lane-passed, which is the suite's exit code on the current code. The lane report still computes the test lint (B1–B7), scenario coverage and mutation, and the step's summary shows them, but none of them decides anything. #137 left that to a follow-up, to be judged by the measurement it took.

What the measurement says

These are runs 22 to 41 on production, recorded in openspec/changes/archive/2026-10-03-test-lanes-hold-on-the-verdict/design.md.

Check Findings over runs 22–41, with the lint as deployed Ready to decide?
B1, B2 (only a spy, only existence) 1, which is the one weak test: expect(parseState(…)).not.toBeNull() as the only check of "accepts an empty seat" yes
B3, B4, B5, B7 0 yes
B6 (each test the run wrote or changed carries a spec tag) 118, all exact; run 41 had 0 after the question below
Scenario coverage 14 to 55 untested scenarios in runs 22 to 37, before tags were counted; 0 in runs 39 and 41 likely
Mutation not run on any of these runs not measured

To decide first: does a test a run only edited owe a tag?

B6 judges the tests a run added or changed. In runs 36 to 39, 27 of its 62 findings were older tests that the run only edited, for instance by adding a prop to every render. There are two options:

  • Judge only the tests a run added. A run never answers for its predecessors' tags.
  • Keep judging edited tests too. The untagged backlog gets tagged as files are touched, at the cost of a retry for the run.

Proposal

  1. Settle B6's scope, as above.
  2. Make a retry able to fix what the lint flags. Today the briefing says "3 ban violation(s) on disk: B2, B6", and the details are only in the report's ban. It should name each finding with its file, line and rule, as it already does for a malformed report's schema errors.
  3. Switch agent.test.unit and agent.test.e2e to suites-green:<lane>. That check already re-derives the lint over the run's own lines (lint_base), and checks coverage and mutation. A failing verdict is tried again twice, briefed with every reason, then goes to a person.
  4. On the next scratch run, watch how many retries the lint causes and whether they converge.

Related: #137, and #125 (the verification steps).

Since #137, `agent.test.unit` and `agent.test.e2e` complete on `lane-passed`, which is the suite's exit code on the current code. The lane report still computes the test lint (B1–B7), scenario coverage and mutation, and the step's summary shows them, but none of them decides anything. #137 left that to a follow-up, to be judged by the measurement it took. ## What the measurement says These are runs 22 to 41 on production, recorded in `openspec/changes/archive/2026-10-03-test-lanes-hold-on-the-verdict/design.md`. | Check | Findings over runs 22–41, with the lint as deployed | Ready to decide? | |---|---|---| | B1, B2 (only a spy, only existence) | 1, which is the one weak test: `expect(parseState(…)).not.toBeNull()` as the only check of "accepts an empty seat" | yes | | B3, B4, B5, B7 | 0 | yes | | B6 (each test the run wrote or changed carries a spec tag) | 118, all exact; run 41 had 0 | after the question below | | Scenario coverage | 14 to 55 untested scenarios in runs 22 to 37, before tags were counted; 0 in runs 39 and 41 | likely | | Mutation | not run on any of these runs | not measured | ## To decide first: does a test a run only edited owe a tag? B6 judges the tests a run added *or changed*. In runs 36 to 39, 27 of its 62 findings were older tests that the run only edited, for instance by adding a prop to every `render`. There are two options: - **Judge only the tests a run added.** A run never answers for its predecessors' tags. - **Keep judging edited tests too.** The untagged backlog gets tagged as files are touched, at the cost of a retry for the run. ## Proposal 1. Settle B6's scope, as above. 2. Make a retry able to fix what the lint flags. Today the briefing says "3 ban violation(s) on disk: B2, B6", and the details are only in the report's `ban`. It should name each finding with its file, line and rule, as it already does for a malformed report's schema errors. 3. Switch `agent.test.unit` and `agent.test.e2e` to `suites-green:<lane>`. That check already re-derives the lint over the run's own lines (`lint_base`), and checks coverage and mutation. A failing verdict is tried again twice, briefed with every reason, then goes to a person. 4. On the next scratch run, watch how many retries the lint causes and whether they converge. Related: #137, and #125 (the verification steps).
Author
Owner

Thinking on the open question (should a test a run only edited still need a tag?), from a review of the hello world baseline in #139.

What a tag is for. A // spec: tag gives traceability: which scenarios have no test, and which tests belong to no scenario. It does not show that a test is meaningful. A test that asserts nothing useful carries a tag as easily as a good one, and a rule that checks for a tag rewards adding one. That is the same proxy trap #139 found with the "107 # on screen" oracle, which passed seven times without checking anything real.

So:

  1. Treat tags as a coverage report (is every scenario tested at all?), not as a measure of test quality, and don't block on it beyond the narrow case below.
  2. Require a tag only on tests the run adds. A test it only edited keeps whatever it had. This stops runs being blocked for gaps they did not create (27 of the 62 findings in runs 36 to 39).
  3. Put the blocking weight on what measures meaningfulness: the weak-assertion rules (B1, B2), which this issue already says are ready to block, and the mutation check, which has never run on these runs. Mutation (do the tests fail when the code is broken?) is the closest thing the system has to "meaningful".

One-time audit of the older untagged tests, separate from the rule. Run mutation on them first, so "meaningful" is measured, then review the survivors and decide each one:

  • meaningful and maps to a scenario: add the tag;
  • meaningful but maps to no scenario: keep it and note the spec gap;
  • frivolous: delete it.

The audit needs to say that deleting and "keep but flag" are acceptable outcomes. Otherwise the easy way to quiet the lint is to tag everything.

Not yet checked: how many of the older untagged tests fall into each group, and whether this repository's own tests have the same gap as the scratch app's.

Thinking on the open question (should a test a run only edited still need a tag?), from a review of the hello world baseline in #139. **What a tag is for.** A `// spec:` tag gives traceability: which scenarios have no test, and which tests belong to no scenario. It does not show that a test is meaningful. A test that asserts nothing useful carries a tag as easily as a good one, and a rule that checks for a tag rewards adding one. That is the same proxy trap #139 found with the "107 `#` on screen" oracle, which passed seven times without checking anything real. **So:** 1. Treat tags as a coverage report (is every scenario tested at all?), not as a measure of test quality, and don't block on it beyond the narrow case below. 2. Require a tag only on tests the run **adds**. A test it only edited keeps whatever it had. This stops runs being blocked for gaps they did not create (27 of the 62 findings in runs 36 to 39). 3. Put the blocking weight on what measures meaningfulness: the weak-assertion rules (B1, B2), which this issue already says are ready to block, and the mutation check, which has never run on these runs. Mutation (do the tests fail when the code is broken?) is the closest thing the system has to "meaningful". **One-time audit of the older untagged tests**, separate from the rule. Run mutation on them first, so "meaningful" is measured, then review the survivors and decide each one: - meaningful and maps to a scenario: add the tag; - meaningful but maps to no scenario: keep it and note the spec gap; - frivolous: delete it. The audit needs to say that deleting and "keep but flag" are acceptable outcomes. Otherwise the easy way to quiet the lint is to tag everything. Not yet checked: how many of the older untagged tests fall into each group, and whether this repository's own tests have the same gap as the scratch app's.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
cmoriarty/braid#138
No description provided.