The test lint and scenario coverage decide nothing, so a run can merge flagged tests and untested scenarios #138
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Since #137,
agent.test.unitandagent.test.e2ecomplete onlane-passed, which is the suite's exit code on the current code. The lane report still computes the test lint (B1–B7), scenario coverage and mutation, and the step's summary shows them, but none of them decides anything. #137 left that to a follow-up, to be judged by the measurement it took.What the measurement says
These are runs 22 to 41 on production, recorded in
openspec/changes/archive/2026-10-03-test-lanes-hold-on-the-verdict/design.md.expect(parseState(…)).not.toBeNull()as the only check of "accepts an empty seat"To decide first: does a test a run only edited owe a tag?
B6 judges the tests a run added or changed. In runs 36 to 39, 27 of its 62 findings were older tests that the run only edited, for instance by adding a prop to every
render. There are two options:Proposal
ban. It should name each finding with its file, line and rule, as it already does for a malformed report's schema errors.agent.test.unitandagent.test.e2etosuites-green:<lane>. That check already re-derives the lint over the run's own lines (lint_base), and checks coverage and mutation. A failing verdict is tried again twice, briefed with every reason, then goes to a person.Related: #137, and #125 (the verification steps).
Thinking on the open question (should a test a run only edited still need a tag?), from a review of the hello world baseline in #139.
What a tag is for. A
// spec:tag gives traceability: which scenarios have no test, and which tests belong to no scenario. It does not show that a test is meaningful. A test that asserts nothing useful carries a tag as easily as a good one, and a rule that checks for a tag rewards adding one. That is the same proxy trap #139 found with the "107#on screen" oracle, which passed seven times without checking anything real.So:
One-time audit of the older untagged tests, separate from the rule. Run mutation on them first, so "meaningful" is measured, then review the survivors and decide each one:
The audit needs to say that deleting and "keep but flag" are acceptable outcomes. Otherwise the easy way to quiet the lint is to tag everything.
Not yet checked: how many of the older untagged tests fall into each group, and whether this repository's own tests have the same gap as the scratch app's.