Labels that name the wrong population: a dead test, a phantom test, and a gate that said every mutant survived - #3194
Merged
Conversation
…ery total Commit f7c1ff5, shipped by this loop an hour ago, carried two defects into master, and the pass that wrote it verified their absence and read a clean answer. The insert anchored on `fn the_query_and_the_marker_read_one_constant()`, whose `#[test]` sits ABOVE it, so the new text landed between the attribute and the function it belonged to. On f7c1ff5: the_query_and_the_marker_read_one_constant DOES NOT RUN cargo test <name> -> 0 passed; 685 filtered out the_freshness_boundary_is_pinned_on_both_sides RUNS TWICE cargo test -- --list prints it on two consecutive lines The check that missed it counted TOTALS. It printed "11 attrs, 11 fns", saw a match, and moved on -- but one function holding two attributes and one holding none leaves both totals unchanged, and so does the suite size, because the phantom fills the seat the dead test left. Two errors that cancel are invisible to every instrument that sums. tri gates tests [--gate] Rule 1: a test attribute followed by another, stepping over doc comments. Rule 2: a function in a #[cfg(test)] module with no test attribute, an assertion in its body, and its name nowhere else in the file. Rule 2's discriminator is the reference count, not the assertion -- a helper is called and appears twice, an orphan is called by nobody and appears once. Measured across all 57 files of cli/: exactly 1 of each, both real, and the three assert-bearing fixtures in test modules correctly silent. Positive control: exit 1 against f7c1ff5's tree, exit 0 against this one. Wired into cli-tri.yml. Both defects repaired here; the recovered test passes. The gate reproduced its own subject four times while being written, each caught by a test or by a comment already in the file: - the first structural test searched the file for a string it also contained, so the mutation it existed to catch made `find` fall through to the test's own body and pass. Now slices at #[cfg(test)]. - orphaned_tests first took "everything after the first #[cfg(test)]". test_module_lines, forty lines above it, already documents that as measurably wrong: five files here keep top-level fns after their test module, gates.rs fifteen. - it recognised only #[test] and only `fn `, so the 13 #[tokio::test] async functions in cli/trios-bridge were invisible in BOTH directions and read as clean -- the cancelling pair again, inside the gate. - the reference count used str::matches, so a fn named `a` is "referenced" by every assert. Now counts whole identifiers. Census re-blessed in this commit: fetches gates.rs:3094 -> 3106, a line shift from the inserted functions. No fetch added or removed. SKILL 537 and 538. Refs #3191
Contributor
|
📓 NotebookLM Notebook linked to this PR
This notebook contains session context, decisions, and artifacts for this work. |
Contributor
…built
Found by a fan-out over the whole CLI hunting the class from SKILL 535 --
a printed count whose label names a different population than the code
counts. 10 candidates, 8 survived adversarial refutation; this was the
strongest, and it lives in the command whose whole subject is whether a
claim was actually tested.
`claims_seen` counted every `# mutant-equivalent:` marker textually,
outside the per-direction loop. Its numerator `claims_broken` comes from
contradicted_claims, which drops any claimed line that is not a mutable
site in the direction being run. Two populations, one sentence:
N equivalence claim(s) in scope, none contradicted.
Each says its mutant cannot die, and each mutant survived.
Measured: all EIGHT markers in tools/ bind to a comparison, an assert, a
def or an assignment. The default operator is `silent`, whose sites are
`return <1..4>` lines only. Not one of the eight is reachable by it, so
every default run called them all in scope and reported that each mutant
survived -- while zero mutants had been built at any of them. The line
printed directly below says a claim is worth only the run that could have
refuted it. That run could not.
The refutation was already in the file, in the comment immediately above
the counter: "the marker names no direction, and every one in the tree
today argues about a comparison". The fact was written down and the next
line was written anyway.
Claims are now partitioned against the union of sites across the
operators actually run. In-scope claims keep the survivor sentence;
out-of-scope claims get their own paragraph naming each and saying no
mutant was built there. Verified live: `--only wp18_conformance_gate.py`
now prints "1 claim(s) NOT TESTED by this run" and "No claim was in scope
for this run", and the same gate under `--boundary` prints "1 equivalence
claim(s) in scope, none contradicted" -- truthfully, because that
operator does build a mutant at line 469.
The surviving mutant was the CALL SITE, not the helper: claims_by_scope
is covered three ways and reverting `claims_seen += in_scope.len()` to
add both halves restores the defect with all three still green. Killed by
a structural test whose needle is split across two literals, because the
first such test written this pass searched the file for a string it also
contained and passed against its own mutant.
SKILL 539.
Refs #3191
Contributor
|
📓 NotebookLM Notebook linked to this PR
This notebook contains session context, decisions, and artifacts for this work. |
Contributor
PR DashboardGenerated at: 2026-09-04 19:22:24 UTC
Summary
Seal Status
|
A concurrent pass landed its own 537, 538 and 539, so mine renumber to 540, 541 and 542 -- resolved BY TITLE, and the one internal cross- reference (my 542 cites my 541) renumbered with them. Two notes on the resolver itself. Master's own 539 is "My resolver committed conflict markers, in a file it never looked at", so this one checks for them before writing. The first version of that check searched for the marker as a plain substring and refused the merge: SKILL.md DOCUMENTS conflict markers, in prose and in code spans, because master's 539 is about them. Anchored to the start of a line the count is 0 before and 0 after. A guard loose enough to match its own subject blocks the work it was meant to protect. tools/census/fetches.txt also conflicted; taken from master and re-blessed here, since both sides had only moved line numbers. Refs #3191
Contributor
|
📓 NotebookLM Notebook linked to this PR
This notebook contains session context, decisions, and artifacts for this work. |
Contributor
PR DashboardGenerated at: 2026-09-04 19:26:33 UTC
Summary
Seal Status
|
… published SKILL 540 cited test_module_lines' doc comment: 'five files in this crate carry real top-level functions AFTER their test module, gates.rs fifteen of them'. Measured today: nine files, gates.rs sixty-eight. The rules differ -- mine counts every top-level fn after the FIRST test module closes, and a file with several test modules has many -- and the crate has grown since that comment was written. The load-bearing claim is unaffected and in fact stronger: 'everything after the first #[cfg(test)]' is wrong, which is why orphaned_tests uses test_module_lines instead. Recording the disagreement rather than repeating a figure I had not measured. Twice this pass a number I did not count reached a commit message, a PR body and the dashboard before an audit caught it. Refs #3191
Contributor
PR DashboardGenerated at: 2026-09-04 19:27:27 UTC
Summary
Seal Status
|
Contributor
|
📓 NotebookLM Notebook linked to this PR
This notebook contains session context, decisions, and artifacts for this work. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
f7c1ff5— merged by this loop an hour ago — carried two defects into master, and the pass thatwrote it verified their absence and read a clean answer.
The insert anchored on
fn the_query_and_the_marker_read_one_constant(), whose#[test]sitsabove it. The new text landed between the attribute and the function it belonged to:
f7c1ff5the_query_and_the_marker_read_one_constantcargo test <name>→0 passed; 685 filtered outthe_freshness_boundary_is_pinned_on_both_sidescargo test -- --listprints it on two consecutive linesThe check that missed it counted totals. It printed
#[test] attrs: 11 fn defs: 11, saw amatch, and moved on. One function holding two attributes and one holding none leaves both totals
unchanged — and so does the suite size, because the phantom fills the seat the dead test left, which
is why
675looked exactly right. Two errors that cancel are invisible to every instrument thatsums.
tri gates tests [--gate]#[cfg(test)]module with no test attribute, an assertion in its body, and itsname nowhere else in the file.
Rule 2's discriminator is the reference count, not the assertion. A helper exists to be called
and appears at least twice; a lost test is called by nobody and appears once. Assertion alone is far
too loose — an earlier attempt that way returned 18 candidates of which 16 were helpers.
Measured across all 57 files of
cli/: rule 1 finds exactly 1, rule 2 exactly 1, both real, andthe three assert-bearing fixtures in test modules are correctly silent. Positive control: exit 1
against
f7c1ff5's tree, exit 0 against this one. Wired intocli-tri.yml. Both defects arerepaired here and the recovered test passes; 695 crate tests.
The gate reproduced its own subject four times while being written
Each caught by a test or by a comment already in the file — none by review:
include_str!("red.rs")for a string italso contained as a literal, so the mutation it existed to catch made
findfall through to thetest's own body and pass. Now slices the source at
#[cfg(test)].orphaned_testsfirst took "everything after the first#[cfg(test)]".test_module_lines, forty lines above, already documents that as measurablywrong. That comment records "five files,
gates.rsfifteen"; measuring it today gives ninefiles and
gates.rssixty-eight — the rules differ and the crate has grown. Both say the sameload-bearing thing; the disagreement is recorded rather than the borrowed figure repeated.
#[test]and onlyfn, so the 13#[tokio::test]async functions incli/trios-bridgewere invisible in both directions and readas clean — the cancelling pair again, inside the gate itself.
str::matches, so a function namedais"referenced" by every
assert. Its own test caught it. Now counts whole identifiers.Also:
tri redcan now saynever greenlast_passwas requested only when the streak read was truncated — a question about whether thepage was full, not about whether the thing ever worked, and 43 of the 50 rows read
1 in a row, sothe majority were never asked. It is now requested for every red row (+1 request per red workflow,
~11% on
trinity-fpga), and the row names the branch, becauseno successis scoped to the branchthe runs were read from:
That mutation is invisible to every value-level test — the difference is a request that is or is not
made — and it survived until a structural test read the call site.
Refs #3191
And the same class, found by fanning out over the whole CLI
A workflow swept every
trimodule and the python loop tools for a printed count whose label namesa different population than the code counts — 10 candidates, 8 survived adversarial refutation. The
strongest is in the command whose entire subject is whether a claim was actually tested.
tri gates mutatereports on# mutant-equivalent:markers and printed:claims_seencounted every marker textually, outside the per-direction loop. Its numeratorclaims_brokencomes fromcontradicted_claims, which drops any claimed line that is not a mutablesite in the direction being run.
Measured: all eight markers in
tools/bind to a comparison, anassert, adefor anassignment. The default operator is
silent, whose sites arereturn <1..4>lines only. Not oneof the eight is reachable by it — so every default run called them all in scope and reported that
each mutant survived, while zero mutants had been built at any of them.
The refutation was already in the file, in the comment immediately above the counter: "the marker
names no direction, and every one in the tree today argues about a comparison". The fact was written
down and the next line was written anyway.
Claims are now partitioned against the union of sites across the operators actually run. Verified
live on
wp18_conformance_gate.py: the default run prints1 claim(s) NOT TESTED by this run/No claim was in scope, and the same gate under--boundaryprints1 equivalence claim(s) in scope, none contradicted— truthfully, because that operator does build a mutant at line 469.The surviving mutant was the call site, not the helper.
claims_by_scopeis covered three ways,and reverting
claims_seen += in_scope.len()to add both halves restores the defect with all threestill green — the same gap, in the same pass, as the
last_passguard above. Killed by a structuraltest whose needle is split across two literals, because the first such test written this pass
searched the file for a string it also contained and passed against its own mutant.
699 crate tests. SKILL 537, 538, 539.