Skip to content

Labels that name the wrong population: a dead test, a phantom test, and a gate that said every mutant survived - #3194

Merged
gHashTag merged 4 commits into
masterfrom
loop/a-dead-test-and-a-phantom-cancel
Sep 4, 2026
Merged

gHashTag merged 4 commits into
masterfrom
loop/a-dead-test-and-a-phantom-cancel

Conversation

@gHashTag

@gHashTag gHashTag commented Sep 4, 2026 •

Copy link
Copy Markdown
Owner

f7c1ff5 — merged by this loop an hour ago — carried two defects into master, and the pass that
wrote it verified their absence and read a clean answer.

The insert anchored on fn the_query_and_the_marker_read_one_constant(), whose #[test] sits
above it. The new text landed between the attribute and the function it belonged to:

on f7c1ff5
the_query_and_the_marker_read_one_constant does not run — cargo test <name> → 0 passed; 685 filtered out
the_freshness_boundary_is_pinned_on_both_sides runs twice — cargo test -- --list prints it on two consecutive lines

The check that missed it counted totals. It printed #[test] attrs: 11 fn defs: 11, saw a
match, and moved on. One function holding two attributes and one holding none leaves both totals
unchanged — and so does the suite size, because the phantom fills the seat the dead test left, which
is why 675 looked exactly right. Two errors that cancel are invisible to every instrument that
sums.

tri gates tests [--gate]

  1. a test attribute followed by another, stepping over doc comments;
  2. a function in a #[cfg(test)] module with no test attribute, an assertion in its body, and its
    name nowhere else in the file.

Rule 2's discriminator is the reference count, not the assertion. A helper exists to be called
and appears at least twice; a lost test is called by nobody and appears once. Assertion alone is far
too loose — an earlier attempt that way returned 18 candidates of which 16 were helpers.

Measured across all 57 files of cli/: rule 1 finds exactly 1, rule 2 exactly 1, both real, and
the three assert-bearing fixtures in test modules are correctly silent. Positive control: exit 1
against f7c1ff5's tree, exit 0 against this one.
Wired into cli-tri.yml. Both defects are
repaired here and the recovered test passes; 695 crate tests.

The gate reproduced its own subject four times while being written

Each caught by a test or by a comment already in the file — none by review:

  • It matched itself. The first structural test searched include_str!("red.rs") for a string it
    also contained as a literal, so the mutation it existed to catch made find fall through to the
    test's own body and pass. Now slices the source at #[cfg(test)].
  • The instrument was already in the file. orphaned_tests first took "everything after the first
    #[cfg(test)]". test_module_lines, forty lines above, already documents that as measurably
    wrong. That comment records "five files, gates.rs fifteen"; measuring it today gives nine
    files and gates.rs sixty-eight
    — the rules differ and the crate has grown. Both say the same
    load-bearing thing; the disagreement is recorded rather than the borrowed figure repeated.
  • Two blind spots that cancelled. It recognised only #[test] and only fn , so the 13
    #[tokio::test] async functions in cli/trios-bridge were invisible in both directions and read
    as clean — the cancelling pair again, inside the gate itself.
  • Substring, not token. The reference count used str::matches, so a function named a is
    "referenced" by every assert. Its own test caught it. Now counts whole identifiers.

Also: tri red can now say never green

last_pass was requested only when the streak read was truncated — a question about whether the
page was full
, not about whether the thing ever worked, and 43 of the 50 rows read 1 in a row, so
the majority were never asked. It is now requested for every red row (+1 request per red workflow,
~11% on trinity-fpga), and the row names the branch, because no success is scoped to the branch
the runs were read from:

50 workflow(s) red on the default branch -- 3 of them in the last 7 days, and 44 have never once been green.
    1 in a row  last run 2026-07-10T03:15  since 2026-07-10T03:15, never green on main   AX7203 Corona Compute ...

That mutation is invisible to every value-level test — the difference is a request that is or is not
made — and it survived until a structural test read the call site.

Refs #3191


And the same class, found by fanning out over the whole CLI

A workflow swept every tri module and the python loop tools for a printed count whose label names
a different population than the code counts
— 10 candidates, 8 survived adversarial refutation. The
strongest is in the command whose entire subject is whether a claim was actually tested.

tri gates mutate reports on # mutant-equivalent: markers and printed:

N equivalence claim(s) in scope, none contradicted.
Each says its mutant cannot die, and each mutant survived.

claims_seen counted every marker textually, outside the per-direction loop. Its numerator
claims_broken comes from contradicted_claims, which drops any claimed line that is not a mutable
site in the direction being run.

Measured: all eight markers in tools/ bind to a comparison, an assert, a def or an
assignment.
The default operator is silent, whose sites are return <1..4> lines only. Not one
of the eight is reachable by it
— so every default run called them all in scope and reported that
each mutant survived, while zero mutants had been built at any of them.

The refutation was already in the file, in the comment immediately above the counter: "the marker
names no direction, and every one in the tree today argues about a comparison"
. The fact was written
down and the next line was written anyway.

Claims are now partitioned against the union of sites across the operators actually run. Verified
live on wp18_conformance_gate.py: the default run prints 1 claim(s) NOT TESTED by this run /
No claim was in scope, and the same gate under --boundary prints 1 equivalence claim(s) in scope, none contradicted — truthfully, because that operator does build a mutant at line 469.

The surviving mutant was the call site, not the helper. claims_by_scope is covered three ways,
and reverting claims_seen += in_scope.len() to add both halves restores the defect with all three
still green — the same gap, in the same pass, as the last_pass guard above. Killed by a structural
test whose needle is split across two literals, because the first such test written this pass
searched the file for a string it also contained and passed against its own mutant.

699 crate tests. SKILL 537, 538, 539.

…ery total

Commit f7c1ff5, shipped by this loop an hour ago, carried two defects
into master, and the pass that wrote it verified their absence and read
a clean answer.

The insert anchored on `fn the_query_and_the_marker_read_one_constant()`,
whose `#[test]` sits ABOVE it, so the new text landed between the
attribute and the function it belonged to. On f7c1ff5:

  the_query_and_the_marker_read_one_constant   DOES NOT RUN
      cargo test <name> -> 0 passed; 685 filtered out
  the_freshness_boundary_is_pinned_on_both_sides   RUNS TWICE
      cargo test -- --list prints it on two consecutive lines

The check that missed it counted TOTALS. It printed "11 attrs, 11 fns",
saw a match, and moved on -- but one function holding two attributes and
one holding none leaves both totals unchanged, and so does the suite
size, because the phantom fills the seat the dead test left. Two errors
that cancel are invisible to every instrument that sums.

  tri gates tests [--gate]

Rule 1: a test attribute followed by another, stepping over doc comments.
Rule 2: a function in a #[cfg(test)] module with no test attribute, an
assertion in its body, and its name nowhere else in the file. Rule 2's
discriminator is the reference count, not the assertion -- a helper is
called and appears twice, an orphan is called by nobody and appears once.
Measured across all 57 files of cli/: exactly 1 of each, both real, and
the three assert-bearing fixtures in test modules correctly silent.
Positive control: exit 1 against f7c1ff5's tree, exit 0 against this one.
Wired into cli-tri.yml.

Both defects repaired here; the recovered test passes.

The gate reproduced its own subject four times while being written, each
caught by a test or by a comment already in the file:
  - the first structural test searched the file for a string it also
    contained, so the mutation it existed to catch made `find` fall
    through to the test's own body and pass. Now slices at #[cfg(test)].
  - orphaned_tests first took "everything after the first #[cfg(test)]".
    test_module_lines, forty lines above it, already documents that as
    measurably wrong: five files here keep top-level fns after their test
    module, gates.rs fifteen.
  - it recognised only #[test] and only `fn `, so the 13 #[tokio::test]
    async functions in cli/trios-bridge were invisible in BOTH directions
    and read as clean -- the cancelling pair again, inside the gate.
  - the reference count used str::matches, so a fn named `a` is
    "referenced" by every assert. Now counts whole identifiers.

Census re-blessed in this commit: fetches gates.rs:3094 -> 3106, a line
shift from the inserted functions. No fetch added or removed.

SKILL 537 and 538.

Refs #3191
@github-actions

github-actions Bot commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

📓 NotebookLM Notebook linked to this PR

This notebook contains session context, decisions, and artifacts for this work.

@github-actions

github-actions Bot commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

PR Dashboard

Generated at: 2026-09-04 19:12:00 UTC

Summary

Status Count
Total Open PRs 13
PRs with Failing Checks 11
PRs with All Checks Green 2
READY 0
FAILING 11
PENDING 0

Seal Status

  • ⚠️ STALE -- sha256(compiler.rs)=9b8875f1c9d4 != manifest seal=87e5cbd3ad94.
    The committed NMSE numbers were certified against an older compiler.rs.
    Run scripts/reseal-check.sh locally for the two-step reseal command (advisory; not a merge gate).

…built

Found by a fan-out over the whole CLI hunting the class from SKILL 535 --
a printed count whose label names a different population than the code
counts. 10 candidates, 8 survived adversarial refutation; this was the
strongest, and it lives in the command whose whole subject is whether a
claim was actually tested.

`claims_seen` counted every `# mutant-equivalent:` marker textually,
outside the per-direction loop. Its numerator `claims_broken` comes from
contradicted_claims, which drops any claimed line that is not a mutable
site in the direction being run. Two populations, one sentence:

    N equivalence claim(s) in scope, none contradicted.
    Each says its mutant cannot die, and each mutant survived.

Measured: all EIGHT markers in tools/ bind to a comparison, an assert, a
def or an assignment. The default operator is `silent`, whose sites are
`return <1..4>` lines only. Not one of the eight is reachable by it, so
every default run called them all in scope and reported that each mutant
survived -- while zero mutants had been built at any of them. The line
printed directly below says a claim is worth only the run that could have
refuted it. That run could not.

The refutation was already in the file, in the comment immediately above
the counter: "the marker names no direction, and every one in the tree
today argues about a comparison". The fact was written down and the next
line was written anyway.

Claims are now partitioned against the union of sites across the
operators actually run. In-scope claims keep the survivor sentence;
out-of-scope claims get their own paragraph naming each and saying no
mutant was built there. Verified live: `--only wp18_conformance_gate.py`
now prints "1 claim(s) NOT TESTED by this run" and "No claim was in scope
for this run", and the same gate under `--boundary` prints "1 equivalence
claim(s) in scope, none contradicted" -- truthfully, because that
operator does build a mutant at line 469.

The surviving mutant was the CALL SITE, not the helper: claims_by_scope
is covered three ways and reverting `claims_seen += in_scope.len()` to
add both halves restores the defect with all three still green. Killed by
a structural test whose needle is split across two literals, because the
first such test written this pass searched the file for a string it also
contained and passed against its own mutant.

SKILL 539.

Refs #3191
@github-actions

github-actions Bot commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

📓 NotebookLM Notebook linked to this PR

This notebook contains session context, decisions, and artifacts for this work.

@gHashTag gHashTag changed the title tri gates tests: a dead test and a phantom test cancel in every total — both shipped in f7c1ff5 Labels that name the wrong population: a dead test, a phantom test, and a gate that said every mutant survived Sep 4, 2026
@github-actions

github-actions Bot commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

PR Dashboard

Generated at: 2026-09-04 19:22:24 UTC

Summary

Status Count
Total Open PRs 13
PRs with Failing Checks 11
PRs with All Checks Green 2
READY 1
FAILING 11
PENDING 0

Seal Status

  • ⚠️ STALE -- sha256(compiler.rs)=9b8875f1c9d4 != manifest seal=87e5cbd3ad94.
    The committed NMSE numbers were certified against an older compiler.rs.
    Run scripts/reseal-check.sh locally for the two-step reseal command (advisory; not a merge gate).

A concurrent pass landed its own 537, 538 and 539, so mine renumber to
540, 541 and 542 -- resolved BY TITLE, and the one internal cross-
reference (my 542 cites my 541) renumbered with them.

Two notes on the resolver itself.

Master's own 539 is "My resolver committed conflict markers, in a file it
never looked at", so this one checks for them before writing. The first
version of that check searched for the marker as a plain substring and
refused the merge: SKILL.md DOCUMENTS conflict markers, in prose and in
code spans, because master's 539 is about them. Anchored to the start of
a line the count is 0 before and 0 after. A guard loose enough to match
its own subject blocks the work it was meant to protect.

tools/census/fetches.txt also conflicted; taken from master and
re-blessed here, since both sides had only moved line numbers.

Refs #3191
@github-actions

github-actions Bot commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

📓 NotebookLM Notebook linked to this PR

This notebook contains session context, decisions, and artifacts for this work.

@github-actions

github-actions Bot commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

PR Dashboard

Generated at: 2026-09-04 19:26:33 UTC

Summary

Status Count
Total Open PRs 12
PRs with Failing Checks 11
PRs with All Checks Green 1
READY 0
FAILING 11
PENDING 0

Seal Status

  • ⚠️ STALE -- sha256(compiler.rs)=9b8875f1c9d4 != manifest seal=87e5cbd3ad94.
    The committed NMSE numbers were certified against an older compiler.rs.
    Run scripts/reseal-check.sh locally for the two-step reseal command (advisory; not a merge gate).

… published

SKILL 540 cited test_module_lines' doc comment: 'five files in this
crate carry real top-level functions AFTER their test module, gates.rs
fifteen of them'. Measured today: nine files, gates.rs sixty-eight.

The rules differ -- mine counts every top-level fn after the FIRST test
module closes, and a file with several test modules has many -- and the
crate has grown since that comment was written. The load-bearing claim is
unaffected and in fact stronger: 'everything after the first #[cfg(test)]'
is wrong, which is why orphaned_tests uses test_module_lines instead.

Recording the disagreement rather than repeating a figure I had not
measured. Twice this pass a number I did not count reached a commit
message, a PR body and the dashboard before an audit caught it.

Refs #3191
@github-actions

github-actions Bot commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

PR Dashboard

Generated at: 2026-09-04 19:27:27 UTC

Summary

Status Count
Total Open PRs 12
PRs with Failing Checks 11
PRs with All Checks Green 1
READY 0
FAILING 11
PENDING 0

Seal Status

  • ⚠️ STALE -- sha256(compiler.rs)=9b8875f1c9d4 != manifest seal=87e5cbd3ad94.
    The committed NMSE numbers were certified against an older compiler.rs.
    Run scripts/reseal-check.sh locally for the two-step reseal command (advisory; not a merge gate).

@github-actions

github-actions Bot commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

📓 NotebookLM Notebook linked to this PR

This notebook contains session context, decisions, and artifacts for this work.

@gHashTag
gHashTag merged commit de8757a into master Sep 4, 2026
38 checks passed
@gHashTag
gHashTag deleted the loop/a-dead-test-and-a-phantom-cancel branch September 4, 2026 19:49
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant