Skip to content

diffbin: a --limit sample is not the corpus — the denominator was the cap, under a paragraph about denominators - #3196

Merged
gHashTag merged 2 commits into
masterfrom
loop/a-sample-is-not-the-corpus
Sep 4, 2026
Merged

gHashTag merged 2 commits into
masterfrom
loop/a-sample-is-not-the-corpus

Conversation

@gHashTag

@gHashTag gHashTag commented Sep 4, 2026

Copy link
Copy Markdown
Owner

Found by the CLI-wide fan-out in #3195, independently by two of the five lenses.

scripts/tri_loop/diffbin.py truncated its file list and then took its denominator from the
truncated list:

files.sort()
if limit:
    files = files[:limit]
...
total = len(files)          # <- AFTER the truncation
print(f"corpus: {len(files)} specs under {corpus}")
print(f"\nMEASURED COVERAGE: {measured}/{total} = {pct:.1f}% of the corpus")

Measured on the real tree: --limit 10 over specs printed corpus: 10 specs under specs.
There are 650 .t27 files there. A 2% sample would report 100.0% of the corpus on a clean run.

And the paragraph printed immediately below that number is about exactly this:

Any sentence of the form 'no regressions' is admissible only with this coverage figure attached
… Coverage below 100% bounds what the run can claim.

The one figure whose job is to bound the claim was the figure the truncation had already destroyed.

After

sample: 10 of 650 specs under specs (scratch excluded)  [--limit 10]
...
MEASURED COVERAGE: 0/10 = 0.0% of the 10 compared
  THIS IS A SAMPLE: 10 of 650 specs (1.5% of the corpus) were compared at all.
  Coverage above is of the sample. No sentence about the corpus is
  admissible from this run.

An untruncated run is unchanged: corpus: 650 specs under specs (scratch excluded).

Verification

scripts/ci/test_a_sample_is_not_the_corpus.py builds its own 25-file fixture with a fake binary
that never produces a verdict — irrelevant, since the subject is the denominator — so it needs no
compiler and runs in loop-tools-gate.yml beside the other compiler-less tests.

exit
against the pre-fix file (git show origin/master:…) 1 — corpus: 4 specs printed for a 25-file tree
against this branch 0
mutant: capture corpus_total after the truncation 1 — killed

The exclusion that cleared this file was true and too narrow

An earlier bounded-read audit named it and let it pass:

cost.py and diffbin.py take --limit N over a LOCAL corpus directory and never touch the API

Every word correct — and an argument about where the data comes from, used to settle a question
about what the label says.
A local --limit truncates the population exactly as thoroughly as a
page boundary, and the printed word "corpus" does not care which one did it.

An exclusion is only as wide as the reason given for it. A file on a list headed "checked"
repels examination in a way an unexamined file does not.

SKILL 543. Refs #3195

Found by the CLI-wide fan-out in #3195, independently by two lenses.

diffbin.py truncated its file list with `files = files[:limit]` and then
took `total = len(files)` from the truncated list, printing:

    corpus: {len(files)} specs under {corpus}
    MEASURED COVERAGE: {measured}/{total} = {pct}% of the corpus

Measured on the real tree: `--limit 10` over `specs` printed
"corpus: 10 specs under specs". There are 650 .t27 files there. A 2%
sample would report "100.0% of the corpus" on a clean run.

The paragraph printed immediately below that number reads "Coverage
below 100% bounds what the run can claim". The one figure whose job is
to bound the claim was the figure the truncation had already destroyed.

The corpus size is now captured before truncation. A sampled run prints
"sample: 10 of 650 specs under specs [--limit 10]", names the coverage
denominator "of the 10 compared", and adds a THIS IS A SAMPLE block where
the number is read rather than only in a header eight lines up. An
untruncated run is unchanged.

scripts/ci/test_a_sample_is_not_the_corpus.py builds its own 25-file
fixture with a fake binary that never produces a verdict -- irrelevant,
since the subject is the denominator -- so it needs no compiler and runs
in loop-tools-gate.yml. Exit 1 against the pre-fix file, exit 0 after,
and moving the capture below the truncation kills it.

An earlier bounded-read audit had cleared this file: "cost.py and
diffbin.py take --limit N over a LOCAL corpus directory and never touch
the API". True -- and an argument about where the data comes from, used
to settle a question about what the label says. An exclusion is only as
wide as the reason given for it.

SKILL 543.

Refs #3195
@github-actions

github-actions Bot commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

📓 NotebookLM Notebook linked to this PR

This notebook contains session context, decisions, and artifacts for this work.

@github-actions

github-actions Bot commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

PR Dashboard

Generated at: 2026-09-04 19:54:56 UTC

Summary

Status Count
Total Open PRs 12
PRs with Failing Checks 11
PRs with All Checks Green 1
READY 0
FAILING 11
PENDING 0

Seal Status

  • ⚠️ STALE -- sha256(compiler.rs)=9b8875f1c9d4 != manifest seal=87e5cbd3ad94.
    The committed NMSE numbers were certified against an older compiler.rs.
    Run scripts/reseal-check.sh locally for the two-step reseal command (advisory; not a merge gate).

Second confirmed finding from the #3195 fan-out.

`tri topic` printed:

    rows searched   759   (open PRs, open issues, last 40 commits, ...)

The parenthetical names the commit window and named no other bound, while
two of the four reads carried caps: `pr list --limit 100` and
`issue list --limit 200`.

Measured on gHashTag/t27 by raising each limit until the count stopped
growing: 12 open PRs, so that cap was slack -- and 509 open issues, so
the command searched 200 and never looked at 309. Raising both to 800
took the same invocation from 759 rows to 1068, and matches from 535 to
569: thirty-four pieces of prior art the tool exists to surface and could
not reach.

Disclosing one bound of three is worse than disclosing none. A reader who
sees "last 40 commits" learns that this command tells you where it stops,
and reasonably concludes the unmarked halves are unbounded.

Both caps are named constants; the request is built FROM the constant and
capped_read compares the row count against the same constant rather than
a second literal. A cap that BOUND is named where the population is
named, with a LOWER BOUND marker; a cap that did not bind is not
mentioned, because an unbound cap is not information.

Lowering ISSUE_CAP back to 200 leaves every test green, and that is
correct: the command would then print "the first 200 open issues ... A
CAP WAS REACHED", which is less complete and still honest. The guarantee
under test is "a cap that binds is named", and all three mutants against
that are killed. Pinning 800 in a test would defend a constant with no
argument behind it.

SKILL 544.

Refs #3195
@github-actions

github-actions Bot commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

📓 NotebookLM Notebook linked to this PR

This notebook contains session context, decisions, and artifacts for this work.

@github-actions

github-actions Bot commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

PR Dashboard

Generated at: 2026-09-04 19:58:37 UTC

Summary

Status Count
Total Open PRs 12
PRs with Failing Checks 11
PRs with All Checks Green 1
READY 0
FAILING 11
PENDING 0

Seal Status

  • ⚠️ STALE -- sha256(compiler.rs)=9b8875f1c9d4 != manifest seal=87e5cbd3ad94.
    The committed NMSE numbers were certified against an older compiler.rs.
    Run scripts/reseal-check.sh locally for the two-step reseal command (advisory; not a merge gate).

@gHashTag
gHashTag merged commit bbc64ab into master Sep 4, 2026
30 checks passed
@gHashTag
gHashTag deleted the loop/a-sample-is-not-the-corpus branch September 4, 2026 19:59
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant