diffbin: a --limit sample is not the corpus — the denominator was the cap, under a paragraph about denominators - #3196
Merged
Conversation
Found by the CLI-wide fan-out in #3195, independently by two lenses. diffbin.py truncated its file list with `files = files[:limit]` and then took `total = len(files)` from the truncated list, printing: corpus: {len(files)} specs under {corpus} MEASURED COVERAGE: {measured}/{total} = {pct}% of the corpus Measured on the real tree: `--limit 10` over `specs` printed "corpus: 10 specs under specs". There are 650 .t27 files there. A 2% sample would report "100.0% of the corpus" on a clean run. The paragraph printed immediately below that number reads "Coverage below 100% bounds what the run can claim". The one figure whose job is to bound the claim was the figure the truncation had already destroyed. The corpus size is now captured before truncation. A sampled run prints "sample: 10 of 650 specs under specs [--limit 10]", names the coverage denominator "of the 10 compared", and adds a THIS IS A SAMPLE block where the number is read rather than only in a header eight lines up. An untruncated run is unchanged. scripts/ci/test_a_sample_is_not_the_corpus.py builds its own 25-file fixture with a fake binary that never produces a verdict -- irrelevant, since the subject is the denominator -- so it needs no compiler and runs in loop-tools-gate.yml. Exit 1 against the pre-fix file, exit 0 after, and moving the capture below the truncation kills it. An earlier bounded-read audit had cleared this file: "cost.py and diffbin.py take --limit N over a LOCAL corpus directory and never touch the API". True -- and an argument about where the data comes from, used to settle a question about what the label says. An exclusion is only as wide as the reason given for it. SKILL 543. Refs #3195
Contributor
|
📓 NotebookLM Notebook linked to this PR
This notebook contains session context, decisions, and artifacts for this work. |
Contributor
Second confirmed finding from the #3195 fan-out. `tri topic` printed: rows searched 759 (open PRs, open issues, last 40 commits, ...) The parenthetical names the commit window and named no other bound, while two of the four reads carried caps: `pr list --limit 100` and `issue list --limit 200`. Measured on gHashTag/t27 by raising each limit until the count stopped growing: 12 open PRs, so that cap was slack -- and 509 open issues, so the command searched 200 and never looked at 309. Raising both to 800 took the same invocation from 759 rows to 1068, and matches from 535 to 569: thirty-four pieces of prior art the tool exists to surface and could not reach. Disclosing one bound of three is worse than disclosing none. A reader who sees "last 40 commits" learns that this command tells you where it stops, and reasonably concludes the unmarked halves are unbounded. Both caps are named constants; the request is built FROM the constant and capped_read compares the row count against the same constant rather than a second literal. A cap that BOUND is named where the population is named, with a LOWER BOUND marker; a cap that did not bind is not mentioned, because an unbound cap is not information. Lowering ISSUE_CAP back to 200 leaves every test green, and that is correct: the command would then print "the first 200 open issues ... A CAP WAS REACHED", which is less complete and still honest. The guarantee under test is "a cap that binds is named", and all three mutants against that are killed. Pinning 800 in a test would defend a constant with no argument behind it. SKILL 544. Refs #3195
Contributor
|
📓 NotebookLM Notebook linked to this PR
This notebook contains session context, decisions, and artifacts for this work. |
Contributor
PR DashboardGenerated at: 2026-09-04 19:58:37 UTC
Summary
Seal Status
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Found by the CLI-wide fan-out in #3195, independently by two of the five lenses.
scripts/tri_loop/diffbin.pytruncated its file list and then took its denominator from thetruncated list:
Measured on the real tree:
--limit 10overspecsprintedcorpus: 10 specs under specs.There are 650
.t27files there. A 2% sample would report100.0% of the corpuson a clean run.And the paragraph printed immediately below that number is about exactly this:
The one figure whose job is to bound the claim was the figure the truncation had already destroyed.
After
An untruncated run is unchanged:
corpus: 650 specs under specs (scratch excluded).Verification
scripts/ci/test_a_sample_is_not_the_corpus.pybuilds its own 25-file fixture with a fake binarythat never produces a verdict — irrelevant, since the subject is the denominator — so it needs no
compiler and runs in
loop-tools-gate.ymlbeside the other compiler-less tests.git show origin/master:…)corpus: 4 specsprinted for a 25-file treecorpus_totalafter the truncationThe exclusion that cleared this file was true and too narrow
An earlier bounded-read audit named it and let it pass:
Every word correct — and an argument about where the data comes from, used to settle a question
about what the label says. A local
--limittruncates the population exactly as thoroughly as apage boundary, and the printed word "corpus" does not care which one did it.
An exclusion is only as wide as the reason given for it. A file on a list headed "checked"
repels examination in a way an unexamined file does not.
SKILL 543. Refs #3195