tools: pin the corpus figures with their instrument - #3506
Merged
Merged
Conversation
Closes #3505 `published_figures.py` pins the spec-side populations. This pins the other side, and it is the side both of that audit's unit conflations came from: `len(x)` was published as 142 -- a DIAGNOSTIC count in the generated C -- inside a sentence about the specs, and `pub const OP_*` as 20, a count of list SITES rather than of declarations. IT REFUSES TO COMPARE ACROSS INSTRUMENTS. The same command gives 20 diagnostics on Apple clang and 50 on ubuntu's gcc for the same input (#3450), and 141 corpus files hit clang's default cap, which made every total measured without the uncap flag a floor (#3448). The pin records the compiler that produced it and `--check` exits 3, CANNOT TELL, rather than reporting a machine difference as a drift. positive control 1: the pin's instrument line set to a gcc string -> exit 3, "CANNOT TELL" positive control 2: a wrong value under the same instrument -> exit 1, and the row is named EVERY ROW CARRIES ITS UNIT -- units, errors, lines -- and each class is pinned by BOTH its diagnostic count and its distinct-line count, because the ratio between them is what says whether a population is real or partly a cascade: expected expression 750 errors / 668 lines cannot use '__auto_type' 385 / 385 use of undeclared identifier 'POS' 412 / 314 use of undeclared identifier 'NEG' 319 / 242 call to undeclared function 'len' 142 / 142 redefinition of 66 / 66 incompatible int-to-pointer 6 / 6 specs seen 651, generating 582, compiling clean 309, errors 10 802 AND THE READER WAS ONE COMMAND FROM PINNING A COLUMN THAT MEASURED NOTHING. The first version keyed distinct lines on the diagnostic's byte offset in the output, so every diagnostic received its own key and the `lines` column came out EXACTLY EQUAL to `errors` on every row -- 750/750, 385/385, 412/412. A column that always equals its neighbour measures nothing, and it was one `--bless` from being recorded as a fact. The line number is captured from the diagnostic now, and the self-check asserts the case that distinguishes them: two errors on one line must count as two errors and ONE line. A reader with a pin, not a fast gate: it rebuilds the corpus, which is minutes. Not wired into Spec Guards for that reason. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
gHashTag
enabled auto-merge (squash)
September 8, 2026 14:49
Contributor
|
📓 NotebookLM Notebook linked to this PR
This notebook contains session context, decisions, and artifacts for this work. |
Contributor
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
tools: pin the corpus figures with their instrument
Closes #3505
published_figures.pypins the spec-side populations. This pins theother side, and it is the side both of that audit's unit conflations
came from:
len(x)was published as 142 -- a DIAGNOSTIC count in thegenerated C -- inside a sentence about the specs, and
pub const OP_*as 20, a count of list SITES rather than of declarations.
IT REFUSES TO COMPARE ACROSS INSTRUMENTS. The same command gives 20
diagnostics on Apple clang and 50 on ubuntu's gcc for the same input
(#3450), and 141 corpus files hit clang's default cap, which made every
total measured without the uncap flag a floor (#3448). The pin records
the compiler that produced it and
--checkexits 3, CANNOT TELL, ratherthan reporting a machine difference as a drift.
positive control 1: the pin's instrument line set to a gcc string
-> exit 3, "CANNOT TELL"
positive control 2: a wrong value under the same instrument
-> exit 1, and the row is named
EVERY ROW CARRIES ITS UNIT -- units, errors, lines -- and each class is
pinned by BOTH its diagnostic count and its distinct-line count, because
the ratio between them is what says whether a population is real or
partly a cascade:
expected expression 750 errors / 668 lines
cannot use '__auto_type' 385 / 385
use of undeclared identifier 'POS' 412 / 314
use of undeclared identifier 'NEG' 319 / 242
call to undeclared function 'len' 142 / 142
redefinition of 66 / 66
incompatible int-to-pointer 6 / 6
specs seen 651, generating 582, compiling clean 309, errors 10 802
AND THE READER WAS ONE COMMAND FROM PINNING A COLUMN THAT MEASURED
NOTHING. The first version keyed distinct lines on the diagnostic's byte
offset in the output, so every diagnostic received its own key and the
linescolumn came out EXACTLY EQUAL toerrorson every row --750/750, 385/385, 412/412. A column that always equals its neighbour
measures nothing, and it was one
--blessfrom being recorded as afact. The line number is captured from the diagnostic now, and the
self-check asserts the case that distinguishes them: two errors on one
line must count as two errors and ONE line.
A reader with a pin, not a fast gate: it rebuilds the corpus, which is
minutes. Not wired into Spec Guards for that reason.
🤖 Generated with Claude Code