Deferred from the self-supplied-evidence spec, which deliberately scoped it out. Recording it as an issue rather than a sentence in an overview, so it is a commitment rather than a note.
The failure
A verdict's scope claim can generalise past the evidence that supports it, even when every piece of that evidence is honest.
The motivating instance, from panopticon PR #112: verdict REQ-034.17.1 asserted that the tests "exercise authenticated shared-client, liveness, and shell-library requests". Three shell libraries in that repository talk to the task service. The tests exercised exactly one.
The single test was not dishonest — it captured real runtime argv with a recording stub, which is precisely the technique the reviewer should want. Nothing about its evidence was self-supplied. What failed was the judgment: a plural category was claimed from a single member, nobody enumerated the set, and the uncovered member turned out to be a live security defect (a shell library still calling the control plane with unauthenticated curl).
Why it is not the self-supplied-evidence class
Those are different mechanisms with the same symptom — a green check that does not establish the property.
- Self-supplied evidence: the check consumed evidence the checker manufactured.
- Scope inflation: the evidence was genuine; the claim generalised beyond it.
The self-supplied-evidence requirements do not catch this and should not be stretched to. Its Rule 1 — name a production failure the test would catch, and prove production can reach it — passes on the motivating instance, correctly, because the test does catch a real failure for the one file it covers.
Why a reviewer question is probably the wrong instrument
Asking a reviewer "is this subject plural, and did you enumerate the members?" invites the same generalisation error it is meant to catch: the reviewer's own enumeration becomes another unverified claim.
What reliably kills scope inflation is mechanical enumeration from the real artifact. A glob that finds three shell files and asserts a property across all three cannot over-claim. Panopticon has a companion task (credential-boundary-guards) building exactly this shape — every route on the real composed app, every shell file matching a glob, every registered harness — with the accompanying rule that a sweep MUST fail loudly when enumeration returns no members, since a glob silently matching zero files is the self-supplied-evidence class wearing a lint's clothes.
So the likely answer here is tooling, done once, rather than judgment repeated per requirement. That should be evaluated rather than assumed.
Possible directions
- A machine-readable way for a requirement to declare that its subject is a set, plus a way to enumerate that set from the real artifact at check time.
- A
check-time warning when a verdict's prose uses plural or categorical nouns that the cited evidence does not cover.
- Guidance steering requirement authors away from unbounded categorical subjects where a set can be enumerated instead.
Acceptance criteria
- A requirement whose subject is a set cannot be satisfied by evidence covering a strict subset of that set without the shortfall being visible.
- Enumeration derives from the real artifact, never from a hand-maintained list — a hand-maintained list decays into exactly the omission this exists to catch.
- Enumeration returning zero members is a loud failure, not a silent pass.
Note
A narrow, near-free piece of this is being folded into the self-supplied-evidence spec separately: a verdict must not describe its scope in terms broader than the evidence it cites. That is a wording constraint on the conclusion, not a third analytical concern, and it would have caught the motivating instance from the verdict text alone. This issue covers the structural version that constraint cannot reach.
Deferred from the
self-supplied-evidencespec, which deliberately scoped it out. Recording it as an issue rather than a sentence in an overview, so it is a commitment rather than a note.The failure
A verdict's scope claim can generalise past the evidence that supports it, even when every piece of that evidence is honest.
The motivating instance, from panopticon PR #112: verdict
REQ-034.17.1asserted that the tests "exercise authenticated shared-client, liveness, and shell-library requests". Three shell libraries in that repository talk to the task service. The tests exercised exactly one.The single test was not dishonest — it captured real runtime argv with a recording stub, which is precisely the technique the reviewer should want. Nothing about its evidence was self-supplied. What failed was the judgment: a plural category was claimed from a single member, nobody enumerated the set, and the uncovered member turned out to be a live security defect (a shell library still calling the control plane with unauthenticated
curl).Why it is not the self-supplied-evidence class
Those are different mechanisms with the same symptom — a green check that does not establish the property.
The
self-supplied-evidencerequirements do not catch this and should not be stretched to. Its Rule 1 — name a production failure the test would catch, and prove production can reach it — passes on the motivating instance, correctly, because the test does catch a real failure for the one file it covers.Why a reviewer question is probably the wrong instrument
Asking a reviewer "is this subject plural, and did you enumerate the members?" invites the same generalisation error it is meant to catch: the reviewer's own enumeration becomes another unverified claim.
What reliably kills scope inflation is mechanical enumeration from the real artifact. A glob that finds three shell files and asserts a property across all three cannot over-claim. Panopticon has a companion task (
credential-boundary-guards) building exactly this shape — every route on the real composed app, every shell file matching a glob, every registered harness — with the accompanying rule that a sweep MUST fail loudly when enumeration returns no members, since a glob silently matching zero files is the self-supplied-evidence class wearing a lint's clothes.So the likely answer here is tooling, done once, rather than judgment repeated per requirement. That should be evaluated rather than assumed.
Possible directions
check-time warning when a verdict's prose uses plural or categorical nouns that the cited evidence does not cover.Acceptance criteria
Note
A narrow, near-free piece of this is being folded into the
self-supplied-evidencespec separately: a verdict must not describe its scope in terms broader than the evidence it cites. That is a wording constraint on the conclusion, not a third analytical concern, and it would have caught the motivating instance from the verdict text alone. This issue covers the structural version that constraint cannot reach.