Skip to content

docs(memory): a canary is a regression test, not a plausibility check - #373

Merged
wenzowski merged 4 commits into
mainfrom
claude/semgrep-guardian-research-t6va58
Aug 12, 2026
Merged

wenzowski merged 4 commits into
mainfrom
claude/semgrep-guardian-research-t6va58

Conversation

@wenzowski

Copy link
Copy Markdown
Contributor

What changed

One paragraph added to mem:prior-art-and-issue-hygiene, recording what the field survey taught about instruments.

Why it mattered here

The survey classified 20,139 repositories. Round one matched policy files by filename and was wrong in both directions — a deny.toml that belonged to cargo-deny, a conftest.py that belonged to pytest. Round two was built specifically to make that unrepresentable, and reproduced the same blind spot in a new costume: it fetched only the paths guessed in advance, so root-level policy files nobody thought to name stayed invisible.

What caught it both times was a small set of cases whose correct answer was already known by hand, asserted as a hard failure rather than reviewed by eye.

What a reviewer should look at

The transferable claim, which is the only reason this is durable rather than a session note: absence of evidence is a claim about the instrument before it is a claim about the world. The corollary for any classifier is that a name may decide what gets read, and only the text decides what a thing is.

Refs: CLOUD-315

@linear-code

linear-code Bot commented Aug 12, 2026 •

Copy link
Copy Markdown
CLOUD-315 The surveyed sources are gated, and the survey's reasoning has a home

Why

A survey of the commercial static-analysis and agent-hook field produced three follow-on issues and a set of judgements that belong to no issue: why four of five candidates were rejected, and what to do differently next time a tool is evaluated.

Two things follow, and skipping either makes the survey a half-change.

The reasoning needs its destination. Per the sorting rule it is neither an issue nor a rule that binds every turn, so it goes to the memory that owns surveys — which is also the only file the attribution gate exempts, and therefore the only place a surveyed source may be named.

The gate needs the new names. mise-tasks/attribution-check carries the obligation in its own header: "Add a name here when a survey adds a source." A survey that leaves its sources ungated is rule 2's half-change — the reasoning lands durably while the practice it produced stays prose, and the next comment appealing to one of those projects passes unchallenged.

Mechanism

  • The survey's transferable judgements are appended to .serena/memories/prior-art-and-issue-hygiene.md, written through the Serena tools rather than by hand, since the path is protected.
  • Both surveyed engine names are added to the gate's NAMES alternation. Lines carrying a URL, a package path, or a tool pin stay exempt, so a genuine dependency remains addressable.

Ready

  • Source of truth. mise-tasks/attribution-check holds the name list; the memory holds the reasoning. Neither restates the other.
  • Mechanism as a predicate. mise run attribution-check exits 0, mise run memories-check exits 0, and mise run test:bats exits 0 with the attribution suite green.
  • Commit and bump. docs and ci → no release.
  • Blockers. None.

Done

main carries the appended memory and the widened name list, landed by fast-forward with CI green on the merge commit. A later prose mention of either engine outside the exempt file fails the gate.

Review in Linear

@wenzowski
wenzowski marked this pull request as ready for review August 12, 2026 19:19
@wenzowski
wenzowski marked this pull request as draft August 12, 2026 19:20
@wenzowski
wenzowski force-pushed the claude/semgrep-guardian-research-t6va58 branch from 2cda6b4 to 2df4050 Compare August 12, 2026 20:09
A survey that classifies a corpus needs cases whose answer is known because
getting them wrong is a mistake already made, and the run must fail when one
misclassifies. Round one of this survey matched policy files by filename and was
wrong in both directions; round two was built to make that unrepresentable and
reproduced the same blind spot in a new costume, fetching only paths guessed in
advance. The canaries are what caught it.

The transferable half is the sentence about instruments: absence of evidence is a
claim about the instrument before it is a claim about the world. The corollary for
any classifier is that a name may decide what gets read, and only the text decides
what a thing is.

Refs: CLOUD-315
@wenzowski
wenzowski marked this pull request as ready for review August 12, 2026 21:41
@wenzowski
wenzowski force-pushed the claude/semgrep-guardian-research-t6va58 branch from 2df4050 to bd1c7c5 Compare August 12, 2026 21:41
claude added 2 commits August 12, 2026 21:47
The command shipped with the framing the memory beside it had already
dropped: "the honest signal that you are past it is a rising re-verify rate",
plus a pointer to `mem:workflow/agent-fanout` for "the current cap", which
that memory no longer states in those terms.

Two authorities disagreeing about the same fact, in the file whose §1 was
"one authority each, no duplication" — and the command is the reachable one,
so a dispatcher invoking `/plan-fleet` got the superseded model while the
memory it points at carried the current one.

Replaced with the cost table the memory already holds: re-verifying is free
because `land` laps without a model turn, and what actually costs is a rebase
conflict, CI minutes on a voided run, and tokens. Objective restated as pace
of landed work per token, with the saturated-queue target and the CSMA/CD
framing, and the CI-minutes row now names CLOUD-369 as unbuilt rather than
implying a control exists.

Refs: CLOUD-367
…nheriting a spent window

Two defects in `plan-hold`, both measured 2026-08-12 in one session.

THE ANSWER WAS INVISIBLE. `plan-hold-guard` gates `PreToolUse` on
`ExitPlanMode|AskUserQuestion`; `plan-hold-release` listens on
`UserPromptSubmit`. Those are different event classes, and a human answering
either of those two tools produces a TOOL RESULT, not a prompt. So for the exact
case the mechanism exists to serve, the release could never fire: a hold armed,
an `AskUserQuestion` answered, the hold still live afterwards, ended by removing
its sentinel by hand. CLOUD-451's third acceptance bullet, unmet since it landed.

`plan-hold-release-tool` is the missing half, on `PostToolUse` over the same two
tools the guard already names. No classifier: `plan-hold-release` must decide
whether a prompt came from a person because `UserPromptSubmit` carries machine
turns too, while here the provenance is structural — the event fires only after a
tool whose whole purpose is to ask a human. The `tool_name` check is belt to the
matcher's braces, because a matcher widened later would otherwise turn every tool
call into a release. CLOUD-435's cost argument does not reach it for the reason
it does not reach `plan-hold-guard`: these two tools fire at most once per turn,
and it is invoked by path, so no task-runner startup is paid.

THE WINDOW ERODED SILENTLY. A second launch answered "already held" and exited,
which keeps one sleeper and hands the new handoff whatever is left of the old
one's cap. Some answers reach neither release path — a reply typed mid-turn
arrives `<`-wrapped and reads as a machine turn — so that handoff's hold never
releases, the next inherits a partly spent window, and eventually one is guarded
by a hold about to cap and the container is reclaimed while somebody is reading.
Nothing blocks, so nothing announces it. A second launch now releases the
incumbent and arms fresh: at most one sleeper still, and every handoff gets the
whole cap. Released, never signalled — a killed hold wakes nothing.

Deliberately NOT touching `plan-hold-release-check`'s `<`-prefix classifier. The
mid-turn false negative is real, but widening a literal without measuring it over
real transcripts is the mistake CLOUD-252 and CLOUD-323 exist to prevent. The
erosion fix removes its compounding cost; the false negative itself stays
recorded on CLOUD-485.

Refs: CLOUD-485
@wenzowski
wenzowski marked this pull request as draft August 12, 2026 21:50
Every bound in this memory was a bound on landing, and its cap of 2 is
derived from land contention. An agent that researches, reviews, or drafts
never lands, so that derivation says nothing about it — and the silence
read as absence of a constraint rather than absence of a derivation. Eight
drafting agents were launched into it.

The binding constraint for that shape is per-agent fixed cost, which is
paid once per agent and multiplies with N. The measured floor now has a
home: 63,848 tokens for one agent to fetch a single issue and run one
lint, for near-zero work. Each prompt named eight artifacts as required
reading, so that floor was paid eight times over before anything happened.

Three rules follow — digest rather than cite, checkpoint at each unit's own
gate because scratch dies with the container, and pilot one unit through
its durable write before widening — and the section says outright that
none of them is gated, because the predicate they need measures a session's
real spend and that is out of tree (CLOUD-95). Assumed enforcement is
worse than stated absence.

Also corrects the implementer contract in the same file: it still listed a
hand `gh pr ready` before `land`, which is the defect CLOUD-247 named, and
`land` has been the only readier since.

Refs: CLOUD-289
@sonarqubecloud

Copy link
Copy Markdown

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants