Skip to content

feat(gate): refuse Done while the issue still has a pull request open (CLOUD-468) - #370

Merged
wenzowski merged 2 commits into
mainfrom
wenzowski/cloud-468-a-merged-pr-is-read-as-its-issues-completion-but-an-issue
Aug 13, 2026
Merged

wenzowski merged 2 commits into
mainfrom
wenzowski/cloud-468-a-merged-pr-is-read-as-its-issues-completion-but-an-issue

Conversation

@wenzowski

@wenzowski wenzowski commented Aug 12, 2026 •

Copy link
Copy Markdown
Contributor

A merged PR is a completion signal for a diff. The board reads it as one for an issue. Those coincide only when the issue has exactly one PR, and nothing checked that they do.

Measured on the board itself, 2026-08-12. CLOUD-420 carries four pull requests — #362, #363, #366, #368 — and #368 was open and draft:

transition at
Todo → In Progress 06:33
In Progress → Done 07:17
Done → In Progress 07:52

A premature Done, reversed by hand 35 minutes later. The reversal is the evidence: the inference was made, acted on, and undone by a human. It was then made independently a second time, in a session reading the same board. Two occurrences of one mechanizable error inside six hours — and two of CLOUD-420's own acceptance clauses were still unmet at the moment it read Done, both living in the PR nobody had noticed was open.

It has since happened a third time, after this issue was filed: CLOUD-110 was set Done at 17:40 on 2026-08-13 and reversed to In Progress at 17:44, with PR #341 open and draft.

This is the threat model rather than a novel class: honest error about a completion signal. Non-negotiable 2 — a rule without a runnable gate is half a change.

Named done-pr-check, not done-check

CLOUD-468 §3 names mise-tasks/done-check. That name was taken while this PR sat unlanded: CLOUD-192's release-containment gate — no release tag contains a Done issue's commits — landed there in f5dd5ea and is already wired into release-plz.yml, AGENTS.md and mem:workflow/board-states. The rebase surfaced it as an add/add conflict on both the task and its suite.

Merging the two was considered and refused on evidence. Their stdin contracts are incompatible in the direction that matters: release-plz.yml pipes Done rows carrying no .pulls, which this gate answers 2 for by design — the clause that stops an unread PR being the cheapest route to Done. Folding them together would turn a green release path red for a reason unrelated to releases, and scoping the rule to "judge only when .pulls is present" would delete the acceptance clause instead. Two questions about one transition, two gates; a caller may run both. The landed-and-wired name keeps it. Recorded on CLOUD-468.

Not the question its neighbours answer

CLOUD-192 is about when the automation fires (merge rather than release); this is whether the transition is licensed at all given the issue's own attachments — an issue with an open PR is not Done on any schedule. graph-check owns the ready frontier and the In Review ⇒ linked PR rule, which is the entry to review rather than the exit from it. landed-check answers "this ref is on main". None of them can see a second PR still in flight.

What it decides, and what it refuses to guess

Arithmetic only: N attached PRs, k open ⇒ not Done. It never judges whether the merged ones did the work — not computable, and would make this a judge rather than a gate (CLOUD-93, non-negotiable 3). The PR filter is claim-check's character for character, so the two agree by construction on what counts as a pull request rather than by two authors agreeing today.

Two decisions worth stating:

  • A missing PR state is exit 2, not a licence. The whole defect is a Done granted over a PR nobody checked, so an absent state must not be the cheapest route to that same outcome.
  • A closed-unmerged PR does NOT refuse. An abandoned or superseded PR is a decided outcome, not work in flight. Refusing on it would block Done forever with no action that could clear it — the shape of a gate that gets bypassed rather than satisfied. The defect named here is an open PR, and only that.

Verification

19 rows in tests/done-pr-check.bats, pure function of stdin — no network, no credential — so they run unconditionally in the gate rather than needing live board data.

Registered in MUTANT_GATES, which the first draft omitted. The five mutations were run by hand and reported 5/5; nothing declared them, so mise run mutant would never have re-run them and the proof would have decayed into a claim. They are #MUTANT rows now — draft-not-open, open-state-ignored, no-pr-licensed, absent-state-licensed, filter-host-unanchored — and the repo-wide run is 25/25 caught.

The last of those is the genuine survivor the first pass found: widening the PR filter to pull/[0-9]+ passed every existing row, because the number capture already requires a leading slash and so rejects how-to-pull/123 on its own. The host restriction was the half nothing exercised — without it a forge link on another host would read as a blocker. That row now exists, and the mutant is caught.

The harness mutates a copy and verifies the tracked file's hash is unchanged on exit. An in-place harness makes a corrupted commit reachable from any concurrent git add -A, which staged a mutant into a pushed commit earlier the same day (recorded on CLOUD-418).

Driven against the live payload, which is also a committed regression row:

$ ./mise-tasks/done-pr-check < cloud-420.json
CLOUD-420 open-pr (#368, draft)
::error:: done-pr-check: not Done — an issue still has a pull request open.

That is the case a human had to catch by hand, now decided by an exit code.

Not wired to a hook yet

This ships as a runnable gate the caller invokes, matching claim-check and graph-check. Wiring it into the board automation is a separate question — the automation is what fires the Done transition, and it is not something this repository controls.

Closes CLOUD-468

Refs: CLOUD-420, CLOUD-418, CLOUD-192

@linear-code

linear-code Bot commented Aug 12, 2026 •

Copy link
Copy Markdown
CLOUD-468 A merged PR is read as its issue's completion, but an issue can carry N PRs — Done is reachable with the work half-landed

Why

"PR #366 is merged" reads as "CLOUD-420 is done". It is not. CLOUD-420 carries four PRs — #362, #363, #366, #368 — and #368 is open and draft. It is the issue's own "below it, for free" bullet: the ready-guard half that refuses an unauthorised hand-ready at zero CI cost, where #366 only stops a run after a matrix has been dispatched and cancelled. Two of CLOUD-420's acceptance clauses are still unmet and both live in #368: the hand-ready is not refused locally, and land-lock-check does not fail a lease body missing branch:.

Measured on the board itself, 2026-08-12. CLOUD-420's stateHistory:

transition at
Todo → In Progress 06:33
In Progress → Done 07:17
Done → In Progress 07:52

A premature Done, reversed 35 minutes later by hand. The reversal is the evidence: the inference was made, acted on, and had to be undone by a human. It was then made independently a second time, in a session reading the same board — two occurrences of one mechanizable error inside six hours.

This is the threat model, not a novel class. Batten exists for honest error: the wrong entity, time, or completion signal. A merged PR is a completion signal for a diff; the board treats it as one for an issue. Those coincide only when the issue has exactly one PR, and nothing checks that they do. Non-negotiable 2 — a rule without a runnable gate is half a change — and there is no gate.

Why the existing gates do not cover it. CLOUD-192 is about when the automation fires (merge rather than release); this is about whether the transition is licensed at all given the issue's own attachments — an issue with an open PR is not Done on any schedule. CLOUD-431 is an agent certifying a Ready block it wrote itself, a different failure at the other end of the lifecycle. graph-check answers the ready frontier, not the completion edge.

Refinement — Ready

  • Source of truth (§1). The issue's own attachments[], plus each attached PR's state — both already in the payloads agents fetch. No new record, and nothing inferred from a commit message or a branch name.
  • Mechanism (§3). mise-tasks/done-check, a pure function of stdin, the same agents fetch, gates decide interface claim-check and graph-check use — so no tracker credential exists in the task, nothing can hang or rate-limit, and the three compose in one pipeline. It reuses claim-check's attachment filter (.url | test("github\\.com/.+/pull/[0-9]+")) rather than growing a second spelling of "is this a PR", so the two agree by construction.
  • Arithmetic only (§3). N attachments, k open or draft ⇒ not Done. It does not judge whether the merged PRs did the work — that is not computable, and a gate that pretended otherwise would be a judge (CLOUD-93).
  • Output (§7). Pointer-only: the issue id, the rule id, and the PR number. Never a PR title, never a body.
  • Exit codes (§6). 0 licensed / 1 refused / 2 unreadable stdin, matching claim-check and graph-check.
  • Deliberately not in scope (§2). Writing the transition, or moving anything. This decides; the caller moves. Same split claim-check already draws between its verdict and the board move it tells the caller to make.

Test obligation

tests/done-check.bats, decision table: every attachment merged ⇒ allow; one open ⇒ refuse naming its number; one draft ⇒ refuse; no attachments ⇒ refuse, since In Review already requires a linked PR and Done cannot need less; a non-PR attachment is ignored rather than counted; unreadable stdin ⇒ exit 2, never allow. Plus the pointer-only property.

Regression fixture: CLOUD-420's real shape — four attachments, three merged, one draft — so the incident that motivated this cannot recur silently. Mutation-checked per CLOUD-418.

Commit / bump (§6): feat(gate) — patch until 0.1.0 regardless of type.

Blockers (§8): none.

Acceptance

  • An issue with any open or draft PR attached is refused Done, and the refusal names the PR number.
  • An issue whose every attached PR is merged is licensed.
  • A non-PR attachment never affects the verdict.
  • The verdict is a pure function of stdin — no network, no credential.
  • The CLOUD-420 payload that produced the premature Done is a committed fixture that refuses.

CLOUD-420 The landing lease is enforced only by the code path that honours it, so an agent that skips `land` still spends a full matrix

Why

CLOUD-393 serialises landing behind a lease and cuts the discarded-CI-run rate. It is enforced entirely inside mise-tasks/land. An agent that runs gh pr ready by hand, pushes to an already-ready PR, or wraps landing in its own logic spends a full four-job matrix without ever touching the lock.

That is the failure this repository names on its front page: "A new rule without a runnable gate is half a change. Prose is feedforward only." The lease is a convention honoured by the cooperating path, and the threat model is the honest agent that does the wrong thing — CLOUD-200 records a session that satisfied gh-guard and then wrapped mise run land in a bespoke retry loop it invented on the spot. Nothing stops the same shape here.

The dominant case is residue, not defiance. Measured 05:17–05:19Z on 2026-08-12: four concurrent pull_request matrices (#356, #354, #332, #302) while the lease changed hands three times, so the lease was being honoured by every session that held it and all four matrices ran anyway. land re-drafts only on red, so a landing interrupted any other way — a container reclaim, a stopped task, a lost lease, a die on a rebase conflict — leaves the PR ready permanently, and every later push to it buys a full matrix with no landing attempt in progress at all. All three of the other PRs were draft=false. The implementation is scoped to that case; an agent that skips land deliberately is the same mechanism and the smaller share.

The enabling gap. The lease identifies a clone (hostname-pid-random), which CI cannot check anything against. Adding the branch it authorises to the lease body is what makes it verifiable by the thing spending the money.

Refinement — Ready

  • Source of truth (§1). The lease ref refs/heads/batten-land-lock — a parentless commit over the empty tree whose body is land-lock / holder: / expires: / nonce: — read server-side by the workflow that is about to spend the runner. This issue adds a fifth line, branch:, and land-lock-check gains its well-formedness assertion rather than a second gate growing beside it. Not a receipt in a clone, which is exactly the artefact an agent skipping land never writes.

  • Mechanism (§3). land-lock gains a read-only verb, authorises <branch> — a pure function of (lease state, branch) reusing the existing observe / expired / released predicates. The alternative, re-parsing the body in workflow YAML, is a second authority for the ref and silently drops the read semantics the lease was pressure-tested into: a per-process observation ref, an unparseable body reading as held, and an unreachable remote that is not "free". Exit 0 run / 3 stop / 2 could not look.

    Each expensive pull_request job gains a first step, before the real checkout: a one-file sparse checkout of mise-tasks/land-lock at main (fetch-depth: 1, ~2s) and one invocation of it. The predicate comes from trunk, never from the head under test — that head has not been graded yet.

lease state action
absent, released, or expired-and-corroborated run — nobody is landing
branch: is this branch run — this is the authorised attempt
branch: names another branch stop, having spent only the spin-up
unreachable, unparseable, or carrying no branch: run, with a warning — fail open

Nothing in the table asks why a push happened, which is what makes it cover the residue case: the precondition is per job rather than per landing, so a push to a PR left ready by an interrupted land is stopped by the same row that stops a hand-typed gh pr ready.

The last row is the design and not a fallback: failing open costs one matrix, while failing closed on an unreadable ref stops every PR in the fleet, and a body minted before this change ages out within one TTL (120s).

  • A step, not a preflight job (§3), and the arithmetic that decides it. In billed minutes the two are closer than a wall-clock reading suggests: a gate job costs 1 job-minute per stopped run, because the expensive jobs never start and a skipped job bills nothing, while a first step in each job costs 7. Against that, the gate taxes every legitimate run one extra job-minute plus ~15-20s of serial spin-up. The step is therefore cheaper exactly while legitimate runs outnumber stopped ones better than 6:1, and lap length sits in the exponent of the collision model this whole line of work exists to shrink, which decides it at parity. That ratio is measurable rather than assumed; it is the number to re-check if the residue rate stays high.

  • THE HAZARD — a stop is a third answer, and the verdict layer has two (§2). A first step cannot make a job skipped; only a job-level if: can, which is the preflight job ruled out above. So the reachable conclusions are success, failure and cancelled, and checks-green reads success or neutral as green, failure and cancelled as red, and a latest-run skip as "not an answer".

    • Success or neutral grades a head green that compiled nothing, and main advances to exactly the SHA that passed. Disqualifying.
    • Failure re-drafts healthy PRs fleet-wide, since land re-drafts on red. Disqualifying.
    • Cancelled is the one that can be made correct. "This run declined to answer" is a third thing the verdict vocabulary lacks, and CLOUD-436's latest-run-per-name grouping is what makes a per-name declination readable at all.

    So: the step self-cancels the run (POST /actions/runs/{id}/cancel, actions: write on that job alone). Reading cancelled as a declination rather than as red is CLOUD-363's change, in flight on fix(tasks): read a cancelled required check as no verdict, not as red #302, which moves it into skipped's not-an-answer bucket in checks-green and in land's graded_runs together. This issue consumes that predicate rather than restating it, and that is why it is a blocker rather than a note: a stop landing first would re-draft healthy PRs fleet-wide, which is the hazard this bullet exists to avoid.

    Two consequences, both load-bearing:

    • Branch protection requires only final, which concludes cancelled too, so an unauthorised head cannot be merged while it waits.
    • Because graded_runs reads a cancelled set as ungraded, the stopped head's next land lap re-fires its ready and mints a real run, so the re-dispatch this would otherwise have to specify arrives with CLOUD-363's second half. It cannot loop: land acquires the lease before it pushes, so a land-driven push is never the one stopped.
  • Humans are unaffected by construction (§2). The "absent or released" row is what preserves them: a person pushing while no agent holds the lease sees no change. They wait only in the case where their run would have been voided anyway, which is a saving rather than a tax. fast-forward.yml is not touched.

  • Below it, for free (§3). Extend ready-guard — which already denies gh pr ready without verify and linear-check receipts — to require a lease.<branch> receipt, written by land-lock acquire into the same receipts directory. Catches the honest mistake at zero CI cost. A fail-fast convenience, not the guarantee: a hook can be unloaded (CLOUD-187) or bypassed, which is the whole reason the server-side check is the load-bearing half.

  • Above it (§3). A ci-local-parity property asserting that every pull_request job which checks the repository out carries the precondition as its first step, so a new job cannot silently omit it. The exemption is computable rather than enumerated: a job with no actions/checkout — final, the fan-in — has no work to protect.

  • Engine gap, per §2 (§3). That property is a predicate over workflow structure, which no rule kind can address, so it extends ci-local-parity instead of landing in batten.toml. The gap is CLOUD-452, linked rather than left as an exception.

Cost (§1), in the unit the invoice uses. This repository is private and every pull_request job is a GitHub-hosted ubuntu-latest runner, so the metered unit is a job-minute rounded up per job, not a second of wall clock. Measured p95 per job, 2026-08-12: ci 429s, darwin-link 97s, cross 82s, commit-lint 65s, msrv 60s, final 14s, action 12s — ~759s of runner time, which bills as 17 job-minutes per full run.

  • The tax on a legitimate run is ~2s in each of 7 jobs, which per-job rounding absorbs: 0 billed minutes, unless a job already sits within 2s of a minute boundary.
  • A stopped run still bills its floor — 7 jobs x 1 minute = 7 job-minutes against the 17 it replaces. That is ~59% of the invoice and ~86% of the wall clock: real, and smaller than a seconds-based reading implies.
  • The cheapest option for the residue case spends nothing at all, because a push to a PR left draft triggers no job: CLOUD-458. This issue is the backstop for what still leaks past that, so it must not be sized as though it carried the volume.

Local execution is the unmetered tier and nothing here moves work onto the metered one. The CI-side check exists because the local one is the half an interrupted session never reaches.

Test obligation

  • tests/land-lock.bats — the four rows above as a pure function of (lease body, branch), plus the two that must not be conflated: unreachable is not absent, and a missing branch: is not a foreign holder.
  • tests/land.bats — the composition rather than the predicate, which CLOUD-363 owns and tests: a head whose required runs were stopped by the precondition reads as ungraded, so the next lap re-fires the ready instead of re-drafting. Mutation-checked per CLOUD-418: with the stop step removed from the fixture workflow, that head grades normally.
  • tests/ci-local-parity.bats — both directions: a checkout-bearing pull_request job without the precondition is refused, and a job with no checkout is not.
  • tests/land-lock-check.bats — a lease body with no branch: line is reported.

Commit / bump (§6): feat(ci) — patch until 0.1.0 regardless of type.

Blockers (§8): blockedBy CLOUD-363 — the stop conclusion is safe only once a cancelled required check reads as no verdict rather than as red, which is in flight on #302. CLOUD-393 landed in #340, so the lease exists, and adding branch: to its body is this issue's own first step; CLOUD-436 landed in #347, supplying the latest-run-per-name grouping the declination sits on.

Acceptance

  • A PR pushed while another branch holds the lease stops within seconds instead of running a matrix.
  • A PR pushed while the lease is free, or while its own branch holds it, runs unchanged.
  • A stopped run neither reds the PR nor leaves its head unrecoverable: the branch's own next lap mints a real run without a hand-minted SHA.
  • A push to a PR left ready by an interrupted land is stopped by the same row as a hand-typed ready, with no landing attempt in progress.
  • An unreachable lease, or one carrying no branch:, runs the matrix rather than stopping the fleet.
  • A new checkout-bearing pull_request job cannot omit the precondition without failing ci-local-parity.
  • land-lock-check fails a lease body missing branch:.

CLOUD-418 A new gate is never shown to fail, so a test that cannot discriminate ships as coverage

Why

This repository's most-repeated failure is a claim nothing exercises. land's refusal branch was dead code for months (CLOUD-235). timeout-check's budgets were placeholders that could not fire (CLOUD-352). A shape rule whose pattern was a program could never match and read as coverage (CLOUD-401). Each was caught after the fact.

It happened again, live, while building the landing lease (CLOUD-393). A concurrency test was written for a real race — observe() reading FETCH_HEAD, which is one file per clone while the heartbeat runs beside held/release in the same checkout. The test was green. Then the buggy version was restored to check the test could catch it, and it passed on the broken code too: every process fetches the same lease ref, so a crossed read yields a different generation of the same lease rather than an observably foreign one. The test asserted nothing.

That was found only because someone chose to mutate and re-run — a discipline nothing asks for and nothing checks. The green suite before that check and the green suite after it were indistinguishable.

Root cause. The obligation is stated as "a rule ships with a runnable gate" — a gate that exists. Nothing requires evidence the gate discriminates. A test that passes on both the fixed and the broken code satisfies every rule this repo currently has.

Scope, deliberately narrow. Not mutation testing over the workspace, which is a research project and a large CI bill. The claim here is about mise-tasks/*-check and the guards — the files whose entire purpose is to refuse — where the mutation is usually a one-line inversion and the suite is bats, so a run is seconds.

Refinement — Ready

  • Source of truth (§1). The gate's own suite, run against a deliberately broken copy of the gate. A pass there is the defect; the verdict is an exit code, not a judgement.
  • Mechanism (§3). Undecided between two, and choosing is what Ready needs:
    • Author-side, checked in. Each gate declares one or more mutant cases — a stated one-line corruption and the test name that must go red. A task runs them, and a mutant nothing catches fails. Costs the repo one small fixture per gate; runs locally, off the landing path.
    • Scheduled sweep. A weekly job applies mechanical mutations to mise-tasks/*-check and reports any whose suite stays green. No per-gate authoring, weaker coverage, and it belongs beside branch-age-check in the hygiene sweep, so no new CI minutes.
  • Deliberately not in scope (§2). Mutation coverage of crates/. Different tooling, different cost, different question.
  • Output (§7). Pointer-only: the gate, the mutant, and the test that failed to notice. Never a diff of the mutated source.

Test obligation

The mechanism must catch the case that motivated it: the FETCH_HEAD mutation of mise-tasks/land-lock against tests/land-lock.bats as it stood before the structural assertion replaced it. That pair is a known-good fixture — a real gate, a real mutant, and a real suite that missed it.

Commit / bump (§6): feat(gate) — patch until 0.1.0 regardless of type.

Blockers (§8): none.

Acceptance

  • Every *-check task has at least one mutation its suite is proven to catch.
  • A gate whose suite passes on a broken copy fails.
  • The land-lock/FETCH_HEAD pair is covered as a regression fixture, so the case that motivated this cannot recur silently.

CLOUD-192 The board transitions to Done on merge, not on release — In Review is never occupied

Why. The board is the observability surface, and its states are defined against a trunk-based model: In Review = landed on main, under post-merge review, pre-release; Done = released. Observed on CLOUD-168, the integration collapses both into one transition fired by the merge.

Measured. PR #103 merged at 19:29; CLOUD-168's state history is Backlog -> In Progress -> Done with Done set at 19:29:03 — merge time. The release (v0.0.15, which does contain the commit) was not cut until 19:31:56, ~3 minutes later. So the issue read Done for three minutes while unreleased, and In Review was never occupied at any point.

The end state converged here only because release-plz happened to release promptly. The transition is keyed on the wrong event, so the failure is latent rather than absent:

  • A release that fails, is delayed, or is held leaves issues sitting in Done while unreleased — the exact "done means landed-and-verified" drift Batten exists to prevent, on Batten's own board.
  • In Review having no occupants means the post-merge review window the trunk-based model depends on has no board representation. Per AGENTS.md we review after merge, before release; a column that is structurally never entered cannot surface that work.

This also supersedes a recorded observation in mem:workflow/board-states, which states "the merge-side transition did not fire" (measured when #100 merged and CLOUD-178 stayed In Progress). It now fires — and overshoots. That memory needs correcting either way, since an agent reading it today will expect to hand-move a landed issue that the integration has already moved.

Acceptance.

  • Merging a PR moves its issue to In Review, not Done.
  • Done is set by the release, not the merge — an issue whose commit is on main but not in any published release is never Done.
  • A test or documented probe demonstrates the two-step path on a real issue (state history shows In Progress -> In Review -> Done).
  • mem:workflow/board-states is updated to record the current, measured behaviour, replacing the "did not fire" note.

Implementation, 2026-08-13.

Re-measured before touching anything. The defect is still live and unchanged from the original CLOUD-168 observation: CLOUD-499's state history reads Todo -> In Progress -> Done, with Done set at 05:13:12 — merge time. Last tag v0.0.62, main 50 commits past it, so ~1.5 days of landed work read as released.

The two halves have different owners, and only one is code.

  • merge -> In Review is a workspace setting — the team's GitHub integration maps the merged event to a status. No repo code and no Linear MCP tool can change it, so nothing in the tree can hold it. It was flipped from Done to In Review (Workflows & automations -> Pull request automations -> "On PR merge, move to…"). The setting alone did not produce the transition — see the probe below. This bullet stays open.
  • release -> Done has no automation available at all, and this is not a gap that a setting closes. Per mise-tasks/released's header: the integration triggers on PR events, and "a release tag now contains this commit" is not one. Performing it needs a Linear credential the repo deliberately does not have.

So the repo ships the half it can hold: the consequence. mise-tasks/done-check refuses a Done that no v* tag reaches — CLOUD-N Done -> In Review, exit 1. It is landed-check's terminal twin (both name In Review, one for a board behind git, one for a board ahead of the release) and composes with released running the other way. It only ever refutes a Done, never confirms one: refs come from commit messages, so a ref inside a tag is weak evidence while a ref nowhere near one is conclusive. A skipped promotion is now a failing check rather than silent drift.

The board was corrected rather than grandfathered. Piping the full Done closure (159 issues) exited 1 and named five: CLOUD-404, CLOUD-484, CLOUD-491, CLOUD-495, CLOUD-499 — each landed, each in no tag. All five moved to In Review, graph-check-adjudicated on their PR attachments first. The same closure now exits 0 (27 not judged, the unlanded channel). That before/after pair is the gate demonstrating it can fail on real data, not only on fixtures.

Fewer than the 50 unreleased commits would suggest, because an issue is judged by its most-released ref: work carrying several PRs, some shipped, passes. That is CLOUD-468's question and this gate deliberately does not answer it.

Notes. If the integration cannot key Done on a release event, the fallback is to stop it at In Review and let the release step promote — a correct-but-manual last transition beats an automatic wrong one, since graph-check already gates In Review => a linked PR attachment.
The probe FAILED, and that is the finding. This issue's own PR (#398) was the demonstration. It merged at 06:27:05 by fast-forward, main at bd99948, all eight required checks green. Read at 06:29:48, 06:35:26 — 8m21s after the merge — the issue was still In Progress with updatedAt unchanged from an edit made before the PR existed. The state history shows no In Review entry at any point.

The automation is not simply slow: mem:workflow/board-states records the open-side transition firing eight seconds after gh pr create, and this same PR's open-side fired normally — linear-code[bot] commented on it at 06:12:20 and attached #398 to this issue. So Linear received the PR, linked it, and did not act on the merge.

What that rules in and out:

  • Not a missing key. The branch was claude/cloud-192-implementation-9l6mcg, carrying cloud-192, and the attachment proves the link resolved to this issue.
  • Not a missing setting, re-confirmed after the fact rather than assumed. The settings page was re-opened and screenshotted at 11:34 local — after the merge, not before it — and "On PR merge, move to…" still read In Review. So the value persisted and was live when feat(done-check): key Done on the release, and refuse a Done no tag contains #398 merged. The obvious explanation, that a dropdown was selected but never saved, is ruled out by measurement.
  • Not the fast-forward merge being invisible. GitHub reports the PR MERGED with mergedAt set, and the timeline carries merged + closed events.

The untested hypothesis is Linear's closing vs contributing PR distinction, which the settings page names in its own copy. #398's body carries Refs: CLOUD-192 and no closing keyword, so it may be registered as contributing — linked, but not driving status. That would also explain why the transition is reachable at all in this workspace while being unreachable for the PRs this repo actually produces, since issue-guard requires the key and nothing requires a closing keyword.

It is not settled by the earlier counter-evidence in the memory (#131 completing on merge with only a Refs: trailer), because that was measured against the old Done mapping; whether the two rows classify PRs the same way is exactly the open question.

RESOLVED: it is the closing keyword. The controlled probe ran on this issue, and the result is unambiguous.

PR body names the issue as merged In Review at delay
#398 Refs: CLOUD-192 06:27:05 never — (8m21s, then hand-moved)
#400 Closes CLOUD-192 06:59:39 06:59:41.483 2 seconds
#404 Closes CLOUD-192 10:50:22 10:50:24.818 2 seconds

#404 is the confirming instance rather than a repeat: it is the PR that ships closing-key-check, so the gate ran against its own body on the way through and its merge exercised the rule it adds. Between #400 and #404 the issue went In Review -> In Progress at 07:17:08 on #404's open event, which is the other direction the settings page configures, so the full In Progress -> In Review cycle is in the state history twice.

One variable. Same repository, same branch name (claude/cloud-192-implementation-9l6mcg, reused), same fast-forward landing, same integration settings, same issue, and the issue was returned to In Progress before #400 so the transition was observable rather than a no-op. So Linear's closing vs contributing split is the cause: a contributing PR is linked and attached but does not drive status, and every PR this repo produces is a contributing one, because issue-guard requires a key and nothing requires the closing form.

That also retires the #131 counter-example from mem:workflow/board-states — a Refs:-only PR completing on merge cannot be reproduced, and it was measured against the old Done mapping.

So the merge-side transition is reachable, and reaching it is a repo-side change: the PR body must name its issue with a closing keyword. That ships with a gate rather than a convention, per non-negotiable 2 — a rule with no runnable mechanism is half a change, and this specific rule is invisible when broken, which is exactly the shape that decays.

Superseded next step (kept for the record): If it moves to In Review and a Refs:-only PR does not, the fix is a repo-side one — issue-guard/land emit the closing form — and it ships with a gate. If neither moves, the merged-event rule is not reaching this workspace's PRs at all and the fallback in the Notes below is the answer: stop at In Review by hand, with done-check (landed) holding the line on the other end.

Status against the four acceptance bullets.

  1. Merging a PR moves its issue to In Review, not Done — holds, twice measured, and now gated rather than conventional: closing-key-check refuses a body that names its issue without closing it, wired into land beside deferral-check.

  2. Done is set by the release, not the merge — holds. done-check refuses a Done that no v* tag reaches; the five issues that were wrongly Done were corrected and the closure now exits 0.

  3. A probe demonstrates the two-step path — holds, completed by this issue's own state history:

    In Progress  07:17:08 -> 10:50:24
    In Review    10:50:24 -> 17:40:43
    Done         17:40:43
    

    Each leg by its own mechanism, and the timings are the evidence. In Review was written by the merge of feat(closing-key-check): a PR that names its issue but never closes it moves nothing #404, two seconds after it, because that body closed the key. Done was written by a release sweep — 6h50m later, and only once v0.0.65 existed to contain the commits. That gap is the whole point of the issue: under the old configuration it was zero, because both were the same event.

    The Done is truthful rather than merely set: v0.0.65 contains 87c99f0, cf3c60f and 97b4819, and done-check exits 0 on it.

  4. mem:workflow/board-states records the measured behaviour — holds, after two wrong intermediate versions of my own (that the merge writes In Review unconditionally, then that it never does). The claim is now conditional: the merge writes it iff the body closes the key.

All four bullets hold, and the issue is Done on the definition it argued for — released, not merely merged. It reached each column by the mechanism that column is supposed to have: In Review by the merge (given a closing body, now gated by closing-key-check), Done by a release sweep nearly seven hours later. done-check is what stops the two collapsing back into one.

Review in Linear

wenzowski pushed a commit that referenced this pull request Aug 12, 2026
…ing it

Measured twice on 2026-08-12, on two branches that touch no part of the landing
loop: PR #354 (a fuzz crate) and PR #370 (a board gate) each lost a full verify
and a lap to `not ok THE DEFECT: a lease sighted before it expired is taken on
the first check after`. Re-run alone, the case passes. It was never the branch
and never the code under test — it was the test's own clock.

The race is in the SETUP, not the measurement, which is what the original
comment missed while correctly calling the row above it "the flake-proof half".
The case pins a 4s TTL and then depends on a wall-clock ordering across separate
process launches. `test:bats` runs under `rush --jobs` (CLOUD-386), and the
process can be descheduled for longer than the whole TTL — at which point the
"sight the live lease" acquire is RIGHT to succeed, and it steals instead of
sighting. The output said so plainly: `took the lease 0s after ... stopped
holding it`. The assertion graded the runner's scheduler.

That is worse than an ordinary flake because it sits in `verify`, on the landing
path, in the suite that gates every branch in the fleet. `land` says "reproduce
and fix locally", nothing reproduces, and the reliable way through is to run
`land` again — the reflexive drive-to-green AGENTS.md forbids, arrived at
honestly.

So the precondition is established rather than assumed, and never asserted
through (CLOUD-249): a sighting acquire that succeeds means the lease had
already expired, so no sighting happened and there is nothing to measure. SETUP
is retried — never the measurement, which would be drive-to-green in the test —
because a plain skip would fire often enough to erase the coverage, this having
raced twice in one day. Three attempts, then a skip naming the reason.

The TTL stays short deliberately. Raising it widens the window without removing
the race, and every second added is paid on every run of the suite.

Both halves of the obligation are checked, since either alone would let the
repair convert a flaky check into one that cannot fail:

  A. Against the pre-fix land-lock (sighting recorded only once expired, steal
     at ~9-12s) the repaired case still goes RED.
  B. With a deschedule past the whole TTL injected before the sighting, it
     SKIPS naming the reason — never fails, never silently passes.

`LAND_LOCK_UNDER_TEST` is added so that harness mutates a COPY: an in-place
mutation makes a corrupted commit reachable from any concurrent `git add -A`,
which staged a mutant into a pushed commit earlier the same day (CLOUD-418).
The harness hashes both tracked files before and after and requires them
byte-identical.

Refs: CLOUD-448
@wenzowski
wenzowski force-pushed the wenzowski/cloud-468-a-merged-pr-is-read-as-its-issues-completion-but-an-issue branch from 31626ca to 931abc8 Compare August 13, 2026 19:46
The table tracks adopted or vendored tools. Three rows survived
tools whose adopt verdicts closed negative: the file-shape linter,
the red-green-refactor/judge prior art, and the hook-file mapping
generator, whose wiring derives from the Harness enum instead. A
mined design creates no dependency and no license obligation.
cargo-deny and ripsecrets stay: one gates today, the other is a
pinned dependency of the secrets rule kind.

Refs: CLOUD-530
Claude-Session: https://claude.ai/code/session_01PKrKv9gwfiB7MKbRzZGZSV
@wenzowski
wenzowski force-pushed the wenzowski/cloud-468-a-merged-pr-is-read-as-its-issues-completion-but-an-issue branch from 931abc8 to 6d6cd5a Compare August 13, 2026 19:56
A merged PR is a completion signal for a DIFF. The board reads it as one for an
ISSUE. Those coincide only when the issue has exactly one PR, and nothing
checked that they do.

Measured on the board itself, 2026-08-12. CLOUD-420 carries four pull requests —
Todo -> In Progress 06:33 -> Done 07:17 -> In Progress 07:52: a premature Done,
reversed by hand 35 minutes later. The same inference was then made
independently by a session reading the same board. Two occurrences of one
mechanizable error inside six hours, and two of CLOUD-420's own acceptance
clauses were still unmet at the moment it read Done — both living in the PR
nobody had noticed was open.

That is the threat model rather than a novel class: honest error about a
completion signal. Non-negotiable 2 says a rule without a runnable gate is half
a change, so this is the gate.

It is not the question its neighbours answer. CLOUD-192 is about WHEN the
automation fires, merge rather than release; this is whether the transition is
licensed at all given the issue's own attachments. `graph-check` owns the ready
frontier and the In Review => linked PR rule, which is the entry to review
rather than the exit from it. Neither can see a second PR still in flight.

Arithmetic only: N attached PRs, k open => not Done. It never judges whether the
merged ones did the work, which is not computable and would make this a judge
rather than a gate. The PR filter is claim-check's, character for character, so
the two agree by construction on what counts as a pull request.

Two decisions worth stating. A missing PR state is exit 2, not a licence: the
whole defect is a Done granted over a PR nobody checked, so an absent state must
not be the cheapest route to that same outcome. And a closed-unmerged PR does
NOT refuse — an abandoned or superseded PR is a decided outcome rather than work
in flight, and refusing on it would block Done forever with no action that could
clear it, which is the shape of a gate that gets bypassed rather than satisfied.

Mutation-proven 5/5, including one genuine survivor: widening the PR filter to
`pull/[0-9]+` passed every row, because the number capture already requires a
leading slash and so rejects `how-to-pull/123` on its own. The host restriction
was the half nothing exercised, so a forge link on another host would have read
as a blocker. That row now exists.

The harness mutates a COPY and verifies the tracked file's hash is unchanged on
exit. An in-place harness makes a corrupted commit reachable from any concurrent
`git add -A`, which staged a mutant into a pushed commit earlier the same day.

CLOUD-420's real payload is a committed regression row, and driving the gate
against it refuses with `CLOUD-420 open-pr (#368, draft)`.

NAMED done-pr-check, not done-check as the issue's Mechanism clause says. That
name landed first for a different predicate (CLOUD-192: no release tag contains a
Done issue's commits) and is already wired into release-plz.yml. Merging the two
was refused on evidence: that caller pipes Done rows carrying no .pulls, which
this gate answers 2 for by design, so folding them together would turn a green
release path red for a reason unrelated to releases.

Registered in MUTANT_GATES, which the first draft omitted — the mutation run that
proved this gate discriminates would not have run again without it.

Refs: CLOUD-468
@wenzowski
wenzowski marked this pull request as ready for review August 13, 2026 20:04
@wenzowski
wenzowski force-pushed the wenzowski/cloud-468-a-merged-pr-is-read-as-its-issues-completion-but-an-issue branch from 6d6cd5a to 346783a Compare August 13, 2026 20:04
@sonarqubecloud

Copy link
Copy Markdown

@wenzowski

Copy link
Copy Markdown
Contributor Author

/fast-forward

@wenzowski
wenzowski merged commit 346783a into main Aug 13, 2026
8 checks passed
@wenzowski
wenzowski deleted the wenzowski/cloud-468-a-merged-pr-is-read-as-its-issues-completion-but-an-issue branch August 13, 2026 20:15
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants