Repository navigation
feat(ci): stop an unauthorised run before it spends a matrix (CLOUD-420) - #366
Conversation
CLOUD-420 The landing lease is enforced only by the code path that honours it, so an agent that skips `land` still spends a full matrix
Why CLOUD-393 serialises landing behind a lease and cuts the discarded-CI-run rate. It is enforced entirely inside That is the failure this repository names on its front page: "A new rule without a runnable gate is half a change. Prose is feedforward only." The lease is a convention honoured by the cooperating path, and the threat model is the honest agent that does the wrong thing — CLOUD-200 records a session that satisfied The dominant case is residue, not defiance. Measured 05:17–05:19Z on 2026-08-12: four concurrent The enabling gap. The lease identifies a clone ( Refinement — Ready
Nothing in the table asks why a push happened, which is what makes it cover the residue case: the precondition is per job rather than per landing, so a push to a PR left ready by an interrupted The last row is the design and not a fallback: failing open costs one matrix, while failing closed on an unreadable ref stops every PR in the fleet, and a body minted before this change ages out within one TTL (120s).
Cost (§1), in the unit the invoice uses. This repository is private and every
Local execution is the unmetered tier and nothing here moves work onto the metered one. The CI-side check exists because the local one is the half an interrupted session never reaches. Test obligation
Commit / bump (§6): Blockers (§8): blockedBy CLOUD-363 — the stop conclusion is safe only once a cancelled required check reads as no verdict rather than as red, which is in flight on #302. CLOUD-393 landed in #340, so the lease exists, and adding Acceptance
CLOUD-393 Landing collisions cost ~2.9 CI runs per landed PR, and the recorded reason for not queueing answers a different question
Why
The measured cost of not queueing. A lap dies whenever The same memory argues this is correct — "collisions are designed for rather than prevented, in the CSMA/CD sense", and "a rising re-verify rate is NOT the stop signal" — and that argument is sound as far as it goes. What it does not do is compare against the alternative, because no alternative was costed. Two alternatives, both preserving exactly-tested-SHA fast-forward:
The honest test, and why this is Backlog rather than Todo. Refinement — Ready
Test obligation If a land lock ships: a lease that expires must be reclaimable, and a holder that dies must not wedge landing. Those two are the gate, and they are exactly the shape Commit / bump (§6): Blockers (§8): none formally, but this should not be pulled before CLOUD-386, CLOUD-390 and CLOUD-392 have landed and Acceptance
Raised from Low to Urgent: the number this issue said it lacked now existsThis issue argued the case for queueing and recorded that the honest test was to shorten the lap first and see whether the case survived. It did not survive. Measured 2026-08-11, 21:31–22:01Z (400
The bot is not the bottleneck — it answered every one of 248 attempts inside 23 seconds. 243 of them were genuine So the recorded position in Mechanism chosen: the land lock, the cheap half of the two options in the body. GitHub merge queue stays out of scope — it is a larger departure from Refinement — Ready
Test obligationTwo assertions the change rests on, both against a real remote rather than a fixture: two concurrent acquires — exactly one exits 0; and a holder whose lease was stolen reports not-held from its pre-comment re-check rather than proceeding. Plus a Commit / bump (§6): Blockers (§8): none. Acceptance
Implementation record — four things the Ready block got wrongAll four were found by testing against the real remote rather than by reasoning, and each is recorded because the wrong version is written down above. 1. The lease is a BRANCH, not
The lease is 2. Renewal is CAS, but not the CAS the plan claimed. I asserted 3. Delete goes through the REST API, not a push. The proxy 403s a delete-push too — the same reason 4. The lease body needs a nonce. Git addresses objects by content, so two mints agreeing on holder and expiry — the same clone renewing twice inside one second — produce the same sha, and pushing a sha the ref already points at is an "up to date" no-op that reports success. That turns a rejected claim into an apparent win. Measured before the fix: a second A defect in
|
| the lease | CAS-only, rolling TTL, no delete, no REST API |
land-lock-check |
the gate, because "stays well-formed" was prose |
| four reliability defects | FETCH_HEAD cross-read, heartbeat fragility, fence margin, clock skew |
| identity + delete refspec | commit-tree needs an identity; an empty mint was a delete |
| the test fix | label the two waits instead of deducing which is which |
the bare wait |
the one that mattered — see below |
The last one is the finding worth carrying forward. land's CI race ended with a bare wait, which blocks on every background job of the shell. This change adds one that by design never exits: the lease heartbeat. So from the moment the lease shipped, land would reach green CI and then block forever. It sat there for five minutes with all eight checks green and the SHA landable, logging nothing.
CLOUD-383 had already named the shape as an intermittent hang when a group kill misses. The heartbeat made it certain. The lease, as first written, could never have landed itself — and that was invisible to every test, because reproducing it requires a never-exiting child, which is precisely what would hang the suite. It is gated structurally now: no bare wait may appear in the task.
It also explains attempts I had attributed to contention and to a moving main. The last several were not that.
What is enforced, and what is not. The lease serialises landing across clones. It does not serialise within one — acquire is re-entrant per clone by design, so two land processes in one checkout both hold it (CLOUD-428, filed after exactly that happened here: three concurrent lands, two heartbeats). And it is honoured only by land itself; an agent that readies by hand still spends a matrix (CLOUD-420). Neither is a defect in what shipped; both are the boundary of it.
Not yet measured: the refusal:success ratio against the 243:5 baseline. The fleet went quiet at 22:01Z, so there is no comparable window yet. That measurement is this issue's acceptance criterion and remains outstanding — it should be taken before this moves to Done, not inferred from the fact that #340 landed in one lap.
CLOUD-363 A `cancelled` required check is read as red, so a run superseded by `land`'s own ready/push race wedges the branch permanently
Why
Measured on PR #293 (CLOUD-346), 2026-08-11. land lap 1 readied the PR and then force-pushed, which is the deliberate order from CLOUD-254 — ready first, so the push's own event is the confirming run. Both events reached the same concurrency: ci-${{ github.ref }} group two seconds apart:
17:53:15 run 31519869739 head 3eba36a (pre-push SHA, the ready_for_review event) -> kept running
17:53:17 run 31519873369 head a3589cc (the force-push event) -> CANCELLED
The run on the SHA that would land was the one cancelled; the run on the stale SHA survived. final then failed on a3589cc because its upstreams were cancelled, and ci-wait reported red:
::error:: CI is not green on a3589cc — final failure, msrv cancelled, ci cancelled,
cross cancelled, darwin-link (aarch64-apple-darwin) cancelled, commit-lint cancelled.
The red is not a verdict. A cancelled run did not judge anything — it is the absence of an answer, exactly like the draft-era skipped that CLOUD-327 taught this repo not to read as a conclusion. But the predicate's catch-all buckets it with failure:
$2 == "skipped" { skipped = ...; next }
$2 == "success" || $2 == "neutral" { graded++; next }
{ graded++; bad = bad sep4 $3 " " $2; sep4 = ", " } # <- cancelled lands hereAnd that wedges the branch, which is the part that makes this urgent. The two rules compose into a trap with no exit:
ci-waitreads the cancelled set as red →landre-drafts the PR and stops (correct, given "red").- Re-running
landcannot recover. HEAD is unchanged, so the push isEverything up-to-date; the verify receipt still holds, so nothing re-proves; andland's ready block only fires when the SHA carries no graded run (CLOUD-255) — the cancelled runs are graded bygraded_runs' reckoning, so the ready is skipped. Lap 1 goes straight back to the same stale reading and re-drafts again.
Observed exactly that: two consecutive land invocations, both terminating on the identical cancelled set, neither able to create a single new check-run. The only recovery available to an agent is to hand-mint a new SHA (git commit --amend), which is a manual step outside the loop land exists to drive.
Ready
cancelled is not an answer. Move it to the same bucket skipped occupies — one place, since CLOUD-346 collapsed the predicate to a single definition.
mise-tasks/checks-green— the catch-all keepsfailure/timed_out/action_required;cancelledjoinsskippedin the "not an answer" set (exit 3). Name it in the same pointer line, distinguishing the two words so a stall is diagnosable:required check(s) with no verdict: ci cancelled, cross skipped.mise-tasks/land—graded_runs(line ~129) must agree, or the compose-trap survives the first fix: a cancelled set must read as "no graded run", so the lap re-fires the ready and a real run replaces it. The two readCI_REQUIRED_CHECKSfrom one place already; they must readcancelledthe same way too.- Consider whether the race itself is worth closing separately — the ready and the push both landing in one concurrency group is what produces the cancellation, and CLOUD-254/CLOUD-255 have iterated on that ordering twice. This issue does not propose re-ordering them: making a cancellation recoverable is the smaller, more general fix, and it holds whatever cancels a run (a concurrent push, a manual cancel, a runner reclaim). Filed as a note, not a task.
Test obligation
tests/checks-green.bats— a required checkcancelledis exit 3 and named; a set mixingcancelledandsuccessis still exit 3, not green;failurestays exit 1.tests/land.bats— a SHA whose required checks are allcancelledre-fires the ready on the next lap rather than re-drafting; the existing lap-count assertions must still hold.
Commit type / bump: fix(tasks). Patch until 0.1.0.
Blockers: none. It sits on top of CLOUD-346's checks-green, which is in flight on PR #293; if that has not landed, this edits ci-wait's awk in the same shape.
Acceptance
- A branch whose required checks are all
cancelledis driven to green by re-runninglandalone, with no hand-minted SHA and nogh run rerun. failureis still red — a cancellation being recoverable must not make a real failure recoverable.
0969db8 to
251090a
Compare
251090a to
088b471
Compare
CLOUD-393's rolling lease serialises landing, but it was enforced entirely inside mise-tasks/land: anything else pushing to an already-ready PR bought a full matrix without ever touching the lock. Measured 2026-08-12 05:17-05:19Z, four concurrent pull_request matrices ran while the lease changed hands three times, every session holding the lease honouring it. The dominant case is residue rather than defiance — land re-drafts only on red, so a landing interrupted any other way leaves the PR ready permanently. So the runner asks too. A first step in every pull_request job that can start immediately fetches mise-tasks/ci-lease-precondition from main, which fetches land-lock from main and asks the one question a runner can ask: does the lease authorise this branch? If not it cancels the run it is standing in. Read from trunk rather than from the head being judged, which is what makes a stale clone unable to dodge the predicate by carrying a stale copy of it. The lease table cannot cover the case that motivated this. An agent running tooling older than 6d3d534 takes no lease, so the ref reads absent and every row authorises it — the lease is evidence of a free queue only among participants that can see the queue. So the precondition also asks whether this head's own mise-tasks/land acquires the lease, content-wise from one API read rather than by ancestry, which needs a deep fetch and answers wrongly after a rebase or a cherry-pick. A head that does not is stopped, and the message names the remedy: fetch, rebase, land. That narrows 420's "humans are unaffected by construction" promise to "humans on a branch newer than the lease" — the remedy is the one rebase they need in order to land anyway, but it is a real change to that guarantee rather than a side effect. Three things the refinement did not have, each found by measurement rather than reading, all detailed on CLOUD-420: final does NOT conclude cancelled, and the obvious remedy for that is worse than the problem. It declares always(), and an always() job runs even when the run is cancelled — observed on run 31566043914, where ci was cancelled and final executed anyway and failed its needs: assertion. So cancelling an unauthorised run reds the one check the host requires, which reads as a reason to skip the job instead. !cancelled() does skip it, and that is a false green: GitHub documents that a job skipped by its own if: "will report its status as 'Success'. It will not prevent a pull request from merging, even if it is a required check." Every lease-stopped run would then present a green required check over a head where nothing was compiled, tested, linted or analysed, and a maintainer's /fast-forward — whose stated safety property is that the exact SHA already has the required checks green — would land it. A false red costs a rebase; a false green lands untested code on main. success is the only conclusion GitHub lets a job report as fine, so a run that was DECLINED has to report something else, and red is the only something else a job that ran can produce. always() therefore stays, now carrying the reason it cannot change, so the next reader who notices that a stopped run reds final does not re-derive the skip and ship it. The misleading half of that red is answered where it belongs: the precondition emits ::error:: annotations naming the lease and the remedy. Those annotations are load-bearing rather than decorative, and they were being swallowed. The runner only reads a line as a workflow command when it begins with :: after leading whitespace is trimmed (actions/runner, ActionCommand.TryParseV2, line-anchored), and the stop message went through a helper that prefixes "lease-precondition: " — putting the token at column 20, where it is ordinary output. A stopped run would have been a cancelled run with a red final, no failed step, and nothing anywhere saying why. Both stop paths now emit at column 0 with the greppable prefix inside the annotation body, and both are pinned by cases that check every occurrence rather than the first. The exemption is needs:, not "has no checkout". final checks out too, so the refinement's rule would have demanded a precondition it cannot use. A fan-in cannot start before its dependencies are terminal, so it can never spend a runner ahead of the cancellation — which is the reason, and reasons survive a job being added where enumerations do not. The repository is private and several checkouts set persist-credentials: false, so the precondition carries its own credential through an http.extraheader — never a userinfo URL, since land-lock prints its remote when it cannot reach it. That is also what lets it run before any checkout exists, which is where it is cheapest. It never exits non-zero. A job that reds before its cancellation lands concludes the run failure rather than cancelled, which reds final and re-drafts the PR: the same failure mode arriving through its own remedy. Stopping is a cancel plus a bounded wait to be killed, and every fail-open row is a plain exit 0 — an unreadable script, an unreachable remote, a refused cancellation, an answer that is neither run nor stop. The asymmetry is the design: a predicate that cannot read itself would stop the whole fleet, where waving one matrix through costs one matrix. Two branch families are exempt, and not as a carve-out for bots. dependabot/* and release-plz-* land through a /fast-forward comment posted by a workflow that fires on workflow_run completed — which a CANCELLED run satisfies. It fires, finds the checks not green, and stops; nothing retries. So cancelling those runs would not save a matrix, it would defer one to the next rebase and add a stall to a landing path that is unattended by design. They are also the population least worth stopping: the release PR is a draft until the debounce readies it, and the debounce fires on 30 minutes of quiet main, which is exactly when the lease is free. ci-local-parity gains property 7 as the sensor, so a job cannot be added without one. Its fixture helper now emits the precondition by default — every existing case rested on a job whose first step was its work, so without that the new property would have reddened all of them for the wrong reason. Mutation-checked in both files: neutering the staleness row reds three cases and no others, reading exit 3 as run reds the acceptance case, exiting non-zero on the stop path reds every stop row, and removing the step from the fixture reds property 7 alone. The suite scrubs the ambient GITHUB_* environment, and that is not housekeeping. Every LEASE_* input has a GITHUB_* fallback, CI sets those and a developer box does not — so a case reaching "unset" by unsetting only the LEASE_ name behaves differently under Actions. One did: `no run id means there is nothing to cancel` went green locally and red in CI, where it fell through to GITHUB_RUN_ID and asked to cancel the very run it was executing in. Caught on a draft, at no CI cost beyond the run that found it, which is what drafts are for. The fallback is real behaviour rather than a mistake — inside a job GITHUB_RUN_ID names exactly the run a stop should cancel — so it is now pinned by its own case instead of merely avoided. Swept the suites for the shape: this was the only one exposed; tests/step-receipt.bats already scrubbed. Property 7 earned itself during this change's own rebase. A `semver` job landed on main while this was in flight, and the gate refused the merge until that job carried the step too — which is the difference between a gate and a paragraph, observed rather than argued. The step vs gate-job arithmetic is also re-derived: a gate job per workflow taxes a legitimate run 4 job-minutes and a stopped one 4, while the step taxes a legitimate run ~0 and a stopped one 7, so the step wins while stopped runs are fewer than 1.33x legitimate ones — not the 6:1 the refinement claimed. The gate job needs no actions: write and no cancel API at all, so the switch stays cheap if that ratio ever inverts. Refs: CLOUD-420, CLOUD-393, CLOUD-363 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01X4tzyT3Q3hXo5QFEESENYP
…signal CLOUD-426 fixed one half of this pair and its scope note named the rest: any other case asserting an outcome via a real race has the same defect, and a sweep for the shape belongs with the fix. This is that shape, in the sibling case, by a different mechanism. The case runs `timeout -k 1 5 "$LAND"` and asserted exit 124. The property it is about is that land was STILL POLLING at 5s, and 124 is not that property. GNU timeout returns 124 when the process dies from the TERM it sent, and 137 when -k had to escalate to KILL because TERM was not serviced within the grace second. land runs `set -m`, installs an EXIT trap and reaps two watcher process groups, so how long that takes is a function of machine load — and the window was one second. Observed inside `mise run verify` with four subagents contending for the box: red there, then 5 of 5 green in isolation, which is exactly the evidence that misleads. Both codes are timeout saying the command did not finish, so both are the property. This is not the widening CLOUD-426 refuses: a land that ends the lap on its own exits with its own status, so neither code can appear, and the case still reds when the poll is broken. Mutation-checked by making the poll break instead of waiting — the assertion fails, naming itself. Swept the suites for the same shape: this is the only instance. main-watch.bats also asserts 137, but it sends the SIGKILL itself, so that outcome is decided by the fixture rather than by scheduling. It rides along with CLOUD-420 rather than taking its own branch because it is what was blocking that branch's gate, and a separate landing for a four-line test change would buy a full matrix to fix a test that costs nothing to run. Refs: CLOUD-464, CLOUD-426, CLOUD-418 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01X4tzyT3Q3hXo5QFEESENYP
088b471 to
6dcd03b
Compare
|
|
/fast-forward |
…once had A bare `gh pr view` answers with the PR associated with this branch in ANY state, so once a branch name has carried a merged PR every later landing on that name binds the merged one. That is the default shape here rather than an edge case: trunk-based development deletes the branch on merge (CLOUD-349) while the session harness pins an agent to one branch name for its whole engagement, so the second landing of any session recycles a merged name. This branch carries eight PRs, seven of them merged. Observed rather than theorised: after #366 merged, a new commit and a new PR #368 on the same name produced `could not re-draft #366`. What kept that from being worse is incidental and worth naming, because it is the reason this is a fix and not a nicety — `redraft` runs before the wait loop and GitHub refuses to re-draft a merged PR, so the run died before reaching the terminal-state read, which treats MERGED as landed and exits 0. A bound-merged PR is one refactor away from reporting a landing that never happened, and a false completion signal is the one thing this repository exists to refuse. `gh pr list --head "$branch" --state open` instead, so "no open PR" falls through to the guard that already existed and names the fix. `branch` moves above the resolution because the query needs it. `// empty` rather than a bare field read: `--jq` prints the string "null" for a missing field, which is not empty and would sail past the guard as a PR number — `land` would then drive a pull request called "null" and block in the wait loop, which is a wedge rather than a failure. The stub had to learn to FILTER before either case could prove anything. First attempt returned the same body whether or not `--state open` was passed, so removing the flag left both cases green — a test that cannot see the defect it names. The real endpoint filters, so the stub now keeps two bodies: what an open-only query returns, and what an unfiltered one returns for a branch whose older PRs merged. Mutation-checked afterwards, which is the only reason this is stated as fact: dropping `--state open` reds both cases, and dropping `// empty` reds the no-open-PR case alone (with a non-empty list the fallback is not reachable, so the other case correctly stays green). Both resolution cases are bounded with `timeout`. The `// empty` mutant does not fail, it HANGS — `land` accepts "null" and waits forever — and an unbounded case wedges the whole file the way CLOUD-434's leaked stubs did. The bound turns a wedge into an ordinary failure the assertion can see. Refs: CLOUD-465, CLOUD-349, CLOUD-418 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01X4tzyT3Q3hXo5QFEESENYP



CLOUD-393's rolling lease serialises landing, but it was enforced entirely inside
mise-tasks/land: anything else pushing to an already-ready PR bought a full matrix without ever touching the lock. Measured 2026-08-12 05:17–05:19Z, four concurrentpull_requestmatrices ran while the lease changed hands three times — every session holding the lease honouring it.This adds the runner's half.
mise-tasks/ci-lease-preconditionruns as the first step of everypull_requestjob that can start immediately (8 jobs across 4 workflows), fetched frommainso a stale clone cannot dodge the predicate by carrying a stale copy of it. It asksland-lock authorises <branch>, and whether this head's ownmise-tasks/landtakes the lease at all — the row the lease table cannot cover, since a clone that takes no lease reads asabsentand would be waved through.Blast radius, measured before landing
11 of 15 open PRs carry a
mise-tasks/landthat never acquires the lease; 8 of those are non-draft. That is the population, not an edge case. On landing, the next push to each is stopped and told to rebase — work those branches already owe, sincemainhas moved past all of them.Five corrections to the refinement, each found by measurement
finaldoes not conclude cancelled — it runs and fails, and the obvious remedy is worse. It declaresalways(), and analways()job executes even when the run is cancelled: observed on run31566043914. So cancelling an unauthorised run reds the one check the host requires, which reads as a reason to skip the job instead.!cancelled()does skip it — and GitHub documents that a job skipped by its ownif:"will report its status as 'Success'. It will not prevent a pull request from merging, even if it is a required check." Every lease-stopped run would then show a green required check over a head where nothing was compiled, tested, linted or analysed. A false red costs a rebase; a false green lands untested code onmain.always()stays, now carrying the reason it cannot change.needs:, not "has no checkout".finalchecks out too; a fan-in cannot start before its dependencies, so it can never spend a runner ahead of the cancel.persist-credentials: false, so the precondition carries its own credential viahttp.extraheader— never a userinfo URL, sinceland-lockprints its remote when unreachable. Probed end-to-end against the real remote.dependabot/*andrelease-plz-*are exempt. A cancelled run iscompleted, so their/fast-forwardlanders fire, find no green, and stop with nothing retrying — cancelling would defer a matrix, not save one.The
::error::annotations are load-bearing rather than decorative, since a stopped run is a cancelled run with a redfinaland no failed step of its own. They were being swallowed: the runner only reads a line as a workflow command when it begins with::after trimming whitespace (actions/runner,ActionCommand.TryParseV2, line-anchored), and the stop message went through a helper that prefixeslease-precondition:, putting the token at column 20. Both stop paths now emit at column 0, pinned by cases that check every occurrence rather than the first.It never exits non-zero: a job that reds before its cancellation lands concludes the run
failurerather thancancelled, which redsfinaland re-drafts the PR — the same failure mode arriving through its own remedy.Verification
tests/ci-lease-precondition.bats, 16 cases. Mutants: neutering the staleness row reds 3 cases and no others; reading exit 3 as run reds the acceptance case; exiting non-zero on the stop path reds every stop row; removing the prefix exemption reds only its case.mise-tasks/ci-local-paritygains property 7 as the sensor, 4 new cases. Its fixture helper now emits the precondition by default — without that, every pre-existing case would have reddened for the wrong reason.semverjob landed onmainmid-flight and the gate refused the merge until that job carried the step too. Verified by mutation against the real workflows.main, where it does not yet exist, and a 404 fails open.Refs: CLOUD-420, CLOUD-393, CLOUD-363
Generated by Claude Code