Skip to content

fix(land): linearize onto the main that is coming, and admit one successor (CLOUD-369) - #369

Merged
wenzowski merged 6 commits into
mainfrom
claude/cloud-369-grooming-fswo8x
Aug 12, 2026
Merged

wenzowski merged 6 commits into
mainfrom
claude/cloud-369-grooming-fswo8x

Conversation

@wenzowski

Copy link
Copy Markdown
Contributor

Stacked on #364 — branched from its head, not from main, so the residue builds on the cancel work rather than racing it in the same two files. #364 stays landable on its own. Rebase this onto main once it lands.

Why

The lease serialises landing correctly, and that is right for cost and wrong for latency: after every merge the queue is empty, and the next branch starts cold — rebase, verify, then a full matrix — before main can move again. Measured on #325: 8 laps, 8 greens, zero commits landed.

Pre-warming is a linearization, not a refresh. A waiter that rebases onto origin/main warms nothing — the branch holding the lease is about to replace that commit, so the waiter is stale again the instant it wins. The main worth linearizing against is the one about to exist.

The economic rule throughout: speculate maximally where it is free, exactly once where it is metered. Local execution — rebase, verify, conflict resolution — costs nothing, so every waiter does it continuously. A CI matrix is metered, so exactly one waiter buys one.

What changed

land-lock gains two advisory fields on the terms branch: already set (CLOUD-420) — read by CI and by waiters, never by a predicate that decides ownership:

  • head: — the commit about to become main, so a waiter can linearize onto the trunk that is coming.
  • next: — one admitted successor, written by the waiter itself via a new reserve verb. Forced rather than chosen: waiters are registered nowhere, so the holder cannot name one. It re-mints the holder's lease with a single field added — same holder id, expiry, branch, head — so the holder keeps holding and mine still answers for it.

authorises admits the holder and its one successor, so the bound is two whatever the fleet size, enforced at the runner rather than agreed between cooperating sessions. One CAS-guarded slot cannot hold two branches. renew/hold carry the reservation across each beat; acquire deliberately clears it, or the bound would drift upward one handover at a time.

land gains four things: speculative linearization onto the holder's head while waiting; a settled bet (won / pending / lost) at the top of every lap; the successor's ready+push without the lease; and a base re-confirmation inside the hold. LAND_LOCK_AGE carries the wait count into acquire, so a branch that has lost repeatedly probes a freed lease sooner — the capture effect is what turned 8 laps and 8 greens into zero landed.

The hazard, and the invariant that answers it

A speculative rebase puts another branch's unlanded commits into this branch's history. Fast-forwarding from there would land somebody else's unmerged work as a side effect of ours — far worse than a cold window. origin/main --is-ancestor HEAD does not catch it: the speculated base is itself a descendant of main, so that check passes for exactly the case that must fail.

So the bet is recorded and settled, never assumed, and settling runs before anything can push. There is no path from a losing bet to a push.

Two defects found by the tests rather than by review, both worth naming:

  • The in-hold check was first written as an ancestry query and had the flaw above. It is now "did main move since this lap rebased" — a sha comparison, the same question main-watch answers.
  • speculate re-bet on every waiting lap, overwriting the undo point with a HEAD that was itself speculative, so an unwind would have restored a tree still carrying another branch's commits. One outstanding bet per base now.

Tests

76/76 in tests/land.bats (11 new), 63/63 in tests/land-lock.bats (20 new). The coverage assertion caught both new families as designed — 16 stopping conditions and 11 lap-ending continues, up from 15 and 8.

The load-bearing cases are the negatives: a waiter that is not admitted stays in draft (without it, "every waiter readies" would pass and spend a matrix per session); a third branch is still refused while two are admitted; a lease whose branch/pr match but whose holder does not is not mine.

Refs: CLOUD-369, CLOUD-420, CLOUD-240


Generated by Claude Code

@linear-code

linear-code Bot commented Aug 12, 2026 •

Copy link
Copy Markdown
CLOUD-369 Admission control is half-landed: the lease serialises laps but empties the queue, starves losers, and leaves voided runs burning

Why

Rescoped 2026-08-12. The admission control this issue was filed for landed while it sat in Todo: land-lock is a CAS lease on refs/heads/batten-land-lock, written only by git push --force-with-lease, with a TTL, a heartbeat, a skew-immune steal clock and an orphan-PID tether. land acquires it before the ready/push pair, so the draft→ready transition — the entire CI bill — happens inside the hold. Three of the five original acceptance bullets are satisfied by it: confirming runs in flight are bounded at one regardless of N; the predicate needs no dispatcher and no shared state beyond one git ref; a loser backs off (jittered exponential) rather than readying anyway. The body below carries only what the lease does not do. The original framing is preserved in the comments and in the section at the foot.

The lease decides who goes first, and nothing more. That leaves three costs the original analysis named and the mechanism does not yet address.

1. The lease serialises whole laps, so the queue empties after every merge. A waiter blocks; when it wins it starts cold — rebase, then verify (~4 min), then ready, then push, then a full confirming matrix. During that window no sibling is ready-and-green behind it. This is the second property the original body called out as the reason this is not simply "ready less often": an admission control that only throttles trades CI minutes for latency, and a saturated queue was the point. land today only detects the adjacent condition — the lease_waits backstop reports "never won the landing lease in N attempts" and points at land-lock-check — which is diagnosis with no remedy behind it.

2. A branch can still starve. acquire's backoff is jittered exponential and a lease wait no longer costs a lap, but a branch that wins the lease and then loses the fast-forward re-enters on identical terms, with its probability of winning unimproved by the number of times it has lost. That is Ethernet's capture effect: aggregate throughput looks healthy — main advancing is the fleet working — while an individual station starves. LAND_MAX_LAPS (default 8) then converts starvation into abandonment: the branch stops having spent 8 verifies and up to 8 confirming matrices, and landed nothing. Measured 2026-08-11 on PR #325: 8 laps, 8 green CIs, 8 lost fast-forwards, zero commits landed.

3. A voided run burns until the next push happens to cancel it. When main-watch wins the CI race, land correctly stops waiting and laps — but the run keeps going. concurrency: cancel-in-progress kills it only when the next push reaches the same group, and the re-priced lease waits make that interval minutes, not seconds. Measured 2026-08-12 05:52Z: the lease read released while two pull_request matrices ran concurrently. CLOUD-240 refused hand-cancelling with a rule that should be read precisely — "supersede your own runs, never someone else's" — which permits cancelling our own run on our own ref and forbids reaching across refs.

The three interact, which is why they are one issue rather than three. Cancelling a voided run (3) without a warm queue (1) trades CI minutes for exactly the idle gap this issue exists to forbid; and a warm queue makes the aging rule (2) cheap, because a waiter that is already rebased-and-verified loses little by waiting longer.

Still not a request for a dispatcher. Coordination between sessions is impossible by construction: no lock beyond the lease ref, no coordinator, no shared state but the board and the remote.

Acceptance

  • A saturated queue, not a quiet one. After any merge, a ready-and-green-pending branch exists whenever at least one sibling had work verified — no idle gap where every sibling sits in draft. A session waiting on the lease rebases onto current origin/main and runs verify while it waits, so acquiring the lease costs only a ready and a push. This spends nothing new: CI does not run on drafts and local execution is free.
  • No branch starves. Admission order is fair or randomised-with-aging, so a branch's probability of admission rises with the time it has waited. No branch spends its lap budget and lands nothing while siblings land repeatedly. Deliberately weaker than FIFO — strict ordering needs a coordinator, which is ruled out; "does not starve" is achievable with a purely local rule.
  • A voided run is cancelled by its own owner. On the main-watch-won path only, the run this lap started on this branch's head is cancelled before lapping. It changes no verdict — the run was void by construction, which is why the lap ended — and it touches no other ref. Any path that could cancel another ref's run fails the gate. Implemented in flight by button-inc/batten#364 (draft, opened 2026-08-12 06:46Z), which scopes the cancel to this lap's head SHA and carries the two-case test — a voided lap cancels, a lap CI answered does not. Whoever picks this issue up inherits that work rather than repeating it; if perf(land): cancel the run main moved under, instead of billing it out (CLOUD-369) #364 lands first, this bullet is satisfied and only the first two remain.
  • The bound is holder + one admitted successor, and still does not depend on N. Raised from one deliberately, and it is the price of the first bullet: a queue is warm only if the next branch's confirming matrix overlaps the current merge instead of starting after it. The successor is admitted on the holder's CI going green — green and holding the lease makes its fast-forward near-certain, so the extra matrix is wasted only when a green holder still fails to land. Every other waiter stays in draft, locally linearized, spending nothing. The predicate stays computable from what a session can already observe, with no new shared state, no second lock, and no dispatcher.
  • Speculate maximally where it is free, exactly once where it is metered. Local execution — rebase, verify, conflict resolution — costs nothing, so every waiting session should be doing it continuously; a CI matrix is metered per minute, so only the branch that is plausibly next buys one. Pre-warming is a linearization operation, not a refresh: a waiter rebases onto the holder's head — the main that is about to exist — not onto the main the holder is about to replace.
  • Byte-stable, pointer-only reporting per non-negotiable rule 4: a count and a PR number, never a diff.

Explicitly rejected

  • A wall-clock timeout or a fixed sleep as the backoff. Ruled out repeatedly (ci-wait, land) — the fix for a race is an exit condition that can fire, never a guessed delay. Aging must be a function of observed state, not of elapsed seconds alone.
  • A WIP cap on sessions. Capping sessions costs pace of landed work, which is the thing being optimised for; capping concurrent readies costs nothing but a short wait.
  • Batching rebase → verify → push → land into one command to close the window. CLOUD-238 measured an agent inferring exactly this and recorded why it is wrong.
  • Cancelling anyone else's run. The narrow permission above is the whole of it.

What the lease already answered

The three questions this issue carried while it sat in Todo, and their answers, so nobody re-derives them:

  1. What is the sense predicate? The lease CAS itself. Holding refs/heads/batten-land-lock is "I am next"; failing to acquire is "I am not". No candidate needed measuring — a CAS cannot admit two.
  2. Who guarantees the queue never empties? Nobody yet. That is the first acceptance bullet above.
  3. Does this belong in land or beside it? In land, acquired before the ready/push pair, so one authority covers the whole window that spends CI.

Refinement — Ready (the residue of an admission control whose lease half already landed)

Refinement gate: Definition of Ready & Done. This body carries only specializations.

  • Source of truth (§1). mise-tasks/land and mise-tasks/land-lock are the one authority for admission; the lease semantics are not re-typed into a memory, a rule file, or this issue. Where this body and the tasks disagree, the tasks win and this body is stale. The warm-queue and aging behaviour extends the existing lease rather than introducing a second lock — a second admission authority is the failure mode, not an implementation option.
  • Computable predicate (§2). Each of the three criteria resolves to a command and an exit code over an object it decides, never a model verdict (non-negotiable rule 3): the warm-queue property is decided by whether a waiter's tree is rebased on the holder's published head — the anticipated next main — with a verify receipt for that exact HEAD, and by re-asserting descendancy inside the hold before any run is bought; a speculation that conflicts falls back to origin/main rather than failing, since a conflict against a base that may never land is information, not an error; the cancel path is decided by the run id and the ref it belongs to. "Plausibly next" is not re-litigated — it is lease ownership.
  • Effect (§3). Sensing and waiting are read. The cancel is a write against this ref's own run and belongs in the effect table as such; it is the only new outward-facing effect this change adds.
  • Output & exit (§5). Pointer-only: a count and a PR number, never a diff and never a run log. Exit codes follow the one contract — 0 admitted or clean, 1 usage, 2 the policy verdict, 3 unreachable — with no per-verb exception.
  • Commit / bump (§6). fix → patch until 0.1.0 (below 0.1.0 release-plz bumps the patch whatever the type says).
  • Test obligation (§7). bats coverage beside tests/land.bats and tests/land-lock.bats, each asserting an exit code, not prose: (a) a session that loses the acquire is rebased-and-verified for its current HEAD by the time it wins, so the ready/push pair is all that remains; (b1) with a green holder and several waiters, exactly one waiter readies and pushes and the rest stay in draft, plus the negative — a holder whose CI is not green admits nobody; (b) a branch that loses repeatedly is admitted within a bounded number of attempts — the anti-starvation property, exercised against a simulated faster-arriving main; (c) the main-watch-won path cancels the run on this branch's head and no other — covered by button-inc/batten#364, 104/104 across both files, mutation-checked; (d) no path can deadlock two waiters into permanent draft.
  • Blockers (§8). blockedBy none — the lease this builds on is landed and on main. relatedTo CLOUD-238 (the terminating edge this races against), CLOUD-240 (the rule that scopes the cancel), CLOUD-399 (the lap and wait budgets this re-prices), CLOUD-363 (a superseded run misread as red, the failure a cancel must not manufacture), and CLOUD-367 (the dispatch pattern that makes N large enough to matter).

CLOUD-420 The landing lease is enforced only by the code path that honours it, so an agent that skips `land` still spends a full matrix

Why

CLOUD-393 serialises landing behind a lease and cuts the discarded-CI-run rate. It is enforced entirely inside mise-tasks/land. An agent that runs gh pr ready by hand, pushes to an already-ready PR, or wraps landing in its own logic spends a full four-job matrix without ever touching the lock.

That is the failure this repository names on its front page: "A new rule without a runnable gate is half a change. Prose is feedforward only." The lease is a convention honoured by the cooperating path, and the threat model is the honest agent that does the wrong thing — CLOUD-200 records a session that satisfied gh-guard and then wrapped mise run land in a bespoke retry loop it invented on the spot. Nothing stops the same shape here.

The dominant case is residue, not defiance. Measured 05:17–05:19Z on 2026-08-12: four concurrent pull_request matrices (#356, #354, #332, #302) while the lease changed hands three times, so the lease was being honoured by every session that held it and all four matrices ran anyway. land re-drafts only on red, so a landing interrupted any other way — a container reclaim, a stopped task, a lost lease, a die on a rebase conflict — leaves the PR ready permanently, and every later push to it buys a full matrix with no landing attempt in progress at all. All three of the other PRs were draft=false. The implementation is scoped to that case; an agent that skips land deliberately is the same mechanism and the smaller share.

The enabling gap. The lease identifies a clone (hostname-pid-random), which CI cannot check anything against. Adding the branch it authorises to the lease body is what makes it verifiable by the thing spending the money.

Refinement — Ready

  • Source of truth (§1). The lease ref refs/heads/batten-land-lock — a parentless commit over the empty tree whose body is land-lock / holder: / expires: / nonce: — read server-side by the workflow that is about to spend the runner. This issue adds a fifth line, branch:, and land-lock-check gains its well-formedness assertion rather than a second gate growing beside it. Not a receipt in a clone, which is exactly the artefact an agent skipping land never writes.

  • Mechanism (§3). land-lock gains a read-only verb, authorises <branch> — a pure function of (lease state, branch) reusing the existing observe / expired / released predicates. The alternative, re-parsing the body in workflow YAML, is a second authority for the ref and silently drops the read semantics the lease was pressure-tested into: a per-process observation ref, an unparseable body reading as held, and an unreachable remote that is not "free". Exit 0 run / 3 stop / 2 could not look.

    Each expensive pull_request job gains a first step, before the real checkout: a one-file sparse checkout of mise-tasks/land-lock at main (fetch-depth: 1, ~2s) and one invocation of it. The predicate comes from trunk, never from the head under test — that head has not been graded yet.

lease state action
absent, released, or expired-and-corroborated run — nobody is landing
branch: is this branch run — this is the authorised attempt
branch: names another branch stop, having spent only the spin-up
unreachable, unparseable, or carrying no branch: run, with a warning — fail open

Nothing in the table asks why a push happened, which is what makes it cover the residue case: the precondition is per job rather than per landing, so a push to a PR left ready by an interrupted land is stopped by the same row that stops a hand-typed gh pr ready.

The last row is the design and not a fallback: failing open costs one matrix, while failing closed on an unreadable ref stops every PR in the fleet, and a body minted before this change ages out within one TTL (120s).

  • A step, not a preflight job (§3), and the arithmetic that decides it. In billed minutes the two are closer than a wall-clock reading suggests: a gate job costs 1 job-minute per stopped run, because the expensive jobs never start and a skipped job bills nothing, while a first step in each job costs 7. Against that, the gate taxes every legitimate run one extra job-minute plus ~15-20s of serial spin-up. The step is therefore cheaper exactly while legitimate runs outnumber stopped ones better than 6:1, and lap length sits in the exponent of the collision model this whole line of work exists to shrink, which decides it at parity. That ratio is measurable rather than assumed; it is the number to re-check if the residue rate stays high.

  • THE HAZARD — a stop is a third answer, and the verdict layer has two (§2). A first step cannot make a job skipped; only a job-level if: can, which is the preflight job ruled out above. So the reachable conclusions are success, failure and cancelled, and checks-green reads success or neutral as green, failure and cancelled as red, and a latest-run skip as "not an answer".

    • Success or neutral grades a head green that compiled nothing, and main advances to exactly the SHA that passed. Disqualifying.
    • Failure re-drafts healthy PRs fleet-wide, since land re-drafts on red. Disqualifying.
    • Cancelled is the one that can be made correct. "This run declined to answer" is a third thing the verdict vocabulary lacks, and CLOUD-436's latest-run-per-name grouping is what makes a per-name declination readable at all.

    So: the step self-cancels the run (POST /actions/runs/{id}/cancel, actions: write on that job alone). Reading cancelled as a declination rather than as red is CLOUD-363's change, in flight on fix(tasks): read a cancelled required check as no verdict, not as red #302, which moves it into skipped's not-an-answer bucket in checks-green and in land's graded_runs together. This issue consumes that predicate rather than restating it, and that is why it is a blocker rather than a note: a stop landing first would re-draft healthy PRs fleet-wide, which is the hazard this bullet exists to avoid.

    Two consequences, both load-bearing:

    • Branch protection requires only final, which concludes cancelled too, so an unauthorised head cannot be merged while it waits.
    • Because graded_runs reads a cancelled set as ungraded, the stopped head's next land lap re-fires its ready and mints a real run, so the re-dispatch this would otherwise have to specify arrives with CLOUD-363's second half. It cannot loop: land acquires the lease before it pushes, so a land-driven push is never the one stopped.
  • Humans are unaffected by construction (§2). The "absent or released" row is what preserves them: a person pushing while no agent holds the lease sees no change. They wait only in the case where their run would have been voided anyway, which is a saving rather than a tax. fast-forward.yml is not touched.

  • Below it, for free (§3). Extend ready-guard — which already denies gh pr ready without verify and linear-check receipts — to require a lease.<branch> receipt, written by land-lock acquire into the same receipts directory. Catches the honest mistake at zero CI cost. A fail-fast convenience, not the guarantee: a hook can be unloaded (CLOUD-187) or bypassed, which is the whole reason the server-side check is the load-bearing half.

  • Above it (§3). A ci-local-parity property asserting that every pull_request job which checks the repository out carries the precondition as its first step, so a new job cannot silently omit it. The exemption is computable rather than enumerated: a job with no actions/checkout — final, the fan-in — has no work to protect.

  • Engine gap, per §2 (§3). That property is a predicate over workflow structure, which no rule kind can address, so it extends ci-local-parity instead of landing in batten.toml. The gap is CLOUD-452, linked rather than left as an exception.

Cost (§1), in the unit the invoice uses. This repository is private and every pull_request job is a GitHub-hosted ubuntu-latest runner, so the metered unit is a job-minute rounded up per job, not a second of wall clock. Measured p95 per job, 2026-08-12: ci 429s, darwin-link 97s, cross 82s, commit-lint 65s, msrv 60s, final 14s, action 12s — ~759s of runner time, which bills as 17 job-minutes per full run.

  • The tax on a legitimate run is ~2s in each of 7 jobs, which per-job rounding absorbs: 0 billed minutes, unless a job already sits within 2s of a minute boundary.
  • A stopped run still bills its floor — 7 jobs x 1 minute = 7 job-minutes against the 17 it replaces. That is ~59% of the invoice and ~86% of the wall clock: real, and smaller than a seconds-based reading implies.
  • The cheapest option for the residue case spends nothing at all, because a push to a PR left draft triggers no job: CLOUD-458. This issue is the backstop for what still leaks past that, so it must not be sized as though it carried the volume.

Local execution is the unmetered tier and nothing here moves work onto the metered one. The CI-side check exists because the local one is the half an interrupted session never reaches.

Test obligation

  • tests/land-lock.bats — the four rows above as a pure function of (lease body, branch), plus the two that must not be conflated: unreachable is not absent, and a missing branch: is not a foreign holder.
  • tests/land.bats — the composition rather than the predicate, which CLOUD-363 owns and tests: a head whose required runs were stopped by the precondition reads as ungraded, so the next lap re-fires the ready instead of re-drafting. Mutation-checked per CLOUD-418: with the stop step removed from the fixture workflow, that head grades normally.
  • tests/ci-local-parity.bats — both directions: a checkout-bearing pull_request job without the precondition is refused, and a job with no checkout is not.
  • tests/land-lock-check.bats — a lease body with no branch: line is reported.

Commit / bump (§6): feat(ci) — patch until 0.1.0 regardless of type.

Blockers (§8): blockedBy CLOUD-363 — the stop conclusion is safe only once a cancelled required check reads as no verdict rather than as red, which is in flight on #302. CLOUD-393 landed in #340, so the lease exists, and adding branch: to its body is this issue's own first step; CLOUD-436 landed in #347, supplying the latest-run-per-name grouping the declination sits on.

Acceptance

  • A PR pushed while another branch holds the lease stops within seconds instead of running a matrix.
  • A PR pushed while the lease is free, or while its own branch holds it, runs unchanged.
  • A stopped run neither reds the PR nor leaves its head unrecoverable: the branch's own next lap mints a real run without a hand-minted SHA.
  • A push to a PR left ready by an interrupted land is stopped by the same row as a hand-typed ready, with no landing attempt in progress.
  • An unreachable lease, or one carrying no branch:, runs the matrix rather than stopping the fleet.
  • A new checkout-bearing pull_request job cannot omit the precondition without failing ci-local-parity.
  • land-lock-check fails a lease body missing branch:.

CLOUD-240 The landing loop still spends CI minutes it does not need, and still wakes a human for work bash can do

Why

CLOUD-238 made land drive the loop. It did not make the loop cheap, and CI minutes are metered while local execution in an agent's sandbox is free. Three leaks and one missing signal, all measurable:

  1. Draft PRs still burn CI. ci.yml and commit-lint.yml gate every job on github.event.pull_request.draft == false — the whole point of "iterate at zero CI cost". zizmor.yml does not: it triggers on any PR touching .github/workflows/**, draft or not. Every push to a draft that edits a workflow spends a runner. Worse, it means re-drafting a PR does not stop all CI, so the obvious response to a red run does not actually close the tap.
  2. A lap's stale CI run is not cancelled. ci.yml has concurrency: cancel-in-progress, so pushing a rebase kills its own in-flight run. commit-lint.yml and zizmor.yml have no concurrency group at all, so their runs from the superseded SHA run to completion for a verdict nobody will read.
  3. A run doomed by main moving is watched to the end. land waits on ci-wait alone. The moment main advances, this lap's SHA can no longer fast-forward — the run in flight is already waste, and every second of it is billed. Nothing notices until CI finishes and the bot refuses.
  4. A red CI leaves the PR ready. land dies and the PR stays non-draft, so the next push — from any source — starts CI again while the failure is still being diagnosed locally.

There is also a feedforward gap behind #1: nothing asserts that everything CI runs is something verify runs. Today the sets happen to match (ci, cross-check, darwin-link, zizmor, commit-lint are all in verify's closure), and that is exactly the state a gate should freeze — the workflow contract already promises "one it runs that you can't is a bug", with no mechanism.

Acceptance

  • No workflow runs a job on a draft PR
  • Every PR-triggered workflow cancels its own superseded in-flight run
  • land watches main concurrently with ci-wait and starts the next lap the moment main moves, instead of waiting out a run whose verdict is already void; the push that follows cancels the stale run
  • The main watch is conditional (If-None-Match), so a quiet main costs no rate limit — the same discipline ci-wait already uses
  • A red CI converts the PR back to draft before land stops, so no further push spends a runner while the failure is diagnosed locally
  • A lap whose HEAD is already verified does not re-run verify
  • A gate fails if a PR-triggered workflow invokes a mise run <task> that verify does not
  • The only thing that wakes a human is a rebase conflict, a failed local verify, or a red CI — never a refusal, never a moved main

Refinement — Ready (make the loop cheap, and make the cheapness a gate)

Refinement gate: Definition of Ready & Done. This body carries only specializations.

  • Source of truth (§1). The workflow files are the authority on what CI spends; [tasks.verify] in mise.toml is the authority on what an agent can run locally. The parity gate compares those two texts and nothing else — no network, no run history.
  • Computable predicate (§2). (a) Every jobs.* in a pull_request-triggered workflow carries the draft guard, and every such workflow carries a concurrency block with cancel-in-progress: true — grep-level over the committed YAML. (b) Every mise run <task> in those workflows appears in [tasks.verify]'s depends or body. (c) main-watch exits 0 when repos/{owner}/{repo}/commits/main reports a SHA other than the one it was given, and blocks otherwise; land races it against ci-wait with wait -n and treats whichever finishes first as the verdict.
  • Effect (§3). main-watch is read. land stays write, and gains one more write — gh pr ready --undo on a red run.
  • Output & exit (§5). Pointer-only, one line per transition. main-watch: main-watch: main moved <old> -> <new> and exit 0. land on red: one ::error:: naming the failing checks plus re-drafted #N. Exit 0 merged, non-zero otherwise.
  • Commit / bump (§6). fix → patch until 0.1.0 (DoR §6: below 0.1.0 release-plz bumps the patch whatever the type says).
  • Test obligation (§7). tests/main-watch.bats over a stubbed gh that answers 304 then a changed SHA, asserting the conditional request carries If-None-Match and that an unchanged main blocks. tests/land.bats gains: main moving mid-wait starts the next lap without waiting for CI; a red CI re-drafts before exiting; an already-verified HEAD skips verify. tests/ci-local-parity.bats over fixture workflow texts: a job with no draft guard fails, a missing concurrency fails, a mise run task absent from verify fails, and the repo's real workflows pass.
  • Blockers (§8). None. Builds on CLOUD-238, which has landed.

Review in Linear

claude added 6 commits August 12, 2026 13:37
The lease bounds confirming runs at one. That is right for cost and wrong
for latency: after every merge the queue is empty, and the next branch
starts cold — rebase, verify, then a full matrix — before `main` can move
again. Measured on PR #325: 8 laps, 8 greens, zero commits landed.

Two advisory fields, on the terms `branch:` already set (CLOUD-420) —
read by CI and by waiters, never by a predicate that decides ownership:

- `head:` names the commit about to become `main`, so a waiter can
  linearize onto the trunk that is COMING rather than onto the one the
  holder is about to replace. Rebasing onto current `origin/main` warms
  nothing; it pays the same staleness earlier.
- `next:` names one admitted successor, so the matrix that overlaps the
  holder's merge is bought instead of started cold afterwards.

The waiter writes `next:` itself through `reserve`, which is forced
rather than chosen: waiters are registered nowhere, so the holder cannot
name one. It re-mints the holder's lease with a single field added —
same holder id, same expiry, same branch, same head — so the holder
keeps holding and `mine` still answers for it. One CAS-guarded slot
cannot hold two branches, so the bound is two whatever the fleet size,
and `authorises` enforces it at the runner rather than by convention.

`renew` and `hold` carry the reservation across each beat, or a holder
would erase it within 30 seconds of a waiter writing it. `acquire`
deliberately does not: a fresh turn whose predecessor's successor was
carried forward would authorise a third branch, then a fourth.

Refs: CLOUD-369, CLOUD-420

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0172rfy7QRfyCEgw8KFYfBVJ
…essor

A waiter that rebases onto `origin/main` warms nothing: the branch
holding the lease is about to replace that commit, so the waiter is
stale again the instant it wins — the cold window paid earlier and no
cheaper. Pre-warming is a linearization, and the base worth linearizing
against is the one the lease publishes as `head:`.

Four changes to the lap, all inside the existing lease:

- Lost the lease: rebase onto the holder's head and re-verify there, so
  this branch's turn costs a ready and a push rather than a rebase, a
  verify and a matrix. Every failure falls back to `origin/main` — a
  conflict against a base that may never land is information, not the
  one real decision the lap stops for.
- The bet is recorded and settled, never assumed. A speculative rebase
  puts another branch's unlanded commits into this history, and
  fast-forwarding from there would land somebody else's work as a side
  effect of ours. `origin/main --is-ancestor HEAD` does not catch it:
  the speculated base is itself a descendant of main. So `spec_base` is
  settled at the top of every lap, before anything can push — won keeps
  the tree, pending keeps it, lost resets it.
- Reserve the successor slot and, if admitted, ready and push without
  the lease so that matrix overlaps the holder's merge. Everything past
  the push needs the lease, so the pass ends there rather than
  duplicating the ready/push pair CLOUD-254 owns.
- Re-confirm inside the hold. `acquire` waits up to a TTL, so the winner
  is at its most stale in the instant it wins; readying there buys a
  matrix the fast-forward will refuse. Asked as "did main move since
  this lap rebased", not as an ancestry query, for the reason above.

`LAND_LOCK_AGE` carries the wait count into `acquire`, so a branch that
has lost repeatedly probes a freed lease sooner — the capture effect is
what turned 8 laps and 8 greens into zero commits landed on #325.

`fetch_main` and `charge_wait` are extracted rather than copied: both
now have two call sites, and a duplicated `die` is a second authority on
what the failure means.

Refs: CLOUD-369

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0172rfy7QRfyCEgw8KFYfBVJ
Two defects, one in the task and one in its harness, and the harness one
is why the task one looked like nine failures.

The task: an admitted successor waits many laps by design, and it
re-entered the ready/push pair on every one of them. Pushing an unchanged
head emits no `synchronize`, so it buys nothing and drops into the
`--undo` re-fire path that exists for a different case entirely. Gated on
the head actually pushed, so a speculation that moves HEAD still buys the
run the new commit needs.

The harness: the git stub answered every `merge-base` from one file. The
lap asks "am I a descendant of main"; the speculation asks "is the
holder's head already in my history" and "did the base I bet on land".
One rc made the second answer yes by accident, so `speculate` returned
early every time and the whole chain below it failed for one reason
upstream of all of them.

The coverage assertion caught both new families as it is meant to: 16
stopping conditions and 12 lap-ending continues, up from 15 and 8.

Refs: CLOUD-369

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0172rfy7QRfyCEgw8KFYfBVJ
… the suite does

Four assertion defects of my own, none in the task under test:

- `call_order` joins with a trailing space, so an equality check against
  "ready push" could never pass. The suite's own convention is a prefix
  match; the once-per-head property is asserted on the push COUNT instead,
  which is what that case is actually about.
- `lock_calls` anchored the whole line, so it counted zero for any verb
  taking an argument — `reserve <branch>` among them.
- The continue counter is 11, not 12: the assertion matches a bare
  `continue` at line start, and I counted with a looser pattern.
- The main-moves lever fired before the bet was placed, so the bet read as
  already-decided and no speculation was ever settled. It now takes the
  acquire number to fire from: bet first, then the world moves under it,
  which is the sequence the unwind exists for.

Refs: CLOUD-369

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0172rfy7QRfyCEgw8KFYfBVJ
…erminate

The lever moved trunk once. The next lap's re-confirmation then passed,
land proceeded to the answer poll, and with no terminal PR state that
poll never returned — the case hung and leaked a main-watch process on
every run. "Main is moving faster than a lap takes" is the condition
under test, so the lever now writes a distinct sha per acquire and the
case gets a wait budget it can actually exhaust.

Refs: CLOUD-369

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0172rfy7QRfyCEgw8KFYfBVJ
A waiter laps repeatedly while the same holder lands, and `speculate`
re-bet on every one of those laps. Each re-bet overwrote `spec_undo`
with a HEAD that was itself speculative — so unwinding would have
restored a tree still carrying another branch's commits, which is
exactly the hazard the undo exists to remove. It also re-minted a sha
per lap and discarded a verify receipt for no gain, since the base had
not changed.

The bet is now placed once per base, and `spec_undo` is only ever this
branch's last non-speculative HEAD. Settling clears both, so a bet
placed after a settled one records the right undo point.

Found by the two unwind cases, which could not reach the lost branch at
all while every lap re-bet: the bet was perpetually pending.

Refs: CLOUD-369

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0172rfy7QRfyCEgw8KFYfBVJ
@wenzowski
wenzowski force-pushed the claude/cloud-369-grooming-fswo8x branch from 8faf69c to 3f9058d Compare August 12, 2026 13:39
@sonarqubecloud

Copy link
Copy Markdown

@wenzowski
wenzowski marked this pull request as ready for review August 12, 2026 13:49
@wenzowski

Copy link
Copy Markdown
Contributor Author

/fast-forward

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants