Repository navigation
feat(sensor): record what a container replacement interrupted (CLOUD-451) - #402
Conversation
CLOUD-451 A turn that hands control to a human ends idle, so the VM is reclaimed while the human is still typing — and their input is destroyed
Waiting on a human is the one legitimate idle turn end, and it is the only one with no mechanism. What happensA session calls So the review step this repo depends on — a human reading a plan before an autonomous agent executes it — is not merely inconvenient, it is anti-correlated with care: the longer the human reads, the likelier their input is destroyed. Why nothing catches it
A hook cannot fix it alone. The hold has to be a real backgrounded tool call the harness knows about; a Mechanism
Refinement — ReadySpecializations only; the shared clauses live in Definition of Ready & Done.
Acceptance
Not in this issueWhether every other idle turn end should arm the same hold. The predicate would be broader and the enforcement point different ( |
1368a1b to
0b37777
Compare
…451) CLOUD-515 removed plan-hold because nobody could say whether occupancy defers the reclaim that destroys a human's typed approval. This collects the evidence that decides it, and nothing else: no deny, no gate, no occupancy, and nothing on the path to a human. Two events are on record and they point opposite ways. 2026-08-13 06:18:54: a container replaced with zero tracked tasks after ~14 minutes idle, disk preserved. 2026-08-12 ~22:12 (CLOUD-491): a container replaced while four tracked tasks were live, one a land doing network I/O every ten seconds. If reclaims only ever hit idle containers, occupancy is the right lever. If they kill active work too, it cannot help. Each is attested exactly once. reclaim-census rides land-lock hold's existing 30s beat, which ticks for exactly as long as a land holds the lease. One append and one sync per beat against a loop already doing a network push per beat. No new process and no new poll. The record is structural, not temporal, reusing CLOUD-491's finding rather than re-deriving it: the last 182s of writes before a replacement were measured not to survive, so a clock reads minutes stale exactly when asked. h per beat, x only where something chose to stop, verdict from the last line's kind. A loop the container killed cannot write an x, so a trailing h is the evidence. Five stop paths carry an x, not four. The four in the hold loop are the obvious ones; the fifth is land's drop_lease, which kills the heartbeat on a NORMAL finish so the loop never runs another statement. Without a record there, every successful landing would afterwards read as "the container died under active work" — the false positive that would wrongly license the mechanism just removed. Written from land rather than the heartbeat's trap, per CLOUD-491: a trap runs on the kill too, and an x from one erases the only distinction here. tests/reclaim-census.bats, 20 cases: every verdict row, and two that are not fixtures because they are the claim — a SIGKILLed loop leaves an h even though its trap runs, and a loop that stops on purpose leaves the matching x. Both declared mutations caught. Decide at 10 recorded replacements. Idle-only means occupancy is the lever and CLOUD-451 earns a mechanism; any that kill active work means it cannot, and the issue belongs upstream with this census as the evidence. Refs: CLOUD-451
0b37777 to
51684b4
Compare
|
|
/fast-forward |
AGENTS.md bans PR-webhook babysitting twice over — this task drives the landing loop "no timeout, no cap, never the PR webhook", and no heartbeat may babysit a PR — and the harness armed a subscription on every PR this repo opened anyway: #397, #402 and #489, three occurrences in five days. On #402 and #489 with no `subscribe_pr_activity` call behind it, so a `permissions.deny` row on the tool would close only the path nobody used. The remedy until now was an agent remembering: prose, therefore feedforward only, therefore exactly the half-change non-negotiable rule 2 refuses. `land` now drops it itself — once before the lap, and again after each ready it fires, since a `pull_request` event is what arms one. `unsubscribe_pr_activity` lives on the session's toolbox MCP server, an http endpoint under `/v2/ccr-sessions/` carrying no bearer token, so a JSON-RPC `tools/call` is a plain POST. Both header values are per-session and account-specific, so nothing here commits them: only the public endpoint SHAPE is tracked, and the volatile halves are read at run time out of the injected client config — the entry chosen by the tool it DECLARES rather than by a server id, which is CLOUD-191's endpoint-anchoring pattern applied to a second consumer. The repository is derived from the remote rather than declared, so no slug is a second authority. Failing open is the posture, not a caveat on it. Off harness there is no config, no subscription and nothing to drop, and the silence is the right answer; on harness, a refused or unreachable call costs the lap nothing, because ending a green landing over a nuisance subscription would be a worse defect than the one being closed. The status is dropped at the call site on purpose. The one thing never silent is a call that did not do what it says. Measured while building this, and it qualifies the mechanism rather than the design: a POST from a task is answered 401 by that endpoint today, so the drop currently reports `could NOT drop #N's webhook subscription (http 401)` instead of firing. Recorded on the issue with the reproduction. The shape is still the one the issue specifies, it costs a landing nothing when refused, and it says so rather than reading as done — which is the difference between this and the deny rule the issue already rejected as a placebo. Ten rows in tests/land.bats cover the test obligation: a subscribed PR is dropped and says so, an absent config and a toolbox without the verb are silent no-ops, and a 500, an unreachable endpoint, a JSON-RPC error and an `isError` result each cost the lap nothing while still being reported. Two `#MUTANT` declarations hold them: dropping the call reddens the first row, and reading the call's status at the call site reddens the fail-open row. Refs: CLOUD-518
AGENTS.md bans PR-webhook babysitting twice over — this loop runs on "no timeout, no cap, never the PR webhook", and no heartbeat may babysit a PR — and the harness arms a subscription on every PR this repo opens anyway: #397, #402 and #489, three occurrences in five days, two of them with no `subscribe_pr_activity` call behind them. So a `permissions.deny` row on the tool closes only the path nobody used, and the remedy until now was an agent remembering, which is prose and therefore feedforward only — the half-change non-negotiable rule 2 refuses. THE ACTOR DESIGN DOES NOT WORK, and this is not it. `land` making the call itself looks reachable: the tool is on the session's toolbox MCP server, an http endpoint under /v2/ccr-sessions/ carrying no bearer token, so a JSON-RPC tools/call looks like a plain POST. Measured 2026-08-18, that POST is answered 401 at both the toolbox and the github endpoint, with the injected config's own header values, and identically when forced through $HTTPS_PROXY: no_proxy carries anthropic.com, so requests to that host bypass the agent proxy and nothing injects a credential; the two headers are routing, not authorization. A first cut of this change shipped that POST behind a fail-open anyway. It removed zero subscriptions while its suite stayed green against a stubbed curl — a mechanism that reads as coverage and is not, which is CLOUD-418's defect rebuilt by hand. Filed as CLOUD-673. So this is `claim-check`'s inversion, the same one `issue-search-check` uses: the agent can do what the task cannot. The session's own tool call succeeds — that is how all three occurrences were remedied — so the agent unsubscribes, `pr-unsubscribed record <pr>` records that it happened for THIS pull request in THIS session from the tool's own answer, and `land` refuses to spend a runner until the record exists. The rule becomes an exit code without pretending to an effect nothing here can produce. Placement and posture: * The check is the FIRST thing `land` does, before the singleton and the lease, so a refusal costs no CI at all and the fix is one tool call away. * Off harness there is no injected config, therefore no session, therefore no subscription — `pr-unsubscribed` passes silently. That fail-open is what makes it safe on the critical path; a gate that cannot look must never become a gate that blocks everything. * Keyed by (session, PR), because a subscription belongs to that pair. A receipt from a previous container attests to nothing about this one, and #489's answer cannot satisfy #490 — the honest error this is built for, since the harness pins a session to one branch name for a whole engagement. * Pointer-only on stdout and in the receipt: the PR, the session and a digest of the answer. Never the answer, which is a message about a webhook stream. The honest limit, stated in the gate's own header: this proves the call was MADE for this PR, not that GitHub's subscription state is empty. Only the API answers that and reaching it is CLOUD-673. The claim receipt has the identical property, accepted there deliberately — the threat model is honest error, not fabrication. Ten rows in tests/pr-unsubscribed.bats cover both verbs: the refusal, the recorded drop, an answer naming the wrong PR, a receipt from another PR and from another session, empty stdin as could-not-look rather than a refusal, off-harness silence, pointer-only output, and bad arguments. Three rows in tests/land.bats cover the landing: the stop spends nothing (no ready, no push, no comment, no verify), the check names the PR being landed, and a passing gate leaves a lap unchanged. Four `#MUTANT` declarations, and the stopping-condition census moves 27 -> 28 — it caught the new stop the moment it was added, which is what it is for. Refs: CLOUD-518
AGENTS.md bans PR-webhook babysitting twice over — this loop runs on "no timeout, no cap, never the PR webhook", and no heartbeat may babysit a PR — and the harness arms a subscription on every PR this repo opens anyway: #397, #402 and #489, three occurrences in five days, two of them with no `subscribe_pr_activity` call behind them. So a `permissions.deny` row on the tool closes only the path nobody used, and the remedy until now was an agent remembering, which is prose and therefore feedforward only — the half-change non-negotiable rule 2 refuses. THE ACTOR DESIGN DOES NOT WORK, and this is not it. `land` making the call itself looks reachable: the tool is on the session's toolbox MCP server, an http endpoint under /v2/ccr-sessions/ carrying no bearer token, so a JSON-RPC tools/call looks like a plain POST. Measured 2026-08-18, that POST is answered 401 at both the toolbox and the github endpoint, with the injected config's own header values, and identically when forced through $HTTPS_PROXY: no_proxy carries anthropic.com, so requests to that host bypass the agent proxy and nothing injects a credential; the two headers are routing, not authorization. A first cut of this change shipped that POST behind a fail-open anyway. It removed zero subscriptions while its suite stayed green against a stubbed curl — a mechanism that reads as coverage and is not, which is CLOUD-418's defect rebuilt by hand. Filed as CLOUD-673. So this is `claim-check`'s inversion, the same one `issue-search-check` uses: the agent can do what the task cannot. The session's own tool call succeeds — that is how all three occurrences were remedied — so the agent unsubscribes, `pr-unsubscribed record <pr>` records that it happened for THIS pull request in THIS session from the tool's own answer, and `land` refuses to spend a runner until the record exists. The rule becomes an exit code without pretending to an effect nothing here can produce. Placement and posture: * The check is the FIRST thing `land` does, before the singleton and the lease, so a refusal costs no CI at all and the fix is one tool call away. * Off harness there is no injected config, therefore no session, therefore no subscription — `pr-unsubscribed` passes silently. That fail-open is what makes it safe on the critical path; a gate that cannot look must never become a gate that blocks everything. * Keyed by (session, PR), because a subscription belongs to that pair. A receipt from a previous container attests to nothing about this one, and #489's answer cannot satisfy #490 — the honest error this is built for, since the harness pins a session to one branch name for a whole engagement. * Pointer-only on stdout and in the receipt: the PR, the session and a digest of the answer. Never the answer, which is a message about a webhook stream. The honest limit, stated in the gate's own header: this proves the call was MADE for this PR, not that GitHub's subscription state is empty. Only the API answers that and reaching it is CLOUD-673. The claim receipt has the identical property, accepted there deliberately — the threat model is honest error, not fabrication. Ten rows in tests/pr-unsubscribed.bats cover both verbs: the refusal, the recorded drop, an answer naming the wrong PR, a receipt from another PR and from another session, empty stdin as could-not-look rather than a refusal, off-harness silence, pointer-only output, and bad arguments. Three rows in tests/land.bats cover the landing: the stop spends nothing (no ready, no push, no comment, no verify), the check names the PR being landed, and a passing gate leaves a lap unchanged. Four `#MUTANT` declarations, and the stopping-condition census moves 27 -> 28 — it caught the new stop the moment it was added, which is what it is for. Refs: CLOUD-518



Collects the evidence that decides whether CLOUD-451 is fixable in this repository at all. A sensor and nothing else — no deny, no gate, no occupancy, nothing on the path to a human.
Why
CLOUD-515 removed
plan-holdbecause nobody could say whether occupancy defers the reclaim that destroys a human's typed approval. Four issues and ~17 hours went into maintaining a mechanism whose premise was never once observed. This is the ordering that should have come first: instrument, read, then decide.Two events are on record and they point opposite ways:
landdoing network I/O every 10 sIf reclaims only ever hit idle containers, occupancy is the right lever and CLOUD-451 earns a mechanism. If they kill active work too, occupancy cannot help and the issue belongs upstream. Each is attested exactly once, which settles nothing.
How, at close to zero cost
mise-tasks/reclaim-censusridesland-lock hold's existing 30 s beat, which ticks for exactly as long as alandholds the lease — the window "active work is in flight". One append and onesync -dper beat, against a loop already doing a network push per beat. No new process, no new poll.That is the difference from what was deleted:
plan-holdwas a process whose only purpose was to be occupancy; this is a record on work that happens anyway.The record is structural, not temporal, reusing CLOUD-491's finding rather than re-deriving it: the last 182 s of writes before a replacement were measured not to survive, so any clock-based reading is minutes stale exactly when asked.
hper beat,xonly where something chose to stop, verdict from the last line's kind. A loop the container killed cannot write anx, so a trailinghis the evidence.The fifth stop path, which the plan missed
Four
xsites are the obvious ones inside the hold loop (holder gone, stalled, lease lost, lease lapsed). The fifth island'sdrop_lease, which kills the heartbeat on a normal finish — the loop never runs another statement, so its last record stays anh. Without a record there, every successful landing would afterwards read as "the container died under active work": the exact false positive that would wrongly license the mechanism just removed.Written from
landrather than the heartbeat's own trap, per CLOUD-491 — a trap runs on the container kill too, and anxfrom one would erase the only distinction the census draws.Verification
tests/reclaim-census.bats, 20 cases: every verdict row, the fail-open rows, and two that are not fixtures because they are the claim — akill -9ed loop leaves anheven though its trap runs, and a loop that stops on purpose leaves the matchingx.mise run mutant— 17 mutations across 11 gates, every one caught, including both new ones.mise run verifyrunning; result reported here before this is readied.Worth flagging for whoever writes the next sensor:
mutantstages the tree withgit ls-files, so a brand-new untracked suite is invisible to it and every mutation reportscase-already-red.git addfirst. The task's header says it works "while a gate and its suite are being written", which is true for modified tracked files and not for new ones.What this does not claim
Nothing about whether occupancy works. That is the question, not the answer.
Decide at 10 recorded replacements — a count, not a date. Idle-only → occupancy is the lever, and enforcement is worth reopening, cheapest form first. Any that kill active work → occupancy cannot prevent those; report the split and escalate with this census as the evidence.
A boundary worth stating, since three attempts have hit it: the user's typed-but-unsent approval text is client state and no repo mechanism can save it. At best this repo reduces how often a reclaim happens.
Refs: CLOUD-451
Generated by Claude Code