Repository navigation
revert(hooks): remove plan-hold, whose occupancy was never shown to work (CLOUD-515) - #397
Conversation
CLOUD-515 `plan-hold` gates every path to a human on occupancy that was never shown to work: remove it, and let a `spanned` reading buy it back
Why
The cost is certain. The benefit has never been observed. CLOUD-491 measured a live hold on 2026-08-12 ~22:12 UTC — the container restarted regardless, killing four tracked tasks including a That reading does not exist. Measured in this clone, 2026-08-13: No hold has ever completed a poll here. The sensor CLOUD-491 shipped to grade CLOUD-451's acceptance has zero data, and CLOUD-451 has been In Review since 2026-08-12 13:15 with its acceptance — "a plan left open for longer than the observed reclaim window survives, with the human's typed input intact" — never met. And the mechanism keeps generating its own defects. CLOUD-485 (the guard and the release listened to different events, so the answer never released the hold), CLOUD-511 (two handoffs in one turn leave the second unguarded), and CLOUD-500's stage 1 are three open follow-ups on a mechanism with no demonstrated effect. The surface is 6 tasks, 3 bats suites, 3 hook registrations and 2 This is the non-negotiable-2 shape read backwards: not a rule shipped without a mechanism, but a mechanism shipped without evidence its object exists, and gates decide rather than estimate. Refinement — Ready
What buys it backNot an argument and not a date. One Acceptance
Not in this issueWhether the reclaim problem is real. It is — CLOUD-451's report from the repo owner stands, and destroying a human's typed approval is the defect. This removes one unvalidated remedy, and CLOUD-500 stage 0 is still the cheapest way to learn whether occupancy is the lever at all. |
…ork (CLOUD-515) `plan-hold-guard` was a PreToolUse deny on `ExitPlanMode|AskUserQuestion`: every handoff to a human paid a refused tool call, a second call to arm a four-hour sleeper, and a retry. The premise was that occupying the container defers the reclaim that destroys a human's typed approval. That premise was measured twice and failed twice. CLOUD-491 recorded a live hold on 2026-08-12 ~22:12 UTC and the container was restarted anyway, killing four tracked tasks including a `land` mid-CI-wait; it reproduced at ~23:45. CLOUD-500 conceded the point and declared its own stage 1 "Ready but not startable" until one `plan-hold-check spanned` = 0 reading exists. No such reading exists. Measured in this clone 2026-08-13: `batten-hold-heartbeat` absent, `batten-boots` carrying one line — no hold has ever completed a poll here. CLOUD-451 has been In Review since 2026-08-12 13:15 with its acceptance never met, and has spawned three open follow-ups (CLOUD-485, CLOUD-511, CLOUD-500's stage 1) on a mechanism with no demonstrated effect. A cost that is certain against a benefit that has never been observed is not a gate; gates decide. Removed: six `mise-tasks/plan-hold*`, three bats suites, three hook registrations (the PreToolUse deny, the UserPromptSubmit release, the whole PostToolUse block), two MUTANT_GATES rows, and the session-start `spanned` report — which reads a heartbeat only the hold writes, so with no hold it can only answer "cannot look". The problem CLOUD-451 names is real and unsolved; this removes an attempt at it, not the finding. The code is one `git revert` from returning, and its price is one `spanned` = 0 reading recorded on CLOUD-500. Refs: CLOUD-515
…LOUD-515)
Removing plan-hold deleted its three bats suites, dropping tests/**/*.bats from
1600 to 1552 and denying at bats-tests-not-deleted. The rule's own no_fix_reason
names the two exits — "restore the tests, or waive the reduction deliberately" —
and restoring is not one: the suites cover code that no longer exists.
Unnarrowed, and not for want of trying. A ratchet finding's path is the glob plus
the counts (rules.rs builds "{glob} {base}->{working}"), because the finding is
about the whole matched set rather than any one file. There is no per-file path a
`path =` could select, and a waiver written against the literal would embed
1600->1552 and lapse the moment either count moved.
It is expected to lapse unused: base is origin/main, so once this lands the floor
becomes the new count on its own.
Refs: CLOUD-515
e6e83a1 to
15b8820
Compare
|
|
/fast-forward |
AGENTS.md bans PR-webhook babysitting twice over — this task drives the landing loop "no timeout, no cap, never the PR webhook", and no heartbeat may babysit a PR — and the harness armed a subscription on every PR this repo opened anyway: #397, #402 and #489, three occurrences in five days. On #402 and #489 with no `subscribe_pr_activity` call behind it, so a `permissions.deny` row on the tool would close only the path nobody used. The remedy until now was an agent remembering: prose, therefore feedforward only, therefore exactly the half-change non-negotiable rule 2 refuses. `land` now drops it itself — once before the lap, and again after each ready it fires, since a `pull_request` event is what arms one. `unsubscribe_pr_activity` lives on the session's toolbox MCP server, an http endpoint under `/v2/ccr-sessions/` carrying no bearer token, so a JSON-RPC `tools/call` is a plain POST. Both header values are per-session and account-specific, so nothing here commits them: only the public endpoint SHAPE is tracked, and the volatile halves are read at run time out of the injected client config — the entry chosen by the tool it DECLARES rather than by a server id, which is CLOUD-191's endpoint-anchoring pattern applied to a second consumer. The repository is derived from the remote rather than declared, so no slug is a second authority. Failing open is the posture, not a caveat on it. Off harness there is no config, no subscription and nothing to drop, and the silence is the right answer; on harness, a refused or unreachable call costs the lap nothing, because ending a green landing over a nuisance subscription would be a worse defect than the one being closed. The status is dropped at the call site on purpose. The one thing never silent is a call that did not do what it says. Measured while building this, and it qualifies the mechanism rather than the design: a POST from a task is answered 401 by that endpoint today, so the drop currently reports `could NOT drop #N's webhook subscription (http 401)` instead of firing. Recorded on the issue with the reproduction. The shape is still the one the issue specifies, it costs a landing nothing when refused, and it says so rather than reading as done — which is the difference between this and the deny rule the issue already rejected as a placebo. Ten rows in tests/land.bats cover the test obligation: a subscribed PR is dropped and says so, an absent config and a toolbox without the verb are silent no-ops, and a 500, an unreachable endpoint, a JSON-RPC error and an `isError` result each cost the lap nothing while still being reported. Two `#MUTANT` declarations hold them: dropping the call reddens the first row, and reading the call's status at the call site reddens the fail-open row. Refs: CLOUD-518
AGENTS.md bans PR-webhook babysitting twice over — this loop runs on "no timeout, no cap, never the PR webhook", and no heartbeat may babysit a PR — and the harness arms a subscription on every PR this repo opens anyway: #397, #402 and #489, three occurrences in five days, two of them with no `subscribe_pr_activity` call behind them. So a `permissions.deny` row on the tool closes only the path nobody used, and the remedy until now was an agent remembering, which is prose and therefore feedforward only — the half-change non-negotiable rule 2 refuses. THE ACTOR DESIGN DOES NOT WORK, and this is not it. `land` making the call itself looks reachable: the tool is on the session's toolbox MCP server, an http endpoint under /v2/ccr-sessions/ carrying no bearer token, so a JSON-RPC tools/call looks like a plain POST. Measured 2026-08-18, that POST is answered 401 at both the toolbox and the github endpoint, with the injected config's own header values, and identically when forced through $HTTPS_PROXY: no_proxy carries anthropic.com, so requests to that host bypass the agent proxy and nothing injects a credential; the two headers are routing, not authorization. A first cut of this change shipped that POST behind a fail-open anyway. It removed zero subscriptions while its suite stayed green against a stubbed curl — a mechanism that reads as coverage and is not, which is CLOUD-418's defect rebuilt by hand. Filed as CLOUD-673. So this is `claim-check`'s inversion, the same one `issue-search-check` uses: the agent can do what the task cannot. The session's own tool call succeeds — that is how all three occurrences were remedied — so the agent unsubscribes, `pr-unsubscribed record <pr>` records that it happened for THIS pull request in THIS session from the tool's own answer, and `land` refuses to spend a runner until the record exists. The rule becomes an exit code without pretending to an effect nothing here can produce. Placement and posture: * The check is the FIRST thing `land` does, before the singleton and the lease, so a refusal costs no CI at all and the fix is one tool call away. * Off harness there is no injected config, therefore no session, therefore no subscription — `pr-unsubscribed` passes silently. That fail-open is what makes it safe on the critical path; a gate that cannot look must never become a gate that blocks everything. * Keyed by (session, PR), because a subscription belongs to that pair. A receipt from a previous container attests to nothing about this one, and #489's answer cannot satisfy #490 — the honest error this is built for, since the harness pins a session to one branch name for a whole engagement. * Pointer-only on stdout and in the receipt: the PR, the session and a digest of the answer. Never the answer, which is a message about a webhook stream. The honest limit, stated in the gate's own header: this proves the call was MADE for this PR, not that GitHub's subscription state is empty. Only the API answers that and reaching it is CLOUD-673. The claim receipt has the identical property, accepted there deliberately — the threat model is honest error, not fabrication. Ten rows in tests/pr-unsubscribed.bats cover both verbs: the refusal, the recorded drop, an answer naming the wrong PR, a receipt from another PR and from another session, empty stdin as could-not-look rather than a refusal, off-harness silence, pointer-only output, and bad arguments. Three rows in tests/land.bats cover the landing: the stop spends nothing (no ready, no push, no comment, no verify), the check names the PR being landed, and a passing gate leaves a lap unchanged. Four `#MUTANT` declarations, and the stopping-condition census moves 27 -> 28 — it caught the new stop the moment it was added, which is what it is for. Refs: CLOUD-518
AGENTS.md bans PR-webhook babysitting twice over — this loop runs on "no timeout, no cap, never the PR webhook", and no heartbeat may babysit a PR — and the harness arms a subscription on every PR this repo opens anyway: #397, #402 and #489, three occurrences in five days, two of them with no `subscribe_pr_activity` call behind them. So a `permissions.deny` row on the tool closes only the path nobody used, and the remedy until now was an agent remembering, which is prose and therefore feedforward only — the half-change non-negotiable rule 2 refuses. THE ACTOR DESIGN DOES NOT WORK, and this is not it. `land` making the call itself looks reachable: the tool is on the session's toolbox MCP server, an http endpoint under /v2/ccr-sessions/ carrying no bearer token, so a JSON-RPC tools/call looks like a plain POST. Measured 2026-08-18, that POST is answered 401 at both the toolbox and the github endpoint, with the injected config's own header values, and identically when forced through $HTTPS_PROXY: no_proxy carries anthropic.com, so requests to that host bypass the agent proxy and nothing injects a credential; the two headers are routing, not authorization. A first cut of this change shipped that POST behind a fail-open anyway. It removed zero subscriptions while its suite stayed green against a stubbed curl — a mechanism that reads as coverage and is not, which is CLOUD-418's defect rebuilt by hand. Filed as CLOUD-673. So this is `claim-check`'s inversion, the same one `issue-search-check` uses: the agent can do what the task cannot. The session's own tool call succeeds — that is how all three occurrences were remedied — so the agent unsubscribes, `pr-unsubscribed record <pr>` records that it happened for THIS pull request in THIS session from the tool's own answer, and `land` refuses to spend a runner until the record exists. The rule becomes an exit code without pretending to an effect nothing here can produce. Placement and posture: * The check is the FIRST thing `land` does, before the singleton and the lease, so a refusal costs no CI at all and the fix is one tool call away. * Off harness there is no injected config, therefore no session, therefore no subscription — `pr-unsubscribed` passes silently. That fail-open is what makes it safe on the critical path; a gate that cannot look must never become a gate that blocks everything. * Keyed by (session, PR), because a subscription belongs to that pair. A receipt from a previous container attests to nothing about this one, and #489's answer cannot satisfy #490 — the honest error this is built for, since the harness pins a session to one branch name for a whole engagement. * Pointer-only on stdout and in the receipt: the PR, the session and a digest of the answer. Never the answer, which is a message about a webhook stream. The honest limit, stated in the gate's own header: this proves the call was MADE for this PR, not that GitHub's subscription state is empty. Only the API answers that and reaching it is CLOUD-673. The claim receipt has the identical property, accepted there deliberately — the threat model is honest error, not fabrication. Ten rows in tests/pr-unsubscribed.bats cover both verbs: the refusal, the recorded drop, an answer naming the wrong PR, a receipt from another PR and from another session, empty stdin as could-not-look rather than a refusal, off-harness silence, pointer-only output, and bad arguments. Three rows in tests/land.bats cover the landing: the stop spends nothing (no ready, no push, no comment, no verify), the check names the PR being landed, and a passing gate leaves a lap unchanged. Four `#MUTANT` declarations, and the stopping-condition census moves 27 -> 28 — it caught the new stop the moment it was added, which is what it is for. Refs: CLOUD-518



Removes
plan-holdentirely — six tasks, three bats suites, three hook registrations, twoMUTANT_GATESrows, and the session-startspannedreport.Why
plan-hold-guardwas aPreToolUsedeny onExitPlanMode|AskUserQuestion. Every handoff to a human paid a refused tool call, a second call to arm a four-hour sleeper, and a retry. The premise: occupying the container defers the reclaim that destroys a human's typed approval.The premise was measured twice and failed twice. CLOUD-491 recorded a live hold on 2026-08-12 ~22:12 UTC — the container was restarted anyway, killing four tracked tasks including a
landmid-CI-wait — and reproduced it at ~23:45. CLOUD-500 conceded the point and gated its own stage 1 on oneplan-hold-check spanned=0reading, calling itself "Ready but not startable" until then.That reading has never happened. Measured in a working clone 2026-08-13:
No hold has completed a poll anywhere. In the ~17 hours since
d6ce33flanded the heartbeat, stage 0 produced not a negative reading but no reading at all — while the mechanism it was meant to grade went on charging a deny on every path to a person.CLOUD-451 has been In Review since 2026-08-12 13:15 with its first acceptance bullet never met, and has spawned three open follow-ups (CLOUD-485, CLOUD-511, CLOUD-500 stage 1). A cost that is certain against a benefit that has never been observed is not a gate.
What is deleted
mise-tasks/plan-hold{,-check,-guard,-release,-release-check,-release-tool}tests/plan-hold{,-check,-guard}.batsExitPlanMode|AskUserQuestionPreToolUseblock, theUserPromptSubmitrelease entry, the wholePostToolUsekeyMUTANT_GATESdropsplan-hold,plan-hold-checksession-start.sh'srecord-boot/spannedblockThe
spannedreport goes with the rest deliberately: it reads a heartbeat only the hold writes, so with no hold it can only ever answer "cannot look", andrecord-boot/batten-bootshas no other consumer.Two comments citing
plan-hold-check's half-written-sentinel distinction (mise-tasks/issue-read-guard,tests/issue-read-guard.bats) are reworded to citealive's corpse-versus-free-lock reading instead, rather than dangling a reference to a deleted file.What is not being said
That the problem is fake. It is real. CLOUD-451 is back in Todo, and an idle handoff turn still risks the reclaim. This removes one unvalidated remedy, not the finding. CLOUD-485's and CLOUD-511's findings were both correct too — they are defects in code whose benefit was never observed, which is why they go with it rather than get fixed again.
Incidentally measured in the session that produced this change: the arm →
ExitPlanMode→ release cycle worked end-to-end (CLOUD-485's fix doing its job). That says the plumbing works. It says nothing about whether occupancy defers a reclaim, which is the claim in question.Buying it back
One
plan-hold-check spanned=0reading, recorded on CLOUD-500 — a hold demonstrably live when a container was replaced. Not an argument and not a date. This commit is onegit revertfrom restoring the whole mechanism, CLOUD-485's and CLOUD-511's fixes included.Verification
git grep 'plan-hold'over the tree returns only the deliberate historical note in.claude/rules/toolchain.md.ExitPlanModeandAskUserQuestionpayloads through.claude/hooks/batten-hook.shexit 0 with no deny (they returnedpermissionDecision: "deny"before)..claude/settings.jsonparses and no longer carries aPostToolUsekey.mise run verifyandmise run mutant— running; results reported on this PR before it is readied.Coordination
PR #396 (CLOUD-511) is open against files this deletes; commented there suggesting it be parked rather than driven to green. CLOUD-500 and CLOUD-511 are In Progress under other sessions, so their states were left alone and flagged by comment rather than yanked.
Refs: CLOUD-515
Generated by Claude Code