Skip to content

revert(hooks): remove plan-hold, whose occupancy was never shown to work (CLOUD-515) - #397

Merged
wenzowski merged 2 commits into
mainfrom
claude/container-hold-plan-issue-kmcpe4
Aug 13, 2026
Merged

wenzowski merged 2 commits into
mainfrom
claude/container-hold-plan-issue-kmcpe4

Conversation

@wenzowski

Copy link
Copy Markdown
Contributor

Removes plan-hold entirely — six tasks, three bats suites, three hook registrations, two MUTANT_GATES rows, and the session-start spanned report.

Why

plan-hold-guard was a PreToolUse deny on ExitPlanMode|AskUserQuestion. Every handoff to a human paid a refused tool call, a second call to arm a four-hour sleeper, and a retry. The premise: occupying the container defers the reclaim that destroys a human's typed approval.

The premise was measured twice and failed twice. CLOUD-491 recorded a live hold on 2026-08-12 ~22:12 UTC — the container was restarted anyway, killing four tracked tasks including a land mid-CI-wait — and reproduced it at ~23:45. CLOUD-500 conceded the point and gated its own stage 1 on one plan-hold-check spanned = 0 reading, calling itself "Ready but not startable" until then.

That reading has never happened. Measured in a working clone 2026-08-13:

.git/batten-hold-heartbeat   → absent
.git/batten-boots            → one line (this boot)

No hold has completed a poll anywhere. In the ~17 hours since d6ce33f landed the heartbeat, stage 0 produced not a negative reading but no reading at all — while the mechanism it was meant to grade went on charging a deny on every path to a person.

CLOUD-451 has been In Review since 2026-08-12 13:15 with its first acceptance bullet never met, and has spawned three open follow-ups (CLOUD-485, CLOUD-511, CLOUD-500 stage 1). A cost that is certain against a benefit that has never been observed is not a gate.

What is deleted

surface files
tasks mise-tasks/plan-hold{,-check,-guard,-release,-release-check,-release-tool}
suites tests/plan-hold{,-check,-guard}.bats
hooks the ExitPlanMode|AskUserQuestion PreToolUse block, the UserPromptSubmit release entry, the whole PostToolUse key
config MUTANT_GATES drops plan-hold,plan-hold-check
sensor session-start.sh's record-boot/spanned block

The spanned report goes with the rest deliberately: it reads a heartbeat only the hold writes, so with no hold it can only ever answer "cannot look", and record-boot/batten-boots has no other consumer.

Two comments citing plan-hold-check's half-written-sentinel distinction (mise-tasks/issue-read-guard, tests/issue-read-guard.bats) are reworded to cite alive's corpse-versus-free-lock reading instead, rather than dangling a reference to a deleted file.

What is not being said

That the problem is fake. It is real. CLOUD-451 is back in Todo, and an idle handoff turn still risks the reclaim. This removes one unvalidated remedy, not the finding. CLOUD-485's and CLOUD-511's findings were both correct too — they are defects in code whose benefit was never observed, which is why they go with it rather than get fixed again.

Incidentally measured in the session that produced this change: the arm → ExitPlanMode → release cycle worked end-to-end (CLOUD-485's fix doing its job). That says the plumbing works. It says nothing about whether occupancy defers a reclaim, which is the claim in question.

Buying it back

One plan-hold-check spanned = 0 reading, recorded on CLOUD-500 — a hold demonstrably live when a container was replaced. Not an argument and not a date. This commit is one git revert from restoring the whole mechanism, CLOUD-485's and CLOUD-511's fixes included.

Verification

  • git grep 'plan-hold' over the tree returns only the deliberate historical note in .claude/rules/toolchain.md.
  • ExitPlanMode and AskUserQuestion payloads through .claude/hooks/batten-hook.sh exit 0 with no deny (they returned permissionDecision: "deny" before).
  • .claude/settings.json parses and no longer carries a PostToolUse key.
  • mise run verify and mise run mutant — running; results reported on this PR before it is readied.

Coordination

PR #396 (CLOUD-511) is open against files this deletes; commented there suggesting it be parked rather than driven to green. CLOUD-500 and CLOUD-511 are In Progress under other sessions, so their states were left alone and flagged by comment rather than yanked.

Refs: CLOUD-515


Generated by Claude Code

@linear-code

linear-code Bot commented Aug 13, 2026 •

Copy link
Copy Markdown
CLOUD-515 `plan-hold` gates every path to a human on occupancy that was never shown to work: remove it, and let a `spanned` reading buy it back

Why

plan-hold buys nothing measurable and charges a hard deny on the one path a session takes to reach a person. The premise — occupy the container with a backgrounded no-op and the reclaim is deferred — has been measured twice and failed twice, and the sensor built to grade it has never returned a single reading.

The cost is certain. plan-hold-guard is a PreToolUse deny on ExitPlanMode|AskUserQuestion. Every handoff now costs a refused tool call, a second tool call to arm a sleeper, and a retry. When it misbehaves the session does not fail — it sits, holding a background slot for up to four hours, doing a test -e and an fsync every five seconds, and is reclaimed anyway.

The benefit has never been observed. CLOUD-491 measured a live hold on 2026-08-12 ~22:12 UTC — the container restarted regardless, killing four tracked tasks including a land mid-CI-wait — and reproduced it at ~23:45. CLOUD-500 states the honest position: the mechanism's own author cannot say whether occupancy is even the right lever, and stage 1 is "Ready but not startable" until at least one spanned = 0 reading exists.

That reading does not exist. Measured in this clone, 2026-08-13:

$ ls .git/batten-hold-heartbeat        → absent
$ cat .git/batten-boots                → 1786599491   (one line: this boot)

No hold has ever completed a poll here. The sensor CLOUD-491 shipped to grade CLOUD-451's acceptance has zero data, and CLOUD-451 has been In Review since 2026-08-12 13:15 with its acceptance — "a plan left open for longer than the observed reclaim window survives, with the human's typed input intact" — never met.

And the mechanism keeps generating its own defects. CLOUD-485 (the guard and the release listened to different events, so the answer never released the hold), CLOUD-511 (two handoffs in one turn leave the second unguarded), and CLOUD-500's stage 1 are three open follow-ups on a mechanism with no demonstrated effect. The surface is 6 tasks, 3 bats suites, 3 hook registrations and 2 MUTANT_GATES rows — all of it maintained against an unvalidated premise.

This is the non-negotiable-2 shape read backwards: not a rule shipped without a mechanism, but a mechanism shipped without evidence its object exists, and gates decide rather than estimate.

Refinement — Ready

  • Source of truth (§1). .claude/settings.json is the authority for which hooks fire; mise-tasks/ is the authority for which tasks exist. Removal is complete when neither names plan-hold.
  • Mechanism as a computable predicate (§2). git grep -l 'plan-hold' -- .claude mise.toml mise-tasks tests returns nothing but incidental prose references, and mise run verify is green. Not a judgement — a grep and an exit code.
  • Effect (§3). Deletion only. No new task, no new gate, no new network reach.
  • Generated artifacts + drift gate (§4). MUTANT_GATES in mise.toml drops plan-hold,plan-hold-check; mise run mutant must stay green, and an entry naming a deleted gate is what would otherwise report unappliable-mutation.
  • Output & exit contract (§5). No contract changes — the removed guard's deny channel simply stops existing.
  • Commit / bump (§6). revert(hooks) — patch until 0.1.0.
  • Test obligation (§7). tests/plan-hold.bats, tests/plan-hold-check.bats and tests/plan-hold-guard.bats are deleted with the code they cover; no ratchet forbids it (checked — no bats-tests-not-deleted gate is wired). mise run test:bats and mise run mutant green after the deletion is the assertion.
  • Blockers (§8). None.

What buys it back

Not an argument and not a date. One plan-hold-check spanned reading of 0 — a hold demonstrably live when a container was replaced — recorded on CLOUD-500. The code is in git history at d6ce33f and its parent; restoring it is a revert, not a rewrite. Until that reading exists, the mechanism is a tax with an unproven benefit and it should not be on the hot path to a human.

Acceptance

  • ExitPlanMode and AskUserQuestion are ungated — no deny, no sleeper, no arming call.
  • No plan-hold* task, suite, hook registration or MUTANT_GATES entry remains.
  • mise run verify and mise run mutant are green.
  • CLOUD-451 is reopened rather than closed: the problem it names is real and unsolved, and this issue removes an attempt at it, not the finding.

Not in this issue

Whether the reclaim problem is real. It is — CLOUD-451's report from the repo owner stands, and destroying a human's typed approval is the defect. This removes one unvalidated remedy, and CLOUD-500 stage 0 is still the cheapest way to learn whether occupancy is the lever at all.

Review in Linear

claude added 2 commits August 13, 2026 06:30
…ork (CLOUD-515)

`plan-hold-guard` was a PreToolUse deny on `ExitPlanMode|AskUserQuestion`: every
handoff to a human paid a refused tool call, a second call to arm a four-hour
sleeper, and a retry. The premise was that occupying the container defers the
reclaim that destroys a human's typed approval.

That premise was measured twice and failed twice. CLOUD-491 recorded a live hold
on 2026-08-12 ~22:12 UTC and the container was restarted anyway, killing four
tracked tasks including a `land` mid-CI-wait; it reproduced at ~23:45. CLOUD-500
conceded the point and declared its own stage 1 "Ready but not startable" until
one `plan-hold-check spanned` = 0 reading exists.

No such reading exists. Measured in this clone 2026-08-13: `batten-hold-heartbeat`
absent, `batten-boots` carrying one line — no hold has ever completed a poll here.
CLOUD-451 has been In Review since 2026-08-12 13:15 with its acceptance never met,
and has spawned three open follow-ups (CLOUD-485, CLOUD-511, CLOUD-500's stage 1)
on a mechanism with no demonstrated effect. A cost that is certain against a
benefit that has never been observed is not a gate; gates decide.

Removed: six `mise-tasks/plan-hold*`, three bats suites, three hook registrations
(the PreToolUse deny, the UserPromptSubmit release, the whole PostToolUse block),
two MUTANT_GATES rows, and the session-start `spanned` report — which reads a
heartbeat only the hold writes, so with no hold it can only answer "cannot look".

The problem CLOUD-451 names is real and unsolved; this removes an attempt at it,
not the finding. The code is one `git revert` from returning, and its price is one
`spanned` = 0 reading recorded on CLOUD-500.

Refs: CLOUD-515
…LOUD-515)

Removing plan-hold deleted its three bats suites, dropping tests/**/*.bats from
1600 to 1552 and denying at bats-tests-not-deleted. The rule's own no_fix_reason
names the two exits — "restore the tests, or waive the reduction deliberately" —
and restoring is not one: the suites cover code that no longer exists.

Unnarrowed, and not for want of trying. A ratchet finding's path is the glob plus
the counts (rules.rs builds "{glob} {base}->{working}"), because the finding is
about the whole matched set rather than any one file. There is no per-file path a
`path =` could select, and a waiver written against the literal would embed
1600->1552 and lapse the moment either count moved.

It is expected to lapse unused: base is origin/main, so once this lands the floor
becomes the new count on its own.

Refs: CLOUD-515
@wenzowski
wenzowski force-pushed the claude/container-hold-plan-issue-kmcpe4 branch from e6e83a1 to 15b8820 Compare August 13, 2026 06:31
@sonarqubecloud

Copy link
Copy Markdown

@wenzowski
wenzowski marked this pull request as ready for review August 13, 2026 06:33
@wenzowski

Copy link
Copy Markdown
Contributor Author

/fast-forward

@wenzowski
wenzowski merged commit 15b8820 into main Aug 13, 2026
15 checks passed
@wenzowski
wenzowski deleted the claude/container-hold-plan-issue-kmcpe4 branch August 13, 2026 06:44
wenzowski added a commit that referenced this pull request Aug 18, 2026
AGENTS.md bans PR-webhook babysitting twice over — this task drives the
landing loop "no timeout, no cap, never the PR webhook", and no heartbeat may
babysit a PR — and the harness armed a subscription on every PR this repo
opened anyway: #397, #402 and #489, three occurrences in five days. On #402
and #489 with no `subscribe_pr_activity` call behind it, so a `permissions.deny`
row on the tool would close only the path nobody used. The remedy until now
was an agent remembering: prose, therefore feedforward only, therefore exactly
the half-change non-negotiable rule 2 refuses.

`land` now drops it itself — once before the lap, and again after each ready it
fires, since a `pull_request` event is what arms one. `unsubscribe_pr_activity`
lives on the session's toolbox MCP server, an http endpoint under
`/v2/ccr-sessions/` carrying no bearer token, so a JSON-RPC `tools/call` is a
plain POST. Both header values are per-session and account-specific, so nothing
here commits them: only the public endpoint SHAPE is tracked, and the volatile
halves are read at run time out of the injected client config — the entry chosen
by the tool it DECLARES rather than by a server id, which is CLOUD-191's
endpoint-anchoring pattern applied to a second consumer. The repository is
derived from the remote rather than declared, so no slug is a second authority.

Failing open is the posture, not a caveat on it. Off harness there is no config,
no subscription and nothing to drop, and the silence is the right answer; on
harness, a refused or unreachable call costs the lap nothing, because ending a
green landing over a nuisance subscription would be a worse defect than the one
being closed. The status is dropped at the call site on purpose. The one thing
never silent is a call that did not do what it says.

Measured while building this, and it qualifies the mechanism rather than the
design: a POST from a task is answered 401 by that endpoint today, so the drop
currently reports `could NOT drop #N's webhook subscription (http 401)` instead
of firing. Recorded on the issue with the reproduction. The shape is still the
one the issue specifies, it costs a landing nothing when refused, and it says
so rather than reading as done — which is the difference between this and the
deny rule the issue already rejected as a placebo.

Ten rows in tests/land.bats cover the test obligation: a subscribed PR is
dropped and says so, an absent config and a toolbox without the verb are silent
no-ops, and a 500, an unreachable endpoint, a JSON-RPC error and an `isError`
result each cost the lap nothing while still being reported. Two `#MUTANT`
declarations hold them: dropping the call reddens the first row, and reading the
call's status at the call site reddens the fail-open row.

Refs: CLOUD-518
wenzowski added a commit that referenced this pull request Aug 18, 2026
AGENTS.md bans PR-webhook babysitting twice over — this loop runs on "no timeout,
no cap, never the PR webhook", and no heartbeat may babysit a PR — and the harness
arms a subscription on every PR this repo opens anyway: #397, #402 and #489, three
occurrences in five days, two of them with no `subscribe_pr_activity` call behind
them. So a `permissions.deny` row on the tool closes only the path nobody used,
and the remedy until now was an agent remembering, which is prose and therefore
feedforward only — the half-change non-negotiable rule 2 refuses.

THE ACTOR DESIGN DOES NOT WORK, and this is not it. `land` making the call itself
looks reachable: the tool is on the session's toolbox MCP server, an http endpoint
under /v2/ccr-sessions/ carrying no bearer token, so a JSON-RPC tools/call looks
like a plain POST. Measured 2026-08-18, that POST is answered 401 at both the
toolbox and the github endpoint, with the injected config's own header values, and
identically when forced through $HTTPS_PROXY: no_proxy carries anthropic.com, so
requests to that host bypass the agent proxy and nothing injects a credential; the
two headers are routing, not authorization. A first cut of this change shipped
that POST behind a fail-open anyway. It removed zero subscriptions while its suite
stayed green against a stubbed curl — a mechanism that reads as coverage and is
not, which is CLOUD-418's defect rebuilt by hand. Filed as CLOUD-673.

So this is `claim-check`'s inversion, the same one `issue-search-check` uses: the
agent can do what the task cannot. The session's own tool call succeeds — that is
how all three occurrences were remedied — so the agent unsubscribes,
`pr-unsubscribed record <pr>` records that it happened for THIS pull request in
THIS session from the tool's own answer, and `land` refuses to spend a runner
until the record exists. The rule becomes an exit code without pretending to an
effect nothing here can produce.

Placement and posture:

* The check is the FIRST thing `land` does, before the singleton and the lease, so
  a refusal costs no CI at all and the fix is one tool call away.
* Off harness there is no injected config, therefore no session, therefore no
  subscription — `pr-unsubscribed` passes silently. That fail-open is what makes
  it safe on the critical path; a gate that cannot look must never become a gate
  that blocks everything.
* Keyed by (session, PR), because a subscription belongs to that pair. A receipt
  from a previous container attests to nothing about this one, and #489's answer
  cannot satisfy #490 — the honest error this is built for, since the harness pins
  a session to one branch name for a whole engagement.
* Pointer-only on stdout and in the receipt: the PR, the session and a digest of
  the answer. Never the answer, which is a message about a webhook stream.

The honest limit, stated in the gate's own header: this proves the call was MADE
for this PR, not that GitHub's subscription state is empty. Only the API answers
that and reaching it is CLOUD-673. The claim receipt has the identical property,
accepted there deliberately — the threat model is honest error, not fabrication.

Ten rows in tests/pr-unsubscribed.bats cover both verbs: the refusal, the recorded
drop, an answer naming the wrong PR, a receipt from another PR and from another
session, empty stdin as could-not-look rather than a refusal, off-harness silence,
pointer-only output, and bad arguments. Three rows in tests/land.bats cover the
landing: the stop spends nothing (no ready, no push, no comment, no verify), the
check names the PR being landed, and a passing gate leaves a lap unchanged. Four
`#MUTANT` declarations, and the stopping-condition census moves 27 -> 28 — it
caught the new stop the moment it was added, which is what it is for.

Refs: CLOUD-518
wenzowski added a commit that referenced this pull request Aug 18, 2026
AGENTS.md bans PR-webhook babysitting twice over — this loop runs on "no timeout,
no cap, never the PR webhook", and no heartbeat may babysit a PR — and the harness
arms a subscription on every PR this repo opens anyway: #397, #402 and #489, three
occurrences in five days, two of them with no `subscribe_pr_activity` call behind
them. So a `permissions.deny` row on the tool closes only the path nobody used,
and the remedy until now was an agent remembering, which is prose and therefore
feedforward only — the half-change non-negotiable rule 2 refuses.

THE ACTOR DESIGN DOES NOT WORK, and this is not it. `land` making the call itself
looks reachable: the tool is on the session's toolbox MCP server, an http endpoint
under /v2/ccr-sessions/ carrying no bearer token, so a JSON-RPC tools/call looks
like a plain POST. Measured 2026-08-18, that POST is answered 401 at both the
toolbox and the github endpoint, with the injected config's own header values, and
identically when forced through $HTTPS_PROXY: no_proxy carries anthropic.com, so
requests to that host bypass the agent proxy and nothing injects a credential; the
two headers are routing, not authorization. A first cut of this change shipped
that POST behind a fail-open anyway. It removed zero subscriptions while its suite
stayed green against a stubbed curl — a mechanism that reads as coverage and is
not, which is CLOUD-418's defect rebuilt by hand. Filed as CLOUD-673.

So this is `claim-check`'s inversion, the same one `issue-search-check` uses: the
agent can do what the task cannot. The session's own tool call succeeds — that is
how all three occurrences were remedied — so the agent unsubscribes,
`pr-unsubscribed record <pr>` records that it happened for THIS pull request in
THIS session from the tool's own answer, and `land` refuses to spend a runner
until the record exists. The rule becomes an exit code without pretending to an
effect nothing here can produce.

Placement and posture:

* The check is the FIRST thing `land` does, before the singleton and the lease, so
  a refusal costs no CI at all and the fix is one tool call away.
* Off harness there is no injected config, therefore no session, therefore no
  subscription — `pr-unsubscribed` passes silently. That fail-open is what makes
  it safe on the critical path; a gate that cannot look must never become a gate
  that blocks everything.
* Keyed by (session, PR), because a subscription belongs to that pair. A receipt
  from a previous container attests to nothing about this one, and #489's answer
  cannot satisfy #490 — the honest error this is built for, since the harness pins
  a session to one branch name for a whole engagement.
* Pointer-only on stdout and in the receipt: the PR, the session and a digest of
  the answer. Never the answer, which is a message about a webhook stream.

The honest limit, stated in the gate's own header: this proves the call was MADE
for this PR, not that GitHub's subscription state is empty. Only the API answers
that and reaching it is CLOUD-673. The claim receipt has the identical property,
accepted there deliberately — the threat model is honest error, not fabrication.

Ten rows in tests/pr-unsubscribed.bats cover both verbs: the refusal, the recorded
drop, an answer naming the wrong PR, a receipt from another PR and from another
session, empty stdin as could-not-look rather than a refusal, off-harness silence,
pointer-only output, and bad arguments. Three rows in tests/land.bats cover the
landing: the stop spends nothing (no ready, no push, no comment, no verify), the
check names the PR being landed, and a passing gate leaves a lap unchanged. Four
`#MUTANT` declarations, and the stopping-condition census moves 27 -> 28 — it
caught the new stop the moment it was added, which is what it is for.

Refs: CLOUD-518
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants