Skip to content

fix(pkl-check): a gate that could not fail, over a package CI refetched every run - #382

Merged
wenzowski merged 3 commits into
mainfrom
claude/ready-issues-handoff-qnra4q
Aug 13, 2026
Merged

wenzowski merged 3 commits into
mainfrom
claude/ready-issues-handoff-qnra4q

Conversation

@wenzowski

Copy link
Copy Markdown
Contributor

CLOUD-406's prevention half, plus a defect found while writing its test.

The package was fetched on every CI run

hk.pkl:14 amends a pkl package uri, which pkl resolves over the network on every evaluation — a second dependency beside the hk binary mise.lock pins, and unpinned at runtime. It works locally because container provisioning warms pkl's default cache at ~/.pkl/cache; nothing warmed it in CI.

Measured on run 31632519615: the fetch of hk@1.54.0.zip failed, hk aborted with exit 134 before reaching a single gate, and the failure grade then wedged the SHA — land re-fires a ready only on a head carrying no graded run. An upstream blip on a release CDN became a red required check on a branch that changed nothing.

One actions/cache step on ~/.pkl/cache, keyed on hk.pkl itself so a version bump changes the key by construction. The ci job only, confirmed rather than assumed:

job reaches hk? why
ci yes hooks (hk check --all) is a dependency of mise run ci and of nothing else
cross, darwin-link no they reach doctor, whose hook probe is skipped under CI
msrv, semver no neither runs hk

A miss behaves as today and then populates, so the exposure narrows to one run per bump or eviction rather than being removed. It cannot be removed here: mise-action's own bootstrap fetch precedes every seam this repository controls, which is why CLOUD-406 carries a recovery half as well.

The gate over it could not fail

pkl-check had no suite, and writing one showed why that mattered. Every path ended in an unconditional exit 0 or in an echo whose status became the task's, with no set -e to stop a failing evaluation reaching them. It passed on a malformed hk.pkl and on an unreachable package alike — while hk.pkl:381 routed its pkl step through it. It now reports pkl's status, from one invocation site so a third exit path cannot grow back.

Test obligation

tests/pkl-check.bats measures both directions over the same command:

cache egress expected
cold denied fails
warm denied exits 0

The first row is load-bearing — without it the second passes on any machine that has ever evaluated hk.pkl (CLOUD-418). Both rows were red before the exit-status fix.

Egress is denied with pkl's own --http-proxy, not with HTTPS_PROXY as CLOUD-406's Ready block specifies. pkl is a native image and ignores the environment variables: the first probe went out over the real network and failed on an unrelated truststore error, which would read as a passing anti-vacuity row while measuring nothing about egress. The flag fails closed in the right direction — if it were ever ignored the fetch would succeed and the cold row would go green. The suite never downloads, so it needs no network of its own.

Not here

The recovery half — one bounded re-fire in land — touches mise-tasks/land and tests/land.bats, held by another session under a standing takeover.

Refs: CLOUD-406

@linear-code

linear-code Bot commented Aug 12, 2026 •

Copy link
Copy Markdown
CLOUD-406 A required job reaches the network before it reaches a gate, so an upstream outage reds a branch that changed nothing — and the `failure` grade then wedges the SHA

Why

Measured on PR #302 (CLOUD-363), 2026-08-11, run 31538830648 on 60c5d79. Two
bats cases failed in CI and pass locally on the identical tree:

not ok 736 the hook is synchronous — async would restore the race it closes
not ok 738 doctor runs inside the synchronous window — after install, before the preflight

The assertion that failed is [ "$status" -eq 0 ] on the hook invocation, and
the reason is in the captured output:

::error:: session-start: mise-install failed — see /tmp/session-start-mise-install.log
mise ERROR Failed to install aqua:cli/cli@2.97: GitHub artifact attestations
  verification failed: Verification failed: Workflow verification failed:
  expected 'cli/cli/.github/workflows/deployment.yml', found certificate identity:
  "https://dotcom.releases.github.com"; Sigstore error: TUF error:
  TUF repository load failed: transport error: timestamp.json not found
::error:: session-start: setup incomplete — expect missing tools or MCP servers

Nothing in #302's diff touches the hook, the toolchain, or mise.toml. The same
suite ran green locally on the same commit (bats tests/session-start.bats,
exit 0), and green in every one of the 13 verify runs land made across three
invocations. The tree is not what changed; Sigstore's TUF metadata endpoint was.

Two more instances, one layer further out again. Measured 2026-08-12 on PR #325 (CLOUD-367), run 31632519615 on a4dec2f9. Both died in job setup, before any repository code ran at all:

ci:           Eval error: HTTP fetch failed for
              https://github.com/jdx/hk/releases/download/v1.54.0/hk@1.54.0.zip
              thread 'main' panicked at src/settings.rs:101:30:
              Failed to load configuration: Failed to read config file: .../hk.pkl
              Aborted (core dumped)                                    -> exit 134

darwin-link:  curl -fsSL .../mise-v2026.8.5-linux-x64.tar.zst
              curl: (56) Connection died, tried 5 times before giving up -> exit 1

final then went red on its needs, correctly. Evaluating hk.pkl resolves the hk pkl package over the network on every run — mise.lock pins the hk binary, but the pkl package it imports is a second dependency resolved per evaluation, and unpinned at runtime.

Three independent instances now, across two layers — inside a bats case, and in job setup before any repository code — is what makes this a class rather than weather.

And the failure does not just cost a matrix: it wedges the branch. This is the part the original body did not carry, and it is why this is no longer No priority. A failure is a graded run, and land re-fires a ready only on a head carrying no graded run. So once the blip grades the SHA, that SHA can never buy another run through the loop. Measured on #325: the next land won the lease, pushed nothing (HEAD unchanged), skipped the ready re-fire because the head is graded, re-read the same stale red, and stopped with

::error:: land: CI is red on a4dec2f9. A red run on a verified branch means verify
          and CI disagree — fix the mismatch locally, then run land again.

which is the correct message for a real disagreement and exactly wrong here: verify and CI do not disagree, because CI never evaluated anything. Recovery took rerun_failed_jobs against the run id — a manual step outside the loop land exists to drive, exactly as CLOUD-363 records for its own non-answer.

This is the lock-check shape, one layer out. .claude/rules/toolchain.md
already records the principle from CLOUD's earlier round on it: a gate whose*
*verdict comes from a remote API is not testing the commit — it answers "did
upstream change", and fails a branch for drift it did not cause. lock-check
was split into lock-complete (committed bytes, no network) and a scheduled
currency job for exactly this reason. These two cases re-introduce the same
coupling by the back door: they execute the hook end-to-end, and the hook's
first real act is a network install with attestation verification.

The cost is not theoretical. It converts an unrelated third-party outage into a
red required check, which land correctly reads as red — so it re-drafts the
PR and stops, and the branch pays a full lap (rebase + verify + a CI run) to
learn the same thing again once upstream recovers.

What the tests are actually about

Neither case's stated property needs the install to succeed:

  • 736 asserts the hook does not emit {"async": true} — a property of what it
    prints and of its source text.
  • 738 asserts doctor runs after mise install and before the preflight — a
    property of ORDER, and the suite already stubs mise to short-circuit
    run container-preflight and run doctor while exec-ing the real binary
    for everything else.

So the ordering fixture exists; it simply lets mise install through to the
network. The exit-status assertion is what couples them to upstream
availability, and it is the assertion neither case is really making.

The per-instance fixes

For the session-start.bats instance the choice belongs to whoever owns that file:

  1. Extend the existing mise stub to intercept bare install the way it
    already intercepts run doctor, so the ordering is asserted from $CALLS
    with no network at all. Cheapest, and matches the fixture's own design.
  2. Keep one case that genuinely exercises the install, and mark it so a
    provisioning failure is distinguishable from a hook defect — the split
    lock-check took.

Either way the property to preserve is the one CLOUD-218 bought: doctor runs
inside the synchronous window, after mise install and before the preflight.

Blockers: none.

Acceptance

  • A required check is never red because an upstream host was unreachable. The branch's verdict depends on the committed bytes, not on who is up.
  • tests/session-start.bats passes with no network reachable, while a hook that stops calling mise install, or calls doctor outside the window, still fails it — the fix must not be "delete the assertion".
  • hk.pkl evaluates without a live fetch of its own pkl package, on the same footing as every other pinned tool; a version bump stays a deliberate, committed change.
  • Whatever remains network-dependent fails in a way a caller can tell from a verdict — today the abort is exit 134, which is not in the 0/1/2/3 contract at all.

Files: tests/session-start.bats and mise-tasks/session-start (only if the fix needs a seam) for the first instance; whatever pins and caches the pkl package for the second — mise.toml, mise.lock, and the workflow's setup steps.

Commit type / bump: fix(tests) for the session-start half, ci for the toolchain-caching half. Neither touches crates/, so release-plz cuts nothing.


Refinement — Ready (a required job's verdict stops depending on who is up)

Refinement gate: Definition of Ready & Done. This body carries only specializations.

  • Source of truth (§1). The rule already exists and is not re-typed here: .claude/rules/toolchain.md records that a gate whose verdict comes from a remote API is not testing the commit, which is why lock-check was split into lock-complete (committed bytes, no network) and a scheduled currency job. This issue extends that settled rule from tasks to required-job setup, and each per-instance fix lives with its own files rather than in a new shared authority.

  • Computable predicate (§2). Two, one per half, each a command and an exit code over an object it decides.

    Prevention — hk.pkl evaluates with no network. hk.pkl:14 amends a pkl package URI (package://github.com/jdx/hk/releases/download/v1.54.0/hk@1.54.0#/Config.pkl, and line 16 imports from the same), which pkl resolves on every evaluation — a second dependency beside the hk binary mise.lock already pins. The predicate: with egress denied, pkl eval hk.pkl exits 0. Deny it by pointing HTTPS_PROXY/HTTP_PROXY at a closed local port rather than by dropping capabilities — portable, unprivileged, and it fails closed if the proxy vars are ever ignored, because the fetch then succeeds and the anti-vacuity case in §7 catches it. Home: mise-tasks/pkl-check, extended. It already owns "does hk.pkl evaluate", is already wired into the gate (hk.pkl:385), and its own header already names this coupling — "pkl resolves amends 'package://…' over the network on a cold cache". No new authority. The cache key derives from the hk version, whose agreement with the amends URL hk-version already gates, so nothing new has to be kept honest.

    Recovery — one bounded re-fire. Distinguishing a pre-gate abort from a real verdict is not available and must not be built: both grade failure, the failing step is the same step (mise run ci, inside which hk panics), and the only remaining discriminator is log text — prose matching, which this repo refuses. The predicate needs no such distinction. land, holding the lease, on a head where all three hold — (a) a green verify receipt exists for this exact HEAD, (b) the required set is red, (c) no re-fire receipt exists for this HEAD — re-fires the ready once and records a SHA-keyed re-fire receipt under .git/batten-receipts/, the same shape ready-guard's receipts already use. A second land on the same head stops exactly as today. Pure filesystem state plus checks-green's existing verdict: no log read, no conclusion heuristic, no new event, and compatible with CLOUD-420 because it happens inside the lease land already holds.

    Why the re-fire is not drive-to-green. Re-running land on an unchanged red head is a guaranteed dead end today — the push moves nothing, the ready is skipped because the head is graded, the same stale red is re-read. This buys one more opinion, bounded at one, only when someone deliberately re-ran land. Condition (a) is what keeps it honest: it fires only where local proof and CI actually disagree, which is precisely the case land's current error message describes and then offers no move for.

  • Effect (§3). No command surface and no effect-table row. Test fixtures and CI configuration only.

  • Output & exit (§5). Unchanged for the passing case. Note for whoever implements: the abort arrives as exit 134, outside the 0/1/2/3 table, so a caller cannot distinguish "could not look" from a verdict by exit code today. Any reporting stays pointer-only — the package coordinate and the failing host, never fetched bytes.

  • Commit / bump (§6). fix → patch until 0.1.0. The prevention half alone is ci and would release nothing; the recovery and session-start halves are fix, and below 0.1.0 every release-worthy type collapses to a patch. Nothing under crates/ changes either way, so release-plz cuts a release only if a crate change rides along.

  • Test obligation (§7). Each case asserts an exit code, never prose.

    Prevention. (a) With egress denied and the cache warm, pkl eval hk.pkl exits 0. (b) With the pkl module cache purged and egress denied, the same command fails. (b) is the load-bearing case: without it the test passes on any warm machine forever and proves nothing, which is exactly the "a new gate is never shown to fail" defect CLOUD-418 landed a discipline against. Both directions or neither.

    Recovery, in tests/land.bats. (c) A red head carrying a green verify receipt and no re-fire receipt re-fires the ready once and writes the receipt. (d) The same head on a second land stops, as today — the bound, and the case a later "simplification" would quietly remove. (e) A red head with no verify receipt never re-fires, so the re-fire cannot become an unconditional retry. The existing lap-count assertions must still hold.

    Session-start instance. The suite passes with egress unavailable, and the anti-vacuity case holds: a hook that stops calling mise install, or calls doctor outside the synchronous window, still fails. The property CLOUD-218 bought must survive the fix.

  • Blockers (§8). None. relatedTo CLOUD-363 — the same shape, a non-answer graded as terminal with no exit through the loop; its cancelled half has landed in checks-green, and this is the failure-from-an-abort half it deliberately does not cover. Also CLOUD-481 (folded in here), CLOUD-476, CLOUD-218, and CLOUD-369/CLOUD-399 where question 3 overlaps.

Two halves, landable separately

  1. Prevention — mise-tasks/pkl-check, mise.toml, and the workflow's cache step. Unowned, independently landable, and it removes the common case. Do this one first.
  2. Recovery — mise-tasks/land and tests/land.bats. These files are held by another session (claude/landing-process-conflicts-3ogq6s under a standing takeover of the landing loop). Coordinate before starting; do not open a competing branch against them.

On the three questions this body used to carry, so nobody re-derives them. What is the predicate is answered in §2. Can prevention ever be total — no, and the fourth instance above is why: mise-action's own bootstrap fetch precedes every seam this repository controls, so no in-repo gate, egress-blocked suite, or pkl cache can reach it. That is what makes the recovery half required rather than conditional. Is a sanctioned re-fire wanted — yes, in the bounded form §2 specifies; the two questions had converged, and one design closes both.

Filed from the CLOUD-363 session, which does not own these files.


Another instance, 2026-08-12, and it shows the wedge half as well as the red half.

PR #374, head 4649085. The msrv job died before reaching any gate:

[command]/usr/bin/curl -fsSL https://github.com/jdx/mise/releases/download/v2026.8.5/mise-v2026.8.5-linux-x64.tar.zst
curl: (22) The requested URL returned error: 503
##[error]The process '/usr/bin/curl' failed with exit code 22

That is mise-action fetching mise itself from GitHub's release CDN. Nothing in the diff was reached, nothing was judged, and msrv graded failure. final then graded failure because msrv is in its needs:, so a 503 on a tarball download produced a red on the one check-run branch protection requires.

The compounding part. The failure grade stuck to the SHA, exactly as this issue says. But the branch did not merely go red — it went red invisibly: checks-green was simultaneously polling on a zizmor that had graded skipped, so ci-wait reported "not an answer yet" over a set that already contained two failures (recorded on CLOUD-436). land therefore never reached its red-stop, and held the fleet landing lease for ~1 hour with another branch admitted behind it.

So an upstream 503 on a tool download cost: one wedged SHA, one starved sibling, and an hour of held lease. The two defects are independent and each is survivable alone; together they turn a transient upstream outage into a fleet stall.

Worth noting for the fix's shape: the failure is in mise-action's own bootstrap, before any repository code runs, so no in-repo gate ordering can catch it. The remedies that reach it are retry/caching at the action step or treating a pre-gate infrastructure failure as a distinct conclusion from a graded red — which is the same distinction CLOUD-470 is drawing on the agent-facing side.

Review in Linear

@wenzowski
wenzowski marked this pull request as ready for review August 12, 2026 23:41
claude added 3 commits August 13, 2026 00:18
…ed (CLOUD-491)

The hold occupies the container while a human reads a plan and left no record of
having done so, so a replaced container could not be told from an exited hold or
from one never armed. All three look identical from inside, which is why
CLOUD-451's acceptance could be asserted but never checked.

The obvious sensor does not work here, and it was measured rather than reasoned
about. Every write that survived the 2026-08-12 23:45 replacement stops at
23:42:15 - four MCP servers' logs, Serena's log, target/.rustc_info.json,
batten-holds - while the container kept serving tool calls until ~23:45. The last
182s of writes did not survive. So `epoch >= btime - poll` reads three minutes
stale exactly when it is asked, and a tolerance wide enough to cover that also
swallows a hold that legitimately exited minutes earlier.

The record is structural instead: `h` per poll, `x` only where the hold chose to
stop. A hold killed by the container going down never reaches an `x`, so an `h`
in final position is the evidence and no clock is involved. The `x` writes are
deliberately on the two terminal paths and never in the exit trap, which runs on
the kill too and would report every reclaimed hold as one that stopped on purpose.

Recording each boot separately is what tells a preserved disk from a fresh one:
the residue probe that would have done it does not work, since a batten-* dir's
mtime updates on every add and remove and reads as post-boot on a disk that
demonstrably survived.

plan-hold-check gains its own suite, following the convention claim-check,
land-lock-check and stop-posture-check keep. That is also what makes the
predicate mutation-checkable at all: mutant derives its suite as tests/$gate.bats.
Both gates join MUTANT_GATES; the three declared mutations are caught.
…ed every run

CLOUD-406's prevention half, plus a defect found while writing its test.

THE PACKAGE WAS FETCHED ON EVERY CI RUN. `hk.pkl:14` amends a pkl PACKAGE uri,
which pkl resolves over the network on every evaluation — a second dependency
beside the `hk` binary `mise.lock` pins, and unpinned at runtime. It works
locally because container provisioning warms pkl's default cache at
`~/.pkl/cache`; nothing warmed it in CI. Measured on run 31632519615: the fetch
of `hk@1.54.0.zip` failed, hk aborted with exit 134 before reaching a single
gate, and the `failure` grade then wedged the SHA, because `land` re-fires a
ready only on a head carrying no graded run. An upstream blip on a release CDN
became a red required check on a branch that changed nothing.

One `actions/cache` step on `~/.pkl/cache`, keyed on `hk.pkl` itself so a version
bump changes the key by construction. The `ci` job only, confirmed rather than
assumed: `hooks` (`hk check --all`) is a dependency of `mise run ci` and of
nothing else, `cross-check` and `darwin-link` reach `doctor` whose hook probe is
skipped under CI, and `msrv`/`semver` run neither. A miss behaves as today and
then populates, so the exposure narrows to one run per bump or eviction rather
than being removed — it cannot be removed here, because `mise-action`'s own
bootstrap fetch precedes every seam this repository controls.

THE GATE OVER IT COULD NOT FAIL. `pkl-check` had no suite, and writing one showed
why that mattered: every path ended in an unconditional `exit 0` or in an `echo`
whose status became the task's, with no `set -e` to stop a failing evaluation
reaching them. It passed on a malformed `hk.pkl` and on an unreachable package
alike, while hk.pkl:381 routed its `pkl` step through it — the presence of a gate
is what stops anyone looking (CLOUD-418). It now reports pkl's status, from one
invocation site so a third exit path cannot grow back.

`tests/pkl-check.bats` measures both directions over the same command: cold cache
plus egress denied fails, warm cache plus egress denied exits 0. The first row is
load-bearing — without it the second passes on any machine that has ever
evaluated `hk.pkl`. Both rows were red before the exit-status fix.

Egress is denied with pkl's own `--http-proxy`, NOT with `HTTPS_PROXY` as
CLOUD-406's Ready block specifies. pkl is a native image and ignores the
environment variables: the first probe went out over the real network and failed
on an unrelated truststore error, which reads as a passing anti-vacuity row while
measuring nothing about egress. The flag fails closed in the right direction — if
it were ever ignored the fetch would succeed and the cold row would go green. The
suite never downloads, so it needs no network of its own; the warm row copies the
ambient cache and skips with a diagnostic when there is nothing to copy.

The recovery half — one bounded re-fire in `land` — is deliberately not here.
Those files are held by another session under a standing takeover.

Refs: CLOUD-406
…r checks can be superseded

A required check whose workflow omits `types:` defaults to
`[opened, synchronize, reopened]`. Both jobs are draft-gated, so a PR created as
a draft mints a `skipped` run on `opened` — and then no event remains that could
replace it. `land` readies before pushing, and its push moves nothing whenever
HEAD is already on the remote, so `ready_for_review` is the only event left and
neither workflow was listening. `checks-green` correctly refuses to read a skip
as an answer, so `ci-wait` polls unbounded while `land` holds the fleet lease.

`ci.yml` and `commit-lint.yml` already declare the full set, which is why this
was invisible: on any head where a lap actually pushed, `synchronize` covered it.

MEASURED AS A DEADLOCK, NOT A WEDGED BRANCH. #382 and #385 were simultaneously
in the same state — every required check SUCCESS, `zizmor` SKIPPED with exactly
one run on the SHA — and `main` frozen at 795847b. That freeze is what made it
self-sustaining: with nothing to rebase onto, each lap's push was
`Everything up-to-date`, so the one event zizmor does subscribe to never fired.
Each PR needed the other to move first. Nothing inside the loop breaks that; a
commit does, which is why this fix and the unwedge are the same act.

Isolated on head 1789dfb: `ci.yml` carries a `skipped` run at 23:40:18 (the
draft `opened` event) AND a `success` at 23:41:08. `zizmor.yml` carries only the
skip. No push separates the two — the ready is the only event between them.

It costs no runner. Both jobs keep their
`if: github.event.pull_request.draft == false` guard, so every draft event still
skips; the only new firing is the single transition that needs one. This is not
a trade against the draft economy — it is the event that economy was missing.

`test.yml`'s header claimed a safety property that is false in this case:
"`ci-wait` reads an absent path-filtered check as absent rather than pending, so
a required check that produces no run here does not hang a landing." True for
ABSENT; a draft `opened` event produces PRESENT-AND-SKIPPED, which does hang it.
Corrected rather than left documenting the bug as a guarantee.

THE GATE FOR THIS PROPERTY IS NOT HERE, DELIBERATELY. CLOUD-503 specifies it in
`ci-local-parity`, on the seam at :282-287 that already couples
`CI_REQUIRED_CHECKS` to job names in both directions — the third arm asks whether
a required check can be REFRESHED, where the first two ask whether it can exist.
PR #385 is readied and rewrites that task and its suite (+1323), and this change
now lands first, so carrying the gate here would hand another session a merge
conflict in a readied PR. CLOUD-503 stays In Progress with that half specified
and its CLOUD-418 both-directions obligation written; neither workflow file here
is touched by #385, so this half conflicts with nothing.

Refs: CLOUD-503
@wenzowski
wenzowski force-pushed the claude/ready-issues-handoff-qnra4q branch from 1789dfb to 49b89dd Compare August 13, 2026 01:30
@sonarqubecloud

Copy link
Copy Markdown

@wenzowski

Copy link
Copy Markdown
Contributor Author

/fast-forward

@wenzowski
wenzowski merged commit 49b89dd into main Aug 13, 2026
10 checks passed
@wenzowski
wenzowski deleted the claude/ready-issues-handoff-qnra4q branch August 13, 2026 01:37
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants