Repository navigation
fix(pkl-check): a gate that could not fail, over a package CI refetched every run - #382
Conversation
CLOUD-406 A required job reaches the network before it reaches a gate, so an upstream outage reds a branch that changed nothing — and the `failure` grade then wedges the SHA
Why Measured on PR #302 (CLOUD-363), 2026-08-11, run 31538830648 on The assertion that failed is Nothing in #302's diff touches the hook, the toolchain, or Two more instances, one layer further out again. Measured 2026-08-12 on PR #325 (CLOUD-367), run 31632519615 on
Three independent instances now, across two layers — inside a bats case, and in job setup before any repository code — is what makes this a class rather than weather. And the failure does not just cost a matrix: it wedges the branch. This is the part the original body did not carry, and it is why this is no longer No priority. A which is the correct message for a real disagreement and exactly wrong here: This is the The cost is not theoretical. It converts an unrelated third-party outage into a What the tests are actually about Neither case's stated property needs the install to succeed:
So the ordering fixture exists; it simply lets The per-instance fixes For the
Either way the property to preserve is the one CLOUD-218 bought: Blockers: none. Acceptance
Files: Commit type / bump: Refinement — Ready (a required job's verdict stops depending on who is up) Refinement gate: Definition of Ready & Done. This body carries only specializations.
Two halves, landable separately
On the three questions this body used to carry, so nobody re-derives them. What is the predicate is answered in §2. Can prevention ever be total — no, and the fourth instance above is why: Filed from the CLOUD-363 session, which does not own these files. Another instance, 2026-08-12, and it shows the wedge half as well as the red half. PR #374, head That is The compounding part. The So an upstream 503 on a tool download cost: one wedged SHA, one starved sibling, and an hour of held lease. The two defects are independent and each is survivable alone; together they turn a transient upstream outage into a fleet stall. Worth noting for the fix's shape: the failure is in |
…ed (CLOUD-491) The hold occupies the container while a human reads a plan and left no record of having done so, so a replaced container could not be told from an exited hold or from one never armed. All three look identical from inside, which is why CLOUD-451's acceptance could be asserted but never checked. The obvious sensor does not work here, and it was measured rather than reasoned about. Every write that survived the 2026-08-12 23:45 replacement stops at 23:42:15 - four MCP servers' logs, Serena's log, target/.rustc_info.json, batten-holds - while the container kept serving tool calls until ~23:45. The last 182s of writes did not survive. So `epoch >= btime - poll` reads three minutes stale exactly when it is asked, and a tolerance wide enough to cover that also swallows a hold that legitimately exited minutes earlier. The record is structural instead: `h` per poll, `x` only where the hold chose to stop. A hold killed by the container going down never reaches an `x`, so an `h` in final position is the evidence and no clock is involved. The `x` writes are deliberately on the two terminal paths and never in the exit trap, which runs on the kill too and would report every reclaimed hold as one that stopped on purpose. Recording each boot separately is what tells a preserved disk from a fresh one: the residue probe that would have done it does not work, since a batten-* dir's mtime updates on every add and remove and reads as post-boot on a disk that demonstrably survived. plan-hold-check gains its own suite, following the convention claim-check, land-lock-check and stop-posture-check keep. That is also what makes the predicate mutation-checkable at all: mutant derives its suite as tests/$gate.bats. Both gates join MUTANT_GATES; the three declared mutations are caught.
…ed every run CLOUD-406's prevention half, plus a defect found while writing its test. THE PACKAGE WAS FETCHED ON EVERY CI RUN. `hk.pkl:14` amends a pkl PACKAGE uri, which pkl resolves over the network on every evaluation — a second dependency beside the `hk` binary `mise.lock` pins, and unpinned at runtime. It works locally because container provisioning warms pkl's default cache at `~/.pkl/cache`; nothing warmed it in CI. Measured on run 31632519615: the fetch of `hk@1.54.0.zip` failed, hk aborted with exit 134 before reaching a single gate, and the `failure` grade then wedged the SHA, because `land` re-fires a ready only on a head carrying no graded run. An upstream blip on a release CDN became a red required check on a branch that changed nothing. One `actions/cache` step on `~/.pkl/cache`, keyed on `hk.pkl` itself so a version bump changes the key by construction. The `ci` job only, confirmed rather than assumed: `hooks` (`hk check --all`) is a dependency of `mise run ci` and of nothing else, `cross-check` and `darwin-link` reach `doctor` whose hook probe is skipped under CI, and `msrv`/`semver` run neither. A miss behaves as today and then populates, so the exposure narrows to one run per bump or eviction rather than being removed — it cannot be removed here, because `mise-action`'s own bootstrap fetch precedes every seam this repository controls. THE GATE OVER IT COULD NOT FAIL. `pkl-check` had no suite, and writing one showed why that mattered: every path ended in an unconditional `exit 0` or in an `echo` whose status became the task's, with no `set -e` to stop a failing evaluation reaching them. It passed on a malformed `hk.pkl` and on an unreachable package alike, while hk.pkl:381 routed its `pkl` step through it — the presence of a gate is what stops anyone looking (CLOUD-418). It now reports pkl's status, from one invocation site so a third exit path cannot grow back. `tests/pkl-check.bats` measures both directions over the same command: cold cache plus egress denied fails, warm cache plus egress denied exits 0. The first row is load-bearing — without it the second passes on any machine that has ever evaluated `hk.pkl`. Both rows were red before the exit-status fix. Egress is denied with pkl's own `--http-proxy`, NOT with `HTTPS_PROXY` as CLOUD-406's Ready block specifies. pkl is a native image and ignores the environment variables: the first probe went out over the real network and failed on an unrelated truststore error, which reads as a passing anti-vacuity row while measuring nothing about egress. The flag fails closed in the right direction — if it were ever ignored the fetch would succeed and the cold row would go green. The suite never downloads, so it needs no network of its own; the warm row copies the ambient cache and skips with a diagnostic when there is nothing to copy. The recovery half — one bounded re-fire in `land` — is deliberately not here. Those files are held by another session under a standing takeover. Refs: CLOUD-406
…r checks can be superseded A required check whose workflow omits `types:` defaults to `[opened, synchronize, reopened]`. Both jobs are draft-gated, so a PR created as a draft mints a `skipped` run on `opened` — and then no event remains that could replace it. `land` readies before pushing, and its push moves nothing whenever HEAD is already on the remote, so `ready_for_review` is the only event left and neither workflow was listening. `checks-green` correctly refuses to read a skip as an answer, so `ci-wait` polls unbounded while `land` holds the fleet lease. `ci.yml` and `commit-lint.yml` already declare the full set, which is why this was invisible: on any head where a lap actually pushed, `synchronize` covered it. MEASURED AS A DEADLOCK, NOT A WEDGED BRANCH. #382 and #385 were simultaneously in the same state — every required check SUCCESS, `zizmor` SKIPPED with exactly one run on the SHA — and `main` frozen at 795847b. That freeze is what made it self-sustaining: with nothing to rebase onto, each lap's push was `Everything up-to-date`, so the one event zizmor does subscribe to never fired. Each PR needed the other to move first. Nothing inside the loop breaks that; a commit does, which is why this fix and the unwedge are the same act. Isolated on head 1789dfb: `ci.yml` carries a `skipped` run at 23:40:18 (the draft `opened` event) AND a `success` at 23:41:08. `zizmor.yml` carries only the skip. No push separates the two — the ready is the only event between them. It costs no runner. Both jobs keep their `if: github.event.pull_request.draft == false` guard, so every draft event still skips; the only new firing is the single transition that needs one. This is not a trade against the draft economy — it is the event that economy was missing. `test.yml`'s header claimed a safety property that is false in this case: "`ci-wait` reads an absent path-filtered check as absent rather than pending, so a required check that produces no run here does not hang a landing." True for ABSENT; a draft `opened` event produces PRESENT-AND-SKIPPED, which does hang it. Corrected rather than left documenting the bug as a guarantee. THE GATE FOR THIS PROPERTY IS NOT HERE, DELIBERATELY. CLOUD-503 specifies it in `ci-local-parity`, on the seam at :282-287 that already couples `CI_REQUIRED_CHECKS` to job names in both directions — the third arm asks whether a required check can be REFRESHED, where the first two ask whether it can exist. PR #385 is readied and rewrites that task and its suite (+1323), and this change now lands first, so carrying the gate here would hand another session a merge conflict in a readied PR. CLOUD-503 stays In Progress with that half specified and its CLOUD-418 both-directions obligation written; neither workflow file here is touched by #385, so this half conflicts with nothing. Refs: CLOUD-503
1789dfb to
49b89dd
Compare
|
|
/fast-forward |



CLOUD-406's prevention half, plus a defect found while writing its test.
The package was fetched on every CI run
hk.pkl:14amends a pkl package uri, which pkl resolves over the network on every evaluation — a second dependency beside thehkbinarymise.lockpins, and unpinned at runtime. It works locally because container provisioning warms pkl's default cache at~/.pkl/cache; nothing warmed it in CI.Measured on run 31632519615: the fetch of
hk@1.54.0.zipfailed, hk aborted with exit 134 before reaching a single gate, and thefailuregrade then wedged the SHA —landre-fires a ready only on a head carrying no graded run. An upstream blip on a release CDN became a red required check on a branch that changed nothing.One
actions/cachestep on~/.pkl/cache, keyed onhk.pklitself so a version bump changes the key by construction. Thecijob only, confirmed rather than assumed:cihooks(hk check --all) is a dependency ofmise run ciand of nothing elsecross,darwin-linkdoctor, whose hook probe is skipped under CImsrv,semverA miss behaves as today and then populates, so the exposure narrows to one run per bump or eviction rather than being removed. It cannot be removed here:
mise-action's own bootstrap fetch precedes every seam this repository controls, which is why CLOUD-406 carries a recovery half as well.The gate over it could not fail
pkl-checkhad no suite, and writing one showed why that mattered. Every path ended in an unconditionalexit 0or in anechowhose status became the task's, with noset -eto stop a failing evaluation reaching them. It passed on a malformedhk.pkland on an unreachable package alike — whilehk.pkl:381routed itspklstep through it. It now reports pkl's status, from one invocation site so a third exit path cannot grow back.Test obligation
tests/pkl-check.batsmeasures both directions over the same command:The first row is load-bearing — without it the second passes on any machine that has ever evaluated
hk.pkl(CLOUD-418). Both rows were red before the exit-status fix.Egress is denied with pkl's own
--http-proxy, not withHTTPS_PROXYas CLOUD-406's Ready block specifies. pkl is a native image and ignores the environment variables: the first probe went out over the real network and failed on an unrelated truststore error, which would read as a passing anti-vacuity row while measuring nothing about egress. The flag fails closed in the right direction — if it were ever ignored the fetch would succeed and the cold row would go green. The suite never downloads, so it needs no network of its own.Not here
The recovery half — one bounded re-fire in
land— touchesmise-tasks/landandtests/land.bats, held by another session under a standing takeover.Refs: CLOUD-406