Skip to content

fix(tests): stop session-start's ordering cases depending on who is up - #389

Merged
wenzowski merged 1 commit into
mainfrom
claude/ready-issues-handoff-qnra4q
Aug 13, 2026
Merged

wenzowski merged 1 commit into
mainfrom
claude/ready-issues-handoff-qnra4q

Conversation

@wenzowski

Copy link
Copy Markdown
Contributor

CLOUD-406's third instance — the last half of that issue not already landed.

The coupling was one line

Measured on run 31538830648: Sigstore's TUF metadata endpoint was down, mise install failed, and two cases went red on a tree that touched neither the hook nor the toolchain.

Their properties are order (doctor runs after install, before the preflight — CLOUD-218) and output (the hook emits no {"async": true} — CLOUD-196). Neither asserts anything about the install succeeding. [ "$status" -eq 0 ] was the only coupling, and the hook's step helper sets fail=1 and continues, so every call lands in $CALLS either way.

The stub now intercepts bare install the way it already intercepts run doctor and run container-preflight. Cases that genuinely need a real install opt back in.

The old rationale was half true, and the false half preserved the bug

The header claimed the real exec is what keeps "the two lockfile assertions" non-vacuous:

assertion needs a real install?
greps the hook's source for MISE_LOCKFILE=false mise install no — touches no toolchain
"running the hook leaves the tracked lockfile untouched" yes — a stubbed install writes no lockfile, so it could not observe the CLOUD-223 residue

So this is the issue's option 2, not option 1: keep exactly one case exercising the real install (plus the end-to-end green case) and stop everything else depending on it.

Establish the precondition, never retry the measurement

real_install_or_skip runs mise install directly, before the hook. A failure there is a statement about the container's egress, so it skips with that reason. Past it, the hook's own install cannot fail for a provisioning reason — so a red from the case that follows is a real defect. That is the discriminator option 2 asks for, and the same split lock-check took one layer up. Idempotent and warm, so the second install costs milliseconds.

Shown able to fail, in every direction the acceptance names

mutation result
egress denied (mise install forced to exit 1) 8 pass; cases 7 and 9 skip with the reason named. The two cases that went red on run 31538830648 now pass.
hook stops calling mise install cases 4 and 8 redden
doctor moved outside the synchronous window case 4 reddens

The property CLOUD-218 bought survives, which the issue was explicit about: deleting the assertion is not the fix.

The recovery half needs no commit

CLOUD-406 §2 states the pre-gate-abort discriminator "is not available and must not be built". It was built — CLOUD-404 landed mise-tasks/nonverdict-scan plus absorbed_transient/charge_transient on land's red arm, which scans each failed run, re-runs the failed jobs when none reached a verdict, refunds the lap and charges a bounded counter. Stronger than §2's receipt heuristic on every axis §2 cared about. Recorded on the issue so it reads as superseded rather than unimplemented.

Refs: CLOUD-406

@linear-code

linear-code Bot commented Aug 13, 2026 •

Copy link
Copy Markdown
CLOUD-406 A required job reaches the network before it reaches a gate, so an upstream outage reds a branch that changed nothing — and the `failure` grade then wedges the SHA

Why

Measured on PR #302 (CLOUD-363), 2026-08-11, run 31538830648 on 60c5d79. Two
bats cases failed in CI and pass locally on the identical tree:

not ok 736 the hook is synchronous — async would restore the race it closes
not ok 738 doctor runs inside the synchronous window — after install, before the preflight

The assertion that failed is [ "$status" -eq 0 ] on the hook invocation, and
the reason is in the captured output:

::error:: session-start: mise-install failed — see /tmp/session-start-mise-install.log
mise ERROR Failed to install aqua:cli/cli@2.97: GitHub artifact attestations
  verification failed: Verification failed: Workflow verification failed:
  expected 'cli/cli/.github/workflows/deployment.yml', found certificate identity:
  "https://dotcom.releases.github.com"; Sigstore error: TUF error:
  TUF repository load failed: transport error: timestamp.json not found
::error:: session-start: setup incomplete — expect missing tools or MCP servers

Nothing in #302's diff touches the hook, the toolchain, or mise.toml. The same
suite ran green locally on the same commit (bats tests/session-start.bats,
exit 0), and green in every one of the 13 verify runs land made across three
invocations. The tree is not what changed; Sigstore's TUF metadata endpoint was.

Two more instances, one layer further out again. Measured 2026-08-12 on PR #325 (CLOUD-367), run 31632519615 on a4dec2f9. Both died in job setup, before any repository code ran at all:

ci:           Eval error: HTTP fetch failed for
              https://github.com/jdx/hk/releases/download/v1.54.0/hk@1.54.0.zip
              thread 'main' panicked at src/settings.rs:101:30:
              Failed to load configuration: Failed to read config file: .../hk.pkl
              Aborted (core dumped)                                    -> exit 134

darwin-link:  curl -fsSL .../mise-v2026.8.5-linux-x64.tar.zst
              curl: (56) Connection died, tried 5 times before giving up -> exit 1

final then went red on its needs, correctly. Evaluating hk.pkl resolves the hk pkl package over the network on every run — mise.lock pins the hk binary, but the pkl package it imports is a second dependency resolved per evaluation, and unpinned at runtime.

Three independent instances now, across two layers — inside a bats case, and in job setup before any repository code — is what makes this a class rather than weather.

And the failure does not just cost a matrix: it wedges the branch. This is the part the original body did not carry, and it is why this is no longer No priority. A failure is a graded run, and land re-fires a ready only on a head carrying no graded run. So once the blip grades the SHA, that SHA can never buy another run through the loop. Measured on #325: the next land won the lease, pushed nothing (HEAD unchanged), skipped the ready re-fire because the head is graded, re-read the same stale red, and stopped with

::error:: land: CI is red on a4dec2f9. A red run on a verified branch means verify
          and CI disagree — fix the mismatch locally, then run land again.

which is the correct message for a real disagreement and exactly wrong here: verify and CI do not disagree, because CI never evaluated anything. Recovery took rerun_failed_jobs against the run id — a manual step outside the loop land exists to drive, exactly as CLOUD-363 records for its own non-answer.

This is the lock-check shape, one layer out. .claude/rules/toolchain.md
already records the principle from CLOUD's earlier round on it: a gate whose*
*verdict comes from a remote API is not testing the commit — it answers "did
upstream change", and fails a branch for drift it did not cause. lock-check
was split into lock-complete (committed bytes, no network) and a scheduled
currency job for exactly this reason. These two cases re-introduce the same
coupling by the back door: they execute the hook end-to-end, and the hook's
first real act is a network install with attestation verification.

The cost is not theoretical. It converts an unrelated third-party outage into a
red required check, which land correctly reads as red — so it re-drafts the
PR and stops, and the branch pays a full lap (rebase + verify + a CI run) to
learn the same thing again once upstream recovers.

What the tests are actually about

Neither case's stated property needs the install to succeed:

  • 736 asserts the hook does not emit {"async": true} — a property of what it
    prints and of its source text.
  • 738 asserts doctor runs after mise install and before the preflight — a
    property of ORDER, and the suite already stubs mise to short-circuit
    run container-preflight and run doctor while exec-ing the real binary
    for everything else.

So the ordering fixture exists; it simply lets mise install through to the
network. The exit-status assertion is what couples them to upstream
availability, and it is the assertion neither case is really making.

The per-instance fixes

For the session-start.bats instance the choice belongs to whoever owns that file:

  1. Extend the existing mise stub to intercept bare install the way it
    already intercepts run doctor, so the ordering is asserted from $CALLS
    with no network at all. Cheapest, and matches the fixture's own design.
  2. Keep one case that genuinely exercises the install, and mark it so a
    provisioning failure is distinguishable from a hook defect — the split
    lock-check took.

Either way the property to preserve is the one CLOUD-218 bought: doctor runs
inside the synchronous window, after mise install and before the preflight.

Blockers: none.

Acceptance

  • A required check is never red because an upstream host was unreachable. The branch's verdict depends on the committed bytes, not on who is up.
  • tests/session-start.bats passes with no network reachable, while a hook that stops calling mise install, or calls doctor outside the window, still fails it — the fix must not be "delete the assertion".
  • hk.pkl evaluates without a live fetch of its own pkl package, on the same footing as every other pinned tool; a version bump stays a deliberate, committed change.
  • Whatever remains network-dependent fails in a way a caller can tell from a verdict — today the abort is exit 134, which is not in the 0/1/2/3 contract at all.

Files: tests/session-start.bats and mise-tasks/session-start (only if the fix needs a seam) for the first instance; whatever pins and caches the pkl package for the second — mise.toml, mise.lock, and the workflow's setup steps.

Commit type / bump: fix(tests) for the session-start half, ci for the toolchain-caching half. Neither touches crates/, so release-plz cuts nothing.


Refinement — Ready (a required job's verdict stops depending on who is up)

Refinement gate: Definition of Ready & Done. This body carries only specializations.

  • Source of truth (§1). The rule already exists and is not re-typed here: .claude/rules/toolchain.md records that a gate whose verdict comes from a remote API is not testing the commit, which is why lock-check was split into lock-complete (committed bytes, no network) and a scheduled currency job. This issue extends that settled rule from tasks to required-job setup, and each per-instance fix lives with its own files rather than in a new shared authority.

  • Computable predicate (§2). Two, one per half, each a command and an exit code over an object it decides.

    Prevention — hk.pkl evaluates with no network. hk.pkl:14 amends a pkl package URI (package://github.com/jdx/hk/releases/download/v1.54.0/hk@1.54.0#/Config.pkl, and line 16 imports from the same), which pkl resolves on every evaluation — a second dependency beside the hk binary mise.lock already pins. The predicate: with egress denied, pkl eval hk.pkl exits 0. Deny it by pointing HTTPS_PROXY/HTTP_PROXY at a closed local port rather than by dropping capabilities — portable, unprivileged, and it fails closed if the proxy vars are ever ignored, because the fetch then succeeds and the anti-vacuity case in §7 catches it. Home: mise-tasks/pkl-check, extended. It already owns "does hk.pkl evaluate", is already wired into the gate (hk.pkl:385), and its own header already names this coupling — "pkl resolves amends 'package://…' over the network on a cold cache". No new authority. The cache key derives from the hk version, whose agreement with the amends URL hk-version already gates, so nothing new has to be kept honest.

    Recovery — one bounded re-fire. Distinguishing a pre-gate abort from a real verdict is not available and must not be built: both grade failure, the failing step is the same step (mise run ci, inside which hk panics), and the only remaining discriminator is log text — prose matching, which this repo refuses. The predicate needs no such distinction. land, holding the lease, on a head where all three hold — (a) a green verify receipt exists for this exact HEAD, (b) the required set is red, (c) no re-fire receipt exists for this HEAD — re-fires the ready once and records a SHA-keyed re-fire receipt under .git/batten-receipts/, the same shape ready-guard's receipts already use. A second land on the same head stops exactly as today. Pure filesystem state plus checks-green's existing verdict: no log read, no conclusion heuristic, no new event, and compatible with CLOUD-420 because it happens inside the lease land already holds.

    Why the re-fire is not drive-to-green. Re-running land on an unchanged red head is a guaranteed dead end today — the push moves nothing, the ready is skipped because the head is graded, the same stale red is re-read. This buys one more opinion, bounded at one, only when someone deliberately re-ran land. Condition (a) is what keeps it honest: it fires only where local proof and CI actually disagree, which is precisely the case land's current error message describes and then offers no move for.

  • Effect (§3). No command surface and no effect-table row. Test fixtures and CI configuration only.

  • Output & exit (§5). Unchanged for the passing case. Note for whoever implements: the abort arrives as exit 134, outside the 0/1/2/3 table, so a caller cannot distinguish "could not look" from a verdict by exit code today. Any reporting stays pointer-only — the package coordinate and the failing host, never fetched bytes.

  • Commit / bump (§6). fix → patch until 0.1.0. The prevention half alone is ci and would release nothing; the recovery and session-start halves are fix, and below 0.1.0 every release-worthy type collapses to a patch. Nothing under crates/ changes either way, so release-plz cuts a release only if a crate change rides along.

  • Test obligation (§7). Each case asserts an exit code, never prose.

    Prevention. (a) With egress denied and the cache warm, pkl eval hk.pkl exits 0. (b) With the pkl module cache purged and egress denied, the same command fails. (b) is the load-bearing case: without it the test passes on any warm machine forever and proves nothing, which is exactly the "a new gate is never shown to fail" defect CLOUD-418 landed a discipline against. Both directions or neither.

    Recovery, in tests/land.bats. (c) A red head carrying a green verify receipt and no re-fire receipt re-fires the ready once and writes the receipt. (d) The same head on a second land stops, as today — the bound, and the case a later "simplification" would quietly remove. (e) A red head with no verify receipt never re-fires, so the re-fire cannot become an unconditional retry. The existing lap-count assertions must still hold.

    Session-start instance. The suite passes with egress unavailable, and the anti-vacuity case holds: a hook that stops calling mise install, or calls doctor outside the synchronous window, still fails. The property CLOUD-218 bought must survive the fix.

  • Blockers (§8). None. relatedTo CLOUD-363 — the same shape, a non-answer graded as terminal with no exit through the loop; its cancelled half has landed in checks-green, and this is the failure-from-an-abort half it deliberately does not cover. Also CLOUD-481 (folded in here), CLOUD-476, CLOUD-218, and CLOUD-369/CLOUD-399 where question 3 overlaps.

Two halves, landable separately

  1. Prevention — mise-tasks/pkl-check, mise.toml, and the workflow's cache step. Unowned, independently landable, and it removes the common case. Do this one first.
  2. Recovery — mise-tasks/land and tests/land.bats. These files are held by another session (claude/landing-process-conflicts-3ogq6s under a standing takeover of the landing loop). Coordinate before starting; do not open a competing branch against them.

On the three questions this body used to carry, so nobody re-derives them. What is the predicate is answered in §2. Can prevention ever be total — no, and the fourth instance above is why: mise-action's own bootstrap fetch precedes every seam this repository controls, so no in-repo gate, egress-blocked suite, or pkl cache can reach it. That is what makes the recovery half required rather than conditional. Is a sanctioned re-fire wanted — yes, in the bounded form §2 specifies; the two questions had converged, and one design closes both.

Filed from the CLOUD-363 session, which does not own these files.


Another instance, 2026-08-12, and it shows the wedge half as well as the red half.

PR #374, head 4649085. The msrv job died before reaching any gate:

[command]/usr/bin/curl -fsSL https://github.com/jdx/mise/releases/download/v2026.8.5/mise-v2026.8.5-linux-x64.tar.zst
curl: (22) The requested URL returned error: 503
##[error]The process '/usr/bin/curl' failed with exit code 22

That is mise-action fetching mise itself from GitHub's release CDN. Nothing in the diff was reached, nothing was judged, and msrv graded failure. final then graded failure because msrv is in its needs:, so a 503 on a tarball download produced a red on the one check-run branch protection requires.

The compounding part. The failure grade stuck to the SHA, exactly as this issue says. But the branch did not merely go red — it went red invisibly: checks-green was simultaneously polling on a zizmor that had graded skipped, so ci-wait reported "not an answer yet" over a set that already contained two failures (recorded on CLOUD-436). land therefore never reached its red-stop, and held the fleet landing lease for ~1 hour with another branch admitted behind it.

So an upstream 503 on a tool download cost: one wedged SHA, one starved sibling, and an hour of held lease. The two defects are independent and each is survivable alone; together they turn a transient upstream outage into a fleet stall.

Worth noting for the fix's shape: the failure is in mise-action's own bootstrap, before any repository code runs, so no in-repo gate ordering can catch it. The remedies that reach it are retry/caching at the action step or treating a pre-gate infrastructure failure as a distinct conclusion from a graded red — which is the same distinction CLOUD-470 is drawing on the agent-facing side.

Review in Linear

CLOUD-406's third instance, and the last half of it not already landed.

Measured on run 31538830648: Sigstore's TUF metadata endpoint was down, so
`mise install` failed, and two cases went red on a tree that touched neither the
hook nor the toolchain. Their properties are ORDER (doctor runs after install,
before the preflight — CLOUD-218) and OUTPUT (the hook emits no `{"async": true}`
— CLOUD-196). Neither asserts anything about the install succeeding. The
`[ "$status" -eq 0 ]` line was the only coupling, and it is load-bearing for
neither: the hook's `step` helper sets `fail=1` and CONTINUES, so every call is
recorded in `$CALLS` whether the install worked or not.

The stub now intercepts bare `install` the way it already intercepts
`run doctor` and `run container-preflight`, and the cases that genuinely need a
real one opt back in.

THE OLD RATIONALE WAS HALF TRUE, AND THE FALSE HALF IS WHAT COST A RED CHECK.
The header claimed the real exec is what keeps "the two lockfile assertions"
non-vacuous. One of them greps the hook's source text for
`MISE_LOCKFILE=false mise install` and touches no toolchain at all. The other —
"running the hook leaves the tracked lockfile untouched" — genuinely does need
one, because a stubbed install writes no lockfile and so could not observe the
CLOUD-223 residue. So this is the issue's option 2 rather than option 1: keep
exactly one case exercising the real install, plus the end-to-end green case, and
stop everything else depending on it.

ESTABLISH THE PRECONDITION, NEVER RETRY THE MEASUREMENT. `real_install_or_skip`
runs `mise install` directly before the hook. A failure THERE is a statement
about the container's egress, so it skips with that reason; past it, the hook's
own install cannot fail for a provisioning reason, so a red from the case that
follows is a real defect. That is the discriminator option 2 asks for, and the
same split `lock-check` took one layer up. Idempotent and warm, so the second
install costs milliseconds.

Shown able to fail, in every direction the acceptance names:

  egress denied (`mise install` forced to exit 1)
    -> 8 pass, and cases 7 and 9 SKIP with the reason named.
       The two cases that went red on run 31538830648 now pass.

  hook stops calling `mise install`
    -> cases 4 and 8 redden.

  doctor moved outside the synchronous window
    -> case 4 reddens.

So the property CLOUD-218 bought survives, which is the half the issue was
explicit about: deleting the assertion is not the fix.

The recovery half of CLOUD-406 needs no commit — CLOUD-404 landed it as
`nonverdict-scan` plus `absorbed_transient`/`charge_transient` on `land`'s red
arm, in a stronger form than §2 specified, and §2's premise that the
discriminator "must not be built" is now false. Recorded on the issue.

Refs: CLOUD-406
@wenzowski
wenzowski marked this pull request as ready for review August 13, 2026 03:01
@wenzowski
wenzowski force-pushed the claude/ready-issues-handoff-qnra4q branch from e075273 to 9a9a524 Compare August 13, 2026 03:01
@sonarqubecloud

Copy link
Copy Markdown

@wenzowski

Copy link
Copy Markdown
Contributor Author

/fast-forward

@wenzowski
wenzowski merged commit 9a9a524 into main Aug 13, 2026
8 checks passed
@wenzowski
wenzowski deleted the claude/ready-issues-handoff-qnra4q branch August 13, 2026 03:10
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants