Repository navigation
fix(tests): stop session-start's ordering cases depending on who is up - #389
Conversation
CLOUD-406 A required job reaches the network before it reaches a gate, so an upstream outage reds a branch that changed nothing — and the `failure` grade then wedges the SHA
Why Measured on PR #302 (CLOUD-363), 2026-08-11, run 31538830648 on The assertion that failed is Nothing in #302's diff touches the hook, the toolchain, or Two more instances, one layer further out again. Measured 2026-08-12 on PR #325 (CLOUD-367), run 31632519615 on
Three independent instances now, across two layers — inside a bats case, and in job setup before any repository code — is what makes this a class rather than weather. And the failure does not just cost a matrix: it wedges the branch. This is the part the original body did not carry, and it is why this is no longer No priority. A which is the correct message for a real disagreement and exactly wrong here: This is the The cost is not theoretical. It converts an unrelated third-party outage into a What the tests are actually about Neither case's stated property needs the install to succeed:
So the ordering fixture exists; it simply lets The per-instance fixes For the
Either way the property to preserve is the one CLOUD-218 bought: Blockers: none. Acceptance
Files: Commit type / bump: Refinement — Ready (a required job's verdict stops depending on who is up) Refinement gate: Definition of Ready & Done. This body carries only specializations.
Two halves, landable separately
On the three questions this body used to carry, so nobody re-derives them. What is the predicate is answered in §2. Can prevention ever be total — no, and the fourth instance above is why: Filed from the CLOUD-363 session, which does not own these files. Another instance, 2026-08-12, and it shows the wedge half as well as the red half. PR #374, head That is The compounding part. The So an upstream 503 on a tool download cost: one wedged SHA, one starved sibling, and an hour of held lease. The two defects are independent and each is survivable alone; together they turn a transient upstream outage into a fleet stall. Worth noting for the fix's shape: the failure is in |
CLOUD-406's third instance, and the last half of it not already landed.
Measured on run 31538830648: Sigstore's TUF metadata endpoint was down, so
`mise install` failed, and two cases went red on a tree that touched neither the
hook nor the toolchain. Their properties are ORDER (doctor runs after install,
before the preflight — CLOUD-218) and OUTPUT (the hook emits no `{"async": true}`
— CLOUD-196). Neither asserts anything about the install succeeding. The
`[ "$status" -eq 0 ]` line was the only coupling, and it is load-bearing for
neither: the hook's `step` helper sets `fail=1` and CONTINUES, so every call is
recorded in `$CALLS` whether the install worked or not.
The stub now intercepts bare `install` the way it already intercepts
`run doctor` and `run container-preflight`, and the cases that genuinely need a
real one opt back in.
THE OLD RATIONALE WAS HALF TRUE, AND THE FALSE HALF IS WHAT COST A RED CHECK.
The header claimed the real exec is what keeps "the two lockfile assertions"
non-vacuous. One of them greps the hook's source text for
`MISE_LOCKFILE=false mise install` and touches no toolchain at all. The other —
"running the hook leaves the tracked lockfile untouched" — genuinely does need
one, because a stubbed install writes no lockfile and so could not observe the
CLOUD-223 residue. So this is the issue's option 2 rather than option 1: keep
exactly one case exercising the real install, plus the end-to-end green case, and
stop everything else depending on it.
ESTABLISH THE PRECONDITION, NEVER RETRY THE MEASUREMENT. `real_install_or_skip`
runs `mise install` directly before the hook. A failure THERE is a statement
about the container's egress, so it skips with that reason; past it, the hook's
own install cannot fail for a provisioning reason, so a red from the case that
follows is a real defect. That is the discriminator option 2 asks for, and the
same split `lock-check` took one layer up. Idempotent and warm, so the second
install costs milliseconds.
Shown able to fail, in every direction the acceptance names:
egress denied (`mise install` forced to exit 1)
-> 8 pass, and cases 7 and 9 SKIP with the reason named.
The two cases that went red on run 31538830648 now pass.
hook stops calling `mise install`
-> cases 4 and 8 redden.
doctor moved outside the synchronous window
-> case 4 reddens.
So the property CLOUD-218 bought survives, which is the half the issue was
explicit about: deleting the assertion is not the fix.
The recovery half of CLOUD-406 needs no commit — CLOUD-404 landed it as
`nonverdict-scan` plus `absorbed_transient`/`charge_transient` on `land`'s red
arm, in a stronger form than §2 specified, and §2's premise that the
discriminator "must not be built" is now false. Recorded on the issue.
Refs: CLOUD-406
e075273 to
9a9a524
Compare
|
|
/fast-forward |



CLOUD-406's third instance — the last half of that issue not already landed.
The coupling was one line
Measured on run 31538830648: Sigstore's TUF metadata endpoint was down,
mise installfailed, and two cases went red on a tree that touched neither the hook nor the toolchain.Their properties are order (doctor runs after install, before the preflight — CLOUD-218) and output (the hook emits no
{"async": true}— CLOUD-196). Neither asserts anything about the install succeeding.[ "$status" -eq 0 ]was the only coupling, and the hook'sstephelper setsfail=1and continues, so every call lands in$CALLSeither way.The stub now intercepts bare
installthe way it already interceptsrun doctorandrun container-preflight. Cases that genuinely need a real install opt back in.The old rationale was half true, and the false half preserved the bug
The header claimed the real exec is what keeps "the two lockfile assertions" non-vacuous:
MISE_LOCKFILE=false mise installSo this is the issue's option 2, not option 1: keep exactly one case exercising the real install (plus the end-to-end green case) and stop everything else depending on it.
Establish the precondition, never retry the measurement
real_install_or_skiprunsmise installdirectly, before the hook. A failure there is a statement about the container's egress, so it skips with that reason. Past it, the hook's own install cannot fail for a provisioning reason — so a red from the case that follows is a real defect. That is the discriminator option 2 asks for, and the same splitlock-checktook one layer up. Idempotent and warm, so the second install costs milliseconds.Shown able to fail, in every direction the acceptance names
mise installforced to exit 1)mise installThe property CLOUD-218 bought survives, which the issue was explicit about: deleting the assertion is not the fix.
The recovery half needs no commit
CLOUD-406 §2 states the pre-gate-abort discriminator "is not available and must not be built". It was built — CLOUD-404 landed
mise-tasks/nonverdict-scanplusabsorbed_transient/charge_transientonland's red arm, which scans each failed run, re-runs the failed jobs when none reached a verdict, refunds the lap and charges a bounded counter. Stronger than §2's receipt heuristic on every axis §2 cared about. Recorded on the issue so it reads as superseded rather than unimplemented.Refs: CLOUD-406