Skip to content

perf(ci): bind the warm job's install list, which is worth 163 of its 171 seconds - #621

Merged
wenzowski merged 1 commit into
mainfrom
claude/ci-performance-degradation-rplznx
Aug 21, 2026
Merged

wenzowski merged 1 commit into
mainfrom
claude/ci-performance-degradation-rplznx

Conversation

@wenzowski

@wenzowski wenzowski commented Aug 21, 2026 •

Copy link
Copy Markdown
Contributor

The warm cycle CLOUD-840 landed costs 358s, and its rust-cache step is 171s of that — which read as the cache being slow on Windows. It isn't.

From the step's own log, second warm cycle:

15:52:07  step starts
15:54:50  "Cache Configuration" printed      <- 163s of silence
15:54:50  Cache hit
15:54:51 → 15:54:54  download 182MB          <- 2.7s, 47-103 MB/s
15:54:54 → 15:54:58  tar extract             <- 3.6s

The cache work is about eight seconds. The rest is rust-cache's pre-flight, which shells out to cargo and rustc to compose its key before emitting a line — and with mise's auto-install on, each of those calls can enter the install/verify path.

rust.yml's windows job restores the identical 182MB entry in 12s. The only difference between the two jobs is the two variables this adds.

The regression is mine

CLOUD-812 set both variables in ci.yml, commit-lint.yml and zizmor.yml and stopped there — correctly, since release-plz.yml had no job that ran cargo. Adding one without them re-opened it two commits later.

A gap this exposes, filed rather than widened

ci-tools-check's absence direction is scoped to pull_request workflows. That is the right call for the property it enforces — a push-triggered workflow stands in for nothing in verify — but it means a cargo job on a push workflow can run with auto-install on and nothing goes red. Widening a measured, mutation-gated property to cover a second case is not a change to make in passing, so it gets its own row.

What it does to the trade

Stated as a prediction, not a result: the warm cycle should fall from 358s to roughly 195s, ~6.5 billed minutes, against the 4.4 a PR's first run saves. Still net negative — but about two billed minutes per merge rather than 7.6, which is a different decision: 2 billed minutes to take 133s off every PR's critical path is defensible where 7.6 was not.

The measurement is the next merge's warm cycle, and it decides keep-or-revert.

Also recorded, not acted on

Of the warm job's 150s compile, ~127s is compiling batten itself — and rust-cache runs cache-workspace-crates: false, so that output is discarded from the cache the job exists to fill. Only the dependency artifacts matter, and they are done before batten starts. There is no first-class cargo flag for "dependencies only", so if the trade is still marginal after this, that is where the next ~2 billed minutes are.

Verification

ci-tools-check, ci-local-parity and zizmor green locally.

DO-NOT-CLOSE CLOUD-840


Generated by Claude Code

Summary by CodeRabbit

  • Chores
    • Updated release automation settings to prevent automatic installation of development tools during task execution and environment setup.
    • Applied the same settings to cache-warming processes for more predictable release workflows.

@linear-code

linear-code Bot commented Aug 21, 2026 •

Copy link
Copy Markdown
CLOUD-840 Every PR compiles the Rust tree cold: the caches exist, are byte-identical across PRs, and no PR can read another's — so `windows` writes 182MB per run that nothing ever restores

Why

CLOUD-813 measured cargo test against cargo nextest run on windows-latest and got 499s against 501s — two seconds. A runner 22% faster at executing this suite on Linux cannot move the job carrying 50.6% of the bill, which is only possible if execution is a small share of that step. The ~8m20s is compilation, and that is what this row is about.

Measured 2026-08-21 from the Actions cache API, not inferred from step timings:

  • The Windows rust cache key is stable and identical across pull requests — v0-rust-windows-windows-Windows_NT-x64-d44ea756-e905c9ae — present on refs/pull/572, 601, 604, 607 and 608, each about 182MB. Five copies of the same bytes.
  • It is absent from refs/heads/main. GitHub scopes a cache read to the run's own ref plus the base branch, so no pull request can read another's copy and none has a base-branch copy to inherit.
  • So every Windows run restores in 7s — a miss — compiles cold, and then spends 70s writing 182MB that only that pull request could ever read. The branch is deleted at merge.
  • The only rust caches on main belong to perf, because perf.yml runs on a schedule. rust.yml and ci.yml trigger on pull_request only, so theirs never land there.
  • The repository holds 10.64 GB against GitHub's 10 GB cap — 6.69 GB rust, 3.95 GB mise across 40 entries. Eviction is live and LRU, so the duplicates are evicting each other.

This reverses a recorded decision, and that decision named the conditions.

CLOUD-176 considered base-branch warming and recommended against it: about 178s of job-time per merge to save about 60s on one pull request's first run, "roughly 3× the job-time it saves." That costing lists four legs — ci, cross and darwin-link twice — all Linux, all billed at 1×. It predates the windows job entirely, and the arithmetic does not survive its arrival.

It also wrote down when to revisit, and both conditions now hold:

  • "PRs routinely land on their first CI run" — land merges on the first green lap, which is the ordinary case rather than the exception.
  • "the cold-vs-warm gap grows well past ~15s per job" — the Windows gap is most of 8m20s, at 2× billing.

Recording the reversal rather than quietly re-filing, because CLOUD-176's recommendation is written down and someone will find it.

Refinement — Ready (warm the base branch, and get under the cap so the warm entry survives)

  • Source of truth (§1). The Actions cache API for what exists and where, and the windows job's own step timings before and after for what it buys. Neither is a number a human types, and the cost side is measured on the runner that pays it rather than inferred from Linux.
  • Computable predicate (§2). After warming, a v0-rust-windows-… entry exists on refs/heads/main, and the first run of a fresh pull request restores it rather than reporting a miss. The gate is the restore, not the wall clock: a run that restores and is still slow is a different finding, and a run that is fast because nothing changed is not evidence.
  • No new trigger, and AGENTS.md is not bent (§2). release-plz.yml already runs on push: branches: [main] and is the only workflow that does. CLOUD-176 identifies it as the attachment point. The rule this repository holds is that CI does not re-run on main — a warm job compiles to fill a cache and asserts nothing, so it is not a second verdict on an already-tested SHA.
  • The cap is part of the work, not a footnote (§2). A cache the pull requests never read is pure eviction pressure on the one they would. 3.95 GB of mise tool caches across 40 entries and roughly 0.9 GB of duplicate Windows rust caches are the bulk of the overflow. Warming without reclaiming that space buys an entry that is evicted before it is read.
  • Deliberately not in scope (§2). The runner swap — CLOUD-813, settled. The toolchain-install overhead — CLOUD-812. The paths filter's omissions — CLOUD-828.
  • Effect (§3). write — CI configuration only. No verb, no command surface, no change to what any gate proves.
  • Output and exit (§5). Unchanged; a warm job asserts nothing and must not be able to red a landing.
  • Commit / bump (§6). ci(cache) — no bump. Nothing under crates/.
  • Test obligation (§7). The cost and the benefit both stated as measurements before the change is kept: the warm job's own billed minutes, against the first-run delta on a fresh pull request measured against the 499s baseline on record. A warm cycle that costs more than it saves is the outcome CLOUD-176 predicted and must be recorded as such rather than absorbed.
  • Blockers (§8). None. relatedTo CLOUD-176 (the decision this reverses, and whose revisit clause authorises it), CLOUD-813 (whose negative result is what redirects the effort here), CLOUD-812 (the overhead layer of the same ledger), CLOUD-398 (the job graph half).

Acceptance

  • A v0-rust-windows-… cache exists on refs/heads/main.
  • The first Rust run of a fresh pull request reports a cache hit rather than a 7s miss, shown against a run that did not.
  • Total cache size is under the 10 GB cap, so the warm entry is not evicted before it is read.
  • The warm cycle's billed minutes and the per-run saving are both recorded, whichever way the trade falls.
  • The Windows job stops writing 182MB on every run that nothing restores.

Review in Linear

@coderabbitai

coderabbitai Bot commented Aug 21, 2026 •

Copy link
Copy Markdown

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 15782171-46d4-4a2c-9aa9-8520fe76a42b

📥 Commits

Reviewing files that changed from the base of the PR and between 6741eab and b3c2210.

📒 Files selected for processing (1)
  • .github/workflows/release-plz.yml

Included review availability: Your plan provides up to 10 included reviews per hour; 7 remain after this review.


📝 Walkthrough

Walkthrough

The release-plz workflow now disables mise automatic tool installation during task execution and mise exec. The settings apply at the workflow level, including the Windows cache-warming job.

Changes

Release-plz mise configuration

Layer / File(s) Summary
Disable mise automatic installation
.github/workflows/release-plz.yml
The workflow adds environment settings that disable automatic mise installation for tasks and mise exec. The settings also apply to the Windows cache-warming job.

Estimated code review effort: 1 (Trivial) | ~5 minutes

Merge Risk: ⚪ Minimal · up to b3c22

This PR makes a localized CI workflow change to avoid unnecessary tool installation during cache setup. No actionable merge-blocking risk remains beyond normal checks and review.

Suggested reviewers: claude

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly identifies the CI performance change and the warm job's install-list binding, which matches the pull request objective.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check. Docstring coverage is scoped to functions touched by this diff. Analyzed 0 functions across 0 files. (1 skipped: 1 unsupported.)
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch claude/ci-performance-degradation-rplznx

Comment @coderabbitai help to get the list of available commands.

@wenzowski
wenzowski marked this pull request as ready for review August 21, 2026 17:20
@wenzowski
wenzowski force-pushed the claude/ci-performance-degradation-rplznx branch from b3c2210 to 671a639 Compare August 21, 2026 18:02
… 171 seconds

CLOUD-840's warm cycle costs 358s and its `rust-cache` step is 171s of that,
which read as the cache being slow on Windows. It is not. From the step's own
log, second warm cycle:

  15:52:07  step starts
  15:54:50  "Cache Configuration" printed      <- 163s of silence
  15:54:50  Cache hit
  15:54:51 -> 15:54:54  download 182MB          <- 2.7s, 47-103 MB/s
  15:54:54 -> 15:54:58  tar extract             <- 3.6s

The cache work is about eight seconds. The rest is rust-cache's pre-flight,
which shells out to `cargo` and `rustc` to compose its key before printing
anything — and with mise's auto-install on, every one of those calls can enter
the install/verify path.

`rust.yml`'s `windows` job restores the identical 182MB entry in 12s. The only
difference between the two jobs is the two variables this commit adds.

This is CLOUD-812 in a third workflow, and the regression is mine. That issue
set both variables in `ci.yml`, `commit-lint.yml` and `zizmor.yml` and stopped
there, correctly: `release-plz.yml` had no job that ran cargo. Adding one
without them re-opened it two commits later.

`ci-tools-check` structurally cannot catch this. Its absence direction is scoped
to `pull_request` workflows — the right call for the property it enforces, since
a push-triggered workflow stands in for nothing in `verify` — but it means a
cargo job on a push workflow can run with auto-install on and nothing goes red.
Filed rather than widened here; widening a measured, mutation-gated property to
cover a second case is not a change to make in passing.

What this does to the trade, stated as a prediction rather than a result: the
warm cycle should fall from 358s to roughly 195s, about 6.5 billed minutes,
against the 4.4 a PR's first run saves. Still net negative, but around two
billed minutes per merge rather than 7.6 — which is a different decision. The
measurement is the next merge's warm cycle.

Refs: CLOUD-840
@wenzowski
wenzowski force-pushed the claude/ci-performance-degradation-rplznx branch from 671a639 to 092e2d8 Compare August 21, 2026 18:16
@sonarqubecloud

Copy link
Copy Markdown

@wenzowski

Copy link
Copy Markdown
Contributor Author

/fast-forward

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant