Skip to content

fix(ci): nightly-verify times out in the build phase and reports nothing when it does #1000

Description

@inureyes

Problem / Background

nightly-verify has two independent defects: its cargo test step never finishes inside the job budget, and when the job is killed by that timeout nobody is told. All figures below were measured on 2026-08-02.

Defect 1: the cargo test step never finishes inside the 180-minute budget

.github/workflows/nightly-verify.yml sets timeout-minutes: 180. A job killed by that timeout is reported by GitHub with conclusion cancelled, not failure, so in the run list it reads as "someone cancelled it". That is why it has gone unexamined.

The last four runs of the workflow:

Run Trigger Created Ended Conclusion Duration
30733081983 workflow_dispatch 2026-08-02T04:51:14Z 2026-08-02T07:51:39Z cancelled 3h00m
30712399425 schedule 2026-08-01T18:23:19Z 2026-08-01T21:23:48Z cancelled 3h00m
30655703583 schedule 2026-07-31T18:34:43Z (short) failure fail-fast at cli_help_consistency, issue #962
30570981898 schedule 2026-07-30T18:34:13Z 2026-07-30T21:34:38Z cancelled 3h00m

Three of the last four hit exactly 3h00m. The only run that produced a verdict did so because cargo test is fail-fast and it died early on the harness bug fixed in #996. The workflow has never completed a full pass.

Per-step timings for run 30733081983, from gh api repos/lablup/mlxcel/actions/runs/30733081983/jobs:

cargo fmt:    04:51:26 -> 04:51:30   (4s,   success)
cargo clippy: 04:51:30 -> 04:52:37   (67s,  success)
cargo test:   04:52:37 -> 07:51:28   (179m, cancelled)

The time goes to the build, not to the tests. The run log contains zero Running tests/... lines and zero test result: lines, and the processes the runner terminated at the timeout were ld and clang:

Terminate orphan process: pid (68682) (ld)
Terminate orphan process: pid (68678) (clang)

The 67-second clippy is the load-bearing contrast: the persistent CARGO_TARGET_DIR at $HOME/.cargo-target/mlxcel is warm. But cargo clippy runs clippy-driver and leaves metadata rather than linkable artifacts, so cargo test still has to codegen the library and link roughly 75 integration-test binaries, each against the large MLX static library, on a shared runner. That is where the three hours go.

For calibration, the same suite on the M1 Ultra development machine completes in well under an hour with a warm target/. So this is about the runner and the clippy-then-test artifact split, not about the test count being unreasonable.

Defect 2: a timed-out run notifies nobody

The workflow has a Report a red main step whose purpose, stated in the file's own header comment, is that "a failed scheduled run files (or comments on) a GitHub issue", because "a red Actions run that nobody opens is the same blind spot in a new place".

In run 30733081983 that step's conclusion is skipped. Its condition matches failure, and a timeout surfaces as cancelled, so the notification path does not fire. Three silent timeouts in four days is precisely the blind spot the workflow was created to close.

Why this matters now

The workflow exists (see its header comment) because two deterministically failing tests reached main and sat there unnoticed, one of them for 26 days. #962 fixed the harness bug that was breaking it, and #953 created it. Neither is sufficient while the job cannot finish and cannot report.

Proposed Solution

These are options to evaluate and measure, not a decided design.

  • Drop cargo clippy from the same job, or reorder so the test build is not paying for clippy having warmed a different artifact set. Measure whether running cargo test --no-run first, or running clippy after test, changes the total.
  • Split fmt/clippy and test into separate jobs, so a slow test build does not consume the whole budget and so the three signals stay independent (the file already states it wants them independent).
  • Raise timeout-minutes, but only with a measurement showing what a full pass actually costs. Raising it blindly trades a silent timeout for a longer silent timeout.
  • Reduce link cost: fewer and larger integration-test binaries, or a -C link-arg / linker choice, or cargo nextest, which builds the same binaries but may schedule them better. Consolidating test binaries is a real refactor with its own review cost.
  • Fix the notification independently of the performance work, and first, since it is small: make the reporting step fire on cancelled and on timeout, not only on failure. That alone converts a silent failure into a visible one and is worth landing even if the timeout takes longer to solve.

Acceptance Criteria

  • nightly-verify completes a full pass within its budget on the self-hosted runner, demonstrated by at least one green run, with the measured duration recorded in the PR.
  • A run that times out or is cancelled reports somewhere a human will see it, verified by an actual run and not only by reading the condition.
  • The chosen approach is justified with before and after timings from the real runner, not from the development machine.
  • If the fix trades coverage for time (a narrower selector, a split job), what is no longer covered is stated explicitly in the workflow's header comment, which is unusually thorough and should stay accurate.

Refs

#953 created the workflow. #962 and #996 cover the harness bug that masked this. #997 is a flaky test found in the same suite.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

priority:highHigh prioritystatus:doneCompletedtype:bugBug fixes, error corrections, or issue resolutions

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions