Problem / Background
nightly-verify has two independent defects: its cargo test step never finishes inside the job budget, and when the job is killed by that timeout nobody is told. All figures below were measured on 2026-08-02.
Defect 1: the cargo test step never finishes inside the 180-minute budget
.github/workflows/nightly-verify.yml sets timeout-minutes: 180. A job killed by that timeout is reported by GitHub with conclusion cancelled, not failure, so in the run list it reads as "someone cancelled it". That is why it has gone unexamined.
The last four runs of the workflow:
| Run |
Trigger |
Created |
Ended |
Conclusion |
Duration |
| 30733081983 |
workflow_dispatch |
2026-08-02T04:51:14Z |
2026-08-02T07:51:39Z |
cancelled |
3h00m |
| 30712399425 |
schedule |
2026-08-01T18:23:19Z |
2026-08-01T21:23:48Z |
cancelled |
3h00m |
| 30655703583 |
schedule |
2026-07-31T18:34:43Z |
(short) |
failure |
fail-fast at cli_help_consistency, issue #962 |
| 30570981898 |
schedule |
2026-07-30T18:34:13Z |
2026-07-30T21:34:38Z |
cancelled |
3h00m |
Three of the last four hit exactly 3h00m. The only run that produced a verdict did so because cargo test is fail-fast and it died early on the harness bug fixed in #996. The workflow has never completed a full pass.
Per-step timings for run 30733081983, from gh api repos/lablup/mlxcel/actions/runs/30733081983/jobs:
cargo fmt: 04:51:26 -> 04:51:30 (4s, success)
cargo clippy: 04:51:30 -> 04:52:37 (67s, success)
cargo test: 04:52:37 -> 07:51:28 (179m, cancelled)
The time goes to the build, not to the tests. The run log contains zero Running tests/... lines and zero test result: lines, and the processes the runner terminated at the timeout were ld and clang:
Terminate orphan process: pid (68682) (ld)
Terminate orphan process: pid (68678) (clang)
The 67-second clippy is the load-bearing contrast: the persistent CARGO_TARGET_DIR at $HOME/.cargo-target/mlxcel is warm. But cargo clippy runs clippy-driver and leaves metadata rather than linkable artifacts, so cargo test still has to codegen the library and link roughly 75 integration-test binaries, each against the large MLX static library, on a shared runner. That is where the three hours go.
For calibration, the same suite on the M1 Ultra development machine completes in well under an hour with a warm target/. So this is about the runner and the clippy-then-test artifact split, not about the test count being unreasonable.
Defect 2: a timed-out run notifies nobody
The workflow has a Report a red main step whose purpose, stated in the file's own header comment, is that "a failed scheduled run files (or comments on) a GitHub issue", because "a red Actions run that nobody opens is the same blind spot in a new place".
In run 30733081983 that step's conclusion is skipped. Its condition matches failure, and a timeout surfaces as cancelled, so the notification path does not fire. Three silent timeouts in four days is precisely the blind spot the workflow was created to close.
Why this matters now
The workflow exists (see its header comment) because two deterministically failing tests reached main and sat there unnoticed, one of them for 26 days. #962 fixed the harness bug that was breaking it, and #953 created it. Neither is sufficient while the job cannot finish and cannot report.
Proposed Solution
These are options to evaluate and measure, not a decided design.
- Drop
cargo clippy from the same job, or reorder so the test build is not paying for clippy having warmed a different artifact set. Measure whether running cargo test --no-run first, or running clippy after test, changes the total.
- Split fmt/clippy and test into separate jobs, so a slow test build does not consume the whole budget and so the three signals stay independent (the file already states it wants them independent).
- Raise
timeout-minutes, but only with a measurement showing what a full pass actually costs. Raising it blindly trades a silent timeout for a longer silent timeout.
- Reduce link cost: fewer and larger integration-test binaries, or a
-C link-arg / linker choice, or cargo nextest, which builds the same binaries but may schedule them better. Consolidating test binaries is a real refactor with its own review cost.
- Fix the notification independently of the performance work, and first, since it is small: make the reporting step fire on
cancelled and on timeout, not only on failure. That alone converts a silent failure into a visible one and is worth landing even if the timeout takes longer to solve.
Acceptance Criteria
Refs
#953 created the workflow. #962 and #996 cover the harness bug that masked this. #997 is a flaky test found in the same suite.
Problem / Background
nightly-verifyhas two independent defects: itscargo teststep never finishes inside the job budget, and when the job is killed by that timeout nobody is told. All figures below were measured on 2026-08-02.Defect 1: the
cargo teststep never finishes inside the 180-minute budget.github/workflows/nightly-verify.ymlsetstimeout-minutes: 180. A job killed by that timeout is reported by GitHub with conclusioncancelled, notfailure, so in the run list it reads as "someone cancelled it". That is why it has gone unexamined.The last four runs of the workflow:
cli_help_consistency, issue #962Three of the last four hit exactly 3h00m. The only run that produced a verdict did so because
cargo testis fail-fast and it died early on the harness bug fixed in #996. The workflow has never completed a full pass.Per-step timings for run 30733081983, from
gh api repos/lablup/mlxcel/actions/runs/30733081983/jobs:The time goes to the build, not to the tests. The run log contains zero
Running tests/...lines and zerotest result:lines, and the processes the runner terminated at the timeout wereldandclang:The 67-second clippy is the load-bearing contrast: the persistent
CARGO_TARGET_DIRat$HOME/.cargo-target/mlxcelis warm. Butcargo clippyrunsclippy-driverand leaves metadata rather than linkable artifacts, socargo teststill has to codegen the library and link roughly 75 integration-test binaries, each against the large MLX static library, on a shared runner. That is where the three hours go.For calibration, the same suite on the M1 Ultra development machine completes in well under an hour with a warm
target/. So this is about the runner and the clippy-then-test artifact split, not about the test count being unreasonable.Defect 2: a timed-out run notifies nobody
The workflow has a
Report a red mainstep whose purpose, stated in the file's own header comment, is that "a failed scheduled run files (or comments on) a GitHub issue", because "a red Actions run that nobody opens is the same blind spot in a new place".In run 30733081983 that step's conclusion is
skipped. Its condition matches failure, and a timeout surfaces ascancelled, so the notification path does not fire. Three silent timeouts in four days is precisely the blind spot the workflow was created to close.Why this matters now
The workflow exists (see its header comment) because two deterministically failing tests reached
mainand sat there unnoticed, one of them for 26 days. #962 fixed the harness bug that was breaking it, and #953 created it. Neither is sufficient while the job cannot finish and cannot report.Proposed Solution
These are options to evaluate and measure, not a decided design.
cargo clippyfrom the same job, or reorder so the test build is not paying for clippy having warmed a different artifact set. Measure whether runningcargo test --no-runfirst, or running clippy after test, changes the total.timeout-minutes, but only with a measurement showing what a full pass actually costs. Raising it blindly trades a silent timeout for a longer silent timeout.-C link-arg/ linker choice, orcargo nextest, which builds the same binaries but may schedule them better. Consolidating test binaries is a real refactor with its own review cost.cancelledand on timeout, not only onfailure. That alone converts a silent failure into a visible one and is worth landing even if the timeout takes longer to solve.Acceptance Criteria
nightly-verifycompletes a full pass within its budget on the self-hosted runner, demonstrated by at least one green run, with the measured duration recorded in the PR.Refs
#953 created the workflow. #962 and #996 cover the harness bug that masked this. #997 is a flaky test found in the same suite.