fix(ci): run the Metal test gate single threaded, like the CUDA one - #1210
Merged
Merged
Conversation
The 2026-08-16 nightly died with `signal: 11, SIGSEGV` in the mlxcel-core test binary. It published no panic and no `test result` line, so `--no-fail-fast` had nothing to collect and cargo reported a failed target with nothing to read. #1092 proposed measuring `--jobs 1` next, on the reading that cargo runs several test binaries concurrently. It does not. Measured on cargo 1.97.1 with a three-crate probe workspace, the build completes in full before the first test binary starts and each binary finishes before the next begins, so `--jobs` bounds only the build and can have no effect at the time a test is running. The concurrency is inside one binary. libtest defaults to one test thread per logical CPU, and the macOS crash report from the local repro on an 18-core M5 Max has 18 libtest workers live at the fault, every one of them in an MLX-backed cache test, two inside `iokit_user_client_trap` and two inside the allocator, faulting on an address in no mapped region. That is the CUDA abort of #1048 on the other backend, and `verify-test-cuda` already carries the flag that bounds it. Serializing is close to free, because the work serializes on the one Metal device whether or not the host threads do. On an M5 Max at 5dfcb39, warm, whole workspace, 101 binaries and 8128 tests: 69.17s parallel against 76.39s serialized. The two large members nearly cancel, mlxcel-core costing +23s while the root suite gains 12s. The nightly budgets 180 minutes for a step that spends its time in the build. `make test-fast` has passed `--test-threads=1` on macOS since #809, so the gate now agrees with the edit-test loop rather than diverging from it. No macOS counterpart to the CUDA guard test: that suite aborts every parallel run, while this one crashes rarely, and a hard guard would break `cargo test -p mlxcel-core --lib`, which is three times faster parallel and nearly always succeeds. Refs #1092
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
The macOS merge gate ran its whole test suite with libtest's default thread
count, one thread per logical CPU, while the CUDA gate has been serialized
since #1048. That is what took
mainred on 2026-08-16: themlxcel-corebinary died with
signal: 11, SIGSEGV, published no panic and notest resultline, and cargo reported a failed target with nothing to read.This adds
--test-threads=1toverify-testand documents the measurementbehind it.
Related issues
Closes #1092
Type of change
fix— bug fixchore— build, CI, dependencies, release infrastructuredocs— documentation onlyWhat the evidence says
--jobs 1is not the lever, and the reasoning that pointed at it waswrong. The issue thread proposed measuring
--jobs 1next, on the readingthat cargo runs several test binaries concurrently. It does not. Measured on
cargo 1.97.1 against a three-crate probe workspace whose tests sleep and print
timestamps: the build completes in full before the first test binary starts
(a 5s build script in crate
cfinishes 446ms before cratea's test begins),and each binary finishes before the next begins (
aends at t+3.010s,bstarts at t+3.339s).
--jobsbounds the build, which is over by the time anytest runs. The Makefile's existing claim that cargo "runs them one at a time"
was the correct one.
The concurrency is inside a single binary. The macOS crash report for the
local repro (
mlxcel_core-7a69ce25e4ef37da-2026-08-17-154444.ips, UUID matchedagainst the binary) has 18 named libtest worker threads live at the fault on an
18-core M5 Max, which is exactly
hw.logicalcpu. Every one of them is in anMLX-backed cache test: 10 in
cache::paged_batch_decode, 5 incache::paged_detach, 2 incache::detach, 1 inautotune. Two are insideiokit_user_client_trap, two inside the allocator. The fault isEXC_BAD_ACCESS/KERN_INVALID_ADDRESSat an address in no mapped region.That is the shape #1048 already documented on CUDA, where the fix was
--test-threads=1. The crashing tests are not new (they landed in #988 and#1004), so this is a probability, not a regression.
Cost
Measured on M5 Max at
5dfcb390,[profile.test-fast], whole workspace,101 binaries, 8128 tests, both arms warm:
--test-threads=1+7.2s, on a
cargo teststep the nightly budgets 180 minutes for and whosetime goes to the build rather than to running tests.
It is that cheap because the work already serializes on the one Metal device.
The two large members pull in opposite directions and nearly cancel:
mlxcel-corecosts +23s serialized (10.2s to 33.2s) while the root suitegains 12s (23.5s to 11.6s), thread contention across 5695 tests being worse
than running them in a row.
One trap worth recording, because it inverted the first measurement: a cold
first run pays roughly 50s of one-time Metal shader compilation. Run
parallel, serial, paralleland compare only the warm arms.make test-fasthas passed--test-threads=1on macOS since #809, so the gatenow agrees with the edit-test loop rather than diverging from it.
What this deliberately does not add
No macOS counterpart to the
the_cuda_test_suite_must_run_single_threadedguard. The CUDA suite aborts on every parallel run, so failing by name costs
nothing. The Metal suite crashes rarely, and a hard guard would break
cargo test -p mlxcel-core --lib, which is three times faster parallel andnearly always succeeds. Narrowed hand-runs are meant to stay parallel.
Test plan
make verify-teston M5 Max: 74.41s, flag applied, same verdict set asbefore the change.
agree on 17 failures; those 17 are this machine's pre-existing local reds
(13 in the root lib, 4 numeric-tolerance failures in
mlxcel-core), andthe nightly runner is green on all of them.
make helprenders the changed target line.This cannot prove the SIGSEGV will not recur, only that the concurrency it
needs is gone. The next nightly is the check.
Not fixed here
tests/granite4_vision_parity.rs::text_only_forward_produces_finite_logitsfails 6/40 parallel and 4/40 serialized against a real
granite-4.0-3b-vision-4bitcheckpoint. Independent of thread count, so itis a separate defect and gets its own issue. It does not reach the nightly,
which carries no weights.
server worker paths as well. Serializing the tests does not answer that.