Skip to content

fix(ci): run the Metal test gate single threaded, like the CUDA one - #1210

Merged
inureyes merged 1 commit into
mainfrom
fix/issue-1092-serialize-metal-test-gate
Aug 18, 2026
Merged

inureyes merged 1 commit into
mainfrom
fix/issue-1092-serialize-metal-test-gate

Conversation

@inureyes

Copy link
Copy Markdown
Member

Summary

The macOS merge gate ran its whole test suite with libtest's default thread
count, one thread per logical CPU, while the CUDA gate has been serialized
since #1048. That is what took main red on 2026-08-16: the mlxcel-core
binary died with signal: 11, SIGSEGV, published no panic and no
test result line, and cargo reported a failed target with nothing to read.
This adds --test-threads=1 to verify-test and documents the measurement
behind it.

Related issues

Closes #1092

Type of change

  • fix — bug fix
  • chore — build, CI, dependencies, release infrastructure
  • docs — documentation only

What the evidence says

--jobs 1 is not the lever, and the reasoning that pointed at it was
wrong.
The issue thread proposed measuring --jobs 1 next, on the reading
that cargo runs several test binaries concurrently. It does not. Measured on
cargo 1.97.1 against a three-crate probe workspace whose tests sleep and print
timestamps: the build completes in full before the first test binary starts
(a 5s build script in crate c finishes 446ms before crate a's test begins),
and each binary finishes before the next begins (a ends at t+3.010s, b
starts at t+3.339s). --jobs bounds the build, which is over by the time any
test runs. The Makefile's existing claim that cargo "runs them one at a time"
was the correct one.

The concurrency is inside a single binary. The macOS crash report for the
local repro (mlxcel_core-7a69ce25e4ef37da-2026-08-17-154444.ips, UUID matched
against the binary) has 18 named libtest worker threads live at the fault on an
18-core M5 Max, which is exactly hw.logicalcpu. Every one of them is in an
MLX-backed cache test: 10 in cache::paged_batch_decode, 5 in
cache::paged_detach, 2 in cache::detach, 1 in autotune. Two are inside
iokit_user_client_trap, two inside the allocator. The fault is
EXC_BAD_ACCESS / KERN_INVALID_ADDRESS at an address in no mapped region.

That is the shape #1048 already documented on CUDA, where the fix was
--test-threads=1. The crashing tests are not new (they landed in #988 and
#1004), so this is a probability, not a regression.

Cost

Measured on M5 Max at 5dfcb390, [profile.test-fast], whole workspace,
101 binaries, 8128 tests, both arms warm:

test threads wall clock
default (18) 69.17s
--test-threads=1 76.39s

+7.2s, on a cargo test step the nightly budgets 180 minutes for and whose
time goes to the build rather than to running tests.

It is that cheap because the work already serializes on the one Metal device.
The two large members pull in opposite directions and nearly cancel:
mlxcel-core costs +23s serialized (10.2s to 33.2s) while the root suite
gains 12s (23.5s to 11.6s), thread contention across 5695 tests being worse
than running them in a row.

One trap worth recording, because it inverted the first measurement: a cold
first run pays roughly 50s of one-time Metal shader compilation. Run
parallel, serial, parallel and compare only the warm arms.

make test-fast has passed --test-threads=1 on macOS since #809, so the gate
now agrees with the edit-test loop rather than diverging from it.

What this deliberately does not add

No macOS counterpart to the the_cuda_test_suite_must_run_single_threaded
guard. The CUDA suite aborts on every parallel run, so failing by name costs
nothing. The Metal suite crashes rarely, and a hard guard would break
cargo test -p mlxcel-core --lib, which is three times faster parallel and
nearly always succeeds. Narrowed hand-runs are meant to stay parallel.

Test plan

  • make verify-test on M5 Max: 74.41s, flag applied, same verdict set as
    before the change.
  • Three A/B/A workspace runs, failure sets compared. The two warm runs
    agree on 17 failures; those 17 are this machine's pre-existing local reds
    (13 in the root lib, 4 numeric-tolerance failures in mlxcel-core), and
    the nightly runner is green on all of them.
  • make help renders the changed target line.
  • Real-checkpoint validation: not applicable, no runtime code changes.

This cannot prove the SIGSEGV will not recur, only that the concurrency it
needs is gone. The next nightly is the check.

Not fixed here

  • tests/granite4_vision_parity.rs::text_only_forward_produces_finite_logits
    fails 6/40 parallel and 4/40 serialized against a real
    granite-4.0-3b-vision-4bit checkpoint. Independent of thread count, so it
    is a separate defect and gets its own issue. It does not reach the nightly,
    which carries no weights.
  • Whether driving MLX eval from several host threads is a hazard for the
    server worker paths as well. Serializing the tests does not answer that.

The 2026-08-16 nightly died with `signal: 11, SIGSEGV` in the mlxcel-core
test binary. It published no panic and no `test result` line, so
`--no-fail-fast` had nothing to collect and cargo reported a failed
target with nothing to read.

#1092 proposed measuring `--jobs 1` next, on the reading that cargo runs
several test binaries concurrently. It does not. Measured on cargo
1.97.1 with a three-crate probe workspace, the build completes in full
before the first test binary starts and each binary finishes before the
next begins, so `--jobs` bounds only the build and can have no effect at
the time a test is running.

The concurrency is inside one binary. libtest defaults to one test
thread per logical CPU, and the macOS crash report from the local repro
on an 18-core M5 Max has 18 libtest workers live at the fault, every one
of them in an MLX-backed cache test, two inside `iokit_user_client_trap`
and two inside the allocator, faulting on an address in no mapped
region. That is the CUDA abort of #1048 on the other backend, and
`verify-test-cuda` already carries the flag that bounds it.

Serializing is close to free, because the work serializes on the one
Metal device whether or not the host threads do. On an M5 Max at
5dfcb39, warm, whole workspace, 101 binaries and 8128 tests: 69.17s
parallel against 76.39s serialized. The two large members nearly cancel,
mlxcel-core costing +23s while the root suite gains 12s. The nightly
budgets 180 minutes for a step that spends its time in the build.

`make test-fast` has passed `--test-threads=1` on macOS since #809, so
the gate now agrees with the edit-test loop rather than diverging from
it. No macOS counterpart to the CUDA guard test: that suite aborts every
parallel run, while this one crashes rarely, and a hard guard would
break `cargo test -p mlxcel-core --lib`, which is three times faster
parallel and nearly always succeeds.

Refs #1092
@inureyes
inureyes merged commit b32059c into main Aug 18, 2026
8 checks passed
@inureyes
inureyes deleted the fix/issue-1092-serialize-metal-test-gate branch August 18, 2026 08:45
@inureyes inureyes self-assigned this Aug 27, 2026
@inureyes inureyes added status:done Completed type:bug Bug fixes, error corrections, or issue resolutions priority:high High priority labels Sep 6, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

priority:high High priority status:done Completed type:bug Bug fixes, error corrections, or issue resolutions

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[nightly-verify] main is red

1 participant