Repository navigation
build: add a fast test profile to cut edit-test iteration - #812
Merged
Merged
Conversation
[profile.release] is tuned for shipping binaries (fat LTO, codegen-units = 1, strip = true), which makes every incremental rebuild of the ~390k-line main crate take 4 to 6 minutes measured on the 2026-07-17 implementation chain, since fat LTO forces a whole-program link across all ~439 locked crates and codegen-units = 1 serializes the crate's own codegen onto one core. A typical issue cycle pays 5 to 8 such rebuilds.
Add [profile.test-fast] (inherits release, lto = "thin", codegen-units = 16, strip = false, incremental = true, opt-level = 3 kept) for local and agent edit-test iteration. Thin LTO was chosen over disabling LTO entirely because the workspace is split across crates (mlxcel, mlxcel-core, mlxcel-surgery) and the hot MLX FFI call sites cross that boundary; thin LTO keeps cross-crate inlining there while dropping the fat-LTO whole-program link that dominates release build time.
Add Makefile test-fast (platform-aware like release/release-cuda) and test-fast-cuda targets wrapping cargo test --profile test-fast, plus check-fast, a FILTER variable for narrowing runs, and a help-text example. Document the profile in docs/installation.md ("Fast iteration builds") and point to it from CONTRIBUTING.md's build/test section, explicit that it is for iteration only and not for anything shipped, benchmarked, or quoted as representative performance.
Measured on this change: a cold build (all ~439 deps plus the MLX CUDA C++ tree under the new profile for the first time) completed in 293s; touching src/server/chat_request.rs and rebuilding the same narrow test target (models::minimax_m3, server::chat_request) took 19s, versus the existing 4 to 6 minute release-profile baseline, a roughly 13 to 19x incremental speedup. Parity check under test-fast: models::minimax_m3 (10) + server::chat_request (68) = 78 passed, 0 failed; mlxcel-core sampling (57) + cache::ring (4) = 61 passed, 0 failed. Both match the release-profile pass counts from today's chain runs on the same code.
Closes #809
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Add
[profile.test-fast]to Cargo.toml plus Makefile and doc support so local and agent edit-test iteration no longer pays the[profile.release]fat-LTO +codegen-units = 1cost, which is measured at 4 to 6 minutes per incremental rebuild of the ~390k-line main crate today.What changed
Cargo.toml: new[profile.test-fast](inherits = "release",lto = "thin",codegen-units = 16,strip = false,incremental = true,opt-level = 3kept), with a comment explaining why it exists and when not to use it, mirroring the[profile.release]comment style.Makefile: newtest-fast(platform-aware, mirrorsrelease/RELEASE_FEATURE_FLAG),test-fast-cuda(mirrorsrelease-cuda, explicit--features cuda), andcheck-fasttargets, all wrappingcargo test/cargo check --profile test-fast; aFILTERvariable for narrowing runs; amake helpexample line.docs/installation.md: new "Fast iteration builds" subsection explaining the tradeoff and givingmake test-fast/make test-fast-cuda/FILTER=usage.CONTRIBUTING.md: points contributors attest-fastfor iteration in the existing build/test step, while keeping the--releasecommands as the pre-PR gate.Why
lto = "thin"and notfalseThe workspace is split across crates (
mlxcel,mlxcel-core,mlxcel-surgery), and the hot MLX FFI call sites cross thatmlxcel/mlxcel-coreboundary.lto = "thin"keeps cross-crate inlining there while dropping the whole-program fat-LTO link that dominates release build time;lto = falsewould wall off that boundary from inlining entirely. The measured incremental number below already lands well inside the "seconds, not minutes" target with thin LTO, so there was no measured case for dropping further tofalseand risking a numerics regression in MLX-heavy tests.Measured timings (this session, Linux/CUDA,
CARGO_TARGET_DIRpointed at the shared target dir)cargo test --profile test-fast --features cuda --lib -- models::minimax_m3 server::chat_request, all ~439 deps plus the MLX CUDA C++ tree compiled under the new profile for the first time): 293s (4m53s) wall clock, 343 crates compiled.touch src/server/chat_request.rs, same narrow test target): 19s wall clock (17.46s of that is thecargocompile step itself).[profile.release]incremental rebuild of the same crate is measured at 4 to 6 minutes (240 to 360s).Parity check
Ran the representative narrow test set from the issue under
test-fastand compared against today's release-profile pass counts (recorded from today's implementation chain runs on the same code atmain80a05f9, not rerun here to avoid clobbering a busy shared build directory):cargo test --profile test-fast --features cuda --lib -- models::minimax_m3 server::chat_request: 78 passed, 0 failed (10minimax_m3+ 68chat_request), matching the release-profile counts.cargo test --profile test-fast --features cuda -p mlxcel-core -- sampling cache::ring: 61 passed, 0 failed (57 sampling + 4cache::ring), matching the release-profile counts.Test plan
cargo test --profile test-fast --features cuda --lib -- models::minimax_m3 server::chat_request(78 passed, 0 failed)cargo test --profile test-fast --features cuda -p mlxcel-core -- sampling cache::ring(61 passed, 0 failed).rsfiles touched, socargo fmt --all -- --checkis not applicable to this changeCloses #809