Skip to content

build: add a fast test profile to cut edit-test iteration - #812

Merged
inureyes merged 2 commits into
mainfrom
build/issue-809-test-fast-profile
Jul 17, 2026
Merged

inureyes merged 2 commits into
mainfrom
build/issue-809-test-fast-profile

Conversation

@inureyes

Copy link
Copy Markdown
Member

Summary

Add [profile.test-fast] to Cargo.toml plus Makefile and doc support so local and agent edit-test iteration no longer pays the [profile.release] fat-LTO + codegen-units = 1 cost, which is measured at 4 to 6 minutes per incremental rebuild of the ~390k-line main crate today.

What changed

  • Cargo.toml: new [profile.test-fast] (inherits = "release", lto = "thin", codegen-units = 16, strip = false, incremental = true, opt-level = 3 kept), with a comment explaining why it exists and when not to use it, mirroring the [profile.release] comment style.
  • Makefile: new test-fast (platform-aware, mirrors release/RELEASE_FEATURE_FLAG), test-fast-cuda (mirrors release-cuda, explicit --features cuda), and check-fast targets, all wrapping cargo test/cargo check --profile test-fast; a FILTER variable for narrowing runs; a make help example line.
  • docs/installation.md: new "Fast iteration builds" subsection explaining the tradeoff and giving make test-fast / make test-fast-cuda / FILTER= usage.
  • CONTRIBUTING.md: points contributors at test-fast for iteration in the existing build/test step, while keeping the --release commands as the pre-PR gate.

Why lto = "thin" and not false

The workspace is split across crates (mlxcel, mlxcel-core, mlxcel-surgery), and the hot MLX FFI call sites cross that mlxcel / mlxcel-core boundary. lto = "thin" keeps cross-crate inlining there while dropping the whole-program fat-LTO link that dominates release build time; lto = false would wall off that boundary from inlining entirely. The measured incremental number below already lands well inside the "seconds, not minutes" target with thin LTO, so there was no measured case for dropping further to false and risking a numerics regression in MLX-heavy tests.

Measured timings (this session, Linux/CUDA, CARGO_TARGET_DIR pointed at the shared target dir)

  • Cold build (cargo test --profile test-fast --features cuda --lib -- models::minimax_m3 server::chat_request, all ~439 deps plus the MLX CUDA C++ tree compiled under the new profile for the first time): 293s (4m53s) wall clock, 343 crates compiled.
  • Incremental rebuild (touch src/server/chat_request.rs, same narrow test target): 19s wall clock (17.46s of that is the cargo compile step itself).
  • Baseline for comparison: today's [profile.release] incremental rebuild of the same crate is measured at 4 to 6 minutes (240 to 360s).
  • Incremental speedup: roughly 13x to 19x, well above the 3 to 5x expected in the issue.

Parity check

Ran the representative narrow test set from the issue under test-fast and compared against today's release-profile pass counts (recorded from today's implementation chain runs on the same code at main 80a05f9, not rerun here to avoid clobbering a busy shared build directory):

  • cargo test --profile test-fast --features cuda --lib -- models::minimax_m3 server::chat_request: 78 passed, 0 failed (10 minimax_m3 + 68 chat_request), matching the release-profile counts.
  • cargo test --profile test-fast --features cuda -p mlxcel-core -- sampling cache::ring: 61 passed, 0 failed (57 sampling + 4 cache::ring), matching the release-profile counts.

Test plan

  • cargo test --profile test-fast --features cuda --lib -- models::minimax_m3 server::chat_request (78 passed, 0 failed)
  • cargo test --profile test-fast --features cuda -p mlxcel-core -- sampling cache::ring (61 passed, 0 failed)
  • Cold and incremental timings recorded above
  • No .rs files touched, so cargo fmt --all -- --check is not applicable to this change

Closes #809

[profile.release] is tuned for shipping binaries (fat LTO, codegen-units = 1, strip = true), which makes every incremental rebuild of the ~390k-line main crate take 4 to 6 minutes measured on the 2026-07-17 implementation chain, since fat LTO forces a whole-program link across all ~439 locked crates and codegen-units = 1 serializes the crate's own codegen onto one core. A typical issue cycle pays 5 to 8 such rebuilds.

Add [profile.test-fast] (inherits release, lto = "thin", codegen-units = 16, strip = false, incremental = true, opt-level = 3 kept) for local and agent edit-test iteration. Thin LTO was chosen over disabling LTO entirely because the workspace is split across crates (mlxcel, mlxcel-core, mlxcel-surgery) and the hot MLX FFI call sites cross that boundary; thin LTO keeps cross-crate inlining there while dropping the fat-LTO whole-program link that dominates release build time.

Add Makefile test-fast (platform-aware like release/release-cuda) and test-fast-cuda targets wrapping cargo test --profile test-fast, plus check-fast, a FILTER variable for narrowing runs, and a help-text example. Document the profile in docs/installation.md ("Fast iteration builds") and point to it from CONTRIBUTING.md's build/test section, explicit that it is for iteration only and not for anything shipped, benchmarked, or quoted as representative performance.

Measured on this change: a cold build (all ~439 deps plus the MLX CUDA C++ tree under the new profile for the first time) completed in 293s; touching src/server/chat_request.rs and rebuilding the same narrow test target (models::minimax_m3, server::chat_request) took 19s, versus the existing 4 to 6 minute release-profile baseline, a roughly 13 to 19x incremental speedup. Parity check under test-fast: models::minimax_m3 (10) + server::chat_request (68) = 78 passed, 0 failed; mlxcel-core sampling (57) + cache::ring (4) = 61 passed, 0 failed. Both match the release-profile pass counts from today's chain runs on the same code.

Closes #809
@inureyes inureyes added type:enhancement New features, capabilities, or significant additions priority:medium Medium priority area:core mlxcel-core: MLX FFI, primitives, KV cache, layers status:review Under review labels Jul 17, 2026
@inureyes inureyes added status:done Completed and removed status:review Under review labels Jul 17, 2026
@inureyes
inureyes merged commit 653bb1c into main Jul 17, 2026
5 checks passed
@inureyes
inureyes deleted the build/issue-809-test-fast-profile branch August 4, 2026 12:07
@inureyes inureyes self-assigned this Aug 31, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:core mlxcel-core: MLX FFI, primitives, KV cache, layers priority:medium Medium priority status:done Completed type:enhancement New features, capabilities, or significant additions

Projects

None yet

Development

Successfully merging this pull request may close these issues.

build: add a fast test profile to cut edit-test iteration from minutes to seconds

1 participant