Summary
[profile.release] in Cargo.toml is tuned for shipping binaries (fat lto = true, codegen-units = 1, opt-level = 3, strip = true). All local and agent test iteration currently runs through this profile because target/debug is usually cold (the MLX C++ tree rebuild makes the debug profile expensive to warm). Add a dedicated fast profile for test/dev iteration.
Motivation (measured)
Measured facts from the 2026-07-17 implementation chain:
- The main crate is a single ~390k-line crate, so any source edit recompiles the whole crate on one core (
codegen-units = 1 blocks parallel codegen), followed by a fat-LTO link across ~439 locked crates. Measured at 4 to 6 minutes per incremental rebuild.
- A typical issue cycle (developer edit-test iterations plus review-stage fixes) pays 5 to 8 such rebuilds, i.e. 25 to 45 minutes of pure compile time per small issue.
- All-targets test runs multiply the LTO link per test binary.
Implementation plan
- Add
[profile.test-fast] to Cargo.toml:
inherits = "release"
lto = "thin" (or false if thin-LTO still dominates link time; measure both)
codegen-units = 16
strip = false
incremental = true
- keep
opt-level = 3 (MLX-heavy tests need optimized numerics for tolerable runtimes; drop to 2 only if measurements show parity)
- Add a Makefile target
test-fast (and optionally check-fast) wrapping cargo test --profile test-fast --features cuda, with a narrow-filter example in the help text.
- Document in the contributor/build docs which profile to use for iteration vs shipping.
- Parity check: run a representative narrow test set (e.g.
models::minimax_m3, server::chat_request, mlxcel-core sampling and cache::ring) under both profiles and confirm identical pass/fail results; record the measured rebuild-time comparison (cold and incremental) in the PR body.
Acceptance criteria
Note: repository CI currently runs no test job, so this is purely a local/agent iteration improvement with no CI changes required.
Summary
[profile.release]in Cargo.toml is tuned for shipping binaries (fatlto = true,codegen-units = 1,opt-level = 3,strip = true). All local and agent test iteration currently runs through this profile becausetarget/debugis usually cold (the MLX C++ tree rebuild makes the debug profile expensive to warm). Add a dedicated fast profile for test/dev iteration.Motivation (measured)
Measured facts from the 2026-07-17 implementation chain:
codegen-units = 1blocks parallel codegen), followed by a fat-LTO link across ~439 locked crates. Measured at 4 to 6 minutes per incremental rebuild.Implementation plan
[profile.test-fast]to Cargo.toml:inherits = "release"lto = "thin"(orfalseif thin-LTO still dominates link time; measure both)codegen-units = 16strip = falseincremental = trueopt-level = 3(MLX-heavy tests need optimized numerics for tolerable runtimes; drop to 2 only if measurements show parity)test-fast(and optionallycheck-fast) wrappingcargo test --profile test-fast --features cuda, with a narrow-filter example in the help text.models::minimax_m3,server::chat_request, mlxcel-core sampling andcache::ring) under both profiles and confirm identical pass/fail results; record the measured rebuild-time comparison (cold and incremental) in the PR body.Acceptance criteria
[profile.test-fast]exists in Cargo.toml and builds with--features cudatest-fasttarget exists and runs tests under the new profileNote: repository CI currently runs no test job, so this is purely a local/agent iteration improvement with no CI changes required.