Fuse SIMD overflow validation into the pack and dequant passes - #5
Merged
Merged
Conversation
The vectorized U8 packer and HP coefficient scaler each validated their input in a separate full pass before the vector loop: the packer re-read every sample to check the output-bias addition, and the scaler widened every coefficient to i64 to check the product. Both checks now run inside the existing vector loop by tracking the observed sample range with fearless_simd 1.0's lane max/min reductions, then comparing it with the representable range once. Dequantization compares against [i32::MIN / step, i32::MAX / step], which is exact for every nonzero step. The packer bounds only its final row, since row starts increase monotonically. Out-of-range input still returns the same error; any wrapped values already written are discarded with the failed call. Scalar paths saturate or wrap so debug builds cannot overflow before the check. Microbenchmarks on an M4 Pro (NEON): 256-coefficient dequantization 96.4 -> 21.0 ns per call, and 1-megapixel luma U8 packing 361 -> 133 us.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The vectorized U8 packer and HP coefficient scaler from #3 each validated their input in a separate full pass before the vector loop: the packer re-read every sample to check the output-bias addition, and the scaler widened every coefficient to
i64to check the product. Both checks now run inside the existing vector loop by tracking the observed sample range with fearless_simd 1.0'sreduce_max/reduce_min, then comparing it with the representable range once. Dequantization uses[i32::MIN / step, i32::MAX / step], which is exact for every nonzero step, and the packer bounds only its final row because row starts increase monotonically.Out-of-range input still returns the same error. Wrapped values already written are discarded with the failed call, and the scalar paths saturate or wrap so debug builds cannot overflow before the check. New tests place boundary and out-of-range values in both the vector body and the scalar tail for steps up to
u32::MAX.Performance (M4 Pro, NEON, temporary microbenchmark not committed): 256-coefficient dequantization 96.4 → 21.0 ns per call; 1-megapixel luma U8 packing 361 → 133 µs. Full-file decode time is unchanged within noise (±5%), because entropy decoding dominates those files.
Validation:
cargo test --workspace --all-featuresin debug and release (258 passed each);cargo clippy --workspace --all-targets --all-features -- -D warnings, pluscargo clippy --target x86_64-unknown-linux-gnu -p jxr-native --all-targets -- -D warningsfor the AVX2 path; T.834/T.835 conformance CPU 517/517 and Metal 517/517. The AVX2 path was compiled and linted but not executed locally; the Ubuntu CI job runs it.