Skip to content

Fuse SIMD overflow validation into the pack and dequant passes - #5

Merged
jcwal1516 merged 1 commit into
mainfrom
perf/simd-fused-validation
Sep 27, 2026
Merged

jcwal1516 merged 1 commit into
mainfrom
perf/simd-fused-validation

Conversation

@jcwal1516

Copy link
Copy Markdown
Member

The vectorized U8 packer and HP coefficient scaler from #3 each validated their input in a separate full pass before the vector loop: the packer re-read every sample to check the output-bias addition, and the scaler widened every coefficient to i64 to check the product. Both checks now run inside the existing vector loop by tracking the observed sample range with fearless_simd 1.0's reduce_max/reduce_min, then comparing it with the representable range once. Dequantization uses [i32::MIN / step, i32::MAX / step], which is exact for every nonzero step, and the packer bounds only its final row because row starts increase monotonically.

Out-of-range input still returns the same error. Wrapped values already written are discarded with the failed call, and the scalar paths saturate or wrap so debug builds cannot overflow before the check. New tests place boundary and out-of-range values in both the vector body and the scalar tail for steps up to u32::MAX.

Performance (M4 Pro, NEON, temporary microbenchmark not committed): 256-coefficient dequantization 96.4 → 21.0 ns per call; 1-megapixel luma U8 packing 361 → 133 µs. Full-file decode time is unchanged within noise (±5%), because entropy decoding dominates those files.

Validation: cargo test --workspace --all-features in debug and release (258 passed each); cargo clippy --workspace --all-targets --all-features -- -D warnings, plus cargo clippy --target x86_64-unknown-linux-gnu -p jxr-native --all-targets -- -D warnings for the AVX2 path; T.834/T.835 conformance CPU 517/517 and Metal 517/517. The AVX2 path was compiled and linted but not executed locally; the Ubuntu CI job runs it.

The vectorized U8 packer and HP coefficient scaler each validated their
input in a separate full pass before the vector loop: the packer
re-read every sample to check the output-bias addition, and the scaler
widened every coefficient to i64 to check the product.

Both checks now run inside the existing vector loop by tracking the
observed sample range with fearless_simd 1.0's lane max/min reductions,
then comparing it with the representable range once. Dequantization
compares against [i32::MIN / step, i32::MAX / step], which is exact for
every nonzero step. The packer bounds only its final row, since row
starts increase monotonically. Out-of-range input still returns the
same error; any wrapped values already written are discarded with the
failed call. Scalar paths saturate or wrap so debug builds cannot
overflow before the check.

Microbenchmarks on an M4 Pro (NEON): 256-coefficient dequantization
96.4 -> 21.0 ns per call, and 1-megapixel luma U8 packing 361 -> 133 us.
@jcwal1516
jcwal1516 merged commit 4d10f81 into main Sep 27, 2026
2 checks passed
@jcwal1516
jcwal1516 deleted the perf/simd-fused-validation branch September 28, 2026 06:05
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant