Release JXR 0.2.0 with CPU and Metal improvements - #10
Merged
Merged
Conversation
PacketBitReader::read_bits consumed one bit per loop iteration with a bounds-checked byte index, and the three prefix-code decoders (entropy VLCs, CBPHP, and YUV DC/LP patterns) read one bit at a time while linearly searching their code tables and recomputing the maximum code length on every call. Reads of up to 57 bits now take one big-endian 64-bit window load, with the end-of-buffer padding and error construction moved into cold out-of-line paths so the hot path inlines. Every normative prefix table is compiled into an eight-bit lookup; construction is const, so an ambiguous or oversized table fails compilation. Error behavior is unchanged: truncated packets report UnexpectedEnd at the bit where the serial decoder stopped, and unmatched full-length prefixes report InvalidVlc at the code start. Exhaustive tests compare the lookup decoders with the serial search for every 16-bit input, offset, and truncation, and the reader with a bit-serial reference for every width and offset.
The vectorized U8 packer and HP coefficient scaler each validated their input in a separate full pass before the vector loop: the packer re-read every sample to check the output-bias addition, and the scaler widened every coefficient to i64 to check the product. Both checks now run inside the existing vector loop by tracking the observed sample range with fearless_simd 1.0's lane max/min reductions, then comparing it with the representable range once. Dequantization compares against [i32::MIN / step, i32::MAX / step], which is exact for every nonzero step. The packer bounds only its final row, since row starts increase monotonically. Out-of-range input still returns the same error; any wrapped values already written are discarded with the failed call. Scalar paths saturate or wrap so debug builds cannot overflow before the check. Microbenchmarks on an M4 Pro (NEON): 256-coefficient dequantization 96.4 -> 21.0 ns per call, and 1-megapixel luma U8 packing 361 -> 133 us.
The generic CPU packers interpreted the output format per sample: every sample re-derived its bias, rounding, and postscaling from the request, every pixel re-converted the crop origin and re-matched the channel layout, and every pixel copied its primaries with a variable-length memmove. The packed RGB555/RGB565/RGB101010/RGBE path repeated the same per-sample scale resolution. scale_integer_component is now "resolve a ChannelScale, then apply it", and the ordered and packed-color packers resolve each channel's scale once per image through that same implementation, so the fast path and the reference cannot diverge. Resolution failures are kept and reported by the first sample that uses the channel, preserving the previous error behavior. Crop conversion and layout decisions are hoisted out of the pixel loops, and primaries are copied with fixed-size moves.
Every reconstruction kernel widened each add, subtract, and multiply to 64-bit, which Apple GPUs emulate, and returned through a branch after each operation. Arithmetic now uses exact 32-bit overflow predicates (sign-bit tests for add/sub, mulhi for quantizer products, range tests for the times-three steps) accumulated into a sticky flag that each kernel tests once before storing. The transforms become straight-line code; per-phase status codes are unchanged, and results derived from an overflowed intermediate are never stored. The HP kernel ran one threadgroup per macroblock with thread 0 applying HP prediction serially between two barriers, leaving half of each SIMD group idle for 16-block luma macroblocks. It now runs one thread per 4x4 block in a flat grid; each thread accumulates its own prediction chain in the normative order, so every partial-sum overflow check matches the serial traversal. Block rows are stored as aligned int4 writes, and plane ABI construction now rejects sample planes that are not four-sample aligned. The output kernels converted color, including chroma upsampling, once per output channel. Each pixel is now loaded and converted once, and premultiplied stores scale alpha once per pixel. Chroma upsampling computes the weighted average exactly in 32 bits by splitting each operand into 8q + r, and unsigned premultiplication uses 32-bit division (65535^2 + 32767 fits in u32).
Overlap schedules were rebuilt on the CPU and uploaded into a newly allocated MTLBuffer for every image and plane on every submission, and their work items baked in absolute sample offsets. Schedules are now built plane-relative, uploaded once per plane geometry into a bounded LRU cache on the runtime, and rebased in the kernel with a per-dispatch base offset. The host still rejects any rebased index beyond the u32 device ABI. Small output descriptor arrays are passed with setBytes instead of a shared-buffer allocation per image. Dense batches encoded every image into its own command buffer on one queue; because all of them write one tracked allocation, Metal serialized them. Dense and caller-destination batches now group up to 16 images, bounded by the batch scratch budget, into concatenated- descriptor command buffers. Batch outputs carry byte offsets for this. Host decodes of 16- and 32-bit formats copied the shared output into a Vec<u8> and then into the typed vector; they now convert once from the mapped allocation. Resident readback reuses one lazily created queue per session instead of creating a command queue per call. JxrDecoder::decode reuses its routing plan for Metal and CUDA preparation instead of planning the request twice. With these host costs removed, eight images per concurrent batch command measured higher 128-tile pipeline throughput than two on the 16-core M4 Pro, so the batch width is now eight.
The previous commit unintentionally included workspace crate version bumps to 0.2.0 and the matching lockfile entries. Version changes belong to a release PR, so this restores every manifest and Cargo.lock to main.
jxr-core re-exports j2k_core::{BackendKind, BackendRequest, Rect}, so
the J2K 0.10 -> 0.11 upgrade changes public types. Every published crate
that exposes them moves to 0.2.0; jxr-math stays at 0.1.0.
docs/releasing.md: restore the published 0.1.2 record (jxr-mpsgraph 0.1.1
shipped with j2k-mpsgraph-support =0.11.0) and add the 0.2.0 section.
… integration/jxr-0.2.0
…tegration/jxr-0.2.0
…tegration/jxr-0.2.0
…o integration/jxr-0.2.0
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Release JXR 0.2.0 with the J2K 0.11 public-type upgrade and the integrated CPU entropy decoding, SIMD output packing, Metal kernel, and host-submission improvements from #5–#9. All seven affected public crates move to 0.2.0; jxr-math remains at its published 0.1.0 version.
Pin j2k-core, j2k-metal-support, and j2k-mpsgraph-support to the published 0.11.2 release used by wsi-rs. Both the workspace and fuzz lockfiles resolve J2K from crates.io, without sibling-checkout overrides.
Package contents have been reviewed and the jxr-core publication dry run passed. Hosted CI and CPU/Metal hardware validation run on this PR; CUDA hardware validation runs from main after merge. Publish in the dependency order documented in docs/releasing.md only after the release checks pass.