Skip to content

Make quantization independent of the architecture (fixed histogram order, consistent NEON distance) - #131

Open
dustinkirkland wants to merge 1 commit into
ImageOptim:mainfrom
dustinkirkland:deterministic-histogram-order
Open

dustinkirkland wants to merge 1 commit into
ImageOptim:mainfrom
dustinkirkland:deterministic-histogram-order

Conversation

@dustinkirkland

Copy link
Copy Markdown

Fixes #130.

The same image quantizes to a different palette on x86_64 and aarch64 (details and measurements in #130). Two causes, three small changes:

  1. Histogram built in a fixed order. Histogram::finalize_builder drained the colour HashMap in iteration order, and everything after that is order-sensitive (cluster bucketing by position, the f64 weight sum, median cut's tie-breaking, k-means' 256-item chunks). The hasher is deterministic, but hashbrown's probe-group width is 16 bytes with SSE2 on x86_64 and 8 bytes with NEON on aarch64, so colliding keys land in different buckets and the map iterates in a different order per architecture. Sort temp by colour before the cluster pass. One sort of at most max_histogram_entries items.
  2. Posterize rehash in key order. init_posterize_bits re-inserts the old map with coarser keys; when several keys collapse, the last one inserted wins, which again depended on iteration order. Collect, sort by key, then extend. Last-wins semantics are kept (summing the boosts of collapsed entries would arguably be better, but that is a behaviour change and left for you to decide).
  3. NEON f_pixel::diff adds in the same order as the other paths. The x86_64, scalar and WASM paths compute (r + g) + b; the NEON path used vpaddq_f32 and computed r + (g + b). Float addition is not associative, and the refinement loop amplifies the ulp differences in the metric into a different palette for some images (visible with --quality 0-100). Use the same association; it is also one instruction fewer.

Verification (pngquant 3.0.3 rebuilt with these changes, x86_64 native and aarch64 under qemu-user, comparing decoded pixels and palette sizes): before, 4 of 5 noto-emoji sample bitmaps differed between the architectures at --quality 85-95 and one more at --quality 0-100; after, all 11 cases (5 images × 2 quality settings, plus a 262,144-colour image at --speed 10 that goes through the posterize rehash) are identical on both. Each change was isolated: sorting the histogram alone fixed the 85-95 cases (and made the x86_64 output of the UN flag byte-for-pixel the previous aarch64 output); the NEON association alone explained the remaining 0-100 case (an x86_64 build with NEON's order reproduced the aarch64 output exactly). cargo test passes on both.

Output on a given architecture changes for images where either effect mattered: same quality target and algorithm, a fixed tie order and a consistent metric.

The same image quantized to a different palette on x86_64 and aarch64
(ImageOptim#130). Two causes:

- The histogram was drained from a HashMap in iteration order, and the
  cluster bucketing, the f64 weight sum, median cut tie-breaking and
  k-means chunking are all order-sensitive. The hasher is deterministic,
  but hashbrown probes 16-byte groups with SSE2 on x86_64 and 8-byte
  groups with NEON on aarch64, so the map iterates in a different order
  per architecture. Sort the entries by colour before the cluster pass,
  and re-insert the posterized entries in key order (last-wins on
  collapsing keys otherwise depended on the same order).

- The NEON f_pixel::diff summed the channel terms as r + (g + b) while
  the x86_64, scalar and WASM paths sum (r + g) + b; float addition is
  not associative, and the refinement loop amplifies the difference in
  the metric. Add in the same order.

With both, x86_64 and aarch64 produce identical pixels and palettes for
every image tested at --quality 85-95 and 0-100, including one that goes
through the posterize rehash. Output on one architecture changes for
images where either effect mattered: same quality target and algorithm,
a fixed tie order and a consistent metric.

Fixes ImageOptim#130.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Quantization result depends on HashMap iteration order, so the same image quantizes differently on x86_64 and aarch64

1 participant