Skip to content

Quantization result depends on HashMap iteration order, so the same image quantizes differently on x86_64 and aarch64 #130

Description

@dustinkirkland

pngquant --speed 1 --quality 85-95 (and any speed ≤ 7) produces a different palette and different pixels for the same input on x86_64 and on aarch64, with the same pngquant/libimagequant version. It is fully deterministic on each architecture (run to run, and independent of RAYON_NUM_THREADS), just not across them.

Found while building Noto Color Emoji in a Linux distribution: 92 of the 4,004 CBDT bitmaps differed between the two architectures' builds.

Cause

Histogram::finalize_builder (src/hist.rs) turns the colour hash map into the item list with

temp.extend(self.hashmap.values().map(|&(boost, color)| { ... }));

and everything after that is order-sensitive: the cluster bucketing assigns items[next_index] in that order, total_perceptual_weight is an f64 sum in that order, median cut's select_nth_unstable_by_key / sort_unstable_by_key resolve ties by position, and k-means chunks the slice with par_chunks_mut(256). U32Hasher makes the hash deterministic, but the iteration order of a HashMap also depends on the table's probe-group width, which is 16 bytes with hashbrown's SSE2 group on x86_64 and 8 bytes with the NEON group on aarch64 (hashbrown/src/control/group/{sse2,neon}.rs). Colliding keys therefore land in different buckets on the two architectures, the map iterates in a different order, and the small order-dependent differences are amplified by the refinement iterations into a different palette.

Evidence

  • pngquant 3.0.3 packaged from source on both architectures, UN flag image from noto-emoji (136x128 RGBA): palette of 111 colours on x86_64, 120 on aarch64; other images 151 vs 142, 201 vs 219, 35 vs 37. Identical at --speed 8 and above (no k-means iterations, no feedback trials) and with a small fixed palette (pngquant 16).

  • Not the SIMD f_pixel::diff: changing the x86_64 horizontal add to NEON's association order (r + (g + b)) changes nothing.

  • Not libm: f32 powf/expf/logf fingerprints are identical on both.

  • Adding one line after the temp.extend(...):

    temp.sort_unstable_by_key(|t| (t.color.r, t.color.g, t.color.b, t.color.a, t.cluster_index));

    makes the x86_64 build produce, for the UN flag, exactly the pixels and 120-colour palette the unpatched aarch64 build produces, and the aarch64 build with the same one-line change produces the same pixels and palettes as that patched x86_64 build for every sample (UN flag, U+1F92B, U+1F307, U+1F9D4+1F3FE, U+1F600). So with a deterministic histogram order the two architectures agree exactly; without it they do not.

Suggested fix

Sort temp by colour (or iterate the map in key order) before the cluster pass, as above; it costs one sort of the histogram (at most max_histogram_entries items) and makes the output a pure function of the input across platforms. Happy to send that as a PR if you would take it.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions