pngquant --speed 1 --quality 85-95 (and any speed ≤ 7) produces a different palette and different pixels for the same input on x86_64 and on aarch64, with the same pngquant/libimagequant version. It is fully deterministic on each architecture (run to run, and independent of RAYON_NUM_THREADS), just not across them.
Found while building Noto Color Emoji in a Linux distribution: 92 of the 4,004 CBDT bitmaps differed between the two architectures' builds.
Cause
Histogram::finalize_builder (src/hist.rs) turns the colour hash map into the item list with
temp.extend(self.hashmap.values().map(|&(boost, color)| { ... }));
and everything after that is order-sensitive: the cluster bucketing assigns items[next_index] in that order, total_perceptual_weight is an f64 sum in that order, median cut's select_nth_unstable_by_key / sort_unstable_by_key resolve ties by position, and k-means chunks the slice with par_chunks_mut(256). U32Hasher makes the hash deterministic, but the iteration order of a HashMap also depends on the table's probe-group width, which is 16 bytes with hashbrown's SSE2 group on x86_64 and 8 bytes with the NEON group on aarch64 (hashbrown/src/control/group/{sse2,neon}.rs). Colliding keys therefore land in different buckets on the two architectures, the map iterates in a different order, and the small order-dependent differences are amplified by the refinement iterations into a different palette.
Evidence
-
pngquant 3.0.3 packaged from source on both architectures, UN flag image from noto-emoji (136x128 RGBA): palette of 111 colours on x86_64, 120 on aarch64; other images 151 vs 142, 201 vs 219, 35 vs 37. Identical at --speed 8 and above (no k-means iterations, no feedback trials) and with a small fixed palette (pngquant 16).
-
Not the SIMD f_pixel::diff: changing the x86_64 horizontal add to NEON's association order (r + (g + b)) changes nothing.
-
Not libm: f32 powf/expf/logf fingerprints are identical on both.
-
Adding one line after the temp.extend(...):
temp.sort_unstable_by_key(|t| (t.color.r, t.color.g, t.color.b, t.color.a, t.cluster_index));
makes the x86_64 build produce, for the UN flag, exactly the pixels and 120-colour palette the unpatched aarch64 build produces, and the aarch64 build with the same one-line change produces the same pixels and palettes as that patched x86_64 build for every sample (UN flag, U+1F92B, U+1F307, U+1F9D4+1F3FE, U+1F600). So with a deterministic histogram order the two architectures agree exactly; without it they do not.
Suggested fix
Sort temp by colour (or iterate the map in key order) before the cluster pass, as above; it costs one sort of the histogram (at most max_histogram_entries items) and makes the output a pure function of the input across platforms. Happy to send that as a PR if you would take it.
pngquant --speed 1 --quality 85-95(and any speed ≤ 7) produces a different palette and different pixels for the same input on x86_64 and on aarch64, with the same pngquant/libimagequant version. It is fully deterministic on each architecture (run to run, and independent ofRAYON_NUM_THREADS), just not across them.Found while building Noto Color Emoji in a Linux distribution: 92 of the 4,004 CBDT bitmaps differed between the two architectures' builds.
Cause
Histogram::finalize_builder(src/hist.rs) turns the colour hash map into the item list withand everything after that is order-sensitive: the cluster bucketing assigns
items[next_index]in that order,total_perceptual_weightis an f64 sum in that order, median cut'sselect_nth_unstable_by_key/sort_unstable_by_keyresolve ties by position, and k-means chunks the slice withpar_chunks_mut(256).U32Hashermakes the hash deterministic, but the iteration order of aHashMapalso depends on the table's probe-group width, which is 16 bytes with hashbrown's SSE2 group on x86_64 and 8 bytes with the NEON group on aarch64 (hashbrown/src/control/group/{sse2,neon}.rs). Colliding keys therefore land in different buckets on the two architectures, the map iterates in a different order, and the small order-dependent differences are amplified by the refinement iterations into a different palette.Evidence
pngquant 3.0.3 packaged from source on both architectures, UN flag image from noto-emoji (136x128 RGBA): palette of 111 colours on x86_64, 120 on aarch64; other images 151 vs 142, 201 vs 219, 35 vs 37. Identical at
--speed 8and above (no k-means iterations, no feedback trials) and with a small fixed palette (pngquant 16).Not the SIMD
f_pixel::diff: changing the x86_64 horizontal add to NEON's association order (r + (g + b)) changes nothing.Not libm: f32
powf/expf/logffingerprints are identical on both.Adding one line after the
temp.extend(...):makes the x86_64 build produce, for the UN flag, exactly the pixels and 120-colour palette the unpatched aarch64 build produces, and the aarch64 build with the same one-line change produces the same pixels and palettes as that patched x86_64 build for every sample (UN flag, U+1F92B, U+1F307, U+1F9D4+1F3FE, U+1F600). So with a deterministic histogram order the two architectures agree exactly; without it they do not.
Suggested fix
Sort
tempby colour (or iterate the map in key order) before the cluster pass, as above; it costs one sort of the histogram (at mostmax_histogram_entriesitems) and makes the output a pure function of the input across platforms. Happy to send that as a PR if you would take it.