Skip to content

Vectorized hashing for hash aggregation code #26

Description

@Dandandan

Updating the hash aggregate implementation to use vectorized hashing should give a decent speed up to queries that are dependant on fast hash aggregate implementations.

Currently keys are generated of type Vec<u8> and are hashed row-by-row which causes

  • more memory usage
  • slow re-hashing of the backing hashmap
  • type un-aware hashing for simple primitive values

The implementation should also solve hash collisions, so the original should be able to be compared with the values.

There is some WIP code here apache/arrow#9213 which can be used as a starting point / to continue from.

Activity

  1. jorgecarleitao commented on Apr 21, 2021

    @jorgecarleitao
    Member

    On arrow2, I took the following approach:

    1. Expose an hash kernel
    2. use hash_hasher

    My hypothesis is that if we move hashing to arrow instead of doing item by item, we likely gain a lot. However, to do that, we need to tell our hasher to not re-hash the keys that it receives, thus the hash_hasher.

  2. Dandandan commented on Apr 21, 2021

    @Dandandan
    ContributorAuthor

    Thanks @jorgecarleitao !

    Yeah I would agree, would be great to have a hashing kernel or maybe some basic primitives to build one easily using arrow. Something like hash_hasher looks cool too and is actually very similar to the one used in the hash join (hashbrown hashmap + IdHashBuilder which just uses the identity function) . Looking at the code of hash_hasher, it's something similar done as in the hash join (IdHashBuilder), but seems that it should be doing a bit more work (e.g. the "hash combiner" works over bytes) and hashbrown is also slightly faster. I believe because of more inlining as the standard library one uses / exports the same crate..

    For this PR I was thinking to move the code to hash_utils

  3. jorgecarleitao commented on Apr 21, 2021

    @jorgecarleitao
    Member

    Perfect, we have the same understanding of the problem, then :)

    Yeah, I used hash_hasher to generalize the Dictionary builder to arbitrary types, but as long as we have the concepts right, we can change it; I agree that hashbrown + IdHashBuilder is more performant. 👍

  4. xudong963 commented on Apr 23, 2022

    @xudong963
    Member

    The issue seems stale?

  5. added a commit that references this issue on Jan 12, 2023
  6. Dandandan commented on Mar 13, 2025

    @Dandandan
    ContributorAuthor

    This is solved

  7. added a commit that references this issue on Aug 9, 2025
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions