Skip to content

Manage group values and states by blocks in aggregation #11931

Description

@Rachelint

Is your feature request related to a problem or challenge?

Now we manage the group values and the aggregation states by a single big vector growing constantly.
This solution is simple to impl, but really leads to some extra cpu cost according to the cpu profile.
Maybe we should manage them by blocks like duckdb.

Describe the solution you'd like

It may be a big work, I want to finish it through following steps:

  • Sketch the total procedure.
  • Impl the block based group values management in GroupValuesRows.
  • Impl the block based group values management in other GroupValues impls.
  • Impl the block based states management in different GroupAccumulator impls.

The general design is similar as #7065 , but introduce it into GroupValues, not only GroupAccumulators.

Describe alternatives you've considered

No response

Additional context

The cpu cost flamegraph:
https://github.com/Rachelint/drawio-store/blob/main/cpucosts0811.png

Activity

  1. Rachelint commented on Aug 10, 2024

    @Rachelint
    ContributorAuthor

    take

  2. 2010YOUY01 commented on Aug 11, 2024

    @2010YOUY01
    Contributor

    Is the plan to manage group values and states in two different kinds of blocks, or a unified block?

    There are so many good optimizations in the aggregation code now, they make the implementation a bit hard to understand, I was thinking managing group values + states in a single block could also be a good cleanup

  3. Rachelint commented on Aug 11, 2024

    @Rachelint
    ContributorAuthor

    Is the plan to manage group values and states in two different kinds of blocks, or a unified block?

    There are so many good optimizations in the aggregation code now, they make the implementation a bit hard to understand, I was thinking managing group values + states in a single block could also be a good cleanup

    Plan to make them two kinds of blocks, because group values is the inner struct in GroupValues, and states is the one in different GroupAccumulator, seems hard to manage them in a same place.

    The design is similar as #7065 , but introduce it into GroupValues, not only GroupAccumulators.

    The reason why doing it is according to the cpu flamegraph, the growing of the big single group values in GroupValues and states in respecitve GroupAccumulator seems really cost cpu(due to copy caused by growing).

  4. Rachelint commented on Aug 12, 2024

    @Rachelint
    ContributorAuthor

    The sketch's detailed design

    1. When will the blocked method triggered?

    • It should not be streaming aggregation, because steaming depends on the excact Emit::First(n) mode, and it is too expansive to impl it in blocked method.
    • The blocked GroupValues will be triggered, if we found the usedGroupValues impl support it.
    • The blocked GroupAccumulator will be only triggered, when all the used GroupAccumulators support blocked, and the used GroupValues supports blocked too.

    2. Introduce new emit modes used in blocked method

    It can support emit multiple blocks in GroupValuess and GroupAccumulators now:

    /// Describes how many rows should be emitted during grouping.
    #[derive(Debug, Clone, Copy)]
    pub enum EmitTo {
        /// Emit all groups
        All,
        /// Emit only the first `n` groups and shift all existing group
        /// indexes down by `n`.
        ///
        /// For example, if `n=10`, group_index `0, 1, ... 9` are emitted
        /// and group indexes '`10, 11, 12, ...` become `0, 1, 2, ...`.
        First(usize),
        /// Emit all groups managed by blocks
        AllBlocks,
        /// Emit only the first `n` group blocks,
        /// similar as `First`, but used in blocked `GroupValues` and `GroupAccumulator`.
        ///
        /// For example, `n=3`, `block size=4`, finally 12 groups will be returned.
        FirstBlocks(usize),
    }
    

    For incrementally development for blocked method for so many detailed GroupValues and GroupAccumulator impls. This sketch pr did a lot of compatibility works, and combinations are allowed:

    • Single GroupAccumulator + single GroupAccumulator
    • Blocked GroupValues + single GroupAccumulator
    • Blocked GroupValues + blocked GroupAccumulator

    3. Introduce GroupIndices to do communication between GroupValues and GroupAccumulator

    One of the problem is how to let GroupAccumulator know the if group indices is flat or blocked?
    I introduce GroupIndices to make it, but it indeed leads to api change for GroupAccumulator.

    pub enum GroupIndices<'a> {
        Flat(&'a [u64]),
        Blocked(&'a [u64]),
    }
    
    #[derive(Debug, Clone, Copy, PartialEq, Eq)]
    pub enum GroupIndicesType {
        Flat,
        Blocked,
    }
    
    impl GroupIndicesType {
        pub fn typed_group_indices<'a>(&self, indices: &'a [u64]) -> GroupIndices<'a> {
            match self {
                GroupIndicesType::Flat => GroupIndices::Flat(indices),
                GroupIndicesType::Blocked => GroupIndices::Blocked(indices),
            }
        }
    }
    
  5. jayzhan211 commented on Aug 12, 2024

    @jayzhan211
    Contributor

    /// For example, n= 10, block size=4, n will be aligned to 12,
    /// and finally 3 blocks will be returned.
    FirstBlocks(usize)

    I think emitting with "n" blocks is much more straightforward. n = 3, block size = 4. emit 3 * 4 = 12 elements

  6. Rachelint commented on Aug 12, 2024

    @Rachelint
    ContributorAuthor

    /// For example, n= 10, block size=4, n will be aligned to 12,
    /// and finally 3 blocks will be returned.
    FirstBlocks(usize)

    I think emitting with "n" blocks is much more straightforward. n = 3, block size = 4. emit 3 * 4 = 12 elements

    It seems indeed more clear! I have switched to this in codes.

  7. Rachelint commented on Aug 12, 2024

    @Rachelint
    ContributorAuthor

    @alamb
    I have finished the draft framework for blocked aggregation intermediate management, and we can incrementally impl blocked method for different GroupAccumulators and GroupValues on it. Minding have a quick look?

    The general design can see:
    #11931 (comment)

    And the pr is here:
    #11943

    cc @jayzhan211 @2010YOUY01 @JasonLi-cn

  8. Dandandan commented on Mar 7, 2026

    @Dandandan
    Contributor

    Implementing this aggregation approach #20773 would also contribute to keeping the state small per aggregation in the partial aggregate (and hopefully improves performance). The approach could be repeated in final / final partitioned aggregation (not sure why it is not done in the paper (maybe it is not as "free" as it is in partial aggregation step).

    While still a large feature to implement, I think implementing it might be more local to aggregation (no changes needed to groupvalues / accumulators, etc...)

  9. Dandandan commented on Mar 16, 2026

    @Dandandan
    Contributor

    This PR
    #20964

    I think shows that we probably can do the block-based aggregation in a couple of smaller steps, something like:

    • Change group accumulators to allocate fixed-size blocks
    • Change group by values to allocate fixed-size blocks
    • Integrate indices with new fixed-size block allocation
  10. Rachelint commented on Mar 17, 2026

    @Rachelint
    ContributorAuthor

    This PR #20964

    I think shows that we probably can do the block-based aggregation in a couple of smaller steps, something like:

    • Change group accumulators to allocate fixed-size blocks
    • Change group by values to allocate fixed-size blocks
    • Integrate indices with new fixed-size block allocation

    Yes, I introduce the similar Blocks<T> in #15591 , but only used in accumulator, it should indeed also used in group values

  11. alamb commented on Aug 27, 2026

    @alamb
    Contributor

    I filed a ticket / epic to track the idea of blocked state management here

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

enhancementNew feature or request

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions