Skip to content

GGML direct conv2d support #739

Description

@netrunnereve

Since GGML now has direct conv2d support for CPU and Vulkan we might want to try it out here and see if it helps. Compared to im2col this uses less memory and should run faster on GPUs that don't have matrix cores.

As a quick test I naively switched all instances of ggml_conv_2d in the code with ggml_conv_2d_direct and replaced the ggml directory with the one from llama.cpp. Right now it generates images fine on CPU (it's a bit slower than im2col) but it fails with a segfault on Vulkan.

Activity

  1. Green-Sky commented on Jul 24, 2025

    @Green-Sky
    Contributor

    Upstream was just synced with the vulkan code. @leejet just missed the update by a few hours 😆

    edit: There is also support for opencl. I kinda miss cuda support...

  2. Green-Sky commented on Jul 25, 2025

    @Green-Sky
    Contributor

    Just tried the vulkan backend and yes, it sadly fails.

        case GGML_OP_CONV_2D:
            if (src0->type == GGML_TYPE_F32 && src1->type == GGML_TYPE_F32 && dst->type == GGML_TYPE_F32 &&
                ggml_is_contiguous(src0) && ggml_is_contiguous(src1) && ggml_is_contiguous(dst)) {
                return ctx->device->pipeline_conv2d_f32;
            }
            return nullptr;
    

    I think the inputs are contiguous, but they are not all f32...

    (gdb) p *src0
    $5 = {type = GGML_TYPE_F16,
    
    (gdb) p *src1
    $6 = {type = GGML_TYPE_F32,
    
  3. Green-Sky commented on Jul 25, 2025

    @Green-Sky
    Contributor

    I hacked together a patch for ggml. Will post numbers later.

    ggml-org/llama.cpp#14872

  4. daniandtheweb commented on Jul 25, 2025

    @daniandtheweb
    Contributor

    I've just tried the changes on my laptop and unfortunately there's no gains for "old" APUs. The ryzen 7 4700u on my laptop using vulkan for a 512*512 image on SD 1.5 goes from 3.78s/it on master to 4.24s/it with the new direct implementation.
    I replaced all instances of ggml_conv_2d with their direct counterpart and pasted the ggml folder from the PR.
    I hope that at least discrete GPUs can get a speedup from this.

  5. Green-Sky commented on Jul 25, 2025

    @Green-Sky
    Contributor

    Yes, I don't expect it to be faster in vulkan, since it is a scalar implementation. Especially when you compare it against im_col+mul with cm1/cm2.

  6. daniandtheweb commented on Jul 25, 2025

    @daniandtheweb
    Contributor

    No coopmat at all on the 4700U, it's a small Vega GPU that uses some ram as vram, so it's going to require a lot of optimization to have some gains. Pytorch takes about half of the time to generate so there's still some margin to improve.

  7. Green-Sky commented on Jul 25, 2025

    @Green-Sky
    Contributor

    Oh I see.

    SD1 (CyberRealistic_V9_FP16)

    512x768

    method diffusion buffer sampling speed vae buffer vae speed
    direct 1219.65mb 1.11s/it 960mb 1.15s
    direct w fa 1217.87mb 1.11s/it 960mb 1.16s
    im2col+mul 1220.07mb 1.01it/s 2496mb 2.08s
    im2col+mul w fa 1218.29mb 1.01it/s 2496mb 2.10s

    Quick testing reveals big wins for VAE, I was sure the memory would go down alot, but I actually did not expect the VAE to be faster too.

    But yes, sampling is ~10% slower in this test on my RTX 2070 mobile.

    edit: compiled with coopmat1

  8. stduhpf commented on Jul 25, 2025

    @stduhpf
    Contributor

    VAE is very slow on Vulkan right now, and the compute buffer very easily exceeds the 4GB allocation limit. For that alone I think these changes are worth the slower diffusion.

  9. Green-Sky commented on Jul 25, 2025

    @Green-Sky
    Contributor
    direct im2col
    Image Image

    Only a slight difference too.

  10. Green-Sky commented on Jul 25, 2025

    @Green-Sky
    Contributor

    VAE is very slow on Vulkan right now, and the compute buffer very easily exceeds the 4GB allocation limit. For that alone I think these changes are worth the slower diffusion.

    We don't need to switch it for the diffuser. We can make the VAE use direct and the diffusion model continue use im2col. At least until it becomes better.

  11. daniandtheweb commented on Jul 25, 2025

    @daniandtheweb
    Contributor

    VAE decoding goes from 20s on master to 6s with these changes on my laptop so I also agree that it's better only implementing it only on the VAE for now.

  12. Green-Sky commented on Jul 25, 2025

    @Green-Sky
    Contributor

    flux.1 light

    (hyper-8step lora pruned and applied)

    f16 (v)ae

    768x768

    method diffusion buffer sampling speed vae buffer vae speed
    direct 1105.07mb 5.47s/it 1440mb 1.74s
    direct fa 505.07mb 4.98s/it 1440mb 1.77s
    im2col+mul 1105.07mb 5.46s/it 3744mb 3.34s
    im2col+mul fa 505.07mb 4.92s/it 3744mb 3.34s

    Flux is expectedly not affected by convolution in the diffuser.
    VAE behaves the same.

    Bonus: vae cpu numbers

    8 threads

    method time
    direct 74.47s
    im2col+mul fa 71.93s
  13. rmatif commented on Jul 25, 2025

    @rmatif
    Contributor

    I've also added the conv2d op to OpenCL: ggml-org/ggml@7ce741c it seems the previous sync missed it, as this was upstreamed to ggml just a couple of hours after the sync

  14. daniandtheweb commented on Jul 25, 2025

    @daniandtheweb
    Contributor

    I've never touched the vae.hpp file before but changing the Conv2d operations in there to Conv2dDirect operations (leaving the diffusion step with the old Conv2d) doesn't perform like changing every operation in ggml-extend.hpp to Conv2dDirect (10s for vae decoding vs 6s when changing everything). Does anyone know if there's any other call do Conv2d in the vae process that I'm missing other than the ones in vae.hpp?

    EDIT: Found it in common.hpp

  15. netrunnereve commented on Jul 25, 2025

    @netrunnereve
    Author

    With a q6_K SDXL Lightning on my RX 470 generating a 512x512 image:

    Direct:

      |==================================================| 4/4 - 1.09s/it
    [INFO ] stable-diffusion.cpp:1806 - sampling completed, taking 4.37s
    [INFO ] stable-diffusion.cpp:1814 - generating 1 latent images completed, taking 4.37s
    [INFO ] stable-diffusion.cpp:1817 - decoding 1 latents
    [INFO ] stable-diffusion.cpp:1827 - latent 1 decoded, taking 1.40s
    [INFO ] stable-diffusion.cpp:1831 - decode_first_stage completed, taking 1.40s
    [INFO ] stable-diffusion.cpp:2079 - generate_image completed in 6.04s
    

    Indirect:

      |==================================================| 4/4 - 1.01it/s
    [INFO ] stable-diffusion.cpp:1806 - sampling completed, taking 4.01s
    [INFO ] stable-diffusion.cpp:1814 - generating 1 latent images completed, taking 4.01s
    [INFO ] stable-diffusion.cpp:1817 - decoding 1 latents
    [INFO ] stable-diffusion.cpp:1827 - latent 1 decoded, taking 6.91s
    [INFO ] stable-diffusion.cpp:1831 - decode_first_stage completed, taking 6.91s
    [INFO ] stable-diffusion.cpp:2079 - generate_image completed in 11.19s
    

    This is great. The direct version uses around 1GB less memory during decoding and it's way faster there, while at the same time it's a bit slower at sampling. So yeah it's probably a good idea to use the direct mode only for the vae.

  16. 10 remaining items

  17. daniandtheweb commented on Jul 27, 2025

    @daniandtheweb
    Contributor

    This is where they implemented the op in ggml JingXuuu/ggml@d85c985 but if I remember correctly there's been a winograd PR some time ago that got stuck because of licensing issues.

    EDIT: Nevermind, it's because it was based on openCNN, this seems not based on it at all

  18. rmatif commented on Jul 27, 2025

    @rmatif
    Contributor

    I tested their implementation a while ago, it was CPU-only, and the compute buffer size was larger than with im2col, and the VAE buffer remains the same. Performance varied depending on the CPU.

    I started working on a Winograd implementation on my own, but the performance gains were minimal, which demotivated me from continuing. Additionally, at high resolutions, the conv2d operation gets dominated by the attention blocks, so it’s not even worth optimizing in the context of sdcpp

  19. Green-Sky commented on Jul 28, 2025

    @Green-Sky
    Contributor

    Ok, the vulkan patch is now merged and synced. So you can now more easily test it by just pulling latest ggml.
    (and patching conv2d calls to use _direct instead...)

  20. daniandtheweb commented on Jul 28, 2025

    @daniandtheweb
    Contributor

    @Green-Sky @netrunnereve if you're okay with it and don't plan to do it yourself I may open a PR later today with the changes required to run the VAE stage in direct mode given the fact that it doesn't seem to provide any downside.

  21. Green-Sky commented on Jul 28, 2025

    @Green-Sky
    Contributor

    @Green-Sky @netrunnereve if you're okay with it and don't plan to do it yourself I may open a PR later today with the changes required to run the VAE stage in direct mode given the fact that it doesn't seem to provide any downside.

    You might have to work some magic to see if a given backend supports the OP. And/Or provide a command line switch.

  22. daniandtheweb commented on Jul 28, 2025

    @daniandtheweb
    Contributor

    For now I've opened the PR as a draft so it's easy to test it. If I come up with some non tricky way to enable it only for the compatible backends I'll mark it as ready.

  23. netrunnereve commented on Jul 28, 2025

    @netrunnereve
    Author

    You might have to work some magic to see if a given backend supports the OP. And/Or provide a command line switch.

    Even if the backend supports the OP it's too early to tell if it's going to be faster for everyone especially with coopmat and all that. I think we should make it optional like flash attention for now.

  24. daniandtheweb commented on Jul 28, 2025

    @daniandtheweb
    Contributor

    Even if the backend supports the OP it's too early to tell if it's going to be faster for everyone especially with coopmat and all that. I think we should make it optional like flash attention for now.

    I have added a Cmake option to manually enable the direct mode, just like the already present SD_FAST_SOFTMAX.

  25. Green-Sky commented on Jul 28, 2025

    @Green-Sky
    Contributor

    Even if the backend supports the OP it's too early to tell if it's going to be faster for everyone especially with coopmat and all that. I think we should make it optional like flash attention for now.

    I have added a Cmake option to manually enable the direct mode, just like the already present SD_FAST_SOFTMAX.

    Which is what flash attention was, then silently fell under the bus, got abandoned and stopped working, until I made it a proper runtime flag. :)

    Similar to --diffusion-fa, we should do --diffusion-conv-direct and --vae-conv-direct or similar.

  26. daniandtheweb commented on Jul 28, 2025

    @daniandtheweb
    Contributor

    I can try and make a proper runtime flag if that sounds better.
    If it ends up not being the favorite method I can always revert to the previous commit and use cmake instead.

  27. leejet commented on Jul 29, 2025

    @leejet
    Owner
    static void ggml_compute_forward_conv_2d_impl(const ggml_compute_params * params,
                                                  const ggml_tensor *         kernel,  // [KW, KH, IC, OC]
                                                  const ggml_tensor *         src,     // [W, H, C, N]
                                                  ggml_tensor *               dst,     // [OW, OH, OC, N]
                                                  ggml_type                   kernel_type) {
    
        GGML_ASSERT(ggml_is_contiguous(kernel));
        GGML_ASSERT(kernel_type == GGML_TYPE_F16 || kernel_type == GGML_TYPE_F32);
        GGML_ASSERT(kernel->type == kernel_type);
    
        const ggml_type_traits * traits = ggml_get_type_traits(kernel_type);
    
        const int32_t stride_x   = dst->op_params[0];
        const int32_t stride_y   = dst->op_params[1];
        const int32_t pad_x      = dst->op_params[2];
        const int32_t pad_y      = dst->op_params[3];
        const int32_t dilation_x = dst->op_params[4];
        const int32_t dilation_y = dst->op_params[5];
    
        const int64_t c_in  = src->ne[2];
        const int64_t c_out = kernel->ne[3];
        GGML_ASSERT(c_in == kernel->ne[2]);
    
        const int64_t src_w = src->ne[0];
        const int64_t src_h = src->ne[1];
        const int64_t knl_w = kernel->ne[0];
        const int64_t knl_h = kernel->ne[1];
        const int64_t dst_w = dst->ne[0];
        const int64_t dst_h = dst->ne[1];
    
        const float * src_data = (float *) src->data;
        void  * knl_data       = kernel->data;
        float * dst_data       = (float *) dst->data;
    
        const int64_t knl_n           = knl_w * knl_h * c_in;
        const int64_t patch_total     = dst->ne[3] * dst_w * dst_h;
    
        const int64_t space_per_patch   = knl_n * traits->type_size + c_out * sizeof(float);
        const int64_t batch_size        = params->wsize / space_per_patch;
        const int64_t patches_per_batch = batch_size > 8 ? (batch_size / 8) * 8 : batch_size;
        const int64_t batch_n           = (patch_total + patches_per_batch - 1) / patches_per_batch;
    
        GGML_ASSERT(patches_per_batch > 0 && batch_size >= 1);
    
        void * tmp = params->wdata;
    
        for (int64_t batch_i = 0; batch_i < batch_n; ++batch_i) {
    
            const int64_t patch_start_batch = batch_i * patches_per_batch;
            const int64_t patch_end_batch   = std::min(patch_start_batch + patches_per_batch,
                                                  patch_total);
            const int64_t patch_n           = patch_end_batch - patch_start_batch;
    
            const int64_t patch_per_thread  = (patch_n + params->nth - 1) / params->nth;
            const int64_t patch_start       = patch_start_batch + params->ith * patch_per_thread;
            const int64_t patch_end         = std::min(patch_start + patch_per_thread, patch_end_batch);
    
            //im2col for a patch
            for (int64_t p = patch_start; p < patch_end; ++p) {
                const int64_t  batch_n     =  p / (dst_w * dst_h);
                const int64_t  src_x       = (p / dst_w) % dst_h;
                const int64_t  src_y       =  p % dst_w;
    
                const float * src_base = (const float *)((const char *)src_data + batch_n * src->nb[3]);
                char *        dst_row  = (char *) tmp + (p % patches_per_batch) * knl_n * traits->type_size;
    
                for (int64_t ic = 0; ic < c_in; ++ic) {
                    for (int64_t ky = 0; ky < knl_h; ++ky) {
                        for (int64_t kx = 0; kx < knl_w; ++kx) {
                            const int64_t sy = src_x * stride_y + ky * dilation_y - pad_y;
                            const int64_t sx = src_y * stride_x + kx * dilation_x - pad_x;
    
                            int64_t dst_idx = ic * (knl_h * knl_w) + ky * knl_w + kx;
    
                            float src_val;
                            if (sy < 0 || sy >= src_h || sx < 0 || sx >= src_w) {
                                src_val = 0.0f;
                            } else {
                                const float * src_ptr = (const float *)((const char *)src_base + sx * src->nb[0] + sy * src->nb[1] + ic * src->nb[2]);
                                src_val               = *src_ptr;
                            }
    
                            char * element_ptr = dst_row + dst_idx * traits->type_size;
                            if (kernel_type == GGML_TYPE_F32) {
                                *(float *) element_ptr = src_val;
                            } else if (kernel_type == GGML_TYPE_F16) {
                                *(ggml_fp16_t *) element_ptr = GGML_CPU_FP32_TO_FP16(src_val);
                            }
                        }
                    }
                }
            }   // patches handled by this thread
    
            ggml_barrier(params->threadpool);
    
            float * gemm_output = (float *) ((char *) tmp + patches_per_batch * knl_n * traits->type_size);
    
            GGML_ASSERT(gemm_output + patch_n * c_out <= (float*)tmp + params->wsize);
    
            // GEMM: patches[patch_n, knl_n] × kernel[knl_n, c_out] = output[patch_n, c_out]
            ggml_call_mul_mat(kernel_type, params, patch_n, c_out, knl_n, tmp, knl_data, gemm_output);
    
            ggml_barrier(params->threadpool);
    
    
            //permute back [OC, N, OH, OW] to [N, OC, OH, OW]
            const int64_t permute_per_thread = (patch_n + params->nth - 1) / params->nth;
            const int64_t permute_start = params->ith * permute_per_thread;
            const int64_t permute_end = std::min(permute_start + permute_per_thread, patch_n);
    
            for (int64_t i = permute_start; i < permute_end; ++i) {
                const int64_t p       = patch_start_batch + i;
                const int64_t batch_n = p / (dst_w * dst_h);
                const int64_t dst_y   = (p / dst_w) % dst_h;
                const int64_t dst_x   = p % dst_w;
    
                for (int64_t oc = 0; oc < c_out; ++oc) {
                    const float value = gemm_output[i * c_out + oc];
                    float * dst_ptr = (float *)((char *)dst_data + dst_x * dst->nb[0] + dst_y * dst->nb[1] + oc * dst->nb[2] + batch_n * dst->nb[3]);
                    *dst_ptr = value;
                }
            }
        }
    }
    

    ggml_conv2d_direct seems to be a combination of im2col and gemm, with data divided into small batches.

  28. Green-Sky commented on Jul 29, 2025

    @Green-Sky
    Contributor

    ggml-org/llama.cpp#14933

    Please check vulkan performance against this pr. Looks like a slight uplift on my card now.

  29. daniandtheweb commented on Jul 29, 2025

    @daniandtheweb
    Contributor

    The numbers on that PR look very promising.

    stable diffusion.cpp im2col, sd 1.5, 512x512, 20 step: diffusion 3.83 s/it, vae decode 20s
    
    stable-diffusion.cpp conv2d master, sd 1.5, 512x512, 20 step: diffusion 4.39 s/it, vae decode 6.18s
    
    stable-diffusion.cpp conv2d PR, sd 1.5, 512x512, 20 step: diffusion 4.08 s/it, vae decode 5.96s
    
  30. FSSRepo commented on Aug 25, 2025

    @FSSRepo
    Contributor

    Hey, I already had a similar approach to this one #221 , but for CUDA, taking advantage of the tensor cores (merge im2col and gemm). I abandoned the PR due to lack of time and also lack of interest

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions