Repository navigation
GGML direct conv2d support #739
Description
Activity
Upstream was just synced with the vulkan code. @leejet just missed the update by a few hours 😆
edit: There is also support for opencl. I kinda miss cuda support...
Reacted by EveJust tried the vulkan backend and yes, it sadly fails.
case GGML_OP_CONV_2D: if (src0->type == GGML_TYPE_F32 && src1->type == GGML_TYPE_F32 && dst->type == GGML_TYPE_F32 && ggml_is_contiguous(src0) && ggml_is_contiguous(src1) && ggml_is_contiguous(dst)) { return ctx->device->pipeline_conv2d_f32; } return nullptr;I think the inputs are contiguous, but they are not all f32...
(gdb) p *src0 $5 = {type = GGML_TYPE_F16, (gdb) p *src1 $6 = {type = GGML_TYPE_F32,I hacked together a patch for ggml. Will post numbers later.
I've just tried the changes on my laptop and unfortunately there's no gains for "old" APUs. The ryzen 7 4700u on my laptop using vulkan for a 512*512 image on SD 1.5 goes from 3.78s/it on master to 4.24s/it with the new direct implementation.
I replaced all instances of ggml_conv_2d with their direct counterpart and pasted the ggml folder from the PR.
I hope that at least discrete GPUs can get a speedup from this.Yes, I don't expect it to be faster in vulkan, since it is a scalar implementation. Especially when you compare it against im_col+mul with cm1/cm2.
No coopmat at all on the 4700U, it's a small Vega GPU that uses some ram as vram, so it's going to require a lot of optimization to have some gains. Pytorch takes about half of the time to generate so there's still some margin to improve.
Oh I see.
SD1 (CyberRealistic_V9_FP16)
512x768
method diffusion buffer sampling speed vae buffer vae speed direct 1219.65mb 1.11s/it 960mb 1.15s direct w fa 1217.87mb 1.11s/it 960mb 1.16s im2col+mul 1220.07mb 1.01it/s 2496mb 2.08s im2col+mul w fa 1218.29mb 1.01it/s 2496mb 2.10s Quick testing reveals big wins for VAE, I was sure the memory would go down alot, but I actually did not expect the VAE to be faster too.
But yes, sampling is ~10% slower in this test on my RTX 2070 mobile.
edit: compiled with coopmat1
VAE is very slow on Vulkan right now, and the compute buffer very easily exceeds the 4GB allocation limit. For that alone I think these changes are worth the slower diffusion.
Reacted by Erik Scholz and dsignariusVAE is very slow on Vulkan right now, and the compute buffer very easily exceeds the 4GB allocation limit. For that alone I think these changes are worth the slower diffusion.
We don't need to switch it for the diffuser. We can make the VAE use direct and the diffusion model continue use im2col. At least until it becomes better.
Reacted by stduhpf and DanieleVAE decoding goes from 20s on master to 6s with these changes on my laptop so I also agree that it's better only implementing it only on the VAE for now.
Reacted by Erik Scholzflux.1 light
(hyper-8step lora pruned and applied)
f16 (v)ae
768x768
method diffusion buffer sampling speed vae buffer vae speed direct 1105.07mb 5.47s/it 1440mb 1.74s direct fa 505.07mb 4.98s/it 1440mb 1.77s im2col+mul 1105.07mb 5.46s/it 3744mb 3.34s im2col+mul fa 505.07mb 4.92s/it 3744mb 3.34s Flux is expectedly not affected by convolution in the diffuser.
VAE behaves the same.Bonus: vae cpu numbers
8 threads
method time direct 74.47s im2col+mul fa 71.93s Reacted by GreenShadowsI've also added the conv2d op to OpenCL: ggml-org/ggml@7ce741c it seems the previous sync missed it, as this was upstreamed to ggml just a couple of hours after the sync
Reacted by Erik ScholzI've never touched the vae.hpp file before but changing the Conv2d operations in there to Conv2dDirect operations (leaving the diffusion step with the old Conv2d) doesn't perform like changing every operation in ggml-extend.hpp to Conv2dDirect (10s for vae decoding vs 6s when changing everything). Does anyone know if there's any other call do Conv2d in the vae process that I'm missing other than the ones in vae.hpp?
EDIT: Found it in common.hpp
With a q6_K SDXL Lightning on my RX 470 generating a 512x512 image:
Direct:
|==================================================| 4/4 - 1.09s/it [INFO ] stable-diffusion.cpp:1806 - sampling completed, taking 4.37s [INFO ] stable-diffusion.cpp:1814 - generating 1 latent images completed, taking 4.37s [INFO ] stable-diffusion.cpp:1817 - decoding 1 latents [INFO ] stable-diffusion.cpp:1827 - latent 1 decoded, taking 1.40s [INFO ] stable-diffusion.cpp:1831 - decode_first_stage completed, taking 1.40s [INFO ] stable-diffusion.cpp:2079 - generate_image completed in 6.04sIndirect:
|==================================================| 4/4 - 1.01it/s [INFO ] stable-diffusion.cpp:1806 - sampling completed, taking 4.01s [INFO ] stable-diffusion.cpp:1814 - generating 1 latent images completed, taking 4.01s [INFO ] stable-diffusion.cpp:1817 - decoding 1 latents [INFO ] stable-diffusion.cpp:1827 - latent 1 decoded, taking 6.91s [INFO ] stable-diffusion.cpp:1831 - decode_first_stage completed, taking 6.91s [INFO ] stable-diffusion.cpp:2079 - generate_image completed in 11.19sThis is great. The direct version uses around 1GB less memory during decoding and it's way faster there, while at the same time it's a bit slower at sampling. So yeah it's probably a good idea to use the direct mode only for the vae.
Reacted by Erik Scholz10 remaining items
This is where they implemented the op in ggml JingXuuu/ggml@d85c985 but if I remember correctly there's been a winograd PR some time ago that got stuck because of licensing issues.
EDIT: Nevermind, it's because it was based on openCNN, this seems not based on it at all
I tested their implementation a while ago, it was CPU-only, and the compute buffer size was larger than with im2col, and the VAE buffer remains the same. Performance varied depending on the CPU.
I started working on a Winograd implementation on my own, but the performance gains were minimal, which demotivated me from continuing. Additionally, at high resolutions, the conv2d operation gets dominated by the attention blocks, so it’s not even worth optimizing in the context of sdcpp
Ok, the vulkan patch is now merged and synced. So you can now more easily test it by just pulling latest ggml.
(and patchingconv2dcalls to use_directinstead...)@Green-Sky @netrunnereve if you're okay with it and don't plan to do it yourself I may open a PR later today with the changes required to run the VAE stage in direct mode given the fact that it doesn't seem to provide any downside.
Reacted by Erik Scholz and Eve@Green-Sky @netrunnereve if you're okay with it and don't plan to do it yourself I may open a PR later today with the changes required to run the VAE stage in direct mode given the fact that it doesn't seem to provide any downside.
You might have to work some magic to see if a given backend supports the OP. And/Or provide a command line switch.
Reacted by DanieleFor now I've opened the PR as a draft so it's easy to test it. If I come up with some non tricky way to enable it only for the compatible backends I'll mark it as ready.
You might have to work some magic to see if a given backend supports the OP. And/Or provide a command line switch.
Even if the backend supports the OP it's too early to tell if it's going to be faster for everyone especially with coopmat and all that. I think we should make it optional like flash attention for now.
Reacted by DanieleEven if the backend supports the OP it's too early to tell if it's going to be faster for everyone especially with coopmat and all that. I think we should make it optional like flash attention for now.
I have added a Cmake option to manually enable the direct mode, just like the already present SD_FAST_SOFTMAX.
Even if the backend supports the OP it's too early to tell if it's going to be faster for everyone especially with coopmat and all that. I think we should make it optional like flash attention for now.
I have added a Cmake option to manually enable the direct mode, just like the already present SD_FAST_SOFTMAX.
Which is what flash attention was, then silently fell under the bus, got abandoned and stopped working, until I made it a proper runtime flag. :)
Similar to
--diffusion-fa, we should do--diffusion-conv-directand--vae-conv-director similar.I can try and make a proper runtime flag if that sounds better.
If it ends up not being the favorite method I can always revert to the previous commit and use cmake instead.Reacted by Erik Scholzstatic void ggml_compute_forward_conv_2d_impl(const ggml_compute_params * params, const ggml_tensor * kernel, // [KW, KH, IC, OC] const ggml_tensor * src, // [W, H, C, N] ggml_tensor * dst, // [OW, OH, OC, N] ggml_type kernel_type) { GGML_ASSERT(ggml_is_contiguous(kernel)); GGML_ASSERT(kernel_type == GGML_TYPE_F16 || kernel_type == GGML_TYPE_F32); GGML_ASSERT(kernel->type == kernel_type); const ggml_type_traits * traits = ggml_get_type_traits(kernel_type); const int32_t stride_x = dst->op_params[0]; const int32_t stride_y = dst->op_params[1]; const int32_t pad_x = dst->op_params[2]; const int32_t pad_y = dst->op_params[3]; const int32_t dilation_x = dst->op_params[4]; const int32_t dilation_y = dst->op_params[5]; const int64_t c_in = src->ne[2]; const int64_t c_out = kernel->ne[3]; GGML_ASSERT(c_in == kernel->ne[2]); const int64_t src_w = src->ne[0]; const int64_t src_h = src->ne[1]; const int64_t knl_w = kernel->ne[0]; const int64_t knl_h = kernel->ne[1]; const int64_t dst_w = dst->ne[0]; const int64_t dst_h = dst->ne[1]; const float * src_data = (float *) src->data; void * knl_data = kernel->data; float * dst_data = (float *) dst->data; const int64_t knl_n = knl_w * knl_h * c_in; const int64_t patch_total = dst->ne[3] * dst_w * dst_h; const int64_t space_per_patch = knl_n * traits->type_size + c_out * sizeof(float); const int64_t batch_size = params->wsize / space_per_patch; const int64_t patches_per_batch = batch_size > 8 ? (batch_size / 8) * 8 : batch_size; const int64_t batch_n = (patch_total + patches_per_batch - 1) / patches_per_batch; GGML_ASSERT(patches_per_batch > 0 && batch_size >= 1); void * tmp = params->wdata; for (int64_t batch_i = 0; batch_i < batch_n; ++batch_i) { const int64_t patch_start_batch = batch_i * patches_per_batch; const int64_t patch_end_batch = std::min(patch_start_batch + patches_per_batch, patch_total); const int64_t patch_n = patch_end_batch - patch_start_batch; const int64_t patch_per_thread = (patch_n + params->nth - 1) / params->nth; const int64_t patch_start = patch_start_batch + params->ith * patch_per_thread; const int64_t patch_end = std::min(patch_start + patch_per_thread, patch_end_batch); //im2col for a patch for (int64_t p = patch_start; p < patch_end; ++p) { const int64_t batch_n = p / (dst_w * dst_h); const int64_t src_x = (p / dst_w) % dst_h; const int64_t src_y = p % dst_w; const float * src_base = (const float *)((const char *)src_data + batch_n * src->nb[3]); char * dst_row = (char *) tmp + (p % patches_per_batch) * knl_n * traits->type_size; for (int64_t ic = 0; ic < c_in; ++ic) { for (int64_t ky = 0; ky < knl_h; ++ky) { for (int64_t kx = 0; kx < knl_w; ++kx) { const int64_t sy = src_x * stride_y + ky * dilation_y - pad_y; const int64_t sx = src_y * stride_x + kx * dilation_x - pad_x; int64_t dst_idx = ic * (knl_h * knl_w) + ky * knl_w + kx; float src_val; if (sy < 0 || sy >= src_h || sx < 0 || sx >= src_w) { src_val = 0.0f; } else { const float * src_ptr = (const float *)((const char *)src_base + sx * src->nb[0] + sy * src->nb[1] + ic * src->nb[2]); src_val = *src_ptr; } char * element_ptr = dst_row + dst_idx * traits->type_size; if (kernel_type == GGML_TYPE_F32) { *(float *) element_ptr = src_val; } else if (kernel_type == GGML_TYPE_F16) { *(ggml_fp16_t *) element_ptr = GGML_CPU_FP32_TO_FP16(src_val); } } } } } // patches handled by this thread ggml_barrier(params->threadpool); float * gemm_output = (float *) ((char *) tmp + patches_per_batch * knl_n * traits->type_size); GGML_ASSERT(gemm_output + patch_n * c_out <= (float*)tmp + params->wsize); // GEMM: patches[patch_n, knl_n] × kernel[knl_n, c_out] = output[patch_n, c_out] ggml_call_mul_mat(kernel_type, params, patch_n, c_out, knl_n, tmp, knl_data, gemm_output); ggml_barrier(params->threadpool); //permute back [OC, N, OH, OW] to [N, OC, OH, OW] const int64_t permute_per_thread = (patch_n + params->nth - 1) / params->nth; const int64_t permute_start = params->ith * permute_per_thread; const int64_t permute_end = std::min(permute_start + permute_per_thread, patch_n); for (int64_t i = permute_start; i < permute_end; ++i) { const int64_t p = patch_start_batch + i; const int64_t batch_n = p / (dst_w * dst_h); const int64_t dst_y = (p / dst_w) % dst_h; const int64_t dst_x = p % dst_w; for (int64_t oc = 0; oc < c_out; ++oc) { const float value = gemm_output[i * c_out + oc]; float * dst_ptr = (float *)((char *)dst_data + dst_x * dst->nb[0] + dst_y * dst->nb[1] + oc * dst->nb[2] + batch_n * dst->nb[3]); *dst_ptr = value; } } } }ggml_conv2d_direct seems to be a combination of im2col and gemm, with data divided into small batches.
Reacted by Erik ScholzPlease check vulkan performance against this pr. Looks like a slight uplift on my card now.
Reacted by DanieleThe numbers on that PR look very promising.
stable diffusion.cpp im2col, sd 1.5, 512x512, 20 step: diffusion 3.83 s/it, vae decode 20sstable-diffusion.cpp conv2d master, sd 1.5, 512x512, 20 step: diffusion 4.39 s/it, vae decode 6.18sstable-diffusion.cpp conv2d PR, sd 1.5, 512x512, 20 step: diffusion 4.08 s/it, vae decode 5.96sReacted by Erik Scholz and GreenShadowsHey, I already had a similar approach to this one #221 , but for CUDA, taking advantage of the tensor cores (merge im2col and gemm). I abandoned the PR due to lack of time and also lack of interest


Since GGML now has direct conv2d support for CPU and Vulkan we might want to try it out here and see if it helps. Compared to im2col this uses less memory and should run faster on GPUs that don't have matrix cores.
As a quick test I naively switched all instances of
ggml_conv_2din the code withggml_conv_2d_directand replaced the ggml directory with the one from llama.cpp. Right now it generates images fine on CPU (it's a bit slower than im2col) but it fails with a segfault on Vulkan.