Repository navigation
ggml-opencl: add opt-in Adreno xmem F16xF32 GEMM for prefill - #22755
Conversation
|
Hi @happyyzy, thanks for your contribution! Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:
Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below. |
|
Thanks for the note. I closed the other open PR (#22117) and will focus on this smaller xmem GEMM PR first. |
|
Thank you - this is much easier. Will take a closer look in the next few days. |
|
I was able to reproduce your results on A840 with Qwen3-1.7B-f16. A840With
build: f79069d9d (9050) Without
build: f79069d9d (9050) X2-90With xmem,
build: 2a52ab71a (9072)
build: 2a52ab71a (9072) Without xmem,
build: 2a52ab71a (9072)
build: 2a52ab71a (9072) |
|
Thanks, addressed both comments: the xmem program is now local to |
|
Just checking in on this PR. The requested review changes have been addressed, and the maintainer-side A840 / X2-90 benchmark results look positive. Please let me know if there is anything else I should adjust on my side. |
|
Thanks, addressed the naming comments: the kernel file is now |
* ggml-opencl: add Adreno xmem F16xF32 GEMM for prefill * ggml-opencl: address Adreno xmem review comments * ggml-opencl: align xmem gemm kernel naming --------- Co-authored-by: Your Name <your@email.com>
* ggml-opencl: add Adreno xmem F16xF32 GEMM for prefill * ggml-opencl: address Adreno xmem review comments * ggml-opencl: align xmem gemm kernel naming --------- Co-authored-by: Your Name <your@email.com> (cherry picked from commit a9883db)
* ggml-opencl: add Adreno xmem F16xF32 GEMM for prefill * ggml-opencl: address Adreno xmem review comments * ggml-opencl: align xmem gemm kernel naming --------- Co-authored-by: Your Name <your@email.com>
* ggml-opencl: add Adreno xmem F16xF32 GEMM for prefill * ggml-opencl: address Adreno xmem review comments * ggml-opencl: align xmem gemm kernel naming --------- Co-authored-by: Your Name <your@email.com>
* ggml-opencl: add Adreno xmem F16xF32 GEMM for prefill * ggml-opencl: address Adreno xmem review comments * ggml-opencl: align xmem gemm kernel naming --------- Co-authored-by: Your Name <your@email.com>
Summary
This PR adds an opt-in Adreno xmem GEMM path for OpenCL prefill matmul.
Scope:
GGML_OPENCL_USE_ADRENO_KERNELSGGML_OPENCL_ADRENO_XMEM_GEMM=1F16 x F32 -> F32GGML_OP_MUL_MATN > 1, so token-generation / GEMV decode is not routed through this pathThe implementation keeps the existing ggml tensor layout externally and uses a small bridge around the xmem GEMM:
The generic OpenCL matmul path remains unchanged unless the new runtime opt-in is set.
Results
Tested on Adreno 830 with OpenCL:
Qwen2.5 1.5B F16
Before, baseline OpenCL:
After, with
GGML_OPENCL_ADRENO_XMEM_GEMM=1:Prefill improved from
204.98 tok/sto356.19 tok/s, about1.74x.Qwen2.5 3B F16
Before, baseline OpenCL:
After, with
GGML_OPENCL_ADRENO_XMEM_GEMM=1:Prefill improved from
101.26 tok/sto163.90 tok/s, about1.62x.Decode is intentionally unchanged. Decode-only profiling confirmed that token generation stays on the existing OpenCL path (
adreno_xmemcount = 0).Correctness
Checked end-to-end generation with the xmem path enabled on Qwen2.5 1.5B F16 and Qwen2.5 3B F16. Both models produced normal decode output.
Notes
This path depends on Qualcomm Adreno OpenCL subgroup constant-load extensions and is therefore guarded behind the existing Adreno kernel build option plus an explicit runtime environment variable.
Requirements