Skip to content

CUDA : use mmvf for tiny-N F32/F16 weights at decode batch sizes - #28875

Closed
douyamv wants to merge 1 commit into
ggml-org:masterfrom
douyamv:pr/mmvf-tiny-n-f32
Closed

douyamv wants to merge 1 commit into
ggml-org:masterfrom
douyamv:pr/mmvf-tiny-n-f32

Conversation

@douyamv

@douyamv douyamv commented Sep 14, 2026

Copy link
Copy Markdown

Small F32/F16 weight matrices with N ≤ 64 rows (hyper-connection inject/gate projections [K, 4], GDN β/α
projections [2560, 32] in qwen4exp) were dispatched to cuBLAS for decode-sized batches: a TF32 GEMM plus a
split-K reduction kernel, ~8 µs and two launches per call, 4 calls per layer. ggml_cuda_should_use_mmvf already
accepted N ≤ 16 for these types; this raises the limit to 64. The vector kernel does them in one launch with full
FP32 accumulation and is not slower for any of the shapes we measured on sm_80.

🤖 Generated with Claude Code

Matrices with N <= 64 rows (hyper-connection inject/gate projections [K,4], GDN beta/alpha [2560,32]) went
through cuBLAS (TF32 GEMM + split-K reduce = 2 launches, ~8 us) for every decode step. The vector kernel handles
them in one launch with full FP32 accumulation.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@douyamv
douyamv requested a review from a team as a code owner September 14, 2026 01:34
@github-actions github-actions Bot added ggml changes relating to the ggml tensor library for machine learning CUDA Related to the CUDA backend labels Sep 14, 2026
@ggml-gh-bot

ggml-gh-bot Bot commented Sep 14, 2026

Copy link
Copy Markdown

Hi @douyamv, thanks for your contribution!

Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:

  • PR Template not respected: Please respect the template when creating a new pull request. Make sure to fill out all required sections.

  • Multiple open PRs from a new contributor: We limit new contributors (those without a previously merged PR) to 1 open PR at a time. You currently have 5 open PRs.

  • AI-generated content: While code is allowed to be generated by AI, please write the PR description and commit messages on your own without the help of AI.


Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below.

@ggml-gh-bot ggml-gh-bot Bot added the draft PR will be changed to draft by github-actions bot label Sep 14, 2026
@github-actions
github-actions Bot marked this pull request as draft September 14, 2026 01:39
@github-actions github-actions Bot removed the draft PR will be changed to draft by github-actions bot label Sep 14, 2026
@JohannesGaessler

Copy link
Copy Markdown
Contributor

According to the llama.cpp AI usage policy:

It is strictly prohibited to use AI to write your posts for you (bug reports, feature requests, pull request descriptions, Github discussions, responding to humans, ...).

@ggml-org ggml-org temporarily blocked douyamv Sep 17, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants