Conversation
Matrices with N <= 64 rows (hyper-connection inject/gate projections [K,4], GDN beta/alpha [2560,32]) went through cuBLAS (TF32 GEMM + split-K reduce = 2 launches, ~8 us) for every decode step. The vector kernel handles them in one launch with full FP32 accumulation. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
|
Hi @douyamv, thanks for your contribution! Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:
Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below. |
|
According to the llama.cpp AI usage policy:
|
Small F32/F16 weight matrices with N ≤ 64 rows (hyper-connection inject/gate projections
[K, 4], GDN β/αprojections
[2560, 32]inqwen4exp) were dispatched to cuBLAS for decode-sized batches: a TF32 GEMM plus asplit-K reduction kernel, ~8 µs and two launches per call, 4 calls per layer.
ggml_cuda_should_use_mmvfalreadyaccepted N ≤ 16 for these types; this raises the limit to 64. The vector kernel does them in one launch with full
FP32 accumulation and is not slower for any of the shapes we measured on sm_80.
🤖 Generated with Claude Code