Skip to content

perf(kernel): keep 16-byte kv loads in dsv4 sparse attention and unroll gated pool - #633

Merged
jason-fxz merged 1 commit into
mainfrom
perf/dsv4-sparse-attn-loads
Oct 8, 2026
Merged

jason-fxz merged 1 commit into
mainfrom
perf/dsv4-sparse-attn-loads

Conversation

@jason-fxz

@jason-fxz jason-fxz commented Oct 8, 2026 •

Copy link
Copy Markdown
Collaborator

Part of #629: fixes the Triton 3.8 slowdowns of DeepSeek-V4 sparse attention and of the compressor's gated pool.

Call paths:

  • DSV4SparseAttnBackend.attend -> sparse_attn_paged -> _sparse_attn_paged_kernel (prefill) or _sparse_attn_paged_splitk_kernel (decode)
  • Compressor.decode_step (DeepSeek-V4) and Compressor.forward_decode (DeepSeek-V4.1) -> gated_pool -> _gated_pool_kernel

Changes:

  • sparse_attn_paged: the per-column pool base tl.where(is_win, win_ptr, cmp_ptr) is wrapped in tl.multiple_of(..., 16) in both kernels, and the wrapper asserts that both pools are 16-byte aligned. On Triton 3.8 the KV gather otherwise compiles to 2-byte ld.global loads (64 in the PTX) and 82048 B of shared memory; with the hint it compiles to 16-byte cp.async and 67968 B under both versions. Under Triton 3.6 the new prefill kernel compiles to the same SASS as the old one on sm_80, sm_86, sm_89, sm_90, sm_100 and sm_120.
  • gated_pool: both loops over R use tl.range(R, loop_unroll_factor=2). On Triton 3.8 the max loop issues 1 load per iteration in SASS (4 on 3.6); with the change it issues 8 under both versions.

The pools passed to sparse_attn_paged are allocated per layer with torch.zeros in DSV4PagedKVCache._alloc_buffers. A misaligned pool, which old code accepted, now fails the assert.

Kernel time

µs, triton.testing.do_bench_cudagraph, best of two medians, GPU used by the benchmark only. old = main 377a3bc. 3.6 = torch 2.11.0 / Triton 3.6.0, 3.8 = torch 2.14.1 / Triton 3.8.0 (#630). H100 80GB HBM3 (driver 580.95.05) and RTX PRO 6000 Blackwell Server Edition (sm_120, driver 595.91.07). DeepSeek-V4 shapes: 64 heads, head_dim 512, 128-token window, index top-k 512; r = compress ratio.

shape GPU old@3.6 old@3.8 new@3.6 new@3.8
sparse prefill r4 S8192 H100 11615.80 15629.07 11565.03 11399.31
RTX PRO 6000 9932.09 12595.64 9931.07 9665.91
sparse decode r4 B1 at position 32767 H100 17.46 21.19 17.34 16.83
RTX PRO 6000 16.39 18.92 16.39 15.74
gated_pool B1 R8 D512 H100 2.54 3.26 2.00 1.99
RTX PRO 6000 2.23 2.82 2.02 2.02
gated_pool B1 R128 D512 H100 21.09 36.89 17.03 17.05
RTX PRO 6000 16.58 31.32 18.62 18.59

Over the 11 sparse attention shapes, new@3.8 is 0.94 to 1.00 of old@3.6 on H100 and 0.95 to 1.005 on the RTX PRO 6000. gated_pool at R = 8 is 0.78 to 0.81 of old@3.6 on H100 and 0.89 to 0.91 on the RTX PRO 6000, under both versions. gated_pool at R = 128 (the KV compressor of the ratio-128 layers) is 0.81 to 0.83 of old@3.6 on H100, and 1.12 to 1.13 on the RTX PRO 6000 under both versions (1.89 to 1.99 for old@3.8); three separate runs on the RTX PRO 6000 gave the same ratios.

All shapes, H100
case old36 new36 old38 new38 new36/old36 old38/old36 new38/old36
sparse prefill r4 S2048 2922.06 2922.38 3911.67 2887.86 1.000 1.339 0.988
sparse prefill r4 S8192 11615.80 11565.03 15629.07 11399.31 0.996 1.346 0.981
sparse prefill r4 S4096@28672 5850.94 5846.37 7931.12 5762.74 0.999 1.356 0.985
sparse prefill r128 S8192 3683.63 3707.31 4902.27 3672.89 1.006 1.331 0.997
sparse prefill r0 S8192 2580.88 2578.67 3409.59 2561.59 0.999 1.321 0.993
sparse decode r4 B1 pos32767 17.46 17.34 21.19 16.83 0.993 1.214 0.964
sparse decode r128 B1 pos32767 16.83 16.83 20.55 15.98 1.000 1.220 0.949
sparse decode r0 B1 pos32767 13.58 13.55 16.81 13.37 0.998 1.238 0.984
sparse decode r4 B8 pos32767 22.24 22.23 26.30 20.81 1.000 1.182 0.936
sparse decode r128 B8 pos32767 18.42 18.33 21.87 17.73 0.995 1.188 0.963
sparse decode r0 B8 pos32767 14.03 14.03 17.09 13.89 1.000 1.218 0.990
gated_pool B1 R8 D128 2.53 1.99 3.24 1.97 0.788 1.281 0.780
gated_pool B1 R8 D512 2.54 2.00 3.26 1.99 0.787 1.284 0.786
gated_pool B1 R128 D128 18.39 15.22 35.58 15.23 0.827 1.934 0.828
gated_pool B1 R128 D512 21.09 17.03 36.89 17.05 0.808 1.749 0.809
gated_pool B8 R8 D128 2.56 2.03 3.34 2.03 0.796 1.308 0.795
gated_pool B8 R8 D512 2.59 2.09 3.34 2.08 0.805 1.290 0.803
gated_pool B8 R128 D128 18.60 15.45 36.14 15.46 0.830 1.943 0.831
gated_pool B8 R128 D512 19.91 16.44 36.82 16.43 0.826 1.849 0.825
All shapes, RTX PRO 6000
case old36 new36 old38 new38 new36/old36 old38/old36 new38/old36
sparse prefill r4 S2048 2217.05 2216.76 2966.29 2206.47 1.000 1.338 0.995
sparse prefill r4 S8192 9932.09 9931.07 12595.64 9665.91 1.000 1.268 0.973
sparse prefill r4 S4096@28672 5006.00 5006.10 6303.13 4901.00 1.000 1.259 0.979
sparse prefill r128 S8192 3225.48 3221.67 4022.34 3143.03 0.999 1.247 0.974
sparse prefill r0 S8192 2309.92 2305.40 2817.86 2283.15 0.998 1.220 0.988
sparse decode r4 B1 pos32767 16.39 16.39 18.92 15.74 1.000 1.154 0.960
sparse decode r128 B1 pos32767 15.91 15.91 18.42 15.27 1.000 1.158 0.960
sparse decode r0 B1 pos32767 13.29 13.29 16.33 13.30 1.000 1.229 1.001
sparse decode r4 B8 pos32767 17.15 16.89 19.55 16.64 0.985 1.140 0.970
sparse decode r128 B8 pos32767 16.75 16.09 18.56 15.85 0.960 1.108 0.946
sparse decode r0 B8 pos32767 13.33 13.34 16.40 13.39 1.000 1.230 1.004
gated_pool B1 R8 D128 2.13 1.91 2.75 1.91 0.894 1.290 0.897
gated_pool B1 R8 D512 2.23 2.02 2.82 2.02 0.905 1.266 0.907
gated_pool B1 R128 D128 15.54 17.43 30.87 17.52 1.122 1.986 1.127
gated_pool B1 R128 D512 16.58 18.62 31.32 18.59 1.123 1.889 1.121
gated_pool B8 R8 D128 2.07 1.84 2.71 1.84 0.886 1.310 0.889
gated_pool B8 R8 D512 2.15 1.92 2.76 1.92 0.893 1.285 0.896
gated_pool B8 R128 D128 15.87 17.86 31.56 17.82 1.126 1.989 1.123
gated_pool B8 R128 D512 16.83 18.87 32.27 18.93 1.121 1.918 1.125

Correctness and tests

  • For all 19 shapes, the sha256 of the full output is the same for old and new under both Triton versions, on both GPUs.
  • Old vs new on H100 under both versions is bit-identical over the 19 table shapes plus 24 edge cases: gated_pool with R = 1/3/5/7/9/16/127/129, D = 300/100/64, bf16 and fp32 output and strided views; sparse attention with 20/16/128 heads, windows of 100 and 50, top-k widths that are not tile multiples, zero counts, the ratio-0 alias and pool views with a row offset.
  • tests/kernels/test_dsv4_sparse_attn.py, tests/dsv4/test_dsv4_attn_backend.py, tests/kernels/test_dsv4_indexer_decode.py, tests/kvcache/test_dsv4_pool.py, tests/scheduler/test_dsv4_generic_manager.py, the DeepSeek-V4.1 kernel, attention and pool tests, and tests/models/deepseek_v41: 172 passed under both versions on H100 and on the RTX PRO 6000.

End to end

DeepSeek-V4-Flash on H100 with the default (hybrid) MoE strategy, AIME25 problem 1 (single sampled runs; old@3.8 is main with the dependency change of #630):

old@3.6 old@3.8 new@3.8
answer 70 70 70
completion tokens 348 568 568
decode tok/s 27.4 28.8 28.3
TTFT, 2k-token prompt (ms) 2987 3016 3024

Not covered: end to end on the RTX PRO 6000 (no weights there), timing on other GPUs (sm_80, sm_86, sm_89 and sm_100 were only compiled), ROCm. DeepSeek-V4.1 sparse attention (dsv41/sparse_attn.py) is not changed.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant