Skip to content

perf(kernel): read dsv41 indexer key rows as int32 words - #634

Merged
jason-fxz merged 1 commit into
mainfrom
perf/dsv41-indexer-logits
Oct 8, 2026
Merged

jason-fxz merged 1 commit into
mainfrom
perf/dsv41-indexer-logits

Conversation

@jason-fxz

Copy link
Copy Markdown
Collaborator

Part of #629: fixes the Triton 3.8 slowdown of the DeepSeek-V4.1 indexer (_indexer_logits_packed_kernel, up to +102% on H100).

Call path: Indexer (DeepSeek-V4.1, prefill and decode) -> DSV41SparseAttnBackend.indexer_logits -> indexer_logits_packed -> _indexer_logits_packed_kernel. Full mode runs on the Full indexer layers, candidate mode on the Reindex layers.

On Triton 3.8 (sm_90) the [64 x 128] packed-row gather tile compiles with its threads along the channel axis (threadsPerWarp [1, 32], order [1, 0]) instead of along the rows ([32, 1], order [0, 1]) as on 3.6. Registers go from 165 to 200 and the loop from 2735 to 3981 SASS instructions.

Changes:

  • For the row format and size the model uses (fp4_e8m0_b32, head_dim 128: IDX_FMT and DeepSeek-V4.1's index_head_dim), _score_tile loads the key rows with a new _load_fp4_e8m0_rows: each row is read as 17 int32 words (16 words of e2m1 codes, 1 word of ue8m0 scales), the nibbles are unpacked in registers with tl.join, e2m1 is decoded with integer bit operations and an exact multiply by 2^126, and the scale is computed once per word. Other formats and head dims keep load_rows.
  • The wrapper asserts that the key pool is 4-byte aligned.

The gather tile now compiles to the same layout under both Triton versions; on sm_90 the kernel uses 128 registers and 1208 (3.6) / 1065 (3.8) SASS instructions in the loop.

Kernel time

µs, triton.testing.do_bench_cudagraph, best of two medians, GPU used by the benchmark only. old = main 377a3bc. 3.6 = torch 2.11.0 / Triton 3.6.0, 3.8 = torch 2.14.1 / Triton 3.8.0 (#630). H100 80GB HBM3 (driver 580.95.05) and RTX PRO 6000 Blackwell Server Edition (sm_120, driver 595.91.07). DeepSeek-V4.1-Flash shapes: 32 index heads, head_dim 128.

shape GPU old@3.6 old@3.8 new@3.6 new@3.8
full prefill, 2k chunk at 128k context H100 43971.67 88527.48 18142.69 17157.43
RTX PRO 6000 24792.20 35210.15 10533.40 10225.83
full prefill, 8k prompt H100 5359.19 9434.84 2291.21 2172.03
RTX PRO 6000 3124.66 4465.73 1369.60 1323.40
full decode B1, 128k context H100 43.57 64.53 17.90 16.38
RTX PRO 6000 25.76 31.89 12.57 11.98
candidate decode B1, 128k context H100 8.28 8.94 4.56 4.51
RTX PRO 6000 7.76 7.06 4.30 4.06

Over all 13 shapes, relative to old@3.6:

GPU old@3.8 new@3.6 new@3.8
H100 1.06 to 2.01 0.30 to 0.55 0.29 to 0.54
RTX PRO 6000 0.91 to 1.49 0.40 to 0.57 0.38 to 0.53
All shapes, H100
case old36 new36 old38 new38 new36/old36 old38/old36 new38/old36
full prefill n2048 start0 r2 [encoder Full, prompt 2k] 366.13 167.00 573.92 156.38 0.456 1.568 0.427
full prefill n8192 start0 r2 [encoder Full, prompt 8k] 5359.19 2291.21 9434.84 2172.03 0.428 1.760 0.405
full prefill n2048 start129024 r2 [encoder Full, chunk 2k at ctx 128k] 43971.67 18142.69 88527.48 17157.43 0.413 2.013 0.390
full prefill n128 start8064 r1 [decoder Full bounded, ctx 8k] 423.28 180.50 569.83 163.01 0.426 1.346 0.385
full prefill n128 start130944 r1 [decoder Full bounded, ctx 128k] 6721.91 2850.54 9147.68 2639.19 0.424 1.361 0.393
cand prefill n128 start8064 NC16384 [decoder Reindex bounded, ctx 8k] 667.73 213.40 904.15 207.36 0.320 1.354 0.311
cand prefill n128 start130944 NC16384 [decoder Reindex bounded, ctx 128k] 730.86 216.59 941.82 211.15 0.296 1.289 0.289
full decode B1 ctx32768 r2 staged1M [encoder Full] 11.68 5.97 17.35 5.47 0.511 1.485 0.468
full decode B4 ctx32768 r2 staged1M [encoder Full] 43.61 17.87 66.06 16.39 0.410 1.515 0.376
full decode B1 ctx131072 r2 staged1M [encoder Full] 43.57 17.90 64.53 16.38 0.411 1.481 0.376
full decode B1 ctx131072 r1 staged1M [decoder Full] 89.15 34.72 96.41 31.37 0.389 1.081 0.352
cand decode B1 ctx131072 NC16384 [decoder Reindex] 8.28 4.56 8.94 4.51 0.551 1.080 0.544
cand decode B4 ctx131072 NC16384 [decoder Reindex] 25.29 9.37 26.71 9.17 0.371 1.056 0.363
All shapes, RTX PRO 6000
case old36 new36 old38 new38 new36/old36 old38/old36 new38/old36
full prefill n2048 start0 r2 [encoder Full, prompt 2k] 243.27 106.60 328.15 101.38 0.438 1.349 0.417
full prefill n8192 start0 r2 [encoder Full, prompt 8k] 3124.66 1369.60 4465.73 1323.40 0.438 1.429 0.424
full prefill n2048 start129024 r2 [encoder Full, chunk 2k at ctx 128k] 24792.20 10533.40 35210.15 10225.83 0.425 1.420 0.412
full prefill n128 start8064 r1 [decoder Full bounded, ctx 8k] 361.67 145.47 333.74 139.52 0.402 0.923 0.386
full prefill n128 start130944 r1 [decoder Full bounded, ctx 128k] 5751.62 2288.80 5303.69 2197.98 0.398 0.922 0.382
cand prefill n128 start8064 NC16384 [decoder Reindex bounded, ctx 8k] 328.43 145.49 489.53 138.81 0.443 1.491 0.423
cand prefill n128 start130944 NC16384 [decoder Reindex bounded, ctx 128k] 376.64 150.44 494.17 142.88 0.399 1.312 0.379
full decode B1 ctx32768 r2 staged1M [encoder Full] 9.68 5.51 11.47 5.15 0.569 1.185 0.532
full decode B4 ctx32768 r2 staged1M [encoder Full] 24.96 11.60 30.85 10.91 0.465 1.236 0.437
full decode B1 ctx131072 r2 staged1M [encoder Full] 25.76 12.57 31.89 11.98 0.488 1.238 0.465
full decode B1 ctx131072 r1 staged1M [decoder Full] 43.00 19.82 47.05 18.79 0.461 1.094 0.437
cand decode B1 ctx131072 NC16384 [decoder Reindex] 7.76 4.30 7.06 4.06 0.555 0.910 0.523
cand decode B4 ctx131072 NC16384 [decoder Reindex] 15.72 7.18 17.95 6.82 0.457 1.142 0.434

Correctness and tests

  • For all 13 shapes the written output is bit-identical for old and new under both Triton versions, on both GPUs (inputs checked by md5 across the four runs).
  • Every code byte against every scale byte (256 x 256) dequantizes to the same bf16 bits as old under Triton 3.8 on H100. Under 3.6, code 0x8 gives -0.0 where old gives +0.0; none of the outputs above changes.
  • The e2m1 decode of code 1 multiplies an fp32 subnormal by 2^126. On NVIDIA it compiles to a non-ftz mul.f32 (checked in the PTX). Triton's AMD backend sets denormal-fp-math-f32="ieee" by default in 3.6 and 3.8; this was not run on ROCm.
  • tests/kernels/test_dsv41_indexer.py, test_dsv41_sparse_attn.py, test_dsv41_topk.py, test_dsv41_pack.py, tests/attention/test_dsv41_backend.py, tests/kvcache/test_dsv41_pool.py and tests/models/deepseek_v41: 94 passed under both versions on H100 and on the RTX PRO 6000. The model tests use index_head_dim 32 and run the unchanged path; test_dsv41_indexer.py and test_dsv41_backend.py run head_dim 128.

Not covered: end to end (no DeepSeek-V4.1 weights were available), timing on other GPUs (sm_80, sm_86, sm_89 and sm_100 were only compiled), ROCm.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant