Skip to content

build(deps): move to torch 2.14.1, triton 3.8 and sglang-kernel 0.4.9 - #630

Merged
jason-fxz merged 1 commit into
mainfrom
build/torch-2.14
Oct 10, 2026
Merged

jason-fxz merged 1 commit into
mainfrom
build/torch-2.14

Conversation

@jason-fxz

@jason-fxz jason-fxz commented Oct 7, 2026 •

Copy link
Copy Markdown
Collaborator

Move the CUDA stack to torch 2.14.1, Triton 3.8 and sglang-kernel 0.4.9. sglang-kernel 0.4.9 requires torch==2.14.1, and torch 2.14.1 requires Triton 3.8. Tracking issue: #629.

  • pyproject.toml: torch>=2.14,<2.15, torchvision>=0.29,<0.30, triton>=3.6,<3.9, sglang-kernel==0.4.9. The build requirement, the kernel-cache build requirement and the torch pin in scripts/ci/manylinux-build.sh move with it.
  • setup.py: build the C++ extensions as C++20. With C++17 the torch 2.14 headers stop with #error C++20 or later compatible compiler is required. feat(rocm): run on RDNA3/RDNA4 with torch 2.14/2.15 and ROCm 7.14/10 #587 makes the same change.
  • PyPI's torch 2.14.1 is the cu130 build (it requires cuda-toolkit==13.0.3), so the install paths do not change. Comments that name the old versions are updated.

Triton 3.8 makes some kernels slower (#629). The fixes are on main and work under Triton 3.6 and 3.8: #631 (NVFP4 MoE decode), #632 (Triton extend attention and Gemma-4 sliding-window decode), #633 (DeepSeek-V4 sparse attention and gated pool), #634 (DeepSeek-V4.1 indexer) and #653 (block-FP8 MoE decode and gpt-oss MXFP4 GEMV). Under Triton 3.6, #632 makes three long-prefix extend shapes slower on H100 (gpt-oss full attention P28672 T4096: 1.37x of the kernel before #632); under Triton 3.8 they are faster than before. The regressions in the #629 plan that are not fixed are listed there.

Tested on H100 80GB (driver 580.95.05) with this branch rebased onto main 2d8b9cc. A = main on torch 2.11.0 / Triton 3.6.0 / sglang-kernel 0.4.5 / flashinfer 0.6.17, B = this branch on torch 2.14.1 / Triton 3.8.0 / sglang-kernel 0.4.9 / flashinfer 0.6.18.post1.

  • uv pip compile --extra accel resolves torch 2.14.1+cu130, triton 3.8.0, sglang-kernel 0.4.9 and flashinfer 0.6.18.post1.
  • The C++ extensions build with gcc 11.4 as C++20.
  • pytest tests -m "not slow": 2036 passed, 218 skipped and 1 failed on both A and B. The failure is a local checkpoint stub (apodex/Apodex-1.1-mini-NVFP4 has an index but no weights).

End to end on H100, A -> B. AIME is AIME25 problem 1 with greedy decoding and a 16384-token limit. TTFT and decode: an 8k-token and a 32k-token prompt, then 256 greedy tokens, median of 3. 1 and 8 requests: concurrent greedy decode on a ~2.5k-token prompt, 256 tokens each, median aggregate tok/s of 3 rounds after a warm-up round. Default MoE strategy unless marked; --max-running-requests 8 for the rows with AIME and --kv-reserve-tokens 40960 for Flash-Next and MiniMax. The last three rows ran only the long-prompt probe.

Model MoE / attention AIME, A / B TTFT 8k ms TTFT 32k ms Decode after 8k, tok/s 1 request, tok/s 8 requests, tok/s
openai/gpt-oss-20b offload / triton 70 / 70 540 -> 544 2569 -> 2318 233.6 -> 240.5 260.1 -> 259.7 713.1 -> 697.5
openai/gpt-oss-120b offload / triton 70 / 70 1260 -> 1263 5071 -> 4742 164.4 -> 165.6 182.2 -> 178.2 478.4 -> 474.0
Qwen/Qwen3.6-35B-A3B-FP8 offload / fa 70 (limit) / 70 1209 -> 1210 3040 -> 3044 197.2 -> 205.2 198.8 -> 207.2 756.4 -> 743.6
nvidia/Qwen3.6-35B-A3B-NVFP4 fused (forced) / fa 70 / 70 1619 -> 1264 6323 -> 4868 242.2 -> 244.3 245.2 -> 247.1 541.2 -> 626.5
nvidia/Gemma-4-26B-A4B-NVFP4 hybrid / triton 70 / 70 2194 -> 1718 9089 -> 7094 122.0 -> 121.8 129.0 -> 128.2 313.4 -> 336.1
Qwen/Qwen3-30B-A3B (bf16) hybrid / fa 70 / 70 2166 -> 2166 5407 -> 5412 92.8 -> 94.8 97.8 -> 97.6 233.8 -> 233.5
nvidia/Qwen3.6-27B-NVFP4 dense / fa 70 / 70 863 -> 858 3502 -> 3496 88.7 -> 91.9 89.9 -> 93.3 488.9 -> 481.3
RadixArk/Qwen3.8-27B-NVFP4 dense / fa 70 / 70 856 -> 852 3490 -> 3477 88.6 -> 91.9 89.9 -> 93.3 485.5 -> 478.1
Qwen3.8-27B-FP8 dense / fa 70 / 70 935 -> 935 3893 -> 3880 59.3 -> 70.4 60.0 -> 71.3 339.7 -> 333.4
nvidia/Gemma-4-31B-IT-NVFP4 dense / triton 70 / 70 1262 -> 1251 8602 -> 8511 55.6 -> 55.9 58.3 -> 58.4 332.0 -> 327.6
RedHatAI/Muse-Glimmer-30B-NVFP4 dense / triton 70 / none (limit) 766 -> 767 3349 -> 3310 81.2 -> 86.7 85.0 -> 90.0 620.6 -> 595.1
DeepSeek-V4-Flash-0731 hybrid / dsv4_sparse - / - 6195 -> 6012 19796 -> 19636 29.3 -> 30.8 - -
nvidia/Qwen3.8-Flash-Next-NVFP4 hybrid / qsa_sparse - / - 4894 -> 4164 15394 -> 12534 73.9 -> 75.1 - -
nvidia/MiniMax-M2.5-NVFP4 hybrid / fa - / - 8823 -> 6647 36293 -> 27607 49.8 -> 49.7 - -

B is 3% or more slower than A in one place: Muse-Glimmer-30B-NVFP4 at 4 and 8 concurrent requests, 315 -> 302 and 621 -> 595 tok/s. On main 0781324, an nsys trace of the decode CUDA graph put most of this difference in _nvfp4_gemm_kernel at small M, which #629 lists and which is not fixed.

Muse-Glimmer-30B-NVFP4 on B reached the token limit without a final answer. In the same run on main 0781324 (on 10-08), the Triton 3.6 run reached the limit and the Triton 3.8 run answered 70 in 1289 tokens. Qwen3.6-35B-A3B-FP8 on A reached the limit before it finished the final answer; the probe found 70 in its output.

Not covered:

  • ROCm. With torch>=2.14 the ROCm images with an older torch no longer satisfy the pin. feat(rocm): run on RDNA3/RDNA4 with torch 2.14/2.15 and ROCm 7.14/10 #587 sets the range to <2.16 and rewrites docs/install_amd.md, which still names the torch 2.11 image; the two PRs need one torch range.
  • A fresh resolve picks nvidia-cutlass-dsl 4.8.0; the runs above used 4.7.1.
  • The release wheel build in pytorch/manylinux2_28-builder:cuda13.0 was not run.
  • Windows, TP > 1, GPUs other than H100.

Follow-up: the torch < 2.12 check for the sm_89 _scaled_mm workaround in fp8_pertensor_linear.py is always false with this floor.

@jason-fxz
jason-fxz merged commit 74a3877 into main Oct 10, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant