Repository navigation
build(deps): move to torch 2.14.1, triton 3.8 and sglang-kernel 0.4.9 - #630
Merged
Merged
Conversation
9 of 18 tasks
This was referenced Oct 8, 2026
jason-fxz
marked this pull request as ready for review
October 10, 2026 00:23
jason-fxz
force-pushed
the
build/torch-2.14
branch
from
October 10, 2026 19:17
6e92be7 to
295fb1b
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Move the CUDA stack to torch 2.14.1, Triton 3.8 and sglang-kernel 0.4.9. sglang-kernel 0.4.9 requires
torch==2.14.1, and torch 2.14.1 requires Triton 3.8. Tracking issue: #629.pyproject.toml:torch>=2.14,<2.15,torchvision>=0.29,<0.30,triton>=3.6,<3.9,sglang-kernel==0.4.9. The build requirement, the kernel-cache build requirement and the torch pin inscripts/ci/manylinux-build.shmove with it.setup.py: build the C++ extensions as C++20. With C++17 the torch 2.14 headers stop with#error C++20 or later compatible compiler is required. feat(rocm): run on RDNA3/RDNA4 with torch 2.14/2.15 and ROCm 7.14/10 #587 makes the same change.cuda-toolkit==13.0.3), so the install paths do not change. Comments that name the old versions are updated.Triton 3.8 makes some kernels slower (#629). The fixes are on
mainand work under Triton 3.6 and 3.8: #631 (NVFP4 MoE decode), #632 (Triton extend attention and Gemma-4 sliding-window decode), #633 (DeepSeek-V4 sparse attention and gated pool), #634 (DeepSeek-V4.1 indexer) and #653 (block-FP8 MoE decode and gpt-oss MXFP4 GEMV). Under Triton 3.6, #632 makes three long-prefix extend shapes slower on H100 (gpt-oss full attention P28672 T4096: 1.37x of the kernel before #632); under Triton 3.8 they are faster than before. The regressions in the #629 plan that are not fixed are listed there.Tested on H100 80GB (driver 580.95.05) with this branch rebased onto
main2d8b9cc. A =mainon torch 2.11.0 / Triton 3.6.0 / sglang-kernel 0.4.5 / flashinfer 0.6.17, B = this branch on torch 2.14.1 / Triton 3.8.0 / sglang-kernel 0.4.9 / flashinfer 0.6.18.post1.uv pip compile --extra accelresolves torch 2.14.1+cu130, triton 3.8.0, sglang-kernel 0.4.9 and flashinfer 0.6.18.post1.pytest tests -m "not slow": 2036 passed, 218 skipped and 1 failed on both A and B. The failure is a local checkpoint stub (apodex/Apodex-1.1-mini-NVFP4has an index but no weights).End to end on H100, A -> B. AIME is AIME25 problem 1 with greedy decoding and a 16384-token limit. TTFT and decode: an 8k-token and a 32k-token prompt, then 256 greedy tokens, median of 3. 1 and 8 requests: concurrent greedy decode on a ~2.5k-token prompt, 256 tokens each, median aggregate tok/s of 3 rounds after a warm-up round. Default MoE strategy unless marked;
--max-running-requests 8for the rows with AIME and--kv-reserve-tokens 40960for Flash-Next and MiniMax. The last three rows ran only the long-prompt probe.B is 3% or more slower than A in one place: Muse-Glimmer-30B-NVFP4 at 4 and 8 concurrent requests, 315 -> 302 and 621 -> 595 tok/s. On
main0781324, an nsys trace of the decode CUDA graph put most of this difference in_nvfp4_gemm_kernelat small M, which #629 lists and which is not fixed.Muse-Glimmer-30B-NVFP4 on B reached the token limit without a final answer. In the same run on
main0781324 (on 10-08), the Triton 3.6 run reached the limit and the Triton 3.8 run answered 70 in 1289 tokens. Qwen3.6-35B-A3B-FP8 on A reached the limit before it finished the final answer; the probe found 70 in its output.Not covered:
torch>=2.14the ROCm images with an older torch no longer satisfy the pin. feat(rocm): run on RDNA3/RDNA4 with torch 2.14/2.15 and ROCm 7.14/10 #587 sets the range to<2.16and rewritesdocs/install_amd.md, which still names the torch 2.11 image; the two PRs need one torch range.nvidia-cutlass-dsl4.8.0; the runs above used 4.7.1.pytorch/manylinux2_28-builder:cuda13.0was not run.Follow-up: the
torch < 2.12check for the sm_89_scaled_mmworkaround infp8_pertensor_linear.pyis always false with this floor.