pytorch/pytorch v2.14.0: PyTorch 2.14.0 Release
<tr><td><strong>A preview of our rewritten NCCL backend for PyTorch</strong>, ported from torchcomms, implementing the full collective contract with nonblocking communicators and eager communicator splitting and advanced features such as fault tolerance and windows designed as a drop-in replacement of existing NCCL c10d backend</td></tr>
Key points
- <tr><td><strong>NVGEMM</strong> brings CuTeDSL-generated CUTLASS kernels to Inductor, with epilogue fusion, scaled and NVFP4 GEMM, and grouped-reduction epilogues autotuned alongside Triton and ATen</td></tr>
- <tr><td><strong><code>torch.switch</code></strong> generalizes <code>torch.cond</code> to multi-way branching, and <code>torch.whileloop</code> can now be captured in a CUDA graph</td></tr>
- <tr><td><strong>Experimental <code>torch.compile</code> support for complex-valued tensors</strong>: Opt-in support decomposes supported complex operations into real and imaginary computations, enabling compiler backends to optimize more complex-number workloads.</td></tr>
- <tr><td><strong>Broader platform support</strong>: ROCm 7.14 wheels are produced from the TheRock pip SDK, Intel XPU adds native graph capture, and Inductor targets Rubin (<code>sm107</code>)</td></tr>
Sources (1)
- [1]pytorch/pytorch v2.14.0: PyTorch 2.14.0 ReleaseGitHub: pytorch/pytorch · Sep 2, 05:40 PM
<tr><td><strong>A preview of our rewritten NCCL backend for PyTorch</strong>, ported from torchcomms, implementing the full collective contract with nonblocking communicators and eager communicator splitting and advanced features such as fault tolerance and windows designed as a drop-in replacement of existing NCCL c10d backend</td></tr>
<tr><td><strong>NVGEMM</strong> brings CuTeDSL-generated CUTLASS kernels to Inductor, with epilogue fusion, scaled and NVFP4 GEMM, and grouped-reduction epilogues autotuned alongside Triton and ATen</td></tr>
Extractive summary: sentences quoted from the sources.