pytorch/pytorch v2.5.0: PyTorch 2.5.0 Release, SDPA CuDNN backend, Flex Attention
We are excited to announce the release of PyTorch® 2.5!
Key points
- This release features a new CuDNN backend for SDPA, enabling speedups by default for users of SDPA on H100s or newer GPUs.
- As well, regional compilation of torch.compile offers a way to reduce the cold start up time for torch.compile by allowing users to compile a repeated nn.Module (e.g. a transformer layer in LLM) without recompilations.
- Finally, TorchInductor CPP backend offers solid performance speedup with numerous enhancements like FP16 support, CPP wrapper, AOT-Inductor mode, and max-autotune mode.
- This release is composed of 4095 commits from 504 contributors since PyTorch 2.4.
Sources (1)
- [1]pytorch/pytorch v2.5.0: PyTorch 2.5.0 Release, SDPA CuDNN backend, Flex AttentionGitHub: pytorch/pytorch · Oct 17, 04:26 PM
We are excited to announce the release of PyTorch® 2.5!
This release features a new CuDNN backend for SDPA, enabling speedups by default for users of SDPA on H100s or newer GPUs.
Extractive summary: sentences quoted from the sources.