Triton
3stories this week
7last 30 days
27all time
Timeline
- Oct 8, 2026 · Open-source release · 1 sourcehuggingface/trl v1.15.0SFT, DPO, KTO, GRPO, RLOO and Distillation now score tokens with a fused LM head: a Triton kernel projects the hidden states through the LM head in tiles and reduces to per-token log-probs and entropy directly, so the [batch, seq, vocab] logits tensor is never built.
- Oct 6, 2026 · Research paper · 1 sourcePHBA: Prefix-State Hybrid Block AttentionIn this work, we introduce Prefix-State Hybrid Block Attention (PHBA), which replaces local sliding-window attention with top-k block-sparse retrieval and couples each retrieved block with a compact prefix state summarizing its preceding context.
- Oct 6, 2026 · Research paper · 1 sourceSpecFold: Folding Multi-Branch Redundancy for Faster Speculative Decoding in Diffusion Language ModelsDiffusion large language models (DLLMs) generate text through iterative block denoising, and multi-branch speculative decoding accelerates this process by verifying a main branch together with multiple draft branches in a single forward pass.
- Oct 2, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.21| Model | Type | Cookbook |
- Oct 1, 2026 · Product / feature launch · 1 sourceunslothai/unsloth v0.1.902-beta: Command Palette + Desktop UI/UXThis release brings faster navigation, shareable run settings, and clearer errors to Unsloth Desktop.
- Sep 30, 2026 · Tutorial / explainer · 1 sourceDeploying an HSTU Generative Recommender with NVIDIA Dynamo-TritonGenerative recommender (GR) systems are emerging as a powerful new approach for large-scale personalization.
- Sep 29, 2026 · Open-source release · 1 sourceNVIDIA/TensorRT-LLM v1.3.0rc29Expose Nemotron-H vision-language LoRA configuration for supported inference paths #19151
- Sep 9, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.29.0MRV2 also gained CUDA graph memory profiling for KV cache auto-sizing (#53306), batch-sharded sampling that cuts per-step logits memory by 1/TP (#50465), prompt embeds (#42963), extracthiddenstates speculation (#49811), padded FULL cudagraph dispatch for uniform decode under spec decode (#53407), and DP-sync skipping before EAGLE/MTP draft prefill (#53694).
- Sep 2, 2026 · Open-source release · 1 sourcepytorch/pytorch v2.14.0: PyTorch 2.14.0 Release<tr><td><strong>A preview of our rewritten NCCL backend for PyTorch</strong>, ported from torchcomms, implementing the full collective contract with nonblocking communicators and eager communicator splitting and advanced features such as fault tolerance and windows designed as a drop-in replacement of existing NCCL c10d backend</td></tr>
- Aug 22, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.18| Model | Type | PRs | Cookbook |
- Aug 10, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.27.0Kimi K3 support with a full stack landing in one release: core model files and kernels (#50089, #50000), Python (#50093) and Rust (#50104) frontends, AttnRes kernels (#50090), DeepGEMM support (#50458), compressed-tensors quantized checkpoints (#50500), DSpark AR fusion (#50242), and an option to shard the shared expert instead of replicating it (#50656).
- Jul 25, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.16DSpark: confidence-driven speculative decoding: A new speculative algorithm.
- Jul 15, 2026 · Open-source release · 1 sourcehuggingface/transformers v5.14.0: Release v5.14.0Inkling is a general-purpose multimodal model that accepts text, image and audio inputs and
- Jul 11, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.25.0Building on quantized-model support from the previous release, MRv2 is now the standard execution path, with new support for EVS (#46535), realtime embeddings (#46762), prefix caching for Mamba hybrid models (#42406), multimodal-prefix bidirectional attention (#46942), and dynamic speculative decoding compatible with full CUDA graphs (#45953).
- Jul 8, 2026 · Open-source release · 1 sourcepytorch/pytorch v2.13.0: PyTorch 2.13.0 Release<tr><td><strong>torchcomms</strong>, a new communications backend for PyTorch Distributed, improves fault tolerance, scalability, and debuggability for large-cluster training.</td></tr>
- Jun 26, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.14Full release notes by category below.
- Jun 18, 2026 · Open-source release · 1 sourcepytorch/pytorch v2.12.1: PyTorch 2.12.1 Release, bug fix releaseThis release is meant to fix the following regressions and silent correctness issues:
- Jun 10, 2026 · Open-source release · 1 sourcehuggingface/transformers v5.11.0: Release v5.11.0DiffusionGemma is engineered to reduce the sequential bottlenecks of standard causal language models by employing an encoder-decoder architecture specifically optimized for inference speed.
- May 15, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.21.0Transformers v4 deprecated: This release formally deprecates transformers v4 support (#40389).
- Apr 6, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.10Piecewise CUDA Graph Enabled by Default: Piecewise CUDA graph capture is now the default execution mode, reducing memory overhead and improving throughput for models with complex control flow patterns: #16331
- Apr 3, 2026 · Open-source release · 1 sourcevllm-project/vllm v0.19.0We recommend using pre-built docker image vllm/vllm-openai:gemma4 for out of box usage.
- Mar 28, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.10rc0Piecewise CUDA Graph Enabled by Default: Piecewise CUDA graph capture is now the default execution mode, reducing memory overhead and improving throughput for models with complex control flow patterns: #16331
- Jan 23, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.8Qwen3-VL-Embedding & Qwen3-VL-Reranker model support: #16635, #16403
- Jan 1, 2026 · Open-source release · 1 sourcesgl-project/sglang v0.5.7[SGLang-Diffusion] Day 0 Support for Qwen-Image-Edit-2509, Qwen-Image-Edit-2511, Qwen-Image-2512 and Qwen-Image-Layered
- Apr 23, 2025 · Open-source release · 1 sourcepytorch/pytorch v2.7.0: PyTorch 2.7.0 Release<td>Torch.Compile support for Torch Function Modes
- Jan 29, 2025 · Open-source release · 1 sourcepytorch/pytorch v2.6.0: PyTorch 2.6.0 ReleaseThis release features multiple improvements for PT2: torch.compile can now be used with Python 3.13; new performance-related knob torch.compiler.setstance; several AOTInductor enhancements.
- Apr 24, 2024 · Open-source release · 1 sourcepytorch/pytorch v2.3.0: PyTorch 2.3: User-Defined Triton Kernels in torch.compile, Tensor Parallelism in DistributedPyTorch 2.3 offers support for user-defined Triton kernels in torch.compile, allowing for users to migrate their own Triton kernels from eager without experiencing performance complications or graph breaks.