AION
Repository / library

Triton

3stories this week
7last 30 days
27all time

Timeline

  1. Oct 8, 2026 · Open-source release · 1 source
    huggingface/trl v1.15.0
    SFT, DPO, KTO, GRPO, RLOO and Distillation now score tokens with a fused LM head: a Triton kernel projects the hidden states through the LM head in tiles and reduces to per-token log-probs and entropy directly, so the [batch, seq, vocab] logits tensor is never built.
  2. Oct 6, 2026 · Research paper · 1 source
    PHBA: Prefix-State Hybrid Block Attention
    In this work, we introduce Prefix-State Hybrid Block Attention (PHBA), which replaces local sliding-window attention with top-k block-sparse retrieval and couples each retrieved block with a compact prefix state summarizing its preceding context.
  3. Oct 6, 2026 · Research paper · 1 source
    SpecFold: Folding Multi-Branch Redundancy for Faster Speculative Decoding in Diffusion Language Models
    Diffusion large language models (DLLMs) generate text through iterative block denoising, and multi-branch speculative decoding accelerates this process by verifying a main branch together with multiple draft branches in a single forward pass.
  4. Oct 2, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.21
    | Model | Type | Cookbook |
  5. Oct 1, 2026 · Product / feature launch · 1 source
    unslothai/unsloth v0.1.902-beta: Command Palette + Desktop UI/UX
    This release brings faster navigation, shareable run settings, and clearer errors to Unsloth Desktop.
  6. Sep 30, 2026 · Tutorial / explainer · 1 source
    Deploying an HSTU Generative Recommender with NVIDIA Dynamo-Triton
    Generative recommender (GR) systems are emerging as a powerful new approach for large-scale personalization.
  7. Sep 29, 2026 · Open-source release · 1 source
    NVIDIA/TensorRT-LLM v1.3.0rc29
    Expose Nemotron-H vision-language LoRA configuration for supported inference paths #19151
  8. Sep 9, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.29.0
    MRV2 also gained CUDA graph memory profiling for KV cache auto-sizing (#53306), batch-sharded sampling that cuts per-step logits memory by 1/TP (#50465), prompt embeds (#42963), extracthiddenstates speculation (#49811), padded FULL cudagraph dispatch for uniform decode under spec decode (#53407), and DP-sync skipping before EAGLE/MTP draft prefill (#53694).
  9. Sep 2, 2026 · Open-source release · 1 source
    pytorch/pytorch v2.14.0: PyTorch 2.14.0 Release
    <tr><td><strong>A preview of our rewritten NCCL backend for PyTorch</strong>, ported from torchcomms, implementing the full collective contract with nonblocking communicators and eager communicator splitting and advanced features such as fault tolerance and windows designed as a drop-in replacement of existing NCCL c10d backend</td></tr>
  10. Aug 22, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.18
    | Model | Type | PRs | Cookbook |
  11. Aug 10, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.27.0
    Kimi K3 support with a full stack landing in one release: core model files and kernels (#50089, #50000), Python (#50093) and Rust (#50104) frontends, AttnRes kernels (#50090), DeepGEMM support (#50458), compressed-tensors quantized checkpoints (#50500), DSpark AR fusion (#50242), and an option to shard the shared expert instead of replicating it (#50656).
  12. Jul 25, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.16
    DSpark: confidence-driven speculative decoding: A new speculative algorithm.
  13. Jul 15, 2026 · Open-source release · 1 source
    huggingface/transformers v5.14.0: Release v5.14.0
    Inkling is a general-purpose multimodal model that accepts text, image and audio inputs and
  14. Jul 11, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.25.0
    Building on quantized-model support from the previous release, MRv2 is now the standard execution path, with new support for EVS (#46535), realtime embeddings (#46762), prefix caching for Mamba hybrid models (#42406), multimodal-prefix bidirectional attention (#46942), and dynamic speculative decoding compatible with full CUDA graphs (#45953).
  15. Jul 8, 2026 · Open-source release · 1 source
    pytorch/pytorch v2.13.0: PyTorch 2.13.0 Release
    <tr><td><strong>torchcomms</strong>, a new communications backend for PyTorch Distributed, improves fault tolerance, scalability, and debuggability for large-cluster training.</td></tr>
  16. Jun 26, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.14
    Full release notes by category below.
  17. Jun 18, 2026 · Open-source release · 1 source
    pytorch/pytorch v2.12.1: PyTorch 2.12.1 Release, bug fix release
    This release is meant to fix the following regressions and silent correctness issues:
  18. Jun 10, 2026 · Open-source release · 1 source
    huggingface/transformers v5.11.0: Release v5.11.0
    DiffusionGemma is engineered to reduce the sequential bottlenecks of standard causal language models by employing an encoder-decoder architecture specifically optimized for inference speed.
  19. May 15, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.21.0
    Transformers v4 deprecated: This release formally deprecates transformers v4 support (#40389).
  20. Apr 6, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.10
    Piecewise CUDA Graph Enabled by Default: Piecewise CUDA graph capture is now the default execution mode, reducing memory overhead and improving throughput for models with complex control flow patterns: #16331
  21. Apr 3, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.19.0
    We recommend using pre-built docker image vllm/vllm-openai:gemma4 for out of box usage.
  22. Mar 28, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.10rc0
    Piecewise CUDA Graph Enabled by Default: Piecewise CUDA graph capture is now the default execution mode, reducing memory overhead and improving throughput for models with complex control flow patterns: #16331
  23. Jan 23, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.8
    Qwen3-VL-Embedding & Qwen3-VL-Reranker model support: #16635, #16403
  24. Jan 1, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.7
    [SGLang-Diffusion] Day 0 Support for Qwen-Image-Edit-2509, Qwen-Image-Edit-2511, Qwen-Image-2512 and Qwen-Image-Layered
  25. Apr 23, 2025 · Open-source release · 1 source
    pytorch/pytorch v2.7.0: PyTorch 2.7.0 Release
    <td>Torch.Compile support for Torch Function Modes
  26. Jan 29, 2025 · Open-source release · 1 source
    pytorch/pytorch v2.6.0: PyTorch 2.6.0 Release
    This release features multiple improvements for PT2: torch.compile can now be used with Python 3.13; new performance-related knob torch.compiler.setstance; several AOTInductor enhancements.
  27. Apr 24, 2024 · Open-source release · 1 source
    pytorch/pytorch v2.3.0: PyTorch 2.3: User-Defined Triton Kernels in torch.compile, Tensor Parallelism in Distributed
    PyTorch 2.3 offers support for user-defined Triton kernels in torch.compile, allowing for users to migrate their own Triton kernels from eager without experiencing performance complications or graph breaks.

Often appears with