AION
Technique

FlashAttention

Also known as: Flash Attention

3stories this week
3last 30 days
14all time

Timeline

  1. Oct 8, 2026 · Open-source release · 1 source
    huggingface/trl v1.15.0
    SFT, DPO, KTO, GRPO, RLOO and Distillation now score tokens with a fused LM head: a Triton kernel projects the hidden states through the LM head in tiles and reduces to per-token log-probs and entropy directly, so the [batch, seq, vocab] logits tensor is never built.
  2. Oct 6, 2026 · Open-source release · 1 source
    huggingface/transformers v5.19.0: Release v5.19.0
    EmbeddingGemma 2 is a multimodal embedding model from Google built on the Gemma 4 architecture.
  3. Oct 6, 2026 · Research paper · 1 source
    Cleave: Scaling Tensor Program Optimization via Decoupled Algebraic Search and Operator Scheduling
    We propose Cleave, an ML compiler built on symbolic decoupling: Cleave discovers transformations by performing superoptimization on a graph with symbolic shapes, and then schedules each resulting graph on concrete shapes.
  4. Aug 10, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.27.0
    Kimi K3 support with a full stack landing in one release: core model files and kernels (#50089, #50000), Python (#50093) and Rust (#50104) frontends, AttnRes kernels (#50090), DeepGEMM support (#50458), compressed-tensors quantized checkpoints (#50500), DSpark AR fusion (#50242), and an option to shard the shared expert instead of replicating it (#50656).
  5. Aug 10, 2026 · Open-source release · 1 source
    huggingface/transformers v5.15.0: Release: v5.15.0
    Muse Glimmer, released today, is Meta’s new multimodal model, especially designed for agentic use cases.
  6. Jul 15, 2026 · Open-source release · 1 source
    huggingface/transformers v5.14.0: Release v5.14.0
    Inkling is a general-purpose multimodal model that accepts text, image and audio inputs and
  7. Jul 11, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.25.0
    Building on quantized-model support from the previous release, MRv2 is now the standard execution path, with new support for EVS (#46535), realtime embeddings (#46762), prefix caching for Mamba hybrid models (#42406), multimodal-prefix bidirectional attention (#46942), and dynamic speculative decoding compatible with full CUDA graphs (#45953).
  8. Jun 29, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.24.0
    MiniMax-M3: Added support for the new MiniMax-M3 model (#45381), with a fast follow-on of BF16/FP8 indexer via MSA (#45892), MXFP4 support (#45896), FP8 sparse GQA (#45744), and extensive AMD/ROCm tuning — mxfp8 MoE/linear on gfx950 (#45725), fp8perchannel for bf16 weights on MI300X (#45854), FP8 KV-cache fix (#45720), and packed-modules mapping (#45794).
  9. Apr 27, 2026 · Open-source release · 1 source
    vllm-project/vllm v0.20.0
    CUDA 13.0 default: Default CUDA wheel on PyPI and vllm/vllm-openai:v0.20.0 image switched to CUDA 13.0; architecture lists and build-args cleaned up (#39878), and CUDA bumped to 13.0.2 to match PyTorch 2.11.0 (#40669).
  10. Apr 6, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.10
    Piecewise CUDA Graph Enabled by Default: Piecewise CUDA graph capture is now the default execution mode, reducing memory overhead and improving throughput for models with complex control flow patterns: #16331
  11. Mar 28, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.10rc0
    Piecewise CUDA Graph Enabled by Default: Piecewise CUDA graph capture is now the default execution mode, reducing memory overhead and improving throughput for models with complex control flow patterns: #16331
  12. Mar 23, 2026 · Open-source release · 1 source
    pytorch/pytorch v2.11.0: PyTorch 2.11.0 Release
    <strong>FlexAttention</strong> now has a <strong>FlashAttention-4</strong> backend on <strong>Hopper</strong> and <strong>Blackwell</strong> GPUs
  13. Jan 23, 2026 · Open-source release · 1 source
    sgl-project/sglang v0.5.8
    Qwen3-VL-Embedding & Qwen3-VL-Reranker model support: #16635, #16403
  14. Oct 17, 2024 · Open-source release · 1 source
    pytorch/pytorch v2.5.0: PyTorch 2.5.0 Release, SDPA CuDNN backend, Flex Attention
    We are excited to announce the release of PyTorch® 2.5!

Often appears with