AION
Open-source releaseLarge Language Models1 source · Jun 10, 2026

huggingface/transformers v5.11.0: Release v5.11.0

DiffusionGemma is engineered to reduce the sequential bottlenecks of standard causal language models by employing an encoder-decoder architecture specifically optimized for inference speed.

Key points

  • During inference, DiffusionGemma leverages multi-canvas sampling, where rather than generating one token at a time, the model iteratively denoises a full block of tokens using a diffusion sampler.
  • This block-autoregressive approach facilitates text generation at higher speeds compared to traditional sequential generation methods.
  • DeepSeek-V3.2-Exp is an experimental model from DeepSeek-AI that introduces DeepSeek Sparse Attention (DSA), a trainable, fine-grained sparse attention mechanism designed to improve training and inference efficiency in long-context scenarios.
  • Fixed model parallel beam search bugs in the Qwen2-VL, Qwen2.5-VL, and Qwen3-VL MoE model families, and added documentation for tensor parallelism support with continuous batching.

Sources (1)

  • [1]huggingface/transformers v5.11.0: Release v5.11.0
    GitHub: huggingface/transformers · Jun 10, 04:32 PM
    DiffusionGemma is engineered to reduce the sequential bottlenecks of standard causal language models by employing an encoder-decoder architecture specifically optimized for inference speed.
    During inference, DiffusionGemma leverages multi-canvas sampling, where rather than generating one token at a time, the model iteratively denoises a full block of tokens using a diffusion sampler.

Extractive summary: sentences quoted from the sources.