huggingface/transformers v5.16.0: Release: v5.16.0
Qwen4-Exp builds on Qwen3.5's hybrid text and multimodal architecture with three key components: GatedResidual (GR), Qwen Sparse Attention (QSA), and Per-Layer Embedding (PLE).
Key points
- GR is a Qwen-developed residual architecture that combines Hyper-Connection with GatedNorm.
- QSA uses multiple query heads to score compressed key blocks, selects the most relevant contiguous token blocks, and keeps the incomplete trailing block uncompressed.
- Combined with Gated DeltaNet, QSA makes Qwen4-Exp the first hybrid architecture to integrate linear and sparse attention, substantially improving inference efficiency for long-context workloads.
- Block attention with Shaw's relative positional embeddings: attention is computed over fixed-size blocks (the sequence is right-padded to a whole number of blocks, with padded frames masked out), using separate bias-free query/key/value projections.
Sources (1)
- [1]huggingface/transformers v5.16.0: Release: v5.16.0GitHub: huggingface/transformers · Aug 26, 12:35 PM
Qwen4-Exp builds on Qwen3.5's hybrid text and multimodal architecture with three key components: GatedResidual (GR), Qwen Sparse Attention (QSA), and Per-Layer Embedding (PLE).
GR is a Qwen-developed residual architecture that combines Hyper-Connection with GatedNorm.
Extractive summary: sentences quoted from the sources.