AION
Open-source releaseLarge Language Models1 source · May 5, 2026

huggingface/transformers v5.8.0: Release 5.8.0

DeepSeek-V4 is the next-generation MoE (Mixture of Experts) language model from DeepSeek that introduces several architectural innovations over DeepSeek-V3.

Key points

  • The architecture replaces Multi-head Latent Attention (MLA) with a hybrid local + long-range attention design, swaps residual connections for Manifold-Constrained Hyper-Connections (mHC), and bootstraps the first few MoE layers with a static token-id → expert-id hash table.
  • This implementation covers DeepSeek-V4-Flash, DeepSeek-V4-Pro, and their -Base pretrained variants, which share the same architecture but differ in width, depth, expert count and weights.
  • Gemma 4 Assistant is a small, text-only model that enables speculative decoding for Gemma 4 models using the Multi-Token Prediction (MTP) method and associated candidate generator.
  • Granite Vision 4.1 is a vision-language model from IBM Research designed for enterprise-grade document data extraction.

Sources (1)

  • [1]huggingface/transformers v5.8.0: Release 5.8.0
    GitHub: huggingface/transformers · May 5, 04:52 PM
    DeepSeek-V4 is the next-generation MoE (Mixture of Experts) language model from DeepSeek that introduces several architectural innovations over DeepSeek-V3.
    The architecture replaces Multi-head Latent Attention (MLA) with a hybrid local + long-range attention design, swaps residual connections for Manifold-Constrained Hyper-Connections (mHC), and bootstraps the first few MoE layers with a static token-id → expert-id hash table.

Extractive summary: sentences quoted from the sources.