AION
Open-source releaseLarge Language Models · Multimodal Models · Image, Video & 3D Generation1 source · Jun 12, 2026

huggingface/transformers v5.12.0: Release v5.12.0

MiniMax-M3-VL is the vision-language member of the MiniMax-M3 family that pairs a CLIP-style vision tower with 3D rotary position embeddings with the MiniMax-M3 text backbone.

Key points

  • The model processes images through a Conv3d patch embedding system and includes specialized components for efficient multimodal understanding and generation.
  • The official weights for PP-OCRv6 are out: PP-OCRv6 is a lightweight OCR system that combines architectural innovation with data-centric optimization.
  • PP-OCRv6: update documentation and slow tests (#46576) by @ zhang-prog
  • Greedy transducer decoding for inference: a blank emission advances the encoder frame by one, a non-blank emission stays on the same frame.

Sources (1)

  • [1]huggingface/transformers v5.12.0: Release v5.12.0
    GitHub: huggingface/transformers · Jun 12, 02:39 PM
    MiniMax-M3-VL is the vision-language member of the MiniMax-M3 family that pairs a CLIP-style vision tower with 3D rotary position embeddings with the MiniMax-M3 text backbone.
    The model processes images through a Conv3d patch embedding system and includes specialized components for efficient multimodal understanding and generation.

Extractive summary: sentences quoted from the sources.

Before this

  1. Jun 10, 2026huggingface/transformers v5.11.0: Release v5.11.0
  2. Jun 9, 2026Introducing Gemma 4 12B: a unified, encoder-free multimodal model
  3. Jun 3, 2026huggingface/transformers v5.10.1: Release v5.10.1
  4. May 20, 2026huggingface/transformers v5.9.0: Release v5.9.0
  5. May 5, 2026huggingface/transformers v5.8.0: Release 5.8.0
  6. Apr 3, 2026vllm-project/vllm v0.19.0

Related