huggingface/transformers v5.12.0: Release v5.12.0
MiniMax-M3-VL is the vision-language member of the MiniMax-M3 family that pairs a CLIP-style vision tower with 3D rotary position embeddings with the MiniMax-M3 text backbone.
Key points
- The model processes images through a Conv3d patch embedding system and includes specialized components for efficient multimodal understanding and generation.
- The official weights for PP-OCRv6 are out: PP-OCRv6 is a lightweight OCR system that combines architectural innovation with data-centric optimization.
- PP-OCRv6: update documentation and slow tests (#46576) by @ zhang-prog
- Greedy transducer decoding for inference: a blank emission advances the encoder frame by one, a non-blank emission stays on the same frame.
Sources (1)
- [1]huggingface/transformers v5.12.0: Release v5.12.0GitHub: huggingface/transformers · Jun 12, 02:39 PM
MiniMax-M3-VL is the vision-language member of the MiniMax-M3 family that pairs a CLIP-style vision tower with 3D rotary position embeddings with the MiniMax-M3 text backbone.
The model processes images through a Conv3d patch embedding system and includes specialized components for efficient multimodal understanding and generation.
Extractive summary: sentences quoted from the sources.
Before this
- Jun 10, 2026huggingface/transformers v5.11.0: Release v5.11.0
- Jun 9, 2026Introducing Gemma 4 12B: a unified, encoder-free multimodal model
- Jun 3, 2026huggingface/transformers v5.10.1: Release v5.10.1
- May 20, 2026huggingface/transformers v5.9.0: Release v5.9.0
- May 5, 2026huggingface/transformers v5.8.0: Release 5.8.0
- Apr 3, 2026vllm-project/vllm v0.19.0