huggingface/transformers v5.10.1: Release v5.10.1
Sorry everyone, this happens when we rush a release!!!
Key points
- Gemma 4 12B Unified is an encoder-free multimodal model with pretrained and instruction-tuned variants.
- Unlike [standard Gemma 4](./gemma4), which uses dedicated encoder towers, Gemma 4 12B Unified projects raw inputs directly into the language model's embedding space through lightweight linear pipelines.
- Key differences from standard Gemma 4:
- Sapiens2 is a family of high-resolution vision transformers pretrained on ~1 billion curated human images, designed for human-centric computer vision tasks including pose estimation, body-part segmentation, surface normal estimation, and pointmap estimation.
Sources (1)
- [1]huggingface/transformers v5.10.1: Release v5.10.1GitHub: huggingface/transformers · Jun 3, 03:37 PM
Sorry everyone, this happens when we rush a release!!!
Gemma 4 12B Unified is an **encoder-free** multimodal model with pretrained and instruction-tuned variants.
Extractive summary: sentences quoted from the sources.