huggingface/transformers v5.9.0: Release v5.9.0
Command A+ is a Mixture-of-Experts (MoE) language model from Cohere that features a hybrid attention pattern combining sliding window and full attention layers.
Key points
- The model incorporates both shared and routed experts and supports a very large context window for processing extensive text sequences.
- HRM-Text is an improved autoregressive language-modeling variant of the Hierarchical Reasoning Model (HRM) that uses a hierarchical recurrent forward pass with two transformer stacks - one for slow, abstract planning (H) and one for fast, detailed computation (L) - reused inside a nested recurrence.
- The textembeds input for SAM3, EdgeTAM, and SAM3-Lite-Text models now expects full text embeddings instead of just pooler outputs, aligning with other models in the library — users must update their inputs accordingly.
- 🚨Fix memory leaks caused by lru decorators in vision models (#45922) by @yonigozlan
Sources (1)
- [1]huggingface/transformers v5.9.0: Release v5.9.0GitHub: huggingface/transformers · May 20, 02:12 PM
Command A+ is a Mixture-of-Experts (MoE) language model from Cohere that features a hybrid attention pattern combining sliding window and full attention layers.
The model incorporates both shared and routed experts and supports a very large context window for processing extensive text sequences.
Extractive summary: sentences quoted from the sources.