huggingface/transformers v5.16.1: Release v5.16.1
This is a special release as we include GLM! (and a few small fixes)
Key points
- GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series.
- GLM-5.3-Flash starts from a newly trained base model, with its architecture and training recipe redesigned around capability and efficiency.
- For the first time in the GLM series, we introduce a hybrid architecture combining sparse and linear attention, sharply reducing long-context serving costs while preserving precise long-context capabilities.
- Together with our latest 30T-token multimodal pre-training corpus, these changes enable GLM-5.3-Flash to deliver more intelligence with less compute.
Sources (1)
- [1]huggingface/transformers v5.16.1: Release v5.16.1GitHub: huggingface/transformers · Aug 26, 02:50 PM
This is a special release as we include GLM! (and a few small fixes)
GLM-5.3-Flash, the first **natively multimodal model** in the GLM-5 series.
Extractive summary: sentences quoted from the sources.