huggingface/transformers v5.17.0: Release 5.17.0
Hy4-Preview is a 780B-parameter mixture-of-experts language model that activates 49B parameters per
Key points
- Each MoE layer holds 256 routed experts plus one always-active shared expert and routes every
- The context window is 1M tokens.
- DeepSeek Sparse Attention (DSA) selects indextopk keys per query with a lightweight indexer.
- Independent Hyper-Connections (iHC) replace the plain residual path with hcmult parallel
Sources (1)
- [1]huggingface/transformers v5.17.0: Release 5.17.0GitHub: huggingface/transformers · Sep 9, 03:42 PM
Hy4-Preview is a 780B-parameter mixture-of-experts language model that activates 49B parameters per
Each MoE layer holds 256 routed experts plus one always-active shared expert and routes every
Extractive summary: sentences quoted from the sources.