Shared Low-rank Basis Factorization for Data-free Mixture-of-Experts Compression
Mixture-of-Experts (MoE) large language models decouple capacity from compute through sparse routing, but their large parameter count creates storage and serving challenges.
ProofPaper ↗
Key points
- We analyze three MoE compression families: expert pruning, expert merging, and weight reconstruction, and derive structural error bounds showing that pruning and merging can incur non-vanishing errors tied to routing and expert heterogeneity.
- In contrast, weight reconstruction avoids these structural costs by preserving expert structure and routing.
- Motivated by the analysis, we propose Shared Low-rank Basis Factorization (SLBF), a data-free weight reconstruction method that uses rank-$k$ bases shared among experts, enabling richer cross-expert sharing, faster convergence, and lower reconstruction error.
- Across five MoE architectures spanning 16B to 122B parameters, SLBF consistently outperforms methods from all three compression families.
Sources (1)
- [1]Shared Low-rank Basis Factorization for Data-free Mixture-of-Experts CompressionarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 03:01 AM
Mixture-of-Experts (MoE) large language models decouple capacity from compute through sparse routing, but their large parameter count creates storage and serving challenges.
We analyze three MoE compression families: expert pruning, expert merging, and weight reconstruction, and derive structural error bounds showing that pruning and merging can incur non-vanishing errors tied to routing and expert heterogeneity.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 6, 2026[AINews] Reflection Beam - 501B-A23B American Open Model
- Oct 5, 2026Introducing GLM 5.3 on Amazon Bedrock
- Oct 5, 2026Sharing AI progress in mathematics
- Sep 28, 2026Holo4: powering generalist computer-use agents
- Jun 10, 2026DiffusionGemma: 4x faster text generation
- Jun 9, 2026Introducing Gemma 4 12B: a unified, encoder-free multimodal model