Smoothing the Top-k Exposure Boundary for Sparse Mixture-of-Experts
Sparse Mixture-of-Experts models scale parameter capacity efficiently while maintaining a fixed compute budget per token.
Key points
- However, traditional training paradigms enforce a static choice of top-$k$ experts, which converts a continuous routing distribution into a rigid step function.
- This constraint introduces a brittle boundary where highly competitive experts are arbitrarily separated into full-supervision and zero-feedback zones based on minor score fluctuations.
- To address this issue, we propose Elastic Expert Routing, which stochastically samples the active expert budget from a localized discrete distribution centered at $k$.
- During supervised fine-tuning, elastic routing improves downstream macro-averages on OLMoE-1B-7B and Qwen3-30B-A3B by $+0.84$ and $+2.02$ points, respectively.
Sources (1)
- [1]Smoothing the Top-k Exposure Boundary for Sparse Mixture-of-ExpertsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 09:31 AM
Sparse Mixture-of-Experts models scale parameter capacity efficiently while maintaining a fixed compute budget per token.
However, traditional training paradigms enforce a static choice of top-$k$ experts, which converts a continuous routing distribution into a rigid step function.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 8, 2026DivMoE: Fine-Grained MoE Upcycling via Cross-Domain Expert Composition
- Oct 8, 2026VibeEdit: Image Editing with Canvas Instructions
- Oct 8, 2026SpaceCast-Bench: Evaluating Predictive Spatial Reasoning in Vision-Language Models
- Oct 8, 2026Opera: A Verbal Critic Framework for Long-horizon Coding Agents
- Oct 7, 2026Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
- Oct 6, 2026huggingface/transformers v5.19.0: Release v5.19.0