AION
Research paperTraining & Scaling1 source · Oct 8, 2026

Smoothing the Top-k Exposure Boundary for Sparse Mixture-of-Experts

Sparse Mixture-of-Experts models scale parameter capacity efficiently while maintaining a fixed compute budget per token.

Key points

  • However, traditional training paradigms enforce a static choice of top-$k$ experts, which converts a continuous routing distribution into a rigid step function.
  • This constraint introduces a brittle boundary where highly competitive experts are arbitrarily separated into full-supervision and zero-feedback zones based on minor score fluctuations.
  • To address this issue, we propose Elastic Expert Routing, which stochastically samples the active expert budget from a localized discrete distribution centered at $k$.
  • During supervised fine-tuning, elastic routing improves downstream macro-averages on OLMoE-1B-7B and Qwen3-30B-A3B by $+0.84$ and $+2.02$ points, respectively.

Sources (1)

  • [1]Smoothing the Top-k Exposure Boundary for Sparse Mixture-of-Experts
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 09:31 AM
    Sparse Mixture-of-Experts models scale parameter capacity efficiently while maintaining a fixed compute budget per token.
    However, traditional training paradigms enforce a static choice of top-$k$ experts, which converts a continuous routing distribution into a rigid step function.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 8, 2026DivMoE: Fine-Grained MoE Upcycling via Cross-Domain Expert Composition
  2. Oct 8, 2026VibeEdit: Image Editing with Canvas Instructions
  3. Oct 8, 2026SpaceCast-Bench: Evaluating Predictive Spatial Reasoning in Vision-Language Models
  4. Oct 8, 2026Opera: A Verbal Critic Framework for Long-horizon Coding Agents
  5. Oct 7, 2026Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
  6. Oct 6, 2026huggingface/transformers v5.19.0: Release v5.19.0

Related