AION
Research paperTraining & Scaling · Efficiency & Inference · Hardware & Compute1 source · Oct 7, 2026

Expert Coupling in MoE Pretraining: Reducing All-to-All Overhead with Correlated Placement and Token Shuffling

Mixture-of-Experts (MoE) layers replace the feed-forward block of a Transformer with E expert networks, and each token is routed to k of these experts.

Key points

  • Under expert parallelism (EP) the experts are distributed across GPUs, and every MoE layer runs all-to-all collectives in the forward and backward passes to dispatch tokens to their experts and then combine the results.
  • On a cluster with 8 AMD Instinct MI300X GPUs per node, these collectives can take 45% of the training step at EP32 with top-2 routing and 60% with top-6 routing.
  • We find that early in pretraining routers have already learned to assign tokens to experts in correlated patterns, both within a layer and across layers.
  • Token shuffling applies when sequence parallelism shards tokens across the EP group.

Sources (1)

Extractive summary: sentences quoted from the sources.