Expert Coupling in MoE Pretraining: Reducing All-to-All Overhead with Correlated Placement and Token Shuffling
Mixture-of-Experts (MoE) layers replace the feed-forward block of a Transformer with E expert networks, and each token is routed to k of these experts.
Key points
- Under expert parallelism (EP) the experts are distributed across GPUs, and every MoE layer runs all-to-all collectives in the forward and backward passes to dispatch tokens to their experts and then combine the results.
- On a cluster with 8 AMD Instinct MI300X GPUs per node, these collectives can take 45% of the training step at EP32 with top-2 routing and 60% with top-6 routing.
- We find that early in pretraining routers have already learned to assign tokens to experts in correlated patterns, both within a layer and across layers.
- Token shuffling applies when sequence parallelism shards tokens across the EP group.
Sources (1)
- [1]Expert Coupling in MoE Pretraining: Reducing All-to-All Overhead with Correlated Placement and Token ShufflingarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 03:26 AM
Mixture-of-Experts (MoE) layers replace the feed-forward block of a Transformer with E expert networks, and each token is routed to k of these experts.
Under expert parallelism (EP) the experts are distributed across GPUs, and every MoE layer runs all-to-all collectives in the forward and backward passes to dispatch tokens to their experts and then combine the results.
Extractive summary: sentences quoted from the sources.