Megatron-LM
Also known as: Megatron
2stories this week
2last 30 days
2all time
Timeline
- Oct 7, 2026 · Research paper · 1 sourceExpert Coupling in MoE Pretraining: Reducing All-to-All Overhead with Correlated Placement and Token ShufflingMixture-of-Experts (MoE) layers replace the feed-forward block of a Transformer with E expert networks, and each token is routed to k of these experts.
- Oct 6, 2026 · Opinion / analysis · 1 sourceScale Bitwise-Deterministic Pretraining with NVIDIA Megatron CoreBitwise determinism makes large-scale pretraining easier to debug, validate, and resume reproducibly.