AION
Repository / library

Megatron-LM

Also known as: Megatron

2stories this week
2last 30 days
2all time

Timeline

  1. Oct 7, 2026 · Research paper · 1 source
    Expert Coupling in MoE Pretraining: Reducing All-to-All Overhead with Correlated Placement and Token Shuffling
    Mixture-of-Experts (MoE) layers replace the feed-forward block of a Transformer with E expert networks, and each token is routed to k of these experts.
  2. Oct 6, 2026 · Opinion / analysis · 1 source
    Scale Bitwise-Deterministic Pretraining with NVIDIA Megatron Core
    Bitwise determinism makes large-scale pretraining easier to debug, validate, and resume reproducibly.

Often appears with