ResearchResearch paperLarge Language Models · Training & Scaling · Efficiency & Inference1 source · Oct 6, 2026

Early Memory Selection for Balanced Adam

We propose a method for choosing the shared memory parameter $β1=β2=β$ in Adam from a short pilot training.

Key points

  • The selected $β$ remains fixed during the subsequent full training.
  • A local model of Adam's normalized direction balances sampling variability against the delay introduced by averaging past gradients.
  • This balance gives a cubic memory rule, whose two coefficients are estimated from gradient probes at a few pilot checkpoints.
  • With a 200-update pilot and sixteen probe gradients at each of four checkpoints, a seed-matched retrospective evaluation on eleven vision and language workloads reduces mean relative validation gap by 40.7% and worst-quarter mean gap by 44.3% against the grid representative of shared $β=0.95$.

Sources (1)

  • [1]Early Memory Selection for Balanced Adam
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 04:22 PM
    We propose a method for choosing the shared memory parameter $β_1=β_2=β$ in Adam from a short pilot training.
    The selected $β$ remains fixed during the subsequent full training.

Extractive summary: sentences quoted from the sources.

Related