Early Memory Selection for Balanced Adam
We propose a method for choosing the shared memory parameter $β1=β2=β$ in Adam from a short pilot training.
ProofPaper ↗
Key points
- The selected $β$ remains fixed during the subsequent full training.
- A local model of Adam's normalized direction balances sampling variability against the delay introduced by averaging past gradients.
- This balance gives a cubic memory rule, whose two coefficients are estimated from gradient probes at a few pilot checkpoints.
- With a 200-update pilot and sixteen probe gradients at each of four checkpoints, a seed-matched retrospective evaluation on eleven vision and language workloads reduces mean relative validation gap by 40.7% and worst-quarter mean gap by 44.3% against the grid representative of shared $β=0.95$.
Sources (1)
- [1]Early Memory Selection for Balanced AdamarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 04:22 PM
We propose a method for choosing the shared memory parameter $β_1=β_2=β$ in Adam from a short pilot training.
The selected $β$ remains fixed during the subsequent full training.
Extractive summary: sentences quoted from the sources.