Rare Gate Disagreements Can Limit Plasticity: When Gradient Flow Mispredicts Finite-Batch SGD
We show that it can mispredict finite-batch stochastic gradient descent (SGD) qualitatively, and we trace the discrepancy to a specific mechanism.
ProofPaper ↗
Key points
- Population gradient flow is a common tool for reasoning about how neural networks adapt, including after pretraining.
- In a two-unit ReLU regression, a source task drives the two neurons toward positive proportionality and a target task rewards separating them.
- After source training for time $T$, gradient flow recovers on the target in time linear in $T$.
- Bounding the cumulative probability of sampling the wedge along the exact online recursion, without a diffusion approximation, shows that recovery with fixed probability from an identical source-gradient-flow checkpoint, within $e^{c/η}$ updates, requires $Nb \gtrsim e^{λT}$ target samples and batch size $b \gtrsim ηe^{λT}$, where $N$ counts updates and $λ$ is the weight decay.
Sources (1)
- [1]Rare Gate Disagreements Can Limit Plasticity: When Gradient Flow Mispredicts Finite-Batch SGDarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 08:22 AM
We show that it can mispredict finite-batch stochastic gradient descent (SGD) qualitatively, and we trace the discrepancy to a specific mechanism.
Population gradient flow is a common tool for reasoning about how neural networks adapt, including after pretraining.
Extractive summary: sentences quoted from the sources.