ResearchResearch paperTraining & Scaling1 source · Oct 8, 2026

Rare Gate Disagreements Can Limit Plasticity: When Gradient Flow Mispredicts Finite-Batch SGD

We show that it can mispredict finite-batch stochastic gradient descent (SGD) qualitatively, and we trace the discrepancy to a specific mechanism.

Key points

  • Population gradient flow is a common tool for reasoning about how neural networks adapt, including after pretraining.
  • In a two-unit ReLU regression, a source task drives the two neurons toward positive proportionality and a target task rewards separating them.
  • After source training for time $T$, gradient flow recovers on the target in time linear in $T$.
  • Bounding the cumulative probability of sampling the wedge along the exact online recursion, without a diffusion approximation, shows that recovery with fixed probability from an identical source-gradient-flow checkpoint, within $e^{c/η}$ updates, requires $Nb \gtrsim e^{λT}$ target samples and batch size $b \gtrsim ηe^{λT}$, where $N$ counts updates and $λ$ is the weight decay.

Sources (1)

Extractive summary: sentences quoted from the sources.

Related