Global Exponential Convergence of Two-Layer Linear Network Training
We prove global exponential (linear) convergence with an explicit rate in the rich scaling for wide two-layer linear networks trained with smooth Polyak-Lojasiewicz predictor losses.
ProofPaper ↗
Key points
- Gradient flow in the factors closes exactly in terms of a finite-dimensional Bures flow of the neuron law covariance, in which the predictor dynamics are preconditioned by hidden covariance blocks.
- Mean-field conservation laws provide uniform spectral lower bounds on the hidden preconditioning blocks when the initial covariance satisfies a spectral support gap condition.
- For an initial covariance $Σ0 = σ^2 Id$, the loss converges to the global minimum with linear rate at least $4σ^2κ$, where $κ$ is the PL constant.
- We establish stability of this rate under finite-width sampling, as well as global convergence of factor gradient descent for an explicit stepsize interval depending on smoothness, the initial loss, and conserved spectral margins.
Sources (1)
- [1]Global Exponential Convergence of Two-Layer Linear Network TrainingarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 03:13 AM
We prove global exponential (linear) convergence with an explicit rate in the rich scaling for wide two-layer linear networks trained with smooth Polyak-Lojasiewicz predictor losses.
Gradient flow in the factors closes exactly in terms of a finite-dimensional Bures flow of the neuron law covariance, in which the predictor dynamics are preconditioned by hidden covariance blocks.
Extractive summary: sentences quoted from the sources.