Spectral Weight Decay: Inducing Low-Rank Structure in Neural Network Weights
We introduce spectral weight decay, a post-step decoupled nuclear-norm update that applies additive rather than multiplicative spectral shrinkage.
Key points
- Standard weight decay treats each weight matrix as a vector and ignores its spectral structure.
- We connect the update to approximate proximal descent and show that its sensitivity to update order can exceed that of conventional $\ell2$ weight decay near rank deficiency.
- Across LLaMA models with $124$M to $500$M parameters, spectral weight decay lowers effective rank and improves SVD-LLM compression at matched validation loss.
- Under fixed-horizon training with $60%$ label noise, it also improves final mean clean-test accuracy over matched $\ell2$ regularization by up to $17.8$ points on MNIST and $4.6$ points across four BERT-base tasks.
Sources (1)
- [1]Spectral Weight Decay: Inducing Low-Rank Structure in Neural Network WeightsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 11:26 AM
We introduce spectral weight decay, a post-step decoupled nuclear-norm update that applies additive rather than multiplicative spectral shrinkage.
Standard weight decay treats each weight matrix as a vector and ignores its spectral structure.
Extractive summary: sentences quoted from the sources.