AION
Research paperTraining & Scaling · Efficiency & Inference · Large Language Models1 source · Oct 8, 2026

Spectral Weight Decay: Inducing Low-Rank Structure in Neural Network Weights

We introduce spectral weight decay, a post-step decoupled nuclear-norm update that applies additive rather than multiplicative spectral shrinkage.

Key points

  • Standard weight decay treats each weight matrix as a vector and ignores its spectral structure.
  • We connect the update to approximate proximal descent and show that its sensitivity to update order can exceed that of conventional $\ell2$ weight decay near rank deficiency.
  • Across LLaMA models with $124$M to $500$M parameters, spectral weight decay lowers effective rank and improves SVD-LLM compression at matched validation loss.
  • Under fixed-horizon training with $60%$ label noise, it also improves final mean clean-test accuracy over matched $\ell2$ regularization by up to $17.8$ points on MNIST and $4.6$ points across four BERT-base tasks.

Sources (1)

  • [1]Spectral Weight Decay: Inducing Low-Rank Structure in Neural Network Weights
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 11:26 AM
    We introduce spectral weight decay, a post-step decoupled nuclear-norm update that applies additive rather than multiplicative spectral shrinkage.
    Standard weight decay treats each weight matrix as a vector and ignores its spectral structure.

Extractive summary: sentences quoted from the sources.