Eigenvalues of the Hessian in Deep Learning: The Origin of Symmetry and Its Breaking
Hessian spectra at trained models in deep learning exhibit a persistent pattern: eigenvalues organize into distinct clusters, including a large bulk near zero and a few isolated outliers.
ProofPaper ↗
Key points
- This paper shows that a natural account of these spectral phenomena emerges when the original setting is understood as a departure from a nearby, otherwise hidden, highly symmetric reference.
- Modifications, including changes to the architecture, data distribution, or parameter metric, expose a nearby reference configuration whose Hessian exhibits rich invariances-ones not accounted for by weight symmetries.
- The framework is developed in some generality, with a detailed analysis of three-layer ReLU networks and applications to convolutional, graph, and transformer models, as well as to the NTK.
- The same mechanism is further shown to yield analogous spectral structures in layerwise Hessians and the Gauss-Newton matrix.
Sources (1)
- [1]Eigenvalues of the Hessian in Deep Learning: The Origin of Symmetry and Its BreakingarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 12:08 PM
Hessian spectra at trained models in deep learning exhibit a persistent pattern: eigenvalues organize into distinct clusters, including a large bulk near zero and a few isolated outliers.
This paper shows that a natural account of these spectral phenomena emerges when the original setting is understood as a departure from a nearby, otherwise hidden, highly symmetric reference.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 7, 2026Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
- Oct 6, 2026EmbeddingGemma 2: an open, lightweight multimodal embedding model
- Oct 5, 2026LiquidAI/d1-omni-600M
- Oct 5, 2026MC-Sparse: Deconstructing and Closing the Dense-Sparse Attention Gap in Diffusion Transformers
- Sep 30, 2026Cloudflare/clef-flash
- Sep 29, 2026microsoft/AesCode-32B