ResearchResearch paperInterpretability · Efficiency & Inference · Training & Scaling1 source · Oct 7, 2026

Eigenvalues of the Hessian in Deep Learning: The Origin of Symmetry and Its Breaking

Hessian spectra at trained models in deep learning exhibit a persistent pattern: eigenvalues organize into distinct clusters, including a large bulk near zero and a few isolated outliers.

Key points

  • This paper shows that a natural account of these spectral phenomena emerges when the original setting is understood as a departure from a nearby, otherwise hidden, highly symmetric reference.
  • Modifications, including changes to the architecture, data distribution, or parameter metric, expose a nearby reference configuration whose Hessian exhibits rich invariances-ones not accounted for by weight symmetries.
  • The framework is developed in some generality, with a detailed analysis of three-layer ReLU networks and applications to convolutional, graph, and transformer models, as well as to the NTK.
  • The same mechanism is further shown to yield analogous spectral structures in layerwise Hessians and the Gauss-Newton matrix.

Sources (1)

  • [1]Eigenvalues of the Hessian in Deep Learning: The Origin of Symmetry and Its Breaking
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 12:08 PM
    Hessian spectra at trained models in deep learning exhibit a persistent pattern: eigenvalues organize into distinct clusters, including a large bulk near zero and a few isolated outliers.
    This paper shows that a natural account of these spectral phenomena emerges when the original setting is understood as a departure from a nearby, otherwise hidden, highly symmetric reference.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 7, 2026Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
  2. Oct 6, 2026EmbeddingGemma 2: an open, lightweight multimodal embedding model
  3. Oct 5, 2026LiquidAI/d1-omni-600M
  4. Oct 5, 2026MC-Sparse: Deconstructing and Closing the Dense-Sparse Attention Gap in Diffusion Transformers
  5. Sep 30, 2026Cloudflare/clef-flash
  6. Sep 29, 2026microsoft/AesCode-32B

Related