AION
Research paperTraining & Scaling · Large Language Models · Interpretability1 source · Oct 7, 2026

Shared Gaussianization: What Gaussian Regularizers Certify About Contrastive Learning, and What They Miss

What can a distribution-matching regularizer such as SIGReg in LeJEPA certify about contrastive learning?

Key points

  • We study shared Gaussianization (SG), a characteristic-function Gaussianity test on the average of two normalized views, scaled by an independent $χd$ radius.
  • With an explicit alignment term, a rotation-invariant uniformity test gives a linear bound if and only if its spectrum dominates that of InfoNCE's kernel $e^{βu^\top v}$; SG's own test does, Gaussian kernels $e^{-γ\|u-v\|^2}$ qualify exactly when $γ\ge β/2$, and moment matching never does.
  • At finite batch size, an off-diagonal U-statistic removes a plug-in bias toward misalignment.
  • In controlled latent-variable models, pure SG retains per-view style, an alignment weight above the measured gain removes it, and for LeJEPA at three batch sizes the measured gain separates the encoders that retain style from those that do not.

Sources (1)

Extractive summary: sentences quoted from the sources.