Shared Gaussianization: What Gaussian Regularizers Certify About Contrastive Learning, and What They Miss
What can a distribution-matching regularizer such as SIGReg in LeJEPA certify about contrastive learning?
Key points
- We study shared Gaussianization (SG), a characteristic-function Gaussianity test on the average of two normalized views, scaled by an independent $χd$ radius.
- With an explicit alignment term, a rotation-invariant uniformity test gives a linear bound if and only if its spectrum dominates that of InfoNCE's kernel $e^{βu^\top v}$; SG's own test does, Gaussian kernels $e^{-γ\|u-v\|^2}$ qualify exactly when $γ\ge β/2$, and moment matching never does.
- At finite batch size, an off-diagonal U-statistic removes a plug-in bias toward misalignment.
- In controlled latent-variable models, pure SG retains per-view style, an alignment weight above the measured gain removes it, and for LeJEPA at three batch sizes the measured gain separates the encoders that retain style from those that do not.
Sources (1)
- [1]Shared Gaussianization: What Gaussian Regularizers Certify About Contrastive Learning, and What They MissarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 03:53 PM
What can a distribution-matching regularizer such as SIGReg in LeJEPA certify about contrastive learning?
We study shared Gaussianization (SG), a characteristic-function Gaussianity test on the average of two normalized views, scaled by an independent $χ_d$ radius.
Extractive summary: sentences quoted from the sources.