$σ$Transfer: Uncertainty Transfer from Small to Large Networks under $μ\mathrm{P}$
Reliable predictive uncertainty in Laplace approximations depends critically on the prior precision, yet selecting it requires a posterior sweep that is prohibitively expensive for neural networks with billions of parameters.
Key points
- Under the Maximal Update Parametrization ($μP$), we derive a rescaling of the prior covariance that makes the selected precision stable as model width grows.
- This leads to $σTransfer$: we select the precision on a smaller model and zero-shot transfer it to the much larger model, i.e., without searching for the precision on the larger model at all.
- We show convergence of the prior kernel, posterior covariance, selected precision, and posterior-derived decisions under explicit conditions, and verify $σTransfer$ across regression, image classification, and Transformer readouts.
- For example, measured precision-sweep speedups reach $\sim 5000\times$ when transferring from width 128 to 4096 on MNIST, at a target-NLL degradation of $0.002$; transferring from a public 1B to 7B model gives a median search speedup of $\sim 2.3\times$ (up to $\sim 330\times$), with a mean measured target-NLL increase below $10^{-4}$ across ten tasks.
Sources (1)
- [1]$σ$Transfer: Uncertainty Transfer from Small to Large Networks under $μ\mathrm{P}$arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 10:40 AM
Reliable predictive uncertainty in Laplace approximations depends critically on the prior precision, yet selecting it requires a posterior sweep that is prohibitively expensive for neural networks with billions of parameters.
Under the Maximal Update Parametrization ($μP$), we derive a rescaling of the prior covariance that makes the selected precision stable as model width grows.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 8, 2026LEGO: A Lifting-Free Approach for Exocentric-to-Egocentric Video Generation
- Oct 8, 2026One Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed Experts
- Oct 8, 2026Scaling to Tens of Thousands of Test-Time Iterations with Loop-Native Attention Residuals
- Oct 7, 2026Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
- Oct 6, 2026EmbeddingGemma 2: an open, lightweight multimodal embedding model
- Oct 6, 2026huggingface/transformers v5.19.0: Release v5.19.0