AION
Research paperTraining & Scaling · Efficiency & Inference · Large Language Models1 source · Oct 8, 2026

$σ$Transfer: Uncertainty Transfer from Small to Large Networks under $μ\mathrm{P}$

Reliable predictive uncertainty in Laplace approximations depends critically on the prior precision, yet selecting it requires a posterior sweep that is prohibitively expensive for neural networks with billions of parameters.

Key points

  • Under the Maximal Update Parametrization ($μP$), we derive a rescaling of the prior covariance that makes the selected precision stable as model width grows.
  • This leads to $σTransfer$: we select the precision on a smaller model and zero-shot transfer it to the much larger model, i.e., without searching for the precision on the larger model at all.
  • We show convergence of the prior kernel, posterior covariance, selected precision, and posterior-derived decisions under explicit conditions, and verify $σTransfer$ across regression, image classification, and Transformer readouts.
  • For example, measured precision-sweep speedups reach $\sim 5000\times$ when transferring from width 128 to 4096 on MNIST, at a target-NLL degradation of $0.002$; transferring from a public 1B to 7B model gives a median search speedup of $\sim 2.3\times$ (up to $\sim 330\times$), with a mean measured target-NLL increase below $10^{-4}$ across ten tasks.

Sources (1)

  • [1]$σ$Transfer: Uncertainty Transfer from Small to Large Networks under $μ\mathrm{P}$
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 10:40 AM
    Reliable predictive uncertainty in Laplace approximations depends critically on the prior precision, yet selecting it requires a posterior sweep that is prohibitively expensive for neural networks with billions of parameters.
    Under the Maximal Update Parametrization ($μP$), we derive a rescaling of the prior covariance that makes the selected precision stable as model width grows.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 8, 2026LEGO: A Lifting-Free Approach for Exocentric-to-Egocentric Video Generation
  2. Oct 8, 2026One Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed Experts
  3. Oct 8, 2026Scaling to Tens of Thousands of Test-Time Iterations with Loop-Native Attention Residuals
  4. Oct 7, 2026Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
  5. Oct 6, 2026EmbeddingGemma 2: an open, lightweight multimodal embedding model
  6. Oct 6, 2026huggingface/transformers v5.19.0: Release v5.19.0

Related