Sparsifying Stochasticity, Not Capacity: Partial Stochasticity via Deep Weight Factorization of Prior Scales
Bayesian neural networks need not be fully stochastic to be universal conditional density approximators, but it remains open which parameters should be stochastic.
ProofPaper ↗
Key points
- We learn this split by applying deep weight factorization to the prior scales, which are the standard deviations of the parameter priors, while fitting the functional prior to a Gaussian process with a maximum mean discrepancy objective.
- A parameter whose prior scale falls below a cutoff becomes deterministic and is optimized during inference, so the regularizer sparsifies stochasticity rather than capacity.
- We give a certificate for universal conditional density approximation that is checkable in linear time, together with a minimal repair when it fails.
- We further show that the common hybrid scheme of sampling some parameters and optimizing the others is stochastic approximation for a type-II maximum a posteriori objective, and that coupled step sizes can leave a tracking error that does not vanish as the step size shrinks.
Sources (1)
- [1]Sparsifying Stochasticity, Not Capacity: Partial Stochasticity via Deep Weight Factorization of Prior ScalesarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 11:45 AM
Bayesian neural networks need not be fully stochastic to be universal conditional density approximators, but it remains open which parameters should be stochastic.
We learn this split by applying deep weight factorization to the prior scales, which are the standard deviations of the parameter priors, while fitting the functional prior to a Gaussian process with a maximum mean discrepancy objective.
Extractive summary: sentences quoted from the sources.