Diffusion Removes Langevin's Conditioning Dependence: A Sharp Gaussian Analysis
Despite their empirical success, why diffusion models overcome the bottlenecks of classical score-based samplers remains unclear.
Key points
- In this work, we leverage Gaussian distributions to isolate this phenomenon.
- We establish 2-Wasserstein convergence bounds for optimized hyperparameters, showing that diffusion processes achieve a sampling error of $O(\sqrt{dλ{\max}}\log N/N)$, where $d$ is the dimension, $N$ the number of sampling steps, and $λ{\max}$ the largest eigenvalue of the target covariance matrix.
- Unadjusted and underdamped Langevin dynamics suffer from an additional $\sqrtκ$ factor, where $κ$ is the condition number.
- By contrast, in the learning phase, we show that estimating the unnoised score by gradient descent leads to essentially the same estimator as estimating a noisy score, which suggests that the benefits of noising do not come from the learning phase.
Sources (1)
- [1]Diffusion Removes Langevin's Conditioning Dependence: A Sharp Gaussian AnalysisarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 02:36 PM
Despite their empirical success, why diffusion models overcome the bottlenecks of classical score-based samplers remains unclear.
In this work, we leverage Gaussian distributions to isolate this phenomenon.
Extractive summary: sentences quoted from the sources.