AION
Research paperInterpretability · Large Language Models1 source · Oct 8, 2026

Memorization and Malign Generalization in Conditional Diffusion Models with Random Features

Conditional diffusion models generate diverse, novel, and high-quality samples under prescribed conditions.

Key points

  • In this work, we analyze a random-feature conditional score model in the high-dimensional proportional limit, deriving asymptotic expressions for training and test losses.
  • By decomposing the test loss, we show that in the overparameterized regime, increasing model width improves prediction of the condition-dependent mean while reducing within-condition prediction variance, a phenomenon we term "malign generalization.
  • Furthermore, analyzing the training loss reveals that more informative conditions lead to memorization of training samples at smaller widths.
  • These theoretical findings are supported by experiments with U-Net architectures on realistic data.

Sources (1)

Extractive summary: sentences quoted from the sources.