AION
Research paperLarge Language Models1 source · Oct 6, 2026

Enhancing Diffusion Language Models with Autoregressive Post-Training Weights

Diffusion language models (dLLMs) have emerged as a promising alternative to autoregressive (AR) language models, offering flexible token-update orders and parallel decoding.

Key points

  • Recent dLLMs are often initialized from pretrained AR models before diffusion conversion in order to inherit their learned representations.
  • In this work, we show that these existing AR post-training weight updates can instead be effectively recycled to enhance diffusion models.
  • Based on these findings, we propose A2D, a simple training-free framework for enhancing diffusion models with existing AR post-training resources.
  • A2D can transfer capabilities from AR post-trained models to diffusion base models, and further improve already post-trained diffusion models by composing AR and diffusion post-training updates.

Sources (1)

  • [1]Enhancing Diffusion Language Models with Autoregressive Post-Training Weights
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 10:31 AM
    Diffusion language models (dLLMs) have emerged as a promising alternative to autoregressive (AR) language models, offering flexible token-update orders and parallel decoding.
    Recent dLLMs are often initialized from pretrained AR models before diffusion conversion in order to inherit their learned representations.

Extractive summary: sentences quoted from the sources.