Enhancing Diffusion Language Models with Autoregressive Post-Training Weights
Diffusion language models (dLLMs) have emerged as a promising alternative to autoregressive (AR) language models, offering flexible token-update orders and parallel decoding.
Key points
- Recent dLLMs are often initialized from pretrained AR models before diffusion conversion in order to inherit their learned representations.
- In this work, we show that these existing AR post-training weight updates can instead be effectively recycled to enhance diffusion models.
- Based on these findings, we propose A2D, a simple training-free framework for enhancing diffusion models with existing AR post-training resources.
- A2D can transfer capabilities from AR post-trained models to diffusion base models, and further improve already post-trained diffusion models by composing AR and diffusion post-training updates.
Sources (1)
- [1]Enhancing Diffusion Language Models with Autoregressive Post-Training WeightsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 10:31 AM
Diffusion language models (dLLMs) have emerged as a promising alternative to autoregressive (AR) language models, offering flexible token-update orders and parallel decoding.
Recent dLLMs are often initialized from pretrained AR models before diffusion conversion in order to inherit their learned representations.
Extractive summary: sentences quoted from the sources.