Omni-Diffusion-Distill: Few-Step Distillation of Unified Multimodal Diffusion Large Language Models
Unified multimodal diffusion large language models (dLLMs) offer a single architecture for both image generation and multimodal understanding, but their iterative decoding requires tens to hundreds of forward passes.
Key points
- We introduce Omni-Diffusion-Distill, a unified two-stage distillation framework that retains strong generation and understanding capabilities while substantially reducing the inference cost of a unified multimodal dLLM.
- Omni-Diffusion-Distill aligns the distillation of both generation and understanding, for both images and text, in the discrete token space.
- We further remedy two sources of degradation in unified distillation with a pairwise collision penalty that reduces repetition under parallel text decoding, and entropy-matched guidance that prevents entropy collapse caused by fitting the sharpened teacher distribution in image generation.
- Omni-Diffusion-Distill achieves state-of-the-art trade-offs between decoding efficiency and generation and understanding performance for multimodal dLLMs, reducing image generation from 128 to 8 decoding steps and multimodal understanding from 512 to 64, giving 18.2x and 21.2x wall-clock speedups.
Sources (1)
- [1]Omni-Diffusion-Distill: Few-Step Distillation of Unified Multimodal Diffusion Large Language ModelsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 11:24 PM
Unified multimodal diffusion large language models (dLLMs) offer a single architecture for both image generation and multimodal understanding, but their iterative decoding requires tens to hundreds of forward passes.
We introduce Omni-Diffusion-Distill, a unified two-stage distillation framework that retains strong generation and understanding capabilities while substantially reducing the inference cost of a unified multimodal dLLM.
Extractive summary: sentences quoted from the sources.