Two Halves are More than One: Phase-wise Velocity Distillation for Fast and High-Quality Image Generation
Recent diffusion-based image generation backbones have grown substantially in scale, making the network inference cost increase rapidly.
Key points
- While diffusion distillation techniques can reduce the number of inference steps, high-quality image generation within a single full-backbone-forward compute budget remains challenging.
- To address this issue, we propose Phase-wise Velocity Distillation (PVD), which partitions the generation timeline into a coarse and a fine phase, and models the transition within each phase via the average velocity.
- We show that the use of two half-sized phase-specific experts outperforms a single full-size monolithic student.
- On more complex text-to-image (T2I) tasks, PVD-distilled models (Stable Diffusion 3.5-Medium, FLUX.1-dev, Qwen-Image) produce results competitive with their multi-step teachers, significantly outperforming prior distillation methods.
Sources (1)
- [1]Two Halves are More than One: Phase-wise Velocity Distillation for Fast and High-Quality Image GenerationarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 10:03 AM
Recent diffusion-based image generation backbones have grown substantially in scale, making the network inference cost increase rapidly.
While diffusion distillation techniques can reduce the number of inference steps, high-quality image generation within a single full-backbone-forward compute budget remains challenging.
Extractive summary: sentences quoted from the sources.