Dino Forcing Flow Models: Do not denoise what you can predict
Co-denoising pretrained representations such as DINO can substantially improve the training speed and quality of flow matching models, but it introduces a second denoising trajectory and requires carefully designed schedules.
Key points
- We propose a simpler alternative: predict the pretrained representation directly, then condition the model on its own prediction.
- This removes the need for a second ODE and any representation-specific denoising schedules, while retaining the benefits of representation guidance.
- On ImageNet, it outperforms the state of the art in latent space at 2x fewer epochs than prior methods; in pixel space, it improves FID over comparable prior methods by more than 20%.
- These results support a simple principle: do not denoise what you can predict.
Sources (1)
- [1]Dino Forcing Flow Models: Do not denoise what you can predictarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 11:41 AM
Co-denoising pretrained representations such as DINO can substantially improve the training speed and quality of flow matching models, but it introduces a second denoising trajectory and requires carefully designed schedules.
We propose a simpler alternative: predict the pretrained representation directly, then condition the model on its own prediction.
Extractive summary: sentences quoted from the sources.