Conditional Residual Prediction: Improving Autoregressive Video Diffusion without a Bidirectional Teacher
Causal video diffusion models generate video autoregressively, which suits streaming, interactive, and long-video generation.
Key points
- We instead train a causal model from an image-model initialization, with no bidirectional video model at any stage.
- On this path, we find that a causal model trained on ground-truth history becomes strongly dependent on it, so that at inference errors in its own generated history propagate forward.
- We propose Conditional Residual Prediction (CRP), a simple recipe for reducing a model's reliance on a condition: the model first predicts the target without the condition, and the condition may only add a residual on top of this prediction.
- Scaling this recipe, we train Optica, a 2B-parameter causal video model that autoregressively generates 5-second 480p videos and reaches 82.78 on VBench with only about 15M training videos.
Sources (1)
- [1]Conditional Residual Prediction: Improving Autoregressive Video Diffusion without a Bidirectional TeacherarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 08:25 AM
Causal video diffusion models generate video autoregressively, which suits streaming, interactive, and long-video generation.
We instead train a causal model from an image-model initialization, with no bidirectional video model at any stage.
Extractive summary: sentences quoted from the sources.