AION
Research paperImage, Video & 3D Generation1 source · Oct 8, 2026

Conditional Residual Prediction: Improving Autoregressive Video Diffusion without a Bidirectional Teacher

Causal video diffusion models generate video autoregressively, which suits streaming, interactive, and long-video generation.

Key points

  • We instead train a causal model from an image-model initialization, with no bidirectional video model at any stage.
  • On this path, we find that a causal model trained on ground-truth history becomes strongly dependent on it, so that at inference errors in its own generated history propagate forward.
  • We propose Conditional Residual Prediction (CRP), a simple recipe for reducing a model's reliance on a condition: the model first predicts the target without the condition, and the condition may only add a residual on top of this prediction.
  • Scaling this recipe, we train Optica, a 2B-parameter causal video model that autoregressively generates 5-second 480p videos and reaches 82.78 on VBench with only about 15M training videos.

Sources (1)

Extractive summary: sentences quoted from the sources.