AION
Research paperImage, Video & 3D Generation1 source · Oct 8, 2026

WorldAlign: Decoupled 4D Reward for World-Consistent Video Generation

Faithful visual world simulation requires generated videos to maintain 4D world consistency, encompassing both static and dynamic consistency.

Key points

  • Geometry-aware post-training offers a promising way to improve world consistency.
  • To address these limitations, we introduce WorldAlign, a decoupled 4D reward framework that semantically separates static regions and dynamic subjects and provides feedback by aligning each with a world prior suited to its assumptions.
  • For static regions, WorldAlign aligns static geometry with a geometric world prior through semantically guided masked reprojection, enabling more reliable static-consistency evaluation; an auxiliary camera-motion reward discourages nearly static solutions.
  • These results support decoupled world-prior alignment for more faithful visual world simulation.

Sources (1)

  • [1]WorldAlign: Decoupled 4D Reward for World-Consistent Video Generation
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 05:35 PM
    Faithful visual world simulation requires generated videos to maintain 4D world consistency, encompassing both static and dynamic consistency.
    Geometry-aware post-training offers a promising way to improve world consistency.

Extractive summary: sentences quoted from the sources.