WorldAlign: Decoupled 4D Reward for World-Consistent Video Generation
Faithful visual world simulation requires generated videos to maintain 4D world consistency, encompassing both static and dynamic consistency.
Key points
- Geometry-aware post-training offers a promising way to improve world consistency.
- To address these limitations, we introduce WorldAlign, a decoupled 4D reward framework that semantically separates static regions and dynamic subjects and provides feedback by aligning each with a world prior suited to its assumptions.
- For static regions, WorldAlign aligns static geometry with a geometric world prior through semantically guided masked reprojection, enabling more reliable static-consistency evaluation; an auxiliary camera-motion reward discourages nearly static solutions.
- These results support decoupled world-prior alignment for more faithful visual world simulation.
Sources (1)
- [1]WorldAlign: Decoupled 4D Reward for World-Consistent Video GenerationarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 05:35 PM
Faithful visual world simulation requires generated videos to maintain 4D world consistency, encompassing both static and dynamic consistency.
Geometry-aware post-training offers a promising way to improve world consistency.
Extractive summary: sentences quoted from the sources.