ResearchResearch paperRobotics & Embodied AI · Multimodal Models · Image, Video & 3D Generation1 source · Oct 7, 2026

Video Prediction Policy 2: Predict Better, Act Better

We introduce Video Prediction Policy 2 (VPP2), a WAM that enables strong zero-shot generalization in both video prediction and action generation.

Key points

  • World action models (WAMs) have emerged as an important class of generalist robot policies, aiming to transfer video prediction priors to action learning.
  • However, we find that existing WAMs frequently produce incorrect motion predictions in open-ended environment, leading to erroneous actions.
  • We attribute this limitation to two factors: (1) base video models are not optimized for manipulation, and (2) naively incorporating action components into video models can substantially degrade their generalization capabilities.
  • Finally, we introduce action module via a mixture-of-transformers (MoT) architecture to learn implicit inverse dynamics model.

Sources (1)

  • [1]Video Prediction Policy 2: Predict Better, Act Better
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 03:40 PM
    We introduce Video Prediction Policy 2 (VPP2), a WAM that enables strong zero-shot generalization in both video prediction and action generation.
    World action models (WAMs) have emerged as an important class of generalist robot policies, aiming to transfer video prediction priors to action learning.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 7, 2026Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
  2. Oct 6, 2026EmbeddingGemma 2: an open, lightweight multimodal embedding model
  3. Oct 5, 2026LiquidAI/d1-omni-600M
  4. Oct 5, 2026MC-Sparse: Deconstructing and Closing the Dense-Sparse Attention Gap in Diffusion Transformers
  5. Sep 30, 2026Cloudflare/clef-flash
  6. Sep 29, 2026microsoft/AesCode-32B

Related