ResearchResearch paperRobotics & Embodied AI · Image, Video & 3D Generation1 source · Oct 7, 2026

STRIKE: Learning Visual State Transitions for Physical World Modeling

We propose STRIKE, a framework that separates visual state transition learning from dense video generation.

Key points

  • Physical world modeling requires predicting how interactions change a scene, not merely generating coherent motion.
  • We construct event-aligned supervision by extracting observed states from training videos and pairing them with transition descriptions and temporal offsets.
  • An image-based transition model learns to predict the next scene configuration from the current image, a local transition specification, and elapsed time.
  • These results support learned visual state transitions as an effective intermediate representation for physical world modeling.

Sources (1)

  • [1]STRIKE: Learning Visual State Transitions for Physical World Modeling
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 06:11 AM
    We propose STRIKE, a framework that separates visual state transition learning from dense video generation.
    Physical world modeling requires predicting how interactions change a scene, not merely generating coherent motion.

Extractive summary: sentences quoted from the sources.

Related