WOVEN: Weaving Visual World Modeling into Multimodal LLMs
We therefore introduce WOVEN, a training source and benchmark for visual transition reasoning that organizes transition supervision by scene, action, and reasoning type, using diverse, realistic rollouts from video-pretrained generative models: 36,076 examples across 20 scene types, 5 action types, and 8 reasoning types.
Key points
- Multimodal large language models (MLLMs) struggle with spatial, embodied, physical, and temporal reasoning.
- We hypothesize that these failures reflect a shared deficit in visual transition reasoning, and test whether this capability can serve as a shared training primitive, one that different models can learn from different supervision sources and reuse across different tasks, with a systematic training recipe.
- Controlled comparisons further yield a training recipe for visual world modeling, validated prospectively on held-out benchmarks: select supervision by the reasoning operation it teaches rather than by the actions, scenes, or domains it shows, and prefer larger changes to the visual state for robustness.
- Our work establishes visual transition reasoning as a reusable foundation for systematic visual world-model training in MLLMs.
Sources (1)
- [1]WOVEN: Weaving Visual World Modeling into Multimodal LLMsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 05:52 PM
We therefore introduce WOVEN, a training source and benchmark for visual transition reasoning that organizes transition supervision by scene, action, and reasoning type, using diverse, realistic rollouts from video-pretrained generative models: 36,076 examples across 20 scene types, 5 action types, and 8 reasoning types.
Multimodal large language models (MLLMs) struggle with spatial, embodied, physical, and temporal reasoning.
Extractive summary: sentences quoted from the sources.