AION
Research paperMultimodal Models1 source · Oct 8, 2026

WOVEN: Weaving Visual World Modeling into Multimodal LLMs

We therefore introduce WOVEN, a training source and benchmark for visual transition reasoning that organizes transition supervision by scene, action, and reasoning type, using diverse, realistic rollouts from video-pretrained generative models: 36,076 examples across 20 scene types, 5 action types, and 8 reasoning types.

Key points

  • Multimodal large language models (MLLMs) struggle with spatial, embodied, physical, and temporal reasoning.
  • We hypothesize that these failures reflect a shared deficit in visual transition reasoning, and test whether this capability can serve as a shared training primitive, one that different models can learn from different supervision sources and reuse across different tasks, with a systematic training recipe.
  • Controlled comparisons further yield a training recipe for visual world modeling, validated prospectively on held-out benchmarks: select supervision by the reasoning operation it teaches rather than by the actions, scenes, or domains it shows, and prefer larger changes to the visual state for robustness.
  • Our work establishes visual transition reasoning as a reusable foundation for systematic visual world-model training in MLLMs.

Sources (1)

  • [1]WOVEN: Weaving Visual World Modeling into Multimodal LLMs
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 05:52 PM
    We therefore introduce WOVEN, a training source and benchmark for visual transition reasoning that organizes transition supervision by scene, action, and reasoning type, using diverse, realistic rollouts from video-pretrained generative models: 36,076 examples across 20 scene types, 5 action types, and 8 reasoning types.
    Multimodal large language models (MLLMs) struggle with spatial, embodied, physical, and temporal reasoning.

Extractive summary: sentences quoted from the sources.