AION
Research paperRobotics & Embodied AI1 source · Oct 6, 2026

OpenWAM: An Open Framework for Composable World-Action Models

We introduce OPENWAM, an open world-action modeling framework built around a common causal robot-video foundation and configurable video-action interaction.

Key points

  • World-action models (WAMs) couple future prediction with robot control, yet existing systems often vary the video backbone, interaction structure, supervision, and inference procedure simultaneously, making their design choices difficult to compare.
  • Starting from Wan2.2-5B, we perform causal robot-video pretraining on over 10,000 hours of video, then integrate an action expert through a shared Mixture-of-Transformers architecture that supports joint, video-then-action, action-then-video, and decoupled generation.
  • When only the video predictor is adapted to a new task, a frozen local-context inverse dynamics model trained on counterfactual data and demonstrations achieves 84.0% mean success across four held-out LIBERO-90 tasks, compared with 47.0% for a full-context inverse model and 21.5% for a local-context model trained only on demonstrations.
  • OPENWAM provides a common testbed for comparing WAM interaction designs and for studying dynamics learning from video data beyond successful demonstrations.

Sources (1)

  • [1]OpenWAM: An Open Framework for Composable World-Action Models
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 08:02 AM
    We introduce OPENWAM, an open world-action modeling framework built around a common causal robot-video foundation and configurable video-action interaction.
    World-action models (WAMs) couple future prediction with robot control, yet existing systems often vary the video backbone, interaction structure, supervision, and inference procedure simultaneously, making their design choices difficult to compare.

Extractive summary: sentences quoted from the sources.