ResearchResearch paperRobotics & Embodied AI1 source · Oct 6, 2026

IronMan: Information-Constrained Video-Action Learning for Robot Manipulation

Therefore, we introduce IronMan (Information-constRained videO-actioN learning for robot MANipulation), a robust video-action learning framework built on the information bottleneck principle.

Key points

  • Video Action Models (VAMs) couple visual dynamics modeling with action generation for robot manipulation.
  • However, video representations are not naturally suited to action generation, as exposing the action policy to excessive visual detail can impair its generalization ability.
  • The core principle of this framework is to impose information constraints that suppress irrelevant visual information while preserving action-relevant dynamics cues.
  • IronMan employs a dynamics-aware bottleneck that distills noisy, entangled one-step video features into compact world representations.

Sources (1)

  • [1]IronMan: Information-Constrained Video-Action Learning for Robot Manipulation
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 08:30 AM
    Therefore, we introduce IronMan (Information-constRained videO-actioN learning for robot MANipulation), a robust video-action learning framework built on the information bottleneck principle.
    Video Action Models (VAMs) couple visual dynamics modeling with action generation for robot manipulation.

Extractive summary: sentences quoted from the sources.

Related