IronMan: Information-Constrained Video-Action Learning for Robot Manipulation
Therefore, we introduce IronMan (Information-constRained videO-actioN learning for robot MANipulation), a robust video-action learning framework built on the information bottleneck principle.
ProofPaper ↗
Key points
- Video Action Models (VAMs) couple visual dynamics modeling with action generation for robot manipulation.
- However, video representations are not naturally suited to action generation, as exposing the action policy to excessive visual detail can impair its generalization ability.
- The core principle of this framework is to impose information constraints that suppress irrelevant visual information while preserving action-relevant dynamics cues.
- IronMan employs a dynamics-aware bottleneck that distills noisy, entangled one-step video features into compact world representations.
Sources (1)
- [1]IronMan: Information-Constrained Video-Action Learning for Robot ManipulationarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 08:30 AM
Therefore, we introduce IronMan (Information-constRained videO-actioN learning for robot MANipulation), a robust video-action learning framework built on the information bottleneck principle.
Video Action Models (VAMs) couple visual dynamics modeling with action generation for robot manipulation.
Extractive summary: sentences quoted from the sources.