Humanoid World Action Model With Joint State--Action Generation
We propose HWAM, a Humanoid World Action Model with joint state--action generation, which makes the robot's post-execution proprioceptive state an explicit prediction target.
Key points
- Humanoid robots are a promising platform for general-purpose manipulation.
- Recent Vision-Language-Action (VLA) policies learn actions directly from multimodal observations, while World Action Models (WAMs) further incorporate future visual prediction to improve action generation.
- However, in hierarchical humanoid systems, VLA and WAM policies output reference actions that are subsequently realized through whole-body control, robot dynamics, balance, and contact.
- By jointly generating reference actions and their realized body states, HWAM directly incorporates supervision of executed motion into action learning.
Sources (1)
- [1]Humanoid World Action Model With Joint State--Action GenerationarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 02:23 PM
We propose HWAM, a Humanoid World Action Model with joint state--action generation, which makes the robot's post-execution proprioceptive state an explicit prediction target.
Humanoid robots are a promising platform for general-purpose manipulation.
Extractive summary: sentences quoted from the sources.