Being-M0.7: A Latent World-Action Model for Humanoid Robots
We present Being-M0.7, a latent world-action model that transfers visual-motion priors learned from mixed-modality human data to humanoid control through pre-training, robot mid-training, and action post-training.
ProofPaper ↗
Key points
- Humanoid loco-manipulation requires coordinated locomotion and manipulation informed by future scene evolution and whole-body motion, yet learning these capabilities is constrained by scarce robot demonstrations.
- We curate a corpus from more than 10,000 hours of raw human-centric data, integrating video-only, motion-only, and paired video-motion streams to learn complementary visual dynamics and whole-body kinematic structure.
- Joint prediction of future latent visual states and motion encourages visual representations to encode future kinematics.
- During action post-training, an action expert combines visual predictive representations from the frozen, adapted prior with current images and proprioception through gated cross-attention, grounding predictive context in executable whole-body commands.
Sources (1)
- [1]Being-M0.7: A Latent World-Action Model for Humanoid RobotsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 05:49 AM
We present Being-M0.7, a latent world-action model that transfers visual-motion priors learned from mixed-modality human data to humanoid control through pre-training, robot mid-training, and action post-training.
Humanoid loco-manipulation requires coordinated locomotion and manipulation informed by future scene evolution and whole-body motion, yet learning these capabilities is constrained by scarce robot demonstrations.
Extractive summary: sentences quoted from the sources.