ResearchResearch paperRobotics & Embodied AI1 source · Oct 8, 2026

Being-M0.7: A Latent World-Action Model for Humanoid Robots

We present Being-M0.7, a latent world-action model that transfers visual-motion priors learned from mixed-modality human data to humanoid control through pre-training, robot mid-training, and action post-training.

Key points

  • Humanoid loco-manipulation requires coordinated locomotion and manipulation informed by future scene evolution and whole-body motion, yet learning these capabilities is constrained by scarce robot demonstrations.
  • We curate a corpus from more than 10,000 hours of raw human-centric data, integrating video-only, motion-only, and paired video-motion streams to learn complementary visual dynamics and whole-body kinematic structure.
  • Joint prediction of future latent visual states and motion encourages visual representations to encode future kinematics.
  • During action post-training, an action expert combines visual predictive representations from the frozen, adapted prior with current images and proprioception through gated cross-attention, grounding predictive context in executable whole-body commands.

Sources (1)

  • [1]Being-M0.7: A Latent World-Action Model for Humanoid Robots
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 05:49 AM
    We present Being-M0.7, a latent world-action model that transfers visual-motion priors learned from mixed-modality human data to humanoid control through pre-training, robot mid-training, and action post-training.
    Humanoid loco-manipulation requires coordinated locomotion and manipulation informed by future scene evolution and whole-body motion, yet learning these capabilities is constrained by scarce robot demonstrations.

Extractive summary: sentences quoted from the sources.

Related