Cross-Embodiment Robot Foundation World Models with Latent Actions
We introduce the Latent Action-Conditioned Robot World Model (LAC-WM), which operates within a learned unified latent action space shared across diverse embodiments.
Key points
- The diversity of robot embodiments and action spaces makes it challenging to build robot world models that generalize across different embodiments.
- We compare LAC-WM with an Explicit Action-Conditioned World Model (EAC-WM), which conditions on explicit motion labels.
- Our results show that explicit action conditioning leads to disjoint action representations across embodiments, limiting downstream performance when adapting to new robots.
- These results highlight the importance of a unified action space for efficient cross-embodiment learning, addressing a key challenge in robotics.
Sources (1)
- [1]Cross-Embodiment Robot Foundation World Models with Latent ActionsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 07:51 PM
We introduce the Latent Action-Conditioned Robot World Model (LAC-WM), which operates within a learned unified latent action space shared across diverse embodiments.
The diversity of robot embodiments and action spaces makes it challenging to build robot world models that generalize across different embodiments.
Extractive summary: sentences quoted from the sources.