Dex-One2Many: Learning Dexterous Manipulation from a Single Human Demonstration
To address these coupled generalization and exploration challenges, we present Dex-One2Many, a real-to-sim-to-real framework that learns a generalizable dexterous manipulation policy from a single human video.
Key points
- While learning dexterous manipulation from a single human video offers a promising alternative to costly robot demonstrations, many recent methods predominantly imitate demonstrated motions.
- Such strict motion matching often limits generalization to initial object poses, goal poses, and grasps not shown in the video.
- Our key insight is to abstract the video into sequential scene graphs that guide RL, enabling efficient exploration while preserving broad generalizability.
- Trained entirely in simulation, Dex-One2Many transfers zero-shot to a real multi-fingered hand.
Sources (1)
- [1]Dex-One2Many: Learning Dexterous Manipulation from a Single Human DemonstrationarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 05:59 PM
To address these coupled generalization and exploration challenges, we present Dex-One2Many, a real-to-sim-to-real framework that learns a generalizable dexterous manipulation policy from a single human video.
While learning dexterous manipulation from a single human video offers a promising alternative to costly robot demonstrations, many recent methods predominantly imitate demonstrated motions.
Extractive summary: sentences quoted from the sources.