OmniHOI: Dexterous Hand-Object Interaction from Monocular Human Video
We present OmniHOI, a pipeline that turns an RGB video of hand-object interaction into an interaction-faithful trajectory on dexterous hands.
ProofPaper ↗
Key points
- Monocular videos of human manipulation provide abundant dexterous demonstrations, yet reconstructing hand-object interaction from a single view and transferring it to robot hands remain difficult, limiting their direct use for robot execution.
- The key idea is to enforce physical consistency using the evidence available at each stage: image evidence during reconstruction, contact geometry during retargeting, and dynamics during physics-in-the-loop refinement.
- Across 150 motion-capture trajectories transferred to each of five dexterous hands with 6 to 22 DoF, we achieve 39-89% success, compared with at most 31% for prior transfer methods.
- Its trajectories also execute on a real bimanual robot across diverse tasks.
Sources (1)
- [1]OmniHOI: Dexterous Hand-Object Interaction from Monocular Human VideoarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 07:59 PM
We present OmniHOI, a pipeline that turns an RGB video of hand-object interaction into an interaction-faithful trajectory on dexterous hands.
Monocular videos of human manipulation provide abundant dexterous demonstrations, yet reconstructing hand-object interaction from a single view and transferring it to robot hands remain difficult, limiting their direct use for robot execution.
Extractive summary: sentences quoted from the sources.