ResearchResearch paperRobotics & Embodied AI1 source · Oct 7, 2026

OmniHOI: Dexterous Hand-Object Interaction from Monocular Human Video

We present OmniHOI, a pipeline that turns an RGB video of hand-object interaction into an interaction-faithful trajectory on dexterous hands.

Key points

  • Monocular videos of human manipulation provide abundant dexterous demonstrations, yet reconstructing hand-object interaction from a single view and transferring it to robot hands remain difficult, limiting their direct use for robot execution.
  • The key idea is to enforce physical consistency using the evidence available at each stage: image evidence during reconstruction, contact geometry during retargeting, and dynamics during physics-in-the-loop refinement.
  • Across 150 motion-capture trajectories transferred to each of five dexterous hands with 6 to 22 DoF, we achieve 39-89% success, compared with at most 31% for prior transfer methods.
  • Its trajectories also execute on a real bimanual robot across diverse tasks.

Sources (1)

  • [1]OmniHOI: Dexterous Hand-Object Interaction from Monocular Human Video
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 07:59 PM
    We present OmniHOI, a pipeline that turns an RGB video of hand-object interaction into an interaction-faithful trajectory on dexterous hands.
    Monocular videos of human manipulation provide abundant dexterous demonstrations, yet reconstructing hand-object interaction from a single view and transferring it to robot hands remain difficult, limiting their direct use for robot execution.

Extractive summary: sentences quoted from the sources.

Related