ResearchResearch paperRobotics & Embodied AI1 source · Oct 7, 2026

RLHND: Video Foundation Models as Physically Grounded Hand Trackers for Robot Learning

To this end, we propose RLHND, a video foundation model-based hand tracking model that jointly estimates hand pose and realistic tactile information from monocular egocentric videos.

Key points

  • Recently, approaches that leverage human video datasets for robot policy training have become increasingly prevalent.
  • RLHND turns the pre-trained Cosmos 3 video diffusion backbone into a deterministic clip-level feature extractor via clean-latent conditioning, carrying its learned priors on hand motion and hand-object interaction into tracking.
  • For pose estimation, RLHND (i) predicts hand poses with anatomically plausible joint angles and (ii) enables optional conditioning on the shape parameter to maintain consistent hand shape within the same video and even across videos recorded by the same actor.
  • Moreover, we demonstrate the utility of RLHND for robot learning through retargeting results and real-world robot experiments.

Sources (1)

  • [1]RLHND: Video Foundation Models as Physically Grounded Hand Trackers for Robot Learning
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 05:11 AM
    To this end, we propose RLHND, a video foundation model-based hand tracking model that jointly estimates hand pose and realistic tactile information from monocular egocentric videos.
    Recently, approaches that leverage human video datasets for robot policy training have become increasingly prevalent.

Extractive summary: sentences quoted from the sources.

Related