ResearchResearch paperImage, Video & 3D Generation · Robotics & Embodied AI1 source · Oct 7, 2026

Video-Conditioned Generative Joint 2D-3D Hand Motion Recovery

To address this problem, we propose JoHan, a unified generative framework that recovers hand motion directly from video sequences without relying on intermediate per-frame pose predictions.

Key points

  • Recovering faithful 3D hand motion from video remains challenging due to frequent occlusions and incomplete visual observations, which make frame-wise pose estimates unreliable and temporally inconsistent.
  • Trained from scratch, our model jointly generates aligned 2D and 3D local hand pose sequences by learning their temporal dynamics and cross-representation correspondence.
  • The generated 2D trajectories exploit direct spatial and temporal cues from the 2D images to guide the following generative 3D motion reconstruction, while the learned motion prior promotes temporal consistency.
  • Their learned 2D-3D correspondence further enables recovery of the hand's global position and orientation relative to the camera.

Sources (1)

  • [1]Video-Conditioned Generative Joint 2D-3D Hand Motion Recovery
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 05:52 PM
    To address this problem, we propose JoHan, a unified generative framework that recovers hand motion directly from video sequences without relying on intermediate per-frame pose predictions.
    Recovering faithful 3D hand motion from video remains challenging due to frequent occlusions and incomplete visual observations, which make frame-wise pose estimates unreliable and temporally inconsistent.

Extractive summary: sentences quoted from the sources.

Related