Video-Conditioned Generative Joint 2D-3D Hand Motion Recovery
To address this problem, we propose JoHan, a unified generative framework that recovers hand motion directly from video sequences without relying on intermediate per-frame pose predictions.
ProofPaper ↗
Key points
- Recovering faithful 3D hand motion from video remains challenging due to frequent occlusions and incomplete visual observations, which make frame-wise pose estimates unreliable and temporally inconsistent.
- Trained from scratch, our model jointly generates aligned 2D and 3D local hand pose sequences by learning their temporal dynamics and cross-representation correspondence.
- The generated 2D trajectories exploit direct spatial and temporal cues from the 2D images to guide the following generative 3D motion reconstruction, while the learned motion prior promotes temporal consistency.
- Their learned 2D-3D correspondence further enables recovery of the hand's global position and orientation relative to the camera.
Sources (1)
- [1]Video-Conditioned Generative Joint 2D-3D Hand Motion RecoveryarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 05:52 PM
To address this problem, we propose JoHan, a unified generative framework that recovers hand motion directly from video sequences without relying on intermediate per-frame pose predictions.
Recovering faithful 3D hand motion from video remains challenging due to frequent occlusions and incomplete visual observations, which make frame-wise pose estimates unreliable and temporally inconsistent.
Extractive summary: sentences quoted from the sources.