4D-HOF: Hand-Object Flow Matching for Feed-Forward 4D Interaction Reconstruction
We introduce 4D-HOF, a feed-forward framework that reconstructs 4D hand-object interactions from coarse but informative estimates produced by vision foundation models.
ProofPaper ↗
Key points
- Existing methods for 4D hand-object reconstruction often rely on costly per-sequence optimization, while generative approaches typically synthesize interactions from random noise, which can lead to unstable interaction prediction.
- Concretely, we learn a conditional flow matching model that transports foundation-model-derived hand-object states toward an interaction manifold, allowing the model to correct errors in translation, rotation, and alignment in a feed-forward manner.
- By training the generative model on diverse datasets, 4D-HOF generalizes robustly to challenging in-the-wild scenarios.
- Experiments on out-of-domain benchmarks show that 4D-HOF achieves state-of-the-art performance, producing more stable and accurate 4D hand-object reconstructions.
Sources (1)
- [1]4D-HOF: Hand-Object Flow Matching for Feed-Forward 4D Interaction ReconstructionarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 05:59 PM
We introduce 4D-HOF, a feed-forward framework that reconstructs 4D hand-object interactions from coarse but informative estimates produced by vision foundation models.
Existing methods for 4D hand-object reconstruction often rely on costly per-sequence optimization, while generative approaches typically synthesize interactions from random noise, which can lead to unstable interaction prediction.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 6, 2026Denoising Hierarchical Representations: Joint Continuous Diffusion for Language Modeling
- Oct 6, 2026FlowCF: Sparse Counterfactual Explanations for Mixed-Type Tabular Data using Flow Matching
- Oct 6, 2026Cylindrical Geodesic Flow Matching for Quasiperiodic Physiological Signal Transformation
- Oct 6, 2026Sensor Geometry as a Flow-Matching Prior for Multi-Channel Brain Signals
- Oct 6, 2026StairVLA: Stage-Aware Hierarchical Action Generation for Vision-Language-Action Models
- Oct 6, 2026ESP: Energy-Score Policy for One-Step Multimodal Action Generation