ResearchResearch paperRobotics & Embodied AI · Image, Video & 3D Generation · Multimodal Models1 source · Oct 6, 2026

4D-HOF: Hand-Object Flow Matching for Feed-Forward 4D Interaction Reconstruction

We introduce 4D-HOF, a feed-forward framework that reconstructs 4D hand-object interactions from coarse but informative estimates produced by vision foundation models.

Key points

  • Existing methods for 4D hand-object reconstruction often rely on costly per-sequence optimization, while generative approaches typically synthesize interactions from random noise, which can lead to unstable interaction prediction.
  • Concretely, we learn a conditional flow matching model that transports foundation-model-derived hand-object states toward an interaction manifold, allowing the model to correct errors in translation, rotation, and alignment in a feed-forward manner.
  • By training the generative model on diverse datasets, 4D-HOF generalizes robustly to challenging in-the-wild scenarios.
  • Experiments on out-of-domain benchmarks show that 4D-HOF achieves state-of-the-art performance, producing more stable and accurate 4D hand-object reconstructions.

Sources (1)

  • [1]4D-HOF: Hand-Object Flow Matching for Feed-Forward 4D Interaction Reconstruction
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 05:59 PM
    We introduce 4D-HOF, a feed-forward framework that reconstructs 4D hand-object interactions from coarse but informative estimates produced by vision foundation models.
    Existing methods for 4D hand-object reconstruction often rely on costly per-sequence optimization, while generative approaches typically synthesize interactions from random noise, which can lead to unstable interaction prediction.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 6, 2026Denoising Hierarchical Representations: Joint Continuous Diffusion for Language Modeling
  2. Oct 6, 2026FlowCF: Sparse Counterfactual Explanations for Mixed-Type Tabular Data using Flow Matching
  3. Oct 6, 2026Cylindrical Geodesic Flow Matching for Quasiperiodic Physiological Signal Transformation
  4. Oct 6, 2026Sensor Geometry as a Flow-Matching Prior for Multi-Channel Brain Signals
  5. Oct 6, 2026StairVLA: Stage-Aware Hierarchical Action Generation for Vision-Language-Action Models
  6. Oct 6, 2026ESP: Energy-Score Policy for One-Step Multimodal Action Generation

Related