GenIA: Generative Reconstruction with Test-Time Input Alignment
We introduce GenIA, a framework for test-time input-aligned generation that grounds SAM3D's generative prior in geometric and photometric observations without retraining the foundation model.
ProofPaper ↗
Key points
- Reconstructing complete 3D object assets from monocular or sparse multi-view observations remains challenging.
- Generative 3D foundation models can complete object geometry beyond the observed views, but their predictions may not faithfully reproduce the observed geometry, appearance, or pose.
- We improve object pose by deriving translation and scale from geometry while retaining the learned rotation prior, and align appearance through visibility-biased attention, cross-observation fusion, and differentiable rendering guidance during denoising.
- Across synthetic and real benchmarks, GenIA improves pose prediction and object reconstruction from monocular, multi-view, and dynamic inputs, outperforming recent optimization-based, per-frame image-to-3D, and video-to-4D methods.
Sources (1)
- [1]GenIA: Generative Reconstruction with Test-Time Input AlignmentarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 05:38 PM
We introduce GenIA, a framework for test-time input-aligned generation that grounds SAM3D's generative prior in geometric and photometric observations without retraining the foundation model.
Reconstructing complete 3D object assets from monocular or sparse multi-view observations remains challenging.
Extractive summary: sentences quoted from the sources.