ResearchResearch paperImage, Video & 3D Generation1 source · Oct 8, 2026

GenIA: Generative Reconstruction with Test-Time Input Alignment

We introduce GenIA, a framework for test-time input-aligned generation that grounds SAM3D's generative prior in geometric and photometric observations without retraining the foundation model.

Key points

  • Reconstructing complete 3D object assets from monocular or sparse multi-view observations remains challenging.
  • Generative 3D foundation models can complete object geometry beyond the observed views, but their predictions may not faithfully reproduce the observed geometry, appearance, or pose.
  • We improve object pose by deriving translation and scale from geometry while retaining the learned rotation prior, and align appearance through visibility-biased attention, cross-observation fusion, and differentiable rendering guidance during denoising.
  • Across synthetic and real benchmarks, GenIA improves pose prediction and object reconstruction from monocular, multi-view, and dynamic inputs, outperforming recent optimization-based, per-frame image-to-3D, and video-to-4D methods.

Sources (1)

  • [1]GenIA: Generative Reconstruction with Test-Time Input Alignment
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 05:38 PM
    We introduce GenIA, a framework for test-time input-aligned generation that grounds SAM3D's generative prior in geometric and photometric observations without retraining the foundation model.
    Reconstructing complete 3D object assets from monocular or sparse multi-view observations remains challenging.

Extractive summary: sentences quoted from the sources.

Related