Pose-Free Feed-Forward 3D Inpainting via Learnable Mask Attention and Support Token Refinement
To overcome this limitation, we present FreeInpaint, a novel feed-forward framework that generates complete and 3D-consistent scenes directly from unposed multi-view images with masked regions.
ProofPaper ↗
Key points
- 3D scene inpainting aims to recover missing or occluded regions in edited 3D scenes, while ensuring geometric and textural consistency.
- At its core, FreeInpaint extends a 3D foundation model to propagate masked regions from a reference view to other unposed views, bridging 3D reconstruction and scene inpainting while preserving the model's native ability to recover camera poses and scene geometry.
- To address this, we introduce a Learnable Mask Attention mechanism that preserves the spatial anchoring of reliable observations while allowing masked regions to progressively absorb useful context in deeper layers.
- Therefore, we propose a Support Token Refinement strategy, which injects diffusion-generated support evidence as confidence-weighted auxiliary tokens to refine under-observed regions while preserving the original spatial anchor.
Sources (1)
- [1]Pose-Free Feed-Forward 3D Inpainting via Learnable Mask Attention and Support Token RefinementarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 12:38 PM
To overcome this limitation, we present FreeInpaint, a novel feed-forward framework that generates complete and 3D-consistent scenes directly from unposed multi-view images with masked regions.
3D scene inpainting aims to recover missing or occluded regions in edited 3D scenes, while ensuring geometric and textural consistency.
Extractive summary: sentences quoted from the sources.