ResearchResearch paperImage, Video & 3D Generation · Computer Vision1 source · Oct 8, 2026

Pose-Free Feed-Forward 3D Inpainting via Learnable Mask Attention and Support Token Refinement

To overcome this limitation, we present FreeInpaint, a novel feed-forward framework that generates complete and 3D-consistent scenes directly from unposed multi-view images with masked regions.

Key points

  • 3D scene inpainting aims to recover missing or occluded regions in edited 3D scenes, while ensuring geometric and textural consistency.
  • At its core, FreeInpaint extends a 3D foundation model to propagate masked regions from a reference view to other unposed views, bridging 3D reconstruction and scene inpainting while preserving the model's native ability to recover camera poses and scene geometry.
  • To address this, we introduce a Learnable Mask Attention mechanism that preserves the spatial anchoring of reliable observations while allowing masked regions to progressively absorb useful context in deeper layers.
  • Therefore, we propose a Support Token Refinement strategy, which injects diffusion-generated support evidence as confidence-weighted auxiliary tokens to refine under-observed regions while preserving the original spatial anchor.

Sources (1)

  • [1]Pose-Free Feed-Forward 3D Inpainting via Learnable Mask Attention and Support Token Refinement
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 12:38 PM
    To overcome this limitation, we present FreeInpaint, a novel feed-forward framework that generates complete and 3D-consistent scenes directly from unposed multi-view images with masked regions.
    3D scene inpainting aims to recover missing or occluded regions in edited 3D scenes, while ensuring geometric and textural consistency.

Extractive summary: sentences quoted from the sources.

Related