ResearchResearch paperImage, Video & 3D Generation1 source · Oct 3, 2026

NAMVIS: Next-Scale Autoregressive Multi-View Image Synthesis

We introduce NAMVIS, a diffusion-free framework that reformulates multi-view image synthesis as geometry-conditioned next-scale autoregression.

Key points

  • Sparse-view novel view synthesis is a central problem in 3D content creation, but diffusion-based approaches remain limited by iterative denoising, making multi-view generation expensive at inference time.
  • Instead of generating target views through repeated denoising, NAMVIS predicts discrete visual tokens through a small number of coarse-to-fine scale steps, while sampling all tokens within each scale and across target views in parallel.
  • To anchor this generation process to explicit camera geometry, we propose Multi-scale Projective Pose Encoding, which injects source and target camera transformations into both target-view self-attention and source-to-target cross-attention at every resolution.
  • NAMVIS further combines global conditioning with dense geometry-aware cross-attention, enabling the model to preserve source-view appearance while maintaining target-view consistency.

Sources (1)

  • [1]NAMVIS: Next-Scale Autoregressive Multi-View Image Synthesis
    Hugging Face Daily Papers · Oct 3, 12:00 AM
    We introduce NAMVIS, a diffusion-free framework that reformulates multi-view image synthesis as geometry-conditioned next-scale autoregression.
    Sparse-view novel view synthesis is a central problem in 3D content creation, but diffusion-based approaches remain limited by iterative denoising, making multi-view generation expensive at inference time.

Extractive summary: sentences quoted from the sources.

Related