DensiTok: Making Feed-Forward 3D Gaussian Splatting See More Views Than It Is Given
Feed-forward 3D Gaussian Splatting (3DGS) reconstructs a scene in a single forward pass, replacing per-scene optimization with a network trained across many scenes.
Key points
- We present DensiTok, a plug-in module for pretrained feed-forward 3DGS models that densifies their internal geometry tokens directly, making a frozen backbone behave as though it had observed many more views than it was given.
- DensiTok compresses those tokens into a compact latent space, completes the latents of the unobserved viewpoints in a single flow-matching step conditioned on camera geometry, and decodes them back into tokens that the original reconstruction heads.
- Completion in a low-dimensional latent space requires no image synthesis or additional encoder passes.
- Across three pretrained backbones and two benchmarks, DensiTok consistently improves sparse-view reconstruction and recovers much of the gap to dense-view reconstruction.
Sources (1)
- [1]DensiTok: Making Feed-Forward 3D Gaussian Splatting See More Views Than It Is GivenarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 08:29 AM
Feed-forward 3D Gaussian Splatting (3DGS) reconstructs a scene in a single forward pass, replacing per-scene optimization with a network trained across many scenes.
We present DensiTok, a plug-in module for pretrained feed-forward 3DGS models that densifies their internal geometry tokens directly, making a frozen backbone behave as though it had observed many more views than it was given.
Extractive summary: sentences quoted from the sources.