Position Forcing: Self-Conditioning 3D Generation
Recent single-stage 3D generative models commonly adopt VecSet representations, encoding 3D shapes as unordered sets of latent tokens.
Key points
- We observe that, despite the absence of explicit positional conditioning, VecSet tokens retain recoverable spatial correspondences.
- Building on this observation, we propose Position Forcing, a position-based self-conditioning framework.
- During denoising, Position Forcing recovers token positions from the current clean latent estimate, quantizes them at progressively finer resolutions according to the denoising stage, and feeds the resulting positional encodings back into the diffusion Transformer.
- Experiments demonstrate that Position Forcing achieves strong performance among single-stage 3D generative methods and outperforms several competitive multi-stage approaches.
Sources (1)
- [1]Position Forcing: Self-Conditioning 3D GenerationarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 04:23 PM
Recent single-stage 3D generative models commonly adopt VecSet representations, encoding 3D shapes as unordered sets of latent tokens.
We observe that, despite the absence of explicit positional conditioning, VecSet tokens retain recoverable spatial correspondences.
Extractive summary: sentences quoted from the sources.