GATOR: Generative and Agentic 3D Object Reconstruction From Casual Images
We present GATOR, a generative and agentic framework that recovers textured object assets and their scene-relative pose from one or more images.
ProofPaper ↗
Key points
- Reconstructing complete, scene-aligned 3D objects from casual images requires integrating sparse, uncertain observations and inferring surfaces hidden by occlusions.
- Our local modality mixer couples patch-aligned RGB, target-mask, and pointmap features before cross-view reasoning, preserving scene context while distinguishing the target from its surroundings.
- The generated asset initializes a multimodal agent, providing instance-specific geometry and pose for targeted structural and texture refinement through an observation-guided edit-render-review loop.
- Across synthetic objects, cluttered tabletops, and indoor scenes, GATOR achieves strong geometric and appearance fidelity while recovering scene-relative pose from sparse observations.
Sources (1)
- [1]GATOR: Generative and Agentic 3D Object Reconstruction From Casual ImagesarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 04:13 AM
We present GATOR, a generative and agentic framework that recovers textured object assets and their scene-relative pose from one or more images.
Reconstructing complete, scene-aligned 3D objects from casual images requires integrating sparse, uncertain observations and inferring surfaces hidden by occlusions.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 8, 2026Nvidia, Samsung back $90M round for AI agent startup Nous Research
- Oct 8, 2026Microsoft event debuts new AI-friendly hardware and Windows changes
- Oct 7, 2026Scaling Decision Optimization to 100 Million Variables and Beyond with mPDLP in NVIDIA cuOpt
- Oct 7, 2026Faster Scientific Image Analysis with NVIDIA cuPhoton
- Oct 7, 2026Google rolls out improved SynthID AI content detector, now available globally
- Oct 6, 2026Introducing Mistral Large 4