LEGO: A Lifting-Free Approach for Exocentric-to-Egocentric Video Generation
Generating an egocentric video from a single exocentric recording is a challenging case of novel view synthesis, as the two cameras share little overlap and much of the target view is unobserved.
Key points
- Current state-of-the-art methods reconstruct the scene explicitly by estimating depth, lifting the video into a point cloud, and re-rendering it from the egocentric camera to condition a video diffusion model.
- This deterministic mapping assigns each pixel to a single reprojected location, which preserves texture but translates depth errors into misplaced content.
- We ask what a video diffusion model should receive as its condition and propose a lifting-free answer: a learned view synthesizer, an LVSM-style transformer fine-tuned to render the egocentric view directly without depth, point clouds, or reprojection, resolving cross-view correspondence internally.
- We argue that this trade-off suits a diffusion generator, whose denoising training excels at restoring detail, so an effective condition should prioritize structural alignment over sharpness.
Sources (2)
- [1]LEGO: A Lifting-Free Approach for Exocentric-to-Egocentric Video GenerationHugging Face Daily Papers · Oct 8, 12:00 AM
Generating an egocentric video from a single exocentric recording is a challenging case of novel view synthesis, as the two cameras share little overlap and much of the target view is unobserved.
Current state-of-the-art methods reconstruct the scene explicitly by estimating depth, lifting the video into a point cloud, and re-rendering it from the egocentric camera to condition a video diffusion model.
- [2]LEGO: A Lifting-Free Approach for Exocentric-to-Egocentric Video GenerationarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 05:58 PM · same content
Extractive summary: sentences quoted from the sources.