AION
Research paperImage, Video & 3D Generation · Multimodal Models · Computer Vision2 sources · Oct 8, 2026

LEGO: A Lifting-Free Approach for Exocentric-to-Egocentric Video Generation

Generating an egocentric video from a single exocentric recording is a challenging case of novel view synthesis, as the two cameras share little overlap and much of the target view is unobserved.

Key points

  • Current state-of-the-art methods reconstruct the scene explicitly by estimating depth, lifting the video into a point cloud, and re-rendering it from the egocentric camera to condition a video diffusion model.
  • This deterministic mapping assigns each pixel to a single reprojected location, which preserves texture but translates depth errors into misplaced content.
  • We ask what a video diffusion model should receive as its condition and propose a lifting-free answer: a learned view synthesizer, an LVSM-style transformer fine-tuned to render the egocentric view directly without depth, point clouds, or reprojection, resolving cross-view correspondence internally.
  • We argue that this trade-off suits a diffusion generator, whose denoising training excels at restoring detail, so an effective condition should prioritize structural alignment over sharpness.

Sources (2)

Extractive summary: sentences quoted from the sources.