ResearchResearch paperImage, Video & 3D Generation · Computer Vision1 source · Oct 6, 2026

Depth-to-RGB: Repurposing a Frozen Depth Estimator for Geometry-Guided Compositing

Our Depth-to-RGB (D2R) framework predicts composite depth for a scene not yet observed in the RGB inputs.

Key points

  • Reference-based object compositing inserts or replaces an object using a background image, a reference image, and a 2D compositing mask.
  • It learns reference-conditioned corrections to a frozen depth estimator using encoder features of paired completed scenes as targets.
  • The unchanged decoder maps the corrected representation to the intended scene's depth, which a separately trained renderer holds fixed during RGB synthesis.
  • We also introduce AnyInsertion++ with paired in-distribution and category-disjoint splits to evaluate generalization beyond compositing training categories.

Sources (1)

Extractive summary: sentences quoted from the sources.

Related