ResearchResearch paperRobotics & Embodied AI1 source · Oct 7, 2026

ECHO: Embodied Camera Observations of Human Object Carrying

To support this task, we present Embodied Camera observations of Human Object carrying (ECHO), a large-scale synthetic dataset that pairs dense RGB-D scans of indoor scenes with recordings of an embodied human carrying everyday objects to context-appropriate destinations.

Key points

  • Embodied and assistive agents must do more than recognize objects: they must reason about where an object belongs given the layout of an environment and the habits of the people who live in it.
  • Existing RGB-D scan datasets reconstruct static rooms without human activity, while human-object-interaction datasets capture motion without a navigable, fully reconstructed scene or a ground-truth notion of an object's natural destination.
  • We introduce contextual object placement as a benchmark task: predicting an object's destination during an observed object-carrying episode.
  • ECHO is the first publicly available dataset to combine reconstructed scenes, human activity, natural language, and contextual-placement annotations.

Sources (1)

  • [1]ECHO: Embodied Camera Observations of Human Object Carrying
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 05:15 PM
    To support this task, we present Embodied Camera observations of Human Object carrying (ECHO), a large-scale synthetic dataset that pairs dense RGB-D scans of indoor scenes with recordings of an embodied human carrying everyday objects to context-appropriate destinations.
    Embodied and assistive agents must do more than recognize objects: they must reason about where an object belongs given the layout of an environment and the habits of the people who live in it.

Extractive summary: sentences quoted from the sources.

Related