ECHO: Embodied Camera Observations of Human Object Carrying
To support this task, we present Embodied Camera observations of Human Object carrying (ECHO), a large-scale synthetic dataset that pairs dense RGB-D scans of indoor scenes with recordings of an embodied human carrying everyday objects to context-appropriate destinations.
ProofPaper ↗
Key points
- Embodied and assistive agents must do more than recognize objects: they must reason about where an object belongs given the layout of an environment and the habits of the people who live in it.
- Existing RGB-D scan datasets reconstruct static rooms without human activity, while human-object-interaction datasets capture motion without a navigable, fully reconstructed scene or a ground-truth notion of an object's natural destination.
- We introduce contextual object placement as a benchmark task: predicting an object's destination during an observed object-carrying episode.
- ECHO is the first publicly available dataset to combine reconstructed scenes, human activity, natural language, and contextual-placement annotations.
Sources (1)
- [1]ECHO: Embodied Camera Observations of Human Object CarryingarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 05:15 PM
To support this task, we present Embodied Camera observations of Human Object carrying (ECHO), a large-scale synthetic dataset that pairs dense RGB-D scans of indoor scenes with recordings of an embodied human carrying everyday objects to context-appropriate destinations.
Embodied and assistive agents must do more than recognize objects: they must reason about where an object belongs given the layout of an environment and the habits of the people who live in it.
Extractive summary: sentences quoted from the sources.