RobotWorld: Benchmarking Multimodal Agents for Robot Use Across Diverse Tasks and Embodiments
To investigate this, we introduce RobotWorld, a challenging simulation testbed for robot use: turning instructions and observations into physical task execution through robot interfaces.
ProofPaper ↗
Key points
- General-purpose agents increasingly write code, use tools, and complete complex digital tasks, raising the question of how far these capabilities carry into the physical world.
- Furthermore, we find that current agents can construct sophisticated perception and control workflows, including image segmentation, camera calibration, spatial estimation, and dynamics-based computation.
- These capabilities, however, do not consistently compose into successful behaviour: agents lose task-relevant object states despite reaching commanded poses, fail to correct ineffective actions, recover too late, or mistake unfinished tasks for completion.
- By linking these outcomes to execution behaviour, RobotWorld provides both a rigorous proving ground and an empirical account of the remaining capability gaps, thereby establishing concrete targets for training and designing more reliable physical-world agents.
Sources (1)
- [1]RobotWorld: Benchmarking Multimodal Agents for Robot Use Across Diverse Tasks and EmbodimentsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 04:55 PM
To investigate this, we introduce RobotWorld, a challenging simulation testbed for robot use: turning instructions and observations into physical task execution through robot interfaces.
General-purpose agents increasingly write code, use tools, and complete complex digital tasks, raising the question of how far these capabilities carry into the physical world.
Extractive summary: sentences quoted from the sources.