Agentic RSR: Real-to-Sim-to-Real through Scene Reconstruction and Execution-Grounded Robot Policies
We present Agentic Real-to-Sim-to-Real (Agentic RSR), a framework that links scene reconstruction, policy development, and real-robot execution through the same manipulation task.
ProofPaper ↗
Key points
- A simulation of a real robot workspace must preserve task-relevant interactions, while policies developed in it must operate on observations available to the real robot.
- Given a workspace video, a task description, and a known robot model, an agent recovers metric scale, iteratively refines the scene using visual feedback, and checks task-relevant interactions in MuJoCo.
- A coding agent then develops an executable policy, progressing from privileged object poses to visual observations and randomized simulation.
- Across 18 reconstructed scenes involving two robots, the mean four-view Depth MAE against reference depth estimates is 0.1057 m, the mean Lab $ΔE{76}$ is 11.04, and the mean grayscale SSIM is 0.6990.
Sources (1)
- [1]Agentic RSR: Real-to-Sim-to-Real through Scene Reconstruction and Execution-Grounded Robot PoliciesarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 05:37 PM
We present Agentic Real-to-Sim-to-Real (Agentic RSR), a framework that links scene reconstruction, policy development, and real-robot execution through the same manipulation task.
A simulation of a real robot workspace must preserve task-relevant interactions, while policies developed in it must operate on observations available to the real robot.
Extractive summary: sentences quoted from the sources.