A Balanced Data Diet: Addressing the Exploration Bottleneck in Mega-Scale RL for Robot Control
While sim-to-real reinforcement learning (RL) has proven to be a useful tool for this goal, current RL pipelines depend on engineering-heavy, per-task structural priors such as shaped rewards and demonstrations.
ProofPaper ↗
Key points
- General-purpose robots must perform a wide range of tasks from agile locomotion to dexterous manipulation.
- To mitigate this, we introduce Success Guided Sampling (SGS), a simple adaptive sampler that concentrates RL training on task configurations around the frontier of the policy's capabilities.
- Across experiments using up to $2^{20}$ (over one million) parallel environments, SGS enables RL to solve challenging multi-terrain quadruped locomotion and contact-rich assembly tasks that prior methods fail to solve.
- Finally, we distill the learned manipulation policies into RGB-based policies and demonstrate zero-shot transfer to several challenging assembly tasks on real hardware.
Sources (1)
- [1]A Balanced Data Diet: Addressing the Exploration Bottleneck in Mega-Scale RL for Robot ControlarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 05:59 PM
While sim-to-real reinforcement learning (RL) has proven to be a useful tool for this goal, current RL pipelines depend on engineering-heavy, per-task structural priors such as shaped rewards and demonstrations.
General-purpose robots must perform a wide range of tasks from agile locomotion to dexterous manipulation.
Extractive summary: sentences quoted from the sources.