FOCUS: From Privileged States to RGB-D with Controlled Modality Switching and Representation Alignment
Vision-based reinforcement learning for robotic manipulation is sample-inefficient because RGB-D observations are high-dimensional and noisy.
Key points
- Privileged state information available in simulation can accelerate training, but its absence at test time creates a train-test modality gap.
- We propose FOCUS, a single-stage PPO framework that trains the critic on privileged state while automatically regulating whether the actor collects rollouts from RGB-D or privileged-state latents.
- Regulation is driven by the KL divergence between the action distributions induced by the two modalities, while representation alignment encourages consistent action selection across them.
- As they align, RGB-D exposure increases, shifting on-policy training toward the RGB-D inputs used at test time.
Sources (1)
- [1]FOCUS: From Privileged States to RGB-D with Controlled Modality Switching and Representation AlignmentarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 02:45 AM
Vision-based reinforcement learning for robotic manipulation is sample-inefficient because RGB-D observations are high-dimensional and noisy.
Privileged state information available in simulation can accelerate training, but its absence at test time creates a train-test modality gap.
Extractive summary: sentences quoted from the sources.