AION
Research paperRobotics & Embodied AI1 source · Oct 8, 2026

FOCUS: From Privileged States to RGB-D with Controlled Modality Switching and Representation Alignment

Vision-based reinforcement learning for robotic manipulation is sample-inefficient because RGB-D observations are high-dimensional and noisy.

Key points

  • Privileged state information available in simulation can accelerate training, but its absence at test time creates a train-test modality gap.
  • We propose FOCUS, a single-stage PPO framework that trains the critic on privileged state while automatically regulating whether the actor collects rollouts from RGB-D or privileged-state latents.
  • Regulation is driven by the KL divergence between the action distributions induced by the two modalities, while representation alignment encourages consistent action selection across them.
  • As they align, RGB-D exposure increases, shifting on-policy training toward the RGB-D inputs used at test time.

Sources (1)

Extractive summary: sentences quoted from the sources.