When Listening Becomes Easier: Scrubbing Visual Cues for Shortcut-Free VLAs
Shortcut learning is a prevalent issue in robot learning.
Key points
- We focus on vision-language-action (VLA) models and discover that different vision-language model backbones exhibit substantially different levels of susceptibility to visual shortcut learning.
- We find that model behavior correlates with our proposed representation-level metric, action margin, which requires no policy rollouts.
- To boost models' attention to language, we introduce task scrubbing, a new domain-adversarial training method that decreases models' likelihood of using visual shortcuts and improves VLAs' generalization.
- Experiments in both simulation and the real world across multiple VLAs and visual cues show that task scrubbing improves out-of-distribution robustness and often eliminates visual shortcut learning.
Sources (1)
- [1]When Listening Becomes Easier: Scrubbing Visual Cues for Shortcut-Free VLAsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 09:11 PM
Shortcut learning is a prevalent issue in robot learning.
We focus on vision-language-action (VLA) models and discover that different vision-language model backbones exhibit substantially different levels of susceptibility to visual shortcut learning.
Extractive summary: sentences quoted from the sources.