Targeted Modality Dropout for Real-Robot Manipulation Robust to Intermittent Vision Loss
In this paper, we introduce Targeted Modality Dropout (TMD), in which the dependence on each modality is estimated using attention and the most dominant modality is selectively dropped.
ProofPaper ↗
Key points
- Imitation learning policies that integrate multiple sensory modalities are prone to overreliance on a dominant modality, such as vision, during training, which can disrupt policy execution when that modality is lost at inference time.
- This is combined with entropy regularization over the dependence distribution.
- Through real-robot evaluation using a bimanual manipulator, we show that under vision loss the success rate of the baseline policy drops substantially, whereas TMD sustains task execution.
- In contrast, a conventional dropout that selects the dropped modality at random, without the entropy regularization, fails on many tasks even without vision loss.
Sources (1)
- [1]Targeted Modality Dropout for Real-Robot Manipulation Robust to Intermittent Vision LossarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 07:07 AM
In this paper, we introduce Targeted Modality Dropout (TMD), in which the dependence on each modality is estimated using attention and the most dominant modality is selectively dropped.
Imitation learning policies that integrate multiple sensory modalities are prone to overreliance on a dominant modality, such as vision, during training, which can disrupt policy execution when that modality is lost at inference time.
Extractive summary: sentences quoted from the sources.