Do Vision-Language-Action Models Understand Instructions? A Mechanistic Interpretability Study on Language Grounding
Vision-Language-Action models are designed to generalise across environments and task descriptions, raising the question of whether their action generation actually depends on the language instruction, or whether they largely rely on visual cues and superficial correlations.
Key points
- Robustness to variance in the visual and linguistic observation space is critical for real-world deployment, yet VLAs lack explicit grounding modules and instead rely on the intrinsic language grounding capabilities of their Vision-Language model backbones.
- For this reason, we conduct a controlled mechanistic interpretability study on the language grounding capabilities of two state-of-the-art Vision-Language-Action models, $π{0.5}$ and GR00T N1.7, by applying activation and attribution patching to the residual stream of the action generation modules.
- We systematically corrupt the task instruction of input samples of the LIBERO benchmark following five strategies: synonym replacement, semantic scaling, directional corruption, random object substitution, and empty string.
- Our experiments find that both models are comparatively insensitive to abstract rephrasing and to referencing non-existent objects, but react strongly to empty task descriptions and, especially, to directional language.
Sources (1)
- [1]Do Vision-Language-Action Models Understand Instructions? A Mechanistic Interpretability Study on Language GroundingarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 02:48 PM
Vision-Language-Action models are designed to generalise across environments and task descriptions, raising the question of whether their action generation actually depends on the language instruction, or whether they largely rely on visual cues and superficial correlations.
Robustness to variance in the visual and linguistic observation space is critical for real-world deployment, yet VLAs lack explicit grounding modules and instead rely on the intrinsic language grounding capabilities of their Vision-Language model backbones.
Extractive summary: sentences quoted from the sources.