DIVA: Dual-Space Intent-Aware Visual Attenuation for Vision-Language-Action Policies
We introduce DIVA, a Dual-Space Intent-Aware Visual Attenuation module with an anchor-then-attenuate design.
ProofPaper ↗
Key points
- Vision-language-action (VLA) policies typically feed dense visual patch tokens into a language-action backbone, preserving scene context but offering no explicit mechanism to regulate how strongly different visual tokens influence policy computation.
- DIVA combines high-level task intent with low-level visual evidence to estimate patch-wise relevance anchors, then applies them in two complementary spaces: it reweights projected visual tokens before backbone entry and persistently attenuates low-relevance visual states within the backbone.
- DIVA preserves the full visual token sequence and requires no external grounding supervision.
- Real-world experiments further show consistent gains under task-irrelevant visual perturbations, supporting the robustness of intent-aware visual attenuation beyond simulation.
Sources (1)
- [1]DIVA: Dual-Space Intent-Aware Visual Attenuation for Vision-Language-Action PoliciesarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 09:37 PM
We introduce DIVA, a Dual-Space Intent-Aware Visual Attenuation module with an anchor-then-attenuate design.
Vision-language-action (VLA) policies typically feed dense visual patch tokens into a language-action backbone, preserving scene context but offering no explicit mechanism to regulate how strongly different visual tokens influence policy computation.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 6, 2026Adapting Vision-Language-Action Models to Unknown Visual Disruptions During Execution
- Oct 6, 2026StairVLA: Stage-Aware Hierarchical Action Generation for Vision-Language-Action Models
- Oct 6, 2026CARE: Certifying Acceleration for Vision-Language-Action Inference
- Oct 2, 2026FastOPD: On-Policy Distillation for Lightweight VLA Deployment
- Jul 30, 2026Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration
- Jul 28, 2026Gemini Robotics 2 brings whole body intelligence to robots