VLA-ACL: Action-Consistent Visual Token Pruning for Efficient Vision-Language-Action Models
We introduce VLA-ACL (Action Consistency Learning), which learns a lightweight visual token pruning policy through action-level supervision while keeping the base VLA model entirely frozen.
ProofPaper ↗
Key points
- Vision-Language-Action (VLA) models achieve strong robotic manipulation performance but incur high computational costs from processing long token sequences at every control step, limiting real-time deployment.
- Visual token pruning offers a direct solution, as visual patches dominate the input sequence and contain considerable redundancy.
- Experiments on LIBERO and real-world manipulation tasks show that VLA-ACL prunes up to 87.5% of visual tokens while retaining competitive performance, reduces computation by up to 75%, and achieves a 1.5x inference speedup.
- These results establish a stronger performance-efficiency trade-off than existing frozen-VLA pruning methods and demonstrate the value of action-level supervision for visual token selection.
Sources (1)
- [1]VLA-ACL: Action-Consistent Visual Token Pruning for Efficient Vision-Language-Action ModelsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 10:47 AM
We introduce VLA-ACL (Action Consistency Learning), which learns a lightweight visual token pruning policy through action-level supervision while keeping the base VLA model entirely frozen.
Vision-Language-Action (VLA) models achieve strong robotic manipulation performance but incur high computational costs from processing long token sequences at every control step, limiting real-time deployment.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 6, 2026CARE: Certifying Acceleration for Vision-Language-Action Inference
- Oct 6, 2026Frozen Models, Evolving Expertise: Model-Agnostic Learning from Deployment Experience for Multimodal Medical AI
- Sep 30, 2026Expanding AI Storage Access with NVIDIA cuObject and the NVIDIA SCADA Server SDK
- Sep 28, 2026openai/openai-python v3.20.0
- Jul 30, 2026Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration
- Jun 10, 2026DiffusionGemma: 4x faster text generation