PAIR: Bridging Perception and Action in Vision-Language-Action Models
Vision-language-action (VLA) models map visual observations and language instructions to continuous robot actions.
ProofPaper ↗
Key points
- We introduce PAIR, a framework that learns a shared perception-action representation between these two spaces.
- During training, a Masked Action Autoencoder encodes expert action chunks into horizon-aligned Action Latent Tokens.
- A Bridge Module extracts task-relevant features from the current visual-language representations.
- These results support a shared intermediate representation as a useful interface between perception and action in continuous-action VLAs.
Sources (1)
- [1]PAIR: Bridging Perception and Action in Vision-Language-Action ModelsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 07:11 PM
Vision-language-action (VLA) models map visual observations and language instructions to continuous robot actions.
We introduce PAIR, a framework that learns a shared perception-action representation between these two spaces.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 6, 2026Adapting Vision-Language-Action Models to Unknown Visual Disruptions During Execution
- Oct 6, 2026StairVLA: Stage-Aware Hierarchical Action Generation for Vision-Language-Action Models
- Oct 6, 2026CARE: Certifying Acceleration for Vision-Language-Action Inference
- Oct 2, 2026FastOPD: On-Policy Distillation for Lightweight VLA Deployment
- Jul 30, 2026Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration
- Jul 28, 2026Gemini Robotics 2 brings whole body intelligence to robots