ResearchResearch paperRobotics & Embodied AI · Multimodal Models1 source · Oct 6, 2026

PAIR: Bridging Perception and Action in Vision-Language-Action Models

Vision-language-action (VLA) models map visual observations and language instructions to continuous robot actions.

Key points

  • We introduce PAIR, a framework that learns a shared perception-action representation between these two spaces.
  • During training, a Masked Action Autoencoder encodes expert action chunks into horizon-aligned Action Latent Tokens.
  • A Bridge Module extracts task-relevant features from the current visual-language representations.
  • These results support a shared intermediate representation as a useful interface between perception and action in continuous-action VLAs.

Sources (1)

  • [1]PAIR: Bridging Perception and Action in Vision-Language-Action Models
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 07:11 PM
    Vision-language-action (VLA) models map visual observations and language instructions to continuous robot actions.
    We introduce PAIR, a framework that learns a shared perception-action representation between these two spaces.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 6, 2026Adapting Vision-Language-Action Models to Unknown Visual Disruptions During Execution
  2. Oct 6, 2026StairVLA: Stage-Aware Hierarchical Action Generation for Vision-Language-Action Models
  3. Oct 6, 2026CARE: Certifying Acceleration for Vision-Language-Action Inference
  4. Oct 2, 2026FastOPD: On-Policy Distillation for Lightweight VLA Deployment
  5. Jul 30, 2026Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration
  6. Jul 28, 2026Gemini Robotics 2 brings whole body intelligence to robots

Related