AION
Research paperRobotics & Embodied AI1 source · Oct 6, 2026

ViDAL: A Visual Dynamics-Grounded Action Latent Space for Vision-Language-Action Models

We introduce ViDAL, a Visual Dynamics-grounded Action Latent Space that anchors continuous action latents in the future visual dynamics of the scene.

Key points

  • Vision-Language-Action (VLA) models have become a central paradigm for robot policy learning, which predict actions in three forms: raw action chunks, discrete action tokens, or continuous action latents.
  • However, existing action representations primarily model action trajectories, with limited consideration of the visual dynamics induced by these actions.
  • Specifically, ViDAL learns action latent space by training an Action Variational Autoencoder (Action VAE) to reconstruct action chunks while aligning its latent with future scene dynamics.
  • Empirically, ViDAL outperforms competitive baselines on LIBERO with 98.1% average success, improves a multi-task $π{0.5}$ policy on RoboTwin 2.0 from 54.3% to 65.5% (Clean) and from 33.2% to 43.1% (Random) success rates over 50 dual-arm tasks, and yields 20.0% and 23.4% absolute success-rate gains on real-world single-arm Franka and dual-arm ARX robot platforms.

Sources (1)

  • [1]ViDAL: A Visual Dynamics-Grounded Action Latent Space for Vision-Language-Action Models
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 11:02 AM
    We introduce ViDAL, a Visual Dynamics-grounded Action Latent Space that anchors continuous action latents in the future visual dynamics of the scene.
    Vision-Language-Action (VLA) models have become a central paradigm for robot policy learning, which predict actions in three forms: raw action chunks, discrete action tokens, or continuous action latents.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 6, 2026Adapting Vision-Language-Action Models to Unknown Visual Disruptions During Execution
  2. Oct 6, 2026StairVLA: Stage-Aware Hierarchical Action Generation for Vision-Language-Action Models
  3. Oct 6, 2026CARE: Certifying Acceleration for Vision-Language-Action Inference
  4. Oct 2, 2026FastOPD: On-Policy Distillation for Lightweight VLA Deployment
  5. Jul 30, 2026Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration
  6. Jul 28, 2026Gemini Robotics 2 brings whole body intelligence to robots

Related