AION
Technique

Vision-language-action models

Also known as: VLA, vision-language-action

51stories this week
52last 30 days
54all time

Timeline

  1. Oct 8, 2026 · Research paper · 1 source
    VersaCamVLA: Camera-Configurable VLA Policies for Robotic Manipulation
    To overcome these limitations, we propose VersaCamVLA, a camera-configurable framework that decouples camera-set representation from action learning.
  2. Oct 8, 2026 · Research paper · 1 source
    VioLA: Learning Generalist Humanoid Control Policies from Human Data
    We introduce VioLA, a generalist humanoid policy that predicts body and hand motion latents instead of joint commands.
  3. Oct 8, 2026 · Research paper · 2 sources
    Embodied Turing Machines: Stateful Code for Robot Recursive Self-Improvement
    We propose a different view: the embodied world is an Embodied Turing Machine, whose tape is the robot and environment state and rules are the policy.
  4. Oct 8, 2026 · Research paper · 1 source
    PLaW-VLA: Predictive Latent World Modeling for Vision-Language-Action Policies
    Learning to predict how the world evolves can provide vision-language-action (VLA) policies with predictive context for long-horizon control, but its effectiveness depends on what future representation is modeled and how it conditions action generation.
  5. Oct 8, 2026 · Research paper · 1 source
    Residual Modeling Closes the Regression and Generative Policy Gap in Robot Learning
    We revisit this gap from the perspective of statistical modeling: how action-prediction residuals shape policy optimization.
  6. Oct 8, 2026 · Research paper · 1 source
    RESETTLE: Robotic Recovery through Disagreement-Triggered Retrieval and Efficient Corrective Control
    To address these challenges, we introduce RESETTLE(Robotic rEcovery through diSagrEement-Triggered reTrievaL and Efficient Corrective Control), a model-agnostic framework that provides computationally efficient recovery at the action-execution interface of frozen robot policies.
  7. Oct 8, 2026 · Research paper · 1 source
    ContourVLA: A Closed-Loop Perception-Action Contour Policy for Generalized Referring Expression Segmentation
    We introduce ContourVLA, a vision-language-action policy that recasts GRES as a closed-loop visuomotor process, in which editable contours serve as explicit policy states that condition multimodal perception and are updated by geometric action chunks.
  8. Oct 8, 2026 · Research paper · 1 source
    Recompose and Refine Latent Reasoning Flows for Vision-Language-Action Models
    Latent reasoning enables vision-language-action (VLA) models to transform multimodal observations into task-relevant internal states before generating continuous robot actions.
  9. Oct 8, 2026 · Research paper · 1 source
    ManiUnit: A Manipulation Skill Dataset and Benchmark for Long-Horizon Tasks
    Long-horizon mobile manipulation requires a robot to navigate multi-room environments and execute a sequence of manipulation skills under a single natural language instruction.
  10. Oct 8, 2026 · Research paper · 1 source
    Humanoid World Action Model With Joint State--Action Generation
    We propose HWAM, a Humanoid World Action Model with joint state--action generation, which makes the robot's post-execution proprioceptive state an explicit prediction target.
  11. Oct 8, 2026 · Research paper · 1 source
    REACT: Rolling Denoising and Dual Decoupling for Reactive Robot Control with VLA Models
    Flow-based vision-language-action (VLA) models generate action chunks for temporally coherent robot motion, but chunked control creates a fundamental closed-loop trade-off: long chunks provide smooth execution, whereas frequent replanning improves reactivity at the cost of action discontinuities.
  12. Oct 8, 2026 · Research paper · 1 source
    CAPABLE: Capability-Aware Policy Adaptation via Behavioral Latent Encoding
    We introduce CAPABLE, a unified capability-aware adaptation framework for frozen VLAs that integrates self-supervised capability inference with residual reinforcement learning.
  13. Oct 8, 2026 · Research paper · 1 source
    Tell Robot What Not to Do: A Negation Understanding Perspective
    To this end, we propose NegaAlign, a parameter-efficient, plug-and-play framework that extends pretrained VLAs to follow negated instructions through image-language supervision alone.
  14. Oct 8, 2026 · Research paper · 1 source
    PathTime-VLA: Path-Time Decoupling for Factorized Post-Training of Vision-Language-Action Policies
    Vision-Language-Action (VLA) policies typically predict actions at fixed time intervals, coupling the route a robot follows with its execution pace.
  15. Oct 8, 2026 · Research paper · 1 source
    WARP-VLA: Wrist-Camera Adaptation for View-Robust Policy Execution in Vision-Language-Action Models
    To address this, we propose WARP-VLA, a camera-view robust VLA for diverse wrist camera configurations.
  16. Oct 8, 2026 · Research paper · 1 source
    RoboAware: Learning to Coordinate Embodied Skills from Counterfactual Outcomes
    We present RoboAware, which builds on coding agents' skill orchestration by learning only a state-conditioned responsibility coordinator from counterfactual outcomes.
  17. Oct 8, 2026 · Research paper · 1 source
    Experience-Guided Initiation Search for Learned Skills in Skill Composition
    We propose EVIS, an Experience-Guided and Behavior-Validated Initiation Search framework for discovering reliable initiation configurations under limited target interaction.
  18. Oct 8, 2026 · Research paper · 1 source
    Rewiring Semantics, Dynamics, and Control: A Simple yet Effective Action-Centric Tri-Stream Transformer
    Vision-Language-Action (VLA) models have emerged as a prominent framework for complex robotic manipulation, building on the strong semantic understanding of pretrained Vision-Language Models (VLMs).
  19. Oct 8, 2026 · Research paper · 1 source
    SimVLA: Zero-Shot Sim-to-Real VLA Learning for Mobile Manipulation
    We introduce SimVLA, an end-to-end framework that trains VLAs entirely on synthetic simulation data without teleoperation for mobile manipulation.
  20. Oct 8, 2026 · Research paper · 1 source
    VGGTWorld-VLA: Intent-Conditioned 3D World Evolution for Autonomous Driving
    We propose VGGTWorld-VLA, an intention-conditioned extension of VGGT-World for controllable 3D world evolution in autonomous driving.
  21. Oct 7, 2026 · Research paper · 1 source
    When Listening Becomes Easier: Scrubbing Visual Cues for Shortcut-Free VLAs
    Shortcut learning is a prevalent issue in robot learning.
  22. Oct 7, 2026 · Research paper · 1 source
    NavGPT-3: Harnessing Context in a Hierarchical Navigation Runtime
    We present NavGPT-3, a harness that connects the two models, with an OS-like runtime built above it: reasoning, acting, and monitoring run as threads with their own context, tools, and permissions, while the runtime schedules them and decides which thread controls the robot's motion, so that the robot can react to sudden real-world events through interruption and thread switching.
  23. Oct 7, 2026 · Research paper · 1 source
    Rephrase Before You Act: Characterizing and Mitigating Language Sensitivity in Vision-Language-Action Models
    Vision-language-action models (VLAs) are strikingly sensitive to instruction phrasing and do not inherit the language robustness of the vision-language models they are built on.
  24. Oct 7, 2026 · Research paper · 2 sources
    Q-Learning with Scalar Adjoint Matching
    Adjoint matching offers a principled way to update the flow model itself by propagating value information from the final action back to each flow step, but it requires a vector--Jacobian product through the policy at every step, a cost that grows with the number of flow steps and the policy size.
  25. Oct 7, 2026 · Research paper · 1 source
    Explicit Geometric Chain-of-Thought for Vision-Language-Action in Autonomous Driving
    Vision-language-action (VLA) models have emerged as a promising paradigm for autonomous driving.
  26. Oct 7, 2026 · Research paper · 1 source
    OpenViTac: Learning and Benchmarking Visuo-Tactile Policies in a Unified Sim-and-Real Framework
    To address this gap, we introduce OpenViTac, a visuo-tactile manipulation benchmark for evaluating robot policies across simulation and the real world.
  27. Oct 7, 2026 · Research paper · 1 source
    Do Vision-Language-Action Models Understand Instructions? A Mechanistic Interpretability Study on Language Grounding
    Vision-Language-Action models are designed to generalise across environments and task descriptions, raising the question of whether their action generation actually depends on the language instruction, or whether they largely rely on visual cues and superficial correlations.
  28. Oct 7, 2026 · Research paper · 1 source
    Many Ways to Succeed: Diversity-Driven RL Fine-Tuning for VLA Generalization
    Reinforcement learning (RL) fine-tuning improves vision-language-action (VLA) policies through closed-loop experience, yet generalization beyond the fine-tuning distribution remains limited.
  29. Oct 7, 2026 · Research paper · 1 source
    Juno: Taming Predictive Latents for Vision-Language-Action Models
    Joint-embedding predictive architectures (JEPAs) predict masked or future observations in representation space, offering a natural source of predictive latents for vision-language-action (VLA) models.
  30. Oct 7, 2026 · Research paper · 1 source
    YUBI-STAG: Contact and Semantic-Rich Alignment for VLAs via Automated Video-Language Grounding
    We introduce YUBI-STAG, a framework for Spatio-Temporal Annotation and Grounding that automatically enriches manipulation demonstrations with interaction-rich semantics to align pretrained VLAs with fine-grained manipulation language.

Often appears with