Vision-language-action models
Also known as: VLA, vision-language-action
51stories this week
52last 30 days
54all time
Timeline
- Oct 8, 2026 · Research paper · 1 sourceVersaCamVLA: Camera-Configurable VLA Policies for Robotic ManipulationTo overcome these limitations, we propose VersaCamVLA, a camera-configurable framework that decouples camera-set representation from action learning.
- Oct 8, 2026 · Research paper · 1 sourceVioLA: Learning Generalist Humanoid Control Policies from Human DataWe introduce VioLA, a generalist humanoid policy that predicts body and hand motion latents instead of joint commands.
- Oct 8, 2026 · Research paper · 2 sourcesEmbodied Turing Machines: Stateful Code for Robot Recursive Self-ImprovementWe propose a different view: the embodied world is an Embodied Turing Machine, whose tape is the robot and environment state and rules are the policy.
- Oct 8, 2026 · Research paper · 1 sourcePLaW-VLA: Predictive Latent World Modeling for Vision-Language-Action PoliciesLearning to predict how the world evolves can provide vision-language-action (VLA) policies with predictive context for long-horizon control, but its effectiveness depends on what future representation is modeled and how it conditions action generation.
- Oct 8, 2026 · Research paper · 1 sourceResidual Modeling Closes the Regression and Generative Policy Gap in Robot LearningWe revisit this gap from the perspective of statistical modeling: how action-prediction residuals shape policy optimization.
- Oct 8, 2026 · Research paper · 1 sourceRESETTLE: Robotic Recovery through Disagreement-Triggered Retrieval and Efficient Corrective ControlTo address these challenges, we introduce RESETTLE(Robotic rEcovery through diSagrEement-Triggered reTrievaL and Efficient Corrective Control), a model-agnostic framework that provides computationally efficient recovery at the action-execution interface of frozen robot policies.
- Oct 8, 2026 · Research paper · 1 sourceContourVLA: A Closed-Loop Perception-Action Contour Policy for Generalized Referring Expression SegmentationWe introduce ContourVLA, a vision-language-action policy that recasts GRES as a closed-loop visuomotor process, in which editable contours serve as explicit policy states that condition multimodal perception and are updated by geometric action chunks.
- Oct 8, 2026 · Research paper · 1 sourceRecompose and Refine Latent Reasoning Flows for Vision-Language-Action ModelsLatent reasoning enables vision-language-action (VLA) models to transform multimodal observations into task-relevant internal states before generating continuous robot actions.
- Oct 8, 2026 · Research paper · 1 sourceManiUnit: A Manipulation Skill Dataset and Benchmark for Long-Horizon TasksLong-horizon mobile manipulation requires a robot to navigate multi-room environments and execute a sequence of manipulation skills under a single natural language instruction.
- Oct 8, 2026 · Research paper · 1 sourceHumanoid World Action Model With Joint State--Action GenerationWe propose HWAM, a Humanoid World Action Model with joint state--action generation, which makes the robot's post-execution proprioceptive state an explicit prediction target.
- Oct 8, 2026 · Research paper · 1 sourceREACT: Rolling Denoising and Dual Decoupling for Reactive Robot Control with VLA ModelsFlow-based vision-language-action (VLA) models generate action chunks for temporally coherent robot motion, but chunked control creates a fundamental closed-loop trade-off: long chunks provide smooth execution, whereas frequent replanning improves reactivity at the cost of action discontinuities.
- Oct 8, 2026 · Research paper · 1 sourceCAPABLE: Capability-Aware Policy Adaptation via Behavioral Latent EncodingWe introduce CAPABLE, a unified capability-aware adaptation framework for frozen VLAs that integrates self-supervised capability inference with residual reinforcement learning.
- Oct 8, 2026 · Research paper · 1 sourceTell Robot What Not to Do: A Negation Understanding PerspectiveTo this end, we propose NegaAlign, a parameter-efficient, plug-and-play framework that extends pretrained VLAs to follow negated instructions through image-language supervision alone.
- Oct 8, 2026 · Research paper · 1 sourcePathTime-VLA: Path-Time Decoupling for Factorized Post-Training of Vision-Language-Action PoliciesVision-Language-Action (VLA) policies typically predict actions at fixed time intervals, coupling the route a robot follows with its execution pace.
- Oct 8, 2026 · Research paper · 1 sourceWARP-VLA: Wrist-Camera Adaptation for View-Robust Policy Execution in Vision-Language-Action ModelsTo address this, we propose WARP-VLA, a camera-view robust VLA for diverse wrist camera configurations.
- Oct 8, 2026 · Research paper · 1 sourceRoboAware: Learning to Coordinate Embodied Skills from Counterfactual OutcomesWe present RoboAware, which builds on coding agents' skill orchestration by learning only a state-conditioned responsibility coordinator from counterfactual outcomes.
- Oct 8, 2026 · Research paper · 1 sourceExperience-Guided Initiation Search for Learned Skills in Skill CompositionWe propose EVIS, an Experience-Guided and Behavior-Validated Initiation Search framework for discovering reliable initiation configurations under limited target interaction.
- Oct 8, 2026 · Research paper · 1 sourceRewiring Semantics, Dynamics, and Control: A Simple yet Effective Action-Centric Tri-Stream TransformerVision-Language-Action (VLA) models have emerged as a prominent framework for complex robotic manipulation, building on the strong semantic understanding of pretrained Vision-Language Models (VLMs).
- Oct 8, 2026 · Research paper · 1 sourceSimVLA: Zero-Shot Sim-to-Real VLA Learning for Mobile ManipulationWe introduce SimVLA, an end-to-end framework that trains VLAs entirely on synthetic simulation data without teleoperation for mobile manipulation.
- Oct 8, 2026 · Research paper · 1 sourceVGGTWorld-VLA: Intent-Conditioned 3D World Evolution for Autonomous DrivingWe propose VGGTWorld-VLA, an intention-conditioned extension of VGGT-World for controllable 3D world evolution in autonomous driving.
- Oct 7, 2026 · Research paper · 1 sourceWhen Listening Becomes Easier: Scrubbing Visual Cues for Shortcut-Free VLAsShortcut learning is a prevalent issue in robot learning.
- Oct 7, 2026 · Research paper · 1 sourceNavGPT-3: Harnessing Context in a Hierarchical Navigation RuntimeWe present NavGPT-3, a harness that connects the two models, with an OS-like runtime built above it: reasoning, acting, and monitoring run as threads with their own context, tools, and permissions, while the runtime schedules them and decides which thread controls the robot's motion, so that the robot can react to sudden real-world events through interruption and thread switching.
- Oct 7, 2026 · Research paper · 1 sourceRephrase Before You Act: Characterizing and Mitigating Language Sensitivity in Vision-Language-Action ModelsVision-language-action models (VLAs) are strikingly sensitive to instruction phrasing and do not inherit the language robustness of the vision-language models they are built on.
- Oct 7, 2026 · Research paper · 2 sourcesQ-Learning with Scalar Adjoint MatchingAdjoint matching offers a principled way to update the flow model itself by propagating value information from the final action back to each flow step, but it requires a vector--Jacobian product through the policy at every step, a cost that grows with the number of flow steps and the policy size.
- Oct 7, 2026 · Research paper · 1 sourceExplicit Geometric Chain-of-Thought for Vision-Language-Action in Autonomous DrivingVision-language-action (VLA) models have emerged as a promising paradigm for autonomous driving.
- Oct 7, 2026 · Research paper · 1 sourceOpenViTac: Learning and Benchmarking Visuo-Tactile Policies in a Unified Sim-and-Real FrameworkTo address this gap, we introduce OpenViTac, a visuo-tactile manipulation benchmark for evaluating robot policies across simulation and the real world.
- Oct 7, 2026 · Research paper · 1 sourceDo Vision-Language-Action Models Understand Instructions? A Mechanistic Interpretability Study on Language GroundingVision-Language-Action models are designed to generalise across environments and task descriptions, raising the question of whether their action generation actually depends on the language instruction, or whether they largely rely on visual cues and superficial correlations.
- Oct 7, 2026 · Research paper · 1 sourceMany Ways to Succeed: Diversity-Driven RL Fine-Tuning for VLA GeneralizationReinforcement learning (RL) fine-tuning improves vision-language-action (VLA) policies through closed-loop experience, yet generalization beyond the fine-tuning distribution remains limited.
- Oct 7, 2026 · Research paper · 1 sourceJuno: Taming Predictive Latents for Vision-Language-Action ModelsJoint-embedding predictive architectures (JEPAs) predict masked or future observations in representation space, offering a natural source of predictive latents for vision-language-action (VLA) models.
- Oct 7, 2026 · Research paper · 1 sourceYUBI-STAG: Contact and Semantic-Rich Alignment for VLAs via Automated Video-Language GroundingWe introduce YUBI-STAG, a framework for Spatio-Temporal Annotation and Grounding that automatically enriches manipulation demonstrations with interaction-rich semantics to align pretrained VLAs with fine-grained manipulation language.