ResearchResearch paperRobotics & Embodied AI · Reinforcement Learning · Safety & Alignment1 source · Oct 8, 2026

CAPABLE: Capability-Aware Policy Adaptation via Behavioral Latent Encoding

We introduce CAPABLE, a unified capability-aware adaptation framework for frozen VLAs that integrates self-supervised capability inference with residual reinforcement learning.

Key points

  • Vision-language-action (VLA) policies assume the embodiment on which they were trained and can fail when a joint fault changes how commanded actions are physically executed.
  • Existing fault-recovery methods often require task-specific retraining, fault labels, explicit diagnosis, or privileged embodiment information.
  • CAPABLE infers capability, how much of the commanded motion each joint actually realizes and how that motion contributes to end-effector behavior, online from command-response history and kinematics using a temporal encoder shared across joints, Jacobian grounding, cross-joint attention, and self-supervised physical prediction.
  • The resulting representation conditions a residual policy that adds bounded corrections to the VLA arm action without fault labels or faulty-joint identifiers.

Sources (1)

  • [1]CAPABLE: Capability-Aware Policy Adaptation via Behavioral Latent Encoding
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 01:46 PM
    We introduce CAPABLE, a unified capability-aware adaptation framework for frozen VLAs that integrates self-supervised capability inference with residual reinforcement learning.
    Vision-language-action (VLA) policies assume the embodiment on which they were trained and can fail when a joint fault changes how commanded actions are physically executed.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 8, 2026Rewiring Semantics, Dynamics, and Control: A Simple yet Effective Action-Centric Tri-Stream Transformer
  2. Oct 8, 2026Embodied Turing Machines: Stateful Code for Robot Recursive Self-Improvement
  3. Oct 7, 2026Explicit Geometric Chain-of-Thought for Vision-Language-Action in Autonomous Driving
  4. Oct 7, 2026Do Vision-Language-Action Models Understand Instructions? A Mechanistic Interpretability Study on Language Grounding
  5. Oct 7, 2026Many Ways to Succeed: Diversity-Driven RL Fine-Tuning for VLA Generalization
  6. Oct 7, 2026Q-Learning with Scalar Adjoint Matching

Related