AION
Research paperReinforcement Learning1 source · Oct 7, 2026

Many Ways to Succeed: Diversity-Driven RL Fine-Tuning for VLA Generalization

Reinforcement learning (RL) fine-tuning improves vision-language-action (VLA) policies through closed-loop experience, yet generalization beyond the fine-tuning distribution remains limited.

Key points

  • Our analysis reveals a selective reshaping of exploration: RL contracts behavior globally, yet diversifies successful trajectories, elicits success with fewer rollouts, and covers more of the latent task-valid solution space than supervised fine-tuning.
  • Inspired by this, we introduce DRIVE (Diversity-driven RL fIne-tuning for VLA gEneralization), which turns successful-behavior diversity into an explicit RL objective.
  • DRIVE groups rollouts under matched task conditions, compares their trajectories with temporal alignment, and derives a success-conditioned intrinsic reward from relative behavioral diversity.
  • Across LIBERO-Plus, ManiSkill3, and RoboTwin 2.0, DRIVE improves the average out-of-domain (OOD) performance over vanilla RL fine-tuning by 5.3 points on $π0$ and 2.0 points on $π{0.5}$.

Sources (1)

  • [1]Many Ways to Succeed: Diversity-Driven RL Fine-Tuning for VLA Generalization
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 12:26 PM
    Reinforcement learning (RL) fine-tuning improves vision-language-action (VLA) policies through closed-loop experience, yet generalization beyond the fine-tuning distribution remains limited.
    Our analysis reveals a selective reshaping of exploration: RL contracts behavior globally, yet diversifies successful trajectories, elicits success with fewer rollouts, and covers more of the latent task-valid solution space than supervised fine-tuning.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 7, 2026TempoBridge: Language-Guided Tempo Control for Vision-Language-Action Policies
  2. Oct 7, 2026Q-Learning with Scalar Adjoint Matching
  3. Oct 7, 2026Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
  4. Oct 6, 2026Co-Evolving Robot Orchestrators and Policies through Deployment
  5. Oct 6, 2026VLA-ACL: Action-Consistent Visual Token Pruning for Efficient Vision-Language-Action Models
  6. Jun 10, 2026DiffusionGemma: 4x faster text generation

Related