AION
Research paperReinforcement Learning · Robotics & Embodied AI2 sources · Oct 8, 2026

ViSkill: Reinforcing VLM Agents with Evolving Visual-Native Skills

We propose ViSkill, a visual-native skill learning framework that encodes successful interactions as composite visual skill cards directly accessible to VLM agents.

Key points

  • Skill-augmented agents improve sample efficiency by distilling successful trajectories into reusable strategies.
  • Yet most existing approaches remain text-centric, linearizing spatial layouts and action-state correspondences into language that loses critical geometric structure.
  • Recent efforts have begun incorporating visual evidence, but construct and update skills separately from policy optimization, leaving their mutual improvement underexplored.
  • An optional cold-start mechanism further accelerates early-stage learning.

Sources (2)

Extractive summary: sentences quoted from the sources.