ViSkill: Reinforcing VLM Agents with Evolving Visual-Native Skills
We propose ViSkill, a visual-native skill learning framework that encodes successful interactions as composite visual skill cards directly accessible to VLM agents.
Key points
- Skill-augmented agents improve sample efficiency by distilling successful trajectories into reusable strategies.
- Yet most existing approaches remain text-centric, linearizing spatial layouts and action-state correspondences into language that loses critical geometric structure.
- Recent efforts have begun incorporating visual evidence, but construct and update skills separately from policy optimization, leaving their mutual improvement underexplored.
- An optional cold-start mechanism further accelerates early-stage learning.
Sources (2)
- [1]ViSkill: Reinforcing VLM Agents with Evolving Visual-Native SkillsHugging Face Daily Papers · Oct 8, 12:00 AM
We propose ViSkill, a visual-native skill learning framework that encodes successful interactions as composite visual skill cards directly accessible to VLM agents.
Skill-augmented agents improve sample efficiency by distilling successful trajectories into reusable strategies.
- [2]ViSkill: Reinforcing VLM Agents with Evolving Visual-Native SkillsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 05:44 PM · same content
Extractive summary: sentences quoted from the sources.