VICO: Visual Environments Co-Evolving for Vision-Language Model Reasoning
Reinforcement learning with verifiable rewards (RLVR) has become a standard recipe for post-training vision-language models (VLMs),
Key points
- As the actor improves, fixed tasks drift out of its learning frontier: many
- We propose VICO, a co-evolutionary framework in which an actor
- By shifting from human-labeled supervision to image-editing co-evolution,
- VICO offers a scalable path beyond static-corpus RLVR for visual reasoning.
Sources (1)
- [1]VICO: Visual Environments Co-Evolving for Vision-Language Model ReasoningarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 06:39 PM
Reinforcement learning with verifiable rewards (RLVR) has become a standard recipe for post-training vision-language models (VLMs),
As the actor improves, fixed tasks drift out of its learning frontier: many
Extractive summary: sentences quoted from the sources.