VideoEvolve: Co-Evolving Memory and Retrieval for Long Video Understanding
To address this issue, we propose VideoEvolve, a novel self-evolving framework that jointly evolves memory and retrieval for long video understanding.
ProofPaper ↗
Key points
- Long video understanding increasingly relies on external memory to organize massive visual streams into compact representations.
- Specifically, starting from a coarse low-frame-rate overview, VideoEvolve couples a Memory Evolver for selective memory augmentation with a Retrieval Evolver for adaptive retrieval over the evolving memory.
- Furthermore, VideoEvolve introduces Capability-Aware Evolution Feedback (CEF) to alleviate downstream feedback from over-specializing memory to a fixed set of training questions, shifting training toward underdeveloped yet learnable video capabilities.
- By integrating Agentic RL with BEF and CEF, VideoEvolve transforms downstream reasoning experience into transferable capability updates, providing a concrete path from static long-video systems toward experience-driven, self-improving multimodal intelligence.
Sources (1)
- [1]VideoEvolve: Co-Evolving Memory and Retrieval for Long Video UnderstandingarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 02:51 PM
To address this issue, we propose VideoEvolve, a novel self-evolving framework that jointly evolves memory and retrieval for long video understanding.
Long video understanding increasingly relies on external memory to organize massive visual streams into compact representations.
Extractive summary: sentences quoted from the sources.