ResearchResearch paperLarge Language Models · Robotics & Embodied AI · Efficiency & Inference1 source · Oct 7, 2026

VideoEvolve: Co-Evolving Memory and Retrieval for Long Video Understanding

To address this issue, we propose VideoEvolve, a novel self-evolving framework that jointly evolves memory and retrieval for long video understanding.

Key points

  • Long video understanding increasingly relies on external memory to organize massive visual streams into compact representations.
  • Specifically, starting from a coarse low-frame-rate overview, VideoEvolve couples a Memory Evolver for selective memory augmentation with a Retrieval Evolver for adaptive retrieval over the evolving memory.
  • Furthermore, VideoEvolve introduces Capability-Aware Evolution Feedback (CEF) to alleviate downstream feedback from over-specializing memory to a fixed set of training questions, shifting training toward underdeveloped yet learnable video capabilities.
  • By integrating Agentic RL with BEF and CEF, VideoEvolve transforms downstream reasoning experience into transferable capability updates, providing a concrete path from static long-video systems toward experience-driven, self-improving multimodal intelligence.

Sources (1)

  • [1]VideoEvolve: Co-Evolving Memory and Retrieval for Long Video Understanding
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 02:51 PM
    To address this issue, we propose VideoEvolve, a novel self-evolving framework that jointly evolves memory and retrieval for long video understanding.
    Long video understanding increasingly relies on external memory to organize massive visual streams into compact representations.

Extractive summary: sentences quoted from the sources.

Related