RACER: Reflective Agent Coupling Query Interpretation and Tool-Based Retrieval for Frame Selection in Long Video Understanding
To address them, we propose RACER, a training-free reflective agentic framework that decomposes long-video frame selection into query interpretation driven by a lightweight Vid-LLM and evidence localization supported by an embedding model serving as a retrieval tool.
ProofPaper ↗
Key points
- Video large language models (Vid-LLMs) excel at diverse video-language tasks by reasoning over selected frames.
- However, frame selection for long videos remains challenging, as it requires retrieving relevant frames distributed across segments from a large candidate pool given complex queries.
- This paper investigates dominant approaches to long-video frame selection from a task-decomposition perspective, identifying two key challenges: the Query Comprehension Gap in similarity-based methods and the Interpretation--Selection Gap in judgment-based methods.
- Experiments across multiple benchmarks show that RACER consistently improves long video understanding.
Sources (1)
- [1]RACER: Reflective Agent Coupling Query Interpretation and Tool-Based Retrieval for Frame Selection in Long Video UnderstandingarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 06:21 PM
To address them, we propose RACER, a training-free reflective agentic framework that decomposes long-video frame selection into query interpretation driven by a lightweight Vid-LLM and evidence localization supported by an embedding model serving as a retrieval tool.
Video large language models (Vid-LLMs) excel at diverse video-language tasks by reasoning over selected frames.
Extractive summary: sentences quoted from the sources.