ResearchResearch paperLarge Language Models · Multimodal Models1 source · Oct 6, 2026

RACER: Reflective Agent Coupling Query Interpretation and Tool-Based Retrieval for Frame Selection in Long Video Understanding

To address them, we propose RACER, a training-free reflective agentic framework that decomposes long-video frame selection into query interpretation driven by a lightweight Vid-LLM and evidence localization supported by an embedding model serving as a retrieval tool.

Key points

  • Video large language models (Vid-LLMs) excel at diverse video-language tasks by reasoning over selected frames.
  • However, frame selection for long videos remains challenging, as it requires retrieving relevant frames distributed across segments from a large candidate pool given complex queries.
  • This paper investigates dominant approaches to long-video frame selection from a task-decomposition perspective, identifying two key challenges: the Query Comprehension Gap in similarity-based methods and the Interpretation--Selection Gap in judgment-based methods.
  • Experiments across multiple benchmarks show that RACER consistently improves long video understanding.

Sources (1)

Extractive summary: sentences quoted from the sources.

Related