VAMR: Multi-Question Agentic Reasoning for Efficient Long-Form Video Understanding
We introduce VAMR (Video Agent for Multi-Question Reasoning), which coordinates all questions about a video through one shared tool-use trajectory.
Key points
- Long-form video understanding often involves multiple questions about different aspects of the same recording.
- Yet existing video agents typically process each question through an isolated tool-use trajectory.
- Question-conditioned visual perception retrieves fine-grained clues for several questions in one call, while layered multi-question memory integrates reusable context into a shared video story and preserves separate evidence for individual questions.
- After supervised fine-tuning initializes this interaction protocol, we propose question-horizon policy optimization (\qhpo) to optimize shared trajectories in which questions progress and finish at different rounds.
Sources (1)
- [1]VAMR: Multi-Question Agentic Reasoning for Efficient Long-Form Video UnderstandingarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 03:26 AM
We introduce VAMR (Video Agent for Multi-Question Reasoning), which coordinates all questions about a video through one shared tool-use trajectory.
Long-form video understanding often involves multiple questions about different aspects of the same recording.
Extractive summary: sentences quoted from the sources.