AION
Research paperLarge Language Models · Multimodal Models · Robotics & Embodied AI1 source · Oct 8, 2026

VAMR: Multi-Question Agentic Reasoning for Efficient Long-Form Video Understanding

We introduce VAMR (Video Agent for Multi-Question Reasoning), which coordinates all questions about a video through one shared tool-use trajectory.

Key points

  • Long-form video understanding often involves multiple questions about different aspects of the same recording.
  • Yet existing video agents typically process each question through an isolated tool-use trajectory.
  • Question-conditioned visual perception retrieves fine-grained clues for several questions in one call, while layered multi-question memory integrates reusable context into a shared video story and preserves separate evidence for individual questions.
  • After supervised fine-tuning initializes this interaction protocol, we propose question-horizon policy optimization (\qhpo) to optimize shared trajectories in which questions progress and finish at different rounds.

Sources (1)

Extractive summary: sentences quoted from the sources.