OneSearch-VL: Unified Multimodal Deep Research Agent for Image and Video
We introduce OneSearch-VL, a unified agent centered on the Visually Grounded Evidence Graph (VGEG), which encodes these dependencies as a shared task-level reference for data construction, process supervision, and operation-level evaluation.
Key points
- Single-image, multi-image, and video deep research require different visual operations but share a workflow of visual grounding, external retrieval, and fact composition.
- Our VGEG-based data engine constructs and verifies multi-image and video questions and filters expert trajectories.
- We further derive the Evidence-aware Visual-Grounded Rubric reward (EVGR) from VGEG annotations to supervise evidence traceability and visual grounding during RL.
- For fine-grained evaluation, we construct OneSearch-MI-Bench and OneSearch-Video-Bench, organizing questions by the research operations encoded in their VGEGs.
Sources (2)
- [1]OneSearch-VL: Unified Multimodal Deep Research Agent for Image and VideoHugging Face Daily Papers · Oct 8, 12:00 AM
We introduce OneSearch-VL, a unified agent centered on the Visually Grounded Evidence Graph (VGEG), which encodes these dependencies as a shared task-level reference for data construction, process supervision, and operation-level evaluation.
Single-image, multi-image, and video deep research require different visual operations but share a workflow of visual grounding, external retrieval, and fact composition.
- [2]OneSearch-VL: Unified Multimodal Deep Research Agent for Image and VideoarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 05:52 PM · same content
Extractive summary: sentences quoted from the sources.