AION
Research paperMultimodal Models · Computer Vision1 source · Oct 8, 2026

GroundSight at GroundLM 2026 Shared Tasks: GoldenViewVQA

We present CoVeR-VQA, a training-free multi-stage verification and correction framework for grounded multi-view VQA.

Key points

  • GoldenViewVQA requires models to jointly answer driving-scene questions and identify the camera view containing the supporting visual evidence, making precise evidence localization as important as answer correctness.
  • Starting from GPT-5.6 zero-shot predictions, CoVeR-VQA progressively applies view-specific verification with Gemini-3.6-Flash, prior-guided joint verification with Claude-Opus-5, and cross-split group-level verification that exploits semantically filtered question groups from shared multi-view scenes and validation-derived prior knowledge.
  • On the official GoldenViewVQA test set, the four-stage CoVeR-VQA pipeline achieves 84.75% Joint Accuracy, improving the GPT-5.6 zero-shot baseline by 13.56 percentage points, while reaching 94.92% Answer Accuracy and 86.44% View Accuracy.
  • Our analysis shows that supporting-view localization remains the primary source of residual errors, highlighting the importance of explicit evidence verification for reliable multi-view multimodal reasoning.

Sources (1)

  • [1]GroundSight at GroundLM 2026 Shared Tasks: GoldenViewVQA
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 07:34 AM
    We present CoVeR-VQA, a training-free multi-stage verification and correction framework for grounded multi-view VQA.
    GoldenViewVQA requires models to jointly answer driving-scene questions and identify the camera view containing the supporting visual evidence, making precise evidence localization as important as answer correctness.

Extractive summary: sentences quoted from the sources.