GroundSight at GroundLM 2026 Shared Tasks: GoldenViewVQA
We present CoVeR-VQA, a training-free multi-stage verification and correction framework for grounded multi-view VQA.
Key points
- GoldenViewVQA requires models to jointly answer driving-scene questions and identify the camera view containing the supporting visual evidence, making precise evidence localization as important as answer correctness.
- Starting from GPT-5.6 zero-shot predictions, CoVeR-VQA progressively applies view-specific verification with Gemini-3.6-Flash, prior-guided joint verification with Claude-Opus-5, and cross-split group-level verification that exploits semantically filtered question groups from shared multi-view scenes and validation-derived prior knowledge.
- On the official GoldenViewVQA test set, the four-stage CoVeR-VQA pipeline achieves 84.75% Joint Accuracy, improving the GPT-5.6 zero-shot baseline by 13.56 percentage points, while reaching 94.92% Answer Accuracy and 86.44% View Accuracy.
- Our analysis shows that supporting-view localization remains the primary source of residual errors, highlighting the importance of explicit evidence verification for reliable multi-view multimodal reasoning.
Sources (1)
- [1]GroundSight at GroundLM 2026 Shared Tasks: GoldenViewVQAarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 07:34 AM
We present CoVeR-VQA, a training-free multi-stage verification and correction framework for grounded multi-view VQA.
GoldenViewVQA requires models to jointly answer driving-scene questions and identify the camera view containing the supporting visual evidence, making precise evidence localization as important as answer correctness.
Extractive summary: sentences quoted from the sources.