Visual Evidence Under Cross-Examination: Evaluating and Controlling Decision-Level Evidence Use in Vision-Language Models
Vision-language models increasingly reason through crops, regions, and tool-produced observations.
ProofPaper ↗
Key points
- We study candidate-bound visual contribution: valid evidence should help, invalidating its supporting relation should remove its additional effect, and valid rebinding should redirect that effect to the newly supported candidate.
- We introduce CROSS-Bench, a benchmark of 28,000 decision problems, with matched invalidation and rebinding tests on a dedicated evaluation subset.
- Our RIVET interface preserves evidence identity and uncertainty, composes a candidate-conditioned response, and separately controls its strength.
- Shared-evidence experiments show that task accuracy and evidence ownership can diverge.
Sources (1)
- [1]Visual Evidence Under Cross-Examination: Evaluating and Controlling Decision-Level Evidence Use in Vision-Language ModelsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 06:48 AM
Vision-language models increasingly reason through crops, regions, and tool-produced observations.
We study candidate-bound visual contribution: valid evidence should help, invalidating its supporting relation should remove its additional effect, and valid rebinding should redirect that effect to the newly supported candidate.
Extractive summary: sentences quoted from the sources.