ResearchResearch paperLarge Language Models · Multimodal Models · Interpretability1 source · Oct 7, 2026

Visual Evidence Under Cross-Examination: Evaluating and Controlling Decision-Level Evidence Use in Vision-Language Models

Vision-language models increasingly reason through crops, regions, and tool-produced observations.

Key points

  • We study candidate-bound visual contribution: valid evidence should help, invalidating its supporting relation should remove its additional effect, and valid rebinding should redirect that effect to the newly supported candidate.
  • We introduce CROSS-Bench, a benchmark of 28,000 decision problems, with matched invalidation and rebinding tests on a dedicated evaluation subset.
  • Our RIVET interface preserves evidence identity and uncertainty, composes a candidate-conditioned response, and separately controls its strength.
  • Shared-evidence experiments show that task accuracy and evidence ownership can diverge.

Sources (1)

Extractive summary: sentences quoted from the sources.

Related