ResearchResearch paperInterpretability · Large Language Models · Computer Vision1 source · Oct 7, 2026

InstanceBench: Diagnosing Referential Reasoning and Target Identity in Referring Expression Segmentation

Referring Expression Segmentation (RES) links natural-language descriptions to pixel-level object masks.

Key points

  • Yet standard evaluation provides limited insight into instance-level referential reasoning: it does not systematically distinguish referential logics, test target preservation across valid grounding paths, or separate target-selection from mask-generation errors.
  • We introduce InstanceBench, an instance-centered diagnostic benchmark comprising 6,194 images, 9,264 target instances, and 25,077 human-verified expressions.
  • Each target-centric expression set (TCES) fixes the image and target mask while pairing a minimal expression with a same-target variant that uses another valid cue or grounding path.
  • Identity-aware metrics measure target retention and set-level success while separating selection from mask-generation errors.

Sources (1)

  • [1]InstanceBench: Diagnosing Referential Reasoning and Target Identity in Referring Expression Segmentation
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 05:34 AM
    Referring Expression Segmentation (RES) links natural-language descriptions to pixel-level object masks.
    Yet standard evaluation provides limited insight into instance-level referential reasoning: it does not systematically distinguish referential logics, test target preservation across valid grounding paths, or separate target-selection from mask-generation errors.

Extractive summary: sentences quoted from the sources.

Related