ResearchResearch paperMultimodal Models · Large Language Models · Interpretability1 source · Oct 8, 2026

SAGE: Sink-Aware Guided Emphasis for Visual Grounding in Vision-Language Decoders

Building on this insight, we propose SAGE (Sink-Aware Guided Emphasis), a lightweight intervention that steers decoder attention away from PIS and toward query-dependent regions of interest (ROIs) using token-aligned ROI masks derived from standard vision backbones such as CLIP, ViT, and DINOv3.

Key points

  • Recent large vision-language models (VLMs) pair a visual encoder with a large language model (LLM) and perform well on diverse image-text tasks, yet their reliability is often limited by decoder attention pathologies that suppress visual evidence and exacerbate hallucinations.
  • In this paper, we revisit visual attention sinks and uncover a structured, layer-dependent behavior: across prompts, early and late decoder layers exhibit prompt-invariant attention collapse onto the same few image regions, which we term PIS (Prompt-Invariant Sinks), whereas mid layers become prompt-conditioned and drive vision-language alignment.
  • This split suggests that treating sinks as a uniform effect is incomplete.
  • Evaluated on diverse vision-encoder + decoder-only LLM VLM families, SAGE improves visual grounding, reduces hallucinations, and yields consistent gains across public downstream vision-language benchmarks, including fine-grained visual discrimination settings where localized evidence is crucial, when instantiated with backbone-derived ROI masks.

Sources (1)

  • [1]SAGE: Sink-Aware Guided Emphasis for Visual Grounding in Vision-Language Decoders
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 08:18 AM
    Building on this insight, we propose SAGE (Sink-Aware Guided Emphasis), a lightweight intervention that steers decoder attention away from PIS and toward query-dependent regions of interest (ROIs) using token-aligned ROI masks derived from standard vision backbones such as CLIP, ViT, and DINOv3.
    Recent large vision-language models (VLMs) pair a visual encoder with a large language model (LLM) and perform well on diverse image-text tasks, yet their reliability is often limited by decoder attention pathologies that suppress visual evidence and exacerbate hallucinations.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 8, 2026Dissecting Representation Structure in Vision Transformers: A Rigorous Architectural Study
  2. Oct 8, 2026MCL: Meta Convolution Layer
  3. Oct 7, 2026On the Necessity of Attention-FFN Split in Vision Transformers
  4. Oct 7, 2026Agentic AI-Assisted Modeling for Production Scheduling: Assessment in Constraint Programming
  5. Oct 7, 2026ORCA: Hunting Compositional Failures in Text-to-Image Diffusion
  6. Oct 6, 2026huggingface/transformers v5.19.0: Release v5.19.0

Related