Beyond Visual Enhancement: Adaptive Multi-Context Steering to Mitigate LVLM Hallucinations
Hallucination remains a significant challenge in Large Vision-Language Models (LVLMs).
Key points
- This raises a fundamental question: Can LVLMs dynamically regulate the contributions of different context sources to suppress hallucinations?
- In this work, we investigate and quantify how LVLMs coordinate multiple context sources during decoding and examine how this intrinsic behavior can guide hallucination mitigation.
- We find that LVLMs exhibit an intrinsic vision-attending tendency that can guide adaptive visual steering, while textual contexts can also contribute to hallucination mitigation.
- Motivated by these findings, we propose AIMS (Adaptive Information Multi-source Steering), a lightweight training-free framework that adaptively coordinates visual, prefilled textual, and generated contexts during decoding.
Sources (1)
- [1]Beyond Visual Enhancement: Adaptive Multi-Context Steering to Mitigate LVLM HallucinationsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 01:08 PM
Hallucination remains a significant challenge in Large Vision-Language Models (LVLMs).
This raises a fundamental question: Can LVLMs dynamically regulate the contributions of different context sources to suppress hallucinations?
Extractive summary: sentences quoted from the sources.