From Suppression to Repair: Mitigating Object Hallucination in Large Vision-Language Models via Localized Distribution Alignment
Object hallucination remains a major obstacle for large vision-language models (LVLMs) to generate reliable content.
Key points
- An intuitive mitigation strategy is to suppress hallucination-related components in hidden representations.
- In this paper, we propose ResOT, a training-free method that repairs representations at inference time through localized distribution alignment.
- Within this subspace, ResOT uses Gaussian optimal transport (OT) to align the hallucinated distribution with the faithful one.
- Experiments on three representative LVLMs show that ResOT substantially reduces object hallucination while improving image caption quality and multimodal performance across multiple benchmarks.
Sources (1)
- [1]From Suppression to Repair: Mitigating Object Hallucination in Large Vision-Language Models via Localized Distribution AlignmentarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 12:25 PM
Object hallucination remains a major obstacle for large vision-language models (LVLMs) to generate reliable content.
An intuitive mitigation strategy is to suppress hallucination-related components in hidden representations.
Extractive summary: sentences quoted from the sources.