AION
Research paperLarge Language Models · Speech & Audio1 source · Oct 6, 2026

The Labeling Problem in Hallucination Detection Benchmarks: An Empirical Evaluation

In recent years, several methods for detecting when large language models (LLMs) hallucinate have been developed.

Key points

  • Then, some automated labeling strategy is used to label these answers as hallucinated or not by comparing them with the reference answers in the dataset.
  • We study this potential criterion mismatch using 900 human-labeled question-answer pairs spanning three commonly used QA datasets and three generator models, with labels targeting answer-level factual correctness.
  • We evaluate lexical similarity metrics, a reference-entailment NLI baseline, and seven LLM judges under controlled prompt variants as automated labelers.
  • Many strategies also exhibit strong directional error biases, and for most judge-generator pairs, replacing a faithfulness-oriented prompt with a factual-correctness prompt improves agreement with human annotations and reduces false-positive dominance, indicating that automated hallucination labels depend strongly on how the target criterion is specified.

Sources (1)

  • [1]The Labeling Problem in Hallucination Detection Benchmarks: An Empirical Evaluation
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 09:19 AM
    In recent years, several methods for detecting when large language models (LLMs) hallucinate have been developed.
    Then, some automated labeling strategy is used to label these answers as hallucinated or not by comparing them with the reference answers in the dataset.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 6, 2026Rethinking Faithfulness in LLMs: A Pairwise Context-Sensitive Perspective
  2. Sep 30, 2026Gemini 4 Argon: our next era of frontier intelligence

Related