PHRBench: A Behavioral Evaluation of Post-Hallucination Reasoning in LLMs
Existing studies of post-hallucination reasoning (PHR) mainly characterize changes in final outcomes and aggregate reasoning dynamics, leaving how models resolve hallucinated premises at the response level insufficiently understood.
Key points
- Hallucinated information can propagate through multi-stage LLM systems and become part of the context for subsequent reasoning.
- In this work, we introduce PHRBench, a controlled benchmark for behaviorally structured PHR across four domains and 18 large language models.
- Across 4820 controlled instances, we find that successful recovery remains relatively rare and is associated with more frequent belief updates along the reasoning trajectory.
- These findings provide a behavioral view of post-hallucination reasoning, characterizing how LLMs resolve erroneous context and when successful recovery is likely to occur.
Sources (1)
- [1]PHRBench: A Behavioral Evaluation of Post-Hallucination Reasoning in LLMsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 05:25 PM
Existing studies of post-hallucination reasoning (PHR) mainly characterize changes in final outcomes and aggregate reasoning dynamics, leaving how models resolve hallucinated premises at the response level insufficiently understood.
Hallucinated information can propagate through multi-stage LLM systems and become part of the context for subsequent reasoning.
Extractive summary: sentences quoted from the sources.