AION
Research paperLarge Language Models · Retrieval, RAG & Search1 source · Oct 8, 2026

RAG-Stress: Probing the Limits of Evidence Reliance in Retrieval-Augmented Generation

We introduce RAG-Stress, a controlled diagnostic protocol for examining the limits of evidence reliance in retrieval-augmented generation.

Key points

  • Following retrieved evidence does not guarantee factual correctness: misleading evidence can induce a model to replace an answer it previously gave correctly.
  • The protocol holds the question and reference answer fixed, edits one assertion to support a designated incorrect answer, and crosses two source priority policies with three positions of the answer span within the evidence text.
  • We measure misleading rate (MR) on each model's subset of questions answered correctly without retrieval, alongside clean accuracy on the full evaluation set.
  • We evaluate fifteen systems spanning API models, open models, and search agents trained with reinforcement learning on TriviaQA-RC, HotpotQA, and SearchQA, with additional English and Chinese MedQA evaluations.

Sources (1)

  • [1]RAG-Stress: Probing the Limits of Evidence Reliance in Retrieval-Augmented Generation
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 03:40 AM
    We introduce RAG-Stress, a controlled diagnostic protocol for examining the limits of evidence reliance in retrieval-augmented generation.
    Following retrieved evidence does not guarantee factual correctness: misleading evidence can induce a model to replace an answer it previously gave correctly.

Extractive summary: sentences quoted from the sources.