AION
Research paperRetrieval, RAG & Search · Large Language Models1 source · Oct 7, 2026

Relevance Is Not Sufficiency: What Actually Closes the Evidence Gap in Long-Term Memory QA

Drawing on the distinction between relevance and sufficiency in legal evidence scholarship, we recast memory retrieval as constructing a sufficient memory set.

Key points

  • LLM agents that interact with a user across many sessions accumulate histories that exceed their context window, so they store past interactions in an external memory and answer each question from a small set of retrieved records.
  • Existing memory systems rank records by lexical or embedding relevance, yet the top-ranked memories can each be relevant while jointly omitting a complementary fact that the answer requires, especially for multi-session and temporal questions.
  • Specifically, we first apply Formal Concept Analysis for Memory Selection (FCA-MS) to decompose the question into information requirements and select a compact candidate subset that jointly covers them.
  • Experiments on LoCoMo and LongMemEval-S show that BFR outperforms same-store adaptations of recent agent-memory systems in both answer quality and evidence coverage.

Sources (1)

  • [1]Relevance Is Not Sufficiency: What Actually Closes the Evidence Gap in Long-Term Memory QA
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 03:09 AM
    Drawing on the distinction between relevance and sufficiency in legal evidence scholarship, we recast memory retrieval as constructing a sufficient memory set.
    LLM agents that interact with a user across many sessions accumulate histories that exceed their context window, so they store past interactions in an external memory and answer each question from a small set of retrieved records.

Extractive summary: sentences quoted from the sources.