AION
Research paperLarge Language Models1 source · Oct 7, 2026

Arctic Questions, Missing Answers: A Dataset and Benchmark for LLM Abstention in Arctic Science

We introduce ArcticQA, a dataset of 194 questions derived from primary Arctic research, with automated checks of answer support and distractor contradiction against source evidence.

Key points

  • Large language models (LLMs) should abstain from scientific multiple-choice questions when no option is valid, but frequent abstention alone does not demonstrate sensitivity to answer availability.
  • We further develop ArcticAbstain, a paired benchmark comparing answer-present and answer-absent conditions, with the correct answer replaced by a distractor in the latter and an explicit abstention option in both.
  • We evaluate eight models from the Gemini, Claude, and ChatGPT families at high reasoning effort, with three trials per condition, yielding 9,312 recorded responses.
  • Answer-present abstention rates range from 0.0% to 63.0%, whereas replacing the correct answer increases abstention by 5.05 percentage points on average.

Sources (1)

  • [1]Arctic Questions, Missing Answers: A Dataset and Benchmark for LLM Abstention in Arctic Science
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 04:58 AM
    We introduce ArcticQA, a dataset of 194 questions derived from primary Arctic research, with automated checks of answer support and distractor contradiction against source evidence.
    Large language models (LLMs) should abstain from scientific multiple-choice questions when no option is valid, but frequent abstention alone does not demonstrate sensitivity to answer availability.

Extractive summary: sentences quoted from the sources.