Arctic Questions, Missing Answers: A Dataset and Benchmark for LLM Abstention in Arctic Science
We introduce ArcticQA, a dataset of 194 questions derived from primary Arctic research, with automated checks of answer support and distractor contradiction against source evidence.
Key points
- Large language models (LLMs) should abstain from scientific multiple-choice questions when no option is valid, but frequent abstention alone does not demonstrate sensitivity to answer availability.
- We further develop ArcticAbstain, a paired benchmark comparing answer-present and answer-absent conditions, with the correct answer replaced by a distractor in the latter and an explicit abstention option in both.
- We evaluate eight models from the Gemini, Claude, and ChatGPT families at high reasoning effort, with three trials per condition, yielding 9,312 recorded responses.
- Answer-present abstention rates range from 0.0% to 63.0%, whereas replacing the correct answer increases abstention by 5.05 percentage points on average.
Sources (1)
- [1]Arctic Questions, Missing Answers: A Dataset and Benchmark for LLM Abstention in Arctic SciencearXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 04:58 AM
We introduce ArcticQA, a dataset of 194 questions derived from primary Arctic research, with automated checks of answer support and distractor contradiction against source evidence.
Large language models (LLMs) should abstain from scientific multiple-choice questions when no option is valid, but frequent abstention alone does not demonstrate sensitivity to answer availability.
Extractive summary: sentences quoted from the sources.