ResearchResearch paperLarge Language Models1 source · Oct 6, 2026

Penalty-Framed No-Valid-Option MCQA: Analyzing LLM Abstention under Invalid Choices

Multiple-choice question answering (MCQA) is commonly used to evaluate large language models under the assumption that one of the provided options is correct, typically using answer-selection accuracy.

Key points

  • Using the mathematics subset of MMLU-Pro, we remove the labeled correct option, allow models to either choose a remaining option or output ABSTAIN, and penalize invalid forced-choice responses.
  • We further introduce correct-conditioned analysis, evaluating abstention only on instances that the model originally answered correctly.
  • Experiments show that high MCQA accuracy does not fully guarantee abstention reliability: even under explicit no-valid-option-aware instructions and penalty-based scoring, models still produce invalid forced-choice responses for a subset of originally correct instances.
  • These results show that penalty-framed no-valid-option MCQA reveals an aspect of model reliability not captured by standard answer-selection accuracy.

Sources (1)

  • [1]Penalty-Framed No-Valid-Option MCQA: Analyzing LLM Abstention under Invalid Choices
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 11:04 AM
    Multiple-choice question answering (MCQA) is commonly used to evaluate large language models under the assumption that one of the provided options is correct, typically using answer-selection accuracy.
    Using the mathematics subset of MMLU-Pro, we remove the labeled correct option, allow models to either choose a remaining option or output ABSTAIN, and penalize invalid forced-choice responses.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 6, 2026Qwen3.8 Flash Next Reasoning Modes: Off vs Low vs Medium vs Xhigh

Related