ResearchResearch paperLarge Language Models1 source · Oct 6, 2026

Knowing When Not to Answer: Cross-Domain and Multi-Turn Generalization of Latent Underspecification Signals

Unanswerability is linearly decodable from hidden states, but it is unclear which of its forms share a representation and whether the signal is useful in dialogue.

Key points

  • Large language models routinely answer questions that cannot be answered from the information given, and in dialogue they answer before enough has been said.
  • We contribute a turn-labeled multi-turn benchmark (423 conversations, 1,661 labeled turn-states) and an evaluation harness with a simulated user who answers clarifying questions, and use them with six datasets and six open-weight LLMs to test how far probes for unanswerability carry.
  • Probes transfer robustly between datasets that share a ground of unanswerability: missing information in math (AUROC 0.77-0.97) and in a passage (SQuAD 2.0<->MuSiQue, 0.77-0.90).
  • Single-turn probes fail zero-shot to detect when a conversation becomes answerable; in-structure probes recover it, but no better than a bag-of-words classifier.

Sources (1)

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 6, 2026Frozen Models, Evolving Expertise: Model-Agnostic Learning from Deployment Experience for Multimodal Medical AI
  2. Sep 30, 2026Expanding AI Storage Access with NVIDIA cuObject and the NVIDIA SCADA Server SDK
  3. Sep 28, 2026openai/openai-python v3.20.0
  4. Sep 24, 2026openai/openai-python v3.19.2
  5. Jul 15, 2026huggingface/transformers v5.14.0: Release v5.14.0
  6. Jun 10, 2026DiffusionGemma: 4x faster text generation

Related