Knowing When Not to Answer: Cross-Domain and Multi-Turn Generalization of Latent Underspecification Signals
Unanswerability is linearly decodable from hidden states, but it is unclear which of its forms share a representation and whether the signal is useful in dialogue.
ProofPaper ↗
Key points
- Large language models routinely answer questions that cannot be answered from the information given, and in dialogue they answer before enough has been said.
- We contribute a turn-labeled multi-turn benchmark (423 conversations, 1,661 labeled turn-states) and an evaluation harness with a simulated user who answers clarifying questions, and use them with six datasets and six open-weight LLMs to test how far probes for unanswerability carry.
- Probes transfer robustly between datasets that share a ground of unanswerability: missing information in math (AUROC 0.77-0.97) and in a passage (SQuAD 2.0<->MuSiQue, 0.77-0.90).
- Single-turn probes fail zero-shot to detect when a conversation becomes answerable; in-structure probes recover it, but no better than a bag-of-words classifier.
Sources (1)
- [1]Knowing When Not to Answer: Cross-Domain and Multi-Turn Generalization of Latent Underspecification SignalsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 02:18 PM
Unanswerability is linearly decodable from hidden states, but it is unclear which of its forms share a representation and whether the signal is useful in dialogue.
Large language models routinely answer questions that cannot be answered from the information given, and in dialogue they answer before enough has been said.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 6, 2026Frozen Models, Evolving Expertise: Model-Agnostic Learning from Deployment Experience for Multimodal Medical AI
- Sep 30, 2026Expanding AI Storage Access with NVIDIA cuObject and the NVIDIA SCADA Server SDK
- Sep 28, 2026openai/openai-python v3.20.0
- Sep 24, 2026openai/openai-python v3.19.2
- Jul 15, 2026huggingface/transformers v5.14.0: Release v5.14.0
- Jun 10, 2026DiffusionGemma: 4x faster text generation