You Changed Your Mind, The Model Didn't: Demystifying Intent in Multi-Turn Dialogue
To systematically study language model behavior under evolving user intent, we introduce Intent-Eval, a controlled benchmark spanning tool actions, code, databases, and mathematics.
ProofPaper ↗
Key points
- When a large language model handles a multi-turn task and a user proposes a change but ultimately rejects it, the model should continue as if nothing changed.
- We find a surprising failure: merely mentioning a rejected change can derail task execution, even when the user's final intent remains unchanged.
- Across diverse tasks, models are vulnerable to both rejected proposals and superseded requirements, consistent with mentioned-as-in-effect confusion: conversational content is treated as active requirements even after it has been rejected or replaced.
- Building on this insight, we propose Intent-OPSD, a decision-conditioned on-policy self-distillation framework with Teacher and Student initialized from the same model.
Sources (1)
- [1]You Changed Your Mind, The Model Didn't: Demystifying Intent in Multi-Turn DialogueHugging Face Daily Papers · Oct 5, 12:00 AM
To systematically study language model behavior under evolving user intent, we introduce Intent-Eval, a controlled benchmark spanning tool actions, code, databases, and mathematics.
When a large language model handles a multi-turn task and a user proposes a change but ultimately rejects it, the model should continue as if nothing changed.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 2, 2026FastOPD: On-Policy Distillation for Lightweight VLA Deployment
- Sep 29, 2026[AINews] AMD buys World Labs for $8.2B, as Atlas solves sparse reconstruction problem for robotics, design and more
- Sep 29, 2026Salt++: Context-Aligned Post-Training for Few-Step Streaming Multimodal Generation
- Sep 28, 2026Notes on NVIDIA Nemotron
- Aug 22, 2026sgl-project/sglang v0.5.18
- Aug 10, 2026huggingface/transformers v5.15.0: Release: v5.15.0