UXBench Pro: Benchmarking Personalized User Experience in Multi-Turn Dialogue Interactions
In this paper, we present UXBench Pro, comprising 1{,}000 test instances derived from real user interactions across 12 task scenarios and 82 domains.
ProofPaper ↗
Key points
- Evaluating user experience (UX) with automated computational methods has gained increasing attention, supported by empirical evidence from UXBench.
- To provide richer evaluation insights, we introduce a dual-perspective paradigm that combines a personalized User Reward Model (URM) for third-person judgment with Sim4Eval, a user simulator that enables multi-turn interactions and provides first-person evaluation across four cognitive state dimensions.
- To assess the reliability of these based evaluators, we further introduce two meta-benchmarks, URMBench and USimBench, that evaluate how faithfully they reproduce real human preferences and behaviors.
- Extensive experiments reveal seven key findings that highlight the importance of user modeling and multi-perspective evaluation, offering a fresh perspective on user-centric benchmarking and motivating personalized model optimization.
Sources (1)
- [1]UXBench Pro: Benchmarking Personalized User Experience in Multi-Turn Dialogue InteractionsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 10:15 AM
In this paper, we present UXBench Pro, comprising 1{,}000 test instances derived from real user interactions across 12 task scenarios and 82 domains.
Evaluating user experience (UX) with automated computational methods has gained increasing attention, supported by empirical evidence from UXBench.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 8, 2026DreamTrue: Action-Faithful Robot World Model with Counterfactual Post-Training
- Oct 7, 2026I would rather quit NLP than read another paper like this: The rise of antithesis in NLP papers
- Oct 7, 2026BoT-GRPO: Efficient Process-Reward RL for Reasoning via Bag-of-Token Aggregation
- Oct 7, 2026Judging in Latent Space: Efficient Generative Reward Modeling via Semantics-Preserving Compression
- Oct 6, 2026Verify Less, Evolve More: Training Idea-Level Critics for Verification-Efficient ML Evolving Agents
- Oct 6, 2026Reinforcement Learning for Hierarchical Reasoning Rewards: Minimax-Optimal Rates with Transformers