MIMESIS: Learning User Simulators as Training Environments for Interactive Agents
We introduce MIMESIS, a purpose-built user simulator trained on human conversations with explicit reasoning supervision and 13 realistic behavioral patterns derived from real user interactions.
Key points
- Training and evaluating interactive language agents typically requires rich user interactions, yet collecting human feedback is expensive and difficult to scale.
- Simulated users offer a scalable alternative, but they must both resemble real user behavior and provide useful learning experiences for agents.
- Across eight environments, training with MIMESIS yields better agent performance than training with GPT-5.5 under all nine unseen user simulators, demonstrating stronger generalization to new user simulators.
- Moreover, we propose Coached On-Policy Self-Distillation (CSD), which leverages simulator-generated private reasoning traces and subsequent utterances as feedback on how well the agent addresses user needs.
Sources (2)
- [1]MIMESIS: Learning User Simulators as Training Environments for Interactive AgentsHugging Face Daily Papers · Oct 7, 12:00 AM
We introduce MIMESIS, a purpose-built user simulator trained on human conversations with explicit reasoning supervision and 13 realistic behavioral patterns derived from real user interactions.
Training and evaluating interactive language agents typically requires rich user interactions, yet collecting human feedback is expensive and difficult to scale.
- [2]MIMESIS: Learning User Simulators as Training Environments for Interactive AgentsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 05:37 AM · same content
Extractive summary: sentences quoted from the sources.
Before this
- Oct 6, 2026CM-DPO: Constraint-Margin Direct Preference Optimization for LLM Planning
- Oct 6, 2026WorldSolver: Can LLM Agents Simulate the Physical Dynamics via Solver Generation?
- Oct 6, 2026Natural Language Questions as an Interface for Knowledge Graphs: QRAKEN Graph Distillation and Semantic Self-Healing
- Oct 5, 2026Sharing AI progress in mathematics
- Oct 4, 2026nerkyor/Qwen3.8-27B-Coder390-EfficientThink-Opus5.5-GPT6Astra-Grok4.7-DSV4Pro-K3-SFT-RLOO-MTP-DFlash2
- Sep 30, 2026pydantic/pydantic-ai v2.52.0: v2.52.0 (2026-09-29)
