AION
Research paperReinforcement Learning · Large Language Models · Robotics & Embodied AI2 sources · Oct 7, 2026

MIMESIS: Learning User Simulators as Training Environments for Interactive Agents

We introduce MIMESIS, a purpose-built user simulator trained on human conversations with explicit reasoning supervision and 13 realistic behavioral patterns derived from real user interactions.

Key points

  • Training and evaluating interactive language agents typically requires rich user interactions, yet collecting human feedback is expensive and difficult to scale.
  • Simulated users offer a scalable alternative, but they must both resemble real user behavior and provide useful learning experiences for agents.
  • Across eight environments, training with MIMESIS yields better agent performance than training with GPT-5.5 under all nine unseen user simulators, demonstrating stronger generalization to new user simulators.
  • Moreover, we propose Coached On-Policy Self-Distillation (CSD), which leverages simulator-generated private reasoning traces and subsequent utterances as feedback on how well the agent addresses user needs.

Sources (2)

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 6, 2026CM-DPO: Constraint-Margin Direct Preference Optimization for LLM Planning
  2. Oct 6, 2026WorldSolver: Can LLM Agents Simulate the Physical Dynamics via Solver Generation?
  3. Oct 6, 2026Natural Language Questions as an Interface for Knowledge Graphs: QRAKEN Graph Distillation and Semantic Self-Healing
  4. Oct 5, 2026Sharing AI progress in mathematics
  5. Oct 4, 2026nerkyor/Qwen3.8-27B-Coder390-EfficientThink-Opus5.5-GPT6Astra-Grok4.7-DSV4Pro-K3-SFT-RLOO-MTP-DFlash2
  6. Sep 30, 2026pydantic/pydantic-ai v2.52.0: v2.52.0 (2026-09-29)

Related