AION
Research paperLarge Language Models · Robotics & Embodied AI · Agents & Tool Use1 source · Oct 6, 2026

ToolRACER: A Robust Agentic Conversation Emulation Resource for Agent Training and Evaluation

We present ToolRACER, a synthetic data generation pipeline that coordinates user, assistant and tool emulation models to generate and validated multi-turn interactions between a user and an agent.

Key points

  • Task-oriented conversational agents remain fragile under real world conversation scenarios as they rarely follow a predictable script, especially when users exhibit non-cooperative behavior.
  • Using \sysn, we construct ToolRACERBench a robust multi-turn conversation benchmark spanning six domains, ranging over 55 varied personas, generating a validated corpus of 5.6K conversation trajectories, with approximately 66% of conversations containing failure-prone conversation scenarios.
  • We inject adversarial behaviors, producing validated conversational interaction trajectories that capture realistic, robust scenarios.
  • Models trained on ToolRACERBench improve end to end agentic accuracy across $τ^2$-bench and ACEBench, demonstrating significant gains when mixed with in-domain dataset in small language models for agent capability tasks.

Sources (1)

  • [1]ToolRACER: A Robust Agentic Conversation Emulation Resource for Agent Training and Evaluation
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 10:06 PM
    We present ToolRACER, a synthetic data generation pipeline that coordinates user, assistant and tool emulation models to generate and validated multi-turn interactions between a user and an agent.
    Task-oriented conversational agents remain fragile under real world conversation scenarios as they rarely follow a predictable script, especially when users exhibit non-cooperative behavior.

Extractive summary: sentences quoted from the sources.