ToolRACER: A Robust Agentic Conversation Emulation Resource for Agent Training and Evaluation
We present ToolRACER, a synthetic data generation pipeline that coordinates user, assistant and tool emulation models to generate and validated multi-turn interactions between a user and an agent.
Key points
- Task-oriented conversational agents remain fragile under real world conversation scenarios as they rarely follow a predictable script, especially when users exhibit non-cooperative behavior.
- Using \sysn, we construct ToolRACERBench a robust multi-turn conversation benchmark spanning six domains, ranging over 55 varied personas, generating a validated corpus of 5.6K conversation trajectories, with approximately 66% of conversations containing failure-prone conversation scenarios.
- We inject adversarial behaviors, producing validated conversational interaction trajectories that capture realistic, robust scenarios.
- Models trained on ToolRACERBench improve end to end agentic accuracy across $τ^2$-bench and ACEBench, demonstrating significant gains when mixed with in-domain dataset in small language models for agent capability tasks.
Sources (1)
- [1]ToolRACER: A Robust Agentic Conversation Emulation Resource for Agent Training and EvaluationarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 10:06 PM
We present ToolRACER, a synthetic data generation pipeline that coordinates user, assistant and tool emulation models to generate and validated multi-turn interactions between a user and an agent.
Task-oriented conversational agents remain fragile under real world conversation scenarios as they rarely follow a predictable script, especially when users exhibit non-cooperative behavior.
Extractive summary: sentences quoted from the sources.