Loud Failures, Quiet Failures: Fault Detection and Recovery in Tool-Using Language Model Agents
Prior work has shown that language models over-trust tool outputs that fail silently; we ask how that over-trust plays out across the stages of failure handling in multi-turn agents.
Key points
- Tool-using agents are usually scored on whether they finish a task while the tools work.
- Deployments are less forgiving: services time out, endpoints disappear, parameter names change, and results come back well formed but wrong.
- Wrapping the executable environments of an established function-calling benchmark in a fault-injection layer, we inject one of four typed faults at a controlled point in the trajectory and record whether the agent notices, changes plan, recovers the task, or repeats itself.
- Because agents are stochastic, two fault-free runs of the same task end in the same state only 63.3% of the time; against that baseline, only a missing tool clearly lowers recovery (39.9%), while timeouts, schema drift, and corruption stay within run-to-run variation.
Sources (1)
- [1]Loud Failures, Quiet Failures: Fault Detection and Recovery in Tool-Using Language Model AgentsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 01:29 PM
Prior work has shown that language models over-trust tool outputs that fail silently; we ask how that over-trust plays out across the stages of failure handling in multi-turn agents.
Tool-using agents are usually scored on whether they finish a task while the tools work.
Extractive summary: sentences quoted from the sources.