AION
Research paperAgents & Tool Use · Safety & Alignment · Large Language Models1 source · Oct 7, 2026

Loud Failures, Quiet Failures: Fault Detection and Recovery in Tool-Using Language Model Agents

Prior work has shown that language models over-trust tool outputs that fail silently; we ask how that over-trust plays out across the stages of failure handling in multi-turn agents.

Key points

  • Tool-using agents are usually scored on whether they finish a task while the tools work.
  • Deployments are less forgiving: services time out, endpoints disappear, parameter names change, and results come back well formed but wrong.
  • Wrapping the executable environments of an established function-calling benchmark in a fault-injection layer, we inject one of four typed faults at a controlled point in the trajectory and record whether the agent notices, changes plan, recovers the task, or repeats itself.
  • Because agents are stochastic, two fault-free runs of the same task end in the same state only 63.3% of the time; against that baseline, only a missing tool clearly lowers recovery (39.9%), while timeouts, schema drift, and corruption stay within run-to-run variation.

Sources (1)

Extractive summary: sentences quoted from the sources.