When Tools Lie: Reliability of Mathematical Agents Under Corrupted Tool Feedback
Mathematical problem solving often requires deterministic computational steps that agents delegate to tools and implicitly trust.
ProofPaper ↗
Key points
- Yet tools can fail silently, returning plausible but incorrect results.
- How well can agents detect and correct corrupted tool call outputs?
- We study this through a controlled corruption framework where a hidden interceptor replaces tool call results with plausible incorrect information on targeted problems.
- We evaluate agents across 31 problems under four verification designs including no verification (baseline), mandatory same-context reflection, optional fresh-context verification, and optional structural verification.
Sources (1)
- [1]When Tools Lie: Reliability of Mathematical Agents Under Corrupted Tool FeedbackarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 10:28 AM
Mathematical problem solving often requires deterministic computational steps that agents delegate to tools and implicitly trust.
Yet tools can fail silently, returning plausible but incorrect results.
Extractive summary: sentences quoted from the sources.