How Do Agentic LLMs Decide to Call Tools? A Tool-Call Vector Shaped by Suppression
To obtain such a variable, we propose a method that converts complex agentic prompts into minimal contrastive pairs in which a single request verb determines the tool-call decision: replacing an execution-verb (e.g., write) with an analysis-verb (e.g., discuss) reliably flips the decision, suggesting it is mediated by a compact internal state.
Key points
- Tool calling, invoking external tools on demand, is central to agentic LLMs, yet the mechanism that decides whether a model calls a tool or responds directly remains poorly understood.
- Agentic prompts are long and heavily scaffolded, combining role instructions, tool schemas, format templates, and the user's request across hundreds of tokens, creating a noisy, highly entangled context in which no single controllable variable for mechanistic analysis is obvious.
- We trace the decision to a vector, $μΔ$, that is both causally necessary and sufficient and generalizes beyond the discovery prompts to native multi-turn $τ^2$-Bench trajectories and verb-free requests.
- Behavioral ablations show that the scaffold establishes a tool-call prior; Transcoder decomposition then reveals that analysis verbs suppress this prior through features signaling that tool use is unnecessary, whereas execution verbs largely leave it intact.
Sources (1)
- [1]How Do Agentic LLMs Decide to Call Tools? A Tool-Call Vector Shaped by SuppressionarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 08:01 AM
To obtain such a variable, we propose a method that converts complex agentic prompts into minimal contrastive pairs in which a single request verb determines the tool-call decision: replacing an execution-verb (e.g., write) with an analysis-verb (e.g., discuss) reliably flips the decision, suggesting it is mediated by a compact internal state.
Tool calling, invoking external tools on demand, is central to agentic LLMs, yet the mechanism that decides whether a model calls a tool or responds directly remains poorly understood.
Extractive summary: sentences quoted from the sources.