Training Advisors for LLM Agents from Task Outcomes
We introduce Caddie, a method for training critics to provide natural-language analysis and advice as agents work through a task.
Key points
- Large language model agents tackle multi-step tasks by interleaving reasoning and tool calls with observations from the environment.
- Prior work has shown that natural-language feedback can help these agents revise their decisions during task execution.
- Trained on multi-hop question answering with a single base model, our Qwen3-4B critic improves success rates across four base models of different scales and architectures, including three not used during critic training.
- Our results show that agents can decide when to seek help from a critic at inference time and that outcome-based critic training can produce guidance that transfers across base models and task domains.
Sources (1)
- [1]Training Advisors for LLM Agents from Task OutcomesarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 11:14 AM
We introduce Caddie, a method for training critics to provide natural-language analysis and advice as agents work through a task.
Large language model agents tackle multi-step tasks by interleaving reasoning and tool calls with observations from the environment.
Extractive summary: sentences quoted from the sources.