AION
Research paperLarge Language Models1 source · Oct 7, 2026

Training Advisors for LLM Agents from Task Outcomes

We introduce Caddie, a method for training critics to provide natural-language analysis and advice as agents work through a task.

Key points

  • Large language model agents tackle multi-step tasks by interleaving reasoning and tool calls with observations from the environment.
  • Prior work has shown that natural-language feedback can help these agents revise their decisions during task execution.
  • Trained on multi-hop question answering with a single base model, our Qwen3-4B critic improves success rates across four base models of different scales and architectures, including three not used during critic training.
  • Our results show that agents can decide when to seek help from a critic at inference time and that outcome-based critic training can produce guidance that transfers across base models and task domains.

Sources (1)

  • [1]Training Advisors for LLM Agents from Task Outcomes
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 11:14 AM
    We introduce Caddie, a method for training critics to provide natural-language analysis and advice as agents work through a task.
    Large language model agents tackle multi-step tasks by interleaving reasoning and tool calls with observations from the environment.

Extractive summary: sentences quoted from the sources.