Self-Evolve With a Reference:Anchored Training of Tool-Integrated Agents
Self-evolving tool-integrated agents learn from tasks and feedback generated within their own training loop.
Key points
- A Curriculum Agent generates tasks, while an Executor Agent learns from self-consistency signals through reinforcement learning.
- We propose AnchorLoop, which introduces a frozen copy of the previous iteration's Executor as a historical reference and reuses it on both sides of the training loop.
- Across 13 reasoning benchmarks, AnchorLoop improves over Agent0 by 2.5% on mathematical reasoning and 2.8% on general reasoning tasks.
- These results demonstrate the benefit of introducing a lightweight historical reference into self-evolving tool-integrated agents without external task or answer supervision.
Sources (1)
- [1]Self-Evolve With a Reference:Anchored Training of Tool-Integrated AgentsarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 11:13 AM
Self-evolving tool-integrated agents learn from tasks and feedback generated within their own training loop.
A Curriculum Agent generates tasks, while an Executor Agent learns from self-consistency signals through reinforcement learning.
Extractive summary: sentences quoted from the sources.
