AION
Research paperLarge Language Models · Reinforcement Learning · Robotics & Embodied AI1 source · Oct 6, 2026

Verify Less, Evolve More: Training Idea-Level Critics for Verification-Efficient ML Evolving Agents

As large language models become more powerful, self-evolving agents are able to tackle challenging tasks including AI for machine learning (AI4ML).

Key points

  • In AI4ML, while empirical verification is available, it often requires computationally costly model training and evaluation, limiting the speed and scale of agent evolution.
  • We train the critic models through supervised fine-tuning on high-quality critiques synthesized by Gemini-3.1-Pro, followed by GRPO to further improve their predictive accuracy.
  • Empirically, our critic models outperform Gemini-3.1-Pro in static idea evaluation, and these gains extend to agent inference, continual learning, and policy training.
  • Together, these results show that idea-level critic models help ML agents discover better solutions and learn stronger proposal policies under limited verification budgets.

Sources (1)

Extractive summary: sentences quoted from the sources.