AION
Research paperReinforcement Learning1 source · Oct 8, 2026

AdaptEvo: Adaptive Agent Learning with Evolving Supervision

To address these challenges, we introduce AdaptEvo, a framework for learning under imperfect supervision that couples confidence-adaptive policy optimization with evolving decision knowledge and evaluation rubrics.

Key points

  • Rule-governed contextual decision tasks require models to apply specified rules to case-specific context and evidence.
  • Its Evolution module synthesizes reusable decision knowledge from recurring failures across training cases and refines process rubrics to detect overlooked errors.
  • To support empirical evaluation, we construct an industrial multimodal content moderation dataset comprising a training set and In-Period and Out-of-Period test sets, with the latter collected under changed rules.
  • On Out-of-Period, the policy trained with CA-GRPO retains exact-label accuracy gains over the base model across evaluated checkpoints without injected decision knowledge, while GRPO declines with continued training.

Sources (1)

  • [1]AdaptEvo: Adaptive Agent Learning with Evolving Supervision
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 06:47 AM
    To address these challenges, we introduce AdaptEvo, a framework for learning under imperfect supervision that couples confidence-adaptive policy optimization with evolving decision knowledge and evaluation rubrics.
    Rule-governed contextual decision tasks require models to apply specified rules to case-specific context and evidence.

Extractive summary: sentences quoted from the sources.