AdaptEvo: Adaptive Agent Learning with Evolving Supervision
To address these challenges, we introduce AdaptEvo, a framework for learning under imperfect supervision that couples confidence-adaptive policy optimization with evolving decision knowledge and evaluation rubrics.
Key points
- Rule-governed contextual decision tasks require models to apply specified rules to case-specific context and evidence.
- Its Evolution module synthesizes reusable decision knowledge from recurring failures across training cases and refines process rubrics to detect overlooked errors.
- To support empirical evaluation, we construct an industrial multimodal content moderation dataset comprising a training set and In-Period and Out-of-Period test sets, with the latter collected under changed rules.
- On Out-of-Period, the policy trained with CA-GRPO retains exact-label accuracy gains over the base model across evaluated checkpoints without injected decision knowledge, while GRPO declines with continued training.
Sources (1)
- [1]AdaptEvo: Adaptive Agent Learning with Evolving SupervisionarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 06:47 AM
To address these challenges, we introduce AdaptEvo, a framework for learning under imperfect supervision that couples confidence-adaptive policy optimization with evolving decision knowledge and evaluation rubrics.
Rule-governed contextual decision tasks require models to apply specified rules to case-specific context and evidence.
Extractive summary: sentences quoted from the sources.