RewardWeaver: Long-Horizon Interactive Learning for Language Agents via Self-Evolving Reward Adaptation
We introduce RewardWeaver, a self-evolving reward adaptation framework for language agents in long-horizon interaction.
Key points
- Reinforcement learning with verifiable rewards (RLVR) has driven substantial progress in domains where task outcomes can be reliably evaluated, but long-horizon interaction remains challenging due to sparse terminal feedback and difficult credit assignment.
- Process rewards provide denser supervision, yet the capabilities most relevant for training can change as the policy evolves: a behavior that is easy to evaluate or frequently deficient need not be the bottleneck currently limiting task success.
- RewardWeaver maintains a validated capability space in which the semantics of admitted Rubrics remain fixed, and closes the loop between policy optimization, task evaluation, failure attribution, and reward adaptation.
- Ablations further demonstrate the importance of dynamic reward allocation, failure-grounded attribution, and stable semantics for admitted capabilities.
Sources (1)
- [1]RewardWeaver: Long-Horizon Interactive Learning for Language Agents via Self-Evolving Reward AdaptationarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 02:04 PM
We introduce RewardWeaver, a self-evolving reward adaptation framework for language agents in long-horizon interaction.
Reinforcement learning with verifiable rewards (RLVR) has driven substantial progress in domains where task outcomes can be reliably evaluated, but long-horizon interaction remains challenging due to sparse terminal feedback and difficult credit assignment.
Extractive summary: sentences quoted from the sources.