Learning to Accumulate Knowledge with Mutual Information
Therefore, we propose Knowledge Weaver, a reinforcement learning framework that trains a language model to curate reusable knowledge from agent trajectories.
Key points
- Large language model (LLM) agents can improve their performance by reusing knowledge distilled from past interactions.
- Effective knowledge accumulation should limit redundant overlap among entries and ensure that new knowledge contributes beyond what the bank already provides.
- Yet training a curator with Group Relative Policy Optimization (GRPO) on standalone task success can reinforce general guidance even when it duplicates existing knowledge.
- We couple feedback inspired by token-wise mutual information (MI) with marginal success rewards to guide knowledge accumulation.
Sources (1)
- [1]Learning to Accumulate Knowledge with Mutual InformationarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 01:20 PM
Therefore, we propose Knowledge Weaver, a reinforcement learning framework that trains a language model to curate reusable knowledge from agent trajectories.
Large language model (LLM) agents can improve their performance by reusing knowledge distilled from past interactions.
Extractive summary: sentences quoted from the sources.