AION
Research paperReinforcement Learning1 source · Oct 7, 2026

Learning to Accumulate Knowledge with Mutual Information

Therefore, we propose Knowledge Weaver, a reinforcement learning framework that trains a language model to curate reusable knowledge from agent trajectories.

Key points

  • Large language model (LLM) agents can improve their performance by reusing knowledge distilled from past interactions.
  • Effective knowledge accumulation should limit redundant overlap among entries and ensure that new knowledge contributes beyond what the bank already provides.
  • Yet training a curator with Group Relative Policy Optimization (GRPO) on standalone task success can reinforce general guidance even when it duplicates existing knowledge.
  • We couple feedback inspired by token-wise mutual information (MI) with marginal success rewards to guide knowledge accumulation.

Sources (1)

  • [1]Learning to Accumulate Knowledge with Mutual Information
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 01:20 PM
    Therefore, we propose Knowledge Weaver, a reinforcement learning framework that trains a language model to curate reusable knowledge from agent trajectories.
    Large language model (LLM) agents can improve their performance by reusing knowledge distilled from past interactions.

Extractive summary: sentences quoted from the sources.