AION
Research paperReinforcement Learning1 source · Oct 8, 2026

Do LLMs Learn from Rewards in Context? : Rethinking the role of reward in In-Context Reinforcement Learning

LLM agents increasingly improve at inference time by accumulating experience in context rather than by updating parameters.

Key points

  • This process is often described as in-context reinforcement learning (ICRL).
  • We study this question in its simplest form, direct ICRL, where the model conditions directly on raw trajectory-reward pairs, and ask whether the reward acts as a learning signal.
  • Through controlled experiments on four benchmarks across six models, we find that the reward is read, but its effect is small: flipping, randomizing, or removing the reward leaves the improvement curve almost unchanged, and this holds even under meta-prompts that explicitly instruct the model to explore, exploit, or reason over rewards.
  • This reframing has implications for agent memory design: ICL factors such as input distribution and demonstrations may matter more than RL elements such as reward shaping and exploration.

Sources (1)

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 8, 2026SFT-as-Context Mitigates Forgetting in Supervised Fine-Tuning
  2. Oct 7, 2026Leaner Transformers Can Easily Learn to Cluster
  3. Oct 6, 2026Spatial Induction Heads: In-Context Learning of Multidimensional Cellular Automata
  4. Oct 6, 2026Towards In-Parameter Memory Augmentation for Large Language Models
  5. Oct 6, 2026Adaptive Mean Estimation by In-Context Learning: A Gradient-Flow Analysis
  6. Oct 6, 2026Adversarially Trained Linear Transformers Are Optimal Robust In-Context Learners for Gaussian Mixtures

Related