Do LLMs Learn from Rewards in Context? : Rethinking the role of reward in In-Context Reinforcement Learning
LLM agents increasingly improve at inference time by accumulating experience in context rather than by updating parameters.
Key points
- This process is often described as in-context reinforcement learning (ICRL).
- We study this question in its simplest form, direct ICRL, where the model conditions directly on raw trajectory-reward pairs, and ask whether the reward acts as a learning signal.
- Through controlled experiments on four benchmarks across six models, we find that the reward is read, but its effect is small: flipping, randomizing, or removing the reward leaves the improvement curve almost unchanged, and this holds even under meta-prompts that explicitly instruct the model to explore, exploit, or reason over rewards.
- This reframing has implications for agent memory design: ICL factors such as input distribution and demonstrations may matter more than RL elements such as reward shaping and exploration.
Sources (1)
- [1]Do LLMs Learn from Rewards in Context? : Rethinking the role of reward in In-Context Reinforcement LearningarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 03:15 AM
LLM agents increasingly improve at inference time by accumulating experience in context rather than by updating parameters.
This process is often described as in-context reinforcement learning (ICRL).
Extractive summary: sentences quoted from the sources.
Before this
- Oct 8, 2026SFT-as-Context Mitigates Forgetting in Supervised Fine-Tuning
- Oct 7, 2026Leaner Transformers Can Easily Learn to Cluster
- Oct 6, 2026Spatial Induction Heads: In-Context Learning of Multidimensional Cellular Automata
- Oct 6, 2026Towards In-Parameter Memory Augmentation for Large Language Models
- Oct 6, 2026Adaptive Mean Estimation by In-Context Learning: A Gradient-Flow Analysis
- Oct 6, 2026Adversarially Trained Linear Transformers Are Optimal Robust In-Context Learners for Gaussian Mixtures
