Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight
Reinforcement learning with verifiable rewards (RLVR) turns agent experience into learning signals primarily through scalar outcome rewards after interaction.
Key points
- For group-relative objectives, however, this signal vanishes when all rollouts receive the same reward, even though their trajectories may reveal useful information about what the task requires and how the agent fails.
- We ask a complementary question: can hindsight teach an agent what it could have anticipated before acting?
- We introduce prospective learning, which uses post-hoc experience to supervise foresight predictions from the pre-interaction view, and instantiate it with Self-Retrospection Distillation (SRD).
- Our results suggest that post-hoc agent experience is useful not only for evaluating or improving behavior, but also for shaping predictive representations before available interaction.
Sources (2)
- [1]Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior ForesightHugging Face Daily Papers · Oct 6, 12:00 AM
Reinforcement learning with verifiable rewards (RLVR) turns agent experience into learning signals primarily through scalar outcome rewards after interaction.
For group-relative objectives, however, this signal vanishes when all rollouts receive the same reward, even though their trajectories may reveal useful information about what the task requires and how the agent fails.
- [2]Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior ForesightarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 10:06 AM · same content
Extractive summary: sentences quoted from the sources.