A Good Self-Teacher Meets the Student Where They Are: Joint On-Policy Learning and Teaching
Reinforcement Learning (RL) from outcome rewards suffers from sparse supervision, particularly on difficult, long-horizon tasks where successful trajectories are rare and costly to generate.
Key points
- On-Policy Distillation (OPD) offers an attractive alternative by providing dense token-level supervision from a stronger teacher along the student's own generations.
- Self-distillation methods further remove the need for a separate teacher model by conditioning the same policy on privileged information to serve as its own teacher.
- We derive a necessary and sufficient condition for the teacher's local distillation update to be a positive multiple of the student's reward gradient.
- Based on this result, we propose Joint On-Policy Learning and Teaching (JOLT), which jointly trains a single policy in two roles: a privileged teacher using a KL-regularized objective, and an unprivileged student using dense on-policy distillation.
Sources (1)
- [1]A Good Self-Teacher Meets the Student Where They Are: Joint On-Policy Learning and TeachingarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 05:19 PM
Reinforcement Learning (RL) from outcome rewards suffers from sparse supervision, particularly on difficult, long-horizon tasks where successful trajectories are rare and costly to generate.
On-Policy Distillation (OPD) offers an attractive alternative by providing dense token-level supervision from a stronger teacher along the student's own generations.
Extractive summary: sentences quoted from the sources.