AION
Research paperReinforcement Learning1 source · Oct 7, 2026

A Good Self-Teacher Meets the Student Where They Are: Joint On-Policy Learning and Teaching

Reinforcement Learning (RL) from outcome rewards suffers from sparse supervision, particularly on difficult, long-horizon tasks where successful trajectories are rare and costly to generate.

Key points

  • On-Policy Distillation (OPD) offers an attractive alternative by providing dense token-level supervision from a stronger teacher along the student's own generations.
  • Self-distillation methods further remove the need for a separate teacher model by conditioning the same policy on privileged information to serve as its own teacher.
  • We derive a necessary and sufficient condition for the teacher's local distillation update to be a positive multiple of the student's reward gradient.
  • Based on this result, we propose Joint On-Policy Learning and Teaching (JOLT), which jointly trains a single policy in two roles: a privileged teacher using a KL-regularized objective, and an unprivileged student using dense on-policy distillation.

Sources (1)

  • [1]A Good Self-Teacher Meets the Student Where They Are: Joint On-Policy Learning and Teaching
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 05:19 PM
    Reinforcement Learning (RL) from outcome rewards suffers from sparse supervision, particularly on difficult, long-horizon tasks where successful trajectories are rare and costly to generate.
    On-Policy Distillation (OPD) offers an attractive alternative by providing dense token-level supervision from a stronger teacher along the student's own generations.

Extractive summary: sentences quoted from the sources.