ResearchResearch paperReinforcement Learning · Large Language Models · Training & Scaling1 source · Oct 8, 2026

When Do We Need On-Policy Distillation? Distilling on Offline Student Rollouts Is Often Better

In this work, we ask a critical research question: Is on-policy sampling always beneficial for distilling arbitrary teacher-student pairs?

Key points

  • On-policy distillation (OPD) has become increasingly popular for transferring teacher capabilities to student models.
  • We show that a simple alternative, Semi-OPD, which distills from offline rollouts generated by the initial student, can often outperform OPD in both accuracy and training efficiency.
  • We further find that the choice between OPD and Semi-OPD depends on the alignment between the initial teacher and student, quantified by an output-token overlap ratio: OPD is beneficial only when the two are highly aligned with high overlap ratios.
  • For misaligned pairs, student rollouts can become increasingly off-policy w.r.t. the teacher as context length grows, weakening the distillation signal.

Sources (1)

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 8, 2026One Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed Experts
  2. Oct 8, 2026SpatialOPSD: Self-Distilling Spatial Intelligence from Verified Coding Agent Traces
  3. Oct 8, 2026Distilling Routed 3D Privilege for Spatial Reasoning in Vision-Language Models
  4. Oct 7, 2026MIMESIS: Learning User Simulators as Training Environments for Interactive Agents
  5. Oct 7, 2026Composing What Each Teacher Learned: Multi-Teacher On-Policy Distillation through Teacher-Relative Shifts
  6. Oct 7, 2026On-Policy Distillation Teaches New Skills but Not New Knowledge

Related