ResearchResearch paperReinforcement Learning · Large Language Models · Training & Scaling1 source · Oct 8, 2026

Why On-Policy Distillation Sometimes Fails: Vanishing Learning Signals

On-policy distillation (OPD) enables effective capability transfer between language models, yet the mechanisms underlying its failures are not fully understood.

Key points

  • Across code generation and mathematical reasoning, OPD with larger-scale teachers exhibits early loss plateaus, with an average final loss reduction of 25.1% after 200 updates, compared with 96.2% for self-RL teachers, obtained by further reinforcement learning (RL) training of the initial student.
  • To understand this difference, we analyze OPD as an idealized continuous-time dynamical system in the small-learning-rate limit.
  • Our training-log diagnostics associate these plateaus with an early decline in a gradient-based learning-signal proxy while substantial loss remains; these measurements do not establish why the underlying gradient weakens.
  • Across runs with and without loss plateaus, we observe small relative parameter changes (0.025-0.098%) and high similarity between the student's representations before and after OPD (linear CKA $>0.98$ across layers).

Sources (1)

  • [1]Why On-Policy Distillation Sometimes Fails: Vanishing Learning Signals
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 04:52 AM
    On-policy distillation (OPD) enables effective capability transfer between language models, yet the mechanisms underlying its failures are not fully understood.
    Across code generation and mathematical reasoning, OPD with larger-scale teachers exhibits early loss plateaus, with an average final loss reduction of 25.1% after 200 updates, compared with 96.2% for self-RL teachers, obtained by further reinforcement learning (RL) training of the initial student.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 8, 2026One Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed Experts
  2. Oct 8, 2026SpatialOPSD: Self-Distilling Spatial Intelligence from Verified Coding Agent Traces
  3. Oct 8, 2026Distilling Routed 3D Privilege for Spatial Reasoning in Vision-Language Models
  4. Oct 7, 2026MIMESIS: Learning User Simulators as Training Environments for Interactive Agents
  5. Oct 7, 2026Composing What Each Teacher Learned: Multi-Teacher On-Policy Distillation through Teacher-Relative Shifts
  6. Oct 7, 2026On-Policy Distillation Teaches New Skills but Not New Knowledge

Related