AION
Research paperReinforcement Learning · Large Language Models · Reasoning & Planning1 source · Oct 8, 2026

DIAL-OPD: Learning More from Fewer Tokens in On-Policy Distillation

We propose DIAL-OPD, a token-selection method that bridges log-probability and probability spaces by weighting reward magnitude with the logarithmic mean of teacher and student probabilities.

Key points

  • On-policy distillation (OPD) supervises student-generated trajectories with token-level teacher signals.
  • Yet we find that training on fewer tokens can outperform full-token OPD, challenging the intuition that more supervision improves learning.
  • Existing disagreement-based criteria ignore probability scale: tokens assigned negligible probability by both models, termed low-low tokens, can receive large log-ratio rewards and hinder learning.
  • Token-level evidence reveals that DIAL-OPD filters high-reward tokens with limited reasoning value while preserving supervision critical to reasoning correctness.

Sources (1)

  • [1]DIAL-OPD: Learning More from Fewer Tokens in On-Policy Distillation
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 10:36 AM
    We propose DIAL-OPD, a token-selection method that bridges log-probability and probability spaces by weighting reward magnitude with the logarithmic mean of teacher and student probabilities.
    On-policy distillation (OPD) supervises student-generated trajectories with token-level teacher signals.

Extractive summary: sentences quoted from the sources.