ResearchResearch paperReinforcement Learning1 source · Oct 6, 2026

On-Policy Distillation with Negative-Policy Rollouts

In this work, we introduce Negative-Policy OPD (NP-OPD), which complements teacher supervision with rollouts from a lower-performing, lower-capability negative policy that serves as a negative reference for the student.

Key points

  • On-policy distillation (OPD) has been widely studied as a post-training method in which a student model obtains token-level supervision from a stronger teacher on its own rollouts.
  • Rather than modifying the distillation reward formulation, NP-OPD introduces the negative policy at the rollout stage, continuously supplying tokens preferred by the negative policy over the teacher so that they remain exposed to teacher supervision throughout training.
  • Through extensive experiments, we show that NP-OPD improves OPD across model scales, generation modes, reasoning domains, and different OPD variants.
  • These results support our design of introducing negative signals through negative-policy rollouts and provide new insight into the role of the rollout policy in OPD.

Sources (1)

  • [1]On-Policy Distillation with Negative-Policy Rollouts
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 07:25 AM
    In this work, we introduce Negative-Policy OPD (NP-OPD), which complements teacher supervision with rollouts from a lower-performing, lower-capability negative policy that serves as a negative reference for the student.
    On-policy distillation (OPD) has been widely studied as a post-training method in which a student model obtains token-level supervision from a stronger teacher on its own rollouts.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 6, 2026Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight
  2. Oct 6, 2026RISED: Rubrics for Agentic Multi-Environment Selection and Self-Distillation
  3. Sep 29, 2026[AINews] AMD buys World Labs for $8.2B, as Atlas solves sparse reconstruction problem for robotics, design and more
  4. Sep 29, 2026Salt++: Context-Aligned Post-Training for Few-Step Streaming Multimodal Generation
  5. Sep 28, 2026Notes on NVIDIA Nemotron
  6. Aug 10, 2026huggingface/transformers v5.15.0: Release: v5.15.0

Related