ResearchResearch paperLarge Language Models · Reinforcement Learning1 source · Oct 6, 2026

Privileged Context as Drift in On-Policy Self-Distillation

On-policy self-distillation (OPSD) trains a language model to match a copy of itself conditioned on privileged context.

Key points

  • Existing work varies what privileged context contains and how it is produced while also changing models, data, and training setups, making the effects of privileged context design difficult to isolate.
  • Motivated by efforts in continual learning to reduce catastrophic forgetting, we study how the choice of privileged context affects policy drift.
  • We train Qwen2.5-7B with OPSD across these nine combinations and three datasets, measuring target-task accuracy, prior-task retention, reverse KL from the base policy, and parameter-update geometry.
  • For continual learning, these findings suggest that privileged context should be treated as part of OPSD's stability design because it is associated with how far and in what direction the policy moves.

Sources (1)

  • [1]Privileged Context as Drift in On-Policy Self-Distillation
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 06:45 AM
    On-policy self-distillation (OPSD) trains a language model to match a copy of itself conditioned on privileged context.
    Existing work varies what privileged context contains and how it is produced while also changing models, data, and training setups, making the effects of privileged context design difficult to isolate.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 6, 2026[AINews] Reflection Beam - 501B-A23B American Open Model
  2. Oct 4, 2026nerkyor/Qwen3.8-27B-Coder390-EfficientThink-Opus5.5-GPT6Astra-Grok4.7-DSV4Pro-K3-SFT-RLOO-MTP-DFlash2
  3. Oct 2, 2026alesha-pro/Qwen3.8-Flash-Next-abliterated-GSQ-RCO-Strata-GGUF
  4. Oct 1, 2026nvidia/PixelUMM
  5. Sep 28, 2026unslothai/unsloth v0.1.900-beta: Laya Decision Models + Library
  6. Aug 22, 2026sgl-project/sglang v0.5.18

Related