ResearchResearch paperLarge Language Models · Reinforcement Learning · Training & Scaling1 source · Oct 7, 2026

SAPD: Step-Aligned Privileged Distillation

We introduce Step-Aligned Privileged Distillation (SAPD), a rollout-free self-distillation method that turns demonstrations into step-aligned distributional supervision.

Key points

  • On-policy post-training can improve large language models by learning from their own trajectories, but requires costly rollout generation.
  • We ask whether fixed demonstrations can support competitive off-policy learning through better supervision.
  • Our premise is that their usefulness depends not only on the training trajectories, but also on whether supervision provides informative preferences among continuations and connects this guidance to the reasoning decision being learned.
  • On mathematical reasoning benchmarks, SAPD outperforms supervised fine-tuning and label smoothing on average while remaining competitive with on-policy reinforcement learning and self-distillation.

Sources (1)

  • [1]SAPD: Step-Aligned Privileged Distillation
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 08:31 AM
    We introduce Step-Aligned Privileged Distillation (SAPD), a rollout-free self-distillation method that turns demonstrations into step-aligned distributional supervision.
    On-policy post-training can improve large language models by learning from their own trajectories, but requires costly rollout generation.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 7, 2026Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
  2. Oct 6, 2026Multi-Objective Aligned Small Language Model Framework for SUD Patient Dialogue Generation
  3. Oct 6, 2026Catastrophic Forgetting in Sequential Thermal Anti-UAV Detection: The Role of Scale-Conditioned Gradient Imbalance
  4. Oct 6, 2026AutodidactWAM: Cross-Modal Self-Distillation from Generated Video to Robot Actions
  5. Oct 6, 2026Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight
  6. Sep 28, 2026Notes on NVIDIA Nemotron

Related