Rethinking Cross-Tokenizer On-Policy Distillation: From Alignment Coverage to Supervision Reliability
In this paper, we examine whether expanding this alignment coverage improves learning.
ProofPaper ↗
Key points
- On-Policy Distillation (OPD) trains a student on its own generations using teacher feedback.
- With different tokenizers, comparing teacher and student predictions requires alignment at both sequence and vocabulary levels.
- Across three heterogeneous teacher--student pairs on mathematical reasoning and code generation, strict 1:1 groups already cover most student-generated tokens despite substantial vocabulary mismatch.
- Our findings motivate a shift from maximizing alignment coverage to prioritizing supervision reliability: compact supervision at strict positions can be more effective than broader coverage that introduces weakly aligned or conflicting training signals.
Sources (1)
- [1]Rethinking Cross-Tokenizer On-Policy Distillation: From Alignment Coverage to Supervision ReliabilityarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 02:37 PM
In this paper, we examine whether expanding this alignment coverage improves learning.
On-Policy Distillation (OPD) trains a student on its own generations using teacher feedback.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 6, 2026Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight
- Oct 6, 2026RISED: Rubrics for Agentic Multi-Environment Selection and Self-Distillation
- Sep 29, 2026[AINews] AMD buys World Labs for $8.2B, as Atlas solves sparse reconstruction problem for robotics, design and more
- Sep 29, 2026Salt++: Context-Aligned Post-Training for Few-Step Streaming Multimodal Generation
- Sep 28, 2026Notes on NVIDIA Nemotron
- Aug 10, 2026huggingface/transformers v5.15.0: Release: v5.15.0