ResearchResearch paperLarge Language Models1 source · Oct 6, 2026

Rethinking Cross-Tokenizer On-Policy Distillation: From Alignment Coverage to Supervision Reliability

In this paper, we examine whether expanding this alignment coverage improves learning.

Key points

  • On-Policy Distillation (OPD) trains a student on its own generations using teacher feedback.
  • With different tokenizers, comparing teacher and student predictions requires alignment at both sequence and vocabulary levels.
  • Across three heterogeneous teacher--student pairs on mathematical reasoning and code generation, strict 1:1 groups already cover most student-generated tokens despite substantial vocabulary mismatch.
  • Our findings motivate a shift from maximizing alignment coverage to prioritizing supervision reliability: compact supervision at strict positions can be more effective than broader coverage that introduces weakly aligned or conflicting training signals.

Sources (1)

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 6, 2026Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight
  2. Oct 6, 2026RISED: Rubrics for Agentic Multi-Environment Selection and Self-Distillation
  3. Sep 29, 2026[AINews] AMD buys World Labs for $8.2B, as Atlas solves sparse reconstruction problem for robotics, design and more
  4. Sep 29, 2026Salt++: Context-Aligned Post-Training for Few-Step Streaming Multimodal Generation
  5. Sep 28, 2026Notes on NVIDIA Nemotron
  6. Aug 10, 2026huggingface/transformers v5.15.0: Release: v5.15.0

Related