AION
Research paperReasoning & Planning · Large Language Models2 sources · Oct 7, 2026

On-Policy Distillation Teaches New Skills but Not New Knowledge

On-policy distillation (OPD) strengthens language-model reasoning, yet whether students acquire new factual knowledge or compositional skill for multi-step reasoning remains unknown.

Key points

  • We separate these capabilities using a controlled synthetic framework that measures the student's initial capabilities and independently controls the teacher's additional facts, compositional skill, or both.
  • Across four models from three families, reverse-KL OPD reliably transfers compositional skill across unseen reasoning structures, but transfers minimal factual knowledge.
  • Decoupling the distillation recipe reveals the source of this asymmetry: replacing reverse KL with forward KL restores factual transfer, whereas student rollouts specifically improve the execution of multi-step reasoning.
  • Experiments on recent factual QA and competition mathematics show a similar asymmetry under reverse-KL OPD, yielding notable reasoning gains without factual memory expansion.

Sources (2)

  • [1]On-Policy Distillation Teaches New Skills but Not New Knowledge
    Hugging Face Daily Papers · Oct 7, 12:00 AM
    On-policy distillation (OPD) strengthens language-model reasoning, yet whether students acquire new factual knowledge or compositional skill for multi-step reasoning remains unknown.
    We separate these capabilities using a controlled synthetic framework that measures the student's initial capabilities and independently controls the teacher's additional facts, compositional skill, or both.
  • [2]On-Policy Distillation Teaches New Skills but Not New Knowledge
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 08:15 AM · same content

Extractive summary: sentences quoted from the sources.