On-Policy Distillation Teaches New Skills but Not New Knowledge
On-policy distillation (OPD) strengthens language-model reasoning, yet whether students acquire new factual knowledge or compositional skill for multi-step reasoning remains unknown.
Key points
- We separate these capabilities using a controlled synthetic framework that measures the student's initial capabilities and independently controls the teacher's additional facts, compositional skill, or both.
- Across four models from three families, reverse-KL OPD reliably transfers compositional skill across unseen reasoning structures, but transfers minimal factual knowledge.
- Decoupling the distillation recipe reveals the source of this asymmetry: replacing reverse KL with forward KL restores factual transfer, whereas student rollouts specifically improve the execution of multi-step reasoning.
- Experiments on recent factual QA and competition mathematics show a similar asymmetry under reverse-KL OPD, yielding notable reasoning gains without factual memory expansion.
Sources (2)
- [1]On-Policy Distillation Teaches New Skills but Not New KnowledgeHugging Face Daily Papers · Oct 7, 12:00 AM
On-policy distillation (OPD) strengthens language-model reasoning, yet whether students acquire new factual knowledge or compositional skill for multi-step reasoning remains unknown.
We separate these capabilities using a controlled synthetic framework that measures the student's initial capabilities and independently controls the teacher's additional facts, compositional skill, or both.
- [2]On-Policy Distillation Teaches New Skills but Not New KnowledgearXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 08:15 AM · same content
Extractive summary: sentences quoted from the sources.