MetaOPD: Meta-Learned Token Weighting for On-Policy Distillation
In this paper, we propose MetaOPD, a bilevel optimization framework that jointly learns the student model and a lightweight token-weighting network.
Key points
- On-policy distillation (OPD) trains a student on its own generated responses using token-level teacher supervision.
- However, uniform weighting overlooks differences in token learning value, while existing weighting methods rely on predefined mappings from prediction signals to token weights.
- The inner objective updates the student through weighted OPD, while the outer objective optimizes the weighting network using validation loss on reference solutions after a virtual student update.
- Experiments on six mathematical reasoning and three out-of-domain datasets, covering two student scales and seven baselines, demonstrate the effectiveness of MetaOPD, with Avg@8/Pass@8 gains over OPD of 1.99/5.97 percentage points for the 0.6B student and 2.25/6.41 points for the 1.7B student.
Sources (1)
- [1]MetaOPD: Meta-Learned Token Weighting for On-Policy DistillationarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 01:58 PM
In this paper, we propose MetaOPD, a bilevel optimization framework that jointly learns the student model and a lightweight token-weighting network.
On-policy distillation (OPD) trains a student on its own generated responses using token-level teacher supervision.
Extractive summary: sentences quoted from the sources.