UP-MOPD: Update Projection in Multi-Teacher On-Policy Distillation
To address this gap, we propose Update Projection for Multi-Teacher On-Policy Distillation (UP-MOPD).
Key points
- On-policy distillation from multiple teachers combines expertise from different domains in a single student, but conflicting gradients can hinder this integration.
- Gradient corrections directly constrain parameter updates under plain SGD.
- UP-MOPD lets the original mixed gradient update the optimizer state and generate a candidate displacement, then projects only violating candidates before they are committed to the parameters.
- In experiments combining medical and general domains, UP-MOPD improves IFEval-loose accuracy late in training by 2.96 points over vanilla M-OPD.
Sources (1)
- [1]UP-MOPD: Update Projection in Multi-Teacher On-Policy DistillationarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 02:13 PM
To address this gap, we propose Update Projection for Multi-Teacher On-Policy Distillation (UP-MOPD).
On-policy distillation from multiple teachers combines expertise from different domains in a single student, but conflicting gradients can hinder this integration.
Extractive summary: sentences quoted from the sources.