ReCal: Calibrating Structured Pruning for On-Policy Distillation Recovery
Structured pruning reduces the deployment cost of reasoning language models, but the resulting capability degradation can hinder subsequent on-policy distillation (OPD) recovery.
Key points
- We propose RECAL, Recovery-Aware Calibration, a simple plug-and-play approach that improves OPD recovery by adjusting calibration before pruning.
- RECAL uses forward KL between an unpruned teacher and a pruned probe to identify teacher-supported predictions disrupted by pruning, then reweights calibration statistics to guide existing pruning criteria toward preserving these predictions.
- Across multiple models and pruning methods, RECAL consistently improves mathematical reasoning after OPD, achieving gains of up to 16.7 percentage points on AIME, alongside improvements in most code-generation comparisons.
- These results demonstrate the value of recovery-aware calibration for improving on-policy distillation recovery of pruned reasoning models.
Sources (1)
- [1]ReCal: Calibrating Structured Pruning for On-Policy Distillation RecoveryarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 06:28 AM
Structured pruning reduces the deployment cost of reasoning language models, but the resulting capability degradation can hinder subsequent on-policy distillation (OPD) recovery.
We propose RECAL, Recovery-Aware Calibration, a simple plug-and-play approach that improves OPD recovery by adjusting calibration before pruning.
Extractive summary: sentences quoted from the sources.