AION
Research paperLarge Language Models · Reasoning & Planning · Efficiency & Inference1 source · Oct 8, 2026

ReCal: Calibrating Structured Pruning for On-Policy Distillation Recovery

Structured pruning reduces the deployment cost of reasoning language models, but the resulting capability degradation can hinder subsequent on-policy distillation (OPD) recovery.

Key points

  • We propose RECAL, Recovery-Aware Calibration, a simple plug-and-play approach that improves OPD recovery by adjusting calibration before pruning.
  • RECAL uses forward KL between an unpruned teacher and a pruned probe to identify teacher-supported predictions disrupted by pruning, then reweights calibration statistics to guide existing pruning criteria toward preserving these predictions.
  • Across multiple models and pruning methods, RECAL consistently improves mathematical reasoning after OPD, achieving gains of up to 16.7 percentage points on AIME, alongside improvements in most code-generation comparisons.
  • These results demonstrate the value of recovery-aware calibration for improving on-policy distillation recovery of pruned reasoning models.

Sources (1)

  • [1]ReCal: Calibrating Structured Pruning for On-Policy Distillation Recovery
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 06:28 AM
    Structured pruning reduces the deployment cost of reasoning language models, but the resulting capability degradation can hinder subsequent on-policy distillation (OPD) recovery.
    We propose RECAL, Recovery-Aware Calibration, a simple plug-and-play approach that improves OPD recovery by adjusting calibration before pruning.

Extractive summary: sentences quoted from the sources.