Distillation for Incrimination and Distillation for Capabilities
Powerful misaligned AI models might recognize alignment evaluations and strategically behave well on them, rendering direct audits uninformative.
ProofPaper ↗
Key points
- However, distilling such a model into a weaker benign student places the teacher in a Distillation Double Bind: if misalignment transfers, the student may conceal it less effectively, revealing evidence about the teacher; if it does not, the student may learn useful capabilities while remaining benign.
- We introduce two distinct distillation approaches, one targeting each outcome.
- Distilling AuditBench's secret-keeping models into their underlying instruction-tuned model produces students that are significantly more likely than their teachers to admit their hidden behavior when asked, suggesting that knowledge of the behavior transferred more readily than the propensity to conceal it.
- Together, these findings demonstrate two ways distillation can be used for AI safety: incriminating misaligned models, and extracting their capabilities without their misalignment.
Sources (1)
- [1]Distillation for Incrimination and Distillation for CapabilitiesarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 11:51 PM
Powerful misaligned AI models might recognize alignment evaluations and strategically behave well on them, rendering direct audits uninformative.
However, distilling such a model into a weaker benign student places the teacher in a Distillation Double Bind: if misalignment transfers, the student may conceal it less effectively, revealing evidence about the teacher; if it does not, the student may learn useful capabilities while remaining benign.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 7, 2026MIMESIS: Learning User Simulators as Training Environments for Interactive Agents
- Oct 7, 2026Composing What Each Teacher Learned: Multi-Teacher On-Policy Distillation through Teacher-Relative Shifts
- Oct 7, 2026On-Policy Distillation Teaches New Skills but Not New Knowledge
- Oct 7, 2026UniSkill: Learning Actor-Aligned Skill Proposals for an Evolving Policy
- Oct 6, 2026Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight
- Oct 6, 2026RISED: Rubrics for Agentic Multi-Environment Selection and Self-Distillation